A Method and Apparatus for Protein Biomolecular Structure Analysis Based on Structural Encoders
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-06
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]本发明提供一种基于结构编码器的蛋白质生物大分子结构分析方法及装置,用以解决现有技术中通过单一特征分析蛋白质生物大分子的结构分析,且基于AlphaFold3无法主动对蛋白质生物大分子的结构进行主动探究并剥离其强大的结构理解能力,导致对蛋白质生物大分子的结构分析不准确、不完整的缺陷,实现基于预先训练的蛋白质结构编码器,蛋白质结构编码器对蛋白质生物大分子的蛋白质生物大分子结构进行特征向量和特征矩阵两个特征方面的结构编码分析和结构理解,得到蛋白质结构分析结果,提升蛋白质分析结果的准确性和完整性
[0020]本发明还提供一种非暂态计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时实现如上述任一种基于结构编码器的蛋白质生物大分子结构分析方法。
Smart Images

Figure CN122575461A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomacromolecule structure encoding technology, and in particular to a method and apparatus for analyzing the structure of protein biomacromolecules based on a structure encoder. Background Technology
[0002] The core technology for feature extraction of biological macromolecules is the structural encoder.
[0003] In existing technologies, the extraction of structural features from biomolecules is primarily based on structural encoders using full-folding models (such as AlphaFold3, a revolutionary AI model for predicting the structure of individual proteins). AlphaFold3's development focuses on generating three-dimensional coordinates from sequence to structure during biomolecule structural encoding, resulting in the inability to effectively explore and decouple its deep structural understanding capabilities. In achieving high-precision structural folding, AlphaFold3 dedicates all its attention mechanisms and computational resources to the final output of three-dimensional atomic coordinates. Under this deeply coupled end-to-end architecture, although AlphaFold3's Pairformer module (the module in AlphaFold3 and its derivatives used to improve model efficiency and performance) has objectively constructed an extremely excellent structural representation with high-order spatial topological insight, existing technologies generally consider AlphaFold3 merely an intermediate state for generating coordinates. This inherent emphasis on "generation" over "representation" has prevented the industry from actively exploring and decoupling its powerful structural understanding capabilities. Secondly, existing structural encoders suffer from overly simplistic feature representation methods. They employ vector representations, typically only explicitly modeling the local geometry and physical information of individual residues. Crucial information such as interactions between residues and their relative spatial positions remains implicitly represented by the numerical changes in these vectors, making them impossible to intuitively and explicitly read and constrain. Conversely, while matrix representations can explicitly model the global spatial relationships between residue pairs, they cannot directly and precisely characterize the local directional and other vector-like equivariant features of individual residue side chains in three-dimensional space. The vector representation methods of existing structural encoders consistently suffer from incomplete feature representation and insufficient modeling accuracy when capturing high-order, complex spatial geometric constraints.
[0004] Therefore, there is an urgent need for a structural encoder-based method for analyzing the structure of protein biomolecules to improve our understanding of their structures. Summary of the Invention
[0005] This invention provides a method and apparatus for analyzing the structure of protein biomolecules based on a structure encoder. It addresses the shortcomings of existing technologies that rely on single-feature analysis of protein biomolecule structure, and the fact that AlphaFold3 cannot actively explore the structure of protein biomolecules and extract their powerful structural understanding capabilities, leading to inaccurate and incomplete structural analysis. The invention utilizes a pre-trained protein structure encoder to perform structural encoding analysis and understanding of the protein biomolecule structure from two aspects: feature vectors and feature matrices, resulting in improved accuracy and completeness of the protein analysis results.
[0006] This invention provides a method for analyzing the structure of protein biomacromolecules based on a structural encoder, comprising the following steps.
[0007] To obtain the structure of protein biomolecules; The protein biomolecular structure is input into the protein structure encoder to obtain the protein structure analysis results output by the protein structure encoder. The protein structure encoder is trained based on the protein biomolecular sample information. The protein structure encoder is a model that obtains the protein coding results by performing structural coding analysis on the feature vectors and feature matrices of the protein biomolecular structure.
[0008] According to the present invention, a method for analyzing the structure of protein biomacromolecules based on a structure encoder is provided. The training of the protein structure encoder includes the following steps: Obtain information from protein biomolecule samples; The protein biomacromolecule sample information is input into the basic protein structure encoder to obtain the first basic feature matrix, the first basic feature vector, the second basic feature vector, and the second basic feature matrix output by the basic protein structure encoder. Based on the second basic feature vector and the second basic feature matrix, knowledge distillation is performed on the first basic feature vector and the first basic feature matrix to obtain the knowledge distillation loss function. The basic protein structure encoder is iteratively optimized based on the knowledge distillation loss function to obtain the protein structure encoder.
[0009] According to the present invention, a method for analyzing the structure of protein biomolecules based on a structure encoder is provided. The protein biomolecule sample information includes the protein biomolecule sample structure and the protein biomolecule sample sequence. The protein biomolecule sample information is input into a basic protein structure encoder to obtain a first basic feature matrix, a first basic feature vector, a second basic feature vector, and a second basic feature matrix output by the basic protein structure encoder, including: The structure of a protein biomolecule sample is input into the basic structure encoding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic structure encoding sub-model; wherein, the basic structure encoding sub-model is a sub-model that encodes the structure of the protein biomolecule sample. The protein biomolecule sample sequence is input into the basic folding sub-model of the basic protein structure encoder to obtain the second basic feature vector and the second basic feature matrix output by the basic folding sub-model; wherein, the basic folding sub-model is a sub-model for encoding the structure of the protein biomolecule sample sequence.
[0010] According to the present invention, a method for analyzing the structure of protein biomacromolecules based on a structure encoder is provided. The structure of a protein biomacromolecule sample is input into the basic structure encoding sub-model of a basic protein structure encoder, and the first basic feature vector and the first basic feature matrix output by the basic structure encoding sub-model are obtained, including: The structure of a protein biomacromolecule sample is input into the structure analysis module of the basic structure coding sub-model of the basic protein structure encoder to obtain the evolutionary information and self-template feature information output by the structure analysis module. Evolutionary information and self-template feature information are input into the basic folding module in the basic structure coding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic folding module. The basic folding module is used to perform high-order space set insight analysis on evolutionary information and self-template feature information.
[0011] According to the present invention, a protein biomacromolecule structure analysis method based on a structure encoder is provided. The structure of a protein biomacromolecule sample is input into the structure analysis module of the basic structure coding sub-model of a basic protein structure encoder. The evolutionary information and self-template feature information output by the structure analysis module are obtained, including: The structure of a protein biomolecule sample is input into the basic graph neural network submodule in the structure analysis module of the basic structure encoding submodel of the basic protein structure encoder, and the evolutionary information output by the basic graph neural network submodule is obtained. The basic graph neural network submodule is a submodule that performs three-dimensional spatial structure analysis on the protein biomolecule sample structure. The structure of a protein biomacromolecule sample is input into the basic self-template submodule of the structure analysis module of the basic structure coding submodel of the basic protein structure encoder, and the self-template feature information output by the basic self-template submodule is obtained. The basic self-template submodule is a submodule that extracts the hidden distance and angle constraints of the protein biomacromolecule sample structure based on the self-attention mechanism, and performs feature transformation on the extracted results.
[0012] According to the present invention, a method for analyzing the structure of protein biomacromolecules based on a structure encoder is provided. This method performs knowledge distillation on a first fundamental feature vector and a first fundamental feature matrix based on a second fundamental feature vector and a second fundamental feature matrix to obtain a knowledge distillation loss function, including: Knowledge distillation loss function ;in, Represents the first fundamental eigenvector. This represents the second fundamental eigenvector. This represents the first fundamental characteristic matrix. This represents the second fundamental characteristic matrix. The square of the vector norm. This represents the square of the matrix norm.
[0013] According to the present invention, a protein biomacromolecule structure analysis method based on a structure encoder is provided, which iteratively optimizes a basic protein structure encoder based on a knowledge distillation loss function to obtain a protein structure encoder, including: Get a protein structure encoding task request; Determine the downstream task network and the true labels of the downstream tasks based on the protein structure coding task request; The protein biomacromolecule sample structure and the real labels of the downstream task are input into the downstream task network to obtain the downstream task loss function of the downstream task network. The basic protein structure encoder is iteratively optimized based on the downstream task loss function and the knowledge distillation loss function to obtain the protein structure encoder; wherein, the downstream task loss function is used to perform end-to-end gradient backpropagation and parameter fine-tuning on the basic protein structure encoder.
[0014] According to the present invention, a protein biomacromolecule structure analysis method based on a structure encoder is provided. The method iteratively optimizes a basic protein structure encoder based on a downstream task loss function and a knowledge distillation loss function to obtain a protein structure encoder, comprising: The joint training loss function is determined based on the downstream task loss function and the knowledge distillation loss function. The basic protein structure encoder is iteratively optimized based on the joint training loss function to obtain the protein structure encoder.
[0015] According to the present invention, a method for analyzing the structure of protein biomacromolecules based on a structure encoder is provided, which determines a joint training loss function based on a downstream task loss function and a knowledge distillation loss function, including: Joint training loss function ;in, This represents the loss function for downstream tasks. Represents a basic protein structure encoder. Indicates the downstream task network. This represents the sample structure of protein biomolecules. Indicates the true label of the downstream task. Indicates hyperparameters, This represents the knowledge distillation loss function.
[0016] According to the present invention, a protein biomacromolecule structure analysis method based on a structure encoder is provided, which iteratively optimizes a basic protein structure encoder based on a joint training loss function to obtain a protein structure encoder, comprising: The basic protein structure encoder is iteratively optimized based on the joint training loss function. During the iterative optimization process, if the joint training loss function is less than or equal to the preset loss function, it is determined that the iterative optimization of the basic protein structure encoder is complete, and the protein structure encoder is obtained.
[0017] According to the present invention, a method for analyzing the structure of protein biomacromolecules based on a structure encoder is provided, which determines the downstream task network and the true labels of the downstream tasks based on the protein structure encoding task request, including: Determine the task type based on the protein structure coding task request; The downstream task network and the real label of the downstream task are determined according to the task type; different task types correspond to different downstream task networks.
[0018] This invention also provides a protein biomacromolecule structure analysis device based on a structure encoder, comprising the following modules: The structure acquisition module is used to acquire the protein biomolecule structure; The structure analysis module is used to input the protein biomolecular structure into the protein structure encoder and obtain the protein structure analysis results output by the protein structure encoder. The protein structure encoder is trained based on the protein biomolecular sample information. The protein structure encoder is a model that obtains the protein coding results by performing structural encoding analysis on the feature vectors and feature matrices of the protein biomolecular structure.
[0019] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the protein biomacromolecule structure analysis methods based on the structure encoder described above.
[0020] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the protein biomacromolecule structure analysis methods based on structure encoders described above.
[0021] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the protein biomacromolecule structure analysis methods based on a structure encoder as described above.
[0022] This invention provides a method and apparatus for protein biomolecule structure analysis based on a structure encoder. The method involves acquiring the structure of a protein biomolecule, inputting this structure into a protein structure encoder, and obtaining the protein structure analysis result output by the encoder. The protein structure encoder is trained based on protein biomolecule sample information and is a model that performs structural encoding analysis of the protein biomolecule structure using feature vectors and feature matrices to obtain the protein encoding result. This invention addresses the shortcomings of existing technologies that analyze protein biomolecule structure using single features, and the limitations of AlphaFold3 in actively exploring and understanding the structure of proteins, leading to inaccurate and incomplete structural analysis. The invention utilizes a pre-trained protein structure encoder to perform structural encoding analysis and understanding of the protein biomolecule structure using both feature vectors and feature matrices, thereby improving the accuracy and completeness of the protein analysis results. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0024] Figure 1 This is a schematic flowchart of the protein biomacromolecule structure analysis method based on a structure encoder provided by the present invention.
[0025] Figure 2 This is one of the schematic diagrams of the training process of the protein structure encoder provided by the present invention.
[0026] Figure 3 This is the second schematic diagram of the training process of the protein structure encoder provided by the present invention.
[0027] Figure 4 This is a schematic diagram of the structure of the protein biomacromolecule structure analysis device based on the structure encoder provided by the present invention.
[0028] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0030] The following is combined with Figure 1 This invention describes a protein biomacromolecule structure analysis method based on a structure encoder. This method is applicable to protein biomacromolecule structure analysis using a structure encoder guided by AlphaFold3. The execution subject of this method can be an electronic device or a protein biomacromolecule structure analysis device based on a structure encoder installed in the electronic device. This protein biomacromolecule structure analysis device based on a structure encoder can be implemented through software, hardware, or a combination of both. Figure 1 This is a schematic flowchart of the protein biomacromolecule structure analysis method based on a structure encoder provided by the present invention, as shown below. Figure 1 As shown, the method includes the following steps: 101 and 102.
[0031] Step 101: Obtain the protein biomolecule structure.
[0032] In this step, the protein biomacromolecule structure refers to the structure after the protein biomacromolecule has been resolved.
[0033] Specifically, the process involves obtaining the protein biomolecules for which protein structure analysis is required, and then analyzing the protein biomolecules to obtain their structures.
[0034] Step 102: Input the protein biomacromolecule structure into the protein structure encoder to obtain the protein structure analysis results output by the protein structure encoder; wherein, the protein structure encoder is trained based on the protein biomacromolecule sample information, and the protein structure encoder is a model that obtains the protein coding results by performing structural coding analysis of the feature vector and feature matrix of the protein biomacromolecule structure.
[0035] Specifically, the protein biomolecular structure is input into the protein structure encoder. Based on the pre-trained protein structure encoder, which can represent the encoding of biomolecular structures and knowledge distillation using feature vectors and feature matrices, the protein biomolecular structure is analyzed, and the protein structure analysis results output by the protein structure encoder are obtained.
[0036] The advantage of this approach is that, in response to the shortcomings of traditional single representation methods that tend to overlook certain aspects when faced with complex spatial geometric constraints, this invention uses vectors and matrices to jointly encode the structure of protein biomolecules, achieving complete structural encoding from local to global and from nodes to relationships, thereby improving the accuracy of protein structure analysis results.
[0037] In one specific embodiment, the training of the protein structure encoder includes the following steps: acquiring protein biomolecule sample information; inputting the protein biomolecule sample information into the basic protein structure encoder to obtain a first basic feature matrix, a first basic feature vector, a second basic feature vector, and a second basic feature matrix output by the basic protein structure encoder; performing knowledge distillation on the first basic feature vector and the first basic feature matrix based on the second basic feature vector and the second basic feature matrix to obtain a knowledge distillation loss function; and iteratively optimizing the basic protein structure encoder according to the knowledge distillation loss function to obtain the protein structure encoder.
[0038] In this step, the first basic feature matrix refers to the residue-level feature matrix output by the basic structure coding sub-model during training, and the first basic feature vector refers to the residue-level feature vector output by the basic structure coding sub-model in the basic protein structure encoder; the second basic feature matrix refers to the feature matrix output by the basic folding sub-model in the basic protein structure encoder, and the second basic feature vector refers to the feature vector output by the basic folding sub-model in the basic protein structure encoder.
[0039] Specifically, the protein biomolecule sample information is input into the basic protein structure encoder. Based on the basic structure encoding sub-model and basic folding sub-model in the basic protein structure encoder, the protein biomolecule sample information is structurally encoded to obtain the first basic feature matrix, the first basic feature vector, the second basic feature vector, and the second basic feature matrix output by the basic protein structure encoder. Then, using the alignment concept, knowledge distillation is performed on the first basic feature vector and the first basic feature matrix based on the second basic feature vector and the second basic feature matrix to obtain the knowledge distillation loss function. The basic protein structure encoder is iteratively optimized according to the knowledge distillation loss function to obtain the protein structure encoder.
[0040] In one specific embodiment, the protein biomolecule sample information includes the protein biomolecule sample structure and the protein biomolecule sample sequence. Inputting the protein biomolecule sample information into a basic protein structure encoder to obtain a first basic feature matrix, a first basic feature vector, a second basic feature vector, and a second basic feature matrix output by the basic protein structure encoder includes: inputting the protein biomolecule sample structure into a basic structure encoding sub-model of the basic protein structure encoder to obtain a first basic feature vector and a first basic feature matrix output by the basic structure encoding sub-model; wherein, the basic structure encoding sub-model is a sub-model that encodes the structure of the protein biomolecule sample; and inputting the protein biomolecule sample sequence into a basic folding sub-model of the basic protein structure encoder to obtain a second basic feature vector and a second basic feature matrix output by the basic folding sub-model; wherein, the basic folding sub-model is a sub-model that encodes the structure of the protein biomolecule sample sequence.
[0041] In this step, the base folding sub-model can be, for example, AlphaFold3, but this embodiment does not limit it.
[0042] Specifically, the structure of a protein biomolecule sample is input into the basic structure encoding sub-model of the basic protein structure encoder. Based on the basic structure encoding sub-model, the structure of the protein biomolecule sample is encoded to obtain the first basic feature vector and the first basic feature matrix output by the basic structure encoding sub-model. The sequence of a protein biomolecule sample is input into the basic folding sub-model of the basic protein structure encoder. Based on the basic folding model, the sequence of the protein biomolecule sample is encoded to obtain the second basic feature vector and the second basic feature matrix output by the basic folding model.
[0043] In one specific embodiment, the protein biomolecule sample structure is input into the basic structure coding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic structure coding sub-model. This includes: inputting the protein biomolecule sample structure into the structure analysis module in the basic structure coding sub-model of the basic protein structure encoder to obtain the evolutionary information and self-template feature information output by the structure analysis module; inputting the evolutionary information and self-template feature information into the basic folding module in the basic structure coding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic folding module; wherein, the basic folding module is a module used for high-order spatial set insight analysis of the evolutionary information and self-template feature information.
[0044] In this step, the basic folding module in the basic structure encoding sub-model of the basic protein structure encoder is the backbone network of Pairformer (the core module of AlphaFold3 for processing global features), and this embodiment does not limit this.
[0045] In this implementation, Pairformer is a stack of multiple PairformerBlocks (used for iterative optimization of two types of key information). Within each PairformerBlock, the matrix representation first captures the complex path combination logic between residues through triangular multiplication and then uses triangular attention to propagate long-range dependencies in the global topology. Next, this geometrically constrained two-dimensional information is mapped back to a one-dimensional vector representation through a biased attention mechanism. After multiple blocks are stacked, the basic folding module can be repeatedly coupled between one-dimensional residue representations and two-dimensional residue pair representations, thereby simultaneously modeling local geometric constraints and long-range interactions. This embodiment does not limit this approach.
[0046] Specifically, the structure of a protein biomolecule sample is input into the structure analysis module of the basic structure coding sub-model of the basic protein structure encoder to obtain evolutionary information and self-template feature information output by the structure analysis module. The evolutionary information and self-template feature information are then input into the basic folding module of the basic structure coding sub-model of the basic protein structure encoder. Based on the basic folding module, the evolutionary information and self-template feature information are analyzed and processed. The basic folding module is mainly responsible for repeatedly exchanging information between vector representation and matrix representation and gradually refining the global structural context, thereby obtaining the first basic feature vector and the first basic feature matrix output by the basic folding module. Among them, the basic folding module is a module used to perform high-order space set insight analysis on the evolutionary information and self-template feature information.
[0047] In one specific embodiment, the protein biomacromolecule sample structure is input into the structure analysis module of the basic structure coding sub-model of the basic protein structure encoder to obtain evolutionary information and self-template feature information output by the structure analysis module. This includes: inputting the protein biomacromolecule sample structure into the basic graph neural network sub-module of the structure analysis module of the basic structure coding sub-model of the basic protein structure encoder to obtain evolutionary information output by the basic graph neural network sub-module; wherein, the basic graph neural network sub-module is a sub-module for performing three-dimensional spatial structure analysis of the protein biomacromolecule sample structure; and inputting the protein biomacromolecule sample structure into the basic self-template sub-module of the structure analysis module of the basic structure coding sub-model of the basic protein structure encoder to obtain self-template feature information output by the basic self-template sub-module; wherein, the basic self-template sub-module is a sub-module for extracting hidden distance and angle constraints of the protein biomacromolecule sample structure based on a self-attention mechanism and performing feature transformation on the extracted results.
[0048] In this step, the basic graph neural network submodule can be, for example, a graph neural network submodule composed of a Geometric Vector Perceptron-Graph Neural Network (GVP-GNN). The input of GVP-GNN is the protein biomolecule sample structure, and the output is evolutionary information. The basic self-template submodule takes the protein biomolecule sample structure as input and outputs self-template feature information.
[0049] Specifically, the protein biomolecule sample structure is input into the basic graph neural network submodule within the structure analysis module of the basic structure encoding submodel of the basic protein structure encoder. This submodule utilizes a lightweight geometric graph neural network to complete the evolutionary information. Specifically, the protein biomolecule sample structure consists of nodes and edges; nodes represent protein residue features, and edges represent the graph connections and geometric relationships between residues. The basic graph neural network submodule first decouples the node and edge representations into scalar and vector channels: scalar channels encode rotation- and translation-invariant information such as amino acid categories and local environment; vector channels capture rotation- and isotropic features such as spatial orientation and local geometry. Subsequently, the basic graph neural network submodule performs message passing and feature fusion on the graph's adjacency topology. Specifically, each residue node receives information from its neighboring nodes and corresponding edge features, and updates it based on its current representation. After layer-by-layer propagation, the node representation not only integrates the local geometric environment but also gradually encodes a larger structural context. Finally, the basic graph neural network submodule outputs the high-dimensional feature vector corresponding to each residue, which is the evolutionary information. This evolutionary information serves as the input prior to the Pairformer backbone network.
[0050] Meanwhile, inspired by the Template in AlphaFold3 (the Template is a module in AlphaFold3 that implicitly learns spatial relationships through iterative optimization of pairwise information), the structure of protein biomolecule samples is input into the basic self-template submodule in the structure analysis module of the basic structure coding submodel of the basic protein structure encoder. The basic self-template submodule extracts implicit distance and angle constraints based on its internal attention mechanism and transforms the extracted implicit distance and angle constraints into internal self-template features, thereby obtaining the self-template feature information output by the basic self-template submodule. The self-template feature information serves as the input prior of the basic folding module in the basic structure coding submodel of the basic protein structure encoder.
[0051] Furthermore, the evolutionary information and self-template feature information are input into the basic folding module in the basic structure coding sub-model of the basic protein structure encoder. The basic folding module performs high-order space set insight analysis on the evolutionary information and self-template feature information to obtain the first basic feature vector and the first basic feature matrix output by the basic folding module.
[0052] In one specific embodiment, knowledge distillation is performed on the first basic feature vector and the first basic feature matrix based on the second basic feature vector and the second basic feature matrix to obtain a knowledge distillation loss function, including: knowledge distillation loss function. ;in, Represents the first fundamental eigenvector. This represents the second fundamental eigenvector. This represents the first fundamental characteristic matrix. This represents the second fundamental characteristic matrix. The square of the vector norm. This represents the square of the matrix norm.
[0053] Specifically, the second basic feature vector and the second basic feature matrix are used as soft labels. Based on the second basic feature vector and the second basic feature matrix, a lossless knowledge distillation is performed on the first basic feature vector and the first basic feature matrix using a global topological structure deep semantic understanding capability, thus obtaining the knowledge distillation loss function. ;in, Represents the first fundamental eigenvector. This represents the second fundamental eigenvector. This represents the first fundamental characteristic matrix. This represents the second fundamental characteristic matrix. The square of the vector norm. This represents the square of the matrix norm.
[0054] in, This indicates that the first fundamental eigenvector Approximating the second fundamental feature vector ; This indicates that the first fundamental characteristic matrix is... Approximating the second fundamental feature matrix .
[0055] In one specific embodiment, the basic protein structure encoder is iteratively optimized based on the knowledge distillation loss function to obtain the protein structure encoder, including: obtaining a protein structure encoding task request; determining a downstream task network and downstream task real labels based on the protein structure encoding task request; inputting the protein biomolecule sample structure and downstream task real labels into the downstream task network to obtain the downstream task loss function of the downstream task network; iteratively optimizing the basic protein structure encoder based on the downstream task loss function and the knowledge distillation loss function to obtain the protein structure encoder; wherein, the downstream task loss function is used to perform end-to-end gradient backpropagation and parameter fine-tuning of the basic protein structure encoder.
[0056] In this step, the protein structure encoding task request refers to the user's task requirement to encode the structure of a protein. The protein structure encoding task request includes task categories, and different task types correspond to different downstream task networks. This embodiment does not limit this.
[0057] Specifically, the process involves obtaining protein structure encoding task requests; determining downstream task networks and real labels based on these requests; inputting the protein biomolecule sample structure and the real labels into the downstream task network to obtain its downstream task loss function; iteratively optimizing the basic protein structure encoder based on the downstream task loss function and the knowledge distillation loss function to obtain the protein structure encoder; and using the downstream task loss function for end-to-end gradient backpropagation and parameter fine-tuning of the basic protein structure encoder.
[0058] In one specific embodiment, the basic protein structure encoder is iteratively optimized based on the downstream task loss function and the knowledge distillation loss function to obtain the protein structure encoder, including: determining a joint training loss function based on the downstream task loss function and the knowledge distillation loss function; and iteratively optimizing the basic protein structure encoder based on the joint training loss function to obtain the protein structure encoder.
[0059] Specifically, after obtaining the downstream task loss function, a joint training loss function is determined based on the downstream task loss function and the knowledge distillation loss function; the basic protein structure encoder is iteratively optimized based on the joint training loss function to obtain the protein structure encoder.
[0060] The advantage of this setup is that when downstream task networks and real labels are added for training, the trained protein structure encoder can not only guarantee the accuracy of task recognition, but also fully inherit the structural understanding ability of the folding model. This overcomes the technical defects of the single spatial feature expression of biological macromolecules and the inability to independently decouple the capabilities of the folding model.
[0061] In one specific embodiment, a joint training loss function is determined based on the downstream task loss function and the knowledge distillation loss function, including: a joint training loss function. ;in, This represents the loss function for downstream tasks. Represents a basic protein structure encoder. Indicates the downstream task network. This represents the sample structure of protein biomolecules. Indicates the true label of the downstream task. Indicates hyperparameters, This represents the knowledge distillation loss function.
[0062] In one specific embodiment, iterative optimization of the basic protein structure encoder based on the joint training loss function to obtain the protein structure encoder includes: iterative optimization of the basic protein structure encoder based on the joint training loss function; during the iterative optimization process, if the joint training loss function is less than or equal to a preset loss function, it is determined that the iterative optimization of the basic protein structure encoder is completed, and the protein structure encoder is obtained.
[0063] In this step, the preset loss function is a pre-defined function used to judge the joint training loss function. Based on the judgment result, it is determined whether the protein structure encoder has been trained. If the joint training loss function is less than or equal to the preset loss function, it is determined that the iterative optimization of the basic protein structure encoder has been completed, and the protein structure encoder is obtained.
[0064] Specifically, the basic protein structure encoder is iteratively optimized based on the joint training loss function. During the iterative optimization process, if the joint training loss function is less than or equal to the preset loss function, it is determined that the iterative optimization of the basic protein structure encoder is complete, and the protein structure encoder is obtained.
[0065] The advantage of this setup is that, through a series of techniques such as lightweight graph neural networks, self-template mechanisms, knowledge distillation, and downstream loss fine-tuning, the resulting protein structure encoder can fully acquire the ability of the folding model to perceive and model the complex spatial geometric constraints of biomacromolecules.
[0066] In one specific embodiment, determining the downstream task network and the true label of the downstream task based on the protein structure encoding task request includes: determining the task type based on the protein structure encoding task request; determining the downstream task network and the true label of the downstream task corresponding to the task type based on the task type; wherein, different task types correspond to different downstream task networks.
[0067] Specifically, the task type is determined based on the protein structure coding task request; the downstream task network and the real label of the downstream task are determined based on the task type; different task types correspond to different downstream task networks.
[0068] In one specific embodiment, Figure 2 This is one of the schematic diagrams of the training process of the protein structure encoder provided by the present invention, such as... Figure 2 As shown, the training steps include the following: Step 201, Step 202, Step 203, Step 204, Step 205, Step 206 and Step 207.
[0069] Step 201: Obtain protein biomolecule sample information; protein biomolecule sample information includes protein biomolecule sample structure and protein biomolecule sample sequence.
[0070] Step 202: Input the protein biomacromolecule sample structure into the basic graph neural network submodule in the structure analysis module of the basic structure encoding submodel of the basic protein structure encoder, and obtain the evolutionary information output by the basic graph neural network submodule.
[0071] Specifically, the protein biomolecule sample structure is input into the basic graph neural network submodule within the structure analysis module of the basic structure encoding submodel of the basic protein structure encoder. This submodule utilizes a lightweight geometric graph neural network to complete the evolutionary information. Specifically, the protein biomolecule sample structure consists of nodes and edges; nodes represent protein residue features, and edges represent the graph connections and geometric relationships between residues. The basic graph neural network submodule first decouples the node and edge representations into scalar and vector channels: scalar channels encode rotation- and translation-invariant information such as amino acid categories and local environment; vector channels capture rotation- and isotropic features such as spatial orientation and local geometry. Subsequently, the basic graph neural network submodule performs message passing and feature fusion on the graph's adjacency topology. Specifically, each residue node receives information from its neighboring nodes and corresponding edge features, and updates it based on its current representation. After layer-by-layer propagation, the node representation not only integrates the local geometric environment but also gradually encodes a larger structural context. Finally, the basic graph neural network submodule outputs the high-dimensional feature vector corresponding to each residue, which is the evolutionary information. This evolutionary information serves as the input prior to the Pairformer backbone network.
[0072] Step 203: Input the protein biomacromolecule sample structure into the basic self-template submodule of the basic structure coding submodel of the basic protein structure encoder, and obtain the self-template feature information output by the basic self-template submodule.
[0073] Specifically, the structure of the protein biomacromolecule sample is input into the basic self-template submodule of the structure analysis module of the basic structure coding submodel of the basic protein structure encoder. The basic self-template submodule extracts implicit distance and angle constraints based on its internal attention mechanism and transforms the extracted implicit distance and angle constraints into internal self-template features, thereby obtaining the self-template feature information output by the basic self-template submodule. The self-template feature information serves as the input prior of the basic folding module in the basic structure coding submodel of the basic protein structure encoder.
[0074] In one specific embodiment, the execution order of steps 202 and 203 is not fixed. Step 202 can be executed first, followed by step 203; or step 203 can be executed first, followed by step 202; or steps 202 and 203 can be executed simultaneously. This embodiment does not limit this.
[0075] Step 204: Input the evolutionary information and self-template feature information into the basic folding module of the basic structure encoding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic folding module.
[0076] Step 205: Input the protein biomacromolecule sample sequence into the basic folding sub-model of the basic protein structure encoder to obtain the second basic feature vector and the second basic feature matrix output by the basic folding sub-model.
[0077] Specifically, the protein biomolecule sample sequence is input into the basic folding sub-model of the basic protein structure encoder. The protein biomolecule sample sequence is structurally encoded based on the basic folding sub-model, and the second basic feature vector and the second basic feature matrix output by the basic folding sub-model are obtained.
[0078] Step 206: Perform knowledge distillation on the first basic feature vector and the first basic feature matrix based on the second basic feature vector and the second basic feature matrix to obtain the knowledge distillation loss function.
[0079] Specifically, the knowledge distillation loss function ;in, Represents the first fundamental eigenvector. This represents the second fundamental eigenvector. This represents the first fundamental characteristic matrix. This represents the second fundamental characteristic matrix. The square of the vector norm. This represents the square of the matrix norm.
[0080] Step 207: Iteratively optimize the basic protein structure encoder based on the knowledge distillation loss function to obtain the protein structure encoder.
[0081] In one specific embodiment, Figure 3 This is the second schematic diagram of the training process of the protein structure encoder provided by the present invention, as shown below. Figure 3 As shown, specifically, the steps for iteratively optimizing the basic protein structure encoder based on the knowledge distillation loss function to obtain the protein structure encoder include the following: steps 301, 302, 303, 304, and 305.
[0082] Step 301: Obtain the protein structure coding task request.
[0083] Step 302: Determine the task type based on the protein structure coding task request; and determine the downstream task network and downstream task real labels corresponding to the task type; different task types correspond to different downstream task networks.
[0084] For example, the downstream task network could be a diffusion network, which is used to generate 3D coordinates for reconstruction; or the downstream task network could be a lightweight head, used to implement the downstream task. When the acquired task type is one that requires reconstruction, the diffusion network is selected; when the acquired task type is one that requires task output, the lightweight head is selected. This embodiment does not limit this.
[0085] Step 303: Input the protein biomacromolecule sample structure and the real labels of the downstream task into the downstream task network to obtain the downstream task loss function of the downstream task network.
[0086] Step 304: Determine the joint training loss function based on the downstream task loss function and the knowledge distillation loss function.
[0087] Specifically, the joint training loss function ;in, This represents the loss function for downstream tasks. Represents a basic protein structure encoder. Indicates the downstream task network. This represents the sample structure of protein biomolecules. Indicates the true label of the downstream task. Indicates hyperparameters, This represents the knowledge distillation loss function.
[0088] Step 305: Iteratively optimize the basic protein structure encoder based on the joint training loss function to obtain the protein structure encoder.
[0089] Specifically, the basic protein structure encoder is iteratively optimized based on the joint training loss function. During the iterative optimization process, if the joint training loss function is less than or equal to the preset loss function, it is determined that the iterative optimization of the basic protein structure encoder is complete, and the protein structure encoder is obtained.
[0090] In summary, the protein structure encoder trained by this invention uses vectors and matrices to jointly encode the structure of biological macromolecules, simultaneously representing the spatial topological relationships between residue nodes and residue pairs. Given the extremely high requirements for the accuracy and comprehensiveness of feature representation in protein structure encoding tasks (downstream biological computing tasks), the protein structure encoder needs to capture both the local geometric information of residues at the microscopic level and understand the spatial relationship structure of residue pairs at the macroscopic level. To address the shortcomings of traditional single representation methods that are prone to overlooking certain aspects when facing complex spatial geometric constraints, a protein structure encoder using a joint representation of vectors and matrices is proposed, achieving complete structure encoding from local to global and from nodes to relationships, thereby improving the accuracy of protein structure analysis results.
[0091] This invention provides a protein biomolecule structure analysis method based on a structure encoder. The method involves acquiring the structure of a protein biomolecule, inputting this structure into a protein structure encoder, and obtaining the protein structure analysis result output by the encoder. The protein structure encoder is trained based on protein biomolecule sample information and is a model that performs structural encoding analysis on the feature vectors and feature matrices of the protein biomolecule structure to obtain the protein encoding result. This invention addresses the shortcomings of existing technologies that analyze protein biomolecule structure using single features, and the limitations of AlphaFold3 in actively exploring and extracting its powerful structural understanding capabilities, leading to inaccurate and incomplete structural analysis. The invention achieves this by using a pre-trained protein structure encoder, which performs structural encoding analysis and structural understanding on both feature vectors and feature matrices of the protein biomolecule structure to obtain the protein structure analysis result, thus improving the accuracy and completeness of the protein analysis results.
[0092] The following describes the protein biomacromolecule structure analysis device based on the structure encoder provided by the present invention. The protein biomacromolecule structure analysis device based on the structure encoder described below can be referred to in correspondence with the protein biomacromolecule structure analysis method based on the structure encoder described above.
[0093] Figure 4This is a schematic diagram of the protein biomacromolecule structure analysis device based on a structure encoder provided by the present invention, with reference to... Figure 4 As shown, the protein biomacromolecule structure analysis device 400 based on a structure encoder includes: a structure acquisition module 401 and a structure analysis module 402; wherein, The structure acquisition module 401 is used to acquire the protein biomacromolecule structure; The structure analysis module 402 is used to input the protein biomacromolecule structure into the protein structure encoder and obtain the protein structure analysis results output by the protein structure encoder. The protein structure encoder is trained based on the protein biomacromolecule sample information. The protein structure encoder is a model that performs structural encoding analysis of the protein biomacromolecule structure using feature vectors and feature matrices to obtain the protein encoding results.
[0094] In one example embodiment, the device further includes a model training module. The model training module is configured to: acquire protein biomolecule sample information; input the protein biomolecule sample information into a basic protein structure encoder to obtain a first basic feature matrix, a first basic feature vector, a second basic feature vector, and a second basic feature matrix output by the basic protein structure encoder; perform knowledge distillation on the first basic feature vector and the first basic feature matrix based on the second basic feature vector and the second basic feature matrix to obtain a knowledge distillation loss function; and iteratively optimize the basic protein structure encoder according to the knowledge distillation loss function to obtain a protein structure encoder.
[0095] In one example embodiment, the protein biomacromolecule sample information includes the protein biomacromolecule sample structure and the protein biomacromolecule sample sequence.
[0096] In one example embodiment, the model training module inputs protein biomolecule sample information into a basic protein structure encoder to obtain a first basic feature matrix, a first basic feature vector, a second basic feature vector, and a second basic feature matrix output by the basic protein structure encoder. Specifically, it is used to: input the structure of the protein biomolecule sample into the basic structure encoding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic structure encoding sub-model; wherein, the basic structure encoding sub-model is a sub-model that encodes the structure of the protein biomolecule sample; and input the sequence of the protein biomolecule sample into the basic folding sub-model of the basic protein structure encoder to obtain the second basic feature vector and the second basic feature matrix output by the basic folding sub-model; wherein, the basic folding sub-model is a sub-model that encodes the structure of the protein biomolecule sample sequence.
[0097] In one example embodiment, the model training module inputs the protein biomacromolecule sample structure into the basic structure coding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic structure coding sub-model. Specifically, it is used to: input the protein biomacromolecule sample structure into the structure analysis module in the basic structure coding sub-model of the basic protein structure encoder to obtain the evolutionary information and self-template feature information output by the structure analysis module; input the evolutionary information and self-template feature information into the basic folding module in the basic structure coding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic folding module; wherein, the basic folding module is a module used to perform high-order spatial set insight analysis on the evolutionary information and self-template feature information.
[0098] In one example embodiment, the model training module inputs the protein biomacromolecule sample structure into the structure analysis module of the basic structure coding sub-model of the basic protein structure encoder to obtain evolutionary information and self-template feature information output by the structure analysis module. Specifically, it is used to: input the protein biomacromolecule sample structure into the basic graph neural network sub-module of the structure analysis module of the basic structure coding sub-model of the basic protein structure encoder to obtain evolutionary information output by the basic graph neural network sub-module; wherein, the basic graph neural network sub-module is a sub-module for performing three-dimensional spatial structure analysis of the protein biomacromolecule sample structure; and input the protein biomacromolecule sample structure into the basic self-template sub-module of the structure analysis module of the basic structure coding sub-model of the basic protein structure encoder to obtain self-template feature information output by the basic self-template sub-module; wherein, the basic self-template sub-module is a sub-module that extracts the hidden distance and angle constraints of the protein biomacromolecule sample structure based on the self-attention mechanism and performs feature transformation on the extracted results.
[0099] In one example embodiment, the model training module performs knowledge distillation on the first basic feature vector and the first basic feature matrix based on the second basic feature vector and the second basic feature matrix to obtain a knowledge distillation loss function, specifically used for: Knowledge Distillation Loss Function ;in, Represents the first fundamental eigenvector. This represents the second fundamental eigenvector. This represents the first fundamental characteristic matrix. This represents the second fundamental characteristic matrix. The square of the vector norm. This represents the square of the matrix norm.
[0100] In one example embodiment, the model training module iteratively optimizes the basic protein structure encoder based on the knowledge distillation loss function to obtain the protein structure encoder. Specifically, this is used to: obtain a protein structure encoding task request; determine the downstream task network and the true labels of the downstream tasks based on the protein structure encoding task request; input the protein biomolecule sample structure and the true labels of the downstream tasks into the downstream task network to obtain the downstream task loss function of the downstream task network; and iteratively optimize the basic protein structure encoder based on the downstream task loss function and the knowledge distillation loss function to obtain the protein structure encoder. The downstream task loss function is used for end-to-end gradient backpropagation and parameter fine-tuning of the basic protein structure encoder.
[0101] In one example embodiment, the model training module iteratively optimizes the basic protein structure encoder based on the downstream task loss function and the knowledge distillation loss function to obtain the protein structure encoder. Specifically, it is used to: determine the joint training loss function based on the downstream task loss function and the knowledge distillation loss function; and iteratively optimize the basic protein structure encoder based on the joint training loss function to obtain the protein structure encoder.
[0102] In one example embodiment, the model training module determines a joint training loss function based on the downstream task loss function and the knowledge distillation loss function, specifically for: the joint training loss function. ;in, This represents the loss function for downstream tasks. Represents a basic protein structure encoder. Indicates the downstream task network. This represents the sample structure of protein biomolecules. Indicates the true label of the downstream task. Indicates hyperparameters, This represents the knowledge distillation loss function.
[0103] In one example embodiment, the model training module iteratively optimizes the basic protein structure encoder based on the joint training loss function to obtain the protein structure encoder. Specifically, it is used to: iteratively optimize the basic protein structure encoder based on the joint training loss function; during the iterative optimization process, if the joint training loss function is less than or equal to a preset loss function, it is determined that the iterative optimization of the basic protein structure encoder is completed, and the protein structure encoder is obtained.
[0104] In one example embodiment, the model training module determines the downstream task network and the true labels of the downstream task based on the protein structure encoding task request. Specifically, it is used to: determine the task type based on the protein structure encoding task request; and determine the downstream task network and the true labels of the downstream task corresponding to the task type based on the task type; wherein different task types correspond to different downstream task networks.
[0105] The apparatus of this embodiment can be used to execute the method of any embodiment in the side embodiment of the protein biomacromolecule structure analysis method based on structure encoder. Its specific implementation process and technical effects are similar to those in the side embodiment of the protein biomacromolecule structure analysis method based on structure encoder. For details, please refer to the detailed description in the side embodiment of the protein biomacromolecule structure analysis method based on structure encoder, which will not be repeated here.
[0106] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a protein biomacromolecule structure analysis method based on a structure encoder. This method includes: acquiring the protein biomacromolecule structure; inputting the protein biomacromolecule structure into a protein structure encoder to obtain the protein structure analysis result output by the protein structure encoder; wherein the protein structure encoder is trained based on protein biomacromolecule sample information, and the protein structure encoder is a model that performs structural encoding analysis of the protein biomacromolecule structure using feature vectors and feature matrices to obtain the protein encoding result.
[0107] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0108] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the protein biomacromolecule structure analysis method based on the structure encoder provided by the above methods. The method includes: obtaining the protein biomacromolecule structure; inputting the protein biomacromolecule structure into a protein structure encoder to obtain the protein structure analysis result output by the protein structure encoder; wherein, the protein structure encoder is trained based on the protein biomacromolecule sample information, and the protein structure encoder is a model that performs structural encoding analysis of the protein biomacromolecule structure using feature vectors and feature matrices to obtain the protein encoding result.
[0109] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the protein biomacromolecule structure analysis method based on the structure encoder provided by the above methods. The method includes: obtaining the protein biomacromolecule structure; inputting the protein biomacromolecule structure into a protein structure encoder to obtain the protein structure analysis result output by the protein structure encoder; wherein the protein structure encoder is trained based on the protein biomacromolecule sample information, and the protein structure encoder is a model for obtaining the protein coding result by performing structural coding analysis of the feature vector and feature matrix of the protein biomacromolecule structure.
[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0111] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for analyzing the structure of protein biomacromolecules based on a structural encoder, characterized in that, include: To obtain the structure of protein biomolecules; The protein biomacromolecule structure is input into a protein structure encoder to obtain the protein structure analysis result output by the protein structure encoder; wherein, the protein structure encoder is trained based on the protein biomacromolecule sample information, and the protein structure encoder is a model that obtains the protein encoding result by performing structural encoding analysis of the feature vector and feature matrix of the protein biomacromolecule structure.
2. The method for protein biomacromolecule structure analysis based on a structure encoder according to claim 1, characterized in that, The training of the protein structure encoder includes the following steps: Obtain the protein biomolecule sample information; The protein biomacromolecule sample information is input into the basic protein structure encoder to obtain the first basic feature matrix, the first basic feature vector, the second basic feature vector, and the second basic feature matrix output by the basic protein structure encoder. Based on the second basic feature vector and the second basic feature matrix, knowledge distillation is performed on the first basic feature vector and the first basic feature matrix to obtain the knowledge distillation loss function; The basic protein structure encoder is iteratively optimized based on the knowledge distillation loss function to obtain the protein structure encoder.
3. The method for protein biomacromolecule structure analysis based on a structure encoder according to claim 2, characterized in that, The protein biomolecule sample information includes the protein biomolecule sample structure and the protein biomolecule sample sequence; the step of inputting the protein biomolecule sample information into a basic protein structure encoder to obtain the first basic feature matrix, the first basic feature vector, the second basic feature vector, and the second basic feature matrix output by the basic protein structure encoder includes: The protein biomolecule sample structure is input into the basic structure encoding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic structure encoding sub-model; wherein, the basic structure encoding sub-model is a sub-model that encodes the structure of the protein biomolecule sample. The protein biomolecule sample sequence is input into the basic folding sub-model of the basic protein structure encoder to obtain the second basic feature vector and the second basic feature matrix output by the basic folding sub-model; wherein, the basic folding sub-model is a sub-model for structural encoding of the protein biomolecule sample sequence.
4. The method for protein biomacromolecule structure analysis based on a structure encoder according to claim 3, characterized in that, The step of inputting the protein biomacromolecule sample structure into the basic structure encoding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic structure encoding sub-model includes: The structure of the protein biomacromolecule sample is input into the structure analysis module of the basic structure coding sub-model of the basic protein structure encoder to obtain the evolutionary information and self-template feature information output by the structure analysis module. The evolutionary information and the self-template feature information are input into the basic folding module in the basic structure encoding sub-model of the basic protein structure encoder to obtain the first basic feature vector and the first basic feature matrix output by the basic folding module; wherein, the basic folding module is a module used to perform high-order space set insight analysis on the evolutionary information and the self-template feature information.
5. The method for protein biomacromolecule structure analysis based on a structure encoder according to claim 4, characterized in that, The process involves inputting the protein macromolecule sample structure into the structure analysis module of the basic structure encoding sub-model of the basic protein structure encoder to obtain the evolutionary information and self-template feature information output by the structure analysis module, including: The protein biomolecule sample structure is input into the basic graph neural network submodule in the structure analysis module of the basic structure encoding submodel of the basic protein structure encoder to obtain the evolutionary information output by the basic graph neural network submodule; wherein, the basic graph neural network submodule is a submodule that performs three-dimensional spatial structure analysis on the protein biomolecule sample structure. The protein biomacromolecule sample structure is input into the basic self-template submodule of the structure analysis module of the basic structure encoding submodel of the basic protein structure encoder to obtain the self-template feature information output by the basic self-template submodule; wherein, the basic self-template submodule is a submodule that extracts the hidden distance and angle constraints of the protein biomacromolecule sample structure based on the self-attention mechanism, and performs feature transformation on the extracted results.
6. The method for protein biomacromolecule structure analysis based on a structure encoder according to claim 2, characterized in that, The knowledge distillation process, based on the second basic feature vector and the second basic feature matrix, on the first basic feature vector and the first basic feature matrix to obtain the knowledge distillation loss function, includes: The knowledge distillation loss function ;in, This represents the first basic feature vector. This represents the second basic feature vector. This represents the first basic feature matrix. This represents the second basic feature matrix. The square of the vector norm. This represents the square of the matrix norm.
7. The method for protein biomacromolecule structure analysis based on a structure encoder according to claim 2, characterized in that, The step of iteratively optimizing the basic protein structure encoder based on the knowledge distillation loss function to obtain the protein structure encoder includes: Get a protein structure encoding task request; Determine the downstream task network and the true labels of the downstream tasks based on the protein structure encoding task request; The protein biomacromolecule sample structure and the real label of the downstream task are input into the downstream task network to obtain the downstream task loss function of the downstream task network. The basic protein structure encoder is iteratively optimized based on the downstream task loss function and the knowledge distillation loss function to obtain the protein structure encoder; wherein, the downstream task loss function is used to perform end-to-end gradient backpropagation and parameter fine-tuning on the basic protein structure encoder.
8. The method for analyzing the structure of protein biomacromolecules based on a structure encoder according to claim 7, characterized in that, The process of iteratively optimizing the basic protein structure encoder based on the downstream task loss function and the knowledge distillation loss function to obtain the protein structure encoder includes: The joint training loss function is determined based on the downstream task loss function and the knowledge distillation loss function. The basic protein structure encoder is iteratively optimized based on the joint training loss function to obtain the protein structure encoder.
9. The method for analyzing the structure of protein biomacromolecules based on a structure encoder according to claim 8, characterized in that, The determination of the joint training loss function based on the downstream task loss function and the knowledge distillation loss function includes: The joint training loss function ;in, This represents the downstream task loss function. This represents the basic protein structure encoder. This refers to the downstream task network. This represents the structure of the protein biomolecule sample. This indicates the true label of the downstream task. Indicates hyperparameters, This represents the knowledge distillation loss function.
10. The method for analyzing the structure of protein biomacromolecules based on a structure encoder according to claim 8, characterized in that, The step of iteratively optimizing the basic protein structure encoder based on the joint training loss function to obtain the protein structure encoder includes: The basic protein structure encoder is iteratively optimized based on the joint training loss function. During the iterative optimization process, if the joint training loss function is less than or equal to the preset loss function, it is determined that the iterative optimization of the basic protein structure encoder is complete, and the protein structure encoder is obtained.
11. The method for analyzing the structure of protein biomacromolecules based on a structure encoder according to claim 7, characterized in that, The step of determining the downstream task network and the true label of the downstream task based on the protein structure encoding task request includes: The task type is determined based on the protein structure encoding task request; The downstream task network and the real label of the downstream task are determined according to the task type; wherein different task types correspond to different downstream task networks.
12. A method for analyzing the structure of protein biomacromolecules based on a structural encoder, characterized in that, include: The structure acquisition module is used to acquire the protein biomolecule structure; The structure analysis module is used to input the protein biomacromolecule structure into the protein structure encoder to obtain the protein structure analysis result output by the protein structure encoder; wherein, the protein structure encoder is trained based on the protein biomacromolecule sample information, and the protein structure encoder is a model that obtains the protein encoding result by performing structural encoding analysis of the feature vector and feature matrix of the protein biomacromolecule structure.