Protein structure representation method based on vector quantization
Through a vector quantization method, protein data is divided into known and unknown parts, encoded and quantified, and the amino acid residues and spatial structures of unknown parts are generated by FoldGPT, which solves the problem of insufficient information utilization in protein modeling in the prior art, and achieves efficient and accurate protein structure generation.
Patent Information
- Application Number
- CN202510387916.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
Existing protein sequence and structural modeling methods cannot fully utilize the spatial structure information of proteins, it is difficult to efficiently process diversified input data, and there are problems of insufficient accuracy and flexibility in protein folding and reverse folding tasks.
Using a vector quantization method, protein data is divided into known and unknown parts, protein sequence structure coding and soft condition vector quantization are performed, and amino acid residues and spatial structures of unknown parts are generated by fold symbol generation model FoldGPT, and continuous potential representations are quantified through soft condition vector quantizer to simplify the data processing process.
It improves the accuracy and computational efficiency of protein structure generation, enhances the stability and robustness of the model, and can effectively integrate the sequence and structural information of the protein to generate high-precision three-dimensional structures.
Smart Images

Figure CN120279980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of artificial intelligence and life science, and particularly to a method for representing protein structures based on vector quantization. Background Art
[0002] Protein sequence and structure modeling is a core research area in bioinformatics, and its complexity stems from the close relationship between the three - dimensional structure of proteins and the one - dimensional amino acid sequence. In recent years, protein sequence - structure co - modeling has gradually become an important research direction, aiming to simultaneously model the amino acid sequence and three - dimensional structure of proteins and create their joint representation. Existing co - modeling methods mainly include combining pre - trained sequence models with graph neural networks or generating new protein sequences and structures from scratch through co - generative models. These methods effectively integrate the complementary information between sequences and structures and are widely applied, especially in fields such as antibody design, protein folding prediction, and functional research. However, how to further improve the accuracy and flexibility of the model to better handle protein folding and inverse folding problems remains a current challenge.
[0003] Vector quantization technology has played an important role in the representation of protein sequences and structures. Traditional vector quantization methods perform data compression and generation by mapping latent representations to discrete codebook vectors. However, in protein structure modeling, it still faces the dual challenges of accuracy and flexibility. To improve this, existing research has proposed methods such as residual quantization and lookup - free vector quantization, which improve the generation efficiency and accuracy by decomposing the space, iterative quantization, and avoiding the lookup process respectively. However, although these techniques have been successful in other fields, in protein structure modeling, how to handle complex protein geometric structures and retain sufficient structural information still requires further optimization.
[0004] Angle - based structure representation is another important technique in protein modeling. It mainly represents the three - dimensional structure of proteins by converting the geometric structure of proteins into angle features (such as bond angles, torsion angles, etc.) that are not affected by translation and rotation. This method can not only eliminate the interference of geometric transformation but also retain the inherent features of protein structures, making it show high robustness in protein folding and inverse folding tasks. By encoding the angle information of proteins into sequences, existing research enables the effective combination of protein structures with sequence information, thus promoting the generation and optimization of protein structures. However, how to combine these angle representations with sequence information and optimize the performance of the model is an urgent problem to be solved currently.
[0005] In summary, although relevant research results have been achieved in the modeling of protein sequences and structures, existing protein structure prediction methods often fail to fully utilize the spatial structure information of proteins or cannot efficiently process diverse input data. How to accurately integrate protein sequence and structure information, how to improve the accuracy of generation and reconstruction, and how to handle large-scale complex data remain key challenges in current research. Summary of the Invention
[0006] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a protein structure representation method based on vector quantization, to realize a new protein structure language representation, and to be able to effectively generate high-precision three-dimensional protein structures.
[0007] The purpose of the present invention can be achieved through the following technical solutions: A protein structure representation method based on vector quantization, comprising the following steps:
[0008] S1. Divide protein data into known parts and unknown parts, where the known parts include known protein sequences and known protein structures;
[0009] S2. For the known parts, perform protein sequence structure encoding and soft conditional vector quantization processing in sequence to obtain the quantized representation of the known parts;
[0010] S3. Input the quantized representation of the known parts into the folding symbol generation model FoldGPT, and output the predicted quantized representation;
[0011] S4. Perform protein sequence structure decoding on the predicted quantized representation to generate amino acid residues and their spatial structures of the unknown parts.
[0012] Further, the specific process of step S1 is as follows:
[0013] Divide protein sequence and structure data into known parts and unknown parts, where the unknown parts include the residue sequence and structure data to be generated, and I m is the index set of known residues;
[0014] The known parts include the determined residue sequence and structure data, and I k is the index set of known residues.
[0015] Further, in step S2, the protein sequence structure encoding is specifically to encode the protein sequence and structure data respectively and map them to a continuous latent space representation to obtain a continuous latent representation.
[0016] Further, the soft conditional vector quantization process in step S2 specifically uses a soft conditional vector quantizer to quantize the continuous latent representation and convert it into a discrete latent space code.
[0017] Further, the specific process of protein sequence structure encoding in step S2 includes:
[0018] S201. Encode the protein sequence, convert the amino acid sequence S = {s i : 1 ≤ i ≤ n} of the protein into a numerical representation form, where s i is the i-th amino acid;
[0019] S202. Encode the protein structure and encode the three-dimensional spatial structure data of the protein;
[0020] S203. Integrate the sequence information and structure information of the protein to generate a comprehensive latent representation h = q enc (S, A), which includes not only the order information of amino acids but also their spatial conformation information.
[0021] Further, step S201 specifically uses word embedding technology to encode each amino acid residue and map it into a high-dimensional space to obtain the corresponding continuous vector representation.
[0022] Further, step S202 specifically uses a graph neural network model to model the spatial structure of the protein and convert it into a continuous high-dimensional vector representation, retaining the spatial topological information of the structure.
[0023] Further, the specific process of soft conditional vector quantization in step S2 includes:
[0024] S211. Input of latent representation, use the continuous latent representation obtained by encoding protein sequence and structure data as input and pass it to the soft conditional vector quantizer z = Q(h);
[0025] S212. Construct the codebook where Bit(·) is used to compress the hidden layer representation z to a 0 / 1 code of length log2(m), and then use a multi-layer perceptron to map the code vector in the codebook into a d-dimensional vector, that is, v j = ConditionNe(Bit(z, log2(m)));
[0026] S213. Codebook alignment to obtain the quantized hidden layer representation where SoftVQ(·) represents a differentiable vector quantization operation;
[0027] S214. During the training phase, use the decoder to decode the quantized hidden layer representation to obtain the reconstructed structure And perform iterative training in combination with the loss function to minimize the difference between the folded symbol and the original protein structure, where the folded symbol is the potential representation obtained by encoding the protein sequence structure; during the evaluation phase, use Bit(i, log2(m)) as the quantized discrete code
[0028] Furthermore, in the training phase, step S214 specifically adopts the gradient descent method to optimize the error between the reconstructed structure And the real structure A, as well as the error between the reconstructed sequence And the real sequence S
[0029] Furthermore, the specific process of step S3 is as follows
[0030] Input the discrete latent code of the known part Into the folded symbol generation model FoldGPT, and generate the latent code of the unknown part through the autoregressive generation method
[0031] Compared with the prior art, the present invention has the following advantages
[0032] The present invention first divides protein data into known parts and unknown parts, then performs protein sequence structure encoding and soft conditional vector quantization processing on the known parts to obtain the quantized representation of the known parts; then inputs the quantized representation of the known parts into the folded symbol generation model FoldGPT to output the predicted quantized representation; finally, performs protein sequence structure decoding on the predicted quantized representation to generate the amino acid residues and their spatial structures of the unknown parts. Thus, by encoding protein sequence and structure data and mapping them to the latent space representation, and quantizing the continuous latent representation through the soft conditional vector quantization method, it can make full use of the spatial structure information of proteins and effectively simplify the data processing process, not only improving the generation accuracy of protein structures, but also enhancing the computational efficiency in large-scale protein structure prediction
[0033] By jointly encoding protein sequence and structure data and mapping them to a continuous latent space representation, the present invention can effectively combine the sequence features and spatial structure information of proteins, thereby providing a richer and more comprehensive input data representation for subsequent structure generation. Based on this protein sequence and structure data encoding method, a folding symbol is obtained as a joint representation of protein sequence and structure, which can realize the unified representation of protein sequence and three-dimensional structure information. Compared with the limitations of traditional methods that only rely on protein sequence or only rely on structure data, the present invention makes full use of the complementary information of sequence and structure.
[0034] The present invention proposes a soft conditional vector quantizer for quantifying the continuous latent representation of protein sequence and structure data. Through this quantization method, the continuous latent space can be mapped to discrete latent codes, thus simplifying the data representation and making the data more compact and operable. This technical means greatly improves the processing ability of spatial information in the protein structure generation process and enhances the stability and robustness of the model. The soft conditional vector quantization process overcomes the accuracy bottleneck in structure reconstruction of existing methods through soft queries in the entire symbol space, significantly improves the reconstruction quality of protein structures, and supports the efficient generation of protein sequences and structures.
[0035] The present invention uses the folding symbol generation model FoldGPT to generate protein structures. This model adopts an autoregressive generation method. According to the given partial sequence and structure data, the missing amino acid residues and their spatial structures can be generated. The application of the FoldGPT model enables protein structure prediction to not only handle the known parts but also intelligently infer and generate the unknown parts, solving the problem that traditional methods in the prior art cannot handle partially missing or unmodeled structures. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a schematic flow chart of the method of the present invention;
[0037] Figure 2 is a schematic diagram of the protein sequence and structure data encoding process in the present invention;
[0038] Figure 3 is a schematic diagram of the soft conditional vector quantization process in the present invention;
[0039] Figure 4 is a schematic diagram of the protein structure generation process in the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0040] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0041] Embodiment
[0042] As Figure 1 shown, a protein structure representation method based on vector quantization includes the following steps:
[0043] S1. Divide the protein data into a known part and an unknown part, where the known part includes a known protein sequence and a known protein structure;
[0044] S2. For the known part, perform protein sequence structure encoding and soft-conditional vector quantization processing in sequence to obtain a quantized representation of the known part;
[0045] Among them, protein sequence structure encoding is to encode the protein sequence and structure data respectively, and map them to a continuous latent space representation, that is, obtain the folding symbol, so as to realize the unified representation of protein sequence and three-dimensional structure information;
[0046] Soft-conditional vector quantization processing is to use a soft-conditional vector quantizer to quantize the continuous latent representation and convert it into a discrete latent space encoding;
[0047] S3. Input the quantized representation of the known part into the folding symbol generation model FoldGPT, and output the predicted quantized representation;
[0048] S4. Perform protein sequence structure decoding on the predicted quantized representation to generate the amino acid residues and their spatial structures of the unknown part.
[0049] Among them, the protein sequence structure encoding process is as Figure 2 shown, mainly including:
[0050] (1) Encode the protein sequence, and convert the amino acid sequence S = {s i : 1 ≤ i ≤ n} (where s i is the i-th amino acid) into a numerical representation form. In this embodiment, word embedding technology is used to encode each amino acid residue and map it into a high-dimensional space to obtain the corresponding continuous vector representation.
[0051] (2) Encode the protein structure and encode the three-dimensional spatial structure of the protein data. In this embodiment, a graph neural network model is used to model the spatial structure of the protein and convert it into a continuous high-dimensional vector representation, retaining the spatial topological information of the structure.
[0052] (3) Fusion of protein sequence and structure information, fuse the sequence information and structure information of the protein to generate a comprehensive latent representation h = q enc(S, A), that is, the folding symbol. This latent representation not only includes the sequence information of amino acids but also contains its spatial conformation information, which can provide more comprehensive features for the training of subsequent models.
[0053] Subsequently, the latent representation of protein data is quantified using the soft conditional vector quantization method. This process is as Figure 3 shown and mainly includes:
[0054] (1) Latent representation input: The continuous latent representation obtained by encoding protein sequence and structure data is used as the input and passed to the soft conditional vector quantizer z = Q(h). This latent representation is a high-dimensional vector containing the latent information of protein sequence and spatial structure.
[0055] (2) Constructing the codebook where Bit(·) compresses the hidden layer representation z into a 0 / 1 code of length log2(m), and then uses a multi-layer perceptron to map the code vectors in the codebook into d-dimensional vectors, that is, v j = ConditionNe(Bit(z, log2(m))).
[0056] (3) Codebook alignment to obtain the quantized hidden layer representation where SoftVQ(·) represents a differentiable vector quantization operation.
[0057] (4) In the training stage, the decoder is used to decode the quantized hidden layer representation to obtain the reconstructed structure Using the gradient descent method, optimize the error between the reconstructed structure and the real structure A, as well as the error between the reconstructed sequence and the real sequence S;
[0058] In the evaluation stage, Bit(i, log2(m)) is used as the quantized discrete code.
[0059] This embodiment applies the above technical solution. As Figure 4 shown, the protein structure generation process based on the folding symbol generation model FoldGPT mainly includes:
[0060] (1) Divide the protein sequence and structure data into known parts and unknown parts, where: the unknown part includes the residue sequence and structure data to be generated. I m is the index set of known residues;
[0061] The known part includes the determined residue sequence and structure data. Ik is a set of indices of known residues.
[0062] (2) Discretize the protein sequence and structure data using the above soft conditional vector quantization method, and map to discrete latent codes (z1, z2, …, z n ).
[0063] (3) Based on the latent codes of the known part Use the fold symbol generation model FoldGPT to generate the latent codes of the unknown part through autoregressive generation Then, through protein sequence structure decoding, the missing amino acid residues and their spatial structures can be obtained.
[0064] This embodiment also maximizes the conditional probability objective function:
[0065]
[0066] To further optimize the parameters of the fold symbol generation model, thereby improving the accuracy and quality of protein structure generation.
[0067] In summary, this solution consists of three parts: protein sequence and structure data encoding, soft conditional vector quantization, and protein structure generation based on the FoldGPT model. Through these three parts of the design, an efficient protein structure representation can be obtained, and a high-precision protein three-dimensional structure can be effectively generated.
[0068] Specifically, traditional protein modeling methods usually process the protein sequence and structure separately, resulting in the disconnection of the relationship between the protein sequence and structure. However, the protein sequence and structure are closely related in terms of function and function prediction, and the sequence and structure must be considered together to more accurately understand the characteristics and functions of proteins. To solve this problem, this solution proposes a protein structure language representation method based on a fold tokenizer. This method maps the amino acid sequence of the protein and its corresponding three-dimensional structure information to a common discrete symbol space through a fold tokenizer. The fold tokenizer converts the amino acid type of the protein and the corresponding spatial geometric information (such as spatial coordinates, geometric angles, etc.) into discrete symbols to generate fold symbols. These fold symbols not only represent the sequence information of the protein but also contain the geometric features of the protein structure, thus realizing the unified representation of the protein sequence and structure.
[0069] To ensure that the generated folding symbols can effectively retain the information of the original protein sequence and structure, this solution adopts the method of reconstruction loss function. This loss function guarantees the integrity of the protein sequence and structure by minimizing the difference between the folding symbols and the original protein structure. This means that through training the folding tokenizer, the folding symbols can fully reflect the structural characteristics and amino acid sequence characteristics of the protein, ensuring that the generated protein structure can be consistent with the actual three-dimensional structure. The introduction of the reconstruction loss function not only enables the folding symbols to be a good representation of the protein sequence and structure, but also ensures that the generated protein language can gradually recover information close to the real protein structure during the generation process. This enables this solution to achieve good performance in various protein generation, prediction, and reconstruction tasks.
[0070] This solution also introduces the soft conditional vector quantization method to improve the protein structure reconstruction process. Traditional vector quantization methods show certain limitations in dealing with protein structure reconstruction, especially when precise reconstruction of complex structural information is required. To overcome this challenge, this solution proposes the soft conditional vector quantization method, which improves the quality of structure reconstruction by performing "soft queries" throughout the codebook space. The soft conditional vector quantization method is different from traditional hard quantization methods, which can only quantize each vector to the closest code vector. Instead, the soft conditional vector quantization method allows each vector to be weighted according to its relationship with multiple codebook vectors, thus providing more detailed and accurate information during the reconstruction process. This method effectively solves the contradiction between accuracy and generation ability in protein structure reconstruction and improves the protein structure generation and prediction ability.
[0071] After realizing the generation and representation of folding symbols, this solution further proposes a generation model based on folding symbols - the folding symbol generation model FoldGPT. The folding symbol generation model is an autoregressive generation model that can accept the folding symbols of proteins as input and sequentially generate the amino acid sequence and three-dimensional structure of proteins. Different from traditional protein generation methods based on angles or continuous coordinates, the folding symbol generation model directly uses folding symbols as the input of the generation process. It gradually generates the structure and sequence of proteins through an autoregressive mechanism. This process can not only consider the sequence and structural characteristics of proteins simultaneously, but also generate more accurate protein structures through the sequence-structure co-generation model. In practical applications, the folding symbol generation model can be applied to tasks such as protein design, antibody design, and protein folding prediction. In antibody design, the folding symbol generation model can generate new antibody sequences and their corresponding structures, and improve the accuracy and efficiency of antibody design by optimizing the generation process.
[0072] The method of protein structure language representation based on vector quantization proposed in this scheme has significant advantages and broad application prospects. By mapping the sequence and structure information of proteins into a unified discrete space, the folding symbols can simplify the representation of proteins while retaining information, making the protein modeling task more efficient. Especially when dealing with complex and diverse protein data, it shows good scalability and adaptability. It can not only enhance the modeling flexibility of protein sequences and structures, but also be widely applied to multiple biomedical fields such as protein design, function prediction, and antibody design, providing more accurate and efficient prediction capabilities.
Claims
1. A protein structure representation method based on vector quantization, characterized in that, It includes the following steps: S1. Divide the protein data into a known part and an unknown part, where the known part includes the known protein sequence and the known protein structure; S2. For the known part, perform protein sequence structure encoding and soft conditional vector quantization processing in sequence to obtain the quantized representation of the known part; S3. Input the quantized representation of the known part into the fold symbol generation model FoldGPT to output the predicted quantized representation; S4. Perform protein sequence structure decoding on the predicted quantized representation to generate the amino acid residues and their spatial structures of the unknown part.
2. The method for representing protein structures based on vector quantization according to claim 1, wherein The specific process of step S1 is as follows: Divide the protein sequence and structure data into known and unknown parts, where the unknown part includes the residue sequence and structure data to be generated. I m is the index set of known residues; Known part including a determined residue sequence and structure data, I k is a set of indexes of known residues.
3. The method for representing protein structure based on vector quantization according to claim 2, wherein In step S2, the protein sequence structure encoding is specifically to encode the protein sequence and the structure data respectively and map them to a continuous latent space representation to obtain a continuous latent representation.
4. A method for representing protein structures based on vector quantization according to claim 3, wherein, In step S2, the soft conditional vector quantization processing is specifically to use a soft conditional vector quantizer to quantize the continuous latent representation and convert it into a discrete latent space encoding.
5. A method for representing protein structures based on vector quantization according to claim 4, characterized in that The specific process of performing protein sequence structure encoding in step S2 includes: S201. Encode the protein sequence, convert the amino acid sequence S = {s i : 1 ≤ i ≤ n} of the protein into a numerical representation form, where s i is the i-th amino acid; S202. Encode the protein structure and encode the three-dimensional spatial structure data of the protein; S203. Fuse the sequence information and structural information of the protein to generate a comprehensive latent representation h = q enc (S, A), which includes not only the sequential information of amino acids but also their spatial conformation information.
6. A method for representing protein structures based on vector quantization according to claim 5, wherein, In step S201, specifically use the word embedding technology to encode each amino acid residue and map it into a high-dimensional space to obtain the corresponding continuous vector representation.
7. A method for representing protein structures based on vector quantization according to claim 5, characterized in that, In step S202, specifically use a graph neural network model to model the spatial structure of the protein and convert it into a continuous high-dimensional vector representation, retaining the spatial topological information of the structure.
8. A method for representing protein structures based on vector quantization according to claim 5, characterized in that The specific process of performing soft conditional vector quantization processing in step S2 includes: S211. Input of the latent representation, use the continuous latent representation obtained by encoding the protein sequence and the structure data as the input and transfer it to the soft conditional vector quantizer z = Q(h); S212. Construct a codebook Among them, Bit(·) is used to compress the hidden layer representation z into a 0 / 1 code with a length of log2(m), and then use the multi-layer perceptron ConditionNet: Map the code vectors in the codebook into d-dimensional vectors, that is, v j = ConditionNe(Bit(z, log2(m))); S213. Align the codebook to obtain the quantized hidden layer representation where SoftVQ(·) represents the differentiable vector quantization operation; S214. During the training phase, use the decoder to decode the quantized hidden layer representation to obtain the reconstructed structure And perform iterative training in combination with the loss function to minimize the difference between the folded symbol and the original protein structure, where the folded symbol is the potential representation obtained by encoding the protein sequence structure; during the evaluation phase, use Bit(i, log2(m)) as the quantized discrete code.
9. A method for representing protein structures based on vector quantization according to claim 8, characterized in that In the training phase, step S214 specifically uses the gradient descent method to optimize the errors between the reconstructed structure and the real structure A, as well as the errors between the reconstructed sequence and the real sequence S.
10. A method for representing protein structures based on vector quantization according to claim 8, characterized in that, The specific process of step S3 is as follows: The discrete latent codes of the known part are input into the folding symbol generation model FoldGPT, and the latent codes of the unknown part are generated through an autoregressive generation method