Protein multi-modal joint modeling method and system based on cross-modal alignment

Through the cross-modal alignment method, multimodal information of proteins is integrated, features are extracted using geometric graph neural networks and large language models to generate unified invariant feature representations, solving the gap in multimodal information processing in protein analysis and achieving more efficient protein analysis and prediction.

CN120260664AActive Publication Date: 2025-07-04ZHEJIANG LAB

Patent Information

Application Number
CN202510736818.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The prior art lacks a mature framework that can uniformly process protein multimodal information, limiting the accuracy and reliability of protein analysis and prediction.

Method used

Through the cross-modal alignment method, the sequence information, text description information and three-dimensional structural information of proteins are integrated into the large language model, and features are extracted using geometric graph neural network and protein sequence model, and a unified invariant feature representation is generated through the projection module to achieve the coordinated work of multimodal information.

Benefits of technology

It significantly improves the accuracy and reliability of protein analysis prediction, provides a powerful tool for complex protein research, and promotes the in-depth development of protein scientific research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260664A_ABST
    Figure CN120260664A_ABST
Patent Text Reader

Abstract

The invention discloses a protein multi-modal joint modeling method and system based on cross-modal alignment, and aims to realize unified representation and processing of protein text description, sequence information and structural features so as to improve the accuracy and generalization ability of complex protein analysis and prediction tasks and open research tasks. The method comprises the following steps: firstly collecting and preprocessing protein multi-modal data to generate standardized representation, extracting features by using a geometric graph neural network and a protein sequence model, aligning the features through a projection module, transmitting the features to a large language model, performing projection fusion to generate uniform invariant features, and finally finishing a protein analysis and prediction task in combination with the invariant features. According to the method, multi-modal information of the protein can be effectively integrated, efficient alignment and fusion of cross-modal features are achieved, the ability to understand and predict complex characteristics of the protein is improved, and powerful support is provided for precise bioinformatics analysis and biological medicine research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics, and specifically to a method and system for protein multimodal joint modeling based on cross-modal alignment. Background Art

[0002] Proteins are key molecules in life science research, and their functions and behaviors are usually jointly determined by their sequences and three-dimensional structures. Each protein has multiple modalities of information, including sequence information based on UniProt (Universal Protein) files and three-dimensional structure information based on PDB (Program Database) files. For these different modal data, many mature models have been developed. For example, geometric graph neural networks perform well in processing the three-dimensional structures of proteins, while protein sequence models have advantages in parsing and understanding sequence information. However, there is still a lack of a mature framework that can uniformly process this multimodal information.

[0003] In recent years, large language models have achieved great success, demonstrating their powerful capabilities in text understanding and generation. These models can not only process natural language but also effectively integrate and utilize scientific knowledge. However, large language models have deficiencies in understanding and processing three-dimensional spatial data, which limits their application potential in protein analysis.

[0004] The problems of protein analysis and prediction are usually complex and diverse. Establishing a unified framework that can simultaneously process multimodal information is of great significance for comprehensively analyzing protein functions.

[0005] The present invention explores the possibility of integrating large language models, geometric graph deep learning models, and protein sequence models in the field of proteins. By aligning the multimodal representations of these models, a unified framework is proposed to better handle complex protein analysis and prediction problems. Specifically, this framework aims to organically combine the sequence information and three-dimensional structure information of proteins, utilize the advantages of their respective models to provide a more comprehensive protein description. Through model alignment, their collaborative work in protein analysis tasks is achieved, improving the accuracy and reliability of protein analysis and prediction.

[0006] The present invention not only fills the gap in the existing multimodal protein analysis framework but also provides a powerful tool for complex protein research and open problems. This innovative framework is expected to promote the in-depth development of protein science research and have extensive application value in fields such as biomedicine. Summary of the Invention

[0007] Aiming at the problem of the lack of a mature framework for unified processing of multimodal information in the current protein analysis field, the present invention proposes a protein multimodal joint modeling method and system based on cross-modal alignment.

[0008] The purpose of the present invention is achieved through the following technical solutions: The present invention provides a protein multimodal joint modeling method based on cross-modal alignment, comprising the following steps: Step 1: Collect and preprocess the multimodal data of proteins to generate a standardized representation form; Step 2: Use a geometric graph neural network to extract the invariant feature representation and the equivariant feature representation of the three-dimensional structure of proteins, and obtain the feature representation by processing the sequence information of proteins through a protein sequence model ; Step 3: Input the feature representations and into a projection module to generate feature representations aligned with the input space of the large language model and ; Step 4: Transfer the aligned feature representations and as well as the text description information of proteins to the large language model to output the semantic feature representation ; Step 5: Project the semantic feature representation back to the space corresponding to the invariant feature representation through the projection module, and at the same time project the feature representation to the space corresponding to the invariant feature representation , and finally fuse , and to generate a unified invariant feature representation ; Step 6: Combine the invariant feature representation and the three-dimensional equivariant feature representation to complete the analysis and prediction tasks of proteins.

[0009] Furthermore, Step 1 includes: Collect multimodal data of proteins. The data of each protein includes sequence information and text description information from the UniProt database, as well as three-dimensional map structures from the PDB database; Extract the protein sequence information S from the UniProt file, convert the content other than the sequence information in the UniProt file into natural language descriptions that can be understood by the large language model, that is, text description information T, and extract the three-dimensional structure information G of the protein from the PDB file; Preprocess the sequence information, text description information, and three-dimensional structure information to ensure the consistency and usability of the input format; Find the relevant UniProt ID from the annotations of the PDB entry to achieve the paired link between the three-dimensional structure information of the same protein and the sequence and text description information; For proteins lacking UniProt entries, obtain the sequence information of the protein from the PDB entry, and the corresponding text description information is missing; For proteins lacking PDB entries, use a structure prediction tool to predict the three-dimensional structure of the protein; In the subsequent processing, higher weights are assigned to the modal data that is clearly present in the database.

[0010] Furthermore, step 1 further includes: The protein information described in the text uses a unified modular language structure: "The protein structure [protein ID] has a sequence length of [number] amino acids, involving the following chains: [chain], and this protein is named [protein name], originating from the organism [organism]"; Add other characteristic descriptions of the protein to the text description, including statistical information and 3D structure information. For the 3D structure information, select a certain node in the protein structure diagram as the anchor point and describe the relative positions of other nodes with respect to this node; Provide a concise description of the task in the text description to enable the large language model to quickly identify the task objective and release domain-specific knowledge.

[0011] Furthermore, step 2 includes: Input the three-dimensional structure information G of the protein obtained in step 1 into the geometric graph neural network to extract the invariant feature representation of the protein and the equivariant feature representation ; Input the protein sequence information S obtained in step 1 into the protein sequence model to obtain the feature representation .

[0012] Furthermore, step 3 includes: Use the trained projection module to project the feature representations and obtained in step 2 into the input space of the large language model to generate feature representations aligned with the input space of the large language model and ; The projection module is used to ensure that the aligned features can interact with the text description information T in the same semantic space.

[0013] Further, step 4 includes: For the text description information T obtained in step 1, use the tokenizer and embedding layer of the large language model to obtain the corresponding word embedding features. Subsequently, according to the specific task, the obtained word embedding features are concatenated with the feature representation obtained in step 3 and to form a combined feature and input it into the large language model, and combine the knowledge of the large language model to obtain the semantic feature representation .

[0014] Further, step 5 includes: Use the trained projection module to project the semantic feature representation obtained in step 4 back to the space corresponding to , and at the same time project to the space corresponding to , fuse the three aligned features , and to generate a unified invariant feature representation .

[0015] Further, step 6 includes: Combine the feature representation obtained in step 5 and the three-dimensional equivariant feature representation obtained in step 2 , and complete the analysis and prediction task of the protein; for the invariant task of protein function prediction, only use the unified invariant feature representation for prediction; for the task whose core goal is to predict the three-dimensional coordinates, combine the feature representation and the three-dimensional equivariant feature representation for prediction.

[0016] The present invention also provides a cross-modal alignment-based protein multi-modal joint modeling system for implementing the above method, including: A data preprocessing module, which is used to collect and preprocess the multi-modal data of the protein; A feature extraction module, which is used to extract three-dimensional structure features, sequence features and equivariant features from the multi-modal data of the protein and generate a preliminary feature representation; A cross-modal feature alignment module, which is used to align the extracted protein features with the input space of the large language model; A multi-modal semantic fusion module, which is used to fuse the multi-modal features of the protein; An inverse projection and feature fusion module for fusing the semantic features generated by a large language model with the feature representations generated by a geometric graph neural network and a protein sequence model to generate a unified invariant feature representation; A task adaptive prediction module for completing the analysis and prediction tasks of proteins.

[0017] The present invention also provides a protein multi-modal joint modeling device based on cross-modal alignment, including a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, the above-mentioned protein multi-modal joint modeling method based on cross-modal alignment is implemented.

[0018] The beneficial effects of the present invention are as follows: The present invention provides a protein multi-modal joint modeling method based on cross-modal alignment, filling the gap in the existing multi-modal protein analysis framework and providing a powerful tool for complex protein research and open problems. By aligning a geometric graph neural network, a protein sequence model, and a large language model, the present invention realizes their collaborative work in protein analysis tasks. When applied to protein analysis tasks, the present invention significantly improves the accuracy and reliability of protein analysis and prediction. Description of the Drawings

[0019] Figure 1 is a flowchart of the protein multi-modal joint modeling method based on cross-modal alignment proposed by the present invention; Figure 2 is a schematic structural diagram of the protein multi-modal joint modeling system based on cross-modal alignment provided by the present invention; Figure 3 is a schematic structural diagram of the protein multi-modal joint modeling device based on cross-modal alignment provided by the present invention. Detailed Embodiments

[0020] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of systems and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0021] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0022] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0023] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the following further elaborates on the present invention in detail in conjunction with specific implementation cases and with reference to the accompanying drawings.

[0024] See Figure 1 , the present invention proposes a protein multimodal joint modeling method based on cross-modal alignment, and its main steps include: Step 1, collect and preprocess the multimodal data of proteins to generate a standardized representation.

[0025] Specifically, collect the multimodal data of proteins. The data of each protein includes sequence information and text description information from the UniProt database and three-dimensional structure from the PDB database. Extract the sequence information S of the protein from the UniProt file, convert the content other than the sequence information in the UniProt file into a clear and detailed natural language description T (i.e., text description information) that can be understood by the large language model, and extract the three-dimensional structure information G of the protein from the PDB file. Preprocess the sequence information, text description information, and three-dimensional structure information to ensure the consistency and usability of the input format.

[0026] It should be noted that the amino acid sequence itself is very long and is excluded from the text description information.

[0027] Find the relevant UniProt ID from the annotations of the PDB entry to achieve the paired link of the three-dimensional structure information and the sequence and text description information of the same protein.

[0028] For proteins lacking UniProt entries, obtain the sequence information of the protein from the PDB entry, and the corresponding text description information is missing; for proteins lacking PDB entries, use structure prediction tools such as AlphaFold, RoseTTAFold, etc. to predict the three-dimensional structure of the protein.

[0029] In the subsequent processing, higher weights are assigned to the modal data that is clearly present in the database.

[0030] The protein information described in the text uses a unified modular language structure, similar to: "The protein structure [Protein ID] has a sequence length of [number] amino acids, involving the following chains: [chain], and this protein is named [Protein Name], originating from the organism [Organism]"; Regular expressions or natural language processing techniques can be used to convert non-sequence information into structured natural language descriptions.

[0031] It should be noted that other characteristic descriptions, statistical information, and even 3D structure information of the protein can be added to the text description. For 3D structure information, a certain node in the protein structure diagram can be selected as an anchor point to describe the relative positions of other nodes with respect to this node.

[0032] A concise description of the task can be provided in the text description to help the large language model (LLM) quickly identify the task objective and also help the LLM release domain-specific knowledge.

[0033] Here, the anchor point refers to the key reference node in the structure, which can be determined based on the localization of functional sites or the method of selecting key structural nodes. Functional sites include key amino acid residues (such as conserved sequence regions) and catalytic active centers, etc. Key structural nodes are generally obtained through geometric analysis methods.

[0034] Although geometric graph neural networks also process 3D structure information, due to differences in the abstraction levels, it forms the complementarity between multi-modal information. Through semantic, global, and noise-resistant information complementarity, it forms multi-granularity knowledge fusion and cross-modal error correction mechanisms.

[0035] Use bioinformatics tools (such as Biopython) to process PDB files, extract atomic coordinates, bond structures, and spatial relationships, and convert them into graph structure representations.

[0036] Step 2, based on the protein multi-modal data obtained in Step 1, use geometric graph neural networks to extract invariant feature representations of the protein three-dimensional structure and equivariant feature representations , and obtain feature representations by processing the sequence information of the protein through a protein sequence model .

[0037] Specifically, input the three-dimensional structure information G of the protein obtained in Step 1 into the geometric graph neural network, which can be any suitable equivariant model in the field and must satisfy the E(3) symmetry requirement to extract the invariant feature representation of the protein and equivariant feature representations ; Input the protein sequence information S obtained in step 1 into a protein sequence model, such as the ESM (Evolutionary Scale Modeling) model, to obtain a feature representation .

[0038] It is easy to understand that the E(3) symmetry requirement means that when the model processes geometric data in three-dimensional space (such as protein structures), it needs to satisfy symmetry invariance or equivariance with respect to the three-dimensional Euclidean group (Euclidean group E(3)).

[0039] Adjust the hyperparameters of the geometric graph neural network, such as the number of layers and the node feature dimension, to optimize the feature extraction effect.

[0040] Step 3, input the feature representation and into the projection module to generate a feature representation aligned with the input space of the large language model and .

[0041] Specifically, use the trained projection module to project the feature representations and obtained in step 2 into the input space of the large language model to generate feature representations aligned with the input space of the large language model and . The projection module can be a multilayer perceptron (MLP) or other suitable mapping models to ensure that the aligned features can interact with the text description information T in the same semantic space.

[0042] It should be noted that the training of the projection module can be carried out in two steps. In the first step, the projection module is trained separately to ensure that it can retain the key information of the original data. In the second step, the projection modules and the large language model are jointly trained to maximize the complementarity between the modules.

[0043] Step 4, transfer the aligned features and along with information such as the text description of the protein to the large language model to output a semantic feature representation .

[0044] Specifically, for the protein text description information T obtained in step 1, use the tokenizer and embedding layer of the large language model to obtain the corresponding word embedding features. Subsequently, according to the different specific tasks, the obtained word embedding features are combined with the feature representations and Perform splicing in an appropriate manner to form a combined feature input, input the combined feature into a large language model, and obtain a semantic feature representation by integrating the knowledge of the large language model . According to specific task requirements, task-specific prompts can be designed to guide the large language model to generate more accurate semantic features. For example, for the ligand binding site prediction task, the prompt is: "Identify possible ligand binding sites based on the characteristic information of the protein". For the molecular dynamics simulation task, the prompt is: "Predict the trajectory of the next F frames based on the previous trajectory information".

[0045] Adjust the input layer of the large language model so that it can process the combined feature input. The specific adjustment process is as follows: After projection, and have the same dimension as the embedding dimension of the LLM. Use the gated attention mechanism to splice the feature representations and with the text embedding features, not only keeping the dimension of the spliced result unchanged, but also effectively integrating the feature information from different sources.

[0046] Design specific fusion strategies, such as weighted summation or attention mechanism, to optimize the feature fusion effect.

[0047] Step 5, project back to the space corresponding to through the projection module, and at the same time project to the space corresponding to , and finally fuse , and to generate a unified invariant feature representation .

[0048] Specifically, use the trained projection module to project the semantic feature representation obtained in step 4 back to the space corresponding to , and at the same time project to the space corresponding to , and fuse the three aligned features , and to generate a unified invariant feature representation .

[0049] Step 6, combine and the three-dimensional equivariant feature representation to complete the protein analysis and prediction task.

[0050] Specifically, combine the feature representation obtained in step 5 and the three-dimensional equivariant feature representation obtained in step 2, complete the analysis and prediction tasks of proteins. For invariant tasks such as protein function prediction, only use a unified invariant feature representation for prediction. For tasks whose core goal is to predict three-dimensional coordinates, combine the feature representation and three-dimensional equivariant feature representation for prediction.

[0051] It should be noted that during the training process of the entire multimodal model, the parameters of the LLM and the protein sequence model are always frozen, which significantly reduces the computational cost and time overhead. The entire training process is divided into the following 3 stages.

[0052] Stage 1: First, train the geometric graph neural network alone, and then freeze the parameters of the geometric graph neural network.

[0053] Stage 2: Design a projection module to project the feature representations and obtained in step 2 into the input space of the large language model to generate feature representations and aligned with the input space of the large language model, project the semantic feature representation obtained in step 4 back to the space corresponding to to obtain the feature , and at the same time project to the space corresponding to to obtain the feature . Use a contrastive learning loss function to train the projection module so that the projected features and are aligned in the same space as . For the same protein, and as well as and are positive sample pairs. For different proteins, and of other proteins as well as and of other proteins are negative sample pairs.

[0054] and The contrastive loss is defined as:

[0055] and The contrastive loss is defined as:

[0056] The total contrastive loss is:

[0057] Among them, represents the similarity measure between and represents the similarity measure between and is the temperature parameter, and N is the number of negative samples.

[0058] In stage 3, for different downstream tasks, train the method of fusing three-modal features , and . Here, a task-aware readout function is adopted to calculate the attention weights between each modal feature and the task-specific query, and then the three-modal features are weighted to obtain .

[0059] The protein multi-modal joint modeling system based on cross-modal alignment provided by the present invention is described below. The protein multi-modal joint modeling system based on cross-modal alignment described below can be correspondingly referred to the protein multi-modal joint modeling method described above.

[0060] Figure 2 is a schematic structural diagram of the protein multi-modal joint modeling system based on cross-modal alignment provided by the present invention. As Figure 2 shown, a set of protein multi-modal joint modeling systems based on cross-modal alignment provided by the present invention includes: A data preprocessing module for collecting and preprocessing the multi-modal data of proteins; A feature extraction module for extracting three-dimensional structure features, sequence features, and equivariant features from the multi-modal data of proteins to generate a preliminary feature representation; A cross-modal feature alignment module for aligning the extracted protein features with the input space of the large language model; A multi-modal semantic fusion module for fusing the multi-modal features of proteins; A reverse projection and feature fusion module for fusing the semantic features generated by the large language model with the feature representations generated by the geometric graph neural network and the protein sequence model to generate a unified invariant feature representation; A task adaptive prediction module for completing the analysis and prediction tasks of proteins.

[0061] This embodiment effectively aligns the protein multi-modal data and effectively improves the reliability of protein model analysis and prediction.

[0062] Corresponding to the embodiments of the protein multimodal joint modeling method based on cross-modal alignment described above, the present invention also provides an embodiment of a protein multimodal joint modeling device based on cross-modal alignment.

[0063] See Figure 3 , an embodiment of a protein multimodal joint modeling device based on cross-modal alignment provided by an embodiment of the present invention includes a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement the protein multimodal joint modeling method based on cross-modal alignment in the above embodiments.

[0064] An embodiment of a protein multimodal joint modeling device based on cross-modal alignment provided by the present invention can be applied to any device with data processing capabilities. Such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 3 shown, it is a hardware structure diagram of any device with data processing capabilities where a protein multimodal joint modeling device based on cross-modal alignment provided by the present invention is located. In addition to Figure 3 the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities where the device in the embodiment is located usually also includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.

[0065] The specific implementation process of the functions and roles of each unit in the above device can be specifically seen in the implementation process of the corresponding steps in the above method, which will not be elaborated here.

[0066] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present invention's solution. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0067] Corresponding to the embodiments of the foregoing protein multimodal joint modeling method based on cross-modal alignment, an embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the protein multimodal joint modeling method based on cross-modal alignment in the foregoing embodiments is implemented.

[0068] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.

[0069] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0070] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or modifications made in accordance with the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A protein multimodal joint modeling method based on cross-modal alignment, characterized in that, It includes the following steps: Step 1: Collect and preprocess the multimodal data of proteins to generate a standardized representation; Step 2: Use a geometric graph neural network to extract invariant feature representations of the three-dimensional structure of proteins and equivariant feature representations , and obtain feature representations by processing the sequence information of proteins through a protein sequence model ; Step 3: Input the feature representations and into the projection module to generate feature representations aligned with the input space of the large language model and ; Step 4: Transfer the aligned feature representation and along with the text description information of the protein to the large language model to output the semantic feature representation ; Step 5: Project the semantic feature representation back to the space corresponding to the invariant feature representation , while projecting the feature representation to the space corresponding to the invariant feature representation , and finally fusing , and to generate a unified invariant feature representation ; Step 6: Combine the invariant feature representation and the three-dimensional equivariant feature representation to complete the analysis and prediction task of proteins.

2. The protein multimodal joint modeling method based on cross-modal alignment according to claim 1, wherein Step 1 includes: Collect the multimodal data of proteins. The data of each protein includes sequence information and text description information from the UniProt database and three-dimensional map structures from the PDB database; Extract the protein sequence information S from the UniProt file, convert the content other than the sequence information in the UniProt file into a natural language description that can be understood by the large language model, that is, text description information T, and extract the three-dimensional structure information G of the protein from the PDB file; Preprocess the sequence information, text description information, and three-dimensional structure information to ensure the consistency and usability of the input format; Find the relevant UniProt ID from the annotations of the PDB entry to achieve the pairing link between the three-dimensional structure information of the same protein and the sequence and text description information; For proteins lacking UniProt entries, obtain the sequence information of the protein from the PDB entry, and the corresponding text description information is missing; For proteins lacking PDB entries, use a structure prediction tool to predict the three-dimensional structure of the protein; In the subsequent processing, higher weights are given to the modal data that is clearly present in the database.

3. The protein multimodal joint modeling method based on cross-modal alignment according to claim 1, characterized in that, Step 1 further includes: The protein information described in the text uses a unified modular language structure: "The protein structure [protein ID] has a sequence length of [number] amino acids, involving the following chains: [chain], this protein is named [protein name], and is derived from the organism [organism]"; Add other characteristic descriptions of the protein to the text description, including statistical information and 3D structure information. For the 3D structure information, select a certain node in the protein structure diagram as the anchor point and describe the relative positions of other nodes with respect to this node; Provide a concise description of the task in the text description to enable the large language model to quickly identify the task objective and release domain-specific knowledge.

4. The protein multimodal joint modeling method based on cross-modal alignment according to claim 2, wherein Step 2 includes: Input the three-dimensional structure information G of the protein obtained in step 1 into the geometric graph neural network to extract the invariant feature representation of the protein and the equivariant feature representation ; Input the protein sequence information S obtained in step 1 into the protein sequence model to obtain the feature representation .

5. The protein multimodal joint modeling method based on cross-modal alignment according to claim 2, wherein, Step 3 includes: Using the trained projection module, project the feature representation obtained in step 2 and into the input space of the large language model to generate a feature representation aligned with the input space of the large language model and ; the projection module is used to ensure that the aligned features can interact with the text description information T in the same semantic space.

6. The protein multimodal joint modeling method based on cross-modal alignment according to claim 2, wherein Step 4 includes: For the text description information T obtained in step 1, use the tokenizer and embedding layer of the large language model to obtain the corresponding word embedding features. Subsequently, according to the different specific tasks, concatenate the obtained word embedding features with the feature representations obtained in step 3 and to form joint features and input them into the large language model, and combine the knowledge of the large language model to obtain semantic feature representations .

7. The protein multimodal joint modeling method based on cross-modal alignment according to claim 1, wherein Step 5 includes: Using the trained projection module, project the semantic feature representation obtained in step 4 back to the space corresponding to , and at the same time project to the space corresponding to . Then fuse the three aligned features , and to generate a unified invariant feature representation .

8. The protein multimodal joint modeling method based on cross-modal alignment according to claim 1, wherein Step 6 includes: Combined with the feature representation obtained in step 5 and the three-dimensional equivariant feature representation obtained in step 2 , complete the analysis and prediction task of proteins; for the invariant task of protein function prediction, only use the unified invariant feature representation for prediction; for the task whose core goal is to predict three-dimensional coordinates, combine the feature representation and the three-dimensional equivariant feature representation for prediction.

9. A protein multimodal joint modeling system based on cross-modal alignment for implementing the method as claimed in claim 1, characterized in that, It includes: A data preprocessing module for collecting and preprocessing the multimodal data of proteins; A feature extraction module for extracting three-dimensional structure features, sequence features, and equivariant features from the multimodal data of proteins to generate a preliminary feature representation; A cross-modal feature alignment module for aligning the extracted protein features with the input space of the large language model; A multimodal semantic fusion module for fusing the multimodal features of proteins; A reverse projection and feature fusion module for fusing the semantic features generated by the large language model with the feature representations generated by the geometric graph neural network and the protein sequence model to generate a unified invariant feature representation; A task adaptive prediction module for completing the analysis and prediction tasks of proteins.

10. A protein multimodal joint modeling device based on cross-modal alignment, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the processor executes the executable code, it implements the cross-modal alignment-based protein multimodal joint modeling method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Protein expression model pre-training and protein interaction prediction method and device

    CN114333982A

  • Protein reverse folding method and device based on multi-modal pre-training large model

    CN117727365A

  • Cross-modal model training method and device, electronic equipment and storage medium

    CN119091971A

  • DNA binding residue prediction method based on multi-modal protein language model

    CN119418777A

  • Method for Sequence-Based Prediction of Controlled Terms and Generating Protein Sequences from Controlled Terms using Enhanced Large Language Models

    US20240404632A1

Cited By

  • Protein function prediction method, model training method, device, equipment and medium

    CN122224282A