An information processing method, device and computer readable storage medium

CN114974398BActive Publication Date: 2026-09-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110203184.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-23
Publication Date
2026-09-18
Estimated Expiration
2041-02-23

AI Technical Summary

Technical Problem

[0003]在传统的虚拟药物筛选、蛋白质性质分析的过程需要消耗大量的资源,使得研发周期大幅度增加的同时研发费用巨大,造成资源的浪费,因此,将人工智能技术应用于药物筛选中,可以大幅度的减少相关实验所需的时间和费用

Benefits of technology

[0013] This application's embodiments acquire protein sample information and decompose it into amino acid residue graph structure data. The amino acid residue graph structure data is input into a graph neural network, outputting multiple amino acid residue node vectors. Multiple amino acid microenvironment samples are constructed based on the correlation between each amino acid residue in the protein sample information. These amino acid microenvironment samples are clustered to obtain multiple cluster types of amino acid microenvironment sample sets. The multiple cluster types are used as label information and input with the multiple amino acid residue node vectors into a preset classification model for training, resulting in a trained preset classification model. The amino acid residues to be detected are then classified according to the preset classification model. Thus, by representing amino acid residues with vectors and based on self-supervised learning, a model for identifying amino acids is obtained. Compared to existing schemes that manually label protein information, this application can reasonably and effectively perform self-supervised learning, using a large amount of unlabeled protein sample information for model training, greatly improving the efficiency of information processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114974398B_ABST
    Figure CN114974398B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an information processing method and device and a computer readable storage medium. Embodiments of the present application obtain protein sample information, and decompose the protein sample information into amino acid residue graph structure data. The amino acid residue graph structure data is input into a graph neural network to output a plurality of amino acid residue node vectors. A plurality of amino acid microenvironment samples are constructed according to the correlation degree relationship between each amino acid residue in the protein sample information. The amino acid microenvironment samples are clustered to obtain a plurality of sets of amino acid microenvironment samples of clustering types. The plurality of clustering types are taken as label information and input into a preset classification model together with the plurality of amino acid residue node vectors to train the preset classification model to obtain a trained preset classification model. The preset classification model is used to classify a to-be-detected amino acid residue. In this way, the amino acid residues are expressed in vectors, and a model for identifying amino acids is obtained based on self-supervised learning, which greatly improves the efficiency of information processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to an information processing method, apparatus, and computer-readable storage medium. Background Technology

[0002] Artificial Intelligence (AI) is a comprehensive technology within computer science that studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities. AI technology is a multidisciplinary field, encompassing a wide range of areas, including natural language processing and machine learning / deep learning. It is believed that with technological advancements, AI will be applied in more fields and play an increasingly important role.

[0003] Traditional virtual drug screening and protein property analysis processes consume a lot of resources, significantly increasing the research and development cycle and resulting in huge research and development costs and wasting resources. Therefore, applying artificial intelligence technology to drug screening can significantly reduce the time and cost required for related experiments.

[0004] In the process of researching and practicing existing technologies, the inventors of this application have found that in existing technologies, there is relatively little protein information that has been manually annotated, and adding new annotated data requires the knowledge of domain experts, resulting in high information processing costs and low efficiency. Summary of the Invention

[0005] This application provides an information processing method, apparatus, and computer-readable storage medium, which can improve the efficiency of information processing.

[0006] To address the aforementioned technical problems, this application provides the following technical solutions: An information processing method, comprising: Obtain protein sample information and decompose the protein sample information into amino acid residue diagram structure data; The amino acid residue diagram structure data is input into a graph neural network, which outputs multiple amino acid residue node vectors. Multiple amino acid microenvironment samples were constructed based on the correlation between each amino acid residue in the protein sample information; The amino acid microenvironment samples were clustered to obtain multiple cluster types of amino acid microenvironment sample sets. The multiple clustering types are used as label information and the multiple amino acid residue node vectors are input into a preset classification model for training to obtain the trained preset classification model. The amino acid residues to be detected are classified according to the preset classification model after training.

[0007] An information processing apparatus, comprising: An acquisition unit is used to acquire protein sample information and decompose the protein sample information into amino acid residue diagram structure data. The first input unit is used to input the amino acid residue diagram structure data into the graph neural network and output multiple amino acid residue node vectors. The building unit is used to construct multiple amino acid microenvironment samples based on the correlation between each amino acid residue in the protein sample information. Clustering unit, used to cluster the amino acid microenvironment samples to obtain multiple cluster types of amino acid microenvironment sample sets; The second input unit is used to input the multiple clustering types as label information and the multiple amino acid residue node vectors into a preset classification model for training, so as to obtain the trained preset classification model. The classification unit is used to classify the amino acid residues to be detected according to the trained preset classification model.

[0008] In some embodiments, the second input unit is configured to: The target cluster type is selected from the multiple cluster types for labeling based on the amino acid type corresponding to each amino acid residue node vector; The labeled amino acid residue node vectors are input into a preset classification model for training, resulting in a trained preset classification model.

[0009] In some embodiments, the apparatus further includes: The summation unit is used to sum the amino acid residue node vectors belonging to the same protein to obtain the protein representation vector; The third input unit is used to input the amino acid residue node vector and the corresponding protein vector into the preset binary classification model for training, and obtain the trained preset binary classification model. The determination unit is used to determine whether the amino acid residues to be detected belong to the protein to be detected according to the preset binary classification model.

[0010] A computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to perform the steps in the above-described information processing method.

[0011] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps in the information processing method described above.

[0012] A computer program product or computer program includes computer instructions stored in a storage medium. A processor of a computer device reads the computer instructions from the storage medium and executes the computer instructions, causing the computer to perform the steps of the aforementioned information processing method.

[0013] This application's embodiments acquire protein sample information and decompose it into amino acid residue graph structure data. The amino acid residue graph structure data is input into a graph neural network, outputting multiple amino acid residue node vectors. Multiple amino acid microenvironment samples are constructed based on the correlation between each amino acid residue in the protein sample information. These amino acid microenvironment samples are clustered to obtain multiple cluster types of amino acid microenvironment sample sets. The multiple cluster types are used as label information and input with the multiple amino acid residue node vectors into a preset classification model for training, resulting in a trained preset classification model. The amino acid residues to be detected are then classified according to the preset classification model. Thus, by representing amino acid residues with vectors and based on self-supervised learning, a model for identifying amino acids is obtained. Compared to existing schemes that manually label protein information, this application can reasonably and effectively perform self-supervised learning, using a large amount of unlabeled protein sample information for model training, greatly improving the efficiency of information processing. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a schematic diagram of a scenario for the information processing system provided in an embodiment of this application; Figure 2 This is a flowchart illustrating the information processing method provided in an embodiment of this application; Figure 3 This is another flowchart illustrating the information processing method provided in the embodiments of this application; Figure 4a A schematic diagram of a scenario for the information processing method provided in an embodiment of this application; Figure 4b This is another schematic diagram of a scenario for the information processing method provided in the embodiments of this application; Figure 4c This is another schematic diagram of a scenario for the information processing method provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the information processing device provided in the embodiments of this application; Figure 6 This is a schematic diagram of the server structure provided in the embodiments of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] This invention provides an information processing method, apparatus, storage medium, and computer device. The information processing method can be used in an information processing apparatus. The information processing apparatus can be integrated into a computer device, which can be a terminal with information processing capabilities. This terminal can be a smartphone, tablet, laptop, desktop computer, wearable device, in-vehicle computer, etc., but is not limited to these. The computer device can also be a server, which can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0018] Please see Figure 1 The figure illustrates a scenario for information processing provided in this application. As shown, a computer device acquires protein sample information and decomposes the protein sample information into amino acid residue graph structure data. The amino acid residue graph structure data is input into a graph neural network, which outputs multiple amino acid residue node vectors. Multiple amino acid microenvironment samples are constructed based on the correlation between each amino acid residue in the protein sample information. The amino acid microenvironment samples are clustered to obtain multiple cluster types of amino acid microenvironment sample sets. The multiple cluster types are used as label information and input with the multiple amino acid residue node vectors into a preset classification model for training, resulting in a trained preset classification model. The amino acid residues to be detected are classified according to the preset classification model.

[0019] It should be noted that, Figure 1The illustrated information processing scenario diagram is merely an example. The information processing scenario described in the embodiments of this application is intended to more clearly illustrate the technical solution of this application and does not constitute a limitation on the technical solution provided in this application. Those skilled in the art will understand that, with the evolution of information processing and the emergence of new business scenarios, the technical solution provided in this application is equally applicable to similar technical problems.

[0020] The following sections will provide detailed explanations.

[0021] In this embodiment, the description will be from the perspective of an information processing device, which can be integrated into a server that has storage units and is equipped with a microprocessor and has computing capabilities.

[0022] Please see Figure 2 , Figure 2 This is a flowchart illustrating the information processing method provided in an embodiment of this application. The information processing method includes: In step 101, protein sample information is obtained and decomposed into amino acid residue diagram structure data.

[0023] The protein sample information consists of multiple protein molecules, which are macromolecular compounds composed of amino acids. The amino acids that synthesize proteins are linked together in long chains. Due to the vast differences in chain length and arrangement, proteins of all shapes and sizes are formed. All proteins are polymers formed by the linkage of 20 different amino acids.

[0024] It's important to note that when amino acids link together to form proteins, the chemical bond connecting these amino acids is called a peptide bond. (Actually, a peptide bond is a chemical bond formed when two amino acids are linked, where the carboxyl group (-COOH) of one amino acid connects to the amino group (-NH₂) of another amino acid (-NH₂) (-NH₃CO₃). Multiple peptide bonds form a peptide chain, and a protein molecule consists of one or more peptide chains.) When the amino acids that make up a peptide bond combine, some of their groups participate in the formation of the peptide bond and lose a water molecule. Therefore, the amino acid unit that forms the peptide bond is called an amino acid residue, which is the part remaining after the amino acid linked by the peptide bond loses water.

[0025] In existing technologies, when neural networks analyze amino acids in proteins, they are often limited by the small number of amino acid types and the neglect of non-local interactions between amino acid residues, which often results in poor recognition performance.

[0026] To address the aforementioned issues, this application first obtains multiple protein molecules, uses the amino acid residues in these protein molecules as nodes, and generates amino acid residue graph structure data by combining the interactions between each amino acid residue and its neighboring amino acid residues. This amino acid residue graph structure data includes not only the amino acid residue nodes in the protein but also the spatial structural relationships between each amino acid residue node.

[0027] In some embodiments, the step of acquiring protein sample information and decomposing the protein sample information into amino acid residue map structural data includes: (1) Generate multiple amino acid residue nodes based on the amino acid residues in the protein sample information as nodes; (2) Connect the amino acid residue nodes with related relationships to obtain amino acid residue graph structure data.

[0028] In this process, multiple amino acid residue nodes can be generated based on the amino acid residues in the protein sample information. Since there is an interaction between amino acid residues that are close to each other, ignoring the interaction between amino acid residue nodes will lead to a decrease in subsequent recognition efficiency.

[0029] Therefore, the embodiments of this application can obtain the correlation between amino acid residue nodes, determine the amino acid residue nodes with a correlation greater than a preset threshold as amino acid residue nodes with a correlation relationship, and connect the amino acid residue nodes with a correlation relationship by edges. That is, it can also be understood as representing the peptide chains between amino acid residue nodes on the same peptide bond by connecting them by edges, so as to obtain amino acid residue graph structure data that includes not only amino acid residue nodes in the protein, but also the spatial structural relationship between each amino acid residue node. The node feature of the amino acid residue node can be an amino acid category vector, which is a one-hot encoded 20-dimensional vector information, and the edge feature of the connection is the Euclidean distance between the two.

[0030] In some embodiments, the step of connecting related amino acid residue nodes with edges to obtain amino acid residue graph structure data may include: (1) Calculate the spatial distance information between each amino acid residue node; (2) Connect the amino acid residue nodes whose spatial distance information is less than a preset threshold by edge connection.

[0031] Among them, the Euclidean distance between each amino acid side chain heavy atom (CB atom) can be calculated, and the Euclidean distance between amino acid side chain heavy atoms can be used as the spatial distance information between amino acid residue nodes.

[0032] Furthermore, the preset threshold is a critical value for determining whether two substances belong to the same peptide bond, typically 12 angstroms (Å), where Å is a length symbol, and 1 Å is 10 to the power of -10 meters. Therefore, amino acid residue nodes with spatial distance information less than this preset threshold are identified as amino acid residue nodes that interact with each other, and edge connections are established between these amino acid residue nodes with spatial distance information less than the preset threshold.

[0033] In step 102, the amino acid residue diagram structure data is input into the graph neural network, which outputs multiple amino acid residue node vectors.

[0034] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0035] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0036] The solutions provided in this application involve technologies such as machine learning in artificial intelligence, and are specifically illustrated through the following embodiments: In this context, since the local and non-local structural distribution and physicochemical properties of each amino acid node in the graph structure data are affected by the amino acid residues on neighboring nodes, and the closer the relationship between amino acid nodes, the greater the influence, the graph neural network (GCN) can use the features of each amino acid node in the graph structure data and the spatial positional relationship between the corresponding amino acid nodes to perform convolution calculations and output the vector of each amino acid residue node. This amino acid residue node vector not only contains the feature information of the amino acid residue, but also the feature information of the influence of other amino acid residues with edge connections on it. That is, it combines the non-local interaction between amino acid residues, and the feature expression effect is better.

[0037] In step 103, multiple amino acid microenvironment samples are constructed based on the correlation between each amino acid residue in the protein sample information.

[0038] Since a small number of amino acid types can affect the accuracy of identification, this application embodiment can obtain the correlation relationship between each amino acid residue in the protein sample information. This correlation relationship can be an interaction between them. The greater the interaction, the greater the correlation, and the smaller the interaction, the smaller the correlation. In order to enrich the types of amino acids, each amino acid can be taken as the center, and two amino acid residues with a correlation greater than a certain threshold can be taken in the range before and after it to form an amino acid microenvironment sample corresponding to the amino acid at the center, thus expanding the identification dimension of amino acid residues.

[0039] In some implementations, the step of constructing multiple amino acid microenvironment samples based on the correlation between each amino acid residue in the protein sample information may include: (1) Obtain the correlation between each amino acid residue in the protein sample information; (2) Taking each amino acid residue as the center point, select a preset number of amino acid residues with a correlation greater than a preset threshold with respect to the amino acid residues as the center point in the direction relative to the amino acid residues as the center point to construct multiple amino acid microenvironment samples.

[0040] This involves obtaining the correlation degree between each amino acid residue in the protein sample information. This correlation degree represents the degree of interaction between them. The higher the correlation degree, the greater the degree of interaction between them and the closer their Euclidean distance. The lower the correlation degree, the less the degree of interaction between them and the greater their Euclidean distance. The preset threshold can be used to define whether two things can be considered as neighbor nodes. For example, the correlation degree can range from 0 to 1, the preset threshold can be 0.1, and the preset number can be any number such as 1, 2, or 3.

[0041] Based on this, each amino acid residue is used as the center point in sequence. In the direction relative to the amino acid residues used as the center point, which can be the front and back direction, a preset number of amino acid residues with a correlation greater than a preset threshold with the amino acid residues used as the center point are selected to construct multiple amino acid microenvironment samples, thereby expanding the recognition dimension of amino acid residues.

[0042] In step 104, the amino acid microenvironment samples are clustered to obtain multiple cluster types of amino acid microenvironment sample sets.

[0043] Since the expanded amino acid microenvironments are numerous and diverse, while there are only 20 actual amino acid types, in order to construct an appropriately sized number of types (amino acid type library), all amino acid microenvironment samples can be clustered according to similarity. All amino acid microenvironments can be classified into multiple cluster types of amino acid microenvironment sample sets. Each cluster type can be understood as a label, reflecting the local and non-local structural distribution and physicochemical properties of the corresponding amino acid in the protein.

[0044] In some embodiments, the step of clustering the amino acid microenvironment samples to obtain multiple cluster types of amino acid microenvironment sample sets may include: (1) Alignment operation is performed between each amino acid microenvironment sample to obtain aligned amino acid microenvironment samples; (2) Calculate the similarity between the aligned amino acid microenvironment samples to obtain the similarity matrix between the aligned amino acid microenvironment samples; (3) Cluster the aligned amino acid environment samples according to the similarity matrix to obtain a set of amino acid microenvironment samples with multiple cluster types.

[0045] Since the display dimensions of the amino acid microenvironment samples are different, in order to facilitate subsequent comparison, we can first perform an alignment operation between each amino acid microenvironment sample to obtain aligned amino acid microenvironment samples.

[0046] Furthermore, similarity scores can be extracted between the amino acids at the center of the aligned amino acid microenvironment samples to obtain the similarity between the aligned amino acid microenvironment samples, and thus obtain the similarity matrix between the aligned amino acid microenvironment samples.

[0047] Based on the similarity matrix, the aligned amino acid microenvironment samples are clustered and segmented with a certain inter-class spacing to obtain a set of amino acid microenvironment samples with multiple cluster types.

[0048] In step 105, multiple cluster types are used as label information and multiple amino acid residue node vectors are input into a preset classification model for training to obtain the trained preset classification model.

[0049] Specifically, the cluster type of each amino acid residue node vector can be used as the label information for that amino acid residue. The amino acid residue node vectors with label information are input into the preset classification model for training. The preset classification model can output the predicted type based on the amino acid residue node vector. The difference value is compared with the predicted type and the label information. The loss function is adjusted based on the difference value. The process is repeated iteratively to continuously optimize the model until the difference converges, and the trained preset classification model is obtained.

[0050] In step 106, the amino acid residues to be detected are classified according to the trained preset classification model.

[0051] The trained preset classification model can quickly classify the amino acid residues to be detected. Specifically, the node vector of the amino acid residue to be detected can be extracted by the graph neural network. The node vector of the amino acid residue to be detected is input into the preset classification model, which can identify the cluster category corresponding to the amino acid residue to be detected. Based on the cluster category, the local and non-local structural distribution and physicochemical properties of the amino acid residue to be detected in the protein can be quickly determined.

[0052] As described above, this application's embodiments acquire protein sample information and decompose it into amino acid residue graph structure data; input the amino acid residue graph structure data into a graph neural network to output multiple amino acid residue node vectors; construct multiple amino acid microenvironment samples based on the correlation between each amino acid residue in the protein sample information; cluster the amino acid microenvironment samples to obtain multiple cluster types of amino acid microenvironment sample sets; input the multiple cluster types as label information and multiple amino acid residue node vectors into a preset classification model for training to obtain a trained preset classification model; and classify the amino acid residues to be detected according to the preset classification model. Thus, by representing amino acid residues with vectors and based on self-supervised learning, a model for identifying amino acids is obtained. Compared to existing schemes that manually label protein information, this application can reasonably and effectively perform self-supervised learning, using a large amount of unlabeled protein sample information for model training, greatly improving the efficiency of information processing.

[0053] Based on the methods described in the above embodiments, the following examples will provide further detailed explanations.

[0054] In this embodiment, the information processing device will be specifically integrated into a server as an example for explanation. Please refer to the following description for details.

[0055] Please see Figure 3 , Figure 3 Another schematic flowchart illustrating the information processing method provided in this application embodiment. The method flow may include: In step 201, the server obtains protein sample information.

[0056] For a better illustration of the embodiments of this application, please refer to Figure 4a As shown, Figure 4a This is a schematic diagram of a scenario for the information processing method provided in the embodiments of this application. The server can obtain protein sample information (i.e., obtain protein data 11). The protein sample information can use the refined set (4852 proteins) and general set (21382 proteins) of the Protein Data Bank (PDB) and the Protein Data Bank (146836 proteins) as unlabeled training datasets of different sizes.

[0057] In step 202, the server generates multiple amino acid residue nodes based on the amino acid residues in the protein sample information, calculates the spatial distance information between each amino acid residue node, and connects the amino acid residue nodes whose spatial distance information is less than a preset threshold with edges to obtain amino acid residue graph structure data.

[0058] Please continue reading for more details. Figure 4a As shown, the server can generate multiple amino acid residue nodes 121 based on the amino acid residues in the protein data 11 as nodes, calculate the spatial distance information between each amino acid residue node 121, which can be Euclidean distance. The preset threshold is the critical value for determining whether the two belong to the same peptide bond, which is generally 12 angstroms. The amino acid residue nodes with Euclidean distance less than 12 angstroms are connected by edges to generate connection lines 122, thus obtaining the amino acid residue graph structure data 12.

[0059] In step 203, the server inputs the amino acid residue graph structure data into the graph neural network and outputs multiple amino acid residue node vectors.

[0060] Please continue reading for more details. Figure 4a As shown, the server inputs the amino acid residue graph structure 12 into the graph neural network 13. The graph neural network can use the features of each amino acid node in the graph structure data and the spatial position relationship between the corresponding amino acid nodes to perform convolution calculation and output the vector 14 of each amino acid residue node. The amino acid residue node vector not only contains the feature information of the amino acid residue, but also the feature information of the influence of other amino acid residues with edge connection relationship on it. That is, it combines the non-local interaction between amino acid residues, and the feature expression effect is better.

[0061] In step 204, the server obtains the correlation between each amino acid residue in the protein sample information.

[0062] The server can obtain the correlation between each amino acid residue in the protein sample information. The correlation is the degree of interaction between them. The greater the correlation, the greater the degree of interaction and the closer the Euclidean distance between them. The smaller the correlation, the smaller the degree of interaction and the greater the Euclidean distance between them.

[0063] In step 205, the server sequentially uses each amino acid residue as a center point, and in the direction relative to the amino acid residues used as the center point, selects two amino acid residues with a correlation greater than a preset threshold to construct multiple amino acid microenvironment samples.

[0064] Please refer to the following: Figure 4b As shown, Figure 4b This is another scenario illustration of the information processing method provided in the embodiments of this application. In order to expand the recognition dimension of amino acid residues and avoid only 20 amino acid recognition types, the server can sequentially use each amino acid residue in the protein data 11 as the center point, and select two amino acid residues with a correlation greater than a preset threshold with respect to the amino acid residues used as the center point in the direction relative to the amino acid residues used as the center point to construct multiple amino acid microenvironment samples 16.

[0065] In step 206, the server performs an alignment operation between each amino acid microenvironment sample to obtain aligned amino acid microenvironment samples, calculates the similarity between the aligned amino acid microenvironment samples, and obtains a similarity matrix between the aligned amino acid microenvironment samples.

[0066] Please refer to the following: Figure 4c As shown, Figure 4c This is another scenario illustration of the information processing method provided in the embodiments of this application. Multiple amino acid microenvironment samples B (i.e., the microenvironment of amino acids) can be extracted from protein dataset A. It can be seen that since the display dimensions of the amino acid microenvironment samples B are different, in order to facilitate subsequent comparison, an alignment operation can be performed between each amino acid microenvironment sample to obtain aligned amino acid microenvironment samples.

[0067] Furthermore, a set of unit vectors determined by continuous CA atoms can be extracted from the central amino acids in the aligned amino acid microenvironment samples. The similarity score between each pair of amino acid microenvironment samples, ranging from 0 to 10, can be calculated using the root mean square deviation of the unit vector (URMSD). This yields the similarity matrix between the aligned amino acid microenvironment samples.

[0068] In step 207, the server clusters the aligned amino acid environment samples according to the similarity matrix to obtain a set of amino acid microenvironment samples with multiple cluster types.

[0069] Please continue reading for more details. Figure 4c As shown, the server clusters amino acid microenvironment samples B according to the similarity matrix, and performs clustering segmentation of amino acid microenvironment samples with a certain inter-class spacing to obtain a set of amino acid microenvironment samples with multiple cluster types (i.e., the vocabulary C).

[0070] In step 208, the server sequentially selects the target cluster type from multiple cluster types based on the amino acid type corresponding to each amino acid residue node vector for labeling. The labeled amino acid residue node vectors are then input into a preset classification model for training to obtain the trained preset classification model. The amino acid residues to be detected are then classified according to the trained preset classification model.

[0071] The server can determine which cluster type each amino acid residue node vector belongs to, and then label the amino acid residue node vector according to the cluster type it belongs to. The labeled amino acid residue node vector is then input into a preset classification model for training. The preset classification model can output a predicted type based on the amino acid residue node vector. The difference value is compared with the predicted type and the label information. The loss function is adjusted according to the difference value. The process is repeated iteratively to continuously optimize the model until the difference converges, and the trained preset classification model is obtained.

[0072] Furthermore, the server can extract the node vector of the amino acid residue to be detected through a graph neural network. By inputting the node vector of the amino acid residue to be detected into a preset classification model, the cluster category corresponding to the amino acid residue to be detected can be identified. Based on the cluster category, the local and non-local structural distribution and physicochemical properties of the amino acid residue to be detected in the protein can be quickly determined, which facilitates research by researchers in the field and greatly improves the efficiency of information processing.

[0073] In step 209, the server sums the amino acid residue node vectors belonging to the same protein to obtain the protein representation vector. The amino acid residue node vector and the corresponding protein vector are then input into a preset binary classification model for training to obtain the trained preset binary classification model. The preset binary classification model is then used to determine whether the amino acid residue to be detected belongs to the protein to be detected.

[0074] Please continue reading for more details. Figure 4aAs shown. Since a protein molecule is composed of one or more peptide chains, the server can first sum the amino acid residue node vectors 14 on the same peptide chain to obtain the peptide chain-level representation vector. Then, by summing the peptide chain-level representation vectors belonging to the same protein (or summing all amino acid residue node vectors of the same protein), the protein representation vector is obtained.

[0075] The embodiments of this application can be used to learn whether an amino acid comes from a given protein. Specifically, the amino acid residue node vector and its corresponding protein vector can be input into a preset binary classification model for training. The preset binary classification model can output a predicted type based on the amino acid residue node vector and its corresponding protein vector. The difference value is compared with the actual relationship (i.e., labels 0 and 1, where 0 represents not belonging and 1 represents belonging). The loss function is adjusted based on the difference value. The iterative process is repeated to continuously optimize the model until the difference converges, and the trained preset binary classification model is obtained.

[0076] Furthermore, the server can extract the target amino acid residue node vector through a graph neural network. By inputting the target amino acid residue node vector and the target protein vector into a preset binary classification model, it can quickly determine whether the target amino acid belongs to the target protein, which facilitates research by researchers in the field and greatly improves the efficiency of information processing.

[0077] As described above, this application's embodiments acquire protein sample information and decompose it into amino acid residue graph structure data; input the amino acid residue graph structure data into a graph neural network to output multiple amino acid residue node vectors; construct multiple amino acid microenvironment samples based on the correlation between each amino acid residue in the protein sample information; cluster the amino acid microenvironment samples to obtain multiple cluster types of amino acid microenvironment sample sets; input the multiple cluster types as label information and multiple amino acid residue node vectors into a preset classification model for training to obtain a trained preset classification model; and classify the amino acid residues to be detected according to the preset classification model. Thus, by representing amino acid residues with vectors and based on self-supervised learning, a model for identifying amino acids is obtained. Compared to existing schemes that manually label protein information, this application can reasonably and effectively perform self-supervised learning, using a large amount of unlabeled protein sample information for model training, greatly improving the efficiency of information processing.

[0078] Furthermore, embodiments of this application can also quickly identify the affiliation between the amino acid to be detected and the protein to be detected through a pre-trained binary classification model, further improving the efficiency of information processing.

[0079] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an information processing device provided in an embodiment of this application. The information processing device may include an acquisition unit 301, a first input unit 302, a construction unit 303, a clustering unit 304, a second input unit 305, and a classification unit 306, etc.

[0080] The acquisition unit 301 is used to acquire protein sample information and decompose the protein sample information into amino acid residue diagram structure data.

[0081] In some embodiments, the acquisition unit 301 includes: The acquisition subunit is used to obtain protein sample information; A subunit is generated to generate multiple amino acid residue nodes based on the amino acid residues in the protein sample information. The connecting subunit is used to connect related amino acid residue nodes with edges to obtain amino acid residue graph structure data.

[0082] In some embodiments, the connection subunit is used for: Calculate the spatial distance information between each amino acid residue node; By connecting the amino acid residue nodes whose spatial distance information is less than a preset threshold, the amino acid residue graph structure data is obtained.

[0083] The first input unit 302 is used to input the amino acid residue diagram structure data into a graph neural network and output multiple amino acid residue node vectors.

[0084] Construction unit 303 is used to construct multiple amino acid microenvironment samples based on the correlation between each amino acid residue in the protein sample information.

[0085] In some embodiments, the building unit 303 includes: The acquisition subunit is used to obtain the correlation between each amino acid residue in the protein sample information; A subunit is constructed to sequentially select a predetermined number of amino acid microenvironment samples in the direction relative to each amino acid residue as the center point, and the amino acid residue as the center point has a correlation degree greater than a predetermined threshold with the amino acid residue as the center point.

[0086] In some implementations, the building subunit is used for: Using each amino acid residue as a center point, two amino acid residues with a correlation greater than a preset threshold to the amino acid residues used as the center point are selected in the direction opposite to the amino acid residues used as the center point to construct multiple amino acid microenvironment samples.

[0087] Clustering unit 304 is used to cluster the amino acid microenvironment samples to obtain a set of amino acid microenvironment samples with multiple cluster types.

[0088] In some implementations, the clustering unit 304 is used for: Alignment operations were performed between each amino acid microenvironment sample to obtain aligned amino acid microenvironment samples. The similarity between aligned amino acid microenvironment samples is calculated to obtain the similarity matrix between aligned amino acid microenvironment samples; Based on the similarity matrix, the aligned amino acid environment samples are clustered to obtain multiple cluster types of amino acid microenvironment sample sets.

[0089] The second input unit 305 is used to input the multiple clustering types as label information and the multiple amino acid residue node vectors into a preset classification model for training, so as to obtain the trained preset classification model.

[0090] In some implementations, the second input unit 304 is used for: The target cluster type is selected from the multiple cluster types for labeling based on the amino acid type corresponding to each amino acid residue node vector; The labeled amino acid residue node vectors are input into a preset classification model for training, resulting in a trained preset classification model.

[0091] The classification unit 306 is used to classify the amino acid residues to be detected according to the trained preset classification model.

[0092] In some embodiments, the apparatus further includes: The summation unit is used to sum the amino acid residue node vectors belonging to the same protein to obtain the protein representation vector; The third input unit is used to input the amino acid residue node vector and the corresponding protein vector into the preset binary classification model for training, and obtain the trained preset binary classification model. The determination unit is used to determine whether the amino acid residues to be detected belong to the protein to be detected according to the preset binary classification model.

[0093] The specific implementation of each of the above units can be found in the previous embodiments, and will not be repeated here.

[0094] As described above, in this embodiment, the acquisition unit 301 acquires protein sample information and decomposes it into amino acid residue graph structure data; the first input unit 302 inputs the amino acid residue graph structure data into a graph neural network and outputs multiple amino acid residue node vectors; the construction unit 303 constructs multiple amino acid microenvironment samples based on the correlation between each amino acid residue in the protein sample information; the clustering unit 304 clusters the amino acid microenvironment samples to obtain multiple cluster types of amino acid microenvironment sample sets; the second input unit 305 uses multiple cluster types as label information and multiple amino acid residue node vectors as input to a preset classification model for training to obtain a trained preset classification model; and the classification unit 306 classifies the amino acid residues to be detected according to the preset classification model. Thus, by representing amino acid residues with vectors and based on self-supervised learning, a model for identifying amino acids is obtained. Compared to existing schemes that manually label protein information, this application can reasonably and effectively perform self-supervised learning, using a large amount of unlabeled protein sample information for model training, greatly improving the efficiency of information processing.

[0095] This application also provides a computer device, such as... Figure 6 As shown, it illustrates a schematic diagram of the server structure involved in an embodiment of this application. Specifically: The computer device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 6 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: Processor 401 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines. By running or executing software programs and / or modules stored in memory 402, and by calling data stored in memory 402, it performs various functions of the computer device and processes data, thereby performing overall detection of the computer device. Optionally, processor 401 may include one or more processing cores; optionally, processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into processor 401.

[0096] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and information processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the server, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0097] The computer equipment also includes a power supply 403 that supplies power to the various components. Optionally, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0098] The computer device may also include an input unit 404, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0099] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the computer device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402, thereby implementing the various method steps provided in the foregoing embodiments, as follows: Protein sample information is acquired and decomposed into amino acid residue graph structure data. This amino acid residue graph structure data is then input into a graph neural network, which outputs multiple amino acid residue node vectors. Multiple amino acid microenvironment samples are constructed based on the correlation between each amino acid residue in the protein sample information. These amino acid microenvironment samples are clustered to obtain multiple cluster types of amino acid microenvironment sample sets. The multiple cluster types are used as label information, along with the multiple amino acid residue node vectors, and input into a preset classification model for training, resulting in a trained preset classification model. Finally, the amino acid residues to be detected are classified according to the trained preset classification model.

[0100] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed description of the information processing method above, which will not be repeated here.

[0101] As described above, the computer device in this embodiment can acquire protein sample information and decompose it into amino acid residue graph structure data; input the amino acid residue graph structure data into a graph neural network to output multiple amino acid residue node vectors; construct multiple amino acid microenvironment samples based on the correlation between each amino acid residue in the protein sample information; cluster the amino acid microenvironment samples to obtain multiple cluster types of amino acid microenvironment sample sets; input the multiple cluster types as label information and multiple amino acid residue node vectors into a preset classification model for training to obtain a trained preset classification model; and classify the amino acid residues to be detected according to the preset classification model. Thus, by representing amino acid residues with vectors and based on self-supervised learning, a model for identifying amino acids is obtained. Compared with existing schemes that manually label protein information, this application can reasonably and effectively perform self-supervised learning, using a large amount of unlabeled protein sample information for model training, greatly improving the efficiency of information processing.

[0102] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0103] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the information processing methods provided in embodiments of this application. For example, the instructions can execute the following steps: Protein sample information is acquired and decomposed into amino acid residue graph structure data. This amino acid residue graph structure data is then input into a graph neural network, which outputs multiple amino acid residue node vectors. Multiple amino acid microenvironment samples are constructed based on the correlation between each amino acid residue in the protein sample information. These amino acid microenvironment samples are clustered to obtain multiple cluster types of amino acid microenvironment sample sets. The multiple cluster types are used as label information, along with the multiple amino acid residue node vectors, and input into a preset classification model for training, resulting in a trained preset classification model. Finally, the amino acid residues to be detected are classified according to the trained preset classification model.

[0104] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the above embodiments.

[0105] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0106] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0107] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the information processing methods provided in the embodiments of this application, the beneficial effects that any of the information processing methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0108] The above provides a detailed description of an information processing method, apparatus, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. An information processing method, characterized in that, include: Obtain protein sample information and decompose the protein sample information into amino acid residue diagram structure data; The amino acid residue diagram structure data is input into a graph neural network, which outputs multiple amino acid residue node vectors. Multiple amino acid microenvironment samples are constructed based on the correlation between each amino acid residue in the protein sample information; the correlation relationship is used to characterize the magnitude of the correlation, the correlation degree is used to characterize the degree of interaction between amino acid residues, and the correlation is negatively correlated with the spatial distance between amino acid residues. The amino acid microenvironment samples were clustered to obtain multiple cluster types of amino acid microenvironment sample sets. For each amino acid residue node vector, the clustering type of the corresponding amino acid residue is determined, and the label is assigned to the amino acid residue node vector based on the clustering type to obtain label information; The node vectors of each amino acid residue and the corresponding tag information are input into a preset classification model for training to obtain the trained preset classification model. The amino acid residues to be detected are classified according to the trained preset classification model; The step of decomposing the protein sample information into amino acid residue map structure data includes: Based on the amino acid residues in the protein sample information as nodes, multiple amino acid residue nodes are generated, and amino acid residue nodes with a correlation greater than a preset threshold are identified as amino acid residue nodes with a correlation relationship. By connecting the related amino acid residue nodes with edges, the amino acid residue graph structure data is obtained; the amino acid residue graph structure data includes the amino acid residue nodes and the spatial structural relationships between each amino acid residue node.

2. The information processing method according to claim 1, characterized in that, The step of connecting related amino acid residue nodes with edges includes: Calculate the spatial distance information between each amino acid residue node; Edge connections are made between amino acid residue nodes whose spatial distance information is less than a preset threshold.

3. The information processing method according to claim 1, characterized in that, The step of constructing multiple amino acid microenvironment samples based on the correlation between each amino acid residue in the protein sample information includes: To obtain the correlation between each amino acid residue in the protein sample information; Using each amino acid residue as a center point, within a predetermined range in the direction before and after the amino acid residue as the center point, a predetermined number of amino acid residues with a correlation greater than a predetermined threshold with the amino acid residue as the center point are selected to construct multiple amino acid microenvironment samples.

4. The information processing method according to claim 3, characterized in that, The step of selecting a predetermined number of amino acid residues with a correlation greater than a predetermined threshold to the amino acid residues used as the center point to construct multiple amino acid microenvironment samples includes: Multiple amino acid microenvironment samples were constructed by selecting two amino acid residues with a correlation greater than a preset threshold with the amino acid residues used as the center point.

5. The information processing method according to claim 1, characterized in that, The step of clustering the amino acid microenvironment samples to obtain multiple cluster types of amino acid microenvironment sample sets includes: Alignment operations were performed between each amino acid microenvironment sample to obtain aligned amino acid microenvironment samples. The similarity between aligned amino acid microenvironment samples is calculated to obtain the similarity matrix between aligned amino acid microenvironment samples; Based on the similarity matrix, the aligned amino acid environment samples are clustered to obtain multiple cluster types of amino acid microenvironment sample sets.

6. The information processing method according to any one of claims 1 to 5, characterized in that, The method further includes: The protein representation vector is obtained by summing the amino acid residue node vectors belonging to the same protein. The amino acid residue node vector and the corresponding protein representation vector are input into a preset binary classification model for training, and the trained preset binary classification model is obtained. The preset binary classification model is used to determine whether the amino acid residues to be detected belong to the protein to be detected.

7. An information processing device, characterized in that, include: An acquisition unit is used to acquire protein sample information and decompose the protein sample information into amino acid residue diagram structure data. The first input unit is used to input the amino acid residue diagram structure data into the graph neural network and output multiple amino acid residue node vectors. A construction unit is used to construct multiple amino acid microenvironment samples based on the correlation relationship between each amino acid residue in the protein sample information; the correlation relationship is used to characterize the magnitude of the correlation, the correlation degree is used to characterize the degree of interaction between amino acid residues, and the correlation degree is negatively correlated with the spatial distance between amino acid residues. Clustering unit, used to cluster the amino acid microenvironment samples to obtain multiple cluster types of amino acid microenvironment sample sets; The second input unit is used to determine the clustering type of the corresponding amino acid residues for each amino acid residue node vector, and to label the amino acid residue node vector based on the clustering type to obtain label information; and to input each amino acid residue node vector and the corresponding label information into a preset classification model for training to obtain a trained preset classification model. A classification unit is used to classify the amino acid residues to be detected according to the trained preset classification model; The acquisition unit includes: The acquisition subunit is used to obtain protein sample information; A subunit is generated to generate multiple amino acid residue nodes based on the amino acid residues in the protein sample information, and to identify amino acid residue nodes with a correlation greater than a preset threshold as amino acid residue nodes with a correlation relationship. The connecting subunit is used to connect amino acid residue nodes with related relationships to obtain amino acid residue graph structure data; the amino acid residue graph structure data includes amino acid residue nodes and the spatial structural relationships between each amino acid residue node.

8. The information processing apparatus according to claim 7, characterized in that, The connection subunit is used for: Calculate the spatial distance information between each amino acid residue node; By connecting the amino acid residue nodes whose spatial distance information is less than a preset threshold, the amino acid residue graph structure data is obtained.

9. The information processing apparatus according to claim 7, characterized in that, The building unit includes: The acquisition subunit is used to obtain the correlation between each amino acid residue in the protein sample information; The subunit is constructed by sequentially selecting a predetermined number of amino acid microenvironment samples with a correlation greater than a predetermined threshold with each amino acid residue as the center point within a predetermined range in the direction before and after the amino acid residue as the center point.

10. The information processing apparatus according to claim 7, characterized in that, The clustering unit is used for: Alignment operations were performed between each amino acid microenvironment sample to obtain aligned amino acid microenvironment samples. The similarity between aligned amino acid microenvironment samples is calculated to obtain the similarity matrix between aligned amino acid microenvironment samples; Based on the similarity matrix, the aligned amino acid environment samples are clustered to obtain multiple cluster types of amino acid microenvironment sample sets.

11. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the information processing method according to any one of claims 1 to 6.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the information processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Protein binding site prediction method, device and equipment and storage medium

    CN107563150A

  • Protein activity prediction device based on multi-view classification model

    CN110993037A