Method, apparatus, device and medium for extracting protein attribute features
By extracting and representing protein attribute features through centroid coordinates and spatial vectors, the method enhances the efficiency and accuracy of machine learning models in drug molecule screening, addressing the inefficiencies of current models.
Patent Information
- Application Number
- CN202310843414.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-07-10
AI Technical Summary
The existing machine learning algorithms take a long time in the process of drug molecule screening, and the existing feature extraction methods fail to make full use of protein attribute features, resulting in insufficient prediction accuracy.
By extracting small molecule ligands, protein pockets and protein molecular data from compound data, calculating center of gravity coordinates and spatial vectors, multi-class protein attribute features are generated, and used as input vectors for machine learning models to improve feature description capabilities.
It effectively improves the prediction accuracy of machine learning models in the molecular screening process, shortens screening time, and improves screening efficiency.
Smart Images

Figure CN117059158B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics and data processing, and particularly relates to a method, device, equipment and medium for extracting protein attribute features. Background Art
[0002] In the process of drug research and development, drug molecule screening is the most difficult and time-consuming link. A traditional drug molecule screening process usually refers to the process of selecting a small molecule ligand in a molecular database, and performing calculations, scoring, and evaluation on the compound obtained after docking it with the target (protein) pocket. In this process, calculations need to be strictly performed according to physical formulas from various interactions between the ligand and the receptor (target pocket), such as calculations and scoring from aspects such as geometric shape, electrode polarity, and hydrogen bonds. Therefore, it takes a long time to calculate the compound obtained after docking a small molecule ligand with the target. In order to find a suitable small molecule ligand to obtain a qualified compound, a large number of the above screening and calculation processes need to be carried out in a huge molecular database, so this process is very time-consuming.
[0003] With the wide application of machine learning algorithms, in order to solve the problem of time-consuming in the molecular screening process, a machine learning model is adopted using existing docking and evaluation experience. Through the machine learning model, a rough search and evaluation are first carried out from the molecular database, and then the search range can be effectively narrowed. For machine learning algorithms, a very important step is to obtain the input embedding (a distributed representation method, that is, the original input data is distributedly represented as a combination of a series of multi-dimensional features, which is called the input vector in the following description). The input vector consists of multiple or multiple groups of features. Features with appropriate quantity and comprehensive physical meaning can enable the machine learning model to obtain better learning effects. Currently, when using machine learning algorithms for drug molecule screening, usually the molecular vector converted from the entire candidate compound molecule is used as the input vector; or as in the Chinese patent invention with the publication number CN 114283899 A and the name "A Method for Training a Molecular Binding Model, a Molecular Screening Method and Device", the adjacency matrix of the molecule is used as the feature of the input vector for the model; or as in the "Drug Molecule Screening Method and System" with the publication number CN114530210 A, the molecular orientation information descriptor is used as the input vector of the model. The molecular orientation information descriptor recorded therein is a plurality of one-dimensional vectors representing the molecular bond characteristics of the drug molecule.
[0004] In summary, in order to apply machine learning algorithms in the molecular screening process, the industry has made various attempts and efforts, using a variety of features that characterize the properties of molecules or proteins as training data for machine learning models to train the machine learning models. Due to the complexity of drug molecules, based on the existing technology, there is still a large room for improving the role of machine learning models in the molecular screening process. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method, apparatus, device and medium for extracting protein property features, so as to improve the prediction accuracy of machine learning models in the molecular screening process.
[0006] According to one aspect of the present invention, the present invention provides a method for extracting protein property features, including the following steps:
[0007] Extract a plurality of candidate compound data from a compound data file.
[0008] Separate small molecule ligand data, protein pocket data and protein molecule data from each candidate compound data respectively, to obtain a plurality of small molecule ligand data, a plurality of protein pocket data and a plurality of protein molecule data.
[0009] Calculate the centroid coordinates of each type of atom in each small molecule ligand data and the centroid coordinates of each type of atom in each protein molecule data respectively.
[0010] Convert each small molecule ligand data, each protein pocket data and each protein molecule data into corresponding molecular vectors respectively, to obtain a plurality of small molecule ligand molecular vectors, a plurality of protein pocket molecular vectors and a plurality of protein molecule vectors.
[0011] Align the sequences of a plurality of protein molecule data respectively to obtain a plurality of protein molecule sequences.
[0012] Extract the central carbon coordinates of each amino acid from a plurality of protein molecule sequences respectively, and form a protein amino acid central carbon coordinate sequence.
[0013] Based on the atomic coordinates in each small molecule ligand data and the atomic coordinates in each protein pocket data in each candidate compound data, generate spatial vectors that characterize the spatial structures of small molecule ligands and protein pockets respectively.
[0014] The first set composed of the centroid coordinates of each type of atom in each small molecule ligand, the second set composed of the centroid coordinates of each type of atom in each protein molecule, the third set composed of each small molecule ligand molecular vector, the fourth set composed of each protein pocket molecular vector, the fifth set composed of each protein molecular vector, the sixth set composed of each spatial vector representing the spatial structures of the small molecule ligand and the protein pocket, and the seventh set composed of the protein amino acid central carbon coordinate sequence are used as different types of protein attribute features, where the protein attribute features are used to form the input vector of a machine learning model, and the machine learning model is used to predict the score of a preset scoring item in the molecular screening process.
[0015] Among them, the steps of generating the spatial vector representing the spatial structures of the small molecule ligand and the protein pocket include:
[0016] Generate a small molecule ligand atom matrix based on the atoms and atom coordinates in the small molecule ligand data.
[0017] Taking each small molecule ligand atom in the small molecule ligand atom matrix as a calculation object, calculate the spatial distance between the calculation object and each protein pocket atom respectively.
[0018] Take the protein pocket atom with the minimum distance as the target protein pocket atom corresponding to the calculation object, and obtain the target protein pocket atom coordinates.
[0019] Generate a protein pocket atom matrix based on the target protein pocket atom and the target protein pocket atom coordinates.
[0020] Obtain a first coordinate vector based on the small molecule ligand atom matrix, and obtain a second coordinate vector based on the protein pocket atom matrix. Each coordinate vector includes a coordinate x vector, a coordinate y vector, and a coordinate z vector respectively.
[0021] Taking any two vectors in the first coordinate vector as independent variables, and taking the corresponding two vectors in the second coordinate vector as reference quantities, calculate a coordinate value of the spatial vector respectively based on a preset transformation kernel function.
[0022] According to another aspect of the present invention, the present invention further provides an apparatus for extracting protein attribute features, including a data extraction module, a data separation module, a centroid coordinate calculation module, a vector conversion module, a coordinate sequence extraction module, and a spatial vector generation module. Among them, the data extraction module is used to extract multiple candidate compound data from a compound data file; the data separation module is used to separately separate small molecule ligand data, protein pocket data, and protein molecule data from each candidate compound data to obtain multiple small molecule ligand data, multiple protein pocket data, and multiple protein molecule data; the centroid coordinate calculation module is used to calculate the centroid coordinates of each type of atom in each small molecule ligand data and the centroid coordinates of each type of atom in each protein molecule data respectively; the vector conversion module is used to convert each small molecule ligand data, each protein pocket data, and each protein molecule data into corresponding molecular vectors respectively to obtain multiple small molecule ligand molecular vectors, multiple protein pocket molecular vectors, and multiple protein molecule vectors; the coordinate sequence extraction module is used to perform sequence alignment on the multiple protein molecule data of multiple candidate compounds; extract the central carbon coordinates of each amino acid from the aligned protein molecule sequences and form a protein amino acid central carbon coordinate sequence; the spatial vector generation module is used to generate spatial vectors representing the spatial structures of small molecule ligands and protein pockets respectively based on the atomic coordinates in each small molecule ligand data and each protein pocket data in each candidate compound data; among them, the first set composed of the centroid coordinates of each type of atom in each small molecule ligand, the second set composed of the centroid coordinates of each type of atom in each protein molecule, the third set composed of each small molecule ligand molecular vector, the fourth set composed of each protein pocket molecular vector, the fifth set composed of each protein molecule vector, the sixth set composed of each spatial vector representing the spatial structures of small molecule ligands and protein pockets, and the seventh set composed of the protein amino acid central carbon coordinate sequence are used as different types of protein attribute features, where the protein attribute features are used to constitute the input vector of a machine learning model, and the machine learning model is used to predict the score of a preset scoring item in the molecular screening process.
[0023] Among them, the steps for the spatial vector generation module to generate spatial vectors representing the spatial structures of small molecule ligands and protein pockets include:
[0024] Generate a small molecule ligand atom matrix based on the atoms and atomic coordinates in the small molecule ligand data.
[0025] Taking each small molecule ligand atom in the small molecule ligand atom matrix as a calculation object, calculate the spatial distance between the calculation object and each protein pocket atom respectively.
[0026] Take the protein pocket atom with the minimum distance as the target protein pocket atom corresponding to the calculation object and obtain the target protein pocket atom coordinates.
[0027] Generate a protein pocket atomic matrix based on the target protein pocket atoms and their coordinates.
[0028] Obtain a first coordinate vector based on the small molecule ligand atomic matrix, and obtain a second coordinate vector based on the protein pocket atomic matrix. Each coordinate vector includes a coordinate x vector, a coordinate y vector, and a coordinate z vector respectively.
[0029] Using any two vectors in the first coordinate vector as independent variables, and the corresponding two vectors in the second coordinate vector as reference quantities, calculate a coordinate value of the spatial vector respectively based on a preset transformation kernel function.
[0030] According to another aspect of the present invention, the present invention also provides a computing device, including a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, a method for extracting protein attribute features is implemented.
[0031] According to another aspect of the present invention, the present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, a method for extracting protein attribute features is implemented.
[0032] The multiple types of protein attribute features extracted by the method for extracting protein attribute features provided by the embodiments of the present invention can effectively describe the protein molecular features, and thus effectively improve the prediction accuracy of the machine learning model trained using the protein attribute features. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] For a clearer description of the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings in the embodiments of the present invention.
[0034] Figure 1 is a flowchart of a method for extracting protein attribute features according to an embodiment of the present invention;
[0035] Figure 2 is a flowchart of a method for calculating a spatial vector representing the spatial structure of a small molecule ligand and a protein pocket according to an embodiment of the present invention;
[0036] Figure 3 Principle block diagram of a device for extracting protein attribute features according to an embodiment of the present invention;
[0037] Figure 4 is a curve graph drawn based on the data in Table 1 according to an embodiment of the present invention;
[0038] Figure 5 is a curve graph drawn based on the data in Table 2 according to an embodiment of the present invention;
[0039] Figure 6 is a curve graph plotted based on the data in Table 3 according to an embodiment of the present invention;
[0040] Figure 7 is a curve graph plotted based on the data in Table 4 according to an embodiment of the present invention;
[0041] Figure 8 is a flowchart of a molecular docking scoring processing method applied to the molecular screening process according to an embodiment of the present invention;
[0042] Figure 9 is a schematic block diagram of a molecular docking scoring processing device applied to the molecular screening process according to an embodiment of the present invention; and
[0043] Figure 10 is a schematic diagram of the hardware structure principle of a computing device according to an embodiment of the present invention. Detailed Embodiments
[0044] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that the purpose of providing these embodiments is to make the principles and spirit of the present invention clearer and more thorough, so that those skilled in the art can better understand and then implement the principles and spirit of the present invention. The exemplary embodiments provided herein are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments herein without creative efforts shall fall within the scope of protection of the present invention.
[0045] The principles and spirit of the present invention will be elaborated in detail below with reference to several exemplary or representative embodiments of the present invention.
[0046] First, relevant contents such as concepts and technical terms involved in the embodiments of the present invention will be briefly described.
[0047] PDB (Protein Data Bank) is a standard bioinformatics data file format that records the structural information of biological macromolecules in text format, such as the chemical composition and atomic coordinates of biological macromolecules. Each line in the file is called a record, and a PDB file usually includes various different types of records and is arranged in a specific order to describe a specific structure. The types of records are, for example, the title part, the primary structure, heteroatoms, secondary structure... the coordinate part, and so on.
[0048] The FASTA format is a bioinformatics data format used to record nucleic acid sequences or peptide sequences, in which nucleic acids or amino acids are presented in single-letter codes.
[0049] Protein is an important component that makes up all cells and tissues in the human body. It is a large biomolecule or macromolecule composed of one or more amino acid residues, with a three-dimensional structure. In bioinformatics, it can have multiple representation methods depending on the scenario, such as the coordinate-based 3D structure representation method or the amino acid chain representation method.
[0050] A small molecule ligand is a substance that can bind to a receptor to produce a certain physiological effect. In bioinformatics, it is generally represented by several atoms or an atomic sequence.
[0051] Protein Binding Pockets (abbreviated as Pocket) refer to the cavities on the surface or inside of a protein that are suitable for binding ligands. The amino acid residues around the pocket determine its shape, position, physicochemical properties, and function. In bioinformatics, it is represented by several atoms or an atomic sequence that make up the amino acid residues.
[0052] Mol2vec is a vector conversion algorithm that converts molecular structures into molecular vectors.
[0053] The Fourier kernel function is a kernel function based on the Fourier transform that maps two original data through the mapping function Z into a low-dimensional Euclidean space.
[0054] Sequence alignment is a processing method that compares and aligns two or more sequences to find similarities. The structures of the elements in the same column of the aligned sequences are similar.
[0055] Clustal is a class of bioinformatics software tools for multiple sequence alignment. The more commonly used versions are ClustalW / ClustalX.
[0056] SHAP (SHapley Additive exPlanations) is a tool based on game theory methods for explaining the output of machine learning models and analyzing the importance of attributes.
[0057] The Root Mean Square Deviation (RMSD) reflects the degree of deviation of data from the average value. In protein structure determination, modeling, structure alignment, and molecular dynamics simulations, RMSD is used to measure the degree of deviation of atoms from the aligned positions. The larger the RMSD value, the larger the spatial range of the atomic movement and the smaller the steric hindrance of the atom. When the RMSD value gradually tends to be constant, it indicates that the system has reached energy equilibrium.
[0058] The present invention uses the compound data that has completed pocket search and determined the small molecule ligand to be docked as the data source, and extracts the descriptive features of the compound molecules from multiple different aspects. In one embodiment, the data source is the data of one or more compounds recorded in the PDB file. Usually, the PDB file has a corresponding description file, which records the location of the data of each compound in the PDB file. Based on the location recorded in the description file, the data belonging to the same compound can be separated from the PDB file. In the following description, the compound used for feature extraction is called the candidate compound. Each line of information in the PDB file is called a record. A PDB file usually includes various different types of records and is arranged in a specific order to describe a specific structure. Usually, the record types in the PDB file include the title part, the primary structure, the secondary structure, the hetero factor... the coordinate part, etc. Among them, a part of the content in the coordinate part is marked as "ATOM", which is the atom of the standard residue. Each record describes the atom name, residue name, rectangular coordinates (unit: angstrom), occupancy, temperature factor, etc. of each atom in the standard residue (amino acid and nucleic acid). The following is a section of data in the PDB file.
[0059]
[0060]
[0061] Each line (record) starts with the record type ATOM; field 2 in the record is the atom serial number; field 3 is the atom type, such as C or O in the above table, referring to carbon or oxygen respectively, which is represented by the element name, and the number in the atom type indicates the serial number of the atom in the amino acid; field 4 represents the residue type, for example, MOL in the above table represents a small molecule ligand, GLU, SER, ALA, etc. represent amino acids, and so on; field 5 is the residue serial number; field 6 includes three data fields, which are the x, y, and z coordinates of the atom respectively; field 7 is the occupancy; field 8 refers to the energy value (also known as the temperature factor, B value). Field 9 is the short name of the atom type. Usually, each PDB file also includes a corresponding description file, which records the positions of the small molecule ligand data, the protein pocket data, and the protein molecule data of each candidate compound. In addition, as can be seen from the PDB file data mentioned above, data records of the same type are sorted continuously. As shown in the above table, the atom serial numbers start from 1 to 51 for the small molecule ligand atoms. The data starting with residue names such as "GLY", "ALA", "VAL", "LEU", "ILE", "PRO", "PHE", "TYR", "TRP", "SER", "THR", "CYS", "MET", "ASN", "GLN", "ASP", "GLU", "LYS", "ARG", "HIS", etc. are protein molecule data. As shown in the above table, the atom serial numbers start from 52 for the protein molecule data. The pocket positions are given in the description file. For example, the pocket position description content provided in the description file is "'LEU298', 'LEU301', 'ALA302'...", from which it can be obtained that the residue LEU at the 298th position has a pocket, the residue LEU at the 301st position has a pocket, and the residue ALA at the 302nd position has a pocket.
[0062] Figure 1 It is a flowchart of a method for extracting protein attribute features according to an embodiment of the present invention. Based on the candidate compound data of the above structure, the method for extracting protein attribute features provided in this embodiment includes the following steps:
[0063] Step S11: Separate the small molecule ligand data, protein pocket data, and protein molecule data from each candidate compound data respectively. For example, when the description file of the PDB file records the positions of the small molecule ligand data, protein pocket data, and protein molecule data of each candidate compound, separate the small molecule ligand data, protein pocket data, and protein molecule data of each candidate compound from the data file according to the positions recorded in the description file. If the description file of the PDB file does not record the position of the small molecule ligand data, search in the PDB file using the small molecule ligand identifier, such as "MOL" in the aforementioned PDB file, as a keyword, so as to extract the small molecule ligand data. Of course, in different PDB files, the small molecule ligand identifier may also be other identifiers, such as "MG", "ZN", "PDB (a short name for a small molecule ligand)", etc. Refer to the specific PDB file for the specific identifier. When the description file lacks the description of the protein pocket position, some algorithms (such as the fpocket algorithm) or tools can be used to calculate the PDB file to obtain potential binding sites and thus identify the protein pocket. Some current professional platforms or tools can implement the fpocket algorithm. Therefore, when the description file lacks the description of the protein pocket position, input the PDB file into the fpocket algorithm platform or tool, and the content of the protein pocket part can be obtained. Similar tools also include SiteHound and P2Rank.
[0064] Step S12: Calculate the centroid coordinates of each type of atom in each small molecule ligand data and the centroid coordinates of each type of atom in each protein molecule data respectively. Extract the coordinates of all types of atoms from each small molecule ligand data, and then calculate the centroid coordinates of various types of atoms. The first set composed of the centroid coordinates of each type of atom in each small molecule ligand is integrated into the attribute feature data set as a kind of feature. Similarly, the second set composed of the centroid coordinates of each type of atom in each protein molecule is integrated into the attribute feature data set as a kind of feature.
[0065] In this step, in the small molecule ligand data, extract the same type of atoms and their coordinates, such as the coordinates of all hydrogen atoms, the coordinates of all oxygen atoms, the coordinates of all carbon atoms, etc. Then, calculate the centroid coordinates of each type of atom according to the following formula (1-1) to obtain the centroid coordinates of various types of atoms.
[0066] A = (a1 + a2 + …… + a n ) / n (1-1)
[0067] Among them, A is the centroid coordinate, a1……a n Are the coordinates of one atom of the same type and different atoms respectively, and n is the number of atoms.
[0068] Specifically, the calculation formula for the barycentric coordinate x is shown in (1-2):
[0069] x ij = (x1 + x2 + …… x n ) / n (1-2)
[0070] The calculation formula for the barycentric coordinate y is shown in (1-3):
[0071] y ij = (y1 + y2 + …… y n ) / n (1-3)
[0072] The calculation formula for the barycentric coordinate z is shown in (1-4):
[0073] z ij = (z1 + z2 + …… z n ) / n (1-4)
[0074] Wherein, when i is L, it represents a small molecule ligand molecule, and when i is P, it represents a protein molecule; j is the atomic type, which is replaced by the element name. For example, when j is O, it is an oxygen atom, and when j is H, it is a hydrogen atom.
[0075] In an optional step S13, calculate the difference in the barycentric coordinates of the like atoms of each small molecule ligand and each protein molecule, and use the difference in the barycentric coordinates of the like atoms as a type of protein attribute feature, so as to further show the differences in the spatial structures between the small molecule ligand and the protein molecule.
[0076] Step S14, generate multiple types of molecular vectors. Specifically, it includes:
[0077] For the small molecule ligand data, use the Mol2voc algorithm or the PCM2vec algorithm or the MorganFP algorithm to convert each small molecule ligand data into a vector, named here as the small molecule ligand molecular vector, and store the third set composed of each small molecule ligand molecular vector into the attribute feature set.
[0078] For the protein pocket data and the protein molecule data, since each protein pocket and each protein molecule are composed of multiple standard amino acids, there are already trained vectors in the industry. Search the molecular vector library in the industry to find the corresponding vectors, so as to obtain the corresponding protein pocket vector and protein molecule vector. Store the fourth set composed of each protein pocket molecular vector and the fifth set composed of each protein molecule vector into the attribute feature set respectively.
[0079] Optionally, each small molecule ligand data and each protein pocket data are combined together to obtain a set of combined data; the Mol2voc algorithm, the PCM2vec algorithm, or the MorganFP algorithm is used to convert each set of combined data into a combined molecular vector, simply referred to as a combine vector, and stored in the attribute feature set.
[0080] The dimensions of the various molecular vectors described above are the same.
[0081] Step S15: Generate a spatial vector representing the spatial structures of the small molecule ligand and the protein pocket. Specifically, based on the atomic coordinates in each small molecule ligand data and each protein pocket data in each candidate compound data, a spatial vector representing the spatial structures of the small molecule ligand and the protein pocket is generated. Refer to Figure 2 , Figure 2 which is a flowchart of a method for calculating a spatial vector representing the spatial structures of a small molecule ligand and a protein pocket according to an embodiment of the present invention. The method for calculating a spatial vector representing the spatial structures of a small molecule ligand and a protein pocket includes the following steps:
[0082] Step S151: Generate a small molecule ligand atomic coordinate matrix L. That is, based on the atoms and their coordinates in the small molecule ligand data, a small molecule ligand atomic coordinate matrix L is generated. Among them, the small molecule ligand atomic coordinate matrix L has three rows, which respectively represent the coordinate x, coordinate y, and coordinate z of the atomic coordinates, and the columns are small molecule ligand atoms, and the number of columns is the number of small molecule ligand atoms.
[0083] Step S152: Generate a protein pocket atomic coordinate matrix W. Specifically, taking each small molecule ligand atom in the small molecule ligand atomic coordinate matrix L as a calculation object, the spatial distance between the calculation object and each protein pocket atom in the protein pocket data is calculated respectively, and the protein pocket atom with the smallest distance is determined as the target protein pocket atom corresponding to the calculation object, and the target protein pocket atom coordinates are obtained. Similarly, the corresponding protein pocket atoms of other small molecule ligand atoms in the small molecule ligand atomic coordinate matrix L are calculated until all small molecule ligand atoms are calculated, so as to obtain the target protein pocket atoms and their coordinates with the same number as the number of small molecule ligand atoms, and then a protein pocket atomic coordinate matrix W is obtained. Each protein pocket atom in the protein pocket atomic coordinate matrix W is the protein pocket atom with the smallest spatial distance from the corresponding small molecule ligand atom in the small molecule ligand atomic coordinate matrix. The protein pocket atomic coordinate matrix W has three rows, which respectively represent the coordinate x, coordinate y, and coordinate z of the atomic coordinates, and the columns are protein pocket atoms, and the number of columns is the same as the number of columns of the small molecule ligand atomic coordinate matrix.
[0084] Step S153: Obtain the first coordinate vector based on the small molecule ligand atomic coordinate matrix, and obtain the second coordinate vector based on the protein pocket atomic coordinate matrix. Each coordinate vector includes a coordinate x vector, a coordinate y vector, and a coordinate z vector. For example, the first coordinate vector includes the coordinate x vector x L , which is the vector composed of the x rows in the small molecule ligand atomic coordinate matrix; the first coordinate vector includes the coordinate y vector y L , which is the vector composed of the y rows in the small molecule ligand atomic coordinate matrix; the first coordinate vector includes the coordinate z vector z L , which is the vector composed of the z rows in the small molecule ligand atomic coordinate matrix. The second coordinate vector includes the coordinate x vector W x , which is the vector composed of the x rows in the protein pocket atomic coordinate matrix; the second coordinate vector includes the coordinate y vector W y , which is the vector composed of the y rows in the protein pocket atomic coordinate matrix; the second coordinate vector includes the coordinate z vector W z , which is the vector composed of the z rows in the protein pocket atomic coordinate matrix.
[0085] Step S154: Calculate each coordinate value in the spatial vector respectively based on the preset transformation kernel function. Specifically, take any two vectors in the first coordinate vector as independent variables, take the corresponding two vectors in the second coordinate vector as reference quantities, and calculate a coordinate value of the spatial vector respectively based on the preset transformation kernel function. Since the coordinate vector includes three vectors, three coordinate values can be obtained, that is, the spatial vector includes three coordinate values.
[0086] In one embodiment, the preset transformation kernel function is shown in the following formula (2-1):
[0087]
[0088] where a and b are two independent variables, A and B are reference quantities, and D is a constant.
[0089] Based on formula (2-1), take two vectors from the coordinate x vector x L , the coordinate y vector y L , and the coordinate z vector z as independent variables, and the calculated values are used as a coordinate value of the spatial vector. Specifically, the three coordinate values obtained according to the following formulas (2-2), (2-3), and (2-4) constitute the spatial vector Z[Z(x L , y L ), Z(y L , z L ), Z(z L , x L )].
[0090]
[0091] In this embodiment, D is an arbitrary integer greater than or equal to the number of atoms of the small molecule ligand and less than or equal to the number of atoms of the protein pocket.
[0092] Among them, based on the trigonometric function types in the foregoing formula (2-1), they can also be interchanged, that is, the following formula (2-5):
[0093]
[0094] For the convenience of description, the present invention refers to formula (2-1) and formula (2-5) as Fourier kernel functions.
[0095] The sixth set composed of each spatial vector representing the spatial structures of the small molecule ligand and the protein pocket is stored as a type of protein attribute feature in the attribute feature set.
[0096] Since the number of atoms constituting a protein molecule is large and the number of atoms in different protein molecules is different, when describing the atomic spatial structure of a protein, if the coordinate values of each atom are used as features, not only is the data volume huge, but the quantity is not fixed, which does not meet the requirement in machine learning that the dimension numbers of one type of feature in the input vector should be the same. The present invention extracts three numerical values from a PDB file as a type of feature for machine learning through the above processing process. Its dimension is three, which meets the requirement of fixed dimension number. And in the extraction process of the present invention, the protein pocket atoms closest to the small molecule ligand atoms in terms of spatial distance are combined, and the protein pocket atoms with the best binding force to the small molecule ligand atoms are reflected in the data. Therefore, the screening accuracy can be improved during the screening process.
[0097] Step S16, perform sequence alignment. Specifically, the separated multiple small molecule ligand data, protein molecule data, and protein pocket data are respectively converted from the PDB file format to the FASTA format, so as to obtain multiple small molecule ligand sequences, multiple protein molecule sequences, and multiple protein pocket sequences. Then, a sequence alignment tool is used to align the multiple small molecule ligand sequences, multiple protein molecule sequences, and multiple protein pocket sequences respectively. The missing data in the middle is filled with the number 0 during alignment. The alignment tool is, for example, ClustalW / ClustalX. Among them, by aligning the protein molecule sequence and the protein pocket sequence, amino acids with the same or similar physical and chemical properties are in the same column, and the sequence lengths are unified. By aligning the small molecule ligand sequence, atoms with the same or similar physical and chemical properties are in the same column, and the length of the small molecule ligand sequence is unified.
[0098] Step S17: Extract coordinates to obtain a coordinate sequence. Combining with the original PDB file, the central carbon coordinates of each aligned amino acid in the aligned protein molecule sequence and the aligned protein pocket sequence are respectively extracted, and an amino acid central carbon coordinate sequence is formed to obtain a protein amino acid central carbon coordinate sequence and a protein pocket amino acid central carbon coordinate sequence. For the aligned small molecule ligand sequence, the coordinates of each aligned atom are extracted to form a small molecule atom coordinate sequence. Among them, the protein amino acid central carbon coordinate sequence forms the seventh set, which is stored as a type of protein attribute feature in the attribute feature set. Similarly, the protein pocket amino acid central carbon coordinate sequence and the small molecule atom coordinate sequence respectively constitute the eighth set and the ninth set, and are respectively stored as a type of protein attribute feature in the attribute feature set.
[0099] The method for extracting protein attribute features provided by the present invention uses a PDB file in standard format as the data source. For convenience of use, the original PDB file can also be processed, such as identifying and classifying the data in the original PDB file, and reordering it, and then generating a corresponding description file. For example, based on small molecule ligand identifiers such as "MOL", "MG", "ZN", "PDB", etc., the small molecule ligand data is identified from the data file, and the position of the small molecule ligand data in the data file is recorded; based on residue names such as "GLY", "ALA", "VAL", "LEU", "ILE", "PRO", "PHE", "TYR", "TRP", "SER", "THR", "CYS", "MET", "ASN", "GLN", "ASP", "GLU", "LYS", "ARG", "HIS", etc., the protein molecule data is identified from the data file, and the position of the protein molecule data in the data file is recorded; the protein pocket data is identified from the data file based on the protein pocket recognition algorithm, and the position of the protein pocket data in the data file is recorded. When performing the aforementioned feature extraction, the position of various types of data can be determined through the description file.
[0100] Optionally, the data in the attribute feature data set obtained according to the aforementioned method can also be subjected to data cleaning and regularization. For example: correcting the errors that occur during vector conversion by the Mol2vec algorithm, or converting all the converted vectors proportionally into data between -1 and 1, or when there is no data due to a certain protein lacking a certain atom type when calculating the centroid, filling it with the number 0, etc., so that the data in the attribute feature data set better meets the data requirements of machine learning algorithms.
[0101] On the other hand, the present invention also provides a device for extracting protein attribute features. Refer to Figure 3 , Figure 3It is a block diagram of the principle of an apparatus for extracting protein attribute features according to an embodiment of the present invention. The apparatus 10 for extracting protein attribute features in this embodiment includes a data extraction module 101, a data separation module 102, a centroid coordinate calculation module 103, a vector conversion module 104, a coordinate sequence extraction module 105, and a spatial vector generation module 106. Among them, the data extraction module 101 is configured to extract a plurality of candidate compound data from a compound data file. The data separation module 102 is configured to separate small molecule ligand data, protein pocket data, and protein molecule data from each candidate compound data. The centroid coordinate calculation module 103 is configured to calculate the centroid coordinates of each type of atom in each small molecule ligand data and the centroid coordinates of each type of atom in each protein molecule data respectively. The vector conversion module 104 is configured to convert each small molecule ligand data, each protein pocket data, and each protein molecule data into corresponding molecular vectors respectively. The coordinate sequence extraction module 105 is configured to perform sequence alignment on the protein molecule data of a plurality of candidate compounds; extract the central carbon coordinates of each amino acid from the aligned protein molecule sequences, and form a protein amino acid central carbon coordinate sequence. The spatial vector generation module 106 is configured to generate a spatial vector characterizing the spatial structures of the small molecule ligand and the protein pocket based on each atom coordinate in each small molecule ligand data and each atom coordinate in each protein pocket data. The centroid coordinate calculation module 103, the vector conversion module 104, the coordinate sequence extraction module 105, and the spatial vector generation module 106 store the corresponding attribute features obtained after processing into an attribute feature set respectively. The attribute feature set can be a database, a server, a certain computer, or cloud space, or a storage space, one or more data files are opened in these storage devices.
[0102] In a further embodiment, the centroid coordinate calculation module 103 also calculates the difference in centroid coordinates of the same type of atoms of each small molecule ligand and each protein molecule, and stores the difference in centroid coordinates of the same type of atoms as a type of protein attribute feature into the attribute feature set.
[0103] In a further embodiment, the vector conversion module 104 is further configured to merge each small molecule ligand data and each protein pocket data into combined data; convert the combined data into a combined molecular vector; wherein, each combined molecular vector is stored as a type of protein attribute feature into the attribute feature set.
[0104] In a further embodiment, the coordinate sequence extraction module 105 is configured to further perform sequence alignment on the multiple small molecule ligands of the multiple candidate compounds; extract the atomic coordinates constituting the small molecules from the aligned small molecule ligand sequences, and form a small molecule atomic coordinate sequence; wherein, the small molecule atomic coordinate sequence is used as a type of protein attribute feature; and / or, the coordinate sequence extraction module is configured to further perform sequence alignment on the multiple protein pockets of the multiple candidate compounds; extract the central carbon coordinates of each amino acid from the aligned protein pocket sequences, and form a protein pocket amino acid central carbon coordinate sequence; wherein, the protein pocket amino acid central carbon coordinate sequence is stored in the attribute feature set as a type of protein attribute feature.
[0105] In the molecular screening process, a new compound is obtained by docking a small molecule ligand to a pocket of a protein molecule. To evaluate the new compound obtained after molecular docking, multiple evaluation indicators are set, such as geometric complementarity terms, RMSD of the small molecule ligand, RMSD of the side chains after the pocket and the small molecule are combined and bound, interface contact area, van der Waals interaction energy, and electrostatic interaction energy, etc. The scores of each evaluation indicator are calculated using a scoring function, and the new compound obtained after molecular docking is evaluated based on the scores. For convenience of description, the foregoing evaluation indicators are referred to as scoring items. To obtain the scores of each scoring item, corresponding machine learning models are trained based on each scoring item, and features characterizing the characteristics of the molecule or protein are used as the input of the model to predict the scores of the corresponding scoring items.
[0106] An embodiment of the present invention provides a training method for a machine learning model for predicting scoring items of molecular docking in the molecular screening process. Among them, various protein attribute features are obtained from the original compound sample data by using the foregoing method for extracting protein attribute features and stored in the sample feature data set. Features are selected from the sample feature data set based on the scoring items of molecular docking and the applied machine learning model to generate input samples. In one embodiment, the importance of the features in the sample feature data set is analyzed by SHAP, so as to determine the protein attribute features adapted to the scoring items and the machine learning model. Each protein attribute feature is used as a feature of an input vector, and all the determined protein attribute features are combined to form an input sample. The machine learning model is iteratively trained and verified based on the input sample until the requirements are met. Among them, the machine learning model in this embodiment can adopt common ones such as linear regression, support vector machine, convolutional neural network, etc. Since the dimensionality of each type of protein attribute feature extracted by the present invention is fixed, it can be widely applied to the evaluation of various different compounds.
[0107] In one embodiment, taking the RMSD of small molecule ligands as an example of the scoring item, during the training process, the RMSD of a group of small molecule ligands is used as the label column Y. An input sample is input into the machine learning model to obtain the RMSD of the small molecule ligand, which is used as the prediction column Pre-Y. The prediction column Pre-Y is compared with the label column Y to determine whether the difference between the two meets the requirements. For example, when the prediction column Pre-Y is exactly the same as the label column Y, or the difference between the two is less than the threshold, or the prediction value meets a certain accuracy rate, it is considered to meet the requirements. When the requirements are not met, the internal weights of the model are corrected through the model feedback mechanism, and then an input sample is input into the machine learning model again for prediction and comparison. After multiple iterations, a model that meets the requirements will finally be obtained. The validation set is used to verify whether the model truly meets the requirements. If it meets the requirements after verification, the model can be saved at this time for predicting the RMSD of small molecule ligands in molecular docking evaluation in the future.
[0108] Since the features in the input samples used in the training and validation of the present invention are multiple protein attribute features obtained by the method of extracting protein attribute features described above, compared with traditional input samples, the machine learning model trained by the present invention can effectively reduce the RMSD of small molecule ligands. Table 1 shows the data obtained after 1000 iterations of the input samples before and after extracting attribute features using the method of the present invention when training a machine learning model for the scoring item of the RMSD of small molecule ligands. Figure 4 is a curve graph drawn based on the data in Table 1 according to an embodiment of the present invention. Table 2 shows the data obtained after 1000 iterations of the input samples before and after extracting attribute features using the method provided by the present invention when validating a machine learning model for the scoring item of the RMSD of small molecule ligands. Figure 5 is a curve graph drawn based on the data in Table 2 according to an embodiment of the present invention.
[0109] In this embodiment, before extracting attribute features, the centroid coordinates of hydrogen atoms and oxygen atoms in the small molecule ligand molecule are used as the features of the input sample. After extracting attribute features, the centroid coordinates of each type of atom in the small molecule ligand, the centroid coordinates of each type of atom in each protein molecule, the difference in centroid coordinates of the same type of atoms between the small molecule ligand and the protein molecule, each small molecule ligand molecular vector, each protein pocket vector, each protein molecular vector, the combined vector of the small molecule ligand and the protein pocket, the spatial vector of the spatial structure of each small molecule ligand and the protein pocket, and the combined sequence of the protein amino acid central carbon coordinate sequence, the small molecule atom coordinate sequence, and the protein pocket amino acid central carbon coordinate sequence are used as the input sample.
[0110] Table 1
[0111]
[0112] Table 2
[0113]
[0114]
[0115] Through Figure 4 and Figure 5 The data curves shown, it can be clearly seen that after adopting the protein attribute features extracted by the present invention, the RMSD of the small molecule ligand is significantly reduced.
[0116] In another embodiment, taking the RMSD of the side chain after the binding of the pocket and the small molecule ligand as the scoring item as an example, the training and verification processes are the same as those in the previous embodiment, and will not be elaborated here. Table 3 shows the data obtained after 1000 iterations of the input samples before and after extracting the attribute features by the method provided by the present invention when training the machine learning model for the scoring item of the RMSD of the side chain after the binding of the pocket and the small molecule ligand. Figure 6 is a curve graph drawn based on the data in Table 3 according to an embodiment of the present invention. Table 4 shows the data obtained after 1000 iterations of the input samples before and after extracting the attribute features by the method provided by the present invention when verifying the machine learning model for the scoring item of the RMSD of the side chain after the binding of the pocket and the small molecule ligand. Figure 7 is a curve graph drawn based on the data in Table 4 according to an embodiment of the present invention.
[0117] Table 3
[0118]
[0119]
[0120] Table 4
[0121]
[0122] Through Figure 6 and Figure 7 The data curves shown, it can be clearly seen that after adopting the protein attribute features extracted by the present invention, the RMSD of the side chain after the binding of the pocket and the small molecule ligand is significantly reduced.
[0123] The above training and verification iteration times are 1000 times. When the iteration times are increased, the predicted values of the RMSD of the side chain after the binding of the pocket and the small molecule ligand and the RMSD after the docking of the small molecule ligand can still be further reduced. Based on experiments, compared with the traditional input samples, when using the protein features extracted by the method of the present invention as the input samples, the prediction accuracy is greatly improved. For individual scoring items, the prediction accuracy can even be improved by 2 orders of magnitude at a sufficient number of iteration times.
[0124] In another aspect, referring to Figure 8 , Figure 8 is a flowchart of a molecular docking scoring processing method applied to the molecular screening process according to an embodiment of the present invention, specifically including the following steps:
[0125] Step S21, extracting protein attribute features to obtain a feature data set. Specifically, the protein attribute features are extracted from the target compound data by using the aforementioned method for extracting protein features to obtain a feature data set. Among them, the extracted protein attribute features include the centroid coordinates of each type of atom in each small molecule ligand, the centroid coordinates of each type of atom in each protein molecule, the molecular vector of each small molecule ligand, the protein pocket vector, the protein molecule vector, the spatial vector of the spatial structure of each small molecule ligand and protein pocket, and the protein amino acid central carbon coordinate sequence. It may further include the difference in centroid coordinates of the same type of atoms between the small molecule ligand and the protein molecule, the combined molecular vector obtained after the small molecule ligand and the protein pocket are combined, the small molecule atom coordinate sequence, the protein pocket amino acid central carbon coordinate sequence, and so on.
[0126] Step S22, selecting features from the feature data set of the target compound based on the scoring item and the applied machine learning model to generate an input vector. In one embodiment, each type of protein attribute feature obtained above can be used as a feature of the input vector, thereby obtaining the input vector. In another embodiment, based on different scoring items, SHAP can be used to perform importance analysis on the protein attribute features in the sample feature data set, and only some of the protein attribute features obtained after the importance analysis are used to form the input vector.
[0127] Step S23, inputting the input vector into the machine learning model to obtain the predicted score of the scoring item. The machine learning model in this embodiment is trained according to the aforementioned method, so when predicting the scoring item, the prediction accuracy can be effectively improved.
[0128] Referring to Figure 9 , Figure 9 is a principle block diagram of a molecular docking scoring processing device applied to the molecular screening process according to an embodiment of the present invention. The processing device 20 in this embodiment includes a protein attribute feature extraction module 201, an input vector generation module 202, and a prediction module 203. The processing device in this example can be used for model training and can also be used for the scoring processing of molecular docking in the molecular screening process when the model training is completed.
[0129] During model training, the protein attribute feature extraction module 201 extracts protein attribute features from the sample compound data file according to the aforementioned method for extracting protein features to obtain a feature data set. The input vector generation module 202 selects features from the feature data set based on the scoring items of molecular docking and the applied machine learning model to generate input vectors, stores them in the database, and simultaneously sends a notification to the prediction module 203. The prediction module 203 reads a set of input vectors from the database, inputs them into the machine learning model that needs to be trained, obtains prediction values from the machine learning model, compares them with the preset label values, adjusts the parameters and weights of the model based on the comparison results, and re-inputs new input vectors into the machine learning model. After a certain number of iterations, when the prediction values meet the requirements, the iteration stops, and the machine learning model is saved, as Figure 9 shown by the dashed line in
[0130] During actual application, the protein attribute feature extraction module 201 extracts protein attribute features from the target compound data file according to the aforementioned method for extracting protein features to obtain a feature data set. The input vector generation module 202 selects features from the feature data set based on the scoring items of molecular docking and the applied machine learning model to generate input vectors. The prediction module 203 reads the input vectors and retrieves the trained machine learning model, and inputs the input vectors into the trained machine learning model to obtain prediction scores.
[0131] Figure 10 is a schematic diagram of the hardware structure principle of a computing device according to an embodiment of the present invention. As Figure 10 shown, the computing device may include a processor 601 and a memory 602 storing computer program instructions. When the processor 601 executes the computer program instructions, it implements the aforementioned method for extracting protein attribute features, and / or the aforementioned method for training the machine learning model, or / and, the aforementioned method for molecular docking scoring processing.
[0132] Specifically, the aforementioned processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0133] The memory 602, being a computer-readable storage medium, may include a large-capacity memory for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In suitable cases, the memory 602 may include removable or non-removable (or fixed) media. In suitable cases, the memory 602 may be inside or outside the integrated gateway disaster recovery device. In a particular embodiment, the memory 602 is a non-volatile solid-state memory.
[0134] In one example, the computing device may further include a communication interface 603 and a bus 610. Among them, as Figure 10 shown, the processor 601, the memory 602, and the communication interface 603 are connected through the bus 610 and complete communication with each other. The communication interface 603 is mainly used to implement the communication between the modules, devices, units, and / or devices in the embodiments of the present invention. The bus 610 includes hardware, software, or both, and couples the components of the online data flow billing device to each other. By way of example and not limitation, the bus may include an accelerated graphics port (AGP) or other graphics bus, an enhanced industry standard architecture (EISA) bus, a front-side bus (FSB), a hypertransport (HT) interconnect, an industry standard architecture (ISA) bus, an infinite bandwidth interconnect, a low pin count (LPC) bus, a memory bus, a microchannel architecture (MCA) bus, a peripheral component interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a serial advanced technology attachment (SATA) bus, a video electronics standards association local (VLB) bus, or other suitable buses, or a combination of two or more of these. In suitable cases, the bus 610 may include one or more buses. Although the embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.
[0135] The computing device in the embodiments of the present invention may be a server, a personal computer, or other forms of computing devices.
[0136] As described above, the foregoing is only the specific implementation manner of the present invention. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, modules, and units may refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A method for extracting protein attribute features, characterized in that, Including: Extracting a plurality of candidate compound data from a compound data file; Separating small molecule ligand data, protein pocket data, and protein molecule data from each candidate compound data respectively to obtain a plurality of small molecule ligand data, a plurality of protein pocket data, and a plurality of protein molecule data; Calculating the centroid coordinates of each type of atom in each small molecule ligand data and the centroid coordinates of each type of atom in each protein molecule data respectively; Converting each small molecule ligand data, each protein pocket data, and each protein molecule data into corresponding molecular vectors respectively to obtain a plurality of small molecule ligand molecular vectors, a plurality of protein pocket molecular vectors, and a plurality of protein molecule vectors; Aligning the sequences of a plurality of the protein molecule data respectively to obtain a plurality of protein molecule sequences; Extracting the central carbon coordinates of each amino acid from a plurality of the protein molecule sequences respectively and forming a protein amino acid central carbon coordinate sequence; Generating spatial vectors characterizing the spatial structures of the small molecule ligand and the protein pocket respectively based on the atomic coordinates in each small molecule ligand data in each candidate compound data and the atomic coordinates in each protein pocket data; Taking a first set composed of the centroid coordinates of each type of atom in each small molecule ligand, a second set composed of the centroid coordinates of each type of atom in each protein molecule, a third set composed of each small molecule ligand molecular vector, a fourth set composed of each protein pocket molecular vector, a fifth set composed of each protein molecule vector, a sixth set composed of each spatial vector characterizing the spatial structures of the small molecule ligand and the protein pocket, and a seventh set composed of the protein amino acid central carbon coordinate sequence as different types of protein attribute features, wherein the protein attribute features are used to constitute the input vector of a machine learning model, and the machine learning model is used to predict the score of a preset scoring item in the molecular screening process; Wherein, the step of generating the spatial vectors characterizing the spatial structures of the small molecule ligand and the protein pocket includes: Generating a small molecule ligand atom matrix based on the atoms and atomic coordinates in the small molecule ligand data; Taking each small molecule ligand atom in the small molecule ligand atom matrix as a calculation object and calculating the spatial distance between the calculation object and each protein pocket atom respectively; Taking the protein pocket atom with the minimum distance as the target protein pocket atom corresponding to the calculation object and obtaining the target protein pocket atom coordinates; Generating a protein pocket atom matrix based on the target protein pocket atom and the target protein pocket atom coordinates; Obtaining a first coordinate vector based on the small molecule ligand atom matrix and a second coordinate vector based on the protein pocket atom matrix, and each coordinate vector includes a coordinate x vector, a coordinate y vector, and a coordinate z vector respectively; Taking any two vectors in the first coordinate vector as independent variables and taking the corresponding two vectors in the second coordinate vector as reference quantities, and calculating a coordinate value of the spatial vector respectively based on a preset transformation kernel function.
2. The method for extracting protein attribute features according to claim 1, wherein The step of calculating the centroid coordinates of each type of atom in each small molecule ligand includes: Identifying small molecule ligand atom data from the small molecule ligand data in each candidate compound data based on the atom identifier; Based on the atomic chemical element names or abbreviations, identify each type of atom and extract the coordinate data of all atoms of each type. Based on the centroid formula and the coordinate data of all atoms of each type, calculate the centroid coordinates of each type of atom.
3. The method for extracting protein attribute features according to claim 1, wherein After separately calculating the centroid coordinates of each type of atom in each small molecule ligand data and each protein molecule data, the method further includes: calculating the difference in the centroid coordinates of the same type of atoms between each small molecule ligand and each protein molecule, and taking the difference in the centroid coordinates of the same type of atoms as a type of protein attribute feature.
4. The method for extracting protein attribute features according to claim 1, wherein After separately separating the small molecule ligand data, protein pocket data, and protein molecule data from each candidate compound data, the method further includes: Combining a small molecule ligand data and a protein pocket data into combined data; Converting the combined data into a combined molecular vector; Taking the combined molecular vector as a type of protein attribute feature.
5. The method for extracting protein attribute features according to claim 1, wherein After obtaining multiple small molecule ligand data, multiple protein pocket data, and multiple protein molecule data, the method further includes: Separately performing sequence alignment on the multiple small molecule ligand data to obtain multiple small molecule ligand sequences; Separately extracting the atomic coordinates constituting the small molecule ligand from the multiple small molecule ligand sequences and forming a small molecule atomic coordinate sequence; Taking the small molecule atomic coordinate sequence as a type of protein attribute feature; and / or Separately performing sequence alignment on the multiple protein pocket data to obtain multiple protein pocket sequences; Separately extracting the central carbon coordinates of each amino acid from the multiple protein pocket sequences and forming a protein pocket amino acid central carbon coordinate sequence; Taking the protein pocket amino acid central carbon coordinate sequence as a type of protein attribute feature.
6. The method for extracting protein attribute features according to claim 1, characterized in that Separately separating the small molecule ligand data, protein pocket data, and protein molecule data from each candidate compound data includes: Obtaining the description file of the compound data file; Obtaining the positions of the small molecule ligand data, protein pocket data, and / or protein molecule data of each candidate compound recorded in the description file in the compound data file; Based on the positions of the small molecule ligand data, protein pocket data, and / or protein molecule data of each candidate compound recorded in the description file in the compound data file, separating the small molecule ligand data, protein pocket data, and / or protein molecule data of each candidate compound from the compound data file.
7. The method for extracting protein attribute features according to claim 1, wherein Separately separating the small molecule ligand data, protein pocket data, and protein molecule data from each candidate compound data further includes: Based on the small molecule ligand identifier, separately identifying the small molecule ligand data from each candidate compound data and recording the position of the small molecule ligand data in the compound data file; Based on the residue name, separately identifying the protein molecule data from each candidate compound data and recording the position of the protein molecule data in the compound data file; Based on the protein pocket recognition algorithm, protein pocket data is respectively identified from each candidate compound data, and the positions of the protein pocket data in the compound data file are recorded.
8. The method for extracting protein attribute features according to claim 1, wherein The steps for calculating a coordinate value of a spatial vector based on a preset transformation kernel function include: Calculating the first product of the first vector in the first coordinate vector and the corresponding first vector in the second coordinate vector, and the second product of the second vector in the first coordinate vector and the corresponding second vector in the second coordinate vector, respectively; Calculating the first trigonometric function value of the first product and the second trigonometric function value of the second product, respectively; Calculating the third trigonometric function value of the third product of the first trigonometric function value and the second trigonometric function value; and Calculating the fourth product of the third trigonometric function value and a constant factor, and taking the fourth product as a coordinate value of the spatial vector; Wherein, the first trigonometric function and the second trigonometric function are of the same type, and the third trigonometric function and the first trigonometric function are of different types.
9. An apparatus for extracting protein attribute features, characterized in that, Including: A data extraction module configured to extract a plurality of candidate compound data from a compound data file; A data separation module configured to separately isolate small molecule ligand data, protein pocket data, and protein molecule data from each candidate compound data, obtaining a plurality of small molecule ligand data, a plurality of protein pocket data, and a plurality of protein molecule data; A centroid coordinate calculation module configured to calculate the centroid coordinates of each type of atom in each small molecule ligand data and the centroid coordinates of each type of atom in each protein molecule data, respectively; A vector conversion module configured to convert each small molecule ligand data, each protein pocket data, and each protein molecule data into corresponding molecular vectors, obtaining a plurality of small molecule ligand molecular vectors, a plurality of protein pocket molecular vectors, and a plurality of protein molecule vectors; A coordinate sequence extraction module configured to align the sequences of a plurality of the protein molecule data respectively, obtaining a plurality of protein molecule sequences; Extracting the central carbon coordinates of each amino acid from a plurality of the protein molecule sequences respectively, and forming a protein amino acid central carbon coordinate sequence; And A spatial vector generation module configured to generate spatial vectors characterizing the spatial structures of the small molecule ligand and the protein pocket respectively based on the atomic coordinates in each small molecule ligand data and the atomic coordinates in each protein pocket data in each candidate compound data; Taking a first set composed of the centroid coordinates of each type of atom in each small molecule ligand, a second set composed of the centroid coordinates of each type of atom in each protein molecule, a third set composed of each small molecule ligand molecular vector, a fourth set composed of each protein pocket molecular vector, a fifth set composed of each protein molecule vector, a sixth set composed of each spatial vector characterizing the spatial structures of the small molecule ligand and the protein pocket, and a seventh set composed of the protein amino acid central carbon coordinate sequence as different types of protein attribute features, wherein the protein attribute features are used to constitute the input vector of a machine learning model, and the machine learning model is used to predict the score of a preset scoring item in the molecular screening process. Among them, the steps of the spatial vector generation module generating spatial vectors representing the spatial structures of small molecule ligands and protein pockets include: Generating a small molecule ligand atom matrix based on the atoms and atomic coordinates in the small molecule ligand data; Taking each small molecule ligand atom in the small molecule ligand atom matrix as a calculation object, and respectively calculating the spatial distance between the calculation object and each protein pocket atom; Taking the protein pocket atom with the minimum distance as the target protein pocket atom corresponding to the calculation object, and obtaining the target protein pocket atom coordinates; Generating a protein pocket atom matrix based on the target protein pocket atoms and the target protein pocket atom coordinates; Obtaining a first coordinate vector based on the small molecule ligand atom matrix, and obtaining a second coordinate vector based on the protein pocket atom matrix. Each type of coordinate vector includes a coordinate x vector, a coordinate y vector, and a coordinate z vector; Taking any two vectors in the first coordinate vector as independent variables, taking the corresponding two vectors in the second coordinate vector as reference quantities, and respectively calculating a coordinate value of the spatial vector based on a preset transformation kernel function.
10. A computing device, comprising a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for extracting protein attribute features as described in any one of claims 1-8 is implemented.
11. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by the processor, the method for extracting protein attribute features as described in any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Method for training molecule binding model and molecule screening method and device
CN114283899A
Drug molecule screening method and system
CN114530210A
Targeting antiglioma protein and application thereof
CN101824084A
Method and device for model training, drug screening and affinity prediction
CN114333986A