Methods for extracting protein surface features and methods for generating extraction models, and related equipment.

By generating a protein surface feature extraction model and utilizing neural network training and multi-scale surface fingerprint descriptor data, the problems of insufficient computational efficiency and accuracy in existing technologies are solved, and efficient and accurate protein surface feature extraction at multiple scales is achieved.

CN119920317BActive Publication Date: 2025-10-31SUZHOU MERNA THERAPEUTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311433711.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-10-31
Estimated Expiration
2043-10-31

AI Technical Summary

Technical Problem

Existing methods for extracting protein 3D surface interaction fingerprints cannot meet users' diverse needs for computational efficiency and accuracy, resulting in insufficient accuracy in protein surface feature extraction.

Method used

By acquiring a training dataset of protein surface fingerprint descriptors, a protein surface feature extraction model is generated using neural network training. Then, by utilizing surface fingerprint descriptor data at multiple scales and combining chemical properties and geometric structure features, protein surface features are extracted at multiple scales.

Benefits of technology

It enables the extraction of diverse protein surface features based on computational efficiency and accuracy requirements, thereby improving the accuracy of protein surface feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920317B_ABST
    Figure CN119920317B_ABST
Patent Text Reader

Abstract

A method for extracting protein surface features, a method for generating an extraction model, and related equipment are disclosed. The method includes: acquiring a protein surface fingerprint descriptor training dataset, which comprises surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales. The scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and accuracy. A neural network is trained using the protein surface fingerprint descriptor training dataset to obtain a protein surface feature extraction model. The technical solution of this invention enables protein surface feature extraction at multiple scales, meets diverse user requirements for computational efficiency and accuracy, and helps improve the accuracy of protein surface feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics processing technology, and in particular to a method for extracting protein surface features, a method for generating extraction models, and related equipment. Background Technology

[0002] Proteins are the core biological macromolecules of living organisms, responsible for performing almost all cellular processes. Their functions depend primarily on protein complexes assembled in specific ways, while the properties of the protein surface determine the type and strength of interactions with other molecules. Characterizing protein surface interactions is crucial for explaining the operation of biological systems, drug discovery, and disease treatment.

[0003] The molecular surface of a protein is a smooth, compact surface composed of atoms at its boundaries. This surface exhibits chemical and geometric characteristics and is more directly related to the interactions and functions of biomolecules than to its sequence and structure.

[0004] A recently proposed method for extracting molecular surface interaction fingerprints (MaSIF) of proteins uses one or more geodesic convolutional layers to extract input features from protein molecular structure data, including chemical features, geometric features, or information related to complementary proteins, generating higher-level structural descriptors. However, existing methods for extracting protein 3D surface interaction fingerprints cannot meet users' varying demands for computational efficiency and accuracy in the process of extracting protein surface features. Summary of the Invention

[0005] The problem addressed by this invention is to provide a method for extracting protein surface features, a method for generating extraction models, and related equipment. This method enables the extraction of protein surface features at multiple scales, meets the diverse computational efficiency and accuracy requirements of users, and helps improve the accuracy of protein surface feature extraction.

[0006] To address the above problems, embodiments of the present invention provide a method for extracting protein surface features, including...

[0007] A protein surface fingerprint descriptor training dataset is obtained, which includes surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales. The scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and accuracy.

[0008] The protein surface fingerprint descriptor training dataset is used to train a neural network to obtain a protein surface feature extraction model.

[0009] Optionally, obtaining the protein surface fingerprint descriptor training dataset includes:

[0010] Acquire multiple protein molecular structure data;

[0011] The protein surfaces of the acquired protein molecular structure data are meshed to obtain multiple corresponding protein surface meshes.

[0012] Based on the information of the size, shape, and surface complexity of the protein molecular structure data, as well as the computational requirements including computational efficiency and accuracy, the scaling scale of the protein molecular structure data is obtained.

[0013] Based on the scaling scale of the protein molecular structure data, multiple radial surface patches are obtained in the corresponding protein surface grid with each vertex as the center point.

[0014] The chemical properties and geometric structure feature vectors of each radial surface patch, as well as the spatial position information of each radial surface patch, are obtained respectively.

[0015] The chemical properties and geometric feature vectors of each radial surface patch are used as the radial boxes of the corresponding soft pixels, and the spatial position of each radial surface patch is used as the corner box of the corresponding soft pixels. Each radial surface patch is then projected onto a local soft pixel grid.

[0016] Radial bins and corner bins of multiple soft pixels are obtained from the local soft pixel grid to generate the protein surface fingerprint descriptor training dataset.

[0017] Optionally, adjacent radial surface tiles in a plurality of radial surface tiles may partially overlap.

[0018] Optionally, obtaining the chemical properties and geometric features of each of the radial surface patches includes:

[0019] The radial surface patch is divided into multiple overlapping sub-patterns along the radial direction;

[0020] Based on the chemical properties and geometric structure of the vertices within the sub-plot, the feature vector of the sub-plot is obtained;

[0021] Centered on the sub-plot, the sub-plot is expanded in the protein surface grid to obtain a local expansion window. The local expansion window includes the sub-plot and adjacent sub-plots surrounding the sub-plot. The sub-plot is composed of overlapping regions with the adjacent sub-plots surrounding the sub-plot.

[0022] The weight coefficients of each overlapping region within the sub-tile are calculated using a preset first Gaussian kernel within the local expansion window.

[0023] Based on the feature vectors and weight coefficients of each overlapping region, the weighted average feature vector of the overlapping regions within the sub-block is obtained and used as the feature vector of the sub-block.

[0024] The feature vectors of each sub-pattern within the radial surface patch are concatenated to obtain the chemical properties and geometric structure feature vectors of the radial surface patch.

[0025] Optionally, the geometric features include shape index and distance-related curvature;

[0026] The chemical properties include Boltzmann continuous electrostatics, hydrophobicity, and the location of free electrons and proton donors.

[0027] Optionally, the spatial position of the radial surface patch is obtained in the following manner:

[0028] The radial surface patch is multiplied by multiple rotation matrices to obtain multiple corresponding rotation patches;

[0029] The multiple rotated patches are convolved with a preset second Gaussian kernel to obtain multiple corresponding convolutional feature patches.

[0030] The maximum value is extracted from each of the multiple convolutional feature patches to form the geodesic convolution output of the radial surface patch, which is used as the spatial location of the radial surface patch.

[0031] Accordingly, embodiments of the present invention also provide an apparatus for generating a protein surface feature extraction model, comprising:

[0032] The first acquisition unit is adapted to acquire a protein surface fingerprint descriptor training dataset, which includes surface fingerprint descriptor data of multiple protein molecular structure data at a corresponding scaling scale. The scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and computational accuracy.

[0033] The training unit is adapted to train a neural network using the protein surface fingerprint descriptor training dataset to obtain a protein surface feature extraction model.

[0034] Optionally, the first acquisition unit is adapted to acquire multiple protein molecular structure data; to perform meshing processing on the protein surfaces of the acquired multiple protein molecular structure data to obtain multiple corresponding protein surface meshes; to acquire the scaling scale of the protein molecular structure data based on the size, shape, and surface complexity information of the protein molecular structure data, as well as computational requirements including computational efficiency and computational accuracy; to acquire multiple radial surface patches in the corresponding protein surface mesh with each vertex as the center point based on the scaling scale of the protein molecular structure data; to acquire the chemical properties and geometric structure feature vectors of each radial surface patch, as well as the spatial position information of each radial surface patch; to use the chemical properties and geometric structure feature vectors of each radial surface patch as radial boxes of the corresponding soft pixels, and the spatial position of each radial surface patch as corner boxes of the corresponding soft pixels, and to project each radial surface patch onto the local soft pixel mesh; to acquire the radial boxes and corner boxes of multiple soft pixels from the local soft pixel mesh, and to generate the protein surface fingerprint descriptor training dataset.

[0035] Optionally, adjacent radial surface tiles in a plurality of radial surface tiles may partially overlap.

[0036] Optionally, the first acquisition unit is adapted to divide the radial surface patch into multiple overlapping sub-patterns along the radial direction; acquire the feature vector of the sub-pattern based on the chemical property characteristics and geometric structure characteristics of the vertices within the sub-pattern; expand the sub-pattern in the protein surface grid with the sub-pattern as the center to acquire a local expansion window, the local expansion window including the sub-pattern and adjacent sub-patterns surrounding the sub-pattern, the sub-pattern being composed of overlapping regions between the sub-pattern and its adjacent sub-patterns; calculate the weight coefficients of each overlapping region within the sub-pattern using a preset first Gaussian kernel within the local expansion window; acquire the weighted average feature vector of the overlapping regions within the sub-pattern based on the feature vectors and weight coefficients of each overlapping region, as the feature vector of the sub-pattern; and concatenate the feature vectors of each sub-pattern within the radial surface patch to acquire the chemical property and geometric structure feature vectors of the radial surface patch.

[0037] Optionally, the geometric features include shape index and distance-related curvature;

[0038] The chemical properties include Boltzmann continuous electrostatics, hydrophobicity, and the location of free electrons and proton donors.

[0039] Optionally, the first acquisition unit is adapted to multiply the radial surface patch with multiple rotation matrices to obtain multiple corresponding rotation patches; to convolve the multiple rotation patches with a preset second Gaussian kernel to obtain multiple corresponding convolution feature patches; and to extract the maximum value from the multiple convolution feature patches to form the geodesic convolution output of the radial surface patch, which is used as the spatial position of the radial surface patch.

[0040] Accordingly, embodiments of the present invention also provide a method for extracting protein surface features, comprising:

[0041] Obtain the molecular structure data of the protein to be processed;

[0042] The protein surface feature extraction model generated by the method described in any of the preceding claims is used to extract protein surface features from the protein molecular structure data to be processed, thereby obtaining a protein surface fingerprint descriptor of the protein molecular structure data to be processed.

[0043] Accordingly, embodiments of the present invention also provide a protein surface feature extraction device, comprising:

[0044] The second acquisition unit is adapted to acquire the molecular structure data of the protein to be processed.

[0045] The extraction unit is adapted to use a protein surface feature extraction model generated by the method for generating a protein surface feature extraction model as described in any of the preceding claims to extract protein surface features from the protein molecular structure data to be processed, and to obtain a protein surface fingerprint descriptor of the protein molecular structure data to be processed.

[0046] Accordingly, embodiments of the present invention also provide an apparatus, including at least one memory and at least one processor, wherein the memory stores one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method for generating a protein surface feature extraction model as described in any of the preceding claims or the method for extracting protein surface features as described above.

[0047] Accordingly, embodiments of the present invention also provide a storage medium storing one or more computer instructions, the one or more computer instructions being used to implement the method for generating a protein surface feature extraction model as described above or the method for extracting protein surface features as described above.

[0048] Compared with the prior art, the technical solution of the embodiments of the present invention has the following advantages:

[0049] This invention provides a method for generating a protein surface feature extraction model, comprising: acquiring a protein surface fingerprint descriptor training dataset, wherein the protein surface fingerprint descriptor training dataset includes surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales, wherein the scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and computational accuracy; and using the protein surface fingerprint descriptor training dataset to train a neural network to obtain a protein surface feature extraction model.

[0050] In the protein surface feature extraction model generation method provided by this invention, the protein surface fingerprint descriptor training dataset includes surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales. The scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and accuracy. This allows the generated protein surface feature extraction model to extract surface features from input protein molecular structure data with different sizes, shapes, and surface complexities using an appropriate scaling scale, based on computational requirements including computational efficiency and accuracy. This enables the protein surface feature extraction model to achieve multi-scale protein surface feature extraction, meeting diverse user requirements for computational efficiency and accuracy, and helping to improve the accuracy of protein surface feature extraction. Attached Figure Description

[0051] Figure 1 This is a schematic flowchart of an embodiment of the method for generating a protein surface feature extraction model provided by the technical solution of the present invention;

[0052] Figure 2 This is a schematic diagram of the process for obtaining the protein surface fingerprint descriptor training dataset in an embodiment of the present invention;

[0053] Figure 3 This is a schematic diagram of the process of obtaining the protein surface fingerprint descriptor training dataset in an embodiment of the present invention, from obtaining protein molecular structure data to obtaining the chemical properties and geometric structure feature vectors of each radial surface patch;

[0054] Figure 4 This is a schematic diagram of a partially expanded window;

[0055] Figure 5This is a schematic diagram of projecting each of the radial surface patches onto a local soft pixel grid and obtaining radial and corner boxes of multiple soft pixels from the local soft pixel grid to generate the protein surface fingerprint descriptor training dataset, and then training a geometric deep neural network to obtain a protein surface feature extraction model;

[0056] Figure 6 This is a schematic diagram of an embodiment of the apparatus for generating a protein surface feature extraction model provided by the technical solution of the present invention;

[0057] Figure 7 This is a schematic flowchart of the protein surface feature extraction method provided in the embodiments of the present invention;

[0058] Figure 8 This is a schematic diagram of the structure of a protein surface feature extraction device provided in an embodiment of the present invention;

[0059] Figure 9 This is a schematic diagram of an optional hardware structure of a terminal device provided in an embodiment of the present invention. Detailed Implementation

[0060] Currently, methods for extracting protein surface features meet users' different needs for computational efficiency and accuracy.

[0061] To address the technical problem, this invention provides a method for generating a protein surface feature extraction model, comprising: acquiring a protein surface fingerprint descriptor training dataset, wherein the protein surface fingerprint descriptor training dataset includes surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales, wherein the scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and computational accuracy; and using the protein surface fingerprint descriptor training dataset to train a neural network to obtain a protein surface feature extraction model.

[0062] In the protein surface feature extraction model generation method provided by this invention, the protein surface fingerprint descriptor training dataset includes surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales. The scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and accuracy. This allows the generated protein surface feature extraction model to extract surface features from input protein molecular structure data with different sizes, shapes, and surface complexities using an appropriate scaling scale, based on computational requirements including computational efficiency and accuracy. This enables the protein surface feature extraction model to achieve multi-scale protein surface feature extraction, meeting diverse user requirements for computational efficiency and accuracy, and helping to improve the accuracy of protein surface feature extraction.

[0063] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0064] Figure 1 This diagram illustrates a flowchart of an embodiment of the method for generating a protein surface feature extraction model provided by the present invention. See also... Figure 1 A method for generating a protein surface feature extraction model may specifically include the following steps:

[0065] Step S100: Obtain a protein surface fingerprint descriptor training dataset. The protein surface fingerprint descriptor training dataset includes surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales. The scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and computational accuracy.

[0066] Step S200: Use the protein surface fingerprint descriptor training dataset to train a neural network and obtain a protein surface feature extraction model.

[0067] Please continue to refer to this. Figure 1 Execute step S100 to obtain a protein surface fingerprint descriptor training dataset. The protein surface fingerprint descriptor training dataset includes surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales. The scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and computational accuracy.

[0068] Obtaining a protein surface fingerprint descriptor training dataset provides a foundation for subsequent neural network training using the protein surface fingerprint descriptor training dataset to obtain a protein surface feature extraction model.

[0069] Figure 2 A schematic diagram of the process for obtaining a protein surface fingerprint descriptor training dataset is shown in an embodiment of the present invention; Figure 3 This illustration shows a flowchart of the process for obtaining a protein surface fingerprint descriptor training dataset in an embodiment of the present invention, from obtaining protein molecular structure data to obtaining the chemical properties and geometric feature vectors of each radial surface patch. (Refer to the reference...) Figure 2 and Figure 3 The steps for obtaining a training dataset of protein surface fingerprint descriptors may specifically include:

[0070] Step S110: Obtain multiple protein molecular structure data;

[0071] Step S120: The protein surfaces of the acquired multiple protein molecular structure data are respectively processed into meshes to obtain multiple corresponding protein surface meshes;

[0072] Step S130: Based on the information of the size, shape, and surface complexity of the protein molecular structure data, as well as the computational requirements including computational efficiency and accuracy, obtain the scaling scale of the protein molecular structure data;

[0073] Step S140: Based on the scaling scale of the protein molecular structure data, obtain multiple radial surface patches in the corresponding protein surface grid with each vertex as the center point;

[0074] Step S150: Obtain the chemical properties and geometric structure feature vectors of each radial surface patch, as well as the spatial position information of each radial surface patch;

[0075] Step S160: The chemical properties and geometric structure feature vectors of each radial surface patch are used as the radial boxes of the corresponding soft pixels, and the spatial position of each radial surface patch is used as the corner boxes of the corresponding soft pixels. The radial surface patches are then projected onto the local soft pixel grid.

[0076] Step S170: Obtain radial and corner boxes of multiple soft pixels from the local soft pixel grid to generate the protein surface fingerprint descriptor training dataset.

[0077] Please continue to refer to this. Figure 2 and Figure 3 Execute step S110 to obtain multiple protein molecular structure data 10.

[0078] Multiple protein molecular structure data 10 are obtained, which provides a basis for subsequent meshing of the protein surfaces of the multiple protein molecular structure data 10 to obtain the corresponding multiple protein surface meshes 20.

[0079] In some embodiments, the step of obtaining multiple protein molecular structure data 10 includes: obtaining multiple raw protein molecular structure data; preprocessing the multiple raw protein molecular structure data respectively to obtain the multiple protein molecular structure data 10.

[0080] The sources of the multiple raw protein molecular structure data can be set according to actual needs. As an example, multiple raw protein molecular structure data that meet the preset search criteria can be selected from a preset protein crystal database.

[0081] In some embodiments, the protein crystal database includes one or more of the following: the Research Collaboratory for Structural Bioinformatics Protein DataBank (RCSB PDB), the Swiss-Model Repository, the Protein Information Resource (PIR) database, the Structural Classification Protein (SCOP) database, the Protein Structure Initiative Structural Genomics Knowledgebase (PSI-SGKB), the Protein Data Bank Europe (PDBe), and the Critical Assessment of Protein Structure Prediction (CASP) database.

[0082] In some implementations, the search criteria may include at least one of the following, depending on actual needs: type of protein molecule structure, resolution, crystallization conditions, etc.

[0083] In some embodiments, the data format of the raw protein molecular structure data can be a protein database (PDB) file format, a crystallographic information file (CIF) format, etc., and those skilled in the art can set it according to actual needs, without limitation.

[0084] In some embodiments, the preprocessing operations performed on the raw protein molecular structure data include data cleaning, noise reduction, and redundancy removal, so that the preprocessed protein molecular structure data 10 can meet the requirements of subsequent protein surface feature extraction model training.

[0085] In other embodiments, the preprocessing operations performed on the raw protein molecular structure data may also include checking the correctness of the raw protein molecular structure data using a preset format verification tool, and deleting unnecessary atoms such as water molecules and ions from the raw protein molecular structure data, so as to optimize and modify the structure of the raw protein molecular structure data and obtain protein molecular structure data that meets the calculation conditions.

[0086] The number of protein molecular structure data points (10) can be selected based on the actual needs of training the protein surface feature extraction model. It is understandable that, within a preset range, the more protein molecular structure data points (10), the greater the computational load required for subsequent protein surface feature extraction model training, and this relatively helps improve the computational accuracy of the generated protein surface feature extraction model.

[0087] Please continue to refer to this. Figure 2 and Figure 3 In step S120, the protein surfaces of the acquired multiple protein molecular structure data 10 are respectively meshed to obtain multiple corresponding protein surface meshes 20. Among them, the sub-surface mesh 205 is a magnified schematic diagram of a local surface of the multiple protein surface meshes 20.

[0088] The protein surfaces of the acquired multiple protein molecular structure data 10 are respectively meshed to obtain multiple corresponding protein surface meshes 20. This provides a basis for subsequently obtaining the scaling scale of the protein molecular structure data 10 based on information such as the size, shape, and surface complexity of the protein molecular structure data 10, as well as computational requirements including computational efficiency and accuracy. Based on the scaling scale of the protein molecular structure data 10, multiple radial surface patches are obtained in the corresponding protein surface meshes 20 with each vertex as the center point.

[0089] In some embodiments, the step of performing meshing processing on the protein surfaces of the acquired multiple protein molecular structure data 10 to obtain corresponding multiple protein surface meshes 20 includes: protonating the multiple protein molecular structure data 10 to obtain multiple protonated protein molecular structure data; obtaining the solvent-accessible surface area of ​​the multiple protonated protein molecular structure data respectively; and performing triangular discretization processing on the solvent-accessible surface area of ​​the acquired multiple protonated protein molecular structure data to obtain corresponding multiple protein surface meshes 20.

[0090] Protonation of the protein molecular structure data 10 refers to hydrogenating the ligand molecules in the protein molecular structure data 10. As an example, a pre-defined hydrogenation software is used to hydrogenate the ligand molecules in the protein molecular structure data 10. The hydrogenation software can be a simplified (REDUCE) program developed by the Richardson Laboratory at Duke University, etc.

[0091] Solvet accessibility surface area (SASA), also known as access surface area (ASA) or solvent accessibility (SA), refers to the molecular surface area that a solvent can access.

[0092] In some implementations, the solvent-accessible surface area is obtained using a rolling sphere algorithm, and the solvent-accessible surface area is described in square angstroms.

[0093] The density and water probe radius of the protein surface mesh 20 can be set according to actual needs. As an example, the density of the protein surface mesh 20 is 3.0 and the water probe radius is 1.5 angstroms. Of course, the density and water probe radius of the protein surface mesh 20 here are just examples. For different needs, the protein molecular structure data 10 can be meshed into protein surface meshes 20 with different densities and water probe radii, which is not limited here.

[0094] In some embodiments, the step of obtaining a protein surface fingerprint descriptor training dataset after obtaining the corresponding plurality of protein surface grids 20 further includes: performing downsampling processing on the plurality of protein surface grids 20 respectively.

[0095] Downsampling the multiple protein surface grids 20 can reduce the size of the protein surface grids 20, which helps to reduce the computational load of subsequent processing of the protein surface grids 20. Consequently, it can improve the acquisition speed of the protein surface fingerprint descriptor training dataset, and thus improve the generation speed and efficiency of the protein surface feature extraction model.

[0096] In some implementations, the Python geometry processing library (PyMesh) is used to downsample the plurality of protein surface meshes 20 to regularize them to a preset resolution.

[0097] The preset resolution can be set according to actual needs. As an example, the preset resolution is 1.0 to 10.0.

[0098] Please continue to refer to this. Figure 2 and Figure 3 Step S130 is executed, and the scaling scale of the protein molecular structure data 10 is obtained based on the information of the size, shape and surface complexity of the protein molecular structure data 10 and the computational requirements including computational efficiency and computational accuracy.

[0099] Based on the information of the size, shape, and surface complexity of the protein molecular structure data 10, as well as the computational requirements including computational efficiency and accuracy, the scaling scale of the protein molecular structure data 10 is obtained. This provides a basis for subsequently obtaining multiple radial surface patches in the corresponding protein surface grid 20 with each vertex as the center point, according to the scaling scale of the protein molecular structure data 10.

[0100] Based on the size, shape, and surface complexity of the protein molecular structure data 10, as well as computational requirements including computational efficiency and accuracy, a scaling scale for the protein molecular structure data 10 is obtained. Subsequently, using the scaling scale, radial surface patches with geodesic radii corresponding to the scaling scale are obtained in the corresponding protein surface grid 20, with each vertex as the center point. This ensures that the computational efficiency and accuracy of the subsequently trained protein surface feature extraction model can meet the diverse needs of different users for computational efficiency and accuracy when extracting surface features.

[0101] It is understandable that, given a fixed computational efficiency and accuracy in the computational requirements, the larger the protein molecular structure data 10 is, the higher the shape and surface complexity of the protein molecular structure data 10, and the larger the scaling scale of the protein molecular structure data 10 will be; conversely, the smaller the protein molecular structure data 10 is, the lower the shape and surface complexity of the protein molecular structure data 10 will be, and the smaller the scaling scale of the protein molecular structure data 10 will be.

[0102] In some implementations, a preset scaling evaluation model is used to obtain the scaling scale of the protein molecular structure data 10 based on information about the size, shape, and surface complexity of the protein molecular structure data 10, as well as computational requirements including computational efficiency and accuracy.

[0103] Accordingly, a scaling evaluation training dataset can be generated by extracting data on the size, shape, and surface complexity of protein molecular structure data from multiple protein molecular structure data used for evaluation, as well as computational requirements including computational efficiency and accuracy. The scaling evaluation model can then be obtained by training the model using the scaling evaluation training dataset.

[0104] In some embodiments, the volume and mass of the protein molecular structure data 10 are used as indicators of the size of the protein molecular structure data 10, physical measurements such as radius and aspect ratio of the protein molecular structure data 10 are used as indicators of the shape of the protein molecular structure data 10, and the number of concave and convex surfaces and fractal dimensions of the protein surface of the protein molecular structure data 10 are used as indicators of the surface complexity of the protein molecular structure data 10.

[0105] Accordingly, the size, shape, and surface complexity of the protein molecular structure data 10 are each assigned corresponding weights, and a genetic algorithm is used to train the model based on the computational efficiency and accuracy information in the computational requirements, thereby obtaining the scaling evaluation model.

[0106] Please continue to refer to this. Figure 2 and Figure 3 In step S140, based on the scaling scale of the protein molecular structure data 10, multiple radial surface patches are obtained in the corresponding protein surface grid 20 with each vertex as the center point.

[0107] Based on the scaling scale of the protein molecular structure data 10, multiple radial surface patches are obtained in the corresponding protein surface grid 20 with each vertex as the center point, in preparation for the subsequent acquisition of the chemical properties and geometric structure feature vectors of each radial surface patch, as well as the spatial position information of each radial surface patch.

[0108] According to the scaling scale of the protein molecular structure data 10, multiple radial surface patches are obtained in the corresponding protein surface grid 20 with each vertex as the center point. This means that, according to the scaling scale of the protein molecular structure data 10, multiple radial surface patches with corresponding geodesic radii are obtained in the corresponding protein surface grid 20 with each vertex as the center point.

[0109] In other words, the geodesic radius of multiple radial surface patches obtained in the protein surface grid 20 with each vertex as the center point is related to the scaling scale of the corresponding protein molecular structure data 10.

[0110] Specifically, for the same protein molecular structure data 10, the larger the scaling scale of the protein molecular structure data 10, the larger the geodesic radius of the multiple radial surface patches obtained in the corresponding protein surface grid 20 with each vertex as the center point; the smaller the scaling scale of the protein molecular structure data 10, the smaller the geodesic radius of the multiple radial surface patches obtained in the corresponding protein surface grid 20 with each vertex as the center point.

[0111] In some implementations, adjacent radial surface patches in a plurality of radial surface patches partially overlap, such that the plurality of radial surface patches can cover all regions of the protein surface grid 20.

[0112] Please continue to refer to this. Figure 2 and Figure 3 Step S150 is executed to obtain the chemical properties and geometric structure feature vectors of each radial surface patch, as well as the spatial position information of each radial surface patch.

[0113] The chemical properties and geometric structure feature vectors of each radial surface patch, as well as the spatial position information of each radial surface patch, are obtained respectively. This is to prepare for the subsequent projection of each radial surface patch onto the local soft pixel grid by using the chemical properties and geometric structure feature vectors of each radial surface patch as the radial bins of the corresponding soft pixels and the spatial position of each radial surface patch as the corner bins of the corresponding soft pixels.

[0114] In some embodiments, the step of obtaining the chemical property and geometric feature vectors of each of the radial surface patches includes: dividing the radial surface patch into multiple overlapping sub-patterns along the radial direction; obtaining the feature vector of the sub-pattern based on the chemical property and geometric feature of the vertices within the sub-pattern; expanding the sub-pattern in the protein surface grid 20 with the sub-pattern as the center to obtain a local expansion window, the local expansion window including the sub-pattern and adjacent sub-patterns surrounding the sub-pattern, the sub-pattern being composed of overlapping regions between the sub-pattern and its adjacent sub-patterns; calculating the weight coefficients of each overlapping region within the sub-pattern using a preset first Gaussian kernel within the local expansion window; obtaining the weighted average feature vector of the overlapping regions within the sub-pattern based on the feature vectors and weight coefficients of each overlapping region, as the feature vector of the sub-pattern; and concatenating the feature vectors of each sub-pattern within the radial surface patch to obtain the chemical property and geometric feature vector of the radial surface patch.

[0115] The radial surface patch is divided into multiple overlapping sub-patterns along the radial direction, that is, the radial surface patch is divided into multiple overlapping sub-patterns along the direction of the geodesic radius.

[0116] In some embodiments, the step of obtaining the feature vector of the sub-block based on the chemical property characteristics and geometric structure characteristics of the vertices within the sub-block includes: obtaining the chemical property characteristics and geometric structure characteristics of each vertex within the sub-block to form the feature vector of each vertex; and obtaining the feature vector of the sub-block based on the feature vector of each vertex within the sub-block.

[0117] In some implementations, corresponding weights can be assigned to the vertices within the sub-plot, and then the feature vectors of each vertex within the sub-plot can be multiplied by their respective weights and summed to obtain the feature vector of the sub-plot. In other words, the weighted average of the feature vectors of each vertex within the sub-plot is obtained as the feature vector of the sub-plot.

[0118] In other embodiments, the average value of the feature vectors of each vertex within the sub-plot can also be used as the feature vector of the sub-plot. Those skilled in the art can set this according to actual needs, and there are no restrictions here.

[0119] In some embodiments, the geometric features of vertices within the sub-block include shape indexes and distance-dependent curvatures, and the chemical features of vertices within the sub-block include Poisson-Boltzmann continuous electrostatics, hydrophobicity, and the location of free electron and proton donors.

[0120] The shape index is a scalar value defined on each vertex of the protein surface mesh, used to describe the curvature variation information near that vertex.

[0121] In some implementations, the curvature tensor matrix of each vertex is calculated based on the curvature tensor, and then the corresponding shape index value is obtained by performing eigenvalue analysis on the curvature tensor matrix.

[0122] In some implementations, the formula used to calculate the shape index of each vertex within the sub-tile is as follows:

[0123]

[0124] Where I represents the shape index of the vertex, λ1 represents the curvature feature value of the curvature tensor along the preset first principal direction, and λ2 represents the curvature feature value of the curvature tensor along the preset second principal direction. The curvature tensor is calculated from the local curvature at the vertex position and is used to describe the curvature variation information in different directions.

[0125] As can be seen from the above formula (1), when the curvature at the vertex is the same, the value of the shape index I of the vertex is zero; when the curvature at the vertex is different, the value of the shape index I of the vertex is positive.

[0126] Correspondingly, the more complex the shape at the vertex position, the larger the shape index value I of the vertex; the simpler the shape at the vertex position, the smaller the shape index value I of the vertex.

[0127] In some implementations, the shape index value I ranges from -1 to +1. Specifically, a shape index value of -1 indicates that the vertex is a concave point on the protein surface, and a shape index value of +1 indicates that the vertex is a convex point on the protein surface.

[0128] Distance-dependent curvature is a method that correlates curvature with the distance between molecules, enabling a more accurate description of the local geometric features of the surface of protein molecules. Specifically, the distance-dependent curvature of a vertex is obtained by calculating the distance between each vertex on the protein surface grid 20 and its nearest neighbor, and then using the calculated distance as a weight to adjust the curvature of that vertex.

[0129] In some embodiments, the step of obtaining Boltzmann continuous electrostatics includes: obtaining protein molecular structure data 10 in protein database (PDB) file format and charge parameter information in protein molecular structure data 10; converting protein molecular structure data 10 in protein database (PDB) file format and charge parameter information in protein molecular structure data 10 into protein quantitative reporter gene (PQR) format, obtaining information on atomic coordinates and predicted charge values ​​of each vertex in protein molecular structure data 10; and assigning corresponding charges to each vertex in protein surface grid 20 using surface charge values ​​(multivalues) provided in the Adaptive Poisson-Boltzmann solver (APBS) suite for solving continuous electrostatic equations for large biomolecule assemblies, as the Boltzmann continuous electrostatic value of each vertex in protein surface grid 20.

[0130] In some embodiments, the steps of obtaining the positions of free electrons and proton donors include: calculating the positions of free electrons, proton donor-free electrons, and potential hydrogen bond donors on the surface of the protein molecular structure using hydrogen bond potentials, wherein polar hydrogen, nitrogen, or oxygen atoms on the molecular surface closest to the protein molecular structure data 10 are considered as potential hydrogen bond donors or acceptors; and assigning corresponding Gaussian distribution values ​​to each vertex in the protein surface grid 20 based on the orientation between heavy atoms, according to the positions of free electrons, proton donor-free electrons, and potential hydrogen bond donors on the surface of the protein molecular structure, as the positions of free electrons and proton donors at that vertex.

[0131] In some implementations, the corresponding charge values ​​assigned to each vertex in the protein surface mesh 20 range from -1 to +1. Specifically, each vertex in the protein surface mesh 20 is first assigned an initial charge value, wherein the range of the assigned initial charge value is -30 to +30. Then, the initial charge values ​​assigned to each vertex in the protein surface mesh 20 are normalized so that the range of the charge value of each vertex in the protein surface mesh 20 is -1 to +1.

[0132] In some embodiments, the values ​​of the free electron and proton donor positions at each vertex of the protein surface grid 20 range from -1 to +1. Specifically, a value of -1 indicates that the vertex corresponds to the optimal position of a hydrogen bond receiver, while a value of +1 indicates that the vertex corresponds to the optimal position of a hydrogen bond provider.

[0133] Hydrophobicity is important for proteins in their biological functions of interaction.

[0134] In some implementations, the step of obtaining the hydrophobicity of each vertex within a sub-plot includes: calculating the Kyte & Doolittle scale level of each amino acid identity in the amino acid sequence of the protein molecular structure data 10 using the Kyte & Doolittle calculation method; and assigning a corresponding hydrophobicity value to each vertex in the protein surface grid 20 based on the Kyte & Doolittle scale level of the amino acid identity of the nearest atom.

[0135] In some implementations, when assigning a corresponding hydrophobicity scalar value to each vertex of the protein surface grid 20 based on the Kyte & Doolittle scale of the amino acid identity of the nearest atom, the initial hydrophobicity scalar value is first assigned to each vertex of the protein surface grid 20 based on the Kyte & Doolittle scale of the amino acid identity of the nearest atom. Then, the initial hydrophobicity scalar values ​​assigned to each vertex of the protein surface grid 20 are normalized to obtain the hydrophobicity value of each vertex in the protein surface grid 20. The initial hydrophobicity scalar value of each vertex ranges from -4.5 to +4.5 on the original scale. An initial hydrophobicity scalar value of -4.5 indicates that the vertex is hydrophilic, and an initial hydrophobicity scalar value of +4.5 indicates that the vertex is the most hydrophobic.

[0136] Centered on the sub-plot, the sub-plot is expanded within the protein surface grid 20 to obtain a local expansion window. A preset first Gaussian kernel is used to calculate the weight coefficients of each overlapping region within the local expansion window. Then, based on the feature vectors and weight coefficients of each overlapping region, the weighted average feature vector of the overlapping regions within the sub-plot is obtained as the feature vector of the sub-plot. This results in the chemical properties and geometric structure feature vector of the radial surface patch obtained by splicing the feature vectors of each sub-plot within the radial surface patch containing information about the surface features of each sub-plot and its surrounding area.

[0137] For clarity, let's take a rectangular sub-block as an example. See [link / reference] Figure 4 The local expansion window obtained by expanding the sub-plot 401 in the protein surface grid 20 with the sub-plot 401 as the center includes the sub-plot 401 and the adjacent upper left sub-plot 402, upper right sub-plot 403, lower left sub-plot 404 and lower right sub-plot 405.

[0138] Accordingly, sub-block 401 includes an overlapping area 4011 with the adjacent upper left sub-block 402, an overlapping area 4012 with the adjacent upper right sub-block 403, an overlapping area 4013 with the adjacent lower left sub-block 404, and an overlapping area 4014 with the adjacent lower right sub-block 405.

[0139] The above description uses the example of a local expanded window including a sub-tile located at the center and its four adjacent sub-tiles. The number of sub-tiles adjacent to the sub-tile located at the center in the local expanded window, as well as the overlapping area between the adjacent sub-tiles and the sub-tile located at the center, can be set as needed and are not limited here.

[0140] The first Gaussian kernel is a learnable Gaussian kernel, or a learned soft polar coordinate grid, which is used to locally average the features of radial surface patches in the vertex direction and produce a fixed-size output associated with a set of learnable filters.

[0141] In some implementations, the step of obtaining the first Gaussian kernel includes: obtaining a Gaussian kernel function as a basis function of a local geodesic system; assigning corresponding geodesic coordinate positions to vertices in the protein surface grid 20 obtained by meshing the protein molecular structure data 10, wherein the geodesic coordinate positions include information about the radial distance of the vertex; using the radial distance of each vertex in the protein surface grid 20 as a learning parameter, and representing the learning parameter using a learnable weight vector or a neural network; training the Gaussian kernel function using a preset number of training datasets to obtain the first Gaussian kernel.

[0142] In some implementations, the weight coefficients of overlapping regions within each sub-block are normalized values, such that the sum of the weight coefficients of overlapping regions within each sub-block is 1.

[0143] In some embodiments, the step of obtaining the spatial position of the radial surface patch includes: multiplying the radial surface patch with multiple rotation matrices to obtain multiple corresponding rotation patches; convolving the multiple rotation patches with a preset second Gaussian kernel to obtain multiple corresponding convolution feature patches; and extracting the maximum value from the multiple convolution feature patches to form the geodesic convolution output of the radial surface patch, which is used as the spatial position of the radial surface patch.

[0144] Each vertex in the radial surface patch has a corresponding geodesic polar coordinate position, which includes the radial and angular coordinates of the vertex. The radial coordinate represents the geodesic distance between each vertex in the radial surface patch and the center vertex of the radial surface patch, and the angular coordinate represents the angle between each vertex in the radial surface patch and a random direction.

[0145] Each vertex in the radial surface patch has a corresponding geodesic polar coordinate position, thereby adding information about the spatial positional relationship between protein surface features to the trained protein surface feature extraction model.

[0146] In some implementations, the Dijkstra algorithm in the mathematical software Matlab produced by MathWorks is used to calculate an approximate value of the actual geodesic distance, which is then used as the geodesic distance between each vertex in the radial surface patch and the vertex that serves as the center of the radial surface patch.

[0147] In some implementations, the radial surface tiles do not have a canonical orientation, so a random direction in the computation plane is selected as a reference direction, and the angle between each vertex in the computation plane and the reference direction is set as angular coordinates.

[0148] By assigning corresponding geodesic polar coordinate positions to the vertices in the radial surface patch, the protein surface mesh 20 can be tiled on a computational plane, and the position of each vertex can be represented in angular coordinates.

[0149] Since angular coordinates are calculated relative to random directions, it is necessary to calculate information that remains unchanged across different directions. To do this, multiple rotation operations are performed on the radial surface patch, and the maximum value of all rotations is calculated to generate a geodesic convolution output of the radial surface patch, which serves as the spatial position of the radial surface patch.

[0150] Figure 5 This diagram illustrates how radial and corner boxes are used to generate a protein surface fingerprint descriptor training dataset by projecting each radial surface patch onto a local soft pixel grid and obtaining multiple soft pixels from the local soft pixel grid. This dataset is then used to train a geometric deep neural network to obtain a protein surface feature extraction model. Please refer to the references. Figures 2 to 5 In step S160, the chemical properties and geometric feature vectors of each radial surface patch are used as the radial boxes of the corresponding soft pixels, and the spatial position of each radial surface patch is used as the corner boxes of the corresponding soft pixels. Each radial surface patch is then projected onto the local soft pixel grid.

[0151] The chemical properties and geometric structure feature vectors of each radial surface patch are used as radial bins for the corresponding soft pixels, and the spatial position of each radial surface patch is used as the corner bin for the corresponding soft pixels. Each radial surface patch is projected onto a local soft pixel grid, which provides a basis for obtaining radial bins and corner bins of multiple soft pixels from the local soft pixel grid to generate the protein surface fingerprint descriptor training dataset.

[0152] In the manner described above, based on the chemical properties, geometric feature vectors, and spatial positions of the radial surface patches, each radial surface patch can be treated as a soft pixel with radial and corner boxes, thereby projecting the entire protein molecular structure data 10 onto a local soft pixel grid.

[0153] Please continue to refer to this. Figures 2 to 5 Step S170 is executed to obtain radial boxes and corner boxes of multiple soft pixels from the local soft pixel grid, and generate the protein surface fingerprint descriptor training dataset.

[0154] Radial and corner bins of multiple soft pixels are obtained from the local soft pixel grid to generate the protein surface fingerprint descriptor training dataset, which provides a foundation for subsequent neural network training using the protein surface fingerprint descriptor training dataset to obtain a protein surface feature extraction model.

[0155] In some implementations, a preset sliding window is used to acquire corresponding soft pixel blocks on the local soft pixel grid with a preset step size. The soft pixel block includes multiple soft pixels. The radial and corner bins of the multiple soft pixels located within the sliding window are acquired respectively. The radial and corner bins of each soft pixel are used as training data for a protein surface fingerprint descriptor. Thus, a corresponding number of protein surface fingerprint descriptor training data that meet the training requirements are acquired on the local soft pixel grid, forming the protein surface fingerprint descriptor training dataset.

[0156] The sliding window and the step size can be set according to actual needs. In some embodiments, the step size of the sliding window on the local soft pixel grid is smaller than the length of the sliding window, so that the soft pixel blocks in adjacent sliding windows partially overlap.

[0157] Please continue to refer to this. Figure 1 Step S200 is executed, using the protein surface fingerprint descriptor training dataset to train a neural network and obtain a protein surface feature extraction model.

[0158] The protein surface fingerprint descriptor training dataset is used to train a neural network to obtain a protein surface feature extraction model. This model can extract surface features from the protein molecular structure data to be processed according to the user's different needs for computational efficiency and accuracy, thereby meeting the user's diverse needs for computational efficiency and accuracy.

[0159] In some implementations, the protein surface fingerprint descriptor training dataset is used to train a geometric deep neural network to obtain a protein surface feature extraction model.

[0160] In other embodiments, other model training methods can be used to train the protein surface feature extraction model. Those skilled in the art can choose according to actual needs, and no restrictions are imposed here.

[0161] The protein surface feature extraction model generation method in this embodiment of the invention includes a protein surface fingerprint descriptor training dataset comprising surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales. The scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and accuracy. This allows the generated protein surface feature extraction model to extract surface features from input protein molecular structure data of different sizes, shapes, and surface complexities using an appropriate scaling scale, based on computational requirements including computational efficiency and accuracy. This enables multi-scale protein surface feature extraction, meeting diverse user requirements for computational efficiency and accuracy, and contributing to improved accuracy in protein surface feature extraction.

[0162] The protein surface feature extraction model generation method in this embodiment of the invention learns and updates the neighborhood feature information between radial surface patches and the self-feature information of the center of the radial surface patch in the protein molecular structure data at different scaling scales in a coordinated manner, which can effectively improve the robustness and confidence of the protein surface feature extraction model.

[0163] The method for generating a white matter surface feature extraction model in this embodiment of the invention aggregates neighborhood feature information from multiple scaling scales of protein molecular structure data, resulting in excellent anti-interference capabilities. Even if the surface feature dataset of the protein molecular structure data to be processed contains missing data, noise, or unclear information, it does not affect the effectiveness of the generated protein surface feature extraction model. Furthermore, the protein surface feature extraction model generated by the method in this embodiment of the invention is insensitive to hyperparameters; even different parameter combinations will not affect the protein surface fingerprint descriptor results output by the protein surface feature extraction model.

[0164] Accordingly, embodiments of the present invention also provide an apparatus for generating a protein surface feature extraction model.

[0165] Figure 6 This diagram illustrates a structural schematic of an embodiment of the apparatus for generating a protein surface feature extraction model provided by the present invention. See also... Figure 6 A protein surface feature extraction model generation device 60 includes: a first acquisition unit 601, adapted to acquire a protein surface fingerprint descriptor training dataset, wherein the protein surface fingerprint descriptor training dataset includes surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales, wherein the scaling scale is related to the size, shape, and surface complexity information of the protein molecular structure data, as well as computational requirements including computational efficiency and computational accuracy; and a training unit 602, adapted to perform neural network training using the protein surface fingerprint descriptor training dataset to acquire a protein surface feature extraction model.

[0166] The protein surface feature extraction model generation device in this embodiment of the invention can be used to execute the aforementioned protein surface feature extraction model generation method to generate a protein surface feature extraction model for extracting surface features from protein molecular structure data.

[0167] It is understood that other functional structures can also be used to execute the aforementioned method for generating protein surface feature extraction models, and no restrictions are imposed here. For details regarding the apparatus for generating protein surface feature extraction models, please refer to the relevant content regarding the method for generating protein surface feature extraction models mentioned above; it will not be repeated here.

[0168] Accordingly, embodiments of the present invention also provide a method for extracting protein surface features.

[0169] Figure 7 A schematic flowchart of a method for extracting protein surface features according to an embodiment of the present invention is shown. See also Figure 7 A method for extracting protein surface features, specifically including:

[0170] Step S710: Obtain the molecular structure data of the protein to be processed;

[0171] Step S720: Using the protein surface feature extraction model generated by the method described above, protein surface feature extraction is performed on the protein molecular structure data to be processed to obtain the protein surface fingerprint descriptor of the protein molecular structure data to be processed.

[0172] In some embodiments, the protein molecular structure data to be processed is the protein molecular structure data for which protein surface feature extraction is to be performed.

[0173] Accordingly, the protein surface feature extraction model generated by the method described above is used to extract surface features from the protein molecular structure data to be processed, thereby generating a surface fingerprint feature descriptor for the protein molecular structure data. For details on the method for generating the protein surface feature extraction model, please refer to the corresponding description in the preceding sections; it will not be repeated here.

[0174] Accordingly, embodiments of the present invention also provide an apparatus for extracting protein surface features.

[0175] Figure 8 A schematic diagram of a protein surface feature extraction device provided in an embodiment of the present invention is shown. See also Figure 8 A protein surface feature extraction device 80 may specifically include: a second acquisition unit 801, adapted to acquire protein molecular structure data to be processed; and an extraction unit 802, adapted to use a protein surface feature extraction model generated by the protein surface feature extraction model generation method to extract protein surface features from the protein molecular structure data to be processed, and obtain a protein surface fingerprint descriptor of the protein molecular structure data to be processed.

[0176] The protein surface feature extraction device in this embodiment can be used to perform the aforementioned protein surface feature extraction method, or other functional structures can be used to perform the aforementioned protein surface feature extraction method. For details regarding the protein surface feature extraction device, please refer to the aforementioned content on the protein surface feature extraction method; it will not be repeated here.

[0177] This invention also provides an apparatus that can execute the above-described method for generating a protein surface feature extraction model or a method for extracting protein surface features by loading a program.

[0178] An optional hardware structure for the terminal device provided in this embodiment of the invention can be as follows: Figure 9 As shown, it includes: at least one processor 01, at least one communication interface 02, at least one memory 03, and at least one communication bus 04.

[0179] In this embodiment of the invention, the number of processor 01, communication interface 02, memory 03 and communication bus 04 is at least one, and processor 01, communication interface 02 and memory 03 communicate with each other through communication bus 04.

[0180] Communication interface 02 can be an interface for a communication module used for network communication, such as the interface of a GSM module.

[0181] Processor 01 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0182] Memory 03 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0183] The memory 03 stores one or more computer instructions, which are executed by the processor 01 to implement the protein surface feature extraction model generation method or protein surface feature extraction method of the embodiments of the present invention.

[0184] It should be noted that the aforementioned terminal device may also include other devices (not shown) that may not be essential to understanding the content disclosed in the embodiments of the present invention; given that these other devices may not be essential for understanding the content disclosed in the embodiments of the present invention, the embodiments of the present invention will not describe them one by one.

[0185] This invention also provides a storage medium storing one or more computer instructions, which are used to implement the protein surface feature extraction model generation method or protein surface feature extraction method provided in this invention.

[0186] The embodiments of the present invention described above are combinations of elements and features of the present invention. Unless otherwise stated, the elements or features described are optional. Individual elements or features may be practiced without combination with other elements or features. Furthermore, embodiments of the present invention may be constructed by combining some elements and / or features. The order of operations described in the embodiments of the present invention may be rearranged. Some constructions of any embodiment may be included in another embodiment and may be replaced by corresponding constructions of another embodiment. It will be apparent to those skilled in the art that claims in the appended claims that are not expressly referenced to each other may be combined to form embodiments of the present invention, or may be included as new claims in amendments made after the filing of this application.

[0187] Embodiments of the present invention can be implemented by various means, such as hardware, firmware, software, or combinations thereof. In a hardware configuration, the method according to an exemplary embodiment of the present invention can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, etc.

[0188] In firmware or software configuration, embodiments of the present invention can be implemented in the form of modules, processes, functions, etc. Software code can be stored in a memory unit and executed by a processor. The memory unit is located inside or outside the processor and can send data to and receive data from the processor via various known means.

[0189] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is accorded the widest scope consistent with the principles and novel features disclosed herein.

[0190] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.

Claims

1. A method for generating a protein surface feature extraction model, characterized in that, include: A protein surface fingerprint descriptor training dataset is obtained, comprising surface fingerprint descriptor data of multiple protein molecular structure data at corresponding scaling scales. The scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and accuracy. The process includes: acquiring multiple protein molecular structure data; meshing the protein surfaces of the acquired multiple protein molecular structure data to obtain corresponding multiple protein surface meshes; and, based on the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and accuracy, obtaining the protein molecular structure data... The scaling scale is determined; based on the scaling scale of the protein molecular structure data, multiple radial surface patches are obtained in the corresponding protein surface grid with each vertex as the center point; the chemical properties and geometric structure feature vectors of each radial surface patch, as well as the spatial position information of each radial surface patch, are obtained; the chemical properties and geometric structure feature vectors of each radial surface patch are used as radial boxes of the corresponding soft pixels, and the spatial position of each radial surface patch is used as corner boxes of the corresponding soft pixels; each radial surface patch is projected onto the local soft pixel grid; multiple radial boxes and corner boxes of soft pixels are obtained from the local soft pixel grid to generate the protein surface fingerprint descriptor training dataset; The protein surface fingerprint descriptor training dataset is used to train a neural network to obtain a protein surface feature extraction model.

2. The method for generating a protein surface feature extraction model as described in claim 1, characterized in that, There is partial overlap between adjacent radial surface tiles in multiple radial surface tiles.

3. The method for generating a protein surface feature extraction model as described in claim 1, characterized in that, The acquisition of the chemical properties and geometric features of each of the radial surface patches includes: The radial surface patch is divided into multiple overlapping sub-patterns along the radial direction; Based on the chemical properties and geometric structure of the vertices within the sub-plot, the feature vector of the sub-plot is obtained; Centered on the sub-plot, the sub-plot is expanded in the protein surface grid to obtain a local expansion window. The local expansion window includes the sub-plot and adjacent sub-plots surrounding the sub-plot. The sub-plot is composed of overlapping regions with the adjacent sub-plots surrounding the sub-plot. The weight coefficients of each overlapping region within the sub-tile are calculated using a preset first Gaussian kernel within the local expansion window. Based on the feature vectors and weight coefficients of each overlapping region, the weighted average feature vector of the overlapping regions within the sub-block is obtained and used as the feature vector of the sub-block. The feature vectors of each sub-pattern within the radial surface patch are concatenated to obtain the chemical properties and geometric structure feature vectors of the radial surface patch.

4. The method for generating a protein surface feature extraction model as described in claim 3, characterized in that, The geometric features include shape index and distance-related curvature; The chemical properties include Boltzmann continuous electrostatics, hydrophobicity, and the location of free electrons and proton donors.

5. The method for generating a protein surface feature extraction model as described in claim 1, characterized in that, The spatial position of the radial surface patch is obtained in the following way: The radial surface patch is multiplied by multiple rotation matrices to obtain multiple corresponding rotation patches; The multiple rotated patches are convolved with a preset second Gaussian kernel to obtain multiple corresponding convolutional feature patches. The maximum value is extracted from each of the multiple convolutional feature patches to form the geodesic convolution output of the radial surface patch, which is used as the spatial location of the radial surface patch.

6. A device for generating a protein surface feature extraction model, characterized in that, include: The first acquisition unit is adapted to acquire a protein surface fingerprint descriptor training dataset. This dataset includes surface fingerprint descriptor data for multiple protein molecular structure data at corresponding scaling scales. The scaling scale is related to the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and accuracy. The acquisition unit includes: acquiring multiple protein molecular structure data; performing meshing processing on the protein surfaces of the acquired multiple protein molecular structure data to obtain corresponding multiple protein surface meshes; and, based on the size, shape, and surface complexity of the protein molecular structure data, as well as computational requirements including computational efficiency and accuracy, acquiring the protein surface fingerprint descriptor data. The scaling scale of the structural data; based on the scaling scale of the protein molecular structure data, multiple radial surface patches are obtained in the corresponding protein surface grid with each vertex as the center point; the chemical properties and geometric structure feature vectors of each radial surface patch, as well as the spatial position information of each radial surface patch, are obtained; the chemical properties and geometric structure feature vectors of each radial surface patch are used as radial boxes of the corresponding soft pixels, and the spatial position of each radial surface patch is used as corner boxes of the corresponding soft pixels; each radial surface patch is projected onto the local soft pixel grid; multiple radial boxes and corner boxes of soft pixels are obtained from the local soft pixel grid to generate the protein surface fingerprint descriptor training dataset; The training unit is adapted to train a neural network using the protein surface fingerprint descriptor training dataset to obtain a protein surface feature extraction model.

7. The apparatus for generating a protein surface feature extraction model as described in claim 6, characterized in that, There is partial overlap between adjacent radial surface tiles in multiple radial surface tiles.

8. The apparatus for generating a protein surface feature extraction model as described in claim 6, characterized in that, The first acquisition unit is adapted to divide the radial surface patch into multiple overlapping sub-patterns along the radial direction; acquire the feature vector of the sub-pattern based on the chemical property characteristics and geometric structure characteristics of the vertices within the sub-pattern; expand the sub-pattern in the protein surface grid with the sub-pattern as the center to acquire a local expansion window, the local expansion window including the sub-pattern and adjacent sub-patterns surrounding the sub-pattern, the sub-pattern being composed of overlapping regions between the sub-pattern and its adjacent sub-patterns; calculate the weight coefficients of each overlapping region within the sub-pattern using a preset first Gaussian kernel within the local expansion window; acquire the weighted average feature vector of the overlapping regions within the sub-pattern based on the feature vectors and weight coefficients of each overlapping region, as the feature vector of the sub-pattern; and concatenate the feature vectors of each sub-pattern within the radial surface patch to acquire the chemical property and geometric structure feature vectors of the radial surface patch.

9. The apparatus for generating a protein surface feature extraction model as described in claim 8, characterized in that, The geometric features include shape index and distance-related curvature; The chemical properties include Boltzmann continuous electrostatics, hydrophobicity, and the location of free electrons and proton donors.

10. The apparatus for generating a protein surface feature extraction model as described in claim 6, characterized in that, The first acquisition unit is adapted to multiply the radial surface patch with multiple rotation matrices to obtain multiple corresponding rotation patches; to convolve the multiple rotation patches with a preset second Gaussian kernel to obtain multiple corresponding convolution feature patches; and to extract the maximum value from the multiple convolution feature patches to form the geodesic convolution output of the radial surface patch, which is used as the spatial position of the radial surface patch.

11. A method for extracting protein surface features, characterized in that, include: Obtain the molecular structure data of the protein to be processed; The protein surface feature extraction model generated by the method described in any one of claims 1 to 5 is used to extract protein surface features from the protein molecular structure data to be processed, thereby obtaining the protein surface fingerprint descriptor of the protein molecular structure data to be processed.

12. A device for extracting protein surface features, characterized in that, include: The second acquisition unit is adapted to acquire the molecular structure data of the protein to be processed. The extraction unit is adapted to use the protein surface feature extraction model generated by the method for generating the protein surface feature extraction model as described in any one of claims 1 to 5 to extract protein surface features from the protein molecular structure data to be processed, and obtain the protein surface fingerprint descriptor of the protein molecular structure data to be processed.

13. A device, characterized in that, It includes at least one memory and at least one processor, the memory storing one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method for generating a protein surface feature extraction model as claimed in any one of claims 1 to 5 or the method for extracting protein surface features as claimed in claim 11.

14. A storage medium, characterized in that, The storage medium stores one or more computer instructions, which are used to implement the method for generating the protein surface feature extraction model as described in any one of claims 1-5 or the method for extracting protein surface features as described in claim 11.

Citation Information

Patent Citations

  • Protein image classification method and device, equipment and medium

    CN111242922A

  • Protein immunogenicity classifier construction method, prediction method, device and medium

    CN115064217A