Protein solubility prediction method and device, electronic equipment, computer readable storage medium and computer program product

The protein structure graph features are extracted through deep learning models and graph neural networks, and the problem of time-consuming and inaccurate protein solubility prediction is solved, achieving more efficient and accurate prediction of solubility changes, supporting drug design and disease research.

CN120581094APending Publication Date: 2025-09-02TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410239467.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-01
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The prediction methods for protein solubility in the prior art are time-consuming and labor-intensive and not accurate enough, making it difficult to effectively evaluate the impact of protein denaturation on drug effectiveness and safety.

Method used

By obtaining the structural diagram of the original and mutant proteins, using deep learning models such as Alphafold to predict the three-dimensional structure of the protein, combining with the graph neural network to extract amino acid characteristics, predict protein solubility changes, and improve the fine-grainedness and accuracy of the prediction.

Benefits of technology

The prediction method based on protein structure chart improves the accuracy of protein solubility changes, saves computing resources, enhances the interpretability and fine-grainedness of predictions, and can better evaluate protein functional changes in drug design and disease research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581094A_ABST
    Figure CN120581094A_ABST
Patent Text Reader

Abstract

The invention provides a protein solubility prediction method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring a first structural diagram of original protein and a second structural diagram of mutant protein, and performing feature extraction processing on the first structural diagram and the second structural diagram to obtain a plurality of first amino acid features and second amino acid features corresponding to the first amino acid features respectively; and performing solubility prediction processing based on each first amino acid feature and the corresponding second amino acid feature to obtain a sub prediction change value corresponding to each first amino acid feature, and determining a solubility change value between the original protein and the mutant protein according to each sub prediction change value. According to the invention, the accuracy of predicting the protein solubility change can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to artificial intelligence technology, and in particular to a method, device, electronic device, computer-readable storage medium, and computer program product for predicting protein solubility. Background Art

[0002] The amount or level of protein dispersion in a liquid is called the solubility of the protein. Solubility characteristic data needs to be applied in the extraction, separation, and purification of natural proteins, and the degree of protein denaturation can also be evaluated by changes in the protein's solubility behavior.

[0003] Protein solubility may affect its distribution, absorption, metabolism, and excretion in the body. This information is crucial for drug efficacy and safety. In related technologies, drug water solubility studies are primarily conducted through physical and chemical experimental methods, including spectroscopy, chromatography, or X-ray crystallography. These experimental methods are often time-consuming and labor-intensive, and there is currently no effective method to improve the accuracy of predicting changes in protein solubility. Summary of the Invention

[0004] The embodiments of the present application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for predicting protein solubility, which can improve the accuracy of predicting changes in protein solubility.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present invention provides a method for predicting protein solubility, comprising:

[0007] Obtaining a first structure diagram of an original protein and a second structure diagram of a mutant protein, wherein the mutant protein is obtained based on a mutation of the original protein, the first structure diagram includes a plurality of first amino acid nodes, the second structure diagram includes a plurality of second amino acid nodes, and different first amino acid nodes correspond to different second amino acid nodes;

[0008] performing feature extraction processing on the first structure graph to obtain a plurality of first amino acid features;

[0009] performing feature extraction processing on the second structure graph to obtain second amino acid features corresponding to each of the first amino acid features;

[0010] performing solubility prediction processing based on each of the first amino acid features and the corresponding second amino acid features to obtain a sub-prediction change value corresponding to each of the first amino acid features, wherein the sub-prediction change value represents a solubility change between the first amino acid feature and the corresponding second amino acid feature;

[0011] Based on each of the sub-predicted change values, the solubility change value between the original protein and the mutant protein is determined.

[0012] The present invention provides a device for predicting protein solubility, comprising:

[0013] a data acquisition module, configured to acquire a first structure diagram of an original protein and a second structure diagram of a mutant protein, wherein the mutant protein is obtained based on a mutation of the original protein, the first structure diagram includes a plurality of first amino acid nodes, the second structure diagram includes a plurality of second amino acid nodes, and different first amino acid nodes correspond to different second amino acid nodes;

[0014] a feature extraction module, configured to perform feature extraction processing on the first structure graph to obtain a plurality of first amino acid features; and perform feature extraction processing on the second structure graph to obtain a second amino acid feature corresponding to each of the first amino acid features;

[0015] a prediction processing module, configured to perform solubility prediction processing based on each of the first amino acid features and the corresponding second amino acid features, to obtain a sub-prediction change value corresponding to each of the first amino acid features, wherein the sub-prediction change value represents a solubility change between the first amino acid feature and the corresponding second amino acid feature;

[0016] The prediction processing module is used to determine the solubility change value between the original protein and the mutant protein according to each of the sub-prediction change values.

[0017] An embodiment of the present application provides an electronic device, comprising:

[0018] a memory for storing computer-executable instructions or computer programs;

[0019] The processor is configured to implement the protein solubility prediction method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.

[0020] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program, which, when executed by a processor, implements the protein solubility prediction method provided in the embodiment of the present application.

[0021] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the method for predicting protein solubility provided in the embodiment of the present application is implemented.

[0022] The embodiments of the present application have the following beneficial effects:

[0023] Predicting the change in solubility before and after a protein mutation based on the structural diagram of the protein before and after the mutation makes the change more closely related to the protein structure. Compared with the protein sequence prediction scheme in the related art, the prediction method provided by this application has better interpretability. Predicting based on each amino acid before and after the protein mutation saves the computing resources required in the prediction process compared to the prediction scheme based on the overall protein characteristics in the related art. Compared with the scheme in the related art, it has better granularity and can improve the accuracy of predicting the change in solubility before and after the protein mutation. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Schematic diagram of the application mode of the method for predicting protein solubility provided in the examples of the present application;

[0025] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;

[0026] Figure 3A This is a schematic diagram of the first process of the method for predicting protein solubility provided in the embodiments of the present application;

[0027] Figure 3B This is a second flow chart of the method for predicting protein solubility provided in the examples of the present application;

[0028] Figure 3C 3 is a schematic diagram of the third process of the method for predicting protein solubility provided in the examples of the present application;

[0029] Figure 3D 4 is a schematic diagram of a fourth flow chart of a method for predicting protein solubility provided in an embodiment of the present application;

[0030] Figure 3E 5 is a schematic diagram of the fifth flow chart of the method for predicting protein solubility provided in the examples of the present application;

[0031] Figure 3F 1 is a sixth flow chart of the method for predicting protein solubility provided in the examples of the present application;

[0032] Figure 3G 7 is a schematic diagram of the seventh flow chart of the method for predicting protein solubility provided in the examples of the present application;

[0033] Figure 3H 8 is a schematic diagram of the eighth flow chart of the method for predicting protein solubility provided in the examples of the present application;

[0034] Figure 4 This is a ninth flow chart of the method for predicting protein solubility provided in the examples of the present application;

[0035] Figure 5A This is a first principle diagram of the method for predicting protein solubility provided in the examples of the present application;

[0036] Figure 5B This is a second schematic diagram of the method for predicting protein solubility provided in the examples of the present application;

[0037] Figure 6 It is a structural diagram of the graph neural network model provided in the embodiment of the present application.

[0038] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0040] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0041] If similar descriptions of "first / second" appear in the application documents, the following explanation is added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0042] It should be pointed out that the collection and processing of relevant data in this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in practice, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0043] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0045] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0046] 1) Protein solubility, the amount or level of protein dispersion in water. Protein, as an organic macromolecular compound, exists in a dispersed state (colloidal state) in water. Therefore, there is no strict solubility of protein in water. Instead, the amount or level of protein dispersion in water is referred to as protein solubility.

[0047] 2) Protein folding is the process by which proteins acquire their functional structure and conformation. Through the physical process of protein folding, proteins fold from random coils into specific functional three-dimensional structures. When translated from mRNA sequences into linear peptide chains, proteins exist as unfolded polypeptides or random coils.

[0048] 3) The Alphafold model, a deep learning model developed by DeepMind, aims to solve the protein folding problem in biology. The Alphafold model uses deep neural networks and extensive bioinformatics data to predict the three-dimensional structure of proteins. This model can be used to understand disease mechanisms, drug development, and biological research.

[0049] 4) Graph Neural Network (GNN) refers to a general term for algorithms that use neural networks to learn graph-structured data, extract and discover features and patterns in graph-structured data, and meet the needs of graph learning tasks such as clustering, classification, prediction, segmentation, and generation.

[0050] 5) Amino acid residue: The portion of amino acids linked by peptide bonds that remains after water is lost. When the amino acids that make up a polypeptide bind to each other, some of their groups participate in the formation of the peptide bond and lose a molecule of water. Therefore, the amino acid unit in the polypeptide is called an amino acid residue.

[0051] 6) One-hot encoding, also known as single-bit encoding, uses an N-bit state register to encode N states. Each state has its own independent register bit, and at any time, only one of them is valid.

[0052] 7) The Radial Basis Function (RBF) is a real-valued function whose value depends solely on the distance from the origin, i.e., φ(x) = φ(∥x∥). Alternatively, it can be defined based on the distance from a central point c, i.e., φ(x, c) = φ(∥xc∥). Any function that satisfies φ(x) = φ(∥x∥) is called a radial function. The norm is typically the Euclidean distance, but other distance functions can also be used.

[0053] 8) A multilayer perceptron (MLP) is a feedforward artificial neural network model that maps multiple input datasets to a single output dataset. It consists of an input layer, an output layer, and one or more hidden layers. MLPs are a solution to the linearly inseparable problem. MLPs consist of stacking multiple layers of linear classifiers and adding nonlinear activation functions to the intermediate layers (also called hidden layers).

[0054] The embodiments of the present application provide a method for predicting protein solubility, a device for predicting protein solubility, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the accuracy of predicting changes in protein solubility.

[0055] The following describes exemplary applications of electronic devices provided in embodiments of the present application. The electronic devices provided in embodiments of the present application can be implemented as terminal devices, such as laptop computers, tablet computers, desktop computers, set-top boxes, smart TVs, vehicle-mounted terminals, virtual reality (VR) devices, augmented reality (AR) devices, and other types of terminals, and can also be implemented as servers. The following describes exemplary applications of electronic devices implemented as terminal devices or servers.

[0056] refer to Figure 1 , Figure 1 Schematic diagram of the application mode of the method for predicting protein solubility provided in the embodiment of the present application; for example, Figure 1The server 200, the network 300, the terminal device 400 and the database 500 are involved. The terminal device 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0057] In some embodiments, the database 500 stores a large amount of protein-related data. The user is a technician, and the terminal device 400 can be a computer used by the technician.

[0058] For example, a technician sends the protein data to be processed to a server via terminal device 400, or alternatively, a technician sends the protein data to server 200 via terminal device 400. The protein data includes a structural diagram of the protein before and after the mutation. Server 200 invokes the protein solubility prediction method provided in the embodiments of the present application based on the structural diagram of the protein before and after the mutation, obtains a predicted protein solubility change value, and sends the solubility change value to terminal device 400 for the technician to study.

[0059] The protein solubility prediction method provided in the embodiments of this application can be applied in the following application scenarios: 1. Protein drug design, calling the protein solubility prediction method provided in the embodiments of this application to determine the solubility change value of the protein drug. Researchers can evaluate the solubility of drug candidate molecules based on the solubility change value, thereby conducting better drug design. 2. Research on disease-related proteins, calling the protein solubility prediction method provided in the embodiments of this application to determine the solubility change value of the mutant protein that causes disease in the human body compared to the original protein, helping researchers determine the functional changes of the mutant protein and analyze the process of disease occurrence and development.

[0060] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.

[0061] The embodiments of the present application can be implemented through artificial intelligence technology. Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results. In other words, AI is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI is the study of the design principles and implementation methods of various intelligent machines, enabling them to have the capabilities of perception, reasoning, and decision-making. Machine Learning (ML) is a multidisciplinary interdisciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of AI and the fundamental way to make computers intelligent. Its applications are widespread in all areas of AI. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pre-trained models are the latest development in deep learning and integrate the above technologies.

[0062] The embodiments of the present application can be implemented using database technology. A database, in short, can be considered an electronic filing cabinet that stores electronic files, allowing users to add, query, update, and delete data in these files. A "database" is a collection of data that is stored together in a specific manner, can be shared by multiple users, has minimal redundancy, and is independent of applications.

[0063] A database management system (DBMS) is a computer software system designed for managing databases, typically providing basic functions such as storage, retrieval, security, and backup. DBMSs can be categorized by the database model they support, such as relational or XML (Extensible Markup Language); by the type of computer they support, such as server clusters or mobile phones; by the query language they use, such as SQL or XQuery; by performance priorities, such as maximum scale or maximum speed; or by other classification methods. Regardless of the classification method used, some DBMSs are cross-category, for example, supporting multiple query languages ​​simultaneously.

[0064] The embodiments of the present application can also be implemented through cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the rapid development and application of the Internet industry, as well as the promotion of search services, social networks, mobile commerce and open collaboration, each item may have its own hash code identification mark in the future, and all of them need to be transmitted to the background system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.

[0065] See also Figure 2 , Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application, the electronic device may be a server 200, Figure 2 The server 200 shown includes: at least one processor 410, a memory 450, and at least one network interface 420. The various components in the server 200 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .

[0066] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0067] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0068] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0069] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0070] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0071] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB).

[0072] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 A protein solubility prediction device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a data acquisition module 4551, a feature extraction module 4552, and a prediction processing module 4553. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0073] The method for predicting protein solubility provided in the embodiments of the present application will be described in conjunction with the exemplary application and implementation of the terminal device provided in the embodiments of the present application.

[0074] The following describes the protein solubility prediction method provided in the embodiments of the present application. As previously described, the electronic device implementing the protein solubility prediction method of the embodiments of the present application can be a terminal device or a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.

[0075] It should be noted that in the example of protein solubility processing below, the solubility of protein in water is used as an example. Based on the understanding of the following, those skilled in the art can apply the protein solubility prediction method provided in the embodiment of the present application to the prediction processing of changes in protein solubility in other types of liquids.

[0076] See also Figure 3A , Figure 3A This is a flow chart of the method for predicting protein solubility provided in the embodiment of the present application, which is combined with Figure 3A The steps shown are explained. Figure 3A The body of the step is Figure 1 Server 200.

[0077] In step 301, a first structural diagram of the original protein and a second structural diagram of the mutant protein are obtained.

[0078] Here, the mutant protein is obtained based on the mutation of the original protein. The first structure graph includes multiple first amino acid nodes, and the second structure graph includes multiple second amino acid nodes. Different first amino acid nodes correspond to different second amino acid nodes.

[0079] For example, the above correspondence means that the first amino acid is identical to the second amino acid, or that the second amino acid is derived from a mutation of the first amino acid. A structure graph is a two-dimensional node structure graph. Each node in the structure graph represents an amino acid, and the edges between nodes represent the connections between amino acids. Each node is annotated with information about the corresponding amino acid.

[0080] In some embodiments, reference Figure 3B , Figure 3B This is a second flow chart of the method for predicting protein solubility provided in the embodiment of the present application. Step 301 can be performed by Figure 3B Steps 3011 to 3014 are implemented as described below.

[0081] In step 3011, a first amino acid sequence of the original protein and a second amino acid sequence of the mutant protein are obtained.

[0082] For example, an amino acid sequence is used to represent the arrangement of amino acid residues in a polypeptide chain that constitutes the basic structure of a protein. Amino acid sequences can be expressed in text form. For example, the first amino acid sequence of the original protein contains "VAEKAKDE...F...LNNEIRPAH," while the second amino acid sequence of the mutant protein contains "VAEKAKDE...K...LNNEIRPAH." The first amino acid in the first amino acid sequence, represented as F, mutates to the second amino acid, K, forming the second amino acid sequence.

[0083] In step 3012, structure prediction processing is performed based on the first amino acid sequence and the second amino acid sequence to obtain a first three-dimensional structure diagram of the original protein and a second three-dimensional structure diagram of the mutant protein.

[0084] For example, structure prediction processing can be implemented using deep learning models, such as the Alphafold model and the Rosetta system. The Rosetta system is a large-scale system for text detection and recognition in images, capable of predicting the three-dimensional structure of proteins based on text in amino acid sequences. In the present embodiment, the Alphafold model is used as an example for illustration.

[0085] For example, the principles of structure prediction processing for the first amino acid sequence and the second amino acid sequence are the same, and the first amino acid sequence is used as an example for explanation. Based on the text of the first amino acid sequence, a deep learning model (Alphafold model) is called to perform feature extraction processing to obtain dependencies in the amino acid sequence, that is, amino acids in the amino acid sequence that affect other amino acids. The model analyzes the relationship between dependencies and adjacent amino acids, and the connection between at least some non-adjacent amino acids. The biological information data in the protein database is called to predict the distance between amino acids in the protein. Using the predicted distance information, the probabilistic graph model in the AlphaFold model is called to simulate the three-dimensional spatial structure of the protein. During the simulation process, possible protein folding configurations are generated based on the predicted distances between amino acids, dependencies and other constraints. The protein folding configuration is also a three-dimensional protein structure diagram, and each node in the diagram is an amino acid.

[0086] In step 3013, a first structural graph in the form of a geometric graph is constructed with each first amino acid in the first three-dimensional structural graph as a node and the chemical bonds between each first amino acid as an edge.

[0087] For example, a three-dimensional protein structure diagram is converted into a two-dimensional structure diagram to obtain a first structure diagram, in which each node is labeled with information of the corresponding amino acid, and each edge is labeled with information of the association between the amino acids.

[0088] In step 3014, a second structural graph in the form of a geometric graph is constructed with each second amino acid in the second three-dimensional structural graph as a node and the chemical bonds between each second amino acid as edges.

[0089] For example, the principle of step 3014 is the same as that of step 3013, which will not be repeated here.

[0090] In the embodiments of the present application, a deep learning model is called based on the sequence before and after the protein mutation to predict the structure of the protein before and after the mutation. The protein geometric diagram is obtained through the three-dimensional structure of the protein. Based on the two-dimensional structure diagram, various features in the protein can be obtained. Compared with directly performing feature extraction processing on the sequence, the dimension of the features can be improved, the relationship between the solubility change value and the protein structure can be reflected, and the accuracy of the solubility prediction can be improved.

[0091] Continue to refer Figure 3A In step 302, feature extraction processing is performed on the first structure graph to obtain multiple first amino acid features.

[0092] For example, the feature extraction processing for the first structure graph is a graph feature extraction processing performed on each node and edge in the structure graph. For each first amino acid, the node feature corresponding to the node of the first amino acid is extracted, and the feature of the edge between the first amino acid and the adjacent first amino acid is extracted, and the edge feature and the node feature are combined into the first amino acid feature.

[0093] In some embodiments, reference Figure 3C , Figure 3C This is a third flow chart of the method for predicting protein solubility provided in the embodiment of the present application. Step 302 can be performed by Figure 3C Steps 3021 to 3024 are implemented as described below.

[0094] In step 3021, the following processing is performed on each first amino acid node in the first structure graph: feature extraction processing is performed on the first amino acid node to obtain a node feature of the first amino acid.

[0095] For example, the node feature representation is V. Parameter types in node features include: vector features, spatial unit vectors, and one-hot encoding. The following describes the extraction of each parameter.

[0096] In some embodiments, reference Figure 3D , Figure 3D This is a fourth flow chart of the method for predicting protein solubility provided in the embodiment of the present application; Step 3021 can be performed by Figure 3D Steps 30211 to 30214 are implemented as described below.

[0097] In step 30211, a vector feature of the first amino acid node is determined based on the dihedral angles between backbone atoms in the first amino acid node.

[0098] For example, backbone atoms make up the main structure of amino acids, which are also the basic structural units of protein molecules. They include carbon, hydrogen, and nitrogen. The protein backbone in protein complexes can be represented by the backbone atoms of amino acids, which are alpha, carbon, and nitrogen. The coordinates of carbon C and nitrogen N are expressed as Therefore, the protein skeleton can be represented as a graph G = (V, E), and the vector feature representation is in ψ, ω are the dihedral angles calculated from the backbone atoms.

[0099] In step 30212, the spatial unit vector of the first amino acid is determined based on the backbone atoms in the first amino acid node.

[0100] Here, the spatial unit vector represents the extension direction of the first amino acid in space, and the types of spatial unit vectors include: forward unit vector, reverse unit vector and input direction unit vector.

[0101] For example, based on the coordinates of the skeleton atoms above, the space unit vectors are: Reverse unit vector Unit vector C in the input direction βi - The input direction unit vector is calculated by assuming tetrahedral geometry and normalizing it, referring to the following formula (2):

[0102]

[0103] in The unit vector in the input direction, the forward unit vector, and the reverse unit vector clearly define the direction of each amino acid residue, that is, the extension direction of each amino acid residue in space.

[0104] In step 30213, a one-hot encoding of the first amino acid node is determined based on the position of the first amino acid node in the first amino acid sequence of the original protein.

[0105] For example, one-hot encoding uses an N-bit state register to encode N states. Each state is represented by its own register bit. At any given time, only one register bit is active, i.e., only one state is encoded. The position of the first amino acid node currently being processed in the first amino acid sequence of the original protein is marked as 1, and all other amino acid nodes are marked as 0. This forms a one-hot encoding of the first amino acid node currently being processed. This one-hot encoding is used to represent the position of the first amino acid node in the first amino acid sequence of the original protein.

[0106] In step 30214, the vector feature, the spatial unit vector and the one-hot encoding are connected to obtain the node feature of the first amino acid.

[0107] For example, the connection process can be implemented in the following manner: the vector feature, the spatial unit vector and the one-hot encoding are represented as vectors respectively, and the vectors are connected end to end in sequence to obtain the node feature V of the first amino acid.

[0108] Continue to refer Figure 3C In step 3022, feature extraction processing is performed on the edge between the first amino acid node and each adjacent node to obtain the edge feature of the first amino acid.

[0109] For example, the edge feature is represented as E, which is determined based on the edges between an amino acid node and its adjacent amino acid nodes. The edge feature of each amino acid node is obtained by concatenating the features of different edges between the node and multiple adjacent nodes.

[0110] In some embodiments, reference Figure 3E , Figure 3E This is a fifth flow chart of the method for predicting protein solubility provided by the embodiment of the present application; Step 3022 can be performed by Figure 3E Steps 30221 to 30226 are implemented as described below.

[0111] In step 30221, the following processing is performed on the edge between the first amino acid node and each adjacent node: the first position of the first alpha carbon of the first amino acid node and the second position of the second alpha carbon of the adjacent node are obtained.

[0112] For example, the first position and the second position are only used to distinguish the positions of different alpha carbons. The edge between the nodes can be represented as E, where E = {e j→i} i≠j For all i, j, the node v in the representation structure graph j and node v i The edge between nodes v j is node v i is one of the k neighbors of . k, i, and j are all positive integers.

[0113] In step 30222, a directional unit vector of the edge is determined based on a first difference between the second position and the first position.

[0114] For example, assume that the node currently being processed is node v i Continuing with the above example, the first difference between the second position and the first position is characterized as Sure The corresponding unit vector, get the node v j and node v i The direction unit vector of the edge between them.

[0115] In step 30223, a radial basis function is called based on the norm value of the first difference to obtain the encoding distance between the first amino acid node and the adjacent node.

[0116] For example, the radial basis function can be a Gaussian radial basis function, which is a real-valued function whose value depends only on the distance to the origin. Call the Gaussian radial basis function to get the encoding distance.

[0117] In step 30224, a skeleton distance is obtained based on the sinusoidal curve encoding between the second position and the first position.

[0118] For example, the backbone distance is the distance between the first amino acid node and the adjacent node on the backbone of the original protein. The position difference between the second position and the first position is represented as a sine curve to obtain the backbone distance.

[0119] In step 30225, the direction unit vector, the encoding distance and the skeleton distance are connected to obtain the sub-edge feature between the first amino acid node and the adjacent node.

[0120] For example, the connection process can be implemented in the following manner: the direction unit vector, the encoding distance and the skeleton distance are sequentially connected in the first place to obtain the sub-edge feature between the first amino acid node and the adjacent node.

[0121] In step 30226, each sub-edge feature is concatenated to obtain the edge feature of the first amino acid node.

[0122] For example, each sub-edge feature is sequentially concatenated to obtain the edge feature of the first amino acid node.

[0123] Continue to refer Figure 3C In step 3023, the node features and the edge features are connected to obtain the first backbone feature of the first amino acid.

[0124] For example, node features and edge features can be used to describe the backbone of a protein. The first backbone feature of the first amino acid node of the i-th node can be represented as (V i , E i ).

[0125] In step 3024, feature aggregation processing is performed on the first backbone feature and each second backbone feature to obtain a first amino acid feature of the first amino acid.

[0126] For example, the second backbone feature is the backbone feature of each adjacent node of the first amino acid node. The aggregation process can be implemented as follows: the following process is performed on the second backbone feature of each adjacent node: based on the first backbone feature and the second backbone feature, information transfer is performed from the adjacent node to the first amino acid node. The information transfer can be implemented based on a geometric vector perceptron (GVP), and the information transfer results corresponding to each second backbone feature are weighted summed, and the sum of the weighted summed results and the first backbone feature is normalized to obtain the first amino acid feature.

[0127] In some embodiments, reference Figure 3F , Figure 3F This is a sixth flow chart of the method for predicting protein solubility provided by the embodiment of the present application; Step 3024 can be performed by Figure 3E Steps 30241 to 30245 are implemented as described below.

[0128] In step 30241, the following processing is performed on the second trunk feature of each adjacent node: the first trunk feature and the second trunk feature are spliced ​​to obtain a splicing result.

[0129] For example, the concatenation process (concat) is performed in the order of the first backbone feature first and the second backbone feature last. The concatenation result can be characterized as Among them, concat represents the splicing process.

[0130] In step 30242, information transmission processing is performed on the splicing processing result to obtain the message transmission vector of the adjacent node.

[0131] For example, the process of obtaining the message passing vector can be represented by the following formula (3):

[0132]

[0133] g represents the Geometric Vector Perceptron (GVP). Message passing vector is the amino acid feature of the adjacent nodes of the first amino acid node i Transmitted to the first amino acid feature get.

[0134] In step 30243, a first sum between each message passing vector is obtained, and a first product between the first sum and a preconfigured coefficient is obtained.

[0135] For example, the first sum is represented by The preconfigured coefficients can be fixed preset values, and the first product can be represented as The Dropout function is an effective regularization technique in deep learning that can effectively reduce overfitting when training neural networks. The Dropout function randomly discards a small number of node parameters in a neural network, hence the name "node dropout." This technique effectively reduces overfitting.

[0136] In step 30244, a second sum of the first backbone feature and the first product is obtained.

[0137] For example, the second sum is characterized as

[0138] In step 30245, the second sum is normalized to obtain the first amino acid feature of the first amino acid.

[0139] For example, normalization can be achieved through the normalization layer of a graph neural network. Layer Normalization (LayerNorm) is a commonly used normalization technique used for hierarchical normalization operations in neural network models. It normalizes each feature dimension of each sample so that the mean of each feature is 0 and the variance is 1, which helps improve the training effect and generalization ability of the model.

[0140] Step 30245 can be represented by the following formula (4):

[0141]

[0142] Among them, LayerNorm is layer normalization processing.

[0143] Continue to refer Figure 3A In step 303, feature extraction processing is performed on the second structure graph to obtain the second amino acid feature corresponding to each first amino acid feature.

[0144] For example, the principle of extracting the second amino acid feature from the second structure graph is the same as the principle of extracting the first amino acid feature from the first structure graph, and will not be repeated here. The order of executing steps 302 and 303 is not specific. For example, steps 302 and 303 can be executed simultaneously; steps 302 and 303 can be executed serially, with step 302 executed first and step 303 executed later; or steps 302 and 303 can be executed serially, with step 303 executed first and step 302 executed later.

[0145] When the feature extraction process is implemented through a graph neural network model, the feature extraction layer of the graph neural network model includes parallel geometric vector perceptrons, and the parallel geometric vector perceptrons perform feature extraction processing simultaneously, then step 302 and step 303 can be performed synchronously.

[0146] Continue to refer Figure 3A In step 304, solubility prediction processing is performed based on each first amino acid feature and the corresponding second amino acid feature to obtain a sub-prediction change value corresponding to each first amino acid feature.

[0147] Here, the sub-prediction change value represents the solubility change between a first amino acid feature and a corresponding second amino acid feature.

[0148] For example, the solubility prediction process can be implemented by: determining the first amino acid feature and the corresponding second amino acid feature in different splicing methods, performing prediction processing based on the two splicing features, and using the difference between the two prediction processing results as the sub-prediction change value.

[0149] In some embodiments, reference Figure 3G , Figure 3G This is a seventh flow chart of the method for predicting protein solubility provided in the embodiment of the present application; Step 304 can be performed by Figure 3G Steps 3041 to 3045 are implemented as described below.

[0150] In step 3041, the first amino acid feature and the corresponding second amino acid feature are spliced ​​together according to the first order to obtain a first splicing feature.

[0151] Here, the first order includes: the first amino acid feature precedes the second amino acid feature.

[0152] For example, the first splicing feature is characterized as in, is the first amino acid signature of the ith first amino acid of the original protein, is the second amino acid characteristic of the ith second amino acid of the mutant protein.

[0153] In step 3042, a nonlinear activation function is called based on the first splicing feature to perform activation processing to obtain a first prediction value.

[0154] For example, the activation process performed by calling the nonlinear activation function can be implemented by a multilayer perceptron (MLP), which is used to map multiple input data sets to a single output data set. The multilayer perceptron calls the nonlinear activation function in the multilayer linear classifier based on the data of each dimension in the first concatenated feature to map the first concatenated feature to a first predicted value.

[0155] In step 3043, the first amino acid feature and the corresponding second amino acid feature are concatenated according to the second sequence to obtain a second concatenation feature.

[0156] Here, the second order includes: the second amino acid feature precedes the first amino acid feature.

[0157] For example, the second splicing feature is characterized as

[0158] In step 3044, a nonlinear activation function is called to perform activation processing based on the second splicing feature to obtain a second prediction value.

[0159] The principle of step 3044 is the same as that of step 3042 and will not be repeated here. When the multilayer perceptron has a symmetrical structure, steps 3041 to 3042 and steps 3043 to 3044 are executed synchronously.

[0160] In step 3045, the difference between the first predicted value and the second predicted value is used as the sub-prediction change value corresponding to the first amino acid feature.

[0161] For example, the sub-prediction change value can be represented by the following formula (5):

[0162]

[0163] Where MLP is a symmetric multilayer perceptron network. The meaning of formula (5) is based on The first concatenated feature calls the multilayer perceptron to obtain the first prediction value based on The second concatenated feature of the multilayer perceptron is called to obtain the second predicted value, and the first predicted value minus the second predicted value is obtained to obtain x i .

[0164] Continue to refer Figure 3A , in step 305, the solubility change value between the original protein and the mutant protein is determined based on each sub-predicted change value.

[0165] For example, the method for determining the solubility change value can be a weighted summation process for each of the sub-predicted change values. The weight value corresponding to each sub-predicted change value in the weighted summation process can be pre-configured. In actual applications, the weight value can be positively correlated with the importance of the amino acid in the entire protein.

[0166] In some embodiments, step 305 can be implemented by: obtaining a preconfigured weight matrix; obtaining a fourth sum between each sub-prediction change value; and determining the product of the preconfigured weight matrix and the fourth sum as the solubility change value between the original protein and the mutant protein.

[0167] For example, the preconfigured weight matrix is ​​represented as W, and the fourth sum between each sub-prediction change value can be represented as ∑ i x i , then the process of obtaining the solubility change value is represented by the following formula (6):

[0168] f(x i )=W∑ i x i (6)

[0169] Among them, W is the weight matrix of the prediction model, f(x i ) is the predicted value.

[0170] In the examples of this application, predictions are made based on structural features of each amino acid before and after a protein mutation, resulting in better interpretability. Compared to related art approaches that predict based on overall protein features, this approach reduces computational resources required for the prediction process and offers greater granularity, improving the accuracy of predicted solubility changes before and after a protein mutation.

[0171] In some embodiments, the feature extraction process and the solubility prediction process are implemented by a graph neural network model; Figure 6 , Figure 6 Graph neural network model 600 includes a geometric graph vector perception layer 601 and a multi-layer perception layer 602, wherein the geometric graph vector perception layer 601 is used to perform feature extraction processing, and the multi-layer perception layer 602 is used to perform solubility prediction processing.

[0172] The geometric vector perception layer 601 is also known as the geometric vector perceptron (GVP), and the multi-layer perception layer 602 is also known as the multi-layer perceptron (MLP).

[0173] In some embodiments, the geometric graph vector perception layer includes a first geometric graph vector perceptron 6011 and a second geometric graph vector perceptron 6012 of identical structure. The first geometric graph vector perceptron is used to perform feature extraction processing on the first structure graph, and the second geometric graph vector perceptron is used to perform feature extraction processing on the second structure graph. The multi-layer perception layer has a symmetrical structure. For example, the multi-layer perception layer includes two multi-layer perceptrons, and the two multi-layer perceptrons perform activation processing in parallel, and the difference between the two multi-layer perceptrons is used as the solubility change value.

[0174] In some embodiments, reference Figure 3H , Figure 3H This is the eighth flow chart of the method for predicting protein solubility provided in the embodiment of the present application. Figure 3H Steps 3061 to 3064 train the graph neural network model.

[0175] In step 3061, a training sample set is obtained.

[0176] Here, the training sample set includes multiple sample groups, each sample group includes: a first sample structure diagram of the sample original amino acid, a second sample structure diagram of the sample mutated amino acid, and an actual solubility change value between the sample original amino acid and the sample mutated amino acid.

[0177] For example, the sample structure diagram can be obtained from an existing data set, and the second sample structure diagram of the sample mutated amino acid can be obtained by adjusting the first structure diagram of the sample original amino acid.

[0178] In step 3062, based on the first sample structure diagram and the second sample structure diagram, the initialized graph neural network model is called to perform solubility prediction processing to obtain a predicted solubility change value.

[0179] For example, the principle of solubility prediction processing is as follows Figure 3A As shown in steps 301 to 305 in , they will not be repeated here.

[0180] In step 3063, based on the third difference value of each sample group, the loss function value of the graph neural network model is determined.

[0181] Here, the third difference is the difference between the predicted solubility change value and the actual solubility change value of the sample group.

[0182] In some embodiments, step 3063 can be implemented by obtaining the square of the third difference value of each sample group; determining the average value between each square; and using the average value as the loss function value of the graph neural network model. For example, the above process can be represented by the following formula (1):

[0183]

[0184] y i is the true value of the change in protein solubility, and f(x i ) is the predicted value, x i This is the input vector. It is the loss determined based on the difference between the predicted value and the true value.

[0185] In step 3064, the parameters of the graph neural network model are updated based on the loss function value to obtain a trained graph neural network model.

[0186] For example, the principle of parameter update processing is back propagation processing, which iteratively performs back propagation processing on the graph neural network model based on the loss function value. The number of iterations is set according to the actual application scenario, and the trained graph neural network model is obtained after iterative training.

[0187] In the embodiments of the present application, the change in solubility before and after a protein mutation is predicted based on the structural diagram of the protein before and after the mutation, making the change more closely related to the protein structure. Compared with the protein sequence prediction schemes in the related art, the prediction method provided by the present application has better interpretability. The prediction process is based on each amino acid before and after the protein mutation. Compared with the prediction schemes based on the overall protein characteristics in the related art, the prediction process saves the computing resources required in the prediction process. Compared with the schemes in the related art, the prediction process has better granularity and can improve the accuracy of the prediction of the change in solubility before and after the protein mutation.

[0188] Below, an exemplary application of the protein solubility prediction method of the embodiment of the present application in a practical application scenario will be described.

[0189] The study of protein water solubility is helpful for the design and optimization of protein drugs in the embodiments of the present application. The water solubility of proteins may affect the distribution, absorption, metabolism and excretion of proteins in the organism. This information is crucial for the effectiveness and safety of drugs. Traditional drug water solubility studies are mainly carried out through some physical and chemical experimental methods, including spectroscopic methods, chromatographic techniques or X-ray crystallography methods, etc. These experimental methods are usually time-consuming and labor-intensive. Due to the high cost of protein solubility experiments and protein structure analysis experiments, there is almost no protein solubility data containing protein structure, so there is a lack of deep learning research on predicting protein solubility based on protein structure.

[0190] In the related art, there are two main categories of deep learning methods for predicting protein solubility. The first category is protein sequence-based methods, including Ankh and solPredict. Among them, the solPredict method obtains the protein embedding sequence (Embedding) through the Evolutionary Scale Modeling (ESM) model. The Evolutionary Scale Model is a model that uses deep learning and natural language processing techniques to understand and predict protein structure and function, and then uses these embedded sequences to train a prediction model. The first category of methods completely ignores the morphology of proteins in actual scenarios and only extracts protein embedding features based on the properties of amino acids, which lacks interpretability.

[0191] The second category is the traditional structure-based method, which is based on the assumption that the solubility of a protein is mainly determined by the chemical properties of its internal and surface amino acid residues. The overall solubility of a protein is predicted by calculating the solubility scores of the internal and surface regions of the protein. Although the second category of methods can accurately predict the solubility of a protein in some cases, it has potential limitations and shortcomings. It cannot predict the solubility of all types of proteins. For some special types of proteins, such as membrane proteins or proteins that exhibit special solubility in specific environments (such as extreme acid and alkali conditions, high temperatures, etc.), their solubility may not be accurately predicted. It is impossible to take into account all factors that affect solubility. Although the second category of methods takes into account the solubility of the internal and surface regions of the protein, there are many other factors that may also affect the solubility of the protein, such as the aggregation state, hydrophobicity, charge, etc. of the protein. These factors are not fully considered.

[0192] In response to the problems of the prior art, the embodiments of the present application use a geometric graph neural network model based on protein structure to predict changes in protein solubility, predict the protein structure through a deep learning model (for example: Alphafold model), construct a geometric graph of the structure of the protein, use atoms as points and interatomic distances as edge information, and obtain the angle information and dihedral information between the edges. Based on the constructed geometric graph, the graph neural network is called to learn the law of changes in protein solubility caused by changes in protein structure before and after amino acid mutations in proteins. The method for predicting protein solubility provided by the embodiments of the present application can intuitively observe the changes in protein structure, thereby more intuitively explaining the reasons for changes in protein solubility. By analyzing the impact of changes in the structure of the protein before and after the amino acid mutation on its solubility change, the prediction model is made more refined, and the changes in the protein structure generated by the deep learning model also make the prediction model have a certain degree of explanatory power.

[0193] The following is an explanation of the protein solubility prediction method provided in the embodiment of the present application with reference to the accompanying drawings. Figure 4, Figure 4 This is the ninth flow chart of the method for predicting protein solubility provided in the embodiment of the present application. Figure 4 The steps shown are explained. Figure 4 The body of the step is Figure 1 Server 200.

[0194] In step 401, a geometric graph neural network is trained for predicting changes in protein solubility.

[0195] For example, the geometric graph neural network learns the vector representation (embedded sequence) of each residue by learning the proximity of atoms around the target amino acid residue. Based on the trained geometric graph neural network, key amino acid residues that play an important role in protein solubility near the protein surface can be identified. Specifically, for each residue in the protein complex, the network first identifies the importance of its surrounding residues through the graph neural network mechanism, and obtains the embedded sequence of the target amino acid by scoring their total, including information including spatial proximity and physicochemical properties. Therefore, the scoring information represents the encoding environment and interaction characteristics of each residue. The geometric graph neural network is referred to as the prediction model below, and the prediction model includes a geometric vector perceptron layer and a multi-layer perceptron.

[0196] The following describes the process of obtaining a geometric graph neural network. Figure 5A , Figure 5A It is a first principle diagram of the method for predicting protein solubility provided in an embodiment of the present application. The deep learning model 501A performs structural prediction on the amino acid sequence of the mutated protein and the amino acid sequence of the protein before the mutation to obtain a three-dimensional structure diagram of the protein. For example: based on the sequence "VAEKAKDE...K...LNNEIRPAH" of the mutant (MUT) and the sequence "VAEKAKDE...F...LNNEIRPAH" of the wild type (WT), the deep learning model Alphafold is called to obtain the three-dimensional structure of the protein before and after mutation (502A and 503A). Among them, the wild type (WT) is the protein before mutation and the mutant (MUT) is the protein after mutation. The protein geometric diagram (505A and 504A) is obtained through the three-dimensional (3D) structure of the protein, and then the complex interactions between amino acids in the protein complex are learned through the graph neural network (506A and 507A). The graph neural network 506A and the graph neural network 507A can be a geometric vector perceptron (GVP) graph neural network.

[0197] To estimate the impact of a mutation, the protein mutation site is determined in advance and the amino acids surrounding the mutation site are obtained. A neural network model is then used to encode the wild-type (WT) and mutant (MUT) protein complexes to obtain the embedded sequences of the wild-type and mutant forms. The two embedded sequences are compared using the multi-layer perceptron (MLP) neural network layer of the prediction model to predict changes in solubility. The objective function for the change in solubility before and after the protein mutation is then expressed as formula (1):

[0198]

[0199] y i is the true value of the change in protein solubility, and f(x i ) is the predicted value, x i This is the input vector. It is the loss determined based on the difference between the predicted value and the true value.

[0200] The main advantage of the prediction model provided by the embodiment of the present application is that it can learn the effects of mutations from the three-dimensional protein complex structure, rather than just from sequence information. It has a certain degree of interpretability by observing the structure. In addition, by using the graph neural network mechanism, the prediction model can identify and find residue pairs that contribute significantly to solubility, rather than paying equal attention to all residue pairs. This approach may be able to provide a deeper understanding of protein-protein interactions and the effects of mutations, and may be helpful in areas such as drug design and disease research.

[0201] Continue to refer Figure 4 , in step 402, the amino acid features of the wild protein and the mutant protein are extracted respectively.

[0202] For example, the feature extraction method of the wild protein (the original protein mentioned above) and the mutant protein is the same, and the content included in the amino acid feature is explained below.

[0203] The protein backbone in a protein complex can be represented by the backbone atoms of amino acids, namely alpha carbon The coordinates of carbon C and nitrogen N are expressed as Therefore, the protein backbone can be represented as a graph G = (V, E), where each point v i ∈V represents the information of the points in the structure graph. E refers to the characteristics of the edges between the nodes corresponding to the amino acids.

[0204] Among them, the embedding characteristics of amino acid residues Included features include the following:

[0205] (1) Vector features in ψ, ω are the dihedral angles calculated from the backbone atoms.

[0206] (2) Forward unit vector and the reverse unit vector

[0207] (3) Unit vector in the input direction This is calculated by assuming tetrahedral geometry and normalizing according to the following formula (2):

[0208]

[0209] in The unit vector in the input direction, together with the forward and reverse unit vectors, clearly defines the direction of each amino acid residue, that is, the direction in which each amino acid residue extends in space.

[0210] (4) One-hot encoding of protein amino acid sequence.

[0211] Edge E = {e j→i} i≠j For all i, j, the node v in the representation structure graph j and node v i The edge between v j It is v i One of the k neighbors, using Alpha Carbon The distance between them is used to calculate the feature vector Embedding of each edge. Features included are:

[0212] (1) The unit vector

[0213] (2) Coding distance Represented by Gaussian radial basis function

[0214] (3) The encoding of the sinusoidal curve of ji represents the distance on the skeleton.

[0215] In the above description, the feature vector h for each amino acid is the concatenation of the scalar and vector features described above. That is, the scalar and vector in the node features, and the scalar and vector in the edge features, are concatenated into a single feature. Overall, these features are sufficient to fully describe the protein backbone. Based on the determined amino acid features, the features between amino acid nodes can be obtained as follows:

[0216] The information transfer formula from amino acid residue j to amino acid residue i is expressed as the following formula (3) and formula (4):

[0217]

[0218]

[0219] Among them, the concat function represents the merging layer and is used to perform feature splicing.

[0220] Message Passing Vector The message vectors from other residues are transformed through residual connections and layer normalization steps. By exchanging information between residues, the final feature vector of each residue interacting with the surrounding residues is Then the final eigenvector Predictive processing for changes in solubility.

[0221] Layer Normalization (LayerNorm) is a commonly used normalization technique used for layer-by-layer normalization in neural network models. It normalizes each feature dimension of each sample so that each feature has a mean of 0 and a variance of 1, which helps improve the model's training performance and generalization ability.

[0222] Dropout is an effective regularization technique in deep learning that can effectively reduce overfitting when training neural networks. Dropout, also known as node dropout, effectively reduces overfitting by randomly discarding a small number of node parameters in the neural network.

[0223] Continue to refer Figure 4 In step 403, the geometric graph neural network is called to predict the change in solubility before and after the protein mutation.

[0224] Example, reference Figure 5B , Figure 5B This is the second schematic diagram of the protein solubility prediction method provided in an embodiment of the present application. The trained geometric perceptron layers in graph neural networks 506A and 507A output the extracted amino acid features of the protein. Multilayer perceptrons 501B and 502B in the graph neural network process the amino acid features in parallel to predict the change in solubility before and after the protein mutation.

[0225] The above processing can be achieved in the following way: the wild-type (WT) protein complex and the mutant (MUT) protein complex are encoded by the geometric graph neural network described previously, represents the characteristics of the i-th amino acid in the wild-type complex, Represents the characteristics of the corresponding amino acids in the corresponding mutant complex. The amino acid characteristics are input into the symmetric multilayer perceptron network layer to predict the change in solubility between the two protein complexes. The prediction process refers to the following formula (5) and formula (6):

[0226]

[0227] f(x i )=W∑ i x i (6)

[0228] Among them, W is the weight matrix of the prediction model, f(x i ) is the predicted value. MLP is a symmetrical multilayer perceptron network. The meaning of the above formula (5) is based on The first concatenation vector calls the multilayer perceptron to get the first result, based on The second concatenation vector of the multilayer perceptron is called to obtain the second result, and the second result is subtracted from the first result to obtain x i .

[0229] For example, the present embodiment uses the Alphafold model to predict protein structure, and other protein structure prediction tools, such as the Rosetta system, can also be used. At the same time, the embedding features of amino acids learned through the geometric graph neural network can be replaced by an equivariant neural network.

[0230] The Rosetta system is a large-scale system for detecting and recognizing text in images. Text recognition is performed using a fully convolutional model called CTC (because it uses a sequence-to-sequence CTC loss during training), which outputs a sequence of characters. The final convolutional layer predicts the most likely character at each image position of the input word. The CTC fully convolutional model uses a sequence-to-sequence (seq2seq) CTC loss function for model training and outputs a sequence of characters.

[0231] Equivariant neural networks (Equivariant Neural Networks) are a type of neural network architecture designed to process symmetric data and maintain the symmetry of the input data in the output. Unlike traditional neural networks, equivariant neural networks maintain equivariance under transformations of the input data, meaning that transformations of the input data are reflected accordingly in the network output. Through structural design and constraints, equivariant neural networks ensure that transformations of the input data correspond to corresponding transformations of the output data, thereby maintaining symmetry.

[0232] In some embodiments, the protein solubility prediction method provided in the embodiments of the present application can be applied in the following application scenarios:

[0233] In protein engineering, researchers may mutate a protein's amino acid sequence to alter its properties, such as increasing its stability, changing its activity, or modifying its structure. Before making these mutations, predicting changes in protein solubility can help researchers assess the potential impact of these mutations on protein solubility, leading to better protein design.

[0234] In drug design, researchers may need to design protein drugs with specific properties, such as specific activity or target binding ability. In the process of designing these drugs, predicting changes in protein solubility can help researchers evaluate the solubility of drug candidate molecules and thus make better drug designs.

[0235] In disease research, such as genetic disorders or cancer, researchers may discover mutations in a protein's amino acid sequence. Predicting the effects of these mutations on protein solubility can help researchers understand why these mutations affect protein function and thus shed light on the development and progression of diseases.

[0236] In industrial biotechnology, such as biopharmaceuticals and enzymes, large quantities of proteins need to be expressed and produced. If proteins are mutated through genetic engineering to improve their performance, predicting the solubility changes after mutation can provide important information during production, helping to optimize the process.

[0237] Antibody humanization refers to the process of converting non-human antibodies (such as mouse antibodies) into human antibodies to reduce the immune response in the human body and improve therapeutic efficacy. In this process, researchers need to modify the amino acid sequence of the antibody to make it more like a human antibody. However, this modification may affect the stability and solubility of the antibody. Predicting the solubility of antibodies can help researchers evaluate the impact of these modifications on the solubility of antibodies, thereby making better designs. For example, if the prediction results show that a certain modification may reduce the solubility of the antibody, then researchers can choose other modification strategies based on the prediction results.

[0238] The following is combined with experimental data to illustrate the beneficial effects of the prediction of protein solubility provided by the embodiment of the present application: The embodiment of the present application uses the SoluProtMutDB database for experiments, cleans the SoluProtMutDB database, and filters out conflicting data. For example, if the same mutation occurs in the same position of the same protein, the solubility increases in one experiment and decreases in another experiment, the corresponding data will be deleted. Based on the data obtained by screening, the Alphafold model is called to fold the protein sequence, and 125 valid protein structures and corresponding solubility change labels are obtained to form the internal test set of the embodiment of the present application. The results of the internal data set using the model of the embodiment of the present application, the Ankh model, and the CamSol tool are shown in the following table (1). Here, the Pearson correlation coefficient r is used to represent the accuracy of the prediction of the embodiment of the present application. The larger the Pearson correlation coefficient, the more accurate the result prediction and the better the model performance.

[0239] Methods used Test results of the internal test set Embodiments of the present application 0,45 Ankh 0,28 CamSol 0,1

[0240] Table (1)

[0241] The following continues to describe the exemplary structure of the protein solubility prediction device 455 provided in the embodiments of the present application implemented as a software module. In some embodiments, Figure 2 As shown, the software modules in the protein solubility prediction device 455 stored in the memory 450 may include: a data acquisition module 4551, configured to obtain a first structure diagram of the original protein and a second structure diagram of the mutant protein, wherein the mutant protein is obtained based on the mutation of the original protein, the first structure diagram includes multiple first amino acid nodes, the second structure diagram includes multiple second amino acid nodes, and different first amino acid nodes correspond to different second amino acid nodes; a feature extraction module 4552, configured to perform feature extraction processing on the first structure diagram to obtain multiple first amino acid features, and perform feature extraction processing on the second structure diagram to obtain second amino acid features corresponding to each first amino acid feature; a prediction processing module 4553, configured to perform solubility prediction processing based on each first amino acid feature and the corresponding second amino acid feature to obtain a sub-prediction change value corresponding to each first amino acid feature, wherein the sub-prediction change value represents the solubility change between the first amino acid feature and the corresponding second amino acid feature; and a prediction processing module 4553, configured to determine the solubility change value between the original protein and the mutant protein based on each sub-prediction change value.

[0242] In some embodiments, the data acquisition module 4551 is used to obtain a first amino acid sequence of the original protein and a second amino acid sequence of the mutant protein; perform structure prediction processing based on the first amino acid sequence and the second amino acid sequence respectively to obtain a first three-dimensional structure diagram of the original protein and a second three-dimensional structure diagram of the mutant protein; use each first amino acid in the first three-dimensional structure diagram as a node and the chemical bond between each first amino acid as an edge to construct a first structure diagram in the form of a geometric graph; use each second amino acid in the second three-dimensional structure diagram as a node and the chemical bond between each second amino acid as an edge to construct a second structure diagram in the form of a geometric graph.

[0243] In some embodiments, the feature extraction module 4552 is used to perform the following processing on each first amino acid node in the first structure diagram: perform feature extraction processing on the first amino acid node to obtain the node feature of the first amino acid; perform feature extraction processing on the edge between the first amino acid node and each adjacent node to obtain the edge feature of the first amino acid; connect the node feature and the edge feature to obtain the first trunk feature of the first amino acid; perform feature aggregation processing on the first trunk feature and each second trunk feature to obtain the first amino acid feature of the first amino acid, wherein the second trunk feature is the trunk feature of each adjacent node of the first amino acid node.

[0244] In some embodiments, the feature extraction module 4552 is used to determine the vector features of the first amino acid node based on the dihedral angles between the backbone atoms in the first amino acid node; determine the spatial unit vector of the first amino acid based on the backbone atoms in the first amino acid node, wherein the spatial unit vector represents the extension direction of the first amino acid in space; determine the one-hot encoding of the first amino acid node based on the position of the first amino acid node in the first amino acid sequence of the original protein; and connect the vector features, the spatial unit vector, and the one-hot encoding to obtain the node features of the first amino acid.

[0245] In some embodiments, the feature extraction module 4552 is used to perform the following processing on the edge between the first amino acid node and each adjacent node: obtain the first position of the first alpha carbon of the first amino acid node and the second position of the second alpha carbon of the adjacent node; determine the direction unit vector of the edge based on the first difference between the second position and the first position; call the radial basis function based on the norm value of the first difference to obtain the coding distance between the first amino acid node and the adjacent node; obtain the skeleton distance based on the sinusoidal curve coding between the second position and the first position, wherein the skeleton distance is the distance between the first amino acid node and the adjacent node on the skeleton of the original protein; connect the direction unit vector, the coding distance and the skeleton distance to obtain the sub-edge feature between the first amino acid node and the adjacent node; splice each sub-edge feature to obtain the edge feature of the first amino acid node.

[0246] In some embodiments, the feature extraction module 4552 is used to perform the following processing on the second backbone feature of each adjacent node: splicing the first backbone feature and the second backbone feature to obtain a splicing processing result; performing information transmission processing on the splicing processing result to obtain a message transmission vector of the adjacent node; obtaining a first sum between each message transmission vector, and obtaining a first product between the first sum and a preconfigured coefficient; obtaining a second sum of the first backbone feature and the first product; and normalizing the second sum to obtain a first amino acid feature of the first amino acid.

[0247] In some embodiments, the prediction processing module 4553 is used to splice the first amino acid feature and the corresponding second amino acid feature according to a first order to obtain a first splicing feature, wherein the first order includes: the first amino acid feature is before the second amino acid feature; based on the first splicing feature, a nonlinear activation function is called for activation processing to obtain a first prediction value; according to the second order, the first amino acid feature and the corresponding second amino acid feature are spliced ​​to obtain a second splicing feature, wherein the second order includes: the second amino acid feature is before the first amino acid feature; based on the second splicing feature, a nonlinear activation function is called for activation processing to obtain a second prediction value; and the difference between the first prediction value and the second prediction value is used as the sub-prediction change value corresponding to the first amino acid feature.

[0248] In some embodiments, the prediction processing module 4553 is used to obtain a preconfigured weight matrix; obtain a fourth sum between each sub-prediction change value; and determine the product of the preconfigured weight matrix and the fourth sum as the solubility change value between the original protein and the mutant protein.

[0249] In some embodiments, feature extraction processing and solubility prediction processing are implemented through a graph neural network model; the prediction processing module 4553 is used to obtain a training sample set before obtaining the first structure diagram of the original protein and the second structure diagram of the mutant protein, wherein the training sample set includes multiple sample groups, each sample group includes: a first sample structure diagram of the original amino acid of the sample, a second sample structure diagram of the mutant amino acid of the sample, and an actual solubility change value between the original amino acid of the sample and the mutant amino acid of the sample; based on the first sample structure diagram and the second sample structure diagram, call the initialized graph neural network model to perform solubility prediction processing to obtain a predicted solubility change value; based on the third difference value of each sample group, determine the loss function value of the graph neural network model, wherein the third difference value is the difference between the predicted solubility change value and the actual solubility change value of the sample group; based on the loss function value, perform parameter update processing on the graph neural network model to obtain a trained graph neural network model.

[0250] In some embodiments, the prediction processing module 4553 is used to obtain the square of the third difference value of each sample group; determine the average value between each square; and use the average value as the loss function value of the graph neural network model.

[0251] In some embodiments, the graph neural network model includes a geometric graph vector perception layer and a multi-layer perception layer, wherein the geometric graph vector perception layer is used to perform feature extraction processing and the multi-layer perception layer is used to perform solubility prediction processing.

[0252] In some embodiments, the geometric graph vector perception layer includes a first geometric graph vector perception device and a second geometric graph vector perception device having the same structure. The first geometric graph vector perception device is used to perform feature extraction processing for the first structural graph, and the second geometric graph vector perception device is used to perform feature extraction processing for the second structural graph. The structure of the multi-layer perception layer is symmetrical.

[0253] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer program or computer-executable instructions, causing the electronic device to perform the protein solubility prediction method described in the present invention.

[0254] The present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the protein solubility prediction method provided in the present application, for example, Figure 3A A method for predicting protein solubility is shown.

[0255] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.

[0256] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0257] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0258] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0259] In summary, the present invention predicts the change in solubility before and after a protein mutation based on the structural diagram of the protein before and after the mutation, making the change more closely related to the protein structure. Compared with the protein sequence prediction scheme in the related art, the prediction method provided by the present invention has better interpretability. The prediction process is based on each amino acid before and after the protein mutation. Compared with the prediction scheme based on the overall protein characteristics in the related art, the prediction process saves the computing resources required in the prediction process, and has better granularity than the scheme in the related art, which can improve the accuracy of predicting the change in solubility of the protein before and after the mutation.

[0260] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A method for predicting protein solubility, characterized in that The method comprises: Obtaining a first structure diagram of an original protein and a second structure diagram of a mutant protein, wherein the mutant protein is obtained based on a mutation of the original protein, the first structure diagram includes a plurality of first amino acid nodes, the second structure diagram includes a plurality of second amino acid nodes, and different first amino acid nodes correspond to different second amino acid nodes; performing feature extraction processing on the first structure graph to obtain a plurality of first amino acid features; performing feature extraction processing on the second structure graph to obtain second amino acid features corresponding to each of the first amino acid features; performing solubility prediction processing based on each of the first amino acid features and the corresponding second amino acid features to obtain a sub-prediction change value corresponding to each of the first amino acid features, wherein the sub-prediction change value represents a solubility change between the first amino acid feature and the corresponding second amino acid feature; Based on each of the sub-predicted change values, the solubility change value between the original protein and the mutant protein is determined.

2. The method according to claim 1, characterized in that The obtaining of the first structural diagram of the original protein and the second structural diagram of the mutant protein comprises: obtaining a first amino acid sequence of the original protein and a second amino acid sequence of the mutant protein; Performing structure prediction processing based on the first amino acid sequence and the second amino acid sequence to obtain a first three-dimensional structure diagram of the original protein and a second three-dimensional structure diagram of the mutant protein; Constructing a first structural diagram in the form of a geometric graph by using each of the first amino acids in the first three-dimensional structural diagram as a node and using the chemical bonds between each of the first amino acids as edges; A second structural graph in the form of a geometric graph is constructed by taking each of the second amino acids in the second three-dimensional structural graph as a node and taking the chemical bonds between each of the second amino acids as edges.

3. The method according to claim 1, characterized in that The performing feature extraction processing on the first structure graph to obtain a plurality of first amino acid features includes: The following processing is performed for each of the first amino acid nodes in the first structure graph: performing feature extraction processing on the first amino acid node to obtain a node feature of the first amino acid; performing feature extraction processing on the edge between the first amino acid node and each adjacent node to obtain edge features of the first amino acid; Connecting the node features and the edge features to obtain a first backbone feature of the first amino acid; Feature aggregation processing is performed on the first backbone feature and each second backbone feature to obtain a first amino acid feature of the first amino acid, wherein the second backbone feature is a backbone feature of each adjacent node of the first amino acid node.

4. The method according to claim 3, characterized in that The performing feature extraction processing on the first amino acid node to obtain the node feature of the first amino acid includes: determining a vector feature of the first amino acid node based on dihedral angles between backbone atoms in the first amino acid node; Determining a spatial unit vector of the first amino acid based on backbone atoms in the first amino acid node, wherein the spatial unit vector represents an extension direction of the first amino acid in space; determining a one-hot encoding of the first amino acid node based on a position of the first amino acid node in the first amino acid sequence of the original protein; The vector feature, the spatial unit vector, and the one-hot encoding are connected to obtain a node feature of the first amino acid.

5. The method according to claim 3, characterized in that The performing feature extraction processing on the edge between the first amino acid node and each adjacent node to obtain the edge feature of the first amino acid includes: The following processing is performed on the edge between the first amino acid node and each of the adjacent nodes: Obtain a first position of a first alpha carbon of the first amino acid node and a second position of a second alpha carbon of the adjacent node; determining a directional unit vector of the edge based on a first difference between the second position and the first position; Calling a radial basis function based on a norm value of the first difference value to obtain a coding distance between the first amino acid node and the adjacent node; Obtaining a backbone distance based on a sinusoidal encoding between the second position and the first position, wherein the backbone distance is a distance between the first amino acid node and the adjacent node on the backbone of the original protein; Connecting the direction unit vector, the encoding distance, and the skeleton distance to obtain a sub-edge feature between the first amino acid node and the adjacent node; Each of the sub-edge features is concatenated to obtain the edge feature of the first amino acid node.

6. The method according to claim 3, characterized in that The performing feature aggregation processing on the first backbone feature and each second backbone feature to obtain the first amino acid feature of the first amino acid includes: The following processing is performed on the second backbone feature of each adjacent node: Performing a splicing process on the first trunk feature and the second trunk feature to obtain a splicing result; Performing information transfer processing on the splicing processing result to obtain a message transfer vector of the adjacent node; Obtaining a first sum between each of the message passing vectors, and obtaining a first product between the first sum and a preconfigured coefficient; Obtaining a second sum of the first backbone feature and the first product; The second sum is normalized to obtain a first amino acid feature of the first amino acid.

7. The method according to any one of claims 1 to 6, characterized in that The performing solubility prediction processing based on each of the first amino acid features and the corresponding second amino acid features to obtain a sub-prediction change value corresponding to each of the first amino acid features includes: The following processing is performed for each of the first amino acid features: splicing the first amino acid feature and the corresponding second amino acid feature according to a first sequence to obtain a first splicing feature, wherein the first sequence includes: the first amino acid feature is before the second amino acid feature; Calling a nonlinear activation function to perform activation processing based on the first splicing feature to obtain a first prediction value; splicing the first amino acid feature and the corresponding second amino acid feature according to a second sequence to obtain a second splicing feature, wherein the second sequence includes: the second amino acid feature is before the first amino acid feature; Calling a nonlinear activation function to perform activation processing based on the second splicing feature to obtain a second prediction value; The difference between the first predicted value and the second predicted value is used as the sub-prediction change value corresponding to the first amino acid feature.

8. The method according to any one of claims 1 to 6, characterized in that Determining the solubility change value between the original protein and the mutant protein according to each of the sub-predicted change values ​​comprises: Get the preconfigured weight matrix; Obtaining a fourth sum between each of the sub-prediction change values; A product of the preconfigured weight matrix and the fourth sum is determined as a solubility change value between the original protein and the mutant protein.

9. The method according to any one of claims 1 to 6, characterized in that The feature extraction process and the solubility prediction process are implemented by a graph neural network model; Before obtaining the first structural diagram of the original protein and the second structural diagram of the mutant protein, the method further comprises: Obtaining a training sample set, wherein the training sample set includes a plurality of sample groups, each of the sample groups including: a first sample structure diagram of the original amino acid of the sample, a second sample structure diagram of the mutated amino acid of the sample, and an actual solubility change value between the original amino acid of the sample and the mutated amino acid of the sample; Based on the first sample structure graph and the second sample structure graph, calling the initialized graph neural network model to perform solubility prediction processing to obtain a predicted solubility change value; Determining a loss function value of the graph neural network model based on a third difference value of each sample group, wherein the third difference value is a difference between the predicted solubility change value and the actual solubility change value of the sample group; The parameters of the graph neural network model are updated based on the loss function value to obtain the trained graph neural network model.

10. The method according to claim 9, characterized in that Determining the loss function value of the graph neural network model based on the third difference value of each sample group includes: Obtaining the square of the third difference value of each sample group; determining a mean between each of said squares; The average value is used as the loss function value of the graph neural network model.

11. The method according to claim 9, characterized in that The graph neural network model includes a geometric graph vector perception layer and a multi-layer perception layer, wherein the geometric graph vector perception layer is used to perform the feature extraction processing, and the multi-layer perception layer is used to perform the solubility prediction processing.

12. The method according to claim 11, characterized in that The geometric graph vector perception layer includes a first geometric graph vector perceptron and a second geometric graph vector perceptron with the same structure. The first geometric graph vector perceptron is used to perform feature extraction processing on the first structural graph, and the second geometric graph vector perceptron is used to perform feature extraction processing on the second structural graph. The structure of the multi-layer perception layer is symmetrical.

13. A device for predicting protein solubility, characterized in that: The device comprises: a data acquisition module, configured to acquire a first structure diagram of an original protein and a second structure diagram of a mutant protein, wherein the mutant protein is obtained based on a mutation of the original protein, the first structure diagram includes a plurality of first amino acid nodes, the second structure diagram includes a plurality of second amino acid nodes, and different first amino acid nodes correspond to different second amino acid nodes; a feature extraction module, configured to perform feature extraction processing on the first structure graph to obtain a plurality of first amino acid features, and perform feature extraction processing on the second structure graph to obtain a second amino acid feature corresponding to each of the first amino acid features; a prediction processing module, configured to perform solubility prediction processing based on each of the first amino acid features and the corresponding second amino acid features, to obtain a sub-prediction change value corresponding to each of the first amino acid features, wherein the sub-prediction change value represents a solubility change between the first amino acid feature and the corresponding second amino acid feature; The prediction processing module is used to determine the solubility change value between the original protein and the mutant protein according to each of the sub-prediction change values.

14. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the method for predicting protein solubility according to any one of claims 1 to 12 when executing the computer-executable instructions or computer program stored in the memory.

15. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer-executable instructions or computer program are executed by a processor, the method for predicting protein solubility according to any one of claims 1 to 12 is implemented.

16. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the method for predicting protein solubility according to any one of claims 1 to 12 is implemented.