Method and system for processing protein phase separation behavior on basis of machine learning

By obtaining and processing the amino acid sequence and mutation sequence of proteins, combining the protein graph structure prediction model and isolation genetic algorithm, the problem of insufficient time performance of protein phase separation prediction methods in the prior art and the inability to predict mutant proteins is solved, and fast and convenient protein phase separation prediction is achieved.

WO2025131097A1PCT designated stage expired Publication Date: 2025-06-26SHENZHEN UNIV

Patent Information

Application Number
PCT/CN2024/141140
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing protein phase separation prediction methods have poor time performance, are not suitable for massive search tasks, and cannot perform phase separation prediction of brand new proteins after mutated, making it inconvenient for actual use by biologists.

Method used

By obtaining the amino acid sequence of the target protein, the mutated amino acid sequence is obtained based on the preset site, and input it into the trained protein map structure prediction model to output a three-dimensional structural feature map. Then, the three-dimensional structural feature map, amino acid sequence and mutant amino acid sequence are input into the separation genetic algorithm, and the training phase separation ability distinction model is used to construct a fitness function to generate phase separation result information.

Benefits of technology

It quickly obtains the three-dimensional structural feature map of the mutant protein, which is convenient to obtain protein phase separation information, and can perform phase separation and prediction of the brand new protein after mutation, improving the efficiency and convenience of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141140_26062025_PF_FP_ABST
    Figure CN2024141140_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A method and system for processing a protein phase separation behavior on the basis of machine learning, which relates to the field of artificial intelligence. The method comprises: acquiring an amino acid sequence of a target protein, and obtaining a mutant amino acid sequence on the basis of a preset site; inputting the mutant amino acid sequence into a trained protein graph structure prediction model, and outputting a three-dimensional structure feature map of the mutant amino acid sequence; and inputting the three-dimensional structure feature map, the amino acid sequence, and the mutant amino acid sequence into a separation genetic algorithm, and outputting phase separation result information of the mutant amino acid sequence, a fitness function of the separation genetic algorithm being constructed by using a trained phase separation capability differentiation model. By means of said method, it is possible to rapidly obtain the phase separation situation of a protein, and it is also possible to obtain the phase separation situation of a new protein after mutation.
Need to check novelty before this filing date? Find Prior Art

Description

Method and system for processing protein phase separation behavior based on machine learning Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method, system, intelligent terminal and computer-readable storage medium for processing protein phase separation behavior based on machine learning. Background Art

[0002] Protein phase separation occurs when weak interactions between proteins cause them to condense into droplets, creating a sparse phase (LP) and a dense phase (DP) of varying concentrations within the cell. Ultimately, different substances within the cell separate into distinct phases. In recent years, biological research based on protein phase separation has become a hot topic in biology.

[0003] Currently, protein phase separation prediction mostly uses phase separation protein prediction tools to predict the phase separation of natural proteins. However, the current protein phase separation prediction methods have poor time performance, are not suitable for massive search tasks, and cannot predict phase separation of new proteins after mutation, which is not convenient for practical use by biologists.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method and system for processing protein phase separation behavior based on machine learning, aiming to solve the problems that the protein phase separation prediction methods in the existing technology have poor time performance, are not suitable for massive search tasks, and cannot predict phase separation of new proteins after mutation, which is inconvenient for actual use by users.

[0006] In order to achieve the above objectives, the present invention provides a first aspect of a method for processing protein phase separation behavior based on machine learning, wherein the method for processing protein phase separation behavior based on machine learning comprises:

[0007] Obtain the amino acid sequence of the target protein and obtain the mutant amino acid sequence according to the preset site;

[0008] Inputting the mutant amino acid sequence into a trained protein graph structure prediction model to output a three-dimensional structural feature graph of the mutant amino acid sequence;

[0009] The three-dimensional structural feature map, the amino acid sequence and the mutant amino acid sequence are input into a separation genetic algorithm, and phase separation result information of the mutant amino acid sequence is output, wherein the fitness function of the separation genetic algorithm is constructed using the trained phase separation ability differentiation model.

[0010] Optionally, inputting the mutant amino acid sequence into a trained protein graph structure prediction model to output a three-dimensional structural feature graph of the mutant amino acid sequence specifically includes:

[0011] Inputting the mutant amino acid sequence into the trained protein graph structure prediction model;

[0012] Based on the word vector generation model of the trained protein graph structure prediction model, a first three-dimensional structural feature graph of the mutant amino acid sequence is generated, the first three-dimensional structural feature graph is input into the self-attention mechanism for encoding, and the three-dimensional structural feature graph is output.

[0013] Optionally, the protein graph structure prediction model is trained according to the structure prediction results of proteins in a preset database to obtain the trained protein graph structure prediction model.

[0014] Optionally, inputting the three-dimensional structural feature map, the amino acid sequence, and the mutant amino acid sequence into a separation genetic algorithm and outputting phase separation result information of the mutant amino acid sequence specifically includes:

[0015] Inputting the three-dimensional structural feature map, the amino acid sequence and the mutant amino acid sequence into the segregation genetic algorithm;

[0016] The separation genetic algorithm generates phase separation ability evaluation information according to the fitness function, and obtains the phase separation result information according to the phase separation ability evaluation information.

[0017] Optionally, the fitness function of the separation genetic algorithm is constructed using the trained phase separation ability differentiation model, specifically including:

[0018] Using a dual-tower graph contrast learning framework to train the phase separation ability differentiation model to obtain the trained phase separation ability differentiation model;

[0019] The trained phase separation ability differentiation model is used to construct the fitness function of the separation genetic algorithm.

[0020] Optionally, the adopting of a dual-tower graph contrast learning framework to train the phase separation ability differentiation model to obtain the trained phase separation ability differentiation model specifically includes:

[0021] Acquire a preset first auxiliary data set and a second auxiliary data set;

[0022] Simultaneously training the phase separation ability differentiation model based on the first auxiliary data set and the second auxiliary data set;

[0023] When the preset training condition is reached, the training of the phase separation ability differentiation model is stopped, and the phase separation ability differentiation model obtained by the last training is used as the trained phase separation ability differentiation model.

[0024] Optionally, the simultaneously training the phase separation ability differentiation model based on the first auxiliary dataset and the second auxiliary dataset specifically includes:

[0025] Each time the phase separation ability differentiation model is trained simultaneously using the first auxiliary data set and the second auxiliary data set, obtaining a first gradient and a second gradient of a previous training process;

[0026] Adjusting the second gradient in the current training process according to the first gradient and the second gradient in the previous training process to obtain a current second gradient;

[0027] Based on the first gradient, the phase separation ability differentiation model is trained using the first auxiliary data set, and based on the current second gradient, the phase separation ability differentiation model is simultaneously trained using the second auxiliary data set.

[0028] A second aspect of the present invention provides a protein phase separation behavior processing system based on machine learning, wherein the protein phase separation behavior processing system based on machine learning comprises:

[0029] The data acquisition module is used to obtain the amino acid sequence of the target protein and obtain the mutant amino acid sequence according to the preset site;

[0030] a feature generation module, configured to input the mutant amino acid sequence into a trained protein graph structure prediction model and output a three-dimensional structural feature graph of the mutant amino acid sequence;

[0031] A result generation module is used to input the three-dimensional structural feature map, the amino acid sequence and the mutant amino acid sequence into a separation genetic algorithm, and output phase separation result information of the mutant amino acid sequence, wherein the fitness function of the separation genetic algorithm is constructed using the trained phase separation ability differentiation model.

[0032] A third aspect of the present invention provides an intelligent terminal, comprising a memory, a processor, and a machine learning-based protein phase separation behavior processing program stored in the memory and executable on the processor. When the machine learning-based protein phase separation behavior processing program is executed by the processor, the program implements any one of the steps of the machine learning-based protein phase separation behavior processing method.

[0033] A fourth aspect of the present invention provides a computer-readable storage medium, on which is stored a protein phase separation behavior processing program based on machine learning. When the protein phase separation behavior processing program based on machine learning is executed by a processor, it implements any step of the protein phase separation behavior processing method based on machine learning.

[0034] As can be seen from the above, in the scheme of the present invention, the amino acid sequence of the target protein is obtained, and a mutant amino acid sequence is obtained according to the preset site; the mutant amino acid sequence is input into the trained protein graph structure prediction model, and a three-dimensional structural feature graph of the mutant amino acid sequence is output; the three-dimensional structural feature graph, the amino acid sequence and the mutant amino acid sequence are input into the separation genetic algorithm, and the phase separation result information of the mutant amino acid sequence is output, wherein the trained phase separation ability discrimination model is used to construct the fitness function of the separation genetic algorithm.

[0035] Compared with the existing technology, in order to address the problem that the current protein phase separation prediction method has poor time performance and is not suitable for massive search tasks, the present invention uses a trained protein graph structure prediction model to process the mutant amino acid sequence, so that a three-dimensional structural feature map can be quickly obtained. On the basis of being able to obtain the three-dimensional structural feature map of the mutant protein in real time, the separation genetic algorithm in the present invention can conveniently obtain protein phase separation information; at the same time, in order to address the problem that the phase separation of the new protein after mutation cannot be predicted, which is inconvenient for actual use by users, the present invention processes the mutant protein through a separation genetic algorithm based on a phase separation ability differentiation model to construct a fitness function, which can obtain phase separation result information of the mutant amino acid sequence, so that the phase separation of the new protein after mutation can be quickly predicted, which facilitates the tracking of the decisive sequence perturbations related to phase separation, and thus can change the protein phase separation behavior in a predictable manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0037] FIG1 is a schematic flow chart of a method for processing protein phase separation behavior based on machine learning provided in an embodiment of the present invention;

[0038] FIG2 is a schematic diagram of a process framework of a method for processing protein phase separation behavior based on machine learning provided in an embodiment of the present invention;

[0039] FIG3 is a schematic diagram of the workflow of a method for processing protein phase separation behavior based on machine learning provided in an embodiment of the present invention;

[0040] FIG4 is a schematic diagram of the structure of a protein graph structure prediction model provided by an embodiment of the present invention;

[0041] FIG5 is a schematic diagram of a processing flow of a dual-tower graph comparative learning framework provided by an embodiment of the present invention;

[0042] FIG6 is a schematic diagram of the processing flow of the phase separation ability differentiation model provided by an embodiment of the present invention

[0043] FIG7 is a schematic diagram of the components of a protein phase separation behavior processing system based on machine learning provided by an embodiment of the present invention;

[0044] FIG8 is a block diagram of the internal structure of a smart terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration and not limitation to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present invention with unnecessary detail.

[0046] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0047] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0048] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0049] As used in this specification and the appended claims, the term "if" can be interpreted as meaning "when" or "upon" or "in response to determining" or "in response to being classified into," depending on the context. Similarly, the phrase "if it is determined" or "if it is classified into [described condition or event]" can be interpreted as meaning "upon determination" or "in response to determining" or "upon classification into [described condition or event]" or "in response to being classified into [described condition or event]," depending on the context.

[0050] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0051] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0052] Protein phase separation occurs when weak interactions between proteins cause them to condense into droplets, creating a sparse phase (LP) and a dense phase (DP) of varying concentrations within the cell. Ultimately, different substances within the cell separate into distinct phases. In recent years, biological research based on protein phase separation has become a hot topic in biology.

[0053] Currently, protein phase separation prediction mostly uses phase separation protein prediction tools to predict the phase separation of natural proteins. However, the current protein phase separation prediction methods have poor time performance, are not suitable for massive search tasks, and cannot predict phase separation of new proteins after mutation, which is not convenient for practical use by biologists.

[0054] In order to solve at least one of the above-mentioned problems, the present invention obtains the amino acid sequence of a target protein and obtains a mutant amino acid sequence according to preset sites; inputs the mutant amino acid sequence into a trained protein graph structure prediction model, and outputs a three-dimensional structural feature graph of the mutant amino acid sequence; inputs the three-dimensional structural feature graph, the amino acid sequence, and the mutant amino acid sequence into a separation genetic algorithm, and outputs phase separation result information of the mutant amino acid sequence, wherein the trained phase separation ability differentiation model is used to construct a fitness function of the separation genetic algorithm.

[0055] Compared with the existing technology, in order to address the problem that the current protein phase separation prediction method has poor time performance and is not suitable for massive search tasks, the present invention uses a trained protein graph structure prediction model to process the mutant amino acid sequence, so that a three-dimensional structural feature map can be quickly obtained. On the basis of being able to obtain the three-dimensional structural feature map of the mutant protein in real time, the separation genetic algorithm in the present invention can conveniently obtain protein phase separation information; at the same time, in order to address the problem that the phase separation of the new protein after mutation cannot be predicted, which is inconvenient for actual use by users, the present invention processes the mutant protein through a separation genetic algorithm based on a phase separation ability differentiation model to construct a fitness function, which can obtain phase separation result information of the mutant amino acid sequence, so that the phase separation of the new protein after mutation can be quickly predicted, which facilitates the tracking of the decisive sequence perturbations related to phase separation, and thus can change the protein phase separation behavior in a predictable manner.

[0056] Exemplary Methods

[0057] As shown in FIG1 , an embodiment of the present invention provides a method for processing protein phase separation behavior based on machine learning. Specifically, the method for processing protein phase separation behavior based on machine learning includes the following steps:

[0058] Step S100: Obtain the amino acid sequence of the target protein and obtain a mutant amino acid sequence according to a preset site.

[0059] It should be noted that the structure of a protein may change due to mutation, and changes in the protein structure will cause changes in its original phase separation ability. Obtaining the phase separation ability of the mutated protein is extremely important for biological research, but currently there are certain difficulties in obtaining the phase separation ability of the mutated protein. In this application, after obtaining the amino acid sequence of the target protein, the mutant amino acid sequence is obtained according to the preset site. The method for obtaining the amino acid sequence of the target protein is not limited in this application, and the preset site is a pre-selected mutation position, and the mutant amino acid sequence is obtained according to the preset mutation method.

[0060] Step S200: input the mutant amino acid sequence into a trained protein graph structure prediction model, and output a three-dimensional structural feature graph of the mutant amino acid sequence.

[0061] Specifically, in the present application, the mutant amino acid sequence is input into the trained protein graph structure prediction model, and the corresponding three-dimensional structural feature map of the mutant amino acid is output. Among them, the specific structure of the protein graph structure prediction model is shown in Figure 4. The protein graph structure prediction model is referred to as BetaFold. BetaFold first generates a feature matrix of the mutant amino acid sequence by using the biological serialization variant ProtVec of the word vector generation model, that is, the first three-dimensional structural feature map, and then inputs the feature matrix into the self-attention mechanism for encoding and decoding process, that is, the encoding and decoding process is performed in the Transformer architecture, that is, Seq2Seq is performed. After encoding by the self-attention mechanism Self-Attention, the feature vector corresponding to each amino acid in the feature matrix output by the Transformer architecture contains potential correlation information with other amino acids in the sequence, that is, the three-dimensional structural feature map is obtained. Among them, the word vector generation model in the present application adopts Word2Vec.

[0062] In addition, BetaFold pairs the output feature vectors corresponding to each amino acid in the three-dimensional structural feature map as input vectors of a feedforward neural network, uses the feedforward neural network as a classifier to predict whether there is a residue contact relationship between each pair of amino acids, and outputs the direct residue contact relationship of each pair of amino acids. The three-dimensional structural feature map in this application can be further described by the base contact relationship. When the output feature vectors corresponding to each amino acid in the three-dimensional structural feature map are paired, arbitrary pairing is performed. For example, when there are 10 amino acids, there are 10*9 / 2=45 pairs of pairings.

[0063] Furthermore, in this application, the formula of the self-attention mechanism in the protein graph structure prediction model can be expressed as:

[0064] Attention(Q,K,V)=ReLU 2 (Q'((K') T V));

[0065] Among them, Q, K, and V are general representations in the self-attention mechanism, which are the new feature matrices obtained by the input feature matrix through their respective linear transformation functions. Q' and K' respectively represent the mapping of Q and K on the kernel function. ReLU 2 It is a calculation of a ReLU function operation and then taking the square; that is, in this application, the activation function of the feedforward neural network in the Transformer architecture is modified to ReLU 2 This can reduce the computational complexity of the model while ensuring the effectiveness of the high-order polynomial of the eigenvector corresponding to the amino acids in the protein sequence in the model.

[0066] The protein graph structure prediction model is trained according to the protein structure prediction results in a preset database to obtain the trained protein graph structure prediction model.

[0067] Furthermore, the step of inputting the mutant amino acid sequence into a trained protein graph structure prediction model and outputting a three-dimensional structural feature graph of the mutant amino acid sequence specifically includes:

[0068] Inputting the mutant amino acid sequence into the trained protein graph structure prediction model;

[0069] Based on the word vector generation model of the trained protein graph structure prediction model, a first three-dimensional structural feature graph of the mutant amino acid sequence is generated, the first three-dimensional structural feature graph is input into the self-attention mechanism for encoding, and the three-dimensional structural feature graph is output.

[0070] Specifically, in the present application, the trained protein graph structure prediction model can output a three-dimensional structural feature graph containing residue contact relationships for the input mutant amino acid sequence, and the phase separation ability of the mutant amino acid sequence can be further obtained through the three-dimensional structural feature graph.

[0071] Furthermore, in the present application, the protein graph structure prediction model is trained according to the structure prediction results of proteins in a preset database to obtain the trained protein graph structure prediction model.

[0072] Specifically, when training the protein graph structure prediction model, the protein structure data predicted by AlphaFold2 is used to train the protein graph structure prediction model. The prediction speed of AlphaFold2 is slow and does not meet the task requirements of obtaining the protein structure feature map in this application. However, AlphaFold2 has stored most of the natural protein prediction results in the database. Therefore, the protein graph structure prediction model in this application can be trained by the natural protein prediction results in AlphaFold2, thereby overcoming the problem of insufficient training data of the traditional protein graph structure prediction model.

[0073] When the number of training times reaches the preset requirement or the model accuracy reaches the preset threshold, the training is stopped, and the protein graph structure prediction model obtained from the last training is used as the trained protein graph structure prediction model.

[0074] Step S300: input the three-dimensional structural feature map, the amino acid sequence, and the mutant amino acid sequence into a separation genetic algorithm, and output phase separation result information of the mutant amino acid sequence, wherein the fitness function of the separation genetic algorithm is constructed using the trained phase separation ability differentiation model.

[0075] Specifically, in the present application, the fitness function of the separation genetic algorithm is constructed based on the trained phase separation ability differentiation model, and the fitness function is used to output the phase separation result information of the mutant amino acid sequence according to the amino acid sequence and the mutant amino acid sequence. Among them, the three-dimensional structural feature map predicted by the BetaFold model will also be used as an additional input feature of the fitness function in the separation genetic algorithm.

[0076] Furthermore, the step of inputting the three-dimensional structural feature map, the amino acid sequence, and the mutant amino acid sequence into a separation genetic algorithm and outputting phase separation result information of the mutant amino acid sequence specifically includes:

[0077] Inputting the three-dimensional structural feature map, the amino acid sequence and the mutant amino acid sequence into the segregation genetic algorithm;

[0078] The separation genetic algorithm generates phase separation ability evaluation information according to the fitness function, and obtains the phase separation result information according to the phase separation ability evaluation information.

[0079] Specifically, in order to screen out key mutation sites on the amino acid sequence of a protein that may cause a significant change in the liquid-liquid phase separation (LLPS) ability, the present invention combines a phase separation ability differentiation model with a separation genetic algorithm based on a genetic algorithm, which is used by biologists to track the decisive sequence perturbations associated with LLPS, thereby changing the protein phase separation behavior in a predictable manner. This plays a very important role in understanding the role of phase separation in molecular biology in the biological and medical fields. Considering the huge possible space of mutation combinations, the separation genetic algorithm is used in this application to search for the optimal solution, and its fitness function is defined as: fit(x) = || PSDM(x) - PSDM(x org )||;

[0080] Where x is the mutated protein sequence, i.e., the mutated amino acid sequence, org is the original protein sequence, that is, the amino acid sequence of the target protein, and PSDM represents the trained phase separation ability differentiation model. The separation genetic algorithm works in the form of a stream. First, a batch of mutants of the target protein are randomly generated and used as the initial population. Then, a certain number of excellent samples are selected from the population and retained according to the fitness function fit(x); a new population can be generated by crossover and mutation of the retained samples; after a certain number of iterations, the best sample in the population is taken as the optimal solution to the problem, which is the phase separation result information obtained in this application. In addition, modifying a large number of amino acid sites at the same time may also significantly change the functional properties of the protein, but this does not conform to the situation where natural mutations occur. Therefore, in this application, only one or several sites are modified, that is, the modification of the amino acid sites is based on the preset sites, and it should be ensured that the number of mutation sites in each sample does not exceed a small positive integer k. In one embodiment of the present application, k is set to 2.

[0081] Furthermore, the fitness function of the separation genetic algorithm is constructed using the trained phase separation ability differentiation model, specifically including:

[0082] Using a dual-tower graph contrast learning framework to train the phase separation ability differentiation model to obtain the trained phase separation ability differentiation model;

[0083] The trained phase separation ability differentiation model is used to construct the fitness function of the separation genetic algorithm.

[0084] It should be noted that the phase separation ability differentiation model in this application is constructed by combining the self-attention mechanism Self-Attention and the graph self-attention network GAT, as shown in Figure 6, where Q, K, and V are general representations in the self-attention mechanism, and are the new feature matrices obtained by the input feature matrix through their respective linear transformation functions. L represents the sequence length of the input protein, r represents the dimensions of Q and K, and d represents the dimension of V.

[0085] Specifically, in this application, for the phase separation model used to construct the fitness function of the separation genetic algorithm, a dual-tower graph contrastive learning framework is used to train the phase separation ability differentiation model to obtain the trained phase separation ability differentiation model. The dual-tower graph contrastive learning framework T3GCL achieves the purpose of making the model training effect close to the original task by simultaneously performing contrastive learning training on two auxiliary tasks associated with the original task labels. As shown in Figure 5, the dataset of the target task is represented by A, and the two auxiliary task datasets are represented by B and C respectively; the auxiliary datasets B and C are composed of sufficient and well-labeled data, and A is the intersection of the two datasets. The dual-tower graph contrastive learning method allows the model to learn the intrinsic consistency characteristics of the data and ignore other unimportant features. When assuming that the intrinsic consistency characteristics of B and C are FB = {f_1, f_2, f_3} and FC = {f_4, f_5}, then the intrinsic consistency information of A is FA = {f_1, f_2, f_3, f_4, f_5}. When A has a sufficiently strong prior association with B and C, respectively, and satisfies A=B∩C, then the features FB∪C extracted by training on both B and C can be used as an approximation of FA, that is, the features after training on dataset A can be obtained. Therefore, even if A lacks labels related to the target task, its inherent consistency information can still be collaboratively learned by training two auxiliary tasks on B and C simultaneously. And in order to achieve the goal of collaborative learning by training two auxiliary tasks on B and C simultaneously, the gradients of the two auxiliary tasks should be properly balanced. This is because the learning difficulty of the two tasks may be significantly different from each other, which will cause the model to only capture the inherent consistency features of one task while ignoring the other task, so that the trained model cannot achieve the expected effect.

[0086] Furthermore, the phase separation ability differentiation model is trained using a dual-tower graph contrast learning framework to obtain the trained phase separation ability differentiation model, specifically comprising:

[0087] Acquire a preset first auxiliary data set and a second auxiliary data set;

[0088] Simultaneously training the phase separation ability differentiation model based on the first auxiliary data set and the second auxiliary data set;

[0089] When the preset training condition is reached, the training of the phase separation ability differentiation model is stopped, and the phase separation ability differentiation model obtained by the last training is used as the trained phase separation ability differentiation model.

[0090] Specifically, the twin-tower graph contrast learning framework is applied to the present application, and the corresponding first auxiliary data set B is the clinical mutation data set NCBI ClinVar, and the second auxiliary data set C is the natural protein phase separation data set LLPSDB, wherein the NC BI ClinVar data set stores clinical mutant protein data related to neurodegenerative diseases and cancer, and these two diseases are causally related to abnormal protein phase separation, and the LLPSDB data set stores phase separation data of natural proteins. The data set has features that can distinguish whether the protein has phase separation ability. Further, when the twin-tower graph contrast learning framework is applied to the present application, the corresponding data set A is mutant protein phase separation data. The feasibility of the selection of data sets B and C depends on prior knowledge. It is worth noting that in the clinical mutation data set B, that is, the clinical mutation data set NCBI ClinVar, this application only uses missense mutation data, and the amino acid sequence changes between positive samples and negative samples are small, which makes the original contrast learning loss function almost unable to distinguish the difference between positive and negative samples.

[0091] In the phase separation ability differentiation model, the loss function is modified as follows, and the similarity function is expressed as: sim(x,y)=1-arccos(x T y / ||x||||y||) / π;

[0092] The loss function is defined as:

[0093] Where x is the original protein sequence, x+ is the protein sequence with a benign mutation, x- is the protein sequence with a pathogenic mutation, and |||| is the modulus of the calculated vector. For the auxiliary task on the second auxiliary dataset C, during training, this application expects that the phase separation ability discrimination model can learn the invariance of protein phase separation ability. Therefore, x and x+ represent protein sequences with phase separation ability, that is, x is the original protein sequence with phase separation ability, x+ is the protein sequence with a benign mutation with phase separation ability, and x- represents the protein sequence with a pathogenic mutation without phase separation ability.

[0094] Furthermore, the simultaneously training the phase separation ability differentiation model based on the first auxiliary data set and the second auxiliary data set specifically includes:

[0095] Each time the phase separation ability differentiation model is trained simultaneously using the first auxiliary data set and the second auxiliary data set, obtaining a first gradient and a second gradient of a previous training process;

[0096] Adjusting the second gradient in the current training process according to the first gradient and the second gradient in the previous training process to obtain a current second gradient;

[0097] Based on the first gradient, the phase separation ability differentiation model is trained using the first auxiliary data set, and based on the current second gradient, the phase separation ability differentiation model is simultaneously trained using the second auxiliary data set.

[0098] In order to ensure the performance of the model, the present application adjusts the learning rate, i.e., the gradient, of the training process of the first auxiliary dataset and the second auxiliary dataset. The gradients of the training process corresponding to the first auxiliary dataset and the second auxiliary dataset in the previous training process are respectively expressed as grad B and grad C , then the second gradient grad of the training process corresponding to the second auxiliary data set in the current training process C1 It can be calculated by the following formula:

[0099] where τ is the proportionality constant, which is set to 0.05 in this application.

[0100] During the training process, when a preset training condition is reached, the training of the phase separation ability differentiation model is stopped, and the phase separation ability differentiation model obtained by the last training is used as the trained phase separation ability differentiation model.

[0101] The present application further describes the process framework of the protein phase separation behavior processing method based on machine learning through Figure 2. As shown in Figure 2, the protein phase separation behavior processing method based on machine learning is referred to as PScalpel. PScalpel uses BetaFold to quickly calculate the three-dimensional structural feature map of the protein, and uses PSDM as the fitness function of the genetic algorithm GA to process the three-dimensional structural feature map as well as the amino acid sequence and the mutant amino acid sequence, thereby recommending protein mutants that meet the requirements, wherein T is used to 3The trained PSDM model obtained by the GCL method. This application further describes the workflow of the protein phase separation behavior processing method based on machine learning through Figure 3. Specifically, as shown in Figure 3, the protein phase separation behavior processing method based on machine learning is referred to as PScalpel. PScalpel first uses BetaFold to quickly calculate the three-dimensional structural feature map of the protein, and then uses the PSDM model obtained by training the T3GCL method as the fitness function of the genetic algorithm GA to process the three-dimensional structural feature map as well as the amino acid sequence and the mutant amino acid sequence, thereby recommending protein mutants that meet the requirements.

[0102] As can be seen from the above, compared with the existing technology, the current protein phase separation prediction method has poor time performance and is not suitable for massive search tasks. The present invention uses a trained protein graph structure prediction model to process the mutant amino acid sequence, so that a three-dimensional structural feature map can be quickly obtained. On the basis of being able to obtain the three-dimensional structural feature map of the mutant protein in real time, the separation genetic algorithm in the present invention can conveniently obtain protein phase separation information; at the same time, in order to address the current problem that phase separation cannot be predicted for the new protein after mutation, which is not convenient for actual use by users, the present invention processes the mutant protein through a separation genetic algorithm based on a phase separation ability differentiation model to construct a fitness function, which can obtain phase separation result information of the mutant amino acid sequence, so that the phase separation of the new protein after mutation can be quickly predicted, which facilitates tracking of the decisive sequence perturbations related to phase separation, thereby being able to change the protein phase separation behavior in a predictable manner.

[0103] Exemplary devices

[0104] As shown in FIG7 , corresponding to the above-mentioned protein phase separation behavior processing method based on machine learning, an embodiment of the present invention further provides a protein phase separation behavior processing system based on machine learning, and the above-mentioned protein phase separation behavior processing system based on machine learning includes:

[0105] The data acquisition module 71 is used to obtain the amino acid sequence of the target protein and obtain the mutant amino acid sequence according to the preset site;

[0106] A feature generation module 72 is configured to input the mutant amino acid sequence into a trained protein graph structure prediction model and output a three-dimensional structural feature graph of the mutant amino acid sequence;

[0107] The result generation module 73 is used to input the three-dimensional structural feature map, the amino acid sequence and the mutant amino acid sequence into the separation genetic algorithm, and output the phase separation result information of the mutant amino acid sequence, wherein the fitness function of the separation genetic algorithm is constructed using the trained phase separation ability differentiation model.

[0108] It should be noted that the specific structure and implementation of the above-mentioned protein phase separation behavior processing system based on machine learning and its various modules or units can refer to the corresponding description in the above-mentioned method embodiment, and will not be repeated here.

[0109] It should be noted that the division method of each module of the above-mentioned protein phase separation behavior processing system based on machine learning is not unique and is not used as a specific limitation here.

[0110] Based on the above embodiment, the present invention also provides an intelligent terminal, whose principle block diagram can be shown in Figure 8. The above intelligent terminal includes a processor 10, a memory 20, a network interface and a display 30 connected via a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a protein phase separation behavior processing program 40 based on machine learning. The internal memory provides an environment for the operation of the operating system and the protein phase separation behavior processing program based on machine learning in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the protein phase separation behavior processing program based on machine learning is executed by the processor, the steps of any one of the above-mentioned protein phase separation behavior processing methods based on machine learning are implemented. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen.

[0111] Those skilled in the art will understand that the principle block diagram shown in Figure 8 is only a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the smart terminal to which the solution of the present invention is applied. The specific smart terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0112] In one embodiment, a smart terminal is provided, comprising a memory, a processor, and a machine learning-based protein phase separation behavior processing program stored in the memory and executable on the processor. When the machine learning-based protein phase separation behavior processing program is executed by the processor, the program implements the steps of any one of the machine learning-based protein phase separation behavior processing methods provided in the embodiments of the present invention.

[0113] An embodiment of the present invention also provides a computer-readable storage medium, on which a protein phase separation behavior processing program based on machine learning is stored. When the protein phase separation behavior processing program based on machine learning is executed by a processor, the steps of any one of the protein phase separation behavior processing methods based on machine learning provided in the embodiments of the present invention are implemented.

[0114] It should be understood that the sequence numbers of the steps in the above embodiments do not imply a specific order of execution; the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0115] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0116] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0117] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0118] In the embodiments provided herein, it should be understood that the disclosed systems / terminal devices and methods may be implemented in other ways. For example, the system / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units described above is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or omitting or not implementing certain features.

[0119] If the above-mentioned integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The above-mentioned computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the above-mentioned computer program includes computer program code, and the above-mentioned computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The above-mentioned computer-readable medium may include: any entity or device that can carry the above-mentioned computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0120] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for processing protein phase separation behavior based on machine learning, characterized in that: The protein phase separation behavior processing method based on machine learning includes: Obtaining the amino acid sequence of the target protein, and obtaining the mutant amino acid sequence according to the preset site; Inputting the mutant amino acid sequence into a trained protein graph structure prediction model, and outputting a three-dimensional structural feature graph of the mutant amino acid sequence; The three-dimensional structural feature map, the amino acid sequence and the mutant amino acid sequence are input into a separation genetic algorithm, and the phase separation result information of the mutant amino acid sequence is output, wherein the fitness function of the separation genetic algorithm is constructed using a trained phase separation ability differentiation model.

2. The method for processing protein phase separation behavior based on machine learning according to claim 1, characterized in that: The step of inputting the mutant amino acid sequence into a trained protein graph structure prediction model and outputting a three-dimensional structural feature graph of the mutant amino acid sequence specifically includes: Inputting the mutant amino acid sequence into the trained protein graph structure prediction model; Based on the word vector generation model of the trained protein graph structure prediction model, a first three-dimensional structural feature graph of the mutant amino acid sequence is generated, the first three-dimensional structural feature graph is input into the self-attention mechanism for encoding, and the three-dimensional structural feature graph is output.

3. The method for processing protein phase separation behavior based on machine learning according to claim 1, characterized in that: The protein graph structure prediction model is trained according to the protein structure prediction results in a preset database to obtain the trained protein graph structure prediction model.

4. The method for processing protein phase separation behavior based on machine learning according to claim 1, characterized in that: The step of inputting the three-dimensional structural feature graph, the amino acid sequence and the mutant amino acid sequence into a separation genetic algorithm and outputting phase separation result information of the mutant amino acid sequence specifically includes: Inputting the three-dimensional structural feature map, the amino acid sequence and the mutant amino acid sequence into the separation genetic algorithm; The separation genetic algorithm generates phase separation ability evaluation information according to the fitness function, and obtains the phase separation result information according to the phase separation ability evaluation information.

5. The method for processing protein phase separation behavior based on machine learning according to claim 1, characterized in that: The fitness function of the separation genetic algorithm is constructed by using the trained phase separation ability differentiation model, which specifically includes: The phase separation ability differentiation model is trained by using a dual-tower graph comparison learning framework to obtain the trained phase separation ability differentiation model; The trained phase separation ability differentiation model is used to construct the fitness function of the separation genetic algorithm.

6. The method for processing protein phase separation behavior based on machine learning according to claim 5, characterized in that: The adopting of the dual-tower graph contrast learning framework to train the phase separation ability differentiation model to obtain the trained phase separation ability differentiation model specifically includes: Acquire a preset first auxiliary data set and a second auxiliary data set; Simultaneously training the phase separation ability differentiation model based on the first auxiliary data set and the second auxiliary data set; When the preset training condition is reached, the training of the phase separation ability differentiation model is stopped, and the phase separation ability differentiation model obtained by the last training is used as the trained phase separation ability differentiation model.

7. The method for processing protein phase separation behavior based on machine learning according to claim 6, characterized in that: The simultaneously training the phase separation ability differentiation model according to the first auxiliary data set and the second auxiliary data set specifically includes: Each time the phase separation ability differentiation model is trained simultaneously by the first auxiliary data set and the second auxiliary data set, obtaining a first gradient and a second gradient of a previous training process; According to the first gradient and the second gradient of the previous training process, the second gradient of the current training process is adjusted to obtain the current second gradient; Based on the first gradient, the phase separation ability differentiation model is trained through the first auxiliary data set, and based on the current second gradient, the phase separation ability differentiation model is simultaneously trained through the second auxiliary data set.

8. A protein phase separation behavior processing system based on machine learning, characterized in that: The protein phase separation behavior processing system based on machine learning includes: A data acquisition module is used to obtain the amino acid sequence of the target protein and obtain the mutant amino acid sequence according to the preset site; A feature generation module, used for inputting the mutant amino acid sequence into a trained protein graph structure prediction model, and outputting a three-dimensional structural feature graph of the mutant amino acid sequence; The result generation module is used to input the three-dimensional structural feature map, the amino acid sequence and the mutant amino acid sequence into a separation genetic algorithm, and output the phase separation result information of the mutant amino acid sequence, wherein the fitness function of the separation genetic algorithm is constructed using the trained phase separation ability differentiation model.

9. An intelligent terminal, characterized in that: The intelligent terminal includes a memory, a processor, and a protein phase separation behavior processing program based on machine learning stored in the memory and executable on the processor. When the protein phase separation behavior processing program based on machine learning is executed by the processor, the steps of the protein phase separation behavior processing method based on machine learning as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a protein phase separation behavior processing program based on machine learning, and when the protein phase separation behavior processing program based on machine learning is executed by a processor, the steps of the protein phase separation behavior processing method based on machine learning as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and device for training matching model, predicting amino acid sequence and designing medicine

    CN114283878A

  • Construction method and application of protein phase separation characteristic detection model

    CN116863993A

  • Method and system for predicting phase separation driving residues

    CN117012269A

  • Protein phase separation behavior processing method and system based on machine learning

    CN117894370A

  • System for identifying and developing food ingredients from natural sources by machine learning and database mining combined with empirical testing for a target function

    US20220104515A1

Cited By

  • Protein phase separation characteristic prediction method based on artificial intelligence

    CN121789772A