Disease marker structural domain function prediction method
By constructing the domain functional correlation matrix and optimizing model parameters, combined with deep learning and graph neural network, the accuracy and interpretability problems of the functional prediction of disease marker domains in existing methods are solved, and more accurate functional prediction of disease marker is achieved.
Patent Information
- Application Number
- CN202510397106.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
Existing methods for functional prediction of disease marker domains fail to effectively integrate the functional correlation, sequence conservatism and structural stability between domains, resulting in a lack of biological significance and accuracy of the prediction results.
By constructing a domain functional correlation matrix, the proportion coefficient of amino acid conserved sites, the characteristic coefficient of the domain secondary structure and the polarity distribution coefficient of amino acids are calculated, combined with deep learning and graph neural networks, model parameters are optimized to achieve prediction of the domain function of disease marker.
Improve the accuracy and interpretability of the predicted results, better understand the functional mechanism of disease markers, and provide reliable analysis results.
Smart Images

Figure CN120260684A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of disease markers, and more particularly, relates to a method for predicting the function of a disease marker domain. Background Art
[0002] In the fields of biomedical research and clinical diagnosis, the prediction of the function of disease markers is of great significance for disease diagnosis, treatment plan formulation, and drug development. Currently, the mainstream domain function prediction method is an end-to-end prediction method based on deep learning, which directly learns feature representations from raw sequence data using a deep neural network, avoiding the complex feature engineering process. However, there is a prominent problem with existing deep learning-based prediction methods: modeling domains as independent individuals, ignoring the functional relevance between domains, and insufficient consideration of biological prior knowledge such as the conservation and structural stability of domain sequences. This results in the lack of interpretability and biological significance of the prediction results, affecting the accuracy and reliability of the prediction.
[0003] In practical applications, there are often complex functional association networks between domains. For example, some domains may cooperate to complete specific biological functions. In addition, the conservation and structural stability of domain sequences are also important factors affecting their functions. Existing methods cannot effectively integrate these key information, resulting in prediction results that are difficult to meet the needs of biomedical research. Summary of the Invention
[0004] In view of this, the present invention provides a method for predicting the function of a disease marker domain, which can solve the problem that existing methods cannot effectively integrate these key information, resulting in prediction results that are difficult to meet the needs of biomedical research.
[0005] The present invention is implemented as follows:
[0006] The present invention provides a method for predicting the function of a disease marker domain, comprising the following steps: obtaining a target protein sequence and extracting an amino acid composition feature matrix and a physicochemical feature matrix from the domain sequence data, and calculating the amino acid conserved site proportion coefficient; decomposing the feature matrix to obtain a principal component feature matrix and a physicochemical decomposition feature matrix, and calculating the domain secondary structure feature coefficient; calculating a domain functional association matrix based on the matrix; calculating the self-functional contribution value of the domain and the interactive functional contribution value of the domain, and calculating the amino acid polarity distribution coefficient; obtaining a fused feature matrix through processing by a deep learning model and a graph neural network; training to obtain an initial function prediction model; optimizing the model parameters based on the domain sequence feature learning factor, and outputting the prediction result of the disease marker domain function.
[0007] Based on the above technical solutions, the method for predicting the domain function of a disease biomarker of the present invention can be further improved as follows:
[0008] Among them, the steps for obtaining the amino acid composition feature matrix and the physicochemical feature matrix in the domain sequence data include: obtaining the domain sequence data of the target protein sequence through a multiple sequence alignment tool, extracting the amino acid composition feature matrix and the physicochemical feature matrix in the domain sequence data, and calculating the ratio of the number of amino acid conserved sites to the total number of sites in the domain sequence data to obtain the amino acid conserved site proportion coefficient.
[0009] Further, the feature matrix decomposition step includes: performing principal component analysis decomposition on the amino acid composition feature matrix to obtain the amino acid principal component feature matrix, performing singular value decomposition on the physicochemical feature matrix to obtain the physicochemical decomposition feature matrix, and calculating the domain secondary structure feature coefficient based on the α-helix structure tendency and β-sheet structure tendency in the domain sequence data.
[0010] Further, the calculation step of the domain function association matrix includes: calculating the Euclidean distance between each pair of domains based on the amino acid principal component feature matrix and the physicochemical decomposition feature matrix to obtain the domain distance matrix, performing normalization processing on the domain distance matrix to obtain the domain similarity matrix, calculating the functional correlation coefficient between domains based on the domain similarity matrix, and multiplying the functional correlation coefficient by the sequence homology score between domains to obtain the domain function association matrix.
[0011] Further, the calculation step of the functional contribution value includes: calculating the self-functional contribution value of each domain to the target function and the interactive functional contribution value of the domain, where the self-functional contribution value of the domain is obtained based on the primary sequence features of the domain sequence data, the interactive functional contribution value of the domain is obtained based on the domain function association matrix, and calculating the ratio of the number of hydrophilic amino acids to the number of hydrophobic amino acids in the domain sequence data to obtain the amino acid polarity distribution coefficient.
[0012] Further, the deep learning model processing step includes: inputting the domain function association matrix into a deep neural network model, extracting the local features of the domain function association matrix through a convolutional neural network, extracting the long-range dependence features of the domain function association matrix through a long short-term memory network to obtain the domain function feature vector, and performing weighted operations on the domain function feature vector based on the self-functional contribution value of the domain and the interactive functional contribution value of the domain.
[0013] Furthermore, the graph neural network processing step includes: processing the network node correlation matrix based on a graph neural network model, constructing a domain interaction feature matrix, and fusing the weighted domain function feature vector and the domain interaction feature matrix through an attention mechanism to obtain a fusion feature matrix.
[0014] Furthermore, the domain functions of the disease biomarker include one or more of the following functions: domain catalytic activity function, domain ligand binding function, domain protein interaction function, domain transmembrane transport function, and domain signal transduction function.
[0015] Furthermore, the value range of the preset threshold is from 0.001 to 0.1. When the loss function value is less than the preset threshold for 5 consecutive iterative calculations, it is determined that the initial function prediction model reaches the convergence condition.
[0016] Furthermore, the function prediction model optimization step includes: calculating the total domain function contribution degree of each domain feature to the function prediction result, where the total domain function contribution degree is the weighted sum operation result of the domain's own function contribution value and the domain's interaction function contribution value, constructing a domain sequence feature learning factor by linearly combining the amino acid conserved site proportion coefficient, the domain secondary structure feature coefficient, and the amino acid polarity distribution coefficient, and adjusting the feature weight matrix based on the domain sequence feature learning factor.
[0017] Compared with the prior art, a method for predicting the domain functions of a disease biomarker provided by the present invention effectively solves the above technical problems by innovatively introducing a domain function correlation matrix and a domain sequence feature learning factor. This method not only considers the functional correlation relationship between domains but also incorporates biological prior knowledge such as the conservation and structural stability of domain sequences, significantly improving the accuracy and interpretability of the prediction results.
[0018] Specifically, the present invention realizes the quantitative characterization of the functional correlation relationship between domains by constructing a domain function correlation matrix. By calculating the amino acid conserved site proportion coefficient, the domain secondary structure feature coefficient, and the amino acid polarity distribution coefficient, a domain sequence feature learning factor is constructed, effectively capturing the key biological features of the domain sequence. This method combining deep learning and biological prior knowledge enables the model to more accurately understand and predict the functions of domains.
[0019] The present invention also innovatively introduces the concepts of the self - functional contribution value of a domain and the interactive - functional contribution value of domains. By calculating the total functional contribution degree of domains, it realizes the quantitative interpretation of the prediction results. This not only improves the accuracy of prediction but also provides interpretable analysis results for researchers, helping to deeply understand the functional mechanisms of disease markers.
[0020] Therefore, the present invention solves the problem that existing methods cannot effectively integrate these key information, resulting in the prediction results being difficult to meet the requirements of biomedical research. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart of the method provided by the present invention;
[0022] Figure 2 is a scatter plot of the amino acid principal component analysis results in Example 2;
[0023] Figure 3 is a curve graph of the model training loss in Example 2;
[0024] Figure 4 is a radar chart of the functional prediction results in Example 2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0026] As Figure 1 shown, it is a flowchart of a method for predicting the domain function of a disease marker provided by the present invention. This method includes the following steps:
[0027] S10. Obtain the target protein sequence, construct a training data matrix, obtain the domain sequence data of the target protein sequence through a multiple - sequence alignment tool, extract the amino - acid composition feature matrix and the physicochemical feature matrix from the domain sequence data, and calculate the ratio of the number of amino - acid conserved sites to the total number of sites in the domain sequence data to obtain the amino - acid conserved site proportion coefficient;
[0028] S20. Perform principal - component analysis decomposition on the amino - acid composition feature matrix to obtain an amino - acid principal - component feature matrix, perform singular - value decomposition on the physicochemical feature matrix to obtain a physicochemical decomposition feature matrix, and calculate a domain secondary - structure feature coefficient based on the α - helix structure tendency and β - sheet structure tendency in the domain sequence data;
[0029] S30. Based on the amino acid principal component feature matrix and the physicochemical decomposition feature matrix, calculate the Euclidean distance between each pair of domains to obtain a domain distance matrix, normalize the domain distance matrix to obtain a domain similarity matrix, calculate the functional correlation coefficient between domains based on the domain similarity matrix, and multiply the functional correlation coefficient by the sequence homology score between domains to obtain a domain functional association matrix. The value of the domain functional association matrix represents the functional association strength between any two domains.
[0030] S40. Calculate the self-functional contribution value of each domain to the target function and the interactive functional contribution value of the domain, where the self-functional contribution value of the domain is obtained based on the primary sequence features of the domain sequence data, the interactive functional contribution value of the domain is obtained based on the domain functional association matrix, and calculate the ratio of the number of hydrophilic amino acids to the number of hydrophobic amino acids in the domain sequence data to obtain an amino acid polarity distribution coefficient.
[0031] S60. Input the domain functional association matrix into a deep neural network model, extract the local features of the domain functional association matrix through a convolutional neural network, extract the long-range dependence features of the domain functional association matrix through a long short-term memory network to obtain a domain functional feature vector, and perform a weighted operation on the domain functional feature vector based on the self-functional contribution value and the interactive functional contribution value of the domain.
[0032] S60. Process the network node association degree matrix based on a graph neural network model, construct a domain interaction feature matrix, and fuse the weighted domain functional feature vector and the domain interaction feature matrix through an attention mechanism to obtain a fused feature matrix.
[0033] S70. Establish a multi-layer neural network classifier, input the fused feature matrix into the multi-layer neural network classifier for training to obtain an initial function prediction model.
[0034] S80. Calculate the total functional contribution degree of each domain feature to the function prediction result, where the total functional contribution degree of the domain is the result of the weighted sum operation of the self-functional contribution value and the interactive functional contribution value of the domain. Construct a domain sequence feature learning factor from the linear combination of the amino acid conservative site proportion coefficient, the domain secondary structure feature coefficient, and the amino acid polarity distribution coefficient. The domain sequence feature learning factor is used to characterize the conservativeness and structural stability of the domain sequence, and adjust the feature weight matrix based on the domain sequence feature learning factor.
[0035] S90. Calculate the loss function value between the prediction result of the initial functional prediction model and the standard functional annotation data according to the feature weight matrix, and optimize the model parameters of the initial functional prediction model based on the domain sequence feature learning factor, where different weights are assigned to different domains according to the numerical value of the domain sequence feature learning factor. When the loss function value is less than the preset threshold, use the optimized initial functional prediction model to predict the function of the input disease biomarker domain sequence, and output the prediction result of one or more functions among the domain catalytic activity function, domain ligand binding function, domain protein interaction function, domain transmembrane transport function, and domain signal transduction function of the disease biomarker domain.
[0036] Among them, the disease biomarker domain function includes one or more functions among the domain catalytic activity function, domain ligand binding function, domain protein interaction function, domain transmembrane transport function, and domain signal transduction function.
[0037] Among them, the value range of the preset threshold is from 0.001 to 0.1. When the loss function value is less than the preset threshold for 5 consecutive iterative calculations, it is determined that the initial functional prediction model reaches the convergence condition.
[0038] The following describes the specific implementation manners of the above steps in detail:
[0039] The specific implementation manner of step S10 is to obtain the domain sequence data of the target protein sequence through a multiple sequence alignment tool. This step first uses the sequence search tool in the protein sequence database to perform sequence homology search to obtain a sequence set with significant homology to the target protein sequence, where the homology threshold is set to 30%, and the number of searched sequences is not less than 50. Then use the multiple sequence alignment software to perform alignment analysis on the obtained sequence set, and adopt the progressive multiple sequence alignment algorithm. This algorithm first performs pairwise alignment on sequences with higher sequence similarity, and then gradually adds sequences with lower similarity, and finally obtains the complete multiple sequence alignment result. On this basis, extract the amino acid composition features, including the occurrence frequencies of 20 amino acids in the sequence, and construct a 20-dimensional amino acid composition feature vector. At the same time, extract the physicochemical features of amino acids, including properties such as hydrophobicity, polarity, and volume, and construct a physicochemical feature matrix. Finally, by analyzing the conserved sites in the multiple sequence alignment result, calculate the ratio of the number of conserved sites to the total number of sites to obtain the amino acid conserved site proportion coefficient, which reflects the evolutionary conservation degree of the domain sequence. The purpose of this step is to obtain the basic data reflecting the domain sequence features.
[0040] The specific implementation of step S20 is to perform dimensionality reduction on the amino acid composition feature matrix and the physicochemical feature matrix. The amino acid composition feature matrix is decomposed by the principal component analysis method, and the principal components with a cumulative contribution rate reaching 85% are selected as the features after dimensionality reduction. For the physicochemical feature matrix, the singular value decomposition method is used to retain the eigenvectors corresponding to the largest several singular values, and the physicochemical feature matrix after dimensionality reduction is constructed. At the same time, based on the amino acid composition of the domain sequence, the secondary structure features are calculated, and the GOR algorithm is used to predict the tendency of the α-helix and β-sheet structures, and the domain secondary structure feature coefficient is calculated. This coefficient reflects the structural stability of the domain. The purpose of this step is to reduce the dimension of the feature data and extract key feature information.
[0041] The specific implementation of step S30 is to calculate the similarity between domains based on the feature matrix after dimensionality reduction. First, the amino acid principal component feature matrix and the physicochemical decomposition feature matrix are feature fused, and the Euclidean distance is used to calculate the distance between any two domains, and the domain distance matrix is constructed. Then, the maximum-minimum normalization is performed on the distance matrix to obtain the domain similarity matrix with a value range between 0 and 1. Based on this similarity matrix, the Pearson correlation coefficient is used to calculate the functional correlation between domains, and the correlation coefficient is multiplied by the sequence homology score to obtain the domain functional association matrix, where the sequence homology score is calculated using the local alignment algorithm. The purpose of this step is to construct a data matrix reflecting the strength of the functional association between domains.
[0042] The specific implementation of step S40 is to calculate the functional contribution value of the domain. First, based on the primary sequence features of the domain, including sequence length, amino acid composition, physicochemical properties, etc., the support vector regression model is used to calculate the functional contribution value of the domain itself. Then, based on the domain functional association matrix, the random walk algorithm is used to calculate the interactive functional contribution value between domains, where the number of steps of the random walk is set to 100 steps. At the same time, the number of hydrophilic and hydrophobic amino acids in the domain sequence is counted, and their ratio is calculated to obtain the amino acid polarity distribution coefficient, which reflects the hydrophobic property of the domain. The purpose of this step is to quantify the contribution degree of the domain to the target function.
[0043] The specific implementation of step S50 is to use a deep learning model to extract domain functional features. First, the domain functional association matrix is input into a convolutional neural network, and a 3×3 convolutional kernel is used to extract local features. The number of convolutional layers is 3, and the number of convolutional kernels in each layer is 32, 64, and 128 respectively. Then, a long short-term memory network is used to extract long-range dependence features, with the number of network units set to 256 and the time step set to the sequence length. The output features of the two networks are concatenated to obtain a domain functional feature vector. Finally, the feature vector is weighted based on the domain's own functional contribution value and interactive functional contribution value calculated previously, and the weight coefficient is determined through cross-validation. The purpose of this step is to use a deep learning model to extract high-level functional features.
[0044] The specific implementation of step S60 is to use a graph neural network model to process the interaction relationship between domains. First, a graph structure is constructed based on the domain functional association matrix, with each domain as a node in the graph and the association strength as the weight of the edge. Then, a graph convolutional network is used to update the node features, with the number of network layers set to 2 and the output dimension of each layer set to 128 dimensions. Next, an attention mechanism is used to fuse the weighted domain functional feature vector with the domain interaction features extracted by the graph neural network, and the attention weights are learned through the backpropagation algorithm. The purpose of this step is to integrate the functional features and interaction features of the domains.
[0045] The specific implementation of step S70 is to construct a multi-layer neural network classifier. This classifier contains 3 hidden layers, with the number of neurons in each layer being 512, 256, and 128 respectively, and the activation function uses the rectified linear unit function. During the training process, batch normalization technology is used to improve the convergence speed and generalization ability of the model, and the batch size is set to 64. At the same time, weight decay and dropout techniques are used to prevent the model from overfitting, where the weight decay coefficient is set to 0.0001 and the dropout probability is set to 0.5. The purpose of this step is to establish an initial functional prediction model.
[0046] The specific implementation of step S80 is to calculate feature importance and adjust model parameters. First, the domain's own functional contribution value and interactive functional contribution value are weighted and summed to obtain the total functional contribution degree of the domain, and the weight coefficient is determined through grid search. Then, the amino acid conservation site proportion coefficient, domain secondary structure feature coefficient, and amino acid polarity distribution coefficient are linearly combined to construct a domain sequence feature learning factor, and the combination coefficients are optimized through the least squares method. Based on this learning factor, the feature weight matrix of the model is adjusted, and the adjustment amplitude is proportional to the value of the learning factor. The purpose of this step is to optimize the model parameters and improve the prediction accuracy.
[0047] The specific implementation of step S90 is model training and function prediction. First, the cross-entropy loss function is used to calculate the difference between the model prediction result and the standard functional annotation data, and a preset threshold is set to 0.01. When the loss function value is less than this threshold for 5 consecutive iterative calculations, it is considered that the model has converged. Then, different weights are assigned to different domains based on the domain sequence feature learning factor, and the weight value is positively correlated with the learning factor. For the input domain sequence of the disease biomarker, the model first extracts its feature vector, and then calculates the probability distribution of different functional categories through the trained neural network. Specifically, for the catalytic activity function of the domain, the conservation and spatial conformation of functional groups in the sequence are mainly examined; for the ligand-binding function of the domain, the physicochemical properties and conformational flexibility of the binding site are mainly analyzed; for the protein-protein interaction function of the domain, the distribution characteristics of interface residues and surface charge distribution are concerned; for the transmembrane transport function of the domain, the hydrophobicity and topology of the transmembrane segment are analyzed; for the signal transduction function of the domain, the characteristics of the signal peptide sequence and regulatory sites are examined. The model outputs the predicted probability of each function. When the predicted probability of a certain function exceeds 0.5, it is determined that the domain has the corresponding function. During the prediction process, the model simultaneously considers sequence features, structural features, and evolutionary information, and can accurately predict the function of the domain. The purpose of this step is to achieve accurate prediction of the function of the disease biomarker domain.
[0048] Among them, the calculation of the amino acid conserved site proportion coefficient in step S10 is specifically expressed as follows:
[0049]
[0050] In the formula, C cons is the amino acid conserved site proportion coefficient; N cons is the number of conserved sites; N total is the total number of sites; this coefficient reflects the conservation degree of the sequence, and the range is between 0 and 1.
[0051] The principal component analysis decomposition in step S20 is specifically expressed as follows:
[0052] X AA = UΣV T + E pca ;
[0053] In the formula, X AA is the original matrix of amino acid composition characteristics; U is the left singular matrix; Σ is the singular value diagonal matrix; V T is the transpose of the right singular matrix; E pca is the principal component analysis error term, and the range is between 0.01 and 0.1.
[0054] The singular value decomposition of the physicochemical feature matrix is specifically expressed as follows:
[0055] X pc = PΛQ T + E svd ;
[0056] Wherein, X pc is the original matrix of physicochemical characteristics; P is the left singular vector matrix; Λ is the singular value diagonal matrix; Q T is the transpose of the right singular vector matrix; E svd is the singular value decomposition error term, ranging from 0.01 to 0.1.
[0057] The calculation of the secondary structure characteristic coefficient of the domain is specifically expressed as follows:
[0058]
[0059] Wherein, C ss is the secondary structure characteristic coefficient; N α is the number of α - helix structures; N β is the number of β - sheet structures; w α and w β are the weight coefficients of α - helix and β - sheet respectively, and the default values are both 0.5; E ss is the error term, ranging from 0.01 to 0.1.
[0060] The calculation of the Euclidean distance matrix in step S30 is specifically expressed as follows:
[0061]
[0062] Wherein, D ij is the Euclidean distance between domains i and j; x ik and x jk are the values of domains i and j in the k - th characteristic dimension respectively; n is the total number of characteristic dimensions; E d is the distance calculation error term, ranging from 0.01 to 0.1.
[0063] The normalization of the domain similarity matrix is specifically expressed as follows:
[0064]
[0065] Wherein, S ij is the normalized similarity value; min(D) and max(D) are the minimum and maximum values in the distance matrix respectively; E s is the normalization error term, ranging from 0.01 to 0.1.
[0066] The calculation of the function - related coefficient is specifically expressed as follows:
[0067]
[0068] In the formula, R ij is the functional correlation coefficient between domains i and j; and are the characteristic means of domains i and j respectively; E r is the correlation coefficient calculation error term, ranging from 0.01 to 0.1.
[0069] The calculation of the domain functional association matrix is specifically expressed as follows:
[0070] F ij = R ij ·H ij + E f ;
[0071] In the formula, F ij is the functional association strength between domains i and j; H ij is the sequence homology score; E f is the functional association calculation error term, ranging from 0.01 to 0.1.
[0072] The calculation of the amino acid polarity distribution coefficient in step S40 is specifically expressed as follows:
[0073]
[0074] In the formula, C pol is the amino acid polarity distribution coefficient; N hydro is the number of hydrophilic amino acids; N phob is the number of hydrophobic amino acids; E p is the polarity distribution calculation error term, ranging from 0.01 to 0.1.
[0075] The feature extraction of the convolutional neural network in step S50 is specifically expressed as follows:
[0076]
[0077] In the formula, H l is the output of the l-th layer of convolution; c in is the number of input channels; W i is the weight of the convolution kernel; X i is the input feature map; b is the bias term; f is the activation function; * represents the convolution operation; E cnn is the convolution calculation error term, ranging from 0.001 to 0.01.
[0078] The feature extraction of the long short-term memory network is specifically expressed as follows:
[0079] f t = σ(W f·[h t-1 ,x t +b f )+E f ;
[0080] i t =σ(W i ·[h t-1 ,x t +b i )+E i ;
[0081]
[0082] In the formula, f t is the forget gate; i t is the input gate; is the candidate memory cell; C t is the current memory cell; W f ,W i ,W C are the corresponding weight matrices; b f ,b i ,b C are the corresponding bias vectors; h t-1 is the previous hidden state; x t is the current input; σ is the sigmoid function; represents element-wise multiplication; E f ,E i ,E C ,E lstm are the error terms for each calculation step, and the range is between 0.001 and 0.01.
[0083] The weighted operation is specifically expressed as follows:
[0084] P weighted =αF self +βF inter +E w ;
[0085] In the formula, F weighted is the weighted feature vector; F self is the self-functional contribution vector of the domain; F inter is the interactive functional contribution vector of the domain; α, β are weight coefficients, and α + β = 1; E w is the weighted calculation error term, and the range is between 0.01 and 0.1.
[0086] The graph neural network processing in step S60 is specifically expressed as follows:
[0087]
[0088] Where, H (l) is the node feature matrix of the l-th layer; is the adjacency matrix after adding self-loops; is the corresponding degree matrix; W (l) is the learnable weight matrix; σ is the activation function; E gnn is the error term calculated by the graph neural network, with a range between 0.001 and 0.01.
[0089] The specific representation of the attention mechanism fusion is as follows:
[0090]
[0091] F fusion = MultiHead(F weighted , F inter ) + E fusion ;
[0092] Where, Q, K, V are the query, key, and value matrices respectively; d k is the dimension of the key vector; MultHead represents the multi-head attention operation; F fusion is the fused feature matrix; E att , E fusion are the error terms for attention calculation and feature fusion, with a range between 0.01 and 0.1.
[0093] The specific representation of the total functional contribution degree calculation of the domain in step S80 is as follows:
[0094] C total = γC self + δC inter + E c ;
[0095] Where, C total is the total functional contribution degree of the domain; C self is the self-functional contribution value; C inter is the interactive functional contribution value; γ, δ are weight coefficients, satisfying γ + δ = 1; E c is the contribution degree calculation error term, with a range between 0.01 and 0.1.
[0096] The specific representation of the domain sequence feature learning factor calculation is as follows:
[0097] λ = w1C cons + w2C ss + w3C po l + E λ ;
[0098] Where, λ is the sequence feature learning factor; w1, w2, w3 are combined weights, satisfying Eλ Calculate the error term for the learning factor, with the range between 0.01 and 0.1.
[0099] The adjustment of the feature weight matrix is specifically expressed as follows:
[0100]
[0101] In the formula, W adj is the adjusted feature weight matrix; W orig is the original feature weight matrix; M is the adjustment coefficient matrix; E adj is the weight adjustment error term, with the range between 0.01 and 0.1.
[0102] The calculation of the loss function in step S90 is specifically expressed as follows:
[0103]
[0104] In the formula, L is the total loss value; y i is the true label; is the predicted value; N is the number of samples; W j is the model weight; λ k is the feature learning factor; α, β are regularization coefficients; E L is the loss calculation error term, with the range between 0.001 and 0.01.
[0105] The construction principles and meanings of these equations are explained as follows:
[0106] 1. The proportion coefficient of amino acid conserved sites adopts a simple ratio relationship because conservation is a relative concept and needs to refer to the total number of sites for evaluation;
[0107] 2. Principal component analysis and singular value decomposition adopt the form of matrix decomposition, which can effectively reduce the dimension and extract the main feature information. The introduction of the error term takes into account the influence of data noise;
[0108] 3. The secondary structure feature coefficient adopts the form of weighted summation, reflecting the combined contribution of α-helix and β-sheet to the protein structure;
[0109] 4. The Euclidean distance calculation adopts the form of taking the square root of the sum of squares, which can effectively measure the distance between points in a high-dimensional space;
[0110] 5. The convolutional neural network adopts the form of multi-layer convolution, which can effectively extract local feature patterns. Through the sliding of the convolution kernel and the non-linear transformation of the activation function, hierarchical representation of features is achieved;
[0111] 6. The long short-term memory network introduces a gating mechanism, which can effectively process long sequence information. Through the cooperation of the forget gate, input gate, and output gate, it realizes the selective memory and update of historical information;
[0112] 7. The graph neural network adopts a message passing mechanism, which updates the representation of the central node by aggregating the information of neighboring nodes, and can effectively capture the topological relationship between domains;
[0113] 8. The attention mechanism determines the weights of the value vectors by calculating the similarity between the query and the key, realizes the adaptive fusion of different features, and improves the flexibility of feature combination;
[0114] 9. The domain sequence feature learning factor adopts the form of linear combination, comprehensively considering the influences of sequence conservation, structural stability, and polarity distribution, etc.;
[0115] 10. The loss function includes a prediction error term, a weight regularization term, and a feature learning factor regularization term, achieving the balance between the prediction accuracy and generalization ability of the model.
[0116] In step S90, the core principle of being able to output the functions of multiple disease biomarker domains lies in the adoption of a deep learning framework for multi-label classification. This framework realizes the simultaneous prediction of different functions through multiple function prediction branches. First, for each function type, the model establishes dedicated feature extraction modules and prediction modules, which can capture the sequence and structural features related to specific functions. During the training process of the model, a large number of domain sequences with known functions are used as training data. The domains in these sequences may have multiple functions simultaneously. Through the method of multi-label learning, the model learns the correlation relationships between different functions.
[0117] In the prediction of the catalytic activity function of domains, the model mainly focuses on the active site patterns and conserved motifs in the sequence. By analyzing the known catalytically active domains in the training data, the model learns specific amino acid combination patterns, which are usually closely related to enzyme catalytic functions. At the same time, the model also considers the distribution of secondary structure elements, because catalytic activity often requires a specific structural backbone to support the spatial conformation of the active site. Through the multi-layer feature extraction of the deep neural network, the model can identify these complex feature patterns related to catalytic activity.
[0118] For the prediction of the ligand-binding function of domains, the model focused on analyzing the distribution of physicochemical properties on the domain surface. Ligand binding usually occurs in specific pocket regions on the domain surface, which have unique amino acid compositions and spatial arrangements. By learning the characteristics of known ligand-binding domains, the model established a characteristic model of ligand-binding sites. This includes the distribution of hydrophobicity, the distribution of electrostatic potential, and the distribution patterns of hydrogen bond donors and acceptors. At the same time, the model also considered the flexible characteristics of the domain, because ligand binding often requires the domain to have a certain conformational flexibility.
[0119] In the prediction of the protein-protein interaction function of domains, the model analyzed the characteristic distribution of interface residues. Protein interaction interfaces usually have specific amino acid compositions and physicochemical properties. By learning the characteristics of known protein interaction domains, the model established a prediction model of interface characteristics. This includes the distribution of surface hydrophobic patches, the distribution of charged residues, and the geometric characteristics of the interface. Through a deep learning network, the model can capture these complex interface characteristic patterns.
[0120] The prediction of the transmembrane transport function of domains is mainly based on the hydrophobicity analysis of sequences. Transmembrane domains usually have characteristic hydrophobicity patterns, including the arrangement of transmembrane helices and the distribution of residues inside and outside the membrane. By learning the characteristics of known transmembrane transport domains, the model established an identification model of transmembrane segments. At the same time, the model also considered the characteristics of signal peptide sequences and membrane topology, which are crucial for predicting transmembrane transport functions.
[0121] In the prediction of the signal transduction function of domains, the model focused on analyzing the characteristics of regulatory sites and signal sequences. Signal transduction domains usually contain specific phosphorylation sites, protein-binding sites, or other regulatory elements. By learning the sequence characteristics and spatial distribution characteristics of these sites, the model established a prediction model of signal transduction functions. At the same time, the model also considered the dynamic characteristics of the domain, because signal transduction often involves conformational changes and molecular switch mechanisms.
[0122] The model integrated these different types of characteristic information and performed feature extraction and function prediction through a deep learning network. During the prediction process, the model outputs a probability value for each function type. When the probability value exceeds a preset threshold, it is determined that the domain has the corresponding function. This multi-label classification framework allows a domain to have multiple functions simultaneously, which is consistent with the actual biological situation, because many domains do have multifunctional characteristics. The prediction results of the model also considered the cooperative relationships between functions. Certain combinations of functions may be more likely to occur simultaneously, while some functions may be mutually exclusive.
[0123] Through this comprehensive analysis of the deep learning model, step S90 can evaluate the functional characteristics of the domain from multiple perspectives and give reasonable functional prediction results. This prediction method not only considers sequence and structural information but also integrates multi-dimensional features such as evolutionary conservation and physicochemical properties, thus achieving accurate prediction of the function of the disease biomarker domain. The reliability of the prediction results is evaluated through the validation dataset of the model, ensuring the accuracy and reliability of the prediction.
[0124] The following describes the derivation and establishment process of some of the equations.
[0125] 1. Derivation of the amino acid conserved site proportion coefficient (C cons ):
[0126] First, obtain the sequence alignment results through a multiple sequence alignment tool (such as MUSCLE or Clustal Omega);
[0127] Then, according to the alignment results, count the amino acid type distribution at each site;
[0128] For each site, if more than 90% of the sequences have the same amino acid, it is identified as a conserved site;
[0129] Finally, obtain the proportion coefficient by dividing the number of conserved sites by the total number of sites;
[0130] This coefficient reflects the evolutionary conservation degree of the sequence, and the larger the value, the more important the domain.
[0131] 2. Derivation of the principal component analysis decomposition (X AA = U∑V T + E pca ):
[0132] First, construct an amino acid composition feature matrix, where each row represents a domain and each column represents the occurrence frequency of an amino acid;
[0133] Center the matrix: X centered = X AA - mean(X AA );
[0134] Calculate the covariance matrix:
[0135] Solve the eigenvalue equation: (Cov - λI)v = 0;
[0136] Arrange the eigenvalues in descending order to obtain ∑, and the corresponding eigenvectors form U and V;
[0137] Determine the number of principal components through cross-validation, generally retaining the principal components that explain more than 85% of the variance;
[0138] Introduce the error term F pca Consider the information not captured by the principal components.
[0139] 3. Derivation of Singular Value Decomposition of Physicochemical Feature Matrix:
[0140] Construct a physicochemical feature matrix containing 20 physicochemical properties such as hydrophobicity, polarity, volume, etc.;
[0141] Standardize each property:
[0142] Directly perform SVD decomposition: X pc = PΛQ T ;
[0143] Select an appropriate truncation dimension k according to the singular value size;
[0144] P and Q represent the left and right singular vectors respectively, with dimensions m×k and n×k;
[0145] Λ is a diagonal matrix containing the first k largest singular values;
[0146] Error term E svd Reflects the accuracy of the low-rank approximation.
[0147] 4. Derivation of Secondary Structure Feature Coefficient (C ss )
[0148] Use a secondary structure prediction tool (such as PSIPRED) to predict the secondary structure type of each amino acid;
[0149] Count the number of α-helices and β-sheets;
[0150] Weight coefficients w α and w β Are determined through the following steps:
[0151] (1) Collect training data with known structures;
[0152] (2) Use linear regression to determine the optimal weights;
[0153] (3) Verify the rationality of the weights through cross-validation;
[0154] Error term E ss Considers the accuracy of the prediction tool and statistical errors.
[0155] 5. Derivation of Euclidean Distance Matrix (D ij )
[0156] Concatenate the eigenvectors obtained from principal component analysis and singular value decomposition;
[0157] Normalized eigenvector:
[0158] Calculate the Euclidean distance between any two domains;
[0159] Considering the feature importance, introduce a weighting term:
[0160] Weight w k Determined by the feature importance score;
[0161] Error term E d Reflects the uncertainty of distance calculation.
[0162] 6. Derivation of the domain similarity matrix (S ij )
[0163] Perform Min - Max normalization on the distance matrix;
[0164] Introduce a Gaussian kernel function to optimize the similarity calculation:
[0165] σ is the kernel bandwidth parameter, determined by the median estimation method;
[0166] Considering the local structure features, add a locality - sensitive hashing term;
[0167] Error term E s Reflects the precision loss in the normalization process.
[0168] 7. Derivation of the functional correlation coefficient (R ij )
[0169] Based on the definition of the Pearson correlation coefficient;
[0170] Introduce sample weights:
[0171] λ k Is the importance score of the k - th sample;
[0172] Considering the non - linear relationship, add a kernel correlation coefficient term;
[0173] Error term E r Reflects the confidence of the correlation calculation.
[0174] 8. Derivation of the domain - function association matrix (F ij )
[0175] Steps to calculate the sequence homology score H ij :
[0176] (1) Use BLAST for sequence alignment;
[0177] (2) Calculate the alignment coverage:
[0178] (3) Calculate the sequence identity:
[0179] (4) Comprehensive score: H ij = coverage ij × identity ij × (1 - E value );
[0180] Combine the functional correlation coefficient with the homology score: F ij = R ij · H ij ;
[0181] Introduce a non - linear term: F ij = R ij · H ij · (1 + tanh(α· R ij H ij ));
[0182] The error term E f takes into account the uncertainty of homology assessment.
[0183] 9. Derivation of the amino acid polarity distribution coefficient (C pol ):
[0184] Classify according to the physicochemical properties of amino acids: Hydrophilic amino acids: R, K, D, E, N, Q, H; Hydrophobic amino acids: A, V, L, I, P, F, W, M; Neutral amino acids: G, S, T, C, Y;
[0185] Calculate the polarity index:
[0186] Introduce position weights: L is the sequence length;
[0187] Final coefficient:
[0188] Optionally, the relevant parameters of the convolutional neural network feature extraction (H l ):
[0189] Design multi - scale convolutional kernels: where s k is the size of the k - th convolutional kernel, generally taking 3, 5, 7; d is the input feature dimension; Introduce dilated convolution to increase the receptive field: rate l = 2 l; Nonlinear activation: f(x) = max(0, x) + α·min(0, x); α is the negative slope parameter of LeakyReLU; Add residual connection:
[0190] The following details the relevant calculation process of the learning factor:
[0191] 1. Establishment process of the domain sequence feature learning factor (λ):
[0192] First, construct the basic feature vector:
[0193]
[0194] Introduce nonlinear transformation:
[0195]
[0196] Weight optimization objective:
[0197]
[0198] Final learning factor calculation:
[0199]
[0200] In the formula, W transform is the transformation matrix; b transform is the bias vector; β is the coefficient of the nonlinear term; σ is the sigmoid function; E λ is the error term.
[0201] 2. Establishment process of the feature weight matrix adjustment (W adj )
[0202] Construct the adjustment coefficient matrix:
[0203]
[0204] where d ij is the Euclidean distance between features i and j;
[0205] Introduce the adaptive learning rate:
[0206] η ij = η0·(1 + γ·|λ i - λ j |);
[0207] Weight update rule:
[0208]
[0209] Where η0 is the base learning rate; γ is the adaptive coefficient; and t is the number of iterations.
[0210] 3. Establishment process of the loss function (L total )
[0211] Prediction error term:
[0212]
[0213] Feature weight regularization:
[0214]
[0215] Learning factor regularization:
[0216] L factor = β∑ k |λ k - λ k-1 | + μ∑ k |λ k |;
[0217] Smoothness constraint:
[0218]
[0219] Total loss function:
[0220] L total = L pred + L reg + L factor + L smooth + E L ;
[0221] Where α, β, μ, δ are trade-off coefficients; is the gradient of the learning factor.
[0222] 4. The dynamic update mechanism of the learning factor includes:
[0223] Calculating the gradient:
[0224] Momentum term: m t = ρm t-1 + (1 - ρ)g λ ;
[0225] Second moment estimation:
[0226] Bias correction:
[0227] Parameter update:
[0228] Where ρ and θ are the attenuation rates; η is the learning rate; ∈ is the numerical stability term.
[0229] The parameter acquisition method for the above equation is as follows:
[0230] 1. W transform and b transform are obtained by optimizing through backpropagation;
[0231] 2. β is determined through cross-validation and is generally in the range of 0.1 to 0.5;
[0232] 3. The initial value of η0 is set to 0.001 and then dynamically adjusted according to the training process;
[0233] 4. ρ and θ are set to 0.9 and 0.999 respectively, which are the recommended values of the Adam optimizer;
[0234] 5. ∈ is set to 10 -8 to ensure numerical stability.
[0235] The specific effects of the learning factor are as follows:
[0236] 1. The learning factor comprehensively considers sequence conservation, structural stability, and polarity distribution, and can more accurately characterize the characteristics of the domain sequence;
[0237] 2. The dynamic weight adjustment mechanism enables the model to adaptively adjust the importance of different features;
[0238] 3. Multiple regularization constraints ensure the smoothness and sparsity of the learning factor;
[0239] 4. The use of the Adam optimizer improves the convergence speed and stability of the learning factor;
[0240] 5. The loss function design based on the learning factor achieves a balance between prediction accuracy and model complexity;
[0241] 6. The dynamic update mechanism enables the learning factor to be continuously optimized during the training process, improving the adaptability of the model.
[0242] Specifically, the principle of the present invention is: based on deep learning technology, comprehensively extract, high-order fuse, and adaptively optimize the feature information of protein domains, so as to achieve accurate prediction of domain functions. The core technical principle is as follows:
[0243] 1. Comprehensively extract domain feature information. This method first obtains the domain sequence data of the target protein and extracts a feature matrix reflecting amino acid composition and physicochemical properties from it. Through mathematical transformations such as principal component analysis and singular value decomposition, higher-order features reflecting sequence conservation and structural stability are further extracted. At the same time, the tendency features of the domain secondary structure are calculated. The comprehensive acquisition of these feature information lays a foundation for subsequent function prediction.
[0244] 2. Higher-order fusion of domain function relationships. Based on domain features, this method calculates the Euclidean distance, similarity, and functional correlation between any two domains, and constructs a domain function association matrix. Then, through a deep neural network, the local and long-range functional dependence features contained in this matrix are extracted and weighted and fused with the self-function and interactive function contributions of individual domains. Further, this method also uses a graph neural network to capture the topological relationships between domains and adaptively fuse different features through an attention mechanism to generate the final domain function feature representation. This multi-level and multi-modal feature fusion can effectively depict complex domain function relationships.
[0245] 3. Adaptive optimization of the function prediction model. This method introduces a domain sequence feature learning factor, which comprehensively considers characteristics such as sequence conservation, structural stability, and amino acid polarity distribution, and quantifies the importance of different domains in function prediction. Based on this, this method adaptively adjusts the initial function prediction model, improves the prediction performance for known domains, and also enhances the generalization ability for new domains. Through the optimization of the loss function, this method finally obtains a disease biomarker domain function prediction model with excellent performance.
[0246] The following provides a specific Example 1 of the present invention. The specific implementation of each step in this Example 1 is described in detail as follows: The specific implementation of step S10 is to first receive the input of the target protein sequence and construct a training data matrix. This matrix adopts the amino acid primary sequence encoding method, and each amino acid is represented by a 20-dimensional one-hot encoding vector to form a basic feature space. Subsequently, a multiple sequence alignment tool (such as MUSCLE or ClustalW) is used to analyze the target protein sequence, and the evolutionary distance matrix method and progressive alignment strategy are adopted to obtain the domain sequence data of the target protein sequence. Then, an amino acid composition feature matrix is extracted from the domain sequence data, which contains the occurrence frequency and position distribution information of 20 amino acids, and at the same time, a physicochemical feature matrix is extracted, including 20 physicochemical property indexes such as hydrophobicity index, polarity, and volume. Then, the number of amino acid conserved sites in the domain sequence data is counted, the conservation score is calculated based on the evolutionary site entropy and the amino acid substitution matrix BLOSUM62, and according to the formula Calculate the proportion coefficient of amino acid conserved sites, where N cons represents the number of conserved sites, and N total represents the total number of sites. In practice, by default, the determination threshold for conserved sites is set to 0.8, that is, if the identity of a certain site in the multiple sequence alignment result exceeds 80%, it is considered a conserved site. The main purpose of this step is to establish the data basis required for training. Through sequence alignment and feature extraction, multi-dimensional information reflecting the composition and structural characteristics of protein sequences is obtained, providing a reliable feature representation for subsequent function prediction. The feature extraction process uses a method that combines information entropy and substitution matrix, which can consider both the degree of variation of sites and the physicochemical similarity of amino acids simultaneously.
[0247] The specific implementation of step S20 is to first perform principal component analysis on the amino acid composition feature matrix. Using the singular value decomposition algorithm, through the formula X AA = U∪V T + E pca for matrix decomposition, where X AA is the original amino acid composition feature matrix, U is the left singular matrix, ∑ is the diagonal matrix of singular values, V T is the transpose of the right singular matrix, and E pca is the principal component analysis error term. Perform singular value decomposition operation on the physicochemical feature matrix, and obtain the physicochemical decomposition feature matrix according to the formula X pc = PΛQ T + E svd where X pc is the original physicochemical feature matrix, P is the left singular vector matrix, Λ is the diagonal matrix of singular values, Q T is the transpose of the right singular vector matrix, and E svd is the singular value decomposition error term. Based on the domain sequence data, calculate the α-helix structure propensity and β-sheet structure propensity values through the secondary structure prediction algorithm, and calculate the domain secondary structure feature coefficient according to the formula where C ss is the secondary structure feature coefficient, N α is the number of α-helix structures, N β is the number of β-sheet structures, w α and w β are the weight coefficients of α-helix and β-sheet respectively, and the default values are both 0.5. By default, the cumulative explained variance ratio of principal component analysis and singular value decomposition is set above 95% to ensure the integrity of feature information. The main purpose of this step is to extract the main components of sequence composition and structural characteristics through dimensionality reduction and feature transformation, reduce data redundancy and enhance the expression ability of features.
[0248] The specific implementation of step S30 is based on the amino acid principal component feature matrix and the physicochemical decomposition feature matrix. According to the formula calculate the Euclidean distance between domains, where D ij is the Euclidean distance between domains i and j, and x ik and x jk are the values of domains i and j in the k-th feature dimension respectively. Normalize the domain distance matrix and use the formula to obtain the domain similarity matrix. Based on the domain similarity matrix, according to the formula calculate the functional correlation coefficient between domains. Multiply the functional correlation coefficient by the sequence homology score between domains. According to the formula F ij = R ij ·H ij + E f to obtain the domain functional association matrix. By default, set the similarity threshold to 0.7. Domains pairs with values lower than this may have relatively weak functional associations. The main purpose of this step is to quantitatively describe the functional association degree between domains and establish a domain functional relationship network through distance calculation and similarity analysis.
[0249] The specific implementation of step S40 is based on the primary sequence features of the domain sequence data. Use the sequence pattern recognition algorithm to calculate the self-functional contribution value of the domain, and determine the key functional sites through the position weight matrix and sequence conservation analysis. Calculate the interactive functional contribution value of the domain based on the domain functional association matrix, and use the graph theory algorithm to analyze the topological relationship and interaction strength between domains. Count the number of hydrophilic and hydrophobic amino acids in the domain sequence data. According to the formula calculate the amino acid polarity distribution coefficient, where N hydro is the number of hydrophilic amino acids, N hpob is the number of hydrophobic amino acids, and E p is the polarity distribution calculation error term. By default, set the reasonable range of the hydrophilic-hydrophobic ratio between 0.8 and 1.5 to ensure the stability of the domain. The main purpose of this step is to quantify the functional contributions of domains from both the sequence and structure levels, providing an important reference basis for subsequent function prediction.
[0250] The specific implementation of step S50 is to input the domain functional association matrix into a deep neural network model, use a multi-layer convolutional neural network to extract local features, and according to the formula to achieve feature extraction, where H l is the l-th layer convolution output, and W i is the convolution kernel weight. Extract long-range dependence features through the long short-term memory network. According to the formula group f t = σ(W f ·[ht-1 , x t + b f ) + E f , i t = σ(W i ·[h t-1 , x t + b i ) + E i , and realize the extraction of sequence features. Perform weighted fusion on the domain functional feature vectors, and use the attention mechanism to dynamically adjust the feature weights. By default, set the number of convolutional layers to 3 to 5 layers, use 64 to 128 convolutional kernels for each layer, and set the hidden layer dimension of the long short-term memory network to 256. The main purpose of this step is to automatically learn the high-level feature representation of the domain sequence using the deep learning model, and improve the expressive ability and discriminability of the features.
[0251] The specific implementation of step S60 is to process the network node correlation matrix based on the graph neural network model, and perform information transfer and feature aggregation according to the formula , where H (l) is the node feature matrix of the l-th layer. Construct the domain interaction feature matrix, and use the attention mechanism to fuse the weighted domain functional feature vectors and the domain interaction feature matrix. According to the formulas and F fusion = MultiHead(F weighted , F inter ) + E fusion to realize feature fusion. By default, set the number of layers of the graph neural network to 2 to 3 layers, and the number of attention heads to 8. The main purpose of this step is to capture the complex interaction relationships between domains through graph structure modeling and the attention mechanism, and improve the model's understanding ability of domain functions.
[0252] The specific implementation of step S70 is to first construct a multi-layer neural network classifier, and use techniques such as batch normalization and residual connection to improve the training stability and convergence speed of the model. Input the fused feature matrix into the classifier for training, and use the cross-entropy loss function and the Adam optimizer to update the model parameters. By default, set the number of hidden layers to 3 to 5 layers, the number of neurons in each layer to be halved layer by layer from 512, and the dropout rate to 0.5. The main purpose of this step is to establish an initial function prediction model, providing a basis for subsequent model optimization.
[0253] The specific implementation of step S80 is to calculate the total functional contribution degree of the domain, according to the formula C total = γC self + δC inter+E c Perform weighted summation. Linearly combine the amino acid conserved site proportion coefficient, the domain secondary structure feature coefficient, and the amino acid polarity distribution coefficient. According to the formula λ = w1C cons +w2C ss +w3C pol +E λ Construct the domain sequence feature learning factor. Based on this learning factor, adjust the feature weight matrix. According to the formula Realize the adaptive adjustment of weights. By default, set the combined weights w1, w2, and w3 to 0.4, 0.3, and 0.3 respectively. The main purpose of this step is to achieve the adaptive learning of different domain features by the model through the introduction of the feature learning factor.
[0254] The specific implementation of step S90 is to calculate the loss function value between the prediction result of the initial function prediction model and the standard function annotation data according to the feature weight matrix, using the formula Perform the calculation. Optimize the model parameters of the initial function prediction model based on the domain sequence feature learning factor. When the loss function value is less than the preset threshold (by default, set between 0.01 and 0.05), use the optimized model for function prediction. The main purpose of this step is to improve the prediction accuracy of the model through iterative optimization and finally achieve the accurate prediction of the function of the disease biomarker domain.
[0255] To better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention: the domain function prediction of the tumor marker α-hydroxyl glycosidase.
[0256] This embodiment selects α-hydroxylase (AHG) as the research object. This enzyme shows abnormal expression in various tumor tissues and is an important tumor marker. First, obtain the amino acid sequence of human α-hydroxylase. The sequence length is 526 amino acid residues, including a catalytic domain (residues 50 - 320) and a regulatory domain (residues 321 - 480). Through the MUSCLE multiple sequence alignment tool, align the target sequence with 100 homologous sequences in the database to obtain the domain sequence data. The extracted amino acid composition features show that the catalytic domain is rich in cysteine (Cys) and histidine (His), which is consistent with the composition of its catalytic active center. Physicochemical feature analysis shows that the regulatory domain has high hydrophobicity and β-sheet tendency. Through statistical analysis, the number of conserved sites in the catalytic domain is 185, and the total number of sites is 271. Calculate the amino acid conserved site proportion coefficient according to the formula
[0257] Perform principal component analysis on the extracted amino acid composition feature matrix and decompose the matrix through singular value decomposition. The original matrix X AA is an amino acid composition feature matrix of 271×20. After decomposition, X AA = U∑V T + E pca obtains: U is a left singular matrix of 271×6, ∑ is a singular value diagonal matrix of 6×6, V is a right singular matrix of 20×6, where the principal component analysis error term E pca has a value less than 0.001. Select the first 6 principal components with a cumulative contribution rate reaching 95% as the feature representation, as shown in Table 1. The physicochemical feature matrix X pc (271×20) is decomposed by singular value decomposition X pc = PΛQ T + E svd and then retain the first 8 eigenvectors to construct the physicochemical decomposition feature matrix, where the singular value decomposition error term E svd has a value less than 0.0015.
[0258] Table 1 Amino acid principal component analysis results
[0259] Principal component Eigenvalue Contribution rate (%) Cumulative contribution rate (%) PC1 8.426 42.13 42.13 PC2 4.853 24.27 66.40 PC3 2.967 14.83 81.23 PC4 1.584 7.92 89.15 PC5 0.726 3.63 92.78 PC6 0.445 2.23 95.01
[0260] Figure 2 is a scatter plot of the amino acid principal component analysis results, showing the individual contribution rate and cumulative contribution rate from PC1 to PC6. The blue scatter points represent the individual contribution rate of each principal component, and the red solid line represents the cumulative contribution rate. The horizontal axis uses the principal component label with subscript (PC n ), and the vertical axis shows the contribution rate percentage. This figure intuitively shows the degree to which the first six principal components explain the variability of the original data.
[0261] After analyzing through the secondary structure prediction algorithm, substitute the data into the formula to calculate the secondary structure characteristic coefficient of the domain. Among them, the number of α-helix structures N α is 92, the number of β-sheet structures N β is 78, the total number of amino acids N total is 271, the weight coefficients w α and w β both take 0.5, and the error term E ss takes 0.002. Finally, calculate to obtain C ss = 0.314.
[0262] Based on the amino acid principal component feature matrix and the physicochemical decomposition feature matrix, calculate the distance between domains according to the Euclidean distance calculation formula where E dis the distance calculation error term, with a value of 0.001. The calculated domain distance matrix is normalized using the formula to obtain the domain similarity matrix, where E s is the similarity calculation error term, with a value of 0.002. The following Table 2 shows the distance calculation and similarity analysis results between the target domain and the top 5 homologous domains in the database.
[0263] Table 2 Domain distance and similarity analysis results
[0264]
[0265]
[0266] Based on the domain similarity matrix, the formula is used to calculate the functional correlation coefficient, where E r is the correlation coefficient calculation error term, with a value of 0.003. Multiply the functional correlation coefficient by the sequence homology score, and according to the formula F ij = R ij ·H ij + E f to construct the domain functional association matrix, where H ij is the sequence homology score, and E f is the functional association calculation error term, with a value of 0.002.
[0267] Statistical analysis shows that the number of hydrophilic amino acids in the catalytic domain is 156, and the number of hydrophobic amino acids is 115. According to the formula the amino acid polarity distribution coefficient is calculated to be 1.357, where E p is the polarity distribution calculation error term, with a value of 0.001. When constructing the deep neural network model, the formula is used for feature extraction in the convolutional layer, where E cnn is the convolution calculation error term, with a value of 0.003. The calculation process of the long short-term memory network is as follows:
[0268] f t = σ(W f ·[h t-1 , x t + b f )+ E f i t = σ(W i ·[h t-1 , x t + b i )+ E i
[0269] Among them, each error term E f , E i , E C , E lstm is set to 0.002. The model training adopts the following parameter configurations, as shown in Table 3.
[0270] Table 3 Parameter Configuration for Deep Neural Network Model Training
[0271] Parameter type Value Description Learning rate 0.001 Adam optimizer Batch size 32 Training batch Number of training epochs 100 Total number of iterations L2 regularization coefficient 0.0001 Prevent overfitting Momentum coefficient 0.9 Gradient update parameter Number of convolutional kernels 64,96,128 Three - layer increment Number of attention heads 8 Multi - head attention Dropout rate 0.5 Random inactivation ratio
[0272] When processing the network node correlation matrix based on the graph neural network model, the formula is used for information transmission and feature aggregation, where E gnn is the calculation error term of the graph neural network, with a value of 0.002. The feature fusion of the attention mechanism adopts and F fusion = MultiHead(F weighted , F inter ) + E fusion , where E att and E fusion are both set to 0.002.
[0273] Figure 3 is the model training loss curve graph, showing the changing trends of the training loss and the validation loss during the training process. The horizontal axis is the number of training epochs (Nepoch with superscript), and the vertical axis is the value of the loss function (Ltotal with subscript). The green dashed line represents the preset convergence threshold (0.02).
[0274] When calculating the total functional contribution degree of the domain, the formula C total = γC self + δC inter + E c is adopted, where the self-functional contribution weight γ takes 0.6, the interactive functional contribution weight δ takes 0.4, and E c is the total contribution degree calculation error term, with a value of 0.003. The formula for constructing the domain sequence feature learning factor is λ = w1C cons + w2C ss + w3C pol + E λ , and the three combined weights are taken as 0.4, 0.3, and 0.3 respectively. E λ is the learning factor calculation error term, with a value of 0.002.
[0275] The loss function during the model training process is calculated using the formula Among them, α is the L2 regularization coefficient, taking 0.0001, β is the regularization coefficient of the feature learning factor, taking 0.001, and E L is the error term of the loss function calculation, taking 0.001. After 90 rounds of training, the value of the loss function drops to 0.018 and remains stable, meeting the requirement of the preset convergence threshold of 0.02. The functional prediction results of the final model for the catalytic domain of α - hydroxylase glycosidase are shown in Table 4 below.
[0276] Table 4 Functional prediction results and credibility analysis of the catalytic domain
[0277] Function type Prediction probability Confidence level Feature contribution degree Catalytic activity function 0.925 High 0.386 Ligand binding function 0.847 High 0.342 Protein - protein interaction function 0.632 Medium 0.156 Transmembrane transport function 0.156 Low 0.064 Signal transduction function 0.283 Low 0.052
[0278] Figure 4 is the radar chart of the functional prediction results, which shows the prediction probabilities of five functional types in polar coordinate form. Each direction represents a functional type, and the distance from the center represents the prediction probability value.
[0279] The prediction results were analyzed and verified, and the actual functional activity of the catalytic domain was determined by experimental methods. Through enzyme activity determination, ligand - binding experiments, and protein - interaction verification, it was confirmed that the consistency between the prediction results and the experimental data reached 85.6%. Further comparison of this method with existing mainstream prediction methods showed that this method had significant advantages, as shown in Table 5.
[0280] Table 5 Performance comparison and analysis of different prediction methods
[0281] Evaluation index Traditional sequence alignment method Simple machine learning method Method of the present invention Prediction accuracy rate (%) 65.3 72.8 85.6 Computation time (min) 45 28 15 Feature dimension 20 64 256 Generalization ability index 0.682 0.735 0.893 Robustness coefficient 0.625 0.712 0.856
[0282] The technical advantages of this embodiment are mainly reflected in the following aspects: (1) The feature extraction is more comprehensive, and the expression ability is improved by extracting multi - level feature combinations; (2) The domain - association analysis is more accurate, integrating similarity indicators in multiple dimensions of sequence, structure, and function; (3) The prediction model has stronger adaptability, and the dynamic weight adjustment is realized by introducing the feature learning factor; (4) The calculation efficiency is significantly improved, shortening the modeling time by 66.7%. These improvements provide more reliable technical support for the functional prediction research of disease markers.
[0283] It should be noted that the variable explanations involved in the present invention are shown in Table 6 below.
[0284] Table 6 Variable explanation table
[0285]
[0286]
[0287] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. A method for predicting the function of a disease biomarker domain, characterized in that, The steps include: obtaining the target protein sequence and extracting the amino acid composition feature matrix and the physicochemical feature matrix from the domain sequence data, and calculating the amino acid conserved site proportion coefficient; decomposing the feature matrix to obtain the principal component feature matrix and the physicochemical decomposition feature matrix, and calculating the domain secondary structure feature coefficient; calculating the domain functional association matrix based on the matrices; calculating the domain self-functional contribution value and the domain interaction functional contribution value, and calculating the amino acid polarity distribution coefficient; obtaining the fused feature matrix through processing by a deep learning model and a graph neural network; training to obtain an initial functional prediction model; optimizing the model parameters based on the domain sequence feature learning factor, and outputting the disease biomarker domain functional prediction result.
2. The method for predicting the function of a disease marker domain according to claim 1, wherein The steps for obtaining the amino acid composition feature matrix and the physicochemical feature matrix in the domain sequence data include: obtaining the domain sequence data of the target protein sequence through a multiple sequence alignment tool, extracting the amino acid composition feature matrix and the physicochemical feature matrix from the domain sequence data, and calculating the ratio of the number of amino acid conserved sites to the total number of sites in the domain sequence data to obtain the amino acid conserved site proportion coefficient.
3. A method for predicting the function of a disease marker domain according to claim 2, characterized in that The feature matrix decomposition steps include: performing principal component analysis decomposition on the amino acid composition feature matrix to obtain the amino acid principal component feature matrix, performing singular value decomposition on the physicochemical feature matrix to obtain the physicochemical decomposition feature matrix, and calculating the domain secondary structure feature coefficient based on the α-helix structure propensity and β-sheet structure propensity in the domain sequence data.
4. A method for predicting the function of a disease marker domain according to claim 3, wherein The steps for calculating the domain functional association matrix include: calculating the Euclidean distance between each pair of domains based on the amino acid principal component feature matrix and the physicochemical decomposition feature matrix to obtain the domain distance matrix, performing normalization processing on the domain distance matrix to obtain the domain similarity matrix, calculating the functional correlation coefficient between domains based on the domain similarity matrix, and multiplying the functional correlation coefficient by the sequence homology score between domains to obtain the domain functional association matrix.
5. A method for predicting the function of a disease marker domain according to claim 4, characterized in that, The steps for calculating the functional contribution value include: calculating the domain self-functional contribution value and the domain interaction functional contribution value of each domain to the target function, where the domain self-functional contribution value is obtained based on the primary sequence features of the domain sequence data, the domain interaction functional contribution value is obtained based on the domain functional association matrix, and calculating the ratio of the number of hydrophilic amino acids to the number of hydrophobic amino acids in the domain sequence data to obtain the amino acid polarity distribution coefficient.
6. The method for predicting the function of a disease marker domain according to claim 5, wherein The steps for processing by the deep learning model include: inputting the domain functional association matrix into a deep neural network model, extracting the local features of the domain functional association matrix through a convolutional neural network, extracting the long-range dependence features of the domain functional association matrix through a long short-term memory network to obtain the domain functional feature vector, and performing a weighted operation on the domain functional feature vector based on the domain self-functional contribution value and the domain interaction functional contribution value.
7. A method for predicting the function of a disease biomarker domain according to claim 6, characterized in that, The graph neural network processing steps include: processing the network node association degree matrix based on a graph neural network model, constructing a domain interaction feature matrix, and fusing the weighted domain function feature vectors and the domain interaction feature matrix through an attention mechanism to obtain a fusion feature matrix.
8. A method for predicting the function of a disease marker domain according to claim 7, wherein The domain functions of the disease biomarker include one or more of the following functions: domain catalytic activity function, domain ligand binding function, domain protein interaction function, domain transmembrane transport function, and domain signal transduction function.
9. The method for predicting the function of a disease marker domain according to claim 8, wherein The value range of the preset threshold is from 0.001 to 0.
1. When the loss function value is less than the preset threshold for 5 consecutive iterative calculations, it is determined that the initial function prediction model reaches the convergence condition.
10. A method for predicting the function of a disease marker domain according to claim 9, wherein The function prediction model optimization steps include: calculating the total domain function contribution degree of each domain feature to the function prediction result, where the total domain function contribution degree is the weighted sum operation result of the domain's own function contribution value and the domain's interaction function contribution value, constructing a domain sequence feature learning factor from the linear combination of the amino acid conservation site proportion coefficient, the domain secondary structure feature coefficient, and the amino acid polarity distribution coefficient, and adjusting the feature weight matrix based on the domain sequence feature learning factor.