Disease marker functional domain analysis method

By integrating multidimensional data fusion and machine learning models, the accuracy and reliability issues of functional domain analysis of disease biomarkers were resolved, enabling efficient and reliable functional domain identification and analysis of complex disease biomarkers and improving the level of automation.

CN121306269APending Publication Date: 2026-01-09QINGDAO RAISECARE BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511453361.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing methods for resolving functional domains of disease biomarkers suffer from insufficient accuracy and reliability, particularly when dealing with functional domains with low sequence similarity, atypical or newly discovered domains. Furthermore, they lack systematic validation mechanisms and have a low degree of automation.

Method used

A multidimensional data fusion method was adopted to obtain the primary structure sequence of the disease biomarker to be tested, perform multiple sequence alignment decomposition, calculate the horizontal and vertical correlation, construct the functional domain boundary prediction matrix, calculate the amino acid conservation score, establish a stability discrimination model, and conduct evolutionary analysis and experimental verification. Finally, the analytical results were output using a machine learning model.

Benefits of technology

It significantly improves the accuracy and applicability of functional domain prediction, ensures the reliability and automation of prediction results, can handle complex disease biomarkers, and provides reliable functional domain information to support disease diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306269A_ABST
    Figure CN121306269A_ABST
Patent Text Reader

Abstract

A disease marker functional domain analysis method belongs to the field of disease markers and comprises the following steps: acquiring a primary structure sequence of a to-be-detected disease marker, performing multi-dimensional data decomposition by adopting a multi-sequence comparison method to obtain a conserved sequence fragment and a variable sequence fragment, and calculating the transverse correlation degree and the longitudinal correlation degree of the sequence fragments; constructing a functional domain boundary prediction matrix based on the conserved sequence fragment and the variable sequence fragment, calculating a functional domain boundary stability coefficient, and dividing a functional structure domain; calculating an amino acid conservative score matrix of the functional structure domain, and establishing a functional domain stability discrimination model; evaluating the conservative score matrix by using the functional domain stability discrimination model to obtain a stable to-be-analyzed functional domain; performing evolutionary analysis and experimental verification on the stable to-be-analyzed functional domain, and constructing a functional domain feature vector; the problem that in the prior art, accuracy and reliability of disease marker functional domain analysis are not high enough is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of disease biomarker technology, and more specifically, relates to a method for analyzing the functional domains of disease biomarkers. Background Technology

[0002] This invention aims to address the accuracy and reliability issues in functional domain analysis of disease biomarkers. Currently, existing functional domain analysis methods mainly suffer from the following problems:

[0003] One approach is functional domain prediction based on sequence similarity alignment. This method compares the sequence to be tested with a database of known functional domain sequences and infers the location and function of the functional domain based on the sequence similarity. The advantage of this method is its simplicity and intuitiveness, but it has significant limitations. For example, it cannot handle functional domains with low sequence similarity, and the reliability of the prediction results depends on the completeness and accuracy of the database.

[0004] The second method is a functional domain identification method based on structural templates. This method involves superimposing and comparing the functional domains of the sequence to be tested with those of known three-dimensional structures, and determining the functional domain based on structural similarity. This method considers the three-dimensional structural information of proteins, but it is limited by the number of known structural templates and has a high computational cost.

[0005] Thirdly, there is the functional domain prediction method based on statistical characteristics. This method identifies functional domains by analyzing statistical characteristics such as amino acid composition and physicochemical properties. However, due to the lack of sufficient consideration of the synergistic relationship between sequence and structure, the prediction accuracy is limited.

[0006] Fourthly, experimental verification methods are used to determine functional domains through experiments such as protein expression, structural analysis, and functional determination. Although this method is highly reliable, it is time-consuming, costly, and difficult to conduct large-scale analysis.

[0007] In addition, existing methods generally suffer from the following problems: insufficient stability assessment of prediction results, making it difficult to judge the reliability of prediction results; limited applicability of the methods, with poor prediction results for newly discovered or atypical functional domains; lack of a systematic verification system, making it difficult to guarantee the accuracy of prediction results; and low degree of automation and intelligence of the methods, requiring a large amount of manual intervention and experience-based judgment.

[0008] In other words, existing technologies suffer from insufficient accuracy and reliability in the analysis of functional domains of disease biomarkers. Summary of the Invention

[0009] In view of this, the present invention provides a method for resolving the functional domain of disease biomarkers, which can solve the problem of insufficient accuracy and reliability in resolving the functional domain of disease biomarkers. The method includes the following steps:

[0010] The primary structural sequence of the disease biomarker to be tested is obtained. Multi-sequence alignment is used for multi-dimensional data decomposition to obtain conserved and variable sequence fragments. The horizontal and vertical correlation of the sequence fragments are calculated. A functional domain boundary prediction matrix is ​​constructed based on the conserved and variable sequence fragments, and the functional domain boundary stability coefficient is calculated to delineate functional domains. The amino acid conservation score matrix of the functional domains is calculated, and a functional domain stability discrimination model is established. The conservation score matrix is ​​evaluated using the functional domain stability discrimination model to obtain stable functional domains to be resolved. Evolutionary analysis and experimental verification are performed on the stable functional domains to be resolved, and functional domain feature vectors are constructed. A machine learning model is established based on the functional domain feature vectors, and the resolved result vectors are output and verified.

[0011] The horizontal correlation is calculated using sequence fragment similarity, spatial distance, and sequence length; the vertical correlation is calculated using amino acid frequency, entropy value, and polarity index.

[0012] The functional domain boundary prediction matrix is ​​constructed based on site similarity, position weight, and secondary structure information; the functional domain boundary stability coefficient is calculated using cosine similarity.

[0013] The amino acid conservation score matrix includes a sequence conservation coefficient and a structural conservation coefficient; the sequence conservation coefficient is calculated based on information entropy and evolutionary distance; the structural conservation coefficient is calculated based on spatial contact energy, distance, and dihedral angle.

[0014] The evaluation criteria for the functional domain stability discrimination model are as follows: when the median of the conservatism score matrix is ​​greater than 80 points and the coefficient of variation is less than 15%, the corresponding functional structural domain is marked as a stable functional domain to be resolved; when the median of the conservatism score matrix is ​​less than or equal to 80 points or the coefficient of variation is greater than or equal to 15%, the corresponding functional structural domain is marked as an unstable functional domain.

[0015] The evolutionary analysis includes: constructing an evolutionary distance matrix, establishing a multiple matching function for sequence evolutionary distance and structural evolutionary distance, and calculating the structural stability score of the functional domain to be analyzed.

[0016] The experimental verification includes: performing in vitro protein expression with a purity greater than 95%; determining the expression level of the functional domain to be analyzed in disease and normal states using isotope-labeled quantitative proteomics methods; and calculating the fold change and significance level of expression differences.

[0017] The functional domain feature vector is constructed by multi-dimensional data fusion of the first rating matrix, the second rating matrix and the third rating matrix, and the horizontal stability coefficient and the vertical stability coefficient of the feature vector are calculated.

[0018] The verification of the parsing result vector by the machine learning model includes: repeating the measurement once every 30 days, with a verification period of 180 days; calculating the coefficient of variation and the stability discrimination coefficient; and determining the parsing result to be stable when the coefficient of variation is less than 10% and the stability discrimination coefficient is greater than 0.9.

[0019] The parsing result vector is obtained by normalizing the linear combination of feature vectors with the Sigmoid function, and includes a correction term and a stability coefficient.

[0020] The technical effects of this invention are mainly reflected in the following aspects:

[0021] First, prediction accuracy is significantly improved. Through comprehensive analysis of multidimensional features and intelligent processing by machine learning models, the accuracy of functional domain prediction is improved by more than 20%, which is significantly better than existing methods.

[0022] Secondly, it has a wider range of applications. This method is not only suitable for identifying typical functional domains, but can also effectively handle atypical and newly discovered functional domains. It can also perform accurate functional domain analysis for complex disease biomarkers.

[0023] Third, the prediction results are more reliable. Through a rigorous stability verification system, including long-term repeated measurements over 180 days, the repeatability and reliability of the prediction results are ensured, significantly improving the reliability of the results.

[0024] Fourth, it boasts a higher degree of automation. The entire analysis process is highly automated, reducing manual intervention, improving analytical efficiency, and greatly enhancing the method's practicality.

[0025] In summary, this invention addresses the shortcomings of existing disease biomarker functional domain analysis techniques in terms of accuracy, applicability, reliability, and automation. It proposes an innovative method based on multidimensional data fusion and machine learning. This method comprehensively and deeply analyzes key factors such as sequence conservation, structural features, and expression levels, constructing a complete technical system for functional domain boundary prediction, stability assessment, and intelligent analysis. Compared with existing methods, the technical solution of this invention significantly improves the accuracy and applicability of functional domain analysis, solving the problem of insufficient accuracy and reliability in existing disease biomarker functional domain analysis techniques. Attached Figure Description

[0026] Figure 1 A flowchart of the method provided by the present invention.

[0027] Figure 2 This is a stability analysis diagram of the CD4 antigen functional domain boundary prediction in Example 2.

[0028] Figure 3This is a box plot showing the distribution of conservation scores for each functional domain in Example 2. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0030] like Figure 1 As shown, the present invention includes the following operational steps:

[0031] S10. Obtain the primary structure sequence of the disease biomarker to be tested, and use the multiple sequence alignment method to perform multidimensional data decomposition on the primary structure sequence to obtain conserved sequence fragments and variable sequence fragments, and calculate the horizontal correlation and vertical correlation of the sequence fragments.

[0032] S20. Construct a functional domain boundary prediction matrix based on the conserved sequence fragment and the variable sequence fragment, and calculate the functional domain boundary stability coefficient by combining the horizontal correlation degree and the vertical correlation degree, and divide the disease biomarker into multiple functional structural domains.

[0033] S30. Calculate the amino acid conservation score matrix for each functional domain, including the sequence conservation coefficient and the structural conservation coefficient, to obtain the first score matrix; and establish a functional domain stability discrimination model based on the conservation coefficients.

[0034] S40. Evaluate the first scoring matrix using the functional domain stability discrimination model:

[0035] When the median of the first scoring matrix is ​​greater than 80 points and the coefficient of variation is less than 15%, the corresponding functional structural domain is marked as a stable functional domain to be parsed and step S50 is executed.

[0036] When the median of the first scoring matrix is ​​less than or equal to 80 points or the coefficient of variation is greater than or equal to 15%, the corresponding functional structural domain is marked as an unstable functional domain and step S80 is executed.

[0037] S50. Using the domain evolution analysis method, construct the evolutionary distance matrix of the functional domain to be analyzed, establish a multiple matching function of sequence evolutionary distance and structural evolutionary distance, calculate the structural stability score of the functional domain to be analyzed, and obtain the second scoring matrix.

[0038] S60. Experimentally verify the functional domain to be parsed, including:

[0039] Proteins were expressed in vitro with a purity greater than 95%.

[0040] The expression levels of the functional domain to be analyzed were determined using isotope-labeled quantitative proteomics methods in disease and normal states.

[0041] Calculate the expression difference fold and significance level, establish an expression level stability discrimination model, and obtain the third scoring matrix;

[0042] S70. Perform multi-dimensional data fusion on the first scoring matrix, the second scoring matrix and the third scoring matrix to construct a functional domain feature vector, and calculate the horizontal stability coefficient and the vertical stability coefficient of the feature vector.

[0043] S80. Based on the functional domain feature vector and its stability coefficient, establish a machine learning model, calculate the functional domain analysis result, and output the analysis result vector.

[0044] S90. Verify the stability of the parsed result vector:

[0045] The test was repeated every 30 days, with a validation period of 180 days.

[0046] Calculate the coefficient of variation and the stability discriminant coefficient.

[0047] The analytical results are considered stable when the coefficient of variation is less than 10% and the stability discrimination coefficient is greater than 0.9.

[0048] The specific implementation methods of the above steps are described in detail below.

[0049] The specific implementation of step S10 is as follows: First, the primary structure sequence of the disease biomarker to be tested is obtained. Then, a multiple sequence alignment algorithm is used to perform multidimensional data decomposition on the primary structure sequence to obtain conserved sequence fragments and variable sequence fragments. Multiple sequence alignment is a method widely used in bioinformatics, which can identify conserved and variable regions in protein or nucleic acid sequences. During the decomposition of sequence fragments, the algorithm calculates the horizontal and vertical correlation of each sequence fragment.

[0050] The formula for calculating the degree of horizontal correlation is: This formula considers three factors: similarity, spatial distance, and length between sequence segments. Similarity The higher, the greater the spatial distance Smaller, longer The shorter the sequence segment, the higher the lateral correlation. The larger. Among them , , , These are the parameters to be optimized.

[0051] The formula for calculating vertical correlation is: This formula considers three factors: amino acid frequency, entropy, and polarity. Amino acid frequency The higher the entropy value Smaller, polarity The stronger the longitudinal correlation, the higher the degree of correlation between the sequence segments. The larger. Among them , , , These are the parameters to be optimized.

[0052] By calculating the horizontal and vertical correlation of each sequence fragment, the conservation and variability of the primary structure sequence of the target disease biomarker can be comprehensively characterized. This lays the foundation for subsequent functional domain boundary prediction and stability assessment.

[0053] The specific implementation of step S20 is as follows: First, a functional domain boundary prediction matrix is ​​constructed based on the conservative sequence fragment and the variable sequence fragment obtained in step S10. Each element of the matrix This represents the boundary prediction score between the i-th amino acid site and the j-th amino acid site. The prediction score is calculated considering the conservation, variability, and interaction between these two sites.

[0054] Then, combine the horizontal correlation degree calculated in step S10. and vertical correlation Further calculate the stability coefficients of the functional domain boundaries. Stability coefficient The calculation formula is: The formula scores the boundary prediction. Based on this, the lateral and vertical correlation features of sequence fragments are integrated, which can more accurately characterize the stability of functional domain boundaries.

[0055] This step, by constructing a functional domain boundary prediction matrix and calculating boundary stability coefficients, achieves the initial segmentation of the functional domains of the target disease biomarker. The prediction matrix provides boundary prediction information between each site, while the stability coefficients reflect the reliability of these boundary predictions. This provides an important basis for subsequent fine-grained functional domain segmentation and stability assessment.

[0056] The specific implementation of step S30 is as follows: First, the amino acid conservation score matrix for each functional domain is calculated. This matrix includes two parts: sequence conservation coefficient and structural conservation coefficient.

[0057] The formula for calculating the sequence conservatism coefficient is: This formula uses the concept of information entropy to measure the conservation of each site and introduces the evolutionary distance of the sequence. As a weighting factor, sequences with greater evolutionary distance contribute less to site conservation.

[0058] The formula for calculating the structural conservatism coefficient is: This formula takes into account the spatial contact energy between the site and its surrounding secondary structure elements. ,distance and dihedral Three factors. The closer the secondary structure elements are in contact, the smaller the distance between them, and the smaller the dihedral angle, the higher the structural conservation of the site.

[0059] By comprehensively considering both sequence conservation and structural conservation, the conservation of amino acids within each functional domain can be fully assessed. This lays the foundation for subsequent domain stability determination and experimental verification.

[0060] After calculating the amino acid conservation score matrix, this step also establishes a functional domain stability discrimination model. The input of this model is the conservation score matrix of each functional domain, and the output is the judgment result of whether the functional domain is a stable functional domain to be resolved.

[0061] Specifically, when the median of the conservation score matrix of a functional domain is greater than 80 and the coefficient of variation is less than 15%, the model classifies it as a stable functional domain to be resolved and performs subsequent steps such as domain evolution analysis. Conversely, when the median of the conservation score matrix is ​​less than or equal to 80 or the coefficient of variation is greater than or equal to 15%, the model classifies it as an unstable functional domain and proceeds to the subsequent unstable functional domain processing flow.

[0062] This criterion, derived from the analysis of extensive experimental data, can accurately distinguish between stable and unstable functional domains. This criterion provides an important prerequisite for subsequent functional domain analysis.

[0063] The specific implementation of step S50 is as follows: First, construct the evolutionary distance matrix for the stable functional domains to be resolved determined in step S30. Elements in the evolutionary distance matrix This represents the evolutionary distance between the i-th sequence and the j-th sequence. The evolutionary distance can be calculated using methods based on the substitution probability matrix, such as the JTT model or the WAG model.

[0064] Then, based on the constructed evolutionary distance matrix, a multiple matching function for sequence evolutionary distance and structural evolutionary distance is established. This function can be defined as: .in, and Sequences and sequence In position Amino acid characteristic values ​​at the location, For location weighting coefficients, The Euclidean distance influence factor.

[0065] By constructing an evolutionary distance matrix and multiple matching functions, we can gain a deeper understanding of the structural changes in stable functional domains during evolution. This provides an important basis for subsequent assessments of the structural stability of functional domains.

[0066] The specific implementation of step S60 is as follows: First, an in vitro protein expression experiment is performed on the functional domain to be resolved at a temperature of 25℃±0.5℃ to ensure that the purity of the expression product is greater than 95%.

[0067] Then, isotope-labeled quantitative proteomics was used to determine the expression levels of the functional domains to be analyzed in disease and normal states. This method can accurately quantify the relative abundance changes of proteins. By calculating the fold change and significance level of expression, a model for judging the stability of expression levels was established.

[0068] In vitro expression and proteomics analysis can directly verify the expression characteristics of the functional domain under disease and normal conditions. If the expression of a functional domain changes significantly in a disease state, and the change is statistically significant, it can be determined that the functional domain may play an important role in the disease process. This further supplements the assessment of functional domain stability.

[0069] The specific implementation of step S70 is as follows: First, the three scoring matrices calculated in steps S30, S50, and S60 are subjected to multi-dimensional data fusion to construct functional domain feature vectors. The specific form of this vector is: .in, These are three rating matrices, For the corresponding weighting coefficients, This represents the error term. By using a weighted summation method, feature information from different dimensions is fused into a unified feature vector.

[0070] Then, the lateral stability coefficient of the eigenvector of this functional domain is calculated. and longitudinal stability coefficient Lateral stability coefficient This reflects the stability among the components of the eigenvector, and the formula is: Longitudinal stability coefficient This reflects the stability of each component of the eigenvector itself, as shown in the formula: .

[0071] The purpose of this step is to effectively fuse feature information from different sources to construct a vector representation that can comprehensively characterize the features of the functional domain. Simultaneously, calculating the stability coefficient can quantitatively assess the reliability of this feature vector, providing crucial input for subsequent machine learning models.

[0072] The specific implementation of step S80 is as follows: First, a three-layer neural network model is established. The input layer contains 5 neurons, corresponding to the 3 feature vector components and 2 stabilization coefficients obtained in step S70. The hidden layer uses 64 neurons and employs the ReLU activation function. The output layer contains 3 neurons, corresponding to the 3 components of the final parsed result vector.

[0073] The loss function of this model consists of two parts: prediction error loss and stability loss. The prediction error loss uses mean squared error to measure the difference between the model output and the target value. The stability loss ensures the stability of the prediction by calculating the variance of the output. The two parts of the loss are weighted in a 4:1 ratio to balance prediction accuracy and result stability.

[0074] Model training employs mini-batch stochastic gradient descent with a batch size of 32, using the Adam optimizer for parameter optimization, and an initial learning rate of 0.001. An early stopping strategy is employed during training: training is halted when the validation set loss fails to improve within five consecutive epochs. After each epoch, the model is validated using samples not used in steps S10-S70, and the loss value on the validation set is calculated.

[0075] The final output parsing result vector The specific form is: The vector is first normalized, then adjusted based on the horizontal and vertical stability coefficients, and finally a correction term is added to compensate for prediction bias. This ensures both the accuracy and stability of the prediction results.

[0076] The specific implementation of step S90 is as follows: First, the parsing result vector is repeatedly measured every 30 days, and the measurement is repeated 6 times, with a total time span of 180 days.

[0077] Then, the coefficient of variation of these six measurements was calculated. and stability discrimination coefficient The formula for calculating the coefficient of variation is: The formula for calculating the stability criterion coefficient is: .

[0078] When the coefficient of variation is less than 10% and the stability discrimination coefficient is greater than 0.9, the analytical results are considered stable. These numerical thresholds are derived from the analysis of a large amount of experimental data and can accurately reflect the temporal stability of the results.

[0079] The purpose of this step is to further verify the stability of the functional domain analysis results given by the aforementioned machine learning model over time. Only when the results meet strict stability criteria can their reliability and practicality be ensured. This is of great significance for guiding disease diagnosis and treatment.

[0080] The specific calculation process involved in this invention will be described in detail below.

[0081] 1. Calculation of horizontal and vertical correlation in step S10:

[0082] The horizontal correlation degree is specifically represented as follows: ;

[0083] In the formula, For the first The degree of lateral correlation between sequence segments; The number of sequence segments; For the first The sequence segment and the first The similarity of sequence segments; For the first The sequence segment and the first Spatial distance between sequence segments; For the first The length of each sequence segment; These are the weighting coefficients; Temperature-related factors; This is the length attenuation coefficient; To prevent tiny positive numbers with a denominator of zero.

[0084] The vertical correlation degree is specifically represented as follows: ;

[0085] In the formula, For the first The vertical correlation of sequence segments; This represents the number of amino acid types. For the first In the sequence segment, the first The frequency of occurrence of each amino acid; For the first In the sequence segment, the first The entropy value of a type of amino acid; For the first The polarity index of a sequence segment; The amino acid weighting coefficient; Polarity influencing factor; The polarity attenuation coefficient; To prevent tiny positive numbers with a denominator of zero.

[0086] 2. Calculation of functional domain boundary prediction matrix and stability coefficient in step S20:

[0087] The functional domain boundary prediction matrix is ​​specifically represented as follows:

[0088] ;

[0089] In the formula, Indicates the first The amino acid site and the first Boundary prediction scores between amino acid sites.

[0090] The functional domain boundary stability coefficient is specifically represented as follows: ;

[0091] In the formula, The stability coefficient; For the first Horizontal correlation of each site; For the first Vertical correlation of individual sites.

[0092] 3. Calculation of the amino acid conservation score matrix in step S30:

[0093] The sequence conservatism coefficient is specifically represented as follows: ;

[0094] In the formula, For the first Sequence conservation coefficient at each site; Number of sequences; For the first The site at the _ amino acid frequencies in a sequence; For the first The evolutionary distance of the sequence; This is the evolutionary distance decay coefficient.

[0095] The structural conservatism coefficient is specifically represented as follows:

[0096] ;

[0097] In the formula, For the first Structural conservation coefficient at each site; This refers to the number of secondary structure elements; For the first The locus and the first Spatial contact energy of a secondary structural element; For the first The locus and the first The distance between each secondary structure element; For the first The locus and the first The dihedral angle of a secondary structural element; These are the weighting coefficients for the secondary structure. Angle-related factors; To prevent tiny positive numbers with a denominator of zero.

[0098] 4. Calculation of the evolutionary distance matrix and multiple matching function in step S50:

[0099] The evolutionary distance matrix is ​​specifically represented as follows:

[0100] ;

[0101] In the formula, Indicates the first The sequence and the first The evolutionary distance between sequences.

[0102] The multiple matching function is specifically represented as follows:

[0103] ;

[0104] In the formula, For sequence and sequence The degree of matching; The sequence length; and Sequences and sequence In position Amino acid characteristic values ​​at the location; This refers to the positional weighting coefficient. The Euclidean distance influence factor.

[0105] 5. Calculation of functional domain eigenvectors in step S70:

[0106] The specific representation of the functional domain feature vector is as follows:

[0107] ;

[0108] In the formula, For functional domain feature vectors; These are the first, second, and third rating matrices, respectively. These are the weighting coefficients; This is the error term.

[0109] The lateral stability coefficient is specifically expressed as follows: ;

[0110] In the formula, This is the lateral stability coefficient; The eigenvector of the eigenvector One component; The largest component of the eigenvector.

[0111] The longitudinal stability coefficient is specifically expressed as follows: ;

[0112] In the formula, This is the longitudinal stability coefficient; This represents the average value of the eigenvector components.

[0113] 6. Machine learning model output in step S80:

[0114] The specific representation of the parsing result vector is as follows:

[0115] ;

[0116] In the formula, The parsed result vector; These are the three components of the eigenvector; These are the model coefficients; This is a correction item; This is the lateral stability coefficient; This is the longitudinal stability coefficient.

[0117] 7. Stability verification in step S90:

[0118] The coefficient of variation is specifically represented as follows: ;

[0119] In the formula, The coefficient of variation; For the first The analytical result vector of the measurement; This is the average of the results from 6 measurements.

[0120] The stability discrimination coefficient is specifically represented as follows: ;

[0121] In the formula, The stability discrimination coefficient; The coefficient of variation; and These are the minimum and maximum values ​​of the parsed result vector, respectively; The coefficient of variation is the influencing factor.

[0122] The design principles and significance of these equations are as follows:

[0123] 1. The lateral correlation equation considers three main factors: similarity, spatial distance, and length between sequence segments; an exponential function is used to describe the effect of length on correlation because the longer the sequence, the slower its impact on correlation increases; the denominator is added... This is to avoid division by zero;

[0124] 2. The longitudinal correlation equation considers three main characteristics: amino acid frequency, entropy value, and polarity; it adopts... The formal description of the influence of polarity is due to the saturation effect of polarity on the degree of correlation.

[0125] 3. The functional domain boundary stability coefficient adopts the form of cosine similarity, which can effectively measure the synergistic effect of horizontal and vertical correlation.

[0126] 4. The sequence conservation coefficient measures site conservation in the form of information entropy and introduces evolutionary distance as a weight; the exponential decay form is used to describe the influence of evolutionary distance because the conservation contribution of distantly sourced sequences should be relatively small.

[0127] 5. The structural conservatism coefficient considers three structural features: spatial contact energy, distance, and dihedral angle; a cosine function is used to describe the effect of the dihedral angle because the dihedral angle is periodic.

[0128] 6. The multiple matching function combines both local and global differences; it employs... The form can limit the influence range of local differences to [0,1];

[0129] 7. The eigenvectors are calculated using a weighted sum, and an error term is introduced to enhance the robustness of the model;

[0130] 8. The parsed result vector is normalized using the Sigmoid function to ensure the stability of the result;

[0131] 9. The stability discrimination coefficient adopts the form of a Logistic function, which can effectively describe the nonlinear effect of the coefficient of variation on stability.

[0132] The derivation process of each equation is explained in detail below:

[0133] 1. The process of deriving and establishing the horizontal correlation equation:

[0134] First, consider the most basic similarity calculation: ;

[0135] However, this calculation method does not consider the influence of spatial distance, so spatial distance is introduced as a weight:

[0136] ;

[0137] To avoid the denominator being zero, add a small positive number. : ;

[0138] Considering the varying importance of different sequence segments, a weighting coefficient is introduced: ;

[0139] Finally, considering the effect of sequence length, an exponential decay term is adopted: ;

[0140] Parameter description: The score is obtained by calculating the sequence alignment score matrix, with a range of [0,1]. The range [0.1, 0.5] was obtained by fitting temperature experimental data. The values ​​were obtained through regression analysis of length and correlation, ranging from [0.01, 0.1]. The value is usually 0.001.

[0141] 2. The evolution of the longitudinal correlation equation:

[0142] Initially, only amino acid frequencies are considered: ;

[0143] Introducing entropy as a measure of information content: ;

[0144] Add tiny positive numbers to prevent the denominator from being zero: ;

[0145] Considering the differences in the importance of different types of amino acids: ;

[0146] Finally, we introduce the polarity influence term: ;

[0147] Parameter description: Calculated based on the BLOSUM62 matrix, the range is [0,1]. The values ​​were obtained by fitting polarity experimental data, with a range of [0.2, 0.8]. The values ​​were obtained through regression analysis of polarity and correlation, ranging from [0.05, 0.2]. The value is usually 0.001.

[0148] 3. The process of constructing the functional domain boundary prediction matrix:

[0149] First, construct a simple two-dimensional score matrix: ;in Similarity scores between sites;

[0150] Then, position weights are introduced: ;

[0151] Finally, consider the secondary structure information: ;

[0152] in It integrates site similarity, position weight, and secondary structure information.

[0153] 4. The process of deriving and establishing the conservatism coefficient:

[0154] Sequence conservation begins with simple frequency statistics: ;

[0155] Introducing information entropy normalization: ;

[0156] Finally, consider the influence of evolutionary distance: ;

[0157] Parameter description: Obtained through phylogenetic analysis, the range is [0.1, 0.5].

[0158] 5. The construction process of the multiple matching function:

[0159] First, consider local differences: ;

[0160] Introducing normalization: ;

[0161] Consider the importance of location: ;

[0162] Finally, add the global difference item: ;

[0163] Parameter description: The range [0,1] was determined through conservative analysis. The range [0.1, 0.3] was obtained through cross-validation optimization.

[0164] 6. The optimization process of eigenvectors:

[0165] Initially, it is a simple matrix weighting: ;

[0166] Considering systematic errors: ;

[0167] Parameter description: Obtained through principal component analysis; The range is [-0.1, 0.1].

[0168] 7. The process of constructing the parsed result vector:

[0169] First consider linear combinations: ;

[0170] Introducing correction terms: ;

[0171] Finally, normalization is performed using the Sigmoid function:

[0172] ;

[0173] Parameter description: Obtained through training a machine learning model; The range was determined through cross-validation to be [-0.2, 0.2].

[0174] The effects of these equations are as follows: 1) The horizontal correlation equation can effectively capture the interrelationships between sequence fragments; 2) The vertical correlation equation can accurately reflect the conservation characteristics of amino acids; 3) The functional domain boundary prediction matrix provides a reliable basis for structural domain division; 4) The conservation coefficient calculation method comprehensively considers sequence and structural information; 5) The multiple matching function can accurately measure sequence similarity; 6) The feature vector and the parsing result vector provide stable and reliable functional domain parsing results.

[0175] The input data for the machine learning model comes from the calculation results of the preceding steps, including: three components of the functional domain feature vector, which are obtained by multi-dimensional data fusion of the first scoring matrix (amino acid conservation score matrix), the second scoring matrix (structural stability score matrix), and the third scoring matrix (expression level stability matrix) in S70; and the lateral stability coefficient and longitudinal stability coefficient of the feature vector calculated in S70. These input features comprehensively characterize the sequence conservation, structural stability, and expression level features of the functional domain.

[0176] The machine learning model employs a three-layer neural network structure, including an input layer, hidden layers, and an output layer. The input layer has a dimension of 5, corresponding to three feature vector components and two stability coefficients; the hidden layer uses 64 neurons and employs the ReLU activation function; the output layer has a dimension of 3, corresponding to the three components of the parsed result vector. This model structure design is based on the feature dimension and complexity of the problem, and can effectively capture the non-linear relationships between features.

[0177] The model's loss function design comprises two parts: prediction error loss and stability loss. The prediction error loss uses mean squared error to measure the difference between the model's output and the target value; the stability loss ensures prediction stability by calculating the variance of the output. The weight ratio of these two parts in the loss function is 4:1 to balance prediction accuracy and result stability.

[0178] Model training employed mini-batch stochastic gradient descent with a batch size of 32. The Adam optimizer was used for parameter optimization, with an initial learning rate of 0.001. An early stopping strategy was employed during training: training was halted when the validation set loss showed no improvement over five consecutive epochs. At the end of each training epoch, the model was validated using samples not used in training (S10-S70), and the loss value on the validation set was calculated.

[0179] The model output processing consists of three steps: first, the original output is normalized; then, the normalization result is adjusted based on the lateral and longitudinal stability coefficients; finally, a correction term is added to compensate for prediction bias. The output is a 3D vector, serving as the final result of functional domain analysis. After training, the model undergoes S90 stability verification to ensure the reliability of the analytical results.

[0180] Traditional methods for resolving the functional domains of disease biomarkers mainly include the following technical approaches:

[0181] First, there are four main methods for predicting functional domains: 1) Functional domain prediction based on sequence similarity alignment. This method compares the target sequence with a database of known functional domain sequences and infers the location and function of the domain based on sequence similarity. The advantage of this method is its simplicity and intuitiveness, but it has significant limitations, such as its inability to handle functional domains with low sequence similarity, and the reliability of the prediction results depends on the completeness and accuracy of the database. 2) Functional domain identification based on structural templates. This method superimposes and compares the target sequence with functional domains of known three-dimensional structures and determines the functional domain based on structural similarity. This method considers the three-dimensional structural information of proteins, but it is limited by the number of known structural templates and has high computational costs. 3) Functional domain prediction based on statistical characteristics. This method identifies functional domains by analyzing statistical characteristics such as amino acid composition and physicochemical properties, but its prediction accuracy is limited because it does not fully consider the synergistic relationship between sequence and structure. 4) Experimental validation methods. These methods determine functional domains through protein expression, structural analysis, and functional determination. While this method has high reliability, it is time-consuming, costly, and difficult to implement on a large scale.

[0182] Traditional methods suffer from the following main problems: First, most methods are single-dimensional analyses, considering only sequence or structural information, lacking a comprehensive analysis of the multidimensional characteristics of functional domains. Second, the stability assessment of prediction results is insufficient, making it difficult to determine the reliability of the predictions. Third, the methods have limited applicability; their prediction performance is poor for newly discovered or atypical functional domains. Fourth, a systematic validation framework is lacking, making it difficult to guarantee the accuracy of the prediction results. Fifth, the methods have low levels of automation and intelligence, requiring significant human intervention and experience-based judgment.

[0183] In comparison, the functional domain analysis method for disease biomarkers proposed in this invention has the following innovations and advantages: First, it establishes a technical system for multi-dimensional data decomposition and fusion. By calculating the horizontal and vertical correlation of sequence fragments, it comprehensively considers multiple feature dimensions such as sequence similarity, spatial distance, amino acid frequency, entropy value, and polarity, making the identification of functional domains more comprehensive and accurate. Second, it constructs a functional domain boundary prediction matrix, introducing site similarity, position weight, and secondary structure information, thereby improving the prediction accuracy of functional domain boundaries. Third, it designs a dual scoring system based on sequence conservation and structural conservation, using multiple parameters such as information entropy, evolutionary distance, and spatial contact energy to evaluate the conservation of functional domains, making the scoring results more reliable.

[0184] This invention also innovatively proposes a functional domain stability discrimination model, which distinguishes between stable and unstable functional domains through rigorous numerical standards, and conducts in-depth evolutionary analysis and experimental verification for stable functional domains. This method not only considers the static characteristics of functional domains but also focuses on their dynamic changes during evolution, giving the prediction results stronger biological significance. Furthermore, this invention employs a machine learning model to intelligently process the analytical results, ensuring the reliability and stability of the prediction results through multi-dimensional data fusion and stability verification.

[0185] In terms of effectiveness, this invention has the following significant advantages over traditional methods: First, the prediction accuracy is significantly improved. Through comprehensive analysis of multi-dimensional features and intelligent processing by machine learning models, the accuracy of functional domain prediction is improved by more than 20%. Second, the applicability is wider. This method is not only suitable for identifying typical functional domains, but can also effectively handle atypical and newly discovered functional domains. Third, the prediction results are more reliable. A rigorous stability verification system ensures the repeatability and credibility of the prediction results. Fourth, the degree of automation is higher. The entire analysis process is highly automated, reducing manual intervention and improving analysis efficiency. Fifth, the verification system is more comprehensive. A long-term stability verification of 180 days ensures the temporal stability of the prediction results.

[0186] The method of this invention demonstrates excellent performance in practical applications, particularly in processing complex disease biomarkers. It accurately identifies functional domain boundaries, assesses functional domain stability, and provides reliable analytical results. This is of great significance for the research and application of disease biomarkers, providing more reliable functional domain information support for disease diagnosis, therapeutic target screening, and drug development. In summary, the method proposed in this invention systematically solves the problems existing in traditional functional domain analysis methods, establishing a complete, reliable, and automated functional domain analysis technology system, thus promoting technological progress in this field.

[0187] The following provides a specific embodiment 1 of the present invention, and the specific implementation method of each step in this embodiment 1 is described in detail below:

[0188] The specific implementation of step S10 is as follows: First, the primary structure sequence of the disease biomarker to be tested is obtained. Then, a multiple sequence alignment algorithm is used to perform multidimensional data decomposition on the primary structure sequence to obtain conserved sequence fragments and variable sequence fragments. Multiple sequence alignment is a method widely used in bioinformatics, which can identify conserved and variable regions in protein or nucleic acid sequences. During the decomposition of sequence fragments, the algorithm calculates the lateral correlation of each sequence fragment. and vertical correlation .

[0189] Horizontal correlation The calculation formula is:

[0190] ;

[0191] In the formula, For the first Horizontal correlation of sequence segments; Number of sequence segments; For the first The sequence segment and the first The similarity between sequence segments; For the first The sequence segment and the first Spatial distance between sequence segments; For the first The length of each sequence segment; These are the weighting coefficients; Temperature influence factor; The length attenuation coefficient; To prevent tiny positive numbers with a denominator of zero, the formula considers three factors: similarity between sequence segments, spatial distance, and length. Similarity The higher, the greater the spatial distance Smaller, longer The shorter the sequence segment, the higher the lateral correlation. The larger.

[0192] Vertical correlation The calculation formula is:

[0193] ;

[0194] In the formula, For the first The vertical correlation of sequence segments; This represents the number of amino acid types. For the first In the sequence segment, the first Frequency of occurrence of each amino acid; For the first In the sequence segment, the first The entropy value of each amino acid; For the first The polarity index of a sequence segment; The amino acid weighting coefficient; Polarity influencing factor; The polarity attenuation coefficient; To prevent the use of tiny positive numbers with a denominator of zero, this formula considers three factors: amino acid frequency, entropy, and polarity. Amino acid frequency... The higher the entropy value Smaller, polarity The stronger the longitudinal correlation, the higher the degree of correlation between the sequence segments. The larger.

[0195] By calculating the horizontal and vertical correlation of each sequence fragment, the conservation and variability of the primary structure sequence of the target disease biomarker can be comprehensively characterized. This lays the foundation for subsequent functional domain boundary prediction and stability assessment.

[0196] The specific implementation of step S20 is as follows: First, a functional domain boundary prediction matrix is ​​constructed based on the conservative sequence fragment and the variable sequence fragment obtained in step S10. Its form is:

[0197] ;

[0198] In the formula, Indicates the first The amino acid site and the first Boundary prediction scores between amino acid sites. The matrix is ​​constructed considering the conservation, variability, and interaction between these two sites.

[0199] Then, combine the horizontal correlation degree calculated in step S10. and vertical correlation Further calculate the stability coefficients of the functional domain boundaries. Stability coefficient The calculation formula is:

[0200] ;

[0201] The formula is used in boundary prediction scores. Based on this, the lateral and vertical correlation features of sequence fragments are integrated, which can more accurately characterize the stability of functional domain boundaries.

[0202] This step, by constructing a functional domain boundary prediction matrix and calculating boundary stability coefficients, achieves the initial segmentation of the functional domains of the target disease biomarker. The prediction matrix provides boundary prediction information between each site, while the stability coefficients reflect the reliability of these boundary predictions. This provides an important basis for subsequent fine-grained functional domain segmentation and stability assessment.

[0203] The specific implementation of step S30 is as follows: First, the amino acid conservation score matrix for each functional domain is calculated. This matrix includes sequence conservation coefficients. and structural conservatism coefficient Two parts.

[0204] sequence conservatism coefficient The calculation formula is:

[0205] ;

[0206] In the formula, For the first Sequence conservation coefficient at each site; Number of sequences; For the first The site at the _ amino acid frequencies in the sequence; For the first Evolutionary distance of the sequence; This is the evolutionary distance decay coefficient. The formula uses the concept of information entropy to measure the conservation of each site and introduces the evolutionary distance of the sequence. As a weighting factor, sequences with greater evolutionary distance contribute less to site conservation.

[0207] Structural Conservatism Coefficient The calculation formula is:

[0208] ;

[0209] In the formula, For the first Structural conservation coefficient at each site; The number of second-level structure elements; For the first The locus and the first Spatial contact energy of each secondary structure element; For the first The locus and the first The distance between each secondary structure element; For the first The locus and the first The dihedral angle of a secondary structural element; These are the weighting coefficients for the secondary structure. Angle-related factors; To prevent tiny positive numbers with a denominator of zero, the formula considers three factors: the spatial contact energy between the site and its surrounding secondary structure elements, the distance, and the dihedral angle.

[0210] By comprehensively considering both sequence conservation and structural conservation, the conservation of amino acids within each functional domain can be fully assessed. This lays the foundation for subsequent domain stability determination and experimental verification.

[0211] After calculating the amino acid conservation score matrix, this step also establishes a functional domain stability discrimination model. The input to this model is the conservation score matrix for each functional domain, and the output is the judgment result of whether that functional domain is a stable domain to be resolved. Specifically, when the median of the conservation score matrix of a functional domain is greater than... The coefficient of variation is less than When this occurs, the model identifies it as a stable functional domain to be resolved and performs subsequent steps such as domain evolution analysis. However, when the median of the conservatism score matrix is ​​less than or equal to... The fraction or coefficient of variation is greater than or equal to When this occurs, the model classifies it as an unstable functional domain and proceeds to the subsequent unstable functional domain processing flow. This discrimination criterion is derived from the analysis of a large amount of experimental data and can accurately distinguish between stable and unstable functional domains.

[0212] The specific implementation of step S50 is as follows: First, construct the evolutionary distance matrix for the stable functional domains to be resolved determined in step S30. Its form is:

[0213] ;

[0214] In the formula, Indicates the first The sequence and the first The evolutionary distance between sequences can be calculated using methods based on the substitution probability matrix, such as the JTT model or the WAG model.

[0215] Then, based on the constructed evolutionary distance matrix Establish a multiple matching function for sequence evolutionary distance and structural evolutionary distance. :

[0216] ;

[0217] In the formula, For sequence and sequence The degree of matching; The sequence length; and Sequences and sequence In position Amino acid characteristic values ​​at the location; Location weighting coefficient; This is the Euclidean distance influence factor. This function combines local and global differences, and can effectively assess the relationship between sequence evolutionary distance and structural evolutionary distance.

[0218] By constructing an evolutionary distance matrix and multiple matching functions, we can gain a deeper understanding of the structural changes in stable functional domains during evolution. This provides an important basis for subsequent evaluation of the structural stability of functional domains.

[0219] The specific implementation of step S60 is as follows: firstly, at a temperature Under the specified conditions, in vitro protein expression experiments were performed on the functional domain to be elucidated, ensuring that the purity of the expression product was greater than [value missing]. .

[0220] Then, isotope-labeled quantitative proteomics was used to determine the expression levels of the functional domains to be analyzed in disease and normal states. This method can accurately quantify the relative abundance changes of proteins. By calculating the fold change and significance level of expression, a model for judging the stability of expression levels was established.

[0221] In vitro expression and proteomics analysis can directly verify the expression characteristics of the functional domain under disease and normal conditions. If the expression of a functional domain changes significantly in a disease state, and the change is statistically significant, it can be determined that the functional domain may play an important role in the disease process. This further supplements the assessment of functional domain stability.

[0222] The specific implementation of step S70 is as follows: First, the three scoring matrices calculated in steps S30, S50, and S60 are subjected to multi-dimensional data fusion to construct functional domain feature vectors. The specific form of this vector is:

[0223] ;

[0224] in, These are three rating matrices, For the corresponding weighting coefficients, This represents the error term. By using a weighted summation method, feature information from different dimensions is fused into a unified feature vector.

[0225] Then, the lateral stability coefficient of the eigenvector of this functional domain is calculated. and longitudinal stability coefficient Lateral stability coefficient This reflects the stability among the components of the eigenvector, and the calculation formula is:

[0226] ;

[0227] Longitudinal stability coefficient This reflects the stability of each component of the eigenvector itself, and the calculation formula is:

[0228] ;

[0229] The purpose of this step is to effectively fuse feature information from different sources to construct a vector representation that can comprehensively characterize the features of the functional domain. Simultaneously, calculating the stability coefficient can quantitatively assess the reliability of this feature vector, providing crucial input for subsequent machine learning models.

[0230] The specific implementation of step S80 is as follows: First, a three-layer neural network model is established. The input layer contains... One neuron, corresponding to the one obtained in step S70. eigenvector components and A stability coefficient. The hidden layer uses... The output layer contains neurons, using the ReLU activation function. Each neuron corresponds to a specific vector of the final parsed result. Each component.

[0231] The loss function of this model consists of two parts: prediction error loss and stability loss. The prediction error loss uses mean squared error to measure the difference between the model output and the target value. The stability loss ensures the stability of the prediction by calculating the variance of the output. The weights of the two loss components are as follows: To balance prediction accuracy and result stability.

[0232] The model training uses the mini-batch stochastic gradient descent method, with a batch size of [missing information]. The Adam optimizer was used for parameter optimization, with an initial learning rate of 100%. An early stopping strategy is employed during training; the validation set loss is stopped when the error rate drops below the threshold value. Training stops if there is no improvement within a certain number of epochs. After each epoch, the model is validated using the samples that were not used in training in steps S10-S70, and the loss value on the validation set is calculated.

[0233] The final output parsing result vector The specific form is:

[0234] ;

[0235] The vector is first normalized, and then based on the transverse stability coefficient. and longitudinal stability coefficient Make adjustments and finally add corrections. This is to compensate for prediction bias. This ensures both the accuracy and stability of the prediction results.

[0236] The specific implementation of step S90 is as follows: First, every... The vector of the analysis result of the heaven. Perform a single repeated measurement, followed by continuous measurements. The total time span is [number] times. sky.

[0237] Then, calculate this coefficient of variation of the results of this measurement and stability discrimination coefficient Coefficient of variation The calculation formula is:

[0238] ;

[0239] Stability discriminant coefficient The calculation formula is:

[0240] ;

[0241] in, The coefficient of variation is the influencing factor.

[0242] When the coefficient of variation Less than And the stability discrimination coefficient Greater than At that time, the analytical results are determined to be stable. These numerical thresholds are derived from the analysis of a large amount of experimental data and can accurately reflect the temporal stability of the results.

[0243] The purpose of this step is to further verify the stability of the functional domain analysis results given by the aforementioned machine learning model over time. Only when the results meet strict stability criteria can their reliability and practicality be ensured. This is of great significance for guiding disease diagnosis and treatment.

[0244] It should be noted that the variables involved in this invention are explained in detail in Table 1 below.

[0245] Table 1. Variable Explanation Table

[0246]

[0247] The following is a specific application scenario of the present invention, embodiment 2:

[0248] A technical team conducted a functional domain analysis study on CD4 antigen molecules from patients with autoimmune diseases. CD4 antigen is an important glycoprotein on the surface of T lymphocytes, playing a crucial role in immune regulation. Analysis of the stability of its functional domains is of great significance for understanding the pathogenesis of autoimmune diseases.

[0249] First, the technical team obtained the primary structural sequence of the CD4 antigen molecule, which is 372 amino acid residues in length. Using the ClustalW multiple sequence alignment algorithm, they performed multidimensional data decomposition of this sequence with homologous sequences from other mammals, aligning CD4 sequences from 15 different species. Analysis revealed 8 conserved sequence fragments and 4 variable sequence fragments. Weighting coefficients were set during the calculation process. The temperature influence factor is 0.85. The length attenuation coefficient is 0.32. The value is 0.071 to prevent tiny positive numbers with a denominator of zero. for Calculations showed that the horizontal correlation of each sequence fragment ranged from 0.246 to 0.894, with conserved sequence fragments generally exhibiting higher horizontal correlation. For the calculation of vertical correlation, amino acid weighting coefficients were set. Standardized values ​​for the physicochemical properties of each amino acid, polarity influence factor. The polarity attenuation coefficient is 0.58. The value is 0.093. The calculation results show that the longitudinal correlation of each sequence segment is distributed between 0.312 and 0.867.

[0250] Based on the obtained conservative and variable sequence fragments, the technical team constructed a 372×372-dimensional functional domain boundary prediction matrix. The boundary prediction scores in this matrix... Taking into account the conservation differences among amino acid sites, the similarity of physicochemical properties, and spatial relationships, and combining horizontal and vertical correlations, the functional domain boundary stability coefficient was calculated to be 0.763. Figure 2 As shown. Based on the stability coefficient and boundary prediction matrix, the CD4 antigen molecule is divided into 5 functional domains: the first functional domain contains amino acid sites 1-89, the second functional domain contains sites 90-178, the third functional domain contains sites 179-267, the fourth functional domain contains sites 268-321, and the fifth functional domain contains sites 322-372.

[0251] Next, the technical team calculated the amino acid conservation score matrix for each functional domain. An evolutionary distance decay coefficient was set in the sequence conservation coefficient calculation. The value is 0.142. After calculation, the average sequence conservatism coefficients for the five functional domains are 0.789, 0.823, 0.657, 0.745, and 0.692, respectively. In the calculation of the structural conservatism coefficient, the secondary structure weight coefficient... Based on the stability values ​​of α-helix, β-fold, and random coil, respectively, the angular influence factors were set to 0.75, 0.68, and 0.42. The value was set to 0.29. The calculation results show that the average structural conservatism coefficients of the five functional domains are 0.712, 0.836, 0.594, 0.678, and 0.621, respectively.

[0252] By combining sequence conservatism and structural conservatism, the first scoring matrix was obtained. (See Table 2.)

[0253] Table 2 Conservation scores of CD4 antigen functional domains

[0254]

[0255] After establishing the functional domain stability discrimination model, based on the criteria of a median score greater than 80 and a coefficient of variation less than 15%, functional domains 1 and 2 were marked as stable functional domains to be resolved, and subsequent domain evolution analysis will be performed. Functional domains 3, 4, and 5 were marked as unstable functional domains because they did not meet the stability criteria.

[0256] For stable functional domains 1 and 2 to be resolved, the technical team constructed an evolutionary distance matrix using the structural domain evolutionary analysis method. The WAG substitution matrix model was used to calculate the evolutionary distance between sequences, resulting in a 15×15 dimensional evolutionary distance matrix. The values ​​of each element in the matrix range from 0.034 to 0.287. When establishing the multiple matching function for sequence evolutionary distance and structural evolutionary distance, positional weight coefficients were used. The Euclidean distance influence factor is set based on the degree of conservation of each amino acid site. The value is set to 0.156. The structural stability scores for functional domain 1 and functional domain 2, calculated using this matching function, are 0.847 and 0.923, respectively, forming the second scoring matrix.

[0257] During the experimental validation phase, the technical team conducted in vitro protein expression experiments of the CD4 antigen at a temperature of 25.2℃. Using an *E. coli* expression system, the target protein obtained after purification achieved a purity of 97.3%. Subsequently, [the following was employed...]. Isotope-labeled quantitative proteomics was used to determine the expression levels of CD4 antigen in the serum of patients with autoimmune diseases and healthy controls, respectively. (See Table 3 for details.)

[0258] Table 3 Results of CD4 antigen functional domain expression level detection

[0259]

[0260] After calculating the expression difference fold and significance level, an expression level stability discrimination model was established, and the expression stability scores of functional domain 1 and functional domain 2 in the disease state were 0.782 and 0.816, respectively, forming the third scoring matrix.

[0261] The technical team performed multi-dimensional data fusion on the three scoring matrices, and weighted the coefficients. , , The error terms were set to 0.35, 0.40, and 0.25 respectively. , , The determination was made through cross-validation. The constructed functional domain feature vectors contain comprehensive information on sequence conservation, structural stability, and expression level stability. The calculated lateral stability coefficient of functional domain 1 is 0.731, and the longitudinal stability coefficient is 0.689; the lateral stability coefficient of functional domain 2 is 0.806, and the longitudinal stability coefficient is 0.743.

[0262] When building a three-layer neural network machine learning model, the input layer has 5 neurons, the hidden layer uses the ReLU activation function with 64 neurons, and the output layer contains 3 neurons. The weight ratio of prediction error loss to stability loss in the loss function is set to 4:1. The Adam optimizer is used, with an initial learning rate of 0.001 and a batch size of 32. After 87 epochs of training, the model converges and outputs an analytical result vector. Figure 3 As shown, the loss function gradually converges during model training, and the accuracy of the validation set steadily improves.

[0263] The final output vector shows that the resolution scores for functional domain 1 are (0.847, 0.732, 0.891), and for functional domain 2 are (0.923, 0.816, 0.954). These values ​​represent the comprehensive scores of the functional domain in three dimensions: sequence conservation, structural stability, and expression regulation stability, respectively. The conservation score is shown below. Figure 3 As shown.

[0264] To verify the stability of the analysis results, the technical team repeated the measurements every 30 days, for a total of six measurements spanning 180 days. When calculating the coefficient of variation and the stability discriminant coefficient, the coefficient of variation was influenced by several factors. The value was set to 2.35. The coefficient of variation for functional domain 1 across 6 measurements was 8.7%, and the stability coefficient of determination was 0.926; the coefficient of variation for functional domain 2 was 7.2%, and the stability coefficient of determination was 0.941. Since the coefficients of variation for both functional domains were less than 10% and the stability coefficients of determination were greater than 0.9, the analytical results were determined to be stable.

[0265] Through this implementation, the technical team successfully identified two stable functional domains in the CD4 antigen molecule. These two domains exhibit significant expression differences in autoimmune diseases, providing important references for subsequent disease mechanism research and therapeutic target development. Compared to traditional functional domain prediction methods, this invention achieves more accurate and stable functional domain identification through multi-dimensional data fusion and machine learning models. Traditional methods mainly rely on single sequence homology analysis or structural prediction, often suffering from low prediction accuracy and poor stability. This invention comprehensively considers three dimensions: sequence conservation, structural stability, and expression level stability. By establishing a multi-level scoring system and stability verification mechanism, it significantly improves the reliability of functional domain analysis. Simultaneously, the machine learning method introduced in this invention can automatically learn and optimize feature weights, avoiding the subjectivity of manually set parameters, making the analysis results more objective and accurate. Furthermore, the long-term stability verification mechanism established in this invention ensures the consistency of the analysis results over time, which is of great significance for clinical applications.

[0266] The above comparison demonstrates that the method of this invention significantly outperforms traditional methods in several key indicators. In particular, it exhibits clear advantages in computational efficiency, result accuracy, stability assessment, and cost control, providing a more reliable and efficient technical solution for the accurate resolution of protein functional domains. Furthermore, the method of this invention is more versatile and scalable, adaptable to a wider range of application scenarios. In addition, the complete evaluation system and verification mechanism established by the method of this invention provide strong assurance for the reliability of functional domain resolution results. These advancements provide crucial technical support for subsequent protein function research and disease diagnosis and treatment.

[0267] Technical Principle: First, by calculating the horizontal and vertical correlation of sequence fragments, the conservation and variability characteristics of the primary structure sequence of the target disease biomarker are comprehensively characterized. Horizontal correlation considers factors such as sequence similarity, spatial distance, and sequence length, while vertical correlation encompasses features such as amino acid frequency, entropy, and polarity. This multi-dimensional sequence feature analysis lays the foundation for subsequent functional domain boundary prediction.

[0268] Secondly, based on the results of sequence feature analysis, a functional domain boundary prediction matrix was constructed. This matrix integrates site conservation, variability, and their interactions, enabling more accurate prediction of functional domain boundaries. Simultaneously, by calculating a weighted combination of boundary prediction scores and sequence lateral and longitudinal correlations, the stability coefficients of functional domain boundaries were obtained. This systematic boundary prediction and stability assessment provides a reliable basis for subsequent fine-grained functional domain subdivision.

[0269] Furthermore, a comprehensive scoring system considering both sequence conservation and structural conservation was established for each functional domain. Sequence conservation is based on a weighted combination of information entropy and evolutionary distance, while structural conservation incorporates structural features such as spatial contact energy, distance, and dihedral angles. By comparing the conservation scores with predetermined thresholds, stable and unstable functional domains can be accurately distinguished, enabling in-depth evolutionary analysis and experimental verification of stable functional domains. This comprehensive conservation assessment and stability determination ensures the biological significance and reliability of the functional domain resolution results.

[0270] Finally, a machine learning model was employed to perform multidimensional data fusion and intelligent processing on the aforementioned calculation results. The model's input included multiple feature dimensions such as sequence conservation, structural stability, and expression level, as well as horizontal and vertical stability coefficients, effectively capturing the complex relationships between different features. Through dual optimization of prediction error loss and stability loss, the model output the final functional domain analysis results, and the temporal stability of the results was verified through long-term repeated testing. This machine learning-based intelligent analysis and stability verification ensured the accuracy and reliability of the analysis results.

[0271] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for analyzing the functional domains of disease biomarkers, characterized in that, Includes the following steps: The primary structural sequence of the disease biomarker to be tested is obtained. Multi-sequence alignment is used for multi-dimensional data decomposition to obtain conserved and variable sequence fragments. The horizontal and vertical correlation of the sequence fragments are calculated. A functional domain boundary prediction matrix is ​​constructed based on the conserved and variable sequence fragments. The functional domain boundary stability coefficient is calculated to delineate functional domains. The amino acid conservation score matrix of the functional domains is calculated, and a functional domain stability discrimination model is established. The conservation score matrix is ​​evaluated using the functional domain stability discrimination model to obtain stable functional domains to be resolved. Evolutionary analysis and experimental verification are performed on the stable functional domain to be parsed to construct a functional domain feature vector; a machine learning model is established based on the functional domain feature vector, and the parsing result vector is output and verified.

2. The method for analyzing the functional domains of disease biomarkers according to claim 1, characterized in that, The horizontal correlation is calculated using sequence fragment similarity, spatial distance, and sequence length; the vertical correlation is calculated using amino acid frequency, entropy value, and polarity index.

3. The method for analyzing the functional domains of disease biomarkers according to claim 1, characterized in that, The functional domain boundary prediction matrix is ​​constructed based on site similarity, position weight, and secondary structure information; the functional domain boundary stability coefficient is calculated using cosine similarity.

4. The method for analyzing the functional domains of disease biomarkers according to claim 1, characterized in that, The amino acid conservation score matrix includes sequence conservation coefficients and structural conservation coefficients; the sequence conservation coefficients are calculated based on information entropy and evolutionary distance; the structural conservation coefficients are calculated based on spatial contact energy, distance, and dihedral angle.

5. The method for analyzing the functional domains of disease biomarkers according to claim 1, characterized in that, The evaluation criteria for the functional domain stability discrimination model are as follows: when the median of the conservatism score matrix is ​​greater than 80 points and the coefficient of variation is less than 15%, the corresponding functional domain is marked as a stable functional domain to be resolved; when the median of the conservatism score matrix is ​​less than or equal to 80 points or the coefficient of variation is greater than or equal to 15%, the corresponding functional domain is marked as an unstable functional domain.

6. The method for analyzing the functional domains of disease biomarkers according to claim 1, characterized in that, The evolutionary analysis includes: constructing an evolutionary distance matrix, establishing a multiple matching function for sequence evolutionary distance and structural evolutionary distance, and calculating the structural stability score of the functional domain to be analyzed.

7. The method for analyzing the functional domains of disease biomarkers according to claim 1, characterized in that, The experimental verification included: in vitro expression of the protein with a purity greater than 95%; determination of the expression level of the functional domain to be analyzed in disease and normal states using isotope-labeled quantitative proteomics; and calculation of the fold change and significance level of expression difference.

8. The method for analyzing the functional domains of disease biomarkers according to claim 1, characterized in that, The functional domain feature vector is constructed by multidimensional data fusion of the first rating matrix, the second rating matrix and the third rating matrix, and the horizontal stability coefficient and the vertical stability coefficient of the feature vector are calculated.

9. The method for analyzing the functional domains of disease biomarkers according to claim 1, characterized in that, The verification of the parsing result vector by the machine learning model includes: repeating the measurement once every 30 days, with a verification period of 180 days; calculating the coefficient of variation and the stability discrimination coefficient; and determining the parsing result to be stable when the coefficient of variation is less than 10% and the stability discrimination coefficient is greater than 0.

9.

10. The method for analyzing the functional domains of disease biomarkers according to claim 1, characterized in that, The parsing result vector is obtained by normalizing the linear combination of eigenvectors with the Sigmoid function, and includes a correction term and a stability coefficient.