A model for predicting changes in detection signals of proteins after perchloric acid precipitation of serum based on physicochemical properties of proteins, a construction method and applications thereof

By constructing a predictive model based on the physicochemical properties of proteins, the problem of changes in protein detection signals after perchloric acid precipitation was solved, enabling quantitative prediction in the experimental design stage and improving the efficiency and accuracy of serum proteomics research.

CN122435976APending Publication Date: 2026-07-21NATIONAL INSTITUTE OF METROLOGY CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NATIONAL INSTITUTE OF METROLOGY CHINA
Filing Date
2026-06-23
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies struggle to predict the direction and magnitude of protein detection signal changes after perchloric acid precipitation during the experimental design phase. They also lack quantitative prediction methods based on the physicochemical properties of proteins, resulting in researchers lacking quantifiable references when selecting pretreatment strategies. Furthermore, existing models are difficult to adapt into reusable and transferable quantitative models.

Method used

By converting protein amino acid sequences into quantifiable physicochemical properties, and combining feature selection algorithms with machine learning regression models, a predictive model for changes in protein detection signals is constructed, including feature selection and the division of training and validation sets, to achieve quantitative prediction of changes in candidate protein signals.

Benefits of technology

This method enables the prediction of the signal change direction and amplitude of the protein after perchloric acid precipitation based on the amino acid sequence information of the protein to be predicted without conducting actual mass spectrometry experiments. This provides a quantifiable decision-making tool for the design of serum proteomics pretreatment strategies and improves the interpretability and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435976A_ABST
    Figure CN122435976A_ABST
Patent Text Reader

Abstract

The application discloses a model for predicting changes of detection signals of proteins after perchloric acid precipitation of serum based on physicochemical properties of proteins and a construction method and application thereof, and the method comprises the following steps: obtaining detection signal intensities of the same protein in a perchloric acid precipitation treatment group and an untreated group in serum proteomics data through experiments, and calculating a signal change value log2FC as a target variable; calculating physicochemical property characteristics based on a protein amino acid sequence; dividing samples into a training set and a verification set, screening key features on the training set by using a feature screening algorithm, training a machine learning regression model, evaluating model performance on the verification set, and obtaining a prediction model; inputting an identifier of a protein to be predicted, automatically obtaining an amino acid sequence, calculating physicochemical property characteristics, and outputting a log2FC prediction value and a signal change direction determination result through the prediction model. The application can provide a quantitative tool for design of a serum proteomics sample pretreatment strategy and quantitative pre-evaluation of analysis of protein markers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of proteomics data analysis, bioinformatics and machine learning, and in particular to a model, construction method and application for predicting changes in protein detection signals after serum precipitation with perchloric acid based on the physicochemical properties of proteins. Background Technology

[0002] Serum proteomics has important applications in biomarker discovery, disease mechanism research, and protein quantification. However, the abundance dynamic range of human serum proteome is extremely wide, spanning more than ten orders of magnitude. A few high-abundance proteins (such as albumin and immunoglobulins) account for the vast majority of total serum protein, significantly suppressing the detection signal of low-abundance proteins, thus limiting the ability to discover and quantify key biomarkers.

[0003] To reduce the dynamic range of serum proteomics and improve the detectability of low-abundance proteins, researchers have developed various sample pretreatment strategies. Among them, perchloric acid precipitation, as a chemical precipitation strategy, can selectively weaken the interference of some high-abundance proteins, thus increasing the detection chance of some medium- and low-abundance proteins. This method has the advantages of simple operation, low cost, and high throughput, and has been applied in large-scale serum proteomics studies.

[0004] However, the effect of perchloric acid precipitation on the detection signals of different proteins exhibits significant selectivity; the direction and amplitude of signal changes after perchloric acid precipitation of serum vary for different proteins. This selective change is not determined solely by protein abundance but is closely related to the physicochemical properties of proteins, although the underlying mechanisms have not yet been fully elucidated.

[0005] In existing technologies, studies on changes in protein detection signals after perchloric acid precipitation of serum typically rely on post-treatment analysis and interpretation of proteomic data before and after treatment, which has the following shortcomings:

[0006] First, during the experimental design phase, it is difficult to predict the direction and magnitude of the change in the detection signal of a specific protein or a class of candidate biomarkers after perchloric acid precipitation treatment, which leads to a lack of quantifiable reference for researchers when selecting pretreatment strategies.

[0007] Second, existing research is mostly limited to descriptive analysis of specific experimental data, and the model building process lacks a strict division between training and validation sets, making it difficult to form a quantitative model that can be reused and transferred to other protein prediction tasks.

[0008] Third, existing methods lack a technical solution for systematically modeling changes in protein detection signals after perchloric acid precipitation of serum from the perspective of protein physicochemical properties, and therefore cannot provide an effective tool for the pre-evaluation of candidate proteins.

[0009] Machine learning methods can effectively handle high-dimensional nonlinear problems, providing a feasible technical approach for predicting changes in protein detection signals. Therefore, there is an urgent need in this field for a quantitative prediction method based on the physicochemical properties of proteins, so as to pre-evaluate changes in the detection signals of candidate proteins during the experimental design stage. Summary of the Invention

[0010] The purpose of this invention is to provide a model, construction method, and application for predicting changes in protein detection signals after serum precipitation with perchloric acid based on the physicochemical properties of proteins. This invention converts the protein amino acid sequence into quantifiable physicochemical characteristics, combines feature selection algorithms with machine learning regression models, and constructs a predictive model for changes in protein detection signals, achieving quantitative prediction of the direction and magnitude of changes in candidate protein signals.

[0011] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0012] A method for constructing a model based on the physicochemical properties of proteins to predict changes in protein detection signals after perchloric acid precipitation of serum includes the following steps:

[0013] (1) Serum proteomics data after perchloric acid precipitation and serum proteomics data without perchloric acid precipitation were obtained through experiments;

[0014] (2) Match the quantitative intensity of the same protein in the perchloric acid precipitation treatment group and the untreated group, and calculate the target variable log2FC to characterize the change in the detection signal;

[0015] (3) Obtain the amino acid sequence of the protein and calculate the physicochemical properties of the protein based on the amino acid sequence;

[0016] (4) Divide the obtained protein samples into training set and validation set according to a preset ratio. The training set is used for feature selection and model training, and the validation set is used for independent evaluation of model performance.

[0017] (5) The physicochemical properties of the protein are screened on the training set using a feature screening algorithm to obtain a subset of key features after screening;

[0018] (6) Train the machine learning regression model on the training set using the key feature subset and the target variable log2FC, and evaluate the model performance on the validation set to obtain the final prediction model.

[0019] Furthermore, the formula for calculating log2FC is as follows:

[0020]

[0021] Among them, Intensity DepletedIntensity represents the detection signal intensity of the same protein in the perchloric acid precipitation group. NonDepleted The signal intensity of the same protein in the untreated group.

[0022] The physicochemical properties of the protein include at least the positively charged amino acid characteristic PosAA, the negatively charged amino acid characteristic NegAA, and the charge ratio characteristic ChargeRatio, which are calculated using the following formula:

[0023]

[0024]

[0025]

[0026] Where, N R N K N D N E These represent the number of arginine, lysine, aspartic acid, and glutamic acid residues in the protein's amino acid sequence.

[0027] The physicochemical properties of the protein may also include one or more of the following: amino acid number (length), molecular weight (MW), theoretical isoelectric point (pI), stability index (Instability Index), aliphatic index (Aliphatic Index), hydrophobicity index (GRAVY), and aromaticity.

[0028] The physicochemical properties of the protein were obtained using the ExPASy ProtParam tool or an equivalent physicochemical property calculation program based on amino acid sequences.

[0029] The serum proteomics data are derived from one or more serum sample cohorts, which include standard substance serum cohorts, clinical serum cohorts, or combinations thereof.

[0030] When the data comes from multiple independent queues, the training set and validation set can be divided by random partitioning or stratified random sampling based on log2FC value binning. Alternatively, one queue can be used as the training set and the rest as independent validation sets to verify the model's cross-queue generalization ability. The validation set does not participate in feature selection or model training during the entire model building process and is only used for the final independent external evaluation.

[0031] The feature selection algorithm is the Recursive Feature Elimination (RFE) algorithm, and its feature selection process is as follows:

[0032] (1) Train a machine learning regression model based on the initial feature set and calculate the importance score of each feature;

[0033] (2) Remove the features with the lowest importance scores in each iteration to obtain a new feature subset;

[0034] (3) Retrain the model based on the new feature subset and evaluate its performance;

[0035] (4) Repeat steps (2) to (3) until the preset number of features or the optimal performance is reached, and output the optimal feature subset.

[0036] The machine learning regression model includes at least one of the following: XGBoost regression model, random forest regression model, gradient boosting regression model, and support vector regression model.

[0037] This invention provides a predictive model for changes in protein detection signals after serum precipitation with perchloric acid obtained by the above construction method.

[0038] This invention provides an application of the above-mentioned prediction model in predicting changes in protein detection signals after perchloric acid precipitation of serum, characterized by comprising the following steps:

[0039] (1) Obtain the identifier of the protein to be predicted; the protein identifier includes any one or a combination of UniProt accession number, gene symbol, and protein name;

[0040] (2) Based on the identifier, automatically obtain the amino acid sequence of the protein from a protein sequence database; the protein sequence database includes UniProt database, NCBI Protein database or other public or private databases that can provide protein amino acid sequence information;

[0041] (3) The protein physicochemical properties are automatically calculated based on the amino acid sequence, including PosAA, NegAA and ChargeRatio;

[0042] (4) Input the physicochemical properties of the protein into the constructed prediction model and automatically output the predicted log2FC value of the protein to be predicted relative to the untreated condition after serum precipitation with perchloric acid;

[0043] (5) Based on the predicted log2FC value, automatically determine the direction of the detection signal change of the protein to be predicted after serum precipitation with perchloric acid:

[0044] When the log2FC predicted value is > 1, it is judged as a significant up-adjustment, that is, it has a retention or enrichment trend;

[0045] When the predicted value of log2FC is less than -1, it is judged as a significant down-adjustment, that is, it has a weakening or settling trend;

[0046] When |log2FC predicted value|≤1, the change is considered insignificant.

[0047] Furthermore, for multiple proteins to be predicted, the above steps are automatically performed in batches, and all proteins to be predicted are sorted according to the absolute value of the log2FC prediction value, and a candidate protein signal change prediction list is output.

[0048] Compared with the prior art, the outstanding effect of the present invention is as follows:

[0049] (1) This invention establishes for the first time a quantitative prediction model and method for protein detection signal changes after perchloric acid precipitation of serum based on the physicochemical properties of proteins. It can predict the direction and magnitude of signal changes after perchloric acid precipitation treatment based solely on the amino acid sequence information of the protein to be predicted without conducting actual mass spectrometry experiments. This provides a quantifiable decision-making tool for the design of serum proteomics pretreatment strategies.

[0050] (2) This invention explicitly introduces a mechanism for dividing the training set and the validation set, which restricts feature selection and model training to the training set and evaluates model performance on an independent validation set, effectively avoiding information leakage and ensuring the objectivity and generalization ability of the model evaluation results.

[0051] (3) This invention systematically transforms protein sequence information into quantifiable physicochemical properties, and combines recursive feature elimination algorithm with machine learning models such as XGBoost to achieve modeling. The model has good predictive performance and interpretability.

[0052] (4) This invention identifies the dominant role of protein charge-related physicochemical features, represented by PosAA, ChargeRatio and NegAA, in signal change prediction. This discovery is consistent with the chemical characteristics of perchloric acid as a strong acid and perchlorate ion as a typical liquid-dissociating anion in the Hofmeister sequence, providing a quantitative basis for understanding the selective action mechanism of perchloric acid from the perspective of protein physicochemical properties.

[0053] (5) The prediction model constructed by this invention can be reused and can achieve automated prediction based on protein identifiers (such as UniProt accession number). The system can automatically complete the whole process of "identifier input → column acquisition → feature calculation → signal change prediction → result output", and has good usability and scalability.

[0054] (6) The prediction method of the present invention can be applied to the pre-evaluation of candidate biomarkers, helping researchers to determine whether the detection signal of candidate proteins after serum precipitation with perchloric acid may be retained, enriched, weakened or precipitated, thereby guiding the design of serum proteomics sample pretreatment strategies and improving the efficiency and success rate of proteomics research.

[0055] The following description, in conjunction with the accompanying drawings and specific embodiments, further illustrates the model, construction method, and application of the present invention for predicting changes in protein detection signals after perchloric acid precipitation of serum based on the physicochemical properties of proteins. Attached Figure Description

[0056] Figure 1 This is a simplified flowchart of a method for predicting changes in protein detection signals after serum precipitation with perchloric acid based on the physicochemical properties of proteins.

[0057] Figure 2 This is a schematic diagram of the overall process for predicting changes in protein detection signals after serum precipitation with perchloric acid based on the physicochemical properties of proteins.

[0058] Figure 3 The results show the proteomic identification of CRP and NT-proBNP serum samples. A represents the number of CRP and B represents the number of NT-proBNP serum proteins identified; C and D represent the number of FDA-approved protein biomarkers identified in CRP and NT-proBNP serum; Conc1, Conc2, and Conc3 represent different expression levels of CRP and NT-proBNP in serum, respectively.

[0059] Figure 4 The results show the performance of the prediction model based on the XGBoost and RFE algorithms.

[0060] Figure 5 The importance ranking diagram of prediction models based on XGBoost and RFE algorithms when the number of features k=10.

[0061] Figure 6 The feature importance ranking diagram of the best prediction model based on XGBoost and RFE algorithms when the number of features k=3.

[0062] Figure 7 Scatter plot of predicted and true values ​​of the best prediction model based on XGBoost and RFE algorithms when the number of features k=3.

[0063] Figure 8 When the number of features k=3, the prediction results of the best prediction model based on XGBoost and RFE algorithms for FDA approval of protein biomarkers. Detailed Implementation

[0064] likeFigures 1-2 As shown, a model for predicting changes in protein detection signals after serum precipitation with perchloric acid based on the physicochemical properties of proteins and a prediction method based on this model are presented, specifically including:

[0065] I. Dataset Construction

[0066] 1. Acquisition of signal strength data:

[0067] Two independent serum sample cohorts were used as data sources for modeling. Cohort 1 was a serum cohort containing CRP standards (standard numbers GBW09865, GBW09866, and GBW09868), with three CRP concentration gradients of 1.50 mg / L, 12.07 mg / L, and 82.49 mg / L. Cohort 2 was an NT-proBNP serum cohort, with three NT-proBNP concentration gradients of 321 ng / L, 2107 ng / L, and 17125 ng / L. In this embodiment, the NT-proBNP concentration was obtained using chemiluminescence reagents in a fully automated chemiluminescence analyzer manufactured by Mindray Bio-Medical Electronics Co., Ltd. (batch number: 1023051). The fully automated chemiluminescence analyzer model was i3000. Alternatively, it could be measured using commercially available ELISA kits; specific measurement procedures can be found in the instructions of the corresponding brand's kit.

[0068] Each cohort included a non-depleted group and a depleted group (with high-abundance proteins removed by perchloric acid precipitation) for each concentration / state, with multiple biological replicates for each group. The specific experimental procedures are shown below:

[0069] 1.1 Depleted experimental group

[0070] 1) Sample dilution: Take 50 μL of serum sample into a 1.5 mL centrifuge tube, add 450 μL of ultrapure water to dilute and vortex to mix;

[0071] 2) Precipitation: Add 20 μL of 70% perchloric acid and vortex thoroughly to mix; then let stand at 4°C for 20 minutes to ensure that the protein is fully precipitated;

[0072] 3) Collect the supernatant: Centrifuge the sample at 12,000 rpm for 5 minutes, then carefully aspirate the supernatant and transfer it to a new centrifuge tube;

[0073] 4) Neutralization and drying: Add 15% ammonia water for neutralization treatment, and adjust the pH to the range of 7.5~8.0 as the standard for completion of neutralization. Then, concentrate under vacuum at 45°C until the solvent is completely evaporated.

[0074] 5) Denaturation: Add 120 μL of 6 M guanidine hydrochloride solution to the dried sample and vortex thoroughly to dissolve and denature the protein;

[0075] 6) Reduction: Add 13 μL of 50 mM dithiothreitol (DTT) to make the final concentration of DTT in the reaction system about 5 mM. Incubate at room temperature in the dark for 1 h to reduce the disulfide bonds in the protein molecule.

[0076] 7) Alkylation: Add 15 μL of 100 mM iodoacetamide (IAA) to make the final concentration of IAM in the reaction system about 15 mM. Incubate at room temperature for 1 h under dark conditions to perform alkylation treatment to block the thiol group on the cysteine ​​residue.

[0077] 8) Enzyme digestion: Dilute the sample with 0.6 mL of ammonium bicarbonate solution, add trypsin at a protein to trypsin mass ratio of 1:50 (w / w), and digest at 37℃ for 16 hours.

[0078] 9) Solid-phase extraction: 20 μL of 1% trifluoroacetic acid solution was added to the enzymatic hydrolysis product for acidification; Oasis HLB μElution solid-phase microextraction plate was used. The packing material was first activated with methanol, and then equilibrated with an aqueous solution containing 0.1% (v / v) trifluoroacetic acid; after completion, the acidified enzymatic hydrolysis product was added to the microwells for loading. The packing material was briefly cleaned with 500 μL of 0.1% (v / v) trifluoroacetic acid aqueous solution to remove unretained impurities and non-specific binding components, thereby improving the purity of the subsequent elution components; 90% acetonitrile aqueous solution (containing 0.1% trifluoroacetic acid, v / v) was used as the eluent. 20 μL of eluent was added each time, for a total of 3 times. The eluent was collected and concentrated under vacuum until the solvent was completely evaporated.

[0079] 10) The sample was reconstituted using an aqueous solution containing 0.1% formic acid to obtain peptide samples for liquid chromatography-mass spectrometry analysis.

[0080] 1.2 Non-Depleted experimental group

[0081] Replace the 70% perchloric acid in step 2) of 1.1 with pure water, and replace the 15% volumetric ammonia in step 4) of 1.1 with pure water. The remaining conditions and steps are the same as in 1.1.

[0082] 2. Calculation of the change value log2FC of the detected signal:

[0083] For the same protein, the mass spectrometry detection signal intensity of its characteristic peptides in the Depleted and Non-Depleted groups is matched, and the target variable log2FC of the detection signal change is calculated using the following formula:

[0084]

[0085] Among them, Intensity Depleted The intensity of the mass spectrometry signal for the same protein in the perchloric acid precipitation group is [Intensity]. NonDepleted The mass spectrometry signal intensity of the same protein in the untreated group is shown. Taking the ThermoFisher Vanquish Neo ultra-high pressure nano-level liquid chromatography system tandem with the Orbitrap Exploris 480 high-resolution mass spectrometer as an example, the relevant core liquid chromatography and mass spectrometry parameters are shown in Table 1 below:

[0086] Table 1

[0087]

[0088] Mass spectrometry conditions: positive ion mode. The spray voltage was set to 1.9 kV, and the ion transfer tube temperature was 280 °C. The ion source type was a nanoliter spray ionization (NSI) source. The FAIMS external electrode temperature was set to 85 °C, and the compensation voltage (CV) was set to -45 V to perform gas-phase separation of ions before they enter the mass spectrometer, reducing chemical noise and improving signal selectivity. Data acquisition was performed in data-independent acquisition (DIA) mode: Level 1 full scan resolution was set to 120,000, scan range was 400–800 m / z, automatic gain control (AGC) target value was set to 3e6, and maximum ion implantation time was set to 25 ms. Level 2 resolution was set to 45,000, scan mass range was 400–800 m / z. High-energy collisional dissociation (HCD) was used for fragmentation, with the collision energy set to 27%. The AGC target value is set to 2e7, the maximum ion implantation time is set to Auto mode, and the Pre-AccumμLation function is enabled. Each DIA cycle contains 22 scan windows (Loopcount=22) to ensure sufficient sampling depth across the entire liquid phase gradient.

[0089] The results are as follows Figure 3 As shown. From the protein identification results ( Figure 3(A, 3B) In both CRP and NT-proBNP serum, the Depleted group showed higher protein identification coverage at all three concentrations. This result suggests that perchloric acid precipitation in standard serum systems can effectively reduce ionization competition and MS / MS sampling bias caused by high-abundance proteins, allowing for higher detection opportunities of fragmented, low-to-medium abundance proteins that are normally difficult to select due to dynamic range limitations. Figure 3 As shown in C and D, in the statistics of FDA-approved protein biomarkers, both pretreatment methods can detect a certain number of biomarkers, indicating that the experimental procedure has usability and stability in biomarker identification. The specific calculation results of log2FC are shown in Table 2:

[0090] Table 2. Summary of log2FC values ​​and significance of FDA-approved biomarkers in CRP and NT-proBNP serum groups.

[0091]

[0092]

[0093] 3. Calculation of the physicochemical properties of proteins

[0094] The core objective of this study is to construct a prediction system based on an XGBoost regression model to predict protein expression changes across different experimental groups using biological characteristics, particularly predicting the log2FC (logarithmic fold change) between experimental groups. Using log2FC as the target variable, representing protein changes under different experimental conditions (e.g., between the Depleted and Non-Depleted groups), this system can quantitatively describe the impact of perchloric acid pretreatment on protein expression and reveal its role in the structural rearrangement of the proteome within the experimental design.

[0095] 3.1 Protein Feature Selection

[0096] Based on the UniProt accession number of the protein, the amino acid sequence of each protein was obtained from the UniProt database. Using the ExPASy ProtParam tool (https: / / web.expasy.org / protparam / ), the following 10 candidate protein physicochemical properties were calculated according to its standard calculation rules to form the initial feature set S(0):

[0097] (1) Number of amino acids (Length): The total number of residues in the amino acid sequence of a protein;

[0098] (2) Molecular weight (MW): The theoretical molecular weight of the protein calculated based on the sum of the masses of each amino acid residue, in Da;

[0099] (3) Theoretical isoelectric point (pI): The pH value at which the net charge of the protein is zero, calculated based on the Bjellqvist algorithm;

[0100] (4) Negatively charged amino acid characteristics (NegAA):

[0101]

[0102] (5) Positively charged amino acid characteristics (PosAA):

[0103]

[0104] (6) Charge Ratio:

[0105]

[0106] (7) Instability Index: Calculated based on a dipeptide composition weighting method;

[0107] (8) Aliphatic Index: Calculated based on the relative volume percentages of alanine, valine, isoleucine, and leucine;

[0108] (9) Hydrophobicity index (GRAVY): The sum of the Kyte-Doolittle hydrophobicity values ​​of all amino acid residues divided by the total number of amino acids;

[0109] (10) Aromaticity: The sum of the mole fractions of phenylalanine, tryptophan and tyrosine.

[0110] Where, N R N K N D N E These represent the number of arginine, lysine, aspartic acid, and glutamic acid residues in the protein's amino acid sequence.

[0111] The physicochemical properties of proteins can also be obtained using any program or algorithm with calculation rules equivalent to those of the ExPASy ProtParam tool, including but not limited to the ProtParam module in BioPython, the Peptides package in R, and other similar tools.

[0112] 3.2 Target Variable

[0113] In this study, the target variable corresponding to protein characteristics is log2FC, and its calculation formula is as follows:

[0114]

[0115] Among them, Intensity Depleted With Intensity NonDeplete The values ​​represent protein intensity in the Depleted and Non-Depleted groups, respectively. This indicator effectively reflects the relative changes in protein expression under the two treatment conditions.

[0116] 3.3 Evaluation Indicators

[0117] After model construction, multiple evaluation metrics were used to measure the model's predictive accuracy and ensure its strong performance in predicting protein changes across different experimental groups. These metrics included directionality, error tolerance (Tight, Medium, Loose), and R-squared value. 2 MAE and RMSE. Each metric measures the predictive power of the model from a different perspective, helping us to fully understand the model's strengths and potential weaknesses.

[0118] Direction measures the consistency of the sign between the model's predicted log2FC and the actual log2FC. Specifically, the direction metric calculates whether the predicted value matches the actual value, given that the actual log2FC is greater than or less than 0. That is, if the actual value is positive (indicating protein upregulation in the treatment group), the model also predicts protein upregulation; if the actual value is negative (indicating protein downregulation in the treatment group), the model also predicts protein downregulation. The closer the direction value is to 1, the more accurate the model's prediction and the more correctly it reflects the direction of protein expression changes.

[0119] To further evaluate the model's predictive accuracy, three different error tolerance levels were defined: Tight, Medium, and Loose. These metrics measure whether the model's error in predicting protein changes is within a certain tolerable range.

[0120] To further evaluate the model's prediction accuracy, the concept of proportional error is introduced. Proportional error measures the relative error between the predicted full-factor (FC) and the true FC, and is expressed by the following formula:

[0121]

[0122] As shown in the table below, three different scale error tolerance levels are defined: Tight, Medium, and Loose, and the accuracy of the model is evaluated based on the scale error.

[0123] Table 3. Definition of Proportional Error (PE) Tolerance

[0124]

[0125] Tight indicates a very small error; the relative error between the predicted and actual FC values ​​is within a certain range (e.g., a 1.5-fold change), meaning the model's prediction is very accurate. Generally, a small difference in the fold change between the predicted and actual FC values ​​(e.g., less than 1.5 times) is considered "very accurate." Medium indicates an acceptable error; the relative error between the predicted log2 FC value and the actual value is approximately within 2 times. For most applications, this error range is acceptable. Loose indicates a larger error, but still within a tolerable range. Generally, if the relative error between the predicted and actual values ​​is greater than 2 but less than 3 times, it is considered a large error range, but still "acceptable." These error tolerance metrics help us analyze the model's accuracy more meticulously, especially when facing complex protein expression variations, allowing us to determine whether the model can still provide effective predictions within different error tolerance ranges.

[0126] R 2 The coefficient of determination (R²) is a classic metric for measuring the fit of a regression model, representing the proportion of variation in the target variable (log²FC in this study) explained by the model. 2 The value ranges from 0 to 1; the closer to 1, the more variation the model can explain, indicating stronger predictive power. In this study, R... 2 Used to measure the accuracy of XGBoost regression models in predicting changes in protein expression (log2FC). A higher R-value indicates a higher accuracy. 2 The value indicates that the model can capture the patterns of protein expression changes well, and the model shows good fit on both training and validation data. The calculation formula is:

[0127]

[0128] Where: n is the total number of samples;

[0129] y i This represents the measurement result for the i-th sample;

[0130] The average value of the measurement results of i samples;

[0131] Let be the predicted value for the i-th sample.

[0132] MAE (Mean Absolute Error) is the average absolute difference between predicted and actual values, measuring the average error between the model's predicted protein changes and the actual changes. A smaller MAE value indicates higher prediction accuracy. In proteomics research, MAE directly reflects the magnitude of the model's error in predicting different protein expression changes, making it a crucial evaluation metric. By minimizing MAE, the model can maintain high overall prediction accuracy. Its calculation formula is:

[0133]

[0134] Where: n is the total number of samples;

[0135] y i This represents the measurement result for the i-th sample;

[0136] Let be the predicted value for the i-th sample.

[0137] RMSE (Root Mean Square Error), similar to MAE, is used to measure model prediction error. However, unlike MAE, RMSE assigns higher weight to larger errors. A smaller RMSE value indicates a smaller overall error in the model. RMSE better reflects the reliability of model predictions, especially in the presence of large errors, highlighting predictive shortcomings. In proteomics analysis, RMSE helps assess whether the model accurately predicts large protein changes (e.g., changes in log2FC values ​​greater than 1 or less than -1). Its calculation formula is:

[0138]

[0139] Where: n is the total number of samples;

[0140] y i This represents the measurement result for the i-th sample;

[0141] Let be the predicted value for the i-th sample.

[0142] 4. Division of training and validation sets

[0143] The CRP queue data was used as the training set, and the NT-proBNP queue data was used as the validation set. The validation set did not participate in any form of feature selection or model training during the entire model building process, and was only used for the final independent external evaluation to avoid information leakage.

[0144] II. Prediction Model Construction

[0145] 1. Feature Filtering

[0146] On the training set, the Recursive Feature Elimination (RFE) algorithm was used to systematically screen the physicochemical properties of the 10 candidate proteins obtained in step one. The RFE algorithm uses a machine learning regression model as the base learner, recursively removes the least important features, and retrains the model in each iteration and evaluates the contribution of the remaining features to obtain the optimal feature subset.

[0147] In this embodiment, the overall performance of the model was evaluated under different numbers of retained features (k from 1 to 10). The results show that ( Figure 4 When k=3, the model achieves optimal overall performance. The key features selected are PosAA, ChargeRatio, and NegAA. Figure 5 ).

[0148] 2. Training of Machine Learning Regression Models

[0149] The optimal feature subset S∗ obtained from feature selection is used as input, and the log2FC value calculated in step one is used as the target variable. A machine learning regression model is trained on the training set. The machine learning regression model includes, but is not limited to, XGBoost regression, random forest regression, gradient boosting regression, and support vector regression. This embodiment preferably uses the XGBoost regression model, whose hyperparameters are tuned through 5-fold cross-validation within the training set.

[0150] 3. Model Performance Evaluation

[0151] The independent validation set obtained in step one is used to evaluate the performance of the constructed prediction model. The evaluation metrics include the coefficient of determination (R²). 2 The parameters include root mean square error (RMSE), mean absolute error (MAE), and direction consistency. Direction consistency represents the degree of agreement between the predicted and actual values ​​in the upward / downward adjustment direction. 2 The closer the Direction value is to 1, the better; the smaller the RMSE and MAE values ​​are, the better.

[0152] In this embodiment, the performance metric of the optimal model (k=3) on the independent validation set is: R 2 =0.835, MAE =0.934, RMSE =1.280, and directional consistency =81.2%, indicating that the model has good prediction accuracy and directional consistency.

[0153] Feature importance analysis results show that ( Figure 6In the optimal model, the contribution weights of PosAA, ChargeRatio, and NegAA are 0.414, 0.351, and 0.235, respectively, indicating that the physicochemical characteristics related to protein charge properties play a dominant role in predicting changes in protein detection signals after perchloric acid precipitation of serum. This result is consistent with the chemical characteristics of perchloric acid as a strong acid and the perchlorate ion as a typical liquid-dissociating anion in the Hofmeister sequence.

[0154] III. Application of Predictive Models

[0155] 1. Enter the identifier of the protein to be predicted:

[0156] Enter the identifier of the protein to be predicted. The identifier includes, but is not limited to, UniProt accession number, gene symbol, protein name, or any other unique identifier for the protein. For example, for the heart failure biomarker "B-type natriuretic peptide (BNP)," enter its corresponding English name on the UniProt database website (https: / / www.uniprot.org / ) to find its accession number as P16860. Then, enter this accession number on the website https: / / web.expasy.org / protparam / to open the corresponding protein list page.

[0157] 2. Automatically obtain amino acid sequences:

[0158] Based on the input protein identifier, the system automatically retrieves and obtains the amino acid sequence of the protein from a protein sequence database. This protein sequence database includes, but is not limited to, the UniProt database and the NCBI Protein database.

[0159] 3. Automatically calculate the physicochemical properties of proteins:

[0160] Based on the obtained amino acid sequence, the physicochemical property characteristics corresponding to the optimal feature subset S* of the protein to be predicted are automatically calculated according to the method in step 3.1. Taking the heart failure biomarker "B-type natriuretic peptide (BNP)" as an example, the physicochemical properties are calculated using the amino acid sequence (SPKMV QGSGC FGRKM DRISS SSGLGCKVLR RH) or accession number. The physicochemical property characteristics corresponding to its optimal feature subset S* are shown in Table 4.

[0161] Table 4. Characteristic subset S* of BNP, a biomarker for heart failure.

[0162]

[0163] 4. Automatically predict log2FC values:

[0164] Input the feature values ​​calculated in the previous step into the prediction model constructed in step two, and the model will automatically output the log2FC prediction value of the protein to be predicted relative to the untreated condition after serum precipitation with perchloric acid. Taking the heart failure biomarker "B-type natriuretic peptide (BNP)" as an example, the prediction value is "-0.71".

[0165] 5. Automatically determine the direction of signal change:

[0166] When the log2FC prediction value is greater than 1, it is determined that the detection signal of the protein to be predicted is significantly upregulated after serum precipitation with perchloric acid, that is, it has a retention or enrichment trend.

[0167] When the log2FC prediction value is less than -1, it is determined that the detection signal of the protein to be predicted is significantly downregulated after serum precipitation with perchloric acid, that is, it has a weakening or precipitation trend.

[0168] When |log2FC predicted value|≤1, it is determined that the detection signal of the protein to be predicted does not change significantly after serum precipitation with perchloric acid.

[0169] Taking the heart failure biomarker "B-type natriuretic peptide (BNP)" as an example, based on the predicted value, it was determined that the detection signal of BNP did not change significantly after serum precipitation with perchloric acid.

[0170] 6. Batch prediction and output:

[0171] For a list containing multiple protein identifiers to be predicted (e.g., a list of candidate biomarkers), the above five steps are automatically performed in batches, and all proteins to be predicted are sorted by the absolute value of the log2FC predicted value, and a candidate protein signal change prediction report is output. Figure 7 The scatter plot shows the true and predicted values ​​of the k=3 model. It can be seen that the relationship between the true and predicted values ​​is relatively strong, and most data points fall near the red dashed line, indicating that the model has strong predictive ability and can fit protein expression changes well. The red dashed line represents the ideal perfect prediction line; the model's performance is close to this ideal state, proving that it can effectively predict log2FC values ​​and reflect the changing trends of protein across different experimental groups.

[0172] Furthermore, applying the best model to an FDA-approved biomarker dataset and predicting its log2FC variation shows that... Figure 8In the model, the prediction results of most biomarkers are relatively evenly distributed within the log2FC range; the number of biomarkers that are upregulated and downregulated are 101 and 114, respectively, indicating that the model can be well applied to actual biomarker prediction.

[0173] The method of this invention can predict the direction and magnitude of the detection signal change after serum precipitation with perchloric acid based solely on the amino acid sequence information of the protein to be predicted, without performing actual mass spectrometry experiments, thus providing a quantitative tool for the pre-evaluation of candidate proteins in serum proteomics.

[0174] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for constructing a model based on the physicochemical properties of proteins to predict changes in protein detection signals after perchloric acid precipitation of serum, characterized in that, Includes the following steps: (1) Serum proteomics data after perchloric acid precipitation and serum proteomics data without perchloric acid precipitation were obtained through experiments; (2) Match the quantitative intensity of the same protein in the perchloric acid precipitation treatment group and the untreated group, and calculate the target variable log2FC to characterize the change in the detection signal; (3) Obtain the amino acid sequence of the protein and calculate the physicochemical properties of the protein based on the amino acid sequence; (4) Divide the obtained protein samples into training set and validation set according to a preset ratio. The training set is used for feature selection and model training, and the validation set is used for independent evaluation of model performance. (5) The physicochemical properties of the protein are screened on the training set using a feature screening algorithm to obtain a subset of key features after screening; (6) Train the machine learning regression model on the training set using the key feature subset and the target variable log2FC, and evaluate the model performance on the validation set to obtain the final prediction model.

2. The model construction method for predicting changes in protein detection signals after perchloric acid precipitation of serum based on the physicochemical properties of proteins according to claim 1, characterized in that: The formula for calculating log2FC is: ; Among them, Intensity Depleted Intensity represents the detection signal intensity of the same protein in the perchloric acid precipitation group. NonDepleted The signal intensity of the same protein in the untreated group.

3. The method for constructing a model based on the physicochemical properties of proteins to predict changes in protein detection signals after perchloric acid precipitation of serum, as described in claim 1, is characterized in that: The physicochemical properties of the protein include positively charged amino acid characteristics PosAA, negatively charged amino acid characteristics NegAA, and charge ratio characteristics ChargeRatio, which are calculated using the following formula: ; ; ; Where, N R N K N D N E These represent the number of arginine, lysine, aspartic acid, and glutamic acid residues in the protein's amino acid sequence.

4. The model construction method for predicting changes in protein detection signals after perchloric acid precipitation of serum based on the physicochemical properties of proteins according to claim 3, characterized in that: The physicochemical properties of the protein also include one or more of the following: number of amino acids, molecular weight, theoretical isoelectric point, stability index, aliphatic index, hydrophobic index, and aromaticity. The physicochemical properties of the protein were obtained using the ExPASy ProtParam tool or an equivalent physicochemical property calculation program based on amino acid sequences.

5. The method for constructing a model based on the physicochemical properties of proteins to predict changes in protein detection signals after perchloric acid precipitation of serum, as described in claim 1, is characterized in that: The serum proteomics data are derived from one or more serum sample cohorts, which include standard substance serum cohorts, clinical serum cohorts, or combinations thereof. When the data comes from multiple independent queues, the training set and validation set can be divided by random partitioning or stratified random sampling based on log2FC binning. Alternatively, one queue can be used as the training set and the rest as independent validation sets to verify the model's cross-queue generalization ability. The validation set does not participate in feature selection or model training during the entire model building process and is only used for the final independent external evaluation.

6. The method for constructing a model based on the physicochemical properties of proteins to predict changes in protein detection signals after perchloric acid precipitation of serum, as described in claim 1, is characterized in that: The feature selection algorithm is a recursive feature elimination algorithm, and its feature selection process is as follows: (1) Train a machine learning regression model based on the initial feature set and calculate the importance score of each feature; (2) Remove the features with the lowest importance scores in each iteration to obtain a new feature subset; (3) Retrain the model based on the new feature subset and evaluate its performance; (4) Repeat steps (2) to (3) until the preset number of features or the optimal performance is reached, and output the optimal feature subset.

7. The method for constructing a model based on the physicochemical properties of proteins to predict changes in protein detection signals after perchloric acid precipitation of serum, as described in claim 1, is characterized in that: The machine learning regression model includes at least one of the following: XGBoost regression model, random forest regression model, gradient boosting regression model, and support vector regression model.

8. A predictive model for changes in protein detection signals after serum precipitation with perchloric acid, obtained by the construction method according to any one of claims 1-7.

9. The application of the prediction model according to claim 8 in predicting changes in protein detection signals after perchloric acid precipitation of serum, characterized in that, Includes the following steps: (1) Obtain the identifier of the protein to be predicted; the protein identifier includes any one or a combination of UniProt accession number, gene symbol, and protein name; (2) Based on the identifier, automatically obtain the amino acid sequence of the protein from a protein sequence database; the protein sequence database includes UniProt database, NCBI Protein database or other public or private databases that can provide protein amino acid sequence information; (3) The protein physicochemical properties are automatically calculated based on the amino acid sequence, including PosAA, NegAA and ChargeRatio; (4) Input the physicochemical properties of the protein into the constructed prediction model and automatically output the predicted log2FC value of the protein to be predicted relative to the untreated condition after serum precipitation with perchloric acid; (5) Based on the predicted log2FC value, automatically determine the direction of the detection signal change of the protein to be predicted after serum precipitation with perchloric acid: When the log2FC predicted value is > 1, it is judged as a significant up-adjustment, that is, it has a retention or enrichment trend; When the predicted value of log2FC is less than -1, it is judged as a significant down-adjustment, that is, it has a weakening or settling trend; When |log2FC predicted value|≤1, the change is considered insignificant.

10. The application of the prediction model according to claim 9 in predicting changes in protein detection signals after perchloric acid precipitation of serum, characterized in that: For multiple proteins to be predicted, steps (1) to (5) are automatically executed in batches, and all proteins to be predicted are sorted according to the absolute value of the log2FC prediction value, and a candidate protein signal change prediction list is output.