Model construction method for evaluating gastric cancer suffering possibility of subject based on Uniprot database
By standardizing and screening proteomics data in various ways, a gastric cancer prediction model based on the Uniprot database was constructed, which solved the problem of poor diagnostic performance in existing technologies and achieved efficient and low-cost early diagnosis of gastric cancer.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-03-27
AI Technical Summary
Existing gastric cancer diagnostic methods have low sensitivity and specificity, making early detection difficult. Furthermore, deficiencies in protein biomarker screening and data processing methods lead to poor diagnostic results.
Multiple standardization methods were used to process proteomics data, and various algorithms were combined to screen protein biomarkers. A gastric cancer prediction model based on the Uniprot database was constructed, including data preprocessing, model training and optimization. 119 core biomarkers were screened, and 17 logistic regression prediction models were constructed.
It improves the sensitivity and specificity of early diagnosis of gastric cancer, reduces costs, expands the application scope of the model, and improves prediction efficiency.
Smart Images

Figure CN121747951A_ABST
Abstract
Description
Technical Field
[0001] This application mainly relates to the fields of biomedicine, biological research, proteomics and bioinformatics research technology, and gastric cancer prediction. Specifically, it relates to methods, systems, reagent kits, electronic devices and computer-readable media storing computer program code for constructing models to assess the likelihood of subjects developing gastric cancer based on the Uniprot database. Background Technology
[0002] Gastric carcinoma (GC) is a malignant tumor originating from the gastric mucosal epithelium. It ranks fifth in incidence worldwide and third in cancer-related mortality. Unfortunately, because early-stage gastric cancer often lacks specific gastrointestinal manifestations, most patients are diagnosed at an advanced stage with poor prognosis and limited treatment options, resulting in a low five-year survival rate. However, existing biomarkers for gastric cancer diagnosis and prognosis have low sensitivity and specificity. Therefore, current diagnosis relies primarily on barium meal X-rays, gastroscopy, abdominal ultrasound, spiral CT, and positron emission tomography (PET). These methods all have their limitations; for example, imaging techniques struggle to detect small tumors, leading to a high rate of missed diagnoses in early screening. Invasive surgery is inconvenient, and the low acceptance of gastroscopy and colonoscopy further restricts early detection and treatment. Therefore, there is an urgent need to find more sensitive and specific biomarkers and more effective diagnostic methods. Liquid biopsy, a novel detection technology, is gaining increasing attention. Peripheral blood, saliva, urine, or gastric lavage fluid / gastric juice can serve as sources of specific biomarkers, providing important data for the screening and diagnosis of gastric cancer.
[0003] Biomarkers are biochemical indicators that reflect changes or potential changes in the structure, tissues, organs, systems, or functions of cells and subcellular structures. They can be used to determine disease staging, diagnose diseases, or evaluate the safety and efficacy of new drugs or therapies. Among them, protein biomarkers, with their unique advantages, excel in accurately and sensitively screening for early, low-level damage, providing early warning of tumor development and offering clinicians a basis for auxiliary diagnosis.
[0004] However, there is still much room for improvement in the screening and research of protein biomarkers.
[0005] First, most current proteomics-based studies do not adequately consider the effectiveness of protein biomarkers in practical applications, which is one of the important reasons for the failure of protein biomarkers in practical applications.
[0006] Secondly, most current proteomics studies do not differentiate between random and non-random missing values when handling protein theorem values, applying the same treatment method. This approach may lead to the omission of some protein biomarkers that are not highly discriminative for cancer. For example, some proteins are generally expressed in healthy individuals or patients with benign diseases, but are mostly not expressed in cancer patients; these proteins are crucial for cancer differentiation. If a generic method of filling missing values using the mean or median is applied to all these proteins, they are highly likely to be misjudged and filtered out during subsequent screening, thus missing a potentially important biomarker. Although a few studies have attempted to differentiate the sources of missing protein quantitative values and use different missing value filling methods, the methods used in these studies are also difficult to implement.
[0007] Furthermore, current research shows that data processing methods significantly impact the selection of proteomics data. For example, the choice of data standardization method is crucial for the accuracy of subsequent proteomics data analysis. Different literature and research sources recommend varying optimal standardization methods, posing challenges in practical application. In addition, proteomics data possesses two prominent characteristics: sparsity and high dimensionality, which are also among the most challenging issues in data processing.
[0008] Therefore, there is an urgent need to establish a detection method that is highly accurate, rapid, and low-cost in multiple populations to help promote the implementation of routine screening and monitoring of high-risk groups for gastric cancer, and to effectively contribute to the early detection, early diagnosis, and early treatment of gastric cancer, thereby improving the cure rate and survival rate of gastric cancer patients. Summary of the Invention
[0009] Based on in-depth research into existing technologies, the inventors explored the screening of protein biomarkers. Experiments revealed that by comprehensively employing multiple commonly used standardization methods to standardize proteomics data and applying them to protein biomarker screening methods based on partial least squares discriminant analysis, orthogonal partial least squares discriminant analysis, random forest, lasso retrospective, Mann-Whitney U test, Welch's t-test, and Odds ratio, extremely excellent screening results can be obtained.
[0010] To address the aforementioned technical problems, this application provides, in one aspect, a method for constructing a model based on the Uniprot database to assess the likelihood of a subject developing gastric cancer. This method includes obtaining samples from the subject, including multiple serum samples; mass spectrometry data analysis, including performing mass spectrometry data analysis on the multiple serum samples to extract peptide information and obtain sample data; database retrieval, including performing a Uniprot database search on the sample data to identify proteins and obtain a raw protein quantification matrix; anomaly handling, including identifying and removing abnormal samples from the raw protein quantification matrix to obtain a valid protein quantification matrix; and analyzing differentially expressed proteins. This includes: comparing protein expression differences among different groups to obtain a first set of protein biomarkers with significant differences; data standardization processing, including: standardizing the effective protein quantification matrix to eliminate the influence of different dimensions and scales to obtain a standardized protein quantification matrix; and screening of protein biomarkers, including: using at least m of the following methods to screen the standardized protein quantification matrix to obtain at least m candidate protein biomarker sets: LASSO regression analysis, partial least squares discriminant analysis, orthogonal partial least squares discriminant analysis, RFS-based random forest algorithm, Mann-Whitney U test, Welch's t-test, Odds The ratio method and the P-values comprehensive value method are used, where m is a positive integer, 3≤m≤8; the biomarker summary analysis includes: summarizing and analyzing the first protein biomarker set and the m candidate protein biomarker sets, and obtaining the target protein biomarker that exists in at least 4 sets in both the first protein biomarker set and the m candidate protein biomarker sets; the correlation analysis includes: analyzing the correlation between the target protein biomarker and all protein biomarkers in the effective protein quantification matrix, and comparing the first biomarker with the target protein biomarker set whose correlation coefficient is within the first range among all protein biomarkers. A new set P1 is formed; variable structuring includes: dividing the samples in set P1 into gastric cancer group and non-gastric cancer group according to clinical diagnosis, and encoding them respectively to form the dependent variable of the prediction model. The independent variables of the prediction model include the quantitative results of protein, peptide or antibody and the patient's basic information; model training and testing includes: dividing the samples in set P1 into training set and test set, training the prediction model using the training set, and testing the prediction model using the test set. The prediction model is a regression model including the independent variables and the dependent variable; model optimization includes: screening effective protein biomarkers according to at least one of the indicators of p-value, correlation coefficient and AUC value, using the effective protein biomarkers as effective independent variables, and obtaining the optimized model.
[0011] This application proposes a system for assessing the likelihood of a subject having gastric cancer based on the Uniprot database in its second aspect, comprising: a data acquisition module for acquiring sample data of the subject, wherein the sample data includes A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A0X9V9B3, A0A126LAY7, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0A1Z1G4M2, A0 A2U8J8L9, A0A384MEF1, A0A385HVZ2, A0A5C2FTW7, A0A5C2FVW2, A0A5C2FX6 7. A0A5C2FXP5, A0A5C2FYJ4, A0A5C2FYK5, A0A5C2FZZ3, A0A5C2G130, A0A5C2 G1Q3, A0A5C2G1X1, A0A5C2G3M4, A0A5C2G4U1, A0A5C2G6G6, A0A5C2G7F2, A0 A5C2GA97, A0A5C2GBQ5, A0A5C2GCL9, A0A5C2GDA6, A0A5C2GGY3, A0A5C2GJH5 , A0A5C2GJL5, A0A5C2GMN5, A0A5C2GN07, A0A5C2GRZ5, A0A5C2GV43, A0A5C2 GW23, A0A5C2H3N5, A0A5S8K7B6, A0A7I2V2D2, A0A7P0TAB0, A0A7S5BYW5, A0A 7S5BZY2, A0A7S5C366, A0A7S5EWA8, A0A7S5EYK2, A0A7T0LP36, A0A8F0WQF6 , A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A6XNE2, A8K3I0, B2M1S7, B2 R4C5, B2R888, B2RMS9, B4DZH3, B4E1Z4, B4E273, B4E367, B4E3M1, B7WNR0, D 9YZU5, E9KL23, E9PHK0, G3V3A0, H0YAC1, H6VRF8, H7BYG8, H7C517, H9KVD5, L 8E853, O00617, O14746, P00738, P00915, P01031, P02533, P02652, P02671, P02750, P02753, P02763, P04275, P05109, P06727, P0C0L5, P0DJI8, P18428,The sample data includes P27169, P35527, P35900, P35908, P43652, P59666, Q06033, Q2L9S7, Q49A33, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q86YQ4, Q8NHV9, Q9H387, Q9NXP7, and V9H1D9 proteins, or any of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins, stable isotope characteristic peptide segments, and the corresponding nucleic acid quantitative values in serum; a data preprocessing module for preprocessing the sample data, including removing duplicate samples, unifying units, and dividing the sample data into different analysis groups according to the type and quantity of indicators; and a model recommendation module for... For different analysis groups, based on the types and number of indicators they contain, multiple predictive models with larger AUC values are recommended from among the multiple predictive models built using the construction method described above. The model selection module provides users with a selection function to output one or more predictive models from the models recommended by the model recommendation module for risk assessment and calculation. The risk assessment module calculates the corresponding predicted values based on the one or more predictive models output by the model selection module, and provides an appropriate Logit threshold or risk score threshold according to the department or population information corresponding to the sample data. The risk level is then classified as high probability, medium probability, or low probability based on the Logit threshold or risk score threshold.
[0012] This application proposes, in a third aspect, a kit for assessing the likelihood of a subject having gastric cancer based on the Uniprot database, comprising effective protein biomarkers, said effective protein biomarkers being any one or more of the following protein biomarkers: A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A0X9V9B3, A0A126LAY7, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0 A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, A0A385HVZ2, A0A5C2FTW7, A0A5C2FVW 2. A0A5C2FX67, A0A5C2FXP5, A0A5C2FYJ4, A0A5C2FYK5, A0A5C2FZZ3, A0A5C 2G130, A0A5C2G1Q3, A0A5C2G1X1, A0A5C2G3M4, A0A5C2G4U1, A0A5C2G6G6, A 0A5C2G7F2, A0A5C2GA97, A0A5C2GBQ5, A0A5C2GCL9, A0A5C2GDA6, A0A5C2GGY 3. A0A5C2GJH5, A0A5C2GJL5, A0A5C2GMN5, A0A5C2GN07, A0A5C2GRZ5, A0A5C 2GV43, A0A5C2GW23, A0A5C2H3N5, A0A5S8K7B6, A0A7I2V2D2, A0A7P0TAB0, A 0A7S5BYW5, A0A7S5BZY2, A0A7S5C366, A0A7S5EWA8, A0A7S5EYK2, A0A7T0LP 36. A0A8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A6XNE2, A8K 3I0, B2M1S7, B2R4C5, B2R888, B2RMS9, B4DZH3, B4E1Z4, B4E273, B4E367, B4 E3M1, B7WNR0, D9YZU5, E9KL23, E9PHK0, G3V3A0, H0YAC1, H6VRF8, H7BYG8, H 7C517, H9KVD5, L8E853, O00617, O14746, P00738, P00915, P01031, P02533, P02652, P02671, P02750, P02753, P02763, P04275, P05109, P06727, P0C0L5,P0DJI8, P18428, P27169, P35527, P35900, P35908, P43652, P59666, Q06033, Q2L9S7, Q49A33, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q86YQ4, Q8NHV9, Q9H387, Q9NXP7 and V9H1D9, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, or corresponding nucleic acids.
[0013] In its fourth aspect, this application proposes a kit for assessing the likelihood of a subject having gastric cancer based on the Uniprot database, comprising effective protein biomarkers, said effective protein biomarkers including any one or more of P59666 or its antigen / antibody, a single peptide chain, a characteristic peptide segment of the peptide chain, a stable isotopic protein, a stable isotopic characteristic peptide segment, and a corresponding nucleic acid.
[0014] This application provides, in a fifth aspect, an electronic device comprising: a memory for storing instructions executable by a processor; and a processor for executing the instructions to implement the construction method described above.
[0015] In its sixth aspect, this application provides a computer-readable medium storing computer program code that, when executed by a processor, implements the construction method described above.
[0016] This application proposes a marker-based model building method and identifies 119 important markers. Numerous models can be derived through any combination of one or more of these 119 markers. However, due to space limitations, this application selects only 17 representative models as examples to demonstrate the effectiveness of the proposed markers in model building. It is important to emphasize that these 17 models are merely examples and do not limit the breadth and diversity of the markers in practical applications. The marker-based model building method proposed in this application has a wide range of applications and can be applied to multiple fields, providing new ideas and methods for research and practice in related fields.
[0017] The construction method of this application, through the screening of independent variables and the gradual screening of independent variables during the training process of the regression model, enables the model to obtain good prediction results with fewer effective independent variables, thereby improving prediction efficiency and saving costs. This allows the resulting model, reagent kit, system, and electronic equipment to be applied to a wider range of scenarios and populations, possessing significant social value. Based on data from 902 samples, the gastric cancer prediction model and reagent kit established according to the construction method of this application derived at least 17 logistic regression prediction models with AUC values between 0.8 and 1.0 on the test set, demonstrating good predictive performance. Attached Figure Description
[0018] The accompanying drawings are included to provide a further understanding of this application; they are incorporated into and constitute a part of this application. The drawings illustrate embodiments of this application and, together with this specification, serve to explain the principles of this application. In the drawings: Figure 1 This is an exemplary flowchart of a method for constructing a model based on the Uniprot database to assess the likelihood of a subject having gastric cancer according to an embodiment of this application; Figure 2 These are the ROC curve and corresponding AUC results of prediction model 1 in this application embodiment on the training set; Figure 3 These are the ROC curves and corresponding AUC results of the prediction model 1 in this application, which are cross-validated on the training set. Figure 4 These are the ROC curve and corresponding AUC results of prediction model 1 in this application on the test set; Figure 5 This is a system block diagram of an embodiment of the present application for assessing the likelihood of a subject having gastric cancer based on the Uniprot database; Figure 6 This is a system block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely examples or embodiments of this application, intended to provide an intuitive demonstration. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0020] Furthermore, it should be noted that the use of terms such as "first" and "second" to define components is merely for the purpose of distinguishing the corresponding components. Unless otherwise stated, these terms have no special meaning and therefore should not be construed as limiting the scope of protection of this application. In addition, although the terminology used in this application is selected from commonly known and used terms, some terms mentioned in this application's specification may have been chosen by the applicant according to his or her judgment, and their detailed meanings are explained in the relevant sections of this description. Moreover, this application should be understood not only through the actual terms used, but also through the meaning implied by each term.
[0021] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more steps may be removed from these processes.
[0022] As used herein, the term "cancer" refers to the presence of cells exhibiting typical characteristics of cancerous cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rates, and certain characteristic morphological features known in the art. In one embodiment, "cancer" may be gastric cancer or stomach cancer. In one embodiment, "cancer" may include pre-malignant cancer as well as malignant cancer.
[0023] In one embodiment, as those skilled in the art will understand, the methods described herein do not involve steps performed by a physician. Therefore, the results obtained by the methods described herein require consideration of clinical data and other clinical presentations before a final diagnosis by a physician can be provided to the subject. A final diagnosis regarding whether a subject has gastric cancer is within the scope of the physician and is not considered part of this disclosure. Therefore, the terms “determine,” “detect,” and “diagnose” as used herein refer to identifying the probability or likelihood that a subject has a disease (such as gastric cancer) at any stage of development or determining the subject’s susceptibility to developing said disease. In one embodiment, “diagnosis,” “determine,” and “detect” are performed before symptoms appear. In one embodiment, “diagnosis,” “determine,” and “detect” allow a clinician (in conjunction with other clinical presentations) to confirm gastric cancer in a subject suspected of having it.
[0024] As used herein, the term "sample" refers to a sample collected from a subject for the purpose of detecting the types and amounts of gastric cancer markers present therein. Subject samples may be from the circulatory system (i.e., from the blood) or not from the circulatory system (i.e., not from the blood). Subject samples can be any sample containing substances suitable for detecting gastric cancer markers, and their sources include whole blood, bone marrow, pleural fluid, peritoneal fluid, central cerebrospinal fluid, breast milk, urine, tears, sweat, saliva, organ secretions, and lavage fluids from the bronchi, nasal cavity, pharynx, etc.
[0025] In one embodiment, the subject sample is blood, including, for example, whole blood or any part or component thereof. Blood samples suitable for this application can be extracted from any known source including blood cells or components thereof, such as veins, arteries, peripheral tissues, tissues, spinal cord, and the like. For example, the obtained sample can be obtained and processed using known and conventional clinical methods (e.g., procedures for drawing and processing whole blood). In one embodiment, the subject sample is serum. Methods for obtaining serum from blood are well known to those skilled in the art.
[0026] The inventors of this application have discovered that by monitoring whether a sample contains a group of gastric cancer biomarkers, gastric cancer can be diagnosed with high specificity and sensitivity. Especially for early-stage gastric cancer, which was previously difficult to diagnose, the biomarkers of this application also exhibit extremely high specificity and sensitivity.
[0027] As used herein, a “biomarker” or “marker” is a biological molecule that is objectively measured to serve as a characteristic indicator of the physiological state of a biological system. For the purposes of this disclosure, biological molecules include ions, small molecules, peptides, peptide chains, proteins, and peptides and proteins with post-translational modifications, nucleosides, nucleotides, and polynucleotides including RNA and DNA, glycoproteins, lipoproteins, and various covalent or non-covalent modifications of these types of molecules. Biological molecules include any kind of entities that are native to, characteristic of, and / or essential to the function of a biological system. Most biomarkers are polypeptides, although they may also be pre-translational mRNA or modified mRNA representing a gene product expressed as a polypeptide, or may include post-translational modifications of the polypeptide.
[0028] As used herein, "protein biomarker" means the biomarker that contains protein information. In one instance, it refers to the biomarker that contains a protein sequence. Further, in one instance, it refers to a full-length protein, a single peptide chain, a characteristic peptide segment of a peptide chain, and a stable isotopic protein or a stable isotopic characteristic peptide segment thereof.
[0029] The following lists some terms used in the embodiments of this application. Within the scope of this application specification and claims, the relevant terms are defined as follows. Other terms not listed are defined using definitions commonly used in the art, and their meanings are well known to those skilled in the art.
[0030] This application utilizes a quantitative matrix of proteins or antibodies based on clinical samples. Through a series of steps including abnormal feature processing, feature filtering based on the needs of subsequent practical application effectiveness, missing value imputation, abnormal sample identification and processing, data standardization, screening of protein biomarkers, training of logistic regression models, evaluation of the effectiveness of logistic regression models, and evaluation of the predictive performance of logistic regression models, a set of serum protein biomarkers for gastric cancer screening was identified. After further screening, 119 core serum protein biomarkers were obtained. Then, based on these 119 protein biomarkers, 17 logistic regression prediction models were constructed that can effectively distinguish gastric cancer samples from benign gastric disease samples and have predictive capabilities.
[0031] Non-random missing data: This refers to the situation where, due to problems with the sample itself, some protein quantification data cannot be detected, i.e., non-random missing protein quantification data occurs.
[0032] Random missing data: This refers to the situation where, due to random perturbations in the instrument during the actual protein quantification process, some samples cannot be detected in the protein quantification data, i.e., random missing protein quantification data occurs.
[0033] Log2 normalization, also known as Log2 transformation, is a logarithmic transformation of the expression values in the expression matrix, with the base 2.
[0034] Log2+Median Standardization: For the expression values in the expression matrix, first perform Log2 transformation. For each protein or antibody quantitative value of each sample in the transformed matrix, divide the protein or antibody quantitative value by the median value of all proteins or their antigen / antibody quantitative values in that sample, and then multiply by the median value of all proteins or their antigen / antibody quantitative values in all samples.
[0035] Log2+CycLoess standardization: For the expression values in the expression matrix, first perform Log2 transformation, and then use local weighted regression to standardize the transformed expression matrix.
[0036] Log2+Mean standardization: For the expression values in the expression matrix, first perform Log2 transformation. For each protein or antibody quantitative value of each sample in the transformed matrix, divide the protein or antibody quantitative value by the mean of all proteins or their antigen / antibody quantitative values in that sample, and then multiply by the mean of all proteins or their antigen / antibody quantitative values in all samples.
[0037] VSN Standardization: Since VSN standardization is similar to log2 transformation, there is no need to perform log2 transformation first. The expression values in the expression matrix can be directly standardized by variance stabilization.
[0038] Log2+RLR standardization: For the expression values in the expression matrix, first perform Log2 transformation, and then use robust linear regression to standardize the transformed expression matrix.
[0039] Log2+GI standardization: For the expression values in the expression matrix, first perform Log2 transformation. For each protein or antibody quantitative value of each sample in the transformed matrix, divide the protein or antibody quantitative value by the sum of the quantitative values of all proteins or their antigens / antibodies in that sample, and then multiply by the median value of the sum of the quantitative values of all proteins or their antigens / antibodies in all samples.
[0040] Log2+Quantile normalization: For the expression values in the expression matrix, first perform Log2 transformation, then sort each column separately, calculate the average of the sorted matrix to obtain the average vector, and then replace the corresponding average according to the original matrix sorting.
[0041] PLS: Partial Least Squares Regression Analysis, which combines principal component analysis, canonical correlation, and multiple linear regression into one, is a mapping dimensionality reduction method, and therefore can be used for feature screening of small samples.
[0042] True Yang: Correctly predicts the number of positive samples, which is actually a positive sample, and the prediction is also a positive sample.
[0043] True negative: The number of negative samples is correctly predicted, and the actual number of negative samples is also predicted to be negative.
[0044] False positive: The number of positive samples is incorrectly predicted when they are actually negative samples.
[0045] False negative: The number of negative samples is incorrectly predicted when they are actually positive samples.
[0046] Sensitivity: also known as recall or true positive rate, is the number of correctly predicted positive samples / the total number of actual positive samples, or (true positive) / (true positive + false negative).
[0047] Specificity: also known as the true negative rate, which is the number of correctly predicted negative samples / the total number of actual negative samples, or (true negative) / (true negative + false positive).
[0048] Accuracy: also known as positive predictive value, is the number of correctly predicted positive samples / the total number of predicted positive samples, or (true positives) / (true positives + false positives).
[0049] Accuracy: The number of correctly predicted positive and negative samples / the total number of samples, that is, (true positive + true negative) / (true positive + true negative + false positive + false negative).
[0050] F1 value: F1 = 2 (accuracy) Recall) / (Precision + Recall).
[0051] False positive rate: the number of incorrectly predicted positive samples / the total number of actual negative samples, equal to (1 - specificity), or (false positive) / (true negative + false positive).
[0052] Multicollinearity refers to a strong linear correlation among independent variables. Its main problems include: 1. Affecting the accuracy of least squares estimation, increasing the variance of regression coefficients; 2. Inability to accurately estimate the marginal effect of independent variables on the dependent variable; 3. Leading to insignificant regression coefficients for some independent variables; 4. Decreasing the predictive power of the regression equation. Methods for identifying multicollinearity include: 1. Correlation matrix method: observing the correlation coefficients between independent variables; 2. Variance inflation factor method: an excessively large VIF value indicates multicollinearity; 3. Condition number method: an excessively large condition number indicates multicollinearity. Methods for handling multicollinearity include: 1. Increasing the sample size; 2. Removing multicollinear independent variables; 3. Combining independent variables or using principal component analysis; 4. Regularization methods such as ridge regression. Therefore, multicollinearity can adversely affect regression analysis results and needs to be identified and addressed during the modeling process.
[0053] Lasso regression, also known as lasso regularization, compresses the coefficients of variables in a regression model by generating a penalty function to prevent overfitting and address severe multicollinearity. First proposed by Robert Tibshirani in the UK, Lasso regression is now widely used in predictive models. Lasso achieves sparsity by adding a penalty term (L1 regularization) to the objective function, making the weights of many features in the coefficient vector zero. By selecting features corresponding to non-zero coefficients, it can filter out features with the greatest predictive power for the target variable, thereby simplifying the model and improving its generalization ability. In medical research, multiple correlated independent variables often exist. Lasso regression can reduce the impact of multicollinearity on regression results by making the coefficients of correlated independent variables zero based on their correlations.
[0054] Random Forest: Random forest is an ensemble learning method that constructs multiple decision trees and obtains the final prediction result by voting or averaging the outputs of these trees. Each decision tree is built on a randomly selected subset of features and samples, which helps reduce overfitting and improve the model's generalization ability. Random forests can be used not only for classification and regression tasks but also for feature importance evaluation.
[0055] Recursive Feature Elimination (RFE): RFE is a greedy search algorithm that selects features by recursively considering increasingly smaller sets of features. At each step, the model is trained on the current feature set and removes the least important features (e.g., based on the absolute value of the coefficients or some measure of importance in the model output). This process continues until the desired number of features is reached or other stopping conditions are met.
[0056] Random Forest with Automatic Recursive Feature Elimination (RF-RFE): Combining random forests with automatic recursive feature elimination allows for more efficient feature selection. The specific steps are as follows: 1. Train the random forest model: First, train a random forest model using all features. 2. Evaluate feature importance: Determine the importance of each feature using the random forest model's feature importance evaluation mechanism (e.g., based on Gini impurity reduction or out-of-bag accuracy). 3. Recursively eliminate features: Recursively remove the least important features based on their importance, and retrain the random forest model in each iteration. 4. Select the optimal feature subset: Select the subset of features that best performs the model using cross-validation or other validation methods. Advantages include: 1. Reducing the complexity of the model by decreasing unimportant features, thus reducing the risk of overfitting. 2. Selecting the most useful features improves the model's predictive performance. 3. Examining the selected features allows for a better understanding of the data and interpretation of the model's predictions.
[0057] Partial Least Squares Discriminant Analysis (PLS-DA) is a supervised pattern recognition multivariate statistical analysis method that groups multidimensional data into groups based on the discriminant factors to be identified before compression (by pre-setting Y values for target classification and discrimination). This approach helps identify the most relevant variables for grouping while reducing the influence of other factors.
[0058] Orthogonal Partial Least Squares Discriminant Analysis (OPLS-DA): Combining Orthogonal Signal Correction (OSC) and PLS-DA methods, it can decompose the information of the X matrix into two types of information, Y, which are correlated and uncorrelated. By removing the uncorrelated differences, it can screen for differential variables.
[0059] Odds ratio, also known as the odds ratio, is a statistical indicator used to measure the ratio between the probabilities of two different occurrences. It is commonly used to compare the probability of a particular occurrence in two sets of data, helping to determine if there is a correlation between two variables and predict future probabilities. The odds ratio is defined as the ratio between the probability of a certain occurrence and the probability of it not occurring, usually expressed as "a:b" or "a / b", where 'a' represents the number of successes and 'b' represents the number of failures. The formula for calculating the odds ratio is: Odds Ratio = (A / C) / (B / D), where A and B represent the number of times a particular occurrence occurs and does not occur in the two sets of data, respectively, and C and D represent the number of times the same occurrence and does not occur in the other set of data, respectively.
[0060] Welch's t-test, also known as the Welch unequal variances t-test, is an improved version of the t-test. It is primarily used to test whether the population means of two independent samples are equal, especially when the variances or sample sizes are unequal. A fundamental premise of the classic t-test is the assumption that the variances of the two samples are equal. However, this assumption may not always hold in practice. Welch's t-test adjusts the formula to accommodate cases where the variances of the two samples may be unequal. This method makes the two groups comparable by correcting for the degrees of freedom or critical values. The steps of Welch's t-test typically include: calculating the means and variances of the two samples; calculating the t-statistic and degrees of freedom using the Welch's t-test formula; and determining the p-value based on the t-statistic and degrees of freedom to determine whether there is a significant difference between the population means of the two samples. It is worth noting that Welch's t-test eliminates the normality assumption, thus making it suitable for non-normally distributed data. This makes it more flexible in handling various types of data. In general, Welch's t-test is an effective method for comparing whether the population means of two independent samples are equal, especially when the variances may be unequal. In practical applications, it can help researchers more accurately determine whether the differences between two sets of data are significant.
[0061] P-values: This method provides a combined p-value by comprehensively considering the normality of the data. Specifically, it uses the Welch test to derive the result when the data follows a normal distribution, and the Mann-Whitney test when the data does not. Considering that in a set of data, some data for different markers may conform to a normal distribution while others may not, this method can flexibly handle these different situations, making it very suitable for practical applications.
[0062] %IncMSE: increase in MSE means assigning random values to each variable. If the importance is changed, the prediction error will increase, so the increase in error is equivalent to the decrease in accuracy.
[0063] IncNodePurity: increases node purity. Node purity is actually a decrease in RSS (residual sum of squares). An increase in node purity is equivalent to a decrease in the Gini index, meaning that the data or classes within a node are all the same, which is also known as Mean Decrease Gini.
[0064] The Mann-Whitney U test, also known as the Wilcoxon rank-sum test, is a nonparametric statistical test primarily used to compare whether there is a significant difference between the medians of two independent samples. This test is particularly suitable when the two groups of samples do not satisfy the assumptions of normality or homogeneity of variance, and can be seen as a nonparametric alternative to the independent samples t-test. The basic principle of the Mann-Whitney U test is to combine the data from the two samples, sort them according to their numerical values, and calculate the rank of each data point in both samples. Then, by comparing the sum of the ranks of the two samples, it can be determined whether the medians of the two samples are the same. If the medians of the two samples are the same, then the value of the Mann-Whitney U statistic U should be close to (n1...). The median of the two samples is calculated as (n1, n2) / 2, where n1 and n2 are the sizes of the two samples, respectively. If the value of U deviates from this expected value, the hypothesis that the medians of the two samples are the same can be rejected. The results of the Mann-Whitney U test typically report a p-value, which represents the probability of observing the current statistic or a more extreme statistic under the same assumptions. If the p-value is less than the significance level (usually 0.05), the hypothesis that the medians of the two samples are the same can be rejected.
[0065] Logistic Regression (LR) is a commonly used classification model that establishes the relationship between independent variables and categorical dependent variables. The main characteristics of a logistic regression model are: 1. The predicted dependent variables are discrete variables of binary or multi-category classification; 2. The logistic function is used to convert the values of linear combinations of independent variables into probabilities between 0 and 1; 3. It can handle both categorical and continuous variables; 4. Parameter estimation typically uses the maximum likelihood method; 5. It can explain the influencing factors of classification and the weights of each dependent variable; 6. It can calculate classification probabilities and make category predictions. The main steps in establishing a logistic regression model are: (1) Collect data and handle missing values, etc. (2) Select input variables and handle categorical variables. (3) Establish the logistic regression equation. (4) Estimate parameters using maximum likelihood. (5) Evaluate the overall effect of the model. (6) Conduct statistical tests to evaluate the impact of each variable. (7) Establish classification rules using probabilities. (8) Predict new data. The logistic regression model can both quantitatively analyze the effects of variables and be used for classification prediction, making it a very useful classification analysis method.
[0066] The present application will be further described below with reference to the embodiments and accompanying drawings.
[0067] Example 1 Method for constructing a model to assess the likelihood of a subject suffering from gastric cancer based on the Uniprot database.
[0068] Figure 1 This is an exemplary flowchart illustrating a method for constructing a model based on the Uniprot database to assess the likelihood of a subject developing gastric cancer, according to an embodiment of this application. (See reference...) Figure 1 The construction method of this embodiment includes the following steps: Step S1: Obtain samples from the subjects; Step S2: Mass spectrometry data analysis; Step S3: Database retrieval; Step S4: Exception handling; Step S5: Analyze differentially expressed proteins; Step S6: Data standardization processing; Step S7: Screening for protein biomarkers; Step S8: Biomarker summary analysis; Step S9: Correlation analysis; Step S10: Variable structuring; Step S11: Model training and testing; Step S12: Model optimization.
[0069] The following is a detailed explanation of steps S1 to S12 above.
[0070] Step S1 is used to obtain samples from known sources, which consist of a known number of benign gastric disease samples and a known number of gastric cancer samples. The sample sources in step S1 are: Serum samples from 365 individuals without gastric cancer and 537 individuals with stage I-IV gastric cancer were obtained from Shanghai Zhongshan Hospital and used as samples for this application. The collection of these samples followed the ethical standards established by the Ethics Committee of Shanghai Zhongshan Hospital, and informed consent forms were signed.
[0071] In steps S2 and S3, samples from known sources undergo mass spectrometry data analysis and Uniprot database retrieval to obtain protein quantification matrices for these samples. In some embodiments, the mass spectrometry data analysis in step S2 also includes sample preprocessing. For example, the 902 samples are treated as high-abundance, high-abundance mixed samples, and original samples, respectively, and appropriate amounts of protein from each are subjected to SDS-PAGE electrophoresis to assess sample consistency. This aims to exclude potentially contaminated samples. After evaluation, all samples are considered uncontaminated. Protein reduction and alkylation, and enzymatic digestion: Add 35 μL of UA buffer (8M Urea, 150 mM Tris-HCl, pH 8.0) and mix well. Add DTT to a final concentration of 20 mM, react at 37℃ for 2 h, then allow to return to room temperature. Add IAA to a final concentration of 25 mM (50 mM IAA in UA), shake at 600 rpm for 1 min, and incubate at room temperature in the dark for 30 min. Add 150 μL of NH4HCO3 buffer (50 mM), then add 2 μg of Lys-C to the sample and react for 4 h. Finally, add 4 μg of Trypsin and incubate at 37℃ for 16 h. Desalt using a C18 column, and determine peptide concentration at OD280. Then, take 2 μg of peptide from each sample, incorporate an appropriate amount of iRT standard peptide, and perform LC-MS / MS DDA and LC-MS / MS DIA methods for detection. LC-MS / MS analysis details are as follows: 1. All fractions used for DDA library generation were detected on an Easy-nLC 1200 chromatography system-OrbitrapExploris 480 mass spectrometer (Thermo Scientific). Peptides were separated on a C18 analytical column using a linear gradient of buffer B (84% acetonitrile in 0.1% formic acid) at a flow rate of 300 nL / min. MS detection was performed in positive ion mode, with a scan range of 350–1800 m / z. The MS1 scan resolution was 60,000 (@m / z 200), with an AGC (automatic gain control) target of 1e6, a maximum IT of 50 ms, and a dynamic exclusion time of 10.0 s. Twenty ddMS2 scans were performed according to the inclusion list after each complete MS–SIM scan. The isolation window is 1.5 m / z, the MS2 scan resolution is 30000 (@m / z 200), the AGC target is 1e5, the maximum IT is 50 ms, and the collision energy is 30 eV.
[0072] 2. DIA scan analysis
[0073] The chromatographically separated samples were analyzed by DIA mass spectrometry. Ionization mode: positive ion. Primary mass spectrometry scan range: 350-1800 m / z, mass resolution: 120,000 (@m / z 200), AGC target: 3e6, Maximum IT: 30 ms. MS2 used DIA data acquisition mode, with 44 DIA acquisition windows set, mass resolution: 30,000 (@m / z 200), AGC target: 3e6, Maximum IT: auto, MS2 Activation Type: HCD, collision energy: 30 eV, Spectral datatype: profile.
[0074] 3. Mass spectrometry data analysis
[0075] For DDA data, use Spectronaut TM 18. The FASTA sequence database was searched using Biognosys software. Parameter settings: enzyme: trypsin; maximum missed cleavage: 1; fixed modification: carbamoylmethyl (C); dynamic modifications: oxidation (M) and acetyl (protein N item). All reported data are based on a 99% confidence level for protein identification, determined by a false discovery rate (FDR) ≤ 1%.
[0076] DIA data using Spectronaut TM18. Analysis was performed, searching the constructed spectral library. Key parameter settings: retention time prediction type was set to dynamic iRT, MS2 level interference correction was enabled, and cross-run normalization was enabled. All results were filtered based on a Q-value cutoff of 0.01 (equivalent to FDR < 1%).
[0077] The database retrieval process in step S3 is as follows: The obtained mass spectrometry data were processed using Spectronaut software. TM 18) Perform a Uniprot database search, using the software's default parameters for other parameters, to identify proteins and obtain a protein quantification matrix with 902 samples and 6943 proteins.
[0078] In some embodiments of this application, between step S3 and step S4, further data processing steps are included, including: Checking for batch effects: Due to the large sample size, the data was divided into 7 batches for processing and testing. Before proceeding with subsequent analysis, we checked for batch effects between different batches. Principal component analysis was used to determine if batch effects existed.
[0079] Removing batch effects: If batch effects exist, use the R package statTarget and the QC-RFSC method to remove batch effects between samples caused by different batches based on the QC samples. statTarget is the default parameter setting.
[0080] In some embodiments, between steps S3 and S4, a further step is performed: protein filtering. Due to the limitations of the technology itself, quantitative values of protein biomarkers may be randomly missing. To improve the effectiveness of the application, proteins that have quantitative values in a preset number of samples are selected for subsequent protein biomarker screening. The number of qualified proteins is 3611. In one embodiment, this preset number is more than 80%.
[0081] In some embodiments, between step S3 and step S4, the following is further included: Missing value imputation: Both non-randomly missing protein quantification values and randomly missing protein quantification values were imputed using the K-nearest neighbor method.
[0082] The anomaly handling in step S4 specifically includes: using principal component analysis to identify 14 abnormal samples, removing these abnormal samples from the population, and obtaining a protein intensity information matrix with 888 samples and 3611 proteins.
[0083] Step S5, analyzing differentially expressed proteins, includes the following steps: After rigorous processing including steps S1-S3 above, batch effect removal, protein filtering, and missing value imputation, expression levels of the protein biomarkers to be screened in benign gastric lesion samples and gastric cancer samples were obtained, respectively. To further explore the expression differences of these proteins in the two samples, in some embodiments, two statistical analysis methods were used: t-test and Deseq2. T-test-based differentially expressed protein analysis initially screened out proteins with significant differences between the two groups of samples, forming a protein biomarker set ①, i.e., set G1. Deseq2-based differentially expressed protein analysis further considered the variability between samples and sequencing depth, providing more refined and reliable information on differentially expressed proteins, thus forming a protein biomarker set ②, i.e., set G2. Set ① contains a total of 46 proteins, namely A0A024R035, A0A024R0R6, A0A024R6P0, A0A024RCW3, A0A0S2Z3X8, A0A1Z1G4M2, A0A2U8J8L9, A0A384MDQ7, A0A5C2G2R4, A0A5C2GE26, A0A5C2GHJ3, A0A5C2GHK4, A0A5C2H2D3, A0A7I2V389, A0A7S5C3P9, A0A7S5C6L0, A2VCK8, A4D1J9, B4DZH3, B4E273, and B4E3M. 1. B7WNR0, B7Z347, G3V3A0, H6VRF8, H7BYG8, K7EM04, L8E853, O14746, P00738, P02533, P02671, P02750, P02763, P04275, P05109, P0DJI8, P18428, P35900, P59666, Q02985, Q06033, Q49A33, Q4VXF1, Q7Z4J2 and Q7Z664 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.Set ② contains a total of 57 proteins, namely P0DJI8, P02750, A0A024R035, A0A1Z1G4M2, P18428, P02671, Q06033, P59666, P04275, A0A024R6P0, A4D1J9, P02763, A2VCK8, H7BYG8, P05109, G3V3A0, and O14. 746, L8E853, A0A024RCW3, A0A024QZH6, Q4VXF1, A0A7S5C3P9, B7WNR0, H6VRF8, B4E3M 1. P00738, Q7Z664, A0A0S2Z3X8, B4DZH3, P35527, B4E273, Q7Z4J2, A0A024R0R6, P0253 3. Q49A33, P35900, A0A2U8J8L9, Q02985, A0A1B0GVI3, A0A5C2GHJ3, A0A5C2GQE6, A0A 5C2GE26, A0A384MDQ7, A0A5C2H2D3, A0A140VJU8, A0A5C2GRQ8, P02675, A0A5C2GG03, A 0A7S5C6L0, A0A5C2GHK4, A0A7I2V389, A0A5C2G2R4, B7Z347, A0A140T8Y3, A0A7S5BYR8, A0A5C2FXA4 and A0A5C2G9H4 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0084] Step S6, data standardization, specifically includes the following: Based on preliminary analysis, the inventors found that the combinations of protein biomarkers obtained by different standardization methods may vary significantly. Considering the effectiveness of practical applications, the inventors of this application adopted eight standardization methods: Log2, Log2+Median, Log2+Mean, VSN, Log2+RLR, Log2+GI, Log2+Quantile, and Log2+CycLoess. The protein expression matrices after steps S1-S3, batch effect removal, protein filtering, and missing value imputation were standardized. All eight standardized protein expression matrices were then screened using m of the following methods: LASSO regression, PLS-DA analysis, OPLS-DA analysis, and protein biomarker screening based on RFS random forest. Here, m is a positive integer, 3 ≤ m ≤ 8.
[0085] LASSO regression: The cv.glmnet function in the glmnet package of R language was used to screen protein markers by LASSO regression on the protein expression matrices obtained by the eight standardization methods. The set of protein markers obtained by analyzing at least 5 standardization methods was selected for subsequent analysis. Set ③ contains a total of 309 proteins, namely P19827, P02671, P59666, A4D1J9, A2VCK8, E9KL23, B2RBW9, A0A087WZ31, Q06033, P0DOX7, A6XNE2, A0A087WWT3, A0A140VK00, A0A5C2G410, P04275, A0A5C2GVD4, A0A126LAY7, A0A5C2GAB7, A0A5C2GI12, Q9UL92, A0A7S5BZY8, P02750, A0A5C2FT02, and S6C4R2. A0A5C2FXB2, A0A5C2GTA3, A0A5C2GLU1, A0A7S5BZY2, A0A5C2GLL1, P08697, A0A5C2FWZ4, H6VRF8, A0A5C2G130, P02533, A0A7S5EV71, A0A 7S5C1Q3, A0A5C2GLF5, A0A5C2G9Z6, A0A5C2FUE7, A0A5C2G946, A0A7S5C1R9, A0A5C2GBK0, P35908, A0A5C2GSX9, A0A5C2G9T5, A0A5C2GP9 1. A0A5C2FU05, A0A7S5EVM6, A0A5C2G486, A0A5C2G1C3, P11226, A0A109PW33, A0A5C2GIP8, A0A7S5C4N9, A0A7S5EUF4, A0A5C2G6H5, A0A5 C2GDB9, A0A5C2GC97, A0A5C2H2K5, B4E3L9, Q7M4S4, A0A5C2GWS2, A0A8G1A656, Q2L9S7, A0A5C2GPC0, A0A7S5BY93, A0A5C2G3V5, A0A5C2G 6X8, A0A5C2GM13, A0A5C2GXE0, A0A7S5EUP5, A0A7S5C0E5, A0A5C2FZH4, A0A5C2GW09, A0A5C2G6G6, A0A5C2GGA1, A0A5C2GHP0, A0A7S5EXS 4. A0A5C2GT79, A0A024RAB7, A0A7S5BXY2, A0A5C2GCL9, A0A5C2GQH9, A0A5H1ZRQ7, A0A7S5BZC9, A0A126GVG9, A0A5C2G2L3, A0A5C2FYC1,A0A5C2GE53、A0A5C2FVX0、A0A7S5C3B1、A0A5C2GLC0、A0A5C2GPW8、A0A7T0LNG1、A0A5C2GF53、A0A5C2FUM2、A0A5C2G4F7、A0A5C2GLD3、A0A5C2G2I2、A0A5C2G1Q2、A0A5C2FX21、A0A5C2GHS2、A0A5C2FTN5、A0A7T0LNG9、A0A1B2FK72、A0A5C2G9E4、Q86YL2、A0A5C2GFX4、A0A5C2FWH6、A0A5C2FYP4、A0A5C2H0N4、A0A5C2G455、A0A5C2G3Q6、A0A109PW55、A0A5C2G779、A0A5C2GKT9、A0A5C2GI92、A0N5G4、A0A7S5BZL2、P01833、A0A5C2GHG4、A0A5C2GX17、A0A5C2G9T0、A0A5C2G316、A0A5C2G3A1、A0A0X9TD88、P47992、A0A5C2G6N3、A0A7S5C4F8、A0A5C2GV43、A0A5C2G3I9、A0A5C2FVY2、A0A5C2GYX9、A0A7S5EX91、S6BAQ4、A0A5C2G577、A0A7S5EW62、A0A5C2FUD7、A0A5C2GAN8、A0A5C2GSS8、A0A7S5BYK8、A0A5C2GF59、A0A5C2GJ87、A0A2U8J8W9、A0A5C2GMH6、A0A5C2GXH1、A0A494C037、Q13784、A0A5C2FZW7、A0A5C2G801、A0A5C2GF87、A0A5C2G3Z3、A0A5C2GXM0、Q15408、Q92820、A0A5C2GH77、A0A5C2G252、A0A7S5C0M0、A0A5C2FYE5、A0A5C2G846、A0A0X9T7T4、A0A7T0LNJ3、A0A090N7U9、Q16777、A0A7S5BYR8、A0A5C2GQG6、A0A5C2G3F5、A0A5C2G711、A0A7S5C1F2、A0A5C2GNC4、A0A7S5C4F1、A0A5C2H1X6、A0A5C2GEW6、A0A7S5EU68、A0A5C2G5D4、A0A7S5BZR6、A0A7S5BY99、A0A5C2G0L1、A0A5C2GH73、A0A5C2H0U9、A0A5C2GJ21、A0A5C2GRE3、A0A5C2G0D4、A0A5C2FTG4、A0A5C2FZM5、A0A5C2G1K4、A0A5C2GD56、A0A5C2GER6、B7Z4R3、A0A5C2GR71、A0A5C2GJ77、A0A5C2FZX7、A0A5C2FZS2、A0A5C2FWL8、A0A7S5C0E1、A0A140VJI7、A0A5C2GJC8、A0A7S5EXB6、A0A7S5BY64、A0A5C2GH79、A0A5C2GTD7、A0A5C2G4J4、A0A7S5C0M9、A0A0A0MRS8、A0A5C2GMD4、A0A7S5C2Q8、A0A7S5C4H8、A0A5C2GUV1、A0A7S5C3Z6、A0A5C2G2X5、A0A7S5EWH5、A0A5C2FZA2、B2R701、A0A024R9J3、A0A5C2G3L8、A0A5C2GH94、A0A5C2FXP5、A0A7S5EYK2、Q86TT1、A0A5C2GJB2、A0A5C2GMM6、A8KAP9、A0A0S2Z3Y1、A0N071、A0A5C2GJB9、A0A5C2G9P8、A0A7S5BXQ1、B4E367、V9H1D9、A0A5C2G7L1、A0A7S5EX22、M0QX59、A0A5C2FVW2、P27169、A0A5C2GNX8、A0A7T0LND4、A0A5C2GEX5、A0A7S5EX43、A0A5C2GYN0、A0A5C2GCR5、A0A7S5EXJ5、A0A5C2G4U1、E9PHK0、A0A5C2GQU4、P14151、A0A5C2GM78、A0A5C2GTA5、A0A5C2G401、Q6IQ49、A0A5C2FZ71、A0A7S5C051、A0A5C2G0A9、A0A5C2G3P7、A0A5C2GVR0、A0A5C2GKW4、A0A7S5BYD6、A0A5C2GM50、M0R1M6、O14746、A0A5C2GJH5、A0A5C2G1Q3、A0A024R930、A0A5C2GH71、A0A5C2GQM2、A0A5C2G827、H7BYG8、A0A5C2G1M4、A0A5C2FZ01、A0A7S5EWA8、P07360、A8K1K1、Q5CZ93、A0A5C2G2K3、H0YAC1、A0A5S8K7B6、A0A7S5C336、A0A5C2G3M3、A0A5C2G818、B3KUH0、Q8NHV9、Q4TZM4、A0A5C2GJK5, A0A5C2GJY0, A0A7S5C036, A0A3G1HEF1, A0A5C2FWH3, A0A5C2G115, A0A5C2GFZ5, A0A5C2F XE7, A0A7S5C2Y8, A0A5C2GW40, A0A5C2GSP8, A0A5C2GTX8, A0A5C2FVF4, A0A7S5BZS1, A0A5C2GXR3, A0A7 S5C5T5, A0A5C2FVB0, Q6N092, A0A0A0MRZ9, A0A7S5C385, A0A5C2H1L1, P02652, A0A5C2GB45, A0A5C2H0L1, A0A0G2JMS6, A0A023T424, A8K3I0, and P43652, or their antigens / antibodies, or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0086] Partial Least Squares Discriminant Analysis (PLS-DA): Protein expression matrices obtained by eight standardization methods were screened for protein biomarkers using PLS-DA, and protein biomarkers that appeared in the results of at least five standardization methods were selected to form a set of protein biomarkers. Set ④ contains a total of 900 proteins or peptides, namely P0DJI8, O00617, A0A384MDQ7, Q86W24, A0A0G2JMS6, Q4TZM4, A0A024QZH6, P59666, A0A024R980, B7Z3V1, A0N071, A0A0A0MTQ8, A0A5C2GYJ4, A0A5C2FTX2, Q49A33, A0A7P0TAB0, C9JEE0, A0A0K2BMD8, P18428, A0A5C2G0Y8, B4DZH3, P02743, D9ZGG2, and A0A7P0T7. A 0A5C2G787, A0A5C2GE25, A0A5C2GAN3, A0A8I5KT93, A0A494C037, P05109, Q86YQ4, A0A5C2GE15, A0A384MR03, P02750, B2R6W1, A0A5C2G8H8 A 0A024R6P0, A0A5C2G3I5, Q86TT2, A0A0U1RR27, A0A5C2GPF4, Q96JA8, A0A024R035, P02763, A0A5C2GMY3, Q06033, A0A1W6IYL8, G3V3A0, A0A 7S5BY74, A0A024R6T9, A0A024RBU2, Q8W5, A0A5C2GGG9, H9KVD5, A0A7S5C6B0, A0A5C2G7P4, A7L8C6, P0C0L4, A0A7S5EY08, B2RA39, A0A7T0 PXT4, A0A5C2G6Q8, A0A140KFU0, A0A5C2G3A1, A0A1W2PPX9, A0A7S5C0E2, A0A0D9SF05, C8C504, A0A5C2GJ24, A0A385HVZ2, A7E2C5, Q7Z5J4,A0A7S5EUJ3, A0A5C2G7S3, Q9H387, A0A5C2GA89, A0A7S5EV15, A0A7T0LP36, A0A5C2G648, Q13790, A0A7S5EUR0, A0A5C2GEX2, A0A5C2G755, A0A5C2GJ30 A0A5C2G6X8, A0A7S5C1Z6, A0A5C2GRD0, O75636, P43251, D6RBM5, A0A024R4F9, A0A5C2GVM5, A0A5C2FYC5, A0A2U8J8S3, A0A5C2GFT6, A0A2U8J8L9, H7C5 17, B4E1Z4, A0A5C2GCD2, A0A590UK17, P19652, P0C0L5, A0A7S5C2Y9, A0N5G5, A0A2U8J993, A0A7S5C4U0, A0A1L7B5J3, A0A5C2GIJ6, A0A7S5C041, A0A5C 2GDW0, B2RMS9, A0A5C2FXI7, A0A8G1A656, P04275, A0A5C2FYQ6, A0A5C2GAB7, A0A7P0T8D1, A0A5C2GL46, A0A5C2GVL1, A0A0S2Z5S9, A0A7S5C5T5, P0DOX7 、A0A024RDY9、A0A7T0PXV6、A0A2U8J987、P07360、A0A5C2GXH1、A0A5C2GHP0 、A0A0S2Z3Y1、B2R888、B2M1S7、A0A7S5BYI5、A0A5C2GGY4、A0A5C2GVX4、A0A 5C2FT02、P00915、B4E3M1、O00187、A0A5C2G1W1、P00742、A0A5C2GJI3、A0A1 W2PQX5、A0A5C2GRW3、A0A5C2GEW6、P11226、B4DIJ4、A0A7S5BZS9、A0A5C2FU N2、A0A7S5EUL2、A0A140VK00、A0A024R930、P07357、A2RRP1、A0A5C2G0F5、P 08603、A0A5C2GLZ3、Q53H26、E9KL23、A6NKX1、A0A5C2G1E1、A0A7S5C173、A0 A4P8J9P8, A0A5C2GBC6, A0A7T0LNJ3, A0A5C2FZ84, A0A3B3IQ51, A8K237, A0A5C2GWD7, A0A5C2G8Y9, O14893, Q03591, B2R5G8, A0A5C2G172, A0A5C2G0A9A0A3B3ISR2、A0A1U9X793、A0A8F0WQF6、A0A5C2GDN1、Q13784、A0A5C2GV43、 Q6N091、A0A7S5EXP5、A0A5C2G0Q5、Q9H213、P01031、A0A5C2GSJ0、A0A5C2GJF 4, A0A5C2G4G0, A0A7S5C0H1, A0A5C2G3Y2, A8K1K1, Q6L981, A0A5C2GN85, A0A494C1P3, A0A024R3L8, B4DR59, B3KS98, A0A5F9ZGX3, P04003, A0A5C2FY76 A0A5C2FX99, A0A7S5C130, A0A7S5C0A9, A0A5C2FZP6, A0A7S5C6L1, A0A7S5C3P9, A0A5C2GA92, A0A5C2GM78, A0A0S2Z3D5, A0A5C2FWH6, A0A126GVG9, Q5T4 I8, A0A5C2GN30, A0A1W6IYK2, A0A075B708, A0A5C2GT92, A0A5C2G4, A0A5C2G699, A0A5C2GA52, A0A5C2FZV3, A0A7S5BZY2, P20851, A0A087WUZ3, A0A5C2G L08、A0A5C2GIK6、A0A5C2G4P2、A0A5C2GTA9、P02790、A0A5C2GTJ5、A0A5C2G 370、A0A7S5C299、A0A5C2GJB2、A0A5C2GFV3、A0A5C2GMH6、A2NV54、P09871、B 4DPR2、A0A346RV53、A0A024R6F8、A8K1B1、Q8NCL6、A0A5C2FZE6、E7ETH0、A0 A5C2GP10、A0A5C2GY49、Q6IQ49、Q96IY4、Q04756、A0A5C2GF13、A0A5C2G252、 Q9Y6R7、A0A5C2G1Q3、A0A5C2GG08、A5PL27、A0A5C2G5A1、G3XAK1、A0A3G1HE F1、A0A090N7U9、A6XNE2、A0A0D9SFG7、L8EBI7、A0A5C2GU97、P35542、Q4VXF1 、A0A140VKF3、A0A5C2G729、A0A5C2G561、A0A5C2G7A2、A0A5C2GVF6、B2RBW9 、A0A5C2GEX5、A0A5C2FYM3、A0A4V1EJ15、A0A5C2H0N4、P13671、A0A5C2GGH7、P22352、A0A7S5EWD6、A0A5C2FYQ1、A0A5C2FYV6、A0A5C2GD51、A0A5C2G919、 A0A5C2GKS6、A0A023T424、A0A5C2GBR0、P10643、A0A5C2FZW7、B4DZQ3、A0A5 C2GL60、A0A5C2GY47、Q9NQ79、A0A5C2GLL5、A0A5C2GDF6、A0A5C2GW00、A0A0 24RAA7、A0A7S5C1Q3、P01024、B4DG83、B4E1C2、C0JYY2、A0A5C2GTA5、A0A5C 2GHU5, A0A384N669, A0A7T0LNB3, M0QX59, P04217, B4E273, A0A384MEC0, A0A5C2GKT3, A0A5C2GT38, A0A5C2GC48, Q6LAM1, A0A7S5C1D2, A0A5C2FUP9, A8 K9G4, A0A5C2GA43, Q56G89, A0A7S5C051, A0A5C2FVC6, A0A7S5EYK2, A0A7S5BYN9, A0A7S5BZH2, A0A5C2GF87, A0A5C2FXP5, A0A5C2G5X1, B4E215, A0A7S5 EX52、A0A7S5C4M5、B2R582、A0A5C2G7A1、P05546、A0A024RAB9、A0A5C2H0U9 、A0A5C2GF54、A0A5C2FUS4、A0A5C2GBQ5、A0A5C2GVX1、A0A7S5C442、A0A5C2 GCT8、A0A7S5EXE1、A0A0S2Z4I5、A0A5C2GH71、A0A5C2FW89、A0A7S5C300、A0 A7S5EXI9、A0A5C2GYW2、A0A7S5EX91、A0A7I2V2D2、A0A0S2Z4L3、A0A5C2G3F 9, A0A5C2GSZ1, A0A7S5EYC3, A0A0X9V9B3, A0A087WZ31, A0A5C2GXN9, A0A5C2GRA7, A0A5C2GTP5, O75882, A0A5C2G5D9, A0A5C2G4T9, B7Z8Q2, A0A7S5BZN 7, Q15408, A0A075B7A1, A0A7S5EXJ5, O00634, A0A5C2FWW2, A0A5C2GK81, A0A7S5C336, A0A1W6IYJ9, P22792, A0A7S5C775, H0Y973, A0A5C2GIZ0, A2J1N7A0A109PS32、P00747、J3KT96、V9HWI6、A0A7S5C0G1、A8K9M5、B4DLR8、A0A5C2GM50、A0A5C2GKE6、A0A5C2FWN1、A0A7S5C3K5、P19827、A0A0C4DGB6、A0A024R032、A0A7S5C137、A0A5C2GF35、B4DUV1、A0A7S5EXZ4、A0A5C2GLU1、D3DNN4、Q9UHG3、A0A5C2G921、A0A5C2G292、A0A5C2GNS8、P80108、A0A5C2GHA6、A0A7S5C3I9、L8E853、A0A5C2GGI4、A0A2Z4LCH4、A0A5C2GH94、A0A5C2GA26、V9H1D9、Q9BYH2、A0A7T0LN44、A0A5C2FZZ3、A0A5C2FYG2、A0A0S2Z4F1、B4E0X1、A0A7S5C0M2、A0A5C2GJC8、A0A5C2G7K4、A0A5C2GYW0、Q9NXP7、P07358、A0A5C2GGY7、A0A5C2FZ99、A0A5C2GGI5、A0A384NKS6、A0A7S5C3X3、A0A140VK24、A0A5C2G2L3、A0A5C2GNC7、A0A7S5BXB3、A0A5C2G3P7、A0A384MEF1、A0A0M3KKW6、P15169、A0A5C2FXF8、Q9UMS6、A0A5C2FX28、A0A5C2GVG6、P00734、A0A5C2G6D1、A0A5C2GGK7、A0A5C2G410、A0A7S5BY02、A0A5C2GH61、I3L145、A0A5C2G9N0、A0A5C2G900、A0A2U8J8X2、A0A5C2FXY6、P35527、A0A5C2G7L1、E9PHK0、A0A5C2G4T8、A0A1B0GX98、A0A5C2G1R6、Q8NHV9、A0A193AUV8、A0A5C2GUV2、A8K1Z4、A0A5C2FYJ4、A0A5C2GTA3、A0A5C2GFC3、A0A5C2FYF1、A0A7S5EX22、A0A2U8J9A8、A0A5C2G9T5、A0A5C2GXE0、A0A5C2G2E8、A0A5C2GMF2、A0A5C2G9T0、A0A024R892、B3KUW8、A0A7S5EVM6、A0A5C2GY06、H0YAC1、A0A5C2G525、A0A5C2GXG7、A0A5C2FWW6、P03950、A0A024R6I9、A0A5C2GCR5、A0A024R8G3、A0A5C2GAJ1、A0A7S5EWG2、A0A5C2GWI4、P02753、A0A024R5X3、P02768、A0A5C2G987、A2RTY6、A0A5C2GGH5、A0A5C2FY35、P00748、A0A0D9SGE8、A8KAP9、Q9NSI6、A0A5C2G2P9、A0A5C2GZU5、A0A5C2GW19、B7Z7M2、A0A5C2FUG7、Q7Z4J2、A0A5C2GJ96、A0A5C2FZM5、A0A5C2GD99、A0A5C2GVA5、A0A5C2H2K5、A0A5C2FZ71、A0A5C2GMG3、A0A5C2GPG2、A0A5C2FVW2、A0A5C2G9I0、A0A5C2GH79、P12236、E7EQY3、A0A1B0GVI3、P47992、A0A5C2G4U1、A0A5C2G3M3、A0A7S5EX43、F8VVF9、A0A7S5C361、F2Z2G4、P48740、Q96NM2、A0A0U1RQQ9、E9PI90、A0A5C2GLH1、A0A1B0GU24、A0A5C2G1D9、A0A5C2GE11、B4DPQ3、B7Z6K7、P14151、A0A7S5EVT4、A0A5C2GR38、B3KUH0、A0A5C2G9Q6、A0A5C2G1Y7、A0A5C2G2K3、A0A5C2GLU3、A0A5C2GX17、A0A5C2GQJ7、A0A286YEY4、A0A5C2G280、A0A7S5BXY2、Q6ZW64、A0A7S5C031、Q15475、D9IWP9、A0A1L2BU44、A0A5C2GQU4、Q2L9S7、P27169、A0A5C2GQE5、A0A5C2GI12、A0A5C2FVW4、A0A7S5C372、A0A5C2FXB2、A0A7S5EW26、A0A5C2GHZ0、P02760、A0A5C2G401、A0A5C2GQE4、A0A5C2G7E0、E7EUW0、A0A5C2GF85、A6XGL1、A0A5C2GU54、A0A5C2FWX3、A0A0S2Z4D9、A0A5C2GIB6、A0A7T0PXH4、A0A7S5C2G0、A0A5C2FVY0、P06727、A0A161I202、A0A5C2GRQ9、A0A7U3R8A8、Q9UGM5、A0A5C2GBF4、A0A067XG54、A0A7S5C4F1、A0A024R6C9、A0A5C2FV95、A8K669、B4DI57、A0A5C2GKG2、A0A5C2G8Y7、A0A7S5C177、A0A7S5C1W9、A0A5C2GLJ1、A0A024RAG6、A0A7S5EU68、A0A5C2GSL5、A0A2U8J9D3、A8K6C9、A0A5C2GSF4、A0A7S5BYD6、A0A7S5EWA8、A0A5C2G6H4、A0A5C2G826、A0A5C2GAY0、A1L4G7、A8K3I0、A0A126LAY7、A0A5C2FZA5、A0A2U8J933、Q8NFI4、Q0ZCI2、A0A5C2GSW5、A0A5C2G368、D6RB81、A0A5C2GNX8、A0A024R6I6、A0A5C2GUL3、K7ER74、A0A7S5EXS6、A0A5C2G792、A0A5C2GAJ3、B7WNR0、Q86SQ4、Q0ZCJ1、A0A024RCW3、A0A140VJU4、A0A5C2G8T6、A0A7S5C2A9、P43652、A0A5C2FZJ0、A0A7S5BZF8、A0A7S5BYB3、Q92496、B7ZMH4、A0A5C2GKW4、A0A7S5EUV7、A0A5C2GE53、A0A024RAB7、A0A5C2G8L6、P35908、A0A5C2FYC6、B7ZM24、H7BYG8、A0A7S5BXW8、A0A5C2GE10、A0A7S5C2N4、P02652、O60271、A0A5C2G3K7、A0A5C2GDW3、A0A5C2FV87、Q6N022、A0A5C2GAI3、A0A024RDM6、A0A7S5C2D6、A0A5C2G3I9、B2R4C5、A0A5C2G7F2、A0A024R3E3、E5RJW4、P02787、A0A7S5C3H9、A0A024QZ94、B0YIW2、A0A024R1G8、Q96N06、B4DZM1、A0A7S5BY99、A0A5C2GMN5、A0A804HI36、A0A7S5BYQ0、A0A5C2GK31、P08697、A0A5C2GBU2、A0A494C0I1、A0A5C2GBH7、A0A5C2FSY5、A0A7S5C1X5、A0A7T0PWW2、Q9Y4F3、A0A5C2GGS8、A0A5C2GF70、B2R5U1、A0A5C2GLZ9、A2J1M5、A6XND0、A0A5C2GXD7、A0A7S5C4Y8、A0A7S5C781、A0A7S5EW62、A0A5C2GZS9、B4DHZ6、A0A5C2G5W9、E7ETY2、A0A5C2G6G6、A0A7S5C4I6、Q5CZ93、B4DI50、A0A7S5C2T1、A0A5C2GW38、A0A5C2GMJ1、A0A5C2G678、A0A024R1G6、A0A1L2BU33、A0A7S5C579、F8WE85、A0A1W2PQB1、A0A5C2GSR8、A0A5C2GNM9、A0A5C2GLW9、A0A5C2G013、A0A5C2GFT7、A0A5C2GGA1、A0A024R3Z8、A0A5C2G9U3、A0A5C2GE58、D6RHD5、A0A5C2G818、A0A5C2GAS1、A0A5C2GAM5、A0A5C2GN18、A0A8I5KTH0、A0A5C2GJH5、A0A0S2Z3Q4、A0A5C2G689、A0A5C2GGI1、Q9H9A7、A0A1B1CYC9、A0A5C2G9Z6、A0A2U8J953、A0A5C2GLL4、A0A024R4P6、B4E0R9、A0A024R944、A0A5C2GYN0、A2NYU9、A0A5C2GK50、Q9UL92、A0A5C2FVX0、Q5T7N2、A0A5C2G2U2、A0A5C2G130、A0A5C2FYY5、A0A7S5C233、A0A7S5C3W8、A0A7U3M537、A0A7S5BYW5、A0A5C2GN07、A0A7S5C4L0、A0A0K0K1H8、A0A024R2T9、A0A140VJI7、B4DZM3、A0A5C2GVX7、B0AZL7、A0A5C2G6X2、A0A7S5EXB6、A0A5C2GPU9、A0A5C2GW63、A0A5C2GG27、A0A5C2GCA5、A0A5C2GAV8、O14746、A0A7T0LNE0、Q5VYV0、P35900、P02533、A0A5C2GFS8、F6KPG5、A8K6Q8、A0A5C2GJB9、A0A4P8J4B8、A0A7S5C2Y4、A0A7S5C5D6、A0A5C2G779、A0A5C2GA64、A0A5C2G0G6、A0A0A0MRS8、A0A5C2GK68、A0A5C2GCZ6、A0A5C2GP91、A0A024R6N9、A0A5C2GE26、A0A5C2GDT4、A0A5C2G228、A0A5C2GSI8、A0A5C2G1I8、A0A1W6IYK0、A0A087WWT3、A0A5C2G509、A0A5C2GE84、A0A5C2G6K5、A0A0S2Z428、A0N7I9、A0A5C2GJ11、A0A7S5C6L0、A0A5C2H1X6、H6VRF8、A0A5C2GUM5、A0A5S8K7B6、A0PJA6、A0A7S5EUT2、Q562R1、A0N5G4、P33908、A0A7S5C007、A0A024R483、A0A024RDU9、A0A7S5BY19、A0A5C2FZ13、A4D1J9、A2VCK8、A0A5C2FU71、P20848、Q86TT1、A0A7T0LNL2、P60709、A0A7S5C014、A0A5C2GL27、A0A5C2G829、A0A5C2H2D3、A0A5C2G8M3、A0A5C2GFU7、P21333、A0A7S5EX37、A2KBC1、A0A7S5BYR8、A0A5C2GCK8、Q8IWU2、A0A5C2GCL9、A0A5C2FX21、A0A7S5C0X4、P02671、A0A140VJU8、A0A5C2GI49、A0A024R3I5、A0A5C2GBK0、Q8WXI4、A0A5C2FVX7、E7EVZ1、A0A5C2GP39、A0A7S5BXC5、A0A5C2H0H1、A0A7S5C215、A0A5C2GNZ7、A0A5C2GMR7、A0A7S5C420、A0A5C2G4L3、A0A5C2GTG1、A0A5C2GZG9、A2IPI4、A0A5C2GH20、A0A5C2FU97、A0A5C2GV09、A0A5C2GT63、A0A024R5Z7、A0A024R8N8、Q6N095、A0A1U9X8X5、A0A5C2G7E7、A0A5C2GCV9、A0A024R462、A0A5C2GD72、A0A5C2G8D9、A0A5C2GRG5、A0A5C2G6W9、A0A5C2GW65、A0A5C2GT29、A0A5C2GAM1、A0A024QZB1、Q7Z664、P02675、P02775、A0A5C2G455、A0A0A0MRJ7、A0A5C2GLN6、A0A7S5EX08 and D3DQH8, or their antigens / antibodies, or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0087] Orthogonal Partial Least Squares Discriminant Analysis (OPLS-DA): Protein expression matrices obtained by eight standardization methods were screened for protein biomarkers using OPLS-DA, and protein biomarkers that appeared in the results of at least five standardization methods were selected to form a set of protein biomarkers. Set ⑤ contains a total of 706 proteins or peptides, namely P0DJI8, O00617, A0A384MDQ7, Q86W24, A0A0G2JMS6, Q4TZM4, A0A024QZH6, P59666, A0A024R980, B7Z3V1, A0N071, A0A0A0MTQ8, A0A5C2GYJ4, Q49A33, A0A7S5C1G4, A0A7P0TAB0, C9JEE0, A0A0K2BMD8, P18428, A0A5C2G0Y8, B4DZH3, P02743, A0A024R3I5, A0A 5C2GBG1, A0A0S2Z3X8, P00738, A0A1W6IYJ7, A0A5C2GAY9, A0A087WU78, A0A5C2G684, A2VCQ3, A0A7S5BZA6, A8K6J9, A0A1Z1G4M2, A0A8I5K T93, A0A494C037, P05109, Q86YQ4, A0A5C2GE15, P02750, B2R6W1, A0A024R0R6, A0A075B6S9, Q68CN4, D9YZU5, A0A5C2GJ77, Q02985, E7EVZ1 , A0A7S5C0E1, B4E367, Q96T46, A0A024R6P0, A0A5C2G3I5, Q86TT2, A0A5C2G1J3, A0A5C2GR94, A0A024R035, P02763, A0A5C2GF59, Q06033, A0A1W6IYL8, G3V3A0, A0A7S5BY74, A0A024RBU2, A0A5C2GGG9, H9KVD5, A0A7S5C6B0, A0A5C2G7P4, A7L8C6, P0C0L4, A0A7S5EY08, B2RA39, A0 A7T0PXT4, A0A5C2G6Q8, A0A140KFU0, A0A1W2PPX9, A0A7S5C0E2, C8C504, A0A5C2GJ24, A0A385HVZ2, A7E2C5, Q7Z5J4, A0A7S5EUJ3, Q9H387 , A0A5C2GA89, A0A7T0LP36, A0A5C2FVG1, Q13790, A0A5C2G2X0, A0A7S5EUR0, A0A5C2GEX2, A0A5C2G755, A0A5C2GJ30, P43251, A0A024R4F9,A0A5C2FYC5、A0A2U8J8S3、A0A5C2GFT6、A0A2U8J8L9、P34096、H7C517、B4E1 Z4、A0A5C2GCD2、A0A5C2GE35、A0A590UK17、P19652、P0C0L5、A0A7S5C4U0、A 0A1L7B5J3, A0A5C2GIJ6, A0A5C2GDW0, B2RMS9, A0A8G1A656, P04275, A0A5C2GNN4, A0A5C2GVL1, A0A7S5C5T5, A0A2U8J987, A0A5C2GL36, B2R888, A0A7S 5C0J8、B2M1S7、A0A7S5BYI5、A0A5C2GVX4、A0A5C2GIP8、P00915、B4E3M1、A0 A5C2G1W1、A0A1W2PQX5、A0A5C2FY23、A0A5C2GEW6、P11226、A0A7S5BXX6、A0 A7S5C048, A0A7S5BZS9, A0A5C2FUN2, A0A140VK00, A0A5C2GMG8, P07357, A0A5C2GLZ3, E9KL23, A6NKX1, A0A7S5C173, A0A5C2GBC6, A0A7T0LNJ3, A0A5C2 FZ84, A0A3B3IQ51, A8K237, A0A5C2G8Y9, O14893, A0A5C2GFN7, Q03591, B2R5G8, A0A5C2GQ45, A0A3B3ISR2, A0A8F0WQF6, A0A5C2GDN1, Q13784, Q6N091 A0A7S5EXP5、A0A5C2G0Q5、P01031、A0A7S5C002、A0A5C2GEM8、A0A5C2GJF4、 A0A5C2G4G0、A0A7S5C0H1、A8K1K1、Q6L981、A0A5C2GN85、A0A494C1P3、A0A7 S5C215, B4DR59, B3KS98, A0A5F9ZGX3, P04003, A0A5C2FY76, A0A7S5C0A9, A0A5C2FZP6, A0A7S5C6L1, A0A7S5C3P9, A0A5C2GA92, A0A5C2GM78, A0A126GV G9, A0A5C2GN30, A0A1W6IYK2, A0A075B708, A0A5C2GT92, A0A5C2G699, A0A5C2FXM3, A0A5C2GA52, A0A7S5BZY2, A0A5C2H1E9, A0A5C2G3J8, A0A5C2GKT9A0A5C2GL08、A0A5C2GIK6、A0A5C2G4P2、P02790、A0A5C2GTJ5、A0A7S5C299、A0A5C2GJB2、A0A5C2GFV3、A2NV54、P09871、A0A024R6F8、A0A5C2GDA7、Q8NCL6、E7ETH0、A0A5C2GP10、A0A5C2GY49、Q6IQ49、A0A5C2GF13、A0A5C2G252、Q9Y6R7、A0A5C2GSQ5、A0A5C2G1Q3、A0A5C2G4L3、A0A5C2GG08、A5PL27、A0A5C2G5A1、G3XAK1、A0A5C2GQL0、A0A3G1HEF1、A6XNE2、A0A0D9SFG7、P35542、Q4VXF1、A0A5C2GDU6、A0A5C2G7A2、B2RBW9、A0A5C2GEX5、A0A5C2FYM3、A0A4V1EJ15、A0A5C2G6B6、P13671、P22352、A0A5C2GD51、A0A5C2GKS6、A0A5C2GFQ5、A0A5C2GBR0、P10643、A0A5C2GY47、Q9NQ79、A0A5C2GLL5、A0A7S5C5Z8、A0A5C2GDF6、A0A5C2GW00、P01024、A0A5C2GTA5、A0A5C2GHU5、A0A384N669、A0A7T0LNB3、P04217、B4E273、A0A384MEC0、A0A5C2GT38、A0A5C2GC48、Q6LAM1、A0A5C2GE65、A0A5C2GSJ6、A0A5C2FUP9、A8K9G4、A0A7S5C051、A0A5C2FVC6、A0A7S5EYK2、A0A7S5BYN9、A0A7S5BZH2、A0A5C2GAU8、A0A5C2FXP5、B4E215、A0A7S5EX52、A0A7S5C4M5、B2R582、A0A024RAB9、A0A5C2H0U9、A0A5C2GF54、A0A5C2GTV7、A0A5C2GBQ5、A0A5C2GVX1、A0A7S5C442、A0A0S2Z4I5、A0A5C2GV93、A0A5C2GH71、A0A5C2FW89、A0A7S5C300、A0A7S5EXI9、A0A7S5EX91、A0A7I2V2D2、A0A7S5EYC3、A0A5C2FU97、A0A0X9V9B3、A0A087WZ31、A0A5C2GXN9、A0A5C2GRA7、A0A5C2GTP5、O75882、A0A5C2G5D9、A0A5C2G4T9、A0A5C2GET4、A0A7S5EXJ5、O00634、A0A5C2FWW2、A0A7S5C336、A0A1W6IYJ9、A0A7S5C775、A0A5C2GPL4、H0Y973、A2J1N7、A0A109PS32、A0A4P8J6J1、A0A7S5C0G1、A8K9M5、A0A5C2GM50、A0A5C2GKE6、A0A5C2FWN1、A0A7S5C1E3、A0A0C4DGB6、B4DUV1、A0A7S5EXZ4、D3DNN4、A0A5C2G921、A0A5C2G292、A0A5C2GNS8、P80108、A0A024R5Z7、A0A5C2GHA6、A0A7S5C3I9、L8E853、A0A5C2GA26、A0A5C2GN71、V9H1D9、Q9BYH2、A0A7T0LN44、A0A5C2FZZ3、A0A5C2FYG2、A0A5C2G1A1、A0A0S2Z4F1、B4E0X1、A0A5C2GAA5、A0A7S5C0M2、A0A5C2GJC8、A0A5C2G5T8、A0A5C2G7K4、A0A5C2GDY1、Q9NXP7、P07358、A0A5C2GGY7、A0A5C2GGI5、A0A140VK24、A0A5C2G2L3、A0A5C2GNC7、A0A5C2G3P7、A0A384MEF1、A0A0M3KKW6、P15169、Q9UMS6、A0A5C2FX28、A0A5C2GVG6、A0A5C2G6D1、A0A7S5BY02、I3L145、A0A2U8J8X2、A0A5C2FXY6、P35527、A0A5C2G7L1、E9PHK0、A0A5C2G4T8、A0A5C2G1R6、Q8NHV9、A0A193AUV8、A0A5C2GUV2、A0A7S5BYJ7、A0A5C2FYJ4、A0A5C2GFC3、A0A5C2FZ01、A0A7S5EX22、A0A2U8J9A8、A0A5C2GXE0、A0A024R8N8、A0A5C2GMF2、A0A5C2G9T0、A0A5C2G0Q1、A0A024R892、B3KUW8、A0A7S5EVM6、A0A5C2GY06、H0YAC1、A0A5C2G525、A0A5C2GXG7、A0A5C2FWW6、A0A024R6I9、A0A5C2GCR5、A0A024R8G3、A0A7S5EWG2、A0A5C2GJQ5、A0A5C2GWI4、P02753、A0A7S5BY47、A0A5C2GGH5、A0A5C2FY35、A0A5C2G2P9、A0A5C2GW19、A0A7S5EV42、Q7Z4J2、A0A5C2GJ96、A0A5C2GPG2、A0A5C2FVW2、A0A5C2G9I0、A0A5C2GX61、A0A1B0GVI3、P47992、A0A5C2G4U1、F8VVF9、A0A7S5C361、F2Z2G4、A0A0U1RQQ9、A0A1B0GU24、A0A5C2GE11、A0A5C2GX13、P14151、A0A7S5EVT4、A0A5C2GR38、A0A5C2G9Q6、A0A5C2G1Y7、A0A5C2G2K3、A0A5C2GX17、A0A5C2GQE6、A0A5C2GQJ7、A0A5C2GJ23、A0A7S5EUE3、A0A286YEY4、S6BGD6、Q6ZW64、A0A7S5C031、Q15475、D9IWP9、A0A1L2BU44、Q2L9S7、P27169、A0A5C2G989、A0A5C2G4A0、A0A5C2FXB2、A0A7S5BXA9、A0A5C2GHZ0、P02760、A0A5C2GDU8、A0A5C2G401、A0A5C2GQE4、A0A5C2G7E0、A0A5C2GF85、A6XGL1、A0A5C2FWX3、A0A7T0PXH4、A0A7S5C2G0、A0A5C2FVY0、P06727、A0A161I202、A0A7U3R8A8、A0A5C2GBF4、A0A067XG54、A0A024R6C9、A0A7P0TAI0、A0A5C2FV95、A8K669、A0A7S5C563、B4DI57、A0A5C2G8Y7、A0A5C2GST8、A0A7S5C177、A0A7S5C1W9、A0A7S5EU68、A0A5C2G772、A0A2U8J9D3、A8K6C9、A0A0C4DH35、A0A7S5EWA8、A0A5C2G6H4、A0A5C2FUM5、A0A5C2G826、A0A5C2GAY0、A8K3I0、A0A126LAY7、A0A5C2FZA5、Q0ZCI2、A0A5C2GSW5、A0A5C2G368、A0A5C2GNX8、A0A024R6I6、A0A7S5EXS6、A0A5C2GAJ3、B7WNR0、Q86SQ4、Q0ZCJ1、A0A024RCW3、A0A5C2G8T6、A0A7S5C2A9、P43652、A0A5C2FZJ0、A0A5C2FWP4、A0A7S5BZF8、A0A7S5C441、A0A7S5BYB3、Q92496、A0A3B0J0F2、A0A5C2G8L6、P35908、A0A5C2FYC6、B7ZM24、H7BYG8、A0A5C2GE10、A0A7S5C2N4、P06702、P02652、A0A5C2GDW3、Q6N022、A0A5C2GAI3、A0A5C2GCV9、A0A5C2G3I9、B2R4C5、A0A5C2G7F2、A0A024R3E3、B0YIW2、A0A024R1G8、B4DZM1、A0A7S5BY99、A0A5C2GMN5、A0A7S5BYQ0、A0A5C2GK31、P08697、A0A5C2GBU2、A0A5C2GX44、A0A5C2GBH7、A0A5C2FSY5、Q9Y4F3、A0A5C2G6W9、A0A5C2GGS8、A0A5C2GF70、A2J1M5、A6XND0、A0A7S5C781、A0A7S5BYT5、A0A5C2GKI1、A0A5C2GZS9、B4DHZ6、A0A5C2GQE2、A0A5C2G5W9、D3DQH8、A0A5C2GH04、A0A5C2G6G6、Q5CZ93、B4DI50、A0A5C2GW38、A0A5C2G678、A0A024R1G6、A0A7S5C579、A0A1W2PQB1、A0A5C2G2K8、A0A5C2GNM9、A0A5C2GLW9、A0A5C2G013、A0A5C2G9U3、A0A5C2GE58、A0A5C2GQT8、A0A5C2G818、A0A5C2GAS1、A0A5C2FZ24、A0A5C2GJH5、A0A5C2G689、B7Z7R8、A0A5C2GGI1、A0A1B1CYC9、A0A2U8J953、A0A024R4P6、B4E0R9、A0A024R944、A0A7S5C0L7、A0A7S5ETA2、A2NYU9、A0A5C2GK50、A0A5C2FVX0、A0A5C2GGG1、A0A5C2G2U2、A0A5C2G130、A0A5C2GG47、A0A5C2FYY5、A0A7S5EYY8、A0A7U3M537、A0A7S5BYW5、A0A5C2GN07、A0A024R2T9、A0A140VJI7、B4DZM3、B0AZL7、A0A5C2G6X2、A0A5C2GPU9、A0A5C2GG27、O14746、Q5VYV0、P35900、P02533、A0A5C2GFS8、A8K6Q8、A0A7S5C2Y4、A0A7S5C5D6、A0A5C2G779、A0A5C2G0G6、A0A0A0MRS8、A0A5C2GK68、A0A5C2GP91、A0A2U8J927、A0A024R6N9、A0A5C2GDT4、A0A5C2GSI8、A0A5C2G3G9、Q7Z351、A0A5C2G509、A0A5C2GE84、A0A5C2G6K5、A0A0S2Z428、A0A5C2GJ11、H6VRF8、B4DPP8、A0A5C2GUM5、A0A5S8K7B6、Q562R1、P33908、A0A024RDU9、A0A7S5BY19、A0A5C2FZ13、A4D1J9、A2VCK8、A0A5C2FU71、P20848、Q86TT1、P60709、A0A5C2G3G8、A0A7S5C014、A0A5C2GL27、A0A5C2G829、P21333、A0A7S5EX37、A2KBC1、A0A5C2FYJ2、A0A5C2GCK8、A0A5C2GCL9、A0A7S5C0X4、P02671、A0A140VJU8、Q9UNU2、A0A5C2G7X1、D9ZGG2、P02775、E7EQ64、A0A0D9SF05、A0A2U8J993、A0A5C2GRW3、A0A5C2GHP0、A0A7S5EVS2、A0A5C2GSJ0、A0A5C2G3E4、P01871、A0A5C2GT19、A0A5C2FZR0、A0A2U8J8I9、A0A5C2G7Y6、A0A5C2G3H4、A0A5C2FUD7、Q96IY4、A0A5C2GL60、A0A5C2GBQ2、A0A5C2GYW0、A0A7T0LND4、A0A5C2G7U8、A0A7S5C3K5、A0A5C2GT63、A0A5C2GSL9、A8KAP9、A0A5C2G2E8、A0A1B0GX98、P00748、A0A5C2FZ71、A0A5C2GAJ1、A0A5C2FZM5、A0A7S5BZC0、Q9UGM5、A0A7S5BY64、A0A5C2G280、A0A5C2GU54、E7ETY2、A0A5C2GWG8、A0A5C2FT87、A0A7S5C2D6, A0A5C2GLU4, A0A1L2BU33, A0A5C2G435, A0A5C2GXD7, A0A5C2GCZ6, A0A5C2GAM1, A0A024QZB1, Q7Z664, P02675, A0A024RDY9, P20851, A0A5C2FVV5, and A0PJA6, or their antigens / antibodies, or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0088] RFS-based random forest: Protein expression matrices obtained by eight normalization methods were screened for protein biomarkers using RFS-based random forest, and a set of protein biomarkers obtained by at least one of the normalization methods was selected for subsequent analysis. Set ⑥ contains a total of 147 proteins, namely A2VCK8, P02750, Q06033, A0A024R035, H7BYG8, E9KL23, P59666, H6VRF8, Q4TZM4, A0A024RCW3, A4D1J9, P04275, O14746, P02533, B7WNR0, A2VCQ3, P35527, D9YZU5, B4E273, A0A0K2BMD8, A0A0S2Z5S9, A0A5C2FXF8, A0A1Z1G4M2, M0QX59, V9H1D9, P18428, A0A0 A 0A5C2G525, A0A7S5EXS6, A0A0C4DH30, A0A5C2GAY9, E7EQY3, O00617, B4E215, A0A384MEF1, B4DZM3, F8VVF9, A0A5C2GTH5, A0A5C2GGH5, A0 A5S8K7B6, A0A126LAY7, A0A5C2G6X2, A0A5C2GDY8, Q7Z5J4, B4DHZ6, P02671, A0A140KFU0, S6C4S0, A0A7S5BY88, A0A0D9SFG7, P27169, A0A 5C2GMN5, P02675, A0A5C2FVS9, A0A0A0MRS8, A0A286YEY4, G3V3A0, B2R4C5, B4DIJ4, A0A5C2G3K7, A0A5C2GSJ2, A0A5C2GCL9, A0A5C2G1Q3, A0A5C2FX55, A0A7S5C4Q8, A0A1B0GVI3, A0A5C2GBQ5, A0A7P0T7Z2, A0A024QZ94, A0A5C2GW09, A7E2C5, A0A7S5EYK2, A0A5C2FU97, A0A5C2G 0A9, Q2L9S7, A0A8F0WQF6, P35900, A0A5C2FUE7, A0A5C2FZX0, A0A5C2FVR7, A0A494C1P3, A0A494C0I1, A0A5C2FX67, A0A384MEC0, H0YAC1,A0A7S5BZH5, A0A5C2GJH5, A0A5C2GKL4, Q6IQ49, A0A2Z4LCH4, B4DP44, P0DOX3, A0 A5C2G4U1, A0A024R0R6, P02652, A6XNE2, A0A5C2GXX2, A0A075B7D4, P80108, P0510 9. P43652, Q7Z4J2, A0A5C2GMM1, A0A5C2G0Q3, A0A0S2Z3F6, Q96JA8, A0A5C2GHU5, Q86TT1, P06727, A0A7T0PWX7, A0A5C2GI12, A0A5C2FXQ9, A0A075B7A1, A0A5C2FVW2 , Q8NHV9, A0A5C2GKW4, A0A5C2FUG7, A0A5C2G772, V9HWI6, A0A5C2FX99, A0A5C2GC K8, A0A5C2GJI3, A0A7S5BZY2, A0A5C2G130, A0A7S5EXD1, A0A090N7U9, A0A5C2FXP5 P00915, A0A0S2Z4I5, B4E367, A0A5C2GCY5, A0A5C2FTW0, A0A7S5BZX6, P35908, A0A7S5EYD3, and A0A140VJU8, or their antigens / antibodies, or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0089] Mann-Whitney U test: The protein expression matrices obtained by the eight standardization methods were screened for protein biomarkers using the Mann-Whitney U test. A set of protein biomarkers with p-values less than 1.00E-10 under all eight standardization methods was selected for subsequent analysis. Set ⑦ contains a total of 94 proteins, namely A0A024QZH6, A0A024R035, A0A024R0R6, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A087WZ31, A0A0K2BMD8, A0A0S2Z3X8, A0A0S2Z4F1, A0A0S2Z4I5, A0A126LAY7, A0A140KFU0, A0A1B0GVI3, A0A1B1CYC9, A0A1W2PQX5, and A0A1Z1G4M2. , A0A2U8J8L9, A0A384MEC0, A0A384MEF1, A0A3B3ISR2, A0A5C2FVY0, A0A5C2G7F2, A0A5C2GD51, A0A5S8K7B6, A0A7I2V2D2, A0A7P0 TAB0, A0A7S5C336, A0A7S5EYK2, A0A8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A7L8C6, A8K3I0, B2R4C5, B2R582, B 2R888, B2RMS9, B4DZH3, B4E1Z4, B4E215, B4E273, B4E3M1, B7WNR0, D3DNN4, D9YZU5, E9KL23, E9PHK0, G3V3A0, H6VRF8, H7BYG8, H7 C517, H9KVD5, I3L145, L8E853, O14746, O14893, P00738, P01031, P02533, P02671, P02750, P02753, P02763, P04275, P05109, P06 727, P09871, P0C0L4, P0DJI8, P18428, P27169, P35527, P35542, P35900, P43652, P59666, P80108, Q03591, Q06033, Q2L9S7, Q49A33, Q4TZM4, Q4VXF1, Q7Z4J2, Q8NHV9, Q9H387, Q9NXP7, Q9UMS6 and V9H1D9 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0090] Welch's t-test: The protein expression matrices obtained by the eight standardization methods were screened for protein biomarkers using Welch's t-test. The set of protein biomarkers whose p-values were all less than 1.00E-10 under the eight standardization methods was selected for subsequent analysis. Set ⑧ contains a total of 84 proteins, namely A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0K2BMD8, A0A0S2Z3X8, A0A0S2Z4F1, A0A0S2Z4I5, A0A126LAY7, A0A1B0GVI3, A0A1B1CYC9, A0A1W2PQX5, A0A1Z1G4M2, and A0A2U8J8. L9, A0A384MEC0, A0A384MEF1, A0A3B3ISR2, A0A5C2FVY0, A0A5C2G7F2, A0A5C2GD51, A0A5S8K7B6, A0A7I2V2D2, A0A 7S5C336, A0A8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A7L8C6, A8K3I0, B2R4C5, B2R582, B2R888, B 2RMS9, B4E1Z4, B4E215, B4E273, B4E3M1, B7WNR0, D3DNN4, D9YZU5, E9KL23, E9PHK0, G3V3A0, H6VRF8, H7BYG8, H7C 517, H9KVD5, L8E853, O14746, P00738, P01031, P02533, P02671, P02750, P02753, P02763, P04275, P05109, P06727 P07358, P09871, P0C0L4, P0DJI8, P18428, P27169, P35527, P35900, P43652, P59666, P80108, Q03591, Q06033, Q4TZM4, Q4VXF1, Q7Z5J4, Q8NHV9, Q9H387, Q9NXP7 and V9H1D9 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0091] Odds ratio: Protein expression matrices obtained by eight normalization methods were screened for protein biomarkers using the Odds ratio. A set of ⑨ consisting of protein biomarkers with p-values less than 1.00E-10 under all eight normalization methods was selected for subsequent analysis. Set ⑨ contains a total of 69 proteins, namely A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A087WZ31, A0A0K2BMD8, A0A0S2Z3X8, A0A0S2Z4F1, A0A0S2Z4I5, A0A1B0GVI3, A0A1W2PQX5, A... 0A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, A0A3B3ISR2, A0A5C2FVY0, A0A5C2G7F2, A0A7I2V2D2, A0A 8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A7L8C6, A8K3I0, B2R4C5, B2R888, B2RM S9, B4E1Z4, B4E273, B7WNR0, D3DNN4, D9YZU5, E9KL23, E9PHK0, G3V3A0, H6VRF8, H7BYG8, H7C517 , L8E853, O14746, P01031, P02533, P02671, P02750, P02753, P02763, P04275, P05109, P06727, P 09871, P0C0L4, P0DJI8, P18428, P35900, P43652, P59666, P80108, Q03591, Q06033, Q4TZM4, Q4VXF1, Q8NHV9, Q9H387, Q9NXP7 and V9H1D9 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0092] P-values: Protein expression matrices obtained by eight standardization methods were used to screen for protein biomarkers. A set ⑩ consisting of protein biomarkers with p-values less than 1.00E-10 under all eight standardization methods was selected for subsequent analysis. Set ⑩ contains a total of 76 proteins, namely A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A087WZ31, A0A0K2BMD8, A0A0S2Z3X8, A0A0S2Z4F1, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0A1Z1G4M2, A0 A2U8J8L9, A0A384MEF1, A0A5C2G7F2, A0A5S8K7B6, A0A7I2V2D2, A0A7P0TAB0, A0A8F0WQF6, A0A8G1A656 , A2VCK8, A2VCQ3, A4D1J9, A5PL27, A7L8C6, A8K3I0, B2R4C5, B2R888, B2RMS9, B4DZH3, B4E1Z4, B4E215, B4E273, B4E3M1, B7WNR0, D3DNN4, D9YZU5, E9KL23, E9PHK0, G3V3A0, H6VRF8, H7BYG8, H7C517, H9KVD5, L 8E853, O14746, P00738, P01031, P02533, P02671, P02750, P02753, P02763, P04275, P05109, P06727, P0 9871, P0C0L4, P0DJI8, P18428, P35527, P35900, P43652, P59666, Q03591, Q06033, Q49A33, Q4TZM4, Q4VXF1, Q8NHV9, Q9H387, Q9NXP7, Q9UMS6 and V9H1D9 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0093] Step S8, the biomarker summary analysis, includes: summarizing and analyzing the biomarkers in sets ① to ⑩, resulting in 11 protein biomarkers appearing in all 10 sets. These 11 proteins constitute the target protein biomarker set G11, which includes P59666, P02750, Q06033, P04275, H7BYG8, O14746, P02533, H6VRF8, A4D1J9, A2VCK8, and P02671, or their antigens. The target protein biomarker set G12 consists of 41 protein biomarkers, including antibodies or their single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins or stable isotope characteristic peptide segments, and corresponding nucleic acids; and antibodies or their single peptide chains or characteristic peptide segments of stable isotopes. The set also includes A0A0G2JMS6, Q4TZM4, P59666, P02750, B4E367, Q06033, P04275, E9KL23, and A0A7S5BZ. Y2, Q6IQ49, A0A5C2G1Q3, A6XNE2, A0A7S5EYK2, A0A5C2FXP5, V9H1D9, E9PHK0, Q8NHV9, H0YAC1, A0 A5C2FVW2, A0A5C2G4U1, Q2L9S7, P27169, A0A7S5EWA8, A0A126LAY7, P43652, P35908, H7BYG8, P026 52. Q5CZ93, A0A5C2GJH5, A0A5C2G130, O14746, P02533, A0A0A0MRS8, H6VRF8, A0A5S8K7B6, A4D1J9, A2VCK8, Q86TT1, A0A5C2GCL9 and P02671 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids;The target protein marker set G13 consists of 51 markers appearing in any 7 of the 10 sets ① to ⑩. This set includes P59666, P02750, Q06033, P04275, E9KL23, V9H1D9, E9PHK0, Q8NHV9, P43652, H7BYG8, O14746, P02533, H6VRF8, A0A5S8K7B6, A4D1J9, A2VCK8, P02671, P0DJI8, A0A024QZH6, A0A0K2BMD8, P18428, A0A0S2Z3X8, P00738, A2VCQ3, A0A1Z1G4M2, and P05109. D9YZU5, A0A024R6P0, A0A024R035, P02763, G3V3A0, A0A2U8J8L9, A0A8G1A656, B4E3M1, A0A8F0WQF6, Q4VXF1, B4E273, L8E853, A0A384MEF1, P35527, P02753, A0A1B0GVI3, P06727, A8K3I0, B7WNR0, A0A024RCW3, B2R4C5, A0A5C2G7F2, A0A024R1G6 and P35900 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and corresponding nucleic acids;After merging sets G12 and G13 and removing duplicates, a total of 74 important biomarkers were obtained. These 74 proteins constitute the target protein biomarker set G14, which includes A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A126LAY7, A0A1B0GVI3, and A0A1Z1. G4M2, A0A2U8J8L9, A0A384MEF1, A0A5C2FVW2, A0A5C2FXP5, A0A5C2G130, A0A5C2G1Q3, A0A5C2G4U1, A0A5C2G7F 2. A0A5C2GCL9, A0A5C2GJH5, A0A5S8K7B6, A0A7S5BZY2, A0A7S5EWA8, A0A7S5EYK2, A0A8F0WQF6, A0A8G1A656, A2 VCK8, A2VCQ3, A4D1J9, A6XNE2, A8K3I0, B2R4C5, B4E273, B4E367, B4E3M1, B7WNR0, D9YZU5, E9KL23, E9PHK0, G3 V3A0, H0YAC1, H6VRF8, H7BYG8, L8E853, O14746, P00738, P02533, P02652, P02671, P02750, P02753, P02763, P04 275, P05109, P06727, P0DJI8, P18428, P27169, P35527, P35900, P35908, P43652, P59666, Q06033, Q2L9S7, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q8NHV9 and V9H1D9 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0094] The correlation analysis in step S9 includes the following: When using machine learning methods such as random forests for feature selection, some important biomarkers may be missed. This is usually because these methods rely on the statistical properties of the data to evaluate the importance of each feature, rather than on direct biological knowledge. Therefore, we perform a correlation analysis on the 41 important biomarkers in set G12 obtained in the previous step with all 3611 biomarkers, and form a new set P1 with the biomarkers whose correlation coefficients with these biomarkers are within a first range and the target protein biomarkers. In one embodiment, the first range is: greater than or equal to 0.65 or less than or equal to -0.65. These first markers, together with the 41 markers, form a new set P1 containing 75 markers, namely Q06033, E9KL23, P02750, P02671, P04275, A2VCK8, A4D1J9, P59666, Q2L9S7, P02533, A0A126LAY7, H0YAC1, A0A0A0MRS8, A0A7S5EYK2, P43652, Q4TZM4, V9H1D9, and A0A5S8K. 7B6, P27169, Q8NHV9, E9PHK0, O14746, H7BYG8, A6XNE2, A0A5C2G130, H6VRF8, A0A5C2GCL9, P35908, A0A7S5BZY2 , A0A5C2G4U1, A0A5C2FVW2, A0A5C2FXP5, A0A5C2GJH5, B4E367, A0A0G2JMS6, A0A7S5EWA8, Q86TT1, Q6IQ49, Q5CZ 93. P02652, A0A5C2G1Q3, A0A024R035, A0A024R6P0, A0A1Z1G4M2, Q06033, P02763, H7C517, P18428, B2RMS9, Q9H 387. A0A1W2PQX5, P0DJI8, B2R888, A0A024QZH6, L8E853, A0A024RCW3, B7WNR0, A0A024R1G6, A0A0X9V9B3, A0A14 0KFU0, D9YZU5, P00915, B2M1S7, A0A0K2BMD8, A0A5C2FYJ4, A0A5C2FZZ3, A0A5C2GBQ5, A0A5C2GMN5, A0A5C2GN07, A0A7I2V2D2, A0A7S5BYW5, P35900, O00617, P06727 and P35527 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments,The corresponding nucleic acids. Simultaneously, we performed correlation analysis between the 74 important biomarkers of set G14 obtained in the previous step and all 3611 biomarkers. Biomarkers with correlation coefficients greater than or equal to 0.65 or less than or equal to -0.65 were grouped with these 74 biomarkers to form a new set P2 containing 119 biomarkers. These biomarkers are A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A0X9V9B3, and A0A126. LAY7, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, A0A385HVZ2, A0A5C2FTW7, A0A5C2FVW2, A0A5C2FX67, A0A5C2FXP 5. A0A5C2FYJ4, A0A5C2FYK5, A0A5C2FZZ3, A0A5C2G130, A0A5C2G1Q3, A0A5C 2G1X1, A0A5C2G3M4, A0A5C2G4U1, A0A5C2G6G6, A0A5C2G7F2, A0A5C2GA97, A 0A5C2GBQ5, A0A5C2GCL9, A0A5C2GDA6, A0A5C2GGY3, A0A5C2GJH5, A0A5C2GJ L5, A0A5C2GMN5, A0A5C2GN07, A0A5C2GRZ5, A0A5C2GV43, A0A5C2GW23, A0A5 C2H3N5, A0A5S8K7B6, A0A7I2V2D2, A0A7P0TAB0, A0A7S5BYW5, A0A7S5BZY2, A0A7S5C366, A0A7S5EWA8, A0A7S5EYK2, A0A7T0LP36, A0A8F0WQF6, A0A8G1A 656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A6XNE2, A8K3I0, B2M1S7, B2R4C5, B2 R888, B2RMS9, B4DZH3, B4E1Z4, B4E273, B4E367, B4E3M1, B7WNR0, D9YZU5, E 9KL23, E9PHK0, G3V3A0, H0YAC1, H6VRF8, H7BYG8, H7C517, H9KVD5, L8E853, O00617, O14746, P00738, P00915, P01031, P02533, P02652, P02671, P02750,P02753, P02763, P04275, P05109, P06727, P0C0L5, P0DJI8, P18428, P27169, P35527, P35900, P35908, P43652, P59666, Q06033, Q2L9S7, Q49A33, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q86YQ4, Q8NHV9, Q9H387, Q9NXP7, V9H1D9 or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids. Some of these 119 proteins showed high correlation, as detailed in Table 1. Higher values indicate higher correlation, suggesting that highly correlated proteins are substitutes for each other and can be used in the modeling process.
[0095] Table 1. Results of correlation analysis among 119 biomarkers and between them and sample groups.
[0096] Step S10, variable structuring, includes: arranging and organizing information such as protein, peptide, or antibody quantification results and patient basic information in a certain order to form multiple independent variables; then dividing the samples into gastric cancer group and non-gastric cancer group according to clinical diagnosis, and encoding them as 1 and 0 respectively to form dependent variables; finally, organizing the processed independent and dependent variables into a two-dimensional table or data matrix, with each column representing a variable and each row representing a sample, to facilitate the subsequent use of this information to build a predictive model.
[0097] Step S11, model training and testing, includes: dividing the data into training and testing sets; constructing a regression model containing independent variables based on the training set; and testing the performance of the regression model based on the testing set. The data includes multiple independent and dependent variables.
[0098] Step S12, model optimization, includes: selecting effective independent variables based on any one of the following: the p-value of each independent variable in the regression analysis results, the correlation coefficient, and the AUC values of the regression model on the training and test sets, to obtain the selected effective independent variables. Specifically, if the p-value of an independent variable is less than 0.05, that effective independent variable is included in the selected effective independent variables.
[0099] In some embodiments, after performing step S12, step S13 may be included: model iteration, which includes repeating the above steps S11 and S12 to construct multiple prediction models and summarize their result data on the training set and test set to evaluate the performance of different prediction models.
[0100] The order and numbering of steps S2 to S12 above are not used to restrict the execution order of each step.
[0101] The inventors of this application conducted research on the screening and simplification of existing gastric cancer biomarkers and discovered that if existing gastric cancer biomarkers are structured into independent variables, and then samples are grouped and coded according to clinical diagnoses to form dependent variables, then variable screening can be performed through correlation analysis and statistical tests between the independent and dependent variables. The ability of an independent variable or gastric cancer biomarker to distinguish different clinical groups is judged by the magnitude of the correlation coefficient between the independent and dependent variables. The larger the correlation coefficient, the stronger the ability of the independent variable or gastric cancer biomarker to distinguish different clinical groups, and the more likely it should be prioritized in constructing gastric cancer prediction models. The significance of the numerical distribution of the independent variable between two clinical groups and the size of the p-value can also be used to determine the distinguishing ability of the independent variable / gastric cancer biomarker for the dependent variable or different clinical groups. Meanwhile, the inventors also discovered that correlation analysis between independent variables or gastric cancer biomarkers can reveal their relationships. Two strongly correlated independent variables or gastric cancer biomarkers can be substituted for each other in model construction or practical application, and only one needs to be selected in model construction, thus achieving model simplification. Furthermore, the inventors found that if a logistic regression model containing all variables, including effective variables, is constructed on the training set, and the model's performance is tested on the test set, the p-values, coefficients, and overall AUC of each variable in the logistic regression results, combined with the aforementioned correlation analysis and statistical test results, can further screen variables, narrowing down the set of effective variables, thereby achieving excellent gastric cancer biomarker screening and model simplification effects. Based on the above findings, this application proposes a method for constructing a model based on the Uniprot database to assess the probability of subjects developing gastric cancer.
[0102] Based on the above steps, this application obtained a total of 17 models, which are described below in conjunction with Examples 2 to 18. The training and test sets used in each model are identical. For each model, the Receiver Operating Characteristic (ROC) curve and the Area Under the Curve (AUC) were used as evaluation metrics, and these metrics were calculated on the training, validation, and test sets. The ROC curve, by comprehensively judging the model's ability to distinguish between positive and negative samples, can intuitively reflect the model's classification performance; while the AUC value can quantify the ROC curve, further evaluating the model's predictive sensitivity and specificity. Performing ROC and AUC analyses can comprehensively evaluate the effectiveness of our constructed predictive models, providing a basis for subsequent model optimization and clinical application.
[0103] Example 2 Establishment and validation of logistic regression prediction model 1
[0104] In Example 2, the effective protein biomarkers are P02750 and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. On the training set, a logistic regression model containing two independent variables, P02750 and P59666, was constructed to predict the probability of a sample having gastric cancer. First, based on the quantitative data of P02750 and P59666 and the sample labels, prediction model 1 was obtained using a logistic regression algorithm. The mathematical expression for prediction model 1 is: P = 1 / (1 + Exp(-(β0 + β1 × P02750 + β2 × P59666))), where P represents the probability of a sample being classified as a positive example. In the formula, the names of the protein biomarker or its single peptide chain, characteristic peptide segments of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segments, and the corresponding nucleic acid represent the standardized content values of the protein biomarker or its single peptide chain, characteristic peptide segments of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segments, and the corresponding nucleic acid. The coefficient β0 ranges from [-35.47, -26.94]; the coefficient β1 ranges from [1.79, 2.50]; and the coefficient β2 ranges from [0.60, 1.05]. Preferably, β0 = -31.21, β1 = 2.14, and β2 = 0.82 yield the optimal prediction results. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of its single peptide chain, characteristic peptide segments, stable isotope protein, or stable isotope characteristic peptide segments in new samples P02750 and P59666 or antibodies, or their single peptide chains, characteristic peptide segments of peptide chains, and stable isotope proteins or stable isotope characteristic peptide segments.
[0105] The ROC curve and corresponding AUC results of prediction model 1 on the training set are as follows: Figure 2As shown in the figure, the ROC curve and corresponding AUC results of the model on the training set show that the overall performance of the model is good, and the AUC value is also high at 0.89. This indicates that the model has good predictive performance in distinguishing between gastric cancer samples and non-gastric cancer samples.
[0106] To comprehensively verify the model's stability and generalization ability, a 5-fold cross-validation method was used for internal cross-validation. Specifically, the training set was divided into five mutually exclusive subsets. Four subsets were used as the training set in each round, and the remaining subset was used as the validation set, for a total of five rounds of training and validation. 5-fold cross-validation can measure the model's stability across different data subsets. The ROC curve and corresponding AUC results for the model's 5-fold cross-validation on the training set are shown below. Figure 3 As shown in the figure, the ROC curve and corresponding AUC results of the 5-fold cross-validation show that, after 5-fold cross-validation, the final model achieved an average AUC of 0.88 and a standard deviation of 0.03 on the validation set, meeting the set performance target. This also proves that the model is stable and effective, providing a guarantee for the practical application of the model.
[0107] After obtaining satisfactory training set results, prediction model 1 is applied to an independent test set for prediction, and various evaluation metrics of the model on the test set are calculated. This process verifies the model's generalization ability from the training set to the test set, examines whether the model overfits on unseen samples, and thus further evaluates the model's generalization performance. The ROC curve and corresponding AUC results of prediction model 1 on the test set are shown below. Figure 4 As shown in the figure, the ROC curve and corresponding AUC results of prediction model 1 on the test set indicate that the overall performance of prediction model 1 on the test set is good, and the corresponding AUC value is also high at 0.85. This indicates that prediction model 1 maintains a good ability to distinguish between gastric cancer samples and non-gastric cancer samples on the test set. Furthermore, the AUC on the test set is not significantly different from the AUC on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, prediction model 1 has good predictive performance on the test set and is suitable for clinical validation and application.
[0108] According to prediction model 1, only two indicators, P02750 and P59666, are needed to accurately predict the likelihood of a subject developing gastric cancer, which is low-cost and highly efficient.
[0109] In Examples 3 to 17, the same method as in Example 2 was used to obtain the ROC curves and corresponding AUC results on the training set, the ROC curves and corresponding AUC results of 5-fold cross-validation, the ROC curves and corresponding AUC results on the test set, and related ROC curve graphs, which were used to evaluate the predictive performance of prediction models 2 to 16. The following will only describe these results without referring to the accompanying figures; identical content will not be repeated.
[0110] Example 3 Establishment and validation of logistic regression prediction model 2
[0111] In Example 3, the effective protein biomarkers are B7WNR0 and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. On the training set, a logistic regression model containing two independent variables, B7WNR0 and P59666, was constructed to predict the probability of a sample having gastric cancer. First, based on the quantitative data of B7WNR0 and P59666 and the sample labels, prediction model 2 was obtained using a logistic regression algorithm. The mathematical expression for prediction model 2 is: P = 1 / (1 + Exp(-(β0 + β1 × B7WNR0 + β2 × P59666))), where P represents the probability of a sample being classified as a positive example. In the formula, the names of the protein biomarker or its single peptide chain, characteristic peptide segments of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segments, and the corresponding nucleic acid represent the standardized content values of the protein biomarker or its single peptide chain, characteristic peptide segments of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segments, and the corresponding nucleic acid. The coefficients β0, β1, and β2 range from [-13.12, -7.61] to [-0.57, -0.28] and [1.00, 1.40]. Preferably, β0 = -10.37, β1 = -0.42, and β2 = 1.20 yield the optimal prediction results. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of B7WNR0 and P59666 or their antibody or its single peptide chain, characteristic peptide segments of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification.
[0112] The ROC curve of prediction model 2 on the training set shows good overall performance and a high AUC value of 0.82, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.02 corresponding to the 5-fold cross-validation ROC curve of prediction model 2 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 2 on the test set also shows good overall performance and a high AUC value of 0.81. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is comparable to that on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0113] Example 4 Establishment and validation of logistic regression prediction model 3
[0114] In Example 4, the effective protein biomarkers are A0A024R6P0 and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. A logistic regression model with two independent variables, A0A024R6P0 and P59666, was constructed on the training set. First, based on the quantitative data of A0A024R6P0 and P59666 and the sample labels, prediction model 3 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 3 is: P = 1 / (1 + Exp(-(β0 + β1 × A0A024R6P0 + β2 × P59666))); where P represents the probability of a sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids refer to the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The coefficient β0 ranges from [-36.64, -27.29]; the coefficient β1 ranges from [1.55, 2.37]; and the coefficient β2 ranges from [0.79, 1.22]. Preferably, β0 is -31.96, β1 is 1.96, and β2 is 1.01, yielding the optimal prediction results. The established logistic regression model can predict the likelihood of a new sample developing gastric cancer based on A0A024R6P0 and P59666, their antibodies, single peptide chains, characteristic peptide segments, stable isotope proteins, or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification results.
[0115] The ROC curve of prediction model 3 on the training set shows good overall performance and a high AUC value of 0.85, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.01 corresponding to the 5-fold cross-validation ROC curve of prediction model 3 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 3 on the test set also shows good overall performance and a high AUC value of 0.81. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0116] Example 5 Establishment and validation of logistic regression prediction model 4
[0117] In Example 5, the effective protein biomarkers are A0A024R035 and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. A logistic regression model with two independent variables, A0A024R035 and P59666, was constructed on the training set. First, based on the quantitative data of A0A024R035 and P59666 and the sample labels, prediction model 4 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 4 is: P = 1 / (1 + Exp(-(β0 + β1 × A0A024R035 + β2 × P59666))); where P represents the probability of a sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids refer to the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The value range of coefficient β0 is [-39.46, -30.21]; the value range of β1 is [1.97, 2.79]; and the value range of coefficient β2 is [0.70, 1.15]. Preferably, β0 is -34.84, β1 is 2.38, and β2 is 0.93, yielding the optimal prediction results. The established logistic regression model can predict the likelihood of a new sample developing gastric cancer based on A0A024R035 and P59666, their antibodies, single peptide chains, characteristic peptide segments, stable isotope proteins, or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification results.
[0118] The ROC curve of prediction model 4 on the training set shows good overall performance and a high AUC value of 0.88, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.02 corresponding to the 5-fold cross-validation ROC curve of prediction model 4 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 4 on the test set also shows good overall performance and a high AUC value of 0.86. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0119] Example 6 Establishment and validation of logistic regression prediction model 5
[0120] In Example 6, the effective protein biomarkers are P59666, P04275, and P02533, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. On the training set, a logistic regression model containing three independent variables—P59666, P04275, and P02533—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of P59666, P04275, and P02533 and the sample labels, the prediction model 5 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 5 is: P = 1 / (1 + Exp(-(β0 + β1 × P59666 + β2 × P04275 + β3 × P02533))); where P represents the probability of a sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids refer to the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The coefficients β0, β1, β2, and β3 range from [-32.50, -24.96] to [0.73, 1.15], [1.08, 1.60], and [0.31, 0.64]. Preferably, β0 is -28.73, β1 is 0.94, β2 is 1.34, and β3 is 0.47, yielding the optimal prediction results. The established logistic regression model can predict the likelihood of gastric cancer in new samples based on the P59666, P04275, and P02533 proteins, their antibodies, single peptide chains, characteristic peptide segments of the peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification results.
[0121] The ROC curve of prediction model 5 on the training set shows good overall performance and a high AUC value of 0.87, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.02 corresponding to the 5-fold cross-validation ROC curve of prediction model 5 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 5 on the test set also shows good overall performance and a high AUC value of 0.85. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0122] Example 7 Establishment and validation of logistic regression prediction model 6
[0123] In Example 7, the effective protein biomarkers are A0A024R035, A2VCK8, and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. On the training set, a logistic regression model containing three independent variables—A0A024R035, A2VCK8, and P59666—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of A0A024R035, A2VCK8, and P59666, and the sample labels, the prediction model 6 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 6 is: P = 1 / (1 + Exp(-(β0 + β1 × A0A024R035 + β2 × A2VCK8 + β3 × P59666))); where P represents the probability of a sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids refer to the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The coefficient β0 ranges from [-41.44, -31.73]; β1 ranges from [1.94, 2.76]; the coefficient β2 ranges from [0.42, 0.86]; and the coefficient β3 ranges from [0.42, 0.86] to [0.42, 0.86]. [0.89]. Preferably, β0 is -36.59, β1 is 2.35, β2 is 0.64, and β3 is 0.65, which yields the optimal prediction results. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the A0A024R035, A2VCK8, and P59666 proteins, their antigens / antibodies, their single peptide chains, characteristic peptide segments of the peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification results.
[0124] The ROC curve of prediction model 6 on the training set shows good overall performance and a high AUC value of 0.90, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the ROC curve for 5-fold cross-validation on the training set for prediction model 6 are 0.89 and 0.02, respectively, achieving the set performance target and demonstrating the model's stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 6 on the test set also shows good overall performance and a high AUC value of 0.88. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0125] Example 8 Establishment and validation of logistic regression prediction model 7
[0126] In Example 8, the effective protein biomarkers are P59666, P04275, P02533, and A0A024R6P0, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. On the training set, a logistic regression model containing four independent variables—P59666, P04275, P02533, and A0A024R6P0—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of P59666, P04275, P02533, and A0A024R6P0, and the sample labels, the prediction model 7 was obtained using a logistic regression algorithm. The mathematical expression for prediction model 7 is: P = 1 / (1 + Exp(-(β0 + β1 × P59666 + β2 × P04275 + β3 × P02533 + β4 × A0A024R6P0))); where P represents the probability of a sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids, represent the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The coefficient β0 ranges from [-42.81, -32.14]; the coefficient β1 ranges from [0.60, 1.05]; and the coefficient β2 ranges from [0.63, [1.17]; the value range of coefficient β3 is [0.26, 0.60]; the value range of coefficient β4 is [0.99, 1.92]. Preferably, β0 is -37.47, β1 is 0.82, β2 is 0.90, β3 is 0.43, and β4 is 1.45 to obtain the optimal prediction result. The established logistic regression model can predict the possibility of gastric cancer in new samples based on the results of quantification of P59666, P04275, P02533 and A0A024R6P0 proteins or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0127] The ROC curve of prediction model 7 on the training set shows good overall performance and a high AUC value of 0.88, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.02 corresponding to the 5-fold cross-validation ROC curve of prediction model 7 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 7 on the test set also shows good overall performance and a high AUC value of 0.84. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0128] Example 9 Establishment and validation of logistic regression prediction model 8
[0129] In Example 9, the effective protein biomarkers are P59666, P04275, P02533, and P00915, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. On the training set, a logistic regression model containing four independent variables—P59666, P04275, P02533, and P00915—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of P59666, P04275, P02533, and P00915 and the sample labels, the prediction model 8 was obtained using a logistic regression algorithm. The mathematical expression for prediction model 8 is: P = 1 / (1 + Exp(-(β0 + β1 × P59666 + β2 × P04275 + β3 × P02533 + β4 × P00915))); where P represents the probability of a sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids refer to the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The coefficient β0 ranges from [-27.47, -19.42]; the coefficient β1 ranges from [0.76, 1.19]; and the coefficient β2 ranges from [1.15, ...]. [1.70]; the value range of coefficient β3 is [0.32, 0.67]; the value range of coefficient β4 is [-0.93, -0.51]. Preferably, β0 is -23.44, β1 is 0.98, β2 is 1.43, β3 is 0.49, and β4 is -0.72 to obtain the optimal prediction results. The established logistic regression model can predict the likelihood of gastric cancer in new samples based on the results of quantification of P59666, P04275, P02533 and P00915 proteins or their antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0130] The ROC curve of prediction model 8 on the training set shows good overall performance with a high AUC value of 0.89, indicating good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.02 corresponding to the 5-fold cross-validation ROC curve of prediction model 8 on the training set meet the set performance target, demonstrating model stability and effectiveness, and providing a guarantee for practical application. The ROC curve of prediction model 8 on the test set also shows good overall performance with a high AUC value of 0.86. This indicates that the model maintains good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0131] Example 10 Establishment and validation of logistic regression prediction model 9
[0132] In Example 10, the effective protein biomarkers are P59666, P04275, P02533, P00915, and P0DJI8, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. On the training set, a logistic regression model containing five independent variables—P59666, P04275, P02533, P00915, and P0DJI8—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of P59666, P04275, P02533, P00915, and P0DJI8, and the sample labels, a predictive model 9 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 9 is: P = 1 / (1 + Exp(-(β0 + β1 × P59666 + β2 × P04275 + β3 × P02533 + β4 × P00915 + β5 × P0DJI8))); where P represents the probability of a sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids, represent the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The coefficient β0 ranges from [-26.49, -18.14]; the coefficient β1 ranges from [0.67, 1.13]; and the coefficient β2 ranges from [0.72, [1.31]; the value range of coefficient β3 is [0.36, 0.73]; the value range of coefficient β4 is [-0.94, -0.49]; the value range of coefficient β5 is [0.19, 0.44]. Preferably, β0 is -22.31, β1 is 0.90, β2 is 1.02, β3 is 0.55, β4 is -0.72, and β5 is 0.31, which yields the optimal prediction results. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of P59666, P04275, P02533, P00915, and P0DJI8 proteins or their antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0133] The ROC curve of prediction model 9 on the training set shows good overall performance with a high AUC value of 0.90, indicating good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.02 corresponding to the 5-fold cross-validation ROC curve of prediction model 9 on the training set meet the set performance targets, demonstrating model stability and effectiveness, and providing assurance for practical application. The ROC curve of prediction model 9 on the test set also shows good overall performance with a high AUC value of 0.86. This indicates that the model maintains good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and suitable for clinical validation and application.
[0134] Example 11 Establishment and validation of logistic regression prediction model 10
[0135] In Example 11, the effective protein biomarkers are any one or more of the five proteins described below, or their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins, stable isotope characteristic peptide segments, and corresponding nucleic acids. On the training set, a logistic regression model containing five independent variables—P59666, P04275, P02533, A0A024R6P0, and H7BYG8—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of P59666, P04275, P02533, A0A024R6P0, and H7BYG8, and the sample labels, the prediction model 10 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 10 is: P = 1 / (1 + Exp(-(β0 + β1 × P59666 + β2 × P04275 + β3 × P02533 + β4 × A0A024R6P0 + β5 × H7BYG8))); where P represents the probability of a sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids refer to the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The coefficient β0 ranges from [-35.55, -23.97]; the coefficient β1 ranges from [0.35, 0.82]; and the coefficient β2 ranges from [0.59, [1.18]; the value range of coefficient β3 is [0.29, 0.66]; the value range of coefficient β4 is [1.08, 2.09]; the value range of coefficient β5 is [-0.92, -0.52]. Preferably, β0 is -29.76, β1 is 0.58, β2 is 0.88, β3 is 0.47, β4 is 1.58, and β5 is -0.72, which yields the optimal prediction results. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of P59666, P04275, P02533, A0A024R6P0 and H7BYG8 proteins or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0136] The ROC curve of prediction model 10 on the training set shows good overall performance with a high AUC value of 0.90, indicating good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.02 corresponding to the 5-fold cross-validation ROC curve of prediction model 10 on the training set meet the set performance targets, demonstrating model stability and effectiveness, and providing assurance for practical application. The ROC curve of prediction model 10 on the test set also shows good overall performance with a high AUC value of 0.87. This indicates that the model maintains good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and suitable for clinical validation and application.
[0137] Example 12 Establishment and validation of logistic regression prediction model 11
[0138] In Example 12, the effective protein biomarkers are any one or more of the 11 proteins described below, or their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. On the training set, a logistic regression model containing 11 independent variables—P59666, P02750, Q06033, P04275, H7BYG8, O14746, P02533, H6VRF8, A4D1J9, A2VCK8, and P02671—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of P59666, P02750, Q06033, P04275, H7BYG8, O14746, P02533, H6VRF8, A4D1J9, A2VCK8, and P02671, and the sample labels, prediction model 11 was obtained using a logistic regression algorithm. The mathematical expression of prediction model 11 is: P = 1 / (1 + Exp(-(β0 + β1 × P59666 + β2 × P02750 + β3 × Q06033 + β4 × P04275 + β5 × H7BYG8 + β6 × O14746 + β7 × P02533 + β8 × H6VRF8 + β9 × A4D1J9 + β10 × A2VCK8 + β11 × P02671); where P represents the probability of a sample being classified as a positive example. In the formula, the name of the protein biomarker or its antigen / antibody or its single peptide chain, characteristic peptide segment of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segment, and the corresponding nucleic acid refers to the standardized content value of the protein biomarker or its antigen / antibody or its single peptide chain, characteristic peptide segment of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segment, and the corresponding nucleic acid. The coefficients β0, β1, β2, β3, β4, and β5 range from [-41.99, -28.30] to [-0.01, 0.57], [0.53, 1.64], [0.14, 1.32], [0.17, 0.83], to [-0.96, -0.30]. The coefficient β6 ranges from [-0.45, 0.16]; the coefficient β7 ranges from [0.06, 0.71]; the coefficient β8 ranges from [-0.33, 0.28]; the coefficient β9 ranges from [-0.31, 0.78]; the coefficient β10 ranges from [0.37, 0.86]; and the coefficient β11 ranges from [0.47, 1.83].Preferably, the optimal prediction results can be obtained by setting β0 to -35.14, β1 to 0.28, β2 to 1.08, β3 to 0.73, β4 to 0.50, β5 to -0.63, β6 to -0.15, β7 to 0.38, β8 to -0.03, β9 to 0.23, β10 to 0.61, and β11 to 1.15. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of the following proteins in a new sample: P59666, P02750, Q06033, P04275, H7BYG8, O14746, P02533, H6VRF8, A4D1J9, A2VCK8, and P02671, or their antigens / antibodies, or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0139] The ROC curve of prediction model 11 on the training set shows good overall performance and a high AUC value of 0.94, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.01 corresponding to the 5-fold cross-validation ROC curve of prediction model 11 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 11 on the test set also shows good overall performance and a high AUC value of 0.91. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0140] Example 13 Establishment and validation of logistic regression prediction model 12
[0141] In Example 13, the effective protein biomarkers are any one or more of the 31 proteins described below, or their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins, stable isotope characteristic peptide segments, and corresponding nucleic acids. A training set containing A0A024R035, Q06033, E9KL23, P02750, P02671, P04275, A2VCK8, A4D1J9, P59666, H7BYG8, A0A8F0WQF6, G3V3A0, B4E215, A0A024RCW3, O14746, P06727, and P05109 was constructed. A logistic regression model with 31 independent variables, P02753, A0A384MEF1, A2VCQ3, E9PHK0, A0A087WZ31, B4E273, Q8NHV9, A0A024R0R6, D9YZU5, P27169, Q7Z4J2, A0A5S8K7B6, B2R4C5, and P80108, was used to predict the likelihood of a sample developing gastric cancer. First, based on A0A024R035, Q06033, E9KL23, P02750, P02671, P04275, A2VCK8, A4D1J9, P59666, H7BYG8, A0A8F0WQF6, G3V3A0, B4E215, A0A024RCW3, O14746, P06727, P0 Quantitative data and sample labels for samples 5109, P02753, A0A384MEF1, A2VCQ3, E9PHK0, A0A087WZ31, B4E273, Q8NHV9, A0A024R0R6, D9YZU5, P27169, Q7Z4J2, A0A5S8K7B6, B2R4C5, and P80108.Prediction model 12 was obtained using the logistic regression algorithm. The mathematical expression of prediction model 12 is: P=1 / (1+Exp(-(β0+β1 ×A0A024R035 + β2 ×Q06033 + β3 × E9KL23 + β4 ×P02750 + β5 × P02671 + β6 ×P04275 + β7 ×A2VCK8 + β8 ×A4D1J9 + β9 ×P59666+ β10 × H7BYG8+ β11 × A0A8F0WQF6 +β12 ×G3V3A0+ β13 × B4E215+ β14 ×A0A024RCW3+ β15 × O14746+ β16 × P06727+ β17 ×P05109+ β18 ×P02753+ β19 ×A0A384MEF1+ β20 ×A2VCQ3+ β21 × E9PHK0+ β22 × A0A087WZ31+β23 ×B4E273+ β24× Q8NHV9+ β25 × A0A024R0R6+ β26 × D9YZU5+ β27 × P27169+β28 ×Q7Z4J2+ β29× A0A5S8K7B6+ β30 × B2R4C5+ β31 × P80108)))); where P represents the probability of a sample being classified as a positive example. In the formula, the name of the protein biomarker or its antigen / antibody or its single peptide chain, characteristic peptide segment of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segment, and the corresponding nucleic acid refers to the standardized content value of the protein biomarker or its antigen / antibody or its single peptide chain, characteristic peptide segment of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segment, and the corresponding nucleic acid.The coefficients β0, β1, β2, β3, β4, β5, β6, β7, β8, β9, and β10 are all within the range of -48.56 to -20.22. The coefficients β0, β1, β2, β3, β4, β5, β6, β7, β8, β9, and β10 are all within the range of -1.38 to -1.25. -0.43]; the value range of coefficient β11 is [0.24, 2.12]; the value range of coefficient β12 is [0.25, 1.23]; the value range of β13 is [-0.09, 0.87]; the value range of coefficient β14 is [0.41, 1.39]; the value range of coefficient β15 is [-0.50, 0.39]; the value range of coefficient β16 is [-1.20, 0.42]; the value range of coefficient β17 is [-0.22, 0.29]; the value range of coefficient β18 is [-1.20, 0.30]; the value range of coefficient β19 is [-0.48, 1.14]; the value range of coefficient β20 is [-1.00, 0.08]; the value range of coefficient β21 is [-1.98, -0.56]; the value range of coefficient β22 is [0.38, 1.78]; the value range of coefficient β23 is [-0.62, 0.02]; the value range of coefficient β24 is [-1.02, 0.28]; the value range of coefficient β25 is [-0.50, 0.08]; the value range of coefficient β26 is [-1.56, -0.83]; the value range of coefficient β27 is [-1.42, 0.08]; the value range of coefficient β28 is [-0.20, 0.33]; the value range of coefficient β29 is [-1.26, -0.30]; the value range of coefficient β30 is [0.28, [1.06]; the value range of coefficient β31 is [-0.52, 0.97]. Preferably, β0 is -34.39, β1 is -0.69, β2 is 0.48, β3 is 1.99, β4 is 0.26, β5 is 0.79, β6 is 0.45, β7 is 0.75, β8 is 0.78, β9 is 0.42, β10 is -0.91, β11 is 1.18, β12 is 0.74, β13 is 0.39, β14 is 0.90, β15 is -0.06, β16 is -0.39, β17 is 0.04, β18 is -0.45, β19 is 0.33, β20 is -0.46, β21 is -1.27, and β22 is 1.08.The optimal prediction results can be obtained with β23 = -0.30, β24 = -0.37, β25 = -0.21, β26 = -1.19, β27 = -0.67, β28 = 0.07, β29 = -0.78, β30 = 0.67, and β31 = 0.23. The established logistic regression model can be based on the new sample's A0A024R035, Q06033, E9KL23, P02750, P02671, P04275, A2VCK8, A4D1J9, P59666, H7BYG8, A0A8F0WQF6, G3V3A0, B4E215, A0A024RCW3, O14746, P06727, P05109, P02753, A0A384MEF1, A The likelihood of a new sample developing gastric cancer is predicted based on the quantification results of proteins or antigens / antibodies of 2VCQ3, E9PHK0, A0A087WZ31, B4E273, Q8NHV9, A0A024R0R6, D9YZU5, P27169, Q7Z4J2, A0A5S8K7B6, B2R4C5, and P80108, or their single peptide chains, characteristic peptide segments, and stable isotope proteins or stable isotope characteristic peptide segments.
[0142] The ROC curve of prediction model 12 on the training set shows good overall performance and a high AUC value of 0.97, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the ROC curve for 5-fold cross-validation on the training set for prediction model 12 are 0.96 and 0.01, respectively, achieving the set performance target and demonstrating the model's stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 12 on the test set also shows good overall performance and a high AUC value of 0.94. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0143] Example 14 Establishment and validation of logistic regression prediction model 13
[0144] In Example 14, the effective protein biomarkers are any one or more of the 36 proteins described below, or their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins, stable isotope characteristic peptide segments, and corresponding nucleic acids. A training set containing B7WNR0, P00738, D9YZU5, H6VRF8, G3V3A0, A0A8F0WQF6, H7C517, A4D1J9, P05109, A5PL27, A0A1W2PQX5, O14746, A0A024RCW3, P02533, Q9NXP7, Q4VXF1, P06727, P0DJI8, P18428, and A0A02 was constructed. A logistic regression model using 36 independent variables—4R6P0, A2VCK8, P59666, A0A1Z1G4M2, H7BYG8, P04275, Q06033, E9KL23, A0A024R035, P02671, P02750, A0A140KFU0, A0A0K2BMD8, V9H1D9, E9PHK0, Q4TZM4, and A2VCQ3—was used to predict the likelihood of a sample developing gastric cancer. Quantitative data of these 36 proteins and sample labels were used to derive the predictive model 13 through the logistic regression algorithm. The mathematical expression of prediction model 13 is: P=1 / (1+Exp(-(β0+β1 ×B7WNR0+ β2 ×P00738 + β3 ×D9YZU5+ β4 × H6VRF8+ β5 × G3V3A0+ β6 ×A0A8F0WQF6+ β7 ×H7C517+ β8 ×A4D1J9+ β9 ×P05109+ β10 × A5PL27+ β11 × A0A1W2PQX5 + β12 ×O14746+ β13 ×A0A024RCW3 + β14 ×P02533+ β15 × Q9NXP7+ β16 × Q4VXF1+ β17 ×P06727+ β18 ×P0DJI8+ β19 ×P18428+ β20 ×A0A024R6P0+ β21 × A2VCK8+ β22 × P59666 +β23 ×A0A1Z1G4M2+ β24 × H7BYG8+ β25 × P04275+ β26 × Q06033 + β27 × E9KL23+ β28×A0A024R035 + β29 ×P02671+ β30 × P02750 + β31 × A0A140KFU0 + β32 ×A0A0K2BMD8+ β33 ×V9H1D9+ β34 ×E9PHK0 + β35 × Q4TZM4 + β36 × A2VCQ3))));Where P represents the probability of a sample being classified as a positive example, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids in the formula refer to the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The coefficients β0, β1, β2, β3, β4, β5, and β6 range from -49.40 to -22.93. 1.86]; the value range of coefficient β7 is [-1.43, 0.31]; the value range of coefficient β8 is [-0.75, 0.80]; the value range of coefficient β9 is [-0.11, 0.40]; the value range of coefficient β10 is [-0.72, 1.59]; the value range of coefficient β11 is [-0.36, 0.86]; the value range of coefficient β12 is [-0.54, 0.29]; the value range of coefficient β13 is [0.25, 1.26]; the value range of coefficient β14 is [0.04, 0.82]; the value range of coefficient β15 is [-0.21, 1.51]; the value range of coefficient β16 is [-0.14, 0.59]; the value range of coefficient β17 is [-1.40, -0.40]; the value range of coefficient β18 is [-0.26, 0.27]; the value range of coefficient β19 is [-1.20, 0.28]; the value range of coefficient β20 is [-1.79, 0.24]; the value range of coefficient β21 is [0.36, 0.94]; the value range of coefficient β22 is [-0.09, 0.73]; the value range of coefficient β23 is [-0.02, 1.40]; the value range of coefficient β24 is [-1.35, -0.41]; the value range of coefficient β25 is [0.11, 0.92]; the value range of coefficient β26 is [-0.39, 1.16]; The value range of coefficient β27 is [1.26, 4.25]; The value range of coefficient β28 is [-1.54, 0.32]; The value range of coefficient β29 is [0.15, 1.89]; The value range of coefficient β30 is [-0.59, 1.18]; The value range of coefficient β31 is [-0.54, 0.13]; The value range of coefficient β32 is [-1.70, 1.22]; The value range of coefficient β33 is [-0.81, 0.24];The coefficient β34 ranges from [-1.56, -0.20]; the coefficient β35 ranges from [-0.50, 0.85]; and the coefficient β36 ranges from [-1.18, -0.11]. Preferably, β0 is -36.16, β1 is 0.10, β2 is -0.09, β3 is -0.61, β4 is -0.01, β5 is 0.64, β6 is 0.94, β7 is -0.56, β8 is 0.03, β9 is 0.14, β10 is 0.43, β11 is 0.25, β12 is -0.12, β13 is 0.76, β14 is 0.43, β15 is 0.65, β16 is 0.23, β17 is -0.90, β18 is 0.01, and β19 is... The optimal prediction results are obtained with β20 = -0.46, β20 = -0.77, β21 = 0.65, β22 = 0.32, β23 = 0.69, β24 = -0.88, β25 = 0.51, β26 = 0.38, β27 = 2.76, β28 = -0.61, β29 = 1.02, β30 = 0.30, β31 = -0.20, β32 = -0.24, β33 = -0.28, β34 = -0.88, β35 = 0.17, and β36 = -0.65. The established logistic regression model can be based on the following new samples: B7WNR0, P00738, D9YZU5, H6VRF8, G3V3A0, A0A8F0WQF6, H7C517, A4D1J9, P05109, A5PL27, A0A1W2PQX5, O14746, A0A024RCW3, P02533, Q9NXP7, Q4VXF1, P06727, P0DJI8, P18428, A0A024R6P0, A2VCK8, P59666. The likelihood of a new sample having gastric cancer is predicted based on the following proteins: A0A1Z1G4M2, H7BYG8, P04275, Q06033, E9KL23, A0A024R035, P02671, P02750, A0A140KFU0, A0A0K2BMD8, V9H1D9, E9PHK0, Q4TZM4, and A2VCQ3, or their antigens / antibodies, or their single peptide chains, characteristic peptide segments, and stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification results.
[0145] The ROC curve of prediction model 13 on the training set shows good overall performance and a high AUC value of 0.96, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.01 corresponding to the 5-fold cross-validation ROC curve of prediction model 13 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 13 on the test set also shows good overall performance and a high AUC value of 0.93. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0146] Example 15 Establishment and validation of logistic regression prediction model 14
[0147] In Example 15, the effective protein biomarkers are any one or more of the 41 proteins described below, or their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins, stable isotope characteristic peptide segments, and corresponding nucleic acids. A training set containing A0A0G2JMS6, Q4TZM4, P59666, P02750, B4E367, Q06033, P04275, E9KL23, A0A7S5BZY2, Q6IQ49, A0A5C2G1Q3, A6XNE2, A0A7S5EYK2, A0A5C2FXP5, V9H1D9, E9PHK0, Q8NHV9, H0YAC1, A0A5C2FVW2, A0A5C2G4U1, Q2L9S7, and P271 was constructed. A logistic regression model with 41 independent variables (A0A7S5EWA8, A0A126LAY7, P43652, P35908, H7BYG8, P02652, Q5CZ93, A0A5C2GJH5, A0A5C2G130, O14746, P02533, A0A0A0MRS8, H6VRF8, A0A5S8K7B6, A4D1J9, A2VCK8, Q86TT1, A0A5C2GCL9, P02671) was used to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of the 41 proteins and the sample labels, the predictive model 14 was obtained using the logistic regression algorithm. The mathematical expression of prediction model 14 is: P=1 / (1+Exp(-(β0+β1 ×A0A0G2JMS6+ β2 ×Q4TZM4+ β3 ×P59666+ β4 × P02750+ β5 × B4E367+ β6 ×Q06033+ β7 ×P04275+ β8 ×E9KL23+ β9×A0A7S5BZY2+ β10 × Q6IQ49+ β11 × A0A5C2G1Q3+ β12 ×A6XNE2+ β13 ×A0A7S5EYK2+ β14 ×A0A5C2FXP5+ β15 × V9H1D9 + β16 × E9PHK0+ β17 ×Q8NHV9+ β18 × H0YAC1+ β19 ×A0A5C2FVW2+ β20 ×A0A5C2G4U1+ β21 × Q2L9S7+ β22 ×P27169+β23 ×A0A7S5EWA8+ β24 × A0A126LAY7+ β25 × P43652+ β26 × P35908+ β27 ×H7BYG8+ β28 ×P02652 + β29 ×Q5CZ93+ β30 × A0A5C2GJH5+ β31 × A0A5C2G130+ β32 ×O14746+ β33 ×P02533+ β34 ×A0A0A0MRS8+ β35 × H6VRF8+ β36 ×A0A5S8K7B6+ β37 × A4D1J9+ β38 ×A2VCK8+ β39 ×Q86TT1+ β40 × A0A5C2GCL9+ β41× P02671))));where P represents the probability of a sample being classified as a positive example. In the formula, the name of the protein biomarker or its antigen / antibody or its single peptide chain, characteristic peptide segment of the peptide chain, and its stable isotope protein or stable isotope characteristic peptide segment, and the corresponding nucleic acid refers to the standardized content value of the protein biomarker or its antigen / antibody or its single peptide chain, characteristic peptide segment of the peptide chain, and its stable isotope protein or stable isotope characteristic peptide segment, and the corresponding nucleic acid. The coefficient β0 ranges from [-22.02, 31.86]; the value range of β1 is [-0.78, -0.19]; the value range of coefficient β2 is [-1.36, -0.02]; the value range of coefficient β3 is [0.25, 1.12]; the value range of coefficient β4 is [-0.03, 1.66]; the value range of coefficient β5 is [-1.14, -0.03]; the value range of coefficient β6 is [-0.87, 0.80]; the value range of coefficient β7 is [-0.26, 0.69]; the value range of coefficient β8 is [1.55, 5.00]; the value range of coefficient β9 is [0.17, 0.92]; the value range of coefficient β10 is [-0.70, 0.19]; the value range of coefficient β11 is [-2.95, -0.81]; the value range of coefficient β12 is [0.52, 1.99]; the value range of β13 is [-0.40, 0.63]; the value range of coefficient β14 is [-0.72, 0.27]; the value range of coefficient β15 is [-1.19, 0.03]; the value range of coefficient β16 is [-1.38, 0.27]; the value range of coefficient β17 is [-0.99, -0.06]; the value range of coefficient β18 is [-2.15, 0.11]; the value range of coefficient β19 is [-0.53, 0.03]; the value range of coefficient β20 is [-0.87, -0.08]; the value range of coefficient β21 is [-0.05, [0.61]; the value range of coefficient β22 is [-0.42, 1.66]; the value range of coefficient β23 is [-0.79, -0.05]; the value range of coefficient β24 is [0.10, 0.64]; the value range of coefficient β25 is [-1.74, 0.73]; the value range of coefficient β26 is [0.06, 0.76]; the value range of coefficient β27 is [-1.27, -0.10]; the value range of coefficient β28 is [-0.82,1.82]; the value range of coefficient β29 is [-2.27, -0.31]; the value range of coefficient β30 is [-0.49, 0.27]; the value range of coefficient β31 is [-0.81, 0.07]; the value range of coefficient β32 is [-0.23, 0.84]; the value range of coefficient β33 is [-0.12, 0.90]; the value range of coefficient β34 is [-0.56, 0.66]; the value range of coefficient β35 is [-0.50, 0.37]; the value range of coefficient β36 is [-1.78, -0.49]; the value range of coefficient β37 is [-0.59, 1.23]; the value range of coefficient β38 is [0.20, [0.93]; the value range of coefficient β39 is [-0.52, -0.02]; the value range of coefficient β40 is [0.12, 0.77]; the value range of coefficient β41 is [-0.03, 1.98]. Preferably, β0 is 4.92, β1 is -0.48, β2 is -0.69, β3 is 0.68, β4 is 0.82, β5 is -0.59, β6 is -0.03, β7 is 0.21, β8 is 3.27, β9 is 0.55, β10 is -0.26, β11 is -1.88, β12 is 1.25, β13 is 0.11, β14 is -0.23, β15 is -0.58, β16 is -0.56, β17 is -0.52, β18 is -1.02, β19 is -0.25, β20 is -0.47, and β21 is 0. The optimal prediction result can be obtained with the following values: β22 = 0.62, β23 = -0.42, β24 = 0.37, β25 = -0.51, β26 = 0.41, β27 = -0.68, β28 = 0.50, β29 = -1.29, β30 = -0.11, β31 = -0.37, β32 = 0.30, β33 = 0.39, β34 = 0.05, β35 = -0.06, β36 = -1.13, β37 = 0.32, β38 = 0.56, β39 = -0.27, β40 = 0.44, and β41 = 0.98. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of these 41 proteins or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0148] The ROC curve of prediction model 14 on the training set shows good overall performance and a high AUC value of 0.98, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.01 corresponding to the 5-fold cross-validation ROC curve of prediction model 14 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 14 on the test set also shows good overall performance and a high AUC value of 0.96. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0149] Example 16 Establishment and validation of logistic regression prediction model 15
[0150] In Example 16, the effective protein biomarkers are any one or more of the 74 proteins described below, or their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins, stable isotope characteristic peptide segments, and corresponding nucleic acids. On the training set, a dataset containing A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A126LAY7, A0A1B0GVI3, A0A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, and A... 0A5C2FVW2, A0A5C2FXP5, A0A5C2G130, A0A5C2G1Q3, A0A5C2G4U1, A0A5C2G7F2, A0A5C2GCL9, A0A5C2GJ H5, A0A5S8K7B6, A0A7S5BZY2, A0A7S5EWA8, A0A7S5EYK2, A0A8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A 4D1J9, A6XNE2, A8K3I0, B2R4C5, B4E273, B4E367, B4E3M1, B7WNR0, D9YZU5, E9KL23, E9PHK0, G3V3A0, H 0YAC1, H6VRF8, H7BYG8, L8E853, O14746, P00738, P02533, P02652, P02671, P02750, P02753, P02763, P A logistic regression model with 74 independent variables (P04275, P05109, P06727, P0DJI8, P18428, P27169, P35527, P35900, P35908, P43652, P59666, Q06033, Q2L9S7, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q8NHV9, and V9H1D9) was used to predict the probability of a sample developing gastric cancer. First, based on the quantitative data of the 74 proteins and the sample labels, the predictive model 15 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 15 is: P = 1 / (1 + Exp(-(β0 + β1 ×A0A024QZH6 + β2 ×A0A024R035 + β3 ×A0A024R1G6 + β4 ×A0A024R6P0 + β5 ×A0A024RCW3 + β6 ×A0A0A0MRS8 + β7 ×A0A0G2JMS6 + β8 ×A0A0K2BMD8 + β9 ×A0A0S2Z3X8 + β10 ×A0A126LAY7 + β11)×A0A1B0GVI3 + β12 ×A0A1Z1G4M2 + β13 ×A0A2U8J8L9 + β14 ×A0A384MEF1 + β15 ×A0A5C2FVW2 + β16 ×A0A5C2FXP5 + β17 ×A0A5C2G130 + β18 ×A0A5C2G1Q3 + β19 ×A0A5C2G4U1 + β20 ×A0A5C2G7F2 + β21 ×A0A5C2GCL9 + β22 ×A0A5C2GJH5 + β23 ×A0A5S8K7B6 + β24 ×A0A7S5BZY2 + β25 ×A0A7S5EWA8 + β26 ×A0A7S5EYK2 + β27 ×A0A8F0WQF6 + β28 ×A0A8G1A656 + β29 ×A2VCK8 + β30 ×A2VCQ3 + β31 ×A4D1J9 + β32 ×A6XNE2 + β33 ×A8K3I0 + β34 ×B2R4C5 + β35 ×B4E273 + β36 ×B4E367 + β37 ×B4E3M1 + β38 ×B7WNR0 + β39 ×D9YZU5 + β40 ×E9KL23 + β41 ×E9PHK0 + β42 ×G3V3A0 + β43 ×H0YAC1 + β44 ×H6VRF8 + β45 ×H7BYG8 + β46 ×L8E853 + β47 ×O14746 + β48 ×P00738 + β49 ×P02533 + β50 ×P02652 + β51 ×P02671 + β52 ×P02750 + β53 ×P02753 + β54 ×P02763 + β55 ×P04275 + β56 ×P05109 + β57 ×P06727 + β58 ×P0DJI8 + β59 ×P18428 + β60 ×P27169 + β61 ×P35527 + β62 ×P35900 + β63 ×P35908 + β64 ×P43652 + β65 ×P59666 + β66 ×Q06033 + β67 ×Q2L9S7 + β68 ×Q4TZM4 + β69 ×Q4VXF1 + β70 ×Q5CZ93 + β71 ×Q6IQ49 + β72 ×Q86TT1 + β73 ×Q8NHV9 + b74×V9H1D9))));where P represents the probability of the sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids refer to the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The range and preferred values of the coefficients β0~β74 are shown in Table 2. The optimal prediction results can be obtained based on the preferred values. The established logistic regression model can predict the likelihood of gastric cancer in new samples based on the quantitative results of these 74 proteins or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids.
[0151] Table 2 shows the range and preferred values of the coefficients for prediction model 15.
[0152] The ROC curve of prediction model 15 on the training set shows good overall performance and a high AUC value of 0.99, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the ROC curve for 5-fold cross-validation on the training set for prediction model 15 are 0.98 and 0.01, respectively, achieving the set performance target and demonstrating the model's stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 15 on the test set also shows good overall performance and a high AUC value of 0.96. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0153] Example 17 Establishment and validation of logistic regression prediction model 16
[0154] In Example 17, the effective protein biomarkers are any one or more of the 74 proteins described below, or their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins, stable isotope characteristic peptide segments, and corresponding nucleic acids. A training set was constructed containing Q06033, E9KL23, P02750, P02671, P04275, A2VCK8, A4D1J9, P59666, Q2L9S7, P02533, A0A126LAY7, H0YAC1, A0A0A0MRS8, A0A7S5EYK2, P43652, Q4TZM4, V9H1D9, A0A5S8K7B6, P27169, Q8NHV9, and E9. PHK0, O14746, H7BYG8, A6XNE2, A0A5C2G130, H6VRF8, A0A5C2GCL9, P35908, A0A7S5BZY2, A0A5C2G4U1, A 0A5C2FVW2, A0A5C2FXP5, A0A5C2GJH5, B4E367, A0A0G2JMS6, A0A7S5EWA8, Q86TT1, Q6IQ49, Q5CZ93, P026 52. A0A5C2G1Q3, A0A024R035, A0A024R6P0, A0A1Z1G4M2, P02763, H7C517, P18428, B2RMS9, Q9H387, A0A 1W2PQX5, P0DJI8, B2R888, A0A024QZH6, L8E853, A0A024RCW3, B7WNR0, A0A024R1G6, A0A0X9V9B3, A0A140 A logistic regression model with 74 independent variables (KFU0, D9YZU5, P00915, B2M1S7, A0A0K2BMD8, A0A5C2FYJ4, A0A5C2FZZ3, A0A5C2GBQ5, A0A5C2GMN5, A0A5C2GN07, A0A7I2V2D2, A0A7S5BYW5, P35900, O00617, P06727, and P35527) was used to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of the 74 proteins and the sample labels, a predictive model was obtained using the logistic regression algorithm. The mathematical expression for prediction model 16 is: P = 1 / (1 + Exp(-(β0 + β1 × Q06033 + β2 × E9KL23 + β3 × P02750 + β4 × P02671 + β5 × P04275 + β6 × A2VCK8 + β7 × A4D1J9 + β8 × P59666 + β9 × Q2L9S7 + β10 × P02533 + β11 × A0A126LAY7 + β12 × H0YAC1 + )β13 ×A0A0A0MRS8+ β14 × A0A7S5EYK2+ β15 × P43652+ β16 × Q4TZM4+ β17 × V9H1D9+ β18 × A0A5S8K7B6+ β19 × P27169+ β20 × Q8NHV9+ β21 × E9PHK0+ β22 × O14746+β23 × H7BYG8+ β24 × A6XNE2+ β25 × A0A5C2G130+ β26 × H6VRF8+ β27 ×A0A5C2GCL9+ β28 × P35908+ β29 × A0A7S5BZY2+ β30 × A0A5C2G4U1+ β31 ×A0A5C2FVW2+ β32 × A0A5C2FXP5+ β33 × A0A5C2GJH5+ β34 × B4E367+ β35 ×A0A0G2JMS6+ β36 × A0A7S5EWA8+ β37 × Q86TT1+ β38 × Q6IQ49+ β39 × Q5CZ93+ β40 × P02652+ β41 × A0A5C2G1Q3+ β42 × A0A024R035+ β43 × A0A024R6P0+ β44 ×A0A1Z1G4M2+ β45 × P02763+ β46 × H7C517+ β47 × P18428+ β48 × B2RMS9+ β49× Q9H387+ β50 × A0A1W2PQX5+ β51 × P0DJI8+ β52 × B2R888+ β53 × A0A024QZH6+ β54 × L8E853+ β55 × A0A024RCW3+ β56 × B7WNR0+ β57 × A0A024R1G6+ β58 ×A0A0X9V9B3+ β59 × A0A140KFU0+ β60 × D9YZU5+ β61 × P00915+ β62 × B2M1S7+ β63 × A0A0K2BMD8+ β64 × A0A5C2FYJ4+ β65 × A0A5C2FZZ3+ β66 × A0A5C2GBQ5+ β67 × A0A5C2GMN5+ β68 × A0A5C2GN07+ β69 × A0A7I2V2D2+ β70 × A0A7S5BYW5+ β71 × P35900+ β72 × O00617+ β73 × P06727+ β74 ×P35527)))); where P represents the probability of the sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids refer to the standardized content values of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids. The range and preferred values of the coefficients β0~β74 are shown in Table 3. The optimal prediction results can be obtained based on the preferred values. The established logistic regression model can predict the likelihood of gastric cancer in new samples based on the quantitative results of these 74 proteins or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acids.
[0155] Table 3 shows the range and preferred values of the coefficients for prediction model 16.
[0156] The ROC curve of prediction model 16 on the training set shows good overall performance with a high AUC value of 0.99, indicating good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the 5-fold cross-validation ROC curve of prediction model 16 on the training set are 0.99 and 0.00, respectively, achieving the set performance target and demonstrating model stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 16 on the test set also shows good overall performance with a high AUC value of 0.94. This indicates that the model maintains good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0157] Example 18 Establishment and validation of logistic regression prediction model 17
[0158] In Example 18, the effective protein biomarkers are any one or more of the 119 proteins described below, or their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. On the training set, a dataset containing A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A0X9V9B3, A0A126LAY7, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, A0A385HVZ2, and A0A5C2 was constructed. FTW7, A0A5C2FVW2, A0A5C2FX67, A0A5C2FXP5, A0A5C2FYJ4, A0A5C2FYK5, A0A5C2FZZ3, A0A5C2G130, A0A5C2G1Q3, A0A5C2G1X1, A0A5C2G3M4 , A0A5C2G4U1, A0A5C2G6G6, A0A5C2G7F2, A0A5C2GA97, A0A5C2GBQ5, A0A5C2GCL9, A0A5C2GDA6, A0A5C2GGY3, A0A5C2GJH5, A0A5C2GJL5, A0A 5C2GMN5, A0A5C2GN07, A0A5C2GRZ5, A0A5C2GV43, A0A5C2GW23, A0A5C2H3N5, A0A5S8K7B6, A0A7I2V2D2, A0A7P0TAB0, A0A7S5BYW5, A0A7S5B ZY2, A0A7S5C366, A0A7S5EWA8, A0A7S5EYK2, A0A7T0LP36, A0A8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A6XNE2, A8K3I0, B2 M1S7, B2R4C5, B2R888, B2RMS9, B4DZH3, B4E1Z4, B4E273, B4E367, B4E3M1, B7WNR0, D9YZU5, E9KL23, E9PHK0, G3V3A0, H0YAC1, H6VRF8, H7BY G8, H7C517, H9KVD5, L8E853, O00617, O14746, P00738, P00915, P01031, P02533, P02652, P02671, P02750, P02753, P02763, P04275, P05109,A random forest model with 119 independent variables (P06727, P0C0L5, P0DJI8, P18428, P27169, P35527, P35900, P35908, P43652, P59666, Q06033, Q2L9S7, Q49A33, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q86YQ4, Q8NHV9, Q9H387, Q9NXP7, and V9H1D9) was used to predict the probability of a sample having gastric cancer. First, based on quantitative data of the 119 proteins or their antigens / antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, the corresponding nucleic acid, and the sample labels, the prediction model 17 was obtained using the random forest algorithm with the R package randomForest. The parameters of the `randomForest` function are set to `ntree = 500`, `mtry = 3`, `importance = TRUE`, and `proximity = TRUE`. The importance values of the 119 markers in the model are shown in Table 4. Table 4 has been sorted from highest to lowest by %IncMSE; the higher the ranking of the marker, the more important it is in the model and the more likely it is to be prioritized in subsequent model optimization.
[0159] The ROC curve of prediction model 17 on the training set shows good overall performance and a high AUC value of 1.00, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The ROC curve of prediction model 17 on the test set also shows good overall performance and a high AUC value of 0.94. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0160] Table 4. Biomarkers and corresponding importance index values for prediction model 17
[0161] The above models involve different indicators and can be applied to multiple different classification groups obtained based on the indicators and their quantities. Preferably, prediction models 1, 2, 3, and 4 can obtain accurate prediction results using a minimum of two indicators, making them simple and low-cost.
[0162] Figure 5 This is a system block diagram illustrating an embodiment of the present application for assessing the likelihood of a subject having gastric cancer based on the Uniprot database. (Reference) Figure 5As shown, the system 500 in this embodiment includes a data acquisition module 510, a data preprocessing module 520, a model recommendation module 530, a model selection module 540, and a risk assessment module 550.
[0163] The data acquisition module 510 is used to acquire sample data of the subjects, including A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A0X9V9B3, A0A126LAY7, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, A0A385HVZ2, and A0A5C2F. TW7, A0A5C2FVW2, A0A5C2FX67, A0A5C2FXP5, A0A5C2FYJ4, A0A5C2FYK5, A0A 5C2FZZ3, A0A5C2G130, A0A5C2G1Q3, A0A5C2G1X1, A0A5C2G3M4, A0A5C2G4U1 , A0A5C2G6G6, A0A5C2G7F2, A0A5C2GA97, A0A5C2GBQ5, A0A5C2GCL9, A0A5C2 GDA6, A0A5C2GGY3, A0A5C2GJH5, A0A5C2GJL5, A0A5C2GMN5, A0A5C2GN07, A0A 5C2GRZ5, A0A5C2GV43, A0A5C2GW23, A0A5C2H3N5, A0A5S8K7B6, A0A7I2V2D2 , A0A7P0TAB0, A0A7S5BYW5, A0A7S5BZY2, A0A7S5C366, A0A7S5EWA8, A0A7S5 EYK2, A0A7T0LP36, A0A8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5P L27, A6XNE2, A8K3I0, B2M1S7, B2R4C5, B2R888, B2RMS9, B4DZH3, B4E1Z4, B4E 273. B4E367, B4E3M1, B7WNR0, D9YZU5, E9KL23, E9PHK0, G3V3A0, H0YAC1, H6 VRF8, H7BYG8, H7C517, H9KVD5, L8E853, O00617, O14746, P00738, P00915, P 01031, P02533, P02652, P02671, P02750, P02753, P02763, P04275, P05109, P06727, P0C0L5, P0DJI8, P18428, P27169, P35527, P35900, P35908, P43652,Any one of the following proteins in serum: P59666, Q06033, Q2L9S7, Q49A33, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q86YQ4, Q8NHV9, Q9H387, Q9NXP7, and V9H1D9, or their antigens / antibodies, or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding quantitative values of nucleic acids.
[0164] The data preprocessing module 520 is used to preprocess the sample data, including removing duplicate samples, filling missing values, and dividing the sample data into different analysis groups.
[0165] The model recommendation module 530 is used to recommend multiple predictive models with larger AUC values from the predictive models built using the construction method described above, based on the types and number of indicators included in different analysis groups. The model formula with the largest AUC value is given priority, and the final result is displayed on the user interface as a list of model formulas, sorted from largest to smallest AUC value.
[0166] The model selection module 540 provides users with a selection function to output one or more predictive models from the predictive models recommended by the model recommendation module 530 for risk assessment and calculation. Users can choose not to make a selection, and the system will default to selecting the model with the highest AUC value.
[0167] The risk assessment module 550 calculates the corresponding predicted value based on one or more prediction models output by the model selection module 540, and provides an appropriate Logit threshold or risk score threshold according to the department or population information corresponding to the sample data. It then classifies the risk level as any one of high probability, medium probability, or low probability based on the Logit threshold or risk score threshold. In some embodiments, the Logit threshold or risk score threshold includes one threshold for binary classification, i.e., dividing into high probability and low probability. In some embodiments, the Logit threshold or risk score threshold includes two thresholds for trigonometric classification, i.e., dividing into high probability, medium probability, and low probability. In other embodiments, further subdivisions can be made, which are not limited in this application.
[0168] System 500 can quantitatively obtain biomarker information from subject samples and input it into a model for gastric cancer risk assessment, avoiding errors in subjective judgment and achieving effective assessment and risk prediction of the likelihood of developing gastric cancer in different populations. Unlike existing gastric cancer screening systems, System 500 can not only quantitatively, automatically, and continuously provide gastric cancer risk levels, improving the accuracy and efficiency of assessment; it also provides a rich set of model formulas containing different types of indicators. Considering that some samples may not have the most complete set of indicators in practical applications, System 500 can also provide the most suitable prediction model in such cases. In addition, System 500 also provides a model recommendation module 530 and a model selection module 540, making the system more flexible and meeting diverse application needs of customers. Furthermore, System 500 can apply the most appropriate threshold for risk assessment based on the department or population from which the sample originates.
[0169] This application also proposes a kit for assessing the likelihood of a subject developing gastric cancer based on the Uniprot database, including a predictive model established by the model construction method described above. Specifically, this predictive model can be integrated into an integrated circuit within the kit. The kit can receive any one or more of the effective protein biomarkers described above and calculate the prediction result based on the integrated predictive model, providing a probability result of the subject developing gastric cancer. Taking predictive model 1 as an example, the kit can receive, for example, P02750 and P59666, and calculate the prediction result based on the built-in predictive model 1 to provide a probability result of the subject developing gastric cancer. The kit can simultaneously include the 17 different predictive models described above to suit different populations, exhibiting very broad applicability.
[0170] This application also proposes a kit for assessing the likelihood of a subject having gastric cancer based on the Uniprot database, used to receive effective protein biomarkers, which are any one or more of the following protein biomarkers: A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A0X9V9B3, A0A126LAY7, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0A1Z1 G4M2, A0A2U8J8L9, A0A384MEF1, A0A385HVZ2, A0A5C2FTW7, A0A5C2FVW2, A0 A5C2FX67, A0A5C2FXP5, A0A5C2FYJ4, A0A5C2FYK5, A0A5C2FZZ3, A0A5C2G130 , A0A5C2G1Q3, A0A5C2G1X1, A0A5C2G3M4, A0A5C2G4U1, A0A5C2G6G6, A0A5C2 G7F2, A0A5C2GA97, A0A5C2GBQ5, A0A5C2GCL9, A0A5C2GDA6, A0A5C2GGY3, A0A 5C2GJH5, A0A5C2GJL5, A0A5C2GMN5, A0A5C2GN07, A0A5C2GRZ5, A0A5C2GV43 , A0A5C2GW23, A0A5C2H3N5, A0A5S8K7B6, A0A7I2V2D2, A0A7P0TAB0, A0A7S5B YW5, A0A7S5BZY2, A0A7S5C366, A0A7S5EWA8, A0A7S5EYK2, A0A7T0LP36, A0A 8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A6XNE2, A8K3I0, B2 M1S7, B2R4C5, B2R888, B2RMS9, B4DZH3, B4E1Z4, B4E273, B4E367, B4E3M1, B 7WNR0, D9YZU5, E9KL23, E9PHK0, G3V3A0, H0YAC1, H6VRF8, H7BYG8, H7C517, H 9KVD5, L8E853, O00617, O14746, P00738, P00915, P01031, P02533, P02652, P02671, P02750, P02753, P02763, P04275, P05109, P06727, P0C0L5, P0DJI8,P18428, P27169, P35527, P35900, P35908, P43652, P59666, Q06033, Q2L9S7, Q49A33, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q86YQ4, Q8NHV9, Q9H387, Q9NXP7 and V9H1D9, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, or corresponding nucleic acids.
[0171] In some embodiments, the effective protein biomarkers that the kit is used to receive are effective protein biomarkers included in any of the prediction models described above.
[0172] This application also includes an electronic device comprising a memory and a processor. The memory stores instructions executable by the processor; the processor executes the instructions to implement the aforementioned method for constructing a model based on the Uniprot database to assess the likelihood of a subject having gastric cancer.
[0173] Figure 6 This is a system block diagram of an electronic device according to an embodiment of this application. (Reference) Figure 6 As shown, the electronic device 600 may include an internal communication bus 601, a processor 602, a read-only memory (ROM) 603, a random access memory (RAM) 604, and a communication port 605. When applied to a personal computer, the electronic device 600 may also include a hard disk 606. The internal communication bus 601 enables data communication between components of the electronic device 600. The processor 602 can make judgments and issue prompts. In some embodiments, the processor 602 may consist of one or more processors. The communication port 605 enables data communication between the electronic device 600 and external devices. In some embodiments, the electronic device 600 can send and receive information and data from a network through the communication port 605. The electronic device 600 may also include different forms of program storage units and data storage units, such as the hard disk 606, the read-only memory (ROM) 603, and the random access memory (RAM) 604, capable of storing various data files used for computer processing and / or communication, as well as possible program instructions executed by the processor 602. The processor executes these instructions to implement the main part of the method. The results processed by the processor are transmitted to the user device through the communication port and displayed on the user interface.
[0174] The above-described model construction method can be implemented as a computer program, stored in the hard disk 606, and loaded into the processor 602 for execution to implement the model construction method of this application.
[0175] This application also includes a computer-readable medium storing computer program code that, when executed by a processor, implements the method for constructing the model described above.
[0176] When a method for constructing a model based on the Uniprot database to assess the likelihood of a subject developing gastric cancer is implemented as a computer program, it can also be stored as an article of manufacture in a computer-readable storage medium. For example, a computer-readable storage medium may include, but is not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic stripes), optical discs (e.g., compact discs (CDs), digital multifunction discs (DVDs)), smart cards, and flash memory devices (e.g., electrically erasable programmable read-only memory (EPROM), cards, sticks, key drives). Furthermore, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media (and / or storage media) capable of storing, containing, and / or carrying code and / or instructions and / or data.
[0177] This application uses specific terms to describe embodiments of the application. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of the application. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics in one or more embodiments of the application can be appropriately combined.
[0178] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of scope in some embodiments of this application are approximate values, in specific embodiments, such values are set as precisely as feasible.
Claims
1. A method for constructing a model based on the Uniprot database to assess the probability of a subject developing gastric cancer, comprising: Obtain samples from the subjects, including multiple serum samples; Mass spectrometry data analysis includes: performing mass spectrometry data analysis on the multiple serum samples to extract peptide information and obtain sample data; Database retrieval includes: performing a Uniprot database search on the sample data to identify proteins and obtain a raw protein quantification matrix; Anomaly handling includes: identifying and removing abnormal samples from the original protein quantification matrix to obtain an effective protein quantification matrix; Analyze differentially expressed proteins, including: comparing the expression differences of proteins among different groups to obtain a set of first protein biomarkers with significant differences; Data standardization processing includes: performing data standardization processing on the effective protein quantification matrix to eliminate the influence of different dimensions and scales, and obtaining a standardized protein quantification matrix; The screening of protein biomarkers includes: screening the standardized protein quantification matrix using at least m of the following methods to obtain at least m candidate protein biomarker sets: LASSO regression analysis, partial least squares discriminant analysis, orthogonal partial least squares discriminant analysis, RFS-based random forest algorithm, Mann-Whitney U test, Welch's t-test, Odds ratio, and P-values, where m is a positive integer and 3 ≤ m ≤ 8; The biomarker aggregation analysis includes: performing a aggregation analysis on the first set of protein biomarkers and the set of m candidate protein biomarkers, and obtaining a set of target protein biomarkers that simultaneously exist in at least four sets in the first set of protein biomarkers and the set of m candidate protein biomarkers from the first set of protein biomarkers and the set of m candidate protein biomarkers. Correlation analysis includes: analyzing the correlation between the target protein biomarker set and all protein biomarkers in the effective protein quantification matrix, and forming a new set P1 with the first biomarker among all protein biomarkers whose correlation coefficient is within the first range and the target protein biomarker set; The variable structuring process includes: dividing the samples in set P1 into gastric cancer group and non-gastric cancer group according to clinical diagnosis, and encoding them respectively to form the dependent variable of the prediction model. The independent variables of the prediction model include the quantitative results of protein, peptide or antibody and the patient's basic information. Model training and testing include: dividing the samples in set P1 into a training set and a test set, training the prediction model using the training set, and testing the prediction model using the test set, wherein the prediction model is a regression model including the independent variable and the dependent variable; Model optimization includes: selecting effective protein biomarkers based on at least one of the indicators of p-value, correlation coefficient, and AUC value, using the effective protein biomarkers as effective independent variables, and obtaining the optimized model.
2. The construction method as described in claim 1, characterized in that, Between the database retrieval step and the exception handling step, there is also a further step; Examining batch effects in the data includes: dividing the sample into multiple batches and using principal component analysis to analyze the batch effects between different batches; To remove batch effects, if batch effects exist, use the R package statTarget and the QC-RFSC method to remove batch effects between samples caused by different batches based on the QC samples. statTarget is the default parameter setting.
3. The construction method as described in claim 1, characterized in that, Between the database retrieval step and the exception handling step, there is also a further step; Protein filtering includes selecting proteins with quantitative values in a preset number of samples for subsequent protein biomarker screening.
4. The construction method as described in claim 1, characterized in that, Between the database retrieval step and the exception handling step, there is also a further step; Missing values were filled, including both non-randomly missing protein quantification values and randomly missing protein quantification values, which were filled using the K-nearest neighbor method.
5. The construction method as described in claim 1, characterized in that, The steps for analyzing differentially expressed proteins include: The t-test method was used to compare protein expression differences among different groups, obtaining the set G1 with significant differences; and The Deseq2 method was used to compare the differences in protein expression among different groups to obtain a set G2 with significant differences. The first set of protein biomarkers includes set G1 and set G2.
6. The construction method as described in claim 1, characterized in that, The data standardization process includes: performing data standardization on the effective protein quantification matrix using one or more of eight standardization methods, including: Log2, Log2+Median, Log2+Mean, VSN, Log2+RLR, Log2+GI, Log2+Quantile, and Log2+CycLoess.
7. The construction method as described in claim 1, characterized in that, Also includes: Model iteration includes repeating the model training and testing steps and the model optimization steps to build multiple prediction models, and evaluating the performance of different prediction models based on the training and testing results of the multiple prediction models.
8. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are any one or more of the following protein biomarkers: A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A0X9V9B3, A0A126LAY7, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, A0A385HVZ2, A0A5C2FTW7. A0A5C2FVW2, A0A5C2FX67, A0A5C2FXP5, A0A5C2FYJ4, A0A5C2FYK5, A0A5C2F ZZ3, A0A5C2G130, A0A5C2G1Q3, A0A5C2G1X1, A0A5C2G3M4, A0A5C2G4U1, A0A5 C2G6G6, A0A5C2G7F2, A0A5C2GA97, A0A5C2GBQ5, A0A5C2GCL9, A0A5C2GDA6, A0A5C2GGY3, A0A5C2GJH5, A0A5C2GJL5, A0A5C2GMN5, A0A5C2GN07, A0A5C2GR Z5, A0A5C2GV43, A0A5C2GW23, A0A5C2H3N5, A0A5S8K7B6, A0A7I2V2D2, A0A7 P0TAB0, A0A7S5BYW5, A0A7S5BZY2, A0A7S5C366, A0A7S5EWA8, A0A7S5EYK2, A 0A7T0LP36, A0A8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A6 XNE2, A8K3I0, B2M1S7, B2R4C5, B2R888, B2RMS9, B4DZH3, B4E1Z4, B4E273, B4 E367, B4E3M1, B7WNR0, D9YZU5, E9KL23, E9PHK0, G3V3A0, H0YAC1, H6VRF8, H 7BYG8, H7C517, H9KVD5, L8E853, O00617, O14746, P00738, P00915, P01031, P 02533, P02652, P02671, P02750, P02753, P02763, P04275, P05109, P06727, P0C0L5, P0DJI8, P18428, P27169, P35527, P35900, P35908, P43652, P59666,Q06033, Q2L9S7, Q49A33, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q86YQ4, Q8NHV9, Q9H387, Q9NXP7, and V9H1D9, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, or characteristic peptide segments of stable isotopes.
9. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P02750 and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 1, and the mathematical expression of prediction model 1 is: P=1 / (1+Exp(-(β0+β1 × P02750 + β2 × P59666))); where the name of the effective protein biomarker refers to the standardized content value of any one or more of the protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-35.47, -26.94]; the coefficient β1 ranges from [1.79, 2.50]; and the coefficient β2 ranges from [0.60, 1.05].
10. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are B7WNR0 and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 2, and the mathematical expression of prediction model 2 is: P=1 / (1+Exp(-(β0+β1 ×B7WNR0 + β2 × P59666))); where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-13.12, -7.61]; the coefficient β1 ranges from [-0.57, -0.28]; and the coefficient β2 ranges from [1.00, 1.40].
11. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are A0A024R6P0 and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 3, and the mathematical expression of prediction model 3 is: P=1 / (1+Exp(-(β0+β1 ×A0A024R6P0 + β2 × P59666))); where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-36.64, -27.29]; the coefficient β1 ranges from [1.55, 2.37]; and the coefficient β2 ranges from [0.79, 1.22].
12. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are A0A024R035 and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 4, and the mathematical expression of prediction model 4 is: P=1 / (1+Exp(-(β0+β1 ×A0A024R035 + β2 × P59666))); where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-39.46, -30.21]; the coefficient β1 ranges from [1.97, 2.79]; and the coefficient β2 ranges from [0.70, 1.15].
13. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P59666, P04275, and P02533, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 5, and the mathematical expression of prediction model 5 is: P=1 / (1+Exp(-(β0+β1 ×P59666+ β2 × P04275 + β3 × P02533))); wherein, the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, and corresponding nucleic acid; the coefficient β0 ranges from [-32.50, -24.96]; the coefficient β1 ranges from [0.73, 1.15]; and the coefficient β2 ranges from [1.08, 1.60]; The value range of coefficient β3 is [0.31, 0.64].
14. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are A0A024R035, A2VCK8, and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 6, and the mathematical expression of prediction model 6 is: P=1 / (1+Exp(-(β0+β1 ×A0A024R035 + β2 × A2VCK8 + β3 ×P59666))); wherein, the name of the effective protein biomarker refers to the standardized content value of any one or more of the protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, and corresponding nucleic acid, and the coefficient β0 ranges from [-41.44, -31.73]; the coefficient β1 ranges from [1.94, 2.76]; and the coefficient β2 ranges from [0.42, [0.86]; the value range of coefficient β3 is [0.42, 0.89].
15. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P59666, P04275, P02533, and A0A024R6P0, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 7, and the mathematical expression of prediction model 7 is: P=1 / (1+Exp(-(β0+β1 ×P59666 + β2 × P04275 + β3 ×P02533 + β4 × A0A024R6P0))); wherein, the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, and corresponding nucleic acid; the coefficient β0 ranges from [-42.81, -32.14]; the coefficient β1 ranges from [0.60, [1.05]; the value range of coefficient β2 is [0.63, 1.17]; the value range of coefficient β3 is [0.26, 0.60]; the value range of coefficient β4 is [0.99, 1.92].
16. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P59666, P04275, P02533, and P00915, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 8, and the mathematical expression of prediction model 8 is: P=1 / (1+Exp(-(β0+β1 ×P59666 + β2 × P04275 + β3 × P02533+ β4 × P00915))); wherein, the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, and corresponding nucleic acid; the coefficient β0 ranges from [-27.47, -19.42]; the coefficient β1 ranges from [0.76, 1.19]; the value range of coefficient β2 is [1.15, 1.70]; the value range of coefficient β3 is [0.32, 0.67]; the value range of coefficient β4 is [-0.93, -0.51].
17. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P59666, P04275, P02533, P00915, and P0DJI8, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 9, and the mathematical expression of prediction model 9 is: P = 1 / (1 + Exp(-(β0 + β1 × P59666 + β2 × P04275 + β3 × P02533 + β4 × P00915 + β5 × P0DJI8))); wherein, the name of the effective protein biomarker refers to the standardized content value of any one or more of the protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, and corresponding nucleic acid, and the coefficient β0 ranges from [-26.49, -18.14]; the value range of β1 is [0.67, 1.13]; the value range of coefficient β2 is [0.72, 1.31]; the value range of coefficient β3 is [0.36, 0.73]; the value range of coefficient β4 is [-0.94, -0.49]; the value range of coefficient β5 is [0.19, 0.44].
18. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P59666, P04275, P02533, A0A024R6P0, and H7BYG8, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 10, and the mathematical expression of prediction model 10 is: P = 1 / (1 + Exp(-(β0 + β1 × P59666 + β2 × P04275 + β3 × P02533 + β4 × A0A024R6P0 + β5 × H7BYG8))); wherein, the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, and corresponding nucleic acid, and the coefficient β0 ranges from [-35.55, -23.97]; the value range of β1 is [0.35, 0.82]; the value range of coefficient β2 is [0.59, 1.18]; the value range of coefficient β3 is [0.29, 0.66]; the value range of coefficient β4 is [1.08, 2.09]; the value range of coefficient β5 is [-0.92, -0.52].
19. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P59666, P02750, Q06033, P04275, H7BYG8, O14746, P02533, H6VRF8, A4D1J9, A2VCK8, and P02671, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids. The optimized model is prediction model 11, and the mathematical expression of prediction model 11 is: P = 1 / (1 + Exp(-(β0 + β1 × P59666 + β2 × P02750 + β3 × Q06033 + β4 × P04275 + β5 × H7BYG8 + β6 × O14746 + β7 × P02533 + β8 × H6VRF8 + β9) ×A4D1J9+ β10 × A2VCK8+ β11 × P02671))); where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-41.99, -28.30]; β1 ranges from [-0.01, 0.57]; β2 ranges from [0.53, 1.64]; β3 ranges from [0.14, 1.32]; β4 ranges from [0.17, 0.83]; β5 ranges from [-0.96, -0.30]; and β6 ranges from [-0.45, [0.16]; the value range of coefficient β7 is [0.06, 0.71]; the value range of coefficient β8 is [-0.33, 0.28]; the value range of coefficient β9 is [-0.31, 0.78]; the value range of coefficient β10 is [0.37, 0.86]; the value range of coefficient β11 is [0.47, 1.83].
20. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are A0A024R035, Q06033, E9KL23, P02750, P02671, P04275, A2VCK8, A4D1J9, P59666, H7BYG8, A0A8F0WQF6, G3V3A0, B4E215, A0A024RCW3, O14746, P06727, P05109, P02753, A0A384MEF1, A2VCQ3, E9PHK0, A0A0 87WZ31, B4E273, Q8NHV9, A0A024R0R6, D9YZU5, P27169, Q7Z4J2, A0A5S8K7B6, B2R4C5, and P80108, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, or corresponding nucleic acids, wherein the optimization model is prediction model 12, and the mathematical expression of prediction model 12 is: P=1 / (1+Exp(-(β0+β1)) ×A0A024R035 + β2 ×Q06033 + β3 × E9KL23 +β4 ×P02750 + β5 × P02671 + β6 ×P04275 + β7 ×A2VCK8 + β8 ×A4D1J9 + β9 ×P59666+ β10 × H7BYG8+ β11 × A0A8F0WQF6 +β12 ×G3V3A0+ β13 × B4E215+ β14 ×A0A024RCW3+ β15 × O14746+ β16 × P06727+ β17 ×P05109+ β18 ×P02753+ β19 ×A0A384MEF1+ β20 ×A2VCQ3+ β21 × E9PHK0+ β22 × A0A087WZ31+β23 ×B4E273+ β24× Q8NHV9+ β25 × A0A024R0R6+ β26 × D9YZU5+ β27 × P27169+β28 ×Q7Z4J2+ β29× A0A5S8K7B6+ β30 × B2R4C5+ β31 × P80108)))); Wherein, the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid; the coefficient β0 ranges from [-48.56, -20.22]; the coefficient β1 ranges from [-1.59, [0.21]; the value range of coefficient β2 is [-0.29, 1.25]; the value range of coefficient β3 is [0.58, 3.40]; the value range of coefficient β4 is [-0.54, 1].07]; The value range of coefficient β5 is [-0.11, 1.70]; The value range of coefficient β6 is [0.06, 0.84]; The value range of coefficient β7 is [0.45, 1.05]; The value range of coefficient β8 is [-0.03, 1.58]; The value range of coefficient β9 is [0.00, 0.84]; The value range of coefficient β10 is [-1.38, -0.43]; The value range of coefficient β11 is [0.24, 2.12]; The value range of coefficient β12 is [0.25, 1.23]; The value range of coefficient β13 is [-0.09, 0.87]; The value range of coefficient β14 is [0.41, 1.39]; The value range of coefficient β15 is [-0.50, 0.39]; the value range of coefficient β16 is [-1.20, 0.42]; the value range of coefficient β17 is [-0.22, 0.29]; the value range of coefficient β18 is [-1.20, 0.30]; the value range of coefficient β19 is [-0.48, 1.14]; the value range of coefficient β20 is [-1.00, 0.08]; the value range of coefficient β21 is [-1.98, -0.56]; the value range of coefficient β22 is [0.38, 1.78]; the value range of coefficient β23 is [-0.62, 0.02]; the value range of coefficient β24 is [-1.02, 0.28]; the value range of coefficient β25 is [-0.50, [0.08]; the value range of coefficient β26 is [-1.56, -0.83]; the value range of coefficient β27 is [-1.42, 0.08]; the value range of coefficient β28 is [-0.20, 0.33]; the value range of coefficient β29 is [-1.26, -0.30]; the value range of coefficient β30 is [0.28, 1.06]; the value range of coefficient β31 is [-0.52, 0.97].
21. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are B7WNR0, P00738, D9YZU5, H6VRF8, G3V3A0, A0A8F0WQF6, H7C517, A4D1J9, P05109, A5PL27, A0A1W2PQX5, O14746, A0A024RCW3, P02533, Q9NXP7, Q4VXF1, P06727, P0DJI8, P18428, A0A024R6P0, A2VCK8, P59666, A0A1Z1G4M2, and H7BYG8. P04275, Q06033, E9KL23, A0A024R035, P02671, P02750, A0A140KFU0, A0A0K2BMD8, V9H1D9, E9PHK0, Q4TZM4 and A2VCQ3, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids, wherein the optimization model is prediction model 13, and the mathematical expression of prediction model 13 is: P=1 / (1+Exp(-(β0+β1)) ×B7WNR0+ β2 ×P00738 + β3 × D9YZU5+ β4 × H6VRF8+ β5 × G3V3A0+ β6×A0A8F0WQF6+ β7 ×H7C517+ β8 ×A4D1J9+ β9 ×P05109+ β10 × A5PL27+ β11 ×A0A1W2PQX5 + β12 ×O14746+ β13 × A0A024RCW3 + β14 ×P02533+ β15 × Q9NXP7+ β16 × Q4VXF1+ β17 ×P06727+ β18 × P0DJI8+ β19 ×P18428+ β20 ×A0A024R6P0+ β21 × A2VCK8+ β22 × P59666 +β23 ×A0A1Z1G4M2+ β24 × H7BYG8+ β25 × P04275+β26 × Q06033 + β27 × E9KL23+ β28 ×A0A024R035 + β29 ×P02671+ β30 × P02750+ β31 × A0A140KFU0 + β32 × A0A0K2BMD8+ β33 ×V9H1D9+ β34 ×E9PHK0 + β35 ×Q4TZM4 + β36 × A2VCQ3))));Wherein, the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-49.40, -22.93]; β1 ranges from [-0.26, 0.46]; β2 ranges from [-0.41, 0.23]; β3 ranges from [-2.26, 1.05]; β4 ranges from [-0.38, 0.35]; β5 ranges from [0.19, 1.09]; β6 ranges from [0.02, 1.86]; β7 ranges from [-1.43, 0.31]; and β8 ranges from [-0.75, ...]. [0.80]; the value range of coefficient β9 is [-0.11, 0.40]; the value range of coefficient β10 is [-0.72, 1.59]; the value range of coefficient β11 is [-0.36, 0.86]; the value range of coefficient β12 is [-0.54, 0.29]; the value range of coefficient β13 is [0.25, 1.26]; the value range of coefficient β14 is [0.04, 0.82]; the value range of coefficient β15 is [-0.21, 1.51]; the value range of coefficient β16 is [-0.14, 0.59]; the value range of coefficient β17 is [-1.40, -0.40]; the value range of coefficient β18 is [-0.26, [0.27]; the value range of coefficient β19 is [-1.20, 0.28]; the value range of coefficient β20 is [-1.79, 0.24]; the value range of coefficient β21 is [0.36, 0.94]; the value range of coefficient β22 is [-0.09, 0.73]; the value range of coefficient β23 is [-0.02, 1.40]; the value range of coefficient β24 is [-1.35, -0.41]; the value range of coefficient β25 is [0.11, 0.92]; the value range of coefficient β26 is [-0.39, 1.16]; the value range of coefficient β27 is [1.26, 4.25]; the value range of coefficient β28 is [-1.54, [0.32]; the value range of coefficient β29 is [0.15, 1.89]; the value range of coefficient β30 is [-0.59, 1.18]; the value range of coefficient β31 is [-0.54, 0.13]; the value range of coefficient β32 is [-1.70, 1.22]; the value range of coefficient β33 is [-0.81, 0.24]; the value range of coefficient β34 is [-1.56, -0.20]; the value range of coefficient β35 is [-0.50, 0.85]; the value range of coefficient β36 is [-1.18, -0.11].
22. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are A0A0G2JMS6, Q4TZM4, P59666, P02750, B4E367, Q06033, P04275, E9KL23, A0A7S5BZY2, Q6IQ49, A0A5C2G1Q3, A6XNE2, A0A7S5EYK2, A0A5C2FXP5, V9H1D9, E9PHK0, Q8NHV9, H0YAC1, A0A5C2FVW2, A0A5C2G4U1, Q2L9S7, P27169, A0A7S5EWA8, A0A126LAY7, P43652, and P3590.
8. H7BYG8, P02652, Q5CZ93, A0A5C2GJH5, A0A5C2G130, O14746, P02533, A0A0A0MRS8, H6VRF8, A0A5S8K7B6, A4D1J9, A2VCK8, Q86TT1, A0A5C2GCL9, and P02671, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, or corresponding nucleic acids, wherein the optimization model is prediction model 14, and the mathematical expression of prediction model 14 is: P=1 / (1+Exp(-(β0+β1)) ×A0A0G2JMS6+ β2 ×Q4TZM4+ β3 × P59666+ β4 × P02750+ β5 × B4E367+ β6 ×Q06033+ β7 ×P04275+ β8 ×E9KL23+ β9 ×A0A7S5BZY2+ β10 × Q6IQ49+ β11 ×A0A5C2G1Q3+ β12 ×A6XNE2+ β13 × A0A7S5EYK2+ β14 ×A0A5C2FXP5+ β15 × V9H1D9+ β16 × E9PHK0+ β17 ×Q8NHV9+ β18 × H0YAC1+ β19 ×A0A5C2FVW2+ β20 ×A0A5C2G4U1+ β21 × Q2L9S7+ β22 ×P27169 +β23 ×A0A7S5EWA8+ β24 × A0A126LAY7+β25 × P43652+ β26 × P35908+ β27 × H7BYG8+ β28 ×P02652 + β29 ×Q5CZ93+ β30× A0A5C2GJH5+ β31 × A0A5C2G130+ β32 × O14746+ β33 ×P02533+ β34 ×A0A0A0MRS8+ β35 × H6VRF8+ β36 × A0A5S8K7B6+ β37 × A4D1J9+ β38×A2VCK8+ β39 ×Q86TT1+ β40 × A0A5C2GCL9+ β41 × P02671)))); where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-22.02, 31.86]; the coefficient β1 ranges from [-0.78, -0.19]; the coefficient β2 ranges from [-1.36, -0.02]; the coefficient β3 ranges from [0.25, 1.12]; the coefficient β4 ranges from [-0.03, 1.66]; and the coefficient β5 ranges from [-1.14, -0.03]; the value range of coefficient β6 is [-0.87, 0.80]; the value range of coefficient β7 is [-0.26, 0.69]; the value range of coefficient β8 is [1.55, 5.00]; the value range of coefficient β9 is [0.17, 0.92]; the value range of coefficient β10 is [-0.70, 0.19]; the value range of coefficient β11 is [-2.95, -0.81]; the value range of coefficient β12 is [0.52, 1.99]; the value range of coefficient β13 is [-0.40, 0.63]; the value range of coefficient β14 is [-0.72, 0.27]; the value range of coefficient β15 is [-1.19, 0.03]; the value range of coefficient β16 is [-1.38, 0.27]; the value range of coefficient β17 is [-0.99, -0.06]; the value range of coefficient β18 is [-2.15, 0.11]; the value range of coefficient β19 is [-0.53, 0.03]; the value range of coefficient β20 is [-0.87, -0.08]; the value range of coefficient β21 is [-0.05, 0.61]; the value range of coefficient β22 is [-0.42, 1.66]; the value range of coefficient β23 is [-0.79, -0.05]; the value range of coefficient β24 is [0.10, 0.64]; the value range of coefficient β25 is [-1.74, 0.73]; the value range of coefficient β26 is [0.06, 0.76]; the value range of coefficient β27 is [-1.27, -0.10]; the value range of coefficient β28 is [-0.82, 1.82]; the value range of coefficient β29 is [-2.27, -0.31]; the value range of coefficient β30 is [-0.49, 0.27]; the value range of coefficient β31 is [-0.81, 0.07]; the value range of coefficient β32 is [-0.23, 0.84]; the value range of coefficient β33 is [-0.12, 0.90]; the value range of coefficient β34 is [-0.56,[0.66]; the value range of coefficient β35 is [-0.50, 0.37]; the value range of coefficient β36 is [-1.78, -0.49]; the value range of coefficient β37 is [-0.59, 1.23]; the value range of coefficient β38 is [0.20, 0.93]; the value range of coefficient β39 is [-0.52, -0.02]; the value range of coefficient β40 is [0.12, 0.77]; the value range of coefficient β41 is [-0.03, 1.98].
23. A system for assessing the likelihood of a subject developing gastric cancer based on the Uniprot database, characterized in that, include: The data acquisition module is used to acquire sample data from the subjects. The sample data includes A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A0X9V9B3, A0A126LAY7, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, A0A385HVZ2, A0A5C2FTW7, A 0A5C2FVW2, A0A5C2FX67, A0A5C2FXP5, A0A5C2FYJ4, A0A5C2FYK5, A0A5C2FZ Z3, A0A5C2G130, A0A5C2G1Q3, A0A5C2G1X1, A0A5C2G3M4, A0A5C2G4U1, A0A5C 2G6G6, A0A5C2G7F2, A0A5C2GA97, A0A5C2GBQ5, A0A5C2GCL9, A0A5C2GDA6, A 0A5C2GGY3, A0A5C2GJH5, A0A5C2GJL5, A0A5C2GMN5, A0A5C2GN07, A0A5C2GRZ 5. A0A5C2GV43, A0A5C2GW23, A0A5C2H3N5, A0A5S8K7B6, A0A7I2V2D2, A0A7P 0TAB0, A0A7S5BYW5, A0A7S5BZY2, A0A7S5C366, A0A7S5EWA8, A0A7S5EYK2, A 0A7T0LP36, A0A8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A6 XNE2, A8K3I0, B2M1S7, B2R4C5, B2R888, B2RMS9, B4DZH3, B4E1Z4, B4E273, B4 E367, B4E3M1, B7WNR0, D9YZU5, E9KL23, E9PHK0, G3V3A0, H0YAC1, H6VRF8, H 7BYG8, H7C517, H9KVD5, L8E853, O00617, O14746, P00738, P00915, P01031, P 02533, P02652, P02671, P02750, P02753, P02763, P04275, P05109, P06727, P0C0L5, P0DJI8, P18428, P27169, P35527, P35900, P35908, P43652, P59666,The following proteins, or their antigens / antibodies, single peptide chains, characteristic peptide segments, stable isotope proteins, stable isotope characteristic peptide segments, and corresponding nucleic acids, may be identified as Q06033, Q2L9S7, Q49A33, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q86YQ4, Q8NHV9, Q9H387, Q9NXP7, and V9H1D9 proteins; The data preprocessing module is used to preprocess the sample data, including removing duplicate samples, unifying units, and dividing the sample data into different analysis groups according to the type and quantity of indicators. The model recommendation module is used to recommend multiple prediction models with larger AUC values from multiple prediction models built using the construction method described in claim 1, based on the types and number of indicators contained in different analysis groups. The model selection module is used to provide users with a selection function to output one or more prediction models from the prediction models recommended by the model recommendation module for risk assessment and calculation. The risk assessment module is used to calculate the corresponding predicted value based on one or more prediction models output by the model selection module, and to provide an appropriate Logit threshold or risk score threshold according to the department or population information corresponding to the sample data, and to classify the risk level as any one of high probability, medium probability, and low probability according to the Logit threshold or risk score threshold.
24. A kit for assessing the likelihood of a subject having gastric cancer based on the Uniprot database, characterized in that, The method is used to receive effective protein biomarkers, which are any one or more of the following protein biomarkers: A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A0X9V9B3, A0A126LAY7, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, A0A385HVZ 2. A0A5C2FTW7, A0A5C2FVW2, A0A5C2FX67, A0A5C2FXP5, A0A5C2FYJ4, A0A5C 2FYK5, A0A5C2FZZ3, A0A5C2G130, A0A5C2G1Q3, A0A5C2G1X1, A0A5C2G3M4, A0 A5C2G4U1, A0A5C2G6G6, A0A5C2G7F2, A0A5C2GA97, A0A5C2GBQ5, A0A5C2GCL 9. A0A5C2GDA6, A0A5C2GGY3, A0A5C2GJH5, A0A5C2GJL5, A0A5C2GMN5, A0A5C2 GN07, A0A5C2GRZ5, A0A5C2GV43, A0A5C2GW23, A0A5C2H3N5, A0A5S8K7B6, A0 A7I2V2D2, A0A7P0TAB0, A0A7S5BYW5, A0A7S5BZY2, A0A7S5C366, A0A7S5EWA 8. A0A7S5EYK2, A0A7T0LP36, A0A8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4 D1J9, A5PL27, A6XNE2, A8K3I0, B2M1S7, B2R4C5, B2R888, B2RMS9, B4DZH3, B4 E1Z4, B4E273, B4E367, B4E3M1, B7WNR0, D9YZU5, E9KL23, E9PHK0, G3V3A0, H 0YAC1, H6VRF8, H7BYG8, H7C517, H9KVD5, L8E853, O00617, O14746, P00738, P 00915, P01031, P02533, P02652, P02671, P02750, P02753, P02763, P04275, P05109, P06727, P0C0L5, P0DJI8, P18428, P27169, P35527, P35900, P35908,P43652, P59666, Q06033, Q2L9S7, Q49A33, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q86YQ4, Q8NHV9, Q9H387, Q9NXP7 and V9H1D9, or their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids.
25. A kit for assessing the likelihood of a subject having gastric cancer based on the Uniprot database, characterized in that, For receiving effective protein biomarkers, said effective protein biomarkers include any one or more of P59666 or its antigen / antibody, a single peptide chain, a characteristic peptide segment of the peptide chain, a stable isotopic protein, a stable isotopic characteristic peptide segment, and a corresponding nucleic acid.
26. The kit as described in claim 25, characterized in that, The effective protein biomarkers also include any one or more of P02750, B7WNR0, A0A024R6P0, A0A024R035, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids.
27. The kit as claimed in claim 25, characterized in that, The effective protein biomarkers are P59666, P04275, and P02533, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids.
28. The kit as claimed in claim 25, characterized in that, The effective protein biomarkers are A0A024R035, A2VCK8, and P59666, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids.
29. The kit as described in claim 25, characterized in that, The effective protein biomarkers are P59666, P04275, P02533 and A0A024R6P0, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids.
30. The kit according to claim 25, characterized in that, The effective protein biomarkers are P59666, P04275, P02533 and P00915, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids.
31. The kit according to claim 25, characterized in that, The effective protein biomarkers are P59666, P04275, P02533, P00915, and P0DJI8, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids.
32. The kit as described in claim 25, characterized in that, The effective protein biomarkers are P59666, P04275, P02533, A0A024R6P0 and H7BYG8, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids.
33. The kit as described in claim 25, characterized in that, The effective protein biomarkers are P59666, P02750, Q06033, P04275, H7BYG8, O14746, P02533, H6VRF8, A4D1J9, A2VCK8 and P02671, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids.
34. The kit as described in claim 25, characterized in that, The effective protein biomarkers are A0A024R035, Q06033, E9KL23, P02750, P02671, P04275, A2VCK8, A4D1J9, P59666, H7BYG8, A0A8F0WQF6, G3V3A0, B4E215, A0A024RCW3, O14746, P06727, P05109, P02753, and A0A38. 4MEF1, A2VCQ3, E9PHK0, A0A087WZ31, B4E273, Q8NHV9, A0A024R0R6, D9YZU5, P27169, Q7Z4J2, A0A5S8K7B6, B2R4C5 and P80108, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, or corresponding nucleic acids.
35. The kit according to claim 25, characterized in that, The effective protein biomarkers are B7WNR0, P00738, D9YZU5, H6VRF8, G3V3A0, A0A8F0WQF6, H7C517, A4D1J9, P05109, A5PL27, A0A1W2PQX5, O14746, A0A024RCW3, P02533, Q9NXP7, Q4VXF1, P06727, P0DJI8, P18428, A0A024R6P0, and A2VCK8. P59666, A0A1Z1G4M2, H7BYG8, P04275, Q06033, E9KL23, A0A024R035, P02671, P02750, A0A140KFU0, A0A0K2BMD8, V9H1D9, E9PHK0, Q4TZM4 and A2VCQ3, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, or corresponding nucleic acids.
36. The kit as described in claim 25, characterized in that, The effective protein biomarkers are A0A0G2JMS6, Q4TZM4, P59666, P02750, B4E367, Q06033, P04275, E9KL23, A0A7S5BZY2, Q6IQ49, A0A5C2G1Q3, A6XNE2, A0A7S5EYK2, A0A5C2FXP5, V9H1D9, E9PHK0, Q8NHV9, H0YAC1, A0A5C2FVW2, A0A5C2G4U1, Q2L9S7, P27169, and A0A7S5EWA8. A0A126LAY7, P43652, P35908, H7BYG8, P02652, Q5CZ93, A0A5C2GJH5, A0A5C2G130, O14746, P02533, A0A0A0MRS8, H6VRF8, A0A5S8K7B6, A4D1J9, A2VCK8, Q86TT1, A0A5C2GCL9, and P02671, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, or corresponding nucleic acids.
37. The kit according to claim 25, characterized in that, The effective protein biomarkers are A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A126LAY7, A0A1B0GVI3, A0A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, and A0A5C2FV. W2, A0A5C2FXP5, A0A5C2G130, A0A5C2G1Q3, A0A5C2G4U1, A0A5C2G7F2, A0A5C2GCL9, A0A5C2GJH5, A0A5S8K 7B6, A0A7S5BZY2, A0A7S5EWA8, A0A7S5EYK2, A0A8F0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A6XNE2, A8K3I0, B2R4C5, B4E273, B4E367, B4E3M1, B7WNR0, D9YZU5, E9KL23, E9PHK0, G3V3A0, H0YAC1, H6VRF8, H7 BYG8, L8E853, O14746, P00738, P02533, P02652, P02671, P02750, P02753, P02763, P04275, P05109, P0672 7. P0DJI8, P18428, P27169, P35527, P35900, P35908, P43652, P59666, Q06033, Q2L9S7, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q8NHV9 and V9H1D9, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, or corresponding nucleic acids.
38. The kit according to claim 25, characterized in that, The effective protein biomarkers are Q06033, E9KL23, P02750, P02671, P04275, A2VCK8, A4D1J9, P59666, Q2L9S7, P02533, A0A126LAY7, H0YAC1, A0A0A0MRS8, A0A7S5EYK2, P43652, Q4TZM4, V9H1D9, A0A5S8K7B6, P27169, Q8NHV9, E9PHK0, O1 4746, H7BYG8, A6XNE2, A0A5C2G130, H6VRF8, A0A5C2GCL9, P35908, A0A7S5BZY2, A0A5C2G4U1, A0A5C2FVW2, A0A5C2FXP5, A0A5C2GJH5, B4E367, A0A0G2JMS6, A0A7S5EWA8, Q86TT1, Q6IQ49, Q5CZ93, P02652, A0A5C2G1Q3 , A0A024R035, A0A024R6P0, A0A1Z1G4M2, P02763, H7C517, P18428, B2RMS9, Q9H387, A0A1W2PQX5, P0DJI8, B 2R888, A0A024QZH6, L8E853, A0A024RCW3, B7WNR0, A0A024R1G6, A0A0X9V9B3, A0A140KFU0, D9YZU5, P00915 B2M1S7, A0A0K2BMD8, A0A5C2FYJ4, A0A5C2FZZ3, A0A5C2GBQ5, A0A5C2GMN5, A0A5C2GN07, A0A7I2V2D2, A0A7S5BYW5, P35900, O00617, P06727 and P35527, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, and corresponding nucleic acids.
39. The kit according to claim 25, characterized in that, The effective protein biomarkers are A0A024QZH6, A0A024R035, A0A024R1G6, A0A024R4F9, A0A024R6P0, A0A024RCW3, A0A0A0MRS8, A0A0G2JMS6, A0A0K2BMD8, A0A0S2Z3X8, A0A0X9V9B3, A0A126LAY7, A0A140KFU0, A0A1B0GVI3, A0A1W2PQX5, A0A1Z1G4M2, A0A2U8J8L9, A0A384MEF1, A0A385HVZ2, A0A5C2FTW7, A0A5C2FVW2, and A0A5C2. FX67, A0A5C2FXP5, A0A5C2FYJ4, A0A5C2FYK5, A0A5C2FZZ3, A0A5C2G130, A0 A5C2G1Q3, A0A5C2G1X1, A0A5C2G3M4, A0A5C2G4U1, A0A5C2G6G6, A0A5C2G7F 2. A0A5C2GA97, A0A5C2GBQ5, A0A5C2GCL9, A0A5C2GDA6, A0A5C2GGY3, A0A5C 2GJH5, A0A5C2GJL5, A0A5C2GMN5, A0A5C2GN07, A0A5C2GRZ5, A0A5C2GV43, A0 A5C2GW23, A0A5C2H3N5, A0A5S8K7B6, A0A7I2V2D2, A0A7P0TAB0, A0A7S5BYW 5. A0A7S5BZY2, A0A7S5C366, A0A7S5EWA8, A0A7S5EYK2, A0A7T0LP36, A0A8F 0WQF6, A0A8G1A656, A2VCK8, A2VCQ3, A4D1J9, A5PL27, A6XNE2, A8K3I0, B2M 1S7, B2R4C5, B2R888, B2RMS9, B4DZH3, B4E1Z4, B4E273, B4E367, B4E3M1, B7W NR0, D9YZU5, E9KL23, E9PHK0, G3V3A0, H0YAC1, H6VRF8, H7BYG8, H7C517, H9 KVD5, L8E853, O00617, O14746, P00738, P00915, P01031, P02533, P02652, P 02671, P02750, P02753, P02763, P04275, P05109, P06727, P0C0L5, P0DJI8, P18428, P27169, P35527, P35900, P35908, P43652, P59666, Q06033, Q2L9S7,Q49A33, Q4TZM4, Q4VXF1, Q5CZ93, Q6IQ49, Q86TT1, Q86YQ4, Q8NHV9, Q9H387, Q9NXP7, and V9H1D9, or any one or more of their antigens / antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins, stable isotopic characteristic peptide segments, or corresponding nucleic acids.
40. An electronic device, characterized in that, include: Memory is used to store instructions that can be executed by the processor; A processor for executing the instructions to implement the construction method as described in any one of claims 1-22.
41. A computer-readable medium storing computer program code that, when executed by a processor, implements the construction method as described in any one of claims 1-22.