Model construction method for evaluating gastric cancer suffering possibility of subject based on Swiss-prot database
By comprehensively utilizing various standardized methods and algorithms to screen protein biomarkers, a gastric cancer prediction model based on the Swiss-prot database was constructed, which solved the problems of insufficient sensitivity and specificity of existing diagnostic methods and achieved efficient diagnosis and screening of early gastric cancer.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-03-27
AI Technical Summary
Existing methods for diagnosing gastric cancer have unsatisfactory sensitivity and specificity, making it difficult to detect gastric cancer in its early stages. Furthermore, deficiencies exist in the screening of protein biomarkers and data processing, affecting diagnostic effectiveness.
We employed a combination of standardization methods to process proteomics data and combined multiple algorithms to screen protein biomarkers. We constructed a gastric cancer prediction model based on the Swiss-prot database, including data standardization, anomaly handling, differentially expressed protein analysis, and correlation analysis, to screen out effective protein biomarkers.
It improves the accuracy and efficiency of gastric cancer diagnosis, enabling accurate and rapid gastric cancer screening in various populations, promoting early detection and treatment, and improving cure and survival rates.
Smart Images

Figure CN121747952A_ABST
Abstract
Description
Technical Field
[0001] This application mainly relates to the fields of biomedicine, biological research and bioinformatics research technology, gastric cancer prediction, proteomics, etc., specifically to models, systems, reagent kits, electronic devices and computer-readable media storing computer program code for assessing the likelihood of subjects developing gastric cancer based on the Swiss-prot database. Background Technology
[0002] Gastric carcinoma (GC), a malignant tumor originating from the gastric mucosal epithelium, ranks fifth in incidence worldwide and is the third leading cause of cancer death. Because early-stage gastric cancer often lacks clear gastrointestinal symptoms, most patients are not diagnosed until they reach the middle or late stages, where prognosis is poor and treatment options are limited, significantly lowering the five-year survival rate. However, existing biomarkers for gastric cancer diagnosis and prognostic assessment have unsatisfactory sensitivity and specificity, limiting their effectiveness in clinical application. Therefore, the diagnosis of gastric cancer currently relies primarily on imaging techniques, including barium X-ray, gastroscopy, abdominal ultrasound, spiral CT, and positron emission tomography (PET). However, these methods all have their limitations. For example, imaging techniques struggle to detect smaller tumors, leading to a high rate of missed diagnoses in early screening. Invasive surgery is inconvenient, and the acceptance rate of gastroscopy and colonoscopy is low among the vast majority of people. This is one of the important reasons for the persistently high mortality rate of gastrointestinal tumors in my country.
[0003] In recent years, liquid biopsy, as a novel detection technology, has gradually attracted widespread attention. It utilizes bodily fluids such as peripheral blood, saliva, urine, or gastric lavage fluid / gastric juice as sources of specific biomarkers, providing new possibilities for the screening and diagnosis of gastric cancer. Biomarkers are a class of biochemical indicators used to label changes or potential changes in the structure, tissues, organs, systems, or functions of cells and subcellular structures, playing a crucial role in disease diagnosis, staging, and evaluation of new therapies. Among them, protein biomarkers have unique advantages in accurately and sensitively screening early, low-level lesions, providing important evidence for early tumor warning and clinical auxiliary diagnosis.
[0004] Nevertheless, the screening and research of protein biomarkers still face many challenges.
[0005] First, current proteomics-based research often lacks sufficient consideration when evaluating the effectiveness of protein biomarkers in practical applications. This deficiency is one of the important reasons for the poor performance of protein biomarkers in real-world applications. Second, many proteomics-based studies often fail to distinguish between random and non-random missing values when processing protein quantification values, instead using the same processing method. This approach may lead to the neglect of some protein biomarkers that are highly discriminative against cancer. For example, a protein that is widely expressed in healthy individuals or patients with benign diseases, but whose expression is significantly reduced or absent in cancer patients, is of great significance for cancer differentiation. However, if missing values for such proteins are uniformly filled with the mean or median, this key biomarker may be incorrectly excluded during subsequent protein biomarker screening, resulting in the loss of important information. Although a few studies have attempted to use different missing value filling methods based on the source of missing protein quantification values, these methods are often difficult to implement in practice.
[0006] Furthermore, existing research indicates that data processing methods significantly impact the selection of proteomics data. For example, the choice of data standardization method is crucial for subsequent proteomics data analysis. However, the optimal standardization methods recommended by different literature or research sources vary considerably, posing a challenge to the consistency of data processing. Simultaneously, the two most prominent characteristics of proteomics data—sparseness and high dimensionality—make data processing particularly complex and challenging. These two characteristics not only increase the difficulty of data processing but may also affect the accuracy and reliability of the analytical results.
[0007] Therefore, it is hoped that a detection method with high accuracy, speed and low cost in multiple populations can be established to help promote the implementation of routine screening and monitoring of high-risk groups for gastric cancer, and effectively help the early detection, early diagnosis and early treatment of gastric cancer, thereby improving the cure rate and survival rate of gastric cancer patients. Summary of the Invention
[0008] Based on existing technologies, the inventors conducted a systematic and in-depth study on the screening of protein biomarkers. Through exploration, they discovered that comprehensively applying various commonly used standardization methods to standardize proteomics data can effectively improve data quality and reliability. Subsequently, the inventors employed algorithms based on partial least squares discriminant analysis, orthogonal partial least squares discriminant analysis, random forest, and lasso retrospective to screen protein biomarkers from various standardized data. This comprehensive screening strategy not only fully utilizes the advantages of various algorithms but also complements each other, thereby improving the accuracy and efficiency of screening. Practical verification shows that this comprehensive screening method achieves extremely excellent screening results.
[0009] To address the aforementioned technical problems, this application provides, in one aspect, a method for constructing a model based on the Swiss-Prot database to assess the likelihood of a subject developing gastric cancer. The method includes: obtaining samples from the subject, the samples comprising multiple serum samples; mass spectrometry data analysis, including: performing mass spectrometry data analysis on the multiple serum samples to extract peptide information and obtain sample data; database retrieval, including: performing a Swiss-Prot database search on the sample data to identify proteins and obtain a raw protein quantification matrix; anomaly handling, including: identifying and removing abnormal samples from the raw protein quantification matrix to obtain an effective protein quantification matrix; and analyzing differentially expressed proteins, including: comparing... The expression differences of proteins within the same group are analyzed to obtain a first set of protein biomarkers with significant differences; data standardization processing includes: standardizing the effective protein quantification matrix to eliminate the influence of different dimensions and scales to obtain a standardized protein quantification matrix; protein biomarker screening includes: using at least m of the following methods to screen the standardized protein quantification matrix to obtain at least m candidate protein biomarker sets: LASSO regression analysis, partial least squares discriminant analysis, orthogonal partial least squares discriminant analysis, and RFS-based random forest algorithm, where m is a positive integer, 3≤m≤4; biomarker summary analysis includes: analyzing the expression differences of proteins within the same group ...ers; protein biomarker summary analysis includes: analyzing the expression differences of proteins within the same group to obtain at least m candidate protein biomarkers; protein biomarker summary analysis includes: analyzing the expression differences of proteins within the same group to obtain at least m candidate protein biomarkers; protein biomarker summary analysis includes: analyzing the expression differences of proteins within the same group to obtain at least m candidate protein biomarkers; protein biomarker summary analysis includes: analyzing the expression differences of proteins within the same group to The protein biomarker set and the m candidate protein biomarker sets are summarized and analyzed. From the first protein biomarker set and the m candidate protein biomarker sets, a target protein biomarker set that exists in at least four sets simultaneously in the first protein biomarker set and the m candidate protein biomarker sets is obtained. Correlation analysis includes: analyzing the correlation between the target protein biomarker set and all protein biomarkers in the effective protein quantification matrix, and forming a new set P1 with the first biomarker among all protein biomarkers whose correlation coefficient is within the first range and the target protein biomarker set. Variable structuring processing includes: structuring set P1. The samples in set P1 are divided into gastric cancer group and non-gastric cancer group according to clinical diagnosis, and are coded respectively to form the dependent variable of the prediction model. The independent variables of the prediction model include the quantitative results of protein, peptide or antibody and the patient's basic information. Model training and testing include: dividing the samples in set P1 into training set and test set, training the prediction model using the training set, and testing the prediction model using the test set. The prediction model is a regression model including the independent variable and the dependent variable. Model optimization includes: screening effective protein biomarkers according to at least one of the indicators of p-value, correlation coefficient and AUC value, using the effective protein biomarkers as effective independent variables, and obtaining the optimized model.
[0010] This application proposes a system for assessing the likelihood of a subject having gastric cancer based on the Swiss-prot database in a second aspect, comprising: a data acquisition module for acquiring sample data of the subject, wherein the sample data includes A0A075B6K4, A0PJY2, A6NFD8, A6QL64, O15033, O15553, O75152, O75460, O75683, O94822, O94885, O95236, O95490, P00738, P00751, P00915, P01009, P01011, P01185, P01242, P01701, P02042, P02100, and P0253. 3. P02538, P02671, P02748, P02750, P02763, P04275, P05109, P05452, P0 6727, P08571, P08575, P0DJI8, P18065, P18428, P31150, P35527, P41235 , P53367, P59666, P59923, P60709, P62328, P62906, P68871, P69891, P69 905, Q01538, Q06033, Q09666, Q12851, Q14156, Q15047, Q15293, Q16280, Q Any one of 2TBA0, Q4G0S7, Q4VXF1, Q53HC9, Q5H9J7, Q5JY77, Q5VU43, Q66K14, Q68DV7, Q6ZVL6, Q8IVV2, Q8IXQ9, Q8IZF2, Q8TA94, Q8TEW8, Q92797, Q92824, Q93034, Q96N87, Q96RV3, Q9H0G5, Q9H5Y7, Q9H6K5, Q9NQV5, Q9NQX4, Q9P0M6, Q9UBF2, Q9UBP8, Q9UFC0, Q9UHG0, Q9UKN7, Q9UN37, Q9UPX8, and Q9Y5E1 ; or a single peptide chain, a characteristic peptide segment of the peptide chain, and its stable isotopic protein or stable isotopic characteristic peptide segment, or an antigen / antibody and any one or more of the single peptide chain, characteristic peptide segment of the peptide chain, and its stable isotopic protein or stable isotopic characteristic peptide segment, and corresponding nucleic acids; a data preprocessing module, used to preprocess the sample data, including removing duplicate samples, unifying units, and dividing the sample data into different analysis groups according to the type and number of indicators; a model recommendation module, used to recommend multiple prediction models with larger AUC values from the prediction models established using the construction method described above, based on the type and number of indicators contained in different analysis groups;The model selection module provides users with a selection function to output one or more predictive models from the predictive models recommended by the model recommendation module for risk assessment and calculation. The risk assessment module calculates the corresponding predicted values based on the one or more predictive models output by the model selection module, and provides an appropriate Logit threshold or risk score threshold according to the department or population information corresponding to the sample data. The risk level is then classified as high probability, medium probability, or low probability based on the Logit threshold or risk score threshold.
[0011] This application, in its third aspect, discloses a kit for assessing the likelihood of a subject having gastric cancer based on the Swiss-prot database, used to receive effective protein biomarkers, wherein the effective protein biomarkers are any one or more of the following protein biomarkers: A0A075B6K4, A0PJY2, A6NFD8, A6QL64, O15033, O15553, O75152, O75460, O75683, O94822, O94885, O95236, O95490, P00738, P0075 1. P00915, P01009, P01011, P01185, P01242, P01701, P02042, P02100, P02533, P02538, P02671, P02748, P02750, P027 63. P04275, P05109, P05452, P06727, P08571, P08575, P0DJI8, P18065, P18428, P31150, P35527, P41235, P53367, P59 666, P59923, P60709, P62328, P62906, P68871, P69891, P69905, Q01538, Q06033, Q09666, Q12851, Q14156, Q15047, Q1 5293, Q16280, Q2TBA0, Q4G0S7, Q4VXF1, Q53HC9, Q5H9J7, Q5JY77, Q5VU43, Q66K14, Q68DV7, Q6ZVL6, Q8IVV2, Q8IXQ9, Q 8IZF2, Q8TA94, Q8TEW8, Q92797, Q92824, Q93034, Q96N87, Q96RV3, Q9H0G5, Q9H5Y7, Q9H6K5, Q9NQV5, Q9NQX4, Q9P0M6, Q9UBF2, Q9UBP8, Q9UFC0, Q9UHG0, Q9UKN7, Q9UN37, Q9UPX8, Q9Y5E1, or a single peptide chain, characteristic peptide segments of the peptide chain, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0012] In a fourth aspect, this application provides an electronic device comprising: a memory for storing instructions executable by a processor; and a processor for executing the instructions to implement the construction method described above.
[0013] In a fifth aspect, this application provides a computer-readable medium storing computer program code that, when executed by a processor, implements the construction method described above.
[0014] The gastric cancer prediction model established according to the construction method of this application is based on data from 903 samples, deriving 41 logistic regression prediction models with AUC values between 0.8 and 0.9 on the test set. The construction method of this application, through meticulous screening of independent variables and progressive screening of independent variables during the training process of the regression model, enables the model to obtain good prediction results with only a few effective independent variables, significantly improving prediction efficiency and saving costs. Therefore, the constructed model, reagent kit, system, and electronic equipment can be applied to a wider range of scenarios and populations, demonstrating extremely high social value. Attached Figure Description
[0015] The accompanying drawings are included to provide a further understanding of this application; they are incorporated into and constitute a part of this application. The drawings illustrate embodiments of this application and, together with this specification, serve to explain the principles of this application. In the drawings: Figure 1 This is an exemplary flowchart of a method for constructing a model based on the Swiss-prot database to assess the likelihood of a subject having gastric cancer according to an embodiment of this application. Figure 2 This is a correlation coefficient matrix diagram of the various markers involved in the embodiments of this application; Figure 3 These are the ROC curve and corresponding AUC results of prediction model 1 in this application embodiment on the training set; Figure 4 These are the ROC curves and corresponding AUC results of the prediction model 1 in this application, which are cross-validated on the training set. Figure 5 These are the ROC curve and corresponding AUC results of prediction model 1 in this application on the test set; Figure 6 This is a system block diagram of an embodiment of the present application for assessing the likelihood of a subject having gastric cancer based on the Swiss-prot database; Figure 7 This is a system block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings involved in the description of the embodiments will be briefly introduced below. It should be understood that the drawings described below are only some examples or embodiments of this application and are not exhaustive of all possibilities. For those skilled in the art, this application can be applied to other similar scenarios based on these drawings without creative effort. Unless obvious from the context or otherwise explained, the same reference numerals in the drawings represent the same structures or operations.
[0017] Furthermore, it should be noted that the use of terms such as "first" and "second" to define components is merely for the purpose of distinguishing the corresponding components. Unless otherwise stated, these terms have no special meaning and therefore should not be construed as limiting the scope of protection of this application. In addition, although the terminology used in this application is selected from commonly known and used terms, some terms mentioned in this application's specification may have been chosen by the applicant according to his / her judgment, and their detailed meanings are explained in the relevant sections of this description. Furthermore, the understanding of this application should not be limited to the actual terms used, but should delve into the meaning implied by each term to fully grasp the substantive content of this application.
[0018] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more steps may be removed from these processes.
[0019] As used herein, the term "cancer" refers to the presence of cells exhibiting typical characteristics of cancerous cells, such as uncontrolled proliferation, immortality, metastatic potential, rapid growth and proliferation rates, and certain characteristic morphological features known in the art. In one embodiment, "cancer" may be gastric cancer or stomach cancer. In one embodiment, "cancer" may include pre-malignant cancer as well as malignant cancer.
[0020] In one embodiment, as those skilled in the art will understand, the methods described herein do not involve steps performed by a physician. Therefore, the results obtained by the methods described herein require consideration of clinical data and other clinical presentations before a final diagnosis by a physician can be provided to the subject. A final diagnosis regarding whether a subject has gastric cancer is within the scope of the physician and is not considered part of this disclosure. Therefore, the terms “determine,” “detect,” and “diagnose” as used herein refer to identifying the probability or likelihood that a subject has a disease (such as gastric cancer) at any stage of development or determining the subject’s susceptibility to developing said disease. In one embodiment, “diagnosis,” “determine,” and “detect” are performed before symptoms appear. In one embodiment, “diagnosis,” “determine,” and “detect” allow a clinician (in conjunction with other clinical presentations) to confirm gastric cancer in a subject suspected of having it.
[0021] As used herein, the term "sample" refers to a sample collected from a subject for the purpose of detecting the types and amounts of gastric cancer markers present therein. Subject samples may be from the circulatory system (i.e., from the blood) or not from the circulatory system (i.e., not from the blood). Subject samples can be any sample containing substances suitable for detecting gastric cancer markers, and their sources include whole blood, bone marrow, pleural fluid, peritoneal fluid, central cerebrospinal fluid, breast milk, urine, tears, sweat, saliva, organ secretions, and lavage fluids from the bronchi, nasal cavity, pharynx, etc.
[0022] In one embodiment, the subject sample is blood, including, for example, whole blood or any portion or component thereof. Blood samples suitable for use in this invention can be extracted from any known source including blood cells or components thereof, such as veins, arteries, peripheral tissues, tissues, spinal cord, and the like. For example, the obtained sample can be obtained and processed using known and conventional clinical methods, such as procedures for drawing and processing whole blood. In one embodiment, the subject sample is serum. Methods for obtaining serum from blood are well known to those skilled in the art.
[0023] This invention discovers that by monitoring whether a sample contains a set of gastric cancer biomarkers, gastric cancer can be diagnosed with high specificity and sensitivity. Especially for early-stage gastric cancer, which was previously difficult to diagnose, the biomarkers of this invention exhibit extremely high specificity and sensitivity.
[0024] As used herein, a “biomarker” or “marker” is a biological molecule that is objectively measured to serve as a characteristic indicator of the physiological state of a biological system. For the purposes of this disclosure, biological molecules include ions, small molecules, peptides, peptide chains, proteins, and peptides and proteins with post-translational modifications, nucleosides, nucleotides, and polynucleotides including RNA and DNA, glycoproteins, lipoproteins, and various covalent or non-covalent modifications of these types of molecules. Biological molecules include any kind of entities that are native to, characteristic of, and / or essential to the function of a biological system. Most biomarkers are polypeptides, although they may also be pre-translational mRNA or modified mRNA representing a gene product expressed as a polypeptide, or may include post-translational modifications of said polypeptide.
[0025] As used herein, "protein biomarker" means the biomarker that contains protein information. In one instance, it refers to the biomarker that contains a protein sequence. Further, in one instance, it refers to a full-length protein, a single peptide chain, a characteristic peptide segment of a peptide chain, and a stable isotopic protein or a stable isotopic characteristic peptide segment thereof.
[0026] The following lists some terms used in the embodiments of the present invention. Within the scope of the specification and claims of this invention, the relevant terms are defined as follows. Other terms not listed are defined using common definitions in the art, and their meanings are well known to those skilled in the art.
[0027] This invention utilizes a protein quantification matrix based on clinical samples. Through a series of steps including abnormal feature processing, feature filtering based on the needs of subsequent practical applications, missing value imputation, abnormal sample identification and processing, data standardization, screening of protein biomarkers, training of logistic regression models, evaluation of the effectiveness of logistic regression models, and evaluation of the predictive performance of logistic regression models, a set of serum protein biomarkers for gastric cancer screening was identified. After further screening, 92 core serum protein biomarkers were obtained. Based on these 92 protein biomarkers, 14 logistic regression prediction models were constructed that can effectively distinguish gastric cancer samples from benign gastric disease samples and have predictive capabilities.
[0028] Non-random missing data: This refers to the situation where, due to problems with the sample itself, some protein quantification data cannot be detected, i.e., non-random missing protein quantification data occurs.
[0029] Random missing data: This refers to the situation where, due to random perturbations in the instrument during the actual protein quantification process, some samples cannot be detected in the protein quantification data, i.e., random missing protein quantification data occurs.
[0030] Principal component analysis (PCA) is a commonly used preprocessing method for linearly reducing the dimensionality of data. Its goal is to use variance to measure the variability of the data and project high-dimensional data with significant variability into a low-dimensional space for representation, thus making it applicable to anomaly detection.
[0031] K-nearest neighbor imputation is a well-known method for imputing missing values. It uses the combined information of multiple nearest neighbors of a sample with missing values to impute the missing values.
[0032] Log2 normalization, also known as Log2 transformation, is a logarithmic transformation of the expression values in the expression matrix, with the base 2.
[0033] Log2+Median Standardization: For the expression values in the expression matrix, first perform Log2 transformation. For each protein or antibody quantitative value of each sample in the transformed matrix, divide the protein or antibody quantitative value by the median value of all proteins or their antigen / antibody quantitative values in that sample, and then multiply by the median value of all proteins or their antigen / antibody quantitative values in all samples.
[0034] Log2+CycLoess standardization: For the expression values in the expression matrix, first perform Log2 transformation, and then use local weighted regression to standardize the transformed expression matrix.
[0035] Log2+Mean standardization: For the expression values in the expression matrix, first perform Log2 transformation. For each protein or its antigen / antibody quantitative value in each sample in the transformed matrix, divide it by the mean of all proteins or their antigen / antibody quantitative values in that sample, and then multiply it by the mean of all proteins or their antigen / antibody quantitative values in all samples.
[0036] VSN Standardization: Since VSN standardization is similar to log2 transformation, there is no need to perform log2 transformation first. The expression values in the expression matrix can be directly standardized by variance stabilization.
[0037] Log2+RLR standardization: For the expression values in the expression matrix, first perform Log2 transformation, and then use robust linear regression to standardize the transformed expression matrix.
[0038] Log2+GI standardization: For the expression values in the expression matrix, first perform Log2 transformation. For each protein or antibody quantitative value of each sample in the transformed matrix, divide the protein or antibody quantitative value by the sum of the quantitative values of all proteins or their antigens / antibodies in that sample, and then multiply by the median value of the sum of the quantitative values of all proteins or their antigens / antibodies in all samples.
[0039] Log2+Quantile normalization: For the expression values in the expression matrix, first perform Log2 transformation, then sort each column separately, calculate the average of the sorted matrix to obtain the average vector, and then replace the corresponding average according to the original matrix sorting.
[0040] PLS: Partial Least Squares Regression Analysis, which combines principal component analysis, canonical correlation, and multiple linear regression into one, is a mapping dimensionality reduction method, and therefore can be used for feature screening of small samples.
[0041] True Yang: Correctly predicts the number of positive samples, which is actually a positive sample, and the prediction is also a positive sample.
[0042] True negative: The number of negative samples is correctly predicted, and the actual number of negative samples is also predicted to be negative.
[0043] False positive: The number of positive samples is incorrectly predicted when they are actually negative samples.
[0044] False negative: The number of negative samples is incorrectly predicted when they are actually positive samples.
[0045] Sensitivity: also known as recall or true positive rate, is the number of correctly predicted positive samples / the total number of actual positive samples, or (true positive) / (true positive + false negative).
[0046] Specificity: also known as the true negative rate, which is the number of correctly predicted negative samples / the total number of actual negative samples, or (true negative) / (true negative + false positive).
[0047] Accuracy: also known as positive predictive value, is the number of correctly predicted positive samples / the total number of predicted positive samples, or (true positives) / (true positives + false positives).
[0048] Accuracy: The number of correctly predicted positive and negative samples / the total number of samples, that is, (true positive + true negative) / (true positive + true negative + false positive + false negative).
[0049] F1 value: F1 = 2 (accuracy) Recall) / (Precision + Recall).
[0050] False positive rate: the number of incorrectly predicted positive samples / the total number of actual negative samples, equal to (1 - specificity), or (false positive) / (true negative + false positive).
[0051] Correlation analysis, or correlation analysis, is a statistical method used to study whether a relationship exists between two or more random variables. Its main purpose is to determine whether a statistical correlation or dependency exists between two or more variables and to quantify the degree and form of this correlation. Basic methods of correlation analysis include linear correlation, rank correlation, and distance correlation.
[0052] Correlation coefficient: In correlation analysis, this statistic quantifies the strength of the relationship between variables, reflecting the intensity of the linear correlation between two variables. A larger absolute value indicates a stronger linear correlation; values closer to 0 indicate a weaker correlation. Commonly used correlation coefficients include Pearson's correlation coefficient and Spearman's correlation coefficient. This study refers to the Pearson correlation coefficient.
[0053] p-value, short for Probability Value, represents the probability of observed data occurring within the hypothesis space. Specifically, the p-value represents the probability of obtaining data equal to or more extreme than the observed result when the null hypothesis is true. Generally, if the p-value is very small, for example, less than 0.01, it means the result is highly unlikely to be a random event under the null hypothesis, so the null hypothesis is rejected, meaning the result is statistically significant. If the p-value is large, for example, greater than 0.05, the null hypothesis cannot be rejected, meaning the result is not statistically significant. The smaller the p-value, the higher the statistical significance of the result. Commonly used significance thresholds are 0.05 and 0.01. Therefore, the p-value reflects the probability of observing the current result under the premise that the null hypothesis is true, and is an important basis for judging whether the hypothesis test result is significant. The smaller the p-value, the more significant the result.
[0054] The t-test, also known as the t-test, is a statistical method used to test whether there is a significant difference between the means of two samples. The basic idea of the t-test is to construct a hypothesis, calculate the t-statistic, determine the p-value based on the t-distribution, and finally use the p-value to determine whether the null hypothesis is true.
[0055] Multicollinearity refers to a strong linear correlation among independent variables. Its main problems include: 1. Affecting the accuracy of least squares estimation, increasing the variance of regression coefficients; 2. Inability to accurately estimate the marginal effect of independent variables on the dependent variable; 3. Leading to insignificant regression coefficients for some independent variables; 4. Decreasing the predictive power of the regression equation. Methods for identifying multicollinearity include: 1. Correlation matrix method: observing the correlation coefficients between independent variables; 2. Variance inflation factor method: an excessively large VIF value indicates multicollinearity; 3. Condition number method: an excessively large condition number indicates multicollinearity. Methods for handling multicollinearity include: 1. Increasing the sample size; 2. Removing multicollinear independent variables; 3. Combining independent variables or using principal component analysis; 4. Regularization methods such as ridge regression. Therefore, multicollinearity can adversely affect regression analysis results and needs to be identified and addressed during the modeling process.
[0056] Lasso regression, also known as lasso regularization, compresses the coefficients of variables in a regression model by generating a penalty function to prevent overfitting and address severe multicollinearity. First proposed by Robert Tibshirani in the UK, Lasso regression is now widely used in predictive models. Lasso achieves sparsity by adding a penalty term (L1 regularization) to the objective function, making the weights of many features in the coefficient vector zero. By selecting features corresponding to non-zero coefficients, it can filter out features with the greatest predictive power for the target variable, thereby simplifying the model and improving its generalization ability. In medical research, multiple correlated independent variables often exist. Lasso regression can reduce the impact of multicollinearity on regression results by making the coefficients of correlated independent variables zero based on their correlations.
[0057] Random Forest: Random forest is an ensemble learning method that constructs multiple decision trees and obtains the final prediction result by voting or averaging the outputs of these trees. Each decision tree is built on a randomly selected subset of features and samples, which helps reduce overfitting and improve the model's generalization ability. Random forests can be used not only for classification and regression tasks but also for feature importance evaluation.
[0058] Recursive Feature Elimination (RFE): RFE is a greedy search algorithm that selects features by recursively considering increasingly smaller sets of features. At each step, the model is trained on the current feature set and removes the least important features (e.g., based on the absolute value of the coefficients or some measure of importance in the model output). This process continues until the desired number of features is reached or other stopping conditions are met.
[0059] Random Forest with Automatic Recursive Feature Elimination (RF-RFE): Combining random forests with automatic recursive feature elimination allows for more efficient feature selection. The specific steps are as follows: 1. Train the random forest model: First, train a random forest model using all features. 2. Evaluate feature importance: Determine the importance of each feature using the random forest model's feature importance evaluation mechanism (e.g., based on Gini impurity reduction or out-of-bag accuracy). 3. Recursively eliminate features: Recursively remove the least important features based on their importance, and retrain the random forest model in each iteration. 4. Select the optimal feature subset: Select the subset of features that best performs the model using cross-validation or other validation methods. Advantages include: 1. Reducing the complexity of the model by decreasing unimportant features, thus reducing the risk of overfitting. 2. Selecting the most useful features improves the model's predictive performance. 3. Examining the selected features allows for a better understanding of the data and interpretation of the model's predictions.
[0060] Partial Least Squares Discriminant Analysis (PLS-DA) is a supervised pattern recognition multivariate statistical analysis method that groups multidimensional data into groups based on the discrepancies to be identified before compression (by pre-setting Y values for target classification and discrimination). This approach helps identify the most relevant variables for grouping while reducing the influence of other factors.
[0061] Orthogonal Partial Least Squares Discriminant Analysis (OPLS-DA): Combining Orthogonal Signal Correction (OSC) and PLS-DA methods, it can decompose the information of the X matrix into two types of information, Y, which are correlated and uncorrelated. By removing the uncorrelated differences, it can screen for differential variables.
[0062] Logistic Regression (LR) is a commonly used classification model that establishes the relationship between independent variables and categorical dependent variables. The main characteristics of a logistic regression model are: 1. The predicted dependent variables are discrete variables of binary or multi-category classification; 2. The logistic function is used to convert the values of linear combinations of independent variables into probabilities between 0 and 1; 3. It can handle both categorical and continuous variables; 4. Parameter estimation typically uses the maximum likelihood method; 5. It can explain the influencing factors of classification and the weights of each dependent variable; 6. It can calculate classification probabilities and make category predictions. The main steps in establishing a logistic regression model are: (1) Collect data and handle missing values, etc. (2) Select input variables and handle categorical variables. (3) Establish the logistic regression equation. (4) Estimate parameters using maximum likelihood. (5) Evaluate the overall effect of the model. (6) Conduct statistical tests to evaluate the impact of each variable. (7) Establish classification rules using probabilities. (8) Predict new data. The logistic regression model can both quantitatively analyze the effects of variables and be used for classification prediction, making it a very useful classification analysis method.
[0063] ROC curve and AUC value: Both are standards used to measure the performance of a classifier. The ROC (Receiver Operating Characteristic) curve has the false positive rate on the horizontal axis and the true positive rate on the vertical axis. The plotted curve should lie above the y=x line. The area under the ROC curve is the AUC value. The larger the AUC, the better the classification performance of the classifier (such as the logistic regression model).
[0064] The present application will be further described below with reference to the embodiments and accompanying drawings.
[0065] Example 1: Method for constructing a model to assess the probability of a subject having gastric cancer based on the Swiss-prot database
[0066] Figure 1 This is an exemplary flowchart illustrating a method for constructing a model based on the Swiss-prot database to assess the likelihood of a subject developing gastric cancer, according to an embodiment of this application. (See reference...) Figure 1 The construction method of this embodiment includes the following steps: Step S1: Obtain samples from the subjects; Step S2: Mass spectrometry data analysis; Step S3: Database retrieval; Step S4: Exception handling; Step S5: Analyze differentially expressed proteins; Step S6: Data standardization processing; Step S7: Screening for protein biomarkers; Step S8: Biomarker summary analysis; Step S9: Correlation analysis; Step S10: Variable structuring; Step S11: Model training and testing; Step S12: Model optimization.
[0067] The following details steps S1 to S12. Step S1 involves obtaining samples from known sources, consisting of a known number of benign gastric disease samples and a known number of gastric cancer samples. The samples in Step S1 were obtained from 365 serum samples from non-gastric cancer individuals and 537 serum samples from stage I-IV gastric cancer individuals obtained from Shanghai Zhongshan Hospital, serving as the samples for this invention. The collection of these samples followed the ethical standards established by the Ethics Committee of Shanghai Zhongshan Hospital, and informed consent forms were signed.
[0068] In steps S2 and S3, mass spectrometry data analysis and Swiss-prot database searches are performed on samples from known sources to obtain protein quantification matrices for samples from known sources.
[0069] Mass spectrometry data was analyzed in depth using mass spectrometry software to accurately identify peptide sequence information. These peptide sequences were then compared and searched against the Swiss-prot protein database to confirm their corresponding proteins. Based on the relative abundance of peptides, the identified proteins were quantitatively analyzed, thereby constructing a protein quantification matrix for samples from known sources.
[0070] In some embodiments, the step of obtaining sample data in step S2 includes: Sample preparation 1. Protein Extraction and Quantification: High-abundance proteins in serum were separated using a multiple affinity removal column. High-abundance and low-abundance proteins were collected separately, and the high-abundance and low-abundance fractions were desalted and concentrated using 5 kDa ultrafiltration tubes. SDT buffer (4% SDS, 100 mM Tris-HCl pH 7.6) was added, and the mixture was boiled for 15 minutes and centrifuged at 14,000 g for 20 minutes. The supernatant was then quantified using a BCA protein assay kit. Samples were stored at -80°C.
[0071] 2. Peptide Extraction: DTT (final concentration 10 mM) was added to the samples and mixed at 600 rpm for 1.5 hours (37°C). After the samples cooled to room temperature, IAA was added to the mixture to a final concentration of 20 mM, and incubated in the dark for 30 minutes. Next, the samples were transferred to a 10 kDa filter. The filter was washed three times with 100 μL UA buffer, and then twice with 100 μL 25 mM NH4HCO3 buffer. Finally, trypsin was added to the samples and incubated at 37°C for 15–18 hours. The obtained peptides were collected as the filtrate.
[0072] The peptides were desalted using a C18 column, and their concentration was determined at OD280. Then, 2 μg of peptide was extracted from each sample, incorporated with an appropriate amount of iRT standard peptide, and analyzed by LC-MS / MS DDA and LC-MS / MS DIA methods.
[0073] LC-MS / MS analysis
[0074] 1. All fractions used for DDA library generation were analyzed on an Easy-nLC 1200 chromatography system - OrbitrapExploris 480 mass spectrometer (Thermo Scientific). Peptides were separated on a C18 analytical column using a linear gradient of buffer B (84% acetonitrile in 0.1% formic acid) at a flow rate of 300 nL / min. MS was performed in positive ion mode, with a scan range of 350–1800 m / z, an MS1 scan resolution of 60,000 (@m / z 200), and an AGC (automatic gain control) target value of 1e. 6 The maximum IT is 50 ms, and the dynamic exclusion time is 10.0 s. After each complete MS-SIM scan, 20 ddMS2 scans are performed based on the include list. The isolation window is 1.5 m / z, the MS2 scan resolution is 30000 (@m / z 200), and the AGC target is 1e. 5 The maximum IT is 50 ms, and the collision energy is 30 eV.
[0075] 2. DIA scan analysis
[0076] The chromatographically separated sample was analyzed by DIA mass spectrometry. Ionization mode: positive ion. Primary mass spectrometry scan range: 350-1800 m / z, mass spectral resolution: 120,000 (@m / z 200), AGC target: 3e 6Maximum IT: 30 ms. MS2 uses DIA data acquisition mode, with 44 DIA acquisition windows set, mass spectrometry resolution: 30,000 (@m / z 200), AGC target: 3e 6 Maximum IT: auto, MS2 Activation Type: HCD, Collision Energy: 30 eV, Spectral datatype: profile.
[0077] 3. Mass spectrometry data analysis
[0078] For DDA data, use Spectronaut TM 18. The FASTA sequence database was searched using Biognosys software. Parameter settings: enzyme: trypsin; maximum missed cleavage: 1; fixed modification: carbamoylmethyl (C); dynamic modifications: oxidation (M) and acetyl (protein N item). All reported data are based on a 99% confidence level for protein identification, determined by a false discovery rate (FDR) ≤ 1%.
[0079] DIA data using Spectronaut TM 18. Analysis was performed, searching the constructed spectral library. Key parameter settings: retention time prediction type was set to dynamic iRT, MS2 level interference correction was enabled, and cross-run normalization was enabled. All results were filtered based on a Q-value cutoff of 0.01 (equivalent to FDR < 1%).
[0080] In some embodiments, the database retrieval in step S3 includes: processing the obtained mass spectrometry data using Spectronaut software. TM 18) A search of the Swiss-prot database was performed. During the search, the software's default parameters were used to identify proteins, ultimately yielding a quantitative matrix containing 902 samples and 4384 proteins. This matrix provides detailed quantitative information on the proteins in the samples.
[0081] In some embodiments, between step S3 and step S4, there are further data processing steps, including: Checking for batch effects: Due to the large sample size, the data was divided into 7 batches for peptide extraction and analysis. Before proceeding with further analysis, it was crucial to carefully examine and assess whether batch effects existed between the different batches. Principal component analysis revealed the presence of batch effects.
[0082] Removing batch effects: Use the R package statTarget and the QC-RFSC method to remove batch effects between samples caused by different batches based on the QC samples. statTarget is set to the default parameters.
[0083] In some embodiments, between steps S3 and S4, protein filtering is further included. Due to technological limitations, quantitative values of protein biomarkers may be randomly missing. To improve the effectiveness of the application, proteins with quantitative values in a predetermined number of samples are selected for subsequent protein biomarker screening; the number of qualified proteins is 1196. In one embodiment, this predetermined number is 80% or more.
[0084] In some embodiments, between step S3 and step S4, the method further includes: filling missing values: for both non-randomly missing protein quantification values and randomly missing protein quantification values, the K-nearest neighbor method is used to calculate the filling.
[0085] Step S4, the anomaly handling, specifically includes: using principal component analysis to identify 13 abnormal samples, which are then removed from the population, resulting in a protein abundance information matrix with 889 samples and 1196 proteins. Step S5, the differentially expressed protein analysis, includes: after the rigorous processing described in steps S1-S3 above, including batch effect removal, protein filtering, and missing value imputation, expression levels of the protein biomarkers to be screened in benign gastric lesion samples and gastric cancer samples, respectively, are obtained. To further explore the expression differences of these proteins in the two samples, two classic statistical analysis methods are used: t-test and Deseq2. T-test-based differential protein analysis initially screened out proteins with significant differences between the two groups of samples, forming a protein biomarker set ①, i.e., set G1. Deseq2-based differential protein analysis further considers the variability between samples and sequencing depth, providing more refined and reliable information on differentially expressed proteins, thus forming a protein biomarker set ②, i.e., set G2. Set ① contains a total of 60 proteins, namely A6NFD8, A6QL64, O15075, O15553, O75152, O75683, O94822, O94885, O95490, P00738, P01011, P02042, P02100, P02533, P02671, P02748, P02750, P04275, P05109, P0DJI8, P18065, P18428, P25092, P33527, P41235, P53367, P59666, P59923, P60709, P62328, P62906, and P68871. P69905, Q02218, Q06033, Q09666, Q15047, Q15293, Q16280, Q2TBA0, Q4G0S7, Q4VXF1, Q5H9J7, Q5VU43, Q66K14, Q68DV7, Q8IVV2, Q8IZF2, Q93034, Q99683, Q9BXU1, Q9H5Y7, Q9NQV5, Q9P0M6, Q9UBP8, Q9UKN7, Q9UL25, Q9UN37, Q9UPX8 and Q9Y5E1 or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.Set ② contains a total of 70 proteins, namely P02750, O75152, P02748, P59666, P02671, Q06033, P01011, Q9UN37, P04275, P62328, Q15293, O94822, P62906, O15553, Q68DV7, P60709, Q93034, P 41235, P05109, Q9NQX4, Q09666, P59923, A6NFD8, Q4G0S7, Q66K14, Q15047, P0267 9. Q5VU43, Q9NQV5, Q9Y5E1, Q9UPX8, P18065, P18428, P53367, P00738, Q16280, P3 1150, Q9P0M6, P02100, Q9BZG1, Q9UKN7, P69905, Q4VXF1, Q8IZF2, P02533, O95490 , Q9H5Y7, P35527, Q2TBA0, Q9UL25, O75683, O94885, Q5H9J7, Q12756, P33527, Q8T EW8, P02042, A6QL64, Q8IVV2, Q9UBP8, Q99683, P25092, Q6UWP8, O14867, O15061, Q15700, O15075, P04075 and P78527 or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0086] Step S6, data standardization, specifically includes the following: Based on preliminary analysis, the inventors discovered that the combinations of protein biomarkers obtained by different standardization methods may vary significantly. Considering the effectiveness of practical applications, the inventors of this invention employed eight standardization methods (Log2, Log2+Median, Log2+Mean, VSN, Log2+RLR, Log2+GI, Log2+Quantile, and Log2+CycLoess) to standardize the protein expression matrices after steps S1-S3, including batch effect removal, protein filtering, and missing value imputation. All eight standardized protein expression matrices were then screened using m of the following methods: LASSO regression, PLS-DA analysis, OPLS-DA analysis, and protein biomarker screening based on the RFS-based random forest algorithm. Here, m is a positive integer, 3 ≤ m ≤ 4.
[0087] LASSO regression: The cv.glmnet function in the glmnet package of R language was used to screen protein markers by LASSO regression on the protein expression matrices obtained by the eight standardization methods. The set of protein markers obtained by analyzing at least seven standardization methods was selected for subsequent analysis. Set ③ contains a total of 197 proteins, namely P02749, P01008, P19827, P02671, P62328, P01009, Q93034, P01024, P60709, P02787, P59666, P02745, P00751, Q8IWZ5, P98170, P00746, O95573, Q9H0G5, P41235, O43361, A0A075B6K4, A0A075B6S5, O15553, Q04118, O15056, P02538, Q9NY93, Q9NQX4, Q02 388, P63167, Q9HDC9, Q99996, Q96A09, Q4G0F5, O43147, Q8NBJ4, P01701, Q8TEW8, Q9NZ43, P08637, P02788, P21817, A6NHZ5, O60239, P1 8065, A2A3K4, A0A087WSZ0, Q8NE65, Q8NDH2, P12236, Q6P9F7, Q8IVF5, Q3KQU3, Q12851, P41250, P02452, Q5THJ4, P01700, Q9Y6D9, P239 75. Q9UMZ3, O43526, P02533, Q06141, Q9H0J4, P42263, O75683, Q9UBF2, P42356, P11226, Q9UMS6, Q8IZT6, Q9H7N4, Q15349, Q96PX9, Q92 499, Q07890, O75116, Q8TDW0, Q5TGY3, Q9UJ41, Q8IVV2, O15504, P48735, P23634, P0DOX2, Q76LX8, O75369, Q9Y5R2, P33908, P04275, A1 YPR0, Q8TDX5, Q5TAT6, P08729, P69905, Q9BXL7, Q7KZ85, O14936, Q15848, Q8WXH0, Q9UKL3, Q8NB15, P68363, Q9NUQ3, P05455, A0A0A0MR Z9, Q9NQ79, Q02218, Q7Z4L5, A6NM28, Q9UHG0, Q9BYB0, Q9Y692, Q99714, P46940, Q8NG48, P08709, A0A0A0MS15, Q9Y5S2, Q92824, Q9Y5W8,Q9UBP8, P35916, O95236, A0A0A0MS14, Q8IZJ1, P04211, Q6IPT4, O15015, O75694, Q9GZW8, A0A075B6I6, Q99728, Q96KN2, Q96M20, P02100, P62736, O60293, P01042, Q5T1A1, Q9NX52, P 36955, A0A075B6H7, Q6IC98, Q68DE3, Q16610, A4QPB2, Q6P9G9, Q8WZ42, A0A0B4J1Y8, Q9HC I7, Q9NS15, Q9HBM1, Q9H497, P04406, Q08380, Q13670, P43121, Q14896, P05452, P01714, Q6 ZVL6, Q8TA94, P35443, Q8N283, P00915, Q8NHQ9, Q8IVF2, A0A087WW87, A4UGR9, Q6ZRR7, P0 6702, P01780, Q8TBZ0, A0A0B4J1V0, P02743, Q8IWS0, A0A0C4DH30, P10244, Q99969, Q16280 P09172, A5LHX3, Q8IXQ9, P02753, P07358, P01242, Q92954, P07360, P13473, O94822, Q14156, A0A075B6S2, P02647, Q96SU4 and P02774 or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0088] Partial Least Squares Discriminant Analysis (PLS-DA): Protein expression matrices obtained by eight normalization methods were screened for protein biomarkers using PLS-DA, and protein biomarkers that appeared simultaneously in the results of all eight normalization methods were selected to form a set of protein biomarkers. Set ④ contains a total of 225 proteins or peptides, namely P02750, P62328, P02748, O75152, P02671, P01009, P59666, Q06033, Q9UN37, Q93034, P01011, O94822, P04275, P0DJI8, P02763, P62906, P05155, P60709, P00450, Q9H0G5, Q16280, Q9NQX4, A0A1B0GUA7, P41235, O15553, Q9UHG0, P68871, P02753, P00751, P6 9905, Q15293, Q8IXQ9, Q9BZG1, P06396, P05452, Q68DV7, P00740, P06727, P02100, P06276, Q9H6K5, A6NFD8, P08571, A0A075B6K4, P43652 , Q9UFC0, P53367, O95236, Q14624, P05109, Q9NXP7, P02749, P61626, Q38SD2, Q03591, P29622, Q8TA94, P01031, P02538, Q09666, P08575, P 18065, O75882, P00915, P80108, P27918, P02533, P00738, Q5VU43, P51884, A0A075B6S2, A0PJY2, Q9UPX8, P01242, P51570, P02656, P0264 7. Q96SU4, P49747, P0C0L4, Q15047, Q9Y5E1, Q92824, Q8TC76, Q8IWZ5, P31150, P59923, Q9UBF2, P06681, Q4G0S7, P07360, P27169, P35542, P13671, Q9UKN7, P35443, Q14156, P02652, P01008, P35858, P02743, P36776, Q8NCA9, P02751, P18428, P00736, P01024, Q9P2D7, P04003, Q1 6552, Q5JY77, P07357, Q4VXF1, P17936, P19827, Q9NYC9, A0A0B4J1V6, Q8IUY3, Q66K14, P06312, P15848, Q9NQV5, Q9BYW2, Q9P0W8, Q3KQU3,Q9UBP8, P19823, Q6IPT4, O75144, Q8NBJ4, P05156, Q15542, P69891, Q14520, P43121, P51164, Q9BXL7, P07358, P02 042, P03952, P25311, P02774, P01780, Q562R1, P33908, Q96KN2, P01344, P01701, O95573, Q9UMS6, Q8IVV2, Q9H5Y7, P35908, Q8IZF2, P01042, Q9UK55, O75460, Q6IC98, O95490, O94967, Q5ST30, Q5TAT6, Q6ZVL6, Q6P387, Q9NR34, Q9N V72, Q86UK7, Q9NZP8, O75683, O15033, Q2TBA0, A0A075B6N3, P98170, O43166, P13473, P00747, P05543, Q9Y5P4, P35 527, P33981, Q13790, Q96HS1, Q9H7N4, Q9BXR6, Q92797, Q13200, Q99742, Q9UL25, Q05469, Q96PD5, Q9Y250, A1YPR0 , P82930, P23142, P02765, Q9P0M6, A0A0C4DH30, P07359, P28347, Q96LK0, A6NM28, Q86T24, Q9UN81, P10909, A4QPB2 P09172, Q9BTV5, Q6P9F7, P00746, P11226, O95817, Q8TEW8, Q96N87, A0A0C4DH73, Q4G0F5, O14795, Q9P2P1, O94885, A0A075B6J9, P02788, P36980, Q96RV3, Q02388, P42263, and A5LHX3, or a single peptide chain, characteristic peptide segments of the peptide chain, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0089] Orthogonal Partial Least Squares Discriminant Analysis (OPLS-DA): OPLS-DA was applied to the protein expression matrices obtained by eight standardization methods to screen for protein biomarkers. To ensure the stability and reliability of the results, proteins that appeared in the results of all eight standardization methods were specifically selected to form the final set of protein biomarkers. Set ⑤ contains a total of 181 proteins or peptides, namely P0DJI8, Q9Y5R5, A6NFD8, Q9UMZ3, Q9Y5E1, Q5TAT6, Q13733, Q9H7N4, Q8NBJ4, Q68DV7, O75460, P41235, O15553, P31150, P98170, Q92820, P82930, P59923, P28347, Q9UFC0, P00738, P69905, P59666, P68871, P02042, Q15542, Q5VU43, Q8TA94, and P02743. , P18428, P69891, P05109, Q9NQV5, Q9UKN7, Q9UPX8, Q96RV3, Q5ST30, O75152, A0PJY2, Q12756, Q8IXQ9, P02750, Q14156, P61626, O150 33. P01011, Q15293, Q9UN37, Q6P9F7, Q9BYW2, Q09666, P02763, P02748, Q92824, Q5JY77, Q06033, Q16560, P08571, P32119, Q9UN81, P0 0751, P43121, Q9P0M6, Q4G0S7, P05156, P00915, Q15047, P51570, P0C0L4, Q9H5Y7, Q9BXL7, P11226, Q9NXP7, P04275, P08575, P01009, P25311, Q8TC76, Q86T24, P06681, Q16552, P33981, Q8TEW8, P0C0L5, P00736, P35542, Q8IZF2, Q9UBP8, O00469, Q9H0G5, P00740, Q9BTV 5. Q86UK7, Q8IUY3, O95490, P01031, A0A0B4J1V6, Q66K14, Q14624, Q03591, Q13790, P13671, P02100, P26927, Q9UK55, Q9UHG0, A2A3K4 , A0A075B6S2, P51884, P01780, Q4VXF1, Q96KN2, P06396, Q9BZG1, Q9P0W8, P49747, P36776, Q9Y6N6, P00450, Q3KQU3, P05543, P27918,P53367, P05155, Q9UMS6, Q8IWZ5, Q9NX52, P06312, Q9UBF2, O75882, Q6IPT4, A0A1B0GUA7, P23142, Q9NZP8, Q38SD2, P07359, P01242, P06276, A 1YPR0, Q96HS1, Q9NYC9, A4QPB2, Q9HBG6, P36980, P35527, P80108, P03952, P01701, P62906, P01042, Q9Y5S2, P02749, O94885, P02753, P29622 P01344, P27169, Q96PY5, O14795, P05452, Q2TBA0, Q6ZVL6, P49756, P18065, P09172, O94822, P15848, O75144, A0A0C4DH30, A0A075B6J9, P06727, P02647, P02652, Q16280, Q8NCA9, Q99742, A0A075B6K4, O75683, P02538, O95573, and P17936, or a single peptide chain, characteristic peptide segments of the peptide chain, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0090] RFS-based Random Forest: Protein expression matrices obtained through eight normalization methods were used to screen for protein biomarkers using RFS-based random forest. Subsequently, a set of protein biomarkers analyzed under at least six normalization methods was selected for further analysis. Set ⑥ contains a total of 53 proteins, namely P62328, O94822, P01009, P02748, Q93034, O75152, P59666, P02750, P68871, P62906, Q8IXQ9, P41235, P02538, Q06033, P05452, Q8TA94, Q9UN37, P53367, P02100, P69905, Q9H0G5, Q14156, P04275, Q9NQX4, P02533, Q9UHG0, A0PJY2, Q16280, and Q86T24. A0A1B0GUA7, P06396, Q6ZVL6, P00915, Q9UFC0, O15553, P33981, Q9UBF2, P27918, P01011, Q13670, P05155, P02652, P19320, A0A075B6K4, Q92824, Q16552, Q6IC98, Q14520, P69891, P35527, Q15047, Q8NCA9 and P01242 or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, and the corresponding nucleic acids.
[0091] Step S8, the biomarker summary analysis, includes: summarizing and analyzing the biomarkers in sets ① to ⑥, yielding: 8 target protein biomarkers appearing in all 6 sets, namely O94822, P59666, P41235, P02100, P69905, P04275, Q16280, and O15553; and 15 target protein biomarkers appearing in any 5 of the 6 sets, including P6232. 8. P02748, Q93034, O75152, P02750, P62906, Q06033, Q9UN37, P53367, P02533, P01011, Q15047, P18065, O75683, Q9UBP8; 45 target protein biomarkers appearing in any 4 of the 6 sets, namely P01009, P68871, Q8IXQ 9. P02538, P05452, Q8TA94, Q9H0G5, Q14156, Q9NQX4, Q9UHG0, Q6ZVL6, P00915, Q9UBF2, A0A 075B6K4, Q92824, P35527, P01242, P02671, P60709, Q8TEW8, Q8IVV2, P0DJI8, Q15293, Q68D V7, A6NFD8, P05109, Q09666, P00738, Q5VU43, Q9UPX8, Q9Y5E1, P59923, Q4G0S7, Q9UKN7, P18428, Q4VXF1, Q66K14, Q9NQV5, P02042, Q9H5Y7, Q8IZF2, O95490, Q2TBA0, Q9P0M6, O94885. A total of 68 important target protein biomarkers were obtained.
[0092] The correlation analysis in step S9 includes the following: When using machine learning methods such as random forests for feature selection, some important biomarkers may be missed. This is usually because these methods rely on the statistical properties of the data to evaluate the importance of each feature, rather than directly utilizing biological knowledge. Therefore, we perform a correlation analysis between the 68 important target protein biomarkers obtained in the previous step and all 1196 total protein biomarkers. The biomarkers with correlation coefficients within the first range to the 68 target protein biomarkers are then grouped with these 68 biomarkers to form a new set P1 containing 92 biomarkers. These markers are A0A075B6K4, A0PJY2, A6NFD8, A6QL64, O15033, O15553, O75152, O75460, O75683, O94822, O94885, O95236, O95490, P00738, P00751, P00915, P01009, P01011, P01185, P01242, P01701, P02042, and P02100. , P02533, P02538, P02671, P02748, P02750, P02763, P04275, P05109, P05452, P06727, P08571, P08575, P 0DJI8, P18065, P18428, P31150, P35527, P41235, P53367, P59666, P59923, P60709, P62328, P62906, P688 71. P69891, P69905, Q01538, Q06033, Q09666, Q12851, Q14156, Q15047, Q15293, Q16280, Q2TBA0, Q4G0S7 , Q4VXF1, Q53HC9, Q5H9J7, Q5JY77, Q5VU43, Q66K14, Q68DV7, Q6ZVL6, Q8IVV2, Q8IXQ9, Q8IZF2, Q8TA94, Q 8TEW8, Q92797, Q92824, Q93034, Q96N87, Q96RV3, Q9H0G5, Q9H5Y7, Q9H6K5, Q9NQV5, Q9NQX4, Q9P0M6, Q9UBF2, Q9UBP8, Q9UFC0, Q9UHG0, Q9UKN7, Q9UN37, Q9UPX8, and Q9Y5E1, or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments. Some of these 92 proteins exhibit high correlations; detailed results are attached. Figure 2 The darker the color, the higher the correlation. Proteins with high correlation have a substitution relationship and can substitute for each other during the modeling process.
[0093] Step S10, variable structuring, includes: arranging the quantitative results of proteins, peptides, or antibodies in set P1, along with the patient's basic information, in a specific order to form multiple independent variables; then, based on the patient's clinical diagnosis, dividing the samples into a gastric cancer group and a non-gastric cancer group, and encoding them as 1 and 0 respectively, forming the dependent variable. Finally, these processed independent and dependent variables are organized into a two-dimensional table or data matrix, where each column represents a variable and each row represents an independent sample, facilitating the subsequent use of this information to build a predictive model.
[0094] Step S11, model training and testing, includes: dividing the data into training and testing sets; then, using the training set data to construct a regression model containing independent variables; and finally, using the testing set data to test the performance of the regression model. The data includes multiple independent and dependent variables.
[0095] Step S12, model optimization, includes: filtering effective independent variables based on any one of the following: the p-value of each independent variable in the regression analysis results, the correlation coefficient, and the AUC value of the regression model on the training and test sets, to obtain the filtered effective independent variables. Specifically, when the p-value of an independent variable is less than 0.05, that effective independent variable is included in the filtered effective independent variables.
[0096] In some embodiments, after performing step S12, step S13 may be included: model iteration, which includes repeating the above steps S11 and S12 to build multiple models and summarize their result data on the training set and test set to evaluate the performance of different models.
[0097] The order and numbering of steps 2 to S12 above are not used to restrict the execution order of each step.
[0098] The inventors of this application conducted research on the screening and simplification of existing gastric cancer biomarkers and discovered that if existing gastric cancer biomarkers are structured into independent variables, and then samples are grouped and coded according to clinical diagnoses to form dependent variables, then variable screening can be performed through correlation analysis and statistical tests between the independent and dependent variables. The ability of an independent variable or gastric cancer biomarker to distinguish different clinical groups is judged by the magnitude of the correlation coefficient between the independent and dependent variables. The larger the correlation coefficient, the stronger the ability of the independent variable or gastric cancer biomarker to distinguish different clinical groups, and the more likely it should be prioritized in constructing gastric cancer prediction models. The significance of the numerical distribution of the independent variable between two clinical groups and the size of the p-value can also be used to determine the distinguishing ability of the independent variable / gastric cancer biomarker for the dependent variable or different clinical groups. Meanwhile, the inventors also discovered that correlation analysis between independent variables or gastric cancer biomarkers can reveal their relationships. Two strongly correlated independent variables or gastric cancer biomarkers can be substituted for each other in model construction or practical application, and only one needs to be selected in model construction, thus achieving model simplification. Furthermore, the inventors found that if a logistic regression model containing all variables, including effective variables, is constructed on the training set, and the model's performance is tested on the test set, the p-values, coefficients, and overall AUC of each variable in the logistic regression results, combined with the aforementioned correlation analysis and statistical test results, can further screen variables, narrowing down the set of effective variables, thereby achieving excellent gastric cancer biomarker screening and model simplification effects. Based on the above findings, this application proposes a method for constructing a model based on the Swiss-prot database to assess the probability of subjects developing gastric cancer.
[0099] Based on the above steps, this application obtained a total of 14 models, which are described below in conjunction with Examples 2 to 15. It should be noted that the training and test sets used in each model are the same. For each model, the Receiver Operating Characteristic (ROC) curve and the Area Under the Curve (AUC) were used as evaluation metrics, and these metrics were calculated on the training, validation, and test sets. The ROC curve, by comprehensively judging the model's ability to distinguish between positive and negative samples, can intuitively reflect the model's classification performance; while the AUC value can quantify the ROC curve and further evaluate the model's predictive sensitivity and specificity. Performing ROC and AUC analyses can comprehensively evaluate the effectiveness of our constructed predictive models, providing a basis for subsequent model optimization and clinical application.
[0100] Example 2: Establishment and Validation of Logistic Regression Prediction Model 1
[0101] On the training set, a logistic regression model with two independent variables, P01009 and P59666, was constructed to predict the probability of a sample having gastric cancer. First, based on the quantitative data of P01009 and P59666 and the sample labels, prediction model 1 was obtained through logistic regression. The mathematical expression of prediction model 1 is: P=1 / (1+Exp(-(β0+β1 ×P01009 + β2 × P59666))), where P represents the probability of the sample being classified as a positive example, the name of the protein biomarker in the formula refers to the standardized content value of the protein biomarker, the coefficient β0 ranges from [-54.02, -36.56]; the coefficient β1 ranges from [2.41, 4.02]; and the coefficient β2 ranges from [0.81, 1.37]. Preferably, β0 = -45.29, β1 = 3.22, and β2 = 1.09 yield the optimal prediction results. The established logistic regression model can predict the likelihood of a new sample developing gastric cancer based on quantitative results of any one or more of the following: proteins or antibodies, single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins or stable isotope characteristic peptide segments, and nucleic acids of P01009 and P59666.
[0102] The ROC curve and corresponding AUC results of prediction model 1 on the training set are as follows: Figure 3 As shown in the figure, the ROC curve and corresponding AUC results of the model on the training set show that the overall performance of the model is good, and the AUC value is also high at 0.87. This indicates that the model has good predictive performance in distinguishing between gastric cancer samples and non-gastric cancer samples.
[0103] To comprehensively verify the model's stability and generalization ability, a 5-fold cross-validation method was used for internal cross-validation. Specifically, the training set was divided into five mutually exclusive subsets. Four subsets were used as the training set in each round, and the remaining subset was used as the validation set, for a total of five rounds of training and validation. 5-fold cross-validation can measure the model's stability across different data subsets. The ROC curve and corresponding AUC results for the model's 5-fold cross-validation on the training set are shown below. Figure 4 As shown in the figure, the ROC curve and corresponding AUC results of the 5-fold cross-validation show that, after 5-fold cross-validation, the final model achieved an average AUC of 0.87 and a standard deviation of 0.02 on the validation set, meeting the set performance target. This also proves that the model is stable and effective, providing a guarantee for the practical application of the model.
[0104] After obtaining satisfactory training set results, prediction model 1 is applied to an independent test set for prediction, and various evaluation metrics of the model on the test set are calculated. This process verifies the model's generalization ability from the training set to the test set, examines whether the model overfits on unseen samples, and thus further evaluates the model's generalization performance. The ROC curve and corresponding AUC results of prediction model 1 on the test set are shown below. Figure 5 As shown in the figure, the ROC curve and corresponding AUC results of prediction model 1 on the test set indicate that the overall performance of prediction model 1 on the test set is good, and the corresponding AUC value is also high at 0.86. This shows that prediction model 1 maintains a good ability to distinguish between gastric cancer samples and non-gastric cancer samples on the test set. Furthermore, the AUC on the test set is comparable to that on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, prediction model 1 has good predictive performance on the test set and can be clinically validated and applied.
[0105] According to prediction model 1, it is only necessary to obtain two indicators, P01009 and P59666, or a single peptide chain, characteristic peptide segments of the peptide chain, and its stable isotope protein or stable isotope characteristic peptide segments, and the corresponding nucleic acid of the subject to make an accurate prediction of the likelihood of the subject having gastric cancer. This method is low-cost and highly efficient.
[0106] In Examples 3-15, the same method as in Example 2 was used to obtain the ROC curves and corresponding AUC results on the training set, the ROC curves and corresponding AUC results of 5-fold cross-validation, the ROC curves and corresponding AUC results on the test set, and related ROC curve graphs, which were used to evaluate the predictive performance of prediction models 2-14. The following will only describe these results without referring to the accompanying figures; identical content will not be repeated.
[0107] Example 3: Establishment and Validation of Logistic Regression Prediction Model 2
[0108] On the training set, a logistic regression model containing two independent variables, O94822 and P62328, was constructed to predict the probability of a sample having gastric cancer. First, based on the quantitative data of O94822 and P62328 and the sample labels, prediction model 2 was obtained using the logistic regression algorithm. The mathematical expression of prediction model 2 is: P = 1 / (1 + Exp(-(β0 + β1 × O94822 + β2 × P62328))), where P represents the probability of the sample being classified as a positive example, the name of the protein biomarker in the formula refers to the standardized content value of that protein biomarker, and the coefficients β0, β1, and β2 range from [-3.37, 2.68], [-1.14, -0.68], to [0.83, 1.37]. Preferably, β0 = -0.34, β1 = -0.91, and β2 = 1.10 yields the optimal prediction result. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the protein or antibody of O94822 and P62328 or its single peptide chain, characteristic peptide segments of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification results.
[0109] The ROC curve of prediction model 2 on the training set shows good overall performance and a high AUC value of 0.85, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the ROC curve for 5-fold cross-validation on the training set for prediction model 2 are 0.84 and 0.04, respectively, achieving the set performance target and demonstrating the model's stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 2 on the test set also shows good overall performance and a high AUC value of 0.87. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is comparable to that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0110] Example 4: Establishment and Validation of Logistic Regression Prediction Model 3
[0111] On the training set, a logistic regression model containing three independent variables, O94822, P62328, and P00915, was constructed to predict the probability of a sample developing gastric cancer. First, based on the quantitative data of O94822, P62328, and P00915 and the sample labels, prediction model 3 was obtained through the logistic regression algorithm. The mathematical expression for prediction model 3 is: P = 1 / (1 + Exp(-(β0 + β1 × O94822 + β2 × P62328 + β3 × P00915))); where P represents the probability of a sample being classified as a positive example, the name of the protein biomarker in the formula refers to the standardized content value of that protein biomarker, and the coefficients β0, β1, β2, and β3 are in the range of [1.50, 9.17], [-1.12, -0.64], [0.83, 1.39], and [-0.90, -0.40]. Preferably, β0 is 5.33, β1 is -0.88, β2 is -0.88, and β3 is -0.65, which yields the optimal prediction result. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of the O94822, P62328, and P00915 proteins or their antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0112] The ROC curve of prediction model 3 on the training set shows good overall performance and a high AUC value of 0.86, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the ROC curve for 5-fold cross-validation on the training set for prediction model 3 are 0.86 and 0.04, respectively, achieving the set performance target and demonstrating the model's stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 3 on the test set also shows good overall performance and a high AUC value of 0.89. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is comparable to that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0113] Example 5: Establishment and Validation of Logistic Regression Prediction Model 4
[0114] On the training set, a logistic regression model containing three independent variables, Q14156, P62328, and Q16280, was constructed to predict the probability of a sample developing gastric cancer. First, based on the quantitative data of O94822, P62328, and P00915 and the sample labels, prediction model 4 was obtained through the logistic regression algorithm. The mathematical expression for prediction model 4 is: P = 1 / (1 + Exp(-(β0 + β1 × Q14156 + β2 × P62328 + β3 × Q16280))), where P represents the probability of a sample being classified as a positive example. In the formula, the name of the protein biomarker refers to its standardized content value. The coefficients β0, β1, β2, and β3 range from [5.36, 15.19] to [-0.89, -0.31], β2, and β3 respectively. Preferably, β0 = 10.28, β1 = -0.60, β2 = 1.13, and β3 = -1.34, yielding the optimal prediction result. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of Q14156, P62328, and Q16280 proteins or their antibodies or their single peptide chains, characteristic peptide segments of peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0115] The ROC curve of prediction model 4 on the training set shows good overall performance and a high AUC value of 0.88, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the ROC curve for 5-fold cross-validation on the training set for prediction model 4 are 0.88 and 0.03, respectively, achieving the set performance target and demonstrating the model's stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 4 on the test set also shows good overall performance and a high AUC value of 0.86. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is comparable to that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0116] Example 6: Establishment and Validation of Logistic Regression Prediction Model 5
[0117] On the training set, a logistic regression model containing three independent variables, P01009, P59666, and Q92824, was constructed to predict the probability of a sample developing gastric cancer. First, based on the quantitative data of O94822, P62328, and P00915, and the sample labels, prediction model 5 was obtained through the logistic regression algorithm. The mathematical expression for prediction model 5 is: P = 1 / (1 + Exp(-(β0 + β1 × P01009 + β2 × P59666 + β3 × Q92824))), where P represents the probability of a sample being classified as a positive example. In the formula, the name of the protein biomarker refers to its standardized content value. The coefficients β0, β1, β2, and β3 range from [-51.02, -33.02], [2.78, 4.53], [0.73, 1.31], to [-0.83, -0.41]. Preferably, β0 = -42.02, β1 = 3.65, β2 = 1.02, and β3 = -0.62, yielding the optimal prediction result. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of P01009, P59666, and Q92824 proteins or their antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification.
[0118] The ROC curve of prediction model 5 on the training set shows good overall performance and a high AUC value of 0.88, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.02 corresponding to the 5-fold cross-validation ROC curve of prediction model 5 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 5 on the test set also shows good overall performance and a high AUC value of 0.89. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is comparable to that on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0119] Example 7: Establishment and Validation of Logistic Regression Prediction Model 6
[0120] On the training set, a logistic regression model containing three independent variables, P01009, P59666, and P02533, was constructed to predict the probability of a sample developing gastric cancer. First, based on the quantitative data of P01009, P59666, and P02533, as well as the sample labels, the prediction model 6 was obtained through the logistic regression algorithm. The mathematical expression for prediction model 6 is: P = 1 / (1 + Exp(-(β0 + β1 × P01009 + β2 × P59666 + β3 × P02533))), where P represents the probability of a sample being classified as a positive example. In the formula, the name of the protein biomarker refers to its standardized content value. The coefficients β0, β1, and β2 range from [-56.40, -38.41], [2.37, 4.00], [0.75, 1.32], to [0.12, 0.56]. Preferably, β0 = -47.41, β1 = 3.18, β2 = 1.03, and β3 = 0.34, yielding the optimal prediction result. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of P01009, P59666, and P02533 proteins or their antibodies or their single peptide chains, characteristic peptide segments of peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification.
[0121] The ROC curve of prediction model 6 on the training set shows good overall performance and a high AUC value of 0.88, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the ROC curve for 5-fold cross-validation on the training set for prediction model 6 are 0.87 and 0.05, respectively, achieving the set performance target and demonstrating the model's stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 6 on the test set also shows good overall performance and a high AUC value of 0.88. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is comparable to that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0122] Example 8: Establishment and Validation of Logistic Regression Prediction Model 7
[0123] On the training set, a logistic regression model containing five independent variables—P01009, P59666, P00915, Q14156, and Q9UBF2—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of P01009, P59666, P00915, Q14156, and Q9UBF2, as well as the sample labels, the predictive model 7 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 7 is: P=1 / (1+Exp(-(β0+β1 ×P01009 + β2 × P59666 + β3 × P00915+β4 × Q14156+ β5 × Q9UBF2))), where P represents the probability of a sample being classified as a positive example. In the formula, the name of the protein biomarker refers to the standardized content value of that protein biomarker. The coefficients β0, β1, β2, and β5 range from [-51.68, -29.98] to [2.88, 4.80], [0.96, 1.62], [-1.36, -0.71], [-1.22, -0.50], and [0.33, 0.95]. Preferably, the optimal prediction results are obtained when β0 is -40.83, β1 is 3.84, β2 is 1.29, β3 is -1.03, β4 is -0.86, and β5 is 0.64. The established logistic regression model can predict the likelihood of gastric cancer in new samples based on the P01009, P59666, P00915, Q14156, and Q9UBF2 proteins or their antibodies or their single peptide chains, characteristic peptide segments of the peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification results.
[0124] The ROC curve of prediction model 7 on the training set shows good overall performance and a high AUC value of 0.93, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.01 corresponding to the 5-fold cross-validation ROC curve of prediction model 7 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 7 on the test set also shows good overall performance and a high AUC value of 0.91. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0125] Example 9: Establishment and Validation of Logistic Regression Prediction Model 8
[0126] On the training set, a logistic regression model containing seven independent variables—P62328, Q14156, Q16280, P59666, P01009, Q9NQX4, and Q6ZVL6—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of P62328, Q14156, Q16280, P59666, P01009, Q9NQX4, and Q6ZVL6, as well as the sample labels, the predictive model 8 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 8 is: P = 1 / (1 + Exp(-(β0 + β1 × P62328 + β2 × Q14156 + β3 × Q16280 + β4 × P59666 + β5 × P01009 + β6 × Q9NQX4 + β7 × Q6ZVL6))), where P represents the probability of a sample being classified as a positive example. In the formula, the name of the protein biomarker refers to its standardized content value. The coefficients β0, β1, β2, and β3 range from [-38.76, -14.57] to [0.45, 1.09], β2, β3, and β4 range from [-1.16, -0.41] to [-1.63, -0.70], respectively. [1.15]; the value range of coefficient β5 is [1.87, 3.97]; the value range of coefficient β6 is [0.59, 1.56]; the value range of coefficient β7 is [-1.04, -0.36]. Preferably, the optimal prediction result can be obtained when β0 is -26.66, β1 is 0.77, β2 is -0.79, β3 is -1.16, β4 is 0.78, β5 is 2.92, β6 is 1.08, and β7 is -0.70. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of P62328, Q14156, Q16280, P59666, P01009, Q9NQX4 and Q6ZVL6 proteins or their antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0127] The ROC curve of prediction model 8 on the training set shows good overall performance and a high AUC value of 0.94, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.01 corresponding to the 5-fold cross-validation ROC curve of prediction model 8 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 8 on the test set also shows good overall performance and a high AUC value of 0.93. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is comparable to that on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0128] Example 10: Establishment and Validation of Logistic Regression Prediction Model 9
[0129] On the training set, a logistic regression model containing eight independent variables—P62328, Q14156, Q16280, P59666, P01009, Q9NQX4, Q6ZVL6, and P00915—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of P62328, Q14156, Q16280, P59666, P01009, Q9NQX4, Q6ZVL6, and P00915, as well as the sample labels, the prediction model 9 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 9 is: P = 1 / (1 + Exp(-(β0 + β1 × P62328 + β2 × Q14156 + β3 × Q16280 + β4 × P59666 + β5 × P01009 + β6 × Q9NQX4 + β7 × Q6ZVL6 + β8 × P00915))), where P represents the probability of a sample being classified as a positive example. In the formula, the name of the protein biomarker refers to its standardized content value. The coefficients β0, β1, β2, and β3 range from [-34.50, -8.76] to [0.42, 1.09], β2, β3, and β4 range from [-1.71, -0.72] to [0.52, 1.09]. [1.31]; the value range of coefficient β5 is [2.12, 4.37]; the value range of coefficient β6 is [2.12, 4.37]; the value range of coefficient β7 is [0.49, 1.51]; the value range of coefficient β8 is [-1.35, -0.64]. Preferably, the optimal prediction result can be obtained when β0 is -21.63, β1 is 0.75, β2 is -0.78, β3 is -1.21, β4 is 0.92, β5 is 3.24, β6 is 1.00, β7 is -0.65, and β8 is -0.99. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of the proteins P62328, Q14156, Q16280, P59666, P01009, Q9NQX4, Q6ZVL6, and P00915, or their antibodies, or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0130] The ROC curve of prediction model 9 on the training set shows good overall performance with a high AUC value of 0.95, indicating good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.01 corresponding to the 5-fold cross-validation ROC curve of prediction model 9 on the training set meet the performance target, demonstrating model stability and effectiveness, and providing assurance for practical application. The ROC curve of prediction model 9 on the test set also shows good overall performance with a high AUC value of 0.93. This indicates that the model maintains good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and suitable for clinical validation and application.
[0131] Example 11: Establishment and Validation of Logistic Regression Prediction Model 10
[0132] On the training set, a logistic regression model containing eleven independent variables—P62328, O94822, Q93034, P59666, P41235, P02100, P69905, P04275, P02533, Q16280, and O15553—was constructed to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of P62328, O94822, Q93034, P59666, P41235, P02100, P69905, P04275, P02533, Q16280, and O15553, as well as the sample labels, a prediction model 10 was obtained using the logistic regression algorithm. The mathematical expression for prediction model 10 is: P = 1 / (1 + Exp(-(β0 + β1 × P62328 + β2 × O94822 + β3 × Q93034 + β4 × P59666 + β5 × P41235 + β6 × P02100 + β7 × P69905 + β8 × P04275 + β9 × P02533 + β10 × Q16280 + β11 × O15553))), where P represents the probability of a sample being classified as a positive example. In the formula, the names of the protein biomarker or its antigen / antibody or its single peptide chain, characteristic peptide segments of the peptide chain, and its stable isotope protein or stable isotope characteristic peptide segments refer to the standardized content value of the corresponding nucleic acid for that protein biomarker or its antigen / antibody or its single peptide chain, characteristic peptide segments of the peptide chain, and its stable isotope protein or stable isotope characteristic peptide segments. The coefficients β0, β1, β2, β3, β4, and β5 range from [-19.30, 2.94] to [0.40, 1.03], [-0.89, -0.27], [-0.19, 0.94], [0.11, 0.97], and [0.27, 1.34] respectively. The coefficient β6 ranges from [-1.09, 0.13]; the coefficient β7 ranges from [-1.18, -0.08]; the coefficient β8 ranges from [0.38, 1.28]; the coefficient β9 ranges from [0.01, 0.59]; the coefficient β10 ranges from [-1.77, -0.78]; and the coefficient β11 ranges from [-0.47, 0.73]. Preferably, the optimal prediction result can be obtained when β0 is -8.18, β1 is 0.72, β2 is -0.58, β3 is 0.38, β4 is 0.54, β5 is 0.81, β6 is -0.48, β7 is -0.63, β8 is 0.83, β9 is 0.30, β10 is -1.27, and β11 is 0.13.The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of the following proteins in a new sample: P62328, O94822, Q93034, P59666, P41235, P02100, P69905, P04275, P02533, Q16280, and O15553, or their antibodies, or their single peptide chains, characteristic peptide segments of the peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0133] The ROC curve of prediction model 10 on the training set shows good overall performance with a high AUC value of 0.95, indicating good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.94 for the 5-fold cross-validation ROC curve of prediction model 10 on the training set meet the performance target, demonstrating model stability and effectiveness, and providing assurance for practical application. The ROC curve of prediction model 10 on the test set also shows good overall performance with a high AUC value of 0.93. This indicates that the model maintains good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model demonstrates good predictive performance on the test set and is suitable for clinical validation and application.
[0134] Example 12: Establishment and Validation of Logistic Regression Prediction Model 11
[0135] On the training set, a logistic regression model containing 26 independent variables, namely P01009, Q9H0G5, Q16280, Q14156, P59666, Q9NQX4, P62328, P01242, P69905, P41235, O15553, Q6ZVL6, P04275, A0A075B6K4, Q9UBF2, P02100, Q8TA94, Q93034, Q92824, P00915, Q6IC98, O94822, P02533, P05452, P02538, and Q8IXQ9, was constructed to predict the probability of a sample having gastric cancer. First, based on the quantitative data and sample labels of P01009, Q9H0G5, Q16280, Q14156, P59666, Q9NQX4, P62328, P01242, P69905, P41235, O15553, Q6ZVL6, P04275, A0A075B6K4, Q9UBF2, P02100, Q8TA94, Q93034, Q92824, P00915, Q6IC98, O94822, P02533, P05452, P02538, and Q8IXQ9, a prediction model 11 was obtained using a logistic regression algorithm. The mathematical expression of prediction model 11 is: P=1 / (1+Exp(-(β0+β1 ×P01009 + β2 ×Q9H0G5 + β3 ×Q16280+ β4 × Q14156+ β5 × P59666+ β6 ×Q9NQX4+ β7 ×P62328+ β8 ×P01242+ β9×P69905+ β10 × P41235+ β11 × O15553+β12 ×Q6ZVL6+ β13 × P04275 + β14 ×A0A075B6K4+ β15 × Q9UBF2+ β16 × P02100+ β17 ×Q8TA94+ β18 ×Q93034+ β19 ×Q92824+ β20 ×P00915+ β21 × Q6IC98+ β22 × O94822+β23 ×P02533+ β24 ×P05452 + β25 × P02538+ β26 × Q8IXQ9)))), where P represents the probability of a sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments refer to the standardized content value of the corresponding nucleic acid of the protein biomarker or its antigen / antibody or its single peptide chain, characteristic peptide segments of peptide chains, and stable isotope proteins or stable isotope characteristic peptide segments. The coefficient β0 ranges from [-35.05, 8].
[05] ; the value range of β1 is [0.63, 4.13]; the value range of coefficient β2 is [0.14, 2.30]; the value range of coefficient β3 is [-1.81, -0.52]; the value range of coefficient β4 is [-1.61, -0.55]; the value range of coefficient β5 is [0.28, 1.38]; the value range of coefficient β6 is [0.20, 1.43]; the value range of coefficient β7 is [0.40, 1.17]; the value range of coefficient β8 is [-1.61, 0.06]; the value range of coefficient β9 is [-1.83, 0.62]; the value range of coefficient β10 is [-0.27, 1.36] The coefficients β11, β12, β13, β14, β15, β16, β17, β18, β19, β20, and β21 are all within the range of -1.16 and -0.60 respectively. The coefficients β11, β12, β13, β14, β15, β16, β17, β18, β19, β11, and β21 are all within the range of -1.28 and -0.26 respectively. [0.86]; the value range of coefficient β22 is [-0.70, 0.16]; the value range of coefficient β23 is [-0.25, 0.68]; the value range of coefficient β24 is [-1.35, 0.92]; the value range of coefficient β25 is [-0.27, 0.55]; the value range of coefficient β26 is [-1.07, 1.21]. Preferably, β0 is -13.50, β1 is 2.38, β2 is 1.22, β3 is -1.16, β4 is -1.08, β5 is 0.83, β6 is 0.82, β7 is 0.78, β8 is -0.77, β9 is -0.61, β10 is 0.54, β11 is -0.51, β12 is -0.48, and β13 is 0. 0.46, β14 is -0.42, β15 is 0.42, β16 is -0.37, β17 is -0.37, β18 is 0.32, β19 is -0.32, β20 is -0.28, β21 is 0.27, β22 is -0.27, β23 is 0.22, β24 is -0.21, β25 is 0.14, and β26 is 0.The optimal prediction result can be obtained at 07. The established logistic regression model can predict the likelihood of gastric cancer in new samples based on the following parameters: P01009, Q9H0G5, Q16280, Q14156, P59666, Q9NQX4, P62328, P01242, P69905, P41235, O15553, Q6ZVL6, P04275, A0A075B6K4, Q9UBF2, P02100, Q8TA94, Q93034, Q92824, P00915, Q6IC98, O94822, P02533, P05452, P02538, and Q8IXQ9 proteins, their antigens / antibodies, their single peptide chains, characteristic peptide segments of peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid quantification results.
[0136] The ROC curve of prediction model 11 on the training set shows good overall performance and a high AUC value of 0.97, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of 0.01 corresponding to the 5-fold cross-validation ROC curve of prediction model 11 on the training set meet the set performance target, demonstrating the model's stability and effectiveness, and providing a guarantee for its practical application. The ROC curve of prediction model 11 on the test set also shows good overall performance and a high AUC value of 0.95. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and the AUC on the test set is not significantly different from that on the training set, indicating that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0137] Example 13: Establishment and Validation of Logistic Regression Prediction Model 12
[0138] On the training set, a dataset containing Q68DV7, P41235, P02748, P01009, O75152, Q06033, P02750, P02763, P01011, Q9H0G5, P60709, P08575, Q12851, P06727, Q9H6K5, O95236, P00915, and P... A logistic regression model with 32 independent variables (69905, Q9UFC0, P02100, P68871, Q8IXQ9, P62906, Q5JY77, A0PJY2, O15553, O94822, P04275, P59666, Q16280, Q8TA94, and Q92824) was used to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data and sample labels of Q68DV7, P41235, P02748, P01009, O75152, Q06033, P02750, P02763, P01011, Q9H0G5, P60709, P08575, Q12851, P06727, Q9H6K5, O95236, P00915, P69905, Q9UFC0, P02100, P68871, Q8IXQ9, P62906, Q5JY77, A0PJY2, O15553, O94822, P04275, P59666, Q16280, Q8TA94, and Q92824, a prediction model 12 was obtained using a logistic regression algorithm. The mathematical expression of prediction model 12 is: P=1 / (1+Exp(-(β0+β1 ×Q68DV7 + β2 ×P41235 + β3 × P02748+ β4 × P01009+ β5 × O75152+ β6 ×Q06033+ β7 ×P02750+ β8 ×P02763+ β9 ×P01011+ β10 × Q9H0G5+ β11 × P60709 + β12 ×P08575+ β13 × Q12851 + β14 ×P06727+ β15 × Q9H6K5+ β16 × O95236+ β17 ×P00915+ β18 × P69905+ β19 ×Q9UFC0+ β20 ×P02100+ β21 × P68871+ β22 × Q8IXQ9 +β23 ×P62906+ β24 ×Q5JY77 + β25 × A0PJY2+ β26 × O15553+ β27 × O94822+ β28 × P04275 + β29 ×P59666+ β30 × Q16280 + β31 × Q8TA94 + β32 × Q92824)))), where,P represents the probability of a sample being classified as a positive example. In the formula, the names of the protein biomarker, its antigen / antibody, its single peptide chain, characteristic peptide segments of the peptide chain, and its stable isotope protein or stable isotope characteristic peptide segments represent the standardized content value of the corresponding nucleic acid for that protein biomarker, its antigen / antibody, its single peptide chain, characteristic peptide segments of the peptide chain, its stable isotope protein or stable isotope characteristic peptide segments, and the corresponding nucleic acid. The coefficients β0, β1, β2, β3, β4, β5, and β6 range from -40.87 to -2.20. 1.03]; The value range of coefficient β7 is [-0.86, 1.94]; The value range of coefficient β8 is [-0.51, 1.62]; The value range of coefficient β9 is [-1.58, 1.26]; The value range of coefficient β10 is [-0.17, 2.12]; The value range of coefficient β11 is [0.88, 2.69]; The value range of coefficient β12 is [-0.54, 2.32]; The value range of coefficient β13 is [-0.16, 0.90]; The value range of coefficient β14 is [-1.43, 3.58]; The value range of coefficient β15 is [-2.23, 1.81]; The value range of coefficient β16 is [-1.48, 1.04]; The value range of coefficient β17 is [-0.82, 0.96]; the value range of coefficient β18 is [-2.74, 1.06]; the value range of coefficient β19 is [-1.26, 1.50]; the value range of coefficient β20 is [-1.02, 0.36]; the value range of coefficient β21 is [-3.98, 1.00]; the value range of coefficient β22 is [-0.53, 2.15]; the value range of coefficient β23 is [-0.59, 0.64]; the value range of coefficient β24 is [-0.45, 0.39]; the value range of coefficient β25 is [-0.61, 0.74]; the value range of coefficient β26 is [-1.12, 0.42]; the value range of coefficient β27 is [-1.15, [0.03]; the value range of coefficient β28 is [-0.08, 1.01]; the value range of coefficient β29 is [-0.11, 0.92]; the value range of coefficient β30 is [-3.39, -0.08]; the value range of coefficient β31 is [-1.31, 0.06]; the value range of coefficient β32 is [-0.95, 0.04]. Preferably, β0 is -21.53, β1 is -0.58, β2 is 0.51, β3 is -1.17, β4 is 2.12, and β5 is 0.69.β6 is 0.05, β7 is 0.54, β8 is 0.56, β9 is -0.16, β10 is 0.97, β11 is 1.78, β12 is 0.89, β13 is 0.37, β14 is 1.07, β15 is -0.21, β16 is -0.22, β17 is 0.07, β18 is -0.84, β19 is 0.12, and β20 is... The optimal prediction results are obtained with β21 = -1.49, β22 = 0.81, β23 = 0.02, β24 = -0.03, β25 = 0.06, β26 = -0.35, β27 = -0.56, β28 = 0.47, β29 = 0.41, β30 = -1.74, β31 = -0.62, and β32 = -0.46. The established logistic regression model can be based on the new sample's Q68DV7, P41235, P02748, P01009, O75152, Q06033, P02750, P02763, P01011, Q9H0G5, P60709, P08575, Q12851, P06727, Q9H6K5, O95236, P00915, P69905, Q9UFC0, P02100 The likelihood of a new sample developing gastric cancer is predicted by using the following proteins: P68871, Q8IXQ9, P62906, Q5JY77, A0PJY2, O15553, O94822, P04275, P59666, Q16280, Q8TA94, and Q92824, or their antigens / antibodies, or their single peptide chains, characteristic peptide segments, and stable isotope proteins or stable isotope characteristic peptide segments, along with the corresponding nucleic acid quantification results.
[0139] The ROC curve of prediction model 12 on the training set shows good overall performance and a high AUC value of 0.97, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the ROC curve for 5-fold cross-validation on the training set for prediction model 12 are 0.96 and 0.01, respectively, achieving the set performance target and demonstrating the model's stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 12 on the test set also shows good overall performance and a high AUC value of 0.93. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0140] Example 14: Establishment and Validation of Logistic Regression Prediction Model 13
[0141] On the training set, a dataset containing P62328, O94822, P01009, P02748, Q93034, O75152, P59666, P02750, P68871, P62906, Q8IXQ9, P41235, P02538, Q06033, P05452, Q8TA94, and Q9UN3 was constructed. 7. P53367, P02100, P69905, Q9H0G5, Q14156, P04275, Q9NQX4, P02533, Q9UHG0, Q16280, Q6ZVL6, P00915, O15553, Q9UBF2, P01011, A0A075B6K4, Q92824, P35527 , Q15047, P01242, P02671, P60709, Q8TEW8, P18065, O75683, Q8IVV2, Q9UBP8, P 0DJI8, Q15293, Q68DV7, A6NFD8, P05109, Q09666, P00738, Q5VU43, Q9UPX8, Q9Y A logistic regression model with 68 independent variables (5E1, P59923, Q4G0S7, Q9UKN7, P18428, Q4VXF1, Q66K14, Q9NQV5, P02042, Q9H5Y7, Q8IZF2, O95490, Q2TBA0, Q9P0M6, and O94885) was used to predict the likelihood of a sample developing gastric cancer. First,Based on P62328, O94822, P01009, P02748, Q93034, O75152, P59666, P02750, P68871, P62906, Q8IXQ9, P41235, P02538, Q06033, P05452, Q8TA94, Q9UN37, P53367, P02100, P69905, Q9H0G5, Q14156, P04275, Q9NQX4, P02533, Q9UHG0, Q16280, Q6ZVL6, P00915, O15553, Q9UBF2, P01011, A0A075B6K4, Q92824, P35 Quantitative data and sample labels for samples 527, Q15047, P01242, P02671, P60709, Q8TEW8, P18065, O75683, Q8IVV2, Q9UBP8, P0DJI8, Q15293, Q68DV7, A6NFD8, P05109, Q09666, P00738, Q5VU43, Q9UPX8, Q9Y5E1, P59923, Q4G0S7, Q9UKN7, P18428, Q4VXF1, Q66K14, Q9NQV5, P02042, Q9H5Y7, Q8IZF2, O95490, Q2TBA0, Q9P0M6, and O94885.Predictive model 13 was obtained using the logistic regression algorithm. The mathematical expression of prediction model 13 is: P=1 / (1+Exp(-(β0+ β1 ×P62328 + β2 ×O94822 + β3 × P01009+ β4 × P02748+ β5× Q93034+ β6 ×O75152+ β7 ×P59666+ β8 ×P02750+ β9 ×P68871+ β10 × P62906+β11 × Q8IXQ9 + β12 ×P41235+ β13 × P02538 + β14 ×Q06033+ β15 × P05452+ β16 × Q8TA94+ β17 ×Q9UN37+ β18 × P53367+ β19 ×P02100+ β20 ×P69905+ β21 ×Q9H0G5+ β22 × Q14156 +β23 ×P04275+ β24 × Q9NQX4 + β25 × P02533+ β26 ×Q9UHG0+ β27 × Q16280+ β28 × Q6ZVL6 + β29 ×P00915+ β30 × O15553 + β31 ×Q9UBF2 + β32 × P01011+β33 ×A0A075B6K4 + β34 ×Q92824 + β35 × P35527+ β36× Q15047+ β37 × P01242+ β38 ×P02671+ β39 ×P60709+ β40 ×Q8TEW8+ β41 ×P18065+ β42× O75683+ β43 × Q8IVV2+ β44 ×Q9UBP8+ β45× P0DJI8+ β46 ×Q15293+ β47 × Q68DV7+ β48 × A6NFD8+ β49×P05109+ β50 × Q09666+ β51×P00738+ β52×Q5VU43+ β53 × Q9UPX8+ β54 × Q9Y5E1+β55 ×P59923+ β56 × Q4G0S7+ β57 ×Q9UKN7+ β58 × P18428+ β59 × Q4VXF1+ β60 × Q66K14+ β61 ×Q9NQV5+ β62 ×P02042+ β63 × Q9H5Y7+ β64 × Q8IZF2+ β65 × O95490+ β66 ×Q2TBA0+ β67 ×Q9P0M6+ β68 × O94885)))), where P represents the probability that the sample is classified as a positive example.In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments refer to the standardized content values of the corresponding nucleic acids of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and stable isotopic proteins or stable isotopic characteristic peptide segments. The range and preferred values of the coefficients β0 to β68 are shown in Table 1. The optimal prediction results can be obtained based on the preferred values. The established logistic regression model can be based on the new samples P62328, O94822, P01009, P02748, Q93034, O75152, P59666, P02750, P68871, P62906, Q8IXQ9, P41235, P02538, Q06033, P05452, Q8TA94, Q9UN37, P5336 7. P02100, P69905, Q9H0G5, Q14156, P04275, Q9NQX4, P02533, Q9UHG0, Q16280, Q6ZVL6 , P00915, O15553, Q9UBF2, P01011, A0A075B6K4, Q92824, P35527, Q15047, P01242, P026 71. P60709, Q8TEW8, P18065, O75683, Q8IVV2, Q9UBP8, P0DJI8, Q15293, Q68DV7, A6NFD 8. P05109, Q09666, P00738, Q5VU43, Q9UPX8, Q9Y5E1, P59923, Q4G0S7, Q9UKN7, P18428, The likelihood of a new sample developing gastric cancer is predicted based on the quantitative results of the following proteins: Q4VXF1, Q66K14, Q9NQV5, P02042, Q9H5Y7, Q8IZF2, O95490, Q2TBA0, Q9P0M6, and O94885, or their antigens / antibodies, or their single peptide chains, characteristic peptide segments, and stable isotope proteins or stable isotope characteristic peptide segments.
[0142] Table 1. Range and preferred values of coefficients for prediction model 13
[0143] The ROC curve of prediction model 13 on the training set shows good overall performance and a high AUC value of 0.99, indicating that the model has good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the ROC curve for 5-fold cross-validation on the training set for prediction model 13 are 0.98 and 0.01, respectively, achieving the set performance target and demonstrating the model's stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 13 on the test set also shows good overall performance and a high AUC value of 0.95. This indicates that the model maintains a good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set, and that the AUC on the test set is not significantly different from that on the training set, suggesting that the model is not overfitting and has good generalization ability. In summary, the model's predictive performance on the test set is good and it is suitable for clinical validation and application.
[0144] Example 15: Establishment and Validation of Logistic Regression Prediction Model 14
[0145] On the training set, a dataset containing A0A075B6K4, A0PJY2, A6NFD8, A6QL64, O15033, O15553, O75152, O75460, O75683, O94822, O94885, O95236, O95490, P00738, P00751, P00915, P01009, P01011, P01185, P01242, P01701, P02042, and P... 02100, P02533, P02538, P02671, P02748, P02750, P02763, P04275, P05109, P05452, P06727, P08571, P08 575, P0DJI8, P18065, P18428, P31150, P35527, P41235, P53367, P59666, P59923, P60709, P62328, P62906 , P68871, P69891, P69905, Q01538, Q06033, Q09666, Q12851, Q14156, Q15047, Q15293, Q16280, Q2TBA0, Q 4G0S7, Q4VXF1, Q53HC9, Q5H9J7, Q5JY77, Q5VU43, Q66K14, Q68DV7, Q6ZVL6, Q8IVV2, Q8IXQ9, Q8IZF2, Q8T A logistic regression model with 92 independent variables (A94, Q8TEW8, Q92797, Q92824, Q93034, Q96N87, Q96RV3, Q9H0G5, Q9H5Y7, Q9H6K5, Q9NQV5, Q9NQX4, Q9P0M6, Q9UBF2, Q9UBP8, Q9UFC0, Q9UHG0, Q9UKN7, Q9UN37, Q9UPX8, and Q9Y5E1) was used to predict the likelihood of a sample developing gastric cancer. First, based on the quantitative data of the 92 proteins and the sample labels, the predictive model 14 was obtained using the logistic regression algorithm.The mathematical expression for prediction model 14 is: P = 1 / (1 + Exp(-(β0 + β1 × A0A075B6K4 + β2 × A0PJY2 + β3 × A6NFD8 + … + β88 × Q9UHG0 + β89 × Q9UKN7 + β90 × Q9UN37 + β91 × Q9UPX8 + β92 × Q9Y5E1)))), where P represents the probability of a sample being classified as a positive example. In the formula, the names of protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments refer to the standardized content values of the corresponding nucleic acids of the protein biomarkers or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments. The range and preferred values of the coefficients β0 to β92 are shown in Table 2. The optimal prediction result can be obtained based on the preferred values. The established logistic regression model can predict the likelihood of a new sample having gastric cancer based on the results of quantification of these 92 proteins or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, their stable isotope proteins or stable isotope characteristic peptide segments, and the corresponding nucleic acid.
[0146] Table 2. Range and preferred values of coefficients for prediction model 14
[0147] The ROC curve of prediction model 14 on the training set shows good overall performance with a high AUC value of 1.00, indicating good predictive performance in distinguishing between gastric cancer and non-gastric cancer samples. The mean AUC and standard deviation of the 5-fold cross-validation ROC curve of prediction model 14 on the training set are 0.99 and 0.00, respectively, achieving the set performance target and demonstrating model stability and effectiveness, thus ensuring its practical application. The ROC curve of prediction model 14 on the test set also shows good overall performance with a high AUC value of 0.91. This indicates that the model maintains good ability to distinguish between gastric cancer and non-gastric cancer samples on the test set. However, the difference of 0.09 between the test set AUC and the training set AUC suggests overfitting. This is due to strong correlations among multiple proteins in the 92 protein markers, indicating collinearity and slightly lower generalization ability, although the performance on the test set still meets the target. In summary, the model's predictive performance on the test set is good and suitable for clinical validation and application.
[0148] The above models involve different indicators and can be applied to multiple different classification groups obtained based on the indicators and quantities. Preferably, prediction models 1 and 2 can obtain accurate prediction results using a minimum of two indicators, which is simple and low-cost.
[0149] Figure 6 This is a system block diagram illustrating an embodiment of this application for assessing the likelihood of a subject having gastric cancer based on the Swiss-prot database. (Reference) Figure 6 As shown, the system in this embodiment includes a data acquisition module 610, a data preprocessing module 620, a model recommendation module 630, a model selection module 640, and a risk assessment module 650.
[0150] The data acquisition module 610 is used to acquire sample data from the subjects, including A0A075B6K4, A0PJY2, A6NFD8, A6QL64, O15033, O15553, O75152, O75460, O75683, O94822, O94885, O95236, O95490, P00738, P00751, P00915, P01009, P01011, P01185, P01242, and P01701. P02042, P02100, P02533, P02538, P02671, P02748, P02750, P02763, P04275, P05109, P05452, P06727, P08571, P 08575, P0DJI8, P18065, P18428, P31150, P35527, P41235, P53367, P59666, P59923, P60709, P62328, P62906, P6 8871, P69891, P69905, Q01538, Q06033, Q09666, Q12851, Q14156, Q15047, Q15293, Q16280, Q2TBA0, Q4G0S7, Q4V XF1, Q53HC9, Q5H9J7, Q5JY77, Q5VU43, Q66K14, Q68DV7, Q6ZVL6, Q8IVV2, Q8IXQ9, Q8IZF2, Q8TA94, Q8TEW8, Q927 97, Q92824, Q93034, Q96N87, Q96RV3, Q9H0G5, Q9H5Y7, Q9H6K5, Q9NQV5, Q9NQX4, Q9P0M6, Q9UBF2, Q9UBP8, Q9UFC0, Q9UHG0, Q9UKN7, Q9UN37, Q9UPX8 and Q9Y5E1 proteins or their antigens / antibodies or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotope proteins or stable isotope characteristic peptide segments, any one of the quantitative values of these proteins in serum.
[0151] The data preprocessing module 620 is used to preprocess the sample data, including removing duplicate samples, filling missing values, and dividing the sample data into different analysis groups.
[0152] The model recommendation module 630 is used to recommend multiple predictive models with larger AUC values from the predictive models built using the construction method described above, based on the types and number of indicators included in different analysis groups. The model formula with the largest AUC value is given priority, and the final result is displayed on the user interface as a list of model formulas, sorted from largest to smallest AUC value.
[0153] The model selection module 640 provides users with a selection function to output one or more predictive models from the predictive models recommended by the model recommendation module 630 for risk assessment and calculation. Users can choose not to make a selection, and the system will default to selecting the model with the highest AUC value.
[0154] The risk assessment module 650 calculates the corresponding predicted value based on one or more predictive models output by the model selection module 640, and provides an appropriate Logit threshold or risk score threshold according to the department or population information corresponding to the sample data. It then classifies the risk level as any one of high probability, medium probability, or low probability based on the Logit threshold or risk score threshold. In some embodiments, the Logit threshold or risk score threshold includes one threshold for binary classification, i.e., dividing into high probability and low probability. In some embodiments, the Logit threshold or risk score threshold includes two thresholds for trigonometric classification, i.e., dividing into high probability, medium probability, and low probability. In other embodiments, further subdivisions can be made, which are not limited in this application.
[0155] System 600 can quantitatively obtain biomarker information from subject samples and input it into a model for gastric cancer risk assessment, avoiding errors in subjective judgment and achieving effective assessment and risk prediction of the likelihood of developing gastric cancer in different populations. Unlike existing gastric cancer screening systems, System 600 can not only quantitatively, automatically, and continuously provide gastric cancer risk levels, improving the accuracy and efficiency of assessment; it also provides a rich set of model formulas containing different types of indicators. Considering that some samples may not have the most complete set of indicators in practical applications, System 600 can also provide the most suitable prediction model in such cases. In addition, System 600 also provides a model recommendation module 630 and a model selection module 640, making the system more flexible and meeting diverse application needs of customers. Furthermore, System 600 can apply the most appropriate threshold for risk assessment based on the department or population from which the sample originates.
[0156] This application also proposes a kit for assessing the likelihood of a subject developing gastric cancer based on the Swiss-prot database, including a predictive model established by the model construction method described above. Specifically, this predictive model can be integrated into an integrated circuit within the kit. The kit can receive any one or more of the effective protein biomarkers described above and calculate the prediction result based on the integrated predictive model, providing a probability result of the subject developing gastric cancer. Taking predictive model 1 as an example, the kit can receive, for example, O94822 and P62328, and calculate the prediction result based on the built-in predictive model 1 to provide a probability result of the subject developing gastric cancer. The kit can simultaneously include the 41 different predictive models 1 described above to be applicable to different populations, exhibiting very broad applicability.
[0157] This application also proposes a kit for assessing the likelihood of a subject developing gastric cancer based on the Swiss-prot database, used to receive effective protein biomarkers, which are any one or more of the following protein biomarkers: A0A075B6K4, A0PJY2, A6NFD8, A6QL64, O15033, O15553, O75152, O75460, O75683, O94822, O94885, O95236, O95490, P00738, P00751, P0 0915, P01009, P01011, P01185, P01242, P01701, P02042, P02100, P02533, P02538, P02671, P02748, P02750, P02763, P04275, P05109, P05452, P06727, P08571, P08575, P0DJI8, P18065, P18428, P31150, P35527, P41235, P53367, P59666 , P59923, P60709, P62328, P62906, P68871, P69891, P69905, Q01538, Q06033, Q09666, Q12851, Q14156, Q15047, Q152 93. Q16280, Q2TBA0, Q4G0S7, Q4VXF1, Q53HC9, Q5H9J7, Q5JY77, Q5VU43, Q66K14, Q68DV7, Q6ZVL6, Q8IVV2, Q8IXQ9, Q8 IZF2, Q8TA94, Q8TEW8, Q92797, Q92824, Q93034, Q96N87, Q96RV3, Q9H0G5, Q9H5Y7, Q9H6K5, Q9NQV5, Q9NQX4, Q9P0M6, Q9UBF2, Q9UBP8, Q9UFC0, Q9UHG0, Q9UKN7, Q9UN37, Q9UPX8, Q9Y5E1, or a single peptide chain, a characteristic peptide segment of the peptide chain, and its stable isotope protein or stable isotope characteristic peptide segment, and the corresponding nucleic acid.
[0158] In some embodiments, the effective protein biomarkers that the kit is used to receive are effective protein biomarkers included in any of the prediction models described above.
[0159] This application also includes an electronic device comprising a memory and a processor. The memory stores instructions executable by the processor; the processor executes the instructions to implement the aforementioned method for constructing a model based on the Swiss-prot database to assess the likelihood of a subject developing gastric cancer.
[0160] Figure 7 This is a system block diagram of an electronic device according to an embodiment of this application. (Reference) Figure 7 As shown, the electronic device 700 may include an internal communication bus 701, a processor 702, a read-only memory (ROM) 703, a random access memory (RAM) 704, and a communication port 705. When applied to a personal computer, the electronic device 700 may also include a hard disk 706. The internal communication bus 701 enables data communication between the components of the electronic device 700. The processor 702 can make judgments and issue prompts. In some embodiments, the processor 702 may consist of one or more processors. The communication port 705 enables data communication between the electronic device 700 and external devices. In some embodiments, the electronic device 700 can send and receive information and data from a network through the communication port 705. The electronic device 700 may also include different forms of program storage units and data storage units, such as the hard disk 706, the read-only memory (ROM) 703, and the random access memory (RAM) 704, capable of storing various data files used for computer processing and / or communication, as well as possible program instructions executed by the processor 702. The processor executes these instructions to implement the main part of the method. The results processed by the processor are transmitted to the user device through the communication port and displayed on the user interface.
[0161] The above-described model construction method can be implemented as a computer program, stored in the hard disk 706, and loaded into the processor 702 for execution to implement the model construction method of this application.
[0162] This application also includes a computer-readable medium storing computer program code that, when executed by a processor, implements the method for constructing the model described above.
[0163] When a method for constructing a model based on the Swiss-prot database to assess the likelihood of a subject developing gastric cancer is implemented as a computer program, it can also be stored as an article of manufacture in a computer-readable storage medium. For example, computer-readable storage media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic stripes), optical discs (e.g., compact discs (CDs), digital multifunction discs (DVDs)), smart cards, and flash memory devices (e.g., electrically erasable programmable read-only memory (EPROM), cards, sticks, key drives). Furthermore, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media (and / or storage media) capable of storing, containing, and / or carrying code and / or instructions and / or data.
[0164] It should be understood that the embodiments described above are merely illustrative. The embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or any combination thereof. For hardware implementation, the processor may be implemented within one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, and / or other electronic units designed to perform the functions described herein, or combinations thereof.
[0165] Some aspects of this application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The aforementioned hardware or software may be referred to as a "data block," "module," "engine," "unit," "component," or "system." The processor may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, or combinations thereof. Furthermore, aspects of this application may manifest as computer products residing in one or more computer-readable media, including computer-readable program code. For example, computer-readable media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic tapes, etc.), optical discs (e.g., compressed CDs, digital multifunction DVDs, etc.), smart cards, and flash memory devices (e.g., cards, sticks, key drives, etc.).
[0166] A computer-readable medium may contain a propagated data signal containing computer program code, for example, on baseband or as part of a carrier wave. This propagated signal may take various forms, including electromagnetic, optical, and so on, or suitable combinations thereof. A computer-readable medium can be any computer-readable medium other than a computer-readable storage medium, which can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. The program code located on the computer-readable medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, radio frequency signals, or similar media, or any combination of the above media.
[0167] This application uses specific terms to describe embodiments of the application. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of the application. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics in one or more embodiments of the application can be appropriately combined.
[0168] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of scope in some embodiments of this application are approximate values, in specific embodiments, such values are set as precisely as feasible.
Claims
1. A method for constructing a model based on the Swiss-prot database to assess the probability of a subject developing gastric cancer, comprising: Obtain samples from the subjects, including multiple serum samples; Mass spectrometry data analysis includes: performing mass spectrometry data analysis on the multiple serum samples to extract peptide information and obtain sample data; Database retrieval includes: performing a Swiss-prot database search on the sample data to identify proteins and obtain the original protein quantification matrix; Anomaly handling includes: identifying and removing abnormal samples from the original protein quantification matrix to obtain an effective protein quantification matrix; Analyze differentially expressed proteins, including: comparing the expression differences of proteins among different groups to obtain a set of first protein biomarkers with significant differences; Data standardization processing includes: performing data standardization processing on the effective protein quantification matrix to eliminate the influence of different dimensions and scales, and obtaining a standardized protein quantification matrix; The screening of protein biomarkers includes: screening the standardized protein quantification matrix using at least m of the following methods to obtain at least m candidate protein biomarker sets: LASSO regression analysis, partial least squares discriminant analysis, orthogonal partial least squares discriminant analysis, and RFS-based random forest algorithm, where m is a positive integer, 3≤m≤4; The biomarker aggregation analysis includes: performing a aggregation analysis on the first set of protein biomarkers and the set of m candidate protein biomarkers, and obtaining a set of target protein biomarkers that simultaneously exist in at least four sets in the first set of protein biomarkers and the set of m candidate protein biomarkers from the first set of protein biomarkers and the set of m candidate protein biomarkers. Correlation analysis includes: analyzing the correlation between the target protein biomarker set and all protein biomarkers in the effective protein quantification matrix, and forming a new set P1 with the first biomarker among all protein biomarkers whose correlation coefficient is within the first range and the target protein biomarker set; The variable structuring process includes: dividing the samples in set P1 into gastric cancer group and non-gastric cancer group according to clinical diagnosis, and encoding them respectively to form the dependent variable of the prediction model. The independent variables of the prediction model include the quantitative results of protein, peptide or antibody and the patient's basic information. Model training and testing include: dividing the samples in set P1 into a training set and a test set, training the prediction model using the training set, and testing the prediction model using the test set, wherein the prediction model is a regression model including the independent variable and the dependent variable; Model optimization includes: selecting effective protein biomarkers based on at least one of the indicators of p-value, correlation coefficient, and AUC value, using the effective protein biomarkers as effective independent variables, and obtaining the optimized model.
2. The construction method as described in claim 1, characterized in that, Between the database retrieval step and the exception handling step, there is also a further step; Examining batch effects in the data includes: dividing the sample into multiple batches and using principal component analysis to analyze the batch effects between different batches; To remove batch effects, if batch effects exist, use the R package statTarget and the QC-RFSC method to remove batch effects between samples caused by different batches based on the QC samples. statTarget is the default parameter setting.
3. The construction method as described in claim 1, characterized in that, Between the database retrieval step and the exception handling step, there is also a further step; Protein filtering includes selecting proteins with quantitative values in a preset number of samples for subsequent protein biomarker screening.
4. The construction method as described in claim 1, characterized in that, Between the database retrieval step and the exception handling step, there is also a further step; Missing values were filled, including both non-randomly missing protein quantification values and randomly missing protein quantification values, which were filled using the K-nearest neighbor method.
5. The construction method as described in claim 1, characterized in that, The steps for analyzing differentially expressed proteins include: The t-test method was used to compare protein expression differences among different groups, obtaining the set G1 with significant differences; and The Deseq2 method was used to compare the differences in protein expression among different groups to obtain a set G2 with significant differences. The first set of protein biomarkers includes set G1 and set G2.
6. The construction method as described in claim 1, characterized in that, The data standardization process includes: performing data standardization on the effective protein quantification matrix using one or more of eight standardization methods, including: Log2, Log2+Median, Log2+Mean, VSN, Log2+RLR, Log2+GI, Log2+Quantile, and Log2+CycLoess.
7. The construction method as described in claim 1, characterized in that, Also includes: Model iteration includes repeating the model training and testing steps and the model optimization steps to build multiple prediction models, and evaluating the performance of different prediction models based on the training and testing results of the multiple prediction models.
8. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P01009 and P59666, or their single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins or stable isotope characteristic peptide segments, and corresponding nucleic acids, any one or more of these. The optimization model is prediction model 1, and the mathematical expression of prediction model 1 is: P=1 / (1+Exp(-(β0+β1 × P01009 + β2 × P59666))), where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: the protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-54.02, -36.56]; the coefficient β1 ranges from [2.41, 4.02]; and the coefficient β2 ranges from [0.81, 1.37].
9. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are O94822 and P62328, or their single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins or stable isotope characteristic peptide segments, and corresponding nucleic acids, any one or more of these. The optimized model is prediction model 2, and the mathematical expression of prediction model 2 is: P=1 / (1+Exp(-(β0+β1 ×O94822 + β2 × P62328))), where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: the protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-3.37, 2.68]; the coefficient β1 ranges from [-1.14, -0.68]; and the coefficient β2 ranges from [0.83, 1.37].
10. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are O94822, P62328, and P00915, or any one or more of their single peptide chains, characteristic peptide segments of the peptide chains, stable isotope proteins or stable isotope characteristic peptide segments, or corresponding nucleic acids. The optimized model is prediction model 3, and the mathematical expression of prediction model 3 is: P=1 / (1+Exp(-(β0+β1 ×O94822 + β2 × P62328 + β3 × P00915))), where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: the protein biomarker, single peptide chain, characteristic peptide segment of the peptide chain, stable isotope protein, stable isotope characteristic peptide segment, or corresponding nucleic acid. The coefficient β0 ranges from [1.50, 9.17]; the coefficient β1 ranges from [-1.12, -0.64]; and the coefficient β2 ranges from [0.83, 1.39]; The value range of coefficient β3 is [-0.90, -0.40].
11. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are Q14156, P62328, and Q16280, or any one or more of their single peptide chains, characteristic peptide segments of the peptide chains, stable isotope proteins or stable isotope characteristic peptide segments, or corresponding nucleic acids. The optimized model is prediction model 4, and the mathematical expression of prediction model 4 is: P=1 / (1+Exp(-(β0+β1 ×Q14156 + β2 × P62328 + β3 × Q16280))), where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: the protein biomarker, single peptide chain, characteristic peptide segment of the peptide chain, stable isotope protein, stable isotope characteristic peptide segment, or corresponding nucleic acid. The coefficient β0 ranges from [5.36, 15.19]; the coefficient β1 ranges from [-0.89, -0.31]; and the coefficient β2 ranges from [0.85, 1.42]; The value range of coefficient β3 is [-1.68, -1.00].
12. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P01009, P59666, and Q92824, or any one or more of their single peptide chains, characteristic peptide segments of the peptide chains, stable isotopic proteins or stable isotopic characteristic peptide segments, or corresponding nucleic acids. The optimized model is prediction model 5, and the mathematical expression of prediction model 5 is: P=1 / (1+Exp(-(β0+β1 ×P01009 + β2 × P59666 + β3 × Q92824))), where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: the protein biomarker, single peptide chain, characteristic peptide segment of the peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, or corresponding nucleic acid. The coefficient β0 ranges from [-51.02, -33.02]; the coefficient β1 ranges from [2.78, 4.53]; and the coefficient β2 ranges from [0.73, 1.31]; The value range of the coefficient β3 is [-0.83, -0.41].
13. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P01009, P59666, and P02533, or any one or more of their single peptide chains, characteristic peptide segments of peptide chains, stable isotope proteins or stable isotope characteristic peptide segments, or corresponding nucleic acids. The optimized model is prediction model 6, and the mathematical expression of prediction model 6 is: P=1 / (1+Exp(-(β0+β1 ×P01009 + β2 × P59666 + β3 × P02533))), where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: the protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, or corresponding nucleic acid. The coefficient β0 ranges from [-56.40, -38.41]; the coefficient β1 ranges from [2.37, 4.00]; and the coefficient β2 ranges from [0.75, 1.32]; The value range of coefficient β3 is [0.12, 0.56].
14. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P01009, P59666, P00915, Q14156, and Q9UBF2, or any one or more of their single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins or stable isotopic characteristic peptide segments, or corresponding nucleic acids. The optimized model is prediction model 7, and the mathematical expression of prediction model 7 is: P = 1 / (1 + Exp(-(β0 + β1 × P01009 + β2 × P59666 + β3 × P00915 + β4 × Q14156 + β5 × Q9UBF2))), where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: the protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotopic protein, stable isotopic characteristic peptide segment, or corresponding nucleic acid. The coefficient β0 ranges from -5 to 1.
68. -29.98]; the value range of β1 is [2.88, 4.80]; the value range of coefficient β2 is [0.96, 1.62]; the value range of coefficient β3 is [-1.36, -0.71]; the value range of coefficient β4 is [-1.22, -0.50]; the value range of coefficient β5 is [0.33, 0.95].
15. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are any one or more of the following: P62328, Q14156, Q16280, P59666, P01009, Q9NQX4, and Q6ZVL6, or a single peptide chain, a characteristic peptide segment of the peptide chain, a stable isotope protein or a stable isotope characteristic peptide segment, or the corresponding nucleic acid. The optimized model is prediction model 8, and the mathematical expression of prediction model 8 is: P = 1 / (1 + Exp(-(β0 + β1 × P62328 + β2 × Q14156 + β3 × Q16280 + β4 × P59666 + β5 × P01009 + β6 × Q9NQX4 + β7) ×Q6ZVL6))), wherein the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-38.76, -14.57]; the coefficient β1 ranges from [0.45, 1.09]; the coefficient β2 ranges from [-1.16, -0.41]; the coefficient β3 ranges from [-1.63, -0.70]; the coefficient β4 ranges from [0.41, 1.15]; the coefficient β5 ranges from [1.87, 3.97]; the coefficient β6 ranges from [0.59, 1.56]; and the coefficient β7 ranges from [-1.04, -0.36].
16. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are any one or more of the following: P62328, Q14156, Q16280, P59666, P01009, Q9NQX4, Q6ZVL6, and P00915, or a single peptide chain, a characteristic peptide segment of the peptide chain, a stable isotope protein or a stable isotope characteristic peptide segment, or the corresponding nucleic acid. The optimized model is prediction model 9, and the mathematical expression of prediction model 9 is: P = 1 / (1 + Exp(-(β0 + β1 × P62328 + β2 × Q14156 + β3 × Q16280 + β4 × P59666 + β5 × P01009 + β6 × Q9NQX4 + β7 × Q6ZVL6 + β8 × P00915), wherein the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-34.50, -8.76]; β1 ranges from [0.42, 1.09]; β2 ranges from [-1.18, -0.39]; β3 ranges from [-1.71, -0.72]; β4 ranges from [0.52, 1.31]; β5 ranges from [2.12, 4.37]; β6 ranges from [2.12, 4.37]; and β7 ranges from [0.49, ...]. 1.51]; The value range of the coefficient β8 is [-1.35, -0.64].
17. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are any one or more of the following: P62328, O94822, Q93034, P59666, P41235, P02100, P69905, P04275, P02533, Q16280, and O15553, or a single peptide chain, a characteristic peptide segment of the peptide chain, a stable isotope protein or a stable isotope characteristic peptide segment, or a corresponding nucleic acid. The optimized model is prediction model 10, and the mathematical expression of prediction model 10 is: P = 1 / (1 + Exp(-(β0 + β1 × P62328 + β2 × O94822 + β3 × Q93034 + β4 × P59666 + β5 × P41235 + β6 × P02100 + β7 × P69905 + β8 × P04275 + β9 ×P02533+ β10 × Q16280+ β11 ×O15553))), where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-19.30, 2.94]; β1 ranges from [0.40, 1.03]; β2 ranges from [-0.89, -0.27]; β3 ranges from [-0.19, 0.94]; β4 ranges from [0.11, 0.97]; and β5 ranges from [0.27, 1.34]; the value range of coefficient β6 is [-1.09, 0.13]; the value range of coefficient β7 is [-1.18, -0.08]; the value range of coefficient β8 is [0.38, 1.28]; the value range of coefficient β9 is [0.01, 0.59]; the value range of coefficient β10 is [-1.77, -0.78]; the value range of coefficient β11 is [-0.47, 0.73].
18. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are P01009, Q9H0G5, Q16280, Q14156, P59666, Q9NQX4, P62328, P01242, P69905, P41235, O15553, Q6ZVL6, P04275, A0A075B6K4, Q9UBF2, P02100, Q8TA94, Q93034, and Q92824. P00915, Q6IC98, O94822, P02533, P05452, P02538, Q8IXQ9, or any single peptide chain, characteristic peptide segment of the peptide chain, and its stable isotope protein or stable isotope characteristic peptide segment, or the corresponding nucleic acid, etc., the optimization model is prediction model 11, and the mathematical expression of prediction model 11 is: P=1 / (1+Exp(-(β0+β1)) ×P01009 + β2 ×Q9H0G5 + β3 ×Q16280+ β4 × Q14156+ β5 × P59666+ β6 ×Q9NQX4+ β7 ×P62328+ β8 ×P01242+ β9×P69905+ β10 × P41235+ β11 × O15553+β12 ×Q6ZVL6+ β13 × P04275 + β14 ×A0A075B6K4+ β15 × Q9UBF2+ β16 × P02100+ β17 ×Q8TA94+ β18 ×Q93034+ β19 ×Q92824+ β20 ×P00915+ β21 × Q6IC98+ β22 × O94822+β23 ×P02533+ β24 ×P05452 + β25 × P02538+ β26 × Q8IXQ9)))), where the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid. The coefficient β0 ranges from [-35.05, 8.05]; β1 ranges from [0.63, 4.13]; β2 ranges from [0.14, 2.30]; β3 ranges from [-1.81, -0.52]; β4 ranges from [-1.61, -0.55]; and β5 ranges from [0.28, 1.38]. The coefficient β6 ranges from [0.20, 1.43]; the coefficient β7 ranges from [0.40, 1.17]; the coefficient β8 ranges from [-1.61, 0.06]; the coefficient β9 ranges from [-1.83, 0.62]; and the coefficient β10 ranges from [-0.27, 1].36]; the value range of coefficient β11 is [-1.28, 0.26]; the value range of coefficient β12 is [-0.91, -0.05]; the value range of coefficient β13 is [-0.07, 0.99]; the value range of coefficient β14 is [-1.02, 0.18]; the value range of coefficient β15 is [-0.09, 0.93]; the value range of coefficient β16 is [-1.09, 0.34]; the value range of coefficient β17 is [-0.78, 0.03]; the value range of coefficient β18 is [-0.33, 0.98]; the value range of coefficient β19 is [-0.70, 0.05]; the value range of coefficient β20 is [-1.16, [0.60]; the value range of coefficient β21 is [-0.32, 0.86]; the value range of coefficient β22 is [-0.70, 0.16]; the value range of coefficient β23 is [-0.25, 0.68]; the value range of coefficient β24 is [-1.35, 0.92]; the value range of coefficient β25 is [-0.27, 0.55]; the value range of coefficient β26 is [-1.07, 1.21].
19. The construction method as described in claim 1, characterized in that, The effective protein biomarkers are Q68DV7, P41235, P02748, P01009, O75152, Q06033, P02750, P02763, P01011, Q9H0G5, P60709, P08575, Q12851, P06727, Q9H6K5, O95236, P00915, P69905, Q9UFC0, P02100, P68871, Q8IXQ9, and P629. 06, Q5JY77, A0PJY2, O15553, O94822, P04275, P59666, Q16280, Q8TA94, Q92824, or any one or more of their single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins or stable isotopic characteristic peptide segments, or corresponding nucleic acids, wherein the optimization model is prediction model 12, and the mathematical expression of prediction model 12 is: P=1 / (1+Exp(-(β0+β1×Q68DV7)) + β2 ×P41235 + β3 × P02748+ β4 × P01009+ β5 × O75152+ β6 ×Q06033+ β7 ×P02750+ β8 ×P02763+ β9 ×P01011+ β10 × Q9H0G5+ β11 × P60709 +β12 ×P08575+ β13 × Q12851 + β14 ×P06727+ β15 × Q9H6K5+ β16 × O95236+ β17×P00915+ β18 × P69905+ β19 ×Q9UFC0+ β20 ×P02100+ β21 × P68871+ β22 ×Q8IXQ9 +β23 ×P62906+ β24 × Q5JY77 + β25 × A0PJY2+ β26 × O15553+ β27 ×O94822+ β28 × P04275 + β29 ×P59666+ β30 × Q16280 + β31 × Q8TA94 + β32 ×Q92824)))), wherein, the name of the effective protein biomarker refers to the standardized content value of any one or more of the following: protein biomarker, single peptide chain, characteristic peptide segment of peptide chain, stable isotope protein, stable isotope characteristic peptide segment, and corresponding nucleic acid; the coefficient β0 ranges from [-40.87, -2.20]; the coefficient β1 ranges from [-1.44, 0.28]; the coefficient β2 ranges from [-0.34, 1.35]; the value range of coefficient β3 is [-2.34, 0.00]; the value range of coefficient β4 is [0.05, 4.19]; the value range of coefficient β5 is [-0.49, 1.87]; the value range of coefficient β6 is [-0.93, 1.03]; the value range of coefficient β7 is [-0.86, 1.94]; the value range of coefficient β8 is [-0.51, 1.62]; the value range of coefficient β9 is [-1.58, 1.26]; the value range of coefficient β10 is [-0.17, 2.12]; the value range of coefficient β11 is [0.88, 2.69]; the value range of coefficient β12 is [-0.54, 2.32]; the value range of coefficient β13 is [-0.16, 0.90]; the value range of coefficient β14 is [-1.43, 3.58]; the value range of coefficient β15 is [-2.23, 1.81]; the value range of coefficient β16 is [-1.48, 1.03 ...6 is [-1.48, 1.03]; the value range of coefficient β17 is [-0.86, 1.94]; the value range of coefficient β8 is [-0.51, 1.62]; the value range of coefficient β9 is [-1.58, 1.26]; the value range of coefficient β10 is [-0.17, 2.12]; the value range of coefficient β11 is [0.88, 2.69]; the value range of coefficient β12 is [-0.54, 2.32]; the value range of coefficient β13 is [-0.16, 0.90]; the value range of coefficient β14 is [-1.43, 3.58]; the value range 1.04]; the value range of coefficient β17 is [-0.82, 0.96]; the value range of coefficient β18 is [-2.74, 1.06]; the value range of coefficient β19 is [-1.26, 1.50]; the value range of coefficient β20 is [-1.02, 0.36]; the value range of coefficient β21 is [-3.98, 1.00]; the value range of coefficient β22 is [-0.53, 2.15]; the value range of coefficient β23 is [-0.59, 0.64]; the value range of coefficient β24 is [-0.45, 0.39]; the value range of coefficient β25 is [-0.61, 0.74]; the value range of coefficient β26 is [-1.12, [0.42]; the value range of coefficient β27 is [-1.15, 0.03]; the value range of coefficient β28 is [-0.08, 1.01]; the value range of coefficient β29 is [-0.11, 0.92]; the value range of coefficient β30 is [-3.39, -0.08]; the value range of coefficient β31 is [-1.31, 0.06]; the value range of coefficient β32 is [-0.95, 0.04].
20. A system for assessing the likelihood of a subject developing gastric cancer based on the Swiss-prot database, characterized in that, include: The data acquisition module is used to acquire sample data from the subjects, including serum protein quantification: A0A075B6K4, A0PJY2, A6NFD8, A6QL64, O15033, O15553, O75152, O75460, O75683, O94822, O94885, O95236, O95490, P00738, P00751, P00915, P01009, P01011, P01185, P01242, P01701, P02042, P02 100, P02533, P02538, P02671, P02748, P02750, P02763, P04275, P05109, P05452, P06727, P08571, P08575, P0DJI8, P18 065, P18428, P31150, P35527, P41235, P53367, P59666, P59923, P60709, P62328, P62906, P68871, P69891, P69905, Q01 538, Q06033, Q09666, Q12851, Q14156, Q15047, Q15293, Q16280, Q2TBA0, Q4G0S7, Q4VXF1, Q53HC9, Q5H9J7, Q5JY77, Q5V U43, Q66K14, Q68DV7, Q6ZVL6, Q8IVV2, Q8IXQ9, Q8IZF2, Q8TA94, Q8TEW8, Q92797, Q92824, Q93034, Q96N87, Q96RV3, Q9H Any one of 0G5, Q9H5Y7, Q9H6K5, Q9NQV5, Q9NQX4, Q9P0M6, Q9UBF2, Q9UBP8, Q9UFC0, Q9UHG0, Q9UKN7, Q9UN37, Q9UPX8, and Q9Y5E1; or a single peptide chain, a characteristic peptide segment of the peptide chain, and its stable isotopic protein or stable isotopic characteristic peptide segment, or an antigen / antibody and any one or more of the single peptide chain, characteristic peptide segment of the peptide chain, its stable isotopic protein or stable isotopic characteristic peptide segment, and the corresponding nucleic acid; The data preprocessing module is used to preprocess the sample data, including removing duplicate samples, unifying units, and dividing the sample data into different analysis groups according to the type and quantity of indicators. The model recommendation module is used to recommend multiple prediction models with larger AUC values from the prediction models built using the construction method described in claim 1, based on the types and number of indicators included in different analysis groups. The model selection module provides users with a selection function to output one or more prediction models from the prediction models recommended by the model recommendation module for risk assessment and calculation. The risk assessment module is used to calculate the corresponding predicted value based on one or more prediction models output by the model selection module, and to provide an appropriate Logit threshold or risk score threshold according to the department or population information corresponding to the sample data, and to classify the risk level as any one of high probability, medium probability, and low probability according to the Logit threshold or risk score threshold.
21. A kit for assessing the likelihood of a subject having gastric cancer based on the Swiss-prot database, characterized in that, For receiving effective protein biomarkers, said effective protein biomarkers are any one or more of the following protein biomarkers: A0A075B6K4, A0PJY2, A6NFD8, A6QL64, O15033, O15553, O75152, O75460, O75683, O94822, O94885, O9523 6. O95490, P00738, P00751, P00915, P01009, P01011, P01185, P01242, P01701, P02042, P02100, P02533, P02538, P02671, P02748, P02750, P02763, P04275, P05109, P05452, P06727, P08571, P08575, P0DJI8, P1 8065, P18428, P31150, P35527, P41235, P53367, P59666, P59923, P60709, P62328, P62906, P68871, P698 91. P69905, Q01538, Q06033, Q09666, Q12851, Q14156, Q15047, Q15293, Q16280, Q2TBA0, Q4G0S7, Q4VXF1 , Q53HC9, Q5H9J7, Q5JY77, Q5VU43, Q66K14, Q68DV7, Q6ZVL6, Q8IVV2, Q8IXQ9, Q8IZF2, Q8TA94, Q8TEW8, Q 92797, Q92824, Q93034, Q96N87, Q96RV3, Q9H0G5, Q9H5Y7, Q9H6K5, Q9NQV5, Q9NQX4, Q9P0M6, Q9UBF2, Q9UBP8, Q9UFC0, Q9UHG0, Q9UKN7, Q9UN37, Q9UPX8, Q9Y5E1, or a single peptide chain, a characteristic peptide segment of the peptide chain, and its stable isotope protein or stable isotope characteristic peptide segment, and the corresponding nucleic acid.
22. The kit according to claim 21, characterized in that, The effective protein markers are any one or more of P01009 and P59666, or their single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins or stable isotopic characteristic peptide segments, and corresponding nucleic acids.
23. The kit according to claim 21, characterized in that, The effective protein markers are any one or more of O94822 and P62328, or their single peptide chains, characteristic peptide segments of peptide chains, and their stable isotopic proteins or stable isotopic characteristic peptide segments, or their corresponding nucleic acids.
24. The kit according to claim 21, characterized in that, The effective protein markers are any one or more of O94822, P62328 and P00915, or a single peptide chain, a characteristic peptide segment of the peptide chain, a stable isotope protein or a stable isotope characteristic peptide segment, or a corresponding nucleic acid.
25. The kit according to claim 21, characterized in that, The effective protein biomarkers are any one or more of Q14156, P62328 and Q16280, or their single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins or stable isotopic characteristic peptide segments, and corresponding nucleic acids.
26. The kit according to claim 21, characterized in that, The effective protein markers are any one or more of P01009, P59666 and Q92824, or their single peptide chains, characteristic peptide segments of peptide chains, stable isotopic proteins or stable isotopic characteristic peptide segments, and corresponding nucleic acids.
27. The kit according to claim 21, characterized in that, The effective protein markers are any one or more of P01009, P59666 and P02533, or a single peptide chain, a characteristic peptide segment of the peptide chain, a stable isotope protein or a stable isotope characteristic peptide segment, or a corresponding nucleic acid.
28. The kit according to claim 21, characterized in that, The effective protein markers are any one or more of P01009, P59666, P00915, Q14156 and Q9UBF2, or a single peptide chain, a characteristic peptide segment of the peptide chain, a stable isotope protein or a stable isotope characteristic peptide segment, or a corresponding nucleic acid.
29. The kit according to claim 21, characterized in that, The effective protein markers are any one or more of P62328, Q14156, Q16280, P59666, P01009, Q9NQX4 and Q6ZVL6, or a single peptide chain, a characteristic peptide segment of the peptide chain, and a stable isotope protein or a stable isotope characteristic peptide segment, or a corresponding nucleic acid.
30. The kit according to claim 21, characterized in that, The effective protein markers are any one or more of the following: P62328, Q14156, Q16280, P59666, P01009, Q9NQX4, Q6ZVL6 and P00915, or a single peptide chain, a characteristic peptide segment of the peptide chain, a stable isotope protein or a stable isotope characteristic peptide segment, or a corresponding nucleic acid.
31. The kit according to claim 21, characterized in that, The effective protein biomarkers are any one or more of the following: P62328, O94822, Q93034, P59666, P41235, P02100, P69905, P04275, P02533, Q16280, and O15553, or a single peptide chain, a characteristic peptide segment of the peptide chain, a stable isotope protein or a stable isotope characteristic peptide segment, or a corresponding nucleic acid.
32. The kit according to claim 21, characterized in that, The effective protein biomarkers are Q68DV7, P41235, P02748, P01009, O75152, Q06033, P02750, P02763, P01011, Q9H0G5, P60709, P08575, Q12851, P06727, Q9H6K5, O95236, P00915, P69905, and Q9UFC. 0, P02100, P68871, Q8IXQ9, P62906, Q5JY77, A0PJY2, O15553, O94822, P04275, P59666, Q16280, Q8TA94 and Q92824, or a single peptide chain, a characteristic peptide segment of the peptide chain, and a stable isotopic protein or a stable isotopic characteristic peptide segment, or any one or more of the corresponding nucleic acids.
33. The kit according to claim 21, characterized in that, The effective protein markers are any one or more of the following: P01009, Q9H0G5, Q16280, Q14156, P59666, Q9NQX4, P62328, P01242, P69905, P41235, O15553, Q6ZVL6, P04275, A0A075B6K4, Q9UBF2, P02100, Q8TA94, Q93034, Q92824, P00915, Q6IC98, O94822, P02533, P05452, P02538 and Q8IXQ9, or a single peptide chain, a characteristic peptide segment of the peptide chain, and a stable isotope protein or a stable isotope characteristic peptide segment, or a corresponding nucleic acid.
34. The kit according to claim 21, characterized in that, The effective protein biomarkers are P62328, O94822, P01009, P02748, Q93034, O75152, P59666, P02750, P68871, P62906, Q8IXQ9, P41235, P02538, Q06033, P05452, Q8TA94, Q9UN37, and P5336.
7. P02100, P69905, Q9H0G5, Q14156, P04275, Q9NQX4, P02533, Q9UHG0, Q16280, Q6 ZVL6, P00915, O15553, Q9UBF2, P01011, A0A075B6K4, Q92824, P35527, Q15047, P01 242. P02671, P60709, Q8TEW8, P18065, O75683, Q8IVV2, Q9UBP8, P0DJI8, Q15293, Q68DV7, A6NFD8, P05109, Q09666, P00738, Q5VU43, Q9UPX8, Q9Y5E1, P59923, Q4G0S 7. Q9UKN7, P18428, Q4VXF1, Q66K14, Q9NQV5, P02042, Q9H5Y7, Q8IZF2, O95490, Q2TBA0, Q9P0M6 and O94885, or a single peptide chain, a characteristic peptide segment of the peptide chain, and a stable isotopic protein or a stable isotopic characteristic peptide segment, or any one or more of the corresponding nucleic acids.
35. The kit according to claim 21, characterized in that, The effective protein biomarkers are A0A075B6K4, A0PJY2, A6NFD8, A6QL64, O15033, O15553, O75152, O75460, O75683, O94822, O94885, O95236, O95490, P00738, P00751, P00915, P01009, P01011, P01185, P01242, P01701, P02042, P02100, P 02533, P02538, P02671, P02748, P02750, P02763, P04275, P05109, P05452, P06727, P08571, P08575, P0DJI 8. P18065, P18428, P31150, P35527, P41235, P53367, P59666, P59923, P60709, P62328, P62906, P68871, P69 891, P69905, Q01538, Q06033, Q09666, Q12851, Q14156, Q15047, Q15293, Q16280, Q2TBA0, Q4G0S7, Q4VXF1, Q53HC9, Q5H9J7, Q5JY77, Q5VU43, Q66K14, Q68DV7, Q6ZVL6, Q8IVV2, Q8IXQ9, Q8IZF2, Q8TA94, Q8TEW8, Q9279 7. Q92824, Q93034, Q96N87, Q96RV3, Q9H0G5, Q9H5Y7, Q9H6K5, Q9NQV5, Q9NQX4, Q9P0M6, Q9UBF2, Q9UBP8, Q9UFC0, Q9UHG0, Q9UKN7, Q9UN37, Q9UPX8, Q9Y5E1, or any single peptide chain, characteristic peptide segment of the peptide chain, and its stable isotope protein or stable isotope characteristic peptide segment, or the corresponding nucleic acid, one or more of these.
36. An electronic device, characterized in that, include: Memory is used to store instructions that can be executed by the processor; A processor for executing the instructions to implement the construction method as described in any one of claims 1-19.
37. A computer-readable medium storing computer program code that, when executed by a processor, implements the construction method as described in any one of claims 1-19.
Citation Information
Cited By
Coronary artery disease early prediction model construction method, system, equipment and medium
CN122117431A