Machine learning based plasma protein marker colorectal cancer risk prediction system, method and storage medium

CN122822341APending Publication Date: 2026-09-25ANHUI MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611018820.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]当前结直肠癌风险预测模型的构建存在以下缺陷:(一)是特征筛选多依赖单变量分析,未有效处理蛋白间的共线性和冗余性,导致标志物组合不够精炼,检测成本高且模型泛化能力不足;(二)是建模方法多采用逻辑回归等线性模型,难以捕捉血浆蛋白与疾病状态之间的复杂非线性关系,预测准确率受限;(三)缺乏严格的超参数调优与外部验证,泛化能力不足,模型构建常忽略系统性的超参数搜索和独立外部验证,易造成过拟合

Benefits of technology

(1)通过limma差异分析和LASSO回归的双重筛选,有效剔除冗余蛋白和共线性干扰,在保证预测精度的前提下将关键蛋白数量压缩至最少,大幅降低检测成本和复杂度;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822341A_ABST
    Figure CN122822341A_ABST
Patent Text Reader

Abstract

The application discloses a kind of plasma protein marker colorectal cancer risk prediction system based on machine learning, including data acquisition module, data processing module, model training module and prognosis risk assessment module, data acquisition module: for collecting the plasma protein expression data of screening population;Data processing module: it is connected with data acquisition module, for the plasma protein expression data is preprocessed, further extract key protein marker combination;Model training module: it is connected with data processing module, for using the expression data of key protein marker combination and the disease label of corresponding sample, trains machine learning model based on XGBoost integrated learning algorithm, and exports the risk assessment model of well-trained.Risk prediction assessment module: it is connected with model training module, for receiving the plasma protein expression data of individual to be evaluated, input the risk assessment model of well-trained, output the risk probability value of the individual to be evaluated suffering from colorectal cancer.The prediction system of the application can effectively assist the accuracy of diagnosis of suffering from colorectal cancer, improve the precision of patient survival outcome prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent medical technology, and in particular to a machine learning-based system for predicting the risk of colorectal cancer using plasma protein biomarkers. Background Technology

[0002] Colorectal cancer (CRC) is a significant public health problem, ranking as the third most common cancer and the second leading cause of cancer death. Notably, while CRC is curable in its early, localized stages, approximately 25% of patients are diagnosed with metastatic disease, resulting in poorer treatment outcomes and prognoses. However, colonoscopy, the gold standard for early CRC screening, is invasive, time-consuming, and expensive. There is an urgent need for a non-invasive biomarker for early screening and diagnosis of CRC, enabling large-scale population screening with a simple and accurate method.

[0003] Plasma proteins, including classic circulating proteins and tissue "leakage" proteins, are valuable biomarkers for various diseases. Studies have shown that alterations in plasma proteins are closely related to the risk of colorectal cancer and have the potential to become biomarkers for early screening of colorectal cancer, thereby establishing risk prediction models for high-risk populations. With the rapid development of big data and information technology, artificial intelligence methods centered on machine learning are being more widely applied in clinical practice for early disease warning. Machine learning, with its advantages in high-throughput data processing, has been widely used in constructing tumor prognostic models and has shown great potential.

[0004] The current construction of colorectal cancer risk prediction models has the following defects: (i) Feature selection relies heavily on univariate analysis, which fails to effectively handle collinearity and redundancy among proteins, resulting in insufficiently refined biomarker combinations, high detection costs, and insufficient model generalization ability; (ii) Modeling methods mostly adopt linear models such as logistic regression, which are difficult to capture the complex nonlinear relationship between plasma proteins and disease status, thus limiting prediction accuracy; (iii) There is a lack of rigorous hyperparameter tuning and external validation, resulting in insufficient generalization ability. Model construction often ignores systematic hyperparameter search and independent external validation, which easily leads to overfitting. Summary of the Invention

[0005] To address the technical problems in the prior art, the present invention provides the following technical solution: Firstly, this invention provides a machine learning-based system for predicting colorectal cancer risk using plasma protein biomarkers. The system comprises four modules: a data acquisition module, used to collect plasma samples from a screened population and obtain plasma protein expression levels for each sample through a standardized proteomics process; a data processing module, connected to the data acquisition module, used to preprocess the plasma protein expression data, perform differential expression analysis using the limma algorithm, screen for differentially expressed proteins associated with colorectal cancer, and further extract key protein biomarker combinations from the differentially expressed proteins using the LASSO regression algorithm; a model training module, connected to the data processing module, used to train a machine learning model based on the XGBoost ensemble learning algorithm using the expression levels of the key protein biomarker combinations and the corresponding disease labels of the samples, and output the trained risk assessment model; and a risk prediction and assessment module, connected to the model training module, used to receive plasma protein expression data of the individual to be assessed, input it into the trained risk assessment model, and output the individual's risk probability value for colorectal cancer.

[0006] The data acquisition module includes the following proteomics workflow: plasma protein extraction, peptide digestion, data-dependent acquisition (DDA) library construction, chromatographic fractionation, liquid chromatography-tandem mass spectrometry (LC-MS / MS) data acquisition, and database retrieval. This workflow can obtain relative or absolute quantitative information of proteins in each sample, forming an initial plasma protein expression matrix.

[0007] The data processing module is configured to execute the following processing flow: Preprocessing: Missing values ​​in the protein expression matrix were imputed and the matrix was standardized to eliminate systematic errors.

[0008] Differential Expression Analysis (limma Algorithm): The limma (Linear Models for Microarray Data) algorithm was used for differential expression analysis. This algorithm addresses the characteristics of plasma proteomics—small sample size, high dimensionality, and high noise—by fitting a linear model and introducing an empirical Bayesian method to stabilize variance estimation. A modified t-statistic is then used for significance inference, forming the core technical means for accurately screening candidate biomarkers from massive protein datasets. The core advantage of limma lies in its empirical Bayesian contraction strategy: by estimating a common prior variance distribution for all proteins, the estimated variance of each protein is contracted towards the global mean, effectively avoiding false positives or missed detections caused by unstable variance estimations of individual proteins under small sample conditions. Simultaneously, the modified t-statistic calculated based on the contracted variance gains augmented degrees of freedom, maintaining reasonable statistical power even with only a few dozen samples per group. It is particularly suitable for differential analysis of small-sample, high-dimensional proteomics data.

[0009] For each protein g Its expression level y g Modeled as: ; in X To design a matrix that indicates the group information of the samples, These are the coefficients to be estimated. By comparing the matrices... C Constructing hypothesis testing = 0 (i.e., the protein showed no difference between the two groups), the alternative hypothesis is ≠ 0 (meaning the protein differs between the two groups). A single modeling step allows for flexible extraction of comparison results between any two groups, avoiding the tediousness and information waste of repeatedly performing two-group t-tests.

[0010] To determine whether to reject the null hypothesis, a test statistic needs to be constructed. First, the effect size needs to be estimated. ,in This is the least squares estimate of the coefficients. In pairwise comparison scenarios, this effect size is equal to the difference between the two group means (on a logarithmic scale). The standard error of the effect size is: ; in , is a scalar determined by the design matrix and the comparison matrix; For protein g The residual standard deviation.

[0011] For each protein, estimate its own variance. Compared with the prior variance estimated from all proteins We perform a weighted average to obtain the variance after shrinkage: ; in, These are the residual degrees of freedom of the protein itself. These are the prior degrees of freedom. The variance after contraction. Between and The specific degree of contraction is determined by the degrees of freedom of both. This operation brings the extreme values ​​of variance estimation back to a reasonable range. For proteins with abnormally small variance due to random errors, the variance is appropriately increased to avoid false positives; for proteins with abnormally large variance, the variance is appropriately decreased to avoid missed detections.

[0012] Using the standard deviation after shrinkage Replace the original standard deviation The modified t-statistic was calculated as follows: ; This statistic, when the null hypothesis is true, approximately follows a set of degrees of freedom. The t-distribution is used. Increased degrees of freedom allow the variance of each protein to converge towards the global mean even with small sample sizes of only a few dozen samples per group, effectively avoiding false positives caused by unstable variance estimation of individual proteins in small samples. Further calculation of fold change (FC) is performed, and the result is adjusted to |log2FC|>1.5. P Using <0.05 as the criterion, a set of differentially expressed proteins that are significantly associated with cancer was screened.

[0013] LASSO Regression Feature Selection: To further compress features and eliminate collinear proteins, this invention uses the LASSO algorithm based on logistic regression to select differentially expressed proteins. LASSO achieves variable selection by applying an L1 regularization term to the negative log-likelihood loss function of logistic regression, causing some coefficients to shrink precisely to zero. Its objective function is: ; Where n is the total number of training samples, i.e. the number of plasma samples used for model construction; p The number of candidate proteins that entered the LASSO screening; For disease labeling; For the sample i Expression vectors on differentially expressed proteins; For the intercept term; This is a vector of regression coefficients; Let λ be the coefficient corresponding to the j-th protein, and λ be the regularization parameter that controls the intensity of the penalty.

[0014] The core of LASSO's ability to precisely shrink some coefficients to zero lies in the mathematical properties of the L1 norm, specifically the L1 norm penalty term. A proportional shrinkage is applied to each coefficient. This characteristic is reflected in the soft thresholding operation during the optimization process, denoted as the... j The updated value of each coefficient without penalty is Then the solution after adding L1 penalty is explicitly given by the soft threshold function: ; For the first j The regression coefficient estimate for each protein is calculated using the following rule: when | When, the coefficient shrinks towards zero. Units; when | At this point, the coefficients are directly set to zero, rather than approaching a very small non-zero value. This zeroing operation is a unique feature of L1 regularization, stemming from the non-differentiability of the absolute value function at zero. This allows LASSO to produce sparse solutions, automatically removing proteins with weak contributions from the model. As the regularization parameter λ increases, the threshold rises, more coefficients are compressed to zero, and the number of retained proteins gradually decreases.

[0015] This invention employs 10-fold cross-validation for parameter selection. Specifically, the training set is randomly divided into 10 equal parts. Each time, 9 parts are used as the training subset to fit the LASSO model, and the remaining part is used as the validation subset to evaluate prediction performance. This process is repeated until each data set has been validated once. Given candidate values, the mean prediction bias is calculated across all 10 validations. Prediction bias is measured using binomial deviation. ; Where n is the number of samples in the validation subset; This represents the true disease label for the i-th sample; The predicted probability of the model for this sample is obtained by mapping the linear predicted value to the [0,1] interval using the Sigmoid function. The binomial bias measures the degree of agreement between the predicted probability and the true label; a smaller value indicates a more accurate model prediction. The cross-validation average bias is selected. Value (denoted as) Under this parameter, proteins with non-zero coefficients in the model constitute the final key protein biomarker set, while proteins with coefficients compressed to zero are removed in subsequent modeling. Model training module The model training module is configured to train the XGBoost model using the combination of the key protein biomarkers. The specific process of training the XGBoost model includes the following steps: Data partitioning: To ensure the model's generalization ability, a stratified sampling method is used to divide the complete dataset containing key protein expression levels and disease labels into a training set and a validation set according to a set ratio (e.g., 7:3), so that the proportion of colorectal cancer and normal samples in the two sets is consistent.

[0016] Feature engineering and format conversion: The sparse.model.matrix function is used to convert categorical variables (if any) to a sparse matrix format, remove the intercept term, and construct the feature matrix X_train. Then, the xgb.DMatrix function is used to encapsulate the feature matrix and labels into the XGBoost-specific data format dtrain. The validation set data is constructed in the same way as dtest.

[0017] Automatic hyperparameter tuning: A strategy combining adaptive cross-validation and grid search is employed to find the optimal hyperparameter combination. A hyperparameter grid is defined, encompassing candidate values ​​for the following parameters: number of iterations (nrounds), maximum tree depth (max_depth), learning rate (eta), minimum loss reduction for node splits (gamma), feature subsampling ratio (colsample_bytree), minimum leaf node weights (min_child_weight), and sample subsampling ratio (subsample). Using 10-fold adaptive cross-validation as the search method and minimizing log loss as the objective, all parameter combinations are automatically trained and evaluated to select the optimal hyperparameters.

[0018] Final Model Training and Formula: With a determined optimal hyperparameter configuration, an XGBoost binary classification model is trained on the training set dtrain using the xgb.train function. XGBoost (eXtreme Gradient Boosting) is an ensemble algorithm based on gradient boosting decision trees, which minimizes a regularized objective function by progressively adding decision trees. This invention employs binary logistic loss, and the model is an additive model with K trees. ; in, This represents the predicted score output by the model for the i-th sample. Let be the feature vector of the i-th sample, which is the expression level of the sample in the combination of key protein biomarkers; The total number of decision trees (number of iterations); f k This is the k-th regression tree; Space for all possible trees. Add one new tree in each training round. f t To minimize the following regularization objective: ; in Let n be the objective function in the t-th iteration; n is the total number of training samples. This represents the true disease label for the i-th sample; The cross-entropy loss function; For the first t The predicted value for round 1; is a regularization term used to control the complexity of the t-th tree and prevent overfitting; T is the number of leaf nodes in the t-th tree; Let be the weight score of the j-th leaf node; γ and λ are regularization hyperparameters, controlling the number of leaf nodes and the weight magnitude, respectively. The loss function is approximated using a second-order Taylor expansion, and a greedy algorithm is used to split the tree structure node by node. After training, the model's final prediction probability for sample i is transformed into the output probability using a logistic function: ; The linear prediction (log-odds) of the model for the i-th sample is obtained by summing the scores of the corresponding leaf nodes of all K trees; This represents the risk value of sample i for colorectal cancer. After training, the model object is saved, which contains the structure, splitting features, and leaf node scores of all decision trees, constituting the trained risk assessment model.

[0019] Performance Evaluation: The trained model was used to predict the risk probability of colorectal cancer for both the training and validation sets. The area under the receiver operating characteristic (ROC) curve (AUC) was used to evaluate the model's discriminative ability. The ROC curve, with the false positive rate (1-specificity) on the horizontal axis and the true positive rate (sensitivity) on the vertical axis, reflects the model's overall classification performance at different decision thresholds. The AUC value ranges from 0.5 to 1.0; the closer the AUC is to 1.0, the stronger the model's ability to distinguish between cancerous and normal samples.

[0020] Risk prediction and assessment module The risk prediction and assessment module is configured to: perform the same proteomics process as the data acquisition module on the plasma sample of the individual to be assessed to obtain the expression levels of the key protein biomarker combination; process the data according to the same standardization rules as the training data (e.g., use the mean and standard deviation of the histones in the training set for Z-score transformation); input the processed data into the trained XGBoost model, and output the colorectal cancer risk probability through model calculation; set a threshold to classify the probability into high risk and low risk, and generate an assessment report.

[0021] Secondly, this application provides a method for constructing a colorectal cancer risk prediction model based on machine learning plasma protein biomarkers, which is executed by some or all components of the colorectal cancer risk prediction system based on machine learning plasma protein biomarkers provided in the first aspect. The method includes: obtaining plasma protein expression data of samples through a proteomics process; preprocessing the plasma protein expression data, performing differential expression analysis using the limma algorithm, and screening for differentially expressed proteins associated with colorectal cancer; extracting key protein biomarker combinations from the differentially expressed proteins using the LASSO regression algorithm; and training a risk assessment model using the expression data of the key protein biomarker combinations and the corresponding disease labels of the samples as inputs, and using the XGBoost ensemble learning algorithm.

[0022] Thirdly, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method for constructing a machine learning-based plasma protein biomarker colorectal cancer risk prediction model provided in the second aspect.

[0023] Through the above technical solution, the present invention has the following beneficial effects: (1) By using limma difference analysis and LASSO regression for dual screening, redundant proteins and collinear interference are effectively eliminated, and the number of key proteins is reduced to a minimum while ensuring prediction accuracy, thus significantly reducing detection costs and complexity. (2) The XGBoost ensemble learning algorithm can fully capture the nonlinear interaction between protein biomarkers and has higher prediction accuracy and robustness than traditional linear models such as logistic regression. (3) The entire model training process incorporates rigorous stratified sampling, adaptive hyperparameter tuning and cross-validation to ensure the generalization ability and stability of the model.

[0024] It is understood that the prediction system, method, and readable storage medium provided above have the same beneficial effects. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0026] Figure 1 This is a block diagram of a machine learning-based plasma protein biomarker system for predicting colorectal cancer risk.

[0027] Figure 2 The results were screened using LASSO regression.

[0028] Figure 3 The AUC value is used to predict the model in the training set.

[0029] Figure 4 The AUC value of the prediction model in the test set. Detailed Implementation

[0030] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0031] The following detailed description, in conjunction with the accompanying drawings, through specific embodiments and application scenarios, illustrates the machine learning-based plasma protein biomarker colorectal cancer risk prediction system and model construction method provided by this invention.

[0032] Example 1 Please refer to the appendix. Figure 1 The machine learning-based plasma protein biomarker colorectal cancer risk prediction system provided in the embodiments of the present invention includes the following four modules: Data acquisition module: Used to collect plasma samples from the screened population and obtain plasma protein expression data for each sample through a standardized proteomics process.

[0033] Data processing module: Connected to the data acquisition module, it is used to preprocess the plasma protein expression data, perform differential expression analysis using the limma algorithm, screen out differentially expressed proteins associated with colorectal cancer, and then use the LASSO regression algorithm to further extract key protein biomarker combinations from the differentially expressed proteins.

[0034] Model training module: Connected to the data processing module, it is used to train a machine learning model based on the XGBoost ensemble learning algorithm using the expression level data of the key protein biomarker combination and the disease labels of the corresponding samples, and output the trained risk assessment model.

[0035] Risk prediction and assessment module: connected to the model training module, used to receive plasma protein expression data of the individual to be assessed, input the trained risk assessment model, and output the risk probability value of the individual having colorectal cancer.

[0036] Example 2 The present invention also provides a specific method for colorectal cancer risk prediction based on the machine learning-based plasma protein biomarker colorectal cancer risk prediction system according to Embodiment 1. The risk prediction process performed by this method mainly includes several parts such as system construction, model training, and risk assessment application.

[0037] Specifically, this embodiment is based on plasma samples from 50 patients with pathologically confirmed colorectal cancer (experimental group) and 100 healthy controls (control group) to construct the system.

[0038] The plasma sample data used in this embodiment must be subject to strict standard control. The inclusion criteria are: (1) all patients are diagnosed with colorectal cancer by postoperative pathology or tissue biopsy; (2) patients have not received radiotherapy, chemotherapy or radiotherapy-chemotherapy before surgery; (3) relevant clinical pathological information can be obtained; (4) patients do not have other tumor diseases at the same time; (5) the current medical history, personal history, family history and physical examination data are detailed and complete.

[0039] Exclusion criteria: (1) Exclude other colorectal tumors that are not colorectal adenocarcinoma, undifferentiated carcinoma, adenosquamous carcinoma, or squamous cell carcinoma, such as carcinoid, neuroendocrine carcinoma, malignant melanoma, malignant lymphoma, etc.; (2) Exclude multiple primary colorectal cancers, familial adenomatous polyposis, or other malignant tumors that are present or have been present in the past.

[0040] The control group consisted of healthy individuals from the same cohort, matched for the same sex and age ± 5 years.

[0041] The prediction system in this embodiment can be further subdivided into the following steps when making a specific colorectal cancer risk prediction based on the above sample data: Step 1: Data Collection Each sample was processed using a standardized procedure: total plasma protein was extracted from the lysate, and trypsin was used to digest it into peptides. Liquid chromatography-mass spectrometry (LC-MS / MS) was performed using data-dependent acquisition mode (DDA). Quantitative information of each protein was obtained by searching databases (such as the UniProt human protein database) to construct an original matrix containing 150 samples and several protein expression levels.

[0042] Step 2: Data Processing The protein expression matrix was imported into the R language environment. First, missing values ​​were imputed using K-nearest neighbor interpolation, and quantile standardization was performed. Then, differential expression analysis was performed using the limma package: a design matrix was constructed with cancer group vs. normal group, and a linear model was fitted. P Using a criterion of <0.05 and |log2FC|>1.5, 503 differentially expressed proteins were selected. Then, LASSO logistic regression was performed using the glmnet package, and lambda.min was selected through 10-fold cross-validation. Finally, 33 key plasma protein biomarkers were obtained through compression. Specific screening results are shown below. Figure 2 As shown.

[0043] Step 3: Model Training Dataset partitioning: Using the createDataPartition function of the caret package, the 100 samples were divided into a training set (70 cases) and a validation set (30 cases) in a 7:3 ratio, with disease status as the stratification variable.

[0044] Feature engineering: Using the expression levels of seven key proteins as features, a sparse matrix (or ordinary numerical matrix if containing only continuous variables) is constructed using `sparse.model.matrix`. The intercept column is removed to obtain the training feature matrix `train_x` and the validation feature matrix `val_x`. The label vectors are converted to 0 / 1 values ​​(normal = 0, cancer = 1), and `xgb.DMatrix` objects `dtrain` and `dtest` are constructed.

[0045] Hyperparameter grid search: Initial parameters are set as follows: objective = "binary:logistic", eval_metric = "logloss", booster = "gbtree". The grid to be searched is defined as follows: nrounds = c(50,100,200), max_depth = c(3,5,6), eta = c(0.01,0.1,0.3), gamma = c(0,0.1,0.5), colsample_bytree = c(0.8,1), min_child_weight = c(1,3), subsample = c(0.8,1). The train function from the caret package is used, with xgbTree as the method, and 10-fold adaptive cross-validation (adaptive_cv) as the control parameter. After the search, the optimal hyperparameter combination is obtained as follows: max_depth = 5, eta = 0.1, gamma = 0, colsample_bytree = 0.8, min_child_weight = 3, subsample = 0.8, nrounds = 100.

[0046] Final Model Training and Evaluation: Using the optimal parameters described above, the final model xgb_model_caret.bst was trained using xgb.train. After training, risk probabilities were predicted on both the training and validation sets, and the ROC curve was calculated using the pROC package. The results are as follows: Figure 3-4 As shown: the training set AUC = 0.74 and the validation set AUC = 0.71, indicating that the model has excellent discriminative power and generalization performance. The model is saved as a deployable file.

[0047] Step 4: Application of Risk Assessment For a plasma sample from a candidate to be evaluated, protein quantification data were obtained following the procedure described in step one, and the expression levels of the 33 key proteins were extracted. The data were Z-score standardized using the mean and standard deviation of the 33 proteins in the training set. The transformed feature vector was input into the loaded XGBoost model. The model partitioned the sample to the corresponding leaf node according to the decision path of each tree, accumulated the scores, and transformed them through a logistic function to output the risk probability value.

[0048] Example 3 This application also provides a corresponding method for constructing a colorectal cancer risk prediction model based on machine learning plasma protein biomarkers, which is used in conjunction with the aforementioned Embodiment 1 and Embodiment 2. The method mainly includes the process of constructing a colorectal cancer risk prediction model, specifically including the following steps: S1, obtain plasma protein expression data of samples through proteomics workflow; S2, the plasma protein expression data are preprocessed, and differential expression analysis is performed using the limma algorithm to screen for differentially expressed proteins associated with colorectal cancer; S3, use the LASSO regression algorithm to extract a combination of key protein biomarkers from the differentially expressed proteins; S4. Using the expression levels of the key protein biomarkers and the corresponding disease labels of the samples as input, the risk assessment model is trained using the XGBoost ensemble learning algorithm.

[0049] The final trained risk assessment model can receive plasma protein expression data of the individual to be assessed as input and output the risk probability value of the individual having colorectal cancer, which can effectively improve the accuracy of predicting patient survival outcomes.

[0050] In summary, the machine learning-based plasma protein biomarker colorectal cancer risk prediction system and model construction method provided in the above embodiments of this application effectively eliminate redundant proteins and collinearity interference through dual screening using limma differential analysis and LASSO regression. This minimizes the number of key proteins while maintaining prediction accuracy, significantly reducing detection costs and complexity. The XGBoost ensemble learning algorithm effectively captures the nonlinear interactions between protein biomarkers, exhibiting higher prediction accuracy and robustness than traditional linear models such as logistic regression. The entire model training process incorporates rigorous stratified sampling, adaptive hyperparameter tuning, and cross-validation to ensure the model's generalization ability and stability.

[0051] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0052] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0053] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0054] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0055] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0056] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0057] If the integrated module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0058] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A machine learning-based system for predicting colorectal cancer risk using plasma protein biomarkers, characterized in that, include: The data acquisition module is used to collect plasma samples from the screened population and obtain plasma protein expression data for each sample through a standardized proteomics process. The data processing module, connected to the data acquisition module, is used to preprocess and perform differential expression analysis on the plasma protein expression data, screen out differentially expressed proteins related to colorectal cancer, and extract key protein biomarker combinations from the differentially expressed proteins using a feature selection algorithm. The model training module, connected to the data processing module, is used to train a machine learning model using the expression level data of the key protein biomarker combination and the disease labels of the corresponding samples, and output the trained risk assessment model. The risk prediction and assessment module, connected to the model training module, is used to receive plasma protein expression data of the individual to be assessed, input the trained risk assessment model, and output the risk probability value of the individual having colorectal cancer.

2. The machine learning-based plasma protein biomarker colorectal cancer risk prediction system according to claim 1, characterized in that, The differential expression analysis in the data processing module uses the limma algorithm, and the feature selection algorithm uses the LASSO regression algorithm.

3. The machine learning-based plasma protein biomarker colorectal cancer risk prediction system according to claim 2, characterized in that, The limma algorithm fits a linear model and uses an empirical Bayesian method to adjust the variance, ensuring that |log2FC|>1.5 and after adjustment... P <0.05 is the standard for screening differentially expressed proteins, and FC is the fold change. The LASSO regression algorithm shrinks some regression coefficients to zero by applying an L1 regularization term to the negative log-likelihood loss function of logistic regression, and selects the optimal regularization parameter through cross-validation to retain proteins with non-zero coefficients as the combination of key protein biomarkers.

4. The machine learning-based plasma protein biomarker colorectal cancer risk prediction system according to claim 1, characterized in that, The machine learning model in the model training module is the XGBoost ensemble learning model.

5. The machine learning-based plasma protein biomarker colorectal cancer risk prediction system according to claim 4, characterized in that, The model training module is configured to perform the following operations: A stratified sampling method was used to divide the dataset containing the combined expression levels of the key protein biomarkers and disease labels into a training set and a validation set according to a set ratio; the training data was constructed using the xgb.DMatrix format. An adaptive cross-validation and grid search strategy is adopted to automatically search for and determine the optimal hyperparameter combination with the goal of minimizing log loss. The optimal hyperparameter combination is used to train an XGBoost binary classification model to obtain the trained risk assessment model. The trained model is used to predict the risk probability of the validation set samples, and the model performance is evaluated by the area under the receiver operating characteristic curve.

6. The machine learning-based plasma protein biomarker colorectal cancer risk prediction system according to claim 5, characterized in that, The search range for the optimal hyperparameter combination includes the number of iterations, maximum tree depth, learning rate, minimum loss reduction for node splits, feature subsampling ratio, minimum weight of leaf nodes, and sample subsampling ratio.

7. The machine learning-based plasma protein biomarker colorectal cancer risk prediction system according to claim 1, characterized in that, The risk prediction and assessment module is further configured to: standardize the plasma protein expression data of the individual to be assessed using the mean and standard deviation of the key protein markers in the training set using Z-score, and then input the data into the trained risk assessment model.

8. The machine learning-based plasma protein biomarker colorectal cancer risk prediction system according to claim 1, characterized in that, The proteomics workflow includes, in sequence: plasma protein extraction, peptide digestion, data-dependent acquisition and library construction, chromatographic fractionation, liquid chromatography-tandem mass spectrometry data acquisition, and database retrieval, to obtain the plasma protein expression matrix for each sample.

9. A method for constructing a colorectal cancer risk prediction model based on machine learning plasma protein biomarkers, applied to the prediction system as described in any one of claims 1-8, characterized in that, include: Plasma protein expression data of samples were obtained through a proteomics workflow; The plasma protein expression data were preprocessed, and differential expression analysis was performed using the limma algorithm to screen for differentially expressed proteins associated with colorectal cancer. The LASSO regression algorithm was used to extract key protein biomarker combinations from the differentially expressed proteins; the expression levels of the key protein biomarker combinations and the disease labels of the corresponding samples were used as inputs to train a risk assessment model using the XGBoost ensemble learning algorithm.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method for constructing a colorectal cancer risk prediction model based on machine learning as described in any one of claims 1-8.