Biomarker combination, kit, system and application thereof for predicting liver metastasis of colorectal cancer

By screening out combinations based on plasma protein biomarkers and combining machine learning models, kits and systems for predicting liver metastasis risks in colorectal cancer have been developed, which solves the problem of difficult to effectively predict liver metastasis in colorectal cancer in the prior art, and achieves accurate prediction of liver metastasis risks in colorectal cancer patients and individualized treatment decisions.

CN118465282BActive Publication Date: 2025-05-06WESTLAKE LAB OF LIFE SCI & BIOMEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410618368.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-05-06
Estimated Expiration
2044-05-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict whether there is liver metastasis in patients with colorectal cancer, especially the lack of specificity in early diagnosis, which leads to difficult treatment decision-making.

Method used

By screening out combinations of plasma protein biomarkers, including P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B, kits and systems for predicting liver metastasis risks in colorectal cancer were developed in combination with machine learning models (such as random forest algorithms) and clinical features.

Benefits of technology

Accurate prediction of the risk of liver metastasis in patients with colorectal cancer is achieved, unnecessary chemotherapy side effects in low-risk patients, and improved the efficiency and accuracy of the diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118465282B_ABST
    Figure CN118465282B_ABST
Patent Text Reader

Abstract

This invention provides a combination of biomarkers, kits, systems, and applications for predicting liver metastasis in colorectal cancer. The invention establishes a model based on six biomarkers (six protein features) and a combined model of six biomarkers plus three clinical features to predict the risk parameters of liver metastasis in colorectal cancer patients. The areas under the AUC curves for both models reach 0.899 and 0.904, respectively, with a negative predictive value approaching 95%. Therefore, this invention can efficiently, accurately, cost-effectively, and early predict whether colorectal cancer patients have a risk of liver metastasis while precisely excluding cases with low liver metastasis risk, thereby effectively guiding clinical decision-making and forming personalized precision treatment plans.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biomedicine, and specifically, relates to a biomarker combination, a kit, a system and applications thereof for predicting colorectal cancer liver metastasis. Background Art

[0002] Colorectal cancer (CRC) is the third most common cancer in the world and the second most deadly cancer (according to the WHO). Its survival rate is closely related to the diagnosis period of colorectal cancer and whether metastasis occurs. The main target organ of hematogenous metastasis of colorectal cancer is the liver, and liver metastasis is also the main cause of death in colorectal cancer patients. Therefore, early diagnosis and treatment of liver metastasis is one of the key aspects to improve the prognosis of colorectal cancer.

[0003] At present, clinical diagnosis of colorectal cancer patients with liver metastasis is usually performed by regular postoperative review of serum CEA, CA19-9 and other tumor markers to indicate the possibility of recurrence or distant metastasis, and routine preoperative or regular postoperative review of liver ultrasound, enhanced MRI and abdominal enhanced CT and other imaging examinations to find evidence of liver imaging metastasis. If it is still impossible to judge, serum AFP, liver ultrasound angiography, liver cell-specific contrast agent enhanced MRI examination can be added, or the diagnosis can be confirmed when liver biopsy puncture is performed in patients with liver metastases. In addition, colorectal cancer has been carried out related gene testing, immunohistochemistry, etc., including RAS, BRAF testing, mismatch repair gene MMR, microsatellite instability MSI testing, UGT1A1 testing, HER2 testing, etc., to provide a basis for clinical decision-making for patient treatment. However, the above is mainly related to the medication and prognosis of colorectal cancer, and is not a "liver metastasis-related" gene test.

[0004] However, CEA and CA19-9 lack specificity for early diagnosis of colorectal cancer liver metastasis, and gene detection can only provide reference for prognosis and medication decisions after the disease is diagnosed to a certain extent. In addition, existing gene detection and microRNA markers such as miR-6084 are still in the substantive examination stage. As for proteins that best reflect the functional status of the human body, protein markers related to colorectal cancer liver metastasis are still relatively rare.

[0005] Therefore, there is an urgent need for a technical means to combine protein markers to detect whether colorectal cancer patients have liver metastasis to guide the clinical diagnosis and treatment process. Summary of the invention

[0006] One object of the present invention is to provide a use of a reagent for detecting the expression level of a biomarker in the preparation of a product for predicting the risk of liver metastasis in a subject with colorectal cancer.

[0007] Another object of the present invention is to provide a biomarker combination for predicting the risk of liver metastasis in a subject with colorectal cancer.

[0008] Still another object of the present invention is to provide a kit for predicting the risk of liver metastasis in a subject with colorectal cancer.

[0009] It is yet another object of the present invention to provide two different systems for predicting the risk of liver metastasis in a subject with colorectal cancer.

[0010] It is yet another object of the present invention to provide two different modeling approaches for predicting the risk of liver metastasis in subjects with colorectal cancer.

[0011] It is intended to provide an application, biomarker combination, kit, prediction system and modeling method as described above, to screen patients with liver metastasis of colorectal cancer in a relatively non-invasive manner only through the blood markers of the subjects, or blood markers combined with clinical indicators, so as to implement a more precise treatment policy, including accurately excluding patients with low risk of liver metastasis, and for some patients with low risk of liver metastasis of T2-T3, for whom the necessity of postoperative adjuvant chemotherapy is still controversial, postoperative adjuvant chemotherapy can be prudently exempted in clinical decision-making, and the prediction indicators and other clinical review items provided by the present invention can be regularly monitored to avoid unnecessary chemotherapy side effects for these patients, and even damage the body's original immune barrier. In addition, the present invention successfully screened out plasma-based protein biomarkers that can accurately predict the risk of liver metastasis of colorectal cancer, which will also help explore the pathogenesis of this disease and is of great significance for the efficient, accurate and low-cost diagnosis of patients with liver metastasis of colorectal cancer.

[0012] In a first aspect, the present invention provides an application of a reagent for detecting the expression level of biomarkers in the preparation of a product for predicting the risk of liver metastasis in subjects with colorectal cancer, wherein the biomarkers include P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B.

[0013] Information on the biomarkers of the present invention can be found in the Uniprot database (https: / / www.uniprot.org).

[0014] In a second aspect, the present invention provides a kit for predicting the risk of liver metastasis in subjects with colorectal cancer, comprising reagents for detecting the expression levels of biomarkers, wherein the biomarkers include P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B.

[0015] In a third aspect, the present invention provides a system for predicting the risk of liver metastasis of a subject with colorectal cancer, the system comprising a processor and a display. The processor is configured to: obtain the expression levels of the following six biomarkers of the subject: P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1, and P05160_F13B; based on the expression levels of the six biomarkers obtained, use a first machine learning model to predict the risk parameters of colorectal cancer liver metastasis in the subject; and cause the display to present the predicted risk parameters of colorectal cancer liver metastasis in the subject.

[0016] In some embodiments, in the system provided in the third aspect of the present invention, the first machine learning model is constructed based on the random forest algorithm with the expression levels of the six biomarkers as feature information.

[0017] In a fourth aspect, the present invention provides another system for predicting the risk of liver metastasis of a subject with colorectal cancer, the system comprising a processor and a display. The processor is configured to: obtain the expression levels of the following six biomarkers of the subject: P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B; obtain the clinical characteristics of the subject including CEA, CA19-9 and age; based on the expression levels of the acquired six biomarkers and the expression of clinical characteristics, use a second machine learning model to predict the risk parameters of colorectal cancer liver metastasis in the subject; and cause the display to present the predicted risk parameters of colorectal cancer liver metastasis in the subject.

[0018] In some embodiments, in the system provided in the fourth aspect of the present invention, the second machine learning model is obtained by screening the expression levels of the six biomarkers and clinical characteristics including gender, age, CEA, CA19-9 and CA125 based on the random forest algorithm, and is constructed with the expression levels of the six biomarkers and the expression amounts of clinical characteristics including CEA, CA19-9 and age as feature information.

[0019] In some embodiments, in the system provided by the third aspect or the fourth aspect of the present invention, the processor is further configured to: predict the risk parameters of colorectal cancer liver metastasis in the subject before imaging reveals that the subject has liver metastasis, or when the subject has no other known clinical evidence of liver metastasis, and based on the predicted risk parameters of colorectal cancer liver metastasis in the subject, enable the display to present the correspondence between the predicted risk parameters and at least one preset threshold value, as well as the corresponding diagnosis and treatment recommendations when the predicted risk parameters are higher than the preset threshold value, wherein close monitoring of liver metastasis is recommended when the predicted risk parameters are higher than the first threshold value.

[0020] In some embodiments, the system provided in the third aspect or the fourth aspect of the present invention also includes a liquid chromatography device and a mass spectrometry analysis device, which is configured to: obtain a sample of the plasma collected from the subject, analyze and quantify it through targeted proteomics, and detect the expression levels of 6 biomarkers therefrom.

[0021] In a fifth aspect, the present invention provides a modeling method for predicting the risk of liver metastasis in a subject with colorectal cancer, the modeling method comprising the following steps. A group of plasma proteome data including colorectal cancer patients without liver metastasis and healthy people is obtained as a non-metastasis sample group, and a group of plasma proteome data of colorectal cancer patients with liver metastasis is obtained as a metastasis sample group, wherein each sample in the non-metastasis sample group and the metastasis sample group includes the expression level of each protein feature in the plasma proteome and the annotation information of whether the corresponding patient has colorectal cancer liver metastasis; the non-metastasis sample group and the metastasis sample group are divided into a training set and a test set; based on the training set, the features are screened using t-test and P value, and each protein feature in the plasma proteome that meets the preset expression level difference conditions is determined as the optimal marker combination; based on the expression level of the optimal marker combination of each sample in the training set, a random forest algorithm is used to construct and train an application prediction model, and the application prediction model is tested based on the test set to evaluate the prediction performance, and the application prediction model whose prediction performance exceeds the performance index threshold is used to give a prediction result about the risk of liver metastasis based on the expression level of the optimal marker combination in the plasma of subjects with colorectal cancer.

[0022] In a sixth aspect, the present invention provides another modeling method for predicting the risk of liver metastasis in subjects with colorectal cancer, the modeling method comprising the following steps. Obtain a group of plasma proteome data including colorectal cancer patients without liver metastasis and healthy people as a non-metastasis sample group, and obtain a group of plasma proteome data of colorectal cancer patients with liver metastasis as a metastasis sample group, wherein each sample in the non-metastasis sample group and the metastasis sample group includes the expression level of each protein feature in the plasma proteome and the annotation information of whether the corresponding patient has colorectal cancer liver metastasis; obtain clinical characteristic information of the patient corresponding to each sample in the non-metastasis sample group and the metastasis sample group, wherein the clinical characteristic information includes at least gender, age, CEA, CA19-9 and CA125; divide the non-metastasis sample group and the metastasis sample group into a training set and a test set; based on the training set, use t-test and P The invention relates to a method for screening features based on the value of the screening feature, and determining each protein feature in the plasma proteome whose expression level difference meets the preset conditions as the optimal marker combination; constructing and training an application prediction model using a random forest algorithm based on the expression level and clinical feature information of the optimal marker combination of each sample in the training set, and determining the optimal clinical feature combination based on the feature importance of each clinical feature in the predictive performance of the application prediction model; testing the application prediction model based on the test set to evaluate the predictive performance, and using the application prediction model whose predictive performance exceeds the performance index threshold to provide a prediction result on the risk of liver metastasis based on the expression level of the optimal marker combination in the plasma of subjects with colorectal cancer and the expression amount of the optimal clinical feature combination.

[0023] In some embodiments, in the modeling method provided in the fifth aspect or the sixth aspect of the present invention, the preset conditions include: the P value is less than 0.05.

[0024] In some embodiments, in the modeling method provided in the fifth aspect or the sixth aspect of the present invention, the optimal marker combination includes the following 6 protein features: P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B.

[0025] In some embodiments, in the modeling method provided in the fifth aspect or the sixth aspect of the present invention, the optimal clinical feature combination includes CEA, CA19-9 and age.

[0026] The present invention establishes an association between 6 protein features (or 6 protein features combined with 3 clinical features) and the risk of liver metastasis in subjects with colorectal cancer and a corresponding prediction model. The acquisition of these protein feature information and clinical feature information is low-invasive for the subject, wherein the protein feature information is obtained by using the subject's plasma sample (usually collected and prepared in other medical procedures for colorectal cancer) by means of proteomics analysis, SWATH mass spectrometry analysis, etc., and the clinical features should have been obtained during the subject's standardized diagnosis and treatment and follow-up process. Therefore, the feature information used in the present invention has very good compatibility with existing colorectal cancer-related medical procedures. In addition, the negative predictive value of the prediction model of the present invention is close to 95%, thereby being able to efficiently and accurately predict whether liver metastasis occurs in patients with colorectal cancer, especially being able to accurately exclude situations with low risk of liver metastasis, reduce unnecessary treatment and waste of medical resources, and alleviate the psychological pressure of these patients on metastasis. Therefore, the present invention can conduct efficient, accurate and low-cost screening for colorectal cancer patients as early as possible without the need for additional invasive examinations, thereby predicting the risk of liver metastasis in patients, guiding clinical decision-making, and forming an individualized and accurate diagnosis and treatment plan. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In the drawings, which are not necessarily drawn to scale, the same reference numerals may describe similar parts in different views. The same reference numerals with letter suffixes or different letter suffixes may represent different instances of similar parts. The accompanying drawings generally illustrate various embodiments by way of example and not limitation, and together with the specification and claims, are used to illustrate the disclosed embodiments. When appropriate, the same reference numerals are used throughout the drawings to refer to the same or similar parts. Such embodiments are illustrative and are not intended to be exhaustive or exclusive embodiments of the present apparatus or method.

[0028] Figure 1 A schematic diagram showing some components of a system for predicting the risk of liver metastasis in a subject with colorectal cancer according to an embodiment of the present invention.

[0029] Figure 2 A schematic diagram showing part of the components of another system for predicting the risk of liver metastasis in a subject with colorectal cancer according to an embodiment of the present invention.

[0030] Figure 3 A flow chart showing a modeling method for predicting the risk of liver metastasis in a subject with colorectal cancer according to an embodiment of the present invention.

[0031] Figure 4 A flow chart showing another modeling method for predicting the risk of liver metastasis in a subject with colorectal cancer according to an embodiment of the present invention.

[0032] Figure 5 A schematic diagram showing the importance of six protein features as the optimal marker combination screened according to an embodiment of the present invention.

[0033] Figure 6 A schematic diagram showing the importance of 6 protein features and clinical features screened according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0034] The terms used herein are only used to explain the specific embodiments and are not used to limit the present invention. A singular expression includes its plural expression unless it is explicitly stated or it is obvious from the context that it is not intended to do so.

[0035] Although various modifications can be made to the present invention and the present invention can have various forms, specific examples will be described in detail and explained below. However, it should be understood that these are not intended to limit the present invention to specific disclosures, and the present invention includes all modifications, equivalents or substitutions thereto without departing from the spirit and technical scope of the present invention.

[0036] The present invention will be further described in detail below through examples, but the scope of the present invention is not limited to the examples.

[0037] In the first aspect, the present invention provides an application of a reagent for detecting the expression level of biomarkers in the preparation of a product for predicting the risk of liver metastasis in subjects with colorectal cancer, wherein the biomarkers include P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B.

[0038] In the present invention, the above biomarkers are all proteins. In the names of the above biomarkers, the characters before "_" are the Uniprot Accession ID of the protein, and the characters after "_" are the gene name of the protein.

[0039] In a second aspect, the present invention provides a kit for predicting the risk of liver metastasis in a subject with colorectal cancer, comprising reagents for detecting the expression levels of biomarkers, wherein the biomarkers include P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B.

[0040] Figure 1 FIG. 2 is a schematic diagram showing a partial composition of a system for predicting the risk of liver metastasis of a subject with colorectal cancer according to an embodiment of the present invention. Figure 1As shown, system 100 includes a processor 101 and a display 102. The processor 101 can be, for example, a processing device including one or more general processing devices, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor 101 can also be one or more special processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a system on a chip (SoC), etc. The display 102 can use LED, OLED, etc., which will not be repeated here.

[0041] In some embodiments, the processor 101 can be configured to obtain the expression levels of the following six biomarkers of the subject 103: P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B.

[0042] In some embodiments, the processor 101 can also be configured to: predict the risk parameters of colorectal cancer liver metastasis of subject 103 based on the expression levels of the acquired 6 biomarkers using the first machine learning model; and enable the display 102 to present the predicted risk parameters of colorectal cancer liver metastasis of the subject, wherein the risk parameters include but are not limited to being expressed as a percentage, such as "risk of colorectal cancer liver metastasis: 90%", etc. The actual risk parameter can also be other numerical ranges, and the content and method of presentation can also be set as needed, which will not be repeated here. As an example only, the first machine learning model in the processor 101 can be constructed based on the random forest algorithm with the expression levels of the 6 biomarkers as feature information, and the specific model construction method will be described in detail below. In addition, the first machine learning model can also be constructed based on other algorithms such as neural networks, decision trees, support vector machines, etc., using the expression levels of the above 6 biomarkers as feature information, and the specific method is not limited by the present invention.

[0043] In some embodiments, the processor 101 is further configured to: predict risk parameters for colorectal cancer liver metastasis in the subject before imaging reveals that the subject has liver metastasis, or when the subject has no other known clinical evidence of liver metastasis, and based on the predicted risk parameters for colorectal cancer liver metastasis in the subject, enable the display to present a correspondence between the predicted risk parameters and at least one preset threshold, as well as corresponding diagnosis and treatment recommendations when the predicted risk parameters are higher than the preset threshold, wherein close monitoring of liver metastasis is recommended when the predicted risk parameters are higher than the first threshold.

[0044] Among them, close monitoring can further include: 1) shortening the time interval for postoperative review of colorectal cancer patients who have not yet developed liver metastasis, so as to detect possible liver metastasis early; 2) using more sophisticated imaging methods to check whether the patient has micro-metastases in the liver that may be missed in routine preoperative examinations before primary surgery. If micro-metastases are found, the patient's liver metastases can be reassessed to see whether they are initially resectable to avoid missing the opportunity for concurrent surgery, or directly perform primary surgery for patients who are unresectable and should undergo neoadjuvant therapy; 3) For patients with early clinical stage (such as stage I) and no high-risk recurrence factors (such as insufficient number of lymph nodes dissected), that is, patients who originally had no adjuvant treatment plan after primary surgery, if the predicted risk parameter is higher than the first threshold, it is recommended that the clinic consider the necessity of postoperative adjuvant chemotherapy.

[0045] As an example only, two risk parameter thresholds, low and high, can be set based on the clinical experience of medical workers. When the predicted risk parameter is lower than the low threshold, it can be considered that the risk of colorectal cancer liver metastasis in the subject is very low, and the reexamination time interval can be specified according to clinical follow-up experience, and some patients (for example, T3N0M0 patients without other high-risk recurrence factors, which means that the cancer has a TNM stage of 3, a large primary tumor, but no lymph node involvement and distant metastasis) can be exempted from postoperative adjuvant chemotherapy, etc.; when the predicted risk parameter is between the low threshold and the high threshold, other medical means should be combined for follow-up. Further investigation; when the predicted risk parameter is higher than the high threshold (corresponding to the above-mentioned first threshold, the specific value can be set to 50%, for example), it means that the risk of colorectal cancer liver metastasis in the subject is extremely high. The prediction results are for the reference of doctors, so that they can judge whether it is necessary to conduct detailed and other risky reexamination items in advance (for example, the original setting of reexamination of abdominal enhanced CT in 2 months can be advanced to 1 month, and consider adding liver enhanced MRI); for abnormal lesions that cannot exclude metastasis when the image reexamination is found during the reexamination, it can also help clinicians make a tendency diagnosis of malignant considerations if it is higher than the high threshold. The present invention does not specifically limit the specific setting of the threshold and the specific application of the prediction results by medical workers.

[0046] In some embodiments, the system 100 of the present invention may also include a liquid chromatography device and a mass spectrometry device (not shown), which is configured to obtain a sample of plasma collected from a subject, analyze and quantify by targeted proteomics (MRMHR), and detect the expression levels of 6 biomarkers therefrom. For example, for example, a liquid phase ekspertTMnanoLC 415system (Eksigent, Dublin, CA, USA) and a mass spectrometer TripleTOF 5600-MS / MS (SCIEX, CA, USA) may be used to perform targeted proteomics analysis and quantification. For example, a specific proteomics analysis method may be performed as follows:

[0047] Step 1: Experimental batch design and quality control sample preparation

[0048] Experimental design is performed for clinical cohort samples, taking into account the different clinical characteristics of the population, and using any applicable proteomic big data design and quality control tools to make the samples as evenly dispersed as possible to minimize batch effects. In addition, QC (quality control) samples can be prepared while processing samples, that is, mixed samples of equal volume of experimental samples, and the system stability can be evaluated throughout the experiment.

[0049] Step 2: Proteomic analysis

[0050] Plasma samples were denatured in lysis buffer containing 20 μL of 8M urea (dissolved in 100 mM ABB (Ammonium bicarbonate, NH4HCO3)), and then incubated at 32°C for 45 minutes with 10 mM tris(2-carboxyethyl)phosphine (TCEP), followed by reduction and alkylation of protein lysates in 40 mM iodoacetamide (IAA) in the dark for 60 minutes. After further dilution with 76 μL of 100 mM ABB, the samples were digested at 32°C for 4 hours at a ratio of 1:50 for trypsin: substrate, followed by another digestion at 32°C for 12 hours with the same ratio of trypsin. The reaction was stopped by adding 10% trifluoroacetic acid (TFA) to a final concentration of 1%. The digested peptides were washed and desalted using a desalting column, then spun dry, re-dissolved, and diluted to a uniform concentration for subsequent mass spectrometry analysis.

[0051] Step 3: SWATH mass spectrometry and data analysis

[0052] Each sample was analyzed online using a liquid phase ekspertTM nanoLC 415 system (Eksigent, Dublin, CA, USA) and TripleTOF5600-MS / MS (SCIEX, CA, USA) in data-independent acquisition (DIA, SWATH) mode.

[0053] Specifically, the sample was first loaded onto a preloaded column (3 μm, 10mm*0.3mmi.d.), and then the sample loaded on the preloaded column was flushed to the analytical column (3μm, 150mm*0.3mm), the analysis time was 20 minutes, the LC gradient was 5% to 32% buffer B (buffer A was 2% ACN (acetonitrile), 98% H2O (containing 0.1% FA (formic acid)), buffer B was 98% ACN (containing 0.1% FA). All reagents were MS grade. In terms of mass spectrometry parameters, the range of TOF MS was 350-1250m / z. The precursor ion was subjected to secondary fragmentation of MS / MS, which ranged from 100-1500m / z. The accumulation time of each isolation window was set to 30ms, and the total time period was 1.9s.

[0054] Next, data analysis was performed. Specifically, mass spectrometry data can be analyzed using DIA-NN (version 1.8) and the protein spectral library (PMID: 34783559). The enzyme was set to trypsin with two missing cleavage tolerances. The static modification was cysteine ​​aminomethylation (+57.021464), and the variable modification was methionine oxidation (+15.994915). The peptide length range, parent ion m / z range, and fragment ion m / z range were set to 7–30, 300–1800, and 200–1800, respectively. A robust LC (high accuracy) quantitative strategy and RT-dependent run normalization were used, with a parent ion false discovery rate (FDR) of 1%. In addition, different sample search libraries were set to be unrelated, and MBR (Match between runs) was not used. The default parameters were used for other parameters. The mass spectrometry data of each protein were obtained according to the above method.

[0055] In addition, quality control analysis was performed. From the correlation data of plasma proteome QC samples, it can be seen that the median correlation of plasma proteome QC samples is 98.42%, indicating that the experimental data has high consistency and repeatability. Next, PCA dimensionality reduction analysis of plasma proteome data found that there was no batch effect between different batches after correction.

[0056] Then, protein difference analysis is performed. For example, the condition P value can be set to be less than 0.05 to screen the differential proteins between patients with colorectal cancer liver metastasis and colorectal cancer patients without liver metastasis, and the screened proteins are subjected to the next step of targeted verification.

[0057] Step 4: MRM-HR (MRM high resolution) targeted proteome analysis

[0058] Peptides with unsatisfactory spectra were removed, and 81 peptides corresponding to the screened differential proteins were subjected to targeted proteomic detection, with the liquid phase gradient and SWATH acquisition consistent. Targeted mass spectrometry data were analyzed using Skyline, and retention times were predicted using CiRT. Both MS1 and MS2 were set to "TOF mass spectrometer", with resolutions of 30,000 and 15,000, respectively, and the "target" acquisition method was defined in MS / MS. Quality control analysis was then performed, and the specific method was similar to the quality control analysis in step three.

[0059] Based on steps one to four above, the expression levels of various proteins serving as biomarkers can be detected from the plasma samples of the subjects, and the differential proteomes of patients with colorectal cancer liver metastasis and colorectal cancer patients without liver metastasis can be screened out, thereby obtaining the expression levels (i.e., expression amounts) of the six protein features serving as biomarkers as described above.

[0060] Figure 2 A schematic diagram of the partial components of another system for predicting the risk of liver metastasis in a subject with colorectal cancer according to an embodiment of the present invention is shown. Specifically, Figure 2 Corresponding to the system described in the fourth aspect of the present invention.

[0061] like Figure 2 As shown, the system 200 includes a processor 201 and a display 202. Figure 1 Similar to the processor 101 in , the processor 201 may also be a processing device including one or more general-purpose processing devices, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor 101 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a system on a chip (SoC), etc. The display 102 may use LED, OLED, etc., which will not be described in detail here.

[0062] In some embodiments, the processor 201 can be configured to not only obtain the expression levels of the six biomarkers P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B of the subject 103, but also obtain the clinical characteristic expression of the subject 103 including CEA, CA19-9 and age. For CEA and CA19-9, their expression levels can be the expression level values ​​obtained by inspection and measurement, or they can be the expression levels after normalization processing, etc., and the present invention does not limit this. For age, a specific age value can be directly used, or the age can be divided into stages according to certain rules, such as youth, middle-aged, elderly, etc., and the present invention does not limit this.

[0063] Among them, CEA (Carcinoembryonic Antigen) is a protein and also an early discovered tumor marker, commonly used in the diagnosis and monitoring of lung cancer, gynecological tumors, thyroid cancer, colorectal cancer and other digestive system tumors. An increase in CEA can indicate the possibility of a tumor, but it cannot determine the specific type of tumor.

[0064] CA19-9 (Carbohydrate antigen 19-9, a carbohydrate antigen) is an oligosaccharide tumor-associated antigen and a new tumor marker. It is a glycolipid on the cell membrane with a molecular weight greater than 1000kD. It is the most sensitive marker for pancreatic cancer reported so far. It exists in the form of salivary mucin in serum and is distributed in the pancreas, gallbladder, liver, intestine of normal fetuses and the pancreas and bile duct epithelium of normal adults. It is a gastrointestinal tumor-associated antigen that exists in the blood circulation.

[0065] Further, the processor 201 is also configured to predict the risk parameters of colorectal cancer liver metastasis in the subject 103 based on the expression levels of the acquired 6 biomarkers and the expression amounts of clinical characteristics using the second machine learning model, and to enable the display 202 to present the predicted risk parameters of colorectal cancer liver metastasis in the subject 103. The specific content and method of presentation can be set as needed and will not be described in detail here.

[0066] and Figure 1 Similar to the system 100 in FIG. 1 , the system 200 may also include a liquid chromatography device and a mass spectrometry device 204, which are configured to obtain a sample of the collected plasma of the subject, perform targeted proteomics analysis and quantification, and detect the expression levels of the six biomarkers therefrom. Figure 1 An example has been given in , so I will not repeat it here.

[0067] Similar to processor 101, processor 201 can also be further configured to: predict the risk parameters of colorectal cancer liver metastasis in subject 103 before imaging finds that subject 103 has liver metastasis, or when subject 103 has no other known clinical evidence of liver metastasis, and based on the predicted risk parameters of colorectal cancer liver metastasis in subject 103, enable the display 202 to present the corresponding relationship between the predicted risk parameters and at least one preset threshold value, as well as the corresponding diagnosis and treatment recommendations when the predicted risk parameters are higher than the preset threshold value, wherein when the predicted risk parameters are higher than the first threshold value, close monitoring of liver metastasis is recommended, and specific measures for close monitoring of liver metastasis have been combined Figure 1 A detailed description is not given here.

[0068] In some embodiments, the system 200 may also include a liquid chromatography device and a mass spectrometry device 204, which are configured to obtain a sample of the plasma collected from the subject 103, perform targeted proteomics analysis and quantification, and detect the expression levels of 6 biomarkers therefrom.

[0069] In some embodiments, the second machine learning model in the processor 202 is obtained by screening the expression levels of the six biomarkers and clinical characteristics including gender, age, CEA, CA19-9 and CA125 based on the random forest algorithm, and is constructed with the expression levels of the six biomarkers and the expression amounts of clinical characteristics including CEA, CA19-9 and age as feature information.

[0070] In the present invention, the above 6 biomarkers are all proteins, and currently clinically it is believed that the main functions of various proteins are as follows.

[0071] "AFM" ("P43652|AFM_HUMAN"), as a carrier of hydrophobic molecules in body fluids, is essential for the solubility and activity of Wnt family members and may transport vitamin E in body fluids. Currently, there is no direct article reporting association with colorectal cancer liver metastasis. Therefore, the present invention successfully identified the correlation between this protein and colorectal cancer liver metastasis without the aid of prior knowledge, and through more in-depth analysis, it is believed that this correlation is likely to be caused by the fact that the protein encoded by this gene is developmentally regulated, expressed in the liver and secreted into the blood.

[0072] "CP" ("P00450|CP_HUMAN"), this protein is a metalloprotein that can bind most of the copper in plasma and participate in the peroxidation of Fe(II) transferrin to Fe(III) transferrin. This protein is a potential prognostic marker for cholangiocarcinoma. Insufficient CP expression is associated with a poor prognosis in adrenocortical carcinoma (ACC). At the same time, ceruloplasmin is a potential (prognostic) marker for pancreatic ductal adenocarcinoma (PDAC) in CA19-9 negative patients. There is currently no report on its association with colorectal cancer liver metastasis. The present invention successfully identified the correlation between this protein and colorectal cancer liver metastasis, which was not revealed by other prior arts before the present invention, and the effect was unexpected.

[0073] "AZGP1" ("P25311|AZGP1_HUMAN"), i.e., P25311_AZGP1, has been reported in the literature to be associated with cancer diagnosis and prognosis (PMID: 31632499). It was found that AZGP1 was more highly expressed in colorectal cancer tissues with liver metastasis than in colorectal cancer tissues without metastasis, and that high expression of AZGP1 was associated with poor prognosis. The present invention is consistent with this cognition.

[0074] "MDGA2" ("Q7Z553|MDGA2_HUMAN"), this protein is involved in the regulation of presynaptic assembly, the regulation of synaptic membrane adhesion and the differentiation of spinal motor neurons. There is no report related to liver metastasis of colorectal cancer. Therefore, the correlation between this protein and liver metastasis of colorectal cancer is not revealed by other prior arts before the present invention, and the effect is unexpected.

[0075] "APOA1" ("P02647|APOA1_HUMAN"), the protein is apolipoprotein A1, apolipoprotein is the main protein component of high-density lipoprotein (HDL) in plasma. The preproprotein is proteolytically processed to produce a mature protein, which promotes the efflux of cholesterol from tissues to the liver for excretion and is a cofactor of lecithin cholesterol acyltransferase (LCAT). There are currently no reports related to liver metastasis of colorectal cancer, therefore, the correlation between this protein and liver metastasis of colorectal cancer is not revealed by other prior arts before the present invention, and the effect is also unexpected.

[0076] "F13B" ("P05160|F13B_HUMAN"), which is the B subunit of coagulation factor XIII. Coagulation factor XIII is the last zymogen activated in the blood coagulation cascade. There are reports in the literature that the activation peptide of coagulation factor XIII (AP-F13A1) can be used as a new biomarker for screening colorectal cancer (PMID: 29657559). Without relying on this prior knowledge, the present invention independently identifies the correlation between this protein and liver metastasis of colorectal cancer.

[0077] Table 1 below shows the performance metrics of the above systems 100 and 200 on the test set. Among them, system 100 only uses 6 protein features as the optimal biomarkers, while system 200 uses 6 biomarkers and clinical features to predict whether subjects with colorectal cancer have liver metastasis.

[0078] Table 1 Performance Metrics of Systems 100 and 200 on the Test Set

[0079]

[0080] The meanings and calculation methods of each performance metric in Table 1 are as follows:

[0081] Accuracy (Acc): It represents the proportion of the number of samples correctly predicted by the system in the total number of samples. The specific calculation method is as follows:

[0082]

[0083] Among them, TP is the true positive, TN is the true negative, FP is the false positive, and FN is the false negative.

[0084] Area Under Curve (AUC): It is the area covered under the ROC (Receiver Operating Characteristic) curve, which is a reasonable criterion for comparing the prediction performance of the system. The AUC value is not greater than 1, and the larger the value, the better the system performance. Generally, when 0.5 < AUC < 1, it indicates that the system prediction performance is better than random prediction.

[0085] Sensitivity (Sensitivity or Recall): It represents the proportion of samples that are actually positive and are correctly predicted as positive by the system. It focuses on the coverage of positive samples by the system. The specific calculation method is as follows:

[0086]

[0087] Specificity (TNR): It indicates the proportion of samples that are actually counterexamples that are correctly predicted by the system as counterexamples. The specific calculation method is as follows:

[0088]

[0089] PPV Positive Predictive Value (PPV): It indicates the proportion of samples that are actually positive (true positive) among the samples predicted by the system as positive. It focuses on the possibility that the subjects predicted as positive are actually positive. The higher the value, the more reliable the samples predicted by the system as positive. The specific calculation method is as follows:

[0090]

[0091] NPV Negative Predictive Value (NPV): It indicates the proportion of samples that are actually negative examples (true negatives) among the samples that are predicted as negative examples by the system. It focuses on the possibility that the subjects predicted as negative examples are actually negative examples. The higher the value, the more reliable the samples predicted as negative examples by the system are. The specific calculation method is as follows:

[0092]

[0093] As can be seen from Table 1, whether it is system 100 or system 200, the area under the AUC curve is close to or exceeds 0.9, indicating that the prediction performance of both systems is very good, and is significantly better than the system using only the clinical feature CEA and the system using only the clinical feature CA19-9, wherein the area under the AUC curve of the system using only the clinical feature CEA is 0.781, and the area under the AUC curve of the system using only the clinical feature CA19-9 is 0.681. The prediction performance indicators of the above two systems using only clinical features cannot meet clinical applications well. In addition, the NPV negative predictive values ​​of the systems 100 and 200 of the present invention are both close to 0.95, indicating that when they are used to exclude low-risk colorectal cancer liver metastases, the prediction results are relatively accurate and reliable, thereby avoiding unnecessary waste of medical resources and patient side effects caused by excessive treatment and anxiety about possible metastases.

[0094] In some embodiments, plasma used to obtain the expression levels of six biomarkers of the subject can be obtained simultaneously when the subject undergoes a routine blood test, and the required amount of plasma is small, and the subject's blood loss or blood collection trauma is not increased. Clinical characteristics including CEA, CA19-9 and age can also be obtained during routine examinations without the need for additional blood collection. In other words, whether it is the expression level of the above six protein markers or the expression of the three clinical characteristics, they can be well compatible with the hospital's existing examinations of the subjects, and can effectively assist clinical indicators, to a certain extent reduce unnecessary enhanced CT and enhanced MRI examinations for low-risk patients and the number of chemotherapy for some low-risk patients, reduce the missed diagnosis of some tiny liver metastases, reduce the psychological pressure, contrast agent risks, radiation and body invasion of the subjects, reduce the workload of doctors, and are conducive to promotion in the medical and health system.

[0095] In addition, both CEA and CA19-9 are tumor markers, but in clinical practice, the increase in the values ​​of the two markers cannot be used alone to determine whether a tumor exists, and is generally used as an important indicator for monitoring postoperative recurrence and metastasis of tumor patients. Among them, the increase in CA19-9 is more often used as an important indicator of pancreatic cancer in clinical practice, and there is no report that the above two tumor markers are used as diagnostic indicators for colorectal cancer liver metastasis. Therefore, the inventor successfully revealed the association between these two tumor markers and colorectal cancer liver metastasis through experiments and modeling, especially when CEA, CA19-9 and age are combined with the 6 protein features in the present invention, the accuracy of colorectal cancer liver metastasis prediction and other performance indicators can be further improved, which is not revealed by other existing technologies before the present invention, and the effect is unexpected.

[0096] As mentioned above, the system of the present invention can predict the risk parameters of colorectal cancer liver metastasis in subjects based on the expression levels of 6 biomarkers using a first machine learning model, and the first machine learning model can be constructed based on a random forest algorithm using the expression levels of the 6 biomarkers as feature information. Therefore, the present invention also provides a modeling method for predicting the risk of liver metastasis in subjects with colorectal cancer. Figure 3 FIG. 4 is a flow chart showing a modeling method for predicting the risk of liver metastasis in a subject with colorectal cancer according to an embodiment of the present invention. Figure 3 As shown, the modeling method includes steps 301 to 304.

[0097] First, in step 301, a group of plasma proteome data of colorectal cancer patients without liver metastasis can be obtained as a non-metastasis sample group, and a group of plasma proteome data of colorectal cancer patients with liver metastasis can be obtained as a metastasis sample group, wherein each sample in the non-metastasis sample group and the metastasis sample group includes the expression level of each protein feature in the plasma proteome and the annotation information of whether the corresponding patient has colorectal cancer liver metastasis.

[0098] Next, in step 302, the non-transferred sample group and the transferred sample group are divided into a training set and a test set. As an example only, the samples in each sample group can be randomly divided into a training set and a test set, or a cross-validation method such as a five-fold or three-fold cross-validation method can be used to divide the samples into the training set and the test set, so that even when there are few samples, after the samples in the group are divided into the training set and the test set, problems such as overfitting of the model will not occur. The present invention does not limit the specific method of dividing the sample set.

[0099] Then, in step 303, based on the training set, t-tests such as Welch's t-test and P-values ​​are used to screen features, and protein features in the plasma proteome whose expression level differences meet preset conditions are determined as the optimal marker combination.

[0100] Figure 5 A schematic diagram showing the importance of six protein features as the optimal marker combination screened according to an embodiment of the present invention. Figure 5 The horizontal axis is the importance of the feature, and the vertical axis shows the importance of each protein feature (represented by its corresponding gene name) decreasing from top to bottom. Figure 5 It can be seen that in the trained application prediction model, the importance of the six proteins in predicting the risk of colorectal cancer liver metastasis is ranked.

[0101] After the optimal marker combination is screened out in the above manner, in step 304, a random forest algorithm is used to construct and train an application prediction model based on the expression level of the optimal marker combination of each sample in the training set.

[0102] The base learner of random forest is decision tree, and its essential attribute is ensemble learning method. By adding sample perturbations by sampling with replacement, as the number of learners increases, random forest will converge to a lower generalization error, improving the accuracy and stability of prediction. The following is an example of building and training an application prediction model based on random forest algorithm.

[0103] First, the data is cleaned and filled, and the data is randomly divided into training data and test data. Model training is performed in the training data, and parameter optimization is performed (for example, through the randomForest package (R language v4.7-1.1)), using a combination of random search and cross-validation. First, a random search is performed to explore various parameter combinations, with the number of trees ntree set to 100-2000 and the minimum size of the terminal node Minimum size of terminal node 1-10). Finally, the model with the best performance is selected, ntree=500, Minimum size of terminal node=1. This set of parameter configurations helps achieve a high degree of fit to the training data, and at the same time ensures that the model has good generalization ability through three-fold cross-validation.

[0104] The number of trees ntree in the optimal parameters is set to 500. Usually more trees will improve performance and make predictions more stable, but it will also slow down the calculation speed. Through testing, the optimal parameters are selected to achieve a balance between stability and speed. By building 500 independent decision trees, each tree will randomly extract samples and features for learning and construction. Due to the introduction of two randomness, random forests are not prone to overfitting (random samples, random features), and the introduction of this randomness helps to improve the generalization ability of the model.

[0105] At the same time, each internal node represents a "test" of an attribute (for example, whether colorectal cancer has metastasized or not), each branch represents the result of the test, and each leaf node represents a class label (decided after calculating all attributes). Nodes without child nodes are leaves. The minimum size of terminal node is 1, that is, each terminal node has at least one sample, and the model can explore the rules in the data more comprehensively.

[0106] The classification tree uses the Gini index or out-of-bag error to select the optimal feature and determine the optimal binary cut point of the feature, that is, how much the classification error rate can be reduced by dividing the data through this feature. The Gini coefficient can be used to evaluate the importance of each clinical and protein feature in predicting patients with colorectal cancer liver metastasis and those without metastasis (the importance can be ranked), and to select important clinical features and protein biomarkers related to liver metastasis, which is more conducive to improving the interpretability of the model and carrying out clinical transformation.

[0107] The calculation method of the Gini coefficient (GI) is shown in formula (1), where GI represents the Gini coefficient, C categories (i.e., colorectal cancer liver metastasis and no colorectal cancer liver metastasis), P qcIt represents the proportion of category c in node q. Intuitively speaking, it is the probability that two samples are randomly selected from node q and their category labels are inconsistent.

[0108]

[0109] After the above-mentioned application prediction model based on the random forest algorithm is trained, it is necessary to further test the application prediction model based on the test set to evaluate the prediction performance, and use the application prediction model whose prediction performance exceeds the performance index threshold to give a prediction result about the risk of liver metastasis based on the expression level of the optimal marker combination in the plasma of subjects with colorectal cancer. Among them, the above-mentioned performance index thresholds include but are not limited to the corresponding thresholds of the aforementioned performance indicators such as Acc, AUC, Sensitivity, TNR, PPV, NPV, etc., and can also be other performance indicators defined according to specific application scenarios, or composite performance indicators and their corresponding thresholds designed according to specific algorithms based on the above-mentioned various performance indicators, etc., and the present invention does not limit this. In particular, in the case where there are multiple application prediction models that meet the threshold requirements, the optimal application prediction model can be selected according to the prediction performance indicators concerned by different application scenarios. Therefore, through the test set test, it can be ensured that the trained application prediction model has strong robustness and generalization ability, which is not only applicable to samples in the training set, but also to new samples that have not participated in the training. It can be seen that according to Figure 3 The application prediction model constructed and trained by the method in can be used as the first machine learning model described in the aforementioned embodiment, and is used to predict the risk parameters of colorectal cancer liver metastasis in subjects based on the expression levels of 6 biomarkers.

[0110] In addition, the present invention also provides another modeling method for predicting the risk of liver metastasis in subjects with colorectal cancer, for example corresponding to a second machine learning model. Figure 4 FIG. 2 is a flow chart showing another modeling method for predicting the risk of liver metastasis in a subject with colorectal cancer according to an embodiment of the present invention. Figure 4 As shown, the modeling method includes steps 401 to 406.

[0111] First, in step 401, a group of plasma proteomic data of colorectal cancer patients without liver metastasis is obtained as a non-metastasis sample group, and a group of plasma proteomic data of colorectal cancer patients with liver metastasis is obtained as a metastasis sample group, wherein each sample in the non-metastasis sample group and the metastasis sample group includes the expression level of each protein feature in the plasma proteome and the annotation information of whether the corresponding patient has colorectal cancer liver metastasis.

[0112] Then, in step 402, the clinical characteristic expression of each sample corresponding to the patient in the non-metastasis sample group and the metastasis sample group is obtained, wherein the clinical characteristics at least include gender, age, CEA, CA19-9 and CA125. It is worth noting that there is no time sequence constraint between step 401 and step 402.

[0113] In step 403, the non-transfer sample group and the transfer sample group are divided into a training set and a test set. The specific method of dividing the training set and the test set is similar to the description in the above step 302, and will not be repeated here.

[0114] In step 404, based on the training set, t-test methods such as Welch's t-test and P value are used to screen features, and each protein feature in the plasma proteome that satisfies the preset expression level difference condition is determined as the optimal marker combination.

[0115] In step 405, based on the expression level and clinical feature information of the optimal marker combination of each sample in the training set, a random forest algorithm is used to construct and train an application prediction model, and the optimal clinical feature combination is determined based on the feature importance of each clinical feature in the prediction performance of the application prediction model. It is worth noting that when determining the optimal clinical feature combination, it is not based on isolated screening of each clinical feature alone, but the optimal marker combination determined in step 404 is used together for the construction and training of the application prediction model, thereby determining the optimal clinical feature combination when the optimal marker combination is combined with the clinical feature information. It can be seen that according to Figure 4 The application prediction model constructed and trained by the method in can be used as the second machine learning model described in the aforementioned embodiment, and is used to predict the risk parameters of colorectal cancer liver metastasis in subjects in a combined manner based on the expression levels of 6 biomarkers and the expression amounts of clinical characteristics.

[0116] How to use the random forest algorithm to model and train on the training set and combine Figure 3 The description is similar and will not be repeated here.

[0117] Figure 6 A schematic diagram showing the importance of 6 protein features and clinical features screened according to an embodiment of the present invention is shown. Figure 6 The horizontal axis is the importance of the feature, and the vertical axis shows the importance of each protein feature (indicated by its corresponding gene name) or clinical feature (Age indicates age) from top to bottom. Figure 6It can be seen that when the five clinical characteristics of gender, age, CEA, CA19-9 and CA125 are included in the model construction and training, according to the unified ranking of the importance of the six proteins and the five clinical characteristics, CEA, CA19-9 and age are finally selected as Figure 4 The applied prediction model constructed by the modeling method in was used for clinical characteristics of colorectal cancer liver metastasis risk prediction together with 6 protein features, while the other two clinical characteristics, gender and CA125 (not shown), were not included in the applied prediction model because of their lower importance in the unified ranking.

[0118] In step 406, the applied prediction model is tested based on the test set to evaluate the prediction performance, and the applied prediction model whose prediction performance exceeds a threshold is used to provide a prediction result on the risk of liver metastasis based on the expression level of the optimal marker combination and the expression amount of the optimal clinical feature combination in the plasma of the subject with colorectal cancer. Figure 4 The modeling method in this paper not only utilizes the six protein features, but also the clinical characteristics of the subjects. Figure 3 The applied prediction model shown here is distinguished from the one that only uses 6 protein features, which can be called a joint prediction model.

[0119] In addition, Figure 3 or Figure 4 In the modeling method in, the preset condition of the difference in expression level may include, for example, a P value less than 0.05, that is, P-value<0.05, where P-value is the P value. The specific calculation method is known to those skilled in the art and will not be described in detail here.

[0120] As an example only, follow Figure 3 or Figure 4 The optimal marker combination obtained by the modeling method includes the following 6 protein features: P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B.

[0121] In addition, Figure 4 In the modeling method, the optimal clinical feature combination obtained by screening includes CEA, CA19-9 and age. As an example only, multiple clinical features including but not limited to gender, age, CEA, CA19-9 and CA125 can be included as candidates, and the optimal clinical feature combination can be screened using the modeling method described in the embodiment of the present invention.

[0122] In proposing Figure 3 and Figure 4In the practice of the modeling method in, different P values ​​were used, from loose to strict, to obtain the models of 6 proteins, 4 proteins, and 1 protein (P values ​​were 0.05, 0.04, and 0.025, respectively), and the performance of the models predicted by using different numbers of protein combinations was compared. According to the comparison results, it was finally determined that the 6 protein features in the present invention were used as the optimal marker combination for model prediction. The following Table 2 shows the performance comparison of the area under the AUC curve of the application prediction model and the joint prediction model using different numbers of protein features.

[0123] Table 2 Performance indicators of different models on the test set

[0124]

[0125] In Table 1:

[0126] 1 protein: P05160_F13B;

[0127] The four proteins were: P02647_APOA1, P05160_F13B, P25311_AZGP1, and Q7Z553_MDGA2;

[0128] The six proteins were finally determined as the optimal marker combination: P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B.

[0129] It can be clearly seen from Table 2 that the performance of the optimal marker combination of 6 protein features is better than the performance of the application prediction model / combined prediction model using 1 protein or 4 proteins. Therefore, the 6 protein features screened out by the present invention are used as the optimal marker combination, which greatly improves the performance of colorectal cancer liver metastasis prediction.

[0130] In addition, it can be seen from Table 2 that the performance of the joint prediction model using the 6 optimal marker combinations and clinical characteristics is better than the application prediction model using only the 6 optimal marker combinations. It can be seen that when the subjects' CEA, CA19-9 and age clinical characteristics are available, taking them into consideration and using the joint prediction model to predict the risk of colorectal cancer liver metastasis can further improve the prediction performance of the model and system.

[0131] The above description is intended to be illustrative rather than restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. For example, a person of ordinary skill in the art can use other embodiments when reading the above description. In addition, in the above-mentioned specific embodiments, various features can be grouped together to simplify the present invention. This should not be interpreted as an intention that a disclosed feature that is not required to be protected is necessary for any claim. On the contrary, the subject matter of the present invention may be less than all the features of a specific disclosed embodiment. Thus, the claims are incorporated into the specific embodiments as examples or embodiments, wherein each claim is independently used as a separate embodiment, and it is considered that these embodiments can be combined with each other in various combinations or arrangements. The scope of the present invention should be determined with reference to the claims and the full scope of equivalent forms granted by these claims.

Claims

1. Use of a reagent for detecting the expression level of a biomarker in the preparation of a product for predicting the risk of liver metastasis in a subject with colorectal cancer, wherein: The biomarkers are a combination of P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1 and P05160_F13B.

2. A system for predicting the risk of liver metastasis in a subject with colorectal cancer, the system comprising a processor and a display, The processor is configured as follows: Obtain the expression levels of the following six biomarkers of the subjects: P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1, and P05160_F13B; Based on the expression levels of the six biomarkers obtained, using a first machine learning model to predict the risk parameters of colorectal cancer liver metastasis in the subject; and The display is caused to present the predicted risk parameter of colorectal cancer liver metastasis in the subject.

3. The system according to claim 2, characterized in that The first machine learning model is constructed based on the random forest algorithm with the expression levels of the six biomarkers as feature information.

4. A system for predicting the risk of liver metastasis in a subject with colorectal cancer, the system comprising a processor and a display, The processor is configured as follows: Obtain the expression levels of the following six biomarkers of the subjects: P43652_AFM, P00450_CP, P25311_AZGP1, Q7Z553_MDGA2, P02647_APOA1, and P05160_F13B; Obtain the clinical characteristics of the subjects including CEA, CA19-9 and age; Based on the expression levels of the six biomarkers and the expression amounts of clinical characteristics obtained, a second machine learning model is used to predict the risk parameters of colorectal cancer liver metastasis in the subject; and The display is caused to present the predicted risk parameter of colorectal cancer liver metastasis in the subject.

5. The system according to claim 2 or 4, characterized in that: The processor is further configured to: predict risk parameters for colorectal cancer liver metastasis in the subject before imaging reveals that the subject has liver metastasis, or when the subject has no other known clinical evidence of liver metastasis, and based on the predicted risk parameters for colorectal cancer liver metastasis in the subject, enable the display to present a correspondence between the predicted risk parameters and at least one preset threshold, as well as corresponding diagnosis and treatment recommendations when the predicted risk parameters are higher than the preset threshold, wherein close monitoring of liver metastasis is recommended when the predicted risk parameters are higher than the first threshold.

6. The system according to claim 2 or 4, characterized in that: It also includes a liquid chromatography device and a mass spectrometry analysis device, which are configured to obtain a sample of the plasma collected from the subject, analyze and quantify it through targeted proteomics, and detect the expression levels of 6 biomarkers therefrom.

7. The system according to claim 4, characterized in that The second machine learning model is constructed based on the random forest algorithm, with the expression levels of the six biomarkers and the expression amounts of clinical characteristics including CEA, CA19-9 and age as feature information.

Citation Information

Patent Citations

  • Application of SLC14A1 as marker in preparation of product for evaluating colorectal cancer liver metastasis risk and / or prognosis condition

    CN117330752A

  • Subcutaneous Delivery of Messenger RNA

    US20210220449A1