Bile duct cancer biomarker combination
The diagnostic model constructed through biomarker combination and MLP algorithm solves the problem of early diagnosis of cholangiocarcinoma, realizes non-invasive and highly sensitive cholangiocarcinoma detection, and improves the accuracy and safety of early screening.
Patent Information
- Application Number
- CN202510766763.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing cholangiocarcinoma diagnosis technology has problems such as limitations in early screening, insufficient sensitivity of serum markers, and difficulty in pathological diagnosis, making it difficult to achieve non-invasive and high-sensitivity early diagnosis.
A diagnostic model was constructed by combining biomarker combinations (IL6, VTCN1, PPY, CEACAM5, CPXM1, FUS, CLMP, KDR, LGALS7) combined with a multi-layer perceptron (MLP) algorithm, and the diagnosis of cholangiocarcinoma was performed by detecting the biomarker expression level in the sample.
It improves the sensitivity and specificity of early diagnosis of cholangiocarcinoma, provides a non-invasive and efficient diagnostic method, and overcomes the limitations of traditional methods.
Smart Images

Figure CN120275636A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedicine, and particularly relates to a biomarker combination for cholangiocarcinoma. Background Art
[0002] Cholangiocarcinoma is a malignant tumor originating from the biliary tract system. Currently, the diagnosis of cholangiocarcinoma (CCA) mainly relies on blood biomarkers (such as CA19-9), imaging examinations (such as ultrasound, CT, MRI), endoscopic retrograde cholangiopancreatography (ERCP), and tissue biopsy.
[0003] There are significant challenges in the early clinical diagnosis of cholangiocarcinoma, which are specifically manifested in the following aspects: 1. Limitations in early screening: More than 60% of patients have progressed to the advanced stage at the initial diagnosis, and the existing detection system has insufficient recognition efficiency for early lesions. The sensitivity of imaging examinations (such as ultrasound, CT) for micro-lesions (<2 cm) is only 35 - 48%, which is prone to missed diagnosis. 2. Defects in the efficacy of serum biomarkers: Although the carbohydrate antigen 19-9 (CA19-9) widely used clinically has auxiliary diagnostic value, its sensitivity is only 72%, and it is prone to false positives in diseases such as benign biliary obstruction and pancreatitis. The lack of specificity restricts its independent diagnostic value. 3. Clinical limitations of pathological diagnosis: Although histopathology is the gold standard for diagnosis, there are two dilemmas: a. Technical feasibility: Due to the complex anatomical location of hilar and distal cholangiocarcinoma, the success rate of endoscopic biopsy or cytological brushing is less than 40%; b. Patient adaptability: Patients in the advanced or cachectic state are difficult to tolerate invasive operations, and the risk of complications increases (such as a bleeding rate >15%).
[0004] The above defects result in the existing diagnostic system being difficult to balance precision, safety, and universality, and there is an urgent need to establish a non-invasive and highly sensitive diagnostic scheme. Summary of the Invention
[0005] To make up for the deficiencies of the existing technology, the present invention provides a diagnostic model for distinguishing cholangiocarcinoma patients from non-cholangiocarcinoma individuals (including patients with benign biliary diseases and healthy people).
[0006] To achieve the above object, the present invention adopts the following technical solutions: The present invention provides a combination of biomarkers capable of diagnosing or predicting cholangiocarcinoma, and the biomarkers include: IL6, VTCN1, PPY, CEACAM5, CPXM1, FUS, CLMP, KDR, and LGALS7.
[0007] The biomarker IL6 used in the present invention refers to interleukin 6, and its Gene ID in NCBI is 3569. The biomarker VTCN1 used in the present invention refers to V-set domain containing T cell activation inhibitor 1, and its Gene ID in NCBI is 79679. The biomarker PPY used in the present invention refers to pancreatic polypeptide, and its Gene ID in NCBI is 5539. The biomarker CEACAM5 used in the present invention refers to CEA cell adhesion molecule 5, and its Gene ID in NCBI is 1048. The biomarker CPXM1 used in the present invention refers to carboxypeptidase X, M14 family member 1, and its Gene ID in NCBI is 56265. The biomarker FUS used in the present invention refers to FUS RNA binding protein, and its Gene ID in NCBI is 2521. The biomarker CLMP used in the present invention refers to CXADR like cell adhesion molecule, and its Gene ID in NCBI is 79827. The biomarker KDR used in the present invention refers to kinase insert domain receptor, and its Gene ID in NCBI is 3791. The biomarker LGALS7 used in the present invention refers to galectin 7, and its Gene ID in NCBI is 3963.
[0008] The present invention provides the use of a reagent for detecting a combination of the aforementioned biomarkers in a sample in the preparation of a product for diagnosing or predicting cholangiocarcinoma.
[0009] Furthermore, the sample is from a tissue sample, primary or cultured cells or cell lines, cell supernatant, cell lysate, platelets, serum, plasma, vitreous humor, lymph fluid, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tumor lysate, tissue culture medium, or tissue extract.
[0010] Furthermore, the sample is from a tissue sample, platelets, serum, plasma, whole blood, tissue culture medium, or tissue extract.
[0011] Furthermore, the sample is plasma.
[0012] Furthermore, the sample is from a human or non-human mammal.
[0013] Further, the sample is from a human.
[0014] In some specific embodiments, suitable mammals falling within the scope of the present invention include vertebrates, specifically including but not limited to any member of the subphylum Chordata, including primates, and including monkey species, rodents, lagomorphs, cattle, sheep, goats, pigs, horses, dogs, cats, avians (e.g., chickens, turkeys, ducks, geese, companion birds (such as canaries, budgerigars, etc.)), marine mammals, reptiles, and fish.
[0015] Further, the product includes reagents, kits, primers, test strips, chips, probes.
[0016] Further, the kit includes a reagent for extracting biomarkers and a reagent for quantitatively detecting biomarkers.
[0017] Further, the reagents for detecting biomarkers are: primers for specifically amplifying biomarkers, probes for specifically recognizing metabolites, and binders for specifically binding metabolite proteins.
[0018] Further, the binder includes an antibody for specifically binding a biomarker protein, an antibody functional fragment, a conjugated antibody.
[0019] In some embodiments, the chip can be configured such that a detectable output is provided only when the concentration of one or more of these biomarkers exceeds a threshold, the threshold being selected to distinguish the concentration and / or amount and / or expression level of the biomarkers indicating a control subject from the concentration and / or amount and / or expression level indicating that the subject has cholangiocarcinoma.
[0020] The content of biomarkers can also extend to mean proteins or transcripts. In some embodiments, the biomarker content refers to an increase in the content by the following percentages: at least 5%; at least 10%; at least 20%; at least 30%; at least 40%; at least 50%; at least 60%; at least 70%; at least 80%; at least 90%; at least 100%; at least 110%; at least 120%; at least 130%; at least 140%; at least 150%; or more. In other embodiments, the biomarker refers to a decrease in the content by the following percentages: at least 5%; at least 10%; at least 20%; at least 30%; at least 40%; at least 50%; at least 60%; at least 70%; at least 80%; at least 90%; at least 100% (i.e., the biomarker is absent). The biomarkers are expressed at statistically significantly different contents (i.e., the p-value is less than 0.05 and / or the q-value is less than 0.10, as determined using a t-test, Welch's T-test, or Wilcoxon rank sum test).
[0021] The present invention provides a product for diagnosing or predicting cholangiocarcinoma, the product comprising a reagent for detecting the combination of the aforementioned biomarkers.
[0022] Furthermore, the reagent comprises a reagent for detecting mRNA and / or a reagent for detecting protein.
[0023] In some embodiments, the reagent for detecting mRNA includes, but is not limited to: a reaction reagent for visualizing the amplicon corresponding to the primer, such as a reagent for visualizing the amplicon by agarose gel electrophoresis, enzyme-linked gel method, chemiluminescence method, in situ hybridization method, fluorescence detection method, etc.; an RNA extraction reagent; a reverse transcription reagent; a cDNA amplification reagent; a standard for preparing a standard curve; a positive control. In some embodiments, the reagent for detecting protein includes, but is not limited to: a blocking solution, an antibody dilution solution, a washing buffer, a color development termination solution.
[0024] The present invention provides an application of a biomarker in a sample in constructing a diagnostic / predictive model for cholangiocarcinoma, the diagnostic / predictive model being established by an ensemble learning method, and the biomarker being the combination of the aforementioned biomarkers.
[0025] Furthermore, the ensemble learning method includes linear regression algorithm, support vector machine algorithm, nearest neighbor / k-nearest neighbor algorithm, logistic regression algorithm, decision tree algorithm, k-means algorithm, random forest algorithm, naive Bayes algorithm, dimensionality reduction algorithm, gradient boosting algorithm, MLP algorithm.
[0026] Furthermore, the ensemble learning method is the MLP algorithm.
[0027] Furthermore, the sample is from a tissue sample, primary or cultured cells or cell lines, cell supernatant, cell lysate, platelets, serum, plasma, vitreous humor, lymph fluid, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tumor lysate, tissue culture medium, or tissue extract.
[0028] Furthermore, the sample is from a tissue sample, platelets, serum, plasma, whole blood, tissue culture medium, or tissue extract.
[0029] Furthermore, the diagnostic / predictive model diagnoses / predicts whether a subject has cholangiocarcinoma or is at risk of cholangiocarcinoma by calculating the AUC value of each sample and comparing it with the optimal threshold obtained from the ROC curve generated by the prediction results of the model established by the ensemble learning method on the test set.
[0030] Furthermore, the optimal threshold is 0.498.
[0031] Further, the judgment of the diagnosis / prediction model on the result is based on the following criteria: When the AUC value in the sample calculated by the diagnosis / prediction model is greater than or equal to the optimal threshold, the diagnosis / prediction model outputs the result as cholangiocarcinoma or high risk of cholangiocarcinoma; When the AUC value in the sample calculated by the diagnosis / prediction model is less than the optimal threshold, the diagnosis / prediction model outputs the result as non-cholangiocarcinoma or low risk of cholangiocarcinoma.
[0032] In some embodiments, the logarithmic function for associating a biomarker or combination of biomarkers with a disease preferably uses an algorithm developed and obtained by applying statistical methods. In some embodiments, the sample directly taken from the subject is not further processed. For example, blood can be obtained from the peripheral circulatory system of the subject.
[0033] The multi-layer perceptron (MLP) algorithm used in the present invention is a supervised learning model based on the feed-forward neural network structure. Its core feature is to realize the non-linear mapping of data through the fully connected architecture of the input layer, hidden layer and output layer. The input layer receives the original feature vector, the hidden layer extracts high-order features layer by layer through activation functions (such as ReLU or Sigmoid), the output layer generates the final prediction result, and the training process uses the backpropagation algorithm to optimize the weight parameters. Typical applications include image classification (such as the accuracy rate of the MNIST dataset reaching 98.2%) and regression analysis. Its performance advantage is reflected in enhancing the feature expression ability through the multi-layer structure and overcoming the linear limitation of the single-layer perceptron.
[0034] The present invention provides a diagnosis / prediction model as described above. The diagnosis / prediction model includes an ensemble learning module for constructing an MLP structure capable of diagnosing or determining whether it is a diagnosis / prediction of cholangiocarcinoma.
[0035] Further, the diagnosis / prediction model includes a result judgment module for using the MLP structure to judge the result.
[0036] Further, the diagnosis / prediction model includes a detection module for detecting the AUC value in each sample.
[0037] The present invention provides a system or device for judging or predicting whether a subject has cholangiocarcinoma. The system or device includes: The computer imports the expression level data of the combination of the aforementioned biomarkers detected in the sample collected from the subject into the aforementioned diagnosis and prediction model, and obtains the result of the computer's judgment on whether the subject has cholangiocarcinoma according to the result output.
[0038] In some embodiments, the system or device performs the following operations: 1. Collect peripheral blood using an EDTA anticoagulant tube; 2. Extract plasma after centrifuging at 4 degrees Celsius and 3000 rpm for 15 minutes; 3. Detect the abundances of 9 target proteins using the Olink platform; 4. Input the obtained expression values into a pre-trained MLP classification model; 5. Output the predicted probability as the auxiliary diagnosis result.
[0039] In the context of the present invention, the term "sample" refers to a composition obtained from or derived from a subject (e.g., an individual of interest) that contains cells and / or other molecular entities to be characterized and / or identified according to, for example, physical, biochemical, chemical, and / or physiological characteristics. A "sample" can be any biological sample isolated from a subject. The term "sample" includes any biological sample that can be extracted from a subject, whether untreated, treated, diluted, or concentrated.
[0040] The term "ensemble learning method" as used in the present invention refers to an algorithm that gives a computer the ability to learn without being explicitly programmed, including algorithms that learn from data and make predictions on data. The ensemble learning methods used in the embodiments disclosed herein can include (but are not limited to) random forest (RF), least absolute shrinkage and selection operator (LASSO) logistic regression, regularized logistic regression, XGBoost, decision tree learning, artificial neural network (ANN), deep neural network (DNN), support vector machine, rule-based machine learning, etc. Algorithms such as linear regression or logistic regression can be used as part of the ensemble learning method process.
[0041] The terms "object", "patient", "subject", "test subject", or "individual" as used in the present invention refer to any subject in need of treatment or prevention, particularly a vertebrate subject, and even more particularly a mammalian subject.
[0042] The term "comprising / including" as used in the present invention means that a composition or method of one or more named elements or steps is open-ended, meaning that the named elements or steps are required, but other elements or steps can be added within the scope of the composition and method. To avoid prolixity, it should also be understood that any composition or method described as "comprising / including" one or more named elements or steps also describes the corresponding more limited composition or method "consisting essentially of the same named elements or steps", meaning that the composition or method includes the named required elements or steps and may also include other elements or steps that do not substantially affect the basic and novel features of the composition or method.
[0043] The term "reagent" as used in the present invention should not be construed narrowly, but should be extended to small molecules, protein molecules such as peptides, polypeptides and proteins, and compositions containing them, as well as genetic molecules such as RNA, DNA and their mimetics and chemical analogs, and cellular agents.
[0044] The term "diagnosis" as used in the present invention refers to the identification or classification of a molecular or pathological state, disease or disorder.
[0045] The terms "level of expression" or "expression level" as used in the present invention are generally used interchangeably and generally refer to the amount of a biomarker in a sample. "Expression" generally refers to a process by which information (e.g., coding genes and / or epigenetics) is converted into a structure present and operative in a cell.
[0046] The term "chip" as used in the present invention may refer to a solid substrate having a generally planar surface to which an adsorbent is attached. The surface of the biochip may contain a plurality of addressable locations, each of which may be bound with an adsorbent.
[0047] The term "probe" as used in the present invention refers to a molecule that can bind to a specific sequence or subsequence or other part of another molecule.
[0048] The term "module" as used in the present invention refers to the computer program logic for providing a specified function. Thus, a module can be implemented in hardware, firmware and / or software.
[0049] Advantages and beneficial effects of the present invention: Collect plasma samples from cholangiocarcinoma patients and control groups, detect 384 proteins in plasma using Olink Proximity Extension Assay (PEA) technology, screen features using LASSO regression, calculate the individual disease probability by combining with the MLP algorithm, and finally construct a high-performance cholangiocarcinoma diagnosis model based on 9 proteins. Description of the drawings
[0050] Figure 1 It is the ROC curve of 9-PCM, CA19-9 and CEA in the training set.
[0051] Figure 2 It is the ROC curve of 9-PCM, CA19-9 and CEA in the validation set.
[0052] Figure 3 It is the ROC curve of 9-PCM, CA19-9 and CEA in the independent test set. Detailed implementation manners
[0053] The following provides definitions of some terms used in this specification. Unless otherwise specified, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0054] It should be noted that, without conflict, the embodiments and features in the embodiments of this application may be combined with each other. In addition, the terms used herein are for the purpose of describing specific embodiments only and should not limit the scope of the present invention, as the scope of the present invention is limited only by the scope defined by the appended claims. The present invention will be described in detail below in conjunction with embodiments.
[0055] Receiver operating characteristic curve (ROC curve for short): It is a curve plotted with sensitivity (true positive rate) as the ordinate and 1 - specificity (false positive rate) as the abscissa according to a series of different binary classification methods (cutoff values or decision thresholds). The area under the ROC curve is an important index of test accuracy. The larger the area under the ROC curve, the greater the diagnostic value of the test.
[0056] The term "AUC" is an abbreviation for "area under the curve". Specifically, it refers to the area under the receiver operating characteristic (ROC) curve. The ROC curve is a curve of the true positive rate and false positive rate at different possible cut - off points of a diagnostic test. It shows that the trade - off between sensitivity and specificity depends on the selected cut - off point (any increase in sensitivity will be accompanied by a decrease in specificity). The area under the ROC curve (AUC) is a measure of the accuracy of a diagnostic test (the larger the area, the better; the optimal value is 1; the ROC curve of a random test lies on the diagonal line and the area is 0.5; refer to J.P. Egan. (1975) Signal Detection Theory and ROC Analysis, Academic Press, New York).
[0057] The present invention will be further elaborated below in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are presented by way of example and are not intended to limit the present invention. Without departing from the scope of the present invention, the main features of the present invention can be used in various embodiments.
[0058] Embodiment 1. Data collection From January 2018 to January 2024, a total of 168 plasma samples from patients with cholangiocarcinoma, 32 from the benign control group (including those pathologically diagnosed with cholangitis, choledocholithiasis, choledochal cyst, choledochal papilloma, etc.), and 66 from healthy volunteers were collected from two independent tertiary hospitals, with a total of 266 samples from the three groups. 70% of the samples were randomly divided into the training set, and 30% into the validation set; in addition, 32 plasma samples from newly diagnosed cholangiocarcinoma patients, 8 from the benign control group, and 14 from healthy volunteers in this region during March to September 2024 were collected, totaling 54 samples as an independent test set.
[0059] The inclusion criteria for the cholangiocarcinoma group were as follows: 1. Aged 18 - 90 years, pathologically diagnosed as a patient with malignant cholangiocarcinoma; 2. No history of other malignancies; 3. Not undergoing anti-tumor treatment before sample collection; 4. Voluntarily participating in the study and signing the informed consent form. The exclusion criteria were: 1. Having other malignancies.
[0060] The inclusion criteria for the benign cholangiopathy group were as follows: 1. Aged 18 - 90 years, pathologically diagnosed with benign cholangiopathy (including cholangitis, choledocholithiasis, choledochal cyst, choledochal papilloma); 2. Voluntarily participating in the study and signing the informed consent form; the exclusion criteria were: diagnosed with malignancy or diagnosed with malignancy within 1 year of follow-up after enrollment.
[0061] The inclusion criteria for the healthy population were as follows: 1. Aged 18 - 90 years, those who underwent physical examinations at the local hospital physical examination center; 2. Voluntarily participating in the study and signing the informed consent form. The exclusion criteria were: 1. Having a history of malignancy or diagnosed with malignancy within 1 year of follow-up after enrollment.
[0062] All plasma samples were collected using EDTA anticoagulant tubes, centrifuged at 3000 rpm for 15 minutes at 4°C to separate the plasma, and detected using the Olink® Oncology II Panel platform. This platform has the advantages of high sensitivity and high repeatability for detecting low-abundance proteins and can detect 384 tumor-related proteins.
[0063] Specific partial information of the collected samples is shown in Table 1.
[0064] Table 1 2. Model Construction In the training set, the numbers of cholangiocarcinoma and non-cholangiocarcinoma samples were 123 and 63 respectively; in the validation set, they were 45 and 35 respectively. Based on the detection results of the Olink platform, LASSO regression analysis was performed on 384 plasma proteins, and 56 candidate proteins were initially screened out; subsequently, the multi-layer perceptron (MLP) algorithm was combined with the stepwise forward regression method to further optimize the features, and finally 9 plasma proteins with diagnostic value were determined as the model input variables. These 9 proteins include: IL6, VTCN1, PPY, CEACAM5, CPXM1, FUS, CLMP, KDR, LGALS7.
[0065] Model algorithm expression: The model uses the expression values of 9 proteins to form a vector , constructs an MLP structure, including an input layer, a hidden layer, and an output layer, and finally outputs the risk probability of cholangiocarcinoma : Among them: : Activation function; : Sigmoid function; : Weights from the input layer to the hidden layer; : Weights from the hidden layer to the output layer; : Bias term.
[0066] The loss function uses the binary cross-entropy function: The model training uses the cross-entropy loss function as the optimization objective, and the Adam optimizer is used for parameter update, with a learning rate of 0.001. The mini-batch method is used during training, the batch size is set to 32, the maximum number of training epochs is 100, and an early stopping mechanism (Early Stopping) is set. If the validation set loss does not decrease within 10 consecutive epochs, the training is terminated early to prevent overfitting. The classification decision threshold is 0.498, and when the model outputs it is determined as cholangiocarcinoma.
[0067] 3. Model Parameter and Performance Comparison The cholangiocarcinoma diagnosis model (9-PCM) constructed in the present invention is based on the multi-layer perceptron (MLP) algorithm, including an input layer (9 nodes), two hidden layers (structures with 30 neurons and 10 neurons), and an output layer (1 node, with the activation function being Sigmoid). Adam optimizer is used for training, the learning rate is 0.001, the batch size is 32, the maximum number of training epochs is set to 200, and an early stopping strategy is set. The loss function is the binary cross-entropy loss function, and the optimization objective is to minimize the prediction error of the classification of cholangiocarcinoma (1) and non-cholangiocarcinoma (0). The model outputs a cholangiocarcinoma risk score, and the classification threshold is 0.498. If P≥0.498, it is determined as cholangiocarcinoma.
[0068] True hyperparameters: L2 penalty (regularization term) parameter α = 0.01 Solver: "solver" = "adam" Adam stability value ϵ = 1×10^(-8) Optimization tolerance "tol" = 0.0001 Hidden layer size "hidden_layer_sizes" = (30, 10) Initial learning rate "learning_rate_init" = 0.001 Maximum number of iterations "max_iter" = 200 Model performance comparative analysis: In the training set (123 cases of cholangiocarcinoma vs 63 cases of non-cholangiocarcinoma), validation set (45 cases vs 35 cases), and independent test set (32 cases vs 22 cases), the model of the present invention is compared with the traditional blood tumor markers CA19-9 and CEA in multiple indicators. Specifically, as shown in Table 2, the ROC curves of the training set, validation set, and independent test set are as Figure 1 , Figure 2 , Figure 3 shown.
[0069] Table 2 It can be seen from the above results that the 9-PCM model shows better diagnostic performance than the traditional tumor markers CA19-9 and CEA in the training set, validation set, and independent test set, especially in terms of sensitivity and F1-score, and is more suitable for the early detection and high-risk screening of cholangiocarcinoma.
[0070] The description of the above embodiments is only for understanding the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications will also fall within the protection scope of the claims of the present invention.
Claims
1. A combination of biomarkers capable of diagnosing or predicting cholangiocarcinoma, the biomarkers including: IL6, VTCN1, PPY, CEACAM5, CPXM1, FUS, CLMP, KDR, and LGALS7.
2. Use of a reagent for detecting the combination of biomarkers according to claim 1, which can diagnose or predict cholangiocarcinoma, in the preparation of a product for diagnosing or predicting cholangiocarcinoma.
3. According to the use of claim 2, the sample is from a tissue sample, primary or cultured cells or cell line, cell supernatant, cell lysate, platelet, serum, plasma, vitreous humor, lymph, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tumor lysate, tissue culture medium, or tissue extract; Preferably, the sample is from a tissue sample, platelet, serum, plasma, whole blood, tissue culture medium, or tissue extract; Preferably, the sample is from a human or non-human mammal; Preferably, the sample is from a human.
4. According to the use of claim 2, the product includes reagents, kits, primers, test strips, chips, probes; Preferably, the kit includes reagents for extracting biomarkers and reagents for quantitative detection of biomarkers.
5. A product for diagnosing or predicting cholangiocarcinoma, the product includes a reagent for detecting the combination of biomarkers according to claim 1; Preferably, the reagent includes a reagent for detecting mRNA and / or a reagent for detecting protein.
6. Use of a biomarker in a sample in the construction of a diagnostic / predictive model for cholangiocarcinoma, the diagnostic / predictive model uses an ensemble learning method to establish the model, and the biomarker is the combination of biomarkers according to claim 1.
7. According to the use of claim 6, the ensemble learning method includes linear regression algorithm, support vector machine algorithm, nearest neighbor / k-nearest neighbor algorithm, logistic regression algorithm, decision tree algorithm, k-means algorithm, random forest algorithm, naive Bayes algorithm, dimensionality reduction algorithm, gradient boosting algorithm, MLP algorithm; Preferably, the ensemble learning method is the MLP algorithm; Preferably, the sample is from a tissue sample, primary or cultured cells or cell line, cell supernatant, cell lysate, platelet, serum, plasma, vitreous humor, lymph, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tumor lysate, tissue culture medium, or tissue extract; Preferably, the sample is from a tissue sample, platelet, serum, plasma, whole blood, tissue culture medium, or tissue extract; Preferably, the diagnostic / predictive model diagnoses / predicts whether a person has cholangiocarcinoma or is at risk of cholangiocarcinoma by calculating the AUC value of each sample and comparing it with the best threshold obtained from the ROC curve generated by the prediction results of the model established by the ensemble learning method on the test set; Preferably, the best threshold is 0.498; Preferably, the judgment of the result of the diagnostic / predictive model is based on the following criteria: When the AUC value in the sample calculated by the diagnostic / predictive model is greater than or equal to the optimal threshold, the output result of the diagnostic / predictive model is cholangiocarcinoma or high risk of cholangiocarcinoma; When the AUC value in the sample calculated by the diagnostic / predictive model is less than the optimal threshold, the output result of the diagnostic / predictive model is non-cholangiocarcinoma or low risk of cholangiocarcinoma.
8. A diagnostic / predictive model according to claim 6 or 7, wherein the diagnostic / predictive model comprises an ensemble learning module for constructing an MLP structure capable of performing diagnosis or determining whether it is a diagnosis / prediction of cholangiocarcinoma.
9. The diagnostic / predictive model according to claim 8, wherein the diagnostic / predictive model comprises a result judgment module for using the MLP structure to judge the result; Preferably, the diagnostic / predictive model comprises a detection module for detecting the AUC value in each sample.
10. A system or device for judging or predicting whether a subject has cholangiocarcinoma, the system or device comprising: The computer imports the expression level data of the combination of biomarkers described in claim 1 detected in the sample collected from the subject into the diagnostic and predictive model of claim 8 or 9, and obtains the result of the computer's judgment on whether the subject has cholangiocarcinoma according to the result output.
Citation Information
Patent Citations
Probe set and kit for detecting whole exons of extended genetic diseases and application of probe set
CN110499364A
Bile duct cancer molecular typing gene marker and application thereof
CN116240282A
Screening and diagnosing method for extrahepatic cholangiocarcinoma
CN118652957A
New biliary tract cancer biomarker
JP2012237685A
Biomarker for diagnosis of extrahepatic bile duct carcinoma, intrahepatic bile duct carcinoma, or gallbladder carcinoma
US20180080935A1