A biomarker panel for cholangiocarcinoma
The cholangiocarcinoma diagnostic model constructed by the biomarker combination IL6, VTCN1, PPY, CEACAM5, CPXM1, FUS, CLMP, KDR, LGALS7 and MLP algorithms solves the problem of early diagnosis in the prior art and realizes high sensitivity and non-invasive cholangiocarcinoma detection.
Patent Information
- Application Number
- CN202510766763.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing cholangiocarcinoma diagnosis technology has limitations in early screening, insufficient serum marker efficacy, and difficulty in pathological diagnosis. It is difficult to balance accuracy and universality, and non-invasive and high-sensitivity diagnostic solutions are urgently needed.
The biomarker combination IL6, VTCN1, PPY, CEACAM5, CPXM1, FUS, CLMP, KDR, and LGALS7 were used, and combined with an integrated learning method, especially a multi-layer perceptron (MLP) algorithm, a cholangiocarcinoma diagnosis model was constructed, and the biomarker combination expression level in the sample was detected for diagnosis.
It improves the sensitivity and specificity of early diagnosis of cholangiocarcinoma, reduces the rate of misdiagnosis, and provides a non-invasive and efficient diagnostic tool suitable for cholangiocarcinoma detection in human and mammalian samples.
Smart Images

Figure CN120275636B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedicine, and specifically relates to a bile duct cancer biomarker combination. Background Art
[0002] Cholangiocarcinoma is a malignant tumor originating from the biliary system. Currently, the diagnosis of cholangiocarcinoma (CCA) mainly relies on blood markers (such as CA19-9), imaging examinations (such as ultrasound, CT, MRI), endoscopic retrograde cholangiopancreatography (ERCP) and tissue biopsy.
[0003] Early clinical diagnosis of cholangiocarcinoma presents significant challenges, specifically the following: 1. Limitations of early screening: Over 60% of patients have already progressed to advanced stages at the time of initial diagnosis, and existing detection systems are inefficient in identifying early lesions. Imaging studies (such as ultrasound and CT) have a sensitivity of only 35-48% for small lesions (<2 cm), which can easily lead to missed diagnoses. 2. Deficiencies in serum biomarkers: While the widely used carbohydrate antigen 19-9 (CA19-9) test has auxiliary diagnostic value, its sensitivity is only 72%, and it is prone to false positives in conditions such as benign biliary obstruction and pancreatitis. Its lack of specificity limits its independent diagnostic value. 3. Clinical limitations of pathological confirmation: Although histopathology is the gold standard for diagnosis, it faces two challenges: a. Technical feasibility: Due to the complex anatomical location of hilar and distal cholangiocarcinoma, the success rate of endoscopic biopsy or cytology brushing is less than 40%; b. Patient adaptability: Patients with advanced stages or cachexia have difficulty tolerating invasive procedures, which increase the risk of complications (e.g., bleeding rates >15%).
[0004] The above-mentioned defects make it difficult for the existing diagnostic system to balance accuracy, safety and universality, and there is an urgent need to establish a non-invasive and highly sensitive diagnostic solution. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the present invention provides a diagnostic model for distinguishing bile duct cancer patients from non-bile duct cancer individuals (including patients with benign biliary diseases and healthy people).
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] The present invention provides a combination of biomarkers capable of diagnosing or predicting cholangiocarcinoma, wherein the biomarkers include IL6, VTCN1, PPY, CEACAM5, CPXM1, FUS, CLMP, KDR, and LGALS7.
[0008] The biomarker IL6 used in the present invention refers to interleukin 6, NCBI Gene ID: 3569. The biomarker VTCN1 used in the present invention refers to V-set domain containing T cell activation inhibitor 1, NCBI Gene ID: 79679. The biomarker PPY used in the present invention refers to pancreatic polypeptide, NCBI Gene ID: 5539. The biomarker CEACAM5 used in the present invention refers to CEA cell adhesion molecule 5, NCBI Gene ID: 1048. The biomarker CPXM1 used in the present invention refers to carboxypeptidase X, M14 family member 1, NCBI Gene ID: 56265. The biomarker FUS used in the present invention refers to FUS RNA binding protein, NCBI Gene ID: 2521. The biomarker CLMP used in the present invention refers to CXADR-like cell adhesion molecule, NCBI Gene ID: 79827. The biomarker KDR used in the present invention refers to kinase insert domain receptor, NCBI Gene ID: 3791. The biomarker LGALS7 used in the present invention refers to galectin 7, NCBI Gene ID: 3963.
[0009] The present invention provides use of a reagent for detecting the aforementioned combination of biomarkers in a sample in the preparation of a product for diagnosing or predicting bile duct cancer.
[0010] Furthermore, the sample is from a tissue sample, primary or cultured cells or cell lines, cell supernatant, cell lysate, platelets, serum, plasma, vitreous humor, lymph fluid, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tumor lysate, tissue culture fluid, or tissue extract.
[0011] Furthermore, the sample is from a tissue sample, platelets, serum, plasma, whole blood, tissue culture fluid, or tissue extract.
[0012] Furthermore, the sample is plasma.
[0013] Furthermore, the sample is from a human or a non-human mammal.
[0014] Furthermore, the sample is from a human.
[0015] In some specific embodiments, suitable mammals falling within the scope of the present invention include vertebrates, specifically including, but not limited to, any member of the subphylum Chordata, including primates, and including monkey species, rodents, Leporidae, bovines, ovines, goats, porcines, equines, dogs, cats, avians (e.g., chickens, turkeys, ducks, geese, companion birds (e.g., canaries, budgies, etc.)), marine mammals, reptiles, and fish.
[0016] Furthermore, the products include reagents, kits, primers, test strips, chips, and probes.
[0017] Furthermore, the kit includes reagents for extracting biomarkers and reagents for quantitative detection of biomarkers.
[0018] Furthermore, the reagents for detecting biomarkers include: primers for specifically amplifying biomarkers, probes for specifically recognizing metabolites, and binding agents for specifically binding to metabolite proteins.
[0019] Furthermore, the binding agent includes antibodies, antibody functional fragments, and conjugated antibodies that specifically bind to biomarker proteins.
[0020] In some embodiments, the chip can be configured to provide a detectable output only when the concentration of one or more of these biomarkers exceeds a threshold value, wherein the threshold value is selected or distinguishes the concentration and / or amount and / or expression level of the biomarker indicative of a control subject from the concentration and / or amount and / or expression level indicative of a subject having cholangiocarcinoma.
[0021] The level of a biomarker can also be extended to mean a protein or transcript. In some embodiments, the level of a biomarker refers to an increase in the following percentage: at least 5%; at least 10%; at least 20%; at least 30%; at least 40%; at least 50%; at least 60%; at least 70%; at least 80%; at least 90%; at least 100%; at least 110%; at least 120%; at least 130%; at least 140%; at least 150%; or more. In other embodiments, the level of a biomarker refers to a decrease in the following percentage: at least 5%; at least 10%; at least 20%; at least 30%; at least 40%; at least 50%; at least 60%; at least 70%; at least 80%; at least 90%; at least 100% (i.e., the biomarker is absent). The biomarkers are expressed at statistically significant different levels (ie, p-value less than 0.05 and / or q-value less than 0.10, as determined using a t-test, Welch T-test, or Wilcoxon rank sum test).
[0022] The present invention provides a product for diagnosing or predicting bile duct cancer, which comprises a reagent for detecting the combination of the aforementioned biomarkers.
[0023] Furthermore, the reagents include reagents for detecting mRNA and / or reagents for detecting protein.
[0024] In some embodiments, the reagents for detecting mRNA include, but are not limited to: reagents for visualizing the amplicons corresponding to the primers, such as reagents for visualizing the amplicons by agarose gel electrophoresis, enzyme-linked gel electrophoresis, chemiluminescence, in situ hybridization, fluorescence detection, etc.; RNA extraction reagents; reverse transcription reagents; cDNA amplification reagents; standards for preparing a standard curve; and positive controls. In some embodiments, the reagents for detecting proteins include, but are not limited to, blocking buffer, antibody diluent, wash buffer, and color development stop buffer.
[0025] The present invention provides an application of biomarkers in a sample in constructing a diagnosis / prediction model for bile duct cancer. The diagnosis / prediction model adopts an integrated learning method to establish the model, and the biomarker is a combination of the biomarkers described above.
[0026] Furthermore, the ensemble learning method includes linear regression algorithm, support vector machine algorithm, nearest neighbor / k-nearest neighbor algorithm, logistic regression algorithm, decision tree algorithm, k-means algorithm, random forest algorithm, naive Bayes algorithm, dimensionality reduction algorithm, gradient boosting algorithm, and MLP algorithm.
[0027] Furthermore, the ensemble learning method is an MLP algorithm.
[0028] Furthermore, the sample is from a tissue sample, primary or cultured cells or cell lines, cell supernatant, cell lysate, platelets, serum, plasma, vitreous humor, lymph fluid, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tumor lysate, tissue culture fluid, or tissue extract.
[0029] Furthermore, the sample is from a tissue sample, platelets, serum, plasma, whole blood, tissue culture fluid, or tissue extract.
[0030] Furthermore, the diagnostic / predictive model diagnoses / predicts whether a patient has bile duct cancer or is at risk of bile duct cancer by calculating the AUC value of each sample and comparing it with the optimal threshold value obtained by the ROC curve generated by the prediction results of the model established based on the ensemble learning method on the test set.
[0031] Furthermore, the optimal threshold is 0.498.
[0032] Furthermore, the diagnosis / prediction model determines the results based on the following criteria:
[0033] When the AUC value of the sample calculated by the diagnosis / prediction model is greater than or equal to the optimal threshold, the diagnosis / prediction model outputs the result as bile duct cancer or high risk of bile duct cancer;
[0034] When the AUC value in the sample calculated by the diagnosis / prediction model is less than the optimal threshold, the diagnosis / prediction model outputs a result of non-cholangiocarcinoma or low risk of cholangiocarcinoma.
[0035] In some embodiments, the logarithmic function used to associate a biomarker or combination of biomarkers with a disease preferably uses an algorithm developed and derived by applying statistical methods. In some embodiments, the sample is taken directly from the subject without further processing, for example, blood can be obtained from the subject's peripheral circulatory system.
[0036] The multi-layer perceptron (MLP) algorithm used in the present invention is a supervised learning model based on a feedforward neural network structure. Its core feature is to achieve nonlinear mapping of data through a fully connected architecture of input layer, hidden layer and output layer. The input layer receives the original feature vector, the hidden layer extracts high-order features layer by layer through activation functions (such as ReLU or Sigmoid), and the output layer generates the final prediction result. The training process uses the backpropagation algorithm to optimize the weight parameters. Typical applications include image classification (such as the MNIST data set with an accuracy of 98.2%) and regression analysis. Its performance advantage is reflected in the enhanced feature expression capability through a multi-layer structure to overcome the linear limitations of a single-layer perceptron.
[0037] The present invention provides a diagnosis / prediction model as described above, which includes an integrated learning module for constructing an MLP structure capable of diagnosing or determining whether it is bile duct cancer.
[0038] Furthermore, the diagnosis / prediction model includes a result judgment module for performing result judgment using an MLP structure.
[0039] Furthermore, the diagnosis / prediction model includes a detection module for detecting the AUC value in each sample.
[0040] The present invention provides a system or device for determining or predicting whether a subject has bile duct cancer, the system or device comprising:
[0041] The computer imports the expression level data of the aforementioned combination of biomarkers detected in the sample collected from the subject into the aforementioned diagnostic prediction model, and based on the result output, obtains the result of the computer determining whether the subject has bile duct cancer.
[0042] In some embodiments, the system or device performs the following operations: 1. collects peripheral blood using EDTA anticoagulant tubes; 2. extracts plasma after centrifugation at 3000 rpm for 15 minutes at 4 degrees Celsius; 3. detects the abundance of 9 target proteins using the Olink platform; 4. inputs the obtained expression values into a pre-trained MLP classification model; 5. outputs the predicted probability as an auxiliary diagnosis result.
[0043] As used herein, the term "sample" refers to a composition obtained or derived from a subject (e.g., an individual of interest) that contains cells and / or other molecular entities to be characterized and / or identified based on, for example, physical, biochemical, chemical, and / or physiological characteristics. A "sample" can be any biological specimen isolated from a subject. The term "sample" includes any biological sample that can be extracted from a subject, whether unprocessed, processed, diluted, or concentrated.
[0044] The term "ensemble learning method" used in the present invention refers to an algorithm that gives a computer the ability to learn without being explicitly programmed, including an algorithm that learns from data and makes predictions about the data. The ensemble learning method used in the embodiments disclosed herein may include (but is not limited to) random forest (RF), least absolute shrinkage and selection operator (LASSO) logistic regression, regularized logistic regression, XGBoost, decision tree learning, artificial neural network (ANN), deep neural network (DNN), support vector machine, rule-based machine learning, etc. Algorithms such as linear regression or logistic regression can be used as part of the ensemble school method process.
[0045] The terms "subject," "patient," "subject," "subject," or "individual" as used herein refer to any subject in need of treatment or prevention, particularly a vertebrate subject, and even more particularly a mammalian subject.
[0046] As used herein, the term "comprising" refers to a composition or method that is open-ended, meaning that the named elements or steps are required, but that other elements or steps may be added within the scope of the composition or method. To avoid redundancy, it should also be understood that any composition or method described as "comprising" one or more named elements or steps also describes a corresponding, more limited composition or method that "consists essentially of the same named elements or steps," meaning that the composition or method includes the named required elements or steps and may further include other elements or steps that do not materially affect the basic and novel characteristics of the composition or method.
[0047] The term "agent" used in the present invention should not be interpreted narrowly, but should be extended to small molecules, protein molecules, such as peptides, polypeptides and proteins, and compositions containing them, as well as genetic molecules, such as RNA, DNA and their mimetics and chemical analogs and cellular agents.
[0048] The term "diagnosis" as used herein refers to the identification or classification of a molecular or pathological state, disease or condition.
[0049] As used herein, the terms "level of expression" or "expression level" are generally used interchangeably and generally refer to the amount of a biomarker in a sample. "Expression" generally refers to the process by which information (e.g., encoding genes and / or epigenetic information) is converted into structures present and operable in a cell.
[0050] The term "chip" as used in the present invention may refer to a solid substrate having a generally planar surface to which an adsorbent is attached. The surface of a biochip may comprise a plurality of addressable locations, each of which may be bound to an adsorbent.
[0051] The term "probe" as used in the present invention refers to a molecule that can bind to a specific sequence or subsequence or other portion of another molecule.
[0052] The term "module" used in the present invention refers to computer program logic for providing a specified functionality. Therefore, a module can be implemented in hardware, firmware and / or software.
[0053] Advantages and beneficial effects of the present invention:
[0054] Plasma samples from cholangiocarcinoma patients and the control group were collected, and 384 plasma proteins were detected using Olink non-proximal extension analysis (PEA) technology. LASSO regression was used to screen features, and the individual probability of disease was calculated using the MLP algorithm. Finally, a high-performance cholangiocarcinoma diagnostic model based on 9 proteins was constructed. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is the ROC curve of the training set 9-PCM, CA19-9 and CEA.
[0056] Figure 2 is the ROC curve of the validation set 9-PCM, CA19-9 and CEA.
[0057] Figure 3 is the ROC curve of the independent test set 9-PCM, CA19-9 and CEA. DETAILED DESCRIPTION
[0058] The following provides definitions of some terms used in this specification. Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0059] It should be noted that, unless there is a conflict, the embodiments and features described in the embodiments of this application may be combined with each other. In addition, the terms used herein are only used to describe specific embodiments and should not limit the scope of the present invention, as the scope of the present invention is limited only to the scope of the appended claims. The present invention will be described in detail below with reference to the embodiments.
[0060] The receiver operating characteristic (ROC) curve is a curve plotted based on a series of different binary classification methods (cutoff values or decision thresholds), with sensitivity (true positive rate) on the vertical axis and 1-specificity (false positive rate) on the horizontal axis. The area under the ROC curve is an important indicator of test accuracy; the larger the area under the ROC curve, the greater the diagnostic value of the test.
[0061] The term "AUC" is an abbreviation for "area under the curve." It specifically refers to the area under the receiver operating characteristic (ROC) curve. The ROC curve is a plot of the true positive rate versus the false positive rate for different possible cut points for a diagnostic test. This shows that the trade-off between sensitivity and specificity depends on the cut point chosen (any increase in sensitivity is accompanied by a decrease in specificity). The area under the ROC curve (AUC) is a measure of the accuracy of a diagnostic test (the larger the area, the better; an optimal value is 1; the ROC curve for a randomized test lies on the diagonal with an area of 0.5; see J.P. Egan (1975) Signal Detection Theory and ROC Analysis, Academic Press, New York).
[0062] The present invention will be further described below with reference to specific examples. It should be understood that the specific embodiments described herein are presented by way of example and are not intended to limit the present invention. The main features of the present invention may be applied to various embodiments without departing from the scope of the present invention.
[0063] Example
[0064] 1. Data Collection
[0065] From January 2018 to January 2024, plasma samples were collected from 168 patients with cholangiocarcinoma at two independent tertiary hospitals. Furthermore, 32 benign controls (including those diagnosed with cholangitis, bile duct stones, bile duct cysts, and biliary papilloma by pathology) and 66 healthy volunteers were collected for a total of 266 subjects. 70% of these samples were randomly divided into a training set and 30% into a validation set. Furthermore, 54 samples, totaling 32 newly diagnosed cholangiocarcinoma patients, 8 benign controls, and 14 healthy volunteers from the region between March and September 2024, were collected as an independent test set.
[0066] Inclusion criteria for the cholangiocarcinoma group were: 1. Patients aged 18-90 years with a pathological diagnosis of bile duct malignancy; 2. No history of other malignancies; 3. No anti-cancer treatment before sample collection; 4. Voluntary participation in the study and signing of informed consent. Exclusion criteria were: 1. Patients with concurrent other malignancies.
[0067] The inclusion criteria for the benign bile duct disease group were: 1. aged 18-90 years, pathologically diagnosed with benign bile duct disease (including cholangitis, bile duct stones, bile duct cysts, and bile duct papilloma); 2. voluntary participation in the study and signing of the informed consent form; the exclusion criteria were: diagnosis of malignant tumor or diagnosis of malignant tumor within 1 year of follow-up after enrollment.
[0068] Inclusion criteria for healthy individuals were: 1. Aged 18-90 years, undergoing a physical examination at a local hospital; 2. Voluntary participation in the study and signing of an informed consent form. Exclusion criteria were: 1. History of malignancy or diagnosis of malignancy within 1 year of enrollment.
[0069] All plasma samples were collected in EDTA-anticoagulant tubes and centrifuged at 3000 rpm at 4°C for 15 minutes to separate the plasma. Plasma was then analyzed using the Olink® Oncology II Panel platform, which offers high sensitivity and reproducibility for low-abundance proteins and can detect 384 tumor-related proteins.
[0070] The specific information of the collected samples is shown in Table 1.
[0071] Table 1
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084] 2. Model construction
[0085] The training set consisted of 123 cholangiocarcinoma and 63 non-cholangiocarcinoma samples; the validation set consisted of 45 and 35 samples, respectively. Based on the Olink platform's detection results, LASSO regression analysis was performed on 384 plasma proteins, initially identifying 56 candidate proteins. Subsequently, a multilayer perceptron (MLP) algorithm combined with stepwise forward regression was used to further optimize features, ultimately identifying nine plasma proteins with diagnostic value as model input variables. These nine proteins include IL6, VTCN1, PPY, CEACAM5, CPXM1, FUS, CLMP, KDR, and LGALS7.
[0086] Model algorithm expression:
[0087] The model is composed of 9 protein expression values into a vector , construct an MLP structure, including an input layer, a hidden layer and an output layer, and finally output the risk probability of bile duct cancer :
[0088]
[0089] in:
[0090] : activation function;
[0091] : Sigmoid function;
[0092] : weights from the input layer to the hidden layer;
[0093] : The weight from the hidden layer to the output layer;
[0094] : Bias term.
[0095] The loss function uses the binary cross entropy function:
[0096]
[0097] The model training uses the cross entropy loss function as the optimization target, and uses the Adam optimizer for parameter update with a learning rate of 0.001. The mini-batch method is used during training, with a batch size of 32 and a maximum number of training rounds of 100. An early stopping mechanism is also set. If the validation set loss does not decrease within 10 consecutive rounds, the training is terminated early to prevent overfitting. The classification decision threshold is 0.498. When the model output It was diagnosed as bile duct cancer.
[0098] 3. Model parameters and performance comparison
[0099] The bile duct cancer diagnosis model (9-PCM) constructed in the present invention is based on the multi-layer perceptron (MLP) algorithm and includes an input layer (9 nodes), two hidden layers (structures of 30 neurons and 10 neurons), and an output layer (1 node, with a Sigmoid activation function). The Adam optimizer is used for training, with a learning rate of 0.001, a batch size of 32, a maximum number of training rounds of 200, and an early stopping strategy. The loss function is a binary cross-entropy loss function, and the optimization objective is to minimize the prediction error between bile duct cancer (1) and non-bile duct cancer (0) classification. The model outputs a bile duct cancer risk score with a classification threshold of 0.498. If P ≥ 0.498, it is judged as bile duct cancer.
[0100] Actual hyperparameters:
[0101] L2 penalty (regularization term) parameter α=0.01
[0102] Solver: "solver"="adam"
[0103] Adam stability value ϵ=1×10^(-8)
[0104] Optimization tolerance "tol" = 0.0001
[0105] Hidden layer size "hidden_layer_sizes" = (30, 10)
[0106] Initial learning rate "learning_rate_init" = 0.001
[0107] Maximum number of iterations "max_iter" = 200
[0108] Comparative analysis of model performance:
[0109] In the training set (123 cases of cholangiocarcinoma vs 63 cases of non-cholangiocarcinoma), validation set (45 cases vs 35 cases) and independent test set (32 cases vs 22 cases), the model of the present invention was compared with traditional blood tumor markers CA19-9 and CEA in multiple indicators, as shown in Table 2. The ROC curves of the training set, validation set and independent test set are shown in Table 2. Figure 1 、 Figure 2 、 Figure 3 shown.
[0110] Table 2
[0111]
[0112] The above results show that the 9-PCM model demonstrates superior diagnostic performance to traditional tumor markers CA19-9 and CEA in the training set, validation set, and independent test set, especially in terms of sensitivity and F1-score, making it more suitable for early detection and high-risk screening of cholangiocarcinoma.
[0113] The above embodiments are only provided for understanding the method and core concept of the present invention. It should be noted that, without departing from the principles of the present invention, a number of improvements and modifications may be made to the present invention by a person skilled in the art, and such improvements and modifications shall fall within the scope of protection of the claims of the present invention.
Claims
1. A combination biomarker capable of diagnosing or predicting cholangiocarcinoma, wherein the combination biomarker is: IL6, VTCN1, PPY, CEACAM5, CPXM1, FUS, CLMP, KDR, and LGALS7.
2. Use of a reagent for detecting the combined biomarker capable of diagnosing or predicting bile duct cancer as claimed in claim 1 in a sample in the preparation of a product for diagnosing or predicting bile duct cancer.
3. The use according to claim 2, wherein the sample is selected from tissue samples, primary or cultured cells or cell lines, cell supernatants, cell lysates, platelets, serum, plasma, vitreous humor, lymph fluid, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, tumor lysates, tissue culture fluid, or tissue extracts.
4. The use according to claim 3, wherein the sample is selected from tissue sample, platelet, serum, plasma, whole blood, tissue culture fluid, or tissue extract. The use according to claim 3 , wherein the sample is from a human or a non-human mammal. The use according to claim 5 , wherein the sample is from a human.
7. The use according to claim 2, wherein the product comprises a reagent, a kit, or a chip.
8. The use according to claim 7, wherein the reagent comprises a primer, a test paper, or a probe.
9. The use according to claim 7, wherein the kit comprises reagents for extracting the combined biomarker and reagents for quantitative detection of the combined biomarker.
10. A product for diagnosing or predicting bile duct cancer, comprising a reagent for detecting the combined biomarker of claim 1. The product according to claim 10 , wherein the reagent comprises a reagent for detecting mRNA of the combination biomarker and / or a reagent for detecting protein of the combination biomarker.
12. Use of the combined biomarker of claim 1 in constructing a diagnostic / predictive model for cholangiocarcinoma, wherein the diagnostic / predictive model is established using an ensemble learning method.
13. The application according to claim 12, wherein the ensemble learning method comprises a linear regression algorithm, a support vector machine algorithm, a nearest neighbor / k-nearest neighbor algorithm, a logistic regression algorithm, a decision tree algorithm, a k-means algorithm, a random forest algorithm, a naive Bayes algorithm, a dimensionality reduction algorithm, a gradient boosting algorithm, or an MLP algorithm. The application according to claim 13 , wherein the ensemble learning method is an MLP algorithm.
15. The use according to claim 13, wherein the diagnostic / predictive model diagnoses / predicts whether a patient has bile duct cancer or is at risk of bile duct cancer by calculating the AUC value for each sample and comparing it with the optimal threshold value obtained by the ROC curve generated by the prediction results of the model established based on the ensemble learning method on the test set. The use according to claim 15 , wherein the optimal threshold is 0.
498.
17. The use according to claim 16, wherein the diagnosis / prediction model judges the result based on the following criteria: When the AUC value of the sample calculated by the diagnosis / prediction model is greater than or equal to the optimal threshold, the diagnosis / prediction model outputs the result as bile duct cancer or high risk of bile duct cancer; When the AUC value in the sample calculated by the diagnosis / prediction model is less than the optimal threshold, the diagnosis / prediction model outputs a result of non-cholangiocarcinoma or low risk of cholangiocarcinoma.
18. A diagnostic / predictive model product according to any one of claims 12 to 17, comprising an integrated learning module for constructing an MLP structure capable of diagnosing or determining whether a patient has bile duct cancer.
19. The diagnosis / prediction model product according to claim 18, comprising a result judgment module for performing result judgment using an MLP structure.
20. The diagnostic / predictive model product according to claim 19, comprising a detection module for detecting an AUC value in each sample.
21. A system for determining or predicting whether a subject has bile duct cancer, the system comprising: The computer imports the expression level data of the combination biomarker according to claim 1 detected in the sample collected from the subject into the diagnostic prediction model of any one of claims 18 to 20, and obtains the result of the computer judging whether the subject has bile duct cancer based on the result output.
22. A device for determining or predicting whether a subject has bile duct cancer, the device comprising: The computer imports the expression level data of the combination biomarker according to claim 1 detected in the sample collected from the subject into the diagnostic prediction model of any one of claims 18 to 20, and obtains the result of the computer judging whether the subject has bile duct cancer based on the result output.
Citation Information
Patent Citations
Probe set and kit for detecting whole exons of extended genetic diseases and application of probe set
CN110499364A
Bile duct cancer molecular typing gene marker and application thereof
CN116240282A