Method, device, medium and program product for predicting renal cell carcinoma of von Hippel-Lindau syndrome based on DAMs

By constructing a predictive model based on DAMs, and using statistical analysis of plasma and tissue samples to screen out specific metabolites, the problem of early detection of VHL syndrome-related tumors is solved, the accuracy of diagnosis rate and risk assessment is improved, and a personalized treatment strategy is provided.

CN118737430BActive Publication Date: 2025-06-24PEKING UNIVERSITY FIRST HOSPITAL (PEKING UNIVERSITY FIRST CLINICAL MEDICAL COLLEGE)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410729239.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2025-06-24
Estimated Expiration
2044-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to detect and predict VHL syndrome-related tumors, especially renal cell carcinoma, leading to delayed diagnosis and treatment challenges.

Method used

By constructing a predictive model based on DAMs, specific metabolites were screened out using statistical analysis of plasma and tissue samples to construct a model for predicting the risk of renal cancer in VHL syndrome.

Benefits of technology

It significantly improves the early diagnosis rate and risk assessment of renal cancer in VHL syndrome, provides personalized treatment strategies, and improves patient care and treatment effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118737430B_ABST
    Figure CN118737430B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, medium and program product for constructing a prediction model for VHL syndrome renal cancer, and also discloses a method, device, medium and program product for predicting VHL syndrome renal cancer based on DAMs, and a method, device, medium and program product for predicting non-renal organ diseases based on DAMs, which relate to the field of intelligent medicine. The risk of a subject suffering from VHL syndrome renal cancer and non-renal organ diseases is predicted through the biomarker DAMs, so as to strengthen the early diagnosis plan of the diseases and improve the risk assessment strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent medicine, and more particularly, to a method, device, medium and program product for predicting renal cell carcinoma of von Hippel-Lindau (VHL) syndrome based on DAMs. Background Art

[0002] Von Hippel-Lindau (VHL) syndrome is a rare autosomal dominant genetic disease caused by mutations in the VHL gene on chromosome 3, with a genetic penetrance of over 90% by the age of 70. The basic mechanism of VHL syndrome involves dysfunction of the VHL protein, leading to elevated levels of substrates such as HIF-α, which triggers the activation of various carcinogenic factors and plays a central role in increasing the risk of various tumors, including central nervous system hemangioblastoma (CHB), renal cell carcinoma (RCC), retinal hemangioblastoma (RA), pheochromocytoma (PHEO), pancreatic cysts or tumors (PCT), as well as tumors in the reproductive system (GS) and endolymphatic sac.

[0003] In VHL patients, RCC and CHB are the main causes of death. The presence of these complications significantly increases the risk of death from VHL disease, making the prognosis of this genetic disease particularly complex and challenging. Compared with sporadic RCC, VHL-RCC usually appears earlier, mainly between the ages of 30 and 50, and approximately 70% of VHL patients may develop RCC in their 60s. The metastatic progression of RCC is the main cause of death in VHL disease patients. Although nephron-sparing surgery is currently recommended clinically, patients often face a high recurrence rate after surgery, and multiple surgeries also accelerate the end-stage of the disease, thus significantly increasing the physical and economic burden on patients. Therefore, in patients diagnosed with VHL disease, the goals of surgical intervention are different from those in sporadic cases. Surgical intervention focuses more on protecting renal function and reducing the risk of metastasis rather than simply aiming to completely remove the tumor.

[0004] Studies have shown that key factors such as age of onset, family history, and initial symptoms significantly affect the survival of patients. Therefore, it is necessary to explore whether there are new biomarkers that can detect and predict VHL-related tumors, especially renal cell carcinoma, at an early stage. The identification of such biomarkers is crucial for strengthening the early diagnosis plan for renal cell carcinoma and improving the risk assessment strategy in this unique patient population. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a method, device, medium and program product for predicting VHL syndrome renal cell carcinoma and non-renal organ diseases based on DAMs, which can predict the risk of a subject suffering from VHL syndrome renal cell carcinoma and non-renal organ diseases through the biomarker DAMs, strengthen the early diagnosis plan of the diseases and improve the risk assessment strategy.

[0006] The first aspect of the present application discloses a method for constructing a prediction model for VHL syndrome renal cell carcinoma, the method comprising:

[0007] S1: Obtain plasma samples of the plasma training set and tissue samples of the tissue training set; the plasma samples include plasma samples of the VHL-RCC patient group and the healthy group, and the tissue samples include VHL-RCC tissues and corresponding para-carcinoma tissue samples;

[0008] S2: Perform statistical analysis on the plasma samples of the VHL-RCC patient group and the healthy group, and screen out the first DAMs from the statistical analysis results;

[0009] S3: Perform statistical analysis on the VHL-RCC tissues and the corresponding para-carcinoma tissue samples, and screen out the third DAMs from the statistical analysis results;

[0010] S4: Take the intersection of the second DAMs and the fourth DAMs to obtain N VHL-RCC-related DAMs; N is a natural number greater than or equal to 1;

[0011] S5: Screen out M DAMs from the VHL-RCC-related DAMs, and use an algorithm to construct a model for the M DAMs to obtain the prediction model.

[0012] In some embodiments, the statistical analysis in S2 includes:

[0013] Perform differential analysis on the plasma samples of the VHL-RCC patient group and the healthy group to obtain a differential analysis result;

[0014] Perform variable analysis and Mann-Whitney test analysis on the differential analysis result to obtain a test analysis result;

[0015] Screen out the DAMs in the test analysis result and the differential analysis result according to the first criterion to obtain the second DAMs, and the second DAMs are denoted as the first DAMs;

[0016] Optionally, the statistical analysis in S2 further includes: calculating the AUC value of a single DAM in the second DAMs, and screening out the first DAMs from the first DAMs according to the criterion that the AUC is greater than the first threshold;

[0017] Optionally, the machine learning algorithm in S5 is an ensemble learning; the ensemble learning includes any one or more of the following: bagging, random forest, boosting, stacking, preferably random forest;

[0018] Optionally, the method for screening M DAMs in S5 includes: performing importance scoring on the VHL-RCC related DAMs, and selecting the top M DAMs with descending scores;

[0019] Optionally, M is a natural number greater than or equal to 1, preferably 10; M is less than N.

[0020] In some embodiments, the statistical analysis in S3 includes:

[0021] Performing differential analysis on the VHL-RCC tissue and the corresponding adjacent tissue samples to obtain a differential analysis result;

[0022] Performing variable analysis and Mann-Whitney test analysis on the differential analysis result to obtain a test analysis result;

[0023] Screening DAMs in the test analysis result and the differential analysis result according to a first criterion to obtain fourth DAMs, and the fourth DAMs are denoted as the third DAMs;

[0024] Optionally, the statistical analysis in S3 further includes: calculating the AUC value of a single DAM in the fourth DAMs, and screening the third DAMs from the fourth DAMs according to the criterion that the AUC is greater than a first threshold;

[0025] Optionally, the method further includes: when processing the plasma sample and the tissue sample using a mass spectrometer, operating in two ionization modes, namely positive ion mode POS and negative ion mode NEG.

[0026] In some embodiments, the DAMs include N2,N2-Dimethylguanosine and any one or more of the following: PC(16:0 / 16:0), Cysteine-S-sulfate, gamma-Glutamylalanine, 1-deoxy-1-(N6-lysino)-D-fructose, Montecristin, PE(22:4(7Z,10Z,13Z,16Z) / 14:0), PC(16:1(9Z) / P-18:1(11Z)), 1-Kestose, N2-gamma-Glutamylglutamine;

[0027] Optionally, the method for differential analysis includes any one or more of the following: PCA, PLS-DA, OPLS-DA; preferably OPLS-DA;

[0028] Optionally, the variable analysis is univariate analysis;

[0029] Optionally, the first criterion includes: VIP≥1 and P value≤0.05;

[0030] Optionally, the first threshold is 0.65.

[0031] In some embodiments, the method further includes: using regression analysis to analyze the relationship between the concentrations of VHL-RCC related DAMs and the onset age of the training set samples, to obtain a DAMs-age related model;

[0032] Optionally, the regression analysis is Cox regression analysis.

[0033] The second aspect of the present application discloses a method for predicting von Hippel-Lindau syndrome renal cell carcinoma based on DAMs, the method includes:

[0034] Obtaining a plasma sample of a subject;

[0035] Detecting the concentration of the target DAM in the plasma sample;

[0036] Comparing the concentration of the target DAM with a control value to determine the result of the risk of the subject having von Hippel-Lindau syndrome renal cell carcinoma;

[0037] Optionally, the method further includes determining the onset time of the subject based on the concentration of the target DAM;

[0038] Optionally, the method further includes: giving a suggestion on whether the subject needs close monitoring based on the concentration of the target DAM;

[0039] Optionally, the method further includes: obtaining the age of the subject; determining the result of the risk of the subject having von Hippel-Lindau syndrome renal cell carcinoma based on the age;

[0040] Optionally, the method further includes: determining the result of the risk of the subject having von Hippel-Lindau syndrome renal cell carcinoma based on the age and the concentration of the target DAM;

[0041] Optionally, the target DAMs include N2,N2-Dimethylguanosine and any one or more of the following: PC(16:0 / 16:0), Cysteine-S-sulfate, gamma-Glutamylalanine, 1-deoxy-1-(N6-lysino)-D-fructose, Montecristin, PE(22:4(7Z,10Z,13Z,16Z) / 14:0), PC(16:1(9Z) / P-18:1(11Z)), 1-Kestose, N2-gamma-Glutamylglutamine;

[0042] Optionally, the method further includes: inputting the target DAMs into a prediction model to calculate the AUC value of the subject. If the AUC value is greater than or equal to a second threshold, a result that the subject has a high risk of VHL syndrome renal cancer is obtained; if the AUC value is less than the second threshold, a result that the subject has a low risk of VHL syndrome renal cancer is obtained.

[0043] The third aspect of the present application discloses a method for predicting non-renal organ diseases based on DAMs, the method including:

[0044] Obtaining a plasma sample of a subject;

[0045] Detecting the concentration of DAM markers in the plasma sample;

[0046] Judging the result of the risk of the subject having non-renal organ diseases based on the concentration of the DAM markers;

[0047] Optionally, if the non-renal organ disease is CHB, judging the result of the high or low risk of the subject having CHB based on the concentration of DAM markers of any one or more of hypoxanthine, lauroyl carnitine, and 4-dihydroxy-1-pyrrolidine propionamide;

[0048] Optionally, if the level concentration of hypoxanthine increases and the concentrations of lauroyl carnitine and 4-dihydroxy-1-pyrrolidine propionamide decrease, a result that the subject has a high risk of early-stage CHB is obtained;

[0049] Optionally, if the non-renal organ disease is PCT, judging the result of the high or low risk of the subject having PCT based on the concentration of trehalose;

[0050] Optionally, if the plasma concentration of trehalose decreases, a result that the subject has a high risk of PCT is obtained.

[0051] A fourth aspect of the present application discloses a computer device, which includes: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method described in the first aspect or the second aspect or the third aspect of the present application.

[0052] A fifth aspect of the present application discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect or the second aspect or the third aspect of the present application.

[0053] A sixth aspect of the present application discloses a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the steps of the method described in the first aspect or the second aspect or the third aspect of the present application.

[0054] The present application has the following beneficial effects:

[0055] 1. Based on the current situation that VHL-RCC is often underdiagnosed or misdiagnosed due to its rarity, the present application innovatively discloses a method for constructing a VHL syndrome renal cancer prediction model. This method performs LC-MS sequencing on the plasma of VHL patients with RCC and compares it with the peripheral blood of healthy individuals; at the same time, it also analyzes the RCC tumor tissues of VHL patients and uses the matched AN tissues as a control, and takes the intersection of the two to accurately identify the metabolites uniquely related to VHL renal cell carcinoma patients. Moreover, this metabolite shows excellent performance in an independent test set and significantly improves the diagnostic rate, thereby providing a deeper understanding of the unique metabolic characteristics of the disease; the above method can specifically identify the plasma DAMs mainly affected by VHL-RCC and eliminate the potential confounding effects of other VHL syndrome-related lesions (such as pheochromocytoma, pancreatic tumors, reproductive system tumors) on the plasma metabolite levels. When detecting DAMs, differential variables are screened by OPLS-DA, which is a supervised discriminant analysis statistical method that can overcome the defect of being insensitive to variables with relatively small correlations when the unsupervised dimensionality reduction method PCA is executed.

[0056] 2. Based on the fact that the renal lesions in the existing VHL disease show a wide range and diverse lesion types, which significantly increase the risk of renal cell carcinoma, the present application creatively solves the key problem of "whether there are specific metabolic biomarkers that can distinguish VHL renal cell carcinoma patients in a wider population?"

[0057] 3. Monitoring the metabolites disclosed in this application can significantly improve patient care and treatment outcomes, thereby providing in-depth understanding of the metabolic complexity of VHL-RCC. This study not only clarifies the path to better treat this rare cancer, but also sets a new standard for metabolomics research on VHL-RCC. It can achieve timely and effective intervention through enhanced early detection, and guide the formulation of personalized and targeted treatment strategies for VHL-RCC patients;

[0058] 4. When predicting the disease risk of a subject based on DAMs in this application, in addition to comparing the concentrations of DAMs and the control concentration, the AUC value of each subject is calculated based on DAMs, and compared with the optimal threshold (the threshold with the largest Youden index) obtained from the ROC curve generated by the prediction results of the model on the test set to determine whether the predicted sample is VHL-RCC. This prediction method is more accurate. It mainly takes into account that the method used in model construction is randomForest. Since it is essentially an ensemble learning method that makes predictions by constructing multiple decision trees, it is not as accurate as predicting using the calculated AUC value compared to representing the model with an explicit formula similar to linear regression. In this application, the optimal threshold calculated for the validation set is 0.788, but this optimal threshold is not specifically limited and will change with the change of the training set samples. After calculating the prediction probability for clinical samples, it is compared with this optimal threshold. If it is greater than it, it is a high risk; if it is lower than it, it is a low risk.

[0059] 5. When performing mass spectrometry analysis in this application, both positive ion mode (POS) and negative ion mode (NEG) ionization methods are used simultaneously, which can enhance the metabolite coverage range and thus achieve excellent detection performance. In subsequent data analysis, the datasets of the two ionization modes are analyzed separately to ensure the comprehensiveness and accuracy of the metabolic profile. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0061] Figure 1 is a schematic flowchart of the method provided in the first aspect of the embodiment of the present invention;

[0062] Figure 2 is a schematic flowchart of the method provided in the second aspect of the embodiment of the present invention;

[0063] Figure 3 is a schematic flowchart of the method provided in the third aspect of the embodiment of the present invention;

[0064] Figure 4 It is a schematic diagram of the device for predicting VHL syndrome renal cancer based on DAMs provided by an embodiment of the present invention;

[0065] Figure 5 It is a schematic diagram of the device for constructing a prediction model for VHL syndrome renal cancer provided by an embodiment of the present invention;

[0066] Figure 6 It is a schematic diagram of the computer device provided by an embodiment of the present invention;

[0067] Figure 7 It is a schematic diagram of the architecture of an exemplary computing device provided by an embodiment of the present invention;

[0068] Figure 8 It is a schematic diagram of the storage medium provided by an embodiment of the present invention;

[0069] Figure 9 A framework schematic diagram showing the research idea of this application;

[0070] Figure 10 It shows the quality control result diagrams of plasma and tissue samples, where A: Positive ion PCA score diagram of plasma samples; B: Negative ion PCA score diagram of plasma samples. Green dots represent QC samples, and blue dots represent formal experimental samples; C: One-dimensional distribution diagram of positive ion mode PCA-X of QC samples of plasma samples; D: One-dimensional distribution diagram of negative ion mode PCA-X of QC samples of plasma samples; E: Positive ion PCA score diagram of tissue samples; F: Negative ion PCA score diagram of tissue samples. One-dimensional distribution diagrams of positive (G) and negative (H) ion mode PCA-X of QC samples of plasma samples; QC, quality control; PCA, principal component analysis;

[0071] Figure 11 It shows the DAM distribution diagrams in plasma and tissue samples, where A: Scatter score diagram of OPLS-DA analysis of plasma samples; B: Scatter score diagram of OPLS-DA analysis of plasma samples; C: Scatter score diagram of OPLS-DA analysis of tissue samples; C: Scatter score diagram of OPLS-DA analysis of tissue samples; E: Arrangement distribution diagram of R2 and Q2 of plasma samples in POS mode; F: Arrangement distribution diagram of R2 and Q2 of plasma samples in NEG mode; G: Arrangement distribution diagram of R2 and Q2 of tissue samples in POS mode; F: Arrangement distribution of R2 and Q2 of tissue samples in NEG mode;

[0072] Figure 12 It shows the result diagrams of the construction of the VHL-RCC patient diagnosis model based on the training set, where A: Venn diagram; B: ROC curve diagram of the diagnosis model (generated by the top 10 DAMs) calculated using the training set;

[0073] Figure 13 Show the AUC value result graphs of 23 common DAMs; among them, A: POS ion mode; B: NEG ion mode;

[0074] Figure 14 Show the VHL-RCC specific importance score result graphs of 23 common DAMs;

[0075] Figure 15 Show the ROC curve graphs of the test cohort;

[0076] Figure 16 Show the result graphs related to N2,N2-dimethylguanosine, where A: result graphs of N2,N2-dimethylguanosine expression levels in plasma and tissues; B: KM curve analysis graph; C: univariate and multivariate Cox regression result graphs;

[0077] Figure 17 Show the result graphs of the correlation between the plasma concentration of metabolites and non-renal organ lesions, where A: there is a correlation between the plasma concentration of xanthine and the early onset of CHB; B: there is a correlation between the plasma concentration of dodecanoyl carnitine and the early onset of CHB; C: the correlation between the plasma concentration of 4-dihydroxy-2-hydroxymethyl-1-pyrrolidine propionamide (4D2h1p) and the early onset of CHB; D: there is a correlation between the plasma concentration of trehalose and the early onset of PCT;

[0078] Figure 18 Show the expression analysis of 23 DAMs in plasma and tissue samples; DAMs, differentally abundant metabolites; VHL, von Hippel-Lindau disease; POS, positive ionization mode; NEG, negative ionization mode; VIP, variable importance in the projection; AUC, area under the curve; HC, healthy control; FC, fold change. Detailed implementation manners

[0079] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0080] In some processes described in the specification, claims and the above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the sequence, and do not limit that "first" and "second" are of different types.

[0081] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0082] Figure 2 It is a schematic flowchart of a method for predicting VHL syndrome renal cancer based on DAMs provided by an embodiment of the present invention. Specifically, the method includes the following steps:

[0083] 101: Obtain a plasma sample of the subject;

[0084] In some embodiments, the term "subject" or "test subject" or "test sample" used herein refers to any animal (e.g., mammal), including but not limited to humans, non-human primates, rodents, etc., which will be the recipient of a particular treatment. Generally, the terms "subject" and "patient" may be used interchangeably herein when referring to human subjects. Preferably, the subject is a human.

[0085] In some embodiments, the subject is a patient clinically used for prognostic evaluation.

[0086] 102: Detect the concentration of the target DAM in the plasma sample;

[0087] 103: Compare the concentration of the target DAM with a control value to determine the result of the risk of the subject suffering from VHL syndrome renal cancer;

[0088] In some embodiments, if N2,N2-Dimethylguanosine, gamma-Glutamylalanine, PC(16:1(9Z) / P-18:1(11Z)), 1-deoxy-1-(N6-lysino)-D-fructose,

[0089] If the concentration of any one or more of N2-gamma-Glutamylglutamine and Cysteine-S-sulfate is lower than the control value, it is concluded that the subject has a high risk of renal cancer with VHL syndrome; if the concentration of any one or more of Montecristin, PE(22:4(7Z,10Z,13Z,16Z) / 14:0), 1-Kestose, and PC(16:0 / 16:0) is higher than the control value, it is concluded that the subject has a high risk of renal cancer with VHL syndrome;

[0090] In some embodiments, the method further includes determining the onset time of the subject based on the concentration of the target DAM; optionally, the method further includes: giving a suggestion on whether the subject needs intensive monitoring based on the concentration of the target DAM; as Figure 16 shown in B, the lower the concentration of this metabolite, the more likely it is that the patient will develop renal cancer at a younger stage, indicating that such patients need more intensive monitoring.

[0091] In some embodiments, the method further includes: obtaining the age of the subject; determining the result of the subject's risk of renal cancer with VHL syndrome based on the age; as Figure 16 shown in C, the HR of the birth year is greater than 1, that is, the older the age, the more likely it is to develop renal cancer.

[0092] In some embodiments, the method further includes: determining the result of the subject's risk of renal cancer with VHL syndrome based on the age and the concentration of the target DAM.

[0093] In some embodiments, the target DAM includes N2,N2-Dimethylguanosine (N2, N2-dimethylguanosine) and any one or more of the following: PC(16:0 / 16:0), Cysteine-S-sulfate, gamma-Glutamylalanine, 1-deoxy-1-(N6-lysino)-D-fructose, Montecristin, PE(22:4(7Z,10Z,13Z,16Z) / 14:0), PC(16:1(9Z) / P-18:1(11Z)), 1-Kestose, N2-gamma-Glutamylglutamine;

[0094] In some embodiments, the method further includes: inputting the target DAMs into a prediction model to calculate the AUC value of the subject; if the AUC value is greater than or equal to a second threshold, obtaining the result that the subject has a high risk of VHL syndrome renal cancer; if the AUC value is less than the second threshold, obtaining the result that the subject has a low risk of VHL syndrome renal cancer.

[0095] The method for constructing the prediction model includes:

[0096] S1: Obtaining plasma samples of a plasma training set and tissue samples of a tissue training set; the plasma samples include plasma samples of a VHL-RCC patient group and a healthy group, and the tissue samples include VHL-RCC tissues and corresponding para-cancerous tissue samples;

[0097] S2: Performing statistical analysis on the plasma samples of the VHL-RCC patient group and the healthy group, and screening out the first DAMs (166 + 72) from the statistical analysis results;

[0098] S3: Performing statistical analysis on the VHL-RCC tissues and the corresponding para-cancerous tissue samples, and screening out the third DAMs (158 + 83) from the statistical analysis results;

[0099] S4: Taking the intersection of the second DAMs and the fourth DAMs to obtain N VHL-RCC related DAMs; N is a natural number greater than or equal to 1;

[0100] S5: Screening out M DAMs from the VHL-RCC related DAMs, and using an algorithm to construct a model for the M DAMs to obtain the prediction model.

[0101] In some embodiments, the statistical analysis in S2 includes:

[0102] Performing differential analysis on the plasma samples of the VHL-RCC patient group and the healthy group to obtain differential analysis results;

[0103] Performing variable analysis and Mann-Whitney test analysis on the differential analysis results to obtain test analysis results;

[0104] Screening out DAMs in the test analysis results and the differential analysis results according to a first criterion to obtain the second DAMs (1306 + 1211), and the second DAMs are denoted as the first DAMs;

[0105] Optionally, the statistical analysis in S2 further includes: calculating the AUC value of a single DAM in the second DAMs, and screening out the first DAMs from the first DAMs according to the criterion that the AUC is greater than a first threshold;

[0106] Optionally, the machine learning algorithm in S5 is an ensemble learning; the ensemble learning includes any one or more of the following: bagging, random forest, boosting, stacking, preferably random forest;

[0107] Optionally, the screening method for the M DAMs in S5 includes: performing importance scoring on the VHL-RCC related DAMs, and selecting the M DAMs with scores from high to low;

[0108] Optionally, M is a natural number greater than or equal to 1, preferably 10; M is less than N.

[0109] In some embodiments, the statistical analysis in S3 includes:

[0110] Performing differential analysis on the VHL-RCC tissue and the corresponding adjacent tissue samples to obtain a differential analysis result; optionally, the differential analysis method includes any one or more of the following: PCA, PLS-DA, OPLS-DA; preferably OPLS-DA;

[0111] Performing variable analysis and Mann-Whitney test analysis on the differential analysis result to obtain a test analysis result; optionally, the variable analysis is univariate analysis;

[0112] Screening DAMs in the test analysis result and the differential analysis result according to a first criterion to obtain the fourth DAMs (1043 + 1338), and the fourth DAMs are denoted as the third DAMs; optionally, the first criterion includes: VIP≥1 and P value≤0.05;

[0113] Optionally, the statistical analysis in S3 further includes: calculating the AUC value of a single DAM in the fourth DAMs, and screening the third DAMs from the fourth DAMs according to the criterion that AUC is greater than a first threshold; optionally, the first threshold is 0.65.

[0114] Optionally, the method further includes: when processing the plasma sample and the tissue sample using a mass spectrometer, operating in two ionization modes, positive ion mode POS and negative ion mode NEG.

[0115] In some embodiments, the DAMs include N2,N2-Dimethylguanosine and any one or more of the following: PC(16:0 / 16:0), Cysteine-S-sulfate, gamma-Glutamylalanine, 1-deoxy-1-(N6-lysino)-D-fructose, Montecristin, PE(22:4(7Z,10Z,13Z,16Z) / 14:0), PC(16:1(9Z) / P-18:1(11Z)), 1-Kestose, N2-gamma-Glutamylglutamine;

[0116] In some embodiments, the method further includes: analyzing the relationship between the levels of VHL-RCC-related DAMs and the onset age in the training set using regression analysis to obtain a DAMs level-onset age-related model;

[0117] Optionally, the regression analysis is Cox regression analysis.

[0118] In the present invention, the indicators for prognosis include objective remission rate, overall survival rate (OS), progression-free survival (PFS), objective response rate (ORR), time to progress (TTP), disease-free survival (DFS), time to treatment failure (TTF), response rate (RR), complete response (CR), partial response (PR). The term "prognosis" is well recognized in the art and includes predictions regarding the likely course or development of a disease, particularly regarding the likelihood of disease remission, disease recurrence, tumor recurrence, metastasis, and death. "Good prognosis" means the likelihood that a patient with cancer, particularly pancreatic cancer, remains disease-free (i.e., cancer-free). "Poor prognosis" means the likelihood of recurrence or reoccurrence, metastasis, or death of a potential cancer or tumor.

[0119] Figure 3 It is a schematic flowchart of a method for predicting non-renal organ diseases based on DAMs provided by an embodiment of the present invention. Specifically, the method includes the following steps:

[0120] 201: Obtain a plasma sample from a subject;

[0121] 202: Detect the concentration of DAM markers in the plasma sample;

[0122] 203: Based on the concentration of the DAM marker, determine the result of the risk of the subject suffering from non-renal organ diseases;

[0123] In some embodiments, if the non-renal organ disease is CHB, based on the concentration of DAM markers of any one or several of hypoxanthine, lauroyl carnitine, and 4-dihydroxy-1-pyrrolidine propionamide, determine the result of the high or low risk of the subject suffering from CHB;

[0124] Optionally, if the concentration of hypoxanthine increases and the concentrations of lauroyl carnitine and 4-dihydroxy-1-pyrrolidine propionamide decrease, obtain the result of the subject having a high risk in the early stage of CHB;

[0125] In some embodiments, if the non-renal organ disease is PCT, based on the concentration of trehalose, determine the result of the high or low risk of the subject suffering from PCT;

[0126] Optionally, if the plasma concentration of trehalose decreases, obtain the result of the subject having a high risk of PCT.

[0127] During the research process, as Figure 9 shown, the determination of DAM markers for non-renal organ diseases is obtained through the analysis of metabolites in the plasma of the training set samples and does not involve tissues.

[0128] I. Information on plasma samples of the plasma training set and tissue samples of the tissue training set

[0129] Patients diagnosed with VHL syndrome at the First Hospital of Peking University, the only international VHL consortium clinical care center in China, participated in this study. The inclusion criteria are as follows:

[0130] (1) Genetic confirmation of VHL gene point mutations or fragment deletions by Sanger sequencing or next-generation sequencing;

[0131] (2) Clinically diagnosed with VHL syndrome, and at least one family member has undergone genetic testing;

[0132] (3) Radiological evidence (ultrasound, CT, MRI) suggesting renal occupancy and suspected renal cell carcinoma;

[0133] (4) No history of radiotherapy, radical surgery, or any palliative surgical intervention before participating in the study.

[0134] Peripheral blood samples were collected from patients diagnosed with VHL-RCC and healthy individuals from the same institution for comparison. RCC tumors and AN samples (adjacent normal tissues) were taken from VHL patients undergoing surgical treatment for renal cancer. The collection of these samples was an integral part of the study, allowing subsequent comparative analysis between the plasma of VHL patients and healthy subjects, as well as between tumor and AN samples. Patient demographic data, including gender, age, family history, generation of onset, age of onset, and pathological type, were obtained through the clinical registration system or by directly consulting the patients and their relatives. The age of onset was defined as the age at which the patient first presented symptoms related to VHL manifestations.

[0135] This project followed the principles and spirit of the Declaration of Helsinki and was approved by the Medical Ethics Committee of Peking University First Hospital in Beijing, China. Informed consent was obtained from all participants after they had a full understanding of the research process and potential outcomes.

[0136] II. Methods

[0137] 1. Sample pretreatment and preservation

[0138] Whole blood samples from all study subjects were collected on an empty stomach in the morning to minimize the effects of food and circadian variations on the plasma levels of low-molecular-weight metabolites. These samples were then placed in tubes containing ethylenediaminetetraacetic acid (EDTA) as an anticoagulant. Subsequently, the samples were centrifuged at 3500 rpm for 10 minutes to separate the plasma. After centrifugation, all samples were systematically cataloged and stored at -80 °C for future experimental analysis. RCC tissues and corresponding AN samples were collected immediately after partial or radical nephrectomy. The samples were quickly rinsed with PBS to remove any surface blood clots, snap-frozen in liquid nitrogen, and subsequently transferred to an -80 °C freezer for long-term storage pending further sequencing analysis.

[0139] 2. Metabolite extraction and LC-MS / MS analysis

[0140] To extract metabolites, both plasma and tissue samples underwent similar chemical processes. First, 50 μL of each sample was transferred to an EP tube. For tissue samples, a homogenization or pulverization step was also performed before further processing to ensure effective cell lysis and metabolite release. Subsequently, 200 μL of a 1:1 acetonitrile-methanol mixture containing isotope standards was added to each sample (either plasma or tissue). The mixture was then vortexed for 30 seconds and sonicated in an ice-water bath for 10 minutes. After that, the samples were incubated at -40 °C for 1 hour to promote protein precipitation, and then centrifuged at 12,000 rpm for 15 minutes at 4 °C. Then, the supernatants of plasma and tissue samples were collected for subsequent analysis. Finally, equal volumes of supernatants from all processed samples were mixed together to prepare quality control (QC) samples. LC-MS / MS analysis was performed on a UHPLC system (Vanquish, Thermo Fisher Scientific) connected to a Q Exactive HFX Orbitrap mass spectrometer (Thermo). This platform used an electrospray ionization source and operated in two ionization modes: positive ion mode (POS) and negative ion mode (NEG). When applied to metabolomics, using both ionization methods simultaneously could improve metabolite coverage, thus achieving excellent detection performance. In subsequent data analysis, the datasets of the two ionization modes were analyzed separately to ensure the comprehensiveness and accuracy of metabolic analysis. Chromatographic separation was carried out using a 2.1 mm × 100 mm UPLC BEH Amide column (1.7 μm particles). Elution was performed with a mobile phase consisting of 25 mmol / L ammonium acetate and 25 ammonia water (pH adjusted to 9.75) as solvent A and acetonitrile as solvent B. Samples were kept at 4 °C in the autosampler, and the injection volume for analysis was set to 3 μL. The QE HFX mass spectrometer was selected and controlled by Xcalibur software (Thermo) because it was capable of obtaining MS / MS spectra through the information-dependent acquisition (IDA) mode. In this configuration, the software continuously analyzed the full-scan MS spectra. The ESI source parameters were: sheath gas flow rate 30 Arb, auxiliary gas flow rate 25 Arb, capillary temperature 350 °C, full MS resolution 60000, MS / MS resolution 7500, and collision energies in the NCE mode were 10, 30, 60 respectively; the spray voltage for instrument operation was 3.6 kV in the positive mode and -3.2 kV in the negative mode.

[0141] 3. Data Processing and Model Development

[0142] The raw data was converted to the mzXML format using ProteoWizard software. In the subsequent processing stages, including peak identification, extraction, alignment, and integration, custom-developed programs based on the R language and the XCMS platform were utilized. For the metabolite annotation step, an in-house developed MS2 database (BiotreeDB) was used, and the annotation threshold was set at 0.3. Peak filtering was performed to eliminate noise and filter out biases based on the relative standard deviation or coefficient of variation. The criteria for peak retention included no more than 50% missing values in the peak area data for a single group or all groups. Missing values in the raw data were addressed by imputation, replacing them with half of the detected minimum value. Data normalization was performed using an internal standard (IS) to ensure consistency. Subsequently, orthogonal projections to latent structures discriminant analysis (OPLS-DA) was utilized for more efficient analysis. The data was log-transformed and UV-scaled using SIMCA software (version 16.0.2, Sartorius Stedim DataAnalytics AB, Umea, Sweden). Initially, an OPLS-DA model focusing on the first principal component was constructed. Seven-fold cross-validation was used to evaluate the quality of the model. The effectiveness of the model was evaluated based on the R 2 Y (interpretability of the model for the categorical variable Y) and Q2 (predictive ability of the model). Finally, a permutation test was conducted by randomly changing the order of the categorical variable Y multiple times to generate various random Q2 values to further validate the effectiveness of the model. After model construction, a permutation test was performed to verify the statistical significance and predictive ability of the model. This test involved randomly shuffling the class labels and reconstructing the model multiple times to evaluate whether the performance of the original model was significantly better than these random iterations, thus confirming its robustness and reliability.

[0143] 4. Identification of Differentially Abundant Metabolites (DAMs)

[0144] Based on the criteria of VIP (Variable Importance in Projection) ≥ 1 and P-value ≤ 0.05, DAMs were identified using the VIP from the first principal component projection and the P-value obtained from the Mann-Whitney test. To identify metabolites with high sensitivity and high specificity, a 5-fold cross-validation method was adopted for Logistic regression modeling to analyze metabolomics data. For each fold, the model was trained on a subset of the data and validated on a separate test set, and the area under the curve (AUC) for each metabolite was calculated. The average AUC value for each metabolite across all folds was calculated. This process evaluated the effectiveness of each metabolite in differentiating different conditions (such as VHL patients and control groups). Metabolites with an AUC greater than 0.65 were selected as DAMs for subsequent analysis.

[0145] 5. Quantification and Statistical Analysis

[0146] To accurately distinguish healthy individuals from VHL patients with RCC using plasma metabolites, a diagnostic model was developed through a series of analyses. Plasma group subjects were randomly divided into a training group and a test group at a ratio of 4:1. The training group included 48 VHL patients and 24 healthy volunteers, while the independent test group included 13 VHL patients and 7 healthy volunteers (Table S3).

[0147] Table S3

[0148]

[0149]

[0150] Using the randomForest package in R, importance scores of 23 RCC-specific abundant metabolites (DAMs) shared by plasma and tissue were calculated in the training cohort. The top 10 DAMs ranked by important scores were selected to establish a diagnostic model based on the training cohort. The efficiency of the model was evaluated using the area under the ROC curve, which was calculated by the pROC package in R.8T to mitigate overfitting, and the predictive performance of the model was further evaluated using the randomForest package on the independent test set.

[0151] To identify DAMs associated with tumor risk related to the age of VHL patients, the relationship between DAM levels and the age of onset of various VHL manifestations was investigated. In this study, the plasma cohort of VHL patients was randomly divided into a training set and a validation set at a ratio of 4:1, enabling univariate and multivariate Cox regression analyses to examine the associations of interest. The predictive accuracy of the risk prediction model was evaluated using the concordance index (c-index). To further explore DAMs associated with the onset of RCC, the intersection of DAMs in plasma and tissue was re-analyzed using Cox.

[0152] The research method of the present invention is as Figure 9 shown.

[0153] II. Results

[0154] 1. Clinical and genetic characteristics

[0155] Table 1 and Table 2 respectively describe a cohort of 61 VHL patients with RCC and 31 healthy subjects through non-targeted metabolomics analysis of plasma samples. The VHL patient cohort showed a balanced gender composition (49.2% male and 50.8% female). The onset of VHL-related symptoms mainly occurred before the age of 30, accounting for 62.3%. The most common initially affected organ was renal cell carcinoma (42.6%), followed by chronic hepatitis B (CBH) (26.2%). A detailed assessment of organ involvement showed that 70.5% of RCC patients also frequently had PCT, and 54.1% of patients frequently had CHB. In Table 2, chi-square analysis showed that there were no statistically significant differences in the gender or year of birth distribution between the healthy individual group and the VHL patient group. Table 3 details the clinical profiles of tissue samples from 19 VHL-RCC patients. This group was predominantly male (63.2%), and most diagnoses occurred before the age of 30 (57.9%). Most had a family history of VHL (68.4%). Almost all patients showed type I VHL (89.5%), and RCC occurred in all cases, with PCT accounting for 63.2%.

[0156] Table 1. Clinical Characteristics of VHL Syndrome Renal Cancer Patients in Non-Targeted Plasma Metabolomics Analysis

[0157]

[0158]

[0159] Table 2. Comparison of Clinical Characteristics between Patients and Healthy Donors in Non-Targeted Plasma Metabolomics Analysis

[0160]

[0161] Table 3. Clinical Characteristics of VHL Syndrome Renal Cancer Patients in Non-Targeted Tissue Metabolomics Analysis

[0162]

[0163]

[0164] 2. Quality Control (QC) of Data

[0165] In plasma samples, the initial numbers of identified POS and NEG precursor molecules were 5,457 and 5,281, respectively, which decreased to 4,122 and 3,862 after the processing and filtering steps mentioned in the method section. Similarly, in tissue samples, following the same processing protocol, the counts of POS and NEG precursors decreased from 4,801 to 3,576 and from 5,187 to 4,076, respectively. In theory, the QC samples should be the same, but variations in substance extraction and analytical detection may lead to differences between them. The smaller these differences are, the higher the stability of the method and the higher the quality of the data. Figure 10 (A,B) and Figure 10 (E,F) respectively show significant clustering of plasma and tissue QC samples, which emphasizes the robustness and stability of the method. Figure 10 (C,D) and Figure 10 (G,H) indicate that all QC samples are within the ±2STD (standard deviation) range, indicating reliable quality of the experimental data.

[0166] 3. DAM Distribution in Plasma and Tissue Samples

[0167] Metabolites detected from plasma and tissue samples were all used for OPLS-DA, which can effectively reduce model complexity and enhance interpretability, maximizing the visualization of differences between groups. Scatter score plots of OPLS-DA analysis of plasma and tissue samples in POS and NEG ionization modes are shown as Figure 11 A-D. These plots illustrate significant differences between groups, while the differences within groups are relatively insignificant.

[0168] To evaluate the effectiveness of our model and prevent overfitting, a permutation test (n = 200 permutations) was conducted. The permutation distributions of R2 and Q2 for plasma and tissue samples in POS and NEG modes are depicted in the histograms ( Figure 11 E-H). The R2 value of the original model (representing the explained variance) was significantly higher than that of the permutation model, and the actual value was at the far right of the distribution. This indicates that the model explained the variance within the data to a certain extent, which was unlikely to be accidental. Similarly, the Q2 value representing the prediction accuracy of the model was significantly better than the permutation results. The actual Q2 was at the extreme of the permutation distribution, exceeding the 95th percentile of the permutation Q2 values. This distinct contrast emphasizes the prediction reliability of the model and indicates its high effectiveness in terms of prediction results based on metabolomic data.

[0169] Comparative analysis between the subsequent two groups was performed using univariate analysis and the Mann-Whitney U test. Metabolites with a P-value less than 0.05 in the test were intersected with metabolites having a VIP value greater than 1 in the OPLS-DA model. This method generated a total of 1306 POSDAM and 1211 NEGDAM for the plasma group. Similarly, for the tissue group, 1043 POSDAM and 1338 NEGDAM were identified. For each comparison, a heatmap was constructed to show the top 10 upregulated and downregulated DAM in VHL patients and the tumor group.

[0170] To further refine the selection of DAM with higher specificity and sensitivity, DAM with an AUC greater than 0.65 from plasma and tissue sources were filtered for subsequent analysis. After this selection process, the plasma group retained 166 DAM in POS and 72 DAM in NEG, while the tissue group retained 158 DAM in POS and 83 DAM in NEG.

[0171] 4. DAMs Specifically Associated with VHL-RCC

[0172] To identify DAMs specifically associated with VHL patients with RCC, DAM with an AUC greater than 0.65 from plasma and tissue sources were intersected to obtain common DAM. The Venn diagram illustrates the number of intersections between different groups ( Figure 12 A in Figure 13 ), showing that among 238 DAM from plasma and 241 DAM from tissue, 23 DAM are common. The lollipop chart shows the AUC values of these 23 common DAM in the POS ( Figure 13 A in

[0173] 5. Predictive Ability of DAMs (Differentially Abundant Metabolites) for VHL-RCC

[0174] To accurately identify VHL patients with RCC in a wide population, a diagnostic model was constructed using the top 10 DAM selected from 23 common pools according to the VHL-RCC specific importance score calculated using the randomForest software package ( Figure 14)。The top 10 DAMs are PC(16:0 / 16:0), Cysteine-S-sulfate, gamma-Glutamylalanine, 1-deoxy-1-(N6-lysino)-D-fructose, Montecristin, PE(22:4(7Z,10Z,13Z,16Z) / 14:0), PC(16:1(9Z) / P-18:1(11Z)), 1-Kestose, N2-gamma-Glutamylglutamine, N2,N2-Dimethylguanosine. This model demonstrated high efficiency in differentiating healthy individuals from VHL patients with RCC in the training cohort( Figure 12 B) in

[0175] In addition, the effectiveness of this model was confirmed in an independent test cohort, showing excellent validation metrics (AUC = 0.989, 95% CI = 0.959 - 1.000, Figure 15 )。Among these 10 DAMs, we identified N2,N2-Dimethylguanosine as a potential predictor of RCC development in VHL patients. Its level in plasma was significantly downregulated in VHL-RCC patients compared to healthy individuals (P < 0.001). In the comparison between RCC tumors and AN tissues of VHL patients, the level of N2,N2-Dimethylguanosine in tumor tissues was significantly decreased (P < 0.001)( Figure 16 A) in Figure 18 As shown in

[0176] for the analysis of 23 DAMs in plasma and tissue samples, the DAM concentrations of N2,N2-Dimethylguanosine, gamma-Glutamylalanine, PC(16:1(9Z) / P-18:1(11Z)), 1-deoxy-1-(N6-lysino)-D-fructose, N2-gamma-Glutamylglutamine, Cysteine-S-sulfate were lower than the control values in the training set samples of patients with VHL syndrome renal cancer, and the DAM concentrations of Montecristin, PE(22:4(7Z,10Z,13Z,16Z) / 14:0), 1-Kestose, PC(16:0 / 16:0) were higher than the control values.KM curve analysis showed a significant correlation between the plasma level of N2,N2-dimethylguanosine and the occurrence of renal cell carcinoma in patients ( Figure 16 B in). A lower level of N2,N2-dimethylguanosine indicated an earlier onset of RCC (P = 0.0013). Univariate and multivariate Cox regression analyses both confirmed that N2,N2-dimethylguanosine and the patient's birth year could independently predict the onset time of RCC in VHL patients ( Figure 16 C in). The c-index of the internal validation model was 0.81, indicating a high prediction accuracy.

[0177] 6. Plasma DAM levels and the occurrence of non-renal organ lesions

[0178] Consistent with the background information, VHL patients often present with lesions in multiple organs simultaneously. Our aim was to investigate whether certain plasma-derived DAMs could serve as predictors for the onset of non-renal organ diseases such as CHB, PCT, PHEO, RA, and GS. Kaplan-Meier analysis found a significant correlation between the plasma concentrations of hypoxanthine, lauroyl carnitine, and 4-dihydroxy-2-hydroxymethyl-1-pyrrolidine propionamide (4D2h1p) and the early onset of CHB ( Figure 17 A-C in). An increased plasma level of hypoxanthine (P = 0.0038), as well as decreased levels of lauroyl carnitine (P = 0.044) and 4D2h1p (P = 0.045), indicated an earlier occurrence of CHB. Additionally, a lower plasma concentration of trehalose was associated with an earlier onset of PCT (P < 0.029) ( Figure 17 D in). Due to the small sample sizes of patients with PHEO, RA, and GS, no significant DAMs for these conditions were found. The full names of CHB, PCT, PHEO, RA, and GS in Chinese and English are as follows: central nervous system hemangioblastomas (CHB), renal cell carcinomas (RCC), retinal haemangioblastoma (RA), pheochromocytomas (PHEO), pancreatic cysts or tumors (PCT), tumors in the genital system (GS); central nervous system hemangioblastomas (CHB), renal cell carcinomas (RCC), retinal haemangioblastoma (RA), pheochromocytomas (PHEO), pancreatic cysts or tumors (PCT), tumors in the genital system (GS).

[0179] Figure 6 is a schematic diagram of a computer device provided by an embodiment of the present invention, as Figure 6As shown, the device may include: one or more processors, and one or more memories; wherein, computer-readable code is stored in the memory, and when the computer-readable code is run by the one or more processors, the above-described method can be executed.

[0180] The processor in this embodiment may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, operations and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.

[0181] Generally speaking, various exemplary embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0182] For example, the method or device according to the embodiments of the present disclosure may also be implemented by means of Figure 7 the architecture of the computing device 3000 shown. As Figure 7 shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for the processing and / or communication of the method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 7 the architecture shown is only exemplary, and when implementing different devices, one or more components shown in the Figure 7 computing device may be omitted according to actual needs.

[0183] An embodiment of the present invention also provides a computer-readable storage medium, such as Figure 8As shown, it is a schematic diagram of a storage medium provided by an embodiment of the present invention. Computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the methods according to the embodiments of the present disclosure described with reference to the above drawings can be executed. The computer-readable storage medium in the embodiments of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memories of the methods described herein are intended to include but not be limited to these and any other suitable types of memories. It should be noted that the memories of the methods described herein are intended to include but not be limited to these and any other suitable types of memories.

[0184] The embodiments of the present disclosure also provide a computer program product or a computer program, which implements the steps of the above method when executed by a processor.

[0185] In some embodiments, the present embodiment also discloses a device for predicting VHL syndrome renal cancer based on DAMs, as Figure 4 shown, the device includes:

[0186] A first acquisition module 301, configured to acquire a plasma sample of a subject;

[0187] A first concentration detection module 302, configured to detect the concentration of a target DAM in the plasma sample;

[0188] A first judgment module 303, configured to compare the concentration of the target DAM with a control value to judge the result of the risk that the subject has VHL syndrome renal cancer; specifically, if N2,N2-Dimethylguanosine, gamma-Glutamylalanine, PC(16:1(9Z) / P-18:1(11Z)), 1-deoxy-1-(N6-lysino)-D-fructose,

[0189] If the concentration of any one or more of the target DAMs in N2-gamma-Glutamylglutamine and Cysteine-S-sulfate is lower than the control value, the result is that the subject has a high risk of VHL syndrome renal cancer; if the concentration of any one or more of the target DAMs in Montecristin, PE(22:4(7Z,10Z,13Z,16Z) / 14:0), 1-Kestose, and PC(16:0 / 16:0) is higher than the control value, the result is that the subject has a high risk of VHL syndrome renal cancer;

[0190] In some embodiments, the device further includes: a second acquisition module, configured to acquire the age of the subject; a second judgment module, configured to judge the result of the subject's risk of VHL syndrome renal cancer based on the age;

[0191] In some embodiments, the device further includes: a third judgment module, configured to judge the result of the subject's risk of VHL syndrome renal cancer based on the age and the concentration of the target DAM;

[0192] In some embodiments, the device further includes: a calculation module, configured to input the target DAMs into a prediction model to calculate the AUC value of the subject. If the AUC value is greater than or equal to a second threshold, the result is that the subject has a high risk of VHL syndrome renal cancer; if the AUC value is less than the second threshold, the result is that the subject has a low risk of VHL syndrome renal cancer.

[0193] In some embodiments, the present embodiment also discloses a device for constructing a VHL syndrome renal cancer prediction model, as Figure 5 shown, the device includes:

[0194] A fourth acquisition module 401, configured to acquire a plasma sample of the subject;

[0195] A second concentration detection module 402, configured to detect the concentration of DAM markers in the plasma sample;

[0196] A fourth judgment module 403, configured to judge the result of the subject's high or low risk of non-renal organ diseases based on the concentration of the DAM markers;

[0197] In some embodiments, the device further includes a CHB judgment module. If the non-renal organ disease is CHB, judge the result of the subject's high or low risk of CHB based on the concentration of any one or more of the DAM markers in hypoxanthine, lauroyl carnitine, and 4-dihydroxy-1-pyrrolidine propionamide; if the concentration of hypoxanthine increases and the concentrations of lauroyl carnitine and 4-dihydroxy-1-pyrrolidine propionamide decrease, the result is that the subject has a high risk of early CHB;

[0198] In some embodiments, the device further includes a PCT determination module, which determines the result of whether the subject has a high or low risk of PCT based on the trehalose concentration if the non-renal organ disease is PCT; if the plasma concentration of trehalose decreases, it is obtained that the subject has a high risk of PCT.

[0199] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0200] Generally speaking, the various example embodiments of the present disclosure can be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or a controller or other computing devices, or some combination thereof.

[0201] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0202] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0203] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0204] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0205] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. A method for constructing a VHL syndrome renal cancer prediction model, characterized in that: The method comprises: S1: Obtain plasma samples of a plasma training set and tissue samples of a tissue training set; the plasma samples include plasma samples of a VHL-RCC patient group and a healthy group, and the tissue samples include VHL-RCC tissues and corresponding adjacent cancer tissue samples; S2: Statistical analysis was performed on the plasma samples of the VHL-RCC patient group and the healthy group, and the first DAMs were screened out from the statistical analysis results; S3: Statistical analysis was performed on VHL-RCC tissues and corresponding adjacent paracancerous tissue samples, and the third DAMs were screened out from the statistical analysis results; S4: taking the intersection of the first DAMs and the third DAMs to obtain N VHL-RCC related DAMs; N is a natural number greater than or equal to 1; S5: Screening out M DAMs from the VHL-RCC-related DAMs, building a model for the M DAMs using an algorithm, and obtaining the prediction model.

2. The method for constructing a VHL syndrome renal cancer prediction model according to claim 1, characterized in that: The statistical analysis in S2 includes: Performing differential analysis on the plasma samples of the VHL-RCC patient group and the healthy group to obtain differential analysis results; Performing variable analysis and Mann-Whitney test analysis on the difference analysis results to obtain test analysis results; The DAMs in the test analysis results and the difference analysis results are screened according to the first standard to obtain second DAMs, which are recorded as the first DAMs.

3. The method for constructing a VHL syndrome renal cancer prediction model according to claim 2, characterized in that: The statistical analysis in S2 further includes: calculating the AUC value of a single DAM in the second DAMs, and screening the first DAMs from the second DAMs based on a criterion that the AUC is greater than a first threshold.

4. The method for constructing a VHL syndrome renal cancer prediction model according to claim 1, characterized in that: The machine learning algorithm in S5 is ensemble learning; the ensemble learning includes any one or more of the following: self-service aggregation, random forest, boosting method, and stacking method.

5. The method for constructing a VHL syndrome renal cancer prediction model according to claim 1, characterized in that: The screening method of the M DAMs in S5 includes: scoring the importance of the VHL-RCC-related DAMs, and selecting M DAMs with scores from high to low.

6. The method for constructing a VHL syndrome renal cancer prediction model according to claim 5, characterized in that: M is a natural number greater than or equal to 1, and M is less than N.

7. The method for constructing a VHL syndrome renal cancer prediction model according to claim 6, characterized in that: M is 10.

8. The method for constructing a VHL syndrome renal cancer prediction model according to claim 1, characterized in that: The statistical analysis in S3 includes: Performing differential analysis on the VHL-RCC tissue and the corresponding adjacent cancer tissue samples to obtain differential analysis results; Performing variable analysis and Mann-Whitney test analysis on the difference analysis results to obtain test analysis results; The DAMs in the test analysis results and the difference analysis results are screened according to the first standard to obtain fourth DAMs, which are recorded as the third DAMs.

9. The method for constructing a VHL syndrome renal cancer prediction model according to claim 8, characterized in that: The statistical analysis in S3 further includes: calculating the AUC value of a single DAM in the fourth DAMs, and screening the third DAMs from the fourth DAMs based on a criterion that the AUC is greater than a first threshold.

10. The method for constructing a VHL syndrome renal cancer prediction model according to claim 1, characterized in that: The method further includes: when processing the plasma sample and the tissue sample using a mass spectrometer, operating in two ionization modes: a positive ion mode POS and a negative ion mode NEG.

11. The method for constructing a VHL syndrome renal cancer prediction model according to claim 1, characterized in that: The DAMs include N2,N2-Dimethylguanosine and any one or more of the following: PC (16:0 / 16:0), Cysteine-S-sulfate, gamma-Glutamylalanine, 1-deoxy-1-(N6-lysino)-D-fructose, Montecristin, PE (22:4 (7Z, 10Z, 13Z, 16Z) / 14:0), PC (16:1 (9Z) / P-18:1 (11Z)), 1-Kestose, and N2-gamma-Glutamylglutamine.

12. The method for constructing a VHL syndrome renal cancer prediction model according to claim 2, characterized in that: The difference analysis method includes any one or more of the following: PCA, PLS-DA, OPLS-DA.

13. The method for constructing a VHL syndrome renal cancer prediction model according to claim 12, characterized in that: The method of difference analysis is OPLS-DA.

14. The method for constructing a VHL syndrome renal cancer prediction model according to claim 2, characterized in that: The variable analysis was univariate analysis.

15. The method for constructing a VHL syndrome renal cancer prediction model according to claim 2, characterized in that: The first criterion includes: VIP ≥ 1 and P value ≤ 0.

05.

16. The method for constructing a VHL syndrome renal cancer prediction model according to claim 3, characterized in that: The first threshold is 0.

65.

17. The method for constructing a VHL syndrome renal cancer prediction model according to claim 1, characterized in that: The method further comprises: using regression analysis to analyze the relationship between the concentration of the VHL-RCC-related DAMs and the onset age of training set samples to obtain a DAMs-age correlation model.

18. The method for constructing a VHL syndrome renal cancer prediction model according to claim 17, characterized in that: The regression analysis is Cox regression analysis.

19. A method for predicting VHL syndrome renal cancer based on DAMs, characterized in that: The method comprises: Obtain plasma samples from subjects; Detecting the concentration of the target DAM in the plasma sample; the target DAM includes N2,N2-Dimethylguanosine and any one or more of the following: PC (16:0 / 16:0), Cysteine-S-sulfate, gamma-Glutamylalanine, 1-deoxy-1-(N6-lysino)-D-fructose, Montecristin, PE (22:4 (7Z, 10Z, 13Z, 16Z) / 14:0), PC (16:1 (9Z) / P-18:1 (11Z)), 1-Kestose, N2-gamma-Glutamylglutamine; The concentration of the target DAM is compared with a control value to determine the risk of the subject having VHL syndrome renal cancer.

20. The method for predicting VHL syndrome renal cancer based on DAMs according to claim 19, characterized in that: The method further includes: determining the onset time of the subject based on the concentration of the target DAM.

21. The method for predicting VHL syndrome renal cancer based on DAMs according to claim 19, characterized in that: The method further includes advising the subject whether close monitoring is required based on the concentration of the target DAM.

22. The method for predicting VHL syndrome renal cancer based on DAMs according to claim 19, characterized in that: The method also includes: obtaining the age of the subject; and determining the risk of the subject having VHL syndrome renal cancer based on the age.

23. The method for predicting VHL syndrome renal cancer based on DAMs according to claim 22, characterized in that: The method further includes: determining the result of the subject's risk of having VHL syndrome renal cancer based on the age and the concentration of the target DAM.

24. The method for predicting VHL syndrome renal cancer based on DAMs according to claim 19, characterized in that: The method also includes: inputting the target DAMs into the prediction model to calculate the AUC value of the subject, and if the AUC value is greater than or equal to a second threshold, obtaining a result that the subject is at high risk of VHL syndrome renal cancer; if the AUC value is less than the second threshold, obtaining a result that the subject is at low risk of VHL syndrome renal cancer.

25. A method for predicting non-renal organ diseases based on DAMs, characterized in that: The method comprises: Obtain plasma samples from subjects; detecting the concentration of the DAM marker in the plasma sample; The result of judging the risk of the subject suffering from a non-renal organ disease based on the concentration of the DAM marker; if the non-renal organ disease is CHB, judging the high or low risk of the subject suffering from CHB based on the concentration of any one or more of the DAM markers of hypoxanthine, dodecanoylcarnitine and 4-dihydroxy-1-pyrrolidinepropionamide; if the hypoxanthine level concentration increases, and the dodecanoylcarnitine and 4-dihydroxy-1-pyrrolidinepropionamide concentrations decrease, a result is obtained that the subject is at high risk for early CHB; and / or if the non-renal organ disease is PCT, judging the high or low risk of the subject suffering from PCT based on the trehalose concentration; if the plasma concentration of trehalose decreases, a result is obtained that the subject is at high risk for PCT.

26. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 25.

27. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 25 are implemented.

28. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 25 are implemented.