A screening method for tumor metabolic markers
By combining the BIO-FIRE algorithm with the NMF and Boruta algorithms in a metabolomics approach, the problems of sample size and standardization in tumor diagnosis in existing technologies have been solved. This approach enables multi-center studies of metabolic biomarkers and is applicable to various clinical scenarios, including efficient diagnosis of early gastric cancer.
Patent Information
- Application Number
- CN202510372340.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-03-27
AI Technical Summary
Existing metabolomics research in tumor diagnosis suffers from problems such as small sample size, insufficient multi-center validation, lack of unified standards for detection technologies and data analysis, and lack of biological interpretability for individual metabolite screening strategies, resulting in poor research comparability and limitations in biomarker screening.
The BIO-FIRE algorithm was adopted, combined with unsupervised functional module analysis and feature importance screening. Metabolite functional modules were clustered through nonnegative matrix factorization (NMF), importance scores were calculated using the Boruta algorithm, and metabolic biomarkers were selected based on the principle of recursive optimization. Broad-based and targeted metabolomics technologies were integrated to conduct large-scale multicenter studies.
It enhances the biological significance and statistical superiority of tumor metabolic marker screening, ensures the specificity and reproducibility of markers, and is suitable for various clinical applications, including efficient diagnosis of early gastric cancer and gastric cancer that is negative for traditional tumor markers.
Smart Images

Figure CN120221049B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biological medicine, in particular to a screening method of a metabolic marker combination and its application in diagnosing tumors. BACKGROUND
[0002] Metabolomics is generally divided into non-targeted and targeted research strategies. Targeted metabolomics focuses on the quantitative analysis of specific metabolites, has high sensitivity and specificity, and can be used for quantitative verification of known markers and clinical application transformation. While widely targeted metabolomics, as a new generation of targeted metabolomics technology, combines the advantages of targeted metabolomics and non-targeted metabolomics detection technology, and can realize the quantification of hundreds of metabolites (see Non-Patent Literature A Novel Integrated Method for Large-Scale Detection, Identification, and Quantification of Widely Targeted Metabolites: Application in the Study of Rice Metabolomics. [J] Volume 6, Issue 6, November 2013, Pages 1769-1780).
[0003] In recent years, non-invasive cancer diagnosis methods based on plasma metabolomics have attracted widespread attention. Metabolites are the end products of biochemical activities in the body, and their dynamic changes can directly reflect the metabolic reprogramming of cancer cells, host-tumor interactions, and the occurrence and development process of diseases. Compared with genes and proteins, the fluctuations of metabolites are more significant, and plasma metabolomics analysis has the advantages of simple operation, relatively low cost, and good repeatability, so it has become an important research direction for non-invasive cancer diagnosis. However, existing metabolomics research still faces many challenges in tumor diagnosis applications. First, many studies are based on small-scale patient cohorts, lack of multi-center, large-sample verification, and affect the generalization ability of the model. Second, different studies use detection techniques and data analysis processes lack uniform standards, metabolomics data are easily affected by sampling, detection platform and analysis method, etc., leading to poor comparability between different studies. In addition, traditional biomarker screening strategies mainly focus on the statistical significance of single metabolites, ignoring the systematic role of metabolic functional modules, and the selected markers have certain limitations in biological interpretation. SUMMARY
[0004] To solve the problems of the prior art, the present application provides a biomarker identification and optimization algorithm based on functional and importance-based recursive enhancement (BIO-FIRE) by combining unsupervised functional module analysis, feature importance screening and recursive optimization strategy. The algorithm includes three key steps: 1) functional module clustering and explanatory enhancement; 2) potential biomarker identification based on feature importance; and 3) biomarker determination based on recursive optimization principle. On the basis of combining biological interpretability, statistical importance and model simplicity, the algorithm forms an efficient and optimized biomarker screening process, which not only improves the biological significance of the screened biomarkers, but also guarantees their superiority in statistical and model simplicity.
[0005] The present application also establishes a large-scale, multi-center clinical research system, integrates wide-target and targeted metabolomics technology, and develops a new algorithm (BIO-FIRE) based on biological functional modules and recursive optimization strategy, thereby obtaining a robust screening method for tumor metabolic biomarkers.
[0006] The specific scheme is as follows:
[0007] In a first aspect of the present application, a screening method for tumor metabolic biomarkers is provided, which comprises:
[0008] S1) screening and obtaining preliminary screening metabolites from samples of tumor patient groups and non-tumor control groups, and performing qualitative and quantitative analysis;
[0009] S2) performing feature extraction using a biomarker identification and optimization algorithm based on functional and importance-based recursive enhancement (BIO-FIRE) to further determine tumor metabolic biomarkers from the preliminary screening metabolites.
[0010] The BIO-FIRE algorithm comprises:
[0011] A) using a non-negative matrix factorization method (NMF) to perform functional module clustering on the preliminary screening metabolites;
[0012] B) using a Boruta algorithm to perform feature screening by calculating importance scores;
[0013] C) Based on the recursive optimization principle, the importance score of each marker calculated in B) is ranked from high to low, and the metabolic marker subset is selected from the first one in turn until the metabolic marker subset that can maximize the coverage of the most functional modules is selected, that is, the tumor metabolic marker.
[0014] Preferably, S1) comprises constructing a tumor-specific metabolite database and determining differential metabolites by using extensive targeted metabolomics, and further screening and quantifying the preliminary screening metabolites by using targeted metabolomics.
[0015] Preferably, S1) comprises:
[0016] 1) Constructing a tumor-specific metabolite database and determining differential metabolites by using extensive targeted metabolomics; preferably, the sample is detected by using quadrupole time-of-flight mass spectrometry (QTOF) in full scan mode to obtain metabolite ion pair signals, and further converting the metabolite ion pair signals to tandem quadrupole complex linear ion trap mass spectrometry (QTRAP6500) to obtain a tumor-specific metabolite database and determine differential metabolites;
[0017] 2) Screening and quantifying the preliminary screening metabolites by using targeted metabolomics; preferably, the detection range of the targeted metabolomics includes the differential metabolites in step 1), and further preferably includes clinically and literarily recorded tumor metabolic markers; preferably, the literarily recorded tumor metabolic markers include tumor metabolic markers reported more than 5 times in the literature or recognized tumor metabolic markers. Preferably, the differential metabolites are separated by using liquid chromatography, and the preliminary screening metabolites are quantified by using mass spectrometry.
[0018] Further preferably, S1) comprises: detecting the sample by using quadrupole time-of-flight mass spectrometry (QTOF) in full scan mode to obtain metabolite ion pair signals, and further converting the metabolite ion pair signals to tandem quadrupole complex linear ion trap mass spectrometry (QTRAP6500) to obtain a tumor-specific metabolite database and determine differential metabolites; and then separating the differential metabolites by using liquid chromatography and quantifying the preliminary screening metabolites by using mass spectrometry.
[0019] Preferably, the preliminary screening metabolites include amino acids and their derivatives, bile acids and their conjugates, purine / pyrimidine metabolites, steroid hormones, lipids (fatty acids, phospholipids, sphingolipids), biogenic amines, organic acids, sugars, flavonoids / alkaloids and other substances, which are related to amino acid metabolism, bile acid metabolism, nucleotide metabolism, lipid metabolism, hormone metabolism, drug metabolism and microbial metabolism, etc.
[0020] In one specific embodiment of the present application, the liquid chromatography and mass spectrometry conditions used in the extensive targeted metabolomics are as follows:
[0021] Chromatographic column: Waters ACQUITY UPLC HSS T3 C18, column temperature 40 °C;
[0022] Mass spectrometry conditions: electrospray ionization (ESI) temperature 500 °C, mass spectrometry voltage 5500 V (positive) or -4500 V (negative), ion source gas I (GS I) 55 psi, gas II (GS II) 60 psi, curtain gas (CUR) 25 psi, collision-activated dissociation (CAD) parameters set to high. In a triple quadrupole (Qtrap), each ion pair is scanned in MRM mode according to optimized declustering potential (DP) and collision energy (CE).
[0023] Mass spectrometry conditions: electrospray ionization (ESI) temperature 500 °C, mass spectrometry voltage 5500 V (positive) or -4500 V (negative), ion source gas I (GS I) 55 psi, gas II (GS II) 60 psi, curtain gas (CUR) 25 psi, collision-activated dissociation (CAD) parameters set to high. In a triple quadrupole (Qtrap), each ion pair is scanned in MRM mode according to optimized declustering potential (DP) and collision energy (CE).
[0024] In one embodiment of the present application, the liquid chromatography and mass spectrometry conditions used in the targeted metabolomics are:
[0025] Chromatographic column: Waters ACQUITY UPLC HSS T3 C18, column temperature 40 °C;
[0026] Mass spectrometry conditions: electrospray ionization (ESI) temperature 500 °C, mass spectrometry voltage 5500 V (positive) or -4500 V (negative), ion source gas I (GS I) 55 psi, gas II (GS II) 60 psi, curtain gas (CUR) 25 psi, collision-activated dissociation (CAD) parameters set to high. In a triple quadrupole (Qtrap), each ion pair is scanned in MRM mode according to optimized declustering potential (DP) and collision energy (CE).
[0027] Preferably, the screening method further comprises linear discriminant analysis (LDA) grouping of ion pair signals in the constructed tumor-specific metabolite database before targeted metabolomics detection, and screening of differential metabolic signals using statistical model and P value and FC value of univariate statistical T detection. Preferably, FDR < 0.05, FC > 1.2 or FC < 0.80, and VIP > 1 are set as the standard of significant difference.
[0028] Preferably, the grouping is into a tumor patient group and a non-tumor control group.
[0029] Further preferably, the tumor group comprises different subtypes or different development stages of the same tumor.
[0030] Preferably, the BIO-FIRE algorithm comprises:
[0031] A) using NMF to perform functional module clustering on the preliminary screening metabolites, and performing functional naming on the obtained modules according to KEGG annotation information;
[0032] B) using Boruta algorithm to calculate importance scores to perform feature screening on the functional module data extracted in A), to obtain potential metabolic markers;
[0033] C) based on the principle of recursive optimization, ranking the importance scores of each marker calculated in B) from high to low, and selecting metabolic marker subsets from the first one in turn until a metabolite subset is selected that can maximize the coverage of all functional modules covered by the potential metabolic markers in B), i.e., tumor metabolic markers.
[0034] Preferably, the NMF introduces a KL divergence optimization strategy to strengthen the sparse correlation between metabolites and functional modules. Preferably, the objective function is:
[0035]
[0036] wherein X represents the original metabolite data matrix, with a dimension of n x m (n: number of metabolites, m: number of samples); W represents the weight matrix, with a dimension of n x r (r: number of functional modules), representing the contribution of metabolites to functional modules; H represents the activation matrix, with a dimension of r x m, representing the distribution of functional modules in samples; represents the square of the Frobenius norm, used to measure the reconstruction error of the matrix λ1, λ2 represent L1 regularization parameters, controlling the sparsity of W and H; λ3 represents the KL divergence regularization parameter, adjusting the strength of the sparsity constraint; represents a sparsity prior matrix (such as generated by thresholding or low-rank template), guiding the sparse distribution of the weight matrix W; denotes the Kullback-Leibler divergence, measuring the difference between the weight matrix W and the prior sparse matrix ; i, j denote the row and column indices of the matrix (i: metabolite index, j: functional module or sample index).
[0037] The NMF comprises: determining the optimal rank by multi-rank synchronous calculation combined with residual sum of squares inflection point analysis; preferably, multi-rank synchronous calculation (the rank range can be set to 3 to 8) combined with residual sum of squares inflection point analysis to finally determine the optimal decomposition rank (preferably, rank = 5).
[0038] Preferably, for candidate rank r ∈ {3, 4, …, 8}, the NMF decomposition is calculated in parallel:
[0039]
[0040] wherein r represents the candidate decomposition rank (functional module number), and the value range is 3 to 8; W r, H r denotes the weight matrix and the activation matrix when the rank is r; X is the original metabolite data matrix (the row represents the metabolite, and the column represents the sample); denotes the Frobenius norm, which is used to measure the error between X and WH; RSS represents the residual sum of squares (Residual Sum of Squares), that is used to evaluate the fitting effect of different ranks; the inflection point is the point where the slope changes the most in the RSS curve, which is used to automatically select the optimal rank (such as r = 5).
[0041] Preferably, the features that explain the top 20% of the functional modules are screened (preferably, the maximum contribution principle is adopted).
[0042] Preferably, the functional module overlap rate is controlled to be below 15%. Further preferably, the Jaccard similarity is used to control the functional module overlap rate to be below 15%.
[0043] Further preferably, the NMF comprises screening the features that explain the top 20% of the functional modules by using the maximum contribution principle, and / or realizing the biological annotation of the functional modules through the temporal trajectory analysis and the pathway enrichment analysis, and controlling the functional module overlap rate to be below 15% by using the Jaccard similarity.
[0044] Preferably, the function of the maximum contribution principle for screening the features is:
[0045] Selected j = {i | W i,j ≥ Quantile(W :,j , 0.8)}
[0046] where Selected j denotes the set of metabolites retained in functional module j (top 20%); W i,j denotes the weight value of metabolite i in functional module j; W:,j denotes the jth column of the weight matrix (all metabolite weights of functional module j); Quantile(W:,j,0.8) denotes the 80th quantile of weights in functional module j, only metabolites higher than this value are retained.
[0047] Preferably, the function for controlling the functional module overlap rate is controlled to be below 15%:
[0048]
[0049] where W j ,W k denotes the weight vector of functional module j and functional module k; Supp(W j ) denotes the set of metabolites with non-zero weights in functional module j (i.e. {i | W i,j > 0}); Supp(W k ) denotes the set of metabolites with non-zero weights in functional module k; ∩, ∪ denotes the intersection and union operators of sets; 0.15 denotes the maximum allowed overlap rate between functional modules (Jaccard similarity threshold).
[0050] In the Boruta algorithm, the split criterion of feature Xj at decision tree node t is set as:
[0051]
[0052] where j denotes the feature index (such as the number of genes or modules); t denotes the current node in the decision tree; denotes the Gini impurity of node t; C denotes the total number of classes of the target variable (in a classification task); p(c | t) denotes the proportion of samples in node t belonging to class c; N denotes the total number of samples in node t; N L ,N R denotes the number of samples in the left and right child nodes after splitting; t L , t R denotes the left and right child nodes; ΔGini(j, t): the Gini impurity reduction amount brought by feature j at node t.
[0053] The tumor is a hematological tumor or a solid tumor, preferably gastric cancer.
[0054] In one embodiment of the present application, the tumor is gastric cancer, and the functional modules include glutamine metabolism functional module, lipid metabolism functional module, host-microorganism co-metabolism functional module, metabolic network bridge functional module, and branched-chain amino acid metabolism functional module.
[0055] The sample is a body fluid. Preferably, the body fluid is selected from blood, plasma or serum.
[0056] In one embodiment of the present application, the sample is plasma.
[0057] The number of tumor metabolic markers is 10-15, preferably 12.
[0058] In one embodiment of the present application, the tumor metabolic markers include nicotinamide, taurine, phthalic acid monoethylhexyl ester, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (±) 12-hydroxy-5Z,8Z,10E,14Z-eicosatetraenoic acid ((±) 12-HETE), 5'-deoxy-5'-methylthioadenosine, chondritic acid, hypoxanthine and azelaic acid.
[0059] In a second aspect of the present application, a tumor metabolic marker obtained by the screening method is provided.
[0060] In a third aspect of the present application, a computer readable storage medium or a device comprising a computer readable storage medium is provided, wherein the computer readable storage medium stores a computer program, and the computer program comprises a BIO-FIRE algorithm, and the BIO-FIRE algorithm comprises:
[0061] A) using non-negative matrix factorization method (NMF) to perform functional module clustering on the preliminary screening metabolites;
[0062] B) using Boruta algorithm to perform feature screening by calculating importance scores;
[0063] C) based on the principle of recursive optimization, ranking the importance scores of each marker calculated in B) from high to low, and selecting the metabolic marker subset from the first one in turn until the metabolic marker subset that can maximize the coverage of the most functional modules is selected, i.e., the tumor metabolic markers.
[0064] The tumor is a hematological tumor or a solid tumor; preferably, the tumor is gastric cancer.
[0065] Preferably, the functional modules include glutamine metabolism functional module, lipid metabolism functional module, host-microorganism co-metabolism functional module, metabolic network bridge functional module, and branched-chain amino acid metabolism functional module.
[0066] In a fourth aspect, the present application provides a use of the computer readable storage medium or the device comprising the computer readable storage medium in screening tumor markers.
[0067] In a fifth aspect, the present application provides a method for constructing a tumor diagnosis model, which comprises:
[0068] i) collecting the concentration detection results of the metabolic markers of the tumor patient group and the non-tumor control group;
[0069] ii) constructing a tumor diagnosis model according to the information collected in step i);
[0070] Alternatively, the method for constructing a tumor diagnosis model comprises:
[0071] I) collecting the sample of the subject and detecting the concentration of the metabolic markers;
[0072] II) conducting clinical diagnosis on the subject and dividing the subject into the tumor patient group and the non-tumor control group;
[0073] III) constructing a tumor diagnosis model according to the detection results of step I) and the diagnosis results of step II).
[0074] The metabolic markers are the metabolic markers screened by the screening method of the first aspect. Preferably, the metabolic markers comprise nicotinamide, taurine, phthalic acid monoethylhexyl ester, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycy-tyl-phenylalanine, N-phenylacetyl-L-glutamine, (±) 12-hydroxy-5Z,8Z,10E,14Z-eicosatetraenoic acid ((±) 12-HETE), 5'-deoxy-5'-methylthioadenosine, cantharidic acid, hypoxanthine and azelaic acid.
[0075] Preferably, the tumor is a hematological tumor or a solid tumor. Further preferably, the tumor is gastric cancer. The method for clinical diagnosis comprises, but is not limited to, electronic gastroscopy.
[0076] The sample is a body fluid; preferably, the body fluid is selected from blood, plasma or serum.
[0077] The method for detecting the concentration of the metabolic markers in the sample comprises one or more than two of nuclear magnetic resonance spectroscopy, mass spectrometry, chromatography, high performance liquid chromatography-tandem mass spectrometry, Fourier transform ion cyclotron resonance, ion mobility spectrometry, electrochemical detection, Raman spectroscopy or radio-labeling.
[0078] The algorithm used in the construction of the diagnostic model is a machine learning model, preferably including one or more of logistic regression (LR), k-nearest neighbor algorithm (KNN), naive Bayes (NB), support vector machine (SVM Radial), random forest (RF), XGBoost, TreeNet, gradient boosting machine (GBM), LASSO, or neural network (NNET).
[0079] In a sixth aspect of the present application, a diagnostic model obtained by the construction method of the fifth aspect is provided.
[0080] In a seventh aspect of the present application, a use of metabolic markers in the preparation of a product for diagnosing gastric cancer is provided, wherein the metabolic markers include nicotinamide, taurine, monoethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycy-tyl-phenylalanine, N-phenylacetyl-L-glutamine, (±) 12-hydroxy-5Z,8Z,10E,14Z-eicosatetraenoic acid ((±) 12-HETE), 5'-deoxy-5'-methylthioadenosine, chondrillin, hypoxanthine, and azelaic acid.
[0081] The tumor is a hematological tumor or a solid tumor. Preferably, the tumor is gastric cancer, such as early gastric cancer, gastric cancer negative for traditional tumor markers, gastric cancer with a tumor length less than 4 cm, gastroesophageal junction cancer, etc.
[0082] The metabolic marker is a metabolic marker in a body fluid. Preferably, the body fluid is selected from blood, plasma, or serum; more preferably, the metabolic marker is in plasma.
[0083] The product includes a reagent for detecting the metabolic marker, preferably, the reagent detects the concentration of the metabolic marker.
[0084] Preferably, the method for detecting the metabolic marker includes one or more of nuclear magnetic resonance spectroscopy, mass spectrometry, chromatography, high-performance liquid chromatography-tandem mass spectrometry, Fourier transform ion cyclotron resonance, ion mobility spectrometry, electrochemical detection, Raman spectroscopy, or radioactive labeling.
[0085] In a specific embodiment of the present application, the method for detecting the metabolic marker is high-performance liquid chromatography-tandem mass spectrometry.
[0086] Preferably, the product includes a diagnostic model, a kit, a test paper, a chip, or a device.
[0087] Preferably, the product further includes a reagent for sample processing and / or a standard of the metabolic marker.
[0088] In an eighth aspect of the present application, there is provided a metabolic marker for diagnosing gastric cancer, the metabolic marker comprising nicotinamide, taurine, phthalic acid monoethylhexyl ester, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (±) 12-hydroxy-5Z,8Z,10E,14Z-eicosatetraenoic acid ((±) 12-HETE), 5'-deoxy-5'-methylthioadenosine, wounding acid, hypoxanthine and azelate.
[0089] The metabolic marker is a metabolic marker in a body fluid. Preferably, the body fluid is selected from blood, plasma or serum; more preferably, the metabolic marker is in plasma.
[0090] In a ninth aspect of the present application, there is provided a method for diagnosing a tumor, the method comprising detecting the concentration of a metabolic marker in a sample of a subject.
[0091] The metabolic marker comprises nicotinamide, taurine, phthalic acid monoethylhexyl ester, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (±) 12-hydroxy-5Z,8Z,10E,14Z-eicosatetraenoic acid ((±) 12-HETE), 5'-deoxy-5'-methylthioadenosine, wounding acid, hypoxanthine and azelate.
[0092] Preferably, the sample is a body fluid. Further preferably, the body fluid is selected from blood, plasma or serum.
[0093] The tumor is a hematological tumor or a solid tumor. Preferably, the tumor is gastric cancer.
[0094] The method comprises detecting the concentration of the metabolic marker by one or more of nuclear magnetic resonance spectroscopy, mass spectrometry, chromatography, high performance liquid chromatography-tandem mass spectrometry, Fourier transform ion cyclotron resonance, ion mobility spectrometry, electrochemical detection, Raman spectroscopy or radiolabeling.
[0095] Preferably, the method comprises inputting the detected concentration of the metabolic marker into a constructed machine learning model, outputting a probability value, comparing the output probability value with a threshold value, and determining whether the subject has the tumor.
[0096] The threshold value is obtained from prior experiments, i.e., by determining the difference in the metabolic marker between a gastric cancer patient group and a non-gastric cancer control group, training a machine learning model, and determining the threshold value.
[0097] In a tenth aspect of the present application, there is provided a diagnostic system for a tumor, the diagnostic system comprising a device, unit or module for determining whether a subject has the tumor based on the concentration of a metabolic marker.
[0098] Preferably, the diagnostic system comprises:
[0099] a data detecting device, unit or module for detecting the concentration of the metabolic marker in the sample;
[0100] a data input device, unit or module for inputting the concentration data of the metabolic marker;
[0101] a data analyzing device, unit or module for analyzing the tumor risk based on the concentration data of the metabolic marker;
[0102] a data output device, unit or module for outputting the analysis result of whether the individual has a tumor.
[0103] The analysis employs a machine learning model, which can be one or more than two of logistic regression (LR), k-nearest neighbor algorithm (KNN), naive Bayes (NB), support vector machine (SVM Radial), random forest (RF), XGBoost, TreeNet, gradient boosting machine (GBM), LASSO or neural network (NNET).
[0104] The metabolic marker comprises nicotinamide, taurine, phthalic acid monoethylhexyl ester, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycy-l-phenylalanine, N-phenylacetyl-L-glutamine, (±) 12-hydroxy-5Z,8Z,10E,14Z-eicosatetraenoic acid ((±) 12-HETE), 5'-deoxy-5'-methylthioadenosine, wounding acid, hypoxanthine and azelaic acid.
[0105] In an eleventh aspect, the present application provides a device comprising the diagnostic system of the tenth aspect.
[0106] The term "subject" can be a human or non-human animal, including "patient", "suspected patient" and "healthy individual", etc., which does not indicate a specific age or gender, covering adult or neonatal subjects and fetuses. The non-human animal can be a wild animal, a zoo animal, an economic animal, a pet, a laboratory animal, etc. Specifically, the non-human animal includes but is not limited to pig, cow, sheep, horse, donkey, fox, raccoon dog, mink, camel, dog, cat, rabbit, mouse (such as rat, mouse, guinea pig, hamster, gerbil, chinchilla, squirrel) or monkey, etc.
[0107] The term "diagnosis" refers to finding out whether a patient has a disease or a condition in the past, at the time of diagnosis or in the future, or finding out the progress or possible progress of the disease in the future.
[0108] The present application has the following beneficial effects:
[0109] 1. Large-scale, multi-center data support: This application integrates large-scale clinical samples from multiple authoritative medical centers to ensure the robustness and clinical applicability of the research results;
[0110] 2. Broad-target and targeted metabolomics joint analysis: Broad-target metabolomics is used to screen potential candidates, and targeted metabolomics is used for quantitative verification to ensure the specificity and repeatability of metabolites;
[0111] 3. Systematic metabolic marker screening strategy: This application develops the BIO-FIRE algorithm, uses non-negative matrix factorization (NMF) to mine functional modules of metabolites, and combines the Boruta feature selection algorithm and recursive optimization principle to determine the least marker combination covering the most functional modules, improving the comprehensiveness and efficiency of marker screening. In particular, the inventors of this application used targeted metabolomics combined with RFE-RF machine learning algorithm for feature extraction in previous studies (see patent document 2025103171683), which did not use a screening strategy covering functional modules, although the AUC value of the screened marker combination was high, but the number was large, and the cost increased.
[0112] 4. Suitable for various clinical application scenarios: The screening method of this application is not only suitable for routine gastric cancer screening, but also can provide an efficient diagnosis scheme in complex clinical scenarios such as early gastric cancer, traditional tumor marker negative gastric cancer, and gastroesophageal junction cancer.
[0113] 5. This application breaks through the limitations of existing metabolomics research and constructs an efficient, stable, and generalizable metabolomics marker screening and diagnosis process for non-invasive screening and precise diagnosis of gastric cancer, and provides an extensible technical platform for the field of metabolic research. BRIEF DESCRIPTION OF DRAWINGS
[0114] In the following, embodiments of the present application will be described in detail with reference to the accompanying drawings, in which:
[0115] Figure 1 The broad-target metabolomics detection results are shown, wherein, Figure 1 A is a heat map covering 3327 metabolic signals, wherein HC represents the healthy group, BGD represents the benign disease group, EGC represents the early gastric cancer group, and AGC represents the advanced gastric cancer group, Figure 1 B is a linear discriminant analysis (LDA) plot, wherein HC represents the healthy group, BGD represents the benign disease group, EGC represents the early gastric cancer group, and AGC represents the advanced gastric cancer group, Figure 1 C shows the differential metabolic signal screening process taking the comparison group of gastric cancer (GC) vs non-gastric cancer (NGC) as an example, wherein VIP represents the projection variable importance, FC represents the fold change, and FDR represents the false discovery rate, Figure 1D is the differential metabolic signals of 10 comparison groups and their union set, in which, HC represents the healthy group, GC represents the gastric cancer group, BGD represents the benign disease group, EGC represents the early gastric cancer group, and AGC represents the advanced gastric cancer group;
[0116] Figure 2 The screening and identification process of differential metabolic ion pairs is shown;
[0117] Figure 3 Various evaluation indexes of NMF under different decomposition ranks are shown;
[0118] Figure 4 The heat map of different clustering methods is shown, in which, Figure 4 A is the hclust clustering heat map, Figure 4 B is the Kmean clustering heat map, Figure 4 C is the NMF clustering heat map;
[0119] Figure 5 Part of the results during the running of the Boruta algorithm is shown, in which, Figure 5 A reflects the rigor of the Boruta algorithm in feature selection, Figure 5 B reflects the comprehensive coverage of important metabolites selected by the Boruta algorithm on functional modules;
[0120] Figure 6 The detailed process of determining markers using the recursive optimization principle is shown;
[0121] Figure 7 The importance score ranking of 12 modeling metabolic markers (A) and the expression difference between the gastric cancer (GC) group and the non-gastric cancer (NGC) group (B) are shown;
[0122] Figure 8 The column chart of the functional module coverage rate of metabolic markers selected by the Bortua algorithm, the LASSO algorithm and the RF algorithm and the proportion of new spectral substances is shown;
[0123] Figure 9 The performance of the sum of P values and the sum of absolute values of logFC of metabolic markers selected by the Bortua algorithm, the LASSO algorithm and the RF algorithm after FDR correction is shown;
[0124] Figure 10 The modeling robustness of metabolic markers selected by the Bortua algorithm, the LASSO algorithm and the RF algorithm is shown;
[0125] Figure 11The ROC curve of the gastric cancer metabolic marker diagnostic model constructed by 10 machine learning algorithms in the test set (A) and the corresponding AUC value are shown; the radar chart shows the performance comparison of the metabolic marker diagnostic model constructed by 10 machine learning algorithms and 3 traditional tumor marker diagnostic models in the test set (B), and the evaluation indexes include accuracy, sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV); the ROC curve (C) and AUC value, accuracy, sensitivity and specificity column chart (D) of the random forest (RF) model of PMB-P12, CEA, CA199 and CA724;
[0126] Figure 12 The diagnostic performance of the metabolite combination covering 4 functional modules and covering a single functional module is shown;
[0127] Figure 13 The discrimination ability of 10 diagnostic models of 12 metabolic markers in complex clinical scenarios is shown, wherein Figure A is gastric cancer with negative traditional tumor markers, Figure B is gastric cancer with a diameter less than 4 cm, Figure C is very early stage (stage IA) gastric cancer, and Figure D is gastroesophageal junction cancer. DETAILED DESCRIPTION
[0128] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0129] It should be noted that the methods used in the present application are conventional methods unless otherwise specified, and the reagents used in the present application are commercially available products unless otherwise specified.
[0130] Example 1: Construction of gastric cancer metabolic atlas and identification of differential metabolites based on extensive targeted metabolomics technology
[0131] 1 Collection of plasma samples in the discovery cohort
[0132] The present embodiment collects peripheral venous blood plasma from 267 non-gastric cancer subjects (including 100 healthy people and 167 patients with benign gastric disease) and 202 gastric cancer patients (including 134 early gastric cancer and 68 advanced gastric cancer) from two research centers (Beijing Friendship Hospital, Hubei Provincial People's Hospital) after obtaining the consent of the patients as a discovery cohort. Among them, the diagnostic criteria for gastric cancer patients are confirmed by postoperative pathology; The samples in the non-gastric cancer group include healthy people without gastric diseases after physical examination and patients with benign gastric diseases including chronic gastritis, erosive gastritis, gastric ulcer, gastric polyp, etc. after hospital examination. All gastric cancer patients and non-gastric cancer group subjects have no history of other malignant tumors, no other major systemic diseases, and no chronic disease history of long-term medication. The samples of healthy people (HC) and benign disease (BGD) patients are included as non-gastric cancer group (NGC), and the samples of early (EGC) and advanced gastric cancer (AGC) patients are included as gastric cancer group (GC). All blood samples are collected in the morning on an empty stomach, then centrifuged to separate the plasma, and immediately stored in a -80°C refrigerator. During the study, the plasma samples were taken out according to the experimental requirements, thawed and then analyzed.
[0133] 2. Construction of gastric cancer-specific metabolite ion pair database
[0134] Database construction process: In order to study the relationship between plasma metabolome and gastric cancer, we used 20% of the plasma samples in the discovery cohort (balanced sample size in each group) to collect gastric cancer plasma metabolite ion characteristics widely using different mass spectrometry detection techniques (QTOF and QTRAP) and metabolite databases, and constructed a gastric cancer plasma-specific metabolome database. First, QTOF high-resolution instrument was used to detect gastric cancer plasma samples in full scan mode to ensure as many metabolite ion characteristics as possible were detected. However, due to the characteristics of QTOF instrument, its detection sensitivity for low-concentration metabolites is low, making it difficult to capture low-content substance signals. Therefore, the detected ion pair signals were further transferred to QTRAP6500 instrument, and AB SCIEX QTRAP6500 MIM-IDA-EPI scan mode was used to supplement the detection of low-concentration metabolites. Combined with all the above metabolite ion characteristics, local standard database metabolite information, and literature metabolite information, after optimization and deduplication, a gastric cancer plasma-specific metabolome database was constructed.
[0135] 3. Plasma extensive targeted metabolomics detection
[0136] 1) Sample pretreatment
[0137] The sample collected in step 1 was taken out from the -80℃ refrigerator, thawed on ice until there was no ice in the sample (all subsequent operations were required to be performed on ice); after the sample was thawed, it was vortexed for 10 s and mixed, 50 μL of the sample was taken and added to the corresponding numbered centrifuge tube; 300 μL of pure methanol internal standard extraction solution (containing 100 ppm concentration of L-phenylalanine internal standard) was added; vortexed for 5 min, stood for 24 h, and then centrifuged at 12000 r / min, 4℃ for 10 min; 270 μL of supernatant was concentrated for 24 h; then 100 μL of resuspension solution composed of acetonitrile and water in a volume ratio of 1:1 was added for LC-MS / MS analysis. 20 μL of each sample was mixed to form a quality control sample (QC), and one was collected every 15 samples.
[0138] 2) Sample metabolite detection and analysis
[0139] The liquid chromatography conditions were determined as follows:
[0140] Chromatographic column: Waters ACQUITY UPLC HSS T3 C18 1.8 μm, 2.1 mm*100 mm; column temperature was 40℃; sample size was 2 μL.
[0141] Mobile phase: A phase was 0.1% acetic acid aqueous solution, B phase was 0.1% acetic acid acetonitrile solution. The elution gradient program was: 0 min, the volume ratio of A phase to B phase was 95:5; 11.0 min, the volume ratio of A phase to B phase was 10:90; 12.0 min, the volume ratio of A phase to B phase was 10:90; 12.1 min, the volume ratio of A phase to B phase was 95:5; 14.0 min, the volume ratio of A phase to B phase was 95:5 V / V. The flow rate was 0.4 mL / min.
[0142] The mass spectrometry conditions were determined as follows:
[0143] Electrospray ion source (ESI) temperature 500℃, mass spectrometry voltage 5500V (positive) or -4500V (negative), ion source gas I (GS I) 55 psi, gas II (GS II) 60 psi, curtain gas (CUR) 25 psi, collision-activated dissociation (CAD) parameter setting was high.
[0144] In the triple quadrupole (Qtrap), each ion pair was scanned and detected in MRM mode according to the optimized declustering potential (DP) and collision energy (CE).
[0145] The samples were analyzed and detected according to the determined liquid chromatography conditions and mass spectrometry conditions: for all samples in the above discovery cohort, a metabolomics method combining enhanced ion scan mass spectrometry (MIM-EPI) and time-of-flight mass spectrometry (TOF) with multiple reflection monitoring acquisition mode was used, and a local standard database was integrated to construct a gastric cancer plasma metabolite database, to obtain the original mass spectrometry data of each plasma sample.
[0146] 3) Chromatogram peak area preprocessing and integration
[0147] Based on the constructed gastric cancer serum specific metabolite database, the metabolites of the samples were subjected to mass spectrometry qualitative and quantitative analysis. Different molecular weight metabolites can be separated by liquid chromatography. The characteristic ions of each substance are screened by the multiple reaction monitoring mode (MRM) of the triple quadrupole, and the signal intensity (CPS) of the characteristic ions is obtained in the detector. The MultiQuant3.0.3 software is used to open the sample off-line mass spectrometry file, and the chromatographic peak integration and correction work is carried out. The peak area (Area) of each chromatographic peak represents the relative content of the corresponding substance, and the peaks with S / N>5 and retention time deviation not more than 0.2 min are reserved. Finally, all chromatographic peak area integration data are saved.
[0148] 4) Experimental quality control
[0149] By overlapping and analyzing the total ion flow chart of different quality control QC sample mass spectrometry detection and analysis, the repeatability of metabolite extraction and detection, i.e. technical repeatability, can be judged. The high stability of the instrument provides an important guarantee for the repeatability and reliability of the data. The CV value, i.e. the coefficient of variation, is the ratio of the standard deviation of the original data to the average of the original data, which can reflect the degree of data dispersion. Using the empirical cumulative distribution function (Empirical Cumulative Distribution Function, ECDF) can analyze the frequency of substances with CV less than the reference value. The higher the proportion of substances with low CV value in the QC sample, the more stable the experimental data is. The proportion of substances with CV value less than 0.5 in the QC sample is higher than 85%, indicating that the experimental data is stable. The proportion of substances with CV value less than 0.3 in the QC sample is higher than 75%, indicating that the experimental data is very stable. At the same time, the change of L-phenylalanine internal standard CV value in the detection process is monitored. The change of internal standard CV value is less than 20%, indicating that the instrument stability in the detection process is good.
[0150] 4) Construction of plasma metabolite profile of gastric cancer patients and identification of differential metabolites
[0151] (1) Data preprocessing
[0152] Step 1 above detected a total of 3327 ion pairs. The peak area integration data after normalization was pre-processed, and according to the 80% rule, a metabolite was considered to be detectable when it was detected in at least 4 / 5 samples in a group. According to this rule, 83 metabolites were robustly detected in all 928 samples. For a small number of undetected metabolites (<1 / 5 samples), the minimum value detected in the group was filled in half to facilitate subsequent statistical analysis.
[0153] (2) Analysis of wide-target metabolomics detection results
[0154] The heat map analysis results covering 3327 metabolic signals showed that the healthy control group (HC), benign gastric disease group (BGD), early gastric cancer group (EGC) and advanced gastric cancer group (AGC) showed significant differential expression on different metabolic signals. Figure 1 A) In order to effectively distinguish the above four groups, the present embodiment further uses supervised learning linear discriminant analysis (LDA) to analyze the data. As shown in Figure 1 B, the overall difference between groups can be distinguished by linear combination of all compound characteristics. Each point in the figure represents a sample, and the closer the distance between sample points, the closer the compound content. The results show that the sample points of the HC, BGD, EGC and AGC groups show a clear clustering trend, and the sample points between the HC group and the other three groups show good separation.
[0155] In order to more comprehensively mine the differences between the four groups to determine the most differential metabolic signals, the four groups were further compared pairwise and added, a total of 10 comparison groups were set, and the peak area integration data was used to analyze the differential metabolites between the two groups. At the same time, the P value and FC value of the T test of the multivariate statistical model and the VIP value of the single variable statistics were used to screen the differential metabolic signals, and the FDR<0.05, FC>1.2 or FC<0.80, VIP>1 were set as the significant difference standard Figure 1 C), the differential metabolic signals of the 10 comparison groups were screened and their union was taken Figure 1 D), and finally 1097 differential metabolic signals were obtained for subsequent ion pair spectral analysis and substance identification.
[0156] (3) Spectral analysis and substance identification of differential metabolic signals
[0157] To improve the confidence of targeted detection of metabolites, 792 signals with high-resolution data were selected from the 1097 differential metabolic signals screened in the early stage for accurate spectral resolution. The spectral resolution process mainly includes two parts: ①Molecular formula fitting and compound matching: First, unknown ions are fitted with molecular formula according to their accurate mass-to-charge ratio by peakview2.2 (AB Sciex) software, with a mass-to-charge ratio range of less than 10ppm. After comprehensive evaluation, the appropriate molecular formula is selected, and the Chemspider database (www.Chemspider.com) and Pubchem database are used to obtain multiple possible compounds corresponding to the molecular formula. Combined with biological and biochemical information, the most likely candidate compound is selected, and the secondary matching of the unknown substance is fitted with the fomular pane fragment by peakview, with a matching rate of >70%. ②Standard verification: To confirm the identity of the compound, standard verification is purchased. The unknown metabolites are confirmed according to the following four standards: column retention time, isotope pattern, MS fragmentation (collision-induced dissociation MS / MS), and accurate mass. The detailed results of this process are shown in Table 1. Figure 2 After de-duplication, 84 differential metabolites identified by standard (MSI Level 1) were finally obtained. The specific information of the 84 differential metabolites is shown in Table 1.
[0158] Table 1
[0159]
[0160]
[0161]
[0162]
[0163] Example 2: Determination of preliminary screening metabolites based on targeted metabolomics technology
[0164] 1 Collection of plasma samples of modeling cohort
[0165] In this embodiment, 460 non-gastric cancer subjects (including 216 healthy people and 244 patients with benign gastric diseases) and 468 gastric cancer patients (including 247 early gastric cancer patients and 221 advanced gastric cancer patients) from four research centers (Beijing Friendship Hospital, Cancer Hospital of Chinese Academy of Medical Sciences, Hubei Provincial People's Hospital, Tongji Hospital of Tongji Medical College of Huazhong University of Science and Technology) were collected as the modeling cohort. The patient inclusion and exclusion criteria are the same as in Example 1. The blood collection time was in the morning on an empty stomach. After centrifugation, all plasma samples were stored in a -80℃ refrigerator. The plasma samples were taken out for subsequent analysis after thawing.
[0166] 2 Targeted metabolomics detection of plasma samples
[0167] (1) Determine the detection range of target
[0168] Based on the 84 differential metabolites identified in the previous screening, after systematic retrieval of gastric cancer metabolomics research and bibliometrics analysis, 32 literature metabolites with a reported frequency greater than 5 or recognized tumor metabolic markers were summarized, and after deduplication, 102 metabolites to be detected were finally obtained. These metabolites mainly cover aminos acids and their derivatives, bile acids and their conjugates, purine / pyrimidine metabolites, steroid hormones, lipids (fatty acids, phospholipids, sphingolipids), biogenic amines, organic acids, sugars, flavonoids / alkaloids, etc. Multiple substances are involved in amino acid metabolism, bile acid metabolism, nucleotide metabolism, lipid metabolism, hormone metabolism, drug metabolism and microbial metabolism, etc. physiological and pathological processes, reflecting the metabolic state of the body.
[0169] (2) Sample pretreatment
[0170] Take the sample collected in step 1 from the -80°C refrigerator and thaw on ice until completely free of ice crystals; after thawing the sample, vortex for 10 seconds to mix thoroughly. Take 50 μL of sample, add 150 μL of extraction solution (containing 100 ppm concentration of isotope internal standard), then vortex for 3 minutes, centrifuge at 12000 rpm, 4°C for 10 minutes, and store in a -20°C refrigerator overnight. The next day, centrifuge at 12000 rpm, 4°C for 5 minutes, and transfer 170 μL of supernatant to a 96-well plate. After completing the protein precipitation treatment, seal the plate for LC-MS / MS analysis. Take another 20 μL of each sample to prepare a quality control sample (QC), and collect one every 15 samples.
[0171] (3) Determine the detection conditions for detection
[0172] Targeted quantitative analysis uses T3 column for metabolite separation, and the specific detection conditions are as follows:
[0173] T3 column liquid chromatography conditions: The chromatographic column is Waters ACQUITY UPLC HSS T3 C18 1.8 μm, 2.1 mm x 100 mm, the column temperature is set to 40°C, and the injection volume is 2 μL. The mobile phase A is 0.1% acetic acid in water, and the mobile phase B is 0.1% acetic acid in acetonitrile; the elution gradient is: 0 min, A:B = 95:5; 11.0 min, A:B = 10:90; 12.0 min, A:B = 10:90; 12.1 min, A:B = 95:5; 14.0 min, A:B = 95:5, the flow rate is set to 0.4 mL / min.
[0174] Mass spectrometry conditions: The temperature of the electrospray ion source (ESI) was set to 500 °C, and the mass spectrometry voltage was 5500 V (positive ion mode) and -4500 V (negative ion mode), respectively. The ion source gas I (GS I) pressure was 55 psi, the gas II (GS II) pressure was 60 psi, and the curtain gas (CUR) pressure was 25 psi. The collision-induced ionization (CAD) was set to high. In the triple quadrupole (Qtrap), the optimized declustering voltage (DP) and collision energy (CE) were used for MRM mode scanning to detect the signal of each ion pair.
[0175] (4) Spectrum peak area preprocessing and integration
[0176] The mass spectrometry data was processed by MultiQuant 3.0.3 software, and the mass spectrometry peaks of the target substances in each sample were integrated and corrected according to the retention time and peak shape information of the standard, so as to ensure the accuracy of qualitative and quantitative analysis. All samples were subjected to qualitative and quantitative analysis processing, and the peak area of each chromatographic peak reflected the relative concentration of the substance. By substituting the linear equation and calculation formula, the qualitative and quantitative analysis results of the target substances in all samples were finally obtained.
[0177] (5) Metabolite concentration calculation
[0178] Standard solutions with different concentrations were prepared, with the concentration range from 0.01 ng / mL to 500 ng / mL, and the mass spectrometry peak intensity data corresponding to each concentration standard was obtained. By drawing the standard curve between the concentration ratio of external standard and internal standard (Concentration Ratio) and the area ratio (Area Ratio), the concentration of each substance was further analyzed. The peak area ratio value detected in all samples was substituted into the linear equation of the standard curve, and the concentration calculation formula was applied. In the MultiQuant 3.0.3 software, the dilution factor was set to 3, and the concentration (ng / mL) of the substance in the sample was finally calculated, so as to obtain the content data of the substance in the sample.
[0179] (6) Experimental quality control
[0180] The reproducibility of metabolite extraction and detection, i.e. technical replicates, can be evaluated by total ion chromatogram overlay analysis of mass spectrometry data of different quality control (QC) samples. The stability of the instrument is crucial to ensure the reliability and reproducibility of the data. The coefficient of variation (CV), i.e. the ratio of the standard deviation to the mean of the raw data, is used to reflect the dispersion of the data. The frequency of the CV values of substances lower than the reference value can be analyzed by the empirical cumulative distribution function (ECDF). If the proportion of substances with low CV values in the QC samples is high, it indicates that the experimental data is relatively stable. When the CV values of the QC samples are all less than 0.3, it indicates that the data stability is good; if the proportion of substances with CV values less than 0.2 in the QC samples is more than 90%, it indicates that the data is very stable. In addition, the CV value change of the monitoring isotope internal standard is monitored. If the internal standard CV change is less than 20%, it indicates that the instrument maintains good stability during detection.
[0181] Data preprocessing was performed on the qualitative and quantitative results obtained in step 2. According to the 80% rule, when a metabolite is detected in at least 4 / 5 of the samples in a group, the metabolite is considered to be detectable. According to this rule, 83 metabolites were robustly detected in all 928 samples. For a small number of undetected metabolites (<1 / 5 samples), the minimum value detected in the group was filled in half to facilitate subsequent statistical analysis.
[0182] Example 3 Screening of gastric cancer metabolic markers based on BIO-FIRE algorithm
[0183] Algorithm development: This embodiment develops a biomarker identification and optimization algorithm based on functional and feature importance recursive enhancement (Biomarker Identification and Optimization Algorithm via Functional and Importance-based Recursive Enhancement, BIO-FIRE), and applies it to the screening of gastric cancer markers. Taking the data of 83 metabolites of 928 samples in Example 2 as an example, the implementation process of the algorithm mainly includes 3 steps:
[0184] I. Functional module clustering and interpretive enhancement
[0185] The main purpose of this step is to extract functional modules with biological significance from metabolic data, and to provide a solid functional framework and data support for subsequent biomarker screening and optimization through functional module clustering and interpretive enhancement. By accurately dividing the functional units of different metabolic processes, this method aims to reveal the key regulatory factors in metabolic pathways, thereby improving the overall biological interpretability and clinical application value.
[0186] Non-negative Matrix Factorization (NMF) is chosen as the core algorithm for this step, mainly based on its natural adaptability to non-negative data and additive interpretation ability. Metabolome data is essentially non-negative values, and NMF can naturally maintain the physical and biological significance of the data when processing such data. The decomposition results not only reveal the potential functional modules in the data, but also quantitatively reflect the contribution of each metabolite to different functional modules through the output weight matrix. Compared with traditional clustering methods, NMF can more effectively avoid information overlap and ensure the orthogonality between functional modules, thereby enhancing the independence and accuracy of functional interpretation.
[0187] In the NMF algorithm, the objective function is usually based on the reconstruction error minimization problem of matrix decomposition. NMF attempts to decompose a non-negative matrix X (usually a metabolite concentration matrix) into the product of two non-negative matrices W and H:
[0188] X ≈ W × H
[0189] Where: X is the original metabolite data matrix (rows represent metabolites, columns represent samples); W is the weight matrix of genes / metabolites to functional modules (representing the contribution of each metabolite in different functional modules); H is the activation matrix of functional modules (representing the weighted distribution of each sample on different functional modules).
[0190] NMF optimizes W and H by minimizing the reconstruction error:
[0191]
[0192] Where, Frobenius norm is used to measure the error between X and WH.
[0193] In addition, sparsity constraints or regularization terms can be introduced, such as:
[0194] L1 regularization (improve the sparsity of functional modules): Encourage elements in W or H to approach zero, so that each metabolite is mainly attributed to a small number of functional modules;
[0195] L2 regularization (prevent overfitting): Prevent individual metabolites from contributing too much to a functional module, improve simulation stability, and in general, the objective function is usually written as:
[0196]
[0197] Where:
[0198] λ1, λ2 are regularization parameters; ||W||1 and ||H||1 control sparsity, ensure the orthogonality between functional modules, and reduce the overlap of metabolites in multiple functional modules.
[0199] In specific implementation, the application improves the traditional NMF by adopting dynamic rank selection and consensus enhancement strategy. First, through multi-rank synchronous calculation (the rank range in the application is set to 3 to 8) combined with residual sum of squares inflection point analysis, the optimal decomposition rank (rank = 5 in the application) is finally determined, and the specific process is shown in Figure 3 . At the same time, 100 independent runs are performed for each candidate rank to generate a consensus matrix to verify the stability of the functional module. Further, the KL divergence (Kullback-Leibler Divergence) optimization strategy is adopted to strengthen the sparse association between metabolites and functional modules, and the top 20% features with the highest contribution to the functional module are screened out by the maximum contribution principle to eliminate noise interference. Finally, through the analysis of the time sequence trajectory and the pathway enrichment analysis, the biological annotation of the functional module is realized, and the Jaccard similarity is used to control the functional module overlap rate below 15%, ensuring the independence and accuracy of each functional module. The specific improvements are as follows:
[0200] 1. Dynamic rank selection and consensus enhancement strategy
[0201] The traditional NMF needs to manually specify the rank, while the application automatically determines the optimal rank through multi-rank synchronous calculation and stability verification.
[0202] (1) Multi-rank synchronous calculation and inflection point analysis
[0203] Function form: for candidate rank r ∈ {3, 4, …, 8}, parallel calculation of NMF decomposition:
[0204]
[0205] Where r represents the candidate decomposition rank (the number of functional modules), and the value range is 3 to 8; W r, H r represents the weight matrix and the activation matrix when the rank is r; X is the original metabolic data matrix (rows represent metabolites, and columns represent samples); represents the Frobenius norm, which is used to measure the error between X and WH; RSS represents the residual sum of squares (Residual Sum of Squares), which is used to evaluate the fitting effect of different ranks; the inflection point is the point where the slope changes the most in the RSS curve, which is used to automatically select the optimal rank (such as r = 5).
[0206] (2) Consensus matrix verifies the stability of the functional module
[0207] Function form: for each candidate rank r, 100 independent runs are performed to generate a consensus matrix Cr:
[0208]
[0209] where C r denotes the consensus matrix with rank r, dimension n x n, representing the frequency of metabolites co-occurring in the same functional module; N denotes the number of independent runs (N = 100); W r (k) denotes the weight matrix obtained in the kth independent run; Cluster k denotes the clustering function in the kth run, which converts W r (k) into a binary association matrix (1 represents that the metabolites belong to the same functional module).
[0210] Stability evaluation: select the optimal rank by uniformity and consistency of C.
[0211] 2. Improvement of objective function: introduce KL divergence optimization strategy
[0212] The improved objective function is as follows:
[0213]
[0214] where X denotes the original metabolite data matrix, dimension n x m (n: number of metabolites, m: number of samples); W denotes the weight matrix, dimension n x r (r: number of functional modules), representing the contribution of metabolites to functional modules; H denotes the activation matrix, dimension r x m, representing the distribution of functional modules in samples; denotes the square of the Frobenius norm, used to measure the matrix reconstruction error denotes the L1 regularization parameter, controlling the sparsity of W and H; λ3 denotes the KL divergence regularization parameter, adjusting the strength of sparsity constraint; denotes the sparsity prior matrix (such as generated by thresholding or low-rank template), guiding the sparse distribution of the weight matrix W; denotes the Kullback-Leibler divergence, measuring the difference between the weight matrix W and the prior sparse matrix ; i, j denote the row and column indices of the matrix (i: metabolite index, j: functional module or sample index).
[0215] 3. Feature screening and functional module independence constraint
[0216] (1) Maximum contribution principle to screen features
[0217] Function form: after column normalization of the weight matrix W, the top 20% of metabolites in each column (functional module) are retained, and the rest are removed
[0218] In addition to low contribution metabolites, improve the interpretability of functional modules:
[0219] Selected j = {i | W i,j ≥ Quantile(W :,j , 0.8)}
[0220] where Selected j represents the set of metabolites retained in functional module j (top 20%); W i,j represents the weight value of metabolite i in functional module j; W:,j represents the jth column of the weight matrix (all metabolite weights of functional module j); Quantile(W:,j, 0.8) represents the 80th percentile of weights in functional module j, only metabolites higher than this value are retained.
[0221] (2) Functional module overlap rate control
[0222] Function form: constrain functional module independence by Jaccard similarity:
[0223]
[0224] where W j , W k represent the weight vectors of functional module j and functional module k; Supp(W j ) represents the set of metabolites with non-zero weights in functional module j (i.e., {i | W i,j > 0}); Supp(W k ) represents the set of metabolites with non-zero weights in functional module k; ∩, ∪ represent the intersection and union operators of sets; 0.15 represents the maximum allowed overlap rate between functional modules (Jaccard similarity threshold).
[0225] Results: 83 metabolites were classified using unsupervised NMF clustering analysis, and the resulting clustered modules were functionally named according to KEGG annotation information, resulting in five metabolic functional modules: glutamine metabolism functional module (NMF cluster 1), lipid metabolism functional module (NMF cluster 2), host-microbial co-metabolism functional module (NMF cluster 3), metabolic network bridge functional module (NMF cluster 4), and branched-chain amino acid metabolism functional module (NMF cluster 5).
[0226] In addition, the nonnegative matrix factorization (NMF) algorithm was compared with two common unsupervised clustering methods—hierarchical clustering (hclust) and K-means clustering—and their respective clustering performance was systematically evaluated. The specific results are as follows: (1) Hierarchical clustering (hclust): The data distribution between different categories is severely unbalanced, and the number of metabolites in some categories is sparse, which cannot provide sufficient statistical support, thus limiting the effective screening of subsequent functional modules. Figure 4 A). (2) K-means clustering: The discriminative power between clusters is low, and there is significant data overlap within some clusters. Figure 4 B). (3) NMF algorithm: In contrast, the NMF algorithm can efficiently divide all metabolites into five functional modules, with significant differences between the functional modules and a relatively balanced distribution of metabolite numbers, thus providing a more reliable statistical basis for subsequent analysis. Figure 4 C). II. Identification of Latent Biomarkers Based on Feature Importance
[0227] The main goal of this step is to achieve accurate identification of potential biomarkers based on feature importance, thereby ensuring that the subsequently constructed biomarker subset has high interpretability and predictive power. By screening out features that significantly contribute to the response variable, this method can eliminate noise and redundant information, providing solid data support for building robust biomarker ensembles.
[0228] The Boruta algorithm was chosen for its unique feature selection mechanism and nonlinear modeling capabilities. Boruta constructs a null distribution model by generating randomly perturbed shadow features, and its iterative decision-making process (i.e., feature importance scores must consistently surpass shadow features) avoids false positives in single random forest calculations. Furthermore, its fully correlated feature preservation strategy captures features with low individual importance but significant combined effects, making it more suitable than methods like LASSO for revealing synergistic interactions between metabolites and providing a more comprehensive perspective for biomarker research.
[0229] In its implementation, this invention uses the `Boruta::Boruta()` function from the `Boruta` package in R to perform feature filtering on the functional module data pre-extracted by NMF. The parameters `ntree = 500` and `pValue = 0.001` are set to ensure high robustness and stability during the iteration process. The following is the specific training method and mathematical implementation process of the Boruta algorithm:
[0230] 1. Feature importance measurement (Gini importance)
[0231] Let the splitting criterion for feature Xj at decision tree node t be:
[0232]
[0233] where j represents the feature index (e.g. the number of a gene or functional module); t represents the current node in the decision tree; Gini impurity of node t; C represents the total number of classes of the target variable (in classification tasks); p(c|t) represents the proportion of samples in node t belonging to class c; N represents the total number of samples in node t; t L , t R represent left and right child nodes; N L , N R represent the number of samples in the left and right child nodes after splitting; ΔGini(j, t): the reduction of Gini impurity brought by feature j in node t.
[0234] The cumulative importance of feature Xj in the entire tree is:
[0235]
[0236] where, represents the cumulative importance of feature Xj in a single decision tree; t represents a node in the decision tree; Splits(j) represents the set of all nodes that use feature Xj for splitting in the entire tree; ΔGini(j, t) represents the reduction of Gini impurity brought by feature Xj when node t is split.
[0237] The importance score of the random forest is finally obtained by averaging all the trees.
[0238] 2. Shadow feature generation mechanism
[0239] For the original feature matrix X ∈ R n×p , the shadow feature matrix
[0240]
[0241] where, represents the generated shadow feature matrix; X represents the original feature matrix (the result after NMF extraction); :, k represents the kth column (all rows) of the matrix; Permute(·) represents a random permutation operation on the column vector; k represents the feature index (1 to p).
[0242] 3. Wilcoxon signed rank test
[0243] For each feature X j , calculate the difference in importance between it and the largest shadow feature:
[0244]
[0245] Construct statistics:
[0246]
[0247] where Importance j represents the importance score of original feature j; represents the importance score of shadow feature k;d j represents the importance difference between feature j and the best shadow feature; p represents the total number of variables or features in the dataset; R represents the number of iterations (total rounds of algorithm running);d j (i) represents the difference value of feature j in the i-th iteration; sign(·) represents the sign function (the sign of the value); rank(|·|) represents the rank ordering of the absolute value; W represents the Wilcoxon signed rank statistic.
[0248] 4. Bonferroni-corrected significance threshold
[0249] For m features to be tested, the adjusted significance level is:
[0250]
[0251] where α represents the original significance level (user-set value, here 0.001); m represents the number of simultaneous hypothesis tests (equal to the total number of features p); p represents the number of original features (dimension after NMF extraction); α adj represents the corrected significance threshold. Used to ensure that the overall error rate (Family-Wise Error Rate, FWER) does not exceed 0.001.
[0252] 5. Iteration termination condition
[0253] Define the feature state set where S represents the state set of all features; S j represents the state of the j-th feature; j represents the feature index; p represents the total number of features.
[0254] When the following condition is met for t consecutive iterations:
[0255]
[0256] the algorithm is terminated.
[0257] where s j (t)State of feature j at iteration t (Confirmed / Rejected / Tentative); p represents the total number of features; II(·) represents the indicator function (1 if condition is true, 0 otherwise); t represents the threshold of the number of iterations without state change (default 20).
[0258] Results: This combination of parameters not only greatly reduces the influence of random errors, but also filters out features closely related to the response variable through statistical significance test, laying a foundation for subsequent optimal marker subset construction and clinical application. Figure 5 To show part of the results during the running of the Boruta algorithm, Figure 5 A reflects the rigor of feature selection by the Boruta algorithm (judged as "important" when the importance of the original feature is significantly higher than all shadow features, "irrelevant" when significantly lower than all shadow features, and "tentative" when it cannot be clearly judged, and the confidence is continuously improved through iteration), Figure 5 B reflects the comprehensive coverage of important metabolites filtered by the Boruta algorithm to functional modules.
[0259] Further filter key metabolic markers using the Boruta feature selection algorithm, and finally obtain 57 potential metabolic markers (covering NMF clusters 2-5), and sort them according to feature importance scores.
[0260] III. Marker determination based on recursive optimization principle
[0261] The main goal of this step is to reasonably simplify and optimize the number of metabolic marker combinations. Combining the NMF clustering modules obtained in step one and the importance scores obtained in step two, gradually select markers. The specific process is as follows: Sort each marker according to its importance score, and add markers one by one in descending order until the selected marker set can maximize the coverage of all functional modules. In this process, the new marker added each time will further improve the coverage of functional modules through its synergistic effect with existing markers, ensuring that each selected marker has biological interpretability and can supplement previously selected markers.
[0262] Once the combination of selected markers can achieve the predetermined maximum functional module coverage, i.e., meet the maximum value of the number of clusters, the marker selection process stops. The core logic of this step is to maximize the explanatory coverage of functional modules by minimizing the number of markers, thereby improving the simplicity and diagnostic performance of the model. The detailed process of determining markers using the recursive optimization principle is shown in Figure 6 .
[0263] The results show that: when the first 5 metabolic markers are included, 1 NMF cluster can be covered; when the first 6 markers are included, 2 clusters can be covered; when the first 7 markers are included, 3 clusters can be covered; and when the first 12 markers are included, 4 clusters can be covered. Therefore, the first 12 potential metabolic markers are finally selected as the final modeling markers in this embodiment, including: nicotinamide, 9, 12, 13-trihydroxy-octadecenoic acid, phthalic acid monoethylhexyl ester, taurine, sphingosine, chondritic acid, 5'-deoxy-5'-methylthioadenosine, (±) 12-hydroxy-5Z, 8Z, 10E, 14Z-eicosatetraenoic acid ((±) 12-HETE), hypoxanthine, glycyl-phenylalanine, azelaic acid and N-phenylacetyl-L-glutamine. The specific information of the above 12 metabolic markers is shown in Table 2.
[0264] Table 2
[0265] Serial No. Chinese Name Molecular Formula CAS No. 1 Nicotinamide [C6H6N2O] 98-92-0 2 Taurine [C2H7NO3S] 107-35-7 3 Phthalic acid monoethylhexyl ester C 16 H 22 O4]]> 4376-20-9 4 Sphingosine C 18 H 37 NO2]]> 123-78-4 5 9,12,13-trihydroxy-octadecenoic acid C 18 H 34 O5]]> 97134-11-7 6 Glycyl-phenylalanine C 11 H 14 N2O3]]> 3321-03-7 7 N-phenylacetyl-L-glutamine C 13 H 16 N2O4]]> 28047-15-6 8 (+ / -) 12-HETE C 20 H 32 O3]]> 71030-37-0 9 5'-deoxy-5'-methylthioadenosine [C 11 H 15 N5O3S]]> 2457-80-9 10 Traumatic acid C 12 H 20 O4]]> 6402-36-4 11 Hypoxanthine [C5H4N4O] 68-94-0 12 Azeleic acid [C9H 16 O4]]> 123-99-9
[0266] In addition, by statistical analysis, the expression levels of 12 metabolites in the GC group (gastric cancer group) and the NGC group (non-gastric cancer group) were compared, and the results showed that the expression differences of all metabolites were statistically significant (P<0.001). Among them, 5'-deoxy-5'-methylthioadenosine and N-phenylacetyl-L-glutamine were significantly up-regulated in the plasma of the GC group, while the remaining 10 metabolic markers were significantly down-regulated in the plasma of the GC group. Figure 7 The above results show that the combination of the 12 metabolites can be used as potential metabolic markers for the diagnosis of GC, providing new biological basis for the early detection and accurate diagnosis of gastric cancer.
[0267] Overall, the present application integrates algorithmic parameter innovation and biological interpretation depth, and upgrades NMF and Boruta from a tool level application to a biology-oriented analysis framework, forming an efficient and optimized marker screening process, which not only improves the biological interpretability of the screened markers, but also ensures their statistical importance and model simplicity.
[0268] Example 4: Comparison of results of feature importance screening algorithm
[0269] In Example 3, Boruta algorithm and two common feature importance sorting algorithms, LASSO and RF, were compared. First, taking the 928 samples and 83 metabolites in Example 2 as an example, according to the importance scores calculated by each algorithm, the top 30 metabolites with the highest importance scores were selected as candidate metabolic markers, and the following comparisons were made from multiple dimensions:
[0270] (1) Overall and novelty of candidate markers
[0271] ①NMF functional module coverage: NMF divides all metabolites into 5 functional modules, and the metabolites screened by LASSO only cover 3 functional modules (60%); while the metabolites screened by Boruta and RF algorithms cover 4 functional modules (80%), indicating that the latter two have higher sensitivity and accuracy in reflecting the complexity and diversity of each functional module in the underlying metabolic network. Figure 8
[0272] ②Newly identified metabolite ratio: Among the top 30 metabolites screened by each algorithm, the metabolites screened by Boruta algorithm are all the differential metabolites identified by the early broad-target metabolome profiling (the newly identified metabolite ratio is 100%), while the metabolites screened by LASSO and RF algorithms have 3 from the later literature supplement (the newly identified metabolite ratio is 90%), indicating that Boruta algorithm can more accurately extract the core metabolites obtained from experimental data, ensuring the objectivity and reproducibility of the screening results. Figure 8
[0273] (2) Difference significance of candidate markers
[0274] By comparing the performance of metabolites screened by each algorithm in statistical indicators (such as FDR-corrected P-value sum and logFC absolute value sum), the results show that the metabolites screened by Boruta are more significantly different, with the smallest FDR-P-value sum and the largest logFC absolute value sum, supporting its advantage in grouping differentiation from a statistical point of view. Figure 9
[0275] (3) Distribution balance of marker combination and modeling robustness
[0276] Based on the 30 candidate metabolic markers screened by each algorithm, the final modeling marker combination was determined by recursive optimization principle, and the following comparisons were made on the three marker combinations:
[0277] ① Distribution balance: In terms of the up-regulation and down-regulation ratio of metabolites, the up-regulation and down-regulation ratio of metabolites in the LASSO algorithm screening results is 0:6 (no up-regulation metabolites), the ratio of RF algorithm screening results is 1:8 (containing 1 up-regulation metabolite), and the ratio of Boruta screening results is 1:5 (containing 2 up-regulation metabolites), indicating that Boruta algorithm has a relative advantage in the balance of up-regulation and down-regulation distribution.
[0278] In this embodiment, 153 non-gastric cancer subjects and 156 gastric cancer patients from four research centers (Beijing Friendship Hospital, Cancer Hospital, Chinese Academy of Medical Sciences, Hubei Provincial People's Hospital, Tongji Hospital, Tongji Medical College, Huazhong University of Science and Technology) were collected as an independent validation cohort. The inclusion and exclusion criteria of all subjects, targeted detection methods of plasma samples and data preprocessing procedures were consistent with those of Example 1 to ensure the comparability of the data and the reliability of the analysis.
[0279] 2. Model robustness: The modeling cohort in Example 2 was divided into training and test sets in a ratio of 3:1. The RF model was used to build diagnostic models for the above three combinations of metabolic markers on the training set, and the models were validated on the independent validation cohort. The results showed that the combination of 12 markers selected by the Boruta algorithm had the highest AUC value, accuracy and specificity in the test set compared to the other two algorithms, indicating that this marker combination had a clear advantage in modeling robustness. Figure 10 ).
[0280] Example 5: Establishment and validation of gastric cancer diagnosis model based on multiple machine learning algorithms
[0281] 1. Model establishment and performance comparison based on 10 machine learning algorithms
[0282] To improve the early diagnosis efficiency of gastric cancer and optimize the modeling method of multi-marker combination, this embodiment used 10 advanced machine learning algorithms (RF, TreeNet, GBM, XGBoost, SVM Radial, LASSO, LR, KNN, NNET, NB) to model the 12 metabolic markers selected by the above BIO-FIRE algorithm (Table 2).
[0283] In the data analysis process, the train function in the caret package was used for model training, and the method parameter was set to select different machine learning algorithms. Hyperparameter optimization used trainControl(method="cv", number=10, classProbs=TRUE) for 10-fold cross-validation to avoid model overfitting. When dividing the data set, in order to ensure the generalization ability of the model, this embodiment divided the modeling cohort in Example 2 into training and test sets in a ratio of 3:1, used the training set for modeling, and evaluated the model performance in the test set.
[0284] Model performance was evaluated by receiver operating characteristic (ROC) curve and area under curve (AUC). The results showed that the gastric cancer diagnosis models established by 10 machine learning algorithms performed well in distinguishing gastric cancer and non-gastric cancer patients, with AUC of the test set ranging from 0.877 to 0.921, sensitivity ranging from 0.783 to 0.879, and specificity ranging from 0.795 to 0.923, indicating that the models had good stability and generalization ability Figure 11 A). To select the optimal machine learning algorithm, the sensitivity, specificity, accuracy, F1 value, positive predictive value, and negative predictive value of the models constructed by the 10 machine learning algorithms were compared using radar charts, and the results showed that the comprehensive performance of the RF model was the best Figure 11 B), so it was selected as the final gastric cancer diagnosis model.
[0285] 2Independent verification of the best diagnosis model based on 12 metabolic markers
[0286] To further evaluate the stability and generalization ability of the best diagnosis model based on 12 metabolic markers constructed in this embodiment, the independent verification cohort in Example 4 was used for verification.
[0287] In the independent verification cohort, the best-performing random forest (RF) model was externally verified, and the results showed that the AUC (area under curve) of the model could reach 0.95, the diagnostic accuracy was 90.3%, and both the sensitivity and specificity were over 90%, which had a significant advantage over the diagnostic performance of traditional tumor diagnostic markers Figure 11 C-D), indicating that the model still had excellent discrimination ability and good clinical application potential in independent populations, and could provide reliable auxiliary decision support for early screening and accurate diagnosis of gastric cancer.
[0288] Example 6: Effect of the number of functional modules covered on the diagnosis performance of gastric cancer
[0289] To further confirm the advantage of the metabolic marker combination covering multiple functional modules in the metabolic marker combination determined in this application, the diagnostic performance of the metabolic marker combination covering multiple functional modules and the metabolic marker combination covering a single functional module was compared. The metabolic marker combination covering 4 functional modules (nicotinamide + phloretic acid + 5'-deoxy-5'-methylthioadenosine + N-phenylacetyl-L-glutamine) and the metabolic marker combination covering a single functional module (9,12,13-trihydroxy-octadecenoic acid + 2-aminoethanesulfonic acid / taurine + hypoxanthine + azelaic acid) were modeled using the RF algorithm (modeled according to the method of Example 5), and the samples were selected from the independent verification cohort described in Example 4, and the results are shown in Figure 12 .
[0290] Example 7: Application of 12-metabolite combinations in complex clinical scenarios
[0291] In clinical practice, early diagnosis and subtype identification of gastric cancer still face many challenges, especially in the following complex or special diagnostic scenarios: (1) very early gastric cancer (stage IA gastric cancer): due to low tumor burden, traditional imaging and serum markers often have difficulty in accurately identifying, at the same time, this part of patients is the absolute indication of endoscopic resection of early gastric cancer, early diagnosis and treatment is of great significance to improve the survival rate of gastric cancer. (2) gastric cancer with negative traditional tumor markers (CA19-9, CA72-4, CEA): these patients cannot be effectively screened by existing serological means and are easily missed; (3) gastric cancer with tumor diameter <4 cm: smaller tumors may not be obvious in imaging or traditional biomarker detection, and are easily missed; (4) gastroesophageal junction cancer: this subtype has anatomical and biological characteristics that overlap with both gastric cancer and esophageal cancer, making diagnosis more difficult.
[0292] To verify the diagnostic value of the 12-metabolite combinations (Table 2) constructed in this embodiment in the above complex clinical scenarios, this embodiment also collected the clinical information of patients in the internal training set, internal test set in Example 2, and independent verification cohort in Example 4, and respectively took the patients with very early gastric cancer (stage IA gastric cancer) in the internal training set, internal test set, and independent verification cohort as the modeling or verification cohort for very early gastric cancer (stage IA gastric cancer), gastric cancer patients with negative traditional tumor markers as the modeling or verification cohort for traditional tumor marker-negative gastric cancer, gastric cancer patients with tumor diameter <4 cm as the modeling or verification cohort for gastric cancer with tumor diameter <4 cm, and gastroesophageal junction cancer patients as the modeling or verification cohort for gastroesophageal junction cancer, and used 10 machine learning algorithms to model in the internal training set and verify in the internal test set and independent verification cohort to ensure the comparability and stability of the results.
[0293] In the independent verification cohort, the RF diagnostic models based on 12 metabolites all showed high discrimination ability, and the specific AUC results of the independent verification cohort are shown in Table 7. Figure 13 .
Claims
1. A method for screening tumor metabolic markers, characterized in that, The screening method includes: S1) Initial screening metabolites were obtained from samples from cancer patients and non-cancer control groups, and their qualitative and quantitative analysis was performed. S2) Feature extraction was performed using a biomarker identification and optimization algorithm based on recursive enhancement of function and feature importance (BIO-FIRE) to further identify tumor metabolic biomarkers from the initial screening metabolites; The BIO-FIRE algorithm includes: A) Functional module clustering of the initially screened metabolites was performed using a nonnegative matrix factorization method; The nonnegative matrix factorization method described above introduces a KL divergence optimization strategy to enhance the sparse association between metabolites and functional modules; The nonnegative matrix factorization method includes: determining the optimal rank by combining multi-rank synchronous calculation with residual sum of squares inflection point analysis; B) The Boruta algorithm performs feature selection by calculating importance scores; The Boruta algorithm sets features Xj In decision tree nodes t The splitting criterion is: in, j Indicates feature index; t This represents the current node in the decision tree; Represents a node t The impurity of the gin; C This represents the total number of categories for the target variable; p(c|t) represents the node t The medium sample belongs to the category c The proportion; N represents a node t The total number of samples; N L , N R This represents the number of samples in the left and right child nodes after the split; t L , t R Indicates the left and right child nodes; Δ Gini ( j , t ):feature j At the node t The resulting reduction in impurity of the gin; C) Based on the principle of recursive optimization, the importance scores of each biomarker calculated in B) are sorted from high to low, and the subset of metabolic biomarkers to be included are selected sequentially starting from the first one, until the subset of metabolites that can maximize the coverage of the most functional modules is selected, which is the tumor metabolic biomarker.
2. The screening method according to claim 1, characterized in that, S1) includes: 1) Construct a tumor-specific metabolite database using broad-based targeted metabolomics and identify differentially expressed metabolites; 2) Targeted metabolomics screening to obtain and quantify the initial metabolites.
3. The screening method according to claim 2, characterized in that, The broad-targeted metabolomics uses quadrupole time-of-flight mass spectrometry in full-scan mode to detect metabolite ion pair signals in samples. The ion pair signals of metabolites are then transferred to tandem quadrupole composite linear ion trap mass spectrometry to obtain a tumor-specific metabolite database and identify differentially expressed metabolites.
4. The screening method according to claim 2, characterized in that, The targeted metabolomics detection scope includes the differentially expressed metabolites in step 1).
5. The screening method according to claim 2, characterized in that, The targeted metabolomics uses liquid chromatography to separate differential metabolites and mass spectrometry to quantify metabolites through initial screening. The liquid chromatography is performed using a T3 column, and the mass spectrometry is performed using a T3 column.
6. The screening method according to claim 1, characterized in that, The objective function of the nonnegative matrix factorization method that introduces the KL divergence optimization strategy is: in: X This represents the original metabolic data matrix, with dimensions of [missing information]. n × m ,in n Represents the amount of metabolites. m Represents the number of samples; W This represents the weight matrix, with dimension 1. n × r ,in r The number of functional modules represents the contribution of metabolites to the functional modules; H This represents the activation matrix, with dimension 1. r × m , representing the distribution of functional modules in the sample; The square of the Frobenius norm is used to measure the error in matrix reconstruction. ; λ1 and λ2 represent L1 regularization parameters, controlling... W and H sparsity; λ3 represents the KL divergence regularization parameter, which adjusts the strength of the sparsity constraint; Represents the sparsity prior matrix, guiding weight matrix W sparse distribution; This represents the Kullback-Leibler divergence, which measures the weight matrix. W With prior sparse matrix Differences; i,j Represents the row and column indices of the matrix, where i Represents a metabolite index. j Represents a functional module or sample index.
7. The screening method according to claim 6, characterized in that, The sparse prior matrix is generated through thresholding or a low-rank template.
8. The screening method according to claim 1, characterized in that, Determining the optimal rank involves decomposing candidate rank r∈{3,4,…,8} using a parallel computation of nonnegative matrix factorization methods. in, r This represents the candidate decomposition rank, with a value ranging from 3 to 8; W r, H r Indicates rank as r The weight matrix and activation matrix at time; X is the original metabolic data matrix, where rows represent metabolites and columns represent samples; This represents the Frobenius norm, used to measure the error between X and WH; RSS represents the sum of squared residuals, i.e. , used to evaluate the fitting effect of different ranks; The inflection point is the point in the RSS curve where the slope changes the most, and it is used to automatically select the optimal rank.
9. The screening method according to claim 8, characterized in that, The optimal rank r=5.
10. The screening method according to claim 1, characterized in that, The feature index includes the gene or module number.
11. The screening method according to any one of claims 1-10, characterized in that, The tumor in question is either a hematologic tumor or a solid tumor.
12. The screening method according to claim 11, characterized in that, The tumor is gastric cancer, and the functional modules include a glutamine metabolism module, a lipid metabolism module, a host-microbe co-metabolism module, a metabolic network bridging module, and a branched-chain amino acid metabolism module.
13. The screening method according to any one of claims 1-10, characterized in that, The sample was bodily fluid.
14. The screening method according to claim 13, characterized in that, The bodily fluids mentioned are selected from blood, plasma, or serum.
15. The screening method according to any one of claims 1-10, characterized in that, The number of tumor metabolic markers is 10-15.
16. The screening method according to claim 15, characterized in that, The tumor metabolic markers consist of 12 items.
17. The screening method according to claim 16, characterized in that, The tumor metabolic markers include nicotinamide, taurine, monoethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecanoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (±)12-hydroxy-5Z,8Z,10E,14Z-eicosatetraenoic acid ((±)12-HETE), 5'-deoxy-5'-meththioadenosine, tuyereic acid, hypoxanthine, and azelaic acid.
18. A computer-readable storage medium or an apparatus comprising a computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including a BIO-FIRE algorithm, the BIO-FIRE algorithm comprising: A) Functional module clustering of the initially screened metabolites was performed using a nonnegative matrix factorization method; The nonnegative matrix factorization method described above introduces a KL divergence optimization strategy to enhance the sparse association between metabolites and functional modules; The nonnegative matrix factorization method includes: determining the optimal rank by combining multi-rank synchronous calculation with residual sum of squares inflection point analysis; B) The Boruta algorithm performs feature selection by calculating importance scores; The Boruta algorithm sets features Xj In decision tree nodes t The splitting criterion is: in, j Indicates feature index; t This represents the current node in the decision tree; Represents a node t The impurity of the gin; C This represents the total number of categories for the target variable; p(c|t) represents the node t The medium sample belongs to the category c The proportion; N represents a node t The total number of samples; N L , N R This represents the number of samples in the left and right child nodes after the split; t L , t R Indicates the left and right child nodes; Δ Gini ( j , t ):feature j At the node t The resulting reduction in impurity of the gin; C) Based on the principle of recursive optimization, the importance scores of each biomarker calculated in B) are sorted from high to low, and the subset of metabolic biomarkers to be included are selected sequentially starting from the first one, until the subset of metabolites that can maximize the coverage of the most functional modules is selected, which is the tumor metabolic biomarker.
19. A method for constructing a tumor diagnostic model, characterized in that, The construction method includes: i) Collect the concentration detection results of metabolic markers in the cancer patient group and the non-cancer control group; ii) Construct a tumor diagnostic model based on the information collected in step i); Alternatively, the construction method may include: I) Collect samples from the subjects and detect the concentrations of metabolic markers; II) Clinical diagnosis was performed on the subjects, and the subjects were divided into a cancer patient group and a non-cancer control group; III) Construct a tumor diagnostic model based on the detection results of step I) and the diagnostic results of step II); The metabolic biomarker is the metabolic biomarker obtained by screening using any of the screening methods described in claims 1-17.
20. The construction method according to claim 19, characterized in that, The tumor in question is either a hematologic tumor or a solid tumor.
21. The construction method according to claim 20, characterized in that, The tumor in question is stomach cancer.
22. The construction method according to claim 19, characterized in that, The sample was a bodily fluid.
23. The construction method according to claim 22, characterized in that, The bodily fluids mentioned are selected from blood, plasma, or serum.
24. The construction method according to claim 19, characterized in that, The diagnostic model was constructed using a machine learning model.
25. The construction method according to claim 24, characterized in that, The machine learning model includes one or more of the following: logistic regression, k-nearest neighbor algorithm, Naive Bayes, support vector machine, random forest, XGBoost, TreeNet, gradient boosting machine, LASSO, or neural network.
26. The application of a metabolic marker in the preparation of products for diagnosing gastric cancer, characterized in that, The metabolic markers include nicotinamide, taurine, monoethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecanoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (±)12-hydroxy-5Z,8Z,10E,14Z-eicosatetraenoic acid ((±)12-HETE), 5'-deoxy-5'-meththioadenosine, tuyereic acid, hypoxanthine, and azelaic acid.
27. The application according to claim 26, characterized in that, The metabolic markers are metabolic markers in plasma.
Citation Information
Patent Citations
Metabolic marker for gastric cancer diagnosis or monitoring and screening method and application thereof
CN119936223A
Application of metabolic marker combination in diagnosis of gastric cancer
CN120253932A