Application of metabolic marker combination in diagnosis of gastric cancer

Through targeted metabolic markers and building machine learning models, the existing gastric cancer diagnosis methods are solved, and non-invasive and efficient early diagnosis of gastric cancer is achieved.

CN120253932APending Publication Date: 2025-07-04BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510317168.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing gastric cancer diagnosis methods such as electronic gastroscopy, imaging and tumor marker detection have high invasiveness, high cost or insufficient sensitivity and specificity, making it difficult to achieve early and efficient gastric cancer screening.

Method used

Targeted metabolomics technology combined with recursive feature elimination method and random forest algorithm were used to screen out a set of metabolic markers, including dihydro-D-sphingosine and (+/-)12-hydroxyecosattetradeonic acid, to construct a machine learning model for diagnosis of gastric cancer.

Benefits of technology

It improves the sensitivity and specificity of gastric cancer diagnosis, provides a non-invasive and efficient diagnostic tool that can accurately identify gastric cancer patients and reduce the rate of missed diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120253932A_ABST
    Figure CN120253932A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biological medicine, and particularly relates to application of a metabolic marker combination in diagnosis of gastric cancer. According to the method, specific metabolites in plasma or tissue of a gastric cancer patient are accurately quantified and analyzed through a targeted metabonomics technology, an optimal metabolic marker combination is developed through machine learning, and a novel non-invasive, efficient and accurate gastric cancer detection method is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedicine technology, and particularly relates to the application of a combination of metabolic markers in the preparation of a product for diagnosing gastric cancer. Background Art

[0002] Gastric cancer (GC), as one of the common malignant tumors, refers to an epithelial-derived malignant tumor originating from the stomach. Gastric cancer can be affected by various factors. Under the action of various non-cancerous diseases of the stomach (such as gastric polyps, gastric adenomas, chronic gastritis, and gastric ulcers, etc.), Helicobacter pylori infection, as well as poor eating habits, environment, and genetics, etc., the occurrence of gastric cancer can be caused. Therefore, early diagnosis and treatment of gastric cancer will help reduce the mortality rate caused by gastric cancer, reduce the treatment cost, and obtain better treatment effects.

[0003] Currently, the clinical examination methods for gastric cancer mainly include electronic gastroscopy examination, imaging examination, and tumor marker detection. Among them, electronic gastroscopy examination refers to the examination through a gastroscope, which can directly observe the specific situation in the patient's stomach. Once ulcers or atrophic gastritis are found in the stomach, further pathological examination can be carried out for diagnosis. As an intuitive diagnostic method, electronic gastroscopy examination has a high accuracy rate and is the gold standard for the diagnosis of gastric cancer. However, its high invasiveness, high cost, and resource dependence limit its wide application in large-scale screening, especially in areas with limited medical resources. Imaging examinations mainly include techniques such as abdominal CT, abdominal ultrasound, and PET examination, etc., which can elaborate on the changes in the condition from multiple angles such as morphological characteristics, density, and enhancement patterns. Compared with electronic gastroscopy examination, imaging examination can quickly display information such as the size, shape, and whether there is metastasis of the lesion area. It is a non-invasive detection method and can be used as a diagnostic method for the initial diagnosis of cancer, which helps to diagnose the degree of gastric cancer differentiation, pathological type, TNM stage, and evaluate the chemotherapy efficacy, etc. But ultimately, it still needs to be combined with pathological examination, so it can only be used as an auxiliary diagnostic means for gastric cancer. Tumor marker detection provides a relatively non-invasive screening option and is widely used in the clinical diagnosis of gastric cancer, such as tumor markers such as CA72-4, CEA, and CA19-9. By detecting specific proteins or genetic markers in the blood, this method can preliminarily screen for potential tumors without physically entering the body. Although this method is simple to operate and has a low cost, its effect in balancing sensitivity and specificity is poor, which may lead to a high missed diagnosis rate. Although these traditional methods are widely used clinically, their effects in early tumor screening are still limited. This current situation has given rise to an urgent need to develop non-invasive, highly sensitive, and highly specific diagnostic tools to achieve early detection and accurate diagnosis of gastric cancer.

[0004] Blood metabolic markers have been proven to directly reflect the phenotypic characteristics of cancer and play a key role in the occurrence and progression of gastric cancer. Therefore, they are expected to become potential markers for identifying high-risk populations. Compared with molecular markers upstream of pathways such as nucleic acids and proteins, metabolic molecules, as downstream biomarkers, can more directly reflect the phenotypic changes of diseases. Before the occurrence of diseases or during the stage when risk factors gradually accumulate, the body's endogenous substances will produce corresponding metabolic reactions, thus more sensitively capturing specific metabolic characteristics caused by the interaction of genetic and environmental factors. In addition, the types of metabolites are much fewer than the corresponding genes and proteins, and the slight changes in their expression can often be amplified at the metabolic level, enabling detection when there are subtle changes in metabolic characteristics and improving the sensitivity of detection. Summary of the Invention

[0005] In order to further improve the diagnostic efficacy of gastric cancer and provide a more suitable screening method for gastric cancer metabolic markers, the present invention uses targeted metabolomics technology combined with RFE-RF algorithm feature extraction to screen out a group of optimal metabolic markers. The specific scheme is as follows:

[0006] In a first aspect of the present invention, there is provided the use of a metabolic marker in the preparation of a product for diagnosing gastric cancer, wherein the metabolic marker comprises dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE).

[0007] Preferably, the metabolic marker comprises dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine.

[0008] More preferably, the metabolic marker further comprises one or more of taurine, mono-2-ethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, 5'-deoxy-5'-methylthioadenosine, traumatic acid, lysophosphatidylethanolamine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, caffeine, or creatinine.

[0009] Further preferably, the metabolic markers include nicotinamide, taurine, mono-2-ethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), 5'-deoxy-5'-methylthioadenosine, traumatic acid, indole-3-carboxaldehyde, dihydro-D-sphingosine, lysophosphatidylethanolamine, hypoxanthine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, p-cresol sulfate ammonium salt, caffeine, inosine, and creatinine.

[0010] The metabolic markers are metabolic markers in body fluids and / or tissues.

[0011] Preferably, the body fluids are selected from blood, plasma, and serum.

[0012] Preferably, the tissue is gastrointestinal tissue, preferably cancer tissue or para-cancer tissue.

[0013] The product includes a reagent for detecting metabolic markers. Preferably, the reagent detects the concentration of metabolic markers.

[0014] Preferably, the methods for detecting metabolic markers include one or more of nuclear magnetic resonance spectroscopy, mass spectrometry, chromatography, high performance liquid chromatography-tandem mass spectrometry, Fourier transform ion cyclotron resonance, ion mobility spectrometry, electrochemical detection, Raman spectroscopy, or radioactive labeling.

[0015] Preferably, the product includes a diagnostic model, a kit, a test strip, a chip, or a device.

[0016] Preferably, the product further includes a reagent for sample processing and / or a standard for metabolic markers.

[0017] In the second aspect of the present invention, a method for constructing a gastric cancer diagnostic model is provided. The construction method includes:

[0018] i) Collecting the concentration detection results of metabolic markers in a gastric cancer patient group and a non-gastric cancer control group;

[0019] ii) Constructing a gastric cancer diagnostic model based on the information collected in step i);

[0020] Alternatively, the construction method includes:

[0021] I) Collecting samples from subjects and detecting the concentration of metabolic markers;

[0022] II) Clinically diagnose the subjects and divide the subjects into a gastric cancer patient group and a non-gastric cancer control group;

[0023] III) Construct a gastric cancer diagnosis model based on the test results in step I) and the diagnosis results in step II).

[0024] The metabolic markers include dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE). Preferably, the metabolic markers include dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine.

[0025] More preferably, the metabolic markers further include one or more of taurine, mono(2-ethylhexyl) phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, 5'-deoxy-5'-methylthioadenosine, traumatic acid, lysophosphatidylethanolamine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, caffeine, or creatinine.

[0026] More preferably, the metabolic markers include nicotinamide, taurine, mono(2-ethylhexyl) phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), 5'-deoxy-5'-methylthioadenosine, traumatic acid, indole-3-carboxaldehyde, dihydro-D-sphingosine, lysophosphatidylethanolamine, hypoxanthine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, p-cresol sulfate ammonium salt, caffeine, inosine, and creatinine.

[0027] Preferably, the sample is body fluid and / or tissue.

[0028] More preferably, the body fluid is selected from blood, plasma, and serum.

[0029] More preferably, the tissue is gastrointestinal tissue, preferably cancer tissue or adjacent cancer tissue.

[0030] The non-gastric cancer control group includes healthy people or patients with benign gastric diseases such as chronic gastritis, erosive gastritis, gastric ulcer, and gastric polyps.

[0031] The concentration of metabolic markers in the detection sample includes one or more of nuclear magnetic resonance spectroscopy, mass spectrometry, chromatography, high performance liquid chromatography-tandem mass spectrometry, Fourier transform ion cyclotron resonance, ion mobility spectrometry, electrochemical detection, Raman spectroscopy, or radioactive labeling for detecting the concentration of metabolic markers.

[0032] The methods for clinical diagnosis include, but are not limited to, electronic gastroscopy.

[0033] Preferably, the algorithm used for constructing the diagnostic model is a machine learning model, preferably including gradient boosting machine.

[0034] In the third aspect of the present invention, a gastric cancer diagnostic model obtained by the construction method as described in the second aspect is provided.

[0035] In the fourth aspect of the present invention, an application of a metabolic marker in constructing a diagnostic model for gastric cancer is provided.

[0036] The metabolic marker includes dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE).

[0037] Preferably, the metabolic marker includes dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine.

[0038] More preferably, the metabolic marker further includes one or more of taurine, mono(2-ethylhexyl) phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, 5'-deoxy-5'-methylthioadenosine, traumatic acid, lysophosphatidylethanolamine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, caffeine, or creatinine.

[0039] More preferably, the metabolic marker includes nicotinamide, taurine, mono(2-ethylhexyl) phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), 5'-deoxy-5'-methylthioadenosine, traumatic acid, indole-3-carboxaldehyde, dihydro-D-sphingosine, lysophosphatidylethanolamine, hypoxanthine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, p-cresol sulfate ammonium salt, caffeine, inosine, and creatinine.

[0040] In a fifth aspect of the present invention, a metabolic marker for diagnosing gastric cancer is provided. The metabolic marker includes dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE).

[0041] Preferably, the metabolic marker includes dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine.

[0042] More preferably, the metabolic marker further includes one or more of taurine, mono(2-ethylhexyl) phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, 5'-deoxy-5'-methylthioadenosine, traumatic acid, lysophosphatidylethanolamine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, caffeine, or creatinine.

[0043] More preferably, the metabolic marker includes nicotinamide, taurine, mono(2-ethylhexyl) phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), 5'-deoxy-5'-methylthioadenosine, traumatic acid, indole-3-carboxaldehyde, dihydro-D-sphingosine, lysophosphatidylethanolamine, hypoxanthine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, p-cresol sulfate ammonium salt, caffeine, inosine, and creatinine.

[0044] In a sixth aspect of the present invention, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the following operations are implemented: obtaining the concentration of the metabolic marker in a sample to be tested, inputting the concentration of the metabolic marker into a constructed machine learning model, outputting a probability value, and comparing the output probability value with a threshold to determine whether the subject has gastric cancer. The metabolic marker includes dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE).

[0045] Preferably, the metabolic markers include dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine. More preferably, they are the metabolic markers of the fifth aspect above.

[0046] In a seventh aspect of the present invention, there is provided a device comprising a computer-readable storage medium.

[0047] In an eighth aspect of the present invention, there is provided a method for screening gastric cancer metabolic markers, the screening method comprising:

[0048] 1) Collect samples from a gastric cancer patient group and non-gastric cancer controls to establish a gastric cancer-specific metabolite library;

[0049] 2) Screen metabolites through differential ion pair screening and ion pair deconvolution screening and supplement common gastric cancer metabolites to obtain pre-screened metabolites;

[0050] 3) Use a T3 column for liquid chromatography separation of metabolites and mass spectrometry for metabolite quantification;

[0051] 4) Screen for differential metabolites based on differential significance analysis using the Wilcoxon rank sum test, where the p-value is corrected for false discovery rate;

[0052] 5) Perform feature extraction on the screened differential metabolites using a machine learning algorithm of recursive feature elimination combined with random forest (RFE-RF) to further determine a metabolite subset from the screened differential metabolites;

[0053] Preferably, the pre-screened metabolites include amino acids and their derivatives, organic acids and their derivatives, nucleosides, nucleotides and their derivatives, alkaloids and their derivatives, acylcarnitines, bile acids and their derivatives, amines, carbohydrates and their derivatives, and other functional metabolites (such as including organic oxygen compounds, steroids, fatty acid derivatives).

[0054] Preferably, the metabolite subset includes dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), preferably nicotinamide, taurine, mono-2-ethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), 5'-deoxy-5'-methylthioadenosine, traumatic acid, indole-3-carboxaldehyde, dihydro-D-sphingosine, lysophosphatidylethanolamine, hypoxanthine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, p-cresol sulfate ammonium salt, caffeine, inosine, and creatinine;

[0055] Preferably, the RFE-RF uses the rfe function in the caret package to implement the RFE process.

[0056] More preferably, the recursive steps of the RFE include:

[0057] ① Training a random forest model using all features; ② Calculating the importance score I_j of each feature; ③ Removing the feature with the lowest importance score; ④ Repeating steps ①-③ until the number of remaining features reaches the set target number of features;

[0058] Preferably, the target number of features is obtained from the results of the change in model accuracy with the number of features.

[0059] In a specific embodiment of the present invention, the target number of features is 20-30.

[0060] In a ninth aspect of the present invention, a method for diagnosing gastric cancer is provided, and the method includes detecting the concentration of a metabolic marker in a sample of a subject.

[0061] The metabolic marker includes dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE).

[0062] Preferably, the metabolic marker includes dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine.

[0063] Further preferably, the metabolic markers further include one or more of taurine, monoethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, 5'-deoxy-5'-methylthioadenosine, traumatic acid, dihydro-D-sphingosine, lysophosphatidylethanolamine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, caffeine or creatinine.

[0064] Further preferably, the metabolic markers include nicotinamide, taurine, monoethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), 5'-deoxy-5'-methylthioadenosine, traumatic acid, indole-3-carboxaldehyde, dihydro-D-sphingosine, lysophosphatidylethanolamine, hypoxanthine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, p-cresol sulfate ammonium salt, caffeine, inosine and creatinine.

[0065] Preferably, the sample is body fluid and / or tissue.

[0066] Further preferably, the body fluid is selected from blood, plasma, and serum.

[0067] Further preferably, the tissue is gastrointestinal tissue, preferably cancer tissue or adjacent tissue of cancer.

[0068] The method includes one or more of nuclear magnetic resonance spectroscopy, mass spectrometry, chromatography, high performance liquid chromatography-tandem mass spectrometry, Fourier transform ion cyclotron resonance, ion mobility spectrometry, electrochemical detection, Raman spectroscopy or radioactive labeling to detect the concentration of metabolic markers.

[0069] Preferably, the method includes inputting the concentration of the detected metabolic markers into a constructed machine learning model, outputting a probability value, and comparing the output probability value with a threshold to determine whether the subject has gastric cancer.

[0070] The threshold is obtained from previous experiments, that is, by experiments and data analysis of the difference degree of metabolic markers between the gastric cancer patient group and the non-gastric cancer control group, training the machine learning model to determine the threshold.

[0071] If the metabolic marker described in this application has a difference or significant difference from the threshold (the difference is statistically significant, for example, p < 0.05, p < 0.01, p < 0.001, p < 0.0001), it is determined as a disease.

[0072] In a tenth aspect of the present invention, there is provided a diagnostic system for gastric cancer, and the diagnostic system includes a device, unit or module for judging whether a subject has gastric cancer according to the concentration of a metabolic marker.

[0073] Preferably, the diagnostic system includes:

[0074] A data detection device, unit or module for detecting the concentration of a metabolic marker in a sample;

[0075] A data input device, unit or module for inputting the concentration data of the metabolic marker;

[0076] A data analysis device, unit or module for analyzing the gastric cancer risk based on the concentration data of the metabolic marker;

[0077] A data output device, unit or module for outputting the analysis result of whether an individual has gastric cancer.

[0078] The algorithm used in the analysis is a machine learning model, preferably including Gradient Boosting Machine (GBM).

[0079] The metabolic marker includes dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE).

[0080] Preferably, the metabolic marker includes dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine and N-phenylacetyl-L-glutamine.

[0081] More preferably, the metabolic marker further includes one or more of taurine, mono-2-ethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, 5'-deoxy-5'-methylthioadenosine, traumatic acid, lysophosphatidylethanolamine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, caffeine or creatinine.

[0082] Further preferably, the metabolic markers include nicotinamide, taurine, monoethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), 5'-deoxy-5'-methylthioadenosine, traumatic acid, indole-3-carboxaldehyde, dihydro-D-sphingosine, lysophosphatidylethanolamine, hypoxanthine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, p-cresol sulfate ammonium salt, caffeine, inosine, and creatinine.

[0083] In an eleventh aspect of the present invention, there is provided a device, which comprises the diagnostic system described in the tenth aspect.

[0084] In a twelfth aspect of the present invention, there is provided an application of the metabolic markers described in the fifth aspect in the preparation of a product for differentiating gastric cancer from non-gastric cancer.

[0085] In a thirteenth aspect of the present invention, there is provided a method for studying the metabolic mechanism of gastric cancer, which includes analyzing the concentration trend of the metabolic markers described in the fifth aspect in a sample.

[0086] The metabolic markers include dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE).

[0087] Preferably, the metabolic markers include dihydro-D-sphingosine, (+ / -)12-HETE, indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine.

[0088] The analysis includes evaluating the fold change of the metabolite in the gastric cancer patient group and the non-gastric cancer control group.

[0089] Preferably, the tissue is gastrointestinal tissue, preferably cancer tissue or adjacent cancer tissue.

[0090] The term "subject" can be a human or a non-human animal, including "patient", "suspected patient", and "healthy individual", etc. This term does not indicate a specific age or gender, and covers adult or neonatal subjects as well as fetuses. The non-human animal can be a wild animal, a zoo animal, an economic animal, a pet, a laboratory animal, etc. Specifically, the non-human animal includes but is not limited to pigs, cows, sheep, horses, donkeys, foxes, raccoons, minks, camels, dogs, cats, rabbits, rats (such as rats, mice, guinea pigs, hamsters, gerbils, chinchillas, squirrels), or monkeys, etc.

[0091] The term "diagnosis" refers to ascertaining whether a patient has a disease or disorder in the past, at the time of diagnosis, or in the future, or to ascertaining the progression of a disease or its likely future progression.

[0092] The present invention has the following beneficial effects:

[0093] (1) This application uses targeted metabolomics technology, combines the corrected P-value and the recursive elimination method with the random forest algorithm to accurately screen out gastric cancer-specific plasma metabolite markers. Compared with traditional non-targeted metabolomics methods, the present invention has significant advantages in the confidence of metabolite identification and the accuracy of quantification.

[0094] (2) The metabolites screened in this application are derived from a self-constructed gastric cancer plasma-specific metabolite library and combined with literature-derived metabolites. Compared with the existing research methods mainly based on network public databases, the present invention has a relative advantage in metabolite novelty.

[0095] (3) This application has identified 8 metabolite markers with both plasma and tissue differences. Through in-depth research on these 8 core interactive metabolites, the complexity of gastric cancer cell metabolic reprogramming and the interaction between the local tumor microenvironment and the systemic metabolic state can be further revealed.

[0096] (4) The 8 metabolite markers with both plasma and tissue differences in this application and 25 plasma metabolite markers are used for the diagnosis of gastric cancer, with extremely high specificity, sensitivity, and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings, wherein:

[0098] Figure 1 It shows a visualization analysis diagram of plasma differential metabolites between gastric cancer patients and non-gastric cancer control groups;

[0099] Figure 2 It shows a GO (Gene Ontology Enrichment Analysis) enrichment diagram;

[0100] Figure 3 It shows a curve graph of the change of model accuracy with the number of features;

[0101] Figure 4A It shows a receiver operating characteristic curve (ROC) of the diagnostic efficacy of PMB-P25 for gastric cancer in an independent validation cohort, where CEA is carcinoembryonic antigen, and CA724 and CA199 are carbohydrate antigens;

[0102] Figure 4B It shows the performance display of the PMB-P25 diagnostic model in an independent validation cohort;

[0103] Figure 5 The confusion matrix of applying the PMB-P25 diagnostic model in the independent validation cohort is shown as follows;

[0104] Figure 6 The receiver operating characteristic curve (ROC) of PMB-P8 in the independent validation cohort is shown as follows;

[0105] Figure 7 The receiver operating characteristic curve (ROC) of the GBM gastric cancer diagnostic model with one metabolic biomarker, namely dihydro-D-sphingosine, in the independent validation cohort is shown as follows;

[0106] Figure 8 The receiver operating characteristic curve (ROC) of the GBM gastric cancer diagnostic model with one metabolic biomarker, namely (+ / -)12-HETE, in the independent validation cohort is shown as follows;

[0107] Figure 9 The receiver operating characteristic curve (ROC) of the GBM gastric cancer diagnostic model with two metabolic biomarkers, namely dihydro-D-sphingosine + (+ / -)12-HETE, in the independent validation cohort is shown as follows. Detailed implementation manners

[0108] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.

[0109] It should be noted that the methods used in the present invention are all conventional methods unless otherwise specified, and the reagents used in the present invention are all commercially available products unless otherwise specified.

[0110] The key instruments used in the examples are shown in Table 1.

[0111] Table 1

[0112] Name Model Brand HPLC-MS / MS QTRAP 6500+ SCIEX Centrifuge 5424R Eppendorf Centrifugal concentrator CentriVap LABCONCO Vortex mixer VORTEX-5 Kyllin-Be11

[0113] The experimental reagents used in the examples are shown in Table 2.

[0114] Table 2

[0115]

[0116]

[0117] Example 1 Screening of plasma metabolic biomarkers for gastric cancer

[0118] 1 Plasma sample collection

[0119] After obtaining the consent of the patients, peripheral venous blood plasma samples were collected from 460 non-gastric cancer subjects (including healthy individuals and patients with benign gastric diseases) and 468 gastric cancer patients (including early and advanced gastric cancer) from four research centers (Beijing Friendship Hospital, Cancer Hospital of the Chinese Academy of Medical Sciences, Hubei Provincial People's Hospital, Tongji Hospital Affiliated to Tongji Medical College of Huazhong University of Science and Technology). Among them, the diagnosis of gastric cancer patients was confirmed by postoperative pathology; the non-gastric cancer group samples included healthy individuals without gastric diseases after physical examination and patients with benign gastric diseases such as chronic gastritis, erosive gastritis, gastric ulcer, gastric polyps, and benign gastric tumors after hospital examination. All gastric cancer patients and non-gastric cancer group subjects had no history of other malignant tumors, no other major systemic diseases (including metabolic diseases such as severe kidney diseases or liver diseases), and no chronic medical history of long-term medication. Samples of healthy individuals and patients with benign diseases were included as the non-gastric cancer group (NGC), and samples of early and advanced gastric cancer patients were included as the gastric cancer group (GC).

[0120] All blood samples were collected in the early morning fasting state of the subjects, and then centrifuged to separate plasma, which was immediately stored in a -80 °C refrigerator. During the research process, plasma samples were taken out and thawed according to experimental requirements for subsequent analysis.

[0121] 2 Targeted metabolomics detection of plasma samples

[0122] 2.1 Determination of the targeted detection range

[0123] Based on the self-built plasma-specific metabolite library of gastric cancer in the early stage, 84 metabolites were determined through differential ion pair screening and ion pair deconvolution (MSI Level 1). At the same time, 32 common cancer-specific metabolites reported more than 5 times in the gastric cancer metabolome literature were supplemented. After removing duplicates, 102 metabolites were finally obtained, covering nine major substance categories, namely: amino acids and their derivatives, organic acids and their derivatives, nucleosides, nucleotides and their derivatives, alkaloids and their derivatives, acylcarnitines, bile acids and their derivatives, amines, carbohydrates and their derivatives, other functional metabolites (including organic oxygen compounds, steroids, fatty acid derivatives, etc.). Standard products of 102 metabolites to be detected were prepared respectively.

[0124] 2.2 Plasma metabolite extraction

[0125] Take out the sample collected in Step 1 from an -80°C refrigerator and thaw it on ice until there are no ice crystals left; after the sample is thawed, vortex it for 10 seconds to mix well. Take 50 μL of the sample, add 150 μL of the extraction solution (the extraction solution contains an isotope internal standard with a concentration of 100 ppm), then vortex and mix for 3 minutes, centrifuge at 12000 rpm and 4°C for 10 minutes, and leave it to stand overnight at a low temperature in a -20°C refrigerator. The next day, centrifuge at 12000 rpm and 4°C for 5 minutes, take 170 μL of the supernatant, and transfer it to a 96-well plate in sequence. After completing the protein precipitation treatment, seal the well plate for LC-MS / MS analysis. Additionally, take 20 μL from each sample and mix them to prepare a quality control sample (QC), and perform one collection at intervals of every 15 samples.

[0126] 2.3 Chromatography and Mass Spectrometry Acquisition Conditions

[0127] Targeted quantitative analysis uses a T3 column for metabolite separation, and the specific detection conditions are as follows:

[0128] T3 column liquid chromatography conditions: The chromatographic column is Waters ACQUITY UPLC HSS T3 C18 1.8 μm, 2.1 mm × 100 mm, the column temperature is set at 40°C, and the injection volume is 2 μL. Mobile phase A is an aqueous solution containing 0.1% acetic acid, and mobile phase B is an acetonitrile solution containing 0.1% acetic acid; the elution gradient is: 0 minutes, A:B = 95:5; 11.0 minutes, A:B = 10:90; 12.0 minutes, A:B = 10:90; 12.1 minutes, A:B = 95:5; 14.0 minutes, A:B = 95:5, and the flow rate is set at 0.4 mL / min.

[0129] Mass spectrometry conditions: The temperature of the electrospray ionization source (ESI) is set at 500°C, and the mass spectrometry voltages are 5500 V (positive ion mode) and -4500 V (negative ion mode) respectively. The pressure of ion source gas I (GS I) is 55 psi, the pressure of gas II (GS II) is 60 psi, and the pressure of the curtain gas (CUR) is 25 psi. The collision-induced ionization (CAD) is set to high. In the triple quadrupole (Qtrap), the optimized declustering voltage (DP) and collision energy (CE) are used for MRM mode scanning to detect the signals of each ion pair.

[0130] 2.4 Processing and Integration of Spectral Peak Areas

[0131] The mass spectrometry data was processed by MultiQuant 3.0.3 software. According to the retention time and peak shape information of the standards, the mass spectrometry peaks of the target substances in each sample were integrated and corrected to ensure the accuracy of qualitative and quantitative analysis. All samples were subjected to qualitative and quantitative analysis. The peak area of each chromatographic peak reflected the relative concentration of the substance. By substituting into the linear equation and calculation formula, the qualitative and quantitative analysis results of the target substances in all samples were finally obtained.

[0132] 2.5 Calculation of metabolite concentration

[0133] Standard solutions with different concentrations were prepared, with the concentration range from 0.01 ng / mL to 500 ng / mL, and the mass spectrometry peak intensity data corresponding to each concentration of the standard were obtained. By plotting the standard curve between the concentration ratio (Concentration Ratio) of the external standard to the internal standard and the area ratio (Area Ratio), the concentration of each substance was further analyzed. The peak area ratio detected in all samples was substituted into the linear equation of the standard curve, and the concentration calculation was performed using the calculation formula. The dilution factor was set to 3 in MultiQuant 3.0.3 software, and finally the concentration (ng / mL) of the substance in the sample was calculated, thereby obtaining the content data of the substance in the sample.

[0134] 2.6 Experimental quality control

[0135] By performing total ion current chromatogram overlay analysis on the mass spectrometry data of different quality control (QC) samples, the repeatability of metabolite extraction and detection, i.e., technical replicates, can be evaluated. The stability of the instrument is crucial for ensuring the reliability and repeatability of the data. The coefficient of variation (CV), which is the ratio of the standard deviation to the mean of the raw data, is used to reflect the discreteness of the data. The empirical cumulative distribution function (ECDF) can be used to analyze the frequency of CV values of substances below the reference value. If the proportion of substances with low CV values in the QC samples is high, it indicates that the experimental data is relatively stable. When the CV values of all QC samples are less than 0.3, it shows that the data stability is good; if the proportion of substances with CV values less than 0.2 in the QC samples exceeds 90%, it indicates that the data is very stable. In addition, monitoring the change of the CV value of the isotopic internal standard, if the change of the internal standard CV is less than 20%, it indicates that the instrument maintains good stability during the detection process.

[0136] 3 Plasma metabolite biomarker screening

[0137] 3.1 Data preprocessing

[0138] Perform data preprocessing on the qualitative and quantitative results obtained in step 2. According to the 80% rule, when a metabolite is detected in at least 4 / 5 of the samples in a certain group, the metabolite is considered detectable. According to this rule, 83 metabolites were robustly detected in all 928 samples. For the few metabolites that were not detected (<1 / 5 samples), they were filled with 1 / 2 of the lowest detection value in that group for subsequent statistical analysis.

[0139] 3.2 Screening and annotation of differential metabolites

[0140] Perform a differential significance analysis based on the Wilcoxon rank-sum test on the concentrations of 83 metabolites in 928 gastric cancer and non-gastric cancer patients. To control the false positive problem caused by multiple comparisons, the p-values were corrected by the false discovery rate (FDR). FDR correction is a method widely used in high-throughput data analysis, which can effectively control the false positive rate and ensure the reliability of the screening results. By setting the FDR threshold to 0.05, a total of 68 metabolites with statistically significant differences were screened out, and the fold change (FC) of each metabolite was further calculated to quantify the change degree of metabolite concentrations between the gastric cancer group and the non-gastric cancer group. The FDR-P values and FC results of all plasma differential metabolites are shown in Table 3.

[0141] Table 3

[0142]

[0143]

[0144]

[0145]

[0146] 3.3 Visual analysis of differential metabolites

[0147] Use the t-distributed Stochastic Neighbor Embedding (t-SNE) algorithm to perform a visual analysis of 68 differential metabolites in gastric cancer and non-gastric cancer populations. t-SNE is a non-linear dimensionality reduction technique that can effectively map high-dimensional data to a low-dimensional space while maintaining the local similarity between data points, thus realizing the visualization of data. As Figure 1As shown, the t-SNE algorithm was used to reduce the dimensionality of 68 differential metabolites, and the results were visualized in a two-dimensional space. The abscissa and ordinate in the figure represent the two main components of t-SNE (t-SNE component 1 and t-SNE component 2), respectively. The points of different colors represent gastric cancer patients and non-gastric cancer control groups. It can be observed that based on the content characteristics of 68 differential metabolites, gastric cancer patients and non-gastric cancer control groups formed a relatively obvious separation in the two-dimensional space after t-SNE dimensionality reduction ( Figure 1 ). This indicates that by comprehensively using the concentration data of these differential metabolites, it is possible to effectively distinguish between gastric cancer and non-gastric cancer populations.

[0148] 3.4 Enrichment analysis of gastric cancer-related metabolic pathways of differential metabolites

[0149] This example relates to a method for enrichment analysis of gastric cancer-related metabolic pathways based on 68 differential metabolites. The method includes using the KEGG database to perform enrichment analysis on the differential metabolites to identify their potential biological significance in gastric cancer. The analysis results revealed 25 significantly enriched metabolic pathways, including but not limited to key pathways such as "Biosynthesis of valine, leucine and isoleucine", "Caffeine metabolism", "Alanine, aspartate and glutamate metabolism", etc., and the results are as Figure 2 . The enrichment analysis determines the significance of the metabolic pathway by calculating -log10(P value), and evaluates the proportion of differential metabolites in a specific pathway through the enrichment rate. The enrichment analysis results show that the "Biosynthesis of valine, leucine and isoleucine" pathway has the highest significance, suggesting that it may play a key role in gastric cancer metabolic reprogramming. In addition, the significant enrichment of other metabolic pathways including "Citric acid cycle (TCA cycle)", "Ammonia metabolism" and "Glyoxylate and dicarboxylate metabolism" further confirms the importance of amino acid metabolism and energy metabolism in the development of gastric cancer.

[0150] 3.5 Feature extraction in the process of mining metabolic markers

[0151] A machine learning algorithm combining Recursive Feature Elimination (RFE) and Random Forest (RF) is used for feature extraction to screen out the most valuable metabolic markers for gastric cancer diagnosis from a large number of metabolites. RFE is a classic feature selection method that gradually removes the features with the least contribution to the model and finally selects the optimal feature subset. As an ensemble learning algorithm, RF has strong robustness to high-dimensional data, can effectively handle non-linear relationships, and provides feature importance scores to help identify key metabolites. In the specific implementation of RFE-RF, the rfe function in the caret package is used to implement the RFE process. The recursive process of RFE can be expressed as follows: ① Train a random forest model using all features. ② Calculate the importance score \(I_j\) of each feature. ③ Remove the feature with the lowest importance score. ④ Repeat steps 1-3 until the number of remaining features reaches the set target number of features. The random forest is selected as the base model, and the importance of features is evaluated through the rfFuncs function set. To ensure the accuracy and reliability of model evaluation, 10-fold cross-validation is used as the model evaluation method, and the parameters of cross-validation are set through the rfeControl function, specifically 10 repeated cross-validations. The sequence of the size of the feature subset is specified as: sizes = {1, 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30}. These values cover various combinations from a single feature to 30 features. Through fine-grained selection of the number of features, the optimal number of features can be found more precisely, avoiding missing the best feature combination due to too large a step size. The results and performance index diagrams of the RFE process are as Figure 3 shown, demonstrating the change of model accuracy with the number of features.

[0152] Through the RFE-RF process, the influence of different numbers of features on the model performance is systematically evaluated, and the optimal number of features is finally determined to be 25. This result indicates that a combination of 25 plasma metabolic biomarkers (Plasma Metabolic Biomarker - 25Metabolite Panel, PMB-P25) can provide the best classification performance while avoiding the overfitting problem. Through this rigorous feature extraction process, a high-quality feature set is provided for the subsequent construction of the diagnostic model, ensuring the reliability and effectiveness of the selected metabolic markers.

[0153] PMB-P25 contains the following 25 metabolites: nicotinamide, taurine, mono(2-ethylhexyl) phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (+ / -) 12-HETE, 5'-deoxy-5'-methylthioadenosine, traumatic acid, indole-3-carboxaldehyde, dihydro-D-sphingosine, lysophosphatidylethanolamine, hypoxanthine, 13(R)-HODE, DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, p-cresol sulfate ammonium salt, caffeine, inosine, and creatinine. The information of the selected metabolic markers is shown in Table 4.

[0154] Table 4

[0155]

[0156]

[0157] Example 2 Construction of Gastric Cancer Diagnosis Model

[0158] This example relates to a gastric cancer diagnosis model constructed based on the Gradient Boosting Machine (GBM) algorithm. The GBM algorithm effectively handles complex non-linear relationships by gradually optimizing the loss function and shows strong robustness to outliers and noise. During the model construction process, the following parameter settings were adopted: the number of decision trees (n.trees) was set to 150 to enhance the learning ability of the model; the maximum depth of each tree (interaction.depth) was set to 3 to avoid overfitting of the model; the learning rate (shrinkage) was set to 0.1 to control the contribution of each tree to the final model; the minimum number of samples required at the leaf nodes (n.minobsinnode) was set to 10 to ensure the generalization ability of the model to data.

[0159] To avoid overfitting of the model, 928 subjects were divided into an internal training set and an internal test set according to a ratio of 3:1 before modeling. The create Data Partition function was used for the division, and a fixed seed was set to ensure the repeatability of the results. By training the GBM model on the internal training set and validating it on the internal test set, the optimal cutoff values were determined respectively. The optimal cutoff values for the internal training set and the internal test set were 0.522 and 0.322 respectively. The cutoff values were obtained by evaluating the Receiver Operating Characteristic Curve (ROC) of the model, ensuring that the model had optimal diagnostic accuracy.

[0160] Verification of Gastric Cancer Diagnosis Model in Example 3 and Comparison of Diagnostic Efficacy with Traditional Tumor Markers

[0161] To verify the effect of the PMB-P25 combined with the GBM algorithm in the diagnosis of gastric cancer, in this example, 153 non-gastric cancer subjects and 156 gastric cancer patients' peripheral blood plasma from four research centers (Beijing Friendship Hospital, Cancer Hospital of the Chinese Academy of Medical Sciences, Hubei Provincial People's Hospital, Tongji Hospital Affiliated to Tongji Medical College of Huazhong University of Science and Technology) were collected at different time periods as an independent validation cohort. The detailed steps of plasma sample processing and targeted metabolomics detection were the same as those in Example 1. The targeted detection range of the independent validation cohort only covered 25 metabolic markers. At the same time, taking the traditional gastrointestinal tumor-related markers CEA (carcinoembryonic antigen), CA19-9 (carbohydrate antigen), and CA72-4 (carbohydrate antigen) as references, the diagnostic efficacy of the PMB-P25 combined with the GBM algorithm shown in Table 4 in the independent validation cohort for gastric cancer was verified. The area under the ROC curve (AUC) was 0.965, which had clinical diagnostic significance, and the diagnostic performance of the model was significantly higher than that of traditional tumor markers. The results are shown in Figure 4. In addition, the confusion matrix of applying the PMB-P25 diagnostic model in the independent validation cohort is as Figure 5 shown (the corresponding optimal cutoff value is 0.40). Table 5 details the comparison of various performance indicators when PMB-P25 and traditional tumor markers were used to establish the gastric cancer diagnosis model.

[0162] Table 5

[0163]

[0164] Verification of Gastric Cancer Plasma Differential Metabolites in Tissue Samples and Annotation of Core Interactive Metabolites in Example 4

[0165] In this example, cancer and adjacent tissue samples from 30 gastric cancer patients from the Cancer Hospital of the Chinese Academy of Medical Sciences were collected, and targeted metabolomics detection covering 68 plasma differential metabolites was performed, aiming to provide histological evidence directly reflecting the changes in metabolites in the tumor microenvironment, so as to provide deeper insights into the biological significance of plasma metabolic markers and the metabolic heterogeneity of gastric cancer. In this example, an optimized tissue sample collection method was adopted for metabolomics analysis. This method included negotiating with a pathologist to determine the sampling plan, ensuring the timely collection of samples, and controlling the cold ischemia time within 30 minutes. Cancer tissue and adjacent tissue were collected simultaneously during sample collection. The sample transfer process was maintained at 28°C to avoid contamination. In a biosafety cabinet, the samples were cut into small pieces 0.5 cm thick using a sterile blade, and the blade was replaced to avoid cross-contamination, and it was ensured that the fresh weight of the tissue was not less than 50 mg. The cut sample pieces were quickly frozen and stored to maintain the stability and quality of the samples, thus providing a reliable sample basis for metabolomics detection.

[0166] In the tissue sample validation stage, considering factors such as the validation objective setting, the biological intuitiveness of the fold change, and the uniqueness of the tissue samples, an absolute value of logFC greater than 0.263 was selected as the threshold to perform differential validation on the concentrations of 68 plasma differential metabolites in the tissue samples. Through this criterion, 24 tissue differential metabolites were screened out, including key substances such as dihydro-D-sphingosine, uridine, and 3-hydroxytetradecanoic acid. In order to identify the core interactive metabolites in the PMB-P25 diagnostic model - that is, those metabolites that showed differences in both plasma and tissue samples and were suitable as markers after feature extraction, in this embodiment, an intersection analysis was performed on the metabolites included in the PMB-P25 model and the newly screened 24 tissue differential metabolites, and finally 8 core interactive metabolites (PMB-P8) were determined. The concentration multiples of these metabolites in plasma and tissue are shown in Table 6. Using these core interactive metabolites, a model was built through the GBM model in an independent validation cohort (see Example 3). The AUC value of the obtained model was 0.908, the specificity was 0.882, and the sensitivity was 0.833, showing excellent diagnostic performance (see Figure 6 ). This finding not only further streamlined the list of potential biomarkers but also facilitated the future development of diagnostic kits, thus confirming the potential of these markers as a diagnostic tool for gastric cancer.

[0167] Table 6

[0168]

[0169] It is worth noting that when analyzing the concentration trends of these 8 core interactive metabolites in plasma and tissues, it was observed that 4 metabolites (nicotinamide, hypoxanthine, p-cresol sulfate ammonium salt, N-phenylacetyl-L-glutamine) showed consistent regulation in the plasma and tissues of gastric cancer patients, while the other 4 metabolites (dihydro-D-sphingosine, (+ / -)12-HETE, indole-3-carboxaldehyde, inosine) showed opposite regulation trends. The biological explanation for this phenomenon may involve the reprogramming of energy metabolism, nucleotide synthesis, and amino acid utilization by gastric cancer cells to adapt to rapid proliferation and cope with microenvironmental stress. The downregulated metabolites may represent substances that are accelerated in utilization or conversion to support tumor growth, while the upregulated metabolites may be "by-products" or response products of the activation of tumor-specific metabolic pathways, and their changes are reflected in both local tissues and the circulatory system. Conversely, those metabolites that are downregulated in plasma but upregulated in tissues may reflect the activation of specific metabolic pathways within the tumor, and the downregulation in plasma may be due to dilution, clearance in the blood circulation, and the results of the body's metabolic regulation mechanisms. This significant asymmetry between the local tumor microenvironment and the systemic metabolic state suggests that when using metabolites as biomarkers, the influence of sample source and local and systemic metabolic differences must be considered.

[0170] Example 5 Validation of a Gastric Cancer Diagnostic Model Containing Different Numbers of Metabolic Markers

[0171] In this example, dihydro-D-sphingosine and (+ / -)12-HETE, which have the largest absolute value of logFC, that is, the largest downregulation difference, were selected from the 25 plasma metabolic markers screened in Example 1 and modeled individually and in combination (modeled according to the method in Example 2) to verify the diagnostic efficacy for gastric cancer. The samples were selected from the independent validation cohort in Example 3.

[0172] The results are as Figures 7-9 shown, and the specificity and sensitivity data in the independent validation cohort are shown in Table 7.

[0173] Table 7

[0174]

[0175] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solutions of the present invention, and these simple modifications all fall within the protection scope of the present invention.

Claims

1. Use of a metabolic marker in the preparation of a product for diagnosing gastric cancer, characterized in that, The metabolic markers include dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE).

2. The application according to claim 1, wherein The metabolic markers include dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine; Preferably, the metabolic markers further include one or more of taurine, mono-2-ethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, 5'-deoxy-5'-methylthioadenosine, traumatic acid, lysophosphatidylethanolamine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, caffeine, or creatinine; More preferably, the metabolic markers include nicotinamide, taurine, mono-2-ethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), 5'-deoxy-5'-methylthioadenosine, traumatic acid, indole-3-carboxaldehyde, dihydro-D-sphingosine, lysophosphatidylethanolamine, hypoxanthine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, p-cresol sulfate ammonium salt, caffeine, inosine, and creatinine.

3. The application according to claim 1 or 2, characterized in that The metabolic markers are metabolic markers in body fluids and / or tissues; Preferably, the body fluid is selected from blood, plasma, and serum; Preferably, the tissue is gastrointestinal tissue, preferably cancer tissue or adjacent cancer tissue.

4. The application according to any one of claims 1-3, characterized in that The product includes a reagent for detecting metabolic markers. Preferably, the reagent detects the concentration of metabolic markers; Preferably, the methods used to detect metabolic markers include one or more of nuclear magnetic resonance spectroscopy, mass spectrometry, chromatography, high performance liquid chromatography-tandem mass spectrometry, Fourier transform ion cyclotron resonance, ion mobility spectrometry, electrochemical detection, Raman spectroscopy, or radioactive labeling; Preferably, the product includes a diagnostic model, a kit, a test strip, a chip, or a device.

5. A method for constructing a gastric cancer diagnosis model, characterized in that, The construction method includes: i) Collecting the concentration detection results of metabolic markers in a gastric cancer patient group and a non-gastric cancer control group; ii) Constructing a gastric cancer diagnostic model based on the information collected in step i); Alternatively, the construction method includes: I) Collecting samples from subjects and detecting the concentration of metabolic markers; II) Clinically diagnosing the subjects and dividing the subjects into a gastric cancer patient group and a non-gastric cancer control group; III) Constructing a gastric cancer diagnostic model based on the detection results in step I) and the diagnostic results in step II); The metabolic markers described above include dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE); Preferably, the metabolic markers include dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine; More preferably, the metabolic markers further include one or more of taurine, mono(2-ethylhexyl) phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, 5'-deoxy-5'-methylthioadenosine, traumatic acid, dihydro-D-sphingosine, lysophosphatidylethanolamine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, caffeine, or creatinine.

6. The construction method according to claim 5, wherein The sample is body fluid and / or tissue; Preferably, the body fluid is selected from blood, plasma, and serum; Preferably, the tissue is gastrointestinal tissue, preferably cancer tissue or adjacent cancer tissue.

7. The construction method according to any one of claims 5-6, characterized in that, The algorithm used for constructing the diagnostic model is a machine learning model, preferably including gradient boosting machine.

8. A metabolic marker for diagnosing gastric cancer, characterized in that, The metabolic markers described above include dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE); Preferably, the metabolic markers include dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine.

9. A computer-readable storage medium or a device comprising a computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the following operations are implemented: obtaining the concentration of metabolic markers in a sample to be tested, inputting the concentration of metabolic markers into the constructed machine learning model, outputting a probability value, and comparing the output probability value with a threshold value to determine whether the subject has gastric cancer. The metabolic markers include dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE); Preferably, the metabolic markers include dihydro-D-sphingosine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), indole-3-carboxaldehyde, inosine, nicotinamide, p-cresol sulfate ammonium salt, hypoxanthine, and N-phenylacetyl-L-glutamine.

10. A screening method for gastric cancer metabolic markers, characterized in that, The screening method includes: 1) Collect samples from a gastric cancer patient group and a non-gastric cancer control group to establish a gastric cancer-specific metabolite library; 2) Screen metabolites through differential ion pair screening and ion pair deconvolution screening and supplement common gastric cancer metabolites to obtain preliminarily screened metabolites; 3) Use a T3 column for liquid chromatography separation of metabolites and mass spectrometry for metabolite quantification; 4) Screen for differential metabolites based on the differential significance analysis of Wilcoxon rank sum test, wherein the p-value is corrected for false discovery rate; 5) Feature extraction is performed on the selected differential metabolites using a machine learning algorithm of recursive feature elimination combined with random forest (RFE-RF) to further determine a metabolite subset from the selected differential metabolites; Preferably, the initially screened metabolites include amino acids and their derivatives, organic acids and their derivatives, nucleosides, nucleotides and their derivatives, alkaloids and their derivatives, acylcarnitines, bile acids and their derivatives, amines, sugars and their derivatives, and other functional metabolites (such as including organic oxygen compounds, steroids, fatty acid derivatives); Preferably, the metabolite subset includes dihydro-D-sphingosine and / or (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), preferably nicotinamide, taurine, mono-2-ethylhexyl phthalate, sphingosine, 9,12,13-trihydroxy-octadecenoic acid, glycyl-phenylalanine, N-phenylacetyl-L-glutamine, (+ / -)12-hydroxyeicosatetraenoic acid ((+ / -)12-HETE), 5'-deoxy-5'-methylthioadenosine, traumatic acid, indole-3-carboxaldehyde, dihydro-D-sphingosine, lysophosphatidylethanolamine, hypoxanthine, 13(R)-hydroxyoctadecadienoic acid (13(R)-HODE), DL-3-phenyllactic acid, azelaic acid, piperine, 2,6-dihydroxybenzoic acid, L-lactic acid, salicylic acid, p-cresol sulfate ammonium salt, caffeine, inosine and creatinine; Preferably, the RFE-RF uses the rfe function in the caret package to implement the RFE process.

Citation Information

Cited By

  • Screening method of tumor metabolism markers

    CN120221049A

  • A screening method for tumor metabolic markers

    CN120221049B