Biomarker combination, application thereof and kit for gastric precancerous lesion detection
The gastric precancerous lesion detection model constructed by combining biomarkers and 5hmC high-throughput sequencing technology solves the problems of insufficient invasiveness and accuracy of existing diagnostic tools, and realizes non-invasive and accurate detection of gastric precancerous lesions.
Patent Information
- Application Number
- CN202511626340.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2025-12-12
AI Technical Summary
Existing technologies lack diagnostic tools for precancerous lesions of the stomach that can be non-invasive, convenient, and highly accurate. Traditional serological markers have unsatisfactory sensitivity and specificity, and gastroscopy has invasive and subjective issues.
A combination of biomarkers, including ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1, and NCOR2, was used in conjunction with 5hmC high-throughput sequencing technology and multiple machine learning algorithms to construct a diagnostic model. By detecting the 5-hydroxymethylcytosine modification level of extravesicular DNA in peripheral blood, a model for detecting precancerous lesions of the stomach was constructed.
It significantly improves the accuracy and objectivity of the diagnosis of precancerous lesions of the stomach, provides a non-invasive and accurate detection method, can accurately distinguish precancerous lesions of the stomach from healthy people, and avoids diagnostic bias caused by differences in physician experience.
Smart Images

Figure CN121109593A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of biological medicine, and particularly relates to a biomarker combination, application thereof, and a kit for detecting precancerous lesions of gastric cancer. BACKGROUND
[0002] Gastric cancer is one of the malignant tumors with high morbidity and mortality worldwide. Precancerous lesions of gastric cancer, such as chronic atrophic gastritis and intestinal metaplasia, are key pathological stages in the development of gastric cancer. Effective screening and intervention of the precancerous lesions are the core strategies to reduce the incidence of gastric cancer.
[0003] Currently, the main method for diagnosing precancerous lesions of gastric cancer (PLGC) in clinics is gastroscopy. However, this method is an invasive operation, which not only has high cost, but also may cause discomfort, bleeding, infection and other complications to patients. Moreover, the accuracy of the diagnostic results of this method depends on the experience of the operating physician to a certain extent, and is subjective.
[0004] In order to overcome the shortcomings of gastroscopy, researchers have tried to develop non-invasive diagnostic methods. However, the diagnostic sensitivity and specificity of traditional serological markers, such as pepsinogen, are not ideal, and it is difficult to meet the needs of precise screening.
[0005] In recent years, disease diagnosis based on epigenetic modifications of circulating nucleic acids in body fluids, such as 5-hydroxymethylcytosine modification, has shown great potential. The existing technology discloses a general technical framework for obtaining circulating deoxyribonucleic acid in a body fluid sample, detecting the 5-hydroxymethylcytosine modification level, and outputting a diagnostic result combined with a calculation model. However, for the specific disease of precancerous lesions of gastric cancer, the existing technology has not yet provided a specific combination of 5-hydroxymethylcytosine biomarkers with sufficient diagnostic performance, resulting in the lack of a diagnostic tool that can balance non-invasiveness, convenience and high accuracy. SUMMARY
[0006] Based on the above technical background, the main purpose of the present application is to provide a biomarker combination, application thereof, and a kit for detecting precancerous lesions of gastric cancer, in order to overcome the shortcomings in the prior art.
[0007] To achieve the aforementioned purposes, the technical solutions adopted by the present application comprise: The first aspect of the present application is to provide a biomarker combination, which comprises the following biomarkers: ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1 and NCOR2.
[0008] The second aspect of the present application provides use of the biomarker combination of the first aspect of the present application in the preparation of a kit for the detection of precancerous lesions of gastric cancer.
[0009] The third aspect of the present application provides a kit for the detection of precancerous lesions of gastric cancer, comprising detection reagents for detecting the presence or content, expression level of each biomarker in a biological sample, wherein the biomarker combination comprises the following biomarkers: ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1 and NCOR2.
[0010] Preferably, the kit is used for detecting the expression amount of the biomarkers in the biological sample.
[0011] Preferably, detecting the expression amount of the biomarkers in the biological sample comprises the following steps: Step 1, obtaining a biological sample of the individual to be tested; Step 2, obtaining exosome DNA from the biological sample; detecting the exosome DNA to obtain the level data of the biomarkers.
[0012] The above steps are described in detail as follows.
[0013] In step 1, the biological sample is a body fluid sample.
[0014] Preferably, the biological sample is a peripheral blood sample.
[0015] In step 2, the exosome DNA is obtained by the following steps: separating peripheral blood plasma and extracting plasma extracellular vesicle DNA.
[0016] The exosome DNA is subjected to 5hmC high-throughput sequencing to obtain the level data of each biomarker in the biomarker combination.
[0017] Preferably, the exosome DNA is subjected to 5hmC-Seal high-throughput sequencing to obtain the level data of the biomarker combination.
[0018] Preferably, the biomarker combination comprises the following biomarkers: ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1 and NCOR2. The biomarker parameters for diagnosing precancerous lesions of gastric cancer are shown in Table 1.
[0019] Table 1
[0020] inputting the level data of the biomarker into a trained gastric precancerous lesion detection model, and outputting a detection result of the individual having the gastric precancerous lesion from the gastric precancerous lesion detection model.
[0021] The input variable of the detection model is the content data of each biomarker in the biomarker combination.
[0022] The present application selects the combination of ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1 and NCOR2 as specific biomarkers for gastric precancerous lesions, and constructs a diagnostic model by combining 5hmC high-throughput sequencing technology and multi-machine learning algorithm, which significantly improves the accuracy and objectivity of gastric precancerous lesion diagnosis. The model shows high accuracy, high area under the curve (AUC), high sensitivity and high specificity in the training group and the validation group, with an accuracy of 1.000 in the training group and 0.886 in the validation group, which fully proves its excellent detection performance and stable generalization ability, and can effectively distinguish gastric precancerous lesions from healthy patients. At the same time, the present application creates a PLGC non-invasive detection model based on the 5hmC landscape in plasma evDNA. The biomarker combination consists of nine 5hmC markers, showing high sensitivity and specificity. The research results show that the 5hmC expression profile is one of the most promising tools for early detection and accurate diagnosis of PLGC.
[0023] The determination method of the biomarker content data is 5hmC-Seal high-throughput sequencing method.
[0024] The gastric precancerous lesion detection model is constructed by ensemble learning strategy, and the construction method further comprises: inputting the 5-hydroxymethylcytosine modification level data into a plurality of machine learning models respectively to generate respective prediction probability matrices; splicing the plurality of prediction probability matrices to form a new feature matrix; and inputting the new feature matrix into a meta-classifier model to output the final detection result from the meta-classifier model.
[0025] Preferably, the plurality of machine learning models comprises a neural network model, a random forest model and a stochastic gradient descent model; and the meta-classifier model is a logistic regression model.
[0026] More preferably, the neural network model adopts a structure comprising a plurality of convolutional layers; the random forest model is set to have 300 decision trees; and the stochastic gradient descent model uses a logarithmic loss function and adopts elastic net for regularization.
[0027] The detection model can significantly improve the accuracy and objectivity of gastric precancerous lesion diagnosis, and avoid the diagnosis bias caused by the experience difference of doctors.
[0028] Specifically, the method for constructing the gastric precancerous lesion detection model comprises the following steps: S1, collecting peripheral blood samples of gastric precancerous lesion patients and healthy people, separating plasma and extracting plasma extracellular vesicle DNA; S2, high-throughput sequencing of the plasma extracellular vesicle DNA to obtain sequencing data; S3, finding biomarkers through filtering and screening processes; S4, using the screened biomarkers as features, using machine learning algorithms to construct a gastric precancerous lesion diagnosis model.
[0029] In step S2, preferably, 5hmC high-throughput sequencing is performed on the plasma extracellular vesicle DNA to obtain 5hmC sequencing data.
[0030] In step S3, the specific steps for finding biomarkers through filtering and screening processes include: S31, removing 5hmC peak information that only appears in ≤10 samples; S32, using DEseq2 software to retain 5hmC peak regions with reads>50, and screening the differential markers according to FoldChange≥0.4 and p-value<0.05; S33, using Mfuzz function to find trend expression differential markers with disease progression; S34, selecting the biomarkers selected by any two of linear discriminant analysis, logistic regression and random forest algorithm through joint screening of linear discriminant analysis, logistic regression and random forest algorithm.
[0031] Preferably, the joint screening of linear discriminant analysis, logistic regression and random forest algorithm in step S34 specifically includes: (1) Linear Discriminant Analysis (LDA): lda = LinearDiscriminantAnalysis() lda.fit(features, target) coefficients = lda.coef_.ravel() top_gene_indices = np.abs(coefficients).argsort()[-cxl:][::-1] LDA_features = features.columns[top_gene_indices] (2) Logistic Regression: logreg_clf = LogisticRegression(solver='liblinear') logreg_clf.fit(features, target) logreg_coefficients = logreg_clf.coef_[0] threshold = np.sort(np.abs(logreg_coefficients))[-cxl] log_features = features.columns[np.abs(logreg_coefficients)>=threshold].tolist() (3) Random Forest: rf_clf = RandomForestClassifier(n_estimators=100, random_state=42) rf_clf.fit(features, target) importances = rf_clf.feature_importances_ sorted_feature_importance = pd.DataFrame({ 'Feature': features.columns, 'Importance': importances }).sort_values(by='Importance', ascending=False) rf_features = sorted_feature_importance.head(cxl)['Feature'].tolist() Integrate features from three methods, and retain only features selected by at least two methods.
[0032] combined_features = set(rfe_features).union (set (LDA_features)).union (set(log_features)) final_features = [feat for feat in combined_features if sum([feat inlst for lst in [rfe_features, LDA_features, log_features]])>= 2] The biomarker obtained by screening is a biomarker combination of ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1 and NCOR2.
[0033] The step S4 specifically comprises: S41, input the biomarker data obtained by screening into a neural network model, a random forest model and a stochastic gradient descent model respectively, and generate corresponding probability matrices; S42, splice the probability matrices to form a feature matrix, and input the feature matrix into a meta-classifier to judge the precancerous lesion of gastric cancer according to the output probability, wherein the meta-classifier is a logistic regression model, and an L-BFGS optimization algorithm is adopted.
[0034] More preferably, the step S41 specifically comprises: The neural network model adopts a convolutional layer structure of 256x128x64x32x16x8x4x1, the activation function is ReLu, the loss function is binary cross entropy, and the training epoch is 200; The random forest model is set to have 300 decision trees; The stochastic gradient descent uses a logarithmic loss function, the regularization method is elastic net, and the L1 and L2 regularization ratio is 0.5.
[0035] The third aspect is to provide a kit for detecting precancerous lesions of gastric cancer, the kit comprising: reagents for detecting the presence or content / expression level of a biomarker combination in a biological sample.
[0036] The biomarker combination comprises the following biomarkers: ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1 and NCOR2.
[0037] The present application has the following beneficial effects: (1) The detection result accuracy is high: the application determines, for the first time, a nine-gene biomarker combination of ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1 and NCOR2 as a specific 5-hydroxymethylcytosine marker for diagnosing precancerous lesions of gastric cancer through systematic bioinformatics screening. The detection kit and detection model constructed based on the specific biomarker combination have extremely high accuracy, sensitivity and specificity, can accurately distinguish between precancerous lesion patients and healthy people, and solve the problem of insufficient accuracy of existing non-invasive markers.
[0038] (2) The detection process is non-invasive and safe: the application only needs to obtain body fluid samples such as peripheral blood of the patient, which belongs to minimally invasive detection. Compared with the invasive gastroscopy, the application greatly improves the acceptance and compliance of the patient and reduces the discomfort and potential risk in the screening process.
[0039] (3) The detection result is objective and reliable: the final detection result of the precancerous lesion of gastric cancer is output by the data-driven detection model, which avoids the diagnostic bias caused by the difference in subjective experience of doctors, so that the detection result is more objective, standardized and repeatable.
[0040] (4) The detection method has wide application prospect: the kit detection method based on the biomarker combination provided by the application has clear operation process and easy-to-obtain detection samples, is suitable for large-scale population health examination and early screening, and is helpful to promote early detection, early diagnosis and early treatment of precancerous lesions of gastric cancer, and has high detection accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 Fig. 2 shows the ROC curve of the diagnostic model in the training group and the validation group of the precancerous lesion diagnosis model of the application; Figure 2 Fig. 3 shows the confusion matrix and related accuracy performance parameters of the training group of the precancerous lesion diagnosis model of the application; Figure 3 Fig. 4 shows the confusion matrix and related accuracy performance parameters of the test group of the precancerous lesion diagnosis model of the application. DETAILED DESCRIPTION
[0042] The application will be described in detail below, and the features and advantages of the application will become clearer and more explicit with these descriptions.
[0043] The 5hmC-Seal high-throughput sequencing method used in the precancerous lesion detection model construction method of the application is explained as follows: 5hmC-Seal is a high-throughput sequencing method based on 5hmC. This method is based on improved chemical glycosylation labeling combined with second-generation high-throughput sequencing technology to obtain the distribution information of 5hmC on genomic DNA.
[0044] Due to the high sensitivity of the chemical labeling method, the input DNA can be as low as 1-10 ng, which can be fragmented genomic DNA, cfDNA and other small fragments. According to the requirements of second-generation sequencing, the DNA fragments are polished at both ends, and then an A tail is connected at the 3' end. By AT-specific connection, a sequencing Y-shaped adapter is connected at both ends of each DNA fragment, which contains index information and amplification primer sequences to distinguish samples. Then, the 5hmC labeling step is carried out. First, UDP-6-N3-Glc is added, and under certain conditions, 5hmC on the DNA is all reacted to N3-5ghmC. Then, DBCO-PEG4-Biotin is added to make N3-5ghmC all connected to biotin. Finally, through the specific binding of biotin-magnetic beads, all DNA fragments containing 5hmC sites are screened. After PCR amplification and purification, the 5hmC-based DNA library construction is completed. Through Fragment Analyzer quality control analysis of each sample DNA band size and distribution, and after accurate quantification of the library by qPCR, high-throughput sequencing is carried out by Illumina Nextseq 500 sequencer to obtain the base sequence of all DNA fragments in the library.
[0045] Embodiments The present application is further illustrated by the following specific examples, which are only intended to illustrate the present application and not to limit the scope of the present application.
[0046] Embodiment 1 This embodiment provides a method for constructing a gastric precancerous lesion diagnosis model. Peripheral blood samples of 67 gastric precancerous lesion patients and 67 healthy people are collected. The peripheral blood samples are randomly divided into a training group and a validation group according to a ratio of 2:1.
[0047] Diagnostic criteria: (1) Age between 20 and 70 years old, regardless of gender; (2) In accordance with the diagnostic criteria of the "Guidelines for the Diagnosis and Treatment of Chronic Gastritis in China" (2022, Shanghai), gastric precancerous lesions were evaluated by at least two relevant specialists through endoscopy and pathological evaluation; (3) Informed consent, voluntary participation in the study, and signing of the consent form. Exclusion criteria: (1) Patients with severe liver and kidney function impairment, blood system diseases, autoimmune diseases, endocrine disorders, or other serious primary diseases that affect life expectancy; (2) Patients with malignant tumors, acute infections, or other major diseases; (3) Pregnant women, women who have had abortions or are breastfeeding; (4) Patients who cannot or do not wish to cooperate with the collection of relevant information due to illness or other reasons.
[0048] All sample collections were carried out with the informed consent of the patients. Subsequently, the model was constructed according to the following steps: S1, Collect 3-4 mL of fasting peripheral blood from gastric precancerous lesion patients and healthy people in the morning, and extract DNA from extracellular vesicles in plasma; S2, Take extracellular vesicle DNA.
[0049] S21, Obtain whole blood samples by conventional venipuncture and collect them into cell-free DNA collection tubes (Roche). Store the tubes at ambient temperature of 15°C to 25°C.
[0050] S22, Within 24 hours, centrifuge the whole blood at 4°C for 15 minutes at 1350 x g, then at 4°C for 5 minutes at 13,500 x g, and then store at -80°C for future use. Extracellular vehicle (EV) extraction is performed according to the exosome extraction kit (H-Wayen), including incubation with reagents and extraction by centrifugation at 10,000 x g and 3,000 x g.
[0051] S23, Extraction and purification of extracellular vesicle DNA is performed according to the rapid DNA extraction kit (ZYMO). The process includes the addition of BioFluid&Cell buffer, proteinase K, genomic binding buffer, g-DNA wash buffer, and incubation and centrifugation at 12,000 x g for extraction and purification. The extracted extracellular vesicle DNA should be stored at -20°C.
[0052] S24, 5hmC library is constructed by using the efficient hmC-Seal technology, according to the kit's protocol, using KAPA Hyper Prep kit (KAPA Biosystems) for end repair and 3'-adenylation of evDNA extracted from plasma, followed by Illumina-compatible adapter ligation.
[0053] S25, Glycosylation of ligated evDNA by incubation in a solution containing HEPES buffer (pH 8.0), MgCl2, UDP-6-N3-Glc and b-glucosyltransferase (NEB). Subsequently, DBCO-PEG4-biotin (click chemistry tool) was added and incubated again.
[0054] S26, Purification of DNA using DNA Clean & Concentrator kit (ZYMO). Purified DNA was incubated with streptavidin beads (Life Technologies) in a buffer containing Tris pH 7.5, EDTA, NaCl and Tween 20, followed by washing. All bundling and washing steps were performed at room temperature with gentle rotation.
[0055] S27, Beads were then resuspended in RNase-free water and subjected to 14-16 cycles of PCR amplification. PCR products were purified using AMPure XP magnetic beads (Beckman).
[0056] S28, Library concentration was measured using Qubit 3.0 fluorometer (Life Technologies). Libraries that met the quantification criteria were then subjected to paired-end 150 bp high-throughput sequencing on the NovaSeq6000 platform.
[0057] S3, Finding biomarkers by filtering and screening process; S31, Removing 5hmC peak information that only appears in <10 samples; S32, Using DEseq2 software to retain 5hmC peak regions with reads number >50, screening the differential biomarkers with FoldChange >0.4 and p-value <0.05; S33, Using Mfuzz function to find trend expression differential markers with disease progression; S34, Selecting biomarkers that are selected by any two of the algorithms by joint screening of linear discriminant analysis, logistic regression and random forest algorithm.
[0058] Linear discriminant analysis (LDA): lda = LinearDiscriminantAnalysis() lda.fit(features, target) coefficients = lda.coef_.ravel() top_gene_indices = np.abs(coefficients).argsort()[-cxl:][::-1] LDA_features = features.columns[top_gene_indices] Logistic Regression: logreg_clf = LogisticRegression(solver='liblinear') logreg_clf.fit(features, target) logreg_coefficients = logreg_clf.coef_[0] threshold = np.sort(np.abs(logreg_coefficients))[-cxl] log_features = features.columns[np.abs(logreg_coefficients)>=threshold].tolist() Random Forest: rf_clf = RandomForestClassifier(n_estimators=100, random_state=42) rf_clf.fit(features, target) importances = rf_clf.feature_importances_ sorted_feature_importance = pd.DataFrame({ 'Feature': features.columns, 'Importance': importances }).sort_values(by='Importance', ascending=False) rf_features = sorted_feature_importance.head(cxl)['Feature'].tolist() # Integrate features from three methods and retain only biomarkers selected by at least two algorithms.
[0059] combined_features = set(rfe_features).union (set (LDA_features)).union (set(log_features)) final_features = [feat for feat in combined_features if sum([feat inlst for lst in [rfe_features, LDA_features, log_features]])>= 2] Select the first 30 difference markers selected by at least two of the above three algorithms, and finally determine the biomarker combination as: ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1 and NCOR2.
[0060] S4, using the screened biomarkers as features, using machine learning algorithms to construct a gastric precancerous lesion diagnosis model; S41, input the data of the screened 9 biomarkers into the following model respectively, and generate a probability matrix: Among them, the neural network adopts the convolution layer structure of 256*128*64*32*16*8*4*1, the activation function is ReLu, the loss function is binary cross entropy, and the training epoch=200; The random forest sets the number of decision trees to 300; The random gradient descent uses a logarithmic loss function and elastic net regularization (L1:L2=0.5).
[0061] S42, the probability matrix of the above three models is spliced into a feature matrix, which is input into a unary classifier (logistic regression model, using L-BFGS optimization algorithm), and gastric precancerous lesions are judged according to the output probability (probability> 0.5 is determined as gastric precancerous lesions).
[0062] The scheme provided by the application, the gastric precancerous lesion detection model constructed by the above method, the model shows a high accuracy of 1.000 (95% CI: 0.970-1.000) in the training queue, the AUC value is 1.000, the accuracy is 1.000, the sensitivity is 1.000, and the specificity is 1.000, see Figure 1 、 Figure 2The model was verified by the verification group, and the accuracy of the verification set was 0.886 (95% CI: 0.784-0.999), the AUC value was 0.963, the accuracy was 0.886, the sensitivity was 0.955, the specificity was 0.818, and a total of 4 samples were predicted incorrectly, such as Figure 3 .
[0063] The above results show that the diagnostic model based on 5hmC can withstand external queue verification, indicating that the model has certain reliability, and the diagnostic model can be used to diagnose gastric precancerous lesion patients.
[0064] The above results show that the diagnostic model constructed by the above 9 biomarkers can significantly distinguish gastric precancerous lesion patients, and can significantly improve the accuracy and objectivity of gastric precancerous lesion diagnosis. The present application realizes precise diagnosis of gastric precancerous lesions based on 5hmC markers of plasma extracellular vesicle DNA, and provides a molecular level basis, which can realize rapid detection of batch samples.
[0065] Example 2 The embodiment provides a gastric precancerous lesion detection model, which is constructed by the construction method in Example 1.
[0066] The input variable of the model is the content data of each biomarker in the biomarker combination.
[0067] The determination method of the biomarker content data is 5hmC-Seal high-throughput sequencing method. Example 3 The embodiment provides a gastric precancerous lesion detection method based on epigenetic markers, which comprises the following steps: Step 1, obtaining a biological sample of a to-be-detected individual; the biological sample is a peripheral blood sample, Step 2, obtaining extracellular vesicle DNA from the biological sample; the extracellular vesicle DNA is obtained by the S2 step in Example 1.
[0068] The extracellular vesicle DNA is detected by 5hmC-Seal high-throughput sequencing to obtain the level data of the biomarker combination, and the biomarker combination comprises ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1 and NCOR2. Step 3, inputting the data of the biomarker combination into the trained detection model (the detection model described in Example 2), and outputting the detection result of the to-be-detected individual suffering from gastric precancerous lesion by the detection model.
[0069] The present application is described in detail above with reference to specific embodiments and exemplary examples, but these are not to be understood as limiting the present application. It is understood by a person skilled in the art that various equivalent substitutions, modifications or improvements can be made to the technical solutions of the present application and the embodiments thereof without departing from the spirit and scope of the present application, and these all fall within the scope of the present application. The scope of protection of the present application is defined by the appended claims.
Claims
1. A biomarker combination, characterized in that, The biomarker combination comprises the following biomarkers: ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1, and NCOR2.
2. Use of the biomarker combination of claim 1 in the preparation of a kit for the detection of precancerous lesions of gastric cancer.
3. A kit for detection of precancerous lesions of gastric cancer, characterized by, The kit comprises detection reagents for detecting the presence or content, expression level of each biomarker in the biomarker combination in a biological sample, wherein the biomarker combination comprises the following biomarkers: ARHGEF16, MEGF6, CASP9, EPB41, TMEM39B, SGIP1, GNG12-AS1, WIPF1, and NCOR2.
4. The kit of claim 3, wherein The kit is used for detecting the expression amount of biomarkers in a biological sample.
5. The kit of claim 4, wherein The biological sample is a body fluid sample.
6. The kit of claim 4, wherein Detecting the expression amount of biomarkers in a biological sample comprises the following steps: Step 1: obtaining a biological sample of the individual to be tested; Step 2: obtaining exovesicle DNA from the biological sample; detecting the exovesicle DNA to obtain the level data of biomarkers.
7. The kit of claim 6, wherein In step 2, The 5hmC high-throughput sequencing is used for the exovesicle DNA to obtain the level data of the biomarker combination.
8. The kit of claim 6, wherein The level data of the biomarkers is input into the trained detection model, and the detection result of the individual to be tested suffering from precancerous lesions of gastric cancer is output by the detection model.
9. The kit of claim 8, wherein The input variable of the detection model is the content data of each biomarker in the biomarker combination.