Detection marker for squamous cell carcinoma and application of detection marker in diagnostic model and reagent
Through large sample size collection and wide target metabolomics technology, a high sensitivity and high specificity HNSCC diagnostic model was constructed, which solved the problems of small sample size and limited metabolites coverage in the existing technology, and achieved early diagnosis with high accuracy.
Patent Information
- Application Number
- CN202510509875.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-15
AI Technical Summary
In the prior art, early diagnosis methods for squamous cell carcinoma of head and neck have problems with small sample size and limited metabolites coverage, resulting in poor model stability and low diagnostic accuracy.
By systematically collecting metabolite data in patients' serum based on large sample sizes, combining broad-target metabolomics technology, key metabolites combinations were screened out, and diagnostic models were constructed using the LC-MS/MS platform and LASSO feature selection algorithm, including metabolites such as cystine, L-glutamic acid, hypoxanthine, para-aminobenzoic acid, acetylcysteine, choline and L-arginine, and a high sensitivity and high specificity were established.
It improves the accuracy of early diagnosis and model stability of squamous cell carcinoma, can detect characteristic metabolic changes in HNSCC early, and provides strong early diagnosis support.
Smart Images

Figure CN120432011A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a new use of a known compound, in particular to a detection marker for squamous cell carcinoma screening, and the use of the marker in diagnostic models and the preparation of diagnostic kits. Background Art
[0002] Head and neck squamous cell carcinoma (HNSCC) is one of the most common malignant tumors worldwide, commonly occurring in the oral cavity, pharynx, and larynx. Due to its insidious early symptoms, most patients are already in the advanced stage of the disease at the time of diagnosis, severely impacting treatment efficacy and survival prognosis. Therefore, developing a sensitive, specific, and non-invasive early diagnosis model is a critical issue that needs to be addressed in the current clinical diagnosis and treatment of HNSCC.
[0003] Currently, early diagnosis of HNSCC relies primarily on imaging, tissue biopsy, and the detection of some tumor markers. However, these methods all have limitations. For example, imaging has poor ability to identify small lesions, tissue biopsy is invasive and unsuitable for widespread screening, and commonly used molecular markers (such as EGFR and p53) have low sensitivity and specificity, making them difficult to meet the high clinical standards for early diagnosis.
[0004] With the development of metabolomics technology, researchers have begun to try to assist in disease diagnosis by detecting changes in metabolites in cancer patients. Metabolomics is a high-throughput technology that can systematically analyze changes in small molecule metabolites in the body, and can reflect the disease status from a holistic level. This technology is mainly divided into three categories: non-targeted metabolomics, targeted metabolomics, and broad-target metabolomics. Among them, non-targeted metabolomics is suitable for discovering novel metabolites, but its qualitative and quantitative accuracy is poor; targeted metabolomics has high quantitative accuracy, but is limited to the preset metabolite range and lacks flexibility. In contrast, broad-target metabolomics combines the advantages of both, not only covering a wider range of metabolite types, but also having good quantitative capabilities, and is particularly suitable for screening and model building of complex disease-related metabolite combinations.
[0005] Previous studies have shown that HNSCC patients have a series of characteristic metabolite changes, and some metabolite combinations have been attempted to be used in the establishment of diagnostic models. However, most of these studies are based on small clinical data samples, and the metabolic detection technologies used are mainly targeted or untargeted metabolomics, with limited metabolite coverage. The clinical stability and scalability of the models remain challenges. Summary of the Invention
[0006] One objective of the present invention is to provide a squamous cell carcinoma detection marker and to establish a diagnostic model closely related to the metabolism of head and neck squamous cell carcinoma (HNSCC) based on broad-target metabolomics to address the following technical defects in the prior art: small sample size leading to poor model stability, limited metabolite coverage, and other technical defects.
[0007] Another object of the present invention is to provide an early screening model for squamous cell carcinoma to improve diagnostic accuracy and stability.
[0008] Another object of the present invention is to provide a squamous cell carcinoma metabolite detection reagent for early screening, providing strong support and basis for early diagnosis and treatment of the disease.
[0009] Application of an organism metabolic marker in constructing a diagnostic model for squamous cell carcinoma.
[0010] Application of another biological metabolic marker in constructing a diagnostic model for head and neck squamous cell carcinoma.
[0011] The present invention systematically collects metabolite data in patient serum based on a large sample size, and combines it with broad-target metabolomics technology to screen out key metabolite combinations with statistical significance and disease relevance from hundreds of metabolites, thereby establishing a highly sensitive and specific HNSCC metabolic diagnostic model suitable for early clinical screening and auxiliary diagnosis.
[0012] The present invention uses an LC-MS / MS platform for metabolite analysis, combined with the LASSO feature selection algorithm, and a computer system to screen key metabolites from patients diagnosed with head and neck squamous cell carcinoma. These metabolites, such as cystine, L-glutamate, hypoxanthine, para-aminobenzoic acid, acetylcysteine, choline, glycerophosphocholine, and L-arginine, are closely associated with the development and progression of HNSCC. Using these metabolites, the model's diagnostic accuracy is significantly improved, enabling earlier detection of characteristic metabolic changes in HNSCC, thus providing strong support for early diagnosis.
[0013] Another application of an organism metabolic marker in constructing a diagnostic model for (head and neck) squamous cell carcinoma, wherein the metabolic marker comprises p-aminobenzoic acid and at least one of the following substances:
[0014] Group 1: one or more of hypoxanthine and choline;
[0015] Group 2: one or more of glycerophosphocholine, L-glutamic acid and L-arginine;
[0016] Group III: cystine; and
[0017] Group 4: Acetylcysteine.
[0018] A system for diagnosing early-stage squamous cell carcinoma, comprising the following units:
[0019] A detection unit, comprising a biological metabolic marker detection module;
[0020] an analyzing unit, which uses the result of the marker in the sample detected by the detecting unit as an input item for analysis; and
[0021] An evaluation unit that outputs the risk of squamous cell carcinoma for the individual corresponding to the sample.
[0022] According to the classification threshold of the model, judgment is made. If the sample marker is above the classification threshold, it is judged that the risk of the sample suffering from squamous cell carcinoma is high. If the sample marker is near the classification threshold, it is judged that the risk of the sample suffering from squamous cell carcinoma needs to be further observed. If the sample marker is below the classification threshold, the risk of the sample suffering from squamous cell carcinoma is low.
[0023] In this study, the classification cutoff value (cut-off value) was set to 0.5. This fixed value, determined by a machine learning algorithm during the diagnostic model training phase, represents the demarcation point between high-risk and low-risk HNSCC. Once the model is built, this cutoff value remains consistent across all subsequent sample tests, eliminating the need for regeneration each time, ensuring consistent test results and model stability.
[0024] Specifically, after obtaining the metabolic marker detection information in the sample serum, the detection information is input into the diagnostic model to generate a comprehensive score (risk score), which is compared with a fixed classification threshold (0.5) to achieve a classification judgment of the HNSCC risk of the sample. In order to improve the discriminative power and clinical interpretability of the model, the concept of "threshold deviation (deviation from cut-off)" is further introduced to describe the degree of deviation of the sample score relative to the classification threshold. The evaluation rules are as follows:
[0025] When the score is higher than the classification threshold (0.5) + 0.059, it is judged as high risk, indicating that the sample has a high risk of HNSCC;
[0026] When the score is within the classification threshold (0.5) ± 0.059, it is judged to be in the critical risk zone and requires further follow-up or the combination of other clinical information to assist in the judgment;
[0027] When the score is lower than the classification threshold (0.5)-0.059, it is judged as low risk and the sample has a low risk of HNSCC.
[0028] Among them, 0.059 is the tolerance threshold range set in the model training stage (determined based on the standard deviation of the HNSCC score distribution in the training dataset), which can be used as the confidence interval setting for model stability judgment.
[0029] The detection information of the biological metabolic markers detected by the detection unit is obtained from the information obtained by using the kit to detect the biological sample. The metabolic markers detected by the kit include p-aminobenzoic acid and at least one of the following substances:
[0030] Group 1: one or more of hypoxanthine and choline;
[0031] Group 2: one or more of glycerophosphocholine, L-glutamic acid and L-arginine;
[0032] Group III: cystine; and
[0033] Group 4: Acetylcysteine.
[0034] The kit of the present invention comprises reagents and standards required for detecting at least one of cystine, L-glutamic acid, hypoxanthine, p-aminobenzoic acid, acetylcysteine, choline, glycerophosphocholine and L-arginine.
[0035] The diagnostic model provided by the present invention, based on high-throughput, broad-target metabolomics data and combined with statistical and machine learning algorithms, has been validated to demonstrate exceptionally high accuracy. The AUC value for the training set was 1, while the AUC values for test sets 1 and 2 were 0.99 and 0.9958, respectively, demonstrating the model's consistency and high accuracy across diverse datasets. An AUC value close to 1 indicates that the model can accurately distinguish HNSCC patients from healthy controls, demonstrating higher diagnostic accuracy than traditional methods. The selected metabolites are closely associated with the disease, resulting in high diagnostic accuracy and robustness.
[0036] The model provided by this invention uses a rigorous quality control process to ensure the reliability of metabolomics data. By inserting QC samples every 10 samples to monitor instrument stability, and applying batch correction and biomass normalization to eliminate batch effects and biomass variability, this process significantly improves data stability and accuracy, providing assurance for subsequent analysis and model building.
[0037] The diagnostic model provided by the present invention has a standardized process for establishing a model, obtains metabolic markers that are correlated with diseases, improves accurate diagnostic capabilities, and has the potential to be widely used in early screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is an overall flow chart of the method for constructing an early diagnosis model for head and neck squamous cell carcinoma involved in the present invention;
[0039] Figure 2 Schematic diagram of the process for diagnostic model construction and validation;
[0040] Figure 3 Figure 2 shows the performance evaluation results of the early diagnosis model for head and neck squamous cell carcinoma (HNSCC); A is the ROC curve of the training dataset (Train_data) and test dataset 1 (Test_data1); B is the ROC curve of test dataset 2 (Test_data2); and C is a bar chart showing the ranking of metabolite feature importance.
[0041] Figure 4 Figure 2 is a comparison of the abundance of eight metabolites in the serum of patients at various stages of HNSCC and healthy controls; A is the relative abundance (Z-score) of cystine between various stages of HNSCC and healthy controls (N), B is the relative abundance (Z-score) of L-glutamic acid between various stages of HNSCC and healthy controls (N), C is the relative abundance (Z-score) of hypoxanthine between various stages of HNSCC and healthy controls (N), and D is the relative abundance (Z-score) of p-aminobenzoic acid. Figure 1 is the relative abundance (Z-score) of acetylcysteine in each stage of HNSCC and healthy controls (N), E is the relative abundance (Z-score) of acetylcysteine in each stage of HNSCC and healthy controls (N), F is the relative abundance (Z-score) of choline in each stage of HNSCC and healthy controls (N), G is the relative abundance (Z-score) of glycerophosphocholine in each stage of HNSCC and healthy controls (N), and H is the relative abundance (Z-score) of L-arginine in each stage of HNSCC and healthy controls (N);
[0042] Figure 5 The following are the prediction results of the HNSCC early diagnosis model; A is the prediction result of the training data, B is the prediction result of the test data set 1, and C is the prediction result of the test data set 2; the red and blue scatter points represent HNSCC patients and healthy controls, respectively, and the yellow horizontal line is the classification threshold predicted by the model. The areas above the horizontal line should be patients, and the areas below the horizontal line should be healthy people. DETAILED DESCRIPTION
[0043] The technical solution of the present invention is described in detail below with reference to the accompanying drawings. The embodiments of the present invention are intended only to illustrate the technical solution of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solution of the invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solution of the present invention, and all such modifications or equivalents should be included in the scope of the claims of the present invention.
[0044] Example 1 Construction of an early diagnosis model for HNSCC based on broad-target metabolomics
[0045] The steps of this embodiment are as follows: Figure 1 As shown, including:
[0046] Step 1: Sample collection
[0047] A total of 293 pathologically confirmed head and neck squamous cell carcinoma (HNSCC) patients and 282 healthy controls (HC) were included in this example.
[0048] Cohort 1 was from the Shanghai Ninth People's Hospital affiliated to Shanghai Jiao Tong University School of Medicine / Shanghai Oral and Maxillofacial Tumor Tissue Sample and Bioinformatics Database Professional Technical Service Platform, consisting of 498 subjects, including 248 HNSCC patients and 250 healthy controls.
[0049] Cohort 2 also originated from the Shanghai Ninth People's Hospital affiliated to Shanghai Jiao Tong University School of Medicine / Shanghai Oral and Maxillofacial Tumor Tissue Sample and Bioinformatics Database Professional Technical Service Platform, and included 77 subjects, including 45 HNSCC patients and 32 healthy controls.
[0050] Step 2: Serum sample collection
[0051] All subjects underwent fasting venous blood collection after an overnight fast using inert separation gel coagulation tubes. The samples were centrifuged at 3000 g for 10 minutes at 4°C, and the supernatant (serum) was collected and stored at -80°C.
[0052] Step 3: Metabolite Extraction
[0053] 40 μL of serum from each sample was mixed with 280 μL of ice-cold methanol:acetonitrile (1:1, v:v), vortexed for 1 minute, and centrifuged at 13,000 rpm for 15 minutes at 4°C. The supernatant was collected and placed in an injection vial. For quality control (QC) samples, 5 μL of plasma from each sample was processed identically to the serum samples and subsequently injected for analysis.
[0054] Step 4: Targeted metabolomics analysis
[0055] Samples were analyzed using an LC-MS / MS platform. Samples were randomly injected (1 μL injection volume, stored at 4°C) and separated on an ACQUITY UPLC HSS T3 column (2.1 × 100 mm, 1.8 μm). The mobile phase consisted of 0.01% formic acid in water (phase A) and acetonitrile (phase B). The gradient elution program was 100% phase A from 0 to 2 minutes, followed by 100% to 5% phase A until 14 minutes, a 2-minute hold at 5% phase A, and finally a return to 100% phase A over 4 minutes, for a total of 20 minutes. The flow rate was 0.2 mL / min and the column temperature was 30°C. Mass spectrometry detection was performed on an AB QTRAP 4500 system in scheduled MRM mode to monitor 420 water-soluble metabolites. The data were processed by SCIEX OS1.6 software, and 185 stable metabolites were screened out by applying the 80% rule (detected in ≥80% of samples in each group). A small number of missing values were filled with the baseline value of 1000 to ensure the reliability of subsequent analysis.
[0056] Step 5: Data correction and preprocessing
[0057] This example establishes a strict quality control (QC) and two-stage normalization strategy to ensure the reliability of metabolomics data. QC samples are inserted into the analysis batch every 10 test samples to monitor instrument stability. Normalization is divided into two stages, namely
[0058] Batch correction: Calculate metabolite-specific correction factors (QCall / QCadj) based on the global mean (QCall) and the neighboring local mean (QCadj) of QC samples to dynamically eliminate systematic bias and batch effects.
[0059] Biomass normalization: The corrected metabolite abundance is converted into a relative ratio of total ion current (TIC) to eliminate the interference of biomass differences.
[0060] Step 6: Construction and validation of diagnostic models
[0061] Based on LASSO feature selection and machine learning algorithm (see Figure 2 ) and constructed a diagnostic prediction model. The study cohort (cohort 1, n = 498) was divided into a training set (n = 348) and a test set (n = 150) through random stratified sampling. The diagnostic model constructed from the training set was validated in the test set and cohort 2 samples. The final diagnostic model was composed of eight metabolites: choline, p-aminobenzoic acid, L-glutamic acid, glycophosphocholine, hypoxanthine, cystine, L-arginine, and acetylcysteine.
[0062] Figure 3 Figure A shows the ROC curve for cohort 2. The AUC value for the training dataset is 1, indicating perfect predictive ability of the model on the training dataset. The AUC value for test dataset 1 (Test_data1) is 0.99, demonstrating excellent performance of the model on an independent test dataset. Figure 3 B shows the ROC curve of test data set 2 (Test_data2), with an AUC value of 0.9958, further verifying the consistency and high accuracy of the model on different data sets. Figure 3 C shows the important metabolites associated with HNSCC diagnosis and their importance scores.
[0063] Figure 4 Figure 2 shows the expression differences of the eight key metabolites screened between HNSCC patients and healthy controls. The p-value was 2.455e-16, indicating that the abundance of cystine was significantly higher in HNSCC patients than in healthy controls. The p-value was 1.177e-17, indicating that the abundance of l-glutamic acid was significantly different between HNSCC stage groups and healthy controls. The p-value was 9.835e-27, indicating that the abundance of hypoxanthine showed significant changes between early-stage and late-stage HNSCC patients. The p-value was 5.886e-27, indicating that the abundance of p-aminobenzoic acid was significantly different between HNSCC patients and healthy controls. The p-value was 6.286e-09, indicating that the abundance of p-aminobenzoic acid varied significantly between patients at different stages. The p-value was 2.532e-20, indicating that the abundance of choline was significantly different between healthy controls and HNSCC patients. The p-value was 4.477e-09, indicating that the abundance of glycocerophosphocholine was significantly different between healthy controls and patients with HNSCC at all stages. The p-value was 5.418e-22, indicating that the abundance of L-arginine was significantly different between HNSCC at all stages and healthy controls.
[0064] The prediction results of the HNSCC early diagnosis model, such as Figure 5 As shown, the metabolites in 8 of this example were applied to test data sets 1 and 2. The results showed that according to the classification threshold predicted by the model, patients should be above the horizontal line, the vast majority of healthy people are below the threshold horizontal line, and only individuals are above the horizontal line, which are patients with HNSCC. The vast majority of patients are above the threshold horizontal line, and only individuals are below the horizontal line, which are classified as healthy people. The AUC value of the training set is 1, and the AUC values of test sets 1 and 2 are 0.99 and 0.9958, respectively, indicating the consistency and high accuracy of the model on different data sets. An AUC value close to 1 means that the model can accurately distinguish HNSCC patients from healthy controls, and the diagnostic accuracy is higher than that of traditional methods.
Claims
1. Application of an organism metabolic marker in constructing a diagnostic model for squamous cell carcinoma.
2. The use according to claim 1, characterized in that The metabolite markers are screened from hundreds of metabolites by systematically collecting patient blood metabolite data based on a large sample size and combining broad-target metabolomics technology.
3. The use according to claim 2, characterized in that The metabolite markers are determined by performing metabolite analysis on an LC-MS / MS platform combined with a LASSO feature selection algorithm.
4. The use according to claim 1, characterized in that The metabolic marker is selected from at least one of cystine, L-glutamic acid, hypoxanthine, p-aminobenzoic acid, acetylcysteine, choline, glycerophosphocholine and L-arginine.
5. The use according to claim 1, characterized in that The metabolic markers include p-aminobenzoic acid and at least the following group of substances: Group 1: one or more of hypoxanthine and choline; Group 2: one or more of glycerophosphocholine, L-glutamic acid and L-arginine; Group III: cystine; and Group 4: Acetylcysteine.
6. A system for diagnosing early-stage squamous cell carcinoma, characterized in that: The following units are included: A detection unit, comprising a biological metabolic marker detection module; an analysis unit, which uses the results of the markers in the sample detected by the detection unit as input items for analysis; as well as An evaluation unit, which outputs the risk of squamous cell carcinoma for the individual corresponding to the sample; The biological metabolic markers include p-aminobenzoic acid and at least the following group of substances: Group 1: one or more of hypoxanthine and choline; Group 2: one or more of glycerophosphocholine, L-glutamic acid and L-arginine; Group III: cystine; and Group 4: Acetylcysteine.
7. The system according to claim 6, characterized in that The detection information of the biological metabolic marker detected by the detection unit comes from the information obtained by using the kit to detect the biological sample.
8. The system according to claim 6, wherein: According to the classification threshold of the model, judgment is made. If the sample marker is above the classification threshold, it is judged that the risk of the sample suffering from squamous cell carcinoma is high. If the sample marker is near the classification threshold, it is judged that the risk of the sample suffering from squamous cell carcinoma needs to be further observed. If the sample marker is below the classification threshold, the risk of the sample suffering from squamous cell carcinoma is low.
9. Use of an organism metabolic marker in the preparation of a diagnostic kit for squamous cell carcinoma, wherein the metabolic marker comprises p-aminobenzoic acid and at least one of the following substances: Group 1: one or more of hypoxanthine and choline; Group 2: one or more of glycerophosphocholine, L-glutamic acid and L-arginine; Group III: cystine; and Group 4: Acetylcysteine.
10. A kit, characterized in that The invention comprises reagents for detecting at least one of cystine, L-glutamic acid, hypoxanthine, p-aminobenzoic acid, acetylcysteine, choline, glycerophosphocholine and L-arginine, and standard substances.
Citation Information
Patent Citations
Diagnosis marker suitable for early-stage esophageal squamous cell cancer diagnosis and screening method of diagnosis marker
CN105044361A
Joint marker for diagnosing bladder cancer, kit and application
CN109709220A
Application of metabolic marker in preparation of oral cancer diagnostic kit
CN114965801A
Head and neck squamous cell carcinoma organoid and culture method, culture medium and application thereof
CN119506215A
Plasma metabolite composition for auxiliary diagnosis of subclinical mastitis of dairy cow and application of plasma metabolite composition
CN119555943A