Squamous cell carcinoma detection markers and their use in diagnostic models and reagents

By using broad-target metabolomics technology to screen key metabolite combinations and establish diagnostic models, the problems of small sample size and limited metabolite coverage in the early diagnosis of head and neck squamous cell carcinoma have been solved, achieving highly accurate and stable early diagnosis.

CN120432011BActive Publication Date: 2025-11-04SHANGHAI NINTH PEOPLES HOSPITAL SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510509875.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-11-04
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In existing technologies, early diagnostic methods for head and neck squamous cell carcinoma suffer from small sample sizes and limited metabolite coverage, resulting in poor model stability and difficulty in meeting the diagnostic requirements of high sensitivity and high specificity.

Method used

Using broad-target metabolomics technology and combining large-sample serum metabolite data, key metabolite combinations were screened using the LC-MS/MS platform and LASSO feature selection algorithm to establish a diagnostic model. These metabolites included cystine, L-glutamate, hypoxanthine, para-aminobenzoic acid, acetylcysteine, choline, and glycerophosphate choline. The diagnostic model was constructed and detected using a kit.

Benefits of technology

The model improved the diagnostic accuracy and stability of squamous cell carcinoma. Its AUC value was close to 1 on different datasets, showing high accuracy and consistency. It can identify the characteristic metabolic changes of squamous cell carcinoma at an early stage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120432011B_ABST
    Figure CN120432011B_ABST
Patent Text Reader

Abstract

The application of a biological metabolism marker in constructing an early diagnosis model of squamous cell carcinoma, the metabolism marker is screened and determined from hundreds of metabolites by systematically collecting metabolite data in the serum of patients on the basis of a large sample size and combining wide-target metabolomics technical means. It is verified that the consistency and high accuracy of the model on different data sets. The AUC value close to 1 means that the model can accurately distinguish HNSCC patients from healthy controls, and the diagnostic accuracy is higher than that of traditional methods. The selected metabolites are closely related to the disease, the diagnostic accuracy is high, and the model is robust.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a new use of a known compound, in particular to a marker for detecting squamous cell carcinoma, and use of the marker in a diagnostic model and preparation of a diagnostic kit. BACKGROUND

[0002] Head and neck squamous cell carcinoma (HNSCC) is one of the most common malignant tumors worldwide, commonly occurring in the oral cavity, throat, and larynx. Due to its insidious early symptoms, most patients are in the advanced stage of the disease at the time of diagnosis, which seriously affects the treatment effect and survival prognosis. Therefore, developing a sensitive, specific, and non-invasive early diagnostic model is a key problem that needs to be solved in the current clinical diagnosis and treatment of HNSCC.

[0003] Currently, the early diagnosis of HNSCC mainly relies on imaging examination, tissue biopsy, and detection of some tumor markers, but these methods have limitations. For example, imaging has poor recognition ability for small lesions, tissue biopsy is invasive and not suitable for widespread screening, and commonly used molecular markers (such as EGFR, p53, etc.) have low sensitivity and specificity, making it difficult to meet the high standard requirements of early diagnosis in clinical practice.

[0004] With the development of metabolomics technology, researchers have begun to try to detect changes in metabolites in tumor patients to assist in disease diagnosis. Metabolomics is a high-throughput technology that can systematically analyze changes in small molecule metabolites in the body, and can reflect the disease state from the whole body level. This technology is mainly divided into non-targeted metabolomics, targeted metabolomics, and broad-targeted metabolomics. Among them, non-targeted metabolomics is suitable for discovering novel metabolites, but its qualitative and quantitative accuracy is poor; targeted metabolomics has high quantitative accuracy, but is limited by the pre-set metabolite range and lacks flexibility. In contrast, broad-targeted metabolomics combines the advantages of both, covering a wider range of metabolites and having good quantitative ability, making it particularly suitable for screening and model building of complex disease-related metabolite combinations.

[0005] Previous studies have shown that HNSCC patients have a series of characteristic changes in metabolites, and some metabolite combinations have been tried for the establishment of diagnostic models. However, these studies are mostly based on small sample sizes of clinical data, and the metabolite detection techniques used are mainly targeted or non-targeted metabolomics, with limited metabolite coverage, and the clinical stability and generalizability of the models still face challenges. SUMMARY

[0006] An object of the present application is to provide a squamous cell carcinoma detection marker, to establish a diagnostic model closely related to head and neck squamous cell carcinoma (HNSCC) metabolism based on broad-target metabolomics, so as to solve the technical defects of the prior art, such as poor model stability and limited metabolite coverage range caused by small sample size.

[0007] Another object of the present application is to provide a squamous cell carcinoma early screening model to improve diagnostic accuracy and stability.

[0008] Still another object of the present application is to provide a squamous cell carcinoma metabolite detection reagent for early screening, which provides strong support and basis for early diagnosis and treatment of diseases.

[0009] Application of a biological metabolism marker in constructing a squamous cell carcinoma diagnostic model.

[0010] Application of another biological metabolism marker in constructing a head and neck squamous cell carcinoma diagnostic model.

[0011] The present application screens key metabolite combinations with statistical significance and disease correlation from hundreds of metabolites by systematically collecting metabolite data in patient serum based on a large sample size and combining broad-target metabolomics technical means, and establishes a HNSCC metabolic diagnostic model with high sensitivity, strong specificity and suitable for clinical early screening and auxiliary diagnosis.

[0012] The present application screens key metabolites from head and neck squamous cell carcinoma confirmed patients by LC-MS / MS platform for metabolite analysis, combining LASSO feature selection algorithm and with the aid of a computer system, such as at least one of cystine, L-glutamic acid, hypoxanthine, p-aminobenzoic acid, acetylcysteine, choline, glycerophosphocholine and L-arginine closely related to the occurrence and progression of HNSCC. Through these metabolites, the diagnostic accuracy of the model is significantly improved, which can detect the characteristic metabolic changes of HNSCC earlier, thereby providing strong support for early diagnosis.

[0013] Application of another biological metabolism marker in constructing a (head and neck) squamous cell carcinoma diagnostic model, the metabolism marker includes p-aminobenzoic acid, and at least one of the following groups of substances:

[0014] First group: one or more of hypoxanthine and choline;

[0015] Second group: one or more of glycerophosphocholine, L-glutamic acid and L-arginine;

[0016] Third group: cystine; and

[0017] Fourth group: acetylcysteine.

[0018] A system for diagnosing early squamous cell carcinoma, characterized in comprising the following units:

[0019] A detection unit comprising a biological metabolic marker detection module;

[0020] An analysis unit that analyzes the results of the markers in the sample detected by the detection unit as input items; and

[0021] An evaluation unit that outputs the risk level of the individual corresponding to the sample suffering from squamous cell carcinoma.

[0022] According to the classification threshold of the model, if the sample marker is above the classification threshold, it is judged that the sample has a higher risk of suffering from squamous cell carcinoma, if the sample marker is near the classification threshold, it is judged that the sample needs to be further observed, and if the sample marker is below the classification threshold, the risk of suffering from squamous cell carcinoma is lower.

[0023] In the present application, the classification threshold (cut-off value) is set to 0.5, which is a fixed value determined by the machine learning algorithm during the training of the diagnostic model, representing the dividing point between high-risk and low-risk HNSCC. After the model is built, this threshold remains consistent in all subsequent sample detection, without the need to regenerate each time, ensuring the consistency of the detection results and the stability of the model.

[0024] Specifically, after obtaining the metabolic marker detection information in the sample serum, the detection information is input into the diagnostic model to generate a comprehensive score value (risk score), which is compared with the fixed classification threshold (0.5) to realize the classification judgment of the sample HNSCC risk. In order to improve the discriminant power and clinical interpretability of the model, the concept of "deviation from cut-off" is further introduced to describe the degree of deviation of the sample score relative to the classification threshold. The judgment rule is as follows:

[0025] When the score is higher than the classification threshold (0.5) + 0.059, it is judged as high risk, indicating that the sample has a higher risk of suffering from HNSCC;

[0026] When the score is in the classification threshold (0.5) ± 0.059 interval, it is judged as a risk critical zone, which needs to be further followed up or assisted by other clinical information;

[0027] When the score is lower than the classification threshold (0.5) - 0.059, it is judged as low risk, and the sample has a lower risk of suffering from HNSCC.

[0028] Wherein, 0.059 is the tolerance threshold range set in the model training stage (determined based on the standard deviation of the training data set HNSCC score distribution), which can be used as a confidence interval to set the stability of the model.

[0029] The detection information of the biological metabolism markers detected by the detection unit is obtained by using the kit to detect the biological sample. The metabolic markers detected by the kit include p-aminobenzoic acid, and at least one of the following groups:

[0030] The first group: one or more of hypoxanthine and choline;

[0031] The second group: one or more of glycerophosphocholine, L-glutamic acid and L-arginine;

[0032] The third group: cystine; and

[0033] The fourth group: acetylcysteine.

[0034] The kit of the present application comprises reagents and standards required for detecting at least one of cystine, L-glutamic acid, hypoxanthine, p-aminobenzoic acid, acetylcysteine, choline, glycerophosphocholine and L-arginine.

[0035] It has been verified that the diagnostic model provided by the present application is based on high-throughput broad-target metabolome data, combined with statistical and machine learning algorithms, and the diagnostic model constructed has extremely high accuracy. The AUC values of the training set, test set 1 and test set 2 are 1, 0.99 and 0.9958 respectively, indicating the consistency and high accuracy of the model on different data sets. The AUC value close to 1 means that the model can accurately distinguish between HNSCC patients and healthy controls, and the diagnostic accuracy is higher than that of traditional methods. The selected metabolites are closely related to the disease, the diagnostic accuracy is high, and the model is robust.

[0036] The model provided by the present application adopts a strict quality control process to ensure the reliability of the metabolomics data. By inserting QC samples every 10 samples, the instrument stability is monitored, and batch correction and biomass standardization are applied to eliminate batch effects and biomass differences, which greatly improves the stability and accuracy of the data, providing a guarantee for subsequent analysis and model construction.

[0037] The diagnostic model provided by the present application has a standardized establishment process, has correlation between the metabolic markers and the disease, improves the accurate diagnosis ability, and has the potential to be widely applied to early screening. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 The overall flowchart of the head and neck squamous cell carcinoma early diagnosis model construction method of the present application;

[0039] Figure 2 Flowchart of the procedure for diagnostic model construction and validation;

[0040] Figure 3 Figure of the performance evaluation results of the early diagnosis model for head and neck squamous cell carcinoma (HNSCC); wherein, A is the ROC curve of the training data set (Train_data) and test data set 1 (Test_data1), B is the ROC curve of test data set 2 (Test_data2), and C is the column chart of the importance ranking of metabolite features;

[0041] Figure 4 Figure of the comparison of the abundance of 8 metabolites in patients at each stage of HNSCC and healthy controls; wherein, A is the relative abundance (Z-score) graph of cystine (Cystine) between HNSCC at each stage and healthy controls (N), B is the relative abundance (Z-score) graph of L-glutamic acid (L-Glutamic acid) between HNSCC at each stage and healthy controls (N), C is the relative abundance (Z-score) graph of hypoxanthine (Hypoxanthine) between HNSCC at each stage and healthy controls (N), D is the relative abundance (Z-score) graph of p-aminobenzoic acid (p-Aminobenzoic acid) between HNSCC at each stage and healthy controls (N), E is the relative abundance (Z-score) graph of acetylcysteine (Acetylcysteine) between HNSCC at each stage and healthy controls (N), F is the relative abundance (Z-score) graph of choline (Choline) between HNSCC at each stage and healthy controls (N), G is the relative abundance (Z-score) graph of glycerophosphocholine (Glycerophosphocholine) between HNSCC at each stage and healthy controls (N), and H is the relative abundance (Z-score) graph of L-arginine (L-Arginine) between HNSCC at each stage and healthy controls (N);

[0042] Figure 5 Figure of the prediction results of the early diagnosis model for HNSCC; wherein, A is the prediction results of the training data, B is the prediction results of test data set 1, and C is the prediction results of test data set 2; the red and blue scattered points respectively represent HNSCC patients and healthy controls, and the yellow horizontal line is the classification threshold of the model prediction, and the patients should be above the horizontal line and the healthy population should be below the horizontal line. DETAILED DESCRIPTION

[0043] The technical solutions of the present application are described in detail below with reference to the drawings. The embodiments of the present application are only used to illustrate the technical solutions of the present application and not to limit the same. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, and all of them should be covered in the scope of the claims of the present application.

[0044] Example 1: Construction of HNSCC early diagnosis model based on broad-target metabolomics

[0045] The steps of this embodiment are shown as follows: Figure 1

[0046] Step 1: Sample collection

[0047] This embodiment includes 293 patients with head and neck squamous cell carcinoma (HNSCC) diagnosed by pathology and 282 healthy controls (HC).

[0048] Queue 1 comes from the Ninth People's Hospital Affiliated to Shanghai Jiao Tong University School of Medicine / Professional Technical Service Platform of Oral and Maxillofacial Tumor Tissue Sample and Bioinformatics Database, which contains 498 subjects, including 248 HNSCC patients and 250 healthy controls.

[0049] Queue 2 also comes from the Ninth People's Hospital Affiliated to Shanghai Jiao Tong University School of Medicine / Professional Technical Service Platform of Oral and Maxillofacial Tumor Tissue Sample and Bioinformatics Database, which contains 77 subjects, including 45 HNSCC patients and 32 healthy controls.

[0050] Step 2: Serum sample collection

[0051] All subjects were collected with fasting venous blood after overnight fasting, and sampling was performed using inert separation gel coagulation tubes. The samples were centrifuged at 3000g for 10 minutes at 4°C, and the supernatant (serum) was collected and stored at -80°C.

[0052] Step 3: Metabolite extraction

[0053] 40 μL of serum of each sample was mixed with 280 μL of ice-cold methanol: acetonitrile (1:1, v:v), vortexed for 1 minute, and then centrifuged at 13000 rpm for 15 minutes at 4°C. The supernatant was collected and placed in an injection vial. For quality control (QC) samples, 5 μL of plasma of each sample was treated the same as the serum sample, and then injection analysis was performed.

[0054] Step 4: Targeted metabolomics analysis

[0055] ​The samples were analyzed based on LC-MS / MS platform. The samples were injected in random order (injection volume 1 μL, stored at 4°C), separated by ACQUITY UPLC HSS T3 column (2.1 x 100 mm, 1.8 μm) with mobile phase of water containing 0.01% formic acid (phase A) and acetonitrile (phase B), gradient elution program was 100% A phase for 0-2 min, then 100%→5% A phase, to 14 min, 5% A phase for 2 min, finally A phase back to 100% in 4 min, a total of 20 min, flow rate 0.2 mL / min, column temperature 30°C. Mass spectrometry detection used AB QTRAP 4500 system, 420 water-soluble metabolites were monitored in Scheduled MRM mode. The data were processed by SCIEX OS1.6 software, 185 stable metabolites were screened out by 80% rule (≥80% of samples were detected in each group), a small amount of missing values were filled with baseline value 1000 to ensure the reliability of subsequent analysis.

[0056] Step 5: Data correction and pretreatment

[0057] This example established a strict quality control (QC) and two-stage standardization strategy to ensure the reliability of metabolomics data. QC samples were inserted into the analysis batch every 10 test samples to monitor instrument stability. Standardization was divided into two stages, namely

[0058] Batch correction: Based on the global mean of QC samples (QCall) and the adjacent local mean (QCadj), metabolite-specific correction factors (QCall / QCadj) were calculated to dynamically eliminate systematic bias and batch effects.

[0059] Biomass standardization: The corrected metabolite abundance was converted into the relative proportion of total ion current (TIC) to eliminate the interference of biomass differences.

[0060] Step 6: Construction and verification of diagnostic model

[0061] Based on LASSO feature selection and machine learning algorithm (see Figure 2 ), a diagnostic prediction model was constructed. The study cohort (cohort 1, n=498) was divided into training set (n=348) and test set (n=150) by random stratified sampling. The diagnostic model constructed from the training set was verified in the test set and cohort 2 samples, respectively. The final diagnostic model consisted of eight metabolites, namely: Choline, p. Aminobenzoic. acid, L. Glutamic. acid, Glycerophosphocholine, Hypoxanthine, Cystine, L. Arginine, Acetylcysteine.

[0062] Figure 3 A shows the ROC curves for queue 2. The AUC value for the training dataset is 1, indicating the model's perfect predictive ability on the training set. The AUC value for test dataset 1 (Test_data1) is 0.99, demonstrating the model's excellent performance on the independent test dataset. Figure 3 B shows the ROC curve for test dataset 2 (Test_data2), with an AUC value of 0.9958, further validating the model's consistency and high accuracy across different datasets. Figure 3 C shows the important metabolites and their importance scores that are relevant to the diagnosis of HNSCC.

[0063] Figure 4 This graph shows the difference in expression levels of the eight key metabolites screened between HNSCC patients and healthy controls. A p-value of 2.455e-16 indicates that the abundance of Cystine was significantly higher in HNSCC patients than in healthy controls. A p-value of 1.177e-17 indicates a significant difference in the abundance of L-Glutamic acid across different HNSCC stages compared to healthy controls. A p-value of 9.835e-27 shows a significant variation in the abundance of Hypoxanthine between early and late-stage HNSCC patients. A p-value of 5.886e-27 indicates a significant difference in the abundance of p-Aminobenzoic acid between HNSCC patients and healthy controls. A p-value of 6.286e-09 shows significant variations in abundance across different stages of the disease. A p-value of 2.532e-20 indicates a significant difference in the abundance of Choline between healthy controls and HNSCC patients. The p-value was 4.477e-09, indicating a significant difference in the abundance of Glycerophosphocholine between healthy controls and patients at each stage of HNSCC. The p-value was 5.418e-22, indicating an extremely significant difference in the abundance of L-Arginine between each stage of HNSCC and the healthy control group.

[0064] Predictive results of the HNSCC early diagnosis model, such as Figure 5 As shown, the eight metabolites from this embodiment were applied to test dataset 1 and test dataset 2. The results showed that, according to the classification threshold predicted by the model, those above the horizontal line were patients, the vast majority of healthy individuals were below the threshold, and only one individual was above the threshold, classifying them as HNSCC patients. The training set AUC value was 1, while the AUC values ​​for test datasets 1 and 2 were 0.99 and 0.9958, respectively, indicating the model's consistency and high accuracy across different datasets. An AUC value close to 1 means the model can accurately distinguish between HNSCC patients and healthy controls, demonstrating higher diagnostic accuracy than traditional methods.

Claims

1. The application of a reagent for detecting metabolic markers in organisms in constructing a diagnostic model of head and neck squamous cell carcinoma, characterized in that... The aforementioned biomarkers for metabolism are cystine, L-glutamic acid, hypoxanthine, para-aminobenzoic acid, acetylcysteine, choline, glycerophosphate choline, and L-arginine.

2. The application according to claim 1, characterized in that... The aforementioned biomarkers for metabolism are obtained by systematically collecting blood metabolite data from patients on a large sample basis and screening from hundreds of metabolites using broad-target metabolomics techniques.

3. The application according to claim 2, characterized in that... The aforementioned biomarkers were determined by metabolite analysis using an LC-MS / MS platform, combined with the LASSO feature selection algorithm.

4. A system for diagnosing early-stage squamous cell carcinoma of the head and neck, characterized in that, Includes the following units: The detection unit includes a module for detecting metabolic biomarkers in organisms; The analysis unit takes the results of the markers in the sample detected by the detection unit as input for analysis; as well as The assessment unit outputs the risk level of squamous cell carcinoma for the individual corresponding to the sample. The aforementioned biomarkers for metabolism are cystine, L-glutamic acid, hypoxanthine, para-aminobenzoic acid, acetylcysteine, choline, glycerophosphate choline, and L-arginine.

5. The system according to claim 4, characterized in that, The detection information of the biological metabolic markers detected by the detection unit comes from the information obtained by using the reagent kit to detect biological samples.

6. The system according to claim 4, characterized in that, The risk of squamous cell carcinoma is assessed based on the model's classification threshold. If the sample marker is above the classification threshold, the risk is considered high. If the sample marker is near the classification threshold, the risk needs further observation. If the sample marker is below the classification threshold, the risk is considered low.

7. The application of a reagent for detecting metabolic markers in the preparation of a diagnostic kit for head and neck squamous cell carcinoma, wherein the metabolic markers are cystine, L-glutamic acid, hypoxanthine, para-aminobenzoic acid, acetylcysteine, choline, glycerophosphate choline, and L-arginine.

Citation Information

Patent Citations

  • Diagnosis marker suitable for early-stage esophageal squamous cell cancer diagnosis and screening method of diagnosis marker

    CN105044361A

  • Plasma metabolite composition for auxiliary diagnosis of subclinical mastitis of dairy cow and application of plasma metabolite composition

    CN119555943A