Method for assessing risk of thyroid cancer and cervical lymph node metastasis

By acquiring peripheral blood single-cell metabolomics signals using MicroCyESI-MS, and combining genetic algorithm screening and multi-level probability aggregation, the problem of non-invasive assessment of the risk of thyroid cancer and cervical lymph node metastasis was solved, achieving high-accuracy risk assessment and improved model performance.

CN122117355APending Publication Date: 2026-05-29CHINA UNIV OF GEOSCIENCES (WUHAN)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF GEOSCIENCES (WUHAN)
Filing Date
2026-01-13
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Current detection methods cannot accurately and non-invasively assess the risk of thyroid cancer and its cervical lymph node metastasis, especially lacking molecular-level markers at the single-cell level, and existing feature screening and classification models are difficult to reliably distinguish between tumor and non-tumor states.

Method used

Single-cell metabolomics signals from peripheral blood immune cells were obtained using microspray label-free mass cytometry (MicroCyESI-MS). A two-dimensional matrix of single-cell metabolomics signal intensity was constructed, and hierarchical aggregation was performed. Key metabolic features were screened using a genetic algorithm, and a multi-level probabilistic aggregation mechanism was used to generate patient-level risk scores.

Benefits of technology

It achieves highly accurate non-invasive risk assessment of thyroid cancer and cervical lymph node metastasis, with an AUC value above 0.92, significantly improving assessment performance and can be widely used for peripheral blood single-cell metabolic data analysis of various tumors and immune-related diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117355A_ABST
    Figure CN122117355A_ABST
Patent Text Reader

Abstract

The application discloses a thyroid cancer and neck lymph node metastasis risk assessment method, relates to the cross technical field of bioinformatics and metabolomics, and mainly comprises the following steps: converting single-cell metabolomics signals into a metabolic feature matrix with cells as rows and metabolites as columns, constructing a patient-level metabolic dataset based on patient identification information, preprocessing the patient-level metabolic dataset to obtain a preprocessed patient-level metabolic dataset, screening features to obtain key metabolic features, and performing hierarchical summary prediction to obtain individual risk assessment results. The thyroid cancer and neck lymph node metastasis risk assessment method can non-invasively and accurately predict the risk of thyroid cancer and neck lymph node metastasis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of bioinformatics and metabolomics, and more specifically, to a method for assessing the risk of thyroid cancer and cervical lymph node metastasis. Background Technology

[0002] Thyroid cancer is a common malignant tumor of the endocrine system, and its incidence is increasing year by year. Clinically, cervical lymph node metastasis is an important indicator affecting the staging and prognosis of thyroid cancer. Existing detection methods mainly include B-ultrasound imaging, fine needle aspiration (FNA) and intraoperative pathology, but these methods generally have the following problems: high invasiveness - requiring puncture or surgical sampling, resulting in poor patient compliance; limited detection resolution - imaging and pathological methods cannot reveal metabolic differences at the single-cell level; lack of molecular-level non-invasive markers - peripheral blood detection is mostly limited to circulating DNA or proteome signals, which cannot accurately reflect the state of immune cells. In recent years, the rapid development of single-cell mass spectrometry (SMS) technology has made it possible to obtain intracellular metabolic signals at single-cell resolution. Among them, micro-spray label-free mass cytometry (MicroCyESI-MS, Analytical Chemistry, 97(33):18382–18391) can achieve real-time high-throughput detection of metabolites at the level of a single immune cell. This technology can stably detect signals of multiple metabolites, including lipids, sterols, and sphingomyelins, at the single-cell level in peripheral blood. However, single-cell metabolomics data are typically characterized by high noise, high dimensionality, and significant individual variability. Currently, there is no publicly available method for predicting individual patient risk using single-cell metabolomics data while considering cellular heterogeneity. Existing feature screening and classification models struggle to stably distinguish between tumor and non-tumor states when processing such data.

[0003] How to combine single-cell metabolic detection with stratified probability aggregation to achieve non-invasive assessment of the risk of thyroid cancer and its cervical lymph node metastasis at the peripheral blood level is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] The purpose of this invention is to provide a method for assessing the risk of thyroid cancer and cervical lymph node metastasis, which can predict the risk of thyroid cancer and cervical lymph node metastasis with high accuracy and non-invasiveness.

[0005] This invention provides a method for assessing the risk of thyroid cancer and cervical lymph node metastasis, comprising the following steps: S1: Acquire single-cell metabolomics signal data from multiple cells, wherein the single-cell metabolomics signal data consists of mass spectrometry characteristic peaks generated by all metabolites; S2: Perform matrix transformation on the single-cell metabolomics signal data to construct a two-dimensional matrix of single-cell metabolomics signal intensity; S3: Based on patient identification information, hierarchical aggregation is performed on the two-dimensional matrix of single-cell metabolomics signal intensity, and a patient-level metabolic dataset is constructed by combining patient category labels; the patient identification information is used to associate all single-cell data from the same patient, the hierarchical aggregation is to statistically integrate all single-cell metabolomics signal data corresponding to the same patient to obtain patient-level metabolic feature data, and the patient-level metabolic dataset is a data matrix containing metabolic feature data of all patients and corresponding patient category labels; S4: Preprocess the patient-level metabolic dataset to obtain a preprocessed patient-level metabolic dataset; S5: Perform feature screening on the preprocessed patient-level metabolic dataset to obtain key metabolic features; S6: Perform hierarchical summary prediction based on the key metabolic characteristics to obtain individual risk assessment results.

[0006] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for assessing the risk of thyroid cancer and cervical lymph node metastasis.

[0007] The method for assessing the risk of thyroid cancer and cervical lymph node metastasis provided by this invention has the following beneficial effects: This invention utilizes single-cell metabolomics signals from peripheral blood immune cells obtained by microspray label-free mass cytometry (MicroCyESI-MS). After multi-level feature screening and optimization, a cell-level classification model is established, and a multi-level probabilistic aggregation mechanism is employed to generate patient-level risk scores, enabling non-invasive assessment of the risk of thyroid cancer and cervical lymph node metastasis. The ratios of phospholipids, sphingomyelins, and sterols screened by this invention can significantly distinguish papillary thyroid cancer patients from control individuals. This invention requires only peripheral blood lymphocytes for assessment, achieving non-invasive detection. It utilizes a multi-level aggregation mechanism to achieve stable patient-level risk output through single-cell probability distribution aggregation. It employs a genetic algorithm to effectively screen highly relevant feature combinations, reducing redundant feature interference and achieving feature evolution optimization. The invention achieves an AUC (Area Under the Receiver Operating Characteristic) of 0.92 in assessing the positive risk of papillary thyroid carcinoma and 0.79 in assessing the risk of cervical lymph node metastasis, demonstrating significant performance improvements. Combining single-cell mass spectrometry detection, feature screening, and machine learning model construction, this invention has good versatility. Besides being used for data-driven prediction and assessment of the risk of thyroid cancer and cervical lymph node metastasis, it can also be extended to peripheral blood single-cell metabolic data analysis for various tumors and immune-related diseases. Attached Figure Description

[0008] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a flowchart of the risk assessment method for thyroid cancer and cervical lymph node metastasis provided by the present invention; Figure 2 This is a schematic diagram of the overall technical solution of the risk assessment method for thyroid cancer and cervical lymph node metastasis provided by the present invention; Figure 3 This invention provides , , and A schematic diagram showing the distribution of the four indicators in patients with non-papillary thyroid carcinoma, papillary thyroid carcinoma without lymph node metastasis, and papillary thyroid carcinoma with lymph node metastasis. Figure 4 This is a schematic diagram of the multi-level probability aggregation mechanism provided by the present invention; Figure 5 This is a schematic diagram of the AUC curve for risk assessment of papillary thyroid carcinoma provided by the present invention; Figure 6 This is a schematic diagram of the feature selection and genetic algorithm optimization lineage provided by the present invention; Figure 7 This is a schematic diagram of the AUC curve for assessing the risk of lymph node metastasis in patients with papillary thyroid carcinoma, provided by the present invention. Detailed Implementation

[0009] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0010] Figure 1 A schematic diagram of the risk assessment method for thyroid cancer and cervical lymph node metastasis according to this embodiment is shown. In this embodiment, the risk assessment method for thyroid cancer and cervical lymph node metastasis includes the following steps: S1: Acquire single-cell metabolomics signal data from multiple cells.

[0011] It should be noted that single-cell metabolomics signal data refers to the mass spectrometry characteristic peaks produced by all metabolites, including identifiable lipid characteristic peaks and characteristic peaks of all metabolic features.

[0012] As an exemplary embodiment, in step S1, peripheral blood samples of the subject are collected, immune cells are isolated, and the metabolic profile of a single cell is acquired by mass spectrometry metabolic signal acquisition technology with single-cell resolution to obtain single-cell metabolomics signal data of the cell.

[0013] As an example, each sample collects 150 to 250 cell signals and detects more than 600 characteristic peaks, mainly lipids, of metabolites.

[0014] As an exemplary embodiment, the immune cells are peripheral blood lymphocytes, and the single-cell metabolomics signal acquisition technology is microspray label-free mass cytometry (MicroCyESI-MS) or other equivalent single-cell mass spectrometry detection technology.

[0015] S2: Perform matrix transformation on the single-cell metabolomics signal data to construct a two-dimensional matrix of single-cell metabolomics signal intensity; S3: Based on patient identification information, hierarchical aggregation is performed on the two-dimensional matrix of single-cell metabolomics signal intensity, and a patient-level metabolic dataset is constructed by combining patient category labels; the patient identification information is used to associate all single-cell data from the same patient, the hierarchical aggregation is to statistically integrate all single-cell metabolomics signal data corresponding to the same patient to obtain patient-level metabolic feature data, and the patient-level metabolic dataset is a data matrix containing metabolic feature data of all patients and corresponding patient category labels.

[0016] In one exemplary embodiment, the patient-level metabolic dataset includes a two-dimensional matrix of cells and metabolic features for each patient. The two-dimensional matrix is ​​arranged with cell numbers as rows and metabolic features as columns, and the matrix elements are the mass spectrometry signal intensities of each metabolic feature under the corresponding cell number.

[0017] As an exemplary embodiment, in the above steps, the original signal is converted into a "cellular signal". A two-dimensional matrix of metabolic features is used, labeled according to the subject's ID. Each cell represents a row in the matrix, and the metabolic features represent columns, forming a hierarchical data structure for patients and constructing a patient-level metabolic dataset.

[0018] S4: Preprocess the patient-level metabolic dataset to obtain a preprocessed patient-level metabolic dataset.

[0019] In one exemplary embodiment, the preprocessing includes missing value completion, outlier identification and removal, quantile truncation, and standardization.

[0020] In one exemplary embodiment, the method for identifying abnormal samples is an unsupervised anomaly identification model.

[0021] In one exemplary embodiment, the unsupervised anomaly identification model is an isolated forest algorithm or a density anomaly detection algorithm.

[0022] In one exemplary embodiment, the pollution rate of the isolated forest algorithm is 0.02; In one exemplary embodiment, the method for filling in missing values ​​is median padding; In one exemplary embodiment, the truncation ratio is 1% above and below. In one exemplary embodiment, the standardization operation includes Z-score normalization and a logarithmic transformation to base 10; As an exemplary embodiment, in step S4, missing value completion, abnormal sample identification and removal, quantile truncation and standardization operations are performed on the matrix data to identify abnormal metabolic signals at the single-cell level, improve model stability, reduce detection errors, reduce noise interference and stabilize feature distribution.

[0023] S5: Perform feature screening on the preprocessed patient-level metabolic dataset to obtain key metabolic features.

[0024] In one exemplary embodiment, the feature filtering step includes: High-information-content candidate features are selected from the preprocessed patient-level metabolic dataset using an information content filtering method. The variance inflation factor is calculated based on the high-information-content candidate features, and high-collinearity features are removed based on the variance inflation factor to obtain the candidate features after removing high-collinearity features. Based on the candidate features after removing highly collinear features, nonparametric tests or equivalent significance detection methods are used to screen out differential features and obtain key metabolic features.

[0025] In one exemplary embodiment, the feature selection step further includes: optimizing the feature subset using a genetic algorithm based on the differential features, and iteratively evolving using cross-validation AUC as the fitness function to obtain the key metabolic features of the optimal feature combination.

[0026] In one exemplary embodiment, the information filtering method is a mutual information calculation method, a correlation coefficient method, or an information gain method; The threshold for removing highly collinear features is 10; The nonparametric test method is the Mann-Whitney U test method; In one exemplary embodiment, the probability that the null hypothesis of the Whitney U test is true is 0.05; In one exemplary embodiment, the key metabolic characteristic is a ratio index formed by phospholipids, sphingomyelins, sterols and polyunsaturated lipid metabolites, which is used to reflect cell membrane composition, signal transduction activity and energy metabolism status.

[0027] In one exemplary embodiment, the key metabolic characteristics include , , , ,in, This represents the ratio of the sum of the signal intensities of phosphatidylcholine and sphingomyelin to the signal intensity of 10-pentadien-1-ol. This represents the ratio of the sphingosine base signal intensity to the sum of the signal intensities of ether-bonded phosphatidic acid and phosphatidylinositol. This represents the ratio of the total signal intensity of phospholipid metabolites to the total signal intensity of sterol metabolites; This represents the ratio of the sum of saturated lipid signal intensities to the sum of polyunsaturated lipid signal intensities.

[0028] It should be noted that the key metabolic features selected in the model include the following ratio indicators: , P , ; PEN: Signal intensity of 10-pentadecadien-1-ol; SPL: The sum of the signal intensities of ether-bonded phosphatidic acid (PAO-42:4) and ether-bonded phosphatidylinositol (PIO-34:0 and PIO-36:0); SPB: SPB17:0; The sum of signal strengths; PL: The sum of signal intensities of all detected phospholipid metabolites; ST: The sum of signal intensities of all detected sterol metabolites; PUL: The sum of signal intensities of all detected polyunsaturated lipids; SL: The sum of signal intensities of all detected saturated lipids; This refers to the ratio of PC+SM to PEN. It refers to the ratio of SPB to SPB. It refers to the ratio of PL to ST. It refers to the ratio of SL to PUL.

[0029] In this embodiment: PC+SM is the sum of the signal intensities of phosphatidylcholine and sphingomyelin, specifically the sum of the signal intensities of PC34:0 (phosphatidylcholine with 34 carbons and no double bonds), PC34:1 (phosphatidylcholine with 1 double bond), PC34:2 (phosphatidylcholine with 2 double bonds), PC34:3 (phosphatidylcholine with 3 double bonds), PC36:2 (phosphatidylcholine with 2 double bonds), PC36:3 (phosphatidylcholine with 3 double bonds) and SM34:1;O2 (sphingomyelin molecule with 34 carbons, 1 double bond, and two oxygen atoms). PEN is the sum of the signal intensities of 10-pentadecadien-1-ol; SPB represents the signal intensity of the sphingosine base (SPB17:0;O2); SPL is the sum of the signal intensities of ether-bonded phosphatidic acid and ether-bonded phosphatidylinositol, that is, the sum of the signal intensities of PAO-42:4, PIO-34:0 and PIO-36:0; PL is the sum of the signal intensities of all detected phospholipid metabolites, i.e., the sum of the signal intensities of PSO-28:1, PCO-36:3, PCO-36:2, PCO-36:1, PCO-32:0, PC40:7, PC38:6, PC38:3, PC38:2, PC36:4, PC36:3, PC36:2, PC36:1, PC36:0, PC34:4, PC34:3, PC34:2, PC34:1, PC34:0, PC32:2, PC32:0, PAO-42:4, PAO-42:2, PAO-42:0, PA41:6, PIO-36:0, and PIO-34:0. ST is the sum of the signal intensities of all detected sterol metabolites, specifically the sum of the signal intensities of ST30:3;O,ST28:2;O4,ST27:2;O;Hex,ST26:1;O5,ST24:5;O3,ST24:2;O4,ST24:1;O5,ST22:1;O3 and ST21:1;O2. SL is the sum of the signal intensities of all detected saturated lipids, namely, the sum of the signal intensities of WE22:0, TG24:0, SPB17:0; O2, PIO-36:0, PIO-34:0, PCO-32:0, PC36:0, PC34:0, PC32:0, PAO-42:0, NAE22:0, NAE20:0, NAE17:0, NAE16:0, NAE12:0, MG18:0, MG17:0, MG16:0, MG16:0, LPC18:0, LPC16:0, LPC14:0 and LPAO-16:0. PUL is the sum of signal intensities of all detected polyunsaturated lipids, namely: WE22:3,WE20:3,WE20:2,WE19:6;O2,WE13:2,TGO-58:8,TGO-58:7,TGO-56:8,TGO-56:7,TGO-56:6,TGO-56:5,TG54:5,TG54:4,TG54:3,TG52:4,TG52:3,TG52:2,TG50:3,TG50:2,TG48:5,ST30:3;O,ST28:2;O4,ST27:2;O;Hex,ST24:5;O3,ST24:2;O4,SM42:3;O2,SM42 :2;O2,SM40:2;O2,SM36:2;O2,SM34:2;O2,PCO-36:3,PCO-36:2,PC40:7,PC38:6,PC38:3,PC38:2,PC36:4,PC36:3,PC36:2,PC34:4,PC34:3,PC34:2,PC32:2,PAO-42:4,PAO-42:2,PA41:6,PA41:3,NAE24:4,NAE18:2,MG18:2,LPA19:3,FOH23:2;O3,FOH19:2;O3,DG44:7,DG38:5 and CAR16:2;The sum of the signal strengths of O; Cer is the sum of the signal intensities of C2-Ceramide and Dihydroceramide; SM is the sum of the signal intensities of SM34:2; O2, SM34:1; O2, SM36:2; O2, SM36:1, SM40:2; O2, SM42:3; O2, SM42:2; O2.

[0030] As an exemplary embodiment, in step S5, key metabolic features related to the disease state are screened based on statistical analysis, information content evaluation, or equivalent feature selection methods. This screening process is specifically adapted to the high-noise features of single-cell metabolomics signals, and a cell-level classification model is established based on the features. As an exemplary embodiment, the feature selection step includes the following sub-steps: (1) selecting candidate features with high information content based on mutual information, correlation coefficient or information gain; (2) eliminating highly collinear features through variance inflation factor analysis; (3) selecting differential features using nonparametric tests or equivalent significance detection methods; and (4) inputting the retained features into a nonlinear classification model for training.

[0031] As an exemplary embodiment, the key metabolic features include the following ratio indices calculated from single-cell metabolomics signals: , , , Wherein: PEN represents 10-pentadecadien-1-ol; SPL represents the sphingomyelin subgroup composed of PAO-42:4, PIO-34:0, and PIO-36:0; SPB represents SPB17:0; PL represents all detected phospholipid metabolites; ST represents all detected sterol metabolites; PUL represents all detected polyunsaturated lipid metabolites; SL represents all detected saturated lipid metabolites; Cer represents C2-ceramide and dihydroceramide; SM represents all detected sphingomyelin metabolites. These ratios are used to construct a single-cell metabolic feature matrix of immune cells and serve as input features for the model to predict the occurrence of papillary thyroid carcinoma in subjects.

[0032] S6: Perform hierarchical summary prediction based on the key metabolic characteristics to obtain individual risk assessment results.

[0033] In one exemplary embodiment, the hierarchical aggregation prediction includes: Based on the key metabolic features, cell-level prediction results are obtained using a cell-level classification model. Based on the cell-level prediction results, a hierarchical summary is performed to obtain a patient-level disease risk score. Individual risk assessment results are obtained based on the patient-level disease risk score and the set threshold.

[0034] In one exemplary embodiment, the cell-level classification model is a nonlinear classifier based on an ensemble learning algorithm, wherein the nonlinear classifier determines the optimal parameters through cross-validation and Bayesian optimization.

[0035] In one exemplary embodiment, the nonlinear classifier is an XGBoost model; In one exemplary embodiment, the patient-level disease risk score includes a positive proportion, which is calculated using the following formula: , in, The positive rate; To predict the number of single cells that will test positive; The number of single cells predicted to be negative; In one exemplary embodiment, the formula for calculating the patient-level disease risk score is as follows: , in, Assess the disease risk level of patients. This is an aggregation function used to reflect the probability distribution characteristics of single cells; For the first The positive predictive probability of a single cell.

[0036] In one exemplary embodiment, the set threshold is determined based on maximizing the classification evaluation metric on the validation set.

[0037] As an exemplary embodiment, in step S5, the cell-level prediction results are aggregated by individual, wherein the aggregation is based on the statistical distribution of the prediction probability of each single cell, and a patient-level disease risk score is generated by weighted average, distribution integral or quantile function, and the individual risk assessment result is output according to a set threshold. As an exemplary embodiment, the positive rate P in the hierarchical aggregation prediction step is calculated according to the formula... Calculation, where To predict the number of single cells that test positive, The number of single cells predicted as negative is given; a positive prediction result is output when P exceeds a threshold T, otherwise a negative result is output. The threshold T is automatically determined based on maximizing the classification evaluation index on the validation set.

[0038] As an exemplary embodiment, the hierarchical aggregation prediction step is based on the statistical aggregation of single-cell hierarchical prediction probabilities, and aggregates the positive prediction probabilities of each single cell. Perform weighted average or distribution analysis to generate patient-level disease risk scores. ,in For aggregation functions used to reflect the probability distribution characteristics of single cells, when Exceeding the threshold Output a positive prediction result if the prediction is positive, otherwise output a negative result.

[0039] In one exemplary embodiment, the method for assessing the risk of thyroid cancer and cervical lymph node metastasis further includes: outputting an analysis report that includes individual risk assessment results and model performance indicators.

[0040] As an exemplary embodiment, in the results output step, an analysis report including patient-level prediction results and model performance indicators is output, wherein the performance indicators include at least one area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, or specificity.

[0041] It should be noted that the aggregation rule in step S6 constitutes a multi-level summary prediction mechanism, which is used to reduce the influence of correlation between single cells and achieve stable discrimination at the individual level.

[0042] In some embodiments, the above-described risk assessment methods for thyroid cancer and cervical lymph node metastasis can also be implemented in the following ways.

[0043] like Figure 2 As shown, in this embodiment, the risk assessment method for thyroid cancer and cervical lymph node metastasis includes the following steps: lymphocyte collection; data acquisition: collecting peripheral blood, isolating immune cells, and performing single-cell metabolomics mass spectrometry testing; data structure, i.e., single-cell metabolic data matrix construction: generating "cell..." A two-dimensional matrix of metabolic features; data preprocessing and feature selection: missing value completion, anomaly handling, feature selection / optimization; feature selection and model building; hierarchical summarization and risk assessment, assessment result output; details are as follows: Step 1. Data Acquisition Peripheral blood samples were collected from the subjects, and immune cells (preferably peripheral blood lymphocytes) were isolated. Metabolic profiles of individual cells were acquired using microspray label-free mass cytometry (MicroCyESI-MS). Signals from 150 to 250 cells were collected from each sample, detecting over 600 characteristic peaks, primarily lipid-based metabolites.

[0044] Step 2. Data Structure Construction Transform the raw signal into a "cellular" The data is presented as a two-dimensional matrix of metabolic characteristics, labeled according to the subject's number. Each cell represents a row in the matrix, and the metabolic characteristics represent columns, forming a hierarchical data structure for patients.

[0045] Step 3. Data Preprocessing Missing values ​​were filled using the median; outlier identification was based on the Isolation Forest algorithm (contamination rate 0.02); features were truncated at the upper and lower 1% quantiles; normalization methods included Z-score normalization and logarithmic transformation to base 10. This step was used to identify abnormal metabolic signals at the single-cell level and improve model stability.

[0046] Step 4. Feature Selection and Modeling The following screening process was performed on the matrix features: mutual information calculation was used to select the top 1000 features with the highest information content; variance inflation factor (VIF) was calculated to remove features with high collinearity (threshold 10); the Mann-Whitney U test was used to screen features with significant differences (p<0.05); and a genetic algorithm (GA) was used to optimize the feature subset, with the area under the ROC curve (AUC) of cross-validation used as the fitness function for iterative evolution. Approximately 90 optimal feature combinations were finally obtained and input into a cell-level classification model (XGBoost) to identify papillary thyroid carcinoma.

[0047] like Figure 3 As shown, , , and The distribution of the four indicators in patients with non-papillary thyroid carcinoma, papillary thyroid carcinoma without lymph node metastasis, and papillary thyroid carcinoma with lymph node metastasis; each of the four indicators showed significant differences in the three groups.

[0048] Step 5. Hierarchical Summary and Risk Assessment The cell-level model outputs a classification label for each single cell. ,in Indicates a positive result. The result is negative. Patient-level risk score. It is calculated based on the classification results of all cells, and its definition is as follows:

[0049] in, This refers to the total number of single cells detected in this patient. To predict the number of cells that will be positive. When (threshold) When the F1 score is maximized in the validation set, the assessment result of "high risk of individual positivity" is output; when At that time, the assessment result of "low risk of individual positivity" is output.

[0050] like Figure 4 The diagram illustrates a multi-level probabilistic aggregation mechanism, showcasing the case of volunteer number 53. Each volunteer independently predicts the classification result for each cell, and then a cell voting strategy (based on a proportion threshold) is employed. The individual classification result of the subject is determined. If the ratio of cells classified as PTC to cells classified as NTC exceeds a certain threshold, the subject is determined to be PTC; otherwise, the subject is determined to be NTC (non-PTC). The single-cell classification result was 215 PTC cells and 43 NTC cells, which were collectively classified as PTC based on the above voting strategy.

[0051] This multi-level aggregation mechanism based on single-cell tag proportions can comprehensively reflect the abnormal proportions of cell populations at the individual level, reduce the impact of heterogeneity among single cells on the overall discrimination results, and thus obtain a more stable patient-level risk assessment.

[0052] Step 6. Output Results Output a patient risk assessment report, including predicted probability, classification results, and model performance metrics (AUC, accuracy, sensitivity, specificity).

[0053] like Figure 5 The figure shows the AUC curve for risk assessment of papillary thyroid carcinoma. Figure 6 The diagram illustrates feature selection and genetic algorithm optimization. After 40 generations of model genetic evolution, 84 features were selected for risk assessment of lymph node metastasis in papillary thyroid carcinoma, with an AUC value consistently maintained at 0.79. Figure 7 The figure shows the AUC curve for assessing the risk of lymph node metastasis in papillary thyroid carcinoma.

[0054] The method of this invention achieved an average AUC of 0.92 using the four indicators mentioned above in the risk assessment of papillary thyroid carcinoma, and an average AUC of 0.79 using 958 metabolic characteristics in the risk assessment of lymph node metastasis of papillary thyroid carcinoma.

[0055] It should be noted that the key metabolic features selected in the model include the following ratio indicators: , P , Where: PEN: 10-Pentadecadien-1-ol signal strength; SPL: sum of signal strengths from PAO-42:4, PIO-34:0, and PIO-36:0; SPB: SPB17:0; Sum of signal intensities; PL: Sum of signal intensities of all detected phospholipid metabolites; ST: Sum of signal intensities of all detected sterol metabolites; PUL: Sum of signal intensities of all detected polyunsaturated lipids; SL: Sum of signal intensities of all detected saturated lipids; This refers to the ratio of PC+SM to PEN. It refers to the ratio of SPB to SPB. It refers to the ratio of PL to ST. It refers to the ratio of SL to PUL. In this embodiment: PC+SM is the sum of the signal strengths of PC34:0, PC34:1, PC34:2, PC34:3, PC36:2, PC36:3 and SM34:1; O2; PEN is the sum of the signal strengths of 10-Pentadecadien-1-ol; SPB is the signal strength of SPB17:0; O2; SPL is the sum of the signal strengths of PAO-42:4, PIO-34:0 and PIO-36:0; PL is PSO-28:1, PCO-36:3, PCO-36:2, PCO-36:1, PCO-32:0, PC40:7, PC38:6, PC38:3, PC38:2, PC36:4, PC36:3, PC36:2, PC36:1, PC36:0, PC34:4, PC34:3, PC34:2, PC34:1. The sum of signal strengths for PC34:0, PC32:2, PC32:0, PAO-42:4, PAO-42:2, PAO-42:0, PA41:6, PIO-36:0, and PIO-34:0; ST is the sum of signal strengths for ST30:3;O, ST28:2;O4, ST27:2;O;Hex, ST26:1;O5, ST24:5;O3, ST24:2;O4, ST24:1;O5, ST22:1;O3, and ST21:1;O2; SL is the sum of signal strengths for WE22:0, TG24:0, SPB17:0;O2, PIO-36:0, PIO-34:0, PCO-32:0, PC36:0, PC34:0, PC32:0, PAO-42:0, NAE22:0, and NAE20:0. The sum of signal strengths for NAE17:0, NAE16:0, NAE12:0, MG18:0, MG17:0, MG16:0, MG16:0, LPC18:0, LPC16:0, LPC14:0, and LPAO-16:0; PUL is the sum of signal strengths for WE22:3, WE20:3, WE20:2, WE19:6; O2, WE13:2, TGO-58:8, TGO-58:7, TGO-56:8, TGO-56:7, TGO-56:6, TGO-56:5, TG54:5, TG54:4, TG54:3, TG52:4, TG52:3, TG52:2, TG50:3, TG50:2, TG48:5. ST30:3;O,ST28:2;O4, ST27:2;O;Hex, ST24:5;O3, ST24:2;O4, SM42:3;O2, SM42:2;O2, SM40:2;O2, SM36:2;O2, SM34:2;O2, PCO-36:3, PCO-36:2, PC40:7,PC38:6, PC38:3, PC38:2, PC36:4, PC36:3,PC36:2, PC34:4, PC34:3, PC34:2, PC32:2, PAO-42:4, PAO-42:2, PA41:6, PA41:3,NAE24:4, NAE18:2, MG18:2, LPA19:3, FOH23:2;O3, FOH19:2;O3, DG44:7, The sum of the signal intensities of DG38:5 and CAR16:2;O; Cer is the sum of the signal intensities of C2-Ceramide and Dihydroceramide; SM is the sum of the signal intensities of SM34:2;O2, SM34:1;O2, SM36:2;O2, SM36:1, SM40:2;O2, SM42:3;O2, and SM42:2;O2.

[0056] These ratios comprehensively reflect cell membrane fluidity, signal transduction activity, and energy metabolism status. When used as a risk assessment for papillary thyroid carcinoma, the AUC value was 0.92.

[0057] In some embodiments, the above-described risk assessment methods for thyroid cancer and cervical lymph node metastasis can also be implemented in the following ways.

[0058] 1: Single-cell metabolic profile acquisition Sixty-six patients with pathologically confirmed papillary thyroid carcinoma (PTC) and 26 healthy controls, totaling 92 subjects, were selected. PTC patients were further divided into a non-metastatic group (NNM, n=20) and a metastatic group (LNM, n=33) based on cervical lymph node status. Peripheral blood lymphocytes were isolated from peripheral blood samples and analyzed using microspray label-free mass cytometry (MicroCyESI-MS). The detection conditions were positive ion mode, with a mass detection range of 100–1200 m / z. An average of approximately 200 single-cell signals were detected per sample, totaling approximately 18,805 cells analyzed. The detected metabolic features mainly included phospholipids, sphingomyelins, sterols, and polyunsaturated lipid metabolites. After single-cell peak screening and mass-to-charge ratio alignment, "cell" data were generated. A two-dimensional matrix of metabolic features is used as input data for subsequent feature selection and model building.

[0059] The key features obtained through screening include ratios of phospholipids, sphingomyelins, sterols, and polyunsaturated lipid metabolites, mainly including... , , and Four indicators (for reference) Figure 3 These metabolic ratios reflect cell membrane composition, lipid signaling, and energy metabolism status, effectively distinguishing papillary thyroid carcinoma patients from healthy controls. The model built based on these four indicators outputs a classification label for each single cell. ( Indicates a positive result. (Indicating a negative result), and further calculate the patient-level risk score R, which is defined as:

[0060] Where n is the total number of single cells detected in the patient. To predict the number of positive cells; The threshold T is determined by maximizing the F1 score in the validation set, when At that time, the output of the assessment result of "high risk of individual positivity" is displayed. At that time, output the assessment result of "low risk of individual positivity" (refer to...). Figure 4 These four indicators distinguished papillary thyroid carcinoma patients from healthy controls with an AUC value of 0.92 (reference). Figure 5 This multi-level aggregation mechanism based on single-cell tag proportions can comprehensively reflect the abnormal proportions of cell populations at the individual level, reduce the impact of inter-cell heterogeneity on the overall discrimination results, and thus obtain a more stable patient-level risk assessment.

[0061] 2: Data Preprocessing and Filtering Patients with PTC were selected from the data matrix in section 1 and divided into a non-lymph node metastasis group (NNM, n=20) and a lymph node metastasis group (LNM, n=33) based on the cervical lymph node status. The following preprocessing and feature screening process was performed on the single-cell metabolic data matrix: 1. Missing Value and Abnormal Signal Processing: Missing values ​​in the feature matrix were filled with half of the minimum value. An isolated forest algorithm was used to identify and remove abnormal cell signals, with a contamination rate of 0.02 and a removal rate of approximately 2%.

[0062] 2. Feature standardization and distribution correction: In accordance with the experimental procedure, log transformation and Pareto scaling were performed on all metabolic features to balance different feature dimensions and reduce the influence of high variance peaks, thus ensuring the stability of the modeling process.

[0063] 3. Feature Selection and Optimization: Mutual Information was used to calculate the information gain between each feature and the disease label, initially retaining the top 1000 features. The Variance Inflation Factor (VIF) was calculated for the retained features, with a threshold of 10 to eliminate highly collinear features. The Mann-Whitney U test was used to screen for significantly different features (p<0.05). The resulting candidate features were then input into a Genetic Algorithm (GA) for feature subset optimization. The GA used the cross-validation AUC (Area Under the Receiver Operating Characteristic) as the fitness function, with a population size of 35, retaining the top 10% of highly fit individuals, and a mutation rate of 0.1. After approximately 40 generations, the AUC stabilized at 0.79 (reference). Figure 6 and Figure 7 This indicates that the feature subset converges and the model has the best discriminative performance.

[0064] The method of this invention can be extended to the analysis of peripheral blood single-cell metabolic data in other tumors (such as breast cancer, lung cancer, etc.) or immune system diseases. The detection platform is not limited to MicroCyESI-MS, but can be replaced by mass spectrometry systems with single-cell resolution such as matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS) and secondary ion mass spectrometry (SIMS).

[0065] This embodiment provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for assessing the risk of thyroid cancer and cervical lymph node metastasis.

[0066] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for assessing the risk of thyroid cancer and cervical lymph node metastasis, characterized in that, Includes the following steps: S1: Acquire single-cell metabolomics signal data from multiple cells, wherein the single-cell metabolomics signal data consists of mass spectrometry characteristic peaks generated by all metabolites; S2: Perform matrix transformation on the single-cell metabolomics signal data to construct a two-dimensional matrix of single-cell metabolomics signal intensity; S3: Based on patient identification information, hierarchical aggregation is performed on the two-dimensional matrix of single-cell metabolomics signal intensity, and a patient-level metabolic dataset is constructed by combining patient category labels; the patient identification information is used to associate all single-cell data from the same patient, the hierarchical aggregation is to statistically integrate all single-cell metabolomics signal data corresponding to the same patient to obtain patient-level metabolic feature data, and the patient-level metabolic dataset is a data matrix containing metabolic feature data of all patients and corresponding patient category labels; S4: Preprocess the patient-level metabolic dataset to obtain a preprocessed patient-level metabolic dataset; S5: Perform feature screening on the preprocessed patient-level metabolic dataset to obtain key metabolic features; S6: Perform hierarchical summary prediction based on the key metabolic characteristics to obtain individual risk assessment results.

2. The method for assessing the risk of thyroid cancer and cervical lymph node metastasis according to claim 1, characterized in that, The patient-level metabolic dataset includes a two-dimensional matrix of cells and metabolic features for each patient. The two-dimensional matrix is ​​arranged with cell numbers as rows and metabolic features as columns, and the matrix elements are the mass spectrometry signal intensities of each metabolic feature under the corresponding cell number.

3. The method for assessing the risk of thyroid cancer and cervical lymph node metastasis according to claim 1, characterized in that, The feature selection steps include: High-information-content candidate features are selected from the preprocessed patient-level metabolic dataset using an information content filtering method. The variance inflation factor is calculated based on the high-information-content candidate features, and high-collinearity features are removed based on the variance inflation factor to obtain the candidate features after removing high-collinearity features. Based on the candidate features after removing highly collinear features, nonparametric tests or equivalent significance detection methods are used to screen out differential features and obtain key metabolic features.

4. The method for assessing the risk of thyroid cancer and cervical lymph node metastasis according to claim 1, characterized in that, The key metabolic features include at least one of the following metabolic signal ratios: the ratio of the sum of phosphatidylcholine and sphingomyelin signal intensities to the signal intensity of 10-pentadien-1-ol, the ratio of sphingosine base signal intensity to the sum of ether-bonded phosphatidic acid and phosphatidylinositol signal intensities, the ratio of the sum of phospholipid metabolite signal intensities to the sum of sterol metabolite signal intensities, and the ratio of the sum of saturated lipid signal intensities to the sum of polyunsaturated lipid signal intensities.

5. The method for assessing the risk of thyroid cancer and cervical lymph node metastasis according to claim 1, characterized in that, The hierarchical aggregation prediction includes: Based on the key metabolic features, cell-level prediction results are obtained using a cell-level classification model. Based on the cell-level prediction results, a hierarchical summary is performed to obtain a patient-level disease risk score. Individual risk assessment results are obtained based on the patient-level disease risk score and the set threshold.

6. The method for assessing the risk of thyroid cancer and cervical lymph node metastasis according to claim 5, characterized in that, The cell-level classification model is a nonlinear classifier based on an ensemble learning algorithm. The optimal parameters of the nonlinear classifier are determined through cross-validation and Bayesian optimization.

7. The method for assessing the risk of thyroid cancer and cervical lymph node metastasis according to claim 5, characterized in that, The patient-level disease risk score includes a positive rate, and the formula for calculating the positive rate is as follows: , in, The positive rate; To predict the number of single cells that will test positive; The number of single cells predicted to be negative.

8. The method for assessing the risk of thyroid cancer and cervical lymph node metastasis according to claim 5, characterized in that, The formula for calculating the patient-level disease risk score is as follows: , in, Assess the disease risk level of patients. This is an aggregation function used to reflect the probability distribution characteristics of single cells; For the first The positive predictive probability of a single cell.

9. The method for assessing the risk of thyroid cancer and cervical lymph node metastasis according to claim 1, characterized in that, The method for assessing the risk of thyroid cancer and cervical lymph node metastasis also includes outputting an analysis report that includes individual risk assessment results and model performance indicators.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the risk assessment method for thyroid cancer and cervical lymph node metastasis as described in any one of claims 1-9.