Metabolism-oriented clustering-based tumor-associated macrophage metabolism typing method
By constructing a core set of metabolic genes and optimizing their weights using a metabolism-guided clustering method, the problems of insufficient algorithm robustness and automation in TAMs metabolic typing were solved, achieving more accurate tumor-associated macrophage typing and providing a basis for clinical intervention and treatment strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies for metabolic typing of tumor-associated macrophages (TAMs) suffer from poor algorithm robustness, excessive human intervention, and a lack of data-driven automated frameworks. This results in insufficient accuracy and reliability of typing results, difficulty in handling complex metabolic data and high-dimensional single-cell data, and impacts clinical decision-making.
A metabolism-guided clustering approach was adopted. By constructing a core set of metabolic genes, weight optimization and entropy change assessment were performed to generate tumor-associated macrophage metabolic typing results. KNN entropy estimation and loss function were used to optimize the weights. Gaussian noise was combined to simulate the tumor microenvironment. Community detection algorithm was used for clustering to generate a refined metabolic typing.
It improves the accuracy and robustness of TAM metabolic typing, enabling more precise identification of subgroups, applicable to multiple tumor types, providing clinical intervention targets and strategies, and enhancing the sensitivity to metabolic characteristics and the interpretability of typing results.
Smart Images

Figure CN121768477A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical artificial intelligence, specifically to a method for metabolic typing of tumor-associated macrophages based on metabolism-guided clustering. Background Technology
[0002] Tumor-associated macrophages (TAMs) are key components of the tumor microenvironment, influencing tumor development, metastasis, and immune evasion by regulating metabolic pathways. In cancer research, particularly for cholangiocarcinoma or other solid tumors, TAM metabolic profiling has become an important tool for assessing tumor prognosis and personalized treatment.
[0003] Traditionally, the classification of TAMs has primarily relied on inflammatory or immunosuppressive phenotypes, such as the M1 / M2 binary model. However, these methods fail to adequately capture the complexity of metabolic pathways, such as glycolysis, oxidative phosphorylation, and purine metabolism, which are closely related to the functional polarization of TAMs. Existing research indicates that metabolic reprogramming not only defines cellular energy states but also regulates immunosuppressive mechanisms and treatment resistance. For example, current research has found that purine-dominated TAM subsets are associated with poor responses to immune checkpoint blockade. Metabolic analysis of TAMs mainly relies on gene expression profiling and metabolomics data, but existing techniques have significant limitations in clustering and typing, leading to insufficient accuracy and reliability. Conventional clustering techniques often rely on broad transcriptomic patterns or predefined markers, resulting in insufficient resolution of subtle metabolic adaptations.
[0004] In existing technologies, metabolic typing of TAMs typically employs traditional clustering algorithms, such as k-means or hierarchical clustering. These methods group cell types based on simple distance metrics (e.g., Euclidean distance). However, these algorithms often overlook the complexity and nonlinear relationships of metabolic networks, making the typing results highly susceptible to data noise, sample heterogeneity, and parameter selection. For example, when processing high-throughput sequencing data, traditional clustering is prone to overfitting, especially in scenarios where metabolic pathways (such as glycolysis, lipid metabolism, and amino acid metabolism) are intertwined. This not only reduces the accuracy of typing but may also lead to biases in clinical decision-making. In high-dimensional single-cell data, prioritizing the processing of metabolic genes presents challenges, as traditional methods may overlook the intersections of metabolic and functional pathways, such as regulatory factors like Trem2 or HIF1α.
[0005] Another key issue is that metabolomics data often exhibit high dimensionality and sparsity. Traditional algorithms fail to effectively integrate metabolic network topology and gene expression information, resulting in a lack of interpretability for metabolic-related functions in the typing results. For example, in cholangiocarcinoma research, the M1 / M2 polarization state of TAMs is closely related to metabolic reprogramming, but existing methods struggle to accurately distinguish these subtypes. Existing metabolic typing techniques are often limited to laboratory environments and lack the ability to mine and interpret high-throughput single-cell data. Commonly used clustering algorithms in the literature are overly sensitive to parameters (e.g., the choice of k value significantly affects clustering results) and rely heavily on penalty terms (e.g., regularization coefficients), further emphasizing the need for more robust optimization strategies. These limitations slow the progress of cancer research and increase medical costs and patient risks.
[0006] In summary, the shortcomings of existing TAMs metabolic typing techniques are mainly reflected in: (1) poor algorithm robustness, which cannot effectively handle complex metabolic data; (2) excessive human subjective intervention, resulting in low reproducibility of results; and (3) lack of data-driven automated framework, making it difficult to achieve clinical translation. Therefore, there is an urgent need for a new method that can improve the accuracy, objectivity and practicality of typing through metabolic-guided clustering and advanced optimization techniques. Summary of the Invention
[0007] The purpose of this invention is to provide a metabolic typing method for tumor-associated macrophages (TAMs) based on metabolism-guided clustering. This method aims to construct an automated, data-driven process that enhances the sensitivity to metabolic characteristics through weight optimization and entropy change assessment, thereby providing more accurate subgroup identification for tumor treatment and enabling refined metabolic typing of TAMs. This method is not only applicable to TAMs but can also be extended to other tumor-associated cell types, filling a gap in existing technologies, thus deepening our understanding of the tumor microenvironment and providing new targets and strategies for clinical intervention.
[0008] To achieve the above-mentioned objectives, an embodiment provides a method for metabolic typing of tumor-associated macrophages based on metabolism-guided clustering, comprising the following steps: Collect and organize single-cell RNA sequencing datasets for analysis, and identify and extract tumor-associated macrophage populations from the organized single-cell RNA sequencing datasets; Based on the extracted tumor-associated macrophage population, a core metabolic gene set is constructed, and then the weights of the core metabolic genes are allocated and optimized. Based on the optimized weights, metabolic-guided clustering is performed to generate tumor-associated macrophage metabolic typing results and output the typing categories.
[0009] In one embodiment, the Seurat objects in the single-cell RNA sequencing dataset undergo preliminary quality control to filter out low-quality cells and low-expression genes; wherein, low-quality cells are damaged, dead, or uncaptured cells; and low-expression genes are genes that are expressed at extremely low levels or undetectable in most cells. Based on the filtered Seurat objects, the NormalizeData function was used for standardization, the CPM method was used for gene expression level normalization, the ScaleData function was used for data scaling, principal component analysis was used for dimensionality reduction, and the top 50 principal components were selected for metabolism-guided clustering analysis.
[0010] Furthermore, the indicators for identifying low-quality cells include: total number of UMIs, gene book, mitochondrial gene ratio, and ribosomal gene ratio; the indicators for identifying low-expression genes include: number of expressing cells and average expression level.
[0011] In one embodiment, identifying and extracting a tumor-associated macrophage population from the processed single-cell RNA sequencing dataset includes: Based on known cell surface markers or cell type-specific gene expression profiles, cell type annotation is performed on cells in the processed single-cell RNA sequencing dataset to identify and extract tumor-associated macrophage populations.
[0012] In one embodiment, the construction of a core metabolic gene set based on the extracted tumor-associated macrophage population includes: Based on the extracted tumor-associated macrophage population, a core set of metabolic genes was constructed by focusing on metabolism-related genes through database queries, literature searches, and gene list integration. Specifically, the database queries involved consulting one or more metabolic pathway databases, such as KEGG and Reactome; the literature searches involved screening for metabolism-related genes expressed in the tumor-associated macrophage population and functionally related to it, based on existing literature research; and the gene list integration involved combining the results of the database queries and literature searches to construct a list containing the core metabolism-related genes, which included information on the target genes and their related metabolic pathways.
[0013] In one embodiment, the core metabolic gene set includes: It covers key metabolic pathways closely related to the function of tumor-associated macrophage populations, including glycolysis, oxidative phosphorylation, lipid metabolism, amino acid metabolism, and purine and pyrimidine metabolism. And transcription factors that regulate key metabolic pathways.
[0014] In one example, the weight allocation and optimization of core metabolic genes includes: Weighting factors were assigned to the expression values of core metabolic genes, and cluster entropy was calculated using KNN entropy estimation. And establish a loss function to find the cluster entropy that makes the clustering entropy... The optimal weighting factor, serving as the optimal weight for core metabolic genes, is calculated using the following formula: , , in, This is the entropy calculation term, which is calculated using the k-nearest neighbor algorithm and is used to measure the weight factor. The disorder of the clustering results The value range is [0.2, 20]. For the first Local conditional entropy of a cell For the number of neighbors, The value range is [3, 15]. To interact with cells The normalization factor related to the density of the local region. Total number of cells; The coefficient for the penalty term; The penalty term is the L2 regularization term. The entropy value is the term used for entropy calculation.
[0015] In one instance, the entropy calculation term is obtained in the following way: In the dimensionality-reduced cell space, a k-nearest neighbor graph is constructed using the FindNeighbors function. For each cell, the probability distribution of its k neighbors in its respective cluster is calculated. ; The local entropy value of the cell was calculated based on the entropy formula. We calculate the entropy by weighting or averaging the local entropy values of all cells.
[0016] This invention calculates cluster entropy for global evaluation of clustering performance and combines it with regularization terms for automatic parameter optimization (such as...). The value range is [0.01, 0.1] to improve robustness. Cluster entropy, as an evaluation index of clustering quality, helps to optimize weights. and Furthermore, the entropy value is calculated for each cell, reflecting the clustering distribution of its local neighbors, by traversing the weights. (like The range of values is [0.2, 20]) to minimize the loss function, thereby finding the optimal weight to minimize the cluster entropy, achieving robust capture of metabolic heterogeneity, and ensuring better subgroup separation.
[0017] In one embodiment, the metabolic-guided clustering based on optimized weights generates tumor-associated macrophage metabolic typing results, outputting typing categories, including: Extract metabolic feature vectors containing glucose metabolism indicators and lipid metabolism indicators; Based on the optimized weights, principal component analysis is used to reduce the dimensionality of the weighted gene expression data. Then, the FindNeighbors function is used to construct the nearest neighbor graph of cells based on the cell distance in the dimensionality-reduced space. Based on the nearest neighbor graph and community detection algorithm, cells are divided into different clusters. The silhouette coefficient is used to evaluate the clustering quality and help determine the optimal clustering resolution. The tumor-associated macrophage metabolic typing results are generated and the typing categories are output.
[0018] In one embodiment, when using silhouette coefficients to evaluate clustering quality, the penalty term coefficient is automatically selected using a Bayesian optimization algorithm. , The value range is [0.01, 0.1].
[0019] In one embodiment, constructing the core metabolic gene set further includes enhancing the metabolic data to simulate variations in the tumor microenvironment, wherein the enhancement operation includes applying Gaussian noise to the metabolic data, wherein the noise standard deviation is... They also performed metabolic pathway simulations to model different stages of tumor development.
[0020] In one embodiment, the output classification categories are visualized, and the visualization types include: UMAP or t-SNE plots, heatmaps, violin plots, and dot plots; The UMAP or t-SNE diagram is used to reduce the dimension of clustered cells to a two-dimensional space and visualize them, reflecting the similarity between cells; The heatmap is used to display the expression levels of important genes in different clusters; The violin plot and dot plot are used to show the expression distribution of a single gene in different clusters.
[0021] Compared with the prior art, the beneficial effects of the present invention include at least the following: The metabolic typing method for tumor-associated macrophages (TAMs) based on metabolism-guided clustering provided by this invention eliminates technical bias by collecting and organizing single-cell RNA sequencing datasets for analysis, identifying and extracting TAM populations from them; constructing a core metabolic gene set based on the TAM population, and then weighting and optimizing the core metabolic genes to enhance their contribution to cluster analysis; performing metabolism-guided clustering based on the optimized weights to generate metabolic typing results, thereby enhancing the expression characteristics of metabolic genes and more accurately distinguishing TAM subpopulations; verifying the typing results through GSEA enrichment analysis or GO enrichment analysis, and outputting the typing categories. This method effectively overcomes the challenges faced by traditional methods in processing high-dimensional single-cell data, and is not only applicable to TAMs but can also be extended to other tumor-associated cell types, filling a gap in existing technologies. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0023] Figure 1 This is a flowchart illustrating the metabolic typing method for tumor-associated macrophages based on metabolism-guided clustering provided in an embodiment of the present invention.
[0024] Figure 2 This is a system framework diagram for tumor-associated macrophage metabolic typing provided in an embodiment of the present invention.
[0025] Figure 3 This is a schematic diagram illustrating the weight allocation of core metabolic genes provided in an embodiment of the present invention.
[0026] Figure 4 The loss curve is provided for an embodiment of the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the following description is provided in conjunction with the accompanying drawings and... The embodiments further illustrate the present invention in detail. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of the invention.
[0028] The examples detail the specific process of tumor-associated macrophage metabolic typing based on metabolism-guided clustering, aiming to provide a reproducible cell classification algorithm and guide subsequent research on tumor metabolism-related cell subpopulations. Figure 1 As shown, the present invention proposes a method for metabolic typing of tumor-associated macrophages based on metabolism-guided clustering, comprising the following steps: S1. Collect and organize the single-cell RNA sequencing dataset for analysis, and identify and extract tumor-associated macrophage populations from the organized single-cell RNA sequencing dataset.
[0029] Combination Figure 2 The system framework for tumor-associated macrophage metabolic typing, as shown, includes data input, weight optimization, and clustering result output. First, for data preparation and processing, this method requires collecting and organizing single-cell RNA sequencing (scRNA-seq) datasets for analysis. As an example, this embodiment uses an scRNA-seq dataset from a mouse tumor model. The preprocessed Seurat object serves as the input for analysis, containing a gene expression matrix that has undergone preliminary quality control, standardization, and dimensionality reduction. The preprocessing steps include quality control of the raw scRNA-seq data using the Seurat package, filtering for low-quality cells and low-expression genes. Low-quality cells are damaged, dead, or uncaptured cells; low-expression genes are genes with extremely low expression levels or undetectable in most cells. The indicators for identifying low-quality cells include: total UMI count, detected gene book, mitochondrial gene ratio, and ribosomal gene ratio. The indicators for identifying low-expression genes include: number of expressing cells and average expression level.
[0030] Then, the data was standardized using the NormalizeData function, gene expression levels were normalized using the CPM method, and the data was scaled using the ScaleData function to eliminate technical differences between cells. Principal Component Analysis (PCA) was used for dimensionality reduction, and the top 50 principal components (PCs) were selected for subsequent cluster analysis. Finally, cell type annotation was performed on the cells in the dataset based on known cell surface markers or cell type-specific gene expression profiles to identify and extract TAMs (Transformer Ingredients and Metrics), such as by screening based on the expression levels of marker genes like CD68 and CD163.
[0031] S2. Based on the extracted tumor-associated macrophage population, a core metabolic gene set is constructed. Then, the weights of these core metabolic genes are assigned and optimized. Metabolic-guided clustering is performed based on the optimized weights to generate tumor-associated macrophage metabolic typing results. The typing results are validated using GSEA enrichment analysis or GO enrichment analysis, and the typing categories are output. Details are as follows: First, a core set of metabolic genes was constructed to cover key metabolic pathways closely related to the function of TAMs. This construction process included database searching, literature retrieval, and gene list integration. Database searching involved consulting metabolic pathway databases such as KEGG and Reactome to screen for genes related to glycolysis, oxidative phosphorylation, lipid metabolism, amino acid metabolism, and purine and pyrimidine metabolism. Literature retrieval involved combining literature research to screen for metabolism-related genes expressed in TAMs and associated with their function. Gene list integration involved combining the results of database searching and literature retrieval to construct a list containing core metabolism-related genes, including target genes and their associated metabolic pathway information.
[0032] Next, the weighting and optimization of metabolic genes are performed. First, the expression values of core metabolic genes are multiplied by a weighting factor. The value of the weighting factor can be adjusted based on prior experience, for example, set between 0.2 and 20. KNN entropy estimation is then used to calculate the clustering entropy. And establish a loss function to find the cluster entropy that makes the clustering entropy... The optimal weighting factor, serving as the optimal weight for core metabolic genes, is calculated using the following formula: , , in, This is the entropy calculation term, which is calculated using the k-nearest neighbor algorithm and is used to measure the weight factor. The disorder of the clustering results The value range is [0.2, 20]. For the first Local conditional entropy of a cell For the number of neighbors, To interact with cells The normalization factor related to the density of the local region. Total number of cells; The coefficient for the penalty term; The penalty term, also known as the L2 regularization term, consists of entropy and weight regularization terms, where... The entropy value is the entropy value of the entropy calculation term based on the k-NN algorithm. This is the penalty term coefficient (e.g., 0.005), used to control the impact of the weights on the loss function. By iterating through different weight values (e.g., 0.2, 0.5, 1, 2, 4, 6, 8, 10, 15, 20), the loss corresponding to each weight value is calculated, and the loss curve is plotted. The weight that minimizes the loss function is selected as the optimal weight. Figure 3 By allocating weights, the amplification effect of pathways such as purine metabolism is highlighted.
[0033] In the cell space after PCA dimensionality reduction, the FindNeighbors function is used to construct a k-nearest neighbor graph for each cell, where k can be set between 3 and 15. For each cell, the probability distribution of its k neighbors within its cluster is calculated, and then the entropy value is calculated using the entropy formula: ,in The local entropy value. For cells Among the neighbors, the one who belongs to the first The proportion of cells in each cluster. Optimize weights. The calculation process includes: initialization The learning rate is set to 0.5 and adjusted iteratively via gradient descent, where the loss function is the entropy difference. The number of iterations does not exceed 1000, and the learning rate is dynamically adjusted from 0.001 to 0.01, adaptively based on the descent rate of the evaluation exponent.
[0034] In mouse data testing, weight 6 showed a minimum value in the loss curve, indicating that this weight minimizes the entropy of the clustering results and achieves the best subgroup separation effect.
[0035] After weight allocation and optimization, cluster analysis was performed. Metabolic feature vectors containing glucose and lipid metabolism indicators were extracted. Dimensionality reduction of the weighted gene expression data was performed using PCA. Then, the FindNeighbors function was used to construct a nearest neighbor graph of cells based on cell distances in the PCA space. Using the FindClusters function, cells were divided into different clusters based on the nearest neighbor graph and the Louvain algorithm (or Leiden algorithm). The clustering resolution could be adjusted based on experimental data; for example, the initial resolution could be set to 0.5. The Silhouette Score was used to evaluate clustering quality, ensuring that score fluctuations did not exceed 10%. The Silhouette Score decreased sharply when w∈[0.7,0.8] and showed robustness to values of k (5≤k≤15), but... (0.01≤ ≤0.1) is highly sensitive, therefore it is automatically selected using the Bayesian optimization algorithm. To maximize robustness, the optimal clustering resolution is determined, tumor-associated macrophage metabolic typing results are generated, and the typing categories are output.
[0036] To better understand the metabolic heterogeneity of TAMs, the clustering results need to be visualized, including UMAP or t-SNE dimensionality reduction, heatmaps, violin plots, and dot plots.
[0037] Clustered cells are reduced to a two-dimensional space and visualized. Cell positions on UMAP or t-SNE plots reflect their similarity. Heatmaps are generated to show the expression levels of important genes in different clusters, and F1-Scores are calculated as performance metrics, for example, to demonstrate the expression differences of glycolysis-related genes in different clusters. Violin plots or dot plots are used to show the expression distribution of individual genes in different clusters, for example, showing the high expression of the APRT gene in the TREM2+ cluster. The irGSEA.score is used to calculate metabolic pathway enrichment scores to assess the enrichment of metabolic pathways in different clusters, and trajectory inference tools (such as Monocle or Slingshot) are used to analyze cell differentiation trajectories or state transitions between different clusters. The performance metric calculation formula is: where TP, FP, and FN are true positives, false positives, and false negatives, respectively, and the threshold is dynamically set based on cross-validation. Clustering optimization employs a distributed computing framework, such as Apache Spark, supporting parallel processing of multiple samples; the KNN entropy estimation module supports GPU computation to handle large-scale data, with computation time not exceeding 5 seconds per batch. It integrates machine learning libraries, such as Scikit-learn, to implement variations of the KNN algorithm and connects to clinical databases via an API. The output of genotyping results includes an alert function; alert thresholds are trained based on historical data, for example, activating when the proportion of M2-type TAMs exceeds 60%. When the genotyping results indicate high-risk TAMs, a notification is triggered, and suggested treatment options are provided.
[0038] This method can be extended to other tumor types, such as breast cancer. By adjusting the weights of metabolic features, the extension steps include metastasis learning, transferring data from TAMs to new tumor type data, pre-training weights, and accuracy calculation based on 10-fold cross-validation to ensure accuracy for k and Robustness.
[0039] To validate the biological significance of metabolic typing, it is necessary to combine it with other omics data and experimental verification. For example, GSEA can be used to assess the enrichment of metabolic pathways in different clusters. Figure 4By comparing the gene expression profiles of different clusters with known metabolic pathway gene sets, it is possible to identify which clusters are enriched for specific metabolic pathways. GO enrichment analysis is then used to analyze the biological processes enriched in each cluster. Association analyses are performed between different clusters and clinical outcomes (such as patient survival and treatment response). Furthermore, experimental validation is conducted to detect the expression of key metabolic markers in different clusters; for example, flow cytometry is used to detect the expression of proteins such as GLUT1, HK2, and ARG1 in different clusters. In a specific example, the scRNA-seq dataset A from a mouse tumor model was used, with its preprocessed Seurat object being Allobj. The list of metabolic genes was loaded from the MetaboRefGenes_Mouse.csv file, and a weight of 6 was experimentally determined to be optimal. Based on the weighted gene expression data, the Louvain algorithm was used for clustering with a clustering resolution of 0.5, and the RunUMAP function was used to reduce the dimensionality of the clustered cells to two-dimensional space for visualization. Experimental results showed that cells in the TREM2+ cluster expressed high levels of glycolysis-related proteins such as GLUT1 and HK2, and exhibited high ARG1 activity. This method can be used to identify TAMs subsets associated with tumor progression, immunosuppression, and treatment resistance, providing a basis for developing targeted therapeutic strategies against specific TAM metabolic subsets and improving the efficacy of immunotherapy.
[0040] The specific embodiments described above illustrate the technical solution and beneficial effects of the present invention in detail. It should be understood that the above description is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for metabolic subtyping of tumor-associated macrophages based on metabolic-oriented clustering, characterized by, The method comprises the following steps: Collect and organize a single-cell RNA sequencing data set for analysis, identify and extract a tumor-associated macrophage population from the organized single-cell RNA sequencing data set; Based on the extracted tumor-associated macrophage population, a core metabolic gene set is constructed, and weight distribution and optimization of the core metabolic genes are performed; based on the optimized weight, metabolic-oriented clustering is performed to generate a tumor-associated macrophage metabolic typing result, and the typing result is verified by GSEA enrichment analysis or GO enrichment analysis, and the typing category is output.
2. The method of claim 1, wherein the metabolic profiling of tumor-associated macrophages is based on metabolic-oriented clustering. The method of collecting and organizing a single-cell RNA sequencing data set for analysis comprises the following steps: Performing preliminary quality control on Seurat objects in the single-cell RNA sequencing data set, filtering low-quality cells and low-expression genes; wherein the low-quality cells are damaged, dead or failed to be successfully captured; the low-expression genes are genes with extremely low expression or undetectable expression in most cells; Based on the filtered Seurat objects, standardization is performed using the NormalizeData function, gene expression normalization is performed using the CPM method, then data scaling is performed using the ScaleData function, dimensionality reduction is performed using principal component analysis, and the first 50 principal components are selected for metabolic-oriented clustering analysis.
3. The method of claim 2, wherein the metabolic profiling of tumor-associated macrophages is based on metabolic-oriented clustering. The method of identifying and extracting a tumor-associated macrophage population from the organized single-cell RNA sequencing data set comprises the following steps: According to known cell surface markers or cell type-specific gene expression profiles, the cells in the organized single-cell RNA sequencing data set are annotated by cell type, and the tumor-associated macrophage population is identified and extracted.
4. The method of claim 1, wherein, The core metabolic gene set comprises: It covers key metabolic pathways such as glycolysis, oxidative phosphorylation, lipid metabolism, amino acid metabolism, and purine and pyrimidine metabolism that are closely related to the functions of the tumor-associated macrophage population; And transcription factors that regulate key metabolic pathways.
5. The method of claim 1, wherein, The weight distribution and optimization of the core metabolic genes comprise: The expression values of the core metabolic genes are assigned a weight factor, and a clustering entropy is calculated using KNN entropy estimation A loss function is established to find the optimal weight factor that makes the clustering entropy optimal, which is the optimal weight of the core metabolic genes, and the calculation formula is as follows: , , wherein, is an entropy computation term, which is computed by k-nearest neighbor algorithm, for measuring the confusion degree of the clustering result in the weight factor is a confusion degree of the clustering result, is in the range of [0.2, 20], is the local condition entropy of the th cell, is the number of neighbors, is in the range of [3, 15], is a normalization factor related to the density of the local area where the cell is located, is the total number of cells; is a penalty term coefficient; is a penalty term, i.e., an L2 regularization term, is an entropy value of the entropy computation term.
6. The method of claim 5, wherein the metabolic profiling of tumor-associated macrophages is based on metabolic-oriented clustering. The entropy calculation term is obtained by the following method: Construct k-nearest neighbor graph in reduced cell space using FindNeighbors function, for each cell, compute the distribution probability of its k neighbors in its assigned cluster ; and calculating a local entropy value of the cell based on an entropy formula The local entropy values of all cells are weighted or averaged to obtain an entropy calculation term, wherein, is the proportion of cells belonging to the th cluster among the neighboring cells of the cell.
7. The method of claim 1, wherein, The method of performing metabolic-oriented clustering based on the optimized weight to generate a tumor-associated macrophage metabolic typing result and output the typing category comprises the following steps: Extract a metabolic feature vector containing sugar metabolism indicators and lipid metabolism indicators; Based on the optimized weight, the dimensionality of the weighted gene expression data is reduced using principal component analysis, then the nearest neighbor graph of cells is constructed based on the cell distance in the reduced space using the FindNeighbors function, the cells are divided into different clusters based on the nearest neighbor graph and community detection algorithm, and the clustering quality is evaluated using the silhouette coefficient to assist in determining the best clustering resolution, a tumor-associated macrophage metabolic typing result is generated, and the typing category is output.
8. The method for metabolic typing of tumor-associated macrophages based on metabolism-guided clustering according to claim 7, characterized in that, When using silhouette coefficients to evaluate clustering quality, the penalty term coefficient is automatically selected using a Bayesian optimization algorithm. , The value range is [0.01, 0.1].
9. The method of claim 1, wherein, In constructing the core metabolic gene set, the metabolic data is also augmented for modeling tumor microenvironment variations, where the augmentation operation includes applying Gaussian noise to the metabolic data, where the noise standard deviation and metabolic pathway simulation is performed to model different tumor stages.
10. The method of claim 1, wherein, The output typing category is visualized, and the visualization types include UMAP or t-SNE graphs, heat maps, violin plots, and dot plots; The UMAP or t-SNE graph is used to reduce the clustered cells to a two-dimensional space and visualize them, reflecting the similarity between cells; The heat map is used to display the expression levels of important genes in different clusters; The violin plot and dot plot are used to display the expression distribution of individual genes in different clusters.