Method, device and equipment for identifying cancer focus and storage medium

By collecting and analyzing single-cell transcriptome data from various cancer types, a cancer lesion prediction model and a hexagonal system were constructed, solving the problem of cancer lesion identification relying on reference data in existing technologies. This enabled the accurate identification of cancer lesions and the frontal zone of tumor invasion, and promoted the discovery of new targets for tumor immunotherapy.

CN121768489APending Publication Date: 2026-03-31WUHAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing spatial transcriptomics datasets rely on accurate reference data for cancer lesion identification, and obtaining high-quality normal tissue samples presents challenges, affecting the accuracy of analytical tools and their clinical application.

Method used

By collecting single-cell transcriptome data from various cancer types, filtering and dimensionality-reducing clustering were performed to extract epithelial cell characteristic gene sets, construct cancer lesion prediction models, and combine pathological annotations and malignancy scores to identify the spatial location of cancer lesions. Furthermore, the hexagonal system of spatial transcriptomics was used to identify the tumor invasion front region.

Benefits of technology

It enables rapid and accurate identification of cancerous lesion locations and the forefront of tumor invasion, providing a new perspective for the study of tumor immune escape mechanisms and promoting the discovery of immunotherapy targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768489A_ABST
    Figure CN121768489A_ABST
Patent Text Reader

Abstract

The invention discloses a cancer focus identification method, device and equipment and a storage medium, and relates to the technical field of biomedicine, and the method comprises the steps: collecting single cell transcriptome data of various cancer types, and carrying out the dimension reduction clustering of the filtered single cell transcriptome data; carrying out cell annotation on clustered cell groups, extracting epithelial cell groups, carrying out inter-group differential expression gene analysis, obtaining significantly up-regulated genes in each cancer type, and constructing a generic cancer cell characteristic gene set; carrying out pathological annotation on pathological sections of the spatial transcriptome data, and carrying out malignant scoring on each spot by utilizing a generic cancer cell characteristic gene set; and based on the pathological annotation label and the malignant score of each spot, constructing a cancer focus prediction model, and predicting a cancer focus spatial position in the spatial transcriptome data. According to the invention, the position of the cancerous lesion in the space transcriptome can be rapidly identified, and the tumor invasion leading edge area can be identified based on the cancerous lesion space positioning and the hexagonal system of the space transcriptomics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical technology, specifically to a method, apparatus, device, and storage medium for identifying cancerous lesions. Background Technology

[0002] Cancer is a complex disease that seriously threatens human life and health. It is a polygenic disease influenced by both heredity and environment. Its pathogenic mechanisms are complex and its treatment is difficult, making it a key focus of biomedical research. The difficulty in treating cancer stems from the proliferative and invasive capabilities and significant heterogeneity of cancer cells, accompanied by a highly complex microenvironment system. The tumor microenvironment (TME) is composed of various resident or infiltrating host cells (e.g., malignant cells, immune cells, and stromal cells) and non-cellular components (e.g., secretory factors, extracellular matrix proteins). All these components have important influences on tumor occurrence, progression, and metastasis, and are closely related to the response to immune checkpoint blockade (ICB) therapy. The tumor invasion frontier (TIF) is an ecological niche composed of the outermost layer of malignant cells in solid tumors and spatially adjacent non-malignant cells. Cellular interactions and molecular network regulation between the cancerous lesion and the tumor invasion frontier have a profound impact on TME remodeling.

[0003] Single-cell transcriptomics (scRNA-seq) technology has developed rapidly in recent years, deepening our understanding of tumorigenesis mechanisms. Through single-cell level analysis, researchers can reveal the heterogeneity and functional states of different cell types within the tumor microenvironment, thus gaining a deeper understanding of tumorigenesis, development, and immune escape mechanisms. However, a significant limitation of single-cell transcriptomics is the loss of spatial location information of cells. Cell differentiation states, functional characteristics, and intercellular signaling networks are highly dependent on their spatial distribution within tissues and their local microenvironment. The lack of spatial information makes it difficult for researchers to comprehensively analyze intercellular interactions within the tumor microenvironment and their impact on tumorigenesis and development, thereby limiting more in-depth research into tumorigenesis mechanisms.

[0004] To address this challenge, spatial transcriptomics emerged. The advent of spatial transcriptomics (ST) technology provides spatial location information for tumor development and progression, and enables the study of molecular characteristics across the boundaries between cancer and normal tissues. Compared to clustering methods in scRNA-seq analysis, ST sequencing requires a more comprehensive and holistic approach to assess gene expression, spatial location, and histological information, significantly enhancing researchers' understanding of cancer cell growth. The rapid development of ST technology has provided unprecedented spatial resolution for tumor research, deepening our understanding of tumor heterogeneity and the complexity of the microenvironment, and opening new avenues for developing more effective treatment strategies.

[0005] In spatial transcriptomics research, the definition and identification of pathological niches, particularly within tumor regions, are crucial. Identifying the spatial localization of tumor lesions within the tumor microenvironment (TME) allows for further analysis of interactions between spatially adjacent cells and the identification of potential therapeutic targets. While advancements in spatial transcriptomics (ST) have expanded the scope of tissue spatial landscape analysis, characterizing tumor spatial niches across cancer types remains challenging. Current analytical tools for identifying cancer lesions on ST datasets exhibit certain limitations, a significant common characteristic being their reliance on precise reference data. For instance, in copy number variation (CNV)-based analyses, normal cells must be specified as reference data to distinguish genomic differences between tumor and normal cells. However, obtaining high-quality normal tissue samples in real-world clinical settings often presents numerous challenges. Furthermore, the acquisition of normal tissue samples can be limited by individual patient variability, tissue heterogeneity, and sample preservation conditions, all of which affect the accuracy and reliability of reference data. These limitations not only hinder the widespread application of spatial transcriptomics analysis tools but also impede their translational potential in clinical diagnosis and treatment. Summary of the Invention

[0006] This application provides a method, apparatus, device, and storage medium for identifying cancerous lesions, which can quickly identify the location of cancerous lesions in the spatial transcriptome and identify the tumor invasion front region based on the spatial localization of cancerous lesions and the hexagonal system of spatial transcriptomics.

[0007] In a first aspect, embodiments of this application provide a method for identifying cancerous lesions, the method comprising: Single-cell transcriptome data from various cancer types were collected, filtered according to set criteria, and then dimensionality-reduced clustering was performed on the filtered single-cell transcriptome data. Cell annotation was performed on each cell group after dimensionality reduction and clustering, and the group annotated as epithelial cells was extracted. Differential gene expression analysis was performed on epithelial cells of different cancer types to obtain genes that were significantly upregulated in each cancer type, so as to construct a pan-cancer cell characteristic gene set. Pathological annotation was performed on the pathological sections of the spatial transcriptome data, and malignancy scores were assigned to each spot in the spatial transcriptome data using the pan-cancer cell characteristic gene set. Based on the pathological annotation tags and the malignancy score of each spot, a cancer lesion prediction model is constructed, and the spatial location of cancer lesions in the spatial transcriptome data is predicted according to the cancer lesion prediction model.

[0008] In conjunction with the first aspect, in one implementation, predicting the spatial location of cancer lesions in spatial transcriptome data based on the cancer lesion prediction model includes: Based on the constructed cancer lesion prediction model, variable genes in new spatial transcriptome data were detected, and multiple principal components were selected for analysis to perform unsupervised clustering. Based on the malignancy score of each input spot, a prediction is made. If the proportion of predicted malignant cells in a cluster is greater than a set value, the cluster is defined as a cancerous lesion. If the proportion of predicted non-malignant cells in a cluster exceeds the proportion of malignant cells and reaches a set value, the cluster is defined as a non-cancer lesion.

[0009] In conjunction with the first aspect, in one implementation, if the proportion of predicted malignant cells in a cluster is greater than 75%, the cluster is defined as a cancerous lesion; if the proportion of predicted non-malignant cells in a cluster exceeds the proportion of malignant cells by 75%, the cluster is defined as a non-cancer lesion.

[0010] In conjunction with the first aspect, in one implementation, it further includes: Based on the spatial location of cancer lesions and the hexagonal system of spatial transcriptomics, the tumor invasion front region was identified.

[0011] In conjunction with the first aspect, in one implementation, the hexagonal system based on the spatial location of the cancer lesion and spatial transcriptomics identifies the tumor invasion front region, including: Within the cancer lesions identified by the cancer lesion prediction model, it is determined whether the number of lesions around each lesion is 6; If a spot is surrounded by 6 spots, then the spot is located within the cancerous lesion area. If a spot is surrounded by less than 6 spots, then the spot is located in the area at the forefront of tumor invasion.

[0012] In conjunction with the first aspect, in one implementation, the filtering based on set criteria includes: Calculate the mitochondrial content in each cell, filter out cells with fewer than 300 and / or more than 6000 expressed genes, and retain only genes expressed in at least three cells.

[0013] In conjunction with the first aspect, in one embodiment, the differential gene expression analysis of epithelial cells from different cancer types to obtain genes that are significantly upregulated in each cancer type includes: Based on the screening rules: minimum expression ratio min.pct=0.1, change threshold logfc.threshold=0, and statistical significance level p<0.05, differential expression analysis algorithm was used to identify epithelial cell-specific upregulated genes in each cancer type.

[0014] Secondly, embodiments of this application provide a device for identifying cancerous lesions, the device for identifying cancerous lesions comprising: The acquisition module is used to collect single-cell transcriptome data of various cancer types, filter the data according to set criteria, and perform dimensionality reduction and clustering on the filtered single-cell transcriptome data. The module is used to annotate each cell group after dimensionality reduction and clustering, extract the annotated epithelial cell group, perform differential gene expression analysis between groups of epithelial cells of different cancer types, and obtain genes that are significantly upregulated in each cancer type to construct a pan-cancer cell characteristic gene set. The scoring module is used to perform pathological annotation on pathological sections of spatial transcriptome data and to score the malignancy of each spot in the spatial transcriptome data using the pan-cancer cell characteristic gene set. The prediction module constructs a cancer lesion prediction model based on the pathological annotation labels and the malignancy score of each spot, and predicts the spatial location of cancer lesions in the spatial transcriptome data according to the cancer lesion prediction model.

[0015] Thirdly, embodiments of this application provide a device for identifying cancerous lesions. The device for identifying cancerous lesions includes a processor, a memory, and a program for identifying cancerous lesions stored in the memory and executable by the processor. When the program for identifying cancerous lesions is executed by the processor, it implements the steps of the above-described method for identifying cancerous lesions.

[0016] Fourthly, a computer-readable storage medium storing a program for identifying cancerous lesions, wherein when the program for identifying cancerous lesions is executed by a processor, it implements the steps of the method for identifying cancerous lesions described above.

[0017] The beneficial effects of the technical solutions provided in this application include at least the following: The method for identifying cancer lesions in this application involves collecting single-cell transcriptome data from multiple cancer types, filtering the data according to set criteria, and performing dimensionality reduction clustering on the filtered single-cell transcriptome data. Cell annotation is performed on each cell group after dimensionality reduction clustering, extracting groups annotated as epithelial cells. Differential gene expression analysis is conducted between groups of epithelial cells from different cancer types to obtain genes that are significantly upregulated in each cancer type, thereby constructing a pan-cancer cell characteristic gene set. Pathological annotation is performed on pathological sections of the spatial transcriptome data. Using the pan-cancer cell characteristic gene set, each spot in the spatial transcriptome data is scored for malignancy. Based on the pathological annotation labels and the malignancy score of each spot, a cancer lesion prediction model is constructed. Based on the cancer lesion prediction model, the spatial location of cancer lesions in the spatial transcriptome data is predicted.

[0018] Therefore, this application combines single-cell and spatial transcriptomics sequencing technologies with machine learning algorithms to rapidly identify cancer lesion locations in the spatial transcriptome, aiding in precise localization and analysis. This provides a method for tissue regional comparison in spatial transcriptomics. Furthermore, based on the identified spatial localization of cancer lesions and the hexagonal system of spatial transcriptomics, this application further identifies the tumor invasion front region, providing a new perspective for the study of tumor immune escape mechanisms and promoting the discovery of new targets for immunotherapy. Attached Figure Description

[0019] Figure 1 This is a schematic flowchart of an embodiment of the method for identifying cancerous lesions according to this application; Figure 2 The data in this application are the number of scRNA-seq (single-cell transcriptome) data for different cancer types. Figure 3 This is a dimension-reduced clustering diagram of epithelial cell data for different cancer types in the embodiments of this application; Figure 4 This is a volcano map of differentially expressed genes in tumor epithelial cells of different cancer types in the embodiments of this application; Figure 5 This is a petal diagram of the intersection of differentially expressed genes in tumor epithelial cells of different cancer types in the embodiments of this application; Figure 6 This is a pathological annotation diagram of H&E sections in the tumor spatial transcriptomics data in the embodiments of this application; Figure 7 This is a spatial score of the pan-cancer cell characteristic gene set of tumor spatial transcriptomics data in the embodiments of this application; Figure 8 This refers to the location of the cancer lesion identified by the cancer lesion prediction model in the embodiments of this application; Figure 9 These are the AUC (area under the curve) values ​​of the cancer lesion prediction model in this application embodiment for cancer lesion prediction in 7 cancer types; Figure 10 This refers to the accuracy of the cancer lesion prediction model in this application embodiment in predicting cancer lesions in 7 cancer types; Figure 11 This is a comparison of the prediction accuracy of the cancer lesion prediction model in this application embodiment with existing tools (yellow represents this model). Figure 12 This is a schematic diagram of tumor invasion front region identification based on the spatial transcriptomics hexagonal system in the embodiments of this application; Figure 13 This is the result of tumor invasion front region identification in the embodiments of this application; Figure 14 This is a comparison of invasion feature scores for different regions identified in the embodiments of this application; Figure 15 This is a structural block diagram of an embodiment of the device for identifying cancerous lesions according to this application; Figure 16 This is a schematic diagram of the hardware structure of the device for identifying cancerous lesions involved in the embodiments of this application. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0022] In one aspect, embodiments of this application provide a method for identifying cancerous lesions.

[0023] In one embodiment, reference is made to Figure 1 , Figure 1 This is a schematic flowchart illustrating an embodiment of the method for identifying cancerous lesions according to this application. Figure 1 As shown, methods for identifying cancerous lesions include: S1. Collect single-cell transcriptome data of various cancer types, filter them according to set criteria, and perform dimensionality reduction and clustering on the filtered single-cell transcriptome data. S2. Perform cell annotation on each cell group after dimensionality reduction and clustering, extract the groups annotated as epithelial cells, perform differential gene expression analysis between groups of epithelial cells of different cancer types, obtain genes that are significantly upregulated in each cancer type, and construct a pan-cancer cell characteristic gene set. The purpose of steps S1 and S2 is to identify a pan-cancer cell characteristic gene set based on scRNA-seq. Specifically, this involves the following steps: (1) Cell quality control of scRNA-seq data This embodiment collected single-cell transcriptome data for 10 cancer types (data shown in Table 1). It is understood that the number of cancer types can be reasonably set as needed, and this embodiment does not impose any limitations.

[0024] Table 1: Details of Single-Cell Transcriptome Data for 10 Cancer Types

[0025] To remove low-quality cells, we first used scRNA-seq data expression matrices (an expression matrix is ​​a table where each row represents a gene and each column represents a cell; each element in the matrix represents the expression level of a specific gene in a specific cell) as input, and created objects using the R package Seurat. The mitochondrial content of each cell was calculated using the "PercentageFeatureSet" function. Then, the "subset" function was used to filter cells according to the following criteria: cells with fewer than 300 and / or more than 6000 expressed genes were removed, and only genes expressed in at least three cells were retained, resulting in 766,030 cells. See the results below. Figure 2 .

[0026] (2) Dimensionality reduction clustering of scRNA-seq data Next, the data was normalized using the "NormalizeData" function, and the top 3000 highly variable genes were detected using the "FindVariableFeatures" function. Then, based on these 3000 genes, the dimensionality of the scRNA-seq data was reduced using the "RunPCA" function, and 30 principal components were selected for subsequent analysis. Furthermore, to eliminate batch processing effects between samples, soft k-means clustering was performed using the R package Harmony. Subsequently, a KNN graph was constructed in the principal component space based on Euclidean distance using the "FindNeighbors" function, and the edge weights between any two cells were defined based on the shared overlap of adjacent regions. Finally, the "FindClusters" function was used to identify cell clusters based on the shared nearest neighbor (SNN) method, with a resolution set to 0.5.

[0027] (3) Cell annotation and epithelial cell extraction of scRNA-seq data The "FindAllMarkers" function was used to retrieve the specific genes expressed by each cell population. Next, cell type annotation was performed based on these specific genes; for example, cell populations highly expressing CD3D and CD3E were annotated as T cells, those highly expressing CD79A and MZB1 as B cells, and those highly expressing EPCAM and KRT19 as epithelial cells. Subsequently, the "subset" function was used to extract all epithelial cell populations for subsequent analysis. Results are shown below. Figure 3 .

[0028] (4) Differential analysis of epithelial cells based on scRNA-seq data For all extracted epithelial cells, the "FindMarkers" function was used to sequentially obtain the specific genes that were significantly upregulated in epithelial cells for each cancer type, with the following parameters: minimum expression ratio min.pct = 0.1, change threshold logfc.threshold = 0, and statistical significance level p < 0.05 (see results). Figure 4 Furthermore, intersection analysis identified 209 genes that were significantly upregulated across 10 cancer types, constructing a pan-cancer cell characteristic gene set. Results are shown below. Figure 5 .

[0029] S3. Pathological annotation of pathological sections of spatial transcriptome data, and malignancy score of each spot in spatial transcriptome data using the pan-cancer cell characteristic gene set. S4. Based on the pathological annotation labels and the malignancy score of each spot, construct a cancer lesion prediction model, and predict the spatial location of cancer lesions in the spatial transcriptome data according to the cancer lesion prediction model.

[0030] In steps S3 and S4, the main objective is to annotate spatial transcriptome data and construct cancer lesion prediction models. Specifically, this includes the following steps: (1) Pathological annotation of H&E sections from spatial transcriptome (ST) data This embodiment collects spatial transcriptome data for seven types of cancer (data shown in Table 2).

[0031] Table 2: Details of Spatial Transcriptome Data for 7 Cancer Types

[0032] To establish a gold standard for cancer lesion identification, the location of cancer lesions in each slide was selected based on the morphological characteristics of H&E-stained pathological sections using spatial transcriptome data. Results are shown below. Figure 6 .

[0033] (2) Spatial transcriptome data annotation of cancer lesions The raw expression matrix, H&E slices, and spatial coordinate information of the spatial transcriptome data were read using the R package Seurat, and Seurat objects were created. To ensure the integrity of the spatial regions, all detected spots were included in the downstream analysis. Subsequently, each spot in the spatial transcriptome was classified based on the pathological annotation of H&E to annotate the spatial location of the tumor lesion (Mal), and the remaining locations were defined as adjacent normal (nMal).

[0034] (3) Pan-oncogene set scoring of spatial transcriptome data Using the aforementioned pan-cancer cell characteristic gene set, the VISION algorithm was employed to score malignancy based on the raw count of each spot in each sample. This algorithm first requires a dimensionality-reduced space, performs PCA (Principal Component Analysis), and selects the top 30 PCA components. Then, a K-Nearest Neighbors (KNN) graph is constructed between cells, followed by calculation of the expression of the pan-cancer cell characteristic gene set to obtain a gene set score. The autocorrelation statistical method Geary-C is further used to calculate the consistency assessment between adjacent neighbors. This is primarily to account for the influence of sample-level metrics (UMI per cell), and the final score is corrected based on its mean and standard deviation. Subsequently, the malignancy score is added to the Seurat object of the spatial transcriptome using the "AddMetaData" function, and the score is mapped to spatial tissue slices using the "SpatialFeaturePlot" function. The results are shown below. Figure 7 .

[0035] (4) Construction of cancer lesion prediction model Using the leave-one-out method, one sample from each cancer type was reserved for subsequent testing, while the remaining samples were used for model training (data shown in Table 3).

[0036] Table 3: Sample information for constructing cancer lesion prediction models

[0037] Using the malignancy score of the pan-cancer characteristic gene set (Intersection_gene_set) for each spot in spatial transcriptomics as the feature variable, a logistic regression-based machine learning method was employed. The logistic regression model was constructed using the "glm" function, specifying the use of a binomial distribution (family = binomial). The "summary" function was then called to view the detailed results of the model, allowing for analysis of regression coefficients, standard error, z-value, p-value, and other information to help interpret the model's fit. (Data is shown in Table 4.)

[0038] Table 4: Summary of Cancer Follicle Prediction Models

[0039] It is worth noting that the cancer lesion prediction model takes the malignancy score of the spots in (3) as input and outputs whether it is a tumor (malignant cell) in terms of 0 and 1. The model uses the logistic regression algorithm and the trained model can be used to predict tumor regions for new samples.

[0040] After constructing the cancer lesion prediction model, cancer lesions and invasion front regions can be identified from spatial transcriptome data. Specifically, this includes: (1) Prediction of cancer lesions using spatial transcriptome data Based on the constructed cancer lesion prediction model, it was applied to new spatial transcriptomic data, and then classified according to probability values ​​to achieve automatic prediction of the spatial location of cancer lesions (Mal, Malignant). The "SCTransform" function was further used for the standardization process of spatial transcriptomics data. The FindVariableFeatures function was used to detect the top 3000 highly variable genes, and 30 principal components were selected for subsequent analysis. The FindClusters function was used for unsupervised clustering. For each cluster result, the prediction result was corrected according to the following rules: if the proportion of predicted malignant cells in the cluster is greater than 75%, the cluster is defined as a cancer lesion (Mal); if the proportion of predicted non-malignant cells in the cluster is greater than 75%, the cluster is defined as a non-cancer lesion (nMal). The corrected prediction results showed high consistency with pathological annotations. Results are shown below. Figure 8 .

[0041] (2) Performance testing of tumor prediction models To evaluate the predictive performance of this model, spatial transcriptomics data of seven cancer types collected in this application were used for validation. By calculating the AUC (Area Under the Curve) values ​​of the model in different cancer types, it was found that the AUC of this model exceeded 0.9 in six cancer types, indicating that the model has extremely high predictive accuracy and reliability in these cancer types. Results are shown below. Figure 9 .

[0042] (3) Accuracy test of tumor prediction model To evaluate the accuracy of the tumor prediction model, spatial transcriptomics data for seven cancer types collected in this application were used as a validation set. These samples included pre-completed pathological annotation information. The differences between the prediction results and the pathological annotation results were then calculated, and the accuracy of the prediction results was statistically analyzed. The prediction results showed that the model achieved a prediction accuracy of over 90% in five of the cancer types. See the results below. Figure 10 .

[0043] (4) Comparison of multiple methods To evaluate the predictive performance of the tumor prediction model compared to existing tools, spatial transcriptomics data for seven cancer types described in this application were used as a reference to assess the predictive accuracy of different methods. Using pathologically confirmed cancerous lesion regions as the standard, the prediction results of our model (My model), Cotrazm, SpaCET, and Ikarus were compared across the spatial transcriptomics data for the seven cancer types, and the predictive accuracy of each method was further calculated. The results show that our model (My model) exhibits superior cancerous lesion prediction ability across all seven cancer types, with the best prediction results in four of them. See [link to results]. Figure 11 .

[0044] (5) Identification principle of the frontal region of tumor invasion Based on the spatial localization of cancer lesions and the hexagonal system characteristics of spatial transcriptomics, the tumor invasion front region was further identified.

[0045] The detailed process is as follows: ①The structural features of 10X Visium spatial transcriptomics are: each spot is surrounded by 6 spots, see example... Figure 12 As shown: ② Calculate whether the number of spots around each spot identified by the cancer lesion prediction model is 6; if the number of spots around a spot is equal to 6, then the spot is located within the cancer lesion area; if the number of spots around a spot is less than 6, then the spot is located in the tumor invision frontier (TIF). ③ Finally, each spatial transcriptome sample was divided into three regions: malignant (Mal), tumor invision frontier (TIF), and non-malignant (nMal). (6) Identification results of the tumor invasion front region Based on the spatial location of cancer lesions predicted by the cancer lesion prediction model, the identification of the tumor invasion front region was ultimately achieved. Results are shown in [link to results]. Figure 13 .

[0046] (7) The degree of invasion of the tumor in the leading edge region Based on the predicted tumor invasion front region, the degree of invasion in different regions was analyzed. Invasion scoring results showed that the invasion score in the tumor invasion front region was significantly higher than that in other regions. See results below. Figure 14 .

[0047] In summary, the method for identifying cancer lesions in this application involves collecting single-cell transcriptome data from multiple cancer types, filtering the data according to set criteria, and performing dimensionality reduction clustering on the filtered single-cell transcriptome data. Cell annotation is performed on each cell group after dimensionality reduction clustering, extracting groups annotated as epithelial cells. Differential gene expression analysis is conducted between groups of epithelial cells from different cancer types to obtain genes that are significantly upregulated in each cancer type, thereby constructing a pan-cancer cell characteristic gene set. Pathological annotation is performed on pathological sections of the spatial transcriptome data. Using the pan-cancer cell characteristic gene set, each spot in the spatial transcriptome data is scored for malignancy. Based on the pathological annotation labels and the malignancy score of each spot, a cancer lesion prediction model is constructed. Based on the cancer lesion prediction model, the spatial location of cancer lesions in the spatial transcriptome data is predicted.

[0048] Therefore, this application combines single-cell and spatial transcriptomics sequencing technologies with machine learning algorithms to rapidly identify cancer lesion locations in the spatial transcriptome, aiding in precise localization and analysis. This provides a method for tissue regional comparison in spatial transcriptomics. Furthermore, based on the identified spatial localization of cancer lesions and the hexagonal system of spatial transcriptomics, this application further identifies the tumor invasion front region, providing a new perspective for the study of tumor immune escape mechanisms and promoting the discovery of new targets for immunotherapy.

[0049] Secondly, embodiments of this application also provide a device for identifying cancerous lesions.

[0050] In one embodiment, reference is made to Figure 15 , Figure 15 This is a schematic diagram of the functional modules of an embodiment of the device for identifying cancerous lesions according to this application. Figure 15 As shown, the device for identifying cancerous lesions includes: an acquisition module, a construction module, a scoring module, and a prediction module.

[0051] The acquisition module is used to collect single-cell transcriptome data of various cancer types, filter the data according to set criteria, and perform dimensionality reduction and clustering on the filtered single-cell transcriptome data. The module is used to annotate each cell group after dimensionality reduction and clustering, extract the annotated epithelial cell group, perform differential gene expression analysis between groups of epithelial cells of different cancer types, and obtain genes that are significantly upregulated in each cancer type to construct a pan-cancer cell characteristic gene set. The scoring module is used to perform pathological annotation on pathological sections of spatial transcriptome data and to score the malignancy of each spot in the spatial transcriptome data using the pan-cancer cell characteristic gene set. The prediction module constructs a cancer lesion prediction model based on the pathological annotation labels and the malignancy score of each spot, and predicts the spatial location of cancer lesions in the spatial transcriptome data according to the cancer lesion prediction model.

[0052] Further, in one embodiment, the prediction module predicts the spatial location of cancer lesions in the spatial transcriptome data based on the cancer lesion prediction model, including: Based on the constructed cancer lesion prediction model, variable genes in new spatial transcriptome data were detected, and multiple principal components were selected for analysis to perform unsupervised clustering. Based on the malignancy score of each input spot, a prediction is made. If the proportion of predicted malignant cells in a cluster is greater than a set value, the cluster is defined as a cancerous lesion. If the proportion of predicted non-malignant cells in a cluster exceeds the proportion of malignant cells and reaches a set value, the cluster is defined as a non-cancer lesion.

[0053] Furthermore, in one embodiment, if the proportion of predicted malignant cells in a cluster is greater than 75%, the cluster is defined as a cancerous lesion; if the proportion of predicted non-malignant cells in a cluster exceeds the proportion of malignant cells by 75%, the cluster is defined as a non-cancer lesion.

[0054] Furthermore, in one embodiment, the prediction module is also used for: Based on the spatial location of cancer lesions and the hexagonal system of spatial transcriptomics, the tumor invasion front region was identified.

[0055] Furthermore, in one embodiment, the prediction module identifies the tumor invasion front region based on the spatial location of the cancer lesion and a hexagonal system of spatial transcriptomics, including: Within the cancer lesions identified by the cancer lesion prediction model, it is determined whether the number of lesions around each lesion is 6; If a spot is surrounded by 6 spots, then the spot is located within the cancerous lesion area. If a spot is surrounded by less than 6 spots, then the spot is located in the area at the forefront of tumor invasion.

[0056] Furthermore, in one embodiment, the acquisition module performs filtering according to set criteria, including: Calculate the mitochondrial content in each cell, filter out cells with fewer than 300 and / or more than 6000 expressed genes, and retain only genes expressed in at least three cells.

[0057] Furthermore, in one embodiment, the construction module performs differential gene expression analysis on epithelial cells of different cancer types to obtain genes that are significantly upregulated in each cancer type, including: Based on the screening rules: minimum expression ratio min.pct=0.1, change threshold logfc.threshold=0, and statistical significance level p<0.05, differential expression analysis algorithm was used to identify epithelial cell-specific upregulated genes in each cancer type.

[0058] The functions of each module in the above-mentioned device for identifying cancerous lesions correspond to the steps in the above-mentioned method embodiment for identifying cancerous lesions, and their functions and implementation processes will not be described in detail here.

[0059] Thirdly, embodiments of this application provide a device for identifying cancerous lesions. The device for identifying cancerous lesions can be a personal computer (PC), a laptop computer, a server, or other device with data processing capabilities.

[0060] Reference Figure 16 , Figure 16 This is a schematic diagram of the hardware structure of a device for identifying cancerous lesions involved in an embodiment of this application. In this embodiment, the device for identifying cancerous lesions may include a processor, a memory, a communication interface, and a communication bus.

[0061] The communication bus can be of any type and is used to interconnect the processor, memory, and communication interface.

[0062] Communication interfaces include input / output (I / O) interfaces, physical interfaces, and logical interfaces used for interconnecting internal devices within the cancer lesion identification device, as well as interfaces used for interconnecting the cancer lesion identification device with other devices (such as other computing devices or user equipment). Physical interfaces can be Ethernet interfaces, fiber optic interfaces, ATM interfaces, etc.; user equipment can be displays, keyboards, etc.

[0063] Memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0064] The processor can be a general-purpose processor, which can call a program for identifying cancer lesions stored in memory and execute the method for identifying cancer lesions provided in the embodiments of this application. For example, the general-purpose processor can be a central processing unit (CPU). The method executed when the program for identifying cancer lesions is called can be referred to in the various embodiments of the method for identifying cancer lesions in this application, and will not be repeated here.

[0065] Those skilled in the art will understand that Figure 16 The hardware structure shown does not constitute a limitation of this application and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0066] Fourthly, embodiments of this application also provide a readable storage medium.

[0067] The present application has a program for identifying cancerous lesions stored on a readable storage medium, wherein when the program for identifying cancerous lesions is executed by a processor, it implements the steps of the method for identifying cancerous lesions as described above.

[0068] The method implemented when the procedure for identifying cancerous lesions is executed can be referred to in various embodiments of the method for identifying cancerous lesions in this application, and will not be repeated here.

[0069] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0070] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.

[0071] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.

[0072] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.

[0073] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.

[0074] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.

[0075] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for identifying cancerous lesions, characterized in that, The method for identifying cancer foci comprises: Collecting single-cell transcriptome data of multiple cancer types, filtering according to a set standard, and performing dimension reduction clustering on the filtered single-cell transcriptome data; Performing cell annotation on each cell cluster after dimension reduction clustering, extracting clusters annotated as epithelial cells, performing inter-group differential expression gene analysis on epithelial cells of different cancer types, obtaining genes significantly up-regulated in each cancer type, and constructing a pan-cancer cell feature gene set; Performing pathological annotation on pathological sections of spatial transcriptome data, using the pan-cancer cell feature gene set to perform malignant scoring on each spot in the spatial transcriptome data; Based on the pathological annotation label and the malignant score of each spot, a cancer focus prediction model is constructed, and the spatial position of the cancer focus in the spatial transcriptome data is predicted according to the cancer focus prediction model.

2. The method of identifying a cancerous lesion of claim 1, wherein, The method for predicting the spatial position of the cancer focus in the spatial transcriptome data according to the cancer focus prediction model comprises: According to the constructed cancer focus prediction model, detect the variable genes of the new spatial transcriptome data, and select multiple principal components for analysis to perform unsupervised clustering; According to the input malignant score of each spot, if the proportion of predicted malignant cells in the cluster is greater than a set value, the cluster is defined as a cancer focus, and if the proportion of predicted non-malignant cells in the cluster exceeds the proportion of malignant cells by a set value, the cluster is defined as a non-cancer focus.

3. The method for identifying cancer foci according to claim 2, wherein: If the proportion of predicted malignant cells in the cluster is greater than 75%, the cluster is defined as a cancer focus, and if the proportion of predicted non-malignant cells in the cluster exceeds the proportion of malignant cells by 75%, the cluster is defined as a non-cancer focus.

4. The method of identifying a cancerous lesion of claim 2, wherein, Further comprising: Identifying the tumor invasion front area based on the spatial position of the cancer focus and the hexagonal system of spatial transcriptomics.

5. The method of identifying a cancerous lesion of claim 4, wherein, The method for identifying the tumor invasion front area based on the spatial position of the cancer focus and the hexagonal system of spatial transcriptomics comprises: Within the cancer focus identified by the cancer focus prediction model, determine whether the number of spots around each spot is 6; If the number of spots around a spot is equal to 6, the spot is located in the cancer focus area, and if the number of spots around a spot is less than 6, the spot is located in the tumor invasion front area.

6. The method of identifying a cancerous lesion of claim 1, wherein, The filtering according to the set standard comprises: Calculating the mitochondrial content in each cell, filtering out cells with less than 300 and / or more than 6,000 expressed genes, and only retaining genes expressed in at least three cells.

7. The method of identifying a cancerous lesion of claim 1 wherein, The inter-group differential expression gene analysis on epithelial cells of different cancer types to obtain genes significantly up-regulated in each cancer type comprises: Using a differential expression analysis algorithm to identify epithelial cell-specific up-regulated genes in each cancer type based on a screening rule: minimum expression ratio min.pct=0.1, change threshold logfc.threshold=0, and statistical significance level p<0.

05.

8. A device for identifying cancerous lesions, characterized in that, The device for identifying cancer foci comprises: The collection module is used for collecting single-cell transcriptome data of multiple cancer types, filtering according to a set standard, and dimensionally reducing and clustering the filtered single-cell transcriptome data; The construction module is used for cell annotation of each cell cluster after dimensional reduction and clustering, extraction of a cluster annotated as epithelial cells, inter-group differential expression gene analysis of epithelial cells of different cancer types, acquisition of genes significantly up-regulated in each cancer type, and construction of a pan-cancer cell feature gene set; The scoring module is used for pathological section annotation of spatial transcriptome data, malignant scoring of each spot in the spatial transcriptome data by using the pan-cancer cell feature gene set; The prediction module is used for constructing a cancer focus prediction model based on the pathological annotation label and the malignant score of each spot, and predicting the spatial position of a cancer focus in the spatial transcriptome data according to the cancer focus prediction model.

9. A device for identifying cancerous lesions, characterized in that, The device for identifying a cancer focus comprises a processor, a memory, and a program for identifying a cancer focus stored on the memory and executable by the processor, wherein the program for identifying a cancer focus, when executed by the processor, implements the steps of the method for identifying a cancer focus according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for identifying a cancer focus, wherein the program for identifying a cancer focus, when executed by a processor, implements the steps of the method for identifying a cancer focus according to any one of claims 1 to 7.