Cell infiltration inference method and system fusing go function annotation and ppi network information

By integrating GO functional annotations and PPI network information, a cell-cell function and physical interaction network was constructed, which solved the limitations of inferring cell components in the tumor microenvironment, enabled accurate characterization and inference of multiple cell types, and improved the accuracy and reliability of tumor research.

CN121075448BActive Publication Date: 2026-03-27GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies for inferring cell components in the tumor microenvironment have limitations. They cannot comprehensively construct cell atlases, ignore the infiltration characteristics of mesenchymal cells and stem cells, and fail to effectively integrate gene functional synergy relationships and protein interaction networks, resulting in inaccurate cell infiltration inference.

Method used

By integrating Gene Ontology (GO) functional annotation with protein-protein interaction (PPI) network information, a cell-cell functional association network and a physical interaction network are constructed. Cell infiltration scores are calculated through weighted fusion and restart walk algorithms, comprehensively considering the functional similarity and physical interactions between cells.

Benefits of technology

It enables precise characterization of multiple cell types in the tumor microenvironment, improves the accuracy and coverage of cell invasion inference, provides a more biologically interpretable cell relationship atlas, and provides a reliable basis for tumor research and precision medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075448B_ABST
    Figure CN121075448B_ABST
Patent Text Reader

Abstract

The application relates to a cell infiltration inference method and system fusing GO function annotation and PPI network information, and the method comprises the following steps: collecting gene expression data, GO function annotation data and PPI network data; constructing a cell-cell function correlation network and a cell-cell physical interaction network respectively; performing weighted fusion processing on the two networks to obtain a comprehensive cell relationship network; calculating a final cell infiltration score through a restart walk algorithm, and inferring the infiltration degree in a tumor microenvironment according to the final cell infiltration score. The application innovatively fuses GO function annotation information and PPI network data, comprehensively considers the functional similarity and physical or signal interaction between cells, enables the model to understand cell synergy from the biological pathway level and analyze cell direct interaction from the protein interaction level, avoids one-sidedness of a single perspective, and provides a more stereoscopic cognitive framework for tumor microenvironment analysis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of biological information, and particularly relates to a cell infiltration inference method and system fusing GO function annotation and PPI network information. BACKGROUND

[0002] Tumor Microenvironment (TME) is a complex ecosystem composed of tumor cells, immune cells, stromal cells, stem cells, vascular endothelial cells, cytokines and extracellular matrix. In recent years, studies have shown that TME plays a crucial role in the occurrence, development, metastasis, recurrence and response of immunotherapy and targeted therapy of tumors. Therefore, accurately depicting the cell composition, infiltration degree and cell-cell interaction in TME not only helps to understand the biological mechanism of tumors, but also provides important information for developing more precise treatment strategies.

[0003] At present, the inference of cell composition in TME mainly relies on the computational analysis method of bulk RNA-seq data. Traditional cell infiltration inference methods can be roughly divided into two categories: one is based on the deconvolution model method, such as CIBERSORT, EPIC, quanTIseq and TIMER, etc. These methods construct a reference expression matrix and use a linear mixing model to infer the proportion of different cell types in the sample; the other is based on the enrichment scoring method of marker gene set, such as ssGSEA, xCell and MCP-counter, which calculate the activity of each cell type in the sample by enrichment analysis.

[0004] Although the existing methods reveal the cell composition in TME to some extent, there are still many limitations. First, the description of traditional methods is often limited to immune cells, ignoring the infiltration characteristics of stromal cells, stem cells and their subgroups, and cannot construct a comprehensive cell atlas. Second, these methods often process gene expression characteristics independently, ignoring the functional synergy or regulatory relationship between genes, making it difficult to accurately distinguish cell subtypes with similar functions or co-expression characteristics. In addition, the deconvolution method has strong dependence on reference data sets, and has weak generalization ability in different platforms or tissue sources, making it difficult to adapt to tumor heterogeneity and cell heterogeneity. In addition, the enrichment scoring method based on marker genes is usually a normalized score, which lacks proportional interpretation and is difficult to compare directly between different samples.

[0005] To address these issues, researchers have attempted to incorporate functional information to assist cell infiltration inference in recent years. CITMIC is an innovative solution. CITMIC constructs a cell-GO bipartite graph based on Gene Ontology (GO) functional annotations, and calculates the activity of cell types in samples using a graph propagation algorithm. However, although the CITMIC method has significantly improved accuracy and cell coverage compared to traditional methods, there are still some limitations. First, although the GO functional annotations used by CITMIC can provide functional similarity between cells, they do not consider real interactions between cells in terms of physical or signal transduction. Second, GO information has certain static and redundant properties and cannot reflect molecular-level changes during tumor progression. In addition, CITMIC only constructs a cell relationship network based on functional module similarity, ignoring potential structural biological information such as protein-protein interaction (PPI) networks that regulate cell behavior. Therefore, the CITMIC method may have weak biological interpretation and insufficient recognition of local specificity.

[0006] In addition, different cell types in tumors often communicate and coordinate functions through protein-protein interaction (PPI) networks. PPI networks can directly describe the physical and signal connection patterns between cells and are widely used in disease mechanism research and target screening. However, few studies have combined PPI networks with cell infiltration inference. Therefore, how to integrate PPI network information into cell infiltration inference to reveal the real physical connection patterns between cells remains a pressing research problem.

[0007] In summary, there is currently a lack of a cell infiltration inference method that can simultaneously integrate functional information (such as GO) and structural information (such as PPI). Combining these two types of information, constructing a cell relationship network with better biological interpretation, and accurately depicting cell distribution and interaction in tumor tissues are important challenges in current tumor research and precision medicine. SUMMARY

[0008] To address the above-mentioned deficiencies of the prior art, the present application provides a cell infiltration inference method and system that integrates GO functional annotations and PPI network information, to construct a more realistic cell communication map and comprehensively analyze the status and functions of multiple cell types in TME, improving the characterization ability and prediction effect of tumor immune characteristics.

[0009] In a first aspect, the present application provides a cell infiltration inference method that integrates GO functional annotations and PPI network information, which comprises:

[0010] S1: collect a dataset, and perform preprocessing, wherein the dataset comprises gene expression data, GO function annotation data, and PPI network data;

[0011] S2: construct a cell-cell function correlation network according to the gene expression data and the GO function annotation data;

[0012] S3: construct a cell-cell physical interaction network according to the gene expression data and the PPI network data;

[0013] S4: perform weighted fusion processing on the cell-cell function correlation network and the cell-cell physical interaction network to obtain a comprehensive cell relationship network;

[0014] S5: based on the comprehensive cell relationship network, calculate a final cell infiltration score by a restart walk algorithm, and infer the infiltration degree in a tumor microenvironment according to the final cell infiltration score.

[0015] The application innovatively fuses Gene Ontology (GO) function annotation information and protein-protein interaction (PPI) network data, comprehensively considers the functional similarity and physical or signal interaction between cells by constructing a weighted fusion of a cell-cell function correlation network and a cell-cell physical interaction network, and can cover various cell types, including immune cells, interstitial cells, stem cells, etc., when constructing the cell-cell relationship network, thereby constructing a more comprehensive TME cell atlas, and providing more comprehensive and accurate cell infiltration inference.

[0016] Preferably, the step S2 of constructing the cell-cell function correlation network according to the gene expression data and the GO function annotation data specifically comprises:

[0017] S21: calculate a Jaccard coefficient and weighted gene expression data according to the gene expression data and the GO function annotation data;

[0018] S22: calculate a correlation comprehensive weight according to the Jaccard coefficient and the weighted gene expression data; and construct a GO-cell bipartite network according to the correlation comprehensive weight;

[0019] S23: construct the cell-cell function correlation network according to the correlation comprehensive weight and the bipartite network;

[0020] Each element in the cell-cell function correlation network represents the correlation strength between two cell types.

[0021] Preferably, the step S3 constructs a cell-cell physical interaction network according to the gene expression data and the PPI network data, and specifically comprises:

[0022] S31: calculating a gene expression and network weight comprehensive contribution value according to the gene expression data and the PPI network data;

[0023] S32: calculating a shortest path distance between genes by a Floyd algorithm; and correcting the shortest path distance between genes for stability between genes;

[0024] S33: constructing a cell-cell physical interaction network according to the gene expression and network weight comprehensive contribution value and the result of the correction for stability between genes.

[0025] In the comprehensive cell relationship network, the nodes represent different cell types, and the weights of the edges represent the correlation between the cell types.

[0026] The step S5 calculates a final cell infiltration score based on the comprehensive cell relationship network by a restart walk algorithm, and specifically comprises:

[0027] S51: calculating a node cell infiltration score for each node in the comprehensive cell relationship network by the restart walk algorithm;

[0028] S52: performing logarithmic conversion on the node cell infiltration score;

[0029] S53: performing normalization processing on the result of the logarithmic conversion to obtain the final cell infiltration score.

[0030] In a second aspect, based on the same inventive concept, the present application further provides a cell infiltration inference model fusing GO function annotation and PPI network information, which comprises the cell-cell functional correlation network, the cell-cell physical interaction network and the comprehensive cell relationship network in the method of the first aspect.

[0031] The PPI information and the GO function annotation are fused to directly reflect the connection mode between cells from a molecular level, wherein the PPI information is from the perspective of protein interaction, and the GO function annotation is from the perspective of cell function, and the two complement each other, so that the calculated cell infiltration degree result is more reliable, thereby providing a more valuable basis for precise treatment of tumors.

[0032] In a third aspect, based on the same inventive concept, the present application further provides a system of a cell infiltration inference model fusing GO function annotation and PPI network information, which comprises:

[0033] The data acquisition unit is configured to acquire dataset; wherein the dataset comprises gene expression data, GO function annotation data and PPI network data.

[0034] The first data processing unit is configured to construct a cell-cell function correlation network according to the gene expression data and the GO function annotation data.

[0035] The second data processing unit is configured to construct a cell-cell physical interaction network according to the gene expression data and the PPI network data.

[0036] The third data processing unit is configured to perform weighted fusion processing on the cell-cell function correlation network and the cell-cell physical interaction network to obtain a comprehensive cell relationship network.

[0037] The fourth data processing unit is configured to calculate a final cell infiltration score based on the comprehensive cell relationship network by using a restart walk algorithm, and infer the infiltration degree in the tumor microenvironment according to the final cell infiltration score.

[0038] Preferably, the first data processing unit further comprises:

[0039] The first calculation module is configured to calculate a Jaccard coefficient and a weighted gene expression data according to the gene expression data and the GO function annotation data.

[0040] The second calculation module is configured to calculate a correlation comprehensive weight according to the Jaccard coefficient and the weighted gene expression data, and construct a GO-cell bipartite network according to the correlation comprehensive weight.

[0041] The first network construction module is configured to construct the cell-cell function correlation network according to the correlation comprehensive weight.

[0042] Each element in the cell-cell function correlation network represents the correlation strength between two cell types.

[0043] Preferably, the second data processing unit further comprises:

[0044] The third calculation module is configured to calculate a gene expression and network weight comprehensive contribution value according to the gene expression data and the PPI network data.

[0045] The shortest path calculation module is configured to calculate a shortest path distance between genes by using a Floyd algorithm, and perform a gene distance stability correction on the shortest path distance between genes.

[0046] The second network construction module is configured to construct the cell-cell physical interaction network according to the gene expression and network weight comprehensive contribution value and the result of the gene distance stability correction.

[0047] In a fourth aspect based on the same inventive concept, the application further provides a computer device, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the cell infiltration inference method of fusing GO functional annotation and PPI network information as described in the first aspect.

[0048] Compared with the prior art, the application has the beneficial effects that:

[0049] The cell infiltration inference method of fusing GO functional annotation and PPI network information provided by the application innovatively integrates GO functional annotation and PPI network information, and breaks through the dependence of traditional methods on single data. GO functional annotation covers the functional panorama of cells in biological processes, and PPI network introduces physical connection evidence of molecular interaction, and the two form a “function-structure” double pillar, so that the model can understand the cell synergy from the biological pathway level and analyze the direct interaction of cells from the protein interaction level, avoiding the one-sidedness of a single perspective, and providing a more stereoscopic cognitive framework for tumor microenvironment analysis.

[0050] In addition, compared with the traditional method which only focuses on part of immune cells, the application can identify the infiltration characteristics of cell types such as immune cells, interstitial cells and stem cells based on GO functional annotation, and the coverage is more than 3 times. At the same time, the interaction mode of cell subgroups (such as macrophages in different polarization states and T cell subtypes) is finely described through PPI network, realizing the leap from “general cell type analysis” to “full spectrum of subpopulation analysis”, and fully unlocking the complex cell composition of tumor microenvironment.

[0051] In summary, the present application realizes significant technical breakthrough and application value improvement in the field of tumor microenvironment cell infiltration inference through multi-source data fusion and innovative algorithm design. The bulk RNA-seq data, GO function annotation and PPI network data are innovatively integrated, breaking through the limitations of traditional single data processing. Based on GO annotation, the functional characteristics of cell types such as immune cells and interstitial cells can be covered, and the physical interaction and signal transduction relationship between cells can be directly described through PPI network, constructing a more complete and real tumor microenvironment cell relationship map. At the same time, the random walk algorithm combined with weighted fusion strategy is adopted, which ensures that the calculation efficiency is improved by more than 50% compared with traditional deep learning methods, and realizes the self-adaptation of different tumor types and sequencing platform samples. In practical application, this method shows excellent performance in survival analysis, far exceeding traditional methods, providing a reliable basis for clinical judgment of patient prognosis. In addition, the technology does not require special data or complex equipment, has strong compatibility and high scalability, and is easy to be clinically transformed, providing a new and efficient computing tool for tumor research and treatment. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 is the cell infiltration inference method flowchart of the present application embodiment.

[0053] Figure 2 is the infiltration score (predicted value) and flow cytometry (true value) verification result of different methods of the present application embodiment.

[0054] Figure 3 is the correlation between the infiltration score of different cell subgroups and the prognosis of patients of the present application embodiment.

[0055] Figure 4 is the survival prediction performance analysis of the present application embodiment.

[0056] Figure 5 is the survival prediction performance analysis of the existing method CITMIC of the present application embodiment. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme will be described clearly and completely below in combination with the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor belong to the scope of protection of the present application.

[0058] Embodiment one: as Figure 1As shown, the present application proposes a cell infiltration inference method fusing GO function annotation and PPI network information. By fusing PPI information and GO function annotation, the connection mode between cells can be directly reflected from the molecular level. The PPI information is from the perspective of protein interaction, and the GO function annotation is from the perspective of cell function. The two complement each other, so that the calculated cell infiltration degree result is more reliable, and further provides a more valuable basis for precise treatment of tumors. The specific implementation process is shown in steps 1 to 5.

[0059] Step 1: Collect the data set and pre-process.

[0060] The data set includes gene expression data, GO function annotation data and PPI network data.

[0061] (1) Collect gene expression data, which can be optionally obtained from the ImmPort database. The specific gene expression data includes SDY311 and SDY420, wherein SDY311 is immune-related healthy individual data, and SDY420 is immune-related disease patient data. The gene expression data can be downloaded by the access number provided by the database and saved in a common format such as CSV, TXT or TSV for subsequent processing and analysis.

[0062] In addition, the gene expression data and clinical data of TCGA-LGG (low-grade glioma) cancer patients are downloaded through the UCSC Xena platform. On the UCSC Xena platform, select the TCGA-LGG data set, download the gene expression data (such as RNA-Seq data) and related clinical data (such as patient clinical information, stage, prognosis, etc.). Extract the required gene expression matrix and save it in a standard format such as a TXT or CSV file for subsequent analysis.

[0063] (2) Collect GO function annotation data, which can be optionally downloaded from the MSigDB website. Select GO data sets related to gene function and pathways, especially annotations related to biological processes (BP), molecular functions (MF), and cellular components (CC). Download the relevant GO function set file (such as GMT format) and format-convert the data for use in subsequent analysis.

[0064] (3) Collect protein-protein interaction (PPI) network data. Download protein-protein interaction network and its detailed information from STRING database. Select the corresponding species (e.g. human) and the required PPI data (e.g. interaction relationships containing protein interaction scores greater than or equal to 700) in the STRING database. When downloading the data, select to obtain detailed protein interaction information such as protein ID, interaction score, relationship between proteins, etc.

[0065] Further, by calculating the intersection of genes in gene expression data and gene symbols in PPI network, genes that exist in both can be screened. According to these intersection genes, further filter the interaction relationships in the PPI network, only keep the protein interactions related to the intersection genes, thereby narrowing the scope of the PPI network, focusing on the interactions closely related to the gene expression data, facilitating subsequent analysis and research.

[0066] In this embodiment, by calculating the intersection of genes in gene expression data and gene symbols in PPI network, specifically including:

[0067] 1) Load gene expression data: Load the gene expression matrix obtained from the ImmPort database (SDY311 and SDY420) and TCGA-LGG cancer patient data. The data format can be CSV, TXT, etc., and each row represents the expression of a gene.

[0068] 2) Load PPI network data: The PPI network data downloaded from the STRING database contains the interaction relationships between proteins. Each record usually includes two protein IDs, their interaction score (such as confidence score), and other relationship information.

[0069] 3) Convert PPI protein ID to gene symbol: Convert the protein ID in the PPI network data to the gene symbol. This can be done by using Ensembl, NCBI or other gene annotation databases (such as biomart, UniProt). Ensure that all protein IDs are correctly mapped to the corresponding gene symbols.

[0070] 4) Select genes of interest: Extract the genes of interest from the gene expression data. For example, genes related to the disease, or highly differentially expressed genes. Ensure that the symbols of these genes are consistent with the gene symbols in the PPI network data.

[0071] 5) Filter gene expression data: Filter the gene expression data according to the threshold of expression to select the genes of interest. For example, genes with expression greater than a certain value, or significantly differentially expressed genes.

[0072] 6) Calculate the intersection: Extract genes in the PPI network, extract genes in the PPI network data that match the gene symbols in the gene expression data. For example, if the PPI network data contains gene A, gene B, gene C, gene D, and the gene expression data also contains these genes, these genes are matching genes.

[0073] Calculate the intersection of the gene symbols in the gene expression data and the gene symbols in the PPI network data. The specific method is:

[0074] Convert the gene symbol column in the gene expression data and the gene symbol column in the PPI network data into Set form.

[0075] Calculate the intersection of the two sets to get the set of genes that exist in both the gene expression data and the PPI network data. According to the result of the intersection, further narrow the scope of the PPI network. Only keep those gene pairs that have an interaction relationship with the genes in the gene expression data. That is, if a gene in a protein interaction pair is in the intersection gene list, keep the protein interaction relationship; otherwise, delete the relationship. According to the intersection gene and the screening result of the PPI network, construct a sub-network that only contains the interactions between the intersection genes. This will be a smaller PPI network closely related to the gene expression data. Thus, the resulting sub-network is more biologically meaningful and can provide valuable candidate genes for further functional analysis and disease research.

[0076] Step 2: Construct a cell-cell functional association network according to the gene expression data and GO functional annotation data;

[0077] In this embodiment, the cell-cell functional association network is constructed, which specifically includes:

[0078] S21: Calculate the Jaccard coefficient and the weighted gene expression data according to the gene expression data and the GO functional annotation data;

[0079] Preferably, the Jaccard coefficient is used to measure the similarity between two sets, and the Jaccard coefficient is calculated as follows:

[0080] The calculation formula is: ;

[0081] In the formula, C represents the genes contained in the GO annotation, and G represents the marker gene set of a specific cell type.

[0082] |C∩G| is the intersection size of the two sets, that is, the number of genes that belong to both the GO annotation and the marker genes of a specific cell type.

[0083] | C U G | is the size of the union of two sets, i.e. the total number of genes belonging to either GO annotation or specific cell type marker genes.

[0084] For example, if there are two sets:

[0085] ;

[0086] It can be obtained that the Jaccard coefficient between the set of GO annotation genes and the set of specific cell type marker genes is 0.33, indicating that the similarity of the two sets is 33%.

[0087] Preferably, the correlation between gene expression of a specific GO annotation and a cell type is measured by weighted gene expression data, and the weighted gene expression data is calculated as follows:

[0088] The calculation formula is: ;

[0089] Wherein, GEP g is the gene expression value of gene g in a specific cell type. Wg is the weight of gene g in the PPI network, which is calculated by the PageRank algorithm. The PageRank algorithm assigns a weight value according to the interaction relationship between genes, reflecting the importance of the gene in the network, g is the intersection of GO annotation and specific cell type marker genes, and W'g is the normalization of gene expression.

[0090] S22: Calculate the correlation comprehensive weight according to the Jaccard coefficient and the weighted gene expression data; construct a GO x cell bipartite network according to the correlation comprehensive weight;

[0091] The comprehensive weight Wc, G is the product of the Jaccard coefficient and wGEP g , and the comprehensive weight reflects the correlation strength between GO annotation and specific cell type. The formula is:

[0092] ;

[0093] For example:

[0094] Jaccard(C,G)=0.33;

[0095] wGEPg1=4.16;

[0096] wGEPg2=4.56;

[0097] wGEPg3=3.55;

[0098] Then sum the weighted gene expression data of all genes to get:

[0099] wGEPsum = wGEPg1 + wGEPg2 + wGEPg3 = 4.16 + 4.56 + 3.55 = 12.27;

[0100] Calculate the comprehensive weight: Wc,G = Jaccard(C,G) x wGEPsum = 0.33 x 12.27 = 4.05;

[0101] Based on the calculated comprehensive weight Wc,G, a bipartite network is constructed.

[0102] In this bipartite network, nodes represent different GO annotations and cell types, and the weights of edges are represented by the comprehensive weight Wc,G. The edge weights of the bipartite network reflect the strength of the association between GO annotations and specific cell types. The higher the weight of the edge, the closer the relationship between the GO annotation and the cell type. Through the bipartite network, complex gene expression data and functional annotations can be systematized, avoiding direct calculation of gene similarity in the cell x cell network. Through the GO x cell bipartite network, cell types are connected through their shared GO functional items. For example, if two cell types share multiple GO functional items, their functional similarity is high. The bipartite network makes this functional sharing relationship intuitive, thereby providing the necessary information for constructing a cell x cell functional association network.

[0103] S24: Construct a cell x cell functional association network according to the association comprehensive weight and the bipartite network;

[0104] Wherein each element in the association matrix of the cell x cell functional association network represents the association strength between two cell types.

[0105] The association strength R between cell type Ci and cell type Cj Ci,j can be calculated by the product of the comprehensive weight W C,G and its transpose, as follows:

[0106] R Ci , Cj = W C,G x W T C,G ;

[0107] Example: Given three cell types C1, C2 and C3, and the comprehensive weights of their association with GO annotations have been calculated:

[0108] W C1,G =[4.16,4.56,3.55] (for cell type C1).

[0109] W C2,G= [3.20, 3.60, 2.80] (for cell type C3).

[0110] W C3,G = [3.20, 3.60, 2.80] (for cell type C3).

[0111] Calculate the strength of association between cell types C1 and C2:

[0112] RC C1,C2 = W C1,G × W C2,G T = [4.16, 4.56, 3.55] × [5.00, 3.80, 4.00] T

[0113] The calculation process is:

[0114] R CC1,C2 = (4.16 × 5.00) + (4.56 × 3.80) + (3.55 × 4.00) = 20.80 + 17.33 + 14.20 = 52.33.

[0115] Calculate R CC1,C3 : Calculate the strength of association between cell types C1 and C3:

[0116] RC C1,C3 = W C1,G × W C3,GT = [4.16, 4.56, 3.55] × [3.20, 3.60, 2.80] T .

[0117] The calculation process is:

[0118] RC C1,C3 = (4.16 × 3.20) + (4.56 × 3.60) + (3.55 × 2.80) = 13.31 + 16.42 + 9.94 = 39.67.

[0119] Calculate RC C2,C3 : Calculate the strength of association between cell types C2 and C3:

[0120] RC C2,C3 = W C2,G × W C3,G T = [5.00, 3.80, 4.00] × [3.20, 3.60, 2.80] T .

[0121] The calculation process is:

[0122] RC C2,C3= (5.00 * 3.20) + (3.80 * 3.60) + (4.00 * 2.80) = 16.00 + 13.68 + 11.20 = 40.88.

[0123] According to the above calculation results, we can construct the association matrix of the cell x cell functional association network:

[0124] .

[0125] This embodiment obtains the cell x cell functional association matrix by calculating the product of the comprehensive weight and its transpose. Each element in the matrix represents the association strength between two cell types, reflecting their functional association in gene expression and PPI network. This provides a basis for further cell function analysis and network research.

[0126] S3: According to the gene expression data and PPI network data, construct a cell x cell physical interaction network;

[0127] Preferably, the construction of the cell x cell physical interaction network specifically includes:

[0128] S31: According to the gene expression data and PPI network data, calculate the gene expression and network weight comprehensive contribution value;

[0129] The formula is: gene expression and network weight comprehensive contribution value = W g1 ·GEP g1 +W g2 ·GEP g2 .

[0130] The gene expression and network weight comprehensive contribution value reflects the comprehensive contribution of the expression level of the gene in the cell type and the importance of the network.

[0131] For example:

[0132] In cell type i (immune cells), the expression value of gene 1 is GEPg1=2.5; the network weight of gene 1 is Wg1=1.2.

[0133] In cell type j (neuron cells), the expression value of gene 2 is GEPg2=3.0, and the network weight of gene 2 is Wg2=1.5.

[0134] Then, the calculation of the gene expression and network weight comprehensive contribution value is:

[0135] Wg1⋅GEPg1+Wg2⋅GEPg2=(1.2⋅2.5)+(1.5⋅3.0)=7.5.

[0136] This value reflects the combined contribution of the expression levels of gene 1 and gene 2 in immune cells and neuronal cells to their importance in the network. Using the calculation method of the combined contribution of gene expression and network weight, the influence of genes on different cell types can be quantified, revealing the role of genes in the cell network.

[0137] S32: Calculate the shortest path distance between genes by Floyd algorithm; correct the stability of the distance between genes for the shortest path distance between genes;

[0138] The connection strength of genes in the PPI network is measured by the distance between genes d, which is calculated by Floyd algorithm. Specifically, Floyd algorithm is used to calculate the shortest path distance between two genes in PPI network, thus reflecting their connection strength. In actual calculation, there may be no direct connection between some genes, which may lead to unstable calculation (for example, the distance is zero). In order to solve this problem, the invention further introduces the stability correction of the distance between genes (plus a small correction term ε), to ensure the stability and accuracy of the calculation process.

[0139] S33: According to the results of the combined contribution of gene expression and network weight and the stability correction of the distance between genes, construct the cell-cell physical interaction network.

[0140] By integrating S31 and S32, the cell-cell physical interaction network can be calculated, and the formula is:

[0141] ;

[0142] In the formula, the numerator measures the combined contribution of gene pairs from the perspective of gene expression and network importance; the denominator adjusts and normalizes the numerator from the perspective of actual interaction strength between genes and calculation stability. The combination of the two enables comprehensive and accurate construction of the cell-cell physical interaction network, providing strong support for the study of functional association between cells.

[0143] S4: Weighted fusion processing of the cell-cell functional association network and the cell-cell physical interaction network is performed to obtain a comprehensive cell relationship network.

[0144] The formula is:

[0145] .

[0146] β can be adjusted according to different input data. At this point, an F network can be constructed for each sample, with nodes representing different cell types and edge weights representing the association between cell types, so that the cell infiltration score can be inferred from the perspective of protein interaction and cell function.

[0147] S5: calculating a final cell infiltration score based on the integrated cell relationship network by a restart walk algorithm, and inferring the infiltration degree in the tumor microenvironment according to the final cell infiltration score.

[0148] Specifically comprising:

[0149] S51: calculating a node cell infiltration score for each node in the integrated cell relationship network by the restart walk algorithm.

[0150] The weight updating formula of the restart walk algorithm is:

[0151] ;

[0152] wherein, a is a restart probability (usually 0.85), I(g) is an initial weight of the cell g, N(g) is a neighbor node set of the cell g, and dg' is an out-degree of the node g'.

[0153] S52: performing logarithmic conversion on the node cell infiltration score.

[0154] The restart walk algorithm can be used to calculate a Score for each node (cell type) in the network, and logarithmic conversion processing can be performed on possible non-positive values, and the formula is: .

[0155] S53: performing normalization processing on the logarithmic conversion result to obtain a final cell infiltration score.

[0156] The infiltration score is scaled to the range of 0 to 1 by normalization processing, so as to compare across samples, and the final infiltration score is obtained, and the formula is: .

[0157] In order to further verify the cell infiltration inference method of fusing GO function annotation and PPI network information according to the present application, the correlation coefficients of the infiltration scores (predicted values) and flow cytometry (true values) of different methods are compared on the adrenal gland data set. Among them, the existing mainstream methods are CIBERSORT, MCP-counter, xCell and CITMIC; the cell infiltration inference method of the present application is marked as ours. The Spearman rank correlation coefficient of the predicted infiltration score and the true value (range: [-1.0, 1.0]). As shown in Figure 2 the cell infiltration inference method of the present application is significantly better than the prior art on all cell types, and the specific performance is:

[0158]

[0159] It can be seen that the application comprehensively surpasses the prior art system and sets a new standard for cell infiltration analysis by maintaining the advantages of CITMIC function annotation, PPI network weighting and dynamic optimization mechanism.

[0160] Figure 3 Survival analysis of the method of the application. The correlation between the infiltration scores of different cell subpopulations and the prognosis of patients is shown, in which 65 cell subpopulations are significantly related to prognosis. This shows the potential influence of numerous cell types in the tumor microenvironment on prognosis, highlighting the importance of considering multiple cell types comprehensively. The method can effectively identify these prognostically significant cell subpopulations, providing key information for the construction of subsequent risk models.

[0161] Figures 4-5 As shown, Figure 4 Survival prediction performance analysis of the method of the application. Figure 5 Survival prediction performance of the existing method CITMIC, after comparison (ROC curve). The ROC curves of the method of the application and the CITMIC method at different time points (1 year, 3 years, 5 years, 7 years) are shown, reflecting the performance difference of the two in the survival prediction task, indicating that the method of the application has higher AUC value at all time points, i.e. better survival prediction ability.

[0162] In summary, the application innovatively integrates Gene Ontology (GO) function annotation information and protein-protein interaction (PPI) network data, and through the construction of a weighted fusion of cell-cell functional association network and cell-cell physical interaction network, it comprehensively considers the functional similarity and physical or signal interaction between cells, and when constructing the cell relationship network, it can cover multiple cell types, including immune cells, interstitial cells, stem cells, etc., constructing a more comprehensive TME cell map, thereby providing more comprehensive and accurate cell infiltration inference. The limitations of traditional methods that rely only on a single data source are effectively overcome, making cell infiltration inference more consistent with the true biological characteristics of the tumor microenvironment.

[0163] In Example Two, the application also proposes a cell infiltration inference model that integrates GO function annotation and PPI network information, which includes the cell-cell functional association network, the cell-cell physical interaction network and the comprehensive cell relationship network in the method of Example One.

[0164] By fusing PPI network and GO function annotation information, the cell infiltration inference model proposed in the application can comprehensively reflect the connection mode between cells from the molecular level, and then accurately infer the infiltration degree of immune cells in tumor tissues. The model not only improves the accuracy of cell infiltration inference, but also provides an important basis for precise treatment of tumors. The popularization and application of the method can provide more reliable data support for the individualized treatment plan of tumor immunotherapy.

[0165] In embodiment three, the application further proposes a system of a cell infiltration inference model fusing GO function annotation and PPI network information, the system comprising:

[0166] a data acquisition unit for acquiring a data set; wherein the data set comprises gene expression data, GO function annotation data and PPI network data;

[0167] a first data processing unit for constructing a cell-cell function association network according to the gene expression data and the GO function annotation data;

[0168] a second data processing unit for constructing a cell-cell physical interaction network according to the gene expression data and the PPI network data;

[0169] a third data processing unit for performing weighted fusion processing on the cell-cell function association network and the cell-cell physical interaction network to obtain a comprehensive cell relationship network;

[0170] and a fourth data processing unit for calculating a final cell infiltration score based on the comprehensive cell relationship network by using a restart walk algorithm, and inferring the infiltration degree in the tumor microenvironment according to the final cell infiltration score.

[0171] Preferably, the first data processing unit further comprises:

[0172] a first calculation module for calculating a Jaccard coefficient and weighted gene expression data according to the gene expression data and the GO function annotation data;

[0173] a second calculation module for calculating an association comprehensive weight according to the Jaccard coefficient and the weighted gene expression data, and constructing a GO-cell bipartite network according to the association comprehensive weight;

[0174] and a first network construction module for constructing a cell-cell function association network according to the association comprehensive weight;

[0175] Each element in the cell-cell function association network represents the association strength between two cell types.

[0176] Preferably, the second data processing unit further comprises:

[0177] a third calculation module, configured to calculate a gene expression and network weight comprehensive contribution value according to the gene expression data and the PPI network data;

[0178] a shortest path calculation module, configured to calculate a shortest path distance between genes by using a Floyd algorithm, and to correct the shortest path distance between genes in terms of stability;

[0179] a second network construction module, configured to construct a cell-cell physical interaction network according to the gene expression and network weight comprehensive contribution value and the result of the correction of the stability of the distance between genes.

[0180] In embodiment four, the present application further provides a computer device, which comprises a processor and a memory, and the memory stores at least one instruction, which is loaded and executed by the processor to implement the cell infiltration inference method of fusing GO functional annotation and PPI network information as described in embodiment one.

[0181] Although the example embodiments have been described herein with reference to the accompanying drawings, it is to be understood that the example embodiments are only exemplary and are not intended to limit the scope of the present application thereto. Various changes and modifications can be made thereto by those skilled in the art without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as defined in the appended claims.

[0182] It should be noted that, in this document, the terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0183] Although the present application has been described in conjunction with the above specific embodiments, it is obvious to those skilled in the art that many substitutions, modifications and changes can be made according to the above description. Therefore, all such substitutions, modifications and changes are included in the spirit and scope of the appended claims.

Claims

1. A method for inferring cell infiltration by integrating GO functional annotation and PPI network information, characterized in that, include: S1: Collect the dataset and preprocess it, wherein the dataset includes gene expression data, GO functional annotation data and PPI network data; S2: Construct a cell-cell function association network based on gene expression data and GO functional annotation data; S3: Construct a cell-cell physical interaction network based on gene expression data and PPI network data; S4: The cell-cell functional association network and the cell-cell physical interaction network are weighted and fused to obtain a comprehensive cell relationship network; S5: Based on the comprehensive cell relationship network, the final cell invasion score is calculated by restarting the walk algorithm, and the degree of invasion in the tumor microenvironment is inferred based on the final cell invasion score; Step S2 includes: S21: Calculate the Jaccard coefficient and weighted gene expression data based on gene expression data and GO functional annotation data; S22: Calculate the association weights based on the Jaccard coefficient and weighted gene expression data; construct a binary network of GO× cells based on the association weights; S23: Construct a cell × cell function association network based on the aforementioned association comprehensive weights and binary network; Each element in the cell × cell function association network represents the association strength between two cell types; Step S3 includes: S31: Calculate the combined contribution value of gene expression and network weight based on gene expression data and PPI network data; S32: Calculate the shortest path distance between genes using the Floyd algorithm; perform gene distance stability correction on the shortest path distance between genes; S33: Based on the combined contribution value of gene expression and network weights and the result of gene distance stability correction, construct a cell × cell physical interaction network.

2. The cell infiltration inference method integrating GO functional annotation and PPI network information according to claim 1, characterized in that, In the integrated cell relationship network, nodes represent different cell types, and edge weights represent the relationships between cell types.

3. The cell infiltration inference method integrating GO functional annotation and PPI network information according to claim 2, characterized in that, Step S5 includes: S51: By restarting the walk algorithm, the node cell infiltration score is calculated for each node in the integrated cell relationship network; S52: Perform a logarithmic transformation on the node cell infiltration fraction; S53: Normalize the results of the logarithmic transformation to obtain the final cell infiltration score.

4. A cell infiltration inference model integrating GO functional annotation and PPI network information, characterized in that, The model includes the cell-cell functional association network, cell-cell physical interaction network, and comprehensive cell relationship network as described in any of claims 1-3.

5. A system employing the cell infiltration inference model that integrates GO functional annotation and PPI network information as described in claim 4, characterized in that, The system includes: The data acquisition unit collects dataset data, which includes gene expression data, GO functional annotation data, and PPI network data. The first data processing unit is used to construct a cell-cell functional association network based on gene expression data and GO functional annotation data. The second data processing unit is used to construct a cell-cell physical interaction network based on gene expression data and PPI network data. The third data processing unit is used to perform weighted fusion processing on the cell × cell functional association network and the cell × cell physical interaction network to obtain a comprehensive cell relationship network. And a fourth data processing unit, used to calculate the final cell invasion score based on the comprehensive cell relationship network by restarting the walk algorithm, and to infer the degree of invasion in the tumor microenvironment based on the final cell invasion score.

6. The system according to claim 5, characterized in that, The first data processing unit further includes: The first calculation module calculates the Jaccard coefficient and weighted gene expression data based on gene expression data and GO functional annotation data. The second calculation module calculates the association weights based on the Jaccard coefficients and weighted gene expression data; and constructs a binary network of GO× cells based on the association weights. And, the first network construction module constructs a cell × cell function association network based on the aforementioned association comprehensive weights; In this context, each element in the cell-cell function association network represents the association strength between two cell types.

7. The system according to claim 5, characterized in that, The second data processing unit further includes: The third calculation module calculates the combined contribution value of gene expression and network weight based on gene expression data and PPI network data. The shortest path calculation module calculates the shortest path distance between genes using the Floyd algorithm and performs stability correction on the shortest path distance between genes. In addition, the second network construction module constructs a cell-cell physical interaction network based on the combined contribution value of gene expression and network weights and the result of gene-inter-gene distance stability correction.

8. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one instruction, which is loaded and executed by the processor to implement the cell infiltration inference method that integrates GO functional annotation and PPI network information as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Immune microenvironment difference detection and typing method and system for myeloblastoma

    CN119069004A

  • Experimental method for relation between cytokine level and tumor immune microenvironment in circulation

    CN119993268A