Graph network construction method and device, electronic equipment and storage medium

By constructing a graph network of genes and cells and utilizing feature screening and machine learning methods, the problem of the inability to qualitatively and quantitatively analyze gene-cell interactions in existing technologies has been solved. This enables quantitative indication and intensity assessment of gene-cell interactions, providing greater support for clinical analysis.

CN116230090BActive Publication Date: 2026-04-28SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
Filing Date
2023-02-06
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies are insufficient for qualitative and quantitative analysis of gene-cell interactions, cannot predict new gene-cell interaction pairs, and cannot provide quantitative indicators of interaction strength.

Method used

By acquiring genes and cells from multiple samples, performing feature screening and evaluation, constructing gene sample graph networks and cell sample graph networks, and using machine learning and graph network edge prediction methods, inferring other possible gene-cell interaction pairs, and updating the interaction relationships through factor graph networks.

Benefits of technology

It enables qualitative and quantitative analysis of gene-cell interactions, indicating the existence and strength of interactions, and providing more information for clinical analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116230090B_ABST
    Figure CN116230090B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a kind of based on gene and cell relationship's graph network construction method and device, the method includes: obtaining gene and cell in multiple samples, respectively carry out feature screening to gene and cell in each sample, obtain the gene and cell for indicating the characteristics of each sample;According to the gene and cell for indicating the characteristics of each sample, gene sample graph network and cell sample graph network are respectively constructed, and the importance of each node feature in the two sample graph networks is calculated respectively;According to the gene and cell that there is interaction relationship, and the importance of each node feature in the two sample graph networks, construct initial factor graph network;According to the adjacency matrix of factor graph network, whether there is other gene and cell that exists interaction relationship in factor graph network is predicted, and factor graph network is updated based on the gene and cell predicted.The present application solves the problem that gene and cell interaction relationship cannot be qualitatively and quantitatively analyzed in related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical omics information, and in particular to a method, apparatus, electronic device, and storage medium for constructing graph networks based on gene-cell relationships. Background Technology

[0002] Tumorigenesis is a complex and dynamic process, comprising three stages: initiation, progression, and metastasis. The physiological state of the tumor microenvironment (TME) is closely related to each step of tumorigenesis. The TME contains tumor cells, immune cells, stromal cells, fibroblasts, adipocytes, and inflammatory cells, among others. These cells interact with each other; for example, T helper cells support CD8+ cells. In addition to the cells within the TME, genes also play a crucial role in its development. Cell-cell interactions within the TME are mediated by gene-encoded proteins—cytokines, such as chemokines and growth factors. In other words, factors secreted by some cells act on other cells, leading to changes in the TME and thus influencing disease progression.

[0003] Current research includes many studies on genes or cells in the tumor microenvironment, such as gene co-expression module analysis to establish the correlation between gene modules and tumor characteristics, or analysis of different omics information of the same gene to identify cancer microenvironment characteristics, such as somatic mutations, copy number variations, and methylation, or inference of the degree of cell invasion in the tumor microenvironment based on gene expression levels, such as CIBERSORT and Xcell, and analysis of tumor microenvironment characteristics based on the degree of cell invasion.

[0004] However, current technologies mostly focus on genes or cells, rarely considering the interactions between genes and cells. Exploring the correlation of gene-cell interactions is equally important. At the same time, it is impossible to predict new gene-cell interaction pairs based on known gene-cell relationships. This not only fails to indicate whether there is a gene-cell interaction relationship, but also fails to quantitatively provide the strength of the interaction, thus failing to provide more information for exploring the progression and characteristics of the tumor microenvironment.

[0005] Therefore, there is an urgent need for a graph network construction method based on gene-cell relationships that can qualitatively and quantitatively analyze the interaction between genes and cells. Summary of the Invention

[0006] The embodiments of the present invention provide a method, apparatus, electronic device and storage medium for constructing graph networks based on gene-cell relationships, in order to solve the problem in related technologies that cannot qualitatively and quantitatively analyze the interaction relationship between genes and cells.

[0007] The technical solution adopted in this invention is as follows:

[0008] According to one aspect of the present invention, a method for constructing a graph network based on gene-cell relationships is provided. The method includes: acquiring genes and cells from multiple samples; performing feature screening on the genes and cells in each sample to obtain genes and cells used to indicate the features of each sample; constructing a gene sample graph network and a cell sample graph network based on the genes and cells used to indicate the features of each sample, and calculating the importance of the features of each node in the gene sample graph network and the cell sample graph network respectively; constructing an initial factor graph network based on the genes and cells with interaction relationships, and the features and importance of each node in the gene sample graph network and the cell sample graph network; predicting whether there are other genes and cells with interaction relationships in the factor graph network based on the adjacency matrix of the factor graph network, and updating the factor graph network based on the predicted genes and cells with interaction relationships.

[0009] According to one aspect of the present invention, a graph network construction apparatus based on gene-cell relationships is provided, the apparatus comprising: a data preprocessing module for acquiring genes and cells from multiple samples, performing feature screening on the genes and cells in each sample to obtain genes and cells indicating the characteristics of each sample; a factor characteristic evaluation module for constructing a gene sample graph network and a cell sample graph network based on the genes and cells indicating the characteristics of each sample, and calculating the importance of the features of each node in the gene sample graph network and the cell sample graph network respectively; a factor graph network establishment module for constructing an initial factor graph network based on genes and cells with interaction relationships, and the features and importance of each node in the gene sample graph network and the cell sample graph network; and a path prediction module for predicting whether there are other genes and cells with interaction relationships in the factor graph network based on the adjacency matrix of the factor graph network, and updating the factor graph network based on the predicted genes and cells with interaction relationships.

[0010] According to one aspect of the present invention, an electronic device includes a processor and a memory, wherein the memory stores computer-readable instructions, which, when executed by the processor, implement the gene-cell relationship-based graph network construction method as described above.

[0011] According to one aspect of the present invention, a storage medium having a computer program stored thereon, which, when executed by a processor, implements the graph network construction method based on gene-cell relationships as described above.

[0012] According to one aspect of the present invention, a computer program product includes a computer program stored in a storage medium, a processor of a computer device reads the computer program from the storage medium, and the processor executes the computer program such that the computer device, when executed, implements the graph network construction method based on gene-cell relationships as described above.

[0013] The above technical solution realizes the prediction of new gene-cell interaction pairs based on known interaction relationships between gene and cell, which not only indicates whether there is an interaction relationship between genes and cells, but also quantitatively gives the interaction strength based on the gene-cell relationship graph network construction method.

[0014] Specifically, this method characterizes the correlation between genes and cells by identifying the features and importance of known gene-cell pairs with interactions, and uses machine learning to predict edges in graph networks, thereby inferring other gene-cell pairs with interactions. This provides more information for the analysis of practical clinical problems and solves the problem of the inability to qualitatively and quantitatively analyze the interaction between genes and cells in existing technologies.

[0015] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0017] Figure 1 This is a flowchart illustrating a graph network construction method based on gene-cell relationships according to an exemplary embodiment;

[0018] Figure 2 yes Figure 1 A flowchart of step 110 in one embodiment corresponds to the following example;

[0019] Figure 3 yes Figure 1 A flowchart of step 130 in one embodiment corresponds to the following example;

[0020] Figure 4 yes Figure 1 A flowchart of step 130 in one embodiment corresponds to the following example;

[0021] Figure 5 yes Figure 1 A flowchart of step 150 in one embodiment corresponds to the following example;

[0022] Figure 6 yes Figure 1A flowchart of step 170 in one embodiment corresponds to the following example;

[0023] Figure 7 This is a block diagram illustrating a graph network construction device based on gene-cell relationships according to an exemplary embodiment;

[0024] Figure 8 yes Figure 7 Flowchart of the device in the application scenario corresponding to the embodiment;

[0025] Figure 9 This is a hardware structure diagram of an electronic device according to an exemplary embodiment;

[0026] Figure 10 This is a block diagram illustrating an electronic device according to an exemplary embodiment.

[0027] The accompanying drawings have illustrated specific embodiments of the present invention, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the inventive concept in any way, but rather to illustrate the concept of the invention to those skilled in the art by referring to specific embodiments. Detailed Implementation

[0028] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0029] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0030] Most existing technologies focus on genes or cells, rarely considering the interactions between genes and cells. However, exploring the correlations between gene-cell interactions is equally important. The complexity of the disease environment necessitates the exploration of more undiscovered gene-cell correlations to provide more information for understanding the progress and characteristics of the disease microenvironment.

[0031] Meanwhile, existing related technologies often only consider the characteristics of genes or cells and use relatively simple methods to calculate correlations. These methods can only obtain correlations roughly and cannot go further to obtain the degree of correlation.

[0032] In addition, incomplete sample labels limit the consideration of different characteristics of genes and cells. Although the biological characteristics of genes or cells are often judged in the research process based on their expression differences in different samples, such as when studying whether a gene or cell is related to inflammation, the conclusion is drawn by judging whether there is a significant difference in the expression value of a certain gene or cell in inflammatory and non-inflammatory samples. However, in some large databases, the lack of some tumor feature labels or the "unknown" label of the samples limits the consideration of different characteristics of genes and cells.

[0033] As can be seen from the above, the relevant technologies still have the limitation of being unable to qualitatively and quantitatively analyze the interaction between genes and cells.

[0034] Therefore, the gene-cell relationship-based graph network construction method provided in this application characterizes the correlation between genes and cells by knowing the characteristics and importance of gene-cell pairs that interact, and uses machine learning to predict the edges of the graph network, thereby inferring other gene-cell pairs that also interact. This provides more information for the analysis of actual clinical problems. The gene-cell relationship-based graph network construction method is applicable to gene-cell relationship-based graph network construction devices, which can be deployed on electronic devices configured with the von Neumann architecture, such as desktop computers, laptops, servers, etc.

[0035] Please see Figure 1 This application provides a method for constructing a graph network based on the relationship between genes and cells. This method is applicable to electronic devices, such as desktop computers, laptops, servers, etc.

[0036] In the following method embodiments, for ease of description, the execution subject of each step of the method is an electronic device, but this does not constitute a specific limitation.

[0037] like Figure 1 As shown, the method may include the following steps:

[0038] Step 110: Obtain genes and cells from multiple samples, and perform feature screening on the genes and cells in each sample to obtain genes and cells used to indicate the characteristics of each sample.

[0039] The genes in the sample are one or more genes depending on the research context, such as mRNA, somatic mutations, methylation, or copy number variations, without any specific limitations.

[0040] In one possible implementation, feature screening refers to removing genes that are lowly expressed in the sample, genes that are irrelevant to the research objective, genes that are not differentially expressed in the control environment, etc., without any limitation here.

[0041] Specifically, such as Figure 2 As shown, step 110 may include the following steps:

[0042] Step 210: Obtain genes from multiple samples, and screen the genes of each sample to obtain genes used to indicate sample characteristics.

[0043] In one possible implementation, the genes in each sample are screened, including removing genes with low expression and genes with irrelevant expression, etc., without limitation. Here, low-expressed genes refer to genes whose expression value is below a certain threshold in the sample, and genes with irrelevant expression refer to genes whose expression is not related to the research purpose.

[0044] Step 230: Obtain the cell composition data of each sample based on the gene expression profile of each sample, and assess cell infiltration based on the cell composition data.

[0045] In one possible implementation, deconvolution is used to probe the local from the global perspective, and the composition data of cells in each sample is obtained based on the gene expression profile of each sample, such as the specific immune cell composition of solid tumors, etc., which is not limited here.

[0046] Infiltration refers to the invasion of abnormal cells into human tissues or the appearance of cells that should not be present under normal circumstances, as well as the phenomenon of certain diseased tissues spreading to the surrounding areas. The results of cell infiltration assessment indicate the changes in cells caused by the spread of disease.

[0047] Step 250: Screen the cells based on the cell infiltration assessment results to obtain cells used to indicate sample characteristics.

[0048] Through the above process, it is ensured that the genes and cells used in the subsequent process of this invention can be used to indicate sample characteristics, thus ensuring the accuracy of the subsequently constructed sample graph network and consequently the accuracy of the constructed factor graph network.

[0049] Step 130: Construct gene sample graph networks and cell sample graph networks based on the genes and cells used to indicate the characteristics of each sample, and calculate the importance of the features of each node in the graph networks.

[0050] In this sample graph network, each node corresponds to multiple samples. The features of genes and cells used to indicate the characteristics of the samples are used as the features of each node in the sample graph network. In one possible implementation, the node features of the gene sample graph network are gene features, and the node features of the cell sample graph network are cell features.

[0051] Specifically, such as Figure 3 As shown, step 130, which involves constructing a gene sample map network based on the genes and cells used to indicate the characteristics of each sample, may include the following steps:

[0052] Step 310: Calculate the similarity between unlabeled samples and the similarity between unlabeled samples and labeled samples.

[0053] In one possible implementation, a subset of samples carries a tag indicating vascular invasion. In this case, a connection is established between samples carrying the vascular invasion tag for the purpose of studying vascular invasion. The samples may carry multiple tags or may not carry any tags.

[0054] In one possible implementation, the similarity between unlabeled samples and between unlabeled and labeled samples can be achieved using algorithms such as cosine similarity, Euclidean distance, Mahalanobis distance, Manhattan distance, Chebyshev distance, and Jaccard index, without any specific limitations here.

[0055] Step 330: Using samples as nodes, establish paths between different nodes based on the calculated similarity and whether the labels carried by the samples are the same.

[0056] In one possible implementation, establishing paths between different nodes is related to the similarity calculated in step 310. For example, cosine similarity can be used as the similarity between unlabeled samples. Assuming the cosine similarity is greater than 0.75, then a path can be established between the unlabeled samples.

[0057] In one possible implementation, establishing a path between different nodes is related to whether the samples carry the same label. Specifically, when the research objective involves multiple labels carried by a sample, it can be set so that only samples carrying the same label can establish a connection, while samples not carrying the same label cannot establish a connection. Of course, in other embodiments, the method of establishing a path between different nodes can be flexibly adjusted according to the research objective, and this does not constitute a specific limitation.

[0058] Step 350 yields the gene sample map network and the cell sample map network.

[0059] Specifically, using the features of genes and cells that indicate the characteristics of each sample as node features, gene sample graph networks with gene features as node features and cell sample graph networks with cell features as node features are obtained respectively.

[0060] In one possible implementation, genes and cells used to indicate sample characteristics are used as features of each node in the gene sample graph network and cell sample graph network, respectively. For example, somatic mutations, copy number variations, and methylation of genes are used as node features. Cells are not limited to immune cells, but may also include fibroblasts, stromal cells, etc., without limitation here.

[0061] Through the above process, this embodiment can establish connections based on the biometric similarity between samples, and obtain a sample graph network to provide a basis for the subsequent construction of a factor graph network based on gene-cell relationships.

[0062] It is worth mentioning that, based on the features of each node in the above sample graph network, semi-supervised learning is performed using graph neural networks to infer the labels of samples that do not carry labels in the above sample graph network. This overcomes the problem of incomplete sample labels in the existing technology and can complete the labels of the samples, so as to facilitate subsequent research and analysis of the samples.

[0063] Specifically, such as Figure 4 As shown, step 130, which calculates the importance of each node feature in the gene sample map network and the cell sample map network respectively, may include the following steps:

[0064] Step 410: Calculate the correlation between nodes in the gene sample graph network using a pixel-level interpretation algorithm to obtain the importance of each node's features in the gene sample graph network.

[0065] Step 430: Calculate the correlation between nodes in the cell sample graph network using a pixel-level interpretation algorithm to obtain the importance of each node's features in the cell sample graph network.

[0066] One possible implementation utilizes the LRP pixel-level interpretation algorithm to evaluate the importance of features at each node. The LRP algorithm is a general interpretation method for nonlinear classification architectures; its core idea is to decompose the output function of a specific target into a set of relevance scores and redistribute them to neurons in the previous layer.

[0067] The specific rules for propagating importance are as follows:

[0068]

[0069] Among them, Ri and R j Let a and j represent the correlations between nodes i and j, respectively. ∑j traverses all the parent nodes connected to node i. i It is the output or activation of node i; w ij The weights of nodes i and j are represented by , and l represents the number of propagation layers.

[0070] Through the above process, the correlation between each node can be obtained, which reflects the importance of the node in the classification task. Among them, for graph neural networks with different research purposes, the important node features obtained by the LRP algorithm and the importance of the node features will be different. For example, the role of some node features is related to immune escape in the tumor microenvironment, while others are related to angiogenesis.

[0071] Step 150: Construct an initial factor graph network based on the interaction relationships between genes and cells, as well as the characteristics and importance of each node in the gene sample graph network and the cell sample graph network.

[0072] In this embodiment, the construction of the initial factor graph network may include the following steps: using genes or cells as nodes, establishing paths between different nodes based on genes and cells with interactive relationships; calculating the weights of each path in the factor graph network based on the features and importance of each node in the gene sample graph network and the cell sample graph network; the node features and importance of the gene sample graph network are gene features and gene importance, and the node features and importance of the cell sample graph network are cell features and cell importance; the initial factor graph network is constructed from each node, the paths between nodes, and their weights. Here, nodes in the factor graph network represent genes or cells, node features are gene features or cell features, paths connecting two nodes represent genes and cells with interactive relationships, and path weights represent the interaction strength between genes and cells with interactive relationships.

[0073] In one possible implementation, known gene-cell interactions refer to inhibitory or activating effects, such as HLA-B / C molecules acting as inhibitory receptors for natural killer (NK) cells in lung adenocarcinoma, etc., without being limited here.

[0074] In one possible implementation, the factor graph network only contains cell-gene connections, so it can be viewed as a bipartite graph with only G = (W, E, R), and the encoder is [Z]. u Z v ]=f(X u X v M1, ..., M R ),in This indicates the rating level type: 1 means there is a connection between genes and cells, which means there is weight; 0 means there is no connection between genes and cells, which means there is no weight. The decoder is... This represents the embedding function in genes and cells, returning a weight matrix of dimension N. u ×N v .

[0075] Specifically, such as Figure 5 As shown, the calculation process of the weights of each path in a factor graph network may include the following steps:

[0076] Step 510: For genes and cells that have interaction relationships, determine the characteristics and importance of each gene and cell.

[0077] Specifically, the characteristics and importance of genes and cells refer to gene features and gene importance, as well as cell features and cell importance. Determining the accurate gene features and gene importance, and cell features and cell importance, prepares the groundwork for subsequent calculations of the weights in the factor graph network.

[0078] Step 530: Determine whether the difference between gene importance and cell importance is less than a threshold.

[0079] The threshold can be flexibly adjusted according to the actual needs of the application scenario, and is not limited here. For example, if the threshold is 0.5, then in step 530, it is determined whether the difference between gene importance and cell importance is less than 0.5.

[0080] If the difference between gene importance and cell importance is less than the threshold, proceed to step 570.

[0081] Conversely, if the difference between gene importance and cell importance is not less than the threshold, then proceed to step 550.

[0082] Step 550: Update gene importance and cell importance based on the difference between gene importance and cell importance.

[0083] The updated calculation process is as follows:

[0084]

[0085] Among them, w g w c These represent the importance of genes and cells, respectively, w′ g It is the updated gene importance, w′ c It is the importance of the updated cells.

[0086] Through the above process, the importance of genes and cells will not be extremely high or low, so that the weights of each path in the factor graph network will not be affected by the extreme value of a certain value. The representation method of the weights of each path in the factor graph network in this invention does not favor the importance of features or genes and cells. We avoid the situation where the importance of a gene to a certain tumor microenvironment feature is extremely high (or extremely low) while the importance of a cell is extremely low (or extremely high), so that the synergistic effect of genes and cells calculated in the end is quite high.

[0087] Step 570: Calculate the weights of each path in the factor graph network based on gene characteristics, cell characteristics, gene importance, and cell importance.

[0088] If the difference between gene importance and cell importance is less than a threshold, then the unupdated gene importance and cell importance determined in step 510 are used to calculate the path weight.

[0089] Conversely, if the difference between gene importance and cell importance is not less than the threshold, then the updated gene importance and cell importance from step 550 will be used in the path weight calculation.

[0090] Specifically, the calculation process is as follows:

[0091] Where w g w c Satisfy |w g -w c | < 0.5.

[0092] otherwise,

[0093]

[0094] Among them, w g&c E represents the weight of each path in the factor graph network. g E c These are gene characteristics and cellular characteristics, respectively. Here, gene characteristics refer to gene expression levels (mutation rates, etc.), and cellular characteristics refer to the degree of cell invasion. g w c These are the importance of genes and the importance of cells, respectively. g It is the updated gene importance, w′ c It is the importance of the updated cells.

[0095] Step 170: Based on the adjacency matrix of the factor graph network, predict whether there are other genes and cells with interaction relationships in the factor graph network, and update the factor graph network based on the predicted genes and cells with interaction relationships.

[0096] Specifically, such as Figure 6As shown, step 170 may include the following steps:

[0097] Step 610: Standardize the features of each node in the factor graph network according to the adjacency matrix of the factor graph network, and sort the standardization results.

[0098] In one possible implementation, the adjacency matrix A represents the connection relationships between genes and cells that have an interaction relationship. It only includes connection relationships between genes and cells that are known to have an interaction relationship. The node features are the expression values ​​of genes and cells in each sample, such as the degree of infiltration and mutation rate. The nodes are ranked by calculating the dot product of the embeddings of the features of the nodes connected by the same path, thereby representing the gene-cell interaction. The dot product of the embeddings of the node features is used to represent the standardized features of each node.

[0099] Step 630: Based on the sorting results, predict whether there are still genes and cells with interaction relationships in the factor graph network, establish paths between the predicted genes and cells with interaction relationships, and obtain new paths in the factor graph network.

[0100] The conditions for predicting unknown gene-cell interactions based on the sorting results can be flexibly adjusted according to the actual needs of the application scenario, and are not limited here. For example, the condition could be that if the nodes in the top 30% of the sorting are considered to have interactions, then new paths would be established between these nodes.

[0101] Step 650: Calculate the weight of the new path based on the weight of each path in the factor graph network, and update the factor graph network based on the new path and its weight.

[0102] In one possible implementation, a graph neural network is used to predict the weights of the new paths, so that the factor graph network can be updated based on the new paths and their weights.

[0103] In other words, in the process of predicting whether there are still interactions between genes and cells in the factor graph network, not only is the interaction between genes and cells considered with the adjacency matrix, but also the strength of the gene-cell interaction is considered through the path weight.

[0104] It should be further noted that, in one possible implementation, the calculation of the labels of unlabeled samples, the importance of the features of each node in the sample graph network, the prediction of unknown gene-cell interactions based on the adjacency matrix of the initial factor graph network, and the updating of the factor graph network based on the prediction results are achieved through a graph network model. This graph network model is a trained graph neural network capable of predicting the interactions between nodes in both the sample graph network and the factor graph network.

[0105] Through the above process, the embodiments of the present invention overcome the problem in the prior art that it is impossible to qualitatively and quantitatively analyze the interaction relationship between genes and cells. It not only considers the characteristics of genes and cells, but also the importance of genes and cells, thereby improving the accuracy and effectiveness of prediction results. Therefore, the embodiments of the present invention not only indicate whether there is an interaction relationship between genes and cells, but also quantitatively give the interaction strength, thereby realizing accurate graph network construction based on gene-cell relationship, and providing more information for the analysis of actual clinical problems.

[0106] This invention establishes a graph network of known gene-cell interactions to predict new gene-cell pair interactions. In this gene-cell graph network, we consider the characteristics and importance of genes and cells to determine the strength of gene-cell interactions. During the strength characterization process, our characterization method does not favor expression values ​​or the relative importance of genes and cells. We avoid situations where a gene is extremely important (or extremely low important) for a particular tumor microenvironment feature while the cell is extremely low (or extremely high important), ensuring that the final calculated gene-cell synergistic effects are relatively high.

[0107] The following are embodiments of the apparatus described in this application, which can be used to execute the graph network construction method based on gene-cell relationships involved in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments of the graph network construction method based on gene-cell relationships involved in this application.

[0108] Please see Figure 7 In one exemplary embodiment, a graph network construction device 700 based on gene-cell relationships is provided.

[0109] The device 700 includes, but is not limited to: a data preprocessing module 710, a factor characteristic evaluation module 730, a factor graph network establishment module 750, and a path prediction module 770.

[0110] The data preprocessing module 710 is used to acquire genes and cells from multiple samples, perform feature screening on the genes and cells in each sample, and obtain genes and cells used to indicate the characteristics of each sample.

[0111] The factor characteristic evaluation module 730 is used to construct gene sample graph networks and cell sample graph networks based on genes and cells used to indicate the characteristics of each sample, and to calculate the importance of the features of each node in the gene sample graph network and cell sample graph network, respectively.

[0112] The factor graph network building module 750 is used to construct an initial factor graph network based on the interaction between genes and cells, as well as the characteristics and importance of each node in the gene sample graph network and the cell sample graph network.

[0113] The path prediction module 770 is used to predict whether there are other genes and cells with interaction relationships in the factor graph network based on the adjacency matrix of the factor graph network, and to update the factor graph network based on the predicted genes and cells with interaction relationships.

[0114] In one exemplary embodiment, Figure 8 The flowchart shows a graph network construction device based on gene-cell relationship in an application scenario. The graph network construction device based on gene-cell relationship includes a data preprocessing module 810, a factor characteristic evaluation module 830, a factor graph network establishment module 850, and a path prediction module 870.

[0115] Specifically, the data preprocessing module 810 acquires genes and cells from multiple samples, performs feature screening on the genes and cells in each sample, and obtains genes and cells used to indicate the characteristics of each sample. The factor characteristic evaluation module 830 constructs a gene sample graph network and a cell sample graph network based on the genes and cells used to indicate the characteristics of each sample, and calculates the importance of the features of each node in the gene sample graph network and the cell sample graph network, respectively. The factor graph network building module 850 constructs an initial factor graph network based on the genes and cells with interaction relationships, as well as the features and importance of each node in the gene sample graph network and the cell sample graph network. The path prediction module 870 predicts whether there are other genes and cells with interaction relationships in the factor graph network based on the adjacency matrix of the factor graph network, and updates the factor graph network based on the predicted genes and cells with interaction relationships.

[0116] It should be noted that the above-described graph network construction device based on gene-cell relationships is only illustrated by the division of the above-described functional modules when constructing a graph network based on gene-cell relationships. In actual applications, the above functions can be assigned to different functional modules as needed. That is, the internal structure of the graph network construction device based on gene-cell relationships will be divided into different functional modules to complete all or part of the functions described above.

[0117] Furthermore, the embodiments of the graph network construction device based on gene-cell relationship and the graph network construction method based on gene-cell relationship provided in the above embodiments belong to the same concept. The specific way in which each module performs operations has been described in detail in the method embodiments, and will not be repeated here.

[0118] Figure 9 A schematic diagram of the structure of an electronic device according to an exemplary embodiment is shown.

[0119] It should be noted that this electronic device is merely an example adapted to this application and should not be construed as providing any limitation on the scope of use of this application. Furthermore, this electronic device should not be interpreted as requiring or depending on any specific feature. Figure 9 One or more components of the exemplary electronic device 2000 shown.

[0120] The hardware structure of electronic devices 2000 can vary significantly due to differences in configuration or performance, such as... Figure 9 As shown, the electronic device 2000 includes: a power supply 210, an interface 230, at least one memory 250, and at least one central processing unit (CPU) 270.

[0121] Specifically, power supply 210 is used to provide operating voltage for various hardware devices on electronic device 2000.

[0122] Interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices.

[0123] Of course, in other examples adapted in this application, interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input / output interface 235, and at least one USB interface 237, etc. Figure 9 As shown, this does not constitute a specific limitation.

[0124] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored on it include the operating system 251, application programs 253, and data 255, etc., and the storage method can be temporary storage or permanent storage.

[0125] The operating system 251 is used to manage and control the various hardware devices and application programs 253 on the electronic device 2000, so as to enable the central processing unit 270 to perform calculations and processing on the massive data 255 in the memory 250. It can be Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0126] Application 253 is a computer program that performs at least one specific task based on operating system 251, and may include at least one module ( Figure 9 (Not shown), each module may contain a computer program for the electronic device 2000. For example, a graph network construction device based on gene-cell relationships can be considered as an application program 253 deployed on the electronic device 2000.

[0127] Data 255 can be photos, pictures, etc. stored on a disk, or samples, etc., stored in memory 250.

[0128] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read computer programs stored in the memory 250, thereby enabling the computation and processing of massive amounts of data 255 in the memory 250. For example, a graph network construction method based on gene-cell relationships can be implemented by the central processing unit 270 reading a series of computer programs stored in the memory 250.

[0129] Furthermore, this application can also be implemented through hardware circuits or a combination of hardware circuits and software. Therefore, the implementation of this application is not limited to any specific hardware circuit, software, or combination thereof.

[0130] Please see Figure 10 This application provides an electronic device 4000, which may include: a desktop computer, a laptop computer, a server, etc.

[0131] exist Figure 10 The electronic device 4000 includes at least one processor 4001, at least one communication bus 4002, and at least one memory 4003.

[0132] The processor 4001 and memory 4003 are connected, for example, via a communication bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.

[0133] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0134] The communication bus 4002 may include a path for transmitting information between the aforementioned components. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0135] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0136] The memory 4003 stores a computer program, and the processor 4001 reads the computer program stored in the memory 4003 through the communication bus 4002.

[0137] When the computer program is executed by the processor 4001, it implements the graph network construction method based on gene-cell relationship in the above embodiments.

[0138] Furthermore, this application provides a storage medium storing a computer program, which, when executed by a processor, implements the gene-cell relationship-based graph network construction method described in the above embodiments.

[0139] This application provides a computer program product comprising a computer program stored in a storage medium. A processor of a computer device reads the computer program from the storage medium and executes the computer program, causing the computer device to perform the gene-cell relationship-based graph network construction method described in the above embodiments.

[0140] Compared with related technologies, the beneficial effects of the present invention are:

[0141] 1. This invention provides a graph network construction method based on gene-cell relationships, which enables the prediction of new gene-cell interaction pairs based on known gene-cell relationships. It not only indicates whether genes and cells have an interaction relationship, but also quantitatively provides the interaction strength. This method is particularly suitable for tumor microenvironment-related analysis.

[0142] 2. This invention proposes a gene-cell pair interaction relationship, which is displayed through a factor graph network. It considers both the characteristics and importance of genes and cells. In the process of constructing the sample graph network, samples with missing labels are supplemented, overcoming the shortcomings of sample label supplementation in the prior art.

[0143] 3. In the prediction of gene-cell interaction relationships, this invention considers not only the node features in the graph network, i.e., the features of genes and cells, but also the intensity of gene-cell interaction. New possible gene-cell pairs are predicted using the dot product of genes and cells, and the interaction intensity of new gene-cell pairs is indicated by weight matrix completion.

[0144] 4. This invention can be extended to other factors that interact in actual biological processes but for which no further correlations remain to be discovered, and to networks such as PPI networks, where interactions between genes can still be explored beyond the database. Furthermore, this invention provides customized gene-cell mapping networks for different types and characteristics of tumor microenvironments, incorporating gene and cell characteristics of single samples, which is beneficial for providing more useful information for personalized clinical research.

[0145] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0146] The above description is merely a preferred exemplary embodiment of the present invention and is not intended to limit the implementation of the present invention. Those skilled in the art can easily make corresponding modifications or alterations based on the main concept and spirit of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of protection claimed in the claims.

Claims

1. A method for constructing graph networks based on gene-cell relationships, characterized in that, The method includes: Genes and cells from multiple samples are obtained, and feature screening is performed on the genes and cells in each sample to obtain genes and cells used to indicate the characteristics of each sample; Gene sample graph network and cell sample graph network are constructed based on the genes and cells used to indicate the characteristics of each sample, and the importance of each node feature in the gene sample graph network and cell sample graph network is calculated respectively. Using genes or cells as nodes, pathways are established between different nodes based on the interaction between genes and cells; For genes or cells that have interactions, if the difference between gene importance and cell importance is less than a threshold, the weights of each pathway are calculated based on gene characteristics, cell characteristics, and gene importance and cell importance. Otherwise, update gene importance and cell importance based on the difference between gene importance and cell importance, and calculate the weight of each path based on gene characteristics, cell characteristics, and the updated gene importance and cell importance. An initial factor graph network is constructed from each node, the paths between nodes, and their weights. Based on the adjacency matrix of the factor graph network, it is predicted whether there are other genes and cells with interaction relationships in the factor graph network, and the factor graph network is updated based on the predicted genes and cells with interaction relationships.

2. The method as described in claim 1, characterized in that, The process of acquiring genes and cells from multiple samples, and performing feature screening on the genes and cells in each sample to obtain genes and cells used to indicate the characteristics of each sample includes: Genes from multiple samples are obtained, and the genes from each sample are screened to obtain genes used to indicate sample characteristics; Based on the expression profiles of the genes in each sample, the composition data of the cells in each sample are obtained, and the cell infiltration is assessed based on the cell composition data. The cells are screened based on the infiltration assessment results to obtain cells used to indicate sample characteristics.

3. The method as described in claim 1, characterized in that, The construction of gene sample map networks and cell sample map networks based on genes and cells used to indicate the characteristics of each sample includes: Calculate the similarity between unlabeled samples and between unlabeled and labeled samples; Using samples as nodes, paths are established between different nodes based on the calculated similarity and whether the labels carried by the samples are the same. Using the features of genes and cells that indicate the characteristics of each sample as node features, gene sample graph networks with gene features as node features and cell sample graph networks with cell features as node features are obtained respectively.

4. The method as described in claim 1, characterized in that, The calculation of the importance of each node feature in the gene sample map network and the cell sample map network, respectively, includes: The correlation between nodes in the gene sample graph network is calculated using a pixel-level interpretation algorithm to obtain the importance of the features of each node in the gene sample graph network; The correlation between nodes in the cell sample graph network is calculated using a pixel-level interpretation algorithm to obtain the importance of each node's features in the cell sample graph network.

5. The method as described in claim 1, characterized in that, The step of predicting whether there are other genes and cells with interaction relationships in the factor graph network based on the adjacency matrix of the factor graph network, and updating the factor graph network based on the predicted genes and cells with interaction relationships, includes: The features of each node in the factor graph network are standardized based on the adjacency matrix of the factor graph network, and the standardization results are sorted. Based on the sorting results, predict whether there are still genes and cells with interaction relationships in the factor graph network, establish paths between the predicted genes and cells with interaction relationships, and obtain new paths in the factor graph network. The weight of the newly added path is calculated based on the weight of each path in the factor graph network, and the factor graph network is updated based on the newly added path and its weight.

6. A graph network construction device based on gene-cell relationships, characterized in that, The device includes: The data preprocessing module is used to obtain genes and cells from multiple samples, perform feature screening on the genes and cells in each sample, and obtain genes and cells used to indicate the characteristics of each sample. The factor characteristic evaluation module is used to construct gene sample graph networks and cell sample graph networks based on genes and cells used to indicate the characteristics of each sample, and to calculate the importance of the features of each node in the gene sample graph network and cell sample graph network, respectively. The factor graph network construction module is used to establish paths between different nodes, using genes or cells as nodes, based on genes and cells with interaction relationships. For genes or cells with interaction relationships, if the difference between gene importance and cell importance is less than a threshold, the weight of each path is calculated based on gene features, cell features, and gene and cell importance. Otherwise, the gene and cell importance are updated based on the difference between gene and cell importance, and the weight of each path is calculated based on gene features, cell features, and the updated gene and cell importance. The initial factor graph network is constructed from each node, the paths between nodes, and their weights. The path prediction module is used to predict whether there are other genes and cells with interaction relationships in the factor graph network based on the adjacency matrix of the factor graph network, and to update the factor graph network based on the predicted genes and cells with interaction relationships.

7. An electronic device, characterized in that, include: The system includes at least one processor, at least one memory, and at least one communication bus, wherein the memory stores a computer program, and the processor reads the computer program from the memory via the communication bus. When the computer program is executed by the processor, it implements the graph network construction method based on gene-cell relationships as described in any one of claims 1 to 5.

8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the graph network construction method based on gene-cell relationships as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for analyzing active pathway in single cell multi-omics based on graph neural network

    CN115240772A