Method, device and equipment for identifying intercell communication relationship and storage medium

By acquiring cellular gene expression data and constructing hypothetical models, the causal relationships between cells are identified, solving the problem of insufficient causal explanatory power of prior databases in identifying intercellular communication relationships, and realizing the accurate identification and evaluation of intercellular communication relationships.

CN120808884BActive Publication Date: 2025-12-16HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511254422.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-12-16
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

In existing technologies, prior databases do not have a rigorous communication representation of receptor and ligand co-occurrence, making it difficult to accurately explain the actual regulatory or signal transduction processes of ligand and receptor, thus making the identification of intercellular communication relationships lack causal explanation.

Method used

By acquiring gene expression data from different cell types, constructing different hypothetical models using covariate gene sets, determining the causal relationship between source cell genes and target cell genes, and calculating communication strength, we can identify causal regulatory relationships between cells and eliminate interference within the target cell.

Benefits of technology

It enables causal explanation of intercellular communication relationships, accurately assesses the overall regulatory effects among cell populations, and improves the accuracy of intercellular communication relationship identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808884B_ABST
    Figure CN120808884B_ABST
Patent Text Reader

Abstract

The application discloses a kind of intercellular communication relationship identification method, device, equipment and storage medium, involve bioinformatics technical field, the causal relationship of source cell gene to target cell gene is used to quantify the direct causal contribution between cells, can let intercellular communication relationship have the explanatory power of causal level, and then accurately assess the overall regulation effect between cell population.The method comprises: obtaining gene expression data of different cell types, different cell types at least include source cell and target cell;For the candidate gene of target cell, use covariate gene set to construct different hypothesis models, to determine the causal relationship of source cell gene to target cell gene by different hypothesis models;According to the causal relationship of source cell gene to target cell gene, the communication intensity of source cell to target cell is calculated, to identify the causal regulation relationship between cells by communication intensity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and in particular to a method and device for identifying cell communication relationship, equipment and storage medium. BACKGROUND

[0002] Cell communication is a basic process for maintaining homeostasis, development and immune regulation of multicellular organisms. With the development of single-cell RNA sequencing technology, researchers can analyze the interaction between cells at the single-cell level.

[0003] In order to mine the potential communication mechanism between cells, the related technology can identify the communication relationship between cells through a cell communication analysis tool. Specifically, in the process of identifying the cell communication relationship, the cell communication analysis tool analyzes the expression pattern of ligand and receptor pairs in the single-cell transcription array, and then infers the communication relationship between cells according to the prior database formed by the receptor and ligand pairs. By comparing the co-expression relationship of ligands and receptors between different cell types, the potential communication pairs between cells are predicted. However, this inference method is highly dependent on the prior database, and the co-occurrence of receptors and ligands in the prior database does not have a rigorous communication expression, which makes it difficult to accurately explain the actual regulatory or signal transduction process of ligands and receptors, and the process of identifying the communication relationship between cells lacks causal explanation. SUMMARY

[0004] Therefore, the present application provides a method and device for identifying cell communication relationship, equipment and storage medium, which mainly aims to solve the problem that the co-occurrence of receptors and ligands in the prior database does not have a rigorous communication expression, which makes it difficult to accurately explain the actual regulatory or signal transduction process of ligands and receptors, and the process of identifying the communication relationship between cells lacks causal explanation.

[0005] In a first aspect, a method for identifying cell communication relationship is provided, the method comprising:

[0006] Obtaining gene expression data of different cell types, the gene expression data being the expression amount of each gene in the cell's own genome, and the different cell types including at least a source cell and a target cell;

[0007] For the candidate genes of the target cell, different hypothesis models are constructed using a set of covariate genes to determine the causal relationship between the source cell genes and the target cell genes through the different hypothesis models, the set of covariate genes being a set of genes in the target cell that are significantly related to the candidate genes;

[0008] According to the causal correlation between the source cell genes and the target cell genes, a communication strength of the source cell to the target cell is calculated to identify a communication relationship between cells through the communication strength.

[0009] Further, the gene expression data of different cell types is obtained, including:

[0010] Single-cell expression data of different cell types is obtained, the single-cell expression data including source cell data and target cell data, and the single-cell expression data being in the form of a single-cell transcriptome expression matrix to describe the expression amount of genes in cells;

[0011] The single-cell expression data is preprocessed to obtain gene expression data meeting a set format requirement, the preprocessing including at least one of format conversion, quality control, and standardization.

[0012] Further, before the candidate genes for the target cell are used to construct different hypothesis models using a covariate gene set to determine the causal correlation between the source cell genes and the target cell genes through the different hypothesis models, the method further includes:

[0013] For the candidate target genes of the target cell, a correlation coefficient between the candidate target genes and other genes of the target cell is calculated;

[0014] According to the correlation coefficient, a covariate gene set significantly correlated with the candidate target genes in the target cell is screened.

[0015] Further, the candidate genes for the target cell are used to construct different hypothesis models using a covariate gene set to determine the causal correlation between the source cell genes and the target cell genes through the different hypothesis models, including:

[0016] The candidate genes for the target cell are used as gene variables of the target cell, and different sample data sets are constructed using the covariate gene set to train the gene variables of the target cell, to obtain a first hypothesis model and a second hypothesis model;

[0017] During the training of the first hypothesis model and the second hypothesis model, a first prediction error of the first hypothesis model obtained through multiple iterations and a second prediction error of the second hypothesis model obtained through multiple iterations are respectively acquired, the first prediction error being used to represent a prediction error of the target cell genes under regression modeling without the source cell genes, and the second prediction error being used to represent a prediction error of the target cell genes under regression modeling with the source cell genes;

[0018] determine a causal correlation between the source cell gene and the target cell gene according to the first prediction error and the second prediction error.

[0019] Further, the candidate gene of the target cell is taken as a gene variable of the target cell, and different sample data sets are constructed using the covariate gene set to train a model of the gene variable of the target cell, to obtain a first hypothesis model and a second hypothesis model, including:

[0020] The candidate gene of the target cell is taken as a gene variable of the target cell, and regression fitting is performed on the gene variable of the target cell using the covariate gene set as a sample data set, to obtain a first hypothesis model.

[0021] The candidate gene of the target cell is taken as a gene variable of the target cell, and regression modeling is performed on the gene variable of the target cell using the covariate gene set and the gene variable of the source cell as a sample data set, to obtain a second hypothesis model.

[0022] Further, the determination of the causal correlation between the source cell gene and the target cell gene according to the first prediction error and the second prediction error includes:

[0023] The causal strength score of the source cell gene to the target cell gene is calculated according to the first prediction error and the second prediction error.

[0024] If the causal strength score is greater than a set value, it is determined that the source cell gene has a causal effect on the target cell gene.

[0025] If the causal strength score is less than or equal to a set value, it is determined that the source cell gene does not have a causal effect on the target cell gene.

[0026] Further, the calculation of the communication strength of the source cell to the target cell according to the causal correlation between the source cell gene and the target cell gene includes:

[0027] According to the causal correlation between the source cell gene and the target cell gene, a cross-cell-type gene regulation network is constructed, the gene regulation network uses an adjacency matrix to describe the causal correlation between the source cell gene and the target cell gene, if the source cell gene has a causal effect on the target cell gene, the corresponding element in the adjacency matrix is assigned a first value, if the source cell gene does not have a causal effect on the target cell gene, the corresponding element in the adjacency matrix is assigned a second value.

[0028] The communication intensity of the source cell to the target cell is calculated by performing a weighted sum calculation on all elements in the adjacency matrix, so as to identify the communication relationship between cells through the communication intensity.

[0029] In a second aspect, a device for identifying a communication relationship between cells is provided, and the device comprises:

[0030] An acquisition unit is configured to acquire gene expression data of different cell types, the gene expression data being expression amounts of respective genes in a genome of a cell, and the different cell types at least including a source cell and a target cell;

[0031] A construction unit is configured to construct different hypothetical models using a set of covariate genes for a candidate gene of the target cell, so as to determine a causal correlation relationship of a source cell gene to a target cell gene through the different hypothetical models, the set of covariate genes being a set of genes in the target cell that are significantly correlated with the candidate gene;

[0032] An identification unit is configured to calculate a communication intensity of the source cell to the target cell according to the causal correlation relationship of the source cell gene to the target cell gene, so as to identify the communication relationship between cells through the communication intensity.

[0033] Further, the acquisition unit is specifically configured to:

[0034] Acquire single-cell expression data of different cell types, the single-cell expression data including source cell data and target cell data, and the single-cell expression data being in a form of a single-cell transcriptome expression matrix to describe expression amounts of genes in cells;

[0035] Preprocess the single-cell expression data to obtain gene expression data meeting a set format requirement, and the preprocessing includes at least one of format conversion, quality control, and standardization.

[0036] Further, the device further comprises:

[0037] A calculation unit is configured to calculate correlation coefficients between a candidate target gene and other genes of the target cell for the candidate target gene of the target cell, before the construction of different hypothetical models using a set of covariate genes for the candidate gene of the target cell, so as to determine a causal correlation relationship of a source cell gene to a target cell gene through the different hypothetical models.

[0038] A screening unit is configured to screen a set of covariate genes that are significantly correlated with the candidate target gene in the target cell according to the correlation coefficients.

[0039] Further, the construction unit comprises:

[0040] constructing a module, configured to use the candidate gene of the target cell as a gene variable of the target cell, use the covariate gene set to construct different sample data sets for model training of the gene variable of the target cell, and obtain a first hypothesis model and a second hypothesis model;

[0041] an acquisition module, configured to acquire a first prediction error of the first hypothesis model obtained through multiple iterations and a second prediction error of the second hypothesis model obtained through multiple iterations in the process of training the first hypothesis model and the second hypothesis model, the first prediction error being used to represent a prediction error of the target cell gene under regression modeling without the source cell gene, and the second prediction error being used to represent a prediction error of the target cell gene under regression modeling with the source cell gene;

[0042] a determination module, configured to determine a causal correlation between the source cell gene and the target cell gene according to the first prediction error and the second prediction error.

[0043] Further, the constructing module is specifically configured to:

[0044] use the covariate gene set as the gene variable of the target cell for regression fitting to obtain the first hypothesis model;

[0045] use the covariate gene set as the gene variable of the target cell for regression fitting to obtain the first hypothesis model;

[0046] Further, the determination module is specifically configured to:

[0047] calculate a causal strength score of the source cell gene to the target cell gene according to the first prediction error and the second prediction error;

[0048] if the causal strength score is greater than a set value, it is determined that the source cell gene has a causal effect on the target cell gene;

[0049] if the causal strength score is less than or equal to the set value, it is determined that the source cell gene has no causal effect on the target cell gene.

[0050] Further, the recognition unit is specifically configured to:

[0051] construct a gene regulation network across cell types according to the causal correlation of the source cell genes to the target cell genes, the gene regulation network using an adjacency matrix to describe the causal correlation of the source cell genes to the target cell genes, if there is a causal action relationship between the source cell genes and the target cell genes, the corresponding element in the adjacency matrix is assigned a first value, if there is no causal action relationship between the source cell genes and the target cell genes, the corresponding element in the adjacency matrix is assigned a second value;

[0052] perform weighted sum calculation on all elements in the adjacency matrix to obtain the communication strength of the source cell to the target cell, so as to identify the communication relationship between cells through the communication strength.

[0053] In a third aspect, a storage medium having a computer program stored thereon is provided, the program being executed by a processor to implement the above-mentioned method for identifying the communication relationship between cells.

[0054] In a fourth aspect, a device for identifying the communication relationship between cells is provided, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, the processor executing the program to implement the above-mentioned method for identifying the communication relationship between cells.

[0055] By the above technical solutions, the method, device, equipment and storage medium for identifying the communication relationship between cells provided by the present application are compared with the current way of identifying the communication relationship between cells through a priori database. The present application acquires gene expression data of different cell types, the gene expression data being the expression amount of each gene in the cell's own genome, and different cell types at least including source cells and target cells. For candidate genes of the target cells, different hypothesis models are constructed using a set of covariate genes to determine the causal correlation of the source cell genes to the target cell genes through different hypothesis models, the set of covariate genes being a set of genes in the target cells that are significantly related to the candidate genes. According to the causal correlation of the source cell genes to the target cell genes, the communication strength of the source cell to the target cell is calculated to identify the causal regulation relationship between cells through the communication strength. The whole process excludes the regulatory interference of the source cell to the target cell by controlling the set of covariate genes expressed in the background of the target cells, so as to accurately acquire the causal correlation of the source cell genes to the target cell genes, and use the causal correlation of the source cell genes to the target cell genes to quantify the direct causal contribution between cells, so that the communication relationship between cells has causal interpretation, and the overall regulation effect between cell populations can be accurately evaluated.

[0056] The above description is only a summary of the technical solutions of the present application. In order to make the technical means of the present application more clearly understood and implemented according to the content of the description, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0057] The accompanying drawings, which are included to provide a further understanding of the present application, form a part of the present application and illustrate the illustrative embodiments of the present application and its description serve to explain the present application. The accompanying drawings do not constitute an undue limitation on the present application. In the drawings:

[0058] Figure 1 is a flowchart of the method for identifying the intercell communication relationship in an embodiment of the present application;

[0059] Figure 2 is a flowchart of the method for identifying the intercell communication relationship in an embodiment of the present application; Figure 1 is a flowchart of a specific embodiment of step 101 in the method;

[0060] Figure 3 is a flowchart of the method for identifying the intercell communication relationship in another embodiment of the present application;

[0061] Figure 4 is a flowchart of the method for identifying the intercell communication relationship in another embodiment of the present application; Figure 1 is a flowchart of a specific embodiment of step 102 in the method;

[0062] Figure 5 is a flowchart of a specific embodiment of step 103 in the method; Figure 1

[0063] is a flowchart of the method for identifying the intercell communication relationship in another embodiment of the present application; Figure 6

[0064] is a structural diagram of the device for identifying the intercell communication relationship in an embodiment of the present application; Figure 7

[0065] is a structural diagram of the device for identifying the intercell communication relationship in an embodiment of the present application; Figure 8 DETAILED DESCRIPTION

[0066] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0067] ​Related technologies utilize cell communication analysis tools to identify intercellular communication relationships. Specifically, in this process, these tools analyze the expression patterns of ligand-receptor pairs in single-cell transcriptomes and then infer intercellular communication relationships based on a priori database of receptor-ligand pairs. By comparing the co-expression relationships of ligands and receptors across different cell types, potential intercellular communication pairs can be predicted. However, this inference method is highly dependent on the priori database, which does not provide a rigorous representation of receptor-ligand co-occurrence. This makes it difficult to accurately explain the actual regulatory or signal transduction processes involved, thus lacking causal explanatory power in the identification of intercellular communication relationships.

[0068] To address this problem, this embodiment provides a method for identifying intercellular communication relationships, such as... Figure 1 As shown, it includes the following steps:

[0069] 101. Obtain gene expression data for different cell types.

[0070] In this embodiment, gene expression data refers to the expression levels of each gene in the cell's own genome, and different cell types include at least source cells and target cells. Generally speaking, source cells and target cells are cell types with a conversion relationship or alignment relationship.

[0071] As one method of acquisition, it can be obtained through experimental detection. First, cell separation technology is used to separate different cell types. Intermediate products of gene expression are extracted from the separated cells. Gene expression information is analyzed based on the intermediate products of gene expression to obtain gene expression data for different cell types. Here, it is necessary to conduct experimental detection on the source cells and target cells separately to obtain the gene expression data of the source cells and the gene expression data of the target cells accordingly.

[0072] As another way to obtain the data, it is possible to search public databases. Since public databases contain a large amount of gene expression data from different species and cell types, gene expression data of different cell types can be retrieved from public databases based on cell type. Here, it is necessary to search public databases separately for source cells and target cells to obtain the gene expression data of source cells and target cells respectively.

[0073] Furthermore, to remove noise and contamination, and to facilitate subsequent cell analysis, the gene expression data can be quality controlled and standardized after acquisition to obtain structured gene expression data. Typically, gene expression data is in matrix form; after quality control and standardization, the gene and sample dimensions in the gene expression matrices of the source and target cells are consistent.

[0074] 102. For the candidate gene of the target cell, different hypothesis models are constructed using the covariate gene set to determine the causal relationship between the source cell gene and the target cell gene through the different hypothesis models.

[0075] In this embodiment, the covariate gene set is a gene set in the target cell that is significantly correlated with the candidate gene. Generally, gene expression data is mainly used to reflect the expression of genes in the cell. For any gene in the target cell, the covariate gene set can be obtained by analyzing the significant correlation between genes, i.e., the covariate gene set. Here, a suitable correlation index can be selected to quantify the correlation between the expressions of two genes, and the covariate gene set can be screened in the target cell according to the correlation between the expressions of the two genes.

[0076] It can be understood that after obtaining the covariate gene set, in order to quantify the regulation of the source cell gene on the target cell gene, different hypothesis models can be constructed on the basis of controlling the internal interference of the target cell. Here, the different hypothesis models include but are not limited to a hypothesis model without regulation and a hypothesis model with direct regulation. For the hypothesis model without regulation, it is assumed that the source cell gene does not regulate the target cell gene, i.e., the gene expression variable of the target cell is only determined by the covariate gene set in the target gene. For the hypothesis model with direct regulation, it is assumed that the source cell gene regulates the target cell gene directly. The direct regulation can be indirectly realized by relying on the covariate gene set in the target cell, or can not rely on the covariate gene set, i.e., the gene expression variable of the target cell is introduced on the basis of the covariate gene set in the target gene.

[0077] In this embodiment, the difference between the different hypothesis models is whether the gene expression variable of the target cell contains the source cell gene. This difference is mainly used to capture the direct regulation signal that does not depend on the covariate gene set in the target cell. That is, by comparing the fitting errors of different hypothesis models, it can be verified whether the cross-cell signal of the source cell gene is necessary to explain the variation of the target cell gene on the basis of controlling the internal interference of the target cell, i.e., to quantify and infer the causal relationship between the source cell gene and the target cell gene.

[0078] Specifically, if part of the variation of the target cell gene cannot be fully explained by the covariate gene set in the target cell, at this time, the direct regulation hypothesis model can capture this part of the intercellular signal due to the inclusion of the source cell gene, and the explanation ability of the target cell gene variable is better than that of the hypothesis model without regulation, which is manifested as a small fitting error, indicating that the source cell has a causal relationship with the target cell. On the contrary, if all the variations of the target cell can be explained by the covariate gene set in the target cell, the direct effect of the source cell gene is redundant, at this time, the fitting error of the direct regulation hypothesis model and the hypothesis model without regulation is close, and even due to the introduction of redundant parameters in the direct regulation hypothesis model, fitting noise is generated, so that the fitting effect of the hypothesis model without regulation is better, indicating that the source cell does not have a causal relationship with the target cell.

[0079] 103. According to the causal relationship between the source cell gene and the target cell gene, the communication strength of the source cell to the target cell is calculated to identify the communication relationship between the cells through the communication strength.

[0080] It can be understood that the causal regulation strength of a single gene pair cannot reflect the communication strength of the source cell to the target cell, and the communication strength of the source cell to the target cell needs to comprehensively reflect the overall causal regulation ability of the source cell to the target cell. That is, the communication strength of the source cell to the target cell is comprehensively considered the causal regulation relationship of the gene pairs formed by the source cell and the target cell and the contribution weight of the gene pairs.

[0081] In this embodiment, the source cell gene to the target cell gene can be quantified as a causal pair, and different values can be assigned to each gene pair according to the causal relationship. Correspondingly, the communication strength of the source cell to the target cell can be obtained by weighted integration or weighted average of multiple causal pair assignments. For example, the gene set in the source cell is S, and the gene set in the target cell is T. For each significant causal pair , the causal strength is , and correspondingly, the communication strength C of the source cell to the target cell can be defined as: .

[0082] In actual application scenarios, if the importance of the target gene in cell function is different, for example, the regulation significance of key transcription factors, signal molecules and ordinary genes is different, a gene weight can be further introduced. At this time, the communication strength of the source cell to the target cell can be obtained by weighted integration or weighted average of multiple causal pair assignments on the basis of the gene weight. Continue to illustrate the above example, the weight of the target gene is , wherein is the proportion of the expression amount of the target gene to the total expression of the target cell, or The core degree in the pathway. Correspondingly, the communication intensity C of the source cell to the target cell can be corrected as: .

[0083] It can be understood that the communication intensity of the source cell to the target cell is an important quantitative index for judging the intercellular communication relationship. Generally speaking, the greater the communication intensity of the source cell to the target cell, the closer the communication relationship between the source cell and the target cell, and vice versa. The smaller the communication intensity of the source cell to the target cell, the weaker the communication relationship between the source cell and the target cell.

[0084] The intercellular communication relationship identification method provided by the embodiment of the application is compared with the current way of identifying the intercellular communication relationship through a priori database. The application obtains gene expression data of different cell types, the gene expression data being the expression amount of each gene in the genome of the cell itself, and the different cell types at least including the source cell and the target cell. For the candidate gene of the target cell, a different hypothesis model is constructed using a set of covariate genes to determine the causal relationship of the source cell gene to the target cell gene through the different hypothesis models, and the set of covariate genes being a set of genes in the target cell that are significantly related to the candidate gene. According to the causal relationship of the source cell gene to the target cell gene, the communication intensity of the source cell to the target cell is calculated to identify the causal regulation relationship between the cells through the communication intensity. The whole process excludes the regulatory interference of the source cell to the target cell by controlling the set of covariate genes expressed in the background of the target cell, so as to accurately obtain the causal relationship of the source cell gene to the target cell gene. The causal relationship of the source cell gene to the target cell gene is used to quantify the direct causal contribution between the cells, so that the intercellular communication relationship has causal explanation, and the overall regulation effect between the cell populations can be accurately evaluated.

[0085] In actual application scenarios, in order to remove low-quality cells and genes and reduce technical bias, when obtaining gene expression data of different types, the different cell expression data needs to be preprocessed. Specifically, as shown in Figure 2 , step 101 includes the following steps:

[0086] 201, obtaining single cell expression data of different cell types.

[0087] 202, preprocessing the single cell expression data to obtain gene expression data meeting the set format requirements.

[0088] The single-cell expression data includes source cell data and target cell data, and the single-cell expression data is in the form of a single-cell transcriptome expression matrix, which describes the expression of genes in cells. That is, whether it is source cell data or target cell data, the basic format of the expression matrix is consistent, and the gene expression can be quantified by forming a two-dimensional table of genes and cells. The elements in the expression matrix represent the expression of genes in the corresponding cells.

[0089] It can be understood that in the scenario of intercellular communication analysis, the source cell and the target cell are functionally divided according to the signal transmission direction, and the difference in the expression matrix is reflected in the cell type rather than the matrix structure. That is, the source cell and the target cell have the same format in the expression matrix, and the definitions of rows and columns are consistent, and the elements in the matrix are gene expression data. The difference is the expression characteristics of the genes in the cells.

[0090] In order to convert the original cell data into high-quality, standardized and format-compliant gene expression matrix, the single-cell expression data can be preprocessed to obtain gene expression data that meets the set format requirements. In this embodiment, the preprocessing includes at least one of format conversion, quality control and standardization. Generally, the original expression matrix of different cell types may have differences, and the expression matrix of different cell types can be converted to the same format by format conversion. Since there may be low-quality cells or low-expression genes in the original expression matrix, in order to filter low-quality cells and genes, comprehensive quality judgment can be made on different types of cell data through quality control, and cells and / or genes with low comprehensive quality are filtered. Further, the sequencing depth of different cells is different, in order to reduce technical noise, the expression between different cells can be made comparable by standardization. Common standardization methods include, but are not limited to, total expression normalization, logarithmic conversion and scaling, etc. The specific method can be selected according to the data type.

[0091] In the actual application scenario of intercellular communication analysis, the covariate gene set is the key to control interference and improve analysis reliability. In order to accurately exclude the interference in the target biological signal and ensure that the analysis result can accurately reflect the causal relationship between cells, further, as shown in Figure 3 Before step 102, the method further includes the following steps:

[0092] 301. For the candidate target gene of the target cell, calculate the correlation coefficient between the candidate target gene and other genes of the target cell.

[0093] 302. According to the correlation coefficient, screen a covariate gene set that is significantly related to the candidate target gene in the target cell.

[0094] In the embodiment, the candidate target gene is any gene in the target cell, and the consistency of the expression change trend of the candidate target gene and other genes can be quantified by the correlation coefficient, so as to determine whether there is co-regulation or functional association between the genes. For positively correlated genes A and B, when the expression of gene A increases, the expression of gene B also increases. For negatively correlated genes A and B, when the expression of gene A increases, the expression of gene B decreases. For genes A and B without significant correlation, the expression change of genes A and B has no significant correlation, and the functional association is low.

[0095] Specifically, for the candidate target gene of the target cell, the expression vector of the candidate target gene can be extracted from the matrix expression of the target cell, and then the expression vectors of other genes are extracted in sequence. According to the expression vector of the candidate target gene and the expression vectors of other genes, the correlation coefficient between the candidate target gene and other genes is calculated. The process can use the following formula:

[0096]

[0097] wherein, is the covariance, , is the expression vector of the two genes.

[0098] It can be understood that the correlation coefficient is a statistical quantity for quantifying the correlation degree of the expression pattern of the candidate target gene and other genes. The absolute value of the correlation coefficient is close to 1, indicating that the expression change of the two genes is more synchronous. The absolute value of the correlation coefficient is close to 0, indicating that the correlation is weaker. Here, a coefficient threshold can be set. Correspondingly, for the genes with a correlation coefficient greater than the coefficient threshold, the genes can be added to the covariant gene set as the genes significantly correlated with the candidate target gene.

[0099] It should be noted that if the candidate target gene is a set, the correlation between each gene in the set and all other genes can be calculated, and then the results are integrated by taking the mean value, voting method, etc. to screen out the gene set related to the whole candidate target gene set, and obtain the covariant gene set.

[0100] In actual application scenarios, in order to cope with the multi-dimensional complexity of cell-cell communication, the source cell and the target cell are often disturbed by multiple non-communication direct related factors. These factors can be captured by the covariant gene set. At this time, for multiple potential patterns existing between cells, different hypothetical models can be constructed. Specifically, as shown in FIG. 1, step 102 includes the following steps: Figure 4

[0101] ​401. Using the candidate genes of the target cell as the gene variables of the target cell, construct different sample datasets using the covariate gene set to train the model on the gene variables of the target cell, and obtain the first hypothesis model and the second hypothesis model.

[0102] 402. During the training of the first hypothetical model and the second hypothetical model, the first prediction error obtained by the first hypothetical model after multiple iterations and the second prediction error obtained by the second hypothetical model after multiple iterations are obtained respectively.

[0103] 403. Based on the first prediction error and the second prediction error, determine the causal relationship between the source cell gene and the target cell gene.

[0104] In this embodiment, the first prediction error represents the prediction error of the target cell gene under regression modeling without including the source cell gene. In this case, only the covariate gene set is used as the sample dataset in the model. Specifically, candidate genes of the target cell are used as the gene variables of the target cell, and the covariate gene set is used as the sample dataset of the target cell gene variables for regression fitting to obtain the first hypothetical model.

[0105] The first hypothetical model described above can be expressed using the following formula:

[0106]

[0107] in, For the genetic variables of the target cell, For the set of covariate genes, This represents the prediction error of the first hypothetical model.

[0108] In this embodiment, the second prediction error represents the prediction error of the target cell gene under regression modeling that includes the source cell gene. In this case, in addition to the covariate gene set, the source cell gene is also included as a sample dataset applied to the model. Specifically, candidate genes of the target cell are used as the target cell's gene variables, and the source cell's gene variables are introduced as a sample dataset based on the covariate gene set to perform regression modeling on the target cell's gene variables, resulting in the second hypothesis model.

[0109] The second hypothesis model described above can be expressed using the following formula:

[0110]

[0111] in, For the genetic variables of the target cell, For the set of covariate genes, Genetic variables of the source cell, is the prediction error of the second hypothetical model.

[0112] Specifically in the process of model training, the sample data set can be divided into a training set and a test set, the training set is used for regression, and the test set is used for error judgment. The division method can use K-fold cross validation, which effectively avoids the randomness bias of single division, making the average test error of the two models more reliable. In order to ensure the comparability of the results, the two hypothetical models need to be compared under the same division method, the same number of iterations, and the same error index.

[0113] It can be understood that the first prediction error and the second prediction error usually correspond to the prediction performance of two different hypothetical models. By comparing the error difference between the two, it can be inferred whether the source cell gene has a causal effect on the target cell gene. If the source cell gene has a causal relationship with the target cell gene, the second hypothetical model with the source cell gene should be able to more accurately predict the state of the target cell gene, and its prediction error should be significantly lower than the first hypothetical model without the source cell gene. On the contrary, if the source cell gene does not have a causal relationship with the target cell gene, the prediction ability of the second hypothetical model with the source cell gene does not improve compared to the first prediction model without the source cell gene. Specifically, in the process of determining the causal relationship between the source cell gene and the target cell gene according to the first prediction error and the second prediction error, the causal strength score of the source cell gene to the target cell gene can be calculated according to the first prediction error and the second prediction error; if the causal strength score is greater than a set value, it is determined that the source cell gene has a causal effect on the target cell gene; if the causal strength score is less than or equal to the set value, it is determined that the source cell gene does not have a causal effect on the target cell gene.

[0114] The above causal strength score can be represented by the following formula:

[0115]

[0116] wherein, is the prediction error of the target cell gene under a regression model that does not contain the source cell gene , is the prediction error of the target cell gene under a regression model that contains the source cell gene .

[0117] Correspondingly, if the prediction error is significantly reduced in the regression model with the introduction of the source cell gene , that is, , it is considered that the source cell gene has a causal effect on the target cell gene ; on the contrary, it is considered that the source cell gene target cell genes There is a causal relationship.

[0118] In actual application scenarios, cell-to-cell communication usually relies on the interaction between cells. For cells with causal correlation, communication can be achieved. Therefore, the causal correlation of the source cell gene to the target cell gene can be used to reflect the communication ability between cells. Specifically, as shown in Figure 5 Step 103 includes the following steps:

[0119] 501. According to the causal correlation of the source cell gene to the target cell gene, a gene regulation network across cell types is constructed.

[0120] 502. Weighted sum calculation is performed on all elements in the adjacency matrix to obtain the communication strength of the source cell to the target cell, so as to identify the communication relationship between cells through the communication strength.

[0121] In this embodiment, the gene regulation network uses an adjacency matrix to describe the causal correlation of the source cell gene to the target cell gene. If the source cell gene has a causal relationship with the target cell gene, the corresponding element in the adjacency matrix is assigned a first value. If the source cell gene has no causal relationship with the target cell gene, the corresponding element in the adjacency matrix is assigned a second value. The assignment of elements in the adjacency matrix can refer to the following formula:

[0122]

[0123] Further, in order to quantitatively evaluate the causal regulation ability of the source cell to the target cell, all elements in the critical matrix can be weighted and summed. Specifically, the existence of the causal relationship of the source cell gene to the target cell gene can be regarded as a directed edge, and the corresponding weighted sum process is to count all existing directed edges in the adjacency matrix. The result obtained by counting is the communication strength of the source cell to the target cell. The following formula can be used:

[0124]

[0125] Wherein, is the adjacency matrix representation of the gene regulation network, The larger the corresponding value is, the stronger the causal regulation ability of the source cell to the target cell is, and the stronger the communication relationship between cells is.

[0126] In actual application scenarios, the identification process of the cell-to-cell communication relationship can refer to Figure 6 After obtaining the gene expression data of the target cell and the source cell, the candidate target gene in the target cell is compared with other genes in the target cell Correlation calculation to screen a set of covariate genes significantly correlated with the candidate target gene, and cross-validation predictability according to the set of covariate genes , respectively construct different hypothesis models, obtain the causal correlation relationship of the source cell gene to the target cell gene through the prediction error of the different hypothesis models, construct a cross-cell type gene regulation network according to the causal correlation relationship, the gene regulation network is described using an adjacency matrix, if the source cell gene has a causal relationship with the target cell gene, the corresponding element in the adjacency matrix is assigned a value of 1, if the source cell gene has no causal relationship with the target cell gene, the corresponding element in the adjacency matrix is assigned a value of 0, and further weighted sum of all elements in the adjacency matrix is obtained Communication strength of source cell to target cell , accordingly, through the communication strength of the source cell to the target cell, the communication relationship between cells can be identified.

[0127] Further, as a specific implementation of the above method, the embodiment of the present application provides an intercellular communication relationship identification device, as shown in Figure 7 , the device comprises an acquisition unit 61, a construction unit 62 and an identification unit 63.

[0128] The acquisition unit 61 is used to acquire gene expression data of different cell types, the gene expression data being the expression amount of each gene in the cell's own genome, and the different cell types at least including source cells and target cells;

[0129] The construction unit 62 is used to construct different hypothesis models for the candidate gene of the target cell using a set of covariate genes, to determine the causal correlation relationship of the source cell gene to the target cell gene through the different hypothesis models, and the set of covariate genes being a set of genes in the target cell significantly correlated with the candidate gene;

[0130] The identification unit 63 is used to calculate the communication strength of the source cell to the target cell according to the causal correlation relationship of the source cell gene to the target cell gene, to identify the communication relationship between cells through the communication strength.

[0131] Compared with the prior art of identifying the intercellular communication relationship by using a prior database, the device for identifying the intercellular communication relationship provided by the embodiment of the application acquires gene expression data of different cell types, the gene expression data is the expression amount of each gene in the genome of the cell itself, and the different cell types at least include a source cell and a target cell; for a candidate gene of the target cell, different hypothesis models are constructed by using a set of covariate genes to determine the causal correlation relationship of the source cell gene to the target cell gene through the different hypothesis models, the set of covariate genes is a set of genes in the target cell that are significantly related to the candidate gene; and the communication strength of the source cell to the target cell is calculated according to the causal correlation relationship of the source cell gene to the target cell gene, so as to identify the causal regulation relationship between the cells through the communication strength. The whole process excludes the regulation interference of the source cell to the target cell by controlling the set of covariate genes expressed in the background of the target cell, so as to accurately acquire the causal correlation relationship of the source cell gene to the target cell gene, and the causal correlation relationship of the source cell gene to the target cell gene is used to quantify the direct causal contribution between the cells, so that the intercellular communication relationship has the explanatory power at the causal level, and the overall regulation effect between the cell populations can be accurately evaluated.

[0132] In a specific application scenario, the acquisition unit is specifically configured to:

[0133] Acquire single-cell expression data of different cell types, the single-cell expression data including source cell data and target cell data, and the single-cell expression data being in the form of a single-cell transcriptome expression matrix to describe the expression amount of a gene in a cell.

[0134] Preprocess the single-cell expression data to obtain gene expression data meeting a set format requirement, the preprocessing including at least one of format conversion, quality control and standardization.

[0135] In a specific application scenario, the device further includes:

[0136] The calculation unit is configured to, before the step of constructing different hypothesis models by using a set of covariate genes for the candidate gene of the target cell to determine the causal correlation relationship of the source cell gene to the target cell gene through the different hypothesis models, calculate the correlation coefficient between the candidate target gene and other genes of the target cell for the candidate target gene of the target cell.

[0137] The screening unit is configured to screen a set of covariate genes that are significantly related to the candidate target gene in the target cell according to the correlation coefficient.

[0138] In a specific application scenario, the construction unit includes:

[0139] constructing a module, configured to use the candidate gene of the target cell as a gene variable of the target cell, use the covariate gene set to construct different sample data sets for model training on the gene variable of the target cell, and obtain a first hypothesis model and a second hypothesis model;

[0140] an acquisition module, configured to acquire a first prediction error of the first hypothesis model obtained through multiple iterations and a second prediction error of the second hypothesis model obtained through multiple iterations in the process of training the first hypothesis model and the second hypothesis model, the first prediction error being used to represent a prediction error of the target cell gene under regression modeling without the source cell gene, and the second prediction error being used to represent a prediction error of the target cell gene under regression modeling with the source cell gene;

[0141] a determination module, configured to determine a causal correlation between the source cell gene and the target cell gene according to the first prediction error and the second prediction error.

[0142] In a specific application scenario, the constructing module is specifically configured to:

[0143] use the covariate gene set as the gene variable of the target cell for regression fitting to obtain the first hypothesis model;

[0144] introduce the gene variable of the source cell as a sample data set on the basis of the covariate gene set to perform regression modeling on the gene variable of the target cell to obtain the second hypothesis model.

[0145] In a specific application scenario, the determination module is specifically configured to:

[0146] calculate a causal strength score of the source cell gene on the target cell gene according to the first prediction error and the second prediction error;

[0147] if the causal strength score is greater than a set value, it is determined that the source cell gene has a causal effect on the target cell gene;

[0148] if the causal strength score is less than or equal to the set value, it is determined that the source cell gene has no causal effect on the target cell gene.

[0149] In a specific application scenario, the recognition unit is specifically configured to:

[0150] According to the causal correlation of the source cell gene to the target cell gene, a gene regulation network across cell types is constructed, the gene regulation network uses an adjacency matrix to describe the causal correlation of the source cell gene to the target cell gene, if the source cell gene has a causal action relationship to the target cell gene, the corresponding element in the adjacency matrix is assigned a first value, if the source cell gene has no causal action relationship to the target cell gene, the corresponding element in the adjacency matrix is assigned a second value;

[0151] All elements in the adjacency matrix are weighted and summed to obtain the communication strength of the source cell to the target cell, so as to identify the communication relationship between cells through the communication strength.

[0152] Based on the above method as shown in Figures 1-5 Accordingly, the embodiments of the present application also provide a storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for identifying the communication relationship between cells as shown in Figures 1-5

[0153] Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.

[0154] Based on the above method as shown in Figures 1-5 and the virtual device embodiment as shown in Figure 7 In order to achieve the above-mentioned purpose, the embodiments of the present application also provide an entity device for identifying the communication relationship between cells, which can be a computer, a smart phone, a tablet computer, a smart watch, a server, or a network device, etc. The entity device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the above-mentioned method for identifying the communication relationship between cells as shown in Figures 1-5

[0155] Optionally, the entity device can also include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface can include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. The optional user interface can also include a USB interface, a card reader interface, etc. The network interface can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.

[0156] In the exemplary embodiments, refer to Figure 8 ​​The entity device includes a communication bus, a processor, a memory, and a communication interface, and can further include an input / output interface and a display device, wherein the respective functional units can communicate with each other through the bus. The memory stores a computer program, and the processor is configured to execute the program stored in the memory to perform the method for identifying the intercellular communication relationship in the above embodiment.

[0157] Those skilled in the art can understand that the structure of the entity device for identifying the intercellular communication relationship provided in the embodiment does not constitute a limitation on the entity device, and can include more or fewer components, or combine certain components, or different component arrangements.

[0158] The storage medium can further include an operating system and a network communication module. The operating system is a program for managing hardware and software resources of the entity device for identifying the intercellular communication relationship, and supports the running of information processing programs and other software and / or programs. The network communication module is configured to realize communication between components in the storage medium, and communication with other hardware and software in the information processing entity device.

[0159] Those skilled in the art can clearly understand from the above description of the embodiments that the present application can be implemented by means of software and a necessary general hardware platform, or by hardware. By applying the technical solutions of the present application, compared with the existing manner, the present application can exclude the regulatory interference of the source cell on the target cell by controlling the covariate gene set expressed in the target cell, so as to accurately obtain the causal relationship of the source cell gene to the target cell gene, and use the causal relationship of the source cell gene to the target cell gene to quantify the direct causal contribution between cells, so that the intercellular communication relationship has causal interpretation, and the overall regulatory effect between cell populations can be accurately evaluated.

[0160] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or flows in the drawings are not necessarily required for implementing the present application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be changed and located in one or more devices different from the implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0161] The above serial numbers of the present application are only for description, and do not represent the advantages and disadvantages of the implementation scenario. The above disclosure is only several specific implementation scenarios of the present application, but the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.

Claims

1. A method for identifying intercellular communication relationships, characterized in that, include: Gene expression data of different cell types are obtained, wherein the gene expression data is the expression level of each gene in the cell's own genome, and the different cell types include at least source cells and target cells; For candidate genes in the target cells, different hypothetical models are constructed using a set of covariate genes to determine the causal relationship between source cell genes and target cell genes. The set of covariate genes is a set of genes in the target cells that are significantly related to the candidate genes. The different hypothetical models assume whether the source cell genes have a regulatory effect on the target cell genes, including a no-regulation hypothetical model and a direct-regulation hypothetical model. The no-regulation hypothetical model assumes that the source cell genes do not have a regulatory effect on the target cell genes, while the direct-regulation hypothetical model assumes that the source cell genes have direct regulation on the target cell genes. Based on the causal relationship between the source cell genes and the target cell genes, the communication strength between the source cell and the target cell is calculated, so as to identify the communication relationship between cells through the communication strength.

2. The method according to claim 1, characterized in that, The acquisition of gene expression data for different cell types includes: Single-cell expression data of different cell types are obtained. The single-cell expression data includes source cell data and target cell data. The single-cell expression data describes the expression level of genes in cells using the form of a single-cell transcriptome expression matrix. The single-cell expression data is preprocessed to obtain gene expression data that meets the set format requirements. The preprocessing includes at least one of format conversion, quality control, and standardization.

3. The method according to claim 1, characterized in that, Before constructing different hypothetical models using a set of covariate genes for candidate genes in the target cell, and determining the causal relationship between source cell genes and target cell genes through the different hypothetical models, the method further includes: For the candidate target genes of the target cells, calculate the correlation coefficient between the candidate target genes and other genes of the target cells; Based on the correlation coefficient, a set of covariate genes that are significantly correlated with the candidate target gene are screened in the target cells.

4. The method according to claim 1, characterized in that, For the candidate genes of the target cell, different hypothetical models are constructed using a covariate gene set to determine the causal relationship between source cell genes and target cell genes through the different hypothetical models, including: Using candidate genes of the target cell as the genetic variables of the target cell, different sample datasets are constructed using the covariate gene set to train the model on the genetic variables of the target cell, thereby obtaining the first hypothesis model and the second hypothesis model. During the training of the first hypothetical model and the second hypothetical model, the first prediction error obtained by the first hypothetical model after multiple iterations and the second prediction error obtained by the second hypothetical model after multiple iterations are obtained respectively. The first prediction error is used to represent the prediction error of the target cell gene under regression modeling without source cell genes, and the second prediction error is used to represent the prediction error of the target cell gene under regression modeling with source cell genes. Based on the first prediction error and the second prediction error, the causal relationship between the source cell gene and the target cell gene is determined.

5. The method according to claim 4, characterized in that, The process of using candidate genes of the target cell as gene variables of the target cell, constructing different sample datasets using the covariate gene set, and training the model on the gene variables of the target cell to obtain a first hypothesis model and a second hypothesis model includes: Using the candidate genes of the target cells as the gene variables of the target cells, and using the covariate gene set as the sample dataset, regression fitting is performed on the gene variables of the target cells to obtain the first hypothesis model; Using candidate genes of the target cell as the gene variables of the target cell, and introducing gene variables of the source cell as a sample dataset based on the covariate gene set, regression modeling of the gene variables of the target cell is performed to obtain the second hypothesis model.

6. The method according to claim 4, characterized in that, The step of determining the causal relationship between source cell genes and target cell genes based on the first prediction error and the second prediction error includes: Based on the first prediction error and the second prediction error, calculate the causal strength score of the source cell gene to the target cell gene. If the causal strength score is greater than the set value, it is determined that there is a causal relationship between the source cell gene and the target cell gene. If the causal strength score is less than or equal to the set value, it is determined that there is no causal relationship between the source cell gene and the target cell gene.

7. The method according to any one of claims 1-6, characterized in that, The step of calculating the communication strength between the source cell and the target cell based on the causal relationship between the source cell genes and the target cell genes, in order to identify the intercellular communication relationship through the communication strength, includes: Based on the causal relationship between the source cell gene and the target cell gene, a gene regulation network across cell types is constructed. The gene regulation network uses an adjacency matrix to describe the causal relationship between the source cell gene and the target cell gene. If there is a causal relationship between the source cell gene and the target cell gene, the corresponding element in the adjacency matrix is ​​assigned a first value. If there is no causal relationship between the source cell gene and the target cell gene, the corresponding element in the adjacency matrix is ​​assigned a second value. The communication strength between the source cell and the target cell is obtained by weighted summation of all elements in the adjacency matrix, and the communication relationship between cells is identified by the communication strength.

8. A device for identifying intercellular communication relationships, characterized in that, include: An acquisition unit is used to acquire gene expression data of different cell types, wherein the gene expression data is the expression level of each gene in the cell's own genome, and the different cell types include at least source cells and target cells; The construction unit is used to construct different hypothetical models for candidate genes in the target cell using a set of covariate genes, so as to determine the causal relationship between source cell genes and target cell genes through the different hypothetical models. The set of covariate genes is the set of genes in the target cell that are significantly related to the candidate genes. The different hypothetical models are based on whether the source cell genes have a regulatory effect on the target cell genes, including a hypothetical model of no regulation and a hypothetical model of direct regulation. The hypothetical model of no regulation assumes that the source cell genes do not have a regulatory effect on the target cell genes, and the hypothetical model of direct regulation assumes that the source cell genes have direct regulation on the target cell genes. The identification unit is used to calculate the communication strength between the source cell and the target cell based on the causal relationship between the source cell genes and the target cell genes, so as to identify the communication relationship between cells through the communication strength.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cell communication analysis method and system

    CN112466403A

  • Intercellular signal transduction inference method and device based on optimal transmission of knowledge graph, equipment and medium

    CN119339796A