Method, device and equipment for identifying communication relationship between cells and storage medium

By acquiring cell gene expression data and constructing a hypothesis model, we can identify the causal relationship between cells, solve the problem of strong dependence on prior databases, and achieve causal explanation and quantification of intercellular communication relationships.

CN120808884AActive Publication Date: 2025-10-17HANGZHOU INST FOR ADVANCED STUDY UCAS

Patent Information

Application Number
CN202511254422.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-17
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

The prior database in the existing technology does not have a rigorous communication expression for the co-occurrence of receptors and ligands, and it is difficult to accurately explain the actual regulation or signal transduction process of ligands and receptors, resulting in a lack of causal explanatory power in the identification of intercellular communication relationships.

Method used

By obtaining gene expression data of different cell types and using covariate gene sets to construct different hypothesis models, the causal relationship between source cell genes and target cell genes is determined, and the communication intensity is calculated to identify the communication relationship between cells and eliminate the interference of the target cell expression background.

Benefits of technology

It achieves the causal-level explanatory power of intercellular communication relationships, accurately evaluates the overall regulatory effects between cell populations, and provides quantification of the direct causal contribution of intercellular communication relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808884A_ABST
    Figure CN120808884A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a device and equipment for identifying a communication relationship between cells, and a storage medium, and relates to the technical field of bioinformatics, direct causal contribution between the cells is quantified by using a causal association relationship of a source cell gene to a target cell gene, the communication relationship between the cells can have explanatory force at a causal level, and the identification accuracy of the communication relationship between the cells is improved. Further, the overall regulation effect among the cell populations is accurately evaluated. The method comprises the following steps: acquiring gene expression data of different cell types, wherein the different cell types at least comprise source cells and target cells; aiming at the candidate genes of the target cell, constructing different hypothesis models by using the covariable gene set, and determining a causal association relationship between the source cell gene and the target cell gene through the different hypothesis models; and calculating the communication intensity of the source cell to the target cell according to the causal association relationship of the source cell gene to the target cell gene, so as to identify the causal regulation relationship between the cells through the communication intensity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and in particular to a method and device for identifying cell communication relationship, equipment and storage medium. BACKGROUND

[0002] Cell communication is a basic process for maintaining homeostasis, development and immune regulation of multicellular organisms. With the development of single-cell RNA sequencing technology, researchers can analyze the interaction between cells at the single-cell level.

[0003] In order to mine the potential communication mechanism between cells, the related technology can identify the communication relationship between cells through a cell communication analysis tool. Specifically, in the process of identifying the cell communication relationship, the cell communication analysis tool analyzes the expression pattern of ligand and receptor pairs in the single-cell transcription array, and then infers the communication relationship between cells according to the prior database formed by the receptor and ligand pairs. By comparing the co-expression relationship of ligands and receptors between different cell types, the potential communication pairs between cells are predicted. However, this inference method is highly dependent on the prior database, and the co-occurrence of receptors and ligands in the prior database does not have a rigorous communication expression. It is difficult to accurately explain the actual regulatory or signal transduction process of ligands and receptors, so the process of identifying the communication relationship between cells lacks causal explanation. SUMMARY

[0004] Therefore, the present application provides a method and device for identifying cell communication relationship, equipment and storage medium, which mainly aims to solve the problem that the co-occurrence of receptors and ligands in the prior database does not have a rigorous communication expression, and it is difficult to accurately explain the actual regulatory or signal transduction process of ligands and receptors, so the process of identifying the communication relationship between cells lacks causal explanation.

[0005] In a first aspect, a method for identifying cell communication relationship is provided, the method comprising: obtaining gene expression data of different cell types, the gene expression data being the expression amount of each gene in the cell's own genome, and the different cell types including at least a source cell and a target cell; constructing different hypothesis models using a set of covariate genes for candidate genes of the target cell, to determine the causal relationship between the source cell genes and the target cell genes through the different hypothesis models, the set of covariate genes being a set of genes in the target cell that are significantly related to the candidate genes; calculating the communication strength of the source cell to the target cell according to the causal relationship between the source cell genes and the target cell genes, to identify the communication relationship between cells through the communication strength.

[0006] Further, the gene expression data of different cell types is obtained, including: Obtaining single-cell expression data of different cell types, the single-cell expression data including source cell data and target cell data, the single-cell expression data being in the form of a single-cell transcriptome expression matrix to describe the expression amount of genes in cells; Preprocessing the single-cell expression data to obtain gene expression data meeting the set format requirements, the preprocessing including at least one of format conversion, quality control and standardization.

[0007] Further, before the candidate genes for the target cells are used to construct different hypothesis models using a set of covariate genes to determine the causal correlation between source cell genes and target cell genes, the method further comprises: For the candidate target genes of the target cells, calculating the correlation coefficients between the candidate target genes and other genes of the target cells; According to the correlation coefficients, screening a set of covariate genes significantly correlated with the candidate target genes in the target cells.

[0008] Further, the candidate genes for the target cells are used to construct different hypothesis models using a set of covariate genes to determine the causal correlation between source cell genes and target cell genes, including: Using the candidate genes of the target cells as gene variables of the target cells, using the set of covariate genes to construct different sample data sets for model training of the gene variables of the target cells to obtain a first hypothesis model and a second hypothesis model; In the process of training the first hypothesis model and the second hypothesis model, a first prediction error of the first hypothesis model obtained through multiple iterations and a second prediction error of the second hypothesis model obtained through multiple iterations are respectively obtained, the first prediction error being used to represent the prediction error of the target cell genes under regression modeling without including source cell genes, and the second prediction error being used to represent the prediction error of the target cell genes under regression modeling including source cell genes; According to the first prediction error and the second prediction error, the causal correlation between source cell genes and target cell genes is determined.

[0009] Further, using the candidate genes of the target cells as gene variables of the target cells, using the set of covariate genes to construct different sample data sets for model training of the gene variables of the target cells to obtain a first hypothesis model and a second hypothesis model, including: using the candidate gene of the target cell as a gene variable of the target cell, performing regression fitting on the gene variable of the target cell using the covariate gene set as sample data set, to obtain a first hypothesis model; using the candidate gene of the target cell as a gene variable of the target cell, introducing the gene variable of the source cell as a sample data set on the basis of the covariate gene set, to perform regression modeling on the gene variable of the target cell, to obtain a second hypothesis model.

[0010] Further, the determining of the causal correlation between the source cell gene and the target cell gene according to the first prediction error and the second prediction error comprises: calculating a causal strength score of the source cell gene to the target cell gene according to the first prediction error and the second prediction error; if the causal strength score is greater than a set value, it is determined that the source cell gene has a causal effect on the target cell gene; if the causal strength score is less than or equal to the set value, it is determined that the source cell gene has no causal effect on the target cell gene.

[0011] Further, the calculating of the communication strength of the source cell to the target cell according to the causal correlation between the source cell gene and the target cell gene, to identify the communication relationship between cells through the communication strength, comprises: According to the causal correlation between the source cell gene and the target cell gene, a cross-cell type gene regulation network is constructed, the gene regulation network uses an adjacency matrix to describe the causal correlation between the source cell gene and the target cell gene, if the source cell gene has a causal effect on the target cell gene, the corresponding element in the adjacency matrix is assigned a first value, if the source cell gene has no causal effect on the target cell gene, the corresponding element in the adjacency matrix is assigned a second value; performing weighted summation calculation on all elements in the adjacency matrix to obtain the communication strength of the source cell to the target cell, to identify the communication relationship between cells through the communication strength.

[0012] In a second aspect, a device for identifying a communication relationship between cells is provided, the device comprising: an acquisition unit configured to acquire gene expression data of different cell types, the gene expression data being the expression amount of each gene in the cell's own genome, the different cell types at least including a source cell and a target cell; constructing units, configured to construct different hypothesis models using a covariate gene set for a candidate gene of the target cell, to determine a causal relationship of a source cell gene to a target cell gene through the different hypothesis models, the covariate gene set being a gene set significantly related to the candidate gene in the target cell; identifying units, configured to calculate a communication strength of a source cell to a target cell according to the causal relationship of the source cell gene to the target cell gene, to identify a communication relationship between cells through the communication strength.

[0013] Further, the obtaining units are specifically configured to: obtain single-cell expression data of different cell types, the single-cell expression data including source cell data and target cell data, the single-cell expression data being in a form of a single-cell transcriptome expression matrix to describe expression amounts of genes in cells; perform preprocessing on the single-cell expression data to obtain gene expression data meeting set format requirements, the preprocessing including at least one of format conversion, quality control, and standardization.

[0014] Further, the apparatus further includes: computing units, configured to, before the constructing units construct different hypothesis models using a covariate gene set for a candidate gene of the target cell, to determine a causal relationship of a source cell gene to a target cell gene through the different hypothesis models, calculate correlation coefficients between the candidate target gene and other genes of the target cell for the candidate target gene of the target cell; screening units, configured to screen a covariate gene set significantly related to the candidate target gene in the target cell according to the correlation coefficients.

[0015] Further, the constructing units include: constructing modules, configured to use the covariate gene set to construct different sample data sets for model training on gene variables of the target cell, with the candidate gene of the target cell as the gene variable of the target cell, to obtain a first hypothesis model and a second hypothesis model; obtaining modules, configured to, in a process of training the first hypothesis model and the second hypothesis model, respectively obtain a first prediction error of the first hypothesis model obtained through multiple iterations and a second prediction error of the second hypothesis model obtained through multiple iterations, the first prediction error being used to represent a prediction error of a target cell gene under regression modeling without a source cell gene, and the second prediction error being used to represent a prediction error of the target cell gene under regression modeling with the source cell gene; determining a causal correlation between the source cell gene and the target cell gene according to the first prediction error and the second prediction error.

[0016] Further, the constructing module is specifically configured to: regression fitting is performed on the gene variable of the target cell using the candidate gene of the target cell as the gene variable of the target cell and using the covariate gene set as the sample data set, to obtain a first hypothesis model; regression modeling is performed on the gene variable of the target cell using the candidate gene of the target cell as the gene variable of the target cell and introducing the gene variable of the source cell as the sample data set on the basis of the covariate gene set, to obtain a second hypothesis model.

[0017] Further, the determining module is specifically configured to: a causal strength score of the source cell gene to the target cell gene is calculated according to the first prediction error and the second prediction error; if the causal strength score is greater than a set value, it is determined that the source cell gene has a causal effect on the target cell gene; if the causal strength score is less than or equal to the set value, it is determined that the source cell gene does not have a causal effect on the target cell gene.

[0018] Further, the identifying unit is specifically configured to: a gene regulatory network across cell types is constructed according to the causal correlation between the source cell gene and the target cell gene, the gene regulatory network using an adjacency matrix to describe the causal correlation between the source cell gene and the target cell gene, if the source cell gene has a causal effect on the target cell gene, a corresponding element in the adjacency matrix is assigned a first value, and if the source cell gene does not have a causal effect on the target cell gene, a corresponding element in the adjacency matrix is assigned a second value; a weighted sum of all elements in the adjacency matrix is calculated to obtain a communication strength of the source cell to the target cell, so as to identify the communication relationship between cells through the communication strength.

[0019] In a third aspect, a storage medium having a computer program stored thereon is provided, the program being executed by a processor to implement the above-mentioned method for identifying the communication relationship between cells.

[0020] In a fourth aspect, a device for identifying the communication relationship between cells is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, the processor implementing the above-mentioned method for identifying the communication relationship between cells when executing the program.

[0021] By means of the technical scheme, the method, device, equipment and storage medium for identifying the intercellular communication relationship provided by the application can compare with the prior art of identifying the intercellular communication relationship by using a prior database. The application obtains gene expression data of different cell types. The gene expression data is the expression amount of each gene in the genome of the cell itself. The different cell types at least include a source cell and a target cell. For a candidate gene of the target cell, a different hypothesis model is constructed by using a covariate gene set to determine the causal correlation relationship between the source cell gene and the target cell gene by using the different hypothesis model. The covariate gene set is a gene set in the target cell that is significantly related to the candidate gene. According to the causal correlation relationship between the source cell gene and the target cell gene, the communication strength of the source cell to the target cell is calculated to identify the causal regulation relationship between the cells by using the communication strength. The whole process excludes the regulation interference of the source cell to the target cell by controlling the covariate gene set expressed in the background of the target cell, so that the causal correlation relationship between the source cell gene and the target cell gene is accurately obtained. The causal correlation relationship between the source cell gene and the target cell gene is used to quantify the direct causal contribution between the cells, so that the intercellular communication relationship has the explanatory power at the causal level, and the overall regulation effect between the cell populations can be accurately evaluated.

[0022] The above description is only a summary of the technical scheme of the application. In order to more clearly understand the technical means of the application, the specific embodiments of the application can be implemented according to the content of the description, and in order to make the above and other purposes, characteristics and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS

[0023] The drawings described herein are used to provide further understanding of the application, and form a part of the application. The schematic embodiments of the application and the description thereof are used to explain the application, and do not constitute an improper limitation on the application. In the drawings: Figure 1 is a flowchart of the method for identifying the intercellular communication relationship in an embodiment of the application; Figure 2 is a flowchart of the method for identifying the intercellular communication relationship in an embodiment of the application; Figure 1 is a flowchart of a specific embodiment of step 101 in the method; Figure 3 is a flowchart of the method for identifying the intercellular communication relationship in another embodiment of the application; Figure 4 is a flowchart of the method for identifying the intercellular communication relationship in another embodiment of the application; Figure 1 is a flowchart of a specific embodiment of step 102 in the method; Figure 5 is a flowchart of a specific embodiment of step 103 in the method; Figure 1 is a flowchart of a specific embodiment of step 103 in the method; Figure 6Another embodiment of the application is a flowchart of a method for identifying intercellular communication relationships. Figure 7 is a structural schematic diagram of an apparatus for identifying intercellular communication relationships in an embodiment of the application. Figure 8 is a structural schematic diagram of an apparatus of a computer device provided in an embodiment of the application. DETAILED DESCRIPTION

[0024] The application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that the embodiments in the application and the features in the embodiments can be combined with each other without conflict.

[0025] The related art can identify intercellular communication relationships through a cell communication analysis tool. Specifically, in the process of identifying intercellular communication relationships, the cell communication analysis tool parses expression patterns of ligand-receptor pairs in single-cell transcriptional arrays, and then infers intercellular communication relationships according to a prior database formed by receptors and ligand pairs. By comparing the co-expression relationships of ligands and receptors between different cell types, potential communication pairs between cells are predicted. However, this inference method has strong dependence on the prior database, and the co-occurrence of receptors and ligands in the prior database does not have strict communication expression. It is difficult to accurately explain the actual existence of the regulation or signal transduction process of ligands and receptors, so that the process of identifying intercellular communication relationships lacks causal explanation.

[0026] To solve this problem, the present embodiment provides a method for identifying intercellular communication relationships, as shown in Figure 1 includes the following steps: 101. Obtain gene expression data of different cell types.

[0027] In the present embodiment, the gene expression data is the expression amount of each gene in the cell's own genome, and the different cell types at least include source cells and target cells. Generally, the source cells and the target cells are cell types having a conversion relationship or an alignment relationship.

[0028] As one way of obtaining, the different cell types can be separated by cell separation technology, and the intermediate product of gene expression can be extracted in the separated cells. The gene expression information is analyzed according to the intermediate product of gene expression, and the gene expression data of different cell types is obtained. Here, the source cells and the target cells need to be detected by experiments respectively, and the gene expression data of the source cells and the gene expression data of the target cells are obtained correspondingly.

[0029] As another acquisition manner, the gene expression data can be acquired by querying a public database. Since a large amount of gene expression data from different species and different cell types is collected in the public database, the gene expression data of different cell types can be searched in the public database according to the cell types. Here, the gene expression data of the source cell and the gene expression data of the target cell are searched in the public database respectively, and are obtained correspondingly.

[0030] Further, in order to eliminate noise and pollution, it is beneficial to subsequent cell analysis process. After the gene expression data is acquired, quality control and standardization processing can be performed on the gene expression data to obtain structured gene expression data. Generally, the gene expression data is in a matrix form. After the quality control and standardization processing, the gene and sample dimensions in the gene expression matrix of the source cell and the target cell are consistent.

[0031] 102. For the candidate gene of the target cell, different hypothesis models are constructed using the covariate gene set to determine the causal correlation between the source cell gene and the target cell gene by the different hypothesis models.

[0032] In this embodiment, the covariate gene set is a gene set in the target cell that is significantly related to the candidate gene. Generally, the gene expression data is mainly used to reflect the expression of genes in the cell. For any gene in the target cell, the covariate gene set that is co-variable in the expression pattern can be obtained by analyzing the significant correlation between genes. Here, a suitable correlation index can be selected to quantify the correlation degree of the expression of two genes, and the covariate gene set is screened in the target cell according to the correlation degree of the expression of two genes.

[0033] It can be understood that after the covariate gene set is acquired, in order to quantify the regulation of the source cell gene on the target cell gene, different hypothesis models can be constructed on the basis of controlling the internal interference of the target cell. Here, the different hypothesis models include but are not limited to a hypothesis model without regulation and a hypothesis model with direct regulation, which assumes whether the source cell gene regulates the target cell gene. For the hypothesis model without regulation, it is assumed that the source cell gene does not regulate the target cell gene, that is, the gene expression variable of the target cell is only determined by the covariate gene set in the target gene. For the hypothesis model with direct regulation, it is assumed that the source cell gene regulates the target cell gene directly. The direct regulation can be indirectly realized by relying on the covariate gene set in the target cell, or can not rely on the covariate gene set, that is, the gene expression variable of the target cell is introduced into the source cell gene on the basis of the covariate gene set in the target gene.

[0034] In the embodiment, the difference between different hypothesis models is whether the gene expression variable of the target cell includes the source cell gene, which is mainly used to capture the direct regulation signal independent of the covariant gene set in the target cell. That is, by comparing the fitting errors of different hypothesis models, it can be verified whether the cross-cell signal of the source cell gene is necessary to explain the variation of the target cell gene on the basis of controlling the interference in the target cell, that is, to quantify and infer the causal relationship between the source cell gene and the target cell gene.

[0035] Specifically, if part of the variation of the target cell gene cannot be explained by the covariant gene set in the target cell, at this time, the direct regulation hypothesis model can capture this part of the intercellular signal due to the inclusion of the source cell gene, and the explanation ability of the target cell gene variable is better than that of the hypothesis model without regulation, which is manifested as a small fitting error, which indicates that the source cell has a causal relationship with the target cell. On the contrary, if all the variations of the target cell can be explained by the covariant gene set in the target cell, the direct effect of the source cell gene is redundant, at this time, the fitting error of the direct regulation hypothesis model and the hypothesis model without regulation is close, and even because the direct regulation hypothesis model introduces redundant parameters to generate fitting noise, the fitting effect of the hypothesis model without regulation is better, which indicates that the source cell does not have a causal relationship with the target cell.

[0036] 103、According to the causal relationship between the source cell gene and the target cell gene, the communication strength of the source cell to the target cell is calculated to identify the communication relationship between the cells through the communication strength.

[0037] It can be understood that the causal regulation strength of a single gene pair cannot reflect the communication strength of the source cell to the target cell, and the communication strength needs to comprehensively reflect the overall causal regulation ability of the source cell to the target cell. That is, the communication strength of the source cell to the target cell is comprehensively considered the causal regulation relationship of the source cell and the target cell formed by the gene pair and the contribution weight of the gene pair.

[0038] In the embodiment, the source cell gene to the target cell gene can be quantified as a causal pair, and different values can be given to each gene pair according to the causal relationship. Correspondingly, the communication strength of the source cell to the target cell can be obtained by weighted integration or weighted average of multiple causal pair values. For example, the gene set in the source cell is S, the gene set in the target cell is T, and for each significant causal pair , the causal strength is , and correspondingly, the communication strength C of the source cell to the target cell can be defined as: .

[0039] In actual application scenarios, if the importance of the target gene in cell function is different, for example, the regulatory significance of key transcription factors, signal molecules and ordinary genes is different, a gene weight can be further introduced, at this time, the communication strength of the source cell to the target cell can be obtained by weighted integration or weighted average of multiple causal pairs based on the gene weight. Continue to illustrate with the example in the above text, the weight of the target gene is , wherein is the proportion of the expression amount of the target gene in the total expression of the target cell, or the core degree in the pathway. Correspondingly, the communication strength C of the source cell to the target cell can be modified as: .

[0040] It can be understood that the communication strength of the source cell to the target cell is an important quantitative index for judging the communication relationship between cells. Generally speaking, the greater the communication strength of the source cell to the target cell, the closer the communication relationship between the source cell and the target cell, and vice versa, the smaller the communication strength of the source cell to the target cell, the weaker the communication relationship between the source cell and the target cell.

[0041] The method for identifying the communication relationship between cells provided by the embodiments of the present application can be compared with the current method for identifying the communication relationship between cells by using prior database. The present application obtains gene expression data of different cell types, the gene expression data being the expression amount of each gene in the genome of the cell itself, and different cell types at least including a source cell and a target cell. For candidate genes of the target cell, different hypothesis models are constructed using a set of covariate genes to determine the causal correlation relationship between the source cell genes and the target cell genes through different hypothesis models, and the set of covariate genes being a set of genes in the target cell that are significantly related to the candidate genes. According to the causal correlation relationship between the source cell genes and the target cell genes, the communication strength of the source cell to the target cell is calculated to identify the causal regulation relationship between cells through the communication strength. The whole process excludes the regulatory interference of the source cell to the target cell by controlling the set of covariate genes expressed in the background of the target cell, so as to accurately obtain the causal correlation relationship between the source cell genes and the target cell genes, and use the causal correlation relationship between the source cell genes and the target cell genes to quantify the direct causal contribution between cells, so that the communication relationship between cells has causal explanation, and the overall regulation effect between cell populations can be accurately evaluated.

[0042] In actual application scenarios, in order to remove low-quality cells and genes and reduce technical bias, when obtaining different types of gene expression data, different cell expression data need to be preprocessed. Specifically, as shown in Figure 2 , step 101 includes the following steps: 201、obtaining single-cell expression data of different cell types.

[0043] 202、preprocessing the single-cell expression data to obtain gene expression data meeting the set format requirements.

[0044] In the single-cell expression data, the source cell data and the target cell data are in the form of a single-cell transcriptome expression matrix, which describes the expression of genes in cells. That is, the basic format of the expression matrix is consistent for both source cell data and target cell data, and the gene expression can be quantified by forming a two-dimensional table of genes and cells. In the expression matrix, the rows represent genes, the columns represent cells, and the elements represent the expression of genes in the corresponding cells.

[0045] It can be understood that in the scenario of intercellular communication analysis, the source cell and the target cell are functionally divided according to the direction of signal transmission, and the difference in their expression matrices is reflected in the cell type rather than the matrix structure. That is, the format of the source cell and the target cell in the expression matrix is consistent, and they follow the same format, the definition of rows and columns is consistent, and the elements in the matrix are all gene expression data. The difference is the expression characteristics of the genes in the cells.

[0046] In order to convert the original cell data into high-quality, standardized and format-compliant gene expression matrix, the single-cell expression data can be preprocessed to obtain gene expression data meeting the set format requirements. In this embodiment, the preprocessing includes at least one of format conversion, quality control and standardization. Generally, the original expression matrix of different cell types may have differences, and the expression matrix of different cell types can be converted to the same format through format conversion. Since there may be low-quality cells or low-expression genes in the original expression matrix, in order to filter low-quality cells and genes, comprehensive quality judgment can be made on different types of cell data through quality control, and cells and / or genes with low comprehensive quality are filtered. Further, the sequencing depth of different cells is different, in order to reduce technical noise, the expression between different cells can be made comparable through standardization. Common standardization methods include, but are not limited to, total expression normalization, logarithmic conversion and scaling, etc. The specific method can be selected according to the data type.

[0047] In the actual application scenario of intercellular communication analysis, the covariate gene set is the key to control interference and improve analysis reliability. In order to accurately exclude the interference in the target biological signal and ensure that the analysis result can accurately reflect the causal relationship between cells, further, as shown in Figure 3 Before step 102, the method further includes the following steps: 301、For the candidate target gene of the target cell, the correlation coefficient between the candidate target gene and other genes of the target cell is calculated.

[0048] 302、According to the correlation coefficient, a set of covariant genes significantly related to the candidate target gene is screened in the target cell.

[0049] In this embodiment, the candidate target gene is any gene in the target cell. The consistency of the expression change trend of the candidate target gene and other genes can be quantified by the correlation coefficient, so as to determine whether there is a co-regulation or functional association between the genes. For positively correlated genes A and B, when gene A expression increases, gene B also increases. For negatively correlated genes A and B, when gene A expression increases, gene B decreases. For genes A and B with no significant correlation, the expression change of genes A and B has no obvious correlation, and the functional association is low. Specifically, for the candidate target gene of the target cell, the expression vector of the candidate target gene is extracted in the matrix expression of the target cell, and then the expression vectors of other genes are extracted in turn. According to the expression vector of the candidate target gene and the expression vectors of other genes, the correlation coefficient between the candidate target gene and other genes is calculated. This process can use the following formula:

[0050] wherein, is the covariance, , is the expression vector of the two genes.

[0051] It can be understood that the correlation coefficient is a statistical quantity quantifying the correlation degree of the expression pattern of the candidate target gene and other genes. The closer the absolute value of the correlation coefficient is to 1, the more synchronous the expression change of the two genes is. The closer the absolute value of the correlation coefficient is to 0, the weaker the correlation is. Here, a coefficient threshold can be set. Correspondingly, for the genes with a correlation coefficient greater than the coefficient threshold, the genes can be added to the covariant gene set as genes significantly related to the candidate target gene.

[0052] It should be noted that if the candidate target gene is a set, the correlation between each gene in the set and all other genes can be calculated, and then the results can be integrated by taking the mean value, voting method, etc. to screen a gene set related to the entire candidate target gene set, and obtain the covariant gene set.

[0053] In actual application scenarios, in order to cope with the multidimensional complexity of cell-to-cell communication, the source cell and the target cell are often disturbed by multiple non-communication direct related factors. These factors can be captured by the covariant gene set. At this time, for multiple potential patterns existing between cells, different hypothetical models can be constructed. Specifically, for example, Figure 4As shown, step 102 includes the following steps: 401. Using the candidate genes of the target cell as the gene variable of the target cell, the covariate gene set is used to construct different sample data sets for model training of the gene variable of the target cell, to obtain a first hypothesis model and a second hypothesis model.

[0054] 402. In the process of training the first hypothesis model and the second hypothesis model, a first prediction error obtained by the first hypothesis model after multiple iterations and a second prediction error obtained by the second hypothesis model after multiple iterations are respectively acquired.

[0055] 403. According to the first prediction error and the second prediction error, the causal relationship between the source cell gene and the target cell gene is determined.

[0056] In this embodiment, the first prediction error is used to represent the prediction error of the target cell gene under regression modeling without including the source cell gene, at this time, only the covariate gene set is applied to the model as the sample data set. Specifically, using the candidate genes of the target cell as the gene variable of the target cell, the covariate gene set is used as the sample data set for regression fitting of the gene variable of the target cell, to obtain a first hypothesis model.

[0057] The above first hypothesis model can be represented by the following formula:

[0058] wherein, is the gene variable of the target cell, is the covariate gene set, is the prediction error of the first hypothesis model.

[0059] In this embodiment, the second prediction error is used to represent the prediction error of the target cell gene under regression modeling including the source cell gene, at this time, in addition to the covariate gene set, the source cell gene is also added together as the sample data set applied to the model. Specifically, using the candidate genes of the target cell as the gene variable of the target cell, introducing the gene variable of the source cell on the basis of the covariate gene set as the sample data set together for regression modeling of the gene variable of the target cell, to obtain a second hypothesis model.

[0060] The above second hypothesis model can be represented by the following formula:

[0061] wherein, is the gene variable of the target cell, is the covariate gene set, is the gene variable of the source cell, the prediction error of the second hypothetical model.

[0062] Specifically in the process of model training, the sample data set can be divided into a training set and a test set, the training set is used for regression, and the test set is used for error judgment. The division method can use K-fold cross validation, which effectively avoids the randomness bias of single division, making the average test error of the two models more reliable. In order to ensure the comparability of the results, the two hypothetical models need to be compared under the same division method, the same number of iterations, and the same error index.

[0063] It can be understood that the first prediction error and the second prediction error usually correspond to the prediction performance of two different hypothetical models. By comparing the error difference between the two, it can be inferred whether the source cell gene has a causal effect on the target cell gene. If the source cell gene has a causal relationship with the target cell gene, the second hypothetical model with the source cell gene should be able to more accurately predict the state of the target cell gene, and its prediction error should be significantly lower than the first hypothetical model without the source cell gene. On the contrary, if the source cell gene does not have a causal relationship with the target cell gene, the prediction ability of the second hypothetical model with the source cell gene does not improve compared to the first prediction model without the source cell gene. Specifically, in the process of determining the causal relationship between the source cell gene and the target cell gene according to the first prediction error and the second prediction error, the causal strength score of the source cell gene to the target cell gene can be calculated according to the first prediction error and the second prediction error; if the causal strength score is greater than a set value, it is determined that the source cell gene has a causal effect on the target cell gene; if the causal strength score is less than or equal to the set value, it is determined that the source cell gene does not have a causal effect on the target cell gene.

[0064] The above causal strength score can be represented by the following formula:

[0065] wherein, is the prediction error of the target cell gene under the regression model without the source cell gene , is the prediction error of the target cell gene under the regression model with the source cell gene .

[0066] Correspondingly, if the prediction error is significantly reduced in the regression model with the source cell gene , that is, , it is considered that the source cell gene has a causal effect on the target cell gene ; on the contrary, it is considered that the source cell gene Target cell genes There is no causal relationship.

[0067] In actual application scenarios, intercellular communication usually relies on the interaction between cells. For cells with causal correlation, communication can be achieved. Therefore, the causal correlation of the source cell gene to the target cell gene can be used to reflect the communication ability between cells. Specifically, as shown in Figure 5 Step 103 includes the following steps: 501. Construct a gene regulation network across cell types according to the causal correlation of the source cell gene to the target cell gene.

[0068] 502. Perform weighted summation calculation on all elements in the adjacency matrix to obtain the communication strength of the source cell to the target cell, so as to identify the communication relationship between cells through the communication strength.

[0069] In this embodiment, the gene regulation network uses an adjacency matrix to describe the causal correlation of the source cell gene to the target cell gene. If the source cell gene has a causal relationship with the target cell gene, the corresponding element in the adjacency matrix is assigned a first value. If the source cell gene has no causal relationship with the target cell gene, the corresponding element in the adjacency matrix is assigned a second value. The assignment of elements in the adjacency matrix can be referred to the following formula:

[0070] Further, in order to quantitatively evaluate the causal regulation ability of the source cell to the target cell, all elements in the critical matrix can be weighted and summed. Specifically, the existence of the causal relationship of the source cell gene to the target cell gene can be regarded as a directed edge, and the corresponding weighted summation process is to count all existing directed edges in the adjacency matrix. The result obtained by counting is the communication strength of the source cell to the target cell. The following formula can be used:

[0071] Wherein, is the adjacency matrix representation of the gene regulation network, The larger the corresponding value is, the stronger the causal regulation ability of the source cell to the target cell is, and the stronger the communication relationship between cells is.

[0072] In actual application scenarios, the identification process of the intercellular communication relationship can be referred to Figure 6 After obtaining the gene expression data of the target cell and the source cell, the candidate target gene in the target cell and other genes in the target cell are correlation calculation is performed to screen a set of covariate genes that are significantly related to the candidate target gene. Then, cross-validation predictability is performed according to the set of covariate genes , respectively construct different hypothetical models, obtain the causal correlation relationship of the source cell gene to the target cell gene through the prediction error of the different hypothetical models, construct a cross-cell type gene regulation network according to the causal correlation relationship, the gene regulation network is described using an adjacency matrix, if the source cell gene has a causal relationship with the target cell gene, the corresponding element in the adjacency matrix is assigned a value of 1, if the source cell gene has no causal relationship with the target cell gene, the corresponding element in the adjacency matrix is assigned a value of 0, and further weighted sum of all elements in the adjacency matrix is obtained , and correspondingly, the communication relationship between cells can be identified through the communication strength of the source cell to the target cell.

[0073] Further, as a specific implementation of the above method, the embodiment of the present application provides an intercellular communication relationship identification device, as shown in Figure 7 , the device comprises an acquisition unit 61, a construction unit 62 and an identification unit 63.

[0074] The acquisition unit 61 is configured to acquire gene expression data of different cell types, the gene expression data being the expression amount of each gene in the cell's own genome, and the different cell types at least including a source cell and a target cell; The construction unit 62 is configured to use a covariate gene set to construct different hypothetical models for the candidate gene of the target cell, so as to determine the causal correlation relationship of the source cell gene to the target cell gene through the different hypothetical models, and the covariate gene set being a gene set in the target cell which is significantly related to the candidate gene; The identification unit 63 is configured to calculate the communication strength of the source cell to the target cell according to the causal correlation relationship of the source cell gene to the target cell gene, so as to identify the communication relationship between cells through the communication strength.

[0075] Compared with the prior art of identifying the intercellular communication relationship by using a prior database, the device for identifying the intercellular communication relationship provided by the embodiment of the application acquires gene expression data of different cell types, the gene expression data being the expression amount of each gene in the genome of the cell itself, and the different cell types at least including a source cell and a target cell; for a candidate gene of the target cell, different hypothesis models are constructed by using a covariate gene set to determine the causal correlation relationship of the source cell gene to the target cell gene through the different hypothesis models, the covariate gene set being a gene set that is significantly related to the candidate gene in the target cell; and the communication strength of the source cell to the target cell is calculated according to the causal correlation relationship of the source cell gene to the target cell gene to identify the causal regulation relationship between the cells through the communication strength. The whole process excludes the regulation interference of the source cell to the target cell by controlling the covariate gene set expressed in the background of the target cell, so that the causal correlation relationship of the source cell gene to the target cell gene is accurately acquired, the direct causal contribution between the cells is quantified by using the causal correlation relationship of the source cell gene to the target cell gene, the intercellular communication relationship can be explained in the causal level, and the overall regulation effect between the cell populations can be accurately evaluated.

[0076] In a specific application scenario, the acquisition unit is specifically configured to: acquire single-cell expression data of different cell types, the single-cell expression data including source cell data and target cell data, and the single-cell expression data being in the form of a single-cell transcriptome expression matrix to describe the expression amount of a gene in a cell; perform preprocessing on the single-cell expression data to obtain gene expression data meeting a set format requirement, the preprocessing including at least one of format conversion, quality control, and standardization.

[0077] In a specific application scenario, the device further includes: a calculation unit configured to, before the different hypothesis models are constructed by using the covariate gene set for the candidate gene of the target cell to determine the causal correlation relationship of the source cell gene to the target cell gene, calculate a correlation coefficient between the candidate target gene and other genes of the target cell for the candidate target gene of the target cell. a screening unit configured to screen a covariate gene set that is significantly related to the candidate target gene in the target cell according to the correlation coefficient.

[0078] In a specific application scenario, the construction unit includes: constructing a module, configured to use the candidate genes of the target cell as gene variables of the target cell, and use the covariate gene set to construct different sample data sets for model training of the gene variables of the target cell, to obtain a first hypothesis model and a second hypothesis model; an acquisition module, configured to acquire a first prediction error of the first hypothesis model obtained through multiple iterations and a second prediction error of the second hypothesis model obtained through multiple iterations in the process of training the first hypothesis model and the second hypothesis model, the first prediction error being used to represent a prediction error of the target cell gene under regression modeling without the source cell gene, and the second prediction error being used to represent a prediction error of the target cell gene under regression modeling with the source cell gene; a determination module, configured to determine a causal correlation between the source cell gene and the target cell gene according to the first prediction error and the second prediction error.

[0079] In a specific application scenario, the constructing module is specifically configured to: use the candidate genes of the target cell as gene variables of the target cell, and use the covariate gene set as gene variables of the target cell of the sample data set for regression fitting, to obtain the first hypothesis model; use the candidate genes of the target cell as gene variables of the target cell, and introduce gene variables of the source cell as sample data set on the basis of the covariate gene set for regression modeling of the gene variables of the target cell, to obtain the second hypothesis model.

[0080] In a specific application scenario, the determination module is specifically configured to: calculate a causal strength score of the source cell gene to the target cell gene according to the first prediction error and the second prediction error; if the causal strength score is greater than a set value, it is determined that the source cell gene has a causal effect relationship to the target cell gene; if the causal strength score is less than or equal to the set value, it is determined that the source cell gene does not have a causal effect relationship to the target cell gene.

[0081] In a specific application scenario, the recognition unit is specifically configured to: construct a cross-cell-type gene regulation network according to the causal correlation between the source cell gene and the target cell gene, the gene regulation network using an adjacency matrix to describe the causal correlation between the source cell gene and the target cell gene, if the source cell gene has a causal effect relationship to the target cell gene, a corresponding element in the adjacency matrix is assigned a first value, and if the source cell gene does not have a causal effect relationship to the target cell gene, a corresponding element in the adjacency matrix is assigned a second value. The communication intensity of the source cell to the target cell is calculated by performing a weighted sum calculation on all elements in the adjacency matrix.

[0082] Based on the method as described above, Figures 1-5 Accordingly, the embodiments of the present application also provide a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for identifying the communication relationship between cells as described above. Figures 1-5

[0083] Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, etc.) and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.

[0084] Based on the method as described above, Figures 1-5 and Figure 7 In order to achieve the above-mentioned purpose, the embodiments of the present application also provide an entity device for identifying the communication relationship between cells, which can be a computer, a smart phone, a tablet computer, a smart watch, a server, or a network device, etc. The entity device includes a storage medium and a processor; the storage medium is used to store a computer program; and the processor is used to execute the computer program to implement the method for identifying the communication relationship between cells as described above. Figures 1-5

[0085] Optionally, the entity device can also include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface can include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. The optional user interface can also include a USB interface, a card reader interface, etc. The network interface can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.

[0086] In the exemplary embodiments, referring to Figure 8 The above-mentioned entity device includes a communication bus, a processor, a memory, and a communication interface, and can also include an input / output interface and a display device. The various functional units can communicate with each other through the bus. The memory stores a computer program, and the processor is used to execute the program stored in the memory to execute the method for identifying the communication relationship between cells in the above-mentioned embodiments.

[0087] ​​Those skilled in the art can understand that the entity device structure for identifying the intercellular communication relationship provided by the embodiment does not constitute a limitation on the entity device, and can include more or fewer components, or combine certain components, or different component arrangements.

[0088] The storage medium can also include an operating system and a network communication module. The operating system is a program for managing hardware and software resources of the entity device for identifying the intercellular communication relationship, and supports the running of information processing programs and other software and / or programs. The network communication module is used to realize communication between components inside the storage medium, and communication with other hardware and software in the information processing entity device.

[0089] Through the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and necessary general hardware platforms, or by hardware. By applying the technical solutions of the present application, compared with the existing manner, the present application can exclude the regulatory interference of the source cell on the target cell by controlling the covariate gene set expressed in the target cell, so as to accurately obtain the causal relationship of the source cell gene to the target cell gene, and quantize the direct causal contribution between cells using the causal relationship of the source cell gene to the target cell gene, which can make the intercellular communication relationship have causal interpretation, and then accurately evaluate the overall regulatory effect between cell populations.

[0090] Those skilled in the art can understand that the drawings are only a schematic diagram of a preferred implementation scenario, and the modules or flows in the drawings are not necessarily necessary for implementing the present application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be changed and located in one or more devices different from the implementation scenario. The modules of the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0091] The above application serial numbers are only for description, and do not represent the advantages and disadvantages of the implementation scenario. The above disclosure is only a few specific implementation scenarios of the present application, but the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.

Claims

1. A method for identifying intercellular communication relationships, characterized in that: include: Obtaining gene expression data of different cell types, wherein the gene expression data is the expression level of each gene in the cell's own genome, and the different cell types include at least source cells and target cells; Constructing different hypothesis models for the candidate genes of the target cells using covariate gene sets to determine the causal relationship between the source cell genes and the target cell genes through the different hypothesis models, wherein the covariate gene set is a set of genes in the target cells that are significantly correlated with the candidate genes; According to the causal relationship between the source cell gene and the target cell gene, the communication strength between the source cell and the target cell is calculated, so as to identify the communication relationship between the cells through the communication strength.

2. The method according to claim 1, characterized in that The obtaining of gene expression data of different cell types includes: Acquire single-cell expression data of different cell types, wherein the single-cell expression data includes source cell data and target cell data, and the single-cell expression data describes the expression level of genes in cells in the form of a single-cell transcriptome expression matrix; The single-cell expression data is preprocessed to obtain gene expression data that meets set format requirements, wherein the preprocessing includes at least one of format conversion, quality control, and standardization.

3. The method according to claim 1, characterized in that Before constructing different hypothesis models for the candidate genes of the target cell using the covariate gene set to determine the causal relationship between the source cell gene and the target cell gene through the different hypothesis models, the method further includes: For the candidate target gene of the target cell, calculating the correlation coefficient between the candidate target gene and other genes of the target cell; A covariate gene set significantly correlated with the candidate target gene is screened in the target cell according to the correlation coefficient.

4. The method according to claim 1, wherein The method of constructing different hypothesis models for the candidate genes of the target cell using a covariate gene set to determine the causal relationship between the source cell gene and the target cell gene through the different hypothesis models includes: Taking the candidate genes of the target cells as the gene variables of the target cells, constructing different sample data sets using the covariate gene set to perform model training on the gene variables of the target cells, and obtaining a first hypothesis model and a second hypothesis model; During the training of the first hypothesis model and the second hypothesis model, respectively obtaining a first prediction error of the first hypothesis model obtained after multiple iterations and a second prediction error of the second hypothesis model obtained after multiple iterations, wherein the first prediction error is used to represent the prediction error of the target cell gene under regression modeling that does not include the source cell gene, and the second prediction error is used to represent the prediction error of the target cell gene under regression modeling that includes the source cell gene; A causal relationship between the source cell gene and the target cell gene is determined based on the first prediction error and the second prediction error.

5. The method according to claim 4, characterized in that The method uses the candidate genes of the target cells as the gene variables of the target cells, constructs different sample data sets using the covariate gene set to perform model training on the gene variables of the target cells, and obtains a first hypothesis model and a second hypothesis model, including: Using the candidate genes of the target cells as the gene variables of the target cells and using the covariate gene set as the gene variables of the target cells in the sample data set to perform regression fitting to obtain a first hypothesis model; The candidate genes of the target cells are used as the gene variables of the target cells. On the basis of the covariate gene set, the gene variables of the source cells are introduced as the sample data set to jointly perform regression modeling on the gene variables of the target cells to obtain a second hypothesis model.

6. The method according to claim 4, characterized in that Determining the causal relationship between the source cell gene and the target cell gene based on the first prediction error and the second prediction error includes: Calculating a causal strength score of the source cell gene on the target cell gene according to the first prediction error and the second prediction error; If the causal strength score is greater than the set value, it is determined that the source cell gene has a causal effect on the target cell gene; If the causal strength score is less than or equal to the set value, it is determined that there is no causal effect relationship between the source cell gene and the target cell gene.

7. The method according to any one of claims 1 to 6, characterized in that The calculating the communication strength between the source cell and the target cell based on the causal relationship between the source cell gene and the target cell gene, so as to identify the communication relationship between the cells through the communication strength, includes: Constructing a gene regulatory network across cell types based on the causal relationship between the source cell gene and the target cell gene, wherein the gene regulatory network uses an adjacency matrix to describe the causal relationship between the source cell gene and the target cell gene, and if the source cell gene has a causal effect on the target cell gene, the corresponding element in the adjacency matrix is ​​assigned a first value; if the source cell gene does not have a causal effect on the target cell gene, the corresponding element in the adjacency matrix is ​​assigned a second value; A weighted sum calculation is performed on all elements in the adjacency matrix to obtain the communication strength of the source cell to the target cell, so as to identify the communication relationship between the cells through the communication strength.

8. A device for identifying intercellular communication relationships, characterized in that: include: an acquisition unit, configured to acquire gene expression data of different cell types, wherein the gene expression data is the expression level of each gene in the cell's own genome, and the different cell types include at least source cells and target cells; a construction unit configured to construct different hypothesis models for the candidate genes of the target cells using covariate gene sets, so as to determine the causal relationship between the source cell genes and the target cell genes through the different hypothesis models, wherein the covariate gene set is a gene set in the target cells that is significantly correlated with the candidate genes; The identification unit is used to calculate the communication strength of the source cell to the target cell based on the causal relationship between the source cell gene and the target cell gene, so as to identify the communication relationship between the cells through the communication strength.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Cell communication analysis method and system

    CN112466403A

  • Cell communication network identification method and device, equipment and storage medium

    CN114722988A

  • Cell group correlation network construction method and device, equipment and storage medium

    CN116564418A

  • Cell communication prediction method based on single cell data

    CN117612598A

  • Inter-cell communication inference method based on single cell transcriptome data

    CN118918962A

Cited By

  • Cell annotation method and device, electronic equipment and storage medium

    CN121415883A