Method for establishing gene co-expression network and related device

By extracting features from the single-cell gene expression matrix and constructing a gene co-expression network, the problem of complexity in the construction of gene co-expression network caused by high dimensionality and sparseness of single-cell data is solved, and the ability to predict accuracy and identify cell type-specific modules is improved.

CN119323985BActive Publication Date: 2025-06-13AGRICULTURAL GENOMICS INSTITUTE AT SHENZHEN CHINESE ACADEMY OF AGRICULTURAL SCIENCES (SHENZHEN BRANCH GUANGDONG LABORATORY FOR LINGNAN MODERN AGRICULTURE)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411856779.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-06-13
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

The high dimension, high noise and high sparsity of single-cell sequencing data increase the complexity of the construction of gene co-expression networks, leading researchers to face huge challenges in inferring gene expression from single-cell data.

Method used

By obtaining the target gene expression matrix, feature extraction is performed to obtain the gene characteristic matrix, and then a cell type-specific co-expression network or global gene co-expression network is constructed based on the gene characteristic matrix and the target gene expression matrix.

Benefits of technology

This method can avoid the problem of predicting the co-expression of single-cell genes directly based on sparse single-cell sequencing data, which improves the accuracy of single-cell gene co-expression prediction, and predicts more significantly related cell type-specific gene modules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119323985B_ABST
    Figure CN119323985B_ABST
Patent Text Reader

Abstract

The present application provides a method for establishing a gene co-expression network and related devices. Among them, the method for establishing a gene co-expression network includes: obtaining a target gene expression matrix; performing feature extraction on the target gene expression matrix to obtain a first gene feature matrix; constructing a global gene co-expression network according to the first gene feature matrix, and / or extracting a second gene feature matrix of different cell types according to the first gene feature matrix and the target gene expression matrix, and constructing a cell type-specific co-expression network according to the second gene feature matrix of different cell types. It solves the problem that traditional single-cell gene co-expression methods directly predict based on high-dimensional and highly sparse single-cell transcriptome sequencing data, resulting in a high false positive rate. It not only improves the accuracy of single-cell gene co-expression prediction, but also can predict more significantly related cell type-specific gene modules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method for establishing a gene co-expression network and related devices. Background Art

[0002] In the biological field, gene co-expression analysis is crucial for understanding complex biological processes. Among them, single-cell sequencing technology can analyze gene expression at single-cell resolution and reveal the heterogeneity between cells. However, due to the high-dimensionality, high-noise, and high-sparsity of single-cell sequencing data, it increases the complexity of network construction, making it a huge challenge for researchers to infer gene expression from single-cell data. Summary of the Invention

[0003] In view of this, this application provides a method for establishing a gene co-expression network and related devices to solve at least some of the above technical problems.

[0004] On the one hand, this application provides a method for establishing a gene co-expression network, including: obtaining a target gene expression matrix; the elements in the target gene expression matrix are used to measure the expression level of a certain gene in a certain cell; performing feature extraction on the target gene expression matrix to obtain a first gene feature matrix; the elements in the first gene feature matrix are used to measure the eigenvalue of the genes in the target gene expression matrix in different dimensions; constructing a corresponding cell type-specific co-expression network according to the first gene feature matrix and the target gene expression matrix; and / or, constructing a global gene co-expression network according to the first gene feature matrix.

[0005] In some embodiments, constructing a cell type-specific co-expression network according to the first gene feature matrix and the target gene expression matrix includes: determining a second gene feature matrix for each cell type according to the target gene expression matrix and the first gene feature matrix; each row in the second gene feature matrix is the eigenvalue of the genes in the first gene feature matrix in different dimensions, and each column is the eigenvalue of different genes in the same dimension; obtaining a corresponding cell type-specific module according to the second gene feature matrix of each cell type; performing hierarchical clustering on the genes in each cell type-specific module to obtain a corresponding cell type-specific co-expression network.

[0006] In some embodiments, determining the second gene feature matrix of each cell type according to the target gene expression matrix and the first gene feature matrix includes: obtaining the gene expression matrix of differentially expressed genes included in each cell type according to the target gene expression matrix; calculating the correlation between the gene expression matrix of differentially expressed genes included in each cell type and the first gene feature matrix to form a cell-embedded gene correlation matrix corresponding to the cell type; wherein each row in the cell-embedded gene correlation matrix represents the importance degree of different features in the corresponding cell, and each column represents the importance degree of the same feature in the corresponding cell; for each cell type, determining the second gene feature matrix of the corresponding cell type according to the cell-embedded gene correlation matrix of the corresponding cell type and the first gene feature matrix.

[0007] In some embodiments, determining the second gene feature matrix of the corresponding cell type according to the cell-embedded gene correlation matrix of the corresponding cell type and the first gene feature matrix includes: extracting features with a set dimension from the first gene feature matrix according to the cell-embedded gene correlation matrix as the features of the corresponding cell type; forming the second gene feature matrix according to the extracted features of the corresponding cell type and each gene included in the first gene feature matrix.

[0008] In some embodiments, obtaining the corresponding cell type-specific module according to the second gene feature matrix of each cell type includes: calculating the cosine similarity between each gene in the second gene feature matrix; performing clustering processing on the second gene feature matrix according to the cosine similarity between each gene to form a first sub-cell type-specific module and a second sub-cell type-specific module; calculating a first score of the first sub-cell type-specific module and a second score of the second sub-cell type-specific module; comparing the first score and the second score; wherein, in response to the first score being greater than the second score, determining the first sub-cell type-specific module as the corresponding cell type-specific module; in response to the first score being less than the second score, determining the second sub-cell type-specific module as the corresponding cell type-specific module; in response to the first score being equal to the second score, determining either the first sub-cell type-specific module or the second sub-cell type-specific module as the corresponding cell type-specific module.

[0009] In some embodiments, calculating the first score of the first sub-cell type specific module and the second score of the second sub-cell type specific module includes: calculating the first relative expression level of each gene in the first sub-cell type specific module in different cells; and calculating the first score according to each of the first relative expression levels and the number of genes included in the first sub-cell type specific module; and calculating the second relative expression level of each gene in the second sub-cell type specific module in different cells; and calculating the second score according to each of the second relative expression levels and the number of genes included in the second sub-cell type specific module.

[0010] In some embodiments, constructing the global gene co-expression network according to the first gene feature matrix includes: calculating the co-expression relationship between each gene according to the first gene feature matrix to obtain a gene co-expression correlation matrix; performing hierarchical clustering on the genes in the gene co-expression correlation matrix to construct the global gene co-expression network; the global gene co-expression network has a tree-like structure, wherein each branch of the tree-like structure includes a gene co-expression module; the gene co-expression module contains a group of genes with similar expression patterns.

[0011] In some embodiments, each row of the first gene feature matrix represents the feature values of the same gene in different dimensions, and each column represents the feature values of different genes in the same dimension; calculating the co-expression relationship between each gene according to the first gene feature matrix to obtain a gene co-expression correlation matrix includes: calculating the cosine similarity between each gene in the first gene feature matrix; the cosine similarity is used to measure the co-expression relationship between genes; obtaining the gene co-expression correlation matrix according to the obtained cosine similarity between each gene.

[0012] In some embodiments, performing hierarchical clustering on the genes in the gene co-expression correlation matrix to construct the global gene co-expression network includes: setting the number of genes included in the minimum gene module, and gradually processing the data in the gene co-expression correlation matrix according to the dynamic shear tree strategy to form at least one gene co-expression module; recording each obtained gene co-expression module to construct the global gene co-expression network.

[0013] In some embodiments, obtaining the target gene expression matrix includes: obtaining the single-cell transcriptome sequencing (scRNA-seq) data set of at least one biological sample, and generating a sub-expression matrix of the genes included in the corresponding biological sample according to each scRNA-seq data set; preprocessing at least one of the sub-expression matrices to obtain the target gene expression matrix.

[0014] In some embodiments, the preprocessing includes at least one of the following: quality control, screening for highly variable genes, and data normalization.

[0015] In some embodiments, performing feature extraction on the target gene expression matrix to obtain a first gene feature matrix includes: using a feature extraction unit included in a trained gene clustering model to perform feature extraction on the target gene expression matrix to obtain the first gene feature matrix; wherein, the feature extraction unit includes an input layer, a convolutional layer, an activation function layer, a pooling layer, and a fully connected layer, the pooling layer is a spatial pyramid pooling layer, and the output result of the pooling layer is used as the input of the fully connected layer; the convolutional layer includes N, and the output result of the i-th convolutional layer is the input of the (i + 1)-th convolutional layer, and the output result of the N-th convolutional layer is used as the input of the pooling layer after being processed by the activation function layer, where N and i are positive integers, and i is less than N.

[0016] On the one hand, the present application provides a device for establishing a gene co-expression network, including:

[0017] An acquisition module, configured to acquire a gene expression matrix; elements of the gene expression matrix are used to measure the expression level of a certain gene in a certain single cell;

[0018] An extraction module, configured to perform feature extraction on the target gene expression matrix to obtain a first gene feature matrix; elements in the first gene feature matrix are used to measure eigenvalue features of genes in the target gene expression matrix in different dimensions;

[0019] A construction module, configured to construct a corresponding cell type-specific co-expression network according to the first gene feature matrix and the gene expression matrix; and / or construct a global gene co-expression network according to the first gene feature matrix.

[0020] On the one hand, the present application provides a device for establishing a gene co-expression network, including a processor, configured to call a program in a memory, so that the device executes the method described in any one of the foregoing items.

[0021] On the one hand, the present application provides a chip, including a processor, configured to call a program from a memory, so that a device installed with the chip executes the method described in any one of the foregoing items.

[0022] On the one hand, the present application provides a computer-readable storage medium, on which a program is stored, and the program enables a computer to execute the method described in any one of the foregoing items.

[0023] An embodiment of the present application provides a method for establishing a gene co-expression network and related devices. Among them, the method for establishing the co-expression network includes: obtaining a target gene expression matrix; the elements in the target gene expression matrix are used to measure the expression level of a certain gene in a certain cell; performing feature extraction on the target gene expression matrix to obtain a first gene feature matrix; the elements in the first gene feature matrix are used to measure the eigenvalue of the genes in the target gene expression matrix in different dimensions; constructing a corresponding cell type-specific co-expression network according to the first gene feature matrix and the target gene expression matrix; and / or constructing a global gene co-expression network according to the first gene feature matrix. The method for establishing a gene co-expression network provided by the embodiment of the present application obtains a first gene feature matrix by performing feature extraction on the target gene expression matrix, and then constructs a global gene co-expression network according to the first gene feature matrix, and / or constructs a corresponding cell type-specific co-expression network according to the first gene feature matrix and the target gene expression matrix. In this way, using gene features to construct a global gene co-expression network and / or a cell type-specific co-expression network can avoid the problem of high false positives caused by directly predicting single-cell gene co-expression based on sparse single-cell sequencing data, not only improving the accuracy of single-cell gene co-expression prediction, but also being able to predict more significantly related cell type-specific gene modules. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 The figure shows a schematic flowchart of a method for establishing a gene co-expression network provided by an embodiment of the present application.

[0025] Figure 2 The figure shows an exemplary schematic diagram of a target gene expression matrix provided by an embodiment of the present application.

[0026] Figure 3 The figure shows a schematic structural diagram of a feature extraction unit in a gene clustering model provided by an embodiment of the present application.

[0027] Figure 4 The figure shows an exemplary schematic diagram of a first gene feature matrix provided by an embodiment of the present application.

[0028] Figure 5 The figure shows a schematic flowchart of a method for constructing a cell type-specific co-expression network provided by an embodiment of the present application.

[0029] Figure 6 The figure shows a schematic flowchart of a process for determining a corresponding cell type-specific module provided by an embodiment of the present application.

[0030] Figure 7 The figure shows a schematic flowchart of a method for constructing a global gene co-expression network provided by an embodiment of the present application.

[0031] Figure 8 The following is a schematic structural diagram of a gene co-expression network establishment device provided by an embodiment of the present application.

[0032] Figure 9 The following is a device for performing gene co-expression network establishment provided by an embodiment of the present application. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.

[0034] As described above, gene co-expression analysis is crucial for understanding complex biological processes. For example, in the study of microarray and conventional transcriptome sequencing (i.e., RNA-seq) data, gene co-expression analysis has shown excellent performance in gene module detection, disease discovery, and detection of biologically relevant genes. In particular, the development of single-cell RNA sequencing (i.e., scRNA-seq) technology has enabled the analysis of gene expression at single-cell resolution, revealing cell-to-cell heterogeneity. Currently, multiple methods or metrics have been proposed to infer gene co-expression relationships from scRNA-seq data, including but not limited to Pearson correlation coefficient, Spearman correlation coefficient, GENIE3, PIDC, minet, sclink, and glasso. When inferring gene co-expression from scRNA-seq data, these aforementioned metrics or methods consider different association measures. For example, the currently popular weighted gene co-expression network analysis (WGCNA) uses Pearson as the co-expression correlation metric, and this method is mainly applied to the construction of conventional transcriptome gene co-expression networks. PIDC and Minet use mutual information (MI) as the co-expression measurement method. GENIE3 is based on an ensemble of trees method to estimate the correlation between genes. sclink modifies the Gaussian graphical model by explicitly modeling in single-cell data for gene co-expression estimation. However, due to the high dimensionality and high sparsity of scRNA-seq data, the aforementioned metrics or methods have low accuracy in constructing gene co-expression networks. In addition, the aforementioned metrics or methods perform poorly in dealing with large-scale transcriptome integration data with multiple sources and high heterogeneity and cannot identify gene modules with relatively high correlations. Also, for methods of inferring cell-type specific co-expression networks from large datasets, such as sclink and CS-CORE, genes highly expressed in cell types are directly screened for network construction, but these methods will miss some genes with low expression in cell types, resulting in inaccurate final constructed networks. Therefore, there is an urgent need to develop new computational methods to effectively extract and integrate gene co-expression information from these large scRNA-seq data to improve the accuracy of gene co-expression prediction. And improve the current way of constructing gene co-expression networks specific to cell types from large samples to infer more accurate cell-type specific networks and reveal their potential gene co-expression patterns.

[0035] To this end, the embodiments of the present application provide a method for establishing a co-expression network. By obtaining a target gene expression matrix and performing feature extraction on the target gene expression matrix to obtain a first gene feature matrix, then, using the first gene feature matrix to construct a global gene co-expression network, and / or using the first gene feature matrix and the target gene expression matrix to construct a corresponding cell type-specific co-expression network. In this way, using gene features to construct a global gene co-expression network and / or a cell type-specific co-expression network can avoid the problem of high false positives caused by directly predicting single-cell gene co-expression based on sparse single-cell sequencing data, not only improving the accuracy of single-cell gene co-expression prediction, but also being able to predict more significantly related cell type-specific gene modules.

[0036] The embodiments described in the present application will be elaborated in detail below with reference to the accompanying drawings.

[0037] As Figure 1 shown, it shows a schematic flowchart of a method for establishing a gene co-expression network provided by the embodiments of the present application. In Figure 1 it, the method for establishing a gene co-expression network may include:

[0038] Step 101: Obtain a target gene expression matrix; the elements of the target gene expression matrix are used to measure the expression level of a certain gene in a certain single cell.

[0039] It should be noted that so-called gene expression refers to the process of synthesizing functional gene products from the genetic information of genes. Among them, gene expression products are usually proteins, and all known life uses gene expression to synthesize macromolecules of life. The gene expression process may include the following three main steps: transcription, RNA processing, and translation. The so-called gene expression matrix may be a core data structure in bioinformatics and transcriptomics research, and it can be generated by high-throughput sequencing technologies such as scRNA-seq datasets. The rows of this gene expression matrix represent genes, and the columns represent cells. Among them, each element (and / or cell) in the gene expression matrix contains the expression level of the corresponding gene in the corresponding cell. An example of a gene expression matrix is Figure 2 shown.

[0040] In some embodiments, for step 101, it may include: obtaining scRNA-seq datasets of at least one biological sample, and generating sub-expression matrices of genes included in the corresponding biological samples according to each scRNA-seq dataset;

[0041] Performing preprocessing on at least one of the sub-expression matrices to obtain the target gene expression matrix.

[0042] It should be noted that the target gene expression matrix mentioned in this application can be a matrix composed of the expression of multiple genes extracted from at least one biological sample. In other words, the genes or cells in the target gene expression matrix can come from at least one biological sample. The biological sample can be any organism with genetic genes, such as mice, certain tissues of humans (such as cortical tissue, peripheral blood mononuclear cells, etc.). Among them, after each biological sample undergoes scRNA-seq, an scRNA-seq dataset will be obtained; each scRNA-seq dataset will generate a corresponding sub-expression matrix.

[0043] After obtaining each sub-expression matrix, preprocess the at least one sub-expression matrix to obtain the target gene expression matrix. Among them, the preprocessing may include at least one of the following: quality control, screening of highly variable genes, and data normalization.

[0044] It should be noted that the so-called quality control may mean that since the sequencing quality of the scRNA-seq dataset obtained after scRNA-seq will seriously affect subsequent analysis, therefore, the obtained scRNA-seq dataset can be subjected to quality control to filter out genes or cells that do not meet the conditions. For example, in some embodiments, the purpose of this quality control is mainly to remove low-quality cells and genes. Therefore, the scanpy in the python package can be used to perform quality control on each obtained scRNA-seq dataset. Among them, scanpy is a python package dedicated to processing single-cell data. That is, the filter_cells function and filter_genes function of the scanpy package remove the cells and genes with low expression levels in each scRNA-seq dataset to obtain the filtered expression data.

[0045] Here, the so-called screening of highly variable genes may mean selecting genes with large expression differences between different cells from the scRNA-seq dataset after the aforementioned quality control for research. For example, in this application, 1000 highly variable genes can be screened out from the scRNA-seq dataset that has passed quality control for analysis.

[0046] Here, the so-called normalization may mean that in order to avoid the disappearance or explosion of gradients during neural network training, the data of the screened highly variable genes can be normalized to limit the data of the screened highly variable genes within a suitable range. For example, the min-max normalization method can be used to map the data of the screened highly variable genes to between 0 and 1. In this way, the training speed of the gene clustering model can be accelerated and the training efficiency can be improved.

[0047] After the above preprocessing, a subsequent gene expression matrix is obtained. As described above Figure 2 The following shows an exemplary obtained gene expression matrix, where each column represents all genes corresponding to a cell, and each row represents all cells corresponding to a gene. For example Figure 2 the gene expression data in it includes a total of n cells, each cell includes m genes, and both n and m are positive integers. Each element represents the expression level of a gene in a cell, and the value of each element is between 0 and 1.

[0048] Step 102: Extract features from the target gene expression matrix to obtain a first gene feature matrix; the elements in the first gene feature matrix are used to measure the eigenvalue of the genes in the target gene expression matrix in different dimensions.

[0049] In some embodiments, for step 102, it may include: using the feature extraction unit included in the trained gene clustering model to extract features from the target gene expression matrix to obtain the first gene feature matrix; wherein, the feature extraction unit includes an input layer, a convolutional layer, an activation function layer, a pooling layer and a fully connected layer, the pooling layer is a spatial pyramid pooling layer, and the output result of the pooling layer is used as the input of the fully connected layer; there are N convolutional layers, and the output result of the i-th convolutional layer is the input of the (i + 1)-th convolutional layer, and the output result of the N-th convolutional layer is processed by the activation function layer and then used as the input of the pooling layer, where N and i are positive integers, and i is less than N.

[0050] That is to say, this feature extraction unit can be implemented by a deep neural network. An exemplary implementation is that this feature extraction unit can be a convolutional neural network (CNN, Convolutional Neural Network). For example, this feature extraction unit can be an AlexNet network.

[0051] Exemplarily, as Figure 3 shown, this feature extraction unit may include an input layer (input), a convolutional layer (Cnov), an activation function layer (Relu), a pooling layer (Pool) and a fully connected layer (FC). Among them, exemplarily, this feature extraction unit (such as the convolutional layer in the feature extraction unit) may meet one or more of the following: the number of convolutional kernels is 2 n, the size of the convolutional kernel is 3*1, the stride (i.e., the stride of the convolution) is set to 1, and the padding is set to 1; where n is a positive integer. Here, the convolutional kernel (which can also be called a filter or a feature detector) is a small matrix that slides (or convolves) over the input data (such as an image) and calculates the sum of the element-wise products of the convolutional kernel and the input data at each position. This process produces a new two-dimensional array, usually called a feature map or an activation map, which represents the intensity and position of certain features in the input data. Among them, the stride can refer to the interval at which the convolutional kernel moves over the input data, that is, the convolutional kernel matrix itself. Padding can refer to adding additional values to the edges of the input data to keep the input and output sizes consistent. It should be noted that in the embodiments of the present application, the number of convolutional kernels, the size of the convolutional kernel, the stride of the convolution, the padding, the number of convolutional layers, etc. in the feature extraction unit can also be changed according to actual needs.

[0052] In some embodiments, the pooling layer in the feature extraction unit can be a spatial pyramid pooling layer (SPPNet, Spatial Pyramid Pooling Net). Exemplarily, the output result of the pooling layer can be used as the input of the fully connected layer. In this way, by converting gene expression matrices of different sizes into fixed-size feature vectors, data of different sizes can be processed without changing the gene clustering model, thereby improving the adaptability and generalization of the gene clustering model. That is to say, the trained gene clustering model can be applicable to gene expression matrices of different sizes.

[0053] Exemplarily, the dimension of the output result of SPPNet can be set to 44 dimensions, so that the dimensions of the output results of gene expression matrices of different sizes after being processed by SPPNet are all 44 dimensions. Then, the output result of SPPNet is input into the fully connected layer for processing to obtain a gene feature matrix. In this way, the ability of the gene clustering model to extract features under different sample sizes can be improved, thereby improving the adaptability and generalization of the gene clustering model.

[0054] In some embodiments, except for the last convolutional layer in the feature extraction unit, there may be no pooling layer behind other convolutional layers. For example, the feature extraction unit may include N convolutional layers. The output result of the i-th convolutional layer can be the input of the (i + 1)-th convolutional layer. The output result of the N-th convolutional layer can be the input of the pooling layer after being processed by the activation function layer, where N and i are positive integers, and i is less than N. In this way, the computational amount of the gene clustering model can be effectively reduced.

[0055] After the feature extraction unit in the above-mentioned trained gene clustering model extracts features from the target gene expression matrix, a first gene feature matrix is obtained. The elements in the first gene feature matrix represent the expression intensity of genes in specific dimensions in the target gene expression matrix. That is, the elements in the first gene feature matrix are used to measure the eigenvalue of a certain gene in different dimensions in the target gene expression matrix.

[0056] The first gene feature matrix obtained in step 102 can be as Figure 4 shown. In Figure 4 , each row is the eigenvalue of the same gene in different dimensions, and each column is the eigenvalue of different genes in the same dimension. Among them, for example, Figure 4 the first gene feature matrix in totally includes p dimensions, that is, each gene contains eigenvalues in p dimensions. There are m genes in each dimension, and each gene corresponds to an eigenvalue. Both p and m are positive integers. In some embodiments, p is less than n (n is the number of cells).

[0057] Step 103: Construct a corresponding cell type-specific co-expression network according to the first gene feature matrix and the target gene expression matrix; and / or construct a global gene co-expression network according to the first gene feature matrix.

[0058] It should be noted that after obtaining the first gene feature matrix, a global gene co-expression network and / or a corresponding cell type-specific co-expression network can be obtained according to requirements.

[0059] In some embodiments, for the construction of the corresponding cell type-specific co-expression network according to the first gene feature matrix and the target gene expression matrix in step 103, as Figure 5 shown, it may include:

[0060] Step 501: Determine the second gene feature matrix of each cell type according to the target gene expression matrix and the first gene feature matrix; each row in the second gene feature matrix is the eigenvalue of the gene in the first gene feature matrix in different dimensions, and each column is the eigenvalue of different genes in the same dimension;

[0061] Step 502: Obtain the corresponding cell type-specific module according to the second gene feature matrix of each cell type;

[0062] Step 503: Perform hierarchical clustering on the genes in each cell type-specific module to obtain the corresponding cell type-specific co-expression network.

[0063] In some embodiments, for step 501, it may include: obtaining the gene expression matrix of the differentially expressed genes contained in each cell type according to the target gene expression matrix; calculating the correlation between the gene expression matrix of the differentially expressed genes contained in each cell type and the first gene feature matrix to form a cell-embedded gene correlation matrix corresponding to the cell type; wherein, each row in the cell-embedded gene correlation matrix represents the importance degree of different features in the corresponding cell, and each column represents the importance degree of the same feature in the corresponding cell; for each cell type, determining the second gene feature matrix of the corresponding cell type according to the cell-embedded gene correlation matrix of the corresponding cell type and the first gene feature matrix.

[0064] Here, the so-called cell type may refer to a cell population classified based on specific morphology, function, structure, and molecular markers. In biology, the diversity of cell types is the basis of the complexity of organisms. Among them, there are different cell types according to different classification methods. For example, classified by morphological structure, it may include epithelial cell populations, muscle cell populations, nerve cell populations, and so on. Another example is classified by function, it may include: photoreceptor cell populations, olfactory cell populations, taste cell populations, and so on. The so-called differentially expressed gene (Differential Gene Expression) may refer to a gene whose gene expression level changes significantly under different conditions (such as different tissue types, different developmental stages, different environmental pressures, or different pathological states, etc.), that is, a gene with a relatively large expression difference. In some embodiments, a threshold may be set to screen differentially expressed genes. For example, when the gene expression similarity is less than a certain threshold, it is defined as a differentially expressed gene; another example is that if the gene expression similarity is greater than a certain threshold, it is defined as an expression-similar gene, and the rest are differentially expressed genes, and so on.

[0065] That is to say, for step 501, first obtain the gene expression matrix of the differentially expressed genes corresponding to each cell type from the target gene expression matrix; then, calculate the correlation between the gene expression matrix of the differentially expressed genes corresponding to each cell type and the first gene feature matrix calculated above to obtain the cell-embedded gene correlation matrix corresponding to each cell type. Each row in the cell-embedded gene correlation matrix represents the importance degree of different features in the corresponding cell, and each column represents the importance degree of the same feature in the corresponding cell; then, for each cell type, determine the second gene feature matrix of the corresponding cell type according to the cell-embedded gene correlation matrix of the corresponding cell type and the first gene feature matrix.

[0066] Among them, in some embodiments, the gene expression matrix of the differentially expressed genes contained in each cell type obtained according to the target gene expression matrix may refer to the gene expression matrix of the differentially expressed genes contained in each cell type screened from the target gene expression matrix.

[0067] Among them, in some embodiments, calculating the correlation between the gene expression matrix of the differentially expressed genes contained in each cell type and the first gene feature matrix to form a cell-embedded gene correlation matrix for the corresponding cell type may refer to evaluating the correlation between them by calculating the correlation coefficient (such as Pearson correlation coefficient) of the gene pairs contained in the corresponding cell type through the gene expression matrix of the differentially expressed genes contained in the cell type and the first gene feature matrix. Then, according to the correlation, combining the gene expression matrix of the differentially expressed genes contained in each cell type and the correlation matrix of the co-expression network, the cell-embedded gene correlation matrix can be obtained.

[0068] Among them, in some embodiments, for each cell type, determining the second gene feature matrix of the corresponding cell type according to the cell-embedded gene correlation matrix of the corresponding cell type and the first gene feature matrix may include: extracting features with a set dimension from the first gene feature matrix according to the cell-embedded gene correlation matrix as the features of the corresponding cell type; and forming the second gene feature matrix according to the extracted features of the corresponding cell type and each gene contained in the first gene feature matrix.

[0069] That is, after obtaining the cell-embedded gene correlation matrix, the first gene feature matrix is combined to determine the second gene feature matrix. Specifically, features with a set dimension are extracted from the first gene feature matrix according to the cell-embedded gene correlation matrix as the features of the corresponding cell type. For example, 100 dimensions are extracted as the features of the corresponding cell type. Then, according to the extracted features of the corresponding cell type and each gene contained in the first gene feature matrix, the second gene feature matrix is formed. For example, if the first gene feature matrix contains 1000 genes, then the dimension of the obtained second gene feature matrix is (1000, 100), where 1000 is the number of genes and 100 is the feature dimension.

[0070] After that, in some embodiments, for step 502, as Figure 6 shown, it may include:

[0071] Step 601: Calculate the cosine similarity between each gene in the second gene feature matrix;

[0072] Step 602: Perform clustering processing on the second gene feature matrix according to the cosine similarity between the respective genes to form a first cell type-specific module and a second cell type-specific module;

[0073] Step 603: Calculate a first score of the first cell type-specific module and a second score of the second cell type-specific module;

[0074] Step 604: Compare the first score and the second score; wherein, in response to the first score being greater than the second score, determine the first cell type-specific module as the corresponding cell type-specific module; in response to the first score being less than the second score, determine the second cell type-specific module as the corresponding cell type-specific module; in response to the first score being equal to the second score, determine either the first cell type-specific module or the second cell type-specific module as the corresponding cell type-specific module.

[0075] It should be noted that for the first two steps in step 502: the manner of obtaining the first cell type-specific module and the second cell type-specific module is similar to the steps of steps 701 and 702 described below, and the detailed details will not be elaborated here. Only one point needs to be explained. The number of clusters is set to 2 here, that is, clustering forms the first cell type-specific module and the second cell type-specific module.

[0076] In some embodiments, for calculating the first score of the first cell type-specific module and the second score of the second cell type-specific module in step 502, it may include: calculating the first relative expression level of each differentially expressed gene in the first cell type-specific module in different cells; and calculating the first score according to each first relative expression level and the number of genes included in the first cell type-specific module; and calculating the second relative expression level of each differentially expressed gene in the second cell type-specific module in different cells; and calculating the second score according to each second relative expression level and the number of genes included in the second cell type-specific module.

[0077] It should be noted that the so-called cell type specificity refers to the differences in gene expression levels among different cell types, and these differences reflect the functions and characteristics of different cell types. The gene expression patterns of cell type specificity are of great significance for understanding cell development, tissue function, and disease mechanisms.

[0078] The embodiments of the present application define a new cell type-specific score to screen cell type-specific modules, and the specific steps are as follows:

[0079] First, calculate the relative expression levels of genes in different cells. For example, the relative expression level of gene g is calculated by dividing the expression level of gene g in cell c by the sum of the expression levels of gene g in all cells of that cell type. The calculation formula is as follows:

[0080] (1)

[0081] Wherein, represents the expression level of gene g in cell c, and C represents the total number of cells contained in a certain cell type. represents the relative expression level of gene g in cell c.

[0082] Second, calculate the cell type-specific scores in different sub-cell type-specific modules, that is, the first score or the second score. Specifically, the sum of the cell type-specific scores of all genes in the sub-cell type-specific module is divided by the number of genes in the sub-cell type-specific module to obtain the cell type-specific score of the sub-cell type-specific module. The calculation formula is as follows:

[0083] (2)

[0084] Wherein, represents the expression level of gene g in cell type t, which is also the sum of the relative expression levels of the genes calculated above in different cells. ∣ m ∣ represents the total number of genes in module m. represents the cell type-specific score of the sub-cell type-specific module, that is, the first score or the second score. That is, both the first score and the second score are calculated according to the above two steps.

[0085] After obtaining the first score of the first sub-cell type-specific module and the second score of the second sub-cell type-specific module, compare the magnitudes of the first score and the second score. When the first score is greater than the second score, determine the first sub-cell type-specific module as the corresponding cell type-specific module; if the first score is less than the second score, determine the second sub-cell type-specific module as the corresponding cell type-specific module. If the first score is equal to the second score, determine either the first sub-cell type-specific module or the second sub-cell type-specific module as the corresponding cell type-specific module. That is, compare the cell type-specific scores of different sub-cell type-specific modules and select the sub-cell type-specific module with the higher score as the specific module of this cell type.

[0086] After that, hierarchical clustering is performed on the screened cell type-specific modules to obtain a co-expression network corresponding to the cell type specificity. Among them, the co-expression network corresponding to the cell type specificity refers to a network formed by a gene set with similar gene expression patterns in a specific cell type. Such a network can help understand the uniqueness and functions of different cell types at the gene expression level.

[0087] In some other embodiments, the so-called Global Gene Co-expression Network can be a bioinformatics analysis method used to study the interactions and co-expression patterns between genes. This method can help us understand the functions and regulatory networks of genes in different biological processes. In the Global Gene Co-expression Network, each gene is represented as a node in the network, and the edges (connection lines) between the nodes represent the co-expression relationships between genes. These relationships are usually determined by calculating the correlation coefficients of gene expression data, such as Pearson correlation coefficient or Spearman correlation coefficient. The weight of the edges in the network can reflect the strength of the co-expression relationship. The higher the weight, the more similar the expression patterns of the two genes. That is, a co-expression network between each gene in the gene expression matrix is obtained according to the first gene feature matrix.

[0088] In some embodiments, for the construction of the global gene co-expression network according to the first gene feature matrix in step 103, as Figure 7 shown, it may include:

[0089] Step 701: Calculate the co-expression relationships between each gene according to the first gene feature matrix to obtain a gene co-expression correlation matrix;

[0090] Step 702: Perform hierarchical clustering on the genes in the gene co-expression correlation matrix to construct the global gene co-expression network; the global gene co-expression network has a tree-like structure, where each branch of the tree-like structure includes a gene co-expression module; the gene co-expression module contains a group of genes with similar expression patterns.

[0091] That is, constructing the global gene co-expression network based on the first gene feature may include the following two steps: calculating the co-expression relationship between each gene in the aforementioned target gene co-expression matrix according to the first gene feature matrix obtained previously to obtain a gene co-expression correlation matrix, that is, step 701; then, performing hierarchical clustering on the genes in the gene co-expression correlation matrix to obtain the global gene co-expression network, that is, step 702. Among them, the so-called hierarchical clustering is a commonly used clustering analysis method. It does not require specifying the number of clusters in advance, but constructs a clustering tree composed of hierarchies (called a dendrogram) by gradually merging or splitting. That is to say, the global gene co-expression network has a tree-like structure, and each branch of the tree-like structure includes a gene co-expression module; the gene co-expression module includes a group of genes with relatively high similarity in expression patterns. In implementation, a group of genes with relatively high similarity in expression patterns can be determined by a threshold.

[0092] Here, each row of the first gene feature matrix represents the feature values of the same gene in different dimensions, and each column represents the feature values of different genes in the same dimension; step 701 may include: calculating the cosine similarity between genes in the first gene feature matrix; the cosine similarity is used to measure the co-expression relationship between genes; obtaining the gene co-expression correlation matrix according to the obtained cosine similarity between each gene.

[0093] That is, the cosine similarity is used to calculate the co-expression relationship between genes to obtain the gene co-expression correlation matrix. The so-called cosine similarity is a similarity measurement method for measuring the angle between two vectors. It is determined by calculating the ratio of the dot product of two vectors and the product of their norms. The range of cosine similarity is from -1 to 1, where 1 indicates that the vectors have exactly the same direction, -1 indicates exactly the opposite direction, and 0 indicates that the vectors are orthogonal, that is, there is no correlation. In this application, the two vectors may refer to any two rows in the first gene feature matrix.

[0094] Among them, the calculation formula of cosine similarity is as follows:

[0095] Cosine Similarity = (A ⋅ B) / (||A|| * ||B||).

[0096] Where: A ⋅ B is the dot product of vectors A and B; ||A|| is the norm (or norm) of vector A; ||B|| is the norm (or norm) of vector B.

[0097] In some embodiments, for step 602, it may include: setting the number of genes included in the minimum gene module, and gradually processing the data in the gene co-expression correlation matrix according to the dynamic shear tree strategy to form at least one gene co-expression module; recording each obtained gene co-expression module to construct the global gene co-expression network.

[0098] Here, setting the number of genes included in the minimum gene module may be the smallest unit of the specified gene module. For example, if the number of genes included in the set minimum gene module is 10, then the clustered gene module contains at least 10 genes. The dynamic shear tree strategy is a method used to identify gene modules in weighted gene co-expression network analysis (WGCNA). Therefore, for step 602, a global gene co-expression network can be constructed according to the dynamic shear tree strategy.

[0099] The gene co-expression network establishment method provided by the embodiments of the present application constructs a global gene co-expression network or a cell type-specific co-expression network by converting a high-dimensional and sparse single-cell gene expression matrix into a low-dimensional gene feature matrix, realizing matrix decoupling to avoid directly constructing a network based on sparse single-cell data and improving the accuracy of co-expression calculation. That is, the embodiments of the present application provide a decoupled feature learning framework (DeepCSCN) to construct a cell type-specific gene co-expression network (that is, the aforementioned cell type-specific co-expression network). Among them, DeepCSCN can infer a cell type-specific co-expression network from a large-scale sample by decoupling cell type features (that is, separating the feature dimensions related to different cell types). DeepCSCN has the following innovations and advantages: (1) Effective feature extraction: DeepCSCN uses a deep clustering model (that is, GeneCluster, that is, the aforementioned trained gene clustering model) to extract gene features, solves the sparsity problem of scRNA-seq data, and improves the accuracy of gene correlation calculation. (2) Higher accuracy: Compared with existing methods, DeepCSCN has higher co-expression prediction accuracy and can identify more cell type-specific gene modules. (3) Global to local network construction: DeepCSCN first constructs a global gene co-expression network using the embedded features of all genes; then decouples the embeddings of specific cell types to construct a cell type-specific co-expression network.

[0100] To understand the present application, the following provides examples for illustration.

[0101] Among them, the process of the gene co-expression network establishment method can be as follows:

[0102] (1) Data collection and preprocessing

[0103] In this application, we obtained a total of 8 single-cell datasets from Single Cell Protal, namely 4 mouse cortex datasets and 4 human peripheral blood mononuclear cell (PBMC) datasets. In the following description, we mainly take the human PBMC1 drop-seq dataset as an example. This dataset includes 9 cell types, a total of 10,802 cells and 22,792 genes. First, this dataset is initially filtered for genes and cells using the Python package scanpy, then 1,000 highly variable genes are selected, and the original count matrix is normalized to obtain a gene expression matrix. Then, the data in the gene expression matrix is input into GeneCluster for training to obtain the first gene feature matrix.

[0104] (2) Model training to obtain the first gene feature matrix

[0105] The data in the preprocessed gene expression matrix is input into the GeneCluster model for feature extraction to obtain the first gene feature matrix, where the dimension of this first gene feature matrix is (1000, 2048).

[0106] (3) Construction of the global gene co-expression network

[0107] The gene feature matrix of the PBMC1 drop-seq dataset extracted by the GeneCluster model is used to calculate the correlation matrix between genes based on cosine similarity, and hierarchical clustering is performed on the correlation matrix to obtain the global gene co-expression modules of the PBMC1 drop-seq dataset. Finally, through experiments, a total of 25 global gene co-expression modules are identified for this dataset. Then we associate different gene modules with cell types and find that module 1 is significantly correlated with CD16+ monocyte and Dendritic cell, module 2 is significantly correlated with Megakaryocyte, module 3 is significantly correlated with Cytotoxic T cell, Dendritic cell and Natural killer cell, module 5 is significantly correlated with Plasmacytoid dendritic cell, and module 8 is significantly correlated with B cell.

[0108] (4) Construction of the cell type-specific co-expression network

[0109] Although the global gene co-expression network can associate gene modules with cell types, it cannot accurately assign each module to a specific cell type. The construction of the cell type-specific network for the PBMC1 drop-seq dataset is mainly divided into three steps. First, we first generate a correlation matrix to quantify the importance of gene representations in identifying different cell types. This is achieved by first extracting the differentially expressed genes in different cell types of this dataset, calculating the correlation between the gene expression matrix of these differentially expressed genes and the gene feature matrix, and obtaining the correlation matrix of cell-embedding, with dimensions (10802, 2048); Second, extract the top 100-dimensional features of a specific cell type (such as B cells, with a cell count of 477) as the features of B cells, with dimensions (477, 100). Calculate the cosine similarity based on the gene features (1000, 100) of all genes in the specific cell type B cells, and perform clustering. Use the cell type-specific score to screen the B cell-specific module. Finally, the B cell-specific module we screened contains a total of 182 genes. Then, cluster the genes of this module again (182, 100), and finally, a total of 4 specific B cell co-expression modules are identified. Run GO enrichment analysis on each module, and it is found that module 1 is mainly enriched in "MHC protein complex binding", module 2 is mainly enriched in "B cell activation", and module 3 is mainly enriched in "immunoglobulin receptor binding functions". Apply the same dataset to sclink and CS-CORE for comparison. In the identification of B cell gene co-expression modules, CS-CORE identified a total of 4 B cell function modules, but only one module is related to B cells; sclink did not identify any B cell-related function modules. Therefore, compared with other cell type-specific network construction methods, DeepCSCN can identify more and more relevant functional pathways.

[0110] (5)Comparison of different gene co-expression methods

[0111] We compared different gene co-expression methods, including pearson, spearman, GENIE3, minet, PIDC, sclink, glasso, and CS-CORE. Also taking the PBMC1 drop-seq dataset as an example, for different co-expression methods, we used AUPRC and AUROC as evaluation metrics. AUROC is a curve graph of the false positive rate (FDR) and the true positive rate (TPR), and AUPRC is the area under the curve formed by the recall rate and the precision rate. We first screened 400 highly variable genes for comparison, calculated the gene correlation matrix of different co-expression methods, and based on this matrix, calculated the AUPRC and AUROC scores of different methods. It was found that in the PBMC1 drop-seq dataset, the AUPRC scores of pearson, spearman, GENIE3, minet, PIDC, sclink, glasso, CS-CORE, and DeepCSCN were 0.50, 0.48, 0.40, 0.07, 0.48, 0.51, 0.18, 0.20, and 0.60 respectively. Their AUROC scores were 0.89, 0.89, 0.74, 0.30, 0.88, 0.89, 0.54, 0.59, and 0.91 respectively. Therefore, when DeepCSCN was compared with other methods, DeepCSCN had the highest AUPRC and AUROC scores, effectively improving the accuracy of co-expression prediction.

[0112] As Figure 8 shown, the embodiment of the present application provides a schematic structural diagram of a gene co-expression network establishment device. In Figure 8 it, the co-expression network establishment device 800 may include:

[0113] An acquisition module 801, configured to acquire a target gene expression matrix; the elements in the target gene expression matrix are used to measure the expression level of a certain gene in a certain cell;

[0114] An extraction module 802, configured to perform feature extraction on the target gene expression matrix to obtain a first gene feature matrix; the elements in the first gene feature matrix are used to measure the eigenvalue of the genes in the target gene expression matrix in different dimensions;

[0115] A construction module 803, configured to construct a corresponding cell type-specific co-expression network according to the first gene feature matrix and the target gene expression matrix; and / or, construct a global gene co-expression network according to the first gene feature matrix.

[0116] In some embodiments, the construction module 803 is further configured to: determine a second gene feature matrix for each cell type according to the target gene expression matrix and the first gene feature matrix; each row in the second gene feature matrix is the eigenvalue of the gene in the first gene feature matrix in different dimensions, and each column is the eigenvalue of different genes in the same dimension; obtain a corresponding cell type-specific module according to the second gene feature matrix of each cell type; perform hierarchical clustering on the genes in each cell type-specific module to obtain a corresponding cell type-specific co-expression network.

[0117] In some embodiments, the construction module 803 is further configured to: obtain a gene expression matrix of differentially expressed genes included in each cell type according to the target gene expression matrix; calculate the correlation between the gene expression matrix of differentially expressed genes included in each cell type and the first gene feature matrix to form a cell-embedded gene correlation matrix corresponding to the cell type; wherein, each row in the cell-embedded gene correlation matrix represents the importance degree of different features in the corresponding cell, and each column represents the importance degree of the same feature in the corresponding cell; for each cell type, determine the second gene feature matrix of the corresponding cell type according to the cell-embedded gene correlation matrix of the corresponding cell type and the first gene feature matrix.

[0118] In some embodiments, the construction module 803 is further configured to: extract features of a set dimension from the first gene feature matrix according to the cell-embedded gene correlation matrix as the features of the corresponding cell type; form the second gene feature matrix according to the extracted features of the corresponding cell type and each gene included in the first gene feature matrix.

[0119] In some embodiments, the construction module 803 is further configured to: calculate the cosine similarity between each gene in the second gene feature matrix; perform clustering processing on the second gene feature matrix according to the cosine similarity between each gene to form a first sub-cell type-specific module and a second sub-cell type-specific module; calculate a first score of the first sub-cell type-specific module and a second score of the second sub-cell type-specific module; compare the first score and the second score; wherein, in response to the first score being greater than the second score, determine the first sub-cell type-specific module as the corresponding cell type-specific module; in response to the first score being less than the second score, determine the second sub-cell type-specific module as the corresponding cell type-specific module; in response to the first score being equal to the second score, determine the first sub-cell type-specific module or the second sub-cell type-specific module as the corresponding cell type-specific module.

[0120] In some embodiments, the building module 803 is further configured to: calculate the first relative expression level of each gene in the first sub-cell type specific module in different cells; calculate the first score according to each of the first relative expression levels and the number of genes included in the first sub-cell type specific module; calculate the second relative expression level of each gene in the second sub-cell type specific module in different cells; and calculate the second score according to each of the second relative expression levels and the number of genes included in the second sub-cell type specific module.

[0121] In some embodiments, the building module 803 is further configured to: calculate the co-expression relationship between genes according to the first gene feature matrix to obtain a gene co-expression correlation matrix; perform hierarchical clustering on the genes in the gene co-expression correlation matrix to construct the global gene co-expression network; the global gene co-expression network has a tree structure, wherein each branch of the tree structure includes a gene co-expression module; the gene co-expression module includes a set of genes with similar expression patterns.

[0122] In some embodiments, each row of the first gene feature matrix represents the eigenvalue of the same gene in different dimensions, and each column represents the eigenvalue of different genes in the same dimension; the building module 803 is further configured to: calculate the cosine similarity between genes in the first gene feature matrix; the cosine similarity is used to measure the co-expression relationship between genes; and obtain the gene co-expression correlation matrix according to the obtained cosine similarity between genes.

[0123] In some embodiments, the building module 803 is further configured to: set the number of genes included in the minimum gene module, and gradually process the data in the gene co-expression correlation matrix according to the dynamic shear tree strategy to form at least one gene co-expression module; record each obtained gene co-expression module to construct the global gene co-expression network.

[0124] In some embodiments, the obtaining module 801 is specifically configured to: obtain a single-cell transcriptome sequencing (scRNA-seq) dataset of at least one biological sample, and generate a sub-expression matrix of genes included in the corresponding biological sample according to each scRNA-seq dataset; preprocess at least one of the sub-expression matrices to obtain the target gene expression matrix.

[0125] In some embodiments, the extraction module 802 is specifically configured to: extract features from the target gene expression matrix by using the feature extraction unit included in the trained gene clustering model to obtain the first gene feature matrix; wherein, the feature extraction unit includes an input layer, a convolutional layer, an activation function layer, a pooling layer, and a fully connected layer, the pooling layer is a spatial pyramid pooling layer, and the output result of the pooling layer is used as the input of the fully connected layer; there are N convolutional layers, and the output result of the i-th convolutional layer is the input of the (i + 1)-th convolutional layer, and the output result of the N-th convolutional layer is processed by the activation function layer and then used as the input of the pooling layer, where N and i are positive integers, and i is less than N.

[0126] It should be noted that the gene co-expression network establishment device provided in the embodiments of the present application is a device for implementing the foregoing co-expression network establishment method. Therefore, the technical features appearing here have been described in detail in the above description of the co-expression network establishment method and can be understood by referring to the previous description. For the sake of saving space, they will not be elaborated here.

[0127] On the one hand, the present application provides a gene co-expression network establishment device, including a processor, which is used to call a program in a memory so that the device executes the method described in any one of the foregoing.

[0128] On the one hand, the present application provides a chip, including a processor, which is used to call a program from a memory so that a device installed with the chip executes the method described in any one of the foregoing.

[0129] On the one hand, the present application provides a computer-readable storage medium, on which a program is stored, and the program causes a computer to execute the method described in any one of the foregoing.

[0130] Figure 9 It is a schematic block diagram of a device 900 according to an embodiment of the present application. Figure 9 The illustrated device 900 includes a memory 901, a processor 902, a communication interface 903, and a bus 904. Among them, the memory 901, the processor 902, and the communication interface 903 are communicatively connected to each other through the bus 904.

[0131] The memory 901 may be a graphics processing unit (GPU, Graphics Processing Unit) storage system, a read-only memory (ROM, Read Only Memory), a static storage device, a dynamic storage device, and / or a random access memory (RAM, Random Access Memory). The memory 901 may store a program. When the program stored in the memory 901 is executed by the processor 902, the processor 902 is used to execute the steps of the method of the embodiments of the present application. For example, it may execute the foregoingFigure 1 , Figure 5 , Figure 6 , Figure 7 and other steps of the embodiments shown above.

[0132] The processor 902 may be a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an application specific integrated circuit (ASIC), and / or one or more integrated circuits for executing relevant programs to implement the method of the method embodiment of the present application.

[0133] The processor 902 may also be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the method of the embodiment of the present application may be completed by the integrated logic circuit in the hardware of the processor 902 and / or instructions in the form of software.

[0134] The above-mentioned processor 902 may also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit, a field programmable gate array (FPGA), and / or other programmable logic devices, discrete gates, and / or transistor logic devices, discrete hardware components. It can implement and / or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor and / or the processor may also be any conventional processor, etc.

[0135] The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by a hardware decoding processor, and / or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, and / or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 901, and the processor 902 reads the information in the memory 901 and combines its hardware to complete the functions required to be executed by the units included in each device of the embodiment of the present application, and / or execute the methods of the method embodiments of the present application. For example, it may execute Figure 1 , Figure 5 , Figure 6 and Figure 7 and other steps / functions of the embodiments shown above.

[0136] The communication interface 903 may use, but is not limited to, transceiver devices such as transceivers to implement communication between the device 900 and other devices or communication networks.

[0137] The bus 904 may include a path for transmitting information between various components of the device 900 (e.g., the memory 901, the processor 902, the communication interface 903).

[0138] It should be understood that the device 900 shown in the embodiments of the present application may be a processor or a chip for executing the methods described in the embodiments of the present application.

[0139] It should be understood that in the embodiments of the present application, the processor may be a graphics processing unit (GPU, Graphics Processing Unit), and the processor may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), and / or other programmable logic devices, discrete component gate circuits, and / or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor and / or the processor may also be any conventional processor, etc.

[0140] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.

[0141] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0142] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0143] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined and / or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0144] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place and / or may be distributed to multiple network units. Some and / or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0145] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0146] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, and / or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, and / or other programmable devices. The computer instructions can be stored in a computer-readable storage medium and / or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be read by a computer and / or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital versatile disc (DVD)), and / or a semiconductor medium (such as a solid state disk (SSD), etc.).

[0147] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for establishing a gene co-expression network, characterized in that: include: Obtain the target gene expression matrix; The elements in the target gene expression matrix are used to measure the expression level of a gene in a cell; Performing feature extraction on the target gene expression matrix to obtain a first gene feature matrix; the elements in the first gene feature matrix are used to measure the feature values ​​of genes in the target gene expression matrix in different dimensions; Obtaining a gene expression matrix of differentially expressed genes contained in each cell type according to the target gene expression matrix; Calculating the correlation between the gene expression matrix of the differentially expressed genes contained in each cell type and the first gene feature matrix to form a cell-embedded gene correlation matrix of the corresponding cell type; wherein each row in the cell-embedded gene correlation matrix represents the importance of different features in the corresponding cell, and each column represents the importance of the same feature in the corresponding cell; For each cell type, determining a second gene feature matrix of the corresponding cell type according to the cell-embedded gene correlation matrix of the corresponding cell type and the first gene feature matrix; each row in the second gene feature matrix is ​​a feature value of a gene in the first gene feature matrix in different dimensions, and each column is a feature value of different genes in the same dimension; Obtaining a corresponding cell type-specific module according to the second gene feature matrix of each cell type; The genes in each cell type-specific module were hierarchically clustered to obtain the corresponding cell type-specific co-expression network.

2. The method according to claim 1, characterized in that The determining the second gene feature matrix of the corresponding cell type according to the cell-embedded gene correlation matrix of the corresponding cell type and the first gene feature matrix comprises: Extracting features of a set dimension from the first gene feature matrix as features of the corresponding cell type according to the cell-embedded gene correlation matrix; The second gene feature matrix is ​​formed according to the extracted features of the corresponding cell type and the genes included in the first gene feature matrix.

3. The method according to claim 1, characterized in that The step of obtaining a corresponding cell type specific module according to the second gene feature matrix of each cell type comprises: Calculating the cosine similarity between each gene in the second gene feature matrix; Performing clustering processing on the second gene feature matrix according to the cosine similarity between the genes to form a first sub-cell type specific module and a second sub-cell type specific module; calculating a first score for the first sub-cell type-specific module and a second score for the second sub-cell type-specific module; Compare the first score and the second score; wherein, in response to the first score being greater than the second score, determine the first sub-cell type-specific module as the corresponding cell type-specific module; in response to the first score being less than the second score, determine the second sub-cell type-specific module as the corresponding cell type-specific module; in response to the first score being equal to the second score, determine the first sub-cell type-specific module or the second sub-cell type-specific module as the corresponding cell type-specific module.

4. The method according to claim 3, characterized in that The calculating a first score of the first sub-cell type-specific module and a second score of the second sub-cell type-specific module comprises: Calculating a first relative expression level of each gene in the first sub-cell type-specific module in different cells; and calculating the first score according to each of the first relative expression levels and the number of genes included in the first sub-cell type-specific module; and calculating a second relative expression level of each gene in the second sub-cell type-specific module in different cells; and calculating the second score according to each of the second relative expression levels and the number of genes included in the second sub-cell type-specific module.

5. The method according to claim 1, characterized in that The method further comprises: Calculate the co-expression relationship between each gene according to the first gene feature matrix to obtain a gene co-expression correlation matrix; The genes in the gene co-expression correlation matrix are hierarchically clustered to construct a global gene co-expression network; the global gene co-expression network has a tree structure, wherein each branch of the tree structure includes a gene co-expression module; the gene co-expression module includes a group of genes with similar expression patterns.

6. The method according to claim 5, characterized in that Each row of the first gene feature matrix represents the feature value of the same gene in different dimensions, and each column represents the feature values ​​of different genes in the same dimension; The step of calculating the co-expression relationship between the genes according to the first gene feature matrix to obtain a gene co-expression correlation matrix includes: Calculating the cosine similarity between genes in the first gene feature matrix; the cosine similarity is used to measure the co-expression relationship between genes; The gene co-expression correlation matrix is ​​obtained according to the cosine similarity between the obtained genes.

7. The method according to claim 5, characterized in that The step of performing hierarchical clustering on the genes in the gene co-expression correlation matrix to construct the global gene co-expression network comprises: The number of genes included in the minimum gene module is set, and the data in the gene co-expression correlation matrix is ​​gradually processed according to the dynamic cutting tree strategy to form at least one gene co-expression module; each obtained gene co-expression module is recorded to construct the global gene co-expression network.

8. The method according to claim 1, characterized in that The step of obtaining the target gene expression matrix comprises: Obtaining a single-cell transcriptome sequencing scRNA-seq dataset of at least one biological sample, and generating a sub-expression matrix of genes included in the corresponding biological sample according to each scRNA-seq dataset; Preprocessing is performed on at least one of the sub-expression matrixes to obtain the target gene expression matrix.

9. The method according to claim 8, characterized in that The preprocessing includes at least one of the following: quality control, screening of highly variable genes, and data normalization.

10. The method according to claim 1, characterized in that The step of extracting features from the target gene expression matrix to obtain a first gene feature matrix includes: The feature extraction unit included in the trained gene clustering model is used to extract features from the target gene expression matrix to obtain the first gene feature matrix; wherein the feature extraction unit includes an input layer, a convolution layer, an activation function layer, a pooling layer and a fully connected layer, the pooling layer is a spatial pyramid pooling layer, and the output result of the pooling layer is used as the input of the fully connected layer; the convolution layer includes N layers, and the output result of the i-th convolution layer is the input of the i+1-th convolution layer, and the output result of the N-th convolution layer is processed by the activation function layer as the input of the pooling layer, wherein N and i are positive integers, and i is less than N.

11. A gene co-expression network establishment device, characterized in that: include: An acquisition module is used to obtain the target gene expression matrix; The elements of the target gene expression matrix are used to measure the expression level of a gene in a cell; An extraction module is used to extract features from the target gene expression matrix to obtain a first gene feature matrix; The elements in the first gene feature matrix are used to measure the characteristic values ​​of genes in the target gene expression matrix in different dimensions; A construction module is used to obtain a gene expression matrix of differentially expressed genes contained in each cell type according to the target gene expression matrix; calculate the correlation between the gene expression matrix of differentially expressed genes contained in each cell type and the first gene feature matrix to form a cell-embedded gene correlation matrix of the corresponding cell type; wherein each row in the cell-embedded gene correlation matrix represents the importance of different features in the corresponding cell, and each column represents the importance of the same feature in the corresponding cell; for each cell type, a second gene feature matrix of the corresponding cell type is determined according to the cell-embedded gene correlation matrix of the corresponding cell type and the first gene feature matrix; each row in the second gene feature matrix is ​​the feature value of genes in the first gene feature matrix in different dimensions, and each column is the feature value of different genes in the same dimension; obtain a corresponding cell type-specific module according to the second gene feature matrix of each cell type; hierarchically cluster the genes in each cell type-specific module to obtain a corresponding cell type-specific co-expression network.

12. A gene co-expression network establishment device, characterized in that: The device comprises a processor, configured to call a program from a memory so as to enable the device to execute the method according to any one of claims 1 to 10.

13. A chip, characterized in that: The device comprises a processor, which is used to call a program from a memory, so that a device equipped with the chip executes the method according to any one of claims 1 to 10.

14. A computer-readable storage medium, characterized in that: A program is stored thereon, and the program enables a computer to execute the method according to any one of claims 1 to 10.

15. A computer program product, characterized in that The method comprises a program for causing a computer to execute the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Cell communication network identification method and device, equipment and storage medium

    CN114722988A

  • Gene clustering model training method and gene clustering method and device

    CN116978456A