Neutrophil differentiation trajectory inference method and system
Through single-cell transcriptome sequencing technology and deep learning method, combined with dynamic correlation and Bayesian information criterion, a causal model was constructed to infer the differentiation trajectory of neutrophils, solving the problem of causal regulation in the existing technology and achieving higher-precision analysis of cell differentiation trajectory.
Patent Information
- Application Number
- CN202510460330.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The prior art relies on pseudo-time analysis in the inference of cell differentiation trajectory, ignoring the dynamic regulation of causality, resulting in limited prediction accuracy, making it difficult to comprehensively and accurately depict the dynamic picture of cell differentiation and the regulatory logic behind it.
Single-cell transcriptome sequencing technology was used to obtain gene expression data of neutrophils, and key genes were screened through the dynamic correlation feature selection algorithm, combined with deep learning method for feature mapping, multiple alternative causal models were constructed, and model scores were evaluated using Bayesian information criterion, and the best causal model was screened to infer the differentiation trajectory.
It improves the accuracy and efficiency of inferring cell differentiation trajectory, can have a deeper understanding of the differentiation process and regulatory mechanism of neutrophils, and provides a high-resolution and high-precision cell-level data basis.
Smart Images

Figure CN119993281A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cell differentiation trajectory analysis, and in particular, to a neutrophil differentiation trajectory inference method and system. Background Art
[0002] Cell differentiation, a core process in biology, describes how cells from the same source gradually evolve into cell populations with different morphologies, structures, and functions. It not only causes cells to differ in the spatial dimension, but also marks a significant change in the temporal dimension from their original state. The essence of differentiation lies in the precise regulation of the genome in time and space, that is, the selective expression of specific genes, some genes are activated, while others are silenced. This process ultimately leads to the synthesis of iconic proteins. It is worth noting that cell differentiation is often a one-way and irreversible journey.
[0003] For bioinformatics researchers, in-depth exploration of cell differentiation trajectories is a valuable way to tap into the value of bioinformatics, reveal intercellular heterogeneity, and unlock new biological insights (such as discovering rare cell types). It is widely used to analyze the single-cell gene expression dynamics of various cellular processes including differentiation, proliferation, and even carcinogenic transformation, providing a strong impetus for the innovation of biological research and medical technology.
[0004] However, most of the current traditional differentiation trajectory inference methods rely only on pseudo-time analysis, ignoring the key factor of dynamic regulation of causality, resulting in limited prediction accuracy of differentiation paths and key nodes, making it difficult to fully and accurately depict the dynamic picture of cell differentiation and the regulatory logic behind it.
[0005] Currently, no effective solution has been proposed for the problems in the related technologies. Summary of the invention
[0006] In view of this, the present invention provides a neutrophil differentiation trajectory inference method and system to solve the above-mentioned problems.
[0007] In order to solve the above problems, the specific technical solutions adopted by the present invention are as follows: According to one aspect of the present invention, a method for inferring neutrophil differentiation trajectories is provided, and the method for inferring neutrophil differentiation trajectories comprises the following steps: S1. Based on single-cell transcriptome sequencing technology, the gene expression data of neutrophils is obtained, and the gene expression data of neutrophils is processed by feature selection algorithm of dynamic correlation to obtain the gene expression data set; S2. Perform feature mapping on the gene expression dataset based on deep learning method, and build multiple candidate causal models based on the mapping results; and use Bayesian information criterion to evaluate the scores of the candidate models, and select the candidate model with the highest score as the best causal model according to the scoring results; S3. Constructing a differentiation trajectory network based on the causal relationship network in the best model, and inferring the differentiation trajectory of neutrophils based on the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory; Preferably, based on the single-cell transcriptome sequencing technology, the gene expression data of neutrophils is obtained, and the gene expression data of neutrophils is subjected to feature selection processing by a feature selection algorithm of dynamic correlation to obtain the gene expression data set, which comprises the following steps: S11. Collect gene expression data of neutrophils by single-cell transcriptome sequencing technology, construct a gene expression matrix, and preprocess the gene expression matrix; obtain gene characteristics of neutrophils based on the preprocessing results; S12, based on the gene features of neutrophils and the differentiation state of the cells, a clustering algorithm is used to generate a class label for each gene feature; S13. Determine the dynamic correlation weight of the gene feature by calculating the conditional mutual information of each gene feature to the class label and combining it with the dynamic correlation weight algorithm; select the gene features of granulocytes in an iterative manner according to the dynamic correlation weight of the gene feature, and construct a standard gene expression data set based on the selection results.
[0008] Preferably, the dynamic correlation weight of the gene feature is determined by calculating the conditional mutual information of each gene feature to the class label and combining the dynamic correlation weight algorithm; the gene features of granulocytes are selected in an iterative manner according to the dynamic correlation weight of the gene feature, and a standard gene expression data set is constructed according to the selection results, including the following steps: S131, initialize the gene expression data set, and construct a candidate feature set based on all the gene features of neutrophils; S132. For the gene features in the candidate feature set, calculate the conditional mutual information value between each gene feature and the class label, select the gene feature with the largest conditional mutual information value, remove it from the candidate feature set to the gene expression data set, and update the gene expression data set and the candidate feature set; S133, for the updated gene expression dataset and candidate feature set, calculating the additional newly added information of each gene feature in the candidate feature set to each gene feature in the gene expression dataset; S134, calculating the dynamic relevance weight of each gene feature in the current candidate feature set using a dynamic relevance weight algorithm according to the additional newly added information of each gene feature in the candidate feature set to each gene feature in the gene expression data set; S135. According to the dynamic correlation weight, select the gene feature with the largest weight value from the candidate feature set, add it to the gene expression data set, and remove it from the candidate feature set. Repeat steps S133-S135 until the number of gene features in the gene expression data set reaches a preset threshold, and obtain the final gene expression data set.
[0009] Preferably, for the updated gene expression dataset and candidate feature set, the calculation formula for calculating the additional new information amount of each gene feature in the candidate feature set to each gene feature in the gene expression dataset is: In the formula, extra ( X K , R ) represents each gene feature in the candidate feature set X K Gene expression datasets R Each gene feature X R The amount of additional information added; ∣ R ∣ represents the gene expression dataset R The number of genetic features in Y Represents the class label; I ( X R ; Y ) indicates gene characteristics X R and class labels Y The mutual information value between ; I ( X R ; Y ∣ X K ) indicates that in the gene feature X K Under the candidate conditions, gene features X R and class labels Y The conditional mutual information value between ; H ( X R ) indicates gene characteristics X R Information entropy of H ( Y ) represents the class label Y Information entropy.
[0010] Preferably, according to the additional additional information amount of each gene feature in the candidate feature set to each gene feature in the gene expression data set, the dynamic relevance weight algorithm is used to calculate the dynamic relevance weight of each gene feature in the current candidate feature set: In the formula, DW ( X K , R ) represents each gene feature in the current candidate feature set X K Dynamic relevance weight of extra ( X K , R ) represents each gene feature in the candidate feature set X K Gene expression datasets R Each gene feature X R The amount of additional information added; I ( X K ; Y ∣ X R ) indicates gene characteristics X K and class labels Y The conditional mutual information value between .
[0011] Preferably, feature mapping is performed on the gene expression data set based on a deep learning method, and multiple candidate causal models are constructed based on the mapping results; and the scores of the candidate models are evaluated using the Bayesian information criterion, and the candidate model with the highest score is selected as the best causal model according to the scoring results, which includes the following steps: S21. Use the autoencoder in the deep learning model to perform feature mapping on the gene features in the gene expression dataset, and build multiple alternative causal models based on the mapping results and in combination with the Granger causality analysis method; S22. Calculating the likelihood of the alternative causal model according to the Bayesian information criterion, and scoring the alternative causal model according to the likelihood of the alternative causal model to obtain a scoring result; S23. Based on the scoring results of each candidate causal model, select the model with the highest score as the best causal model, and analyze the causal structure through compression representation.
[0012] Preferably, the autoencoder in the deep learning model is used to perform feature mapping processing on the gene features in the gene expression data set, and based on the mapping results, multiple candidate causal models are constructed in combination with the Granger causality analysis method, including the following steps: S211, based on the pre-trained autoencoder, mapping the gene features in the gene expression dataset to a low-dimensional latent space through the encoder part in the autoencoder to obtain a low-dimensional feature representation; S212. Use low-dimensional feature representation in low-dimensional latent space and Granger causality analysis to identify causal relationships between genes, and generate causal network through causal integration; S213. Construct multiple alternative causal models based on the causal network.
[0013] Preferably, using low-dimensional feature representation in low-dimensional latent space, combined with Granger causality analysis method to identify the causal relationship between genes, and generating a causal relationship network through causal relationship integration includes the following steps: S2121. Pseudo-time sorting of neutrophils in low-dimensional latent space by trajectory inference method. S2122. According to the pseudo-time ordering of the single-cell data, the low-dimensional features of each gene are reordered to obtain a low-dimensional latent space sequence; S2123. Perform Granger causality analysis on low-dimensional latent space sequences to identify causal relationships between low-dimensional features; S2124. Through the decoder of the autoencoder, the causal relationship in the low-dimensional latent space is mapped back to the high-dimensional gene space to generate a gene causal network.
[0014] Preferably, the likelihood of the candidate causal model is calculated according to the Bayesian information criterion, and the candidate causal model is scored according to the likelihood of the candidate causal model, and obtaining the scoring result includes the following steps: S221. For each candidate causal model, calculating a maximum likelihood estimate of the candidate causal model under the gene expression data set by maximizing the likelihood function; S222. Calculate the total number of free parameters based on the model structure, and calculate the score of each alternative causal model based on the Bayesian Information Criterion formula.
[0015] Preferably, constructing a differentiation trajectory network according to the causal relationship network in the best model, and inferring the differentiation trajectory of neutrophils according to the differentiation trajectory network, and obtaining the prediction result of the neutrophil differentiation trajectory includes the following steps: S31. Integrate gene causal relationships according to the causal structure in the optimal causal model and generate a causal relationship network; S32, incorporate pseudo-time information into the causal network to construct a differentiation trajectory network containing dynamic regulation and differentiation information; S33. Use the differentiation trajectory network to perform dynamic simulation, infer the differentiation path and key nodes, and output the prediction results of the differentiation trajectory.
[0016] According to one aspect of the present invention, a neutrophil differentiation trajectory inference system is provided, the neutrophil differentiation trajectory inference system comprising: a gene expression data set construction module, an optimal causal model construction module and a differentiation trajectory inference module, and the gene expression data set construction module, the optimal causal model construction module and the differentiation trajectory inference module are sequentially connected; A gene expression data set construction module is used to obtain the gene expression data of neutrophils based on single-cell transcriptome sequencing technology, and perform feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation to obtain a gene expression data set; The best causal model construction module is used to perform feature mapping on the gene expression data set based on the deep learning method, and to construct multiple candidate causal models based on the mapping results; and to evaluate the scores of the candidate models using the Bayesian information criterion, and to select the candidate model with the highest score as the best causal model based on the scoring results; The differentiation trajectory inference module is used to construct a differentiation trajectory network based on the causal relationship network in the best model, and to infer the differentiation trajectory of neutrophils based on the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory.
[0017] The beneficial effects of the present invention are: 1. The present invention obtains the gene expression data of neutrophils through single-cell transcriptome sequencing technology, providing a high-resolution and high-precision cellular-level data basis for analysis, which is helpful for a deeper understanding of the differentiation process of neutrophils. The gene expression data are processed by a feature selection algorithm of dynamic correlation, which can screen out key genes closely related to neutrophil differentiation, reduce redundant information, and improve the accuracy and efficiency of subsequent analysis. A differentiation trajectory network is constructed based on the causal relationship network in the optimal model, and differentiation trajectory inference is performed, which can obtain the prediction results of the neutrophil differentiation trajectory, and help to reveal the dynamic process and regulatory mechanism of neutrophil differentiation.
[0018] 2. The present invention uses single-cell transcriptome sequencing technology to obtain gene expression data of neutrophils at the single cell level, and pre-processes the gene expression matrix to remove noise, correct deviations, and improve data quality. In combination with the differentiation state of the cells, a class label is generated for each gene feature through a clustering algorithm, which helps to classify genes with similar expression patterns into one category. By calculating the conditional mutual information of each gene feature for the class label and combining it with a dynamic correlation weight algorithm, the correlation between each gene feature and the differentiation state of neutrophils can be accurately evaluated. The gene features of neutrophils are selected in an iterative manner, and a standard gene expression data set is constructed according to the selection results. The gene expression data set can be gradually optimized, the gene features most relevant to neutrophil differentiation are retained, and redundant and noise information is removed.
[0019] 3. The present invention uses an autoencoder to perform feature mapping on gene expression data, maps high-dimensional data to a low-dimensional latent space, effectively reduces the dimension of the data while retaining key information. Combined with the Granger causality analysis method, the causal relationship between genes is identified in the low-dimensional latent space, avoiding the complexity and noise interference of high-dimensional data, and improving the accuracy of causal identification. The Bayesian information criterion is used to score the alternative causal models, and the fit and complexity of the model are comprehensively considered to ensure that the selected model has a good fitting effect and is not too complex. The differentiation trajectory network is used for dynamic simulation, which can intuitively display the path and key nodes of cell differentiation, providing a powerful tool for understanding the mechanism of cell differentiation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 is a flow chart of a method for inferring neutrophil differentiation trajectory according to an embodiment of the present invention; Figure 2 It is a principle block diagram of a neutrophil differentiation trajectory inference system according to an embodiment of the present invention.
[0021] In the figure: 1. Gene expression dataset construction module; 2. Optimal causal model construction module; 3. Differentiation trajectory inference module. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0023] According to an embodiment of the present invention, a method and system for inferring neutrophil differentiation trajectory are provided.
[0024] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1 As shown, according to one embodiment of the present invention, a method for inferring neutrophil differentiation trajectory is provided, and the method for inferring neutrophil differentiation trajectory comprises the following steps: S1. Based on single-cell transcriptome sequencing technology, the gene expression data of neutrophils is obtained, and the gene expression data of neutrophils is processed by feature selection algorithm of dynamic correlation to obtain the gene expression data set; As a preferred embodiment, based on single-cell transcriptome sequencing technology, the gene expression data of neutrophils is obtained, and the gene expression data of neutrophils is subjected to feature selection processing by a feature selection algorithm of dynamic correlation to obtain a gene expression data set, including the following steps: S11. Collect gene expression data of neutrophils by single-cell transcriptome sequencing technology, construct a gene expression matrix, and preprocess the gene expression matrix; obtain gene characteristics of neutrophils based on the preprocessing results; It should be noted that single-cell transcriptome sequencing technology can sequence the cell transcriptome at the single-cell level, that is, detect the differential expression of cells at a higher resolution. The core of this technology is the ability to successfully distinguish different single-cell transcripts in a sample.
[0025] First, neutrophils need to be extracted from the sample tissue and prepared into a single cell suspension. Then, these single cells are sequenced using single-cell transcriptome sequencing technology to obtain the gene expression data of each cell.
[0026] After sequencing is completed, the data obtained needs to be processed to construct a gene expression matrix. The rows of the gene expression matrix usually represent genes, the columns represent cells, and the values in the matrix represent the number of transcripts of each gene in each single cell.
[0027] After constructing the gene expression matrix, a series of preprocessing steps are required to improve the accuracy and reliability of subsequent analysis. Preprocessing steps usually include: Quality control: Eliminate low-quality cells and low-expression genes to reduce the impact of noise on subsequent analysis. This can be achieved by setting certain thresholds, such as eliminating cells with fewer than a certain number of expressed genes or genes with expression levels below a certain threshold.
[0028] Normalization: Since the sequencing depth may vary between different cells or genes, the gene expression matrix needs to be normalized to eliminate this difference. Commonly used normalization methods include RPK (reads per million reads) or TPM (transcripts per million transcripts).
[0029] Missing value processing: There may be missing values in the gene expression matrix, which may be caused by insufficient sequencing depth or too low gene expression. Cells or genes with missing values can be directly deleted and then filled with the average expression or median of the gene.
[0030] Specifically, no matter which filling method is used, the accuracy of the gene expression data after filling needs to be verified, especially whether the filled data can reflect the true expression level of key genes. This is a very important link, and the following are several verification methods: a. Biological validation: verify the expression of key genes after filling through experiments, perform RNA-Seq validation or real-time quantitative PCR (qPCR) validation to ensure the accuracy and biological rationality of the filling data. This is the most direct way to verify whether the filling method is reasonable.
[0031] b. Compare the filled data with other data processing methods (such as the results based on different missing value filling strategies) to check whether they are consistent. Different filling methods may affect the final analysis results, so comparing the results of different processing methods can help verify the rationality of data processing.
[0032] c. Simulation verification, by introducing missing values into a known data set (such as a simulated data set) and testing different filling methods to evaluate the effect of filling. If the filling method can effectively restore the missing parts in the simulated data, it means that the method may be effective in the actual data.
[0033] Pseudo-time sorting: Pseudo-time sorting is performed on the gene expression matrix to determine the time point sequence of the dynamic changes in gene expression.
[0034] After preprocessing, the gene features of neutrophils were further analyzed based on the gene expression matrix. The gene features include: By comparing gene expression data of neutrophils under different conditions or at different time points, differentially expressed genes can be identified that play important roles in neutrophil differentiation, development, or function.
[0035] By analyzing the gene co-expression relationship in the gene expression matrix, a gene co-expression network can be constructed. This network can reveal the interactions and regulatory relationships between genes and help understand the gene expression regulation mechanism of neutrophils.
[0036] By clustering the gene expression matrix, neutrophils can be divided into different subpopulations. Each subpopulation may have unique gene expression characteristics and biological functions, which helps to deeply understand the heterogeneity of neutrophils.
[0037] S12, based on the gene features of neutrophils and the differentiation state of the cells, a clustering algorithm is used to generate a class label for each gene feature; Specifically, the characteristic gene set extracted from the neutrophil gene expression data divides the cells into different differentiation states (such as early differentiation, mid-differentiation, and late differentiation) according to pseudo-time sorting or known biological knowledge. Each cell has a differentiation state label to guide the grouping analysis of gene features.
[0038] For each differentiation state, the gene expression submatrix of all cells in that state is extracted, and the extracted gene expression submatrix is subjected to dimensionality reduction processing; The K-means clustering algorithm is used to group gene features (i.e., divide genes into different classes), and each gene is assigned a class label to indicate its clustering category in this differentiation state.
[0039] For example, suppose there are 10 genes (Gene1 to Gene10) and 30 neutrophils (Cell1 to Cell30), which have been divided into three differentiation states: early, middle and late.
[0040] 1. Data preparation: There is a gene expression matrix with 30 rows (cells) × 10 columns (genes).
[0041] 2. Division of cell differentiation states: Through pseudo-time sorting or biological knowledge, Cell1-Cell10 is divided into early differentiation states, Cell11-Cell20 is divided into middle differentiation states, and Cell21-Cell30 is divided into late differentiation states.
[0042] 3. Extract sub-matrix: Early differentiation state submatrix: contains gene expression data of Cell1-Cell10.
[0043] Mid-stage differentiation state submatrix: contains gene expression data of Cell11-Cell20.
[0044] Late differentiation state submatrix: contains gene expression data of Cell21-Cell30.
[0045] 4. Dimensionality reduction: Apply the PCA algorithm to each submatrix to reduce the dimension of the gene expression data to 2D or 3D space.
[0046] 5. K-means clustering: K-means clustering was applied to the early differentiation state submatrix, and the 10 genes were clustered into 3 categories. It was assumed that Gene1-Gene3 were clustered into category 1, Gene4-Gene6 were clustered into category 2, and Gene7-Gene10 were clustered into category 3.
[0047] Similarly, K-means clustering was performed on the mid- and late-differentiation state submatrices, and class labels were assigned to each gene.
[0048] 6. Class label generation: For the early differentiation state, Gene1-Gene3 were labeled as “Class 1 (early)”, Gene4-Gene6 were labeled as “Class 2 (early)”, and Gene7-Gene10 were labeled as “Class 3 (early)”.
[0049] For the intermediate and late differentiation states, class labels were assigned to genes in the same manner.
[0050] S13. Determine the dynamic correlation weight of the gene feature by calculating the conditional mutual information of each gene feature to the class label and combining it with the dynamic correlation weight algorithm; select the gene features of granulocytes in an iterative manner according to the dynamic correlation weight of the gene feature, and construct a standard gene expression data set based on the selection results.
[0051] As a preferred implementation, the dynamic correlation weight of the gene feature is determined by calculating the conditional mutual information of each gene feature to the class label and combining it with the dynamic correlation weight algorithm; according to the dynamic correlation weight of the gene feature, the gene feature of granulocytes is selected in an iterative manner, and a standard gene expression data set is constructed according to the selection result, including the following steps: S131, initialize the gene expression data set, and construct a candidate feature set based on all the gene features of neutrophils; It should be noted that by creating an empty gene expression dataset (usually represented as an empty set or empty matrix), this dataset will be used to store the gene features finally selected.
[0052] Based on all the gene features of neutrophils, a candidate feature set containing all these features is constructed. This set can be represented as a list or a matrix, where each row represents a gene feature and each column represents the expression value of the feature in different cells.
[0053] S132. For the gene features in the candidate feature set, calculate the conditional mutual information value between each gene feature and the class label, select the gene feature with the largest conditional mutual information value, remove it from the candidate feature set to the gene expression data set, and update the gene expression data set and the candidate feature set; It should be noted that for each gene feature in the candidate feature set, the conditional mutual information value between it and the class label (such as differentiation state) is calculated. The conditional mutual information can be calculated by statistical methods or information theory formulas. It measures the amount of additional information that the gene feature can provide given the class label. After calculating the conditional mutual information values of all candidate features, find the gene feature with the largest conditional mutual information value. This feature is considered to be the most relevant to the class label in the current candidate feature set, and therefore is most likely to be useful for classification or prediction tasks.
[0054] The selected gene features are removed from the candidate feature set and added to the gene expression dataset. The updated gene expression dataset contains one or more (depending on the number of iterations) gene features that are highly correlated with the class label. At the same time, the updated candidate feature set contains the remaining gene features that have not yet been selected.
[0055] Specifically, by calculating the conditional mutual information between each candidate gene feature and the class label, the gene with the highest mutual information is selected as the preliminary feature set. The initial number of gene features is set according to experimental requirements. For example, the top 10, 20 or 30 gene features that are most relevant to the class label are selected.
[0056] S133. For the updated gene expression dataset and candidate feature set, calculate the additional amount of new information of each gene feature in the candidate feature set on each gene feature in the gene expression dataset; (i.e., the influence of the feature-based on each feature-based in the gene expression dataset).
[0057] As a preferred implementation, for the updated gene expression data set and candidate feature set, the calculation formula for calculating the additional new information amount of each gene feature in the candidate feature set to each gene feature in the gene expression data set is: In the formula, extra ( X K , R ) represents each gene feature in the candidate feature set X KGene expression datasets R Each gene feature X R The additional amount of information added; R ∣ represents the gene expression dataset R The number of genetic features in Y represents the class label; I ( X R ; Y ) indicates gene characteristics X R and class labels Y The mutual information value between ; I ( X R ; Y ∣ X K ) indicates that in the gene feature X K Under the candidate conditions, gene features X R and class labels Y The conditional mutual information value between ; H ( X R ) indicates gene characteristics X R Information entropy of H ( Y ) represents the class label Y Information entropy.
[0058] S134, calculating the dynamic relevance weight of each gene feature in the current candidate feature set using a dynamic relevance weight algorithm according to the additional newly added information of each gene feature in the candidate feature set to each gene feature in the gene expression data set; As a preferred implementation, according to the additional additional information of each gene feature in the candidate feature set to each gene feature in the gene expression data set, the dynamic relevance weight algorithm is used to calculate the dynamic relevance weight of each gene feature in the current candidate feature set: In the formula, DW ( X K , R ) represents each gene feature in the current candidate feature set X K Dynamic relevance weight of extra ( X K , R ) represents each gene feature in the candidate feature set X KGene expression datasets R Each gene feature X R The amount of additional information added; I ( X K ; Y ∣ X R ) indicates gene characteristics X K and class labels Y The conditional mutual information value between .
[0059] S135. According to the dynamic correlation weight, select the gene feature with the largest weight value from the candidate feature set, add it to the gene expression data set, and remove it from the candidate feature set. Repeat steps S133-S135 until the number of gene features in the gene expression data set reaches a preset threshold, and obtain the final gene expression data set.
[0060] Specifically, the selection of initial gene features should not be too much, because this will lead to redundancy in the data set, affecting the efficiency of subsequent calculations and the complexity of the model. By calculating the preliminary mutual information between the candidate feature set and the class label, the first few most relevant gene features are selected as the initial set to avoid redundancy. Usually, 5-10 gene features as the initial selection is a reasonable range.
[0061] The number of iterations directly affects the convergence speed and calculation time of the model. The most relevant gene features are selected in each iteration, but the number of selected features in each iteration is limited, so there will be a balance point to ensure that the accuracy of the gene expression dataset is maximized when the model converges, while avoiding too many iterations. The use of dynamic relevance weights can help select the most valuable gene features more accurately in each iteration and reduce unnecessary iterations. The number of iterations can be limited by setting a threshold, for example, stopping the iteration when a certain number of features is reached or the amount of information added by the features is lower than a certain value.
[0062] Specifically, whether to continue iteration can also be determined by monitoring the information gain of the selected gene features in each iteration; if the information gain of the newly added gene features in a certain round for the data set is very small, it means that the existing gene features are sufficient to express the correlation of the class labels, and the iteration can be terminated at this time.
[0063] For example, you can set the following stop conditions: Threshold condition 1: When the information gain of the newly added gene features in each iteration is less than a certain threshold (such as 1% or 0.5%), the model is considered to have converged and the iteration can be terminated.
[0064] Threshold condition 2: When the information gain is less than the set threshold in multiple consecutive iterations (for example, the information gain is less than 0.5% for three consecutive rounds), the model is considered to be stable and the iteration is stopped.
[0065] Threshold condition 3: When the number of selected features reaches the preset maximum number, the iteration is terminated regardless of how the information gain changes.
[0066] S2. Perform feature mapping on the gene expression dataset based on deep learning method, and build multiple candidate causal models based on the mapping results; and use Bayesian information criterion to evaluate the scores of the candidate models, and select the candidate model with the highest score as the best causal model according to the scoring results; As a preferred embodiment, feature mapping is performed on a gene expression dataset based on a deep learning method, and multiple candidate causal models are constructed based on the mapping results; and the scores of the candidate models are evaluated using the Bayesian information criterion, and the candidate model with the highest score is selected as the best causal model according to the scoring results, including the following steps: S21. Use the autoencoder in the deep learning model to perform feature mapping on the gene features in the gene expression dataset, and build multiple alternative causal models based on the mapping results and in combination with the Granger causality analysis method; As a preferred implementation, the autoencoder in the deep learning model is used to perform feature mapping processing on the gene features in the gene expression data set, and based on the mapping results, multiple alternative causal models are constructed in combination with the Granger causality analysis method, including the following steps: S211, based on the pre-trained autoencoder, mapping the gene features in the gene expression dataset to a low-dimensional latent space through the encoder part in the autoencoder to obtain a low-dimensional feature representation; It should be noted that an autoencoder usually consists of an encoder and a decoder. The encoder is responsible for mapping the input data (here is the gene expression dataset) to a low-dimensional latent space (also called potential space or feature space), and the decoder is responsible for mapping the data in the low-dimensional latent space back to the original high-dimensional space.
[0067] In practical applications, the autoencoder is usually pre-trained, which means that it has learned how to encode and decode effectively on a certain dataset. Pre-training ensures that the autoencoder can capture the key features in the data and map them to a low-dimensional latent space. Using the encoder part of the pre-trained autoencoder, each gene feature in the gene expression dataset is mapped to the low-dimensional latent space. The low-dimensional feature representation is the representation of the gene expression data in the low-dimensional latent space. It is usually a vector in which each element represents the value of the original data on a certain low-dimensional feature.
[0068] Specifically, temporal features are an important component of gene expression datasets, especially in tasks such as cell development, disease progression, and drug response, where gene expression is often related to time.
[0069] For the encoder part, LSTM or TCN can be used as the encoder structure. These models can map each time step (each moment of gene expression) to a low-dimensional latent space and capture the dependencies between time steps; For the decoder part: The decoder attempts to reconstruct the original gene expression data from the low-dimensional latent space. In time series data, the decoder design can use LSTM, GRU or TCN to perform reverse mapping of time steps to restore the original gene expression sequence.
[0070] The training process of the autoencoder optimizes the model by minimizing the error between the input data and the reconstructed data. Specifically, it includes: Setting training objectives: The goal of the training process is to minimize the reconstruction error, that is, mapping the gene expression data to a low-dimensional space through the encoder and restoring it back to the original data space through the decoder. Commonly used reconstruction error metrics include mean squared error (MSE) or cross entropy loss (if the data is sparse or discrete).
[0071] Self-supervised learning: Autoencoders are essentially a self-supervised learning method that captures the latent characteristics of the data by letting the model learn on its own how to reconstruct the output from the input.
[0072] Optimization algorithm: Use gradient descent (such as Adam optimizer) to optimize the model, and update the model parameters through backpropagation during training.
[0073] Regularization: To prevent the model from overfitting, techniques such as L2 regularization, dropout, or early stopping can be used to limit the complexity of the model.
[0074] In addition, the training of autoencoders usually relies on a large amount of data, and the data set type is single-cell RNA-seq data: in gene expression data, single-cell RNA-seq is a more common data type, which provides the expression of each cell on different genes. If the data set has time series information (such as sampling the same cell population at different time points), then time-series autoencoders (such as LSTM or TCN) are more suitable.
[0075] S212. Use low-dimensional feature representation in low-dimensional latent space and Granger causality analysis to identify causal relationships between genes, and generate causal network through causal integration; As a preferred implementation, using low-dimensional feature representation in low-dimensional latent space, combined with Granger causality analysis method to identify the causal relationship between genes, and through causal integration, generating a causal relationship network includes the following steps: S2121. Pseudo-time sorting of neutrophils in low-dimensional latent space by trajectory inference method. It should be noted that in single-cell data, cell differentiation is usually a dynamic and continuous process. Pseudo-time sorting can infer the relative order of each cell in the differentiation process based on the changing trajectory of gene expression. Common pseudo-time sorting algorithms include: Monocle: Based on principal component analysis (PCA) or dimensionality reduction results.
[0076] PAGA: Combining cluster graphs with local connectivity to infer pseudotime.
[0077] Slingshot: Fitting curves to infer trajectories.
[0078] Specifically, the implementation steps of pseudo-time sorting include: Select a suitable pseudo-time sorting algorithm (such as Monocle, PAGA, or Slingshot) and determine the starting and branching points of the trajectory in the low-dimensional space according to the algorithm requirements.
[0079] According to the path of the trajectory or graph network, a pseudo time value of each cell is calculated, which represents the relative position of the cell in the differentiation process.
[0080] S2122. According to the pseudo-time ordering of the single-cell data, the low-dimensional features of each gene are reordered to obtain a low-dimensional latent space sequence; It should be noted that, for each cell, its feature representation in the low-dimensional latent space is extracted (such as obtained by the encoder part of the autoencoder), and the low-dimensional features of each cell are reordered according to the result of pseudo-time sorting, that is, the feature vector of each cell is arranged in the order of pseudo-time values.
[0081] For each gene, its low-dimensional features in all cells are combined in pseudo-time order to form a low-dimensional latent space sequence of the gene, which reflects the expression changes of the gene during cell differentiation.
[0082] S2123. Perform Granger causality analysis on low-dimensional latent space sequences to identify causal relationships between low-dimensional features; It should be noted that Granger Causality Analysis is a statistical method based on time series data, which is used to determine whether a time series (cause variable) has a predictive effect on another time series (result variable), that is, whether there is a causal relationship. In biology, this method can be used to infer the regulatory relationship between genes.
[0083] Specifically, Granger causality analysis of low-dimensional latent space sequences includes the following steps: Determine the parameters in the Granger causality analysis model, such as the lag order (i.e., how many past time points are considered to predict the current time point). The choice of lag order can be determined by model selection criteria (such as AIC, BIC, etc.); Using the linear Granger causality model, Granger causality analysis is performed on low-dimensional latent space sequences. This usually involves calculating the predictive power of the cause variable on the outcome variable and evaluating whether this predictive power is significantly higher than the random level; According to the results of Granger causality analysis, the causal relationship between low-dimensional features is identified. If the cause variable has a significant predictive effect on the result variable, it is considered that there is a causal relationship between them. Verify the identified causal relationship, for example, by using cross-validation, Bootstrap and other methods to evaluate the stability and reliability of the results; combine biological knowledge and literature reports to explain the biological significance of the identified causal relationship.
[0084] S2124. Through the decoder of the autoencoder, the causal relationship in the low-dimensional latent space is mapped back to the high-dimensional gene space to generate a gene causal network.
[0085] It should be noted that in the low-dimensional latent space, the causal relationships between genes are identified through methods such as Granger causality analysis. These causal relationships are represented in the form of interactions between low-dimensional features; The identified low-dimensional causal relationships are used as input and mapped back to the high-dimensional gene space through the decoder part of the autoencoder. The decoder converts the low-dimensional causal relationships into causal relationships between high-dimensional gene expression data based on the learned data distribution and feature mapping relationship.
[0086] In the high-dimensional gene space, a gene causal network is constructed based on the causal relationship output by the decoder. The nodes in the network represent genes, and the edges represent the causal relationship between genes. Graph theory tools or visualization software can be used to display and analyze this network.
[0087] S213. Construct multiple alternative causal models based on the causal network.
[0088] Specifically, constructing multiple alternative causal models based on the causal network is an important step in the research, which helps to deeply understand and analyze the interaction mechanism between variables. Constructing multiple alternative causal models based on the causal network includes the following steps: 1. Clarify goals and assumptions Goal setting: First, clarify the goal of building a causal model, such as explaining a specific biological phenomenon.
[0089] Hypothesis formulation: Based on the causal network and existing knowledge, multiple hypotheses are proposed about the interactions between variables. These hypotheses will serve as the basis for constructing alternative causal models.
[0090] 2. Choosing a Modeling Method Statistical methods: such as regression analysis, path analysis, structural equation modeling, etc. These methods can help us quantify the causal relationship between variables and evaluate the goodness of fit of the model.
[0091] Machine learning methods: such as Bayesian networks, causal graph models, etc. These methods can handle more complex causal networks and automatically learn the dependencies between variables.
[0092] Expert system: Combines the knowledge and experience of domain experts to build a rule-based causal model. This approach is suitable for situations where domain knowledge is abundant but data is insufficient.
[0093] 3. Constructing Alternative Causal Models Model design: Based on the proposed hypothesis and the selected modeling method, design multiple alternative causal models. Each model should contain a set of variables and the causal relationships between them.
[0094] S22. Calculating the likelihood of the alternative causal model according to the Bayesian information criterion, and scoring the alternative causal model according to the likelihood of the alternative causal model to obtain a scoring result; As a preferred implementation, the likelihood of the candidate causal model is calculated according to the Bayesian information criterion, and the candidate causal model is scored according to the likelihood of the candidate causal model. Obtaining the scoring result includes the following steps: S221. For each candidate causal model, calculating a maximum likelihood estimate of the candidate causal model under the gene expression data set by maximizing the likelihood function; It should be noted that for each alternative causal model Mi , its likelihood function L ( Mi ) indicates that in the model Mi Under the observation, the current gene expression dataset D Probability: L ( Mi )= P (D ∣ Mi ); Maximize the likelihood function through optimization algorithms (such as gradient descent, Newton's method, etc.) L ( Mi ), the calculation formula is: G =argmax θ L ( Mi ) G Representation Model Mi The parameter estimation under the maximum likelihood is θ Representation Model Mi A collection of parameters.
[0095] S222. Calculate the total number of free parameters based on the model structure, and calculate the score of each alternative causal model based on the Bayesian Information Criterion formula.
[0096] It should be noted that the Bayesian Information Criterion (BIC) is a model selection criterion used to evaluate the trade-off between model complexity and goodness of fit: In the formula, L max( Mi ) represents the model Mi The maximum likelihood estimate of k Representation Model Mi The total number of free parameters, n Represents the number of samples (usually the number of cells in a gene expression dataset).
[0097] For each model Mi , the number of all free parameters in the statistical model, including: coefficients of linear regression models, mixing ratios in mixed models, edge weights in network models, etc.
[0098] By calculating the maximum likelihood estimate of each candidate causal model and combining it with the total number of free parameters of the model, the Bayesian Information Criterion (BIC) is used to score the candidate models. The smaller the BIC value, the lower the complexity of the model while fitting the data, and therefore the greater the possibility of being selected as the best causal model. This process provides a quantitative standard for model selection, which helps to screen the most appropriate causal model from multiple candidate models.
[0099] S23. Based on the scoring results of each candidate causal model, select the model with the highest score as the best causal model, and analyze the causal structure through compression representation.
[0100] It should be noted that the goal of compressed representation is to map the complex high-dimensional causal network into a more concise form, retaining the core causal relationship for easy interpretation. The specific steps include: The causal edges with higher weights or significant significance are screened out from the best model to generate a simplified causal network.
[0101] The random forest method was used to evaluate the causal importance of each feature (gene) and further screen the key genes.
[0102] Combined with dimensionality reduction algorithms (such as PCA and t-SNE), the causal network is visualized to display the key causal paths.
[0103] S3. Constructing a differentiation trajectory network based on the causal relationship network in the best model, and inferring the differentiation trajectory of neutrophils based on the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory; As a preferred embodiment, a differentiation trajectory network is constructed according to the causal relationship network in the optimal model, and the differentiation trajectory of neutrophils is inferred according to the differentiation trajectory network. The prediction result of the neutrophil differentiation trajectory is obtained, which includes the following steps: S31. Integrate gene causal relationships according to the causal structure in the optimal causal model and generate a causal relationship network; It should be noted that after completing the evaluation and screening of the candidate causal models, an optimal causal model is obtained, and all gene causal relationships are extracted from the optimal causal model. These causal relationships are usually represented in the form of directed edges, where the starting point represents the cause gene and the end point represents the result gene. In the causal network, each gene is represented as a node, and the attributes of the node can include the name of the gene, expression level, regulatory effect, etc. According to the extracted causal relationship, directed edges are added between the nodes, and the attributes of the edge can include the strength, direction, confidence, etc. of the causal relationship. These attributes help us to have a deeper understanding of the regulatory relationship between genes.
[0104] S32, incorporate pseudo-time information into the causal network to construct a differentiation trajectory network containing dynamic regulation and differentiation information; It should be noted that in biological research, especially in single-cell transcriptome analysis, since it is impossible to directly obtain the actual time information of cell differentiation, researchers often use algorithms (such as Monocle, Slingshot, etc.) to infer the differentiation order of cells. This inferred time information is called pseudo-time. Pseudo-time information can help understand the dynamic changes of cells during differentiation.
[0105] In biological networks, dynamic regulation usually refers to changes in gene expression, protein interactions, etc. over time. These changes are crucial for understanding processes such as cell differentiation, development, and disease occurrence.
[0106] Cell differentiation is a fundamental process in biology that involves the transformation of cells from one type to another. During differentiation, cells undergo a series of changes in gene expression that can be used to construct a cell's differentiation trajectory.
[0107] To construct a differentiation trajectory network containing dynamic regulation and differentiation information, the following steps can be followed: Data collection and processing: Collect transcriptome data of cells at different differentiation stages and perform preprocessing (such as denoising, normalization, etc.).
[0108] Pseudo-time inference: Use appropriate algorithms (such as Monocle, Slingshot, etc.) to perform pseudo-time inference on transcriptome data to obtain the order of cell differentiation.
[0109] Causal network construction: Based on transcriptome data and pseudo-time information, causal inference methods (such as Granger causality test, dynamic Bayesian network, etc.) are used to construct a causal network.
[0110] Dynamic regulation analysis: In a causal network, analyze changes in gene expression, protein interactions, etc. over time to identify key dynamic regulatory events.
[0111] Differentiation trajectory network construction: Combine pseudo-time information and the results of dynamic regulation analysis to construct a differentiation trajectory network containing dynamic regulation and differentiation information. This network should be able to clearly show the dynamic changes and regulatory mechanisms of cells during differentiation.
[0112] S33. Use the differentiation trajectory network to perform dynamic simulation, infer the differentiation path and key nodes, and output the prediction results of the differentiation trajectory.
[0113] Specifically, the differentiation trajectory network is a network model built based on single-cell transcriptome data to describe the dynamic changes and regulatory relationships of cells during differentiation. The network usually contains nodes (representing cells or cell types) and edges (representing differentiation relationships or regulatory relationships between cells).
[0114] Based on the differentiation trajectory network, dynamic network analysis methods are used to simulate the dynamic process of cell differentiation. Through dynamic simulation, the possible paths of cell differentiation from one type to another are inferred, and network analysis methods (such as centrality analysis, community detection, etc.) are used to identify key nodes in the differentiation trajectory network.
[0115] The prediction results of differentiation trajectories are usually displayed in graphical form, such as differentiation trajectory trees, differentiation trajectory diagrams, etc. These graphics can clearly show the cell differentiation path, key nodes, and the regulatory relationship between each node.
[0116] like Figure 2 As shown, according to another embodiment of the present invention, a neutrophil differentiation trajectory inference system is provided, the neutrophil differentiation trajectory inference system comprising: a gene expression data set construction module 1, an optimal causal model construction module 2 and a differentiation trajectory inference module 3, and the gene expression data set construction module 1, the optimal causal model construction module 2 and the differentiation trajectory inference module 3 are sequentially connected; Gene expression data set construction module 1 is used to obtain gene expression data of neutrophils based on single-cell transcriptome sequencing technology, and perform feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation to obtain a gene expression data set; The best causal model construction module 2 is used to perform feature mapping processing on the gene expression data set based on the deep learning method, and to construct multiple candidate causal models based on the mapping results; and to evaluate the scores of the candidate models using the Bayesian information criterion, and to select the candidate model with the highest score as the best causal model according to the scoring results; The differentiation trajectory inference module 3 is used to construct a differentiation trajectory network according to the causal relationship network in the optimal model, and to infer the differentiation trajectory of neutrophils according to the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory.
[0117] In summary, with the help of the above-mentioned technical scheme of the present invention, the present invention obtains the gene expression data of neutrophils through single-cell transcriptome sequencing technology, provides a high-resolution and high-precision cellular level data basis for analysis, and helps to have a deeper understanding of the differentiation process of neutrophils. The gene expression data is processed by the feature selection algorithm of dynamic correlation, which can screen out key genes closely related to neutrophil differentiation, reduce redundant information, and improve the accuracy and efficiency of subsequent analysis. The differentiation trajectory network is constructed based on the causal relationship network in the optimal model, and the differentiation trajectory is inferred, which can obtain the prediction results of the neutrophil differentiation trajectory, which helps to reveal the dynamic process and regulatory mechanism of neutrophil differentiation. The present invention uses single-cell transcriptome sequencing technology to obtain gene expression data of neutrophils at the single cell level, pre-processes the gene expression matrix, removes noise, corrects deviations, and improves data quality. Combined with the differentiation state of the cell, a class label is generated for each gene feature through a clustering algorithm, which helps to classify genes with similar expression patterns into one class. By calculating the conditional mutual information of each gene feature to the class label and combining it with a dynamic correlation weight algorithm, the correlation between each gene feature and the differentiation state of the neutrophil can be accurately evaluated. The gene features of the neutrophils are selected in an iterative manner, and a standard gene expression data set is constructed according to the selection results. The gene expression data set can be gradually optimized, the gene features most relevant to the differentiation of neutrophils are retained, and redundant and noise information is removed. The present invention uses an autoencoder to perform feature mapping on gene expression data, maps high-dimensional data to a low-dimensional latent space, effectively reduces the dimension of the data while retaining key information, and combines the Granger causality analysis method to identify the causal relationship between genes in a low-dimensional latent space, avoiding the complexity and noise interference of high-dimensional data, and improving the accuracy of causal identification. The Bayesian information criterion is used to score the alternative causal models, and the fit and complexity of the model are comprehensively considered to ensure that the selected model has a good fitting effect and is not too complex. The differentiation trajectory network is used for dynamic simulation, which can intuitively display the path and key nodes of cell differentiation, providing a powerful tool for understanding the mechanism of cell differentiation.
[0118] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0119] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for inferring neutrophil differentiation trajectory, characterized in that: include: S1. Based on single-cell transcriptome sequencing technology, the gene expression data of neutrophils is obtained, and the gene expression data of neutrophils is processed by feature selection algorithm of dynamic correlation to obtain the gene expression data set; S2. Perform feature mapping on the gene expression dataset based on deep learning method, and build multiple candidate causal models based on the mapping results; and use Bayesian information criterion to evaluate the scores of the candidate models, and select the candidate model with the highest score as the best causal model according to the scoring results; S3. Constructing a differentiation trajectory network based on the causal relationship network in the best model, and inferring the differentiation trajectory of neutrophils based on the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory; The S1 includes: S11. Collect gene expression data of neutrophils by single-cell transcriptome sequencing technology, construct a gene expression matrix, and preprocess the gene expression matrix; obtain gene characteristics of neutrophils based on the preprocessing results; S12, based on the gene features of neutrophils and the differentiation state of the cells, a clustering algorithm is used to generate a class label for each gene feature; S13. Determine the dynamic correlation weight of the gene feature by calculating the conditional mutual information of each gene feature to the class label and combining it with the dynamic correlation weight algorithm; select the gene features of granulocytes in an iterative manner according to the dynamic correlation weight of the gene feature, and construct a standard gene expression data set based on the selection results.
2. A neutrophil differentiation trajectory inference method according to claim 1, characterized in that: The method comprises the following steps: calculating the conditional mutual information of each gene feature to the class label and combining the dynamic correlation weight algorithm to determine the dynamic correlation weight of the gene feature; selecting the gene features of granulocytes in an iterative manner according to the dynamic correlation weight of the gene feature, and constructing a standard gene expression data set according to the selection results: S131, initialize the gene expression data set, and construct a candidate feature set based on all the gene features of neutrophils; S132. For the gene features in the candidate feature set, calculate the conditional mutual information value between each gene feature and the class label, select the gene feature with the largest conditional mutual information value, remove it from the candidate feature set to the gene expression data set, and update the gene expression data set and the candidate feature set; S133, for the updated gene expression dataset and candidate feature set, calculating the additional newly added information of each gene feature in the candidate feature set to each gene feature in the gene expression dataset; S134, calculating the dynamic relevance weight of each gene feature in the current candidate feature set using a dynamic relevance weight algorithm according to the additional newly added information of each gene feature in the candidate feature set to each gene feature in the gene expression data set; S135. According to the dynamic correlation weight, select the gene feature with the largest weight value from the candidate feature set, add it to the gene expression data set, and remove it from the candidate feature set. Repeat steps S133-S135 until the number of gene features in the gene expression data set reaches a preset threshold, and obtain the final gene expression data set.
3. A neutrophil differentiation trajectory inference method according to claim 2, characterized in that: For the updated gene expression dataset and candidate feature set, the calculation formula for calculating the additional new information amount of each gene feature in the candidate feature set to each gene feature in the gene expression dataset is: In the formula, extra ( X K , R ) represents each gene feature in the candidate feature set X K Gene expression datasets R Each gene feature X R The amount of additional information added; ∣ R ∣ represents the gene expression dataset R The number of genetic features in Y represents the class label; I ( X R ; Y ) indicates gene characteristics X R and class labels Y The mutual information value between ; I ( X R ; Y ∣ X K ) indicates that in the gene feature X K Under the candidate conditions, gene features X R and class labels Y The conditional mutual information value between ; H ( X R ) indicates gene characteristics X R Information entropy of H ( Y ) represents the class label Y Information entropy.
4. A method for inferring neutrophil differentiation trajectory according to claim 3, characterized in that: The calculation formula for calculating the dynamic relevance weight of each gene feature in the current candidate feature set using the dynamic relevance weight algorithm based on the additional additional information of each gene feature in the candidate feature set to each gene feature in the gene expression data set is: In the formula, DW ( X K , R ) represents each gene feature in the current candidate feature set X K Dynamic relevance weight of extra ( X K , R ) represents each gene feature in the candidate feature set X K Gene expression datasets R Each gene feature X R The amount of additional information added; I ( X K ; Y ∣ X R ) indicates gene characteristics X K and class labels Y The conditional mutual information value between .
5. A method for inferring neutrophil differentiation trajectory according to claim 1, characterized in that: The method of performing feature mapping on the gene expression data set based on the deep learning method, and constructing multiple candidate causal models based on the mapping results; and evaluating the scores of the candidate models using the Bayesian information criterion, and selecting the candidate model with the highest score as the best causal model according to the scoring results includes the following steps: S21. Use the autoencoder in the deep learning model to perform feature mapping on the gene features in the gene expression dataset, and build multiple alternative causal models based on the mapping results and in combination with the Granger causality analysis method; S22. Calculating the likelihood of the alternative causal model according to the Bayesian information criterion, and scoring the alternative causal model according to the likelihood of the alternative causal model to obtain a scoring result; S23. Based on the scoring results of each candidate causal model, select the model with the highest score as the best causal model, and analyze the causal structure through compression representation.
6. A method for inferring neutrophil differentiation trajectory according to claim 5, characterized in that: The method of using the autoencoder in the deep learning model to perform feature mapping processing on the gene features in the gene expression data set, and constructing multiple alternative causal models based on the mapping results and in combination with the Granger causality analysis method includes the following steps: S211, based on the pre-trained autoencoder, mapping the gene features in the gene expression dataset to a low-dimensional latent space through the encoder part in the autoencoder to obtain a low-dimensional feature representation; S212. Use low-dimensional feature representation in low-dimensional latent space and Granger causality analysis to identify causal relationships between genes, and generate causal network through causal integration; S213. Construct multiple alternative causal models based on the causal network.
7. A method for inferring neutrophil differentiation trajectory according to claim 6, characterized in that: The method of using low-dimensional feature representation in low-dimensional latent space, combining with Granger causality analysis method to identify causal relationships between genes, and generating a causal relationship network through causal relationship integration includes the following steps: S2121. Pseudo-time sorting of neutrophils in low-dimensional latent space by trajectory inference method. S2122. According to the pseudo-time ordering of the single-cell data, the low-dimensional features of each gene are reordered to obtain a low-dimensional latent space sequence; S2123. Perform Granger causality analysis on low-dimensional latent space sequences to identify causal relationships between low-dimensional features; S2124. Through the decoder of the autoencoder, the causal relationship in the low-dimensional latent space is mapped back to the high-dimensional gene space to generate a gene causal network.
8. A method for inferring neutrophil differentiation trajectory according to claim 6, characterized in that: The method of calculating the likelihood of the candidate causal model according to the Bayesian information criterion, and scoring the candidate causal model according to the likelihood of the candidate causal model to obtain the scoring result comprises the following steps: S221. For each candidate causal model, calculating a maximum likelihood estimate of the candidate causal model under the gene expression data set by maximizing the likelihood function; S222. Calculate the total number of free parameters based on the model structure, and calculate the score of each alternative causal model based on the Bayesian Information Criterion formula.
9. A method for inferring neutrophil differentiation trajectory according to claim 1, characterized in that: The step of constructing a differentiation trajectory network according to the causal relationship network in the best model, and inferring the differentiation trajectory of neutrophils according to the differentiation trajectory network to obtain a prediction result of the neutrophil differentiation trajectory comprises the following steps: S31. Integrate gene causal relationships according to the causal structure in the optimal causal model and generate a causal relationship network; S32, incorporate pseudo-time information into the causal network to construct a differentiation trajectory network containing dynamic regulation and differentiation information; S33. Use the differentiation trajectory network to perform dynamic simulation, infer the differentiation path and key nodes, and output the prediction results of the differentiation trajectory.
10. A neutrophil differentiation trajectory inference system, used to implement the neutrophil differentiation trajectory inference method according to any one of claims 1 to 9, characterized in that: The neutrophil differentiation trajectory inference system comprises: a gene expression data set construction module, an optimal causal model construction module and a differentiation trajectory inference module, and the gene expression data set construction module, the optimal causal model construction module and the differentiation trajectory inference module are connected in sequence; The gene expression data set construction module is used to obtain the gene expression data of neutrophils based on the single-cell transcriptome sequencing technology, and perform feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation to obtain a gene expression data set; The optimal causal model construction module is used to perform feature mapping processing on the gene expression data set based on the deep learning method, and to construct multiple candidate causal models based on the mapping results; and to evaluate the scores of the candidate models using the Bayesian information criterion, and to select the candidate model with the highest score as the optimal causal model according to the scoring results; The differentiation trajectory inference module is used to construct a differentiation trajectory network based on the causal relationship network in the best model, and to infer the differentiation trajectory of neutrophils based on the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory.
Citation Information
Patent Citations
Method and system for constructing dynamic gene regulation network and computer equipment
CN119317963A
Cell differentiation trajectory inference method based on deep hyperbolic manifold embedding
CN119479828A
Single-cell regulatory multi-omics with deep mitochondrial mutation profiling (redeem)
WO2024238350A1
Cited By
Causal relationship prediction method and device for single cell data, equipment and medium
CN120877835A