A method and system for inferring neutrophil differentiation trajectories
Through the method of single-cell transcriptome sequencing and deep learning combined with Bayesian information criterion, a neutrophil differentiation trajectory network was constructed, which solved the problem of limited prediction accuracy caused by causal neglect in traditional methods, and achieved efficient and accurate disclosure of differentiation trajectory prediction and regulatory mechanism.
Patent Information
- Application Number
- CN202510460330.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing traditional differentiation trajectory inference method relies on pseudo-time analysis and ignores the dynamic regulation of causal relationships, resulting in limited prediction accuracy of cell differentiation paths and key nodes, making it difficult to comprehensively and accurately depict the dynamic picture of cell differentiation and the regulatory logic behind it.
Gene expression data of neutrophils were obtained through single-cell transcriptome sequencing technology, and the characteristic selection algorithm of dynamic correlation was used to screen key genes. The causal model was constructed based on deep learning and Bayesian information criterion, and a differentiation trajectory network was constructed for differentiation trajectory inference.
It improves the accuracy and efficiency of differentiation trajectory prediction, reveals the dynamic process and regulatory mechanism of neutrophil differentiation, provides a high-resolution cellular-level data basis, reduces redundant information, and reduces noise interference.
Smart Images

Figure CN119993281B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cell differentiation trajectory analysis, and in particular, to a neutrophil differentiation trajectory inference method and system. Background Art
[0002] Cell differentiation, a core process in biology, describes how cells from the same source gradually evolve into cell populations with different morphologies, structures, and functions. It not only causes cells to differ in the spatial dimension, but also marks a significant change in the temporal dimension from their original state. The essence of differentiation lies in the precise regulation of the genome in time and space, that is, the selective expression of specific genes, some genes are activated, while others are silenced. This process ultimately leads to the synthesis of iconic proteins. It is worth noting that cell differentiation is often a one-way and irreversible journey.
[0003] For bioinformatics researchers, in-depth exploration of cell differentiation trajectories is a valuable way to tap into the value of bioinformatics, reveal intercellular heterogeneity, and unlock new biological insights (such as discovering rare cell types). It is widely used to analyze the single-cell gene expression dynamics of various cellular processes including differentiation, proliferation, and even carcinogenic transformation, providing a strong impetus for the innovation of biological research and medical technology.
[0004] However, most of the current traditional differentiation trajectory inference methods rely only on pseudo-time analysis, ignoring the key factor of dynamic regulation of causality, resulting in limited prediction accuracy of differentiation paths and key nodes, making it difficult to fully and accurately depict the dynamic picture of cell differentiation and the regulatory logic behind it.
[0005] Currently, no effective solution has been proposed for the problems in the related technologies. Summary of the invention
[0006] In view of this, the present invention provides a neutrophil differentiation trajectory inference method and system to solve the above-mentioned problems.
[0007] In order to solve the above problems, the specific technical solutions adopted by the present invention are as follows:
[0008] According to one aspect of the present invention, a method for inferring neutrophil differentiation trajectories is provided, and the method for inferring neutrophil differentiation trajectories comprises the following steps:
[0009] S1. Based on single-cell transcriptome sequencing technology, the gene expression data of neutrophils is obtained, and the gene expression data of neutrophils is processed by feature selection algorithm of dynamic correlation to obtain the gene expression data set;
[0010] S2. Perform feature mapping processing on the gene expression dataset based on deep learning methods, and construct multiple alternative causal models based on the mapping results; and use the Bayesian information criterion to evaluate the scores of the alternative models, and select the alternative model with the highest score as the best causal model according to the score results;
[0011] S3. Construct a differentiation trajectory network according to the causal relationship network in the best model, and infer the differentiation trajectory of neutrophils according to the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory;
[0012] Preferably, based on single-cell transcriptome sequencing technology, obtain the gene expression data of neutrophils, and perform feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation degree, and the steps for obtaining the gene expression dataset are as follows:
[0013] S11. Collect the gene expression data of neutrophils through single-cell transcriptome sequencing technology, construct a gene expression matrix, and preprocess the gene expression matrix; obtain the gene characteristics of neutrophils according to the preprocessing results;
[0014] S12. According to the gene characteristics of neutrophils, combined with the differentiation state of cells, generate class labels for each gene characteristic through a clustering algorithm;
[0015] S13. Determine the dynamic correlation degree weights of gene characteristics by calculating the conditional mutual information of each gene characteristic with respect to the class label and combining with the dynamic correlation degree weight algorithm; select the gene characteristics of neutrophils in an iterative manner according to the dynamic correlation degree weights of the gene characteristics, and construct a standard gene expression dataset according to the selection results.
[0016] Preferably, determine the dynamic correlation degree weights of gene characteristics by calculating the conditional mutual information of each gene characteristic with respect to the class label and combining with the dynamic correlation degree weight algorithm; select the gene characteristics of neutrophils in an iterative manner according to the dynamic correlation degree weights of the gene characteristics, and the steps for constructing a standard gene expression dataset are as follows:
[0017] S131. Initialize the gene expression dataset, and construct a candidate feature set according to all the gene characteristics of neutrophils;
[0018] S132. For the gene characteristics in the candidate feature set, calculate the conditional mutual information value of each gene characteristic with the class label, and select the gene characteristic with the largest conditional mutual information value, remove it from the candidate feature set to the gene expression dataset, and update the gene expression dataset and the candidate feature set;
[0019] S133. For the updated gene expression dataset and candidate feature set, calculate the additional new information content of each gene feature in the candidate feature set with respect to each gene feature in the gene expression dataset;
[0020] S134. According to the additional new information content of each gene feature in the candidate feature set with respect to each gene feature in the gene expression dataset, use the dynamic correlation weight algorithm to calculate the dynamic correlation weight of each gene feature in the current candidate feature set;
[0021] S135. According to the dynamic correlation weight, select the gene feature with the largest weight value from the candidate feature set, add it to the gene expression dataset, and remove it from the candidate feature set. Repeat steps S133 - S135 until the number of gene features in the gene expression dataset reaches a preset threshold, and obtain the final gene expression dataset.
[0022] Preferably, for the updated gene expression dataset and candidate feature set, the calculation formula for the additional new information content of each gene feature in the candidate feature set with respect to each gene feature in the gene expression dataset is:
[0023]
[0024] In the formula, extra ( X K , R ) represents the additional new information content of each gene feature X K in the candidate feature set with respect to each gene feature R in the gene expression dataset X R ;
[0025] ∣ R ∣ represents the number of gene features in the gene expression dataset R ;
[0026] Y represents the class label;
[0027] I ( X R ; Y ) represents the mutual information value between the gene feature X R and the class label Y ;
[0028] I ( X R ; Y ∣ X K ) represents the conditional entropy of the class label given the gene featureX K Under the candidate conditions, the gene features X R and the class labels Y The conditional mutual information value between;
[0029] H ( X R ) represents the information entropy of the gene feature X R ;
[0030] H ( Y ) represents the information entropy of the class label Y ;
[0031] Preferably, according to the additional new information amount of each gene feature in the candidate feature set to each gene feature in the gene expression dataset, the calculation formula for the dynamic correlation weight of each gene feature in the current candidate feature set using the dynamic correlation weight algorithm is:
[0032]
[0033] In the formula, DW ( X K , R ) represents the dynamic correlation weight of each gene feature X K in the current candidate feature set;
[0034] extra ( X K , R ) represents the additional new information amount of each gene feature X K in the candidate feature set to each gene feature R in the gene expression dataset X R ;
[0035] I ( X K ; Y ∣ X R ) represents the conditional mutual information value between the gene feature X K and the class label Y ;
[0036] Preferably, perform feature mapping processing on the gene expression dataset based on the deep learning method, and construct multiple alternative causal models based on the mapping results; and use the Bayesian information criterion to evaluate the scores of the alternative models, and select the alternative model with the highest score as the best causal model, including the following steps:
[0037] S21. Use the autoencoder in the deep learning model to perform feature mapping processing on the gene features in the gene expression dataset. Based on the mapping results, combine with the Granger causality analysis method to construct multiple alternative causal models;
[0038] S22. Calculate the likelihood of the alternative causal models according to the Bayesian information criterion, and score the alternative causal models according to the likelihood of the alternative causal models to obtain the scoring results;
[0039] S23. According to the scoring results of each alternative causal model, select the model with the highest score as the best causal model, and analyze the causal structure through compressed representation.
[0040] Preferably, use the autoencoder in the deep learning model to perform feature mapping processing on the gene features in the gene expression dataset. Based on the mapping results, combine with the Granger causality analysis method to construct multiple alternative causal models, including the following steps:
[0041] S211. Based on the pre-trained autoencoder, map the gene features in the gene expression dataset to a low-dimensional hidden space through the encoder part in the autoencoder to obtain a low-dimensional feature representation;
[0042] S212. Use the low-dimensional feature representation in the low-dimensional hidden space, combine with the Granger causality analysis method to identify the causal relationships between genes, and generate a causal relationship network through causal relationship integration;
[0043] S213. Construct multiple alternative causal models according to the causal relationship network.
[0044] Preferably, use the low-dimensional feature representation in the low-dimensional hidden space, combine with the Granger causality analysis method to identify the causal relationships between genes, and generate a causal relationship network through causal relationship integration, including the following steps:
[0045] S2121. In the low-dimensional hidden space, perform pseudo-time sorting on neutrophils through the trajectory inference method;
[0046] S2122. According to the pseudo-time sorting of the single-cell data, reorder the low-dimensional features of each gene to obtain a low-dimensional hidden space sequence;
[0047] S2123. Perform Granger causality analysis on the low-dimensional hidden space sequence to identify the causal relationships between the low-dimensional features;
[0048] S2124. Through the decoder of the autoencoder, map the causal relationships in the low-dimensional latent space back to the high-dimensional gene space to generate a gene causal network.
[0049] Preferably, calculating the likelihood of alternative causal models according to the Bayesian information criterion, and scoring the alternative causal models according to the likelihood of the alternative causal models, the steps for obtaining the scoring result include:
[0050] S221. For each alternative causal model, calculate the maximum likelihood estimate of the alternative causal model under the gene expression dataset by maximizing the likelihood function;
[0051] S222. Calculate the total number of free parameters according to the model structure, and calculate the score of each alternative causal model according to the Bayesian information criterion formula.
[0052] Preferably, constructing a differentiation trajectory network according to the causal relationship network in the best model, and inferring the differentiation trajectory of neutrophils according to the differentiation trajectory network, the steps for obtaining the prediction result of the neutrophil differentiation trajectory include:
[0053] S31. Integrate the gene causal relationships according to the causal structure in the best causal model, and generate a causal relationship network;
[0054] S32. Incorporate the pseudotime information into the causal relationship network to construct a differentiation trajectory network containing dynamic regulation and differentiation information;
[0055] S33. Use the differentiation trajectory network for dynamic simulation, infer the differentiation path and key nodes, and output the prediction result of the differentiation trajectory.
[0056] According to one aspect of the present invention, there is provided a system for inferring neutrophil differentiation trajectory, the system for inferring neutrophil differentiation trajectory includes: a gene expression dataset construction module, an optimal causal model construction module, and a differentiation trajectory inference module, and the gene expression dataset construction module, the optimal causal model construction module, and the differentiation trajectory inference module are sequentially connected;
[0057] The gene expression dataset construction module is configured to obtain the gene expression data of neutrophils based on single-cell transcriptome sequencing technology, and perform feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation degree to obtain a gene expression dataset;
[0058] The optimal causal model construction module is configured to perform feature mapping processing on the gene expression dataset based on deep learning method, and construct multiple alternative causal models based on the mapping result; and evaluate the scores of the alternative models using the Bayesian information criterion, and select the alternative model with the highest score as the optimal causal model according to the scoring result;
[0059] A differentiation trajectory inference module, which is used to construct a differentiation trajectory network according to the causal relationship network in the best model, and infer the differentiation trajectory of neutrophils according to the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory.
[0060] The beneficial effects of the present invention are as follows:
[0061] 1. The present invention obtains the gene expression data of neutrophils through single-cell transcriptome sequencing technology, providing a high-resolution and high-precision cell-level data basis for analysis, which helps to understand the differentiation process of neutrophils more deeply. The feature selection algorithm with dynamic correlation is used to process the gene expression data, which can screen out the key genes closely related to neutrophil differentiation, reduce redundant information, and improve the accuracy and efficiency of subsequent analysis. Based on the causal relationship network in the best model, a differentiation trajectory network is constructed and differentiation trajectory inference is carried out, and the prediction result of the neutrophil differentiation trajectory can be obtained, which helps to reveal the dynamic process and regulatory mechanism of neutrophil differentiation.
[0062] 2. Through single-cell transcriptome sequencing technology, the present invention can obtain the gene expression data of neutrophils at the single-cell level. Pretreating the gene expression matrix can remove noise and correct biases, improving the quality of the data. Combining with the differentiation state of cells, class labels are generated for each gene feature through a clustering algorithm, which helps to group genes with similar expression patterns into one category. By calculating the conditional mutual information of each gene feature with respect to the class label and combining with the dynamic correlation weight algorithm, the correlation between each gene feature and the neutrophil differentiation state can be accurately evaluated. By iteratively selecting the gene features of neutrophils and constructing a standard gene expression data set according to the selection results, the gene expression data set can be gradually optimized, retaining the gene features most relevant to neutrophil differentiation and removing redundant and noise information.
[0063] 3. The present invention uses an autoencoder to perform feature mapping on the gene expression data, mapping high-dimensional data to a low-dimensional latent space, effectively reducing the dimension of the data while retaining key information. Combining with the Granger causality analysis method, the causal relationship between genes is identified in the low-dimensional latent space, avoiding the complexity and noise interference of high-dimensional data and improving the accuracy of causal identification. Using the Bayesian information criterion to score alternative causal models, comprehensively considering the goodness of fit and complexity of the model, ensures that the selected model has both a good fitting effect and is not overly complex. Using the differentiation trajectory network for dynamic simulation can intuitively display the path and key nodes of cell differentiation, providing a powerful tool for understanding the mechanism of cell differentiation. Brief Description of the Drawings
[0064] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0065] Figure 1 is a flowchart of a method for inferring the neutrophil differentiation trajectory according to an embodiment of the present invention;
[0066] Figure 2 is a schematic block diagram of a system for inferring the neutrophil differentiation trajectory according to an embodiment of the present invention.
[0067] In the figure:
[0068] 1. Gene expression dataset construction module; 2. Optimal causal model construction module; 3. Differentiation trajectory inference module. Detailed implementation manners
[0069] In order to enable those skilled in the art of this technology to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts should fall within the protection scope of this application.
[0070] According to an embodiment of the present invention, a method and system for inferring the neutrophil differentiation trajectory are provided.
[0071] Now, the present invention will be further described in conjunction with the drawings and specific implementation manners. As Figure 1 shown, according to an embodiment of the present invention, a method for inferring the neutrophil differentiation trajectory is provided. The method for inferring the neutrophil differentiation trajectory includes the following steps:
[0072] S1. Based on single-cell transcriptome sequencing technology, obtain the gene expression data of neutrophils, and perform feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation degree to obtain a gene expression dataset;
[0073] As a preferred implementation manner, based on single-cell transcriptome sequencing technology, obtaining the gene expression data of neutrophils, and performing feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation degree to obtain a gene expression dataset includes the following steps:
[0074] S11. Collect gene expression data of neutrophils through single-cell transcriptome sequencing technology, construct a gene expression matrix, and preprocess the gene expression matrix; obtain the gene characteristics of neutrophils according to the preprocessing results.
[0075] It should be noted that single-cell transcriptome sequencing technology can sequence the cell transcriptome at the single-cell level, that is, detect the differential expression of cells at a higher resolution level. The core of this technology lies in being able to successfully distinguish different single-cell transcripts in the sample.
[0076] First, neutrophils need to be extracted from the sample tissue and prepared into a single-cell suspension. Then, use single-cell transcriptome sequencing technology to sequence these single cells to obtain the gene expression data of each cell.
[0077] After sequencing, the obtained data needs to be processed to construct a gene expression matrix. The rows of the gene expression matrix usually represent genes, the columns represent cells, and the values in the matrix represent the number of transcripts of each gene in each single cell.
[0078] After constructing the gene expression matrix, a series of preprocessing steps need to be carried out to improve the accuracy and reliability of subsequent analysis. The preprocessing steps usually include:
[0079] Quality control: Remove low-quality cells and low-expressed genes to reduce the impact of noise on subsequent analysis. This can be achieved by setting certain thresholds, such as removing cells with fewer expressed genes than a certain value or genes with expression levels below a certain threshold.
[0080] Normalization: Since there may be differences in the sequencing depth between different cells or different genes, it is necessary to perform normalization processing on the gene expression matrix to eliminate this difference. Commonly used normalization methods include RPK (reads per million reads) or TPM (transcripts per million transcripts), etc.
[0081] Missing value processing: There may be missing values in the gene expression matrix, which may be caused by insufficient sequencing depth or too low gene expression levels. Cells or genes with missing values can be directly deleted, and then a filling method can be used for processing, such as filling with the average expression level or median of the gene.
[0082] Specifically, regardless of the filling method used, the accuracy of the filled gene expression data needs to be verified, especially whether the filled data can reflect the true expression level of key genes. This is a very important link. The following are several verification methods:
[0083] a. Biological verification: Verify the expression of key genes after filling through experiments, and conduct RNA-Seq verification or real-time quantitative PCR (qPCR) verification to ensure the accuracy and biological rationality of the filled data. This is the most direct method to verify whether the filling method is reasonable.
[0084] b. Multi-method comparison: Compare the filled data with other data processing methods (such as the results based on different missing value filling strategies) to check for consistency. Different filling methods may affect the final analysis results, so comparing the results of different processing methods can help verify the rationality of data processing.
[0085] c. Simulation verification: Introduce missing values into a known dataset (such as a simulated dataset) and test different filling methods to evaluate the filling effect. If the filling method can effectively recover the missing part of the simulated data, it indicates that the method may be effective in the application of actual data.
[0086] Pseudotime ordering processing: Perform pseudotime ordering on the gene expression matrix to determine the temporal order of gene expression dynamics.
[0087] After preprocessing, further analyze the gene characteristics of neutrophils based on the gene expression matrix. Gene characteristics include:
[0088] By comparing the gene expression data of neutrophils under different conditions or at different time points, differentially expressed genes can be identified. These genes play important roles in the differentiation, development, or function of neutrophils.
[0089] By analyzing the gene co-expression relationships in the gene expression matrix, a gene co-expression network can be constructed. This network can reveal the interactions and regulatory relationships between genes, helping to understand the gene expression regulation mechanism of neutrophils.
[0090] By performing clustering analysis on the gene expression matrix, neutrophils can be divided into different subpopulations. Each subpopulation may have unique gene expression characteristics and biological functions, contributing to a deeper understanding of the heterogeneity of neutrophils.
[0091] S12. According to the gene characteristics of neutrophils and combined with the differentiation state of cells, generate class labels for each gene characteristic through a clustering algorithm;
[0092] Specifically, from the set of characteristic genes extracted from neutrophil gene expression data, according to pseudotime ordering or known biological knowledge, cells are divided into different differentiation states (such as early differentiation, mid-differentiation, late differentiation), and each cell has a differentiation state label, which is used to guide the grouped analysis of gene characteristics.
[0093] For each differentiation state, extract the gene expression submatrix of all cells in that state, and perform dimensionality reduction on the extracted gene expression submatrix;
[0094] Use the K-means clustering algorithm to group the gene features (i.e., divide the genes into different classes), and each gene is assigned a class label indicating its clustering category in that differentiation state.
[0095] For example, assume there are 10 genes (Gene1 to Gene10) and 30 neutrophils (Cell1 to Cell30), and these cells have been divided into three differentiation states: early, middle, and late.
[0096] 1. Data preparation: There is a gene expression matrix with 30 rows (cells) × 10 columns (genes).
[0097] 2. Cell differentiation state division: By pseudotime ordering or biological knowledge, divide Cell1 - Cell10 into the early differentiation state, Cell11 - Cell20 into the middle differentiation state, and Cell21 - Cell30 into the late differentiation state.
[0098] 3. Extract submatrix:
[0099] Early differentiation state submatrix: Contains the gene expression data of Cell1 - Cell10.
[0100] Middle differentiation state submatrix: Contains the gene expression data of Cell11 - Cell20.
[0101] Late differentiation state submatrix: Contains the gene expression data of Cell21 - Cell30.
[0102] 4. Dimensionality reduction: Apply the PCA algorithm to each submatrix to reduce the gene expression data to 2D or 3D space.
[0103] 5. K-means clustering:
[0104] Apply K-means clustering to the early differentiation state submatrix, and cluster the 10 genes into 3 classes. Assume Gene1 - Gene3 are clustered into class 1, Gene4 - Gene6 are clustered into class 2, and Gene7 - Gene10 are clustered into class 3.
[0105] Similarly, perform K-means clustering on the middle and late differentiation state submatrices, and assign class labels to each gene.
[0106] 6. Class label generation:
[0107] For the early differentiation state, Gene1 - Gene3 are labeled as "Class 1 (early)", Gene4 - Gene6 are labeled as "Class 2 (early)", and Gene7 - Gene10 are labeled as "Class 3 (early)".
[0108] For the mid - stage and late - stage differentiation states, class labels are assigned to the genes in the same way.
[0109] S13. Determine the dynamic relevance weights of gene features by calculating the conditional mutual information of each gene feature with respect to the class labels and combining with the dynamic relevance weight algorithm; select the gene features of neutrophils iteratively according to the dynamic relevance weights of gene features, and construct a standard gene expression dataset based on the selection results.
[0110] As a preferred implementation, determine the dynamic relevance weights of gene features by calculating the conditional mutual information of each gene feature with respect to the class labels and combining with the dynamic relevance weight algorithm; select the gene features of neutrophils iteratively according to the dynamic relevance weights of gene features, and constructing a standard gene expression dataset based on the selection results includes the following steps:
[0111] S131. Initialize the gene expression dataset and construct a candidate feature set according to all gene features of neutrophils;
[0112] It should be noted that by creating an empty gene expression dataset (usually represented as an empty set or empty matrix), this dataset will be used to store the finally selected gene features.
[0113] According to all gene features of neutrophils, construct a candidate feature set containing all these features. This set can be represented as a list or matrix, where each row represents a gene feature and each column represents the expression value of the feature in different cells.
[0114] S132. For the gene features in the candidate feature set, calculate the conditional mutual information value of each gene feature with the class labels, and select the gene feature with the largest conditional mutual information value, remove it from the candidate feature set to the gene expression dataset, and update the gene expression dataset and the candidate feature set;
[0115] It should be noted that for each gene feature in the candidate feature set, calculate its conditional mutual information value with the class labels (such as the differentiation state). The conditional mutual information can be calculated by statistical methods or information - theoretic formulas. It measures the additional information that the gene feature can provide given the class labels. After calculating the conditional mutual information values of all candidate features, find the gene feature with the largest conditional mutual information value, and this feature is considered to be the most relevant to the class labels in the current candidate feature set, and thus the most likely to be useful for classification or prediction tasks.
[0116] Remove the selected gene features from the candidate feature set and add them to the gene expression dataset. The updated gene expression dataset contains one or more (depending on the number of iterations) gene features that are highly correlated with the class label. At the same time, the updated candidate feature set contains the remaining unselected gene features.
[0117] Specifically, by calculating the conditional mutual information between each candidate gene feature and the class label, select the genes with the highest mutual information as the preliminary feature set. The number of initial gene features is set according to experimental requirements. For example, select the top 10, 20, or 30 gene features that are most relevant to the class label.
[0118] S133. For the updated gene expression dataset and candidate feature set, calculate the additional new information amount of each gene feature in the candidate feature set for each gene feature in the gene expression dataset; (i.e., based on the influence of the feature on each feature in the gene expression dataset).
[0119] As a preferred embodiment, for the updated gene expression dataset and candidate feature set, the calculation formula for the additional new information amount of each gene feature in the candidate feature set for each gene feature in the gene expression dataset is:
[0120]
[0121] In the formula, extra ( X K , R ) represents the additional new information amount of each gene feature X K in the candidate feature set for each gene feature R in the gene expression dataset X R ; ∣ R ∣ represents the number of gene features in the gene expression dataset R ; Y represents the class label; I ( X R ; Y ) represents the mutual information value between the gene feature X R and the class label Y ; I ( X R ; Y ∣ X K ) represents that under the candidate condition of the gene feature X K in the gene featureX R and class label Y The conditional mutual information value between; H ( X R ) represents the gene feature X R Information entropy; H ( Y ) represents the class label Y Information entropy.
[0122] S134. According to the additional new information amount of each gene feature in the candidate feature set for each gene feature in the gene expression dataset, use the dynamic correlation weight algorithm to calculate the dynamic correlation weight of each gene feature in the current candidate feature set;
[0123] As a preferred implementation, according to the additional new information amount of each gene feature in the candidate feature set for each gene feature in the gene expression dataset, the calculation formula for using the dynamic correlation weight algorithm to calculate the dynamic correlation weight of each gene feature in the current candidate feature set is:
[0124]
[0125] In the formula, DW ( X K , R ) represents the dynamic correlation weight of each gene feature in the current candidate feature set X K ; extra ( X K , R ) represents the additional new information amount of each gene feature in the candidate feature set X K for each gene feature in the gene expression dataset R in X R ; I ( X K ; Y ∣ X R ) represents the conditional mutual information value between the gene feature X K and the class label Y ;
[0126] S135. Select the gene feature with the largest weight value from the candidate feature set according to the dynamic relevance weight, add it to the gene expression data set, and remove it from the candidate feature set. Repeat steps S133 - S135 until the number of gene features in the gene expression data set reaches the preset threshold, and the final gene expression data set is obtained.
[0127] Specifically, the selection of initial gene features should not be excessive, as this will lead to data set redundancy and affect the efficiency of subsequent calculations and the complexity of the model. By calculating the initial mutual information between the candidate feature set and the class label, select the first few most relevant gene features as the initial set to avoid redundancy. Usually, it is a reasonable range to select 5 - 10 gene features as the initial choice.
[0128] The number of iterations directly affects the convergence speed and calculation time of the model. The most relevant gene features are selected in each iteration, but the number of features selected in each iteration is limited. Therefore, there is a balance point to ensure that when the model converges, the accuracy of the gene expression data set is maximized while avoiding excessive iteration times. The use of dynamic relevance weight can help to more accurately select the most valuable gene features in each iteration and reduce unnecessary iterations. The number of iterations is limited by setting a threshold, for example, stop iterating when a certain number of features are reached or the increased information amount of the features is lower than a certain value.
[0129] Specifically, it is also possible to determine whether to continue iterating by monitoring the information gain of the selected gene features in each iteration; if the information gain of the newly added gene features in a certain round is very small for the data set, it means that the existing gene features are sufficient to express the relevance of the class label, and at this time, the iteration can be terminated.
[0130] For example: the following stopping conditions can be set:
[0131] Threshold condition 1: When the information gain of the newly added gene features in each iteration is less than a certain threshold (such as 1% or 0.5%), it is considered that the model has converged and the iteration can be terminated.
[0132] Threshold condition 2: When the information gain is less than the set threshold in consecutive multiple rounds of iteration (such as the information gain is less than 0.5% in consecutive 3 rounds), it is considered that the model has tended to be stable and the iteration is stopped.
[0133] Threshold condition 3: When the number of selected features reaches the preset maximum number, regardless of how the information gain changes, the iteration is terminated.
[0134] S2. Perform feature mapping processing on the gene expression data set based on the deep learning method, and construct multiple alternative causal models based on the mapping results; and use the Bayesian information criterion to evaluate the scores of the alternative models, and select the alternative model with the highest score as the best causal model according to the score results;
[0135] As a preferred embodiment, perform feature mapping processing on the gene expression dataset based on the deep learning method, and construct multiple alternative causal models based on the mapping results; and use the Bayesian information criterion to evaluate the scores of the alternative models, and select the alternative model with the highest score as the best causal model, including the following steps:
[0136] S21. Use the autoencoder in the deep learning model to perform feature mapping processing on the gene features in the gene expression dataset, and construct multiple alternative causal models based on the mapping results and in combination with the Granger causality analysis method;
[0137] As a preferred embodiment, use the autoencoder in the deep learning model to perform feature mapping processing on the gene features in the gene expression dataset, and construct multiple alternative causal models based on the mapping results and in combination with the Granger causality analysis method, including the following steps:
[0138] S211. Based on the pre-trained autoencoder, map the gene features in the gene expression dataset to a low-dimensional latent space through the encoder part in the autoencoder to obtain a low-dimensional feature representation;
[0139] It should be noted that an autoencoder usually consists of an encoder and a decoder. The encoder is responsible for mapping the input data (here it is the gene expression dataset) to a low-dimensional latent space (also called the latent space or feature space), and the decoder is responsible for mapping the data in the low-dimensional latent space back to the original high-dimensional space.
[0140] In practical applications, the autoencoder is usually pre-trained, which means that it has learned how to encode and decode effectively on a certain dataset. Pre-training can ensure that the autoencoder can capture the key features in the data and map them to the low-dimensional latent space. Use the encoder part of the pre-trained autoencoder to map each gene feature in the gene expression dataset to the low-dimensional latent space. The low-dimensional feature representation is the representation form of the gene expression data in the low-dimensional latent space. It is usually a vector, where each element represents the value of the original data on a certain low-dimensional feature.
[0141] Specifically, the temporal feature is an important part of the gene expression dataset. Especially in tasks such as cell development processes, disease progressions, and drug responses, gene expression is often related to time.
[0142] For the encoder part, LSTM or TCN, etc. can be used as the structure of the encoder. These models can map each time step (each moment of gene expression) to a low-dimensional latent space and capture the dependencies between time steps;
[0143] For the decoder part: The decoder attempts to reconstruct the original gene expression data from the low-dimensional latent space. In time series data, the design of the decoder can use LSTM, GRU, or TCN for reverse mapping of time steps to restore the original gene expression sequence.
[0144] Among them, the training process of the autoencoder optimizes the model by minimizing the error between the input data and the reconstructed data. Specifically, it includes:
[0145] Setting the training objective: The goal of the training process is to minimize the reconstruction error, that is, to map the gene expression data to the low-dimensional space through the encoder and then restore it to the original data space through the decoder. Commonly used reconstruction error metrics include mean squared error (MSE) or cross-entropy loss (if the data is sparse or discrete).
[0146] Self-supervised learning: The autoencoder is essentially a self-supervised learning method. It enables the model to learn how to reconstruct the output from the input by itself, thereby capturing the latent features of the data.
[0147] Optimization algorithm: Use gradient descent (such as the Adam optimizer) to optimize the model, and update the model parameters through backpropagation during training.
[0148] Regularization: To prevent the model from overfitting, techniques such as L2 regularization, dropout, or early stopping can be used to limit the complexity of the model.
[0149] In addition, the training of the autoencoder usually relies on a large amount of data. The dataset type is single-cell RNA-seq data: In gene expression data, single-cell RNA-seq is a relatively common data type, which provides the expression of each cell on different genes. If the dataset has time series information (such as sampling the same cell population at different time points), then a temporal autoencoder (such as LSTM or TCN) is more appropriate.
[0150] S212. Use the low-dimensional feature representation in the low-dimensional latent space, combine with the Granger causality analysis method to identify the causal relationships between genes, and generate a causal relationship network through causal relationship integration;
[0151] As a preferred implementation, using the low-dimensional feature representation in the low-dimensional latent space, combining with the Granger causality analysis method to identify the causal relationships between genes, and generating a causal relationship network through causal relationship integration includes the following steps:
[0152] S2121. In the low-dimensional latent space, through the trajectory inference method, perform pseudo-time ordering on neutrophils;
[0153] It should be noted that in single-cell data, cell differentiation is usually a dynamic and continuous process. Pseudotime ordering can infer the relative order of each cell during the differentiation process based on the change trajectory of gene expression. Common pseudotime ordering algorithms include, for example:
[0154] Monocle: Based on principal component analysis (PCA) or dimensionality reduction results.
[0155] PAGA: Infer pseudotime by combining the clustering graph and local connectivity.
[0156] Slingshot: Infer the trajectory by fitting a curve.
[0157] Specifically, the implementation steps of pseudotime ordering include:
[0158] Select a suitable pseudotime ordering algorithm (such as Monocle, PAGA, or Slingshot), and determine the starting point and branching points of the trajectory according to the algorithm requirements in the low-dimensional space.
[0159] Calculate the pseudotime value of each cell according to the path of the trajectory or graph network. The pseudotime value of each cell represents the relative position of the cell during the differentiation process.
[0160] S2122. According to the pseudotime ordering of the single-cell data, reorder the low-dimensional features of each gene to obtain a low-dimensional hidden space sequence;
[0161] It should be noted that for each cell, extract its feature representation in the low-dimensional hidden space (such as obtained through the encoder part of the autoencoder), and according to the result of the pseudotime ordering, reorder the low-dimensional features of each cell, that is, arrange the feature vectors of each cell in the order of the pseudotime value.
[0162] For each gene, combine its low-dimensional features in all cells in the order of pseudotime to form a low-dimensional hidden space sequence of this gene, which reflects the expression change of this gene during cell differentiation.
[0163] S2123. Perform Granger causality analysis on the low-dimensional hidden space sequence to identify the causal relationship between low-dimensional features;
[0164] It should be noted that Granger Causality Analysis is a statistical method based on time series data, used to determine whether a time series (cause variable) has a predictive effect on another time series (result variable), that is, whether there is a causal relationship. In biology, this method can be used to infer the regulatory relationship between genes.
[0165] Specifically, performing Granger causality analysis on the low-dimensional latent space sequence includes the following steps:
[0166] Determine the parameters in the Granger causality analysis model, such as the lag order (i.e., considering how many past time points of data to predict the data at the current time point). The selection of the lag order can be determined by model selection criteria (such as AIC, BIC, etc.);
[0167] Use the linear Granger causality model to perform Granger causality analysis on the low-dimensional latent space sequence. This usually involves calculating the predictive ability of the cause variable for the result variable and evaluating whether this predictive ability is significantly higher than the random level;
[0168] Based on the results of the Granger causality analysis, identify the causal relationships between the low-dimensional features. If the cause variable has a significant predictive effect on the result variable, it is considered that there is a causal relationship between them;
[0169] Verify the identified causal relationships. For example, use methods such as cross-validation and Bootstrap to evaluate the stability and reliability of the results; combine biological knowledge and literature reports to explain the biological significance of the identified causal relationships.
[0170] S2124. Through the decoder of the autoencoder, map the causal relationships in the low-dimensional latent space back to the high-dimensional gene space to generate a gene causal network.
[0171] It should be noted that in the low-dimensional latent space, causal relationships between genes are identified through methods such as Granger causality analysis. These causal relationships are represented in the form of interactions between low-dimensional features;
[0172] Take the identified low-dimensional causal relationships as input and map them back to the high-dimensional gene space through the decoder part of the autoencoder. The decoder will convert the low-dimensional causal relationships into causal relationships between high-dimensional gene expression data according to the learned data distribution and feature mapping relationships.
[0173] In the high-dimensional gene space, construct a gene causal network based on the causal relationships output by the decoder. The nodes in the network represent genes, and the edges represent the causal relationships between genes. Graph theory tools or visualization software can be used to display and analyze this network.
[0174] S213. Construct multiple alternative causal models based on the causal relationship network.
[0175] Specifically, constructing multiple alternative causal models based on the causal relationship network is an important step in the research, which helps to deeply understand and analyze the interaction mechanism between variables. Constructing multiple alternative causal models based on the causal relationship network includes the following steps:
[0176] I. Define the objectives and assumptions
[0177] Objective setting: First, clearly define the objective of constructing a causal model, such as explaining a specific biological phenomenon, etc.
[0178] Hypothesis formulation: Based on the causal relationship network and existing knowledge, propose multiple hypotheses about the interactions between variables. These hypotheses will serve as the basis for constructing alternative causal models.
[0179] II. Select a modeling method
[0180] Statistical methods: Such as regression analysis, path analysis, structural equation modeling, etc. These methods can help us quantify the causal relationships between variables and evaluate the goodness of fit of the model.
[0181] Machine learning methods: Such as Bayesian networks, causal graph models, etc. These methods can handle more complex causal relationship networks and automatically learn the dependencies between variables.
[0182] Expert systems: Combine the knowledge and experience of domain experts to construct rule-based causal models. This method is applicable when there is rich domain knowledge but insufficient data.
[0183] III. Construct alternative causal models
[0184] Model design: According to the proposed hypotheses and the selected modeling method, design multiple alternative causal models. Each model should include a set of variables and the causal relationships between them.
[0185] S22. Calculate the likelihood of the alternative causal models according to the Bayesian information criterion, and score the alternative causal models based on their likelihoods to obtain a scoring result;
[0186] As a preferred implementation, calculating the likelihood of the alternative causal models according to the Bayesian information criterion and scoring the alternative causal models based on their likelihoods to obtain a scoring result includes the following steps:
[0187] S221. For each alternative causal model, calculate the maximum likelihood estimate value of the alternative causal model under the gene expression dataset by maximizing the likelihood function;
[0188] It should be noted that for each alternative causal model Mi , its likelihood function L ( Mi ) represents the probability of observing the current gene expression dataset Mi under the model D : L ( Mi ) = P ( D ∣Mi )
[0189] Maximize the likelihood function by optimizing algorithms (such as gradient descent, Newton's method, etc.) L ( Mi ) and the calculation formula is:
[0190] G = argmax θ L ( Mi )
[0191] G represents the parameter estimation of the model Mi under maximum likelihood, θ represents the parameter set of the model Mi .
[0192] S222. Calculate the total number of free parameters according to the model structure, and calculate the score of each alternative causal model according to the Bayesian information criterion formula.
[0193] It should be noted that the Bayesian information criterion (BIC) is a model selection criterion used to evaluate the trade-off between model complexity and goodness of fit:
[0194]
[0195] In the formula, L max( Mi ) represents the maximum likelihood estimate value of the model Mi , k represents the total number of free parameters of the model Mi , n represents the number of samples (usually the number of cells in the gene expression dataset).
[0196] For each model Mi , count the number of all free parameters in the model, including: the coefficients of the linear regression model, the mixing ratios in the mixture model, the edge weights in the network model, etc.
[0197] By calculating the maximum likelihood estimate value of each alternative causal model and combining the total number of free parameters of the model, score the alternative models using the Bayesian information criterion (BIC). The smaller the BIC value, the lower the complexity of the model while fitting the data, so the greater the possibility of being selected as the best causal model. This process provides a quantitative standard for model selection, which helps to screen out the most suitable causal model from multiple alternative models.
[0198] S23. According to the scoring results of each alternative causal model, select the model with the highest score as the best causal model, and analyze the causal structure through compressed representation.
[0199] It should be noted that the goal of compressed representation is to map a complex high-dimensional causal network into a more concise form, retaining the core causal relationships for easy interpretation. Specifically, it includes the following steps:
[0200] Filter out the causal relationship edges with higher weights or significance from the best model to generate a simplified causal relationship network.
[0201] Use the random forest method to evaluate the causal importance of each feature (gene) and further screen for key genes.
[0202] Combine dimensionality reduction algorithms (such as PCA, t-SNE) to visualize the causal network and display the key causal paths.
[0203] S3. Construct a differentiation trajectory network based on the causal relationship network in the best model, and infer the differentiation trajectory of neutrophils according to the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory;
[0204] As a preferred implementation, constructing a differentiation trajectory network based on the causal relationship network in the best model and inferring the differentiation trajectory of neutrophils according to the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory includes the following steps:
[0205] S31. Integrate gene causal relationships according to the causal structure in the best causal model and generate a causal relationship network;
[0206] It should be noted that after the evaluation and screening of alternative causal models are completed, a best causal model is obtained. All gene causal relationships are extracted from the best causal model. These causal relationships are usually represented in the form of directed edges, where the starting point represents the cause gene and the ending point represents the result gene. In the causal relationship network, each gene is represented as a node, and the attributes of the node can include the name, expression level, regulatory role, etc. of the gene. According to the extracted causal relationships, directed edges are added between the nodes, and the attributes of the edges can include the strength, direction, confidence, etc. of the causal relationship. These attributes help us understand the regulatory relationships between genes more deeply.
[0207] S32. Incorporate pseudotime information into the causal relationship network to construct a differentiation trajectory network containing dynamic regulation and differentiation information;
[0208] It should be noted that in biological research, especially in single-cell transcriptome analysis, since the actual time information of cell differentiation cannot be directly obtained, researchers often use algorithms (such as Monocle, Slingshot, etc.) to infer the differentiation order of cells, and this inferred time information is called pseudotime. Pseudotime information can help understand the dynamic changes of cells during differentiation.
[0209] In biological networks, dynamic regulation usually refers to changes in gene expression, protein interactions, etc. over time. These changes are crucial for understanding processes such as cell differentiation, development, and disease occurrence.
[0210] Cell differentiation is a fundamental process in biology, involving the transformation of cells from one type to another. During differentiation, cells undergo a series of gene expression changes, which can be used to construct the differentiation trajectories of cells.
[0211] To construct a differentiation trajectory network that includes dynamic regulation and differentiation information, the following steps can be taken:
[0212] Data collection and processing: Collect transcriptome data of cells at different differentiation stages and perform preprocessing (such as denoising, normalization, etc.).
[0213] Pseudotime inference: Use appropriate algorithms (such as Monocle, Slingshot, etc.) to perform pseudotime inference on the transcriptome data to obtain the order of cell differentiation.
[0214] Causal relationship network construction: Based on the transcriptome data and pseudotime information, use causal inference methods (such as Granger causality test, dynamic Bayesian network, etc.) to construct a causal relationship network.
[0215] Dynamic regulation analysis: In the causal relationship network, analyze the changes in gene expression, protein interactions, etc. over time to identify key dynamic regulation events.
[0216] Differentiation trajectory network construction: Combine the pseudotime information and the results of dynamic regulation analysis to construct a differentiation trajectory network that includes dynamic regulation and differentiation information. This network should be able to clearly show the dynamic changes and regulatory mechanisms of cells during differentiation.
[0217] S33. Use the differentiation trajectory network for dynamic simulation, infer the differentiation path and key nodes, and output the prediction results of the differentiation trajectory.
[0218] Specifically, the differentiation trajectory network is a network model constructed based on single-cell transcriptome data, used to describe the dynamic changes and regulatory relationships of cells during differentiation. This network usually includes nodes (representing cells or cell types) and edges (representing the differentiation relationships or regulatory relationships between cells).
[0219] Based on the differentiation trajectory network, use dynamic network analysis methods to simulate the dynamic process of cell differentiation. Through dynamic simulation, infer the possible paths for cells to differentiate from one type to another, and use network analysis methods (such as centrality analysis, community detection, etc.) to identify the key nodes in the differentiation trajectory network.
[0220] The prediction results of the differentiation trajectories are usually presented in a graphical manner, such as differentiation trajectory trees, differentiation trajectory graphs, etc. These graphs can clearly show the paths of cell differentiation, key nodes, and the regulatory relationships between the nodes.
[0221] As Figure 2 shown, according to another embodiment of the present invention, a neutrophil differentiation trajectory inference system is provided. The neutrophil differentiation trajectory inference system includes: a gene expression dataset construction module 1, an optimal causal model construction module 2, and a differentiation trajectory inference module 3, and the gene expression dataset construction module 1, the optimal causal model construction module 2, and the differentiation trajectory inference module 3 are sequentially connected;
[0222] The gene expression dataset construction module 1 is used to obtain the gene expression data of neutrophils based on single-cell transcriptome sequencing technology, and perform feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation degree to obtain a gene expression dataset;
[0223] The optimal causal model construction module 2 is used to perform feature mapping processing on the gene expression dataset based on deep learning, and construct multiple alternative causal models based on the mapping results; and use the Bayesian information criterion to evaluate the scores of the alternative models, and select the alternative model with the highest score as the optimal causal model according to the score results;
[0224] The differentiation trajectory inference module 3 is used to construct a differentiation trajectory network according to the causal relationship network in the optimal model, and perform differentiation trajectory inference on neutrophils according to the differentiation trajectory network to obtain the prediction results of the neutrophil differentiation trajectory.
[0225] In summary, by means of the above technical solutions of the present invention, the present invention obtains gene expression data of neutrophils through single-cell transcriptome sequencing technology, providing a high-resolution and high-precision data basis at the cellular level for analysis, which helps to understand the differentiation process of neutrophils more deeply. The gene expression data is processed using a feature selection algorithm with dynamic relevance, which can screen out key genes closely related to neutrophil differentiation, reduce redundant information, and improve the accuracy and efficiency of subsequent analysis. Based on the causal relationship network in the best model, a differentiation trajectory network is constructed and differentiation trajectory inference is performed, and a prediction result of the neutrophil differentiation trajectory can be obtained, which helps to reveal the dynamic process and regulatory mechanism of neutrophil differentiation. Through single-cell transcriptome sequencing technology, the present invention can obtain gene expression data of neutrophils at the single-cell level. Pretreating the gene expression matrix can remove noise, correct biases, and improve the quality of the data. Combining the differentiation state of cells, class labels are generated for each gene feature through a clustering algorithm, which helps to group genes with similar expression patterns into one category. By calculating the conditional mutual information of each gene feature with respect to the class label and combining with the dynamic relevance weight algorithm, the correlation between each gene feature and the neutrophil differentiation state can be accurately evaluated. The gene features of neutrophils are selected iteratively, and a standard gene expression data set is constructed according to the selection results, which can gradually optimize the gene expression data set, retain the gene features most relevant to neutrophil differentiation, and remove redundant and noise information. The present invention uses an autoencoder to perform feature mapping on gene expression data, mapping high-dimensional data to a low-dimensional latent space, effectively reducing the dimension of the data while retaining key information. Combining with Granger causality analysis, causal relationships between genes are identified in the low-dimensional latent space, avoiding the complexity and noise interference of high-dimensional data and improving the accuracy of causal identification. The Bayesian information criterion is used to score alternative causal models, comprehensively considering the goodness of fit and complexity of the models, ensuring that the selected models have both good fitting effects and are not overly complex. Using the differentiation trajectory network for dynamic simulation can visually display the path and key nodes of cell differentiation, providing a powerful tool for understanding the mechanism of cell differentiation.
[0226] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0227] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for inferring neutrophil differentiation trajectories, characterized in that, The method for inferring the neutrophil differentiation trajectory includes the following steps: S1. Based on single-cell transcriptome sequencing technology, obtain the gene expression data of neutrophils, and perform feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation degree to obtain a gene expression data set; S2. Based on deep learning method, perform feature mapping processing on the gene expression data set, and based on the mapping result, construct multiple alternative causal models; and use the Bayesian information criterion to evaluate the scores of the alternative models, and select the alternative model with the highest score as the best causal model according to the score result; S3. Construct a differentiation trajectory network according to the causal relationship network in the best model, and infer the differentiation trajectory of neutrophils according to the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory; The step of obtaining the gene expression data of neutrophils based on single-cell transcriptome sequencing technology and performing feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation degree to obtain a gene expression data set includes the following steps: S11. Collect the gene expression data of neutrophils through single-cell transcriptome sequencing technology, construct a gene expression matrix, and perform preprocessing on the gene expression matrix; obtain the gene characteristics of neutrophils according to the preprocessing result; S12. According to the gene characteristics of neutrophils and combined with the differentiation state of cells, generate class labels for each gene characteristic through a clustering algorithm; S13. Calculate the conditional mutual information of each gene characteristic with respect to the class label, and combined with the dynamic correlation degree weight algorithm, determine the dynamic correlation degree weight of the gene characteristic; according to the dynamic correlation degree weight of the gene characteristic, select the gene characteristics of neutrophils in an iterative manner, and construct a standard gene expression data set according to the selection result; The step of calculating the conditional mutual information of each gene characteristic with respect to the class label, and combined with the dynamic correlation degree weight algorithm, determine the dynamic correlation degree weight of the gene characteristic; according to the dynamic correlation degree weight of the gene characteristic, select the gene characteristics of neutrophils in an iterative manner, and construct a standard gene expression data set according to the selection result includes the following steps: S131. Initialize the gene expression data set, and construct a candidate feature set according to all the gene characteristics of neutrophils; S132. For the gene characteristics in the candidate feature set, calculate the conditional mutual information value of each gene characteristic with the class label, and select the gene characteristic with the largest conditional mutual information value, remove it from the candidate feature set to the gene expression data set, and update the gene expression data set and the candidate feature set; S133. For the updated gene expression data set and candidate feature set, calculate the additional new information amount of each gene characteristic in the candidate feature set with respect to each gene characteristic in the gene expression data set; S134. According to the additional new information amount of each gene characteristic in the candidate feature set with respect to each gene characteristic in the gene expression data set, calculate the dynamic correlation degree weight of each gene characteristic in the current candidate feature set by using the dynamic correlation degree weight algorithm; S135. Select the gene feature with the largest weight value from the candidate feature set according to the dynamic relevance weight, add it to the gene expression data set, and remove it from the candidate feature set. Repeat steps S133 - S135 until the number of gene features in the gene expression data set reaches the preset threshold, and obtain the final gene expression data set; For the updated gene expression data set and candidate feature set, the calculation formula for the additional new information amount of each gene feature in the candidate feature set for each gene feature in the gene expression data set is: In the formula, extra ( X K , R ) represents the additional new information of each gene feature in the candidate feature set X K for each gene feature R in the gene expression dataset X R ; | R | represents the number of gene features R in the gene expression dataset; Y Indicates a class label; I ( X R ; Y ) represents a gene feature X S and class label Y the mutual information value between them; I ( X R ; Y ∣ X K ) represents a gene feature X R and a class label Y for the conditional mutual information value between; H ( X R ) represents the gene feature X R of the information entropy; H ( Y ) represents the class label Y of the information entropy.
2. The neutrophil differentiation trajectory inference method according to claim 1, wherein According to the additional new information amount of each gene feature in the candidate feature set for each gene feature in the gene expression data set, the calculation formula for the dynamic relevance weight of each gene feature in the current candidate feature set using the dynamic relevance weight algorithm is: In the formula, DW ( X K , R ) represents the dynamic correlation weight of each gene feature in the current candidate feature set X K ; extra ( X K , R ) represents each gene feature in the candidate feature set X K for the gene expression dataset R each gene feature in X R the additional new information content; I ( X K ; Y ∣ X R ) represents a gene feature X K and the class label Y The conditional mutual information value between them.
3. The neutrophil differentiation trajectory inference method according to claim 1, wherein Based on the deep learning method, perform feature mapping processing on the gene expression data set, and based on the mapping results, construct multiple alternative causal models; and use the Bayesian information criterion to evaluate the scores of the alternative models, and select the alternative model with the highest score as the best causal model according to the score results, including the following steps: S21. Use the autoencoder in the deep learning model to perform feature mapping processing on the gene features in the gene expression data set. Based on the mapping results, and combine the Granger causality analysis method to construct multiple alternative causal models; S22. Calculate the likelihood of the alternative causal models according to the Bayesian information criterion, and score the alternative causal models according to the likelihood of the alternative causal models to obtain the score results; S23. According to the score results of each alternative causal model, select the model with the highest score as the best causal model, and analyze the causal structure through compressed representation.
4. The method for inferring the neutrophil differentiation trajectory according to claim 3, wherein Using the autoencoder in the deep learning model to perform feature mapping processing on the gene features in the gene expression data set, and based on the mapping results, and combine the Granger causality analysis method to construct multiple alternative causal models, including the following steps: S211. Based on the pre-trained autoencoder, map the gene features in the gene expression data set to the low-dimensional hidden space through the encoder part in the autoencoder to obtain the low-dimensional feature representation; S212. Use the low-dimensional feature representation in the low-dimensional hidden space, combine the Granger causality analysis method to identify the causal relationships between genes, and generate a causal relationship network through causal relationship integration; S213. According to the causal relationship network, construct multiple alternative causal models.
5. The neutrophil differentiation trajectory inference method according to claim 4, characterized in that Using the low-dimensional feature representation in the low-dimensional hidden space, combine the Granger causality analysis method to identify the causal relationships between genes, and generate a causal relationship network through causal relationship integration, including the following steps: S2121. In the low-dimensional hidden space, perform pseudo-time sorting on neutrophils through the trajectory inference method; S2122. According to the pseudo-time sorting of the single-cell data, reorder the low-dimensional features of each gene to obtain the low-dimensional hidden space sequence; S2123. Perform Granger causality analysis on the low-dimensional hidden space sequence to identify the causal relationships between the low-dimensional features; S2124. Through the decoder of the autoencoder, map the causal relationships in the low-dimensional latent space back to the high-dimensional gene space to generate a gene causal network.
6. The method for inferring the neutrophil differentiation trajectory according to claim 5, wherein The steps of calculating the likelihood of alternative causal models according to the Bayesian information criterion and scoring the alternative causal models according to the likelihood of the alternative causal models to obtain the scoring result include the following steps: S221. For each alternative causal model, calculate the maximum likelihood estimate value of the alternative causal model under the gene expression dataset by maximizing the likelihood function; S222. Calculate the total number of free parameters according to the model structure, and calculate the score of each alternative causal model according to the Bayesian information criterion formula.
7. A method for inferring the neutrophil differentiation trajectory according to claim 1, characterized in that The steps of constructing a differentiation trajectory network according to the causal relationship network in the best model and inferring the differentiation trajectory of neutrophils based on the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory include the following steps: S31. Integrate the gene causal relationships according to the causal structure in the best causal model and generate a causal relationship network; S32. Incorporate the pseudotime information into the causal relationship network to construct a differentiation trajectory network containing dynamic regulation and differentiation information; S33. Use the differentiation trajectory network for dynamic simulation, infer the differentiation path and key nodes, and output the prediction result of the differentiation trajectory.
8. A neutrophil differentiation trajectory inference system for implementing the neutrophil differentiation trajectory inference method described in any one of claims 1-7, characterized in that, The neutrophil differentiation trajectory inference system includes: a gene expression dataset construction module, a best causal model construction module, and a differentiation trajectory inference module, and the gene expression dataset construction module, the best causal model construction module, and the differentiation trajectory inference module are connected in sequence; The gene expression dataset construction module is used to obtain the gene expression data of neutrophils based on single-cell transcriptome sequencing technology, and perform feature selection processing on the gene expression data of neutrophils through a feature selection algorithm of dynamic correlation degree to obtain a gene expression dataset; The best causal model construction module is used to perform feature mapping processing on the gene expression dataset based on the deep learning method, and construct multiple alternative causal models based on the mapping result; and use the Bayesian information criterion to evaluate the scores of the alternative models, and select the alternative model with the highest score as the best causal model according to the scoring result; The differentiation trajectory inference module is used to construct a differentiation trajectory network according to the causal relationship network in the best model, and infer the differentiation trajectory of neutrophils based on the differentiation trajectory network to obtain the prediction result of the neutrophil differentiation trajectory.
Citation Information
Patent Citations
Method and system for constructing dynamic gene regulation network and computer equipment
CN119317963A
Cell differentiation trajectory inference method based on deep hyperbolic manifold embedding
CN119479828A