Interpretable machine learning-based plant cell type prediction method and apparatus
By using an interpretable machine learning approach to preprocess and feature-select plant single-cell data, a multi-classification model is constructed, and the LightGBM algorithm is used for cell type prediction. This solves the problem of low accuracy in plant single-cell data analysis and achieves efficient and accurate cell type identification and marker gene mining.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE INST OF BIOTECHNOLOGY OF THE CHINESE ACAD OF AGRI SCI
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-21
AI Technical Summary
Existing plant single-cell data analysis methods have low accuracy and significant limitations, making it difficult to accurately classify and identify cell types.
Using an interpretable machine learning approach, we acquired single-cell data from the target plant, preprocessed it, screened and normalized it, constructed a multi-classification model, used the LightGBM algorithm to predict cell types, and identified specific marker genes through SHAP value analysis.
It improves the accuracy and efficiency of plant cell type prediction, identifies cell type-specific marker genes, and solves the problems of low accuracy and large limitations in existing technologies.
Smart Images

Figure CN2024135725_21052026_PF_FP_ABST
Abstract
Description
Plant cell type prediction method and device based on interpretable machine learning
[0001] Cross-references
[0002] This application claims priority to Chinese Patent Application No. 2024116074278, filed on November 12, 2024, entitled “Plant Cell Type Prediction Method and Device Based on Explainable Machine Learning”, the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] This invention relates to the field of artificial intelligence technology, and in particular to a method and apparatus for predicting plant cell types based on interpretable machine learning. Background Technology
[0004] Unlike higher animals, higher plant cells have high plasticity, and differentiated plant bodies can acquire totipotency and regenerate complete fertile plants.
[0005] In single-cell data analysis, accurate cell classification is the foundation of data analysis. Based on statistical methods, specific marker genes for different cell populations are identified. Finally, experimental methods require specific fluorescent labeling of the marker genes for cell types. However, the above methods lack accuracy and have significant limitations.
[0006] It is evident that the single-cell data analysis methods for plants in related technologies suffer from low accuracy and significant limitations. Summary of the Invention
[0007] This invention provides a plant cell type prediction method and apparatus based on interpretable machine learning, which solves the technical problems of low accuracy and large limitations in related single-cell data analysis methods for plants, and achieves efficient and accurate prediction of cell types for multiple organs of various plants.
[0008] This invention provides a plant cell type prediction method based on interpretable machine learning, comprising the following steps.
[0009] Single-cell data of the target organ of the target plant are obtained; the single-cell data are preprocessed to obtain standardized single-cell data; feature screening is performed on the standardized single-cell data to obtain hypervariable genes, wherein the expression level of the hypervariable gene in the target cell is higher than a first expression threshold, and the expression level of the hypervariable gene in other cells is lower than the first expression threshold; normalization is performed based on the single-cell data and the hypervariable genes to obtain a cell gene matrix; a target multi-classification model corresponding to the target organ of the target plant is determined; the cell gene matrix is input into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model.
[0010] According to the present invention, a plant cell type prediction method based on interpretable machine learning, after inputting the cell gene matrix into the target multi-classification model and obtaining the predicted cell type output by the target multi-classification model, the method further includes: performing SHAP value analysis on the target multi-classification model to obtain specific marker genes corresponding to the predicted cell type.
[0011] According to the present invention, a plant cell type prediction method based on interpretable machine learning is provided. The preprocessing of the single-cell data to obtain standardized single-cell data includes: determining the original single-cell RNA sequencing data; performing quality control processing on the original single-cell RNA sequencing data to obtain filtered single-cell RNA sequencing data, wherein the quality control processing includes: removing low-expression genes and removing mitochondrial genes, wherein the expression level of the low-expression genes is below a second expression threshold; and standardizing the sequencing depth and gene expression levels of the filtered single-cell RNA sequencing data to obtain standardized single-cell data.
[0012] According to the present invention, a plant cell type prediction method based on interpretable machine learning further includes, before obtaining single-cell data of target organs of a target plant, the method further includes: obtaining single-cell data of different organs of multiple plant species; preprocessing and annotating the single-cell data to obtain cell type labels for each single-cell data; performing feature filtering on the single-cell data based on the cell type labels to obtain a target number of hypervariable genes for each single-cell data; normalizing the single-cell data and the hypervariable genes for each single-cell data to obtain a cell gene matrix for each single cell; classifying the cell gene matrix of each single cell according to different organs of the multiple plant species to obtain a training dataset for each organ of each plant species, wherein the training dataset includes: the cell gene matrix and cell type labels corresponding to the cell gene matrix; and constructing multi-classification models corresponding to each organ of each plant species based on the training dataset for each organ of each plant species and the target machine learning algorithm.
[0013] According to the present invention, a plant cell type prediction method based on interpretable machine learning, before constructing a multi-classification model corresponding to each organ of each plant species based on a training dataset of each organ of each plant species and a target machine learning algorithm, the method further includes: obtaining a test dataset for plant cell type prediction, wherein the test dataset includes a test cell gene matrix and cell type labels corresponding to the test cell gene matrix; inputting the test cell gene matrix into a plurality of preset machine learning algorithms to obtain the test predicted cell type and test time output by each of the plurality of machine learning algorithms; comparing the test predicted cell type output by each machine learning algorithm with the cell type labels to obtain the test accuracy of each machine learning algorithm; and determining a target machine learning algorithm based on the test time, the test accuracy, and the model robustness of each machine learning algorithm.
[0014] According to the present invention, a plant cell type prediction method based on interpretable machine learning is provided, wherein the target machine learning algorithm is the LightGBM algorithm.
[0015] This invention also provides a plant cell type prediction device based on interpretable machine learning, comprising the following modules: an acquisition module for acquiring single-cell data of a target organ of a target plant; a preprocessing module for preprocessing the single-cell data to obtain standardized single-cell data; a screening module for feature screening of the standardized single-cell data to obtain hypervariable genes in the single-cell data, wherein the expression level of the hypervariable gene in the target cell is higher than a first expression threshold, and the expression level of the hypervariable gene in other cells is lower than the first expression threshold; a normalization module for normalizing the single-cell data and the hypervariable genes to obtain a cell gene matrix; a determination module for determining a target multi-classification model corresponding to the target organ of the target plant; and a prediction module for inputting the cell gene matrix into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model.
[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the plant cell type prediction method based on interpretable machine learning as described above.
[0017] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the plant cell type prediction method based on interpretable machine learning as described above.
[0018] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the plant cell type prediction method based on interpretable machine learning as described above.
[0019] The plant cell type prediction method and apparatus based on interpretable machine learning provided by this invention can eliminate noise and bias in single-cell data by preprocessing the data, thereby improving the accuracy and reliability of the data. Through feature screening, highly variable genes with significantly different expression levels in single-cell data can be identified, which is of great significance for distinguishing different cell types. At the same time, by setting a first expression threshold, genes that are highly expressed only in target cells can be further screened, thereby improving the accuracy and specificity of feature selection. Based on the normalization processing of single-cell data and highly variable genes, the expression differences between different samples can be eliminated, making the data more comparable and easier to analyze. By determining the target multi-classification model corresponding to the target organ of the target plant, the cell gene matrix can be specifically input into the model to classify cell types more efficiently and accurately, and identify the predicted cell types in the target organs of the target plant. Thus, it solves the technical problems of low accuracy and large limitations in related single-cell data analysis methods for plants. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0021] Figure 1 is a flowchart illustrating the plant cell type prediction method based on interpretable machine learning provided by this invention.
[0022] Figure 2 is a schematic diagram of the overall process of the plant cell type prediction method based on interpretable machine learning provided by the present invention.
[0023] Figure 3 is a schematic diagram comparing the prediction performance of the model based on the LightGBM algorithm provided by this invention.
[0024] Figure 4 is a schematic diagram of the plant cell type prediction device based on interpretable machine learning provided by the present invention.
[0025] Figure 5 is a schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0027] Single-cell sequencing technology is rapidly developing, becoming a powerful method for studying gene expression in complex multicellular organisms. Compared to bulk sequencing, single-cell sequencing can identify rare cell populations and reveal state transitions at different developmental stages, which are difficult to capture using traditional methods. Therefore, solving scientific problems at the single-cell level is currently a key research focus.
[0028] Unlike higher animals, higher plants exhibit high cell fate plasticity, and differentiated plant bodies can acquire totipotency, regenerating complete fertile plants. Currently, many single-cell sequencing technologies are applied to model organisms such as Arabidopsis thaliana, rice, and maize, and this research is particularly important for plant studies. In single-cell data analysis, accurate cell classification is fundamental. While unsupervised clustering algorithms are used to classify cells, and statistical methods are used to identify specific marker genes for different cell populations, the final experimental methods require cell type-specific fluorescent labeling of these marker genes. However, these methods lack accuracy and have significant limitations.
[0029] Therefore, starting with various single-cell omics and machine learning methods, standardizing and unifying the data processing, and modeling and analyzing multi-omics data of different single-cell populations are of great significance for understanding cell heterogeneity and identifying new cell subpopulations and marker genes.
[0030] The main technical problem to be solved by this invention is to perform intelligent prediction of plant single-cell data and to discover new plant cell type-specific marker genes.
[0031] To address the aforementioned technical challenges, this invention integrates existing plant single-cell data. After re-integrating and analyzing the data, a machine learning model is constructed for each tissue of each species to accurately predict cell type. Furthermore, cell type-specific marker genes are selected within the model based on feature importance for subsequent experimental validation.
[0032] Optionally, the plant cell type prediction method based on interpretable machine learning in this embodiment can be executed by a server, by a terminal device, or by both a server and a terminal device. For example, the plant cell type prediction method based on interpretable machine learning in this embodiment can be executed by a server.
[0033] Figure 1 is a flowchart illustrating the plant cell type prediction method based on interpretable machine learning provided by the present invention. As shown in Figure 1, the method includes the following:
[0034] Step 101: Obtain single-cell data of the target organ of the target plant.
[0035] It should be noted that the single-cell data is extracted from the target organ of the target plant. It is understood that the target plant and the target organ are predetermined.
[0036] In this embodiment of the invention, single-cell RNA sequencing (scRNA-seq) data from specific organs of the target plant are collected; single-cell sequencing technologies, such as 10x Genomics and SMART-seq2, are used to sequence the cells of the target organ to obtain single-cell data containing a large amount of cellular gene expression information.
[0037] Step 102: Preprocess the single-cell data to obtain standardized single-cell data.
[0038] The main purpose of standardizing single-cell data is to eliminate technical bias, batch effects, and improve data consistency and comparability.
[0039] In this embodiment of the invention, based on indicators such as sequencing depth and number of gene detections, cells of low quality are identified and removed; genes with extremely low expression levels in all cells are removed, which may be generated due to sequencing noise.
[0040] Since the sequencing depth may vary among different cells, it is necessary to divide the gene expression level of each cell by the sequencing depth of that cell (i.e., the total number of reads or UMIs) to eliminate the influence of sequencing depth on the expression level; and to perform a logarithmic transformation (such as log2 or log10) on the normalized expression level to reduce the influence of extreme values and make the data closer to a normal distribution.
[0041] If the data comes from different experimental batches, a specific algorithm is used to correct for batch effects.
[0042] In some embodiments, the expression level of each gene in single-cell data is converted into a Z-score, which is the mean of the gene across all cells, and then divided by the standard deviation to eliminate differences in gene expression levels.
[0043] Step 103: Perform feature screening on the standardized single-cell data to obtain hypervariable genes in the single-cell data. Among them, the expression level of hypervariable genes in target cells is higher than the first expression threshold, and the expression level of hypervariable genes in other cells is lower than the first expression threshold.
[0044] Feature screening was performed on standardized single-cell data to identify hypervariable genes (i.e. genes that are expressed in a specific cell type in a way that is significantly different from those in other cell types).
[0045] A first expression threshold is set based on the overall distribution of the data or the expression characteristics of a specific cell type. The first expression threshold is used to distinguish between highly expressed and poorly expressed genes.
[0046] For each gene in the standardized single-cell data, its expression level is compared in the target cell and other cells. Genes whose expression level in the target cell is higher than a first expression threshold and whose expression level in other cells is lower than the first expression threshold are selected as hypervariable genes.
[0047] Step 104: Normalize the single-cell data and highly variable genes to obtain the cell gene matrix.
[0048] In this embodiment of the invention, the expression data of these hypervariable genes in all cells are extracted to obtain hypervariable gene expression data;
[0049] The hypervariable gene expression data were normalized and organized into a cell gene matrix; where the rows of the cell gene matrix represent genes, the columns represent cells, and the values in the matrix represent the expression level of genes in cells.
[0050] Here, the purpose of normalization is to eliminate the differences in expression levels between different genes, so that the expression data of different genes can be compared on the same scale.
[0051] For example, commonly used normalization methods include Z-score standardization, min-max normalization, or quantile normalization.
[0052] Step 105: Determine the target multi-classification model corresponding to the target organ of the target plant.
[0053] To perform cell type classification more efficiently and accurately, a target multi-classification model corresponding to the target organ of the target plant is determined from multiple pre-set multi-classification models for different plant organs of different plant species.
[0054] Step 106: Input the cell gene matrix into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model.
[0055] In this embodiment of the invention, the cell gene matrix is passed as input data to the target multi-classification model. For example, the prediction function of the multi-classification model is called, and the cell gene matrix is used as the parameter of the function. The target multi-classification model takes the received cell gene matrix as input, performs internal calculations, including feature extraction, classification decision and other steps, and outputs prediction results. The prediction results include a vector or list of predicted cell types corresponding to single cell data.
[0056] This invention provides a method for obtaining single-cell data of target organs of a target plant. The single-cell data is preprocessed to obtain standardized single-cell data. Feature screening is performed on the standardized single-cell data to identify hypervariable genes, where the expression level of these genes in target cells is higher than a first expression threshold, while their expression level in other cells is lower than the first expression threshold. Normalization is performed based on the single-cell data and the hypervariable genes to obtain a cell gene matrix. A target multi-classification model corresponding to the target organ of the target plant is determined. The cell gene matrix is input into the target multi-classification model to obtain the predicted cell type output by the model. This method enables targeted input of the cell gene matrix into the model for more efficient and accurate cell type classification, identifying the predicted cell type in the target organ of the target plant. This solves the technical problems of low accuracy and significant limitations in related single-cell data analysis methods for plants.
[0057] According to the plant cell type prediction method based on interpretable machine learning provided by the present invention, after inputting the cell gene matrix into a target multi-classification model and obtaining the predicted cell type output by the target multi-classification model, the method further includes:
[0058] SHAP value analysis was performed on the target multi-classification model to obtain specific marker genes corresponding to the predicted cell types.
[0059] In this embodiment of the invention, based on the feature importance of the target multi-classification model and the analysis of SHAP importance, marker genes specific to each cell type are screened out.
[0060] In some embodiments, SHAP value calculation analysis is performed on each constructed model to identify genes with high importance for cell type differentiation, and the contribution of each gene in the prediction of a specific cell type is quantified based on the SHAP value corresponding to each cell type. The final results are output in tabular form.
[0061] SHAP (SHapley Additive exPlanations) value analysis is a method for interpreting the prediction results of machine learning models, applicable to classification and regression tasks. Through SHAP value analysis, we can understand the contribution of each feature (in this case, genes) to the model's prediction results, thereby identifying specific marker genes corresponding to the predicted cell type.
[0062] For example, depending on the type of the target multi-class classification model (such as a tree-based model, a linear model, etc.), a suitable SHAP interpreter is selected for initialization; using the initialized SHAP interpreter, SHAP values are calculated for the cell gene matrix. This will generate a SHAP value for each gene in each cell sample, representing the gene's contribution to the model's prediction results.
[0063] By analyzing the magnitude and sign of SHAP values, specific marker genes corresponding to the predicted cell type can be identified. Generally, genes with larger SHAP values and clearly defined positive and negative signs are more likely to be specific marker genes.
[0064] Through the embodiments of the present invention, SHAP value analysis is used to interpret the target multi-classification model and identify specific marker genes corresponding to the predicted cell types. These specific marker genes are used to understand the biological characteristics of cell types.
[0065] According to the present invention, a plant cell type prediction method based on interpretable machine learning preprocesses single-cell data to obtain standardized single-cell data, including:
[0066] Determine the raw single-cell RNA sequencing data for single-cell data;
[0067] The raw single-cell RNA sequencing data was subjected to quality control processing to obtain filtered single-cell RNA sequencing data. The quality control processing included: removing low-expression genes and removing mitochondrial genes. The expression level of low-expression genes was lower than the second expression threshold.
[0068] The sequencing depth and gene expression levels of the filtered single-cell RNA sequencing data were standardized to obtain standardized single-cell data.
[0069] The second expression threshold is preset and used to distinguish between low-expression genes and high-expression genes. For example, the gene expression level is compared with the second expression threshold, and genes with expression levels higher than the threshold are retained, while low-expression genes with expression levels lower than the threshold are removed.
[0070] In this embodiment of the invention, data collection and preprocessing includes: collecting single-cell data and marker gene data from multiple plant species.
[0071] Quality control was performed on the collected raw single-cell RNA sequencing data to remove low-quality cells, low-expression genes, and mitochondrial genes, ensuring the reliability of subsequent analysis data. The filtered data was then standardized to eliminate technical errors, making gene expression levels comparable across cells with different sequencing depths and library sizes, thus better reflecting differences at the gene level.
[0072] Hypervariable genes were screened for those that were highly expressed in some cells and lowly expressed in others. These genes are typically the most important for distinguishing cell types and were therefore used in subsequent analyses. Further normalization analysis was performed using the hypervariable genes, and the resulting matrix was used for subsequent modeling analysis.
[0073] Furthermore, we integrated single-cell data from different species at the organ level. Single-cell data from different species and the same organ from different sources were merged and batch-corrected using Harmony with the data source as the variable.
[0074] Finally, the integrated data was re-annotated by cell type. Appropriate cell type labels were assigned to each cell based on known marker genes and cell characteristics from the literature. This process may include manual review and correction to ensure the accuracy of the annotation.
[0075] According to the plant cell type prediction method based on interpretable machine learning provided by the present invention, before obtaining single-cell data of the target organ of the target plant, the method further includes:
[0076] Obtain single-cell data of different organs from multiple plant species;
[0077] Preprocessing and annotation of single-cell data yields cell type labels for each single-cell dataset;
[0078] Feature filtering of single-cell data based on cell type labels yields a target number of hypervariable genes for each single-cell dataset.
[0079] The cell gene matrix for each single cell is obtained by normalizing the data of each single cell and the hypervariable genes of each single cell.
[0080] The cell gene matrix of each single cell is classified according to different organs of multiple plant species to obtain the training dataset for each organ of each plant species. The training dataset includes: cell gene matrix and cell type label corresponding to cell gene matrix.
[0081] Based on the training dataset and target machine learning algorithm for each organ of each plant species, a multi-classification model corresponding to each organ of each plant species is constructed.
[0082] In this embodiment of the invention, single-cell RNA sequencing data are collected from different organs of multiple plant species. This data is typically obtained using high-throughput sequencing technology and contains gene expression information for each cell.
[0083] The raw single-cell data undergoes preprocessing, including noise removal, batch effect correction, and removal of low-quality cells. Cell type annotation is then performed on each single-cell data point using known gene markers or databases, resulting in a cell type label for each cell. For example, this can be achieved by comparing gene expression profiles with characteristic gene expression profiles of known cell types.
[0084] Based on cell type tags, feature filtering is performed on single-cell data for each organ of each plant species to identify hypervariable genes (i.e., genes with significant expression differences between different cell types) in each organ. This can be achieved, for example, through differential expression analysis, analysis of variance, and other methods.
[0085] Based on the analysis requirements, determine the number of hypervariable genes to be retained for each single-cell dataset. The target number can be determined based on the characteristics of the data, the purpose of the analysis, and the limitations of computing resources.
[0086] Based on each single-cell dataset and its hypervariable genes, normalization is performed to eliminate differences in gene expression levels, making the data more suitable for subsequent machine learning algorithms. Normalization methods include, but are not limited to, z-score normalization and min-max normalization.
[0087] The normalized data is organized into a cell gene matrix, where each row represents a gene, each column represents a cell, and the values in the matrix represent the expression level of the gene in the cell.
[0088] The cell gene matrices are classified according to different organs of multiple plant species to obtain training datasets for each organ of each plant species. The datasets include cell gene matrices and their corresponding cell type labels.
[0089] Choose an appropriate machine learning algorithm, such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Tree (GBDT), or deep learning algorithm.
[0090] Using the training dataset for each organ of each plant species, a multi-classification model corresponding to each organ is trained so that the multi-classification model can predict the cell type based on the cell gene matrix.
[0091] The model's performance is evaluated using metrics such as cross-validation, accuracy, recall, and F1 score. Based on the evaluation results, the model is optimized to improve its predictive accuracy.
[0092] Through the embodiments of the present invention, a multi-classification model can be constructed for each organ of each plant species, thereby achieving accurate classification of single-cell data and accurate prediction of cell types.
[0093] According to the plant cell type prediction method based on interpretable machine learning provided by the present invention, before constructing a multi-classification model corresponding to each organ of each plant species based on the training dataset and target machine learning algorithm of each organ of each plant species, the method further includes:
[0094] Obtain a test dataset for plant cell type prediction, wherein the test dataset includes a test cell gene matrix and cell type labels corresponding to the test cell gene matrix;
[0095] The test cell gene matrix is input into multiple preset machine learning algorithms to obtain the test predicted cell type and test time output by each machine learning algorithm.
[0096] The test prediction cell type output by each machine learning algorithm is compared with the cell type label to obtain the test accuracy of each machine learning algorithm;
[0097] The target machine learning algorithm is determined based on the test time, test accuracy, and model robustness of each machine learning algorithm.
[0098] In this embodiment of the invention, a test dataset for plant cell type prediction is collected. The test dataset includes a test cell gene matrix, i.e., gene expression data for each test cell, and cell type labels corresponding to these gene matrices, i.e., the true cell type of each test cell.
[0099] Choose a series of pre-defined machine learning algorithms, such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Tree (e.g., XGBoost, LightGBM), and Neural Networks (e.g., deep learning models). These algorithms have already undergone preliminary training or optimization.
[0100] The gene matrix of the test cells is input into these pre-defined machine learning algorithms. Each algorithm makes a prediction based on the input data and outputs the predicted cell type.
[0101] While inputting test data, record the test time required for each algorithm to complete a prediction. The test time is used to evaluate the computational efficiency of the algorithm and its feasibility in practical applications.
[0102] The predicted cell type output by each machine learning algorithm is compared with the cell type labels in the test dataset, for example, by calculating evaluation metrics such as accuracy, recall, and F1 score; based on the comparison results, the test accuracy of each machine learning algorithm is calculated.
[0103] It should be noted that, in addition to test accuracy, the test time (i.e., computational efficiency) and model robustness (i.e., the stability of the model's performance under different datasets and conditions) of each machine learning algorithm also need to be considered. Based on the comprehensive evaluation results of test accuracy, test time, and model robustness, the optimal target machine learning algorithm is determined.
[0104] Through the embodiments of the present invention, the performance of multiple machine learning algorithms on plant cell type prediction tasks can be systematically evaluated, and the target machine learning algorithm can be determined by combining a comprehensive evaluation of accuracy, testing time and model robustness.
[0105] According to the present invention, a plant cell type prediction method based on interpretable machine learning is provided, wherein the target machine learning algorithm is the LightGBM algorithm.
[0106] The following describes an example of the plant cell type prediction method based on interpretable machine learning provided by this invention in a practical application scenario, specifically including the following steps.
[0107] Referring to Figure 2, which is a schematic diagram of the overall process of the plant cell type prediction method based on interpretable machine learning provided by the present invention, the method includes: data processing, model training, and downstream analysis.
[0108] As shown in Figure 2, data processing includes input single-cell RNA sequencing, i.e., scRNA (Genes), followed by feature selection of the input data to identify top genes with highly variable expression, and the construction of a cell gene matrix based on the top genes and cells. In the cell gene matrix, the rows represent genes, the columns represent cells, and the values in the matrix represent the expression level of the genes in the cells.
[0109] The model training specifically includes the following modules: Logistic Regression, Support Vector Machine, CatBoost, Neural Networks, and LightGBM.
[0110] Downstream analysis specifically includes predicting cell type and identifying marker genes.
[0111] Step 1, Data Collection and Preprocessing: Collect single-cell data and marker gene data from multiple plant species.
[0112] Quality control was performed on the collected raw single-cell RNA sequencing data to remove low-quality cells, low-expression genes, and mitochondrial genes, ensuring the reliability of subsequent analysis data.
[0113] The filtered data was standardized to eliminate technical errors, making the gene expression levels of cells with different sequencing depths and library sizes comparable, thus better reflecting differences at the gene level.
[0114] We screen for hypervariable genes that are highly expressed in some cells and lowly expressed in others. These genes are often the most important for distinguishing cell types and are therefore used for subsequent analysis.
[0115] Further normalization analysis of the data was performed using highly variable genes, and the resulting matrix was used for subsequent modeling analysis.
[0116] Simultaneously, single-cell data from different species are integrated at the organ level. Single-cell data from different species and the same organ from different sources are merged and batch-corrected using Harmony with the data source as the variable.
[0117] Finally, we re-annotated the integrated data by cell type. Based on known marker genes and cell characteristics in the literature, we assigned appropriate cell type labels to each cell. This process may involve manual review and correction to ensure the accuracy of the annotation.
[0118] Step 2: Calculate 5,000 hypervariable genes for each data point based on the upstream standard process and use them as input features for the model.
[0119] The upstream standard process here is the data collection and preprocessing in step 1 above.
[0120] Step 3, Multi-class Learning Model Construction: By comparing the classification performance of multiple machine learning models, the optimal model was selected based on computational accuracy, computational time cost, and model robustness. The LightGBM algorithm was ultimately adopted, which effectively reduces overfitting and improves the model's accuracy and generalization ability. Models were then constructed to predict cell types for each organ of each species.
[0121] Referring to Figure 3, which is a schematic diagram comparing the prediction performance of the model based on the LightGBM algorithm provided by the present invention, it includes a comparison of the area under the curve (AUC) of the model between different cell types. The horizontal axis includes: Seurat prediction (prediction_Seurat), singleR prediction (prediction_singleR), scPred prediction (prediction_scPred), and scMarker prediction (prediction_scMarker).
[0122] As shown in Figure 3, the cell types include: root endodermis, sclerenchyma, pericycle, root cortex, meristematic cell, exodermis, root hair, root epidermis, root cap, root stele, protoxylem, phloem, and metaxylem.
[0123] Meristematic cells, exodermis, root hairs, epidermal cells, root tip cap, main axis, and secondary xylem and phloem.
[0124] For each species, single-cell data (a matrix of cells × genes) of each organ are modeled and classified to obtain the cell type of each cell (e.g., epidermis, column, outer epidermis, etc.). The marker genes for each cell type in different organs are obtained by ranking the importance of features according to the multi-classification model. The obtained genes can be used for downstream experimental verification and other work.
[0125] It should be noted that single-cell technology has only been applied to a limited number of plant species, approximately 30, and only a few species (rice, maize, Arabidopsis thaliana) have had multiple organs studied; most other species have only one organ. In this embodiment of the invention, 25 species and approximately 2.5 million cells are involved. The term "each species" as used herein refers to the collected species data (i.e., a total of 25 species).
[0126] Step 4, Marker Gene Mining: Based on the analysis of the model's feature importance and SHAP importance, specific marker genes for each cell type are screened out.
[0127] SHAP value calculation and analysis were performed on each constructed model to identify genes with high importance for cell type differentiation. The contribution of each gene in the prediction of a specific cell type was quantified based on the SHAP value corresponding to each cell type. The final results were output to the local environment in tabular form.
[0128] Through the above methods, this invention can predict cell types in multiple organs of various plants, and has the ability to discover cell type-specific marker genes and good cross-species applicability, providing a new technical means for plant single-cell research.
[0129] The plant cell type prediction device based on interpretable machine learning provided by the present invention will be described below. The plant cell type prediction device based on interpretable machine learning described below can be referred to in correspondence with the plant cell type prediction method based on interpretable machine learning described above.
[0130] Referring to Figure 4, which is a schematic diagram of the plant cell type prediction device based on interpretable machine learning provided by the present invention.
[0131] The acquisition module 401 is used to acquire single-cell data of the target organ of the target plant;
[0132] The preprocessing module 402 is used to preprocess single-cell data to obtain standardized single-cell data;
[0133] The screening module 403 is used to perform feature screening on standardized single-cell data to obtain hypervariable genes in the single-cell data, wherein the expression level of hypervariable genes in target cells is higher than a first expression threshold, and the expression level of hypervariable genes in other cells is lower than the first expression threshold.
[0134] Normalization module 404 is used to normalize based on single-cell data and highly variable genes to obtain a cell gene matrix.
[0135] Module 405 is used to determine the target multi-classification model corresponding to the target organ of the target plant;
[0136] The prediction module 406 is used to input the cell gene matrix into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model.
[0137] Specifically, the plant cell type prediction device based on interpretable machine learning provided by the present invention can realize all the method steps implemented in the above-described plant cell type prediction method embodiment based on interpretable machine learning, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0138] Figure 5 is a schematic diagram of the physical structure of the electronic device provided by the present invention. As shown in Figure 5, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. The processor 510, communication interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a plant cell type prediction method based on interpretable machine learning. This method includes: acquiring single-cell data of a target organ of a target plant; preprocessing the single-cell data to obtain standardized single-cell data; performing feature screening on the standardized single-cell data to obtain hypervariable genes in the single-cell data, wherein the expression level of the hypervariable gene in the target cell is higher than a first expression threshold, and the expression level of the hypervariable gene in other cells is lower than the first expression threshold; normalizing the single-cell data and the hypervariable genes to obtain a cell gene matrix; determining a target multi-classification model corresponding to the target organ of the target plant; and inputting the cell gene matrix into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model.
[0139] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0140] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the plant cell type prediction method based on interpretable machine learning provided by the above methods. The method includes: acquiring single-cell data of a target organ of a target plant; preprocessing the single-cell data to obtain standardized single-cell data; performing feature screening on the standardized single-cell data to obtain hypervariable genes in the single-cell data, wherein the expression level of the hypervariable genes in the target cells is higher than a first expression threshold, and the expression level of the hypervariable genes in other cells is lower than the first expression threshold; normalizing the single-cell data and the hypervariable genes to obtain a cell gene matrix; determining a target multi-classification model corresponding to the target organ of the target plant; and inputting the cell gene matrix into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model.
[0141] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the plant cell type prediction method based on interpretable machine learning provided by the above methods. This method includes: acquiring single-cell data of a target organ of a target plant; preprocessing the single-cell data to obtain standardized single-cell data; performing feature filtering on the standardized single-cell data to obtain hypervariable genes in the single-cell data, wherein the expression level of the hypervariable genes in the target cells is higher than a first expression threshold, and the expression level of the hypervariable genes in other cells is lower than the first expression threshold; normalizing the single-cell data and the hypervariable genes to obtain a cell gene matrix; determining a target multi-classification model corresponding to the target organ of the target plant; and inputting the cell gene matrix into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model.
[0142] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0143] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Industrial applicability
[0145] This invention provides a method and apparatus for predicting plant cell types based on interpretable machine learning. The method includes: acquiring single-cell data of a target organ of a target plant; preprocessing the single-cell data to obtain standardized single-cell data; performing feature screening on the standardized single-cell data to obtain hypervariable genes, wherein the expression level of the hypervariable genes in the target cells is higher than a first expression threshold, and the expression level of the hypervariable genes in other cells is lower than the first expression threshold; normalizing the single-cell data and the hypervariable genes to obtain a cell gene matrix; determining a target multi-classification model corresponding to the target organ of the target plant; and inputting the cell gene matrix into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model. This invention enables efficient and accurate prediction of cell types for multiple organs of various plants, and has good economic value and application prospects.
Claims
1. A plant cell type prediction method based on interpretable machine learning, characterized by, include: Obtain single-cell data of the target organ of the target plant; The single-cell data is preprocessed to obtain standardized single-cell data; Feature screening is performed on the standardized single-cell data to obtain hypervariable genes from the single-cell data, wherein the expression level of the hypervariable gene in the target cell is higher than a first expression threshold, and the expression level of the hypervariable gene in other cells is lower than the first expression threshold. Based on the single-cell data and the highly variable genes, a cell gene matrix is obtained by normalization. Determine the target multi-classification model corresponding to the target organ of the target plant; The cell gene matrix is input into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model.
2. The plant cell type prediction method based on interpretable machine learning according to claim 1, characterized in that, After inputting the cell gene matrix into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model, the method further includes: SHAP value analysis was performed on the target multi-classification model to obtain specific marker genes corresponding to the predicted cell types. 3.The plant cell type prediction method based on interpretable machine learning according to claim 1, characterized in that, The preprocessing of the single-cell data to obtain standardized single-cell data includes: Determine the raw single-cell RNA sequencing data of the single-cell data; The raw single-cell RNA sequencing data is subjected to quality control processing to obtain filtered single-cell RNA sequencing data. The quality control processing includes: removing low-expression genes and removing mitochondrial genes, wherein the expression level of the low-expression genes is lower than a second expression threshold. The sequencing depth and gene expression levels of the filtered single-cell RNA sequencing data are standardized to obtain standardized single-cell data. 4.The plant cell type prediction method based on interpretable machine learning according to claim 1, wherein, Prior to acquiring single-cell data of the target organ of the target plant, the method further includes: Obtain single-cell data of different organs from multiple plant species; The single-cell data is preprocessed and annotated to obtain the cell type label for each single-cell data; Based on the cell type label, feature filtering is performed on the single-cell data to obtain a target number of hypervariable genes for each single-cell data. Based on the data of each single cell and the hypervariable genes of each single cell, the cell gene matrix of each single cell is obtained by normalization. The cell gene matrix of each single cell is classified according to different organs of the multiple plant species to obtain a training dataset for each organ of each plant species. The training dataset includes: the cell gene matrix and cell type labels corresponding to the cell gene matrix. Based on the training dataset and target machine learning algorithm for each organ of each plant species, a multi-classification model corresponding to each organ of each plant species is constructed.
5. The plant cell type prediction method based on interpretable machine learning according to claim 4, characterized in that, Before constructing a multi-classification model corresponding to each organ of each plant species based on the training dataset and target machine learning algorithm for each organ of each plant species, the method further includes: Obtain a test dataset for plant cell type prediction, wherein the test dataset includes a test cell gene matrix and cell type labels corresponding to the test cell gene matrix; The test cell gene matrix is input into multiple preset machine learning algorithms to obtain the test predicted cell type and test time output by each machine learning algorithm. The test predicted cell type output by each machine learning algorithm is compared with the cell type label to obtain the test accuracy of each machine learning algorithm; The target machine learning algorithm is determined based on the test time, test accuracy, and model robustness of each machine learning algorithm.
6. The plant cell type prediction method based on interpretable machine learning according to claim 4, characterized in that, The target machine learning algorithm is the LightGBM algorithm.
7. A plant cell type prediction apparatus based on interpretable machine learning, characterized by, include: The acquisition module is used to acquire single-cell data of the target organ of the target plant; The preprocessing module is used to preprocess the single-cell data to obtain standardized single-cell data; A screening module is used to perform feature screening on the standardized single-cell data to obtain hypervariable genes in the single-cell data, wherein the expression level of the hypervariable gene in the target cell is higher than a first expression threshold, and the expression level of the hypervariable gene in other cells is lower than the first expression threshold. The normalization module is used to normalize the single-cell data and the hypervariable genes to obtain the cell gene matrix. A determination module is used to determine the target multi-classification model corresponding to the target organ of the target plant; The prediction module is used to input the cell gene matrix into the target multi-classification model to obtain the predicted cell type output by the target multi-classification model.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the plant cell type prediction method based on interpretable machine learning as described in any one of claims 1 to 6. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the plant cell type prediction method based on interpretable machine learning as described in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the plant cell type prediction method based on interpretable machine learning as described in any one of claims 1 to 6.