Cell proportion determination method and device, computer device, readable storage medium and program product
Patent Information
- Application Number
- CN202510310195.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2026-09-15
AI Technical Summary
然而,来源不同的参考数据存在批次效应,最终计算得到的推断结果可能存在很大不同,参考数据的质量对细胞比例推断结果的准确性产生较大影响
[0053]The aforementioned cell proportion determination method, apparatus, computer equipment, computer-readable storage medium, and computer program product can acquire a first cell data subset and a second cell data subset. These subsets are obtained by randomly partitioning an integrated cell dataset. The integrated cell dataset includes cell datasets from multiple batches of the target tissue structure and contains multiple cells of different actual cell types. Then, based on the first cell in the first cell data subset, the distribution of characteristic genes for various cell types in the integrated cell dataset is determined. Based on the characteristic gene distribution, the predicted cell type of the second cell in the second cell data subset is determined. From these second cells, target cells whose predicted cell types match the actual cell types are selected. Finally, based on each target cell, the cell proportion of each cell type in the target tissue structure is determined. In this embodiment, by determining the distribution of characteristic genes reflecting cell types in different batches of data based on the first cell, determining the predicted cell type of the second cell based on this distribution, and selecting target cells whose predicted cell types match the actual cell types, representative cells from multiple batches of cell datasets can be quickly screened for each cell type, reducing batch effects between different cell datasets and thus improving the accuracy of the cell proportion determined based on multiple cell datasets.
Smart Images

Figure CN122761998A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for determining cell proportions. Background Technology
[0002] Abnormal pathological changes in biological tissues are often accompanied by changes in the proportion of cells in the tissue. By determining the process of these changes, we can help to reveal the molecular mechanisms of abnormal pathological changes in tissues and identify the targets of related chemical substances in the tissues.
[0003] In related technologies, reference data and relevant cell proportion determination models can be used to determine the proportion of different cell types in a tissue. However, reference data from different sources exhibits batch effects, which can lead to significant differences in the final calculated inference results. The quality of the reference data has a substantial impact on the accuracy of the cell proportion inference results. Therefore, it is evident that related technologies struggle to obtain accurate and reliable cell proportions based on reference data from diverse sources. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for determining cell proportions to improve the accuracy of cell proportion determination, in order to address the above-mentioned technical problems.
[0005] In a first aspect, this application provides a method for determining cell proportions, including:
[0006] Obtain a first subset of cell data and a second subset of cell data; the first subset of cell data and the second subset of cell data are obtained by randomly dividing the integrated cell dataset, wherein the integrated cell dataset includes cell datasets from multiple batches of the target tissue structure, and the integrated cell dataset includes multiple cells of different actual cell types;
[0007] Based on the first cell in the first cell data subset, determine the distribution of characteristic genes of various cell types in the integrated cell dataset;
[0008] Based on the distribution of the characteristic genes, the predicted cell type of the second cell in the second cell data subset is determined, and target cells that match the predicted cell type with the actual cell type are obtained from multiple second cells;
[0009] Based on each of the target cells, determine the proportion of each cell type in the target tissue structure.
[0010] In one embodiment, determining the distribution of characteristic genes for various cell types in the integrated cell dataset based on the first cell in the first cell data subset includes:
[0011] Based on the first cell in the first cell data subset, determine the distinguishability of the cell types by the characteristic genes associated with each cell type;
[0012] Based on the distinguishability of the cell types by the characteristic genes associated with each cell type, the distribution of characteristic genes of various cell types in the integrated cell dataset is determined.
[0013] In one embodiment, determining the predicted cell type of the second cell in the second cell data subset based on the distribution of the characteristic genes includes:
[0014] The trained cell classification model determines the probability that the second cell belongs to each of the cell types based on the cell genes contained in the second cell and the distinguishability of the cell types associated with the various cell types.
[0015] The predicted cell type of the second cell is determined based on the classification probability of the second cell belonging to each of the aforementioned cell types.
[0016] In one embodiment, obtaining target cells from a plurality of second cells whose predicted cell type matches the actual cell type of the second cells includes:
[0017] From the plurality of second cells, identify a plurality of type-matching cells whose predicted cell type is the same as the actual cell type of the second cells;
[0018] For each cell type, select matching cells belonging to that cell type, and among the selected matching cells, select cells whose target classification probability satisfies the probability condition as target cells; the target classification probability is the classification probability of the matching cell belonging to that cell type.
[0019] Multiple target cells are obtained based on the target cells of each described cell type.
[0020] In one embodiment, the actual cell type of the cells in the integrated cell dataset is determined by the following steps:
[0021] Obtain multiple batches of cell datasets corresponding to the target tissue structure;
[0022] Batch effect correction is performed on the multiple batches of cell datasets, and cells in the batch-corrected cell datasets are clustered according to cell characteristics to obtain multiple cell clusters;
[0023] For each cell cluster, characteristic genes of cells in the cell cluster are determined, and reference cell annotations of cells in the cell cluster are determined based on the characteristic genes of the cell cluster. The actual cell type of cells in the cell cluster is determined based on the reference cell annotations.
[0024] In one embodiment, the cell ratio is determined based on the target cells and a target cell ratio determination model; the target cell ratio determination model is obtained through the following steps:
[0025] Obtain multiple candidate cell proportions to determine the model;
[0026] Based on the test results of the model on the sample cell dataset according to the cell proportions of each candidate, the performance index information of the cell proportion determination model of each candidate is determined; the sample cell dataset includes: a first cell dataset with the same data features as at least one batch of cell datasets, and a second cell dataset with different data features from at least one batch of cell datasets;
[0027] From the multiple candidate cell proportion determination models, the target cell proportion determination model whose performance index information meets the performance conditions is determined.
[0028] In one embodiment, determining the performance metrics of each candidate cell proportion determination model based on the test results of the model on the sample cell dataset according to the cell proportions of each candidate model includes:
[0029] Based on the multiple candidate cell proportion determination models and the first cell dataset, determine the first performance index information of each candidate cell proportion determination model;
[0030] Cell proportion determination models that meet the performance conditions based on the first performance index information are selected. Based on the selected cell proportion determination models and the second cell dataset, the second performance index information of the selected cell proportion determination models is determined.
[0031] The step of determining the target cell proportion determination model from multiple candidate cell proportion determination models, wherein the performance index information satisfies the performance conditions, includes:
[0032] From the selected cell proportion determination models, the model whose second performance index information meets the performance conditions is selected as the target cell proportion determination model.
[0033] Secondly, this application also provides a cell ratio determination device, comprising:
[0034] The dataset acquisition module is used to acquire a first cell data subset and a second cell data subset; the first cell data subset and the second cell data subset are obtained by randomly dividing the integrated cell dataset, wherein the integrated cell dataset includes cell datasets from multiple batches of the target tissue structure, and the integrated cell dataset includes multiple cells of different actual cell types;
[0035] The gene pattern determination module is used to determine the distribution of characteristic genes of various cell types in the integrated cell dataset based on the first cell in the first cell data subset.
[0036] The target cell screening module is used to determine the predicted cell type of the second cell in the second cell data subset based on the distribution of the characteristic genes, and to obtain target cells from multiple second cells whose predicted cell type matches the actual cell type.
[0037] The cell proportion determination module is used to determine the cell proportion of each cell type in the target tissue structure based on each target cell.
[0038] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0039] Obtain a first subset of cell data and a second subset of cell data; the first subset of cell data and the second subset of cell data are obtained by randomly dividing the integrated cell dataset, wherein the integrated cell dataset includes cell datasets from multiple batches of the target tissue structure, and the integrated cell dataset includes multiple cells of different actual cell types;
[0040] Based on the first cell in the first cell data subset, determine the distribution of characteristic genes of various cell types in the integrated cell dataset;
[0041] Based on the distribution of the characteristic genes, the predicted cell type of the second cell in the second cell data subset is determined, and target cells that match the predicted cell type with the actual cell type are obtained from multiple second cells;
[0042] Based on each of the target cells, determine the proportion of each cell type in the target tissue structure.
[0043] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0044] Obtain a first subset of cell data and a second subset of cell data; the first subset of cell data and the second subset of cell data are obtained by randomly dividing the integrated cell dataset, wherein the integrated cell dataset includes cell datasets from multiple batches of the target tissue structure, and the integrated cell dataset includes multiple cells of different actual cell types;
[0045] Based on the first cell in the first cell data subset, determine the distribution of characteristic genes of various cell types in the integrated cell dataset;
[0046] Based on the distribution of the characteristic genes, the predicted cell type of the second cell in the second cell data subset is determined, and target cells that match the predicted cell type with the actual cell type are obtained from multiple second cells;
[0047] Based on each of the target cells, determine the proportion of each cell type in the target tissue structure.
[0048] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0049] Obtain a first subset of cell data and a second subset of cell data; the first subset of cell data and the second subset of cell data are obtained by randomly dividing the integrated cell dataset, wherein the integrated cell dataset includes cell datasets from multiple batches of the target tissue structure, and the integrated cell dataset includes multiple cells of different actual cell types;
[0050] Based on the first cell in the first cell data subset, determine the distribution of characteristic genes of various cell types in the integrated cell dataset;
[0051] Based on the distribution of the characteristic genes, the predicted cell type of the second cell in the second cell data subset is determined, and target cells that match the predicted cell type with the actual cell type are obtained from multiple second cells;
[0052] Based on each of the target cells, determine the proportion of each cell type in the target tissue structure.
[0053] The aforementioned cell proportion determination method, apparatus, computer equipment, computer-readable storage medium, and computer program product can acquire a first cell data subset and a second cell data subset. These subsets are obtained by randomly partitioning an integrated cell dataset. The integrated cell dataset includes cell datasets from multiple batches of the target tissue structure and contains multiple cells of different actual cell types. Then, based on the first cell in the first cell data subset, the distribution of characteristic genes for various cell types in the integrated cell dataset is determined. Based on the characteristic gene distribution, the predicted cell type of the second cell in the second cell data subset is determined. From these second cells, target cells whose predicted cell types match the actual cell types are selected. Finally, based on each target cell, the cell proportion of each cell type in the target tissue structure is determined. In this embodiment, by determining the distribution of characteristic genes reflecting cell types in different batches of data based on the first cell, determining the predicted cell type of the second cell based on this distribution, and selecting target cells whose predicted cell types match the actual cell types, representative cells from multiple batches of cell datasets can be quickly screened for each cell type, reducing batch effects between different cell datasets and thus improving the accuracy of the cell proportion determined based on multiple cell datasets. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0055] Figure 1 This is a flowchart illustrating a method for determining cell proportions in one embodiment;
[0056] Figure 2 This is a flowchart illustrating a step in screening target cells in one embodiment;
[0057] Figure 3a This is a schematic diagram of the test results of one model in one embodiment;
[0058] Figure 3b This is a schematic diagram illustrating the test results of another model in one embodiment;
[0059] Figure 4 This is a flowchart illustrating another method for determining cell proportions in one embodiment;
[0060] Figure 5This is a schematic diagram of the framework of a cell ratio determination method in one embodiment;
[0061] Figure 6 This is a structural block diagram of a cell ratio determination device in one embodiment;
[0062] Figure 7 This is an internal structural diagram of a computer device in one embodiment;
[0063] Figure 8 This is an internal structural diagram of another computer device in one embodiment. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0065] To enable those skilled in the art to better understand this application, the relevant technologies are first introduced below.
[0066] Abnormal pathological changes in biological tissues are often accompanied by changes in the proportion of cells in the tissue. By determining the process of these changes, we can help to reveal the molecular mechanisms of abnormal pathological changes in tissues and identify the targets of related chemical substances in the tissues.
[0067] Traditional bulk RNA sequencing (BRNA-seq) technology has accumulated a wealth of data in large databases such as The Cancer Genome Atlas (TCGA) and Genotype-Tissue Expression (GTEx), covering a range of important abnormal tissue samples. Among related techniques, flow cytometry processing of BRNA-seq data or processing of single-cell datasets can yield reference data for the tissue samples. Furthermore, relevant cell proportion determination models can be used to analyze and infer changes in cell proportions and related gene expression regulation during the progression of abnormal tissue pathologies. In practice, it has been found that cell proportion determination models based on reference data demonstrate better inference performance compared to other cell proportion determination models.
[0068] However, batch effects exist in reference data from different sources, which can lead to significant differences in the final calculated inferences. The quality of the reference data severely affects the accuracy of cell proportion inference, making the acquisition of reliable and high-quality reference data crucial for cell proportion inference. Batch effects refer to variables introduced during sample processing due to technical factors. These variables are unrelated to the biological variables being studied and have no actual biological significance, but can still have a significant impact on the research results. In some examples, batch effects may originate from multiple stages of the experimental procedure, such as sample collection, processing, library construction, and sequencing. Variations in reagents, equipment, operators, or experimental conditions used between different batches can all lead to batch effects.
[0069] Taking the case where the proportion of cells in a Bulk RNA-seq sample tissue is unknown as an example, it is often difficult to determine whether the inferred cell proportion is reliable. For example, if the actual proportion of cell type A in the sample tissue is 10% (assuming this proportion is unknown), when using reference data from different sources or batches for inference, the obtained proportion of cell type A may be 20% or 30%, etc., which may deviate from the correct cell proportion and result in significant differences in the inference results.
[0070] It is evident that the relevant technologies struggle to obtain accurate and reliable cell proportions based on reference data from different sources.
[0071] Therefore, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for determining cell ratios to address the aforementioned technical problems.
[0072] In one embodiment, such as Figure 1 As shown, a method for determining cell ratios is provided. This embodiment illustrates the application of this method to a server. It is understood that this method can also be applied to a terminal, or to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0073] S101, Obtain the first cell data subset and the second cell data subset; The first cell data subset and the second cell data subset are obtained by randomly dividing the integrated cell dataset. The integrated cell dataset includes cell datasets from multiple batches of the target tissue structure and includes multiple cells of different actual cell types.
[0074] The target tissue structure can be a tissue structure for which cell proportion analysis is to be performed, such as one or more of cells, organs, or systems.
[0075] In practice, multiple batches of cell datasets targeting a specific tissue structure can be acquired. These multiple batches of cell datasets can refer to multiple cell datasets with differing data content; that is, for the same target tissue structure, multiple cell datasets can represent relevant information about that tissue structure, and the data content of these multiple cell datasets differs. For example, multiple batches of cell datasets can include cell datasets from different sources, such as those provided by different laboratories or other institutions, or cell datasets from the same source but provided at different times, for example, the same data provider providing cell datasets for the target tissue structure in June and November. In some examples, multiple batches of cell datasets can also be reference data from different batches.
[0076] Cell datasets can characterize the cell types contained in a target tissue structure and the characteristic genes of each cell type. Specifically, the data in a cell dataset can record the various cell types present in the target tissue structure and the number of cells of each cell type. The number of cells collected for different cell types can be relatively high and balanced (e.g., the number of cells of each cell type reaches a first threshold and the difference in the number of cells of different cell types is less than a second threshold). In addition, for each cell type, characteristic genes whose gene expression changes significantly among various cell types can be recorded. In some examples, the cell dataset can be a single-cell dataset, which refers to a collection of cell data obtained through single-cell sequencing technology, such as single-cell RNA sequencing (scRNA-seq) or single-cell ATAC sequencing (scATAC-seq). A single-cell dataset records the gene expression of a single cell. Taking a single-cell dataset as reference data, the reference data can record each cell in the target tissue structure and the gene expression of each cell in the form of a feature matrix. In the feature matrix, rows can represent genes or gene expression levels, and columns can represent individual cells.
[0077] After obtaining multiple batches of cell datasets from the target tissue, these batches can be integrated and processed to obtain an integrated cell dataset. It can be understood that the target tissue structure includes cells of various cell types, and the integrated cell dataset also includes multiple cells of different cell types. For ease of distinction, the actual cell type to which each cell belongs is referred to as the actual cell type of that cell.
[0078] While improving the accuracy of cell proportion inference results in tissues often involves improvements to the algorithm itself, the inventors discovered that batch-to-batch variations in reference data significantly impact the accuracy of cell proportion inference. To address this, this application optimizes the cell data used in the cell proportion inference process by integrating cell datasets from multiple batches.
[0079] In this embodiment, after obtaining the integrated cell dataset, the cell dataset can be randomly divided into different cell data subsets. The subset used to determine or identify common information across multiple batches of cell datasets is called the first cell data subset, and the subset used to verify the identified common information is called the second cell data subset. By randomly dividing the integrated cell dataset, the cell distribution in the first and second cell data subsets can match the cell distribution in the integrated cell dataset, preventing cells from being overly concentrated in a particular cell data subset and improving the accuracy of the identified common information and the accuracy of the verification results.
[0080] In some examples, the first subset of cell data can be called the training set, and the second subset of cell data can be called the test set. It is understood that the first and / or second subsets of cell data can be one or more. For example, one first subset of cell data and multiple second subsets of cell data can be set to perform multiple verifications based on the multiple second subsets. Alternatively, multiple first subsets of cell data and one second subset of cell data can be set, allowing for the identification of common information and mutual verification of identification results based on multiple first subsets, thereby improving the accuracy of common information identification. Of course, multiple first subsets of cell data and multiple second subsets of cell data can also be set.
[0081] S102, Based on the first cell in the first cell data subset, determine the distribution of characteristic genes of various cell types in the integrated cell dataset.
[0082] For ease of distinction, each cell in the first cell data subset is referred to as a first cell. In this step, since the first cell data subset records the actual cell type of each first cell and the characteristic genes of each first cell, by analyzing the first cells in the first cell data subset, the distribution of characteristic genes of various cell types in the first cell data subset can be determined. The distribution of characteristic genes can characterize the characteristic genes possessed by various cell types and the degree of distribution of each characteristic gene in various cell types.
[0083] Since the first subset of cell data is obtained by randomly partitioning the integrated cell dataset, the distribution of characteristic genes determined based on the first cell can be considered as the distribution of characteristic genes for various cell types in the integrated cell dataset. It can be understood that in this step, by analyzing the first cell in the first subset of cell data, the common characteristic gene distribution across multiple batches of cell datasets can be identified, and the common characteristic gene distribution pattern across multiple batches of cell datasets can be determined through the first subset of cell data. For each cell dataset in multiple batches of cell datasets, it can be assumed that the various cell types in that cell dataset satisfy this characteristic gene distribution.
[0084] S103. Based on the distribution of characteristic genes, determine the predicted cell type of the second cell in the second cell data subset, and obtain the target cell from multiple second cells whose predicted cell type matches the actual cell type.
[0085] For ease of distinction, the cells in the second cell data subset will be referred to as the second cell in this step.
[0086] Specifically, after obtaining the characteristic gene distribution of various cell types, since the characteristic gene distribution of each cell type is different and the characteristic gene distribution of different cell types is different, in this step, the cell type of multiple second cells can be predicted based on the characteristic gene distribution, and the cell type of each of the multiple second cells can be obtained. For easy differentiation, the cell type predicted based on the characteristic gene distribution is called the predicted cell type.
[0087] Then, cells whose predicted cell type matches the actual cell type can be selected from multiple second cells and designated as target cells. Since the distribution of characteristic genes reflects the common characteristic gene distribution across multiple batches of cell datasets, this step allows for the selection of target cells that meet the characteristic gene distribution criteria from multiple second cells. This enables the selection of representative cells through self-consistency. Specifically, for a given cell type, when a cell meets the characteristic gene distribution criteria for that cell type, because the characteristic gene distribution is common information across multiple batches of cell datasets, the cell will not only be identified as a cell of that cell type in its own batch of cell datasets but also as a matching cell type in other batches of cell datasets. Furthermore, because the predicted cell type matches the actual cell type, the known cell type information in the corresponding batch of cell datasets and the cell type information predicted based on the characteristic gene distribution can be cross-validated, ensuring the reliability and accuracy of the final cell type determined for the target cell. This effectively and accurately improves the representativeness of the final target cell across multiple batches of cell datasets and reduces batch effects.
[0088] S104, based on each target cell, determines the proportion of each cell type in the target tissue structure.
[0089] After obtaining the target cells, their relevant data can be used as reference data. Then, based on the target cells and a method using the reference data, the proportion of each cell type in the target tissue structure can be determined. The reference data-based method refers to using expression data of known cell types to infer the cell proportions. In some examples, deconvolution algorithms and the expression data of the target cells can be used to determine the proportion of each cell type in the target tissue structure. Of course, those skilled in the art can also use other reference data-based methods to determine the cell proportions.
[0090] In the aforementioned method for determining cell proportions, a first subset of cell data and a second subset of cell data can be obtained. These subsets are obtained by randomly partitioning an integrated cell dataset. The integrated cell dataset includes cell datasets from multiple batches of the target tissue structure and contains multiple cells of different actual cell types. Then, based on the first cell in the first subset of cell data, the distribution of characteristic genes for various cell types in the integrated cell dataset is determined. Based on the distribution of characteristic genes, the predicted cell type of the second cell in the second subset of cell data is determined. From these second cells, target cells whose predicted cell types match the actual cell types are selected. Finally, based on each target cell, the cell proportions of each cell type in the target tissue structure are determined. In this embodiment, by determining the distribution of characteristic genes reflecting cell types in different batches of data based on the first cell, determining the predicted cell type of the second cell based on this distribution, and selecting target cells whose predicted cell types match the actual cell types, representative cells from multiple batches of cell datasets can be quickly screened for each cell type, reducing batch effects between different cell datasets and thus improving the accuracy of cell proportions determined based on multiple cell datasets.
[0091] In one embodiment, in step S102, determining the distribution of characteristic genes of various cell types in the integrated cell dataset based on the first cell in the first cell data subset may include the following steps:
[0092] Based on the first cell in the first cell data subset, determine the distinguishability of cell types by the characteristic genes associated with each cell type; based on the distinguishability of cell types by the characteristic genes associated with each cell type, determine the distribution of characteristic genes of each cell type in the integrated cell dataset.
[0093] Specifically, each first cell in the first cell data subset can carry a cell annotation indicating the actual cell type of the first cell; for ease of distinction, this cell annotation is also called the actual cell annotation. The first cell belonging to each cell type can be determined through each actual cell annotation.
[0094] Furthermore, based on the first cell of each cell type, the distinguishing power of characteristic genes associated with each cell type can be determined. These characteristic genes, also known as marker genes, are genes specifically expressed in a known cell type; they can be understood as marker genes used to define that cell type. Identifying one or more characteristic genes carried by a cell helps to determine the cell type. Characteristic genes may be expressed in cells of one or more specific cell types, while their expression levels are low or absent in other cell types.
[0095] It is understandable that when the same characteristic gene can be expressed by multiple cell types, the distinguishing power of the characteristic gene in defining a certain cell type may differ. Therefore, in this embodiment, for each cell type, the distinguishing power of the Miyi characteristic gene in that cell type can be determined. The higher the distinguishing power of the characteristic gene, the more effectively it can characterize that cell type. The degree of distinguishing effect of the same characteristic gene varies in different cell types. For example, the CD3D and CD3E genes in T cells can effectively characterize T cells, so the distinguishing power of CD3D and CD3E genes in T cells is greater than their distinguishing power in other cell types.
[0096] Furthermore, the distribution of characteristic genes for various cell types in the integrated cell dataset can be determined by combining the discriminative power of characteristic genes associated with different cell types. In some embodiments, a pre-trained machine learning model can be used to determine the distribution of characteristic genes for each cell type. For example, the discriminative power of characteristic genes for each cell type can be determined by a logistic regression model. In one example, this discriminative power is also referred to as the weight of the characteristic genes. The logistic regression model can obtain a weighted gene list, which includes the characteristic genes for each cell type and the weight of each characteristic gene for each cell type.
[0097] In this embodiment, by using the first cell in the first cell data subset, the characteristic genes of various cell types in the integrated cell dataset and their distinguishability to cell types can be determined, and the distribution of characteristic genes can be determined accordingly. This helps to more accurately characterize cell types and identify the cell types of each second cell in the second cell data subset, thereby more accurately determining representative cells from the second cell dataset that can represent the characteristics of a certain cell type in each batch of cell datasets.
[0098] In one embodiment, determining the predicted cell type of a second cell in a second cell subset based on the distribution of characteristic genes may include the following steps:
[0099] The trained cell classification model determines the probability of the second cell belonging to each cell type based on the cell genes contained in the second cell and the distinguishability of cell types by characteristic genes associated with various cell types; based on the probability of the second cell belonging to each cell type, the predicted cell type of the second cell is determined.
[0100] In a specific implementation, a cell classification model can be pre-trained to predict the cell type based on the cellular genes contained in the cell and the pre-obtained distinguishing feature genes associated with various cell types. In one embodiment, the cell classification model can perform correlation analysis between the second cell and various cell types based on the gene expression data of cellular genes in the second cell and the analyzed distinguishing feature genes associated with various cell types, calculate the correlation coefficient between the second cell and various cell types, and determine the classification probability based on the correlation coefficient. The classification probability can be positively correlated with the correlation coefficient.
[0101] In some examples, the cell classification model can be a machine learning model, such as a Support Vector Machine (SVM), a convolutional neural network, a Bayesian classification model, or a logistic regression model. Taking the SVM model as an example, in a cell classification task, the SVM model can classify cells into different types based on cell gene expression data (such as the expression data of each cell gene in a second cell). Through training, the SVM model can learn the characteristic gene distribution of different cell types and classify new cells based on this information, outputting the classification probability of each cell type. For example, if the target tissue structure has N (N≥2) cell types, then for each second cell, the cell classification model can output N classification probabilities for that second cell. Each classification probability can represent the probability that the second cell belongs to a certain cell type.
[0102] In some optional embodiments, for each second cell, the cell classification model can assign a cell annotation representing the predicted cell type based on the classification probability of the second cell belonging to various cell types. For ease of distinction, in this embodiment, the cell annotation determined based on the classification probability is referred to as the predicted cell annotation. For example, if a second cell is predicted to belong to a T cell based on various classification probabilities, a predicted cell annotation containing characteristic genes such as "CD3D" and "CD3E" can be assigned to the second cell (or "T cell" can be directly used as the predicted cell annotation), thereby representing the predicted cell type of the second cell.
[0103] In this embodiment, a trained cell classification model determines the classification probability of a second cell based on the cell genes contained in the second cell and the distinguishability of cell types by characteristic genes associated with various cell types. Then, the corresponding predicted cell type is determined. This method can quickly predict the cell type of each second cell based on the characteristic gene expression patterns summarized from multiple batches of cell datasets, thereby improving the screening efficiency of target cells.
[0104] In one embodiment, such as Figure 2 As shown, in step S103, obtaining the target cell from multiple second cells whose predicted cell type matches the actual cell type of the second cells may include the following steps:
[0105] S201, from multiple second cells, identify multiple type-matching cells whose predicted cell type is the same as the actual cell type of the second cells.
[0106] In an exemplary embodiment, each cell in the integrated cell dataset may have a pre-defined cell annotation, which is referred to as a reference cell annotation for easy differentiation from other cell annotations. After obtaining multiple second cells, the actual cell type of each second cell can be determined based on the reference cell annotation carried by each second cell. This allows for the identification of second cells whose predicted cell type matches the actual cell type. For ease of differentiation, these second cells are referred to as type-matching cells.
[0107] S202, for each cell type, select matching cells belonging to each cell type, and among the selected matching cells, select the cells whose target classification probability meets the probability condition as target cells; the target classification probability is the classification probability of the matching cells belonging to the cell type.
[0108] It is understandable that after identifying multiple matching cells, these cells can include cells from at least one cell type. For each cell type, we can identify individual matching cells belonging to that cell type from the multiple matching cells and obtain the target classification probability of that matching cell. Specifically, during cell classification, we can obtain the classification probability of each second cell belonging to each cell type. In this step, for a matching cell of a certain cell type, we can determine the classification probability of that cell type belonging to that cell type, i.e., the target classification probability, from the classification probabilities of that matching cell belonging to each cell type. Taking matching cell 1 of the T cell type as an example, during cell classification, its probability of belonging to the T cell type is 85%, its probability of belonging to the B cell type is 5%, and its probability of belonging to the macrophage type is 10%. Therefore, the probability of 85% can be used as the target classification probability of matching cell 1.
[0109] Furthermore, for each cell type, the target classification probability of each matching cell type can be compared with a preset probability condition. Cells whose target classification probability meets the preset probability condition are selected as target cells, while those whose target classification probability does not meet the preset probability condition are removed. In this way, the target cells for each cell type can be obtained. For example, the m (m≥2) matching cells with the highest target classification probability can be selected as target cells. Alternatively, matching cells with target classification probabilities exceeding a preset threshold can be selected as target cells.
[0110] S203, based on the target cells of each cell type, yields multiple target cells.
[0111] For example, each target cell of each cell type can be used as multiple target cells for subsequent determination of cell proportions.
[0112] In this embodiment, by identifying multiple type-matching cells, and then selecting cells whose target classification probability meets the probability condition for each type-matching cell under each cell type as target cells, it is possible to prioritize type-matching cells with higher credibility as target cells under the corresponding cell type, effectively improving the feasibility of target cells and optimizing the quality of reference data.
[0113] In one embodiment, prior to step S101, the actual cell type of the cells in the integrated cell dataset can be determined through the following steps:
[0114] S11, Obtain multiple batches of cell datasets corresponding to the target tissue structure.
[0115] In some examples, the acquired cell datasets from multiple batches can be preprocessed before subsequent steps are performed, such as data cleaning, to perform quality control on the original multiple batches of cell datasets and remove low-quality data.
[0116] S12 performs batch effect correction on multiple batches of cell datasets, and clusters the cells in the batch-corrected cell datasets based on cell characteristics to obtain multiple cell clusters.
[0117] In practical applications, batch effect correction can be performed after obtaining multiple batches of cell datasets. In some exemplary embodiments, one or more methods can be used for correction, such as global models (e.g., ComBat), linear embedding models (e.g., Seurat, Harmony, FastMNN, etc.), graph-based models (e.g., BBKNN), and deep learning models.
[0118] Then, based on cell characteristics, cells in multiple batches of cell datasets after batch correction can be clustered to obtain multiple cell clusters. Cells in the same cell cluster can have the same or similar cell characteristics, while cells in different clusters have different cell characteristics.
[0119] In some embodiments, the Leiden algorithm for graph clustering can be used to cluster multiple batches of cell datasets after batch correction. Specifically, before performing Leiden clustering, a graph can be constructed to represent the relationships between cells. This graph is typically constructed based on similarity measures between cells (such as Euclidean distance, cosine similarity, etc.), where nodes represent cells and edges represent similarity relationships between cells. Then, the constructed graph is used as input, and the Leiden algorithm is applied for clustering. The Leiden algorithm iteratively optimizes the clustering results until a stable state is reached. In each iteration, the algorithm attempts to assign nodes (cells) to different clusters and calculates a measure of clustering quality (such as modularity). Finally, the Leiden algorithm outputs a clustering result where each cell is assigned to a cell cluster. Each cell cluster represents a different cell population or type from multiple batches of cell datasets.
[0120] S13. For each cell cluster, determine the characteristic genes of the cells in the cell cluster, determine the reference cell annotation of the cells in the cell cluster based on the characteristic genes of the cell cluster, and determine the actual cell type of the cells in the cell cluster based on the reference cell annotation.
[0121] After obtaining multiple cell clusters, for each cell cluster, the characteristic genes of each cell in the cluster can be determined based on the original cell annotations carried by each cell. The original cell annotations are the cell annotations for the cells in the batch of cell datasets to which the cells belong. Then, based on the typical characteristic genes in the cell cluster, a reference cell annotation is determined for each cell in that cluster. Each cell in the cluster is renamed according to this reference cell annotation, and the actual cell type of the cells in the cluster is identified based on the renamed reference cell annotations.
[0122] In practice, data from individual cells is often very complex and may contain a lot of noise. If cell annotation is performed directly based on the expression of a single cell, the annotation results may be affected by noise, leading to incorrect cell annotations. In this embodiment, after batch correction of multiple batches of cell datasets, clustering is performed to group cells with similar cell characteristics from different batches of cell datasets into one category. Subsequently, reference cell annotations for cells in the clusters are determined based on the characteristic genes in the cell clusters. This helps to accurately annotate cells based on the typical characteristic genes of multiple cells in the same cluster, reducing the impact of noise caused by individual cells. At the same time, clustering can enhance the cell recognition ability of the subsequently generated reference cell annotations. Since the similarity between cells (such as gene expression patterns) is often more obvious within the same cluster, clustering before generating reference cell annotations can help find the potential structures between cell types, identify similar or identical characteristic genes in the same cell cluster, and identify characteristic genes that differ between different cell clusters. This allows for more accurate generation of reference cell annotations for each cell type, effectively optimizing the cell annotations for each cell.
[0123] In one embodiment, the cell ratio is determined based on multiple target cells and a pre-defined target cell ratio determination model. The target cell ratio determination model is obtained through the following steps:
[0124] S21, Obtain multiple candidate cell proportion determination models.
[0125] The candidate cell proportion determination model can be any model that determines the cell proportion based on reference data. In some examples, multiple candidate cell proportion determination models may include various deconvolution algorithms, such as one or more of the following: Damped weighted least squares (DWLS), Fast and Robust Deconvolution of Expression Profiles (FARDEEP), Multi-Subject Single Cell Deconvolution (MuSiC), Non-negative least squares (nnls), Robust linear regression (RLR), and Epigenetic Dissection of Intra-Sample Heterogeneity (EpiDISH).
[0126] S22, determine the test results of the model on the sample cell dataset based on the cell proportions of each candidate, and determine the performance index information of the model for determining the cell proportions of each candidate; the sample cell dataset includes: a first cell dataset with data features that are the same as those of at least one batch of cell datasets, and a second cell dataset with data features that are different from those of at least one batch of cell datasets.
[0127] Specifically, after obtaining multiple candidate cell proportion determination models, various models and sample cell datasets can be used to test the models under self-reference and cross-reference conditions (such as model benchmarking).
[0128] Specifically, the sample cell dataset includes a first cell dataset and a second cell dataset. The first cell dataset refers to a cell dataset whose data characteristics are the same as those of at least one batch of cell datasets. For example, a cell dataset with the same source or experimental conditions as at least one batch of cell datasets can be used as the first cell dataset. The second cell dataset refers to a second cell dataset whose data characteristics are different from those of at least one batch of cell datasets. For example, a cell dataset with a different source or experimental conditions than at least one batch of cell datasets can be used as the second cell dataset.
[0129] By testing various candidate cell proportion determination models using the first and second cell datasets, we can obtain the model performance and corresponding performance metrics for multiple candidate cell proportion determination models facing similar and different cell types. The performance metrics reflect the accuracy of the candidate cell proportion determination models. In some examples, these metrics may include the Pearson correlation coefficient and the root mean square error (RMSE). The Pearson correlation coefficient measures the strength of the linear relationship between two variables (the input and output of the cell proportion determination model). The RMSE assesses the model's prediction error; it represents the square root of the mean of the squares of the differences between the predicted and true values. A smaller RMSE indicates a smaller prediction error and a prediction result closer to the true value.
[0130] S23, from multiple candidate cell proportion determination models, determine the target cell proportion determination model whose performance index information meets the performance conditions.
[0131] Then, after obtaining the performance index information of multiple candidate cell proportion determination models for the first cell dataset and the second cell dataset, the performance index information of each candidate cell proportion determination model can be compared, and the model whose performance index information meets the performance conditions can be selected as the target cell proportion determination model.
[0132] In this embodiment, on the one hand, by testing with a first cell dataset whose data features are the same as at least one batch of cell datasets, the accuracy of different cell proportion determination models in processing similar data can be verified, ensuring that the subsequently selected target cell proportion determination model can provide reliable results when processing conventional or expected data. On the other hand, by testing with a second cell dataset whose data features are different from at least one batch of cell datasets, the performance of candidate cell proportion determination models on unseen data can be evaluated, which helps to screen out models that can maintain good performance in a wider range of data, ensuring the generalization ability of the selected target cell proportion determination model and effectively improving the reliability of the target cell proportion determination model in the face of different cell data.
[0133] In one embodiment, determining the performance metrics of the model based on the test results of the sample cell dataset according to the cell proportions of each candidate may include the following steps:
[0134] Based on multiple candidate cell proportion determination models and a first cell dataset, the first performance index information of each candidate cell proportion determination model is determined; cell proportion determination models whose first performance index information meets the performance conditions are selected; and based on the selected cell proportion determination models and a second cell dataset, the second performance index information of the selected cell proportion determination models is determined.
[0135] In specific implementation, the first cell data and the second cell data in the sample cell dataset are cell datasets for the target tissue structure. In some examples, the target tissue structure involved in testing the cell ratio determination model may be different from the target tissue structure involved in actually determining the cell ratio. For example, the sample cell dataset may be cell datasets corresponding to one or more different tissue structures, and the integrated cell dataset may be a cell dataset for the tissue structure for which cell ratio identification is actually required.
[0136] In this embodiment, the candidate cell proportions can be tested using the first cell dataset and the second cell dataset respectively to determine the model.
[0137] In an exemplary embodiment, a first cell dataset can be input into each candidate cell proportion determination model to obtain the first cell proportion output by each candidate cell proportion determination model for the target tissue structure, and then the first performance index information of each candidate cell proportion determination model can be determined based on the first cell proportion.
[0138] Subsequently, cell proportion determination models that meet the performance conditions based on the first performance indicator information can be selected. The second cell dataset is then input into the selected cell proportion determination models. Based on the second cell proportion output by the models, the second performance indicator information for each of the selected cell proportion determination models is determined. The indicator type of the second performance indicator information is the same as that of the first performance indicator information.
[0139] Accordingly, determining the target cell proportion determination model from multiple candidate cell proportion determination models that meets the performance conditions can include the following steps:
[0140] From the selected cell proportion determination models, the model whose second performance index information meets the performance conditions is selected as the target cell proportion determination model.
[0141] In one example, 44 simulated first-cell datasets (simulated bulk data) were generated. These first-cell datasets were consistent with the reference data source used in the actual testing process. The model was then tested on 25 deconvolution algorithms and the 44 datasets under self-reference conditions. The test results are as follows: Figure 3a As shown, Figure 3a The x-axis (Studies) corresponds to the labels of 44 datasets, and the y-axis (Methods) corresponds to the names of 25 deconvolution algorithms. The intersection of the x-axis and y-axis indicates the Pearson correlation coefficient and root mean square error of the corresponding model. Then, based on the Pearson correlation coefficient and root mean square error, the six best-performing deconvolution algorithms can be selected, namely DWLS, FARDEEP, MuSiC, nnls, RLR, and EpiDISH.
[0142] Then, 44 simulated second-cell datasets (simulated bulk data) can be generated. These second-cell datasets differ from the reference data used in the actual testing process. The top 6 performing algorithms are then tested under cross-reference conditions, and the results are as follows: Figure 3b As shown, Figure 3b In the equations a to e, the test results correspond to five tissues: pancreas, lung, peripheral blood mononuclear cells (PBMCs), brain, and spleen. Finally, the optimal deconvolution algorithm is selected based on the same evaluation criteria.
[0143] from Figure 3b It is evident that while the performance of the DWLS algorithm alone is superior to other algorithms, it is affected by batch effects. However, the DWLS algorithm, when combined with the optimized reference data from this application (corresponding to...), ... Figure 3b According to Opt. Ref), it can guarantee an accuracy of at least 0.80 in the simulation data.
[0144] In this embodiment, by first screening each candidate cell proportion determination model based on the first cell dataset, and then further screening the cell proportion determination models obtained from the initial screening based on the second cell dataset, it is possible to screen out models that can maintain good performance in a wider range of data while ensuring that the cell proportion determination models perform well on commonly used cell datasets. This saves testing time while ensuring that the selected target cell proportion determination model has a certain generalization ability.
[0145] To enable those skilled in the art to better understand the above steps, the following example illustrates the embodiments of this application, but it should be understood that the embodiments of this application are not limited thereto.
[0146] like Figure 4 As shown, this example includes steps S401 to S409.
[0147] S401, Obtain multiple batches of cell datasets corresponding to the target tissue structure.
[0148] For example, such as Figure 5 As shown, three or more different batches of cell datasets can be obtained.
[0149] S402 performs batch effect correction on multiple batches of cell datasets, and clusters the cells in the batch-corrected cell datasets based on cell characteristics to obtain multiple cell clusters.
[0150] S403: For each cell cluster, determine the characteristic genes of the cells in the cell cluster, determine the reference cell annotation of the cells in the cell cluster based on the characteristic genes of the cell cluster, and determine the actual cell type of the cells in the cell cluster based on the reference cell annotation.
[0151] Specifically, after obtaining multiple cell datasets, these datasets can be integrated to optimize cell annotation. In this example, Harmony is used for batch effect correction, multiple cell datasets with cell annotations are integrated, and then clustered using the Leiden algorithm. The cells in the clustering results are then renamed according to typical marker genes, thus obtaining an integrated cell dataset carrying reference cell annotations.
[0152] S404, randomly divides the integrated cell dataset into a first cell data subset and a second cell data subset.
[0153] like Figure 5 As shown, the integrated cell dataset can be randomly divided into a training set and a test set.
[0154] S405, the logistic regression model obtains a weighted gene list for each cell type based on the first subset of cell data.
[0155] S406, the cell classification model determines the predicted cell type of the second cell in the second cell data subset based on the weighted gene list.
[0156] S407, Identify type-matching cells. For each cell type, select the 100 cells with the highest target classification probability as target cells based on the target classification probability. Obtain subsequent reference data based on the 100 target cells of each cell type.
[0157] S408 uses sample cell data to screen multiple deconvolution algorithms and determine the optimal deconvolution algorithm, such as the DWLS algorithm.
[0158] S409 uses the optimal deconvolution algorithm combined with optimal reference data to infer the proportion of cells in the bulk sample to be tested.
[0159] Compared to previous studies that neglected the impact of reference data quality on cell proportion inference, this embodiment constructs the SCCAF-D framework for inferring cell proportion in test samples by combining optimized reference data with the best deconvolution algorithm. This framework can significantly reduce the impact of batch effects on cell proportion inference. Based on the framework provided in this application, the inventors conducted relevant experiments:
[0160] In Experiment 1, the inventors conducted tests based on bulk RNA-seq data simulated from single-cell data. The tests included 20 simulated datasets from five tissues (pancreas, lung, peripheral blood mononuclear cells, brain, and spleen) under cross-reference conditions. The results showed that the SCCAF-D framework of this application outperformed other algorithms, with a Pearson correlation coefficient exceeding 0.80.
[0161] In Experiment 2, the inventors tested bulk RNA-seq data with known cell proportions and used the SCCAF-D framework to infer cell proportion information in four known cell proportions of real PBMC bulk data. The test results of the SCCAF-D framework were better than other algorithms, with a Pearson correlation coefficient of over 0.80.
[0162] In Experiment 3, the inventors validated the technology in biological scenarios with prior experience, such as real-world applications with known cell proportion changes in types 2 diabetes (T2D) and idiopathic pulmonary fibrosis (IPF). They found that SCCAF-D can restore changes in cell proportions during disease progression. In T2D, the method described in this application showed a decrease in the proportion of beta cells, which was inversely proportional to the concentration of glycated hemoglobin (HbA1c). In IPF, the method described in this application showed that, compared to the control group, the proportions of type I alveolar epithelial cells (AT1) and type II alveolar epithelial cells (AT2) decreased, while the proportions of basal cells (Basal cells) and fibroblasts (Fibroblast cells) increased.
[0163] Experiments demonstrate that the algorithm framework proposed in this application can guarantee an accuracy of over 0.75 using optimal reference data. Furthermore, for large databases such as TCGA and GTEx, it can fully utilize the large amount of valuable samples accumulated previously, realizing the value of the data, exploring changes in cell proportions during abnormal tissue development, and providing effective reference information for studying the occurrence and development of atypical tissue abnormalities.
[0164] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0165] Based on the same inventive concept, this application also provides a cell ratio determination device for implementing the cell ratio determination method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the cell ratio determination device provided below can be found in the limitations of the cell ratio determination method described above, and will not be repeated here.
[0166] In one exemplary embodiment, such as Figure 6 As shown, a cell ratio determination device is provided, comprising:
[0167] The dataset acquisition module 601 is used to acquire a first cell data subset and a second cell data subset; the first cell data subset and the second cell data subset are obtained by randomly dividing the integrated cell dataset, wherein the integrated cell dataset includes cell datasets from multiple batches of the target tissue structure, and the integrated cell dataset includes multiple cells of different actual cell types;
[0168] The gene pattern determination module 602 is used to determine the distribution of characteristic genes of various cell types in the integrated cell dataset based on the first cell in the first cell data subset.
[0169] The target cell screening module 603 is used to determine the predicted cell type of the second cell in the second cell data subset based on the distribution of the characteristic genes, and to obtain target cells from multiple second cells whose predicted cell type matches the actual cell type.
[0170] The cell proportion determination module 604 is used to determine the cell proportion of each cell type in the target tissue structure based on each target cell.
[0171] In one embodiment, the gene pattern determination module 602 is configured to:
[0172] Based on the first cell in the first cell data subset, determine the distinguishability of the cell types by the characteristic genes associated with each cell type;
[0173] Based on the distinguishability of the cell types by the characteristic genes associated with each cell type, the distribution of characteristic genes of various cell types in the integrated cell dataset is determined.
[0174] In one embodiment, the target cell screening module 603 is used for:
[0175] The trained cell classification model determines the probability that the second cell belongs to each of the cell types based on the cell genes contained in the second cell and the distinguishability of the cell types associated with the various cell types.
[0176] The predicted cell type of the second cell is determined based on the classification probability of the second cell belonging to each of the aforementioned cell types.
[0177] In one embodiment, the target cell screening module 603 is used for:
[0178] From the plurality of second cells, identify a plurality of type-matching cells whose predicted cell type is the same as the actual cell type of the second cells;
[0179] For each cell type, select matching cells belonging to that cell type, and among the selected matching cells, select cells whose target classification probability satisfies the probability condition as target cells; the target classification probability is the classification probability of the matching cell belonging to that cell type.
[0180] Multiple target cells are obtained based on the target cells of each described cell type.
[0181] In one embodiment, the dataset acquisition module 601 is configured to:
[0182] Obtain multiple batches of cell datasets corresponding to the target tissue structure;
[0183] Batch effect correction is performed on the multiple batches of cell datasets, and cells in the batch-corrected cell datasets are clustered according to cell characteristics to obtain multiple cell clusters;
[0184] For each cell cluster, characteristic genes of cells in the cell cluster are determined, and reference cell annotations of cells in the cell cluster are determined based on the characteristic genes of the cell cluster. The actual cell type of cells in the cell cluster is determined based on the reference cell annotations.
[0185] In one embodiment, the cell ratio is determined based on the target cells and a target cell ratio determination model; the device further includes a model screening module, the model screening module being used for:
[0186] Obtain multiple candidate cell proportions to determine the model;
[0187] Based on the test results of the model on the sample cell dataset according to the cell proportions of each candidate, the performance index information of the cell proportion determination model of each candidate is determined; the sample cell dataset includes: a first cell dataset with the same data features as at least one batch of cell datasets, and a second cell dataset with different data features from at least one batch of cell datasets;
[0188] From the multiple candidate cell proportion determination models, the target cell proportion determination model whose performance index information meets the performance conditions is determined.
[0189] In one embodiment, the model filtering module is used for:
[0190] Based on the multiple candidate cell proportion determination models and the first cell dataset, determine the first performance index information of each candidate cell proportion determination model;
[0191] Cell proportion determination models that meet the performance conditions based on the first performance index information are selected. Based on the selected cell proportion determination models and the second cell dataset, the second performance index information of the selected cell proportion determination models is determined.
[0192] The step of determining the target cell proportion determination model from multiple candidate cell proportion determination models, wherein the performance index information satisfies the performance conditions, includes:
[0193] From the selected cell proportion determination models, the model whose second performance index information meets the performance conditions is selected as the target cell proportion determination model.
[0194] Each module in the aforementioned cell ratio determination device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0195] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores multiple batches of cell datasets. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a cell proportion determination method.
[0196] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 8As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements a cell ratio determination method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0197] Those skilled in the art will understand that Figure 7 and Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0198] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0199] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0200] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0201] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0202] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0203] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0204] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for determining cell ratio, characterized in that, The method includes: Obtain a first subset of cell data and a second subset of cell data; the first subset of cell data and the second subset of cell data are obtained by randomly dividing the integrated cell dataset, wherein the integrated cell dataset includes cell datasets from multiple batches of the target tissue structure, and the integrated cell dataset includes multiple cells of different actual cell types; Based on the first cell in the first cell data subset, determine the distribution of characteristic genes of various cell types in the integrated cell dataset; Based on the distribution of the characteristic genes, the predicted cell type of the second cell in the second cell data subset is determined, and target cells that match the predicted cell type with the actual cell type are obtained from multiple second cells; Based on each of the target cells, determine the proportion of each cell type in the target tissue structure.
2. The method according to claim 1, characterized in that, The step of determining the distribution of characteristic genes of various cell types in the integrated cell dataset based on the first cell in the first cell data subset includes: Based on the first cell in the first cell data subset, determine the distinguishability of the cell types by the characteristic genes associated with each cell type; Based on the distinguishability of the cell types by the characteristic genes associated with each cell type, the distribution of characteristic genes of various cell types in the integrated cell dataset is determined.
3. The method according to claim 2, characterized in that, The step of determining the predicted cell type of the second cell in the second cell data subset based on the distribution of the characteristic genes includes: The trained cell classification model determines the probability that the second cell belongs to each of the cell types based on the cell genes contained in the second cell and the distinguishability of the cell types associated with the various cell types. The predicted cell type of the second cell is determined based on the classification probability of the second cell belonging to each of the aforementioned cell types.
4. The method according to claim 3, characterized in that, The step of obtaining target cells from a plurality of second cells whose predicted cell type matches the actual cell type of the second cells includes: From the plurality of second cells, identify a plurality of type-matching cells whose predicted cell type is the same as the actual cell type of the second cells; For each cell type, select matching cells belonging to that cell type, and among the selected matching cells, select cells whose target classification probability satisfies the probability condition as target cells; the target classification probability is the classification probability of the matching cell belonging to that cell type. Multiple target cells are obtained based on the target cells of each described cell type.
5. The method according to any one of claims 1 to 4, characterized in that, The actual cell types of the cells in the integrated cell dataset were determined through the following steps: Obtain multiple batches of cell datasets corresponding to the target tissue structure; Batch effect correction is performed on the multiple batches of cell datasets, and cells in the batch-corrected cell datasets are clustered according to cell characteristics to obtain multiple cell clusters. For each cell cluster, characteristic genes of cells in the cell cluster are determined, and reference cell annotations of cells in the cell cluster are determined based on the characteristic genes of the cell cluster. The actual cell type of cells in the cell cluster is determined based on the reference cell annotations.
6. The method according to any one of claims 1 to 4, characterized in that, The cell ratio is determined based on the target cells and a target cell ratio determination model; the target cell ratio determination model is obtained through the following steps: Obtain multiple candidate cell proportions to determine the model; Based on the test results of the model on the sample cell dataset according to the cell proportions of each candidate, the performance index information of the cell proportion determination model for each candidate is determined. The sample cell dataset includes: a first cell dataset with data features identical to those of at least one batch of cell datasets, and a second cell dataset with data features different from those of at least one batch of cell datasets; From the multiple candidate cell proportion determination models, the target cell proportion determination model whose performance index information meets the performance conditions is determined.
7. The method according to claim 6, characterized in that, The step of determining the performance metrics of each candidate cell proportion determination model based on the test results of the model on the sample cell dataset according to the cell proportion of each candidate model includes: Based on the multiple candidate cell proportion determination models and the first cell dataset, determine the first performance index information of each candidate cell proportion determination model; Cell proportion determination models that meet the performance conditions based on the first performance index information are selected. Based on the selected cell proportion determination models and the second cell dataset, the second performance index information of the selected cell proportion determination models is determined. The step of determining the target cell proportion determination model from multiple candidate cell proportion determination models, wherein the performance index information satisfies the performance conditions, includes: From the selected cell proportion determination models, the model whose second performance index information meets the performance conditions is selected as the target cell proportion determination model.
8. A cell ratio determination device, characterized in that, The device includes: The dataset acquisition module is used to acquire a first cell data subset and a second cell data subset; the first cell data subset and the second cell data subset are obtained by randomly dividing the integrated cell dataset, wherein the integrated cell dataset includes cell datasets from multiple batches of the target tissue structure, and the integrated cell dataset includes multiple cells of different actual cell types; The gene pattern determination module is used to determine the distribution of characteristic genes of various cell types in the integrated cell dataset based on the first cell in the first cell data subset. The target cell screening module is used to determine the predicted cell type of the second cell in the second cell data subset based on the distribution of the characteristic genes, and to obtain target cells from multiple second cells whose predicted cell type matches the actual cell type. The cell proportion determination module is used to determine the cell proportion of each cell type in the target tissue structure based on each target cell.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.