Proliferative cell proportion evaluation method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202380094804.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-28
- Publication Date
- 2025-10-03
AI Technical Summary
In the prior art, when evaluating the proportion of proliferating cells, due to the subjectivity of the interpretation of the results, artificial errors are easily caused, resulting in errors in the evaluation results.
By obtaining transcriptome data and immunofluorescence data of training malignant cells from the transcriptome data of the sample, screening the model characteristic genes based on the evaluation model, obtaining the target transcriptome data, and entering the second evaluation model for model training, obtaining the trained Evaluate the model.
The evaluation results are not required to be interpreted manually, which reduces the chance of artificial error and reduces the possibility of errors in the evaluation results of proliferating cell proportion.
Smart Images

Figure CN120752703A_ABST
Abstract
Description
Method and device for evaluating the proportion of proliferating cells, electronic device, and storage medium Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a method and device for evaluating the proportion of proliferating cells, an electronic device, and a storage medium. Background Art
[0002] Cell proliferation is an important vital characteristic of organisms. Detecting cell proliferation can be used to assess the health of normal cells, measure responses to toxic damage, or serve as a prognostic and diagnostic tool for various cancers. An important basis for detecting cell proliferation is the assessment of the proportion of proliferating cells.
[0003] In related technologies, the proportion of proliferating cells is mainly assessed through immunohistochemistry (IHC) staining of Ki67 (antigen KI-67). However, when assessing the proportion of proliferating cells through Ki67 IHC, the Ki67 IHC results need to be manually interpreted. Due to the subjectivity of the result interpretation, human errors are easily caused, which may lead to errors in the assessment results of the proliferating cell proportion.
[0004] Summary of the Invention
[0005] The present disclosure provides a method and apparatus, electronic device, and storage medium for evaluating the proportion of proliferating cells. The primary purpose of the present disclosure is to address the problem in related technologies whereby subjective interpretation of the results during evaluation of the proportion of proliferating cells can easily lead to human error, resulting in potential errors in the evaluation results.
[0006] According to a first aspect of the present disclosure, a method for training a model for evaluating the proportion of proliferating cells is provided, comprising:
[0007] Obtaining transcriptome data of malignant cells for training and obtaining immunofluorescence data for training from transcriptome data of samples;
[0008] Obtaining cell classification labels for the training malignant cells based on the transcriptome data and the training immunofluorescence data of the training malignant cells;
[0009] screening model feature genes based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model;
[0010] Based on the model characteristic genes, obtaining target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells;
[0011] The target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells are input into a second evaluation model for model training to obtain a trained second evaluation model.
[0012] Optionally, obtaining transcriptome data of malignant cells for training from transcriptome data of samples includes:
[0013] Identifying training endothelial cells using a first preset algorithm, and determining first training label information corresponding to the training endothelial cells;
[0014] Based on the first training label information and the training epithelial cell adhesion molecule, performing cell annotation on the training reference data using a second preset algorithm to obtain training epithelial cells, wherein the training reference data is obtained by combining human primary cell atlas data and genetic blueprint project data;
[0015] inferring the copy number variation in the training epithelial cells and the training reference epithelial cells using a third preset algorithm to obtain a copy number variation matrix of the training epithelial cells and a copy number variation matrix of the training reference epithelial cells, wherein the training reference epithelial cells are known normal epithelial cells;
[0016] merging the copy number variation matrix of the training epithelial cells and the copy number variation matrix of the training reference epithelial cells to obtain a merged training copy number variation matrix;
[0017] A clustering calculation is performed on the merged copy number variation matrix for training based on a preset clustering algorithm to obtain transcriptome data of the malignant cell data for training.
[0018] Optionally, obtaining cell classification labels for training malignant cells based on the transcriptome data and training immunofluorescence data of the training malignant cells includes:
[0019] registering the training immunofluorescence data with the training malignant cell transcriptome data to obtain registered training immunofluorescence data;
[0020] Pooling the registered training immunofluorescence data according to a preset rule to obtain pooled training immunofluorescence data;
[0021] The pooled training immunofluorescence data is binarized using a fourth preset algorithm to obtain second label information for proliferating / non-proliferating cells for training, and cell classification labels for the training malignant cells are determined.
[0022] Optionally, screening model feature genes based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model includes:
[0023] Calculating, based on the first evaluation model, a contribution value of each gene in the first evaluation model according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells, wherein the contribution value is the difference between an expected value of the first evaluation model and a predicted value after deleting a gene;
[0024] Based on the first evaluation model, each gene is sorted according to the contribution value to obtain the model characteristic gene.
[0025] Optionally, the acquiring target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells based on the model characteristic genes includes:
[0026] Eliminating a preset number of genes according to the model characteristic genes, and obtaining the model characteristic genes after elimination according to the area under the curve score of the first evaluation model;
[0027] Based on the eliminated model characteristic genes, target transcriptome data of the training malignant cells are obtained from the transcriptome data of the training malignant cells.
[0028] According to a second aspect of the present disclosure, a method for evaluating the proportion of proliferating cells is provided, comprising:
[0029] Obtain transcriptome data of malignant cells from transcriptome data of samples;
[0030] Based on the model characteristic genes, obtaining target transcriptome data of the malignant cells from the transcriptome data of the malignant cells;
[0031] The target transcriptome data of the malignant cells are input into a trained second evaluation model to obtain an evaluation result of the proliferating cell ratio, wherein the trained second evaluation model is trained by the target transcriptome data of the training malignant cells obtained according to the model characteristic genes.
[0032] Optionally, obtaining transcriptome data of malignant cells from transcriptome data of a sample includes:
[0033] Identifying endothelial cells using a first preset algorithm and determining first label information corresponding to the endothelial cells;
[0034] Based on the first label information and epithelial cell adhesion molecules, performing cell annotation on reference data using a second preset algorithm to obtain epithelial cells, wherein the reference data is obtained by combining human primary cell atlas data and genetic blueprint project data;
[0035] inferring the copy number variation in the epithelial cells and the reference epithelial cells using a third preset algorithm to obtain a copy number variation matrix of the epithelial cells and a copy number variation matrix of the reference epithelial cells, wherein the reference epithelial cells are known normal epithelial cells;
[0036] merging the copy number variation matrix of the epithelial cells and the copy number variation matrix of the reference epithelial cells to obtain a merged copy number variation matrix;
[0037] Clustering calculation is performed on the merged copy number variation matrix based on a preset clustering algorithm to obtain the transcriptome data of the malignant cells.
[0038] Optionally, before obtaining target transcriptome data of malignant cells from the transcriptome data of the malignant cells based on the model feature genes, the method further includes:
[0039] Obtaining transcriptome data of malignant cells for training and obtaining immunofluorescence data for training from transcriptome data of samples;
[0040] Obtaining cell classification labels for the training malignant cells based on the transcriptome data and the training immunofluorescence data of the training malignant cells;
[0041] screening model feature genes based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model;
[0042] Based on the model characteristic genes, obtaining target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells;
[0043] The target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells are input into a second evaluation model for model training to obtain the trained second evaluation model.
[0044] Optionally, screening model feature genes based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model includes:
[0045] Calculating, based on the first evaluation model, a contribution value of each gene in the first evaluation model according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells, wherein the contribution value is the difference between an expected value of the first evaluation model and a predicted value after deleting a gene;
[0046] Based on the first evaluation model, each gene in the training set is sorted according to the contribution value to obtain the model characteristic gene.
[0047] According to a third aspect of the present disclosure, a training device for an evaluation model of a proliferating cell ratio is provided, comprising:
[0048] A first acquisition unit is configured to acquire transcriptome data of malignant cells for training and immunofluorescence data for training from transcriptome data of a sample;
[0049] a second acquiring unit, configured to acquire cell classification labels of the training malignant cells based on the transcriptome data and the training immunofluorescence data of the training malignant cells;
[0050] a screening unit, configured to screen model characteristic genes according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model;
[0051] a third acquisition unit, configured to acquire target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells based on the model characteristic genes;
[0052] The training unit is configured to input the target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells into a second evaluation model for model training to obtain a trained second evaluation model.
[0053] Optionally, the first acquiring unit includes:
[0054] an identification module, configured to identify the training endothelial cells using a first preset algorithm and determine first training label information corresponding to the training endothelial cells;
[0055] an annotation module for performing cell annotation on training reference data using a second preset algorithm based on the first training label information and the training epithelial cell adhesion molecule to obtain training epithelial cells, wherein the training reference data is obtained by combining human primary cell atlas data and genetic blueprint project data;
[0056] an inference module, configured to infer copy number variations in the training epithelial cells and the training reference epithelial cells using a third preset algorithm to obtain a copy number variation matrix of the training epithelial cells and a copy number variation matrix of the training reference epithelial cells, wherein the training reference epithelial cells are known normal epithelial cells;
[0057] a merging module, configured to merge the copy number variation matrix of the training epithelial cells and the copy number variation matrix of the training reference epithelial cells to obtain a merged training copy number variation matrix;
[0058] A clustering module is used to perform clustering calculation on the merged copy number variation matrix for training based on a preset clustering algorithm to obtain transcriptome data of the malignant cell data for training.
[0059] Optionally, the second acquiring unit includes:
[0060] a registration module, configured to register the training immunofluorescence data with the training malignant cell transcriptome data to obtain registered training immunofluorescence data;
[0061] A pooling module, configured to pool the registered training immunofluorescence data according to a preset rule to obtain pooled training immunofluorescence data;
[0062] a processing module, configured to perform binarization processing on the pooled training immunofluorescence data using a fourth preset algorithm to obtain second label information for training of proliferating / non-proliferating cells;
[0063] A determination module is used to determine the cell classification labels of the training malignant cells.
[0064] Optionally, the screening unit includes:
[0065] a calculation module, configured to calculate, based on the first evaluation model and according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells, a contribution value of each gene in the first evaluation model, wherein the contribution value is the difference between an expected value of the first evaluation model and a predicted value after deleting a gene;
[0066] A sorting module is used to sort each gene in the training set according to the contribution value based on the first evaluation model to obtain the model characteristic gene.
[0067] Optionally, the third acquiring unit includes:
[0068] A elimination module is used to eliminate a preset number of genes according to the model characteristic genes, and obtain the model characteristic genes after elimination according to the area under the curve score of the first evaluation model;
[0069] An acquisition module is used to acquire target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells based on the eliminated model characteristic genes.
[0070] According to a fourth aspect of the present disclosure, a device for evaluating the proportion of proliferating cells is provided, comprising:
[0071] a first acquiring unit, configured to acquire transcriptome data of malignant cells from transcriptome data of a sample;
[0072] a second acquisition unit, configured to acquire target transcriptome data of the malignant cells from the transcriptome data of the malignant cells based on the model characteristic genes;
[0073] An input unit is used to input the target transcriptome data of the malignant cells into a trained second evaluation model to obtain an evaluation result of the proliferating cell ratio, wherein the trained second evaluation model is trained by the target transcriptome data of the training malignant cells obtained based on the model characteristic genes.
[0074] Optionally, the first obtaining unit includes:
[0075] an identification module, configured to identify endothelial cells using a first preset algorithm and determine first label information corresponding to the endothelial cells;
[0076] an annotation module for performing cell annotation on reference data using a second preset algorithm based on the first tag information and epithelial cell adhesion molecules to obtain epithelial cells, wherein the reference data is obtained by combining human primary cell atlas data and genetic blueprint project data;
[0077] an inference module, configured to infer the copy number variation in the epithelial cells and the reference epithelial cells using a third preset algorithm to obtain a copy number variation matrix of the epithelial cells and a copy number variation matrix of the reference epithelial cells, wherein the reference epithelial cells are known normal epithelial cells;
[0078] a merging module, configured to merge the copy number variation matrix of the epithelial cells and the copy number variation matrix of the reference epithelial cells to obtain a merged copy number variation matrix;
[0079] A clustering module is used to perform clustering calculation on the merged copy number variation matrix based on a preset clustering algorithm to obtain transcriptome data of the malignant cell data.
[0080] Optionally, the device further includes:
[0081] a third acquisition unit, configured to acquire transcriptome data of malignant cells for training and immunofluorescence data for training from the transcriptome data of the sample;
[0082] a fourth acquiring unit, configured to acquire cell classification labels of the training malignant cells based on the transcriptome data and the training immunofluorescence data of the training malignant cells;
[0083] a screening unit, configured to screen model characteristic genes according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model;
[0084] a fifth acquisition unit, configured to acquire target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells based on the model characteristic genes;
[0085] The training unit is configured to input the target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells into a second evaluation model for model training to obtain the trained second evaluation model.
[0086] Optionally, the screening unit includes:
[0087] a calculation module, configured to calculate, based on the first evaluation model and according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells, a contribution value of each gene in the first evaluation model, wherein the contribution value is the difference between an expected value of the first evaluation model and a predicted value after deleting a gene;
[0088] A sorting module is used to sort each gene in the training set according to the contribution value based on the first evaluation model to obtain the model characteristic gene.
[0089] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:
[0090] at least one processor; and
[0091] a memory communicatively connected to the at least one processor; wherein,
[0092] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.
[0093] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.
[0094] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the first aspect above.
[0095] The present disclosure provides a method and device for evaluating the proportion of proliferating cells, an electronic device, and a storage medium. The method comprises obtaining transcriptome data of training malignant cells and training immunofluorescence data from the transcriptome data of a sample; obtaining cell classification labels of the training malignant cells based on the transcriptome data of the training malignant cells and the training immunofluorescence data; screening model feature genes based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on a first evaluation model; obtaining target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells based on the model feature genes; and inputting the target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells into a second evaluation model for model training to obtain a trained second evaluation model. Compared with related technologies, the present disclosure uses an evaluation model to evaluate the proportion of proliferating cells, eliminating the need for manual interpretation of the evaluation results, reducing the probability of human error, and reducing the possibility of errors in the evaluation results of the proliferating cell proportion.
[0096] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0098] FIG1 is a flow chart of a method for training a model for evaluating the proportion of proliferating cells according to an embodiment of the present disclosure;
[0099] FIG2 is a flow chart of a method for obtaining transcriptome data of malignant cells for training provided by an embodiment of the present disclosure;
[0100] FIG3 is a result diagram of a cell annotation provided by an embodiment of the present disclosure;
[0101] FIG4 is a diagram showing the copy number variation results of an epithelial cell used for inference training provided by an embodiment of the present disclosure;
[0102] FIG5 is a diagram showing a clustering result of malignant / non-malignant cells provided by an embodiment of the present disclosure;
[0103] FIG6 is a flow chart of a method for obtaining cell classification labels of malignant cells for training provided by an embodiment of the present disclosure;
[0104] FIG7 is a diagram of immunofluorescence data for training after registration provided by an embodiment of the present disclosure;
[0105] FIG8 is a graph of pooled immunofluorescence data for training provided by an embodiment of the present disclosure;
[0106] FIG9 is a diagram showing a result of determining a proliferating / non-proliferating cell label according to an embodiment of the present disclosure;
[0107] FIG10 is a schematic diagram of a flow chart of a method for obtaining model characteristic genes provided by an embodiment of the present disclosure;
[0108] FIG11 is a schematic flow chart of a method for acquiring target transcriptome data of malignant cells for training provided by an embodiment of the present disclosure;
[0109] FIG12 is a graph showing the AUC (Area under curve) scores of a training set and a validation set in a first evaluation model provided by an embodiment of the present disclosure;
[0110] FIG13 is a graph showing the AUC (Area under curve) scores of another training set and validation set in the first evaluation model provided by an embodiment of the present disclosure;
[0111] FIG14 is a graph showing the AUC (Area under curve) scores of another training set and validation set in the first evaluation model provided by an embodiment of the present disclosure;
[0112] FIG15 is a graph showing the area under the AUC and precision-recall curves provided by an embodiment of the present disclosure;
[0113] FIG16 is another graph of AUC and area under the precision-recall curve provided by an embodiment of the present disclosure;
[0114] FIG17 is a schematic flow chart of a method for evaluating the proportion of proliferating cells provided in an embodiment of the present disclosure;
[0115] FIG18 is a diagram showing an evaluation result of a proliferating cell ratio according to an embodiment of the present disclosure;
[0116] FIG19 is a diagram showing Ki67 immunohistochemistry results of adjacent sections provided by an embodiment of the present disclosure;
[0117] FIG20 is a gene comparison diagram provided by an embodiment of the present disclosure;
[0118] FIG21 is a schematic structural diagram of a training device for a proliferating cell ratio evaluation model provided by an embodiment of the present disclosure;
[0119] FIG22 is a schematic structural diagram of a training device for another proliferating cell ratio evaluation model provided by an embodiment of the present disclosure;
[0120] FIG23 is a schematic structural diagram of a device for evaluating the proportion of proliferating cells provided by an embodiment of the present disclosure;
[0121] FIG24 is a schematic structural diagram of another device for evaluating the proportion of proliferating cells provided by an embodiment of the present disclosure;
[0122] FIG25 is a schematic block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0123] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0124] The following describes a method and apparatus for evaluating the proportion of proliferating cells, an electronic device, and a storage medium according to embodiments of the present disclosure with reference to the accompanying drawings.
[0125] FIG1 is a flow chart of a method for training a model for evaluating the proportion of proliferating cells according to an embodiment of the present disclosure.
[0126] As shown in Figure 1, the method includes the following steps:
[0127] Step 101 : Acquire transcriptome data of malignant cells for training and acquire immunofluorescence data for training from transcriptome data of a sample.
[0128] In the disclosed embodiments, the transcriptome data of the malignant cells used for training and the immunofluorescence data used for training include but are not limited to: matrix data, image data, etc., and the transcriptome data of the sample include but are not limited to: cancer tissue sections to be studied, spatial transcriptome databases, immunofluorescence image databases, etc.
[0129] Step 102 : obtaining cell classification labels of the training malignant cells based on the transcriptome data and the training immunofluorescence data of the training malignant cells.
[0130] In the disclosed embodiment, the cell classification labels of the malignant cells for training include but are not limited to: matrix data with cell classification labels, image data with cell classification labels, etc., wherein each cell data of the cell classification labels of the malignant cells for training carries a label of proliferating / non-proliferating cells.
[0131] Step 103 : Screening model feature genes based on the first evaluation model according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells.
[0132] In the disclosed embodiments, the first evaluation model includes, but is not limited to, an XGBoost model, and the model feature genes are genes associated with proliferating cells. The model feature genes can improve the accuracy of the evaluation of proliferating cells. After the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells are combined, each data point in the transcriptome data of the training malignant cells carries the cell classification label of the malignant cells.
[0133] It should be noted that the first evaluation model must be trained before determining the model's characteristic genes. After the training, the first evaluation model outputs the transcriptome data as input, resulting in a predicted value, and the label corresponding to the transcriptome data as the expected value. The training method includes, but is not limited to, dividing the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells into a training set and a validation set; inputting the training set and the validation set into the first evaluation model; based on the first evaluation model, obtaining a training predicted value and a contribution value based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells; and optimizing the first evaluation model according to a preset optimization algorithm based on the training predicted value and contribution value to obtain a trained first evaluation model.
[0134] Step 104 : acquiring target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells based on the model characteristic genes.
[0135] In the disclosed embodiment, the model characteristic genes are used to search and match the transcriptome data of the training malignant cells, and the data corresponding to the model characteristic genes in the transcriptome data of the training malignant cells is the target transcriptome data of the training malignant cells.
[0136] Step 105 : Input the target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells into a second evaluation model for model training to obtain a trained second evaluation model.
[0137] In the disclosed embodiment, when performing model training, the second evaluation model needs to be optimized by a preset optimization algorithm and a preset optimization objective function to obtain a trained second evaluation model, wherein the preset optimization algorithm is: a custom-selected optimization algorithm, such as: the Bayesian hyperparameter optimization algorithm in the hyperopt software, etc., and the preset optimization objective function is: a custom-selected optimization objective function, such as: AUC (Area under curve, area under the curve) score, etc. Wherein, after the target transcriptome data of the malignant cells for training and the cell classification labels of the malignant cells for training are combined, each data in the target transcriptome data of the malignant cells for training carries the cell classification label of the malignant cells. It should be noted here that the second evaluation model and the first evaluation model are models of the same type, but with different parameters and structures.
[0138] The present disclosure provides a method for estimating the proportion of proliferating cells. The method comprises obtaining transcriptome data of training malignant cells and training immunofluorescence data from the transcriptome data of a sample; obtaining cell classification labels for the training malignant cells based on the transcriptome data and the training immunofluorescence data of the training malignant cells; screening model feature genes based on the transcriptome data and the cell classification labels of the training malignant cells using a first evaluation model; obtaining target transcriptome data for the training malignant cells from the transcriptome data of the training malignant cells based on the model feature genes; and inputting the target transcriptome data and the cell classification labels of the training malignant cells into a second evaluation model for model training to obtain a trained second evaluation model. Compared to related art techniques, the present disclosure uses an evaluation model to assess the proportion of proliferating cells, eliminating the need for human interpretation of the evaluation results, reducing the likelihood of human error, and minimizing the possibility of errors in the evaluation results of the proliferating cell proportion.
[0139] In one implementation of the present disclosure, as a refinement of step 101, regarding obtaining transcriptome data of malignant cells for training from transcriptome data of a sample, the present disclosure provides a flow chart of a method for obtaining transcriptome data of malignant cells for training, as shown in FIG2 , comprising the following steps:
[0140] Step 201 : Identify training endothelial cells using a first preset algorithm, and determine first training label information corresponding to the training endothelial cells.
[0141] In the embodiment of the present disclosure, the first preset algorithm is: a user-selected algorithm for identifying endothelial cells for training, such as: the algorithm in the AUCell software, etc. Specifically, the embodiment of the present disclosure does not limit the first preset algorithm.
[0142] Identifying training endothelial cells can be accomplished by, but is not limited to, using the following method: using AUCell software to score and identify endothelial cells in each bin50 based on the endothelial cell gene set. Bin50 refers to data for bin50 (50×50 DNB bins, 25 μm). For details about the endothelial cell gene set and bin50 (50×50 DNB bins, 25 μm), please refer to the relevant art for details, and will not be detailed here.
[0143] It should be noted here that after annotating endothelial cells, ie, after identifying training endothelial cells and determining the first training label information corresponding to the training endothelial cells, the accuracy of annotating other cell types can be improved.
[0144] Step 202 , based on the first training label information and epithelial cell adhesion molecules, perform cell annotation on the training reference data using a second preset algorithm to obtain training epithelial cells, wherein the training reference data is obtained by combining human primary cell atlas data and gene blueprint project data.
[0145] In the embodiment of the present disclosure, the second preset algorithm is: a custom-selected algorithm for performing cell annotation on reference data, such as: the algorithm in the SingleR software, etc. Specifically, the embodiment of the present disclosure does not limit the second preset algorithm.
[0146] Regarding cell annotation of training reference data, it can be achieved by, but not limited to, the following methods: in the case of annotating training endothelial cells, that is, determining the first label information for training, the training reference data is annotated with cells using SingleR software. Cells that express epithelial cell adhesion molecules and can be annotated by SingleR software are training epithelial cells. After determining and annotating the training epithelial cells, the training epithelial cells in the training reference data are removed, and the remaining cell types are annotated to complete the annotation of all cell types, which facilitates more intuitive distinction between cell types. The results of cell annotation are shown in Figure 3, where epithelial_cells are training epithelial cells.
[0147] The expression of the epithelial cell adhesion molecules can be intuitively observed in the reference data. The human primary cell atlas data in the reference data include but are not limited to: expression profile data of smooth muscle cells, epithelial cells, fibroblasts, etc. The gene blueprint project data in the reference data include but are not limited to: expression profile data of neutrophils, NK cells, CD8T cells, CD4T cells, B cells, neutrophils, macrophages and dendritic cells, etc.
[0148] Step 203: Inferring the copy number variation in the training epithelial cells and the training reference epithelial cells using a third preset algorithm to obtain a copy number variation matrix of the training epithelial cells and a copy number variation matrix of the training reference epithelial cells, wherein the training reference epithelial cells are known normal epithelial cells.
[0149] In the embodiment of the present disclosure, the third preset algorithm is: a custom-selected algorithm for inferring copy number variation, for example: an algorithm in the InferCNV software, etc. Specifically, the embodiment of the present disclosure does not limit the third preset algorithm.
[0150] After using the InferCNV software to infer copy number variation, a cell copy number variation result graph can also be obtained, as shown in Figure 4. Figure 4 shows the copy number variation result graph of the training epithelial cells inferred using the InferCNV software. It should be noted that the copy number variation matrix of the training epithelial cells and the copy number variation matrix of the training reference epithelial cells can be expressed in the form of, but not limited to, behavioral cell data and the copy number variation matrix of genetic data.
[0151] Step 204 : Merge the copy number variation matrix of the training epithelial cells and the copy number variation matrix of the training reference epithelial cells to obtain a merged training copy number variation matrix.
[0152] In the embodiments of the present disclosure, the matrix merging method includes but is not limited to: merging the copy number variation matrix of the training epithelial cells and the copy number variation matrix of the training reference epithelial cells by splicing them in the gene dimension, that is, merging them in a column-by-column splicing manner to obtain behavioral cell data, which is listed as an expression matrix of gene data.
[0153] Step 205 : performing clustering calculation on the merged copy number variation matrix for training based on a preset clustering algorithm to obtain transcriptome data of the malignant cells for training.
[0154] In the embodiment of the present disclosure, the preset clustering algorithm is: a user-selected clustering algorithm, for example: kmeans (k-means) algorithm, etc. Specifically, the embodiment of the present disclosure does not limit the preset clustering algorithm.
[0155] Acquiring the transcriptome data of the training malignant cells can be achieved by, but is not limited to, the following method: clustering the merged training copy number variation matrix using the kmeans algorithm, setting the number of clusters to 2. The transcriptome data of non-malignant cells clustered with the training reference epithelial cells is the transcriptome data of non-malignant cells, and the remaining data is the transcriptome data of the training malignant cells. The clustering result is shown in FIG5 , which is a spatial in situ image of malignant / non-malignant cells. The left image is a spatial in situ image of malignant / non-malignant cells. To more intuitively distinguish malignant / non-malignant cells, the right image is also output: a separated malignant / non-malignant cell image, where Malignant_cells represents malignant cells and Non-Malignant_cells represents non-malignant cells.
[0156] It should be noted here that in order to distinguish non-malignant cell data from malignant cell data, they need to be clustered into two categories, so the number of clusters must be set to 2.
[0157] In one implementation of the present disclosure, as a refinement of step 102, regarding obtaining cell classification labels for training malignant cells based on the transcriptome data and training immunofluorescence data of the training malignant cells, the present disclosure provides a flow chart of a method for obtaining cell classification labels for training malignant cells, as shown in FIG6 , including the following steps:
[0158] Step 601 : aligning the training immunofluorescence data with the training malignant cell transcriptome data to obtain aligned training immunofluorescence data.
[0159] In the disclosed embodiment, the training immunofluorescence data and the training malignant cell transcriptome data include but are not limited to: matrix data, image data, etc., and the acquisition locations of the training immunofluorescence data and the training malignant cell transcriptome data include but are not limited to: cancer tissue sections to be studied, etc., wherein the registered training immunofluorescence data is shown in Figure 7.
[0160] Step 602 : Pooling the registered training immunofluorescence data according to a preset rule to obtain pooled training immunofluorescence data.
[0161] In the embodiment of the present disclosure, the preset rule is: a user-selected pooling rule, for example: bin50 pooling, etc. Specifically, the embodiment of the present disclosure does not limit the preset rule.
[0162] Regarding the pooling of the registered training immunofluorescence data, it can be implemented in the following way but is not limited to: bin50 pooling is performed on the registered training immunofluorescence data, that is, the pixel values in the 50*50 pixel box are merged, and the merged pixel value is specified as the pixel value of the preset percentile in the box, wherein the preset percentile is a custom-set percentile, for example: 75th percentile, 85th percentile, etc. Specifically, the setting of the percentile is not limited in the embodiment of the present disclosure.
[0163] As shown in FIG8 , FIG8 is an image obtained by performing bin50 pooling on the registered training immunofluorescence data, that is, the pooled training immunofluorescence data. When performing bin50 pooling, the merged pixel value is specified to be the pixel value at the 75th percentile within the box.
[0164] Step 603 : Binarize the pooled training immunofluorescence data using a fourth preset algorithm to obtain second label information for training of proliferating / non-proliferating cells, and determine the cell classification labels for the training malignant cells.
[0165] In the embodiment of the present disclosure, the fourth preset algorithm is: a custom-selected binarization processing algorithm, for example: the mean color threshold algorithm in ImageJ software, etc. Specifically, the embodiment of the present disclosure does not limit the fourth preset algorithm.
[0166] Please refer to Figure 9, which shows the result of determining the proliferating / non-proliferating cell labels. As shown in Figure 10, gray represents proliferating cells and black represents non-proliferating cells. The proliferating / non-proliferating cell labels, i.e., the second label information for training the proliferating / non-proliferating cells, are the cell classification labels for the training malignant cells.
[0167] In one implementation of the embodiment of the present disclosure, as a refinement of the above step 103, the embodiment of the present disclosure provides a flow chart of a method for obtaining model characteristic genes, as shown in FIG10 , comprising the following steps:
[0168] Step 1001: Based on the first evaluation model, the contribution value of each gene in the first evaluation model is calculated according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells. The contribution value is the difference between the expected value of the first evaluation model and the predicted value after deleting a gene.
[0169] In the embodiment of the present disclosure, a custom feature selection algorithm is required, such as the shapRFECV algorithm in the probatus software, where shapRFECV is shap (SHapley Additive exPlanation) RFE (Recursive feature elimination) CV (Cross Validation). Specifically, the embodiment of the present disclosure does not limit this.
[0170] The transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells need to be divided into a training set and a validation set. This allocation is random. For example, 70% of the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells are divided into the training set, and 30% of the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells are divided into the validation set; 75% of the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells are divided into the training set, and 25% of the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells are divided into the validation set, etc.
[0171] It should be noted here that both the training set and the validation set need to be input into the first evaluation model to obtain the performance of the model. Regarding the contribution value, it is calculated by the feature selection algorithm. In order to facilitate understanding of the contribution value, an example is provided for illustration: before deleting a gene, the score predicted by the first evaluation model (the output of the first evaluation model) is 100, and after deleting the gene, the score predicted by the first evaluation model is 80, then the score difference of 20 is the contribution value of the gene. Among them, the smaller the difference, that is, the smaller the contribution value of a gene, the lower the importance of the gene, and the larger the difference, that is, the larger the contribution value of a gene, the higher the importance of the gene.
[0172] It should be noted that the contribution value can only be calculated after the first evaluation model has been trained. After the first evaluation model is trained, the output result obtained after inputting the transcriptome data is the predicted value, and the label corresponding to the transcriptome data is the expected value. The training method includes, but is not limited to, dividing the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells into a training set and a validation set; inputting the training set and the validation set into the first evaluation model; based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells, obtaining a training predicted value and a contribution value; and optimizing the first evaluation model according to a preset optimization algorithm based on the training predicted value and contribution value to obtain a trained first evaluation model.
[0173] Step 1002: Based on the first evaluation model, sort each gene in the training set according to the contribution value to obtain the model feature gene.
[0174] In the disclosed embodiment, in order to more intuitively determine the contribution value of each gene, each gene in the training set needs to be sorted according to the contribution value, and all genes after sorting are the model feature genes.
[0175] In one implementation of the present disclosure, as a refinement of step 104, the present disclosure provides a flowchart of a method for acquiring target transcriptome data of malignant cells for training, as shown in FIG11 , comprising the following steps:
[0176] Step 1101: Eliminate a preset number of genes based on the model characteristic genes, and obtain the model characteristic genes after elimination based on the area under the curve score of the first evaluation model.
[0177] In an embodiment of the present disclosure, in order to obtain the model characteristic genes after the elimination, it can be achieved by, but not limited to, the following method: according to the model characteristic genes obtained by the sorting result, the genes with low contribution values of the preset step length are eliminated, and the retained genes are determined according to the AUC scores of the training set and the validation set in the first evaluation model, and the elimination process is repeated for a preset number of rounds to obtain the performance of the first evaluation model under different numbers of genes, wherein the preset number of rounds is: the number of rounds determined according to the actual application situation, and the preset step size is: the step size set by the user, indicating that according to the sorting result, the gene with the smallest contribution value of the preset step size is deleted each time, for example: the preset step size is 1000, indicating that 1000 genes with the smallest contribution value are deleted each time. Specifically, with respect to the preset step size and the preset number of rounds, the embodiment of the present disclosure does not impose any restrictions.
[0178] To facilitate understanding of the implementation process of the disclosed embodiments, an example is provided for illustration: Gene selection was first performed with a step size of 1000. Based on the AUC score image shown in FIG12 , 871 genes were determined to remain as input genes for the second round of elimination. The second round of elimination was performed with a step size of 50. Based on the AUC score image shown in FIG13 , 171 genes were determined to remain as input genes for the third round of elimination. The third round of elimination was performed with a step size of 1. Based on the AUC score image shown in FIG14 , 65 genes were ultimately determined to remain as model feature genes after elimination.
[0179] Step 1102 : Based on the eliminated model feature genes, target transcriptome data of the training malignant cells are obtained from the transcriptome data of the training malignant cells.
[0180] In the disclosed embodiment, the model characteristic genes after elimination are paired with the transcriptome data of the training malignant cells, and the data containing the model characteristic genes after elimination is the target transcriptome data of the training malignant cells.
[0181] In one implementable manner of the embodiment of the present disclosure, as a refinement of the above-mentioned step 105, in order to facilitate understanding of the implementation process of the above-mentioned step 105, an example is provided for illustration: the 65 model characteristic genes obtained in the above-mentioned step 1005 are used as the input features of the model, and the Bayesian hyperparameter algorithm in the hyperopt software is used to optimize the evaluation model to obtain the trained second evaluation model, and the AUC score is used as the optimization objective function to verify the trained second evaluation model. The AUC and area under the precision-recall curve (AUPRC) of the trained second evaluation model in the verification set reached 0.8 and 0.78, respectively, as shown in Figures 15 and 16, proving that the trained second evaluation model has high accuracy.
[0182] Corresponding to the training method of the above-mentioned model for evaluating the proportion of proliferating cells, FIG17 is a flow chart of a method for evaluating the proportion of proliferating cells provided by an embodiment of the present disclosure. As shown in FIG17 , the method includes the following steps:
[0183] Step 1701 : Acquire transcriptome data of malignant cells from transcriptome data of a sample.
[0184] In the disclosed embodiments, the transcriptome data of the malignant cells include but are not limited to: matrix data, image data, etc., and the transcriptome data of the samples include but are not limited to: cancer tissue sections to be studied, spatial transcriptome databases, immunofluorescence image databases, etc.
[0185] Step 1702 : obtaining target transcriptome data of malignant cells from the transcriptome data of the malignant cells based on the model feature genes.
[0186] In the disclosed embodiment, the model characteristic genes are used to search and match the transcriptome data of the malignant cells, and the data corresponding to the model characteristic genes in the transcriptome data of the malignant cells are the target transcriptome data of the malignant cells.
[0187] In step 1703, the target transcriptome data of the malignant cells are input into a trained second evaluation model to obtain an evaluation result of the proliferating cell ratio. The trained second evaluation model is trained by the target transcriptome data of the malignant cells for training obtained based on the model characteristic genes.
[0188] In the embodiment of the present disclosure, the evaluation results of the immunofluorescence plus spatial transcriptome data are shown in Figure 18. Figure 18 is an evaluation result diagram of the proportion of proliferating cells provided in the embodiment of the present disclosure, wherein dark gray represents non-proliferating cells, light gray represents proliferating cells, and Predict_prop is the proportion of proliferating cells.
[0189] In order to verify the performance of the method described in the present disclosure, as shown in Figure 19, the present disclosure also provides a Ki67 immunohistochemistry result graph of an adjacent slice for comparison with the evaluation result graph of the proliferative cell ratio in Figure 18 above. The comparison results show that the Ki67 immunohistochemistry result graph of the adjacent slice is relatively consistent with the evaluation result graph of the proliferative cell ratio, verifying the feasibility of the clinical application of the evaluation model.
[0190] To further illustrate the importance of genes in the evaluation model, this disclosure displays the top 20 genes by contribution value, as shown in Figure 20. Figure 20 shows genes known to be associated with proliferating cells, such as MALAT1, TPT1, and COL1A1. This indicates that the evaluation model described in this disclosure is biologically interpretable and that other genes may play important roles in proliferating cells and could potentially serve as potential therapeutic targets.
[0191] In one implementation of the embodiment of the present disclosure, obtaining transcriptome data of malignant cells from transcriptome data of a sample includes:
[0192] Identifying endothelial cells using a first preset algorithm and determining first label information corresponding to the endothelial cells;
[0193] Based on the first label information and epithelial cell adhesion molecules, performing cell annotation on reference data using a second preset algorithm to obtain epithelial cells, wherein the reference data is obtained by combining human primary cell atlas data and genetic blueprint project data;
[0194] inferring the copy number variation in the epithelial cells and the reference epithelial cells using a third preset algorithm to obtain a copy number variation matrix of the epithelial cells and a copy number variation matrix of the reference epithelial cells, wherein the reference epithelial cells are known normal epithelial cells;
[0195] merging the copy number variation matrix of the epithelial cells and the copy number variation matrix of the reference epithelial cells to obtain a merged copy number variation matrix;
[0196] Clustering calculation is performed on the merged copy number variation matrix based on a preset clustering algorithm to obtain the transcriptome data of the malignant cells.
[0197] Specifically, regarding the implementation process of this embodiment, please refer to the description of steps 201-205 above, so they will not be described here one by one.
[0198] In one implementation of the embodiment of the present disclosure, before obtaining target transcriptome data of malignant cells from the transcriptome data of the malignant cells based on the model feature genes, the method further includes:
[0199] Obtaining transcriptome data of malignant cells for training and obtaining immunofluorescence data for training from transcriptome data of samples;
[0200] Obtaining cell classification labels for the training malignant cells based on the transcriptome data and the training immunofluorescence data of the training malignant cells;
[0201] screening model feature genes based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model;
[0202] Based on the model characteristic genes, obtaining target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells;
[0203] The target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells are input into a second evaluation model for model training to obtain the trained second evaluation model.
[0204] Specifically, regarding the implementation process of this embodiment, please refer to the description of steps 201-205 above, so they will not be described here one by one.
[0205] In one implementation of the disclosed embodiment, screening model feature genes based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model includes:
[0206] Calculating, based on the first evaluation model, a contribution value of each gene in the first evaluation model according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells, wherein the contribution value is the difference between an expected value of the first evaluation model and a predicted value after deleting a gene;
[0207] Based on the first evaluation model, each gene is sorted according to the contribution value to obtain the model characteristic gene.
[0208] Regarding the implementation process of this embodiment, please refer to the description of steps 1001-1002 above, so they will not be described one by one here.
[0209] In summary, the embodiments of the present disclosure can achieve the following effects:
[0210] The disclosed embodiment uses an evaluation model to evaluate the proportion of proliferating cells, eliminating the need for manual interpretation of the evaluation results, thereby reducing the probability of human error and the possibility of errors in the evaluation results of the proliferating cell ratio.
[0211] Corresponding to the aforementioned proliferating cell ratio estimation model training method and proliferating cell ratio estimation method, the present invention also provides a proliferating cell ratio estimation model training device and a proliferating cell ratio estimation device. Since the device embodiments of the present invention correspond to the aforementioned method embodiments, details not disclosed in the device embodiments can be referred to the aforementioned method embodiments and will not be further described in this invention.
[0212] FIG21 is a schematic diagram of the structure of a training device for a proliferating cell ratio evaluation model provided by an embodiment of the present disclosure. As shown in FIG21 , the device includes:
[0213] A first acquiring unit 211 is configured to acquire transcriptome data of malignant cells for training and immunofluorescence data for training from transcriptome data of a sample;
[0214] A second acquiring unit 212 is configured to acquire cell classification labels of the training malignant cells based on the transcriptome data and the training immunofluorescence data of the training malignant cells;
[0215] A screening unit 213 is configured to screen model feature genes based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model;
[0216] A third acquisition unit 214 is configured to acquire target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells based on the model characteristic genes;
[0217] The training unit 215 is configured to input the target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells into the second evaluation model for model training to obtain a trained second evaluation model.
[0218] The present disclosure provides an evaluation device for the proportion of proliferating cells. The device obtains transcriptome data of training malignant cells and training immunofluorescence data from the transcriptome data of a sample; obtains cell classification labels of the training malignant cells based on the transcriptome data and the training immunofluorescence data of the training malignant cells; screens model feature genes based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on a first evaluation model; obtains target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells based on the model feature genes; and inputs the target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells into a second evaluation model for model training to obtain a trained second evaluation model. Compared with related technologies, the present disclosure embodiment uses an evaluation model to evaluate the proportion of proliferating cells, eliminating the need for manual interpretation of the evaluation results, reducing the probability of human error and the possibility of errors in the evaluation results of the proportion of proliferating cells.
[0219] Furthermore, in a possible implementation of the embodiment of the present disclosure, as shown in FIG22 , the first acquiring unit 211 includes:
[0220] An identification module 2111 is configured to identify training endothelial cells using a first preset algorithm and determine first training label information corresponding to the training endothelial cells;
[0221] An annotation module 2112 is configured to perform cell annotation on the training reference data using a second preset algorithm based on the first training label information and the training epithelial cell adhesion molecule to obtain training epithelial cells, wherein the training reference data is obtained by combining human primary cell atlas data and genetic blueprint project data;
[0222] an inference module 2113, configured to infer copy number variations in the training epithelial cells and the training reference epithelial cells using a third preset algorithm to obtain a copy number variation matrix of the training epithelial cells and a copy number variation matrix of the training reference epithelial cells, wherein the training reference epithelial cells are known normal epithelial cells;
[0223] a merging module 2114, configured to merge the copy number variation matrix of the training epithelial cells and the copy number variation matrix of the training reference epithelial cells to obtain a merged training copy number variation matrix;
[0224] The clustering module 2115 is configured to perform clustering calculation on the merged copy number variation matrix for training based on a preset clustering algorithm to obtain transcriptome data of the malignant cell data for training.
[0225] Furthermore, in a possible implementation of the embodiment of the present disclosure, as shown in FIG22 , the second acquiring unit 212 includes:
[0226] a registration module 2121 for registering the training immunofluorescence data with the training malignant cell transcriptome data to obtain registered training immunofluorescence data;
[0227] A pooling module 2122 is configured to pool the registered training immunofluorescence data according to a preset rule to obtain pooled training immunofluorescence data;
[0228] A processing module 2123 is configured to perform binarization processing on the pooled training immunofluorescence data using a fourth preset algorithm to obtain second label information for training of proliferating / non-proliferating cells;
[0229] The determination module 2124 is configured to determine the cell classification labels of the training malignant cells.
[0230] Furthermore, in a possible implementation of the embodiment of the present disclosure, as shown in FIG22 , the screening unit 213 includes:
[0231] a calculation module 2131 configured to calculate, based on the first evaluation model and the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells, a contribution value of each gene in the first evaluation model, wherein the contribution value is the difference between an expected value of the first evaluation model and a predicted value after deleting a gene;
[0232] The sorting module 2132 is used to sort each gene in the training set according to the contribution value based on the first evaluation model to obtain the model feature gene.
[0233] Furthermore, in a possible implementation of the embodiment of the present disclosure, as shown in FIG22 , the third acquiring unit 214 includes:
[0234] A elimination module 2141 is configured to eliminate a preset number of genes based on the model characteristic genes, and obtain the model characteristic genes after elimination based on the area under the curve score of the first evaluation model;
[0235] The acquisition module 2142 is configured to acquire target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells based on the eliminated model feature genes.
[0236] FIG23 is a schematic diagram of the structure of a device for evaluating the proportion of proliferating cells provided by an embodiment of the present disclosure, as shown in FIG23 , comprising:
[0237] A first acquiring unit 231 is configured to acquire transcriptome data of malignant cells from transcriptome data of a sample;
[0238] A second acquisition unit 232 is configured to acquire target transcriptome data of the malignant cells from the transcriptome data of the malignant cells based on the model characteristic genes;
[0239] The input unit 233 is used to input the target transcriptome data of the malignant cells into the trained second evaluation model to obtain an evaluation result of the proliferating cell ratio. The trained second evaluation model is trained by the target transcriptome data of the training malignant cells obtained based on the model characteristic genes.
[0240] Furthermore, in a possible implementation of the embodiment of the present disclosure, as shown in FIG24 , the first acquiring unit 231 includes:
[0241] an identification module 2311, configured to identify endothelial cells using a first preset algorithm and determine first label information corresponding to the endothelial cells;
[0242] An annotation module 2312 is configured to perform cell annotation on reference data using a second preset algorithm based on the first tag information and epithelial cell adhesion molecules to obtain epithelial cells, wherein the reference data is obtained by combining human primary cell atlas data and genetic blueprint project data;
[0243] an inference module 2313, configured to infer the copy number variation in the epithelial cell and the reference epithelial cell using a third preset algorithm to obtain a copy number variation matrix of the epithelial cell and a copy number variation matrix of the reference epithelial cell, wherein the reference epithelial cell is a known normal epithelial cell;
[0244] a merging module 2314, configured to merge the copy number variation matrix of the epithelial cells and the copy number variation matrix of the reference epithelial cells to obtain a merged copy number variation matrix;
[0245] The clustering module 2315 is configured to perform clustering calculation on the merged copy number variation matrix based on a preset clustering algorithm to obtain transcriptome data of the malignant cell data.
[0246] Furthermore, in a possible implementation of the embodiment of the present disclosure, as shown in FIG24 , the apparatus further includes:
[0247] A third acquiring unit 234 is configured to acquire transcriptome data of malignant cells for training and immunofluorescence data for training from the transcriptome data of the sample;
[0248] A fourth acquiring unit 235 is configured to acquire a cell classification label of the training malignant cells based on the transcriptome data and the training immunofluorescence data of the training malignant cells;
[0249] A screening unit 236 is configured to screen model feature genes based on the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model;
[0250] a fifth acquiring unit 237 configured to acquire target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells based on the model characteristic genes;
[0251] The training unit 238 is configured to input the target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells into the second evaluation model for model training to obtain the trained second evaluation model.
[0252] Furthermore, in a possible implementation of the embodiment of the present disclosure, as shown in FIG24 , the screening unit 236 includes:
[0253] a calculation module 2361 configured to calculate, based on the first evaluation model and the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells, a contribution value of each gene in the first evaluation model, wherein the contribution value is the difference between an expected value of the first evaluation model and a predicted value after deleting a gene;
[0254] The sorting module 2362 is used to sort each gene in the training set according to the contribution value based on the first evaluation model to obtain the model feature gene.
[0255] It should be noted that the above explanation of the method embodiment is also applicable to the device of the embodiment of the present disclosure, and the principles are the same, which is no longer limited in the embodiment of the present disclosure.
[0256] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0257] FIG25 shows a schematic block diagram of an example electronic device 2500 that can be used to implement an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0258] As shown in FIG25 , the device 2500 includes a computing unit 2501 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 2502 or a computer program loaded from a storage unit 2508 into a RAM (Random Access Memory) 2503. Various programs and data required for the operation of the device 2500 may also be stored in the RAM 2503. The computing unit 2501, the ROM 2502, and the RAM 2503 are connected to each other via a bus 2504. An I / O (Input / Output) interface 2505 is also connected to the bus 2504.
[0259] Various components in device 2500 are connected to I / O interface 2505, including: an input unit 2506, such as a keyboard, mouse, etc.; an output unit 2507, such as various types of displays, speakers, etc.; a storage unit 2508, such as a magnetic disk, optical disk, etc.; and a communication unit 2509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 2509 allows device 2500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0260] The computing unit 2501 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 2501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 2501 performs the various methods and processes described above, such as the method for assessing the proportion of proliferating cells. For example, in some embodiments, the method for assessing the proportion of proliferating cells can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 2508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 2500 via the ROM 2502 and / or the communication unit 2509. When the computer program is loaded into the RAM 2503 and executed by the computing unit 2501, one or more steps of the method described above can be performed. Alternatively, in other embodiments, the computing unit 2501 may be configured to execute the aforementioned method for evaluating the proportion of proliferating cells in any other appropriate manner (eg, by means of firmware).
[0261] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0262] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0263] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0264] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0265] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0266] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0267] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0268] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0269] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A training method for an evaluation model of a proliferating cell ratio, characterized in that: include: Acquire transcriptome data of malignant cells for training from transcriptome data of samples and acquire immunofluorescence data for training; Acquire a cell classification label for the training malignant cells according to the transcriptome data of the training malignant cells and the immunofluorescence data for training; Based on the first evaluation model, the model feature genes are selected according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells; Based on the model characteristic genes, obtaining target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells; The target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells are input into a second evaluation model for model training to obtain a trained second evaluation model.
2. The method according to claim 1, characterized in that The step of obtaining transcriptome data of malignant cells for training from transcriptome data of samples includes: Identifying training endothelial cells using a first preset algorithm, and determining first training label information corresponding to the training endothelial cells; Based on the first label information for training and the epithelial cell adhesion molecule for training, performing cell annotation on the training reference data by a second preset algorithm to obtain the training epithelial cells, wherein the training reference data is obtained by combining the human primary cell atlas data and the gene blueprint project data; Inferring the copy number variation in the training epithelial cells and the training reference epithelial cells by a third preset algorithm to obtain a copy number variation matrix of the training epithelial cells and a copy number variation matrix of the training reference epithelial cells, wherein the training reference epithelial cells are known normal epithelial cells; Merging the copy number variation matrix of the training epithelial cells and the copy number variation matrix of the training reference epithelial cells to obtain a merged training copy number variation matrix; Clustering calculation is performed on the merged copy number variation matrix for training based on a preset clustering algorithm to obtain transcriptome data of the malignant cell data for training.
3. The method according to claim 2, characterized in that The step of obtaining the cell classification labels of the training malignant cells according to the transcriptome data of the training malignant cells and the training immunofluorescence data comprises: Aligning the training immunofluorescence data with the training malignant cell transcriptome data to obtain aligned training immunofluorescence data; Pooling the registered training immunofluorescence data according to a preset rule to obtain pooled training immunofluorescence data; The pooled training immunofluorescence data is binarized by a fourth preset algorithm. The second label information for training of proliferating / non-proliferating cells is obtained, and the cell classification labels of the training malignant cells are determined.
4. The method according to any one of claims 1 to 3, characterized in that The method of screening the model characteristic genes based on the first evaluation model according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells includes: Based on the first evaluation model, calculating the contribution value of each gene in the first evaluation model according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells, wherein the contribution value is the difference between the expected value of the first evaluation model and the predicted value after deleting a gene; Based on the first evaluation model, each gene is sorted according to the contribution value to obtain the model characteristic gene.
5. The method according to claim 4, characterized in that The obtaining target transcriptome data of the malignant cells for training from the transcriptome data of the malignant cells for training based on the model characteristic genes comprises: Eliminating a preset number of genes according to the model characteristic genes, and obtaining the model characteristic genes after elimination according to the area under the curve score of the first evaluation model; Based on the eliminated model characteristic genes, target transcriptome data of the malignant cells for training are obtained from the transcriptome data of the malignant cells for training.
6. A method for evaluating the proportion of proliferating cells, characterized in that: include: Obtaining transcriptome data of malignant cells from transcriptome data of samples; Based on the model characteristic genes, obtaining target transcriptome data of the malignant cells from the transcriptome data of the malignant cells; The target transcriptome data of the malignant cells are input into a trained second evaluation model to obtain an evaluation result of the proliferating cell ratio, wherein the trained second evaluation model is trained by the target transcriptome data of the training malignant cells obtained according to the model characteristic genes.
7. The method according to claim 6, characterized in that The step of obtaining transcriptome data of malignant cells from transcriptome data of samples includes: Identifying endothelial cells using a first preset algorithm, and determining first label information corresponding to the endothelial cells; Based on the first label information and the epithelial cell adhesion molecules, the reference data is annotated with cells by a second preset algorithm to obtain epithelial cells, wherein the reference data is obtained by combining human primary cell atlas data and gene blueprint project data; Inferring the copy number variation in the epithelial cells and the reference epithelial cells by a third preset algorithm to obtain a copy number variation matrix of the epithelial cells and a copy number variation matrix of the reference epithelial cells, wherein the reference epithelial cells are known normal epithelial cells; The copy number variation matrix of the epithelial cells and the copy number variation matrix of the reference epithelial cells are The different matrices are merged to obtain a merged copy number variation matrix; Clustering calculation is performed on the merged copy number variation matrix based on a preset clustering algorithm to obtain the transcriptome data of the malignant cells.
8. The method according to claim 6, characterized in that Before obtaining target transcriptome data of malignant cells from the transcriptome data of the malignant cells based on the model feature genes, the method further comprises: Acquire transcriptome data of malignant cells for training from transcriptome data of samples and acquire immunofluorescence data for training; Acquire a cell classification label for the training malignant cells according to the transcriptome data of the training malignant cells and the immunofluorescence data for training; Based on the first evaluation model, the model feature genes are selected according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells; Based on the model characteristic genes, obtaining target transcriptome data of the training malignant cells from the transcriptome data of the training malignant cells; The target transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells are input into a second evaluation model for model training to obtain the trained second evaluation model.
9. The method according to claim 8, characterized in that The method of screening the model characteristic genes based on the first evaluation model according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells includes: Based on the first evaluation model, calculating the contribution value of each gene in the first evaluation model according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells, wherein the contribution value is the difference between the expected value of the first evaluation model and the predicted value after deleting a gene; Based on the first evaluation model, each gene is sorted according to the contribution value to obtain the model characteristic gene.
10. A training device for an evaluation model of a proliferating cell ratio, characterized in that: include: A first acquisition unit is used to acquire transcriptome data of malignant cells for training and acquire immunofluorescence data for training from the transcriptome data of the sample; A second acquisition unit, configured to acquire a cell classification label of the training malignant cells according to the transcriptome data of the training malignant cells and the training immunofluorescence data; A screening unit, configured to screen model feature genes according to the transcriptome data of the training malignant cells and the cell classification labels of the training malignant cells based on the first evaluation model; A third acquisition unit is used to acquire target transcriptome data of the malignant cells for training from the transcriptome data of the malignant cells for training based on the model feature genes; A training unit is used to convert the target transcriptome data of the training malignant cells and the training The cell classification labels of the malignant cells are input into the second evaluation model for model training to obtain a trained second evaluation model.
11. A device for evaluating the proportion of proliferating cells, characterized in that: include: A first acquisition unit, used for acquiring transcriptome data of malignant cells from transcriptome data of the sample; A second acquisition unit is used to acquire target transcriptome data of malignant cells from the transcriptome data of the malignant cells based on the model characteristic genes; The input unit is used to input the target transcriptome data of the malignant cells into the trained second evaluation model to obtain an evaluation result of the proliferating cell ratio, wherein the trained second evaluation model is trained by the target transcriptome data of the training malignant cells obtained according to the model characteristic genes.
12. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5 or the method of any one of claims 6-9.
13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5 or the method according to any one of claims 6-9.
14. A computer program product, characterized in that The method comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 5 or the method according to any one of claims 6 to 9.