Data processing apparatus and method for predicting patient immunotherapy responsiveness

By constructing an immunotherapy response prediction model based on the random forest algorithm and utilizing the gene expression levels of cell subsets in single-cell sequencing data, the problem of accuracy in predicting patient immunotherapy response was solved, enabling personalized medicine and precise formulation of treatment plans.

CN120260905BActive Publication Date: 2025-12-26PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510282890.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-12-26
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict patient responses to immunotherapy, and there is a lack of effective data processing methods to screen for features that aid in prediction.

Method used

An immunotherapy response prediction model was constructed using the random forest algorithm. The expression levels of genes in cell subpopulations from single-cell sequencing data were used as input. The prediction model was constructed using machine learning algorithms, and the Shapley Value and LIME methods were used to interpret feature importance. Features were iteratively screened to reduce dimensionality.

Benefits of technology

It enables precise prediction of patients' immune therapy response, promotes personalized medicine, improves treatment effectiveness, reduces side effects, lowers medical costs, and provides more detailed analysis of the role of genes in cell subpopulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260905B_ABST
    Figure CN120260905B_ABST
Patent Text Reader

Abstract

The application discloses a data processing device and method for predicting the responsiveness of patients to immunotherapy in the field of bioinformatics. The technical problem to be solved by the application is how to accurately predict the responsiveness of tumor patients to immunotherapy. The application takes the expression amount characteristics of genes in cell subpopulations with different expressions in single-cell sequencing data of immunotherapy responsive subjects and non-responsive subjects as input, takes whether the patient responds to immunotherapy as output, and uses a random forest algorithm to construct an immunotherapy response prediction model. By inputting the expression amount of genes in cell subpopulations in single-cell sequencing data of the test subject into the model, the immunotherapy response result of the test subject can be predicted. The application can make more accurate treatment plans for tumor patients, reduce unnecessary side effects, reduce medical costs, and can be applied to individualized medical treatment of immunotherapy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of bioinformatics, and particularly relates to a data processing device and method for predicting the responsiveness of a patient to immunotherapy. BACKGROUND

[0002] Immunotherapy is a new therapy in cancer treatment that has attracted much attention, however, the responsiveness of different patients to immunotherapy has large individual differences. In order to exert the effect of immunotherapy and reduce the cost of treatment, it is an important problem in tumor immunotherapy to study the responsiveness of patients to immunotherapy and its mechanism.

[0003] With the rapid development of single-cell sequencing technology, the quality and quantity of tumor immune microenvironment single-cell sequencing data sets have made great progress. By studying patient single-cell sequencing data, the expression changes of different cell types in the disease process can be more comprehensively understood, providing new possibilities for individualized treatment. However, the emergence of a large amount of data also brings new challenges to data analysis. With the reduction of sequencing cost, the scale and complexity of biological data have greatly increased, and the influence of data filtering, noise reduction and other preprocessing on experimental results has become increasingly large, and reasonable data processing methods have gradually become a necessary means to obtain scientific conclusions. With the increase of data size, the field of machine learning has also developed rapidly, from statistical learning to large models based on neural networks, the technological innovation in the field of machine learning has also provided new tools for biological data processing.

[0004] In the field of immunotherapy for cancer, an important challenge is how to accurately predict the responsiveness of tumor patients to immunotherapy and screen features that are helpful for predicting the responsiveness to immunotherapy, so as to help clinical prediction. Some machine learning algorithms can use single-cell sequencing data to predict the responsiveness of patients, however, which data is more important for prediction is still a challenging problem.

[0005] The immune microenvironment refers to the complex network of immune cells, molecules, and structures present within tissues or organs. It plays a crucial role in regulating immune responses and maintaining tissue homeostasis. The immune microenvironment is highly dynamic and can change significantly depending on factors such as tissue type, physiological state, and the presence of disease. At its core are various types of immune cells, including lymphocytes (such as T cells, B cells, and natural killer cells), antigen-presenting cells (such as dendritic cells and macrophages), and other specialized cells (such as mast cells and granulocytes). These cells interact with each other and with non-immune cells within the tissue (such as epithelial cells, fibroblasts, and endothelial cells), forming a complex network of cellular interactions. In addition to playing a role in host defense and tissue repair, the immune microenvironment also plays a key role in the outcome of various diseases, including cancer, autoimmune diseases, and chronic inflammation. For example, in cancer, the immune microenvironment can promote or suppress tumor growth and progression, depending on the balance between pro-inflammatory and anti-tumor immune responses. Understanding the complexity of the immune microenvironment is crucial for developing new therapeutic strategies aimed at modulating immune responses for effective disease treatment. Research on the immune microenvironment can help harness the power of the immune system to combat disease and restore tissue homeostasis. SUMMARY

[0006] The technical problem to be solved by the present application is how to predict the responsiveness of patients to immunotherapy and / or how to accurately predict the responsiveness of disease patients (subjects) to immunotherapy.

[0007] To solve the above technical problem, the present application first provides a data processing device comprising a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement a method for predicting immunotherapy response, which can comprise the following steps:

[0008] S1) Data receiving: receiving the expression amount of genes in cell subpopulations in single-cell sequencing data of a subject to be tested;

[0009] S2) Data processing: substituting the expression amount of genes in cell subpopulations in S1) into an immunotherapy response prediction model to obtain a prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to a method comprising the following steps: receiving the expression amount M of genes in cell subpopulations in single-cell sequencing data of immunotherapy-responsive subjects and the expression amount N of genes in cell subpopulations in single-cell sequencing data of immunotherapy-non-responsive subjects; using the expression amount of genes in cell subpopulations that are different between the expression amount M and the expression amount N as input and using the immunotherapy response result as output, constructing an immunotherapy response model using a machine learning algorithm;

[0010] S3) Data output: outputting the prediction result of the immunotherapy response of the subject to be tested.

[0011] In the data processing device, the machine learning algorithm can be a random forest algorithm.

[0012] In the data processing device, the subject can be a tumor patient, or a person suffering from one or more diseases. The immunotherapy response result can be response or non-response to immunotherapy.

[0013] The data processing device described above can be a computer device.

[0014] To solve the above technical problems, the present application also provides a method for constructing an immunotherapy response model, which can comprise taking the expression amount of a gene in a cell subpopulation that is different between single-cell sequencing data of an immunotherapy response subject and single-cell sequencing data of an immunotherapy non-response subject as input, taking an immunotherapy response result as output, and constructing an immunotherapy response model using a machine learning algorithm.

[0015] The expression amount of the gene in the cell subpopulation described above can be the expression amount of a cell subpopulation and a specific gene of the cell subpopulation. The specific gene of the cell subpopulation is a gene that is expressed in each cell of the cell subpopulation.

[0016] In the method, the machine learning algorithm can be a random forest algorithm. The subject can be a tumor patient, or a person suffering from one or more diseases. The immunotherapy response result can be response or non-response to immunotherapy.

[0017] To solve the above technical problems, the present application also provides a device for predicting or assisting in predicting the immunotherapy response of a subject, which can comprise the following modules:

[0018] A1) Data receiving module: for receiving the expression amount of a gene in a cell subpopulation in single-cell sequencing data of a subject to be tested;

[0019] A2) Data processing module: for substituting the expression amount of the gene in the cell subpopulation in A1) into an immunotherapy response prediction model to obtain a prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to a method comprising the following steps: receiving the expression amount M of a gene in a cell subpopulation in single-cell sequencing data of an immunotherapy response subject and the expression amount N of a gene in a cell subpopulation in single-cell sequencing data of an immunotherapy non-response subject; taking the expression amount of the gene in the cell subpopulation that is different between the expression amount M and the expression amount N as input, taking an immunotherapy response result as output, and constructing an immunotherapy response model using a machine learning algorithm;

[0020] A3) data output module: for outputting the prediction result of the immune therapy response of the subject to be tested.

[0021] In the above device, the machine learning algorithm can be a random forest algorithm. The subject can be a tumor patient, and the subject can also be a person with one or more diseases. The immune therapy response result can be responsive or non-responsive to immune therapy.

[0022] The expression amount of the gene in the above-mentioned cell subpopulation can be the expression amount of the cell subpopulation and a specific gene of the cell subpopulation. The specific gene of the cell subpopulation is a gene that is expressed in each cell of the cell subpopulation.

[0023] To solve the above technical problems, the present application also provides a method for predicting or assisting in predicting the immune therapy response of a subject, which can include the following steps:

[0024] B1) data receiving: receiving the expression amount of the gene in the cell subpopulation in the single-cell sequencing data of the subject to be tested;

[0025] B2) data processing: substituting the expression amount of the gene in the cell subpopulation in B1) into an immune therapy response prediction model to obtain the prediction result of the immune therapy response of the subject to be tested; the immune therapy response prediction model is constructed according to a method comprising the following steps: receiving the expression amount of the gene in the cell subpopulation in the single-cell sequencing data of the immune therapy responsive subject M and the expression amount of the gene in the cell subpopulation in the single-cell sequencing data of the immune therapy non-responsive subject N; using the machine learning algorithm to construct the immune therapy response model by taking the expression amount of the gene in the cell subpopulation that is different between the expression amount M and the expression amount N as input and taking the immune therapy response result as output;

[0026] B3) data output: outputting the prediction result of the immune therapy response of the subject to be tested.

[0027] In the above method, the machine learning algorithm can be a random forest algorithm. The subject can be a tumor patient, and the subject can also be a person with one or more diseases. The immune therapy response result can be responsive or non-responsive to immune therapy.

[0028] To solve the above technical problems, the present application also provides a computer readable storage medium storing a computer program, which can make a computer execute the following steps:

[0029] C1) data receiving: receiving the expression amount of the gene in the cell subpopulation in the single-cell sequencing data of the subject to be tested;

[0030] C2) data processing: inputting the expression amount of the gene under the cell subpopulation in C1) into an immunotherapy response prediction model to obtain a prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to a method comprising the following steps: receiving the expression amount of the gene under the cell subpopulation in the single-cell sequencing data of the immunotherapy response subject and the expression amount of the gene under the cell subpopulation in the single-cell sequencing data of the immunotherapy non-response subject; using the expression amount of the gene under the cell subpopulation that is different between the expression amount M and the expression amount N as input, and using the immunotherapy response result as output, constructing an immunotherapy response model using a machine learning algorithm;

[0031] C3) data output: outputting the prediction result of the immunotherapy response of the subject to be tested.

[0032] In the above computer-readable storage medium, the machine learning algorithm can be a random forest algorithm.

[0033] The subject described above can be a tumor patient, and the subject can also be a person suffering from one or more diseases. The immunotherapy response result can be response or non-response to immunotherapy.

[0034] The expression amount of the gene under the cell subpopulation described above can be the expression amount of the cell subpopulation and a specific gene of the cell subpopulation. The specific gene of the cell subpopulation is the same gene expressed in each cell of the cell subpopulation.

[0035] The method described above can not include the step of obtaining a biological sample from an animal body. The method can not be directed to a living human body or animal body, but only to data. The method can be an information processing method in which all steps are implemented by a computer or other data processing device.

[0036] The above application or method is a non-disease diagnosis application or method. The above application or method is not directly aimed at obtaining a disease diagnosis result or health condition of a living human body or animal body.

[0037] The above application or method is a non-disease treatment application or method. The above application or method is not aimed at restoring or obtaining health or reducing pain of a living human body or animal body.

[0038] The application establishes a model for predicting the responsiveness of patients to immunotherapy and explores the key factors affecting the responsiveness of immunotherapy. First, the application uses random forest algorithm, linear support vector machine and deep neural network to establish a prediction model, respectively taking the gene expression of each gene in all cells or the expression of the gene in each cell subpopulation as the feature, to establish a treatment responsiveness model of the patient, and compare the performance. The application uses Shapley Value and LIME and other explainable methods, or uses the characteristics of the model to calculate the importance of the features, and according to the importance, the features are screened in a loop to reduce the feature dimension for training.

[0039] The beneficial effects of the application are:

[0040] The application establishes a model for predicting the responsiveness of patients to immunotherapy, and analyzes the importance of input features. Unlike traditional differential gene analysis, the importance obtained by using explainable machine learning methods does not ignore the influence of the interaction between features.

[0041] By establishing a prediction model for the immunotherapy responsiveness of patients, the application will promote the individualized medical treatment of immunotherapy. Different patients have different responses to immunotherapy, and by analyzing the cell subpopulation-gene expression level features, a more accurate treatment plan can be developed for the patient, the treatment effect is improved, unnecessary side effects are reduced, the survival rate and quality of life of the patient are improved, and the medical cost is also reduced, which brings long-term economic benefits to the patient and the medical system.

[0042] The application develops a prediction model based on machine learning algorithm, taking the expression level of genes in cell subpopulations as input, and taking whether the patient responds to immunotherapy as output, and using the expression level of genes in cell subpopulations as input, which is more detailed and more helpful to analyze the mechanism of specific genes in specific cell subpopulations than directly using gene expression level. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 is a flowchart of the application.

[0044] Figure 2 is a table of the prediction performance of different models when using gene expression level as input. Among them, RF represents random forest, SVM represents support vector machine, and DL represents deep learning model; ACC and AUC are the average performance of the model after training on randomly divided data sets (50 times random division).

[0045] Figure 3 is the flowchart of the application for screening input features in a loop.

[0046] Figure 4 is the importance ranking of cell subpopulation-genes (top 20).

[0047] Figure 5 ROC curve of the model trained by the expression data of genes in cell subpopulation.

[0048] Figure 6 Shapley-Value of cell subpopulation-genes. The positive direction of the horizontal axis represents response, red represents high expression level, and blue represents low expression level.

[0049] Figure 7 Importance ranking of SBM correlation graph for different inputs. Raw represents cell subpopulation-gene expression level data, Ratio represents cell subpopulation proportion data, ‘+’ represents data splicing, ‘*’ (i.e. Raw*Ratio) represents multiplying the expression of genes in cell subpopulation by the proportion of corresponding cell subpopulation, I represents using the same method after changing the random seed, Random represents multiple average values compared with random sequences. Raw*Ratio+Ratio represents multiplying the expression of genes in cell subpopulation by the proportion of corresponding cell subpopulation, and then splicing the expression of genes in cell subpopulation.

[0050] Figure 8 Performance comparison of different models when using the expression of genes in cell subpopulation as input. RF represents random forest, SVM represents support vector machine, and DL represents deep learning model. ACC represents accuracy, i.e. the number of correctly predicted samples / total number of samples; AUC represents the area under the ROC curve. ACC and AUC are both the average performance of the model after training on randomly divided data sets (100 times random division). DETAILED DESCRIPTION

[0051] The present application will be further described in conjunction with the specific embodiments. The examples given are only to illustrate the present application, and are not intended to limit the scope of the present application. The examples provided below can serve as a guide for further improvement by those skilled in the art, and do not constitute any limitation on the present application in any way.

[0052] The experimental methods in the following examples are all routine methods, and are performed according to the techniques or conditions described in the literature in the art or according to the product instructions, unless otherwise specified. The materials, reagents, etc. used in the following examples can be obtained from commercial channels, unless otherwise specified.

[0053] The following examples use Python 3 software to process data, and the experimental results are expressed as mean ± standard deviation, Mann-Whitney test is used, P<0.05 (*) indicates significant difference, P<0.01 (**) indicates extremely significant difference, and P<0.001 (***) indicates extremely significant difference.

[0054] Some algorithms used in the embodiments of the application are as follows:

[0055] LIME (Local Interpretable Model-Agnostic Explanations) is a machine learning model explanation tool that explains the decision of a black box model by generating a locally interpretable model. Machine learning models are often difficult to explain due to their complex internal structure and parameters. LIME does not attempt to understand the entire model, but rather explains the contribution of features to the prediction result at the feature level, which to some extent reflects the importance of the features. The method first perturbs the features of a specific sample to generate some samples; then uses these perturbed samples as inputs to the model to be explained to obtain a set of prediction results; then uses Lasso regression or tree-based feature importance techniques to select a subset of features with the highest correlation with the prediction results; then trains an interpretable model, such as linear regression and decision tree, on the perturbed samples of these features to predict the model output; finally, the original model's prediction for this sample is explained by the interpretable model.

[0056] Shapley Value is a concept in cooperative game theory, which also plays a role in explaining the prediction of machine learning models. It calculates the average marginal contribution of each feature to the model prediction by traversing all feature combinations, and assigns a relative importance score to each feature. The idea behind this method is that features may interact in complex ways, and their contribution to the prediction result should be jointly evaluated rather than individually evaluated. Like the LIME method, Shapley Value calculates the contribution of features to the model prediction for a single sample, so to calculate the feature importance for multiple samples, the feature importance for a single sample is averaged.

[0057] SBM is a method for measuring the similarity between rankings. It selects the top k elements of two rankings, calculates the intersection size / k, and traverses k to get the average, and finally gets the similarity size of the two rankings.

[0058] The data used in the embodiments of the present application are 2.42 million single-cell RNA-Seq sequencing data of 550 patients of 13 cancer types in a pan-cancer immune microenvironment (public data collected from the GEO database, and the PMIDs of the relevant literature are: 27124452; 30250229; 30388455; 30388456; 31227543; 31359002; 31588021; 32405063; 32497499; 32949350; 33711272; 33723257; 33861994; 33958794; 34290408; 34653364; 34653365; 34836966; 35108529; 35121991 and 36719749).

[0059] Embodiment 1. Construction and comparative analysis of different immune therapy response prediction models.

[0060] The present application uses the random forest algorithm, the linear support vector machine (SVM) and the deep neural network (DL) to establish a prediction model of the patient's response to the immunotherapy, specifically, the model takes the features as the input and the patient's response as the output. The features refer to the gene expression of each gene in all cells or the expression of each cell subpopulation.

[0061] The random forest algorithm can directly generate a set of feature importance values after training; for the linear support vector machine, the present application uses the absolute value of its weight as the feature importance value; for the deep neural network, the present application uses the Shapley Value and the LIME method to calculate the feature importance.

[0062] In order to more intuitively show the influence of feature changes on the response prediction, the present application also uses the Shapley Value to show the feature importance value of the prediction model based on the random forest algorithm and the linear support vector machine. The Shapley Value and the LIME method measure the influence of a set of feature changes on the model output to reflect the importance of the features, so the directly obtained feature importance value is individual level and has positive and negative. In order to obtain the importance value at the data set level, the present application takes the absolute value of the importance of all individual units and sums them up, thereby obtaining the importance ranking of the features.

[0063] 1. Establishing an immune therapy response prediction model with gene expression level as input.

[0064] In order to explore whether the expression level of each gene is related to the immune therapy response of the patient, the present application takes the average value of the expression of each gene in all cells in the single-cell sequencing data as the input for model training.

[0065] The present application uses random forest, support vector machine and deep learning method to train immunotherapy response model respectively, and the performance of each model is as follows Figure 2 as shown in the table, where the performance of support vector machine is the best, and the performance of deep learning is significantly lower than the other two methods, because the feature dimension of the data set is much higher than the sample size, and deep learning uses matrix multiplication to transfer information between layers by its nature, which is likely to overfit on a small number of features. The model based on decision tree or support vector machine is more suitable for this scenario, so the performance is better.

[0066] Since the support vector machine algorithm has better performance Figure 2 , the present application uses the support vector machine algorithm to train repeatedly, and the weights represent the importance of the genes. After 50 rounds of training on different divided data sets, the mean value of the importance ranking is obtained, and the ranking is used as the feature importance ranking.

[0067] The gene expression level is taken as input, and the average ranking of GNL Y gene in the obtained feature importance value is the highest, but since the gene expression has cell subpopulation specificity, directly taking the average value of the expression of each gene in all cells as input loses this part of the difference, and it is difficult to analyze the relationship between specific genes and immunotherapy responsiveness (helpful for prediction but not helpful for specific mechanism research, and may not have biological significance). Therefore, studying the gene expression difference at the cell subpopulation level is of great help to more detailed analysis.

[0068] 2. Establish an immunotherapy responsiveness prediction model with the expression level of genes under cell subpopulation as input.

[0069] 2.1 Model input feature screening.

[0070] Since the number of combinations of genes under cell subpopulation (cell subpopulation-gene, i.e. cell subpopulation and a specific gene of the cell subpopulation) in the sequencing data is as high as 3.8×10 6 , and the sample size is much smaller than this order of magnitude, directly applying the original data for model training not only greatly reduces the algorithm efficiency, but also may lead to overfitting of the model on a small number of features. Therefore, the present application first uses Mann-Whitney test to screen out 1.3×10 5 genes under cell subpopulation with significantly different expression in patients with immunotherapy response and immunotherapy non-response to obtain differentially expressed cell subpopulation-gene combinations, and then use the following method to screen the features used in the final model. In addition, since Shapley Value and LIME are based on feature perturbation, the high dimension greatly increases the computational complexity. In order to avoid these negative effects, the present application needs to pre-screen the features to reduce the feature dimension for training.

[0071] Firstly, the present application uses Mann-Whitney U test to screen cell subpopulation-gene combinations with significant difference in expression amount between patients with immune therapy response and patients without immune therapy response, and the expression amount of each cell subpopulation-gene combination and the genes in the combination (the expression amount of the genes under the cell subpopulation) is taken as the pre-input feature of the model, and the initialization average importance value of the pre-input feature is obtained by random initialization.

[0072] Then, the present application cyclically samples part of the pre-input features as input to train the random forest model, and at the same time, according to the obtained model, the average importance value of the corresponding feature is updated. Finally, the features used for model training are sorted according to the importance value. Figure 3 )。

[0073] The feature screening is divided into multiple stages, and the number of sampled features decreases by stage, and each stage is cycled 100 times. In each cycle, the features are sampled according to the probability distribution shown in the following formula (1):

[0074]

[0075] In formula (1), w i is the average importance value of the i-th feature, and all w i The initial value is set to 1. After training the model, the average importance of each feature is updated using exponential moving average as follows:

[0076]

[0077] In formula (2), X i represents an indicator variable indicating whether the i-th feature is selected, v i is the importance value of the i-th feature in the model in this cycle, and a is the decay factor.

[0078] After multiple rounds of training, the top 200 features (important features) in the importance value are selected as the input of the final model.

[0079] After the cyclic screening, the present application obtains a set of importance ranking Figure 4 In subsequent experiments, the present application trains the model using the top 200 features.

[0080] 2.2 Only use the expression level data of the genes under the cell subpopulation.

[0081] Receiver Operating Characteristic (ROC) curve is a tool for evaluating the performance of a classification model. It plots the True Positive Rate (TPR) on the y-axis and the False Positive Rate (FPR) on the x-axis, reflecting the performance of the model at different thresholds. The ROC curve of the model trained using cell subpopulation-gene expression level data is relatively full, indicating good performance. Figure 5

[0082] Figure 6 It is shown that high expression of PRSS23, GNLY, and ZEB2 in the CD8+ Temra cell subpopulation often leads the model to predict non-response, and high expression of CXCR4 in the ISG15+ macrophage subpopulation often leads the model to predict response. In the importance ranking of all cell subpopulation-gene expression levels, the gene importance ranking of some subpopulations is generally in the front, such as the CD8+ Temra cell subpopulation, and perhaps the proportion of subpopulations has an impact.

[0083] 2.3 Cell subpopulation proportion is not a determining factor for the ranking of some subpopulations.

[0084] In the experiment, the present application found that the gene importance ranking of some cell subpopulations was generally high, such as TE07-CD8-Temra-CXCR1 and M13-Mph-ISG15. In order to study whether the ranking of genes in these cell subpopulations is caused by the proportion of cell subpopulations, the present application multiplies the cell subpopulation-gene expression level data and the cell subpopulation proportion data, splices them, and performs a second round of screening experiment, obtaining different importance rankings, and the SBM correlation is shown in Figure 7 .

[0085] In fact, after multiplying the cell subpopulation-gene expression level data and the corresponding cell subpopulation proportion, the cell subpopulation-gene expression quantity is equivalent to the sum of gene expression quantities in a specific cell subpopulation in an individual divided by the number of cells in that cell population, which actually reflects the gene expression quantity at the cell subpopulation level, to some extent, introducing the influence of cell subpopulation proportion, because the number of gene expression quantities in the cell subpopulation with high proportion will also be biased, therefore, the difference in gene expression level of the cell subpopulation with large difference in proportion between responders and non-responders may be enlarged, so that the model prediction is more focused on these features.

[0086] ​Due to the calculation method of SBM, the SBM mean value of a large number of random sequences and a certain sequence converges to 0.5, the importance ranking obtained by using Raw*Ratio (multiplying the expression of the gene in the cell subpopulation by the proportion of the corresponding cell subpopulation) and Raw*Ratio+Ratio data (multiplying the expression of the gene in the cell subpopulation by the proportion of the corresponding cell subpopulation, and then splicing the expression data of the gene in the cell subpopulation) and the SBM value of Raw data are all close to 0.5, which indicates that the differences are large; the SBM values between the sequences obtained by Raw and Raw+Ratio data, the SBM values between the sequences obtained by Raw*Ratio and Raw*Ratio+Ratio data are close to the SBM values of the original sequence and the sequence obtained after changing the random seed, which indicates that the sequences are relatively close, and to some extent, it reflects that the importance ranking of cell subpopulation-gene expression level is not caused by the proportion of cell subpopulation.

[0087] 2.4 Comparison of performance and importance ranking of different methods

[0088] In addition to the random forest model directly outputting feature importance and using Shapley Value to explain the importance of input features in the random forest model, the present application also attempts to use the weight of a linear support vector machine to represent feature importance, use a support vector machine combined with Shapley Value or LIME method to represent feature importance, use a deep learning model combined with LIME method to represent feature importance, and the like, and the prediction performance of the model is shown in Table 1. Figure 8 Among them, the random forest algorithm (RF in Figure 8 ) performs best.

[0089] The present application has been described in detail above. For those skilled in the art, without departing from the purpose and scope of the present application, and without unnecessary experiments, the present application can be implemented in a wider range under the same parameters, concentrations and conditions. Although the present application gives a special embodiment, it should be understood that further improvements can be made to the present application. In summary, according to the principle of the present application, the present application intends to include any changes, uses or improvements of the present application, including changes made by conventional techniques known in the art, which are outside the scope disclosed in the present application.

Claims

1. A data processing apparatus, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement a method for predicting an immunotherapy response, the method comprising the following steps: S1) data receiving: receiving the expression amounts of genes in cell subpopulations in single-cell sequencing data of a subject to be tested; S2) data processing: inputting the expression amounts of genes in cell subpopulations in S1) into an immunotherapy response prediction model to obtain a prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to a method comprising the following steps: receiving expression amounts of genes in cell subpopulations in single-cell sequencing data of an immunotherapy-responsive subject and expression amounts of genes in cell subpopulations in single-cell sequencing data of an immunotherapy-non-responsive subject; using the expression amounts of genes in cell subpopulations that are different between the expression amount M and the expression amount N as input, and using the immunotherapy response result as output, an immunotherapy response model is constructed using a machine learning algorithm; The construction of the immunotherapy response model further comprises model input feature screening, and the model input features are obtained according to the following method: First, the Mann-Whitney U test is used to screen cell subpopulation-gene combinations with significant differences in expression amounts between immunotherapy-responsive patients and immunotherapy-non-responsive patients, each cell subpopulation-gene combination and the expression amounts of genes in the combination are used as pre-input features of the model, and the initialization average importance values of the pre-input features are obtained by random initialization; Then, part of the pre-input features are cyclically sampled as input to train a random forest model, and at the same time, the average importance values of the corresponding features are updated according to the obtained model; finally, the features used for model training are obtained by sorting according to the importance values; The feature screening is divided into multiple stages, and the number of sampled features decreases by stage; each stage is cycled 100 times; in each cycle, the features are sampled according to the probability distribution shown in the following formula (1): In formula (1), w i is the average importance value of the i-th feature, and all w i The initial value is set to 1; after training the model, the average importance of each feature is updated using the exponential moving average according to the following formula (2): In formula (2), X i is an indicator variable representing whether the i-th feature is selected, v i is the importance value of the i-th feature to the model in this cycle, and a is a decay factor. After multiple rounds of training, the top 200 features in terms of importance values are selected as the input of the final model; S3) data output: outputting the prediction result of the immunotherapy response of the subject to be tested.

2. The data processing apparatus according to claim 1, characterized in that: The subject is a tumor patient.

3. A method of constructing a model for predicting immunotherapy response, characterized by: The method comprises using the expression amounts of genes in cell subpopulations that are different between single-cell sequencing data of an immunotherapy-responsive subject and single-cell sequencing data of an immunotherapy-non-responsive subject as input, and using the immunotherapy response result as output, an immunotherapy response model is constructed using a machine learning algorithm; The method further comprises model input feature screening, and the model input features are obtained according to the following method: First, the Mann-Whitney U test is used to screen cell subpopulation-gene combinations with significant differences in expression amounts between immunotherapy-responsive patients and immunotherapy-non-responsive patients, each cell subpopulation-gene combination and the expression amounts of genes in the combination are used as pre-input features of the model, and the initialization average importance values of the pre-input features are obtained by random initialization; Then, part of the pre-input features are cyclically sampled as input to train a random forest model, and at the same time, the average importance values of the corresponding features are updated according to the obtained model; finally, the features used for model training are obtained by sorting according to the importance values; The feature screening is divided into multiple stages, the number of features sampled decreases by stage, and each stage is looped 100 times; in each loop, the features are sampled according to the probability distribution shown in the following formula (1): In formula (1), w i is the average importance value of the i-th feature, and all w i The initial value is set to 1; after training the model, the average importance of each feature is updated using the exponential moving average according to the following formula (2): In formula (2), X i is an indicator variable representing whether the i-th feature is selected, v i is the importance value of the i-th feature to the model in this cycle, and a is a decay factor. After multiple rounds of training, the top 200 features in terms of importance value are selected as the input of the final model.

4. A device for predicting or assisting in predicting the immunotherapy response of a subject, characterized in that: The device comprises the following modules: A1) Data receiving module: for receiving the expression amount of genes under cell subpopulation in single cell sequencing data of the to-be-tested subject; A2) Data processing module: for substituting the expression amount of genes under cell subpopulation in A1) into an immunotherapy response prediction model to obtain a prediction result of the immunotherapy response of the to-be-tested subject; the immunotherapy response prediction model is constructed according to a method comprising the following steps: receiving the expression amount of genes under cell subpopulation in single cell sequencing data of an immunotherapy response subject M and the expression amount of genes under cell subpopulation in single cell sequencing data of an immunotherapy non-response subject N; using the expression amount of genes under cell subpopulation that are different between the expression amount M and the expression amount N as input, and using the immunotherapy response result as output, an immunotherapy response model is constructed using a machine learning algorithm; The construction of the immunotherapy response model further comprises model input feature screening, and the model input features are obtained according to the following method: First, the Mann-Whitney U test is used to screen cell subpopulation-gene combinations with significant differences in expression amount between immunotherapy response and immunotherapy non-response patients, each cell subpopulation-gene combination and the expression amount of genes in the combination are used as pre-input features of the model, and the initial average importance value of the pre-input features is obtained by random initialization; Then, part of the pre-input features are sampled as input to train the random forest model, and at the same time, the average importance value of the corresponding features is updated according to the obtained model; finally, the features used for model training are sorted according to the importance value; The feature screening is divided into multiple stages, the number of features sampled decreases by stage, and each stage is looped 100 times; in each loop, the features are sampled according to the probability distribution shown in the following formula (1): In formula (1), wi is the average importance value of the i th feature, and the initial value of all wi is set to 1; after training the model, the average importance of each feature is updated using the exponential moving average according to the following formula (2): In formula (2), Xi represents an indicator variable indicating whether the i th feature is selected, vi is the importance value of the i th feature in the model in this loop, and a is a decay factor; After multiple rounds of training, the top 200 features in terms of importance value are selected as the input of the final model. A3) Data output module: for outputting the prediction result of the immunotherapy response of the to-be-tested subject.

5. A method of predicting or aiding in the prediction of the immunotherapy response of a subject, characterized in that: The method comprises the following steps: B1) Data receiving: receiving the expression amount of genes under cell subpopulation in single cell sequencing data of the to-be-tested subject; B2) data processing: substituting the expression amount of the genes under the cell subpopulation of B1) into an immunotherapy response prediction model to obtain a prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to a method comprising the following steps: receiving the expression amount of the genes under the cell subpopulation in the single-cell sequencing data of the immunotherapy response subject and the expression amount of the genes under the cell subpopulation in the single-cell sequencing data of the immunotherapy non-response subject; using the expression amount of the genes under the cell subpopulation with differences in the expression amount M and the expression amount N as input, and using the immunotherapy response result as output, an immunotherapy response model is constructed using a machine learning algorithm; The construction of the immunotherapy response model further comprises model input feature screening, and the model input features are obtained according to the following method: First, the Mann-Whitney U test is used to screen the cell subpopulation-gene combinations with significant differences in expression amount between the immunotherapy response and the immunotherapy non-response patients, and each cell subpopulation-gene combination and the expression amount of the genes in the combination are used as the pre-input features of the model. The initialization average importance value of the pre-input features is obtained by random initialization; Then, part of the pre-input features are cyclically sampled as input to train the random forest model, and at the same time, the average importance value of the corresponding features is updated according to the obtained model. Finally, the features used for model training are sorted according to the importance value to obtain the features used for model training; The feature screening is divided into multiple stages, and the number of features sampled decreases by stage. Each stage is cycled 100 times. In each cycle, the features are sampled according to the probability distribution shown in the following formula (1): In formula (1), w i is the average importance value of the i-th feature, and all w i The initial value is set to 1; after training the model, the average importance of each feature is updated using the exponential moving average according to the following formula (2): In formula (2), X i is an indicator variable representing whether the i-th feature is selected, v i is the importance value of the i-th feature to the model in this cycle, and a is a decay factor. After multiple rounds of training, the top 200 features in the importance value are selected as the input of the final model; B3) data output: outputting the prediction result of the immunotherapy response of the subject to be tested.

6. A computer readable storage medium storing a computer program, wherein the computer program causes a computer to execute the following steps: C1) data receiving: receiving the expression amount of the genes under the cell subpopulation in the single-cell sequencing data of the subject to be tested; C2) data processing: substituting the expression amount of the genes under the cell subpopulation of C1) into an immunotherapy response prediction model to obtain a prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to a method comprising the following steps: receiving the expression amount of the genes under the cell subpopulation in the single-cell sequencing data of the immunotherapy response subject and the expression amount of the genes under the cell subpopulation in the single-cell sequencing data of the immunotherapy non-response subject; using the expression amount of the genes under the cell subpopulation with differences in the expression amount M and the expression amount N as input, and using the immunotherapy response result as output, an immunotherapy response model is constructed using a machine learning algorithm; The construction of the immunotherapy response model further comprises model input feature screening, and the model input features are obtained according to the following method: Firstly, the cell subpopulation-gene combinations with significant differences in expression between patients with immune therapy response and patients without immune therapy response are screened using the Mann-Whitney U test, and each cell subpopulation-gene combination and the expression of genes in the combination are taken as pre-input features of the model, and the initialization average importance value of the pre-input features is obtained by random initialization; Then, part of the pre-input features are sampled as input to train the random forest model, and at the same time, the average importance value of the corresponding features is updated according to the obtained model; finally, the features used for model training are sorted according to the importance value to obtain the features used for model training; The feature screening is divided into multiple stages, and the number of features sampled decreases by stage, and each stage is cycled 100 times; in each cycle, the features are sampled according to the probability distribution shown in the following formula (1): In formula (1), w i is the average importance value of the i-th feature, and all w i The initial value is set to 1; after training the model, the average importance of each feature is updated using the exponential moving average according to the following formula (2): In formula (2), X i is an indicator variable representing whether the i-th feature is selected, v i is the importance value of the i-th feature to the model in this cycle, and a is a decay factor. After multiple rounds of training, the top 200 features in the importance value are selected as the input of the final model; C3) data output: output the prediction result of the immune therapy response of the subject to be tested.

Citation Information

Patent Citations

  • Transcriptome-based PD-1 therapy treatment effect prediction system

    CN111755073A

  • Treatment response prediction method and system, storage medium and terminal

    CN119541847A