Data processing device and method for predicting immunotherapy responsiveness of patient

By constructing an immunotherapy response prediction model based on random forest algorithm, the expression of genes under the cell subpopulations in single-cell sequencing data is solved in the problem of immunotherapy response prediction in tumor patients, and individualized treatment and cost optimization are achieved.

CN120260905AActive Publication Date: 2025-07-04PEKING UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510282890.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-04
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the response of tumor patients to immunotherapy, resulting in poor treatment effects and high cost.

Method used

The random forest algorithm was used to construct an immunotherapy response prediction model, and the expression of genes under the cell subpopulations in single-cell sequencing data was used as input to predict the patient's response to immunotherapy.

Benefits of technology

By carefully analyzing the cell subpopulation-gene expression levels, individualized treatment plans are provided to improve treatment effects, reduce side effects, and reduce medical costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260905A_ABST
    Figure CN120260905A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing device and method for predicting immunotherapy responsiveness of a patient in the field of bioinformatics. The invention aims to solve the technical problem of how to accurately predict the response of a tumor patient to immunotherapy. According to the method, expression quantity characteristics of genes under cell subpopulations with different expressions in single cell sequencing data of a subject with immunotherapy response and a subject without response serve as input, whether a patient responds to immunotherapy or not serves as output, and a random forest algorithm is used for constructing an immunotherapy response prediction model; by inputting the expression quantity of the gene under the cell subset in single cell sequencing data of a to-be-detected subject into the model, an immunotherapy response result of the to-be-detected subject can be predicted and obtained. According to the invention, a more accurate treatment scheme can be formulated for tumor patients, unnecessary side effects are reduced, the medical cost is reduced, and the method can be applied to individualized medical treatment of immunotherapy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics, and particularly relates to a data processing device and method for predicting the immunotherapy responsiveness of patients. Background Art

[0002] Immunotherapy is a new therapy that has attracted much attention in cancer treatment. However, different patients have significant individual differences in their responsiveness to immunotherapy. To exert the effect of immunotherapy and reduce the treatment cost, studying the responsiveness of patients to immunotherapy and its mechanism is an important issue in tumor immunotherapy.

[0003] With the rapid development of single-cell sequencing technology, there has been a huge leap in both the quality and quantity of single-cell sequencing datasets of the tumor immune microenvironment. By studying the single-cell sequencing data of patients, it is possible to more comprehensively understand the expression changes of different cell types during the disease process, providing new possibilities for individualized treatment. However, the emergence of a large amount of data has also brought new challenges to data analysis. With the reduction of sequencing costs, the scale and complexity of biological data have increased significantly, and the influence of preprocessing such as data screening and noise reduction on experimental results has become increasingly large. A reasonable data processing method has gradually become an essential means to obtain scientific conclusions. Along with the increase in data scale, the field of machine learning has also developed rapidly. From statistical learning to large models based on neural networks, the technological innovation in the field of machine learning has also provided new tools for biological data processing.

[0004] In the field of cancer immunotherapy, there is an important challenge, that is, how to accurately predict the response of tumor patients to immunotherapy and screen features that are helpful for predicting immunotherapy response, so as to assist clinical prediction. There are already some machine learning algorithms that can use single-cell sequencing data to predict the responsiveness of patients. However, which data is more important for prediction is still a challenging problem.

[0005] The immune microenvironment refers to the complex network of immune cells, molecules, and structures present within tissues or organs. It plays a crucial role in regulating immune responses and maintaining tissue homeostasis. The immune microenvironment is highly dynamic and can change significantly depending on factors such as tissue type, physiological state, and the presence of disease. At its core are various types of immune cells, including lymphocytes (such as T cells, B cells, and natural killer cells), antigen-presenting cells (such as dendritic cells and macrophages), and other specialized cells (such as mast cells and granulocytes). These cells interact with each other and with non-immune cells within the tissue (such as epithelial cells, fibroblasts, and endothelial cells) to form a complex network of cell interactions. In addition to its role in host defense and tissue repair, the immune microenvironment also plays a key role in shaping the outcomes of various diseases, including cancer, autoimmune diseases, and chronic inflammation. For example, in cancer, the immune microenvironment can either promote or inhibit tumor growth and progression, depending on the balance between pro-inflammatory and anti-tumor immune responses. Understanding the complexity of the immune microenvironment is crucial for developing new treatment strategies aimed at modulating immune responses to effectively treat diseases. Research on the immune microenvironment can help harness the power of the immune system to combat diseases and restore tissue homeostasis. Summary of the Invention

[0006] The technical problem to be solved by the present invention is how to predict the responsiveness of a patient to immunotherapy and / or how to accurately predict the responsiveness of a disease patient (subject) to immunotherapy.

[0007] To solve the above technical problem, the present invention first provides a data processing device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement a method for predicting immunotherapy response, and the method may include the following steps:

[0008] S1) Data reception: Receive the expression levels of genes under cell subsets in the single-cell sequencing data of the subject to be tested.

[0009] S2) Data processing: Substitute the expression levels of genes under the cell subsets in S1) into the immunotherapy response prediction model to calculate and obtain the prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to the method including the following steps: Receive the expression levels M of genes under cell subsets in the single-cell sequencing data of subjects responsive to immunotherapy and the expression levels N of genes under cell subsets in the single-cell sequencing data of subjects non-responsive to immunotherapy; Use the expression levels of genes under the cell subsets with differences between the expression levels M and the expression levels N as inputs, and the immunotherapy response result as the output, and use a machine learning algorithm to construct an immunotherapy response model.

[0010] S3) Data output: Output the prediction result of the immunotherapy response of the subject to be tested.

[0011] In the above data processing device, the machine learning algorithm can be a random forest algorithm.

[0012] In the above data processing device, the subject can be a cancer patient, or can also be a person suffering from one or more diseases. The immunotherapy response result can be a response or no response to immunotherapy.

[0013] The data processing device described above can be a computer device.

[0014] To solve the above technical problems, the present invention also provides a method for constructing a model for predicting immunotherapy response. The method may include using the expression levels of genes in cell subsets with differences in the single-cell sequencing data of subjects with a response to immunotherapy and the single-cell sequencing data of subjects without a response to immunotherapy as inputs, using the immunotherapy response result as an output, and constructing an immunotherapy response model using a machine learning algorithm.

[0015] The expression level of a gene under the above cell subset can be the expression level of a cell subset and a specific gene of this cell subset. The specific gene of this cell subset is the same gene expressed in each cell of this cell subset.

[0016] In the above method, the machine learning algorithm can be a random forest algorithm. The subject can be a cancer patient, or can also be a person suffering from one or more diseases. The immunotherapy response result can be a response or no response to immunotherapy.

[0017] To solve the above technical problems, the present invention also provides a device for predicting or assisting in predicting the immunotherapy response of a subject. The device may include the following modules:

[0018] A1) Data receiving module: used to receive the expression levels of genes in cell subsets in the single-cell sequencing data of the subject to be tested;

[0019] A2) Data processing module: used to substitute the expression levels of genes in the cell subsets described in A1) into the immunotherapy response prediction model to calculate and obtain the prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to the method including the following steps: receiving the expression levels M of genes in cell subsets in the single-cell sequencing data of subjects with a response to immunotherapy and the expression levels N of genes in cell subsets in the single-cell sequencing data of subjects without a response to immunotherapy; using the expression levels of genes in cell subsets with differences between the expression levels M and the expression levels N as inputs, using the immunotherapy response result as an output, and constructing an immunotherapy response model using a machine learning algorithm;

[0020] A3) Data output module: used to output the prediction result of the immunotherapy response of the subject to be tested.

[0021] In the above device, the machine learning algorithm can be a random forest algorithm. The subject can be a cancer patient, or the subject can also be a person suffering from one or more diseases. The immunotherapy response result can be a response or no response to immunotherapy.

[0022] The expression level of the gene under the above-mentioned cell subset can be the expression level of the cell subset and a specific gene of the cell subset. The specific gene of the cell subset is the same gene expressed in each cell of the cell subset.

[0023] To solve the above technical problems, the present invention also provides a method for predicting or assisting in predicting the immunotherapy response of a subject, and the method may include the following steps:

[0024] B1) Data reception: Receive the expression level of the gene under the cell subset in the single-cell sequencing data of the subject to be tested;

[0025] B2) Data processing: Substitute the expression level of the gene under the cell subset described in B1) into the immunotherapy response prediction model to calculate and obtain the prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to the method including the following steps: Receive the expression level M of the gene under the cell subset in the single-cell sequencing data of the subject with a response to immunotherapy and the expression level N of the gene under the cell subset in the single-cell sequencing data of the subject without a response to immunotherapy; Use the expression level of the gene under the cell subset with differences between the expression level M and the expression level N as the input, and the immunotherapy response result as the output, and use a machine learning algorithm to construct an immunotherapy response model;

[0026] B3) Data output: Output the prediction result of the immunotherapy response of the subject to be tested.

[0027] In the above method, the machine learning algorithm can be a random forest algorithm. The subject can be a cancer patient, or the subject can also be a person suffering from one or more diseases. The immunotherapy response result can be a response or no response to immunotherapy.

[0028] To solve the above technical problems, the present invention also provides a computer-readable storage medium storing a computer program, and the computer program can cause a computer to execute the following steps:

[0029] C1) Data reception: Receive the expression level of the gene under the cell subset in the single-cell sequencing data of the subject to be tested;

[0030] C2) Data processing: Substitute the expression levels of genes under the cell subsets described in C1) into the immunotherapy response prediction model to calculate the predicted result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to the method including the following steps: Receive the expression levels M of genes under cell subsets in the single-cell sequencing data of subjects with a response to immunotherapy and the expression levels N of genes under cell subsets in the single-cell sequencing data of subjects without a response to immunotherapy; Use the expression levels of genes under cell subsets with differences between the expression level M and the expression level N as inputs, and the immunotherapy response result as the output, and use a machine learning algorithm to construct an immunotherapy response model;

[0031] C3) Data output: Output the predicted result of the immunotherapy response of the subject to be tested.

[0032] In the above computer-readable storage medium, the machine learning algorithm can be a random forest algorithm.

[0033] The subject described above can be a cancer patient, and the subject can also be a person suffering from one or more diseases. The immunotherapy response result can be a response or no response to immunotherapy.

[0034] The expression level of the gene under the cell subset described above can be the expression level of a cell subset and a specific gene of the cell subset. The specific gene of the cell subset is the same gene expressed in each cell of the cell subset.

[0035] The above method may not include the step of obtaining a biological sample from an animal body. The methods can all be information processing methods that do not take a living human body or animal body as the object, but only take data as the object. The methods can all be information processing methods in which all steps are implemented by a data processing device such as a computer.

[0036] The above application or method is a non-disease diagnosis application or method. The above application or method does not directly aim to obtain the disease diagnosis result or health condition of a living human body or animal body.

[0037] The above application or method is an application or method for non-disease treatment purposes. The above application or method does not aim to restore or obtain health or reduce pain for a living human body or animal body.

[0038] The present invention establishes a model for predicting the immunotherapy responsiveness of patients and explores the key factors affecting immunotherapy response. First, the present invention uses the random forest algorithm, linear support vector machine, and deep neural network to establish a prediction model, taking the gene expression levels of each patient's genes in all cells or the gene expression levels under each cell subset as features, establishing a treatment responsiveness model for patients, and comparing their performances. The present invention uses interpretability methods such as Shapley Value and LIME, or utilizes the characteristics of the model, calculates the importance of features, and cyclically screens features according to the importance to reduce the feature dimension for training.

[0039] Advantages of the present invention:

[0040] The present invention establishes a model for predicting the immunotherapy responsiveness of patients and analyzes the importance of input features. Different from traditional differential gene analysis, the importance obtained by using interpretable machine learning methods does not ignore the influence brought by the interaction between features.

[0041] By establishing a prediction model for immunotherapy responsiveness specific to patients, the present invention will promote personalized medicine for immunotherapy. Different patients have significant differences in their responses to immunotherapy. By analyzing the cell subset-gene expression level features, a more precise treatment plan can be formulated for patients, improving the treatment effect, reducing unnecessary side effects, improving the survival rate and quality of life of patients, and also reducing medical costs, bringing long-term economic benefits to patients and the medical system.

[0042] The present invention develops a prediction model based on machine learning algorithms, taking features such as the gene expression levels under cell subsets as input and whether the patient responds to immunotherapy as output. Using the gene expression levels under cell subsets as input is finer-grained than directly using gene expression levels and is more helpful for analyzing the mechanism of action of specific genes in specific cell subsets. Brief Description of the Drawings

[0043] Figure 1 is a schematic diagram of the process of the present invention.

[0044] Figure 2 is a prediction performance table of different models when taking gene expression levels as input. Among them, RF represents random forest, SVM represents support vector machine, and DL represents deep learning model; both ACC and AUC are the averages of the performances of the models after training on randomly divided datasets (50 random divisions).

[0045] Figure 3 is the process of cyclically screening input features of the present invention.

[0046] Figure 4 is the importance ranking (top 20) of cell subset-genes.

[0047] Figure 5 It is the ROC curve of a model trained using the gene expression data of cell subsets.

[0048] Figure 6 It is the Shapley-Value of cell subset-gene. The positive direction of the horizontal axis represents having a response, red indicates high-level expression, and blue indicates low-level expression.

[0049] Figure 7 It is the SBM correlation graph of the feature importance ranking for different inputs. Raw represents the cell subset-gene expression level data, Ratio represents the cell subset ratio data, ‘+’ means concatenating the data, ‘*’ (i.e., Raw*Ratio) means multiplying the gene expression amount under the cell subset by the corresponding cell subset proportion, I means calculating using the same method after changing the random seed, and Random means taking the average after comparing with the random sequence multiple times. Raw*Ratio+Ratio means multiplying the gene expression amount under the cell subset by the corresponding cell subset proportion and then concatenating it with the cell subset-gene expression level data.

[0050] Figure 8 It is the comparison of the prediction performance of different models when using the gene expression amount under cell subsets as input. RF represents Random Forest, SVM represents Support Vector Machine, and DL represents Deep Learning Model. ACC represents accuracy, that is, the number of correctly predicted samples / the total number of samples; AUC represents the area under the ROC curve. Both ACC and AUC are the averages of the performances of the models after training on randomly divided datasets (100 random divisions). Specific Embodiments

[0051] The present invention will be further described in detail below in conjunction with specific embodiments. The provided embodiments are only for clarifying the present invention and not for limiting the scope of the present invention. The following provided embodiments can be used as a guide for those of ordinary skill in the art to make further improvements and do not constitute any limitation to the present invention in any way.

[0052] The experimental methods in the following embodiments are all conventional methods unless otherwise specified, and are carried out according to the techniques or conditions described in the literature in this field or according to the product specifications. The materials, reagents, etc. used in the following embodiments can be obtained from commercial channels unless otherwise specified.

[0053] In the following embodiments, Python 3 software was used to process the data. The experimental results are expressed as mean ± standard deviation. The Mann-Whitney test was used, and P < 0.05 (*) indicates significant difference, P < 0.01 (**) indicates extremely significant difference, and P < 0.001 (***) indicates extremely significant difference.

[0054] Some of the algorithms used in the embodiments of the present invention are as follows:

[0055] LIME (Local Interpretable Model-Agnostic Explanations) is a machine learning model interpretation tool that explains the decisions of black-box models by generating locally interpretable models. Machine learning models are often difficult to interpret due to their complex internal structures and parameters. LIME does not attempt to understand the entire model but, at the feature level, explains the contribution of features to the prediction results, which to some extent reflects the importance of features. This method first perturbs the features of a specific sample to generate some samples; then these perturbed samples are used as the input to the model to be explained, and a set of prediction results are obtained; then, techniques such as Lasso regression or tree-based feature importance are used to select a subset of features that are most relevant to the prediction results; then, an interpretable model for predicting the model output, such as linear regression and decision trees, is trained on the perturbed samples of these features; finally, through this interpretable model, the prediction of the original model for this sample is explained.

[0056] The Shapley Value is a concept in cooperative game theory and also plays a role in explaining the predictions of machine learning models. It calculates the average marginal contribution of each feature to the model prediction by traversing all feature combinations, thereby assigning a relative importance score to each feature. The idea behind this method is that features may interact in complex ways, and their contributions to the prediction results should be evaluated jointly rather than individually. Similar to the LIME method, the Shapley Value calculates the contribution of features to the model prediction on a single sample. Therefore, to calculate the feature importance of multiple samples, the feature importance on a single sample needs to be averaged.

[0057] SBM is a method for measuring the similarity between rankings. It selects the top k elements of two rankings, calculates the size of their intersection / k, traverses k and takes the average, and finally obtains the similarity size between the two rankings.

[0058] In the embodiments of the present invention, the data used is 2.42 million single-cell RNA-Seq sequencing data of 550 patients with 13 cancer types under the pan-cancer immune microenvironment (public data collected from the GEO database, and the PMIDs of relevant literatures are: 27124452; 30250229; 30388455; 30388456; 31227543; 31359002; 31588021; 32405063; 32497499; 32949350; 33711272; 33723257; 33861994; 33958794; 34290408; 34653364; 34653365; 34836966; 35108529; 35121991 and 36719749).

[0059] Example 1. Construction and comparative analysis of different immune therapy responsiveness prediction models.

[0060] In the present invention, prediction models for patients' responsiveness to immune therapy are established using the Random Forest algorithm, linear Support Vector Machine (SVM), and Deep Learning (DL) respectively. Specifically, the model takes features as input and whether the patient responds as output. Among them, features refer to the gene expression levels of each gene of the patient in all cells or the gene expression levels under each cell subset.

[0061] After the training of the Random Forest algorithm is completed, a set of feature importance values can be directly generated; for the linear Support Vector Machine, the present invention uses the absolute value of its weight as the feature importance value; for the Deep Learning, the present invention uses the Shapley Value and the LIME method to calculate its feature importance.

[0062] To more intuitively show the impact of feature changes on responsiveness prediction, the present invention also uses the Shapley Value for the prediction models based on the Random Forest algorithm and the linear Support Vector Machine to display their feature importance values. The Shapley Value and the LIME method reflect the importance of features by measuring the impact of a certain set of feature changes of a sample on the model output. The directly obtained feature importance values are at the individual level and have positive and negative values. To obtain the importance values at the dataset level, the present invention sums the absolute values of the importance of all individual units to obtain the importance ranking of features.

[0063] 1. Establish an immune therapy responsiveness prediction model with gene expression levels as input.

[0064] To explore whether the expression levels of each gene are related to the patients' responsiveness to immune therapy, the present invention uses the average value of the expression levels of each gene in all cells in the single-cell sequencing data as input for model training.

[0065] The present invention uses random forest, support vector machine and deep learning methods to train the immunotherapy response model respectively. The performance of each model is as Figure 2 shown. Among them, the performance of the support vector machine is the best, while the performance of deep learning is significantly lower than that of the other two methods. This is because the dimensionality of the dataset features is much higher than the sample size. Since deep learning essentially uses matrix multiplication to transfer information between layers, it is very likely to overfit on a few features. Models based on decision trees or support vector machines are more suitable for this scenario, so the performance is better.

[0066] Since the performance of the support vector machine algorithm is better ( Figure 2 ), the present invention uses the support vector machine algorithm to repeat the training, represents the importance of genes with its weights, and after 50 rounds of training on different divided datasets, calculates the mean value of the importance rankings, and uses this ranking as the feature importance ranking.

[0067] Taking the gene expression level as the input, among the obtained feature importance values, the GNLY gene has the highest average ranking. However, due to the cell subset specificity of gene expression, directly taking the average value of the expression levels of each gene in all cells as the input loses this part of the difference, and it is difficult to analyze the relationship between specific genes and immunotherapy responsiveness (it is helpful for prediction but not very helpful for the study of the specific mechanism and may not have biological significance). Therefore, studying the gene expression level differences at the cell subset level is very helpful for more detailed analysis.

[0068] 2. Establish an immunotherapy responsiveness prediction model with the expression levels of genes under cell subsets as the input.

[0069] 2.1 Model input feature screening.

[0070] Since the number of combinations of genes under cell subsets (cell subset - gene, that is, a cell subset and a specific gene of this cell subset) in the sequencing data is as high as 3.8×10 6 pieces, while the number of samples is much smaller than this magnitude. Directly applying the original data for model training not only greatly reduces the algorithm efficiency, but also may cause the model to overfit on a few features. Therefore, the present invention first uses the Mann-Whitney test to screen out 1.3×10 5 genes under cell subsets with significantly different expression levels in patients with immunotherapy response and non-response to obtain differentially expressed cell subset - gene combinations, and then uses the following method to cycle through the screening of the features finally used by the model. In addition, since Shapley Value and LIME are based on feature perturbation, the high dimensionality greatly increases their computational complexity. To avoid these negative impacts, the present invention needs to pre-screen the features to reduce the dimensionality of the features used for training.

[0071] First, the present invention uses the Mann-Whitney test to screen cell subset-gene combinations with significantly different expression levels in patients with response to immunotherapy and patients without response to immunotherapy, and takes each cell subset-gene combination and the expression level of the genes in this combination (the expression level of the genes under the cell subset) as the pre-input features of the model, and obtains the initial average importance value of the pre-input features through random initialization).

[0072] After that, the present invention cyclically samples some of the pre-input features as inputs to train a random forest model. At the same time, according to the obtained model, the average importance value of the corresponding features is updated. Finally, the features used for model training are sorted according to the importance value ( Figure 3 ).

[0073] Feature screening is divided into multiple stages, and the number of sampled features decreases according to the stage. Each stage is cycled 100 times. In each cycle, the features are sampled according to the probability distribution shown in the following formula (1):

[0074]

[0075] In formula (1), w i is the average importance value of the i-th feature, and all w i are initially set to 1. After training the model, the average importance of each feature is updated using exponential moving average according to the following formula (2):

[0076]

[0077] In formula (2), X i represents an indicator variable for whether the i-th feature is selected, v i is the importance value of the i-th feature in this cycle of the model, and α is the decay factor.

[0078] After multiple rounds of training, the top 200 features (important features) ranked by importance value are selected as the inputs of the final model.

[0079] After cyclic screening, the present invention obtains a set of importance rankings ( Figure 4 ), and in subsequent experiments, the present invention uses the top 200 features among them to train the model.

[0080] 2.2 Only use the expression level data of the genes under the cell subset.

[0081] The ROC (Receiver Operating Characteristic) curve is a tool for evaluating the performance of classification models. It plots the True Positive Rate as the vertical axis and the False Positive Rate as the horizontal axis, reflecting the performance of the model at different thresholds. The model trained using cell subpopulation-gene expression level data has a fuller ROC curve, indicating that its performance is good ( Figure 5 ).

[0082] Figure 6 It shows that if the expression of PRSS23, GNLY, and ZEB2 genes in the CD8+Temra cell subset is high, the model tends to predict no response; if the expression of CXCR4 gene in the ISG15+ macrophage subset is high, the model tends to predict response. In the importance ranking of all cell subsets-gene expression levels (quantities), the gene importance ranking of some subsets is generally high, such as the CD8+Temra cell subset, and perhaps the proportion of the subsets has an impact on this.

[0083] 2.3 The proportion of cell subsets is not the determining factor for the top ranking of genes in some subsets.

[0084] In the experiment, the present invention found that the gene importance rankings in some cell subpopulations were generally high, such as TE07-CD8-Temra-CXCR1, M13-Mph-ISG15, etc. In order to study whether the high ranking of genes in these cell subpopulations is caused by the cell subpopulation ratio, the present invention multiplied and spliced ​​the cell subpopulation-gene expression data and the cell subpopulation ratio data, and then performed a circular screening experiment again to obtain different importance rankings. The SBM correlation is shown in Figure 7 .

[0085] In fact, the cell subpopulation-gene expression amount after multiplying the cell subpopulation-gene expression level data with the corresponding cell subpopulation proportion is equivalent to the sum of the gene expression amounts in a specific cell subpopulation in an individual divided by the number of cells in the cell population, which actually reflects the gene expression amount at the cell subpopulation level. To some extent, it introduces the influence of the cell subpopulation proportion, because the number of gene expression counts in the cell subpopulation with a higher proportion will also be more. Therefore, the cell subpopulations with a large difference in proportion between immunotherapy responders and non-responders may have a larger difference in gene expression levels, thereby making the model prediction more focused on these characteristics.

[0086] Due to the calculation method of SBM, the means of a large number of random sequences and the SBM of a certain sorted sequence converge to 0.5. The SBM values of the two sets of importance rankings and the Raw data obtained using Raw*Ratio (multiplying the gene expression level under a cell subset by the corresponding cell subset proportion) and Raw*Ratio+Ratio data (multiplying the gene expression level under a cell subset by the corresponding cell subset proportion and then concatenating with the gene expression level data under the cell subset) are all close to 0.5, indicating a large difference; while the SBM values between the sequences obtained from the two sets of Raw and Raw+Ratio data, and the SBM values between the sequences obtained from the two sets of Raw*Ratio and Raw*Ratio+Ratio data are all close to the SBM values of the sequences obtained after changing the random seed and the original sequence, indicating that their sequences are relatively close, to some extent reflecting that the top ranking of the importance of cell subset - gene expression level is not caused by the cell subset proportion.

[0087] 2.4 Compare the performance and importance rankings of different methods.

[0088] In addition to directly outputting feature importance by the random forest model and using Shapley Value to explain the feature importance input in the random forest model, the present invention also attempts to represent feature importance by the weights of a linear support vector machine, represent feature importance by combining a support vector machine with Shapley Value or LIME method, represent feature importance by combining a deep learning model with LIME method, etc. The prediction performance of the model is shown in Figure 8 , where the random forest algorithm ( Figure 8 RF represents in

[0089] The above has detailed the present invention. For those skilled in the art, without departing from the purpose and scope of the present invention and without the need for unnecessary experiments, the present invention can be implemented within a relatively wide range under equivalent parameters, concentrations, and conditions. Although the present invention gives specific embodiments, it should be understood that the present invention can be further improved. In short, according to the principle of the present invention, this application intends to include any changes, uses, or improvements to the present invention, including changes made using conventional techniques known in the art that are outside the scope disclosed in this application.

Claims

1. A data processing device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement a method for predicting immunotherapy response, the method comprising the following steps: S1) Data reception: Receiving the expression levels of genes under cell subsets in the single-cell sequencing data of the subject to be tested; S2) Data processing: Substituting the expression levels of genes under the cell subsets in S1) into an immunotherapy response prediction model to calculate and obtain the prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to a method comprising the following steps: Receiving the expression levels M of genes under cell subsets in the single-cell sequencing data of subjects with a response to immunotherapy and the expression levels N of genes under cell subsets in the single-cell sequencing data of subjects without a response to immunotherapy; Using the expression levels of genes under the cell subsets that are different between the expression levels M and the expression levels N as inputs and the immunotherapy response result as an output, and constructing an immunotherapy response model using a machine learning algorithm; S3) Data output: Outputting the prediction result of the immunotherapy response of the subject to be tested.

2. The data processing device according to claim 1, wherein: The machine learning algorithm is a random forest algorithm.

3. The data processing device according to claim 1 or 2, characterized in that: The subject is a cancer patient.

4. A method for constructing a model for predicting immunotherapy response, characterized in that: The method includes using the expression levels of genes under cell subsets that are different between the single-cell sequencing data of subjects with a response to immunotherapy and the single-cell sequencing data of subjects without a response to immunotherapy as inputs and the immunotherapy response result as an output, and constructing an immunotherapy response model using a machine learning algorithm.

5. A device for predicting or assisting in predicting a subject's immunotherapy response, characterized in that: The device includes the following modules: A1) Data reception module: Used to receive the expression levels of genes under cell subsets in the single-cell sequencing data of the subject to be tested; A2) Data processing module: Used to substitute the expression levels of genes under the cell subsets in A1) into an immunotherapy response prediction model to calculate and obtain the prediction result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to a method comprising the following steps: Receiving the expression levels M of genes under cell subsets in the single-cell sequencing data of subjects with a response to immunotherapy and the expression levels N of genes under cell subsets in the single-cell sequencing data of subjects without a response to immunotherapy; Using the expression levels of genes under the cell subsets that are different between the expression levels M and the expression levels N as inputs and the immunotherapy response result as an output, and constructing an immunotherapy response model using a machine learning algorithm; A3) Data output module: Used to output the prediction result of the immunotherapy response of the subject to be tested.

6. The device according to claim 5, wherein: The machine learning algorithm is a random forest algorithm.

7. A method for predicting or assisting in predicting a subject's response to immunotherapy, characterized in that: The method includes the following steps: B1) Data reception: Receiving the expression levels of genes under cell subsets in the single-cell sequencing data of the subject to be tested; B2) Data processing: Substitute the gene expression levels in the cell subsets described in B1) into the immunotherapy response prediction model to calculate the predicted result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to the method including the following steps: Receive the gene expression levels M in the cell subsets in the single-cell sequencing data of subjects with a response to immunotherapy and the gene expression levels N in the cell subsets in the single-cell sequencing data of subjects without a response to immunotherapy; Use the gene expression levels in the cell subsets with differences between the expression levels M and the expression levels N as inputs, and the immunotherapy response result as the output, and use a machine learning algorithm to construct an immunotherapy response model; B3) Data output: Output the predicted result of the immunotherapy response of the subject to be tested.

8. The method according to claim 7, wherein: The machine learning algorithm is the random forest algorithm.

9. A computer-readable storage medium storing a computer program, the computer program causing the computer to perform the following steps: C1) Data reception: Receive the gene expression levels in the cell subsets in the single-cell sequencing data of the subject to be tested; C2) Data processing: Substitute the gene expression levels in the cell subsets described in C1) into the immunotherapy response prediction model to calculate the predicted result of the immunotherapy response of the subject to be tested; the immunotherapy response prediction model is constructed according to the method including the following steps: Receive the gene expression levels M in the cell subsets in the single-cell sequencing data of subjects with a response to immunotherapy and the gene expression levels N in the cell subsets in the single-cell sequencing data of subjects without a response to immunotherapy; Use the gene expression levels in the cell subsets with differences between the expression levels M and the expression levels N as inputs, and the immunotherapy response result as the output, and use a machine learning algorithm to construct an immunotherapy response model; C3) Data output: Output the predicted result of the immunotherapy response of the subject to be tested.

10. The computer-readable storage medium according to claim 9, wherein: The machine learning algorithm is the random forest algorithm.

Citation Information

Patent Citations

  • Transcriptome-based PD-1 therapy treatment effect prediction system

    CN111755073A

  • Non-small cell lung cancer PD-1 immunotherapy response prediction method aiming at non-disease diagnosis or treatment

    CN116334225A

  • Treatment response prediction method and system, storage medium and terminal

    CN119541847A

  • Biomarkers for Therapy Response After Immunotherapy

    US20240336971A1