Method for acquiring biomarker for predicting effect of immune checkpoint inhibitor, biomarker, determination device, determination method, learning model, and method for generating learning model
By training a model with NLRC5 and additional genes, the method enhances the accuracy of predicting immune checkpoint inhibitor efficacy, addressing the limitations of existing biomarkers.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HOKKAIDO UNIVERSITY
- Filing Date
- 2025-11-21
- Publication Date
- 2026-05-28
AI Technical Summary
Existing biomarkers for predicting the efficacy of immune checkpoint inhibitors, such as CTLA4, PD-1, PD-L1, and PD-L2, have insufficient accuracy in determining therapeutic effects, necessitating the development of more reliable predictive factors.
A method involving training a model with the expression levels of the NLRC5 gene and additional genes to output an evaluation value for the efficacy of immune checkpoint inhibitors, using gene expression analysis from patient groups, and combining these genes to form a biomarker for accurate prediction.
The method enables high-accuracy prediction of the efficacy of immune checkpoint inhibitors, utilizing a biomarker and learning model that accurately determines the effectiveness of treatment.
Smart Images

Figure JP2025040732_28052026_PF_FP_ABST
Abstract
Description
Method for obtaining biomarker for predicting effect of immune checkpoint inhibitor, biomarker, determination device, determination method, learning model, and method for generating learning model
[0001] The present invention relates to a method for obtaining a biomarker for predicting the effect of an immune checkpoint inhibitor, a biomarker, a determination device, a determination method, a learning model, and a method for generating a learning model.
[0002] Immune checkpoint inhibitors are revolutionary cancer therapeutics that eliminate cancer cells by activating the human immune system. On the other hand, immune checkpoint inhibitors have issues such as being very expensive, having a risk of severe side effects, and having a limited number of patients who show a therapeutic effect. Therefore, it is preferable to predict the effect of an immune checkpoint inhibitor before treatment and perform treatment with an immune checkpoint inhibitor only on cancer patients who may have a therapeutic effect. From this perspective, the development of predictive factors (biomarkers) that can serve as indicators for predicting the treatment of immune checkpoint inhibitors is underway.
[0003] For example, as conventional biomarkers, CTLA4, PD-1, PD-L1, and PD-L2, which are targets of immune checkpoint inhibitors, have been studied. However, the reliability of treatment prediction based on these biomarkers is not as high as initially expected, and it has been found that it is difficult to predict the therapeutic effect by simply measuring them.
[0004] Based on such a situation, Patent Document 1 focuses on the NLRRC5 gene (ENSG00000140853) as a new biomarker. Patent Document 1 also shows that conventional biomarkers such as CTLA4, PD-1, PD-L1, and PD-L2 can be used in combination with the NLRRC5 gene.
[0005] U.S. Patent Publication No. 2017 / 0321285
[0006] As described above, although it has been shown as the prior art to use the NLRRC5 gene in combination with conventional biomarkers, the accuracy of predicting the therapeutic effect is not yet sufficient.
[0007] The object of the present invention is to provide a method for obtaining a biomarker for predicting the efficacy of immune checkpoint inhibitors with sufficiently high accuracy in predicting therapeutic effects, a biomarker, a determination device, a determination method, a learning model, and a method for generating a learning model.
[0008] A method for obtaining a biomarker for predicting the effect of an immune checkpoint inhibitor according to the first aspect of the present invention involves training a model that takes the expression levels of two or more genes, including the NLRC5 gene, as input and outputs an evaluation value regarding the effect of an immune checkpoint inhibitor, based on the effect of administering an immune checkpoint inhibitor to a patient group and the expression levels of the NLRC5 gene and several genes different from the NLRC5 gene, as shown by gene expression analysis of the patient group, and obtaining a biomarker consisting of two or more genes, including the NLRC5 gene, to be used in the trained model.
[0009] A method for obtaining a biomarker for predicting the effect of an immune checkpoint inhibitor according to a second aspect of the present invention comprises a first step and a second step. In the first step, one or more first genes with a high correlation of expression with the NLRC5 gene are obtained based on gene expression analysis of a patient group. In the second step, a model is trained that takes the expression levels of two or more genes, including any one of the one or more first genes, as input and outputs an evaluation value regarding the effect of the immune checkpoint inhibitor, and a biomarker consisting of two or more genes, including any one of the one or more first genes and any of the plurality of second genes, is obtained to be used in the trained model. In training the model in the second step, the effect of administering an immune checkpoint inhibitor to the patient group, the expression levels shown in the gene expression analysis for each of the one or more first genes, and the expression levels shown in the gene expression analysis for each of the plurality of second genes, which are different from any of the one or more first genes and also different from the NLRC5 gene, are used.
[0010] A third aspect of the present invention relates to a method for obtaining biomarkers for predicting the effectiveness of immune checkpoint inhibitors. Based on the effects of administering immune checkpoint inhibitors to a patient group and the expression levels of a plurality of first genes consisting of two or more genes as shown in Table 3 below, and a plurality of second genes different from any of the plurality of first genes, as shown by gene expression analysis of the patient group, a model is trained that takes the expression levels of two or more genes including any of the plurality of first genes as input and outputs an evaluation value regarding the effectiveness of the immune checkpoint inhibitor, and a biomarker consisting of two or more genes including any of the plurality of first genes and any of the plurality of second genes is obtained for use in the trained model.
[0011] A biomarker for predicting the efficacy of an immune checkpoint inhibitor according to a fourth aspect of the present invention includes the NLRC5 gene and a gene indicated by any of the IDs in Table 2 described below.
[0012] A biomarker for predicting the efficacy of an immune checkpoint inhibitor according to a fifth aspect of the present invention includes the NLRC5 gene and two genes A and B, indicated by two IDs in any of #1 to #250 in Table 1 described below.
[0013] The determination device for predicting the effect of an immune checkpoint inhibitor according to the sixth aspect of the present invention calculates an evaluation value regarding the effect of the immune checkpoint inhibitor based on the expression level of each gene included in the biomarker.
[0014] A method for predicting the effectiveness of an immune checkpoint inhibitor according to a seventh aspect of the present invention determines the effectiveness of the immune checkpoint inhibitor based on the expression levels of each gene included in the biomarker.
[0015] The learning model for predicting the effectiveness of immune checkpoint inhibitors according to the eighth aspect of the present invention accepts input of the expression levels of each gene included in the biomarker and causes the computer to function to calculate and output an evaluation value regarding the effectiveness of the immune checkpoint inhibitor.
[0016] A method for generating a learning model for predicting the effects of immune checkpoint inhibitors according to the ninth aspect of the present invention uses the effects of administering immune checkpoint inhibitors to a group of patients and the expression levels of each gene shown by gene expression analysis of that group of patients as training data.
[0017] According to the method for obtaining biomarkers for predicting the effects of immune checkpoint inhibitors according to the first to third aspects of the present invention, it is possible to obtain biomarkers that can predict the effects with high accuracy. Furthermore, according to the biomarkers for predicting the effects of immune checkpoint inhibitors according to the fourth and fifth aspects of the present invention, it is possible to predict the effects with high accuracy. Furthermore, according to the biomarkers for predicting the effects of immune checkpoint inhibitors according to the fourth and fifth aspects of the present invention, it is possible to predict the effects with high accuracy. According to the determination device, determination method, and learning model for predicting the effects of immune checkpoint inhibitors according to the sixth to eighth aspects of the present invention, it is possible to predict the effects with high accuracy by using the biomarker according to the fourth or fifth aspect. According to the method for generating a learning model according to the ninth aspect, it is possible to appropriately generate a learning model according to the eighth aspect.
[0018] This is a block diagram showing the schematic configuration of a determination device according to a first embodiment, which is one embodiment of the present invention. This is a block diagram showing the functional configuration of the determination device in Figure 1. This is a flowchart showing the process of predicting the effect of an immune checkpoint inhibitor using the determination device in Figure 1. This is a flowchart showing the method for obtaining a biomarker used in the determination device in Figure 1. This is a graph showing the ROC curve related to predicting the effect of an immune checkpoint inhibitor using a biomarker according to one embodiment of the present invention. This is a block diagram showing the schematic configuration of a determination device according to a second embodiment, which is another embodiment of the present invention. This is a block diagram showing the schematic configuration of a determination device according to a third embodiment, which is another embodiment of the present invention.
[0019] <First Embodiment> [Configuration of the Determination Device] A determination device 1 for predicting the effect of an immune checkpoint inhibitor (hereinafter referred to as "determination device 1") according to the first embodiment, which is one embodiment of the present invention, will be described with reference to the drawings.
[0020] The determination device 1 is a device that predicts the effect of an immune checkpoint inhibitor. This prediction uses a biomarker consisting of three genes. The biomarkers in this embodiment are, for example, the NLRC5 gene (ENSG00000140853), ENSG00000162576, and ENSG00000100342 (see Table 1 #1 below). In this embodiment, genes are represented using ensembleID. ensembleID is a number that identifies a gene and is used in "Ensembl". "Ensembl" is a gene database project being advanced at EBI (European Molecular Biology Laboratory, European Bioinformatics Institute) (see URL https: / / asia.ensembl.org / index.html).
[0021] The determination device 1 makes predictions based on the expression levels of each gene included in the biomarker. The expression levels of each gene are obtained from data (hereinafter referred to as "analysis data") showing the results of RNA sequencing analysis (corresponding to the "gene expression analysis" of the present invention) performed on the biological sample of the target patient using a next-generation sequencer. The details of the determination device 1 will be described below.
[0022] As shown in Figure 1, the determination device 1 comprises hardware such as a calculation unit 10 including a CPU 11 and a memory unit 12, a data input unit 13, and a display 14, as well as software consisting of program data stored in the memory unit 12. This software can be distributed via download from the Internet or by various recording media such as USB (Universal Serial Bus) memory. The memory unit 12 consists of memory devices such as ROM and RAM, and storage devices. The data input unit 13 consists of a device that accepts input of analysis data. Examples of such devices include a USB port for connecting to a USB memory and receiving analysis data recorded in the USB memory, and a network port for connecting to the Internet and acquiring analysis data via the Internet. The display 14 has a screen that displays characters, images, etc., and outputs various information to the user by displaying these on the screen. The functions of the determination device 1 may be realized by a computer system equipped with multiple computers. For example, the functions of the calculation unit 10 and the other functions may be realized by different computers. In this case, the computer responsible for the functions of the arithmetic unit 10 and the computers responsible for the other functions are connected via a communication network such as the Internet, and the processing by the functions of the arithmetic unit 10 and other computers is executed as a whole computer system while these computers communicate the necessary data with each other. As a result, the same functions as the determination device 1 are realized as a whole computer system.
[0023] Program data stored in the storage device of the memory unit 12 is transferred to the RAM, and the CPU executes a series of processes indicated by the program data transferred to the RAM. Through this cooperation between hardware and software, the various functions of the determination device 1 described below are realized.
[0024] As shown in Figure 2, the calculation unit 10 includes a preprocessing unit 21 and an evaluation value calculation unit 22, which are realized through the above-mentioned cooperation between hardware and software. The preprocessing unit 21 performs preprocessing on the analysis data (Fastq file). The preprocessing includes counting the expression level of each RNA from the base sequence and quality score shown in the analysis data, integrating new patient data with existing data, dealing with potential problems such as duplicate genes and missing values, removing batch effects, normalization and log2 conversion, and conversion to Z-score. The data after preprocessing by the preprocessing unit 21 shows the normalized expression level of each gene as a Z-score.
[0025] The evaluation value calculation unit 22 extracts data indicating the expression levels of each gene included in the biomarker from the data after preprocessing by the preprocessing unit 21. Then, the evaluation value calculation unit 22 performs calculations based on a model that has been trained to calculate an evaluation value indicating the predicted effect of the immune checkpoint inhibitor from the expression levels of each gene included in the biomarker. The evaluation value has a magnitude from 0 to 1, for example, and the larger the value, the higher the probability that the immune checkpoint inhibitor will be effective.
[0026] The calculation unit 10 displays the evaluation value calculated by the evaluation value calculation unit 22 on the display 14. At this time, in addition to displaying a number indicating the evaluation value, the method of judgment based on the evaluation value and whether or not the evaluation value indicates an effect may also be displayed on the screen of the display 14 as reference information. Whether or not the evaluation value indicates an effect may be, for example, determined to be effective when the evaluation value exceeds a predetermined threshold, and ineffective when the evaluation value is below the predetermined threshold. The predetermined threshold may be, for example, 0.5. Instead of, or in addition to, displaying these on the display 14, these may be output by recording data to a USB memory, printing on paper, transmitting to other devices via the internet, etc.
[0027] The model in this embodiment uses a random forest. A random forest is one of the machine learning algorithms that uses decision trees. Other algorithms such as deep neural networks and deep forests may also be used.
[0028] As described above, the biomarker according to this embodiment includes the NLRC5 gene, as well as other genes, ENSG00000162576 and ENSG00000100342. The other genes are genes acquired during the learning of the evaluation value calculation unit 22 and are optimized to enable highly accurate prediction when used in combination with the NLRC5 gene. ENSG00000162576 and ENSG00000100342 correspond to the combination of genes that simultaneously achieve the highest accuracy and the highest test result value in Table 1 described later. In place of this combination of genes, a combination in Table 1 that has an accuracy of 0.71 or higher may be used. Among these, a combination in which one or both of the accuracy and / or test result value is 0.75 or higher is preferred. Among these, a combination in which one or both of the accuracy and / or test result value is 0.78 or higher is more preferred. Among these, a combination in which one or both of the accuracy and / or test result value is 0.79 or higher is even more preferred. Furthermore, among these, combinations in which either the accuracy or the test result value, or both, are 0.80 or higher are more preferable.
[0029] [Prediction Method] A method for predicting the effect of an immune checkpoint inhibitor using the judgment device 1 will be explained with reference to Figure 3. First, a biological sample from the target patient is obtained (S1). The biological sample is obtained, for example, during a medical examination or surgery. Next, RNA sequencing analysis is performed on the biological sample obtained in S1 using a next-generation sequencer (S2). Next, the analysis data output from the next-generation sequencer in S2 is input to the judgment device 1 through the data input unit 13 (S3). Next, the pre-processing unit 21 of the judgment device 1 performs pre-processing on the analysis data input in S3 (S4). Next, the evaluation value calculation unit 22 of the judgment device 1 calculates an evaluation value according to a trained model using data indicating the expression levels of each biomarker gene included in the pre-processed data obtained in S4 (S5). The evaluation value calculated in S5 is displayed on the display 14 (S6).
[0030] [Method for Obtaining Biomarkers] The method for obtaining biomarkers used in the evaluation value calculation unit 22 of the determination device 1 (corresponding to the "Method for Obtaining Biomarkers for Predicting the Efficacy of Immunotherapy Checkpoint Inhibitors" of the present invention) will be explained with reference to Figure 4. This method uses a computer equipped with hardware such as a CPU and memory devices. A supercomputer equipped with many CPUs may be used as this computer. In this case, the functions of the supercomputer can be used from a terminal such as a desktop PC connected to the supercomputer via a communication network such as the Internet. Alternatively, many computers may be connected via a communication network, and the calculations may be distributed and executed by these many computers. Software such as program data for biomarker acquisition is stored in the memory devices, etc. Through the cooperation of this hardware and software, the computer executes the processing described later. The software can be distributed by downloading it via the Internet or by various recording media such as USB (Universal Serial Bus) memory.
[0031] First, training data to be used to train the model is prepared (S11). The training data is obtained from clinical trials conducted on each of several patient groups. Each of the patient groups consists of multiple cancer patients, with no overlap between groups. RNA sequencing analysis is performed on biological samples obtained from each patient in each group. In addition, each patient is administered an immune checkpoint inhibitor, and it is determined whether or not the administration was effective. This determination is made by a physician based on predetermined criteria. The training data consists of data showing the gene expression levels (explanatory variables) in each patient in each group, and data showing the effect of the administration of the immune checkpoint inhibitor (dependent variable). In this embodiment, the training data from one of the patient groups is used as training data, and the training data from the other patient groups is used as test data.
[0032] Next, the computer is made to execute processes S12 to S16. As described above, this is done through the cooperation of the computer's hardware and software.
[0033] First, from all the genes related to the training data prepared in S11, combinations consisting of the NLRC5 gene and two other genes that have not been trained are selected (S12). Next, a random forest model is trained based on the training data related to the genes selected in S12 from the training data prepared in S11 (S13). Specifically, the model is trained to calculate evaluation values from the expression levels of the genes selected in S12, and to reflect the effect of administering immune checkpoint inhibitors as shown in the training data. Along with the training, the accuracy of the trained model is calculated. The accuracy indicates how well the trained model reflects the correct answer shown in the training data. Specifically, it indicates how well the predicted effect based on the evaluation value calculated by the model for the training data matches the correct answer shown in the training data regarding the presence or absence of the effect of administering immune checkpoint inhibitors. In this case, the presence or absence of the effect shown by the evaluation value is based on whether the evaluation value exceeds a threshold (e.g., 0.5). For example, the accuracy is given in the range of 0 to 1, and the larger the value, the higher the degree of agreement. Furthermore, a confidence level of 0 indicates a complete mismatch, while a confidence level of 1 indicates a complete match.
[0034] Next, the trained model obtained in S13 is tested based on test data related to the genes selected in S12 from the training data prepared in S11 (S14). Specifically, the test data related to the genes selected in S12 is input into the trained model, and it is evaluated whether the evaluation value output from the trained model reflects the effect of administering immune checkpoint inhibitors as indicated by the test data. This evaluation is performed, for example, by evaluating for each patient whether the presence or absence of the effect of administering immune checkpoint inhibitors as indicated by the evaluation value matches the presence or absence of the effect of administering immune checkpoint inhibitors as indicated by the test data. Then, the total number of matching patients is divided by the total number of patients to calculate a value for each patient group, and this value is taken as the test result value for that patient group.
[0035] Next, it is determined whether all gene combinations have been trained or not. If it is determined that there are still untrained combinations (S15, No), the process from S12 is executed. If it is determined that there are no untrained combinations (S15, Yes), the training accuracy in S13 and the test results in S14 are output as a data file (S16). Next, a suitable combination of genes and a model are determined based on the data file from S16 (S17). For example, a gene combination with the highest level of training accuracy and good agreement with the test results is extracted. Then, a trained model is obtained for that gene combination. Obtaining the model here may be done by repeating the process in S3 for the extracted gene combination, or the trained model may be saved to memory or storage each time S13 is executed in the series of processes in Figure 4, and the model related to the extracted gene combination is retrieved from memory or storage in S17.
[0036] According to the biomarker acquisition method of this embodiment, it is possible to acquire biomarkers that can predict effects with high accuracy. Furthermore, the determination device 1, the model used in the determination device 1, and the determination method using them enable the prediction of effects with high accuracy.
[0037] [First Embodiment] An embodiment of the method for obtaining a biomarker according to the first embodiment will be described below.
[0038] In relation to S1 in Figure 4, the clinical trial results from the following three studies were used.
[0039] 1) Distinct Immune Cell Populations Define Response to Anti-PD-1 Monotherapy and Anti-PD-1 / Anti-CTLA-4 Combined Therapy Gide, Tuba N. et al. Cancer Cell, Volume 35, Issue 2, 238 - 255.e6
[0040] 2) Genomic and Transcriptomic Features of Response to Anti-PD-1 Therapy in Metastatic Melanoma Hugo, Willy et al. Cell, Volume 165, Issue 1, 35 - 44
[0041] 3) Tumor and Microenvironment Evolution during Immunotherapy with Nivolumab Riaz, Nadeem et al. Cell, Volume 171, Issue 4, 934 - 949.e16
[0042] In any of the studies 1) to 3) above (hereinafter referred to as "Study 1" to "Study 3"), a clinical trial was conducted in which an anti-PD-1 inhibitor was administered to a group of patients consisting of multiple patients with skin cancer, and the efficacy / non-efficacy of the administration of the anti-PD-1 inhibitor was determined for each patient. In addition, RNA sequence analysis (transcriptome analysis) was performed on each patient, and the expression level of each gene was obtained. The patient groups are independent among Studies 1 to 3. Based on the results of the clinical trial shown in Study 1 (hereinafter referred to as "Trial 1"), training data was generated. Also, based on the results of the clinical trials shown in Studies 2 and 3 (hereinafter referred to as "Trial 2" and "Trial 3"), test data was generated.
[0043] Based on the training data and test data obtained as described above, the computer was made to execute the processes of S12 to S16. Table 1 below shows a part of the training accuracy and test result values thus obtained. Genes A and B represent two genes other than the NLRRC5 gene in each combination consisting of three genes used for training. Table 1 shows the results extracted for combinations of Genes A and B with a relatively high training accuracy of 0.71 or more. In any combination, the test result value is 0.71 or more.
[0044] [Table 1]
[0045] Also, for each combination of genes included in Table 1, a ROC (Receiver Operating Characteristic) curve was obtained. That is, for each patient group in Tests 1 and 2, while setting the threshold value serving as the determination criterion to various sizes from the evaluation value output from the model based on the expression levels of each combination of genes, the presence or absence of the effect of administering an immune checkpoint inhibitor was determined. Then, the sensitivity and specificity obtained by comparing the determination result with the correct answer were plotted. FIG. 5 is an example thereof, and shows the ROC curve obtained for the combination #1 in Table 1. FIG. 5(a) is a curve related to the training data (Test 1), and FIG. 5(b) is a curve related to the test data (Test 2). The horizontal axis of each graph represents the specificity, and the vertical axis represents the sensitivity. The numerical value in the lower right of the graph is the AUC (Area Under the Curve).
[0046] <Second Embodiment> A determination device 101 for predicting the effect of an immune checkpoint inhibitor according to a second embodiment, which is another embodiment of the present invention (hereinafter referred to as the "determination device 101") will be described with reference to FIG. 6. Since this embodiment has many points in common with the first embodiment, hereinafter, the description of the common points will be omitted as appropriate, and the same reference numerals will be used for the common configurations.
[0047] In this embodiment, the difference from the first embodiment is the combination of genes of the biomarker used for determination. The biomarker according to the second embodiment consists of a total of two genes, namely the NLRRC5 gene and one gene other than the NLRRC5 gene. As the latter gene, any one is selected from Table 2 described later. For example, ENSG00000105176 having the highest accuracy in Table 2 is selected. Instead of this gene, other genes in Table 2 may be used. Among them, genes having an accuracy of 0.74 or more, corresponding to approximately the top 10% in Table 2, are preferable. Among them, combinations having an accuracy of 0.75 or more, corresponding to approximately the top 4% in Table 2, are more preferable.
[0048] In the determination device 101 according to this embodiment, the difference from the determination device 1 according to the first embodiment is that a calculation unit 110 is used instead of a calculation unit 10. Furthermore, in the calculation unit 110, the difference from the calculation unit 10 according to the first embodiment is that a storage unit 112 is used instead of a storage unit 12.
[0049] The memory unit 112 stores program data different from that of the first embodiment. This program data can be distributed via download from the Internet or by various recording media such as USB (Universal Serial Bus) memory. When predicting the effect of an immune checkpoint inhibitor, the calculation unit 210 performs processing according to the program data in the memory unit 112. This program data is a modified version of the program data according to the first embodiment, with content adapted to the case where the biomarker consists of the NLRC5 gene and one other gene. Therefore, the calculation unit 110 performs processing S3 to S6 in Figure 3 with content adapted to the case where the biomarker consists of the NLRC5 gene and one other gene. Furthermore, when acquiring the biomarker, processing S11 to S17 in Figure 4 is performed with content adapted to the combination consisting of the NLRC5 gene and one other gene.
[0050] According to the biomarker acquisition method of this embodiment, it is possible to acquire biomarkers that can predict effects with high accuracy. Furthermore, the determination device 101, the model used in the determination device 101, and the determination method using them enable the prediction of effects with high accuracy.
[0051] [Second Embodiment] Below, an embodiment of the method for obtaining biomarkers according to the second embodiment will be described. As described above, the processing in S11 to S17 of Figure 4 was modified to correspond to combinations consisting of the NLRC5 gene and one gene other than the NLRC5 gene, and the processing was performed. In obtaining the training data, the results of the clinical trial in skin cancer related to Study 1 of the first embodiment were used. As a result, trained models for each gene and the accuracy of each training shown in Table 2 were obtained. Table 2 shows the results extracted for combinations of genes A and B that had a relatively high training accuracy of 0.70 or higher.
[0052] [Table 2]
[0053] <Third Embodiment> A determination device 201 for predicting the effect of an immune checkpoint inhibitor (hereinafter referred to as "determination device 201") according to a third embodiment, which is another embodiment of the present invention, will be described with reference to Figure 7. Since this embodiment is similar in many respects to the second embodiment, the explanation of the common points will be omitted as appropriate below, and the same reference numerals will be used for common components.
[0054] In this embodiment, the difference from the second embodiment lies in the combination of biomarker genes used for determination. The biomarker according to the second embodiment consists of two genes. One of the two genes (corresponding to the "first gene" of the present invention) is a gene different from the NLRC5 gene and is a gene with a high correlation of expression with the NLRC5 gene (hereinafter referred to as the "surrogate gene"). The other of the two genes (corresponding to the "second gene" of the present invention) is optimized in combination with the first gene to enable prediction of the effect with high accuracy.
[0055] In the determination device 201 according to this embodiment, the difference from the determination device 101 according to the second embodiment is that a calculation unit 210 is used instead of the calculation unit 110. Furthermore, in the calculation unit 210, the difference from the calculation unit 110 according to the second embodiment is that a storage unit 212 is used instead of the storage unit 112.
[0056] The memory unit 212 stores program data different from that of the second embodiment. This program data can be distributed via download from the Internet or by various recording media such as USB (Universal Serial Bus) memory. In predicting the effect of an immune checkpoint inhibitor, the calculation unit 210 performs processing according to the program data in the memory unit 212. This program data is a modified version of the program data according to the first embodiment, with the content changed to reflect the case where the biomarker consists of the above-mentioned substitute gene and the other gene. Therefore, the calculation unit 110 performs processing S3 to S6 in Figure 3, with the content changed to reflect the case where the biomarker consists of the above-mentioned substitute gene and the other gene.
[0057] In determining the substitute gene, the Pearson correlation coefficient is calculated for a given patient group between the expression levels of the NLRC5 gene and other genes. Based on a comparison of the calculation result with a predetermined threshold, it is determined whether the correlation is high or low, and the gene with a high correlation is designated as the substitute gene. Furthermore, in obtaining the biomarker, the processes S11 to S17 in Figure 4 are modified according to the combination of the substitute gene and other genes.
[0058] According to the biomarker acquisition method of this embodiment, it is possible to acquire biomarkers that can predict effects with high accuracy. Furthermore, the determination device 101, the model used in the determination device 101, and the determination method using them enable the prediction of effects with high accuracy.
[0059] [Third Embodiment] Below, an embodiment of the method for determining surrogate genes according to the third embodiment will be described. For a specific group of patients suffering from skin cancer related to Study 1 of the first embodiment, Pearson's correlation coefficient was calculated between the expression level of the NLRC5 gene and the expression levels of other genes. Table 3, which shows a part of the results, is an extraction of genes that had a relatively high correlation coefficient of 0.70 or higher. Note that "NA" in Table 3 means "not available" and corresponds to genes for which a name has not been determined.
[0060] [Table 3]
[0061] As a substitute gene, ENSG00000125347 (IRF1 gene) from Table 3 #2 was extracted, and the processing shown in Figure 4 was performed, modifying the content according to the combination of two genes including the IRF1 gene. As a result, 4670 genes with a training accuracy of 0.7 or higher were obtained as the other genes paired with the IRF1 gene, along with the models associated with each gene.
[0062] Furthermore, as an alternative gene, ENSG00000198851 (CD3E gene) from Table 3 #37 was extracted, and the processing shown in Figure 4 was modified to reflect the combination of two genes including the CD3E gene. As a result, 2572 genes with a training accuracy of 0.7 or higher were obtained as the other genes paired with the CD3E gene, along with models associated with each gene.
[0063] <Modifications> The above describes preferred embodiments of the present invention, but the present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the means for solving the problem.
[0064] For example, in each of the embodiments described above, RNA sequencing analysis is performed as the gene expression analysis. Alternatively, other gene expression analysis methods may be used. For example, microarrays or real-time PCR (Polymerase Chain Reaction) methods may be used. Furthermore, the nCounter® Analysis System (NanoString®) may be used.
[0065] Furthermore, the biomarker acquisition methods according to each embodiment described above (for example, the method shown in Figure 4) are also applicable to checkpoint inhibitors other than anti-PD-1 inhibitors, namely anti-CTLA-4 inhibitors and anti-PD-L1 inhibitors. That is, the method in Figure 4 can be implemented using the results of clinical trials using anti-CTLA-4 inhibitors or anti-PD-L1 inhibitors, and training data obtained from the gene expression levels from RNA sequencing analysis of each patient in the patient group and the effects of administering anti-CTLA-4 inhibitors or anti-PD-L1 inhibitors to each patient. Based on the accuracy of training, test results, AUC, etc. for each combination of genes obtained in this way, one of the gene combinations can be extracted as a biomarker, and by using the trained model related to the extracted combination, it is possible to predict the effects of checkpoint inhibitors other than anti-PD-1 inhibitors, similar to the embodiments described above. Moreover, anti-LAG3 inhibitors and anti-TIM3 inhibitors treat cancer using a similar principle to conventional checkpoint inhibitors such as anti-PD-1 inhibitors, in that they ultimately induce T cell activation. In this sense, the above-described embodiments are also applicable to anti-LAG3 inhibitors and anti-TIM3 inhibitors, and their effects can be predicted.
[0066] Furthermore, in each of the above-described embodiments, the results of clinical trials on a group of patients suffering from skin cancer were used. Moreover, it is possible to predict the effect of checkpoint inhibitors even when the present invention is applied to other cancers. For example, the method shown in Figure 4 can be implemented using the results of clinical trials on a group of patients suffering from other cancers, and training data obtained from the gene expression levels from RNA sequencing analysis of each patient in the patient group and the effect of administering checkpoint inhibitors to each patient. Specifically, the inventors have confirmed that when the method shown in Figure 4 is implemented based on the results of clinical trials on patient groups suffering from renal cancer and lung cancer, respectively, the prediction accuracy comparable to that of the first to third embodiments above can be obtained.
[0067] 1, 101, 201 Judgment device 10, 110, 210 Calculation unit 11 CPU 12, 112, 212 Storage unit 21 Preprocessing unit 22 Evaluation value calculation unit
Claims
1. A method for obtaining a biomarker for predicting the effect of an immune checkpoint inhibitor, characterized by training a model that takes the expression levels of two or more genes, including the NLRC5 gene, as input and outputs an evaluation value regarding the effect of an immune checkpoint inhibitor, based on the effect of administering an immune checkpoint inhibitor to a patient group and the expression levels of the NLRC5 gene and several genes different from the NLRC5 gene, as shown by gene expression analysis of that patient group, and obtaining a biomarker consisting of two or more genes, including the NLRC5 gene, to be used in the trained model.
2. A method for obtaining a biomarker for predicting the effect of an immune checkpoint inhibitor, comprising: a first step of obtaining one or more first genes that have a high correlation in expression with the NLRC5 gene based on gene expression analysis of a patient group; and a second step of training a model that takes the expression levels of two or more genes, including any one of the one or more first genes, as input and outputs an evaluation value regarding the effect of an immune checkpoint inhibitor, and obtaining a biomarker consisting of two or more genes, including any one of the one or more first genes and any of the plurality of second genes, to be used in the trained model, wherein the training of the model in the second step uses the effect of administering an immune checkpoint inhibitor to the patient group, the expression levels shown in the gene expression analysis for each of the one or more first genes, and the expression levels shown in the gene expression analysis for each of the plurality of second genes that are different from any of the one or more first genes and also different from the NLRC5 gene.
3. A method for obtaining a biomarker for predicting the effect of an immune checkpoint inhibitor, characterized by training a model that takes the expression levels of two or more genes, including one or more of the aforementioned first genes, as input and outputs an evaluation value regarding the effect of an immune checkpoint inhibitor, based on the effect of administering an immune checkpoint inhibitor to a patient group and the expression levels of one or more first genes consisting of genes included in List 1 and a plurality of second genes different from any of the aforementioned one or more first genes, as shown by gene expression analysis of the patient group, and obtaining a biomarker consisting of two or more genes, including one or more of the aforementioned first genes and a plurality of the aforementioned second genes, to be used in the trained model. (List 1) ENSG00000105176 ENSG00000125629 ENSG00000168394 ENSG00000138079 ENSG00000073008 ENSG00000163328 ENSG00000179256 ENSG00000135747 ENSG00000163508 ENSG00000135439 ENSG00000089012 ENSG00000162613 ENSG00000172116 ENSG00000235459 ENSG00000188868 ENSG00000136866 ENSG00000074657 ENSG00000260007 ENSG00000278362 ENSG00000143390 ENSG00000099904 ENSG00000079308 ENSG00000115966 ENSG00000257246 ENSG00000243364 ENSG00000134824 ENSG00000171097 ENSG00000111796 ENSG00000211778 ENSG00000125877 ENSG00000008394 ENSG00000223787 ENSG00000163564 ENSG00000274741 ENSG00000108298 ENSG00000240764 ENSG00000126001 ENSG00000104894 ENSG00000053254 ENSG00000158796 ENSG00000196071ENSG00000160051 ENSG00000141401 ENSG00000271503 ENSG00000099139 ENSG00000007545 ENSG00000107731 ENSG00000185245 ENSG00000048405 ENSG00000143412 ENSG00000162775 ENSG00000143256 ENSG00000058272 ENSG00000196557 ENSG00000108506 ENSG00000172292 ENSG00000229689 ENSG00000156920 ENSG00000196504 ENSG00000212994 ENSG00000164484 ENSG00000133639 ENSG00000249141 ENSG00000123066 ENSG00000144395 ENSG00000283761 ENSG00000154217 ENSG00000138674 ENSG00000145743 ENSG00000102935 ENSG00000186310 ENSG00000099942 ENSG00000163032 ENSG00000182372 ENSG00000005379 ENSG00000144229 ENSG00000167508 ENSG00000139131 ENSG00000196503 ENSG00000243649 ENSG00000185418 ENSG00000242485 ENSG00000117153 ENSG00000276293 ENSG00000237651 ENSG00000230037 ENSG00000113749 ENSG00000167895 ENSG00000181016 ENSG00000272391 ENSG00000024422 ENSG00000125895 ENSG00000185739 ENSG00000197142 ENSG00000151575 ENSG00000166526 ENSG00000177374 ENSG00000112739 ENSG00000123595 ENSG00000242852 ENSG00000117713 ENSG00000172366 ENSG00000196793ENSG00000077063 ENSG00000163344 ENSG00000275498 ENSG00000128965 ENSG00000157064 ENSG00000188215 ENSG00000213088 ENSG00000135218 ENSG00000067992 ENSG00000089351 ENSG00000118922 ENSG00000181036 ENSG00000198435 ENSG00000142089 ENSG00000138411 ENSG00000141519 ENSG00000113407 ENSG00000185215 ENSG00000100938 ENSG00000226479 ENSG00000231292 ENSG00000169245 ENSG00000126353 ENSG00000147606 ENSG00000162723 ENSG00000178772 ENSG00000226318 ENSG00000276114 ENSG00000118515 ENSG00000143842 ENSG00000100523 ENSG00000134755 ENSG00000162373 ENSG00000213799 ENSG00000267952 ENSG00000164935 ENSG00000181222 ENSG00000100580 ENSG00000134962 ENSG00000205045 ENSG00000267561 ENSG00000125459 ENSG00000176165 ENSG00000189060 ENSG00000228623 ENSG00000050730 ENSG00000156463 ENSG00000211807 ENSG00000105472 ENSG00000133606 ENSG00000182901 ENSG00000276061 ENSG00000110448 ENSG00000133466 ENSG00000197050 ENSG00000140105 ENSG00000166900 ENSG00000079739 ENSG00000137073 ENSG00000137496 ENSG00000274808 ENSG00000103226ENSG00000135048 ENSG00000188822 ENSG00000196263 ENSG00000257624 ENSG00000274000 ENSG00000101331 ENSG00000114861 ENSG00000255292 ENSG00000164736 ENSG00000070669 ENSG00000155324 ENSG00000138380 ENSG00000160752 ENSG00000237403 ENSG00000187595 ENSG00000244004 ENSG00000251493 ENSG00000069764 ENSG00000141564 ENSG00000148399 ENSG00000170956 ENSG00000167676 ENSG00000196119 ENSG00000164221 ENSG00000174206 ENSG00000198146 ENSG00000221978 ENSG00000101474 ENSG00000250644 ENSG00000288640 ENSG00000179593 ENSG00000285152 ENSG00000165178 ENSG00000116985 ENSG00000169758 ENSG00000163006 ENSG00000100342 ENSG00000055917 ENSG00000103510 ENSG00000138382 ENSG00000100084 ENSG00000183718 ENSG00000184661 ENSG00000092439 ENSG00000128654 ENSG00000197744 ENSG00000275118 ENSG00000113356 ENSG00000277633 ENSG00000049768 ENSG00000122188 ENSG00000123219 ENSG00000284431 ENSG00000033800 ENSG00000119640 ENSG00000135144 ENSG00000172269 ENSG00000154133 ENSG00000166261 ENSG00000169418 ENSG00000170144 ENSG00000230769ENSG00000254087 ENSG00000250366 ENSG00000129682 ENSG00000185960 ENSG00000136541 ENSG00000116161 ENSG00000166206 ENSG00000231841 ENSG00000158006 ENSG00000231925 ENSG00000170989 ENSG00000206505 ENSG00000241258 ENSG00000144214 ENSG00000256771 ENSG00000264655 ENSG00000081014 ENSG00000125968 ENSG00000143324 ENSG00000168785 ENSG00000187601 ENSG00000197857 ENSG00000029639 ENSG00000119703 ENSG00000120549 ENSG00000125999 ENSG00000129173 ENSG00000248383 ENSG00000274575 ENSG00000171827 ENSG00000280789 ENSG00000099783 ENSG00000132541 ENSG00000186479 ENSG00000141568 ENSG00000176435 ENSG00000100450 ENSG00000188064 4. The method for obtaining a biomarker according to any one of items 1 to 3, characterized in that the model is based on a random forest.
5. A biobarker for predicting the efficacy of immune checkpoint inhibitors, characterized by comprising the NLRC5 gene and a gene indicated by any of the IDs in List 2. (List 2) ENSG00000140853 ENSG00000125347 ENSG00000179583 ENSG00000172575 ENSG00000140368 ENSG00000096996 ENSG00000026950 ENSG00000123329 ENSG00000221963 ENSG00000160185 ENSG00000122122 ENSG00000143851 ENSG00000111679 ENSG00000168071 ENSG00000005844 ENSG00000117228 ENSG00000136286 ENSG00000008517 ENSG00000111801 ENSG00000009790 ENSG00000110934 ENSG00000183484 ENSG00000102879 ENSG00000079263 ENSG00000168394 ENSG00000164691 ENSG00000100385 ENSG00000100368 ENSG00000276986 ENSG00000105851 ENSG00000234487 ENSG00000154451 ENSG00000105122 ENSG00000232367 ENSG00000197057 ENSG00000013725 ENSG00000198851 ENSG00000205045 ENSG00000184922 ENSG00000146192 ENSG00000023892 ENSG00000127084 ENSG00000281614 ENSG00000180644 ENSG00000167984 ENSG00000153563 ENSG00000107099 ENSG00000286030 ENSG00000198286 ENSG00000116824 ENSG00000128815 ENSG00000142347 ENSG00000182866 ENSG00000115085 ENSG00000147168 ENSG00000186517 ENSG00000205784ENSG00000197646 ENSG00000138964 ENSG00000134516 ENSG00000081237 ENSG00000204264 ENSG00000110324 ENSG00000163519 ENSG00000180096 ENSG00000141968 ENSG00000121594 ENSG00000050730 ENSG00000134470 ENSG00000132274 ENSG00000100055 ENSG00000276849 ENSG00000229474 ENSG00000167208 ENSG00000164674 ENSG00000136250 ENSG00000120217 ENSG00000082074 ENSG00000123338 ENSG00000162739 ENSG00000125637 ENSG00000288169 ENSG00000175463 ENSG00000153283 ENSG00000159618 ENSG00000183918 ENSG00000057657 ENSG00000185811 ENSG00000140968 ENSG00000025708 ENSG00000134242 ENSG00000159753 ENSG00000254838 ENSG00000158517 ENSG00000167286 ENSG00000179144 ENSG00000015285 ENSG00000110448 ENSG00000205744 ENSG00000211751 ENSG00000004468 ENSG00000019582 ENSG00000173821 ENSG00000181847 ENSG00000117560 ENSG00000117091 ENSG00000160791 ENSG00000077150 ENSG00000198624 ENSG00000172215 ENSG00000089012 ENSG00000173193 ENSG00000163600 ENSG00000163219 ENSG00000198821 ENSG00000172116 ENSG00000180353 ENSG00000168961 ENSG00000135426ENSG00000120899 ENSG00000211799 ENSG00000182179 ENSG00000077420 ENSG00000115165 ENSG00000165178 ENSG00000141506 ENSG00000227191 ENSG00000164483 ENSG00000173762 ENSG00000185215 ENSG00000285048 ENSG00000133106 ENSG00000182487 ENSG00000070190 ENSG00000187764 ENSG00000083799 ENSG00000196329 ENSG00000128340 ENSG00000130475 ENSG00000186470 ENSG00000077984 ENSG00000139192 ENSG00000106948 ENSG00000073861 ENSG00000124203 ENSG00000115956 ENSG00000160654 ENSG00000110876 ENSG00000160255 ENSG00000168404 ENSG00000136167 ENSG00000174946 ENSG00000103522 ENSG00000220517 ENSG00000113263 ENSG00000137193 ENSG00000271503 ENSG00000152969 ENSG00000160219 ENSG00000167207 ENSG00000197142 ENSG00000157303 ENSG00000115415 ENSG00000262418 ENSG00000010030 ENSG00000163564 ENSG00000145649 ENSG00000196684 ENSG00000124256 ENSG00000101082 ENSG00000164307 ENSG00000104894 ENSG00000147138 ENSG00000132530 ENSG00000010671 ENSG00000122188 ENSG00000067066 ENSG00000173208 ENSG00000104856 ENSG00000115607 ENSG00000137474ENSG00000160593 ENSG00000109943 ENSG00000100453 ENSG00000109684 ENSG00000095585 ENSG00000185669 ENSG00000172543 ENSG00000089692 ENSG00000181036 ENSG00000137496 ENSG00000076662 ENSG00000152766 ENSG00000027075 ENSG00000225492 ENSG00000133321 ENSG00000213402 ENSG00000105967 ENSG00000133943 ENSG00000135148 ENSG00000132109 ENSG00000187862 ENSG00000128284 ENSG00000131401 ENSG00000135899 ENSG00000156234 ENSG00000177409 ENSG00000121895 ENSG00000152229 ENSG00000121281 ENSG00000155849 ENSG00000065675 ENSG00000104951 ENSG00000133574 ENSG00000096968 ENSG00000138496 ENSG00000155926 ENSG00000162645 ENSG00000173020 ENSG00000142185 ENSG00000137078 ENSG00000066294 ENSG00000043462 ENSG00000147443 ENSG00000137628 ENSG00000185862 ENSG00000131979 ENSG00000248672 ENSG00000042980 ENSG00000288199 ENSG00000228163 ENSG00000282928 ENSG00000267312 ENSG00000162654 ENSG00000167895 ENSG00000211694 ENSG00000145779 ENSG00000049249 ENSG00000074370 ENSG00000121807 ENSG00000211795 ENSG00000118308 ENSG00000015133ENSG00000139193 ENSG00000205220 ENSG00000137101 ENSG00000213809 ENSG00000131203 ENSG00000068724 ENSG00000092929 ENSG00000163874 ENSG00000266094 ENSG00000093072 ENSG00000180448 ENSG00000205436 ENSG00000133805 ENSG00000121858 ENSG00000149781 ENSG00000130755 ENSG00000028137 ENSG00000161405 ENSG00000275302 ENSG00000116852 ENSG00000163492 ENSG00000100336 ENSG00000023445 ENSG00000160796 ENSG00000122224 ENSG00000128604 ENSG00000176083 ENSG00000143185 ENSG00000100351 ENSG00000182162 ENSG00000277734 ENSG00000211779 ENSG00000196358 ENSG00000158985 ENSG00000066336 ENSG00000240891 ENSG00000163840 ENSG00000086730 ENSG00000234518 ENSG00000120280 ENSG00000100342 ENSG00000143119 ENSG00000189350 ENSG00000119686 ENSG00000130487 ENSG00000179715 ENSG00000211789 ENSG00000138378 ENSG00000100450 ENSG00000211689 ENSG00000002549 ENSG00000100906 ENSG00000181631 ENSG00000155307 ENSG00000122223 ENSG00000211786 ENSG00000213203 ENSG00000059378 ENSG00000235568 ENSG00000135077 ENSG00000275824 ENSG00000175857ENSG00000196189 ENSG00000239713 ENSG00000160856 ENSG00000282711 ENSG00000229754 ENSG00000118971 ENSG00000158714 ENSG00000196405 ENSG00000155629 ENSG00000159496 ENSG00000028277 ENSG00000111537 ENSG00000168421 ENSG00000089639 ENSG00000074706 ENSG00000138755 ENSG00000172794 ENSG00000273686 ENSG00000254087 ENSG00000047365 ENSG00000000938 ENSG00000118503 ENSG00000095370 ENSG00000166501 ENSG00000183813 ENSG00000142512 ENSG00000140678 ENSG00000282657 ENSG00000075884 ENSG00000197721 ENSG00000211787 ENSG00000198771 ENSG00000169403 ENSG00000078589 ENSG00000114127 ENSG00000211803 ENSG00000110848 ENSG00000250254 ENSG00000100365 ENSG00000156127 ENSG00000174123 ENSG00000081320 ENSG00000166428 ENSG00000136541 ENSG00000113088 ENSG00000171115 ENSG00000211802 ENSG00000169442 ENSG00000211817 ENSG00000169245 ENSG00000100298 ENSG00000090554 ENSG00000115267 ENSG00000135114 ENSG00000091592 ENSG00000211807 ENSG00000186265 ENSG00000185187 ENSG00000275565 ENSG00000139626 ENSG00000137767 ENSG00000211794ENSG00000027869 ENSG00000165025 ENSG00000108405 ENSG00000102034 ENSG00000211675 ENSG00000211792 ENSG00000003400 ENSG00000113249 ENSG00000070915 ENSG00000150637 ENSG00000243811 ENSG00000111348 6. A biomarker for predicting the efficacy of an immune checkpoint inhibitor, characterized by comprising the NLRC5 gene and two genes A and B, indicated by two IDs in any of #1 to #250 of List 3. (List 3) 7. The biomarker according to claim 5 or 6, characterized in that it is for predicting the effect of an anti-PD-1 inhibitor or an anti-CTLA-4 inhibitor.
8. A determination device for predicting the effect of an immune checkpoint inhibitor, characterized by calculating an evaluation value regarding the effect of the immune checkpoint inhibitor based on the expression level of each gene contained in the biomarker described in claim 5 or 6.
9. The determination device according to claim 8, characterized in that it comprises a model trained to take the expression levels of each gene as input and output the evaluation value.
10. The determination device according to claim 9, characterized in that the model is based on a random forest.
11. A method for predicting the effectiveness of an immune checkpoint inhibitor, characterized by calculating an evaluation value regarding the effectiveness of the immune checkpoint inhibitor based on the expression levels of each gene included in the biomarker described in claim 5 or 6.
12. A learning model for predicting the effect of an immune checkpoint inhibitor, characterized in that it accepts input of the expression levels of each gene included in the biomarker described in claim 5 or 6, and functions a computer to calculate and output an evaluation value regarding the effect of the immune checkpoint inhibitor.
13. A method for generating a learning model according to claim 12, characterized in that the effects of administering immune checkpoint inhibitors to a group of patients and the expression levels of each gene shown by gene expression analysis of that group of patients are used as learning data.
Citation Information
Patent Citations
Method for predicting response to immune checkpoint inhibitor
WO2022059785A1
Biomarker for predicting response to immune checkpoint inhibitor
WO2023022200A1
Method of predicting non-small cell lung cancer (NSCLC) patient drug response or time until death or cancer progression from circulating tumor DNA (CTDNA) utilizing signals from both baseline ctdna level and longitudinal change of ctdna level over time
WO2024107599A1