Immune checkpoint blocking response prediction method based on neoantigen quality

Through the prediction method based on the quality of neoantigens, the ICB response prediction model is constructed using multimodal features, which solves the problem of small application scope and low accuracy of the existing ICB response prediction methods, and achieves higher prediction accuracy and scope of application.

CN120072035AActive Publication Date: 2025-05-30HARBIN INST OF TECH

Patent Information

Application Number
CN202510145037.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The existing immune checkpoint blockade (ICB) response prediction methods have small scope of application and low accuracy, and cannot fully reflect the natural characteristics of neoantigens, resulting in limited accuracy and scope of application of ICB therapy response prediction.

Method used

Using a prediction method based on neoantigen mass, multimodal features were obtained by obtaining tumor tissue whole exon data and pre-processing to obtain multimodal features, including HLA evolution divergence and immunogenic neoantigen load, and a prediction model was constructed to predict clinical response data.

Benefits of technology

It improves the accuracy and scope of application of ICB response prediction, can more accurately quantify the number of high-quality neoantigens, and supports personalized immunotherapy decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072035A_ABST
    Figure CN120072035A_ABST
Patent Text Reader

Abstract

The invention discloses an immune checkpoint blocking response prediction method based on neoantigen quality, and relates to the technical field of computational biology and biological information. The problems that an existing immune checkpoint blocking (ICB) response prediction method is small in application range and low in accuracy are solved. The method comprises the following steps: acquiring clinical data of a subject and whole exon data of a tumor tissue; preprocessing the whole exon data of the tumor tissue, and obtaining a training data set in combination with clinical data; the immune checkpoint blocking response prediction model predicts clinical response data of a subject according to input data of each training sample in the training data set, and determines an optimal prediction threshold value of the model according to a prediction effect; and predicting clinical response data of the current subject to be tested by using the immune checkpoint blocking response prediction model under the optimal prediction threshold NCuff. The ICB response prediction method is mainly used for ICB response prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computational biology and bioinformatics technology. Background Art

[0002] In recent years, immune checkpoint blockade (ICB) therapy has revolutionized the treatment of patients with advanced cancer. These inhibitors, including antibodies targeting programmed death-1 (anti-PD-1, such as nivolumab and pembrolizumab) or cytotoxic T-lymphocyte-associated antigen 4 (anti-CTLA-4, such as ipilimumab), can reverse T regulatory cell-mediated immunosuppression. However, durable clinical benefits are limited to a small subset of patients. Therefore, researchers are seeking predictive biomarkers for ICB treatment response, such as tumor mutational burden (TMB), T cell infiltration level, PD-L1 immunohistochemistry, somatic genomic features, and immunophenotype. The generation of tumor neoantigens that escape central tolerance of T cells in the immune system is crucial for anti-tumor immunity. TMB can partially reflect the generation of neoantigens in patients due to genomic coding mutations. Generally, biomarkers for predicting ICB response are useful but have limitations.

[0003] Currently, many ICB response prediction methods based on neoantigen immunogenicity have been developed. However, the neoantigen immunogenicity of patients calculated by these methods cannot comprehensively reflect the natural characteristics due to HLA binding affinity or epitope similarity arrangement, and cannot directly explain the potential mechanism of anti-tumor response, resulting in low accuracy and limited applicability of ICB response prediction. Considering that multimodal methods are of great value for determining the efficacy of ICB treatment, and the quality of tumor-specific neoantigens mainly depends on their natural characteristics, such as HLA presentation, T cell cross-reactivity, and T cell recognition. Therefore, a comprehensive functional model is needed to meet the patient stratification for ICB treatment. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems of small applicability and low accuracy of existing immune checkpoint blockade (ICB) response prediction methods; the present invention provides an immune checkpoint blockade response prediction method based on neoantigen quality.

[0005] An immune checkpoint blockade response prediction method based on neoantigen quality, the method comprising the following steps:

[0006] S1. Obtain the clinical response data and tumor tissue whole exome data of subjects undergoing immune checkpoint blockade treatment; the subjects are cancer patients, and the clinical response data are responders or non-responders;

[0007] S2. Preprocess the tumor tissue whole exome data:

[0008] Process the whole-exome data of the tumor tissues of each subject to obtain the multimodal features of each subject; according to the multimodal features of each subject, obtain the HLA evolutionary divergence and immunogenic neoantigen load of the subject, and use them as the input data of the training samples. At the same time, use the obtained clinical response data of the subject as the response label of the training samples to obtain the training dataset;

[0009] The multimodal features of each subject include the wild peptide and mutant peptide of the neoantigen at the tumor mutation site, the HLA binding affinity of the wild peptide and mutant peptide of the neoantigen, and the HLA allele typing;

[0010] The immunogenic neoantigen load of the subject is the number of high-quality immunogenic neoantigens in the subject's body. Among them, when the quality Q of the neoantigen in the subject's body is greater than 1, it is defined as a high-quality immunogenic neoantigen;

[0011] S3. The immune checkpoint blockade response prediction model at each candidate prediction threshold within the threshold interval predicts the clinical response data of the subject according to the input data of each training sample in the training dataset, and combines the response label of the training sample to determine the model prediction effect at the candidate prediction threshold;

[0012] Determine the optimal prediction threshold N of the model according to the model prediction effects at all candidate prediction thresholds Cutoff ;

[0013] S4. Preprocess the whole-exome data of the tumor tissues of the current subject to be tested to obtain the HLA evolutionary divergence and immunogenic neoantigen load of the current subject to be tested;

[0014] Optimal prediction threshold N Cutoff The immune checkpoint blockade response prediction model under predicts the clinical response data of the current subject to be tested according to the HLA evolutionary divergence and immunogenic neoantigen load of the current subject to be tested.

[0015] Preferably, the implementation method of processing the whole-exome data of the tumor tissues of each subject in step S2 to obtain the multimodal features of each subject includes:

[0016] Check the quality of the whole-exome data of the tumor tissues, filter out low-quality fragments, and obtain qualified sequencing data;

[0017] Align the qualified sequencing data to the human reference genome data, delete the repetitive sequences, and obtain the detectable sequences; perform somatic single nucleotide variant identification on the detectable sequences, screen out the tumor mutation sites, and obtain the wild peptide and mutant peptide of the neoantigen at each tumor mutation site according to each tumor mutation site;

[0018] Determine the HLA binding affinities of the wild-type peptides and mutant peptides of the neoantigens at the positions of each tumor mutation site;

[0019] Also determine the HLA allele typing based on the whole exome data of the tumor tissues of the subjects.

[0020] Preferably, the implementation manner of obtaining the HLA evolutionary divergence and the immunogenic neoantigen load of the subjects according to the multimodal features of each subject in step S2 includes:

[0021] Determine the HLA evolutionary divergence using the Grantham distance algorithm according to the HLA allele typing in the multimodal features of each subject;

[0022] Determine the HLA presentation A of the neoantigen according to the HLA binding affinities of the wild-type peptides and mutant peptides of each neoantigen in the multimodal features of each subject;

[0023] Calculate the T cell cross-reactivity C of the neoantigen according to the wild-type peptides and mutant peptides of each neoantigen in the multimodal features of each subject;

[0024] Obtain the immunogenicity score R of each neoantigen according to the mutant peptides and HLA allele typing of each neoantigen in the multimodal features of each subject;

[0025] Obtain the quality Q of each neoantigen as Q = R×(A + C) according to A, C, and R;

[0026] According to the values of the quality Q of each neoantigen, count the number of high-quality immunogenic neoantigens in the subject's body and use it as the immunogenic neoantigen load of the subject.

[0027] Preferably, and are respectively the HLA binding affinities of the wild-type peptide and mutant peptide of the neoantigen.

[0028] Preferably, wherein,

[0029] d i is the amino acid substitution position scaling factor when the i-th amino acid in the wild-type peptide p WT of the neoantigen mutates to the mutant peptide p MT ;

[0030] is the value at the row and column where

[0031] is located in the position-dependent amino acid substitution matrix M; WTThe i-th amino acid in is the mutant peptide p MT The i-th amino acid in, where i = 1, 2... N and N is the length of the wild-type peptide or mutant peptide.

[0032] Preferably, according to the mutant peptides of each neoantigen and the HLA allele typing in the multimodal characteristics of each subject, the immunogenicity score R of each neoantigen is obtained. The specific implementation method includes:

[0033] Based on the mutant peptide of the neoantigen, the neoantigen sequence is obtained;

[0034] Based on the HLA allele typing, the HLA pseudo-sequence is obtained;

[0035] The encoding methods of the neoantigen sequence and the HLA pseudo-sequence are converted into an encoding matrix. The features of the neoantigen and HLA in the encoding matrix are extracted by a feature encoder based on deep learning or machine learning, and the extracted features are concatenated to obtain combined features;

[0036] The combined features are decoded, and finally the immunogenicity score R of the neoantigen is obtained.

[0037] Preferably, the immune checkpoint blockade response prediction model includes an immune neoantigen quality score calculation module and a threshold comparison module;

[0038] The immune neoantigen quality score calculation module calculates the immune neoantigen quality score ICBNQ of the subject according to the HLA evolutionary divergence and the immunogenic neoantigen burden of the subject; ICBNQ = HED × lg(INB + 1);

[0039] where HED is the HLA evolutionary divergence and INB is the immunogenic neoantigen burden;

[0040] The threshold comparison module is used to determine the quality score ICBNQ. When the quality score ICBNQ is greater than or equal to the optimal prediction threshold N Cutoff the subject is determined to be a responder and output as the predicted clinical response data; otherwise, the subject is determined to be a non-responder and output as the predicted clinical response data.

[0041] Preferably, in step S3, the implementation method of determining the model prediction effect under each candidate prediction threshold is:

[0042] Taking the Youden index J of the prediction model as the prediction effect, where

[0043] J = (sensitivity + specificity - 1);

[0044] sensitivity = TP / (TP + FN) × 100%;

[0045] specificity = TN / (TN + FP) × 100%;

[0046] sensitivity is the sensitivity of the prediction model, specificity is the specificity of the prediction model, TP is the total number of true positives, TN is the total number of true negatives, FP is the total number of false positives, and FN is the total number of false negatives;

[0047] True positive means that the prediction model predicts the clinical response data as a responder and the response label is a responder;

[0048] True negative means that the prediction model predicts the clinical response data as a non-responder and the response label is a non-responder;

[0049] False positive means that the prediction model predicts the clinical response data as a responder and the response label is a non-responder;

[0050] False negative means that the prediction model predicts the clinical response data as a non-responder and the response label is a responder.

[0051] Preferably, in step S3, according to the model prediction effects under all candidate prediction thresholds, the optimal prediction threshold N of the model is determined Cutoff The implementation method is:

[0052] Take the candidate prediction threshold corresponding to the maximum value of the model prediction effect as the optimal prediction threshold N of the model Cutoff .

[0053] The immune checkpoint blockade response prediction device based on neoantigen quality includes a storage device, a processor, and a computer program stored in the storage device and executable on the processor. The processor executes the computer program to implement the immune checkpoint blockade response prediction method based on neoantigen quality as described.

[0054] The beneficial effects brought by the present invention are:

[0055] The present invention provides an immune checkpoint blockade response prediction method based on neoantigen quality, which can accurately quantify the number of high-quality neoantigens in a subject and perform quantitative analysis on the number of high-quality neoantigens in the subject. The present invention is beneficial to improving the accuracy of immune checkpoint blockade response prediction, significantly improving the precision and application scope of the prediction of the response to immune checkpoint blockade therapy, and providing a scientific basis for personalized decision-making in immunotherapy. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a flowchart of the immune checkpoint blockade response prediction method based on neoantigen quality;

[0057] Figure 2 Schematic diagram of the principle for obtaining HLA evolutionary divergence and immunogenic neoantigen load;

[0058] Figure 3 Schematic diagram of the principle for obtaining the immunogenicity score of neoantigens;

[0059] Figure 4 It is the position-dependent amino acid substitution matrix diagram provided by the present invention;

[0060] Figure 5 It is the performance diagram of the immune checkpoint blockade response prediction model of the present invention on test data; wherein, Figure 5 a is the performance diagram for predicting subjects with non-small cell lung cancer, Figure 5 b is the performance diagram for predicting subjects with melanoma, Figure 5 c is the performance diagram for predicting subjects with unknown cancer. Detailed implementation manners

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0063] Detailed implementation manner 1. Refer to Figure 1 and Figure 2 To illustrate this implementation manner, this implementation manner provides an immune checkpoint blockade response prediction method based on neoantigen quality, and the method includes the following steps:

[0064] S1. Obtain the clinical response data and tumor tissue whole exome data of subjects undergoing immune checkpoint blockade treatment; the subjects are tumor patients, the clinical response data is responders or non-responders, and the responders are patients who have been cured, and the non-responders are patients who have not been cured;

[0065] S2. Preprocess the tumor tissue whole exome data:

[0066] Process the tumor tissue whole exome data of each subject to obtain the multi-modal features of each subject; according to the multi-modal features of each subject, obtain the HLA evolutionary divergence and immunogenic neoantigen load of the subject, and use them as the input data of the training sample. At the same time, use the obtained clinical response data of this subject as the response label of the training sample to obtain the training data set;

[0067] The multimodal features of each subject include the wild-type peptide and mutant peptide of neoantigens at the tumor mutation sites, the HLA binding affinities of the wild-type peptide and mutant peptide of neoantigens, and HLA allele typing;

[0068] The immunogenic neoantigen load of the subject is the number of high-quality immunogenic neoantigens in the subject's body. Among them, when the quality Q of the neoantigen in the subject's body is >1, it is defined as a high-quality immunogenic neoantigen;

[0069] S3. The immune checkpoint blockade response prediction model at each candidate prediction threshold within the threshold range predicts the clinical response data of the subject according to the input data of each training sample in the training dataset, and combines the response label of the training sample to determine the model prediction effect at the candidate prediction threshold;

[0070] Determine the optimal prediction threshold N of the model according to the model prediction effects at all candidate prediction thresholds Cutoff ;

[0071] S4. Preprocess the whole exome data of the tumor tissue of the current subject to be tested to obtain the HLA evolutionary divergence and immunogenic neoantigen load of the current subject to be tested;

[0072] Optimal prediction threshold N Cutoff The immune checkpoint blockade response prediction model under predicts the clinical response data of the current subject to be tested according to the HLA evolutionary divergence and immunogenic neoantigen load of the current subject to be tested.

[0073] In this embodiment, the whole exome data of the tumor tissue, that is, the DNA data (fasta file) of the subject, is combined with the clinical data of the subject to obtain a training dataset, which improves the integrity and matching degree of important information in the prediction of individualized immune checkpoint blockade (ICB) treatment response, making the prediction more accurate and targeted.

[0074] By constructing a multimodal model framework based on antigen quality, multimodal features are obtained from the whole exome data of the tumor tissue of each subject, and deep immunogenic information in the massive antigen data is captured, enabling the prediction task of ICB treatment to share the same model framework under different cancer backgrounds;

[0075] By selecting the optimal prediction threshold N Cutoff , the sensitivity of the model to individual tumor patients can be increased, thereby improving the accuracy of ICB treatment response prediction.

[0076] In specific applications, the clinical response data and the whole-exome data of tumor tissues may include, but are not limited to: the data corresponding to 919 subjects were used to construct the training dataset. These samples were collected from multiple references in the prior art, specifically involving 6 cancer types, namely: cutaneous melanoma, with the number of training samples n = 359; non-small cell lung cancer, with the number of training samples n = 519; head and neck squamous cell carcinoma, with the number of training samples n = 12; bladder cancer, with the number of training samples n = 27; anal cancer, with the number of training samples n = 1; sarcoma, with the number of training samples n = 1.

[0077] Further, referring to Figure 3 , the implementation manner of processing the whole-exome data of tumor tissues of each subject in step S2 to obtain the multimodal features of each subject includes:

[0078] S211. Check the quality of the whole-exome data (fasta file) of tumor tissues, filter out low-quality fragments, and obtain qualified sequencing data;

[0079] Specifically, use a tool (such as FastQC) to check the quality of the original data for the fasta file, which may specifically include the length of the sequence, GC content, contamination situation, etc. Filter out the sequences with a quality score lower than Q30 to obtain qualified sequencing data;

[0080] S212. Align the qualified sequencing data to the human reference genome data and delete the duplicate sequences to obtain detectable sequences;

[0081] For the qualified sequencing data, use Burrows-Wheeler Aligner (BWA, http: / / bio-bwa.sourceforge.net) version 0.7.17 to align the original fastq file with the reference human genome GRCh38, and use the Picard (v 2.26.6) tool to delete the duplicate sequences to obtain detectable sequences;

[0082] S213. Perform somatic single nucleotide variant identification on the detectable sequences, screen out the positions of tumor mutation sites, and obtain the wild-type peptides and mutant peptides of neoantigens at each tumor mutation site according to the positions of each tumor mutation site;

[0083] The screening of the positions of tumor mutation sites can be achieved by the prior art. Specifically, the MuTect2 tool can be used for somatic single nucleotide variant identification. Candidate tumor mutations will only be selected under the following conditions:

[0084] 1) Detected more than 5 times in high-quality reads;

[0085] 2) The variant allele frequency (VAF) is greater than 1%;

[0086] 3) not observed in multiple single nucleotide polymorphism databases (e.g., 1000Genomes Project with version phase 3, dbSNP version 137, and COSMIC database);

[0087] Specifically, to obtain the wild-type peptide and mutant peptide of the neoantigen at each tumor mutation site, the GATK4 pipeline in the Terra cloud platform (https: / / app.terra.bio / ) can be used to obtain the wild-type peptide and mutant peptide of the neoantigen.

[0088] S214. Determine the HLA binding affinity of each wild-type peptide and mutant peptide of the neoantigen at each tumor mutation site. The specific tool that can be used is NetMHCpan-4.1;

[0089] S215. Also determine the HLA allele typing based on the whole exome data of the tumor tissue of the subject. The specific tool that can be used is Optitype.

[0090] See Figure 2 , further, the implementation manner of obtaining the HLA evolutionary divergence and immunogenic neoantigen load of the subject according to the multimodal features of each subject in step S2 includes:

[0091] Determine the HLA evolutionary divergence using the Grantham distance algorithm according to the HLA allele typing in the multimodal features of each subject;

[0092] Determine the HLA presentation of the neoantigen according to the HLA binding affinity of each wild-type peptide and mutant peptide of the neoantigen in the multimodal features of each subject wherein, and are the HLA binding affinities of the wild-type peptide and mutant peptide of the neoantigen respectively;

[0093] Calculate the T cell cross-reactivity C of the neoantigen according to each wild-type peptide and mutant peptide of the neoantigen in the multimodal features of each subject. Specifically,

[0094] wherein,

[0095] d i is the amino acid substitution position scaling factor when the i-th amino acid in the wild-type peptide p WT of the neoantigen mutates to the mutant peptide p MT ;

[0096] is in the position-dependent amino acid substitution matrix M The value at the row and column positions;

[0097] is the wild peptide p WT at the i-th amino acid, is the mutant peptide p MT at the i-th amino acid, where i = 1, 2... N and N is the length of the wild peptide or mutant peptide.

[0098] Based on the mutant peptides of each neoantigen and HLA allele typing, the immunogenicity score R of each neoantigen is obtained;

[0099] Based on A, C, and R, the quality Q of each neoantigen is obtained as Q = R × (A + C);

[0100] Among them, R is the immunogenicity score of the neoantigen, which is used to characterize the ability to activate T cells and trigger immunogenicity. A is HLA presentation, which is used to characterize the ability of the neoantigen to be presented to T cells. C is T cell cross-reactivity, which is used to characterize the ability of T cells to recognize the mutant peptide and wild peptide of the neoantigen;

[0101] Based on the value of the quality Q of each neoantigen, the number of high-quality immunogenic neoantigens in the subject's body is counted and used as the immunogenic neoantigen load of the subject.

[0102] In this preferred embodiment, the specific implementation means of obtaining the HLA evolutionary divergence and immunogenic neoantigen load of the subject are given. Its advantage is to improve the accuracy of immune checkpoint blockade prediction and meet patient stratification by capturing individualized HLA evolutionary divergence and immunogenic neoantigen load. The specific form of the position-dependent amino acid substitution matrix M is also given during specific application. See Figure 4 .

[0103] See Figure 3 , and further given, the specific implementation means of obtaining the immunogenicity score R of each neoantigen based on the wild peptides of each neoantigen and HLA allele typing in the multi-modal features of each subject include:

[0104] Based on the mutant peptide of the neoantigen, the neoantigen sequence is obtained;

[0105] Based on HLA allele typing, the HLA pseudo-sequence is obtained;

[0106] The neoantigen sequence and HLA pseudo-sequence are converted into a coding matrix through coding methods such as Blosum62 or One-hot. The features of the neoantigen and HLA in the coding matrix are extracted through a feature encoder based on deep learning or machine learning (such as convolutional neural network, Transformer encoder, long short-term memory artificial neural network, etc.), and the extracted features are concatenated to obtain joint features;

[0107] Decode the combined features. For example, use a fully connected network for decoding to finally obtain the neoantigen immunogenicity score R.

[0108] The specific process of obtaining the neoantigen immunogenicity score R given in this preferred embodiment improves the credibility of the neoantigen immunogenicity score by using a deep learning or machine learning model.

[0109] Furthermore, the specific composition of the immune checkpoint blockade response prediction model is given. The immune checkpoint blockade response prediction model includes an immune neoantigen quality score calculation module and a threshold comparison module;

[0110] The immune neoantigen quality score calculation module calculates the immune neoantigen quality score ICBNQ of the subject according to the HLA evolutionary divergence and the immunogenic neoantigen load of the subject.

[0111] ICBNQ = HED × lg(INB + 1);

[0112] where HED is the HLA evolutionary divergence and INB is the immunogenic neoantigen load;

[0113] The threshold comparison module is used to determine the quality score ICBNQ. When the quality score ICBNQ is greater than or equal to the optimal prediction threshold N Cutoff it is determined that the subject responds to ICB treatment, the subject is a responder, and it is output as the predicted clinical response data; otherwise, it is determined that the subject does not respond to ICB treatment, the subject is a non-responder, and it is output as the predicted clinical response data.

[0114] Furthermore, in step S3, the implementation method of determining the model prediction effect under each candidate prediction threshold is:

[0115] Take the Youden index J of the prediction model as the prediction effect, where

[0116] J = (sensitivity + specificity - 1);

[0117] sensitivity = TP / (TP + FN) × 100%;

[0118] specificity = TN / (TN + FP) × 100%;

[0119] sensitivity is the sensitivity of the prediction model, specificity is the specificity of the prediction model, TP is the total number of true positives, TN is the total number of true negatives, FP is the total number of false positives, and FN is the total number of false negatives;

[0120] A true positive is when the prediction model predicts the clinical response data as a responder and the response label is also a responder;

[0121] A true negative is when the prediction model predicts the clinical response data as a non-responder and the response label is also a non-responder;

[0122] A false positive is when the prediction model predicts the clinical response data as a responder and the response label is a non-responder;

[0123] A false negative is when the prediction model predicts the clinical response data as a non-responder and the response label is a responder.

[0124] Furthermore, in step S3, according to the model prediction effects under all candidate prediction thresholds, the optimal prediction threshold N of the model is determined Cutoff The implementation method is:

[0125] Take the candidate prediction threshold corresponding to the maximum value of the model prediction effect as the optimal prediction threshold N of the model Cutoff .

[0126] Specific implementation manner two: The immune checkpoint blockade response prediction device based on neoantigen quality described in this implementation manner includes a storage device, a processor, and a computer program stored in the storage device and executable on the processor. The processor executes the computer program to implement the immune checkpoint blockade response prediction method based on neoantigen quality as described in specific implementation manner one.

[0127] Verification test:

[0128] The effectiveness of the present invention is demonstrated through the following verification test, specifically:

[0129] Using the immune checkpoint blockade response prediction method based on neoantigen quality provided by the present invention, the prediction performance on the test set is as Figure 5 shown.

[0130] Figure 5 represents the receiver operating characteristic curve (ROC) of the test samples, where the ordinate is the true positive rate (TPR) and the abscissa is the false positive rate (FPR). The area under the receiver operating characteristic curve (AUC) in the figure is the evaluation index of the prediction performance.

[0131] Figure 5 In, the AUC result for predicting subjects with non-small cell lung cancer is 0.85, and the 95% confidence interval is (0.69, 0.96), see Figure 5 a; the AUC result for predicting subjects with melanoma is 0.86, and the 95% confidence interval is (0.71, 0.97), see Figure 5b; The AUC result for predicting subjects with unknown cancers was 0.75, and the 95% confidence interval was (0.58, 0.89), see Figure 5 c; It can be seen that the method of the present invention has high accuracy and is also applicable to unknown cancers.

[0132] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Accordingly, it should be understood that many modifications may be made to the exemplary embodiments, and other arrangements may be designed, without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that the different dependent claims and the features described herein may be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a particular embodiment may be used in other described embodiments.

Claims

1. A method for predicting immune checkpoint blockade response based on neoantigen quality, characterized in that: The method comprises the following steps: S1. Obtain clinical response data and whole exome data of tumor tissues of subjects treated with immune checkpoint blockade; the subjects are tumor patients, and the clinical response data are responders or non-responders; S2. Preprocessing of tumor tissue whole exome data: The whole exome data of the tumor tissue of each subject is processed to obtain the multimodal characteristics of each subject; based on the multimodal characteristics of each subject, the HLA evolution divergence and immunogenic neoantigen load of the subject are obtained and used as the input data of the training sample. At the same time, the clinical response data of the subject obtained is used as the response label of the training sample to obtain the training data set; The multimodal characteristics of each subject include wild-type peptides and mutant peptides of the neoantigen at the tumor mutation site, HLA binding affinity of the wild-type peptides and mutant peptides of the neoantigen, and HLA allele typing; The immunogenic neoantigen load of the subject is the number of high-quality immunogenic neoantigens in the subject, wherein when the quality Q of the neoantigen in the subject is greater than 1, it is defined as a high-quality immunogenic neoantigen; S3. The immune checkpoint blockade response prediction model at each candidate prediction threshold within the threshold interval predicts the clinical response data of the subject based on the input data of each training sample in the training data set, and determines the model prediction effect at the candidate prediction threshold in combination with the response label of the training sample; According to the model prediction effect under all candidate prediction thresholds, determine the optimal prediction threshold N of the model Cutoff ; S4. Preprocessing the whole exome data of the tumor tissue of the current subject to be tested to obtain the HLA evolution divergence and immunogenic new antigen load of the current subject to be tested; The best prediction threshold N Cutoff The immune checkpoint blockade response prediction model below predicts the clinical response data of the current candidate subject based on the HLA evolution divergence and immunogenic neoantigen load of the current candidate subject.

2. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 1, characterized in that: In step S2, the whole exome data of the tumor tissue of each subject is processed to obtain the multimodal features of each subject, and the implementation method includes: Check the quality of the whole exome data of tumor tissue, filter out low-quality fragments, and obtain sequencing data of qualified quality; Compare the qualified sequencing data to the human reference genome data, delete the repeated sequences, and obtain the detectable sequences; identify the somatic single nucleotide variations of the detectable sequences, screen out the tumor mutation point positions, and obtain the wild type peptide and mutant peptide of the new antigen at each tumor mutation point position according to the position of each tumor mutation point; Determine the HLA binding affinity of the wild peptide and mutant peptide of each neoantigen based on the wild peptide and mutant peptide of each neoantigen at the location of each tumor mutation point; HLA allele typing was also determined based on whole exome data from the subjects' tumor tissues.

3. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 1, characterized in that: In step S2, the implementation method of obtaining the HLA evolution divergence and immunogenic neoantigen load of the subject according to the multimodal characteristics of each subject includes: According to the HLA allele typing in the multimodal characteristics of each subject, the Grantham distance algorithm was used to determine the HLA evolution divergence; Determine the HLA presentation A of the neoantigen based on the HLA binding affinity of the wild-type peptide and the mutant peptide of each neoantigen in the multimodal characteristics of each subject; Calculate the T cell cross-reactivity C of the new antigen based on the wild type peptide and mutant peptide of each new antigen in the multimodal characteristics of each subject; According to the mutant peptide and HLA allele typing of each neoantigen in the multimodal characteristics of each subject, the immunogenicity score R of each neoantigen was obtained; According to A, C and R, the mass of each new antigen is obtained as Q = R × (A + C); According to the quality Q value of each new antigen, the number of high-quality immunogenic new antigens in the subject's body is counted and used as the subject's immunogenic new antigen load.

4. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 3, characterized in that: and These are the HLA binding affinities of the neoantigen wild-type peptide and mutant peptide, respectively.

5. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 3, characterized in that: in, d i The wild type peptide p WT The ith amino acid in the peptide is mutated to form the mutant peptide p MT Scaling factor for amino acid substitution position; is the position-dependent amino acid substitution matrix M The line The value of the column; Wild peptide p WT The ith amino acid in The mutant peptide p MT The i-th amino acid in is, i = 1, 2...N, N is the length of the wild-type peptide or mutant peptide.

6. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 3, characterized in that: According to the mutant peptide and HLA allele typing of each neoantigen in the multimodal characteristics of each subject, the immunogenicity score R of each neoantigen is obtained. The specific implementation method includes: According to the mutant peptide of the new antigen, the new antigen sequence is obtained; according to the HLA allele typing, the HLA pseudo sequence is obtained; The encoding method of the new antigen sequence and HLA pseudo sequence is converted into a coding matrix, and the features of the new antigen and HLA in the coding matrix are extracted through a feature encoder based on deep learning or machine learning, and the extracted features are spliced ​​to obtain joint features; The combined features are decoded and the neoantigen immunogenicity score R is finally obtained.

7. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 1, characterized in that: The immune checkpoint blockade response prediction model includes an immune neoantigen quality score calculation module and a threshold comparison module; The immune new antigen quality score calculation module calculates the immune new antigen quality score ICBNQ of the subject according to the HLA evolution divergence and immunogenic new antigen load of the subject; ICBNQ = HED × lg (INB + 1); Among them, HED is the HLA evolutionary divergence, and INB is the immunogenic neoantigen load; The threshold comparison module is used to judge the quality score ICBNQ. When the quality score ICBNQ is greater than or equal to the best prediction threshold N Cutoff When the subject is determined to be a responder, the subject is output as the predicted clinical response data; otherwise, the subject is determined to be a non-responder, and the subject is output as the predicted clinical response data.

8. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 1, characterized in that: In step S3, the implementation method of determining the model prediction effect under each candidate prediction threshold is as follows: The Youden index J of the prediction model is used as the prediction effect, where: J=(sensitivity+specificity-1); sensitivity=TP / (TP+FN)×100%; specificity=TN / (TN+FP)×100%; sensitivity is the sensitivity of the prediction model, specificity is the specificity of the prediction model, TP is the total number of true positives, TN is the total number of true negatives, FP is the total number of false positives, and FN is the total number of false negatives; A true positive is when the prediction model predicts that the clinical response data is a responder and the response label is a responder; A true negative is when the prediction model predicts that the clinical response data is a non-responder and the response label is a non-responder; A false positive is when the prediction model predicts that the clinical response data is a responder, and the response label is a non-responder; A false negative is when the prediction model predicts that the clinical response data is a non-responder and the response label is a responder.

9. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 1, characterized in that: In step S3, the optimal prediction threshold N of the model is determined based on the model prediction results under all candidate prediction thresholds. Cutoff The implementation is: Take the candidate prediction threshold corresponding to the maximum value of the model prediction effect as the model's optimal prediction threshold N Cutoff .

10. An immune checkpoint blockade response prediction device based on neoantigen quality, comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, characterized in that: The processor executes a computer program to implement the method for predicting immune checkpoint blockade response based on neoantigen quality as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Prediction method and device for tumor neoantigen load

    CN113053458A

  • Tumor sample immune editing state quantitative method and application thereof

    CN114141304A

  • Tumor neoantigen characteristic analysis and immunogenicity prediction tool and application thereof

    CN114446389A

  • Method for screening personalized tumor neoantigen by fusing multiple omics data and application thereof

    CN118866076A

  • Data-driven immune checkpoint blockade therapy response prediction

    WO2024181928A1

Cited By

  • Immune checkpoint blocking response prediction method and system based on diffusion model

    CN120413084A