Neoadjuvant immunotherapy response prediction method based on neoantigen quality
By constructing a prediction model based on neoantigen quality and combining deep learning and machine learning techniques, the accuracy and applicability issues of existing ICB response prediction methods have been addressed, achieving more efficient ICB treatment response prediction and improving the accuracy and coverage of personalized treatment.
Patent Information
- Application Number
- CN202510145037.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Existing methods for predicting immune checkpoint blockade responses have limited applicability and low accuracy, failing to fully reflect the natural characteristics of neoantigen immunogenicity, leading to inaccurate predictions of ICB treatment responses.
By acquiring clinical response data and whole-exome data of tumor tissue from subjects receiving immune checkpoint blockade therapy, multimodal features, including HLA evolutionary divergence and immunogenic neoantigen load, were obtained. A prediction model based on neoantigen quality was constructed, and features were extracted using deep learning or machine learning. The prediction threshold was optimized by combining the Youden index to improve prediction accuracy.
It significantly improves the accuracy and applicability of ICB treatment response prediction, provides a scientific basis for personalized immunotherapy decisions, and enhances the ability to quantitatively analyze the number of high-quality neoantigens.
Smart Images

Figure CN120072035B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computational biology and bioinformatics. BACKGROUND
[0002] In recent years, immune checkpoint blockade (ICB) therapy has revolutionized the treatment of patients with advanced cancer. These inhibitors, including antibodies targeting programmed death-1 (anti-PD-1, such as nivolumab and pembrolizumab) or cytotoxic T-lymphocyte-associated antigen 4 (anti-CTLA-4, such as ipilimumab), can reverse T regulatory cell-mediated immunosuppression. However, durable clinical benefits are limited to a small subset of patients. Therefore, researchers are looking for predictive biomarkers of ICB treatment response, such as tumor mutation burden (TMB), T cell infiltration level, PD-L1 immunohistochemistry, somatic genomic features, and immune phenotype. The generation of tumor neoantigens that evade central tolerance by T cells in the immune system is crucial for anti-tumor immunity. TMB can partially reflect the situation of patients generating neoantigens due to genomic coding mutations. In general, biomarkers that predict ICB response are useful, but have limitations.
[0003] Currently, many ICB response prediction methods based on neoantigen immunogenicity are developed. However, the neoantigen immunogenicity of patients calculated by these methods cannot fully reflect the natural characteristics due to HLA binding affinity or epitope similarity arrangement, cannot directly explain the potential mechanism of anti-tumor response, resulting in low accuracy and limited scope of ICB response prediction. It is of great value to consider multi-modal methods for determining the effectiveness of ICB treatment, and the quality of tumor-specific neoantigens mainly depends on their natural properties, such as HLA presentation, T cell cross-reactivity, and T cell recognition. Therefore, a comprehensive functional model is needed to meet the patient stratification of ICB treatment. SUMMARY
[0004] The purpose of the present application is to solve the problem of small scope and low accuracy of existing immune checkpoint blockade (ICB) response prediction methods; the present application provides an immune checkpoint blockade response prediction method based on neoantigen quality.
[0005] The immune checkpoint blockade response prediction method based on neoantigen quality, the method comprising the following steps:
[0006] S1, obtaining clinical response data and tumor tissue whole exome data of a subject treated by immune checkpoint blockade; the subject is a tumor patient, and the clinical response data is a responder or a non-responder;
[0007] S2, preprocessing the tumor tissue whole exome data:
[0008] The tumor tissue whole-exon data of each subject is processed to obtain multi-modal features of each subject; according to the multi-modal features of each subject, the HLA evolutionary divergence and immunogenic neoantigen load of the subject are obtained and used as input data of the training sample, and the clinical response data of the subject obtained is used as the response label of the training sample to obtain a training data set;
[0009] The multi-modal features of each subject include wild peptides and mutant peptides of neoantigens at tumor mutation point positions, HLA binding affinities of wild peptides and mutant peptides of neoantigens, and HLA allelic typing;
[0010] The immunogenic neoantigen load of the subject is the number of high-quality immunogenic neoantigens in the body of the subject, wherein when the quality Q of the neoantigen in the body of the subject is greater than 1, the high-quality immunogenic neoantigen is defined;
[0011] S3, the immune checkpoint blockade response prediction model under each candidate prediction threshold in the threshold interval predicts the clinical response data of the subject according to the input data of each training sample in the training data set, and determines the model prediction effect under the candidate prediction threshold in combination with the response label of the training sample;
[0012] According to the model prediction effect under all candidate prediction thresholds, the best prediction threshold N of the model is determined Cutoff ;
[0013] S4, the tumor tissue whole-exon data of the current subject to be tested is preprocessed to obtain the HLA evolutionary divergence and immunogenic neoantigen load of the current subject to be tested;
[0014] The immune checkpoint blockade response prediction model under the best prediction threshold N Cutoff predicts the clinical response data of the current subject to be tested according to the HLA evolutionary divergence and immunogenic neoantigen load of the current subject to be tested.
[0015] Preferably, the implementation mode of obtaining the multi-modal features of each subject by processing the tumor tissue whole-exon data of each subject in step S2 comprises:
[0016] The quality of the tumor tissue whole-exon data is checked, low-quality fragments are filtered, and qualified sequencing data is obtained;
[0017] The qualified sequencing data is compared to human reference genome data, and repeated sequences are deleted to obtain detectable sequences; somatic single nucleotide variation identification is performed on the detectable sequences, tumor mutation point positions are screened out, and wild peptides and mutant peptides of neoantigens at each tumor mutation point position are obtained according to each tumor mutation point position;
[0018] According to the wild peptide and the mutant peptide of the neoantigen at each tumor mutation point position, the HLA binding affinity of each wild peptide and mutant peptide of the neoantigen is determined.
[0019] According to the whole exon data of the tumor tissue of the subject, the HLA allele typing is determined.
[0020] Preferably, the implementation of step S2 of determining the HLA evolutionary divergence and the immunogenic neoantigen load of the subject according to the multi-modal features of each subject comprises:
[0021] According to the HLA allele typing in the multi-modal features of each subject, the HLA evolutionary divergence is determined using the Grantham distance algorithm;
[0022] According to the HLA binding affinity of each wild peptide and mutant peptide of the neoantigen in the multi-modal features of each subject, the HLA presentation A of the neoantigen is determined;
[0023] According to the wild peptide and the mutant peptide of each neoantigen in the multi-modal features of each subject, the T cell cross-reactivity C of the neoantigen is calculated;
[0024] According to the mutant peptide and the HLA allele typing of each neoantigen in the multi-modal features of each subject, the immunogenic score R of each neoantigen is obtained;
[0025] According to A, C and R, the quality Q of each neoantigen is obtained, Q=R×(A+C);
[0026] According to the value of the quality Q of each neoantigen, the number of high-quality immunogenic neoantigens in the subject is counted and taken as the immunogenic neoantigen load of the subject.
[0027] Preferably, And are the HLA binding affinities of the wild peptide and the mutant peptide of the neoantigen.
[0028] Preferably, wherein,
[0029] d i is the amino acid substitution position scaling factor when the i th amino acid in the wild peptide p WT is mutated to the mutant peptide p MT .
[0030] is the value at the row and column of the position-dependent amino acid substitution matrix M .
[0031] is the wild peptide p WT the i-th amino acid in the wild peptide, the i-th amino acid in the mutant peptide p MT the i-th amino acid in the wild peptide or the mutant peptide, i = 1, 2, …, N, N is the length of the wild peptide or the mutant peptide.
[0032] Preferably, according to the mutant peptide of each neoantigen and the HLA allele typing in the multi-modal features of each subject, obtain the immunogenicity score R of each neoantigen, and the specific implementation manner includes:
[0033] According to the mutant peptide of the neoantigen, obtain the neoantigen sequence;
[0034] According to the HLA allele typing, obtain the HLA pseudo sequence;
[0035] The encoding mode of the neoantigen sequence and the HLA pseudo sequence is converted into an encoding matrix, the features of the neoantigen and the HLA in the encoding matrix are extracted by a feature encoder based on deep learning or machine learning, and the joint features are obtained by splicing the extracted features;
[0036] The joint features are decoded, and finally the immunogenicity score R of the neoantigen is obtained.
[0037] Preferably, the immune checkpoint blockade response prediction model includes an immune neoantigen quality score calculation module and a threshold comparison module;
[0038] The immune neoantigen quality score calculation module calculates the immune neoantigen quality score ICBNQ of the subject according to the HLA evolutionary divergence and the immunogenicity neoantigen load of the subject; ICBNQ = HED x lg(INB + 1);
[0039] Wherein, HED is the HLA evolutionary divergence, and INB is the immunogenicity neoantigen load;
[0040] The threshold comparison module is used to determine the quality score ICBNQ, when the quality score ICBNQ is greater than or equal to the best prediction threshold N Cutoff , it is determined that the subject is a responder, and it is output as the predicted clinical response data; otherwise, it is determined that the subject is a non-responder, and it is output as the predicted clinical response data.
[0041] Preferably, in step S3, the implementation manner of determining the model prediction effect under each candidate prediction threshold is:
[0042] The Youden index J of the prediction model is taken as the prediction effect, wherein,
[0043] J = (sensitivity + specificity - 1);
[0044] sensitivity = TP / (TP + FN) x 100%;
[0045] specificity = TN / (TN+FP) x 100%;
[0046] sensitivity is the sensitivity of the prediction model, specificity is the specificity of the prediction model, TP is the total number of true positives, TN is the total number of true negatives, FP is the total number of false positives, and FN is the total number of false negatives;
[0047] true positive is that the prediction model predicts that the clinical response data is a responder, and the response label is a responder;
[0048] true negative is that the prediction model predicts that the clinical response data is a non-responder, and the response label is a non-responder;
[0049] false positive is that the prediction model predicts that the clinical response data is a responder, and the response label is a non-responder;
[0050] false negative is that the prediction model predicts that the clinical response data is a non-responder, and the response label is a responder.
[0051] Preferably, in step S3, the best prediction threshold N of the model is determined according to the model prediction effect under all candidate prediction thresholds Cutoff The implementation mode is:
[0052] The candidate prediction threshold corresponding to the maximum value of the model prediction effect is taken as the best prediction threshold N of the model Cutoff .
[0053] The immune checkpoint blockade response prediction device based on neoantigen quality comprises a storage device, a processor, and a computer program stored in the storage device and executable on the processor, and the processor executes the computer program to realize the immune checkpoint blockade response prediction method based on neoantigen quality.
[0054] The present application has the following beneficial effects:
[0055] The present application provides an immune checkpoint blockade response prediction method based on neoantigen quality, which can accurately quantify the number of high-quality neoantigens in the body of a subject, and quantitatively analyze the number of high-quality neoantigens in the body of a subject. The present application is beneficial to improve the accuracy of immune checkpoint blockade response prediction. Significantly improve the accuracy and scope of immune checkpoint blockade therapy response prediction, and provide scientific basis for individualized decision-making of immunotherapy. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 is a flowchart of the immune checkpoint blockade response prediction method based on neoantigen quality;
[0057] Figure 2 a schematic diagram for obtaining HLA evolutionary divergence and immunogenic neoantigen load;
[0058] Figure 3 a schematic diagram for obtaining neoantigen immunogenicity score;
[0059] Figure 4 a position-dependent amino acid substitution matrix provided by the present application;
[0060] Figure 5 a performance diagram of the immune checkpoint blockade response prediction model of the present application on test data; wherein, Figure 5 a performance diagram for predicting subjects with non-small cell lung cancer, Figure 5 b a performance diagram for predicting subjects with melanoma, Figure 5 c a performance diagram for predicting subjects with unknown cancer. DETAILED DESCRIPTION
[0061] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.
[0062] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0063] DETAILED DESCRIPTION Figure 1 and Figure 2 The present embodiment provides an immune checkpoint blockade response prediction method based on neoantigen quality, which comprises the following steps:
[0064] S1, obtaining clinical response data and tumor tissue whole exome data of a subject treated by immune checkpoint blockade; the subject is a tumor patient, the clinical response data is a responder or a non-responder, and the responder is a cured patient and the non-responder is an uncured patient;
[0065] S2, preprocessing the tumor tissue whole exome data:
[0066] The tumor tissue whole exome data of each subject is processed to obtain the multi-modal features of each subject; the HLA evolutionary divergence and immunogenic neoantigen load of the subject are obtained according to the multi-modal features of each subject, and are used as input data of the training sample, at the same time, the obtained clinical response data of the subject are used as response label of the training sample, to obtain a training data set;
[0067] The multi-modal features of each subject include wild and mutant peptides of neoantigens at tumor mutation point positions, HLA binding affinities of wild and mutant peptides of neoantigens, and HLA allelic typing;
[0068] The immunogenic neoantigen load of the subject is the number of high-quality immunogenic neoantigens in the body of the subject, wherein when the quality Q of the neoantigens in the body of the subject is greater than 1, the high-quality immunogenic neoantigens are defined;
[0069] S3, the immune checkpoint blockade response prediction model under each candidate prediction threshold in the threshold interval predicts the clinical response data of the subject according to the input data of each training sample in the training data set, and determines the model prediction effect under the candidate prediction threshold in combination with the response label of the training sample;
[0070] According to the model prediction effect under all candidate prediction thresholds, the best prediction threshold N of the model is determined Cutoff ;
[0071] S4, the tumor tissue whole exome data of the current subject to be tested is preprocessed to obtain the HLA evolutionary divergence and the immunogenic neoantigen load of the current subject to be tested;
[0072] The immune checkpoint blockade response prediction model under the best prediction threshold N Cutoff predicts the clinical response data of the current subject to be tested according to the HLA evolutionary divergence and the immunogenic neoantigen load of the current subject to be tested.
[0073] In the embodiment, the tumor tissue whole exome data, i.e., the DNA data (fasta file) of the subject, is combined with the clinical data of the subject to obtain the training data set, which improves the perfection and matching degree of important information in individualized immune checkpoint blockade (ICB) treatment response prediction, so that the prediction accuracy is higher and more targeted.
[0074] By constructing a multi-modal model framework based on antigen quality, multi-modal features are obtained from the tumor tissue whole exome data of each subject, deep immunogenic information in massive antigen data is captured, and the ICB treatment prediction task can share the same model framework under different cancer backgrounds;
[0075] By selecting the best prediction threshold N Cutoff , the sensitivity of the model to individual tumor patients can be increased, thereby improving the accuracy of ICB treatment response prediction.
[0076] In specific applications, the clinical response data and the tumor tissue whole exome data can include, but are not limited to: the data corresponding to 919 subjects are used to construct a training data set. These samples are collected from a plurality of references in the prior art, and specifically relate to 6 cancer types, specifically: skin melanoma, the number of training samples n = 359; non-small cell lung cancer, the number of training samples n = 519; head and neck squamous cell carcinoma, the number of training samples n = 12; bladder cancer, the number of training samples n = 27; anal cancer, the number of training samples n = 1; sarcoma, the number of training samples n = 1.
[0077] Further, referring to Figure 3 , the implementation of the step S2 of processing the tumor tissue whole exome data of each subject to obtain the multi-modal features of each subject includes:
[0078] S211, checking the quality of the tumor tissue whole exome data (fasta file), filtering low-quality fragments to obtain qualified sequencing data;
[0079] Specifically, the quality of the raw data is checked using a tool (such as FastQC) for the fasta file, which can specifically include the length of the sequence, the GC content, the pollution condition, etc. Sequences with a quality score lower than Q30 are filtered to obtain qualified sequencing data;
[0080] S212, comparing the qualified sequencing data to human reference genome data and deleting repetitive sequences to obtain detectable sequences;
[0081] For the qualified sequencing data, the original fastq file is aligned to the reference human genome GRCh38 using Burrows-Wheeler Aligner (BWA, http: / / bio-bwa.sourceforge.net) version 0.7.17, and the repetitive sequences are deleted using the Picard (v 2.26.6) tool to obtain detectable sequences;
[0082] S213, identifying somatic single nucleotide variations for the detectable sequences, screening out tumor mutation point positions, and obtaining wild peptides and mutant peptides of neoantigens at each tumor mutation point position according to the tumor mutation point positions;
[0083] The screening of the tumor mutation point positions can be achieved by the prior art, and specifically, the MuTect2 tool can be used for somatic single nucleotide variation identification. A candidate tumor mutation will only be selected under the following conditions:
[0084] 1) detected more than 5 times in high-quality reads;
[0085] 2) the variant allele frequency (VAF) is greater than 1%.
[0086] 3) not observed in multiple single nucleotide polymorphism databases (e.g., 1000 Genomes Project with version phase 3, dbSNP version 137, and COSMIC database);
[0087] Specifically, the wild peptide and the mutant peptide of each tumor mutation point position of the neoantigen can be obtained by using the GATK4 pipeline in the Terra cloud platform (https: / / app.terra.bio / ).
[0088] S214, according to the wild peptide and the mutant peptide of each tumor mutation point position of the neoantigen, the HLA binding affinity of each wild peptide and mutant peptide of the neoantigen is determined, and the tool used can be NetMHCpan-4.1;
[0089] S215, further according to the whole exon data of the tumor tissue of the subject, the HLA allele typing is determined, and the tool that can be used is Optitype.
[0090] Referring to Figure 2 , further, the implementation manner of obtaining the HLA evolutionary divergence and the immunogenic neoantigen load of the subject according to the multi-modal features of each subject in step S2 comprises:
[0091] According to the HLA allele typing in the multi-modal features of each subject, the HLA evolutionary divergence is determined by using the Grantham distance algorithm;
[0092] According to the HLA binding affinity of each wild peptide and mutant peptide of the neoantigen in the multi-modal features of each subject, the HLA presentation of the neoantigen is determined wherein, and are the HLA binding affinities of the wild peptide and the mutant peptide of the neoantigen, respectively;
[0093] According to each wild peptide and mutant peptide of the neoantigen in the multi-modal features of each subject, the T cell cross-reactivity C of the neoantigen is calculated; specifically,
[0094] wherein,
[0095] d i is the amino acid substitution position scaling factor when the i th amino acid in the wild peptide p WT of the neoantigen is mutated into the mutant peptide p MT ;
[0096] is the position-dependent amino acid substitution matrix M in the value at the row, the column;
[0097] the i-th amino acid in the wild-type peptide p WT the i-th amino acid in the wild-type peptide p the i-th amino acid in the mutant peptide p MT the i-th amino acid in the wild-type peptide p, i = 1, 2, …, N, N is the length of the wild-type peptide or the mutant peptide.
[0098] According to the mutant peptide of each neoantigen and the HLA allelic typing, the immunogenicity score R of each neoantigen is obtained;
[0099] According to A, C and R, the quality Q = R x (A + C) of each neoantigen is obtained;
[0100] wherein R is the immunogenicity score of the neoantigen, which is used to represent the ability of activated T cells to initiate immunogenicity, A is the HLA presentation, which is used to represent the ability of the neoantigen to be presented to T cells, and C is the T cell cross-reactivity, which is used to represent the ability of the T cells to recognize the mutant peptide and the wild-type peptide of the neoantigen.
[0101] According to the value of the quality Q of each neoantigen, the number of high-quality immunogenic neoantigens in the subject is counted, and the number is taken as the immunogenic neoantigen load of the subject.
[0102] The preferred embodiment gives a specific implementation method for obtaining the HLA evolutionary divergence and the immunogenic neoantigen load of the subject, which has the advantage of improving the accuracy of immune checkpoint blockade prediction and meeting patient stratification by capturing individualized HLA evolutionary divergence and immunogenic neoantigen load. The specific form of the position-dependent amino acid substitution matrix M is also given in the specific application, see Figure 4 .
[0103] See Figure 3 , it is further given that according to the wild-type peptide and the HLA allelic typing of each neoantigen in the multi-modal features of each subject, the specific implementation method for obtaining the immunogenicity score R of each neoantigen includes:
[0104] According to the mutant peptide of the neoantigen, a neoantigen sequence is obtained;
[0105] According to the HLA allelic typing, an HLA pseudo-sequence is obtained;
[0106] The neoantigen sequence and the HLA pseudo-sequence are converted into an encoding matrix through encoding methods such as Blosum62 or One-hot, the features of the neoantigen and the HLA in the encoding matrix are extracted through a feature encoder based on deep learning or machine learning (such as a convolutional neural network, a Transformer encoder, a long short-term memory artificial neural network, etc.), and the extracted features are spliced to obtain joint features;
[0107] The joint features are decoded, for example, using a fully connected network to obtain a neoantigen immunogenicity score R.
[0108] The specific process for obtaining the neoantigen immunogenicity score R given in the preferred embodiment improves the reliability of the neoantigen immunogenicity score by using a deep learning or machine learning model.
[0109] Further, the specific composition of the immune checkpoint blockade response prediction model is given, which includes an immune neoantigen quality score calculation module and a threshold comparison module.
[0110] The immune neoantigen quality score calculation module calculates the immune neoantigen quality score ICBNQ of the subject according to the HLA evolutionary divergence and the immunogenic neoantigen load of the subject.
[0111] ICBNQ = HED x lg(INB + 1);
[0112] Wherein, HED is the HLA evolutionary divergence, and INB is the immunogenic neoantigen load.
[0113] The threshold comparison module is used to determine the quality score ICBNQ. When the quality score ICBNQ is greater than or equal to the optimal prediction threshold N Cutoff , it is determined that the subject responds to ICB treatment, and the subject is a responder, and it is output as the predicted clinical response data; otherwise, it is determined that the subject does not respond to ICB treatment, and the subject is a non-responder, and it is output as the predicted clinical response data.
[0114] Further, in step S3, the implementation of determining the model prediction effect under each candidate prediction threshold is:
[0115] The Youden index J of the prediction model is used as the prediction effect, wherein,
[0116] J = (sensitivity + specificity - 1);
[0117] sensitivity = TP / (TP + FN) x 100%;
[0118] specificity = TN / (TN + FP) x 100%;
[0119] sensitivity is the sensitivity of the prediction model, specificity is the specificity of the prediction model, TP is the total number of true positives, TN is the total number of true negatives, FP is the total number of false positives, and FN is the total number of false negatives.
[0120] True positive is that the prediction model predicts the clinical response data as a responder, and the response label is a responder;
[0121] True negative is that the prediction model predicts the clinical response data as a non-responder, and the response label is a non-responder;
[0122] False positive is that the prediction model predicts the clinical response data as a responder, and the response label is a non-responder;
[0123] False negative is that the prediction model predicts the clinical response data as a non-responder, and the response label is a responder.
[0124] Further, in step S3, the best prediction threshold N of the model is determined according to the model prediction effect under all candidate prediction thresholds Cutoff The implementation manner is:
[0125] The candidate prediction threshold corresponding to the maximum value of the model prediction effect is taken as the best prediction threshold N of the model Cutoff .
[0126] Specifically, the new antigen quality-based immune checkpoint blockade response prediction device in the embodiment comprises a storage device, a processor, and a computer program stored in the storage device and executable on the processor, and the processor executes the computer program to realize the new antigen quality-based immune checkpoint blockade response prediction method in the embodiment one.
[0127] Verification test:
[0128] The effectiveness of the present application is proved by the following verification test, specifically:
[0129] The prediction performance on the test set by using the new antigen quality-based immune checkpoint blockade response prediction method provided by the present application is shown in Figure 5 .
[0130] Figure 5 The receiver operating characteristic curve (ROC) of the test sample is shown, where the vertical coordinate is the true positive rate (TPR), and the horizontal coordinate is the false positive rate (FPR). The area (AUC) under the receiver operating characteristic curve in the figure is an evaluation index of the prediction performance.
[0131] Figure 5 In the embodiment, the AUC result of predicting the subjects with non-small cell lung cancer is 0.85, and the 95% consistency interval is (0.69, 0.96), see Figure 5 a; the AUC result of predicting the subjects with melanoma is 0.86, and the 95% consistency interval is (0.71, 0.97), see Figure 5b; the AUC result of prediction for unknown cancer is 0.75, and the 95% consistency interval is (0.58, 0.89), see Figure 5 c; it can be seen that the method has high accuracy, and is also suitable for unknown cancer.
[0132] While the application has been described with reference to particular embodiments, it will be understood that the examples are for illustration only and that the principles and applications of the present application can be employed in other arrangements, modifications and designs without departing from the spirit and scope of the present application as defined in the appended claims. It will be understood that the features described with reference to individual embodiments can be used in other described embodiments.
Claims
1. A method for predicting immune checkpoint blockade response based on neoantigen quality, characterized in that, The method includes the following steps: S1. Obtain clinical response data and whole-exome data of tumor tissue from subjects undergoing immune checkpoint blockade therapy; the subjects are cancer patients, and the clinical response data are from responders or non-responders; S2. Preprocessing of whole-exome data from tumor tissue: The whole exome data of tumor tissue of each subject were processed to obtain the multimodal characteristics of each subject; based on the multimodal characteristics of each subject, the HLA evolutionary divergence and immunogenic neoantigen load of the subject were obtained and used as the input data of the training sample. At the same time, the clinical response data of the subject obtained was used as the response label of the training sample to obtain the training dataset. The multimodal characteristics of each subject include wild-type and mutant peptides of neoantigens at the tumor mutation site, HLA binding affinity of wild-type and mutant peptides of neoantigens, and HLA allele typing. The subject's immunogenic neoantigen load is the number of high-quality immunogenic neoantigens in the subject's body. When the mass of neoantigens in the subject's body Q>1, it is defined as a high-quality immunogenic neoantigen. S3. Within the threshold range, the immune checkpoint blockade response prediction model predicts the clinical response data of the subject based on the input data of each training sample in the training dataset, and determines the model prediction effect under the candidate prediction threshold by combining the response label of the training sample. Based on the model's prediction performance at all candidate prediction thresholds, determine the optimal prediction threshold N for the model. Cutoff ; S4. Preprocess the whole exon data of the tumor tissue of the current subject to obtain the HLA evolutionary divergence and immunogenic neoantigen load of the current subject; Optimal prediction threshold N Cutoff An immune checkpoint blockade response prediction model is used to predict the clinical response data of the current subject based on the HLA evolutionary divergence and immunogenic neoantigen load.
2. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 1, characterized in that, Step S2 involves processing the whole-exome data of tumor tissue from each subject to obtain the multimodal characteristics of each subject. The methods for achieving this include: The quality of whole exome sequencing data of tumor tissue is checked, low-quality fragments are filtered out, and qualified sequencing data is obtained. The qualified sequencing data was compared with the human reference genome data, and repetitive sequences were removed to obtain detectable sequences. Somatic single nucleotide variants were identified from the detectable sequences to screen for tumor mutation sites. Based on the location of each tumor mutation site, the wild-type peptide and mutant peptide of the neoantigen at each tumor mutation site were obtained. Based on the wild-type and mutant peptides of the neoantigens at the locations of each tumor mutation site, the HLA binding affinity of each neoantigen wild-type peptide and mutant peptide was determined. HLA allele typing was also determined based on the whole exon data of the subjects' tumor tissue.
3. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 1, characterized in that, Step S2, based on the multimodal characteristics of each subject, obtains the subject's HLA evolutionary divergence and immunogenic neoantigen load through the following methods: Based on the HLA allele typing in the multimodal characteristics of each subject, the Grantham distance algorithm was used to determine the HLA evolutionary divergence. Based on the HLA binding affinity of each neoantigen wild-type peptide to the mutant peptide in the multimodal characteristics of each subject, the HLA presentation A of the neoantigen was determined. Based on the wild-type and mutant peptides of each neoantigen in the multimodal characteristics of each subject, the T-cell cross-reactivity C of the neoantigen is calculated. Based on the mutant peptides and HLA allele typing of each neoantigen in the multimodal characteristics of each subject, the immunogenicity score R of each neoantigen was obtained; Based on A, C, and R, the mass of each new antigen is Q = R × (A + C); Based on the Q value of each neoantigen, the number of high-quality immunogenic neoantigens in the subject's body is counted and used as the subject's immunogenic neoantigen load.
4. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 3, characterized in that, and The HLA binding affinity of the new antigen wild-type peptide and the mutant peptide are respectively.
5. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 3, characterized in that, in, d i wild-type peptide p, a neoantigen WT The i-th amino acid is mutated into the mutant peptide p. MT Scaling factor for amino acid substitution positions; For position-dependent amino acid substitution matrix M The row The value at the column; wild peptide p WT The i-th amino acid, mutant peptide p MT The i-th amino acid in the peptide, i = 1, 2, ..., N, where N is the length of the wild-type peptide or the mutant peptide.
6. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 3, characterized in that, Based on the mutant peptides and HLA allele genotypes of each neoantigen in the multimodal characteristics of each subject, the immunogenicity score R of each neoantigen is obtained. Specific implementation methods include: The neoantigen sequence was obtained based on the mutant peptide of the neoantigen; the HLA pseudo-sequence was obtained based on HLA allele typing. The neoantigen sequence and HLA pseudo-sequence encoding methods are converted into encoding matrices. Features of the neoantigen and HLA in the encoding matrix are extracted by a feature encoder based on deep learning or machine learning, and the extracted features are concatenated to obtain joint features. The combined features are decoded to obtain the neoantigen immunogenicity score R.
7. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 1, characterized in that, The immune checkpoint blockade response prediction model includes an immune neoantigen quality score calculation module and a threshold comparison module; The immune neoantigen quality score calculation module calculates the subject's immune neoantigen quality score ICBNQ based on the subject's HLA evolutionary divergence and immunogenic neoantigen load; ICBNQ = HED × lg(INB + 1); Wherein, HED represents HLA evolutionary divergence, and INB represents the neoantigen load of immunogenicity; The threshold comparison module is used to determine the quality score ICBNQ. When the quality score ICBNQ is greater than or equal to the optimal prediction threshold N, the threshold is determined. Cutoff If the subject is deemed a responder, this is output as the predicted clinical response data; otherwise, the subject is deemed a non-responder, and this is output as the predicted clinical response data.
8. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 1, characterized in that, In step S3, the method for determining the model prediction performance under each candidate prediction threshold is as follows: The Youden index J of the prediction model is used as the prediction performance, where... J=(sensitivity+specificity-1); sensitivity=TP / (TP+FN)×100%; specificity=TN / (TN+FP)×100%; Sensitivity refers to the sensitivity of the prediction model, specificity refers to the specificity of the prediction model, TP is the total number of true positives, TN is the total number of true negatives, FP is the total number of false positives, and FN is the total number of false negatives. A true positive result is when the predictive model predicts that the clinical response data is from a responder and the response label is also from a responder. A true negative result is when the predictive model predicts that the clinical response data is non-responder and the response label is non-responder. A false positive occurs when the predictive model predicts that the clinical response data is from a responder, but the response label is from a non-responder. A false negative is when the predictive model predicts that the clinical response data is non-responder but the response label is responder.
9. The method for predicting immune checkpoint blockade response based on neoantigen quality according to claim 1, characterized in that, In step S3, the optimal prediction threshold N of the model is determined based on the model prediction performance under all candidate prediction thresholds. Cutoff The implementation method is as follows: The candidate prediction threshold corresponding to the maximum value of the model's prediction performance is taken as the optimal prediction threshold N of the model. Cutoff .
10. An immune checkpoint blockade response prediction device based on neoantigen quality, comprising a storage device, a processor, and a computer program stored in the storage device and executable on the processor, characterized in that, The processor executes a computer program to implement the immune checkpoint blockade response prediction method based on neoantigen quality as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Tumor sample immune editing state quantitative method and application thereof
CN114141304A
Tumor neoantigen characteristic analysis and immunogenicity prediction tool and application thereof
CN114446389A