A screening method and screening device for neoantigens
By combining cross-attention modules and multi-head attention mechanisms with star-shaped topological feature fusion, and using feedforward neural networks to predict the binding affinity of peptides to HLA and TCR, the problem of low accuracy in neoantigen screening in existing technologies is solved, and highly reliable neoantigen screening is achieved, which is particularly suitable for patients with advanced colorectal cancer.
Patent Information
- Application Number
- CN202511043003.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-07-28
AI Technical Summary
Existing technologies struggle to accurately predict the binding affinity of peptide sequences to HLA and TCR, especially when dealing with novel sequences not present in the training data. This results in low prediction accuracy and negatively impacts the screening performance of neotumor antigens.
Digital processing was performed using a cross-attention module and a multi-head attention mechanism, combined with star-shaped topological feature fusion. Feedforward neural networks were used to predict the binding affinity of peptides to HLA and TCR. Considering the annotation of variant sites and differential gene expression, highly reliable new antigens were screened.
It improves the accuracy and reliability of neoantigen screening, and is particularly suitable for patients with advanced colorectal cancer of low microsatellite instability, enhancing the efficacy of immunotherapy against specific tumors.
Smart Images

Figure CN120954512B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biomedical technology, specifically to a method and apparatus for screening novel antigens. Background Technology
[0002] The identification of tumor neoantigens is a major challenge in advancing immunotherapy. Gene mutations produced by tumor cells are not expressed in healthy tissues and are highly immunogenic, but the process from gene mutation in tumor cells to recognition by mature T cell lineages involves complex biological processes. Even when mutated genes produce abnormal proteins, only partially hydrolyzed peptides can be presented to the cell surface and recognized by immune cells. Therefore, identifying the correct and effective tumor neoantigens is a pressing issue. While the mutated genomes of cancer patients are highly specific, each type of tumor possesses representative gene mutation characteristics that distinguish it from other tumors.
[0003] With the advent and application of next-generation sequencing (NGS) and artificial intelligence (AI) technologies, identifying patient-specific neoantigens for personalized vaccines has become possible. Neoantigens are primarily tumor-specific short peptides (epitopes) generated by somatic mutations in tumors. These antigens first bind to human leukocyte antigens (HLA) and are presented on the surface of tumor cells as peptide-HLA complexes (pMHC), which can then be recognized by T-cell receptors (TCRs) to trigger an anti-tumor immune response. Neoantigens exhibit strong immunogenicity because they are not present in normal tissues and are not affected by T-cell selection or host central tolerance. Therefore, neoantigens represent a valuable source of targets for T-cell-based cancer immunotherapy. Among the vast number of mutant peptides, only a limited fraction are likely to elicit a robust anti-tumor immune response; therefore, accurate identification of immunogenic neoantigens is crucial.
[0004] With the outstanding performance of deep learning in many fields such as natural language processing and computer vision, various deep learning frameworks have been used to predict peptide binding to TCRs, such as convolutional neural networks (CNNs), long short-term memory neural networks (LSTMs), Transformers, and prediction methods using pre-trained data as embeddings. While these models show good results, their performance degrades significantly when generalized to novel sequences not present in the training data, resulting in lower prediction accuracy. Therefore, accurately predicting the binding of previously unseen peptides to TCRs remains a significant challenge. Summary of the Invention
[0005] This application addresses the aforementioned technical problems in the existing technology. The purpose of this application is to provide a method and apparatus for screening neoantigens, capable of predicting the binding affinity of polypeptide sequences to HLA and TCR. It considers both the binding affinity of the polypeptide sequence to HLA and the binding affinity of TCR to the HLA-polypeptide complex. Combined with variant site annotation and differential gene expression, the screening achieves high accuracy and reliability, and can effectively screen for neoantigens targeting specific tumors.
[0006] According to the first aspect of this application, a method for screening neoantigens is provided. The method, via a processor, includes: acquiring a peptide sequence and an HLA pseudo-sequence corresponding to an HLA typing based on a target object; digitizing the peptide sequence and the HLA pseudo-sequence, and inputting the digitized results into a cross-attention module, calculating attention scores using a multi-head attention mechanism to obtain peptide sequence interaction features and HLA sequence interaction features respectively; using the HLA sequence interaction features as the central node and the peptide sequence interaction features as multiple satellite nodes, performing feature fusion using a star topology, and performing classification prediction based on the feature fusion results to screen out peptides and HLA complexes pMHC whose prediction results exceed a threshold; acquiring the target object antigen-specific TCR sequence, and digitizing the TCR sequence and pMHC respectively, inputting the digitized results into the cross-attention module to obtain fused interaction features after deep interaction between the TCR sequence features and the pMHC features; and predicting the affinity between the TCR sequence and pMHC based on the fused interaction features using a feedforward neural network.
[0007] According to a second aspect of this application, a method for screening neoantigens in patients with advanced colorectal cancer of low microsatellite instability is provided. The screening method, via a processor, includes: acquiring a peptide sequence and an HLA pseudo-sequence corresponding to an HLA subtype obtained based on a target individual, wherein the target individual is a patient with advanced colorectal cancer and a molecular subtype of low microsatellite instability; digitizing the peptide sequence and the HLA pseudo-sequence, and inputting the digitized results into a cross-attention module, calculating attention scores using a multi-head attention mechanism to obtain peptide sequence interaction features and HLA sequence interaction features respectively; and using the HLA pseudo-sequence... Sequence interaction features are used as central nodes, and peptide sequence interaction features are used as multiple satellite nodes. Feature fusion is performed using a star topology, and classification prediction is performed based on the feature fusion results to screen out peptides and HLA complex pMHC whose prediction results exceed a threshold. The target antigen-specific TCR sequence is obtained, and the TCR sequence and pMHC are digitized separately. The digitized results are then input into a cross-attention module to obtain the fused interaction features after deep interaction between the TCR sequence features and the pMHC features. Based on the fused interaction features, a feedforward neural network is used to predict the affinity between the TCR sequence and pMHC.
[0008] According to a third aspect of this application, a neoantigen screening device is provided, the screening device including a processor configured to perform steps of implementing the neoantigen screening method described in various embodiments of this application.
[0009] According to the fourth aspect of this application, a computer-readable storage medium is provided, on which a computer program or instructions are stored, which, when executed by a processor, implement the neoantigen screening method described in various embodiments of this application.
[0010] Compared with the prior art, the beneficial effects of the embodiments of this application are as follows:
[0011] The neoantigen screening method provided in this application involves digitizing peptide sequences and HLA pseudo-sequences, and inputting the digitized results into a cross-attention module. A multi-head attention mechanism is used to calculate attention scores, yielding peptide sequence interaction features and HLA sequence interaction features. Using the HLA sequence interaction features as the central node and the peptide sequence interaction features as multiple satellite nodes, a star topology is employed for feature fusion. Classification prediction is then performed based on the feature fusion results to screen peptides with HLA complex pMHC whose prediction results exceed a threshold. The TCR sequence is obtained, and both the TCR sequence and pMHC are digitized. The digitized results are then input into the cross-attention module to obtain fused interaction features resulting from deep interaction between the TCR sequence features and the pMHC features. Based on these fused interaction features, a feedforward neural network is used to predict the affinity between the TCR sequence and pMHC. Thus, by comprehensively considering the binding force between peptide sequences and HLA pseudo-sequences, as well as the binding force between TCR and peptide-HLA complex pMHC, and combining variant site annotation and differential gene expression, the screened neoantigens exhibit higher reliability and are suitable for specific tumor populations.
[0012] Furthermore, the screening method provided in this application is particularly applicable to patients with advanced low-microsatellite instability colorectal cancer, and is used for screening and applying neoantigens in this population.
[0013] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above description and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0014] In drawings that are not necessarily drawn to scale, the same reference numerals may describe similar parts in different views. Similar reference numerals with different letter suffixes may indicate different examples of similar components. The drawings generally illustrate various embodiments by way of example rather than limitation, and are used together with the specification and claims to illustrate the disclosed embodiments. Such embodiments are illustrative and exemplary, and are not intended to be exhaustive or exclusive embodiments of the method, apparatus, system, or non-transitory computer-readable medium having instructions for implementing the method.
[0015] Figure 1 A flowchart of a method for screening neoantigens according to an embodiment of this application is shown.
[0016] Figure 2 A schematic diagram of the structure of a novel antigen screening method according to an embodiment of this application is shown.
[0017] Figure 3 A schematic diagram showing the AUC comparison between the screening method for neoantigens according to embodiments of this application and the model in related literature (CN117976074B) is illustrated. Detailed Implementation
[0018] To enable those skilled in the art to better understand the technical solutions of this application, the application will be described in detail below with reference to the accompanying drawings and specific embodiments. The embodiments of this application will be further described in detail below with reference to the accompanying drawings and specific examples, but these are not intended to limit the scope of this application.
[0019] The terms "first," "second," and similar words used in this application do not indicate any order, quantity, or importance, but are merely used for distinction. The terms "including" or "comprising," etc., used in this application mean that the element preceding the word encompasses the elements listed after the word, and do not exclude the possibility of encompassing other elements. In this application, the arrows shown in the figures for each step are merely examples of the execution order, not limitations. The technical solution of this application is not limited to the execution order described in the embodiments. The steps in the execution order can be combined, broken down, or rearranged, as long as the logical relationship of the executed content is not affected.
[0020] All terms used in this application (including technical or scientific terms) have the same meaning as understood by one of ordinary skill in the art to which this application pertains, unless otherwise specifically defined. It should also be understood that terms defined in general dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and not as idealized or highly formalized, unless expressly defined herein. Technologies and equipment known to one of ordinary skill in the art may not be discussed in detail, but where appropriate, such technologies and equipment should be considered part of the specification.
[0021] Figure 1 A flowchart of a neoantigen screening method according to an embodiment of this application is shown. The neoantigen screening method is executed by a processor as follows: steps S101-S105. The arrows shown in the figure for each step are merely examples of the execution order and are not limitations. The technical solution of this application is not limited to the execution order described in the embodiment. The steps in the execution order can be combined, decomposed, or rearranged, as long as the logical relationship of the execution content is not affected.
[0022] The processor can be a processing device that includes one or more general-purpose processing devices, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor can be a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor can also be one or more special-purpose processing devices, such as an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), a Digital Signal Processor (DSP), a System-on-Chip (SoC), etc. That is to say, the screening methods described in the various embodiments of this application can all be integrated into a smart chip or a System-on-Chip (SoC).
[0023] In step S101, the peptide sequence obtained based on the target object and the HLA pseudo-sequence corresponding to the HLA typing are obtained.
[0024] Specifically, the target group may be cancer patients, especially patients with advanced colorectal cancer and a molecular subtype of low microsatellite instability.
[0025] Relevant tumor research data and results are recorded in multiple databases, including The Cancer Genome Atlas (TCGA), the Gene Expression Omnibus (GEO), ArrayExpress, the Fudan Data Portal for Cancer Genomics (FUCC), and clinical sequencing data from hospitals. Tumor research data for the target population can be obtained by consulting relevant literature and these databases.
[0026] Specifically, the download links for sequencing data can be obtained and downloaded using sra-tools (v3.1.1) and mwget (v0.2.0). Then, the raw tumor research data can be converted to FASTQ format using parallel-fastq-dump (v0.6.7). The tumor research data in FASTQG format can be quality-controlled and filtered using fastp (v0.23.4). Next, data alignment can be performed using BWA (v0.7.18), aligning the filtered valid data to the reference genome. Finally, GATK (v4.6.0.0) software can be used to filter, correct, detect variants, and perform quality control and deduplication on the alignment results.
[0027] VEP (v112.0) software can be used to annotate and convert the format of variant sites. Next, Python scripts can be used to statistically analyze and filter the annotated sites for variant detection. After a series of filtering steps, the true mutation events are obtained. Then, the mutation results are annotated to obtain pathogenicity information, and differentially expressed genes are detected. For example, STAR (v2.7.11b) software can be used for data alignment, followed by stringtie (v2.2.3) software to calculate expression levels, and the DESeq2 (v1.44.0) R package for differential expression analysis. Target genes are screened using Python scripts. The Bio module can be used for analysis, extracting corresponding amino acid sequences from the transcript sequence database based on site coordinates and amino acid changes, and modifying the variant results. All possibilities containing the variant amino acid sequence can be enumerated at lengths of 8-11 amino acids to generate a polypeptide sequence of a specified length after mutation.
[0028] This embodiment is for illustrative purposes only and does not constitute a limitation on any specific solution.
[0029] In some embodiments, HLA types with high population frequency in the target population can be screened based on tumor research data and databases, and their corresponding pseudo-sequences can be generated.
[0030] In some embodiments, for patients with advanced colorectal cancer and a molecular subtype of low microsatellite instability, the HLA subtype includes HLA-A01:01, HLA-A02:01, HLA-A02:06, HLA-A02:07, HLA-A03:01, HLA-A11:01, HLA-A23:01, HLA-A24:02, HLA-A26:01, HLA-A29:02, HLA-A30:01, HLA-A31:01, HLA-A32:01, HLA-A33:03, HLA-A68:01, HLA-B07:02, HLA-B08:01, HLA-B13:02, HLA-B15:01, HLA-B 18:01, HLA-B27:05, HLA-B35:01, HLA-B40:01, HLA-B44:02, HLA-B44:03, HLA -B46:01, HLA-B51:01, HLA-B57:01, HLA-B58:01, HLA-C01:02, HLA-C02:02, HL A-C03:02, HLA-C03:03, HLA-C03:04, HLA-C04:01, HLA-C05:01, HLA-C06:02, H LA-C07:01, HLA-C07:02, HLA-C08:01, HLA-C08:02, HLA-C12:03 and HLA-C16:01.
[0031] In step S102, the polypeptide sequence and HLA pseudo-sequence are digitized, and the digitized results are input into the cross-attention module. Attention scores are calculated using a multi-head attention mechanism to obtain the polypeptide sequence interaction features and HLA sequence interaction features, respectively.
[0032] In some embodiments, the digitization process includes digitizing the polypeptide sequence and HLA pseudo-sequence based on an amino acid lexicon.
[0033] Specifically, an amino acid terminology dictionary is a mapping tool used to convert amino acid sequences into computer-processable numerical or symbolic representations. For example, Python libraries have built-in mapping tables for single-letter / three-letter abbreviations of amino acids, and command-line tools can also be used to support sequence format conversion and terminology encoding. Furthermore, amino acid terminology dictionaries can be obtained based on physicochemical property databases; for instance, the AAindex database contains over 500 physicochemical property parameters for amino acids, which can be used to construct multidimensional feature vectors.
[0034] like Figure 2It can digitally encode peptide sequences and HLA pseudo-sequences, and digitally map the sequences according to an amino acid lexicon dictionary to obtain digitally processed sequences. Further operations such as embedding can be used to convert the integer codes into high-dimensional vector feature representations, which are then input into the cross-attention module.
[0035] In other embodiments, the peptide sequence and HLA pseudo-sequence are digitally encoded, amino acids are represented by integers, and negative samples are generated to improve the model's generalization ability.
[0036] This is merely an illustrative example and does not constitute a specific limitation on digital processing, nor does it exclude other methods that can achieve digital processing.
[0037] In some embodiments of this application, the binding affinity of peptide-HLA-TCR is predicted based on the Star-Transformer framework.
[0038] In this embodiment, a cross-attention module is used to achieve information interaction between different sequences. During the cross-attention calculation, a multi-head attention mechanism is employed to decompose the attention calculation into multiple "heads," capturing the interaction patterns of different subspaces in parallel. The outputs of all heads are then concatenated and projected back to the original dimension through a linear layer. The multi-head attention mechanism integrates multiple dimensions of peptide-HLA interactions, capturing sequence interaction features from multiple dimensions. The cross-attention module enables bidirectional information flow between peptides and HLA pseudo-sequences, simulating the biomolecular docking process.
[0039] For example, the cross-attention module calculates the attention weights between the two based on a multi-head attention mechanism, capturing the interaction patterns between the peptide and the HLA pseudo-sequence. Therefore, the peptide sequence interaction features and the HLA sequence interaction features integrate the association information of the two. For instance, the peptide sequence interaction features include both the peptide's own sequence information and its association information with the HLA pseudo-sequence. The HLA sequence interaction features include both the HLA pseudo-sequence's own sequence information and its association information with the peptide sequence.
[0040] like Figure 2 The peptide sequence interaction features and HLA sequence interaction features will then be input into subsequent modules such as star topology and fully connected layers.
[0041] In some embodiments, a sparse connection strategy is used in cross-attention computation to retain connections with high attention weights, thereby reducing computational complexity.
[0042] Specifically, after calculating the attention weights, only connections with high weights are retained. For example, a threshold can be set so that connections with weights below the threshold are directly disconnected. Connections with high weights are likely to correspond to key sites where peptides and HLA pseudo-sequences interact, such as anchor residues or binding groove sites.
[0043] Furthermore, by retaining connections with high attention weights, only those connections with high weights can be computed, thereby reducing computational load, lowering computational complexity, and improving computational efficiency.
[0044] In some embodiments, the sparse connection strategy includes focusing only on amino acid sites that are biologically likely to form key interactions when calculating the attention between a polypeptide sequence and an HLA pseudo-sequence, or using a sparse attention mechanism to limit the scope of attention calculation.
[0045] Specifically, when calculating the attention between a peptide sequence and an HLA pseudo-sequence, focusing only on amino acid sites that may form key interactions biologically is a strategy based on prior biological knowledge to optimize attention calculation. This effectively reduces computational complexity while preserving key information.
[0046] For example, based on prior knowledge, the anchoring residue positions and types corresponding to different HLA pseudo-sequence alleles can be determined. When calculating attention, only the amino acids at these anchoring residue positions on the peptide sequence and the amino acid sites in the HLA pseudo-sequence that interact with these anchoring residues are considered. For the screened key amino acid sites, the attention weights between them are calculated to determine the importance and interaction strength of these key amino acid sites in the binding process. For other non-key sites on the peptide sequence and HLA pseudo-sequence, the attention between them is not calculated, thereby reducing the computational load, improving computational efficiency, and allowing the prediction model to focus more on capturing information that plays a key role in binding, thus improving the accuracy of the prediction model in predicting the affinity between peptides and HLA.
[0047] Sparse attention mechanisms can limit the computational scope in two ways: one is by pre-setting a sparse pattern, which forcibly limits the computational scope of attention based on prior knowledge or algorithm design; the other is a dynamic selection mechanism, such as the model automatically learning which positions have more important interactions and retaining only high-weight connections. For example, during model training, the model can automatically identify amino acid sites that contribute significantly to HLA-peptide binding and only calculate the attention for these sites.
[0048] This is provided as an example only and does not constitute a limitation on any specific solution.
[0049] In step S103, the HLA sequence interaction features are used as the central node and the polypeptide sequence interaction features are used as multiple satellite nodes. Feature fusion is performed using a star topology, and classification prediction is performed based on the feature fusion results to screen out polypeptide and HLA complex pMHC whose prediction results exceed the threshold.
[0050] Specifically, using HLA sequence interaction features as the central node, interaction features of all peptide sequences are collected, and attention weights are calculated. The central node then transmits the aggregated information back to each satellite node, updating the peptide feature representation. By changing the connection structure between nodes from fully connected to a star topology, the model's complexity is reduced, focusing on multiple amino acids in the peptide and HLA pseudo-sequences, thereby enhancing the ability to identify key amino acid positions with stronger binding forces. Utilizing a star topology for feature fusion effectively captures local relationships and long-term dependencies. Attention weights between the central node and each satellite node are calculated, and the node state is updated to capture the interactions between amino acids.
[0051] like Figure 2 The feature fusion results are input into a fully connected layer. The fully connected layer and a classifier then perform classification and prediction to determine the binding affinity between the peptide and the HLA pseudo-sequence, such as whether a stable pMHC complex can be formed. Furthermore, high-confidence pMHC sequences are selected. For example, a threshold of 0.7 is set; if the prediction result is greater than 0.7, a stable pMHC complex is considered to be formed.
[0052] Specifically, during the training of the prediction model, the cross-entropy loss function can be used to evaluate the predictive ability of the model. The Adam optimizer can be used for parameter updates, and the learning rate can be dynamically adjusted to improve convergence speed. Dropout technology is used to prevent overfitting, and early stopping is used to monitor the performance on the validation set. After training, the performance of the prediction model is evaluated using metrics such as accuracy, recall, F1 score, and AUC value until the optimal prediction results are obtained, at which point the optimal parameters of the prediction model are determined.
[0053] Returning to the embodiments of this application, in step S104, the target antigen-specific TCR sequence is obtained, and the TCR sequence and pMHC are digitized respectively. The results of the digitization are then input into the cross-attention module to obtain the fusion interaction feature after deep interaction between the TCR sequence feature and the pMHC feature.
[0054] Specifically, CDR3 sequences of TCR structures with high frequency among the target population can be screened based on content such as the VDJdb database. This is only an example and does not constitute a limitation on specific solutions.
[0055] In this embodiment, the binding affinity between the peptide and the HLA pseudo-sequence is further predicted based on the predicted binding affinity between the TCR and pMHC. Specifically, the TCR sequence and pMHC are digitized separately. The digitization process can be similar to that described above, or other digitization methods can be used, as long as the amino acids can be converted into integers that can be recognized and processed by a computer. Further details will not be elaborated here.
[0056] In some embodiments, such as Figure 2 After digitizing the TCR sequence, a sliding transformer is used to divide the TCR sequence into windows of a preset length, so that attention can be calculated independently for each window using the cross-attention module, which facilitates the identification of the combination relationship between TCR and pMHC.
[0057] For example, long TCR sequences can be segmented by sliding at preset lengths (e.g., 10 amino acids) to obtain multiple local windows (e.g., if the TCR sequence length is 30, the window length is 10, and the sliding step size is 5, then 4 windows are obtained). For each window, cross-attention is calculated separately with pMHC to focus on the interactions of local regions.
[0058] When TCR recognizes pMHC, it does not scan the entire sequence at once, but rather binds to antigen peptides in localized regions. The sliding window design mimics the local scanning process of the TCR receptor on the pMHC surface, which is beneficial for focusing on key regions.
[0059] By dynamically modeling the interactions between different amino acids using the sliding window technique, the prediction model can better capture local features and incorporate the principles of interaction from physics to enhance the understanding of the interactions between amino acid residues.
[0060] Specifically, such as Figure 2 The digitized results are input into a cross-attention module to simulate the immune recognition process of the pMHC-TCR triad. For example, the TCR sequence acts as the "query," and pMHC as both the "key" and "value," calculating the attention weights of the TCR sequence on different regions of the pMHC. Simultaneously, pMHC also pays inverse attention to key sites in the TCR sequence. Through this multi-head attention mechanism, multi-dimensional interaction information between pMHC and TCR is fused, including hydrophobic interactions, charge complementarity, and spatial conformation matching, to generate a fusion interaction feature of the pMHC-TCR triad. This fusion interaction feature includes at least information on the strength of pMHC-TCR binding and potential signals for immune activation, such as whether it triggers T cell antigen recognition or signal transduction.
[0061] After receiving the fused interactive feature vector, the feedforward neural network will further complete the prediction of immune response, such as whether T cells are activated and the strength of the immune response, thereby predicting the affinity between the TCR sequence and pMHC.
[0062] Thus, combinations of peptides and HLA complexes with scores exceeding a threshold are selected as binding complexes pMHC based on prediction results. Furthermore, by predicting the affinity between TCR and pMHC, the group of complexes with the highest prediction scores is selected as the most binding complexes and can be used as the final candidate neoantigens for cancer treatment.
[0063] This application embodiment trains multiple related tasks simultaneously to improve the generalization ability and prediction accuracy of the prediction model. Multiple self-attention layers are employed, with inter-layer connections achieved through residual connections and layer normalization. Each self-attention mechanism focuses on different parts of the input sequence, capturing long-distance dependencies. Cross-entropy loss is chosen as the loss function, and the prediction model parameters are updated using the backpropagation algorithm. Training is performed using the Adam optimizer, early stopping is employed to prevent overfitting, and the prediction performance is evaluated using metrics such as ROC curves and AUC values. Cross-validation is used to ensure the stability of the prediction model. By comprehensively predicting the binding affinity of the peptide to the HLA pseudo-sequence and the binding affinity of TCR and pMHC, a peptide-HLA-pMHC triad prediction model is constructed, which improves the accuracy of the final prediction results and yields effective neoantigens.
[0064] Based on the neoantigen screening method provided in the embodiments of this application, a peptide-HLA-pMHC triad prediction model is obtained. Using the trained peptide-HLA-pMHC triad prediction model, the prediction accuracy is greatly improved. For example... Figure 3 Using the same peptide and HLA sequences as samples, and based on the screening method provided in this application, the average AUC of the peptide-HLA-pMHC triplet prediction model is 0.955, while the average AUC of the Star-Transformer model in the literature (publication number CN117976074B) is 0.939. The AUC value of the screening method provided in this application is 1.6% higher than that in the literature. Therefore, the peptide-HLA-pMHC triplet prediction model provided in this application is superior to the model in the literature, and the screening method provided in this application significantly improves the accuracy of HLA molecule and peptide sequence affinity prediction.
[0065] In particular, for patients with advanced colorectal cancer and a low-grade microsatellite instability molecular subtype, specific HLA typing methods (HLA-A01:01, HLA-A02:01, HLA-A02:06, HLA-A02:07, HLA-A03:01, HLA-A11:01, HLA-A23:01, HLA-A24:02, HLA-A26:01, HL) are used. A-A29:02, HLA-A30:01, HLA-A31:01, HLA-A32:01, HLA-A33:03, HLA-A68:01, HLA-B07 :02, HLA-B08:01, HLA-B13:02, HLA-B15:01, HLA-B18:01, HLA-B27:05, HLA-B35:01, HL Based on the screening method provided in the embodiments of this application, new antigens for patients with advanced colorectal cancer and a molecular subtype of low microsatellite instability can be screened more accurately. A-B40:01, HLA-B44:02, HLA-B44:03, HLA-B46:01, HLA-B51:01, HLA-B57:01, HLA-B58:01, HLA-C01:02, HLA-C02:02, HLA-C03:02, HLA-C03:03, HLA-C03:04, HLA-C04:01, HLA-C05:01, HLA-C06:02, HLA-C07:01, HLA-C07:02, HLA-C08:01, HLA-C08:02, HLA-C12:03, and HLA-C16:01 can be screened.In other words, in some embodiments of this application, a method for screening neoantigens in patients with advanced colorectal cancer of low microsatellite instability is provided. This screening method, via a processor, includes: acquiring a peptide sequence and an HLA pseudo-sequence corresponding to an HLA subtype obtained based on a target individual, wherein the target individual is a patient with advanced colorectal cancer and a molecular subtype of low microsatellite instability; digitizing the peptide sequence and the HLA pseudo-sequence, and inputting the digitized results into a cross-attention module, calculating attention scores using a multi-head attention mechanism to obtain peptide sequence interaction features and HLA sequence interaction features, respectively. The method involves: using the HLA sequence interaction features as the central node and the peptide sequence interaction features as multiple satellite nodes, performing feature fusion using a star topology, and performing classification prediction based on the feature fusion results to screen out peptides and HLA complexes pMHC whose prediction results exceed a threshold; obtaining the TCR sequence, and digitizing the TCR sequence and pMHC respectively, and inputting the digitized results into a cross-attention module to obtain the fused interaction features after deep interaction between the TCR sequence features and the pMHC features; and using a feedforward neural network based on the fused interaction features to predict the affinity between the TCR sequence and pMHC.
[0066] The HLA typing includes HLA-A01:01, HLA-A02:01, HLA-A02:06, HLA-A02:07, HLA-A03:01, HLA-A11:01, HLA-A23:01, HLA-A24:02, HLA-A26:01, HLA-A29:02, H H LA-B35:01, HLA-B40:01, HLA-B44:02, HLA-B44:03, HLA-B46:01, HLA-B51:01, HLA-B57:01, HLA-B58:01, HLA-C01:02, HLA-C02:02, HLA-C03:02, HLA-C03:03, HLA-C03:04, HLA-C04:01, HLA-C05:01, HLA-C06:02, HLA-C07:01, HLA-C07:02, HLA-C08:01, HLA-C08:02, HLA-C12:03 and HLA-C16:01.
[0067] Based on this specific HLA typing, neoantigen peptides suitable for patients with advanced colorectal cancer and a low microsatellite instability molecular typing can be screened more accurately.
[0068] In some embodiments of this application, a neoantigen screening apparatus is provided, the screening apparatus including a processor configured to perform steps of implementing the neoantigen screening methods described in various embodiments of this application.
[0069] In other embodiments of this application, a SOC chip is provided, the SOC chip including a processor configured to perform steps of implementing the neoantigen screening method described in various embodiments of this application.
[0070] According to embodiments of this application, a computer-readable storage medium is also provided, on which a computer program or instructions are stored, which, when executed by a processor, implement the steps of the neoantigen screening method described in various embodiments of this application.
[0071] The aforementioned computer-readable storage media may be such as read-only memory (ROM), random access memory (RAM), phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), electrically erasable programmable read-only memory (EEPROM), other types of random access memory (RAM), flash drives or other forms of flash memory, cache, registers, static memory, optical disc read-only memory (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape cassette or other magnetic storage devices, or any other possible non-transitory medium used to store information or instructions that can be accessed by computer equipment.
[0072] According to embodiments of this application, a computer program product is also provided, the computer program product including computer instructions for causing a computer to perform steps in the neoantigen screening methods as described in various embodiments of this application.
[0073] This application describes various operations or functions that can be implemented as software code or instructions, or defined as software code or instructions. Such content can be directly executable source code or differential code (“incremental” or “patch” code) (“object” or “executable” form). The software code or instructions can be stored in a computer-readable storage medium and, when executed, can cause a machine to perform the described functions or operations, and include any mechanism for storing information in a machine-accessible form, such as recordable or non-recordable media (e.g., read-only memory (ROM), random access memory (RAM), disk storage media, optical storage media, flash memory devices, etc.).
[0074] Furthermore, although exemplary embodiments have been described herein, their scope includes any and all embodiments based on this application that have equivalent elements, modifications, omissions, combinations (e.g., schemes involving intersections of various embodiments), adaptations, or alterations. Elements in the claims will be interpreted broadly based on the language used in the claims and are not limited to the examples described in this specification or during the implementation of this application, which will be interpreted as non-exclusive. Therefore, this specification and examples are intended to be considered illustrative only, and the true scope and spirit are indicated by the full scope of the following claims and their equivalents.
[0075] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. Other embodiments may be used by those skilled in the art upon reading the above description. Furthermore, in the above detailed description, various features may be grouped together to simplify the application. This should not be construed as an intention that a disclosed feature not claimed is necessary for any claim. Rather, the subject matter of the application may be less than all the features of a particular disclosed embodiment. Thus, the claims are incorporated herein by reference as examples or embodiments, wherein each claim is an independent, separate embodiment, and these embodiments are contemplated as being combined with each other in various combinations or arrangements. The scope of this application should be determined by reference to the appended claims and the full scope of their equivalents.
[0076] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.
Claims
1. A method for screening neoantigens, characterized in that, The filtering method is transmitted via a processor and includes: Obtain the peptide sequence and HLA pseudo-sequence corresponding to HLA typing based on the target object; The peptide sequence and HLA pseudo-sequence are digitized, and the digitized results are input into the cross-attention module. Attention scores are calculated using a multi-head attention mechanism to obtain peptide sequence interaction features and HLA sequence interaction features, respectively. Using the HLA sequence interaction features as the central node and the polypeptide sequence interaction features as multiple satellite nodes, feature fusion is performed using a star topology, and classification prediction is performed based on the feature fusion results to screen out polypeptide and HLA complex pMHC whose prediction results exceed the threshold. The target antigen-specific TCR sequence is obtained, and the TCR sequence and pMHC are digitized respectively. The results of the digitization are then input into the cross-attention module to obtain the fusion interaction feature after the deep interaction between the TCR sequence feature and the pMHC feature. Based on the aforementioned fusion interaction features, the affinity between TCR sequences and pMHC is predicted using a feedforward neural network.
2. The screening method according to claim 1, characterized in that, The target population consists of patients with advanced colorectal cancer and a molecular subtype of low microsatellite instability.
3. The screening method according to claim 1 or 2, characterized in that, The HLA classification includes: HLA-A01:01, HLA-A02:01, HLA-A02:06, HLA-A02:07, HLA-A03:01, HLA-A11:01, HLA-A23:01, HLA-A24:02, HLA-A26:01, HLA-A29:02, HLA-A30:01, HLA-A31:01, HLA-A32:01, HLA-A33:03, HLA-A68:01, HLA-B07:02, HLA-B08:01, HLA-B13:02, HLA-B15:01, HLA-B18:01, HLA-B27:05, HLA-B35:01, HLA-B40:01, HLA-B44:02, HLA-B44:03, HLA-B46:01, HLA-B51:01, HLA-B57:01, HLA-B58:01, HLA-C01:02, HLA-C02:02, HLA-C03:02, HLA-C03:03, HLA-C03:04, HLA-C04:01, HLA-C05:01, HLA-C06:02, HLA-C07:01, HLA-C07:02, HLA-C08:01, HLA-C08:02, HLA-C12:03 and HLA-C16:
01.
4. The screening method according to claim 1, characterized in that, The digitization process includes digitizing the polypeptide sequence and HLA pseudo-sequence based on an amino acid morpheme dictionary.
5. The screening method according to claim 1, characterized in that, In cross-attention computation, a sparse connection strategy is used to retain connections with high attention weights.
6. The screening method according to claim 5, characterized in that, The sparse connection strategy includes focusing only on amino acid sites that are biologically likely to form key interactions when calculating the attention between a polypeptide sequence and an HLA pseudo-sequence, or using a sparse attention mechanism to limit the scope of attention calculation.
7. The screening method according to claim 1, characterized in that, After digitizing the TCR sequence, a sliding transformer is used to divide the TCR sequence into windows of a preset length, so that attention can be calculated independently for each window using the cross-attention module.
8. A method for screening neoantigens in a population of patients with advanced colorectal cancer of low microsatellite instability, characterized in that, The filtering method is transmitted via a processor and includes: Obtain the peptide sequence and HLA pseudo-sequence corresponding to the HLA subtype obtained based on the target object, wherein the target object is a patient with advanced colorectal cancer and a molecular subtype of low microsatellite instability; The peptide sequence and HLA pseudo-sequence are digitized, and the digitized results are input into the cross-attention module. Attention scores are calculated using a multi-head attention mechanism to obtain peptide sequence interaction features and HLA sequence interaction features, respectively. Using the HLA sequence interaction features as the central node and the polypeptide sequence interaction features as multiple satellite nodes, feature fusion is performed using a star topology, and classification prediction is performed based on the feature fusion results to screen out polypeptide and HLA complex pMHC whose prediction results exceed the threshold. The target antigen-specific TCR sequence is obtained, and the TCR sequence and pMHC are digitized respectively. The results of the digitization are then input into the cross-attention module to obtain the fusion interaction feature after the deep interaction between the TCR sequence feature and the pMHC feature. Based on the aforementioned fusion interaction features, the affinity between TCR sequences and pMHC is predicted using a feedforward neural network.
9. A screening device for a new antigen, characterized in that, The screening device includes a processor configured to perform steps of the method for screening neoantigens according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the method for screening neoantigens as described in any one of claims 1-7.
Citation Information
Patent Citations
MHC molecule and antigen epitope affinity determination method, model training method and device
CN117976074B
Method for screening tumor neoantigen based on HLA typing and structure
CN110675913A
Method and device for predicting affinity of TCR and antigen complex, equipment and medium
CN117877568A