A method, system, device and readable storage medium for screening of multi-source tumor antigens

By constructing a multi-source antigen integration screening framework, we can deeply explore microbial antigens, hidden antigens, and mutated neoantigens, which solves the problems of single antigen source coverage and platform fragmentation in existing technologies, and realizes the effective evaluation and screening of multi-target synergistic therapies.

CN122135771APending Publication Date: 2026-06-02HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
Filing Date
2026-02-26
Publication Date
2026-06-02

Smart Images

  • Figure CN122135771A_ABST
    Figure CN122135771A_ABST
Patent Text Reader

Abstract

This application discloses a method, system, device, and readable storage medium for screening multi-source tumor antigens. The screening method includes the following steps: acquiring high-throughput sequencing data; data processing to identify non-host and non-vector-derived sequences and host sequences, generating candidate peptides for microbial antigens, cryptogenic antigens, and mutant neoantigens, respectively; performing HLA binding / presentation prediction and immunogenicity prediction on the candidate peptides for microbial antigens, cryptogenic antigens, and mutant neoantigens, uniformly classifying and prioritizing them, and screening tumor antigens from the candidate antigen peptides based on a comprehensive score. This application constructs a multi-source antigen integrated screening framework, breaking through the limitation of single antigen source coverage, realizing in-depth mining of microbial antigens, cryptogenic antigens, and mutant neoantigens, filling the gap in high-specificity target screening; and realizing the synergistic effect assessment between multi-source antigen targets, providing a key tool for the development of "multivalent" and "multispecific" immunotherapy regimens.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the fields of bioinformatics and tumor immunology, specifically relating to a screening method for multi-source tumor antigens, as well as a screening system, computer equipment, and computer-readable storage medium constructed based on the screening method. Background Technology

[0002] The efficacy of tumor immunotherapy hinges on the precise identification of tumor-specific antigens. Ideal tumor antigens should possess high tumor specificity, widespread intratumoral expression, and good immunogenicity, thereby effectively killing tumor cells while avoiding off-target damage to normal tissues. Therefore, screening and identifying high-quality tumor antigens is a core foundation for developing cancer therapeutics and achieving precision immunotherapy.

[0003] Currently, existing technologies mainly rely on genome sequencing to find mutated neoantigens or targeted screening for specific types (such as microbial antigens or known cryptic antigens). However, these methods still have the following prominent drawbacks: (1) Limited antigen source coverage: Existing technologies mostly focus on single or partial categories of antigens, lacking an integrated framework; this makes the selected antigen library one-sided and difficult to find enough targets. (2) Incomplete microbial antigen and cryptic antigen mining: The coverage of microbial antigen and cryptic antigen sources is not comprehensive, and important high-specificity target sources are easily missed. (3) Fragmented technology platforms: Screening and validation platforms for different antigen types are independent of each other, and the processes and standards are not unified, which makes it impossible to efficiently and in parallel analyze and evaluate multi-source antigens, hindering the rational design of multi-target synergistic therapies.

[0004] Therefore, there is an urgent need in this field for a tumor antigen screening method that can systematically integrate antigens from multiple sources and fully cover the sources of hidden antigens. Summary of the Invention

[0005] In view of this, the primary objective of this application is to provide a screening method for multi-source tumor antigens. This screening method constructs an integrated screening framework for multi-source antigens, which breaks through the limitation of single antigen source coverage and enables in-depth mining of microbial antigens, cryptogenic antigens, and mutated neoantigens, filling the gap in the screening of high-specificity targets. Furthermore, by unifying the evaluation criteria for multi-source antigens, it breaks down the analytical barriers caused by platform fragmentation and realizes the evaluation of synergistic effects among multi-source antigen targets, thereby providing a key tool for the development of "multivalent" and "multispecific" immunotherapy regimens.

[0006] To achieve the above objectives, this application adopts the following technical solution: One aspect of this application discloses a screening method for multi-source tumor antigens. This screening method systematically integrates the screening of multi-source antigens, enabling in-depth mining of microbial antigens, cryptogenic antigens, and newly mutated antigens, and realizing the evaluation and screening of multi-source antigens within a unified framework.

[0007] The main steps of the screening method in this application are as follows: S1. Obtain high-throughput sequencing data of tumor tissue and its paired samples, wherein the high-throughput sequencing data includes WES or WGS data, and WTS or RNA-seq data.

[0008] In this application, "tumor tissue" refers to a collection of cells that has been pathologically confirmed and originates from a malignant growth or tumor lesion in the subject's body. This tissue is used as the sample to be analyzed in this application.

[0009] In this application, the term "tumor" specifically refers to a solid tumor, which is a non-hematologic malignancy resulting from malignantly transformed cells forming a localized space-occupying mass. It encompasses all malignant tumors originating from epithelial tissue, mesenchymal tissue, neuroectoderm, and other non-hematopoietic stem cell origins. Specifically, it includes, but is not limited to: lung cancer, breast cancer, gastric cancer, colorectal cancer, liver cancer, pancreatic cancer, kidney cancer, melanoma, osteosarcoma, and neuroblastoma. It should be understood that the solid tumors described in this application include resectable, locally advanced, and metastatic solid tumors, as well as primary solid tumors and their metastases. However, this term does not include leukemia, lymphoma, myeloma, or other fluid-filled or diffuse tumors.

[0010] In this application, the term "paired sample" includes at least one of adjacent tissue, normal tissue, or peripheral blood derived from the same subject.

[0011] In this application, the term "adjacent tissue" refers to solid tissue that is anatomically adjacent to tumor tissue but does not exhibit typical malignant cytological features upon pathological evaluation. In some examples, it typically refers to the non-tumor solid area within 1 to 2 cm of the tumor margin; however, in certain specific examples, this range may be adjusted according to organ type, tumor growth pattern, or the pathologist's judgment, including but not limited to 0.5 cm to 3 cm. It is understood that adjacent tissue serves as a baseline sample for comparison with tumor tissue to identify tumor-specific molecular variations.

[0012] In this application, the term "normal tissue" refers to tissue derived from the same subject and confirmed by pathological or clinical evaluation to be free of malignant tumor cells.

[0013] In this application, "WES" stands for Whole Exome Sequencing. It is a method that uses sequence capture or targeted enrichment strategies to specifically capture all protein-coding regions (exons) and flanking splice sites from a subject's genomic DNA and then performs high-throughput sequencing.

[0014] In this application, "WGS" stands for Whole Genome Sequencing. This technology refers to unbiased, target-free, full-coverage high-throughput sequencing of a subject's genomic DNA.

[0015] In this application, "WTS" stands for Whole Transcriptome Sequencing. This technology refers to a method for reverse transcription and high-throughput sequencing of total RNA or poly(A)-enriched mRNA extracted from a subject's sample to obtain the global gene expression profile and transcript structure information of the sample.

[0016] In this application, "RNA-seq" refers to RNA sequencing, which is a general term for methods that use high-throughput sequencing technology to reverse transcribe, construct libraries, and sequence RNA molecules in a sample. In some specific examples, RNA-seq uses a poly(A) capture strategy to enrich mature mRNA; in other specific examples, RNA-seq uses an rRNA removal strategy to retain non-coding RNA information.

[0017] It should be understood that the WES, WTS, WGS, RNA-seq, and other data described in this application can all be obtained based on methods known in the art, or can be downloaded from well-known biomedical public online databases.

[0018] S2. Process the data from step S1 to determine the host sequence and sequences that are not from the host or the vector.

[0019] This step involves data processing of the raw data to obtain the expected or standardized input results. It should be understood that the data processing methods differ slightly for the three different types of tumor antigens. Specifically, the steps are as follows: S21. Perform quality control on high-throughput sequencing data to obtain clean reads; S22. Align the clean reads to the universal human reference genome and extract the unaligned reads; S23. Align the unaligned reads from step S22 to the telomere-to-telomere human reference genome and extract the unaligned reads as non-host sequences. S24. Align the non-host sequences to the vector contamination reference database and retain the miss reads, which are sequences that are neither from the host nor from the vector. S25. Use the host alignment results and host reads to screen for candidate peptides of cryptogenic antigens and mutated neoantigens.

[0020] In this application, the term "data processing" refers to a series of computer-implemented operations, transformations, filtering, corrections, and analyses performed on raw sequencing data from biological samples. These operations can be performed in a manner well known in the art, and in some specific examples, the data processing includes, but is not limited to: base identification, read alignment to a reference genome, expression quantification, variant detection, copy number inference, and fusion gene identification.

[0021] In this application, the term "quality control" refers to a set of computational operations and statistical tests used to assess and ensure the reliability, accuracy, and completeness of sequencing data and its derived analytical results. This term applies to both the raw data level and the level of aligned data and analytical results.

[0022] In this application, "host sequence" refers to a nucleic acid sequence derived from the genome of the tested organism itself, as opposed to nucleic acid sequences introduced from exogenous sources or contamination sources. In this application, a two-step alignment method with dual references is used. First, the sequence is aligned to a universal human reference genome. Then, the mismatched reads are aligned to a telomere-to-telomere human reference genome to achieve more precise screening of non-host sequences.

[0023] The "universal human reference genome" refers to a digital linear genome sequence jointly constructed and continuously maintained by the international genomics community, serving as a standard reference for human genome sequence alignment and variant identification. This term does not include the personal genome of any specific subject, but rather reflects a standardized representation of sequence characteristics common to the human species. In some specific examples, the universal human reference genome includes, but is not limited to, GRCh37 / hg19, GRCh38 / hg38, and subsequent updated versions. The "telomere-to-telomere human reference genome" refers to a reference genome sequence obtained by using long-read sequencing technologies (including but not limited to PacBio HiFi, Oxford Nanopore) and complementary technologies such as optical mapping and strand sequencing to achieve a gap-free, complete assembly of each chromosome from one telomere to the other. The key technical feature distinguishing this term from the universal human reference genome is that it covers previously difficult-to-resolve highly homologous regions, repetitive fragment regions, centromere satellite regions, and ribosomal DNA arrays, and completes telomere sequence characterization at the base level for all chromosome ends. In some specific examples, the reference genome used is T2T-CHM13. By using a two-step host alignment, host sequence reads can be effectively removed.

[0024] In this application, "vector contamination sequence" refers to a nucleic acid sequence fragment derived from a genetically engineered vector, viral vector, transposon system, or any artificially constructed nucleic acid vector, which is detected as an unexpected exogenous signal in sequencing data. This term includes, but is not limited to: plasmid backbone sequences, multiple cloning site sequences, viral long terminal repeat sequences, lentiviral packaging signals, Cas9 coding sequences in CRISPR editing systems, or sgRNA backbone sequences. In some specific examples, data from non-vector sources is identified by alignment to a vector contamination database. Commonly used vector contamination databases include, but are not limited to: the UniVec database, EMBL vector contamination entries, the National Center for Biotechnology Information's vector contamination database, and user-built custom vector sequence libraries based on experimental protocols.

[0025] In this application, alignment can be performed using BWA-MEM, Bowtie2, Minimap2, or other human genome alignment tools capable of achieving high-precision alignment. A preprocessing strategy involving host dual-reference removal and vector contamination removal is employed to achieve source noise reduction of microbial antigens.

[0026] S3. Based on the sequence data from non-host and non-vector sources, generate candidate peptides for microbial antigens.

[0027] This step involves sequentially classifying microorganisms, tracing protein origins, generating candidate peptides, and constructing evidence chains from sequences that are neither from the host nor from the vector, thereby generating microbial antigen candidate peptides. Specifically, it includes the following steps: S31. Construct a microbial protein library, which includes a genomic reference library for microbial classification / quantification and a protein sequence alignment library for protein tracing. These tumor-related microbial databases can be constructed based on other publicly available high-quality microbial genome or protein databases such as NCBI, GTDB, and IMG, and the species range can be expanded or updated according to specific research needs.

[0028] S32. Perform microbial classification and quantification on sequences that are not from the host or vector, and obtain the abundance at the species / genus level; output the classification assignment information of each read at the read level, and retain the source field to achieve microbial classification and quantification at the nucleic acid level.

[0029] S33. Translate and align reads or their effective sequence fragments that have microbial classification assignments at the nucleic acid level to a microbial protein library, retaining reads whose microbial classifications are consistent at both the nucleic acid and protein levels.

[0030] S34. Threshold filtering and candidate peptide generation are performed on the reads hit in step S33 to obtain microbial antigen candidate peptides.

[0031] As a preferred example, the threshold filtering includes at least: sequence consistency, E-value, and alignment coverage. The specific selection and configuration can be tailored to the experimental objectives and research needs.

[0032] The candidate peptides can be generated using a sliding window method commonly used in the field. The peptide length range can be configured as needed to adapt to the predicted length requirements of HLA-I and HLA-II, and therefore there are no particular limitations.

[0033] S4. Based on WTS or RNA-seq data, construct a combined transcript set that simultaneously covers the first and second transcript sources, and generate cryptogenic candidate peptides from it.

[0034] In this application, the first transcript source is an annotated non-coding transcript source, and the second transcript source is a transcript source other than the annotation obtained from transcript reconstruction. Further, the first transcript source is obtained by extracting the transcript sequence range of annotated non-coding transcripts from a reference annotation obtained based on WTS or RNA-seq reads. The second transcript source is obtained by reconstructing transcripts from WTS or RNA-seq alignment results to obtain a sample-specific transcript set, and then screening for new transcript sequences not covered by the annotation.

[0035] It should be understood that WTS or RNA-seq data should first be aligned to the universal human reference genome to generate a standardized alignment file before further screening. This is a routine step and therefore not described in detail.

[0036] In this application, "reference annotation" refers to a collection of information systematically indexed in a computationally accessible digital file format, on a coordinate system of a general or specific reference genome sequence, detailing the location, structure, type, and identifiers of genomic functional elements. In some specific examples, the reference annotation is derived from public databases, including but not limited to: GENCODE, RefSeq, Ensembl, and UCSC Known Genes.

[0037] Specifically, the transcript sequence ranges of annotated non-coding transcripts are extracted from the reference annotations. Furthermore, transcript reconstruction is performed based on RNA-seq alignment results to obtain a sample-specific transcript set. This set is then compared with the aforementioned reference annotations to screen for new transcripts not covered by the annotations. These new transcripts include those occurring in intergenic regions, introns, or antisense strands.

[0038] The generation of cryptogenic antigens from a combined transcript set consisting of two transcript sources involves the following steps: ORF discovery is performed in the combined transcript set, sORFs are screened and translated to obtain short open reading frame (SEP) encoded peptides; the specific ORF length threshold can be selected as needed.

[0039] The coding nucleic acid sequence corresponding to the SEP is mapped and located using reference genome coordinates to obtain its genomic coordinates, and its coordinate range is annotated to improve the interpretability and verifiability of the source. The annotation is combined with genome annotation and usually includes intergenic regions, introns, antisenses, relative relationships with genes / transcripts, and whether they overlap with protein-coding CDS.

[0040] When the coding sequence of a SEP has multiple positional matches in the reference genome, hard constraint rules are applied, and tumor / control expression is quantified for SEPs that pass the constraints. Based on tumor expression, control expression, and differential expression thresholds, tumor-specific abnormally expressed SEPs are screened as candidate peptides for cryptogenic antigens. This can reduce noise caused by incomplete transcripts, splicing artifacts, or incorporation of coding region-related sequences.

[0041] In some specific examples, the hard constraint rule is that all localization points are non-coding and do not overlap with protein-coding CDS.

[0042] It should be understood that transcript assembly and identification of new transcripts can be performed using transcript assembly software well-known in the art, such as StringTie and Cufflinks, but are not limited to these. Denovo transcript assembly can be performed using tools such as Trinity and rnaSPAdes. Short open reading frame tools can be TransDecoder and ORFfinder. In this application, the coverage of the hidden antigen is improved and the denoising effect is enhanced through double coverage of the hidden antigen and non-coding hard constraints.

[0043] S5. Based on WES or WGS data, generate candidate peptides for mutated neoantigens through somatic cell mutation detection.

[0044] In this step, reliability is guaranteed through "multi-algorithm joint detection + cross-support + normalization processing".

[0045] S51. Based on WES or WGS data, run two or more somatic mutation detection algorithms to obtain a candidate somatic mutation set, and perform normalization processing on the mutation records; the normalization processing includes, but is not limited to, coordinate sorting, reference consistency, insertion / deletion left alignment, splitting multiple allotropic sites and unified representation, thereby reducing the impact of representation differences on the consistency of subsequent annotation and peptide generation.

[0046] S52. Retain mutations supported by at least two somatic mutation detection algorithms to form a high-confidence consensus mutation set.

[0047] S53. Perform functional annotation on the consensus mutations in the high-confidence consensus mutation set in step S52 to obtain protein change information, and generate candidate mutant peptides centered on the mutation sites as candidate mutant neoantigen peptides.

[0048] As a preferred example, a uniform transcript selection rule (e.g., prioritizing protein-coding transcripts that are most highly expressed in tumors and consistent with RNA evidence) is used during mutant peptide generation to reduce peptide inconsistencies caused by multiple transcript annotations.

[0049] There are no specific limitations or requirements for the specific somatic mutation algorithm or tool. Known algorithms in this field, such as GATK Mutect2, Strelka2, VarDict, Octopus, etc., can be used, as long as they can obtain a high-confidence somatic mutation set with controllable quality, sensitivity and specificity in tumor-normal paired samples.

[0050] S6. Perform HLA binding / presentation prediction and immunogenicity prediction on candidate peptides of microbial antigens, candidate peptides of cryptogenic antigens, and candidate peptides of mutant neoantigens, respectively. Integrate the HLA binding / presentation prediction results and immunogenicity scores with source-specific evidence factors to achieve unified grading and priority ranking. Based on the comprehensive score, screen out tumor antigens from the candidate antigen peptides.

[0051] In step S6, the HLA binding / presentation prediction is performed using two or more HLA binding / presentation prediction algorithms, and the output fields are standardized.

[0052] In this application, HLA typing and HLA binding / presentation prediction can both be performed using methods well known in the art. Specific examples include, but are not limited to, at least one of: BigMHC_EL, BigMHC_IM, DeepImmuno, MHCflurry, MHCflurryEL, MHCnuggetsI, MHCnuggetsII, Nnalign, NetMHC, NetMHCIIpan, NetMHCIIpanEL, NetMHCpan, NetMHCpanEL, PickPocket, SMM, or SMMPMBEC.

[0053] Furthermore, in step S6, immunogenicity prediction includes the following steps: a) Constructing a combined sequence: The combined sequence is composed of the amino acid sequence of the candidate peptide and its corresponding HLA amino acid sequence.

[0054] (b) A sequence modeling network is used to encode the joint sequence representation. Subsequently, based on the sequence encoding results and the global physicochemical characteristics of the peptide, an immunogenicity score is output. The sequence modeling network used is not particularly limited and can be selected as needed; in some specific examples, a bidirectional long short-term memory network (BiLSTM) is preferred. The global physicochemical characteristics of the peptide are obtained by introducing residue-level representations driven by physicochemical properties (e.g., dimensionality reduction embedding of the AAindex amino acid physicochemical property index library) and global physicochemical descriptors of the peptide (such as composition, polarity, charge, volume, hydrophobicity, etc.).

[0055] c) Train source-specific immunogenicity models for microbial antigens, cryptogenic antigens, and mutated neoantigens respectively, and predict immunogenicity.

[0056] In step S6, the grading features of the unified grading include at least: immunogenicity score threshold, multi-algorithm integrated binding / presentation summary index, and source-specific evidence. The priority sorting is based on a unified hierarchical system and a comprehensive sorting rule is constructed. Preferred, the source-specific evidence includes: For microbial candidates, at least the following should be considered: species abundance, number of supporting reads, nucleic acid / protein consistency evidence chain, and threshold hit rate. For hidden antigen candidates, at least tumor / control expression, differential expression, non-coding constraints by markers and SEP block annotations should be included; For mutation candidates, at least multi-algorithm consensus support markers, mutation function annotations, tumor expression support, and mutation / wildtype difference indicators should be included.

[0057] In this application, training data are constructed and targeted training / calibration is performed for microbial antigens, cryptogenic antigens and mutant neoantigens respectively, in order to adapt to the statistical differences in sequence composition and physicochemical properties of candidates from different sources and reduce the generalization error of cross-source applications.

[0058] Furthermore, this application constructs a cross-source pairing set for three-source candidates within the same framework, performs sequence similarity calculation and threshold screening, and uses the same fields required for outputting mimicry results, thereby supporting vaccine combination optimization, shared target screening, and the discovery of cross-reaction mechanisms.

[0059] This application enables horizontal comparison of candidate antigen targets from different sources at the same scale through a unified data interface, feature quantification standards, and efficacy evaluation system. This not only significantly reduces the time and computing power costs caused by platform switching and standard adaptation, but also provides quantifiable and iterative technical support for the rational design of multi-target synergistic therapies, such as multi-epitope vaccines and combination T-cell therapies.

[0060] Building upon the aforementioned integrated framework and unified standards, this invention further enables the evaluation of synergistic effects among multi-source antigen targets. By employing cross-source mimicry analysis to predict and rank the overall efficacy of multi-target combinations, it fills a gap in the rational design of multi-target combinations in existing technologies, providing a key tool for the development of next-generation "multivalent" and "multispecific" immunotherapies. Another aspect of this application discloses a screening system for multi-source tumor antigens, wherein when the screening system is in operation, the screening method described in this application is implemented; The screening system includes: The data input module is configured to receive high-throughput sequencing data of tumor tissue and its paired samples; The data preprocessing and sequence identification module is configured to perform quality control and alignment on high-throughput sequencing data, and to identify host sequences and sequences from non-host and non-vector sources. A microbial antigen candidate peptide generation module is configured to generate microbial antigen candidate peptides based on the non-host and non-vector-derived sequence data. The cryptogenic antigen candidate peptide generation module generates cryptogenic antigen candidate peptides based on WTS or RNA-seq data. The mutant neoantigen candidate peptide generation module generates mutant neoantigen candidate peptides based on WES or WGS data and through somatic cell mutation detection. The HLA typing module is configured to use HLA typing tools to obtain sets of HLA class I and II alleles and output an HLA list for subsequent prediction. The HLA binding / presentation prediction module is configured to predict HLA binding / presentation using two or more HLA binding / presentation prediction algorithms; An immunogenicity prediction module is configured to predict the immunogenicity of candidate antigens from three different sources. The unified hierarchical and priority sorting module is configured to establish a unified hierarchical system and construct comprehensive sorting rules.

[0061] In a preferred embodiment, the screening system further includes a cross-three-source molecular mimicry analysis and unified output module, which is configured to construct cross-source pairings for the three-source candidate set, perform similarity calculation and mimicry screening, and provide unified output and statistical summary.

[0062] Another aspect of this application discloses a computer device including a memory, a processor, and a computer program or instructions stored in the memory, wherein when the computer program or instructions are executed by the processor, the screening method for multi-source tumor antigens described in any one of this application is implemented.

[0063] Another aspect of this application discloses a computer-readable storage medium storing a computer program or instructions thereon, characterized in that, when the computer program or instructions are executed by a processor, they implement the multi-source tumor antigen screening method described in this application.

[0064] The beneficial effects of this application are: (1) Construct a multi-source antigen integration screening framework to overcome the limitation of single antigen source coverage. This application is the first to construct an integrated screening system for multi-source antigens, encompassing genetic microbial antigens, latent antigens, and mutated neoantigens. By extracting and fusing characteristics from heterogeneous antigen sources, this application enables parallel screening of broad-category antigen targets, significantly expanding the breadth and diversity of the candidate antigen library, avoiding target omissions due to incomplete source coverage, and providing a complete data foundation for subsequent selection of high-potential targets.

[0065] (2) To achieve in-depth mining of multi-source antigens and fill the gap in screening for highly specific targets. This application achieves in-depth mining of microbial antigens, cryptogenic antigens, and mutated neoantigens, overcoming the bottlenecks of traditional methods in terms of antigen recognition immunogenicity and insufficient antigen quantity. This application can effectively capture antigenic epitopes that are difficult to cover using traditional sequencing strategies, greatly improving the probability of discovering highly specific, low-cross-reactivity targets, and solving the technical problem of long-term missed or inefficient screening of this type of antigen. Attached Figure Description

[0066] Figure 1 This is a flowchart illustrating the multi-source tumor antigen screening method of this application. Detailed Implementation

[0067] The present application will be further illustrated below with reference to specific embodiments. It should be noted that the specific embodiments below are for illustrative purposes only and do not limit the scope of the present application in any way.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein in the specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The processes described in the specification, claims, and drawings of this application include multiple steps appearing in a specific order; however, it should be understood that these steps may not be performed in the order they appear in this application or may be performed in parallel. Step numbers such as S1, S2, etc., are merely used to distinguish different steps and do not themselves represent any execution order. Furthermore, these processes may include more or fewer steps, and these steps may be performed sequentially or in parallel. It should be noted that descriptions such as "first," "second," etc., in this application are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0069] In addition, unless otherwise specified, methods that do not specify conditions or steps are conventional methods.

[0070] Example 1: Screening method for multi-source tumor antigens This embodiment uses a microsatellite stable (MSS) colorectal cancer (CRC) patient sample as an example to illustrate in detail the screening method for multi-source tumor antigens in this application. However, it should be understood that the screening method described in this application is applicable to single patient samples or batch analysis of multiple patient samples.

[0071] The clinical samples were obtained from Hefei Cancer Hospital, Chinese Academy of Sciences, and included tumor tissue and its paired adjacent normal tissue. Tissue samples were obtained using routine clinical methods, rinsed with physiological saline, flash-frozen in dry ice, and then transferred to… All samples were stored at 80 °C until further processing, and all samples were subject to ethical review and informed consent from the patients.

[0072] The following uses a single patient sample as a specific example to illustrate the steps: 1.1 Acquisition of Raw Data To obtain whole exome sequencing (WES) data of colorectal cancer tumor tissue and its paired adjacent normal tissue, as well as whole transcriptome sequencing (WTS) data of tumor tissue and its paired adjacent normal tissue.

[0073] 1.2 Data Processing The raw data is processed to form standardized inputs for subsequent screening steps. This data processing includes steps such as quality control, sequence alignment, and identification.

[0074] 1.2.1 Quality Control (1) Perform connector removal, low-quality base shearing (in this embodiment, low-quality bases are defined as quality values ​​< 20) and shortest length filtering (in this embodiment, reads with a length less than 35 bp are filtered out) on the raw WES and WTS reads to obtain clean reads.

[0075] (2) Output quality control statistics (number of reads, base quality distribution (Q20 / Q30), read length distribution, connector contamination ratio, etc.) and record the processing log.

[0076] It should be noted that during this process, QC reports are generated for both WES and WTS, and these reports are included in the sample audit package.

[0077] 1.2.2 Host Dual Reference Removal (1) Align the clean reads obtained from WTS data in section 1.2.1 to the human reference genome GRCh38 and extract the unaligned reads.

[0078] (2) The unaligned reads in step (1) are further aligned to the telomere-to-telomere human reference genome T2T-CHM13, and only the reads that are not effectively aligned in both steps are retained as non-host sequences.

[0079] It should be noted that the output includes the change in the number of reads, the alignment rate, the non-alignment rate, and key alignment parameters for each step, in order to facilitate auditing and review.

[0080] 1.2.3 Carrier contamination removal (1) Align the non-host sequences in section 1.2.2 to the NCBI UniVec vector contamination reference library, remove the hit reads, and keep the miss reads.

[0081] (2) The unsuccessful reads retained in step (1) are identified as reads that are not from the host and are not from the vector, and are used to screen for subsequent microbial antigen candidate peptides.

[0082] (3) The candidate peptides of hidden antigens and mutant neoantigens are screened using the host alignment results and host reads.

[0083] It should be noted that during this process, the output carrier hit rate, main carrier type statistics, and filtering logs are included.

[0084] 1.3 Generation of candidate peptides for microbial antigens Based on the non-host and non-vector-derived reads obtained in section 1.2, microbial classification, protein tracing, candidate peptide generation, and evidence chain construction were carried out sequentially to complete the generation of microbial antigen candidate peptides.

[0085] 1.3.1 Construction of Microbial Reference Library (1) Construct a list of tumor-related microorganisms and establish species names, NCBI classification numbers (tax IDs) and classification hierarchy information.

[0086] (2) Download the corresponding high-quality reference genome and protein sequences, and construct: a) a genome reference library for microbial classification / quantification; and b) a protein sequence alignment library for protein tracing (including a protein accession and tax ID mapping table).

[0087] It should be noted that during this process, a version number, data source timestamp, and deduplication rules are recorded for each reference library to ensure that the results are repeatable.

[0088] 1.3.2 Microbial classification and quantification at the nucleic acid level (1) Input reads from non-host and non-vector sources into the microbial classification and quantification workflow, and calculate the standardized relative abundance of different taxonomic levels such as species and genus based on the alignment or classification assignment results of the reads in the microbial reference database. The standardized relative abundance is the relative proportion of the contribution value of the reads corresponding to each microbial taxonomic unit after normalization by sequencing depth; when a read is assigned to multiple microbial taxonomic units, its contribution value is weighted according to the number of assignments and / or alignment confidence of the read, and the abundance of each taxonomic unit is obtained by accumulating the contribution values ​​of the corresponding reads.

[0089] (2) Output the classification assignment information (tax id / classification label) of each read at the read level, and retain the traceability fields (read id, alignment position / hit reference entry, etc.) to form evidence of "microbial attribution at the nucleic acid level".

[0090] 1.3.3 Protein-level origin tracing and consistency constraints (core of the evidence chain) (1) The reads (or their effective sequence fragments) with nucleic acid level classification assignment in item 1.3.2 are translated and aligned to the microbial protein library to obtain the protein accession and its tax ID; among them, the translation alignment can be performed using BLASTX (an algorithm for aligning nucleic acid translation to a protein database) or DIAMOND (an accelerated protein sequence alignment tool).

[0091] (2) Perform "consistency check" on the same read: require that the tax ID corresponding to the read at the nucleic acid level (classification assignment / genome hit) is consistent with the tax ID inferred at the protein level (translation alignment hit protein) at the preset classification level.

[0092] (3) Remove or mark those with inconsistencies as low confidence, and record the reason for removal (inconsistency between nucleic acid and protein attribution).

[0093] (4) Output a traceable evidence chain field to record the identifier (read id), nucleic acid level classification number (tax id), translation alignment of the protein accession, and the corresponding protein level classification number (tax id) of each read, so as to support subsequent grading and verification.

[0094] 1.3.4 Strict Threshold Filtering and Candidate Peptide Generation (1) Apply a high consistency threshold filter to the protein hit results. The threshold includes at least: sequence identity, E-value, and alignment coverage. In this embodiment, the sequence identity threshold is set to 100%, the alignment coverage threshold is set to no less than 90%, and the E-value threshold is set to no more than 1e. 5.

[0095] (2) On the protein sequence that has passed the consistency check and meets the threshold, a sliding window is used to generate candidate peptides composed of HLA-I and HLA-II. In this embodiment, the peptide length range is set to 8-10 amino acids to adapt to the predicted length requirement of HLA-I and 15 amino acids to adapt to the predicted length requirement of HLA-II.

[0096] (3) For each candidate peptide, retain at least the following fields: source tax_id, source protein accession, number of supporting reads, species abundance and evidence chain summary (nucleic acid-protein consensus marker and key threshold hit status).

[0097] (4) Output the list of microbial candidate peptides and the chain of evidence, and retain the statistics of the change in the number of candidates before and after filtration for review and quality assessment.

[0098] 1.4 Generation of Cryptoantigen Candidate Peptides Based on WTS data, candidate peptides for cryptogenic antigens were generated from translation products of small open reading frames (sORFs) in non-coding regions using a strategy of "two-branch discovery + reference genome coordinate mapping and annotation + multi-position matching constraints + joint judgment based on expression evidence".

[0099] 1.4.1 RNA Alignment and Basic Input (1) Align WTS reads to the human reference genome GRCh38.p3 to generate a standardized alignment file (e.g., BAM format).

[0100] (2) Prepare reference annotations (e.g., GTF / GFF; GENCODE v23) and inputs required for subsequent transcript reconstruction.

[0101] (3) Output comparison statistics (comparison rate, splicing comparison ratio, duplication rate, etc.) and logs.

[0102] 1.4.2 Annotated non-coding transcript branch (1) Extract the transcript structure and sequence range of the annotated non-coding transcripts from the reference notes in section 1.4.1. In this embodiment, the annotated non-coding transcripts are long non-coding RNAs (lncRNAs).

[0103] (2) Within this range, sample-specific variation information is integrated to obtain transcript sequences that are closer to the sample, in order to reduce reference bias.

[0104] (3) Annotated non-coding transcripts are used to form a set for subsequent ORF discovery.

[0105] 1.4.3 New Transcript Branch (Transcriptional Events Outside of Annotation) (1) Transcript reconstruction was performed based on RNA-seq alignment results to obtain a sample-specific transcript set.

[0106] (2) Compare with the reference notes in section 1.4.1 and screen for transcripts not covered by the reference notes. These uncovered transcripts include, but are not limited to, transcripts from intergenic regions, intronic regions and / or antisense strands.

[0107] (3) Local de novo assembly of candidate regions (an assembly method for reconstructing transcription sequences from reads) to improve transcript integrity.

[0108] 1.4.4 ORF discovery, SEP definition, and reference genome coordinate mapping annotation (1) The new transcript set in section 1.4.3 is merged with the annotated non-coding transcript set in section 1.4.2 to form a combined transcript set, and the source marker (annotated / new transcript) of each transcript is retained.

[0109] (2) Perform open reading frame (ORF) discovery in the combined transcript set, set a minimum ORF length threshold (in this embodiment, the minimum ORF length threshold is set to 10 amino acids, and the maximum ORF length is set to be less than 100 amino acids); screen for short open reading frames (sORF) that meet the aforementioned length range, and translate them to obtain short open reading frame encoded peptides (SEP).

[0110] (3) The coding nucleic acid sequence corresponding to the SEP is mapped / located and verified by the reference genome coordinates (e.g., the coding sequence is aligned or mapped to the reference genome to determine the coordinate range) to obtain its genome coordinates.

[0111] (4) Combine genome annotation to annotate the coordinate range of SEP, including intergenic regions, introns, antisenses, relative relationships with genes / transcriptions, and whether they overlap with protein-coding CDS.

[0112] It should be noted that, given that the short epitope peptides used for HLA prediction are derived from SEP through a sliding window, the genomic coordinates and overlapping annotations are preferably given at the level of the source genome interval determined by SEP (or its corresponding ORF), and are recorded as source information of the derived short epitopes for auditing, verification and tracing purposes.

[0113] 1.4.5 Joint Determination of Multi-mapping Non-coding Constraints and Expressive Evidence (1) When the coding sequence of SEP is not unique in the reference genome, i.e. there are multiple matching positions, a hard constraint rule is executed. In this embodiment, the hard constraint rule requires that all positioning points are non-coding and do not overlap with the protein coding CDS; if any positioning point overlaps with the CDS, it is removed or downgraded, and the reason is recorded.

[0114] (2) Quantify tumor / control expression of SEPs that have passed the constraints, and screen tumor-specific abnormal expression SEPs (aeSEPs) based on tumor expression, control expression and differential expression thresholds.

[0115] (3) Only the candidate short epitope peptides of the hidden antigen generated by aesP are entered into the downstream unified classification, and the expression evidence field and threshold hit information are retained.

[0116] 1.5 Generation of candidate peptides for mutant neoantigens Based on the results of somatic cell mutation detection using WES, candidate peptides of mutated neoantigens are generated, and reliability is guaranteed through "multi-algorithm joint detection + cross-support + standardized processing".

[0117] 1.5.1 Mutation Detection and Standardization Processing (1) The WES data were compared, duplicates were marked and necessary corrections were performed to obtain high-quality comparison results.

[0118] (2) In this embodiment, three somatic mutation detection tools, Mutect2, Strelka2 and VarDict, are run to obtain a set of candidate somatic mutations.

[0119] (3) The candidate somatic mutation set is normalized, including: coordinate sorting, reference consistency, insertion and deletion left alignment (also known as insertion and deletion normalization), and splitting multiple allele sites to ensure that downstream annotation is consistent with peptide generation.

[0120] (4) Output the PASS pass rate, mutation type distribution and change in number before and after normalization for each detection tool.

[0121] 1.5.2 Consensus Mutation Set with Cross-Support of Multiple Algorithms (1) Consensus rule construction: In this embodiment, three independent somatic mutation detection algorithms are used to perform parallel detection on the same tumor-normal sample. Mutation sites covered by at least two algorithms are formed into a high-confidence consensus mutation set to reduce the risk of false positives caused by a single algorithm. The consensus rule (intersection of mutation sites) and the required number of supporting algorithms can be parameterized. In this embodiment, the minimum allele frequency threshold is set to no less than 0.02, and the minimum base quality threshold is set to no less than 20.

[0122] (2) Further noise reduction by combining RNA-seq evidence verification (e.g., tumor sample expression support).

[0123] (3) Output the high-confidence consensus mutation set and the support matrix of each tool for easy review.

[0124] 1.5.3 Unified Selection of Peptide Generation and Transcripts (1) Functionally annotate the consensus mutations in section 1.5.2 to obtain information on protein changes, and generate candidate mutant peptides centered on the mutation sites accordingly.

[0125] (2) Simultaneously generate the corresponding wild-type peptide to support the subsequent calculation and priority ranking of mutation / wild-type difference indicators.

[0126] (3) To reduce peptide inconsistencies caused by the same mutation under different transcript annotations, a unified transcript selection rule is adopted (e.g., priority is given to protein-coding transcripts that are most highly expressed in tumor samples and are consistent with RNA evidence), and the selected transcripts and the reason field are recorded.

[0127] 1.6 HLA typing (1) HLA allele inference was performed on WES or WTS data using HLA typing tools to obtain the HLA class I and class II allele sets.

[0128] (2) Standardize the allele names (unify them to a standard format) and output an HLA list for subsequent prediction.

[0129] 1.7 HLA Binding / Presentation Prediction Module (1) Construct peptide-HLA pairing sets for the three types of candidate peptides (microbial antigens, cryptogenic antigens and mutant neoantigens) with the HLA alleles of the samples in item 1.6.

[0130] (2) Integrate two or more HLA binding / presentation prediction algorithms (including but not limited to affinity prediction and presentation / ligand identification models) and unify the output fields (e.g. affinity, percentile / ranking value, etc.).

[0131] (3) Summarize the HLA pairings of each peptide to form a comprehensive index that can be used for grading, such as best / median affinity, best / median percentile, etc.

[0132] (4) The system can configure model sets and threshold calibers for HLA-I and HLA-II respectively, and retain the calculation details and logs of model output to summary indicators.

[0133] (5) Output: Original results table of each peptide-HLA multi-model, summary index table, and records of abnormal / deleted processing.

[0134] 1.8 Predicting Immunogenicity 1.8.1 Input Construction and Joint Sequence Representation (1) Input for constructing the model for each peptide-HLA pairing: The input consists of a joint sequence composed of the amino acid sequence of the candidate peptide and its corresponding HLA amino acid sequence; wherein, the joint sequence can be constructed by sequential splicing, and a separator or zero vector is set between the peptide and the HLA amino acid sequence to distinguish the source. The HLA amino acid sequence can be a complete protein sequence or a pseudo-sequence (HLA key residue sequence) used for modeling, and the specific form can be configured.

[0135] (2) Amino acid sequence tokenization: A vocabulary containing 20 standard amino acids and necessary special symbols and padding symbols is established, and each amino acid is mapped to a corresponding discrete identifier or continuous embedding index for subsequent neural network processing.

[0136] (3) Physicochemical property-driven residue-level representation: Multidimensional physicochemical properties are extracted from AAindex (amino acid physicochemical property index library) and dimensionality is reduced (e.g., principal component dimensionality reduction) to obtain residue vectors, which are used to initialize or enhance the embedding representation. Among them, the residue vectors can be further optimized as trainable parameters during model training to take into account both biological priors and task adaptability.

[0137] (4) Global physicochemical descriptor of peptide: Calculate global features such as amino acid composition, polarity, volume, charge, hydrophobicity, Boman index, aliphatic index, and isoelectric point, and represent them in a vectorized form. The features can be projected to a unified feature space through linear mapping or embedding layers to be fused with sequence features.

[0138] 1.8.2 Model Structure and Scoring Output (1) Sequence modeling network is used. In this embodiment, a bidirectional long short-term memory network (BiLSTM) is selected to encode the joint sequence representation. The contextual association features between peptides and HLA are captured by the joint modeling of forward and reverse sequence information.

[0139] (2) The sequence encoding result and the global physicochemical features of the peptide are embedded and spliced ​​and then input into the classifier to output the immunogenicity score (e.g., 0-1). The score represents the predicted probability that the corresponding peptide-HLA pair has immunogenicity and can be used for sorting or threshold screening.

[0140] (3) Batch processing and parallel feature generation are used in the inference stage, while maintaining the consistency of the input order, in order to improve the computational efficiency of large-scale candidate peptide evaluation.

[0141] (4) Abnormal handling: When some HLA alleles cannot be mapped to the sequence, the system can assign a neutral score (e.g., 0.5) or mark it as unevaluable according to a preset strategy, and record the reason to avoid bias or interruption of prediction results due to missing information.

[0142] 1.8.3 Source-Specific Training and Calibration (1) Training datasets and labeling systems were constructed for microbial antigens, cryptogenic antigens, and mutated neoantigens, respectively. The labels were derived from experimental validation data, publicly available immunology databases, or quality-controlled compilations of literature.

[0143] (2) Train source-specific immunogenicity models separately (or perform source-specific calibration based on a shared backbone network) to adapt to the statistical differences in sequence composition, length distribution, and physicochemical properties of the three types of antigens. Introduce source-related parameter adjustment or output layer calibration strategies based on the shared sequence coding layer.

[0144] (3) Output: Model version, training configuration, evaluation metrics, and applicable scope description for each source, ensuring reproducible deployment. It also supports selecting the corresponding immunogenicity prediction model from different application scenarios for invocation.

[0145] 1.9 Unified Hierarchy and Priority Sorting Based on the unified fusion of HLA binding / presentation prediction results, immunogenicity scores, and source-specific evidence factors, cross-source comparable grading and ranking are achieved.

[0146] (1) Establish a unified grading system: for example, high / medium / low and subthreshold, and agree on the grading rule of "taking the highest grade if multiple grades are satisfied".

[0147] (2) The hierarchical features include at least: a) Immunogenicity score threshold; b) Combined / presented summary metrics for multi-algorithm integration (preferably requiring both best and median to meet thresholds simultaneously to improve robustness); c) Specific evidence of origin: For microbial candidates, at least the following should be considered: species abundance, number of supporting reads, nucleic acid / protein consistency evidence chain, and threshold hit rate. For hidden antigen candidates, at least tumor / control expression, differential expression, non-coding constraints by markers and SEP block annotations should be included; For mutation candidates, at least the following should be included: multi-algorithm consensus support markers, mutation function annotations, tumor expression support, and mutation / wildtype difference indicators.

[0148] (3) Priority sorting: Based on the unified classification, construct comprehensive sorting rules (weighted, hierarchical sorting or rule tree) and output sorting basis fields (key threshold hit, evidence chain strength, source evidence bonus items, etc.).

[0149] (4) Output: Unified candidate summary table, hierarchical result table, sorting list and interpretable fields.

[0150] Representative peptide examples based on unified grading and comprehensive sorting screening in this embodiment are shown in Tables 1-3: Table 1 Examples of candidate antigenic epitopes from microbial sources

[0151] Table 2 Examples of candidate epitopes for cryptogenic antigens

[0152] Table 3 Examples of candidate antigenic epitopes from which mutations originate

[0153] It should be noted that the threshold parameters and representative peptides shown in Tables 1-3 are only examples of one embodiment of this application. The relevant thresholds can be adjusted according to the sample type, application scenario or model version. All equivalent transformations based on the same technical concept should be included within the protection scope of this application.

[0154] 1.10 Three-Source Molecular Mimicry Analysis and Unified Output Module 1.10.1 Construction of Cross-Source Pairing Sets (1) Construct cross-source pairings for the three-source candidate sets: microorganism-cryptic antigen, microorganism-mutation, and cryptogenic antigen-mutation.

[0155] (2) The pairing is constructed in a “systematic pairing” manner. Pre-screening strategies can be set according to hierarchical thresholds, HLA conditions or candidate size to cover the target candidate range under controllable computational load.

[0156] (3) Record the pairing size, pre-screening conditions and logs.

[0157] 1.10.2 Similarity Calculation and Mimicry Screening (1) Perform sequence similarity calculation on the pairing set, preferably using an interpretable index, such as normalized similarity based on the longest common substring (LCS); other equivalent similarity measures may also be used.

[0158] (2) Set similarity thresholds and allow tiered threshold strategies (different thresholds can be set for different combinations of sources).

[0159] (3) Optional joint conditions: For example, both parties need to meet the minimum level or share / compare HLA conditions, so as to improve the availability of the mimicry combination.

[0160] (4) Output the screening results and statistics of reasons for failure (threshold not reached, HLA conditions not met, etc.).

[0161] 1.10.3 Unified Output and Statistical Summary (1) Unified output of mimicry results: Each mimicry pair / combination should include at least two peptide segments, source category, source information (microbial tax ID / protein accession, hidden antigen SEP / block annotation, mutation site annotation, etc.), similarity value, and grade and sorting position of both parties.

[0162] (2) Statistical summary: number of mimicry pairs, similarity distribution, key node peptides (peptides that participate in mimicry at high frequency), proportion of source combinations, etc.

[0163] It should be noted that this application is not limited to the above-described embodiments. The above embodiments are merely examples, and any embodiments with the same structure and effect as the technical concept within the scope of this application are included in the technical scope of this application. Furthermore, various modifications that can be conceived by those skilled in the art to the embodiments, and other ways of constructing by combining some of the constituent elements of the embodiments, without departing from the spirit of this application, are also included in the scope of this application.

Claims

1. A method for screening multi-source tumor antigens, characterized in that, Includes the following steps: S1. Obtain high-throughput sequencing data of tumor tissue and its paired samples, wherein the high-throughput sequencing data includes WES or WGS data, and WTS or RNA-seq data. S2. Process the data from step S1 to determine the non-host and non-vector-origin sequences and the host sequence; S3. Based on the non-host and non-vector-derived sequence, generate microbial antigen candidate peptides; S4. Based on WTS or RNA-seq data, construct a combined transcript set that simultaneously covers the first transcript source and the second transcript source, and generate cryptogenic candidate peptides from it; wherein, the first transcript source is an annotated non-coding transcript source, and the second transcript source is a transcript source other than the reference annotation obtained from transcript reconstruction; S5. Based on WES or WGS data, generate candidate peptides for mutated neoantigens through somatic cell mutation detection; S6. Perform HLA binding / presentation prediction and immunogenicity prediction on candidate peptides of microbial antigens, candidate peptides of cryptogenic antigens, and candidate peptides of mutant neoantigens, respectively. Integrate the HLA binding / presentation prediction results and immunogenicity scores with source-specific evidence factors to achieve unified grading and priority ranking. Based on the comprehensive score, screen out tumor antigens from the candidate antigen peptides.

2. The screening method as described in claim 1, characterized in that, Step S2 includes the following steps: S21. Perform quality control on high-throughput sequencing data to obtain clean reads; S22. Align the clean reads to the universal human reference genome and extract the unaligned reads; S23. Align the unaligned reads from step S22 to the telomere-to-telomere human reference genome and extract the unaligned reads as non-host sequences. S24. Align the non-host sequences to the vector contamination reference database and retain the miss reads, which are sequences that are neither from the host nor from the vector. S25. Screening of candidate peptides for latent antigens and mutated neoantigens based on host alignment results and host reads; And / or, step S3 includes the following steps: S31. Construct a microbial protein library, which includes a genomic reference library for microbial classification / quantification and a protein sequence alignment library for protein tracing. S32. Perform microbial classification and quantification on sequences that are not from the host or vector, and obtain the abundance at the species / genus level; output the classification assignment information of each read at the read level, and retain the source field to achieve microbial classification and quantification at the nucleic acid level; S33. Translate and align reads or their effective sequence fragments that have microbial classification assignments at the nucleic acid level to the microbial protein library, and retain reads that are consistent with the microbial classification at both the nucleic acid and protein levels. S34. Threshold filtering and candidate peptide generation are performed on the reads hit in step S33 to obtain microbial antigen candidate peptides. Preferably, the threshold filtering includes at least: sequence consistency, E-value, and alignment coverage.

3. The screening method as described in claim 1, characterized in that, The first transcript source is obtained by extracting the transcript sequence range of the annotated non-coding transcripts from WTS or RNA-seq reads based on reference annotations. The second source of transcripts is obtained by reconstructing transcripts from WTS or RNA-seq alignment results to obtain a sample-specific transcript set, from which new transcript sequences not covered by reference annotations are screened.

4. The screening method as described in claim 1, characterized in that, The screening of candidate peptides for cryptogenic antigens in step S4 includes the following steps: ORF discovery was performed in the combined transcript set, sORFs were screened and translated to obtain short open reading frame-encoding peptides (SEPs). The coding nucleic acid sequence corresponding to the SEP is mapped and located using reference genome coordinates to obtain its genome coordinates, and its coordinate range is annotated. When the coding sequence of a SEP has multiple positional matches in the reference genome, hard constraint rules are applied, and tumor / control expression is quantified for SEPs that pass the constraints; tumor-specific abnormally expressed SEPs are screened as candidate peptides for cryptogenic antigens based on tumor expression, control expression, and differential expression thresholds. Preferably, the hard constraint rule is that all positioning points are non-coding and do not overlap with protein-coding CDS.

5. The screening method as described in claim 1, characterized in that, Step S5 includes the following steps: S51. Based on WES or WGS data, run two or more somatic mutation detection algorithms to obtain a candidate somatic mutation set, and standardize the mutation records. S52. Retain mutations supported by at least two somatic mutation detection algorithms to form a high-confidence consensus mutation set; S53. Perform functional annotation on the consensus mutations in the high-confidence consensus mutation set in step S52 to obtain protein change information, and generate candidate mutant peptides centered on the mutation sites as candidate mutant neoantigen peptides.

6. The screening method as described in claim 1, characterized in that, In step S6, the HLA binding / presentation prediction is performed using two or more HLA binding / presentation prediction algorithms, and the output fields are standardized. And / or, in step S6, immunogenicity prediction includes the following steps: a) Constructing a combined sequence: The combined sequence is composed of the amino acid sequence of the candidate peptide and its corresponding HLA amino acid sequence; b) The joint sequence representation is encoded using a sequence modeling network. Then, based on the sequence encoding results and the global physicochemical characteristics of the peptide, an immunogenicity score is output. c) Train source-specific immunogenicity models for microbial antigens, cryptogenic antigens, and mutated neoantigens respectively, and predict immunogenicity; And / or, the grading features of the unified grading include at least: immunogenicity score threshold, multi-algorithm integrated binding / presentation summary index, and source-specific evidence; The priority sorting is based on a unified hierarchical system and a comprehensive sorting rule is constructed. Preferred, the source-specific evidence includes: For microbial candidates, at least the species abundance, the number of supporting reads, and the nucleic acid / protein consistency evidence chain and threshold hit rate should be considered. For hidden antigen candidates, at least tumor / control expression, differential expression, non-coding constraints by markers and SEP block annotations should be included; For mutation candidates, at least multi-algorithm consensus support markers, mutation function annotations, tumor expression support, and mutation / wildtype difference indicators should be included.

7. A screening system for multi-source tumor antigens, characterized in that, When the screening system is running, it implements the screening method as described in any one of claims 1-6; The screening system includes: The data input module is configured to receive high-throughput sequencing data of tumor tissue and its paired samples; The data preprocessing and sequence identification module is configured to perform quality control and alignment of high-throughput sequencing data, and to identify host sequences and sequences from non-host and non-vector sources. A microbial antigen candidate peptide generation module is configured to generate microbial antigen candidate peptides based on the non-host and non-vector-derived sequence data. The cryptogenic antigen candidate peptide generation module generates cryptogenic antigen candidate peptides based on WTS or RNA-seq data. The mutant neoantigen candidate peptide generation module generates mutant neoantigen candidate peptides based on WES or WGS data and through somatic cell mutation detection. The HLA typing module is configured to use HLA typing tools to obtain sets of HLA class I and II alleles and output an HLA list for subsequent prediction. The HLA binding / presentation prediction module is configured to predict HLA binding / presentation using two or more HLA binding / presentation prediction algorithms; An immunogenicity prediction module is configured to predict the immunogenicity of candidate antigens from three sources. The unified hierarchical and priority sorting module is configured to establish a unified hierarchical system and construct comprehensive sorting rules; Preferably, the screening system further includes a cross-three-source molecular mimicry analysis and unified output module, which is configured to construct cross-source pairings for the three-source candidate set, perform similarity calculation and mimicry screening, and provide unified output and statistical summary.

8. A computer device comprising a memory, a processor, and a computer program or instructions stored in the memory, characterized in that, When the computer program or instructions are executed by a processor, they implement the screening method for multi-source tumor antigens as described in any one of claims 1-6.

9. A computer-readable storage medium storing a computer program or instructions thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the screening method for multi-source tumor antigens as described in any one of claims 1-6.

10. An antigenic epitope peptide, characterized in that, The screening method described in any one of claims 1-6 was used to obtain the results. Preferably, the antigenic epitope peptide is at least one of the peptide segments with amino acid sequences as shown in any of SEQ ID NO.1-SEQ ID NO.14.