Tumor neoantigen accurate analysis method and system based on multi-dimensional evidence chain scoring

By identifying anchor points using multidimensional evidence chain scoring and amino acid frequency-weighted BLOSUM62 coding values, the problems of multi-omics data validation and MHC-peptide binding specificity prediction in tumor neoantigen identification have been solved, enabling precise analysis and efficient research of tumor neoantigens.

CN120895097AInactive Publication Date: 2025-11-04南昌大学第一附属医院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511395453.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-11-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies for the identification, validation, and prediction of tumor neoantigens suffer from several problems, including a lack of quantitative assessment of multi-omics data validation, insufficient reliability assessment of immune peptide omics data, inadequate accuracy in predicting MHC-peptide binding specificity, and low intelligence of HLA molecular retrieval systems.

Method used

Using a multidimensional evidence chain scoring method, combined with genomic, transcriptomic, and proteomic data, and employing amino acid frequency-weighted BLOSUM62 coding values ​​to identify anchor points, an intelligent MHC retrieval module was established for accurate tumor neoantigen analysis.

Benefits of technology

It enables precise identification and verification of tumor neoantigens, improves identification accuracy and research efficiency, distinguishes tumor-specific antigens from background noise, and enhances the ease of use and accuracy of MHC retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895097A_ABST
    Figure CN120895097A_ABST
Patent Text Reader

Abstract

The invention discloses a tumor neoantigen accurate analysis method and system based on multi-dimensional evidence chain scoring, and relates to the cross technical field of bioinformatics, computational biology and tumor immunology, and the system comprises a multi-dimensional evidence chain scoring module which is used for evaluating the feasibility score of candidate neoantigens; the antigen extraction and verification module is used for identifying and verifying a tumor specific antigen according to the feasible degree score; the anchoring site identification and prediction module is used for scoring the amino acid of each site by using an amino acid frequency weighted BLOSUM62 coded value, obtaining a specific anchoring site and an anchoring site amino acid residue restriction set, and judging whether a certain peptide fragment is a non-binding peptide or not according to the specific anchoring site and the anchoring site amino acid residue restriction set so as to obtain a judgment result; the intelligent MHC retrieval module is used for enabling a user to input a character string at the front end according to the judgment result to conduct intelligent retrieval so as to obtain a retrieval result, and high-precision recognition and evaluation of the tumor neoantigen are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of bioinformatics, computational biology and tumor immunology, and particularly relates to a tumor neoantigen precise analysis method and system based on multi-dimensional evidence chain scoring. BACKGROUND

[0002] Tumor neoantigen is a peptide segment produced by tumor-specific mutation, presented by major histocompatibility complex (MHC), and recognized by the body's immune system, which is the core target of personalized tumor immunotherapy. Accurate identification and evaluation of tumor neoantigen is crucial for the development of effective tumor vaccines and cell therapy. However, the existing technology has the following key technical problems in the identification, verification and prediction of neoantigens: First, the multi-omics data verification lacks a quantitative evaluation system: the authenticity verification of neoantigens requires the integration of whole exome sequencing (WES), RNA sequencing (RNA-Seq) and immune polypeptidomics. The existing method usually adopts a simple "yes / no" binary judgment, lacking a quantitative evaluation mechanism for incomplete evidence chains. In actual research, many candidate antigens are only detected at part of the omics level. How to scientifically evaluate the credibility of these "partial evidence" antigens is the blind spot of the current technology.

[0003] Second, the credibility evaluation of immune polypeptidomics data is missing: mass spectrometry detection of immune polypeptidomics data contains rich quality control information, such as peptide segment score (Score), mass deviation (PPM), position-specific score (PS, Positional Score), etc. The existing technology fails to fully utilize these information to establish a systematic credibility evaluation system, resulting in a large number of false positive results affecting subsequent analysis.

[0004] Third, the prediction accuracy of MHC-peptide binding specificity is insufficient: the binding of MHC molecules and peptides has position specificity, and the amino acid residues at certain positions (anchor positions) are crucial for binding. The existing prediction methods are mostly based on machine learning black box models, lacking accurate identification and biological interpretation of anchor positions. In addition, the existing methods fail to fully consider the physicochemical property similarity of amino acids, resulting in inaccurate prediction of low-frequency but functionally similar amino acid substitutions.

[0005] Fourth, the intelligence level of HLA molecule retrieval system is low: the HLA naming system is complex, with multiple naming formats and variants. The existing retrieval tools require users to input standardized HLA names, and have poor fault tolerance for format errors or incomplete input, seriously affecting research efficiency. SUMMARY

[0006] The embodiment of the application aims to provide a tumor neoantigen precise analysis method and system based on multi-dimensional evidence chain scoring, which can solve the problem of low recognition accuracy when the prior art recognizes tumor neoantigen.

[0007] To solve the above technical problems, the application is implemented as follows: In a first aspect, the embodiment of the application provides a tumor neoantigen precise analysis system based on multi-dimensional evidence chain scoring, which comprises: A multi-dimensional evidence chain scoring module is configured to evaluate the feasibility score of the candidate neoantigen. An antigen extraction and verification module is configured to identify and verify tumor-specific antigens according to the feasibility score. An anchor site identification and prediction module is configured to score the amino acid of each site using the amino acid frequency weighted BLOSUM62 encoding value, obtain the specific anchor site and anchor site amino acid residue restriction set, and determine whether a peptide segment is a non-binding peptide according to the anchor site and anchor site amino acid residue restriction set, to obtain a judgment result. An intelligent MHC retrieval module is configured to perform intelligent retrieval according to the judgment result and the input string of the user in the front end, to obtain a retrieval result.

[0008] Optionally, the specific steps of evaluating the feasibility score of the candidate neoantigen include: obtaining a genomic dimension evidence score based on the mutation quality, sequencing depth and mutation type of the candidate neoantigen; obtaining a transcriptome dimension evidence score based on the gene expression amount and mutation allele frequency; obtaining a proteome dimension evidence score based on the fusion mass spectrum; obtaining a cross-validation reward score when the candidate neoantigen is detected at multiple omics levels; and weighting the genomic dimension evidence score, the transcriptome dimension evidence score, the proteome dimension evidence score and the cross-validation reward score to obtain the feasibility score of the candidate neoantigen.

[0009] Optionally, the specific steps of scoring the amino acid of each site using the BLOSUM62 encoding value weighted by the frequency of amino acids, obtaining the specific anchor site and the anchor site amino acid residue restriction set, and judging whether a certain peptide segment is a non-binding peptide, include: taking the MHC and the length of the peptide segment as the grouping basis of the MHC-peptide segment data; taking the amino acid with the highest frequency in each position of the peptide segment as the reference amino acid according to the grouping basis; obtaining the BLOSUM62 encoding value of the amino acid in each position of the peptide segment relative to the reference amino acid; multiplying the BLOSUM62 encoding value by the frequency of the amino acid to obtain the weighted BLOSUM62 value of each site of the peptide segment; scoring the position with a weighted BLOSUM62 value higher than the average score of the weighted BLOSUM62 value of all peptide segment positions as an anchor site; in each anchor site, all amino acids are classified into consistent residues and different residues according to whether the BLOSUM62 value is greater than 0, and an anchor site amino acid residue restriction set is established; and establishing a knowledge graph based on the specific anchor site and the anchor site amino acid residue restriction set, and judging whether a certain peptide segment is a non-binding peptide.

[0010] Optionally, the antigen extraction and verification module includes WES data processing, RNA-Seq data processing, and immunopeptidomics data processing, The process of WES data processing includes: Finding the corresponding mutation site in the WES data to obtain the mutation quality score, the sequencing depth of the tumor sample, the mutation type weight, and the tumor purity adjustment coefficient to obtain the genomic evidence score, wherein the genomic evidence score includes: ; Wherein, is the genomic evidence score, denotes the mutation quality score, denotes the sequencing depth of the tumor sample, denotes the mutation type weight, denotes the tumor purity adjustment coefficient; The process of RNA-Seq data processing includes: obtaining the gene expression of TP53 gene, the mutant allele frequency, and the expression fold of tumor and normal cells in the RNA-Seq data to calculate the transcriptome evidence score, and the transcriptome evidence score includes: ; Wherein, is the transcriptome evidence score, denotes the gene expression, denotes the mutant allele frequency, denotes the expression fold of tumor and normal cells; The process of immune polypeptidomics data processing includes: obtaining a proteome score based on a mass spectrometry detection score, a mass deviation, a position specificity score, and a signal intensity ratio of a cancer tissue to a paracancer tissue, wherein the proteome score includes: ; wherein, is the proteome score, represents the mass spectrometry detection score, represents the mass deviation, represents an average value of the position specificity score, represents the signal intensity ratio of the cancer tissue to the paracancer tissue.

[0011] Optionally, the intelligent MHC retrieval module includes: a backend standard MHC nomenclature index module for maintaining and periodically automatically updating a complete standard MHC nomenclature index table based on an authoritative database, performing semantic understanding and standardization processing on a query string input by a user; a nomenclature simplification module for removing “-”, “*”, “:”, “^” symbols from the standard MHC nomenclature and uniformly converting them into capital letters, generating a dictionary corresponding to the standard nomenclature, and being compatible with important abbreviations including HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQB1 and HLA-DRB1 loci; a user input processing module for receiving a query string input by a user, performing non-standard symbol replacement, capital letter conversion and common spelling error intelligent correction on the query string to generate a standardized query string; a multi-parameter dynamic weight ordering module for intelligently ordering a candidate MHC molecule list according to species weight ordering, locus weight ordering and allele frequency weight, so as to preferentially present high-frequency and scientific research or clinically relevant MHC molecules to the user.

[0012] In a second aspect, the embodiments of the present application also provide a tumor neoantigen precise analysis method based on multi-dimensional evidence chain scoring, which includes: evaluating a feasibility score of a candidate neoantigen; identifying and verifying a tumor-specific antigen according to the feasibility score; scoring amino acids at each site using amino acid frequency weighted BLOSUM62 encoding values, obtaining a specific anchor site and an anchor site amino acid residue restricted set, and judging whether a peptide segment is a non-binding peptide according to the same to obtain a judgment result; according to the judgment result, a user inputs a string in the front end for intelligent retrieval to obtain a retrieval result.

[0013] Compared with the prior art, the tumor neoantigen precise analysis system based on multi-dimensional evidence chain scoring provided by the application has the beneficial effects in the following aspects: Firstly, the application breaks through the limitation of traditional binary judgment and establishes a scientific quantitative scoring mechanism. By comprehensively considering the multi-dimensional evidence of genome, transcriptome and proteome, and the multiple quality parameters of mass spectrometry data, the credibility of the candidate neoantigen is accurately quantified.

[0014] Secondly, the amino acid frequency weighted BLOSUM62 value is applied to specific anchor site identification and key residue identification. This method not only can identify the specific anchor site of the peptide segment, but also can identify the functionally similar amino acid group as the consistent residue (i.e. the anchor site amino acid residue restriction set) based on the evolutionary conservation. This makes the algorithm can accurately predict the low frequency but functionally equivalent amino acid substitution, which is mainly used to construct the knowledge graph to quickly and accurately exclude the peptides that cannot be presented by MHC class I molecules.

[0015] Thirdly, by comparing and analyzing the cancer tissue and the adjacent normal tissue, and combining with the allele-specific expression quantification, the application can accurately distinguish the tumor-specific antigens and background noise. This differential analysis strategy greatly improves the specificity of the neoantigen, and provides more reliable target for personalized treatment.

[0016] Fourthly, by semantic analysis and multi-parameter dynamic sorting, the application improves the ease of use and accuracy of MHC retrieval to a new level. The fault tolerance rate of the system to various non-standard inputs reaches more than 95%, which greatly improves the research efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 is an internal structure diagram of a tumor neoantigen precise analysis system based on multi-dimensional evidence chain scoring provided by the first embodiment of the application; Figure 2 is a general architecture diagram of the tumor neoantigen precise analysis platform provided by the first embodiment of the application; Figure 3 is a multi-omics evidence chain quantitative scoring algorithm flowchart provided by the first embodiment of the application; Figure 4 is an anchor site and its allowable amino acid residue identification algorithm schematic diagram based on BLOSUM62 coding provided by the first embodiment of the application; Figure 5 is a tumor-specific neoantigen extraction and verification flowchart provided by the first embodiment of the application; Figure 6 is an intelligent MHC molecule automatic search workflow diagram provided by the first embodiment of the application; Figure 7is a flow chart of a tumor neoantigen precise analysis method based on multi-dimensional evidence chain scoring provided by the second embodiment of the present application. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0019] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.

[0020] The multi-dimensional evidence chain scoring-based tumor neoantigen precise analysis method and system provided by the embodiments of the present application will be described in detail below with reference to the drawings, through specific embodiments and application scenarios.

[0021] Embodiment 1 Please refer to Figure 1 , Figure 1 The internal structure diagram of the multi-dimensional evidence chain scoring-based tumor neoantigen precise analysis system provided by the present application is shown, which includes S1 to S4.

[0022] S1, the multi-dimensional evidence chain scoring module 100 is used for evaluating the feasibility score of the candidate neoantigen.

[0023] Specifically, the specific steps of evaluating the feasibility score of the candidate neoantigen include: obtaining a genomic dimension evidence score based on the mutation quality, sequencing depth and mutation type of the candidate neoantigen; obtaining a transcriptome dimension evidence score based on the gene expression amount and mutation allele frequency; obtaining a proteome dimension evidence score based on fusion mass spectrometry; obtaining a cross-validation reward score when the candidate neoantigen is detected at multiple omics levels at the same time; and obtaining the feasibility score of the candidate neoantigen by weighting the genomic dimension evidence score, the transcriptome dimension evidence score, the proteome dimension evidence score and the cross-validation reward score.

[0024] Further, in step S1, the application innovatively proposes a comprehensive scoring system for evaluating the credibility of candidate neoantigens. The method first identifies tumor-specific mutations from whole exome sequencing data through mutation calling, and generates theoretical peptide sequences based on the mutation sites; from RNA-Seq data, the expression level of the mutant gene is confirmed through quantitative analysis, and the mutant allele frequency (VAF) is calculated; from the immunopeptidomics data, the experimental detected peptide sequences and their quality parameters are extracted. The key innovation lies in the three-dimensional evidence scoring function designed by the application, the specific expression is as follows: ; wherein, represents the comprehensive credibility score of the candidate neoantigen, 、 represents a dynamic weight coefficient, represents a genomic dimension evidence score, represents a transcriptome score, represents a proteomic dimension evidence score, represents a cross-validation reward score.

[0025] S2, the antigen extraction and verification module 200 is used to identify and verify tumor-specific antigens according to the feasibility score.

[0026] Specifically, in the execution step S2, the specific steps of identifying and verifying tumor-specific antigens include three processes of WES data processing, RNA-Seq data processing and immunopeptidomics data processing. This step takes a lung adenocarcinoma patient as an example to describe how to extract and verify tumor-specific neoantigens from the multi-omics data of a lung adenocarcinoma patient.

[0027] Specifically, in the process of WES data processing, the application compares the whole exome sequencing data of tumor tissue and paired blood samples to identify somatic mutations. Take the EGFR gene L858R mutation as an example: mutation position: chr7: 55259515 T>G amino acid change: p.L858R (CTG> CGG) tumor mutation frequency: 42% normal tissue: no mutation detected, and a sliding window algorithm is used to generate candidate peptides, which are: 8mer: RLVVKGSR, LVVKGSRE, VVKGSREP; 9mer: RLVVKGSRE, LVVKGSREP, VVKGSREPV; 10mer: RLVVKGSREP, LVVKGSREPV, VVKGSREPVS.

[0028] For the candidate neoantigen KLLFGYPVYV, the system first looks for the corresponding mutation site in the patient's WES data. Assuming that the peptide segment is derived from the p.R273C mutation of the TP53 gene, the system extracts the following parameters: 1. Mutation quality score Q = 35 (Phred quality score) 2. Sequencing depth Depth = 120x (tumor) / 80x (normal) 3. Mutation type Type = Missense 4. Tumor purity Purity = 0.75.

[0029] Further, according to the mutation quality score Q, the sequencing depth Depth, the mutation type Type, and the tumor purity Purity, the genomic evidence score is calculated, and the expression is: ; wherein, represents the evidence score in the genomic dimension, represents the mutation quality score, represents the sequencing depth of the tumor sample pair, represents the mutation type weight, represents the tumor purity adjustment coefficient.

[0030] Substituting, we get, .

[0031] Further, in the process of RNA-Seq data processing, the present application adopts allele-specific expression analysis to distinguish between wild-type and mutant transcripts, and the high expression of mutant transcripts confirms the existence of the mutation at the RNA level.

[0032] The present application finds the expression of the TP53 gene in the RNA-Seq data: gene expression amount TPM = 45.2, mutant allele frequency VAF = 0.38, and tumor vs. normal expression fold change FC = 3.2.

[0033] Further, the transcriptome evidence score is: ; wherein, represents the evidence score in the transcriptome dimension, represents the gene expression amount, represents the mutant allele frequency, represents the expression fold change of tumor and normal cells. Substituting the calculation can get: ; Further, in the immunopeptidomics data, the present application analyzes the HLA-I class molecule presented peptide segments of the cancer tissue and the para-cancer tissue, as shown in Table 1: Table 1 HLA-I class molecule presented peptide segments of cancer tissue and para-cancer tissue As shown in Table 1, RLVVKGSRE and VVKGSREPV are specifically highly expressed in cancer tissues, while KITDFGLAK (wild-type peptide segment) is expressed similarly in both tissues, confirming the tumor-specific presentation of the mutant peptide segment.

[0034] Further, the present application also designs a multi-parameter fused proteome score, the expression of which is: ; wherein, represents the multi-parameter fused proteome score, represents the mass spectrometry detection score, ranging from 0 to 1, represents the mass deviation, represents the average value of the position-specific score, represents the signal intensity ratio of cancer tissue to paracancer tissue.

[0035] Substituting and calculating can obtain = 0.82 × (1 - 2.3 / 20) × 0.75 × log2(8.5) = 1.76.

[0036] Further, the present application performs cross-validation reward calculation, and the candidate antigen is detected in three omics levels, obtaining the highest cross-validation reward: ; wherein, represents the cross-validation reward score, represents the number of omics layers in which the antigen is detected, ranging from 1 to 3, represents the reward coefficient of each omics layer.

[0037] Further, the weight coefficient is dynamically adjusted according to the data quality to obtain the comprehensive score. In this example, the WES quality is high (α = 0.3), the RNA-Seq coverage is good (β = 0.25), the immune polypeptide omics signal is strong (γ = 0.35), and the cross-validation is complete (δ = 0.1), wherein the value of the comprehensive score is: ; Therefore, the comprehensive score is 1.41, which is much higher than the threshold value 0.8, and the new antigen is determined as a high-confidence candidate. In contrast, the score of the candidate antigen detected only in a single omics is usually lower than 0.5, effectively distinguishing high and low-confidence antigens.

[0038] The application applies a multi-dimensional evidence chain scoring system to calculate the value of RLVVKGSRE as 1.52, indicating that the three-dimensional evidence is complete and has high credibility, the value of VVKGSREPV is 1.38, indicating that no cancer is detected, but other evidence is strong and has high credibility, and the value of the random peptide segment AAAAAAAAAA is 0.12, indicating that there is no any evidence to support it and it can be excluded.

[0039] S3, an anchor site identification and prediction module 300, is configured to score the amino acids of each site using the amino acid frequency weighted BLOSUM62 encoding value, obtain a specific anchor site and an anchor site amino acid residue restriction set, and determine whether a certain peptide segment is a non-binding peptide according to the anchor site identification and prediction module, to obtain a determination result.

[0040] The anchor site identification and prediction module includes five steps, namely steps 1 to step 5: Step 1: data preparation and frequency matrix construction.

[0041] Specifically, 9mer peptides known to bind to HLA-A*24:02 are collected. A 20x9 position-specific amino acid frequency matrix is constructed, and the matrix element F[aa,pos] represents the frequency of amino acid aa at position pos. For example, the frequency distribution of position 2 (P2) is shown in Table 2: Table 2 Amino acid and frequency distribution table Step 2: Calculate the weighted BLOSUM62 score. For each position, the highest frequency amino acid is used as a reference to calculate the weighted BLOSUM62 score of the position.

[0042] Specifically, taking P2 as an example, the reference amino acid is Y: wBlosum(P2) = sum[freq(aa) × BLOSUM62(Y, aa)] =0.7684×7+0.1110×3+0.0222×2+0.0066×2+0.0106×(-1.0)+0.0052×(-1.0)+0.0239×(-1.0)+0.0063×(-1.0)+0.0028×(-2.0)+0.0062×(-2.0)+0.0073×(-2.0)+0.0055×(-2.0)+0.0043×(-2.0)+0.0023×(-2.0)+0.0012×(-2.0)+0.0043×(-2.0)+0.0038×(-3.0)+0.0054×(-3.0)+0.0025×(-3.0)≈5.62.

[0043] Then, repeat the calculation for all positions (9 positions in this example), and get the vector of weighted BLOSUM62 values (wBlosum) for all positions: [-0.25, 5.62, -0.54, -0.71, -0.68, -0.41, -0.27, -0.85, 2.27].

[0044] Step 3: Identify anchor positions.

[0045] Specifically, calculate the mean value of the wBlosum vector in Step 2 = 0.46. The positions with wBlosum > 0.46 are identified as anchor positions: P2 (5.62), P9 (2.27). This is highly consistent with the known binding pattern of HLA-A*24:02 (P2 and P9 are the main anchor points), verifying the accuracy of the algorithm.

[0046] Step 4: Residue classification and rule establishment.

[0047] Specifically, in Step 4, the present application classifies all amino acids at anchor position P2 according to the BLOSUM62 value with reference amino acid Y, and the classification results are as follows: Consistent residues (score ≥ 0): Y (7), F (3), W (2), H (2), with a total frequency of 0.9080; Different residues (score < 0): L (-1.0), Q (-1.0), A (-2.0), R (-2.0), M (-1.0), G (-3.0), T (-2.0), D (-3.0), I (-1.0), K (-2.0), S (-2.0), N (-2.0), P (-3.0), E (-2.0), V (-1.0), C (-2.0), etc.

[0048] Further, after classifying all amino acids, the present application establishes the following prediction rules for the classified amino acids: Prediction rule 1: P2 position must be an aromatic amino acid (Y / F / W / H) to effectively bind with HLA-A*24:02. The residue classification of P9 position is calculated according to the above algorithm; Prediction rule 2: P9 position must be one of hydrophobic amino acids (F / L / I / W / M / Y) to effectively bind with HLA-A*24:02.

[0049] Step 5: Prediction verification.

[0050] Specifically, the application uses established rules to predict new peptide segments: KGFGYPVYV: P2=G (differential residues), predicted as a non-binding peptide. KYFGYPVYF: P2=Y (consistent residues), P9=F (consistent residues), predicted as a peptide to be verified. Because even if the amino acids of the two anchor sites meet the consistent residue rule, there are other conditions and mechanisms for MHC class I molecules to present peptide segments.

[0051] Further, the application innovatively introduces the evolutionary biology BLOSUM62 substitution matrix into MHC-peptide binding prediction, and develops an anchor site recognition and residue feature analysis algorithm. The algorithm first constructs a position-specific amino acid frequency matrix, and then calculates the weighted BLOSUM62 value of each position, and the specific expression is as follows: ; Wherein, represents the weighted BLOSUM62 value of position , represents a specific position in the peptide segment, represents the reference amino acid, which is the amino acid with the highest frequency in the peptide segment position , represents the frequency of amino acid at position , represents the first amino acid, represents the first amino acid, represents the first amino acid, represents the first amino acid,

[0052] In summary, this step can be used to quickly and accurately exclude non-binding peptides, and reduce the range of peptides to be verified. At the same time, due to the lack of negative data in the current neural network training data, this step can also be used to assist in generating large-scale and accurate negative data.

[0053] S4, an intelligent MHC retrieval module 400, used for intelligent retrieval according to the judgment result, a user inputting a character string in the front end to obtain a retrieval result.

[0054] Specifically, in the intelligent MHC retrieval module, the MHC naming structures of data from different sources are different, and even after standardization, there are more than 20,000 MHC names, which are very numerous. When a user inputs a query character string in the front-end input box, non-standardized names may be input, and non-expected results may appear in the candidate box. In view of the problems of existing tools that are complicated to operate and have poor fault tolerance, the application makes related optimizations, and the specific process is as follows: Maintenance of standard MHC nomenclature index: The system backend maintains a complete and standard MHC nomenclature index table, which is regularly and automatically updated based on the authoritative IMGT / HLA database, ensuring the accuracy and timeliness of all MHC allele information.

[0055] Simplification of standard MHC nomenclature index: Remove the “-”, “*”, “:”, “^” symbols from each standard nomenclature, and uniformly process them into uppercase state, and generate a dictionary corresponding to the standard nomenclature.

[0056] Locus name variant compatibility rules: Based on the previous step, add the following dictionary to ensure that important special abbreviations of loci can be retrieved: HLA-DPA1 is abbreviated as DPA or DP, HLA-DPB1 is abbreviated as DPB or DP, HLA-DQA1 is abbreviated as DQA or DQ, HLA-DQB1 is abbreviated as DQB or DQ, HLA-DRB1 is abbreviated as DRB or DR, to ensure comprehensive retrieval.

[0057] User input string processing: When the user inputs the query string in the front-end input box, the system does not directly perform database matching. First, the simplification engine will analyze and standardize the input. Identify and replace non-standard separators that users may use, such as spaces, underscores, hyphens, etc. Replace special strings including “-”, “*”, “:”, “^” symbols into empty strings, and then convert all strings to uppercase state.

[0058] Common input error automatic correction rules: Based on a preset common spelling error knowledge base, the input is intelligently corrected. For example, when the HLA class I locus A, B, C is mistakenly input as D, the system can prompt the user whether it is intended to query the class II locus DR, DQ or DP.

[0059] Further, due to the large number of MHC molecules, after generating the candidate list, there may be many choices, which reduces work efficiency. The present application generates a multi-parameter dynamic weight system to intelligently sort the results instead of alphabetical order, with the following logic: species weight: human HLA molecules have the highest weight, followed by frequently used experimental animals such as mice, followed by other primates, and then other mammals. Ensure that the most relevant results for clinical and scientific research are at the forefront. Locus weight: in human HLA, different loci are given different weights according to clinical importance and research popularity, for example: HLA-A > HLA-B > HLA-C > HLA-DP > HLA-DQ > HLA-DR > others. In mouse MHC, the weights of the loci are H-2K > H-2D > H-2L > H-2A > H-2E > others. Frequency weight: the system stores data from the global Allele Frequencies Net Database, dynamically adjusts the weight according to the frequency of the target HLA in the global or specific population, and ranks the high-frequency alleles in the front. Therefore, when the user inputs "1101", the top of the returned list will be high-frequency human HLA molecules such as "HLA-A*11:01", rather than some rare, non-human HLA molecules, greatly improving user experience and research efficiency.

[0060] Embodiment 2 See Figure 7 The second embodiment of the present application also provides a tumor neoantigen precise analysis method based on multi-dimensional evidence chain scoring, comprising: evaluating the feasibility score of the candidate neoantigen; identifying and verifying tumor-specific antigens according to the feasibility score; scoring the amino acids at each site using amino acid frequency weighted BLOSUM62 encoding values, obtaining specific anchor sites and anchor site amino acid residue restricted sets, and judging whether a peptide segment is a non-binding peptide based on the same to obtain a judgment result; according to the judgment result, the user inputs a string in the front end for intelligent retrieval to obtain a retrieval result.

[0061] The tumor neoantigen precise analysis system based on multi-dimensional evidence chain scoring in the embodiment of the application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device, or a non-mobile electronic device. Exemplarily, the mobile electronic device can represent a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), and the like, and the non-mobile electronic device can represent a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, and the like, which are not specifically limited in the embodiment of the application.

[0062] The tumor neoantigen precise analysis system based on multi-dimensional evidence chain scoring in the embodiment of the application can represent a device with an operating system. The operating system can represent an Android operating system, an ios operating system, or other possible operating systems, which are not specifically limited in the embodiment of the application.

[0063] The tumor neoantigen precise analysis method based on multi-dimensional evidence chain scoring provided in the embodiment of the application can achieve Figures 1 to 6 The processes achieved by the tumor neoantigen precise analysis system based on multi-dimensional evidence chain scoring in the method embodiment are not repeated here to avoid repetition.

[0064] Optionally, the embodiment of the application further provides an electronic device, which includes a processor, a memory, a program or instructions stored on the memory and executable on the processor. When the program or instructions are executed by the processor, each process of the tumor neoantigen precise analysis method based on multi-dimensional evidence chain scoring is achieved, and the same technical effects are achieved. To avoid repetition, each process is not repeated here.

[0065] The embodiment of the application further provides a readable storage medium, which stores a program or instructions. When the program or instructions are executed by a processor, each process of the tumor neoantigen precise analysis method based on multi-dimensional evidence chain scoring is achieved, and the same technical effects are achieved. To avoid repetition, each process is not repeated here.

[0066] The processor represents the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0067] It should be noted that in the present application, the terms "comprising", "containing" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus that includes a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or apparatus. Without more limitations, the element defined by the statement "including a" does not exclude the presence of additional identical elements in the process, method, article or apparatus including the element. In addition, it should be pointed out that the scope of the methods and apparatuses in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but can also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method can be performed in an order different from that described, and various steps can also be added, omitted or combined. In addition, the features described with reference to certain examples can be combined in other examples.

[0068] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and a necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal (which can be a mobile phone, computer, server, air conditioner or network device, etc.) execute the method described in each embodiment of the present application.

[0069] The embodiments of the present application have been described above with reference to the accompanying drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are only illustrative, not limiting. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope of protection of the claims, which are all within the scope of protection of the present application.

Claims

1. A precise analysis system for tumor neoantigens based on multidimensional evidence chain scoring, characterized in that, include: A multidimensional evidence chain scoring module is used to evaluate the feasibility score of candidate neoantigens; An antigen extraction and verification module is used to identify and verify tumor-specific antigens based on the feasibility score. The anchor site identification and prediction module is used to score the amino acids at each site using the amino acid frequency-weighted BLOSUM62 encoding value, obtain the specific anchor site and the restriction set of amino acid residues at the anchor site, and determine whether a peptide is a non-binding peptide based on this, so as to obtain the judgment result. The intelligent MHC retrieval module is used to perform intelligent retrieval based on the judgment result, allowing the user to input a string on the front end to obtain the retrieval result.

2. The tumor neoantigen precision analysis system based on multidimensional evidence chain scoring according to claim 1, characterized in that, The specific steps for evaluating the feasibility score of candidate neoantigens include: Genomic evidence scores are obtained based on the mutation quality, sequencing depth, and mutation type of candidate neoantigens. Transcriptome-level evidence scores are obtained based on gene expression levels and mutation allele frequencies. Evidence scores for proteomics dimensions were obtained based on fusion mass spectrometry. Cross-validation reward scores are obtained when the candidate neoantigen is detected simultaneously at multiple omics levels. The feasibility score of the candidate neoantigen is obtained by weighting the evidence scores of the genome dimension, the transcriptome dimension, the proteome dimension, and the cross-validation reward score.

3. The tumor neoantigen precision analysis system based on multidimensional evidence chain scoring according to claim 1, characterized in that, The specific steps of scoring the amino acids at each site using amino acid frequency-weighted BLOSUM62 coding values ​​to obtain the specific anchoring site and the restriction set of amino acid residues at the anchoring site, and determining whether a peptide is a non-binding peptide based on this, include: MHC and peptide length were used as the basis for grouping MHC-peptide data; Based on the grouping criteria, the amino acid with the highest frequency at each position of the peptide segment is used as the reference amino acid. Obtain the BLOSUM62 encoding value of each amino acid position in the peptide relative to the reference amino acid; The BLOSUM62 encoding value is multiplied by the frequency of occurrence of the amino acid to obtain the BLOSUM62 weighted value for each site of the peptide. Locations with a weighted BLOSUM62 score higher than the average score of the weighted BLOSUM62 scores of all peptide locations are used as anchor points. At each anchor point, all amino acids are classified into consistent residues and differential residues based on whether the BLOSUM62 value is greater than 0, thus establishing a restricted set of amino acid residues at the anchor point. A knowledge graph is built based on specific anchor sites and the restriction set of amino acid residues at the anchor sites, and this graph is used to determine whether a peptide is a non-binding peptide.

4. The tumor neoantigen precision analysis system based on multidimensional evidence chain scoring according to claim 1, characterized in that, The antigen extraction and validation module includes WES data processing, RNA-Seq data processing, and immunopeptide omics data processing. The WES data processing procedure includes: The corresponding mutation sites are located in the WES data to obtain mutation quality scores, sequencing depth of tumor samples, mutation type weights, and tumor purity adjustment coefficients, thereby obtaining a genomic evidence score. This genomic evidence score includes: ; in, Score the genomic evidence. Indicates the quality score of the mutation. Indicates the sequencing depth of the tumor sample. Indicates the mutation type weight. Indicates the tumor purity adjustment factor; The RNA-Seq data processing procedure includes: obtaining the gene expression level, mutant allele frequency, and fold change in expression between tumor and normal cells for the TP53 gene from the RNA-Seq data, and then calculating the transcriptome evidence score. The transcriptome evidence score includes: ; in, Transcriptome evidence score, Indicates gene expression level. Indicates the frequency of mutated alleles. Indicates the fold increase in expression between tumor cells and normal cells; The process of processing the immunopeptidomics data includes: obtaining a proteome score based on the mass spectrometry detection score, mass deviation, location specificity score, and the signal intensity ratio of cancerous tissue to adjacent normal tissue, wherein the proteome score includes: ; in, Scoring the proteome This indicates the mass spectrometry detection score. Indicates quality deviation. This represents the average of the location-specific scores. This represents the ratio of signal intensity between cancerous tissue and adjacent tissue.

5. The tumor neoantigen precision analysis system based on multidimensional evidence chain scoring according to claim 1, characterized in that, The intelligent MHC retrieval module includes: The backend standard MHC naming index module is used to maintain and periodically update the complete standard MHC naming index table based on an authoritative database, and to perform semantic understanding and standardization processing on the query strings entered by users. The naming simplification module is used to remove the symbols "-", "*", ":", and "^" from standard MHC names and convert them to uppercase letters. It generates a dictionary that corresponds to the standard names and is compatible with important abbreviations of loci such as HLA-DPA1, HLA-DPB1, HLA-DQA1, HLA-DQB1, and HLA-DRB1. The user input processing module is used to receive the query string input by the user, perform non-standard symbol replacement, uppercase conversion and intelligent correction of common spelling errors to generate a standardized query string; The multi-parameter dynamic weighted sorting module is used to intelligently sort the candidate MHC molecule list according to species weight, locus weight, and allele frequency weight, so as to present high-frequency and research or clinically relevant MHC molecules to users first.

6. The tumor neoantigen precision analysis system based on multidimensional evidence chain scoring according to claim 5, characterized in that, The specific process of ranking species by weight includes: ranking different species according to their relevance to clinical and scientific research, with the ranking results of species weight from first to last being: humans, mice, other primates, and other mammals; The specific process of ranking the weights of the loci includes: assigning different weights to different loci based on their clinical importance and research popularity. In humans, the weights of the loci, from highest to lowest, are: HLA-A > HLA-B > HLA-C > HLA-DP > HLA-DQ > HLA-DR > others. In mouse MHC, the weights of the loci, from largest to smallest, are: H-2K > H-2D > H-2L > H-2A > H-2E > others.

7. The tumor neoantigen precision analysis system based on multidimensional evidence chain scoring according to claim 1, characterized in that, The multidimensional evidence chain scoring module calculates the comprehensive feasibility score of the candidate neoantigen by dynamically weighting and summing the evidence scores from the genome dimension, transcriptome dimension, proteome dimension, and cross-validation reward scores.

8. The tumor neoantigen precision analysis system based on multidimensional evidence chain scoring according to claim 1, characterized in that, The anchor point identification and prediction module further constructs a knowledge graph based on the identified specific anchor points and amino acid residue restriction sets, which is used to quickly exclude peptides that are not bound to major histocompatibility complex class I molecules.

9. A precise analysis method for tumor neoantigens based on multidimensional evidence chain scoring, characterized in that, include: Assess the feasibility score of candidate neoantigens; Based on the feasibility score, tumor-specific antigens are identified and validated. The amino acids at each site are scored using the frequency-weighted BLOSUM62 coding value to obtain the specific anchoring site and the restriction set of amino acid residues at the anchoring site. Based on this, it is determined whether a peptide is a non-binding peptide to obtain the judgment result. Based on the judgment result, the user enters a string on the front end to perform intelligent retrieval and obtain the search results.