Prediction and application of breast cancer public neoantigen
Patent Information
- Application Number
- CN202510376570.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
Smart Images

Figure BDA0005334323030000121 
Figure HDA0005334323250000011 
Figure HDA0005334323250000012
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technologies, and particularly to the prediction and application of common neoantigens in breast cancer. Background Art
[0002] Breast cancer is a common malignant tumor in women, originating from the uncontrolled proliferation of breast epithelial cells. It is a disease with high heterogeneity at the molecular level. According to the expression levels of ER, PR, HER-2, and Ki-67 in the immunohistochemical indicators of tumor tissues, the molecular subtypes of breast cancer are classified into Luminal A type, Luminal B type, HER-2 overexpression type, and triple-negative breast cancer.
[0003] The molecular markers of different subtypes of breast cancer are different, and the prognoses are also different. Luminal A type has positive estrogen receptor (ER) and progesterone receptor (PR), low expression of Ki-67, and a better prognosis; Luminal B type has positive ER and PR but high Ki-67 or positive HER-2; HER-2 overexpression type has positive HER-2; triple-negative has negative ER, PR, and HER-2, with a high degree of malignancy, a poor prognosis, and different focuses in treatment methods. Determining the molecular subtype of breast cancer is of great significance for subsequent treatment.
[0004] Dendritic cells (DC) can efficiently participate in antigen recognition, uptake, processing, and presentation, inducing the body's immune response. Immature DCs have strong migration ability and antigen capture and processing ability. After DC cells are stimulated by antigens, they mature and migrate to secondary lymphoid organs, bind to T cells with a variety of cell surface proteins, cross-present the antigen protein epitope peptides captured to T cells, and promote their differentiation into antigen-specific cytotoxic T lymphocytes, thereby playing a cellular immune process to recognize and degrade antigenic substances.
[0005] Cancer neoantigens are the products of somatic mutations. These cancer-specific antigens can come from somatic cell mutations (such as point mutations, insertions and deletions, open reading frame changes, etc.). Gene mutations and their products that are commonly present in cancer cells are not present in healthy cells and normal genomes. They are not only specific cancer antigens but also highly immunogenic, meeting the ideal conditions for constructing cancer vaccines.
[0006] With the increasing improvement of the new generation of gene sequencing technology, it is currently possible to identify and characterize a large number of mutated neoantigens. Through gene sequencing of cancer tissues and combined with the prediction of the binding affinity of human leukocyte antigen (HLA) epitopes, it is possible to identify, evaluate, and select candidate cancer neoantigens for patients. Therefore, antigen epitope prediction based on mutated neoantigens will facilitate the development of new anti-breast cancer drugs or the formulation of new immunotherapy strategies. Summary of the Invention
[0007] In view of this, the technical problem to be solved by the present invention is to provide breast cancer antigens and their applications.
[0008] The present invention provides the application of a mutation site as an epitope in the preparation of a breast cancer diagnostic reagent or a drug for treating breast cancer; the mutation site is at least one of PIK3CA p.H1047R, AKT1 p.E17K, KMT2C p.K2797Q, and / or TP53 p.R248Q.
[0009] In the present invention, the breast cancer includes Luminal A type, Luminal B type, HER-2 overexpression type, and triple-negative breast cancer.
[0010] In the present invention, tissue samples of breast cancer patients to be predicted are collected, DNA of the tumor tissue samples and normal tissue samples are respectively extracted, and whole-genome sequencing is respectively performed. The HLA genes of the patients are typed according to the sequencing data of the normal tissue samples. The DNA sequencing data of the tumor tissue samples and normal samples are subjected to quality inspection, low-quality data are removed, and then the filtered data are aligned with the human reference genome GRCh38 version. Then, GATK4 Mutect2 is used to analyze somatic mutations of tumor-normal paired samples, ANNOVAR is used to annotate the mutation sites, and SAMtool is used to obtain the mutant sequences and their flanking sequences. The mutated genes of Luminal A type, Luminal B type, HER-2 overexpression type, and triple-negative breast cancer patients are respectively counted, sorted according to the mutation frequency of the genes, and finally the genes with the highest mutation frequencies in the 4 molecular subtypes are PIK3CA p.H1047R, AKT1 p.E17K, KMT2C p.K2797Q, and TP53 p.R248Q in sequence.
[0011] The mutation site PIK3CA p.H1047R means that the 1047th H in PIK3CA is mutated to R;
[0012] The mutation site AKT1 p.E17K means that the 17th E in AKT1 is mutated to K;
[0013] The mutation site KMT2C p.K2797Q is the mutation of the 2797th K to Q in KMT2C;
[0014] The mutation site TP53 p.R248Q is the mutation of the 248th R to Q in TP53.
[0015] The mutation sites obtained by screening and / or the flanking sequences of the mutation sites can be used for diagnosing breast cancer, especially for classifying breast cancer. Or the mutation sites and / or the flanking sequences of the mutation sites can be used as common neoantigens of breast cancer tumors with different molecular classifications, so as to be used for preparing drugs for treating breast cancer.
[0016] The present invention also provides an epitope peptide, which contains 19-31 amino acid residues and contains at least one of the mutation sites of PIK3CA p.H1047R, AKT1 p.E17K, KMT2C p.K2797Q and / or TP53 p.R248Q.
[0017] The number of amino acid residues in the epitope peptide of the present invention is 19-31, including the mutation sites as described above. In some embodiments, the number of amino acid residues in the epitope peptide is 21-29; more specifically, the number of amino acid residues in the epitope peptide is 23-27. For example, the number of amino acid residues in the epitope peptide is 23, 24, 25, 26 or 27.
[0018] The N-terminal side of the mutation site contains 5-20 amino acid residues, preferably 10-15 amino acids. For example, it contains 10 amino acid residues, 11 amino acid residues, 12 amino acid residues, 13 amino acid residues, 14 amino acid residues or 15 amino acid residues.
[0019] The C-terminal side of the mutation site contains 5-20 amino acid residues, preferably 10-15 amino acids. For example, it contains 10 amino acid residues, 11 amino acid residues, 12 amino acid residues, 13 amino acid residues, 14 amino acid residues or 15 amino acid residues.
[0020] The epitope peptide contains one, two, three or four mutation sites. Screening is carried out according to the affinity of the peptide segment to the HLA molecule. Preferably, the epitope peptide in the present invention has the amino acid sequence as described in any one of SEQ ID NO:1-4. In specific embodiments, the amino acid sequence of the epitope peptide provided by the present invention is as shown in SEQ ID NO:1, or as shown in SEQ ID NO:2, or as shown in SEQ ID NO:3, or as shown in SEQ ID NO:4.
[0021] Furthermore, the present invention also provides an antigen, which comprises at least two of the epitope peptides described above.
[0022] In some embodiments, the antigen of the present invention comprises two epitope peptides, for example:
[0023] It comprises the epitope peptide shown in SEQ ID NO:1 and the epitope peptide shown in SEQ ID NO:2;
[0024] or comprises the epitope peptide shown in SEQ ID NO:1 and the epitope peptide shown in SEQ ID NO:3;
[0025] or comprises the epitope peptide shown in SEQ ID NO:1 and the epitope peptide shown in SEQ ID NO:4;
[0026] or comprises the epitope peptide shown in SEQ ID NO:2 and the epitope peptide shown in SEQ ID NO:3;
[0027] or comprises the epitope peptide shown in SEQ ID NO:2 and the epitope peptide shown in SEQ ID NO:4;
[0028] or comprises the epitope peptide shown in SEQ ID NO:3 and the epitope peptide shown in SEQ ID NO:4.
[0029] In some embodiments, the antigen of the present invention comprises three epitope peptides, for example:
[0030] It comprises the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2 and the epitope peptide shown in SEQ ID NO:3;
[0031] or comprises the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2 and the epitope peptide shown in SEQ ID NO:4;
[0032] or comprises the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:3 and the epitope peptide shown in SEQ ID NO:4;
[0033] or comprises the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:3 and the epitope peptide shown in SEQ ID NO:4.
[0034] In some embodiments, the antigen of the present invention comprises four epitope peptides, which comprises the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:3 and the epitope peptide shown in SEQ ID NO:4.
[0035] In the present invention, the connection order of the epitope peptides in the antigen is not limited. In the present invention, the epitope peptides in the antigen can be directly connected or can contain a spacer sequence, and the present invention does not limit this. For example, the spacer sequence can be AGA. In the specific embodiments of the present invention, there is no spacer sequence between the epitope peptides.
[0036] In specific embodiments, the antigen provided by the present invention is, from the N-terminus to the C-terminus, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:3, and the epitope peptide shown in SEQ ID NO:4 in sequence. More specifically, the amino acid sequences of the antigen are SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, and SEQ ID NO:8 respectively.
[0037] Furthermore, the present invention also provides a fusion protein, which comprises a chemokine and the epitope peptide as described above, or comprises a chemokine and the antigen as described above.
[0038] The fusion protein of the present invention contains at least one chemokine. The chemokines contained in the fusion protein of the present invention include but are not limited to at least one of CCL1, CCL2, CCL3, CCL4, CCL5, CCL13, CCL19, CCL21, CXCL3, CXCL9, CXCL10, CXCL12, CXCL13, CXCL14, CXCL15, CXCL16, or CXCL17. In the embodiments of the present invention, the chemokine is selected from at least one of CCL1, CCL13, and / or CXCL3.
[0039] The fusion protein of the present invention contains at least one epitope peptide and / or at least one antigen. In the fusion protein of the present invention, the position of the chemokine is not limited. Preferably, the chemokine in the fusion protein of the present invention is located at the N-terminus of the fusion protein.
[0040] The fusion protein of the present invention can be used for the construction of DC cell tumor vaccines, can improve the killing ability and targeting ability of immune cells against breast cancer cells, and thus better play an anti-tumor role. The vaccine of the present invention is a protein vaccine, an mRNA vaccine, or a DNA vaccine, and the present invention does not limit this.
[0041] In some embodiments, the fusion protein is, from the N-terminus to the C-terminus, CCL1, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:3, and the epitope peptide shown in SEQ ID NO:4 in sequence;
[0042] Or in sequence as CCL1, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:3;
[0043] Or in sequence as CCL1, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:4;
[0044] Or in sequence as CCL1, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:2;
[0045] Or in sequence as CCL1, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:3;
[0046] Or in sequence as CCL1, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:2;
[0047] Or in sequence as CCL1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:4;
[0048] Or in sequence as CCL1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:3;
[0049] Or in sequence as CCL1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:4;
[0050] Or in sequence as CCL1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:1;
[0051] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:3;
[0052] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:1;
[0053] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:4;
[0054] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:2;
[0055] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:4;
[0056] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:1;
[0057] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2;
[0058] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:1;
[0059] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:3;
[0060] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:2;
[0061] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:3;
[0062] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:1;
[0063] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2;
[0064] Or in sequence CCL1, the epitope peptide shown in SEQ ID NO:4, the epitope peptide shown in SEQ ID NO:3, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:1.
[0065] Furthermore, the present invention also provides a nucleic acid, which comprises at least one of the following:
[0066] I), a nucleic acid encoding the epitope peptide described above;
[0067] II), a nucleic acid encoding the antigen described above;
[0068] III), a nucleic acid encoding the fusion protein described above.
[0069] In the present invention, the nucleic acid can be either DNA or RNA, and the present invention does not make any limitation thereto. The nucleic acid can contain only the coding region or can also contain non-coding regions (such as regulatory sequences, where the regulatory sequences include promoters or transcription terminators, etc.). Specifically, the nucleic acid is DNA, which is cDNA or gDNA or artificially synthesized DNA. The nucleic acid described in the present invention can be single-stranded or double-stranded; from a topological perspective, the morphology of the nucleic acid can be linear or circular. The nucleic acid can also exist in various forms, for example, it can be a part of a vector (such as an expression vector or a cloning vector), or just a nucleic acid fragment. The nucleic acid described in the present invention can be directly obtained from natural sources or can also be prepared with the assistance of relevant means such as recombination, enzymatic methods or chemical techniques.
[0070] In a specific embodiment, in the nucleic acid,
[0071] The nucleic acid sequence encoding the epitope peptide shown in SEQ ID NO: 1 is as shown in SEQ ID NO: 5;
[0072] The nucleic acid sequence encoding the epitope peptide shown in SEQ ID NO: 2 is as shown in SEQ ID NO: 6;
[0073] The nucleic acid sequence encoding the epitope peptide shown in SEQ ID NO: 3 is as shown in SEQ ID NO: 7;
[0074] The nucleic acid sequence encoding the epitope peptide shown in SEQ ID NO: 4 is as shown in SEQ ID NO: 8.
[0075] Furthermore, the present invention also provides an mRNA, which includes: 5’Cap, the nucleic acid as described above, and Poly-A Tail.
[0076] As a feasibility case, 5’Cap is 7-methylguanosine cap (m 7 G cap), anti-reverse cap analog (ARCA), 3'-O-methyl-m 7 G(5')ppp(5')G (abbreviation Cap 0), m 7 G(5')ppp(5')m 2 ′, 7 G (Cap 1), m 7 G(5')ppp(5')m 2 ′, 7 Gm 2 ′, 7 G (Cap 2), AG (a commercial capping analog), SG, N 7 -methylguanosine-triphosphate (N 7 -methylguanosine triphosphate), N 6 -methyladenosine-triphosphate (N 6 -methyladenosine triphosphate), β-S-adenosylmethionine (β-S-adenosylmethionine, which can participate in some capping-related synthesis processes).
[0077] As a feasibility case, the Poly-A Tail contains a polyA tail structure with 20 to 350 adenylate residues. For example, the number of adenylate residues is 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 180, 200, 250, or 300.
[0078] As a feasibility case, the mRNA may also include modified amino acids. For example, pseudouridine (Ψ) modification, N 6 -methyladenosine (m 6 A) modification, or 5-methylcytidine (m 5 C) modification.
[0079] Furthermore, the present invention also provides an expression unit, which comprises the nucleic acid and a promoter as described above.
[0080] In the present invention, the promoter is a prokaryotic promoter or a eukaryotic promoter, and the present invention does not limit this. For example, the promoter may be an araBAD promoter, a CaMV 35S promoter, a CMV promoter, an H1 promoter, a lac promoter, an SP6 promoter, an SV40 promoter, a T7 promoter, a U6 promoter, or a trp promoter.
[0081] In addition to the promoter, the expression unit of the present invention may further include at least one of an enhancer, a transcription start site, a poly(A) signal, or a terminator. The terminator includes at least one of an rpsE terminator, an rpsF terminator, an rpsD terminator, an rpoB terminator, an rplT terminator, an rpsL terminator, a rho-independent terminator, a rho-dependent terminator, a trpA terminator, a lacZ terminator, a CMV terminator, an SV40 terminator, an SP6 terminator, a T4 terminator, a T7 terminator. The enhancer includes at least one of an ACT5C enhancer, a CAGG enhancer, a COPIA enhancer, an EF1A enhancer, an EF1α enhancer, a HARE5 enhancer, a PGK enhancer, a ROSA26 enhancer, an SV-1 enhancer, an SV40 enhancer, a CMV enhancer, a UBC enhancer.
[0082] Furthermore, the present invention also provides a plasmid vector, which comprises the nucleic acid or the expression unit as described above.
[0083] The plasmid vector provided by the present invention is used for the storage and amplification of the nucleic acid or the expression unit, or the expression of the epitope peptide, antigen or fusion protein, and the present invention does not limit this. In specific embodiments, the backbone vector of the plasmid vector is the pColdI vector. In addition, it can also be the pColdII vector, pColdIII vector, pColdTF vector, pET series vectors, pGEX series vectors, pMAL series vectors, pBAD vector, pBADHis vector, pBADmycHis series vectors, pQE series vectors, pTrc99a vector, pTrcHis series vectors, pBV220 vector, pBV221 vector, pBV222 vector, pTXB series vectors, pLLP-ompA vector, pIN-III-ompA vector, pQBI63 vector or pACYCduet-1 vector.
[0084] Furthermore, the present invention also provides a host,
[0085] which is transformed or transfected with the plasmid vector as described above;
[0086] or the nucleic acid as described above is integrated into its genome;
[0087] or the expression unit as described above is integrated into its genome.
[0088] In the present invention, the host is used for the storage and amplification of the plasmid vector or the expression of the epitope peptide, antigen or fusion protein. In specific embodiments, it is 293T cells. Or it can also be CHO cells, COS-7 cells, SF9 cells or Hela cells.
[0089] Furthermore, the present invention also provides a preparation method for the epitope peptide, antigen or fusion protein as described above, which includes culturing the host as described above to obtain an expression product containing the epitope peptide, antigen or fusion protein.
[0090] Furthermore, the present invention also provides a preparation method for the mRNA as described above, which includes culturing the host to obtain template DNA, and through in vitro transcription, capping reaction and tailing reaction, obtaining the mRNA.
[0091] Furthermore, the present invention also provides the application of the epitope peptide, antigen, fusion protein, nucleic acid, mRNA, expression unit, plasmid vector or host as described above in the preparation of drugs for treating breast cancer.
[0092] Furthermore, the present invention also provides a drug for treating breast cancer, which comprises: the epitope peptide as described above, the antigen as described above, the fusion protein as described above, the nucleic acid as described above, the mRNA as described above, the expression unit as described above, the plasmid vector as described above, or the host as described above.
[0093] Furthermore, the present invention also provides a vaccine for treating breast cancer, which comprises the mRNA as described above and an in vivo transfection reagent.
[0094] Furthermore, the present invention also provides a method for treating breast cancer, which comprises administering the drug as described above. The present invention first proposes a method for predicting common immunogenic peptide segments in breast tumors of different molecular subtypes based on the full mining of tumor DNA data, thereby realizing the possibility of providing a common new target for breast cancer immunotherapy and improving the survival rate and prognosis of breast cancer patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. The drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the drawings.
[0096] Figure 1 It is a flowchart for predicting common new antigens of breast cancer in an embodiment of the present invention;
[0097] Figure 2 Showing the mouse immunization strategy;
[0098] Figure 3 Showing the tumor growth curves of each group;
[0099] Figure 4 Showing the quantitative statistics of the number of spots after plate imaging. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0100] The present invention provides breast cancer antigens and their applications. Those skilled in the art can draw on the content of this article and appropriately improve the process parameters to achieve. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art, and they are all considered to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments. Relevant personnel can obviously make changes or appropriate changes and combinations to the methods and applications in this article without departing from the content, spirit and scope of the present invention to implement and apply the technical solutions of the present invention.
[0101] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. For definitions and terms in this field, those skilled in the art can specifically refer to Current Protocols in Molecular Biology (Ausubel). The abbreviations of amino acid residues are the standard three-letter and / or one-letter codes used in the art to refer to one of the 20 common L-amino acids.
[0102] In the present invention, "comprising", "including" and "having" are used interchangeably and are intended to indicate the inclusiveness of the solution, meaning that the solution may contain other elements in addition to the listed elements. At the same time, it should be understood that using "comprising", "including" and "having" to describe in this article also provides the solution of "consisting of...".
[0103] In the present invention, when "and / or" is used herein, it includes the meanings of "and", "or" and "any other combination of all or part of the elements linked by the term".
[0104] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items.
[0105] In the present invention, the method for predicting common tumor neoantigens in breast cancer mainly includes the following steps: obtaining tumor samples and corresponding healthy tissue samples of breast cancer patients with different subtypes, respectively extracting DNA to construct a library and performing whole-genome DNA sequencing; typing the HLA genes of the patients; obtaining the genes with the highest mutation frequency in the above 4 subtypes of breast cancer through whole-genome and mutation-related analysis; predicting the antigenic epitopes of the polypeptide fragments of the mutant genes.
[0106] In the embodiments of the present invention, the amino acid sequences of the fragments involved and the encoded nucleic acid fragments are shown in Table 1:
[0107] Table 1 Amino acid sequences of fragments and encoded nucleic acid fragments
[0108] SEQ ID NO:1 Amino acid sequence of PIK3CA p.H1047R polypeptide SEQ ID NO:2 Amino acid sequence of AKT1 p.E17K polypeptide SEQ ID NO:3 Amino acid sequence of KMT2C p.K2797Q polypeptide SEQ ID NO:4 Amino acid sequence of TP53 p.R248Q polypeptide SEQ ID NO:5 Nucleotide sequence encoding PIK3CA p.H1047R polypeptide SEQ ID NO:6 Nucleotide sequence encoding AKT1 p.E17K polypeptide SEQ ID NO:7 Nucleotide sequence encoding KMT2C p.K2797Q polypeptide SEQ ID NO:8 Nucleotide sequence encoding TP53 p.R248Q polypeptide
[0109] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. Some or all of the steps can be executed in parallel or sequentially. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The test materials used in the present invention are all ordinary commercially available products and can be purchased in the market.
[0110] The present invention will be further described below in conjunction with embodiments:
[0111] Example 1
[0112] As Figure 1 shown, a method for predicting and applying common neoantigens of breast cancer tumors with different molecular subtypes, the method comprising the following steps:
[0113] (1) Collection of sample data:
[0114] Extract the DNA of breast cancer tumor tissue samples and corresponding normal tissue samples from 20 Luminal A-type, 23 Luminal B-type, 25 HER-2 overexpressing-type, and 20 triple-negative breast cancer patients respectively, and perform whole-genome sequencing respectively.
[0115] (2) HLA genotyping:
[0116] Use the BWA software (bwa men–a / iedb / HLA_ABCCDS.fasta raw_data_R1_fastq raw_data_R2_fastq>data_HLA.sam) to align the sequencing data of the quality-controlled normal samples to the HLA allele reference sequences in the Immune epitope database (IEDB), and then use HLAminer() to perform allele typing on the A, B, and C genotypes of HLA class I.
[0117] (3) Detection of mutation sites in breast cancer samples:
[0118] Perform quality inspection on the DNA sequencing data of the tumor tissue sample and normal sample. Use fastp software (fastp--qualified_quality_phred 5--unqualified_percent_limit 50--n_base_limit 15--overlap_len_require 30--overlap_diff_limit 1--overlap_diff_percent_limit 10--length_limit 150--trim_poly_g--thread 8 -i raw_data_R1_fastq–I raw_data_R2_fastq -o data_R1.clean.fastq -O data_R2.clean.fastq) to remove low-quality data. Then use Hisat2 (hisat2 -t--dta -p 24 -x grch38 / genome.fa -1 data_R1.clean.fastq -2 data_R2.clean.fastq–S. / his_results / data.sam--un-conc-gz. / his_results / data.unmap.fq.gz 2>data.align.log) to align the filtered data with the human reference genome GRCh38 version. Use samtools (samtools view -b*.sam>*.bam) software to convert the sam-format file obtained from the above alignment to bam format. Then use GATK4 Mutect2 (gatk--java-options "-Xmx20G" Mutect2 -R reference.fa -I tumor.bam -I normal.bam–normal normal_sample_name--germline-resource af-only-gnomad.vcf.gz--panel-of-normals pon.vcf.gz--f1r2-tar-gz f1r2.tar.gz -O somatic.vcf.gz) to analyze somatic mutations in tumor-normal paired samples.
[0119] (4) Annotation of mutation sites in breast cancer samples:
[0120] Annotate the mutation sites of each of the above - obtained samples using ANNOVAR (table_annovar.pl somatic.vcf humandb / - buildver hg38 - out anno.out - remove - protocol refGene,cytoBand,exac03,avsnp147,dbnsfp30a - operation g,r,f,f,f - nastring. - vcfinput - polish) to obtain the breast cancer gene mutation maps for each patient respectively.
[0121] (5) Statistics and ranking of the mutation gene frequencies of different molecular subtypes of breast cancer:
[0122] Statistically analyze the mutation frequencies of the mutated genes in their respective groups for patients with Luminal A type, Luminal B type, HER - 2 over - expression type, and triple - negative breast cancer respectively:
[0123] Table 1 Mutation gene frequencies of different molecular subtypes of breast cancer (1%)
[0124]
[0125] Rank the mutated genes in each group according to the gene mutation frequency. Finally, the top 4 genes (greater than or equal to 2%) with the highest mutation frequencies in the 4 subtypes are PIK3CA p.H1047R, AKT1 p.E17K, KMT2C p.K2797Q, and TP53 p.R248Q in turn.
[0126] (6) Construct polypeptide sequences for the selected mutated genes above:
[0127] Include the wild - type and mutant amino acid FASTA sequences containing the mutation sites of PIK3CA p.H1047R, AKT1 p.E17K, KMT2C p.K2797Q, and TP53 p.R248Q screened in the previous step. For the FASTA sequences, 12 amino acids are retained before and after the mutated amino acid. If the mutation is at the head or tail of the transcript, 25 amino acids are respectively intercepted after the head or before the tail to construct the FASTA sequence.
[0128] (7) Predict neo - antigens of tumors for the selected mutated genes above
[0129] Use the netMHCpan-4.1 software to predict antigenic polypeptides with high affinity for the selected mutant gene polypeptide sequence through the netMHC algorithm (netMHCpan -f input.fasta -BA -xls -a HLA.file –s –l 8,9,10,11,12 -xlsfile my_NetMHCpan_breast_out.xls).
[0130] PIK3CA p.H1047R antigen nucleic acid sequence: gcgctggaatattttatgaaacagatgaacgatgcgcg ccatggcggctggaccaccaaaatggattggattttt (SEQ ID NO:5)
[0131] PIK3CA p.H1047R antigen amino acid sequence: ALEYFMKQMNDARHGGWTTKM DWIF (SEQ ID NO:1)
[0132] AKT1 p.E17K antigen nucleic acid sequence: gcgattgtgaaagaaggctggctgcataaacgcggcaaatat attaaaacctggcgcccgcgctattttctgctg (SEQ ID NO:6)
[0133] AKT1 p.E17K antigen amino acid sequence: AIVKEGWLHKRGKYIKTWRPRYFLL (SEQ ID NO:2)
[0134] KMT2C p.K2797Q antigen nucleic acid sequence: ctggataaccagtgcgtgagcgtggaaccgaaaaaac aggaacaggaaaacaaaaccctggtgctgagcgataaa (SEQ ID NO:7)
[0135] KMT2C p.K2797Q antigen amino acid sequence: LDNQCVSVEPKKQEQENKTLVLS DK (SEQ ID NO:3)
[0136] TP53 p.R248Q antigen nucleic acid sequence: tatatgtgcaacagcagctgcatgggcggcatgcagcgccg cccgattctgaccattattaccctggaagatagc (SEQ ID NO:8)
[0137] TP53 p.R248Q antigen amino acid sequence: YMCNSSCMGGMQRRPILTIITLEDS (SEQ ID NO: 4)
[0138] Example 2
[0139] Intervention effect of the fusion gene mRNA-form vaccine on the development of syngeneic tumors of mouse transplanted tumor cells MDA-MB-468
[0140] (1) Given that the fusion gene can be normally expressed in mammalian cells. The mRNA vaccines expressing PIK3CA p.H1047R, AKT1 p.E17K, KMT2C p.K2797Q, TP53 p.R248Q separately and the fusion-expressed 4 antigens were prepared by in vitro transcription and encapsulated into lipid nanoparticles with the in vivo transfection reagent in vivo-jetPEI to form mRNA-form vaccines. The specific steps are as follows: First, determine the nucleic acid dosage and injection volume according to experimental requirements, and select an appropriate N / P ratio (usually 6 - 8). Then, dilute the nucleic acid to 1 / 2 of the injection volume with 5% glucose as the final concentration, and at the same time dilute the in vivo-jetPEI to 1 / 2 of the injection volume. Next, add the diluted transfection reagent to the nucleic acid solution, gently mix evenly and incubate at room temperature for 15 minutes to form a stable complex. Before injection, equilibrate the complex at room temperature. The whole operation needs to be carried out in a clean environment to ensure that the nucleic acid concentration does not exceed 0.5 μg / μL. Observe the inhibitory effect of the fusion gene immunity on the growth of MDA-MB-468 transplanted tumor cells after syngeneic transplantation of MDA-MB-468 cells.
[0141] Fusion antigen sequence: ALEYFMKQMNDARHGGWTTKMDWIFAIVKEGWL HKRGKYIKTWRPRYFLLLDNQCVSVEPKKQEQENKTLVLSDKYMCNSSC MGGMQRRPILTIITLEDS (SEQ ID NO: 9)
[0142] (2) After determining the tumorigenesis of MDA-MB-468 cells, perform mRNA intramuscular injection on mice according to the Figure 2 immunization strategy marked on the time axis. Divide them into 5 groups with 5 mice in each group, and inject 25 μg for each mouse. One week after injection, inoculate the MDA-MB-468 tumor cells with the tumorigenesis conditions explored before, observe the tumor formation time, measure the long diameter a and short diameter b of the tumor every two days, calculate the tumor volume according to a×b×b / 2, and draw the tumor growth curve. The results are as Figure 3 shown. The group with the fusion of 4 new antigens can prevent the formation of transplanted tumors in mice, and the vaccine effect is significant.
[0143] Example 3 Detection of the immunogenicity of the fusion antigen vaccine in mice:
[0144] Given the excellent effect of this vaccine in the preventive immunization experiment, we detected the immunogenicity of the mice immunized with the fusion antigen.
[0145] Take the spleen cell suspension of the immunized mice in Example 2 above and use ACK lysate to lyse red blood cells. Resuspend 1×10 5 ~3×10 5 cells and plate them into individual wells of a MAIPS4510 multiplex screening 96-well plate. All wells were previously coated with anti-interferon γ detection antibody, and then add the polypeptides of PIK3CAp.H1047R, AKT1 p.E17K, KMT2Cp.K2797Q, TP53 p.R248Q respectively, and at the same time add positive control polymorphisms and negative control (without polypeptides). After 72 hours, wash the plate and add secondary antibody (BD), and incubate overnight at 4 °C on the plate. Then wash the wells with PBS and add HRP streptavidin. After 1 hour of incubation, use AEC substrate to develop the color of the plate for 5-8 minutes. Then, gently wash the plate under cold tap water. When dry, image the plate using an automated plate reader system (CTL Technology Company) and quantify the number of spots. The results are as Figure 4 shown. The immunogenicity produced by fusing four antigens is the strongest, indicating that the fusion 4-antigen vaccine has the best anti-tumor effect.
[0146] In summary, with the above technical solutions of the present invention, the present invention uses the tumor tissues and normal tissues of Luminal A type, Luminal B type, HER-2 overexpressing type and triple-negative breast cancer patients for whole-genome sequencing, which can accurately analyze the HLA alleles of patients. By using whole-genome and mutation-related analysis, the top 4 genes with mutation frequencies in the above 4 subtypes of breast cancer are PIK3CAp.H1047R, AKT1 p.E17K, KMT2C p.K2797Q and TP53p.R248Q in sequence, and based on the above mutant genes, common immunogenic peptides in different molecular subtype breast tumors are predicted, providing broad-spectrum new targets for breast cancer immunotherapy and improving the survival rate and prognosis of breast cancer patients.
[0147] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. Use of a mutation site as an epitope in the preparation of a breast cancer diagnostic reagent or a drug for treating breast cancer; The mutation site is at least one of PIK3CA p.H1047R, AKT1 p.E17K, KMT2C p.K2797Q, and / or TP53 p.R248Q.
2. An epitope peptide, which contains 19 to 31 amino acid residues and contains at least one of the mutation sites of PIK3CA p.H1047R, AKT1 p.E17K, KMT2C p.K2797Q, and / or TP53 p.R248Q.
3. The epitope peptide according to claim 2, wherein It has the amino acid sequence as described in any one of SEQ ID NO:1 to 4.
4. An antigen, which comprises at least two of the epitope peptides described in claim 2 or 3.
5. The antigen according to claim 4, wherein In sequence from the N-terminus to the C-terminus, it is the epitope peptide shown in SEQ ID NO:1, the epitope peptide shown in SEQ ID NO:2, the epitope peptide shown in SEQ ID NO:3, and the epitope peptide shown in SEQ ID NO:
4.
6. A fusion protein, which comprises a chemokine and the epitope peptide described in claim 2 or 3, or comprises a chemokine and the antigen described in claim 4 or 5; The chemokine is selected from at least one of CCL1, CCL13, and / or CXCL3.
7. The fusion protein according to claim 6, wherein In sequence from the N-terminus to the C-terminus, it is CCL1, the polypeptide shown in SEQ ID NO:1, the polypeptide shown in SEQ ID NO:2, the polypeptide shown in SEQ ID NO:3, and the polypeptide shown in SEQ ID NO:
4.
8. A nucleic acid, which comprises at least one of the following: I), a nucleic acid encoding the epitope peptide described in claim 2 or 3; II), a nucleic acid encoding the antigen described in claim 4 or 5; III), a nucleic acid encoding the fusion protein described in claim 6 or 7.
9. The nucleic acid according to claim 8, wherein the nucleic acid sequence encoding the epitope peptide shown in SEQ ID NO:1 is as shown in SEQ ID NO:5; the nucleic acid sequence encoding the epitope peptide shown in SEQ ID NO:2 is as shown in SEQ ID NO:6; the nucleic acid sequence encoding the epitope peptide shown in SEQ ID NO:3 is as shown in SEQ ID NO:7; the nucleic acid sequence encoding the epitope peptide shown in SEQ ID NO:4 is as shown in SEQ ID NO:
8.
10. mRNA, which comprises: 5’Cap, the nucleic acid described in claim 8 or 9, and Poly-A Tail.
11. An expression unit, which comprises the nucleic acid described in claim 8 or 9 and a promoter.
12. A plasmid vector, which comprises the nucleic acid described in claim 8 or 9 or the expression unit described in claim 11.
13. A host, which is transformed or transfected with the plasmid vector described in claim 12; or the nucleic acid described in claim 8 or 9 is integrated into its genome; or the expression unit described in claim 11 is integrated into its genome.
14. A method for preparing the epitope peptide according to claim 2 or 3, the antigen according to claim 4 or 5, or the fusion protein according to claim 6 or 7, which comprises culturing the host according to claim 13 to obtain an expression product containing the epitope peptide, antigen or fusion protein.
15. A method for preparing the mRNA according to claim 10, which comprises culturing the host according to claim 13 to obtain template DNA, and obtaining the mRNA through in vitro transcription, capping reaction and tailing reaction.
16. Use of the epitope peptide according to claim 2 or 3, the antigen according to claim 4 or 5, the fusion protein according to claim 6 or 7, the nucleic acid according to claim 8 or 9, the mRNA according to claim 10, the expression unit according to claim 11, the plasmid vector according to claim 12, or the host according to claim 13 in the preparation of a medicament for treating breast cancer.
17. A drug for treating breast cancer, comprising: The epitope peptide according to claim 2 or 3, the antigen according to claim 4 or 5, the fusion protein according to claim 6 or 7, the nucleic acid according to claim 8 or 9, the mRNA according to claim 10, the expression unit according to claim 11, the plasmid vector according to claim 12, or the host according to claim 13.
18. A vaccine for treating breast cancer, which comprises the mRNA according to claim 10 and an in vivo transfection reagent.