System and method for identifying HLA variants
A subject-specific reference genome approach improves HLA variant detection by aligning sequence reads to personalized references, addressing false-negative issues and eliminating the need for matched controls, thereby enhancing the understanding of HLA gene variants.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- FOUNDATION MEDICINE INC
- Filing Date
- 2024-03-22
- Publication Date
- 2026-04-21
AI Technical Summary
Existing methods for identifying HLA gene variants face challenges due to their highly polymorphic nature, leading to false-negative alignments and the need for matched normal controls, which complicates the detection and understanding of their biological effects.
Generating a subject-specific reference genome using WGS and WES methods, and employing a targeted sequencing approach that aligns sequence reads to a personalized reference, eliminating the need for matched normal controls and improving variant detection rates.
Enhances the accuracy of HLA variant detection by reducing false negatives and enabling identification without requiring matched normal controls, facilitating better understanding of HLA gene variants and their biological impacts.
Smart Images

Figure 2026512810000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 576,450, filed on March 24, 2023, the content of which is incorporated herein by reference in its entirety.
[0002] The present disclosure generally relates to methods and systems for identifying genetic variants, and more particularly, to methods and systems for identifying variants in the human leukocyte antigen (HLA) gene.
Background Art
[0003] The HLA gene encodes cell - surface proteins that regulate the human immune system. Genetic changes in the HLA gene can cause immunological deficiencies, such as impairment of the immune response to cancer and other diseases, and poor response to immune - checkpoint inhibitors. However, it is difficult to identify variants in the HLA gene and understand their effects. The HLA gene is highly polymorphic. There is a need for improved methods for identifying variants in the HLA gene and determining their biological effects. The present disclosure addresses these needs.
Summary of the Invention
[0004] Methods and systems for identifying variants in HLA genes are disclosed herein. Existing methods for identifying variants in HLA genes are often based on whole-genome sequencing (WGS) or whole-exome sequencing (WES) and often involve aligning sequence reads to a genetic reference genome. Given the highly polymorphic nature of HLA variants, aligning reads derived from HLA genes to a genetic reference genome can result in false-negative sequence alignments. The methods disclosed herein generate a reference genome individualized for the subject using WGS and WES methods, as well as sequencing methods (including, but not limited to, targeted sequencing). Generating a subject-specific reference genome will help mitigate false-negative sequence read alignments resulting from the highly polymorphic nature of HLA variants. In this way, the methods disclosed herein can improve the detection rate of variants in HLA genes.
[0005] Existing methods for identifying variants in HLA genes often require matched normal controls. In other words, existing methods often require comparative samples from the same subject as experimental samples. By comparing experimental samples with matched normal controls from the same subject, it is possible to predict that properties other than the biological property being investigated (e.g., HLA gene variants) are otherwise similar or identical. In this way, HLA gene variants can be identified because their variant sequences differ from those of the matched normal controls. The methods and systems disclosed herein address the aforementioned shortcomings of existing methods. The methods and systems disclosed not only generate individualized, subject-specific reference genomes to identify HLA variants, but also eliminate the need for the use of matched normal controls to identify HLA variants.
[0006] In some embodiments, the process involves providing multiple nucleic acid molecules obtained from a sample of a subject, ligating one or more adapters onto one or more nucleic acid molecules from the multiple nucleic acid molecules, amplifying one or more ligated nucleic acid molecules from the multiple nucleic acid molecules, capturing amplified nucleic acid molecules from the amplified nucleic acid molecules, sequencing the captured nucleic acid molecules by a sequencer to obtain multiple sequence reads representing the captured nucleic acid molecules, receiving sequence read data for the multiple sequence reads in one or more processors, and aligning the sequence read data to a first reference genome using one or more processors to obtain a first multiple A method is disclosed herein that includes identifying a number of HLA alignment reads and unmapped sequence reads; using one or more processors to search an HLA polymorphism database with the first multiple HLA alignment reads and unmapped sequence reads to generate a target-specific reference genome; using one or more processors to align the first multiple HLA alignment reads and unmapped sequence reads to the target-specific reference genome to identify a second multiple HLA alignment reads; and using one or more processors to process the second multiple HLA alignment reads to identify HLA variants.
[0007] In some embodiments, a variant caller can be used to process a second set of HLA alignment reads and identify HLA variants. In any of the embodiments herein, HLA variant identification does not require the use of sequence read data from a matched normal control. In any of the embodiments herein, the HLA polymorphism database is the IPD-IMGT / HLA database. In any of the embodiments herein, the identified HLA variant is the HLA-I variant. In some embodiments, the HLA variant is a variant of HLA-A, HLA-B, and / or HLA-C. In any of the embodiments herein, the identified HLA variant is the HLA-I and / or HLA-II variant.
[0008] In any of the embodiments of this specification, the subject is suspected of having cancer or is determined to have cancer. In some embodiments, cancer is B-cell carcinoma (multiple myeloma), melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, uterine cancer, endometrial cancer, oral cancer, pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine cancer, appendiceal cancer, salivary gland cancer, thyroid cancer, adrenal cancer, osteosarcoma, chondrosarcoma, hematological cancer. Cancers, adenocarcinomas, inflammatory myofibroblastomas, gastrointestinal stromal tumors (GISTs), colon cancers, multiple myeloma (MM), myelodysplastic syndromes (MDS), myeloproliferative disorders (MPDs), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft tissue sarcoma, fibrosarcoma, myxosarcoma, and fatty tissue. Tumors, osteosarcoma, chordoma, angiosarcoma, endosarcoma, lymphangiosarcoma, lymphangioendosarcoma, synoviomas, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, liver cancer, bile duct cancer, choriocarcinoma, seminoma, embryonic carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, These include ependymoma, pineal cell tumor, hemangioblastoma, acoustic neuroblastoma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, aplastic myelogenesis, eosinophilic syndrome, systemic mastocytosis, familial eosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, or carcinoid tumors.
[0009] In some embodiments, cancers include acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpression / amplification), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deficiency), chronic myeloid leukemia, chronic myeloid leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, and colon cancer. Rectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild-type), cryopyrin-associated periodic fever syndrome, cutaneous T-cell lymphoma, dermatofibrosarcoma protuberans, diffuse large B-cell lymphoma, fallopian tube cancer, follicular B-cell non-Hodgkin lymphoma, follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, gastrointestinal stromal tumor, gastrointestinal stromal tumor (KIT+), giant cell tumor of bone, glioblastoma, granulomatosis with polyangiitis, head and neck squamous cell carcinoma, hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, systemic lupus erythematosus, mantle cell lymphoma, medullary thyroid carcinoma, melanoma, BRAF Melanoma with V600 mutation, melanoma with BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman disease, multiple hematological malignancies including Philadelphia chromosome-positive ALL and CML, multiple myeloma, myelofibrosis, non-Hodgkin lymphoma, unresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, non-small cell lung cancer, non-small cell lung cancer (ALK+), non-small cell lung cancer (PD-L1+), non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), non-small cell lung cancer (with BRAF V600E mutation), non-small cell lung cancer (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer (EGFRThis may include ovarian cancer (with T790M mutation), ovarian cancer (with BRCA mutation), pancreatic cancer, neuroendocrine tumors of pancreatic, gastrointestinal, or lung origin, pediatric neuroblastoma, peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, renal cell carcinoma, rheumatoid arthritis, small lymphocytic lymphoma, soft tissue sarcoma, solid tumors (MSI-H / dMMR), squamous cell carcinoma of the head and neck, squamous non-small cell lung cancer, thyroid cancer, thyroid carcinoma, urothelial carcinoma, or primary gammaglobulinemia.
[0010] In some embodiments, the methods disclosed herein may further include treating the subject with anticancer therapy. In some embodiments, the treatment of the subject is carried out after the subject has been determined to have cancer. In some embodiments, the anticancer therapy may include targeted anticancer therapy. In some embodiments, the targeted anticancer therapy may include abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), ado-trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), and amivantamab-vmjw (Rybrevant). ), anastrozole (Arimidex), apalutamide (Erleada), aciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicaptagensilolucel (Yescarta), axitinib (Inlyta), verantamab mahodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), verzutifan (We Lireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexcabutadiene oatlucel (Tecartus), brigutinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Ilaris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), semiprimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex),Daratumumab and hyaluronidase-fihj (Darzalex Faspro), darolutamide (Nubeqa), dasatinib (Sprycel), deniroikin difutitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostallumab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enacidenib mesylate (Idhifa), enco Rafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedra hydrochloride Tinib (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glass-degib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idekabuta genbiculose Abecma, Idelalisib (Zydelig), Imatinib mesylate (Gleevec), Infiglatinib phosphate (Truseltiq), Inotuzumab ozogamicin (Besponsa), Ipilimumab (Yervoy), Isatuximab-irfc (Sarclisa), Ivosidenib (Tibsovo), Ixazomib citrate (Ninlaro), Lanreotide acetate (Somatuline) Depot), lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), lenvatinib mesylate (Lenvima), letrozole (Femara), lysocabate maralucel (Breyanzi), loncustuximab tesillin-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu177-dotete (Lutathera), margetuximab-cmkb (Margenza), midostaurin (Rydapt),Mobocertinib succinate (Exkivity), mogamulizumab-kpkc (Poteligeo), moxetumomab pasdotox-tdfk (Lumoxiti), naxitamab-gqgk (Danyelza), necitumumab (Portrazza), neratinib maleate (Nerlynx), nilotinib (Tasigna), niraparib tosylate monohydrate (Zejula), nivolumab (Opdivo), obinutuzumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), olalatumumab (Lartruvo), osimertinib (Tagrisso), palbociclib (Ibrance), panitumumab (Vect ibix), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pexidartinib hydrochloride (Turalio), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), pralcetinib (Gavreto), radium-223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), lipretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase (Rituxan) Hycela), Romidepsin (Istodax), Rucaparib cansylate (Rubraca), Ruxolitinib phosphate (Jakafi), Sacituzumab Govitecan-hziy (Trodelvy), Cericlib, Selinexol (Xpovio), Serpercatinib (Retevmo), Selumetinib sulfate (Koselugo), Siltuximab (Sylvant), Sirolimus-binding protein particles (Fyarro), Sonidedib (Odom) zo), sorafenib (Nexavar), sotracib (Lumakras), sunitinib (Sutent), tafacitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), teventafusp-tebn (Kimmtrak), temsirolimus (Torisel),Tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tocitumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), toremifene (Fa This may include reston, tucatinib (Tukysa), umbralicib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), bismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), ziv-aflibercept (Zaltrap), or any combination thereof.
[0011] In any of the embodiments of this specification, the method described herein may further include obtaining a sample from a subject. In any of the embodiments of this specification, the sample may include a tissue biopsy sample or a liquid biopsy sample. In any of the embodiments of this specification, the sample may be a liquid biopsy sample and may include blood, plasma, cerebrospinal fluid, sputum, feces, urine, or saliva. In some embodiments, the sample may be a liquid biopsy sample and may include circulating tumor cells (CTCs). In some embodiments, the sample may be a liquid biopsy sample and may include cell-free DNA (cfDNA). In some embodiments, cell-free DNA (cfDNA) or a portion thereof may include circulating tumor DNA (ctDNA). In any of the embodiments of this specification, the plurality of nucleic acid molecules may include a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules. In some embodiments, the tumor nucleic acid molecules may originate from the tumor portion of a xenotissue biopsy sample, and the non-tumor nucleic acid molecules may originate from the normal portion of a xenotissue biopsy sample. In some embodiments, the sample may include a liquid biopsy sample, tumor nucleic acid molecules may be derived from the circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and non-tumor nucleic acid molecules may be derived from the non-tumor cell-free DNA (cfDNA) fraction of the liquid biopsy sample.
[0012] In any of the embodiments described herein, one or more adapters may include amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences. In any of the embodiments described herein, captured nucleic acid molecules may be captured from nucleic acid molecules amplified by hybridization to one or more bait molecules. In some embodiments, one or more bait molecules may include one or more nucleic acid molecules, each containing a region complementary to the region of the captured nucleic acid molecule. In some embodiments, amplification of nucleic acid molecules may include performing polymerase chain reaction (PCR) amplification techniques, non-PCR amplification techniques, or isothermal amplification techniques. In any of the embodiments described herein, sequencing may include the use of massively parallel sequencing (MPS) techniques, whole-genome sequencing (WGS), whole-exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing techniques. In some embodiments, sequencing may include massively parallel sequencing, and massively parallel sequencing techniques may include next-generation sequencing (NGS). In any of the embodiments described herein, the sequencer may include a next-generation sequencer. In any of the embodiments herein, one or more of the sequencing reads may overlap with one or more loci within one or more subgenome segments in the sample.
[0013] In some embodiments, one or more gene loci are 10-20, 10-40, 10-60, 10-80, 10-100, 10-150, 10-200, 10-250, 10-300, 10-350, 10-400, 10-450, 10-500, 20-40, 20-60, 20-80, 20-100, 20-150, 20-200, 20-2 50 loci, 20-300 loci, 20-350 loci, 20-400 loci, 20-500 loci, 40-60 loci, 40-80 loci, 40-100 loci, 40-150 loci, 40-200 loci, 40-250 loci, 40-300 loci, 40-350 loci, 40-400 loci, 40-500 loci, 60-80 loci, 60-100 loci, 60-150 loci, 60-200 loci, 60-250 loci, 60-300 loci, 60-35 0 loci, 60-400 loci, 60-500 loci, 80-100 loci, 80-150 loci, 80-200 loci, 80-250 loci, 80-300 loci, 80-350 loci, 80-400 loci, 80-500 loci, 100-150 loci, 100-200 loci, 100-250 loci, 100-300 loci, 100-350 loci, 100-400 loci, 100-500 loci, 150-200 loci, 150-250 loci, 150-3 This may include loci 00, 150-350, 150-400, 150-500, 200-250, 200-300, 200-350, 200-400, 200-500, 250-300, 250-350, 250-400, 250-500, 300-350, 300-400, 300-500, 350-400, 350-500, or 400-500.
[0014] In any of the embodiments herein, one or more gene loci are ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CA SP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CD KN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNM T3A, DOT1L, EED, EGFR, EMSY(C11orf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESR1, ETV4, ETV5, ETV6, EWSR1, EZH2, E ZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1 , FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4(C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A(MLL),KMT2D(MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MI TF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX 2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDC D1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1 , PRKAR1A, PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, R EL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, or any combination thereof. In any of the embodiments herein, one or more loci may include ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK,This may include CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-1β, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRα, PDGFRβ, PD-L1, PI3Kδ, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, VEGFB, or any combination thereof. In any of the embodiments herein, one or more loci may further include HLA class I or HLA class II genes. In any embodiment of this specification, the method disclosed herein may further include generating a report indicating the presence or absence of a genetic variant using one or more processors.
[0015] In some embodiments, the methods disclosed herein may further include transmitting reports to healthcare providers. In some embodiments, reports may be transmitted via a computer network or peer-to-peer connection.
[0016] A method for identifying HLA variants is disclosed herein, comprising: receiving sequence read data for multiple sequence reads in one or more processors; aligning the sequence read data to a first reference genome using one or more processors to identify a first set of multiple HLA alignment reads and unmapped sequence reads; searching an HLA polymorphism database with the first set of multiple HLA alignment reads and unmapped sequence reads using one or more processors to generate a target-specific reference genome; aligning the first set of multiple HLA alignment reads and unmapped sequence reads to a target-specific reference genome using one or more processors to identify a second set of multiple HLA alignment reads; and processing the second set of multiple HLA alignment reads using one or more processors to identify HLA variants.
[0017] In some embodiments, multiple sequence reads are obtained using a targeted sequencing method that includes targeting HLA loci. In some embodiments, a variant caller can be used to process multiple HLA alignment reads and identify HLA variants. In any of the embodiments herein, HLA variant identification does not require the use of sequence read data from matched normal controls. In any of the embodiments herein, the HLA polymorphism database may be an IPD-IMGT / HLA database. In any of the embodiments herein, the HLA variant may be an HLA-I variant. In some embodiments, the identified HLA-I variant is a variant of HLA-A, HLA-B, and / or HLA-C. In any of the embodiments herein, the HLA variant may be an HLA-I and / or HLA-II variant.
[0018] In some embodiments, methods for identifying HLA variants are disclosed herein, comprising: receiving sequence read data for a plurality of sequence reads in one or more processors; aligning the sequence read data to a first reference genome using one or more processors to identify a first plurality of HLA alignment reads and unmapped sequence reads; aligning the first plurality of HLA alignment reads and unmapped sequence reads to a second reference genome using one or more processors to identify a second plurality of HLA alignment reads; and processing the second plurality of HLA alignment reads using one or more processors to identify HLA variants, wherein the identification does not require the use of sequence read data from matched normal controls. In some embodiments, a variant caller can be used to process the second plurality of HLA alignment reads and identify HLA variants.
[0019] In any of the embodiments herein, the second reference genome may be a target-specific reference genome. In some embodiments, a target-specific reference can be generated by searching an HLA polymorphism database using multiple alignment reads and unmapped sequence reads of multiple first HLAs. In some embodiments, the HLA polymorphism database may be an IPD-IMGT / HLA database. In any of the embodiments herein, the HLA variant may be an HLA-I variant. In any of the embodiments herein, the identified HLA-I variant may be a variant of HLA-A, HLA-B, and / or HLA-C. In any of the embodiments herein, the HLA variant may be an HLA-I and / or HLA-II variant.
[0020] In any embodiment of this specification, the method disclosed herein may further include: using one or more processors to align an HLA variant to an HLA coding reference sequence to annotate the exon-intron boundaries in the HLA variant; using one or more processors to translate the annotated HLA variant to an HLA polypeptide sequence; using one or more processors to align the translated HLA polypeptide sequence to a reference HLA polypeptide sequence to identify amino acid changes in the translated HLA polypeptide sequence; and using one or more processors to categorize the translated HLA polypeptides as containing synonymous mutations, non-start mutations, non-stop mutations, nonsense mutations, or missense mutations.
[0021] In any of the embodiments herein, the HLA variant may be an HLA-I variant. In some embodiments, the identified HLA-I variant may be a variant of HLA-A, HLA-B, and / or HLA-C. In any of the embodiments herein, the HLA variant may be an HLA-I and / or HLA-II variant.
[0022] In any of the embodiments of this specification, searching a polymorphic database may involve using a software tool that predicts the probability that an input sequence corresponds to a known HLA sequence. In some embodiments, the software tool may be Optitype. In any of the embodiments of this specification, the variant caller may include a short variant caller. In any of the embodiments of this specification, the subject may be human.
[0023] In some embodiments, a system is disclosed herein, comprising: one or more processors; a memory communicatively connected to one or more processors and configured to store instructions, wherein when an instruction is executed by one or more processors, the system causes the system to receive sequence read data for a plurality of sequence reads, align the sequence read data to a first reference genome to identify a first plurality of HLA alignment reads and unmapped sequence reads, search an HLA polymorphism database with the first plurality of HLA alignment reads and unmapped sequence reads to generate a target-specific reference genome, align the first plurality of HLA alignment reads and unmapped sequence reads to identify a second plurality of HLA alignment reads, and process the second plurality of HLA alignment reads to identify HLA variants.
[0024] In some embodiments, a variant caller can be used to process a second set of HLA alignment reads and identify HLA variants. In any of the embodiments herein, HLA variant identification does not require the use of sequence read data from a matched normal control. In any of the embodiments herein, the HLA polymorphism database may be an IPD-IMGT / HLA database.
[0025] In any of the embodiments of this specification, the system disclosed herein may further include instructions that, when executed by one or more processors, cause the system to align the HLA variant to the HLA code reference sequence, annotate the exon-intron boundaries in the HLA variant, translate the annotated HLA variant into an HLA polypeptide sequence, align the translated HLA polypeptide sequence to a reference HLA polypeptide sequence, identify amino acid changes in the translated HLA polypeptide sequence, and categorize the translated HLA polypeptide as synonymous, non-start, non-stop, nonsense, or missense.
[0026] In any of the embodiments of this specification, the HLA variant can be an HLA-I variant. In some embodiments, the identified HLA-I variant can be a variant of HLA-A, HLA-B, and / or HLA-C. In any of the embodiments of this specification, the HLA variant can be an HLA-I and / or HLA-II variant.
[0027] In any of the embodiments of this specification, the search of the polymorphism database can include the use of software tools that predict the probability that the input sequence corresponds to a known HLA sequence. In some embodiments, the software tool can be Optitype. In any of the embodiments of this specification, the variant caller can include a short variant caller. In any of the embodiments of this specification, the subject can be human.
[0028] In some aspects, a method for diagnosing a disease, the method including diagnosing that a subject has the disease based on the identification of an HLA variant in a sample from the subject, wherein the HLA variant is identified according to any of the embodiments of this specification, is disclosed herein.
[0029] In some embodiments, a method for selecting an anticancer therapy is disclosed herein, comprising selecting an anticancer therapy for a subject in response to the identification of an HLA variant for a sample from the subject, wherein the HLA variant is identified according to any of the embodiments herein. In some embodiments, a method for treating cancer in a subject is disclosed herein, comprising administering an effective dose of an anticancer therapy to the subject in response to the identification of an HLA variant for a sample from the subject, wherein the HLA variant is identified according to any of the embodiments herein.
[0030] In some embodiments, a method for monitoring cancer progression or recurrence in a subject is disclosed herein, the method comprising: identifying a first HLA variant in a first sample obtained from the subject at a first time point; identifying a second HLA variant in a second sample obtained from the subject at a second time point; and comparing the first HLA variant with the second HLA variant to monitor cancer progression or recurrence, according to one of the embodiments herein. In some embodiments, the second HLA variant for the second sample can be identified according to one of the embodiments disclosed herein.
[0031] In any of the embodiments of this specification, the method disclosed herein may further include selecting an anticancer therapy for a subject in response to cancer progression. In any of the embodiments of this specification, the method disclosed herein may further include applying an anticancer therapy to a subject in response to cancer progression. In any of the embodiments of this specification, the method disclosed herein may further include adjusting an anticancer therapy for a subject in response to cancer progression. In any of the embodiments of this specification, the method disclosed herein may further include adjusting the dosage of an anticancer therapy or selecting a different anticancer therapy in response to cancer progression. In some embodiments, the method disclosed herein may further include applying an adjusted anticancer therapy to a subject. In any of the embodiments of this specification, the first time point may be before the subject receives anticancer therapy, and the second time point may be after the subject receives anticancer therapy.
[0032] In any of the embodiments of this Specification, the subject may have cancer, may be at risk of having cancer, may be routinely screened for cancer, or may be suspected of having cancer. In any of the embodiments of this Specification, cancer may be a solid tumor. In any of the embodiments of this Specification, cancer may be a hematological malignancy. In any of the embodiments of this Specification, anti-cancer therapy may include chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery. In any of the embodiments of this Specification, the methods disclosed herein may further include determining, identifying, or applying HLA variants to a sample as diagnostic values related to the sample. In any of the embodiments of this Specification, the methods disclosed herein may further include generating a genomic profile for the subject based on the identification of HLA variants.
[0033] In some embodiments, the genomic profile of the subject may further include results from comprehensive genomic profiling (CGP) tests, gene expression profiling tests, cancer hotspot panel tests, DNA methylation tests, DNA fragmentation tests, RNA fragmentation tests, or any combination thereof. In any of the embodiments herein, the genomic profile of the subject may further include results from tests based on nucleic acid sequencing. In any of the embodiments herein, the methods disclosed herein may further include selecting, administering, or applying anti-cancer therapy to the subject based on the generated genomic profile. In any of the embodiments herein, HLA variant identification of a sample may be used in determining a proposed treatment for the subject. In any of the embodiments herein, HLA variant identification of a sample may be used when applying or administering treatment to the subject. In any of the embodiments herein, the subject may be human.
[0034] In some embodiments, the Specified Non-Temporary Computer-Readable Storage Medium stores one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by one or more processors of the system, the system is caused to receive sequence read data for a plurality of sequence reads, to align the sequence read data to a first reference genome to identify a first plurality of HLA alignment reads and unmapped sequence reads, to search an HLA polymorphism database with the first plurality of HLA alignment reads and unmapped sequence reads to generate a target-specific reference genome, to align the first plurality of HLA alignment reads and unmapped sequence reads to identify a second plurality of HLA alignment reads, and to process the second plurality of HLA alignment reads to identify an HLA variant.
[0035] In some embodiments, a variant caller can be used to process multiple HLA alignment reads and identify HLA variants. In some embodiments, HLA variant identification does not require the use of sequence read data from matched normal controls. In any of the embodiments herein, the HLA polymorphism database may be an IPD-IMGT / HLA database.
[0036] In any embodiment of this specification, the non-temporary computer-readable storage medium may further include instructions, which, when executed by one or more processors, cause the system to align an HLA variant to an HLA code reference sequence, to annotate the exon-intron boundaries in the HLA variant, to translate the annotated HLA variant to an HLA polypeptide sequence, to align the translated HLA polypeptide sequence to a reference HLA polypeptide sequence, to identify amino acid changes in the translated HLA polypeptide sequence, and to categorize the translated HLA polypeptide as synonymous, nonstart, nonstop, nonsense, or missense.
[0037] In some embodiments, the HLA variant may be an HLA-I variant. In any of the embodiments herein, the identified HLA-I variant may be a variant of HLA-A, HLA-B, and / or HLA-C. In any of the embodiments herein, the HLA variant may be an HLA-I and / or HLA-II variant. In any of the embodiments herein, searching a polymorphic database may involve using a software tool capable of predicting the probability that an input sequence corresponds to a known HLA sequence. In some embodiments, the software tool may be Optitype. In any of the embodiments herein, the variant caller may include a short variant caller. In any of the embodiments herein, the subject may be a human.
[0038] Embedding by reference All publications, patents, and patent applications referenced herein are incorporated herein by whole to the same extent as each individual publication, patent, or patent application is specifically and individually indicated to be incorporated herein by whole. In the event of any conflict between the terminology used herein and the terminology used in the incorporated references, the terminology used herein shall prevail.
[0039] Various aspects of the disclosed methods, devices, and systems are described in detail in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by referring to the following detailed description of exemplary embodiments and the appended drawings. [Brief explanation of the drawing]
[0040] [Figure 1] This disclosure describes non-limiting and exemplary methods for identifying HLA variants according to some embodiments of this disclosure. [Figure 2] Further non-limiting and exemplary methods for identifying HLA variants, according to some embodiments of this disclosure, are shown. [Figure 3] The following are exemplary computing devices or systems that conform to some embodiments of this disclosure. [Figure 4] This specification shows an exemplary computer system or computer network that follows some examples of the systems described herein. [Figure 5] This disclosure provides a non-limiting, exemplary schematic workflow of an algorithm used to identify HLA variants, according to several embodiments of this disclosure. [Modes for carrying out the invention]
[0041] Methods and systems for identifying variants in HLA genes are disclosed herein. Existing methods for identifying variants in HLA genes are often based on whole-genome sequencing (WGS) or whole-exome sequencing (WES) and often involve aligning sequence reads to a genetic reference genome. Given the highly polymorphic nature of HLA variants, aligning reads derived from HLA genes to a genetic reference genome can result in false-negative sequence alignments. The methods disclosed herein generate a reference genome individualized for the subject using WGS and WES methods, as well as sequencing methods (including, but not limited to, targeted sequencing). Generating a subject-specific reference genome will help mitigate false-negative sequence read alignments resulting from the highly polymorphic nature of HLA variants. In this way, the methods disclosed herein can improve the detection rate of variants in HLA genes.
[0042] Existing methods for identifying variants in HLA genes often require matched normal controls. In other words, existing methods often require comparative samples from the same subject as experimental samples. By comparing experimental samples with matched normal controls from the same subject, it is possible to predict that properties other than the biological property being investigated (e.g., HLA gene variants) are otherwise similar or identical. In this way, HLA gene variants can be identified because their variant sequences differ from those of the matched normal controls. The methods and systems disclosed herein address the aforementioned shortcomings of existing methods. The methods and systems disclosed not only generate individualized, subject-specific reference genomes to identify HLA variants, but also eliminate the need for the use of matched normal controls to identify HLA variants.
[0043] In some examples, a disclosed method for identifying HLA variants may include: a) receiving sequence read data for multiple sequence reads in one or more processors; b) aligning the sequence read data to a first reference genome using one or more processors to identify a first set of multiple HLA alignment reads and unmapped sequence reads; c) searching an HLA polymorphism database with the first set of multiple HLA alignment reads and unmapped sequence reads using one or more processors to generate a target-specific reference genome; d) aligning the first set of multiple HLA alignment reads and unmapped sequence reads to a target-specific reference genome using one or more processors to identify a second set of multiple HLA alignment reads; and e) processing the second set of multiple HLA alignment reads using one or more processors to identify HLA variants.
[0044] In some examples, a disclosed method for identifying HLA variants may include: a) receiving sequence read data for multiple sequence reads in one or more processors; b) aligning the sequence read data to a first reference genome using one or more processors to identify a first set of multiple HLA alignment reads and unmapped sequence reads; c) aligning the first set of multiple HLA alignment reads and unmapped sequence reads to a second reference genome using one or more processors to identify a second set of multiple HLA alignment reads; and d) processing the second set of multiple HLA alignment reads using one or more processors to identify HLA variants, wherein the identification does not require the use of sequence read data from matched normal controls.
[0045] In some examples, the disclosed method may further include, for instance, using one or more processors to align an HLA variant to an HLA coding reference sequence to annotate the exon-intron boundaries in the HLA variant; using one or more processors to translate the annotated HLA variant to an HLA polypeptide sequence; using one or more processors to align the translated HLA polypeptide sequence to a reference HLA polypeptide sequence to identify amino acid changes in the translated HLA polypeptide sequence; and using one or more processors to categorize the translated HLA polypeptides as containing synonymous mutations, non-start mutations, non-stop mutations, nonsense mutations, or missense mutations.
[0046] definition Unless otherwise defined, all technical terms used herein have the same meaning as those generally understood by those skilled in the art of the field to which this disclosure pertains.
[0047] Where used herein and in the appended claims, the singular forms "a," "an," and "the" include plural references unless otherwise explicitly indicated by the context. Any reference to "or" herein is intended to include "and / or" unless otherwise specified.
[0048] "About" and "approximately" generally refer to an acceptable degree of error in the measured quantity, taking into account the nature or precision of the measurement. Exemplary degrees of error are within 20% of a given value or range of values, typically within 10%, and more typically within 5%.
[0049] As used herein, the terms “comprising” (and any form or variation of “comprising,” such as “comprise” and “comprises”), “having” (and any form or variation of “having,” such as “have” and “has”), “including” (and any form or variation of “includes,” such as “include”), or “containing” (and any form or variation of “containing,” such as “contains” and “contain”) are inclusive or open-ended and do not preclude additional unlisted additives, components, integers, elements, or method steps.
[0050] As used herein, the terms “individual,” “patient,” or “subject” are interchangeable and refer to any single animal to which treatment is desired, e.g., a mammal (including non-human animals such as dogs, cats, horses, rabbits, zoo animals, cattle, pigs, sheep, and non-human primates). In certain embodiments, the individual, patient, or subject herein is a human.
[0051] The terms “cancer” and “tumor” are used interchangeably herein. These terms refer to the presence of cells that have characteristics typical of cancer-causing cells, such as uncontrolled growth, immortality, metastatic ability, rapid growth and proliferation rates, and certain characteristic morphological features. While cancer cells often take the form of tumors, such cells can exist alone in animals or may be non-tumor-forming cancer cells, such as leukemia cells. These terms include solid tumors, soft tissue tumors, or metastatic lesions. As used herein, the term “cancer” includes both precancerous and malignant cancers.
[0052] As used herein, “treatment” (and its grammatical variations such as “to treat” or “to treat”) refers to a clinical intervention that seeks to alter the natural course of the individual being treated (e.g., administering anticancer drugs or performing anticancer therapy), which may be performed for preventive purposes or in the course of a series of clinicopathological treatments. Desired effects of treatment include, but are not limited to, prevention of disease onset or recurrence, symptom reduction, reduction of any direct or indirect pathological consequences of the disease, prevention of metastasis, slowing the rate of disease progression, improvement or mitigation of the disease state, and remission or an improved prognosis.
[0053] As used herein, the term “subgenome section” (or “subgenome sequence section”) refers to a portion of a genome sequence.
[0054] As used herein, the term “target section” refers to a subgenome section or an expression subgenome section (e.g., a transcription sequence of a subgenome section).
[0055] As used herein, the terms “variant sequence” or “variant” are used interchangeably and refer to nucleic acid sequences modified from the corresponding “normal” or “wild-type” sequences. In some examples, a variant sequence may be a “short variant sequence” (or “short variant”), i.e., a variant sequence less than approximately 50 base pairs in length.
[0056] The terms "allele frequency" and "allele fraction" are used interchangeably herein and refer to the fraction of sequence reads corresponding to a particular allele relative to the total number of sequence reads for a genomic locus.
[0057] The terms “variant allele frequency” and “variant allele fraction” are used interchangeably herein and refer to the fraction of sequence reads corresponding to a particular variant allele relative to the total number of sequence reads at a genomic locus.
[0058] Any headings used herein are for structural purposes only and should not be construed as limiting the subjects described.
[0059] Methods for identifying HLA variants HLA genes (e.g., the HLA-I gene) are crucial for proper human immune system function. Alterations in HLA genes can impair responses to immune checkpoint inhibitors and other medical interventions, and can interfere with the general capabilities of T cells and their recognition of antigens (e.g., cancer-derived antigens). HLA genes can be altered in many different ways (e.g., through copy number changes), which can lead to loss of heterozygosity in the HLA gene. In addition, short variant changes or mutations can occur in HLA genes, which can lead to loss of HLA gene function. These alterations may consist of nonsense mutations that convert amino acid code codons to stop codons in the HLA gene, or missense mutations that convert amino acid code codons to different amino acids, which can result in many biological effects, such as inefficient transcription translation or improper polypeptide folding (all of which can lead to the abolition of HLA function).
[0060] However, because HLA genes are among the majority of polymorphic genes in the human genome, identifying HLA gene variants can be challenging. Certain HLA polymorphisms may be improperly discarded during sequence alignment due to inadequate correspondence between the variant HLA sequence and the genetic reference genome. The method disclosed herein describes a process that addresses these challenges and enables the identification of short variants in HLA genes. The method disclosed herein is novel in that, firstly, it generates a patient-specific reference genome for identifying HLA variants, and secondly, it does not require a matched normal control for identifying HLA variants.
[0061] Figure 1 shows an exemplary schematic diagram illustrating a general process 100 for identifying an HLA variant. Process 100 can be performed, for example, using one or more electronic devices that implement a software platform. In some examples, process 100 is performed using a client-server system, and blocks of process 100 are divided in any way between the server and client devices. In other examples, blocks of process 100 are divided between a server and multiple client devices. Thus, while parts of process 100 are described herein as being performed by specific devices in a client-server system, it should be understood that process 100 is not limited in this way. In other examples, process 100 is performed using only one client device, or only multiple client devices. In process 100, some blocks are combined at will, the order of some blocks is changed at will, and some blocks are omitted at will. In some examples, additional steps may be performed in combination with process 100. Thus, the behavior illustrated (and described in more detail below) is illustrative in nature and should not be considered limiting.
[0062] In some embodiments, a method for identifying HLA variants is disclosed herein, comprising: a) receiving sequence read data for a plurality of sequence reads in one or more processors; b) aligning the sequence read data to a first reference genome using one or more processors to identify a first plurality of HLA alignment reads and unmapped sequence reads; c) searching an HLA polymorphism database with the first plurality of HLA alignment reads and unmapped sequence reads using one or more processors to generate a target-specific reference genome; d) aligning the first plurality of HLA alignment reads and unmapped sequence reads to a target-specific reference genome using one or more processors to identify a second plurality of HLA alignment reads; and e) processing the second plurality of HLA alignment reads using one or more processors to identify HLA variants.
[0063] In step 102 of Figure 1, sequence read data for multiple sequence reads is received.
[0064] As shown, in some examples, multiple sequence reads may be derived from targeted sequencing techniques, e.g., targeted exome sequencing techniques. In some examples, sequence read data may be derived from, for example, whole-genome or whole-exome sequencing techniques, as opposed to targeted exome sequencing techniques, in order to increase the number of genomic features detected (e.g., the number of short variants). Targeted sequencing methods may be based on sequencing by synthetic techniques (e.g., but not limited to next-generation sequencing techniques). In addition, targeted sequencing methods may be based on hybridization techniques (e.g., but not limited to microarrays). Targeted sequencing methods may also be based on avidity sequencing (e.g., but not limited to the method described in Arslan et al., bioRxiv, 2022).
[0065] In some cases, the sample may include a tissue biopsy sample, a fluid biopsy sample, or a normal control. The sample may also be a fluid biopsy sample and may include blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In addition, the sample may be a fluid biopsy sample and may contain circulating tumor cells (CTCs). The sample may also contain cell-free DNA (cfDNA) and / or circulating tumor DNA (ctDNA).
[0066] In some cases, the subject may be a human being.
[0067] In step 104 of Figure 1, multiple sequence reads are aligned to a first reference genome to identify a first set of multiple HLA alignment reads and unmapped sequence reads.
[0068] In some cases, the reference genome (e.g., a first reference genome, but not limited to) may be human reference genome Hg19 or human reference genome Hg38.
[0069] In step 106 of Figure 1, an HLA polymorphism database can be searched with a first set of HLA alignment reads and unmapped sequence reads to generate a target-specific reference genome. In some examples, the HLA polymorphism database may be, for example, the IPD-IMGT / HLA database. In some examples, searching the polymorphism database may involve using a software tool that predicts the probability that an input sequence corresponds to a known HLA sequence. In some examples, the software tool may be, for example, Optitype.
[0070] In step 108 of Figure 1, the first set of HLA alignment reads and unmapped sequence reads are aligned to a target-specific reference genome to identify the second set of HLA alignment reads.
[0071] In step 110 of Figure 1, a second set of HLA alignment reads is processed to identify HLA variants.
[0072] In some cases, a variant caller can be used to process multiple HLA alignment reads and identify HLA variants. In some cases, HLA variant identification does not require the use of sequence read data from matched normal controls. In some cases, the identified HLA variant may be, for example, an HLA-I variant. In some cases, an HLA-I variant may be a variant of HLA-A, HLA-B, and / or HLA-C. In some cases, the identified HLA variant may be an HLA-I and / or HLA-II variant. In some cases, the variant caller may include a short variant caller.
[0073] Variants (e.g., HLA variants, but not limited to them) may include single nucleotide substitutions, insertions, deletions, or combinations thereof. A single nucleotide substitution is a difference of a single nucleotide compared to a reference sequence, where the length of the sequence under investigation remains the same. For example, a transversion from purine (e.g., A) to pyramidine (e.g., T), where the sequence under investigation remains otherwise the same length and is a nucleotide substitution. Single nucleotide substitutions may affect the corresponding translated polypeptide sequence via silent, nonsense, or missense mutations. A silent mutation refers to a substitution in a genetic coding sequence where the ultimately translated polypeptide sequence is identical to the reference polypeptide sequence. A nonsense mutation refers to a substitution in a genetic coding sequence where the coded amino acid in the polypeptide sequence is substituted for a stop codon. Depending on how close the nonsense mutation is to the 5' end of the coding sequence in which it occurs, and other biological factors, a transcript derived from the mutated coding sequence may be degraded via biological processes such as nonsense-dependent degradation. Nonsense mutations can also result in transcripts that evade degradation processes such as nonsense-dependent degradation and can proceed to translation, but the translated polypeptide may be cleaved compared to the reference polypeptide sequence. The cleaved polypeptide sequence may consist of loss-of-function biological effects, such as low phenotype or null changes, for example, due to impaired biophysical stability of the polypeptide composition. In some relatively rare cases, variants may occur in polypeptides with novel phenotype functions, where the polypeptide possesses biological functions not found in the wild-type polypeptide or protein. In other relatively rare cases, genetic variants may result in polypeptides with high phenotype functions, where the polypeptide retains the biological functions found in the wild-type polypeptide or protein but has higher activity (e.g., increased enzyme activity, but not limited to). A missense mutation refers to a substitution in a genetic coding sequence in which the encoded amino acid in the polypeptide sequence is mutated to a different amino acid, and the different amino acid is also different from the reference sequence.Missense mutations can result in loss-of-function changes, such as low-level or null mutations, or gain-of-function changes, such as high-level or novel mutations, to the encoded polypeptide.
[0074] An insertion is a variant with a longer sequence length than the sequence shown in the reference sequence. Conversely, a deletion is a variant with a shorter sequence length than the sequence shown in the reference sequence. An indel refers to a variant that consists of both insertions and deletions. Indels can result in loss-of-function changes, such as low phenotypic changes or null changes, or gain-of-function changes, such as high phenotypic changes or novel phenotypic changes, to the encoded polypeptide. Insertions, deletions, or a combination of both can cause frameshift mutations. Frameshift mutations are mutations that alter the boundaries of trinucleotide codons relative to the reference sequence, and can result in cleaved polypeptides, or in some cases, elongated polypeptides, relative to the reference sequence. Frameshift mutations can result in loss-of-function changes, such as low phenotypic changes or null changes, or gain-of-function changes, such as high phenotypic changes or novel phenotypic changes, to the encoded polypeptide.
[0075] The steps shown in Figure 1 may have additional steps added. For example, the HLA variant can be aligned to an HLA code reference sequence to annotate the exon-intron boundaries in the HLA variant.
[0076] In some cases, for example, annotated HLA variants can be translated into HLA polypeptide sequences.
[0077] In some cases, for example, a translated HLA polypeptide sequence can be aligned to a reference HLA polypeptide sequence to identify amino acid changes in the translated HLA polypeptide sequence.
[0078] In some cases, translated HLA polypeptide sequences can be categorized as containing synonymous mutations, non-start mutations, non-stop mutations, nonsense mutations, or missense mutations.
[0079] Figure 2 shows an exemplary schematic diagram illustrating a general process 200 for identifying an HLA variant. Process 200 can be performed, for example, using one or more electronic devices implementing a software platform. In some examples, process 200 is performed using a client-server system, and blocks of process 200 are divided in any way between the server and client devices. In other examples, blocks of process 200 are divided between a server and multiple client devices. Therefore, while parts of process 200 are described herein as being performed by specific devices in a client-server system, it should be understood that process 200 is not limited in this way. In other examples, process 200 is performed using only one client device, or only multiple client devices. In process 200, some blocks are combined at will, the order of some blocks is changed at will, and some blocks are omitted at will. In some examples, additional steps may be performed in combination with process 200. Therefore, the behavior illustrated (and described in more detail below) is illustrative in nature and should not be considered limiting.
[0080] In some embodiments, methods for identifying HLA variants are disclosed herein, comprising: a) receiving sequence read data for a plurality of sequence reads in one or more processors; b) aligning the sequence read data to a first reference genome using one or more processors to identify a first plurality of HLA alignment reads and unmapped sequence reads; c) aligning the first plurality of HLA alignment reads and unmapped sequence reads to a second reference genome using one or more processors to identify a second plurality of HLA alignment reads; and d) processing the second plurality of HLA alignment reads using one or more processors to identify HLA variants, wherein the identification does not require the use of sequence read data from matched normal controls.
[0081] In step 202 of Figure 2, sequence read data is received for multiple sequence reads.
[0082] As shown, in some examples, multiple sequence reads may be derived from targeted sequencing techniques, e.g., targeted exome sequencing techniques. In some examples, sequence read data may be derived from, for example, whole-genome or whole-exome sequencing techniques, as opposed to targeted exome sequencing techniques, in order to increase the number of genomic features detected (e.g., the number of short variants). Targeted sequencing methods may be based on sequencing by synthetic techniques (e.g., but not limited to next-generation sequencing techniques). In addition, targeted sequencing methods may be based on hybridization techniques (e.g., but not limited to microarrays). Targeted sequencing methods may also be based on avidity sequencing (e.g., but not limited to the method described in Arslan et al., bioRxiv, 2022).
[0083] In some cases, the sample may include a tissue biopsy sample, a fluid biopsy sample, or a normal control. The sample may also be a fluid biopsy sample and may include blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In addition, the sample may be a fluid biopsy sample and may contain circulating tumor cells (CTCs). The sample may also contain cell-free DNA (cfDNA) and / or circulating tumor DNA (ctDNA).
[0084] In some cases, the subject may be a human being.
[0085] In step 204 of Figure 2, the sequence read data is aligned to a first reference genome to identify a first set of HLA alignment reads and unmapped sequence reads.
[0086] As shown, in some cases, the reference genome (e.g., a first reference genome, but not limited to) may be human reference genome Hg19 or human reference genome Hg38.
[0087] In step 206 of Figure 2, the first set of HLA alignment reads and unmapped sequence reads are aligned to a second reference genome to identify a second set of HLA alignment reads.
[0088] In some cases, the second reference genome may be a target-specific reference genome. In some cases, a target-specific reference genome can be generated by searching an HLA polymorphism database using multiple aligned and unmapped sequence reads of the first multiple HLA alignment reads. In some cases, the HLA polymorphism database may be, for example, the IPD-IMGT / HLA database. In some cases, searching the polymorphism database may involve using a software tool that predicts the probability that an input sequence corresponds to a known HLA sequence. In some cases, the software tool may be, for example, Optitype.
[0089] In step 208 of Figure 2, a second set of multiple HLA alignment reads are processed to identify HLA variants, which does not require the use of sequence read data from matched normal controls.
[0090] In some cases, a variant caller can be used to process multiple HLA alignment reads and identify HLA variants. In some cases, the HLA variant may be an HLA-I variant. In some cases, the identified HLA variant may be a variant of HLA-A, HLA-B, and / or HLA-C. In some cases, the HLA variant may be an HLA-I and / or HLA-II variant. In some cases, the variant caller may include a short variant caller.
[0091] As described above, in some examples, HLA variants may consist of single nucleotide substitutions, insertions, deletions, or combinations thereof. In some examples, HLA variants may affect the function of the protein corresponding to the gene sequencing performed within the variant.
[0092] How to use In some examples, the disclosed method involves (i) obtaining a sample from a subject (e.g., a subject suspected of having or determined to have cancer), (ii) extracting nucleic acid molecules (e.g., a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules) from the sample, (iii) ligating one or more adapters (e.g., one or more amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences) to the nucleic acid molecules extracted from the sample, and (iv) carrying out a methylation conversion reaction to, for example, convert unmethylated cytosine to uraci (v) a step of converting to a nucleotide, (v) a step of amplifying the nucleic acid molecule (e.g., using polymerase chain reaction (PCR) amplification techniques, non-PCR amplification techniques, or isothermal amplification techniques), (vi) a step of capturing the nucleic acid molecule from the amplified nucleic acid molecule (e.g., by hybridization to one or more bait molecules, where each bait molecule contains one or more nucleic acid molecules, each containing a region complementary to the region of the captured nucleic acid molecule), (vii) a step of, for example, next-generation (ultra-parallel) sequencing techniques, whole-genome sequencing (WGS) techniques, whole-exome sequencing techniques, targeted sequencing techniques (viii) the steps of sequencing nucleic acid molecules extracted from a sample (or a library proxy derived therefrom) using, for example, a next-generation (super-parallel) sequencer (using, for example, a next-generation (super-parallel) sequencer), nucleic acid sequence data (including, for example, variant data, copy number data, methylation status data of the sequenced nucleic acid molecules, etc.) and other biomarker data modalities, including, but not limited to, proteomics-based biomarker data (e.g., detection of specific polypeptides, e.g., proteins) or fragment-mixes-based biomarker data (e.g., detection of specific attributes related to nucleic acid fragments, e.g., fragment size or fragment end sequence), for example, determining the presence of ctDNA in the sample and / or determining a diagnosis, prognosis, and / or treatment response prediction for the subject, and (ix) generating and displaying a report (e.g., electronic, web-based, or paper report) for the subject (or patient), caregiver, healthcare provider, physician, oncologist, electronic medical record system, hospital, clinic, third-party payer, insurance company, or government agency.The process may further include one or more steps of sending and / or transmitting. In some examples, the report includes output from the method described herein. In some examples, all or part of the report may be displayed in a graphical user interface of an online or web-based healthcare portal. In some examples, the report is transmitted over a computer network or peer-to-peer connection.
[0093] The disclosed method may be used with any of a variety of samples. For example, in some cases the sample may include a tissue biopsy sample, a liquid biopsy sample, or a normal control. In some cases the sample may be a liquid biopsy sample and may include blood, plasma, cerebrospinal fluid, sputum, feces, urine, or saliva. In some cases the sample may be a liquid biopsy sample and may include circulating tumor cells (CTCs). In some cases the sample may be a liquid biopsy sample and may include cell-free DNA (cfDNA). In some cases cell-free DNA (cfDNA) or a portion thereof may include circulating tumor DNA (ctDNA). In some cases the liquid biopsy sample may include a combination of cell-free DNA (cfDNA) and circulating tumor DNA (ctDNA).
[0094] In some cases, the nucleic acid molecules extracted from the sample may consist of a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules. In some cases, the tumor nucleic acid molecules may originate from the tumor portion of the xenotissue biopsy sample, and the non-tumor nucleic acid molecules may originate from the normal portion of the xenotissue biopsy sample. In some cases, the sample may consist of a liquid biopsy sample, and the tumor nucleic acid molecules may originate from the circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and the non-tumor nucleic acid molecules may originate from the non-tumor cell-free DNA (cfDNA) fraction of the liquid biopsy sample.
[0095] In some examples, the disclosed methods for identifying HLA variants may be used to diagnose (or as part of) the presence of a disease or other condition in a subject (e.g., a patient) (e.g., cancer, genetic disorders (such as Down syndrome and fragile X), neurological disorders, or any other disease type in which the detection of a variant, e.g., copy number variation, is associated with diagnosing, treating, or predicting such disease). In some examples, the disclosed methods may be applicable to the diagnosis of any of the various cancers, as described elsewhere in this Spec.
[0096] In some cases, the disclosed methods for identifying HLA variants may be used to predict hereditary disorders of fetal DNA (e.g., for invasive or non-invasive prenatal testing). For example, sequence read data obtained by sequencing fetal DNA extracted from samples obtained using invasive amniocentesis, chorionic villus sampling (cVS), or fetal umbilical cord sampling techniques, or from samples obtained using non-invasive sampling of cell-free DNA (cfDNA) samples (including mixtures of maternal and fetal cfDNA), may be processed according to the disclosed methods to identify variants associated with, for example, Down syndrome (trisomy 21), trisomy 18, trisomy 13, extra copies or deletions of X and Y chromosomes, such as copy number variations.
[0097] In some cases, the disclosed methods for identifying HLA variants may be used to select subjects (e.g., patients) for clinical trials based on the identified HLA variants. In some cases, patient selection for clinical trials based on the identification of specific HLA variants at one or more loci may accelerate the development of targeted therapies and improve the medical outcomes of treatment decisions.
[0098] In some examples, the disclosed methods for identifying HLA variants may be used to select an appropriate therapy or treatment (e.g., anti-cancer therapy or anti-cancer treatment) for the subject. In some examples, the anti-cancer therapy or treatment may include the use of poly(ADP-ribose) polymerase inhibitors (PARPi), platinum compounds, chemotherapy, radiotherapy, targeted therapy, immunotherapy, neoantigen-based therapy, surgery, or any combination thereof.
[0099] In some examples, anticancer therapy or treatment may include targeted anticancer therapy or treatment (e.g., monoclonal antibody-based therapy, enzyme inhibitor-based therapy, antibody-drug conjugate therapy, hormone therapy, and / or targeted radiotherapy) that targets specific molecules necessary for the growth, division, and expansion of cancer cells. In some examples, targeted anticancer therapy or treatment may include abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), ado-trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), and amivantamab-vmjw (Ryb Revant, anastrozole (Arimidex), apalutamide (Erleada), aciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicaptagensilolucel (Yescarta), axitinib (Inlyta), verantamab mahodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), Verzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexcabutadiene oatlucel (Tecartus), brigutinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabom etyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Ilaris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), semiprimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro),Daratumumab (Darzalex), Daratumumab and hyaluronidase-fihj (Darzalex Faspro), Darorutamide (Nubeqa), Dasatinib (Sprycel), Deniroikin difutitox (Ontak), Denosumab (Xgeva), Dinutuximab (Unituxin), Dostallumab-gxly (Jemperli), Durvalumab (Imfinzi), Dubellisib (Copiktra), Elotuzumab (Empliciti), Enasidenib mesylate (Idhifa), Enco Rafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedra hydrochloride Tinib (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glass-degib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idekabuta genbiculose Abecma, Idelalisib (Zydelig), Imatinib mesylate (Gleevec), Infiglatinib phosphate (Truseltiq), Inotuzumab ozogamicin (Besponsa), Ipilimumab (Yervoy), Isatuximab-irfc (Sarclisa), Ivosidenib (Tibsovo), Ixazomib citrate (Ninlaro), Lanreotide acetate (Somatuline) Depot), lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), lenvatinib mesylate (Lenvima), letrozole (Femara), lysocabategemmaralucel (Breyanzi), loncustuximab tesillin-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu177-doteate (Lutathera), margetuximab-cmkb (Margenza),Midostaurin (Rydapt), Mobocertinib succinate (Exkivity), Mogamulizumab-kpkc (Poteligeo), Moxetumomab Pasdotox-tdfk (Lumoxiti), Naxitamab-gqgk (Danyelza), Necitumumab (Portrazza), Neratinib maleate (Nerlynx), Nilotinib (Tasigna), Niraparib tosylate monohydrate (Zejula), Nivolumab (Opdivo), Obinutuzumab (Gazyva), Ofatumumab (Arzerra), Olaparib (Lynparza), Oraratumab (Lartruvo), Osimertinib (Tagrisso), Palbociclib (Ibrance), Panit Mumab (Vectibix), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pexidartinib hydrochloride (Turalio), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), pralcetinib (Gavreto), radium-223 chloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), lipretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase (Rituxan) Hycela), Romidepsin (Istodax), Rucaparib cansylate (Rubraca), Ruxolitinib phosphate (Jakafi), Sacituzumab govitecan-hziy (Trodelvy), Celiclib, Selinexol (Xpovio), Serpercatinib (Retevmo), Selumetinib sulfate (Koselugo), Siltuximab (Sylvant), Sirolimus-binding protein particles (Fyarro), So Nidezib (Odomzo), sorafenib (Nexavar), sotracib (Lumakras), sunitinib (Sutent), tafacitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), thalazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), teventafusp-tebn (Kimmtrak),Temsirolimus (Torisel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tocitumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), tretinoin It may contain fen (Fareston), tucatinib (Tukysa), umbralicib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), bismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), ziv-aflibercept (Zaltrap), or any combination thereof.
[0100] In some cases, anti-cancer therapy or treatment may include immunotherapy (e.g., cancer treatment that acts by stimulating the immune system to fight cancer). In some cases, immunotherapy may include, for example, immune system modulators (e.g., cytokines, e.g., interferon or interleukin), immune checkpoint inhibitors (e.g., anti-PD-1 or anti-PD-L1 antibodies), T-cell transfer therapy (e.g., tumor-infiltrating lymphocyte (TIL) therapy in which lymphocytes extracted from the patient's tumor are selected for their ability to recognize tumor cells and proliferate before reintroduction into the patient, or CAR T-cell therapy in which the patient's T cells are modified to express CAR proteins before reintroduction into the patient), monoclonal antibody-based therapy (e.g., monoclonal antibodies that bind to cell surface markers on cancer cells to facilitate recognition by the immune system), or cancer-treating vaccines (e.g., vaccines based on tumor cells, tumor-associated neoantigens, or dendritic cells, etc., that stimulate the immune system to fight cancer).
[0101] In some cases, as described above, anticancer therapy or treatment may include neoantigen-based therapies. Non-limiting examples of neoantigen-based therapies include T-cell receptor (TCR) engineered T-cell (TCR-T) therapy, chimeric antigen receptor T-cell (CAR-T) therapy, TCR bispecific antibody therapy, and cancer vaccines. TCR-T therapy is produced by genetically engineering a patient's T cells to express T cells specific to the neoantigen of interest, and then injecting them back into the patient. CAR-T therapy is produced by genetically engineering a patient's T cells to express a chimeric antigen receptor molecule containing intracellular signaling and co-signaling domains and an extracellular antigen-binding domain. CAR-T therapy does not always rely on neoantigen presentation, but can be designed to be neoantigen-directed. TCR bispecific antibody therapy is a small, engineered antibody molecule containing a neoantigen-specific TCR at one end and a CD3-directed single-chain variable fragment at the other end. Cancer vaccines may contain RNA molecules, DNA molecules, peptides, or combinations thereof, designed to enhance the immune system's ability to locate and destroy neoantigen-presenting cells.
[0102] In some cases, the disclosed methods for identifying HLA variants may be used in treating a disease (e.g., cancer) in a subject. For example, an effective dose of anticancer therapy or anticancer treatment may be used in response to identifying an HLA variant using any of the methods disclosed herein.
[0103] In some examples, the disclosed method for identifying HLA variants may be used to monitor disease progression or recurrence (e.g., progression or recurrence of cancer or tumor) in a subject. For example, in some examples, the method may be used to identify HLA variants in a first sample obtained from a subject at a first time point, and may also be used for HLA variants in a second sample obtained from a patient at a second time point, and the comparison of the first and second identifications of HLA variants will enable monitoring of disease progression or recurrence. In some examples, the first time point is selected before the subject receives therapy or treatment, and the second time point is selected after the subject receives therapy or treatment.
[0104] In some examples, the disclosed methods may be used to adjust a therapy or treatment for a subject (e.g., anti-cancer therapy or anti-cancer treatment) by, for example, adjusting the therapeutic dose and / or selecting a different therapy in response to a specific change in an HLA variant.
[0105] In some examples, the presence of one or more HLA variants identified using the disclosed method may be used as a prognostic or diagnostic indicator related to the sample. For example, in some examples, the prognostic or diagnostic indicator may include an indicator of the presence of a disease (e.g., cancer) in the sample, an indicator of the probability that a disease (e.g., cancer) is present in the sample, an indicator of the probability that the subject from which the sample originates will develop a disease (e.g., cancer) (i.e., a risk factor), or an indicator of the likelihood that the subject from which the sample originates will respond to a particular therapy or treatment.
[0106] In some examples, the disclosed methods for identifying HLA variants may be implemented as part of a genomic profiling process that includes identifying the presence of variant sequences at one or more loci in a sample derived from a subject, as part of the detection, monitoring, risk factor prediction, or treatment selection for a specific disease, e.g., cancer. In some examples, the variant panel selected for genomic profiling may include the detection of variant sequences at a selected set of loci. In some examples, the variant panel selected for genomic profiling may include the detection of variant sequences at several loci via comprehensive genomic profiling (CGP) (a next-generation sequencing (NGS) approach used to evaluate hundreds of genes (including associated cancer biomarkers) in a single assay). The inclusion of the disclosed methods for identifying HLA variants as part of a genomic profiling process (or the output from the disclosed methods for identifying HLA variants as part of a subject's genomic profile) can improve the validity of, for example, disease detection calls and treatment decisions made based on the genomic profile by independently confirming the presence of one or more HLA variants in a given patient sample.
[0107] In some cases, a genomic profile may include information about the presence of genes (or their variant sequences), copy number variations, epigenetic traits, proteins (or their modifications), and / or other biomarkers in an individual's genome and / or proteome, as well as the individual's corresponding phenotypic traits, and information about the interactions between genetic or genomic traits, phenotypic traits, and environmental factors.
[0108] In some cases, the genome profile in question may include results from comprehensive genome profiling (CGP) studies, nucleic acid sequencing-based studies, gene expression profiling studies, cancer hotspot panel studies, DNA methylation studies, DNA fragmentation studies, RNA fragmentation studies, or any combination thereof.
[0109] In some cases, the method may further include administering or applying a treatment or therapy (e.g., an anticancer agent, anticancer treatment, or anticancer therapy) based on the generated genomic profile. An anticancer agent or anticancer treatment may refer to a compound that is effective in treating cancer cells. Examples of anticancer agents or anticancer therapies include, but are not limited to, alkylating agents, antimetabolites, natural products, hormones, chemotherapy, radiotherapy, immunotherapy, surgery, or therapies configured to target defects in specific cellular signaling pathways, such as defects in the DNA mismatch repair (MMR) pathway.
[0110] sample The disclosed methods and systems may be used with any of a variety of samples (also referred to herein as specimens) containing nucleic acids (e.g., DNA or RNA) collected from a subject (e.g., a patient). Examples of samples include, but are not limited to, tumor samples, tissue samples, biopsy samples (e.g., tissue biopsy, fluid biopsy, or both), blood samples (e.g., peripheral whole blood samples), plasma samples, serum samples, lymph samples, saliva samples, sputum samples, urine samples, gynecological fluid samples, circulating tumor cell (CTC) samples, cerebrospinal fluid (CSF) samples, pericardial fluid samples, pleural fluid samples, ascites (peritoneal fluid) samples, fecal (or stool) samples, or other bodily fluid, secretion, and / or excretory samples (or cell samples derived therefrom). In certain examples, the sample may be a frozen sample or a formalin-fixed paraffin-embedded (FFPE) sample.
[0111] In some cases, specimens may be collected by tissue excision (e.g., surgical excision), needle biopsy, bone marrow biopsy, bone marrow aspiration, skin biopsy, endoscopic biopsy, fine-needle aspiration, oral swab, nasal swab, vaginal swab, or cytological smear, abrasion, lavage, or lavage fluid (such as tubular lavage fluid or bronchoalveolar lavage fluid).
[0112] In some cases, the sample is a liquid biopsy sample and may include, for example, whole blood, plasma, serum, urine, stool, sputum, saliva, or cerebrospinal fluid. In some cases, the sample may be a liquid biopsy sample and may contain circulating tumor cells (CTCs). In some cases, the sample may be a liquid biopsy sample and may contain cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.
[0113] In some examples, the sample may contain one or more pre-malignant or malignant cells. As used herein, pre-malignant tumor refers to cells or tissue that are not yet malignant but are ready to become malignant. In certain examples, the sample may be obtained from a solid tumor, a soft tissue tumor, or a metastatic lesion. In certain examples, the sample may be obtained from a hematological malignancy or a pre-malignant tumor. In other examples, the sample may contain tissue or cells from a surgical margin. In certain examples, the sample may contain tumor-infiltrating lymphocytes. In some examples, the sample may contain one or more non-malignant cells. In some examples, the sample may be a primary tumor or a metastasis (e.g., a metastatic biopsy sample), or a portion thereof. In some examples, the sample may be obtained from the site with the highest percentage of tumor cells (e.g., tumor site) compared to adjacent sites (e.g., sites adjacent to the tumor). In some cases, the sample may be obtained from the site containing the largest tumor lesion (e.g., the site with the largest number of tumor cells as viewed under a microscope) compared to adjacent sites (e.g., sites adjacent to the tumor).
[0114] In some examples, the disclosed method may further include analyzing a primary control (e.g., a normal tissue sample). In some examples, the disclosed method may further include determining whether a primary control is available and, if available, isolating a control nucleic acid (e.g., DNA) from the primary control. In some examples, the sample may include any normal control (e.g., normal adjacent tissue (NAT)) if a primary control is not available. In some examples, the sample may be or may include histologically normal tissue. In some examples, the method includes evaluating a sample, for example, a histologically normal sample (e.g., from a surgical tissue margin) using the method described herein. In some examples, the disclosed method may further include obtaining a partial sample enriched with non-tumor cells by, for example, macro-dissecting non-tumor tissue from the NAT in the sample without a primary control. In some examples, the disclosed method may further include determining that a primary control and NAT are not available and marking the sample for analysis without a matched control.
[0115] In some cases, samples obtained from histologically normal tissue (e.g., otherwise histologically normal tissue margins) may still contain genetic alterations, such as variant sequences, as described herein. Therefore, the method may further include reclassifying samples based on the presence of detected genetic alterations. In some cases, multiple samples (e.g., from different subjects) are processed simultaneously.
[0116] The disclosed methods and systems may be applied to the analysis of nucleic acids extracted from various tissue samples (or disease states thereof), such as solid tissue samples, soft tissue samples, metastatic lesions, or liquid biopsy samples. Examples of tissues include, but are not limited to, connective tissue, muscle tissue, nervous system tissue, epithelial tissue, and blood. Tissue samples may be collected from any organ in the animal or human body. Examples of human organs include, but are not limited to, the brain, heart, lungs, liver, kidneys, pancreas, spleen, thyroid gland, mammary glands, uterus, prostate, large intestine, small intestine, bladder, bone, and skin.
[0117] In some cases, nucleic acids extracted from a sample may contain deoxyribonucleic acid (DNA) molecules. Examples of DNA that may be suitable for analysis by the disclosed methods include, but are not limited to, genomic DNA or fragments thereof, mitochondrial DNA or fragments thereof, cell-free DNA (cfDNA), and circulating tumor DNA (ctDNA). Cell-free DNA (cfDNA) consists of DNA fragments released from normal and / or cancer cells during apoptosis and necrosis, circulating in the bloodstream and / or accumulating in other bodily fluids. Circulating tumor DNA (ctDNA) consists of DNA fragments released from cancer cells and tumors that circulate in the bloodstream and / or accumulate in other bodily fluids.
[0118] In some cases, DNA is extracted from nucleated cells in a sample. In some cases, the sample has low nucleated cell solidity, for example, if the sample consists mainly of red blood cells, diseased cells containing excess cytoplasm, or tissue with fibrosis. In some cases, a sample with low nucleated cell solidity may require more, for example, a larger tissue volume, for DNA extraction.
[0119] In some examples, nucleic acids extracted from a sample may include ribonucleic acid (RNA) molecules. Examples of RNA that may be suitable for analysis by the disclosed methods include, but are not limited to, total cellular RNA, total cellular RNA after depletion of a specific amount of RNA sequence (e.g., ribosomal RNA), cell-free RNA (cfRNA), messenger RNA (mRNA) or fragments thereof, poly(A) tail mRNA fraction of total RNA, ribosomal RNA (rRNA) or fragments thereof, transfer RNA (tRNA) or fragments thereof, and mitochondrial RNA or fragments thereof. In some examples, RNA may be extracted from a sample and converted to complementary DNA, for example, using a reverse transcription reaction. In some examples, cDNA is produced by a random-primed cDNA synthesis method. In other examples, cDNA synthesis is initiated at the poly(A) tail of mature mRNA by priming with an oligo(dT)-containing oligonucleotide. Methods for depletion, poly(A) enrichment, and cDNA synthesis are well known to those skilled in the art.
[0120] In some cases, the sample may contain tumor content (e.g., tumor cells or tumor cell nuclei) or non-tumor content (e.g., immune cells, fibroblasts, and other non-tumor cells). In some cases, the tumor content of the sample may constitute a sample index. In some cases, the sample may contain tumor content with at least 5-50%, 10-40%, 15-25%, or 20-30% tumor cell nuclei. In some cases, the sample may contain tumor content with at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% tumor cell nuclei. In some cases, the percentage of tumor nuclei (sample fraction) is determined (e.g., calculated) by dividing the number of tumor cells in the sample by the total number of all cells in the sample that have nuclei. In some cases, for example, when the sample is a liver sample containing hepatocytes, different tumor content calculations may be required due to the presence of hepatocytes with twice or more nuclei, other DNA content, such as non-hepatocytes, somatic cell nuclei. In some cases, the sensitivity to detecting genetic alterations, such as variant sequences, or the sensitivity to determining microsatellite instability, for example, may depend on the tumor content of the sample. For instance, a sample with a lower tumor content may result in lower sensitivity to detection for a sample of a given size.
[0121] In some examples, as described above, the sample includes nucleic acids (e.g., DNA, RNA (or cDNA derived from RNA), or both) from, for example, a tumor or normal tissue. In certain examples, the sample may further include non-nucleic acid components derived from, for example, a tumor or normal tissue, such as cells, proteins, carbohydrates, or lipids.
[0122] subject In some cases, the sample is obtained (e.g., collected) from a subject (e.g., a patient) who has or is suspected of having a certain condition or disease (e.g., a hyperproliferative disorder or a non-cancerous indicator). In some cases, the hyperproliferative disorder is cancer. In some cases, cancer is a solid tumor or its metastatic form. In some cases, cancer is a blood cancer, e.g., leukemia or lymphoma.
[0123] In some cases, the subject has cancer or is at risk of developing cancer. For example, in some cases, the subject has a genetic predisposition to cancer (e.g., having a gene mutation that increases the baseline risk of developing cancer). In some cases, the subject is exposed to environmental changes (e.g., radiation or chemicals) that increase the risk of developing cancer. In some cases, the subject needs to be monitored for the development of cancer. In some cases, the subject needs to be monitored for cancer progression or regression after treatment with anti-cancer therapy (or anti-cancer treatment), for example. In some cases, the subject needs to be monitored for cancer recurrence. In some cases, the subject needs to be monitored for minimal residual disease (MRD). In some cases, the subject has been treated for cancer or is being treated. In some cases, the subject has not been treated with anti-cancer therapy (or anti-cancer treatment).
[0124] In some cases, the subject (e.g., a patient) is being treated with or has been treated with one or more targeted therapies. In some cases, for example, a post-targeted therapy sample (e.g., a specimen) is obtained (e.g., collected) from a patient who has been previously treated with a targeted therapy. In some cases, the post-targeted therapy sample is a sample obtained after the completion of the targeted therapy.
[0125] In some cases, the patients have not been previously treated with targeted therapy. In some cases, for example, in patients who have not been previously treated with targeted therapy, the samples include excisions, e.g., original excisions, or excisions after recurrence (e.g., after disease recurrence following therapy).
[0126] cancer In some cases, samples are obtained from subjects with cancer. Exemplary cancers include, but are not limited to, B-cell carcinoma (e.g., multiple myeloma), melanoma, breast cancer, lung cancer (such as non-small cell lung cancer or NSCLC), bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, oral or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, and biliary tract cancer. , small intestine or adnexal cancer, salivary gland cancer, thyroid cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, hematological cancer, adenocarcinoma, inflammatory myofibroblastic tumor, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-Hypercytic leukemia Dikin's lymphoma (NHL), soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endosarcoma synoviomas, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocellular carcinoma, cholangiocarcinoma, choriocarcinoma, seminoma, embryonic carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma These include medulloblastoma, craniopharyngioma, ependymoma, pineal cell tumor, hemangioblastoma, acoustic neuroblastoma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, agnogenous myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familial eosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, and cancerous tumors.
[0127] In some cases, cancers include acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpression / amplification), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deficiency), chronic myeloid leukemia, chronic myeloid leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, and colorectal cancer. Cancer, colorectal cancer (dMMR and MSI-H), colorectal cancer (KRAS wild-type), cryopyrin-associated periodic fever syndrome, cutaneous T-cell lymphoma, dermatofibrosarcoma protuberans, diffuse large B-cell lymphoma, fallopian tube cancer, follicular B-cell non-Hodgkin lymphoma, follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, gastrointestinal stromal tumor, gastrointestinal stromal tumor (KIT+), giant cell tumor of bone, glioblastoma, granulomatosis with polyangiitis, head and neck squamous cell carcinoma, hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, systemic lupus erythematosus, mantle cell lymphoma, medullary thyroid carcinoma, melanoma, BRAF Melanoma with V600 mutation, melanoma with BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman disease, multiple hematological malignancies including Philadelphia chromosome-positive ALL and CML, multiple myeloma, myelofibrosis, non-Hodgkin lymphoma, unresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, non-small cell lung cancer, non-small cell lung cancer (ALK+), non-small cell lung cancer (PD-L1+), non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), non-small cell lung cancer (with BRAF V600E mutation), non-small cell lung cancer (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer (EGFRThis includes ovarian cancer (with T790M mutation), ovarian cancer (with BRCA mutation), pancreatic cancer, neuroendocrine tumors of pancreatic, gastrointestinal, or lung origin, pediatric neuroblastoma, peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, renal cell carcinoma, rheumatoid arthritis, small lymphocytic lymphoma, soft tissue sarcoma, solid tumors (MSI-H / dMMR), squamous cell carcinoma of the head and neck, squamous non-small cell lung cancer, thyroid cancer, thyroid carcinoma, urothelial carcinoma, or primary gammaglobulinemia.
[0128] In some cases, cancer is a hematological malignancy (or pre-malignancy). As used herein, hematological malignancy refers to tumors of hematopoietic or lymphoid tissue, such as tumors affecting the blood, bone marrow, or lymph nodes. Exemplary hematological malignancies include leukemia (e.g., acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), hairy cell leukemia, acute monocytic leukemia (AMoL), chronic myelomonocytic leukemia (CMML), juvenile myelomonocytic leukemia (JMML), or macrogranular lymphocytic leukemia), lymphoma (e.g., AIDS-associated lymphoma, cutaneous T-cell lymphoma, Hodgkin lymphoma (e.g., classical Hodgkin lymphoma or nodular lymphocyte-predominant lymphoma) and lymphoma (e.g., AIDS-associated lymphoma, cutaneous T-cell lymphoma, Hodgkin lymphoma (e.g., classical Hodgkin lymphoma or nodular lymphocyte-predominant lymphoma)). This includes, but is not limited to, primary central nervous system lymphomas (Hodgkin lymphoma type 1), mycosis fungoides, non-Hodgkin lymphomas (e.g., B-cell non-Hodgkin lymphomas (e.g., Burkitt lymphoma, small lymphocytic lymphoma / small cell lymphoma (CLL / SLL), diffuse large B-cell lymphoma, follicular lymphoma, immunoblastic large cell lymphoma, progenitor B-lymphoblastic lymphoma, or mantle cell lymphoma) or T-cell non-Hodgkin lymphomas (mycosis fungoides, anaplastic large cell lymphoma, or progenitor T-lymphoblastic lymphoma)), and primary central nervous system lymphomas.
[0129] Nucleic acid extraction and processing DNA or RNA can be extracted from tissue samples, biopsy samples, blood samples, or other bodily fluid samples using any of the various techniques known to those skilled in the art (see, for example, Example 1 of International Patent Application Publication No. 2012 / 092426, Tan, et al. (2009), “DNA, RNA, and Protein Extraction: The Past and The Present”, J. Biomed. Biotech. 2009:574398, the technical literature for the Maxwell® 16 LEV Blood DNA Kit (Promega Corporation, Madison, WI), and the Maxwell 16 Buccal Swab LEV DNA Purification Kit Technical Manual (Promega Literature #TM333, January 1, 2011, Promega Corporation, Madison, WI)). A protocol for RNA isolation is disclosed, for example, in Maxwell® 16 Total RNA Purification Kit Technical Bulletin (Promega Literature #TB351, August 2009, Promega Corporation, Madison, WI).
[0130] A typical DNA extraction procedure includes, for example, (i) collecting a fluid, cell, or tissue sample from which DNA will be extracted; (ii) disrupting the cell membrane (i.e., cell lysis), if necessary, to release DNA and other cytoplasmic components; (iii) treating the fluid or lysed sample with a concentrated salt solution to precipitate proteins, lipids, and RNA, followed by centrifugation to separate the precipitated proteins, lipids, and RNA; and (iv) purifying the DNA from the supernatant to remove any detergents, proteins, salts, or other reagents used during the cell membrane lysis step.
[0131] Cell membrane disruption can be carried out using various mechanical shearing techniques (e.g., French press or fine needle) or ultrasonic disruption techniques. The cell lysis step often involves the use of detergents and surfactants to dissolve lipids, cells, and nuclear membranes. In some examples, the lysis step may further include the use of proteases to disrupt proteins and / or RNases for digestion of RNA in the sample.
[0132] Examples of suitable techniques for DNA purification include, but are not limited to, (i) precipitation in ice-cold ethanol or isopropanol, followed by centrifugation (e.g., DNA precipitation which may be enhanced by increasing ionic strength, such as by adding sodium acetate); (ii) phenol-chloroform extraction, followed by centrifugation to separate the aqueous phase containing nucleic acids from the organic phase containing denatured proteins; and (iii) solid-phase chromatography in which nucleic acids are adsorbed onto a solid phase (e.g., silica or others) depending on the pH and salt concentration of the buffer.
[0133] In some cases, DNA-bound cell and histone proteins can be removed by adding proteases, by precipitating the proteins with sodium acetate or ammonium acetate, or by extraction with a phenol-chloroform mixture prior to the DNA precipitation step.
[0134] In some cases, DNA can be extracted using one of a variety of suitable commercially available DNA extraction and purification kits. Examples include, but are not limited to, the QIAamp (for isolation of genomic DNA from human samples) and DNAeasy (for isolation of genomic DNA from animal or plant samples) kits from Qiagen (Germantown, MD), or the Maxwell® and ReliaPrep® series from Promega (Madison, WI).
[0135] As described above, in some examples, the sample may involve formalin fixation (formaldehyde fixation or paraformaldehyde fixation) and paraffin-embedded (FFPE) tissue preparation. For example, an FFPE sample may be a tissue sample embedded in a substrate, such as an FFPE block. Methods for isolating nucleic acids (e.g., DNA) from formaldehyde-fixed or paraformaldehyde-fixed, paraffin-embedded (FFPE) tissues include, for example, Cronin, et al., (2004) Am J Pathol. 164(1):35-42, Masuda, et al., (1999) Nucleic Acids Res. 27(22):4436-4443, Specht, et al., (2001) Am J Pathol. 158(2):419-429, Ambion RecoverAll (trademark) Total Nucleic Acid Isolation Protocol (Ambion, Cat. No. AM1975, September 2008), Maxwell (registered trademark) 16 FFPE Plus LEV DNA Purification Kit Technical Manual (Promega Literature #TM349, February 2011), and EZNA (registered trademark) FFPE DNA Kit This is disclosed in the Handbook (OMEGA bio-tek, Norcross, GA, product numbers D3399-00, D3399-01, and D3399-02, June 2009) and the QIAamp® DNA FFPE Tissue Handbook (Qiagen, Cat. No. 37625, October 2007). For example, the RecoverAll® Total Nucleic Acid Isolation Kit solubilizes paraffin-embedded samples using xylene at high temperatures and captures nucleic acids by passing them through a glass fiber filter. The Maxwell® 16 FFPE Plus LEV DNA Purification Kit is used with the Maxwell® 16 Instrument to purify genomic DNA from 1-10 μm sections of FFPE tissue.DNA is purified using silica-clad paramagnetic particles (PMPs) and eluted at low elution volumes. The EZNA® FFPE DNA Kit uses a spin column and buffer system for genomic DNA isolation. The QIAamp® DNA FFPE Tissue Kit uses QIAamp® DNA Microtechnology for the purification of genomic and mitochondrial DNA.
[0136] In some examples, the disclosed method may further include determining or obtaining a yield value of nucleic acids extracted from a sample and comparing the determined value to a reference value. For example, if the determined or obtained value is less than the reference value, the nucleic acids may be amplified before proceeding with library construction. In some examples, the disclosed method may further include determining or obtaining a value for the size (or average size) of nucleic acid fragments in the sample and comparing the determined or obtained value to a reference value, for example, a size (or average size) of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 base pairs (bps). In some examples, one or more parameters described herein may be adjusted or selected in response to this determination.
[0137] After isolation, nucleic acids are typically dissolved in a slightly alkaline buffer, such as Tris-EDTA(TE) buffer or ultrapure water. In some cases, isolated nucleic acids (e.g., genomic DNA) can be fragmented or sheared using any of the various techniques known to those skilled in the art. For example, genomic DNA can be fragmented by physical shearing, enzymatic cleavage, chemical cleavage, and other methods well known to those skilled in the art. A method for DNA shearing is described, for example, in Example 4 of International Patent Application Publication No. 2012 / 092426. In some cases, alternative methods to DNA shearing can be used to avoid the ligation step during library preparation.
[0138] Library preparation In some cases, nucleic acids isolated from a sample may be used to construct a library (e.g., the nucleic acid library described herein). In some cases, nucleic acids are fragmented using one of the methods described above, optionally subjected to repair of strand-end damage, optionally ligated to synthesize adapters, primers, and / or barcodes (e.g., amplification primers, sequencing adapters, flow cell adapters, substrate adapters, sample barcodes or indices, and / or unique molecular identifier sequences), sizing selected (e.g., by preparative gel electrophoresis), and / or amplified (e.g., using PCR, non-PCR amplification techniques, or isothermal amplification techniques). In some cases, the fragmented and adapter-ligated nucleic acid group is used without explicit sizing or amplification prior to hybridization-based selection of target sequences. In some cases, nucleic acids are amplified by one of a variety of specific or non-specific nucleic acid amplification methods well known to those skilled in the art. In some cases, nucleic acids are amplified by whole-genome amplification methods, such as random prime strand substitution amplification. Examples of nucleic acid library preparation techniques for next-generation sequencing are described, for example, in van Dijk, et al. (2014), Exp. Cell Research 322:12-20, and in Illumina's genomic DNA sample preparation kit.
[0139] In some examples, the resulting nucleic acid library may contain all or substantially all of the complexity of the genome. In this context, the term “substantially all” actually refers to the possibility that there may be some undesirable loss of genomic complexity during the initial steps of the procedure. The methods described herein are also useful when the nucleic acid library is part of a genome, for example, when the complexity of the genome is reduced by design. In some examples, any selected portion of the genome may be used in conjunction with the methods described herein. For example, in certain embodiments, an entire exome or a subset thereof is isolated. In some examples, the library may contain at least 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5% of genomic DNA. In some examples, the library may consist of cDNA copies of genomic DNA containing at least 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5% of genomic DNA. In certain cases, the amount of nucleic acid used to generate a nucleic acid library may be less than 5 micrograms, less than 1 microgram, less than 500 ng, less than 200 ng, less than 100 ng, less than 50 ng, less than 10 ng, less than 5 ng, or less than 1 ng.
[0140] In some examples, a library (e.g., a nucleic acid library) contains a collection of nucleic acid molecules. As described herein, the nucleic acid molecules in a library may include target nucleic acid molecules (e.g., tumor nucleic acid molecules, reference nucleic acid molecules and / or control nucleic acid molecules, also referred herein to as the first, second and / or third nucleic acid molecules, respectively). The nucleic acid molecules in a library may originate from a single subject or individual. In some examples, a library may contain nucleic acid molecules originating from two or more subjects (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30 or more subjects). For example, two or more libraries originating from different subjects may be combined to form a library having nucleic acid molecules from two or more subjects (the nucleic acid molecules originating from each subject are optionally ligated to a unique sample barcode corresponding to a particular subject). In some examples, the subjects are humans who have or are at risk of having cancer or tumors.
[0141] In some cases, a library (or a portion thereof) may contain one or more subgenome segments. In some cases, a subgenome segment may be a single nucleotide location, for example, a nucleotide location where a variant is associated (positively or negatively) with a tumor phenotype. In some cases, a subgenome segment may contain two or more nucleotide locations. Such examples include sequences of nucleotide locations with lengths of at least 2, 5, 10, 50, 100, 150, 250, or more than 250. A subgenome segment may contain, for example, one or more whole genes (or portions thereof), one or more exons or coding sequences (or portions thereof), one or more introns (or portions thereof), one or more microsatellite regions (or portions thereof), or any combination thereof. A subgenome segment may contain all or part of a naturally occurring nucleic acid molecule, for example, a fragment of a genomic DNA molecule. For example, a subgenome segment may correspond to a fragment of genomic DNA subjected to a sequencing reaction. In some cases, a subgenome segment is a continuous sequence from a genomic source. In some cases, subgenome segments contain non-contiguous sequences within the genome; for example, subgenome segments in cDNA may contain exon-exon junctions formed as a result of splicing. In some cases, subgenome segments contain tumor nucleic acid molecules. In some cases, subgenome segments contain non-tumor nucleic acid molecules.
[0142] Target gene loci for analysis The methods described herein may be used, for example, in combination with, or as part of, a method for evaluating a set of target intervals (e.g., target sequences) from a set of genomic loci (e.g., loci or fragments thereof), as described herein.
[0143] In some examples, the set of genomic loci evaluated by the disclosed method includes a plurality of genes, e.g., that are associated in variant forms with effects on cell division, growth, or survival, or with cancer, e.g., cancer as described herein.
[0144] In some examples, the set of loci evaluated by the disclosed method includes at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, or more than 100 loci.
[0145] In some examples, the selected locus (also referred to herein as the target locus or target sequence) or fragment thereof may include a target interval comprising non-coding sequences, coding sequences, intra-gene regions, or inter-gene regions of the target genome. For example, the target interval may include non-coding sequences or fragments thereof (e.g., promoter sequences, enhancer sequences, 5' untranslated regions (5'UTR), 3' untranslated regions (3'UTR), or fragments thereof), coding sequences of those fragments, exon sequences or fragments thereof, intron sequences or fragments thereof.
[0146] Target capture reagent Methods described herein may involve contacting a nucleic acid library with a plurality of target capture reagents to select and capture a plurality of specific target sequences (e.g., gene sequences or fragments thereof) for analysis. In some examples, the target capture reagent (i.e., a molecule that binds to a target molecule and thereby enables the capture of the target molecule) is used to select the target segment to be analyzed. For example, the target capture reagent may be a bait molecule, e.g., a nucleic acid molecule (e.g., a DNA molecule or RNA molecule), which hybridizes to (i.e., is complementary to) the target molecule and thereby enables the capture of the target nucleic acid. In some examples, the target capture reagent, e.g., the bait molecule (or bait sequence), is a capture oligonucleotide (or capture probe). In some examples, the target nucleic acid is a genomic DNA molecule, an RNA molecule, a cDNA molecule derived from an RNA molecule, a microsatellite DNA sequence, etc. In some examples, the target capture reagent is suitable for solution-phase hybridization of the target. In some examples, the target capture reagent is suitable for solid-phase hybridization of the target. In some examples, target capture reagents are suitable for both solution-phase and solid-phase hybridization of targets. The design and construction of target capture reagents are described in detail, for example, in International Patent Application Publication No. 2020 / 236941, the entire contents of which are incorporated herein by reference.
[0147] The methods described herein provide optimized sequencing of numerous genomic loci (e.g., genes or gene products (e.g., mRNA), microsatellite loci, etc.) from samples from one or more subjects (e.g., cancer tissue samples, liquid biopsy samples, etc.) by appropriate selection of target capture reagents for selecting target nucleic acid molecules to be sequenced. In some examples, the target capture reagent may hybridize to a specific target locus, e.g., a specific target locus or fragment thereof. In some examples, the target capture reagent may hybridize to a specific group of target loci, e.g., a specific group of loci or fragment thereof. In some examples, multiple target capture reagents may be used, including a mixture of target-specific and / or group-specific target capture reagents.
[0148] In some examples, the number of target capture reagents (e.g., bait molecules) in multiple target capture reagents (e.g., bait sets) that come into contact with a nucleic acid library to capture multiple target sequences for nucleic acid sequencing is greater than 10, greater than 50, greater than 100, greater than 200, greater than 300, greater than 400, greater than 500, greater than 600, greater than 700, greater than 800, greater than 900, greater than 1,000, greater than 1,250, greater than 1,500, greater than 1,750, greater than 2,000, greater than 3,000, greater than 4,000, greater than 5,000, greater than 10,000, greater than 25,000, or greater than 50,000.
[0149] In some examples, the total length of the target capture reagent sequence may be approximately 70 to 1000 nucleotides. In one example, the length of the target capture reagent is approximately 100 to 300 nucleotides, 110 to 200 nucleotides, or 120 to 170 nucleotides. In addition to the above, intermediate oligonucleotide lengths of approximately 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800, and 900 nucleotides can be used in the method described herein. In some embodiments, oligonucleotides with approximately 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, or 230 bases can be used.
[0150] In some examples, each target capture reagent sequence may include (i) a target-specific capture sequence (e.g., a locus- or microsatellite locus-specific complementary sequence), (ii) an adapter, primer, barcode, and / or a unique molecular identifier sequence, and (iii) a universal tail at one or both ends. As used herein, the term “target capture reagent” may refer to a target-specific target capture sequence or the entire target capture reagent oligonucleotide containing the target-specific target capture sequence.
[0151] In some examples, the target-specific capture sequence in the target capture reagent is approximately 40 to 1000 nucleotides long. In some examples, the target-specific capture sequence is approximately 70 to 300 nucleotides long. In some examples, the target-specific sequence is approximately 100 to 200 nucleotides long. In yet other examples, the target-specific sequence is approximately 120 to 170 nucleotides long, typically 120 nucleotides long. In addition to the above, target-specific sequences of intermediate lengths, for example, approximately 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800, and 900 nucleotides long, as well as target-specific sequences of lengths between the above lengths, may also be used in the methods described herein.
[0152] In some cases, target capture reagents may be designed to select a target interval containing one or more rearrangements, such as an intron containing a genomic rearrangement. In such cases, the target capture reagent is designed so that repetitive sequences are masked to enhance selection efficiency. In these cases, where the rearrangement has a known ligature sequence, a complementary target capture reagent can be designed for the ligature sequence to enhance selection efficiency.
[0153] In some examples, the disclosed methods may involve the use of target capture reagents designed to capture two or more different target categories, each category having a different target capture reagent design strategy. In some examples, the hybridization-based capture methods and target capture reagent compositions disclosed herein provide capture and homogeneous coverage of a target sequence set, while minimizing coverage of genomic sequences outside the targeted sequence set. In some examples, the target sequences may include an entire exome of genomic DNA or a selected subset thereof. In some examples, the target sequences may include, for example, a large chromosomal region (e.g., an entire chromosome arm). The methods and compositions disclosed herein provide different target capture reagents for achieving different sequencing depth and coverage patterns for a composite target nucleic acid sequence set.
[0154] Typically, DNA molecules are used as target capture reagent sequences, but RNA molecules can also be used. In some examples, the DNA molecular target capture reagent may be single-stranded DNA (ssDNA) or double-stranded DNA (dsDNA). In some examples, RNA-DNA double helix is more stable than DNA-DNA double helix and therefore potentially provides better nucleic acid capture.
[0155] In some examples, the disclosed method includes providing a selected set of nucleic acid molecules (e.g., a library catch) captured from one or more nucleic acid libraries. For example, the method may include providing one or more nucleic acid libraries, each containing multiple nucleic acid molecules (e.g., multiple target nucleic acid molecules and / or reference nucleic acid molecules) extracted from one or more samples from one or more subjects; contacting one or more libraries (e.g., in a solution-based hybridization reaction) with one, two, three, four, five, or more than five target capture reagents (e.g., oligonucleotide target capture reagents) to form a hybridization mixture containing multiple target capture reagent / nucleic acid molecule hybrids; and separating the multiple target capture reagent / nucleic acid molecule hybrids from the hybridization mixture by, for example, contacting the hybridization mixture with a binding entity that enables the separation of the multiple target capture reagent / nucleic acid molecule hybrids from the hybridization mixture, thereby providing a library catch (e.g., a selected or concentrated subgroup of nucleic acid molecules from one or more libraries).
[0156] In some examples, the disclosed method may further include amplifying the library catch (e.g., by performing PCR). In other examples, the library catch is not amplified.
[0157] In some cases, the target capture reagent may be part of a kit that may include instructions, standards, buffers, enzymes, or other reagents as needed.
[0158] Hybridization conditions As described above, the methods disclosed herein may include contacting a library (e.g., a nucleic acid library) with a selection of target capture reagents to contact a library target nucleic acid sequence (i.e., library catch). The contact step may be carried out, for example, by solution-based hybridization. In some examples, the method includes repeating the hybridization step with respect to one or more additional solution-based hybridizations. In some examples, the method further includes subjecting the library catch to one or more additional solution-based hybridizations with the same or different sets of target capture reagents.
[0159] In some examples, the contact step is carried out using a solid support, such as an array. Suitable solid supports for hybridization are described, for example, in Albert, T.Jet al. (2007) Nat. Methods 4(11):903-5, Hodges, E. et al. (2007) Nat. Genet. 39(12):1522-7, and Okou, D. et al. (2007) Nat. Methods 4(11):907-9, the contents of which are incorporated herein by reference in their entirety.
[0160] Hybridization methods that can be adapted for use in the methods described herein are described in the art, for example, in International Patent Application Publication No. 2012 / 092426. Methods for hybridizing a target capture reagent to multiple target nucleic acids are described in detail, for example, in International Patent Application Publication No. 2020 / 236941, the entirety of which is incorporated herein by reference.
[0161] Sequence determination method The methods and systems disclosed herein can be used in combination with, or as part of, a method or system for sequencing nucleic acids (e.g., a next-generation sequencing system) to generate multiple sequence reads that overlap with one or more loci within a subgenome section in a sample, thereby enabling, for example, the determination of gene alleles at multiple loci. The term “next-generation sequencing” (or “NGS”) as used herein may also be referred to as “massively parallel sequencing” (or “MPS”), which involves sequencing the nucleotide sequences of individual nucleic acid molecules (e.g., in single-molecule sequencing) or cloned proxies of individual nucleic acid molecules in a high-throughput manner (e.g., 10¹⁶). 3 , 10 4 , 10 5 , or 10 5 This refers to any sequencing method that determines the sequence of molecules (where more than a certain number of molecules are sequenced simultaneously).
[0162] Next-generation sequencing methods are known in the art and are described, for example, in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46, which is incorporated herein by reference. Other examples of sequencing methods suitable for use when implementing the methods and systems disclosed herein are described, for example, in International Patent Application Publication No. 2012 / 092426. In some examples, sequencing may include, for example, whole-genome sequencing (WGS), whole-exome sequencing, targeted sequencing, or direct sequencing. In some examples, sequencing may be performed using, for example, Sanger sequencing. In some examples, sequencing may include paired-end sequencing techniques that enable sequencing of both ends of a fragment and generate high-quality alignable sequence data for, for example, the detection of genome rearrangements, repetitive sequence elements, gene fusions, and novel transcripts.
[0163] The disclosed methods and systems may be implemented using sequencing platforms such as Roche 454, Illumina Solexa, ABI-SOLiD, ION Torrent, Complete Genomics, Pacific Bioscience, Helicos, and / or Polonator platforms. In some examples, sequencing may include Illumina MiSeq sequencing. In some examples, sequencing may include Illumina HiSeq sequencing. In some examples, sequencing may include Illumina NovaSeq sequencing. Optimized methods for sequencing multiple target genomic loci in nucleic acids extracted from a sample are described in detail, for example, in International Patent Application Publication No. 2020 / 236941, the entire contents of which are incorporated herein by reference.
[0164] In a particular example, the disclosed method is to (a) obtain a library from a sample containing multiple normal and / or tumor nucleic acid molecules; (b) contact the library simultaneously or sequentially with one, two, three, four, five, or more than five target capture reagents under conditions that enable hybridization of the target capture reagent to the target nucleic acid molecules, thereby providing a selected set of captured normal and / or tumor nucleic acid molecules (i.e., library catch); and (c) provide a selected subset of nucleic acid molecules (e.g., library catch) by, for example, contacting the hybridization mixture with a binding entity that enables separation of the target capture reagent / nucleic acid molecule hybrid from the hybridization mixture. The process includes (d) separating the bridging mixture, (e) sequencing the library catch to obtain multiple reads (e.g., sequence reads) that overlap with one or more target segments (e.g., one or more target sequences) from the library catch, which may contain mutations (or alterations), such as somatic or germline mutations, (e) aligning the sequence reads using an alignment method as described elsewhere in this Spec, and / or (f) assigning nucleotide values to nucleotide positions within the target segment from one or more sequence reads (e.g., calling mutations using a Bayesian method or other method as described herein).
[0165] In some examples, obtaining sequence reads for one or more target segments may involve sequencing at least 1, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1,000, at least 1,250, at least 1,500, at least 1,750, at least 2,000, at least 2,250, at least 2,500, at least 2,750, at least 3,000, at least 3,500, at least 4,000, at least 4,500, or at least 5,000 loci, such as genomic loci, genomic loci, microsatellite loci, etc. In some examples, obtaining sequence reads for one or more target intervals may involve sequencing target intervals for any number of loci within the range described in this paragraph, e.g., at least 2,850 loci.
[0166] In some examples, obtaining sequence reads for one or more target segments involves sequencing the target segments using a sequencing method that provides sequence read lengths (or average sequence read lengths) of at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, or 400 bases. In some examples, obtaining sequence reads for one or more target segments may involve sequencing the target segments using a sequencing method that provides sequence read lengths (or average sequence read lengths) of any number of bases within the range described in this paragraph, for example, a sequence read length (or average sequence read length) of 56 bases.
[0167] In some examples, obtaining sequence reads for one or more target intervals may include sequencing with an average coverage (or depth) of at least 100x. In some examples, obtaining sequence reads for one or more target intervals may include sequencing with an average coverage (or depth) of at least 100x, at least 150x, at least 200x, at least 250x, at least 500x, at least 750x, at least 1,000x, at least 1,500x, at least 2,000x, at least 2,500x, at least 3,000x, at least 3,500x, at least 4,000x, at least 4,500x, at least 5,000x, at least 5,500x, or at least 6,000x. In some examples, obtaining sequence reads for one or more target intervals may include sequencing with an average coverage (or depth) of at least 160x, for example, having any value within the range of values described in this paragraph.
[0168] In some examples, obtaining sequence reads for one or more target intervals involves sequencing approximately 90%, 92%, 94%, 95%, 96%, 97%, 98%, or more than 99% of sequenced loci at an average sequencing depth of at least 100× to at least 6,000×. For example, in some examples, obtaining reads for a target interval involves sequencing at least 125× at an average sequencing depth for at least 99% of sequenced loci. In another example, in some examples, obtaining reads for a target interval involves sequencing at least 4,100× at an average sequencing depth for at least 95% of sequenced loci.
[0169] In some cases, the relative abundance of nucleic acid species in a library can be estimated by counting the relative number of occurrences of those congeneral sequences in the data generated by sequencing experiments (e.g., the number of sequence reads for a given congeneral sequence).
[0170] In some examples, the disclosed methods and systems provide nucleotide sequences for a set of target intervals (e.g., loci) as described herein. In certain specific cases, the sequences are provided without using methods that include matched normal controls (e.g., wild-type controls) and / or matched tumor controls (e.g., primary versus metastatic).
[0171] In some examples, as used herein, the level of sequencing depth (e.g., X-fold level of sequencing depth) refers to the number of reads (e.g., unique reads) obtained after the detection and removal of duplicate reads (e.g., PCR duplicate reads). In other examples, duplicate reads are evaluated, for example, to aid in the detection of copy number variation (CNA).
[0172] alignment Alignment is the process of matching a read to a specific location, such as a genomic location or locus. In some cases, NGS reads may be aligned to a known reference sequence (e.g., a wild-type sequence). In some cases, NGS reads may be de novo assembled. Methods for sequence alignment of NGS reads are described, for example, in Trapnell, C. and Salzberg, SLNature Biotech., 2009, 27:455-457. Examples of de novo sequence assembly are described, for example, in Warren R. et al., Bioinformatics, 2007, 23:500-501, Butler J. et al., Genome Res., 2008, 18:810-820, and Zerbino DR and Birney E., Genome Res., 2008, 18:821-829. The optimization of sequence alignment has been described in the Art, for example, in International Patent Application Publication No. 2012 / 092426. Further descriptions of sequence alignment methods are described in detail, for example, in International Patent Application Publication No. 2020 / 236941, the entirety of which is incorporated herein by reference.
[0173] Misalignment (e.g., placement of base pairs from short reads in an inaccurate location within the genome), such as alternative allele reads being shifted from the histogram peak of alternative allele reads, can lead to decreased sensitivity in mutation detection due to sequence context around the actual cancer mutation (e.g., the presence of repetitive sequences). Other examples of sequence contexts that can cause misalignment include short tandem repeats, scattered repetitive sequences, low-complexity regions, insertion-deletion (indels), and paralogs. In cases where a problematic sequence situation arises when no actual mutation is present, misalignment may introduce artifact reads of the “mutant” allele by placing reads of the actual reference genome sequence in the wrong location. Since mutation calling algorithms for multiplex analysis must be sensitive even to low-abundance mutations, sequence misalignment can increase the false-positive detection rate and / or decrease specificity.
[0174] In some examples, the methods and systems disclosed herein may integrate the use of multiple individually tailored alignment methods or algorithms to optimize base call performance in sequencing methods, particularly those that rely on massively parallel sequencing (MPS) of numerous diverse genetic events at numerous diverse genomic loci. In some examples, the disclosed methods and systems may include the use of one or more global alignment algorithms. In some examples, the disclosed methods and systems may include the use of one or more local alignment algorithms.Examples of alignment algorithms that may be used include, but are not limited to, the Burrows-Wheeler Alignment (BWA) software bundle (see, e.g., Li, et al. (2009), “Fast and Accurate Short Read Alignment with Burrows-Wheeler Transform”, Bioinformatics 25:1754-60, Li, et al. (2010), “Fast and Accurate Long-Read Alignment with Burrows-Wheeler Transform”, Bioinformatics epub.PMID:20080505), the Smith-Waterman algorithm (see, e.g., Smith, et al. (1981), “Identification of Common Molecular Subsequences”, J. Molecular Biology 147(1):195-197), and the Striped Smith-Waterman algorithm (see, e.g., Farrar (2007), “Striped Smith-Waterman Speeds Database Searches Six Times Over Other Examples include the SIMD Implementations (see Bioinformatics 23(2):156-161), the Needleman-Wunsch algorithm (Needleman, et al. (1970) “A General Method Applicable to the Search for Similarities in the Amino Acid Sequence of Two Proteins”, J. Molecular Biology 48(3):443-53), or any combination thereof.
[0175] In some examples, the methods and systems disclosed herein may also include the use of sequence assembly algorithms, such as the Arachne sequence determination assembly algorithm (see, for example, Batzoglou, et al. (2002), “ARACHNE: A Whole-Genome Shotgun Assembler”, Genome Res. 12: 177-189).
[0176] In some cases, the alignment method used to analyze sequence reads is not individually customized or adjusted for the detection of different variants (e.g., point mutations, insertions, deletions, etc.) at different genomic loci. In some cases, a different alignment method is used to analyze reads that is individually customized or adjusted for the detection of at least a subset of different variants detected at different genomic loci. In some cases, a different alignment method is used to analyze reads that is individually customized or adjusted for the detection of each different variant at different genomic loci. In some cases, the adjustment may be a function of one or more of the following: (i) the locus being sequenced (e.g., a locus, microsatellite locus, or other target interval), (ii) the tumor type associated with the sample, (iii) the variant being sequenced, or (iv) the characteristics of the sample or target. The selection or use of alignment conditions individually adjusted for several specific target intervals being sequenced allows for optimization of speed, sensitivity, and specificity. This method is particularly effective when read alignment is optimized for a relatively large number of diverse target intervals.
[0177] In some examples, this method involves combining an alignment method optimized for reorganization with other alignment methods optimized for target intervals not associated with reorganization.
[0178] In some examples, the methods disclosed herein further include selecting or using an alignment method for analyzing, e.g., aligning, sequence reads, the alignment method being a function of, or selected accordingly, or optimized for, one or more of the following: (i) tumor type, e.g., tumor type in a sample; (ii) location of the segment to be sequenced (e.g., locus); (iii) type of variant within the segment to be sequenced (e.g., point mutation, insertion, deletion, substitution, copy number mutation (CNV), rearrangement, or fusion); (iv) site to be analyzed (e.g., nucleotide position); (v) type of sample (e.g., a sample as described herein); and / or (vi) adjacent sequences within or near the segment to be evaluated (e.g., according to its expected tendency toward misalignment of the segment due to the presence of repeat sequences within or near the segment).
[0179] In some cases, the methods disclosed herein enable the rapid and efficient alignment of cumbersome reads, e.g., reads having rearrangements. Thus, in some cases where the reads for the target interval include rearrangements, e.g., nucleotide positions with translocations, the methods may include using an alignment method that is appropriately modified and comprises: (i) selecting a rearrangement reference sequence for alignment with the read, such that the rearrangement reference sequence aligns with the rearrangement (in some cases, the reference sequence is not identical to the genomic rearrangement); and (ii) comparing, e.g., aligning the read with the rearrangement reference sequence.
[0180] In some cases, alternative methods may be used to align problematic reads. These methods are particularly effective when the alignment of reads is optimized for a relatively large number of diverse target intervals. For example, a method for analyzing a sample may involve (i) performing a read comparison (e.g., alignment comparison) using a first set of parameters (e.g., by using a first mapping algorithm or by comparison with a first reference sequence) to determine whether the read satisfies a first alignment criterion (e.g., the read can be aligned with the first reference sequence, e.g., with fewer than a certain number of mismatches), and (ii) if the read does not satisfy the first alignment criterion, performing a second alignment comparison using a second set of parameters (e.g., a second... (iii) optionally, determining whether the read satisfies a second criterion (e.g., whether the read can be aligned with the second reference sequence, e.g., with fewer than a certain number of mismatches), which may include determining whether the second set of parameters is more likely, for example, to result in alignment with the read for the variant (e.g., rearrangement, insertion, deletion, or translocation) compared to the first set of parameters.
[0181] In some cases, the alignment of sequence reads in the disclosed method may be combined with mutation calling methods described elsewhere in this Spec. As discussed herein, any decrease in sensitivity for detecting actual mutations may be addressed by evaluating (in a manual or automated manner) the quality of the alignment around the expected mutation site of the gene or genomic locus (e.g., a locus) being analyzed. In some cases, the sites to be evaluated may be obtained from a database of the human genome (e.g., the HG19 human reference genome) or cancer mutations (e.g., COSMIC). Regions identified as problematic can be repaired by alignment optimization (or realignment) using a slower but more accurate alignment algorithm, such as the Smith-Waterman alignment, using an algorithm selected to give better performance in the relevant sequence context. If general alignment algorithms fail to improve the situation, a customized alignment approach may be created, for example, by adjusting the largest different mismatch penalty parameters for genes likely to contain substitutions, adjusting specific mismatch penalty parameters based on specific mutation types common to a particular tumor type (e.g., CaT in melanoma), or adjusting specific mismatch penalty parameters based on specific mutation types common to a particular sample type (e.g., substitutions common to FFPE).
[0182] The decrease in specificity of the evaluated target interval (increase in false positive rate) due to misalignment can be assessed by manually or automatically checking all mutation calls in the sequencing data. Regions found to be prone to false mutation calls due to misalignment can be subjected to the alignment improvements discussed above. If algorithmic improvements are not possible, "mutations" from the problem region can be classified or screened from a panel of target loci.
[0183] Mutation Invocation A base call refers to the raw output of a sequencing device, for example, the determined sequence of nucleotides in an oligonucleotide molecule. A mutation call refers to the process of selecting a nucleotide value, for example, A, G, T, or C, for a given nucleotide position being sequenced. Typically, a sequence read (or base call) for a position will provide two or more values, for example, some reads will indicate T and some will indicate G. A mutation call is the process of assigning the correct nucleotide value, for example, one of those values, to the sequence. Although called a "mutation" call, it can be applied to assign a nucleotide value to any nucleotide position, for example, a position corresponding to a mutant allele, a wild-type allele, an allele not characterized as mutant or wild-type, or a position not characterized by variability.
[0184] In some examples, the disclosed methods may include the use of customized or tuned variant calling algorithms or their parameters to optimize performance when applied to sequencing data, particularly in methods that rely on massively parallel sequencing (MPS) of numerous diverse genetic events at numerous diverse genomic loci (e.g., loci, microsatellite regions, etc.) in a sample, e.g., a subject with cancer. Optimization of variant calling is described in the Art, for example, in International Patent Application Publication No. 2012 / 092426.
[0185] Methods for mutation calling may include one or more of the following: making independent calls based on information at each position in the reference sequence (e.g., examining sequence reads; examining base calls and quality scores; calculating the probabilities of observed bases and quality scores given potential genotypes; and assigning genotypes (e.g., using Bayes' rules)); removing false positives (e.g., using depth thresholds to reject SNPs with read depths much lower or higher than expected; local readjustment to remove false positives due to small indels); and refining calls by performing analysis based on linkage disequilibrium (LD) / complementation.
[0186] Formulas used to calculate genotype likelihoods associated with specific genotypes and locations are described, for example, in Li H. and Durbin R. Bioinformatics, 2010;26(5):589-95. Prior predictions for specific mutations in a particular cancer type can be used when evaluating samples from that cancer type. Such likelihoods can be obtained from publicly available cancer mutation databases, such as the Catalogue of Somatic Mutation in Cancer (COSMIC), HGMD (Human Gene Mutation Database), The SNP Consortium, Breast Cancer Mutation Data Base (BIC), and Breast Cancer Gene Database (BCGD).
[0187] Examples of LD / complementary analysis are described, for example, in Browning, BLand Yu, Z. Am. J. Hum. Genet. 2009, 85(6):847-61. Examples of low-coverage SNP calling methods are described, for example, in Li, Y., et al., Annu. Rev. Genomics Hum. Genet. 2009, 10:387-406.
[0188] After alignment, substitution detection can be performed using a calling method (e.g., a Bayesian mutation calling method), which is applied to each base in the target interval, e.g., to the exons of the gene or other loci being evaluated, and the presence of alternative alleles is observed. This method compares the probability of observing read data in the presence of a mutation with the probability of observing read data in the presence of base call errors only. If this comparison strongly supports the presence of a mutation, the mutation can be called.
[0189] The advantage of Bayesian mutation detection methods is that the comparison between the probability of mutation presence and the probability of base call errors alone can be weighted by prior predictions of the presence of the mutation at that site. If several reads of alternative alleles are observed at frequently mutated sites for a given cancer type, the presence of the mutation may be reliably called even if the amount of evidence for the mutation does not meet the usual threshold. This flexibility can then be used to increase the sensitivity of detecting rarer mutations / lower purity samples or to make the test more robust against reduced read coverage. The likelihood of a random base pair in the genome being mutated in cancer is approximately 1e-6. For example, the likelihood of specific mutations occurring at many sites in a typical polygenic cancer genome panel can be orders of magnitude higher. These likelihoods can be derived from publicly available databases of cancer mutations (e.g., COSMIC).
[0190] Indel calling is the process of finding bases in sequencing data that differ from a reference sequence due to insertions or deletions, typically including associated confidence scores or statistical evidence measures. Methods for indel calling may include steps of identifying candidate indels, calculating genotype likelihood by local realignment, and performing LD-based genotype inference and calling. Typically, Bayesian methods are used to obtain potential indel candidates, which are then tested against reference sequences within a Bayesian framework.
[0191] Algorithms for generating candidate indels are described, for example, in McKenna, A., et al., Genome Res. 2010; 20(9): 1297-303, Ye, K., et al., Bioinformatics, 2009; 25(21): 2865-71, Lunter, G., and Goodson, M., Genome Res. 2011; 21(6): 936-9, and Li, H., et al. (2009), Bioinformatics 25(16): 2078-9.
[0192] Methods for generating indel calls and individual-level genotype likelihoods include, for example, the Dindel algorithm (see Albers CA et al., Genome Res. 2011;21(6):961-73). For example, a Bayesian EM algorithm can be used to analyze reads, perform initial indel calls, generate genotype likelihoods for each candidate indel, and then, for example, perform QCALL (Le SQ and Durbin R. Genome Res. 2011;21(6):952-60). Parameters such as pre-observation predictions of indels can be adjusted based on the size or location of the indels (e.g., increased or decreased).
[0193] Methods have been developed to address limited deviations from 50% or 100% allele frequencies for the analysis of cancer DNA. (See, e.g., SNVMix - Bioinformatics. 2010 March 15;26(6):730-736.) However, the methods disclosed herein allow for consideration of frequencies (or allele fractions) in the range of 1% to 100% (i.e., allele fractions in the range of 0.01 to 1.0), and in particular, the possibility of the presence of mutant alleles at levels less than 50%. This approach is particularly important, for example, for detecting mutations in low-purity FFPE samples of natural (multiclonal) tumor DNA.
[0194] In some cases, the mutation calling method used to analyze sequence reads is not individually customized or tuned for the detection of different mutations at different genomic loci. In some cases, different mutation calling methods are used that are individually customized or tuned for at least a subset of the different mutations detected at different genomic loci. In some cases, different mutation calling methods are used that are individually customized or tuned for each different mutation detected at each different genomic locus. The customization or tuning may be based on one or more of the factors described herein, such as the type of cancer in the sample, the gene or locus in which the sequenced target interval is located, or the sequenced variant. This selection or use of mutation calling methods individually customized or tuned for the number of target intervals to be sequenced allows for optimization of the rate, sensitivity, and specificity of mutation calling.
[0195] In some examples, nucleotide values are assigned to each nucleotide position in X unique target intervals using a unique mutation calling method, where X is at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, and at least 5000 or more. The calling methods are different and can thus be unique, for example, by depending on different Bayesian prior values.
[0196] In some cases, assigning a nucleotide value is a function of the expected value or a value representing a variant at that nucleotide position in a type of tumor, e.g., a read showing a mutation, prior to observation (e.g., in the literature).
[0197] In some examples, the method involves assigning nucleotide values (e.g., mutation calls) to at least 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 nucleotide positions, each assignment being a function of a unique value (in contrast to the values of other assignments) that represents or is either the expected value from prior observations (e.g., literature) of reads showing variants, e.g., mutations, at that nucleotide position in a tumor of a particular type.
[0198] In some examples, assigning a nucleotide value is a function of a set of values that represent the probability of observing a read exhibiting the variant at a particular nucleotide position if the variant is present in the sample at a specific frequency (e.g., 1%, 5%, 10%) and / or if the variant is not present (e.g., observed in the read solely due to a base calling error).
[0199] In some examples, the variant calling method described herein may include (a) obtaining, for each of the X target intervals, (i) a first value that is or represents the expected value prior to (e.g., literature) of observing a variant at that nucleotide position in a tumor of type X, e.g., a read exhibiting the variant; and (ii) a second set of values that represent the likelihood of observing a read exhibiting the variant at that nucleotide position if the variant is present in the sample at a certain frequency (e.g., 1%, 5%, 10%) and / or if the variant is not present (e.g., observed in the read due to a base calling error alone); and (b) assigning a nucleotide value from the read (e.g., calling a variant) to each of the nucleotide positions by weighting the comparison between the values in the second set using the first value, for example by a Bayesian method as described herein, in response to the values, and thereby analyzing the sample.
[0200] Additional descriptions of exemplary nucleic acid sequencing methods, mutation calling methods, and methods for the analysis of genetic variants are provided, for example, in U.S. Patent Nos. 9,340,830, 9,792,403, 11,136,619, 11,118,213, and International Patent Application Publication No. 2020 / 236941, the entire contents of each of which are incorporated herein by reference.
[0201] system Also disclosed herein are systems designed to implement any of the disclosed methods for identifying HLA variants in a sample from a subject. The system may comprise, for example, one or more processors and a memory unit communicatively coupled to the one or more processors and configured to store instructions, wherein when an instruction is executed by one or more processors, the system will: i) receive sequence read data for a plurality of off-target sequence reads; ii) align the sequence read data to a reference genome to identify a plurality of HLA alignment reads; and iii) process the plurality of HLA alignment reads using a variant caller to identify HLA variants.
[0202] In some examples, off-target sequence read data can be identified by aligning sequence reads from targeted sequencing methods to a non-mitochondrial reference genome. In some examples, off-target sequence reads may include sequence reads that align to the mitochondrial genome and sequence reads that do not align to the non-mitochondrial reference genome. In any of the examples herein, the variant caller may include a short variant caller. In any of the examples herein, the system may further include instructions, which, when executed by one or more processors, cause the system to determine allele frequencies for HLA variants based on multiple HLA alignment reads, and to categorize the HLA variants as germ cells, somatic cells, or subclonal or artifacts based on the determined allele frequencies. In some examples, categorization does not require the use of sequence read data from matched normal controls.
[0203] In some examples, the disclosed systems may further include a sequencer, for example, a next-generation sequencer (also known as a massively parallel sequencer). Examples of next-generation (or massively parallel) sequencing platforms include, but are not limited to, Roche / 454's Genome Sequencer (GS) FLX system, Illumina / Solexa's Genome Analyzer (GA), Illumina's HiSeq® 2500, HiSeq® 3000, HiSeq® 4000 and NovaSeq® 6000 sequencing systems, Life / APG's Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, Helicos BioSciences' HeliScope Gene Sequencing system, ThermoFisher Scientific's Ion Torrent Genexus system, or Pacific Biosciences' PacBio® RS system.
[0204] In some examples, the disclosed system may be used to identify HLA variants in any of the various samples described herein (e.g., tissue samples, biopsy samples, blood samples, or liquid biopsy samples derived from a subject).
[0205] In some cases, sequencing data may be processed to identify HLA variants at at least 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 chrM loci.
[0206] In some cases, nucleic acid sequence data are obtained using next-generation sequencing technologies (also known as ultra-parallel sequencing technologies) with read lengths of less than 400, 300, 200, 150, 100, 90, 80, 70, 60, 50, 40, or 30 bases.
[0207] In some cases, HLA variant identification is used to select, initiate, adjust, or terminate treatment for cancer in the subject from which the sample originates (e.g., a patient), as described elsewhere in this specification.
[0208] In some cases, the disclosed system may further include a sample processing and library preparation workstation, a microplate handling robot, a fluid dispensing system, a temperature control module, an environmental control chamber, an additional data storage module, a data communication module (e.g., Bluetooth®, WiFi, intranet, or internet communication hardware and associated software), a display module, one or more local and / or cloud-based software packages (e.g., an instrument / system control software package, a sequencing data analysis software package), or any combination thereof. In some cases, the system may include, or be part of, a computer system or computer network as described elsewhere in this Spec.
[0209] Machine Learning Any of the various machine learning techniques and algorithms (including trained machine learning algorithms, as referred to herein) may be used in the implementation of the disclosed method. For example, a machine learning model may include supervised learning models (i.e., models trained using a set of labeled training data), unsupervised learning models (i.e., models trained using a set of unlabeled training data), semi-supervised learning models (i.e., models trained using a combination of labeled and unlabeled training data), self-supervised learning models, or any combination thereof. In some examples, a machine learning model may include a deep learning model (i.e., a model comprising many layers of coupled “nodes” which may be trained in a supervised, unsupervised, or semi-supervised manner).
[0210] In some examples, the disclosed method may be implemented using one or more machine learning models (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 machine learning models), or a combination thereof.
[0211] In some examples, one or more machine learning models may include statistical methods for analyzing data. Machine learning models may be used for data classification and / or regression. Examples of machine learning models include neural networks, support vector machines, decision trees, ensemble learning (e.g., bagging-based learning, e.g., random forests, and / or boosting-based learning), k-nearest neighbor algorithms, linear regression-based models, and / or logistic regression-based models. Machine learning models may include regularization, e.g., L1 regularization and / or L2 regularization. Machine learning models may include the use of dimensionality reduction techniques (e.g., principal component analysis, matrix decomposition techniques, and / or autoencoders) and / or clustering techniques (e.g., hierarchical clustering, k-means clustering, distribution-based clustering, e.g., Gaussian mixture models, or density-based clustering, e.g., DBSCAN or OPTICS). One or more machine learning models may include repeatedly solving an objective function based on a training dataset (e.g., optimizing). Even if a machine learning model includes models for which closed-form solutions exist (for example, linear regression), it may still be possible to use an iterative solving method.
[0212] In some examples, machine learning models may include artificial neural networks (ANNs), such as deep learning models. For example, one or more machine learning models / algorithms used to implement the disclosed methods may include an ANN, which may include, but is not limited to, any of the various computational motifs / architectures known to those skilled in the art, including feedforward connections (e.g., skip connections), recurrent connections, fully connected layers, convolutional layers, and / or pooling functions (e.g., attention including self-attention). The artificial neural network may include differentiable nonlinear functions trained by backpropagation.
[0213] Artificial neural networks, such as deep learning models, generally include interconnected nodes organized into layers of multiple nodes. For example, an ANN architecture may include at least an input layer, one or more hidden layers (i.e., hidden layers), and an output layer. An ANN or deep learning model may also include any total number of layers (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more than 20 layers in total) and any number of hidden layers (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or more than 20 hidden layers), where the hidden layers function as trainable feature extractors that enable mapping a set of input data to a preferred output value or set of output values. Each layer of a neural network contains multiple nodes (e.g., at least 10, 25, 50, 75, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, or more than 10,000 nodes). A node receives input data (e.g., genomic feature data (e.g., variant sequence data, methylation status data, etc.), non-genomic feature data (e.g., digital pathology image feature data), or other types of input data (e.g., patient-specific clinical data)) from one or more input data nodes, or directly from the output of one or more nodes in previous layers, and performs a specific operation (e.g., summation operation). In some examples, the connection from the input to the node is associated with a weight (or weight coefficient). In some examples, the node receives, for example, all pairs of X of the input. i and their associated weights W iThe products of these may be summed. In some examples, the weighted sum is offset by a bias b. In some examples, the node output may be gated using a threshold or activation function f, where f may be a linear or nonlinear function. The activation function may be other functions such as a rectified linear unit (ReLU) activation function, or a saturated hyperbolic tangent, identity, binary step, logistic, arcTan, soft sine, parametric rectified linear unit, exponential linear unit, soft plus, Bent identity, soft exponential, sine, Gaussian, or sigmoid function, or any combination thereof.
[0214] The weighting functions, bias values, and thresholds, or other computational parameters of a neural network (or other machine learning architecture) may be “taught” or “learned” during the training phase using one or more sets of training data (e.g., 1, 2, 3, 4, 5, or more than 5 sets of training data) and a particular training method configured to solve (e.g., minimize) a loss function. For example, tunable parameters for an ANN (e.g., a deep learning model) may be determined based on input data from a training dataset using an iterative solver (e.g., a gradient-based method, e.g., backpropagation), so that the output values computed by the ANN (e.g., sample classification or disease outcome prediction) are consistent with examples contained in the training dataset. Training the model (i.e., determining the tunable parameters of the model using an iterative solver) may or may not be performed using the same hardware used to deploy the trained model.
[0215] In some examples, the disclosed methods may include retraining one of the machine learning models (e.g., repeatedly retraining a previously trained model using one or more training datasets different from those used to initially train the model). In some examples, retraining a machine learning model may include using a continuous (e.g., online) machine learning model; that is, the model is periodically or continuously updated or retrained based on new training data. The new training data may be provided, for example, by a single deployed local operational system, multiple deployed local operational systems, or multiple deployed geographically distributed operational systems. In some examples, the disclosed methods may use, for example, a pre-trained ANN, which may be fine-tuned according to additional datasets input to the pre-trained ANN.
[0216] Computer systems and networks Figure 3 illustrates an example of a computing device or system according to one embodiment. Device 300 may be a host computer connected to a network. Device 300 may be a client computer or a server. As shown in Figure 3, device 300 may be any suitable type of microprocessor-based device, such as a personal computer, workstation, server, or handheld computing device (portable electronic device, e.g., telephone or tablet). The device may include, for example, one or more processors 310, an input device 320, an output device 330, a memory or storage device 340, a communication device 360, and a nucleic acid sequencer 370. Software 350 residing in the memory or storage device 340 may include, for example, an operating system and software for carrying out the methods described herein. The input device 320 and the output device 330 may generally correspond to those described herein, may be connectable to a computer, or may be integrated with a computer.
[0217] The input device 320 may be any suitable device that provides input, such as a touchscreen, keyboard or keypad, mouse, or voice recognition device. The output device 330 may be any suitable device that provides output, such as a touchscreen, haptic device, or speaker.
[0218] Storage 340 may be any suitable device that provides storage (e.g., electrical, magnetic, or optical memory, including RAM (volatile and non-volatile), cache, hard drive, or removable storage disk). Communication device 360 may include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. Computer components may be connected in any suitable manner, for example, via wired media (e.g., physical system bus 380, Ethernet connection, or any other wired transfer technology) or wirelessly (e.g., Bluetooth®, Wi-Fi®, or any other wireless technology).
[0219] The software module 350 is stored in the storage 340 as executable instructions and can be executed by the processor 310, and may include, for example, a process that embodies the functions of an operating system and / or a method of the disclosure (for example, embodied in the above device).
[0220] The software module 350 may also be stored and / or transferred to any non-temporary computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device (e.g., those described herein), and may fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, the computer-readable storage medium may be any medium, such as storage 340, and may contain or store processes for use by or in connection with an instruction execution system, apparatus, or device. Examples of computer-readable storage media include hard drives, flash drives, and memory units such as distribution modules, which operate as single functional units. Furthermore, the various processes described herein may be embodied as modules configured to operate according to the embodiments and techniques described above. In addition, although processes may be shown and / or described separately, those skilled in the art will understand that the above processes may be routines or modules within other processes.
[0221] The software module 350 may also be propagated by an instruction execution system, apparatus, or device such as those described above, or in any transmission medium for use in connection with them, and may fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, the transmission medium may be any medium that can communicate, propagate, or transmit transmission programming for use by or in connection with the instruction execution system, apparatus, or device. The transmission-readable medium may include, but is not limited to, wired or wireless transmission media of electronic, magnetic, optical, electromagnetic, or infrared.
[0222] Device 300 may be connected to a network (e.g., network 404, shown in Figure 4 and / or described below), which may be any preferred type of interconnected communication system. The network may implement any preferred communication protocol and may be protected by any preferred security protocol. The network may include network links in any preferred configuration that can implement the transmission and reception of network signals, such as wireless network connections (T1 or T3 lines), cable networks, DSL, or telephone lines.
[0223] Device 300 may be implemented using any operating system, for example, an operating system suitable for running over a network. Software module 350 can be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying the functionality of this disclosure may be deployed in different configurations (e.g., in a client / server deployment, or via a web browser as a web-based application or web service). In some embodiments, the operating system is run by one or more processors, for example, processor 310.
[0224] Device 300 may further include a sequencer 370, which can be any suitable nucleic acid sequencing instrument.
[0225] Figure 4 illustrates an example of a computing system according to one embodiment. In system 400, device 300 (e.g., as described above and illustrated in Figure 3) is connected to network 404, which is also connected to device 406. In some embodiments, device 406 is a sequencer. Exemplary sequencing devices may include, but are not limited to, Roche / 454's Genome Sequencer (GS) FLX System, Illumina / Solexa's Genome Analyzer (GA), Illumina's HiSeq® 2500, HiSeq® 3000, HiSeq® 4000, and NovaSeq® 6000 sequencing systems, Life / APG's Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, Helicos BioSciences' HeliScope Gene sequencing system, or Pacific Biosciences' PacBio® RS system.
[0226] Devices 300 and 406 can communicate using an appropriate communication interface over a network 404, such as a local area network (LAN), a virtual private network (VPN), or the internet. In some embodiments, network 404 can be, for example, the internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. Devices 300 and 406 can communicate partially or entirely over wireless or wired communication, such as Ethernet or IEEE 802.11b wireless. Additionally, devices 300 and 406 can communicate over a second network, such as a mobile / cellular network, using a suitable communication interface. Communication between devices 300 and 406 may further include, or communicate with, various servers, such as mail servers, mobile servers, media servers, and telephone servers. In some embodiments, devices 400 and 406 can communicate directly (instead of, or in addition to, communication over network 404) over wireless or wired communication, such as Ethernet or IEEE 802.11b wireless. In some embodiments, devices 300 and 406 communicate either through a direct connection or through communication 408 that can occur over a network (e.g., network 404).
[0227] One or all of devices 300 and 406 are generally programmed to include logic (e.g., HTTP web server logic) accessed from local or remote databases or other sources of data and content, or to format data, in order to provide and / or receive information over network 404 in accordance with the various examples described herein.
[0228] Figure 5 shows a non-restrictive, exemplary flowchart illustrating an implemented method for calling short variants in HLA genes (e.g., HLA-I or HLA-II genes). In short, a BAM file from next-generation sequencing is provided, as shown in process 502, and all reads that align to the HLA region and reads that did not map to the human reference genome are extracted, as shown in process 504. Then, as shown in process 506, and as described in Szolek et al., Bioinformatics, 2014, 30(23):3310-3316, both of these reads are genotyped using Optitype software. Then, as shown in process 508, a target-specific / patient-specific HLA reference sequence is generated using an HLA reference sequence available from the IPD-IMGT / HLA database. Then, as shown in process 510, the HLA sequencing reads are realigned to the patient-specific HLA reference to produce a new patient-specific HLA BAM file. Next, as shown in process 512, the BAM file is fed into a series of scripts to invoke short variants. The variants are output to an HLA VCF file. The VCF file provides information about nucleotide mismatches. To translate the mismatches into coding strand (CDS) effects, a local alignment is performed of an HLA nucleotide reference sequence containing both exons and introns of the gene against an HLA coding reference sequence containing only exons. Both HLA nucleotide reference sequences containing both exons and introns, and HLA coding reference sequences containing only exons, are available from the IPD-IMGT / HLA database. The alignment separates the nucleotide sequence into distinct exons and introns, sorts the nucleotide positions, and can be used to number each exon for the target HLA sequence derived from the subject. Analysis of exons and introns also indicates whether the identified nucleotide mismatches occur in non-coding regions that generally do not have biological effects on protein function.If a mismatch occurs in the coding region, the trinucleotide content can be determined by dividing the CDS position by 3. The reference and alternative sequence trinucleotides can be determined using a standard DNA codon table for constructing protein effects.
[0229] Exemplary Embodiments Exemplary embodiments of the methods and systems described herein include the following: 1. A method, To provide multiple nucleic acid molecules obtained from samples from the target, Ligating one or more adapters onto one or more nucleic acid molecules from multiple nucleic acid molecules, Amplifying one or more ligated nucleic acid molecules from multiple nucleic acid molecules, The process involves capturing amplified nucleic acid molecules from amplified nucleic acid molecules, The sequencer is used to sequence the captured nucleic acid molecule and obtain multiple sequence reads representing the captured nucleic acid molecule. In one or more processors, receiving array read data for multiple array reads, Using one or more processors, sequence read data is aligned to a first reference genome to identify a first set of multiple HLA alignment reads and unmapped sequence reads. Using one or more processors, a database of HLA polymorphisms is searched with a first set of HLA alignment reads and unmapped sequence reads to generate a target-specific reference genome. Using one or more processors, align a first set of multiple HLA alignment reads and unmapped sequence reads to a target-specific reference genome to identify a second set of multiple HLA alignment reads, A method comprising processing a second set of HLA alignment reads using one or more processors to identify HLA variants. 2. The method according to Clause 1, wherein a variant caller is used to process a second set of HLA alignment leads and identify HLA variants. 3. The method of Clause 1 or Clause 2, wherein HLA variant identification does not require the use of sequence read data from a matched normal control. 4. The method described in any one of the clauses 1 to 3, wherein the HLA polymorphism database is the IPD-IMGT / HLA database. 5. The method described in any one of the clauses 1 to 4, wherein the identified HLA variant is an HLA-I variant. 6. The method according to Clause 5, wherein the HLA variant is a variant of HLA-A, HLA-B, and / or HLA-C. 7. The method according to any one of the clauses 1 to 6, wherein the identified HLA variant is an HLA-I and / or HLA-II variant. 8. The method described in any one of the paragraphs 1 to 7, wherein the subject is suspected of having cancer or is determined to have cancer. 9. Cancers include B-cell carcinoma (multiple myeloma), melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine cancer, endometrial cancer, oral cancer, pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine cancer, appendiceal cancer, salivary gland cancer, thyroid cancer, adrenal cancer, osteosarcoma, chondrosarcoma, and hematological cancers. Adenocarcinoma, inflammatory myofibroblastoma, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteosarcoma, spinal Cordoma, angiosarcoma, endosarcoma, lymphangiosarcoma, lymphangioendosarcoma, synoviomas, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, liver cancer, cholangiocarcinoma, choriocarcinoma, seminoma, embryonic carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pineal cell carcinoma The method according to Clause 8, wherein the tumor is a cystoma, hemangioblastoma, acoustic neuroblastoma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, primary myelofibrosis, eosinophilic syndrome, systemic mastocytosis, familial eosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, or carcinoid tumor. 10. Cancers include acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpression / amplification), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deficiency), chronic myeloid leukemia, chronic myeloid leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, Colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild-type), cryopyrin-associated periodic fever syndrome, cutaneous T-cell lymphoma, dermatofibrosarcoma protuberans, diffuse large B-cell lymphoma, fallopian tube cancer, follicular B-cell non-Hodgkin lymphoma, follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, gastrointestinal stromal tumor, gastrointestinal stromal tumor (KIT+), giant cell tumor of bone, glioblastoma, granulomatosis with polyangiitis, head and neck squamous cell carcinoma, hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, systemic lupus erythematosus, mantle cell lymphoma, medullary thyroid carcinoma, melanoma, BRAF Melanoma with V600 mutation, melanoma with BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman disease, multiple hematological malignancies including Philadelphia chromosome-positive ALL and CML, multiple myeloma, myelofibrosis, non-Hodgkin lymphoma, unresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, non-small cell lung cancer, non-small cell lung cancer (ALK+), non-small cell lung cancer (PD-L1+), non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), non-small cell lung cancer (with BRAF V600E mutation), non-small cell lung cancer (with EGFR exon 19 deletion or exon 21 substitution (L858R) mutation), non-small cell lung cancer (EGFRThe method according to Clause 8, including ovarian cancer (with T790M mutation), ovarian cancer (with BRCA mutation), pancreatic cancer, neuroendocrine tumors of pancreatic, gastrointestinal, or lung origin, pediatric neuroblastoma, peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, renal cell carcinoma, rheumatoid arthritis, small lymphocytic lymphoma, soft tissue sarcoma, solid tumor (MSI-H / dMMR), squamous cell carcinoma of the head and neck, squamous non-small cell lung cancer, thyroid cancer, thyroid carcinoma, urothelial carcinoma, or primary gammaglobulinemia. 11. The method according to Clause 10, further comprising treating the subject with anticancer therapy. 12. The method described in Clause 11, wherein the treatment of the subject is performed after the subject has been determined to have cancer. 13. The anticancer therapy is the method described in Clause 11, including targeted anticancer therapy. 14. Targeted anticancer therapies include abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), ado-trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), and anastrozole (Ari). (midex), apalutamide (Erleada), aciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicaptagensilolucel (Yescarta), axitinib (Inlyta), verantamab mahodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), verzutifan (Welireg), bevacizumab (Avast (in), bexarotene (Targretin), vinimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexcabutadiene autolucel (Tecartus), brigutinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinu Mab (Ilaris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), semiprimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (Darzalex)Faspro, darolutamide (Nubeqa), dasatinib (Sprycel), deniroikin difutitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostallumab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enacidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebi) c) Fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glassedegib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idekabuta genbiculucel (Abecma), idelaritinib (Zydelig), imatinib mesylate (Gleevec), infiglatinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguan I131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa), ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (Somatuline)Depot), lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), lenvatinib mesylate (Lenvima), letrozole (Femara), lysocabate gemmaralucel (Breyanzi), loncustuximab tesillin-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu177-doteate (Lutathera), margetuximab-cmkb (Margenza), mid Staurin (Rydapt), Mobocertinib succinate (Exkivity), Mogamulizumab-kpkc (Poteligeo), Moxetumomab Pasdotox-tdfk (Lumoxiti), Naxitamab-gqgk (Danyelza), Necitumumab (Portrazza), Neratinib maleate (Nerlynx), Nilotinib (Tasigna), Niraparib tosylate monohydrate (Zejula), Nivolumab (Opdivo), Obinut Zumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), olalatumumab (Lartruvo), osimertinib (Tagrisso), palbociclib (Ibrance), panitumumab (Vectibix), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pexidaritinib hydrochloride (Turalio), po Latuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), prarlcetinib (Gavreto), radium-223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), lipretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase (Rituxan)Hycela), Romidepsin (Istodax), Rucaparib cansylate (Rubraca), Ruxolitinib phosphate (Jakafi), Sacituzumab Govitecan-hziy (Trodelvy), Cericlib, Selinexol (Xpovio), Serpercatinib (Retevmo), Selumetinib sulfate (Koselugo), Siltuximab (Sylvant), Sirolimus-binding protein particles (Fyarro), Sonidedib (O domzo), sorafenib (Nexavar), sotracib (Lumakras), sunitinib (Sutent), tafacitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), thalazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), teventafusp-tebn (Kimmtrak), temsirolimus (Tori sel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tocitumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), toremifene (Fare The method according to Clause 13, including ston), tucatinib (Tukysa), umbralicib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), bismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), ziv-aflibercept (Zaltrap), or any combination thereof. 15. The method described in any one of the clauses 1 to 14, further comprising obtaining a sample from the subject. 16. The method described in any one of the clauses 1 to 15, wherein the sample includes a tissue biopsy sample or a liquid biopsy sample. 17. The method according to Clause 16, wherein the sample is a liquid biopsy sample and includes blood, plasma, cerebrospinal fluid, sputum, feces, urine, or saliva. 18. The method according to Clause 17, wherein the sample is a liquid biopsy sample and contains circulating tumor cells (CTCs). 19. The method according to Clause 16, wherein the sample is a liquid biopsy sample and includes cell-free DNA (cfDNA) and / or circulating tumor DNA (ctDNA). 20. The method according to any one of the claims 1 to 19, wherein the plurality of nucleic acid molecules include a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules. 21. The method according to Clause 20, wherein tumor nucleic acid molecules are derived from the tumor portion of the heterogeneous tissue biopsy sample, and non-tumor nucleic acid molecules are derived from the normal portion of the heterogeneous tissue biopsy sample. 22. The method according to Clause 20, wherein the sample comprises a liquid biopsy sample, the tumor nucleic acid molecules are derived from the circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and the non-tumor nucleic acid molecules are derived from the non-tumor cell-free DNA (cfDNA) fraction of the liquid biopsy sample. 23. The method according to any one of the claims 1 to 22, wherein one or more adapters include an amplification primer, a flow cell adapter sequence, a substrate adapter sequence, or a sample index sequence. 24. The method according to any one of the claims 1 to 23, wherein the captured nucleic acid molecule is captured from a nucleic acid molecule amplified by hybridization to one or more bait molecules. 25. The method according to Clause 24, wherein one or more bait molecules comprise one or more nucleic acid molecules, and each nucleic acid molecule comprises a region complementary to the region of the captured nucleic acid molecule. 26. The method according to any one of the provisions 1 to 25, wherein the amplification of a nucleic acid molecule involves performing polymerase chain reaction (PCR) amplification techniques, non-PCR amplification techniques, or isothermal amplification techniques. 27. The method described in any one of Clauses 1 to 26, wherein the sequencing includes the use of massively parallel sequencing (MPS) technology, whole-genome sequencing (WGS), whole-exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing technology. 28. Sequencing includes massively parallel sequencing, and massively parallel sequencing techniques include next-generation sequencing (NGS), as described in Clause 26. 29. The method described in any one of the clauses 1 to 28, wherein the sequencer includes a next-generation sequencer. 30. The method according to any one of the claims 1 to 29, wherein one or more of the sequencing reads overlap with one or more gene loci in one or more subgenome segments in the sample. 31.1 or more gene loci, 10-20 gene loci, 10-40 gene loci, 10-60 gene loci, 10-80 gene loci, 10-100 gene loci, 10-150 gene loci, 10-200 gene loci, 10-250 gene loci, 10-300 gene loci, 10-350 gene loci, 10-400 gene loci, 10-450 gene loci, 10-500 gene loci, 20-40 gene loci, 20-60 gene loci, 20-80 gene loci, 20-100 gene loci, 20-150 gene loci, 20-200 gene loci, 20-250 gene loci, 20 ~300 gene loci, 20~350 gene loci, 20~400 gene loci, 20~500 gene loci, 40~60 gene loci, 40~80 gene loci, 40~100 gene loci, 40~150 gene loci, 40~200 gene loci, 40~250 gene loci, 40~300 gene loci, 40~350 gene loci, 40~400 gene loci, 40~500 gene loci, 60~80 gene loci, 60~100 gene loci, 60~150 gene loci, 60~200 gene loci, 60~250 gene loci, 60~300 gene loci, 60~350 gene loci, 60~ 400 loci, 60-500 loci, 80-100 loci, 80-150 loci, 80-200 loci, 80-250 loci, 80-300 loci, 80-350 loci, 80-400 loci, 80-500 loci, 100-150 loci, 100-200 loci, 100-250 loci, 100-300 loci, 100-350 loci, 100-400 loci, 100-500 loci, 150-200 loci, 150-250 loci, 150-300 loci, 15 The method according to Clause 30, comprising 0-350 loci, 150-400 loci, 150-500 loci, 200-250 loci, 200-300 loci, 200-350 loci, 200-400 loci, 200-500 loci, 250-300 loci, 250-350 loci, 250-400 loci, 250-500 loci, 300-350 loci, 300-400 loci, 300-500 loci, 350-400 loci, 350-500 loci, or 400-500 loci. ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, and APC AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCR, BCORL1, BCR, BRAF, BRCA1 BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND 1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH 1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CU L3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGF R, EMSY(C11orf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERC C4, ERG, ERRFI1, ESR1, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, F ANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3 FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1 GABRA6, GATA3, GATA4, GATA6, GID4(C17orf39), GNA11, GNA13, GNAQ, GNAS GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF 1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, K DM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A(MLL), KMT2D(MLL2), KRAS.LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL , MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2 , NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFR A, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, P TCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, REL, RET, RICTOR, RN F43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1 The method described in Clause 30 or Clause 31, including, , SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, or any combination thereof. 33. One or more gene loci are present in ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, I The method according to Clause 30 or Clause 31, including L-1β, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRα, PDGFRβ, PD-L1, PI3Kδ, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, VEGFB, or any combination thereof. 34. The method according to any one of the claims 30 to 33, wherein one or more gene loci further include an HLA class I or HLA class II gene. 35. The method described in any one of the clauses 1 to 33, further comprising generating a report indicating the presence or absence of a genetic variant by one or more processors. 36. The method described in Clause 35, further including sending a report to a healthcare provider. 37. The method described in Clause 36, wherein the report is transmitted via a computer network or peer-to-peer connection. 38. A method for identifying HLA variants, a) One or more processors receive array read data for multiple array reads, b) Aligning sequence read data to a first reference genome using one or more processors to identify a first set of HLA alignment reads and unmapped sequence reads, c) Using one or more processors, a database of HLA polymorphisms is searched with a first set of HLA alignment reads and unmapped sequence reads to generate a target-specific reference genome. d) Using one or more processors, align the first set of multiple HLA alignment reads and unmapped sequence reads to a target-specific reference genome to identify a second set of multiple HLA alignment reads, e) A method comprising processing a second set of HLA alignment reads using one or more processors to identify HLA variants. 39. The method according to Clause 38, wherein multiple sequence reads are obtained using a targeted sequencing method that includes targeting an HLA locus. 40. The method according to Clause 38, wherein a variant caller is used to process a second set of HLA alignment leads and identify HLA variants. 41. The method of Clause 39 or Clause 40, wherein HLA variant identification does not require the use of sequence read data from a matched normal control. 42. The method described in any one of the clauses 39 to 41, wherein the HLA polymorphism database is the IPD-IMGT / HLA database. 43. The method described in any one of the clauses 39 to 42, wherein the identified HLA variant is an HLA-I variant. 44. The method according to Clause 43, wherein the HLA-I variant is a variant of HLA-A, HLA-B, and / or HLA-C. 45. The method according to any one of the clauses 38 to 44, wherein the identified HLA variant is an HLA-I and / or HLA-II variant. 46. A method for identifying HLA variants, a) One or more processors receive array read data for multiple array reads, b) Aligning sequence read data to a first reference genome using one or more processors to identify a first set of HLA alignment reads and unmapped sequence reads, c) Using one or more processors, align the first set of HLA alignment reads and unmapped sequence reads to a second reference genome to identify the second set of HLA alignment reads, d) A method comprising processing a second set of HLA alignment reads using one or more processors to identify HLA variants, wherein the identification does not require the use of sequence read data from matched normal controls. 47. The method according to Clause 46, wherein a variant caller is used to process a second set of HLA alignment leads and identify HLA variants. 48. The method according to Clause 46 or Clause 47, wherein the second reference genome is a target-specific reference genome. 49. The method according to any one of the clauses 46 to 48, wherein a target-specific reference is generated by searching an HLA polymorphism database using multiple alignment reads and unmapped sequence reads of a first set of HLAs. 50. The method described in Clause 49, wherein the HLA polymorphism database is the IPD-IMGT / HLA database. 51. The method described in any one of the clauses 46-50, wherein the HLA variant is an HLA-I variant. 52. The method according to Clause 51, wherein the identified HLA-I variant is a variant of HLA-A, HLA-B, and / or HLA-C. 53. The method according to any one of the clauses 46 to 52, wherein the HLA variant is an HLA-I and / or HLA-II variant. 54. Using one or more processors, align HLA variants to an HLA code reference array and annotate the exon and intron boundaries in the HLA variants, Using one or more processors, the annotated HLA variant is translated into an HLA polypeptide sequence, Using one or more processors, the translated HLA polypeptide sequence is aligned to a reference HLA polypeptide sequence to identify amino acid changes in the translated HLA polypeptide sequence, The method according to any one of the clauses 38 to 53, further comprising using one or more processors to categorize translated HLA polypeptides as containing synonymous mutations, non-start mutations, non-stop mutations, nonsense mutations, or missense mutations. 55. The method described in any one of the clauses 38 to 54, wherein the HLA variant is an HLA-I variant. 56. The method according to Clause 51, wherein the identified HLA-I variant is a variant of HLA-A, HLA-B, and / or HLA-C. 57. The method according to any one of the clauses 38 to 56, wherein the HLA variant is an HLA-I and / or HLA-II variant. 58. The method described in any one of the clauses 38 to 57, which includes the use of a software tool that predicts the probability that a search of a polymorphic database corresponds to a known HLA sequence. 59. The software tool is Optitype as described in Clause 58. 60. The method described in any one of the clauses 38 to 59, wherein the variant caller includes a short variant caller. 61. The method described in any one of the clauses 38 to 60, wherein the subject is a human being. 62. A system, a) One or more processors, b) A memory configured to be communicatively connected to one or more processors and to store instructions, wherein when an instruction is executed by one or more processors, the system i) Receive sequence read data for multiple sequence reads, ii) Align the sequence read data to the first reference genome to identify the first set of HLA alignment reads and unmapped sequence reads. iii) Search the HLA polymorphism database with the first set of HLA alignment reads and unmapped sequence reads to generate a target-specific reference genome. iv) Align the first set of multiple HLA alignment reads and unmapped sequence reads to a target-specific reference genome to identify the second set of multiple HLA alignment reads. v) A system comprising memory for processing a second set of HLA alignment reads to identify HLA variants. 63. The system described in Clause 62, which uses a variant caller to process a second set of HLA alignment leads and identify HLA variants. 64. A system as described in Clause 62 or 63 in which HLA variant identification does not require the use of sequence read data from a matched normal control. 65. A system described in any one of clauses 62-64, wherein the HLA polymorphism database is an IPD-IMGT / HLA database. 66. Further instructions, and when an instruction is executed by one or more processors, the system c) Align the HLA variant with the HLA code reference sequence and annotate the exon and intron boundaries in the HLA variant. d) Translate the annotated HLA variant into an HLA polypeptide sequence. e) Align the translated HLA polypeptide sequence with a reference HLA polypeptide sequence to identify amino acid changes in the translated HLA polypeptide sequence. f) A system described in any one of clauses 62-65 that categorizes translated HLA polypeptides as synonymous, nonstart, nonstop, nonsense, or missense. 67. A system described in any one of the clauses 62 to 66, wherein the HLA variant is the HLA-I variant. 68. The system described in Clause 67, wherein the identified HLA-I variant is a variant of HLA-A, HLA-B, and / or HLA-C. 69. A system as described in any one of the clauses 62 to 68, wherein the HLA variant is an HLA-I and / or HLA-II variant. 70. A system described in any one of clauses 62 to 69, which includes the use of a software tool that predicts the probability that an input sequence corresponds to a known HLA sequence when searching a polymorphic database. 71. The software tool is Optitype, as described in Clause 70. 72. A system described in any one of the clauses 62 to 71, in which the variant caller includes a short variant caller. 73. A method for diagnosing a disease, the method comprising diagnosing that a subject has a disease based on the identification of an HLA variant in a sample from the subject, wherein the HLA variant is identified in accordance with the method described in any one of the clauses 1 to 61. 74. A method for selecting an anticancer therapy, the method comprising selecting an anticancer therapy for a subject in response to the identification of an HLA variant for a sample from a subject, wherein the HLA variant is identified in accordance with the method described in any one of the clauses 1 to 61. 75. A method for treating cancer in a subject, comprising administering an effective dose of anticancer therapy to the subject in response to the identification of an HLA variant in a sample from the subject, wherein the HLA variant is identified in accordance with the method described in any one of the clauses 1 to 61. 76. A method for monitoring the progression or recurrence of cancer in a subject, wherein this method is Identifying a first HLA variant in a first sample obtained from the subject at a first point in time, in accordance with the method described in any one of clauses 1 to 61, A method comprising: identifying a second HLA variant in a second sample obtained from a subject at a second time point; and comparing a first HLA variant with the second HLA variant to monitor for cancer progression or recurrence. 77. The method according to Clause 76, wherein a second HLA variant for a second sample is identified in accordance with the method described in any one of Clauses 1 to 61. 78. The method of Clause 76 or Clause 77, further comprising selecting anti-cancer therapy for a target in response to the progression of cancer. 79. The method of Clause 76 or Clause 77, further comprising performing anticancer therapy in response to the progression of cancer. 80. The method of Clause 76 or Clause 77, further comprising adjusting anticancer therapy for a subject in response to the progression of cancer. 81. The method described in any one of the clauses 78 to 80, further comprising adjusting the dosage of anticancer therapy in response to the progression of cancer, or selecting a different anticancer therapy. 82. The method of Clause 81, further comprising administering a modified anti-cancer therapy to the target. 83. The method described in any one of the paragraphs 76 to 82, wherein the first time point is before the subject receives anticancer therapy, and the second time point is after the subject receives anticancer therapy. 84. The method described in any one of the paragraphs 76-83, relating to whether the subject has cancer, is at risk of having cancer, is routinely screened for cancer, or is suspected of having cancer. 85. The method described in any one of the clauses 76 to 84, wherein the cancer is a solid tumor. 86. The method described in any one of the clauses 76-85, wherein the cancer is a blood cancer. 87. The anti-cancer therapy is a method described in any one of the provisions 78 to 86, including chemotherapy, radiotherapy, immunotherapy, targeted therapy, or surgery. 88. The method of any one of the clauses 1 to 87, further comprising determining, identifying, or applying an HLA variant to a sample as a diagnostic value related to the sample. 89. The method according to any one of the clauses 1 to 88, further comprising generating a genomic profile for a subject based on the identification of an HLA variant. 90. The method according to Clause 89, wherein the target genome profile further includes results from comprehensive genome profiling (CGP) testing, gene expression profiling testing, cancer hotspot panel testing, DNA methylation testing, DNA fragmentation testing, RNA fragmentation testing, or any combination thereof. 91. The method according to Clause 89 or 90, wherein the genome profile of the subject further includes results from a nucleic acid sequencing-based test. 92. The method of any one of the clauses 89 to 91, further comprising selecting, administering, or applying anticancer therapy to a subject based on the generated genome profile. 93. The method described in any one of Clauses 1 to 61, wherein the identification of HLA variants in a sample is used in determining a proposed treatment for the subject. 94. Identification of HLA variants in a sample, as described in any one of Clauses 1 to 61, used when applying or performing treatment on a subject. 95. The method described in any one of the clauses 1 to 61, wherein the subject is a human. 96. A non-temporary computer-readable storage medium that stores one or more programs, wherein one or more programs include instructions, and when an instruction is executed by one or more processors of the system, a) Receive sequence read data for multiple sequence reads, b) Align the sequence read data to the first reference genome to identify the first set of HLA alignment reads and unmapped sequence reads. c) Search the HLA polymorphism database with the first set of HLA alignment reads and unmapped sequence reads to generate a target-specific reference genome. d) Align the first set of multiple HLA alignment reads and unmapped sequence reads to a target-specific reference genome to identify the second set of multiple HLA alignment reads. e) A non-temporary computer-readable storage medium for processing a second set of HLA alignment reads to identify HLA variants. 97. A non-temporary computer-readable storage medium as described in Clause 96, which uses a variant caller to process a second set of HLA alignment reads and identify HLA variants. 98. A non-temporary computer-readable storage medium as described in Clause 96 or 97, in which HLA variant identification does not require the use of sequence read data from a matched normal control. 99. A non-temporary computer-readable storage medium as described in any one of clauses 96 to 98, wherein the HLA polymorphism database is an IPD-IMGT / HLA database. 100. Further instructions, and when an instruction is executed by one or more processors, the system Align the HLA variant to the HLA code reference sequence and annotate the exon and intron boundaries in the HLA variant. The annotated HLA variant is translated into an HLA polypeptide sequence. By aligning the translated HLA polypeptide sequence with a reference HLA polypeptide sequence, we can identify amino acid changes in the translated HLA polypeptide sequence. A non-temporary computer-readable storage medium as described in any one of clauses 96-99, which categorizes translated HLA polypeptides as synonymous, non-start, non-stop, nonsense, or missense. 101. A non-temporary computer-readable storage medium as described in any one of the clauses 96 to 100, wherein the HLA variant is the HLA-I variant. 102. A non-temporary computer-readable storage medium as described in Clause 101, wherein the identified HLA-I variant is a variant of HLA-A, HLA-B, and / or HLA-C. 103. A non-temporary computer-readable storage medium as described in any one of clauses 96 to 102, wherein the HLA variant is an HLA-I and / or HLA-II variant. 104. A non-temporary computer-readable storage medium as described in any one of clauses 96 to 103, which includes the use of software tools to predict the probability that an input sequence corresponds to a known HLA sequence when searching a polymorphic database. 105. A non-temporary computer-readable storage medium as described in Clause 104, which is an Optitype software tool. 106. A non-temporary computer-readable storage medium as described in any one of clauses 96 to 105, wherein the variant caller includes a short variant caller. 107. A non-temporary computer-readable storage medium as described in any one of the clauses 96 to 106, wherein the subject is a human being.
[0230] From the above, specific implementations of the disclosed methods and systems have been illustrated and described, but it should be understood that various modifications can be made thereto and are intended herein. The invention is not intended to be limited by the specific examples provided herein. While the invention has been described with reference to the above specification, the descriptions and illustrations of preferred embodiments herein are not meant to be construed as limiting. Furthermore, it should be understood that all aspects of the invention are not limited to the specific descriptions, configurations, or relative proportions described herein, which depend on various conditions and variables. Various modifications in the forms and details of embodiments of the invention will be apparent to those skilled in the art. Therefore, the invention is also intended to encompass any such modifications, variations, and equivalents.
Claims
1. A method for identifying HLA variants, a) Receiving array read data for multiple array reads in one or more processors, b) Aligning the sequence read data to a first reference genome using one or more processors to identify a first plurality of HLA alignment reads and unmapped sequence reads, c) Using one or more processors, search an HLA polymorphism database with the first plurality of HLA alignment reads and unmapped sequence reads to generate a target-specific reference genome, d) Using one or more processors, align the first plurality of HLA alignment reads and unmapped sequence reads to the target-specific reference genome to identify a second plurality of HLA alignment reads, e) A method comprising processing the second plurality of HLA alignment reads using one or more processors to identify HLA variants.
2. The method according to claim 1, wherein the plurality of sequence reads are obtained using a targeted sequencing method that includes targeting HLA gene loci.
3. The method according to claim 1, wherein a variant caller is used to process the second plurality of HLA alignment leads and identify the HLA variants.
4. The method according to claim 1, wherein the identification of the HLA variant does not require the use of sequence read data from a matched normal control.
5. The method according to claim 1, wherein the HLA polymorphism database is an IPD-IMGT / HLA database.
6. The method according to claim 1, wherein the identified HLA variant is an HLA-I and / or HLA-II variant.
7. A method for identifying HLA variants, a) Receiving array read data for multiple array reads in one or more processors, b) Aligning the sequence read data to a first reference genome using one or more processors to identify a first plurality of HLA alignment reads and unmapped sequence reads, c) Aligning the first plurality of HLA alignment reads and unmapped sequence reads to a second reference genome using one or more of the above processors to identify a second plurality of HLA alignment reads, d) A method comprising using one or more processors to process the second plurality of HLA alignment reads to identify HLA variants, wherein the identification does not require the use of sequence read data from a matched normal control.
8. The method according to claim 7, wherein a variant caller is used to process the second plurality of HLA alignment leads and identify the HLA variants.
9. The method according to claim 7, wherein the second reference genome is a target-specific reference genome.
10. The method according to claim 7, wherein the target-specific reference genome is generated by searching an HLA polymorphism database using the multiple aligned reads and unmapped sequence reads of the first plurality of HLA alignment reads.
11. The method according to claim 10, wherein the HLA polymorphism database is an IPD-IMGT / HLA database.
12. The method according to claim 7, wherein the HLA variant is an HLA-I and / or HLA-II variant.
13. Using one or more of the aforementioned processors, align the HLA variant to an HLA code reference array and annotate the exon and intron boundaries in the HLA variant. Using one or more of the aforementioned processors, the annotated HLA variant is translated into an HLA polypeptide sequence. Using one or more of the aforementioned processors, the translated HLA polypeptide sequence is aligned to a reference HLA polypeptide sequence to identify amino acid changes in the translated HLA polypeptide sequence. The method according to claim 1, further comprising using one or more processors to categorize the translated HLA polypeptides as containing synonymous mutations, non-start mutations, non-stop mutations, nonsense mutations, or missense mutations.
14. The method according to claim 1, wherein the search of the HLA polymorphism database includes the use of a software tool that predicts the probability that an input sequence corresponds to a known HLA sequence.
15. The method according to claim 14, wherein the software tool is OptitePe.
16. The method according to claim 1, wherein the variant caller includes a short variant caller.
17. It is a system, a) One or more processors, b) A memory that is communicably connected to one or more processors and configured to store instructions, wherein when an instruction is executed by one or more processors, the system i) Receive sequence read data for multiple sequence reads, ii) Align the sequence read data with a first reference genome to identify a first set of HLA alignment reads and unmapped sequence reads. iii) Search the HLA polymorphism database with the first set of HLA alignment reads and unmapped sequence reads to generate a target-specific reference genome. iv) Align the first plurality of HLA alignment reads and unmapped sequence reads to the target-specific reference genome to identify a second plurality of HLA alignment reads. v) A system comprising a memory that processes the second plurality of HLA alignment reads to identify an HLA variant.
18. The system according to claim 17, wherein the identification of the HLA variant does not require the use of sequence read data from a matched normal control.
19. The system according to claim 17, wherein the HLA polymorphism database is an IPD-IMGT / HLA database.
20. The instruction further includes, and when the instruction is executed by the one or more processors, the system c) Align the HLA variant with the HLA code reference sequence and annotate the exon and intron boundaries in the HLA variant. d) Translate the annotated HLA variant into an HLA polypeptide sequence. e) Align the translated HLA polypeptide sequence with a reference HLA polypeptide sequence to identify amino acid changes in the translated HLA polypeptide sequence. f) The system according to claim 17, wherein the translated HLA polypeptides are categorized as synonymous, nonstart, nonstop, nonsense, or missense mutations.
21. The system according to claim 17, wherein searching the HLA polymorphism database includes using a software tool to predict the probability that an input sequence corresponds to a known HLA sequence.
22. The method according to claim 1, wherein the identification of the HLA variant in the sample is used in determining a proposed treatment for the subject.