Methods and systems for allele typing
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GUARDANT HEALTH INC
- Filing Date
- 2023-04-06
- Publication Date
- 2026-04-14
AI Technical Summary
Current HLA typing methods, such as SBT, often produce uncertain results due to insufficient sequencing and ambiguous haplotype fading, necessitating a more accurate allelotyping method using sequence information.
A method involving determining known allele sequences, aligning sequence reads from a target chromosome region to these known sequences, and calculating the number of sequence reads aligned to each allele to accurately determine the present alleles at specific loci.
This method provides accurate allelotyping by effectively utilizing sequence information, reducing uncertainties associated with existing methods, and improving the reliability of HLA typing.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 228,681, filed April 7, 2022, and U.S. Provisional Patent Application No. 63 / 487,926, filed March 2, 2023. Both provisional applications are incorporated by reference in their entirety for all purposes. [Background technology]
[0002] background The human leukocyte antigen (HLA) genes, which encode major histocompatibility complex proteins in humans, are located on the short arm of chromosome 6. These encoded HLA proteins are presented on the cell surface and can be classified into two distinct classes. Class I HLA proteins (A, B and C) present intracellular antigens of viral or tumor origin to cytotoxic T lymphocytes. Class II HLA proteins (DR, DQ and DP) present extracellular antigens to helper T cells. HLA genes are highly polymorphic and play important roles in immune-mediated diseases, tumor development processes, survival determination of transplanted organs or tissues, and drug hypersensitivity.
[0003] HLA genotyping is a complex procedure due to the extreme polymorphism of the major histocompatibility complex family. The most polymorphic regions, known as core exons, are exons 2 and 3 of HLA class I genes and exon 2 of HLA class II genes. The sequences of core exons are the most common targets of genotyping because they are believed to be essential determinants of antigen specificity that is beneficial for transplantation. However, many polymorphisms of other exons, introns and UTRs have been identified in population genetic and evolutionary studies, contributing to the creation of HLA nomenclature. Currently, HLA typing is performed using DNA-based methods, including SSP- (sequence-specific primers), SSO- (sequence-specific oligonucleotides) and RFLP-PCR (restriction fragment length polymorphism polymerase chain reaction) as well as sequence-based typing (SBT). SBT produces uncertain results due to insufficient sequencing and ambiguous haplotype phasing.
[0004] Therefore, there is a need for methods for accurate allele typing using sequence information. Summary of the Invention [Means for solving the problem]
[0005] Abstract Methods of allele typing are disclosed. Methods of treatment based on allele typing are disclosed.
[0006] Disclosed is a method that includes determining a plurality of known allele sequences; determining a plurality of sequence reads for a target region of a chromosome of a subject, where the target region of the chromosome comprises one or more genetic loci; aligning the plurality of sequence reads to the plurality of known allele sequences; determining, for each known allele sequence of the plurality of known allele sequences based on the alignment, a number of sequence reads that align to each known allele sequence; and determining, for the one or more genetic loci, the known allele sequences present at the one or more genetic loci based on the number of sequence reads that align to each known allele sequence.
[0007] Disclosed are methods that include determining a plurality of known allele sequences; determining a plurality of sequence reads for a target region of a chromosome of a subject, where the target region of the chromosome comprises one or more loci; aligning the plurality of sequence reads to the plurality of known allele sequences; determining, for each known allele sequence of the plurality of known allele sequences based on the alignment, a number of sequence read families (i.e., number of nucleic acid molecules--a sequence read family can be a group of sequence reads corresponding to a single nucleic acid molecule) that align to each known allele sequence; and determining, for one or more loci, known allele sequences present at the one or more loci based on the number of sequence read families that align to each known allele sequence.
[0008] Disclosed is a method that includes determining a plurality of known human leukocyte antigen (HLA) allele sequences; determining a plurality of sequence reads for a target region of a chromosome of a subject, where the target region of the chromosome comprises one or more loci; aligning the plurality of sequence reads to a plurality of known HLA allele sequences; determining, for each known HLA allele sequence of the plurality of known HLA allele sequences, a number of sequence reads aligned to each known HLA allele sequence based on the alignment; generating one or more supersets of the known HLA allele sequences based on the number of sequence reads aligned to each known HLA allele sequence; and determining, for one or more loci, a known HLA allele sequence present at the one or more loci based on the number of distinct reads in the one or more supersets of the known HLA allele sequences.
[0009] Disclosed is a method that includes determining a plurality of known human leukocyte antigen (HLA) allele sequences; determining a plurality of sequence reads for a target region of a chromosome of a subject, where the target region of the chromosome comprises one or more loci; aligning the plurality of sequence reads to a plurality of known HLA allele sequences; determining, for each known HLA allele sequence of the plurality of known HLA allele sequences based on the alignment, a number of sequence read families aligned to each known HLA allele sequence; generating one or more supersets of the known HLA allele sequences based on the number of sequence read families aligned to each known HLA allele sequence; and determining, for one or more loci, the known HLA allele sequences present at the one or more loci based on the number of distinct read families in the one or more supersets of the known HLA allele sequences.
[0010] A method is disclosed that includes the steps of determining a plurality of known killer cell immunoglobulin-like receptor (KIR) allele sequences; determining a plurality of sequence reads for a target region of a chromosome of a subject, where the target region of the chromosome comprises one or more loci; aligning the plurality of sequence reads to the plurality of known KIR allele sequences; determining, for each known KIR allele sequence of the plurality of known KIR allele sequences based on the alignment, a number of sequence reads aligned to each known KIR allele sequence; generating one or more supersets of the known KIR allele sequences based on the number of sequence reads aligned to each known KIR allele sequence; and determining, for one or more loci, a known KIR allele sequence present at the one or more loci based on the number of distinct reads in the one or more supersets of the known KIR allele sequences.
[0011] A method is disclosed that includes the steps of determining a plurality of known killer cell immunoglobulin-like receptor (KIR) allele sequences; determining a plurality of sequence reads for a target region of a chromosome of a subject, where the target region of the chromosome comprises one or more loci; aligning the plurality of sequence reads to the plurality of known KIR allele sequences; determining, for each known KIR allele sequence of the plurality of known KIR allele sequences based on the alignment, a number of sequence read families aligned to each known KIR allele sequence; generating one or more supersets of the known KIR allele sequences based on the number of sequence read families aligned to each known KIR allele sequence; and determining, for one or more loci, the known KIR allele sequences present at the one or more loci based on the number of distinct read families in the one or more supersets of the known KIR allele sequences.
[0012] Disclosed is a method that includes determining a plurality of known allele sequences; determining a plurality of sequence reads for a target region of a chromosome of a subject, where the target region of the chromosome comprises one or more genetic loci; aligning the plurality of sequence reads to the plurality of known allele sequences; determining, for each known allele sequence of the plurality of known allele sequences based on the alignment, a number of sequence reads that align to each known allele sequence; generating one or more supersets of the known allele sequences based on the number of sequence reads that align to each known allele sequence; and determining, for one or more loci, the known allele sequences present at the one or more loci based on the number of distinct reads in the one or more supersets of known allele sequences.
[0013] Disclosed is a method that includes determining a plurality of known allele sequences; determining a plurality of sequence reads for a target region of a chromosome of a subject, where the target region of the chromosome comprises one or more loci; aligning the plurality of sequence reads to the plurality of known allele sequences; determining, for each known allele sequence of the plurality of known allele sequences, a number of sequence read families aligned to each known allele sequence based on the alignment; generating one or more supersets of the known allele sequences based on the number of sequence read families aligned to each known allele sequence; and determining, for one or more loci, the known allele sequences present at the one or more loci based on the number of distinct read families in the one or more supersets of the known allele sequences.
[0014] In some embodiments, the results of the systems and methods disclosed herein are used as input to generate a report. The report may be in paper or electronic format. For example, the determination of allele type (e.g., allele sequence) as determined by the methods and systems disclosed herein may be displayed directly in such a report.
[0015] Various steps of the methods disclosed herein, or steps performed by the systems disclosed herein, may be performed at the same or different times, in the same or different geographic locations, e.g., countries, and / or by the same or different persons.
[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain embodiments and, together with the description herein, serve to explain certain principles of the methods, computer-readable media and systems disclosed herein. The description provided herein is better understood when read in conjunction with the accompanying drawings, which are included by way of example and not by way of limitation. It is understood that like reference numerals identify like components throughout the drawings, unless otherwise indicated by context. It is also understood that some or all of the drawings may be schematic diagrams for illustrative purposes and do not necessarily depict the exact relative size or position of the elements shown. [Brief description of the drawings]
[0017] [Figure 1-1] FIG. 1A is a flow chart that diagrammatically depicts exemplary method steps for allele typing.
[0018] [Figure 1-2] FIG. 1B is a flow chart that diagrammatically depicts another exemplary method steps for allele typing.
[0019] [Diagram 2] FIG. 2 shows an example of a system for allele typing.
[0020] [Diagram 3] Figure 3 shows an example of a k-mer data structure.
[0021] [Figure 4] FIG. 4 shows an example of a Hasse diagram.
[0022] [Diagram 5] FIG. 5 shows an example of a graph data structure.
[0023] [Figure 6] FIG. 6 shows an example of a graph data structure.
[0024] [Figure 7] FIG. 7 shows an example of a graph data structure.
[0025] [Figure 8] FIG. 8 shows an example of a superset comparison.
[0026] [Figure 9] FIG. 9 shows an example of the method.
[0027] [Figure 10] FIG. 10 shows an example of the method.
[0028] [Figure 11-1] Figure 11 shows examples of different alleles with high homology. The alleles in Figure 11 code for distinct products at the protein level, even though their genomic content is largely identical. For example, two alleles that code for different proteins may in fact differ by only one or two bases. [Figure 11-2] Same as above.
[0029] [Figure 12] Figure 12 shows an example flow chart of the disclosed method. The disclosed method can be configured to run in one or more modes, such as a "build index" mode (for generating a kmer map database) and a "call allele" mode (for calling alleles on the input sample).
[0030] [Figure 13]Figure 13 shows an example of the hierarchical structure of allele names. HLA allele names are encoded by a string containing the gene name (A for HLA-A) and up to four pairs of numbers, e.g. A*01:02:03:04. Outline two allele names: · Match the first two pairs of numbers: the alleles code for the same protein, but they may have different genomic content at the exon level; · Match the first three pairs of numbers: the alleles have the same genomic content at the exon level (their bases within the exons are identical); · Match all four pairs of numbers: the alleles have the same genomic content in both exons and introns.
[0031] [Figure 14] Figure 14 shows a screen capture of the Integrative Genomics Viewer showing one allele being called by Kmerizer and the set of all supporting reads for the called allele. This view contains only sequence reads that are perfect matches to the called allele (without regard to alignment to the reference genome).
[0032] [Figure 15] Figure 15 shows the analytical performance of kmerizer for two simulated sets of 100 samples each, comparable to the sensitivity of HISAT-genotypes. The "normal" set contains 100 simulated samples with reads in 14 HLA genes (HISAT-genotypes do not call genes in DRB3, DRB4 and DRB5) with a simulated error rate of 0.1%. The "difficult" set contains 100 simulated samples with reads in 8 HLA genes selected from a pool of alleles known to share high homology levels with a simulated error rate of 0.1%.
[0033] [Figure 16]Figure 16 shows the analytical performance of Kmerizer on a simulated set of 100 samples. The set contains 100 simulated samples with reads in 17 KIR genes and a simulated error rate of 0.1%.
[0034] [Figure 17] FIG. 17 shows the analytical performance of Kmerizer for 19 samples with known HLA status for four HLA genes (HLA-A, HLA-B, HLA-C, and HLA-DQB1), which was comparable to the performance of HISAT-genotype.
[0035] [Figure 18] FIG. 18 shows a comparison of HLA germline allele prevalence between methylation epigenome detection samples (n=2004) and Allele Frequency Net Database (AFND) genotype data from different geographic regions.
[0036] [Figure 19] FIG. 19 shows the distribution of germline HLA allele counts for which immunogenic neoantigens were predicted.
[0037] [Figure 20] FIG. 20 shows the overall germline HLA allele distribution for all methylated epigenomic detection samples.
[0038] [Figure 21] FIG. 21 shows that using a reference from the IMGT / HLA database, germline allele(s) and homozygous / heterozygous status were randomly assigned based on AFND to generate germline reads, then somatic mutations were randomly generated at exon positions, and then variant reads were generated from the modified reference.
[0039] [Figure 22] FIG. 22 shows the variant allele frequency (AF) distribution of the detected variants.
[0040] [Figure 23] Figure 23 shows TMB versus predicted neoepitope affinity across genes. Spearman correlations listed by group.
[0041] [Figure 24] FIG. 24 shows the distribution of neoantigen binding affinities by cancer type. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0042] I. Definition In order that this disclosure may be more readily understood, certain terms are first defined below. Further definitions for the following terms and other terms may be set forth throughout this specification. In the event that the definitions of terms set forth below conflict with definitions in patent applications or issued patents incorporated by reference, the definitions set forth in this application should be used to understand the meaning of the terms.
[0043] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a method" includes one or more methods and / or steps of the type described herein and / or that would become apparent to one of ordinary skill in the art upon reading this disclosure, and so forth. It is also understood that there is an implicit "about" before temperatures, concentrations, times, numbers of bases or base pairs, coverage, and so forth discussed in this disclosure, and thus, only a few and only a few equivalents are within the scope of this disclosure. In this application, the use of the singular includes the plural unless specifically stated otherwise. Also, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" is not intended to be limiting.
[0044] It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting. Moreover, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs. In describing and claiming the method, computer readable medium, and system, the following terminology and its grammatical variants will be used in accordance with the definitions set forth below.
[0045] About: As used herein, "about" or "approximately" refers to a value or element that is similar to a stated reference value or element. In certain embodiments, the term "about" or "approximately" refers to a range of values or elements that are within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction of the stated reference value or element, unless otherwise stated or otherwise clear from the context (except where such number exceeds 100% of the possible values or elements).
[0046] Adapter: As used herein, "adapter" or "adaptor" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, or less than about 50 nucleotides in length), typically at least partially double-stranded, used to ligate to either or both ends of a given sample nucleic acid molecule. The adapter can include a nucleic acid primer binding site to allow amplification of the nucleic acid molecule flanked by the adapters at both ends, and / or a sequencing primer binding site, including a primer binding site for sequencing applications, such as various next generation sequencing (NGS) applications. The adapter can also include a binding site for a capture probe, such as an oligonucleotide bound to a flow cell support or the like. The adapter can also include a nucleic acid tag, as described herein. The nucleic acid tag is typically positioned relative to the primer and sequencing primer binding sites such that the nucleic acid tag is included in the amplicon and sequencing reads of a given nucleic acid molecule. Adapters of the same or different sequences can be ligated to each end of the nucleic acid molecule. In certain embodiments, the same adaptor is linked to each end of the nucleic acid molecule, except that the nucleic acid tag is different in terms of its sequence.In some embodiments, the adaptor is a Y-shaped adaptor, one end of which is blunt-ended or tailed as described herein for joining to the nucleic acid molecule, and the nucleic acid molecule is also blunt-ended or tailed with one or more complementary nucleotides.In yet another exemplary embodiment, the adaptor is a bell-shaped adaptor, which includes a blunt end or tailed end for joining to the nucleic acid molecule to be analyzed.Other exemplary adaptors include the adaptor with T tail and the adaptor with C tail.
[0047] Administer: As used herein, "administering" or "administering" a therapeutic agent (e.g., an immunological therapeutic agent, a DNA damage response (DDR) inhibitor (e.g., a poly(ADP-ribose) polymerase (PARP) inhibitor (PARPi)) to a subject means giving, applying, or contacting the composition with the subject. Administration can be accomplished by any of a number of routes, including, for example, topical, oral, subcutaneous, intramuscular, intraperitoneal, intravenous, intrathecal, and intradermal routes.
[0048] Align: As used herein, "align", "alignment" and "aligning" in the context of nucleic acids refer to arranging DNA or RNA sequences to identify regions of similarity. Similarity may relate to nucleotide sequence, structural, functional and / or evolutionary relationships between sequences. Alignment of DNA sequences includes alignment of the genomic DNA of one sequence to the genomic DNA of at least one other sequence. Such alignment may exclude non-genomic DNA, such as molecular barcodes and padding bases. For example, the genomic DNA of a sequence read may be aligned to the genomic DNA of a reference DNA sequence, excluding any molecular tag or adapter sequences that may be attached to the sequence read.
[0049] As used herein, a statement that a nucleotide "corresponds to" a nucleotide in a sequence refers to the nucleotide identified upon alignment with the sequence to maximize identity using a standard alignment algorithm, such as the GAP algorithm.
[0050] As used herein, "sequence identity", "sequence homology", or "identity" refers to the number of identical or similar nucleotide bases in an alignment between two or more polynucleotide sequences. In one non-limiting example, "at least 90% identical to" refers to a percent identity of 90-100% to a reference polynucleotide. Identity at a level of 90% or higher refers to the fact that no more than 10% (i.e., 10 out of 100) of the nucleotides in the test polynucleotide differ from those of the reference polynucleotide, assuming for illustrative purposes that test and reference polynucleotide lengths of 100 nucleotides are compared. Such differences may be represented as point mutations randomly distributed over the entire length of the nucleotide sequence, or they may be clustered at one or more locations of variable length up to the maximum allowable, e.g., 10 / 100 nucleotide difference (approximately 90% identity). Differences are defined as nucleic acid substitutions, insertions, or deletions.
[0051] Sequence identity can be determined by sequence alignment of nucleic acid sequences to identify regions of similarity or identity. For purposes herein, sequence identity is generally determined by alignment to identify identical bases. Alignment can be local or global. Matches, mismatches and gaps can be identified between the sequences being compared. A gap is a null nucleotide inserted in the sequence of the aligned sequence, so that identical or similar characters are aligned. In general, there can be internal and terminal gaps. Sequence identity can be determined as the number of identical bases / minimum sequence length x 100 by taking gaps into account. When using gap penalties, sequence identity can be determined without penalty for end gaps (e.g., no penalty is imposed on terminal gaps). Alternatively, sequence identity can be determined as the number of identical positions / (total length of aligned sequences) x 100 without taking gaps into account.
[0052] As used herein, a "global alignment" is an alignment that aligns two sequences from beginning to end, with each base in each sequence being aligned only once. The alignment is generated regardless of whether there is similarity or identity between the sequences. For example, 50% sequence identity based on a "global alignment" means that 50% of the bases are the same in the full sequence alignment of two compared sequences, each 100 nucleotides in length. It will be understood that global alignment can be used to determine sequence identity as well, even if the lengths of the aligned sequences are not the same. Differences at the ends of the sequences are taken into account when determining sequence identity, unless "no end gap penalty" is selected. Generally, global alignment is used for sequences that share significant similarity over most of their length. An exemplary algorithm for performing global alignment includes the Needleman-Wunsch algorithm (Needleman et al. J. Mol. Biol. 48: 443 (1970)). Exemplary programs for performing global alignments are publicly available and include the Global Sequence Alignment Tool available at the National Center for Biotechnology Information (NCBI) website (ncbi.nlm.nih.gov / ), and the program available at deepc2.psi.iastate.edu / aat / align / align.html.
[0053] As used herein, a "local alignment" is an alignment that aligns two sequences, but only aligns the parts of the sequences that share similarity or identity. Thus, a local alignment determines whether a subsegment of one sequence exists in another sequence. If there is no similarity, no alignment will be returned. Local alignment algorithms include BLAST or Smith-Waterman algorithms (Adv. Appl. Math. 2: 482 (1981)). For example, 50% sequence identity based on "local alignment" means that in the full sequence alignment of two compared sequences of any length, a region of similarity or identity of 100 nucleotides in length has 50% of the bases that are the same within the region of similarity or identity.
[0054] Allele: As used herein, "allele" or "allelic variant" refers to a specific genetic variant at a defined genomic location or locus. Allelic variants are usually expressed at a frequency of 50% (0.5) or 100%, depending on whether the allele is heterozygous or homozygous. For example, germline variants are inherited and usually have a frequency of 0.5 or 1. However, somatic variants are acquired variants and usually have a frequency of <0.5. Major and minor alleles of a locus refer to nucleic acids that have within them a locus that are occupied by nucleotides of a reference sequence and variant nucleotides that differ from the reference sequence, respectively. Measurements at a locus can take the form of an allele fraction (AF), which measures the frequency with which an allele is observed in a sample.
[0055] Amplify: As used herein, "amplify" or "amplification" in the context of nucleic acids refers to the production of multiple copies of a polynucleotide, or of a portion of a polynucleotide, typically starting from a small amount of the polynucleotide (e.g., a single polynucleotide molecule), where the amplification product or amplicon is generally detectable. Polynucleotide amplification encompasses a variety of chemical and enzymatic processes.
[0056] Barcode: As used herein, "barcode" in the context of nucleic acids refers to a nucleic acid molecule having a sequence that can serve as an identifier for a molecule (molecular barcode), a partition (partition barcode), or a sample (sample barcode or sample index). For example, individual "barcode" sequences are typically added to DNA fragments during next generation sequencing (NGS) library preparation so that each read can be identified and sorted prior to final data analysis.
[0057] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" refers to nucleic acid that is not contained within a cell or otherwise bound to a cell. In some embodiments, "cell-free nucleic acid" refers to nucleic acid that is not contained within a cell or otherwise bound to a cell at the time of isolation from a subject. Cell-free nucleic acid may include, for example, all non-encapsulated nucleic acid provided from a body fluid from a subject (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acid includes DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-binding RNA (piRNA), long non-coding RNA (long ncRNA), and / or cfDNA and / or cfRNA derived from any fragments thereof. Cell-free nucleic acid may be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acid may be released into body fluids by secretion or cell death processes, such as cell necrosis or apoptosis. Some cell-free nucleic acids are released from cancer cells into body fluids, such as circulating tumor DNA (ctDNA). Others are released from healthy cells. ctDNA can be non-encapsulated tumor-derived fragmented DNA. Another example of cell-free nucleic acid is fetal DNA circulating freely in maternal bloodstream, also called cell-free fetal DNA (cffDNA). Cell-free nucleic acid can have one or more epigenetic modifications, for example, cell-free nucleic acid can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated and / or citrullinated.
[0058] Cellular origin: As used herein, "cellular origin" in the context of cell-free nucleic acid refers to the cell type from which a given cell-free nucleic acid molecule originates or is otherwise generated (e.g., by apoptotic or necrotic processes, etc.). In certain embodiments, for example, a given cell-free nucleic acid molecule may originate from a tumor cell (e.g., a cancerous lung cell, etc.) or a non-tumor or normal cell (e.g., a non-cancerous lung cell, etc.).
[0059] Classifier: As used herein, a "classifier" generally refers to algorithmic computer code that receives test data as input and produces as output a classification of the input data as belonging to one or another class (e.g., DNA Damage Repair Deficient (DDRD) or not DDRD; tumor DNA or non-tumor DNA).
[0060] Contiguous sequence: As used herein, "contiguous sequence" or "contig" refers to a set of contiguous nucleic acid bases (possibly overlapping) that together represent the consensus of a nucleic acid region. It may refer to any contiguous stretch of sequence data for a sequence read, or for a reference interval, or for the assembly of reads and / or reference segments (e.g., assembled reads).
[0061] Copy number variant: As used herein, "copy number variant", "CNV", or "copy number variation" refers to the phenomenon in which sections of the genome are repeated such that the number of repeats in the genome varies among individuals in the population under consideration.
[0062] Coverage: As used herein, the terms "coverage", "total molecule count" or "total allele count" are used interchangeably. These terms refer to the total number of nucleic acid molecules at a particular genomic location in a given sample.
[0063] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to a natural or modified nucleotide with a hydrogen group at the 2' position of the sugar moiety. DNA typically comprises a chain of nucleotides that comprises deoxyribonucleosides that comprise one of the four types of nucleobases, namely, adenine (A), thymine (T), cytosine (C) and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to a natural or modified nucleotide with a hydroxyl group at the 2' position of the sugar moiety. RNA typically comprises a chain of nucleotides that comprises ribonucleosides that comprise one of the four types of nucleobases, namely, A, uracil (U), G and C. As used herein, the term "nucleotide" refers to a natural or modified nucleotide. Certain nucleotide pairs specifically bind to each other in a complementary manner (referred to as complementary base pairing). In DNA, alanine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). In RNA, alanine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand that is composed of nucleotides that are complementary to those in the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "sequence information," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "gene sequence," or "fragment sequence," or "nucleic acid sequencing read" refers to any information or data that indicates the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule (e.g., a whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment) of a nucleic acid such as DNA or RNA.It should be understood that the present teachings contemplate sequence information being obtained using any available type of technique, platform or technology, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion or pH-based detection systems, and electronic signature-based systems.
[0064] Detect: As used herein, "detect," "detecting," or "detection" refers to the act of determining the existence or presence of one or more target nucleic acids (e.g., nucleic acids having a targeted mutation or other marker) in a sample.
[0065] Enriched sample: As used herein, "enriched sample" refers to a sample that is enriched for a specific region of interest. Samples can be enriched by amplifying the region of interest or by using single-stranded DNA / RNA probes or double-stranded DNA probes that can hybridize to the nucleic acid molecules of interest (e.g., SureSelect® probes, Agilent Technologies). In some embodiments, enriched sample refers to a subset or part of the processed sample that is enriched, and the subset or part of the processed sample that is enriched contains a cell-free polynucleotide sample or nucleic acid molecules from the polynucleotide sample.
[0066] Gene: As used herein, "gene" refers to any segment of DNA associated with a biological function. Thus, genes include coding sequences and, optionally, regulatory sequences required for their expression. Genes also optionally include non-expressed DNA segments, for example, forming recognition sequences for other proteins.
[0067] Genomic region: As used herein, "genomic region" refers to a fixed location or section on a chromosome, such as the location of a gene or genomic marker. Exemplary genomic markers include transcription factor binding regions (e.g., CTCF binding regions, etc.), distal regulatory elements (DREs), repetitive elements (e.g., microsatellites, etc.), intron-exon or exon-intron junctions, and transcription start sites (TSSs), etc.
[0068] Indel: As used herein, "indel" refers to a mutation comprising an insertion or deletion of a nucleotide position in the genome of a subject.
[0069] k-mer: As used herein, "k-mer" refers to a string (or substring) of length k contained within a genomic sequence. In some embodiments, the length k can be 100 bp or less, 150 bp or less, 200 bp or less, or 250 bp or less.
[0070] Match: As used herein, "match" means that at least a first value or element is at least approximately equal to at least a second value or element. In certain embodiments, for example, the cellular origin of at least a subset of DNA molecules from a cfDNA sample is determined when there is at least a substantial or approximate match between the test sample distribution of cfDNA fragment characteristics and the reference sample distribution of cfDNA fragment characteristics.
[0071] Mutation: As used herein, "mutation", "nucleic acid variant", "variant", "genetic abnormality" or "genetic lesion" refers to a variation from a known reference sequence, including, for example, mutations such as single nucleotide variants (SNVs), copy number variants or variations (CNVs) / abnormalities, insertions or deletions (indels), truncations, gene fusions, transversions, translocations, frameshifts, duplications, repeat expansions, and epigenetic variants. The mutations may be germline or somatic mutations. In some embodiments, the reference sequence for comparison is the wild-type genomic sequence of the subject's species providing the test sample, typically the human genome. In certain cases, the mutation or variant is a "tumor-associated genetic variant" that causes or at least contributes to tumor formation. The driver mutation may be a mutation associated with cancer or another abnormal biological condition. The presence of a driver mutation may indicate cancer diagnosis, stratification of a subject into subtypes, tumor burden, tumors in tissues or organs, tumor metastasis, therapeutic efficacy, or resistance to treatment.
[0072] Next generation sequencing: As used herein, "next generation sequencing" or "NGS" refers to a massively parallel sequencing technology that has increased throughput compared to older Sanger and capillary electrophoresis-based approaches, e.g., has the ability to simultaneously generate hundreds of thousands of relatively short sequence reads. Some examples of next generation sequencing techniques include, but are not limited to, single base synthesis, sequencing by ligation, and sequencing by hybridization.
[0073] Nucleic acid tag: As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500, about 100, about 50 or about 10 nucleotides in length) that is used to label nucleic acid molecules to distinguish nucleic acids from different samples of different types or that have undergone different treatments (e.g., representing a sample index) or that is used to label nucleic acid molecules to distinguish different nucleic acid molecules in the same sample of different types or that have undergone different treatments (e.g., representing a molecular tag). Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can have the same length or various lengths, as appropriate. Nucleic acid tags can include double-stranded molecules with one or more blunt ends, can include 5' or 3' single-stranded regions (e.g., overhangs), and / or can include one or more other single-stranded regions at other locations within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information such as the source sample, form or processing of a given nucleic acid. Nucleic acid tags can also be used to allow pooling and / or parallel processing of multiple samples containing nucleic acids with different nucleic acid tags and / or sample indexes, which are then deconvoluted by reading the nucleic acid tags. Nucleic acid tags can also be referred to as molecular identifiers or tags, sample identifiers, index tags, and / or barcodes. Additionally or alternatively, nucleic acid tags can be used to distinguish different molecules in the same sample. This includes, for example, uniquely tagging different nucleic acid molecules in a given sample, or non-uniquely tagging such molecules. In the case of non-unique tagging applications, tags with a limited number of different sequences can be used to tag nucleic acid molecules, and thus different molecules can be distinguished, for example, based on the start and / or stop positions where they are located in the reference genome selected in combination with at least one nucleic acid tag.Typically, a sufficient number of different nucleic acid tags are used so that the probability that any two molecules have the same start / stop position and also have the same nucleic acid tag is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% chance).Some nucleic acid tags include multiple molecular identifiers to label samples, forms of nucleic acid molecules within samples, and nucleic acid molecules within forms that have the same start and stop positions.Such nucleic acid tags can be referred to using the exemplary format "A1i", where capital letters indicate sample types, Arabic numerals indicate forms of molecules within samples, and lowercase Roman numerals indicate molecules within forms.
[0074] Polynucleotide: As used herein, "polynucleotide", "nucleic acid", "nucleic acid molecule", or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) joined by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides often range in size from a few monomeric units, e.g., 3-4, to several hundred monomeric units. Whenever a polynucleotide is represented by a sequence of letters, such as "ATGCCTG", it is understood that the nucleotides are in 5'→3' order from left to right, unless otherwise noted, and that in the case of DNA, "A" denotes deoxyadenosine, "C" denotes deoxycytidine, "G" denotes deoxyguanosine, and "T" denotes deoxythymidine. The letters A, C, G, and T may be used to refer to the bases themselves, to nucleosides, or to nucleotides that contain the bases, as is common in the art.
[0075] Prevalence: As used herein, "prevalence" in the context of nucleic acid variants refers to the degree, spread, or frequency to which a given nucleic acid variant is seen or found in a given sample (e.g., a given bodily fluid sample, a given non-bodily fluid sample, etc.) or other population (e.g., a given population of bodily fluid samples, a given population of non-bodily fluid samples, etc.).
[0076] Reference sample: As used herein, a "reference sample" or "reference cfDNA sample" refers to a sample of known composition and / or known to have or lack certain characteristics (e.g., known nucleic acid variant(s), known cellular origin, known tumor fraction and / or known coverage, etc.) that is analyzed together with or compared to a test sample to assess the accuracy of an analytical procedure. A reference sample dataset typically includes at least about 25 to at least about 30,000 or more reference samples. In some embodiments, the reference sample dataset comprises about 50, 75, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,500, 5,000, 7,500, 10,000, 15,000, 20,000, 25,000, 50,000, 100,000, 1,000,000, or more reference samples.
[0077] Reference sequence: As used herein, "reference sequence" or "reference genome" refers to a known sequence used for comparison with an experimentally determined sequence. For example, the known sequence can be a whole genome, a chromosome, or any segment thereof. A reference sequence typically comprises at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000, at least about 10,000, at least about 100,000, at least about 1,000,000, at least about 10,000,000, at least about 1,000,000,000, or more nucleotides. A reference sequence can be aligned with a single continuous sequence of a genome or chromosome, or can comprise discontinuous segments that align with different regions of a genome or chromosome. Exemplary reference sequences include, for example, the human genome, e.g., hG19 and hG38.
[0078] Sample: As used herein, "sample" refers to any biological sample that can be analyzed by the methods and / or systems disclosed herein. In certain aspects of the present disclosure, the sample is a bodily fluid sample, e.g., whole blood or a fraction thereof, lymph, urine, and / or cerebrospinal fluid, among other types of bodily fluids that are sources of acellular (circulating, not contained within or otherwise bound to cells) nucleic acids. In certain implementations, the bodily fluid sample is a plasma sample, which is the fluid portion of whole blood excluding cells such as red blood cells and white blood cells. In some implementations, the bodily fluid sample is a serum sample, i.e., plasma that lacks fibrinogen. In some aspects of the present disclosure, the sample is a "non-bodily fluid sample" or "non-plasma sample", i.e., a biological sample other than a "bodily fluid sample", e.g., a biological sample such as a cell and / or tissue sample that is a source of nucleic acids other than acellular nucleic acids. In some embodiments, the sample may include a body tissue or tissue biopsy.
[0079] Sensitivity: As used herein, "sensitivity" in the context of a given assay or method refers to the ability of the assay or method to detect and distinguish between targeted analytes (e.g., nucleic acid variants) and non-targeted analytes.
[0080] Sequence fragment: As used herein, a "sequence fragment" for constructing an allele k-mer data structure may include dividing a known allele sequence into k-mer amounts. For example, k-mer amounts having a length of about 100 nucleotides to about 200 nucleotides. In an embodiment, the k-mer amount may have a length of 125-175 nucleotides, 130-160 nucleotides, 135-155 nucleotides, 140-150 nucleotides. In an embodiment, the k-mer amount may have a length of 140, 141, 142, 143, 144, or 145 nucleotides. Constructing an allele k-mer data structure may include associating each k-mer with metadata. The metadata may include, for example, the amount of alleles containing the k-mer, as well as an indication of the allele identifier and k-mer start position for each allele containing the k-mer.
[0081] Sequence read: As used herein, "read," "sequence read," or "sequencing read" refers to a sequence of base pairs corresponding to all or a portion of a sequence fragment.
[0082] Sequencing: As used herein, "sequencing" refers to any of several techniques used to determine the sequence (e.g., identity and order of monomeric units) of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Exemplary sequencing methods include targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, coamplification-PCR at lower denaturation temperatures (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, and genomic sequencing. Sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single molecule sequencing, single-base synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof, are included, but are not limited to. In some embodiments, sequencing can be performed by a genetic analyzer, such as a genetic analyzer commercially available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among others.
[0083] Sequence information: As used herein, "sequence information" in the context of a nucleic acid polymer means the order and / or identity of the monomeric units (e.g., nucleotides) in that polymer.
[0084] Single nucleotide variant: As used herein, "single nucleotide variant" or "SNV" refers to a mutation or variation of a single nucleotide present at a specific position in a genome or a given genomic sequence.
[0085] Specificity: As used herein, "specificity" in the context of a diagnostic analysis or assay refers to the degree to which the analysis or assay detects the intended target analyte to the exclusion of other components of a given sample.
[0086] Status: As used herein, "status" in the context of a subject refers to one or more conditions of a given subject, for example, whether or not the subject has cancer.
[0087] Subject: As used herein, "subject" or "test subject" refers to an animal, such as a mammalian species (e.g., human) or avian (e.g., avian) species, or other organisms, such as plants. More specifically, the subject can be a vertebrate, such as a mammal, such as a mouse, a primate, a monkey, or a human. Animals include farm animals (e.g., production cattle, dairy cattle, poultry, horses, and pigs, etc.), sport animals, and companion animals (e.g., pets or support animals). The subject can be a healthy individual, an individual having or suspected of having a disease or predisposition to a disease, or an individual in need of treatment or suspected of needing treatment. The terms "individual" or "patient" are intended to be synonymous with "subject." In some embodiments, the subject is a human having or suspected of having cancer. For example, the subject can be an individual who has been diagnosed with cancer, an individual who will be undergoing cancer treatment, and / or an individual who has undergone at least one cancer treatment. The subject can also be in remission of cancer. As another example, the subject may be an individual diagnosed with an autoimmune disease. As another example, the subject may be a female individual who has been diagnosed or suspected with a disease, such as cancer, an autoimmune disease, who is pregnant or planning to become pregnant. "Reference subject" refers to a subject who is known to have or lack a particular characteristic (e.g., a known cancer or disease state, a known nucleic acid variant(s), a known cell origin, a known tumor fraction, and / or a known coverage, etc.).
[0088] Threshold: As used herein, "threshold" refers to a discretely determined value used to characterize or classify experimentally determined values. In certain embodiments, for example, a "threshold value" refers to a selected value to which a quantitative value is compared to determine that a given target nucleic acid variant is absent at a given locus.
[0089] Value: As used herein, a "value" or "score" generally refers to an entry in a dataset, which can be anything that characterizes the feature to which the value refers. This includes, but is not limited to, a number, a word or phrase, a symbol (e.g., + or -), or a degree. Detailed Description II. Introduction
[0090] Provided herein are methods and systems for determining allele types from a nucleic acid sample obtained from a test subject. In some aspects, the nucleic acid sample can be, but is not limited to, cell-free nucleic acid (cfNA), genomic DNA, or RNA. In some embodiments, the nucleic acid sample can be derived from a specific chromosome and / or from a specific region of a chromosome. In some embodiments, the nucleic acid sample can be derived from all or a portion of the major histocompatibility complex (MHC) region of human chromosome 6. In some embodiments, the nucleic acid sample can be derived from all or a portion of the killer cell immunoglobulin-like receptor (KIR) region of human chromosome 19. In some embodiments, the nucleic acid sample can be derived in whole or in part from both the MHC region of human chromosome 6 and the KIR region of human chromosome 19. a.Major histocompatibility complex (MHC)
[0091] MHC is a group of cell surface molecules encoded by a large family of genes that control most of the immune system in all vertebrates. The primary function of the major histocompatibility complex is to bind peptide fragments derived from pathogens and present them on the cell surface for recognition by appropriate T cells.
[0092] The MHC contains two types of polymorphic MHC genes, class I and class II genes, which code for two groups of structurally distinct but homologous proteins, and other non-polymorphic genes whose products are involved in antigen presentation.
[0093] Human MHC is called human leukocyte antigen (HLA). HLA proteins are encoded by MHC genes. HLA class I antigens include HLA-A, HLA-B, and HLA-C. HLA class II antigens include HLA-DR, HLA-DQ, HLA-DP, HLA-DM, HLA-DOA, and HLA-DOB. HLA (A, B, and C) corresponding to MHC class I present peptides from inside the cell. HLA (DP, DM, DOA, DOB, DQ, and DR) corresponding to MHC class II present antigens from outside the cell to T lymphocytes. MHC molecules mediate the binding of a given T cell receptor to a given antigen, so MHC polymorphism between individuals modulates the T cell response to a given antigen. Clinical applications of HLA typing include vaccine testing, disease association, adverse drug reactions, platelet transfusion, and organ and stem cell transplantation.
[0094] There are numerous alleles at each HLA locus (A1, A2, A3, etc.). Each person inherits one complete set of HLA alleles (haplotype) from each parent, and this combination of encoded proteins makes up that person's HLA type (e.g., different antigens A23, A31, B7, B44, C7, C8, DR4, DR7, DQ2, DQ7, DP2, DP3). There are over 50,000 different known HLA types.
[0095] Certain diseases, especially those of an autoimmune nature, are associated with specific HLA types. The level of association varies from disease to disease, and in general there is a lack of strong match between HLA type and disease. The exact mechanisms underlying most HLA-disease associations are not fully understood, and other genetic and environmental factors may also be involved. Among the most prominent associations are narcolepsy and HLA-DQB1 * 0602 / HLA-DRB1 * 1501, Ankylosing spondylitis and HLA-B27, and Celiac disease and HLA-DQB1 *02. HLA-A1, B8, DR17 haplotypes are frequently associated with autoimmune disorders. Rheumatoid arthritis is associated with a specific sequence of amino acids 66-75 in the DRβ1 chain, which is common to HLA-DR4 and the major DR1 subtypes. Type I diabetes is associated with HLA-DR3,4 heterozygotes, and the absence of asparagine at position 57 on the DQβ1 chain appears to confer susceptibility to the disease.
[0096] HLA-A, HLA-B, and HLA-DR have long been known as major transplantation antigens. Recent clinical data indicate that HLA-C matching also influences the clinical outcome of hematopoietic stem cell transplantation. Allogeneic hematopoietic stem cell transplantation is used to treat hematologic malignancies, severe aplastic anemia, severe congenital immune deficiencies, and selected inherited metabolic diseases. The HLA system is the major histocompatibility barrier in stem cell transplantation, and the degree of HLA matching is predictive of clinical outcome. HLA mismatch between the recipient and stem cell donor represents a risk factor for graft rejection / engraftment failure and acute graft-versus-host disease (GVHD). The latter is caused by immunocompetent donor T cells contained in the stem cell product. Depletion of T cells from donor bone marrow results in a lower incidence of acute GVHD but a higher incidence of the subsequent complications of engraftment failure, graft rejection, malignant disease recurrence (i.e., loss of graft-versus-leukemia effect), failed immune recovery, and Epstein-Barr virus-associated lymphoproliferative disorder. B. Killer cell immunoglobulin-like receptor (KIR)
[0097] Natural killer (NK) cells are part of the innate immune system and are specialized in early defense against infections as well as tumors. NK cells were first discovered as a result of their ability to kill tumor cell targets. Unlike cytotoxic T cells, NK cells can kill targets in a non-major histocompatibility complex (non-MHC) restricted manner. As an important part of the innate immune system, NK cells constitute approximately 10% of the total circulating lymphocytes in the human body.
[0098] Because of their ability to kill other cells, NK cells are normally kept under tight control. All normal cells in the body express MHC class I molecules on their surface. These molecules act as ligands for many of the receptors found on NK cells, thus protecting normal cells from being killed by NK cells. Cells that do not have enough MHC class I on their surface are recognized as "abnormal" by NK cells and are killed. As well as killing abnormal cells, NK cells also initiate a cytokine response.
[0099] Natural killer cells constitute an emergency response force against cancer and viral infections. These specialized white blood cells originate from the bone marrow, circulate in the blood, and collect in the spleen and other lymphoid tissues. NK cells match their activity to a subset of HLA proteins present on the surface of healthy cells but released by cells weakened by viruses and by cancer. When NK cells encounter cells that lack HLA proteins, they attack and destroy them - thus preventing them from further spreading the virus or cancer. NK cells are distinguished from other immune system cells by the rapidity and breadth of their defensive response. Other white blood cells generally get to work more slowly, targeting specific pathogens - cancer, viruses or bacteria - rather than damaged cells.
[0100] The natural killer cell immunoglobulin-like receptor (KIR) gene family is one of several families of receptors that code for important proteins found on the surface of NK cells. A subset of KIR genes, the inhibitory KIRs, interact with HLA class I molecules encoded within the human MHC. Such interactions allow NK cells to communicate with other cells of the body, including normal, virally infected, or cancerous cells. This communication between KIR molecules on NK cells and HLA class I molecules on all other cells helps determine whether a cell in the body is recognized by NK cells as self or non-self. Cells that are deemed to be "non-self" are targeted for killing by NK cells.
[0101] The KIR gene family consists of 16 genes (KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, KIR3DL3, KIR2DP1 and KIR3DP1). The KIR gene cluster is located within a 100-200 kb region of the leukocyte receptor complex (LRC) located on chromosome 19 (19q13.4). The gene complex is thought to have arisen by a gene duplication event that occurred after the evolutionary divergence of mammals and rodents. The KIR genes are arranged in a head-to-tail fashion, with only 2.4 kb of sequence separating the genes, except for one 14 kb sequence between 3DP1 and 2DL4. Because KIR genes arise by gene duplication, they are highly similar in sequence, exhibiting 90-95% identity with each other. Human individuals differ in the number and type of KIR genes they inherit, and KIR genotypes within individuals and ethnic groups can be quite different. At the chromosomal level, two distinct types of KIR haplotypes exist (see Martin et al. Immunogenetics. (2008) December; 60(12):767-774). The A-haplotype does not contain the stimulatory genes other than 2DS4 (2DS and 3DS1), the 2DL5 gene, and the 2DL2 gene. B-haplotypes are more variable in gene content, with different B-haplotypes containing different numbers of stimulatory genes, either one or two 2DL5 genes, etc. (Martin MP, et al. (2008) KIR haplotypes defined by segregation analysis in 59 Centre d'Etude Polymorphisme Humain (CEPH) families. Immunogenetics, December; 60(12):767-774.).
[0102] All KIR proteins are anchored to the cell membrane and have either two or three extracellular immunoglobulin-like domains and a cytoplasmic tail. Nine KIR genes (KIR2DL and KIR3DL) encode proteins with long cytoplasmic tails that contain immunotyrosine-based inhibitory motifs (ITIMs). These KIR proteins can send inhibitory signals to natural killer cells when the extracellular domains contact their ligands. The remaining KIR genes encode proteins with short cytoplasmic tails. These proteins send activating signals through adaptor molecules such as DAP12.
[0103] The KIR receptor structure and the identity of the HLA class I ligand for each KIR receptor are described in Parham P. et al., Alloreactive killer cells: hindrance and help for hematopoietic transplants. Nature Rev. Immunology. (2003)3:108-122. The nomenclature of killer cell immunoglobulin-like receptors (KIRs) refers to the number of extracellular immunoglobulin-like domains (2D or 3D) and the length of the cytoplasmic tail (L for long, S for short). Each immunoglobulin-like domain is depicted as a loop, each immunoreceptor tyrosine-based inhibitory motif (ITIM) in the cytoplasmic tail is depicted as an oblong shape, and each positively charged residue in the transmembrane region is depicted as a diamond shape.
[0104] The strength of interaction of KIRs with their HLA class I ligands can depend on both KIR and HLA sequences. Although the HLA region has been studied for more than 40 years, KIR molecules were first described (as NKB1) in the mid-1990s (Lanier et al. (1995) The NKB1 and HP-3E4 NK cells receptors are structurally distinct glycoproteins and independently recognize polymorphic HLA-B and HLA-C molecules. J. Immunol. April 1; 154(7):3320-3327 and Litwin V. et al. (1994) NKB1: a natural killer cell receptor involved in the recognition of polymorphic HLA-B molecules. J Exp Med. August 1; 180(2):537-543). The first few years of discovery were mainly spent describing the different KIR genes and developing methods to determine individual KIR genotypes. Using these methods, the association of KIR genes with autoimmune disease and recipient survival after allogeneic hematopoietic cell transplantation has been shown (Parham P. (2005) MHC class I molecules and KIRs in human history, health and survival. Nature reviews, March; 5(3):201-214). It is now clear that each KIR gene has more than one sequence; that is, each KIR gene has variable sequences due to single nucleotide polymorphisms (SNPs) and, in some cases, insertions or deletions within the coding sequence. Studies have shown that KIR3DL1 polymorphisms can affect not only the expression level of KIR3DL1 on natural killer cells but also the binding affinity of KIR3DL1 to its ligands.
[0105] Studies designed to investigate the role of KIRs in human disease have shown associations between various KIR genes and viral infections, e.g., CMV, HCV and HIV, autoimmune diseases, cancer and preeclampsia (Parham P. (2005) MHC class I molecules and KIRs in human history, health and survival. Nature reviews, March; 5(3):201-214). In a recent study of genetic susceptibility to Crohn's disease, an inflammatory autoimmune bowel disease, patients heterozygous for KIR2DL2 and KIR2DL3 and homozygous for C2 ligand were found to be susceptible to the disease, whereas C1 ligand was protective. (Hollenbach JA et al. (2009) Susceptibility to Crohn's Disease is mediated by KIR2DL2 / KIR2DL3 heterozygosity and the HLAC ligand. Immunogenetics. October; 61(10):663-71). Other studies of KIR versus unrelated hematopoietic cell transplantation (HCT) for acute myeloid leukemia (AML) have shown significantly higher 3-year overall survival rates and a 30% overall improvement in risk for relapse-free survival in B / x donors compared with A / A donors.(Cooley S, et al. (2009) Donors with group B KIR haplotypes improve relapse-free survival after unrelated hematopoietic cell transplantation for acute myelogenous leukemia. Blood. January 15; 113(3):726-732; and Miller JS, et al. (2007) Missing KIR-ligands is associated with less relapse and increased graft versus host disease (GVHD) following unrelated donor allogeneic HCT. Blood, 109(11):5058-5061). Such studies have been driven by knowledge of the KIR genotypes of patients and controls, but not at the allele level for KIR. Although it has been well demonstrated that certain HLA alleles are important in human disease (e.g., the association of HLA-DR3 and HLA-DR4 with type I diabetes and the association of HLA-DR8 with juvenile rheumatoid arthritis), studies linking KIRs to human disease would be more sophisticated if KIRs could be genotyped at the allele level.
[0106] In an embodiment, the disclosed method and system can be applied to a nucleic acid sample derived from one or more genes of the MHC region of human chromosome 6. In an embodiment, the nucleic acid sample can be derived from one or more genes of the KIR region of human chromosome 19. In an embodiment, the nucleic acid sample may be derived in whole or in part from one or more genes of the MHC region of human chromosome 6 and one or more genes of the KIR region of human chromosome 19. The disclosed method and related aspects can be used to evaluate essentially any number of genes. In some embodiments, for example, the set of genes analyzed as described herein includes at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 100, 1,000, 10,000, or more genes. A non-exhaustive list of genes, one or more of which are optionally selected for evaluation using the methods and related embodiments disclosed herein, is provided in Table 1. Table 1 [Table 1-1] [Table 1-2] III. EXEMPLARY SYSTEMS AND METHODS
[0107] FIG. 1A is a flow chart that generally depicts an example of a technique for allele typing in a cell-free nucleic acid (cfDNA) sample obtained from a test subject. Allele typing can be used to determine one or more alleles present at a chromosomal locus. As shown, method 100A can include, at step 102A, obtaining data. The data can include sequence data, e.g., allele sequence data. For example, step 102A can include obtaining (or otherwise determining, capturing, receiving, etc.) known allele sequences. For example, known allele sequences can be downloaded from the IPD-IMGT / HLA database and / or the IPD-KIR database. By way of example, known HLA allele sequences can be downloaded from ftp.ebi.ac.uk / pub / databases / ipd / imgt / hla and known KIR allele sequences can be downloaded from ftp.ebi.ac.uk / pub / databases / ipd / kir.
[0108] The known allele sequence can be any type of allele sequence. In some embodiments, the known allele sequence can be a known human leukocyte antigen (HLA) allele sequence. In some embodiments, the known allele sequence can be a known killer cell immunoglobulin-like receptor (KIR) allele sequence.
[0109] At step 104A, the data can be preprocessed. For example, step 104A can include constructing an allele k-mer data structure. The allele k-mer data structure can be a database. The allele k-mer data structure can be a flat file. The allele k-mer data structure can be any form of data structure. Constructing the allele k-mer data structure can include dividing the known allele sequence into k-mer amounts. For example, k-mer amounts having a length of about 100 nucleotides to about 200 nucleotides. In an embodiment, the k-mer amount can have a length of 143 nucleotides. Constructing the allele k-mer data structure can include associating each k-mer with metadata. The metadata can include, for example, an indication of the amount of alleles containing the k-mer, as well as an allele identifier and k-mer start position for each allele containing the k-mer.
[0110] In step 106A, sequence processing can be performed. For example, step 106A can include obtaining (or otherwise determining, acquiring, receiving, etc.) sequence read pairs (e.g., test sequence reads) from a cell-free nucleic acid (cfDNA) sample obtained from a test subject. Step 106A can include aligning the test sequence reads with known allele sequences. For example, step 106A can include aligning the test sequence reads with k-mers in an allele k-mer data structure. Sequence processing can determine an allele(s) supported by the test sequence read(s). An allele may be supported by more than one test sequence read. A test sequence read may support more than one allele. In an embodiment, a test sequence read may be found to support an allele if the test sequence read aligns to an allele (e.g., a k-mer of an allele) above a threshold percent identity. For example, the identity threshold percentage can be 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 100%, etc. In an embodiment, the identity threshold percentage can be 100%, requiring a "perfect" match between the test sequence read and the allele (e.g., the k-mer of the allele).
[0111] In an embodiment, step 106A may include determining the number of test sequence read families supporting the allele(s) (e.g., the number of nucleic acid molecules supporting the allele(s)). Each test sequence read may include a barcode. The barcode may identify the nucleic acid molecule (e.g., test sequence read family) to which the test sequence read is associated. In an embodiment, a test sequence read family may be found to support an allele if the test sequence read family aligns to the allele above a threshold percent identity. For example, the threshold percent identity may be 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, and 100%, etc. In an embodiment, the threshold percent identity may be 100%, requiring a "perfect" match between the test sequence read family and the allele (e.g., the k-mer of the allele).
[0112] In step 108A, a clustering operation can be performed. The alleles can be sorted by the number of supporting test sequence reads (or by the number of supporting test sequence read families), and one or more allele supersets can be constructed. The allele superset can be constructed by determining the first allele associated with the highest number of supporting test sequence reads (or associated with the highest number of supporting test sequence read families). The first allele can be the basis of the allele superset. Additional alleles that can be added to the allele superset when a given allele is associated with a supporting test sequence read (or a supporting test sequence read family) are themselves subsets of the supporting test sequence read (or supporting test sequence read family) of the first allele. One or more additional allele supersets can be constructed in a similar manner using alleles that are not incorporated into the allele superset of the first allele.
[0113] An allele superset can be a data structure. An allele superset can be a database. An allele superset can be a flat file. An allele superset can include a representation of a Hasse diagram. A Hasse diagram is a representation of the relationships of the elements of a partially ordered set where the upward direction is implied. Points, or nodes, can represent each element of the partially ordered set, and the nodes can be connected by line segments according to the following rules: 1) If p < q in the partially ordered set, the point corresponding to p appears lower in the drawing than the point corresponding to q; 2) Two points p and q will be connected by a line segment if p is related to q. In one embodiment, the Hasse diagram can be represented as a graph data structure, such as, for example, a directed acyclic graph (DAG). For example, a DAG including a line from node A to node B where node A strictly contains node B, and node C does not exist, such that node A strictly contains node C and node C strictly contains node B.
[0114] In step 110A, an allele can be classified. For example, in step 110A, the allele type can be determined for a given allele. An allele can be classified based on one or more allele supersets. In one embodiment, if only one allele superset is constructed, the first allele of the superset can be classified as the allele present at a locus of a chromosome (e.g., a haploid locus). In one embodiment, if multiple allele supersets are constructed, the first alleles of two supersets having the cumulative maximum number of distinct supporting test sequence reads (or the cumulative maximum number of distinct supporting test sequence read families) can be classified as the alleles present at a locus of a chromosome (e.g., a diploid locus).
[0115] The classification of the allele(s) can be used to directly treat the subject. It may be unknown in advance whether the subject has the disease, or the subject may be known to have the disease. The disease may be cancer. The method may include administering one or more therapies to the subject to treat the disease. The treatment may include administering immunotherapy, administering chemotherapy, administering radiation therapy, or performing surgery to remove all or a portion of the tumor. The method may include assisting in the communication of the determination of the classification of the allele(s) to the subject associated with the test sample.
[0116] In one embodiment, method 100A can be performed using known HLA allele sequences to determine a subject's HLA genotype, and / or method 100A can be performed using known KIR allele sequences to determine a subject's KIR genotype.
[0117] FIG. 1B is a flow chart that generally illustrates an example of a technique for allele typing and / or variant calling in a cell-free nucleic acid (cfDNA) sample obtained from a test subject. Allele typing can be used to determine one or more alleles present at a chromosomal locus. Variant calling can be used to identify the presence of known or unknown variants. Variant calling can be used to characterize the progression of cancer. As shown, method 100B can include, at step 102B, obtaining data. The data can include sequence data, such as allele sequence data and / or decoy sequence data.
[0118] For example, step 102B may include obtaining (or otherwise determining, acquiring, receiving, etc.) a known allele sequence. For example, the known allele sequence may be downloaded from the IPD-IMGT / HLA database and / or the IPD-KIR database. By way of example, the known HLA allele sequence may be downloaded from ftp.ebi.ac.uk / pub / databases / ipd / imgt / hla and the known KIR allele sequence may be downloaded from ftp.ebi.ac.uk / pub / databases / ipd / kir. The known allele sequence may be any type of allele sequence. In an embodiment, the known allele sequence may be a known human leukocyte antigen (HLA) allele sequence. In an embodiment, the known allele sequence may be a known killer cell immunoglobulin-like receptor (KIR) allele sequence.
[0119] For example, step 102B may include obtaining (or otherwise determining, acquiring, receiving, etc.) a decoy sequence. The decoy sequence may be any type of decoy sequence. In an embodiment, the decoy sequence may be one or more of an "alt" sequence, an unprobed known HLA allele sequence, a probed known HLA allele sequence (to address small germline imprecision, 3 vs. 4 digit pairs), a combination of these, etc. Decoy sequences are sequences of genomic material (generally human) that resemble the sequence we want to interrogate (e.g., the region we want to genotype). These are not yet part of the reference because they encode an alternative form from a region or gene (hence the name "alt"). The problem for us is to deploy targeted sequencing, a method where we select only molecules from parts of the genome that match some designated regions (these "designated regions" are called probes, or baits, in our case they are 120 bases long): what happens is that the probes designed to capture molecules from the region of interest sometimes capture molecules from these "alt" sequences instead. We can detect this because in these cases the reads (or read pairs) align better to the decoy than to the human reference.
[0120] In an embodiment, the decoy sequence may include a decoy sequence selected to identify contamination in the test sample. The one or more decoy sequences may include one or more non-human reference sequences. For example, the one or more decoy sequences may include a bovine reference sequence, a rat reference sequence, a microbial reference sequence, and combinations thereof. Any test sequence pair that aligns to a non-human decoy sequence can be used to support the conclusion that the test sample is contaminated with DNA from a source other than the one being tested. The idea is the same as above, and here we simply use the sequence of the suspected contaminant as a "decoy". For example, suppose we suspect that there is some cow DNA in our sample (this happened!), and therefore add the entire cow genome to our decoy list, if we align the read pair to both the human reference and the cow genome (decoy), we find that the contaminant molecule aligns better to cow than to human.
[0121] At step 104B, the data can be preprocessed. For example, step 104B can include constructing an allele k-mer data structure. The allele k-mer data structure can be a database. The allele k-mer data structure can be a flat file. The allele k-mer data structure can be any form of data structure. Constructing the allele k-mer data structure can include dividing the known allele sequence into k-mer amounts. For example, k-mer amounts having a length of about 100 nucleotides to about 200 nucleotides. In an embodiment, the k-mer amount can have a length of 143 nucleotides. Constructing the allele k-mer data structure can include associating each k-mer with metadata. The metadata can include, for example, an indication of the amount of alleles containing the k-mer, as well as an allele identifier and k-mer start position for each allele containing the k-mer.
[0122] For example, step 104B may include constructing a decoy data structure. The decoy data structure may be a database. The decoy data structure may be a flat file. The decoy data structure may be any form of data structure. By structuring the algorithm in this way (i.e., target sequences plus decoy sequences), some flexibility is preserved. The idea is that we can always add any number of as-yet-unknown "problematic" sequences to the decoy, where problematic in this case means sequences that resemble one of our targets (in other words, sequences that are not the target region, but that we may accidentally pick up by our targeted sequencing tech dev).
[0123] In step 106B, sequence processing can be performed. For example, step 106B can include obtaining (or otherwise determining, acquiring, receiving, etc.) sequence reads (e.g., test sequence reads) from a cell-free nucleic acid (cfDNA) sample obtained from a test subject. Step 106B can include aligning the test sequence reads with known allele sequences. For example, step 106B can include aligning the test sequence reads with k-mers in an allele k-mer data structure. Sequence processing can determine an allele(s) supported by the test sequence read(s). An allele may be supported by more than one test sequence read. A test sequence read may support more than one allele. In an embodiment, a test sequence read may be found to support an allele if the test sequence read aligns to an allele (e.g., a k-mer of an allele) above a threshold percent identity. For example, the identity threshold percentage may be 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% and 100%, etc. In an embodiment, the identity threshold percentage may be 100%, requiring a "perfect" match between the test sequence read and the allele (e.g., the k-mer of the allele), where a "perfect" match indicates no mismatches and no indels. In an embodiment, the identity threshold percentage may be less than 100%, requiring an "imperfect" match between the test sequence read and the allele (e.g., the k-mer of the allele), where an "imperfect" match indicates at least one mismatch and / or at least one indel. An indication of the identity percentage may be determined for each alignment and stored for later processing. The result of the alignment may be represented by an alignment score, which will be described in more detail with respect to the alignment component 215. The alignment score may be equal to the sum of the number of mismatches and the number of indels.
[0124] In an embodiment, step 106B may include determining the number of test sequence read families supporting the allele(s) (e.g., the number of nucleic acid molecules supporting the allele(s)). Each test sequence read may include a barcode. The barcode may identify the nucleic acid molecule (e.g., test sequence read family) to which the test sequence read is associated. In an embodiment, a test sequence read family may be found to support an allele if the test sequence read family aligns to the allele above a threshold percent identity. For example, the threshold percent identity may be 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, and 100%, etc. In an embodiment, the threshold percent identity may be 100%, requiring a "perfect" match between the test sequence read family and the allele (e.g., the k-mer of the allele).
[0125] Step 106B may include performing an alignment between the test sequence read and the decoy sequence. For example, step 106B may include performing an alignment between the test sequence read and the decoy sequence in the decoy data structure. The sequence processing may determine the decoy sequence(s) supported by the test sequence read(s). In an embodiment, if the test sequence read aligns to the decoy sequence above a threshold percent identity, the test sequence read may be found to support the decoy sequence. For example, the threshold percent identity may be 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% and 100%, etc. In an embodiment, the threshold percent identity may be 100%, requiring a "perfect" match between the test sequence read and the decoy sequence, with a "perfect" match indicating no mismatches and no indels. An indication of percent identity may be determined for each alignment and stored for later processing. In some embodiments, one or more test sequence reads that align with one or more decoy sequences with 100% identity may be discarded and not used for further processing.In some embodiments, to the extent that the decoy sequence includes one or more non-human sequences, any test sequence reads that match with 100% identity with the non-human decoy sequence may be used to support the identification of the test sample as contaminated.Notifications related to potential contamination may be generated and / or sent.
[0126] The result of the alignment may be represented by an alignment score, which will be explained in more detail with respect to the alignment component 215. The alignment score may be equal to the sum of the number of mismatches and the number of indels.
[0127] In step 108B, the clustering operation can be performed based on the alignment of the test array reads and the known allele arrays. The known alleles can be sorted by the number of supporting test array reads (or by the number of supporting test array read families), and one or more allele supersets can be constructed. The allele superset can be constructed by determining a first allele associated with the highest number of supporting test array reads (or associated with the highest number of supporting test array read families). The first allele can form the basis of the allele superset. Additional alleles that can be added to the allele superset when a given allele is associated with a supporting test array read (or a supporting test array read family) are themselves a subset of the supporting test array reads (or supporting test array read families) of the first allele. One or more additional allele supersets can be constructed in a similar manner using alleles not incorporated into the allele superset of the first allele.
[0128] The allele superset can be a data structure. The allele superset can be a database. The allele superset can be a flat file. The allele superset can include a representation of a Hasse diagram. A Hasse diagram is a representation of the relationships of the elements of a partially ordered set where an upward direction is implied. Points, or nodes, can represent each element of the partially ordered set, and the nodes can be connected by line segments according to the following rules: 1) If p < q in the partially ordered set, the point corresponding to p appears lower in the drawing than the point corresponding to q; 2) Two points p and q will be connected by a line segment if p is related to q. In one embodiment, the Hasse diagram can be represented as a graph data structure, such as a directed acyclic graph (DAG). For example, a DAG including a line from node A to node B where node A strictly contains node B and there is no node C such that node A strictly contains node C and node C strictly contains node B.
[0129] In step 110B, the alleles can be classified. For example, in step 110B, the allele type can be determined for a given allele. The alleles can be classified based on one or more allele supersets. In an embodiment, if only one allele superset is constructed, the first allele of the superset can be classified as the allele present at the chromosomal locus (e.g., haploid locus). In an embodiment, if multiple allele supersets are constructed, the first allele of the two supersets with the cumulative maximum number of distinct supporting test sequence reads (or the cumulative maximum number of distinct supporting test sequence read families) can be classified as the allele present at the chromosomal locus (e.g., diploid locus).
[0130] The classification of the allele(s) can be used to directly treat the subject. It may be unknown in advance whether the subject has the disease, or it may be known that the subject has the disease. The disease may be cancer. The method may include administering one or more therapies to the subject to treat the disease. The treatment may include administering immunotherapy, administering chemotherapy, administering radiation therapy, or performing surgery to remove all or a portion of the tumor. The method may include assisting in communicating the determination of the classification of the allele(s) to the subject associated with the test sample.
[0131] In one embodiment, method 100B can be performed using known HLA allele sequences to determine a subject's HLA genotype, and / or method 100B can be performed using known KIR allele sequences to determine a subject's KIR genotype.
[0132] In step 112B, the alignment scores of the remaining test sequence reads that did not align with 100% identity to either the known allele sequence or the decoy sequence can be compared. The remaining test sequence reads can have an alignment score associated with the known allele sequence (e.g., a germline alignment score) and an alignment score associated with the decoy sequence (e.g., a decoy alignment score). Test sequence read pairs with a germline alignment score less than the decoy alignment score can be discarded. Test sequence read pairs with a germline alignment score greater than the decoy alignment score can be sent for variant calling in step 110B. If the test sequence read pair aligns to two or more known allele sequences (e.g., aligns to two or more entries in a fasta file), the first entry in the fasta file with the highest alignment score is selected as the alignment.
[0133] At step 110B, test sequence read pairs with germline alignment scores greater than the decoy alignment score can be analyzed to determine and / or identify the test sequence read pairs as variants. Variant calling is the process of identifying true differences between sequence reads of a test sample and a reference sequence. Variant calling can be performed as further described below with respect to the variant caller component 219. In an embodiment, the test sequence read pairs can be identified as somatic variants. In an embodiment, the test sequence read pairs can be identified as variants that are candidate variants associated with somatic events. At step 110B, candidate variants can be identified in the test sequence read pairs. In one embodiment, candidate variants can be identified by comparing the test sequence read pairs to a reference sequence of the target region of a reference genome (e.g., human reference genome hg19). Edges of the test sequence read pairs can be aligned to the reference sequence, and the genomic locations of the mismatched edges and the mismatched nucleotide bases adjacent to the edges can be recorded as the locations of the candidate variants. In some embodiments, the genomic position of the left and right side of the mismatched nucleotide base is recorded as the position of the called variant.In addition, candidate variants can be identified based on the sequencing depth of target region.In particular, higher confidence can be obtained in identifying variants in target regions with greater sequencing depth, for example, because a larger number of sequence reads can help to resolve mismatches or other base pair variations between sequences (e.g., using redundancy).
[0134] In an embodiment, the reference sequence used for variant calling can include one or more reference sequences.One or more reference sequences can be selected to identify contamination in test sample.One or more reference sequences can include one or more non-human reference sequences.For example, one or more reference sequences can include bovine reference sequence, rat reference sequence, microbial reference sequence, and combinations thereof, etc.Any test sequence pair identified as non-human variant can be used to support the conclusion that the test sample is contaminated with DNA from a source other than the test subject.
[0135] Clinical applications of HLA typing and / or variant calling include vaccine testing, disease association, adverse drug reactions, platelet transfusions, and organ and stem cell transplantation.
[0136] In some embodiments, the method includes determining the HLA type of the subject as described herein and determining HLA-DRB1 * 04:01, * 04:04 and * 04:08;HLA-DRB1 * 04:05;HLA-DRB1 * 01:01 and * 01:02;HLA-DRB1 * 14:02;HLA-DRB1 * 10:01; and / or HLA-DRB1 * 01:01, * 04:01, * 04:04 and * and determining that the subject is predisposed to developing rheumatoid arthritis (RA) if one or more of: 04:05 is determined to be present.
[0137] In some embodiments, the method includes determining the HLA type of the subject as described herein and determining HLA-DRB1 * 15:01, HLA-DQB1 * 06:02, HLA-DRB1 * 01:08, HLA-DRB1 *03:01, and / or HLA-DRB1 * If the 13:03 allele is determined to be present, determining that the subject is predisposed to developing multiple sclerosis (MS).
[0138] In some embodiments, the method includes determining the HLA type of the subject as described herein, and detecting HLA-DRB1, HLA-DR2 (DRB1 * 15:01), HLA-DR3(DRB1 * 03:01), HLA-DRB1 * 08:01, and / or HLA-DQA1 * If the 01:02 allele is determined to be present, determining that the subject is predisposed to developing systemic lupus erythematosus (SLE).
[0139] In some embodiments, the method includes determining the HLA type of the subject as described herein and determining HLA-DQB1 * 03:02 and / or DQB1 * If 02:01 is determined to be present, determining that the subject is predisposed to developing type 1 diabetes (T1D).
[0140] In some embodiments, the method includes determining the HLA type of the subject as described herein and * 05:01, DQB1 * 02:01, and DRB1 * If the 03:01 allele is determined to be present, determining that the subject is predisposed to developing Sjogren's syndrome (SS).
[0141] In some embodiments, the method includes determining the HLA type of the subject as described herein and determining whether the subject is HLA-DQ2 (HLA-DQA1 * 05:01-DQB1 * 02:01) and / or HLA-DQ8 (DQA1 * 03:01-DQB1 *and determining that the subject is predisposed to developing celiac disease (CD) if a genomic DNA sequence (encoded by 03:02) is determined to be present.
[0142] In one embodiment, the method includes determining the woman's KIR type as described herein, and if the KIR2DL1 allele is determined to be present, determining that the woman is predisposed to developing preeclampsia.
[0143] In an embodiment, both HLA genotype and KIR genotype can be determined for the same subject.The method can include determining that if a certain KIR allele is found in an individual in combination with a certain HLA allele, the subject is predisposed to eliminate HCV infection, slow progression of HIV infection to AIDS, or Crohn's disease.In an embodiment, the method includes determining the subject's KIR type as described herein, and if it is determined that KIR2DL3 allele exists, determining that the subject may eliminate HCV infection.In some embodiments, the method further includes determining the subject's HLA-C genotype, and if it is determined that the combination of KIR2DL3 allele and HLA-C1 allele exists, determining that the subject may eliminate HCV infection.
[0144] In other embodiments, the method includes determining the subject's KIR type as described herein, and if the KIR3DS1 allele is determined to be present, determining that the subject is less likely to progress from HIV infection to AIDS. In some embodiments, the method further includes determining the subject's HLA genotype, and if the combination of the KIR3DS1 allele and the HLA-Bw4 allele is determined to be present, determining that the subject is less likely to progress from HIV infection to AIDS.
[0145] In yet other embodiments, the method includes determining the subject's KIR type as described herein, and if the KIR2DS1 allele is determined to be present, determining that the subject is predisposed to developing an autoimmune disease.
[0146] In yet other embodiments, the method includes determining the subject's KIR type as described herein, and determining that the subject is predisposed to developing Crohn's disease if KIR2DL2 / KIR2DL3 heterozygosity is determined to be present. In some embodiments, the method further includes determining the subject's HLA genotype, and determining that the subject is predisposed to developing Crohn's disease if a combination of KIR2DL2 / KIR2DL3 heterozygosity and HLA-C2 alleles is determined to be present.
[0147] In yet other embodiments, the method includes determining the subject's KIR type as described herein, and if a Group B KIR haplotype is determined to be present, determining whether the subject is a suitable candidate to be a donor in an unrelated hematopoietic cell transplant.
[0148] FIG. 2 illustrates an example of a system 200 for determining allele types and / or variants of a test subject 211 according to an embodiment of the present disclosure. The system 200 can process one or more samples 201 from the subject 211 to generate sequence reads. The system 200 can include a laboratory system 202, a computer system 210, and / or other components. It should be noted that the laboratory system 202 and the computer system 210 can be remote from each other and connected to each other by a computer network (not shown). The laboratory system 202 can include a sample collection and preparation pipeline 203, a sequencing pipeline 205, a sequence read data store 209, and / or other components. The sequencing pipeline 205 can include one or more sequencing devices 207 (shown in FIG. 2 as sequencing devices 207a...n).
[0149] The disclosed method can be used in a wide variety of ways in the manipulation, preparation, identification, quantification and / or analysis of cell-free nucleic acids. As shown in FIG. 2, a sample collection and preparation pipeline 203 can include obtaining a cfDNA reference sample 201 from one or more reference subjects and a cfDNA test sample 211 from a test subject. As described herein, a polynucleotide can include any type of nucleic acid, such as DNA and / or RNA. For example, if a polynucleotide is DNA, it can be genomic DNA, complementary DNA (cDNA), or any other deoxyribonucleic acid. A polynucleotide can also be a cell-free nucleic acid, such as cell-free DNA (cfDNA). For example, a polynucleotide can be circulating cfDNA. Circulating cfDNA can include DNA released from body cells by apoptosis or necrosis. cfDNA released by apoptosis or necrosis can originate from normal (e.g., healthy) body cells. A sample
[0150] The isolation and extraction of cell-free polynucleotides can be performed by collecting samples using various techniques. The sample can be any biological sample isolated from a subject. The sample can include body tissue, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies (e.g., biopsies from known or suspected solid tumors), cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial or extracellular fluid (e.g., fluid from intercellular spaces), gingival fluid, crevicular fluid, bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a body fluid, particularly blood and its fractions, and urine. The nucleic acid can include DNA and RNA, and can be in double-stranded and single-stranded form. Sample may be in the form that is first isolated from subject, or may be further processed to remove or add components such as cells, enrich one component with respect to another, or convert one form of nucleic acid into another, for example, RNA into DNA, or single-stranded nucleic acid into double-stranded nucleic acid.Thus, for example, the body fluid sample for analysis is plasma or serum that contains cell-free nucleic acid, for example, cell-free DNA (cfDNA).
[0151] In some embodiments, the sample volume of bodily fluid taken from the subject depends on the desired read depth for the region to be sequenced. Exemplary volumes are about 0.4-40 ml, about 5-20 ml, about 10-20 ml. For example, the volume can be about 0.5 ml, about 1 ml, about 5 ml, about 10 ml, about 20 ml, about 30 ml, about 40 ml, or more milliliters. The volume of plasma sampled is typically between about 5 ml and about 20 ml.
[0152] Samples may contain various amounts of nucleic acid. Typically, the amount of nucleic acid in a given sample is equivalent to a multiple of genome equivalent. For example, a sample of about 30ng of DNA may contain about 10,000 (104) haploid human genome equivalents and about 200 billion (2x1011) individual polynucleotide molecules in the case of cfDNA. Similarly, a sample of about 100ng of DNA may contain about 30,000 haploid human genome equivalents and about 600 billion individual molecules in the case of cfDNA.
[0153] In some embodiments, the sample contains nucleic acid from different sources, e.g., from cells and from acellular sources (e.g., blood samples, etc.). In some embodiments, the sample contains nucleic acid from a region of genomic DNA that contains alleles that allow for allelic typing, such as HLA alleles or KIR alleles.
[0154] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification typically range from about 1 femtogram (fg) to about 1 microgram (μg), e.g., from about 1 picogram (pg) to about 200 nanograms (ng), from about 1 ng to about 100 ng, from about 10 ng to about 1000 ng. In some embodiments, the sample contains up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Optionally, the amount is at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In certain embodiments, the amount is up to about 1 fg, about 10 fg, about 100 fg, about 1 pg, about 10 pg, about 100 pg, about 1 ng, about 10 ng, about 100 ng, about 150 ng, or about 200 ng of cell-free nucleic acid molecule. In some embodiments, the method includes obtaining from about 1 fg to about 200 ng of cell-free nucleic acid molecule from the sample.
[0155] The cell-free nucleic acids typically have a size distribution between about 100 nucleotides in length and about 500 nucleotides in length, with molecules between about 110 nucleotides in length and about 230 nucleotides in length representing about 90% of the molecules in the sample, with a mode about 168 nucleotides in length and a second minor peak ranging between about 240 and about 440 nucleotides in length. In certain embodiments, the cell-free nucleic acids are about 160 to about 180 nucleotides in length, or about 320 to about 360 nucleotides in length, or about 440 to about 480 nucleotides in length.
[0156] In some embodiments, the cell-free nucleic acids are isolated from the body fluid by a partitioning step, in which the cell-free nucleic acids are separated from intact cells and other insoluble components in the body fluid, if found in solution. In some of these embodiments, the partitioning includes techniques such as centrifugation or filtration. Alternatively, the cells in the body fluid are lysed, and the cell-free and cellular nucleic acids are processed together. Generally, after the addition of buffer and a washing step, the cell-free nucleic acids are precipitated, for example with alcohol. In certain embodiments, an additional cleaning step is used, such as a silica-based column to remove contaminants or salts. For example, non-specific bulk carrier nucleic acids are added throughout the reaction as needed to optimize certain aspects of the exemplary procedure, such as yield. After such processing, the sample typically contains various forms of nucleic acids, including double-stranded DNA, single-stranded DNA and / or single-stranded RNA. If necessary, the single-stranded DNA and / or single-stranded RNA are converted to double-stranded form in order to include them in subsequent processing and analysis steps. Further details regarding cfDNA partitioning and associated analysis of epigenetic modifications, as appropriate for use in performing the methods disclosed herein, are described, for example, in WO2018 / 119452, filed December 22, 2017, which is incorporated by reference. B. Nucleic acid tag
[0157] In certain embodiments, the tag providing the molecular identifier or barcode is incorporated or otherwise joined to the adapter by chemical synthesis, ligation, or overlap extension PCR, among other methods. In some embodiments, the assignment of the unique or non-unique identifier or molecular barcode in the reaction follows the methods and utilizes the systems described in, for example, US Patent Application Publication Nos. 20010053519, 20030152490, 20110160078, and US Patent Nos. 6,582,908, 7,537,898, and 9,598,731, each of which is incorporated by reference.
[0158] The tags are randomly or randomly linked (e.g., ligated) to the sample nucleic acid. In some embodiments, the tags are introduced into the microwells with an expected ratio of identifiers (e.g., combinations of unique and / or non-unique barcodes). For example, the identifiers can be loaded such that about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or more than 1,000,000,000 identifiers are loaded per genomic sample. In some embodiments, the identifiers are loaded such that less than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000 or 1,000,000,000 identifiers are loaded per genomic sample. In certain embodiments, the average number of identifiers loaded per sample genome is less than or equal to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 500, 1000, 5000, 10000, 50,000, 100,000, 500,000, 1,000,000, 10,000,000, 50,000,000, or 1,000,000,000 identifiers per genome sample. Identifiers are generally unique or non-unique.
[0159] One exemplary format uses about 2 to about 1,000,000 different tags ligated to both ends of a target nucleic acid molecule, or about 5 to about 150 different tags, or about 20 to about 50 different tags. For 20-50 x 20-50 tags, a total of 400-2500 tags are created. Such a number of tags is typically sufficient to ensure a high probability (e.g., at least 94%, 99.5%, 99.99%, 99.999%) that different molecules with the same start and stop points will receive different combinations of tags.
[0160] In some embodiments, the identifier is an oligonucleotide of a predetermined, random or semi-random sequence. In other embodiments, multiple barcodes may be used, where the barcodes in the multiple barcodes are not necessarily unique to each other. In these embodiments, the barcodes are generally attached (e.g., by ligation or PCR amplification) to individual molecules, so that the combination of the barcode and the sequence to which it can be attached creates a unique sequence that can be tracked individually. As described herein, detection of the non-uniquely tagged barcodes in combination with sequence data of the first (start) and last (stop) parts of the sequence read typically allows a unique identity to be assigned to a particular molecule. The length or number of base pairs of the individual sequence reads is also used to assign a unique identity to a given molecule, if necessary. As described herein, fragments from a single strand of nucleic acid, whereby a unique identity is assigned, may allow subsequent identification of fragments from the parental and / or complementary strands.
[0161] In some embodiments, nucleic acid molecules (from a sample of polynucleotides) can be tagged with a sample index and / or a molecular barcode (commonly referred to as a "tag"). The tag can be incorporated or otherwise attached to an adapter by chemical synthesis, ligation (e.g., blunt-end ligation or sticky-end ligation), or overlap extension polymerase chain reaction (PCR), among other methods. Such an adapter can ultimately be attached to a target nucleic acid molecule. In other embodiments, one or more rounds of amplification cycles (e.g., PCR amplification) are typically applied to introduce a sample index into a nucleic acid molecule using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures (e.g., multiple microwells in an array). The molecular barcode and / or sample index can be introduced simultaneously or in any sequential order. In some embodiments, the molecular barcode and / or sample index are introduced before and / or after the sequence capture step is performed. In some embodiments, only the molecular barcode is introduced before the probe capture, and the sample index is introduced after the sequence capture step is performed. In some embodiments, both the molecular barcode and the sample index are introduced before the probe-based capture step is performed. In some embodiments, the sample index is introduced after the sequence capture step is performed. In some embodiments, the molecular barcode is incorporated into the nucleic acid molecule (e.g., cfDNA molecule) in the sample via an adaptor by ligation (e.g., blunt-end ligation or sticky-end ligation). In some embodiments, the sample index is incorporated into the nucleic acid molecule (e.g., cfDNA molecule) in the sample by overlap extension polymerase chain reaction (PCR). Typically, the sequence capture protocol involves introducing a single-stranded nucleic acid molecule complementary to a targeted nucleic acid sequence, e.g., a coding sequence of a genomic region, where mutations in such region are associated with a cancer type.
[0162] In some embodiments, the tag can be located at one end or both ends of the sample nucleic acid molecule.In some embodiments, the tag is an oligonucleotide of predetermined or random or semi-random sequence.In some embodiments, the tag can be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 nucleotide in length.The tag can be randomly or randomly linked to the sample nucleic acid.
[0163] In some embodiments, each sample is uniquely tagged with a sample index, or a combination of sample indexes. In some embodiments, each nucleic acid molecule of a sample or subsample is uniquely tagged with a molecular barcode, or a combination of molecular barcodes. In other embodiments, multiple molecular barcodes (e.g., non-unique molecular barcodes) may be used, such that the molecular barcodes in the multiple molecular barcodes are not necessarily unique to each other. In these embodiments, the molecular barcodes are generally attached (e.g., by ligation) to individual molecules, such that the combination of the molecular barcode and the sequence to which it can be attached creates a unique sequence that can be tracked individually. In one embodiment, techniques for distinguishing true genomic alterations from technical errors can be used, as described in Lee, et al., "Accurate Detection of Rare Mutant Alleles by Target Base-Specific Cleavage with the CRISPR / Cas9 System," ACS Synth. Biol. 2021, 10, 6, 1451-1464, May 19, 2021, which is incorporated by reference in its entirety. Detection of the non-unique molecular barcode in combination with endogenous sequence information (e.g., the first (start) and / or last (stop) genomic location / position corresponding to the sequence of the original nucleic acid molecule in the sample, the start and stop genomic positions corresponding to the sequence of the original nucleic acid molecule in the sample, the first (start) and / or last (stop) genomic location / position of the sequence read that maps to the reference sequence, the start and stop genomic positions of the sequence read that maps to the reference sequence, the subsequence at one or both ends of the sequence read, the length of the sequence read, and / or the length of the original nucleic acid molecule in the sample) typically allows for the assignment of a unique identity to a particular molecule. In some embodiments, the start region includes the first 1, first 2, first 5, first 10, first 15, first 20, first 25, first 30, or at least the first 30 base positions of the 5' end of the sequencing read that aligns to the reference sequence.In some embodiments, the end region includes the last 1, last 2, last 5, last 10, last 15, last 20, last 25, last 30, or at least the last 30 base positions of the 3' end of the sequencing read that aligns to the reference sequence. The length or number of base pairs of each sequence read is also used as necessary to assign a unique identity to a given molecule. As described herein, fragments from a single strand of nucleic acid that are assigned a unique identity may allow subsequent identification of fragments from the parental strand and / or complementary strand.
[0164] In certain embodiments, the number of different tags used to uniquely identify the number of molecules in a class, z, is 2 * z, 3 * z, 4 * z, 5 * z, 6 * z, 7 * z, 8 * z, 9 * z, 10 * z, 11 * z, 12 * z, 13 * z, 14 * z, 15 * z, 16 * z, 17 * z, 18 * z, 19 * z, 20 * z or 100 * Either z (for example, the lower limit) and 100,000 * z, 10,000 * z, 1000 * z or 100 *z can be between any (e.g., an upper limit). In some embodiments, molecular barcodes are introduced to molecules in a sample in an expected ratio of a set of identifiers (e.g., combinations of unique or non-unique molecular barcodes). One exemplary format uses about 2 to about 1,000,000 different molecular barcode sequences, or about 5 to about 150 different molecular barcode sequences, or about 20 to about 50 different molecular barcode sequences, ligated to both ends of a target molecule. Alternatively, about 25 to about 1,000,000 different molecular barcode sequences can be used. For example, 20-50 x 20-50 molecular barcode sequences (i.e., one of 20-50 different molecular barcode sequences can be attached to each end of a target molecule) can be used. Such a number of identifiers is typically sufficient to ensure a high probability (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) that different molecules with the same start and stop points will receive different combinations of identifiers. In some embodiments, about 80%, about 90%, about 95%, or about 99% of the molecules have the same combination of molecular barcodes. C. Nucleic acid amplification
[0165] The adaptor-flanked sample nucleic acids are typically amplified by PCR and other amplification methods that use binding of nucleic acid primers to primer binding sites in the adaptors flanking the DNA molecules to be amplified as part of the sample collection and preparation pipeline 203. In some embodiments, the amplification method involves cycles of extension, denaturation and annealing resulting from thermal cycling, or may be isothermal, for example, as in the case of transcription-mediated amplification. Other exemplary amplification methods that are optionally utilized include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence-based replication, among other approaches.
[0166] To introduce molecular tags and / or sample indexes / tags to nucleic acid molecules using conventional nucleic acid amplification methods, one or more rounds of amplification cycles are generally applied. Amplification is typically performed in one or more reaction mixtures. Molecular tags and sample indexes / tags are introduced simultaneously or in any sequential order, as needed. In some embodiments, molecular tags and sample indexes / tags are introduced before and / or after a sequence capture step is performed. In some embodiments, only molecular tags are introduced before probe capture, and sample indexes / tags are introduced after a sequence capture step is performed. In certain embodiments, both molecular tags and sample indexes / tags are introduced before a probe-based capture step is performed. In some embodiments, sample indexes / tags are introduced after a sequence capture step is performed. Typically, sequence capture protocols include introducing a single-stranded nucleic acid molecule complementary to a targeted nucleic acid sequence, e.g., a coding sequence of a genomic region, where mutations in such region are associated with a cancer type. Typically, the amplification reaction produces a plurality of non-uniquely or uniquely tagged nucleic acid amplicons having molecular tags and sample indexes / tags ranging in size from about 200 nucleotides (nt) to about 700 nt, 250 nt to about 350 nt, or about 320 nt to about 550 nt. In some embodiments, the amplicons have a size of about 300 nt. In some embodiments, the amplicons have a size of about 500 nt.
[0167] In some embodiments, amplification may be performed before and / or after enrichment. d. Nucleic acid enrichment
[0168] In some embodiments, sequences are enriched prior to sequencing the nucleic acids as part of the sample collection and preparation pipeline 203. Enrichment is performed to specific target regions or non-specifically ("target sequences") as desired. In some embodiments, targeted regions of interest can be enriched using nucleic acid capture probes ("baits") selected for one or more bait set panels using differential tiling and capture schemes. In some embodiments, targeted regions of interest can be enriched using CRISPR-mediated enrichment. Differential tiling and capture schemes generally use different relative concentrations of bait sets to differentially tile (e.g., at different "resolutions") across genomic sections associated with the baits, subject to a set of constraints (e.g., sequencer constraints, e.g., sequencing load, utilization of each bait, etc.), to capture targeted nucleic acids at a desired level for downstream sequencing. These targeted genomic sections of interest include natural or synthetic nucleotide sequences of nucleic acid constructs as desired. In some embodiments, biotin-labeled beads bearing probes for one or more sections of interest can be used to capture target sequences, and optionally, amplification of those sections can be followed to enrich for the regions of interest.
[0169] Sequence capture typically involves the use of oligonucleotide probes that hybridize to target nucleic acid sequences. In certain embodiments, the probe set strategy involves tiling probes across the section of interest. Such probes can be, for example, about 60 to about 120 nucleotides in length. The set can have a depth of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 20x, 50x or more. In general, the effectiveness of sequence capture depends in part on the length of the sequence within the target molecule that is complementary (or nearly complementary) to the sequence of the probe.
[0170] In some embodiments, probes can be designed to be specific for an allele of interest, so that different alleles from the same gene have an equal chance of being captured.
[0171] In some embodiments, enrichment can be followed by amplification (described above). e. Nucleic acid sequencing
[0172] 2, after extraction and isolation of cfDNA from a sample by a sample collection and preparation pipeline 203, the cfDNA can be sequenced by a sequencing pipeline 205 that includes one or more sequencing devices 207. Sample nucleic acid, pre-amplified or not, optionally adapter-flanked, is generally subjected to sequencing. Optionally, the sequencing method or commercially available format may include, for example, Sanger sequencing, high-throughput sequencing, bisulfite sequencing, pyrosequencing, single base synthesis, single molecule sequencing, nanopore-based sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing (NGS), Single Molecule Sequencing by Synthesis (SMSS) (Helicos), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking; sequencing using PacBio, SOLiD, Ion Torrent or nanopore platform. Sequencing reactions can be carried out in a variety of sample processing units, which may include multiple lanes, multiple channels, multiple wells or other means for processing multiple sample sets substantially simultaneously. The sample processing unit may also include multiple sample chambers to allow multiple runs to be processed simultaneously.
[0173] A sequencing reaction can be performed on one or more nucleic acid fragment types or sections that are known to contain the allele of interest. A sequencing reaction can also be performed on any nucleic acid fragment present in the sample. A sequence reaction can provide sequence coverage of at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome. In other cases, the sequence coverage of the genome can be less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome.
[0174] Multiplex sequencing techniques can be used to carry out simultaneous sequencing reactions.In some embodiments, cell-free polynucleotides are sequenced in at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions.In other embodiments, cell-free polynucleotides are sequenced in less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions.Sequencing reactions are typically carried out sequentially or simultaneously.Subsequent data analysis is generally carried out on all or part of sequencing reactions. In some embodiments, data analysis is performed for at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, data analysis may be performed for less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An exemplary read depth is about 1000 to about 50,000 reads per locus (base position).
[0175] In some embodiments, a nucleic acid population is prepared for sequencing by enzymatically forming blunt ends on double-stranded nucleic acids with single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Exemplary enzymes or catalytic fragments thereof that are used as needed include Klenow large fragment and T4 polymerase. In the 5' overhang, the enzyme typically extends the recessed 3' end on the opposite strand until the 3' end overlaps the 5' end to generate a blunt end. In the 3' overhang, the enzyme generally digests from the 3' end to the 5' end of the opposite strand, and may also digest beyond the 5' end. If the digestion proceeds beyond the 5' end of the opposite strand, an enzyme with the same polymerase activity as used for the 5' overhang can fill the gap. The formation of blunt ends on double-stranded nucleic acids, for example, facilitates adapter binding and subsequent amplification.
[0176] In some embodiments, the population of nucleic acids is subjected to additional processing, such as converting single-stranded nucleic acids to double-stranded nucleic acids and / or converting RNA to DNA. These forms of nucleic acids are also optionally ligated to adaptors and amplified.
[0177] The nucleic acid that is subjected to the process of forming blunt ends described above, whether or not it has been previously amplified, and other nucleic acids in the sample as necessary, can be sequenced to produce sequenced nucleic acids. Sequenced nucleic acids can refer to either the sequence of the nucleic acid (i.e., sequence information) or the nucleic acid that has been sequenced. Sequencing can be performed to obtain sequence data of individual nucleic acid molecules in a sample either directly or indirectly from the consensus sequence of the amplification products of individual nucleic acid molecules in a sample.
[0178] In some embodiments, the double-stranded nucleic acid with single-stranded overhangs in the sample after blunt end formation is ligated at both ends to an adapter containing a barcode, and the nucleic acid sequence as well as the in-line barcode introduced by the adapter are determined by sequencing. The blunt-ended DNA molecule is optionally ligated to the blunt end of an at least partially double-stranded adapter (e.g., a Y-shaped or bell-shaped adapter). Alternatively, the sample nucleic acid and the blunt end of the adapter can be tailed with complementary nucleotides to facilitate ligation (e.g., sticky end ligation).
[0179] A nucleic acid sample is typically contacted with a sufficient number of adapters such that the probability that any two copies of the same nucleic acid will receive the same combination of adapter barcodes from adapters ligated to both ends is low (e.g., <1 or 0.1%). The use of adapters in this manner allows the identification of a family of nucleic acid sequences that have the same start and stop points on the reference nucleic acid and are ligated to the same combination of barcodes. Such a family represents the sequence of the amplification product of the nucleic acid in the sample before amplification. When modified by blunt end formation and adapter ligation, the sequences of the family members can be compiled to obtain the consensus nucleotide(s) or complete consensus sequence of the nucleic acid molecules in the original sample. In other words, the nucleotide occupying a specified position of the nucleic acid in the sample is determined to be the consensus of the nucleotides occupying the corresponding positions in the family member sequences. A family may include sequences of one or both strands of a double-stranded nucleic acid. When a family member includes sequences of both strands from a double-stranded nucleic acid, the sequences of one strand are converted to their complementary strands with the aim of compiling all sequences to obtain the consensus nucleotide(s) or sequence. Some families contain only a single member sequence. In this case, this sequence can be considered as the sequence of the nucleic acid in the sample before amplification. Alternatively, families with only a single member sequence can be excluded from subsequent analysis.
[0180] Additional details regarding nucleic acid sequencing, including the formats and applications described herein, can be found in, e.g., Levy et al., Annual Review of Genomics and Human Genetics, 17: 95-115 (2016); Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364:1-11 (2012); Voelkerding et al., Clinical Chem., 55: 641-658 (2009); MacLean et al., Nature Rev. Microbiol., 7: 287-296 (2009); Astier et al., J Am Chem Soc., 128(5):1705-10 (2006), U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, U.S. Patent No. 6,833,246, U.S. Patent No. 7,115,400, U.S. Patent No. 6,969,488, U.S. Patent No. 5,912,148, U.S. Patent No. 6,130,073, U.S. Patent No. 7,169,560, U.S. Patent No. 7,282,337, U.S. Patent No. 7,482,120, U.S. Patent No. 7,501, Nos. 245, 6,818,395, 6,911,345, 7,501,245, 7,329,492, 7,170,050, 7,302,146, 7,313,308, and 7,476,503, each of which is incorporated by reference in its entirety.
[0181] In some embodiments, the primers used for sequencing can be specific to the adapters added to the ends of DNA fragments, or can be specific to the target region of interest.For example, if the target region of interest is an HLA region, primers for specific HLA genes can be used for sequencing.In some embodiments, primers can be designed to cover the HLA region that contains multiple HLA alleles.In some embodiments, primers can be designed to cover the HLA region that contains a single HLA allele. f. Sequencing Panel
[0182] To increase the likelihood of detecting genomic regions of interest, the DNA section to be sequenced can include a panel of genes or genomic sections that include known genomic regions. Selecting a limited section (e.g., a limited panel) for sequencing can reduce the total sequencing required (e.g., the total amount of nucleotides sequenced). The sequencing panel can target multiple different genes or regions, such as HLA alleles or KIR alleles. Alternatively, DNA can be sequenced without a sequencing panel by whole genome sequencing (WGS) or other unbiased sequencing methods. Examples of suitable panels and targets for use in panels can include allele typing regions, such as HLA allele regions and / or KIR allele regions.
[0183] In some embodiments, a panel is selected that targets multiple different genes or genomic regions (such as, for example, HLA and / or KIR genes), so that the panel can be used to HLA-type or KIR-type a subject. The panel can be selected to limit the region to be sequenced to a fixed number of base pairs. The panel can be selected to sequence a desired amount of DNA. The panel can also be selected to achieve a desired sequence read depth. The panel can be selected to achieve a desired sequence read depth or sequence read coverage for the amount of base pairs sequenced. The panel can be selected to achieve a theoretical sensitivity, theoretical specificity, and / or theoretical accuracy for allele typing in a sample.
[0184] Genes included in this panel include, but are not limited to, HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-N, HLA-P, HLA-S, HLA-T, HLA-U, HLA-V, HLA-W, HLA-D. RA, HLA-DRB1, HLA-DRB2, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-DRB6, HLA-DRB7, HLA-DRB8, HLA-DRB9, HLA-DQA1, HLA-DQB1, HLA-DQA2, HLA-DQB2, HLA-DQB3, HLA-DO The HLA or KIR genes may include one or more of the HLA or KIR genes, such as HLA-A, HLA-DOB, HLA-DMA, HLA-DMB, HLA-DPA1, HLA-DPB1, HLA-DPA2, HLA-DPB2, HLA-DPA3, HFE, TAP1, TAP2, PSMB9, PSMB8, MICB, MICA, MICC, MICD, MICE, KIR2DL1, KIR2DL2 / L3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, and KIR3DL3.
[0185] Probes for detecting panels of regions can include nucleosome recognition probes (e.g., KRAS codons 12 and 13) as well as those for detecting genomic regions of interest (allele typing regions), and such probes can be designed to optimize capture based on analysis of cfDNA coverage and fragment size variations affected by nucleosome binding patterns and GC sequence composition.
[0186] Some examples of genomic locations of interest may be chromosomal locations 6p21 or 19q13. In some embodiments, the genomic locations used in the methods of the present disclosure include at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 of the genes in Table 1.
[0187] In some embodiments, one or more regions in the panel include one or more loci from one or more genes for allele typing. Genomic locations can be selected for inclusion in a sequencing panel based on the presence of genes of interest for allele typing, such as HLA typing and KIR typing.
[0188] The gene included in the panel for sequencing may include the complete transcription region, promoter region, enhancer region, regulatory element, and / or downstream sequence. In some embodiments, only exons may be included in the panel. The panel may include all exons of the selected gene, or may include only one or more of the exons of the selected gene. The panel may include exons from each of a plurality of different genes. The panel may include at least one exon from each of a plurality of different genes.
[0189] At least one full exon from each different gene in the panel of genes can be sequenced. In some embodiments, all of the exons of a gene can be sequenced. The panel that is sequenced can include all or some exons from multiple genes. The panel can include exons from 2-100 different genes, from 2-70 genes, from 2-50 genes, from 2-30 genes, from 2-15 genes, or from 2-10 genes.
[0190] The panel selected may include a varying number of exons. In some embodiments, the panel selected may include all of the exons of the gene. The panel may include 2-3000 exons. The panel may include 2-1000 exons. The panel may include 2-500 exons. The panel may include 2-100 exons. The panel may include 2-50 exons. The panel may include 2-50 exons. The panel may include 300 or fewer exons. The panel may include 200 or fewer exons. The panel may include 100 or fewer exons. The panel may include 50 or fewer exons. The panel may include 40 or fewer exons. The panel may include 30 or fewer exons. The panel may include 25 or fewer exons. The panel may include 20 or fewer exons. The panel may include 15 or fewer exons. The panel may include 10 or fewer exons. The panel may include 9 or fewer exons. The panel may include up to 8 exons. The panel may include up to 7 exons.
[0191] A panel may include one or more exons from a plurality of different genes. A panel may include one or more exons from each of a percentage of a plurality of different genes. A panel may include at least two exons from each of at least 25%, 50%, 75% or 90% of the different genes. A panel may include at least three exons from each of at least 25%, 50%, 75% or 90% of the different genes. A panel may include at least four exons from each of at least 25%, 50%, 75% or 90% of the different genes.
[0192] The size of a sequencing panel can vary. For example, a sequencing panel can be larger or smaller (in terms of nucleotide size) depending on several factors including the total amount of nucleotides sequenced for a particular region within the panel or the number of unique molecules sequenced. A sequencing panel can be 5 kb to 50 kb in size. A sequencing panel can be 10 kb to 30 kb in size. A sequencing panel can be 12 kb to 20 kb in size. A sequencing panel can be 12 kb to 60 kb in size. A sequencing panel can be 50 kb to 10 Mb in size. A sequencing panel can be 500 kb to 5 Mb in size. A sequencing panel may be at least 10kb, 12kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, 50kb, 60kb, 70kb, 80kb, 90kb, 100kb, 110kb, 120kb, 130kb, 140kb, 150kb, 200kb, 250kb, 300kb, 350kb, 400kb, 450kb or 500kb in size. A sequencing panel may be less than 100kb, 90kb, 80kb, 70kb, 60kb or 50kb in size. Sequencing panels can be at least 1 Mb, 2 Mb, 3 Mb, 4 Mb, 5 Mb, 6 Mb, 7 Mb, 8 Mb, 9 Mb or 10 Mb in size.
[0193] The panel selected for sequencing may include at least 1, 5, 10, 15, 20, 25, 30, 40, 50, 60, 80, or 100 genomic locations (e.g., each includes a genomic region of interest). In some cases, genomic locations within the panel are selected with location sizes that are relatively small. In some cases, the regions within the panel have a size of about 10 kb or less, about 8 kb or less, about 6 kb or less, about 5 kb or less, about 4 kb or less, about 3 kb or less, about 2.5 kb or less, about 2 kb or less, about 1.5 kb or less, or about 1 kb or less. In some cases, genomic locations within a panel have a size of about 0.5 kb to about 10 kb, about 0.5 kb to about 6 kb, about 1 kb to about 11 kb, about 1 kb to about 15 kb, about 1 kb to about 20 kb, about 0.1 kb to about 10 kb, or about 0.2 kb to about 1 kb. For example, a region within a panel may have a size of about 0.1 kb to about 5 kb.
[0194] A panel may include one or more locations that include a genomic region of interest from each of one or more genes. In some cases, a panel may include one or more locations that include a genomic region of interest from each of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, a panel may include one or more locations that include a genomic region of interest from each of at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, or 80 genes. In some cases, a panel may include one or more locations that include a genomic region of interest from each of about 1 to about 80, about 1 to about 50, about 3 to about 40, about 5 to about 30, about 10 to about 20 different genes.
[0195] The concentration of the probe or bait used in the panel can be increased (from 2 to 6 ng / μL) to capture more nucleic acid molecules in the sample. The concentration of the probe or bait used in the panel can be at least 2 ng / μL, 3 ng / μL, 4 ng / μL, 5 ng / μL, 6 ng / μL, or higher. The concentration of the probe can be about 2 ng / μL to about 3 ng / μL, about 2 ng / μL to about 4 ng / μL, about 2 ng / μL to about 5 ng / μL, about 2 ng / μL to about 6 ng / μL. The concentration of the probe or bait used in the panel can be from 2 ng / μL or higher to about 6 ng / μL or less. In some cases, this can allow more molecules in the biologic to be analyzed, thereby allowing detection of less frequent alleles.
[0196] In an embodiment, a sequencing pipeline 205 can be utilized to subject the panel to one or more of: whole genome bisulfite sequencing (WGBS) to interrogate genome-wide methylation patterns; whole genome sequencing (WGS); and / or targeted sequencing approaches to interrogate copy number variants (CNVs) and single nucleotide variants (SNVs). g. Sequence analysis pipeline
[0197] In an embodiment, after sequencing, the sequence reads and any associated data can be stored in a sequence data store 209. The sequence reads can be stored in any format. The sequence data store 209 can be local and / or remote to where the sequencing is performed. As shown in FIG. 2, the stored reads can be subjected to a sequence analysis pipeline 212. i. Sequence quality control
[0198] The sequence analysis pipeline 212 may include a sequence quality control (QC) component 213 that may filter sequence reads from the laboratory system 102. The sequence QC component 213 may assign a quality score to one or more sequence reads. The quality score may be a representation of the sequence reads that indicates whether these sequence reads may be useful for subsequent analysis based on a threshold. In some cases, some sequence reads are not of sufficient quality or length to perform a subsequent mapping step. Sequence reads having a quality score of at least 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, 99.99% or 99.999% may be filtered out from the dataset of sequence reads. In other cases, sequence reads assigned a quality score of at least 90%, 95%, 99%, 99.9%, 99.99% or 99.999% may be filtered out from the dataset.
[0199] Sequence reads that meet a specified quality score threshold can be mapped to the reference genome by the sequence QC component 213. After mapping alignment, sequence reads can be assigned a mapping score. The mapping score can be a representation of the sequence reads that are mapped back to the reference sequence, indicating whether each position is uniquely mappable or not. Sequence reads with a mapping score of at least 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, 99.99% or 99.999% can be filtered out of the data set. In other cases, sequence reads that are assigned a mapping score of less than 90%, 95%, 99%, 99.9%, 99.99% or 99.999% can be filtered out of the data set. ii. Preprocessor
[0200] The pre-processor 214 can import / receive data from the analysis data store 218. For example, the pre-processor 214 can import / receive data representing a plurality of known allele sequences, a plurality of test sequence reads, and / or a plurality of decoy sequences. The pre-processor 214 can also be configured to import sequence data from another source (e.g., an external source). For example, the pre-processor 214 can be configured to download a plurality of known allele sequences, for example, from the IPD-IMGT / HLA database and / or the IPD-KIR database.
[0201] The preprocessor 214 can be configured to split the known allele sequence into multiple k-mer sequences. In other examples, k can be about 25 to about 250. For example, k can be 135 or 140. In certain embodiments, k can be 125 to 175 nucleotides in length, 130 to 160 nucleotides in length, 135 to 155 nucleotides in length, 140 to 150 nucleotides in length. In certain embodiments, k can be 140, 141, 142, 143, 144, or 145 nucleotides in length.
[0202] The pre-processor 214 can create a database including the k-mer sequences and the additional data. The pre-processor 214 can create a data structure including the k-mer sequences and the additional data. The data structure can be, for example, a table or a flat file. FIG. 3 shows an example data structure 300. The data structure 300 can include an entry for each k-mer sequence. For each k-mer sequence, the data structure 300 can indicate the number of alleles to which the k-mer is associated, and for each of these alleles, the allele identifier and starting position where the k-mer sequence is found for that allele. The pre-processor 214 can create a database including the decoy sequences and the additional data. The pre-processor 214 can create a data structure including the decoy sequences and the additional data. The data structure can be, for example, a table or a flat file. iii. Alignment components
[0203] The alignment component 215 can import / receive data from the analysis data store 218. For example, the alignment component 215 can import / receive data representing a plurality of known allele sequences, k-mer sequences generated from a plurality of known allele sequences, a plurality of test sequence reads, and / or a plurality of decoy sequences.
[0204] In various embodiments, the alignment component 215 can be configured to align the test sequence read to a reference sequence or another test sequence read. The alignment component 215 can be configured to align the test sequence read to one or more k-mer sequences generated from multiple known allele sequences. The alignment component 215 can be configured to align the test sequence read (e.g., a pair) to one or more decoy sequences.
[0205] An alignment score is a score that indicates the similarity of two sequences determined using an alignment method. In some implementations, the alignment score accounts for the number of edits (e.g., deletions, insertions, and substitutions of characters in a string). In some implementations, the alignment score accounts for the number of matches. In some implementations, the alignment score accounts for both the number of matches and the number of edits. In some implementations, the alignment score is weighted equally by the number of matches and edits. For example, the alignment score can be calculated as: number of matches-number of insertions-number of deletions-number of substitutions. In other implementations, the number of matches and edits can be weighted differently. For example, the alignment score can be calculated as: number of matches x 5-number of insertions x 4-number of deletions x 4-number of substitutions x 6.
[0206] Pairwise alignment generally involves placing one sequence along the target portion, introducing gaps according to an algorithm, scoring the degree to which the two sequences match well, and preferably repeating for various positions along the reference. The highest scoring match is considered to be an alignment and represents an inference of homology between the aligned portions of the sequences. In some embodiments, scoring the alignment of a nucleic acid sequence pair involves setting values for the scores of substitutions and indels. When individual bases are aligned, a match or mismatch contributes to the alignment score by a substitution probability, which may be, for example, 1 for a match and -0.33 for a mismatch. Indels are subtracted from the alignment score by a gap penalty, which may be, for example, -1. The gap penalty and substitution probability may be based on empirical knowledge or inferential assumptions about the degree to which sequences evolve. Their values affect the resulting alignment. In particular, the relationship between the gap penalty and the substitution probability affects whether a substitution or indel is favored in the resulting alignment.
[0207] As an example, the alignment component 215 may utilize a Burrows-Wheeler Aligner (BWA). In general, the length of the test sequence read may be substantially shorter than the length of the k-mer sequence generated from the multiple known allele sequences. The test sequence read and the k-mer sequence may comprise a sequence of symbols. The alignment of the test sequence read and the k-mer sequence may include a limited number of mismatches between the symbols of the test sequence read and the symbols of the k-mer sequence. In general, the test sequence read may be aligned to a portion of the k-mer sequence to minimize the number of mismatches between the test sequence read and the k-mer sequence.
[0208] In certain embodiments, the symbols of the test sequence reads and k-mer sequences may represent the composition of the biomolecule. For example, the symbols may correspond to the identity of nucleotides in a nucleic acid such as RNA or DNA. In some embodiments, the symbols may directly correlate to these subcomponents of the biomolecule. For example, each symbol may represent a single base of a polynucleotide. In other embodiments, each symbol may represent two or more adjacent subcomponents of a biomolecule, such as two adjacent bases of a polynucleotide. In addition, the symbols may represent overlapping sets of adjacent subcomponents, or separate sets of adjacent subcomponents. For example, if each symbol represents two adjacent bases of a polynucleotide, two adjacent symbols representing an overlapping set may correspond to three bases of a polynucleotide sequence, while two adjacent symbols representing separate sets may represent a four-base sequence. Furthermore, the symbols may directly correspond to subcomponents such as nucleotides, or they may correspond to a color call or other indirect measure of the subcomponents. For example, the symbols may correspond to incorporation or non-incorporation for a particular nucleotide flow.
[0209] In an embodiment, the alignment component 215 can be configured to determine test sequence reads that have identical or substantially identical alignment to one or more k-mer sequences.
[0210] Two nucleic acid or polypeptide sequences are said to be "identical" if the sequences of nucleotides or amino acid residues in the two sequences are the same when aligned for maximum correspondence, respectively, as described herein. The term "identical" or percent "identity" in the context of two or more nucleic acid or polypeptide sequences refers to two or more sequences or subsequences that are the same, or have a specified percentage of amino acid residues or nucleotides that are the same, when compared over a comparison window and aligned for maximum correspondence, as measured using one of the sequence comparison algorithms described below or by manual alignment and visual inspection. The term "substantially identical" used in the context of two nucleic acids or polypeptides refers to a sequence that has at least 50% sequence identity with a reference sequence. The percent identity can be any integer between 50% and 100%. Some embodiments include at least: 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% when compared to a reference sequence using a program described herein, e.g., BLAST.
[0211] For sequence comparison, typically, one sequence serves as the reference sequence with which test sequence is compared.When using sequence comparison algorithm, test and reference sequences are input into computer, partial sequence coordinates are designated as necessary, and sequence algorithm program parameters are designated.Default program parameters can be used, or alternative parameters can be designated.Then, sequence comparison algorithm calculates the percent sequence identity of test sequence to reference sequence based on program parameters.
[0212] "Comparison window," as used herein, includes reference to any one segment of a number of contiguous positions selected from the group consisting of 20 to 600, usually about 50 to about 200, more usually about 100 to about 150, within which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned. Methods for aligning sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be carried out, for example, by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the similarity search method of Pearson & Lipman, Proc. Nat'l. Acad. Sci. USA 85:2444 (1988), or by computer implementations of these algorithms (GAP, BESTFIT, FASTA and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.).
[0213] Algorithms suitable for determining percent sequence identity and percent sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al. (1990) J. Mol. Biol. 215: 403-410 and Altschul et al. (1977) Nucleic Acids Res. 25: 3389-3402, respectively. Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information (NCBI) website. This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in a query sequence that match or meet some positive threshold score T when aligned with words of the same length in a database sequence. T is called the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds to initiate searches to find longer HSPs that contain them. Thus, the word hits are extended in both directions along each sequence as far as the cumulative alignment score can be increased. The cumulative score is calculated using the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0) for nucleotide sequences. For amino acid sequences, a scoring matrix is used to calculate the cumulative score. The extension of the word hits in each direction is stopped when the cumulative alignment score falls by X amount from its maximum achieved value; when the cumulative score becomes zero or less due to the accumulation of one or more negatively scoring residue alignments; or when the end of either sequence is reached. The BLAST algorithm parameters W, T and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word size (W) of 28, an expectation (E) of 10, M=1, N=-2, and a comparison of both strands.The BLASTP program for amino acid sequences uses as defaults a word size (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89:10915 (1989)).
[0214] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin & Altschul, Proc. Nat'l. Acad. Sci. USA 90:5873-5787 (1993)). One measure of similarity provided by the BLAST algorithm is the minimum sum probability (P(N)), which is an indication of the probability that a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the minimum sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.01, more preferably less than about 10-5, and most preferably less than about 10-20.
[0215] Nucleic acid or protein sequences that are substantially identical to a reference sequence include "conservatively modified variants". With respect to a particular nucleic acid sequence, conservatively modified variants refer to nucleic acids that encode the same or essentially identical amino acid sequence, or, if the nucleic acid does not encode an amino acid sequence, to an essentially identical sequence. Due to the degeneracy of the genetic code, a large number of functionally identical nucleic acids encode any given protein. For example, the codons GCA, GCC, GCG, and GCU all encode the amino acid alanine. Thus, at any position where alanine is specified by a codon, the codon can be altered to any of the corresponding codons described without altering the encoded polypeptide. Such nucleic acid variations are "silent variants", which are one type of conservatively modified variant. Every nucleic acid sequence herein that encodes a polypeptide also describes every possible silent variant of the nucleic acid. Those skilled in the art will appreciate that each codon in a nucleic acid (except AUG, which is usually the only codon for methionine) can be altered to produce a functionally identical molecule. Accordingly, each silent variation of a nucleic acid that encodes a polypeptide is implicit in each described sequence.
[0216] Once the alignment component 215 has aligned multiple test sequence reads to one or more k-mers generated from known allele sequences, a list of test sequence reads aligned (supported) to the k-mer sequence of that allele may be generated for each allele. In an embodiment, only test sequence reads that align similarly (e.g., no mismatches and no indels) to the k-mer sequence are included in the list. In an embodiment, only test sequence reads that align substantially similarly (e.g., at least one mismatch and / or at least one indel) to the k-mer sequence are included in the list. In an embodiment, the alignment component may discard actual alignments. In an embodiment, a test sequence read may align (similarly or substantially similarly) to multiple alleles. Each test sequence read may be associated with a test sequence read identifier. Thus, for each allele, a list of test sequence read identifiers associated with supporting test sequence reads may be generated. A list of test sequence reads aligned to decoy sequences may also be generated. In some embodiments, only test sequence reads that align to the decoy sequence in a similar manner (e.g., no mismatches and no indels) are included in the list. In some embodiments, only test sequence reads that align to the decoy sequence in a substantially similar manner (e.g., at least one mismatch and / or at least one indel) are included in the list. The alignment component 215 can be configured to discard any test sequence reads that align to the decoy sequence without mismatches and without indels. iv. Superset Components
[0217] The cluster component 216 can import / receive data from the analysis data store 218. For example, the cluster component 216 can import / receive data representing a plurality of known allele sequences, k-mer sequences generated from the plurality of known allele sequences, a plurality of test sequence reads, and results from the alignment component 215.
[0218] In an embodiment, one or more supersets of a plurality of known allele sequences can be computationally generated by constructing one or more graph data structures.The graph data structure can include nodes (also called vertices) that represent known allele sequences, and edges connecting the nodes that indicate that the supporting reads of one node are a subset of the supporting reads of the other node.The graph data structure construction can be parallelized, given the computationally intensive nature of such construction.
[0219] In one embodiment, the graph data structure is stored in a memory subsystem (e.g., FIG. 2, memory 222), which may include pointers to identify the physical location in memory 222 where each vertex is stored. Typically, the nodes in a graph data structure each represent an element in a set, while the edges represent relationships between the elements. Graph data structures may include directed graphs, trees, and / or directed acyclic graphs (DAGs), etc. A directed graph is a graph in which the edges have directions. A tree is a type of directed graph data structure that has a root node and several additional nodes, each of which is either an interior node or a leaf node. The root node and the interior nodes each have one or more "child" nodes, each of which is called the "parent" of its child nodes. Leaf nodes do not have any child nodes. Edges in a tree are conventionally directed from parent to child. In a tree, a node has only one parent. A generalization of trees known as directed acyclic graphs (DAGs) allows nodes to have multiple parents, but does not allow edges to form cycles.
[0220] In an embodiment, the graph data structure (superset) may represent a Hasse diagram. For a given locus of a chromosome, the alleles may be sorted by the number of supporting test sequence reads. The graph data structure may be constructed by determining the first allele associated with the highest number of supporting test sequence reads. The first allele may form the basis (e.g., the highest level node) of the graph data structure. The supporting test sequence reads of the first allele may define a set of supporting test sequence reads. Additional alleles that may be added to the graph data structure when a given allele is associated with a supporting test sequence read are themselves subsets of the set of supporting test sequence reads of the first allele. One or more additional allele supersets may be constructed in a similar manner using alleles that are not incorporated into the allele superset of the first allele. As an example, a given allele may have the highest number of supporting test sequence reads, and each supporting test sequence read may be associated with a test sequence read identifier. In this case, the set may take the form of test sequence read identifiers of the supporting test sequence reads for the first allele. In a basic example, a first allele may be supported by test sequence reads with identifiers "1", "2", "3", and "4". This set of test supported reads can be represented as A={1, 2, 3, 4}. The power set of A, P(A), is the set of all subsets of A. If A={1, 2, 3, 4}, then P(A)={φ, {1}, {2}, {3}, {4}, {1, 2}, {1, 3}, {1, 4}, {2, 3}, {2, 4}, {3, 4}, {1, 2, 3}, {1, 2, 4}, {1, 3, 4}, {2, 3, 4}, {1, 2, 3, 4}}.
[0221] 4 shows a Hasse diagram 400 of P(A), partially ordered by set inclusion, where X→Y means Y⊂X. The root node 402 of the Hasse diagram 400 is a set that contains all of the test support leads 1, 2, 3, and 4. The second level nodes, e.g., node 404, are each associated with a ternary set of test support leads. The third level nodes, e.g., node 406, are each associated with a set containing two test support leads. The fourth level nodes, e.g., node 408, are each associated with a single test support lead. The lowest level nodes is the empty set 410.
[0222] FIG. 5 shows the construction of an example graph data structure 500 as a Hasse diagram to represent the superset. Allele A * 03:01:01:05 represents the known allele sequence with the highest number of supporting test sequence reads. Allele A * 03:01:01:05 may represent the first allele inserted into the graph data structure 500 as node 502. Allele A * 03:01:01:05 may be associated with supporting test sequence reads having test sequence read identifiers {0;1;2;23;1,030;2,340;...}. The remaining known allele sequences may be analyzed to determine which, if any, remaining known allele sequences are associated with supporting test sequence reads having test sequence read identifiers that are subsets of {0;1;2;23;1,030;2,340;...}. In the graph data structure 500, allele A * 03:01:01:07 is associated with a supporting test sequence read having a test sequence read identifier {0;1;2;1,030;...} that is a subset of {0;1;2;23;1,030;2,340;...}, and thus, allele A * 03:01:01:07 is inserted as node 504 connected by an edge to node 502 in graph data structure 500. *03:01:01:01 is associated with a supporting test sequence read having a test sequence read identifier {1;2;2,340;...} that is a subset of {0;1;2;23;1,030;2,340;...}, and thus, allele A * 03:01:01:01 is inserted as node 506 connected by an edge to node 502 in graph data structure 500. * 03:01:01:23 is associated with a supporting test sequence read having a test sequence read identifier {0;2;...} that is a subset of {0;1;2;1,030;...}, and thus, allele A * 03:01:01:23 is inserted as node 508 connected by an edge to node 504 in the graph data structure 500. * 03:01:01:23 is associated with supporting test sequence reads having test sequence read identifiers {0;1;2;...} that are a subset of {0;1;2;1,030;...}, and thus allele A * 03:01:01:23 is inserted as node 510 connected by an edge to node 504 in the graph data structure 500 .
[0223] In the graph data structure 500, allele A * 03:01:01:47 is associated with a supporting test sequence read having a test sequence read identifier {2;2,340;...} that is a subset of {1;2;;2,340...}, and thus allele A * 03:01:01:47 is inserted as node 512 connected by an edge to node 506 in graph data structure 500. * 03:01:01:25 is associated with a supporting test sequence read having a test sequence read identifier {2;...} that is a subset of {0;2;...} and {0;1;2;...}, and thus allele A *03:01:01:25 is inserted as node 514 that is connected by edges to nodes 508 and 510 in the graph data structure 500. * 03:01:01:24 is associated with a supporting test sequence read having a test sequence read identifier {2,340;...} that is a subset of {2;2,340;...}, and thus, allele A * 03:01:01:24 is inserted as node 516 connected by an edge to node 512 in the graph data structure 500 .
[0224] The graph data structure 500 is complete when no further known allele sequences are associated with supporting test sequence reads having test sequence read identifiers that are a subset of node 502. Additional graph data structures can be constructed using the remaining known allele sequences. In an embodiment, once all known allele sequences have been inserted into the graph data structure (superset). The root node (first allele) represents a candidate allele. A candidate allele may be a candidate allele present at the locus.
[0225] In one embodiment, the graph data structure 500 can be traversed to determine known allele sequences associated with a given set of test sequence read identifiers.
[0226] In one embodiment, adjacency techniques are used to store a graph data structure (e.g., representing a superset) in a memory subsystem (e.g., FIG. 2, memory 222), which may include pointers to identify the physical location within memory 222 where each vertex is stored. In one embodiment, the graph data structure is stored in memory 222 using adjacency lists. In some embodiments, there is an adjacency list for each vertex.
[0227] FIG. 6 shows a graph data structure 600, which includes vertex objects 605 and edge objects 609. The known allele sequences and associated identifiers of the supporting test sequence reads are identified as blocks, and these blocks are converted into objects 605 stored in a tangible memory device. The objects 605 are connected to create paths such that there is a path for each of the known allele sequences and associated subsets of identifiers of the supporting test sequence reads. The paths can represent the known allele sequences associated with the subset of identifiers of the supporting test sequence reads. The connections that create the paths can themselves be implemented as objects, such that blocks are represented by vertex objects 605 and connections are represented by edge objects 609. In this manner, the directed graph includes vertex and edge objects stored in a tangible memory device. As used herein, node objects 605 and edge objects 609 refer to objects created using a computer system.
[0228] 6 further illustrates an adjacency list 610 for the vertex 605. The disclosed method and system can use a processor to create a graph data structure 600 that includes vertex objects 605 and edge objects 609 through the use of adjacencies, e.g., adjacency lists or index-free adjacency. Thus, the processor can use index-free adjacency to create a graph data structure 600 in which a vertex 605 includes a pointer to another vertex 605 to which it is connected, the pointer identifying a physical location on the memory device 222 where the connected vertex is stored. The graph data structure 600 can be implemented using adjacency lists such that each vertex or edge stores a list of such objects to which it is adjacent. Each adjacency list includes pointers to specific physical locations in the memory device for the adjacent objects.
[0229] The graph data structure 600 is typically stored on the physical device of the memory subsystem 222 in a manner that provides very fast traversal. In that sense, the bottom portion of Figure 6 depicts objects being stored in specific physical locations on the tangible portion of the memory subsystem 222. Each node 605 is stored in a physical location, and this location is referenced by a pointer in any adjacency list 610 that references that node. Each node 605 has an adjacency list 610 that contains every adjacent node in the graph data structure 600. The entries in the list 610 are pointers to the adjacent nodes.
[0230] In one particular embodiment, there is an adjacency list for each vertex and edge, where the adjacency list for a vertex or edge lists the edges or vertices that the vertex or edge is adjacent to.
[0231] Figure 7 illustrates the use of an adjacency list 710 for each vertex 705 and edge 709. As shown in Figure 7, the disclosed method and system can create a graph data structure 700 using an adjacency list 710 for each vertex and edge, where the adjacency list 710 for a vertex 705 or edge 709 lists the edges or vertices to which that vertex or edge is adjacent. Each entry in the adjacency list 710 is a pointer to an adjacent vertex or edge.
[0232] Each pointer identifies a physical location in the memory subsystem where the adjacent object is stored. In a preferred embodiment, a pointer, or native pointer, can be manipulated as a memory address since it points to a physical location in memory and allows access to the intended data by dereferencing the pointer. That is, a pointer is a reference to data stored somewhere in memory and obtaining that data is a matter of dereferencing the pointer. The feature that separates pointers from other kinds of references is that the value of a pointer is interpreted, at a low or hardware level, as a memory address. Such a graph representation provides a means for fast random access, modification, and retrieval of data.
[0233] In some embodiments, graph object storage is implemented with index-free contiguity, supporting fast random access since every element contains direct pointers to its neighboring elements, thereby eliminating the need for index lookups and making traversal very quick. Index-free contiguity is another example of low-level, or hardware-level, memory references for data retrieval. Specifically, index-free contiguity can be implemented such that the pointers contained within elements are references to physical locations in memory.
[0234] Because technical implementations that use physical memory addressing, such as native pointers, can access and use data in such a lightweight manner without requiring a separate index table or other intervening lookup step, the capabilities of a given computer, such as any modern consumer-grade desktop computer, are expanded to allow full computation of genome-scale graphs (e.g., graph data structure 700 representing a superset of known allele sequences). Thus, by storing graph elements (e.g., nodes and edges) using a library of objects with native pointers, or other implementations that provide index-free adjacency, the capabilities of the technology that provides storage, retrieval, and alignment of sequence information is actually improved, as this uses the computer's physical memory in a special way. v. Allele Caller - Selecting an allele from a superset
[0235] The allele caller 217 can import / receive data from the analysis data store 218. For example, the allele caller 217 can import / receive data representing a plurality of known allele sequences, k-mer sequences generated from a plurality of known allele sequences, a plurality of test sequence reads, results from the alignment component 215, and / or one or more graph data structures (supersets) generated by the cluster component 216.
[0236] The allele caller 217 can be configured to determine the allele type for a given allele. The allele caller 217 can be configured to classify alleles based on one or more graph data structures (supersets).
[0237] In one embodiment, if only one graph data structure (a superset) is constructed (e.g., all known allele sequences are included in a single graph data structure), the allele associated with the root node of the graph data structure (the first allele) can be classified as the allele present at the chromosomal locus (e.g., a haploid locus).
[0238] In an embodiment, when multiple graph data structures (supersets) are constructed, the allele (first allele) associated with the root node of the two supersets having the greatest cumulative number of distinct supporting test sequence reads may be classified as the allele present at the chromosomal locus (e.g., a diploid locus).
[0239] A set operation can be performed on the combination of root nodes to determine the two root nodes with the cumulative maximum number of distinct supported test sequence reads. In one embodiment, a union operation (∪) can be used. As shown in FIG. 8, the cluster configuration component 216 builds three graph data structures, and the resulting root nodes from each graph data structure are shown as root node 802, root node 804, and root node 806. A union operation can be performed between root node 802 and root node 804, between root node 802 and root node 806, and between root node 804 and root node 806. It is contemplated that a root node may be associated with hundreds or thousands of test sequence lead identifiers, but for simplicity of illustration, root node 802 is depicted as associated with a set of nine test sequence lead identifiers comprised in the set {0;1;2;23;1,030;1,223;2,001;2,300;2,340}, root node 804 is depicted as associated with a set of ten test sequence lead identifiers comprised in the set {1;3;5;6;7;21;1,708;2,000;2,001;2,340}, and root node 806 is depicted as associated with a set of eleven test sequence lead identifiers comprised in the set {0;6;7;9;10;11;21;58;1,200;2,000;2,900}. The union operation between node 802 and node 804 results in a set {0;1;2;3;5;6;7;21;23;1,030;1,223;1,708;2,000;2,001;2,300;2,340} containing 16 cumulative distinct test sequence lead identifiers since the union operation discards duplicates. The union operation between node 802 and node 806 results in a set {0;1;2;6;7;9;10;11;21;23;58;1,030;1,200;1,223;2,000;2,001;2,300;2,340;2,900} containing 19 cumulative distinct test sequence lead identifiers since the union operation discards duplicates. The union operation between nodes 804 and 806 results in the set {0;1;3;5;6;7;9;10;11;21;58;1,200;1,708;2,000;2,001;2,340;2,900} containing 17 accumulated distinct test sequence read identifiers since the union operation discards duplicates.The union operation of node 802 and node 806 results in the maximum cumulative number of distinct test sequence read identifiers, and therefore distinct test sequence reads. Based on that, allele caller 217 selects node 802 (DQB1. * 03:02:01:01) and node 806 (DQB1 * A known allele sequence (e.g., a first allele) associated with the locus 03:04:03) can be classified as an allele present at the locus. The allele caller 217 can continue to process the graph data structure / root node for any number of loci. vi. Variant Caller
[0240] The variant caller 219 can import / receive data from the analysis data store 218. For example, the variant caller 219 can import / receive data representing a plurality of sequence reads. The variant caller 219 can import test sequence reads aligned to a decoy sequence and to a known allele with at least one mismatch and / or at least one indel, the test sequence reads having a greater alignment score to the known allele. The test sequence reads can be analyzed to determine one or more variants. The variants can include, for example, single nucleotide variants (SNVs), indels, fusions, and / or copy number variations. Any known technique for variant calling can be used. In an embodiment, nucleotide variations in a sequenced nucleic acid can be determined by comparing the sequenced nucleic acid to a reference sequence. The reference sequence is often a known sequence, for example, a known full or partial genome sequence from a subject (e.g., a full genome sequence of a human subject). The reference sequence can be, for example, hG19 or hG38. The sequenced nucleic acids may represent sequences determined directly for the nucleic acids in the sample, as described above, or may represent a consensus of sequences of the amplification products of such nucleic acids. Comparisons may be made at one or more designated positions on the reference sequence. A subset of sequenced nucleic acids may be identified that includes positions that match designated positions of the reference sequence when the respective sequences are maximally aligned. Within such a subset, the sequenced nucleic acids that include nucleotide variants at designated positions, if any, may be determined; the length of a given cfDNA fragment may be determined based on where its ends (i.e., its 5' and 3' terminal nucleotides) are located in the reference sequence; the offset of the midpoint of the given cfDNA fragment from the midpoint of the genomic region in the cfDNA fragment may be determined; and, if necessary, the ones that include the reference nucleotides, if any, may be determined (i.e., the same ones as in the reference sequence).When the number of sequenced nucleic acids in the subset containing the nucleotide variant exceeds a selected threshold, a variant nucleotide can be called at the designated position. The threshold can be a simple number, such as at least 1, 2, 3, 4, 5, 6, 7, 9, or 10, of sequenced nucleic acids in the subset containing the nucleotide variant, or the threshold can be a ratio, such as at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20, of sequenced nucleic acids in the subset containing the nucleotide variant, among other possibilities. Comparison can be repeated for any designated position of interest in the reference sequence. Sometimes, comparison can be performed for designated positions that occupy at least about 20, 100, 200, or 300 consecutive positions on the reference sequence, such as about 20 to 500, or about 50 to 300 consecutive positions.
[0241] The variant caller 417 can be configured to assign a somatic or germline label to a given variant. Further details regarding classification of variants as somatic or germline, as appropriate for use in practicing the methods disclosed herein, are described, for example, in U.S. Patent Application Publication No. 20210265013A1, filed February 17, 2021, which is incorporated by reference. In an embodiment, the germline status can be verified using one or more germline variant callers. Germline variant callers include HAPLOTYPE CALLER, samtools, and / or freebayes. In an embodiment, the somatic status can be verified using one or more somatic variant callers. Somatic variant callers include SEURAT, STRELKA, and / or MUTECT.
[0242] Any data analyzed, determined, and / or output by the sequence analysis pipeline 212 may be stored in the analysis data store 218. Generally speaking, the processor 220 may implement (and the processor 220 may be programmed with) various components of the sequence analysis pipeline 212, such as the sequence quality control component 213, the pre-processor 214, the alignment component 215, the cluster component 216, the allele caller 217, the variant caller 219, and / or other components. It should be noted that these components of the sequence analysis pipeline 212 may alternatively comprise hardware modules. Although shown separately for convenience, one or more of the various components or instructions, such as the sequence quality control component 213, the pre-processor 214, the alignment component 215, the cluster component 216, the allele caller 217, and / or the variant caller 219, may be integrated with one another.
[0243] Computer system 210 can communicate data with computer system 224 using network 223. For example, computer system 224 can retrieve data from analytical data store 218. Computer system 224 can be configured for determining / classifying the alleles present at a genetic locus. h. Method example
[0244] In an embodiment shown in Figure 9, a method 900 for allele calling is disclosed. In an embodiment, the sequence QC component 213, the preprocessor 214, the alignment component 215, the cluster component 216, the allele caller 217, the variant caller 219, and additional components not shown (e.g., components of the computer system 210), alone and / or in combination, can be configured to access the sequence data store 209 and / or the analysis data store 218 to perform all and / or a portion of the method 900. All or a portion of the method 900 can be performed by a single computing device using parallel computing, on one or more cores, or using multiple computing devices, such as, for example, a distributed computer system.
[0245] The method 900 may include, at 901, determining a plurality of known allele sequences. The plurality of known allele sequences may include a plurality of known human leukocyte antigen (HLA) allele sequences, or a plurality of known killer cell immunoglobulin-like receptor (KIR) region allele sequences. The step of determining the plurality of known allele sequences may include importing sequence data for the known alleles, determining a plurality of k-mer sequences based on the sequence data for the known alleles, and assembling the plurality of k-mer sequences into a data structure including a plurality of entries, where each entry may include the k-mer sequence, a number of alleles that contain the k-mer sequence, an allele identifier for each allele that contains the k-mer sequence, and a start and stop position of the k-mer sequence in each allele.
[0246] The method 900 may include determining a plurality of sequence reads for a target region of a chromosome of the subject at 902. The target region of the chromosome may include one or more gene loci. The method 900 may further include obtaining a sample from the subject and sequencing the sample to obtain a plurality of sequence reads for the target region of the chromosome. The target region may include one or more of the following genes: HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-N, HLA-P, HLA-S, HLA-T, HLA-U, HLA-V, HL. A-W, HLA-DRA, HLA-DRB1, HLA-DRB2, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-DRB6, HLA-DRB7, HLA-DRB8, HLA-DRB9, HLA-DQA1, HLA-DQB1, HLA-DQA2, HLA-DQB 2, HLA-DQB3, HLA-DOA, HLA-DOB, HLA-DMA, HLA-DMB, HLA-DPA1, HLA-DPB1, HLA-DPA2, HLA-DPB2, HLA-DPA3, HFE, TAP1, TAP2, PSMB9, PSMB8, MICB, MICA, MICC, MICD, MICE, KIR2DL1, KIR2DL2 / L3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, or KIR3DL3. The target region may include chromosomal location 6p21 or 19q13. The target region may include at least a portion of the major histocompatibility complex (MHC) region, and the chromosome is human chromosome 6. The target region can include at least a portion of a killer cell immunoglobulin-like receptor (KIR) region and the chromosome is human chromosome 19.
[0247] The method 900 may include, at 903, aligning a plurality of sequence reads to a plurality of known allele sequences.
[0248] The method 900 may include, at 904, determining, for each known allele sequence of the plurality of known allele sequences, a number of sequence reads aligned to each known allele sequence based on the alignment. In some embodiments, the sequence reads aligned to each known allele sequence may be grouped into one or more sequence read families (i.e., based on molecular barcodes and genomic locations). Based thereon, the number of sequence read families aligned to each known allele sequence may be determined.
[0249] Determining the number of sequence reads aligned to each known allele sequence for each known allele sequence of the plurality of known allele sequences based on the alignment may include determining the number of sequence reads aligned with 100% identity to each known allele sequence. The method 900 may include, at 904, determining the number of sequence read families aligned to each known allele sequence. Determining the number of sequence read families aligned to each known allele sequence may include grouping reads based on barcodes.
[0250] The method 900 may include determining, for the one or more loci, a known allele sequence present at the one or more loci based on a number of sequence reads aligned to each known allele sequence, at 905. The method 900 may include determining, for the one or more loci, a known allele sequence present at the one or more loci based on a number of sequence read families aligned to each known allele sequence, at 905. The known allele sequence present at the one or more loci may include a human leukocyte antigen (HLA) type at the locus or a killer cell immunoglobulin-like receptor (KIR) type at the locus.
[0251] Determining, for the one or more loci, a known allele sequence present at the one or more loci based on the number of sequence reads aligned to each known allele sequence may include determining one or more known allele sequences with the highest number of sequence reads aligned. Determining, for the one or more loci, a known allele sequence present at the one or more loci based on the number of sequence read families aligned to each known allele sequence may include determining one or more known allele sequences with the highest number of sequence read families aligned.
[0252] The method 900 may further include determining, for each read of the plurality of sequence reads, one or more known allele sequences to which each read aligns based on the alignment.
[0253] The method 900 may further include determining a length of a portion of each known allele sequence aligned to two or more sequence reads of the plurality of sequence reads.
[0254] The method 900 may further include sorting, for a locus, the known allele sequences present at the locus by the number of sequence reads aligned to each known allele sequence, determining, for a locus, a first known allele sequence with the highest number of sequence reads aligned, inserting the first known allele sequence with the highest number of sequence reads aligned into a superset, determining one or more known allele sequences aligned to reads that are a subset of the reads aligned to the first known allele sequence, and inserting the one or more known allele sequences into the superset. The superset may include a graph data structure. The graph data structure may include a directed acyclic graph. The graph data structure may represent a Hasse diagram.
[0255] The method 900 may further include determining that the locus is associated with a single superset, and determining a first known allele sequence of the single superset as the allele present at the locus.
[0256] Method 900 may further include determining a plurality of supersets for the locus. Method 900 may further include determining two supersets having a cumulative maximum number of distinct reads based on the plurality of supersets for the locus, and determining a first known allele sequence of each of the two supersets as the allele present at the locus.
[0257] The method 900 may further include a step of assisting in communicating to a health care provider the known allelic sequences present at one or more loci.
[0258] Method 900 may further include obtaining a sample from the cell to be transplanted into the subject, and sequencing the sample to obtain a plurality of sequence reads of the target region of the chromosome. Method 900 may further include determining whether the HLA type of the subject matches the HLA type at the locus. Method 900 may further include transplanting the cell into the subject if the HLA type of the subject matches the HLA type at the locus.
[0259] Method 900 may further include transplanting the cells into the subject if the subject's HLA type matches the HLA type at the locus, and administering an agent that reduces the likelihood of transplant rejection in the subject.
[0260] Method 900 may further include determining that the subject is predisposed to developing rheumatoid arthritis (RA), wherein the known allelic sequences present at the one or more loci include HLA-DRB1 * 04:01, * 04:04, and * 04:08;HLA-DRB1 * 04:05;HLA-DRB1 *01:01 and * 01:02;HLA-DRB1 * 14:02;HLA-DRB1 * 10:01; and HLA-DRB1 * 01:01, * 04:01, * 04:04, and * May include 04:05.
[0261] The method 900 may further include determining that the subject is predisposed to developing multiple sclerosis (MS), wherein the known allelic sequences present at the one or more loci include HLA-DRB1, * 15:01, HLA-DQB1 * 06:02, HLA-DRB1 * 01:08, HLA-DRB1 * 03:01, and / or HLA-DRB1 * May include 13:03.
[0262] Method 900 may further include determining that the subject is predisposed to developing systemic lupus erythematosus (SLE), and the known allelic sequences present at one or more loci include HLA-DRB1, HLA-DR2 (DRB1 * 15:01), HLA-DR3(DRB1 * 03:01), HLA-DRB1 * 08:01, and / or HLA-DQA1 * May include 01:02.
[0263] The method 900 may further include determining that the subject is predisposed to developing type 1 diabetes (T1D), wherein the known allelic sequences present at the one or more loci include HLA-DQB1 * 03:02 and / or DQB1 * May comprise rise 02:01.
[0264] The method 900 may further include determining that the subject is predisposed to developing Sjogren's syndrome (SS), wherein the known allelic sequences present at the one or more loci include DQA1, DQA2, DQA3, DQA4, DQA5, DQA6, DQA7, DQA8, DQA9, DQA10, DQA11, DQA12, DQA13, DQA14, DQA15, DQA16, DQA17, DQA18, DQA19 ...9, DQA19, DQA19, DQA11, DQA12, DQA13, DQA14, DQA15, DQA * 05:01, DQB1 * 02:01, and DRB1 * May include 03:01.
[0265] Method 900 may further include determining that the subject is predisposed to developing celiac disease (CD), wherein the known allele sequences present at the one or more loci include HLA-DQ2 (HLA-DQA1 * 05:01-DQB1 * 02:01) and / or HLA-DQ8 (DQA1 * 03:01-DQB1 * 03:02).
[0266] Method 900 may further include determining that the subject is predisposed to developing preeclampsia, and the known allelic sequences present at the one or more loci may include KIR2DL1.
[0267] Method 900 may further include determining that the subject is unlikely to progress from HIV infection to AIDS, where the known allelic sequences present at the one or more loci may include KIR3DS1.
[0268] Method 900 may further include determining that the subject is unlikely to progress from HIV infection to AIDS, and the known allelic sequences present at the one or more loci may include a combination of KIR3DS1 and HLA-Bw4.
[0269] Method 900 may further include determining that the subject is predisposed to developing an autoimmune disease, and the known allelic sequences present at the one or more loci may include KIR2DS1.
[0270] Method 900 may further include determining that the subject is predisposed to developing Crohn's disease, wherein the known allelic sequences present at the one or more loci may include KIR2DL2 / KIR2DL3 heterozygosity.
[0271] Method 900 may further include determining that the subject is predisposed to developing Crohn's disease, and the known allelic sequences present at the one or more loci may include a combination of KIR2DL2 / KIR2DL3 heterozygosity and an HLA-C2 allele.
[0272] Method 900 may further include a step of determining whether the subject is a suitable candidate for donor in an unrelated hematopoietic cell transplant, and the known allelic sequences present at the one or more loci may include a Group B KIR haplotype.
[0273] In an embodiment shown in Figure 10, a method 1000 for allele and / or variant calling is disclosed. In an embodiment, the sequence QC component 213, the preprocessor 214, the alignment component 215, the cluster component 216, the allele caller 217, the variant caller 219, and additional components not shown (e.g., components of the computer system 210), alone and / or in combination, can be configured to access the sequence data store 209 and / or the analysis data store 218 to perform all and / or a portion of the method 900. All or a portion of the method 900 can be performed by a single computing device using parallel computing, on one or more cores, or using multiple computing devices, such as, for example, a distributed computer system.
[0274] Method 1000 may include, at 1001, determining a plurality of pairs of sequence reads for a target region of a chromosome of a subject, the target region of the chromosome comprising one or more loci. Method 1000 may further include obtaining a sample from the subject, and sequencing the sample to obtain a plurality of pairs of sequence reads for the target region of the chromosome. The target region may include one or more of the following genes: HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-N, HLA-P, HLA-S, HLA-T, HLA-U, HLA-V, HL. A-W, HLA-DRA, HLA-DRB1, HLA-DRB2, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-DRB6, HLA-DRB7, HLA-DRB8, HLA-DRB9, HLA-DQA1, HLA-DQB1, HLA-DQA2, HLA-DQB 2, HLA-DQB3, HLA-DOA, HLA-DOB, HLA-DMA, HLA-DMB, HLA-DPA1, HLA-DPB1, HLA-DPA2, HLA-DPB2, HLA-DPA3, HFE, TAP1, TAP2, PSMB9, PSMB8, MICB, MICA, MICC, MICD, MICE, KIR2DL1, KIR2DL2 / L3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, or KIR3DL3. The target region may include chromosome location 6p21 or chromosome location 19q13. The target region may include at least a portion of a major histocompatibility complex (MHC) region, and the chromosome is human chromosome 6. The target region can include at least a portion of a killer cell immunoglobulin-like receptor (KIR) region and the chromosome is human chromosome 19.
[0275] The method 1000 may include, at 1002, generating germline alignments of a plurality of pairs of sequence reads to a plurality of known allele sequences. The plurality of known allele sequences may include a plurality of known human leukocyte antigen (HLA) allele sequences, or a plurality of known human killer cell immunoglobulin-like receptor (KIR) region allele sequences.
[0276] Method 1000 may include, at 1003, generating a decoy alignment of a plurality of pairs of sequence reads to a plurality of decoy allele sequences. The plurality of decoy sequences may include a plurality of non-human sequences. The plurality of non-human sequences may include one or more of a plurality of bovine sequences, a plurality of rat sequences, or a plurality of microbial sequences.
[0277] The step of generating germline alignments of a plurality of pairs of sequence reads to a plurality of known allele sequences may include determining, for a pair of sequence reads among the plurality of pairs of sequence reads, based on the germline alignment, one or more known allele sequences to which each read of the pair of sequence reads aligns without mismatches or indels.
[0278] Method 1000 may further include determining, for each known allele sequence of the plurality of known allele sequences, a number of pairs of sequence reads among the plurality of pairs of sequence reads that align to each known allele sequence based on the germline alignment, and determining, for one or more loci, known allele sequences present at the one or more loci based on the number of pairs of sequence reads that align to each known allele sequence.
[0279] The step of generating decoy alignments of a plurality of pairs of sequence reads to a plurality of decoy allele sequences may include: determining, for a pair of sequence reads among the plurality of pairs of sequence reads, based on the decoy alignments, one or more decoy allele sequences to which each read of the pair of sequence reads aligns without mismatches or indels, and discarding the pair of sequence reads.
[0280] The step of generating decoy alignments of the plurality of pairs of sequence reads to the plurality of decoy allele sequences may include determining, for a pair of sequence reads among the plurality of pairs of sequence reads, one or more non-human decoy sequences to which each read of the pair of sequence reads aligns without mismatches or indels based on the decoy alignments, and identifying the plurality of pairs of sequence reads as originating from a contaminated sample.
[0281] The step of generating germline alignments of a plurality of pairs of sequence reads to a plurality of known allele sequences may include determining, for a pair of sequence reads among the plurality of pairs of sequence reads based on the germline alignment, one or more known allele sequences to which each read of the pair of sequence reads aligns with at least one mismatch or indel, and generating a germline alignment score.
[0282] The step of generating decoy alignments of a plurality of pairs of sequence reads to a plurality of decoy allele sequences may include determining, for a pair of sequence reads among the plurality of pairs of sequence reads, based on the decoy alignments, one or more decoy allele sequences to which each read of the pair of sequence reads aligns with at least one mismatch or indel, and generating a decoy alignment score.
[0283] The step of generating germline alignments of a plurality of pairs of sequence reads to a plurality of known allele sequences may include determining that the pair of sequence reads aligns to at least two allele sequences of the plurality of known allele sequences, and selecting one known allele sequence of the at least two allele sequences.
[0284] The step of generating decoy alignments of a plurality of pairs of sequence reads to a plurality of decoy allele sequences may include determining that the pair of sequence reads aligns to at least two decoy allele sequences of the plurality of decoy allele sequences, and selecting one decoy allele sequence of the at least two decoy allele sequences.
[0285] The method 1000 may include, at 1004, determining a pair of sequence reads among a plurality of pairs of sequence reads with both a germline alignment having at least one mismatch and / or indel and a decoy alignment having at least one mismatch and / or indel, wherein the pair of sequence reads has a germline alignment score that is greater than the decoy alignment score.
[0286] The method 1000 may include, at 1005, identifying a pair of sequence reads as a candidate somatic variant. Identifying a pair of sequence reads as a candidate somatic variant may include identifying a reference variant of the one or more reference variants as a candidate variant based on aligning the pair of sequence reads to the one or more reference variants. Identifying a pair of sequence reads as a candidate somatic variant may include identifying a non-human reference variant of the one or more non-human reference variants as a candidate variant based on aligning the pair of sequence reads to the one or more non-human reference variants, and identifying a plurality of pairs of sequence reads as originating from a contaminated sample. IV. SYSTEMS AND COMPUTER READABLE MEDIUM
[0287] The various process operations and / or methods depicted in the figures may be performed using some or all of the system components described in detail herein, and in some implementations, various operations may be performed in different orders and various operations may be omitted. Additional operations may be performed along with some or all of the operations shown in the depicted flow charts. One or more operations may be performed simultaneously. Thus, the operations as shown (and described in more detail herein) are provided by way of example and therefore should not be considered limiting.
[0288] The method can be computer implemented such that any or all of the operations described herein or in the appended claims, except for the wet chemistry steps, can be performed on a suitable programmed computer. The computer can be a mainframe, personal computer, tablet, smartphone, cloud, online data storage, or remote data storage, etc. The computer can be operated in one or more locations.
[0289] Various operations of the method may utilize information and / or programs to generate results that may be stored on a computer readable medium (e.g., a hard drive, secondary memory, external memory, a server; a database, a portable memory device (e.g., a CD-R, a DVD, a ZIP disk, a flash memory card), etc.
[0290] The disclosure also includes an article of manufacture for analyzing a nucleic acid population, the article of manufacture including a machine-readable medium containing one or more programs that, when executed, implement the steps of the method.
[0291] The present disclosure can be implemented in hardware and / or software. For example, different aspects of the present disclosure can be implemented in either client-side logic or server-side logic. The present disclosure or components thereof can be embodied in fixed-media program components that contain logic instructions and / or data that, when loaded into a suitably configured computing device, cause the device to perform according to the present disclosure. The fixed media containing the logic instructions can be delivered to the reader using the fixed media for physical loading into the reader's computer, or the fixed media containing the logic instructions can be present on a remote server that the reader accesses by a communication medium to download the program components.
[0292] The present disclosure provides a computer control system programmed to implement the methods of the present disclosure. Returning to FIG. 2, the processor 220 may include a single-core or multi-core processor, or multiple processors for parallel processing. The storage device 222 may include random access memory, read-only memory, flash memory, hard disk, and / or other types of storage devices. The computer system 210 may include a communication interface (e.g., a network adapter) for communicating with one or more other systems and peripheral devices, such as cache, other memory, data storage devices, and / or electronic display adapters. The components of the computer system 210 may communicate with each other by an internal communication bus, such as a motherboard. The storage device 222 may be a data storage unit (or data repository) for storing data. The computer system 210 may be operatively coupled to a network 223 ("network") utilizing a communication interface. The network 223 may be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. The network 223 may be a telecommunications and / or data network in some cases. Network 223 may include a local area network. Network 23 may include one or more computer servers, which may enable distributed computing, such as cloud computing. Network 223 may, in some cases, utilize computer system 210 to implement a peer-to-peer network, which may enable devices coupled to computer system 220 to act as clients or servers. Computer system 210 may exchange data with computer system 224 using network 223. For example, computer system 224 may retrieve data from analysis data store 218.
[0293] The processor 220 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the storage device 222. The instructions may be directed to the processor 220, which may then be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the processor 220 may include fetch, decode, execute, and writeback.
[0294] The processor 220 may be part of a circuit, such as an integrated circuit. One or more other components of the system 200 may be included in the circuit. In some cases, the circuit may include an application specific integrated circuit (ASIC).
[0295] The storage device 222 can store files, such as drivers, libraries, and saved programs. The storage device 222 can store user data, such as user preferences and user programs. The computer system 210 may include one or more additional data storage units that are external to the computer system 210, such as those located on a remote server in communication with the computer system 210 by an intranet or the Internet in some cases.
[0296] The computer system 210 can communicate with one or more remote computer systems over a network. For example, the computer system 210 can communicate with a user's remote computer system. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., Apple® iPad®, Samsung® Galaxy Tab), a phone, a smartphone (e.g., Apple® iPhone®, Android®-enabled device, Blackberry®), or a personal digital assistant. A user can access the computer system 210 over a network.
[0297] The methods described herein may be implemented by machine (e.g., a computer processor) executable code stored in an electronic memory location of the computer system 210, such as in the storage device 222. The machine executable or machine readable code may be provided in the form of software (e.g., a computer readable medium). During use, the code may be executed by the processor 220. In some cases, the code may be retrieved from the storage device 222 and stored in the storage device 222 for ready access by the processor 220.
[0298] The code may be pre-compiled and configured for use on a machine having a processor configured to execute the code, or the code may be compiled at run time. The code may be provided in a programming language that can be selected to allow the code to be executed in a pre-compiled or compiled form.
[0299] Aspects of the systems and methods provided herein, such as computer system 210, can be embodied in programming. Various aspects of the technology can be thought of as "products" or "articles of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code can be stored in an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk.
[0300] A "storage" type medium may include any or all of the tangible memory of a computer or processor, or its associated modules, such as various semiconductor memories, tape drives, and disk drives, etc., that can provide non-transitory storage for software programming at any time. All or parts of the software can sometimes be communicated by the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, such as from a management server or host computer to an application server computer platform. Thus, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as those used by wired and optical telephone networks and through various air links, across physical interfaces between local devices. Physical elements that carry such waves, such as wired or wireless links, or optical links, may also be considered as media that carry software. As used herein, unless limited to non-transitory, tangible, storage media, "media" may include other types of (intangible) media.
[0301] The term "storage" medium, e.g., computer or machine "readable medium," refers to any tangible (e.g., physical), non-transitory, medium that participates in providing instructions to a processor for execution.
[0302] Thus, a machine-readable medium such as a computer executable code may take many forms, including but not limited to tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s), such as those that may be used to implement the databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus in a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, or may take the form of acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards, paper tape, any other physical storage media having a pattern of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, a carrier wave transmitting data or instructions, a cable or link transmitting such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0303] The computer system 210 may include or be in communication with an electronic display 935 that includes a user interface (UI) for providing, for example, reports. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0304] The methods and systems of the present disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software upon execution by the processor 220. EXAMPLES
[0305] V. Working Examples Example 1 Disclosed herein are examples of accurate genotyping of HLA and KIR alleles using cfDNA assays and k-mer-based algorithms for immunotherapy.
[0306] HLA and KIR genotypes hold great promise as novel biomarkers for understanding immune checkpoint inhibitors (ICIs) and patient prognosis. Multiple studies have shown that HLA-I heterozygosity and high sequence diversity across alleles positively correlate with response to ICIs. However, the high degree of polymorphism and allele sequence similarity in HLA and KIR poses challenges for accurate allele calling. To address these challenges, we developed Kmerizer, a novel allele caller optimized for short fragments such as germline reads from cfDNA assays. Methods and Results
[0307] HLA and KIR typing using Kmerizer can include HLA and KIR panel coverage including: HLA loci: A, B, C, DPA1, DPB1, DQA1, DQB1, DRB1, DRB3, DRB4, DRB5, E, F, G; and KIR loci: 2DL1, 2DL2, 2DL3, 2DL4, 2DL5, 2DP1, 2DS1, 2DS2, 2DS3, 2DS4, 2DS5, 3DL1, 3DL2, 3DL3, 3DP1, 3DS1.
[0308] HLA and KIR sequences show high homology levels between the allele sets (Figure 11). Kmerizer calls HLA and KIR alleles in a sample based on a set of known alleles. The module kmerizer can be run in two modes: build index mode or call allele mode (Figure 12). In build index mode, running without inputting the sample bam file will cause the algorithm to generate a kmer map index file. In call allele mode, running with inputting the sample bam file will cause the algorithm to call alleles based on the kmer map index file generated in the "build index" mode. Build index generates a kmer map based on a given kmer size and known allele fasta file. The called alleles are based on the sample bam file and a kmer map is generated as done in build index.
[0309] Kmerizer uses the existing HLA naming system to encode alleles as a tree, where each leaf node represents an allele sequence and each leaf node conveys its count to its parent node. Figure 13 shows an example of the hierarchical structure of allele names. HLA allele names are strings that contain a gene name (A for HLA-A) and up to four pairs of numbers, e.g., A *Encoded by 01:02:03:04. Two allele names are outlined: match the first two pairs of numbers - alleles code for the same protein, but they may have different genomic content at the exon level; match the first three pairs of numbers - alleles have the same genomic content at the exon level (their bases in the exons are identical); match all four pairs of numbers - alleles have the same genomic content in both exons and introns. For KIR, two allele names are outlined: match the first three numbers - alleles code for the same protein; match the first five numbers - alleles have the same genomic content at the exon level (their bases in the exons are identical). The system can be configured to provide the most likely candidate alleles based on a desired solution: for example, classify allele types based on second level yields in peptide degradation; classify allele types based on third / fourth level yields in nucleic acids.
[0310] We validated Kmerizer for HLA typing (Figure 15) and KIR typing (Figure 16) using simulated samples. The simulation randomly picked two alleles per locus from the IMGT database and simulated allele reads from NGS with sequencing errors. In a separate evaluation, the simulation tried Kmerizer with allele pairs that had more than 95% sequence identity. We also performed HISAT-genotyping on the same set of samples for method comparison.
[0311] We also performed typing performance of the kmerizer on gDNA sourced from orthogonally validated cell lines. We evaluated the HLA typing performance of the kmerizer using 19 cell line samples sourced from the Fred Hutch Anthropology Reference Panel (Figure 17), with 15 samples sourced from sonicated gDNA and 4 samples sourced from cell line media. Finally, we evaluated typing concordance of the kmerizer on plasma cfDNA collected from healthy donors compared to buffy coat gDNA.
[0312] For 19 cell lines and 12 plasma samples, Kmerizer yielded 100% sensitivity and 98% specificity for HLA typing (A, B, C, and DQB1 loci). For simulated datasets, Kmerizer achieved 99% sensitivity and specificity for all MHC class 1 and class 2 loci, and 90% sensitivity and specificity for KIR loci, for both homozygous and heterozygous pairs. The allele caller Kmerizer also demonstrated a smaller footprint on computational resource needs: one deep-sequenced plasma sample can be processed in less than 2 minutes, roughly 15 times faster than the most commonly used HLA typing tool, HISAT-Genotype, which does not support KIR typing. conclusion
[0313] The fragment size of cfDNA is shorter than that of genomic DNA, and existing state-of-the-art methods such as HISAT-genotyping are limited in their ability to call alleles accurately.A novel algorithm, Kmerizer, has been developed, which can type HLA and KIR alleles in cfDNA-like fragments.
[0314] The Kmerizer achieves high sensitivity and specificity with both simulated data and well-characterized cell lines, and the Kmerizer achieves high concordance with plasma DNA and buffy coat DNA.
[0315] The simplicity of Kmerizer allows transparency between the sequencing data and the final reported alleles, and it has a fast turnaround time and minimal hardware requirements.
[0316] With the increasing use of ICIs, the use of genetic and genomic information to accurately identify patients more likely to respond to ICIs is of paramount importance. The ability to perform HLA / KIR typing in the cfDNA workflow is critical to providing cost-effective, comprehensive cancer care with liquid biopsies.
[0317] Example 2 HLA germline genotypes and somatic mutations hold great promise as novel biomarkers for immune checkpoint inhibitors (ICIs) and for understanding patient prognosis. Multiple studies have shown that HLA heterozygosity or loss-of-function somatic mutations are negatively correlated with ICI response rates. High HLA fidelity can be used to accurately model patients' neo-antigen landscapes, providing additional prognostic insights for these patients, and identifying those who may benefit specifically from IO treatment.
[0318] Additional data are presented using the methods and systems described herein, called Kmerizer, configured to perform HLA germline typing and / or somatic mutation detection from cfDNA input material. These results can be used for in silico neoantigen and patient outcome prediction. Methods and Results
[0319] Kmerizer can leverage the high depth coverage of targeted sequencing to rapidly identify germline alleles by matching k-mers from input reads with k-mers of known HLA and KIR alleles. Somatic variant calling can be performed following realignment of reads to the called germline.
[0320] For example, netMHC-4.01 can be used to combine MHC-I germline allele calls with patient mutation data to generate in silico TCR binding affinity predictions. These predictions can be compared across cohorts to assess the extent to which cancer types and bTMB differ by predicted neo-antigens and TCR binding affinities.
[0321] For 19 cell lines, 12 plasma samples and 8 gDNA samples with confirmed HLA typing information, Kmerizer yielded high sensitivity and specificity for both MHC-I and MHC-II genes based on methylation epigenomic detection cfDNA sequencing data. High accuracy is achieved for class I and II genes for homozygous and heterozygous states (Table 2). Table 2 shows Kmerizer germline allele calling performance on datasets with confirmed HLA typing information. [Table 2]
[0322] HLA germline allele prevalence among methylated epigenomic detection samples is consistent with a reference cohort of similar geographic origin in MHC class I genes (Figure 18). Figure 18 shows a comparison of HLA germline allele prevalence between methylated epigenomic detection samples (n=2004) and Allele Frequency Net Database (AFND) genotype data from different geographic regions (Europe, n=1000; North America, n=496; Northeast Asia, n=1510; Central and South America, n=1463; Southeast Asia, n=5266). Alleles of particular clinical interest are those highlighted above with "X".
[0323] The counts of germline HLA alleles (ic50 rank < 1%; subset for 9mer long neoantigens) for which immunogenic neoantigens were predicted were plotted across the methylated inter-epigenomic detection development samples. The resulting distributions shown in Figure 19 are largely similar to the overall germline HLA allele distributions for all methylated inter-epigenomic detection samples shown in Figure 20, indicating that there is no substantial bias in neoantigen prediction for a particular set of alleles.
[0324] Notably, we compared the counts of alleles with immunogenic neoantigen predictions supported by netMHC with those without predictions (insert table; for methylation epigenomic detection development samples only) and found a significant association between HLA allele group (HLA-A / B / C) and immunogenicity predictions, with higher immunogenicity predicted for HLA-A alleles than for HLA-B and HLA-C alleles across patients.
[0325] The Kmerizer somatic caller achieved >99.99% specificity per base, calculated for MHC class I genes in 48 normal samples. For sensitivity assessment, simulated sequencing data were generated with randomly selected somatic variants at expected allele frequencies (AF) between 0.15 and 0.2% (Figure 21). Figure 21 shows that germline allele(s) and homozygous / heterozygous status were randomly assigned based on AFND using references from the IMGT / HLA database to generate germline reads, then somatic mutations were randomly generated at exon positions, and then variant reads were generated from the modified references.
[0326] The Kmerizer somatic caller achieved 100% specificity and 97.5% aggregate sensitivity for both variant detection and mutant molecule recovery for class I genes (Table 3). Table 3 shows Kmerizer somatic allele calling performance on simulated data. Detected variants had an observed AF range [0.08%, 2.3%], and a median of 0.18% (Figure 22). Figure 22 shows the variant allele frequency (AF) distribution of detected variants. [Table 3]
[0327] Kmerizer leverages germline HLA calling for neo-antigen prediction. Matched germline allele calling and mutation data were used to generate 53,107 neo-antigens in netMHC-4.0 from 533 clinical cancer samples. The binding affinity of the patients' top-ranked neo-antigen TCRs was strongly associated with bTMB scores determined by a methylation epigenomic detection assay (ρ=-0.524, p<0.0001). Grouping of patients by HLA alleles associated with the strongest binding affinity prediction shows that this association is concordant across genes (Figure 23). Figure 23 shows TMB against predicted neo-epitope affinity across genes. Spearman correlations listed by group.
[0328] In the absence of expression data, samples (N=122) with a mean somatic neo-antigen binding affinity greater than the germline affinity were selected.
[0329] Comparison of binding affinity (ic50) and bTMB for the highest predicted neo-antigen binding affinity per patient shows a strong inverse correlation between TMB score and predicted TCR binding (Rhospearman=-0.524, p<0.0001).
[0330] To understand whether this relationship is driven by neoantigen prediction for only one HLA allele or for several HLA alleles, we subsetted patients by the HLA-allele with the highest neoantigen prediction, recalculated the correlations, and found that the same associations held across HLA genes (Figure 23; HLA-A / B / C; plots on the right).
[0331] Furthermore, it was determined that this association persisted across alleles within HLA genes. Focusing on HLA-A alleles, four alleles were found to be associated with >5% of the highest HLA-A neoantigen predictions: HLA-A * 02:01, HLA-A * 03:01, HLA-A * 11:01, and HLA-A * Plotting samples with neoantigens that may have the strongest binding affinity for these alleles against patient bTMB shows that the association of stronger binding affinity with increasing bTMB holds for all four alleles (RhoHLA-A * 02:01=-0.683, p<0.0001;RhoHLA-A * 03:01=-0.485, p<0.005;RhoHLA-A * 11:01=-0.682, p<0.0005;RhoHLA-A * 02:01=-0.453, p=0.103).
[0332] The proportion of immunogenic neoantigens (ic50 rank <2%) differed significantly across sample cancer types (χ2=12.04, p=0.017; Figure 24, inset). Figure 24 shows the distribution of neoantigen binding affinity by cancer type. Comparison of immunogenic neoantigen binding affinity across cancer types generally identified bladder and melanoma cancer samples as having the highest immunogenic potential, and colorectal cancer as immunologically cold (ks test, p<0.005; Figure 24; stronger binding indicated by smaller ic50 scores).
[0333] The dev samples were filtered to only those with predicted neo-antigen TCR median affinity greater than that of the matched germline TCR median affinity. In the absence of expression data to identify expression levels of mutated peptides, this filtering helps subset patients more likely to have neo-antigens that will stimulate a CD8+ T cell response compared to the background "noise" of germline-produced antigens.
[0334] To determine whether there was a cancer type-specific effect on the proportion of immunogenic neo-antigen targets, all predictions were classified as either immunogenic ("strong binding") or non-immunogenic ("probably no CD8+ binding") based on netMHC percentile rank (rank < 1%). The proportion of immunogenic neo-antigen targets per cancer type was compared by chi-square test, and a significant effect was found (Figure 24; X2 = 12.04, p = 0.017; inset table).
[0335] Cumulative density plots of immunogenic neoantigens by cancer type (Figure 24, right side) demonstrate the immune stimulatory potential of each group (affinity is represented by IC50, meaning lower values represent stronger TCR binding potential). These distributions suggest that in this cohort, melanoma and bladder cancer have higher average immunogenic potential, while colorectal cancer has the lowest potential (most likely to be a "non-inflammatory" tumor). conclusion
[0336] The incorporation of Kmerizer into methylation epigenomic cross detection enables accurate HLA germline and somatic detection and neo-antigen prediction. Kmerizer has demonstrated that HLA germline prevalence correlates with published data. Kmerizer has demonstrated HLA somatic specificity in 48 normal samples and sensitivity in simulations. Kmerizer has demonstrated neo-antigen prediction analysis in clinical samples.
[0337] Non-limiting embodiments are described below:
[0338] Embodiment 1: Determining a plurality of known allele sequences; determining a plurality of sequence reads for a target region of a chromosome of a subject, the target region of the chromosome comprising one or more gene loci; aligning the plurality of sequence reads to a plurality of known allele sequences; determining, for each known allele sequence of the plurality of known allele sequences, a number of sequence reads that align to each known allele sequence based on the alignment; and determining, for one or more loci, the known allele sequences present at the one or more loci based on the number of sequence reads aligned to each known allele sequence. The method includes:
[0339] Embodiment 2: The method of embodiment 1, wherein the plurality of known allele sequences comprises a plurality of known human leukocyte antigen (HLA) allele sequences, or a plurality of known killer cell immunoglobulin-like receptor (KIR) region allele sequences.
[0340] Embodiment 3: Obtaining a sample from a subject; and Sequencing the sample to obtain a plurality of sequence reads for the target region of the chromosome. 2. The method of embodiment 1, further comprising:
[0341] Embodiment 4: Determining, for each read of the plurality of sequence reads, one or more known allele sequences to which each read aligns based on the alignment. 2. The method of embodiment 1, further comprising:
[0342] Embodiment 5: The target region is the following genes: HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-N, HLA-P, HLA-S, HLA-T, HLA-U, HLA-V, HLA-W, HLA-DRA, HL A-DRB1, HLA-DRB2, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-DRB6, HLA-DRB7, HLA-DRB8, HLA-DRB9, HLA-DQA1, HLA-DQB1, HLA-DQA2, HLA-DQB2, HLA-DQB3, HLA-DOA 2. The method of embodiment 1, comprising one or more of the following: HLA-DOB, HLA-DMA, HLA-DMB, HLA-DPA1, HLA-DPB1, HLA-DPA2, HLA-DPB2, HLA-DPA3, HFE, TAP1, TAP2, PSMB9, PSMB8, MICB, MICA, MICC, MICD, MICE, KIR2DL1, KIR2DL2 / L3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, or KIR3DL3.
[0343] Embodiment 6: The method of embodiment 1, wherein the target region comprises chromosomal location 6p21 or chromosomal location 19q13.
[0344] Embodiment 7: The method of embodiment 1, wherein the target region comprises at least a portion of a major histocompatibility complex (MHC) region and the chromosome is human chromosome 6.
[0345] Embodiment 8: The method of embodiment 1, wherein the target region comprises at least a portion of a killer cell immunoglobulin-like receptor (KIR) region and the chromosome is human chromosome 19.
[0346] Embodiment 9: The step of determining the sequences of the plurality of known alleles comprises: Searching sequence data for known alleles; determining multiple k-mer sequences based on sequence data for known alleles; and Assembling the multiple k-mer sequences into a data structure that includes multiple entries, where each entry includes the k-mer sequence, a number of alleles that contain the k-mer sequence, an allele identifier for each allele that contains the k-mer sequence, and a start and stop position of the k-mer sequence in each allele. 2. The method of embodiment 1, comprising:
[0347] Embodiment 10: The method of embodiment 1, wherein for each known allele sequence of the plurality of known allele sequences, the step of determining the number of sequence reads aligned to each known allele sequence based on the alignment comprises determining the number of sequence reads aligned with 100% identity to each known allele sequence.
[0348] Embodiment 11: The method of embodiment 1, wherein for one or more loci, determining known allele sequences present at one or more loci based on the number of sequence reads aligned to each known allele sequence comprises determining one or more known allele sequences having the highest number of aligned sequence reads.
[0349] Embodiment 12: The method of embodiment 1, wherein the reads aligned to each known allele sequence are grouped into read families, and the method further comprises determining the number of sequence read families that are aligned to each known allele sequence.
[0350] Embodiment 13: The method of embodiment 12, wherein for one or more loci, the known allele sequences present at the one or more loci are determined based on the number of sequence read families aligned to each known allele sequence.
[0351] Embodiment 14: The method of embodiment 1, further comprising determining a length of a portion of each known allele sequence aligned to two or more sequence reads of the plurality of sequence reads.
[0352]
[0081] Embodiment 15: For a locus, sorting known allele sequences present at the locus by the number of sequence reads aligned to each known allele sequence; determining a first known allele sequence for the locus that has the highest number of aligned sequence reads; inserting the first known allele sequence having the highest number of aligned sequence reads into the superset; determining one or more known allele sequences aligned to the reads that are a subset of the reads aligned to the first known allele sequence; and Inserting one or more known allele sequences into the superset 2. The method of embodiment 1, further comprising:
[0353] Embodiment 16: The method of embodiment 15, wherein the superset includes a graph data structure.
[0354] Embodiment 17: The method of embodiment 16, wherein the graph data structure comprises a directed acyclic graph.
[0355] Embodiment 18: The method of embodiment 16, wherein the graph data structure represents a Hasse diagram.
[0356] Embodiment 19: Determining that the loci are associated with a single superset; and determining the first known allele sequence of the single superset as the allele present at the locus; 16. The method of embodiment 15, further comprising:
[0357] Embodiment 20: The method of embodiment 15, further comprising determining a plurality of supersets for the locus.
[0358]
[0081] Embodiment 21: Determining two supersets with a cumulative maximum number of distinct reads based on multiple supersets for a locus; and determining a first known allele sequence of each of the two supersets as the allele present at the locus; 21. The method of embodiment 20, further comprising:
[0359] Embodiment 22: The method of embodiment 1, further comprising the step of assisting in communicating to a health care provider the known allelic sequences present at one or more loci.
[0360] Embodiment 23: The method of embodiment 1, wherein the known allelic sequence present at the locus comprises a human leukocyte antigen (HLA) type at the locus or a killer cell immunoglobulin-like receptor (KIR) type at the locus.
[0361] Embodiment 24: Obtaining a sample from cells to be transplanted into a subject; and Sequencing the sample to obtain a plurality of sequence reads for the target region of the chromosome. 2. The method of embodiment 1, further comprising:
[0362] Embodiment 25: The method of embodiment 24, further comprising determining whether the subject's HLA type matches the HLA type at the locus.
[0363] Embodiment 26: The method of embodiment 25, further comprising transplanting the cells into the subject if the subject's HLA type matches the HLA type at the locus.
[0364] Embodiment 27: Transplanting the cells into the subject if the subject's HLA type matches the HLA type at the locus; and Administering an agent that reduces the likelihood of transplant rejection in a subject. 26. The method of embodiment 25, further comprising:
[0365] Embodiment 28: The method further comprises determining that the subject is predisposed to developing rheumatoid arthritis (RA), wherein the known allelic sequences present at one or more loci are HLA-DRB1 * 04:01,* 04:04, and * 04:08;HLA-DRB1 * 04:05;HLA-DRB1 * 01:01 and * 01:02;HLA-DRB1 * 14:02;HLA-DRB1 * 10:01; and HLA-DRB1 * 01:01, * 04:01, * 04:04, and * 2. The method of embodiment 1, comprising:
[0366] Embodiment 29: The method further comprises determining that the subject is predisposed to developing multiple sclerosis (MS), wherein the known allelic sequences present at one or more loci are HLA-DRB1 * 15:01, HLA-DQB1 * 06:02, HLA-DRB1 * 01:08, HLA-DRB1 * 03:01, and / or HLA-DRB1 * 13:03.
[0367]
[0081] Embodiment 30: The method further comprises determining that the subject is predisposed to developing systemic lupus erythematosus (SLE), and determining that the known allelic sequences present at one or more loci are HLA-DRB1, HLA-DR2 (DRB1 * 15:01), HLA-DR3(DRB1 * 03:01), HLA-DRB1 * 08:01, and / or HLA-DQA1 * 2. The method of embodiment 1, comprising:
[0368]
[0081] Embodiment 31: The method further comprises determining that the subject is predisposed to developing type 1 diabetes (T1D), wherein the known allelic sequences present at one or more loci are HLA-DQB1 * 03:02 and / or DQB1 * 2. The method of embodiment 1, comprising:
[0369]
[0081] Embodiment 32: The method further comprises determining that the subject is predisposed to developing Sjogren's syndrome (SS), wherein the known allele sequences present at one or more loci are DQA1 * 05:01, DQB1 * 02:01, and DRB1 * 2. The method of embodiment 1, comprising:
[0370] Embodiment 33: The subject is predisposed to developing celiac disease (CD). The method further comprises determining that the subject is predisposed to developing celiac disease (CD), and wherein the known allele sequences present at one or more loci are HLA-DQ2 (HLA-DQA1 * 05:01-DQB1 * 02:01) and / or HLA-DQ8 (DQA1 * 03:01-DQB1 * 2. The method of embodiment 1, comprising the steps of:
[0371] Embodiment 34: The method of embodiment 1, further comprising determining that the subject is predisposed to developing preeclampsia, wherein the known allelic sequences present at the one or more loci include KIR2DL1.
[0372] Embodiment 35: The method of embodiment 1, further comprising determining that the subject is unlikely to progress from HIV infection to AIDS, wherein the known allelic sequences present at the one or more loci include KIR3DS1.
[0373] Embodiment 36: The method of embodiment 1, further comprising determining that the subject is unlikely to progress from HIV infection to AIDS, wherein the known allelic sequences present at the one or more loci comprise a combination of KIR3DS1 and HLA-Bw4.
[0374] Embodiment 37: The method of embodiment 1, further comprising determining that the subject is predisposed to developing an autoimmune disease, wherein the known allelic sequences present at the one or more loci include KIR2DS1.
[0375] Embodiment 38: The method of embodiment 1, further comprising determining that the subject is predisposed to developing Crohn's disease, wherein the known allelic sequences present at the one or more loci comprise KIR2DL2 / KIR2DL3 heterozygosity.
[0376] Embodiment 39: The method of embodiment 1, further comprising determining that the subject is predisposed to developing Crohn's disease, wherein the known allelic sequences present at the one or more loci include a combination of KIR2DL2 / KIR2DL3 heterozygosity and an HLA-C2 allele.
[0377] Embodiment 40: The method of embodiment 1, further comprising determining whether the subject is a suitable candidate for donor in an unrelated hematopoietic cell transplant, wherein the known allelic sequences present at the one or more loci comprise a KIR haplotype of group B.
[0378] Embodiment 41 : Determining the sequences of multiple known human leukocyte antigen (HLA) alleles; determining a plurality of sequence reads for a target region of a chromosome of a subject, the target region of the chromosome comprising one or more gene loci; aligning the plurality of sequence reads to a plurality of known HLA allele sequences; determining, for each known HLA allele sequence of the plurality of known HLA allele sequences, a number of sequence reads aligned to each known HLA allele sequence based on the alignment; generating one or more supersets of known HLA allele sequences based on the number of sequence reads aligned to each known HLA allele sequence; and determining, for one or more loci, the known HLA allele sequences present at the one or more loci based on the number of distinct reads in one or more supersets of known HLA allele sequences; The method includes:
[0379] Embodiment 42: Determining the sequences of multiple known killer cell immunoglobulin-like receptor (KIR) alleles; determining a plurality of sequence reads for a target region of a chromosome of a subject, the target region of the chromosome comprising one or more gene loci; aligning the plurality of sequence reads to a plurality of known KIR allele sequences; determining, for each known KIR allele sequence of the plurality of known KIR allele sequences based on the alignment, a number of sequence reads that align to each known KIR allele sequence; generating one or more supersets of known KIR allele sequences based on the number of sequence reads aligned to each known KIR allele sequence; and determining, for one or more loci, the known KIR allele sequences present at the one or more loci based on the number of distinct reads in the one or more supersets of known KIR allele sequences. The method includes:
[0380] Embodiment 43: Determining the sequences of a plurality of known alleles; determining a plurality of sequence reads for a target region of a chromosome of a subject, the target region of the chromosome comprising one or more gene loci; aligning the plurality of sequence reads to a plurality of known allele sequences; determining, for each known allele sequence of the plurality of known allele sequences based on the alignment, a number of sequence reads that align to each known allele sequence; generating one or more supersets of the known allele sequences based on the number of sequence reads aligned to each known allele sequence; and determining, for one or more loci, the known allele sequences present at the one or more loci based on the number of distinct reads in the one or more supersets of known allele sequences; The method includes:
[0381] Embodiment 44: One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by a processor, cause the processor to perform the method of any of embodiments 1 to 43.
[0382]
[0081] Embodiment 45: A computing device configured to perform the method of any one of embodiments 1 to 43; an output device configured to output a representation of known allelic sequences present at one or more genetic loci; Including, the system.
[0383] Embodiment 46: One or more processors; A memory storing processor-executable instructions that, when executed by one or more processors, cause the apparatus to perform the method according to any one of embodiments 1 to 43. 13. An apparatus comprising:
[0384] Embodiment 47: A method comprising the steps of determining a plurality of sequence reads or read pairs for a target region of a chromosome of a subject, wherein the target region of the chromosome comprises one or more gene loci; generating an alignment of the plurality of sequence reads or read pairs to a plurality of known sequences; generating an alignment of the plurality of sequence reads or read pairs to a plurality of decoy sequences; comparing the alignment of the sequence reads or read pairs to the known sequences with an alignment of the sequence reads or read pairs to the decoy sequences, wherein if the alignment score of the sequence read or read pair to the known sequence is greater than the alignment score of the sequence read or read pair to the decoy sequence, the sequence read or read pair is associated with the known sequence; and identifying the sequence read or read pair as a supporting sequence read or read pair for the sequence of interest.
[0385] Embodiment 48: The method of embodiment 47, further comprising obtaining a sample from the subject and sequencing the sample to obtain a plurality of sequence reads or read pairs for the target region of the chromosome.
[0386] Embodiment 49: The target region is the following genes: HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-N, HLA-P, HLA-S, HLA-T, HLA-U, HLA-V, HLA-W, HLA-DRA, H LA-DRB1, HLA-DRB2, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-DRB6, HLA-DRB7, HLA-DRB8, HLA-DRB9, HLA-DQA1, HLA-DQB1, HLA-DQA2, HLA-DQB2, HLA-DQB3, HLA-DO 48. The method of embodiment 47, comprising one or more of: HLA-A, HLA-DOB, HLA-DMA, HLA-DMB, HLA-DPA1, HLA-DPB1, HLA-DPA2, HLA-DPB2, HLA-DPA3, HFE, TAP1, TAP2, PSMB9, PSMB8, MICB, MICA, MICC, MICD, MICE, KIR2DL1, KIR2DL2 / L3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, or KIR3DL3.
[0387] Embodiment 50: The method of embodiment 47, wherein the target region comprises chromosomal location 6p21 or chromosomal location 19q13.
[0388] Embodiment 51: The method of embodiment 47, wherein the target region comprises at least a portion of a major histocompatibility complex (MHC) region and the chromosome is human chromosome 6.
[0389] Embodiment 52: The method of embodiment 47, wherein the target region comprises at least a portion of a killer cell immunoglobulin-like receptor (KIR) region and the chromosome is human chromosome 19.
[0390] Embodiment 53: The method of embodiment 47, wherein the plurality of known sequences comprises a plurality of known human leukocyte antigen (HLA) allele sequences, or a plurality of known human killer cell immunoglobulin-like receptor (KIR) region allele sequences.
[0391] Embodiment 54: The method of embodiment 47, wherein the multiple decoy sequences comprise multiple non-human sequences.
[0392] Embodiment 55: The method of embodiment 54, wherein the plurality of non-human sequences comprises one or more of a plurality of bovine sequences, a plurality of rat sequences, or a plurality of microbial sequences.
[0393] Embodiment 56: The method of embodiment 47, wherein the step of generating an alignment of the plurality of sequence reads or read pairs to the plurality of known sequences includes determining, based on the alignment, for a sequence read or read pair among the plurality of sequence reads or read pairs, one or more known sequences to which each sequence read, or each read of a read pair, aligns without mismatches or indels.
[0394] Embodiment 57: The method of embodiment 56, further comprising the steps of: determining, for each known sequence of the plurality of known sequences, the number of sequence reads or the number of read pairs aligned to each known sequence based on the alignment; and determining, for one or more loci, the known sequences present at the one or more loci based on the number of sequence reads or the number of read pairs aligned to each known sequence.
[0395] Embodiment 58: The method of embodiment 47, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of decoy sequences includes, based on the alignment, determining, for each sequence read or read pair, one or more decoy sequences to which each sequence read or read pair aligns without mismatches or indels, and discarding the sequence read or read pair.
[0396] Embodiment 59: The method of embodiment 47, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of decoy sequences includes: determining, based on the alignment, for each sequence read or read pair, one or more non-human decoy sequences to which each sequence read or read pair aligns without mismatches or indels; and identifying the sequence read or read pair as originating from a contaminated sample.
[0397] Embodiment 60: The method of embodiment 47, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of known sequences includes determining, based on the alignment, for each sequence read or read pair, one or more known sequences to which each sequence read or read pair aligns with at least one mismatch or indel, and generating an alignment score for the known sequence.
[0398] Embodiment 61: The method of embodiment 47, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of decoy sequences includes determining, based on the alignment, for each sequence read or read pair, one or more decoy sequences to which each sequence read or read pair aligns with at least one mismatch or indel, and generating an alignment score for the decoy sequence.
[0399] Embodiment 62: The method of embodiment 47, wherein the alignment score to the known sequence and the alignment score to the decoy sequence each contain at least one mismatch and / or indel.
[0400] Embodiment 63: The method of embodiment 47, wherein the step of identifying a sequence read or read pair as a supporting sequence read or read pair for the sequence of interest comprises identifying a reference variant among the one or more reference variants as a candidate variant based on aligning the sequence read or read pair to the one or more reference variants.
[0401] Embodiment 64: The method of embodiment 47, wherein the step of identifying the sequence read or read pair as a supporting sequence read or read pair for the sequence of interest comprises identifying a non-human reference variant of the one or more non-human reference variants as a candidate variant based on aligning the sequence read or read pair to the one or more non-human reference variants, and identifying the sequence read or read pair as originating from a contaminated sample.
[0402] Embodiment 65: The method of embodiment 47, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of known sequences includes determining that the read pair aligns to at least two sequences of the plurality of known sequences, and selecting one known sequence of the at least two known sequences.
[0403] Embodiment 66: The method of embodiment 47, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of decoy sequences includes determining that the read pair aligns to at least two decoy sequences of the plurality of decoy sequences, and selecting one decoy sequence of the at least two decoy sequences.
[0404] Embodiment 67: A method comprising the steps of determining a plurality of sequence reads or read pairs for a target region of a chromosome of a subject, wherein the target region of the chromosome comprises one or more loci; generating an alignment of the plurality of sequence reads or read pairs to a plurality of known allele sequences; generating an alignment of the plurality of sequence reads or read pairs to a plurality of decoy allele sequences; comparing the alignment of the sequence reads or read pairs to the known alleles with an alignment of the sequence reads or read pairs to the decoy allele sequences, wherein if the alignment score of the sequence read or read pair to the known allele sequence is greater than the alignment score of the sequence read or read pair to the decoy allele sequence, the sequence read or read pair is associated with a germline allele; and identifying the sequence read or read pair as a supporting sequence read or read pair for a candidate somatic variant.
[0405] Embodiment 68: The method of embodiment 67, further comprising obtaining a sample from the subject and sequencing the sample to obtain a plurality of sequence reads or read pairs for the target region of the chromosome.
[0406] Embodiment 69: The target region is the following genes: HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-N, HLA-P, HLA-S, HLA-T, HLA-U, HLA-V, HLA-W, HLA-DRA, H LA-DRB1, HLA-DRB2, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-DRB6, HLA-DRB7, HLA-DRB8, HLA-DRB9, HLA-DQA1, HLA-DQB1, HLA-DQA2, HLA-DQB2, HLA-DQB3, HLA-DO 68. The method of embodiment 67, comprising one or more of: HLA-A, HLA-DOB, HLA-DMA, HLA-DMB, HLA-DPA1, HLA-DPB1, HLA-DPA2, HLA-DPB2, HLA-DPA3, HFE, TAP1, TAP2, PSMB9, PSMB8, MICB, MICA, MICC, MICD, MICE, KIR2DL1, KIR2DL2 / L3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, or KIR3DL3.
[0407] Embodiment 70: The method of embodiment 67, wherein the target region comprises chromosomal location 6p21 or chromosomal location 19q13.
[0408] Embodiment 71: The method of embodiment 67, wherein the target region comprises at least a portion of a major histocompatibility complex (MHC) region and the chromosome is human chromosome 6.
[0409] Embodiment 72: The method of embodiment 67, wherein the target region comprises at least a portion of a killer cell immunoglobulin-like receptor (KIR) region and the chromosome is human chromosome 19.
[0410] Embodiment 73: The method of embodiment 67, wherein the plurality of known allele sequences comprises a plurality of known human leukocyte antigen (HLA) allele sequences, or a plurality of known human killer cell immunoglobulin-like receptor (KIR) region allele sequences.
[0411] Embodiment 74: The method of embodiment 67, wherein the multiple decoy allele sequences comprise multiple non-human sequences.
[0412] Embodiment 75: The method of embodiment 74, wherein the plurality of non-human sequences comprises one or more of a plurality of bovine sequences, a plurality of rat sequences, or a plurality of microbial sequences.
[0413] Embodiment 76: The method of embodiment 67, wherein the step of generating an alignment of the plurality of sequence reads or read pairs to the plurality of known allele sequences comprises determining, based on the alignment, for a sequence read or read pair among the plurality of sequence reads or read pairs, one or more known allele sequences to which each sequence read, or each read of a read pair, aligns without mismatches or indels.
[0414] Embodiment 77: The method of embodiment 76, further comprising determining, for each known allele sequence of the plurality of known allele sequences, the number of sequence reads or the number of read pairs aligned to each known allele sequence based on the alignment, and determining, for one or more loci, the known allele sequences present at the one or more loci based on the number of sequence reads or the number of read pairs aligned to each known allele sequence.
[0415] Embodiment 78: The method of embodiment 67, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of decoy allele sequences includes determining, based on the alignment, for each sequence read or read pair, one or more decoy allele sequences to which each sequence read or read pair aligns without mismatches or indels, and discarding the sequence read or read pair.
[0416] Embodiment 79: The method of embodiment 67, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of decoy allele sequences includes determining, based on the alignment, for each sequence read or read pair, one or more non-human decoy allele sequences to which each sequence read or read pair aligns without mismatches or indels, and identifying the sequence read or read pair as originating from a contaminated sample.
[0417] Embodiment 80: The method of embodiment 67, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of known allele sequences includes determining, based on the alignment, for the sequence reads or read pairs, one or more known allele sequences to which each sequence read or read pair aligns with at least one mismatch or indel, and generating an alignment score for the known allele sequences.
[0418] Embodiment 81: The method of embodiment 67, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of decoy allele sequences includes determining, based on the alignment, for each sequence read or read pair, one or more decoy allele sequences to which each sequence read or read pair aligns with at least one mismatch or indel, and generating an alignment score for the decoy allele sequence.
[0419] Embodiment 82: The method of embodiment 67, wherein the alignment score to the known allele sequence and the alignment score to the decoy allele sequence each contain at least one mismatch and / or indel.
[0420] Embodiment 83: The method of embodiment 67, wherein the step of identifying a sequence read or read pair as a supporting sequence read or read pair for the sequence of interest comprises identifying a reference variant among the one or more reference variants as a candidate variant based on aligning the sequence read or read pair to the one or more reference variants.
[0421] Embodiment 84: The method of embodiment 67, wherein the step of identifying the sequence read or read pair as a supporting sequence read or read pair for the sequence of interest comprises identifying a non-human reference variant of the one or more non-human reference variants as a candidate variant based on aligning the sequence read or read pair to the one or more non-human reference variants, and identifying the sequence read or read pair as originating from a contaminated sample.
[0422] Embodiment 85: The method of claim 67, wherein the step of generating an alignment of the plurality of sequence reads or read pairs to the plurality of known allele sequences comprises determining that the read pair aligns to at least two of the plurality of known allele sequences, and selecting one known allele sequence of the at least two known allele sequences.
[0423] Embodiment 86: The method of claim 67, wherein the step of generating an alignment of a plurality of sequence reads or read pairs to a plurality of decoy allele sequences includes determining that the read pair aligns to at least two decoy allele sequences of the plurality of decoy allele sequences, and selecting one decoy allele sequence of the at least two decoy allele sequences.
[0424] Embodiment 89: A method comprising the steps of: determining a plurality of pairs of sequence reads of a target region of a chromosome of a subject, wherein the target region of the chromosome comprises one or more loci; generating germline alignments of the plurality of pairs of sequence reads to a plurality of known allele sequences; generating decoy alignments of the plurality of pairs of sequence reads to a plurality of decoy allele sequences; determining pairs of sequence reads among the plurality of pairs of sequence reads by both a germline alignment having at least one mismatch and / or indel and a decoy alignment having at least one mismatch and / or indel, wherein the pairs of sequence reads are with a germline alignment score greater than the decoy alignment score; and identifying the pairs of sequence reads as candidate somatic variants.
[0425] Embodiment 90: The method of embodiment 89, further comprising obtaining a sample from the subject and sequencing the sample to obtain a plurality of pairs of sequence reads for the target region of the chromosome.
[0426] Embodiment 91: The target region is the following genes: HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HLA-K, HLA-L, HLA-N, HLA-P, HLA-S, HLA-T, HLA-U, HLA-V, HLA-W, HLA-DRA, H LA-DRB1, HLA-DRB2, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-DRB6, HLA-DRB7, HLA-DRB8, HLA-DRB9, HLA-DQA1, HLA-DQB1, HLA-DQA2, HLA-DQB2, HLA-DQB3, HLA-DO 90. The method of embodiment 89, comprising one or more of: HLA-DPA1, HLA-DPB1, HLA-DPA2, HLA-DPB2, HLA-DPA3, HFE, TAP1, TAP2, PSMB9, PSMB8, MICB, MICA, MICC, MICD, MICE, KIR2DL1, KIR2DL2 / L3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, or KIR3DL3.
[0427] Embodiment 92: The method of embodiment 89, wherein the target region comprises chromosomal location 6p21 or chromosomal location 19q13.
[0428] Embodiment 93: The method of embodiment 89, wherein the target region comprises at least a portion of a major histocompatibility complex (MHC) region and the chromosome is human chromosome 6.
[0429] Embodiment 94: The method of embodiment 89, wherein the target region comprises at least a portion of a killer cell immunoglobulin-like receptor (KIR) region and the chromosome is human chromosome 19.
[0430] Embodiment 95: The method of embodiment 89, wherein the plurality of known allele sequences comprises a plurality of known human leukocyte antigen (HLA) allele sequences, or a plurality of known human killer cell immunoglobulin-like receptor (KIR) region allele sequences.
[0431] Embodiment 96: The method of embodiment 89, wherein the multiple decoy sequences comprise multiple non-human sequences.
[0432] Embodiment 97: The method of embodiment 96, wherein the plurality of non-human sequences comprises one or more of a plurality of bovine sequences, a plurality of rat sequences, or a plurality of microbial sequences.
[0433] Embodiment 98: The method of embodiment 89, wherein the step of generating germline alignments of a plurality of pairs of sequence reads to a plurality of known allele sequences comprises determining, for a pair of sequence reads among the plurality of pairs of sequence reads, one or more known allele sequences to which each read of the pair of sequence reads aligns without mismatches or indels based on the germline alignment.
[0434] Embodiment 99: The method of embodiment 98, further comprising the steps of: for each known allele sequence of the plurality of known allele sequences, determining a number of pairs of sequence reads among the plurality of pairs of sequence reads that align to each known allele sequence based on the germline alignment; and determining, for one or more loci, known allele sequences present at the one or more loci based on the number of pairs of sequence reads that align to each known allele sequence.
[0435] Embodiment 100: The method of embodiment 89, wherein the step of generating decoy alignments of a plurality of pairs of sequence reads to a plurality of decoy allele sequences includes: determining, for a pair of sequence reads among the plurality of pairs of sequence reads, based on the decoy alignments, one or more decoy allele sequences to which each read of the pair of sequence reads aligns without mismatches or indels; and discarding the pair of sequence reads.
[0436] Embodiment 101: The method of embodiment 89, wherein the step of generating decoy alignments of a plurality of pairs of sequence reads to a plurality of decoy allele sequences includes: determining, for a pair of sequence reads among the plurality of pairs of sequence reads, one or more non-human decoy sequences to which each read of the pair of sequence reads aligns without mismatches or indels based on the decoy alignments; and identifying the plurality of pairs of sequence reads as originating from a contaminated sample.
[0437] Embodiment 102: The method of embodiment 89, wherein the step of generating germline alignments of a plurality of pairs of sequence reads to a plurality of known allele sequences includes determining, for a pair of sequence reads among the plurality of pairs of sequence reads, one or more known allele sequences to which each read of the pair of sequence reads aligns with at least one mismatch or indel based on the germline alignment, and generating a germline alignment score.
[0438] Embodiment 103: The method of embodiment 102, wherein the step of generating decoy alignments of a plurality of pairs of sequence reads to a plurality of decoy allele sequences includes determining, for a pair of sequence reads among the plurality of pairs of sequence reads based on the decoy alignments, one or more decoy allele sequences to which each read of the pair of sequence reads aligns with at least one mismatch or indel, and generating a decoy alignment score.
[0439] Embodiment 104: The method of embodiment 89, wherein the step of identifying the pair of sequence reads as a candidate somatic variant comprises identifying a reference variant among the one or more reference variants as a candidate variant based on aligning the pair of sequence reads to the one or more reference variants.
[0440] Embodiment 105: The method of embodiment 89, wherein the step of identifying pairs of sequence reads as candidate somatic variants comprises identifying a non-human reference variant among the one or more non-human reference variants as a candidate variant based on aligning pairs of sequence reads to the one or more non-human reference variants, and identifying a plurality of pairs of sequence reads as originating from a contaminated sample.
[0441] Embodiment 106: The method of embodiment 89, wherein the step of generating germline alignments of a plurality of pairs of sequence reads to a plurality of known allele sequences comprises determining that the pair of sequence reads aligns to at least two allele sequences of the plurality of known allele sequences, and selecting one known allele sequence of the at least two allele sequences.
[0442] Embodiment 107: The method of claim 89, wherein the step of generating decoy alignments of a plurality of pairs of sequence reads to a plurality of decoy allele sequences includes determining that the pair of sequence reads aligns to at least two decoy allele sequences of the plurality of decoy allele sequences, and selecting one decoy allele sequence of the at least two decoy allele sequences.
[0443] All patent applications, websites, other publications and accession numbers, etc. cited above or below are incorporated by reference in their entirety for all purposes to the same extent as if each individual item was specifically and individually indicated to be incorporated by reference in its entirety for all purposes. Where different versions of a sequence are accompanied by accession numbers at different times, it refers to the version accompanied by the accession number as of the effective filing date of this application. Where applicable, the effective filing date refers to the earlier of the actual filing date or the filing date of the priority application referring to the accession number. Similarly, where different versions of a publication or website, etc. are published at different times, it refers to the version published most recent to the effective filing date of this application, unless otherwise indicated. Features described in connection with separate aspects and embodiments of the invention may be used together and / or interchangeable. Similarly, features described in connection with a single embodiment may also be provided separately or in any suitable subcombination. Any feature, step, element, embodiment or aspect of this disclosure may be used in combination with any other, unless specifically indicated otherwise. Although the present disclosure has been described in some detail by way of illustration and example, for purposes of clarity and understanding, it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims.
Claims
1. Steps to determine multiple known allele sequences; A step of determining multiple sequence reads of a target region of a target chromosome, wherein the target region of the chromosome includes one or more gene loci; A step of aligning the plurality of sequence reads to the plurality of known allele sequences; A step of determining, based on the alignment, the number of sequence reads aligned to each of the plurality of known allele sequences; and Step 1: Determine the known allele sequences present at one or more gene loci based on the number of sequence reads aligned to each known allele sequence. A method that includes this.
2. The method according to claim 1, wherein the plurality of known allele sequences include a plurality of known human leukocyte antigen (HLA) allele sequences or a plurality of known killer cell immunoglobulin-like receptor (KIR) region allele sequences.
3. The step of sequencing the sample obtained from the subject to obtain the plurality of sequence reads of the target region of the chromosome. The method according to claim 1, further comprising:
4. Based on the alignment, the step of determining one or more known allele sequences that each of the plurality of sequence reads aligns with. The method according to claim 1, further comprising:
5. The target region contains the following genes: HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, HLA-G, HLA-H, HLA-J, HL AK, HLA-L, HLA-N, HLA-P, HLA-S, HLA-T, HLA-U, HLA-V, HLA-W, HLA-DRA, HLA-D RB1, HLA-DRB2, HLA-DRB3, HLA-DRB4, HLA-DRB5, HLA-DRB6, HLA-DRB7, HLA-DRB 8, HLA-DRB9, HLA-DQA1, HLA-DQB1, HLA-DQA2, HLA-DQB2, HLA-DQB3, HLA-DOA, H The method according to claim 1, comprising one or more of LA-DOB, HLA-DMA, HLA-DMB, HLA-DPA1, HLA-DPB1, HLA-DPA2, HLA-DPB2, HLA-DPA3, HFE, TAP1, TAP2, PSMB9, PSMB8, MICB, MICA, MICC, MICD, MICE, KIR2DL1, KIR2DL2 / L3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DL1 / S1, KIR3DL2, or KIR3DL3.
6. The method according to claim 1, wherein the target region includes chromosome position 6p21 or chromosome position 19q13.
7. The method according to claim 1, wherein the target region comprises at least a portion of the major histocompatibility complex (MHC) region, and the chromosome is human chromosome 6.
8. The method according to claim 1, wherein the target region comprises at least a portion of a killer cell immunoglobulin-like receptor (KIR) region, and the chromosome is human chromosome 19.
9. The step of determining the plurality of known allele sequences is Searching for sequence data for known alleles; Determining multiple k-mer sequences based on the sequence data for known alleles; and Assembling the multiple k-mer sequences into a data structure that includes multiple entries, each entry including a k-mer sequence, the number of alleles containing the k-mer sequence, an allele identifier for each allele containing the k-mer sequence, and the start and end positions of the k-mer sequence in each allele. The method according to claim 1, including the method described in claim 1.
10. The method according to claim 1, wherein the step of determining the number of sequence reads aligned to each of the plurality of known allele sequences based on the alignment includes determining the number of sequence reads aligned with each known allele sequence with 100% identity.
11. The method according to claim 1, wherein the step of determining the known allele sequences present at one or more loci based on the number of sequence reads aligned to each known allele sequence includes determining one or more known allele sequences having the highest number of aligned sequence reads.
12. The method according to claim 1, wherein the reads aligned to each known allele sequence are grouped into read families, and the method further comprises the step of determining the number of sequence read families aligned to each known allele sequence.
13. The method according to claim 12, wherein, for one or more gene loci, the known allele sequences present at one or more gene loci are determined based on the number of sequence read families aligned to each known allele sequence.
14. The method according to claim 1, further comprising the step of determining the length of each known allele sequence portion aligned to two or more sequence reads from the plurality of sequence reads.
15. For each gene locus, the step is to sort the known allele sequences present at the gene locus by the number of sequence reads aligned to each known allele sequence; For the aforementioned gene locus, the step of determining a first known allele sequence having the highest number of aligned sequence reads; Steps include inserting the first known allele sequence having the highest number of aligned sequence reads into the superset; The steps of determining one or more known allele sequences aligned to a subset of the reads aligned to the first known allele sequence; and Step of inserting one or more known allele sequences into the superset. The method according to claim 1, further comprising:
16. The method according to claim 15, wherein the superset includes a graph data structure.
17. The method according to claim 16, wherein the graph data structure includes a directed acyclic graph.
18. The method according to claim 16, wherein the graph data structure represents a Hasse diagram.
19. The step of determining that the aforementioned locus is associated with a single superset; and The step of determining the first known allele sequence of the single superset as the allele present at the gene locus. The method according to claim 15, further comprising:
20. The method according to claim 15, further comprising the step of determining a plurality of supersets for the gene locus.
21. A step of determining two supersets having the maximum cumulative number of distinct reads based on the plurality of supersets for the gene locus; and The step of determining the first known allele sequence of each of the two supersets as the allele present at the gene locus. The method according to claim 20, further comprising:
22. The method according to claim 1, further comprising the step of assisting in communicating the known allele sequences located at one or more gene loci to a healthcare provider.
23. The method according to claim 1, wherein the known allele sequence present at the gene locus comprises a human leukocyte antigen (HLA) type or a killer cell immunoglobulin-like receptor (KIR) type at the gene locus.
24. The step of sequencing a sample obtained from cells intended to be transplanted into the subject to obtain the plurality of sequence reads of the target region of the chromosome. The method according to claim 1, further comprising:
25. The method according to claim 24, further comprising the step of determining whether the target HLA type matches the HLA type at the gene locus.
26. The method according to claim 25, characterized in that the match between the HLA type of the subject and the HLA type at the gene locus indicates that the cells are intended to be transplanted into the subject.
27. The match between the HLA type of the subject and the HLA type at the gene locus is The cells are to be transplanted into the subject; and The subjects are scheduled to be administered drugs that reduce the likelihood of transplant rejection. The method according to claim 25, characterized by showing