Methods and systems to determine HLA-DPB1 expression
Long-read sequence data analysis predicts DPB1 expression levels, addressing HLA-DPB1 mismatch issues in hematopoietic stem cell transplantation to reduce graft-versus-host disease risk.
Patent Information
- Application Number
- JP2025078596
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-02-27
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-05
AI Technical Summary
Current HLA-DPB1 mismatch in hematopoietic stem cell transplantation often leads to graft-versus-host disease due to the lack of routine genotyping for the rs9277534 marker, which influences expression levels, necessitating a method to predict DPB1 expression accurately.
A method and system using long-read sequence data analysis to predict DPB1 expression levels without requiring rs9277534 sequence data, involving computer-implemented alignment and motif identification in exon 3 of DPB1.
Enables accurate prediction of DPB1 expression levels, reducing the risk of graft-versus-host disease by identifying suitable donor-recipient pairs through HLA typing.
Smart Images

Figure 2025114749000003 
Figure 2025114749000004 
Figure 2025114749000005
Abstract
Description
[Technical Field]
[0001] Priority claim This application claims the benefit of and priority to U.S. Provisional Application No. 62 / 982,286, filed February 27, 2020, which is incorporated herein by reference in its entirety for all purposes.
[0002] Field The present disclosure relates to methods and systems for determining DPB1 expression for HLA typing used to match transplant donors and recipients. [Background technology]
[0003] Hematopoietic stem cell transplantation (HSCT) from unrelated donors can cure various blood disorders, but a high level of donor-recipient HLA compatibility is important for success. Currently, matching of HLA-A, -B, -C, -DRB1, -DRB3, and -DQB1 alleles is the absolute standard. HLA-DPB1 is often considered as well, but genetic distance from other HLA genes leads to frequent mismatches at this locus. However, when mismatched, HLA-DPB1 expression level may play an important role in hematopoietic stem cell transplantation. Studies have shown that donors and recipients with mismatched expression levels for HLA-DPB1 are more likely to develop graft-versus-host disease (GvHD).
[0004] In particular, it has been found that the expression level of HLA-DPB1 can be correlated with an A-to-G single nucleotide polymorphism (SNP), rs9277534, located in the 3' untranslated region (UTR) of DPB1 (Thomas et al., J. Virol., 86:6979-85 (2012)). The "A" allele is associated with weak DPB1 expression, while the "G" allele is associated with strong DPB1 expression (Petersdorf et al., New Engl. J. Med. 373:599-609 (2015)). Research has shown that the rs9277534 A-to-G polymorphism is closely linked to seven specific nucleotide variants in DPB1 exon 3 (see, e.g., Schone et al., Human Immunol., 79:20-27 (2018)).
[0005] The risk of GvHD associated with HLA-DPB1 mismatches is influenced by the HLA-DPB1 rs9277534 expression marker. In recipients of HLA-DPB1 mismatched transplants (e.g., exon 2 mismatches or other mismatches) from donors with low-expression alleles (rs277534A), recipients with high-expression alleles have a higher risk of acute GvHD (Petersdorf et al., (2015)). However, the 3'UTR containing rs9277534 is not currently covered by routine genotyping assays for HLA-DPB1, and expression assays are not routinely performed. Therefore, there is a need to develop methods to characterize DPB1 expression in HSCT transplant patients on sequenced samples as a means to identify suitable donor:recipient pairs. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] Thomas et al., J. Virol., 86:6979-85 (2012) [Non-patent document 2] Petersdorf et al., New Engl. J. Med. 373:599-609 (2015) [Non-patent document 3] Schone et al., Human Immunol., 79:20-27 (2018) Summary of the Invention [Means for solving the problem]
[0007] The embodiments of the present disclosure include a method and system for analyzing long-read sequence data to predict DPB1 expression level.In certain embodiments, the method and system are computer-implemented.A computer program product for implementing the method and / or system of the present disclosure is also disclosed.In one embodiment, the disclosed method and system are useful for determining HSCT donor-recipient compatibility.The method and system do not require determining the sequence of rs9277534 or measuring DPB1 expression as mRNA or protein level.
[0008] For example, in one embodiment, there is provided a computer-implemented method for analyzing sequence data from a subject to predict DPB1 expression levels in the subject, comprising: (a) a computer-implemented step of obtaining a query nucleic acid sequence from a sample of interest using a long-read sequencer; (b) using a computer-implemented alignment program to align a query nucleic acid sequence from the subject to a reference nucleic acid sequence, wherein the reference nucleic acid sequence and the query nucleic acid sequence comprise long-read sequence data for at least exon 3 of DPB1; (b) using a computer-implemented algorithm to identify, based on the aligned query nucleic acid sequence and the reference nucleic acid sequence, whether the query nucleic acid sequence has a sequence characteristic of low levels of DPB1 expression or high levels of DPBI expression; (i) comparing nucleotides in the aligned query nucleic acid sequence with a reference nucleic acid sequence to identify differences between the query nucleic acid sequence and the reference nucleic acid sequence; (ii) determining the identity of a nucleotide to the query nucleic acid sequence as compared to the reference nucleic acid sequence at a defined position in exon 3 of DPB1 based on the identified differences between the query nucleic acid sequence and the reference nucleic acid sequence; (iii) determining whether the query sequence exhibits a sequence characteristic of a weakly expressed motif, a sequence characteristic of a strongly expressed motif, or neither, based on the identity of the nucleotide to the query nucleic acid sequence at a defined position in exon 3; (iv) identifying the subject as having a low expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a weakly expressed motif, or identifying the subject as having a high expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a strongly expressed motif.
[0009] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods or processes disclosed herein.
[0010] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more methods disclosed herein.
[0011] The terms and expressions employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude the features shown and described or equivalents thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, while the present disclosure has been specifically disclosed by embodiments and optional features, it will be understood that variations and modifications of the concepts disclosed herein may be employed by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the present disclosure as defined by the appended claims.
[0012] The present disclosure may be better understood from the following non-limiting figures. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 shows sequences in DPB1 cDNA associated with the rs9277534 A allele and weak expression and the rs9277534 G allele and strong expression, according to an embodiment of the present disclosure. [Figure 2] Figure 2 shows the sequences in the DPB1 cDNA for the rs9277534 G allele and haplotype 01:01:01:01 (SEQ ID NO: 1), which is associated with strong expression, and the DPB1 cDNA for the rs9277534 A allele and haplotype 02:01:02:01 (SEQ ID NO: 2), which is associated with weak expression. [Figure 3] 3 illustrates the use of the disclosed method for identifying exon 3 motifs according to various embodiments of the present disclosure. The query positive strand (Qry+) sequence (SEQ ID NO: 4) has the sequence of a strongly expressed motif (solid rectangle) at positions 20, 27, 52, 87, 234, 242, and 270 of exon 3, as well as other variations (dotted rectangles), and the reference positive strand sequence (Ref+) (SEQ ID NO: 3) has a weakly expressed motif. In this figure, the sequencing alignment starts at the fourth nucleotide of exon 3. [Figure 4]Figure 4 shows the use of the disclosed algorithm to identify exon 3 motifs according to various embodiments of the disclosure. The query positive strand (Qry+) (SEQ ID NO: 5) sequence has the sequence of the weakly expressed motif (solid rectangle) at positions 20, 27, 52, 87, 234, 242, and 270, as well as other variations (dotted rectangle). The reference positive strand sequence (Ref+) (SEQ ID NO: 3) is the same reference sequence as in Figure 3, and also has the weakly expressed motif. In this figure, the sequencing alignment starts at the fourth nucleotide of exon 3. [Figure 5] FIG. 5 shows an embodiment of the disclosed method for determining DPB1 exon 3 motifs associated with strong and weak expression. [Figure 6] FIG. 6 shows another embodiment of the disclosed method for determining DPB1 exon 3 motifs associated with strong and weak expression. [Figure 7] Figure 7 shows a comparison of exon 3 DPB1 34:01 (weakly expressed and Ref+ sequence of Figure 3) (SEQ ID NO: 3) and DPB1 01:01 (strongly expressed) (SEQ ID NO: 6) sequences according to one embodiment of the present disclosure. It can be seen that only positions 7, 20, 27, 52, 87, 234, 242, 270 of exon 3 show differences between the reference (Ref+) and the query (Qry+). The Concise Idiosyncratic Gapped Alignment Report (CIGAR) presented at the top of the figure shows how this alignment is reported in an algorithm for further analysis. [Figure 8] Figure 8 shows a comparison of the DPB1 34:01 (weakly expressed) (SEQ ID NO: 3) and DPB1 02:01 (weakly expressed) (SEQ ID NO: 7) sequences in exon 3 according to one embodiment of the present disclosure. It can be seen that there is an exact match for positions 7, 20, 27, 52, 87, 234, 242, 270 (and other exon positions) in exon 3. The CIGAR string above the figure indicates how this alignment would be reported by the algorithm for further analysis. [Figure 9]9 illustrates the CIGAR output of various sequences analyzed using methods according to various embodiments of the present disclosure, showing five strong and five weak sequences, one of which has a variant at position 243 of DPB1 exon 3. [Figure 10] FIG. 10 shows a sequence alignment for the cs:z::242*ct:59 weak sequence (SEQ ID NO: 8) of FIG. 9 compared to the reference (Ref+) (SEQ ID NO: 3) according to one embodiment of the present disclosure showing a variant at position 243. [Figure 11] FIG. 11 illustrates an exemplary computing device according to various embodiments of the present disclosure. [Figure 12] FIG. 12 shows additional results of methods and systems according to various embodiments of the present disclosure obtained using an RSII SMRTcell. [Figure 13] 13 shows additional results of methods and systems according to various embodiments of the present disclosure obtained using a Sequel SMRTcell. A total of 299 motifs in 88 samples were obtained. [Figure 14] 14 shows a visual display of results for HLA genes DRB1, DRB3, DQB1, and DPB1 according to one embodiment of the present disclosure. The 02:01:026 sequence with 133 reads is characterized as weak, and the 0:3:01:01G sequence with 383 reads has 7 mismatches at defined positions in exon 3 that correlate with highly expressed SNPs and is characterized as strong. [Figure 15] FIG. 15 shows that a total of 31,274 DPB1 sequences were analyzed according to one embodiment of the present disclosure, of which 10,774 were typed as "strong," 20,369 were typed as "weak," and 131 were typed as "undetermined." [Figure 16] FIG. 16 shows a plot of exon 3 positions where unique variants occur relative to a weak reference, according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0014] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.
[0015] In the following description, specific details are set forth to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0016] definition In order that this disclosure may be more readily understood, certain terms are first defined. Additional definitions of the following terms, as well as other terms, are set forth throughout the specification.
[0017] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the present disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. However, any numerical value inherently contains certain errors necessarily resulting from the standard deviation found in their respective testing measurements. Moreover, all ranges disclosed herein should be understood to encompass any and all subranges subsumed therein. For example, a range stated as "1 to 10" should be considered to include any and all subranges between (and including) the minimum value of 1 and the maximum value of 10. That is, all subranges beginning with a minimum value of 1 or more, e.g., 1 to 6.1, and ending with a maximum value of 10 or less, e.g., 5.5 to 10. Furthermore, any reference referred to as "incorporated herein" should be understood to be incorporated in its entirety.
[0018] It is further noted that, as used herein, the singular forms "a," "an," and "the" include plural referents unless expressly and unambiguously limited to one referent. The term "and / or" is generally used to refer to at least one or the other. In some cases, the term "and / or" is used interchangeably with the term "or." The term "including" is used herein to mean, and is used interchangeably with, the phrase "including but not limited to." The term "such as" is used herein to mean, and is used interchangeably with, the phrase "such as, but not limited to."
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0020] Also, as used herein, "at least one" refers to any number from 1 to the entire group. For example, for a list of four variants, the phrase "at least one" is understood to mean 1, 2, 3 or 4 variants. Similarly, for a list of 10 variants, the phrase "at least five" is understood to mean 5, 6, 7, 8, 9 or 10 variants.
[0021] Also, as used herein, "comprising" includes embodiments more particularly defined using the term "consisting of."
[0022] Also, as used herein, the terms "substantially," "approximately," and "about" are defined as being mostly, but not necessarily fully specified (and including), as would be understood by one of ordinary skill in the art. In any disclosed embodiment, the terms "substantially," "approximately," or "about" can be replaced with "within [percentage]" of what is specified, where percentage includes 0.1, 1, 5, and 10%.
[0023] As used herein, when an action is "based on" something, this means that the action is based at least in part on at least a part of the something.
[0024] Activity: As used herein, the term "activity" refers to the expression level of a gene. For example, DPB1 activity refers to the level of DPB1 mRNA and / or protein.
[0025] Allele: As used herein, the term "allele" refers to different versions of the nucleotide sequence of the same genetic locus (e.g., gene).
[0026] Coding sequence: As used herein, the term "coding sequence" refers to a sequence of a nucleic acid or its complement, or a portion thereof, that can be transcribed and / or translated to produce mRNA for a polypeptide or fragment thereof. Coding sequences include exons in genomic DNA or premature primary RNA transcripts, which are joined together by the cell's biochemical machinery to provide mature mRNA. The antisense strand is the complement of such a nucleic acid, from which the coding sequence can be deduced. As used herein, the term "non-coding sequence" refers to a sequence of a nucleic acid or its complement, or a portion thereof, that is not transcribed into amino acids in vivo or with which tRNA interacts to place or attempt to place amino acids. Non-coding sequences include both intron sequences in genomic DNA or premature primary RNA transcripts and gene-related sequences such as promoters, enhancers, and silencers.
[0027] Contig: As used herein, the term "contig" refers to a DNA sequence that represents the consensus sequence of a region of DNA assembled from a set of identical, completely or partially overlapping DNA sequencing reads. In some cases, the sequencing reads are generated from next-generation sequencing reactions.
[0028] Deletion: As used herein, the term "deletion" includes a mutation that removes one or more nucleotides from a naturally occurring nucleic acid.
[0029] Exon: As used herein, the term "exon" refers to a nucleic acid sequence found in mature or processed RNA after other portions of the RNA (e.g., intervening regions known as introns) have been removed by RNA splicing. Thus, an exon sequence generally encodes a protein or a portion of a protein. An intron is a portion of RNA that is removed from the surrounding exon sequence by RNA splicing.
[0030] Expression and expressed RNA: As used herein, expressed RNA is RNA that encodes a protein or polypeptide ("coding RNA"), as well as any other RNA that is transcribed but not translated ("non-coding RNA"). The term "expression" is used herein to refer to the process by which a polypeptide is produced from DNA. This process involves transcription of a gene into mRNA and translation of this mRNA into a polypeptide. Depending on the context, "expression" can refer to the production of RNA, protein, or both. As used herein, "strong" expression of DPB1 includes expression levels that are approximately 1.5- to 2-fold greater than weak expression (see, e.g., Petersdorf et al., (2015)).
[0031] Gene: As used herein, the term "gene" refers to a unit of heredity. Generally, a gene is a segment of DNA that encodes a protein or functional RNA. A gene is a locatable region of a genomic sequence that corresponds to a unit of heredity. A gene may be associated with regulatory regions, transcriptional regions, and / or other functional sequence regions.
[0032] Genotype: As used herein, the term "genotype" refers to the genetic makeup of an organism. More specifically, this term refers to the identity of the alleles present in an individual. "Genotyping" an individual or DNA sample refers to identifying the nature of the two alleles that an individual possesses at known polymorphic sites, in terms of nucleotide bases.
[0033] Heterozygous: As used herein, the term "heterozygous" refers to an individual who has two different alleles of the same gene. As used herein, the term "heterozygous" encompasses "compound heterozygous" or "compound heterozygous mutant." As used herein, the term "compound heterozygous" refers to an individual who has two different alleles. As used herein, the term "compound heterozygous mutant" refers to an individual who has two different copies of an allele, and such alleles are characterized as mutant forms of the gene.
[0034] Homozygous: As used herein, the term "homozygous" refers to an individual who has two copies of the same allele. As used herein, the term "homozygous mutant" refers to an individual who has two copies of the same allele, and such an allele is characterized as a mutant form of the gene.
[0035] Insertion or Addition: As used herein, the term "insertion" or "addition" refers to a change in an amino acid or nucleotide sequence that results in the addition of one or more amino acid residues or nucleotides, respectively, as compared to the naturally occurring molecule.
[0036] Long-read sequencing: As used herein, "long-read sequencing," also known as third-generation sequencing, is a DNA sequencing technology that can determine the nucleotide sequence of long-read sequences of DNA from 10,000 base pairs to 100,000 base pairs at a time. This eliminates the need for DNA shearing and subsequent amplification, which is typically required with other DNA sequencing technologies. Long-read sequencing also allows for unambiguous linkage between exons 2 and 3 of DPB1. Short-read sequencing often loses phasing and cannot link strong or weak motifs to the appropriate exon 2.
[0037] Long-read sequence: As used herein, a "long-read sequence" is a continuous read of a single molecule, such as a PCR amplicon, that allows for the elimination of phasing required by shorter-read techniques.
[0038] Mutation and / or variant: As used herein, the terms "mutation" and "variant" are used interchangeably to describe changes in the sequence of a nucleic acid or protein. As used herein, the term "mutant" refers to a mutant or potentially non-functional form of a gene. This term includes any mutation that renders a gene non-functional, from point mutations to large chromosomal rearrangements, as known in the art.
[0039] Nucleic Acid: As used herein, the term "nucleic acid" refers to a polynucleotide such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The term is used to include single-stranded nucleic acids, double-stranded nucleic acids, mRNA, and RNA and DNA made from nucleotide or nucleoside analogs.
[0040] Polymorphism: As used herein, the term "polymorphism" refers to the coexistence of two or more forms of a gene or portion thereof.
[0041] Query sequence: As used herein, the term "query sequence" or "query" or "Qry" refers to an uncharacterized nucleic acid consensus sequence that is compared to a known sequence for analysis of unknown sequence content. The sequence can be genomic DNA or deoxyribonucleic acid, such as messenger RNA (cDNA) or copy DNA amplified from ribonucleic acid (RNA). In one embodiment, the sequence is genomic DNA.
[0042] Reference sequence: As used herein, the term "reference sequence" or "reference" or "Ref" refers to a known sequence used as a standard for the analysis of a query or unknown sequence. The sequence can be genomic DNA or deoxyribonucleic acid, such as messenger RNA (cDNA) or copy DNA amplified from ribonucleic acid (RNA). In one embodiment, the sequence is genomic DNA.
[0043] Sample: As used herein, the term "sample" refers to any type of suitable biological specimen or sample (e.g., a test sample) from which nucleic acids can be isolated. A biological specimen or sample can be any specimen or sample isolated or obtained from a subject or a portion thereof (e.g., a human subject, a pregnant woman, a fetus). Non-limiting examples of specimens or samples include fluids or tissues from a subject, including, but not limited to, blood or blood products (e.g., serum, plasma, etc.), umbilical cord blood, chorioamniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (e.g., bronchoalveolar, stomach, peritoneal, duct, ear, arthroscope), biopsy samples (e.g., from preimplantation embryos), abdominal paracentesis samples, cells (blood cells, placental cells, embryonic or fetal cells, fetal nucleated cells or fetal cell debris) or portions thereof (e.g., mitochondria, nuclei, extracts, etc.), female genital tract washings, urine, stool, sputum, saliva, nasal mucosa, prostatic fluid, lavage fluid, semen, lymph, bile, tears, sweat, breast milk, milk, etc., or combinations thereof.
[0044] Sense strand vs. antisense strand: As used herein, the term "sense strand" refers to the strand of double-stranded DNA (dsDNA) that contains at least a portion of the coding sequence of a functional protein. As used herein, the term "antisense strand" refers to the strand of dsDNA that is the reverse complement of the sense strand. As used herein, the plus strand or + strand is the sense strand. As used herein, "+" refers to the sense strand and "-" refers to the antisense strand.
[0045] Subject or Individual or Patient: As used herein, the term "subject" or "individual" refers to a human or any non-human animal. A subject or individual may be a patient, which refers to a human presenting to a healthcare provider for diagnosis or treatment of a disease, in some cases the disease requiring a hematopoietic stem cell transplant. Also, as used herein, the term "individual," "subject," or "patient" includes all warm-blooded animals.
[0046] Methods for analyzing sequence data to predict DPB1 expression levels HLA-DPB1 expression has been shown to play an important role in hematopoietic stem cell transplantation. Research has demonstrated that donors and recipients with mismatched HLA-DPB1 expression levels are more likely to develop graft-versus-host disease. An embodiment of the present disclosure includes a method for analyzing sequence data to predict DPB1 expression levels. This method can be used to evaluate donor databases to find matches for individuals who need transplants, or during sequencing for HLA typing in reference laboratories. In one embodiment, the transplant is hematopoietic stem cell transplantation. In certain embodiments, at least some of the steps can be computer-implemented.
[0047] This method addresses the need to determine DPB1 expression levels as part of HLA haplotyping analysis. Analysis of DPB1 sequences for HLA haplotyping is often based solely on DPB1 exon 2 and therefore does not include sequences outside of exon 2 (i.e., exons 1, 3, 4, and 5, upstream and / or downstream sequences, regulatory sequences, epigenetic sequences, introns, etc.). However, there is no suggestion that polymorphisms in exon 2 are associated with DPB1 activity. Instead, it was discovered that the marker rs9277534 in the 3'UTR of DPB1 genetically controls HLA-DPB1 expression, and seven specific base pairs within exon 3 can predict single-nucleotide variants of marker rs9277534 (single-nucleotide variants that determine expression levels). Therefore, HLA-DPB1 expression can be predicted by analyzing these seven clearly defined positions within exon 3 using sequence data obtained by long-read sequencing. Figure 1 shows the DPB1 gene structure with the position and base in exon 3 predicting the rs9277543 marker. Single nucleotide variants in this marker control DPB1 expression levels. Figure 2 shows the exon 3 sequence for DPB1*04:01:01:01, which shows a weakly expressed sequence, and DPB1*01:01:01:01, which shows a strongly expressed sequence.
[0048] In various embodiments, methods are provided that involve determining the sequence of DPB1 exon 3, particularly nucleotide variations associated with DPB1 activity. In one embodiment, a method is disclosed for analyzing seven clearly defined positions within DPB1 exon 3 associated with DPB1 activity. The sequences can be contigs constructed as part of long-read or short-read sequencing experiments. In some embodiments, sequencing is performed by next-generation sequencing (NGS), such as long-read sequencing. Thus, in certain embodiments, the methods may be used to predict HLA-DPB1 expression. The methods may also be used to predict transplant rejection and / or graft-versus-host disease (GvHD). In certain embodiments, the methods and systems use bioinformatics to analyze exon 3 sequences from donor and patient NGS data (e.g., long-read sequences) and predict the expression level of each HLA-DPB1 allele.
[0049] For example, in one embodiment, a method is provided for analyzing long read sequence data from a subject to predict DPB1 expression levels in the subject, the method comprising: (a) obtaining a query nucleic acid sequence from a sample of interest using a long-read sequencer; (b) aligning a query nucleic acid sequence from the subject to a reference nucleic acid sequence using a computer-implemented alignment program, wherein the reference nucleic acid sequence and the query nucleic acid sequence comprise long-read sequence data for at least exon 3 of DPB1; (c) using a computer-implemented algorithm to identify, based on the aligned query nucleic acid sequence and the reference nucleic acid sequence, whether the query nucleic acid sequence has a sequence characteristic of low levels of DPB1 expression or high levels of DPBI expression, wherein the identifying step comprises: (i) comparing nucleotides in the aligned query nucleic acid sequence with a reference nucleic acid sequence to identify differences between the query nucleic acid sequence and the reference nucleic acid sequence; (ii) determining the identity of a nucleotide to the query nucleic acid sequence as compared to the reference nucleic acid sequence at a defined position in exon 3 of DPB1 based on the identified differences between the query nucleic acid sequence and the reference nucleic acid sequence; (iii) determining whether the query sequence exhibits a sequence characteristic of a weakly expressed motif, a sequence characteristic of a strongly expressed motif, or neither, based on the identity of the nucleotide to the query nucleic acid sequence at a defined position in exon 3; (iv) identifying the subject as having a low expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a weakly expressed motif, or identifying the subject as having a high expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a strongly expressed motif.
[0050] In one embodiment, the method further comprises determining the query sequence for the target.For example, the query sequence can be obtained from the long-read sequence data (before step (b)) that is generated as part of next-generation sequencing experiment, such as long-read sequencing.Or, other methods of nucleic acid sequencing can be used.
[0051] In various embodiments, at least some of the steps of the method are computer-implemented. For example, in one embodiment, steps (a), (b), (c), or any combination thereof are computer-implemented. In certain embodiments, the steps of the method are implemented using a long-read sequencer, a computer-implemented alignment program, or a computer-implemented algorithm, as described in detail herein. In some examples, the long-read sequencer can include a computer-implemented alignment program and / or a computer-implemented algorithm. In other examples, the long-read sequencer is separate from the computer-implemented alignment program and / or the computer-implemented algorithm, and the computer-implemented alignment program and / or the computer-implemented algorithm are implemented in one or more dedicated computing devices.
[0052] In one embodiment, the method further comprises identifying individual alleles for the DPB1 expression motif and / or other site of interest to the subject.
[0053] The sequence data may include data in addition to the sequence for exon 3 of DPB1. In one embodiment, the sequence data is long-read sequence data. The data may include genomic sequencing of the DPB1 gene. The data may further include genomic sequencing of the DRB1, DRB3, and / or DQB1 genes. In one embodiment, the sequence data does not include the 3'UTR of DPB1. Also, in one embodiment, the sequence data does not include the sequence of rs9277534. In one embodiment, the data is sequence data from any bone marrow registry, or data generated for submission to a registry or transplant center.
[0054] Alignment may be performed by a variety of different methods. In one embodiment, and in the examples herein, the alignment is performed using the program Minimap2 (Heng Li, Bioinformatics, 34(18), 2018, 3094-3100). Minimap2 is implemented in the C programming language and has both C and Python APIs. It is distributed under the MIT license and is free for both commercial and academic use. As known in the art, Minimap2 follows a seed-chain-align procedure typical of most whole-genome aligners. Essentially, the program collects minimizers for the reference sequence and indexes the minimizers into a hash table. Then, for each query sequence, using a value that is a list of the minimizer's hash and position, Minimap2 takes the query minimizer as a seed, finds an exact match (i.e., anchor) with the reference, and identifies a set of colinear anchors as a chain. For base-level alignments, minimap2 can apply dynamic programming (DP) to generate alignments extending from the ends of the chains and closing the regions between adjacent anchors within the chains (see, e.g., Heng Li, Bioinformatics, 34(18), 2018, 3094-3100). Other methods that can be used include BLASR (v1.MC.rc64; Chaisson and Tesler, 2012), BWA-MEM (v0.7.15; Li, 2013), GraphMap (v0.5.2; Sovic et al., 2016), Kart (v2.2.5; Lin and Hsu, 2017), minialign (v0.5.3; available on the web at github.com), and NGMLR (v0.2.5; Sedlazeck et al., 2018).
[0055] In one embodiment, the disclosed method may further include defining the weakly expressed motif as comprising the following: a G at position 20 of exon 3, a T at position 27 of exon 3, a T at position 52 of exon 3, a G at position 87 of exon 3, a T at position 234 of exon 3, a C at position 242 of exon 3, and a T at position 270 of exon 3. Additionally and / or alternatively, the disclosed method may further include defining the strongly expressed motif as comprising an A at position 20 of exon 3, a C at position 27 of exon 3, a C at position 52 of exon 3, an A at position 87 of exon 3, a C at position 234 of exon 3, a T at position 242 of exon 3, and a C at position 270 of exon 3. The method may further include determining whether the strongly expressed motif is linked to the rs9277534 G allele and / or whether the weakly expressed motif is linked to the rs9277534 A allele. In other embodiments, if the nucleotides at positions 20, 27, 52, 87, 234, 242, and 270 of exon 3 are not characteristic of either a weakly expressed motif or a strongly expressed motif, the disclosed methods can include defining the allele as indeterminate.
[0056] The disclosed methods may further include determining any differences between the query sequence and the reference sequence at other positions in DPB1 exon 3. For example, the reference sequence and / or query sequence may include sequence data for the entire DPB1 gene. Additionally and / or alternatively, the disclosed methods may further include determining any other differences between the query sequence and the reference sequence (e.g., other exons of DPB1 and / or other HLA genes of interest, such as HLA-A, -B, -C, DRB1, DRB3, and / or DQB1). For example, the reference sequence and query sequence may further include sequence data for at least one of DRB1, DRB3, and DQB1. The method may also include determining whether other variants in DPB1 exon 3 and / or other regions of the DPB1 gene are linked to the rs9277534 G allele associated with strong expression and / or the rs9277534 A allele associated with weak expression. In one embodiment, this additional data may be submitted to a caregiver or transplant database. In this way, additional variants associated with DPB1 expression can be evaluated for use as indicators of DPB1 expression levels.
[0057] In one embodiment, additional sequence data is compiled with the sequence of DPB1 exon 3. For example, linking other sequences to the activity profile of DPB1 exon 3 may be useful to further characterize the basis of GvHD in transplant recipients.
[0058] In some embodiments, if the number of differences between the query sequence and the reference sequence at positions other than positions 20, 27, 52, 87, 234, 242, and 270 in DPB1 exon 3 is greater than a predetermined number (e.g., 10), the query sequence is excluded from further analysis. For example, this may be the case when the sequence being analyzed is not actually DPB1, but another sequence present in the NGS data.
[0059] In certain embodiments, the subject is a potential donor for a transplant recipient. In one embodiment, the recipient is a hematopoietic stem cell transplant (HSCT) recipient. The analysis can be used to match potential donors to transplant recipients. Thus, in certain embodiments, the method may further include providing the results to a caregiver to reduce the risk of graft-versus-host disease in the transplant recipient. Alternatively, the method may further include providing the sequence analysis to a database for future distribution to caregivers and / or potential recipients. In certain embodiments, the recipient is a hematopoietic stem cell transplant (HSCT) recipient.
[0060] Thus, the disclosed method can include determining whether a query exhibits a sequence characteristic of a weakly expressed motif, a strongly expressed motif, or neither, thereby assessing the expression level of DPB1 in a subject. Figure 1 shows a comparison of nucleotides at specific positions in DPB1 cDNA. By subtracting the length of exons 1 and 2 (354 nt), the positions for each of these nucleotides in exon 3 are listed (i.e., 20, 26, 52, 87, 234, 242, and 270). Figure 1 shows sequence motifs associated with the rs9277534 A allele (weakly expressed) and / or the rs9277534 G allele (strongly expressed). For example, Figure 2 shows a comparison of exon 2 sequences for HLA haplotype 01:01:01:01, which has a strongly expressed DPB1 exon 3 motif of DPB1, compared with HLA haplotype 02:01:02:01, which has a weakly expressed DPB1 exon 3 motif of DPB1. Numbering is based on the cDNA sequence.
[0061] Figures 3 and 4 show an embodiment for comparing sequence data from two different query samples (Qry) with a reference sequence (Ref) encoding a weak genotype. In the figures, the + symbol indicates the positive (sense or coding) strand (i.e., the strand directly corresponding to the sequence of the mRNA transcript). Thus, Figure 3 illustrates the use of the disclosed algorithm to identify a strong exon 3 motif. The query (Qry+) sequence contains the sequence of a strongly expressed motif (solid rectangle) at positions 20, 27, 52, 87, 234, 242, and 270, as well as other variants (dashed rectangles). In this experiment, the reference sequence (Ref+) contains a weakly expressed motif. Figure 4 illustrates the use of the disclosed algorithm to identify a weak exon 3 motif. The query (Qry+) sequence contains the sequence of a weakly expressed motif (solid rectangle) at positions 20, 27, 52, 87, 234, 242, and 270, as well as other variants (dashed rectangles). In this figure, sequencing starts at the fourth nucleotide of exon 3. In this experiment, the reference sequence (Ref+) has a weakly expressed motif.
[0062] 5 illustrates one embodiment of a computer-implemented and / or algorithmic method 500 that can be used to analyze long sequence read data to predict DPB1 expression. The method can include obtaining 510 sequence data from a database or an HLA sequencing run. In one embodiment, a query nucleic acid sequence is obtained from a subject's sample using a long-read sequencer. Examples of long-read sequencers that can be used by various embodiments include the Pacific Biosciences RSII, Sequel, and Sequel II, and the Oxford Examples of suitable sequencing methods include Nanopore. Long-read sequencing data allows for unambiguous linkage between exons 2 and 3 of DPB1, minimizing phasing loss and the inability to link strong or weak motifs to the appropriate exon 2, making it more applicable to clinical samples of DBI expression. In one embodiment, the database is any bone marrow registry database used for donor selection based on HLA typing. In one embodiment, the query nucleic acid sequence contains long-read sequence data for at least exon 3 of DPB1. Additionally, other sequence data may be included (e.g., other exons of DPB1 and / or other HLA genes of interest, such as HLA-A, -B, -C, DRB1, DRB3, and / or DQB1).
[0063] The method may include step 512 of aligning the reference sequence and the query sequence and describing and / or comparing differences between the two sequences. In one embodiment, both the reference sequence and the query sequence comprise at least exon 3 of DPB1. The method may also include step 514 of determining the identity of a nucleotide to the query sequence at a defined position in exon 3 of DPB1.
[0064] The method may then include determining whether the query exhibits a sequence characteristic of a weakly expressed motif, a strongly expressed motif, or neither, thereby assessing the expression level of the subject's DPB1. Thus, if the query has the sequence of G at position 20 of exon 3, T at position 27 of exon 3, T at position 52 of exon 3, G at position 87 of exon 3, T at position 234 of exon 3, C at position 242 of exon 3, and T at position 270 of exon 3, it is defined as a weakly expressed motif 516. If the query has the sequence of A at position 20 of exon 3, C at position 27 of exon 3, C at position 52 of exon 3, A at position 87 of exon 3, C at position 234 of exon 3, T at position 242 of exon 3, and C at position 270 of exon 3, it is defined as a strongly expressed motif 518. However, if the nucleotides at positions 20, 27, 52, 87, 234, 242 and 270 of exon 3 are not characteristic of either the weakly or strongly expressed motif, the method may include defining the allele as indeterminate 520.
[0065] The method may include the optional step 522 of determining whether there are variants in the query (compared to the reference) at other positions in DPB1 exon 3 or other exons of DPB1 and / or other sequences important to transplant compatibility. The method may include determining whether any of the other variants are linked to either a weak, strong, or uncertain motif 524. The method may further include determining whether a strongly expressed motif is linked to the rs9277534 G allele or a weakly expressed motif is linked to the rs9277534 A allele (not shown in FIG. 5).
[0066] If the number of differences between the query sequence and the reference sequence at positions other than exon 3 positions 20, 27, 52, 87, 234, 242 and 270 in DPB1 exon 3 is greater than a predetermined cutoff (e.g., 5, or 7, or 8, or 10, or 15, or 20, or 25, or 30, or more), the query sequence may be excluded from further analysis 526. For example, this may be the case when the sequence being analyzed is not actually DPB1, but another sequence present in the sequence data.
[0067] The method may include a final step of outputting the results of the analysis. The result is a prediction of DPB1 expression, which may be used in downstream processing or analysis to match potential donors to transplant recipients. Thus, in certain embodiments, the method may further include providing the results to a database or a caregiver 528 to reduce the risk of graft-versus-host disease in transplant recipients. In certain embodiments, the recipient is a hematopoietic stem cell transplant (HSCT) recipient. In certain embodiments, the method may further include using the results of the analysis to match one or more potential donors to one or more transplant recipients. Advantageously, the techniques described herein for analyzing long sequence read data to ultimately predict DPB1 expression reduce computation time, provide more accurate predictions of DPB1 expression, and reduce clinical risks caused by errors in computing / analyzing long sequence read data.
[0068] FIG. 6 shows an example of another embodiment of method 600 of the present disclosure. Thus, as shown in FIG. 6, these steps may include obtaining sequence data from a database or performing HLA sequencing prior to submission to database 610. In one embodiment, a query nucleic acid sequence 605 is obtained from a subject's sample using a long-read sequencer. In one embodiment, the database is any bone marrow registry database used for donor selection based on HLA typing. The next several steps may involve characterizing the query sequence 605. Thus, the method may include step 612 of aligning the reference sequence with the query sequence. Next, the method may include comparing 614 the differences between the reference sequence and the query sequence. The method may further include step 616 of determining the identity of the query sequence at a defined position in DPB1 exon 3 and, optionally, other sequences in other exons of DPB1 and / or other HLA genes of interest, such as HLA-A, -B, -C, DRB1, DRB3, and / or DQB1. The method may also include a step 618 of compiling variants in the query sequence within the exon 3 strong-weak motif and, optionally, outside the exon 3 strong-weak motif. If there are more than a predetermined number of variants (e.g., 10), the sequence may be "dropped" from the analysis 619. In some embodiments, this results in the sequence not being considered predictive of whether the donor will be a suitable donor for the recipient at issue.
[0069] However, if there are fewer than 10 variants outside the strong-weak motif positions, the sequence can be further analyzed 620. This can include identifying the query sequence as having an exon 3 motif that is either weak, strong, or indeterminate 622. In one embodiment, this determination is based on the sequence at positions 20, 27, 52, 87, 234, 242, and 270 of exon 3 as described herein. Analysis can also determine whether other variants in exon 3 or other DPB1 regions are associated with either the weak or strong (or indeterminate) motif (i.e., indicative of genetic linkage) 624.
[0070] At this point, further analysis can be performed 625. For example, the sequence of the exon 3 motif (i.e., strong, weak, or indeterminate) may be linked to the sequence within rs9277534, with the "A" allele hypothesized to be associated with "weak" expression and the "G" allele hypothesized to be associated with strong expression. Thus, the data can be used to determine the sequence of rs9277534. The association of A / G with weak and / or strong expression, respectively, can be confirmed and / or the genotyping analysis can be further refined (i.e., by discovering other variants associated with the rs9277534 A and G alleles, respectively) 630 .
[0071] Finally, the results can be submitted to a third party 635. For example, the results may be added to a database 638 and / or provided to a caregiver 640. In one embodiment, the database may be the same database from which the query (and / or reference) sequence was obtained. Alternatively, the database may be a separate database. In some embodiments, the caregiver may be a physician wishing to find a donor match for a particular recipient.
[0072] Methods for reporting the analysis are also disclosed herein. In one embodiment, the Concise Idiosyncratic Gapped Alignment Report (CIGAR) method is used. In this analysis, a sequence is analyzed from 5' to 3' using a predetermined starting site. Variants are scored based on their nucleotide identity (e.g., A, C, G, or T), the nature of the variation (insertion, deletion, base change), and the location of the variant within the sequence region (e.g., contig). For example, as described above, a strong exon 3 motif for a query sequence can be reported based on the nucleotide identity at exon positions 20, 27, 52, 87, 234, 242, and 270. Alternatively, variants can be identified based on their location in cDNA or genome. Such presentation can be based on the absolute identity of the nucleotides at these positions, or any changes compared to the reference sequence. Thus, the sequence shown in Figure 4 for the weak motif can be represented as having the following sequences at the positions of interest: 20:g, 27:t, 52:t, 87:g, 234:t, 242:c, and 270:t (i.e., a non-CIGAR approach). Alternatively, if the weak motif sequence is the reference, the sequence can be represented as 20:gg, 27:tt, 52:tt, 87:gg, 234:tt, 242:cc, and 270:tt, indicating that there is no variant at the motif position of exon 3 in the reference (first nucleotide of the pair) compared to the query (second nucleotide of the pair). Similarly, the strong sequence in Figure 3 can be represented as 20:ga, 27:tc, 52:tc, 87:ga, 234:tc, 242:ct, and 270:tc (a non-CIGAR approach).
[0073] However, using the CIGAR method, motifs can be reported as any change between a reference sequence and a query sequence, regardless of the actual numbering used to identify the position of a nucleotide in a particular exon, cDNA, and / or genome sequence. This can facilitate the rapid comparison of contigs of sequence data. In one embodiment, the cs CIGAR tag encodes the sequence differences between the entire query sequence and the reference sequence in the short format or long format (the format in which reference sequence data is generally provided).
[0074] In one embodiment, the scoring method is as outlined in Table 1. [Table 1]
[0075] For example, the sequence alignment shown in Table 2 for SEQ ID NO:9 (top sequence) and SEQ ID NO:10 (bottom sequence) is represented as 6-ata:10+gtc:4*at:3, where :[0-9]+ represents an identical block, -ata represents a deletion, +gtc represents an insertion, and *at indicates that the reference base "a" is replaced with "t". [Table 2]
[0076] The algorithm can then analyze the CIGAR to determine whether a strong or weak motif is present in the query sequence. Examples of strong and weak sequences reported using the CIGAR format are shown in Figures 7 and 8, respectively. Thus, for the strong motif 3*cg:12*ga:6*tc:24*tc:34*ga:146*tc:7*ct:27*tc:32 (Figure 7), the CIGAR encoding is as follows: 3 - the number of bases the contig matches with the reference before the variant; starting from the beginning of the reference; cg - the variant is g instead of c; 12 - the number of bases the contig matches before the next variant; the ga variable is a instead of g, and so on. The final number ("32") is the number of bases the contig matches with the reference before the end of the reference. For the weak motif in Figure 8 compared to the weak variant, the CIGAR scoring is cs:z::302, indicating that the query and reference are identical over 302 nucleotides of exon 3. In one embodiment, a sequence may contain an overall match with either the weak motif or the long motif. An example of a weak motif (cs:Z::242*ct:59) with a mismatch at position 243 (i.e., one position removed from motif position 242) is shown in Figures 9 and 10.
[0077] A system for predicting DPB1 expression levels by analyzing long-read sequence data Embodiments of the present disclosure include computerized systems and computer program products for analyzing long read sequence data to predict DPB1 expression levels.
[0078] For example, disclosed is a non-transitory computer-readable storage medium comprising one or more data processors and instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations including analyzing sequence data from a subject to predict DPB1 expression levels in the subject, the operations including: (a) obtaining a query nucleic acid sequence from a sample from the subject using a long-read sequencer; (b) aligning the query nucleic acid sequence from the subject to a reference nucleic acid sequence using a computer-implemented alignment program, wherein the reference nucleic acid sequence and the query nucleic acid sequence comprise long-read sequence data for at least exon 3 of DPB1; and (c) identifying, based on the aligned query nucleic acid sequence and reference nucleic acid sequence, whether the query nucleic acid sequence has a sequence characteristic of low levels of DPB1 expression or high levels of DPB1 expression using a computer-implemented algorithm. and a non-transitory computer-readable storage medium, wherein the steps of comparing and identifying include: (i) comparing nucleotides in the aligned query nucleic acid sequence and reference nucleic acid sequence to identify differences between the query sequence nucleic acid sequence and the reference nucleic acid sequence; (ii) determining the identity of the nucleotide to the query nucleic acid sequence when compared to the reference nucleic acid sequence at a specified position in exon 3 of DPB1 based on the identified differences between the query sequence nucleic acid sequence and the reference nucleic acid sequence; (iii) determining whether the query sequence exhibits a sequence characteristic of a weakly expressed motif, a sequence characteristic of a strongly expressed motif, or neither based on the identity of the nucleotide to the query nucleic acid sequence at the specified position in exon 3; and (iv) identifying the subject as having a low expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a weakly expressed motif, or identifying the subject as having a high expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a strongly expressed motif.
[0079] Also disclosed is a computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform operations including analyzing sequence data from a subject to predict DPB1 expression levels in the subject, the computer-implemented steps of: (a) obtaining a query nucleic acid sequence from a sample of the subject using a long-read sequencer; (b) aligning the query nucleic acid sequence from the subject to a reference nucleic acid sequence using a computer-implemented alignment program, wherein the reference nucleic acid sequence and the query nucleic acid sequence comprise long-read sequence data for at least exon 3 of DPB1; and (c) identifying, based on the aligned query nucleic acid sequence and reference nucleic acid sequence, whether the query nucleic acid sequence has a sequence characteristic of low levels of DPB1 expression or high levels of DPB1 expression using a computer-implemented algorithm. and the identifying step comprises: (i) comparing nucleotides in the aligned query nucleic acid sequence and reference nucleic acid sequence to identify differences between the query sequence nucleic acid sequence and the reference nucleic acid sequence; (ii) determining the identity of the nucleotide to the query nucleic acid sequence when compared to the reference nucleic acid sequence at a specified position in exon 3 of DPB1 based on the identified differences between the query sequence nucleic acid sequence and the reference nucleic acid sequence; (iii) determining whether the query sequence exhibits a sequence characteristic of a weakly expressed motif, a sequence characteristic of a strongly expressed motif, or neither based on the identity of the nucleotide to the query nucleic acid sequence at the specified position in exon 3; and (iv) identifying the subject as having a low expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a weakly expressed motif, or identifying the subject as having a high expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a strongly expressed motif.
[0080] Systems and computer products can perform any of the methods disclosed herein. One or more embodiments described herein can be implemented using program modules, engines, or components. A program module, engine, or component can include a program, a subroutine, a portion of a program, or a software or hardware component that can perform one or more described tasks or functions. As used herein, a module or component can exist on a hardware component independent of other modules or components. Alternatively, a module or component can be a shared element or process of other modules, programs, or machines.
[0081] Figure 11 shows a block diagram of a DPB1 sequence analysis system. As shown in Figure 11, modules, engines, or components (e.g., programs, codes, or instructions) executable by one or more processors can be used to implement various subsystems of the analyzer system according to various embodiments. The modules, engines, or components may be stored on non-transitory computer media. As needed, one or more of the modules, engines, or components can be loaded into system memory (e.g., RAM) and executed by one or more processors of the analyzer system. In the example shown in Figure 11, modules, engines, or components for implementing the methods of the present disclosure are shown.
[0082] Accordingly, Figure 11 illustrates an exemplary computing device 1100 suitable for use in systems and methods according to the present disclosure. The exemplary computing device 1100 includes a processor 1105 that communicates with memory 1110 and other components of the computing device 1100 using one or more communication buses 1115. The processor 1105 is configured to execute processor-executable instructions stored in memory 1110 to perform one or more methods for assessing DPB1 expression levels according to different examples, such as some or all of the exemplary processes 500 or 600 described above with respect to Figures 5 and 6. In this example, the memory 1110 stores processor-executable instructions that provide DPB1 sequence analysis 1120 and expression determination 1125, as discussed herein.
[0083] The computing device 1100 in this example also includes one or more user input devices 1130, such as a keyboard, mouse, touchscreen, microphone, etc., for accepting user input. The computing device 1100 also includes a display 1135 for providing visual output to a user, such as a user interface. The computing device 1100 also includes a communication interface 1140. In some examples, the communication interface 1140 may enable communication using one or more networks, including a local area network ("LAN"), a wide area network ("WAN") such as the Internet, a metropolitan area network ("MAN"), point-to-point or peer-to-peer connections, etc. Communication with other devices may be achieved using any suitable network protocol. For example, one suitable network protocol may include the Internet Protocol ("IP"), Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), or a combination thereof, such as TCP / IP or UDP / IP. [Example]
[0084] Example 1 - Identification of strong and weak motifs The disclosed method has been used to characterize de-identified DNA sequences (i.e., sequences from which personal, registry, or other identifying information has been removed so that sequences are identified by random numbers) from a compilation of HLA typing production data from a bone marrow donor registry. Sequence data from 17,801 previously typed individuals, stored in one FASTQ file per individual, was analyzed with this program. All FASTQ files containing the exact seven-base sequence predicting a G at rs9277534 were typed as having "strong" expression. All FASTQ files containing the exact seven-base sequence predicting an A at rs9277534 were then typed as having "weak" expression. Any sequences that did not exactly match any motif were typed as "undetermined." As shown in Figure 15, a total of 31,274 DPB1 sequences were analyzed, of which 10,774 were typed as "strong," 20,369 were typed as "weak," and 131 were typed as "undetermined." The quality score criteria for failure was 30 (0.001%). Only 0.42% of all HLA-DPB1 sequences tested were marked "undefined," while the remaining 99.58% showed the correct sequence required to predict expression. Additionally, other variants were observed. The quality score criteria for failure was 30 (0.001%). All sequences were mapped to the exon 3 weak reference. The plot in Figure 16 shows the exon 3 locations where unique variants occur relative to the reference, downsampled by approximately 60 samples per strong, weak, and undefined. Vertical lines indicate predicted locations.
[0085] The results for sequences with strong motifs are shown in Figure 3. The query (Qry+) sequence (SEQ ID NO: 4) was found to have the exon 3 sequence of the DPB1 strongly expressed motif (solid rectangle) at positions 20, 27, 52, 87, 234, 242, and 270. Further characterization of the sequence showed that the query had other variants compared to the reference at other positions in the exon (dashed rectangle). In this figure, sequencing begins at the fourth nucleotide of exon 3, and the reference sequence (Ref+) (SEQ ID NO: 3) has a weakly expressed motif.
[0086] Figure 4 shows the use of the disclosed algorithm to identify exon 3 motifs. The query (Qry+) sequence has the sequence of the weakly expressed motif (solid rectangle) at positions 20, 27, 52, 87, 234, 242, and 270, as well as other variations (dashed rectangle). In this figure, sequencing begins at the fourth nucleotide of exon 3. In this experiment, the reference sequence (Ref+) has the weakly expressed motif.
[0087] Figures 7 and 8 show similar experiments performed using different query sequences and a reference sequence (SEQ ID NO: 3) with a weak motif in exon 3. As shown in Figure 7, the query sequence (SEQ ID NO: 6) displays a variant characteristic of the strong motif and is reported in CIGAR notation as cs:Z::38cg:12*ga:6*tc:24*tc:34*ga:146*tc:7*ct:27*tc:32. No additional variants are present. Interestingly, this and other sequence analyses confirmed that the variant at position 242 is a T, not an A as originally reported (Schone et al., 2018). The query sequence (SEQ ID NO: 7) shown in Figure 8 is defined as weakly expressed based on its identity to the reference sequence (SEQ ID NO: 3) containing the "weakly expressed" motif and can be reported as cs:Z::302.
[0088] Figure 9 presents a CIGAR-format summary of five samples characterized as having a strongly expressed motif and five samples with a "weakly" expressed motif. One weakly expressed sample has a T mutation at cDNA position 243, shown as the query (Qry-) sequence in Figure 10 (SEQ ID NO: 8). Additional sequence results showing sequences with strong motifs (with and without additional variants) and weak motifs (with and without additional variants), as well as sequences classified as undetermined, are shown in Figures 12 and 13. This sequence data can then be compiled and analyzed so that information regarding DPB1 expression can be evaluated in conjunction with other HLA typing data to pair potential hematopoietic stem cell donors and recipients.
[0089] A printout of the sequence analysis is shown in FIG.
[0090] In summary, the data indicate that the disclosed HLA-DPB1 expression prediction program can be a very useful tool to aid in the selection of donors for patients in need of hematopoietic stem cell transplantation.
[0091] Example 2 - Embodiment A1. A method for analyzing sequence data from a subject to predict DPB1 expression levels in the subject, comprising: (a) aligning a query sequence from a subject to a reference sequence, wherein the reference sequence and the query sequence comprise at least exon 3 of DPB1; (b) comparing the differences between the query sequence and the reference sequence; (c) determining the nucleotide identity to the query sequence at a defined position in exon 3 of DPB1; and (d) determining whether the query exhibits a sequence characteristic of a weakly expressed motif, a strongly expressed motif, or neither, thereby assessing the expression level of DPB1 in the subject.
[0092] A2. The method of any of the preceding or following embodiments, wherein at least one of steps (a)-(d) is computer-implemented.
[0093] A3. A computer-implemented method for analyzing sequence data from a subject to predict DPB1 expression levels in the subject, comprising: (a) obtaining a query nucleic acid sequence from a sample of interest using a long-read sequencer; (b) aligning a query nucleic acid sequence from the subject to a reference nucleic acid sequence using a computer-implemented alignment program, wherein the reference nucleic acid sequence and the query nucleic acid sequence comprise long-read sequence data for at least exon 3 of DPB1; (c) using a computer-implemented algorithm to identify, based on the aligned query nucleic acid sequence and the reference nucleic acid sequence, whether the query nucleic acid sequence has a sequence characteristic of low levels of DPB1 expression or high levels of DPBI expression, wherein the identifying step comprises: (i) comparing nucleotides in the aligned query nucleic acid sequence with a reference nucleic acid sequence to identify differences between the query nucleic acid sequence and the reference nucleic acid sequence; (ii) determining the identity of a nucleotide to the query nucleic acid sequence as compared to the reference nucleic acid sequence at a defined position in exon 3 of DPB1 based on the identified differences between the query nucleic acid sequence and the reference nucleic acid sequence; (iii) determining whether the query sequence exhibits a sequence characteristic of a weakly expressed motif, a sequence characteristic of a strongly expressed motif, or neither, based on the identity of the nucleotide to the query nucleic acid sequence at a defined position in exon 3; (iv) identifying the subject as having a low expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a weakly expressed motif, or identifying the subject as having a high expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a strongly expressed motif.
[0094] A4. The method of any of the preceding or following embodiments, wherein the defined positions of exon 3 are at positions 20, 27, 52, 87, 234, 242 and 270 from the 5' end of the exon.
[0095] A5. The method of any of the preceding or following embodiments, further comprising defining the weakly expressed motif as comprising a G at position 20 of exon 3, a T at position 27 of exon 3, a T at position 52 of exon 3, a G at position 87 of exon 3, a T at position 234 of exon 3, a C at position 242 of exon 3, and a T at position 270 of exon 3.
[0096] A6. The method of any of the preceding or following embodiments, further comprising defining the highly expressed motif as comprising an A at position 20 of exon 3, a C at position 27 of exon 3, a C at position 52 of exon 3, an A at position 87 of exon 3, a C at position 234 of exon 3, a T at position 242 of exon 3, and a C at position 270 of exon 3.
[0097] A7. The method of any of the preceding or following embodiments, wherein the allele is identified as indeterminate if the nucleotides at positions 20, 27, 52, 87, 234, 242 and 270 in exon 3 are not characteristic of either a weakly expressed motif or a strongly expressed motif.
[0098] A8. The method of any of the preceding or following embodiments, wherein identifying further comprises comparing nucleotides in the aligned query and reference nucleic acid sequences to identify the presence of any differences between the query and reference nucleic acid sequences at positions in exon 3 other than the defined position.
[0099] A9. The method of any of the preceding or following embodiments, wherein identifying further comprises comparing nucleotides in the aligned query and reference nucleic acid sequences to identify the presence of any other differences between the query and reference nucleic acid sequences.
[0100] A10. The method of any of the preceding or following embodiments, wherein if the number of differences between the query nucleic acid sequence and the reference nucleic acid sequence at positions other than the defined position in exon 3 is greater than 10, the query sequence is excluded from further analysis.
[0101] A11. The method of any of the preceding or following embodiments, further comprising determining, optionally by computer-implemented linkage analysis, that the strongly expressed motif in the query nucleic acid sequence is linked to the rs9277534 G allele and / or that the weakly expressed motif in the query nucleic acid sequence is linked to the rs9277534 A allele.
[0102] A12. The method of any of the preceding or following embodiments, wherein the reference nucleic acid sequence and / or query nucleic acid sequence comprises long-read sequence data for the entire DPB1 gene.
[0103] A13. The method of any of the preceding or following embodiments, wherein the reference nucleic acid sequence and / or query nucleic acid sequence further comprises long read sequence data of at least one of DRB1, DRB3, DQB1, DRB4, DRB5, DQA1, or DPA1.
[0104] A14. The method of any of the preceding or following embodiments, further comprising providing the results to a caregiver and / or transplant database to reduce the risk of graft-versus-host disease in the transplant recipient.
[0105] A15. The method of any of the preceding or following embodiments, wherein if the number of differences between the query nucleic acid sequence and the reference nucleic acid sequence at positions other than the defined position in exon 3 exceeds 10, then no results for the query sequence are provided to the caregiver or database.
[0106] A16. The method of any of the preceding or following embodiments, wherein the subject is a potential donor to a hematopoietic stem cell transplant (HSCT) recipient.
[0107] A17. The method of any of the preceding or following embodiments, further comprising determining a query nucleic acid sequence for a subject.
[0108] A18. The method of any of the preceding or following embodiments, wherein the query sequence is obtained from long-read or short-read sequence data generated as part of a next-generation sequencing experiment (prior to step (a)).
[0109] B1. A system comprising: one or more data processors; A non-transitory computer-readable storage medium comprising instructions that, when executed on one or more data processors, cause the one or more data processors to perform operations including analyzing sequence data from a subject to predict DPB1 expression levels in the subject; (a) aligning a query nucleic acid sequence from a subject to a reference nucleic acid sequence, wherein the reference sequence and the query sequence comprise at least exon 3 of DPB1; (b) comparing the differences between the query sequence and the reference sequence; (c) determining the nucleotide identity to the query sequence at a defined position in exon 3 of DPB1; (d) determining whether the query exhibits a sequence characteristic of a weakly expressed motif, a strongly expressed motif, or neither, thereby assessing the expression level of DPB1 in the subject.
[0110] B2. A system comprising: one or more data processors; A non-transitory computer-readable storage medium comprising instructions that, when executed on one or more data processors, cause the one or more data processors to perform operations including analyzing sequence data from a subject to predict DPB1 expression levels in the subject; (a) obtaining a query nucleic acid sequence from a sample of interest using a long-read sequencer; (b) aligning a query nucleic acid sequence from the subject to a reference nucleic acid sequence using a computer-implemented alignment program, wherein the reference nucleic acid sequence and the query nucleic acid sequence comprise long-read sequence data for at least exon 3 of DPB1; (c) using a computer-implemented algorithm to identify, based on the aligned query nucleic acid sequence and the reference nucleic acid sequence, whether the query nucleic acid sequence has a sequence characteristic of low levels of DPB1 expression or high levels of DPBI expression, wherein the identifying step comprises: (i) comparing nucleotides in the aligned query nucleic acid sequence with a reference nucleic acid sequence to identify differences between the query nucleic acid sequence and the reference nucleic acid sequence; (ii) determining the identity of a nucleotide to the query nucleic acid sequence as compared to the reference nucleic acid sequence at a defined position in exon 3 of DPB1 based on the identified differences between the query nucleic acid sequence and the reference nucleic acid sequence; (iii) determining whether the query sequence exhibits a sequence characteristic of a weakly expressed motif, a sequence characteristic of a strongly expressed motif, or neither, based on the identity of the nucleotide to the query nucleic acid sequence at a defined position in exon 3; (iv) identifying the subject as having a low expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a weakly expressed motif, or identifying the subject as having a high expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a strongly expressed motif, and a non-transitory computer-readable storage medium, the system comprising:
[0111] B3. The system of any of the preceding or following embodiments, wherein the operation further comprises defining the weakly expressed motif as comprising a G at position 20 of exon 3, a T at position 27 of exon 3, a T at position 52 of exon 3, a G at position 87 of exon 3, a T at position 234 of exon 3, a C at position 242 of exon 3, and a T at position 270 of exon 3.
[0112] B4. The system of any of the preceding or following embodiments, wherein the operation further comprises defining the highly expressed motif as comprising an A at position 20 of exon 3, a C at position 27 of exon 3, a C at position 52 of exon 3, an A at position 87 of exon 3, a C at position 234 of exon 3, a T at position 242 of exon 3, and a C at position 270 of exon 3.
[0113] B5. The system of any of the preceding or following embodiments, wherein if the nucleotides at positions 20, 27, 52, 87, 234, 242 and 270 in exon 3 are not characteristic of either a weakly expressed motif or a strongly expressed motif, the allele is defined as indeterminate.
[0114] B6. The system of any of the preceding or following embodiments, wherein identifying further comprises comparing nucleotides in the aligned query and reference nucleic acid sequences to identify the presence of any differences between the query and reference nucleic acid sequences at positions in exon 3 other than the defined position.
[0115] B7. The system of any of the preceding or following embodiments, wherein identifying further comprises comparing nucleotides in the aligned query and reference nucleic acid sequences to identify the presence of any other differences between the query and reference nucleic acid sequences.
[0116] B8. The system of any of the preceding or following embodiments, wherein if the number of differences between the query nucleic acid sequence and the reference nucleic acid sequence at positions other than the defined position in exon 3 is greater than 10, the query sequence is excluded from further analysis.
[0117] B9. The system of any of the preceding or following embodiments, wherein the operation further comprises determining whether the strongly expressed motif is linked to the rs9277534 G allele or the weakly expressed motif is linked to the rs9277534 A allele.
[0118] B10. The system of any of the preceding or following embodiments, wherein the operation further comprises determining, optionally by computer-implemented linkage analysis, that a strongly expressed motif in the query nucleic acid sequence is linked to the rs9277534 G allele and / or that a weakly expressed motif in the query nucleic acid sequence is linked to the rs9277534 A allele.
[0119] B11. The system of any of the preceding or following embodiments, wherein the reference and / or query sequence comprises long-read sequence data for the entire DPB1 gene.
[0120] B12. The system of any of the preceding or following embodiments, wherein the reference and / or query sequences further comprise long read sequence data of at least one of DRB1, DRB3, DQB1, DRB1, DRB3, DQB1, DRB4, DRB5, DQA1, or DPA1.
[0121] B13. The system of any of the preceding or following embodiments, wherein the operations further comprise providing the results to a caregiver and / or a database to reduce the risk of graft-versus-host disease in the transplant recipient.
[0122] B14. The system of any of the preceding or following embodiments, wherein if the number of differences between the query nucleic acid sequence and the reference nucleic acid sequence at positions other than the defined position in exon 3 exceeds 10, no results for the query sequence are provided to the caregiver or database.
[0123] B15. The system of any of the preceding or following embodiments, wherein the subject is a potential donor to a hematopoietic stem cell transplant (HSCT) recipient.
[0124] B16. The system of any of the preceding or following embodiments, wherein at least one of steps (a)-(d) is computer-implemented.
[0125] B17. The system of any preceding or following embodiment, wherein the operations further comprise determining a query sequence for the subject.
[0126] B18. The system of any of the preceding or following embodiments, wherein the query sequence is obtained from long-read or short-read sequence data generated as part of a sequencing experiment (prior to step (a)).
[0127] B19. The system of B18, wherein the sequencing experiment is a next-generation sequencing experiment.
[0128] B20. A system comprising: one or more data processors; a non-transitory computer-readable storage medium comprising instructions that, when executed on one or more data processors, cause the one or more data processors to perform operations to implement the method of any of the preceding embodiments.
[0129] C1. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to execute the system and / or perform the method described in any of the preceding or following embodiments.
[0130] C2 A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform operations including analyzing sequence data from a subject to predict DPB1 expression levels in the subject, (a) obtaining a query nucleic acid sequence from a sample of interest using a long-read sequencer; (b) aligning a query nucleic acid sequence from the subject to a reference nucleic acid sequence using a computer-implemented alignment program, wherein the reference nucleic acid sequence and the query nucleic acid sequence comprise long-read sequence data for at least exon 3 of DPB1; (c) using a computer-implemented algorithm to identify, based on the aligned query nucleic acid sequence and the reference nucleic acid sequence, whether the query nucleic acid sequence has a sequence characteristic of low levels of DPB1 expression or high levels of DPBI expression, wherein the identifying step comprises: (i) comparing nucleotides in the aligned query nucleic acid sequence with a reference nucleic acid sequence to identify differences between the query nucleic acid sequence and the reference nucleic acid sequence; (ii) determining the identity of a nucleotide to the query nucleic acid sequence as compared to the reference nucleic acid sequence at a defined position in exon 3 of DPB1 based on the identified differences between the query nucleic acid sequence and the reference nucleic acid sequence; (iii) determining whether the query sequence exhibits a sequence characteristic of a weakly expressed motif, a sequence characteristic of a strongly expressed motif, or neither, based on the identity of the nucleotide to the query nucleic acid sequence at a defined position in exon 3; (iv) identifying the subject as having a low expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a weakly expressed motif, or identifying the subject as having a high expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of a strongly expressed motif.
[0131] C3. The computer program product of any of the preceding or following embodiments, wherein the operation further comprises defining the weakly expressed motif as comprising a G at position 20 of exon 3, a T at position 27 of exon 3, a T at position 52 of exon 3, a G at position 87 of exon 3, a T at position 234 of exon 3, a C at position 242 of exon 3, and a T at position 270 of exon 3.
[0132] C4. The computer program product of any of the preceding or following embodiments, wherein the operation further comprises defining the strongly expressed motif as comprising an A at position 20 of exon 3, a C at position 27 of exon 3, a C at position 52 of exon 3, an A at position 87 of exon 3, a C at position 234 of exon 3, a T at position 242 of exon 3, and a C at position 270 of exon 3.
[0133] C5. The computer program product of any of the preceding or following embodiments, wherein identifying further comprises comparing nucleotides within the aligned query and reference nucleic acid sequences to identify the presence of any differences between the query and reference nucleic acid sequences at positions in exon 3 other than the defined position.
[0134] C6. The computer program product of any of the preceding or following embodiments, wherein identifying further comprises comparing nucleotides within the aligned query and reference nucleic acid sequences to identify the presence of any other differences between the query and reference nucleic acid sequences.
[0135] C7. The computer program product of any of the preceding or following embodiments, wherein if the number of differences between the query nucleic acid sequence and the reference nucleic acid sequence at positions other than the defined position in exon 3 is greater than 10, the query sequence is excluded from further analysis.
[0136] C8. The computer program product of any of the preceding or following embodiments, wherein the operation further comprises determining whether the strongly expressed motif is linked to the rs9277534 G allele or the weakly expressed motif is linked to the rs9277534 A allele.
[0137] C9. The computer program product of any of the preceding or following embodiments, wherein the operation further comprises determining, optionally by computer-implemented linkage analysis, that a strongly expressed motif in the query nucleic acid sequence is linked to the rs9277534 G allele and / or that a weakly expressed motif in the query nucleic acid sequence is linked to the rs9277534 A allele.
[0138] C10. The computer program product of any of the preceding or following embodiments, wherein the reference and / or query sequence comprises long-read sequence data for the entire DPB1 gene.
[0139] C11. The computer program product of any of the preceding or following embodiments, wherein the reference and / or query sequences further comprise long read sequence data of at least one of DRB1, DRB3, DQB1, DRB1, DRB3, DQB1, DRB4, DRB5, DQA1, or DPA1.
[0140] C12. The computer program product of any of the preceding or following embodiments, wherein the operations further include providing results to a caregiver and / or a database to reduce the risk of graft-versus-host disease in the transplant recipient.
[0141] C13. The computer program product of any of the preceding or following embodiments, wherein if the number of differences between the query nucleic acid sequence and the reference nucleic acid sequence at positions other than the defined position in exon 3 exceeds 10, then no results for the query sequence are provided to the caregiver or the database.
[0142] C14. The computer program product of any of the preceding or following embodiments, wherein the subject is a potential donor to a hematopoietic stem cell transplant (HSCT) recipient.
[0143] C15. The computer program product of any of the preceding or following embodiments, wherein at least one of steps (a)-(d) is computer-implemented.
[0144] C16. The computer program product of any of the preceding or following embodiments, wherein the operations further comprise determining a query nucleic acid sequence for the subject.
[0145] C17. The computer program product of any of the preceding or following embodiments, wherein the query sequence is obtained from long read or short read sequence data generated as part of a sequencing experiment (prior to step (a)).
[0146] C18. The computer program product of any of the preceding or following embodiments, wherein the sequencing experiment is a next-generation sequencing experiment.
[0147] Additional Considerations In the above description, specific details are set forth to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments can be practiced without these specific details. For example, circuits may be shown in block diagrams in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0148] The implementation of the techniques, blocks, steps, and means described above can be done in various ways. For example, these techniques, blocks, steps, and means can be implemented in hardware, software, or a combination thereof. In the case of a hardware implementation, the processing unit can be one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLCs), or other devices. The functions described above may be implemented in a programmable logic device (PLD), a field programmable gate array (FPGA), a processor, a controller, a microcontroller, a microprocessor, other electronic unit designed to perform the functions described above, and / or combinations thereof.
[0149] Also, it should be noted that the embodiments may be described as a process that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. While a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Moreover, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination corresponds to a return of the function to the calling function or the main function.
[0150] Furthermore, embodiments may be implemented by hardware, software, scripting languages, firmware, middleware, microcode, hardware description languages, and / or any combination thereof. When implemented in software, firmware, middleware, scripting languages, and / or microcode, the program code or code segments to perform the necessary tasks may be stored in a machine-readable medium such as a storage medium. A code segment or machine-executable instructions may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a script, a class, or any combination of instructions, data structures, and / or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, and / or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, ticket passing, network transmission, etc.
[0151] For a firmware and / or software implementation, the methodologies may be implemented with modules (e.g., procedures, functions, and so on) that perform the functions described herein. Any machine-readable medium tangibly embodying instructions may be used in practicing the methodologies described herein. For example, software code may be stored in memory. The memory may be implemented within the processor or external to the processor. As used herein, the term "memory" refers to any type of long-term, short-term, volatile, non-volatile, or other storage medium and is not limited to any particular type or number of memories or the type of medium on which the memory is stored.
[0152] Additionally, as disclosed herein, the terms "storage medium," "storage," or "memory" can refer to one or more memories for storing data, including read-only memory (ROM), random-access memory (RAM), magnetic RAM, core memory, magnetic disk storage media, optical storage media, flash memory devices, and / or other machine-readable media for storing information. The term "machine-readable medium" includes, but is not limited to, portable or non-removable storage devices, optical storage devices, wireless channels, and / or various other storage media capable of storing instructions and / or data that contain or carry the instructions and / or data.
[0153] While the principles of the present disclosure have been described above in connection with specific apparatus and methods, it is to be clearly understood that this description is made only by way of example and not as a limitation on the scope of the disclosure. The present invention provides, for example, the following items. (Item 1) 1. A computer-implemented method for analyzing long read sequence data from a subject to predict DPB1 expression levels in the subject, comprising: (a) a computer-implemented step of obtaining a query nucleic acid sequence from a sample of said subject using a long-read sequencer; (b) aligning the query nucleic acid sequence from the subject to a reference nucleic acid sequence using a computer-implemented alignment program, wherein the reference nucleic acid sequence and the query nucleic acid sequence comprise long-read sequence data for at least exon 3 of DPB1; (c) using a computer-implemented algorithm to determine whether the query nucleic acid sequence is associated with a low level of DPB1 expression, based on the aligned query nucleic acid sequence and the reference nucleic acid sequence; and a computer-implemented step of identifying whether the gene has a sequence characteristic of current or high levels of DPBI expression, wherein said computer-implemented step of identifying comprises: (i) comparing nucleotides in the aligned query nucleic acid sequence with the reference nucleic acid sequence to identify differences between the query nucleic acid sequence and the reference nucleic acid sequence; (ii) determining the identity of the nucleotide to the query nucleic acid sequence when compared to the reference nucleic acid sequence at a defined position in exon 3 of DPB1 based on the identified differences between the query nucleic acid sequence and the reference nucleic acid sequence; (iii) determining whether the query sequence exhibits a sequence characteristic of a weakly expressed motif, a sequence characteristic of a strongly expressed motif, or neither, based on the identity of the nucleotide at the specified position in exon 3 to the query nucleic acid sequence; (iv) identifying the subject as having a low expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of the weakly expressed motif, or identifying the subject as having a high expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of the strongly expressed motif. (Item 2) 2. The method of claim 1, wherein the defined positions in exon 3 are at positions 20, 27, 52, 87, 234, 242, and 270 from the 5' end of the exon. (Item 3) 3. The method of claim 1 or 2, further comprising defining the weakly expressed motif as comprising a G at position 20 of exon 3, a T at position 27 of exon 3, a T at position 52 of exon 3, a G at position 87 of exon 3, a T at position 234 of exon 3, a C at position 242 of exon 3, and a T at position 270 of exon 3. (Item 4) 3. The method of claim 1 or 2, further comprising defining the highly expressed motif as comprising an A at position 20 of exon 3, a C at position 27 of exon 3, a C at position 52 of exon 3, an A at position 87 of exon 3, a C at position 234 of exon 3, a T at position 242 of exon 3, and a C at position 270 of exon 3. (Item 5) 3. The method of claim 2, wherein if the nucleotides at positions 20, 27, 52, 87, 234, 242 and 270 of exon 3 are not characteristic of either the weakly expressed motif or the strongly expressed motif, the allele is identified as indeterminate. (Item 6) 2. The method of claim 1, wherein the identifying further comprises comparing nucleotides in the aligned query nucleic acid sequence and the reference nucleic acid sequence to identify the presence of any differences between the query nucleic acid sequence and the reference nucleic acid sequence at positions in exon 3 other than the specified position. (Item 7) 7. The method of claim 6, wherein the identifying further comprises comparing nucleotides in the aligned query and reference nucleic acid sequences to identify the presence of any other differences between the query and reference nucleic acid sequences. (Item 8) 2. The method of claim 1, wherein if the number of differences between the query nucleic acid sequence and the reference nucleic acid sequence at positions other than the defined position in exon 3 is greater than 10, the query sequence is excluded from further analysis. (Item 9) The strongly expressed motif in the query nucleic acid sequence is linked to the rs9277534 G allele and / or the weakly expressed motif in the query nucleic acid sequence is linked to the rs92 2. The method of claim 1, further comprising performing a computer-implemented linkage analysis to determine that the 77534 A allele is linked. (Item 10) 2. The method of claim 1, wherein the reference nucleic acid sequence and / or query nucleic acid sequence comprises long-read sequence data for the entire DPB1 gene. (Item 11) 10. The method of claim 9, wherein the reference nucleic acid sequence and / or the query nucleic acid sequence further comprises long-read sequence data for at least one of DRB1, DRB3, and DQB1. (Item 12) 10. The method of claim 1, further comprising providing the results to a caregiver and / or transplant database to reduce the risk of graft-versus-host disease in the transplant recipient. (Item 13) 13. The method of claim 12, wherein if the number of differences between the query nucleic acid sequence and the reference nucleic acid sequence at positions other than the defined positions in exon 3 exceeds 10, the result for the query sequence is not provided to a caregiver or database. (Item 14) 2. The method of item 1, wherein the subject is a potential donor for a hematopoietic stem cell transplant (HSCT) recipient. (Item 15) 1. A system comprising: one or more data processors; a non-transitory computer-readable storage medium comprising instructions that, when executed on the one or more data processors, cause the one or more data processors to perform operations including analyzing sequence data from a subject to predict DPB1 expression levels in the subject; (a) a computer-implemented step of obtaining a query nucleic acid sequence from a sample of said subject using a long-read sequencer; (b) aligning the query nucleic acid sequence from the subject to a reference nucleic acid sequence using a computer-implemented alignment program, wherein the reference nucleic acid sequence and the query nucleic acid sequence comprise long-read sequence data for at least exon 3 of DPB1; (c) using a computer-implemented algorithm to identify, based on the aligned query nucleic acid sequence and the reference nucleic acid sequence, whether the query nucleic acid sequence has a sequence characteristic of low levels of DPB1 expression or high levels of DPBI expression, wherein the computer-implemented identifying step comprises: (i) comparing nucleotides in the aligned query nucleic acid sequence with the reference nucleic acid sequence to identify differences between the query nucleic acid sequence and the reference nucleic acid sequence; (ii) determining the identity of the nucleotide to the query nucleic acid sequence when compared to the reference nucleic acid sequence at a defined position in exon 3 of DPB1 based on the identified differences between the query nucleic acid sequence and the reference nucleic acid sequence; (iii) determining whether the query sequence exhibits a sequence characteristic of a weakly expressed motif, a sequence characteristic of a strongly expressed motif, or neither, based on the identity of the nucleotide at a specified position in exon 3 to the query nucleic acid sequence; (iv) identifying the subject as having a low expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of the weakly expressed motif, or identifying the subject as having a high expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of the strongly expressed motif. (Item 16) 16. The system of claim 15, wherein the operation further comprises defining the weakly expressed motif as comprising a G at position 20 of exon 3, a T at position 27 of exon 3, a T at position 52 of exon 3, a G at position 87 of exon 3, a T at position 234 of exon 3, a C at position 242 of exon 3, and a T at position 270 of exon 3. (Item 17) 16. The system of claim 15, wherein the operation further comprises defining the highly expressed motif as comprising an A at position 20 of exon 3, a C at position 27 of exon 3, a C at position 52 of exon 3, an A at position 87 of exon 3, a C at position 234 of exon 3, a T at position 242 of exon 3, and a C at position 270 of exon 3. (Item 18) 1. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform operations including analyzing sequence data from a subject to predict DPB1 expression levels in said subject; (a) obtaining a query nucleic acid sequence from a sample of the subject using a long-read sequencer; (b) aligning the query nucleic acid sequence from the subject to a reference nucleic acid sequence using a computer-implemented alignment program, wherein the reference nucleic acid sequence and the query nucleic acid sequence comprise long-read sequence data for at least exon 3 of DPB1; and (b) using a computer-implemented algorithm to identify, based on the aligned query nucleic acid sequence and the reference nucleic acid sequence, whether the query nucleic acid sequence has a sequence characteristic of low levels of DPB1 expression or high levels of DPBI expression, wherein the identifying step comprises: (i) comparing nucleotides in the aligned query nucleic acid sequence with the reference nucleic acid sequence to identify differences between the query nucleic acid sequence and the reference nucleic acid sequence; (ii) determining the identity of the nucleotide to the query nucleic acid sequence when compared to the reference nucleic acid sequence at a defined position in exon 3 of DPB1 based on the identified differences between the query nucleic acid sequence and the reference nucleic acid sequence; (iii) determining whether the query sequence exhibits a sequence characteristic of a weakly expressed motif, a sequence characteristic of a strongly expressed motif, or neither, based on the identity of the nucleotide to the query nucleic acid sequence at a defined position in exon 3; (iv) identifying the subject as having a low expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of the weakly expressed motif, or identifying the subject as having a high expression level of DPB1 if the query nucleic acid sequence exhibits a sequence characteristic of the strongly expressed motif. (Item 19) 19. The computer program product of item 18, wherein the operations further comprise defining the weakly expressed motif as comprising a G at position 20 of exon 3, a T at position 27 of exon 3, a T at position 52 of exon 3, a G at position 87 of exon 3, a T at position 234 of exon 3, a C at position 242 of exon 3, and a T at position 270 of exon 3. (Item 20) 19. The computer program product of item 18, wherein the operations further comprise defining the highly expressed motif as comprising an A at position 20 of exon 3, a C at position 27 of exon 3, a C at position 52 of exon 3, an A at position 87 of exon 3, a C at position 234 of exon 3, a T at position 242 of exon 3, and a C at position 270 of exon 3.
Claims
[Claim 1] The invention described in the specification.
Citation Information
Patent Citations
Human leukocyte antigen genotyping methods and determination of HLA haplotype diversity in sample populations
JP2019530476A