Composition and method for detecting head and neck cancer
The method involves measuring the expression levels of specific genes in a sample to detect biomarkers for head and neck squamous cell carcinoma (HNSCC), addressing the need for reliable indicators and aiding in early diagnosis.
Patent Information
- Application Number
- JP2025053460
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-12-20
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-19
AI Technical Summary
There is a need for reliable biomarkers to detect head and neck squamous cell carcinoma (HNSCC) due to the unclear biological mechanisms underlying the disease and the limited availability of effective indicators.
A method for detecting biomarkers associated with HNSCC involves obtaining a sample from an individual and measuring the expression level of specific genes such as CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1, HSD17B6, and human papillomavirus (HPV) E6 or E7 genes, potentially using a panel of biomarkers and normalization genes like KHDRBS1 or RPL30.
This approach enables the detection of biomarkers associated with HNSCC, potentially identifying individuals at risk or with the disease, thereby aiding in early diagnosis and management.
Smart Images

Figure 2025092585000001_ABST
Abstract
Description
Technical Field
[0001] Related Applications This application claims priority based on U.S. Provisional Patent Application No. 62 / 608,296, filed on December 20, 2017. The disclosure of U.S. Provisional Patent Application No. 62 / 608,296 is hereby incorporated by reference in its entirety herein.
Background Art
[0002] Background Head and neck cancer is a common disease. Since most head and neck cancers histologically belong to the squamous cell type, they are classified as head and neck squamous cell carcinoma (HNSCC). HNSCC is the sixth most common cancer in the world and the third most common in developing countries.
[0003] The biological mechanisms underlying HNSCC are unclear, and there are only a few biomarkers, if any, that provide reliable indicators of this condition. Nevertheless, adjusting lifestyle to avoid the induction of symptom onset and / or the promotion of further disease progression would be helpful for individuals susceptible to HNSCC. Therefore, there is a need to develop and evaluate biomarkers for HNSCC.
Summary of the Invention
Means for Solving the Problems
[0004] Summary This disclosure can be embodied in various ways.
[0005] In one embodiment, a method for detecting biomarkers associated with head and neck squamous cell carcinoma (HNSCC) in an individual is disclosed, comprising the steps of obtaining a sample from the individual and measuring the expression level of at least one gene in Table 4 and / or Table 6 in the sample. In one embodiment, a method for detecting biomarkers associated with HNSCC in an individual is disclosed, comprising the steps of obtaining a sample from the individual and measuring the expression level of at least one of the following genes: CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1, and HSD17B6 in the sample. In another embodiment, a method for detecting biomarkers associated with HNSCC in an individual is disclosed, comprising the steps of obtaining a sample from the individual and measuring the expression level of at least one of at least one of the human papillomavirus (HPV) E6 or E7 genes. Further and / or alternatively, the method may include the measurement of at least one normalization (e.g., housekeeping) gene. In one embodiment, the normalization gene may be KHDRBS1. In one embodiment, the normalization gene may be RPL30 or another normalization gene. Alternatively, measurements of the expression of various combinations of these genes may be performed.
[0006] In one embodiment, a panel of a plurality of biomarkers is used. In one embodiment, the present disclosure provides a composition for detecting biomarkers associated with head and neck squamous cell carcinoma (HNSCC) in an individual, comprising a reagent for quantifying the expression level of at least one gene in Table 4 and / or Table 6, and / or at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6, and / or at least one of the HPV E6 and E7 genes. Further and / or alternatively, the composition may include at least one normalization (e.g., housekeeping) gene. In one embodiment, the normalization gene may be KHDRBS1. In one embodiment, the normalization gene may be RPL30 or another normalization gene. In certain embodiments, the composition can include primers and / or probes for any one of these genes, and the primers and / or probes are labeled with a detectable moiety described herein.
[0007] Other embodiments include systems for performing the methods disclosed herein and / or using the compositions disclosed herein.
[0008] Other features, objects, and advantages of the disclosure herein will be apparent in the following detailed description, drawings, and claims. However, it is to be understood that the detailed description, drawings, and claims are provided by way of illustration only and not limitation, showing embodiments of the disclosed methods, compositions, and systems. Various changes and modifications within the scope of the invention will be apparent to those skilled in the art. In certain embodiments, for example, the following are provided: (Item 1) A method for detecting biomarkers associated with head and neck squamous cell carcinoma (HNSCC) in an individual, comprising: obtaining a sample from the individual; and measuring the amount of an expression product from a gene comprising at least one gene in Table 4 and / or Table 6 A method comprising (Item 2) The method according to item 1, wherein the gene comprises at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6. (Item 3) The method according to item 1, further comprising measuring the amount of the expression product from at least one of the HPV E6 gene and / or the HPV E7 gene. (Item 4) The method according to item 1, further comprising measuring the amount of the normalization gene. (Item 5) The method according to item 1, wherein the measuring step comprises measuring mRNA. (Item 6) The method according to item 1, wherein the measuring step comprises an immunoassay. (Item 7) The method according to item 1, comprising measuring the expression of at least four of the genes. (Item 8) The method according to item 1, wherein the sample comprises serum, tissue, FFPE, saliva or plasma. (Item 9) A method for detecting susceptibility to head and neck squamous cell carcinoma (HNSCC) in an individual, comprising: obtaining a sample from the individual; measuring the amount of at least one expression product from at least one of the genes in Table 4 and / or Table 6; comparing the expression of at least one of the genes in Table 4 and / or Table 6 in the sample with a control value of the expression product of the gene, wherein a difference between the gene expression in the individual and the control value indicates that the individual may have HNSCC or is prone to developing HNSCC (i.e., has a high risk of HNSCC). (Item 10) The method according to item 9, wherein the gene comprises at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6. (Item 11) The method according to item 9, further comprising: measuring the amount of at least one expression product of at least one HPV E6 and / or E7 gene; and comparing the expression of the at least one HPV E6 and / or E7 gene in the sample with a control value of the expression. (Item 12) The method according to item 9, further comprising measuring the amount of the normalization gene. (Item 13) The method according to item 9, wherein the measuring comprises measuring mRNA. (Item 14) The method according to item 9, wherein the measuring comprises measuring protein. (Item 15) The method according to item 9, comprising measuring the expression of at least four of the genes. (Item 16) The method according to item 9, wherein the sample comprises serum tissue, FFPE, saliva or plasma. (Item 17) A composition for detecting biomarkers associated with head and neck squamous cell carcinoma (HNSCC) in an individual, comprising a reagent for quantifying the expression level of at least one gene in Table 4 and / or Table 6. (Item 18) The composition according to item 17, wherein the at least one gene comprises at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6. (Item 19) The composition according to item 17, further comprising at least one reagent for quantifying the expression level of at least one of HPV E6 and / or E7. (Item 20) The composition according to item 17, further comprising at least one normalization gene. (Item 21) The composition according to item 17, wherein the reagent detects mRNA. (Item 22) The composition according to item 17, wherein the reagent comprises at least one primer and / or probe for any one of these genes, and the at least one primer and / or probe is labeled with a detectable moiety. (Item 23) The composition according to item 17, wherein the reagent detects protein. (Item 24) A kit comprising a reagent for quantifying the expression level of at least one gene in Table 4 and / or Table 6.
Brief Description of the Drawings
[0009] The present invention can be better understood by referring to the following non-limiting drawings.
[0010]
Figure 1
[0011]
Figure 2
[0012]
Figure 3
[0013]
Figure 4
[0014]
Figure 5
[0015]
Figure 6A
Figure 6B
[0016]
Figure 7
[0017]
Figure 8
[0018]
Figure 9A
Figure 9B
[0019]
Figure 10
[0020]
Figure 11
[0021]
Figure 12
[0022]
Figure 13
[0023]
Figure 14
[0024]
Figure 15
[0025]
Figure 16
[0026]
Figure 17
[0027]
Figure 18
[0028]
Figure 19
[0029]
Figure 20
[0030]
Figure 21
[0031] **Detailed Description** **Terms and Definitions** To facilitate a better understanding of the present disclosure, certain terms are first defined. Further definitions of the following terms and other terms are set forth throughout the specification.
[0032] Numerical ranges and parameters setting forth broad ranges of the present disclosure are approximations, although the numerical values set forth in the specific examples are reported as accurately as possible. However, each numerical value inherently contains certain errors necessarily resulting from the standard deviation found in the respective test measurement. Further, all ranges disclosed herein are to be understood to encompass any and all subranges subsumed therein. For example, the recited range of “1 to 10” should be considered to include any and all subranges between the minimum value of 1 and the maximum value of 10 (including the maximum and minimum values); that is, all subranges starting with a minimum value of 1 or more (e.g., 1 to 6.1) and ending with a maximum value of 10 or less (e.g., 5.5 to 10). Further, any reference cited as “incorporated herein” is to be understood as being incorporated in its entirety.
[0033] It should further be noted that, as used herein, the singular forms “a,” “an,” and “the” include plural referents unless expressly and specifically limited to one referent. The term “and / or” is commonly used to refer to at least one of either. In some instances, the term “and / or” is used synonymously with the term “or.” The term “comprising” is used herein to mean “including, but not limited to,” and is used synonymously therewith. The term “such as” is used herein to mean “such as, but not limited to,” and is used synonymously therewith.
[0034] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Specialists are referred to the definitions and terms of the art, particularly as regards Current Protocols in Molecular Biology (Ausubel).
[0035] Antibody: As used herein, the term "antibody" refers to a polypeptide consisting of one or more polypeptides that substantially encode an immunoglobulin gene or a fragment of an immunoglobulin gene. The recognized immunoglobulin genes include the constant region genes of κ, λ, α, γ, δ, ε, and μ, as well as various immunoglobulin variable region genes. Light chains are generally classified as κ or λ. Heavy chains are generally classified as γ, μ, α, δ, or ε, which in turn define the classes of immunoglobulins, IgG, IgM, IgA, IgD, and IgE, respectively. It is known that the structural unit of a typical immunoglobulin (antibody) contains a tetramer. Each tetramer consists of two identical polypeptide chain pairs, and each pair has one "light" chain (about 25 kD) and one "heavy" chain (about 50 - 70 kD). The N-terminus of each chain defines a variable region of about 100 - 110 or more amino acids that is mainly responsible for antigen recognition. The terms "light chain variable chain" (VL) and "heavy chain variable chain" (VH) refer to these light and heavy chains, respectively. An antibody can be specific for a particular antigen. The antibody or its antigen can be either an analyte or a binding partner. Antibodies exist as intact immunoglobulins or as a number of well-characterized fragments produced by digestion with various peptidases. Thus, for example, pepsin digests the antibody under the disulfide bonds of the hinge region to produce F(ab)'2, which is a dimer of Fab, which is itself a light chain linked to VH-CH1 by a disulfide bond. F(ab)'2 can be reduced under mild conditions to cleave the disulfide bonds in the hinge region, whereby the (Fab')2 dimer can be converted to Fab' monomers. Fab' monomers are essentially Fab with a portion of the hinge region (for a more detailed description of other antibody fragments, see Fundamental Immunology, W.E. Paul, ed., Raven Press, N.Y. (1993)). Although various antibody fragments are defined with respect to the digestion of intact antibodies, one of ordinary skill in the art will understand that such Fab' fragments can be newly synthesized chemically or by using recombinant DNA methods.Accordingly, as used herein, the term "antibody" includes antibody fragments made by modification of whole antibodies or newly synthesized using recombinant DNA methods. In some embodiments, the antibody is a single-chain antibody, e.g., a single-chain Fv (scFv) antibody in which the heavy-chain variable region and the light-chain variable region are joined (either directly or via a peptide linker) to form a continuous polypeptide. A single-chain Fv ("scFv") polypeptide is a covalently linked VH::VL heterodimer and can be expressed from a nucleic acid comprising sequences encoding VH- and VL- that are either directly linked or joined by a peptide-encoding linker. (See, e.g., Huston et al., (1988) Proc. Nat. Acad. Sci. USA, 85:5879-5883, the entire contents of which are incorporated herein by reference). There are several structures for converting the naturally aggregated but chemically separated light and heavy polypeptide chains from antibody V regions into scFv molecules folded into a three-dimensional structure substantially similar to the structure of the antigen-binding site. See, e.g., U.S. Patent Nos. 5,091,513, 5,132,405, and 4,956,778.
[0036] The term "antibody" includes monoclonal antibodies, polyclonal antibodies, synthetic antibodies, and chimeric antibodies, e.g., those generated by combinatorial mutagenesis and phage display. The term "antibody" also includes antibody mimetics or peptide mimics. Peptide mimics are compounds based on or derived from peptides and proteins. The peptide mimics of the present disclosure can generally be obtained by structural modification of known peptide sequences using non-natural amino acids, conformational restriction, isoelectronic substitution, etc.
[0037] Allele: As used herein, the term "allele" refers to different versions of the nucleotide sequence of the same locus (e.g., gene).
[0038] Allele-Specific Primer Extension (ASPE): As used herein, the term "Allele-Specific Primer Extension (ASPE)" refers to a method for detecting mutations using a primer that hybridizes to a corresponding DNA sequence and is extended in response to successful hybridization of the 3'-terminal nucleotide of such primer. Generally, an extension primer having a 3'-terminal nucleotide that forms a perfect match with the target sequence is extended to form an extension product. Modified nucleotides can be incorporated into the extension product, and such nucleotides can effectively label the extension product for detection purposes. Alternatively, the extension primer may instead contain a 3'-terminal nucleotide that forms a mismatch with the target sequence. In this case, primer extension does not occur unless the polymerase used for extension inadvertently has exonuclease activity.
[0039] Amplification: As used herein, the term "amplification" refers to methods known in the art for copying a target nucleic acid and thereby increasing the copy number of a selected nucleic acid sequence. Amplification may be exponential or linear. The target nucleic acid may be either DNA or RNA. Generally, the sequences amplified by this method include "amplicons". Amplification may be achieved in various ways, including but not limited to polymerase chain reaction ("PCR"), transcription-based amplification, isothermal amplification, rolling circle amplification, etc. Amplification may be performed with relatively similar amounts of each primer of a primer pair to generate double-stranded amplicons. However, as is well known in the art, asymmetric PCR may be used to amplify mainly or exclusively single-stranded products (e.g., Poddar, Molec. And Cell. Probes 14:25-32 (2000)). This can be achieved by using each pair of primers and significantly reducing the concentration of one primer of the pair compared to the other (e.g., a 100-fold difference). Amplification by asymmetric PCR is usually linear. One of ordinary skill in the art will understand that different amplification methods may be used together.
[0040] Animal: As used herein, the term "animal" refers to any member of the animal kingdom. In some embodiments, "animal" refers to humans at any stage of development. In some embodiments, "animal" refers to non-human animals at any stage of development. In certain embodiments, the non-human animal is a mammal (e.g., rodents, mice, rats, rabbits, monkeys, dogs, cats, sheep, cows, primates, and / or pigs). In some embodiments, animals include, but are not limited to, mammals, birds, reptiles, amphibians, fish, insects, and / or worms. In some embodiments, the animal may be a transgenic animal, a genetically modified animal, and / or a clone.
[0041] Approximately: As used herein, the term "approximately" or "about" when applied to one or more values of interest refers to a value similar to the stated reference value. In certain embodiments, the term "approximately" or "about" refers to a range of values that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either (greater or lesser) direction from the stated reference value, unless otherwise specified or apparent from the context (so long as such numbers do not exceed 100% of the possible values). The term "about" is used to indicate that the variation in the inherent error of the device, the method used to determine the value, or the variation that exists between samples is included in the value.
[0042] Associated with a syndrome or disease of interest: As used herein, "associated with a syndrome or disease of interest" means that a variant is found more frequently in patients having the syndrome or disease of interest than in asymptomatic or non-disease controls. Generally, the statistical significance of such an association can be determined by assaying multiple patients.
[0043] Biological sample: As used herein, the term "biological sample" or "sample" encompasses any sample obtained from a biological origin. Biological samples can include, by way of non-limiting example, blood, amniotic fluid, serum, plasma, fluid or tissue biopsies, urine, feces, epithelial samples, skin samples, buccal swabs, sperm, amniotic fluid, cultured cells, bone marrow samples and / or chorionic villi. A convenient biological sample can be obtained, for example, by scraping cells from the surface of the buccal side of the oral cavity. The term biological sample encompasses samples that have been processed to release or make available nucleic acids or proteins for the detection described herein. The term biological sample also includes cell-free nucleic acids that may be present in a sample (e.g., plasma or amniotic fluid). For example, a biological sample may include cDNA obtained by reverse transcription of RNA from cells in the biological sample. Biological samples may be obtained from stages of life such as fetus, young adult, adult, etc. Fixed or frozen tissue may also be used.
[0044] Biomarker: As used herein, the term "biomarker" or "marker" refers to one or more nucleic acids, polypeptides and / or other biomolecules (e.g., cholesterol, lipids) that can be used, alone or in combination with other biomarkers, to diagnose a disease or syndrome of interest, or to assist in its diagnosis or prognosis; monitor the progression of a disease or syndrome of interest; and / or monitor the effectiveness of treatment of a syndrome or disease of interest.
[0045] Binder: As used herein, the term "binder" refers to a molecule that can specifically and selectively bind to a second (i.e., different) molecule of interest. The interaction may be non-covalent, for example, as a result of hydrogen bonding, van der Waals interactions, or electrostatic or hydrophobic interactions, or it may be covalent. The term "soluble binder" refers to a binder that is not bound to a solid support (i.e., is not covalently or non-covalently bound).
[0046] Carrier: The term "carrier" refers to a person who has no symptoms but carries a mutation that can be passed on to their children. Generally, with respect to autosomal recessive genetic disorders, a carrier has one allele that causes a disease-causing mutation and a second allele that is normal or not associated with the disease.
[0047] Coding sequence and non-coding sequence: As used herein, the term "coding sequence" refers to a sequence of nucleic acid or its complement, or a portion thereof, that can be transcribed and / or translated to produce mRNA and / or a polypeptide or a fragment thereof that can generate a polypeptide or a fragment thereof. A coding sequence contains exons in genomic DNA or immature primary RNA transcripts, which are joined together by the biochemical machinery of the cell to yield mature mRNA. The antisense strand is the complement of such nucleic acid, and the encoded sequence can be deduced therefrom. As used herein, the term "non-coding sequence" refers to a sequence of nucleic acid or its complement, or a portion thereof, that is not transcribed into an amino acid in vivo, or a sequence that does not interact or attempt to interact with tRNA to place an amino acid. Non-coding sequences include both intron sequences in genomic DNA or immature primary RNA transcripts, and gene-associated sequences such as promoters, enhancers, silencers, etc.
[0048] Complement: As used herein, the terms "complement", "complementary", and "complementarity" refer to the pairing of nucleotide sequences that follows the Watson / Crick pairing rules. For example, the sequence 5'-GCGGTCCCA-3' has a complementary sequence of 5'-TGGGACCGC-3'. The complementary sequence can also be an RNA sequence that is complementary to a DNA sequence. Certain bases not commonly found in natural nucleic acids may be included in complementary nucleic acids including, but not limited to, inosine, 7-deazaguanine, locked nucleic acid (LNA), and peptide nucleic acid (PNA). Complementarity need not be perfect; stable double-stranded molecules may contain mismatched base pairs, modified or unpaired bases. Those skilled in nucleic acid technology can empirically determine duplex stability by considering a number of variables, such as the length of the oligonucleotide, the base composition and sequence of the oligonucleotide, the ionic strength, and the incidence of mismatched base pairs.
[0049] Conserved: As used herein, the term "conserved residue" refers to an amino acid that is the same among multiple proteins having the same structure and / or function. Regions of conserved residues can be important for the structure or function of a protein. Thus, adjacent conserved residues identified in a three-dimensional protein structure can be important for the structure or function of the protein. To find conserved residues, or conserved regions of a 3D structure, sequence comparisons of the same or similar proteins from different species, or from individuals of the same species, can be performed.
[0050] Control: As used herein, the term "control" has the meaning understood in the art as a reference substance against which results are compared. Generally, a control is used to increase the consistency of an experiment by separating variables in order to draw conclusions about such variables. In some embodiments, the control is a reaction or assay that is performed concurrently with the test reaction or assay to obtain a comparison standard. In one experiment, a "test" (i.e., the variable being tested) is applied. In a second experiment, no "control," i.e., the variable being tested, is applied. In some embodiments, the control is a historical control (i.e., a previously performed test or assay, or a control of previously known amounts or results). In some embodiments, the control is or includes a printed or otherwise stored record. The control may be a positive control or a negative control.
[0051] "Control" or "predetermined standard" of a biomarker refers to the level of expression of the biomarker in a healthy subject, or the level of expression of said biomarker in non-diseased or asymptomatic tissue from the same subject. The control or predetermined standard expression level or amount of protein of a given biomarker can be established by prospective and / or retrospective statistical studies using only routine experimentation. Such predetermined standard expression levels and / or protein levels (amounts) can be determined using methods well known to those of ordinary skill in the art. A positive control is a sample (or reagent) that provides a measured signal of a predetermined amount.
[0052] Crude: As used herein, the term "crude," when used in relation to a biological sample, refers to a sample that is in a substantially unpurified state. For example, a crude sample can be a cell lysate or a biopsy tissue sample. The crude sample may be present in a lysed state or as a dried preparation.
[0053] Deletion: As used herein, the term "deletion" encompasses a mutation that removes one or more nucleotides from a naturally occurring nucleic acid.
[0054] Diseases or symptom groups of interest: As used herein, the disease or symptom group of interest is head and neck cancer, and in some embodiments, more specifically HNSCC.
[0055] Detection: As used herein, the terms "detect", "detected" or "detection" include "measure", "measured" or "measurement", and vice versa.
[0056] Detectable moiety: As used herein, the terms "detectable moiety" or "detectable biomolecule" or "reporter" refer to a molecule that can be measured in a quantitative assay. For example, a detectable moiety may include an enzyme used to convert a substrate into a measurable product (e.g., a visible product). Or, the detectable moiety may be a radioisotope that can be quantified. Or, the detectable moiety may be a fluorophore. Or, the detectable moiety may be a luminescent molecule. Or, other detectable molecules may be used.
[0057] Epigenetic: As used herein, epigenetic elements can alter gene expression by mechanisms other than changes in the underlying DNA sequence. Such elements can include paramutations, imprinting, gene silencing, X-chromosome inactivation, position effects, reprogramming, transvection, maternal effects, histone modifications, and elements that regulate heterochromatin.
[0058] Epitope: As used herein, the term "epitope" refers to a fragment or portion of a molecule or molecular complex (e.g., a polypeptide or protein complex) that contacts a specific antibody or antibody-like protein.
[0059] Exon: As used herein, an exon is a nucleic acid sequence found in the mature or processed RNA state after other portions of the RNA (such as intervening regions known as introns) have been removed by RNA splicing. Thus, exon sequences typically encode a protein or a portion of a protein. An intron is a portion of the RNA that is removed from the surrounding exon sequences by RNA splicing.
[0060] Expression and expressed RNA: As used herein, expressed RNA is RNA that encodes a protein or polypeptide (“coding RNA”), any other RNA that is transcribed but not translated (“non-coding RNA”). The term “expression” is used herein to mean the process by which a polypeptide is produced from DNA. This process involves transcription of the gene into mRNA and translation of this mRNA into a polypeptide. Depending on the context in which it is used, “expression” may refer to the production of RNA, protein, or both.
[0061] The measurement of the amount of protein and / or the expression of the biomarkers of the present disclosure can be evaluated by any of a wide variety of well-known methods for detecting the expression of transcribed molecules or their corresponding proteins. Non-limiting examples of such methods include immunological methods for detecting secreted proteins, protein purification methods, protein function or activity assays, nucleic acid hybridization methods, nucleic acid reverse transcription methods, and nucleic acid amplification methods. In certain embodiments, the expression of a marker gene is evaluated using an antibody (e.g., a radiolabeled, chromophore-labeled, fluorophore-labeled, or enzyme-labeled antibody) that specifically binds to a protein corresponding to the marker gene, such as a protein encoded by an open reading frame corresponding to the marker gene, or a protein that has undergone all or part of its normal post-translational modification, an antibody derivative (e.g., an antibody complexed with a substrate or with a protein or ligand of a protein-ligand pair {e.g., biotin-streptavidin}), or an antibody fragment (e.g., a single-chain antibody, an isolated antibody hypervariable domain, etc.). In certain embodiments, the reagent may be labeled directly or indirectly with a detectable substance. The detectable substance may be selected, for example, from the group consisting of radioisotopes, fluorescent compounds, enzymes, and enzyme cofactors. Methods for labeling antibodies are well known in the art.
[0062] In another embodiment, the expression of a marker gene is evaluated by preparing mRNA / cDNA (i.e., transcribed polynucleotides) from cells in a sample and hybridizing the mRNA / cDNA with a reference polynucleotide that is complementary to the polynucleotide containing the marker gene and its fragments. The cDNA can be amplified, if desired, using any of a variety of polymerase chain reaction methods prior to hybridization with the reference polynucleotide; preferably, it is not amplified.
[0063] Family history: As used herein, the term "family history" generally refers to the occurrence of events (e.g., disease-related disorders or mutation carriers) related to the biological relatives of an individual, including the parents and siblings. Family history may also include grandparents and other relatives.
[0064] Flanking: As used herein, the term "flanking" means that a primer hybridizes to a target nucleic acid adjacent to the region of interest to be amplified on the target. A preferred primer is a pair of primers that hybridize 5' (upstream) from the region of interest, one to each strand of the target double-stranded DNA molecule, such that nucleotides are added to the 3' end of the primer by a suitable DNA polymerase. Those skilled in the art will understand that, for example, a primer flanking a mutant sequence actually anneals to the sequence adjacent to the mutant sequence rather than to the mutant sequence itself. In some examples, primers flanking an exon are typically designed to anneal to the sequence adjacent to the exon (e.g., an intron sequence) rather than to the exon sequence itself. However, in some examples, amplification primers may be designed to anneal to the exon sequence.
[0065] Gene: As used herein, a gene is the unit of heredity. Generally, a gene is a portion of DNA that encodes a protein or functional RNA. A gene is a locatable region of the genomic sequence corresponding to the unit of heredity. A gene may be related to regulatory regions, transcriptional regions, and / or other functional sequence regions.
[0066] Genotype: As used herein, the term "genotype" refers to the genetic constitution of an organism. More specifically, this term refers to the identity of the alleles present in an individual. "Genotyping" an individual or a DNA sample refers to determining the nature of the two alleles possessed by the individual at known polymorphic sites with respect to the nucleotide bases.
[0067] Gene regulatory element: As used herein, a gene regulatory element or regulatory sequence is a segment of DNA to which regulatory proteins, such as transcription factors, bind to regulate gene expression. Such regulatory regions are often upstream of the gene being regulated.
[0068] Healthy individual: As used herein, the term "healthy individual" or "control" refers to a subject who has not been diagnosed with the syndrome and / or disease of interest.
[0069] Heterozygous: As used herein, the term "heterozygous" or "HET" refers to an individual who carries two different alleles of the same gene. As used herein, the term "heterozygous" encompasses "compound heterozygous" or "compound heterozygous variant". As used herein, the term "compound heterozygous" refers to an individual who carries two different alleles. As used herein, the term "compound heterozygous variant" refers to an individual who carries two different copies of an allele, such alleles being characterized as mutant forms of the gene.
[0070] Homozygous: As used herein, the term "homozygous" refers to an individual who carries two copies of the same allele. As used herein, the term "homozygous variant" refers to an individual who carries two copies of the same allele, such alleles being characterized as mutant forms of the gene.
[0071] Housekeeping gene or normalization gene: As used herein, a "housekeeping gene" is a gene that is generally constitutively expressed in all cells to provide the basic functions necessary for the maintenance of all cell types. A housekeeping gene or "normalization gene" is measured concomitantly with a gene of interest to account for variability due to sample-to-sample variation. Such sample-to-sample variation can reflect variability in experimental variables including, but not limited to, RNA isolation, reverse transcription, and PCR efficiency. Normalization involves reporting the ratio of the gene of interest to the housekeeping gene. See, for example, Bustin, S.A. et al., 2009. The MIQE guidelines: minimum information for publication of quantitative real-time PCR experiments. Clin Chem, Apr;55(4):611-622.
[0072] Hybridize: As used herein, the terms "hybridize" or "hybridization" refer to the process by which two complementary nucleic acid strands anneal to each other under appropriately stringent conditions. Oligonucleotides or probes suitable for hybridization generally contain 10 to 100 nucleotides in length (e.g., 18 to 50, 12 to 70, 10 to 30, 10 to 24, 18 to 36 nucleotides in length). Nucleic acid hybridization techniques are well known in the art. One of ordinary skill in the art understands how to estimate and adjust the stringency of hybridization conditions such that sequences with at least a desired level of complementarity hybridize stably, while sequences with lower complementarity do not. Examples of hybridization conditions and parameters can be found, for example, in Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Press, Plainview, N.Y.; Ausubel, F.M. et al., 1994, Current Protocols in Molecular See Biology. John Wiley & Sons, Secaucus, N.J.
[0073] Identity or percent identity: As used herein, the terms "identity" or "percent identity" refer to sequence identity between two amino acid sequences or between two nucleic acid sequences. The percent identity can be determined by aligning the two sequences and refers to the number of identical residues (i.e., amino acids or nucleotides) at corresponding positions shared by the sequences being compared. Sequence alignments and comparisons can be performed using standard algorithms in the art (e.g., Smith and Waterman, 1981, Adv. Appl. Math. 2:482-489; Needleman and Wunsch, 1970, J. Mol. Biol. 48:443-453; Pearson and Lipman, 1988, Proc. Natl. Acad. Sci., USA, 85:2444-2448), or by computerized versions of these algorithms as published as BLAST and FASTA (Wisconsin Genetics Software Package Release 7.0, Genetics Computer Group, 575 Science Drive, Madison, WI). Additionally, ENTREZ, available through the National Institutes of Health in Bethesda, Maryland, may be used for sequence comparison. In other instances, commercially available software, such as GenomeQuest, etc., may be used to determine the percent identity. When utilizing the BLAST and BLAST Gap programs, the default parameters of each program (e.g., BLASTN; available at the Internet site of the National Center for Biotechnology Information) can be used. In one embodiment, the percent identity of two sequences can be determined using GCG with a gap weight of 1 such that each amino acid gap is weighted as if it were a single amino acid mismatch between the two sequences. Alternatively, the ALIGN program (version 2.0), which is part of the GCG (Accelrys, San Diego, CA) sequence alignment software package, may be used.
[0074] As used herein, the term "at least 90% identical thereto" includes sequences that are in the range of 90 to 100% identity with the indicated sequence, including all ranges therebetween. Thus, the term "at least 90% identical thereto" includes sequences that are 91, 91.5, 92, 92.5, 93, 93.5, 94, 94.5, 95, 95.5, 96, 96.5, 97, 97.5, 98, 98.5, 99, 99.5 percent identical to the indicated sequence. Similarly, the term "at least 70% identical" includes sequences that are in the range of 70 to 100% identity, including all ranges therebetween. The percent identity is determined using the algorithms described herein.
[0075] Insertion or addition: As used herein, the terms "insertion" or "addition" refer to a change in an amino acid or nucleotide sequence that results in the addition of one or more amino acid residues or nucleotides, respectively, as compared to a naturally occurring molecule.
[0076] In vitro: As used herein, the term "in vitro" refers to events that occur in an artificial environment, such as a test tube or reaction vessel, cell culture, etc., rather than within a multicellular organism.
[0077] In vivo: As used herein, the term "in vivo" refers to events that occur within a multicellular organism, such as a human or non-human animal.
[0078] Isolated: As used herein, the term "isolated" refers to a substance and / or entity that has been separated from at least a portion of the components with which it was associated when initially created (whether in nature and / or in an experimental setting), and / or a substance and / or entity that has been produced, prepared, and / or manufactured by the hand of man. An isolated substance and / or entity may be separated from at least about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, about 98%, about 99%, substantially 100%, or 100% of the other components with which it was initially associated. In some embodiments, an isolated agent is greater than about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% pure, substantially 100%, or 100% pure. As used herein, a substance is "pure" when it is substantially free of other components. As used herein, the term "isolated cell" refers to a cell that is not contained within a multicellular organism.
[0079] Labeled: The terms "labeled" and "labeled with a detectable agent or moiety" are used synonymously herein to indicate that an entity (e.g., a nucleic acid probe, an antibody, etc.) can be measured by detection of a label (e.g., visualized detection such as radioactivity) after binding to another entity (e.g., a nucleic acid, a polypeptide, etc.). A detectable agent or moiety may be selected to be capable of being measured and generating a signal whose intensity is related (e.g., proportional) to the amount of the bound entity. A wide variety of systems for labeling and / or detecting proteins and peptides are known in the art. Labeled proteins and peptides can be prepared by incorporation or conjugation of a label detectable by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical, chemical, or other means. A label or labeled moiety may be directly detectable (i.e., does not require further reaction or manipulation to become detectable, e.g., a fluorophore is directly detectable) or indirectly detectable (i.e., becomes detectable by reaction or binding to another detectable entity, e.g., a hapten is detectable by immunostaining after reaction with an appropriate antibody containing a reporter such as a fluorophore). Suitable detectable agents include, but are not limited to, radioactive nucleotides, fluorophores, chemiluminescent agents, microparticles, enzymes, colorimetric labels, magnetic labels, haptens, molecular beacons, aptamer beacons, and the like.
[0080] MicroRNA: As used herein, microRNA (miRNA) is a short (20 - 24 nucleotides) non-coding RNA involved in the post-transcriptional regulation of gene expression. MicroRNA can affect both the stability and translation of mRNA. For example, microRNA can bind to complementary sequences in the 3’UTR of target mRNA and cause gene silencing. miRNA is transcribed as part of a capped and polyadenylated primary transcript (pri-miRNA), which can be either protein-coding or non-protein-coding, by RNA polymerase II. The primary transcript can be cleaved by the Drosha ribonuclease III enzyme to generate a stem-loop precursor miRNA (pre-miRNA) of approximately 70 nucleotides, which can be further cleaved by cytoplasmic Dicer ribonuclease to generate mature miRNA and an antisense miRNA star (miRNA * ) product. Mature miRNA can be incorporated into the RNA-induced silencing complex (RISC), which can recognize target mRNA through imperfect base pairing with the miRNA and most commonly results in inhibition of translation or destabilization of the target mRNA.
[0081] Multiplex PCR: As used herein, the term “multiplex PCR” refers to the simultaneous amplification of two or more regions, each primed using a separate primer pair.
[0082] Multiplex ASPE: As used herein, the term “multiplex ASPE” refers to an assay that combines multiplex PCR and allele-specific primer extension (ASPE) to detect genetic polymorphisms. Generally, multiplex PCR is used to first amplify a region of DNA that functions as the target sequence for the ASPE primers. See the definition of allele-specific primer extension.
[0083] Mutation and / or Variant: As used herein, the terms mutation and variant are used interchangeably to describe changes in nucleic acid or protein sequences. The term "variant" as used herein refers to a gene that is mutated or in a potentially non-functional form.
[0084] Nucleic Acid: As used herein, "nucleic acid" is a polynucleotide such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). This term is used to include single-stranded nucleic acids, double-stranded nucleic acids, and RNA and DNA made from nucleotide or nucleoside analogs.
[0085] Obtain or Obtaining: As used herein, the term "obtain" or "obtaining" includes directly acquiring or indirectly (i.e., from a third party) receiving.
[0086] Polypeptide or Protein: As used herein, the terms "polypeptide" and / or "protein" refer to a polymer of amino acids, not a specific length. Thus, peptides, oligopeptides, and proteins are included within the definition of polypeptide and / or protein. "Polypeptide" and "protein" are used interchangeably herein to describe protein molecules that may include partial or full-length proteins. The term "peptide" is used to indicate a protein that is shorter than full length or a very short protein, unless the context indicates otherwise.
[0087] As is known in the art, "protein", "peptide", "polypeptide" and "oligopeptide" are chains of amino acids (generally L-amino acids) in which their α-carbons are linked by peptide bonds formed by a condensation reaction between the carboxyl group of the α-carbon of one amino acid and the amino group of the α-carbon of another amino acid. Generally, the amino acids that make up a protein are numbered in order, starting from the amino-terminal residue and increasing in the direction of the carboxyl-terminal residue of the protein. The abbreviations of amino acid residues are the standard three-letter and / or one-letter codes used in the art to denote one of the 20 common L-amino acids.
[0088] As used herein, a "domain" of a polypeptide or protein includes a region along the polypeptide or protein that contains an independent unit. A domain can be defined with respect to structure, sequence, and / or biological activity. In one embodiment, a polypeptide domain can include a region of a protein that folds in a manner substantially independent of the remainder of the protein. Domains can be identified using domain databases such as, but not limited to, PFAM, PRODOM, PROSITE, BLOCKS, PRINTS, SBASE, ISREC PROFILES, SAMRT, and PROCLASS.
[0089] Primer: As used herein, the term "primer" refers to a short single-stranded oligonucleotide capable of hybridizing to a complementary sequence in a nucleic acid sample. Generally, a primer functions as an initiation point for template-dependent DNA synthesis. Deoxyribonucleotides can be added to the primer by DNA polymerase. In some embodiments, the addition of such deoxyribonucleotides to the primer is also known as primer extension. As used herein, the term "primer" includes all forms of primers that can be synthesized, including peptide nucleic acid primers, locked nucleic acid primers, phosphorothioate-modified primers, labeled primers, and the like. A "primer pair" or "primer set" for a PCR reaction generally refers to a set of primers that includes a "forward primer" and a "reverse primer". As used herein, the term "forward primer" refers to a primer that anneals to the antisense strand of dsDNA. A "reverse primer" anneals to the sense strand of dsDNA.
[0090] Genetic polymorphism: As used herein, the term "genetic polymorphism" refers to the coexistence of two or more forms of a gene or a portion thereof.
[0091] Portions and fragments: As used herein, the terms "portion" and "fragment" are used interchangeably to refer to polypeptides, nucleic acids, or other molecular constructs.
[0092] Sample: As used herein, the term "sample" refers to a material obtained from an individual or subject or patient. Samples can be of any biological origin, including all body fluids (e.g., whole blood, plasma, serum, saliva, aqueous humor of the eye, sweat, urine, milk, etc.), tissues or extracts, cells, cell-free nucleic acids, formalin-fixed paraffin-embedded (FFPE) tissues, and the like.
[0093] Sense strand and antisense strand: As used herein, the term "sense strand" refers to a strand of double-stranded DNA (dsDNA) that includes at least a portion of the coding sequence of a functional protein. As used herein, the term "antisense strand" refers to a strand of dsDNA that is the reverse complement of the sense strand.
[0094] Significant difference: As used herein, the term "significant difference" is well within the knowledge of one of ordinary skill in the art and is determined empirically with reference to each particular biomarker. For example, a significant difference in the expression of a biomarker in a subject having a disease or syndrome of interest compared to a healthy subject is a statistically significant difference in the amount of protein.
[0095] Similar or homolog: As used herein, the terms "similar" or "homolog" when referring to an amino acid or nucleotide sequence mean a polypeptide having a certain degree of homology or identity to the wild-type amino acid sequence. The homology comparison can be done by eye or, more generally, with the aid of readily available sequence comparison programs. These commercially available computer programs can calculate the percentage of homology between two or more sequences (e.g., Wilbur, W.J. and Lipman, D.J., 1983, Proc. Natl. Acad. Sci. USA, 80:726-730). For example, homologous sequences can, in alternative embodiments, be interpreted to include amino acid sequences that are at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 98% identical to each other.
[0096] Specific: As used herein, the term "specific," when used in reference to an oligonucleotide primer, refers to an oligonucleotide or primer that is capable of hybridizing to a target of interest under appropriate hybridization or wash conditions and that does not substantially hybridize to nucleic acids that are not of interest. Higher levels of sequence identity are preferred and include at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity. In some embodiments, a specific oligonucleotide or primer includes at least 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 35, 40, 45, 50, 55, 60, 65, 70 bases or more of sequence identity with a portion of the nucleic acid to be hybridized or amplified when the oligonucleotide and the nucleic acid are aligned.
[0097] As is known in the art, the conditions for hybridizing nucleic acid sequences to each other can be described as ranging from low stringency to high stringency. Generally, highly stringent hybridization conditions refer to washing the hybrids in a high temperature, low salt buffer. Hybridization can be performed using a standard hybridization solution in the art, such as 0.5 M NaHPO4, 7% sodium dodecyl sulfate (SDS), filtering the bound DNA at 65°C, washing with 0.25 M NaHPO4, 3.5% SDS, and subsequently washing with 0.1× SSC / 0.1% SDS at a temperature ranging from room temperature to 68°C depending on the length of the probe (e.g., Ausubel, F.M. et al., Short Protocols in Molecular Biology, 4 th(See, e.g., Ed., Chapter 2, John Wiley & Sons, N.Y.). For example, high stringency washing includes washing in 6×SSC / 0.05% sodium pyrophosphate at 37°C for a 14-base oligonucleotide probe, or at 48°C for a 17-base oligonucleotide probe, or at 55°C for a 20-base oligonucleotide probe, or at 60°C for a 25-base oligonucleotide probe, or at 65°C for a nucleotide probe of about 250 nucleotides in length. The nucleic acid probe can be labeled, for example, with radioactive nucleotides by end-labeling with [γ- 32 P]ATP or by incorporation of radiolabeled nucleotides such as [α- 32 P]dCTP by random primer labeling. Alternatively, the probe can be labeled by incorporation of biotinylated or fluorescein-labeled nucleotides, and the probe can be detected using streptavidin or an anti-fluorescein antibody.
[0098] siRNA: As used herein, siRNA (small interfering RNA) is an essentially double-stranded RNA molecule consisting of about 20 complementary nucleotides. siRNA is generated by the degradation of larger double-stranded (ds) RNA molecules. siRNA can suppress gene expression by essentially splitting its corresponding mRNA into two pieces by the interaction between siRNA and mRNA, thereby causing the degradation of mRNA. In addition, siRNA can also interact with DNA to promote chromatin silencing and heterochromatin expansion.
[0099] Subject: As used herein, the term "subject" refers to a human or any non-human animal. A subject can be a patient, which refers to a human who visits a healthcare provider for the diagnosis or treatment of a disease. Humans include prenatal and postnatal forms. Also, as used herein, the terms "individual", "subject" or "patient" include all warm-blooded animals. In one embodiment, the subject is a human. In one embodiment, the individual is a subject with an increased risk of developing HNSCC.
[0100] Substantially: As used herein, the term "substantially" refers to a qualitative state indicating the entire or nearly entire scope or degree of a feature or characteristic of interest. One of ordinary skill in the biological arts will understand that biological and chemical phenomena rarely, if ever, complete and / or proceed completely, or rarely achieve or avoid absolute results. Thus, the term "substantially" is used herein to capture the lack of potential completeness inherent in many biological and chemical phenomena.
[0101] Substantially complementary: As used herein, the term "substantially complementary" refers to two sequences that can hybridize under stringent hybridization conditions. One of ordinary skill in the art will understand that substantially complementary sequences need not hybridize over their entire length. In some embodiments, "stringent hybridization conditions" refers to hybridization conditions that are at least as stringent as the following: hybridization overnight at 42 °C in 50% formamide, 5X SSC, 50 mM NaH2PO4, pH 6.8, 0.5% SDS, 0.1 mg / mL sonicated salmon sperm DNA, and 5X Denhart's solution; washing at 45 °C with 2X SSC, 0.1% SDS; and washing at 45 °C with 0.2X SSC, 0.1% SDS. In some embodiments, stringent hybridization conditions should not permit hybridization of two nucleic acids that differ by more than two bases over a range of 20 contiguous nucleotides.
[0102] Substitution: As used herein, the term "substitution" refers to the replacement of one or more amino acids or nucleotides with different amino acids or nucleotides, respectively, as compared to a naturally occurring molecule.
[0103] ~suffering from: An individual "suffering from" a disease, disorder, and / or condition has been diagnosed with or exhibits one or more symptoms of the disease, disorder, and / or condition.
[0104] Prone to suffering from: An individual "prone to suffering from" a disease, disorder, and / or condition has not been diagnosed with the disease, disorder, and / or condition. In some embodiments, an individual prone to suffering from a disease, disorder, and / or condition may not exhibit symptoms of the disease, disorder, and / or condition. In some embodiments, an individual prone to suffering from a disease, disorder, and / or condition will develop the disease, disorder, and / or condition. In some embodiments, an individual prone to suffering from a disease, disorder, and / or condition will not develop the disease, disorder, and / or condition.
[0105] Solid support: The term "solid support" or "support" means a structure that provides a substrate to which a biomolecule can be attached. For example, the solid support may be a well of an assay (i.e., a microtiter plate, etc.), or the solid support may be a position on an array, or a movable support such as a bead.
[0106] Upstream and downstream: As used herein, the term "upstream" refers to the residue at the N-terminus of the second residue if the molecule is a protein, or the residue 5' of the second residue if the molecule is a nucleic acid. Also, as used herein, the term "downstream" refers to the residue at the C-terminus of the second residue if the molecule is a protein, or the residue 3' of the second residue if the molecule is a nucleic acid. All protein, polypeptide, and peptide sequences disclosed herein are described from the N-terminal amino acid to the C-terminal acid, and all nucleic acid sequences disclosed herein are described from the 5' end of the molecule to the 3' end of the molecule.
[0107] Summary The disclosure of this specification provides novel mutations identified in specific genes that are related to genes of diseases and / or syndromes of interest and can be used for more accurate diagnosis of disorders related to the genes and / or syndromes of interest.
[0108] In some embodiments, the sample contains nucleic acids. In some embodiments, the testing step includes nucleotide sequencing. In some embodiments, the testing step includes hybridization. In some embodiments, the hybridization is performed using one or more oligonucleotide probes specific to the region of the biomarker of interest. In some embodiments, for the detection of mutations, the hybridization is performed under conditions stringent enough not to allow any mismatch of one nucleotide. In some embodiments, the hybridization is performed using a microarray. In some embodiments, the testing step includes restriction enzyme digestion. In some embodiments, the testing step includes PCR amplification. In some embodiments, the testing step includes reverse transcriptase PCR (rtPCR). In some embodiments, the PCR amplification is digital PCR amplification. In some embodiments, the testing step includes primer extension. In some embodiments, the primer extension is single-base primer extension. In some embodiments, the testing step includes performing multiplex allele-specific primer extension (ASPE).
[0109] In some embodiments, the sample contains proteins. In some embodiments, the testing step includes amino acid sequencing. In some embodiments, the testing step includes performing an immunoassay using one or more antibodies that specifically recognize the biomarker of interest. In some embodiments, the testing step includes protease digestion (e.g., trypsin digestion). In some embodiments, the testing step further includes performing 2D gel electrophoresis.
[0110] In some embodiments, the test step includes determining the presence of one or more biomarkers using mass spectrometry. In some embodiments, the mass spectrometry format is selected from matrix-assisted laser desorption / ionization, time-of-flight (MALDI-TOF), electrospray (ES), IR-MALDI, ion cyclotron resonance (ICR), Fourier transform, and combinations thereof.
[0111] In some embodiments, the sample is obtained from cells, tissue (e.g., FFPE tissue), whole blood, oral wash, plasma, serum, urine, feces, saliva, cord blood, chorionic sample, chorionic sample culture, amniotic fluid, amniotic fluid culture, endocervical lavage, or combinations thereof. In certain embodiments, the sample may be either a liquid or a tissue biopsy. In some embodiments, the sample contains cell-free nucleic acids (e.g., DNA) that may be present in a biological sample such as blood, plasma, serum, or amniotic fluid.
[0112] In some embodiments, the test step includes determining the nucleotide and / or amino acid identity at a predetermined position of the biomarker. In some embodiments, the presence of a mutation is determined by comparing the nucleotide and / or amino acid identity at the predetermined position to a control.
[0113] In embodiments, the method may include performing an assay (e.g., nucleic acid sequencing) in a plurality of individuals to determine the associated statistical significance.
[0114] In another aspect, the present disclosure provides reagents for detecting biomarkers of interest, such as, but not limited to, nucleic acid probes that specifically bind to biomarkers (e.g., DNA sequences, mRNA, mutations in proteins), or arrays containing one or more probes that specifically bind to biomarkers. In some embodiments, the present disclosure provides antibodies that specifically bind to biomarkers. In some embodiments, the present disclosure provides kits comprising one or more such reagents. In some embodiments, the one or more reagents are provided in the form of a microarray. In some embodiments, the kit further comprises reagents for primer extension. In some embodiments, the kit further comprises a control indicating a healthy individual. In some embodiments, the kit further comprises instructions regarding a method for determining whether an individual has a syndrome or disease of interest based on the biomarker of interest.
[0115] In some examples, the amount of one or more biomarkers can be detected, in certain embodiments, by: (a) detecting the amount of a polypeptide or protein regulated by the one or more biomarkers; (b) detecting the amount of a polypeptide or protein that regulates the biomarker; or (c) detecting the amount of a metabolite of the biomarker.
[0116] In yet another aspect, the present disclosure provides a computer-readable medium encoding information corresponding to the detection of a biomarker.
[0117] Methods and Compositions for Diagnosing HNSCC Embodiments of the present disclosure include methods and compositions for diagnosing the presence of HNSCC or an increased risk of developing the same. Using the methods and compositions of the present disclosure, genetic information can be obtained from or provided for a subject, and the presence of HNSCC or an increased risk of developing the same in that subject or other subjects can be objectively diagnosed. The methods and compositions can be embodied in a variety of ways.
[0118] In one embodiment, a method for detecting biomarkers associated with head and neck squamous cell carcinoma (HNSCC) in an individual is disclosed, comprising the steps of obtaining a sample from the individual and measuring the expression level of at least one gene in Table 4 and / or Table 6 in the sample. In one embodiment, a method for detecting biomarkers associated with HNSCC in an individual is disclosed, comprising the steps of obtaining a sample from the individual and measuring the expression level of at least one of the following genes: CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1, and HSD17B6. In one embodiment, a method for detecting biomarkers associated with HNSCC in an individual is disclosed, comprising the steps of obtaining a sample from the individual and measuring the expression level of at least one human papillomavirus (HPV) E6 or E7 gene. Alternatively, various combinations of genes may be evaluated. In certain embodiments, the measured value of the expression of any of these genes may be compared to a control value. In various embodiments, the difference between the gene expression in the individual and the control value indicates whether the individual may have HNSCC (i.e., a diagnosis of presence) or is at high risk of developing HNSCC (i.e., high risk). The control value may be derived from one or more samples of normal (non-cancerous tissue) or from a normal (non-cancerous) population. Further, and / or alternatively, the method may include the measurement of at least one normalization (e.g., housekeeping) gene. In one embodiment, the normalization gene may be KHDRBS1. In other embodiments, RPL30 or other normalization genes may be used.
[0119] Furthermore, and / or alternatively, obtaining a sample from an individual; measuring the amount of an expression product from a gene comprising at least one gene in Table 4 and / or Table 6 in the sample; and comparing the expression of at least one gene in Table 4 and / or Table 6 in the sample to a control value of expression, a method for detecting susceptibility to head and neck squamous cell carcinoma (HNSCC) in an individual is disclosed. In various embodiments, a difference between the expression of a gene in an individual and the control value indicates that the individual may have HNSCC (i.e., a diagnosis of presence) or is prone to developing HNSCC (i.e., at high risk). The control value may be derived from one or more samples of normal (non-cancerous tissue) or may be derived from a normal (non-cancerous) population. Furthermore, and / or alternatively, the method may include measurement of at least one normalization (e.g., housekeeping) gene. In one embodiment, the normalization gene may be KHDRBS1. In other embodiments, RPL30 or other normalization genes may be used.
[0120] Furthermore, and / or alternatively, obtaining a sample from an individual; measuring the amount of at least one expression product from at least one gene comprising at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6 in the sample; and comparing the expression of at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6 in the sample with a control value for the expression of each of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6, a method for detecting susceptibility to HNSCC in an individual is disclosed. Furthermore, and / or alternatively, obtaining a sample from an individual; measuring the amount of at least one expression product from at least one gene comprising at least one of HPV E6 and / or HPV E7 in the sample; and comparing the expression of at least one of the HPV E6 and / or HPV E7 expression products in the sample with a control value for the expression of each of HPV E6 and / or HPV E7, a method for detecting susceptibility to HNSCC in an individual is disclosed. In various embodiments, the difference between the gene expression and the control value in an individual indicates that the individual may have HNSCC (i.e., a diagnosis of presence) or is prone to developing HNSCC (i.e., at high risk). The control value may be derived from one or more samples of normal (non-cancerous tissue) or may be derived from a normal (non-cancerous) population. Furthermore, and / or alternatively, the method may include measurement of at least one normalization (e.g., housekeeping) gene. In one embodiment, the normalization gene may be KHDRBS1. In other embodiments, RPL30 or other normalization genes may be used.
[0121] In certain embodiments, measuring comprises measuring RNA (e.g., mRNA). Alternatively, measuring may comprise an immunoassay.
[0122] In some cases, increasing the number of biomarkers improves the statistical power of this method. For example, in certain embodiments, the method may include measuring the expression of at least 2, 3, 4, 5 or more biomarkers. In some examples, at least 4 biomarkers are measured.
[0123] A variety of samples may be assayed. In certain embodiments, the sample includes serum, plasma, saliva or tissue (e.g., FFPE tissue).
[0124] The disclosed methods also include methods for identifying markers associated with head and neck squamous cell carcinoma (HNSCC) in an individual, including identifying at least one marker that is increased or decreased in expression in HNSCC but not expressed in HNSCC disease compared to normal controls. As disclosed herein, such methods may include a statistical evaluation of markers that show differential expression in HNSCC compared to normal tissue. Alternatively, the method may include biomarkers (e.g., mutant genes, copy number differences and translocations, DNA methylation, and / or microRNAs) that discriminate HNSCC from normal based on other biological criteria. Alternatively, other biological aspects of the biomarker may be evaluated.
[0125] As disclosed herein, a variety of methods can be used to measure biomarkers of interest. In one embodiment, measuring includes measuring mRNA. In one embodiment, the measurement includes measuring a peptide or polypeptide biomarker. For example, in one embodiment, the measurement includes an immunoassay. Alternatively, the measurement may include flow cytometry. Alternatively, nucleic acid methods may be used as discussed in detail herein.
[0126] In yet other embodiments, methods of treating HNSCC are disclosed. The methods of treatment can include obtaining a sample from an individual; measuring the amount of an expression product from a gene comprising at least one of the genes in Table 4 and / or Table 6 in the sample; comparing the expression of at least one of the genes in Table 4 and / or Table 6 in the sample to a control value of expression; and treating the individual for HNSCC if a difference between the gene expression and the control value in the individual indicates that the individual may have HNSCC (i.e., a diagnosis of presence) or is at risk of developing HNSCC (i.e., high risk). The control value can be derived from one or more samples of normal (non-cancerous tissue) or from a normal (non-cancerous) population. Further and / or alternatively, the method can include measurement of at least one normalization (e.g., housekeeping) gene. In one embodiment, the normalization gene can be KHDRBS1. In other embodiments, RPL30 or other normalization genes can be used.
[0127] For example, in certain embodiments, the methods of treatment can include obtaining a sample from an individual; measuring the amount of at least one expression product from at least one gene comprising at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1, or HSD17B6 in the sample; comparing the expression of at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1, or HSD17B6 in the sample to a control value of expression for each of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1, or HSD17B6; and treating the individual for HNSCC if a difference between the gene expression and the control value in the individual indicates that the individual may have HNSCC (i.e., a diagnosis of presence) or is at risk of developing HNSCC (i.e., high risk).
[0128] The method of treatment may further include measuring the amount of at least one expression product from a gene containing at least one of HPV E6 and / or HPV E7 in a sample; comparing the expression of at least one of the HPV E6 and / or HPV E7 expression products in the sample with a control value for the expression of each of HPV E6 and / or HPV E7; and treating the individual for HNSCC when the difference between the gene expression and the control value in the individual indicates that the individual may have HNSCC (i.e., a diagnosis of presence) or is at risk of developing HNSCC (i.e., high risk).
[0129] In various embodiments of the method of treatment, the control value may be derived from one or more samples of normal (non-cancerous tissue) or may be derived from a normal (non-cancerous) population. Further, and / or alternatively, the method may include measurement of at least one normalization (e.g., housekeeping) gene. In one embodiment, the normalization gene may be KHDRBS1. In other embodiments, RPL30 or other normalization genes may be used.
[0130] As described above, still other embodiments include compositions for detecting biomarkers associated with HNSCC in an individual. In certain embodiments, the composition includes reagents for quantifying the level of at least one disclosed biomarker in a biological sample. For example, as described in detail herein, the composition may include reagents for measuring mRNA. Or, the composition may include reagents for measuring peptide or polypeptide biomarkers. In one embodiment, the composition includes reagents for performing an immunoassay. Or, the composition may include reagents for performing flow cytometry. Or, as described in detail herein, the composition may include reagents for determining the presence of a particular sequence and / or the expression level of a nucleic acid. As described in detail herein, the reagents may be labeled with a detectable moiety.
[0131] Accordingly, other aspects of the disclosure include compositions for detecting biomarkers associated with HNSCC in an individual, comprising reagents for quantifying the expression level of at least one gene in Table 4 and / or Table 6. Further and / or alternatively, other aspects of the disclosure include compositions for detecting biomarkers associated with HNSCC in an individual, comprising reagents for quantifying the expression level of at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6. Further and / or alternatively, aspects of the disclosure are HPV Compositions for detecting biomarkers associated with HNSCC in an individual, comprising reagents for quantifying the expression level of at least one of HPV E6 and / or HPV E7. Further and / or alternatively, the composition may include reagents for detecting at least one normalization (e.g., housekeeping) gene. In one embodiment, the normalization gene may be KHDRBS1. In other embodiments, RPL30 or other normalization genes may be used. For example, in some embodiments, the reagents detect mRNA. Alternatively, the reagents may detect proteins.
[0132] Accordingly, in certain embodiments, the composition can include primers (e.g., primer pairs) and / or probes for any one of these genes, and the primers and / or probes are labeled with a detectable moiety as described herein. Further and / or alternatively, the primers and / or probes can also include an array to which the primers and / or probes are immobilized on a surface. In other embodiments, the reagents can include reagents for measuring peptides and / or proteins expressed from the disclosed genes. For example, the composition can include reagents for performing an immunoassay. In some embodiments, these reagents can include an array as described in detail herein. As described in detail herein, the reagents can be labeled with a detectable moiety.
[0133] Other embodiments include systems for performing the methods disclosed herein and / or using the compositions disclosed herein. For example, other aspects of the disclosure include kits that include the compositions of the disclosure, or kits for performing the methods of the disclosure. For example, other embodiments include systems such as kits that contain at least some of the compositions disclosed herein and / or reagents for performing the methods disclosed herein. Such systems or kits can include a computer-readable medium that includes instructions and / or other information for performing the methods and / or using the compositions of the disclosure.
[0134] Various types of samples can be used with any of the methods, compositions, or systems disclosed herein. In certain embodiments, the sample includes serum, plasma, saliva, or tissue (e.g., FFPE tissue). Alternatively, other sample types may be used.
[0135] Peptide, polypeptide, and protein assays In certain embodiments, biomarkers of interest are detected at the protein level (or peptide or polypeptide level), i.e., the gene product is analyzed. For example, a protein or fragment thereof can be analyzed by amino acid sequencing or by an immunoassay that uses one or more antibodies that specifically recognize one or more epitopes present on the biomarker of interest or in some instances one or more antibodies specific to a mutation of interest. A protein can be analyzed by protease digestion (e.g., trypsin digestion), and in some embodiments, the digested protein products can be further analyzed by 2D gel electrophoresis.
[0136] Antibody-based detection methods Specific antibodies that recognize biomarkers of interest can be used in any of a variety of methods known in the art. Antibodies against specific epitopes, polypeptides, and / or proteins can be generated using any of a variety of methods known in the art. For example, an animal, typically a mammal (e.g., donkey, mouse, rabbit, horse, chicken, etc.), can be injected with an epitope, polypeptide, or protein that produces an antibody against it, and the antibodies produced by the animal can be collected from the animal. Monoclonal antibodies can also be generated by producing hybridomas that express the antibody of interest in an immortalized cell line.
[0137] In some embodiments, the antibody is labeled with a detectable moiety described herein.
[0138] Antibody detection methods are well known in the art and include, but are not limited to, enzyme-linked immunosorbent assay (ELISA) and Western blot. Some such methods are suitable for performing in an array format.
[0139] For example, in some embodiments, a biomarker of interest is detected using a first antibody (or antibody fragment) that specifically recognizes the biomarker. The antibody may be labeled with a detectable moiety (e.g., a chemiluminescent molecule), an enzyme, or a second binding agent (e.g., streptavidin). Alternatively, a second antibody may be used to detect the first antibody as is known in the art.
[0140] In certain embodiments, the method may further comprise adding a capture support, the capture support comprising at least one capture support binder that recognizes and binds a biomarker to immobilize the biomarker on the capture support. The method may further comprise adding, in certain embodiments, a second binder that can specifically recognize and bind at least a portion of the plurality of binder molecules and / or biomarker on the capture support. In one embodiment, the binder that can specifically recognize and bind at least a portion of the plurality of binder molecules and / or biomarker on the capture support is a soluble binder (e.g., a secondary antibody). The second binder may be labeled (e.g., with an enzyme) such that the binding of the biomarker of interest is measured by adding an enzyme substrate and quantifying the amount of product formed.
[0141] In one embodiment, the capture solid support may be an assay well (i.e., such as a microtiter plate). Alternatively, the capture solid support may be a position on an array, or a mobile support such as a bead. Or the capture support can be a filter.
[0142] In some examples, the biomarker can be complexed with a first binder (e.g., a primary antibody specific for the biomarker and labeled with a detectable moiety) and a second binder (e.g., a secondary antibody or a second primary antibody that recognizes the primary antibody), the second binder forms a complex with a third binder (e.g., biotin), and the third binder can then interact with a capture support (e.g., magnetic beads) having a reagent (e.g., streptavidin) that recognizes the third binder linked to the capture support. The complex (labeled primary antibody:biomarker:second primary antibody-biotin:streptavidin-beads) is then captured using a magnet (e.g., a magnetic probe) to measure the amount of the complex.
[0143] A variety of binders may be used in the methods of the present disclosure. For example, the binder bound to the capture support, or the second antibody, may be either an antibody or an antibody fragment that recognizes a biomarker. Alternatively, the binder may include a protein that binds to a non-protein target (i.e., a protein that specifically binds to a small molecule biomarker, or a receptor that binds to a protein, etc.).
[0144] In certain embodiments, the solid support may be treated with an immobilizing agent. For example, in certain embodiments, a biomarker of interest may be captured on an immobilized surface (i.e., a surface treated to reduce non-specific binding). One such immobilizing agent is BSA. Further, and / or alternatively, when the binder used is an antibody, the solid support may be coated with Protein A, Protein G, Protein A / G, Protein L, or another agent that binds with high affinity to the binder (e.g., an antibody). These proteins bind to the Fc domain of the antibody and can thus direct the binding of an antibody that recognizes one or more proteins of interest.
[0145] Nucleic acid assay In certain embodiments, the biomarkers disclosed herein are detected at the nucleic acid level. In one embodiment, the present disclosure includes a method for diagnosing the presence or increased risk of developing a syndrome or disease of interest (e.g., HNSCC) in a subject.
[0146] This method may include obtaining nucleic acids from a tissue or body fluid sample from the subject, and performing an assay to identify whether there is overexpression of a gene of interest. For example, overexpression of a particular gene product may be quantified using reverse transcriptase PCR (RT-PCR). Alternatively, droplet digital PCR (ddPCR) may be used.
[0147] Alternatively, the method can include obtaining nucleic acids from a tissue or body fluid sample from a subject and performing an assay to identify whether the nucleic acids of the subject have a variant sequence (i.e., a mutation). In certain embodiments, the method can include comparing the variant to known variants associated with a syndrome or disease of interest and determining whether the variant is a variant previously identified as being associated with the syndrome or disease of interest. Alternatively, the method can include identifying the variant as a new variant that has not been previously characterized. If the variant is a new variant, the method can further include performing an analysis to determine whether the mutation is expected to be deleterious to gene expression and / or the function of the protein encoded by the gene. The method can further include using a variant profile (i.e., compilation of mutations identified in a subject) to diagnose the presence of a syndrome or disease of interest or an increased risk of developing a syndrome or disease of interest.
[0148] Nucleic acid analysis can be performed on genomic DNA, messenger RNA, and / or cDNA. Also, in various embodiments, the nucleic acids include genes, RNAs, exons, introns, gene regulatory elements, expressed RNAs, siRNAs, or epigenetic elements. Also, regulatory elements including splice sites, transcription factor binding, A-I editing sites, microRNA binding sites, and functional RNA structural sites can be evaluated for mutations (i.e., variants). Thus, for each of the methods and compositions of the present disclosure, a variant can include a nucleic acid sequence encompassing at least one of the following: (1) an A-to-I editing site; (2) a splice site; (3) a conserved functional RNA structure; (4) a valid transcription factor binding site (TFBS); (5) a microRNA (miRNA) binding site; (6) a polyadenylation site; (7) a known regulatory element; (8) a miRNA gene; (9) a small nucleolar RNA gene encoded in the ROI; and / or (10) an ultra-conserved element across placental mammals.
[0149] In many embodiments, the nucleic acid is extracted from a biological sample. In some embodiments, the nucleic acid is analyzed without amplification. In some embodiments, the nucleic acid is amplified using techniques known in the art (such as the generation of cDNA amplified using polymerase chain reaction (PCR)), and the amplified nucleic acid is then used for subsequent analysis. Multiplex PCR may be used to amplify several amplicons (e.g., from different genomic regions) at once using multiple sets of primer pairs. For example, the nucleic acid can be analyzed by sequencing, hybridization, PCR amplification, restriction enzyme digestion, primer extension, such as single-base primer extension or multiplex allele-specific primer extension (ASPE), or DNA sequencing. In some embodiments, the nucleic acid is amplified such that the size of the amplification product of the wild-type allele is different from the size of the mutant allele. Thus, the presence or absence of a particular mutant allele can be determined by detecting differences in the size of the amplification products, for example, on an electrophoresis gel. For example, deletions or insertions in a gene region may be particularly suitable for the use of size-based approaches.
[0150] A particular example of a nucleic acid analysis method is described in detail below.
[0151] Analysis of mRNA In certain embodiments, mRNA is analyzed using real-time and / or reverse transcriptase PCR using methods and / or commercially available reagents and / or kits known in the art. "Real-time PCR" or rPCR is a method for detecting and measuring the products generated during each cycle of PCR that is proportional to the amount of template nucleic acid before the start of PCR. The information obtained, such as the amplification curve, can be used to determine the presence of the target nucleic acid and / or to quantify the initial amount of the target nucleic acid sequence. The term "real-time PCR" is used to represent a subset of PCR techniques that enable the detection of PCR products throughout the PCR reaction, or in real-time.
[0152] In some examples, rPCR is real-time reverse transcriptase (RT) PCR (rRT-PCR). Alternatively, droplet digital PCR can be used. Reverse transcriptase PCR is used when the starting material is RNA and / or mRNA. The RNA is first transcribed into complementary DNA (cDNA) by reverse transcriptase. In rRT-PCR, the cDNA is then used as a template for the qPCR reaction. rRT-PCR can be performed in a one-step method that combines reverse transcription and PCR in a single tube and buffer using reverse transcriptase together with DNA polymerase. In one-step rRT-PCR, both RNA targets and DNA targets are amplified using sequence-specific targets. The term "quantitative PCR" encompasses all PCR-based techniques that enable the quantitative or semi-quantitative determination of the target nucleic acid sequences initially present.
[0153] The principle of real-time PCR (rPCR) is generally, for example, as described by Held et al., "Real" "Time Quantitative PCR" is described in Genome Research 6:986-994 (1996). Generally, rPCR measures the signal at each amplification cycle. Some rPCR techniques rely on fluorophores that emit a signal at the completion of every amplification cycle. Examples of such fluorophores are fluorescent dyes that emit fluorescence at a predetermined wavelength when bound to double-stranded DNA, such as SYBR Green. Thus, the increase in double-stranded DNA during each amplification cycle leads to an increase in fluorescence intensity due to the accumulation of the PCR product. Another example of a fluorophore used for detection in rPCR is the sequence-specific fluorescent reporter probe described elsewhere in this document. An example of such a probe is the TAQMAN® probe. The use of sequence-specific reporter probes results in the detection of the target sequence with high specificity and enables quantification even in the presence of non-specific DNA amplification. Fluorescent probes can also be used in multiplex assays to detect several genes in the same reaction based on specific probes labeled with different colors. For example, a multiplex assay can use several sequence-specific probes labeled with a variety of fluorophores including, but not limited to, FAM, JA270, CY5.5, and HEX in the same PCR reaction mixture.
[0154] rPCR relies on the detection of measurable parameters such as fluorescence during the PCR reaction. The amount of the measurable parameter is proportional to the amount of the PCR product, thereby enabling the "real-time" observation of the increase in the PCR product. Some rPCR methods enable the quantification of the input DNA template based on the observable progress of the PCR reaction. The analysis and processing of the data are discussed below. A "growth curve" or "amplification curve" in the context of a nucleic acid amplification assay is a graph of a function where the independent variable is the number of amplification cycles and the dependent variable is an amplification-dependent measurable parameter measured at each cycle of amplification, such as the fluorescence emitted by a fluorophore. As discussed above, the amount of the amplified target nucleic acid can be detected using a fluorophore-labeled probe. Generally, the amplification-dependent measurable parameter is the amount of fluorescence emitted by the probe during hybridization or during hydrolysis of the probe by the nuclease activity of the nucleic acid polymerase. The increase in fluorescence emission is measured in real time and is directly related to the increase in target nucleic acid amplification. In some examples, the change in fluorescence (dR n ) is calculated using the formula dR n =R n+ -R n- , where R n+ is the fluorescence emission of the product at each time point and R n- is the baseline fluorescence emission. The dR n values are plotted against the number of cycles to obtain an amplification plot. In a typical polymerase chain reaction, the growth curve includes a segment of exponential growth followed by a plateau, resulting in an S-shaped amplification plot when using a linear scale. The growth curve is characterized by a "crossing point" value or "C p " value. This is also called the "threshold" or "cycle threshold" (C t ) and is the number of cycles at which a measurable parameter of a predetermined magnitude is achieved. For example, when using a fluorophore-labeled probe, the threshold (C t ) is the fluorescence emission (dR nis the number of PCR cycles at which the selected threshold is exceeded, which is generally 10 times the standard deviation of the baseline (however, this threshold level can be changed as desired). C t The lower the value, the faster the completion of amplification, and C t the higher the value, the slower the completion of amplification. When the amplification efficiencies are similar, C t the lower the value, the more starting amount of the target nucleic acid, reflecting that C t the higher the value, the less starting amount of the target nucleic acid. By using a control nucleic acid of known concentration to generate a "standard curve" or a set of "control" C t values at various known concentrations of the control nucleic acid, it becomes possible to determine the absolute amount of the target nucleic acid in the sample by comparing the C t values of the target nucleic acid and the control nucleic acid.
[0155] Allele-specific amplification In some embodiments, for example, when the biomarker for a disease and / or syndrome of interest is a mutation, the biomarker is detected using an allele-specific amplification assay. This approach is variously called specific allele PCR amplification (PASA) (Sarkar et al., 1990 Anal. Biochem. 186:64 - 68), allele-specific amplification (ASA) (Okayama et al., 1989 J. Lab. Clin. Med. 114:105 - 113), allele-specific PCR (ASPCR) (Wu et al., 1989 Proc. Natl. Acad. Sci. USA. 86:2757 - 2760), and amplification refractory mutation system (ARMS) (Newton et al., 1989 Nucleic Acids Res. 17:2503 - 2516). The entire content of each of these references is incorporated herein. This method is applicable to single nucleotide substitutions as well as microdeletions / insertions.
[0156] For example, with respect to amplification methods based on PCR, amplification primers can be designed to distinguish between different alleles (e.g., between wild-type alleles and mutant alleles). Thus, the presence or absence of an amplification product can be used to determine whether a gene mutation is present in a given nucleic acid sample. In some embodiments, allele-specific primers can be designed such that the presence of an amplification product indicates a gene mutation. In some embodiments, allele-specific primers can be designed such that the absence of an amplification product indicates a gene mutation.
[0157] In some embodiments, two complementary reactions are used. One reaction uses primers specific for the wild-type allele (“wild-type specific reaction”), and the other reaction uses primers for the mutant allele (“mutation specific reaction”). The two reactions may use a common second primer. PCR primers specific for a particular allele (e.g., a wild-type allele or a mutant allele) typically exactly match one allele variant of the target, but do not match other allele variants (e.g., a mutant allele or a wild-type allele). The mismatch may be located at or near the 3’ end of the primer, leading to preferential amplification of the perfectly matched allele. Whether an amplification product can be detected from one or both reactions indicates the presence or absence of the mutant allele. If an amplification product is detected only from the wild-type specific reaction, the presence of only the wild-type allele (e.g., homozygosity for the wild-type allele) is indicated. If an amplification product is detected only from the mutation specific reaction, the presence of only the mutant allele (e.g., homozygosity for the mutant allele) is indicated. If an amplification product is detected from both reactions (e.g., a heterozygote) is indicated. In this specification, this approach is also referred to as “allele-specific amplification (ASA)”.
[0158] Allele-specific amplification can also be used to detect duplications, insertions, or translocations by using primers that partially hybridize across the junction. The extent of junction overlap can be varied to allow specific amplification.
[0159] The amplification products can be examined by methods known in the art, which include visualizing (e.g., with one or more dyes) bands of nucleic acids that have migrated through a gel (e.g., by electrophoresis) to separate the nucleic acids by size.
[0160] Allele-Specific Primer Extension In some embodiments, an allele-specific primer extension (ASPE) approach is used to detect gene mutations. ASPE uses allele-specific primers that can distinguish alleles (e.g., a mutant allele and a wild-type allele) in an extension reaction, so that extension products are obtained only in the presence of a specific allele (e.g., a mutant allele or a wild-type allele). The extension products can be, or can be made, detectable, for example, by using labeled deoxyribonucleotides in the extension reaction. Any of a variety of labels are suitable for use in these methods, including but not limited to radiolabels, fluorescent labels, chemiluminescent labels, enzyme labels, etc. The nucleotides are labeled with an entity that can then be (directly or indirectly) conjugated by a biotin molecule that binds to a detectable label, such as a streptavidin-conjugated fluorescent dye. In some embodiments, the reaction is performed multiplexed, e.g., using multiple allele-specific primers in the same extension reaction.
[0161] In some embodiments, the extension product is hybridized to a solid or semi-solid support, such as beads, matrix, gel, etc., among others. For example, the extension product may be tagged with a specific nucleic acid sequence (e.g., included as part of an allele-specific primer), and an "anti-tag" (e.g., a nucleic acid sequence complementary to the tag in the extension product) may be attached to the solid support. The extension product can be captured and detected on the solid support. For example, the beads can be sorted and detected.
[0162] Single nucleotide primer extension In some embodiments, a single nucleotide primer extension (SNuPE) assay is used in which the primer is designed to extend by only one nucleotide. In such a method, the identity of the nucleotide immediately downstream of the 3' end of the primer is known and differs for the mutant allele compared to the wild-type allele. SNuPE can be performed using an extension reaction in which only a specific type of deoxynucleotide is labeled (e.g., labeled dATP, labeled dCTP, labeled dGTP, or labeled dTTP). Thus, the presence of a detectable extension product can be used as an indicator of the identity of the nucleotide at the position of interest (e.g., the position immediately downstream of the 3' end of the primer) and thus as an indicator of the presence or absence of a mutation at that position. SNuPE is described in U.S. Patent No. 5,888,819; U.S. Patent No. 5,846,710; U.S. Patent No. 6,280,947; U.S. Patent No. 6,482,595; U.S. Patent No. 6,503,718; U.S. Patent No. 6,919,174; Piggee, C. et al., Journal of Chromatography A 781 (1997), p. 367-375 (“Capillary Electrophoresis for the Detection of Known Point Mutations by Single-Nucleotide Primer Extension and Laser-Induced Fluorescence Detection”); Hoogendoorn, B. et al., Human Genetics (1999) 104:89-93, (“Genotyping Single Nucleotide Polymorphism by Primer Extension and High Performance Liquid Chromatography”), each of which is incorporated herein by reference in its entirety.
[0163] In some embodiments, primer extension can be combined with mass spectrometry to accurately and rapidly detect the presence or absence of mutations. See U.S. Patent No. 5,885,775 to Haff et al. (Analysis of single nucleotide polymorphisms by mass spectrometry); U.S. Patent No. 7,501,251 to Koster (DNA diagnosis based on mass spectrometry); the teachings of both are incorporated herein by reference. Suitable mass spectrometry formats include, but are not limited to, matrix-assisted laser desorption / ionization, time-of-flight (MALDI-TOF), electrospray (ES), IR-MALDI, ion cyclotron resonance (ICR), Fourier transform, and combinations thereof.
[0164] Oligonucleotide ligation assay In some embodiments, an oligonucleotide ligation assay ("OLA" or "OL") is used. OLA uses two oligonucleotides designed to hybridize to adjacent sequences of a single strand of a target molecule. Generally, one oligonucleotide is biotinylated and the other is detectably labeled, for example, with a streptavidin-binding fluorescent moiety. When exact complementary sequences are found in the target molecule, the oligonucleotides hybridize such that their ends are adjacent, creating a ligation substrate that can be captured and detected. See, for example, Nickerson et al., (1990) Proc. Natl. Acad. Sci. U.S.A. 87:8923-8927, Landegren, U. et al., (1988) Science 241:1077-1080 and U.S. Patent No. 4,998,617. The entire contents of which are incorporated herein by reference in their entirety.
[0165] Hybridization approach In some embodiments, the nucleic acids are analyzed by hybridization using one or more oligonucleotide probes specific for the biomarker of interest under conditions stringent enough to not tolerate single nucleotide mismatches. In certain embodiments, suitable nucleic acid probes can distinguish between normal and mutant genes. Thus, for example, one of ordinary skill in the art can use the probes of the invention to determine whether an individual is homozygous or heterozygous for a particular allele.
[0166] Nucleic acid hybridization techniques are well known in the art. One of ordinary skill in the art understands how to estimate and adjust the stringency of hybridization conditions such that sequences having at least a desired level of complementarity hybridize stably while sequences having lower complementarity do not. Examples of hybridization conditions and parameters can be found, for example, in Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Press, Plainview, N.Y.; Ausubel, F.M. et al., 1994, Current Protocols in Molecular Biology. John Wiley & Sons, Secaucus, N.J.
[0167] In some embodiments, probe molecules that hybridize to mutant or wild-type sequences can be used to detect such sequences in the amplification products by liquid-phase or more preferably solid-phase hybridization. Solid-phase hybridization can be achieved, for example, by adhering the probe to a microchip.
[0168] The nucleic acid probe may comprise ribonucleic acid and / or deoxyribonucleic acid. In some embodiments, the provided nucleic acid probe is an oligonucleotide (i.e., an "oligonucleotide probe"). Generally, an oligonucleotide probe is of a length sufficient to specifically bind to a homologous region of a gene of interest, but short enough that a one-nucleotide difference between the probe and the nucleic acid sample being tested disrupts hybridization. Generally, the size of an oligonucleotide probe varies from approximately 10 to 100 nucleotides. In some embodiments, the oligonucleotide probe is 15-90, 15-80, 15-70, 15-60, 15-50, 15-40, 15-35, 15-30, 18-30, or 18-26 nucleotides in length. As will be appreciated by those skilled in the art, the optimal length of an oligonucleotide probe may depend on the particular method and / or conditions in which the oligonucleotide probe may be used.
[0169] In some embodiments, the nucleic acid probe is useful, for example, as a primer in nucleic acid amplification and / or extension reactions. For example, in certain embodiments, the gene sequence being evaluated for variants includes an exon sequence. In certain embodiments, the exon sequence and additional flanking sequences (e.g., about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55 or more nucleotides of a UTR and / or intron sequence) are analyzed in an assay. Alternatively, an intron sequence or other non-coding region may be evaluated for potentially harmful mutations. Or, portions of these sequences may be used. Such variant gene sequences may include sequences that may contain at least one of the mutations described herein.
[0170] Other embodiments of the disclosure provide isolated gene sequences containing mutations associated with a syndrome and / or disease of interest. Such gene sequences can be used to objectively diagnose the presence or increased risk of developing HNSCC in a subject. In certain embodiments, the isolated nucleic acid can contain a non-variant sequence or a variant sequence of any one or a combination thereof. For example, in certain embodiments, the gene sequence includes an exon sequence. In certain embodiments, the exon sequence and additional flanking sequences (e.g., about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55 or more nucleotides of the UTR and / or intron sequence) are analyzed in an assay. Alternatively, an intron sequence or other non-coding region may be used. Alternatively, portions of these sequences may be used. In certain embodiments, the gene sequence includes an exon sequence from at least one of the biomarker genes disclosed herein.
[0171] In some embodiments, the nucleic acid probe is labeled with a detectable moiety described herein.
[0172] Array The various methods referred to herein can be adapted for use as an array that enables the analysis and / or detection of a set of biomarkers in a single experiment. For example, multiple mutations including biomarkers can be analyzed simultaneously. In particular, methods involving the use of nucleic acid reagents (e.g., probes, primers, oligonucleotides, etc.) are particularly suitable for adaptation to array-based platforms (e.g., microarrays). In some embodiments, the array contains one or more probes specific for detecting mutations in the biomarker of interest.
[0173] In one embodiment, a panel of a plurality of disclosed biomarkers is used. In one embodiment, the present disclosure includes at least one gene in Table 4 and / or Table 6, and / or at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6, and / or a reagent for quantifying the expression level of at least one of the HPV E6 and E7 genes, and includes a composition for detecting biomarkers associated with head and neck squamous cell carcinoma (HNSCC) in an individual. Further, and / or alternatively, the composition may include at least one normalization (e.g., housekeeping) gene. In one embodiment, the normalization gene may be KHDRBS1 and / or RPL30 or other normalization genes. This composition, in certain embodiments, can include primers and / or probes for any one of these genes, and the primers and / or probes are labeled with a detectable moiety described herein.
[0174] DNA sequencing In certain embodiments, the diagnosis of a biomarker of interest is performed by detecting, by sequencing, the sequence, genomic location or arrangement, and / or genomic copy number variation of a nucleic acid or panel of nucleic acids.
[0175] In some embodiments, the method can include obtaining nucleic acids from a tissue or body fluid sample from a subject to obtain a sample nucleic acid sequence of at least one gene and sequencing at least a portion of the nucleic acids. In certain embodiments, the method can include comparing the variant to known variants associated with HNSCC and determining whether the variant is a variant previously identified as being associated with HNSCC. Alternatively, the method can include identifying the variant as a new variant not previously characterized. If the variant is a new variant or in some instances a variant previously characterized (i.e., identified), the method can further include performing an analysis to determine whether the mutation is predicted to be detrimental to the expression of the gene and / or the function of the protein encoded by the gene. The method can further include using the variant profile (i.e., compilation of variants identified in a subject) to diagnose the presence of HNSCC or an increased risk of developing HNSCC.
[0176] For example, in certain embodiments, next generation (ultra-parallel sequencing) can be used. Alternatively, the Sanger base sequencing method can be used. Alternatively, a combination of next generation (ultra-parallel sequencing) and the Sanger base sequencing method can be used. Further and / or alternatively, the sequencing can include at least one single molecule synthesis time decoding method. Thus, in certain embodiments, multiple DNA samples are analyzed in a pool and samples showing variation are identified. Further and / or alternatively, in certain embodiments, multiple DNA samples are analyzed in multiple pools and individual samples showing the same variation in at least two pools are identified.
[0177] One of the conventional methods for performing base sequence determination is by chain termination and gel separation as described by Sanger et al., 1977, Proc Natl Acad Sci U S A, 74:5463-67. Another conventional base sequencing method involves chemical degradation of nucleic acid fragments. See Maxam et al., 1977, Proc. Natl. Acad. Sci., 74:560-564. Also, methods based on base sequencing by hybridization have been developed. See, for example, Harris et al., U.S. Patent Application Publication No. 20090156412. Each of these references is hereby incorporated by reference in its entirety.
[0178] In other embodiments, nucleic acid base sequencing is achieved by massively parallel sequencing of single molecules or groups of substantially identical molecules derived from single molecules (also known as "next-generation sequencing") by amplification by methods such as PCR. Massively parallel sequencing is shown, for example, in Lapidus et al., U.S. Patent No. 7,169,560, Quake et al., U.S. Patent No. 6,818,395, Harris, U.S. Patent No. 7,282,337, and Braslavsky et al., PNAS (USA), 100:3960-3964 (2003), the contents of each of which are hereby incorporated by reference.
[0179] In next-generation sequencing, PCR or whole genome amplification can be performed on the nucleic acid to obtain an amount of nucleic acid sufficient for analysis. In some formats of next-generation sequencing, amplification is not necessary because this method allows evaluation of DNA sequences from non-amplified DNA. Once determined, the sequence and / or genomic location and / or genomic copy number of the nucleic acid from the test sample is compared to a standard reference obtained from one or more individuals who were not known to have HNSCC at the time the sample was taken. Any differences between the sequences and / or genomic locations and / or genomic copy numbers of the nucleic acids from the test sample and the standard reference are considered variants.
[0180] In next-generation (ultra-parallel sequencing), all regions of interest are sequenced together, and the origin of each read sequence is determined by comparison (alignment) with a reference sequence. The regions of interest can be concentrated together in one reaction or combined after being concentrated separately and before base sequencing. In certain embodiments, and as described in more detail in the examples herein, DNA sequences derived from the coding exons of genes included in the assay are enriched by bulk hybridization of randomly fragmented genomic DNA to specific RNA probes. The same adapter sequence is ligated to the ends of all fragments, enabling enrichment of all fragments captured in hybridization by PCR using one primer pair in one reaction. Regions that are not captured as efficiently by hybridization are amplified by PCR using specific primers. Additionally, PCR using specific primers can be used to amplify exons where similar sequences (“pseudo-exons”) exist elsewhere in the genome.
[0181] In certain embodiments where ultra-parallel sequencing is used, the PCR products are ligated to form long stretches of DNA, which are sheared into short fragments (e.g., by acoustic energy). This step ensures that the ends of the fragments are evenly distributed across the region of interest. Subsequently, a stretch of dA nucleotides is added to the 3' end of each fragment, allowing the fragments to bind to a planar surface coated with oligo(dT) primers (a "flow cell"). Next, each fragment can be sequenced by extending the oligo(dT) primer with fluorescently labeled nucleotides. During each sequencing cycle, only one type of nucleotide (A, G, T, or C) is added, and only one nucleotide is incorporated using a chain-terminating nucleotide. For example, during the first sequencing cycle, fluorescently labeled dCTP can be added. This nucleotide is incorporated only into the growing complementary DNA strand that requires C as the next nucleotide. After each sequencing cycle, an image of the flow cell is taken to determine which fragments have extended. DNA strands that incorporated C emit light, while those that did not incorporate C appear dark. The chain termination is reversed to make the growing DNA strand extendable again, and this process is repeated for a total of 120 cycles. The images are converted into a string of bases, generally called "read data," that represents 25-60 bases at the 3' end of each fragment. Next, the read data is compared to a reference sequence of the DNA being analyzed. Since any string of 25 bases typically occurs only once in the human genome, most read data can be "aligned" to a specific location in the human genome. Finally, a consensus sequence for each genomic region is constructed from the available read data and can be compared to the exact sequence of the reference at that position. The differences between the consensus sequence and the reference are called sequence variants.
[0182] Detectable moiety In certain embodiments, specific molecules (e.g., nucleic acid probes, antibodies, etc.) used in accordance with and / or provided by the present invention include one or more detectable entities or moieties, i.e., such molecules are “labeled” with such entities or moieties.
[0183] Any of a wide variety of detectable agents can be used in the practice of the present disclosure. Suitable detectable agents include, but are not limited to: various ligands, radionuclides; fluorescent dyes; chemiluminescent agents (e.g., acridinium esters, stabilized dioxetanes, etc.); bioluminescent agents; spectrally resolvable inorganic fluorescent semiconductor nanocrystals (i.e., quantum dots); microparticles; metal nanoparticles (e.g., gold, silver, copper, platinum, etc.); nanoclusters; paramagnetic metal ions; enzymes; colorimetric labels (e.g., dyes, colloidal gold, etc.); biotin; digoxigenin; haptens; and proteins for which antisera or monoclonal antibodies are available.
[0184] In some embodiments, the detectable moiety is biotin. Biotin can bind to avidin (e.g., streptavidin, etc.), which is generally bound (directly or indirectly) to another moiety that is itself detectable (e.g., a fluorescent moiety).
[0185] Some non-limiting examples of detectable moieties that can be used are described below.
[0186] Fluorescent dyes In certain embodiments, the detectable moiety is a fluorescent dye. A number of known fluorescent dyes with a wide variety of chemical structures and physical properties are suitable for use in the practice of the present disclosure. The fluorescent detectable moiety is stimulated by a laser and the emitted light is captured by a detector. The detector can be a charge-coupled device (CCD) or a confocal microscope that records its intensity.
[0187] Suitable fluorescent dyes include, but are not limited to, fluorescein and fluorescein dyes (e.g., fluorescein isothiocyanate or FITC, naphthofluorescein, 4’,5’-dichloro-2’,7’-dimethoxyfluorescein, 6-carboxyfluorescein or FAM, etc.), hexachlorofluorescein (HEX), carbocyanine, merocyanine, styryl dyes, oxonol dyes, phycoerythrin, erythrosine, eosin, rhodamine dyes (e.g., carboxytetramethylrhodamine or TAMRA, carboxyrhodamine 6G, carboxy-X-rhodamine (ROX), lysamine rhodamine B, rhodamine 6G, rhodamine green, rhodamine red, tetramethylrhodamine (TMR), etc.), coumarin and coumarin dyes (e.g., methoxycoumarin, dialkylaminocoumarin, hydroxycoumarin, aminomethylcoumarin (AMCA), etc.), Q-DOTS, Oregon Green dyes (e.g., Oregon Green 488, Oregon Green 500, Oregon Green 514, etc.), Texas Red, Texas Red-X, SPECTRUM RED, SPECTRUM GREEN, cyanine dyes (e.g., CY-3, CY-5, CY-3.5, CY-5.5, etc.), ALEXA FLUOR dyes (e.g., ALEXA FLUOR 350, ALEXA FLUOR 488, ALEXA FLUOR 532, ALEXA FLUOR 546, ALEXA FLUOR 568, ALEXA FLUOR 594, ALEXA FLUOR 633, ALEXA FLUOR 660, ALEXA FLUOR 680, etc.), BODIPY dyes (e.g., BODIPY FL, BODIPY R6G, BODIPY TMR, BODIPY TR, BODIPY 530 / 550, BODIPY 558 / 568, BODIPY 564 / 570, BODIPY 576 / 589, BODIPY 581 / 591, BODIPY 630 / 650, BODIPY 650 / 665, etc.), IRDyes (e.g., IRD40, IRD 700, IRD 800, etc.), and the like. More examples of fluorescent dyes and methods suitable for conjugating fluorescent dyes to other chemical substances such as proteins and peptides are described, for example, in “The Handbook See the 9th Edition of "Handbook of Fluorescent Probes and Research Products", Molecular Probes, Inc., Eugene, Oregon. Preferred properties of the fluorescent label include high molar absorption coefficient, high fluorescence quantum yield, and photostability. In some embodiments, the labeled fluorophore exhibits absorption and emission wavelengths in the visible region (i.e., between 400 and 750 nm) rather than in the ultraviolet region of the spectrum (i.e., less than 400 nm).
[0188] The detectable moiety can include multiple chemical species such as those in fluorescence resonance energy transfer (FRET). Resonance transfer results in an overall improvement in emission intensity. For example, Ju et al. (1995) Proc. Nat’l, the entire content of which is incorporated herein by reference See Acad. Sci. (USA) 92:4347. To achieve resonance energy transfer, a first fluorescent molecule (“donor” phosphor) absorbs light and transfers it to a second fluorescent molecule (“acceptor” phosphor) through resonance of excited electrons. In one approach, both the donor dye and the acceptor dye can be linked together and attached to an oligonucleotide primer. Methods for linking donor and acceptor dyes to nucleic acids are described, for example, in U.S. Patent No. 5,945,526 to Lee et al., the entire content of which is incorporated herein by reference. Examples of donor / acceptor dye pairs that can be used include, for example, fluorescein / tetramethylrhodamine, IAEDANS / fluorescein, EDANS / DABCYL, fluorescein / fluorescein, BODIPY FL / BODIPY FL, and fluorescein / QSY7 dyes. See, for example, U.S. Patent No. 5,945,526 to Lee et al. Many of these dyes are also commercially available, for example, from Molecular Probes Inc. (Eugene, Oregon). Examples of suitable donor fluorophores include 6-carboxyfluorescein (FAM), tetrachloro-6-carboxyfluorescein (TET), 2'-chloro-7'-phenyl-1,4-dichloro-6-carboxyfluorescein (VIC), and the like.
[0189] Enzyme In certain embodiments, the detectable moiety is an enzyme. Examples of suitable enzymes include, but are not limited to, enzymes used in ELISA, such as horseradish peroxidase, β-galactosidase, luciferase, alkaline phosphatase, and the like. Other examples include β-glucuronidase, β-D-glucosidase, urease, glucose oxidase, and the like. The enzyme can be attached to the molecule using linker groups such as carbodiimide, diisocyanate, glutaraldehyde, and the like.
[0190] Radioisotope In certain embodiments, the detectable moiety is a radioisotope. For example, the molecule may be isotopically labeled (i.e., may contain one or more atoms replaced with atoms having an atomic weight or mass number different from that normally found in nature), or the isotope may be attached to the molecule. Non-limiting examples of isotopes that can be incorporated into a molecule include isotopes of hydrogen, carbon, fluorine, phosphorus, copper, gallium, yttrium, technetium, indium, iodine, rhenium, thallium, bismuth, astatine, samarium, and lutetium (i.e., 3H, 13C, 14C, 18F, 19F, 32P, 35S, 64Cu, 67Cu, 67Ga, 90Y, 99mTc, 111In, 125I, 123I, 129I, 131I, 135I, 186Re, 187Re, 201T1, 212Bi, 213Bi, 21lAt, 153Sm, 177Lu).
[0191] Dendrimer In some embodiments, signal amplification is achieved using a labeled dendrimer as the detectable moiety (see, e.g., Physiol Genomics 3:93-99, 2000), the entire contents of which are hereby incorporated by reference in their entirety. Fluorescently labeled dendrimers are available from Genisphere (Montvale, NJ). These can be chemically conjugated to oligonucleotide primers by methods known in the art.
[0192] System In certain embodiments, the present disclosure provides a system for performing the methods disclosed herein and / or for using the compositions described herein. In certain embodiments, the system may include a kit. Alternatively, the system may include computerized instructions and / or reagents for performing the methods disclosed herein.
[0193] Kit In certain embodiments, the present disclosure provides kits for use in accordance with the methods and compositions disclosed herein. Generally, the kits include one or more reagents for detecting a biomarker of interest. Suitable reagents can include nucleic acid probes and / or antibodies or fragments thereof. In some embodiments, suitable reagents are provided in the form of an array, such as a microarray, or a mutation panel. The kit can further include a reagent that functions as a positive control for a biomarker (i.e., gene) of interest.
[0194] In one embodiment, a panel of the disclosed plurality of biomarkers is used. In one embodiment, the present disclosure provides a kit for detecting biomarkers associated with head and neck squamous cell carcinoma (HNSCC) in an individual, the kit comprising reagents for quantifying the expression levels of at least one gene in Table 4 and / or Table 6, and / or at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1, or HSD17B6, and / or at least one of the HPV E6 and E7 genes. Additionally and / or alternatively, the kit can include at least one normalization (e.g., housekeeping) gene and / or reagents for detecting such a housekeeping gene. In one embodiment, the normalization gene can be KHDRBS1 and / or RPL30 or other normalization genes. In some embodiments, the kit can include a positive control for any of the disclosed biomarkers and / or normalization genes.
[0195] In some embodiments, the provided kit further includes reagents for performing the various detection methods described herein (e.g., RT-PCR, sequencing, hybridization, primer extension, multiplex ASPE, immunoassay, etc.). For example, the kit may optionally include buffers, enzymes, and / or reagents used in the methods described herein for amplifying nucleic acids via, e.g., RT-PCR, primer-directed amplification, performing ELISA experiments, etc. In certain embodiments, the kit may include primers and / or probes for any one of these genes, and the primers and / or probes are labeled with a detectable moiety described herein.
[0196] In some embodiments, the provided kit further includes a control indicating a healthy individual, e.g., a nucleic acid and / or protein sample from an individual without the disease and / or syndrome of interest. Alternatively, the kit may include a positive control containing a known amount of one (or more than one) biomarker gene to be measured. The kit may also include instructions on how to determine whether an individual has the disease and / or syndrome of interest or is at risk of developing the disease and / or syndrome of interest.
[0197] In some embodiments, a computer-readable medium encoding information corresponding to a biomarker of interest is provided. Such a computer-readable medium may be included in the kit of the present invention.
[0198] Method for identifying HNSCC markers Data mining In certain embodiments of the present disclosure, biomarkers are identified using data mining approaches. For example, in some instances, public databases (e.g., PubMed, The Cancer Genome Atlas (TCGA)) can be searched for genes that have been shown to be (directly or indirectly) associated with a particular disease and / or genes that are differentially expressed in cancer compared to normal tissue. Such genes can then be evaluated as biomarkers.
[0199] molecule In certain embodiments, the present disclosure includes a method of identifying a biomarker for a syndrome or disease of interest (i.e., a variant in a nucleic acid sequence that is statistically significantly associated with HNSCC). For example, genes of interest and potential normalization genes can be identified by evaluating gene expression in tissue samples isolated from patients with head and neck cancer using random forest analysis (see, e.g., L. Breiman, “Random Forests” Machine Learning, 2001, 45:5 - 32), as discussed in detail herein. In this approach, a random forest is a combination of tree predictors such that each tree depends on the values of a random vector sampled independently and all trees in the forest have the same distribution.
[0200] Alternatively, the genes and / or genomic regions to be assayed for new markers can be selected based on their importance in biochemical pathways that show genetic linkage and / or biological causality to the syndrome and / or disease of interest. Or, the genes and / or genomic regions to be assayed for a marker can be selected based on genetic linkage to DNA regions that are genetically related to the inheritance of HNSCC in families. Or, the genes and / or genomic regions to be assayed for a marker can be systematically evaluated to cover specific regions of the chromosome that have not yet been evaluated.
[0201] In other embodiments, the gene or genomic region being evaluated for the new marker can be part of a biochemical pathway that may be associated with the development of the syndrome and / or disease of interest (e.g., HNSCC). The variant and / or combination of variants can be evaluated for its clinical significance based on one or more of the following methods. If the variant and / or combination of variants is reported or known to occur more frequently in nucleic acids from subjects having the syndrome and / or disease of interest than in subjects not having it, it is considered to be at least potentially more likely to be associated with the syndrome and / or disease of interest. If the variant and / or combination of variants is reported or known to be transmitted exclusively or preferentially to individuals having the syndrome and / or disease of interest, it is considered to be at least potentially more likely to be associated with the syndrome and / or disease of interest. Conversely, if the variant is found with similar frequency in both populations, the likelihood of its association with the development of the syndrome and / or disease of interest is low.
[0202] If a variant or combination of variants has been reported or is known to have an overall deleterious effect on the function of a protein or biological system in an experimental model system suitable for measuring the function of the protein or biological system, and further, if the variant or combination of variants affects one or more genes known to be associated with the syndrome and / or disease of interest, then it is considered to be at least potentially susceptible to the syndrome and / or disease of interest. For example, if a variant or combination of variants is predicted based on its predicted effect on the sequence and / or structure of a protein or nucleic acid to have an overall deleterious effect on protein or gene expression (i.e., result in a nonsense mutation, frameshift mutation, or splice site mutation, or even a missense mutation), and further, if the variant or combination of variants affects one or more genes known to be associated with the syndrome and / or disease of interest, then it is considered to be at least potentially susceptible to the syndrome and / or disease of interest.
[0203] Also, in certain embodiments, the total number of variants may be important. In a test sample, if one or several variants are detected, individually or in combination, that are evaluated to potentially be associated with at least the syndrome and / or disease of interest, then the individual in whose genetic material this variant or these variants are detected may be diagnosed as having or being at high risk of developing the syndrome and / or disease of interest.
[0204] For example, the disclosure herein provides a method for diagnosing the presence of HNSCC or an increased risk of developing the same in a subject. Such a method may include obtaining nucleic acids from a sample of tissue or body fluid. The method may include determining the expression of at least one gene in both normal and cancerous tissues to identify potential biomarkers of interest. The method may further include sequencing the nucleic acids or determining the genomic location or copy number of the nucleic acids to detect whether there are one or more variants in the nucleic acid sequence or genomic location or copy number. The method may further include evaluating the clinical significance of one or more variants. Such an analysis may include assessing the degree of association of the variant sequence in an affected population (i.e., subjects having the disease). Such an analysis may also include analyzing the degree of effect that the mutation may have on gene expression and / or protein function. The method may also include diagnosing the presence of HNSCC or an increased risk of developing the same based on the evaluation.
Example
[0205] The following examples serve to illustrate certain aspects of the present disclosure. These examples are not intended to be limiting in any way.
[0206] Example 1 - Literature-Based Identification of Potential HNSCC Markers A preliminary literature search was performed to identify markers associated with HNSCC. Table 1 shows the types of markers found and the number of markers, and Table 2 shows the identified potential biomarkers.
[0207]
Table 1
[0208]
Table 2-1
Table 2-2
Table 2-3
[0209] Based on this initial search, it was determined to track markers related to differential gene expression.
[0210] Example 2 - Identification of Biomarkers Using the TCGA Database Data from the Cancer Genome Atlas (TCGA) database were used to identify markers showing differential expression in HNSCC. The TCGA database (RNASeqV2) contains data on 18,379 genes available for differential expression. The data include clinical information (e.g., age, smoking, stage, treatment, and survival); copy number; methylation; gene expression; mutations, and information on microRNA expression.
[0211] The HNSCC data consist of 530 samples from four tumor sites: oral cavity (n = 320; 60.4%), oropharynx (n = 82; 15.5%), larynx (n = 117; 22.1%), and hypopharynx (n = 10; 1.5%). An additional 44 samples are from adjacent normal tissue. Of the total 530 samples, 70 are positive for human papillomavirus (HPV), 279 are HPV negative, and 181 have an unknown or undetermined HPV status.
[0212] Random forest analysis was performed to identify genes that are strong predictors for classifying HNSCC from normal samples. In this analysis, samples with 50% of the data with reported value and genes with less than 2-fold change in expression (Wilcox test, adjusted p-value < 0.001) were discarded. In each round of the analysis, 75% of the samples were used as the training set and 25% of the samples were used as the test set. The data was optimized for kappa and 10-fold cross-validation was performed (repeated 10 times and the performance was averaged). The top 20 strong predictors were identified and the whole process was repeated 4 times. The gene lists obtained from each run are shown in Table 3, and the combination of 36 unique genes is shown in Table 4. The data in Table 3 is shown in order from the highest rank (20) to the lowest (1).
[0213]
Table 3
[0214]
Table 4-1
Table 4-2
[0215] Next, the top 4 genes from each of the 4 analyses were selected to obtain an initial candidate list of 8 unique genes listed in Table 5.
[0216]
Table 5
[0217] The results of several statistical analyses of individual genes (i.e., accuracy, kappa, sensitivity, and specificity) are also described in order of specificity and shown in Figure 1.
[0218] Example 3 - Gene Panel An analysis was performed to determine whether the use of gene panels was expected to improve assay performance. As shown in Figure 2, the use of four or five gene panels should significantly improve gene performance. The panels were constructed by forming a two-marker panel by adding the most informative marker (CAB39L) to the next most informative marker (ADAM12), and then forming a three-marker panel by adding the next most informative marker (NRG2). Next, each of the other six markers was added and the predicted performance was evaluated (Figure 2, upper table). The results showed that a four-marker panel of CAB39L, ADAM12, NRG2, and GRIN2D provided the highest levels of accuracy, kappa value, sensitivity, and specificity. The results for the five-gene panel are shown in the lower table (Figure 2). It was found that the improvement by adding more than four to five genes was minimal (e.g., the graph in Figure 2). Nevertheless, such panels can be useful if there are technical issues with one of the markers identified as one of the top four to five markers.
[0219] Example 4 - Gene Expression by Tumor Site Most HNSCCs are found either in the oral cavity (mouth) or in the oropharynx (throat). An analysis was performed to determine whether the same gene panels developed using the entire HNSCC dataset could also be used to distinguish oral or oropharyngeal HNSCC from normal tissue. The results are shown in Figures 3 and 4 for eight markers, CAB39L, ADAM12, SH3BGRL2, NRG2 (Figure 3); and COL13A1, GRIN2D, LOXL2, and HSD17B6 (Figure 4); using the TCGA dataset including oral and oropharynx as the majority of the TCGA samples (out of a total of 530 samples, HNSCC oral (320) and normal oropharynx (82) = 402 (402 / 530 = 76%), and out of a total of 44 samples, normal oral (30) and oropharynx (3) = 33 (33 / 44 = 75%)).
[0220] For both sites and the larynx, the markers were found to show very different levels of expression in normal and cancerous tissues. In FIGS. 3 and 4, the three data sets (larynx, oral cavity, oropharynx) at the left end of the x-axis are the expression levels of normal tissues, and the four data sets (hypopharynx, larynx, oral cavity, oropharynx) at the right end of the x-axis are the expression levels of cancerous tissues with the hypopharynx data combined. It can be seen that similar distributions were present in normal and HNSCC sites (i.e., the normal levels were similar regardless of tissue, and the HNSCC levels were similar regardless of tissue). For example, CAB39L was found to be expressed at significantly lower levels in HNSCC than in normal tissues for the larynx, oral cavity, and oropharynx, whereas ADAM12 was found to be expressed at significantly higher levels in HNSCC than in normal tissues. FIG. 5 shows a statistical compilation of data showing the results by markers for all HNSCC samples compared to samples from the oral cavity and oropharynx. When all samples were compared to the oral cavity and oropharynx, the changes in accuracy, sensitivity, and specificity were minimal. For all eight markers, the sample set decreased by approximately 25%, the change in the median gene expression level was minimal, and there was a significant difference in distribution (HNSCC vs. normal). For both sets, the Mann-Whitney p-value was less than 0.0001.
[0221] Example 5 - Analysis of differential expression of the TCGA gene set for the median fold expression compared to the overlap rate in expression. The graph at the top of FIG. 6A shows the number of times the markers from 36 genes in Table 4 initially selected by random forest analysis were identified from four repetitions of the random forest analysis compared to the median rank (discrimination ability) of the markers. It can be seen that as the number of times a marker is repeatedly identified from the random forest analysis increases, the median rank tends to increase. The four tables below the graph list the markers grouped by the number of times they were repeatedly identified from the random forest analysis and their median ranks. Combining the graph and the tables gives a measure that can enable the ranking of the 36 genes in Table 4 identified from the random forest analysis, first by the number of times repeated and then by the median rank.
[0222] The left graph in FIG. 6B shows the analysis of differential expression of certain selected HNSCC markers of the present disclosure compared to the entire TCGA HNSCC gene set. The x-axis shows the median fold change in expression as either an increase in gene expression (data points to the right of 0) or a decrease in gene expression (data points to the left of 0). The y-axis shows the overlap rate in gene expression between HNSCC and normal tissues. The disclosed markers in Table 4 (n = 36) show either a significant increase or decrease in gene expression and a very low overlap rate compared to other genes in the database. The entire TCGA HNSCC gene set is RNASeq data using over 80% of the genes in the samples represented (i.e., out of 18,379 genes, 16,161 (88%) had reported values for both HNSCC and normal samples). The x-axis shows the median fold change (= HNSCC / normal). It was found that in HNSCC, 8,352 (52%) genes increased compared to normal, 7,809 (48%) genes decreased, 1,387 genes (8.6%) increased more than 2-fold, and 1,701 (10.5%) decreased more than 2-fold. For an increase in gene expression in HNSCC, the cut-off is the 5th percentile of HNSCC and the 95th percentile of normal (e.g., the inset for GRIN2D in FIG. 6B).
[0223] The dotted line crossing the graph of Figure 6B indicates that the duplication rate in the expression of 9 markers repeated 4 times from the random forest analysis was less than 20% (in the range of 0 - 19%) (i.e., these markers are below the dotted line). Further, the dotted line crossing the graph of Figure 6B indicates that 23 out of the 36 unique markers in Table 4, i.e., 23 / 36 = 64%, have a duplication rate in expression less than 20% because the markers gather below the dotted line. Table 6 lists 45 additional genes with a duplication rate in expression of 20% or less that were not identified by the random forest analysis. Genes with a duplication rate in expression of 20% or less, such as GLT25D1 identified in Figure 6B, can be considered additional biomarkers useful for the classification from normal samples of HNSCC.
[0224]
Table 6-1
Table 6-2
[0225] The data in Figure 6B can be compared with the data in Figures 7 and 8. Figure 7 shows a similar analysis of the median fold expression and the duplication rate in expression from TCGA HNSCC RNASeq data, and this example shows tissue markers and saliva markers identified by a literature search highlighted on the graph. The upper and lower panels show the results of markers identified in both tissue (upper panel) and saliva (lower panel). The specific literature markers in Figure 7 show some evidence of differential expression, but only some markers show a high level of differential expression with a low duplication rate. Based on this analysis, the markers MAL, MMP1, CEP55, CENPA, AURKA, and FOXM1 are thought to be the most informative additional biomarkers and can be included in the disclosed methods and compositions.
[0226] Figure 8 shows a similar analysis of the median fold expression and the rate of duplication in expression from TCGA HNSCC RNASeq data. This example shows the standardized markers used in the tissues (upper panel) and saliva (lower panel) identified by literature search. These markers show little change in expression, and it can be seen that there is a significant overlap between HNSCC and normal. An ideal standardized marker should have minimal variation and an expression level similar to the gene panel of interest.
[0227] Example 6 - Identification of Potential Standardization Genes or Housekeeping Genes The TGCA database was analyzed to identify potential standardization genes using three criteria: (1) the minimum median fold change in expression between HNSCC and normal tissues; (2) the minimum interquartile range (IQR) in both HNSCC and normal (IQR is defined as the gene expression at the 75th quartile / the gene expression at the 25th quartile); and (3) the median of the expression levels close to the gene panel of interest to facilitate an experimental comparison between the potential candidate gene expression and the standardized gene.
[0228] The analysis is summarized in Figure 9A. The left panel shows a plot of the median fold change in gene expression in HNSCC and normal (x-axis) against the average IQR for both normal and cancer cells (=[HNSCC I.Q.R. + normal I.Q.R.] / 2) (y-axis). In this figure, a positive number on the x-axis corresponds to an increase in gene expression in cancer cells compared to normal, and a negative number corresponds to a decrease in gene expression in cancer cells compared to normal. Data from a total of 16,161 genes were analyzed (left panel). This corresponds to all genes having data for both normal and HNSCC in the TCGA database, and thus, 88% (16,151 / 18,379) of the total number of genes in the TCGA database. Again, in this case, it was found that in HNSCC, 8,352 (52%) genes increased and 7,809 (48%) genes decreased compared to normal, 1,387 genes (8.6%) increased more than 2-fold, and 1,701 (10.5%) decreased more than 2-fold.
[0229] The most interesting potential normalization genes are those with a fold change (x-axis) of 0 and an average IQR of 1 (the region enclosed by the circle on the plot). The central plot shows data for genes with a median fold change of less than 2 and an average IQR of less than 2 (n = 7,949 genes). The right panel shows gene expression data for 7,949 candidate normalization genes. Genes with a median expression were considered to be of the greatest interest. Based on this analysis, KHDRBS1 (KH domain-containing, RNA-binding, signal transduction-related protein 1) was identified as a normalization gene of interest. Several other more common normalization genes, such as RPLPO, RPL10, RPL30, GAPDH, etc., were identified in the central panel. Data from the TCGA database for KHDRBS1 are shown in Table 6 below. Figure 9B shows that KHDRBS1 exhibited similar characteristics (low fold change expression and low IQR) across many cancer types.
[0230] The average expression levels of 11.60 - 11.90 are higher than the proposed panel markers in the range of 1.52 (NRG2 in HNSCC) - 11.44 (SH3BGRL2 in normal). Nevertheless, since this is within the range of cancer-specific markers, it should be a good normalization gene. Note that for FFPE samples that may contain substantially degraded RNA, an amplicon length of less than 100 bp is preferred.
[0231]
Table 7
[0232] These data can be compared with data for the well-known housekeeping (normalization gene) GAPDH (Table 7). The median expression of 16.3 - 16.50 is approximately 30 times higher than the gene (SH3BGRL2) showing the highest level of expression in the candidate panel discussed above. Thus, GAPDH may not be very useful as a marker for the HNSCC panel of the present disclosure discussed above.
[0233] Example 7 - Evaluation of Expression Assays An experiment was conducted to compare the expression levels determined from the data of the TCGA database (RNA Seq evaluation of gene expression) with the expression levels measured using droplet digital PCR (ddPCR). The results are shown in Figure 10. This figure shows data demonstrating the reproducibility of ddPCR analysis of ADAM12 and SH3BGRL2 by ddPCR in tongue squamous cell carcinoma (SCC) and normal tissues (oral mucosa) (the table above in Figure 10). Also shown is a comparison of ddPCR data (the table below) and TCGA RNA Seq data (the table in the middle) regarding the gene expression levels of ADAM12 and SH3BGRL2 in cancer compared to normal tissues. It can be seen that when measured by both methods, the expression of ADAM12 in cancer is significantly increased compared to normal, and the expression of SH3BGRL2 in cancer is significantly decreased compared to normal. In these experiments, two aliquots from the same sample were analyzed. For the buccal samples, one of the samples had an expression level that was too low to be accurately measured. The ddPCR values were generally lower than the TCGA RNA Seq data, but the trend was the same for both markers (see Figure 10). The reproducibility was good up to 1 copy / μL.
[0234] Figure 11 shows further digital PCR data for three formalin-fixed paraffin-embedded patient samples (DA1081983; DR1041686; DA0063595) and one URNA control sample derived from cell culture cancer tissue. URNA is a universal human reference RNA available from Agilent (catalog number 740000). It consists of 10 human cancer cell lines that function as a consistent control for standard dataset comparisons. Sampling was performed either in two or three ways according to embodiments of the present disclosure. Again, in this case, a substantial decrease in SH3BGRL2 expression in cancer compared to normal was found. It can be seen that the copy number per 1 μL from the URNA control is much larger compared to the cross-linked and probably fragmented RNA from the FFPE samples, perhaps due to the intact nature of the isolated RNA.
[0235] Figure 12 shows a RNA titration experiment using ddPCR and markers SH3BGRL2 (labeled with FAM) and KHDRBS1 (labeled with HEX). In this experiment, RNA was isolated from FFPE samples (or URNA was used as a positive control) and diluted 2- or 10-fold. cDNA was generated using standard techniques and ddPCR was used to detect the presence of biomarker SH3BGRL2 (amplification sequence labeled with FAM) or the housekeeping gene KHDRBS1 (labeled with HEX). Results for three samples (#1, #3, or #5) are shown. A good correlation is seen between the ratios of the marker and the housekeeping gene at various concentrations, except for very dilute samples (i.e., approaching or less than 1 copy / μL). This indicates that the range of sample concentrations that can be measured using this assay method is appropriate.
[0236] Figure 13 shows the relative abundance of biomarker SH3BGRL2 and housekeeping RNA KHDRBS1 in FFPE samples compared to the positive control URNA (left panel), the ratio of biomarker SH3BGRL2 / KHDRBS1 in cancer cells compared to normal cells and URNA (central panel), and the relative amount of KHDRBS1 to SH3BGRL2 in cancer and normal cells measured using RNASeq (right panel). Again, a consistent pattern is seen for the biomarker regardless of the method of measurement (absolute values may vary).
[0237] Figure 14 shows the measurement of SH3BGRL2 as a singleplex assay (i.e., only SH3BGRL2-FAM generated by PCR was measured in either normal or cancer-derived samples compared to a duplex reaction where both SH3BGRL2 and KHDRBS1 were measured). The results are overall very similar.
[0238] Example 8 - Differential Expression in FFPE Tissue Samples by ddPCR Figure 15 shows the measurement of five biomarkers and the housekeeping gene KHDRBS1 from 22 benign and 8 head and neck cancer FFPE samples. RNA was extracted from FFPE tissues using the Roche High Pure FFPET RNA Isolation Kit essentially according to the manufacturing protocol. The five panels show the copy number / μL from duplex ddPCR of five biomarkers (SH3BGRL2, KRT4, EMP1, LOXL2, and ADAM12) and the housekeeping gene KHDRBS1 from 22 benign and 8 head and neck cancer FFPE samples. KHDRBS1 showed very similar distribution patterns and copy number / μL across all samples and all assays. In contrast, the biomarkers showed various distributions at >3-log 10 copy number / μL. One HNSCC sample was "no call" for both biomarkers and KHDRBS1 across all assays.
[0239] The left graph in Figure 16 shows the expression of the five biomarkers normalized to KHDRBS1. Dividing the biomarker copy number / uL by the copy number / uL of the housekeeping gene KHDRBS1 from the duplex ddPCR reaction gives the normalization or ratio of the biomarker for each sample. The right graph in Figure 16 is the RNASeq expression data of the same biomarkers from the HNSCC TCGA dataset. The table shows the median fold change in expression of each biomarker from the ddPCR experiment compared to the TCGA data. Both datasets show the same genes downregulated in cancer (SH3BGRL2, KRT4, and EMP1) and the same genes upregulated in cancer (LOXL2 and ADAM12). The results of ddPCR are consistent with the TCGA dataset, but the magnitude of the change is different.
[0240] Figure 17. The ddPCR score was developed to separate cancer from normal samples by determining the difference between (the total log of upregulated genes) - (the total log of downregulated genes). When adding biomarkers, the formula becomes (logLOXL2 + logADAM12) - (logSH3BGRL2 + logEMP1 + logKRT4). The left panel shows the ddPCR scores of cancer samples plotted against those of normal samples. When the ddPCR cutoff score > 0.24, the specificity of the assay is 95.4% and the sensitivity is 85.7%. The right plot shows the receiver operator characteristic (ROC) analysis, with AUC = 0.961 and p = 0.0003. The specificity from ddPCR is similar to that of the TCGA dataset (see Figure 2), but the sensitivity is lower from ddPCR than from TCGA, probably due to differences in sample size and gene selection.
[0241] Example 9 - HPV16 E6 and E7 Expression in FFPE HNSCC Tissue Samples Figure 18 shows the correlation between E6 and E7 HPV16 expression and p16 from FFPE HNSCC tissue samples. The upper table shows the HPV16 ddPCR copy numbers / μL of E6, E7, and the housekeeping gene KHDRBS1. The lower table shows the results of ddPCR copy numbers / μL of E6, E7, and the housekeeping gene KHDRBS1 for 10 p16-negative samples. The right plot shows, for 4 p16-positive samples, the normalized ddPCR ratio of the biomarker divided by the housekeeping gene KHDRBS1. Two samples with copy numbers / μL of both HPV16 E6 and E7 greater than "no call" showed normalized ddPCR expression of E7 approximately 5-fold greater than that of E6. Furthermore, all p16-negative samples also had negative (no call) E6 and E7 expression by ddPCR, both in preparation and replication. There was a very good overall agreement between p16 by IHC and E6 or E7 expression by ddPCR (Cohen's kappa = 0.81).
[0242] Example 10 - Isolation of RNA and Gene Expression in Saliva Samples In some cases, saliva can be used as a biological sample. Figure 19 shows the yield of RNA from saliva. Saliva samples were collected using a DNA Genotek CP - 190 Human RNA Collection Device. After sample collection, each sample was thoroughly mixed, incubated at 50 °C for 2 hours, and stored at - 20 °C until processed. To each aliquot of saliva to be processed, a DNA Genotek neutralizing agent solution (catalog number RELONN - 5) at 1 / 10 of the sample volume was added. RNA was purified using a Roche High Pure RNA Paraffin Kit (catalog number 03 270 289 001). The left - hand table shows 15 saliva RNA samples, μg RNA / 2 mL saliva calculated from 250 μL of saliva sample preparation, and the A260 / A280 ratio from each saliva sample. The median μg of isolated RNA was 5.8 μg, and the median A260 / A280 ratio was 2.05. The right - hand scatter plot shows the same data in box - and - whisker format, where the whiskers are the maximum and minimum, and the box around the 75th and 25th percentiles and the line through the median are shown.
[0243] Figure 20 shows the measurement of five biomarkers and the housekeeping (HK) gene RPL30 from the 15 saliva RNA samples of Figure 19. The left panel shows the copy number / μL from duplex ddPCR of five biomarkers (LOXL2, SH3BGRL2, CRISP3, EMP1, and KRT4) and the housekeeping gene RPL30 from 15 saliva samples. The distribution of biomarkers ranged from 0.38 to 650 copies / μL. The right panel should be the housekeeping gene for each duplex ddPCR. The housekeeping gene (RPL30) averaged 8 - 170 copies / μL, and the CV% was 6 - 27% across the five types of duplex ddPCR reactions. One sample averaged 1.3 copies / μL with a CV% of 57%.
[0244] Figure 21 shows the normalized ddPCR from saliva compared to TCGA RNASeq. The left graph shows the expression of biomarkers normalized to the housekeeping gene RPL30. Dividing the copy number / μL of the biomarker by the copy number / μL of the housekeeping gene RPL30 in the duplex ddPCR reaction gives the normalization or ratio of the biomarker for each sample. One sample with an average HK gene of 1.3 copies / uL and samples with biomarkers of "no call" or less than 1 copy / uL are excluded. The range of normalized ddPCR expression was 0.018 - 9.7 = 540-fold. The right graph in Figure 21 is the RNASeq expression data of the same biomarker from "normal" in the HNSCC dataset. The table shows the median fold increase in expression relative to LOXL2 for each biomarker from the ddPCR experiment compared to TCGA data. The median fold increase in expression from saliva by ddPCR tends to be similar to the TCGA (tissue) dataset, but the magnitude of change varies.
[0245] Example 11 - Embodiment The present disclosure includes, but is not limited to, the following embodiments.
[0246] A. A method for detecting a biomarker associated with head and neck squamous cell carcinoma (HNSCC) in an individual, comprising: obtaining a sample from the individual; measuring the amount of an expression product from a gene comprising at least one gene in Table 4 and / or Table 6; and a method comprising the steps of:
[0247] A.2. The method according to any of the preceding paragraphs, wherein the gene comprises at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6.
[0248] The method according to any of the preceding paragraphs, wherein the gene comprises at least four of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 and HSD17B6.
[0249] The method according to any of the preceding paragraphs, further comprising measuring the amount of an expression product from at least one of the HPV E6 gene and / or the HPV E7 gene.
[0250] The method according to any of the preceding paragraphs, wherein the gene consists of at least four of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 and HSD17B6, and an expression product from at least one of the HPV E6 and / or HPV E7 genes.
[0251] The method according to any of the preceding paragraphs, further comprising measuring the amount of a normalization gene such as KHDRBS1, or RPL30, or another normalization gene.
[0252] The method according to any of the preceding paragraphs, wherein the measuring step comprises measuring mRNA.
[0253] The method according to any of the preceding paragraphs, wherein the measuring step comprises an immunoassay.
[0254] The method according to any of the preceding paragraphs, comprising measuring the expression of at least four of the genes.
[0255] The method according to any of the preceding paragraphs, wherein the sample comprises serum, tissue, FFPE, saliva or plasma.
[0256] The method according to any of the preceding paragraphs, comprising comparing the expression level with a control value from a normal population.
[0257] A. A method according to any of the preceding paragraphs, wherein a difference between gene expression and a control value in 12 individuals indicates that the individual may have HNSCC (i.e., is a diagnosis), or is prone to developing HNSCC (i.e., has a high risk).
[0258] B. A method for identifying a marker associated with HNSCC in an individual, comprising identifying at least one marker that is increased or decreased in expression in head and neck squamous cell carcinoma (HNSCC) but is not expressed in HNSCC disease compared to a normal control.
[0259] B.2 The method according to B.1, wherein the gene comprises at least one of at least one of the genes in Table 4 and / or Table 6, and / or at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6.
[0260] B.3 The method according to any of B.1 - B.2, wherein the gene comprises at least 4 of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 and HSD17B6.
[0261] B.4 The method according to any of B.1 - B.3, further comprising measuring the amount of an expression product from at least one of the HPV E6 gene and / or the HPV E7 gene.
[0262] B.5 The method according to any of B.1 - B.4, wherein the gene consists of at least 4 of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 and HSD17B6, and an expression product from at least one of the HPV E6 and / or HPV E7 genes.
[0263] The method according to any one of B.1 to B.5, further comprising measuring the amount of a normalization gene such as KHDRBS1, or RPL30, or another normalization gene.
[0264] The method according to any one of B.1 to B.6, wherein the measuring step comprises measuring mRNA.
[0265] The method according to any one of B.1 to B.7, wherein the measuring step comprises an immunoassay.
[0266] The method according to any one of B.1 to B.8, comprising measuring the expression of at least 4 of the genes.
[0267] The method according to any one of B.1 to B.9, wherein the sample comprises serum, tissue, FFPE, saliva or plasma.
[0268] The method according to any one of B.1 to B.10, wherein the difference between the gene expression in the individual and the control value indicates that the individual may have HNSCC (i.e., is a diagnosis), or is prone to developing HNSCC (i.e., has a high risk).
[0269] C.1 Obtaining a sample from an individual; Measuring the amount of at least one expression product from at least one gene in Table 4 and / or Table 6; Comparing the expression of at least one gene in Table 4 and / or Table 6 in the sample with the control value of the gene expression product A method for detecting susceptibility to head and neck squamous cell carcinoma (HNSCC) in an individual, comprising:
[0270] The method according to C.1, wherein the gene comprises at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6.
[0271] The method according to any one of C.1 to C.2, wherein the gene comprises at least four of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 and HSD17B6.
[0272] The method according to any one of C.1 to C.3, further comprising measuring the amount of an expression product from at least one of the HPV E6 gene and / or the HPV E7 gene.
[0273] The method according to any one of C.1 to C.4, wherein the gene consists of at least four of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 and HSD17B6, and an expression product from at least one of the HPV E6 and / or HPV E7 genes.
[0274] The method according to any one of C.1 to C.5, further comprising measuring the amount of a normalization gene such as KHDRBS1, or RPL30, or another normalization gene.
[0275] The method according to any one of C.1 to C.6, wherein the measuring step comprises measuring mRNA.
[0276] The method according to any one of C.1 to C.7, wherein the measuring step comprises an immunoassay.
[0277] The method according to any one of C.1 to C.8, comprising measuring the expression of at least four of the genes.
[0278] The method according to any one of C.1 to C.9, wherein the sample comprises serum, tissue, FFPE, saliva or plasma.
[0279] The method according to any one of C.1 to C.10, wherein the difference between the gene expression and the control value in 11 individuals indicates that the individual may have HNSCC or is prone to developing HNSCC (i.e., high risk).
[0280] D.1 A composition for detecting biomarkers associated with head and neck squamous cell carcinoma (HNSCC) in an individual, comprising a reagent for quantifying the expression level of at least one gene in Table 4 and / or Table 6.
[0281] D.2 The composition according to any D.1, wherein at least one gene comprises at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6.
[0282] D.3 The composition according to D.1 to D.2, wherein at least one gene comprises at least 4 of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6.
[0283] D.4 The composition according to any one of D.1 to D.3, further comprising at least one reagent for quantifying the expression level of at least one HPV E6 and / or E7.
[0284] D.5 The composition according to any one of D.1 to D.4, wherein at least one gene consists of at least 4 of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6 and at least one HPV E6 and / or E7.
[0285] D.6 The composition according to any one of D.1 to D.5, further comprising at least one reagent for measuring at least one normalization gene such as KHDRBS1 or RPL30, or another normalization gene.
[0286] D.7 A composition according to any one of D.1 to D.6, wherein the reagent detects mRNA.
[0287] D.8 A composition according to any one of D.1 to D.7, wherein the reagent detects protein.
[0288] D.9 A composition according to any one of D.1 to D.8, wherein the reagent comprises at least one primer and / or probe for any one of these genes, and at least one primer and / or probe is labeled with a detectable moiety.
[0289] D.10 A composition according to any one of D.1 to D.9, wherein a difference between gene expression in an individual and a control value indicates that the individual may have HNSCC (i.e., is a diagnosis) or is prone to developing HNSCC (i.e., has a high risk).
[0290] E.1 A kit comprising a composition according to any one of the preceding paragraphs.
[0291] E.2 The kit according to E.1, further comprising instructions for measuring at least one gene and / or for determining whether a value differs from a control value.
[0292] E.3 The kit according to any one of E.1 to E.2, comprising at least one positive control for at least one normalization gene such as KHDRBS1 or RPL30, or at least one positive control for another normalization gene.
[0293] E.4 The kit according to any one of E.1 to E.3, further comprising at least one positive control for any one of the genes in Table 4 and / or Table 6.
[0294] E.5 The kit according to any one of E.1 to E.4, wherein at least one gene comprises at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6.
[0295] Kit according to any one of E.1 - E.5, wherein at least one gene comprises at least four of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6.
[0296] Kit according to any one of E.1 - E.6, wherein at least one gene comprises at least one HPV E6 and / or E7.
[0297] Kit according to any one of E.1 - E.7, wherein at least one gene consists of at least four of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6 and at least one HPV E6 and / or E7.
[0298] Kit according to any one of E.1 - E.8, wherein the reagent comprises at least one primer and / or probe for any one of these genes, and at least one primer and / or probe is labeled with a detectable moiety.
[0299] Kit according to any one of E.1 - E.9, wherein a difference between the gene expression in an individual and a control value indicates that the individual may have HNSCC or is prone to developing HNSCC (i.e., at high risk).
[0300] F.1 Obtaining a sample from an individual; Measuring the amount of the expression product from a gene comprising at least one gene in Table 4 and / or Table 6 in the sample; Comparing the expression of at least one gene in Table 4 and / or Table 6 in the sample with a control value of the expression; When the difference between gene expression and a control value in an individual indicates that the individual may have HNSCC (i.e., a diagnosis of presence) or is likely to develop HNSCC (i.e., high risk), treating the individual for HNSCC and A method for treating HNSCC, including.
[0301] The method according to F.1, wherein the gene comprises at least one of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 or HSD17B6.
[0302] The method according to F.1 to F.2, wherein the gene comprises at least four of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 and HSD17B6.
[0303] The method according to any one of F.1 to F.3, further comprising measuring the amount of an expression product from at least one of the HPV E6 gene and / or the HPV E7 gene.
[0304] The method according to any one of F.1 to F.4, wherein the gene consists of at least four of CAB39L, ADAM12, SH3BGRL2, NRG2, COL13A1, GRIN2D, LOXL2, KRT4, EMP1 and HSD17B6, and an expression product from at least one of the HPV E6 and / or HPV E7 genes.
[0305] The method according to any one of F.1 to F.5, further comprising measuring the amount of a normalization gene such as KHDRBS1, or RPL30, or another normalization gene.
[0306] The method according to any one of F.1 to F.6, wherein the measuring step comprises measuring mRNA.
[0307] The measuring step is the method according to any one of F.1 to F.7, including an immunoassay.
[0308] The method according to any one of F.1 to F.8, including measuring the expression of at least four of the genes.
[0309] The method according to any one of F.1 to F.9, wherein the sample comprises serum, tissue, FFPE, saliva or plasma.
[0310] The method according to any one of F.1 to F.10, including comparing the expression level with a control value from a normal population.
[0311] References and citations to other documents, such as patents, patent applications, patent gazettes, magazines, books, papers, web content, etc., are made throughout this disclosure. All such documents are hereby incorporated by reference in their entirety for all purposes. Various modifications and equivalents of the things described herein will be apparent to those skilled in the art from the entire content of this document, including references to the scientific and patent literature cited herein. The subject matter of this specification includes information, examples, and guidance that can be adapted to the practice of this disclosure in its various embodiments and their equivalents.
Claims
[Claim 1] The invention described in the specification.
Citation Information
Patent Citations
Methods related to the prognosis of head and neck cancer
JP2015521480A
Method for distinguishing between head and neck squamous cell carcinoma and lung squamous cell carcinoma
US20070264644A1
Molecular signature for aggressive squamous cell carcinomas of the head and neck
US20150259751A1
Gene aberration(s) in squamous cell carcinoma of head and neck (HNSCC) and applications thereof
WO2016199107A1