Methods and compositions for the identification of tumor models
By employing deep NGS sequencing to analyze SNP loci, the method addresses the challenges of authenticating cell lines and tumor models, achieving high sensitivity and accuracy in identifying samples and detecting contaminants, thus improving the reliability of biobank samples.
Patent Information
- Application Number
- JP2021555073
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-03-12
- Filing Date
- 2020-03-12
- Publication Date
- 2025-05-19
- Estimated Expiration
- 2040-03-12
AI Technical Summary
Current methods for authenticating cell lines and tumor models, such as STR and SNP profiling, face challenges including contamination, genetic heterogeneity, and logistical burdens, especially in biobanks maintaining diverse in vitro and in vivo models.
A method involving deep next-generation sequencing (NGS) to detect genotypes at multiple human or mouse SNP loci, allowing for accurate identification and authentication of samples, including detection of contaminants and estimation of their ratios, with improved sensitivity and accuracy compared to existing methods.
The method achieves high sensitivity and accuracy in identifying and authenticating samples, with the ability to detect contaminants at levels as low as 1%, and effectively handles complex samples like human-mouse mixtures, thereby enhancing the reliability of biobank samples.
Smart Images

Figure 0007679306000034 
Figure 0007679306000035 
Figure 0007679306000036
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims priority based on Application PCT / CN2019 / 077750 filed on March 12, 2019, the disclosure of which is incorporated herein by reference.
[0002] The present invention generally relates to molecular biology, cancer biology, and animal models.
Background Art
[0003] Cell lines, organoids, xenografts, and allograft models are useful model systems in oncology and other biomedical research. Authentication and characterization of models assist in their proper use and mitigate a series of problems such as misidentification and misuse, cross - contamination, incorrect cancer classification, genomic changes due to long - term culture and genetic drift, etc., all of which have been particularly well - noted in cell lines due to their common use. For example, various studies have reported misidentification / contamination rates of approximately 10 - 40% for cell line banks.
[0004] There are various methods for authenticating cell lines, including cell morphology tests, isoenzymology, cytogenetic analysis (karyotyping and FISH), human lymphocyte antigen (HLA) typing, short tandem repeat (STR) profiling, single nucleotide polymorphism (SNP) typing, and DNA and RNA sequencing (Freedman, L.P. et al. Biotechniques 59, 189 - 90, 192 (2015); Nims, R.W. & Reid, Y. In Vitro Cell Dev Biol Anim 53, 880 - 887 (2017)). Among these techniques, STR profiling is the most widely used, and there are criteria (ASN - 0002) that guide its application when authenticating human cell lines (Almeida, J.L., Cole, K.D. & Plant, A.L. PLoS Biol 14, e1002476 (2016)). For mouse cell lines, a panel of 19 STR markers has also been developed (Zaaijer, S. et al. Elife 6 (2017)). The sensitivity of STR assays for detecting contaminants is approximately 5 - 10% (Yu, M. et al. Nature 520, 307 - 11 (2015)). In recent years, SNP typing has become increasingly used for cell line and biological sample authentication due to its improved accuracy, sensitivity, and reduced cost. SNPs can be profiled by PCR, as well as by next - generation sequencing (NGS) including transcriptome sequencing or RNA - seq, whole exome sequencing (WES), and whole genome sequencing (WGS). Current SNP assays have a detection sensitivity of approximately 3 - 5%. There are also databases that contain other information about STRs, SNPs, and cell lines to facilitate their authentication and characterization.
[0005] In addition to cell lines, organoids and mouse tumor models are widely used in oncology research and drug development. Organoids are in vitro three-dimensional cultures derived from stem cells, primary and engineered tumor samples, and xenografted human tumors that maintain many biological structures and functions. Mouse tumor models are in vivo systems that include patient-derived xenografts (PDX), cell line-derived xenografts (CDX), syngeneic or mouse cell line-derived models, mouse allograft models, etc. Some of these models, such as PDX, can more faithfully capture the histopathology and genomics of primary tumors than cell lines. Similar to cell lines, these tumor models have similar quality control issues, but additional problems exist. In xenograft models, tumors contain human tumor cells and mouse stromal cells, and the latter gradually replace the human counterparts during passage of the model, which, when combined with genomic heterogeneity, differences in implantation sites (subcutaneous and orthotopic), fluctuations in growth, and randomness of dissection, results in the human-mouse genetic composition of tumors varying considerably, such that even samples from the same PDX can range from almost pure human content to mouse content to some extent. Such tumor-host mixing and interference occur in all transplanted tumor models, causing fluctuations in allele frequencies for STR markers and SNPs and thus adversely affecting conventional STR- and SNP-based authentication methods. Large-scale sample authentication is also logistically burdensome and error-prone, especially for biobanks where many types of in vitro and in vivo models are maintained and used simultaneously. Therefore, there is a need to develop new SNP-based assays for identifying and authenticating tumor models.
Summary of the Invention
[0006] In one aspect, the present disclosure provides a method for identifying or authenticating a sample. In one embodiment, the method includes obtaining nucleic acid from the sample; detecting a genotype for the sample at a plurality of human single nucleotide polymorphism (SNP) loci or a plurality of mouse SNP loci; comparing the genotype of the sample to a reference genotype; and determining the identity of the sample. In certain embodiments, the human SNPs are selected from the group shown in Table 1. In certain embodiments, the mouse SNPs are selected from the group shown in Table 2.
[0007] In certain embodiments, the sample is a cell, tissue, organoid, or a combination thereof. In certain embodiments, the sample is a cell line or tumor tissue. In certain embodiments, the sample is derived from a xenograft or allograft tumor model. In certain embodiments, the sample is derived from a patient-derived xenograft (PDX), cell line-derived xenograft (CDX), syngeneic or mouse cell line-derived model, mouse allograft model.
[0008] In certain embodiments, the sample contains contaminants, and the method further includes determining the percentage of contaminants in the sample. In certain embodiments, the method further includes determining the identity of the contaminants.
[0009] In certain embodiments, the detecting step uses next-generation sequencing (NGS) or a sequencing-based SNP array. In certain embodiments, the nucleic acid is barcoded.
[0010] In certain embodiments, the method further includes identifying the gender of the subject from whom the sample was obtained. In certain embodiments, the method further includes identifying the ethnicity of the subject from whom the sample was obtained. In certain embodiments, the method further includes detecting the presence of a virus or mycoplasma in the sample. In certain embodiments, the method further includes determining the strain of immunodeficient mouse from which the sample was obtained.
[0011] In another aspect, the present disclosure provides a method for authenticating a sample comprising human and mouse components. In certain embodiments, the method comprises obtaining nucleic acid from the sample; detecting the genotype of the sample at 100 or more mouse genomic loci, wherein each of the mouse genomic loci has a corresponding homologous human genomic locus, and each of the mouse genomic loci and the corresponding homologous human genomic locus have the same flanking nucleotide sequence; and determining the ratio of mouse components in the sample based on the genotype. In certain embodiments, the mouse genomic loci are selected from Table 6.
[0012] In another aspect, the present disclosure provides a kit for identifying a sample. In certain embodiments, the kit comprises primers for detecting a set of human SNP loci or a set of mouse SNP loci in the sample. In certain embodiments, the kit further comprises an agent for amplifying a DNA fragment containing a human or mouse SNP using these primers.
[0013] In another aspect, the present disclosure provides a microarray for identifying a human or mouse sample. In certain embodiments, the microarray comprises probes for detecting the genotype of the sample at a set of human or mouse SNP loci.
[0014] In yet another aspect, the present disclosure provides a non-transitory computer-readable medium having instructions stored thereon, which when executed by a processor, cause the processor to search for the genotype of a sample at a set of human or mouse SNP loci; compare the genotype of the sample to a reference genotype; and determine the identification of the sample.
[0015] In yet another aspect, the present disclosure provides a method for authenticating a sample comprising a major component and a minor component. In certain embodiments, the method comprises detecting the genotype of the sample at 100 or more SNP loci; determining an SNP heterogeneity ratio for each of the SNP loci according to the formula shown in Table 11; determining a sample heterogeneity ratio based on the SNP heterogeneity ratios for the SNP loci using a mixture Gaussian distribution that models the genotype; and comparing the genotype of the sample to a set of reference genotypes each detected in a reference sample, identifying a reference sample having a reference genotype with the highest identity to the genotype of the sample, and determining that the major component of the sample is the reference sample when (i) the reference genotype is more than 90% identical to the genotype of the sample and the sample heterogeneity ratio is less than 10%, or (ii) the reference genotype is more than 80% identical to the genotype of the sample and the sample heterogeneity ratio is more than 10%, thereby determining the major component of the sample.
[0016] In certain embodiments, the method further comprises determining the minor component of the sample. In certain embodiments, the method further comprises determining the percentages of the major and minor components in the sample.
[0017] The following drawings form a part of this specification and are included to further illustrate certain aspects of the present disclosure. The present disclosure may be better understood by referring to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
Brief Description of the Drawings
[0018]
Figure 1A
Figure 1B
Figure 1C
Figure 2
Figure 3
Figure 4
DETAILED DESCRIPTION OF THE INVENTION
[0019] Before explaining the present disclosure in more detail, it should be understood that the present disclosure is not limited to the specific embodiments described and can, of course, change itself. It should also be understood that the technical terms used herein are for the purpose of describing specific embodiments only and do not mean limitation, because the scope of the present disclosure is limited only by the appended claims.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Any methods and materials similar to or equivalent to those described herein can be used in the practice or testing of the present disclosure, but the preferred methods and materials are described herein.
[0021] All publications and patents cited in this specification are hereby incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference, and incorporated herein by reference to disclose and describe methods and / or materials in connection with the cited publications. Any citation of a publication is for its disclosure prior to the filing date of this application, and this disclosure should not be construed as an admission that this disclosure has no right to antedate such a publication by virtue of prior disclosure. Further, the dates of the provided publications may differ from the actual publication dates which may need to be independently confirmed.
[0022] As will be apparent to those skilled in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein can be readily separated from or combined with any of the features of some other embodiments without departing from the scope or spirit of the disclosure. Any recited method can be carried out in the recited order of events or in any other order that is logically possible.
[0023] Definitions The following definitions are provided to assist the reader. Unless otherwise specified, all technical terms, notations and other scientific, medical or technical terms used in this specification shall have the meanings commonly understood by those of ordinary skill in the chemical and medical arts. In some cases, terms with commonly understood meanings are defined in this specification for clarity and / or ease of reference, and the inclusion of such definitions in this specification should not necessarily be construed as representing a substantial difference in the definition of terms as commonly understood in the art.
[0024] As used in this specification, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise.
[0025] The term "allele" means one of two or more existing genetic variants at a particular polymorphic locus.
[0026] As used herein, the term "amount" or "level" means the amount of the polynucleotide or polypeptide of interest present in the sample. Such amount may be expressed absolutely, i.e., as the total amount of polynucleotide or polypeptide in the sample, or relatively, i.e., as the concentration of polynucleotide or polypeptide in the sample.
[0027] As used herein, the term "cancer" or "tumor" means any disease involving abnormal cell growth and includes all stages and forms of disease affecting any tissue, organ or cell in the body. This term includes all known cancers and neoplastic conditions characterized as malignant, benign, soft tissue or solid, as well as all stages and grades of cancer including pre-metastatic and metastatic cancers. Generally, cancers can be classified by the tissue or organ in which the cancer is located or has originated, as well as by the morphology of the cancerous tissue and cells. As used herein, types of cancer include, but are not limited to, acute lymphoblastic leukemia (ALL), acute myeloid leukemia, adrenocortical carcinoma, anal cancer, astrocytoma, childhood cerebellar or cerebral basal cell carcinoma, bile duct cancer, bladder cancer, bone tumor, brain cancer, cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, chordoma, medulloblastoma, supratentorial primitive neuroectodermal tumor, visual pathway and hypothalamic glioma, breast cancer, Burkitt lymphoma, cervical cancer, chronic lymphocytic leukemia, chronic myeloid leukemia, colon cancer, emphysema, endometrial cancer, chordoma, esophageal cancer, Ewing sarcoma, retinoblastoma, gastric (stomach) cancer, glioma, head and neck cancer, heart cancer, Hodgkin lymphoma, islet cell carcinoma (pancreatic endocrine), Kaposi sarcoma, kidney cancer (renal cell carcinoma), laryngeal cancer, leukemia, liver cancer, lung cancer, neuroblastoma, non-Hodgkin lymphoma, ovarian cancer, pancreatic cancer, laryngeal cancer, prostate cancer, rectal cancer, renal cell carcinoma (kidney cancer), retinoblastoma, Ewing family of tumors, skin cancer, stomach cancer, testicular cancer, laryngeal cancer, thyroid cancer, vaginal cancer.
[0028] As used herein, the "cell" may be a prokaryotic cell or a eukaryotic cell. Prokaryotic cells include, for example, bacteria. Eukaryotic cells include, for example, fungi, plant cells, and animal cells.Types of animal cells (e.g., mammalian cells or human cells) include, for example, cells of the circulatory / immune system or organs (e.g., B cells, T cells (cytotoxic T cells, natural killer T cells, regulatory T cells, helper T cells), natural killer cells, granulocytes (e.g., basophilic granulocytes, eosinophilic granulocytes, neutrophilic granulocytes and hypersegmented neutrophils), monocytes or macrophages, erythrocyte cells (e.g., reticulocytes), mast cells, platelets or megakaryocytes, and dendritic cells), cells of the endocrine system or organs (e.g., thyroid cells (e.g., thyroid epithelial cells, parafollicular cells), parathyroid cells (e.g., parathyroid chief cells, oxyphil cells), adrenal cells (e.g., chromaffin cells) and pineal gland cells (e.g., pinealocytes)), cells of the nervous system or organs (e.g., glioblastoma cells (e.g., astrocytes and oligodendrocytes), microglia, giant neurosecretory cells, stellate cells, Betz cells and pituitary cells (e.g., gonadotropin-secreting cells, adrenocorticotropic hormone-secreting cells, thyroid-stimulating hormone-secreting cells, growth hormone-secreting cells and mammotropic hormone-secreting cells)), cells of the respiratory system or organs (e.g., alveolar cells (type I alveolar cells and type II alveolar cells), Clara cells, goblet cells, alveolar macrophages), cells of the circulatory system or organs (e.g., cardiomyocytes and peripheral cells), cells of the digestive system or organs (e.g., gastric chief cells, parietal cells, goblet cells, Paneth cells, G cells, D cells, ECL cells, I cells, K cells, S cells, enteroendocrine cells, enterochromaffin cells, APUD cells, liver cells (e.g., hepatocytes and Kupffer cells)), cells of the integumentary system or organs (e.g., bone cells (e.g., osteoblasts, osteocytes and osteoclasts), dental cells (e.g., cementoblasts and ameloblasts), cartilage cells (e.g., chondroblasts and chondrocytes), skin / hair cells (e.g., hair follicles, keratinocytes and melanocytes (nevus cells)), muscle cells (e.g., myocytes), adipocytes, fibroblasts and tendon cells), cells of the urinary system or organs (e.g., podocytes, juxtaglomerular cells, mesangial cells within the glomerulus, mesangial cells outside the glomerulus, kidney proximal tubule brush border cells and macula densa cells) and cells of the genital system or organs (e.g., sperm, Sertoli cells, Leydig cells, ova, oocytes).The cell may be a normal and healthy cell or a diseased and unhealthy cell (e.g., a cancer cell). The cell further includes stem cells including mammalian zygotes or embryonic stem cells, fetal stem cells, induced pluripotent stem cells, and adult stem cells. A stem cell is a cell that can differentiate into specialized cell types through the cell division cycle while maintaining an undifferentiated state. Stem cells may be totipotent stem cells, pluripotent stem cells, multipotent stem cells, oligopotent stem cells, and unipotent stem cells, and any of them can be induced from somatic cells. Stem cells can also include cancer stem cells. Mammalian cells may be rodent cells, e.g., mouse, rat, hamster cells. Mammalian cells may be lagomorph cells, e.g., rabbit cells. Mammalian cells may also be primate cells, e.g., human cells. In a specific example, the cell is a cell used in mammalian cell production, e.g., a CHO cell.
[0029] The term "complementary" means the ability of a nucleic acid to form hydrogen bonds with another nucleic acid sequence, either by traditional Watson-Crick or other non-traditional methods. Percent complementarity refers to the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairs) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Fully complementary" means that all of the consecutive residues of a nucleic acid sequence hydrogen bond with the same number of consecutive residues in a second nucleic acid sequence. As used herein, "substantially complementary" means a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or means two nucleic acids that hybridize under stringent conditions.
[0030] In the present disclosure, terms such as "comprises", "comprised", "comprising", "contains", "containing", etc. have the meaning in the United States Patent Law, and it should be noted that these are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. Terms such as "consisting essentially of" and "consists essentially of" have the meaning in the United States Patent Law, and these permit the inclusion of additional components or steps that do not significantly affect the basic and novel characteristics of the claimed invention. The terms "consists of" and "consisting of" have the meaning considered to be in the United States Patent Law, that is, these terms are closed-ended.
[0031] As used herein, the term "contaminant" means a component present in a sample that is different from the major constituent components of the sample or an impurity or other undesirable influence in the sample, such as causing spoiling, contamination, or infection.
[0032] The terms "determine", "evaluate", "assay", "measure", and "detect" can be used synonymously and mean both quantitative and semi-quantitative determinations. When meaning either a quantitative or semi-quantitative determination, the phrases "determine the level of" or "detect" the polynucleotide or polypeptide of interest can be used.
[0033] The term "genome" means the total genetic information carried by an individual organism or cell, as indicated by the complete DNA sequence of its chromosomes.
[0034] The term "hybridizing" means the preferential binding, duplex formation or hybridization of a nucleic acid molecule to a specific nucleotide sequence under stringent conditions. The term "stringent conditions" means conditions under which a probe preferentially hybridizes to its target subsequence in a mixed population (e.g., a cell lysate or DNA preparation from a tissue biopsy) and hybridizes to other sequences to a relatively lesser extent or not at all. "Stringent hybridization" and "stringent hybridization wash conditions" (such as in array, microarray, Southern or Northern hybridization) in the context of nucleic acid hybridization are sequence-dependent and vary under different environmental parameters. General guidelines for nucleic acid hybridization can be found, for example, in Tijssen Laboratory Techniques in Biochemistry and Molecular Biology - Hybridization with Nucleic Acid Probes part I, Ch.2, "Overview of principles of Hybridization and the strategy of Nucleic acid probe assays," (1993) Elsevier, N.Y. Generally, high stringency hybridization and wash conditions are selected to be about 5°C lower than the melting temperature (Tm) of a particular sequence at a defined ionic strength and pH. Tm is the temperature at which 50% of the target sequence hybridizes to a perfectly matched probe (at a defined ionic strength and pH). Very stringent conditions are selected to be equal to the Tm of a particular probe.An example of stringent hybridization conditions for the hybridization of a complementary nucleic acid having more than 100 complementary residues on an array or filter in a Southern or Northern blot is 42 °C using a standard hybridization solution (see, e.g., Sambrook and Russell Molecular Cloning: A Laboratory Manual (3rd ed.) Vol. 1-3 (2001) Cold Spring Harbor Laboratory, Cold Spring Harbor Press, NY). An example of high stringency wash conditions is about 15 minutes at 72 °C with 0.15 M NaCl. An example of stringent wash conditions is about 15 minutes at 65 °C with 0.2× SSC wash. Often, a low stringency wash is performed prior to the high stringency wash to remove background probe signal. For example, an example of moderately stringent wash for a double strand greater than 100 nucleotides is 15 minutes at 45 °C with 1× SSC. For example, an example of low stringency wash for a double strand greater than 100 nucleotides is 15 minutes at 40 °C with 4× SSC to 6× SSC.
[0035] The term "locus" means any segment of a DNA sequence in the genome, defined by chromosomal coordinates in a reference genome known in the art, regardless of biological function. A DNA locus may contain multiple genes or no genes; a DNA locus can be a single base pair or millions of base pairs.
[0036] The terms "nucleic acid" and "polynucleotide" are used synonymously and refer to a polymeric form of nucleotides of either deoxyribonucleotides or ribonucleotides of any length, or analogs thereof. A polynucleotide can have any three-dimensional structure and can perform any known or unknown function. Non-limiting examples of polynucleotides include genes, gene fragments, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, shRNA, single-stranded short or long RNAs, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, regulatory regions, isolated RNA of any sequence, nucleic acid probes, and primers. Nucleic acid molecules can be linear or circular.
[0037] The term "oligonucleotide" refers to a nucleic acid sequence of at least about 5 nucleotides to about 500 nucleotides (e.g., 5, 6, 7, 8, 9, 10, 12, 15, 18, 20, 21, 22, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 100, 125, 150, 175, 200, 250, 300, 350, 400, 450, or 500 nucleotides). In some embodiments, for example, an oligonucleotide can be about 15 nucleotides to about 30 nucleotides, or about 20 nucleotides to about 25 nucleotides, and can be used, for example, as a primer in a polymerase chain reaction (PCR) amplification assay and / or as a probe in a hybridization assay or microarray. Oligonucleotides of the present invention can be natural or synthetic, such as, for example, DNA, RNA, PNA, LNA, modified backbones, etc., as is well known in the art.
[0038] The term "polymorphic locus" refers to a genomic locus at which two or more alleles have been identified.
[0039] As used herein, the term "reference genotype" means a pre-determined genotype of one or more genomic loci present in a reference sample, e.g., a sample with a known identity. The reference genotype functions as a basis for comparison of the genotype of a specific genomic locus present in a test sample and is suitable for use in the methods of the present invention. The reference genotype may vary depending on the nature of the sample and other factors such as the sex, age, ethnicity, etc. of the subject on which such reference sample was established.
[0040] As used herein, the term "sample" or "biological sample" means any cell, tissue, organoid, or any other sample containing one or more nucleic acid molecules of interest. In certain embodiments, the sample is a cell (e.g., normal cell, cancer cell, cell line), tissue (e.g., normal tissue, cancer tissue, xenograft or allograft tissue), organoid, etc.
[0041] The term "single nucleotide polymorphism" or "SNP" means a single nucleotide position in a genomic sequence where two or more alternative alleles are present at a frequency detectable in a population, e.g., >1%. SNPs can be present within the coding sequence of a gene, within the non-coding region of a gene, and / or within the intergenic (e.g., intron) region of a gene. SNPs that are not necessarily present within the protein-coding region can still have an impact on gene splicing, transcription factor binding, and / or the sequence of non-coding RNAs. The SNP nomenclature provided herein refers to the official Reference SNP (rs) identification number assigned to each unique SNP by the National Center for Biotechnological Information (NCBI) and available in the GenBank® database.
[0042] As used herein, the term "subject" means a human or any non-human animal (e.g., mouse, rat, rabbit, dog, cat, cow, pig, sheep, horse or primate). Humans include pre-natal and post-natal forms. In many embodiments, the subject is a human. The subject may be a patient, which means a person who has come to a medical institution for the diagnosis or treatment of a disease. The term "subject" is used interchangeably herein with "individual" or "patient". The subject may or may not be suffering from, or be susceptible to, a disease or disorder, and may or may not exhibit symptoms of a disease or disorder.
[0043] The term "substrate", when used in the context of an array, means a material that can support the relevant assay components (e.g., assay regions, cells, test compounds, etc.). Examples of substrates include, but are not limited to, glass, Si-based materials, functionalized polystyrene, functionalized polyethylene glycol, functionalized organic polymers, nitrocellulose or nylon membranes, paper, cotton, and materials suitable for synthesis. The substrate need not be flat and includes any type of shape including spherical shapes (e.g., beads). The material bound to the substrate can be bound to any part of the substrate (e.g., the inner part of a porous substrate material). Preferred embodiments of the technology of the present invention have nucleic acid probes bound to the substrate. A nucleic acid probe is "bound" to the substrate when it is associated with the substrate via non-random chemical or physical interactions. In some preferred embodiments, the binding is via a covalent bond, such as provided by a linker, for example.
[0044] The term "tumor model", as used herein, means a cell, tissue or animal used for studying the development and progression of cancer and for testing treatments before being administered to humans.
[0045] The term "tumor sample" includes biological samples or samples from biological sources containing one or more tumor cells. Biological samples include body fluids, such as samples of blood, plasma, serum or urine, or samples obtained, for example, by biopsy, from cells, tissues or organs, preferably tumor tissue suspected of containing or consisting essentially of cancer cells.
[0046] SNP for identification of tumor samples Misidentification and contamination of biobank samples (e.g., cell lines) have plagued biomedical research. Short tandem repeat (STR) and single nucleotide polymorphism (SNP) assays are widely used to authenticate biological samples and can detect contamination with sensitivities of 5 - 10% and 3 - 5%, respectively. The present disclosure, in one aspect, provides a method having a sensitivity of ≤1% for detecting contamination. This method can further identify contaminants and estimate the contamination ratio for mixed cell line samples. This method is the most sensitive and accurate method ever reported for cell line authentication. In certain embodiments, this method can also detect interspecies contamination in human - mouse mixed samples, such as xenograft tumors, and accurately estimate the mouse ratio. In certain embodiments, mycoplasma and mollicutes are also among the research targets. In certain embodiments, this multifunctional method simultaneously infers the population structure and gender of human samples. In certain embodiments, thanks to DNA barcoding technology, the method disclosed herein can profile 100 - 200 samples in a single run at a cost per sample comparable to that of conventional STR assays, thereby becoming a truly high - throughput and low - cost tool for maintaining high - quality biobanks.
[0047] The methods and compositions described herein are based in part on the discovery of a set of SNP loci that can be used to identify and authenticate samples obtained from tumor models. In certain embodiments, the tumor model is a human tumor model including primary human tumors, patient-derived xenografts (PDX), human tumor cell lines, human cell line-derived xenografts, and human organoids. In certain embodiments, the SNPs are selected from human SNPs based on RNAseq or whole exome sequencing (WES) data of several human tumor models. The selected human SNPs are located in exon regions of highly expressed genes that are mainly in unlinked disequilibrium (non-LD) blocks across 22 autosomes. Thus, each human tumor model has a unique genotype (i.e., SNP fingerprint) at the selected human SNP loci.
[0048] In certain embodiments, the selected human SNP loci have homology in the mouse genome. When a sample is amplified using primers targeting such human SNP loci, if the sample is mixed with mouse cells or tissues, the nucleotide sequence of the corresponding mouse locus can be generated. Such human SNPs can be used, for example, to estimate the percentage of mouse content in a mixture of human and mouse cells / tissues based on the number of mouse and human reads of these SNPs.
[0049] In certain embodiments, the human SNPs used herein are selected from the group shown in Table 1.
[0050] In certain embodiments, the SNPs include a set of mouse SNPs for identifying and authenticating mouse tumor models such as mouse tumor cell lines. In some embodiments, the mouse SNPs used herein are selected from the group shown in Table 2.
[0051] In certain embodiments, the SNPs further include human SNPs in the sex chromosomes (X chromosome and Y chromosome) for determining the gender of the subject from whom the sample was obtained. In certain embodiments, the sex chromosome SNPs are selected from the group shown in Table 3.
[0052] In certain embodiments, the SNPs further include mouse SNPs that can be used to determine the strain of an immunodeficient mouse from which the sample was obtained. In some embodiments, the SNPs are shown in Table 4.
[0053] Method In one aspect, the present disclosure provides a method for identifying and authenticating a sample.
[0054] In certain embodiments, the methods disclosed herein are for matching a sample to a reference (e.g., a standard cancer cell line). Conventional STR and SNP assays have mainly used genotype-based Tanabe-Masters algorithms and variations thereof. STR assays produce similar signals for many markers. SNP assays often genotype even more SNPs. Thus, the similarity thresholds used by SNP assays to call two samples a match are often higher. However, the matching power of conventional assays can be severely impaired for contaminated samples, even when using about 100 SNPs. In certain embodiments, the methods disclosed herein performed high-depth (3000×) sequencing of 237 SNP sites for human samples and showed 100% accuracy in identifying the sample or the major components of a contaminated sample.
[0055] In certain embodiments, the methods disclosed herein are for detecting contaminants in biological samples. The sensitivity for detecting contaminants in cell lines is about 5-10% for STR assays and 3-5% for SNP assays. However, the performance can be rather unstable, to the extent that contaminants of >20% were not detected even by the 96-SNP assay in a mixture of two unrelated cell lines (Liang-Chu, M.M. et al. PLoS One 10, e0116218 (2015)). In certain embodiments, the methods disclosed herein consistently reach a sensitivity of 2% both by the value and the distinguishable bi / multimodal distribution when using only the heterogeneity ratio. The sensitivity reaches ≦1% when there are contaminants present in the library of reference samples with SNP fingerprints. Since cell lines without contaminants exhibit a level of genetic heterogeneity comparable to that of cell line samples with about 1% contamination due to polyclonality and sequencing errors, such sensitivity is effectively the theoretical detection limit.
[0056] In certain embodiments, the methods disclosed herein are for identifying contaminants. Cross-contamination of cell lines is common in biobanks. The composition of contaminated cultures changes over time due to the different growth rates of the cell lines. Cell lines have different genomics, such as gene mutations, and respond differently to drug treatments, which can cause false results in drug screening. The inventors of the present disclosure constructed a SNP fingerprint library for over 1000 cancer cell lines, whereby contaminated cell lines can be unambiguously identified. Furthermore, the contamination ratio can be accurately estimated. In addition to checking the quality of cell lines, this ability can have other uses, such as monitoring the dynamic composition of two cell lines under biological or chemical interference.
[0057] In addition to intraspecies contamination, in certain embodiments, the methods disclosed herein can accurately detect and quantify interspecies contamination between humans and mice. In certain embodiments, the methods disclosed herein use 108 homologous DNA segments that differ between two species but have identical flanking nucleotide sequences, rather than SNPs, and thus common primers can be designed for unbiased amplification of human and mouse DNA segments. This approach demonstrated perfect performance in a series of mouse-human DNA mixture benchmark samples. The homology-based principle can be used to detect other interspecies contaminations.
[0058] In certain embodiments, the power of the methods disclosed herein comes from several novel features. The first is deep NGS sequencing, which obtains both SNP genotypes and nucleotide frequencies, whereas conventional STR and SNP assays only profile SNP genotypes. Second, separate from SNP profiling, the methods disclosed herein perform targeted sequencing to detect mycoplasma contamination and estimate the mouse-human mixing ratio. Third, a suite of statistical models and algorithms have been developed to utilize deep NGS sequencing data, making the authentication method automated, robust, and objective. Finally, DNA barcode technology is used to enable simultaneous parallel sequencing of 100-200 samples, greatly reducing costs.
[0059] The high-throughput and low-cost methods disclosed herein can be routinely used by biobanks to maintain authenticated high-quality samples. This method can be widely adapted to samples from other species and even microbiomes and can be implemented on any NGS sequencing platform.
[0060] In one embodiment, the method includes obtaining nucleic acid from a sample; detecting a genotype for the sample at a plurality of single SNP loci of humans or mice disclosed herein; comparing the genotype for the sample to a reference genotype detected in a reference sample; and determining an identification of the sample. In certain embodiments, genotypes at 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 or more SNP loci are detected.
[0061] The nucleic acid obtained from the sample can be RNA or DNA. In certain embodiments, the nucleic acid obtained from the sample is genomic DNA isolated from the sample. In certain embodiments, the nucleic acid obtained from the sample is genomic DNA and total RNA or mRNA isolated from the sample. In certain embodiments, the nucleic acid obtained from the sample is amplified, for example, by a PCR reaction or PCR after reverse transcription.
[0062] The genotype for the sample at the SNP locus can be detected based on any suitable method known in the art, such as, but not limited to, sequencing-based methods and hybridization-based methods.
[0063] In certain embodiments, the detecting step includes an amplifying step. In such cases, the detection agent includes at least a pair of primers that can hybridize to a genomic region containing the SNP locus and amplify a polynucleotide sequence surrounding the SNP locus in the presence of polymerase. The primer pair used to amplify the genomic region containing the SNP has sufficient identity or complementarity to at least a portion of the genomic region such that the primer or probe can specifically hybridize to the genomic region or its complementary strand. "Specifically hybridize" as used herein means that the primer or probe can hybridize to the intended sequence under stringent conditions. "Stringent conditions" as used herein means hybridizing at 42 °C in a solution consisting of 5×SSPE, 5×Denhardt's solution, 0.5% SDS and 100 μg / mL denatured salmon sperm DNA, and then washing at 42 °C in a solution containing 0.5×SSC and 0.1% SDS.
[0064] After amplification by a suitable nucleic acid amplification method such as PCR, the sequence or SNP in the amplification product is detected. In certain embodiments, the amplification product has a length of 50 bp to 500 bp. In certain embodiments, the sequence of the SNP in the amplification product is detected using a sequencing-based method, such as a next-generation sequencing (NGS) method. In certain embodiments, the NGS method is used to determine the sequences at a number of SNP loci. In certain embodiments, the NGS method can be used to simultaneously determine the sequences of SNP loci from multiple samples by barcoding the nucleic acids obtained from each sample.
[0065] If the nucleic acid obtained from the sample is RNA, the amplifying step may optionally include a reverse transcription step to produce cDNA of the RNA in the sample. The cDNA is then amplified using primers to enable detection of the presence of the SNP.
[0066] In some embodiments, for example, a microarray is used to detect SNPs in nucleic acids. The microarray consists of a reproducible pattern of capture probes bound to a solid support. Labeled RNA or DNA hybridizes to complementary probes on the array and is then detected by laser scanning. The presence of an SNP can be detected by measuring the intensity of the labeled RNA or DNA that binds to specific probes on the array.
[0067] The synthesis techniques for these arrays using mechanical synthesis methods are described, for example, in U.S. Patent No. 5,384,261. Although flat array surfaces are often used, the arrays can actually be fabricated on surfaces of any shape or on multiple surfaces. The arrays can also be nucleic acids on beads, gels, polymer surfaces, fibers such as optical fibers, glass, or any other suitable substrate; see U.S. Patents Nos. 5,770,358, 5,789,162, 5,708,153, 6,040,193, and 5,800,992. The arrays may be packaged in a manner that enables diagnostic methods or other operations of the comprehensive device.
[0068] The probes and primers necessary to practice the present invention can be synthesized or labeled using well-known techniques. The oligonucleotides used as probes and primers can be chemically synthesized using an automated synthesizer by the solid-phase phosphoramidite triester method first described by Beaucage and Caruthers, Tetrahedron Letts. (1981) 22:1859-1862, as described in Needham-Van Devanter et al, Nucleic Acids Res. (1984) 12:6159-6168.
[0069] In certain embodiments, the method further includes identifying the gender of a subject from whom a sample was obtained by detecting, for example, a sex chromosome SNP selected from the group shown in Table 3. In certain embodiments, the method further includes identifying the ethnicity of a subject from whom a sample was obtained. In certain embodiments, the method further includes determining the strain of an immunodeficient mouse from whom a sample was obtained by detecting, for example, a SNP of a supplier shown in Table 4.
[0070] In certain embodiments, the methods disclosed herein further include detecting common viral infections and mycoplasma contaminants, including hepatitis A / B / C virus (HAV / HBV / HCV), human immunodeficiency virus (HIV), Epstein-Barr virus (EBV), and human papillomavirus (HPV) in tumor models. In certain embodiments, the markers used to detect viral infections and mycoplasma contaminants are shown in Table 5.
[0071] In certain embodiments, the methods disclosed herein can be used to authenticate a sample containing major and minor components. In certain embodiments, the method includes estimating a heterogeneity ratio; determining the major components of the sample; determining the minor components of the sample; and estimating the mixing ratio of the major and minor components.
[0072] In certain embodiments, the heterogeneity ratio can be estimated as follows. There are six informative genotype combinations (Table 11) that can be used to estimate the heterogeneity ratio from deep NGS sequencing data. These exhibit four distinguishable nucleotide frequency patterns. Combinations 1 and 2 yield the same pattern, and the inventors use an averaging formula to calculate the percentage of minor constituent S2, i.e., the heterogeneity ratio. This formula yields an accurate estimate of the ratio when two combinations occur with equal frequency, which should approximate in scenarios where the number of SNPs is large. A similar averaging approach is used for combinations 4 and 5. When the heterogeneity ratio is low, sequencing errors can interfere with the inference of the heterogeneity ratio. To mitigate this, a two-step statistical approach can be used. Assuming a sequencing error of e = 0.001 and a sequencing depth of n (n ≧ 500, SNPs with n < 500 are discarded) at a given SNP site, the probability of observing k incorrect nucleotides follows a binomial distribution with parameters n and e.
[0073]
Number
[0074] For each n, a cumulative density function can be calculated to obtain a threshold h such that the probability of observing more than h incorrect nucleotides out of n nucleotides is less than 0.01. In the sequencing data, any low-frequency nucleotides having a number of reads smaller than the corresponding threshold h are discarded. Then, the expectation-maximization algorithm (R package mclust, version 3.5.3) is used to estimate the parameters of a mixture of Gaussians (having 1 to 3 components) that models the distribution of nucleotide frequencies smaller than the maximum heterogeneity (0.2 was used for all samples in this study). If there is only a single Gaussian component, or only a Gaussian component having the smallest mean that occupies more than 60% of all data points, the median of all data points is taken as the sample heterogeneity ratio; otherwise, the median of the data points in the other Gaussian component is taken as the sample heterogeneity ratio.
[0075] To determine the major constituent in a sample, the genotype at the SNP site is determined using only nucleotides having an allele frequency greater than a threshold that is 10% for the reference sample and 25% for the test sample that may be contaminated. The genotype similarity between the reference sample and the test sample is the percentage of SNPs having the same genotype, excluding SNPs having a sequencing depth of less than 500 in the test sample. The major constituent of the test sample is the reference sample having the highest genotype similarity, and this similarity must be higher than 90% (or 80%) when the heterogeneity ratio of the test sample is <10% (or >10%). Otherwise, the major constituent is not called.
[0076] After estimating the heterogeneity ratio and determining the major components, minor components of the test sample can be determined. For a mixture of a major component and one of the other reference samples (e.g., all cell lines with genomic data), a chimeric genotype having perhaps 1 to 4 nucleotides can be obtained at all SNP sites. The nucleotide frequencies are calculated using the heterogeneity ratio. Similarly, the chimeric genotype of the test sample is obtained. Two chimeric genotypes are considered identical if they carry the same nucleotides and the frequency of each nucleotide is within 3-fold. Then, the genotype similarity between the test sample and each reference sample combined with the major component is calculated. Next, a beta distribution with parameters (α,β) is fitted to the set of all pairwise genotype similarities.
[0077]
Number
[0078] In the equation, Γ(α) is the gamma function and x is the genotype similarity. The parameters are estimated by the R package fitdistrplus (version 3.5.3). From the fitted beta distribution, the probability of observing any genotype similarity greater than a particular value is calculated. A quantile-quantile graph with 99% confidence bands is plotted for all observed genotype similarities for visualization. A reference sample is considered a minor component if (1) it has the highest genotype similarity, (2) its genotype similarity exceeds the 99% confidence upper limit in the quantile-quantile graph, and (3) its p-value in the fitted beta distribution < 1.0E-6.
[0079] The mixing ratios for two reference samples can be estimated as follows. Assume that two components S1 and S2 are mixed at a ratio of Θ for S1 and (1 - Θ) for S2, where 0 ≦ Θ ≦ 1. From the deep NGS sequencing data, the nucleotide frequencies of all n SNPs in both components can be accurately estimated. For an SNP, the four nucleotide frequencies are shown as {A 1 , T 1 , G 1 , C 1} for component S1 and {A 2 , T 2 , G 2 , C 2} for component S2, and these sum to 1. In principle, one of the frequencies is close to 1 if the SNP is homozygous, and both frequencies are close to 0.5 if the SNP is heterozygous. The actual data may have some deviation due to sequencing errors and randomness, as well as the polyclonality of the cell line.
[0080] From the sequencing data of the mixed sample, the actual occurrence rates of the four nucleotides are shown as x = {n A , n T , n G , n C}. The likelihood of such an observation is
[0081]
Number
[0082] The likelihood P Θ (x i ) can be calculated for any SNPi ∈ (1, 2,..., n) using the observed data x i . The likelihood of observing the data X = {x 1 , x 2 ,..., x n} for all SNPs is
[0083]
Number
[0084] Therefore, the log-likelihood is
[0085]
Number
[0086] Θ that maximizes the likelihood can be solved by the stepwise increase of Θ.
[0087] Kit and Microarray In another aspect, the present disclosure provides a kit for use in the foregoing method. This kit can include any or all of the reagents for carrying out the methods described herein. In certain embodiments, this kit includes primers for detecting a group of human SNP loci or a group of mouse SNP loci in a sample. In certain embodiments, this kit further includes primers for detecting sex chromosome SNPs to identify the sex of the subject from whom the sample was obtained. In certain embodiments, this kit further includes primers for detecting ethnicity SNPs to identify the ethnicity of the subject from whom the sample was obtained. In certain embodiments, this kit further includes primers for detecting vendor SNPs to determine the strain of immunodeficient mice from which the sample was obtained. In certain embodiments, this kit further includes primers for detecting viral infection or mycoplasma contamination in a sample.
[0088] In certain embodiments, the kit further comprises an agent for amplifying DNA fragments containing human or mouse SNPs using these primers. Additionally, the kit may include instructional material containing instructions (i.e., protocols) for the practice of the methods provided herein. Instructional material typically includes, but is not limited to, written or printed material. Any medium capable of storing such instructions and with which an end user can communicate such instructions is contemplated in the present invention. Such media include, but are not limited to, electronic storage media (e.g., magnetic disks, tapes, cartridges, chips), optical media (e.g., CDROM), and the like. Such media can include an address to an Internet site that provides such instructional material.
[0089] In another aspect, the present disclosure provides oligonucleotide probes bound to a solid support, such as an array slide or chip, as described, for example, in Eds., Bowtell and Sambrook DNA Microarrays: A Molecular Cloning Manual (2003) Cold Spring Harbor Laboratory Press. The structure of such devices is well known in the art as described, for example, in U.S. Patents and Patent Publications, U.S. Patent No. 5,837,832, PCT Application No. WO95 / 11995, U.S. Patent No. 5,807,522; U.S. Patent Nos. 7,157,229, 7,083,975, 6,444,175, 6,375,903, 6,315,958, 6,295,153 and 5,143,854, 2007 / 0037274, 2007 / 0140906, 2004 / 0126757, 2004 / 0110212, 2004 / 0110211, 2003 / 0143550, 2003 / 0003032 and 2002 / 0041420. Nucleic acid arrays are also summarized in the following references: Biotechnol Annu Rev (2002) 8:85-101; Sosnowski et al. Psychiatr Genet (2002) 12(4):181-92; Heller, Annu Rev Biomed Eng (2002) 4:129-53; Kolchinsky et al., Hum. Mutat (2002) 19(4):343-60; and McGail et al., Adv Biochem Eng Biotechnol (2002) 77:21-42.
[0090] A microarray can be composed of a number of unique single-stranded polynucleotides, each of which is typically either a synthetic antisense polynucleotide or a fragment of cDNA, immobilized on a solid support. Typical polynucleotides are preferably about 6 to 60 nucleotides in length, more preferably about 15 to 30 nucleotides in length, and most preferably about 18 to 25 nucleotides in length. For certain types of arrays or other detection kits / systems, it may be preferable to use oligonucleotides that are only about 7 to 20 nucleotides in length. In other types of arrays, such as those used in conjunction with chemiluminescence detection technology, the preferred probe length may be, for example, about 15 to 80 nucleotides in length, preferably about 50 to 70 nucleotides in length, more preferably about 55 to 65 nucleotides in length, and most preferably about 60 nucleotides in length.
[0091] Methods, systems, and apparatuses implemented by a computer Any of the methods described herein can be implemented in a computer system that includes one or more processors configured to perform the steps, either wholly or in part. Accordingly, embodiments are directed to a computer system configured to perform any of the steps of the methods described herein, which may have different components for performing each step or each group of steps. Although presented as numbered steps, the steps of the methods herein can be performed simultaneously or in a different order. Additionally, some of these steps can be used in conjunction with some of the steps of other methods. Also, all or some of the steps may be optional. Any step of any method can be implemented by modules, circuits, or other means for performing these steps.
[0092] Any of the computer systems referred to in this specification can utilize any suitable number of subsystems. In some embodiments, a computer system includes a single computer device where the subsystems can be components of the computer device. In other embodiments, a computer system can include a number of computer devices, each having internal components and each being a subsystem. The subsystems can be interconnected via a system bus. Additional subsystems can include, for example, a printer, a keyboard, a storage device, a monitor coupled to a display adapter, and others. Peripheral devices and input / output (I / O) devices coupled to an I / O regulator can be connected to the computer system by any number of means known in the art, such as a serial port. For example, a serial port or an external interface (e.g., Ethernet®, Wi-Fi, etc.) can be used to connect the computer system to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus enables the central processing unit to communicate with each subsystem and control the execution of instructions from the system memory or a storage device (e.g., a fixed disk such as a hard drive or an optical disk), and enables information exchange between subsystems. The system memory and / or the storage device can incorporate a computer-readable medium. Any of the data referred to in this specification can be an output from one component to another or an output to a user.
[0093] A computer system can include, for example, a plurality of the same components or subsystems connected together by an external interface or an internal interface. In some embodiments, a computer system, subsystem, or device can communicate with a network. In such cases, one computer can be regarded as a client and another as a server, and each can be part of the same computer system. A client and a server can each include a number of systems, subsystems, or components.
[0094] It should be understood that any embodiment of the present disclosure can be implemented in the form of control logic using hardware (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or computer software having modular or integrated general-purpose programmable processors. As used herein, a processor includes a multi-core processor on the same integrated chip, or a number of processing devices on a single circuit board, or a number of networked processing devices. Based on the disclosure and teachings provided herein, those skilled in the art will recognize and understand other ways and / or methods for implementing the embodiments of the present disclosure using hardware as well as combinations of hardware and software.
[0095] Any of the software components or functions described in this application can be implemented as software code executed by a processor using any suitable computer language, such as Java®, C++, or Perl, for example, using conventional or object-oriented techniques. The software code can be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission, and suitable media include magnetic media such as random access memory (RAM), read-only memory (ROM), hard drive, or floppy disk, or optical media such as compact disc (CD) or digital versatile disc (DVD), flash memory, etc. The computer-readable medium can be any combination of such storage or transmission devices.
[0096] Such programs can also be encoded and transmitted using carrier signals adapted for transmission via wired, optical and / or wireless networks that conform to various protocols including the Internet. Thus, a computer-readable medium according to an embodiment of the present invention can be made using a data signal encoded with such a program. A computer-readable medium encoded with program code may be packaged in a compatible device or provided separately from other devices (e.g., via an Internet download). Any such computer-readable medium may be present on or within a single computer product (e.g., a hard drive, CD or entire computer system), or may be present on or within different computer products within a system or network. The computer system may include a monitor, printer or other appropriate display to provide any of the results referred to herein to the user.
Example
[0097] The following examples are provided to better illustrate the claimed invention and are not to be construed as limiting the scope of the invention. All of the specific compositions, materials and methods described below are, in whole or in part, within the scope of the invention. These specific compositions, materials and methods are not intended to limit the invention, but merely illustrate certain embodiments within the scope of the invention. One of ordinary skill in the art can develop equivalent compositions, materials and methods without exercising inventive faculty and without departing from the scope of the invention. It should be understood that many modifications can be made to the described techniques herein while remaining within the scope of the invention. It is the intention of the inventors that such modifications be included within the scope of the invention.
[0098] [Example 1] Materials and Methods Nucleic Acid Extraction Genomic DNA from cells, PDX, and PDXO was purified using the DNeasy Blood&Tissue Kit (QIAGEN, Cat.69506, CA) according to the manufacturer's instructions. The integrity of the DNA was determined by a 2100 Bioanalyser (Agilent) and quantified using a NanoDrop (Thermo Scientific). One aliquot of high-quality DNA samples (OD260 / 280 = 1.8 - 2.0, OD260 / 230 ≥ 2.0, >1 μg) was used for deep NGS sequencing and WES sequencing. Total RNA from cells, PDX, and PDXO was purified using the RNeasy Mini Kit (QIAGEN, Cat.74106, CA) according to the manufacturer's instructions. The integrity of the total RNA was determined by a 2100 Bioanalyser (Agilent) and quantified using a NanoDrop (Thermo Scientific). One aliquot of high-quality RNA samples (OD260 / 280 = 1.8 - 2.2, OD260 / 230 ≥ 2.0, RIN ≥ 8.0, >1 μg) was used for deep NGS sequencing and RNAseq sequencing.
[0099] Preparation of cell line mixtures Cell line mixtures were prepared by mixing cells from two cell lines at a given ratio. Based on the cell growth rate, the cells were seeded in 15 ml of medium in a T75 to reach a cell confluence of 60% - 80%, and then incubated overnight in a Water Jacketed Incubator (SANYO). The cells were harvested during the logarithmic growth phase, counted using a hemocytometer (Chongguang), and the concentration was calculated. Then, cells from the two cell lines were mixed according to a pre-defined ratio to prepare a cell line mixture, which was subsequently centrifuged at 3,000 rpm for 5 minutes. The supernatant was aspirated, and the cell pellet was stored at -20 °C for DNA extraction. 2 Water Jacketed Incubator (SANYO) overnight. The cells were harvested during the logarithmic growth phase, counted using a hemocytometer (Chongguang), and the concentration was calculated. Then, cells from the two cell lines were mixed according to a pre-defined ratio to prepare a cell line mixture, which was subsequently centrifuged at 3,000 rpm for 5 minutes. The supernatant was aspirated, and the cell pellet was stored at -20 °C for DNA extraction.
[0100] Preparation of human-mouse DNA mixtures A series of mouse-human DNA mixture benchmark samples were prepared by mixing mouse spleen DNA and human genomic DNA (Thermo Scientific, Cat.4312660). Mouse spleen DNA was purified using the DNeasy Blood&Tissue Kit (QIAGEN, Cat.69506, CA) according to the manufacturer's instructions and quantified using NanoDrop (Thermo Scientific). Mouse spleen DNA and human genomic DNA were diluted to 200 ng / μL and then mixed at a predefined ratio. The DNA mixture was later used for deep NGS sequencing.
[0101] Barcoding deep NGS sequencing Using multiplex PCR, a target sequencing library for the Illumina sequencer with a 150 bp paired-end read length (pE150) was prepared. NGS deep sequencing covered 630 amplicons, and their sizes ranged from 160 bp to 260 bp. Genomic DNA was amplified using the IGT-EM808 polymerase mixture (iGene TechBioscience Co., Ltd, incubated at 95 °C for 3 minutes 30 seconds, 98 °C for 20 seconds and 60 °C for 8 minutes for 18 cycles, maintained at 72 °C for 5 minutes) and then purified by AMPure XP beads (Beckman, Cat.A63881).
[0102] Barcoding was performed by a second round of amplification. Briefly, the purified target amplicons were used as templates and added together with the upstream IGT-I5 index (10 μM), downstream IGT-I7 index (10 μM) and polymerase mixture for the PCR reaction. Then this mixture was placed in a thermal cycler for amplification with the following settings: incubated at 95 °C for 3 minutes 30 seconds, 98 °C for 20 seconds, 58 °C for 1 minute and 72 °C for 30 seconds for 9 cycles, fixed at 72 °C for 5 minutes. Then the barcoded library was purified using AMPure XP beads (Beckmen, Cat.A63881).
[0103] After library construction, the concentration of the obtained sequencing library was quantified using the Qubit 3.0 fluorometer dsDNA HS Assay (Thermo Fisher Scientific). The size distribution in the range of 280 bp to 420 bp was analyzed using an Agilent BioAnalyzer 2100 (Agilent). Paired-end sequencing was performed using an Illumina system according to the protocol provided by Illumina for 2×150 bp paired-end sequencing.
[0104] RNAseq and WES sequencing In RNAseq sequencing, a sequencing library focused on mRNA was constructed from total RNA. Poly-A mRNA was purified from total RNA using magnetic beads bound with oligo-dT, and then fragmented by fragmentation buffer. Short fragments were used as templates to synthesize the first-strand cDNA using reverse transcriptase and random primers, and then the second-strand cDNA was synthesized. Subsequently, the synthesized cDNA was subjected to end repair, phosphorylation, and "A" base addition according to the library construction protocol. Then, sequencing adapters were added to both ends of the cDNA fragments. After PCR amplification of the cDNA fragments, the targeted 250-350 bp fragments were purified. After library construction, the concentration of the obtained sequencing library was quantified using the Qubit 3.0 fluorometer dsDNA HS Assay (Thermo Fisher Scientific), and the size distribution was analyzed using the Agilent BioAnalyzer 2100 (Agilent). After library verification, clusters were generated using the Illumina cBOT cluster generation system in combination with the HiSeq PE Cluster Kits (Illumina). Paired-end sequencing was performed using the Illumina system according to the protocol provided by Illumina for 2×150 paired-end sequencing.
[0105] WES was performed by Wuxi Nextcode Co., Ltd. (Shanghai, China). Briefly, genomic DNA was extracted and fragmented to an average size of 180-280 bp. A DNA library was generated according to the paired-end protocol of the manufacturer of Illumina. Exons were captured by Agilent SureSelect Human All Exon V6, and subsequently sequenced by the Illumina NovaSeq platform (Illumina Inc., San Diego, CA, USA) to generate 150 bp paired-end reads.
[0106] Selection and Profiling of SNPs The inventors selected panel SNPs for human sample authentication according to several of the following criteria: 1) the SNPs are in exons; 2) since chromosomal abnormalities including large chromosomal segment deletions and duplications are common in tumors, the SNPs are located on all 22 autosomes and are sufficiently separated from each other; 3) the SNPs are in highly expressed genes; 4) the minor allele frequencies (MAFs) of the SNPs are close to 0.5 in three reference populations of the International HapMap Project, namely, Han Chinese in China (CHB), Yoruba in Nigeria (YRI), and Utah residents with Northern and Western European ancestry from the CEPH collection (CEU).
[0107] Benchmark Samples and Data Two cell line benchmark sample sets were prepared. The first set has 78 samples for three pairs of cell lines including PANC-1 and RT4, MV-4-11 and "LNCaP clone FGC", and CAL27 and Raji. Each pair has 26 samples including two pure cell lines and three replicates for eight mixing ratios by cell count (Supplementary Table S2). The second set has 22 cell lines each contaminated with a second, generally small but unspecified, ratio of a known cell line (Supplementary Table S3).
[0108] Estimating the Heterogeneity Ratio There are six informative genotype combinations (Table 11) that can be used to estimate the heterogeneity ratio from deep NGS sequencing data. These exhibit four distinguishable nucleotide frequency patterns. Combinations 1 and 2 yield the same pattern, and we use an averaging formula to calculate the percentage of minor component S2, i.e., the heterogeneity ratio. This formula yields an accurate estimate of the ratio when two combinations occur with equal frequencies, a scenario that should approach approximation when the number of SNPs is large. A similar averaging approach is used for combinations 4 and 5. When the heterogeneity ratio is low, sequencing errors can interfere with the inference of the heterogeneity ratio. To mitigate this, we use a two-step statistical approach. Assuming a sequencing error of e = 0.001 and a sequencing depth of n (n ≧ 500, any SNPs with n < 500 are discarded) at a given SNP site, the probability of observing k incorrect nucleotides follows a binomial distribution with parameters n and e.
[0109]
Number
[0110] For each n, we calculate the cumulative density function and obtain a threshold h such that the probability of observing more than h incorrect nucleotides out of n nucleotides is less than 0.01. In the sequencing data, any low-frequency nucleotides having a number of reads smaller than the corresponding threshold h are discarded. Next, we use the expectation-maximization algorithm (R package mclust, version 3.5.3 (Team, R.C.R: A language and environment for statistical computing. 3.5.3 edn (R Foundation for Statistical Computing, Vienna, Austria., 2018))) to estimate the parameters of a mixture of Gaussians (with 1 - 3 components) that models the distribution of nucleotide frequencies smaller than the maximum heterogeneity (0.2 was used for all samples in this study). If there is only a single Gaussian component or only the Gaussian component with the smallest mean that occupies more than 60% of all data points, the median of all data points is taken as the sample heterogeneity ratio; otherwise, the median of the data points in the other Gaussian components is taken as the sample heterogeneity ratio.
[0111] Determine the main components of the sample The genotype at the SNP site is determined using only nucleotides with an allele frequency greater than a threshold where the genotype for the reference sample is 10% and for the test sample that may be contaminated is 25%. The genotype similarity between the reference sample and the test sample is the percentage of SNPs with the same genotype, excluding SNPs with a sequencing depth of less than 500 in the test sample. The main component of the test sample is the reference sample with the highest genotype similarity, and this similarity must be higher than 90% (or 80%) if the heterogeneity ratio of the test sample is < 10% (or > 10%). Otherwise, the main component is not called.
[0112] Determine the minor components of the sample After estimating the heterogeneity ratio and determining the major components, we determine the minor components of the test sample. For a mixture of the major component and one of the other reference samples (e.g., all cell lines with genomic data), we obtain chimeric genotypes having probably 1 to 4 nucleotides at all SNP sites. The nucleotide frequencies are calculated using the heterogeneity ratio. Similarly, we obtain the chimeric genotype of the test sample. The two chimeric genotypes are considered identical if they carry the same nucleotides and the frequency of each nucleotide is within 3-fold. Next, we calculate the genotype similarity between the test sample and each reference sample combined with the major component. Then, a beta distribution with parameters (α,β) is fitted to the set of all pairwise genotype similarities.
[0113]
Number
[0114] In the equation, Γ(α) is the gamma function and x is the genotype similarity. The parameters were estimated by the R package fitdistrplus (version 3.5.3). Then, from the fitted beta distribution, we calculated the probability of observing any genotype similarity greater than a specific value. A quantile-quantile graph with 99% confidence bands was plotted for all observed genotype similarities for visualization. A reference sample was considered a minor component if (1) it had the highest genotype similarity, (2) its genotype similarity exceeded the 99% confidence upper limit in the quantile-quantile graph, and (3) its p-value in the fitted beta distribution < 1.0E-6.
[0115] Estimate the mixing ratio of two cell lines A cell line is used to explain the estimation of the mixing ratio for two reference samples. Assume that two cell lines S1 and S2 are mixed at a ratio of Θ for S1 and (1 - Θ) for S2, where 0 ≤ Θ ≤ 1. From the deep NGS sequencing data, the nucleotide frequencies of all n SNPs in both cell lines can be accurately estimated. For an SNP, its four nucleotide frequencies are shown as {A 1 , T 1 , G 1 , C 1} for cell line S1 and {A 2 , T 2 , G 2 , C 2} for cell line S2, and these sum to 1. In principle, one of the frequencies is close to 1 when the SNP is homozygous, and both frequencies are close to 0.5 when the SNP is heterozygous. The actual data may have some deviations due to sequencing errors and randomness, as well as the polyclonality of the cell line.
[0116] From the sequencing data of the mixed sample, the actual occurrence rates of the four nucleotides are shown as x = {n A , n T , n G , n C}. The likelihood of such an observation is
[0117]
Number
[0118] The likelihood P Θ (x i ) can be calculated for any SNPi ∈ (1, 2,..., n) using the observed data x i . The likelihood of observing the data X = {x 1 , x 2 ,..., x n} for all SNPs is
[0119]
Number
[0120] Therefore, the log-likelihood is
[0121]
Number
[0122] Next, Θ that maximizes the likelihood can be solved by a stepwise increase of Θ. The above method can be used similarly for any mixture of two human samples.
[0123] Simulation of cell line mixtures for contaminant detection Simulations were performed for three pairs of cell lines including PANC-1 and RT4, MV-4-11 and "LNCaP clone FGC", and CAL27 and Raji. All six cell lines were profiled by deep NGS sequencing to obtain their SNP fingerprints. The two paired cell lines were mixed in silico, where the ratio of the first cell line was r, and r took the following values: 0.15%, 0.30%, 0.625%, 1.25%, 2.5%, 5%, 10%, 15%, and 20%. For each SNP site, r × n nucleotides were obtained from the first cell line, where n was a random integer between 500 and 5000, and r × n was further distributed among the four nucleotides (A, T, G, C) according to their frequencies in the first cell line. Similarly, (1 - r) × n nucleotides were obtained from the second cell line. Then the ratio was reversed, and thus symmetric sampling was performed for the second cell line at ratio r.
[0124] Estimating mouse ratios from RNAseq and WES datasets Sequencing reads were mapped to the human (hg19) and mouse (mm10) genomes using the mapping tool STAR (Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15 - 21 (2013)) for RNAseq data and BWA (Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754 - 60 (2009)) for WES data, using default parameters. Reads that mapped only to the human genome or had fewer mismatches to the human genome than to the mouse genome were classified as human reads. Mouse reads were assigned similarly. Reads that mapped to both genomes with at most 2 close numbers of mismatches were unclassifiable and discarded. The mouse ratio was the proportion of mouse reads among all retained reads.
[0125] [Example 2] This example illustrates human sample authentication and contamination detection.
[0126] SNP Profiling and Fingerprint An SNP panel was selected to authenticate human samples including cell lines, xenografts, and organoids (Table 1). SNPs were profiled by deep NGS sequencing using an average depth of 3000. Each sample has a unique SNP fingerprint consisting of both nucleotide identity and frequency for all SNPs. Due to genetic drift and heterogeneity, cell lines can have fluctuating SNP fingerprints between passages and between biobanks, thus highlighting that current SNP fingerprints can be profiled for better curation. SNP fingerprints can be generated with reduced accuracy by relatively low-depth NGS data. In this example, the inventors generated SNP fingerprints for 1050 cell lines from RNAseq data profiled by the inventors and CCLE. This serves as a reference.
[0127] The inventors illustrated authentication, characterization, and intra- and inter-species contamination detection using SNP profiling data from deep NGS sequencing for 217 cell line samples, 220 PDXs, and 31 PDX-derived organoid (PDXO) samples. For cell line samples, the inventors tested mixtures of two cell lines at known mixing ratios from serial dilutions and six corresponding pure cell lines (Table 7), mixtures of two cell lines at unknown mixing ratios (Table 8), and 117 unmixed cell lines (Table 9).
[0128] Authentication of Human Samples The identity of the sample, or the major constituent components of the contaminated sample, was determined by its genotypic similarity to a library of reference samples. Among the 217 tested cell line samples, the genotypic similarity between the same cell lines was always >90% with an average of 98.6%, and the lowest was 91.7% for the A-875 cell culture with 16.7% contamination of JEG-3 (Figure 1A, Table 8). In contrast, the genotypic similarity between unrelated cell lines was almost always below 50%. There were still cell lines that were closely related or in the same synonymous group for various reasons, including display errors, contamination, being from the same patient, and one cell line being the parent of another cell line. For example, HCT-15 and HCT-8 were likely to be from the same patient; QGY-7701 was contaminated and it is a HeLa derivative. The genotypic similarity for 16 such cell line pairs in the dataset ranged from 84% to 96% (Table 10). These cell line pairs could be distinguished, except for those that were almost identical such as HLE and HLF. The genotypic similarity between the same models was on average 98.0% (87.2 - 100%) for 220 PDX and 31 PDX-derived organoids (PDXO) samples, and almost all were below 50% between different models.
[0129] Estimation of genetic heterogeneity If the sample is uncontaminated and purely monoclonal diploid, SNP sites are either homozygous or heterozygous, and the observed nucleotide frequencies are close to 1 or 0.5 in deep NGS sequencing data, and this difference comes only from errors and randomness in sequencing. In reality, cell lines may have minor clones, be aneuploid, or be contaminated (contaminants), and thus, we observed not only frequencies deviating from 0.5 and 1 at SNP sites, but also 3 or 4 nucleotides. Such information can be used to estimate the genetic heterogeneity of the sample.
[0130] The dominant clone is the major component of the sample, and the minor clone and contaminants are minor components. Based on the four observed nucleotide frequency patterns, there are six informative genotype combinations of major and minor components that can be used to estimate the SNP heterogeneity ratio (Table 11). An SNP site is informative if it gives rise to one of the four patterns. Subsequently, the sample heterogeneity ratio is estimated from the individual SNP heterogeneity ratios by a statistical modeling approach (see Example 1). Using test samples, the inventors found that cell lines without contamination have, on average, 107 informative SNP sites, while contaminated cell lines have slightly more, 112. On average, the PDX and PDXO models have 156 and 111 informative SNP sites, respectively, which reflects higher genetic heterogeneity and / or mouse contamination in the PDX model.
[0131] Detection and quantification of contamination The inventors detected sample contamination by combining three analyses. First, contaminated samples can have a high heterogeneity ratio, while uncontaminated samples do not. Among the test samples, 115 out of 118 (97.5%) of the presumptively uncontaminated cell lines had a heterogeneity ratio <2% and all had <3% (Figure 1B). In contrast, the inventors observed high heterogeneity ratios for contaminated cell lines; for example, an A-875 cell culture mixed with JEG-3 cells had a heterogeneity ratio of 15.5% (Table 8). As shown above, the heterogeneity ratio is proportional to the contamination ratio (percentage of contaminants) and is thus a good indicator of contamination. Human tumors excised from PDX models contain mouse stroma, and indeed, the inventors observed a higher heterogeneity ratio in PDX tumors (Figure 1B) caused by mouse contamination (Figure 1C). PDXO, as an in vitro culture of PDX, has a significantly lower heterogeneity ratio due to much smaller and often only trace amounts of mouse cells (Figure 1B).
[0132] Contamination was also indicated by a distinguishable right peak in the probability density of the SNP heterogeneity ratio for the samples (Figs. 2A-2F). As contamination and heterogeneity ratio increased, the peak shifted to the right and sometimes split into two peaks. The bimodal / trimodal distribution disappeared or only slightly appeared for cell lines without contamination or with very low contamination ratio (<1%) and heterogeneity ratio (<2%).
[0133] Finally, contaminants can be directly detected by statistical modeling that provides intuitive visualization and rigorous probabilistic measurements (see Example 1, Fig. 3A). In 94 cell line samples each mixed with a different cell line, the inventors were always able to accurately infer the minor contaminant cell line in the cell line when the heterogeneity ratio was ≧2% (Fig. 3B). When the heterogeneity ratio was 1-2% and <1%, the accuracy dropped to about 80% and 50%, respectively. For 8 inconclusive samples, 7 samples were characterized as clean and only 1 was evaluated as having a miscontaminated cell line. Of course, such an inference is only plausible when the contaminated cell line also has a known SNP fingerprint. The inventors detected several contaminated cell lines in our biobank, one example being the cell line "G-292 clone A141B1", which had a high heterogeneity ratio of 7.62% (Fig. 3C) and was contaminated with 6.21% of OCI-AML-2 (Fig. 3D).
[0134] After identifying the contaminated cell lines, the inventors were able to estimate the contamination ratio (i.e., the percentage of the second cell line) using a maximum likelihood approach (see Example 1). Simulation studies showed that the estimated contamination ratio was very close to the known ratio (Figure 3E). The inventors observed a tight linear correlation between the heterogeneity ratio and the contamination ratio (Figure 3F). Thus, as previously discussed, the heterogeneity ratio is a good predictor of contamination and is particularly useful when the contaminant is not a standard cell line. Also in the contaminated samples, although the contaminant sometimes contributes to most of the genetic heterogeneity, it only contributes to a part of it, and as a result, the contamination ratio is generally smaller than the corresponding heterogeneity ratio (see Table 8), and minor deviations were caused by the data processing method.
[0135] In summary, the heterogeneity ratio by its value and distribution is a reliable contamination measure for human samples. Cell line samples with a heterogeneity ratio ≧2% are very likely to be contaminated, and if the contaminant is also another cell line with SNP fingerprint information, its identity can be inferred, and the contamination ratio can be estimated with a sensitivity of ≦1% when measured by the cell or DNA mixing ratio (Tables 7 and 8).
[0136] [Example 3] This example illustrates mouse tumor model authentication.
[0137] A panel of mouse SNPs (see Table 2) was selected to validate 32 syngeneic mouse tumor models commonly used in preclinical immunomodulatory drug development, including 4T1, A20, B16-BL6, B16-F0, B16-F1, B16-F10, C1498, Colon26, CT26WT, E.G7-Ova, EL4, EMT6, H22, Hepa1-6, J558, J774A1, JC, KLN205, L1210, L5178-R, LLC, MBT2, MC38, MPC-11, Neuro-2a, P388D1, P815, Pan02, Renca, RM1, S91 and WEHI164. Most models have six unique SNPs. Colon26 and CT26WT are mouse colon adenocarcinoma models originating from the BALB / c mouse strain, each having 12 SNPs in addition to six common SNPs, for a total of 18 unique SNPs. B16-BL6, B16-F0, B16-F1 and B16-F10 are mouse melanoma cell lines in the C57BL / 6 mouse strain, all derived from B16 and thus sharing high genetic similarity. Specifically, B16 is the parental strain of B16-F0, and in turn, B16-F0 is the parental strain of B16-F1. B16-F10 is the 10th consecutive passage of B16-F0 and is the parental strain of B16-BL6 46 . The inventors first assigned test cell lines to this group using seven common SNPs, then to B16-BL6, B16-F0 and B16-F10, each having six unique SNPs, and if none of the 18 SNPs were observed, the test cell line was assigned to B16-F1. Validation for these models achieved 100% accuracy.
[0138] [Example 4] This example illustrates human-mouse interspecies admixture detection.
[0139] The inventors compared the human hg19 and mouse mm10 genomes and identified a group of segments of 100 - 300 bp (see Table 3). As a result, each segment was significantly different between humans and mice by insertions, deletions, and point mutations (sequence similarity of 31 - 97%), but still had identical flanking sequences such that common primer pairs could be designed. After NGS sequencing, the inventors separated human reads from mouse reads, calculated the mouse ratio for all segments, and used the median of these ratios as the mouse ratio in the human - mouse mixed sample. This method demonstrated extremely high accuracy in a set of benchmark samples where mouse and human DNA were mixed by serial dilution (Figure 4A). The inventors also developed a method to estimate mouse content from RNAseq and WES data (see Example 1). The inventors compared three methods in estimating the mouse ratio in 220 PDX and 31 PDXO models (Figures 4B - C). DNA (for WES and deep NGS sequencing) and RNA (for RNAseq) were extracted and sequenced from the same samples of the models to remove sample variation. PDXO models generally had low mouse content. In PDX models, the mouse ratio accurately estimated from deep NGS sequencing data was the highest, followed by that from RNAseq, and then that from WES. This is mainly due to the fact that the exon capture kit used in WES was designed to enrich human exons and had low hybridization affinity for homologous mouse exons. RNAseq used a species - non - preferential polyA enrichment protocol, but gene expression had large spatiotemporal variability in human tumors and mouse stroma of PDX. In fact, the inventors observed a very strong quadratic relationship for the mouse ratio between deep NGS sequencing data and WES data (R = 0.96, Figure 4D), but a much weaker linear correlation between deep sequencing data and RNAseq data (R = 0.62).
[0140] [Example 5] This example illustrates the detection of mycoplasma in a sample.
[0141] The inventors used a pair of universal primers for the detection of all mycoplasma species, the efficacy of which has been proven, and 11 pairs of primers for the detection of 11 mycoplasmas including A. laidlawii, M. arginine, M. fermentans, M. genitalium, M. hominis, M. hyorhinis, M. orale, M. pneumonia, M. salivarium and U. urealyticum (Molla Kazemiha, V. et al. Cytotechnology 61, 117-24 (2009)). The inventors identified one mycoplasma-contaminated cell line in the biobank by deep NGS sequencing method and subsequently verified it by a mycoplasma detection kit.
[0142] [Example 6] This example illustrates population structure analysis and sex determination.
[0143] Of the panel of SNPs used for human sample authentication, 143 were characterized by the International HapMap Project (International HapMap, C. The International HapMap Project. Nature 426, 789-96 (2003)). The inventors performed population structure analysis of the following three reference populations using fastSTRUCTURE (Raj, A., Stephens, M. & Pritchard, J. K. Genetics 197, 573-89 (2014)): Han Chinese in China (CHB), Yoruba in Nigeria (YRI), and Utah residents with Northern and Western European ancestry from the CEPH collection (CEU). All 406 individuals were unambiguously assigned with high probability. The inventors then profiled 423 PDX models derived from East Asian patients and 634 PDX models derived from Western patients in the United States. All East Asian PDX models, with one exception only, have a dominant CHB composition. Most of the Western PDX models dominantly have a CEU composition, and the rest have a major CHB or YRI composition, or a mixture of two or three of the reference populations. The inventors also used three SNPs on the Y chromosome for sex inference (Table 3), which was always accurate except for tumor samples in which the Y chromosome was lost.
[0144] Although the present disclosure has been particularly shown and described with reference to specific embodiments, some of which are preferred embodiments, it should be understood by those skilled in the art that various changes in form and detail can be made without departing from the spirit and scope of the disclosure as disclosed herein.
[0145] References 1. Identity crisis. Nature 457, 935-6 (2009). 2. American Type Culture Collection Standards Development Organization Workgroup, A.S.N. Cell line misidentification: the beginning of the end. Nat Rev Cancer 10, 441-8 (2010). 3. Capes-Davis, A. et al. Match criteria for human cell line authentication: where do we draw the line? Int J Cancer 132, 2510-9 (2013). 4. Gartler, S.M. Apparent Hela cell contamination of human heteroploid cell lines. Nature 217, 750-1 (1968). 5. Lacroix, M. Persistent use of "false" cell lines. Int J Cancer 122, 1-4 (2008). 6. Lorsch, J.R., Collins, F.S. & Lippincott-Schwartz, J. Cell Biology. Fixing problems with cell lines. Science 346, 1452-3 (2014). 7. Fusenig, N.E., Capes-Davis, A., Bianchini, F., Sundell, S. & Lichter, P. The need for a worldwide consensus for cell line authentication: Experience implementing a mandatory requirement at the International Journal of Cancer. PLoS Biol 15, e2001438 (2017). 8. Yu, M. et al. A resource for cell line authentication, annotation and quality control. Nature 520, 307-11 (2015). 9. Bian, X., Yang, Z., Feng, H., Sun, H. & Liu, Y. A Combination of Species Identification and STR Profiling Identifies Cross-contaminated Cells from 482 Human Tumor Cell Lines. Sci Rep 7, 9774 (2017). 10. Horbach, S. & Halffman, W. The ghosts of HeLa: How cell line misidentification contaminates the scientific literature. PLoS One 12, e0186281 (2017). 11. de Maagd, R.A. et al. Identification of Bacillus thuringiensis delta-endotoxin Cry1C domain III amino acid residues involved in insect specificity. Appl Environ Microbiol 65, 4369-74 (1999). 12. Azari, S., Ahmadi, N., Tehrani, M.J. & Shokri, F. Profiling and authentication of human cell lines using short tandem repeat (STR) loci: Report from the National Cell Bank of Iran. Biologicals 35, 195-202 (2007). 13. Wu, M.L. et al. A 2-yr service report of cell line authentication. In Vitro Cell Dev Biol Anim 49, 743-5 (2013). 14. Masters, J.R. HeLa cells 50 years on: the good, the bad and the ugly. Nat Rev Cancer 2, 315-9 (2002). 15. MacLeod, R.A. et al. Widespread intraspecies cross-contamination of human tumor cell lines arising at source. Int J Cancer 83, 555-63 (1999). 16. Cosme, B. et al. Are your results valid? Cellular authentication a need from the past, an emergency on the present. In Vitro Cell Dev Biol Anim 53, 430-434 (2017). 17. Ye, F., Chen, C., Qin, J., Liu, J. & Zheng, C. Genetic profiling reveals an alarming rate of cross-contamination among human cell lines used in China. FASEB J 29, 4268-72 (2015). 18. Freedman, L.P. et al. The culture of cell culture practices and authentication--Results from a 2015 Survey. Biotechniques 59, 189-90, 192 (2015). 19. Nims, R.W. & Reid, Y. Best practices for authenticating cell lines. In Vitro Cell Dev Biol Anim 53, 880-887 (2017). 20. Almeida, J.L., Cole, K.D. & Plant, A.L. Standards for Cell Line Authentication and Beyond. PLoS Biol 14, e1002476 (2016). 21. Almeida, J.L. et al. Interlaboratory study to validate a STR profiling method for intraspecies identification of mouse cell lines. PLoS One 14, e0218412 (2019). 22. Zaaijer, S. et al. Rapid re-identification of human samples using portable DNA sequencing. Elife 6(2017). 23. Yousefi, S. et al. A SNP panel for identification of DNA and RNA specimens. BMC Genomics 19, 90 (2018). 24. Jobling, M.A. & Gill, P. Encoded evidence: DNA in forensic analysis. Nat Rev Genet 5, 739-51 (2004). 25. Sanchez, J.J. et al. A multiplex assay with 52 single nucleotide polymorphisms for human identification. Electrophoresis 27, 1713-24 (2006). 26. Didion, J.P. et al. SNP array profiling of mouse cell lines identifies their strains of origin and reveals cross - contamination and widespread aneuploidy. BMC Genomics 15, 847 (2014). 27. Liang - Chu, M.M. et al. Human biosample authentication using the high - throughput, cost - effective SNPtrace(TM) system. PLoS One 10, e0116218 (2015). 28. Pengelly, R.J. et al. A SNP profiling panel for sample tracking in whole - exome sequencing studies. Genome Med 5, 89 (2013). 29. Morgan, A.P. et al. The Mouse Universal Genotyping Array: From Substrains to Subspecies. G3 (Bethesda) 6, 263 - 79 (2015). 30. Castro, F. et al. High - throughput SNP - based authentication of human cell lines. Int J Cancer 132, 308 - 14 (2013). 31. El - Hoss, J. et al. A single nucleotide polymorphism genotyping platform for the authentication of patient derived xenografts. Oncotarget 7, 60475 - 60490 (2016). 32. Ruitberg, C.M., Reeder, D.J. & Butler, J.M. STRBase: a short tandem repeat DNA database for the human identity testing community. Nucleic Acids Res 29, 320-2 (2001). 33. van der Meer, D. et al. Cell Model Passports-a hub for clinical, genetic and functional datasets of preclinical cancer models. Nucleic Acids Res 47, D923-D929 (2019). 34. Tuveson, D. & Clevers, H. Cancer modeling meets human organoid technology. Science 364, 952-955 (2019). 35. Day, C.P., Merlino, G. & Van Dyke, T. Preclinical mouse cancer models: a maze of opportunities and challenges. Cell 163, 39-53 (2015). 36. Guo, S. et al. Molecular Pathology of Patient Tumors, Patient-Derived Xenografts, and Cancer Cell Lines. Cancer Res 76, 4619-26 (2016). 37. Khaled, W.T. & Liu, P. Cancer mouse models: past, present and future. Semin Cell Dev Biol 27, 54-60 (2014). 38. Li, Q.X., Feuer, G., Ouyang, X. & An, X. Experimental animal modeling for immuno-oncology. Pharmacol Ther 173, 34-46 (2017). 39. Chao, C. et al. Patient-derived Xenografts from Colorectal Carcinoma: A Temporal and Hierarchical Study of Murine Stromal Cell Replacement. Anticancer Res 37, 3405-3412 (2017). 40. Fasterius, E. & Al-Khalili Szigyarto, C. Analysis of public RNA-sequencing data reveals biological consequences of genetic heterogeneity in cell line populations. Sci Rep 8, 11226 (2018). 41. Ghandi, M. et al. Next-generation characterization of the Cancer Cell Line Encyclopedia. Nature 569, 503-508 (2019). 42. Vermeulen, S.J. et al. Did the four human cancer cell lines DLD-1, HCT-15, HCT-8, and HRT-18 originate from one and the same patient? Cancer Genet Cytogenet 107, 76-9 (1998). 43. Rebouissou, S., Zucman-Rossi, J., Moreau, R., Qiu, Z. & Hui, L. Note of caution: Contaminations of hepatocellular cell lines. J Hepatol 67, 896-897 (2017). 44. Barretina, J. et al. The Cancer Cell Line Encyclopedia enables predictive modelling of anticancer drug sensitivity. Nature 483, 603-7 (2012). 45. Bairoch, A. The Cellosaurus, a Cell-Line Knowledge Resource. J Biomol Tech 29, 25-38 (2018). 46. Molla Kazemiha, V. et al. PCR-based detection and eradication of mycoplasmal infections from various mammalian cell lines: a local experience. Cytotechnology 61, 117-24 (2009). 47. International HapMap, C. The International HapMap Project. Nature 426, 789-96 (2003). 48. Raj, A., Stephens, M. & Pritchard, J.K. fastSTRUCTURE: variational inference of population structure in large SNP data sets. Genetics 197, 573-89 (2014). 49. Masters, J.R. et al. Short tandem repeat profiling provides an international reference standard for human cell lines. Proc Natl Acad Sci U S A 98, 8012-7 (2001). 50. Hideyuki Tanabe, Y.T., Daisuke Minegishi, Miharu Kurematsu, Tohru Masui, Hiroshi Mizusawa. Cell line individualization by STR multiplex system in the cell bank found cross-contamination between ECV304 and EJ-1 / T24. Tiss. Cult. Res. Commun. 18, 329-338 (1999). 51. Team, R.C. R: A language and environment for statistical computing. 3.5.3 edn (R Foundation for Statistical Computing, Vienna, Austria., 2018). 52. Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15-21 (2013). 53. Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754-60 (2009).
[0146]
Table 1-1
Table 1-2
Table 1-3
[0147]
Table 2-1
Table 2-2
Table 2-3
Table 2-4
Table 2-5
Table 2-6
[0148]
Table 3
[0149]
Table 4
[0150]
Table 5
[0151]
Table 6-1
Table 6-2
[0152]
Table 7-1
Table 7-2
[0153]
Table 8-1
Table 8-2
[0154]
Table 9-1
Table 9-2
Table 9-3
[0155]
Table 10
[0156]
Table 11
Claims
1. A method for authenticating a human or mouse sample, comprising: obtaining nucleic acid from said human or mouse sample; (i) detecting the genotype at all of the human single nucleotide polymorphism (SNP) loci set forth in SEQ ID NOs: 1-237 of said human sample, or (ii) detecting the genotype at all of the mouse SNP loci set forth in SEQ ID NOs: 238-436 of said mouse sample; (i) comparing the genotype of the human sample to a reference genotype for all of the human SNPs set forth in SEQ ID NOs: 1-237 in a library of human reference samples of known identity, or (ii) comparing the genotype of the mouse sample to a reference genotype for all of the mouse SNPs set forth in SEQ ID NOs: 238-436 in a library of mouse reference samples of known identity; and (i) determining the authenticity of the human sample by similarity to the genotype of the human reference sample in the library of human reference samples, or (ii) determining the authenticity of the mouse sample by similarity to the genotype of the mouse reference sample in the library of mouse reference samples. A method comprising:
2. 2. The method of claim 1, wherein the human or mouse sample is a cell, tissue, organoid, or a combination thereof.
3. The method of claim 2, wherein the human or mouse sample is a cell line or a tumor tissue.
4. 2. The method of claim 1, wherein the human or mouse sample contains a contaminant, the method further comprising determining the percentage of the contaminant in the sample.
5. The method of claim 1, wherein the human sample contains a contaminant, and the method further comprises determining the identity of the contaminant.
6. The method of claim 1, wherein the detecting step uses next generation sequencing (NGS).
7. The method of claim 1 , wherein the nucleic acid is barcoded.
8. 2. The method of claim 1, further comprising identifying the sex of the subject from which the human or mouse sample was obtained by detecting a sex chromosome SNP selected from the group shown in Table 3.
9. 10. The method of claim 1, further comprising identifying the ethnicity of the subject from which the human sample was obtained.
10. The method of claim 1 (excluding human diagnostic methods), further comprising detecting the presence of a virus or mycoplasma in the human or mouse sample.
11. 2. The method of claim 1, wherein the mouse sample is a sample from a mouse tumor model selected from the group consisting of 4T1, A20, B16-BL6, B16-F0, B16-F1, B16-F10, C1498, Colon26, CT26WT, E.G7-Ova, EL4, EMT6, H22, Hepa1-6, J558, J774A1, JC, KLN205, L1210, L5178-R, LLC, MBT2, MC38, MPC-11, Neuro-2a, P388D1, P815, Pan02, Renca, RM1, S91, and WEHI164.
12. A kit for authenticating a human or mouse sample, comprising: (i) primers for detecting all of the human SNPs represented by SEQ ID NOs: 1 to 237 in a human sample, or (ii) primers for detecting all of the mouse SNPs represented by SEQ ID NOs: 238 to 436 in a mouse sample; and Agent for amplifying DNA fragments containing human or mouse SNPs using said primers Including the kit.
13. 13. The kit of claim 12, further comprising primers for detecting sex chromosome SNPs selected from the group shown in Table 3.
14. The kit of claim 13, further comprising primers for detecting viral infection or mycoplasma contamination in the sample.
15. A microarray for identifying a human or mouse sample, comprising: A microarray comprising probes for detecting (i) all of the human SNPs represented by SEQ ID NOs: 1 to 237 in a human sample, or (ii) all of the mouse SNPs represented by SEQ ID NOs: 238 to 436 in a mouse sample.
16. A non-transitory computer readable medium having instructions stored thereon, the instructions, when executed by a processor, causing the processor to: (i) searching for genotypes at all of the human SNP loci set forth in SEQ ID NOs: 1-237 in a human sample, or (ii) searching for genotypes at all of the mouse SNP loci set forth in SEQ ID NOs: 238-436 in a mouse sample; (i) comparing the genotype of the human sample to a human reference genotype for all of the human SNPs set forth in SEQ ID NOs: 1-237 in a library of human reference samples of known identity, or (ii) comparing the genotype of the mouse sample to a mouse reference genotype for all of the mouse SNPs set forth in SEQ ID NOs: 238-436 in a library of mouse reference samples of known identity; and (i) determining the authenticity of the human sample by similarity to the genotype of the human reference sample in the library of human reference samples; or (ii) determining the authenticity of the mouse sample by similarity to the genotype of the mouse reference sample in the library of mouse reference samples. A computer-readable medium for causing a