Methods and reagents for detecting and evaluating genotoxicity
Patent Information
- Application Number
- JP2023222575
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-09-26
- Filing Date
- 2023-12-28
- Publication Date
- 2026-09-14
- Estimated Expiration
- 2039-02-13
Smart Images

Figure 0007920128000005 
Figure 0007920128000006 
Figure 0007920128000007
Abstract
Description
Technical Field
[0001] Cross-Reference to Related Applications This application claims priority and benefit to U.S. Provisional Patent Application No. 62 / 630,228, filed on February 13, 2018, and U.S. Provisional Patent Application No. 62 / 737,097, filed on September 26, 2018, the disclosures of which are incorporated herein by reference in their entireties. Background Art
[0002] Genotoxicity refers to the destructive property of an agent or process (i.e., a genotoxic agent) that causes damage to genetic material (e.g., DNA, RNA). In germ cell lines, damage to nucleic acid material has the potential to cause heritable germline mutations, while damage to nucleic acid material in somatic cells has the potential to cause somatic mutations. In some cases, such somatic mutations may lead to malignant tumors or other diseases. It has been established that exposure to genotoxic agents can cause such nucleic acid damage directly or indirectly, or in some cases, can directly or indirectly trigger nucleic acid damage. For example, a genotoxic agent may interact directly with genetic material, alter the nucleotide sequence itself or its structure, or result in chemical modification (e.g., adduct formation or breakage), which may induce (or increase the probability of inducing) changes to the nucleotide sequence when the nucleic acid is attempted to be replicated, repaired, or otherwise processed by cellular mechanisms. Genotoxic agents may be naturally occurring chemicals or processes (e.g., coal, radium or UV light), or artificially produced chemicals, processes or therapies (e.g., industrial urethane, X-ray machines, many chemotherapeutic agents, and some forms of gene therapy).
[0003] Other genotoxic substances can indirectly trigger nucleic acid damage by activating cellular pathways that reduce the fidelity of DNA replication. For example, this may be the direct or indirect activation of cell cycle mechanisms that bypass normal checkpoints, or it may be by reducing normal nucleic acid repair (e.g., direct or indirect dysregulation of any one of many nucleic acid repair pathways, including mismatch repair (MMR), nucleotide excision repair (NER), base excision repair (BER), double-strand break repair (DSBR), transcription-coupled repair (TCR), and non-homologous end joining (NHEJ)). Other genotoxic substances can act indirectly by promoting a cellular environment that is itself genotoxic. One example of such an environment is "oxidative stress," which can result from increased production of reactive oxygen species in cells that can cause damage to genetic material (e.g., by stimulation of immune-mediated inflammation), by altering the chemical composition of the sequence itself, or by structurally altering the nucleic acid chain. Another indirect form of genotoxicity is that which is a drug or process that suppresses specific aspects of an organism's immune system. Such a reduction in the immune surveillance mechanism may cause genotoxicity in the organism by enabling the growth of potentially genotoxic microorganisms through one of several mechanisms (for example, by causing inflammation in certain tissues or by accelerating the progression of the cell cycle). Furthermore, such drugs or processes may cause a genotoxic burden on the organism by reducing its normal ability to eliminate cells with genetic abnormalities that are oncogenic, which are otherwise eliminated through this mechanism. Many mechanisms of genotoxicity remain undiscovered.
[0004] Genotoxic substances can originate from a variety of internal and external sources. For example, external (i.e., exogenous) sources may include chemical substances or mixtures of chemical substances (e.g., pharmaceuticals, industrial / manufacturing by-products, chemical waste, cosmetics, household cleaning agents, plasticizers, cigarette smoke, solvents, etc.); heavy metals, airborne particulate matter, pollutants, food, radiation (e.g., photons, e.g., gamma rays, X-rays, particle radiation, or mixtures thereof), physical forces from the natural environment or devices (e.g., magnetic fields, gravitational fields, accelerating forces, etc.); substances produced by other organisms (e.g., viruses, parasites, bacteria, protozoa, fungi), or other naturally occurring organisms (e.g., fungi, plants, animals, bacteria, bacteria, protozoa, etc.). Certain crops themselves (e.g., tobacco) contain genotoxic substances known in their natural form. Staple crops may be contaminated with genotoxic substances during growth (e.g., contamination of irrigation water by industrial waste), harvesting (e.g., inadvertently harvesting crops together with aristocholia, which produces the mutagenic aristolochinic acid), storage (e.g., moist legume and grain silos causing the growth of aspergillus species, which produce the mutagenic aflatoxin), or preparation (e.g., smoking and other preservation methods of meat, which produce many forms of genotoxic substances, or high-temperature cooking of starch, which can produce the mutagenic acrylamide). Some examples of internal (i.e., endogenous) sources may include biochemical processes or their consequences. For example, a chemical agent may be determined to be a genotoxic substance if it is a precursor to a mutagen resulting from metabolic activation. Other examples would include stimulants of inflammatory pathways (e.g., stress, autoimmune diseases), or inhibitors of apoptosis or immune surveillance mechanisms. Regardless of the source, many factors play a role in determining whether a drug or process is potentially genotoxic, mutagenic, or carcinogenic (i.e., cancer-causing).
[0005] In certain applications, the ability to detect and quantify mutagenic processes is crucial for assessing cancer risk in humans and predicting the effects of carcinogen exposure. Similarly, evaluating the potential of chemical compounds or other agents to induce nucleic acid mutations is an essential element of pre-market product safety testing (e.g., pharmaceuticals, cosmetics, food, manufacturing by-products, etc.). Current methods for identifying genotoxic substances are cumbersome, costly, time-consuming (e.g., several years between exposure and symptoms), may not reflect the actual effects in the human body (only for specific model organisms), and in some cases, pinpointing the exact causative agent is difficult. For example, sometimes detecting an increased incidence of disease in a population of subjects (e.g., cancer clusters) is necessary before initiating a search for genotoxic substances (e.g., safety analysis of pharmaceuticals and food, investigations of environmental pollutants, or environmental dumping).
[0006] Conventional measures of somatic mutation in vivo are indirectly estimated from selection-based assays in transgenic animals where the effects are estimated from bacteria, cell cultures, or small artificial reporters with whole-genome effects. Therefore, the assays currently in use are imperfect surrogates for the true genotoxicity potential of a compound in vivo, and these assays, while labor-intensive, provide only a limited amount of information about a compound's mutagenic potential. Many compounds that exhibit mutagenic potential in artificial bacterial systems (i.e., Ames assays) likely do not accurately reflect the true risk in humans, unnecessarily displacing compounds with therapeutic potential in other contexts from development or commercial use. Similarly, some compounds with carcinogenic potential function this way through indirect, undetectable mutagenic mechanisms within bacteria. Such compounds can harm subjects because the risk cannot be adequately recognized early on.
[0007] In vivo mammalian reporter systems, such as transgenic rodent assays (e.g., BigBlue® mouse and rat, and Muta® Mouse), provide a better approximation of human drug efficacy than bacteria. While these are limited in that animals are not complete representatives of humans, mammalian transgenic assays remain valuable for early preclinical safety testing; however, these assays are complex and still to some extent artificial. The BigBlue® assay, for example, relies on a reporter-based system, where some mutations occurring within the multicopy lambda phage transgene can be identified phenotypically after the reporter is recovered by a shuttle vector that is later transfected into bacteria. Not all mutations occurring within the 294BP reporter gene can be detected, as many do not confer a phenotype. The transgene itself is highly condensed and methylated and does not represent the highly variable transcriptional and condensation states of the broad genome. When mutant molecules pass through viral and bacterial mechanisms, they can introduce artificial mutations, and the unique obstacles that occur in each process mean that the allele fraction of the mutation cannot be quantified. Furthermore, the tests require the use of a limited number of specific strains of a species. Also, rodents themselves are not a complete representative of humans. For example, aflatoxin is highly mutagenic in humans, but in mice, its detoxification is facilitated after sexual maturation when certain metabolic enzymes are expressed, and it is not meaningfully carcinogenic. Transgenic rodents are still the current representative standard approved by the U.S. Food and Drug Administration (FDA) and other regulatory agencies as an effective genotoxicity criterion that can be used as a carcinogenic substitute in some test situations, but they are far from being the optimal, widely usable tool for assessing the potential of a compound to cause cancer in humans.
[0008] There is a need for a rapid, flexible, and reliable method that enables the direct measurement of the potential genotoxicity of factors / drugs / environments that cause nucleic acid mutations and damage contributing to specific health risks to which subjects may be exposed (i.e., cancer / malignant tumors / neoplasms, neurotoxicity, neurodegeneration, infertility, birth defects, etc.). This method should be usable at any genomic locus of any tissue type and / or cell type within any type of organism, without requiring clonal selection (as required in representative baseline tests of prior art), while providing information (either presumed or direct) about the mechanism by which oncogenic factors cause mutations or other genotoxic damage in vivo that lead to cancer development or other disease or disorder in the subject / organism or in another organism modeled by that subject / organism.
[0009] If sufficiently accurate and convenient tools possessing these characteristics were available, they would have numerous applications, for example, in preclinical and clinical drug safety testing; in preventing, diagnosing, and treating diseases and disorders associated with genotoxic substances; in detecting and identifying mutation-causing substances / drugs and their mechanisms of action; and in interactions with other industries as a whole (e.g., environmental pollution testing and determination of threshold levels for toxic outbreaks, safety testing of high-throughput consumer products, diagnosis and treatment of patients suspected of toxic exposure, and national security risk assessment of intentional or unintentional releases of genotoxic substances). [Overview of the project]
[0010] This technology relates to methods, systems, and reagent kits for evaluating genotoxicity. In particular, several embodiments of this technology relate to the use of double-strand sequencing to assess the potential genotoxicity of compounds (e.g., chemical compounds) and / or environmental agents (e.g., radiation) in exposed subjects. For example, various embodiments of this technology involve providing double-strand sequencing methods that enable the direct measurement of compound-induced mutations in any genomic context of any organism without requiring arbitrary clonal selection. Further examples of this technology relate to methods for detecting and evaluating genomic in vivo mutagenesis using double-strand sequencing and associated reagents. Various aspects of this technology have numerous applications not only in preclinical and clinical drug safety testing but also in other industries as a whole.
[0011] In one embodiment, the technology includes a method for detecting and quantifying genomic mutations that occur in vivo in a subject after exposure to a mutagens, comprising: (1) double-strand sequencing one or more target double-stranded DNA molecules extracted from a subject exposed to a mutagens; (2) creating an error-corrected consensus sequence for the target double-stranded DNA molecules; (3) identifying the mutation spectrum of the target double-stranded DNA molecules; and (4) calculating the mutant frequency of the target double-stranded DNA molecules by calculating the number of unique mutations per one or more sequenced double-stranded base pairs.
[0012] In another embodiment, the technique includes a method for creating a mutagenic signature of a test compound, comprising: (1) double-strand sequencing of DNA fragments extracted from a living organism (e.g., a test animal) exposed to the test compound; and (2) creating a mutagenic signature of the test compound. The method may further include calculating the mutant frequencies of multiple DNA fragments by calculating the number of unique mutations per sequenced double-strand base pair.
[0013] In another embodiment, the present technology includes a method for evaluating the potential genotoxicity of a compound, comprising: (1) double-strand sequencing a target DNA fragment extracted from a test animal exposed to the compound to create an error-corrected consensus sequence of the target DNA fragment; (2) creating a mutagenic signature of the compound from the error-corrected consensus sequence; and (3) determining whether exposure to the compound yields a mutagenic signature that represents a sufficiently genotoxic compound.
[0014] In another embodiment, the technology includes a kit comprising reagents along with instructions for performing the method disclosed herein for detecting and quantifying genotoxic substances. The kit may further include a computer program product that is installed on an electronic computing device (e.g., a laptop / desktop computer, tablet, etc.) or accessible via a network (a remote server containing a database of subject records and detected genotoxic substances). The computer program product is embodied in a non-temporary, computer-readable medium that, when executed on a computer, performs the steps of the method using the kit disclosed herein for detecting and identifying genotoxic substances.
[0015] In another embodiment, the technology includes a network computer system for identifying or confirming a subject's exposure to at least one genotoxic substance, comprising: (1) a remote server; (2) a plurality of user-computer devices capable of utilizing kits disclosed herein for extracting, amplifying, and sequencing a subject's sample; (3) a third-party database (optional element) containing known genotoxic substance profiles; and (4) a wired or wireless network for transmitting electronic communications between the computer devices, the database, and the remote server. The remote server further comprises (a) a database storing the user's genotoxic substance recording results and records of genotoxic substance profiles (e.g., spectrum, frequency, mechanism of action, etc.), and (b) one or more processors communicably connected to memory and one or more non-temporary, computer-readable recording devices or media containing instructions for the processors, wherein the processors are configured to execute instructions for performing operations including the steps of correcting errors in double-stranded sequencing fragments and calculating the mutation spectrum, mutation frequency and triple mutation spectrum of detected drugs, from which the identity of at least one genotoxic substance can be determined.
[0016] The technology further includes a non-temporary, computer-readable storage medium, which, when executed by one or more processors, includes instructions for performing a method for determining whether a subject has been exposed to at least one genotoxic substance and / or for determining the identity of at least one genotoxic substance, the method comprising the steps of correcting errors in a double-stranded sequencing fragment and calculating the mutation spectrum, variant frequency and tri-spectrum of the detected drug, from which the identity of at least one genotoxic substance is determined.
[0017] The technology further includes a computer-aided method for determining whether a subject has been exposed to at least one genotoxic substance and / or for determining the identity of at least one genotoxic substance, the method comprising the steps of correcting errors in a double-stranded sequencing fragment and calculating the mutation spectrum, variant frequency and tri-spectrum of the detected drug, from which the identity of at least one genotoxic substance is determined.
[0018] In another embodiment, the technology includes methods, systems, and kits for diagnosing and treating subjects exposed to genotoxic substances. Diagnosing includes detecting at least one genotoxic substance to which a subject has been exposed and / or ingested, and treating includes administering a treatment protocol (e.g., a pharmaceutical) to eliminate future exposure and / or ingestion of the genotoxic substance and / or block the biological effects of the genotoxic substance and / or interfere with them in other circumstances.
[0019] In another embodiment, the technology includes methods, computer-based systems, and kits for detecting and identifying carcinogens and their mechanisms of action for preclinical and clinical drug safety testing, in relation to other industries (e.g., toxic environmental pollutants, high-throughput consumer products, and drug safety testing).
[0020] In another embodiment, the technology includes methods, systems, and kits for identifying novel genotoxic substances using error-corrected double-strand sequencing and / or subsequently determining the safety threshold dose (e.g., weight, volume, concentration) and / or safety threshold variant frequency of the genotoxic substance to which a subject may be exposed before the subject is at risk of developing a disease or disorder associated with the genotoxic substance (e.g., used in setting standards for the Environmental Protection Agency, used in diagnosing and treating subjects exposed to genotoxic substances).
[0021] In another embodiment, the present technology provides methods, systems and kits for preventing the development of a mutation-associated disease or disorder in a subject by determining whether the subject has been exposed to a genotoxic substance exceeding a safety threshold level (i.e., the amount of the genotoxic substance and / or the variant frequency of the genotoxic substance and a triple signature), and in such case, providing a prophylactic treatment to prevent, inhibit or suppress the onset of the disease, including said methods, systems and kits.
[0022] One aspect of the present technology includes the ability to detect disease-causing mutations within several days, weeks, months or years after exposure to genotoxic substance-induced mutations, before the onset of full-blown disease. Typically, full-blown disease remains undiagnosed for many years (e.g., 10 to 20 years from asbestos exposure to lung cancer development). The methods and kits disclosed herein enable detection of genomic mutations that cause disease development immediately after exposure, rather than waiting many years until symptoms appear.
[0023] Another aspect of the present technology includes the ability to predict whether a subject has an increased risk of developing a disease or disorder caused by genotoxic substance-induced mutations, from as early as about 2 to 5 days after potential exposure to the genotoxic substance to several years after exposure, and if so, provides prophylactic treatment and periodic screening to detect disease onset at an early stage.
[0024] Another aspect includes a DNA library comprising a plurality of double-stranded isolated genomic DNA fragments, each of which is ligated to one or more desired adapter molecules, and a method for producing said DNA library.
[0025] Another aspect includes a high-throughput method for rapidly screening a plurality of compounds to identify which compounds are genotoxic.
[0026] Another aspect includes a high-throughput method for rapidly screening a plurality of different tissue / cell types from the same subject to determine whether the subject has been exposed to any genotoxic substance.
[0027] Another aspect includes a high-throughput method for rapidly screening a plurality of tissues and cells derived from different subjects to determine the proportion of a population that has been exposed to any genotoxic substance.
[0028] Another aspect includes directly or inferentially determining the "mechanism of action" of a genotoxic substance that causes exposure to the genotoxic substance that results in mutations associated with a particular disease or disorder.
[0029] Other embodiments, aspects and advantages of the present technology are further described in the detailed description below. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Many aspects of the present disclosure can be better understood with reference to the following drawings. The components in the drawings are not necessarily to scale. Instead, emphasis is placed on clearly illustrating the principles of the present disclosure. [Figure 1-1] Figure 1A shows a nucleic acid adapter molecule for use with some embodiments of the present technology, and a double-stranded adapter nucleic acid molecule complex obtained from ligation of the adapter molecule to a double-stranded nucleic acid fragment according to an embodiment of the present technology. [Figure 1-2] Figures 1B and 1C are schematic diagrams illustrating steps of various double-stranded sequencing methods according to embodiments of the present technology. [Figure 2A] Figure 2A is a schematic diagram of various method schemes for using in vivo animal assays to predict human cancer risk of a test compound, including a conventional long-term rodent carcinogenicity assay (left scheme), a conventional transgenic rodent mutagenicity assay with in vitro selection (middle scheme), and a mutagenicity assay via a direct DNA sequencing scheme according to an aspect of the present technology (right scheme). [Figure 2B] Figures 2B and 2C are conceptual diagrams of a method scheme for using double-strand sequencing to evaluate in vitro mutagenesis of a test compound in human cells grown in culture, according to an aspect of this technology. [Figure 2C] Same as above. [Figure 3-1] Figures 3A to 3D are box plots showing the calculated mutant frequencies for double-strand sequencing in the liver and bone marrow after mutagenic treatment, according to one embodiment of the present technology. [Figure 3-2] Same as above. [Figure 3-3] Figure 3E is a plot showing the relative cII variant doubling rate in the BigBlue® cII plaque assay compared to the double-strand sequencing assays shown in Figures 3A-3D, according to one embodiment of this technology. [Figure 3-4] Figure 3F shows the percentage of single nucleotide variants (SNVs) within the cII gene in BigBlue® mouse tissue and individually isolated mutant plaques produced from double-strand sequencing of cII gDNA from BigBlue® mouse tissue, according to one embodiment of this technology. [Figure 3-5] Figures 3G and 3H show the distribution of mutations identified by direct double-strand sequencing of cII across all BigBlue® tissue types and treatment groups, based on codon location and functional consequences, according to one embodiment of the present technology. [Figure 3-6] Same as above. [Figure 4] This is a bar graph showing the mutant frequencies measured by double-strand sequencing in multiple samples from each treatment group, according to one embodiment of this technology. [Figure 5A] Figures 5A and 5B are bar graphs showing the frequency of endogenous gene variants compared to cII transgenes in the liver, as determined by double-strand sequencing according to one embodiment of this technology. [Figure 5B] Same as above. [Figure 5C]Figure 5C is a box plot graph showing the SNV variant frequency (MF) calculated for double-strand sequencing of gene regions in the liver and bone marrow according to one embodiment of the present technology, for the indicated treatment category. [Figure 5D] Figure 5D is a scatter plot showing individual measurements of the aggregated data shown in Figure 5C, according to one embodiment of the present technology. [Figure 6] Figure 6 is a bar graph showing the mutation spectrum measured by double-strand sequencing according to one embodiment of this technology. [Figure 7-1] Figures 7A-7C are graphs showing the trinucleotide mutation spectrum for vehicle control (Figure 7A), benzo[a]pyrene (Figure 7B), and N-ethyl-nitrosourea (Figure 7C) according to one embodiment of this technology. [Figure 7-2] Same as above. [Figure 7-3] Same as above. [Figure 7-4] Same as above. [Figure 7-5] Same as above. [Figure 7-6] Same as above. [Figure 8] Figure 8 is a bar graph showing the frequency of variants in lung, spleen, and blood samples for control and urethane-treated experimental animals according to one embodiment of this technology. [Figure 9] Figure 9 is a bar graph showing the mean minimum point mutant frequency across a group of tissue samples, according to one embodiment of this technology. [Figure 10A] Figure 10A is a box plot graph showing the calculated SNV MFs for double-strand sequencing by gene regions for the lung, spleen, and blood, according to one embodiment of the present technology, for the indicated treatment categories. [Figure 10B] Figure 10B is a scatter plot showing individual measurements of the aggregated data shown in Figure 10A, according to one embodiment of the present technology. [Figure 11] Figure 11 is a bar graph showing the mutation spectrum of urethane and vehicle controls in the tested tissue, measured by double-strand sequencing according to one embodiment of the present technology. [Figure 12-1] Figures 12A and 12B are graphs showing the mutation spectrum (i.e., trinucleotide spectrum) in the context of adjacent nucleotides for a vehicle control, according to one embodiment of the present technology. [Figure 12-2] Same as above. [Figure 12-3] Same as above. [Figure 12-4] Same as above. [Figure 13] Figure 13 shows single nucleotide variant (SNV) spectral strand bias in a urethane-treated sample according to one embodiment of the present technology. [Figure 14] Figure 14 is a graph showing early neobiotic clonal selection of variant allele fractions detected by double-strand sequencing according to one embodiment of this technology. [Figure 15A] Figure 15A is a graph showing SNVs plotted across genomic segments for exons captured from the Ras gene family, including human transgene loci, in a Tg-rasH2 mouse model according to one embodiment of the present technology. [Figure 15B] Figure 15B is a graph showing a single nucleotide variant aligned to exon 3 of a human HRAS transgene, according to one embodiment of this technology. [Figure 16A] Figure 16A is a graphical representation of sequencing data from a representative 400-base pair region of human HRAS in urethane-treated mouse lungs, using conventional DNA sequencing. [Figure 16B] Figure 16B is a graphical representation of sequencing data from a representative 400-base-pair portion of the human HRAS in urethane-treated mouse lung, using double-strand sequencing according to one embodiment of this technology. [Figure 17-1] Figures 17A to 17C are graphs showing the mutation spectrum (i.e., trinucleotide spectrum) in the context of adjacent nucleotides for signature 1 from COSMIC. [Figure 17-2] Same as above. [Figure 17-3] Same as above. [Figure 17-4] Same as above. [Figure 17-5] Same as above. [Figure 17-6] Same as above. [Figure 18] Figure 18 shows unsupervised hierarchical clustering of all 30 publicly available COSMIC signatures and four cohort spectra from Examples 1 and 2, according to one embodiment of the present technology. [Figure 19] Figure 19 is a schematic diagram of a network computer system for use with the method and / or kit disclosed herein for identifying mutagenic and / or nucleic acid damage events resulting from genotoxic substance exposure, according to one embodiment of the present technology. [Figure 20] Figure 20 is a flowchart illustrating a routine for providing double-stranded sequencing consensus sequence data according to one embodiment of the present technology. [Figure 21] Figure 21 is a flowchart illustrating a routine for detecting and identifying mutagenic events resulting from genotoxic substance exposure of a sample, according to one embodiment of this technology. [Figure 22] Figure 22 is a flowchart illustrating a routine for detecting and identifying DNA damage events resulting from genotoxic substance exposure of a sample, according to one embodiment of this technology. [Figure 23] Figure 23 is a flowchart illustrating a routine for detecting and identifying carcinogens or carcinogen exposure in a subject, according to one embodiment of this technology. [Modes for carrying out the invention]
[0031] Specific details of several embodiments of this technology are shown in Figures 1A to 1A. 23The following embodiments are described with reference to the above. These embodiments may include, for example, methods, systems, kits for evaluating genotoxicity. Some embodiments of the Art involve using double-strand sequencing to assess the potential genotoxicity of a certain drug (e.g., a chemical compound) or any other type of exposure (e.g., a radiation source) in an exposed subject, model organism, or model cell culture system. Other embodiments of the Art involve using double-strand sequencing to determine mutation signatures associated with genotoxic drugs. Further embodiments of the Art involve identifying one or more genotoxic drugs to which a subject may have been exposed by comparing the DNA mutation spectrum of a subject with the mutation spectrum of a known mutagenic compound. Further embodiments of the Art involve identifying one or more locations or environments to which a subject may have been exposed by comparing the DNA mutation spectrum of a subject from one or more cell types in one or more tissues with the mutation spectrum of a known environment or compound known to be present in a certain location or environment. Further embodiments of this technology relate to identifying subjects by comparing the subject's DNA mutation spectrum from one or more cell types in one or more tissues with the mutation spectrum of a known individual, or the mutation spectrum of a known location or environment to which that individual is known to have been exposed, or the mutation spectrum of compounds known to be present in such locations or environments. In certain embodiments, genotoxic substances may be evaluated for their carcinogenic potential. Further embodiments include identifying and evaluating the carcinogenic risk arising from mutagenic or non-mutagenic carcinogens by identifying clones having mutations that appear with cancer driver mutations.Further embodiments include identifying and assessing cancer risk arising from mutagenic or non-mutagenic carcinogens by identifying the emergence of clones that have mutations that substantially uniquely label the clone, but whose mutations are not considered cancer drivers (many known as “passenger” or “hitchhiker” mutations) (Salk and Horwitz, Sem Cancer Bio 2010 PMID:20951806). Other embodiments of the technology involve utilizing double-strand sequencing to detect and assess nucleic acid damage (in particular, DNA damage such as adducts) resulting from genotoxic substance exposure or other endogenous genotoxic processes (e.g., aging).
[0032] While many embodiments described herein relate to double-stranded sequencing, other sequencing methods capable of producing error-corrected sequence reads, in addition to those described herein, are within the scope of this technology. Furthermore, other embodiments of this technology may have different configurations, components, or procedures than those described herein. Therefore, those skilled in the art will see that this technology may include other embodiments with further elements, as shown in Figures 1A to 1A. 23 You will understand that other embodiments may be included, while referring to the following, which are shown and described below, and which do not include some of the features described.
[0033] definition To facilitate understanding of this disclosure, certain terms are defined below. Further definitions of the terms below and other terms are provided throughout this specification.
[0034] In this application, unless otherwise clearly stated in the context, the term “one (a)” may be understood to mean “at least one.” In this application, the term “or” may be understood to mean “and / or.” In this application, the terms “comprising” and “including” may be understood to encompass the listed component or process, whether indicated by itself or together with one or more further components or processes. Where a scope is presented herein, both ends are included. In this application, the term “comprise” and its variations, e.g., “comprising” and “comprises,” are not intended to exclude other additives, components, integers, or processes.
[0035] Approximately: When the term “approximately” is used herein in reference to a value, it refers to a value that is similar in terms of the referenced value. In general, a person skilled in the art will understand the relevant degree of variance that is encompassed in terms of that view. For example, in some embodiments, the term “approximately” may encompass values within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or a smaller range of values. For the variance of a single-digit integer value, if the step of a single number in either the positive or negative direction exceeds 25% of that value, it is generally accepted by a person skilled in the art that “approximately” includes at least integer values of 1, 2, 3, 4, or 5 in either the positive or negative direction, which may be greater than or less than 0, depending on the context. This non-limiting example is the hypothesis that, in some circumstances that would be obvious to a person skilled in the art, 3 cents could be considered approximately 5 cents.
[0036] Analogue: As used herein, the term “analogue” refers to a substance that shares one or more specific structural features, elements, components, or parts with a reference substance. Typically, an “analogue” exhibits significant structural similarity to the reference substance, for example, sharing a core or consensus structure, but also having differences in certain distinct ways. In some embodiments, an analogue is a substance that can be produced from a reference substance, for example, by the chemical manipulation of the reference substance. In some embodiments, an analogue is a substance that can be produced by the capabilities of a synthetic process that is substantially similar (for example, sharing several steps) to the synthetic process that produces the reference substance. In some embodiments, an analogue is produced or can be produced by the capabilities of a synthetic process different from the one used to produce the reference substance.
[0037] Biological Samples: As used herein, the terms “biological sample” or “sample” typically refer to a sample obtained from or derived from a biological source of interest (e.g., tissue or organism or cell culture), as described herein. In some embodiments, the source of interest includes organisms such as animals or humans. In other embodiments, the source of interest includes microorganisms such as bacteria, viruses, protozoa or fungi. In further embodiments, the source of interest may be synthetic tissue, organism, cell culture, nucleic acid or other material. In yet another embodiment, the source of interest may be a plant organism. In yet another embodiment, the sample may be an environmental sample, such as a water sample, soil sample, archaeological sample or other sample taken from a non-biological source. In other embodiments, the sample may be a multibiological sample (e.g., a mixed biological sample). In some embodiments, the biological sample is or includes biological tissue or biofluid. In some embodiments, the biological sample may be, or may contain, bone marrow, blood, blood cells, ascites, tissue samples, biopsy samples, or microneedle aspiration samples, body fluids containing cells, free suspension nucleic acids, protein-bound nucleic acids, riboprotein-bound nucleic acids, sputum, saliva, urine, cerebrospinal fluid, peritoneal fluid, pleural fluid, feces, lymph, gynecological body fluids, skin swabs, vaginal swabs, Pap smears, oral swabs, nasal swabs, staining or lavage solutions such as duct lavage or bronchopulmonary lavage, vaginal fluid, aspirates, scrapes, bone marrow specimens, tissue biopsy specimens, fetal tissue or body fluids, surgical specimens, feces, other body fluids, secretions, and / or excretions, and / or cells from these. In some embodiments, the biological sample is or contains cells obtained from an individual. In some embodiments, the obtained cells are or contain cells derived from the individual from which the sample is obtained. In some embodiments, cell derivatives such as organelles or vesicles or exosomes. In certain embodiments, the biological sample is a liquid biopsy obtained from a subject.In some embodiments, the sample is a “primary sample” obtained directly from the source of interest by any suitable means. For example, in some embodiments, the primary biological sample is obtained by a method selected from the group consisting of biopsy (e.g., microneedle aspiration or tissue biopsy), surgery, collection of bodily fluids (e.g., blood, lymph, feces, etc.). In some embodiments, as will be apparent from the context, the term “sample” refers to a preparation obtained by processing the primary sample (e.g., by removing one or more of its components and / or by adding one or more agents thereto), for example, by filtration using a semipermeable membrane. Such a “processed sample” may include, for example, nucleic acids or proteins extracted from the sample or obtained by performing techniques such as mRNA amplification or reverse transcription, isolation and / or purification of specific components on the primary sample.
[0038] Cancer: In one embodiment, a disease or disorder associated with a genotoxic substance is a “cancer” well known to those skilled in the art, generally characterized by abnormal dysregulation of cell growth, and may metastasize. Cancers detectable using one or more embodiments of the present technology include, in a non-limiting number of examples, prostate cancer (i.e., adenocarcinoma, small cell carcinoma), ovarian cancer (e.g., ovarian adenocarcinoma, serous carcinoma or embryonic carcinoma, yolk tumor, teratoma), liver cancer (e.g., HCC or hepatoma, angiosarcoma), plasmacytoma (e.g., multiple myeloma, plasmacytoma, plasmacytoma, amyloidosis, Valdenström macroglobulinemia), colorectal cancer (e.g., colonic adenocarcinoma, colonic mucinous adenocarcinoma, carcinoid, lymphoma and rectal adenocarcinoma, rectal squamous cell carcinoma), Skin cancer), leukemia (e.g., acute myeloid leukemia, acute lymphoblastic leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, acute myeloblastic leukemia, acute promyelocytic leukemia, acute myelomonocytic leukemia, acute monocytic leukemia, acute erythroleukemia and chronic leukemia, T-cell leukemia, Sézary syndrome, systemic mastocytosis, hairy cell leukemia, acute transformation of chronic myeloid leukemia), myelodysplastic syndrome, lymphoma (e.g., diffuse large B-cell lymphoma, cutaneous T-cell lymphoma, peripheral T-cell lymphoma, Hodgkin lymphoma, non-Hodgkin lymphoma) Lymphoma, follicular lymphoma, mantle cell lymphoma, MALT lymphoma, marginal zone cell lymphoma, Richter transformation, double-hit lymphoma, transplant-associated lymphoma, CNS lymphoma, extranodal lymphoma, HIV-associated lymphoma, endemic lymphoma, Burkitt lymphoma, transplant-associated lymphoproliferative neoplasms and lymphocytic lymphomas, etc.), cervical cancer (squamous cell cervical cancer, clear cell carcinoma, HPV-associated carcinoma, cervical sarcoma, etc.), esophageal cancer (esophageal squamous cell carcinoma, adenocarcinoma, certain grades of Barrett's esophagus, esophageal squamous cell carcinoma, adenocarcinoma, esophageal squamous cell carcinoma, esophageal squamous cell carcinoma, adenocarcinoma, certain grades of Barrett's esophagus, esophageal squamous cell carcinoma, esophageal squamous cell carcinoma, adenocar Pathological carcinoma), melanoma (cutaneous melanoma, uveal melanoma, peripheral melanoma, melanin-deficient melanoma, etc.), CNS tumors (e.g., oligodendroglioma, astrocytoma, glioblastoma multiforme, meningioma, schwannoma, craniopharyngioma, etc.), pancreatic cancer (e.g., adenocarcinoma, adenosquamous cell carcinoma, signet ring cell carcinoma, hepatoid carcinoma, colloid carcinoma, islet cell carcinoma, pancreatic neuroendocrine carcinoma, etc.), gastrointestinal stromal tumors, sarcomas (e.g., fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma, angiosarcoma, endothelioma sarcoma, lymphangiosarcoma, lymphoendothelioma sarcoma, leiomyosarcoma,Ewing's sarcoma and rhabdomyosarcoma, spindle cell tumors, etc.), breast cancer (e.g., inflammatory carcinoma, lobe carcinoma, ductal carcinoma, etc.), ER-positive cancer, HER-2-positive cancer, bladder cancer (squamous cell carcinoma of the bladder, small cell carcinoma of the bladder, urothelial carcinoma, etc.), head and neck cancer (e.g., squamous cell carcinoma of the head and neck, HPV-associated squamous cell carcinoma, nasopharyngeal carcinoma, etc.), lung cancer (e.g., non-small cell lung carcinoma, large cell carcinoma, bronchial carcinoma, squamous cell carcinoma, etc.) Epithelial cell carcinoma, small cell lung cancer, etc.), metastatic cancer, oral cancer, uterine cancer (leiomyosarcoma, leiomyoma, etc.), testicular cancer (e.g., seminoma, non-seminoma and embryonic carcinoma, yolk tumor, etc.), skin cancer (e.g., squamous cell carcinoma and basal cell carcinoma, Merkel cell carcinoma, melanoma, cutaneous T-cell lymphoma, etc.), thyroid cancer (e.g., papillary carcinoma, medullary carcinoma, anaplastic thyroid cancer, etc.), stomach cancer, carcinoma in situ, bone cancer, biliary tract cancer, eye cancer, laryngeal cancer, kidney cancer (e.g., renal cell carcinoma, Wilms' tumor, etc.), gastric cancer Cancer, blastoma (e.g., nephroblastoma, medulloblastoma, hemangioblastoma, neuroblastoma, retinoblastoma, etc.), myeloproliferative neoplasm (e.g., polycythemia vera, essential thrombocythemia, myelofibrosis, etc.), chordoma, synovial tumor, mesothelioma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, cystadenocarcinoma, cholangiocarcinoma, choriocarcinoma, epithelial carcinoma, ependymoma, pineal glandoma, acoustic neuroma, schwannoma, meningioma, pituitary adenoma, nerve sheath tumor, small intestine cancer, pheochromocytoma, small cell lung cancer, peritoneal mesothelioma , adenoma due to hyperparathyroidism, adrenal carcinoma, primary cancer of unknown origin, endocrine cancer, penile cancer, urethral cancer, melanoma of the skin or eye, gynecological tumors, solid tumors in children, or neoplasms of the central nervous system, primary mediastinal germ cell tumors, clonal hematopoiesis with undetermined potential, smoldering myeloma, monoclonal immunoglobulinemia of unknown significance, monoclonal B-cell lymphocytosis, low-grade cancer, clonal field defects, precancerous neoplasms, ureteral cancer, autoimmune-related cancers (i.e., ulcerative colitis, primary sclerosing cholangitis, celiac disease), cancers associated with hereditary predisposition (i.e., those with gene deficiencies such as BRCA1, BRCA2, TP53, PTEN, ATM), and various genetic syndromes (e.g., MEN1, MEN2 trisomy 21), and those resulting from exposure to chemicals in the uterus (i.e.,This includes clear cell carcinoma in the offspring of women exposed to diethylstilbestrol [DES].
[0039] Cancer Driver or Cancer Driver Gene: As used herein, “cancer driver” or “cancer driver gene” refers to a genetic lesion that, under appropriate circumstances, may cause cells to undergo malignant transformation. Such genes include tumor suppressor factors (e.g., TP53, BRCA1) that normally suppress malignant transformation but, when mutated in a particular manner, no longer normally suppress malignant transformation. Other driver genes may be oncogenes (e.g., KRAS, EGFR) that, when mutated in a particular manner, constitutively activate or acquire new properties that facilitate cell malignancy. Other mutations found in non-coding regions of the genome may also be cancer drivers. For example, a mutation in the promoter region of the telomerase gene (TERT) may cause overexpression of this gene and thus become a cancer driver. Certain transpositions (e.g., BCR-ABL fusion) may juxtapose one gene region with another gene region and promote tumorigenesis through mechanisms associated with overexpression, loss of repression, or chimeric fusion genes. In a broad sense, gene mutations (or epimutations) that confer a phenotype to a cell that gives it a proliferative, viable, or competitive advantage compared to other cells, or the ability to evolve more stably, can be considered driver mutations. This is in contrast to mutations that lack such characteristics, even if these mutations can occur within the same gene (i.e., synonymous mutations). When such mutations are identified within a tumor, they are generally called passenger mutations because they "hitchhike" along with clonal growth without making a meaningful contribution to growth. As those skilled in the art will recognize, the distinction between driver and passenger is not absolute and should not be interpreted as absolute. Some drivers function only in specific circumstances (e.g., specific tissues), while others may not function without other mutations or epimutations or other factors.
[0040] Control sample: As used herein, “control sample” means a sample isolated in the same manner as the sample being compared, except that the control sample has not been exposed to the drug, environment, or process being evaluated for potential genotoxicity.
[0041] Decision: Many of the methodologies described herein involve a “decision-making” step. Those skilled in the art will understand, upon reading this specification, that such “decisions” can be made or achieved by using any of the various technologies available to those skilled in the art, including, for example, certain technologies explicitly mentioned herein. In some embodiments, the decision-making involves the manipulation of a physical sample. In some embodiments, the decision-making involves the examination and / or manipulation of data or information, for example, using a computer or other processing unit tuned to perform the relevant analysis. In some embodiments, the decision-making involves receiving relevant information and / or material from a source. In some embodiments, the decision-making involves comparing one or more features of a sample or entity with comparable references.
[0042] Double-stranded sequencing (DS): As used herein, “double-stranded sequencing (DS)” in its broadest sense refers to a tag-based error correction method that achieves superior accuracy by comparing sequences from both strands of an individual DNA molecule.
[0043] Genotoxicity: As used herein, the term “genotoxicity” refers to the destructive nature of a drug or process (i.e., a genotoxic substance) that causes damage to genetic material (e.g., DNA, RNA). Polynucleotide damage, the formation of gene mutations, and / or disruption of normal nucleic acid structures resulting directly or indirectly from exposure to a genotoxic substance are aspects of genotoxicity. Subjects exposed to a genotoxic substance may develop a disease or disorder (e.g., cancer) immediately or several years later. In one embodiment, the technology aims to identify contributing events and / or factors (e.g., drugs, processes) that cause genotoxicity in a subject in order to prevent the onset of such disease or disorder, reduce the risk of such onset, and / or counteract its adverse effects. In other embodiments, the initiation of genotoxicity is by design, such as to create diversity within a gene library.
[0044] Genotoxic substances or genotoxic agents or factors: As used herein, the terms “genotoxic substance” or “genotoxic agent or factor” refer, for example, any chemical substance to which a nucleic acid source (e.g., a biological source, a subject) is exposed, and / or ingested, environmentally exposed, and / or any triggering event (endogenous precursor mutation) that causes polynucleotide damage, genomic mutation, or disruption of normal nucleic acid structure. In some embodiments, genotoxic substances have the ability to directly or indirectly (e.g., by triggering a mutagenic precursor), or both, cause the development of a disease or disorder in a subject. Genotoxic factors or agents detectable by this technology include, in non-limiting examples, chemical substances or mixtures of chemical substances (e.g., pharmaceuticals, industrial additives and waste by-products, petroleum distillates, heavy metals, cosmetics, household cleaning agents, airborne particles, food, manufacturing by-products, contaminants, plasticizers, detergents, etc.), and radiation (particle radiation, photons, or both), as well as / or physical forces produced by the natural environment or artificially (e.g., from devices) (e.g., magnetic fields, gravitational fields, accelerating forces, etc.). Genotoxic substances may further include liquid, solid, and / or aerosol formulations, and exposure may occur via any route of administration. Genotoxic agents or factors may be exogenous (e.g., exposure originates from outside the biological source), or in other cases, endogenous with respect to the biological source, or a combination of both. Exogenously derived agents or factors may become genotoxic when such exposure is treated endogenously. In other examples, a drug or factor may become genotoxic when combined with one or more other drugs or factors, and in some cases, they may have a synergistic effect.Further examples of genotoxic factors or agents include organisms that, upon exposure (e.g., through infection of the subject), can directly or indirectly cause nucleic acid damage in the subject, such as, non-limiting examples, schistosomiasis contributing to bladder cancer, HPV contributing to cervical or head and neck cancer, polyomavirus contributing to Merkel cell carcinoma, Helicobacter pylori contributing to gastric cancer, and chronic bacterial infection of skin wounds contributing to squamous cell carcinoma. Further examples of genotoxic agents or factors include organisms that can produce genotoxic agents (e.g., within themselves or secreted), such as, non-limiting examples, aflatoxin from Aspergillus flavus, or aristolochinic acid from the Aristocholia family of plants. Genotoxic factors or drugs detectable using various aspects of this technology may also include endogenous genotoxic substances that cannot be accurately quantified or experimentally controlled, and may include, for example, the effects of stress, inflammation, and therapeutic procedures (e.g., gene therapy, gene editing therapy, stem cell therapy, other cell therapies, pharmaceuticals, radiography, etc.) in non-limiting examples. Endogenous factors may also represent the sum of mutations in the subject's tissues and other genotoxic events that reflect the integrated effects of the subject's exposure.
[0045] Genotoxic substance-related illness or disorder: As used herein, the term “genotoxic substance-related illness or disorder” means any medical condition resulting from genomic mutations or other polynucleotide damage or transpositions in a subject, directly or indirectly caused by exposure to one or more genotoxic substances. Genotoxic substance-related illness or disorder may be cancer-related or not. In addition, polynucleotide damage / transposition or mutation may occur in germ cells or somatic cells. In several cases where germ cells are affected, genotoxic substance-related illness or disorder may manifest (or be at risk of manifesting in other circumstances) in subjects who are offspring of the exposed subject.
[0046] Sufficiently Genotoxic Agents: As used herein, the term “sufficiently genotoxic agent” refers to agents, factors, compounds, or processes identified by the Systems, Methods and Kits of this Technology as having a probability of about 50%, about 40%, about 30%, about 20%, about 10%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.5%, about 0.1%, about 0.01%, about 0.001%, about 0.0001%, about 0.00001%, about 0.000001%, etc. of causing nucleic acid damage or mutation above control background levels more than about 50%. In some embodiments, a sufficiently genotoxic agent refers to an agent, factor, compound, or process identified by the Systems, Methods and Kits of the Technology, which has a probability of causing a certain disease or disorder in subjects exposed to the genotoxic substance of approximately 50%, 40%, 30%, 20%, 10%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.01%, 0.001%, 0.0001%, 0.00001%, etc.
[0047] Growth Inhibition: As used herein, the term “growth inhibition” in cancer means reducing cell growth (e.g., tumor size, cancer cell division rate, etc.) in vivo or in vitro by, for example, about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, or about 99%, or more, as evidenced by the reduction in cell proliferation and / or cell size / mass of cells exposed to the treatment compared to cell proliferation and / or cell size growth in the untreated state. Growth inhibition may also be the result of a treatment that induces apoptosis in cells, induces necrosis in cells, slows the progression of the cell cycle, interferes with cell metabolism, induces cell lysis, or induces several other mechanisms that reduce cell proliferation and / or cell size growth.
[0048] Expression: As used herein, “expression” of a nucleic acid sequence means one or more of the following: (1) creation of an RNA template from a DNA sequence (e.g., by transcription); (2) processing of an RNA transcript (e.g., by splicing, editing, 5' cap formation and / or 3' end formation); (3) translation of RNA into a polypeptide or protein; and / or (4) post-translational modification of a polypeptide or protein.
[0049] Mechanism of Action: As used herein, the term “mechanism of action” refers to the biochemical processes that cause changes in nucleic acids after exposure to a genotoxic substance. In one embodiment, “mechanism of action” refers to the biochemical pathways and / or pathophysiological processes that follow genomic mutations or damage until the disease or disorder is fully developed. In another embodiment, “mechanism of action” includes biochemical pathways and / or physiological processes that occur in a living source after exposure to a genotoxic substance and cause genomic damage (e.g., pre-mutagenic lesions) or mutations. In yet another embodiment, the mechanism of action of a genotoxic agent or process may be inferred from one or more of the following: the nucleotide bases affected, the nucleotide changes introduced, the type of DNA damage introduced, the structural changes introduced, the contiguous nucleotide sequence context of the nucleotide(s) affected, the gene context or sequence(s) affected, the transcriptional state or affected region, the methylation state of the affected region, the protein-binding state or condensation state or chromosomal location of the region affected by exposure to the genotoxic substance.
[0050] Mutation: As used herein, the term “mutation” refers to a change in nucleic acid sequence or structure. Mutations in polynucleotide sequences may include, among complex multinucleotide changes, point mutations (e.g., single-nucleotide mutations), multi-nucleotide mutations, nucleotide deletions, sequence transpositions, nucleotide insertions, and DNA sequence duplications in a sample. Mutations may occur on both strands of a double-stranded DNA molecule as complementary base changes (i.e., true mutations) or as mutations on one strand but not the other (i.e., heteroduplexes), and may be repaired, disrupted, or incorrectly repaired / converted to true double-stranded mutations.
[0051] Variant Frequency: As used herein, the term “variant frequency” may also be called “variant frequency,” and refers to the number of unique variants detected per total number of double-stranded base pairs sequenced. In some embodiments, variant frequency is the frequency of variants in a specific gene only, or in a set of genes, or in a set of genomic targets. In some embodiments, variant frequency may refer to only a specific type of variant (for example, the frequency of A>T variants is calculated as the number of A>T variants per total number of A bases). The frequency with which a variant is introduced into a cellular or molecular assembly can vary, in particular, by genotoxic substances, by the amount or duration of exposure to genotoxic substances, by the age of the subject, over time, by the type of tissue or organism, by the region of the genome, by the type of variant, by the trinucleotide context, and by the heritable genetic background.
[0052] Mutation Signature: As used herein, the terms “mutation signature” and “mutation spectrum or multiple spectrums” refer to a characteristic combination of mutation types resulting from mutation introduction processes such as DNA replication failure, exposure to exogenous and endogenous genotoxic substances, defective DNA repair pathways, and DNA enzyme editing. In one embodiment, the mutation spectrum is constructed by computational pattern matching (e.g., unsupervised hierarchical mutation spectrum clustering).
[0053] Non-cancerous diseases: In another embodiment, a disease or disorder associated with a genotoxic substance is a non-cancerous disease. Instead, it is another type of disease or disorder caused by or resulting from a genomic mutation or damage. In non-limiting examples, such non-cancerous diseases or disorders that can be detected or predicted using one or more embodiments of the Art include diabetes, autoimmune diseases or disorders, infertility, neurodegeneration, progeria, cardiovascular disease, any disease associated with the treatment of another gene-mediated disease (i.e., chemotherapy-mediated neuropathy and chemotherapy-related renal failure such as cisplatin), Alzheimer's / dementia, obesity, heart disease, hypertension, arthritis, psychiatric disorders, other neurological disorders (neurofibromatosis), and multifactorial genetic disorders (e.g., environmentally triggered predispositions).
[0054] Nucleic acid: As used herein, in its broadest sense, refers to any compound and / or substance that is incorporated into or can be incorporated into an oligonucleotide chain. In some embodiments, nucleic acid is a compound and / or substance that is incorporated into or can be incorporated into an oligonucleotide chain via a phosphodiester bond. As will be apparent from the context, in some embodiments, “nucleic acid” refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides), and in some embodiments, “nucleic acid” refers to an oligonucleotide chain comprising individual nucleic acid residues. In some embodiments, “nucleic acid” is or contains RNA, and in some embodiments, “nucleic acid” is or contains DNA. In some embodiments, nucleic acid is one or more native nucleic acid residues, or comprises one or more native nucleic acid residues. In some embodiments, nucleic acid is one or more nucleic acid analogs, or comprises one or more nucleic acid analogs. In some embodiments, nucleic acid analogs differ from nucleic acids in that they do not utilize a phosphodiester backbone. For example, in some embodiments, the nucleic acid is one or more "peptide nucleic acids," contains one or more "peptide nucleic acids," or consists of one or more "peptide nucleic acids," where "peptide nucleic acids" are known in the art, have peptide bonds instead of phosphodiester bonds in their backbone, and are considered to be within the scope of the art. Alternatively, or in addition thereto, in some embodiments, the nucleic acid has one or more phosphorothioate bonds and / or 5'-N-phosphoramidite bonds instead of phosphodiester bonds. In some embodiments, the nucleic acid is one or more natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), contains one or more natural nucleosides, or consists of one or more natural nucleosides.In some embodiments, the nucleic acid is one or more nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C-5 propynylcytidine, C-5 propynyluridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, 0(6)-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof), or comprises one or more nucleoside analogs, or consists of one or more nucleoside analogs. In some embodiments, the nucleic acid contains one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) compared to the sugars in natural nucleic acids. In some embodiments, the nucleic acid has a nucleotide sequence encoding a functional gene product such as RNA or protein. In some embodiments, the nucleic acid contains one or more introns. In some embodiments, the nucleic acid is prepared by one or more of the following: isolation from natural sources, enzymatic synthesis by polymerization based on complementary templates (in vivo or in vitro), reproduction in recombinant cells or systems, and chemosynthesis. In some embodiments, the nucleic acid has a residue length of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 residues, or longer. In some embodiments, nucleic acids are partially or completely single-stranded, and in some embodiments, nucleic acids are partially or completely double-stranded.In some embodiments, the nucleic acid may be a branched chain having a secondary structure. In some embodiments, the nucleic acid has a nucleotide sequence comprising at least one element that encodes a polypeptide or is a complement to a polypeptide-encoding sequence. In some embodiments, the nucleic acid has enzymatic activity. In some embodiments, the nucleic acid performs a mechanical function, for example, in a ribonucleic acid-protein complex or transfer RNA.
[0055] Pharmaceutical composition or formulation: As used herein, the term “pharmaceutical composition” includes a pharmacologically effective amount of an active drug or active agent and a pharmaceutically acceptable carrier. In some examples, various aspects of this technique can be used to evaluate the genotoxicity of a pharmaceutical composition or formulation, or of the active drug or active agent contained therein.
[0056] Polynucleotide Damage: As used herein, the terms “polynucleotide damage” or “nucleic acid damage” refer to damage to the deoxyribonucleic acid (DNA) sequence (”DNA damage”) or ribonucleic acid (RNA damage”) of a subject, directly or indirectly caused by a genotoxic substance (e.g., by metabolites or the induction of processes that cause damage or are mutagenic). Damaged nucleic acids may cause the development of diseases or disorders associated with exposure to the toxic substance in the subject. In some embodiments, the detection of damaged nucleic acids in a subject may serve as an indicator of genotoxic substance exposure. Polynucleotide damage may further include chemical and / or physical modifications of DNA in cells.In some embodiments, damage can be caused by, in non-limiting examples, oxidation, alkylation, deamination, methylation, hydrolysis, hydroxylation, cleavage, intrachain crosslinking, interchain crosslinking, blunt end cleavage, sticky end double-strand cleavage, phosphorylation, dephosphorylation, SUMOylation, glycosylation, deglycosylation, putrescinylation, carboxylation, halogenation, formylation, single-strand gap, damage from heat, damage from drying, damage from UV exposure, damage from gamma radiation, damage from X-rays, damage from ionizing radiation, damage from non-ionizing radiation, damage from heavy ion radiation, damage from nuclear decay, damage from beta radiation, damage from alpha radiation, damage from neutron radiation, damage from proton radiation, damage from cosmic radiation, damage from high pH, damage from low pH, damage from active oxidative species, damage from free radicals, damage from peroxides, damage from hypochlorites, damage from tissue fixation such as formalin or formaldehyde, damage from reactive iron, damage from low ionic states, damage from high ionic states Damage from: unbuffered conditions, damage from nucleases, damage from environmental exposure, damage from fire, damage from mechanical stress, damage from enzymatic degradation, damage from microorganisms, damage from mechanical shearing during preparation, damage from enzymatic fragmentation during preparation, spontaneously occurring damage in vivo, damage occurring during nucleic acid extraction, damage occurring during sequencing library preparation, damage introduced by polymerase, damage introduced during nucleic acid repair, damage occurring during nucleic acid end tailing, damage occurring during nucleic acid ligation, damage occurring during sequencing, damage resulting from mechanical handling of DNA, damage occurring while passing through nanopores, damage occurring as part of aging in an organism, damage resulting from an individual's exposure to chemicals, damage caused by mutagens, damage caused by carcinogens, damage caused by chromosomal aberration-inducing substances, damage resulting from in vivo inflammatory damage due to oxygen exposure, damage resulting from one or more strand breaks, and any combination thereof, or including these.
[0057] Reference: When used herein, a standard or control against which a relative comparison is made is described. For example, in some embodiments, the drug, animal, individual, aggregate, sample, sequence, or value of interest is compared to a reference or control drug, animal, individual, aggregate, sample, sequence, or value, or a representative thereof, in a physical or computer database that may be located in a place or accessible remotely via electronic means. In some embodiments, the reference or control is tested and / or determined substantially concurrently with the test or determination of interest. In some embodiments, the reference or control is a historical reference or control, which may optionally be embodied in a tangible medium. Typically, as will be understood by those skilled in the art, the reference or control is determined or characterized under conditions or circumstances equivalent to those being evaluated. Those skilled in the art will understand when sufficient similarity exists to justify reliance on and / or comparison to a particular possible reference or control. "Reference sample" refers to a sample derived from a subject that is isolated in the same manner as the sample being compared and exposed to the same known amount of the same genotoxic drug, unlike a test subject. The reference sample subject may be genetically identical to or different from the test subject. In addition, the reference sample may be derived from several subjects exposed to the same known amount of genotoxic agent.
[0058] Safety threshold level: As used herein, the term “safety threshold level” refers to the amount (e.g., weight, volume, concentration, mass, molar amount, integral per unit time, etc.) to which a subject may be exposed before a genomic mutation that could cause disease development occurs. For example, the safety threshold level may be zero. In other examples, the level of genotoxic substance exposure may be an acceptable level. The tolerability of an acceptable exposure risk may vary depending on the subject, age, sex, type of tissue, the patient’s health status, and other risk-benefit considerations well known to those skilled in the art.
[0059] Safety threshold variant frequency: As used herein, the term “safety threshold variant frequency” refers to the tolerable rate of mutation caused by a genotoxic agent or process, below which the subject is presumed to have an acceptable risk of developing a disease or disorder associated with the genotoxic substance. The tolerability of the tolerable exposure risk and the resulting mutation rate may vary depending on the subject, age, sex, tissue type, and the patient’s health status.
[0060] Single Molecular Identifier (SMI): As used herein, the terms “single molecular identifier” or “SMI” (which may also be called “tag,” “barcode,” “molecular barcode,” or “unique molecular identifier” or “UMI”) refer to any substance (e.g., nucleotide sequence, nucleic acid molecular feature) that makes it possible to substantially distinguish individual molecules within a larger heterogeneous molecular assembly. In some embodiments, an SMI may be an externally applied SMI, or may include an externally applied SMI. In some embodiments, an externally applied SMI may be a denatured or semi-denatured sequence, or may include a denatured or semi-denatured sequence. In some embodiments, a substantially denatured SMI may also be known as a random unique molecular identifier (R-UMI). In some embodiments, an SMI may include a code (e.g., a nucleic acid sequence) from a pool of known codes. In some embodiments, a given SMI code is known as a defined unique molecular identifier (D-UMI). In some embodiments, an SMI may be an endogenous SMI, or may include an endogenous SMI. In some embodiments, the endogenous SMI may be, or include, information relating to a specific shear point of the target sequence, a feature associated with the end of an individual molecule containing the target sequence, or information relating to a specific sequence located at one end of an individual molecule, adjacent to one end of an individual molecule, or within a known distance from one end of an individual molecule. In some embodiments, the SMI may relate to sequence variations in the nucleic acid molecule caused by random or semi-random damage, chemical modification, enzymatic modification, or other modification of the nucleic acid molecule. In some embodiments, the modification described above may be methylcytosine deamination. In some embodiments, the modification described above may be accompanied by a site of nucleic acid nicks. In some embodiments, the SMI may include both exogenous and endogenous elements. In some embodiments, the SMI may include physically adjacent SMI elements. In some embodiments, multiple SMI elements may be spatially distinct within the molecule.In some embodiments, SMI may be non-nucleic acid. In some embodiments, SMI may include two or more different types of SMI information. Various embodiments of SMI are further disclosed in International Patent Publication WO2017 / 100441, which is incorporated herein by reference in whole.
[0061] Chain-Defining Elements (SDEs): As used herein, the term “chain-defining element” or “SDE” refers to any material that enables the identification of a particular strand of a double-stranded nucleic acid material and thus enables its distinction from other / complementary strands (e.g., any material that makes the image byproducts of each of the two single-stranded nucleic acids obtained from a target double-stranded nucleic acid substantially distinguishable from one another after sequencing or other nucleic acid interrogation). In some embodiments, the SDE may be one or more segments of substantially discomplementary sequences in the adapter sequence, or may include such segments. In certain embodiments, the segment of substantially discomplementary sequences in the adapter sequence may be provided by an adapter molecule including a Y-shape or “loop” shape. In other embodiments, the segment of substantially discomplementary sequences in the adapter sequence may form an unpaired “bubble” in the middle of adjacent complementary sequences in the adapter sequence. In other embodiments, the SDE may encompass nucleic acid modifications. In some embodiments, the SDE may include physical separation to a reaction compartment physically separated from the paired strands. In some embodiments, the SDE may include chemical modifications. In some embodiments, the SDE may be a modified nucleic acid. In some embodiments, the SDE may be related to sequence changes in the nucleic acid molecule caused by random or semi-random damage, chemical modification, enzymatic modification, or other modification of the nucleic acid molecule. In some embodiments, the above modification may be methylcytosine deamination. In some embodiments, the above modification may involve a site on the nucleic acid nick. Various embodiments of the SDE are further disclosed in International Patent Publication WO2017 / 100441, which is incorporated herein by reference in whole.
[0062] Subject: As used herein, the term “subject” means an organism, typically a mammal, e.g., human (including, in some embodiments, prenatal human forms), a non-human animal (e.g., mammals and non-mammals, but not limited to, non-human primates, horses, sheep, dogs, cattle, pigs, chickens, amphibians, reptiles, marine organisms (generally excluding sea monkeys), other model organisms, e.g., insects, flies, etc.), and a transgenic animal (e.g., a transgenic rodent). In some embodiments, the subject is exposed to a genotoxic substance or genotoxic factor or drug, or in other embodiments, the subject is exposed to a potential genotoxic substance. In some embodiments, the subject suffers from a disease, disorder or condition associated with the subject. In some embodiments, the subject suffers from a disease or disorder associated with a genotoxic substance. In some embodiments, the subject is susceptible to a certain disease, disorder or condition. In some embodiments, the subject exhibits one or more symptoms or characteristics of a certain disease, disorder or condition. In some embodiments, the subject exhibits no symptoms or characteristics of a particular disease, disorder, or condition. In some embodiments, the subject has one or more characteristics characteristic of susceptibility to, or risk of, a particular disease, disorder, or condition. In some embodiments, the subject exhibits symptoms or characteristics of a particular disease, disorder, or condition, and in some embodiments, such symptoms or characteristics are associated with a disease or disorder related to a genotoxic substance. In some embodiments, the subject is a patient. In some embodiments, the subject is an individual to whom diagnosis and / or treatment is performed and / or has been performed.In yet another embodiment, the test subject refers to any living biological source or other nucleic acid material, which may be exposed to genotoxic substances, and may include, for example, organisms, cells and / or tissues, e.g., those for in vivo testing, e.g., fungi, protists, bacteria, archaea, viruses, isolated cells in culture, intentionally (e.g., stem cell transplantation, organ transplantation) or unintentionally (i.e., fetal or maternal microchimerism) cells, or isolated nucleic acids or organelles (i.e., mitochondria, chloroplasts, free viral genomes, free plasmids, aptamers, ribozymes, or derivatives or precursors of nucleic acids (i.e., oligonucleotides, dinucleotide triphosphates, etc.)).
[0063] Substantially: As used herein, the term “substantially” refers to a qualitative state indicating the whole or nearly whole degree or extent of the characteristic or property of the subject. Those skilled in the art of biology will understand that it is rare, if any, for biological and chemical phenomena to terminate completely and / or progress toward complete termination, or to achieve or avoid absolute results. Thus, the term “substantially” is used herein to capture the inherent non-termination of many biological and chemical phenomena.
[0064] Therapeutic dose: As used herein, the terms “therapeutic dose,” “pharmacologically effective dose,” or simply “effective dose” refer to the amount of active drug or agent that produces the intended pharmacological, therapeutic, or preventive outcome. In some examples, various aspects of this technique can be used to evaluate or determine an effective dose of an active drug or agent (e.g., an active drug delivered to intentionally induce an event related to genotoxicity).
[0065] Trinucleotide or trinucleotide context: As used herein, the terms “trinucleotide” or “trinucleotide context” refer to a nucleotide in the context of the nucleotide bases immediately preceding and following it in a sequence (for example, a mononucleotide in a combination of three mononucleotides).
[0066] Trinucleotide spectrum or signature: Hereinafter, the term “trinucleotide signature” is used interchangeably with “trinucleotide spectrum,” “triple signature,” and “triple spectrum,” and refers to a variant signature within a trinucleotide context (e.g., one associated with exposure to a genotoxic substance). In one embodiment, a genotoxic substance may have a unique, semi-unique, and / or otherwise identifiable triple spectrum / signature.
[0067] Treatment: As used herein, the term “treatment” means the application or administration of a therapeutic agent to a subject, or to a tissue or cell line isolated from a subject, wherein the subject has a disorder, e.g., a disease, condition, or symptoms of a disease, or a predisposition to a disease, and the purpose is to cure, heal, reduce, alleviate, alter, relieve, improve, enhance, or influence that disease, symptoms of a disease, or predisposition to a disease. In one example, the disorder or disease / condition is a genotoxic disease or disorder. In another example, the disorder or disease / condition is not a genotoxic disease or disorder. In some examples, various aspects of this technique are used to evaluate the genotoxicity of a treatment or potential treatment.
[0068] Selected Embodiments of Double-Stranded Sequence Determination Method and Related Adapters and Reagents Double-strand sequencing is a method for generating error-corrected DNA sequences from double-stranded nucleic acid molecules, originally described in International Patent Publication WO2013 / 142389 and U.S. Patent Nos. 9,752,188 and WO 2017 / 100441, Schmitt et al., PNAS, 2012[1], Kennedy et al., PLOS Genetics, 2013[2], Kennedy et al., Nature Protocols, 2014[3], and Schmitt et al., Nature Methods, 2015[4]. The aforementioned patents, patent applications and publications are incorporated herein by reference in their entirety, respectively. As shown in Figures 1A to 1C, in certain embodiments of this technology, double-strand sequencing can be used to independently sequence both strands of individual DNA molecules in such a manner that derivative sequence reads can be recognized as originating from the same double-stranded nucleic acid parent molecule during large-scale parallel sequencing (MPS) (commonly known as next-generation sequencing (NGS)), and that they can be distinguished from each other as distinct entities after sequencing. The sequence reads obtained from each strand are then compared for the purpose of obtaining an error-corrected sequence of the original double-stranded nucleic acid molecule, known as a double-stranded consensus sequence (DCS). The double-strand sequencing process makes it possible to explicitly confirm that both strands of the original double-stranded nucleic acid molecule appear in the generated sequencing data used to form the DCS.
[0069] In certain embodiments, a method for incorporating a DS may include ligating one or more sequencing adapters to a target double-stranded nucleic acid molecule containing a first-strand target nucleic acid sequence and a second-strand target nucleic acid sequence to generate a double-stranded target nucleic acid complex (e.g., Figure 1A).
[0070] In various embodiments, the resulting target nucleic acid complex may contain at least one SMI sequence, which may be accompanied by externally applied denatured or semi-denatured sequences (e.g., randomized double-stranded tags shown in Figure 1A, identified as α and β in Figure 1A), endogenous information related to specific shear points of the target double-stranded nucleic acid molecule, or a combination thereof. The SMI can make the target nucleic acid molecule substantially distinguishable from several other molecules in the sequenced assembly, either alone or in combination with distinguishing the elements of the nucleic acid fragments they are ligated from. The substantially distinguishable features of the SMI element may be independently possessed by each single strand forming the double-stranded nucleic acid molecule, so that the derivative amplification products of each strand can be recognized after sequencing as originating from the same original substantially unique double-stranded nucleic acid molecule. In other embodiments, the SMI may contain further information and / or may be used in other ways in which the distinguishing function of such molecules is useful, such as those described in the publications cited above. In another embodiment, the SMI element may be incorporated after adapter ligation. In some embodiments, the SMI is essentially double-stranded. In other embodiments, the SMI is essentially single-stranded (for example, the SMI may be on a single-stranded portion of the adapter). In other embodiments, the SMI is essentially a combination of single-stranded and double-stranded components.
[0071] In some embodiments, each double-stranded target nucleic acid sequence complex may further include elements (e.g., SDEs) that make the amplified products of the two single-stranded nucleic acids forming the target double-stranded nucleic acid molecule substantially distinguishable from each other after sequencing. In one embodiment, the SDE may include an asymmetric primer site contained within the sequencing adapter, or in other configurations, the sequence asymmetry may be introduced within the adapter molecule rather than within the primer sequence, resulting in at least a portion of the nucleotide sequences of the first-stranded target nucleic acid sequence complex and the second strand of the target nucleic acid sequence complex being different from each other after amplification and sequencing. In other embodiments, the SMI may include another biochemical asymmetry between the two strands that is different from the canonical nucleotide sequence A, T, C, G, or U, but which translates to a difference of at least one canonical nucleotide sequence in the two amplified and sequenced molecules. In yet another embodiment, the SDE may be a means of physically separating the two strands before amplification, so that the derivative amplification products from the first-strand target nucleic acid sequence and the second-strand target nucleic acid sequence remain substantially physically isolated from each other for the purpose of maintaining distinction between these two sequences. Other such arrangements and methodologies for providing the SDE function that enables the distinction between the first and second strands described above, e.g., those described in the publication cited above, or other methods that achieve the described functional purpose may be utilized.
[0072] After creating a double-stranded target nucleic acid complex comprising at least one SMI and at least one SDE, or if one or both of these elements are subsequently introduced, the complex may be subjected to DNA amplification (e.g., using PCR) or any other biochemical method of DNA amplification (e.g., rolling circle amplification, multiple substitution amplification, isothermal amplification, bridge amplification, or surface-bound amplification), resulting in the production of one or more copies of the first-strand target nucleic acid sequence and one or more copies of the second-strand target nucleic acid sequence (e.g., Figure 1B). Then, the one or more amplified copies of the first-strand target nucleic acid molecule and the one or more amplified copies of the second target nucleic acid molecule are sequenced, preferably using a “next-generation” large-scale parallel DNA sequencing platform (e.g., Figure 1B).
[0073] Sequence reads derived from either the first-strand target nucleic acid molecule or the second-strand target nucleic acid molecule derived from the original double-stranded target nucleic acid molecule may be identified based on a common, related, substantially unique SMI and distinguished from the opposite-strand target nucleic acid molecule by an SDE. In some embodiments, the SMI may be a sequence based on a mathematically based error correction code (e.g., a Hamming code), thereby allowing certain amplification errors, sequencing errors, or SMI synthesis errors to be tolerated for the purpose of relating the sequence of the SMI sequence to the complementary strand of the original double-stranded molecule (e.g., a double-stranded nucleic acid molecule). For example, if the SMI includes a double-stranded exogenous SMI containing 15 base pairs of a completely denatured sequence of canonical DNA bases, then approximately 4 to the power of 15 = 1,073,741,824 SMI variants exist in the set of completely denatured SMIs. If two SMIs are recovered from sequencing data reads that differ by only one nucleotide within the SMI sequences from a set of 10,000 sampled SMIs, the probability of this occurring by chance can be mathematically calculated, and a determination can be made as to whether it is likely that this one-base-pair difference reflects one of the types of errors described above, and whether the SMI sequences have the fact that they originate from the same original double-stranded molecule. In some embodiments, if the SMIs are sequences applied at least partially externally, and their sequence variants are not completely denatured from one another, but are at least partially known sequences, the identity of the known sequences can be designed in such a way that one or more of the types described above does not convert the identity of one known SMI sequence to the identity of another SMI sequence, thereby reducing the probability that one SMI is misinterpreted as another. In some embodiments, this SMI design strategy includes the Hamming coding method or a derivative thereof. Once identified, one or more sequence reads generated from the first-strand target nucleic acid molecule are compared with one or more sequence reads generated from the second-strand target nucleic acid molecule to generate an error-corrected target nucleic acid molecule sequence (e.g., Figure 1C).For example, nucleotide positions where bases from both the first-strand target nucleic acid sequence and the second-strand target nucleic acid sequence match are considered true sequences, while nucleotide positions that do not match between these two strands are recognized as sites of potential technical errors, which may be disregarded, excluded, modified, or otherwise identified. In this way, an error-corrected sequence of the original double-stranded target nucleic acid molecule can be generated (shown in Figure 1C). In some embodiments, the respective sequencing reads produced from the first-strand target nucleic acid molecule and the second-strand target nucleic acid molecule may be grouped separately, and single-stranded consensus sequences may be created for each of these first and second strands. The single-stranded consensus sequences derived from the first-strand target nucleic acid molecule and the second-strand target nucleic acid molecule can then be compared to generate an error-corrected target nucleic acid molecule sequence (e.g., Figure 1C).
[0074] Alternatively, in some embodiments, the mismatched sequences between the two strands can be recognized as potentially biologically derived mismatches in the original double-stranded target nucleic acid molecule. Alternatively, in some embodiments, the mismatched sequences between the two strands can be recognized as potentially DNA synthesis-derived mismatches in the original double-stranded target nucleic acid molecule. Alternatively, in some embodiments, the mismatched sequences between the two strands can be recognized as potentially damaged or modified nucleotide bases present on one or both strands that have been converted into mismatches by an enzymatic process (e.g., DNA polymerase, DNA glycosylase, or another nucleic acid modifying enzyme or chemical process). In some embodiments, this latter finding can be used to infer the presence of nucleic acid damage or nucleotide modification prior to the enzymatic process or chemical treatment.
[0075] In some embodiments, according to aspects of the present technology, sequencing reads produced from the double-strand sequencing steps described herein can be further filtered to exclude sequencing reads from DNA damage molecules (e.g., damaged during storage, transport, during or after tissue or blood extraction, or during or after library preparation). For example, DNA repair enzymes, such as uracil-DNA glycosylase (UDG), formamidepyrimidine-DNA glycosylase (FPG), and 8-oxoguanine-DNA glycosylase (OGG1), can be used to exclude or correct DNA damage (e.g., in vitro or in vivo damage). These DNA repair enzymes are, for example, glycosylases that remove damaged bases from DNA. For example, UDG removes uracil resulting from cytosine deamination (caused by the spontaneous hydrolysis of cytosine), and FPG removes 8-oxoguanine (e.g., common DNA lesions resulting from reactive oxygen species). FPG also possesses lyase activity and can generate single-base gaps at abasic sites. Such abasic sites generally cannot be amplified by subsequent PCR, for example, because polymerase cannot copy the template. Therefore, the use of such DNA damage repair / removal enzymes can effectively remove damaged DNA that does not have true mutations but may not be detected as errors in other situations after sequencing and double-strand sequencing analysis. Errors caused by damaged bases can often be corrected by double-strand sequencing, but rarely, complementary errors can theoretically occur at the same location on both strands, thus reducing the probability of artifacts by reducing damage that increases errors. Furthermore, during library preparation, certain fragments of DNA to be sequenced may be single-stranded from their source or from processing steps (e.g., mechanical DNA shearing). These regions are typically converted to double-stranded DNA during the “end repair” process known in the art, where DNA polymerase and nucleoside substrates are added to the DNA sample and extend the 5' invaginated end.Mutagenic sites of DNA damage in the single-stranded portion of the DNA being copied (i.e., single-stranded 5' overhangs or internal single-stranded nicks or gaps at one or both ends of a DNA double-stranded molecule) may result in errors during the end blunting reaction, causing single-stranded mutations, synthesis errors, or nucleic acid damage sites to be converted into double-stranded forms that may be misinterpreted as true mutations in the final double-stranded consensus sequence, even if the true mutation was not actually present in the original double-stranded nucleic acid molecule. This situation is called a “false double-strand” and can be reduced or prevented by the use of enzymes that disrupt / repair such damage. In other embodiments, this occurrence can be reduced or eliminated by the use of strategies to disrupt or prevent the generation of single-stranded sites in the original double-stranded molecule (e.g., mechanical shearing where nicks or gaps may remain, or the use of specific enzymes used to fragment the original double-stranded nucleic acid material rather than certain other enzymes). In other embodiments, the use of a process to remove the single-stranded portion of the original double-stranded nucleic acid (e.g., a single-stranded specific nuclease such as S1 nuclease or mung bean nuclease) can be utilized for a similar purpose.
[0076] In a further embodiment, false mutations can be excluded by further filtering the sequencing reads produced from the double-strand sequencing process described herein, trimming the ends of the reads most likely to produce false double-strand artifacts. For example, DNA fragmentation can generate single-stranded portions at the ends of double-stranded molecules. These single-stranded portions may be blunted during end repair (e.g., by Klenow or T4 polymerase). In some cases, the polymerase makes copying errors in these repaired regions, leading to the generation of "false double-stranded molecules." These artifacts in library preparation may appear as true mutations when sequenced. Errors resulting from these end repair mechanisms can be excluded or reduced from post-sequencing analysis by trimming the ends of the sequencing reads to exclude mutations that may occur in higher-risk regions, thereby reducing the number of false mutations. In one embodiment, such trimming of sequencing reads can be achieved automatically (e.g., by a normal processing step). In another embodiment, the variant frequency can be evaluated for the fragment terminal region, and if a threshold level of mutation is observed in this fragment terminal region, the sequencing reads may be trimmed before creating double-stranded consensus sequence reads of the DNA fragment.
[0077] As a specific example, in some embodiments, methods are provided herein for creating error-corrected sequence reads of a double-stranded target nucleic acid substance, the method comprising the steps of ligating the double-stranded target nucleic acid substance to at least one adapter sequence to form an adapter-target nucleic acid substance complex, wherein the at least one adapter sequence comprises (a) a denatured or semi-denatured single-molecule identifier (SMI) sequence that uniquely labels each molecule of the double-stranded target nucleic acid substance, and (b) a first nucleotide adapter sequence that tags a first strand of the adapter-target nucleic acid substance complex and a second nucleotide adapter sequence that tags a second strand of the adapter-target nucleic acid substance complex, such that each strand of the adapter-target nucleic acid substance complex has a nucleotide sequence that is uniquely identifiable with respect to its complementary strand. The method may then further include the steps of amplifying each strand of the adapter-target nucleic acid substance complex to generate a plurality of first-strand adapter-target nucleic acid substance amplicons and a plurality of second-strand adapter-target nucleic acid substance amplicons. The method may further include a step of amplifying both the first and second strands described above to provide a first nucleic acid product and a second nucleic acid product. The method may also include a step of sequencing the first and second nucleic acid products, respectively, to generate a plurality of first strand sequence reads and a plurality of second strand sequence reads, and a step of confirming the presence of at least one first strand sequence read and at least one second strand sequence read. The method may further include creating error-corrected sequence reads of the double-stranded target nucleic acid substance by comparing at least one first strand sequence read and at least one second strand sequence read and disregarding non-matching nucleotide positions, or alternatively, removing the compared first strand sequence read and second strand sequence read if the compared first strand sequence read and second strand sequence read have one or more nucleotide positions in which they are non-complementary.
[0078] As further specific examples, in some embodiments, methods for identifying DNA variants from a sample are provided herein, the methods comprising the steps of: ligating both strands of a nucleic acid substance (e.g., a double-stranded target DNA molecule) to at least one asymmetric adapter molecule to form an adapter-target nucleic acid substance complex having a first nucleotide sequence associated with a first strand (e.g., the upper strand) of the double-stranded target DNA molecule and a second nucleotide sequence at least partially complementary to the first nucleotide associated with a second strand (e.g., the lower strand) of the double-stranded target DNA molecule; and amplifying each strand of the adapter-target nucleic acid substance to generate a distinct yet related set of amplified adapter-target nucleic acid products in each strand. The method may further include the steps of sequencing each of a plurality of first-strand adapter-target nucleic acid products and a plurality of second-strand adapter-target nucleic acid products; confirming the presence of at least one amplified sequence read from each strand of the adapter-target nucleic acid substance complex; and comparing at least one amplified sequence read obtained from the first strand with at least one amplified sequence read obtained from the second strand to form a consensus sequence read of a nucleic acid substance (e.g., a double-stranded target DNA molecule) having only nucleotide bases whose sequences on both strands of the nucleic acid substance (e.g., a double-stranded target DNA molecule) match, thereby identifying a variant occurring at a specific position in the consensus sequence read (e.g., compared to the reference sequence) as a true DNA variant.
[0079] In some embodiments, this specification provides a method for creating highly accurate consensus sequences from double-stranded nucleic acid material, the method comprising the steps of: tagging individual double-stranded DNA molecules with adapter molecules to form tagged DNA material, wherein each adapter molecule includes (a) a denatured or semi-denatured single-molecule identifier (SMI) that uniquely labels the double-stranded DNA molecule; and (b) for each tagged DNA molecule, first and second non-complementary nucleotide adapter sequences that distinguish the original upper strand from the original lower strand of each individual DNA molecule in the tagged DNA material; and creating a set of duplicates of the original upper strand of the tagged DNA molecule and a set of duplicates of the original lower strand of the tagged DNA molecule to form amplified DNA material. The method may further include the steps of creating a first single-stranded consensus sequence (SSCS) from the original upper strand duplication and a second single-stranded consensus sequence (SSCS) from the original lower strand duplication; comparing the first SSCS of the original upper strand with the second SSCS of the original lower strand; and creating a highly accurate consensus sequence having only complementary nucleotide bases for both the first SSCS of the original upper strand and the second SSCS of the original lower strand.
[0080] In a further embodiment, this specification provides a method for detecting and / or quantifying DNA damage from a sample containing a double-stranded target DNA molecule, the method comprising the steps of: ligating both strands of each double-stranded target DNA molecule to at least one asymmetric adapter molecule to form a plurality of adapter-target DNA complexes, wherein each adapter-target DNA complex has a first nucleotide sequence associated with a first strand of the double-stranded target DNA molecule and a second nucleotide sequence that is at least partially non-complementary to the first nucleotide associated with a second strand of the double-stranded target DNA molecule; and for each adapter-target DNA complex, the step of amplifying each strand of the adapter-target DNA complex to generate a distinct and still related set of amplified adapter-target DNA amplicons in each strand. The method may further include the steps of sequencing a plurality of first-strand adapter-target DNA amplicons and a plurality of second-strand adapter-target DNA amplicons, confirming the presence of at least one sequence read from each strand of the adapter-target DNA complex, and comparing at least one sequence read obtained from the first strand with at least one sequence read obtained from the second strand to detect and / or quantify nucleotide bases where the sequence read of one strand of the double-stranded DNA molecule does not match the sequence read of the other strand of the double-stranded DNA molecule (e.g., non-complementary bases), thereby enabling the detection and / or quantification of DNA damage sites. In some embodiments, the method may further include the steps of: creating a first single-strand consensus sequence (SSCS) from a first strand adapter-target DNA amprincon and a second single-strand consensus sequence (SSCS) from a second strand adapter-target DNA amprincon; comparing the first SSCS of the original first strand with the second SSCS of the original second strand; and identifying nucleotide bases in which the sequences of the first SSCS and the second SSCS are non-complementary to detect and / or quantify DNA damage associated with a double-stranded target DNA molecule in the sample.
[0081] Single molecular identifier sequence (SMI) According to various embodiments, the methods and compositions provided include one or more SMI sequences on each strand of a nucleic acid substance. The SMIs may be independently present on each single strand obtained from the double-stranded nucleic acid molecule so that the derivative amplification products of each strand can be recognized, after sequencing, as originating from the same original, substantially unique double-stranded nucleic acid molecule. In some embodiments, the SMIs may contain further information and / or may be used in other ways in which the function of distinguishing such molecules is useful, as recognized by those skilled in the art. In some embodiments, the SMI elements may be incorporated before, substantially simultaneously with, or after ligating the adapter sequences to the nucleic acid substance.
[0082] In some embodiments, the SMI sequence may contain at least one denatured or semi-denatured nucleic acid. In other embodiments, the SMI sequence may not be denatured. In some embodiments, the SMI may be a sequence associated with the fragment ends of a nucleic acid molecule (e.g., randomly or semi-randomly sheared ends of the ligated nucleic acid material) or near the fragment ends. In some embodiments, the exogenous sequence may be considered in combination with sequences corresponding to randomly or semi-randomly sheared ends of the ligated nucleic acid material (e.g., DNA) to obtain an SMI sequence that can distinguish single DNA molecules from one another. In some embodiments, the SMI sequence is part of an adapter sequence that is ligated to a double-stranded nucleic acid molecule. In certain embodiments, the adapter sequence containing the SMI sequence is double-stranded such that each strand of the double-stranded nucleic acid molecule contains the SMI after ligation to the adapter sequence. In another embodiment, the SMI sequence is single-stranded before or after ligation to a double-stranded nucleic acid molecule, and the complementary SMI sequence may be created by extending the opposite strand with DNA polymerase to obtain a complementary double-stranded SMI sequence. In yet another embodiment, the SMI sequence is located in a single-stranded portion of the adapter (e.g., an arm of an adapter having a Y-shape). In such embodiments, the SMI may facilitate the grouping of a family of sequence reads derived from the original strands of the double-stranded nucleic acid molecule and, in some cases, can confer a relationship between the original first and second strands of the double-stranded nucleic acid molecule (e.g., all or part of the SMI may be able to be associated by a lookup table). In some embodiments, when the first and second strands are labeled with different SMIs, sequence reads from the two original strands may be associated using one or more endogenous SMIs (fragment-specific features such as sequences associated with or near the fragment ends of a nucleic acid molecule), or by using further molecular tags shared by these two original strands (e.g., barcodes in the double-stranded portion of the adapter), or a combination thereof.In some embodiments, each SMI sequence may contain about 1 to about 30 nucleic acids (e.g., 1, 2, 3, 4, 5, 8, 10, 12, 14, 16, 18, 20, or more denatured or semi-denatured nucleic acids).
[0083] In some embodiments, the SMI can be ligated to one or both of the nucleic acid material and the adapter sequence. In some embodiments, the SMI may be ligated to at least one of the T-overhang, A-overhang, CG-overhang, dehydroxylated base, and blunt end of the nucleic acid material.
[0084] In some embodiments, the SMI sequence may be conceived (or designed) in combination with sequences corresponding to randomly or semi-randomly sheared ends of a nucleic acid substance (e.g., a ligated nucleic acid substance) in order to obtain an SMI sequence that can distinguish single nucleic acid molecules from one another.
[0085] In some embodiments, at least one SMI may be an endogenous SMI (e.g., an SMI related to a shear point (e.g., a fragment end)) using, for example, the shear point itself, or a predetermined number of nucleotides in the nucleic acid material immediately adjacent to the shear point [e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides from the shear point]. In some embodiments, at least one SMI may be an exogenous SMI (e.g., an SMI containing a sequence not found in the target nucleic acid material).
[0086] In some embodiments, the SMI may be an imaging portion (e.g., a portion that is optically detectable in fluorescence or other circumstances) or may include an imaging portion. In some embodiments, such an SMI enables detection and / or quantification without requiring an amplification step.
[0087] In some embodiments, the SMI element may include two or more distinct SMI elements positioned at different locations on the adapter-target nucleic acid complex.
[0088] Various embodiments of SMI are further disclosed in International Patent Publication WO2017 / 100441, which is incorporated herein by reference in its entirety.
[0089] Chain Definition Element (SDE) In some embodiments, each strand of the double-stranded nucleic acid material may further include elements that make the amplified products of the two single-stranded nucleic acids forming the target double-stranded nucleic acid material substantially distinguishable from each other after sequencing. In some embodiments, the SDE may be an asymmetric primer site contained within the sequencing adapter, or may include this site, or in other configurations, the sequence asymmetry may be introduced within the adapter sequence rather than within the primer sequence, resulting in at least a portion of the nucleotide sequences of the first-strand target nucleic acid sequence complex and the second strand of the target nucleic acid sequence complex being different from each other after amplification and sequencing. In other embodiments, the SDE may include another biochemical asymmetry between the two strands that is different from the canonical nucleotide sequence A, T, C, G, or U, but which translates to a difference of at least one canonical nucleotide sequence in the two amplified and sequenced molecules. In yet another embodiment, the SDE may be, or may include, a means for physically separating the two strands before amplification, so that the derivative amplification products from the first-strand target nucleic acid sequence and the second-strand target nucleic acid sequence remain substantially physically isolated from each other for the purpose of maintaining distinction between these two derivative amplification products. Other such arrangements or methodologies may be utilized to provide an SDE function that enables distinction between the first and second strands.
[0090] In some embodiments, the SDE may be capable of forming a loop (e.g., a hairpin loop). In some embodiments, the loop may contain at least one endonuclease recognition site. In some embodiments, the target nucleic acid complex may contain an endonuclease recognition site that facilitates a cleavage event within the loop. In some embodiments, the loop may contain a non-canonical nucleotide sequence. In some embodiments, the contained non-canonical nucleotide sequence may be recognizable by one or more enzymes that facilitate chain cleavage. In some embodiments, the contained non-canonical nucleotide sequence may be targeted by one or more chemical processes that facilitate chain cleavage within the loop. In some embodiments, the loop may contain a modified nucleic acid linker that can be targeted by one or more enzymatic, chemical, or physical processes that facilitate chain cleavage within the loop. In some embodiments, this modified linker is a photocleavable linker.
[0091] Various other molecular tools can function as SMI and SDE. Beyond shear point and DNA-based tagging, single-molecule compartmentation methods that maintain paired strands in physical proximity, or other non-nucleic acid tagging methods, can perform strand-related functions. Similarly, asymmetric chemical labeling of adapter strands in a manner that allows for physical separation of the adapter strand can serve as an SDE. A recently described variation of double-strand sequencing uses bisulfite conversion to convert naturally occurring strand asymmetry in the form of cytosine methylation into sequence differences that distinguish the two strands. While this embodiment has limitations in the types of mutations that can be detected, the concept of utilizing natural asymmetry is noteworthy in that it gives rise to sequencing techniques that can directly detect modified nucleotides. Various embodiments of SDE are further disclosed in International Patent Publication WO2017 / 100441, which is incorporated herein by reference in its entirety.
[0092] Adapters and adapter arrays In various configurations, adapter molecules comprising SMI (e.g., molecular barcode), SDE, primer sites, flow cell sequences, and / or other features are intended for use with many of the embodiments disclosed herein. In some embodiments, the provided adapter may be one or more sequences complementary or at least partially complementary to a PCR primer (e.g., primer site) having at least one of the properties of (1) high target specificity, (2) the ability to be multiplexed, and (3) exhibiting stable amplification with minimal bias, or may contain such sequences.
[0093] In some embodiments, the adapter molecule may be Y-shaped, U-shaped, or hairpin-shaped, may have a bubble (e.g., part of a non-complementary sequence), or may have other features. In other embodiments, the adapter molecule may include a Y-shaped, U-shaped, hairpin-shaped, or bubble. A particular adapter may include a modified nucleotide or atypical nucleotide, a restriction site, or other features for manipulating structure or function in vitro. The adapter molecule may ligate various nucleic acid substances having ends. For example, the adapter molecule may be suitable for ligating T-overhangs, A-overhangs, CG-overhangs, multiple nucleotide overhangs, dehydroxylated bases, or blunt ends of nucleic acid substances, where the molecular ends are dephosphorylated at the 5' of the target or otherwise blocked from conventional ligation. In other embodiments, the adapter molecule may include modifications on the 5' chain of the ligation site to prevent dephosphorylation or other ligation. In the latter two embodiments, such strategies may be useful in preventing dimerization of library fragments or adapter molecules.
[0094] The adapter sequence may mean a single-stranded sequence, a double-stranded sequence, a complementary sequence, a non-complementary sequence, a partially complementary sequence, an asymmetric sequence, a primer-binding sequence, a flow cell sequence, a ligation sequence, or any other sequence provided by the adapter molecule. In certain embodiments, the adapter sequence may mean a sequence used for amplification by complementarity to oligonucleotides.
[0095] In some embodiments, the methods and compositions provided include at least one adapter sequence (e.g., two adapter sequences, one at each of the 5' and 3' ends of the nucleic acid material). In some embodiments, the methods and compositions provided may include two or more adapter sequences (e.g., three, four, five, six, seven, eight, nine, ten, or more). In some embodiments, at least two of the adapter sequences are different from each other (e.g., by sequence). In some embodiments, each adapter sequence is different from the others (e.g., by sequence). In some embodiments, at least one adapter sequence is not at least partially complementary to at least one other adapter sequence (e.g., not complementary by at least one nucleotide).
[0096] In some embodiments, the adapter sequence includes at least one non-standard nucleotide. In some embodiments, the non-standard nucleotide is a debase site, uracil, tetrahydrofuran, 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A), 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G), deoxyinosine, 5'-nitroindole, 5-hydroxymethyl-2'-deoxycytidine, iso-cytosine, 5'-methyl-isocytosine, or isoguanosine, methylated nucleotide, RNA nucleotide, ribose nucleotide, 8-oxo-guanine, photocleavable linker, biotinylated nucleotide, desthiobiotin nucleotide, thiol-modified nucleotide Selected from octides, acridite-modified nucleotides, iso-dC, iso-dG, 2'-O-methyl nucleotides, inosine nucleotide locked nucleic acids, peptide nucleic acids, 5-methyldC, 5-bromodeoxyuridine, 2,6-diaminopurine, 2-aminopurine nucleotides, debasalized nucleotides, 5-nitroindole nucleotides, adenylated nucleotides, azide nucleotides, digoxigenin nucleotides, I-linkers, 5'-hexynyl-modified nucleotides, 5-octadiinyl dU, photocleavable spacers, non-photocleavable spacers, click chemistry compatible modified nucleotides, and any combination thereof.
[0097] In some embodiments, the adapter array includes a portion having magnetic properties (i.e., a magnetic portion). In some embodiments, this magnetic property is paramagnetic. In some embodiments, when a magnetic field is applied, the adapter array including the magnetic portion (e.g., nucleic acid material ligated to an adapter array including the magnetic portion) is substantially separated from an adapter array not including the magnetic portion (e.g., nucleic acid material ligated to an adapter array not including the magnetic portion).
[0098] In some embodiments, at least one adapter array is located 5' relative to the SMI. In some embodiments, at least one adapter array is located 3' relative to the SMI.
[0099] In some embodiments, the adapter sequence may be linked to at least one of the SMI and nucleic acid material via one or more linker domains. In some embodiments, the linker domains may consist of nucleotides. In some embodiments, the linker domains may include at least one modified nucleotide or non-nucleotide molecule (e.g., as described elsewhere in this disclosure). In some embodiments, the linker domains may be loops or may contain loops.
[0100] In some embodiments, the adapter sequences on one or both ends of each strand of the double-stranded nucleic acid material may further include one or more elements that provide an SDE. In some embodiments, the SDE may be an asymmetric primer site contained within the adapter sequence, or may consist of this site.
[0101] In some embodiments, the adapter sequence may be at least one SDE and at least one ligation domain (i.e., a domain modifiable to the activity of at least one ligase, e.g., a domain suitable for ligation to nucleic acid material by the activity of a ligase), or may include these. In some embodiments, from 5' to 3', the adapter sequence may be a primer binding site, an SDE, and a ligation domain, or may include these.
[0102] Various methods for synthesizing double-stranded sequencing adapters have already been described, for example, in U.S. Patent No. 9,752,188, International Patent Publication No. WO2017 / 100441, and International Patent Application No. PCT / US18 / 59908 (filed on November 8, 2018), all of which are incorporated herein by reference in their entirety.
[0103] Primer In several embodiments, one or more primers having at least one of the following characteristics—(1) high target specificity, (2) multiplexability, and (3) exhibiting stable amplification with minimal bias—are intended for use in various embodiments of the present art. Many conventional tests and commercial products design primer mixtures that meet these specific criteria for conventional PCR-CE. However, it should be noted that these primer mixtures are not always suitable for use with MPS. In fact, developing highly multiplexed primer mixtures can be a challenging and time-consuming process. For convenience, both Illumina and Promega have recently developed multiplex-compatible primer mixtures for the Illumina platform that exhibit stable and efficient amplification of a variety of standard and non-standard STR and SNP loci. These kits use PCR to amplify their target regions before sequencing, so that the 5' end of each read in paired-end sequencing data corresponds to the 5' end of the PCR primer used to amplify the DNA. In some embodiments, the methods and compositions provided include primers designed to ensure uniform amplification, which may involve varying reaction concentrations, melting temperatures, and minimizing secondary structure-in-primer / inter-primer interactions. Many techniques have been described for the optimization of highly multiplexed primers for MPS applications. In particular, these techniques are many known as ampliseq methods, as have been well described in the art.
[0104] amplification The methods and compositions provided utilize, in various embodiments, at least one amplification step of amplifying a nucleic acid substance (or a portion thereof, e.g., a specific target region or locus) and forming the amplified nucleic acid substance (e.g., several members of an amplicon product), or are used in at least one amplification step.
[0105] In some embodiments, amplifying a nucleic acid substance includes the step of amplifying the nucleic acid substance derived from the first and second nucleic acid strands of the original double-stranded nucleic acid substance, respectively, using at least one single-stranded oligonucleotide that is at least partially complementary to a sequence present in the first adapter sequence, such that the SMI sequence is at least partially maintained. The amplification step further includes using a second single-stranded oligonucleotide to amplify each of the desired strands, such that the second single-stranded oligonucleotide may (a) be at least partially complementary to a target sequence of interest, or (b) be at least partially complementary to a sequence present in the second adapter sequence, such that the at least one single-stranded oligonucleotide and the second single-stranded oligonucleotide are oriented in a manner that effectively amplifies the nucleic acid substance.
[0106] In some embodiments, amplifying nucleic acid material in a sample may involve amplifying the nucleic acid material in a “tube” (e.g., a PCR tube), in an emulsified droplet, a microchamber, and other examples or other known containers described above.
[0107] In some embodiments, at least one amplification step comprises at least one primer which is at least one non-standard nucleotide, or at least one primer which contains at least one non-standard nucleotide. In some embodiments, the non-standard nucleotide is selected from uracil, methylated nucleotides, RNA nucleotides, ribose nucleotides, 8-oxo-guanine, biotinylated nucleotides, locked nucleic acids, peptide nucleic acids, high Tm nucleic acid variants, allele-distinguishing nucleic acid variants, any other nucleotide or linker variant described elsewhere herein, or any combination thereof.
[0108] It is assumed that amplification reactions suitable for any application may fit into several embodiments, but specifically, in some embodiments, the amplification step may be polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple substitution amplification (MDA), isothermal amplification, Poloni amplification in emulsion, bridge amplification on a surface, surface amplification on a bead or on a hydrogel, and any combination thereof, or may include these.
[0109] In some embodiments, amplifying a nucleic acid substance involves using a single-stranded oligonucleotide that is at least partially complementary to the regions of the adapter sequence on the 5' and 3' ends of each strand of the nucleic acid substance. In some embodiments, amplifying a nucleic acid substance involves using at least one single-stranded oligonucleotide that is at least partially complementary to the target region or target sequence of interest (e.g., a genome sequence, mitochondrial sequence, plasmid sequence, synthetically produced target nucleic acid, etc.) and a single-stranded oligonucleotide that is at least partially complementary to a region of the adapter sequence (e.g., a primer site).
[0110] In general, stable amplification (e.g., PCR amplification) can be highly dependent on reaction conditions. Multiplex PCR may be sensitive to, for example, buffer composition, monovalent or divalent cation concentration, detergent concentration, crowding agent concentration (i.e., PEG, glycerol, etc.), primer concentration, primer Tm, primer design, primer GC content, primer modified nucleotide characteristics, and cycling conditions (i.e., temperature and extension time, as well as temperature change rate). Optimizing buffer conditions can be a difficult and time-consuming process. In some embodiments, the amplification reaction may use at least one of the buffer, primer pool concentration, and PCR conditions according to an already known amplification protocol. In some embodiments, a new amplification protocol may be created, and / or the amplification reaction may be optimized. Specifically, in some embodiments, a PCR optimization kit, e.g., a PCR optimization kit from Promega®, may be used, which contains many pre-formulated buffers partially optimized for various PCR amplifications (e.g., multiplex, real-time, GC-rich, and inhibitor-resistant amplification). These pre-formulated buffers contain various Mg 2+ The concentration, primer concentration, and primer pool ratio can be rapidly replenished. In addition, in some embodiments, various cycling conditions (e.g., thermal cycling) may be evaluated and / or used. When evaluating whether a particular embodiment is suitable for a particular desired application, one or more of the following may be evaluated, among other embodiments: specificity, allele coverage ratio to heterozygous loci, balance between loci, and depth. Measurement of amplification success may include DNA sequencing of the product, evaluation of the product by visualization of fragments performed after gel or capillary electrophoresis or HPLC or other size separation methods, lysis curve analysis using dyes or fluorescent probes that bind to double-stranded nucleic acids, mass spectrometry, or other methods known in the art.
[0111] Depending on the various embodiments, any of the following factors may influence the length of a particular amplification step (e.g., the number of cycles during the PCR reaction). For example, in some embodiments, the nucleic acid material provided may be degraded or otherwise suboptimal (e.g., degraded and / or contaminated). In such cases, a longer amplification step may help ensure that the desired product is amplified to an acceptable degree. In some embodiments, the amplification step may provide an average of 3 to 10 sequenced PCRs from each initial DNA molecule, while in other embodiments, only one copy each for the first and second strands is required. While we do not wish to be bound by any particular theory, too many or too few PCR copies can reduce assay efficiency and ultimately decrease depth. Generally, the number of nucleic acid (e.g., DNA) fragments used in the amplification (e.g., PCR) reaction is the first variable that can be adjusted, as it can determine the number of reads sharing the same SMI / barcode sequence.
[0112] Nucleic acid material kinds Depending on the various embodiments, any of the various nucleic acid materials may be used. In some embodiments, the nucleic acid material may include at least one modification to a polynucleotide within the canonical sugar-phosphate backbone. In some embodiments, the nucleic acid material may include at least one modification to any base within the nucleic acid material. For example, in some non-limiting embodiments, the nucleic acid material is at least one of double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA, peptide nucleic acid (PNA), locked nucleic acid (LNA), or comprises these.
[0113] qualification Depending on the various embodiments, the nucleic acid material may undergo one or more modifications before, substantially simultaneously with, or after any particular step, depending on the application in which the particular method or composition provided is used.
[0114] In some embodiments, the modification may be the repair of at least a portion of the nucleic acid material, or may include the repair of at least a portion of the nucleic acid material. It is assumed that a method suitable for any application of nucleic acid repair will fit into some embodiments, and for this reason, specific exemplary methods and compositions are described below and in the examples.
[0115] As a non-limiting example, in some embodiments, DNA repair enzymes, such as uracil-DNA glycosylase (UDG), formamidepyrimidine DNA glycosylase (FPG), and 8-oxoguanine DNA glycosylase (OGG1), can be used to repair DNA damage (e.g., in vitro DNA damage). As mentioned above, these DNA repair enzymes are glycosylases that remove damaged bases from DNA, for example. For example, UDG removes uracil resulting from cytosine deamination (caused by the spontaneous hydrolysis of cytosine), and FPG removes 8-oxoguanine (e.g., common DNA lesions resulting from reactive oxygen species). FPG also has lyase activity and can create a single base gap at abasic site. Such abasic site cannot be amplified by PCR, for example, because polymerase cannot copy the template. Therefore, the use of such DNA damage repair enzymes can effectively remove damaged DNA that does not have true mutations but may not be detected as an error in other contexts after sequencing and double-strand sequencing analysis.
[0116] As described above, in further embodiments, the sequencing reads produced from the processing steps described herein can be further filtered to remove false mutations by trimming the ends of the reads most likely to produce artifacts. For example, DNA fragmentation can generate single-stranded portions at the ends of double-stranded molecules. These single-stranded portions may be blunted during end repair (e.g., by Klenow). In some cases, polymerases make copying errors in these repaired regions, leading to the generation of "false double-stranded molecules." These artifacts may appear as true mutations when sequenced. Errors resulting from these end repair mechanisms can be excluded from post-sequencing analysis by trimming the ends of the sequencing reads to exclude the mutations that may occur, thereby reducing the number of false mutations. In some embodiments, such trimming of sequencing reads can be achieved automatically (e.g., by a normal processing step). In some embodiments, the variant frequency can be evaluated for the fragment terminal region, and if a threshold level of mutation is observed in this fragment terminal region, the sequencing reads may be trimmed before creating double-stranded consensus sequence reads of the DNA fragment.
[0117] The advanced error correction provided by strand comparison techniques in double-strand sequencing reduces sequencing errors of double-stranded nucleic acid molecules by several orders of magnitude compared to standard next-generation sequencing methods. This error reduction improves sequencing accuracy for almost all types of sequences, but may be particularly well suited to biochemically demanding sequences that are well known in the art to be particularly error-prone. One non-limiting example of such sequences is homopolymers or other microsatellite / short tandem repeats. Another non-limiting example of error-prone sequences that benefit from error correction in double-strand sequencing is, for example, molecules damaged by various chemical exposures that produce error-prone chemidodates during copying by one or more nucleotide polymerases, and also produce single-stranded DNA as molecular ends or as nicks and gaps. In a further embodiment, double-strand sequencing can also be used for the precise detection of a small number of sequence variants within an aggregate of double-stranded nucleic acid molecules. One non-limiting example of this application is the detection of a small number of cancer-derived DNA molecules among a large number of unmutated molecules from non-cancerous tissue in a subject. Another non-limiting application of double-strand sequencing for the detection of rare variants is the early detection of DNA damage resulting from exposure to genotoxic substances. A further non-limiting application of double-strand sequencing is the detection of mutations generated from genotoxic or non-genotoxic carcinogens by focusing on gene clones that emerge alongside driver mutations. Yet another non-limiting application for the precise detection of minority sequence variants is the creation of mutagenic signatures associated with genotoxic substances.
[0118] Identification and evaluation of genotoxicity This technology covers methods, systems, kits, etc., for evaluating genotoxicity. In particular, several embodiments of this technology cover the use of double-strand sequencing to assess the potential genotoxicity of a compound (e.g., a chemical compound) or other drug in a given biological source. For example, various embodiments of this technology include providing a double-strand sequencing method that enables the direct measurement of drug-induced mutations in any genomic context of any organism without requiring clonal selection. Further examples of this technology cover methods for detecting and evaluating in vivo genomic mutations using double-strand sequencing. Various aspects of this technology have many applications not only in preclinical and clinical drug safety testing but also in other industries as a whole. For example, this technology includes a method for detecting very low-frequency mutations that cause disease / disability development several years later, where the mutation occurs as a direct result of exposure to at least one genotoxic substance (e.g., radiation, carcinogen) and / or as a result of endogenous sources such as DNA polymerase errors, free radicals, and depurine reactions. Detection can be performed by testing subjects after recent exposure to genotoxic substances (e.g., within a few days of exposure) and identifying very low-frequency variants using double-strand sequencing. In certain cases, detected very low-frequency variants may be compared to variants known to cause specific diseases or disorders, including diseases / disorders that typically manifest many years after exposure (e.g., lung cancer 20 years after asbestos exposure). Thus, this technique provides a convenient method for identifying the presence of genotoxic substances and victims exposed to them in order to prevent future exposure and to take early medical action. This technique can also be used in various high-throughput screening methods to identify genotoxic substances in order to remove unsafe consumer products, pharmaceuticals, and other industrial / commercial / manufacturing by-products from the market or environment.
[0119] In certain embodiments, the effects of genotoxic substances (e.g., deletions, breaks, and / or transpositions) may cause cancer or other genotoxic substance-related diseases or disorders if the damage does not immediately lead to cell death. For example, nucleic acid damage may be sufficient for a subject to develop a genotoxic substance-related disease or disorder, and / or contribute to the activation or progression of another type of disease or disorder already present in the exposed subject. Regions prone to breaks (called fragile sites) may arise from genotoxic agents (e.g., chemicals such as pesticides or certain chemotherapeutic agents). Some chemicals have the ability to induce fragile sites in chromosomal regions where oncogenes exist, potentially leading to carcinogenic effects. Furthermore, occupational exposure to some mixtures of pesticides, manufacturing compounds, or other hazardous substances is positively correlated with increased genotoxic damage in exposed individuals. For example, investigating the potential for genotoxicity before human exposure is highly desirable for any potential genotoxic substance (e.g., potential drugs, cosmetics, consumer products, industrial / manufactured products or by-products, or other chemical compounds under study). Similarly, in embodiments where exposure to a genotoxic substance is suspected, if the genotoxic substance is identifiable, the subject may receive targeted therapeutic treatment, and / or the genotoxic substance may be removed to prevent future exposure to the subject or other humans.
[0120] The ability to detect the genotoxic effects of potential genotoxic agents or factors and to quantify the potentially obtained mutagenic processes in an efficient manner, both in terms of time and cost, is commercially and medically important. In certain cases, the ability to detect and quantify the mutagenic processes of potential genotoxic substances may be crucial for assessing cancer risk, identifying carcinogens, and predicting the effects of exposure in humans. However, current tools are slow, cumbersome, and / or have limited information they provide. As mentioned above, in vivo studies and mammalian reporter systems, e.g., BigBlue™ mice and rats, are currently available under U.S. Food and Drug Administration (FDA) regulations as effective genotoxicity metrics for determining the potential of compounds to cause DNA damage.
[0121] Figure 2A is a conceptual diagram illustrating various methodologies for evaluating in vivo mutagenesis of potential genotoxic substances (e.g., potential mutagens). In each scheme shown in Figure 2A, test subjects (e.g., BigBlue(trademark) mice, mouse model organisms, rat model organisms, etc.) are exposed to the potential genotoxic substance (e.g., the compound / drug / factor under investigation) using an appropriate route of administration. In one conventional scheme shown on the far left of Figure 2A, a long-term rodent carcinogenicity bioassay observes test animals for the development of neoplastic lesions over a long period (e.g., two years) during or after exposure to various doses of the test substance. Test animals may be administered by oral, dermal, or inhalation exposure, for example, depending on the type of expected human exposure. In the conventional scheme, administration typically lasts about two years. However, the administration parameters (e.g., duration of administration, route of administration, dose level, or parameters of other administration regimens) can be set according to the desired test protocol. Referring to the scheme on the left in Figure 2A, the health characteristics of specific animals are monitored throughout the study, but the primary assessment is made through a complete pathological analysis of the test animals' tissues and organs at the end of the study.
[0122] Another in vivo assay, shown in the central scheme of Figure 2A, utilizes transgenic rodents. After a suitable short-term drug regimen (e.g., over several days or weeks), the test animals are euthanized, desired tissues are collected, and DNA is extracted. Transgenic fragments are isolated from the extracted DNA, the resulting purified plasmids are encapsulated in phages, and E. coli are infected. A conventional transgenic plaque assay is performed to calculate the basic mutant frequency.
[0123] Both schemes described above are slow and provide very limited information regarding the genotoxicity (e.g., mutation introduction) of the potential genotoxic substance being tested. The possibility of directly measuring somatic mutations in a manner unrestricted by genomic loci, tissue, or organism is attractive, but the mutation frequency in normal cells (approximately 10) is not sufficient. -7 ~10 -8 ) has a considerably higher error rate (approximately 10 -3 Therefore, this is currently impossible with standard DNA sequencing methods.
[0124] Large-scale parallel sequencing offers the possibility of comprehensively examining the genome of any organism for the in vivo effects of mutagenic exposure; however, as mentioned above, conventional methods are too inaccurate to detect such mutations (which can occur at levels less than one in a million). For example, the error rate of approximately 0.1% in next-generation sequencing (NGS) generates background noise that interferes with the detection of rare variants and unique molecular profiles or signatures. Some common causes of errors in NGS platforms include PCR enzymes (occurring during amplification), sequencer reads, and DNA damage during processing (e.g., 8-oxoguanine, deaminated cytosine, desalting sites, etc.).
[0125] According to aspects of this technology, the steps of the double-strand sequencing method can produce highly accurate DNA sequencing reads, which can further provide detailed mutant frequencies (e.g., resolving mutations induced by genotoxic substances with a frequency of less than 1 in 1,000,000, providing mutation spectrum data for objectively characterizing different mutagenic processes and inferring mechanisms of action). For example, the scheme on the right shown in Figure 2A includes a method for rapidly detecting and evaluating the genotoxicity of potential genotoxic substances (e.g., potential mutagens) in the same test subjects as the scheme of the prior art, while also providing detailed information on mutant frequencies, mutation spectrum, and genomic context data. Furthermore, double-strand sequencing analysis can sensitively detect mutations at any locus in any tissue from any organism. For example, as shown in Figures 2A and 2B, the double-strand sequencing scheme can be used to evaluate in vitro mutagenesis of a test compound in cells grown in culture (e.g., human cells, rodent cells, mammalian cells, non-mammalian cells, etc.) (Figure 2B), and can be used to evaluate in vivo mutagenesis of a test compound in wild-type rodents (e.g., mice) (Figure 2C). For example, in one embodiment, the technique includes a step of a method comprising exposing a test organism (e.g., rodents, cells grown in culture) to a test compound (e.g., a potential genotoxic substance / mutagen) by an appropriate route of administration (e.g., oral, subcutaneous, topical, aerosol, intramuscular, etc.). In one embodiment, the test organism may be exposed to the test compound for a short period of time (e.g., a single dose, a few minutes, a few hours, less than 24 hours, a few days, 2-6 days, etc.), a moderate period of time (e.g., a few days, 3-12 days, about a week, about two weeks, about a month, about two months, about 3-6 months, etc.), or some other suitable period of time. If the test organism is an animal (e.g., a rodent), for example, as shown in Figure 1A (right-hand scheme) and Figure 1C, the animal may be euthanized and / or the desired tissue may be collected for DNA extraction.For example, in certain embodiments, the test animals may not be euthanized, and one or more blood samples (e.g., at the same or different time points after administration or exposure to the test substance) may be collected from the test animals for DNA extraction. In embodiments in which the animals are euthanized, one or more tissues of interest (e.g., liver, bone marrow, lungs, spleen, blood, etc.) may be collected for DNA extraction. If the test organism contains cells in culture (Figure 1B), all or some of the cells may be collected for DNA extraction.
[0126] Following DNA extraction from collected or acquired biological samples, a DNA library (e.g., a sequencing library) may be prepared. In one embodiment, a method for preparing a DNA library (or other nucleic acid sequencing library) may begin with labeling (e.g., tagging) fragmented double-stranded nucleic acid material (e.g., derived from a DNA sample) with a molecular barcode, in a manner similar to that described above, with respect to a double-stranded sequencing library construction protocol (e.g., as shown in Figure 1A). In some embodiments, the double-stranded nucleic acid material may be fragmented (e.g., cell-free DNA, damaged DNA, etc.). However, in other embodiments, various steps may include fragmentation of the nucleic acid material using mechanical shearing such as ultrasound or other DNA cutting methods (e.g., enzymatic digestion, spraying, etc.). The labeling of the fragmented double-stranded nucleic acid material may include end repair and 3'-dA tailing, and, if required for a particular application, may subsequently include ligating the double-stranded nucleic acid fragments with a double-stranded sequencing-appropriate adapter containing SMI (e.g., as shown in Figure 1A). In other embodiments, the SMI may be endogenous or a combination of exogenous and endogenous sequences, due to the inherently relevant information from both strands of the original nucleic acid molecule.
[0127] Following ligation of the adapter molecule to the double-stranded nucleic acid material, the method may be followed by amplification (e.g., PCR amplification, rolling circle amplification, multiple substitution amplification, isothermal amplification, bridge amplification, surface-bound amplification, etc.) (Figure 1B). In certain embodiments, for example, primers specific to one or more adapter sequences may be used to amplify each strand of nucleic acid material resulting in multiple copies of nucleic acid amplicons derived from each strand of the original double-stranded nucleic acid molecule, each amplicon retaining the SMI associated with the original (Figure 1B). After amplification and related steps for removing reaction byproducts, the target nucleic acid region (e.g., region of interest, locus, etc.) may optionally be enriched using hybridization-based capture, or, in another embodiment, enriched using multiplex PCR with adapter sequences and primers specific to the target nucleic acid region of interest (not shown).
[0128] Following the DNA library preparation and amplification steps, the double-stranded adapter-DNA complex may be sequenced using a standard sequencing method and a suitable large-scale parallel DNA sequencing platform (Figure 1B). After sequencing multiple copies of the first strand and multiple copies of the second strand, the sequencing data may be analyzed using a double-stranded sequencing method as described herein, thereby grouping sequencing reads that share the same exogenous (e.g., adapter sequence) and / or endogenous SMI derived from the first or second strand of the original double-stranded target nucleic acid molecule separately. In some embodiments, grouped sequencing reads from the first strand (e.g., “upper strand”) are used to form the first strand consensus sequence (e.g., single-stranded consensus sequence (SSCS)), and grouped sequencing reads from the second strand (e.g., “lower strand”) are used to form the second strand consensus sequence (e.g., SSCS). Referring again to Figure 1C, the first SSCS and the second SSCS may then be compared to create a double-stranded consensus sequence (DCS) with matching nucleotides between these two strands (for example, a variant or mutation is considered true if it appears in sequencing reads from both strands) (see, for example, Figure 1C). Similarly, in this comparing step, the locations of DCS where nucleotides do not match between the two strands may be further evaluated as potential DNA damage locations (e.g., damage caused by exposure to genotoxic substances).
[0129] Referring again to Figures 2A and 2C, according to aspects of this technique, double-strand sequencing analysis can be further used to accurately quantify the frequency of mutations induced across the genome. For example, aspects of this technique aim to create genotoxicity-related information captured in derivative sequence data, including, for example, mutation spectrum, trinucleotide mutation signatures, information on the functional consequences of specific mutations on proliferation and neoplasm selection, and comparisons with empirically derived genotoxicity-related information associated with known genotoxic substances (e.g., mutation spectrum, trinucleotide mutation signatures).
[0130] This technology further provides a method for detecting at least one genomic mutation in a subject as a result of exposure to a genotoxic substance, comprising: (1) providing a sample from a subject after exposure to a genotoxic substance, the sample containing multiple double-stranded DNA molecules; (2) creating multiple adapter-DNA molecules by ligating asymmetric adapter molecules to individual double-stranded DNA molecules; and (3) for each adapter-DNA molecule, (i) a set of copies of the original first strand of the adapter-DNA molecule and a set of copies of the original second strand of the adapter-DNA molecule. The method includes the steps of: (ii) creating a first strand sequence and a second strand sequence by sequencing a set of copies of the original first strand and a set of copies of the original second strand to provide a first strand sequence and a second strand sequence; (iii) comparing the first strand sequence and the second strand sequence to identify one or more correspondences between the first strand sequence and the second strand sequence; and (4) analyzing one or more correspondences in each adapter-DNA molecule to determine at least one of the following: variant frequency and variant spectrum that are indicative of a particular genotoxic substance, the type of genotoxic substance, and / or mechanism of action. In some embodiments, the variant spectrum is a triple variant spectrum. In other embodiments, determining the triple variant spectrum by analyzing one or more correspondences in each adapter-DNA molecule further includes creating a triple variant signature for a particular genotoxic substance. In certain embodiments, determining the variant frequency includes determining the frequency of the triple / trinucleotide context of the mutated base.
[0131] In some embodiments, the triple mutation signature and / or mutation spectrum are compared with empirically derived information associated with genotoxic substances to determine (if unknown) the type of genotoxic substance to which the subject was exposed, the mechanism of action of the genotoxic substance, the likelihood that the subject will develop a disease or disorder associated with the genotoxic substance, and / or other information associated with other genotoxic substances (e.g., based on similarity and / or difference). For example, a double-strand sequencing trinucleotide spectrum pattern obtained from a subject's exposure to a known or suspected genotoxic substance (e.g., a test genotoxic substance) may be compared with empirically derived trinucleotide spectrum patterns (e.g., those stored in a database) associated with exposure to other known genotoxic substances. In certain embodiments, the double-stranded sequencing trinucleotide spectrum pattern may be substantially similar to one or more empirically derived trinucleotide spectrum patterns, and as a result, the physician may be informed of the identity of the test genotoxic substance, the level of exposure to the test genotoxic substance, the mechanism of action of the test genotoxic substance, etc., based on the similarity to one or more empirically derived trinucleotide spectrum patterns.
[0132] Mutant frequency In some embodiments, the double-strand sequencing analysis step may identify variant frequencies associated with a particular genotoxic substance under various exposure conditions. For example, the variant frequency associated with exposure of a biological sample to a genotoxic substance may vary depending on various factors, including, but not limited to, the organism / subject, the subject's age, the type of genotoxic substance, the duration or level of exposure to the genotoxic substance, the type of tissue, the treatment group, the region of the genome (e.g., genomic locus), the type of mutation, the type of substitution, and the trinucleotide context. In some examples, the variant frequency is measured as the number of unique variants detected per double-stranded base pair being sequenced. In other embodiments, the variant frequency is the rate of new mutations in a single gene or organism over time.
[0133] Mutation spectrum In various embodiments, highly accurate (e.g., error-corrected) sequence reads created using double-strand sequencing may be further analyzed to create a mutation spectrum or signature for a specific genotoxic substance or potential genotoxic substance. In one embodiment, the mutation spectrum or signature includes a characteristic combination of mutation types resulting from the mutagenic process arising from exposure to the genotoxic substance. Such characteristic combinations may include information related to the type of mutation (e.g., changes to nucleic acid sequence or structure). For example, the mutation spectrum may include pattern information regarding the number, location, and context of point mutations (e.g., single nucleotide mutations) in the sample, nucleotide deletions, sequence transpositions, nucleotide insertions, and DNA sequence duplications. In some embodiments, the mutation spectrum may include information relevant to determining the mechanism of action that produces the determined mutation pattern. For example, the mutation spectrum may have the ability to determine whether the mutagenic process was directly caused by exposure to an exogenous or endogenous genotoxic substance, or whether the genotoxic substance exposure was indirectly triggered, in particular, by failure of DNA replication, defective DNA repair pathways, and disruption of DNA enzyme editing. In some embodiments, the mutation spectrum may be constructed by computational pattern matching (e.g., unsupervised hierarchical mutation spectrum clustering, non-negative matrix factorization, etc.).
[0134] Triple mutation spectrum / signature In one embodiment, highly accurate (e.g., error-corrected) sequence reads created using double-strand sequencing may be further analyzed to create a triplicate spectrum (also referred to herein as a trinucleotide spectrum or signature). For example, a mutation spectrum associated with the occurrence of a genotoxic substance and / or exposure to a genotoxic substance may be further analyzed to detect changes or mutations in a single nucleotide within a trinucleotide or trinucleotide context. While not theoretically bound, it is understood that exposure to a genotoxic substance or other processes (e.g., aging) may cause variability and / or specific damage to nucleic acids, depending on the trinucleotide context (e.g., nucleotide bases and the bases immediately surrounding them). In some embodiments, a genotoxic substance may have a unique, semi-unique, and / or other contextually identifiable triplicate spectrum / signature. For example, the trinucleotide spectrum of a first genotoxic substance may primarily contain C·G→A·T mutations and may further have a higher affinity for the CpG site. Such a trinucleotide spectrum is a proposed etiology similar to that primarily caused by exposure to tobacco, where benzo[α]pyrene and other polycyclic aromatic hydrocarbons are known mutagens. In another example, urethane is a genotoxic substance that produces DNA damage in a periodic pattern of T·A→A·T in a 5'-NTG-3' trinucleotide context. Thus, in some embodiments, determining the triple mutation spectrum would be advantageous for identifying genotoxic substance exposure in a subject, determining the genotoxicity of a potential genotoxic substance, and, among other benefits, identifying the mechanism of action of a genotoxic agent or factor.
[0135] Mechanism of action In some embodiments, highly accurate (e.g., error-corrected) sequence reads produced using double-strand sequencing can be used to infer one or more biochemical processes that cause detected changes in nucleic acids after exposure to a particular genotoxic substance. For example, in one embodiment, variant frequencies and mutation spectra (including trinucleotide spectra) produced using a double-strand sequencing method may be compared with empirically or deductively derived information regarding the types of mutations observed and the patterns and biochemical properties associated with the genomic location of gene mutations or DNA damage caused by genotoxic substance exposure. In embodiments in which biochemical pathways and / or pathophysiological processes following detected pregenomic mutations, mutations, or damages are identified, such information may be used to inform, in some embodiments, treatment options (e.g., therapy or prevention) for subjects exposed to genotoxic substances; or, in other embodiments, such information may be used to inform the feasibility of commercialization efforts (e.g., new drugs), cleanup efforts (e.g., environmental toxins or manufacturing by-products); or, in further embodiments, such information may be used to inform whether the tested compound, drug, or factor can be modified to eliminate and / or reduce the genotoxicity associated with that compound, drug, or factor.
[0136] Sources of nucleic acid substances for evaluating genotoxicity As described above, nucleic acid substances are expected to originate from a variety of sources. For example, in some embodiments, nucleic acid substances are provided from a sample derived from at least one subject (e.g., a human or animal subject) or other biological source. In some embodiments, nucleic acid substances are provided from a stored / reserved sample. In some embodiments, the sample may be blood, serum, sweat, saliva, cerebrospinal fluid, mucus, uterine lavage fluid, vaginal swabs, nasal swabs, oral swabs, tissue scraps, hair, fingerprints, urine, feces, vitreous fluid, peritoneal lavage fluid, sputum, bronchial lavage, oral lavage, pleural lavage, gastric lavage, gastric juice, bile, pancreatic duct lavage, bile duct lavage, common bile duct lavage, gallbladder juice, synovial fluid, infected wounds, uninfected wounds, archaeological samples, forensic samples, water samples, tissue samples, food samples, bioreactor samples. The samples are at least one of the following: plant samples, nail clippings, semen, prostatic fluid, fallopian tube lavage, cell-free nucleic acids, intracellular nucleic acids, metagenomics samples, lavage of implanted foreign bodies, nasal lavage, intestinal fluid, epithelial brushing, epithelial lavage, tissue biopsy, autopsy samples, necrotic samples, organ samples, human identification ampoules, artificially created nucleic acid samples, synthetic gene samples, nucleic acid data storage samples, tumor tissue, and any combination thereof, or comprising them. In other embodiments, the samples are at least one of the following: microorganisms, plant organisms, or any collected environmental samples (e.g., water, soil, archaeological samples, etc.), or comprising them. In certain examples further described herein, the nucleic acid material may originate from a biological source exposed to a genotoxic substance or a potentially genotoxic substance. In some examples, the genotoxic substance is a mutagen and / or carcinogen. In one example, the nucleic acid material is analyzed to determine whether the biological source from which the nucleic acid material originates has been exposed to a genotoxic substance.
[0137] For example, compared to other known or conventional toxicity assays such as the Ames test (e.g., a test for mutagenesis in bacteria), in vitro tests in mammalian cell cultures, transgenic rodent assays, Pig-a assays, and two-year in vivo bioassays, double-strand sequencing offers multiple evolutions. For instance, many conventional methods are limited to interrogation of reporter genes as a surrogate for useful information related to the genotoxicity of test drugs / factors (e.g., Ames test, in vitro mammalian cell culture, in vivo transgenic rodent assay), or tests in non-human sources (e.g., Ames test, transgenic rodent assay, Pig-a assay, two-year bioassay), which may require long periods of time to complete the very little information provided (e.g., two-year bioassay in wild-type rodents), or may be very expensive (e.g., transgenic rodent assay, two-year bioassay). In contrast to many of the shortcomings of conventional assays and techniques for screening test drugs / factors for genotoxicity, double-strand sequencing assays are widely deployable, economical, and may be suitable for early and late screening of test drugs / factors. They can be used to provide highly accurate data in a short period (e.g., less than two weeks), and can be used to screen in vitro and in vivo tested samples from any organism / biological source (i.e., including in vivo human samples in particular) or any tissue / organ, evaluate multiple loci, use natural genomes as genotoxicity reporters, and inform the mechanism of action of the determined genotoxic drugs / factors.
[0138] Reagent-based kits Aspects of this technology further encompass kits for carrying out various aspects of double-strand sequencing methods (also referred to herein as “DS kits”). In some embodiments, the kit may include various reagents along with instructions for carrying out one or more of the methods or process steps disclosed herein for nucleic acid extraction, nucleic acid library preparation, amplification (e.g., by PCR), and screening. In one embodiment, the kit may further include a computer program product for analyzing sequencing data (e.g., raw sequencing data, sequencing reads, etc.) (e.g., coded algorithms for execution on a computer, access code to a cloud-based server for executing one or more algorithms), for example, determining, in relation to a sample, variant frequency, variant spectrum, triple variant spectrum, comparison to the variant spectrum of known genotoxic substances, etc., according to aspects of this technology.
[0139] In some embodiments, the DS kit may include reagents or combinations of reagents suitable for performing various aspects of sample preparation (e.g., DNA extraction, DNA fragmentation), nucleic acid library preparation, amplification, and sequencing. For example, the DS kit may optionally include one or more DNA extraction reagents (e.g., buffers, columns, etc.) and / or tissue extraction reagents. Optionally, the DS kit may further include one or more reagents or tools for fragmenting double-stranded DNA by, for example, physical means (e.g., tubes, nebulizer units, etc., to facilitate acoustic shearing or sonication) or enzymatic means (e.g., enzymes and appropriate reaction enzymes for random or semi-random genome shearing). For example, the kit may include a DNA fragmentation reagent for enzymatically fragmenting double-stranded DNA, comprising one or more enzymes for targeted digestion (e.g., restriction endonucleases, CRISPR / Cas endonucleases and RNA guides, and / or other endonucleases), a double-stranded fragmentase cocktail, and single-stranded DNase enzymes (e.g., mung bean nucleases, S1 nucleases) primarily for making DNA fragments double-stranded and / or for disrupting single-stranded DNA, as well as appropriate buffers and solutions to facilitate such enzymatic reactions.
[0140] In one embodiment, the DS kit includes primers and adapters for preparing a nucleic acid sequence library from a sample suitable for performing steps in a double-strand sequencing process to create error-corrected (e.g., highly accurate) sequences of double-stranded nucleic acid molecules in the sample. For example, the kit may include at least one pool of adapter molecules containing single-molecule identifier (SMI) sequences, or a tool (e.g., single-stranded oligonucleotides) for the user to create them. In some embodiments, the pool of adapter molecules includes a suitable number of substantially unique SMI sequences, either alone or in combination with the unique features of the fragments to be ligated, so that multiple nucleic acid molecules in the sample can be substantially uniquely labeled after the attachment of the adapter molecules. Those skilled in the art of molecular tagging will understand that the “suitable” number of SMI sequences can vary by several orders of magnitude depending on various specific factors (e.g., input DNA, type of DNA fragmentation, average size of fragments, complexity and repeatability of sequences sequenced within the genome, etc.). Optionally, the adapter molecule further includes one or more PCR primer binding sites, one or more sequencing primer binding sites, or both. In another embodiment, the DS kit does not include an adapter molecule containing an SMI sequence or barcode, but instead includes a conventional adapter molecule (e.g., a Y-shaped sequencing adapter), and various method steps may utilize endogenous SMI to associate molecular sequence reads. In some embodiments, the adapter molecule is an index adapter and / or contains an index sequence.
[0141] In one embodiment, the DS kit includes a set of adapter molecules, each containing a non-complementary region and / or several other strand-defining elements (SDEs), or a tool (e.g., single-stranded oligonucleotides) for the user to create these. In another embodiment, the kit includes at least one set of adapter molecules, where at least a subset of the adapter molecules each contains at least one SMI and at least one SDE, or a tool for creating these. Further features of primers and adapters for preparing nucleic acid sequencing libraries from samples suitable for performing a dual sequencing process are described above and disclosed in U.S. Patent No. 9,752,188, International Patent Publication No. WO2017 / 100441 and International Patent Application No. PCT / US18 / 59908 (filed November 8, 2018), all of which are incorporated herein by reference in their entirety.
[0142] In addition, the kit may further include DNA quantification materials such as, for example, dyes that bind to DNA, such as SYBR® Green or SYBR® Gold (available from Thermo Fisher Scientific, Waltham, MA) for use with a Qubit fluorometer (e.g., available from Thermo Fisher Scientific, Waltham, MA), or PicoGreen® dye (e.g., available from Thermo Fisher Scientific, Waltham, MA) for use with a suitable fluorescence spectrometer. Other reagents suitable for DNA quantification on other platforms are also envisioned. Further embodiments include kits comprising one or more of the following: nucleic acid size selection reagents (e.g., Solid Phase Reversible Immobilization (SPRI) magnetic beads, gels, columns), columns for target DNA capture using bait / play hybridization, qPCR reagents (e.g., for copy number determination), and / or digital droplet PCR reagents. In some embodiments, the kit may optionally include one or more of the following: library preparation enzymes (ligases, polymerases, endonucleases, e.g., reverse transcriptase for RNA interrogation), dNTPs, buffers, capture reagents (e.g., beads, surfaces, coated tubes, columns, etc.), index primers, amplification primers (PCR primers), and sequencing primers. In some embodiments, the kit may include reagents for evaluating types of DNA damage, such as error-prone DNA polymerases and / or high-fidelity DNA polymerases. Further additives and reagents for PCR or ligation reactions under specific conditions (e.g., genomes / targets rich in high GC) are envisioned.
[0143] In one embodiment, the kit further includes reagents, such as DNA error-correcting enzymes that repair errors in DNA sequences that interfere with the polymerase chain reaction (PCR) process (as opposed to repairing disease-causing mutations). In a non-limiting example, the enzymes include one or more of the following: Examples include uracil-DNA glycosylase (UDG), formamidepyrimidine DNA glycosylase (FPG), 8-oxoguanine DNA glycosylase (OGG1), human aprine / apyrimidine endonuclease (APE1), endonuclease III (Endo III), endonuclease IV (Endo IV), endonuclease V (Endo V), endonuclease VIII (Endo VIII), N-glycosylase / AP lyase NEIL1 protein (hNEIL1), T7 endonuclease I (T7 Endo I), T4 pyrimidine dimer glycosylase (T4 PDG), human single-strand selective monofunctional uracil-DNA glycosylase (hSMUG1), and human alkyladenine DNA glycosylase (hAAG). These can be used to repair DNA damage (e.g., in vitro DNA damage). Some of these DNA repair enzymes are glycosylases that remove damaged bases from DNA. For example, UDG removes uracil resulting from cytosine deamination (caused by the spontaneous hydrolysis of cytosine), and FPG removes 8-oxo-guanine (a common DNA lesion resulting from reactive oxygen species, for example). FPG also possesses lyase activity and can create a single-base gap at abasic site. Such abasic site cannot be amplified by PCR, for example, because polymerase cannot copy the template. Thus, the use of such DNA damage repair enzymes and / or others known in the art listed here can effectively remove damaged DNA that does not have true mutations but may not be detected as an error in other contexts after sequencing and double-strand sequencing analysis.
[0144] The kit may further include appropriate controls, e.g., DNA amplification controls, nucleic acid (template) quantification controls, sequencing controls, nucleic acid molecules derived from a biological source exposed to a known genotoxic substance / mutagen (e.g., DNA extracted from a test animal exposed to a genotoxic substance or from cells grown in culture) and / or nucleic acid molecules derived from a biological source not exposed to a genotoxic substance / mutagen. In another embodiment, the control reagent may include nucleic acids that have been intentionally damaged and / or nucleic acids that have not been damaged or exposed by any damaging agent. In a further embodiment, the kit may also include one or more genotoxic and / or non-genotoxic agents (e.g., compounds) to be delivered in a controlled genotoxicity experiment, and optionally include a protocol for delivering such agents to a subject, tissue, cell, etc. Thus, the kit may include reagents (e.g., test compound, nucleic acid, control sequencing library) suitable for providing controls that give double-strand sequencing results (e.g., expected mutation spectrum / signature) that determine the credibility of the protocol for the test substance (e.g., test compound, potential genotoxic agent or factor, etc.). In one embodiment, the kit includes a container for carrying a test sample (e.g., a blood sample) for analysis to detect mutations in the test sample, so that the pattern and type of mutations indicate which genotoxic substance the test subject was exposed to. In another embodiment, the kit may include a nucleic acid contamination control standard (e.g., a hybridization capture probe having affinity for genomic regions in organisms different from the test organism or the test organism).
[0145] The kit may further include one or more other containers containing materials desirable from a commercial and user standpoint, including PCR and sequencing buffers, diluents, and subject sample extraction tools (e.g., syringes, swabs, etc.), and an accompanying document containing instructions for use. In addition, labels may be provided on the containers, for example, along with the instructions for use as described above, and / or these instructions and / or other information may also be included in the accompanying document included with the kit, and / or via a website address provided therein. The kit may also include laboratory equipment, such as sample tubes, plate sealers, microcentrifuge tube openers, labels, magnetic particle separators, foam inserts, ice packs, dry ice packs, and insulation.
[0146] The kit may further include a computer program product that is installable on an electronic computing device (e.g., a laptop / desktop computer, tablet, etc.) or accessible via a network (e.g., a remote server), wherein the computing device or remote server comprises one or more processors configured to execute instructions for performing operations including a double-stranded sequencing analysis step. For example, the processors may be configured to execute instructions for processing raw sequencing reads or unanalyzed sequencing reads to produce double-stranded sequencing data. In a further embodiment, the computer program product may include a database containing subject or sample records (e.g., information about a particular subject or sample or sample group and empirically derived information about known genotoxic substances). When executed on a computer, the computer program product is embodied in a non-temporary, computer-readable medium that performs the steps of the method disclosed herein (see, for example, Figures 19 and 20). The kit may also include instructions and / or access codes / passwords for accessing a remote server (including a cloud-based server) to upload and download data (e.g., sequencing data, reports, and other data), or software to be installed on the local device. All computational work may be performed on the remote server and accessed by the user / kit user via an internet connection or the like.
[0147] High-throughput genotoxicity screening This technology further includes a high-throughput screening scheme for evaluating the genotoxicity of suspected drugs or factors (e.g., compounds, chemicals, pharmaceuticals, manufactured products or by-products, food substances, environmental factors, etc.). In one embodiment, drugs / factors with unknown genotoxic effects may be screened to determine whether the test drug / factor has a genotoxic effect. In some embodiments, drugs / factors may be screened with a strong desire to eliminate the use of drugs / factors that have a genotoxic effect or exceed a threshold genotoxic effect. For example, drugs / factors that are mutagenic in a manner that could potentially cause diseases or disorders associated with genotoxic substances can be identified so that the drug / factor can be appropriately controlled, excluded, discarded, stored, etc. In some embodiments, carcinogenic drugs / factors can be identified using the high-throughput screening scheme described herein. In another embodiment, drugs / factors with unknown genotoxic effects may be screened with the intention of discovering drugs / factors that have a desired genotoxic effect, in particular a desired genotoxic effect on a target biological source. For example, biological samples from patients with a certain disease or disorder (e.g., cancer) may be used in high-throughput screening schemes to test multiple drugs / factors for desired genotoxic effects that can perturb or destroy cells (e.g., cancer cells). Such screenings can be performed for the discovery of new drugs / therapies and / or for targeted therapies for use in personalized medicine.
[0148] In some embodiments, high-throughput screening refers to screening multiple samples simultaneously and / or in a time-efficient manner. For example, testing a drug or factor for genotoxicity involves exposing a subject (e.g., a biological source) to the test drug or factor (e.g., treating, administering, applying, etc.). Thus, in a high-throughput screening scheme, arrays of biological sources / samples may be treated simultaneously with the same test drug / factor, or, in other embodiments, with multiple test drugs / factors. In specific examples, multiple biological samples (e.g., cells of a human or other organism grown in culture, tissue samples, blood or other bodily fluid samples, cells of a transgenic animal, human cells grown in xenografts, organoids of a living patient, feeder cells, etc.) may be exposed to the test drug / factor substantially simultaneously and under consistent conditions. High-throughput screening may also be used with a biofunctional chip, for example, a 10-organ chip containing blood or tissue samples from the same subject extracted from organs and tissues such as the endocrine system, skin, gastrointestinal tract, lungs, brain, heart, bone marrow, liver, kidneys, and spleen. The use of organ chips for high-throughput screening is well known in the art (e.g., Chan et al. [5]). In other embodiments, genetically modified cell lines (e.g., having missing or defective DNA repair pathways to make such cells more sensitive to the effects of mutagenic or genotoxic damage) may be incorporated into the high-throughput screening scheme.
[0149] In some embodiments, multiple biological samples may be the same or substantially similar (e.g., identical cell lines grown during culture, tissue samples from the same subject and / or the same type of tissue). In other embodiments, one or more of the multiple biological samples may be different. For example, a test drug / factor may be tested for genotoxic effects on different tissue / cell types from the same organism, different organisms, or a combination thereof. In certain examples, a suspected genotoxic drug or factor (e.g., a compound, a pharmaceutical, etc.) may be tested simultaneously on tissue samples from various organs of the same subject (e.g., a 10-organ chip). In some embodiments, high-throughput screening may involve testing multiple test drugs / factors simultaneously. Therefore, each tested sample may have different properties, whether intentionally altered or not (e.g., by cell type, by tissue type, by subject from which cells or tissues are extracted, by tumor, etc.), and / or may undergo different test regimens, which may differ from design to design (e.g., by test drug / factor, by dosage level, by exposure time, etc.), so that multiple samples can be efficiently screened using a high-throughput screening scheme in a manner that provides any desired information.
[0150] Once a biological sample has been exposed and / or the desired exposure regimen has been completed, cells / tissues may be collected from the sample, and DNA may be extracted for the purpose of using double-strand sequencing to evaluate the genotoxic / mutagenic effects of the test drug / factor on the DNA derived from each sample. In some embodiments, cell-free DNA (e.g., that released into a culture medium) may be collected from the biological sample for double-strand sequencing analysis. Further embodiments envisioned by this technique include high-throughput processing of DNA samples to produce double-strand sequencing data for evaluating DNA damage, known or suspected genotoxicity, or carcinogenicity of genotoxic substances.
[0151] The high-throughput screening processes described herein may include automation, for example, through the use of robots to perform one or more of the following steps: experimental treatment of biological samples, DNA extraction, library preparation, amplification (e.g., PCR), and / or DNA sequencing (e.g., using various techniques and devices for large-scale parallel sequencing). High-throughput screening allows for the parallel testing of multiple samples (i.e., different cell types from the same subject, or the same cell type from different subjects) to rapidly screen a large number of samples for genotoxicity-related mutations and / or DNA damage.
[0152] In one embodiment, each system consists of an array of wells, and a microplate containing one sample is moved through the system by robotic transport. In one example, the wells in the microplate may be filled by an automated liquid transport system, and sensors may be used to evaluate the sample in the microplate (e.g., often after incubation time). Laboratory automation software can be used to control all or part of the screening process, thereby ensuring accuracy within the process and reproducibility between processes.
[0153] Environmental / Exogenous Genotoxic Substances Aspects of this technology include, for example, evaluating the genotoxicity of environmental / exogenous drugs / factors by using either the in vivo or in vitro double-strand sequencing screening method described above. Further aspects of this technology include evaluating whether a subject / organism has been exposed to a genotoxic substance within a given environmental area. For example, biological samples (e.g., tissue, blood) may be collected from living organisms or from individuals exposed to an area suspected of contamination under different circumstances to determine, for example, whether a certain area is contaminated. In other embodiments, biological samples may be collected from organisms present in a larger area and evaluated as a screening process to accurately indicate the specific geographical location of a source of contamination by a genotoxic substance (e.g., industrial by-products leaked / released into a water system). Various methods described herein may be used to analyze biological samples (e.g., from a subject) exposed to an environmental area under investigation for the presence of a potential genotoxic substance. In another embodiment, various methods described herein may be used to analyze biological samples taken from a subject suspected of being exposed to a known genotoxic substance in a given environmental area (e.g., a geographical area, residential area, occupational environment, etc.). In accordance with aspects of this technology, the biological sample may be sourced from multiple organisms (e.g., marine organisms, mammals, filter feeders, sentinel organisms, etc.) or a specific species (e.g., a human sample).
[0154] Detectable environmental genotoxic substances also include exposure to mutagenic agents, such as, but not limited to, gamma radiation, X-rays, UV radiation, microwaves, electron emissions, toxic gases, toxic air particles (e.g., inhaled asbestos), and lakes, rivers, streams, groundwater, etc., contaminated with chemical compounds and / or pathogens. Further sources of exogenous genotoxic substances would include, for example, food products, cosmetics, household goods, healthcare products, cooking products and tools, and other manufactured consumer goods.
[0155] Double-strand sequencing results may also be used in combination with other methods to identify the presence of disease-causing contaminants, such as epidemiological studies that initially identify the location of cancer clusters. In some embodiments, the methods disclosed herein may be used to identify specific genotoxic substances that affected members of a cluster. From this data, the source of the genotoxic substance can be determined. In contrast to conventional investigation methods that have information on correlations that have been used conventionally to link a subject's disease or medical condition to a causative event (e.g., exposure to the environment or other exogenous mutagens or carcinogens), double-strand sequencing provides highly accurate, reproducible data (e.g., mutation spectrum and mechanism of action) that can be used to empirically determine the aforementioned causative event (e.g., exposure to a specific mutagen or carcinogen).
[0156] Endogenous genotoxic substances Aspects of this technology include, for example, evaluating the genotoxicity of endogenous drugs / factors (e.g., endogenous genotoxic substances or genotoxic processes) by using either the in vivo or in vitro double-strand sequencing screening method described above. Thus, aspects of this technology include evaluating whether a subject / organism has experienced an endogenous genotoxic substance or genotoxic process that caused DNA damage. For example, biological samples (e.g., tissue, blood) may be collected from a subject (e.g., a patient) to determine, for example, whether the subject has a disease or disorder associated with a genotoxic substance or is at risk of developing such a disease or disorder.
[0157] Endogenous factors may include, in non-limiting examples, biological developments that cause incorrect nucleotide integration, such as errors in DNA polymerase, free radicals, and depurine reactions. Endogenous factors may further include the onset of short- or long-term biological conditions that directly contribute to polynucleotide mutations associated with disease or disorder, such as stress, inflammation, activation of endogenous viruses, autoimmune diseases, environmental exposure, food choices (e.g., carcinogenic foods and beverages), smoking, natural genetic makeup, aging, and neurodegeneration. For example, if a subject has been exposed to high levels of stress for a long period, the subject may be tested by double-strand sequencing for any mutations correlated with stress-related cancers (e.g., leukemia, breast cancer, etc.).
[0158] Endogenous factors may also represent the sum of mutations and other genotoxic events in individual human tissues, reflecting the overall effects of an individual's exposure, and may not be accurately quantifiable or experimentally controlled.
[0159] Method for determining the frequency level of safe variants The level or amount of DNA damage resulting from exposure to a genotoxic substance may vary depending on various factors, including, in addition to various characteristics of the subject (e.g., health level, age, sex, genetic makeup, previous genotoxic substance exposure events), the effectiveness of the genotoxic substance in causing DNA damage (directly or indirectly), the dose or amount of exposure, the route or mode of exposure (e.g., ingestion, inhalation, transdermal absorption, intravenous, etc.), the duration of exposure (e.g., over time), and synergistic or antagonistic effects of other drugs or factors to which the subject is exposed. As described above, exposure to a genotoxic substance may cause polynucleic acid damage, which can be evaluated, for example, by the double-strand sequencing methods described herein, and can be associated with a unique, semi-unique, and / or otherwise contextually identifiable mutagenic spectrum or signature, which may include mutation patterns (e.g., mutation type, variant frequency, identifiable mutations in a trinucleotide context) that are sufficiently similar to mutation patterns associated with known diseases (e.g., distinct genomic mutations for breast cancer). Various aspects of this technology relate to methods for detecting and / or quantifying a variant frequency level that is considered safe, further comprising a method for detecting the safety threshold variant frequency of a genotoxic substance. If the variant frequency in a sample exceeds a safety level, it indicates that the subject has a significantly increased risk of developing the aforementioned disease over time.
[0160] The technology further includes a method for detecting and quantifying genomic mutations that occur in vivo in a subject after exposure to a mutagens, comprising: (1) double-strand sequencing one or more target double-stranded DNA molecules extracted from a subject exposed to a mutagens; (2) creating an error-corrected consensus sequence for the target double-stranded DNA molecules; (3) identifying the mutation spectrum of the target double-stranded DNA molecules; and (4) calculating the mutation frequency of the target double-stranded DNA molecules by calculating the number of unique mutations per sequenced double-stranded base pair. In one embodiment of step (3), the mutation spectrum is a unique profile of the sample and includes a "trunucleotide signature".
[0161] In one embodiment, steps (1) and (2) (a) ligating a double-stranded target nucleic acid molecule to at least one adapter molecule to form an adapter-target nucleic acid complex, wherein at least one adapter molecule includes i. a denatured or semi-denatured single-molecule identifier (SMI) sequence, either alone or in combination with a target nucleic acid shear site that uniquely labels the double-stranded target nucleic acid molecule, and ii. a nucleotide sequence that tags each strand of the adapter-target nucleic acid complex such that each strand of the adapter-target nucleic acid complex has a nucleotide sequence that is uniquely identifiable with respect to its complementary strand, and (b) the adapter-target nucleic acid complex This is achieved by (c) amplifying each strand to produce multiple first-strand adapter-target nucleic acid complex amplicons and multiple second-strand adapter-target nucleic acid complex amplicons; (d) sequencing the adapter-target nucleic acid complex amplicons to produce multiple first-strand sequence reads and multiple second-strand sequence reads; and (e) creating error-corrected sequence reads of the double-stranded target nucleic acid molecule by comparing at least one sequence from the multiple first-strand sequence reads with at least one sequence from the multiple second-strand sequence reads, without taking into account non-matching nucleotide positions (see U.S. Patent No. 9,752,188 B2 and WO 2017 / 100441).
[0162] Method for determining the safe threshold level of genotoxic substances The technology further includes experimental in vitro or in vivo methods for determining the safe level of exposure to a specific genotoxic substance by a subject (concentration amount, such as by weight or volume or mass or integral of units*time), and / or whether a certain compound or other agent (e.g., radio waves from a wireless device) is genotoxic at any exposure level. This determination may depend on first determining the safe threshold variant frequency level. In one embodiment, a sample of a control subject is tested for the genotoxic substance (or lack thereof) and compared to the genotoxic substance profile of a sample of an exposed subject (e.g., multiple mice, or multiple cells from the same subject, one set of which are control cells). The exposed subject ingests a specified predetermined exposure level of the suspected genotoxic substance to determine the safe exposure threshold level before mutations induced by the detected genotoxic substance occur that directly contribute to the development of disease.
[0163] In another embodiment, test subjects (e.g., laboratory animals, in vitro cells, etc.) are exposed to various doses over various periods, and from these conditions, a safe cutout level of genotoxic substance exposure is determined, including (1) at which exposure doses no polynucleotide mutations are observed and / or (2) at which exposure doses polynucleotide mutations are detected, but using the mutation levels found to infer levels of other compounds, which dose-equivalent levels do not cause cancer in the subjects and / or (3) regression analysis of dose-response curves and induced mutations of the genotoxic substance to infer a linear low-dose-response curve and / or (4) what the risk of a given health outcome in a population of subjects is, associated with the detected genotoxic substance frequency / signature.
[0164] Safe exposure threshold levels may be further determined for each species, for example, humans, dogs / cats, and horses. Safe exposure threshold levels may also be further determined for each route of exposure to genotoxic substances. For example, experiments using varying amounts of genotoxic substances may be tested using the double-strand sequencing methods disclosed herein to determine the amount (weight, volume, etc.) and / or frequency of oral, topical, or aerosol ingestion that can cause mutations and trispectral mutations associated with the development of specific diseases.
[0165] And / or, using the experimental double-strand sequencing methods disclosed herein, a threshold dose of genotoxic substance exposure can be determined based on time and / or temperature. For example, the amount (dose) of genotoxic substance absorbed percutaneously can be calculated using percutaneous absorption from a shower or bath in water containing a genotoxic substance, based on the duration of exposure, water temperature, and the concentration of the genotoxic substance in the water.
[0166] Error-corrected double-strand sequencing results, used to identify safety threshold levels for genotoxic substances, may be further combined with other safety threshold data (e.g., existing FDA and EPA levels, Hazardous Substances and Disease Registry levels, U.S. National Toxicology Program guidelines, OECD guidelines, Health Canada guidelines, European Agency guidelines, ILSI / HESI guidelines, etc.) to confirm or adjust established standards.
[0167] Detection and treatment methods While the onset of disease or disability may not be detectable by conventional testing and imaging techniques until many years (e.g., 20 years) have passed since exposure to a genotoxic substance, this technology provides a method for detecting indicators of genotoxic processes that may produce disease-causing mutations or precursors to disease-causing mutations within days, weeks, or months of exposure to a genotoxic substance, for the purpose of prophylactically treating subjects, actively screening subjects for disease (due to a higher risk level), and identifying the presence of genotoxic substances and eliminating them to prevent future exposure.
[0168] If a subject is exposed to a genotoxic substance beyond a threshold safety level, and / or if a subject is determined to have been potentially exposed to an unsafe level of the genotoxic substance (e.g., by a health department identifying hazardous exposure levels), the subject has a significantly increased risk of developing a disease or disorder associated with the genotoxic substance. The subject is then prophylactically treated with agents that block and / or detoxify the genotoxic substance, and / or exposure to the genotoxic substance is reduced or eliminated (e.g., remove the genotoxic substance from the environment or relocate the subject). In addition to or instead of this, the subject undergoes a series of timed diagnostic tests (e.g., blood tests for cancer detection) and / or imaging tests (e.g., CAT, MRI, PET, ultrasound, serum biomarker tests, etc.) to detect whether the subject is developing an early stage of the disease or disorder and is treated in the most effective manner during this time. As a non-limiting example, in the case of aflatoxin or aristolokinate exposure, subjects are likely to be instructed to undergo liver ultrasound every six months, which is the typical schedule for screening patients with chronic hepatitis C, another liver cancer inducer, for hepatocellular carcinoma. Once conventional diagnostic tests well known in the art detect the disease (e.g., cancer), treatment is then initiated (e.g., surgery, chemotherapy, immunotherapy, etc.).
[0169] Methods to provide preventative measures (i.e., to prevent the onset of the disease or reduce the risk of developing it), and / or methods to inhibit cancer growth, and / or methods to eradicate the cancer, would include treatment protocols well known to those skilled in the art and would be tailored to the type of genotoxic substance. While there are currently no treatments to reverse already induced mutations, treatment methods to help subjects excrete specific residual genotoxic substances (e.g., certain heavy metals by chelation) would reduce further genotoxicity.
[0170] In the case of mutagen-induced tumors (e.g., lung cancer in smokers, melanoma in severe UV exposure, oral cancer in tobacco users), the mutational burden in these tumors tends to be high, which explains the increased likelihood of generating more novel antigens and a significantly greater tendency to respond favorably to immunotherapy. Prophylactic administration of immunotherapy, such as with checkpoint inhibitors (i.e., PD1 and PDL1 inhibitors such as nivolumab, pembrolizumab, and atezolizumab, and CTLA4 inhibitors such as ipilizumab), is likely to eradicate tumors initially formed by the subject's immune system. Therefore, the therapeutic use of exposure signature identification, while requiring careful testing in the setting of formal clinical trials, represents a potential disease prevention strategy using prophylactic measures to predict future tumor responsiveness to immunotherapy.
[0171] The detection and treatment methods may further include methods for directly or inferentially determining the mechanism of action of genotoxic substances, which may be used to determine the appropriate course of treatment and / or to monitor drug-resistant variants (see Schmitt et al. [6]).
[0172] If a subject is diagnosed with or detected to have been exposed to at least one genotoxic substance, the subject may be administered a therapeutically effective amount of a pharmaceutical composition to prevent, delay, reduce the effects of, and / or eradicate, a disease or disorder associated with that genotoxic substance. The pharmaceutical composition comprises a therapeutically effective amount of a composition comprising an inhibitor or eradicator of a disease or disorder associated with a genotoxic substance and a pharmaceutically acceptable carrier or salt. The therapeutically effective amount also comprises a therapeutically effective non-toxic dose range of a composition comprising an inhibitor or eradicator of a disease or disorder associated with a genotoxic substance that is effective in producing the intended pharmacological, therapeutic, or preventive outcome.
[0173] Pharmaceutical compositions are formulated for and administered via routes of administration including oral, intravenous, intramuscular, subcutaneous, urethral, rectal, intrathecal, topical, oral, or parenteral administration. Pharmaceutical compositions may be mixed with conventional pharmaceutical carriers and excipients and used in the form of tablets, capsules, pills, liquids, intravenous solutions, beverages, and foods, and contain about 0.1% to about 99.9%, or about 1% to about 98%, or about 5% to about 95%, or about 10% to about 80%, or about 15% to about 60%, or about 20% to about 55% by weight or volume.
[0174] For oral administration, tablets, pills, and capsules may further contain conventional carriers such as binders (e.g., acacia gum, gelatin, polyvinylpyrrolidone, sorbitol, or tragacanth), fillers (e.g., calcium phosphate, glycine, lactose, corn starch, sorbitol, or sucrose), lubricants (e.g., magnesium stearate, polyethylene glycol, silica, or talc), disintegrants (e.g., potato starch), flavorings or colorings, or acceptable wetting agents. Oral liquid preparations may be formulated as aqueous or oily solutions, suspensions, emulsions, syrups, or elixirs, and may contain conventional additives such as suspending agents, emulsifiers, non-aqueous agents, preservatives, colorings, and flavorings.
[0175] For intravenous administration, the pharmaceutical composition can be dissolved or suspended in one of the commonly used intravenous infusion solutions and administered by drip infusion. Examples of intravenous infusion solutions include, but are not limited to, physiological saline or Ringer's solution.
[0176] Pharmaceutical compositions for parenteral administration may be in the form of aqueous or non-aqueous isotonic sterile injection solutions or suspensions. These solutions or suspensions can be prepared from sterile powders or granules containing one or more of the carriers mentioned for use in formulations for oral administration. The compounds can be dissolved in polyethylene glycol, propylene glycol, ethanol, corn oil, benzyl alcohol, sodium chloride and / or various buffers.
[0177] The therapeutic dose may further be calculated based on various factors such as the amount or duration of exposure to the genotoxic substance, the subject's age, weight, sex or race, the stage of the disease or disorder, and other methods well known to those skilled in the art. In one embodiment, a subject is tested when potential or suspected exposure to a genotoxic substance is detected, even if the exposure occurred many years ago. If the subject is diagnosed with exposure above a safety threshold level, the subject is administered the medicinal compound immediately or when symptoms appear. In all embodiments, the genotoxic substance is removed from the subject's environment if possible.
[0178] Experimental example The following chapters provide examples of methods for detecting and evaluating genomic in vivo mutations using double-strand sequencing and related reagents. The following examples illustrate the technique and are provided to assist those skilled in the art in manufacturing and using the technique. The examples are not intended to limit the scope of the technique in any other context.
[0179] Generally, to evaluate the effectiveness of DS for measuring in vivo mutagenesis, a series of mouse experiments were conducted using 62 samples to create 8.2 billion error-corrected bases, and the effects of three mutagens were examined on nine genes from five healthy tissues in two independent animal strains. Double-strand sequencing quantitatively demonstrated the increase in mutant frequency between treated animals to the extent that it varied depending on the specific mutagen, tissue type, and genomic locus, closely reflecting the increase in mutant frequency in representative transgenic rodent assays. In various cases, it was possible to distinguish samples by their treatment group based solely on objective mutation patterns. In some cases, mutagen susceptibility varied up to fourfold between different loci and was not constrained by theory, although the spectrum pattern suggested that this was the result of partially regionally distinct processes, which may include transcription and methylation. In various examples, trinucleotide mutation signatures between SNVs identified by DS at very low frequencies in animals treated with benzo[a]pyrene, a tobacco-related carcinogen, have been shown to be nearly identical to those found between clonal SNVs in the genomes of smoking-related lung cancers in publicly available databases. In some cases, DS was used to identify low-frequency oncogenic driver mutations that expanded clonally under selective pressure as little as four weeks after mutagen therapy. Thus, as demonstrated in the various examples described herein, DS can also be used to directly quantify genotoxic processes and real-time neoplasm evolution, and has diverse applications in mutation biology, toxicology, and cancer risk assessment.
[0180] Example 1 Application of Double-Stranded Sequencing for In Vivo Mutation Analysis in cII Transgenes and Endogenous Genes in BigBlue® Mice This chapter describes an example of directly measuring chemically induced mutations in both cII transgenes and native mouse genes used in the BigBlue® Transgenic Rodent (TGR) Mutation Assay using error-corrected next-generation sequencing (NGS). Currently, the TGR Mutation Assay detects rare cII variants via plaque formation. Standard NGS has a high error rate (10 in sequencing). 3 Because of an error rate of approximately 1 per base, it cannot be used for detecting low-frequency mutations. Error-corrected NGS or double-stranded sequencing have significantly lower error rates (approximately 1 / 10). 8 This enables the detection of extremely rare mutations (bases).
[0181] In this example, the mutant frequency (MF) and spectrum were evaluated in BigBlue® C57BL6 male mice exposed to control, N-ethyl-nitrosourea (ENU), and benzo[a]pyrene (B[a]P) using double-strand sequencing.
[0182] BigBlue® transgenic C57BL / 6 male mice were treated daily by forced oral administration of either a vehicle (olive oil) or B[a]P (50 mg / kg / day) from day 1 to 28, or ENU (40 mg / kg / day in pH 6 buffer) from day 1 to 3 (n=6). Tissue was collected on day 31 and frozen. Liver and bone marrow were analyzed for the mutants. DNA was isolated, and the cII mutant plaques were analyzed for the mutants using the RecoverEase and Transpack methods described by Agilent Technologies. Double-strand sequencing was used to sequence the cII and other endogenous genes for the mutations in the liver and bone marrow.
[0183] The genes to be evaluated and the criteria used to select them are as follows: (1) Polr1c (RNA polymerase), which is universally transcribed in all tissue types; (2) Rho (rhodopsin), which is not expressed in tissues other than the retina; (3) Hp (haptoglobin), which is highly expressed in the liver but hardly expressed elsewhere; (4) Ctnnb1 (β-catenin), the most common mutant gene in human hepatocellular carcinoma; and (5) a 360 bp transgenic reporter gene present in approximately 80 copies in CII:BigBlue® mice.
[0184] Figures 3A–3D are box plots showing the calculated mutant frequencies in the liver and bone marrow after the mutagen treatment described above, using double-strand sequencing (Figures 3A and 3B) and the BigBlue® cII plaque assay (Figures 3C and 3D). MF for double-strand sequencing was based on the total mutants per sequenced double-strand base pair (n=5 mice / group). MF for BigBlue® was calculated as the number of mutant plaques relative to the number of mutant plaque-forming units (n=6 mice / group). As shown, MF measured by double-strand sequencing and MF measured by the conventional BigBlue® cII plaque assay showed similar responses to both mutagens. Bone marrow, containing more rapidly dividing cells, showed higher MF than the liver using both methods.
[0185] Figure 3E shows the relative cII mutant doubling rates in the transgenic rodent assay compared to double-strand sequencing. As described above, MF in the plaque assay is calculated by dividing the number of phenotypically active mutant plaques observed on the selection plate by the total number of plaques formed on the acceptance plate. In the double-strand sequencing assay, MF is calculated by dividing the number of observed mutant base pairs by the total number of base pairs sequenced within the 297 BP cII transgene interval. Despite differences in derivative measurement, the correlation between the double-strand sequencing assay and the BigBlue® cII plaque assay is strong across tissue and mutagen treatments.
[0186] Figure 3F shows the proportion of SNVs within the cII gene in individually isolated mutant plaques produced from BigBlue® mouse tissue and double-strand sequencing of cII gDNA from BigBlue® mouse tissue. SNVs are specified using pyrimidine as the reference. Double-strand sequencing provides the same mutation spectrum from each treatment group as achieved by manual collection of 3510 plaques (p-values > 0.999 for all three using the chi-squared test). The proportion was calculated by dividing the total number of observed SNVs by the number of observed reference bases within the cII interval and normalizing to 1.
[0187] Figure 3G shows the distribution of all mutations identified by direct double-strand sequencing of cII across all BigBlue® tissue types and treatment groups, based on codon location and functional consequences. Figure 3H shows the distribution data for mutations identified among individually collected mutant plaques. Referring to Figures 3G and 3H together, direct double-strand sequencing (Figure 3G) identifies mutations along all genes that produce all effect groups, while mutations from collected mutant plaques (Figure 3H) do not contain synonymous variants and mutations at the non-critical C-terminus and N-terminus of this protein. Although not theoretically constrained, synonymous variants and mutations at the non-critical C-terminus and N-terminus of this protein are not thought to disrupt gene function necessary for selective growth and scoring during plaque assays.
[0188] Figure 4 is a bar graph showing the consistency of MF (microfluid motility) measured by double-strand sequencing within each treatment group. MF, aggregated across all genes, was measured in the liver and bone marrow by double-strand sequencing. The number of unique variants was lower in vehicle control animals (1–13 variants / 1.4 billion base pairs) compared to mutagen-exposed mice (118 variants / 2.6 billion base pairs). Inter-animal MF within a given group was reproducible under all treatment conditions, and the low number of variants in control animals (1–13) highlights the need for deeper sequencing to obtain stable estimates of MF.
[0189] Figures 5A and 5B are bar graphs showing the MF of endogenous genes, measured by double-strand sequencing, compared to cII transgenes in the liver (Figure 5A) and bone marrow (Figure 5B). Each gene (approximately 3–6 kb) was sequenced at a depth of approximately 5000-fold, while the cII gene (approximately 350 bp × 80 copies per genome) was sequenced at a depth of approximately 100K–300K. Mutant frequencies were calculated as described above for Figures 3A–3D. As shown, endogenous genes show a similar increase in MF as cII transgenes. Double-strand sequencing shows that MF is higher in bone marrow than in the liver. While not theoretically constrained, the higher rate of cell division in bone marrow would explain the detection of higher MF levels for both mutagens tested. Furthermore, the differences in the response of endogenous genes shown in Figures 5A and 5B would likely be related to differences in the transcriptional state or chromosomal structure of the endogenous genes.
[0190] Figure 5C is a box plot graph showing the calculated SNV MF for double-strand sequencing by gene regions for liver and bone marrow, and Figure 5D is a scatter plot showing the individual measurements of the aggregated data shown in Figure 5C. The scatter points represent individual measurements, and these measurements are assigned a 95% CI. The box plot in Figure 5C shows all four quartiles for all data points related to its tissue and treatment category. The Y-axis scale is linear and has a magnitude of 10 -7 As shown in Figure 5C, this box plot summarizes the SNV mutation frequencies in liver and bone marrow tissue across four endogenous genes and cII transgenes in the Big Blue® mouse model shown in Figure 5D. The degree of mutagenesis is influenced by the specific mutagens, tissue type, and gene locus.
[0191] Figure 6 is a bar graph showing the mutation spectrum of each test mutagens (e.g., treatment) in the test tissue, as measured by double-strand sequencing. Referring to Figure 6, each mutation segment, aggregated across all genes, calculated for each sample, and grouped by unsupervised hierarchical cluster analysis, indicates that its mutation spectrum is unique to each treatment (e.g., test mutagens). Unsupervised cluster analysis of the encoded data allows for data grouping based on mutation spectrum, showing that ENU samples are easily distinguishable in all tissues by the overwhelming majority of T→C, T→A, and C→T mutations. Similarly, B[a]P samples are distinguishable by C→A and G→T mutations.
[0192] Figures 7A–7C are graphs showing the mutation spectrum (i.e., trinucleotide spectrum) in terms of adjacent nucleotides for vehicle control (7A), B[a]P (7B), and ENU (7C). Mutation signatures in trinucleotide spectrum form provide information about different mechanisms of mutagenesis and / or show mutation patterns specific to particular mutagens. For example, contexts such as CCG and CGC appear to be more susceptible to attack by the tobacco-related carcinogen B[a]P than other contexts (Figure 7B). This signature pattern may be similar to the signature pattern shown by aflatoxin exposure (e.g., the mutagenesis mechanisms may be similar). Figure 7C shows that the alkylating agent ENU has two susceptible contexts corresponding to the IUPAC code GTS (where S+[G][C]) and is a trigger for transition mutations.
[0193] In this example, the mutational load in bone marrow and liver samples treated with ENU and B[a]P was significantly increased compared to controls, comparable to the conventional BigBlue® cII mutant plaque frequency (mutant frequency MF), and similarly varied depending on the tissue type. Spectrum evaluation revealed characteristic patterns of indels and single nucleotide substitutions in each treatment group. Trinucleotide base analysis showed that the adjacent nucleotide context strongly modulates the likelihood of mutation, with the most extreme hotspots being CCG and CGC for B[a]P and GTG and GTC for ENU. Double-strand sequencing was extended to four endogenous genes: Polr1c, rhodopsin, haptoglobin, and β-catenin. Here again, MF was increased in animals exposed to ENU and B[a]P, but significantly varied depending on the genomic locus, which likely reflects the transcriptional state. This embodiment demonstrates that double-strand sequencing is a successful method for detecting mutations in cII transgenes, which are acceptable preclinical safety biomarkers in TGR assays. Furthermore, this embodiment shows that double-strand sequencing can serve as the basis for endogenous cancer-related gene-based risk assessment tools.
[0194] Example 2 Direct Quantification of In Vivo Chemical Mutations in Mammalian Genomes Using Double-Strand Sequencing This chapter describes an example of using double-strand sequencing to determine whether early mutations in cancer driver genes reflect the tumorigenic potential of an experimental mutagen.
[0195] This example investigates the effects of urethane on different mouse tissue types (lung, spleen, and blood) in an FDA-approved mouse model prone to cancer. The mouse model is Tg.rasH2 (Saitoh et al., Oncogene 1990, PMID 2202951). This mouse contains approximately three tandem copies of human Hras with activating enhancer mutations to promote expression on a single hemizygous standalone gene. These mice are susceptible to splenic angiosarcoma and lung adenocarcinoma and are typically used in a 6-month carcinogenicity study to replace a 2-year natural animal study. Tumors observed in these mice typically acquire an activating mutation on one copy of the human Hras proto-oncogene. In addition to these four natural mouse genes (Rho, Hp, Ctnnb1, Polr1c), natural mouse Hras transgenes and natural human Hras transgenes are also analyzed in this example.
[0196] In this example, Tg.rasH2 mice (n=5 / group) were administered a vehicle or a carcinogenic dose of urethane (on days 1, 3, and 5), and euthanized on day 29 for double-strand sequencing of target tissues (lungs, spleen) and whole blood to detect mutations. Endogenous genes (Rho, Hp, Ctnnb1, Polr1c) and the Hras (trans) gene from both native mice and humans were also sequenced.
[0197] Tumors (splenic angiosarcoma, lung adenocarcinoma) were collected at 11 weeks from animals administered urethane (n=5 / group), and whole exome sequencing (WES) was performed to identify characteristic cancer driver mutations (CDMs) in these tumors.
[0198] Figure 8 is a bar graph showing the variant frequency (MF) in lung, spleen, and blood samples for control and urethane-treated experimental animals. In this analysis, all detected unique variants were counted as one mutation and summed for each sample. This was then divided by the total number of sequenced double-stranded bases across the entire capture region. The number of events is indicated above each sample. In total, 3,966,947,832 double-stranded sequenced base pairs were generated across all 30 samples. As shown in Figure 8, mutagenesis was consistent among animals in the same treatment group, and confidence increased with increasing sequencing depth.
[0199] Figure 9 is a bar graph showing the mean minimum point mutant frequency across each tissue sample group (error bars represent a standard deviation of ±1). [Table 1]
[0200] Referring together to Figure 9 and Table 1, the difference between the vehicle control (VC) and the treatment group was highly significant. Welch's t-test (for unequal variances) was used to determine the superiority of the mutant frequency in the mutagen-treated tissue compared to the mutant frequency in the tissue control. The slightly wider confidence interval for blood samples reflects the lower mean sequencing depth in blood VC samples in this particular case. It is expected that this can be corrected using the methods described herein.
[0201] Figure 10A is a box plot graph showing the calculated SNV MF for double-strand sequencing by gene regions for the lung, spleen, and blood for the indicated treatment categories, and Figure 10B is a scatter plot showing the individual measurements of the aggregated data shown in Figure 10A. The scatter points represent individual measurements, and these measurements are assigned a 95% CI. The box plot in Figure 10A shows all four quartiles for all data points for that tissue and treatment category. The Y-axis scale is linear and has a magnitude of 10 -7This is shown. Referring to Figure 10A, this box plot summarizes the SNV mutation frequencies in the lungs, spleen, and blood of the Tg-rasH2 mouse model shown in Figure 10B. The Tg-rasH2 mouse model does not have a cII transgene. The degree of mutagenesis is influenced by the specific mutagens, tissue types, and loci. Figure 11 is a bar graph showing the mutation spectrum of urethane and VC in test tissues, measured by double-strand sequencing. Referring to Figure 11, unsupervised cluster analysis of the encoded data allowed for grouping of data based on the mutation spectrum. This data shows that exposure can be identified by a simple spectrum of nucleotide mutations alone. In other words, if the mutagens are unknown, such mutagens can be newly identified by double-strand sequencing of the DNA of the exposed organisms, based on the nature of the mutation spectrum.
[0202] Figures 12A and 12B are graphs showing the mutation spectrum (i.e., trinucleotide spectrum) in terms of adjacent nucleotides for the vehicle control (12A) and urethane (12B). Mutation signatures in trinucleotide spectrum form provide information about different mechanisms of mutagenesis and / or exhibit mutation patterns specific to a particular mutagens. Thus, the detailed breakdown of each mutation type within its trinucleotide context ("triple signature") reveals a highly unique fingerprint for each treatment group, which coincides with the known signatures of clonal mutations from tumors caused by such exposure. In untreated animals, the C:G→A:T and C:G→G:C mutations resulting from guanine oxidation and deamination of cytosine and 5-me-cytosine, respectively, are known patterns of aging, and these mutations were detected. After urethane treatment, T:A→A:T within the motif "NTG" is shown as the most common mutation.
[0203] Figure 13 shows that single nucleotide variant (SNV) strand bias was observed in Ctnnb1 and Polr1c, but not in the Hp and Rho genomic regions. SNV notation is normalized to the reference nucleotide in the forward direction of the transcribed strand. Individual duplications are shown as points and 95% confidence intervals and have linear segments. All mutation frequencies were corrected for the number of nucleotides of each reference base within the variant calling region. The null hypothesis for the absence of strand bias is that intervariates have equal frequencies. Since C>N and T>N variants have uniform frequencies and G>N and A>N variants have increased frequencies, the bias is evident in Ctnnb1 and Polr1c. Compared to Hp and Rho, it is not constrained by theory, but this difference is thought to be due to transcriptionally linked nucleotide excision repair and the relative expression levels of these genes.
[0204] Figure 14 is a graph showing early neoplastic clonal selection of variant allele fractions detected by double-strand sequencing. The majority of identified mutations occurred in single molecules with very low variant allele fractions (VAFs), e.g., around 1 / 10,000. Some variants were found in multiple molecules in the sample and were identified with considerably higher VAFs.
[0205] Figure 15A is a graph showing single nucleotide variants (SNVs) plotted across genomic segments for exons captured from the Ras gene family, including the human transgene locus, in the Tg-rasH2 mouse model. A singlet is a variant found within a single molecule. A multiplet is an identical variant identified within multiple molecules within the same sampler and may represent a clonal proliferation event. The height of each point corresponds to the variant allele frequency (VAF) of each SNV, and the size of this point corresponds only to the observation of multiplets. The location and relative frequency of human cancer mutation hotspots in the Ras family in COSMIC are shown below each gene. Figure 15B is a graph showing single nucleotide variants (SNVs) aligned to exon 3 of the human HRAS transgene. The central residue at codon number 61 in exon 3 of human HRAS, the most common HRAS cancer-causing hotspot, is highlighted.
[0206] Referring to Figures 15A and 15B together, clusters of T>A transposition were observed at the human cancerous Hras codon 61 hotspot in 4 / 5 of urethane-treated lung samples and 1 / 5 of urethane-treated spleen samples. In particular, four out of five treated lung samples had this mutation at variant allele frequencies of 0.1% to 1.8%. Specifically, these clones had the T>A transposition in the contextual NTG, which is characteristic of urethane mutation introduction (indicating a strong bias in the NTG site in Figure 12B). In addition, two treated spleen samples had mutations in this codon; one at the same location and the other in an adjacent base pair. Observation of 4 / 5 treated lung samples revealed that pathogenic mutations were clonally proliferating by day 29, while very few mutations (as >1 member clones or repetitions in multiple samples, as high VAF multiples in well-established cancer drivers) in other parts of this panel are strong indicators of positive selection immediately after exposure. Furthermore, according to embodiments of this technique, double-strand sequencing methods provide the sensitivity necessary to detect such early neobiotic clonal selection. [Table 2]
[0207] Referring to Table 2, 97.5% of the mutations were identified by a single molecule, 1% by two molecules, and approximately 0.5% by more than two molecules. All four of the highest-level clones produced cancerous mutations in AA61, a recurrent tumor hotspot in human HRAS. The fact that these highest-level clones also appeared in cancerous hotspots further emphasizes the magnitude of the strong selective pressure.
[0208] A much larger amount of DNA was extracted from each sample and converted into sequenced double-stranded molecules. Approximately 5 μg of genomic DNA was obtained from the extracted portion of the tissue sample. This was converted to genomic equivalent and multiplied by 3 to obtain the number of tg.HRAS copies in this extraction. Only about 1 / 3 of this was sequenced, and therefore, approximately 300 times more mutants were present in the original portion of the sampled tissue than in the detected tissue. [Table 3]
[0209] In this example, the selected clone contained over 90,000 cells in the highest allele fraction clone. Calculations showed that, assuming no cell death within 29 days of the test (e.g., from mutation exposure), the doubling time of these cells was approximately 1.8 days, resulting in approximately 2^(29 / 1.8) = approximately 90,000 cells. While not theoretically constrained, this calculated cell doubling rate suggests a high probability of detecting these selected mutations in a short time (e.g., as little as two weeks).
[0210] Figures 16A and 16B are graphical representations of sequencing data from a representative 400-base pair region of human HRAS in urethane-treated mouse lung, using conventional DNA sequencing (Figure 16A) and double-strand sequencing (Figure 16B). Conventional DNA sequencing has an error rate of 0.1% to 1%, obscuring the presence of truly low-frequency mutations. Figure 16A shows conventional sequencing data from a representative 400-base pair region of one gene (human HRAS) in one sample (mouse lung) in this study. Each bar corresponds to a nucleotide position. The height of each bar corresponds to the allele fraction of the non-reference base at that position when sequenced to a depth of more than 100,000 times. All positions appear to mutate with some frequency, and almost all of these are errors. Referring to Figure 16B, it becomes clear that when processed with double-strand sequencing, only one mutation is true.
[0211] The results of the experimental analysis in this example demonstrate that double-strand sequencing quantifies the induction of mutagenesis by urethane very stably and with tight replication confidence intervals. Furthermore, the degree of mutagenesis is tissue-specific, with the lungs being more susceptible to mutagenesis than the spleen and blood. The simple mutagenesis spectrum of urethane exposure is clean, and unbiased clustering allows for distinguishing between groups. The triple mutagenesis spectrum of urethane shows a strong trend toward T→A and T→C mutations within the context of "NTG," and the mutagenesis spectrum is distinguishable from the vehicle control (and other mutagens, see Example 1).
[0212] In addition, mutagenesis in peripheral blood was very similar to that observed in the spleen, suggesting that pre-mortem sampling of peripheral blood may be a viable alternative to autopsy (or biopsy) for some mutagens. Furthermore, this example demonstrated that clear evidence of oncogenic mutation selection in human HRAS transgenes can be demonstrated using double-strand sequencing even at day 29. The mutation spectrum in this hotspot accurately reflected the influence of this known mutagen. Therefore, double-strand sequencing can provide early, accurate data for evaluating early cancer driver mutations as biomarkers of future cancer risk. Although heterogeneity remained at extremely low levels, the removal of foreign contaminants was automatically and reliably performed.
[0213] Example 3 Analysis of mutagen signatures in mammalian genomes using double-strand sequencing. This chapter describes examples of how data generated by double-strand sequencing analysis can be used to construct, compare, and / or identify mutagenic signatures for discriminative mutagens.
[0214] The Catalogue of Somatic Mutations in Cancer (COSMIC) database provides a reference to “mutation signatures,” which are defined as unique combinations of mutation types found to be present in the genome. Somatic mutations are present in all cells in the human body and occur throughout a person's life. Such somatic mutations are the result of multiple mutational processes, including, for example, inherent minor failures in the DNA replication mechanism, exposure to exogenous or endogenous mutagens, enzymatic modification of DNA, and defective DNA repair.
[0215] Figures 17A–17C are graphs showing the mutation spectrum (i.e., trinucleotide spectrum) in the context of adjacent nucleotides for signature 1 (Figure 17A), signature 4 (Figure 17B), and signature 29 (Figure 17C) from COSMIC. Referring to Figure 17A, signature 1 is found in all cancer types and has a proposed etiology of spontaneous deamination of 5-methylcytosine, resulting in a C>T transition at the CpG site. Referring to Figures 17B–17C, signatures 4 and 29 are correlated with smoking and are caused by benzo[a]pyrene, a major mutagen in tobacco. Although the patterns are similar, signature 4 is most frequently observed in lung cancer in smokers, while signature 29 is mainly found in squamous cell esophageal cancer and is most frequently seen in smokers and chewing tobacco users. [Table 4]
[0216] Table 4 provides experimental parameters and data derived from Examples 1 and 2 discussed herein. Figure 18 shows unsupervised hierarchical clustering of all 30 published COSMIC signatures and four cohort spectra from Examples 1 and 2. Clustering was performed using the weighted modeling (WGMA) method and cosine similarity as the metric. In particular, benzo[a]pyrene (BaP) is very similar to both Signatures 4 and 29 and correlates with BaP exposure from tobacco ingestion or inhalation. The vehicle control (VC) is similar to Signature 1, which is a pattern related to the spontaneous deamination of 5-methylcytosine and is thought to represent a mixture of mutagenic effects of both reactive oxygen species and the spontaneous deamination of 5-methylcytosine.
[0217] This example demonstrates that double-strand sequencing can be used to create a mutation spectrum analysis, which can then be compared to or referenced from known mutation signatures for identification and other analyses.
[0218] Appropriate computing environment The following considerations provide a general description of suitable computing environments in which aspects of this disclosure can be implemented. While not required, aspects and embodiments of this disclosure are described in general terms of computer-executable instructions, such as routines, executed by general-purpose computers (e.g., servers or personal computers). Those skilled in the art will understand that this disclosure can be implemented using other computer system configurations, including internet appliances, portable devices, wearable computers, cellular phones or mobile phones, multiprocessor systems, multiprocessor-based appliances or programmable appliances, set-top boxes, network PCs, minicomputers, and mainframe computers. This disclosure may also be embodied in a special-purpose computer or data processor specifically programmed, configured, or constructed to execute one or more executable instructions on a computer, as described in detail below. In practice, the term “computer,” as commonly used herein, refers to any data processor, not just any of the devices described above.
[0219] This disclosure can also be implemented in a distributed computing environment, where tasks or modules are performed by remote processing devices connected via a communication network such as a local area network ("LAN"), a wide area network ("WAN"), or the Internet. In a distributed computing environment, program modules or subroutines may reside in both local and remote memory storage devices. The aspects of this disclosure described below may be stored in firmware on a chip (e.g., an EEPROM chip) or stored or distributed on computer-readable media, including magnetically and optically readable and removable computer disks that are electronically distributed via the Internet or other networks (including wireless networks). Those skilled in the art will understand that parts of this disclosure may reside on a server computer, while corresponding parts may reside on a client computer. Data structures and data transmissions specific to the aspects of this disclosure are also included within the scope of this disclosure.
[0220] An embodiment of a computer (e.g., a personal computer or workstation) may comprise one or more processors connected to one or more user input devices and data storage devices. The computer may also be connected to at least one output device (e.g., a display device) and one or more additional output devices of optional elements (e.g., a printer, plotter, speaker, tactile or olfactory output device, etc.). The computer may be connected to an external computer, for example, by an optional network connection, wireless transceiver, or both.
[0221] Various input devices may include a keyboard and / or a pointing device (e.g., a mouse). Other possible input devices include microphones, joysticks, pens, touchscreens, scanners, digital cameras, video cameras, etc. Further input devices may include sequencers (e.g., large-scale parallel sequencers), fluorescent mirrors, and other experimental equipment. Suitable data storage devices may include any type of computer-readable medium capable of storing computer-accessible data, such as magnetic hard drives and floppy disk drives, optical disk drives, magnetic cassettes, tape drives, flash memory cards, digital video discs (DVDs), Bernoulli cartridges, RAM, ROM, smart cards, etc. In fact, any medium for storing or transmitting computer-readable instructions and data may be used, including connection ports or nodes to networks such as local area networks (LANs), wide area networks (WANs), or the Internet.
[0222] Aspects of this disclosure may be practiced in various other computing environments. For example, a distributed computing environment with a network interface may use one or more user computers in the system, which may include a browser program module that enables the computer to access and exchange data with the Internet, including websites within the World Wide Web portion of the Internet. The user computers may also include other program modules (e.g., an operating system), one or more application programs (e.g., a word processing or spreadsheet application), and so on. The computers may be general-purpose devices that can be programmed to run various types of applications, or they may be dedicated devices that are optimized for or limited to a particular function or set of functions. More importantly, any application program may be used to provide the user with a graphical user interface, as shown in the network browser but as described in detail below. The use of web browsers and web interfaces is used only as examples that are well known herein.
[0223] At least one server computer connected to the Internet or the World Wide Web ("Web") can perform many or all of the functions for receiving, routing, and storing electronic messages (e.g., web pages, data streams, audio signals, and electronic images) as described herein. While the Internet is indicated, private networks such as intranets may actually be preferred in some applications. The network may have a client-server architecture, in which computers may be dedicated computers for providing services to other client computers, or they may have other architectures such as peer-to-peer, in which one or more computers function simultaneously as both servers and clients. A database or a group of databases connected to one or more server computers can store many of the web pages and content exchanged between user computers. A server computer containing one or more databases may employ security measures (e.g., firewall systems, Secure Sockets Layer (SSL), password protection systems, encryption, etc.) to prevent malicious attacks on the system and maintain the integrity of the messages and data stored therein.
[0224] A suitable server computer may include, in particular, a server engine, a web page management element, a content management element, and a database management element. The server engine performs basic processing and operating system-level tasks. The web page management element handles the creation and display or routing of web pages. Users may access the server computer using the URL associated with it. The content management element handles most of the functions of the embodiments described herein. The database management element includes storage and retrieval tasks related to the database, querying the database, reading and writing functions to the database, and storage of data such as video, graphics, and audio signals.
[0225] Many of the functional units described herein are classified as modules to more specifically emphasize their implementation independence. For example, modules may be implemented in software for execution on various types of processors. An identified module of executable code includes, for example, one or more physical or logical blocks of computer instructions, which may be organized as, for example, an object, a procedure, or a function. Identified blocks of computer instructions do not need to be physically located together, but may contain different instructions stored in different locations, and when logically bound together, they constitute a module and achieve the stated purpose of that module.
[0226] The module may also be implemented as a hardware circuit including custom VLSI circuits or off-the-shelf semiconductors such as gate arrays, logic chips, transistors, or other distinct elements. The module may also be implemented as a programmable hardware device such as a field-programmable gate array, programmable array logic, or programmable logic device.
[0227] A module of executable code may be a single instruction or many instructions, and may be distributed across several different code segments, across several memory devices, within different programs. Similarly, operational data may be identified and shown within a module as herein, embodied in any suitable form, and organized into any suitable type of data structure. Operational data may be collected as a single dataset, or distributed across different locations, including different memory devices, or exist at least partially as mere electronic signals to a system or network.
[0228] System for genotoxicity testing The present invention further includes a system (e.g., a networked computer system, a high-throughput automated system, etc.) that processes a sample of a subject and transmits sequencing data to a remote server via a wired or wireless network to determine whether there is any similarity between the sample error-corrected sequence reads (e.g., double-stranded sequence reads, double-stranded consensus sequences, etc.), mutation spectrum, mutation frequency, triple mutation signature, and corresponding data associated with one or more known genotoxic substances.
[0229] As will be described in more detail below, and with respect to the embodiment shown in Figure 19, the computer-based system for genotoxic substances comprises (1) a remote server, (2) a plurality of user electronic computing devices capable of creating and / or transmitting sequencing data, (3) a database (optional elements) containing known genotoxic substance profiles and related information, and (4) a wired or wireless network for transmitting electronic communications between the computing devices, the database and the remote server. The remote server further comprises (a) a database storing user genotoxic substance recording results and records of genotoxic substance profiles (e.g., spectrum, frequency, mechanism of action, etc.), and (b) one or more processors communicably connected to memory and one or more non-temporary computer-readable recording devices or media containing instructions for the processors, wherein the processors are configured to execute the above-mentioned instructions for performing operations including one or more steps described in Figures 20-23.
[0230] In one embodiment, the technology further includes a non-temporary, computer-readable storage medium containing instructions for performing a method to determine whether a subject is exposed to at least one genotoxic substance and / or the identity or properties / characteristics of at least one genotoxic substance, when executed by one or more processors. In a particular embodiment, the method may include one or more steps described in Figures 20-23.
[0231] Further aspects of this technology relate to computer-aided methods for determining whether a subject is exposed to at least one genotoxic substance and / or the identity or properties / characteristics of at least one genotoxic substance. In certain embodiments, this method may include one or more steps described in Figures 20-23.
[0232] Figure 19 is a block diagram of a computer system 1900 for use with the methods and / or kits disclosed herein for identifying mutagenic and / or nucleic acid damage events resulting from exposure to genotoxic substances, with computer program product 1950 installed. While Figure 19 shows various components of the computing system, it is assumed that other or different components known to those skilled in the art (e.g., those described above) may provide a suitable computing environment in which embodiments of this disclosure can be implemented. Figure 20 is a flowchart showing a routine for providing double-stranded sequencing consensus sequence data according to one embodiment of the Art. Figures 21-23 are flowcharts showing various routines for identifying mutagenic and / or nucleic acid damage events resulting from exposure of a sample to a genotoxic substance. According to embodiments of the Art, the methods described with respect to Figures 21-23 can provide sample data including, for example, the sample's mutation spectrum, mutation frequency, triple mutation spectrum, and information derived from a comparison of the sample data with a dataset of known genotoxic substances.
[0233] As shown in Figure 19, the computer system 1900 may comprise a plurality of user computing devices 1902, 1904, a wired or wireless network 1910, and a remote server ("DupSeq® server") 1940 equipped with a processor for analyzing mutagenic and / or nucleic acid damage events resulting from a sample's exposure to a genotoxic substance. In several embodiments, the user computing devices 1902, 1904 can be used to create and / or transmit sequencing data. In one embodiment, the user of computing devices 1902, 1904 may be a user performing other aspects of the Art, such as a step in a double-strand sequencing method for a test sample to assess genotoxicity. In one example, the user of computing devices 1902, 1904 performs a specific double-strand sequencing step using a kit (1, 2) comprising reagents and / or adapters, according to one embodiment of the Art, to interrogate a test sample.
[0234] As shown in the figures, each user computing device 1902, 1904 comprises at least one central processing unit 1906, memory 1907, and a user-to-network interface 1908. In one embodiment, the user devices 1902, 1904 include a desktop, laptop, or tablet computer.
[0235] Although two user computing devices 1902 and 1904 are shown, it is assumed that any number of user computing devices may be included in or connected to other components of system 1900. In addition, computing devices 1902 and 1904 may also represent multiple devices and software used by user (1) and user (2) for amplifying and sequencing samples. For example, computing devices may be sequencers (e.g., Illumina HiSeg®, Ion Torrent PGM, ABI SOLiD® sequencer, PacBio RS, Helicos Heliscope®, etc.), real-time PCR machines (e.g., ABI 7900, Fluidigm BioMark®, etc.), microarray instruments, etc.
[0236] In addition to the components described above, the system 1900 may further include a database 1930 for storing genotoxic substance profiles and related information. For example, the database 1930 accessible by the server 1940 may contain records or collections of mutation spectra, triple mutation spectra / signatures, mechanisms of action, etc., for several known genotoxic substances, and may also contain additional information on the mutation profiles / patterns of each stored genotoxic substance. In a particular example, the database 1930 may be a third-party database 1932 containing genotoxic substance profiles. For example, the Catalogue of Somatic Mutations in Cancer (COSMIC) website contains a collection of “mutation spectra” found as clonal mutations in tumors resulting from exposure to carcinogens (e.g., lung cancer in smokers) [8, 9]. In another embodiment, the database may be a standalone database 1930 (which may or may not be private) hosted separately from server 1940, or the database may be hosted on server 1940 (e.g., database 1970) and may include empirically derived genotoxic substance profiles 1972. In some embodiments, when creating a new test drug / factor profile using system 1900, the data produced from the use of system 1900 and related methods (e.g., the methods described herein, e.g., the methods described in Figures 20-23) may be uploaded to databases 1930 and / or 1970 so that further genotoxic substance profiles 1932, 1972 may be created for future comparative work.
[0237] Server 1940 may be configured to receive, compute, and analyze sequencing data (e.g., raw sequencing files) and related information from user computing devices 1902 and 1904 via network 1910. Sample-specific raw sequencing data may be computed locally using a computer program product / module (sequencing module 1905) installed on devices 1902 and 1904 or accessible from the remote server 1940 via network 1910, or using other sequencing software well known in the art. The raw sequencing data may then be transmitted to the remote server 1940 via network 1910, and user results 1974 may be stored in database 1970. Server 1940 is configured to receive raw sequencing data from database 1970 and also includes a program product / module "DS module" 1912 which is configured to computely produce error-corrected double-stranded sequence reads using the double-stranded sequencing technique disclosed herein. Although DS module 1912 is shown on server 1940, those skilled in the art will understand that DS module 1912 may instead be hosted on devices 1902, 1904, or on another remote server (not shown).
[0238] The remote server 1940 may include at least one central processing unit (CPU) 1960, a user and network interface 1962 (or a server-dedicated computing device interfaced to the server), and a database 1970 (e.g., as described above), the database 1970 including multiple computer files / records for storing mutation profiles 1972 of known and novel genotoxic substances and files / records for storing results (e.g., raw sequencing data, double-strand sequencing data, genotoxicity analysis, etc.) 1974 for samples being tested. The server 1940 further includes computer memory 1911 storing a genotoxic substance computer program product (genotoxic substance module) 1950 according to aspects of the art.
[0239] Computer program product / module 1950, embodied in a non-temporary computer-readable medium and executed on a computer (e.g., server 1940), performs steps of the methods disclosed herein for detecting and identifying genotoxic substances. Another aspect of the disclosure comprises a computer program product / module 1950 including a non-temporary computer-readable medium having computer-readable program code or instructions embodied to enable a processor to perform genotoxicity analyses (e.g., calculated variant frequencies, variant spectrum, triple variant spectrum, genotoxic substance comparison reports, threshold level reports, etc.). These computer program instructions may be loaded into a computer or other programmable device to generate a machine, the instructions executed on the computer or other programmable device creating means for performing the functions or steps described herein. These computer program instructions may also be stored in computer-readable memory or medium that can be instructed to function in a particular manner, the instructions stored in computer-readable memory or medium generating a manufactured article containing instruction means for performing analyses. Computer program instructions may also be loaded into a computer or other programmable device and generate a process performed on the computer such that a series of operations performed on the computer or other programmable device provides a process for instructions performed on the computer or other programmable device to carry out the aforementioned functions or processes.
[0240] Furthermore, the computer program product / module 1950 may be implemented in any suitable language and / or browser. For example, it may be implemented in Python, C, and preferably using an object-oriented high-level programming language, such as Visual Basic, SmallTalk, or C++. The application may be written to be suitable for environments such as Microsoft Windows® environments, including Windows® 98, Windows® 2000, and Windows® NT. In addition, the application may be written for MacIntosh®, SUN®, UNIX, or LINUX environments. Furthermore, the functional processes may be implemented using a universal programming language or a platform-independent programming language. Examples of such multi-platform programming languages include, but are not limited to, Hypertext Markup Language (HTML), Java®, JavaScript®, Flash programming language, Common Gateway Interface / Structured Query Language (CGI / SQL), Practical Extraction Report Language (PERL), AppleScript®, and other system scripting languages, programming languages / structured query languages (PL / SQL). Browsers compatible with Java® or JavaScript®, such as HotJava®, Microsoft® Explorer®, or Netscape®, may be used. If active content web pages are used, they may include Java® applets or Active® controls or other active control technologies.
[0241] The system calls a number of routines. While some of the routines are described herein, one skilled in the art can identify other routines that the system may carry out. Further, the routines described herein can be modified in a variety of ways. For example, the order of the illustrated logic may be reordered, sub-steps may be performed in parallel, illustrated logic may be omitted, other logic may be included, among other things.
[0242] FIGS. 20-23 are flow diagrams showing routines 2000, 2100, 2200, 2300 for detecting and identifying mutagenic events and / or nucleic acid damage events resulting from sample exposure to a genotoxic substance. FIG. 20 is a flow diagram showing routine 2000 for providing double-stranded sequencing data for double-stranded nucleic acid molecules in a sample (e.g., a sample from a genotoxicity assay). Routine 2000 may be called by a computing device (e.g., a client computer or a server computer connected to a computer network). In one embodiment, the computing device includes a sequence data generation unit and / or a sequence module. By way of example, the computing device may call routine 2000 after an operator connects to a user interface that communicates with the computing device.
[0243] Routine 2000 begins in block 2002, and the sequence module receives raw sequence data from the user computing device (block 2004) and creates a sample-specific dataset containing multiple raw sequence reads derived from multiple nucleic acid molecules in the sample (block 2006). In some embodiments, the server can store the sample-specific dataset in a database for later processing. The DS module then receives a request to create double-stranded consensus sequencing data from the raw sequence data in the sample-specific dataset (block 2008). The DS module groups the sequence reads from families representing the original double-stranded nucleic acid molecules (e.g., based on SMI sequences) and compares representative sequences from individual strands with each other (block 2010). In one embodiment, the representative sequences may be one or more sequence reads from each original nucleic acid molecule. In another embodiment, the representative sequences may be single-stranded consensus sequences (SSCS) created from alignment and error correction within the representative strand. In such embodiments, the SSCS from the first strand can be compared with the SSCS from the second strand.
[0244] In block 2012, the DS module identifies nucleotide positions that are complementary between representative strands being compared. For example, the DS module identifies nucleotide positions along the compared (e.g., aligned) sequence reads where nucleotide base calls match. Furthermore, the DS module identifies positions that are not complementary between representative strands being compared (block 2014). Similarly, the DS module can identify nucleotide positions along the compared (e.g., aligned) sequence reads where nucleotide base calls do not match.
[0245] Next, the DS module can provide double-stranded sequencing data for the double-stranded nucleic acid molecules in a sample (block 2016). Such data may be in the form of a double-stranded consensus sequence for each processed sequence read. In one embodiment, the double-stranded consensus sequence may include only nucleotide positions where the representative sequences from each strand of the original nucleic acid molecule match. Therefore, in one embodiment, non-matching positions may be excluded or not taken into account such that the error-corrected double-stranded consensus sequence is a highly accurate sequence read. In another embodiment, the double-stranded sequencing data may include report information for non-matching nucleotide positions, for example when DNA damage can be assessed, so that non-matching nucleotide positions can be further analyzed. Routine 2000 may then proceed to and end at block 2018 。
[0246] Figure 21 is a flow diagram illustrating a routine 2100 for detecting and identifying mutagenic events resulting from exposure of a sample to a genotoxic agent. This routine is shown in FIG. 19 can be invoked by the computing device. Routine 2100 begins at block 2102, where the genotoxin module compares the double-stranded sequencing data from Figure 20 (e.g., after block 2016) with reference sequence information (block 2104), and identifies mutations (e.g., where the subject's sequence differs from the reference sequence) (block 2106). Next, the genotoxin module determines the variant frequency for the sample (block 2108) and generates a mutation spectrum (block 2110). In this way, mutation pattern analysis may be provided with information regarding the type, location and frequency of mutagenic events in the nucleic acid molecules analyzed from the sample. Optionally, the genotoxin module may generate a triple mutation spectrum that provides trinucleotide context and pattern information for analyzing the genotoxic outcome of exposure (block 2112).
[0247] The genotoxic substance module may also optionally compare the variant spectrum and / or triple variant spectrum (if determined) with a set of known genotoxic substance datasets (e.g., those stored in the genotoxic substance profile records in the database) (block 2114) to determine, for example, whether the sample was exposed to a known genotoxic substance, or, in another example, whether the test drug / factor has a genotoxic substance profile similar to that of already known genotoxic substances. Optionally, the genotoxic substance module may determine the possible mechanism of action of the genotoxic substance, partly based on the comparison information (block 2116). The genotoxic substance module can then provide genotoxicity data that may be stored in a sample-specific dataset in the database (block 2118). In some embodiments, not illustrated, the genotoxicity data may be used to create a genotoxic substance profile that will be stored in the database for future comparison work. Routine 2100 may then proceed to block 2120, and terminates at block 2120.
[0248] Figure 22 is a flowchart illustrating routine 2200 for detecting and identifying DNA damage events resulting from genotoxic substance exposure in a sample. This routine is shown in Figure 22. 19 This can be invoked by the computing device. Routine 2200 begins in block 2014 of Figure 20, and in determination block 2202, routine 2200 determines whether a nucleotide position that does not have complementarity is a process error. In various embodiments, the parameters for determining whether a position is mismatched between sequence reads of both strands of the original DNA molecule may be identified by the operator, by known characteristics of DNA damage, by known characteristics of a process error, by the minimum number of sequence reads that represent the mismatch, etc.
[0249] If a nucleotide location is determined to be a process error (as opposed to the site of in vivo DNA damage before DNA extraction), the DS module may exclude or disregard such non-complementary nucleotide locations (block 2204). Routine 2200 may follow block 2016 in Figure 20.
[0250] Returning to and referring back to determination block 2202, if it is determined that the nucleotide position is not a process error, the genotoxic substance module may identify such a nucleotide position that does not have complementarity as a site of potential in vivo DNA damage resulting, for example, from exposure to a genotoxic substance (block 2206). After identification, the genotoxic substance module may generate a DNA damage report associated with a sample-specific dataset in the database (block 2208). In some embodiments, the DNA damage report may be used to infer the mechanism of action of a potential genotoxic substance (not shown). Routine 2200 may continue to block 2016 in Figure 20.
[0251] Figure 23 is a flowchart illustrating routine 2300 for detecting and identifying carcinogens or carcinogen exposure in a subject. Routine 2300 is shown in Figure 19 This can be invoked by the computing device. Routine 2300 starts in block 2302, and the genotoxic substance module receives double-strand sequencing data from Figure 20 (e.g., after block 2016) and optionally genotoxicity data from Figure 21 (e.g., after block 2116) to confirm that the sample has been exposed to the genotoxic substance (block 2304). The genotoxic substance module then identifies variants in the sequence of a target genomic region (e.g., a gene) (block 2306). For example, the genotoxic substance module may analyze the double-strand sequencing data and genotoxicity data at a specific locus (e.g., a cancer driver gene, oncogene, etc.). The genotoxic substance module then calculates the variant allele frequency (VAF) (block 2308).
[0252] In decision block 2310, routine 2300 determines whether the VAF is higher in the test group than in the control group. If the VAF in the test group is not higher in the control group, the genotoxicity module labels the drug as less likely to be a carcinogen (block 2312). Routine 2300 may then proceed to block 2314, or terminate at block 2314. If the VAF is higher in the test group than in the control group, routine 2300 proceeds to decision block 2316, in which routine 2300 determines whether the mutation is non-singlet.
[0253] If the mutation is a singlet, the genotoxicity module characterizes the drug as having a moderate level of suspected carcinogenicity (block 2318). If the mutation is determined to be a non-singlet (i.e., a multiplet), this routine proceeds to determination block 2320, and routine 2300 determines whether the variant is detected in the target gene and whether the variant matches a driver mutation (e.g., a mutation known to cause cancer growth / conversion).
[0254] If the mutation is not a driver mutation, the genotoxicity module characterizes the drug as having a moderate level of suspected carcinogenicity (block 2318). If one or more variants match a driver mutation, the genotoxicity module characterizes the drug as having a high level of suspected carcinogenicity (block 2322).
[0255] For drugs characterized as having a moderate level of suspicion (in block 2318) or drugs characterized as having a high level of suspicion (in block 2318), the genotoxic substance module may assess the safety threshold for carcinogens and / or determine the risk associated with the development of post-exposure genotoxic substance-related disease or impairment in subjects (block 2324). Routine 2300 may then be followed by block 2314, and terminates in block 2314.
[0256] Other processes and routines are also envisioned by this technology. For example, the system (e.g., a genotoxic substance module or other modules) may be configured to analyze genotoxic substance data to determine whether a subject has been exposed to a genotoxic substance, whether a test drug / factor is genotoxic, and what characteristics indicate that a genotoxic substance is mutagenic or carcinogenic. Other processes may include determining whether a subject should be treated prophylactically or therapeutically based on genotoxic substance data derived from a biological sample of a particular subject. For example, once one or more genotoxic substances are identified using this system, the server can determine whether the subject has been exposed to a genotoxic substance above a safe threshold level. If the subject has been exposed to a genotoxic substance above a safe threshold level, prophylactic or suppressive disease treatment may be initiated. Further examples 1. A method for detecting and quantifying genomic mutations that occur in vivo in a subject after exposure to a mutagen, To provide a sample from a subject that contains double-stranded DNA molecules, This involves creating error-corrected sequence reads for each of the multiple double-stranded DNA molecules in the sample, Creating a set of copies of the original first strand of the adapter-DNA molecule and a set of copies of the original second strand of the adapter-DNA molecule. To sequence a set of copies of the original first and second strands to provide the first strand sequence and the second strand sequence, as well as Creating an error-corrected sequence read, which includes comparing a first strand sequence with a second strand sequence and identifying one or more correspondences between the first strand sequence and the second strand sequence. A method comprising determining the mutation spectrum of double-stranded DNA molecules in a sample by analyzing one or more correspondences. 2. The method according to Example 1, further comprising calculating the variant frequency of a target double-stranded DNA molecule by calculating the number of unique variants per sequenced double-stranded base pair. 3. The method according to Example 1, wherein the target double-stranded DNA molecule is extracted from the liver, spleen, blood, lungs, or bone marrow of the subject. 4. The method according to Example 1, wherein the subject was exposed to the mutagen within 30 days prior to the removal of the target double-stranded DNA molecule from the subject. 5. The method described in Example 1, wherein the mutation spectrum is created by unsupervised hierarchical mutation spectrum clustering. 6. The method according to Example 1, wherein the mutation spectrum is a triple mutation spectrum. 7. The method according to Example 1, wherein creating error-corrected sequence reads for each of several double-stranded DNA molecules is a method comprising creating error-corrected sequence reads for one or more target genomic regions. 8. The method according to Example 7, wherein one or more target genomic regions are mutation-prone areas in the genome. 9. The method according to Example 7, wherein one or more target genomic regions are known cancer driver genes. 10. The method according to Example 1, wherein the subject is a transgenic animal and at least some of the target double-stranded DNA molecules contain one or more portions of the transgene. 11. The method according to Example 1, wherein the subject is a non-transgenic animal and the target double-stranded DNA molecule includes an endogenous genomic region. 12. The method according to Example 1, wherein the subject is human and the target double-stranded DNA molecule is extracted from a blood sample taken from a human. 13. A method for creating a mutagenic signature of a test drug, The process involves double-strand sequencing of DNA fragments extracted from test subjects exposed to the test drug, and This includes creating a mutagenic signature for the test drug, and creating such a signature is By calculating the unique number of mutations per sequenced double-stranded base pair, the mutant frequencies of multiple DNA fragments can be calculated, A method comprising determining a mutation pattern of multiple DNA fragments, including the type of mutation, the mutated trinucleotide context, and the genomic distribution of the mutation. 14. The method according to Example 13, further comprising comparing the mutation signature of the test agent to a mutation signature of one or more known genotoxic substances. 15. The method according to Example 13, wherein the mutation signature of the test agent varies based on one or more of tissue type, exposure level to the test agent, genomic region, and subject type. 16. The method according to Example 15, wherein the subject type is human cells grown in culture. 17. The method according to Example 13, wherein the test animal is exposed to the test compound within 30 days before euthanizing the animal. 18. The method according to Example 13, wherein the mutagenic signature is generated by computational pattern matching. 19. The method according to Example 13, wherein the mutation signature is a triplicate mutation signature. 20. The method according to Example 13, wherein double-stranded sequencing of DNA fragments comprises double-stranded sequencing of one or more target genomic regions. 21. The method according to Example 20, wherein the one or more target genomic regions are mutation-prone sites in the genome. 22. The method according to Example 20, wherein the one or more target genomic regions are known cancer driver genes. 23. The method according to Example 13, wherein the test animal is a transgenic animal, and at least some of the DNA fragments comprise one or more portions of a transgene. 24. The method according to Example 13, wherein the test animal is a non-transgenic animal, and the DNA fragments comprise endogenous genomic regions. 25. A method for assessing the genotoxic potential of a test agent, comprising: (a) preparing a sequencing library from a sample comprising a plurality of double-stranded DNA fragments from a biological source exposed to the test agent, wherein the preparing the sequencing library comprises ligating asymmetric adapter molecules to the plurality of double-stranded DNA fragments to generate a plurality of adapter-DNA molecules; and (b) Sequence the first and second strands of the adapter-DNA molecule to provide a first strand sequence read and a second strand sequence read for each adapter-DNA molecule, (c) For each adapter-DNA molecule, compare the first strand sequence read with the second strand sequence read to identify one or more correspondences between the first strand sequence read and the second strand sequence read, (d) Determine the mutation signature of the test drug by analyzing one or more correspondences between the first strand sequence reads and the second strand sequence reads for each adapter-DNA molecule, and determine at least one of the following in the sample: mutation pattern, mutation type, mutation frequency, mutation type distribution, and mutation genomic distribution. (e) Comparing the mutation signature of the test drug with multiple mutation spectrums derived from known genotoxic substances to determine whether the mutation signature is sufficiently similar to the mutation spectrum from known genotoxic substances, or (f) assess whether at least one of the variant frequency, variant type, or distribution of variant types exceeds a safety threshold level, (g) A method comprising determining whether the mutant frequency exceeds a safety threshold mutant frequency. 26. The method of Example 25, wherein the mutation signature of the test drug includes a variant frequency exceeding the safety threshold frequency. 27. The method according to Example 25, wherein the mutation signature of the test drug includes a mutation pattern that is sufficiently similar to known cancer-associated mutation patterns. 28. The method according to Example 25, wherein the biological source is at least one of cells grown in culture, animals, humans, human cell lines, transgenic animals, non-transgenic animals, human tissue samples, or human blood samples. 29. The method of Example 25, wherein the biological source was exposed to the test agent within 30 days prior to the extraction of a sample containing multiple double-stranded DNA fragments. 30. The method according to Example 25, wherein the mutation signature is a triple mutation signature. 31. The method according to Example 25, wherein the method includes relating the first strand sequence read and the second strand sequence read using one or more of the adapter sequence, sequence read length, and original strand information before comparing the first strand sequence read and the second strand sequence read. 32. The method of Example 25, wherein the method further comprises exposing a biological source to a test drug before preparing a sequencing library. 33. The method according to Example 32, wherein the biological source is cancerous tissue or contains cancerous tissue before the biological source is exposed to the test drug. 34. The method of Example 32, wherein the biosource is healthy tissue or contains healthy tissue before the biosource is exposed to the test drug. 35. The method according to Example 25, wherein the sample is a blood sample or contains a blood sample. 36. The method according to Example 25, wherein the sample is a cancer cell line or contains a cancer cell line. 37. The method according to Example 25, wherein the biological source contains cancer cells, and the substance is tested for selective genotoxicity against at least a subset of cancer cells. 38. The method according to Example 37, wherein the substance is a therapeutic compound. 39. The method according to Example 38, wherein, for a subset of cancer cells that have been shown to be sensitive to the selective genotoxicity of a therapeutic compound, the method further comprises determining one or more of the mutant frequency and mutation spectrum for the subset of cancer cells before exposure to the therapeutic compound. 40. The method according to Example 25, wherein the test agent includes food, drugs, vaccines, cosmetic substances, industrial additives, industrial by-products, petroleum distillates, heavy metals, household cleaning agents, airborne particles, manufacturing by-products, pollutants, plasticizers, detergents, radioactive products, tobacco products, chemicals, or biological materials. 41. A method for determining a subject's exposure to a genotoxic agent, This involves comparing the DNA mutation spectrum of the subject with the mutation spectrum of known mutagenic compounds, A method comprising identifying the mutation spectrum of a known mutagenic compound that is most similar to the DNA mutation spectrum of a subject. 42. The method according to Example 41, wherein the DNA mutation spectrum of the subject is evaluated by double-strand sequencing. 43. The method according to Example 41, wherein the DNA mutation spectrum of the subject is created from DNA extracted from the patient's blood. 44. The method according to Example 41, wherein the DNA mutation spectrum of the subject is a triple mutation spectrum. 45. The method of Example 41, further comprising sequencing the DNA of the subject to create a DNA mutation spectrum of the subject. 46. The method of Example 45, wherein sequencing of the subject's DNA includes sequencing one or more known cancer driver genes. 47. A kit that can be used to identify genotoxic substances by error-correcting double-stranded sequencing of double-stranded polynucleotides, the kit is: At least one set of polymerase chain reaction (PCR) primers and at least one set of adapter molecules that can be used in error-corrected double-strand sequencing experiments, A kit comprising instructions on how to use the kit to perform error-corrected double-strand sequencing of DNA extracted from a subject's sample and to identify whether the subject was exposed to at least one genotoxic substance. 48. The kit described in Example 47, wherein the reagents include DNA repair enzymes. 49. The kit as described in Example 47, wherein each adapter molecule in the set of adapter molecules contains at least one single molecule identifier (SMI) sequence and at least one chain-defining element. 50. The kit according to Example 47, further comprising a computer program product embodied in a non-temporary, computer-readable medium, which, when executed on a computer, performs the steps of determining error-corrected double-strand sequencing reads for one or more double-stranded DNA molecules in a sample, and determining the variant frequency, variant spectrum and / or tri-spectrum of at least one genotoxic substance using the error-corrected double-strand sequencing reads. 51. The kit as described in Example 50, wherein the computer program product further determines the mechanism of action of a genotoxic substance in which it mutates the DNA of a subject, and a therapeutic or prophylactic treatment suitable for administration to the subject based on the mechanism of action of the genotoxic substance. 52. A method for diagnosing and treating a subject exposed to a genotoxic substance, a) Whether the subject was exposed to a genotoxic substance, i) Obtaining biological samples from the subject, ii) To provide error-corrected double-strand sequencing reads for multiple double-stranded DNA sequences extracted from a sample. iii) Determining the frequency of DNA sequence variants, the variant spectrum and / or the triple variant spectrum, iv) Determining whether the variant frequency, variant spectrum and / or triple variant spectrum indicate that the subject was exposed to a genotoxic substance, b) A method comprising providing prophylactic and / or therapeutic measures to prevent or inhibit the development of a disease or disorder associated with a genotoxic substance if a subject has been exposed to the genotoxic substance. 53. A method for identifying a threshold level for safe exposure to a genotoxic substance and for providing treatment, a) Determining the threshold level for safe exposure to genotoxic substances, b) Whether the subject was exposed to a genotoxic substance at a level higher than the safe exposure threshold. i) Obtaining biological samples from the subject, ii) To provide error-corrected double-strand sequencing reads for multiple double-stranded DNA sequences extracted from biological samples. iii) Determining the frequency of DNA sequence variants, the variant spectrum and / or the triple variant spectrum, iv) Determining whether the variant frequency, variant spectrum and / or triple variant spectrum indicate that the subject was exposed to a specific genotoxic substance. v) Determining the subject's exposure level to the genotoxic substance by calculating it based on the variant frequency, variant spectrum, and / or triple variant spectrum, c) A method comprising providing prophylactic and / or therapeutic measures to prevent or block the development of a disease or disorder associated with a genotoxic substance if a subject has been exposed to a genotoxic substance above a safe exposure threshold level. 54. A system for detecting and identifying mutagenic and / or nucleic acid damage events resulting from exposure of a sample to a genotoxic substance, A computer network for transmitting information related to sequencing data and genotoxicity data, wherein the information includes one or more of the following: raw sequencing data, double-stranded sequencing data, sample information, and genotoxic substance information. A client computer associated with one or more user computing devices and communicating with a computer network, A database connected to a computer network is used to store records of multiple genotoxic substance profiles and user results. A double-strand sequencing module is configured to communicate with a computer network and receive raw sequencing data and requests from client computers for creating double-strand sequencing data, receive group sequence reads from families representing original double-stranded nucleic acid molecules, and compare representative sequences from individual strands with each other to create double-strand sequencing data. A system comprising: a genotoxic substance module configured to communicate with a computer network, compare double-stranded sequencing data with reference sequence information, identify mutations, and create genotoxic substance data including at least one of the following: variant frequency, mutation spectrum, and triple mutation spectrum. 55. The genotoxic substance profile is a system described in Example 54, which includes a spectrum of genotoxic substance mutations from multiple known genotoxic substances. 56. Example 1 for determining whether a subject is exposed to at least one genotoxic substance and / or determining the identity of at least one genotoxic substance when executed by one or more processors. 46 and 52~ A non-temporary, computer-readable storage medium containing instructions to perform any one of the methods described in 53. 57. A non-temporary, computer-readable storage medium as described in Example 56, further comprising calculating the mutation spectrum, mutation frequency and / or triputimetric mutation spectrum of the detected drug, from which the identity of at least one genotoxic substance is determined. 58. A computer system for performing the method of any one of Examples 1 to 53 for determining whether a subject has been exposed to at least one genotoxic substance and / or the identity of at least one genotoxic substance, wherein the system includes at least one computer comprising a processor, memory, a database, and a non-temporary computer-readable storage medium containing instructions for the processor, and the processor is the one described in Examples 1 to 53. 46 and 52~ A computer system configured to execute instructions to perform an operation including one of the methods described in any of 53. 59. The system described in Example 58, further comprising a network computer system, wherein the network computer system is a. A wired or wireless network, b. Multiple user-controlled computing devices capable of receiving data derived from the use of a kit containing reagents for extracting, amplifying, and producing polynucleotide sequences from a subject's sample, and transmitting the polynucleotide sequences to a remote server via a network. c. A processor comprising memory, a database, and a non-temporary computer-readable storage medium containing instructions for the processor, wherein the processor is, Example 1~ 46 and 52~ A remote server configured to execute instructions to perform an operation including one of the methods described in 53, d. A system in which a remote server can detect and identify mutagenic and / or nucleic acid damage events resulting from a sample's exposure to a genotoxic substance. 60. A third-party database accessible via a database and / or network further includes multiple records containing genotoxicity profiles of known genotoxic substances, one or more of the genotoxicity profiles of at least one subject sample, wherein the genotoxicity profiles include sites of mutation or DNA damage, as described in Example 59. 61. Content is a non-temporary computer-readable medium that causes at least one computer to perform a method for providing double-strand sequencing data for double-stranded nucleic acid molecules in a sample from a genotoxicity screening assay, wherein the method is Receiving raw array data from the user computing device, This involves creating a sample-specific dataset containing multiple raw sequence reads derived from multiple nucleic acid molecules in the sample, This involves grouping sequence reads from a family representing the original double-stranded nucleic acid molecule, where this grouping is based on a shared single-molecule identifier sequence. The first strand sequence reads and the second strand sequence reads from the original double-stranded nucleic acid molecule are compared to identify one or more correspondences between the first strand sequence reads and the second strand sequence reads. A non-temporary, computer-readable medium that provides double-strand sequencing data for double-stranded nucleic acid molecules in a sample. 62. Example 61 A computer-readable medium described in [the text] further comprises identifying nucleotide positions having complementarity between a first sequence read and a second sequence read being compared, and the method further comprises: Identifying, eliminating, or discounting process errors at locations where complementarity exists, A computer-readable medium comprising identifying the remaining non-complementary locations as potential in vivo DNA damage sites resulting from exposure to a genotoxic substance, in addition to the non-complementary locations not identified as process errors. 63. Content is a non-temporary computer-readable medium that causes at least one computer to perform a method for detecting and identifying mutagenic events resulting from exposure of a sample to a genotoxic substance, and the method is Comparing double-stranded sequence data with reference sequence information, Mutations in double-stranded sequence data are identified, where mutations are identified as regions that do not match the reference information. Determining the frequency of variants in double-stranded sequence data, Creating a mutation spectrum from double-stranded sequence data, Creating a triple mutation spectrum from double-stranded sequence data, A non-transient, computer-readable medium, including the comparison of mutational and / or triputimetric mutational spectra with multiple known genotoxic substance datasets. 64. Content is a non-temporary computer-readable medium that causes at least one computer to perform a method for detecting and identifying carcinogens or exposure to carcinogens in a subject, and the method is Identifying sequence variants in a target genomic region using double-stranded sequencing data created from samples from subjects, Calculate the variant allele frequencies (VAF) of the test sample and the control sample, To determine whether VAF is higher in the test group than in the control group, In samples with a higher VAF, it is necessary to determine whether the sequence variant is a non-singlet, In samples with a higher VAF, the goal is to determine whether the sequence variant is a driver mutation, Characterizing samples with non-singlet and / or driver mutations as suspected carcinogens, and including this in non-transient, computer-readable media. 65. Further including, for example, assessing the safety threshold of a carcinogen and / or determining the associated risk of developing a genotoxic substance-related disease or disorder in a subject after exposure. 64 A non-temporary computer-readable medium as described above.
[0257] References The references listed below, and the patents and published patent applications cited in the above specification, are incorporated by reference in their entirety as if they were fully described herein. [1]Schmitt MW, Kennedy SR, Salk JJ, Fox EJ, Hiatt JB, and Loeb LA.Detection of ultra-rare mutations by next-generation sequencing.Proc Natl Acad Sci US A.2012;109(36):14508-14513. [2]Kennedy SR, Salk JJ, Schmitt MW, Loeb LA.Ultra-Sensitive Sequencing Reveals an Age-Related Increase in Somatic Mitochondrial Mutations that are inconsistent with oxidative damage.PLOS Genetics.2013;9(9):1-10. [3]Kennedy SR, Schmitt MW, Fox EJ, Kohrn BF, Salk JJ, Ahn EH, et al. Detecting ultralow-frequency mutations by Duplex Sequencing. Nat Protoc. 2014;9(11):2586-2606. [4] Schmitt MW, Fox EJ, Prindle MJ, Reid-Bayliss KS, True LD, et al.Sequencing small genomic targets with high efficiency and extreme accuracy.Nature Methods.2015;12(5):423-5. [5]Chan CY, Huang PH, Guo F, Ding X, Kapur V, Mai JD, et al. Accelerating drug discovery via organs-on-chips.Lab Chip.2013;12(24):4697-4710. [6]Schmitt MW, Loeb LA, and Salk JJ.The influence of subclonal resistance mutations on targeted cancer therapy.Nat Rev Clin Oncol.2016;13(6):335-347. [7] Salk JJ, Schmitt MW, Loeb L A. Enhancing the accuracy of next-generation sequencing for detecting rare and subclonal mutations.Nature Reviews Genetics.2018.19:269-283.
[0258] conclusion The above-described detailed description of embodiments of the Art is not intended to be exhaustive or to limit the Art to the exact forms described above. Specific embodiments and examples of the Art are described above for illustrative purposes, but various equivalent modifications are possible within the scope of the Art, as will be recognized by those skilled in the art. For example, the steps are presented in a given order, but alternative embodiments may perform the steps in a different order. Further embodiments may also be provided by combining the various embodiments described herein. All references cited herein are incorporated by reference as if they were entirely contained herein.
[0259] From the above description, it should be understood that while certain embodiments of the Art are described herein for illustrative purposes, well-known structures and functions are not shown in detail to avoid unnecessarily obscuring the description of the embodiments of the Art. Where the context allows, singular or plural terms may also include plural or singular terms.
[0260] Furthermore, unless the word “or” is explicitly limited to mean only a single item that is exclusive from other items in relation to a list of two or more items, the use of “or” in such a list shall be interpreted as including (a) any single item in that list, (b) all items in that list, or (c) any combination of items in that list. In addition, the term “comprising” is used throughout to mean including at least one of the listed features, and does not exclude any more than one of the same features and / or other features of additional types. While certain embodiments are described herein for illustrative purposes, it should be understood that various modifications can be made without departing from the Art. Furthermore, while advantages associated with certain embodiments of the Art are described in the context of those embodiments, other embodiments may also demonstrate such advantages, and not all embodiments are required to demonstrate advantages that fall within the scope of the Art. Thus, this disclosure and related art may encompass other embodiments not expressly shown or described herein.
[0261] Product names used in this disclosure are for identification purposes only. All trademarks are the property of their respective owners.
Claims
1. A method for creating a mutagenic signature of a test drug, (a) Desequencing the double-stranded DNA fragments extracted from test subjects exposed to the test drug, and (b) Creating a mutagenic signature for the test drug, The mutation frequency of multiple DNA fragments is calculated by calculating the unique number of mutations per sequenced double-stranded base pair, Determining the mutation patterns of the plurality of DNA fragments, wherein the mutation patterns include the type of mutation, the mutated trinucleotide context, and the genomic distribution of the mutations, and creating the mutation patterns. Includes, Here, the mutagenic signature is a triple mutagenic signature, and The test subjects were exposed to the test drug within 30 days prior to the extraction of DNA fragments from the test subjects. method.
2. The method according to claim 1, further comprising comparing the mutation signature of the test agent with the mutation signatures of one or more known genotoxic substances.
3. The method according to claim 1, wherein the mutation signature of the test drug changes based on one or more of the following: tissue type, level of exposure to the test drug, genomic region, and type of subject.
4. The method according to claim 3, wherein the type of subject is human cells grown during culture.
5. The method according to claim 1, wherein the mutagenic signature is created by computational pattern matching.
6. The method according to claim 1, wherein double-strand sequencing of a DNA fragment includes double-strand sequencing of one or more target genomic regions.
7. The method according to claim 6, wherein one or more target genomic regions include a region in the genome that is prone to mutation.
8. The method according to claim 6, wherein the one or more target genomic regions include known cancer driver genes.
9. The method according to claim 1, wherein the test subject includes a transgenic animal, and at least some of the DNA fragments include one or more portions of an introduced gene.
10. The method according to claim 1, wherein the test subject includes a non-transgenic animal and the DNA fragment includes an endogenous genomic region.
11. A non-temporary computer-readable storage medium, which includes instructions, when executed by one or more processors, for performing the method according to any one of claims 1 to 10 for determining whether a subject has been exposed to at least one genotoxic substance and / or for determining the identity of at least one genotoxic substance.
12. A non-temporary computer-readable storage medium according to claim 11, further comprising calculating the mutation spectrum, mutation frequency and / or triple mutation spectrum of the detected drug, from which the identity of the at least one genotoxic substance is determined.
13. A computer system for performing the method according to any one of claims 1 to 10 for determining whether a subject has been exposed to and / or the identity of at least one genotoxic substance, wherein the system includes at least one computer comprising a processor, memory, a database, and a non-temporary computer-readable storage medium containing instructions for the processor, the processor being configured to execute the instructions for performing the operation according to any one of claims 1 to 10.
14. The network computer system further includes, a. A wired or wireless network, b. Multiple user-computer devices capable of receiving data resulting from the use of a kit containing reagents for extracting, amplifying, and producing polynucleotide sequences from a subject's sample, and transmitting the polynucleotide sequences to a remote server via a network. c. A remote server comprising the processor, memory, a database, and the non-temporary computer-readable storage medium containing instructions for the processor, wherein the processor is configured to execute the instructions to perform an operation including the method according to any one of claims 1 to 12, d. The system according to claim 13, wherein the remote server is capable of detecting and identifying mutagenic events and / or nucleic acid damage events resulting from the exposure of a sample to a genotoxic substance.
15. The system according to claim 14, wherein the database and / or a third-party database accessible via the network further includes a plurality of records, each containing one or more genotoxicity profiles of known genotoxic substances and genotoxicity profiles of at least one subject sample, wherein the genotoxicity profiles include sites of mutation or DNA damage.
Citation Information
Patent Citations
In vitro method for determining genotoxic and non-genotoxic carcinogenicity of a compound.
EP2706123A1
Methods of lowering the error rate of massively parallel DNA sequencing using duplex consensus sequencing
US20150044687A1
Systems and methods for analyzing nucleic acid
WO2016149261A1
Method for evaluating genotoxicity of substance
WO2018150513A1