Methods and processes for non-invasive assessment of chromosomal alterations

By analyzing nucleic acid sequence readings in maternal plasma using next-generation sequencing technology, chromosomal alterations can be identified, solving the problem of non-invasive detection of chromosomal alterations in existing technologies and enabling rapid and economical prenatal diagnosis.

CN111863131BActive Publication Date: 2026-05-15SEQUENOM INC
View PDF 19 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SEQUENOM INC
Filing Date
2014-10-03
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and economically detect and diagnose chromosomal alterations, particularly abnormalities in fetal nucleic acids in maternal plasma, through non-invasive methods, thus affecting the accuracy and safety of prenatal diagnosis.

Method used

Next-generation sequencing technology is used to analyze nucleic acid sequence reads in maternal plasma, identify changes in the mappability of sequence read subsequences, compare the number of sequence reads in the sample with those in the reference, and determine whether there are chromosomal alterations, such as translocations, deletions, inversions, and insertions.

Benefits of technology

It enables non-invasive, rapid, and economical detection of chromosomal alterations, improving the accuracy and safety of prenatal diagnosis and reducing the risks of invasive examinations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_3
    Figure SMS_3
  • Figure SMS_4
    Figure SMS_4
  • Figure SMS_5
    Figure SMS_5
Patent Text Reader

Abstract

Provided herein are methods, processes, systems, machines, and apparatuses for non-invasive assessment of chromosomal alterations.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related patent applications

[0002] This patent application claims the benefit of U.S. Provisional Patent Application 61 / 887,801, filed October 7, 2013, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF CHROMOSOMEALTERATIONS,” inventors Sung K. Kim, Taylor Jacob Jensen, and Mathias Ehrich, SEQ-6074-PV. The entire contents of the aforementioned patent application are incorporated herein by reference, including all text, tables, and figures.

[0003] field

[0004] The technical aspects of this article relate to methods, procedures, machines, and devices for non-invasive assessment of chromosomal alterations. background

[0005] The genetic information of living organisms (such as animals, plants, and microorganisms) and other forms of replicating genetic information (such as viruses) is encoded as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a series of nucleotides or modified nucleotides that represent the primary structure of a chemical or presumed nucleic acid. The complete human genome contains approximately 30,000 genes located on twenty-four (24) chromosomes (see The Human Genome, T. Strachan, BIOS Science Press, 1992). Each gene encodes a specific protein, which, after being expressed through transcription and translation, performs a specific biochemical function in living cells.

[0006] The identification of one or more chromosomal alterations can aid in the diagnosis of specific medical conditions or in determining the underlying causes of those conditions. Identifying chromosomal alterations can also assist in medical decision-making and / or the use of beneficial medical treatments. In some implementations, the identification of one or more chromosomal alterations involves the analysis of cell-free DNA. Cell-free DNA (CF-DNA) consists of DNA fragments derived from cell death and peripheral blood circulation. High concentrations of CF-DNA can indicate certain clinical conditions such as cancer, trauma, burns, myocardial infarction, stroke, sepsis, infection, and other diseases. Furthermore, cell-free fetal DNA (CFF-DNA) can be detected in maternal bloodstream and is used in a variety of non-invasive prenatal diagnostic procedures.

[0007] The presence of fetal nucleic acids in maternal plasma allows for non-invasive prenatal diagnosis through the analysis of maternal blood samples. For example, quantitative abnormalities in fetal DNA in maternal plasma can be associated with a variety of pregnancy-related diseases and genetic disorders related to chromosomal alterations. Therefore, analyzing fetal nucleic acids in maternal plasma can be a useful mechanism for monitoring maternal and infant health.

[0008] Overview

[0009] Some aspects of the present invention provide a system including a memory and one or more microprocessors, wherein the memory includes instructions, and the one or more microprocessors are configured to perform, according to the instructions, a process for determining whether one or more chromosomal alterations are present in nucleic acids of a sample, the process including...

[0010] (a) Characterizing the mappability of multiple sequence read subsequences with respect to sequence reads, wherein each sequence read has multiple sequence read subsequences of different lengths, and the sequence reads are sequence reads of sample nucleic acids.

[0011] (b) Identify a subset of sequence reads in which the mappability of one or more subsequences changes.

[0012] (c) A comparison is generated by comparing the number of sequence readings from the subset of samples identified in (i)(b) with the number of sequence readings from the subset of references identified in (ii)(b); and

[0013] (d) Determine whether one or more chromosomal alterations exist in the sample based on the comparison in (c).

[0014] Some aspects of the present invention also provide a method including a memory and one or more microprocessors, wherein the memory includes instructions, and the one or more microprocessors are configured to perform, according to the instructions, a process for determining whether one or more chromosomal alterations are present in nucleic acids of a sample, the process including...

[0015] (a) Characterizing the mappability of multiple sequence read subsequences with respect to sequence reads, wherein each sequence read has multiple sequence read subsequences of different lengths, and the sequence reads are sequence reads of sample nucleic acids.

[0016] (b) Identify a subset of sequence reads in which the mappability of one or more subsequences changes.

[0017] (c) A comparison is generated by comparing the number of sequence readings from the subset of samples identified in (i)(b) with the number of sequence readings from the subset of references identified in (ii)(b); and

[0018] (d) Determine whether one or more chromosomal alterations exist in the sample based on the comparison in (c).

[0019] Some aspects of the present invention also provide a non-transitory computer-readable storage medium having an executable program stored thereon configured to instruct a microprocessor to perform the following operations.

[0020] (a) Characterizing the mappability of multiple sequence read subsequences with respect to sequence reads, wherein each sequence read has multiple sequence read subsequences of different lengths, and the sequence reads are sequence reads of sample nucleic acids.

[0021] (b) Identify a subset of sequence reads in which the mappability of one or more subsequences changes.

[0022] (c) A comparison is generated by comparing the number of sequence readings from the subset of samples identified in (i)(b) with the number of sequence readings from the subset of references identified in (ii)(b); and

[0023] (d) Determine whether one or more chromosomal alterations exist in the sample based on the comparison in (c).

[0024] Certain technical aspects and implementation methods are further described in the following description, examples, claims and drawings. Attached Figure Description

[0025] The accompanying drawings illustrate embodiments of the present technology but are not limiting. For clarity and convenience, the drawings are not made to scale and in some cases, various aspects may be exaggerated or enlarged to aid in understanding the specific embodiments.

[0026] Figure 1A -C indicates the identification of fetal balanced translocation in maternal plasma. Figure 1A The Circos diagram (Krzywinski M. et al., (2009) Genome Res. 19:1639-45) details the fetal balanced translocation identified between chromosomes 8 and 11. Diagonal lines represent the beginning and end of the sequencing fragment. Chromosomes are highlighted to emphasize band shape and centromere. Figure 1B The Circos plot shows the area of ​​translocations identified in each affected chromosome. Repeated regions within each of these regions are highlighted in black. The lines represent the beginning and end of the sequencing fragment. Figure 1CThis displays a base-level description of individual sequencing reads across translocation breakpoints across various reciprocal translocation events. The sequence of chromosome 8 is indicated by "CHR8" and is located to the right of the label. The sequence of chromosome 11 is indicated by "CHR11" and is located to the right of the label. Vertical dashed lines indicate chromosome breakpoint locations. The label "CHR8 / CHR11" shows the chromosome 8 sequence to the left of the breakpoint location and the chromosome 11 sequence to the right of the breakpoint location. The label "CHR11 / CHR8" shows the chromosome 11 sequence to the left of the breakpoint location and the chromosome 8 sequence to the right of the breakpoint location. Horizontal dashed lines indicate deleted nucleotides.

[0027] Figure 2A -D displays the average MAPQ score of the simulated readout subsequences containing structural rearrangement breakpoints (vertical black lines) at each location. The readout subsequences are generated in single-base increments. Figures 2A-2D Displays the overall mapping confidence of each sequence reading subsequence (false reading) for partner pair 1 (R1) and partner pair 2 (R2), where the breakpoint for a given target fragment length of approximately 140 bp is located at position 10 ( Figure 2A ), 40 Figure 2B ), 70 (Figure 2C) or 120 ( Figure 2D For R1, the average MAPQ score for the spurious read length of 32–100 bp is plotted as a gray square starting from the leftmost position of the segment. For R2, the average MAPQ score for the spurious read length of 32–100 bp is plotted as a reverse black square starting from the rightmost end of the segment. Figure 2C The changes in the average MAPQ of the spurious reads that were confirmed to have high mapping properties became unmapped due to the addition of sequences from different genomic regions.

[0028] Figure 3 The simulated translocations of two regions containing highly unique sequences are displayed. For each simulated translocation event, the average slope (y-axis) of the mapped mass fraction is plotted at all simulated breakpoint locations (x-axis).

[0029] Figure 4 The simulated translocations are displayed for the left region containing highly unique sequences and the right region containing repeating elements. For each simulated translocation event, the average slope (y-axis) of the mapped mass fraction is plotted at all simulated breakpoint locations (x-axis).

[0030] Figure 5A -B displays Mixture B ( Figure 5A ) and the pooled control group ( Figure 5B A translocation (possibly a false positive) was observed between chromosomes 2 and 5. Gray bars represent regions of repeating elements. The left and right coordinates correspond to chromosomes 2 and 5 (hg19), respectively.

[0031] Figure 6 Exemplary implementations of the display system, wherein certain implementations of the technology may be carried out.

[0032] Figure 7 Exemplary implementations of the display filter are shown.

[0033] Figure 8 Exemplary implementations of the display system, wherein certain implementations of the technology may be carried out. Invention Details

[0034] This document provides systems and methods for analyzing polynucleotides in mixtures of nucleic acids, including, for example, methods for determining the presence of chromosomal alterations (translocations, deletions, inversions, insertions). Chromosomal alterations are widespread in populations, causing phenotypic variations within the population. Certain chromosomal alterations can play a role in the onset and progression of various diseases (e.g., cancer), disorders (e.g., structural defects, reproductive disorders), and impairments (e.g., mental disorders). The systems, methods, and products provided by this invention can be used to locate and / or identify chromosomal alterations and can be used for the diagnosis and treatment of diseases, conditions, and impairments associated with certain chromosomal alterations.

[0035] Next-generation sequencing allows for the sequencing of nucleic acids at the whole genome scale using faster and cheaper methods than conventional sequencing. The methods, systems, and products provided herein utilize advanced sequencing technologies to locate and identify chromosomal alterations and / or associated diseases and conditions. The methods, systems, and products provided by this invention generally provide non-invasive assessment of the target genome (e.g., fetal genome) using a blood sample or a portion thereof, and are generally safer, faster, and / or cheaper than more invasive techniques (e.g., amniocentesis, biopsy). In some embodiments, this invention provides a method that includes obtaining sequence reads (also referred to herein as “sequencing reads”) of nucleic acids present in a sample, typically mapped to a reference sequence, identifying certain mapping features of a selected subset of sequence reads, and determining the presence of chromosomal alterations. In some embodiments, systems, machines, apparatus, products, and modules for performing the methods described herein are also provided herein.

[0036] Chromosomal alterations

[0037] This document provides methods and systems for identifying the presence of one or more chromosomal alterations. As used herein, “chromosomal alteration” refers to any insertion, deletion (e.g., deletion), translocation, inversion, and / or fusion of genetic material in one or more human chromosomes. The term “genetic material” refers to one or more polynucleotides. Chromosomal alterations may include, or be, insertions, deletions, and / or translocations of polynucleotides of any length, with non-limiting examples including polynucleotides of at least 10 bp, at least 20 bp, at least 50 bp, at least 100 bp, at least 500 bp, at least 1000 bp, at least 2500 bp, at least 5000 bp, at least 10,000 bp, at least 50,000 bp, at least 100,000 bp, at least 500,000 bp, at least 1 megabase pair (Mbp), at least 5 Mbp, at least 10 Mbp, at least 20 Mbp, at least 50 Mbp, at least 100 Mbp, and at least 150 Mbp. In some implementations, chromosomal alterations include insertions, deletions, or translocations of polynucleotides of the following lengths: approximately 10 bp to approximately 200 Mbp, approximately 20 bp to approximately 200 Mbp, approximately 50 bp to approximately 200 Mbp, approximately 100 bp to approximately 200 Mbp, and approximately 500 Mbp. bp - approximately 200Mbp, approximately 1000bp - approximately 200Mbp, approximately 2500bp - approximately 200Mbp, approximately 5000bp - approximately 200Mbp, approximately 10,000bp - approximately 200Mbp, approximately 50,000bp - approximately 200Mbp, approximately 100,000bp - approximately 200Mbp, approximately 500,000bp - approximately 200Mbp, approximately 1Mbp - approximately 200Mbp, approximately 5Mbp - approximately 200Mbp, approximately 10Mbp - approximately 200Mbp, approximately 20Mbp - approximately 200Mbp, approximately 50Mbp - approximately 200Mbp, approximately 100Mbp - approximately 200Mbp, or approximately 150Mbp - approximately 200Mbp. In some embodiments, chromosomal alterations include insertions, deletions, and / or translocations of the following polynucleotides: about 1% or more of the chromosome, about 2% or more of the chromosome, about 3% or more of the chromosome, about 4% or more of the chromosome, about 5% or more of the chromosome, about 10% or more of the chromosome, about 15% or more of the chromosome, about 20% or more of the chromosome, about 25% or more of the chromosome, or about 30% or more of the chromosome. Non-limiting examples of chromosomal alterations detectable by the methods and / or systems described herein are detailed herein and shown in Table 1 (see below).

[0038] Chromosomal alterations sometimes include the insertion, deletion, and / or translocation of homologous genetic material. Homologous genetic material typically includes any suitable polynucleotides homologous to a human reference genome or a portion thereof. In some embodiments, chromosomal alterations include the insertion, deletion, and / or translocation of heterologous genetic material. As used herein, "heterologous genetic material" refers to genetic material derived from any non-human species. Heterologous genetic material sometimes includes polynucleotides highly homologous to the genome or a portion thereof of any non-human species. Examples of heterologous genetic material include viral genomes or portions thereof. Genuses, families, groups, and species of viruses that can be used as heterologous genetic material include Herpesviridae, Adenoviridae, Papillaviridae, Circoviridae, Circoviridae, Parvoviridae, Reoviridae, A Retroviruses, B Retroviruses, G Retroviruses, D Retroviruses, E Retroviruses, Lentivirals, PUMAviruses, Parvoviruses, Bonaviridae, Circoviruses, and Polyomaviruses.

[0039] In some embodiments, the genetic alteration includes or is a translocation. The term "translocation" herein refers to a chromosomal mutation in which the location of a segment of chromosome is altered. A translocation can be a unidirectional translocation or a reciprocal translocation. A unidirectional translocation involves the transfer of genetic material from one segment of the genome (which is deleted or duplicated) to another segment of the genome (which is inserted). A reciprocal translocation involves the exchange of genetic material between one segment of the genome and another segment of the genome. Translocations can occur intrachromosomally (e.g., intrachromosomal translocation) or between chromosomes (e.g., interchromosomal translocation). A translocation can be a balanced translocation, where the exchange of genetic material does not involve the deletion or gain of genetic material. For example, a balanced translocation is typically an exchange of segment x and segment y, where the length and integrity (e.g., sequence) of segments x and y are maintained during the exchange, and no genetic material other than segments x and / or y is added or removed. In some embodiments, a balanced translocation is a unidirectional translocation, where a genomic segment is inserted and no other genetic material other than the inserted segment is added or removed at the insertion site. In some embodiments, the exchanged one or two polynucleotides in a balanced translocation comprises one or more genetic variations (e.g., SNPs, microinsertions, microdeletions), determined by comparison with a reference genome. This genetic variation is typically located within the translocated polynucleotide (e.g., not at or near the breakpoint), and is generally not a result of the translocation event. In some embodiments, the translocation is an unbalanced translocation, where the exchange of segments x and y includes the deletion of genetic material in x and / or y. In some embodiments, an unbalanced translocation includes the exchange of segments x and y, where genetic material is added to x and / or y (e.g., duplication, insertion). Unbalanced translocations typically involve the gain or deletion of genetic material at the end of the translocated polynucleotide (e.g., at or near the breakpoint of each segment). The presence of a translocation and the location of the translocation breakpoint can be determined by the methods or systems described herein. Determining the presence of a translocation typically involves determining the presence of inserted, deleted, and / or exchanged genetic material (e.g., polynucleotides). In some embodiments, the presence of a translocation is determined by the methods and / or systems described herein.

[0040] In some embodiments, translocation includes or is inversion. Inversion is sometimes referred to herein as "chromosomal inversion." In some embodiments, inversion is the removal of a chromosomal segment and its reintroduction into the same chromosome in the opposite direction (e.g., relative to the 5'→3' DNA strand). In some embodiments, the segment is reintroduced into the chromosome at approximately the same location where it was removed. In some embodiments, the segment is reintroduced into the chromosome at a different location where it was removed. In some embodiments, inversion does not result in the deletion or addition of genetic material, although sometimes phenotypic consequences may occur when the inversion breakpoint occurs in a gene or region controlling gene expression. In some embodiments of inversion, genetic material is deleted or added, and can be detected using the methods described herein.

[0041] In some embodiments, chromosomal alterations include or are insertions. Insertion sometimes refers to “chromosomal insertion.” Insertion is sometimes the insertion of one or more nucleotide base pairs or polynucleotides into the genome or a segment thereof (e.g., a chromosome). In some embodiments, insertion refers to the insertion of a large sequence into the chromosome, typically due to unequal crossing over during mitosis. Insertion sometimes refers to a translocation or a portion thereof. For example, in a unidirectional translocation, a polynucleotide is deleted (deleted) at one location and inserted at another (e.g., insertion). In some embodiments, insertion is independent of a translocation (e.g., an insertion containing viral DNA). In some embodiments, insertion is not a translocation. Insertion is often associated with a translocation. For example, sometimes insertion includes additional genetic material added during an unbalanced translocation. Insertions associated with unbalanced translocations are typically added between the chromosomal breakpoint and one or both ends of the translocated polynucleotide. Insertions associated with unbalanced translocations are referred to herein as “microinsertions.” Microinsertions or portions thereof may include homologous and / or heterologous genetic material. In some embodiments, microinsertions or portions thereof include nucleic acids or polynucleotides of unknown origin and / or unknown homology. Determining the presence of an insertion typically involves determining whether microinsertions and / or translocations (e.g., the presence of translocation polynucleotides) are present.

[0042] In some embodiments, the microinsertion is approximately 1 bp to approximately 10,000 bp, approximately 1 bp to approximately 5,000 bp, approximately 1 bp to approximately 1,000 bp, approximately 1 bp to approximately 500 bp, approximately 1 bp to approximately 250 bp, approximately 1 bp to approximately 100 bp, approximately 1 bp to approximately 50 bp, or approximately 1 bp to approximately 30 bp. In some embodiments, the polynucleotide insertion length is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 bp.

[0043] In some embodiments, the genetic alteration includes or is a deletion. Deletion, as used herein, sometimes refers to chromosomal deletion. Deletion, as used herein, refers to the absence and / or loss of genetic material (e.g., one or more nucleotides, a polynucleotide sequence) expected to be located at a specific location (site) or a specific sequence of the genome based on a reference genome. In some embodiments, deletion refers to the absence of a continuous strand of nucleic acid (e.g., a polynucleotide). In some embodiments, deletion results in the absence of genetic material from the genome. Deletion sometimes refers to a unidirectional translocation or a portion thereof. For example, in a unidirectional translocation, a polynucleotide is deleted (removed) at one location and inserted at another. In some embodiments, deletion is not a translocation. Sometimes deletion is independent of translocation. Deletion sometimes accompanies a translocation. For example, deletion sometimes includes genetic material lost during an unbalanced translocation. The determination of missing and / or lost genetic material due to deletion accompanying an unbalanced translocation is referred to herein as microdeletion. Microdeletion accompanying an unbalanced translocation may be a deletion from one or both ends of the translocated polynucleotide. In some embodiments, microdeletion accompanying an unbalanced translocation is a deletion from one or both ends of the insertion site. Microdeletion may include the absence of genetic material at one or both ends of a breakpoint. In some implementations, determining the presence of microdeletions typically includes determining the loss of genetic material and / or the presence of translocations (e.g., the presence of translocated polynucleotides).

[0044] In some implementations, the length of the microdeletion is approximately 1 bp to approximately 10,000 bp, approximately 1 bp to approximately 5,000 bp, approximately 1 bp to approximately 1,000 bp, approximately 1 bp to approximately 500 bp, approximately 1 bp to approximately 250 bp, approximately 1 bp to approximately 100 bp, approximately 1 bp to approximately 50 bp, or approximately 1 bp to approximately 30 bp. In some implementations, the microdeletion includes deletions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 bp.

[0045] sample

[0046] This document provides systems, methods, and products for analyzing nucleic acids. In some embodiments, nucleic acid fragments in a mixture of nucleic acid fragments are analyzed. The nucleic acid mixture may include two or more types of nucleic acid fragments having different nucleotide sequences, different fragment lengths, different sources (e.g., genomic sources, fetal and maternal sources, cell or tissue sources, cancer and non-cancer sources, tumor and non-tumor sources, sample sources, object sources, etc.) or combinations thereof.

[0047] The nucleic acids or mixtures of nucleic acids used in the systems, methods, and products described herein are often isolated from samples obtained from a subject. The subject can be any living or non-living organism, including but not limited to humans, non-human animals, plants, bacteria, fungi, or protozoa. Any human or non-human animal can be selected, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovids (e.g., cattle), equines (e.g., horses), goats and sheep (e.g., sheep, goats), suidae (e.g., pigs), alpacas (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), bears (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The subject can be male or female (e.g., women, pregnant women). The subject can be of any age (e.g., embryos, fetuses, infants, children, adults).

[0048] Nucleic acids can be isolated from any type of suitable biological sample or specimen (e.g., a test sample). The sample or test sample can be any specimen isolated from or obtained from a subject or a portion thereof (e.g., a human subject, a pregnant female, a fetus). Non-limiting examples of samples include liquids or tissues of the subject, including but not limited to blood or blood products (e.g., serum, plasma, etc.), cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, cerebrospinal fluid, lavage fluid (e.g., bronchoalveolar, gastric, peritoneum, catheter, ear, arthroscopy), biopsy samples (e.g., from pre-implantation embryos, cancer biopsies), intermembranous fluid samples, cells (blood cells, placental cells, embryonic or fetal cells, nucleated fetal cells, or fetal cell remnants) or portions thereof (e.g., mitochondria, nucleus, extracts, etc.), female genital tract lavage fluid, urine, feces, sputum, saliva, nasal mucosa, prostatic fluid, lavage fluid, semen, lymph, bile, tears, sweat, breast milk, mammary gland fluid, etc., or combinations thereof. In some embodiments, the biological sample is a cervical swab from the subject. In some embodiments, the biological sample may be blood, and sometimes plasma or serum. The term "blood" as used herein refers to a blood sample or product from a pregnant woman or a woman being tested for a possible pregnancy. The term encompasses whole blood, blood products, or any portion of blood, such as serum and plasma as conventionally defined, tannins, etc. Blood or portions thereof often include nucleosomes (e.g., maternal and / or fetal nucleosomes). Nucleosomes include nucleic acids and are sometimes cell-free or intracellular. Blood also includes a tannin. The tannin is sometimes separated using a Ficoll gradient. The tannin may include leukocytes (e.g., white blood cells, T cells, B cells, platelets, etc.). In some embodiments, the tannin includes maternal and / or fetal nucleic acids. Blood plasma refers to the portion of whole blood obtained by centrifugation of blood treated with an anticoagulant. Blood serum refers to the liquid aqueous layer retained after a blood sample has clotted. Liquid or tissue samples are typically collected according to standard methods followed in hospitals or clinical practice. In the case of blood, an appropriate amount of peripheral blood (e.g., 3-40 ml) is typically collected and preserved according to standard procedures before or after preparation. The liquid or tissue sample used for nucleic acid extraction can be cell-free (e.g., cell-free). In some embodiments, the liquid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, the sample may contain fetal cells or cancer cells.

[0049] The sample may be a liquid sample. Liquid samples may include extracellular nucleic acids (e.g., circulating cell-free DNA). Non-limiting examples of liquid samples include blood or blood products (e.g., serum, plasma, etc.), cord blood, chorionic villus, amniotic fluid, cerebrospinal fluid, cerebrospinal fluid, lavage fluid (e.g., bronchoalveolar, gastric, peritoneal, catheter, ear, arthroscopy), biopsy samples (e.g., liquid biopsies for cancer detection), intermembranous fluid samples, female reproductive tract lavage fluid, urine, feces, sputum, saliva, nasal mucosa, prostatic fluid, irrigation fluid, semen, lymph, bile, tears, sweat, breast milk, mammary gland fluid, etc., or combinations thereof. In some embodiments, the sample is a liquid biopsy, which typically refers to assessing the presence of disease, or progression or remission of disease (e.g., cancer), in a liquid sample from a subject. Liquid biopsies may be used in conjunction with or as an alternative to solid biopsies (e.g., tumor biopsies). In some examples, extracellular nucleic acids are analyzed in the liquid biopsy.

[0050] The samples are typically heterogeneous, meaning they contain more than one type of nucleic acid material. For example, heterogeneous nucleic acids can include, but are not limited to, (i) cancer and non-cancer nucleic acids, (ii) pathogen and host nucleic acids, (iii) fetal and maternal nucleic acids, and / or more commonly, (iv) mutant and wild-type nucleic acids. Samples can be heterogeneous because they contain more than one cell type, such as fetal and maternal cells, cancer and non-cancer cells, or pathogen and host cells. In some embodiments, a few and a majority of nucleic acid materials are present.

[0051] For the prenatal application of the techniques described herein, liquid or tissue samples may be collected from females of suitable gestational age for testing or from females who are suspected of being pregnant. Suitable gestational age may vary depending on the prenatal test performed. In some embodiments, the pregnant female subject is sometimes in the first trimester, sometimes in the second trimester, or sometimes in the last trimester. In some embodiments, the liquid or tissue is collected from pregnant women whose fetuses are approximately 1–approximately 45 weeks pregnant (e.g., fetal gestational ages 1–4, 4–8, 8–12, 12–16, 16–20, 20–24, 24–28, 28–32, 32–36, 36–40, or 40–44 weeks) and sometimes from approximately 5–approximately 28 weeks pregnant (e.g., fetal gestational ages 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 weeks). In some implementations, fluid or tissue samples are collected from pregnant females during or immediately after childbirth (e.g., vaginal or non-vaginal delivery, such as operative delivery) or 0-72 hours after delivery.

[0052] Obtaining blood samples and DNA extraction

[0053] In some embodiments, the method of the present invention includes isolating, enriching and / or analyzing DNA found in the blood of a subject as a non-invasive means of detecting chromosomal alterations in the subject's genome and / or monitoring the subject's health.

[0054] Obtaining blood samples

[0055] Blood samples can be obtained from subjects of any age (e.g., male or female) using the methods of this invention. Blood samples can be obtained from pregnant women of gestational age suitable for testing using the methods described in this invention. The appropriate age of pregnancy may vary depending on the disease being tested, as described below. Blood collection from subjects (e.g., pregnant women) is typically performed according to standard protocols generally followed in hospitals or clinics. An appropriate amount of peripheral blood is collected, typically 5-50 ml, and preserved according to standard procedures before further preparation. The blood samples can be collected, preserved, or transported in a manner that minimizes degradation of the nucleic acids present in the sample or ensures their quality.

[0056] Preparation of blood samples

[0057] DNA found in the blood of a subject is analyzed using, for example, whole blood, serum, or plasma. Fetal DNA found in maternal blood is analyzed using, for example, whole blood, serum, or plasma. Methods for preparing serum or plasma from blood obtained from a subject (e.g., a maternal subject) are known. For example, blood from a subject (e.g., a pregnant woman) can be placed in a tube containing EDTA to prevent blood clotting or a commercially available product such as Vacutainer SST (Becton Dickinson, Franklin Lake, New Jersey), and plasma can then be obtained from the whole blood by centrifugation. Serum can be obtained, with or without centrifugation after blood clotting. If centrifugation is used, it is typically (but not limited to) performed at a suitable speed (e.g., 1,500–3,000 g). Plasma or serum may undergo additional centrifugation steps before being transferred to a new tube for DNA extraction.

[0058] In addition to the non-cellular portion of whole blood, DNA can also be recovered from the cellular components and enriched in the tannin layer, which can be obtained by centrifuging a woman's whole blood sample and removing the plasma.

[0059] DNA extraction

[0060] There are several known methods for extracting DNA from biological samples, including blood. These can be done using standard DNA preparation methods (e.g., described in Sambrook and Russell, *Molecular Cloning: A Laboratory Manual*, 3rd edition, 2001); or using a variety of commercially available reagents or kits, such as Qiagen's QIAamp Cyclic Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qiaagen, Heldon, Germany), and Genome Prep. TM Blood DNA Separation Kit (Promega, Madison, Wisconsin) and GFX TM The Genomic Blood DNA Purification Kit (Amersham, Piscateway, NJ) can also be used to obtain DNA from a subject's blood sample. Combinations of more than one of these methods can also be used.

[0061] In some embodiments, samples obtained from pregnant female subjects may be enriched or enriched relative to fetal nucleic acids using one or more methods. For example, the differentiation between fetal and maternal DNA may be performed using the compositions and methods described in this invention alone or in combination with other differentiating factors. Examples of such factors include, but are not limited to, single nucleotide differences in chromosomes X and Y, chromosome Y-specific sequences, polymorphisms elsewhere in the genome, size differences between fetal and maternal DNA, and differences in methylation forms between maternal and fetal tissues.

[0062] Other methods for enriching samples with specific nucleic acid material are described in PCT patent application No. PCT / US07 / 69991, filed May 30, 2007; PCT patent application No. PCT / US2007 / 071232, filed June 15, 2007; U.S. Provisional Applications Nos. 60 / 968,876 and 60 / 968,878 (as assigned to the applicant); and PCT patent application No. PCT / EP05 / 012707, filed November 28, 2005, all of which are incorporated herein by reference. In some embodiments, the parent nucleic acid is selectively removed (partially, substantially, almost completely, or completely) from the sample.

[0063] The terms “nucleic acid” and “nucleic acid molecule” are used interchangeably herein. The term refers to nucleic acids in any composite form, derived from, for example: DNA (e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), RNA (e.g., messenger RNA (mRNA), short repressor RNA (siRNA), ribosomal RNA (rRNA), tRNA, microRNA, RNA highly expressed in the fetus or placenta, etc.), and / or DNA or RNA analogs (e.g., containing base analogs, sugar analogs, and / or non-natural backbones, etc.), RNA / DNA hybrids, and polyamide nucleic acids (PNAs), all of which may be in single-stranded or double-stranded form, and unless otherwise specified, may encompass known analogs of natural nucleotides that function in a manner similar to naturally occurring nucleotides. In some embodiments, nucleic acids may be or may be derived from: plasmids, bacteriophages, viruses, autonomously replicating sequences (ARS), centromeres, artificial chromosomes, chromosomes, or other nucleic acids capable of replicating or being replicated in vitro or in host cells, cells, the nucleus of cells, or the cytoplasm of cells. In some implementations, the template nucleic acid may be derived from a single chromosome (e.g., the nucleic acid sample may be derived from a chromosome of a sample obtained from a diploid organism). Unless explicitly defined, the term covers known analogs containing a reference nucleic acid with similar binding properties and metabolized in a manner similar to that of naturally occurring nucleotides. Unless otherwise stated, a particular nucleic acid sequence also includes its conserved modified variants (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences, as well as explicitly indicated sequences. Specifically, degenerate codon substitutions can be obtained by producing a sequence in which the third position of one or more selected (or all) codons is substituted with a mixture of bases and / or deoxyinosine residues. The term nucleic acid is used interchangeably with locus, gene, cDNA, and mRNA encoded by a gene. The term may also include equivalents, derivatives, variants, and analogs of RNA or DNA synthesized from nucleotide analogs, single-stranded ("sense" or "antisense", "positive" or "negative", "positive" reading frame or "reverse" reading frame), and double-stranded polynucleotides. The term “gene” refers to a segment of DNA involved in the production of a polypeptide chain; it includes regions before and after the coding region (leader and tail regions) involved in the transcription / translation of the gene product and the regulation of said transcription / translation, as well as insertion sequences (introns) between individual coding segments (exons).

[0064] Deoxyribonucleotides include deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine. For RNA, the cytosine base is replaced with uracil. Template nucleic acids can be prepared using nucleic acids obtained from the target organism as templates.

[0065] Nucleic acid isolation and processing

[0066] Nucleic acids can be obtained from one or more sample sources (such as cells, serum, plasma, ochre layer, lymph, skin, soil, etc.) using methods known in the art. DNA can be isolated, extracted, and / or purified from biological samples (e.g., from blood or blood products) using any suitable method. Non-limiting examples include methods for DNA preparation (e.g., described in Sambrook and Russell, *Molecular Cloning: A Laboratory Manual*, 3rd edition, 2001); various commercially available reagents or kits, such as Qiagen's QIAamp Cyclic Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qagen, Heldon, Germany), and Genome Prep. TM Blood DNA Separation Kit (Promega, Madison, Wisconsin) and GFX TM Genomic blood DNA purification kit (Amersham, Piscavenge, NJ) or combinations thereof.

[0067] Cell lysis methods and reagents are known in the art and can generally be performed by chemical (e.g., detergents, hypotonic solutions, enzymatic processes, etc., or combinations thereof), physical (e.g., French pressure filtration, sonication, etc.), or electrolytic lysis methods. Any suitable lysis process can be used. For example, chemical methods typically use a lysing agent to disrupt cells and extract nucleic acids from them, followed by treatment with a dissociative salt. Physical methods, such as freezing / thawing followed by grinding, and cell pressure filtration, are also useful. High-salt lysis is also commonly used. For example, alkaline lysis can be employed. The latter method conventionally involves the use of a phenol-chloroform solution, and alternatively, a phenol-chloroform-free method comprising three solutions can be used. In the latter method, one solution may contain 15 mM Tris, pH 8.0; 10 mM EDTA and 100 μg / ml RNase A; the second solution may contain 0.2N NaOH and 1% SDS; and the third solution may contain 3 M KOAc, pH 5.5. These methods can be found in sections 6.3.1–6.3.6 (1989) of the *Current Protocols in Molecular Biology*, published by John Wiley & Sons, Inc., New York, which are included in full in this paper.

[0068] Nucleic acids can also be isolated at different time points than other nucleic acids, with each sample originating from the same or different sources. Nucleic acids can be derived from nucleic acid libraries, such as cDNA or RNA libraries. Nucleic acids can be products of nucleic acid purification or isolation and / or amplification of nucleic acid molecules in a sample. Nucleic acids provided for the methods described herein can comprise nucleic acids from one sample or from two or more samples (e.g., from 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more samples).

[0069] In some embodiments, nucleic acids may include extracellular nucleic acids. As used herein, the term "extracellular nucleic acid" refers to nucleic acids isolated from a substantially cell-free source, and is also referred to as "cell-free" nucleic acids, "circulating cell-free nucleic acids" (e.g., CCF fragments), and / or "cell-free circulating nucleic acids." Extracellular nucleic acids may be present in and obtained from blood (e.g., from human blood, such as from the blood of a pregnant woman). Extracellular nucleic acids typically do not contain detectable cells and may contain cellular elements or cellular remnants. Non-limiting examples of cell-free sources of extracellular nucleic acids include blood, plasma, serum, and urine. As used herein, the term "obtaining circulating cell-free sample nucleic acid" includes obtaining a sample directly (e.g., collecting a sample, such as a test sample) or obtaining a sample from a person who has already collected a sample. Without being theoretically limited, extracellular nucleic acids can be products of apoptosis and cell lysis, which often results in extracellular nucleic acids having a range of lengths (e.g., "ladders").

[0070] In some implementations, extracellular nucleic acids may contain different nucleic acid substances, and are therefore referred to herein as "heterogeneity." For example, the blood serum or plasma of a person with cancer may contain nucleic acids from cancer cells and nucleic acids from non-cancer cells. In another example, the blood serum or plasma of a pregnant female may contain maternal nucleic acids and fetal nucleic acids. In some examples, fetal nucleic acids sometimes account for about 5% to about 50% of all nucleic acids (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49% of the total nucleic acids are fetal nucleic acids). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 500 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 500 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 250 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 250 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 200 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 200 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 150 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 150 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 100 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 100 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 50 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 50 base pairs or less).In some implementations, the majority of the fetal nucleic acid in the nucleic acid is about 25 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% of the fetal nucleic acid length is about 25 base pairs or less).

[0071] In some embodiments, nucleic acids may be provided for performing the methods described herein without processing the nucleic acid-containing sample. In some embodiments, nucleic acids are provided for performing the methods described herein after processing the nucleic acid-containing sample. For example, nucleic acids may be extracted, isolated, purified, partially purified, or amplified from the sample. As used herein, the term "isolation" means the removal of nucleic acids from their original environment (e.g., the natural environment in which nucleic acids are naturally produced or the host cell expressing exogenous nucleic acids), thus altering the nucleic acids from their original environment through human intervention (e.g., "artificial"). As used herein, the term "isolated nucleic acid" refers to nucleic acids removed from an object (e.g., a human object). Isolated nucleic acids may contain fewer non-nucleic acid components (e.g., proteins, lipids) compared to the component content present in the source sample. Compositions containing isolated nucleic acids may be about 50% to more than 99% free of non-nucleic acid components. Compositions containing isolated nucleic acids may be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more than 99% free of non-nucleic acid components. As used herein, the term "purified" refers to nucleic acids containing fewer non-nucleic acid components (e.g., proteins, lipids, carbohydrates) compared to the amount of non-nucleic acid components present prior to the purification process. Compositions containing purified nucleic acids may be approximately 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other non-nucleic acid components. The term "purified" as used herein may also refer to nucleic acids containing fewer nucleic acid substances compared to the sample source from which they are derived. Compositions containing purified nucleic acids may be approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other nucleic acid substances. For example, fetal nucleic acids can be purified from a mixture containing maternal and fetal nucleic acids. In some embodiments, small fragments of fetal nucleic acid (e.g., 30-500 bp fragments) can be purified, or partially purified, from a mixture containing fragments of both fetal and maternal nucleic acids. In some examples, nucleosomes containing smaller fragments of fetal nucleic acid can be purified from a mixture of macronucleosome complexes containing larger fragments of maternal nucleic acid. In some examples, cancer cell nucleic acids can be purified from a mixture of cancer cell and non-cancer cell nucleic acids. In some examples, nucleosomes containing small fragments of cancer cell nucleic acid can be purified from a mixture of macronucleosome complexes containing larger fragments of non-cancer nucleic acids.

[0072] In some embodiments, nucleic acids are sheared or cleaved before, during, or after the method of the present invention. The sheared or cleaved nucleic acids may have a nominal, average, or mean length of about 5 to about 10,000 base pairs, about 100 to about 1,000 base pairs, about 100 to 500 base pairs, or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs. Nucleic acids can be cleaved or split using suitable methods known in the art, and the average, geometric mean, or nominal length of the resulting nucleic acid fragments can be controlled by selecting an appropriate fragment generation method.

[0073] In some embodiments, nucleic acids may be sheared or cleaved by suitable methods, non-limiting examples of which include physical methods (e.g., shearing, sonication, French press filtration, heat, ultraviolet irradiation, etc.), enzyme processing (e.g., enzyme cleavage reagents (e.g., suitable nucleases, suitable restriction enzymes, suitable methylation-sensitive restriction enzymes)), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, alkaline hydrolysis, heat, etc. or combinations thereof), the methods described in U.S. Patent Application Publication 20050112590, etc., or combinations thereof.

[0074] As used herein, “shearing” or “cleavage” refers to a method or condition that allows a nucleic acid molecule (such as a nucleic acid template gene molecule or its amplification product) to be divided into two or more smaller nucleic acid molecules. This shearing or cleavage can be sequence-specific, base-specific, or non-specific, and can be accomplished by any of a variety of methods, reagents, or conditions (including, for example, chemical, enzymatic, or physical shearing, such as physical fragmentation). As used herein, “cleavage product,” “shearing product,” or its grammatical variations, refers to nucleic acid molecules obtained by shearing or cleaving nucleic acids or their amplification products.

[0075] As used herein, the term "amplification" refers to the process of generating amplicon nucleic acids in a processed sample in a linear or exponential manner, the nucleotide sequence of which is identical or substantially identical to the nucleotide sequence of the target nucleic acid or a segment thereof. In some embodiments, the term "amplification" refers to methods including polymerase chain reaction (PCR). For example, the amplification product can contain one or more additional nucleotides than the amplified nucleotide region of the nucleic acid template sequence (e.g., primers can contain "extra" nucleotides, such as transcription initiation sequences, in addition to nucleotides complementary to the nucleic acid template gene molecule, to generate an amplification product containing "extra" nucleotides or nucleotides not corresponding to the amplified nucleotide region of the nucleic acid template gene molecule).

[0076] As used herein, the term "complementary shearing reaction" refers to a shearing reaction on the same nucleic acid using different shearing agents or by altering the shearing specificity of the same shearing agent, thereby producing different shearing patterns of the same target or reference nucleic acid or protein. In some embodiments, nucleic acids may be treated in one or more reaction vessels using one or more specific shearing agents (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more specific shearing agents). As used herein, the term "specific shearing agent" refers to a reagent, sometimes a chemical or enzyme that can cleave nucleic acids at one or more specific sites.

[0077] Before being used in the methods described herein, nucleic acids may be treated to modify certain nucleotides within them. For example, the nucleic acid may be subjected to a treatment that selectively modifies the nucleic acid based on the methylation state of the nucleotides. Furthermore, conditions such as high temperature, ultraviolet radiation, and X-ray radiation can induce variations in the nucleic acid molecular sequence. Nucleic acids may be provided in any suitable form for appropriate sequence analysis.

[0078] Nucleic acids can be single-stranded or double-stranded. For example, single-stranded DNA can be generated by denaturing double-stranded DNA through heating or, for example, treatment with an alkali. In some embodiments, the nucleic acid is a D-loop structure, formed by the invasion of an oligonucleotide or DNA-like molecule, such as a peptide nucleic acid (PNA), into the middle strand of a double-stranded DNA molecule. Adding *E. coli* RecA protein and / or altering the salt concentration (e.g., using methods known in the art) facilitates the formation of D-loops.

[0079] Minority and majority of substances

[0080] At least two different nucleic acid substances may be present in varying amounts in extracellular (e.g., circulating cell-free) nucleic acids, sometimes referring to a minority substance and a majority substance. In some examples, the minority substance of nucleic acids originates from affected cell types (e.g., cancer cells, wasting cells, cells attacked by the immune system). In some embodiments, chromosomal alterations are determined for the minority nucleic acid substance. In some embodiments, chromosomal alterations are determined for the majority nucleic acid substance. The terms "minority" and "majority" are not intended to be strictly defined in any respect. In one aspect, a nucleic acid considered to be a "minority" may, for example, have an abundance of at least about 0.1% to less than 50% of the total nucleic acids in the sample. In some embodiments, the abundance of a minority nucleic acid may be at least about 1% to about 40% of the total nucleic acids in the sample. In some embodiments, the abundance of a minority nucleic acid may be at least about 2% to about 30% of the total nucleic acids in the sample. In some embodiments, the abundance of a minority nucleic acid may be at least about 3% to about 25% of the total nucleic acids in the sample. For example, the abundance of a few nucleic acids can be approximately 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, or 30% of the total nucleic acids in the sample. In some examples, a minority of extracellular nucleic acids sometimes constitutes about 1% to about 40% of all nucleic acids (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40% of the nucleic acids are minority nucleic acids). In some embodiments, the minority nucleic acids are extracellular DNA. In some embodiments, the minority nucleic acids are extracellular DNA derived from apoptotic tissue. In some embodiments, the minority nucleic acids are extracellular DNA derived from tissues affected by cell proliferation disorders. In some embodiments, the minority nucleic acids are extracellular DNA derived from tumor cells. In some embodiments, the minority nucleic acids are extracellular fetal DNA.

[0081] In another embodiment, the abundance of nucleic acids considered "majority" can be, for example, more than 50% to about 99.9% of the total nucleic acids in the sample. In some embodiments, the abundance of majority nucleic acids can be at least about 60% to about 99% of the total nucleic acids in the sample. In some embodiments, the abundance of majority nucleic acids can be at least about 70% to about 98% of the total nucleic acids in the sample. In some embodiments, the abundance of majority nucleic acids can be at least about 75% to about 97% of the total nucleic acids in the sample. For example, the abundance of the majority nucleic acids may be at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acids are extracellular DNA. In some embodiments, the majority nucleic acids are extracellular maternal DNA. In some embodiments, the majority nucleic acids are DNA from healthy tissue. In some embodiments, the majority nucleic acids are DNA from non-tumor cells.

[0082] In some embodiments, the length of a minority of the extracellular nucleic acid is about 500 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the length of the minority nucleic acid is about 500 base pairs or less). In some embodiments, the length of a minority of the extracellular nucleic acid is about 300 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the length of the minority nucleic acid is about 300 base pairs or less). In some embodiments, the length of a minority of the extracellular nucleic acid is about 200 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the length of the minority nucleic acid is about 200 base pairs or less). In some embodiments, the length of a minority of the extracellular nucleic acid is about 150 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the length of the minority nucleic acid is about 150 base pairs or less).

[0083] Cell types

[0084] As used herein, “cell type” refers to a class of cells that can be distinguished from other cell types. Extracellular nucleic acids may include nucleic acids from several different cell types. Non-limiting examples of cell types that can provide nucleic acids to circulating cell-free nucleic acids include hepatocytes (e.g., hepatocytes), lung cells, spleen cells, pancreatic cells, colon cells, skin cells, bladder epithelial cells, eye cells, brain cells, esophageal cancer cells, head cells, neck cells, ovarian cells, testicular cells, prostate cells, placental cells, epithelial cells, endothelial cells, adipocytes, kidney / renal cells, heart cells, muscle cells, blood cells (e.g., leukocytes), central nervous system (CNS) cells, and combinations thereof. In some embodiments, cell types that can provide nucleic acids to the circulating cell-free nucleic acids being analyzed include leukocytes, endothelial cells, and hepatocytes. Different cell types may be screened for the identification and selection of portions of nucleic acid loci, wherein the marker status of cell types in subjects with a medical condition is the same or substantially the same as the marker status of cell types in subjects without said medical condition, as detailed below.

[0085] The specific cell types in objects with the medical condition may sometimes remain the same or substantially the same as those in objects without the medical condition. In a non-limiting example, the number of live or viable cells of a specific cell type may be reduced in cell degeneration conditions, and in objects with the medical condition, the number of live or viable cells may be unchanged or not significantly changed.

[0086] Specific cell types sometimes change as part of a medical condition, exhibiting one or more different properties compared to their original state. In non-limiting examples, a specific cell type may proliferate at an above-normal rate, transform into cells with different morphologies, transform into cells expressing one or more different cell surface markers, and / or become part of a tumor as part of a cancer condition. In embodiments where a specific cell type (i.e., a progenitor cell) changes as part of a medical condition, the labeling status of each of the one or more markers tested is generally the same or substantially the same as that of a specific cell type in an object having the medical condition compared to a specific cell type in an object not having the medical condition. Therefore, the term "cell type" sometimes refers to a cell type in an object not having the medical condition and to a variant of a cell in an object having the medical condition. In some embodiments, "cell type" refers only to a progenitor cell and not a variant produced by a progenitor cell. "Cell type" sometimes refers to a progenitor cell and the variant cell produced by that progenitor cell. In this embodiment, the labeling status of the analyzed markers is generally the same or substantially the same as that of a cell type in an object having the medical condition compared to a cell type in an object not having the medical condition.

[0087] In some implementations, the cell type is a cancer cell. Certain cancer cell types include, for example, leukemia cells (such as acute myeloid leukemia, acute lymphoblastic leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia); cancerous kidney / renal cells (e.g., renal cell carcinoma (clear cell, papillary type 1, papillary type 2, chromophobe, eosinophil, collecting duct), renal adenocarcinoma, adrenoidoma, Wilm's tumor, transitional cell carcinoma); and brain tumor cells (such as lateral auditory neuroma, astrocytoma (Grade I: pilocytic astrocytoma, Grade II: ...). Low-grade astrocytoma, Grade III: Anaplastic astrocytoma, Grade IV: Glioblastoma (GBM), chordoma, central nervous system lymphoma, craniopharyngioma, glioma (brainstem glioma, ependymoma, mixed glioma, optic nerve glioma, subependymal tumor), medulloblastoma, meningioma, brain metastases, oligodendroglioma, pituitary adenoma, primitive neuroectodermal (PNET), schwannoma, juvenile pilocytic astrocytoma (JPA), pineal tumor, rhabdoid tumor).

[0088] Different cell types can be distinguished by any suitable characteristics, including but not limited to one or more different cell surface markers, one or more different morphological features, one or more different functions, one or more different protein (e.g., histone) modifications, and one or more different nucleic acid markers. Non-limiting examples of nucleic acid markers include single nucleotide polymorphisms (SNPs), methylation status of nucleic acid loci, short tandem repeats, insertions (e.g., microinsertions), deletions (microdeletions), and combinations thereof. Non-limiting examples of protein (e.g., histone) modifications include acetylation, methylation, ubiquitination, phosphorylation, SUMOylation, and combinations thereof.

[0089] As used herein, the term "related cell type" refers to a cell type that shares multiple characteristics with other cell types. In related cell types, 75% or more of the cell surface markers are sometimes identical to those of the cell type in question (e.g., approximately 80%, 85%, 90%, or 95% or more of the cell surface markers are identical to those of the related cell type).

[0090] Enrichment and isolation of nucleic acid subgroups

[0091] In some embodiments, nucleic acids (e.g., extracellular nucleic acids) are enriched or relatively enriched against nucleic acid subsets or substances. Nucleic acid subsets may include, for example, fetal nucleic acids, maternal nucleic acids, nucleic acids containing fragments of a specific length or length range, or nucleic acids derived from specific genomic regions (e.g., a single chromosome, a set of chromosomes, and / or certain chromosomal regions). Such enriched samples may be used in conjunction with the methods described herein. The methods described herein sometimes involve isolating, enriching, and analyzing fetal DNA found in maternal blood as a non-invasive means of detecting the presence of maternal and / or fetal chromosomal alterations. Therefore, in some embodiments, the technique includes the additional step of enriching nucleic acid subsets, such as fetal nucleic acids, in a sample. In some embodiments, the methods described herein for determining fetal fractions can also be used to enrich fetal nucleic acids. In some embodiments, maternal nucleic acids are selectively removed (partially, substantially, almost completely, or completely) from the sample. In some embodiments, enriching specific low-copy-number nucleic acids (e.g., fetal nucleic acids) can improve quantitative sensitivity. Methods for enriching specific types of nucleic acids in samples, such as those described below, are incorporated herein by reference in U.S. Patent No. 6,927,028, International Application Publication No. WO2007 / 140417, International Application Publication No. WO2007 / 147063, International Application Publication No. WO2009 / 032779, International Application Publication No. WO2009 / 032781, International Application Publication No. WO2010 / 033639, International Application Publication No. WO2011 / 034631, International Application Publication No. WO2006 / 056480 and International Application Publication No. WO2011 / 143659.

[0092] In some embodiments, a subset of nucleic acid fragments is selected prior to sequencing. In some embodiments, hybridization-based techniques (e.g., using oligonucleotide arrays) may be used to first select nucleic acid sequences from certain chromosomes (e.g., sex chromosomes and / or chromosomes suspected of containing chromosomal alterations). In some embodiments, nucleic acids may be separated by size (e.g., by gel electrophoresis, size exclusion chromatography, or by microfluidic-based methods), while in some examples, fetal nucleic acids may be enriched by selecting those with lower molecular weights (e.g., less than 300 base pairs, less than 200 base pairs, less than 150 base pairs, less than 100 base pairs). In some embodiments, fetal nucleic acids may be enriched by suppressing maternal background nucleic acids (e.g., by adding formaldehyde). In some embodiments, a portion or subset of the preselected nucleic acid fragment set is randomly sequenced. In some embodiments, the nucleic acids are amplified prior to sequencing. In some embodiments, a portion or subset of the nucleic acids is amplified prior to sequencing.

[0093] Nucleic acid library

[0094] In some embodiments, a nucleic acid library is a variety of polynucleotide molecules (e.g., nucleic acid samples) prepared, assembled, and / or modified for a specific process, non-limiting examples of which include immobilization, enrichment, amplification, cloning, detection, and / or use for nucleic acid sequencing on a solid phase (e.g., a solid support, such as a flow cell, beads). In some embodiments, the nucleic acid library is prepared before or during the sequencing process. Nucleic acid libraries (e.g., sequencing libraries) can be prepared using suitable methods known in the art. Nucleic acid libraries can be prepared via targeted or non-targeted preparation processes.

[0095] In some embodiments, the nucleic acid library is modified to include chemical portions (e.g., functional groups) configured to immobilize nucleic acids to a solid support. In some embodiments, the nucleic acid library is modified to include biomolecules (e.g., functional groups) and / or binding pair members configured to immobilize the library to a solid support. Non-limiting examples include thyroxine-binding globulin, steroid-binding proteins, antibodies, antigens, haptens, enzymes, hemagglutinins, nucleic acids, inhibitors, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding proteins, receptors, carbohydrates, oligonucleotides, polynucleotides, complementary nucleic acid sequences, and combinations thereof. Examples of specific binding pairs include, but are not limited to: anti-biotin moieties and biotin moieties; antigenic epitopes and antibodies or their immunologically active fragments; antibodies and haptens; digoxigenin moieties and anti-digoxigenin antibodies; luciferin moieties and anti-luciferin antibodies; operons and inhibitors; nucleases and nucleosides; lectins and polysaccharides; steroids and steroid-binding proteins; active compounds and active compound receptors; hormones and hormone receptors; enzymes and substrates; immunoglobulins and protein A; oligonucleotides or polynucleotides and their corresponding complements; and combinations thereof.

[0096] In some embodiments, the nucleic acid library is modified to include one or more polynucleotides of known composition. Non-limiting examples include identifiers (e.g., tags, index tags), capture sequences, labeled adaptors, restriction enzyme sites, promoters, enhancers, origins of replication, stem-loops, complementary sequences (e.g., primer binding sites, annealing sites), suitable integration sites (e.g., transposons, viral integration sites), modified nucleotides, and combinations thereof. Polynucleotides of known sequences can be inserted at suitable positions, such as the 5′ end, 3′ end, or within the nucleic acid sequence. Polynucleotides of known sequences can be the same or different sequences. In some embodiments, polynucleotides of known sequences are configured to hybridize with one or more oligonucleotides immobilized on a surface (e.g., the surface of a flow cell). For example, a known 5′ sequence of a nucleic acid molecule can hybridize with a first plurality of oligonucleotides, while a known 3′ sequence can hybridize with a second plurality of oligonucleotides. In some embodiments, the nucleic acid library may include chromosome-specific tags, capture sequences, labels, and / or adaptors. In some embodiments, the nucleic acid library includes one or more detectable markers. In some embodiments, one or more detectable markers may be incorporated into the 5′ end, 3′ end, and / or any nucleotide position of the nucleic acid in the library. In some embodiments, the nucleic acid library includes hybridized oligonucleotides. In some embodiments, the hybridized oligonucleotides are labeled probes. In some embodiments, the nucleic acid library includes hybridized oligonucleotide probes before immobilization on a solid phase.

[0097] In some embodiments, the polynucleotide of a known sequence includes a universal sequence. A universal sequence is a specific nucleotide sequence that integrates into two or more nucleic acid molecules or subsets of two or more nucleic acid molecules, wherein the universal sequence is identical with respect to all the molecules or subsets into which it is integrated. Universal sequences are typically designed to hybridize and / or amplify multiple different sequences using a single universal primer complementary to the universal sequence. In some embodiments, two (e.g., a pair) or more universal sequences and / or universal primers are used. Universal primers typically include a universal sequence. In some embodiments, an adaptor (e.g., a universal adaptor) includes a universal sequence. In some embodiments, one or more universal sequences are used to capture, identify, and / or detect multiple nucleic acid substances or subsets thereof.

[0098] In some embodiments of nucleic acid library preparation (e.g., in certain sequencing processes of a synthesis procedure), the size of the nucleic acids is selected and / or fragmented to a length of several hundred base pairs or less (e.g., in library generation preparation). In some embodiments, library preparation is not required (e.g., when using ccfDNA).

[0099] In some implementations, a linker-based library preparation method is used. Non-limiting examples of linker-based library preparation methods and kits include TRUSEQ or ScriptMiner, Illumina, San Diego, CA; KAPA Laboratory Preparation Kit, KAPA Biosystems, Woburn, MA; NEBNext, NEB Biolabs, Evian Pool, MA; and MuSeek, Thermo Fisher Scientific, Waterham, MA. DNA Sample Preparation Kits (Lucigen, Michelton, Wisconsin; PureGenome, EMD Millipore, Bill Ricard, MA, etc.). Ligation-based library preparation methods typically utilize adaptor design, which incorporates an index sequence at the initial ligation step and is commonly used to prepare samples for single-read sequencing, paired-read sequencing, and multiplex sequencing. For example, sometimes nucleic acids (e.g., fragmented nucleic acids or ccfDNA) are end-repaired via fill-in reactions, endonuclease reactions, or combinations thereof. In some embodiments, the resulting blunt-end-repaired nucleic acid can then be extended by a single nucleotide, which is complementary to the 3'-terminal single nucleotide overhang of the adaptor / primer. Any nucleotide can be used for the extended / overhanging nucleotide. In some embodiments, nucleic acid library preparation includes linking adaptor oligonucleotides. Adaptor oligonucleotides are typically complementary to flow cell anchors and are sometimes used to immobilize the nucleic acid library to a solid support, such as the inner surface of a flow cell. In some embodiments, the adaptor oligonucleotide includes an identifier, one or more sequencing primer hybridization sites (e.g., sequences complementary to universal sequencing primers, single-end sequencing primers, paired-end sequencing primers, multiplex sequencing primers, etc.) or combinations thereof (e.g., adaptor / sequencing, adaptor / identifier, adaptor / identifier / sequencing).

[0100] The identifier may be a suitable detectable tag incorporating or conjugating a nucleic acid (e.g., a polynucleotide), which allows for the detection and / or identification of nucleic acids including the identifier. In some embodiments, the identifier is incorporating or conjugating the nucleic acid during sequencing methods (e.g., by polymerase). Non-limiting examples of identifiers include nucleic acid tags, nucleic acid indexes or barcodes, radioactive tags (e.g., isotopes), metallic tags, chemiluminescent tags, phosphorescent tags, fluorescence quenchers, dyes, proteins (e.g., enzymes, antibodies or portions thereof, linkers, members of binding pairs), and combinations thereof. In some embodiments, the identifier (e.g., a nucleic acid index or barcode) is a unique, known, and / or identifiable sequence of a nucleotide or nucleotide analogue. In some embodiments, the identifier is six or more consecutive nucleotides. Many fluorophores with various different excitation and emission spectra are available. Any suitable type and / or number of fluorophores can be used as identifiers. In some embodiments, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, or fifty or more different identifiers are used in the methods described herein (e.g., nucleic acid detection and / or sequencing methods). In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are linked to each nucleic acid in the library. Identifier detection and / or quantification can be performed by suitable methods, devices, or machines, non-limiting examples of which include flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, illuminometer, fluorometer, spectrophotometer, suitable gene chip or microarray analysis, Western blotting, mass spectrometry, chromatography, cellular fluorescence analysis, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning flow cytometry, affinity chromatography, manual batch separation, electric field suspension, suitable nucleic acid sequencing methods and / or nucleic acid sequencing devices (e.g., sequencers, such as sequencers), and combinations thereof.

[0101] In some implementations, transposon-based library preparation methods are used (e.g., EPICENTRE NEXTERA, Epicentre, Madison, Wisconsin). Transposon-based methods typically use in vitro translocation to similar fragments or tagged DNA (often allowing for the inclusion of platform-specific tags and optional barcodes) in a single-tube reaction to prepare a sequencer-ready library.

[0102] In some embodiments, a nucleic acid library or a portion thereof is amplified (e.g., by PCR-based methods). In some embodiments, sequencing methods include amplifying a nucleic acid library. The nucleic acid library may be amplified before or after immobilization onto a solid support (e.g., a solid support in a flow cell). Nucleic acid amplification includes the process of amplifying or increasing (e.g., in the nucleic acid library) the amount of a present nucleic acid template and / or its complement, said process being achieved by generating one or more copies of the template and / or its complement. Amplification may be performed by suitable methods. The nucleic acid library may be amplified by thermal cycling or by isothermal amplification. In some embodiments, rolling circle amplification is used. In some embodiments, amplification occurs on a solid support (e.g., within a flow cell) where a nucleic acid library or a portion thereof is immobilized. In some sequencing methods, the nucleic acid library is added to a flow cell and immobilized by hybridization with an anchor under suitable conditions. Such nucleic acid amplification is generally referred to as solid-phase amplification. In some embodiments of solid-phase amplification, all or part of the amplification product is synthesized by extension starting from immobilized primers. Solid-phase amplification reactions are similar to standard solution-phase amplification, except that at least one of the amplified oligonucleotides (e.g., primers) is immobilized on a solid support.

[0103] In some embodiments, solid-phase amplification includes a nucleic acid amplification reaction comprising only one oligonucleotide primer immobilized on a surface. In some embodiments, solid-phase amplification includes multiple different immobilized oligonucleotide primer materials. In some embodiments, solid-phase amplification may include a nucleic acid amplification reaction comprising one oligonucleotide primer immobilized on a solid surface and a second different oligonucleotide primer in solution. Multiple different immobilized or solution primers may be used. Non-limiting examples of solid-phase nucleic acid amplification reactions include interfacial amplification, bridging amplification, emulsion PCR, WildFire amplification (e.g., U.S. Patent Application US20130012399), and combinations thereof.

[0104] sequencing

[0105] In some embodiments, nucleic acids (e.g., nucleic acid fragments, sample nucleic acids, cell-free nucleic acids) may be sequenced. In some embodiments, a full or nearly full sequence is obtained, and sometimes a partial sequence is obtained. Sequencing, localization, and correlation analysis methods are as described herein or known in the art (e.g., U.S. Patent Application Publication US2009 / 0029377, incorporated herein by reference). Some aspects of such methods are described below.

[0106] Any suitable method for sequencing nucleic acids can be used, non-limiting examples of which include Maxim and Gilbert, chain termination methods, synthesis sequencing, ligation sequencing, mass spectrometry sequencing, microscopy-based techniques, and combinations thereof. In some embodiments, first-generation sequencing technologies, such as Sanger sequencing methods, including automated Sanger sequencing methods (including microfluidic Sanger sequencing), can be used in the methods of the present invention. In some embodiments, other sequencing technologies, including nucleic acid imaging techniques (such as transmission electron microscopy (TEM) and atomic force microscopy (AFM)), are also used herein. In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods typically involve clonal amplification of DNA templates or single DNA molecules, sometimes sequenced in massively parallel in a flow cell. Next-generation (e.g., second and third generation) sequencing technologies (capable of sequencing DNA in massively parallel) can be used in the methods described herein and are collectively referred to herein as “massively parallel sequencing” (MPS). In some embodiments, MPS sequencing methods employ a targeted approach, in which specific chromosomes, genes, or regions of interest produce sequence reads. Specific chromosomes, genes, or regions of interest herein sometimes refer to target genomic regions. In some implementations, non-targeted methods are used, in which most or all nucleic acid fragments (e.g., CCF fragments, CCF DNA, polynucleotides) in the sample are sequenced, amplified, and / or randomly captured.

[0107] MPS sequencing sometimes uses sequencing via synthesis and certain imaging methods. Nucleic acid sequencing technologies that can be used in the methods described herein are synthetic sequencing and reversible terminator-based sequencing (such as Illumina's Genome Analyzer and Genome Analyzer II; HISEQ 2000; HISEQ 2500 (Illumina, San Diego, California)). This technology allows for the parallel sequencing of millions of nucleic acid (such as DNA) fragments. In one embodiment of this sequencing technology, a flow cell containing an optically clear slide with eight individual channels is used, the surface of which is bound with oligonucleotide anchors (such as adaptor primers).

[0108] In some embodiments, synthetic sequencing involves repeatedly adding (e.g., covalently adding) nucleotides to primers or a pre-existing nucleic acid chain in a template-guided manner. Each repeatedly added nucleotide is detected, and the process is repeated multiple times until a sequence of the nucleic acid chain is obtained. The length of the obtained sequence depends in part on the number of addition and detection steps performed. In some embodiments of synthetic sequencing, one, two, three, or more nucleotides of the same type (e.g., A, G, C, or T) are added and detected in a nucleotide addition round. Nucleotides can be added by any suitable method (e.g., enzymatic or chemical). For example, in some embodiments, polymerases or ligases add nucleotides to primers or a pre-existing nucleic acid chain in a template-guided manner. In some embodiments of synthetic sequencing, different types of nucleotides, nucleotide analogs, and / or identifiers are used. In some embodiments, reversible terminators and / or removable (e.g., cleavable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and / or nucleotide analogs are used. In some embodiments, synthetic sequencing includes cutting (e.g., cutting and removing identifiers) and / or washing steps. In some embodiments, the addition of one or more nucleotides is detected by methods described herein or known in the art. Non-limiting examples include any suitable imaging device, a suitable camera, a digital camera, a CCD (charge-coupled device) based imaging device (e.g., a CCD camera), a CMOS (complementary metal-oxide-semiconductor) based imaging device (e.g., a CMOS camera), a photodiode (e.g., a photomultiplier tube), an electron microscope, a field-effect transistor (e.g., a DNA field-effect transistor), an ISFET ion sensor (e.g., a CHEMFET sensor), and combinations thereof. Other sequencing methods that can be used to perform the methods described herein include digital PCR and hybridization sequencing.

[0109] The methods described herein can be used to obtain nucleic acid sequencing reads by employing suitable non-human MPS methods, systems, or technology platforms. Non-limiting examples of MPS platforms include Illumina / Solex / HiSeq (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ), SOLiD, Roche / 454, PACBIO and / or SMRT, HelicosTrue single-molecule sequencing, Ion Torrent-based and Ion semiconductor-based sequencing (e.g., developed by Life Technologies), technologies based on WildFire, 5500, 5500xl W and / or 5500xl W genetic analyzers (e.g., developed and marketed by Life Technologies, US patent application US20130012399); Polony sequencing, Pyro sequencing, massively parallel signature sequencing (MPSS), RNA polymerase (RNAP) sequencing, LaserGen systems and methods, nanopore-based platforms, chemically sensitive field-effect transistor (CHEMFET) arrays, electron microscopy-based sequencing (e.g., developed by ZS Genetics, Halcyon Molecular), nanosphere sequencing, etc.

[0110] Other sequencing methods that can be used to perform the methods described herein include digital PCR and hybridization sequencing. Digital polymerase chain reaction (digital PCR or dPCR) can be used to directly identify and quantify nucleic acids in a sample. In some embodiments, digital PCR can be performed in an emulsion. For example, individual nucleic acids are isolated in, for example, a microfluidic device and each nucleic acid is amplified individually by PCR. Nucleic acids are isolated such that no more than one nucleic acid is contained in each well. In some embodiments, different probes can be used to distinguish multiple alleles (e.g., fetal alleles and maternal alleles). Alleles can be counted to determine copy number.

[0111] In some embodiments, hybridization sequencing can be used. The method involves contacting multiple polynucleotide sequences with multiple polynucleotide probes, each of which is optionally attached to a substrate. In some embodiments, the substrate may be a plane with an array of known nucleotide sequences. The pattern of hybridization with the array can be used to determine the polynucleotide sequences present in the sample. In some embodiments, each probe is attached to a bead (such as a magnetic bead). Hybridization with the bead can be identified and used to identify multiple polynucleotide sequences in the sample.

[0112] In some implementations, nanopore sequencing can be used in the methods described herein. Nanopore sequencing is a single-molecule sequencing technology whereby a single nucleic acid molecule (such as DNA) is directly sequenced as it passes through a nanopore.

[0113] In some embodiments, chromosome-specific sequencing is performed. In some embodiments, chromosome-specific sequencing is performed using DANSR (Digital Analysis of Selected Regions). Digital analysis of selected regions can simultaneously quantify hundreds of sites by using interfering 'bridging' oligonucleotides to form PCR templates through cfDNA-dependent linkage of two site-specific oligonucleotides. In some embodiments, chromosome-specific sequencing is performed by generating a library enriched with chromosome-specific sequences. In some embodiments, sequence reads are obtained only for selected chromosome sets. In some embodiments, sequence reads are obtained only for chromosomes 21, 18, and 13.

[0114] In some embodiments, the sequence module acquires, generates, aggregates, assembles, processes, transforms, manipulates, and / or transfers sequence reads. The sequence module may utilize sequencing techniques known in the art to determine nucleic acid sequences. In some embodiments, the sequence module may align, assemble, fragment, complement, reverse complement, detect errors, or correct errors in sequence reads. In some embodiments, the sequence module provides sequence reads to a mapping module or any other suitable module.

[0115] Sequencing reads

[0116] As used herein, a “reading” (i.e., “a reading”, “sequence reading”) is a short nucleotide sequence generated by any sequencing method described herein or known in the art. Readings can be generated from one end of a polynucleotide fragment (“single-end reading”), while sometimes they are generated from both ends of a polynucleotide fragment (e.g., paired-end reading, double-end reading).

[0117] The length of sequence reads is typically related to the specific sequencing technology. For example, high-throughput methods provide sequence reads ranging in size from tens to hundreds of base pairs (bp). Nanopore sequencing, for instance, provides sequence reads ranging in size from tens to hundreds to thousands of base pairs. In some embodiments, sequence reads are arithmetic mean, median, average, or absolute lengths of approximately 15 bp to approximately 900 bp. In some embodiments, the sequence reads are arithmetic mean, median, average, or absolute lengths of approximately 1000 bp or longer.

[0118] The single-end reading can be of any suitable length. In some embodiments, the nominal, average, arithmetic average, or absolute length of the single-end reading is sometimes about 10 nucleotides to about 1000 consecutive nucleotides, about 10 nucleotides to about 500 consecutive nucleotides, about 10 nucleotides to about 250 consecutive nucleotides, about 10 nucleotides to about 200 consecutive nucleotides, about 10 nucleotides to about 150 consecutive nucleotides, about 15 consecutive nucleotides to about 100 consecutive nucleotides, about 20 consecutive nucleotides to about 75 consecutive nucleotides, or about 30 consecutive nucleotides or about 50 consecutive nucleotides. In some implementations, the nominal, average, arithmetic mean, or absolute length of the single-end readings is about 5, 6, 7, 8, 9, 10, 11, 12, 49, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more nucleotides.

[0119] Paired-end reads can be of any suitable length. In some embodiments, both ends are sequenced at a suitable read length sufficient to map each read (e.g., reads from both ends of a fragment template) to a reference genome. In some embodiments, the nominal, average, arithmetic mean, or absolute length of the paired-end reads is approximately 10 nucleotides to approximately 100 consecutive nucleotides, approximately 10 nucleotides to approximately 75 consecutive nucleotides, approximately 10 nucleotides to approximately 50 consecutive nucleotides, approximately 15 nucleotides to approximately 50 consecutive nucleotides, approximately 15 nucleotides to approximately 40 consecutive nucleotides, approximately 15 consecutive nucleotides to approximately 30 consecutive nucleotides, or approximately 15 consecutive nucleotides to approximately 20 consecutive nucleotides. In some implementations, the nominal, average, arithmetic mean, or absolute length of the paired end readings is about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more nucleotides.

[0120] Readings are typically representations of nucleotide sequences in physiological nucleic acids. For example, sequences are described using ATGC, where "A" represents adenine nucleotide, "T" represents thymine nucleotide, "G" represents guanine nucleotide, and "C" represents cytosine nucleotide. Sequence readings are usually obtained from nucleic acid samples from pregnant females carrying a fetus. Sequence readings from nucleic acid samples from pregnant females carrying a fetus are typically represented by sequence readings from the fetus and / or the fetal mother (e.g., the pregnant female subject).

[0121] Sequence reads obtained from the blood of pregnant females can be reads of a mixture of fetal and maternal nucleic acids. Mixtures of relatively short reads can be transformed, using the methods described herein, into a representation of genomic nucleic acids in the pregnant female and / or fetus. Mixtures of relatively short reads can be transformed into representations of, for example, chromosomal alterations. Reads of a mixture of maternal and fetal nucleic acids can be transformed into representations of a complex chromosome or segment thereof containing features of one or both maternal and fetal chromosomes. In some embodiments, “obtaining” nucleic acid sequence reads from a subject sample, and / or “obtaining” nucleic acid sequence reads from biological samples of one or more reference individuals, can directly involve sequencing nucleic acids to obtain sequence information. In some embodiments, “obtaining” can involve receiving sequence information directly obtained from other nucleic acids.

[0122] It has been observed that circulating cell-free nucleic acid fragments (CCF fragments) obtained from pregnant females generally contain nucleic acid fragments from fetal cells (i.e., fetal fragments) and nucleic acid fragments from maternal cells (i.e., maternal fragments). Sequence reads derived from CCF fragments from the fetus are referred to herein as “fetal reads.” Sequence reads derived from CCF fragments from the genome of a pregnant female carrying a fetus (e.g., the mother) are referred to herein as “maternal reads.” The CCF fragment from which fetal reads are derived is referred to herein as the fetal template, and the CCF fragment from which maternal reads are derived is referred to herein as the maternal template.

[0123] In some embodiments, the length of a polynucleotide fragment (e.g., a polynucleotide template) is determined. The length of the polynucleotide fragment in the sample, or the average or arithmetic mean length of the polynucleotide fragment, can be estimated and / or determined by suitable methods. In some embodiments, the length of the polynucleotide fragment in the sample, or the average or arithmetic mean length of the polynucleotide fragment, is determined using sequencing methods. In some embodiments, the fragment length is determined using a paired-end sequencing platform. Sometimes the length of the fragment template is determined by calculating the differences between the genomic coordinates of each mapped read assigned to paired-end reads. In some embodiments, the fragment length can be determined using sequencing methods, thereby obtaining the complete or substantially complete nucleotide sequence of the fragment. The sequencing methods include platforms that produce relatively long read lengths (e.g., Roche 454, Ion Torrent, Single Molecule (Pacific Biosciences), Real-Time SMRT technology, etc.).

[0124] In some implementations, a subset of readings is selected for analysis, while sometimes, certain portions of the readings are removed from the analysis. In some cases, selecting a subset of readings can enrich the types of nucleic acids (e.g., fetal nucleic acids). Enrichment of readings from fetal nucleic acids, for example, typically improves the accuracy of the methods described herein (e.g., chromosomal alteration detection). However, selecting and removing readings from the analysis typically reduces the accuracy of the methods described herein (e.g., due to increased variability). Therefore, without being theoretically limited, generally speaking, in methods that involve selecting and / or removing readings (e.g., fragments from a specific size range), a trade-off is required between the increased accuracy associated with enrichment of fetal readings and the decreased accuracy associated with a reduced number of reads. In some implementations, the method includes selecting a subset of readings that enrich readings from fetal nucleic acids without significantly reducing the accuracy of the method. Regardless of this apparent trade-off, it has been determined that, as described herein, employing a subset of nucleotide sequence readings (e.g., readings from relatively short fragments) can improve or maintain the accuracy of fetal genetic analysis. For example, in some implementations, about 80% or more of the nucleotide sequence readings may be discarded, and the sensitivity and specificity values ​​may be maintained at values ​​similar to those of a method that does not discard the nucleotide sequence readings.

[0125] In some embodiments, some or all nucleic acids in the sample are enriched and / or amplified before or during sequencing (e.g., non-specific, such as by PCR-based methods). In some embodiments, a specific portion or subset of nucleic acids in the sample is enriched and / or amplified before or during sequencing. In some embodiments, a portion or subset of a preselected set of nucleic acids is randomly sequenced. In some embodiments, nucleic acids in the sample are not enriched and / or amplified before or during sequencing.

[0126] In some embodiments, targeted enrichment, amplification, and / or sequencing methods are used. Targeting methods typically use sequence-specific oligonucleotides to isolate, select, and / or enrich subsets of nucleic acids (e.g., target genomic regions) in a sample for further processing. In some embodiments, the target genomic regions are associated with chromosomal alterations, including but not limited to translocations, insertions, additions, deletions, and / or inversions. In some embodiments, nucleic acid fragments from multiple target genomic regions are sequenced and / or analyzed. Polynucleotides (e.g., CCF DNA) derived from any suitable chromosome, its portions, or combinations thereof can be sequenced and / or analyzed using the methods or systems described herein, either using targeted or non-targeted methods. Non-limiting examples of chromosomes that can be analyzed using the methods or systems described herein include chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, X, and Y. In some embodiments, a library of sequence-specific oligonucleotides is employed to target (e.g., hybridize) one or more nucleic acid groups in a sample. Sequence-specific oligonucleotides and / or primers are typically selective for specific sequences (e.g., unique nucleic acid sequences) present in one or more regions of interest on chromosomes, genes, exons, introns, and / or regulatory regions. Any suitable method or combination of methods can be used to enrich, amplify, and / or sequence one or more subsets of target nucleic acids. In some embodiments, target sequences are isolated and / or enriched by capturing them to a solid phase (e.g., flow cell, beads) using one or more sequence-specific anchors. In some embodiments, target sequences are enriched and / or amplified using sequence-specific primers and / or primer sets via polymerase-based methods (e.g., PCR-based methods, via any suitable polymerase-based extension). Sequence-specific anchors are commonly used as sequence-specific primers.

[0127] In some implementations, partial sequencing of the genome is sometimes expressed as the amount of genome covered by the determined nucleotide sequence (e.g., less than 1 "fold" coverage). When sequencing the genome with approximately 1 fold coverage, the reading represents approximately 100% of the genome's nucleotide sequence. Genome sequencing can also be performed using redundancy, where a given region of the genome can be covered by two or more readings or overlapping readings (e.g., greater than 1 "fold" coverage). In some implementations, the genome is sequenced with a coverage of about 0.1-100 times, about 0.2-20 times, or about 0.2-1 times (e.g., about 0.02-, 0.03-, 0.04-, 0.05-, 0.06-, 0.07-, 0.08-, 0.09-, 0.1-, 0.2-, 0.3-, 0.4-, 0.5-, 0.6-, 0.7-, 0.8-, 0.9-, 1-, 2-, 3-, 4-, 5-, 6-, 7-, 8-, 9-, 10-, 15-, 20-, 30-, 40-, 50-, 60-, 70-, 80-, 90- times).

[0128] In some embodiments, sequence coverage is reduced without significantly reducing the accuracy (e.g., sensitivity and / or specificity) of the methods described herein. A significant reduction in accuracy can be a reduction of about 1% to about 20% compared to a method without reduced sequence read counts. For example, a significant reduction in accuracy can be a reduction of about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, or more. In some embodiments, sequence coverage and / or sequence read counts are reduced by about 50% or more. For example, sequence coverage and / or sequence read counts can be reduced by about 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more. In some embodiments, sequence coverage and / or sequence read counts are reduced by about 60% to about 85%. For example, sequence coverage and / or sequence read counts can be reduced by approximately 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, or 84%. In some embodiments, sequence coverage and / or sequence read counts can be reduced by removing certain sequence reads. In some examples, sequence reads from fragments longer than a specific length (e.g., fragments longer than approximately 160 bases) are removed.

[0129] In some implementations, one or more samples are sequenced during a sequencing run. Nucleic acids from different samples are typically identified by one or more unique identifiers or identifier tags. Sequencing methods typically utilize identifiers that allow for sequence reaction multiplication during sequencing. The sequencing process can be performed using any suitable number of samples and / or unique identifiers (e.g., 4, 8, 12, 24, 48, 96, or more).

[0130] Sequencing processes sometimes use a solid phase, which may include a flow cell on which nucleic acids from a library can be bound and on which reagents can flow and contact the bound nucleic acids. Flow cells sometimes include flow cell channels, and the use of identifiers facilitates the analysis of the number of samples in each channel. Flow cells are typically constructed to retain and / or allow reagent solutions to pass through in an orderly manner with bound analytes. Flow cells are typically planar, optically transparent, usually in the millimeter or sub-millimeter range, and often have channels or pathways in which analyte / reagent interactions occur. In some implementations, the number of samples analyzed in a given flow cell channel depends on the number of unique identifiers used in library preparation and / or probe design. Multiplexing 12 identifiers, for example, may allow the simultaneous analysis of 96 samples (e.g., the number of wells in a 96-well microplate) in an 8-channel flow cell. Similarly, multiplexing 48 identifiers, for example, may allow the simultaneous analysis of 384 samples (e.g., the number of wells in a 384-well microplate) in an 8-channel flow cell. Non-limiting examples of commercially available multiplex sequencing kits include Emindek's multiplex sample preparation oligonucleotide kit and multiplex sequencing primer and PhiX control kit (e.g., Emindek catalog numbers PE-400-1001 and PE-400-1002, respectively).

[0131] Mapped readings

[0132] Sequence reads or portions thereof (e.g., sequence read subsequences) can be mapped to and / or aligned to a reference sequence (e.g., a reference genome) using appropriate methods. The process of aligning one or more reads to a reference genome is called “mapping”. Sometimes the number of sequence reads mapped to a specific nucleic acid region (e.g., a chromosome, a portion thereof, or a segment thereof) can be quantified. Any suitable mapping method (e.g., procedure, algorithm, program, software, module, etc., or a combination thereof) can be used. Some aspects of mapping methods are described below.

[0133] Mapped nucleotide sequence reads (i.e., sequence information of fragments whose physical genomic sites are unknown) can be performed in a variety of ways, typically involving aligning the obtained sequencing reads with matching sequences in a reference genome. In this alignment, the sequence reads are usually compared to a reference sequence; those that are aligned are referred to as "mapped," "mapped sequence reads," or "mapped readings."

[0134] As used herein, the terms "alignment" and "alignment" refer to two or more nucleic acid sequences that can be identified as a match (e.g., 100% identical) or a partial match. Alignment can be performed manually (e.g., by visual observation) or by computer (e.g., software, programs, modules, or algorithms), and non-limiting examples include the ELAND (Effective Local Alignment of Nucleotide Data) computer program, which is part of the Illumina genome analysis workflow. Alignment of sequence readings can be 100% sequence match. In some cases, alignment is less than 100% sequence match (i.e., imperfect match, partial match, partial alignment). In some implementations, alignment is approximately 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75% match. In some implementations, the alignment includes mismatches. In some implementations, the alignment includes 1, 2, 3, 4, or 5 mismatches. Two or more sequences can be aligned using either strand. In some implementations, the nucleic acid sequence is aligned with the reverse complementary strand of another nucleic acid sequence.

[0135] Various computer methods can be used to map sequence reads to portions. Non-limiting examples of computer algorithms that can be used for sequence alignment include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE 1, BOWTIE 2, ELAND, MAQ, probe MATCH, SOAP, or SEQMAP, or variations thereof, or combinations thereof. In some embodiments, sequence reads or portions thereof may be aligned to sequences in a reference genome. In some embodiments, sequence reads may be obtained from and / or aligned to sequences in nucleic acid databases known in the art, including, for example, GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory), and DDBJ (Japan DNA Database). BLAST or similar tools can be used to search for identical sequences against sequence databases.

[0136] In some implementations, reads may be uniquely or non-uniquely mapped to a reference genome. A read that aligns to a single sequence in the reference genome is called a “unique mapping.” A read that aligns to two or more sequences in the reference genome is called a “non-unique mapping.” In some implementations, non-uniquely mapped reads are removed from further analysis (e.g., by filtering). In some implementations, a small degree of mismatch (0-1) may indicate a possible single nucleic acid polymorphism between the reference genome and the mapped reads from individual samples. In some implementations, no mismatches may allow reads to be mapped to a reference sequence.

[0137] As used herein, the term "reference genome" can refer to a genome or part thereof that is specifically known, sequenced, or characterized in whole or in part, of any organism or virus, and can be used as a reference to an identification sequence derived from the subject. For example, reference genomes of human subjects, bacteria, parasites, viruses, and many other organisms are available from the National Center for Biotechnology Information, World Wide Web Uniform Resource Locator: www.ncbi.nlm.nih.gov. In some embodiments, the reference genome is obtained from a reference sample or a set of reference samples. "Genome" refers to the complete genetic information of an organism or virus expressed as a nucleic acid sequence. As used herein, the term "reference sequence" refers to a reference genome or part thereof (e.g., chromosome, gene, conserved region, highly mappable region). Reference sequences are sometimes reference genomes or parts thereof. Reference genomes as used herein are often assembled or partially assembled genome sequences from one or more individuals. In some embodiments, the reference genome is an assembled or partially assembled genome sequence from one or more human individuals. In some embodiments, the reference genome includes sequences aligned to chromosomes. In some embodiments, the reference genome is a viral genome or part thereof. In some implementations, a reference genome of one or more viruses is used to align and / or map nucleic acids (e.g., sequence reads) obtained from human subjects (e.g., human samples).

[0138] In some embodiments, when the sample nucleic acid is derived from a pregnant female, the reference sequence may not be derived from the fetus, the fetal mother, or the fetal father, and is thus referred to herein as an "external reference." In some embodiments, a maternal reference may be prepared and used. When a reference from a pregnant female is prepared based on an external reference ("maternal reference sequence"), readings of DNA from the pregnant female that are substantially free of fetal DNA are typically mapped to and assembled from the external reference sequence. In some embodiments, the external reference is derived from DNA from an individual substantially of the same ethnicity as the pregnant female. The maternal reference sequence may not completely cover the maternal genomic DNA (e.g., it may cover approximately 50%, 60%, 70%, 80%, 90%, or more of the maternal genomic DNA), and the maternal reference may not be a perfect match to the maternal genomic DNA sequence (e.g., the maternal reference sequence may contain multiple mismatches).

[0139] Sequence reads can be mapped via a mapping module or a device including a mapping module, which typically maps reads to a reference genome or a segment thereof. The mapping module can map sequencing reads using suitable methods known in the art or methods described herein. In some embodiments, a mapping module or a device including a mapping module is required to provide mapped sequence reads. The mapping module typically includes suitable mapping and / or alignment procedures or algorithms.

[0140] Inconsistent readings

[0141] In some embodiments, this document provides methods for determining the presence of chromosomal alterations and for identifying breakpoints associated with chromosomal alterations. In some embodiments, methods for identifying breakpoints and / or chromosomal alterations include identifying inconsistent sequence reads. In some embodiments, methods include identifying states of inconsistency in sequences, sequencing reads, and / or sequencing read pairs (e.g., read buddy pairs). The term "inconsistency" herein refers to a state in which (i) a first read or a portion thereof maps to a first position in a reference genome and (ii) a second read or a portion thereof maps to or a second portion of the first read cannot be mapped, includes a lower mappability score, and / or maps to a second position in a reference genome, wherein the first and second positions in the reference genome are discontinuous and / or separated by a distance longer than that of a template polynucleotide fragment from which one or more sequence reads are obtained. In some embodiments, inconsistency refers to a state in which neither read (e.g., a pair of read buddies) of a sequencing read pair can be mapped. Inconsistent sequence reads and / or inconsistent sequence read pairs generally include inconsistency. Inconsistencies in polynucleotide sequences, sequence reads, single-end reads, and paired-end reads (e.g., paired-end reads) can be identified. In some embodiments, methods include identifying inconsistent reads and / or inconsistent read pairs. Inconsistent reads and / or inconsistent read pairs can be identified using suitable sequencing methods. In some embodiments, inconsistent read pairs are identified from paired-end sequencing reads. The terms “paired-end sequencing read” and “paired-end read” are used synonymously herein to refer to a sequencing read pair, wherein each member of the pair originates from the sequencing complementary strand of a polynucleotide fragment. Each read of a paired-end read is referred to herein as a “read buddy”.

[0142] Sequence reads and / or paired end reads are typically mapped to a reference genome using appropriate mapping and / or alignment procedures. Non-restrictive examples include BWA (Li H. and Durbin R. (2009) Bioinformatics 25, 1754–60), Novoalign [Novocraft (2010)], Bowtie (Langmead B, et al., (2009) Genome Biol. 10: R25), SOAP2 (Li R, et al., (2009) Bioinformatics 25, 1966–67), BFAST (Homer N, et al., (2009) PLoS ONE 4, e7767), GASSST (Rizk, G. and Lavenier, D. (2010) Bioinformatics 26, 2534–2540), and MPscan (Rivals E., et al., (2009) Lecture Notes in Computer Science). (5724, 246–260), etc. Sequence readings and / or paired end readings can be mapped and / or aligned using a suitable short reading alignment procedure. Non-restrictive examples of short readout alignment procedures include BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, BWA, CASHX, CUDA-EC, CUSHAW, CUSHAW2, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP, GeneiousAssembler, iSAAC, LAST, MAQ, mrFAST, mrsFAST, MOSAIK, MPscan, Novoalign, NovoalignCS, Novocraft, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RTG, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3, SOCS. SSAHA, SSAHA2, Stampy, SToRM, Subread, Subjunc, Taipan, UGENE, VelociMapper, TimeLogic, XpressAlign, ZOOM, and combinations thereof. Paired end reads are typically mapped to opposite ends of the same polynucleotide segment according to a reference genome. In some embodiments, sequence reads are mapped individually. In some embodiments, read buddies are mapped individually.In some embodiments, information from the two sequence reads (i.e., from each end) is factored within the mapping process. A reference genome is typically used to identify and / or infer nucleic acid sequences located between paired end read partners. The term “inconsistent read pair” herein refers to a pair of end reads containing a read partner pair, where one or both read partners do not clearly map to the same region of the reference genome partially defined by a continuous nucleotide segment. In some embodiments, an inconsistent read pair is a paired end read partner mapped to an unexpected site in the reference genome. Non-limiting examples of unexpected sites in the reference genome include (i) two different chromosomes, (ii) sites separated by a predetermined fragment size (e.g., greater than 300 bp, greater than 500 bp, greater than 1000 bp, greater than 5000 bp, or greater than 10000 bp), (iii) an orientation inconsistent with the reference sequence (e.g., reversed), etc., or combinations thereof. In some embodiments, read partners mapped and / or aligned to two different chromosomes are identified as inconsistent read partners. Read partners mapped and / or aligned to two different chromosomes are referred to herein as “chimeric read partners.” In some embodiments, inconsistent read pairs do not include read mate pairs, wherein a first read mate is mapped to a first chromosome and a second read mate is mapped to a second chromosome, and wherein the first chromosome is different from the second chromosome. In some embodiments, inconsistent read pairs include a first read mate mapped to a first segment of a reference genome (e.g., the first chromosome) and a second read mate mapped to a first segment of a reference genome (e.g., the first chromosome). The term "partial mapping" herein refers to 90% or less, 80% or less, 60% or less, 50% or less, 40% or less, 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, or 5% or less of the nucleotides read. In some embodiments, inconsistent read pairs include a first read mate mapped to a first segment of a reference genome (e.g., the first chromosome) and a second read mate that is unmappable and / or has low mappability (e.g., low mappability score). In some embodiments, inconsistent read pairs include a first read mate and a second read mate mapped to a first segment of a reference genome (e.g., the first chromosome), wherein the mappability of the second read mate or a portion thereof is undetermined. In some embodiments, inconsistent read pairs include a first read partner mapped to a first segment of a reference genome (e.g., a first chromosome) and a second read partner partially mapped to a second segment of the reference genome (e.g., a second chromosome), wherein the first and second segments are different (e.g., different chromosomes). In some embodiments, a subset (e.g., a set) of inconsistent read pairs includes first and second read partners mapped to different chromosomes, and first and second read partners to which one or both cannot be mapped and / or are partially mapped to the same or different chromosomes.In some embodiments, inconsistent read pairs are identified based on the length (e.g., average length, predetermined fragment size) or expected length of the template polynucleotide fragment in the sample. For example, read partners in the sample whose mapped positions are separated by a length greater than the average or expected length of the template polynucleotide fragment are sometimes identified as inconsistent read pairs. Read pairs mapped in opposite directions are sometimes determined by taking the reverse complement of one of the reads and comparing the alignment of the two reads with the same strand of a reference sequence. Inconsistent read pairs can be identified by any suitable method and / or algorithm known in the art or described herein. Inconsistent read pairs can be identified by an inconsistent read pair identification module or by a machine that includes an inconsistent read pair identification module, wherein the inconsistent read pair identification module generally identifies inconsistent read pairs. Non-limiting examples of inconsistent read pair identification modules include SVDetect, Lumpy, BreakDancer, BreakDancerMax, CREST, DELLY, etc., or combinations thereof. In some embodiments, inconsistent read pairs are not identified by algorithms that only identify read partners mapped or aligned to different chromosomes. In some implementations, inconsistent read pairs are identified by an algorithm for identifying a set of terminal read pairs mapped or aligned to different chromosomes, and one or both read pairs that cannot be mapped and / or are partially mapped to terminal read pairs on the same or different chromosomes. In some implementations, an inconsistent read identification module or a machine containing it is required to provide inconsistent read pairs.

[0143] In some implementations, inconsistent read pairs are not clustered or do not undergo clustering analysis. Clustering analysis, as described herein, refers to the process of combining paired end reads based on the mapping positions of one or two read partners within the genome. In some implementations, clustering analysis includes generating subsets of paired reads, wherein one read partner of each read pair is mapped to a first chromosome, and the other read partner of each read pair is mapped to a second chromosome, wherein the first chromosome is different from the second chromosome.

[0144] The term "inconsistent read" in this document refers to a sequence read in which a first portion of the read cannot be clearly mapped to the same region of a reference genome, which is a second portion of the read, wherein the same region of the reference genome is defined as a segment of consecutive nucleotides. In some embodiments, an inconsistent read includes first and second portions mapped to unexpected locations in the reference genome. In some embodiments, an inconsistent read includes a first portion mapped to a first chromosome and a second portion mapped to a second chromosome, wherein the first chromosome and the second chromosome are different. In some embodiments, an inconsistent read includes a portion partially mapped to a first segment of the reference genome (e.g., the first chromosome). In some embodiments, an inconsistent read includes a first portion mapped to a first segment of the reference genome (e.g., the first chromosome) and a second portion that is unmappable and / or has low mappability (e.g., low mappability score). In some embodiments, an inconsistent read includes a first portion and a second portion mapped to a first segment of the reference genome (e.g., the first chromosome), wherein the mappability of the second portion or a portion thereof is undetermined. In some embodiments, an inconsistent read includes a first portion mapped to a first segment of the reference genome (e.g., the first chromosome) and a second portion partially mapped to a second segment of the reference genome (e.g., the second chromosome), wherein the first and second segments are different (e.g., different chromosomes). Inconsistent readings can be identified using suitable methods and / or algorithms known in the art or described herein. Inconsistent readings are sometimes identified through a process characterizing the readings. Inconsistent readings can be identified by an inconsistency reading identification module or by a machine including such a module, where the module typically identifies inconsistent readings. In some embodiments, inconsistent readings are identified by an algorithm that identifies a subset or set of inconsistent readings. In some embodiments, an inconsistency reading identification module or a machine including it is required to provide the inconsistent readings.

[0145] Mappable changes

[0146] In some embodiments, the mappability of one or more reads is characterized. In some embodiments, characterizing the mappability of a read includes characterizing the mappability of multiple sequence read subsequences. The term “characterizing the mappability of multiple sequence read subsequences” sometimes refers herein to “characterizing a read.” Characterizing a read sometimes involves generating multiple sequence read subsequences of a read and mapping each sequence read subsequence to a reference genome. In some embodiments, sequence read subsequences are generated for two read partners of inconsistent read pairs. Sequence read subsequences sometimes refer herein to spurious reads. Sequence read subsequences can be generated by any suitable method. Sequence read subsequences are typically generated by a computer process. In some embodiments, sequence read subsequences are generated and mapped by: (i) mapping the read, (ii) removing one or more bases from the end of the mapped read by a computer process, (iii) mapping the resulting shortened read (i.e., the sequence read subsequence), and (v) repeating (ii) and (iii). In some embodiments, steps (ii) and (iii) are repeated until the read reaches the end. In some embodiments, steps (ii) and (iii) are repeated until the resulting shortened read can no longer be mapped. Sequence readout subsequences can be generated by progressively and / or gradually removing one or more bases from the 3' end or 5' end of the readout. Sequence readout subsequences can be generated by removing bases from the end of the readout (e.g., in step (ii)) at a time, one base at a time, two bases at a time, three bases at a time, four bases at a time, five bases at a time, or combinations thereof. In some embodiments, the lengths of the sequence readout subsequences generated for each readout are different. In some embodiments, the sequence readout subsequences of the readout are subsequences of the continuous nucleotides of the readout. In some embodiments, the sequence readout subsequences of the readout are shorter than the full-length readout, and sometimes the largest sequence readout subsequence is about 1 base, about 2 bases or less, about 3 bases or less, about 4 bases or less, about 5 bases or less, about 6 bases or less, about 7 bases or less, about 8 bases or less, about 9 bases or less, or about 10 bases or less. In some embodiments, each sequence reading subsequence of the reads is about 1 base, about 2 bases or less, about 3 bases or less, about 4 bases or less, about 5 bases or less, about 6 bases or less, about 7 bases or less, about 8 bases or less, about 9 bases or less, or about 10 bases or less than the second largest sequence reading subsequence or read. In some embodiments, each inconsistent read partner's sequence reading subsequence is progressively shorter than the second largest subsequence or read partner by about 1 base, about 2 bases or less, about 3 bases or less, about 4 bases or less, about 5 bases or less, about 6 bases or less, about 7 bases or less, about 8 bases or less, about 9 bases or less, about 10 bases or less, or combinations thereof.The term "sequence readout subsequence" in this document refers to a group of polynucleotide fragments generated by a computer for a readout from the methods described herein. "Sequence readout subsequence" as used herein can also refer to one or more groups of polynucleotide fragments generated by a computer for one or more readouts. "Sequence readout subsequence" as used herein can also refer to and / or include groups from which full-length readouts from which sequence readout subsequences are generated. Sequence readout subsequence is sometimes referred to as "subsequence" herein. When used in the singular, "sequence readout subsequence" and / or "subsequence" as used herein refers to a member of a group of polynucleotide fragments generated by a computer for a readout.

[0147] Sequence readout subsequences can be generated by suitable modules, programs, or methods. In some implementations, sequence readout subsequences are generated by a fragmentation module. Sequence readout subsequences can be generated by suitable mapping modules, programs, or methods. Non-limiting examples include BWA (Li H. and Durbin R. (2009) Bioinformatics 25, 1754–60), Novoalign [Novocraft (2010)], Bowtie (Langmead B, et al., (2009) Genome Biol. 10: R25), SOAP2 (Li R, et al., (2009) Bioinformatics 25, 1966–67), BFAST (Homer N, et al., (2009) PLoSONE 4, e7767), GASSST (Rizk, G. and Lavenier, D. (2010) Bioinformatics 26, 2534–2540), and MPscan (Rivals E., et al., (2009) Lecture Notes in Computer Science). (5724,246–260), etc.

[0148] In some embodiments, some or all of the sequence read subsequences generated for a sample are mapped to a reference genome. In some embodiments, sequence read subsequences of a read are mapped to a combination of one or more reference genomes. Subsequences of one or more reads are typically mapped to a human reference genome. In some embodiments, subsequences are mapped to a human genome and / or a viral genome. In some embodiments, the mappability of some or all of the sequence read subsequences of a read is determined. The term "mappability" refers to a measure of the extent to which a polynucleotide fragment maps to a reference genome. Sometimes mappability includes a mappability score or value. Sometimes a mappability score or value is determined for each sequence read subsequence of a read (e.g., a read companion). Any suitable mappability score or value can be determined for a sequence read subsequence. A mappability score can be any suitable score generated by a known or described mapping module, procedure, or method, or a combination thereof. For example, in some embodiments, the mapping score can be a MAPQ score. In some embodiments, the mappability score includes an alignment score. For example, alignment scores can be generated using a suitable local alignment algorithm (e.g., the Smith-Waterman algorithm), or they can be generated based on the Euclidean distance between two sequences weighted by the number of occurrences in a reference genome. Alignment scores can be generated using any suitable measure of the uniqueness of a quantitative or qualitative polynucleotide sequence. Criteria for determining high, good, acceptable, low, unacceptable, and / or poor mappability scores are known in the art, and they are generally specific to the mapping or alignment procedure used.

[0149] In some embodiments, determining the presence of a chromosomal alteration includes identifying and / or characterizing a mappable change in a sequence read subsequence of the read (e.g., an inconsistent read, a read partner of an inconsistent read pair). In some embodiments, characterizing a read includes identifying and / or characterizing a mappable change in a sequence read subsequence of the read. Mappable changes are sometimes determined among one or more sequence read subsequences. For example, sometimes a mappable change is indicated by one or more subsequences containing one or more nucleotides on a first side of the read containing a first subset with substantially the same or similar mappable fractions, and one or more subsequences containing nucleotides on a second side of the read containing a second subset with mappable fractions substantially different from the first subset. A mappable change identified in a sequence read subsequence of a read sometimes refers to a mappable change in the read itself. For example, a read may contain a mappable change where the sequence read subsequence of the read includes a mappable change. In some embodiments, inconsistencies may be identified in reads for which mappable changes have been determined. In some embodiments, inconsistent reads include mappable changes. In some embodiments, mappable changes may be determined for reads for which inconsistencies have been determined. In some implementations, mappability variations in the readings are not identified and / or determined, and the readings do not include inconsistencies. Sometimes one or both of the reading partners of an inconsistent reading pair include mappability variations. In some implementations, inconsistent reading pairs are identified, wherein one or both of the reading partners of a paired end-reading pair include mappability variations. In some implementations, mappability variations may be determined for one or both of the reading partners of an inconsistent reading pair. In some implementations, one or both of the reading partners of a nested reading pair do not include mappability variations.

[0150] Sometimes determining and / or identifying the mappability variations of a sequence of readings involves determining the relationship between each sequence of readings and suitable characteristics of each sequence of readings (e.g., inconsistent readings, inconsistent reading pairs, reading companions of inconsistent reading pairs). The relationship between one or more readings can be determined. For example, in some embodiments, the relationship between two reading companions of an inconsistent reading pair is determined.

[0151] The term "relationship" in this document refers to a mathematical and / or geometric relationship between two or more variables or values. Non-limiting examples of relationships include mathematical or geometric representations: functions, correlations, distributions, linear or non-linear equations, lines, regressions, fitted regressions, and combinations thereof. In some implementations, determining a relationship involves generating a linear, non-linear, or fitted relationship. In some implementations, the relationship is plotted or graphed.

[0152] Non-limiting examples of suitable characteristics of sequence readout subsequences that can be used to determine relationships include fragment length, identifiers indicating the relative order of the fragments, molecular weight, GC content, etc., or combinations thereof. In some embodiments, determining and / or identifying the mappability variations of sequence readout subsequences for a read includes determining the relationship between the length of each sequence readout subsequence of the read (e.g., one or both of the read companions) and mappability.

[0153] In some embodiments, determining and / or identifying a change in mappability includes determining whether a change exists in one or more coefficients, variables, values, constants, etc., or combinations thereof, that describe and / or quantify the relationship. As used herein, “change” sometimes means “distinction.” Non-limiting examples of coefficients, constants, values, and variables that can be used to determine a change in mappability include slopes (e.g., the slope of a linear, non-linear, or fitted relationship), sums, means, medians, or arithmetic means of coordinates (e.g., x-coordinate or y-coordinate values), intercepts (e.g., y-intercept), maximum values ​​(e.g., maximum peak height), minimum values ​​(e.g., minimum values), line integrals (e.g., area under a curve), etc., or combinations thereof. In some embodiments, a change in mappability is determined mathematically. A change can be determined by appropriate statistical tests of significance (e.g., significant difference), non-limiting examples of which include Wilcoxon tests (e.g., Wilcoxon signed-rank), t-tests, X-ray tests, etc. 2 Verification, etc. In some embodiments, mappability variations are visually identified and / or determined (e.g., from a graph or image). In some embodiments, mappability variations include determining the arithmetic mean, median, or mean of the mappability variations of one or more readings. For example, sometimes mappability variations include determining the arithmetic mean, median, or mean of the mappability variations of the two reading partners of an inconsistent reading pair. In some embodiments, mappability variations include the arithmetic mean, median, or mean slope of a first relation generated for a subsequence of a first reading and a second relation generated for a subsequence of a second reading (e.g., the first and second reading partners of an inconsistent reading pair). Mappability variations can be generated, determined, and / or identified by any suitable module, system, or software. Mappability variations are typically identified and / or determined by a mapping characterization module. A mapping characterization module can generate subsequences, generate relations, characterize the mappability of subsequences, determine mappability variations, receive or generate mappability thresholds, and / or compare mappability variations with mappability thresholds.

[0154] In some implementations, the mapping representation module includes microprocessor instructions (e.g., algorithms) in the form of code and / or source code (e.g., a set of standard or custom scripts) and / or one or more software packages (e.g., statistical software packages) that perform the mapping representation module functions. In some implementations, the mapping representation module includes code (e.g., scripts) written in S or R using a suitable package (e.g., an S package or an R package). For example, the slope of the change in the mapping representation can be calculated in R and may include a script like the one described below:

[0155] lm(y~x)[[“coefficients”]][2]

[0156] Where y is the MAPQ score of each stepwise alignment and x is the length of the stepwise alignment. In some implementations, the mapping characterization module includes and / or uses a suitable statistical software package. Non-limiting examples of statistical software packages include statistical packages in S-plus, Stata, SAS, MATLAB, R, etc., or combinations thereof.

[0157] In some embodiments, a change in the mappability of readings and / or subsets of readings may be determined. In some embodiments, a change in the mappability of readings and / or subsets of readings identified as inconsistent may be determined. In some embodiments, one or more readings (e.g., subsets of readings, subsets of inconsistent reading pairs) are selected and / or identified based on the change in the mappability of a sequence of reading subsequences. In some embodiments, identifying and / or determining a change in mappability includes identifying and / or determining a significant difference (e.g., a statistical difference) in the mappability between two or more subsequences or subsets of readings. In some embodiments, one or more readings (e.g., subsets of readings) are identified and / or selected based on a change in mappability and / or a mappability threshold. In some embodiments, inconsistent reading pairs are selected based on a change in the mappability of one or both of the readings and / or a mappability threshold. In some embodiments, selecting readings includes comparing a change in mappability to a mappability threshold. In some embodiments, selected readings include readings with a change in mappability higher than, lower than, within, outside of, significantly different from, or substantially the same as a mappability threshold. Typically, one or more readings are selected, wherein the mappability variations of one or more readings include numerical values ​​(e.g., quantitative values) that are significantly different from a predetermined mappability threshold or fall outside a predetermined range of values ​​defined by the mappability threshold. Readings (e.g., a subset of readings) may be identified and / or selected according to a reading selection module (e.g., 120). In some embodiments, the reading selection module includes microprocessor instructions (e.g., algorithms) in the form of code and / or source code (e.g., a set of standard or custom scripts) and / or one or more software packages (e.g., statistical software packages) that perform mapping characterization module functions. In some embodiments, the reading selection module includes code (e.g., scripts) written in S or R using a suitable package (e.g., S package, R package). For example, to select readings where the average slope of companion 1 and companion 2 is below a threshold 0, it can be written in R as follows:

[0158] data[data<0]

[0159] The data includes the arithmetic mean slopes of companion 1 and companion 2. In some implementations, the reading selection module includes and / or uses a suitable statistical software package. Non-limiting examples of statistical software packages include S-plus, Stata, SAS, MATLAB, the Tongji package in R, Prism (GraphPad Software, La Jolla, CA), SigmaPlot (Systat Software, San Jose, CA), Microsoft Excel (Raymond, Washington, USA), etc., or combinations thereof.

[0160] Mappability thresholds typically include one or more predetermined values, value limits, and / or value ranges. The term "threshold" or "valence" refers to any number calculated using a set of data that meets the requirements and used as a constraint for selection. Mappability thresholds are typically calculated through mathematical and / or statistical operations on mappability changes.

[0161] In some embodiments, a subset of readings is identified and / or selected, wherein the readings in the subset comprise a slope change of a relationship determined between the mappability of multiple subsequences of each reading and the fragment length. In some embodiments, the slope change indicates the presence of a candidate breakpoint. In some embodiments, a slope significantly greater than or less than 1 (e.g., a mappability threshold of 1) of the relationship generally indicates that the readings include a candidate breakpoint or that a mappability change has been identified in the subset of readings. In some embodiments, the mappability change (e.g., slope) contained in the identified and / or selected readings (e.g., the subset of readings) is approximately 0, about 0.1, about 0.2, about 0.3, about 0.4, about 0.5, about 0.6, about 0.7, about 0.8, about 0.9, or about 1.0 greater than the mappability threshold. In some embodiments, the identified and / or selected readings (e.g., a subset of readings) contain mappability variations (e.g., slopes) that exceed a mappability threshold range of about -0.1 to about 0.1, about -0.2 to about 0.2, about -0.3 to about 0.3, about -0.4 to about 0.4, about -0.5 to about 0.5, about -0.6 to about 0.6, about -0.7 to about 0.7, about -0.8 to about 0.8, about -0.9 to about 0.9, or about -1.0 to about 1.0. In some embodiments, mappability variations (e.g., average, arithmetic mean, or median slope) and / or mappability thresholds are expressed as absolute values. The threshold may be any suitable parameter indicating mappability variation (e.g., the difference between one or more mappability scores, standard deviation, or MAD).

[0162] Characterizing and / or evaluating the mappability variations of reads (e.g., subsequences of reads) can provide one or more breakpoints and / or candidate breakpoints. The term "breakpoint" herein refers to a location between two adjacent base determinations of a read partner, wherein bases on a first side of the breakpoint map to a first chromosomal region and bases on a second side of the breakpoint map to a second chromosomal region, wherein the first and second chromosomal regions are not adjacent according to a reference genome. In some embodiments, the first and second chromosomal regions are on different chromosomes. In some embodiments, the first and second chromosomal regions are on the same chromosome, wherein the first and second chromosomal regions are not adjacent according to a reference genome. In some embodiments, the term "breakpoint" herein refers to a location between two adjacent base determinations of a read partner, wherein bases on a first side of the location map to a reference genome and bases on a second side of the location cannot be mapped (e.g., cannot be mapped at a certain level). In some embodiments, the term "breakpoint" herein refers to a location between two adjacent base determinations of a read partner, wherein bases on a first side of the location map to a human genome and bases on a second side of the location map to heterologous genomic material (e.g., a viral genome). In some embodiments, breakpoints indicate the location and / or position of a chromosomal alteration or a portion thereof. In some embodiments, breakpoint identification is based on the nucleic acid position of a reference genome, where genetic material has been inserted, deleted, and / or exchanged. In some embodiments, when the chromosomal alteration includes insertion or translocation, the breakpoint may indicate the location and / or position of one side of the insertion or translocation. In some embodiments, a first breakpoint of an insertion or translocation is identified in a first reading or a first subset of readings, and a second breakpoint of an insertion or translocation is identified in a second reading or a second subset of readings. The term "candidate breakpoint" herein refers to a reading and / or position that may include a breakpoint. In some embodiments, candidate breakpoints include breakpoints. In some embodiments, candidate breakpoints do not include breakpoints. Candidate breakpoints are typically included in readings and / or subsets of readings identified and / or selected based on mappability variations and / or mappability thresholds.

[0163] In some embodiments, characterizing the mappability of multiple sequence read subsequences generated from the reads includes identifying and / or determining the location and / or position of candidate breakpoints. In some embodiments, the reads are representations of a genome (e.g., maternal genome, fetal genome, or a portion thereof). In some embodiments, reads containing mappings of candidate breakpoints indicate that the candidate breakpoints are located within the genome (e.g., maternal genome, fetal genome, or a portion thereof). In some embodiments, the location and / or position of candidate breakpoints in the reads are determined based on a relationship (e.g., a relationship between mappability and sequence length). In some embodiments, the location and / or position of candidate breakpoints in the reads are determined based on changes in mappability. In some embodiments, identifying and / or determining the location and / or position of candidate breakpoints in the reads includes identifying substantial differences (e.g., statistical differences) in the mappability between two or more sequence read subsequences of the read. In some embodiments, the position of a candidate breakpoint is determined at position x, wherein the mappability value of the sequence read subsequence on a first side of position x is substantially different from the mappability value of the sequence read subsequence on a second side of position x, thereby indicating a candidate breakpoint at position x. In some embodiments, the position of a candidate breakpoint is determined based on slope analysis. For example, a relationship is typically established between the mappability and fragment length of multiple subsequences, the relationship being partially defined by a line, and the line, or a portion thereof, being partially defined by a slope. In the foregoing example, a substantial change in the slope typically indicates the location of a candidate breakpoint (e.g., at position x, where the slope of the sequence on the first side of position x is substantially different from the slope of the sequence on the second side of position x). In some implementations, all readings containing the putative breakpoint (e.g., determined based on mappability changes and / or thresholds) are de novo organized, and the breakpoint is determined. Sometimes breakpoints are determined by comparing readings containing mappability changes with a reference genome. In some implementations, the location of candidate breakpoints and / or breakpoints is identified at a resolution of reading length. For example, mapped readings may include mappability changes that indicate a candidate breakpoint located at a position within a reference genome, wherein the readings are mapped. In some implementations, candidate breakpoints and / or breakpoint locations are identified at a resolution of 150 or fewer bases, 100 or fewer bases, 75 or fewer bases, 50 or fewer bases, 10 or fewer bases, 9 or fewer bases, 8 or fewer bases, 7 or fewer bases, 6 or fewer bases, 5 or fewer bases, 4 or fewer bases, 3 or fewer bases, 2 or fewer bases, or at a single base resolution.

[0164] In some implementations, candidate breakpoints are identified using a breakpoint module. The breakpoint module is typically configured to identify breakpoints using the methods described herein. In some implementations, the breakpoint module includes code (e.g., scripts) written in S or R using a suitable package (e.g., S package, R package). In some implementations, the breakpoint module includes and / or uses a suitable statistical software package. Any suitable de novo tissue, such as SOAP de novo tissue or those listed in Wikipedia (e.g., Wikipedia, Sequence Organization [online], [online 2013-09-25], retrieved from the Internet at World Wide Web Uniform Resource Locator: en.wikipedia.org / wiki / Sequence_assembly), can be used alone or in combination with custom scripts to identify the location of breakpoints. In some implementations, given the locations of mates 1 and 2, each reading is evaluated using R and / or one or more bioconductors in R (World Wide Web Uniform Resource Locator: bioconductor.org) based on its similarity to a human reference genome to determine the breakpoint. To determine the precise location of the break point using the slope, any suitable statistical software package or custom script can be used.

[0165] In some embodiments, one or both of the inconsistent read partners of an inconsistent read pair include substantially similar and / or identical candidate breakpoints. In some embodiments, one read partner of an inconsistent read pair includes a candidate breakpoint, and the other read partner of the pair does not include a candidate breakpoint. In some embodiments, the sequence of the first read partner of an inconsistent read pair overlaps with the sequence of the second read partner of the pair, and both read partners include identical or substantially similar candidate breakpoints. The term “substantially similar breakpoint” (e.g., “substantially similar candidate breakpoint”) refers to breakpoints located at the same or substantially identical locations in a reference genome. Substantially similar breakpoints are sometimes located at different relevant locations on different reads (e.g., typically determined relative to the end of the read), where the locations of the breakpoints on each read are substantially the same as those in the reference genome. Sometimes two or more reads (e.g., inconsistent read partners) include identical and / or substantially similar breakpoints, where the locations of the breakpoints on each read may be the same or different. In some embodiments, substantially similar breakpoints are located at the same locations on different reads. In some embodiments, the sequences of 1, 2, 3, 4, 5, 6, 7, or 8 or more nucleotides (e.g., base determination) on each side of substantially similar breakpoints are substantially identical. In some embodiments, substantially similar breakpoints are located on a first readout and a second readout, wherein the first readout is the reverse complement of the second readout.

[0166] In some implementations, a subset of readings is selected based on a change in mappability, wherein each reading in the selected subset comprises a minimum length of 20 consecutive bases, 21 consecutive bases, 22 consecutive bases, 23 consecutive bases, 24 consecutive bases, 25 consecutive bases, 26 consecutive bases, 27 consecutive bases, 28 consecutive bases, 29 consecutive bases, 30 consecutive bases, 31 consecutive bases, 32 consecutive bases, 33 consecutive bases, 34 consecutive bases, 35 consecutive bases, 36 consecutive bases, 37 consecutive bases, 38 consecutive bases, 39 consecutive bases, 40 consecutive bases, 50 consecutive bases, 60 consecutive bases, 70 consecutive bases, 80 consecutive bases, 90 consecutive bases, or 100 consecutive bases.

[0167] In some implementations, a subset of readings is selected based on a change in mappability, wherein each reading in the selected subset comprises at least about 10 to about 60, 15 to about 50, 15 to about 40, 15 to about 30, 15 to about 25, or about 15 to about 20 consecutive bases on each side of a candidate breakpoint.

[0168] In some embodiments, two or more readings in a sample (e.g., inconsistent reading companions) include substantially similar candidate breakpoints. In some embodiments, two or more, five or more, ten or more, 20 or more, 50 or more, 100 or more, or 1000 or more readings obtained from a sample include the same or substantially similar candidate breakpoints. In some embodiments, readings containing substantially similar candidate breakpoints, inconsistent reading companions, and / or inconsistent reading pairs are aggregated into subsets. In some embodiments, two or more subsets are identified and / or selected, wherein each subset includes readings containing substantially similar candidate breakpoints. In some embodiments, a first subset of readings and a second subset of readings include different breakpoints. Sometimes a first subset of readings containing substantially similar candidate breakpoints contains different breakpoints than a second subset of readings containing substantially similar candidate breakpoints. Typically, either subset of readings contains a candidate breakpoint that differs from the breakpoint of the other subset of readings.

[0169] In some embodiments, the systems or methods described herein are used to obtain and / or generate one or more subsets of readings containing substantially similar candidate breakpoints from a reference. The same or substantially the same methods are used to obtain and / or generate subsets of inconsistent reading companions containing substantially similar candidate breakpoints from the reference and test samples (e.g., test subjects). The term "reference" herein refers to one or more reference objects or reference samples. As used herein, "reference" generally refers to data (e.g., readings, subsets of inconsistent reading companions, selected read arrays) obtained from one or more reference objects or reference samples. Reference objects and / or reference samples are generally known or assumed to have no chromosomal alterations. For example, reference objects and / or reference samples typically do not contain chromosomal alterations. In some embodiments, the reference includes polynucleotide and / or nucleotide readings from specific genomic regions or multiple genomic regions that are not associated with chromosomal alterations.

[0170] Generate comparison

[0171] One or more subsets of readings generated and / or acquired from a test sample can be compared. Typically, a subset of readings from a sample containing substantially similar candidate breakpoints is compared to a subset of readings from a reference containing substantially similar candidate breakpoints. In some embodiments, a subset of readings from the sample is compared to a subset of readings from the reference, wherein the readings from both subsets (i.e., the subset from the sample and the subset from the reference) contain substantially similar candidate breakpoints. In some embodiments, a subset of readings from the sample is compared to a subset of readings from the reference, wherein the readings from both subsets (i.e., the subset from the sample and the subset from the reference) map to the same or substantially the same locations in the reference genome. Readings mapped to "substantially identical" positions in the reference genome refer to readings mapped within the following distances: 100,000 kb or less, 50,000 kb or less, 25,000 kb or less, 10,000 kb or less, 5,000 kb or less, 1,000 kb or less, 500 kb or less, 100 kb or less, 50 kb or less, 25 kb or less, 10 kb or less, 5 kb or less, 1,000 base pairs (bp) or less, 500 bp or less, or within 100 bp or less. Readings mapped to "substantially identical" positions in the reference genome refer to readings mapped within the following distances: 50, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, or 0 bases apart. For example, sometimes a subset of readings from a sample is compared with a subset of readings from a reference, where the two subsets (i.e., the subset from the sample and the subset from the reference) are selected based on a mappability variation, and the readings from the two subsets are mapped to the same or substantially the same locations in the reference genome. In some implementations, selected subsets of readings from the sample and the reference are identified based on a mappability variation, mapped to the same bin or portion of the genome, and compared. The same bin or pre-selected portion of the genome to which the readings are mapped can be of any suitable length. In some embodiments, the same bin or pre-selected portion of the genome to which the readings are mapped is approximately 100,000 kb or less, 50,000 kb or less, 25,000 kb or less, 10,000 kb or less, 5,000 kb or less, 1,000 kb or less, 500 kb or less, 100 kb or less, 50 kb or less, 25 kb or less, 10 kb or less, 5 kb or less, 1,000 base pairs (bp) or less, or approximately 500 bp or less. In some embodiments, the number of readings in one or more subsets is quantified and compared. In some embodiments, the number of readings in one or more subsets of the test sample is compared with the number of readings in one or more reference subsets.

[0172] Subsets of candidate breakpoints containing substantial similarities can be compared using appropriate statistical, geometric, or mathematical methods. In some embodiments, the comparison includes determining whether subsets of readings from the test sample and a reference are identical. In some embodiments, determining whether subsets of readings from the test sample and a reference are identical includes statistical analysis. In some embodiments, a determination is made when subsets of readings are compared and the subsets are substantially identical or substantially different. In some embodiments, a determination is made when the number of readings in the subsets is compared and the number of readings in the first subset and the second subset are statistically different or not statistically different. The terms “statistically different” and “statistically significant” herein refer to statistically significant differences. Statistically significant differences can be assessed using appropriate methods. Non-limiting examples of methods for determining statistical differences include determining and / or comparing Z-scores, distributions, correlations (e.g., correlation coefficients, t-tests, k-tests, etc.), uncertainty values, confidence measures (e.g., confidence intervals, confidence levels, confidence coefficients), and combinations thereof. Calculating and / or comparing distributions may include calculating and / or comparing probability distribution functions (e.g., core density assessment). Calculating and / or comparing distributions may include calculating and / or comparing uncertainty values ​​of two or more distributions. Uncertainty is typically a measure of variance or error and can be any suitable measure of variation or error. Non-limiting examples of uncertainty include standard deviation, standard error, calculated variance, p-value, arithmetic mean absolute deviation (MAD), and combinations thereof.

[0173] In some embodiments, comparisons (e.g., determining statistical differences) include comparing subsets of readings (e.g., the number of readings) with a threshold or range. The terms "threshold" and "threshold value" herein refer to any number calculated using a set of qualitative data (e.g., one or more references) and used as the limit for the determination (e.g., determining the presence of a breakpoint and / or chromosomal alteration). In some embodiments, a threshold is exceeded, and two or more subsets are determined to be statistically different. In some embodiments, a threshold is exceeded, and the test sample (e.g., a subject, such as a fetus) is determined to contain chromosomal alterations. In some embodiments, a threshold is exceeded, and a subset of readings is determined to include a breakpoint. In some embodiments, a quantitative value (e.g., reading count, reading distribution, Z-score, uncertainty value, confidence measure, etc., or combinations thereof) determined for a subset of readings is within or outside a threshold range of values, and the presence of a breakpoint and / or chromosomal alteration is determined. In some embodiments, the threshold or range of values ​​is typically calculated by mathematically and / or statistically manipulating the reading data (e.g., the number of readings from one or more subsets of references and / or test subjects). In some embodiments, the threshold includes an uncertainty value.

[0174] Any suitable threshold or range can be used to determine two significantly different subsets of readings. In some cases, two subsets of readings differing by about 0.01% or more (e.g., 0.01% of one or each subset value) are considered significantly different. Sometimes, two subsets differing by about 0.1% or more are considered significantly different. In some cases, two subsets differing by about 0.5% or more are considered significantly different. Sometimes, two subsets differing by about 0.5, 0.75, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or more than 10% are considered significantly different. Sometimes, two subsets are significantly different and there is no overlap between the subsets and / or no overlap within the range defined by the uncertainty calculated for one or both subsets. In some cases, the uncertainty value (e.g., standard deviation) is expressed as σ. Sometimes, two subsets are significantly different, differing by about 1 or more times the stated uncertainty value (e.g., 1σ). Sometimes two subsets are significantly different, with differences of about 2 or more times the uncertainty (e.g., deviation, standard deviation, MAD, etc.). The confidence level can be approximately 3 or more, approximately 4 or more, approximately 5 or more, approximately 6 or more, approximately 7 or more, approximately 8 or more, approximately 9 or more, or approximately 10 or more times the uncertainty. Sometimes, when the difference between two subsets is approximately 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, or 4.0 times the uncertainty or more, they are significantly different. In some implementations, the confidence level increases with increasing difference between the two subsets. In some cases, the confidence level decreases with decreasing difference between the two subsets and / or increasing uncertainty.

[0175] In some implementations, a statistically significant difference is determined to exist when the deviation (e.g., standard deviation, mean absolute deviation) between the number of readings of the test sample and the reference is less than about 3.5, less than about 3.4, less than about 3.3, less than about 3.2, less than about 3.1, less than about 3.0, less than about 2.9, less than about 2.8, less than about 2.7, less than about 2.6, less than about 2.5, less than about 2.0, less than about 1.5, or less than about 1.0. For example, sometimes a significant difference is determined when the difference between the number of readings of the test sample and the number of readings of the reference is less than 3 measures (e.g., 3σ, 3MAD). In some implementations, a significant difference is determined when the difference between the number of readings of the test sample and the number of readings of the reference is less than 3 measures (e.g., 3σ, 3MAD). In some implementations, a deviation of less than 3 (e.g., 3σ of the standard deviation) between the number of readings of the test sample and the reference generally indicates the absence of chromosomal alteration. The deviation between the number of readings of the test sample and the number of readings of one or more reference objects can be plotted and visualized (e.g., Z-score plot).

[0176] In some embodiments, the comparison includes comparing Z-scores. In some embodiments, the comparison includes comparing the Z-scores of a subset of readings from the test sample with a predetermined threshold, a threshold range, and / or with one or more Z-scores derived from a reference (e.g., a Z-score range). In some embodiments, the Z-scores and / or thresholds determined based on the Z-scores are used to determine that the subset of readings is significantly different from another subset and / or significantly different from a reference. In some embodiments, subsets of readings containing Z-scores below a threshold range and / or within a threshold range (e.g., within an uncertainty level, such as below 3, 2, or 1σ, within a predetermined range) are not significantly different. In some embodiments, subsets of readings containing Z-scores above a threshold range and / or outside a threshold range (e.g., above a predetermined uncertainty level, such as above 2, 2.5, 3, 3.5, 4, 5, or 6σ, outside a predetermined range) are significantly different. In some implementations, the threshold or predetermined value used for comparing Z scores is at least 2.5, at least 2.75, at least 3.0, at least 3.25, at least 3.5, at least 3.75, at least 4.0, at least 4.25, at least 4.5, at least 4.75, at least 5.0, at least 5.25, at least 5.5, at least 5.75, at least 6.0, at least 6.25, at least 6.5, at least 6.75, at least 7.0, at least 7.25, at least 7.5, at least 7.75, at least 8, at least 8.5, at least 9, at least 9.5, or at least 10.

[0177] Comparisons typically involve multivariate analysis. In some implementations, multivariate analysis includes generating and / or comparing heatmaps. In some implementations, heatmaps can be visually compared and visually determined to identify breakpoints and / or chromosomal alterations. Multivariate analysis sometimes involves mathematical operations on two or more data sets (e.g., two or more subsets of readings). For example, sometimes two or more data sets (e.g., the number of readings, Z-scores, uncertainties, and / or coefficients obtained for two or more subsets of data) are added, subtracted, multiplied, divided, and / or standardized.

[0178] The comparisons described herein can be performed via a comparison module (e.g., 130) or via a machine containing a comparison module. In some implementations, the comparison module includes code (e.g., a script) written in S or R using a suitable package (e.g., an S package, an R package). For example, a heatmap can be generated using heatmap 2, an R package described with gplots (gplots[online],[launched 2013-09-25], retrieved from the Internet via the World Wide Web Uniform Resource Locator: cran.r-project.org / web / packages / gplots / gplots.pdf), which can be downloaded from gplots (gplots[online],[launched 2013-09-25], retrieved from the Internet via the World Wide Web Uniform Resource Locator: cran.r-project.org / web / packages / gplots). For example, a heatmap can be generated using heatmap 2 and the script described below.

[0179] heatmap.2(x)

[0180] Where x is a matrix of Z-scores (calculated directly in R) for comparisons of chromosomes A and B between the sample and the control group. In some implementations, the comparison module includes and / or uses a suitable statistical software package.

[0181] Identifying chromosomal alterations

[0182] In some implementations, the presence of a chromosomal alteration is determined. Determining the presence of a chromosomal alteration is sometimes referred to herein as determining or generating a “result” or “making a determination.” In some implementations, the presence of a chromosomal alteration is determined by comparison. The presence of a chromosomal alteration is sometimes determined by comparing one or more selected subsets of inconsistent readings obtained from a sample with those obtained from a reference. In some implementations, the number of inconsistent reading companions containing substantially similar candidate breakpoints identified with respect to the test sample is compared with the number of inconsistent reading companions containing substantially similar candidate breakpoints identified with respect to a reference.

[0183] The absence of chromosomal alterations in the test subject (e.g., a fetus) is determined by comparison. In some embodiments, the absence of chromosomal alterations is determined when a selected subset of inconsistent reading partners from the test sample contains candidate breakpoints that are the same as or substantially similar to a selected subset of inconsistent reading partners from a reference. Sometimes, the absence of chromosomal alterations in the test subject is determined when one or more or all subsets of readings from the test sample differ from (e.g., statistically different) from a reference subset of readings. In some embodiments, determining the absence of chromosomal alterations includes determining, by comparison, that one or more breakpoints are absent from the test sample (e.g., the test subject).

[0184] A comparison determines that a test subject (e.g., a fetus) contains one or more chromosomal alterations. In some embodiments, a chromosomal alteration is determined to exist in the test subject when one or more subsets of readings of the test sample (e.g., a selected subset of readings) differ (e.g., statistically different) from one or more subsets of readings of a reference. Sometimes, a chromosomal alteration is determined to exist in the test subject by identifying a test subject with significantly more readings containing candidate breakpoints or substantially similar candidate breakpoints than a test reference containing candidate breakpoints or substantially similar candidate breakpoints, wherein the candidate breakpoints and / or substantially similar candidate breakpoints in the test sample and the reference are substantially similar. In some embodiments, a chromosomal alteration is determined to exist in the test subject when the candidate breakpoints and / or breakpoints included in a selected subset of inconsistent reading partners of the test sample are significantly different (e.g., statistically different) from the candidate breakpoints in a selected subset of inconsistent reading partners of the reference. In some embodiments, determining the presence of a chromosomal alteration includes identifying one or more breakpoints in the test sample (e.g., the test subject) by comparing candidate breakpoints in the test sample with candidate breakpoints in a reference. In some embodiments, determining the presence of a chromosomal alteration includes identifying a subset of readings from a test sample containing substantially similar breakpoints, wherein a reference (e.g., a subset of readings from a reference sample) does not contain candidate breakpoints substantially similar to the breakpoints identified in the test sample. In some embodiments, determining the presence of a chromosomal alteration includes identifying a first breakpoint and a second breakpoint of a chromosomal alteration (e.g., translocation or insertion) in the test subject. In some embodiments, determining the presence of a chromosomal alteration includes identifying a single breakpoint in the test subject. In some embodiments, determining the presence of a chromosomal alteration includes providing one or more breakpoints associated with the chromosomal alteration identified in the test subject.

[0185] Sometimes breakpoints are selected that include true breakpoints, and sometimes they are not. To avoid being limited by theory, candidate breakpoints are sometimes identified based on mapping artifacts caused by misalignment between two regions of a reading and two different chromosomes or non-proximate locations on chromosomes. These candidate breakpoints do not include true breakpoints. Mapping artifacts and / or misalignment often occur in the test sample or a reference sample (e.g., samples known not to contain chromosomal alterations), causing candidate breakpoints to substantially exclude true breakpoints. In some implementations, the presence of a breakpoint is determined by comparison. For example, candidate breakpoints that do not contain breakpoints and candidate breakpoints containing true breakpoints can typically be identified and / or distinguished from each other by comparing candidate breakpoints in the test sample with candidate breakpoints in a reference. For example, sometimes breakpoints are identified by comparing a subset of readings in the test sample with a subset of readings in a reference, where the two subsets contain substantially similar candidate breakpoints. Typically, the subset of readings in the test sample containing substantially similar candidate breakpoints is determined by comparison to contain true breakpoints, where the reference (e.g., a subset of readings from a reference sample) does not include a subset of readings containing candidate breakpoints substantially similar to those identified in the test sample.

[0186] In some embodiments, the location and / or position of a fracture point is determined based on the location and / or position of a candidate fracture point (e.g., a candidate fracture point containing the fracture point). In some embodiments, the location and / or position of a fracture point is determined by the method described herein for determining the location and / or position of a candidate fracture point. In some embodiments, the location and / or position of fracture points 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more in the test object are determined.

[0187] In some implementations, the location and / or position of a chromosomal alteration is determined based on the location and / or position of one or more breakpoints. In some implementations, a first breakpoint is identified in the test subject, wherein the location and / or position of the first breakpoint indicates the location and / or position of a chromosomal alteration (e.g., translocation, insertion). For example, when the first translocation is located at the end of a chromosome, a single breakpoint of the first translocation can be identified, indicating the location of the first translocation event. In some implementations, a first breakpoint and a second breakpoint are identified in the test sample, wherein the location and / or position of the first and second breakpoints indicate the location and / or position of a chromosomal alteration (e.g., translocation, insertion). For example, when the chromosome contains an insertion, two breakpoints may sometimes be identified (e.g., 5' and 3' breakpoints represent the 5' and 3' sides of the inserted segment), where all sequence reads map to the same strand (e.g., the positive strand) of a reference genome. When the translocation is intrachromosomal (e.g., the translocation involves inserting a segment into the chromosome), two breakpoints are typically identified. For balanced translocations between the first and second chromosomes, when an entire segment is exchanged between the two chromosomes, sometimes one or two breakpoints can be identified on the first chromosome (e.g., 5' and / or 3' breakpoints on the first chromosome) and sometimes one or two breakpoints can be identified on the second chromosome (e.g., 5' and / or 3' breakpoints on the second chromosome), where all sequence reads are mapped to the same strand (e.g., the positive strand) of a reference genome. In some embodiments, the location and / or position of the breakpoint is determined in the test sample, wherein the location and / or position of the breakpoint indicates the location and / or position of the chromosomal deletion.

[0188] The methods described herein can provide a result for determining whether a sample contains chromosomal alterations (e.g., fetal translocations), thereby providing a definitive result regarding the presence of chromosomal alterations (e.g., fetal translocations). The presence of chromosomal alterations can be determined by transforming, analyzing, and / or manipulating sequence reads mapped to a reference genome. In some embodiments, determining the result includes analyzing the nucleic acids of the pregnant female.

[0189] The number of readings available for the test sample can be used to factor any other suitable reference to determine whether chromosomal alterations are present in the test region of the test sample. For example, the number of readings available for the test sample can be used to factor the fetal fraction to determine whether chromosomal aberrations are present. Suitable procedures can be used to quantify the fetal fraction; non-limiting examples include mass spectrometry, sequencing procedures, or combinations thereof.

[0190] In some implementations, the presence of a chromosomal alteration (e.g., translocation) is determined based on a decision region. In some implementations, a determination (e.g., a determination of the presence of a chromosomal alteration, such as a result) is made when a value (e.g., a measurement and / or uncertainty level) or a set of values ​​falls within a predetermined range (e.g., a region, a decision region). In some implementations, the decision region is defined based on a set of values ​​obtained from samples from the same patient. In some implementations, the decision region is defined based on a set of values ​​obtained from the same chromosome or its segments. In some implementations, a decision region based on ploidy determination is defined based on a confidence level (e.g., a high confidence level, such as a low uncertainty level) and / or fetal fraction. In some implementations, the decision region is defined based on a ploidy determination of about 2.0% or more, about 2.5% or more, about 3% or more, about 3.25% or more, about 3.5% or more, about 3.75% or more, or about 4.0% or more, and a fetal fraction. For example, in some implementations, for samples obtained from pregnant females carrying a fetus, a determination that the fetus includes trisomy 21 is made based on a ploidy determination greater than 1.25 and a fetal fraction determination of 2% or more or 4% or more. For example, in some embodiments, for samples obtained from pregnant females carrying a fetus, a determination of euploidy of the fetus is made based on a ploidy determination of less than 1.25 and a fetal fraction of 2% or more, or 4% or greater. In some embodiments, a decision region is defined by confidence levels of approximately 99% or greater, approximately 99.1% or greater, approximately 99.2% or greater, approximately 99.3% or greater, approximately 99.4% or greater, approximately 99.5% or greater, approximately 99.6% or greater, approximately 99.7% or greater, approximately 99.8% or greater, or approximately 99.9% or greater. In some embodiments, a decision region is not used for determination. In some embodiments, a decision is made using a decision region and other data or information. In some embodiments, a decision is made based on the ploidy value without using a decision region. In some embodiments, a decision is made without calculating the ploidy value. In some embodiments, a decision is made based on profiling visual observations (e.g., visual observations at the genomic segment level). The determination can be made, in whole or in part, based on the determinations, values ​​and / or data obtained by the methods described herein, by any suitable method. Non-limiting examples of such methods include mappability variation, mappability threshold, relation, comparison, uncertainty and / or confidence determination, z-score, and combinations thereof.

[0191] In some implementations, the non-decision region is a region where no decision is made. In some implementations, the non-decision region is defined by a value or set of values ​​indicating low precision, high risk, high error, low confidence level, high uncertainty, or combinations thereof. In some implementations, the non-decision region is partially defined by a fetal score of about 5% or less, about 4% or less, about 3% or less, about 2.5% or less, about 2.0% or less, about 1.5% or less, or about 1.0% or less.

[0192] In some embodiments, the methods for determining the presence of chromosomal alterations (e.g., translocations) are performed with an accuracy of at least about 90% to about 100%. For example, the presence of chromosomal alterations can be determined with an accuracy of at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%. In some embodiments, the accuracy of determining the presence or absence of chromosomal alterations is approximately equal to or higher than the accuracy of other methods for determining chromosomal alterations (e.g., karyotype analysis). In some embodiments, the accuracy of determining the presence or absence of chromosomal alterations has a confidence interval (CI) of about 80% to about 100%. For example, the confidence interval (CI) can be approximately 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%.

[0193] In some implementations, one or more of the sensitivity, specificity, and / or confidence level are expressed as percentages. In some implementations, the percentages independently corresponding to each variable exceed about 90% (e.g., about 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99%) or exceed 99% (e.g., about 99.5% or higher, about 99.9% or higher, about 99.95% or higher, about 99.99% or higher)). In some implementations, the coefficient of variation (CV) is expressed as a percentage, sometimes said percentage is about 10% or lower (e.g., about 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1%) or is less than 1% (e.g., about 0.5% or lower, about 0.1% or lower, about 0.05% or lower, about 0.01% or lower)). In some implementations, the probability (e.g., a particular outcome is not due to chance) is expressed as a Z-score, p-value, or the result of a t-test. In some implementations, one or more data processing operations described herein can be used to generate variance, confidence intervals, sensitivity, specificity, etc. (collectively referred to as confidence parameters) for the measurement of the result. Specific examples of generating results and associated confidence levels are described in the Embodiments section and in International Application No. PCT / US12 / 59123 (WO2013 / 052913), the entire text of which is incorporated herein by reference, including all text, tables, equations, and figures.

[0194] As used herein, the term "sensitivity" refers to the number of true positives divided by the sum of the number of true positives and false negatives, where sensitivity (sens) can be in the range of 0 ≤ sens ≤ 1. The term "specificity" as used herein refers to the number of true negatives divided by the sum of the number of true negatives and false negatives, where specificity (spec) can be in the range of 0 ≤ spec ≤ 1. In some embodiments, methods are sometimes chosen where sensitivity and specificity are equal to 1, or 100%, or close to 1 (e.g., about 90% to about 99%). In some embodiments, methods are chosen where sensitivity is equal to 1 or 100%, while in other embodiments, methods are chosen where sensitivity is close to 1 (e.g., sensitivity about 90%, sensitivity about 91%, sensitivity about 92%, sensitivity about 93%, sensitivity about 94%, sensitivity about 95%, sensitivity about 96%, sensitivity about 97%, sensitivity about 98%, or sensitivity about 99%). In some implementations, a method with a specificity of 1 or 100% is selected, while in other implementations, a method with a specificity close to 1 (e.g., about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99%) is selected.

[0195] Ideally, the number of false negatives is zero or close to zero, so that if an object actually has at least one chromosomal alteration, no object is incorrectly identified as not having at least one chromosomal alteration. Conversely, the ability of the prediction algorithm to correctly classify negatives is typically evaluated, which is a complementary measure of sensitivity. Ideally, the number of false positives is zero or close to zero, so that if an object does not have the evaluated chromosomal alteration, no object is incorrectly identified as having at least one chromosomal alteration.

[0196] In some implementations, results are generated after performing one or more of the processing steps described herein. In some implementations, results are generated as a result of one of the processing steps described herein, while in others, results are generated after various statistical and / or mathematical operations are performed on the data set. Results regarding the determination of the presence of chromosomal alterations can be expressed in any form, including but not limited to probabilities (e.g., concession ratio, p-value), likelihood, in-cluster or out-of-cluster values, over- or under-threshold values, values ​​within a range (e.g., threshold range), values ​​with variance or confidence measurements, or risk factors. In some implementations, comparisons between samples allow for the determination of sample characteristics (e.g., allowing for the identification of duplicate samples and / or mixed samples (e.g., mislabeled, combined, etc.)).

[0197] In some implementations, the results include values ​​that are above or below a predetermined threshold or cutoff value (e.g., greater than 1, less than 1), and the uncertainty or confidence level associated with said value. In some implementations, the predetermined threshold or cutoff value is an expected level or a range of expected levels. The results may also describe the assumptions used for data processing. In some implementations, the results include values ​​that fall within or outside a predetermined range of values ​​(e.g., a threshold range), and the associated uncertainty or confidence level of said value within or outside said range. In some implementations, the results include values ​​equal to a predetermined value (e.g., equal to 1, equal to 0), values ​​equal to a predetermined range of values, and the associated uncertainty or confidence level of whether they are equal to or outside that range. The results are sometimes illustrated graphically (e.g., a distribution plot).

[0198] As described above, results can be characterized as true positive, true negative, false positive, or false negative. The term "true positive" as used herein refers to a subject correctly diagnosed as having a chromosomal alteration. The term "false positive" as used herein refers to a subject incorrectly identified as having a chromosomal alteration. The term "true negative" as used herein refers to a subject correctly identified as not having a chromosomal alteration. The term "false negative" as used herein refers to a subject incorrectly identified as not having a chromosomal alteration. Two performance metrics can be calculated for any given method based on the proportion of occurrence: (i) a sensitivity value, typically representing the predicted positive portion correctly identified as positive; and (ii) a specificity value, typically representing the predicted negative portion correctly identified as negative.

[0199] In some embodiments, nucleic acids in a sample are detected for the presence of chromosomal alterations. In some embodiments, a detected or undetected variant may remain in nucleic acids from one source of sample but not from nucleic acids from another source. Non-limiting examples of sources include placental nucleic acids, fetal nucleic acids, maternal nucleic acids, cancer cell nucleic acids, non-cancer cell nucleic acids, and combinations thereof. In non-limiting examples, a specific chromosomal alteration may be detected or undetected if (i) it remains in placental nucleic acids but not in fetal or maternal nucleic acids; (ii) it remains in fetal nucleic acids but not in maternal nucleic acids; or (iii) it remains in maternal nucleic acids but not in fetal nucleic acids. In some embodiments, the presence of a chromosomal alteration (e.g., translocation) in the fetus is determined. In this embodiment, the presence of a chromosomal alteration (e.g., translocation) in the mother is determined.

[0200] Some chromosomal alterations (e.g., translocations, insertions, deletions, inversions) that can be detected by the methods and / or systems described herein are associated with disorders or diseases, and non-limiting examples are shown in Table 1.

[0201] Table 1 - Translocation and Association Disorders

[0202]

[0203]

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225]

[0226]

[0227]

[0228]

[0229]

[0230]

[0231]

[0232]

[0233]

[0234]

[0235]

[0236]

[0237]

[0238]

[0239]

[0240]

[0241] Chromosomal alterations are sometimes associated with medical conditions (e.g., Table 1). A definitive outcome of a chromosomal alteration is sometimes the presence of a condition (e.g., a medical condition), disease, symptom, or abnormality, or includes a definitive outcome of detecting a condition, disease, symptom, or abnormality (non-limiting examples are listed in Table 1). In some embodiments, diagnosis includes assessing the outcome. The determination of the presence of a condition (e.g., a medical condition), disease, symptom, or abnormality using the methods described herein can sometimes be verified separately by other tests (e.g., karyotyping and / or amniocentesis). Data analysis and processing can provide one or more outcomes. The term "outcome" herein can refer to a data processing outcome that is conducive to determining the presence of a chromosomal alteration (e.g., translocation, deletion). In some embodiments, the term "outcome" herein refers to a conclusion that predicts and / or determines the presence of a chromosomal alteration (e.g., translocation, deletion). In some embodiments, the term "outcome" herein refers to a conclusion that predicts and / or determines the risk or likelihood of a subject (e.g., fetus) having a chromosomal alteration (e.g., translocation, deletion). Diagnosis sometimes includes using the outcome. For example, a healthy physician may analyze the outcome and provide a diagnosis based on or in part on the outcome. In some implementations, identifying, detecting, or diagnosing a condition, symptom, or abnormality (e.g., listed in Table 1) includes using a definitive result indicating the presence of a chromosomal alteration. In some implementations, the presence of a chromosomal alteration is determined based on results from inconsistent read pairs, mapping features, and breakpoint identification. In some implementations, the presence of one or more conditions, symptoms, or abnormalities listed in Table 1 is determined using results generated by one or more methods or systems described herein. In some implementations, a diagnosis includes determining the presence of a condition, symptom, or abnormality. A diagnosis typically includes determining a chromosomal alteration as the nature and / or cause of the condition, symptom, or abnormality. In some implementations, the result is not a diagnosis. Results often include one or more numerical values ​​generated using the processing methods described herein, with one or more considerations regarding probability. Considerations of risk or probability may include, but are not limited to, uncertainty, measurement variability, confidence level, sensitivity, specificity, standard deviation, coefficient of variance (CV) and / or confidence level, Z-score, Chi value, Phi value, ploidy value, fitted fetal fraction, area ratio, median level, etc., or combinations thereof. Considerations of probability help determine whether a subject is at risk of chromosomal alteration or has a genetic variation, and definitive results indicating the presence of a genetic disease often include such considerations.

[0242] Results are sometimes phenotypes. Sometimes, results are phenotypes with relevant confidence levels (e.g., uncertainty values, such as a fetus being positive for autism with a 99% confidence level; a pregnant female carrying a male fetus with a 95% confidence level; a test subject being negative for chromosomal alteration-related cancer with a 95% confidence level). Different methods of generating result values ​​can sometimes produce different types of results. Typically, there are four possible scores or determinations based on result values ​​generated using the methods described herein: true positive, false positive, true negative, and false negative. The terms “score,” “rating,” and “determination” as used herein refer to a calculation of the probability of the presence of a specific chromosomal alteration in a subject / sample. Scores can be used to determine, for example, changes, differences, or proportions of sequence reads that can be located corresponding to a chromosomal alteration. For example, with respect to a reference genome, a positive score for selected chromosomal alterations or portions of a dataset can guide the identification of the presence of a chromosomal alteration, which is sometimes associated with medical conditions (such as cancer, autism, etc.). In some embodiments, results include levels, profiles, and / or graphs (such as profile graphs). In those embodiments where results include profiles, appropriate profiles or combinations of profiles may be used for the results. Non-limiting examples of summaries that can be used for the results include z-score summaries, p-value summaries, ÷-value summaries, Values, etc., and their combinations.

[0243] Healthcare professionals or other qualified personnel who receive reports containing one or more results determining the presence of chromosomal alterations can use the data presented in the report to make a determination about the status of the test subject or patient. In some embodiments, healthcare professionals can provide recommendations based on the provided results. In some embodiments, healthcare professionals or qualified personnel can provide test subjects or patients with a determination or score regarding the presence of chromosomal alterations, said determination or score based on one or more result values ​​or relevant confidence parameters provided in the report. In some embodiments, the determination or score is made manually by healthcare professionals or qualified personnel through reports provided by visible observation. In some embodiments, the scoring or determination is made by an automated program (sometimes programmed into software), and the information is provided to the test subject or patient after the accuracy has been reviewed by healthcare professionals or qualified personnel. The term “receive a report” as used herein refers to a written and / or graphic representation containing results obtained through any means of contact, which, after review, is provided to healthcare professionals or other qualified personnel to make a determination regarding the presence of chromosomal alterations in the test subject or patient. The report can be generated via computer or human data input and can be communicated electronically (e.g., from one network address to another address at the same or different physical location via the Internet, via computer, via fax) or by any other method of sending or receiving data (e.g., mail service, courier service, etc.). In some embodiments, the results are transmitted to healthcare professionals in a suitable medium, including but not limited to non-transitory computer-readable storage media and / or oral, archival, or document form. Documents may be, for example, but not limited to, audio files, non-transitory computer-readable files, paper documents, laboratory documents, or medical report documents.

[0244] The term "providing results" as used herein, and its grammatical equivalents, may also refer to any method of obtaining such information, including but not limited to obtaining information from a laboratory (e.g., laboratory documents). Laboratory documents can be generated by performing one or more tests or one or more data processing steps in a laboratory to determine the presence of the medical condition. The laboratory may be located in the same or different locations (e.g., in another country) as the person identified by the laboratory documents as having or not having the medical condition. For example, laboratory documents may be generated at one location and transmitted to another location, where the information will be transmitted to the pregnant female subject. In some embodiments, the laboratory documents may be in tangible or electronic form (e.g., computer-readable form).

[0245] In some implementations, the results may be provided to healthcare professionals, physicians, or qualified individuals in laboratories, and these individuals may make diagnoses based on the results. In some implementations, the results may be provided to healthcare professionals, physicians, or qualified individuals in laboratories, and these individuals may make diagnoses in part based on the results, as well as other data and / or information and other results.

[0246] Healthcare professionals and qualified individuals may provide appropriate advice based on the results provided in this report. Non-limiting examples of advice that can be provided based on the results report include surgery, radiation therapy, chemotherapy, genetic counseling, postnatal treatment options (such as life planning, long-term supportive care, medication, symptomatic treatment), termination of pregnancy, organ transplantation, blood transfusion, or a combination thereof.

[0247] Laboratory personnel (e.g., laboratory administrators) can analyze values ​​that may determine the presence of chromosomal alterations (e.g., number of test sample readings, number of reference readings, level of bias). For narrow or questionable determinations regarding the presence of chromosomal alterations, laboratory personnel can perform the same test again and / or arrange different tests (e.g., karyotyping and / or amniocentesis in cases of fetal chromosomal alterations), using the same or different nucleic acid samples from the test subject.

[0248] Results are typically provided to healthcare professionals (such as laboratory technicians or administrators; physicians or assistants). Results are usually provided by an results module. The results module may include suitable statistical software packages. In some implementations, results are provided via a graphing module. Suitable statistical software typically includes appropriate graphing modules. In some implementations, the results module generates and / or compares Z-scores.

[0249] In some embodiments, the plotting module processes and / or transforms data and / or information into suitable visual media, non-limiting examples of which include charts, graphs, diagrams, etc., or combinations thereof. In some embodiments, the plotting module processes, transforms, and / or transfers data and / or information for presentation on a suitable display (e.g., a monitor, LED, LCD, CRT, etc., or combinations thereof), a printer (e.g., a printed paper display), suitable peripheral devices, or other equipment. In some embodiments, the plotting module provides a visual display of relationships and / or overviews.

[0250] In some implementations, results are provided on the device or its peripheral devices or components. For example, results are sometimes provided on a printer or display. In some implementations, definitive results regarding the presence of chromosomal alterations and / or related diseases or disorders are provided to healthcare professionals in report form, while in other implementations, the report includes displaying the result value and associated confidence parameters. Generally, results are displayed in an appropriate format to help determine the presence of chromosomal alterations and / or medical conditions. Non-limiting examples of formats suitable for reporting and / or displaying data sets or reporting results include numerical data, graphs, 2D graphs, 3D graphs, and 4D graphs, pictures, pictographs, charts, bar graphs, pie charts, line graphs, flowcharts, scatter plots, atlases, bar charts, density plots, function graphs, circuit diagrams, block diagrams, bubble charts, constellation diagrams, outline diagrams, statistical graphs, spider diagrams, Venn diagrams, nodal plots, etc., and combinations thereof. Various examples of result representation are shown in the accompanying drawings and described herein.

[0251] Data filtering and processing

[0252] In some implementations, one or more processing steps may include one or more filtering steps. The term "filtering" herein refers to removing a portion or group of data from consideration and retaining a subset of the data. Sequence readings can be selected based on any suitable criteria, including but not limited to redundant data (such as redundant or overlapping mapping readings), uninformative data, sequences containing excessively high or low frequencies, noisy data, or combinations thereof. The filtering process typically involves removing one or more readings and / or pairs of readings (e.g., inconsistent pairs of readings) from consideration. Reducing the number of readings, pairs of readings, and / or readings containing candidate breakpoints in the dataset used to analyze for chromosomal alterations typically reduces the complexity and / or dimensionality of the dataset and sometimes increases the speed of searching for and / or identifying chromosomal alterations by two or more orders of magnitude.

[0253] In some embodiments, the systems or methods described herein include filtering reads, inconsistent read partners, and / or inconsistent read pairs. Filtering may be performed before or after the following steps: identifying inconsistent read partners, characterizing the mappability of multiple sequence read subsequences, providing mappability variations, selecting a subset of inconsistent read partners, identifying candidate breakpoints, comparing subsets of reads, comparing candidate breakpoints, identifying breakpoints, or comparing the number of inconsistent read partners for a sample and a parameter. Filtering is typically performed before determining the presence of one or more chromosomal alterations.

[0254] Filtering is typically performed by a system or module. The system or module used for filtering herein refers to a filtering module. In some implementations, a filtering module includes code (e.g., a script) written in S or R using a suitable package (e.g., an S package, an R package). For example, a filtering module may include and use one or more SAM tools (SAM tools [online], [launched 2013-09-25], retrieved from the Internet via the World Wide Web Uniform Resource Locator: samtools.sourceforge.net). For example, the sum of all applicable markers can be used to identify a consistent read, where the choice between a consistent or PCR repeat read is "if (bitwiseA == 83 || bitwiseA == 163 || bitwiseA == 99 || bitwiseA == 147 || bitwiseA >= 1024)", where bitwiseA is the sum of all applicable markers in the SAM format file.

[0255] Filters typically accept a set of data (e.g., an array of reads) as input and output as a subset of the data to be filtered (e.g., a subset of the filtered reads). In some implementations, reads removed during the filtering process are typically discarded and / or removed from further analysis (e.g., statistical analysis). Filtering typically involves removing data from the array of reads. In some implementations, filtering includes removing one or both read companions from inconsistent read pairs. In some implementations, filtering includes removing multiple reads from the array of reads. Sometimes the filtering step does not remove data.

[0256] Filtering typically removes, discards, or rejects certain data from a read array based on predetermined condition queries. For example, sometimes a filter accepts input readings from the system or other modules, performs filtering on the accepted readings, and only accepts those input readings that meet the conditions. In some implementations, the filter accepts input readings from the system or other modules, performs filtering, and only removes, discards, or rejects those input readings that meet the conditions. In some implementations, the condition query includes a yes / no or true / false decision. For example, sometimes "true" or "yes" is assigned to one or more readings when the query condition is met, and "false" or "no" is assigned to the readings when the query condition is not met.

[0257] In some implementations, filtering includes removing, rejecting, and / or discarding inconsistent readings (e.g., consistent readings). In some implementations, the conditional query of the filter (e.g., 20) includes determining whether inconsistent paired readings exist. In some implementations, inconsistent paired readings are removed, rejected, and / or discarded. In some implementations, inconsistent paired readings are assigned as "false" or "not" and removed, rejected, and / or discarded by the module. In some implementations, inconsistent readings are not allowed to pass through filter 20. In some implementations, inconsistent readings are deleted, moved to a junk file or temporary file (e.g., 10), or their original data positioning and / or format is preserved. In some implementations, inconsistent paired readings are identified and / or retained in a subset of the output of filtered readings. In some implementations, inconsistent paired readings are identified and fed to another module or filter. In some implementations, inconsistent paired readings are identified, assigned as "true" or "yes," and retained in a subset of the output of filtered readings (e.g., inconsistent reading pairs). In some implementations, inconsistent readings are accepted and passed through filter 20. Readings can be filtered based on the presence of inconsistent and / or non-inconsistent readings using an inconsistency reading identification module (e.g., filter 20). The inconsistency reading identification module may sometimes include filters configured to remove, reject, and / or discard non-inconsistent readings.

[0258] In some implementations, filtering includes removing, rejecting, and / or discarding inaccurate repeats. Repeated reads herein refer to PCR repeats. In some implementations, the conditional query of the filter (e.g., 30) includes determining the presence of PCR repeats. In some implementations, PCR repeats are assigned as "true" or "yes" and are removed, rejected, and / or discarded by the module. In some implementations, PCR repeats are not allowed to pass through filter 30. In some implementations, PCR repeats are deleted, moved to a junk file or temporary file (e.g., 10), or their original data location is preserved and / or appropriate. A representative read of a repeated read group is referred to herein as a "representative read." In some implementations, representative reads and unique reads are maintained in a subset of the output of the filtered reads. Representative reads and unique reads are typically identified and fed to another module or filter. In some implementations, representative reads and unique reads are identified, assigned as "false" or "not," and maintained in a subset of the output of the filtered reads. In some implementations, representative reads and unique reads are accepted into a filter (e.g., 30) and / or pass through a filter (e.g., 30). Readings can be filtered based on PCR repeats using a PCR repeat filter (e.g., filter 30). Filter modules sometimes include a PCR repeat filter.

[0259] In some implementations, filtering includes removing, rejecting, and / or discarding low sequencing quality reads. Low sequencing quality reads are typically reads with PHRED scores equal to or below about 40, about 35, about 30, about 25, about 20, about 15, about 10, or about 5. In some implementations, the conditional query of the filter (e.g., 40) includes determining the presence of low sequencing quality reads. In some implementations, low sequencing quality reads are assigned "true" or "yes" and removed, rejected, and / or discarded by the module. In some implementations, low sequencing quality reads are not allowed to pass through filter 40. In some implementations, low sequencing quality reads are deleted, moved to a junk file or temporary file (e.g., 10), or their original data location is preserved and / or appropriate. Non-low sequencing quality reads (e.g., high sequencing quality reads) are sometimes retained in a subset of the filtered reads output. Non-low sequencing quality reads are typically identified and / or sent to another module or filter. In some implementations, non-low sequencing quality reads are identified, assigned "false" or "not," and retained in a subset of the filtered reads output. In some implementations, non-low sequencing quality reads are accepted into filters (e.g., 40) and / or through filters (e.g., 40). Reads can be filtered according to sequencing quality using sequencing quality filters (e.g., filter 40). Filter modules sometimes include sequencing quality filters.

[0260] In some implementations, filtering includes removing, rejecting, and / or discarding reads when the sequence read subsequence of a read includes mapping discontinuities. Mapping discontinuities typically refer to three or more (e.g., >2) fragments in the sequence read subsequence of a read mapping to different (e.g., non-obvious) locations in a reference genome (e.g., reads containing step-by-step multiple alignments). Mapping discontinuities sometimes refer to silica fragments of reads mapped to (i) different chromosomes (e.g., three or more different chromosomes), (ii) different locations, wherein each location is separated by a predetermined fragment size (e.g., greater than 300 bp, greater than 500 bp, greater than 1000 bp, greater than 5000 bp, or greater than 10,000 bp), (iii) different and / or opposite directions, or combinations thereof. For example, a mapping discontinuity can refer to a sequence read subsequence where two fragments map to opposite directions and a third fragment maps to a different chromosome. In some implementations, the conditional query of the filter (e.g., 60) includes determining whether a read containing mapping discontinuities exists. In some implementations, readings containing mapping discontinuities are assigned as "true" or "yes" and are removed, rejected, and / or discarded by the module. In some implementations, readings containing mapping discontinuities are not allowed to pass through filter 60. In some implementations, readings containing mapping discontinuities are deleted, moved to a junk file or temporary file (e.g., 10), or their original data positions are maintained and / or appropriate. Readings without mapping discontinuities are sometimes retained in a subset of the filtered readings output. Readings without mapping discontinuities are typically identified and / or sent to another module or filter. In some implementations, readings without mapping discontinuities are identified, assigned as "false" or "not," and retained in a subset of the filtered readings output. In some implementations, readings without mapping discontinuities are accepted into a filter (e.g., 60) and / or pass through a filter (e.g., 60). Readings can be filtered based on mapping discontinuities using a mapping discontinuity filter (e.g., filter 60). Filter modules sometimes include mapping discontinuity filters.

[0261] In some embodiments, filtering includes removing, rejecting, and / or discarding readings containing unmappable sequence read subsequences. In some embodiments, filtering includes removing, rejecting, and / or discarding readings containing one or more, more than two, more than three, more than four, more than five, more than six, more than seven, more than eight, more than nine, more than ten, more than eleven, more than twelve, more than thirteen, more than fourteen, or more than fifteen unmappable sequence read subsequences. Unmappable means that the polynucleotide cannot be clearly mapped to a location in the reference genome (e.g., the human reference genome). In some embodiments, the conditional query of the filter (e.g., 70) includes determining whether a reading containing an unmappable sequence read subsequence exists. In some embodiments, readings containing unmappable sequence read subsequences are assigned a "true" or "yes" value and are removed, rejected, and / or discarded by the module. In some embodiments, readings containing unmappable sequence read subsequences are not allowed to pass through filter 70. In some implementations, readings containing unmappable sequence subsequences are deleted, moved to a junk file or temporary file (e.g., 10), or have their original data positioning and / or format preserved. Readings not containing unmappable sequence subsequences are sometimes retained in a subset of the filtered readings output. Readings not containing unmappable sequence subsequences are typically identified and / or sent to another module or filter. In some implementations, readings not containing unmappable sequence subsequences are identified, assigned as "false" or "not," and retained in a subset of the filtered readings output. In some implementations, readings not containing unmappable sequence subsequences are accepted into a filter (e.g., 70) and / or pass a filter (e.g., 70). Readings can be filtered based on unmappable sequence subsequences using a mapping filter (e.g., filter 70). Filter modules sometimes include mapping filters.

[0262] In some embodiments, filtering includes removing, rejecting, and / or discarding reads containing sequence read subsequences mapped to mitochondrial DNA. In some embodiments, filtering includes removing, rejecting, and / or discarding reads containing one or more, more than two, more than three, more than four, more than five, more than six, more than seven, more than eight, more than nine, more than ten, more than eleven, more than twelve, more than thirteen, more than fourteen, or more than fifteen sequence read subsequences mapped to mitochondrial DNA. In some embodiments, the conditional query of the filter (e.g., 80) includes determining the presence of reads containing sequence read subsequences mapped to mitochondrial DNA. In some embodiments, reads containing sequence read subsequences mapped to mitochondrial DNA are assigned as "true" or "yes" and are removed, rejected, and / or discarded by the module. In some embodiments, reads containing sequence read subsequences mapped to mitochondrial DNA are not allowed to pass through filter 80. In some implementations, reads containing sequence subsequences mapped to mitochondrial DNA are deleted, moved to a junk file or temporary file (e.g., 10), or their original data location is maintained and / or appropriate. Reads without sequence subsequences mapped to mitochondrial DNA are sometimes retained in a subset of the filtered reads output. Reads without sequence subsequences mapped to mitochondrial DNA are typically identified and / or sent to another module or filter. In some implementations, reads without sequence subsequences mapped to mitochondrial DNA are identified, assigned as "false" or "non-false," and retained in a subset of the filtered reads output. In some implementations, reads without sequence subsequences mapped to mitochondrial DNA are accepted into a filter (e.g., 80) and / or pass through a filter (e.g., 80). Reads can be filtered based on sequence subsequences mapped to mitochondrial DNA using a mitochondrial filter (e.g., filter 80). Filter modules sometimes include mitochondrial filters.

[0263] In some embodiments, filtering includes removing, rejecting, and / or discarding reads containing sequence read subsequences mapped to centromere DNA. In some embodiments, filtering includes removing, rejecting, and / or discarding reads containing one or more, more than two, more than three, more than four, more than five, more than six, more than seven, more than eight, more than nine, more than ten, more than eleven, more than twelve, more than thirteen, more than fourteen, or more than fifteen sequence read subsequences mapped to centromere DNA. In some embodiments, the filter's conditional query includes determining the presence of reads containing sequence read subsequences mapped to centromere DNA. In some embodiments, reads containing sequence read subsequences mapped to centromere DNA are assigned as "true" or "yes" and are removed, rejected, and / or discarded by the module. In some embodiments, reads containing sequence read subsequences mapped to centromere DNA are not allowed to pass through the filter. In some embodiments, reads containing sequence read subsequences mapped to centromere DNA are deleted, moved to a junk file or temporary file (e.g., 10), or retain their original data positioning and / or format. Reads not containing sequence read subsequences mapped to centromere DNA are sometimes retained in a subset of the filtered reads output. Reads not containing sequence read subsequences mapped to centromere DNA are typically identified and fed into another module or filter. In some embodiments, reads not containing sequence read subsequences mapped to centromere DNA are identified, assigned as "false" or "non-false," and retained in a subset of the filtered reads output. In some embodiments, reads not containing sequence read subsequences mapped to centromere DNA are accepted into filters and / or pass through filters. Reads can be filtered by a centromere DNA filter based on the sequence read subsequence mapped to centromere DNA. Filter modules sometimes include centromere DNA filters.

[0264] In some embodiments, filtering includes removing, rejecting, and / or discarding readings containing a sequence reading subsequence mapped to a repeating element. In some embodiments, filtering includes removing, rejecting, and / or discarding readings containing one or more, more than two, more than three, more than four, more than five, more than six, more than seven, more than eight, more than nine, more than ten, more than eleven, more than twelve, more than thirteen, more than fourteen, or more than fifteen sequence reading subsequences mapped to repeating elements. In some embodiments, the conditional query of the filter (e.g., 110) includes determining whether a reading containing a sequence reading subsequence mapped to a repeating element exists. In some embodiments, readings containing a sequence reading subsequence mapped to a repeating element are assigned "true" or "yes" and are removed, rejected, and / or discarded by the module. In some embodiments, readings containing a sequence reading subsequence mapped to a repeating element are not allowed to pass through filter 110. In some implementations, readings of a sequence reading subsequence mapped to a repeating element are deleted, moved to a junk file or temporary file (e.g., 10), or have their original data positioning and / or format preserved. Readings of a sequence reading subsequence not mapped to a repeating element are sometimes retained in a subset of the filtered readings output. Readings of a sequence reading subsequence not mapped to a repeating element are typically identified and / or sent to another module or filter. In some implementations, readings of a sequence reading subsequence not mapped to a repeating element are identified, assigned as "false" or "not," and retained in a subset of the filtered readings output. In some implementations, readings of a sequence reading subsequence not mapped to a repeating element are accepted into a filter (e.g., 110) and / or pass through a filter (e.g., 110). Readings can be filtered based on the sequence reading subsequence mapped to a repeating element by a repeating element filter (e.g., filter 110). Filter modules sometimes include repeating element filters.

[0265] In some implementations, filtering includes removing, rejecting, and / or discarding readings containing a single occurrence mutation event. In some implementations, the conditional query of the filter (e.g., 100) includes determining whether a single occurrence mutation event exists. As used herein, a "single occurrence mutation event" refers to a reading or inconsistent pair of readings containing a first candidate breakpoint, wherein no other reading obtained in the sample has been identified and / or no substantially similar candidate breakpoint exists (e.g., a candidate breakpoint substantially similar to the first candidate breakpoint). In some implementations, single occurrence mutation events are assigned as "true" or "yes" and are removed, rejected, and / or discarded. In some implementations, single occurrence mutation events are not allowed to pass through filter 100. In some implementations, single occurrence mutation events are deleted, moved to a junk file or temporary file (e.g., 10), or their original data location and / or format is preserved. Readings of non-single occurrence mutation events are sometimes retained in a subset of the output of the filtered readings (e.g., a selected subset). Readings of non-single occurrence mutation events are typically identified and / or sent to another module or filter. Readings of non-single-occurrence mutation events are identified, assigned as "false" or "not," and maintained within a subset (e.g., a selected subset) of the filtered readings output. In some implementations, non-single-occurrence mutation events are accepted into a filter (e.g., 100) and / or pass through a filter (e.g., 100). Readings can be filtered based on the presence or absence of a single-occurrence mutation event using a single-occurrence mutation event filter (e.g., filter 100). Filter modules sometimes include single-occurrence mutation event filters.

[0266] In some implementations, filtering includes removing, rejecting, and / or discarding subsets of sample readings, wherein the subset of sample readings contains breakpoints or candidate breakpoints that are substantially similar to a reference subset of readings. In some implementations, subsets of sample readings containing breakpoints or candidate breakpoints that are substantially similar to a reference subset of readings are assigned as "true" or "yes" and are removed, rejected, and / or discarded. In some implementations, subsets of sample readings containing breakpoints or candidate breakpoints that are substantially similar to a reference subset of readings are not allowed to pass through the filter. In some implementations, subsets of sample readings containing breakpoints or candidate breakpoints that are substantially similar to a reference subset of readings are deleted, moved to a junk file or temporary file (e.g., 10), or have their original data positioning and / or format preserved. Subsets of sample readings containing breakpoints or candidate breakpoints that are substantially not found and / or do not exist in the reference subset of readings are sometimes retained in the output subset of the filtered readings (e.g., the selected subset). A subset of sample readings containing no and / or no fracture points or candidate fracture points found in the referenced subset of readings is typically identified and / or fed into another module or filter. This subset of readings is typically identified, assigned as "false" or "non-false," and retained in the output subset of filtered readings (e.g., the selected subset). In some embodiments, the subset of sample readings containing no and / or no fracture points found in the referenced subset of readings is accepted into a filter and / or passes through a filter. A fracture point filter can filter a subset of sample readings based on the presence or absence of candidate fracture points or fracture points found in the referenced subset of readings. Filter modules sometimes include fracture point filters.

[0267] Systems or methods for determining the presence of chromosomal alterations in a sample may include one or more filtering steps and / or filters. These systems or methods may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, twenty or more, thirty or more, forty or more, or fifty or more filtering steps and / or filters. Filtering steps may be performed before and / or after any method, part of it, or step thereof described herein. Systems described herein may include suitable filters before and / or after any suitable process or module described herein. For example, in Figure 6 In the exemplary system shown, one or more filters and / or filtration steps may be introduced before or after 15, at 150, 151, 152, 153, and / or after 140. The one or more filters and / or filtration steps may be in any suitable order or arrangement. For example, as... Figure 7As shown, the inconsistent read identification module sends inconsistent reads to PCR repeat filter 30; filter 30 sends filtered reads to sequence quantification filter 40; filter 40 sends filtered reads to mapping discontinuity filter 60; filter 60 sends filtered reads to mapping filter 70; filter 70 sends filtered reads to read selection module 120; module 120 sends selected reads to single occurrence mutation event filter 100; filter 100 sends filtered reads to repeat element filter 110; and filter 110 sends filtered reads to comparison module 130. In some embodiments, reads may be filtered once or multiple times in the same filter. Any suitable filter may be incorporated into the methods or systems described herein. Filters and / or filtering methods are sometimes optional and may or may not be used in the methods or systems described herein. For example, filters 30, 40, 60, 70, 80, 90, 100 and / or 110 may be included in or excluded from the methods or systems described herein.

[0268] Any suitable procedure can be used to process the data set described herein. Non-limiting examples of methods suitable for processing data sets include filtering, standardization, weighting, mathematical processing of the data, statistical processing of the data, application of mathematical algorithms, plotting the data to identify patterns or trends for further processing, and combinations thereof. In some implementations, processing the data set described herein can reduce the complexity and / or dimensionality of large and / or complex data sets. Non-limiting examples of complex data sets include sequence readings generated from one or more test subjects and multiple reference subjects of different ages and ethnic backgrounds. In some implementations, the data set can contain thousands to millions of sequence readings from each test subject and / or reference subject.

[0269] In some implementations, data processing can be performed in any number of steps. For example, in some implementations, data can be adjusted and / or processed using only a single processing method, while in other implementations, data can be processed using one or more, five or more, ten or more, or twenty or more processing steps (e.g., one or more processing steps, two or more processing steps, three or more processing steps, four or more processing steps, five or more processing steps, six or more processing steps, seven or more processing steps, eight or more processing steps, nine or more processing steps, ten or more processing steps, eleven or more processing steps, twelve or more processing steps, thirteen or more processing steps, fourteen or more processing steps, fifteen or more processing steps, sixteen or more processing steps, seventeen or more processing steps, eighteen or more processing steps, nineteen or more processing steps, or twenty or more processing steps). In some implementations, the processing step may be the same step repeated two or more times (e.g., filtering two or more times, standardizing two or more times), while in other implementations, the processing step may be two or more different processing steps performed simultaneously or sequentially (e.g., filtering, standardizing; standardizing, monitoring peak height and edges; filtering, standardizing, standardizing against a reference, statistical processing to determine p-values, etc.). In some implementations, any suitable number and / or combination of the same or different processing steps may be used to process sequence readout data to aid in providing results. In some implementations, processing the data set using the standards described herein can reduce the complexity and / or dimensionality of the data set.

[0270] In some implementations, the processing steps may include using one or more statistical algorithms. Any suitable statistical algorithm may be used alone or in combination to analyze and / or process the data set described herein. Any suitable number of statistical algorithms may be used. In some implementations, one or more, five or more, ten or more, or twenty or more statistical algorithms may be used to analyze the data set. Non-limiting examples of statistical algorithms suitable for use with the methods described herein include decision trees, counting null values, multiple comparisons, comprehensive tests, the Behrens-Fischer problem, bootstrapping, Fisher's method combined with a significance test for independence, null hypothesis, Type I error, Type II error, exact test, one-sample Z-test, two-sample Z-test, one-sample t-test, paired t-test, two-sample pooled t-test with equal variances, two-sample unpooled t-test with unequal variances, single proportion Z-test, pooled two-proportion Z-test, unpooled two-proportion Z-test, one-sample chi-square test, two-sample F-test with equal variances, confidence intervals, confidence intervals, significance, meta-analysis, simple linear regression, strong linear regression, or a combination thereof.

[0271] In some implementations, the dataset can be analyzed using multiple (e.g., two or more) statistical algorithms (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, Bagging, neural networks, support vector machine models, random forests, classification tree models, k-nearest neighbors, logistic regression and / or loss smoothing) and / or mathematical and / or statistical operations (e.g., those described herein). In some implementations, using multiple operations can produce an N-dimensional space that can be used to provide results. In some implementations, analyzing the dataset using multiple operations can reduce the complexity and / or dimensionality of the dataset.

[0272] In some implementations, the sequence reads data sets are filtered, standardized, clustered, and technically / weighted. The processed data sets can then be compared and / or analyzed mathematically and / or statistically (e.g., using statistical functions or algorithms). In some implementations, the processed data can be further analyzed and / or compared by calculating Z-scores for one or more selected chromosomes or portions thereof. In some implementations, the processed data sets can be further analyzed and / or compared by calculating p-values. One implementation of the equation for calculating Z-scores is shown in Equation A (Example 1).

[0273] Determine fetal nucleic acid content

[0274] In some embodiments, the amount of fetal nucleic acid in the nucleic acid is determined (e.g., concentration, relative amount, absolute amount, copy number, etc.). In some embodiments, the amount of fetal nucleic acid in the sample is referred to as the "fetal fraction." In some embodiments, the "fetal fraction" refers to the fraction of fetal nucleic acid in circulating cell-free nucleic acid obtained from a sample (e.g., blood sample, serum sample, plasma sample). In some embodiments, the method for determining chromosomal alterations may also include determining the fetal fraction. In some embodiments, the presence or absence of chromosomal alterations is determined based on the fetal fraction (e.g., fetal fraction determination of the sample). The determination of the fetal fraction can be performed by suitable methods, non-limiting examples of which include the methods described below.

[0275] In some embodiments, the methods described herein for determining fragment lengths can be used to determine the fetal fraction. Cell-free fetal nucleic acid fragments are typically shorter than maternally derived nucleic acid fragments (see, for example, Chan et al. (2004) Clin. Chem. 50:88-92; Lo et al. (2010) Sci. Transl. Med. 2:61ra91). Therefore, in some embodiments, the fetal fraction can be determined by counting fragments below a specific length threshold and comparing that count to the amount of total nucleic acid in the sample. Methods for counting nucleic acid fragments of specific lengths are further described below.

[0276] In some embodiments, the content of fetal nucleic acids is determined based on: male fetal-specific markers (e.g., Y chromosome STR markers (e.g., DYS 19, DYS 385, DYS 392 markers); RhD markers in RhD-negative women), allele ratios of polymorphic sequences, or one or more markers specific to fetal nucleic acids but not specific to maternal nucleic acids (e.g., differential epigenetic biomarkers between mother and fetus (e.g., methylation; detailed below), or fetal RNA markers in maternal plasma (see, for example, Lo, 2005, Journal of Histochemistry and Cytochemistry 53(3):293-296)).

[0277] Determining fetal nucleic acid content (e.g., fetal fraction) is sometimes performed using fetal quantification assays (FQAs), as described in U.S. Patent Application Publication 2010 / 0105049, which is incorporated herein by reference. Such assays allow for the detection and quantification of fetal nucleic acids in maternal samples based on the methylation status of nucleic acids in the sample. In some embodiments, the content of fetal nucleic acids in the maternal sample can be determined relative to the total amount of nucleic acids present, thereby providing a percentage of fetal nucleic acids in the sample. In some embodiments, the copy number of fetal nucleic acids in the maternal sample can be determined. In some embodiments, the amount of fetal nucleic acids can be determined in a sequence-specific (or partial-specific) manner, and sometimes with sufficient sensitivity for precise chromosomal dosing analysis (e.g., to detect the presence or absence of fetal chromosomal alterations).

[0278] Fetal Quantification Assay (FQA) can be performed in conjunction with any of the methods described herein. This assay can be performed by any method known in the art and / or as described in U.S. Patent Application Publication 2010 / 0105049, such as methods that can distinguish maternal and fetal DNA based on differential methylation status, and methods that quantify fetal DNA (i.e., determine its content). Methods for distinguishing nucleic acids based on methylation status include, but are not limited to, methylation-sensitive capture (e.g., using the MBD2-Fc fragment, wherein the methylation-binding domain of MBD2 is fused to the Fc fragment of an antibody (MBD-FC) (Gebhard et al. (2006) Cancer Res. 66(12): 6118-28)); methylation-specific antibodies; bisulfite conversion methods, such as MSP (methylation-sensitive PCR), COBRA, methylation-sensitive single nucleotide primer extension (Ms-SNuPE), or Sequenom MassCLEAVE. TM The invention relates to techniques and the application of methylation-sensitive restriction enzymes (e.g., digesting maternal DNA in a maternal sample with one or more methylation-sensitive restriction enzymes to enrich fetal DNA). Methylation-sensitive enzymes can also be used to distinguish nucleic acids based on methylation status, such as preferential or significant cleavage or digestion when their DNA recognition sequences are unmethylated. Thus, unmethylated DNA samples are cleaved into smaller fragments than methylated samples, while highly methylated DNA samples are not cleaved. Unless explicitly stated otherwise, any method for distinguishing nucleic acids based on methylation status can be used in the compositions and methods of the present invention. The amount of fetal DNA can be determined, for example, by introducing one or more competing agents at known concentrations during the amplification reaction. The amount of fetal DNA can also be determined, for example, by RT-PCR, primer extension, sequencing, and / or counting. In some examples, the BEAMing technique described in U.S. Patent Application Publication 2007 / 0065823 can be used to determine the amount of nucleic acid. In some embodiments, a restriction efficacy can be determined and the amount of fetal DNA can be further determined using this efficiency ratio.

[0279] In some embodiments, fetal quantification (FQA) can be determined using the concentration of fetal DNA in a maternal sample, for example by: a) determining the total amount of DNA present in the maternal sample; b) selectively digesting the maternal DNA in the maternal sample with one or more methylation-sensitive restriction enzymes to enrich the fetal DNA; c) determining the amount of fetal DNA from step b); and d) comparing the amount of fetal DNA obtained in step c) with the total amount of DNA obtained in step a) to determine the concentration of fetal DNA in the maternal sample. In some embodiments, the absolute copy number of fetal nucleic acid in the maternal sample can be determined, for example, using mass spectrometry and / or a system utilizing a competitive PCR method for determining the absolute copy number. See, for example, Ding and Cantor (2003) PNAS.USA 100:3059-3064, and U.S. Patent Application Publication 2004 / 0081993, both of which are incorporated herein by reference.

[0280] In some embodiments, fetal fractions can be determined based on the allele ratio of polypeptide sequences (e.g., single nucleotide polymorphisms (SNPs)), for example using the method described in U.S. Patent Application Publication 2011 / 0224087, which is incorporated herein by reference. In this method, nucleotide sequence reads are obtained from a maternal sample, and the fetal fraction is determined by comparing the total number of nucleotide sequence reads mapped to a first allele with the total number of nucleotide sequence reads mapped to a second allele located at a reference polymorphic site (e.g., an SNP) in a reference genome. In some embodiments, fetal alleles are identified, for example, by the relatively small contribution of fetal alleles to a mixture of fetal and maternal nucleic acids in the sample, relative to the larger contribution of maternal nucleic acids to the mixture. Therefore, the relative abundance of fetal nucleic acids in the maternal sample can be determined as a parameter of the total number of unique sequence reads (for each of the two alleles at the polymorphic site) mapped to a target nucleic acid sequence on a reference genome.

[0281] The amount of fetal nucleic acid in extracellular nucleic acids can be quantified and can be used in conjunction with the methods described herein. Therefore, in some embodiments, the methods of the techniques described herein include an additional step of determining the amount of fetal nucleic acid. The amount of fetal nucleic acid in the nucleic acid sample of the subject can be determined before or after processing to prepare the sample nucleic acid. In some embodiments, the amount of fetal nucleic acid in the sample is determined after the sample nucleic acid has been processed and prepared, and used for further evaluation. In some embodiments, the results include decomposing the fetal nucleic acid fraction in the sample nucleic acid into factors (e.g., adjusting data, removing the sample, making a decision, or not making a decision).

[0282] The determination step (e.g., determining the presence of chromosomal alterations) can be performed before, during, or at any point in time within the methods described herein, or after certain methods described herein. For example, to achieve a determination with a given sensitivity or specificity (e.g., determining chromosomal alterations in the fetus), fetal nucleic acid quantification methods can be performed before, during, or after the determination of chromosomal alterations to identify samples with more than about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, or more fetal nucleic acids. In some embodiments, samples determined to have a certain fetal nucleic acid threshold amount (e.g., about 15% or more fetal nucleic acid; or about 4% or more fetal nucleic acid) are further used for analysis, for example, fetal sex or the presence of chromosomal alterations. In some implementations, only samples with a certain fetal nucleic acid threshold amount (e.g., about 15% or more fetal nucleic acid; or about 4% or more fetal nucleic acid) are selected (e.g., selected and informed to the patient) to determine whether chromosomal alterations are present.

[0283] In some embodiments, determining the fetal fraction or the amount of fetal nucleic acid is not necessary for identifying the presence of chromosomal alterations. In some embodiments, identifying the presence of chromosomal alterations does not require sequence differentiation between fetal and maternal DNA. In some embodiments, this is because the additive contribution of maternal and fetal sequences to specific chromosomes, chromosomal portions, or segments is analyzed. In some embodiments, identifying the presence of chromosomal alterations does not rely on prior sequence information that distinguishes fetal and maternal DNA.

[0284] Fetal sex

[0285] In some cases, determining the sex of the fetus in the womb is beneficial. For example, a parent with a family history of one or more sex-linked diseases (e.g., a pregnant female) may wish to determine the sex of the fetus to assess the risk of the fetus inheriting the disease.

[0286] In some implementations, predictions of fetal sex or sex-related disorders can be determined using the methods, systems, machines, devices, or non-transitory computer-readable storage media described herein. Sex determination is typically based on sex chromosomes. Humans have two sex chromosomes, the X and Y chromosomes. The Y chromosome contains the gene SRY, which initiates embryonic development into a male. The Y chromosome in humans and other mammals also contains other genes required for the production of normal sperm.

[0287] In some embodiments, the method for determining fetal sex may further include determining fetal fraction and / or the presence or absence of fetal chromosomal alterations. The determination of the presence or absence of fetal sex can be performed in a suitable manner, non-limiting examples of which include karyotype analysis, amniocentesis, circulating cell-free nucleic acid analysis, cell-free fetal DNA analysis, nucleotide sequence analysis, sequence reading quantification, targeted methods, amplification-based methods, mass spectrometry-based methods, differential methylation-based methods, differential digestion-based methods, polymorphism-based methods, hybridization-based methods (e.g., using probes), etc.

[0288] Medical disorders and medical symptoms

[0289] The methods described herein can be used for any suitable medical disorder or medical condition. Non-limiting examples of medical disorders or medical conditions include cell proliferation disorders and conditions, wasting disorders and conditions, degenerative disorders and conditions, autoimmune disorders and conditions, preeclampsia, chemical or environmental toxicity, liver injury or disease, kidney injury or disease, vascular disease, hypertension, and myocardial infarction.

[0290] In some embodiments, cell proliferation disorders or conditions are cancers of the liver, lungs, spleen, pancreas, colon, skin, bladder, eye, brain, esophagus, head, neck, ovary, testis, prostate, etc., or combinations thereof. Non-limiting examples of cancer include hematopoietic neoplastic disorders involving proliferative / tumor cells of hematopoietic origin (e.g., those arising from the bone marrow, lymphoid, or erythroid lineages, or their progenitor cells) and may arise from poorly differentiated acute leukemia (e.g., erythroblastic leukemia and acute megakaryocytic leukemia). Certain bone marrow disorders include, but are not limited to, acute early myeloid leukemia (APML), acute myeloid leukemia (AML), and chronic myeloid leukemia (CML). Certain malignant lymphoproliferative disorders include, but are not limited to, acute lymphoblastic leukemia (ALL), including B-cell ALL and T-cell ALL, chronic lymphoblastic leukemia (CLL), early lymphoblastic leukemia (PLL), hairy cell leukemia (HLL), and Waldenström macroglobulinemia (WM). Certain forms of malignant lymphoma include, but are not limited to, non-Hodgkin's lymphoma and its variants, peripheral T-cell lymphoma, adult T-cell leukemia / lymphoma (ATL), cutaneous T-cell lymphoma (CTCL), large granular lymphoblastic leukemia (LGF), Hodgkin's disease, and Reedberg's disease. Disorders of cell proliferation are sometimes non-endocrine tumors or endocrine tumors. Examples of non-endocrine tumors include, but are not limited to, adenocarcinoma, acinar cell carcinoma, adenosquamous carcinoma, giant cell tumor, intraductal papillary mucinous neoplasm, mucinous cystadenocarcinoma, pancreatoblastoma, serous cystadenoma, solid and pseudopapillary tumors. Endocrine tumors are sometimes islet cell tumors.

[0291] In some implementations, the consumption disorder or condition, or the degenerative disorder or condition, is cirrhosis, amyotrophic lateral sclerosis (ALS), Alzheimer's disease, Parkinson's disease, multiple system atrophy, atherosclerosis, progressive supranuclear palsy, Tai-Sachs disease, diabetes, heart disease, keratoconus, inflammatory bowel disease (IBD), prostatitis, osteoarthritis, osteoporosis, rheumatoid arthritis, Huntington's disease, chronic traumatic encephalopathy, chronic obstructive pulmonary disease (COPD), tuberculosis, chronic diarrhea, acquired immunodeficiency syndrome (AIDS), superior mesenteric artery syndrome, etc., or combinations thereof.

[0292] In some implementations, the autoimmune disorder or condition is acute disseminated encephalomyelitis (ADEM), Addison's disease, alopecia areata, ankylosing spondylitis, antiphospholipid antibody syndrome (APS), autoimmune hemolytic anemia, autoimmune hepatitis, autoimmune inner ear disease, bullous pemphigoid, gluten intolerance, Chagas disease, chronic obstructive pulmonary disease, Crohn's disease (an idiopathic inflammatory bowel disease "IBD"), dermatomyositis, type 1 diabetes, endometriosis, Goodpasture syndrome, Graves' disease, Guillain-Barré syndrome (GBS), Hashimoto's disease, and other autoimmune disorders. Hidradenitis suppurativa, idiopathic thrombocytopenic purpura, interstitial cystitis, lupus erythematosus, mixed connective tissue disease, scleroderma, multiple sclerosis (MS), myasthenia gravis, somnolence, euromyotonia, pemphigus vulgaris, pernicious anemia, polymyositis, primary biliary cirrhosis, rheumatoid arthritis, schizophrenia, scleroderma, Sjögren's syndrome, temporal arteritis (also known as "giant cell arteritis"), ulcerative colitis (an idiopathic inflammatory bowel disease "IBD"), vasculitis, vitiligo, Wegener's granulomatosis, etc., or combinations thereof.

[0293] Systems, machines, storage media, and interfaces

[0294] Some of the processes and methods described herein would typically not be possible without a computer, microprocessor, software, module, or other machine. The methods described herein are generally computer-executed methods, and one or more portions of the methods are sometimes performed by one or more processors (e.g., microprocessors), a computer, or a microprocessor-controlled device. Implementations of the methods described herein are generally applicable to the same or related processes executed by instructions in the systems, devices, and computer program products described herein. Implementations related to the methods described in this application are generally applicable to the same or related steps performed by a non-transitory computer-readable storage medium on which an executable program is stored, wherein the program provides instructions to a microprocessor to perform the method, or portions thereof. The term “non-transitory” is clearly defined herein, excluding transient, propagating signals (e.g., transmitted signals, electronic transmissions, waves (e.g., carrier waves)). The term “non-transitory computer-readable medium” includes all computer-readable media except transient, propagating signals. In some embodiments, the processes and methods described herein are performed by automated methods. In some embodiments, one or more steps and methods described herein are performed by a processor and / or a computer, and / or by combined memory. In some implementations, the automated method encompasses software, modules, microprocessors, peripherals, and / or machines containing these, wherein the method (i) identifies inconsistent readings, (ii) generates mappability variations, (iii) selects a subset of readings based on the mappability variations, (iv) determines candidate breakpoints, (v) filters readings, (vi) compares readings with substantially similar candidate breakpoints, (vii) determines whether a chromosomal alteration exists, or (viii) performs a combination of the above.

[0295] Sequence reads, inconsistent reads, mappable variations, subsets of reads selected based on mappable variations, subsets of filtered reads, subsets of reads containing similar candidate breakpoints, reference reads, and / or reads of the test subject can be further analyzed and processed to determine the presence of chromosomal alterations. Reads, selected reads, subsets of reads, and quantified reads are sometimes referred to as “data” or “data sets.” In some embodiments, data or data sets can be characterized as one or more properties or variables (e.g., sequence-based [e.g., GC content, specific nucleotide sequence, etc.], function-specific [e.g., expressed gene, oncogene, etc.], location-based [genome-specific, chromosome-specific, etc., and combinations thereof]. In some embodiments, data or data sets can be organized into matrices of two or more dimensions based on one or more properties or variables. Data organized into matrices can be categorized using any suitable property or variable. Non-limiting examples of data in a matrix include data organized by reference candidate breakpoints, candidate breakpoints of the test sample, reference Z-scores, sample Z-scores, and breakpoint locations.

[0296] Machines, software, and interfaces can be used to perform the methods described herein. Using these machines, software, and interfaces, users can access, request, query, or determine options for using specific information, procedures, or methods (such as mapping sequence reads, generating sequence read subsequences, mapping sequence read subsequences, generating relationships, generating mappability changes, selecting subsets of reads, comparing reads, and / or providing results). For example, the information, procedures, or methods may involve implementing statistical analysis algorithms, statistical significance algorithms, statistical algorithms, repeated steps, validation algorithms, and graphical displays. In some implementations, data sets can be input by the user as input information, one or more data sets can be downloaded via any suitable hardware medium (such as flash memory), and / or users can send data sets from one system to another for subsequent processing and / or to provide results (e.g., sending sequence read data from a sequencer to a computer system to locate sequence reads; sending located sequence data to a computer system for processing and generating results and / or reports).

[0297] A system typically includes one or more devices. Each device includes one or more memories, one or more microprocessors, and instructions. When a system includes two or more devices, some or all of the devices may be located in the same location, some or all of the devices may be located in different locations, all of the devices may be located in one location, and / or all of the devices may be located in different locations. When a system includes two or more devices, some or all of the devices may be located in the same location as the user, some or all of the devices may be located in different locations as the user, all of the devices may be located in the same location as the user, and / or all of the devices may be located in one or more different locations as the user.

[0298] Systems sometimes include computing devices or sequencing devices, or computing devices and sequencing devices (i.e., sequencing machines and / or computing machines). The devices described herein are sometimes machines. Sequencing devices are typically configured to receive physiological nucleic acids and generate signals corresponding to the nucleotide bases of the nucleic acids. Sequencing devices are typically “loaded” with a sample containing nucleic acids, and the nucleic acids in the sample loaded into the sequencing device generally undergo a nucleic acid sequencing process. As used herein, the term “loading sequencing device” refers to contacting a portion of the sequencing device (e.g., a flow cell) with a nucleic acid sample, a portion of the sequencing device configured to receive the sample for a nucleic acid sequencing process. In some embodiments, multiple sample nucleic acids are used to load the sequencing device. Variants are sometimes generated by modifying the sample nucleic acids to a form suitable for nucleic acid sequencing (e.g., by ligation (e.g., by adding an adaptor to the end of the sample nucleic acid), amplification, restriction digestion, etc., or combinations thereof). Sequencing devices are typically partially configured to perform a suitable DNA sequencing method, which generates signals (e.g., electronic signals, detector signals, images, etc., or combinations thereof) corresponding to the nucleotide bases of the loaded nucleic acids.

[0299] One or more signals corresponding to each base in a DNA sequence are typically processed and / or converted into base determinations (e.g., specific nucleotide bases, such as guanine, cytosine, thymine, uracil, adenine, etc.) through appropriate processes. The set of base determinations from the loaded nucleic acid is typically processed and / or composed into one or more sequence reads. In embodiments where multiple sample nucleic acids are sequenced simultaneously (i.e., multiplexed), appropriate demultiplexing methods can be used to associate specific reads with the sample nucleic acids from which they originate. Sequence reads can be aligned to a reference genome using appropriate methods, and reads that are partially aligned to the reference genome can be counted, as described herein.

[0300] In the system, sequencing equipment is sometimes associated with and / or includes one or more computing devices. These computing devices are sometimes configured to perform one or more of the following processes: generating base determinations from sequencing equipment signals, assembling reads (e.g., generating reads), demultiplexing reads, aligning reads to a reference genome, counting reads aligned with genomic portions in the reference genome, etc. These computing devices are sometimes configured to perform one or more of the following additional processes: normalizing read counts (e.g., reducing or removing biases), generating one or more determinations (e.g., determining fetal fraction, fetal ploidy, fetal sex, fetal chromosome count, results, presence of genetic variations (e.g., presence of fetal chromosomal aneuploidy (e.g., trisomy 13, 18, and / or 21)) etc.

[0301] In some embodiments, a computing device is associated with a sequencing device, and in some embodiments, this computing device performs most or all of the following processes: generating base determination from the sequencing device signal, assembling reads, demultiplexing reads, aligning reads, and counting reads aligned with genomic portions of a reference genome, normalizing read counts, and generating one or more results (e.g., fetal score, presence of a specific genetic variant). In a later embodiment, one computing device associated with the sequencing device typically includes one or more processors (e.g., microprocessors) and memory with instructions executed by the one or more processors to perform the process. In some embodiments, this computing device may be a single-core or multi-core computing device local to the sequencing device (e.g., located in the same location (e.g., same address, same structure, same layer, same chamber, etc.)). In some embodiments, this computing device is integrated with the sequencing device.

[0302] In some implementations, multiple computing devices in the system are associated with the sequencing device, and a subset of the total processes performed by the system may be distributed or partitioned among specific computing devices in the system. A subset of the total number of processes may be partitioned among two or more computing devices or groups thereof in any suitable combination. In some implementations, a first computing device or group thereof performs base determination from the sequencing device signal, assembles reads, and demultiplexes the reads; a second computing device or group thereof performs alignment and counting of reads mapped to portions of a reference genome; and a third computing device or group thereof performs normalization of the read counts and provides one or more results. In systems comprising two or more computing devices or groups thereof, each specific computing device may include memory, one or more processors, or a combination thereof. Multi-computing device systems sometimes include one or more suitable servers local to the sequencing device, and sometimes include one or more suitable servers not local to the sequencing device (e.g., web servers, instant servers, application servers, remote file servers, cloud servers (e.g., cloud environments, cloud computing)).

[0303] Devices in different system configurations can generate different types of output data. For example, a sequencing device can output base signals, and this base signal output data can be transferred to a computing device that converts the base signal data into base determinations. In some embodiments, base determinations are output data from one computing device and transferred to another computing device for generating sequence reads. In some embodiments, base determinations are not output data from a specific device, but rather from the same device used to receive sequencing device base signals to generate sequence reads. In some embodiments, one device receives sequencing device count signals, generates base determinations, sequence reads, and demultiplexed sequence reads, and outputs the demultiplexed sequence reads of the sample, which can be transferred to another device or group thereof for aligning the sequencing reads with a reference genome. In some embodiments, one device or group thereof can output aligned sequence reads mapped to a portion of a reference genome (e.g., a SAM or BAM file), and this output data can be transferred to a second computing device or group thereof for normalizing the sequence reads (e.g., normalizing the counts of the sequence reads) and generating results (e.g., fetal score and / or the presence of fetal trisomy). Output data from one device can be transferred to a second device in any suitable manner. For example, output processing from one device may sometimes be located on a physical storage device, and that storage device may be transported and connected to the second device to which the output data is transferred. Sometimes output data is stored in a device within a database, and the second device evaluates the output data from the same database.

[0304] In some implementations, this is used to interact with devices (e.g., computing devices, sequencing devices). For example, a user may query software settings, which may then obtain a data set via an internet access point. In other implementations, a programmable microprocessor may be instructed to obtain a suitable data set based on given parameters. The programmable microprocessor may also prompt the user to select one or more data set options selected by the microprocessor based on given parameters. The programmable microprocessor may also prompt the user to select one or more data set options selected by the microprocessor based on information discovered through the internet, other internal or external information, etc. Options may be selected to include one or more data characteristic selections of a method, machine, device (multiple devices, referred to herein in plural form), computer program or non-transitory computer-readable storage medium on which an executable program is stored, one or more statistical algorithms, one or more statistical analysis algorithms, one or more statistical significance algorithms, repeated steps, one or more confirmation algorithms, and one or more graphical displays.

[0305] The systems described herein may include common components of computer systems, such as web servers, laptop systems, desktop systems, handheld systems, personal digital assistants, and computer self-service terminals. The computer system may include one or more input methods, such as a keyboard, touchscreen, mouse, voice recognition, or other methods, to allow users to input data into the system. The system may also include one or more outputs, including but not limited to displays (such as CRTs or LCDs), speakers, fax machines, printers (such as laser, inkjet, impact, black-and-white, or color printers), or other means for providing visual, auditory, and / or hard copy outputs (such as results and / or reports) of information.

[0306] In the system, the inputs and outputs can be connected to a central processing unit, which may contain a microprocessor that executes program instructions, a memory that stores program code and data, and other components. In some embodiments, the processing may be implemented as a single-user system located in a single geographic location. In some embodiments, the processing may be implemented as a multi-user system. In the case of multi-user execution, multiple central processing units can be connected via a network. The network may be local, cover a single room in a part of a building, an entire building, span multiple buildings, across regions, across countries, or globally. The network may be private, owned and controlled by a provider, or may be implemented as a network-based service where users access web pages to enter or retrieve information. Thus, in some embodiments, the system includes one or more machines that can be located or remotely controlled by a user. A user can access more than one machine in one or more locations, and data can be plotted and / or processed in a serial and / or parallel manner. Therefore, any suitable architecture and control can be used to plot and / or process data using multiple machines, such as local networks, remote networks, and / or "cloud" computing platforms.

[0307] In some implementations, the system may include a communication interface. The communication interface enables the transfer of software and data between the computer system and one or more external devices. Non-limiting examples of a communication interface may include a modem, a network interface (e.g., an Ethernet card), a communication port, a PCMCIA slot, and cards. The software and data transferred via the communication interface are typically in the form of signals, which can be electronic, electromagnetic, optical, and / or other signals that can be received by the communication interface. Signals are often provided to the communication interface via channels. Channels often carry signals and can be implemented using wires or cables, optical fibers, telephone lines, mobile phone connections, RF connections, and other communication channels. Therefore, in one embodiment, a communication interface may be used to receive signal information that can be determined by a signal detection module.

[0308] Data can be input by any suitable device and / or method, including but not limited to human input devices or direct data input (DDE) devices. Non-limiting examples of human input devices include keyboards, conceptual keyboards, touchscreens, light pens, mice, trackballs, joysticks, graphic tablets, scanners, digital cameras, video digitizers, and voice recognition devices. Non-limiting examples of DDEs include barcode scanners, magnetic stripe encoding, smart cards, magnetic ink character recognition, optical character recognition, optical mark recognition, and revolving documents.

[0309] In some embodiments, the output of a sequencing device or apparatus can be used as data that can be input via an input device. In some embodiments, located sequence reads can be used as data that can be input via an input device. In some embodiments, nucleic acid fragment size (e.g., length) can be used as data that can be input via an input device. In some embodiments, the output from a nucleic acid capture step (e.g., genomic region source data) can be used as data that can be input via an input device. In some embodiments, a combination of nucleic acid fragment size (e.g., length) and data from a nucleic acid capture step (e.g., genomic region source data) can be used as data that can be input via an input device. In some embodiments, simulated data is generated using a computer-simulated (in silico) method, and said simulated data is used as data that can be input via an input device. The term "computer simulation (in silico)" refers to data (e.g., sequence read subsequences), data manipulation, research, and experimentation performed using a computer. Computer simulation processes include, but are not limited to, mapping sequence reads according to the processes described herein, generating sequence read subsequences, mapping reads and read subsequences, and processing mapped sequence reads.

[0310] The system may include software for running the methods described herein, and the software may include one or more modules (such as a sequencing module, a logic processing module, and a data display management module) for running the methods. As used herein, software refers to computer-readable program instructions that perform computer operations when executed by a computer. One or more microprocessor-executable instructions are sometimes provided as executable code that, when run, causes one or more microprocessors to perform the methods of the present invention.

[0311] The modules described herein may exist in software form, and the instructions (e.g., procedures, routines, subroutines) built into the software may be executed or performed by a microprocessor. For example, a module (e.g., a software module) is a part of a program that performs specific methods and tasks. The term "module" refers to an independent functional unit that can be used in a larger device or software system. A module may include a set of instructions to perform the module's functions via one or more microprocessors. The instructions of a module may be executed using a suitable programming language, suitable software, and / or code written in a suitable language (e.g., computer programming languages ​​known in the art) and / or an operating system, non-limiting examples of which include UNIX, Linux, Oracle, Windows, Ubuntu, ActionScript, C, C++, C#, Haskell, Java, JavaScript, Objective-C, Perl, Python, Ruby, Smalltalk, SQL, Visual Basic, COBOL, Fortran, UML, HTML (e.g., PHP), PGP, G, R, S, etc., or combinations thereof. In some embodiments, the modules described herein include code (e.g., scripts) written in S or R using a suitable package (e.g., an S package or an R package). R, R source code, R programs, R packages, and R archives are available for download from mirror sites (R Comprehensive Archive Network (CRAN) [online], [launched 2013-04-24], retrieved from the Internet via the World Wide Web Uniform Resource Locator: cran.us.r-project.org). CRAN is a global network of FTP and web servers that stores archives and the latest versions of the same R code.

[0312] Modules can transform data and / or information. One or more modules can be used in the methods described herein, and non-limiting examples include sequence modules, mapping modules, inconsistent reading identification modules, fragmentation modules, reading selector modules, mapping characterization modules, breakpoint modules, comparison modules, filtering modules, plotting modules, result modules, etc., or combinations thereof. For example, Figure 6The embodiment shown is an example. The inconsistency reading identification module 15 sends inconsistency readings to the mapping characterization module 50, which is configured to accept inconsistency readings from the inconsistency reading identification module 15. The mapping characterization module 50 can send mapping characterizations to the reading selector module 120, which is configured to accept mapping characterizations from the mapping characterization module 50. The reading selector module 120 can send a selected subset of readings (e.g., pairs of inconsistency readings) to the comparison module 130, which is configured to accept the selected subset of readings from the reading selector module 120. The comparison module 130 can generate a comparison (e.g., comparing the following: (i) the number of inconsistency reading companions of samples associated with the candidate breakpoint and optionally associated with one or more substantially similar breakpoints and (ii) the number of inconsistency reading companions of references associated with the candidate breakpoint and optionally associated with one or more substantially similar breakpoints) and send the comparison to the results module 140, which is configured to accept the comparison. The results module 140 can then determine whether translocation exists in the test object and provide the results to the end user or send the results to another module (e.g., a plotting module). Modules are sometimes microprocessor-controlled. In some embodiments, a module or device including one or more modules gathers, collects, receives, acquires, accesses, retrieves, provides, and / or transfers data and / or information to or from other modules, devices, components, peripherals, or operators. In some embodiments, data and / or information (e.g., sequencing reads) is provided to a module via a device comprising one or more of the following components: one or more flow cells, cameras, detectors (e.g., photodetectors, photocells, electrical detectors (e.g., quadrature amplitude modulation detectors, frequency and phase modulation detectors, phase-locked loop detectors), counters, sensors (e.g., pressure, temperature, volume, flow, weight sensors), fluid manipulation devices, data input devices (e.g., keyboards, mice, scanners, voice recognition software and microphones, styluses, etc.), printers, displays (e.g., LEDs, LCTs, or CRTs), etc., or combinations thereof. For example, sometimes operators of such devices provide constants, thresholds, formulas, or predetermined values ​​to the module. Modules are typically configured to receive data from microprocessors and / or other devices. The module may transfer data and / or information to or from another suitable module or device. Modules are typically configured to transfer data and / or information to or from another suitable module or machine. Modules are operable and / or transform data and / or information. Data and / or information from or transformed from a module may be transferred to another suitable machine and / or module. A device including a module may include at least one microprocessor. A device including a module includes a microprocessor (e.g., one or more microprocessors) capable of performing and / or executing one or more instructions (e.g., procedures, routines, and / or subroutines) of the module. In some embodiments, the module operates with one or more external processors (e.g., internal or external networks, servers, storage devices, and / or storage networks (e.g., the cloud)).

[0313] Data and / or information may be in a suitable form. For example, data and / or information may be digital or analog. In some embodiments, data and / or information may sometimes be packets, bytes, characters, or bits. In some embodiments, data and / or information may be any collected, aggregated, or useful data or information. Non-limiting examples of data and / or information include suitable media, pictures, videos, sounds (e.g., audible or inaudible frequencies), numbers, constants, values, objects, time, functions, instructions, graphs, references, sequences, readings, mapped readings, levels, ranges, thresholds, signals, displays, representations, or transformations thereof. The module may accept or receive data and / or information, transform data and / or information into a second form, and provide or transfer such second form to a device, peripheral device, component, or other module. The module may perform one or more of the following non-restrictive functions: for example, mapping sequence reads, identifying inconsistent read pairs, generating sequence read subsequences, characterizing the mappability of multiple sequence read subsequences, generating mappability variations, generating mappability thresholds, filtering, selecting a subset of inconsistent read partners based on mappability variations and / or mappability thresholds, identifying candidate breakpoints, identifying breakpoints, plotting, generating comparisons (e.g., comparing (i) the number of inconsistent read partners from the sample associated with the candidate breakpoint and optionally one or more substantially similar breakpoints with (ii) the number of inconsistent read partners from the reference associated with the candidate breakpoint and said optionally one or more substantially similar breakpoints) and / or determining results (e.g., determining whether a chromosomal alteration exists). In some embodiments, the microprocessor may execute instructions within the module. In some embodiments, one or more microprocessors are required to execute instructions within the module or group of modules. The module may provide data and / or information to and receive data and / or information from other modules, devices, or sources.

[0314] Computer program products are sometimes materialized on non-transitory computer-readable media, and sometimes physically materialized on non-transitory computer-readable media. Modules are sometimes stored in non-transitory computer-readable media (e.g., disks, drives) or memory (e.g., random access memory). Modules and microprocessors capable of executing instructions from modules may be located within a device or in different devices. Modules and / or microprocessors capable of executing instructions from modules may be located at the same location of the user (e.g., a local network) or at a different location of the user (e.g., a remote network, a cloud system). In embodiments where the method is performed in conjunction with two or more modules, the modules may be located in the same device, one or more modules may be located in different devices in the same physical location, and one or more modules may be located in different devices in different physical locations.

[0315] In some embodiments, the device includes at least one microprocessor for executing instructions within a module. Sequence reads mapped to a reference genome are sometimes accessed via a microprocessor that executes instructions to perform the methods described herein. Sequence reads accessed via a microprocessor may be in the system's memory and may be accessed and placed in the system's memory after access. In some embodiments, the device includes a microprocessor (e.g., one or more microprocessors) capable of performing and / or executing one or more instructions (e.g., procedures, routines, and / or subroutines) of a module. In some embodiments, the device includes multiple microprocessors, such as microprocessors that work in a cooperative and parallel manner. In some embodiments, the device operates with one or more external microprocessors (e.g., internal or external networks, servers, storage devices, and / or storage networks (e.g., the cloud)). In some embodiments, the device includes modules. In some embodiments, the device includes one or more modules. The modules in the device are typically capable of receiving and transferring one or more types of data and / or information from and to other modules. In some embodiments, the device includes peripheral devices and / or components. In some embodiments, the apparatus may include one or more peripheral devices or components that can transmit and receive data and / or information to and from other modules, peripheral devices and / or components. In some embodiments, the apparatus interacts with peripheral devices and / or components that provide data and / or information. In some embodiments, peripheral devices and components assist the apparatus in performing its functions or interact directly with modules. Non-limiting examples of peripheral devices and / or components include suitable computer peripherals, I / O or storage methods or devices, including but not limited to scanners, printers, displays (e.g., monitors, LEDs, LCTs, or CRTs), cameras, microphones, tablet computers (e.g., writing tablets), touchscreens, smartphones, mobile phones, USB I / O devices, USB memory, keyboards, computer mice, digital pens, modems, hard disks, jump engines, flash drives, microprocessors, servers, CDs, DVDs, graphics cards, dedicated I / O devices (e.g., sequence generators, photocells, photoamplifiers, optical readers, sensors, etc.), liquid handling components, network interaction controllers, ROM, RAM, wireless transmission devices (Bluetooth, WiFi, etc.), the World Wide Web (WWW), networks, computers, and / or other modules.

[0316] Software is typically provided on a program product containing program instructions recorded on a non-transitory computer-readable medium, including but not limited to magnetic media (e.g., floppy disks, hard disks, ROMs, and magnetic tapes), optical media (e.g., CD-ROMs, DVDs, etc.), magneto-optical disks, flash drives, RAM, and other such media capable of recording the program instructions. In online execution, servers and websites maintained by an organization can be configured to provide software downloads to remote users, or remote users can access the software remotely using remote systems maintained by the organization. The software can obtain or receive input information. The software may include modules specifically for acquiring or receiving data (such as a data receiving module for receiving sequence reads and / or location reads) and may include modules specifically for processing data (such as processing modules for processing data, such as filters, providing results, and / or reports). The terms "acquiring" and "receiving" input information refer to data (such as sequence reads, location reads) received via computer communication from a local or remote location, manual data input, or any other method of receiving data. Input information may be generated at the same location where it is received, or it may be generated at a different location and transmitted to the receiving location. In some embodiments, the input information is modified before processing (e.g., placed in a processing-friendly form (e.g., a table)).

[0317] In some embodiments, a computer program product is provided, such as a computer program product comprising a non-transitory computer-usable medium containing non-transitory computer-readable program code adapted to run to perform a method comprising: (a) identifying inconsistent read pairs from paired end sequence reads, wherein the paired end sequence reads are reads from cyclic cell-free nucleic acids from a sample of a test subject, thereby identifying inconsistent read partners;

[0318] (b) Characterize the mappability of multiple sequence read subsequences of each sequence read buddy aligned with a reference genome, wherein each sequence read subsequence of each inconsistent read buddy has a different length, thereby providing variations in the mappability of inconsistent read buddies.

[0319] (c) Select a subset of the inconsistent reading companions based on the mappability change in (b), wherein the subset includes readings containing candidate breakpoints;

[0320] (d) For the inconsistent reading companions in the subset selected in (c), compare (i) the number of inconsistent reading companions from the sample associated with the candidate breakpoint and optionally one or more substantially similar breakpoints with (ii) the number of inconsistent reading companions from the reference associated with the candidate breakpoint and said optionally one or more substantially similar breakpoints, thereby generating a comparison; and

[0321] (e) Determine whether one or more chromosomal alterations exist in the sample based on the comparison in (d).

[0322] Software can be used to perform one or more steps of the methods or processes described herein, including but not limited to: identifying inconsistent reads (e.g., 15), generating sequence read subsequences, characterizing the mappability of sequence read subsequences, generating mappability variations (e.g., 50), identifying candidate breakpoints and / or breakpoints, selecting a subset of read partners (e.g., 120), comparing subsets of reads containing similar breakpoints (e.g., 130), filtering (e.g., 20, 30, 40, 50, 70, 80, 90, 100, and 110), data processing, determining the presence of chromosomal alterations (e.g., 140), generating results, and / or providing one or more recommendations based on the generated results, as described in detail below. The term "software" herein refers to a non-transitory computer-readable storage medium having an executable program thereon, wherein the program provides instructions to a microprocessor to perform a function (e.g., a method). In some embodiments, a non-transitory computer-readable storage medium having an executable program thereon instructs a microprocessor to identify inconsistent read pairs from paired end sequence reads, wherein the paired end sequence reads are reads of circulating cell-free nucleic acids from a sample of test subjects, thereby identifying inconsistent read partners. In some embodiments, a non-transient computer-readable storage medium instruction microprocessor having an executable program thereon characterizes the mappability of multiple sequence read subsequences of each sequence read companion aligned with a reference genome, wherein each sequence read subsequence of each inconsistent read companion has a different length, thereby providing mappability variations and candidate breakpoints for inconsistent read companions. In some embodiments, the non-transient computer-readable storage medium instruction microprocessor having an executable program thereon selects a subset of the inconsistent read companions based on mappability variations and / or mappability thresholds. In some embodiments, the non-transient computer-readable storage medium instruction microprocessor having an executable program thereon compares (i) the number of inconsistent read companions from the sample associated with candidate breakpoints and optionally one or more substantially similar breakpoints with (ii) the number of inconsistent read companions from the reference associated with candidate breakpoints and the optionally one or more substantially similar breakpoints. In some embodiments, the non-transient computer-readable storage medium instruction microprocessor having an executable program thereon determines whether the sample has one or more chromosomal alterations. In some embodiments, the non-transient computer-readable storage medium instruction microprocessor having an executable program thereon:

[0323] (a) Identifying inconsistent read pairs from paired end sequence reads, wherein the paired end sequence reads are reads from cyclic cell-free nucleic acids from a sample of the test subject, thereby identifying inconsistent read buddies;

[0324] (b) Characterize the mappability of multiple sequence read subsequences of each sequence read buddy aligned with a reference genome, wherein each sequence read subsequence of each inconsistent read buddy has a different length, thereby providing variations in the mappability of inconsistent read buddies.

[0325] (c) Select a subset of the inconsistent reading companions based on the mappability change in (b), wherein the subset includes readings containing candidate breakpoints;

[0326] (d) For the inconsistent reading companions in the subset selected in (c), compare (i) the number of inconsistent reading companions from the sample associated with the candidate breakpoint and optionally one or more substantially similar breakpoints with (ii) the number of inconsistent reading companions from the reference associated with the candidate breakpoint and said optionally one or more substantially similar breakpoints, thereby generating a comparison; and

[0327] (e) Determine whether one or more chromosomal alterations exist in the sample based on the comparison in (d).

[0328] In some implementations, the software can contain one or more algorithms. Algorithms can be used to process data and / or provide results or reports according to a finite sequence of instructions. Algorithms are often defined instruction lists used to accomplish a task. Starting from a starting state, the instructions can describe a computation performed through a defined series of consecutive states and terminating in a final ending state. Transitions from one state to the next do not need to be deterministic (e.g., some algorithms incorporate arbitrariness). As non-limiting examples, algorithms can be search algorithms, classification algorithms, merge algorithms, numerical algorithms, graphical algorithms, string search algorithms, modeling algorithms, computational geometry (geometry) algorithms, combinatorial algorithms, machine learning algorithms, cryptographic algorithms, data compression algorithms, analytical algorithms, etc. Algorithms can contain one algorithm or a combination of two or more algorithms. Algorithms can be any suitable complexity classification and / or parameterized complexity. Algorithms can be used for computation and / or data processing, and in some implementations can be used in deterministic or probabilistic / predictive methods. Algorithms can be embedded into a computer environment using a suitable programming language (non-limiting examples are C, C++, Java, Perl, R, S, Python, Fortran, etc.). In some implementations, the algorithm can be constructed or improved to include error tolerance, statistical analysis, statistical significance, and / or comparison with other information or data sets (as in applications using neural networks or cluster algorithms).

[0329] In some implementations, several algorithms can be embedded in software for ease of use. In some implementations, these algorithms can be trained on raw data. For various new raw data samples, the trained algorithms can generate representative processed datasets or results. Compared to the processed parent dataset, the processed dataset sometimes has reduced complexity. In some implementations, based on the processed dataset, the implementation of the trained algorithms can be evaluated according to sensitivity and specificity. In some implementations, algorithms with the highest sensitivity and / or specificity can be identified and utilized.

[0330] In some implementations, simulated data can assist in data processing, such as through training or testing an algorithm. In some implementations, simulated data comprises multiple hypothetical samples of different groups of sequence readings. Simulated data may be based on possible expected scenarios in a real population or may be distorted to test algorithms and / or assign correct classifications. Simulated data also refers to “real” data herein. In some implementations, simulation can be performed by a computer program. One possible step using a set of simulated data is to evaluate the confidence level of the identified outcome, such as how well the random sample matches or best represents the original data. One approach is to calculate a probability value (p-value) that assesses the probability that a random sample is better than a selected sample. In some implementations, an empirical model can be evaluated, where it is assumed that at least one sample matches a reference sample (with or without resolved variation). In some implementations, other distributions, such as the Poisson distribution, can be used to define the probability distribution.

[0331] In some implementations, the system may include one or more microprocessors. The microprocessors may be connected to a communication bus. The computer system may include main memory (often random access memory (RAM)) and may also include secondary memory. In some implementations, the memory includes non-transitory computer-readable storage media. Secondary memory may include, for example, hard disk devices and / or removable storage devices, representing floppy disk devices, magnetic tape devices, optical disc devices, memory cards, etc. Removable storage drives frequently read from and / or write to removable storage units. Non-limiting examples of removable storage units include floppy disks, magnetic tapes, optical discs, etc., capable of reading from or writing to removable storage drives. Removable storage units may include non-transitory computer-readable storage media containing computer software and / or data.

[0332] Microprocessors can execute software within a system. In some implementations, microprocessors can be programmed to automatically perform tasks that the user, as described herein, can perform. Therefore, the microprocessor, or the algorithm executed by such a microprocessor, requires little to no monitoring or input from the user (e.g., software can be written to automate the implementation of functions). In some implementations, the processing is so complex that a single individual or group of individuals cannot perform the processing within a sufficiently short timeframe to determine the presence of chromosomal alterations.

[0333] In some implementations, the second memory may include other similar methods that allow computer programs or other instructions to be loaded into the computer system. For example, the system may include removable storage units and interactive devices. Non-limiting examples of such systems may include program modules and module interfaces (such as those found in video game devices), removable storage chips (such as EPROM or PROM), and associated sockets and other removable storage units and interfaces that allow software and data to be transferred from the removable storage unit to the computer system.

[0334] In some embodiments, an entity may generate, map, identify, and use inconsistent read pairs within the methods, systems, machines, apparatuses, or computer program products described herein. In some embodiments, within the methods, systems, machines, apparatuses, or computer program products described herein, sequence reads mapped to a reference genome may sometimes be transferred from one entity to a second entity for its use.

[0335] In some embodiments, an entity generates sequence reads and a second entity maps those sequence reads to a reference genome. The second entity sometimes identifies inconsistent reads and employs them in the methods, systems, machines, or computer program products described herein. In some embodiments, the second entity transfers the mapped reads to a third entity, and the third entity identifies and employs the inconsistent reads in the methods, systems, machines, or computer program products described herein. In some embodiments, the second entity identifies and transfers the inconsistent reads to a third entity, and the third entity employs the identified inconsistent reads in the methods, systems, machines, or computer program products described herein. In embodiments involving a third entity, the third entity is sometimes the same as the first entity. That is, the first entity sometimes transfers sequence reads to a second entity, the second entity may map sequence reads to a reference genome and / or identify inconsistent reads, and the second entity may transfer the mapped and / or inconsistent reads to a third entity. The third entity may sometimes employ the mapped and / or inconsistent reads in the methods, systems, apparatus, or computer program products described herein, wherein the third entity is sometimes the same as the first entity, and sometimes different from the first or second entity.

[0336] In some implementations, an entity obtains blood from a pregnant female, optionally separates nucleic acid blood from the blood (e.g., from plasma or serum), and transfers the blood or nucleic acid to a second entity, which generates sequence reads from the nucleic acid.

[0337] Figure 8 This illustrates a non-limiting example of computing environment 510, in which various systems, methods, algorithms, and data structures described herein can be executed. Computing environment 510 is merely one embodiment of a suitable computing environment and is not intended to limit the use or scope of functionality of the systems, methods, and data structures described herein. Computing environment 510 should also not be construed as any dependency or requirement on any of the components or combinations thereof shown in computing environment 510. In some implementations, [the following may be used] Figure 8 The systems, methods, and data structures described herein are subsets of those described. The systems, methods, and data structures described herein can be operated on by a wide range of computing system environments or configurations for other general or specific purposes. Examples of known suitable computing systems, environments, and / or configurations include, but are not limited to, personal computers, server computers, thin clients, thick clients, handheld or lap devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable client electronics, network PCs, minicomputers, mainframe computers, and computing environments that include any of the aforementioned systems or devices.

[0338] Figure 8 The operating environment 510 includes a conventional computing device, in the form of a computer 520, including a processing unit 521, system memory 522, and a system bus 523 operatively coupling various system components (including system memory 522) to the processing unit 521. There may be only one or more processing units 521, thus the microprocessor of the computer 520 may include a single central processing unit (CPU) or multiple processing units, commonly referred to as a parallel processing environment. The computer 520 may be a conventional computer, a distributed computer, or any other type of computer.

[0339] System bus 523 can be any of several bus structures, including memory bus or memory controller, peripheral bus, and local bus, using any of the various bus architectures. System memory can also be simply referred to as memory, including only read-only memory (ROM) 524 and random access memory (RAM). Basic input / output system (BIOS) 526 is stored in ROM 524. The BIOS contains basic routines, for example, to assist in transferring information between components of computer 520 at startup. Computer 520 may also include hard disk drive interface 527 for reading from and writing to a hard disk (not shown), disk drive 528 for reading from or writing to a removable disk 529, and optical disk drive 530 for reading from or writing to a removable optical disk 531, such as a CD ROM or other optical media.

[0340] Hard disk drive 527, disk drive 528, and optical disk drive 530 are connected to system bus 523 via hard disk drive interface 532, disk drive interface 533, and optical disk drive interface 534, respectively. The drives and their associated computer-readable media provide fixed storage for computer-readable instructions, data structures, program modules, and other data of computer 520. Any type of computer-accessible and data-storing computer-readable media, such as magnetic cartridges, flash memory cards, digital video discs, Bernoulli cylinders, random access memory (RAM), read-only memory (ROM), etc., can be used in this operating environment.

[0341] Many program modules may be stored on a hard disk, disk 529, optical disk 531, ROM 524, or RAM, including an operating system 535, one or more application programs 536, other program modules 537, and program data 538. Users can type commands and information into the personal computer 520 via input devices such as 540 and 542. Other input devices (not shown) may include a microphone, joystick, gamepad, satellite TV antenna, scanner, or the like. These and other input devices are typically connected to the processing unit 521 via a serial port interface 546 coupled to the system bus, but may not be connected via other structures such as a parallel port, game port, or Universal Serial Bus (USB). A monitor 547 or other type of display device is also connected to the system bus 523 via an interface such as a video adapter 548. In addition to the monitor, the computer typically includes other peripheral output devices (not shown) such as speakers and printers.

[0342] Computer 520 can operate in a networked environment, using logical connections to one or more remote computers, such as remote computer 549. These logical connections can be implemented through communication devices coupled to or in part of computer 520 or otherwise. Remote computer 549 can be other computers, servers, routers, network PCs, peer devices, or other common network nodes, and generally includes many or all of the elements described above regarding computer 520, although... Figure 8 Only memory storage device 550 is displayed. Figure 8 The logical connections described include Local Area Networks (LANs) 551 and Wide Area Networks (WANs) 552. These networking environments are common in offices, enterprise-wide computer networks, intranets, and the Internet.

[0343] When used in a LAN networking environment, computer 520 connects to local area network 551 via a network interface or adapter 553, serving as a communication device. When used in a WAN networking environment, computer 520 typically includes a modem 554, a communication device, or other type of communication device for establishing communication over a wide area network 552. Modem 554 can be internal or external and can be connected to system bus 523 via serial port interface 546. In a networking environment, program modules or portions thereof related to computer 520 may be stored in a remote memory storage device. It should be understood that the network connections shown are exemplary, and other means of establishing communication links between computers can be used.

[0344] Implementation methods of certain systems, machines, and computer program products

[0345] Some aspects of the present invention provide a system including a memory and one or more microprocessors, wherein the memory includes instructions, and the one or more microprocessors are configured to perform, according to the instructions, a process for determining whether one or more chromosomal alterations are present in nucleic acids of a sample, the process including...

[0346] (a) Characterizing the mappability of multiple sequence read subsequences with respect to sequence reads, wherein each sequence read has multiple sequence read subsequences of different lengths, and the sequence reads are sequence reads of sample nucleic acids.

[0347] (b) Identify a subset of sequence reads in which the mappability of one or more subsequences changes.

[0348] (c) A comparison is generated by comparing the number of sequence readings from the subset of samples identified in (i)(b) with the number of sequence readings from the subset of references identified in (ii)(b); and

[0349] (d) Determine whether one or more chromosomal alterations exist in the sample based on the comparison in (c).

[0350] Some aspects of the present invention also provide a method including a memory and one or more micro...

Claims

1. A device comprising one or more processors and memory, wherein, The memory contains instructions executable by the one or more processors, and the memory contains nucleic acid sequence reads mapped to a reference genome; wherein the instructions executable by the one or more processors are configured to perform: (a) Identifying inconsistent read pairs from paired end sequence reads, wherein the paired end sequence reads are reads from cyclic cell-free nucleic acids from a sample of the test subject, thereby identifying inconsistent read partners; (b) Characterize the mapability of multiple sequence read subsequences of each sequence read buddy aligned with a reference genome, wherein each sequence read subsequence of each inconsistent read buddy has a different length. (c) Select a subset of the inconsistent reading companions based on the mappability variation, wherein the subset includes readings containing candidate breakpoints; (d) For the inconsistent reading companions in the subset selected in (c), compare (i) the number of inconsistent reading companions from the sample associated with the candidate breakpoint and optionally one or more substantially similar breakpoints with (ii) the number of inconsistent reading companions from the reference associated with the candidate breakpoint and the optionally one or more substantially similar breakpoints, thereby generating a comparison. and (e) In the comparison of (d), the presence of one or more chromosomal alterations in the sample is determined by identifying that the number of inconsistent reading partners from the sample is significantly greater than the number of inconsistent reading partners from the reference, or the absence of one or more chromosomal alterations in the sample is determined if the number of inconsistent reading partners from the sample is not significantly greater than the number of inconsistent reading partners from the reference.

2. The device of claim 1, wherein the one or more chromosomal alterations include chromosomal translocation, chromosomal deletion, chromosomal inversion, or heterologous insertion.

3. The device of claim 1, wherein instructions executable by the one or more processors are configured to determine the location of one or more candidate breakpoints.

4. The device of any one of claims 1 to 3, wherein the characterization in (b) includes a fit relationship between the length of each sequence reading subsequence that generates each inconsistent reading companion and the mappability.

5. The device according to any one of claims 1 to 3, wherein each sequence reading subsequence of each inconsistent read partner is 5 bases or less shorter than the second largest fragment or read partner, or, each sequence reading subsequence of each inconsistent read partner is progressively shorter than the second largest fragment or read partner.

6. The device of claim 4, wherein the mappability variation includes the slope of the fitted relationship.

7. The device as claimed in any one of claims 1 to 3, wherein the selection in (c) is based on a mappability threshold.

8. The device as claimed in any one of claims 1 to 3, comprising instructions executable by one or more processors, configured to filter inconsistent reading companions.

9. The device of claim 8, wherein the filtering includes removing one or both of the inconsistent reading companions.

10. The device of claim 9, wherein the filtering is selected from one or more of the following: (i) removal of low-quality reads, (ii) removal of consistent reads, (iii) removal of PCR replication reads, (iv) removal of reads mapped to mitochondrial DNA, (v) removal of reads mapped to repeat elements, (vi) removal of unmapped reads, (vii) removal of reads containing step-by-step multiple alignments, and (vii) removal of reads mapped to centromeres.

11. The device of claim 9, wherein the filtering includes removing one or more singular mutation events, and / or removing inconsistent reading companions when substantially similar breakpoints exist in the reference.

12. The device according to any one of claims 1 to 3, wherein the location of the break point is identified at single-base resolution.

13. The device as claimed in any one of claims 1 to 3, wherein the presence of a balanced translocation or an unbalanced translocation is determined in (e).

14. The device of any one of claims 1 to 3, wherein determining the presence of translocation in (e) includes identifying, in the comparison in (d), that the number of sequence readings from the sample is significantly greater than that from the reference.

15. The device according to any one of claims 1 to 3, wherein the first fracture point and the second fracture point are identified according to the comparison in (d).

16. The device of claim 15, wherein in (e) the presence of chromosomal alteration is determined based on the first breakpoint and the second breakpoint.

17. The device according to any one of claims 1 to 3, wherein the selection in (c) or the comparison in (d), or the selection in (c) and the comparison in (d), does not include performing cluster analysis.

18. The device as claimed in any one of claims 1 to 3, wherein the comparison in (d) includes determining a confidence level.

19. The device as claimed in claim 18, wherein, Determining the confidence level includes determining the p-value or z-score.

20. The device as claimed in any one of claims 1 to 3, comprising a sequencing machine configured to generate sequence reads, or comprising a machine thereof.

21. The device of any one of claims 1 to 3, wherein the memory comprises sequential readouts, inconsistent readout pairs, inconsistent readout companion subsets, mappability variations, breakpoints, or combinations thereof.

22. The device of any one of claims 1 to 3, wherein the sample is circulating cell-free nucleic acid from a pregnant female carrying a fetus.

23. The device of any one of claims 1 to 3, wherein the sample is circulating cell-free nucleic acid from a subject suffering from or suspected of suffering from a cell proliferation disorder.

24. The device of claim 23, wherein the cell proliferation disorder is cancer.

25. The device according to any one of claims 1 to 3, wherein the presence or absence of one or more chromosomal alterations is determined for a small number of nucleic acid substances.

26. The device of claim 25, wherein the minority nucleic acid material includes fetal nucleic acid or nucleic acid of cancer cells.

27. A non-transitory computer-readable storage medium having thereon an executable program, wherein the program is configured to instruct a microprocessor to perform the following operations: (a) Identifying inconsistent read pairs from paired end sequence reads, wherein the paired end sequence reads are reads from cyclic cell-free nucleic acids from a sample of the test subject, thereby identifying inconsistent read partners; (b) Characterize the mapability of multiple sequence read subsequences of each sequence read buddy aligned with a reference genome, wherein each sequence read subsequence of each inconsistent read buddy has a different length. (c) Select a subset of the inconsistent reading companions based on the mappability variation, wherein the subset includes readings containing candidate breakpoints; (d) For the inconsistent reading companions in the subset selected in (c), compare (i) the number of inconsistent reading companions from the sample associated with the candidate breakpoint and optionally one or more substantially similar breakpoints with (ii) the number of inconsistent reading companions from the reference associated with the candidate breakpoint and the optionally one or more substantially similar breakpoints, thereby generating a comparison. and (e) In the comparison of (d), the presence of one or more chromosomal alterations in the sample is determined by identifying that the number of inconsistent reading partners from the sample is significantly greater than the number of inconsistent reading partners from the reference, or the absence of one or more chromosomal alterations in the sample is determined if the number of inconsistent reading partners from the sample is not significantly greater than the number of inconsistent reading partners from the reference.

28. The storage medium of claim 27, wherein the one or more chromosomal alterations include chromosomal translocation, chromosomal deletion, chromosomal inversion, or heterologous insertion.

29. The storage medium of any one of claims 27 to 28, wherein the program instruction microprocessor determines the location of one or more candidate breakpoints.

30. The storage medium of any one of claims 27 to 28, wherein the characterization in (b) includes a fit relationship between the length of each sequence reading subsequence that produces each inconsistent reading companion and the mappability.

31. The storage medium of any one of claims 27 to 28, wherein the storage medium comprises sequential readouts, inconsistent readout pairs, inconsistent readout companion subsets, mappability variations, breakpoints, or combinations thereof.

32. A system comprising a memory and one or more microprocessors, wherein the memory includes instructions, and the one or more microprocessors are configured to perform, according to the instructions, a process for determining whether one or more chromosomal alterations are present in nucleic acids of a sample, the process including... (a) Characterizing the mappability of multiple sequence readout subsequences with respect to the sequence readouts of the nucleic acid in the sample, wherein: Each sequence readout has multiple sequence readout subsequences, and each sequence readout subsequence has a different length. Furthermore, the sequence readouts are sequence readouts of the sample's nucleic acid. (b) Identify a subset of sequence reads in which the mappability of one or more subsequences changes; (c) Compare the number of sequence readings from the subset of samples identified in (i)(b) with the number of sequence readings from the subset of references identified in (ii)(b) to generate a comparison; and (d) In the comparison of (c), the presence of one or more chromosomal alterations in the sample is determined by identifying that the number of inconsistent reading partners from the sample is significantly greater than the number of inconsistent reading partners from the reference; or, in the comparison of (c), the absence of one or more chromosomal alterations in the sample is determined if the number of inconsistent reading partners from the sample is not significantly greater than the number of inconsistent reading partners from the reference, wherein the one or more chromosomal alterations include chromosomal translocation, chromosomal deletion, chromosomal inversion, or heterologous insertion.

33. A system comprising a sequencing device and one or more computing devices, The sequencing device is configured to generate a signal corresponding to the nucleotide bases of a nucleic acid loaded into the sequencing device, wherein the nucleic acid is either a circulating cell-free nucleic acid from a test sample, or a modified variant of the circulating cell-free nucleic acid loaded into the sequencing device; and The one or more computing devices include a memory and one or more processors, the memory including instructions executable by the one or more processors, and the instructions executable by the one or more processors are configured as follows: Sequence reads are generated from the signal and aligned to a reference genome; (a) Characterizing the mappability of multiple sequence readout subsequences in terms of sequence readouts, where: Each sequence readout has multiple sequence readout subsequences, and each sequence readout subsequence has a different length. Furthermore, the sequence readouts are sequence readouts of the sample's nucleic acid. (b) Identify a subset of sequence reads in which the mappability of one or more subsequences changes. (c) Compare the number of sequence readings from the subset of samples identified in (i)(b) with the number of sequence readings from the subset of references identified in (ii)(b) to generate a comparison; and (d) In the comparison of (c), the presence of one or more chromosomal alterations in the sample is determined by identifying that the number of inconsistent reading partners from the sample is significantly greater than the number of inconsistent reading partners from the reference; or, in the comparison of (c), the absence of one or more chromosomal alterations in the sample is determined if the number of inconsistent reading partners from the sample is not significantly greater than the number of inconsistent reading partners from the reference, wherein the one or more chromosomal alterations include chromosomal translocation, chromosomal deletion, chromosomal inversion, or heterologous insertion.