Method for determining arm aneuploidy score

By combining targeted genome and nucleic acid sequencing technologies with variant calling and copy number variation analysis, the problem of high efficiency and low cost in tumor sample arm aneuploidy scoring was solved, achieving efficient and accurate scoring results while reducing computational resources and storage requirements.

CN122029293APending Publication Date: 2026-05-12LIFE TECHNOLOGIES CORP
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIFE TECHNOLOGIES CORP
Filing Date
2024-10-23
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and cost-effectively determine arm aneuploidy in tumor samples, especially given the high computational and storage requirements in whole-genome sequencing.

Method used

Using a low sample input, the target sites in the tumor sample genome were selectively amplified, sequenced using a nucleic acid sequencing instrument, and combined with variant calling, copy number variation analysis and segmentation algorithms to identify copy number changes in chromosome arms. The OCA Plus group and Oncomine™ Comprehensive Assay Plus were used for scoring.

Benefits of technology

It improves the accuracy and efficiency of aneuploidy scoring in tumor sample arms, reduces computational resources and storage requirements, achieves improvements over whole-genome sequencing methods, and enhances the consistency and reproducibility of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122029293A_ABST
    Figure CN122029293A_ABST
Patent Text Reader

Abstract

A method for determining an arm aneuploidy score in a tumor sample genome includes selectively amplifying a nucleic acid sequence at a particular location in a tumor genome using a targeting panel to generate a sequence read. Next, the genomic location is divided into segments with homogeneous copy numbers based on the logarithmic probability of the hybrid SNP and the CNV logarithmic ratio of the sequence reads. Increased and deleted segments relative to a reference copy number are identified, which segments intersect the corresponding chromosome arms. The cell abundance of these segments is compared to a minimum threshold. The longest section of the arm satisfying the minimum cell abundance is retained and the total base number thereof is summed. The total number is divided by the number of bases in the arm to obtain a score. If the score satisfies a minimum threshold, the segment is filtered based on a multiple change and an increase or absence is determined. Arms determined to be added or missing are counted to produce an arm aneuploidy score.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims the benefit of U.S. Provisional Application No. US63 / 592,732, filed October 24, 2023, pursuant to 35 USC § 119(e). The entire contents of the foregoing application are incorporated herein by reference. Technical Field

[0002] This disclosure relates to methods, systems, and computer-readable media for determining arm aneuploidy scores, and more specifically, to methods, systems, and computer-readable media for determining arm aneuploidy scores of tumor sample genomes using nucleic acid sequencing data from targeted sequencing and next-generation sequencing (NGS) technologies. Attached Figure Description

[0003] Figures 1A and 1B are block diagrams of an example procedure for analyzing the genome of a sample to determine arm aneuploidy scores.

[0004] Figure 2 is a schematic diagram of an exemplary system for reconstructing nucleic acid sequences according to various implementation schemes.

[0005] Figure 3 is an example of a block diagram for an analysis pipeline of signal data obtained from a nucleic acid sequencing instrument. Detailed Implementation

[0006] Based on the teachings and principles of this application, a novel method, system, and non-transitory machine-readable storage medium are provided for determining arm aneuploidy scores by analyzing nucleic acid sequence reads from the genome of tumor samples.

[0007] In various implementations, DNA (deoxyribonucleic acid) can be referred to as a nucleotide chain composed of four types of nucleotides: A (adenine), T (thymine), C (cytosine), and G (guanine), and RNA (ribonucleic acid) is composed of four types of nucleotides: A, U (uracil), G, and C. Certain nucleotide pairs bind specifically to each other in a complementary manner (referred to as complementary base pairing). That is, adenine (A) pairs with thymine (T) (however, in the case of RNA, adenine (A) pairs with uracil (U), and cytosine (C) pairs with guanine (G). When the first nucleic acid chain binds to a second nucleic acid chain composed of nucleotides complementary to those in the first chain, the two chains combine to form a double helix. In various implementations, “nucleic acid sequencing data,” “nucleic acid sequencing information,” “nucleic acid sequence,” “genomic sequence,” “gene sequence,” or “fragment sequence,” “nucleic acid sequence read,” or “nucleic acid sequencing read” refers to any information or data indicating the order of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine / uracil) in DNA or RNA molecules (e.g., whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, fragment, etc.). It should be understood that this instruction considers sequence information obtained using all available techniques, platforms, or technologies, including but not limited to: capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, electronic signature-based systems, etc.

[0008] “Polynucleotide,” “nucleic acid,” or “oligonucleotide” refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogues thereof) linked by nucleoside bonds. Typically, a polynucleotide comprises at least three nucleosides. Oligonucleotides are typically several monomer units in size, for example, ranging from 3-4 to several hundred monomer units. Whenever a polynucleotide (such as an oligonucleotide) is represented as a series of letters, such as “ATGCCTG,” it should be understood that, unless otherwise specified, the nucleotides are arranged in a 5'->3' order from left to right, with “A” representing deoxyadenosine, “C” representing deoxycytidine, “G” representing deoxyguanosine, and “T” representing thymidine. As is standard in the art, the letters A, C, G, and T can be used to refer to the base itself, the nucleoside, or the nucleotide comprising the base.

[0009] The phrase "next-generation sequencing" or NGS refers to sequencing technologies that offer increased throughput compared to traditional Sanger and capillary electrophoresis methods, such as the ability to generate hundreds of thousands of relatively small reads at once. Some examples of next-generation sequencing technologies include (but are not limited to) sequencing synthesis, ligation sequencing, and hybridization sequencing.

[0010] The phrase "genomic variants" refers to a single or group of sequences (in DNA or RNA) that have changed relative to a particular species or subgroup within a particular species due to mutation, recombination / crossover, or gene drift. Examples of genomic variant types include, but are not limited to, single nucleotide polymorphisms (SNPs), copy number variations (CNVs), insertions / deletions (Indels), inversions, etc.

[0011] In various implementations, genomic variants can be detected using nucleic acid sequencing systems and / or the analysis of sequencing data. A sequencing workflow may begin by cutting or digesting a test sample into hundreds, thousands, or millions of smaller fragments, which are then sequenced on a nucleic acid sequencer to provide hundreds, thousands, or millions of sequence reads, such as nucleic acid sequence reads. Each read can then be mapped to a reference or target genome, and in the case of paired fragments, reads can be paired, allowing the exploration of repetitive regions of the genome. The results of mapping and pairing can be used as input to various standalone or integrated genomic variant analysis tools, such as SNPs, CNVs, Indels, inversions, etc.

[0012] The phrase “sample genome” can refer to the whole or part of an organism’s genome.

[0013] As used herein, the term "allelic" refers to a genetic variation associated with a gene or segment of DNA, that is, one of two or more alternative forms of a DNA sequence occupying the same locus.

[0014] As used in this article, the term "locus" refers to a specific location on a chromosome or nucleic acid molecule. Alleles at a locus are located at the same locus on homologous chromosomes.

[0015] As used herein, a “target group” refers to a set of target-specific primers designed to selectively amplify target gene sequences in a sample. In some implementations, the workflow further includes nucleic acid sequencing of the amplified target sequence following selective amplification of at least one target sequence.

[0016] As used herein, "target sequence" or "target gene sequence" and its derivatives refer to any single-stranded or double-stranded nucleic acid sequence that can be amplified or synthesized according to this disclosure, including any nucleic acid sequence suspected or expected to be present in a sample. In some embodiments, the target sequence is present in double-stranded form and includes at least a portion of the specific nucleotide sequence to be amplified or synthesized or its complementary sequence prior to the addition of a target-specific primer or adapter. The target sequence may be contained in nucleic acid that can hybridize with primers used for amplification or synthesis reaction prior to polymerase extension. In some embodiments, the term refers to a nucleic acid sequence whose sequence identity, nucleotide sequence, or position is determined by one or more methods of this disclosure.

[0017] As used herein, “target-specific primer” and its derivatives refer to single-stranded or double-stranded polynucleotides, typically oligonucleotides, comprising at least one sequence that is at least 50% complementary, typically at least 75% or at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98% or at least 99% complementary or identical to at least a portion of a nucleic acid molecule comprising the target sequence. In such cases, the target-specific primer and the target sequence are described as “corresponding” to each other. In some embodiments, the target-specific primer is capable of hybridizing with at least a portion of its corresponding target sequence (or a complementary sequence of the target sequence); such hybridization may optionally be performed under standard hybridization conditions or under stringent hybridization conditions. In some embodiments, the target-specific primer cannot hybridize with the target sequence or its complementary sequence, but is capable of hybridizing with a portion of a nucleic acid strand comprising the target sequence or its complementary sequence. In some embodiments, forward and reverse target-specific primers define a target-specific primer pair that can be used to amplify the target sequence via template-dependent primer extension. Typically, each primer in a target-specific primer pair contains at least one sequence that is substantially complementary to at least a portion of a nucleic acid molecule containing the corresponding target sequence, but less than 50% complementary to at least one other target sequence in the sample. In some embodiments, amplification can be performed using multiple target-specific primer pairs in a single amplification reaction, wherein each primer pair contains a forward target-specific primer and a reverse target-specific primer, each including at least one sequence that is substantially complementary to or substantially identical to the corresponding target sequence in the sample, and each primer pair has a different corresponding target sequence.

[0018] Target groups with low sample input requirements can be used to determine arm aneuploidy scores for tumor samples. Target groups can provide a viable alternative to whole-genome sequencing, which may have higher input sample requirements. In some implementations, the target group may include Oncomine. TMThe Comprehensive Assay Plus, or OCA Plus panel (Thermo Fisher Scientific), detects 502 cancer-related genes. The OCA Plus panel contains approximately 13,473 amplicons, including 1,889 specifically designed to include heterozygous SNPs with high minor allele frequencies and evenly distributed across the genome, including chromosomal arms. Additionally, heterozygous SNPs present in the targeted medical contents of this panel are utilized. The panel's heterozygous SNPs allow for a comprehensive analysis of tumor samples regarding structural alterations and copy number (CN) variations. The OCA Plus panel encompasses 19 p-arms and 23 q-arms. The covered chromosome arms include: 1p, 1q, 2p, 2q, 3p, 3q, 4p, 4q, 5p, 5q, 6p, 6q, 7p, 7q, 8p, 8q, 9p, 9q, 10p, 10q, 11p, 11q, 12p, 12q, 13q, 14q, 15q, 16p, 16q, 17p, 17q, 18p, 18q, 19p, 19q, 20p, 20q, 21p, 21q, 22q, Xp, and Xq. The OCA Plus can use nucleic acids isolated from formaldehyde-fixed paraffin-embedded (FFPE) tumor samples (including fine-needle biopsies), with a recommended amount of 20 ng and a minimum of 10 ng. In some embodiments, the group may include a custom-designed target group or other target groups that include coverage of both the "p" and "q" arms.

[0019] Figures 1A and 1B are block diagrams of an example procedure for analyzing a sample genome to determine arm aneuploidy scores. Nucleic acid sequences at target locations in the tumor sample genome are selectively amplified using a low sample input amount of a target group from the tumor sample, and sequenced using a nucleic acid sequencing instrument, generating multiple nucleic acid sequence reads. These reads are mapped to a reference genome to generate aligned reads. In variant invocation step 102, the processor receives the aligned reads generated from the targeted sequencing of the tumor sample. For example, aligned reads can be retrieved from a file using a BAM file format. The aligned reads may correspond to multiple target locations in the tumor sample genome. Variant invocation step 102 can be configured by one or more variant invocation parameters. Variant invocation step 102 can provide observed variant populations, such as SNPs (single nucleotide polymorphisms), detected in the aligned reads. Furthermore, variant invocation step 102 can determine the log odds of variant allele frequencies for the observed SNP populations. The log-odds ratio is calculated as the natural logarithm of the ratio of the number of sequence reads with the variant allele to the number of sequence reads with the reference allele. In some embodiments, variant detection methods used with this teaching may include one or more features described in U.S. Patent Application Publication No. 2013 / 0345066, published December 26, 2013; U.S. Patent Application Publication No. 2014 / 0296080, published October 2, 2014; and U.S. Patent Application Publication No. 2014 / 0052381, published February 20, 2014, each of which is incorporated herein by reference in its entirety. Other variant detection methods may be used. In various embodiments, the variant caller may be configured to call variants for a sample genome in a manner that... .vcf, .gff or Communication is performed in the form of an .hdf data file. Detected variant information can be transmitted using any file format, as long as the detected variant information can be parsed and / or extracted for analysis. The copy number variation (CNV) step 104 provides an estimate of the copy number of the aligned sequence reads and a CNV log ratio. The CNV log ratio is calculated as the log2 ratio of the copy number estimate to the baseline copy number of each amplicon in the assay. In some embodiments, CNV detection methods used with this teaching may include one or more features described in U.S. Patent Application Publication No. 2018 / 0268103, published September 20, 2018; U.S. Patent Application Publication No. 2014 / 0256571, published September 11, 2014; and U.S. Patent Application Publication No. 2016 / 0103957, published April 14, 2016, each of which is incorporated herein by reference in its entirety. Other CNV detection methods may be used. Segmentation step 106 divides the genome sequence into segments with homogeneous copy numbers using log odds and CNV log ratios. For example, the OCA Plus group provides 1889 amplicones with heterozygous SNPs specifically designed to cover the genome, and the genome is segmented using joint segmentation of CNV log ratios and allele log odds to achieve CN variations. The segmentation algorithm is cyclic binary segmentation, which is designed to use Hotelling T2 statistics to perform joint segmentation of log2 ratios and log odds to detect variation points. (See, for example, R. Shen et al., FACETS: allele-specific copy number and clonal heterogeneity analysis tool for high-throughput DNA sequencing, Nucleic Acids Research, 2016, Vol. 44, No. 16 e131 doi: 10.1093 / nar / gkw520). Optionally, segmentation step 106 may exclude segments with fewer than a minimum number of heterozygous SNPs. For example, the minimum number of SNPs in a segment can be set to a range of 5 to 15 SNPs.

[0020] In step 108, segments exhibiting copy number increases or decreases (referred to herein as increase / deletion segments) across the genome relative to a reference copy number (ref_cn) adjusted for the baseline of the sample derived by informatics methods are identified. Methods for deriving the baseline for CNV detection used in conjunction with this teaching may include one or more features described in U.S. Patent Application Publication No. 2014 / 0256571, published September 11, 2014. In step 110, increase / deletion segments with locations intersecting chromosome arms are identified. In step 112, the span of increase / deletion segments on each arm is tested and merged where possible. Where two segments with consistent copy numbers are separated by segments with different copy numbers, the segments may be merged as follows: a) Locate two segments with the same copy number within the arm, wherein the lengths (in bases) of the two segments are A and C, and a gap of length B (in bases) is located between the segments, wherein the gap contains a small number of amplicones, such as 2 or fewer amplicones, and has a copy number different from that of the flanking segments. b) The gap is merged by connecting the gap to the two flank sections to form a merged section, wherein the length of the merged section is A+B+C.

[0021] The span parameter is the number of amplicons that cover a gap region that can be merged with adjacent long segments. This gap may represent focal events and can be ignored when determining arm-level events. This suppresses the generation of small segments. For example, the span parameter can be set to 2 amplicones. Users can set the value of the span parameter.

[0022] In step 114, the cell abundance of the augmented / deleted segments in each arm is tested to determine if it reaches a minimum level relative to the sample cell abundance. Cell abundance refers to the fraction (or %) of tumor cells in the sample, also known as tumor fraction (TF). Heterozygous SNPs are present in both normal and tumor cells, but with different allele frequencies, mediated by different tumor fractions. The observed copy number level of the sample (which may be a mixture of tumor and normal cells) is calculated using the following formula: CN = TCN TF + 2 (1 – TF) (1) Where CN is the observed copy number, TCN is the tumor copy number, TF is the tumor fraction or cell abundance, and 2 represents the diploid copy number. The tumor copy number TCN is determined using the ratio of the alternative allele to the reference allele of a heterozygous SNP as follows: Alt / Ref = [m TF + 1 (1-TF)] / [n TF + 1 (1-TF)] (2) TCN = m + n (3) Where Alt is the copy number of the substitution allele, Ref is the copy number of the reference allele, m is the major copy number, and n is the minor copy number. Due to sample heterogeneity, segments may have different cell abundances. The ratio of the cell abundance of the augmented / deleted segment to the cell abundance of the entire sample is compared to the threshold segment_cellularity_threshold. If the ratio of the cell abundance of a segment to the cell abundance of the entire sample is less than segment_cellularity_threshold, the segment will not be used in further arm aneuploidy analysis. For the merged segment obtained in step 112, if the ratio of the cell abundance of the merged segment to the cell abundance of the entire sample is less than segment_cellularity_threshold, the merged segment will not be used in further arm aneuploidy analysis. For example, the segment_cellularity_threshold value can be set to 0.2. segment_cellularity_threshold represents the minimum copy number level that can distinguish between tumor segments and normal segments. Segments whose cell abundance relative to the sample cell abundance meets the segment cell abundance threshold can be further used for arm aneuploidy analysis.

[0023] Summation step 120 sums the base counts of the longest added / deleted segments in the arm to obtain the total base count of the longest segment. In step 122, a minimum length threshold is applied to the longest segment as follows: a) Divide the total number of bases in the longest segment by the number of bases in the arm to obtain a fraction; b) Compare the score to the minimum score threshold min_fraction; c) If the score is less than the minimum score threshold, the longest added / missing segment will not be considered for ARM level variation. For example, the minimum fraction threshold min_fraction can be set to 0.9. For min_fraction = 0.9, the length of the longest augmented / missing segment must be at least 90% of the corresponding arm length in order to further analyze arm-level aneuploidy.

[0024] Step 124 of the filtering process applies a minimum fold change filter to the increased segment. The fold change of the increased segment is compared to the minimum increase threshold `min_gain_fd` for the fold difference. If the fold change of the increased segment is greater than or equal to the minimum increase threshold, the segment is considered an increase. The fold change multiplied by 2 is the measured copy number of the segment. For example, if the minimum increase threshold is 1.15, the corresponding minimum copy number threshold for an increase is 2.3. A measured copy number less than 2.3 compared to the reference copy number may indicate noise in the measurement. Step 124 reduces noise by excluding increased segments that do not have a copy number increase higher than the minimum increase threshold of the reference copy number. Therefore, if the calculated copy number is greater than or equal to the minimum threshold higher than the reference copy number, the increased segment can be confirmed to have a copy number increase.

[0025] Missing Filtering Step 126: Apply the maximum fold change filter to the missing segments. Compare the fold change of the missing segment to the maximum missing threshold `max_loss_fd` for fold difference. If the fold change of the missing segment is less than or equal to the maximum missing threshold, the segment is considered missing. For example, the maximum missing threshold for fold difference is 0.85, corresponding to a maximum copy number threshold of 1.7 for the missing segment. A measured copy number greater than 1.7 and less than the reference copy number may indicate noise in the copy number measurement of the missing segment. Missing Filtering Step 126: Remove missing segments in the range 1.7 < (copy number) <= ref_cn. If the calculated copy number is less than or equal to the maximum threshold, the missing segment can be confirmed to have copy number missing.

[0026] Add filtering step 124 and missing filter step 126 to filter out segments that fall within the range of the measured copy number given by the following formula: ref_cn max_loss_fd < copy number of segment < ref_cn min_gain_fd(4) Add filtering step 124 and missing number filtering step 126 using a threshold based on fold change. Some implementations may use a threshold based on copy number. The fold change is multiplied by the reference copy number to obtain the copy number. Some implementations may use a threshold based on the difference between the copy number and the reference copy number.

[0027] The values ​​of the minimum gain threshold (min_gain_fd) and the maximum loss threshold (max_loss_fd) for fold change can be empirically obtained by validating test results with known real data. Known real data can come from test samples using orthogonal methods, such as array-based methods like OncoScan. TM The array method (ThermoFisher Scientific) can be used to examine samples to generate realistic data. The minimum gain threshold (min_gain_fd) and maximum loss threshold (max_loss_fd) for fold differences can be optimized to make the results consistent with the real dataset.

[0028] In step 128, a p-value is calculated by performing a Student's T-test on the original copy number of the amplicones contained in the augmentation / deletion segment. A long augmentation / deletion segment may contain multiple amplicones. The original copy number of these amplicones may vary even if they are classified as a long segment and assigned an integer value for the segment copy number. Amplicons within a long segment may have different read coverages. If the variance of the original copy number of amplicones in a particular segment is high, the confidence level for the augmentation or deletion of that segment is low. The p-value for the original copy number of the amplicones reflects the variance. The p-value is based on the assumption that the mean of the original copy number is equal to the reference copy number. The variance of the original copy number of amplicones within the segment is calculated. The p-value indicates whether the hypothesis is null, in which case there is no change relative to the reference copy number, i.e., there is no augmentation or deletion in the particular segment.

[0029] In step 130, the p-value is compared with a threshold p-value to determine whether the arm containing the segment is considered an addition or a missing segment. If the p-value is less than or equal to the maximum p-value max_call_pval, then the arm containing the segment is considered an addition or a missing segment. For example, the maximum p-value could be 10. -5 The value of .

[0030] The following rules specify the conditions that must be met to determine arm extension, as described in steps 110, 122, 124, and 130: cn_value > ref_cn AND fraction(gain_segment_length) >= min_fractionAND p-value <= max_call_pval AND fold_change >= min_gain_fd (5)

[0031] The following rules specify the conditions that must be met for arm loss determination, as described in steps 110, 122, 126, and 130: cn_value <ref_cn AND fraction(loss_segment_length) >= min_fractionAND p-value <= max_call_pval AND fold_change <= max_loss_fd (6)

[0032] When the above conditions are met, the number of copies of the longest segment in the arm is the total number of copies of that arm.

[0033] In step 132, the number of arms with addition or absence determinations is counted. The total number of arms with addition or absence determinations represents the arm aneuploidy score of the sample. The arm aneuploidy score can be obtained as a percentage by dividing the total number of arms with addition or absence determinations by the number of arms tested in the sample. In step 134, the arm aneuploidy score can be reported to the user in the results display or in a file containing the results.

[0034] The table below shows the results of using OncoScan. TM Compared to array-based methods, the results of testing sample arm aneuploidy using the current method are shown. The tests include those based on OncoScan. TM The array method was used to analyze 58 samples for aneuploidy detection and 81 test runs of samples using the OCA Plus targeting method described in this paper. The percentage of positive agreement (PPA) for aneuploidy determination in the positive arm was calculated as follows: PPA = [(Predicted True Positives) / (Predicted Positives)] 100 (7) The predicted true positives were obtained using OncoScan. TM The test was obtained using an array-based approach, and the predicted positivity was obtained using the current approach with the OCA Plus target group.

[0035] The negative concordance percentage (NPA) for negative arm aneuploidy determination is calculated as follows: NPA = [(Predicted True Negative) / (Predicted Negative)] 100 (8) The predicted true negatives are obtained through OncoScan. TM The test results were obtained, and the predicted negative results were obtained using current methods targeting the OCAPlus group.

[0036] Table 1 shows the PPA and NPA for determining aneuploidy in the positive and negative arms. The results indicate that the orthogonal tests for determining aneuploidy in the positive and negative arms are highly consistent. Table 1.

[0037] Table 2 shows the reproducibility test results of the current method using the OCA Plus targeting group. The concordance rate of duplicate sample pairs was determined for different tumor fractions (TFs), where a given sample was tested twice using the current method. The results show that the concordance rate increases with increasing tumor fraction. Table 2.

[0038] The target set and methods described in this article for determining arm aneuploidy scores offer an improvement over whole-genome sequencing (WGS) technology. Sequence assembly methods must be able to efficiently assemble and / or map large numbers of reads, such as by minimizing the use of computational resources. For example, sequencing a human-sized genome can result in tens or hundreds of millions of reads that need to be assembled before further analysis. Compared to processing WGS data, computational processing of nucleic acid sequence reads from targeted sequencing reduces computational and memory requirements. For WGS, this would cover 3 Gb of tumor genome. The data generated from nucleic acid sequence reads in WGS would require computation and memory storage of both nucleic acid sequence reads and variant data. In contrast, the computational requirements and memory requirements for storing nucleic acid sequence reads and variant data for a target set covering approximately 1 Mb of tumor genome are much smaller.

[0039] In various implementation schemes, a variety of technologies, platforms, or techniques can be used to generate nucleic acid sequence data, including but not limited to: capillary electrophoresis, microarrays, connection-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, electronic signature-based systems, fluorescence-based detection systems, single-molecule methods, etc.

[0040] Various implementations of nucleic acid sequencing platforms (such as nucleic acid sequencers) may include the components shown in the block diagram of Figure 2. According to various implementations, sequencing instrument 200 may include a fluid delivery and control unit 202, a sample processing unit 204, a signal detection unit 206, and a data acquisition, analysis, and control unit 208. Descriptions of various implementations of instruments, reagents, libraries, and methods for next-generation sequencing are given in U.S. Patent Application Publications Nos. 2009 / 0127589 and 2009 / 0026082, the entire contents of which are incorporated herein by reference. Various implementations of instrument 200 may provide automated sequencing for acquiring sequence information from multiple sequences in parallel (e.g., substantially simultaneously).

[0041] In various embodiments, the fluid delivery and control unit 202 may include a reagent delivery system. The reagent delivery system may include a reagent reservoir for storing various reagents. Reagents may include RNA-based primers, forward / reverse DNA primers, oligonucleotide mixtures for ligation sequencing, nucleotide mixtures for synthesis sequencing, optional ECC oligonucleotide mixtures, buffers, washing reagents, blocking reagents, stripping reagents, etc. Furthermore, the reagent delivery system may include a pipetting system or a continuous flow system connecting the sample processing unit to the reagent reservoir.

[0042] In various embodiments, the sample processing unit 204 may include sample chambers, such as flow cells, substrates, microarrays, porous disks, etc. The sample processing unit 204 may include multiple lanes, multiple channels, multiple holes, or other means of substantially simultaneously processing multiple sample sets. Furthermore, the sample processing unit may include multiple sample chambers to enable simultaneous processing of multiple rounds. In a particular embodiment, the system may perform signal detection on one sample chamber and process another sample chamber substantially simultaneously. Additionally, the sample processing unit may include an automated system for moving or manipulating the sample chambers.

[0043] In various embodiments, the signal detection unit 206 may include an imaging or detection sensor. For example, the imaging or detection sensor may include a CCD, CMOS, an ion sensor (such as an ion-sensitive layer covering the CMOS), a current detector, etc. The signal detection unit 206 may include an excitation system to cause a probe (such as a fluorescent dye) to emit a signal. The excitation system may include an illumination source, such as an arc lamp, a laser, a light-emitting diode (LED), etc. In a particular embodiment, the signal detection unit 206 may include optics for transmitting light from the illumination source to the sample or from the sample to the imaging or detection sensor. Alternatively, the signal detection unit 206 may not include an illumination source, such as when a signal is spontaneously generated due to a sequencing reaction. For example, the signal may be generated by the interaction of released portions, such as released ions interacting with an ion-sensitive layer, or pyrophosphate reacting with an enzyme or other catalyst to produce a chemiluminescent signal. In another example, changes in current can be detected without an illumination source when nucleic acids pass through a nanopore.

[0044] In various implementations, the data acquisition, analysis, and control unit 208 can monitor a variety of system parameters. System parameters may include the temperature of various parts of the instrument 200 (such as the sample processing unit or reagent reservoir), the volume of various reagents, the status of various system sub-components (such as manipulators, stepper motors, pumps, etc.), or any combination thereof.

[0045] Those skilled in the art will understand that various implementations of the instrument 200 can be used to practice a variety of sequencing methods, including ligation-based methods, synthetic sequencing, single-molecule methods, nanopore sequencing, and other sequencing technologies.

[0046] In various embodiments, sequencing instrument 200 can determine the sequence of nucleic acids (such as polynucleotides or oligonucleotides). Nucleic acids can comprise DNA or RNA and can be single-stranded, such as ssDNA and RNA, or double-stranded, such as dsDNA or RNA / cDNA pairs. In various embodiments, nucleic acids can comprise or be derived from fragment libraries, paired libraries, ChIP fragments, etc. In a particular embodiment, sequencing instrument 200 can obtain sequence information from a single nucleic acid molecule or from a group of substantially identical nucleic acid molecules.

[0047] In various implementation schemes, the sequencing instrument 200 can output nucleic acid sequencing read data in a variety of different output data file types / formats, including but not limited to: .fasta、 .csfasta、 seq.txt qseq.txt .fastq、 .sff、 prb.txt .sms srs and / or .qv.

[0048] Figure 3 is a block diagram of the analysis pipeline for signal data obtained from a nucleic acid sequencing instrument. The sequencing instrument generates raw data files (DAT or .dat files) during the sequencing run. Signal processing can be applied to the raw data to generate incorporation signal measurement data files, such as 1.wells files, which are transferred to a server FTP location along with run log information. The signal processing step can derive the background signal corresponding to the well. The background signal can be subtracted from the measured signal of the corresponding well. The remaining signal can be fitted using an incorporation signal model to estimate incorporation at each nucleotide flow in each well. The output from the above signal processing is the signal measurement results for each well and each flow, which can be stored in a file, such as a 1.wells file.

[0049] In some implementations, the base calling step may perform phase estimation, normalization, and run a solver algorithm to identify the best partial sequence fit and perform base calling. The base sequence of the sequence read is stored in an unmapped BAM file. The base calling step may generate the total number of reads, the total number of bases, and the average read length as quality control (QC) metrics indicating the quality of base calling. Base calling can be performed by analyzing any suitable signal characteristics (e.g., signal amplitude or intensity). Signal processing and base calling used in conjunction with the teachings of this invention may include one or more features described in U.S. Patent Application Publication No. 2013 / 0090860, published April 11, 2013; U.S. Patent Application Publication No. 2014 / 0051584, published February 20, 2014; and U.S. Patent Application Publication No. 2012 / 0109598, published May 3, 2012, each of which is incorporated herein by reference in its entirety.

[0050] Once the base sequence of a read is determined, the read can be provided to the alignment step, for example, in the form of an unmapped BAM file. The alignment step compares the read to a reference genome to determine the aligned read and associated mapping quality parameters. The alignment step may generate the percentage of mappable reads as a QC metric to indicate alignment quality. The alignment results may be stored in a mapped BAM file. The method for aligning reads used in this teaching may include one or more features described in U.S. Patent Application Publication No. 2012 / 0197623, published August 2, 2012, which is incorporated herein by reference in its entirety.

[0051] The BAM file format structure is described in the "Sequence Alignment / Mapping Format Specification" dated September 12, 2014 (github.com / samtools / hts-specs). As described in this document, a "BAM file" refers to a file compatible with the BAM format. As stated in this document, an "unmapped" BAM file is a BAM file that does not contain aligned sequence read information and mapping quality parameters, and a "mapped" BAM file is a BAM file that contains aligned sequence read information and mapping quality parameters.

[0052] The variant calling step may include detecting single nucleotide polymorphisms (SNPs), insertions and deletions (InDels), multiple nucleotide polymorphisms (MNPs), and complex block substitution events. In various implementations, the variant caller can be configured to call variants for a sample genome. .vcf, .gff or Communication is performed in the form of .hdf data files. Detected variant information can be transmitted using any file format, as long as the detected variant information can be parsed and / or extracted for analysis. Variance detection methods used in conjunction with this teaching may include one or more features described in U.S. Patent Application Publication No. 2013 / 0345066, published December 26, 2013; U.S. Patent Application Publication No. 2014 / 0296080, published October 2, 2014; and U.S. Patent Application Publication No. 2014 / 0052381, published February 20, 2014, each of which is incorporated herein by reference in its entirety.

[0053] According to various exemplary embodiments, hardware and / or software elements that are appropriately configured and / or programmed can be used to perform or implement one or more features of any or more of the above teachings and / or exemplary embodiments. Determining whether to use hardware and / or software elements to implement an embodiment can be based on any number of factors, such as desired computing speed, power level, thermal tolerance, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, and other design or performance limitations.

[0054] Examples of hardware components may include processors, microprocessors, one or more input devices and / or one or more output devices (I / O) (or peripherals) coupled communicatively to the following: local interface circuitry, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), integrated circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), logic gates, registers, semiconductor devices, chipsets, microchips, chipsets, etc. Local interfaces may include, for example, one or more buses or other wired or wireless connections, controllers, buffers (buffers), drivers, repeaters, and receivers, to allow appropriate communication between hardware components. A processor is a hardware device for executing software, particularly software stored in memory. A processor can be any custom or commercially available processor, a central processing unit (CPU), an auxiliary processor among several processors associated with a computer, a semiconductor-based microprocessor (e.g., in the form of a microchip or chipset), a macroprocessor, or any device typically used to execute software instructions. A processor can also represent a distributed processing architecture. I / O devices may include input devices such as keyboards, mice, scanners, microphones, touchscreens, interfaces for various medical devices and / or laboratory instruments, barcode readers, styluses, laser readers, RF device readers, etc. Furthermore, I / O devices may also include output devices such as printers, barcode printers, displays, etc. Finally, I / O devices may further include devices that communicate as both input and output, such as modulators / demodulators (modems; used to access another device, system, or network), radio frequency (RF) or other transceivers, telephone interfaces, bridges, routers, etc.

[0055] Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, programs, software interfaces, application programming interfaces (APIs), instruction sets, computational code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. Software in memory may include one or more independent programs, which may include an ordered list of executable instructions for performing logical functions. Software in memory may include a system for identifying data streams according to the teachings of the present invention and any suitable custom or commercially available operating system (O / S) that controls the execution of other computer programs, such as systems, and provides scheduling, input-output control, file and data management, memory management, communication control, etc.

[0056] According to various exemplary embodiments, a suitablely configured and / or programmable non-transitory machine-readable medium or object capable of storing instructions or instruction sets can be used to perform or implement one or more features of any one or more of the above teachings and / or exemplary embodiments, which, if executed by a machine, can cause the machine to perform the methods and / or operations according to the exemplary embodiments. Such machines may include, for example, any suitable processing platform, computing platform, computing device, processing device, computing system, processing system, computer, processor, scientific or laboratory instrument, etc., and may be implemented using any suitable combination of hardware and / or software. Machine-readable media or objects can include, for example, any suitable type of memory cell, memory device, memory object, memory medium, storage device, storage object, storage medium and / or storage cell, such as memory, removable or non-removable media, erasable or non-erasable media, writable or rewritable media, digital or analog media, hard disk, floppy disk, read-only memory optical disc (CD-ROM), recordable optical disc (CD-R), rewritable optical disc (CD-RW), optical disc, magnetic media, magneto-optical media, removable memory cards or discs, various types of digital versatile optical discs (DVDs), magnetic tape, magnetic tape cassettes, etc., including any media suitable for computers. Memory can include any one or combination of volatile memory elements (e.g., random access memory (RAM, such as DRAM, SRAM, SDRAM, etc.)) and non-volatile memory elements (e.g., ROM, EPROM, EEPROM, flash memory, hard disk drive, magnetic tape, CDROM, etc.). Furthermore, memory can incorporate electronic, magnetic, optical and / or other types of storage media. Memory can have a distributed architecture, in which various components are located remotely to each other but are still accessed via a processor. Instructions may include any suitable type of code implemented using any appropriate high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, encrypted code, etc.

[0057] According to various exemplary implementations, distributed, clustered, remote, or cloud computing resources may be used at least in part to perform or implement one or more features of any one or more of the above teachings and / or exemplary implementations and schemes.

[0058] According to the various exemplary embodiments, the teachings above and / or one or more features of any one or more of the exemplary embodiments can be executed or implemented using a source program, an executable program (object code), a script, or any other entity containing a set of instructions to be executed. When it is a source program, the program can be transformed by a compiler, assembler, interpreter, etc., which may or may not be contained in memory to function correctly with the operating system. Instructions can be written using: (a) an object-oriented programming language with data classes and methods; or (b) a procedural programming language with routines, subroutines, and / or functions, which may include, for example, C, C++, R, Python, Pascal, Basic, Fortran, Cobol, Perl, Java, and Ada.

[0059] According to various exemplary embodiments, one or more of the above exemplary embodiments may include sending, displaying, storing, printing, or outputting information relating to any information, signals, data, and / or intermediate or final results that may be generated, accessed, or used through such exemplary embodiments to a user interface device, computer-readable storage medium, local computer system, or remote computer system. For example, such information sent, displayed, stored, printed, or output may take the form of searchable and / or filterable lists of operations and reports, pictures, tables, charts, graphs, spreadsheets, correlations, sequences, and combinations thereof. Example

[0060] Example 1 is a method for determining arm aneuploidy scoring of a tumor sample genome, comprising: selectively amplifying nucleic acid sequences at target locations in the tumor sample genome by targeting groups to generate multiple nucleic acid sequence reads; and dividing the genomic location into segments with homogeneous copy numbers using the log odds of heterozygous single nucleotide polymorphisms (SNPs) and the log ratio of copy number variations (CNVs) determined for the multiple nucleic acid sequence reads, wherein the heterozygous SNPs The regions are distributed across the genome; regions showing an increase relative to a reference copy number are identified as augmented regions, and regions showing a decrease relative to a reference copy number are identified as deleted regions, wherein the corresponding identified augmented / deleted regions have positions intersecting with the corresponding arms of the chromosome; and for each of the multiple arms, the cellular abundance of the augmented / deleted region is compared with the minimum cellular abundance relative to the sample cellular abundance of the tumor sample; the bases in the longest augmented / deleted region are summed to obtain a total number of bases, wherein the longest augmented / deleted region has at least the minimum cellular abundance; the total number of bases in the longest augmented / deleted region is divided by the number of bases in the arm to obtain a score; if the score has a value at least the minimum score threshold, the longest augmented / deleted region is retained; the retained longest augmented / deleted regions are filtered based on the fold change of augmentation / deletion; a determined augmentation or deletion is determined based on the copy number of the amplicon contained in the retained longest augmented / deleted region; and the arms with determined augmentations or deletions are counted to obtain an arm aneuploidy score for the tumor sample.

[0061] Example 2 includes the subject of Example 1 and further illustrates the steps for determining an addition or deletion, which further include calculating a p-value based on the copy number of the amplicon contained in the longest retained addition / deletion segment.

[0062] Example 3 includes the subject matter of Example 2 and further illustrates that the step of determining the addition or absence of the arm further includes applying a p-value threshold to the p-value to determine the addition or absence of the arm if the p-value is less than or equal to the p-value threshold, wherein if the p-value is greater than the p-value threshold, it is determined that the arm is not added or missing.

[0063] Example 4 includes the subject of Example 1 and further illustrates that the step of comparing cell abundance further includes calculating the ratio of the cell abundance of the increased / deleted segment to the sample cell abundance.

[0064] Example 5 includes the subject of Example 4 and further includes comparing the ratio with a segment cell abundance threshold.

[0065] Example 6 includes the subject of Example 5, and further includes not using the augmented / deleted segment in further analysis if the ratio is less than the segment cell abundance threshold.

[0066] Example 7 includes the subject of Example 1 and further includes, for a given arm, identifying a gap located between two flanking augment / deletion segments, wherein the two flanking augment / deletion segments have the same copy number and the gap has a different copy number than the two flanking augment / deletion segments.

[0067] Example 8 includes the subject of Example 7 and further includes merging the gap with the two flanking added / missing segments by connecting the gap with the two flanking added / missing segments to form a merged segment.

[0068] Example 9 includes the subject of Example 8 and further illustrates the steps of comparing the cell abundance of the increased / deleted segments with the minimum cell abundance applied to the merged segments.

[0069] Example 10 includes the subject of Example 1 and further illustrates the steps of retaining the longest augmented / missing segment based on augmentation / missing fold change filtering, which further includes: for a given augmented segment, applying a minimum fold change filter to the given augmented segment.

[0070] Example 11 includes the subject of Example 10 and further includes a comparison of the multiple change of a given increment segment with a minimum increment threshold.

[0071] Example 12 includes the subject of Example 11, and further includes confirming the increase of a given increase segment if the multiple change of a given increase segment is greater than or equal to the minimum increase threshold.

[0072] Example 13 includes the subject of Example 1 and further illustrates that the step of retaining the longest augmented / missing segment based on augmentation / missing fold change filtering further includes: for a given missing segment, applying the maximum fold change filter to the given missing segment.

[0073] Example 14 includes the subject of Example 13 and further includes a comparison of the multiple variation of a given missing segment with the maximum missing threshold.

[0074] Example 15 includes the subject of Example 14, and further includes confirming the absence of a given missing segment if the multiple change of a given missing segment is less than or equal to the maximum missing threshold.

[0075] Example 16 is a system for determining arm aneuploidy scores of a tumor sample genome, including a processor and a memory communicatively connected to the processor, the processor being configured to execute instructions that, when executed by the processor, cause the system to perform a method comprising: receiving a plurality of nucleic acid sequence reads generated by selectively amplifying nucleic acid sequences at target locations in the genome of a tumor sample using a target group; and dividing the genomic locations into segments with homogeneous copy numbers using the log odds of heterozygous single nucleotide polymorphisms (SNPs) and the log ratio of copy number variations (CNVs) determined for the plurality of nucleic acid sequence reads, wherein heterozygous SNPs... Distributed across the genome; segments showing an increase relative to a reference copy number are identified as augmented segments, and segments showing a decrease relative to a reference copy number are identified as deleted segments, wherein the corresponding identified augmented / deleted segments have positions intersecting with the corresponding arms of the chromosome; and for each of a plurality of arms, the cellular abundance of the augmented / deleted segment is compared with the minimum cellular abundance relative to the sample cellular abundance of the tumor sample; the bases in the longest augmented / deleted segment are summed to obtain a total number of bases, wherein the longest augmented / deleted segment has at least the minimum cellular abundance; the total number of bases in the longest augmented / deleted segment is divided by the number of bases in the arm to obtain a score; if the score has a value at least equal to a minimum score threshold, the longest augmented / deleted segment is retained; the retained longest augmented / deleted segments are filtered based on the fold change of augmentation / deletion; the determined augmentation or deletion is determined based on the copy number of the amplicon contained in the retained longest augmented / deleted segment; and the arms with determined augmentation or deletion are counted to obtain an arm aneuploidy score for the tumor sample.

[0076] Example 17 includes the subject of Example 16 and further illustrates the steps for determining an addition or deletion, which further include calculating a p-value based on the copy number of the amplicon contained in the longest retained addition / deletion segment.

[0077] Example 18 includes the subject matter of Example 17 and further illustrates that the steps for determining an addition or absence of an arm further include applying a p-value threshold to the p-value to determine an addition or absence of an arm if the p-value is less than or equal to the p-value threshold, wherein an addition or absence of an arm is not determined if the p-value is greater than the p-value threshold.

[0078] Example 19 includes the subject of Example 16 and further illustrates the steps of comparing cell abundance, which further include calculating the ratio of cell abundance of the increased / deleted segments to the sample cell abundance.

[0079] Example 20 includes the subject of Example 19 and further includes comparing the ratio with a segment cell abundance threshold.

[0080] Example 21 includes the subject of Example 20, and further includes that if the ratio is less than the segment cell abundance threshold, the increased / deleted segment is not used in further analysis.

[0081] Example 22 includes the subject of Example 16 and further includes, for a given arm, identifying a gap located between two flanking augment / deletion segments, wherein the two flanking augment / deletion segments have the same copy number and the gap has a different copy number than the two flanking augment / deletion segments.

[0082] Example 23 includes the subject of Example 22 and further includes merging the gap with the two flanking added / missing segments by connecting the gap with the two flanking added / missing segments to form a merged segment.

[0083] Example 24 includes the subject of Example 23 and further illustrates the application of the steps of comparing the cell abundance of the added / deleted segments with the minimum cell abundance to the merged segments.

[0084] Example 25 includes the subject of Example 16 and further illustrates the steps of retaining the longest augmented / missing segment based on augmentation / missing fold change filtering. It further includes applying a minimum fold change filter to a given augmented segment.

[0085] Example 26 includes the subject of Example 25 and further includes a comparison of the multiple change of a given increment segment with a minimum increment threshold.

[0086] Example 27 includes the subject of Example 26, and further includes confirming the increase of a given increase segment if the multiple change of a given increase segment is greater than or equal to the minimum increase threshold.

[0087] Example 28 includes the subject of Example 16 and further illustrates that the step of retaining the longest augmented / missing segment based on augmentation / missing fold change filtering further includes: for a given missing segment, applying the maximum fold change filter to the given missing segment.

[0088] Example 29 includes the subject of Example 28 and further includes a comparison of the multiple variation of a given missing segment with the maximum missing threshold.

[0089] Example 30 includes the subject of Example 29, and further includes confirming the absence of the given missing segment if the multiple change of the given missing segment is less than or equal to the maximum missing threshold.

[0090] Example 31 is a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method for determining arm aneuploidy scoring of a tumor sample genome, the method comprising: receiving at the processor a plurality of nucleic acid sequence reads, the plurality of nucleic acid sequence reads being generated by selectively amplifying nucleic acid sequences at target locations in the tumor sample genome using a target group; and dividing the genomic location into segments with homogeneous copy numbers using the log odds of heterozygous single nucleotide polymorphisms (SNPs) and the log ratio of copy number variations (CNVs) determined for the plurality of nucleic acid sequence reads, wherein heterozygous SNPs... Distributed across the genome; segments with an increase relative to a reference copy number are identified as augmented segments, and segments with a decrease relative to a reference copy number are identified as deleted segments, wherein the corresponding identified augmented / deleted segments have positions intersecting with the corresponding chromosome arms; and for each of a plurality of arms, the cellular abundance of the augmented / deleted segment is compared with the minimum cellular abundance relative to the sample cellular abundance of the tumor sample; the bases in the longest augmented / deleted segment are summed to obtain a total number of bases, wherein the longest augmented / deleted segment has at least the minimum cellular abundance; the total number of bases in the longest augmented / deleted segment is divided by the number of bases in the arm to obtain a score; if the score has a value at least equal to a minimum score threshold, the longest augmented / deleted segment is retained; the retained longest augmented / deleted segments are filtered based on the fold change of augmentation / deletion; the determined augmentation or deletion is determined based on the copy number of the amplicon contained in the retained longest augmented / deleted segment; and the arms with determined augmentation or deletion are counted to obtain an arm aneuploidy score for the tumor sample.

[0091] Example 32 includes the subject of Example 31 and further illustrates the steps for determining the addition or deletion, which further include calculating a p-value based on the copy number of the amplicon contained in the longest retained addition / deletion segment.

[0092] Example 33 includes the subject matter of Example 32 and further illustrates that the steps for determining an addition or absence of an arm further include applying a p-value threshold to the p-value to determine an addition or absence of an arm if the p-value is less than or equal to the p-value threshold, wherein an addition or absence of an arm is not determined if the p-value is greater than the p-value threshold.

[0093] Example 34 includes the subject of Example 31 and further illustrates the steps of comparing cell abundance by calculating the ratio of cell abundance of the increased / deleted segments to the sample cell abundance.

[0094] Example 35 includes the subject of Example 34 and further includes comparing the ratio with a segment cell abundance threshold.

[0095] Example 36 includes the subject of Example 35 and further includes that if the ratio is less than the segment cell abundance threshold, the added / deleted segment is not used in further analysis.

[0096] Example 37 includes the subject of Example 31 and further includes, for a given arm, identifying a gap located between two flanking augment / deletion segments, wherein the two flanking augment / deletion segments have the same copy number and the gap has a different copy number than the copy number of the two flanking augment / deletion segments.

[0097] Example 38 includes the subject of Example 37 and further includes merging the gap with the two flanking added / missing segments by connecting the gap with the two flanking added / missing segments to form a merged segment.

[0098] Example 39 includes the subject of Example 38 and further illustrates the steps of comparing the cell abundance of the increased / deleted segment with the minimum cell abundance applied to the merged segment.

[0099] Example 40 includes the subject of Example 31 and further illustrates the step of retaining the longest augmented / missing segment based on augmentation / missing fold change filtering, which further includes: for a given augmented segment, applying a minimum fold change filter to the given augmented segment.

[0100] Example 41 includes the subject of Example 40 and further includes comparing the multiple change of a given increment segment with a minimum increment threshold.

[0101] Example 42 includes the subject of Example 41, and further includes confirming the increase of a given increase segment if the multiple change of a given increase segment is greater than or equal to the minimum increase threshold.

[0102] Example 43 includes the subject of Example 31 and further illustrates that the step of retaining the longest augmented / missing segment based on augmentation / missing fold change filtering further includes: for a given missing segment, applying the maximum fold change filter to the given missing segment.

[0103] Example 44 includes the subject of Example 43 and further includes comparing the multiple change of a given missing segment with the maximum missing threshold.

[0104] Example 45 includes the subject of Example 44, and further includes confirming the absence of the given missing segment if the multiple change of the given missing segment is less than or equal to the maximum missing threshold.

Claims

1. A method for determining arm aneuploidy scores in the genome of a tumor sample, comprising: By selectively amplifying nucleic acid sequences at target locations in the genome of the tumor sample using a targeted approach, multiple nucleic acid sequence reads are generated. The genome is divided into regions with homogeneous copy numbers using the log odds of heterozygous single nucleotide polymorphisms (SNPs) and the log ratio of copy number variations (CNVs) determined for the plurality of nucleic acid sequence reads; wherein the heterozygous SNPs are distributed on the genome. Segments showing an increase relative to the reference copy number are identified as augmented segments, and segments showing a decrease relative to the reference copy number are identified as deleted segments, wherein the corresponding identified augmented / deleted segments have positions that intersect with the corresponding arms of the chromosome; For each of the multiple arms The cell abundance of the increased / deleted segments is compared with the minimum cell abundance relative to the sample cell abundance of the tumor sample; The number of bases in the longest added / deleted segment is summed to obtain the total number of bases, wherein the longest added / deleted segment has at least the minimum cellular abundance; The fraction is obtained by dividing the total number of bases in the longest added / deleted segment by the number of bases in the arm; If the score has a value that is at least the minimum score threshold, then the longest added / missing segment is retained; The longest retained augmented / deleted segment is filtered based on the augmentation / deletion fold change; The determination of an addition or deletion is based on the copy number of the amplicon contained in the longest retained addition / deletion segment; and Arms with identified additions or omissions are counted to determine the arm aneuploidy score of the tumor sample.

2. The method of claim 1, wherein the step of determining the determined augmentation or deletion further comprises calculating a p-value based on the copy number of the amplicon contained in the longest retained augmentation / deletion segment.

3. The method of claim 2, wherein the step of determining the determined addition or absence further comprises applying a p-value threshold to the p-value to determine the addition or absence of the arm if the p-value is less than or equal to the p-value threshold, wherein the addition or absence of the arm is not determined if the p-value is greater than the p-value threshold.

4. The method of claim 1, wherein the step of comparing cell abundance further comprises calculating the ratio of cell abundance of the increased / deleted segment to the sample cell abundance.

5. The method of claim 4, further comprising comparing the ratio with a segment cell abundance threshold.

6. The method of claim 5, further comprising not using the augmented / deleted segment in further analysis if the ratio is less than the segment cell abundance threshold.

7. The method of claim 1, further comprising, for a given arm, identifying a gap located between two flanking augmentation / deletion segments, wherein the two flanking augmentation / deletion segments have the same copy number and the gap has a different copy number than the two flanking augmentation / deletion segments.

8. The method of claim 7, further comprising merging the gap with the two flank added / missing segments by connecting the gap with the two flank added / missing segments to form a merged segment.

9. The method of claim 8, wherein the step of comparing the cell abundance of the added / deleted segment with the minimum cell abundance is applied to the merged segment.

10. The method of claim 1, wherein the step of filtering the longest retained augmented / deleted segment based on the augmentation / deletion fold change further comprises: For a given increment segment, apply a minimum multiple change filter to the given increment segment.

11. The method of claim 10, further comprising comparing the multiple change of the given increment segment with a minimum increment threshold.

12. The method of claim 11, further comprising confirming an increase in the given increase segment if the multiple change of the given increase segment is greater than or equal to the minimum increase threshold.

13. The method of claim 1, wherein the step of filtering the longest retained augmented / deleted segment based on the augmentation / deletion fold change further comprises: For a given missing segment, apply a maximum multiple change filter to the given missing segment.

14. The method of claim 13, further comprising comparing the multiple change of the given missing segment with a maximum missing threshold.

15. The method of claim 14, further comprising confirming the absence of the given missing segment if the multiple change of the given missing segment is less than or equal to the maximum missing threshold.

16. A system for determining arm aneuploidy scores of a tumor sample genome, comprising a processor and a memory communicatively connected to the processor, the processor being configured to execute instructions that, when executed by the processor, cause the system to perform a method, the method comprising: Receive multiple nucleic acid sequence reads, which are generated by selectively amplifying nucleic acid sequences at target locations in the genome of the tumor sample using a target group; The genome is divided into regions with homogeneous copy numbers using the log odds of heterozygous single nucleotide polymorphisms (SNPs) and the log ratio of copy number variations (CNVs) determined for the plurality of nucleic acid sequence reads; Segments showing an increase relative to the reference copy number are identified as augmented segments, and segments showing a decrease relative to the reference copy number are identified as deleted segments, wherein the corresponding identified augmented / deleted segments have positions that intersect with the corresponding arms of the chromosome; For each of the multiple arms The cell abundance of the increased / deleted segments is compared with the minimum cell abundance relative to the sample cell abundance of the tumor sample; The total number of bases in the longest added / deleted segment is summed to obtain the total number of bases, wherein the longest added / deleted segment has at least the minimum cellular abundance. The fraction is obtained by dividing the total number of bases in the longest added / deleted segment by the number of bases in the arm; If the score has a value that is at least the minimum score threshold, then the longest added / missing segment is retained; The longest retained augmented / deleted segment is filtered based on the augmentation / deletion fold change; The determination of an addition or deletion is based on the copy number of the amplicon contained in the longest retained addition / deletion segment; and Arms with identified additions or omissions are counted to determine the arm aneuploidy score of the tumor sample.

17. The system of claim 16, wherein the step of determining the determined addition or deletion further comprises calculating a p-value based on the copy number of the amplicon contained in the longest retained addition / deletion segment.

18. The system of claim 17, wherein the step of determining the determined addition or absence further comprises applying a p-value threshold to the p-value to determine the addition or absence of the arm if the p-value is less than or equal to the p-value threshold, wherein if the p-value is greater than the p-value threshold, it is determined that the arm has not been added or is missing.

19. A non-transitory computer-readable medium storing instructions, which, when executed by a processor, cause the processor to perform a method for determining arm aneuploidy scoring of a tumor sample genome, the method comprising: The processor receives multiple sequence reads, which are generated by selectively amplifying nucleic acid sequences at target locations in the genome of the tumor sample using a target group; The genome is divided into regions with homogeneous copy numbers using the log odds of heterozygous single nucleotide polymorphisms (SNPs) and the log ratio of copy number variations (CNVs) determined for the plurality of nucleic acid sequence reads; Segments showing an increase relative to the reference copy number are identified as augmented segments, and segments showing a decrease relative to the reference copy number are identified as deleted segments, wherein the corresponding identified augmented / deleted segments have positions that intersect with the corresponding arms of the chromosome; For each of the multiple arms The cell abundance of the increased / deleted segments is compared with the minimum cell abundance relative to the sample cell abundance of the tumor sample; The total number of bases in the longest added / deleted segment is summed to obtain the total number of bases, wherein the longest added / deleted segment has at least the minimum cellular abundance. The fraction is obtained by dividing the total number of bases in the longest added / deleted segment by the number of bases in the arm; If the score has a value that is at least the minimum score threshold, then the longest added / missing segment is retained; The longest retained augmented / deleted segment is filtered based on the augmentation / deletion fold change; The determination of an addition or deletion is based on the copy number of the amplicon contained in the longest retained addition / deletion segment; and Arms with identified additions or omissions are counted to determine the arm aneuploidy score of the tumor sample.

20. The non-transitory computer-readable medium of claim 19, wherein the step of determining the determined addition or deletion further comprises calculating a p-value based on the copy number of the amplicon contained in the longest retained addition / deletion segment.