Methods for detection of cancer aneuploidy and selection of subjects for Anti-cancer therapy through plasma whole genome sequencing
Whole genome sequencing, combining long-read and short-read sequencing, addresses the challenge of early cancer detection by phasing SNP alleles to determine B-allele frequency scores, enabling sensitive cancer detection in low tumor fractions through allelic imbalance analysis in cfDNA.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MEMORIAL SLOAN KETTERING CANCER CENT
- Filing Date
- 2025-11-06
- Publication Date
- 2026-05-15
AI Technical Summary
Detecting cancer at an early stage using cell-free DNA (cfDNA) is challenging due to the low amount present in bodily fluids, making it difficult to accurately identify cancer presence.
Utilizing whole genome sequencing, including long-read and short-read sequencing, to phase genomic sequence reads and determine relative B-allele frequency (BAF) scores for selecting subjects for cancer therapy by identifying allelic imbalances in cfDNA, allowing for the phasing of SNP alleles and assessing major and minor alleles to detect cancer through plasma WGS.
Enables sensitive detection of low tumor fractions as low as 5*10^-5, effectively identifying cancer through plasma WGS, even in the absence of matched tumor tissue, by leveraging allelic imbalance in CNV regions across the genome.
Smart Images

Figure US2025054384_15052026_PF_FP_ABST
Abstract
Description
Atty. Dkt. No.: 115872-3343METHODS FOR DETECTION OF CANCER ANEUPLOIDY AND SELECTION OF SUBJECTS FOR ANTI-CANCER THERAPY THROUGH PLASMA WHOLE GENOME SEQUENCINGCROSS REFERENCE TO RELATED APPLICATIONS[00011 The present application claims the benefit of priority to U.S. Provisional Patent Application 63 / 717,681, titled “Methods for Detection of Cancer Aneuploidy and Selection of Subjects for Anti-Cancer Therapy Through Plasma Whole Genome Sequencing,” filed November 7, 2024, which is incorporated by reference in its entirety.BACKGROUND(0002] In a subject with cancer, the tumor cells may release deoxyribonucleic acid (DNA) fragments such as circulating tumor deoxyribonucleic acid (ctDNA) into bodily fluids. ctDNA may be type of cell-free DNA (cfDNA) that is associated with the cancer. It may be, however, difficult to detect the presence of cancer in a subject using ctDNA especially at an early stage, due to the low amount in the bodily fluids of the subject.SUMMARY
[0003] Aspects of the present disclosure are directed to systems and methods for selecting a subject for a cancer therapy based on detection of aneuploidy in cell-free deoxyribonucleic acid (cfDNA) in a bodily fluid sample of the subject. One or more processors coupled with memory may obtain, for a subject at risk of or diagnosed with cancer, a first sequencing dataset comprising a first plurality of genomic sequence reads generated by plasma whole genome sequencing (e.g., short-read sequencing) of cfDNA in a bodily fluid sample from the subject. The whole genome sequencing of genomic DNA (gDNA) may include long-read sequencing, alone or in combination with short-read sequencing. The one or more processors may generate, by phasing the first plurality of genomic sequence reads of the first sequencing dataset, (i) a first portion of the first plurality of genomic sequence reads comprising a first set of genomic sequences of single nucleotide polymorphism (SNP) alleles in the bodily fluid sample and (ii) a second portion of the second plurality of genomic sequence reads comprising a second set of genomic sequences of SNP alleles in the bodily fluid sample. The one or more processors may-1-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 identify, from the first portion of the first plurality of genomic sequence reads and the second portion of the first plurality of genomic sequence reads, a plurality of phase groups. The one or more processors may determine, for each phase group, (i) a first phase score based on a frequency of major alleles for each of the plurality of SNP alleles and (ii) a second phase score based on a frequency of minor alleles of the plurality of SNP alleles in the bodily fluid sample. The one or more processors may generate a relative B-allele frequency (BAF) score as a function of the first phase score and the second phase score for each phase group of the plurality of phase groups. The one or more processors may select the subject for a cancer therapy responsive to the relative BAF score satisfying a threshold. The one or more processors may store, using one or more data structures, an association between the subject and the cancer therapy.
[0004] In some embodiments, the one or more processors may identify, for the subject, (i) a second relative BAF score and (ii) a second cancer therapy administered to the subject. In some embodiments, the one or more processors may determine whether the relative BAF score is different from relative BAF score by the second threshold. In some embodiments, the one or more processors may select the subject for the cancer therapy based on whether the second relative BAF score is different from the relative BAF score by the second threshold. In some embodiments, the one or more processors may exclude a second subject from the cancer therapy responsive to a second relative BAF score not satisfying the threshold. In some embodiments, the one or more processors may store, using one or more data structures, an association between the second subject and exclusion from the cancer therapy.
[0005] In some embodiments, the one or more processors may obtain the sequencing dataset comprising the first plurality of genomic sequence reads of the bodily fluid sample from the subject in accordance with whole genome sequencing (WGS) at a sequence depth. The sequence depth ranging between 5-250x. In some embodiments, the one or more processors may divide each of the first set of genomic sequence reads and the second set of genome sequences for each chromosome into the plurality of windows nonoverlapping with one another. Each window of the plurality of windows may range between 1,000 to 50,000,000 base pairs. Along the length of the window, the phasing of SNPs from long read sequencing allows the SNPs to be grouped together, and the allelic-2-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 imbalance can be assessed through groups of phased SNPs rather than through individual SNPs.
[0006] In some embodiments, the one or more processors may determine a relative positive BAF value corresponding to a first B AF of the respective portion of the first plurality of genomic sequence reads relative to a second BAF corresponding portion of a second plurality of genomic sequence reads that fall within the respective chromosome. This would be considered an empirically observed major allele. In some embodiments, the one or more processors may generate the imbalance measure as at least one of (i) a one- direction T test or a (ii) a bidirectional T-test of the first difference in BAF in the window and the plurality of second differences for the remaining windows in the plurality of windows. In some embodiments, the one or more processors may identify the cancer therapy from a plurality of cancer therapies based on the relative BAF score, wherein the plurality of cancer therapies comprises at least one of neoadjuvant therapy, radiotherapy, immunotherapy, chemotherapy, targeted therapy, or surgery.
[0007] In some embodiments, the one or more processors may identify at least one major allele and at least one minor allele based on an adjusted BAF score generated using a ratio of the first phase score and the second phase score. In some embodiments, the one or more processors may identify, from the plurality of SNP alleles, the major allele and the minor allele each of the plurality of SNP alleles based on a combined fragment score. In some embodiments, the one or more processors may obtain a second sequencing dataset comprising a second plurality of genomic sequence reads generated by whole genome sequencing of gDNA in a normal tissue from the subject. The one or more processors may generate, by phasing the second plurality of genomic sequence reads of the sequencing dataset, (i) a first portion of the second plurality of genomic sequence reads comprising a first set of genomic sequences of single nucleotide polymorphism (SNP) alleles in the normal tissue and (ii) a second portion of the second plurality of genomic sequence reads comprising a second set of genomic sequences of SNP alleles in the normal tissue. The one or more processors may compare (i) the first portion of the first plurality of genomic sequence reads with the first portion of the second plurality of genomic sequence reads and (ii) the second portion of the first plurality of genomic sequence reads with the second-3-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 portion of the second plurality of genomic sequence reads to identify the plurality of phase groups.
[0008] In some embodiments, the one or more processors may apply a filter on at least one window of the plurality of windows, to remove at least one of artifact biases or mapping errors. In some embodiments, the first set of genomic sequences may correspond to a maternal set of genomic sequences of SNP alleles in the bodily fluid sample and the second set of genomic sequences correspond to a paternal set of genomic sequences of SNP alleles in the bodily fluid sample. In some embodiments, the phasing may include at least one of population phasing or familial phasing. In some embodiments, the bodily fluid sample may include at least one of plasma, urine, cerebral spinal fluid, or lung fluid. In some embodiments, the normal tissue may be obtained from the subject. In some embodiments, the cancer may include at least one of lung cancer, brain cancer, head and neck cancer, colon cancer, rectal cancer, uterine cancer, endometrial cancer, stomach cancer, ovarian cancer, cervical cancer, bladder cancer, pancreatic cancer, esophageal cancer, prostate cancer, renal cancer, skin cancer, or breast cancer.
[0009] Aspects of the present disclosure are directed to systems and methods for selecting subjects for cancer therapies based on fragment variation. One or more processors coupled with memory may obtain, for a subject at risk of or diagnosed with cancer, a first sequencing dataset comprising a first plurality of fragments including a first plurality of genomic sequences generated by whole genome sequencing of cfDNA in a bodily fluid sample from the subject. The one or more processors may generate, by phasing the first plurality of fragments comprising the first plurality of genomic sequences of the first sequencing dataset corresponding to a plurality of single nucleotide polymorphisms (SNPs),(i) a first portion of the first plurality of fragments comprising a first set of fragments corresponding to a first allele at each of the plurality of SNPs in the bodily fluid sample and(ii) a second portion of the first plurality of fragments comprising a second set of fragments corresponding to a second allele at each of the plurality of SNPs in the bodily fluid sample. The one or more processors may identify the first portion of the first plurality of fragments and the second portion of the first plurality of fragments, a plurality of phase groups. The one or more processors may determine, for each phase group of the plurality of phase groups, a plurality of lengths for the plurality of fragments corresponding to the plurality of-4-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 SNPs. Each of the plurality of lengths may identify a difference of lengths between the first set of fragments corresponding to a major allele at a respective SNP of the plurality of SNPs and the second set of fragments corresponding to a minor allele at the respective SNP of the plurality of SNPs. The one or more processors may classify, for each phase group of the plurality of phase groups, the plurality of lengths into one of a plurality of regions, each of the plurality of regions defined by a respective minimum length and a respective maximum length. The one or more processors may determine, for each respective region of the plurality of regions in each phase group of the plurality of phase groups, a regional fragment score based on the difference of lengths for each of the plurality of lengths classified into respective region. The one or more processors may generate a combined fragment score as a function of the regional fragment score for each respective region of the plurality of regions in each phase group of the plurality of phase groups. The one or more processors may select the subject for a cancer therapy responsive to the combined fragment score satisfying a threshold. The one or more processors may store, using one or more data structures, an association between the subject and the cancer therapy.[00101 In some embodiments, the one or more processors may determine, for each phase group, (i) a first phase score based on a frequency of major alleles for each of the plurality of SNP alleles and (ii) a second phase score based on a frequency of minor alleles of the plurality of SNP alleles in the bodily fluid sample. In some embodiments, the one or more processors may generate a relative BAF score as a function of the first phase score and the second phase score for each phase group of the plurality of phase groups. In some embodiments, the one or more processors may determine an aggregate score based on the combined fragment score and the relative BAF score. The one or more processors may select the subject responsive to the aggregate score satisfying the threshold. In some embodiments, the one or more processors may identify, from the plurality of SNP alleles, the major allele and the minor allele each of the plurality of SNPs based on at least one of the combined fragment score or the relative BAF score.
[0011] In some embodiments, the one or more processors may generate a second combined fragment score for a second subject as the function of a second regional fragment score for each respective region of a second plurality of regions. The one or more processors may select the second subject for the cancer therapy responsive to the second-5-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 combined fragment score not satisfying the threshold. The one or more processors may store, using the one or more data structures, an association between the second subject and exclusion from the cancer therapy.
[0012] In some embodiments, the one or more processors may classify the plurality of lengths into one of the plurality of regions, the plurality of regions comprising (i) a short fragment enrichment peak region, (ii) a depleted median fragment region, and (iii) a short dinucleosome fragment peak region. In some embodiments, the one or more processors may generate the combined fragment score as the function comprising a difference between (i) a summation of a first regional fragment score corresponding to the short fragment enrichment peak region and a second regional fragment score corresponding to the short dinucleosome fragment peak region and (ii) a third regional fragment score corresponding to the depleted median fragment region.
[0013] In some embodiments, the short fragment enrichment peak region may have the respective minimum length corresponding to 120 base pairs and the respective maximum length corresponding to 160 base pairs. In some embodiments, the depleted median fragment region may have the respective minimum length corresponding to 165 base pairs and the respective maximum length corresponding to 225 base pairs. In some embodiments, the short dinucleosome fragment peak region may have the respective minimum length corresponding to 245 base pairs and the respective maximum length corresponding to 315 base pairs.
[0014] In some embodiments, the one or more processors may obtain a second sequencing dataset comprising a second plurality of fragments including a second plurality of genomic sequences generated by whole genome sequencing of gDNA in a normal tissue from the subject. The one or more processors may generate, by phasing the second plurality of fragments comprising the second plurality of genomic sequences of the second sequencing dataset corresponding to a plurality of single nucleotide polymorphisms (SNPs), (i) a first portion of the plurality of fragments comprising a first set of fragments corresponding to a major allele at each of the plurality of SNPs in the normal tissue and (ii) a second portion of the plurality of fragments comprising a second set of fragments corresponding to a minor allele at each of the plurality of SNPs in the normal tissue. The one or more processors may compare (i) the first portion of the first plurality of fragments-6-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 with the first portion of the second plurality of fragments and (ii) the second portion of the first plurality of fragments with the second portion of the second plurality of fragments to identify the plurality of phase groups.
[0015] In some embodiments, the one or more processors may identify the cancer therapy from a plurality of cancer therapies based on the combined fragment score, wherein the plurality of cancer therapies comprises at least one of neoadjuvant therapy, radiotherapy, immunotherapy, chemotherapy, targeted therapy, or surgery. In some embodiments, the phasing may include at least one of population phasing or familial phasing. In some embodiments, the bodily fluid sample may include at least one of plasma, urine, cerebral spinal fluid, or lung fluid. In some embodiments, the cancer may include at least one of lung cancer, brain cancer, head and neck cancer, colon cancer, rectal cancer, uterine cancer, endometrial cancer, stomach cancer, ovarian cancer, cervical cancer, bladder cancer, pancreatic cancer, esophageal cancer, prostate cancer, renal cancer, skin cancer, or breast cancer.BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The foregoing and other objects, aspects, features, and advantages of the disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which:
[0017] FIG. 1: Identifying allelic imbalance above statistical noise (the noise around BAFs at 0.50) is a function of coverage, plasma tumor fraction, and observed SNPs affected by cancer aneuploidy. Increasing coverage from 25X to 200X decreased the number of required SNPs by a factor of 8, making detection at TF=5*10'5feasible through plasma WGS alone.
[0018] FIG. 2A: Schematic of a haplotype-naive and haplotype-aware hidden Markov models (HMMs) to detect allelic imbalance. State transitions are dependent on copy number and phases.
[0019] FIG. 2B: Schematic of a model to integrate haplotype information obtained from population-based phasing with allele and expression signals to detect copy number variations from single-cell ribonucleic acid (scRNA) sequences.-7-4903-2435-7751.1Atty. Dkt. No.: 115872-3343
[0020] FIGs. 3A and 3B: Graphs of allele frequencies and depth ratios across chromosomes 1-22. For non-cancer samples, the allele frequency is expected to be about 0.5, correlated with a roughly even split between maternal and paternal genetic material sources. Due to the presence of cancer, certain chromosomes are shown with an uneven allele frequency, such as chromosomes 3, 8, 13, 17, 18, 20, and 21, among others. This is due to loss of heterozygosity (LOH) in the affected chromosomes.
[0021] FIG. 4: Graph shows B-allele frequency (BAF) variance differs per sample.
[0022] FIG. 5: Graph shows that population phasing is subject to phase switches.
[0023] FIG. 6: In silico mixing studies of admixed pretreatment and postoperative plasma samples from the HR+ / HER2- breast cancer patient for phased BAF, in which the major allele is empirically selected by mean BAF. Pretreatment plasma (TF = 6.8%) was mixed into matched peripheral blood mononuclear cells in 25 replicates. Admixtures model TFs of 1 O’6— 1 O’2(x-axis). An AUC heatmap demonstrates detection performance at the different admixed TFs vs. TF=0, measured by Z score
[0024] FIG. 7: Coefficient of variation and BAF in pretreatment plasma from a patient with node-positive breast cancer. Phased SNPs are grouped into 1 MB bins. The major allele is empirically selected from the higher mean BAF in each 1 MB group.Coefficient of variation (y-axis, left) is calculated as the coefficient of variation of fragment insert sizes in the major allele minus the coefficient of variation of fragment insert sizes in the minor allele. BAF - 0.5 (right, y-axis) is defined as the major allele fragments / (major + minor allele fragments). Chromosomes are arranged in order on the X-axis. Coefficient of variation overtly tracks BAF in high-burden disease.
[0025] FIG. 8: Jointplot of coefficient of variation (adj coeff, y-axis) and BAF (adj baf, x-axis) in a patient with high burden, node positive breast cancer.
[0026] FIG. 9: Fragment size distributions of major and minor allele at single SNP in an amplification region of a high burden (TF 6.8%) HR+ / HER2- breast cancer sample.
[0027] FIGs. 10A and 10B: Differences in major and minor allele fragment insert sizes in pretreatment and postoperative plasma. Major allele was selected by mean BAF-8-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 value in each phase group. A) Representative example of patterned enrichment in an SNP in a region of aneuploidy in plasma. The major allele in high burden, pretreatment plasma is enriched for short (<150 base pairs) and long (>245 base pairs) fragments, while depleted in median (150-235 base pairs) fragments. B) No patterned enrichment is seen in empirically observed major allele in postoperative plasma.
[0028] FIG. 11: Illustration of fragment score. Fragment score is the positive sum of the short fragment enrichment peak (120-160 base pair insert size) minus the depleted median region (165 - 225 base pairs) + the sum of fragments in the short dinucleosome fragment peak (245-315 base pairs).
[0029] FIG. 12: In silico mixing studies of admixed pretreatment and postoperative plasma samples from the HR+ / HER2- breast cancer patient NY-4787 for phased fragment score, in which the major allele is empirically selected by mean BAF. Pretreatment plasma (TF = 6.8%) was mixed into matched peripheral blood mononuclear cells in 25 replicates. Admixtures model TFs of 1 O’6— 1 O’2(x-axis). An AUC heatmap demonstrates detection performance at the different admixed TFs vs. TF=0, measured by Z score
[0030] FIG. 13: In silico mixing studies of admixed pretreatment and postoperative plasma samples from the HR+ / HER2- breast cancer patient NY-4787 for a combined BAF and fragment score aggregated in phase blocks. For fragment score, the major allele is empirically selected by mean BAF. For BAF, the major allele is empirically selected by fragment score. BAF and fragment scores were combined by standard scaling values from both metrics and adding them together. Pretreatment plasma (TF = 6.8%) was mixed into matched peripheral blood mononuclear cells in 10 replicates. Admixtures model TFs of 10’6-10’2(x-axis). An AUC heatmap demonstrates detection performance at the different admixed TFs vs. TF=0, measured by Z score.
[0031] FIG. 14: Tracking aggregate changes in sample-wide BAF score (y-axis) through neoadjuvant treatment and surgery in patient.
[0032] FIG. 15: ROC analysis performed on fragment score classifier in early-stage breast cancer. Preoperative plasma samples (positive label, n=7) are compared postoperative plasma samples in patients without recurrent disease (negative label, n=30).-9-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 The major allele used in the fragment score is determined by the empirically observed major allele BAF, an independent signal.
[0033] FIG. 16: Barplot of ctDNA detection (Sample fragment score, y-axis) in pre and posttreatment BC samples (n=37) using fragment score classifier. Pretreatment plasma is detected above posttreatment (neoadjuvant therapy and surgery) samples from patients without disease recurrence in long term follow-up. Major allele is nominated by BAF empiric major allele, an independent measure, x-axis is annotated with patient (NY-), plasma timepoint (P0), clinical stage (e.g., Tib), and receptor subtype (TNBC vs. HR+ / HER2-)
[0034] FIG. 17: Barplot of ctDNA detection with sample BAF score (y-axis) in pre and posttreatment BC samples (n=37) using BAF classifier. Pretreatment plasma (n=7) is shaded in grey detected above posttreatment (neoadjuvant therapy and surgery) samples from patients without disease recurrence in long term follow-up. Major allele is nominated by BAF empiric major allele, an independent measure, x-axis is annotated with patient (NY- _), plasma timepoint (P0_), clinical stage (e.g., Tib), and receptor subtype (TNBC vs. HR+ / HER2-)
[0035] FIG. 18: ROC analysis performed on BAF score classifier in early-stage breast cancer. Preoperative plasma samples (positive label, n=7) are compared postoperative plasma samples in patients without recurrent disease (negative label, n=30). The major allele used in the BAF score is determined by the empirically observed major allele fragment score, an independent signal.
[0036] FIG. 19: Bar plot of ctDNA detect in which an adjusted BAF score is used to identify the major allele. Adjusted BAF can be used to empirically select major allele, then evaluate with a completely different metric (e.g., summation of fragment score at a site). Since the signals are unrelated, in regions without aneuploidy empiric BAF calls should amount to random noise in the metric.
[0037] FIG. 20: a) Loss of heterozygosity (LOH) can be measured via changes in BAF of SNPs in cfDNA. b) In silico mixing studies of admixed high and low tumor fraction samples from the colorectal cancer patient CRC-930. Pretreatment plasma (TF = 12%) was mixed into matched peripheral blood mononuclear cells in 25 replicates.-10-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 Admixtures model TFs of 1 O’6— 1 O’3(x-axis). An AUC heatmap demonstrates detection performance at the different admixed TFs vs. TF=0, measured by Z score.
[0038] FIG. 21: Use of tumor-informed BAF classifier to detect MRD in observational TNBC recurrence cohort. Early-stage TNBC patients underwent surgical resection plus neoadjuvant and / or adjuvant chemotherapy. Plasma was sampled intermittently throughout clinical course. Lead-time calculations for ctDNA detection posttherapy (green dot) versus clinical recurrence (red dot). Where available, purple dot shows ctDNA detection after surgery or initiation of chemotherapy.
[0039] FIG. 22: Plasma BAF inference in a neutral (left) and CNV (right) region in BC plasma. Each dot is an individual SNP site. In CNV regions, unphased A and B alleles (top) lead to a mean BAF of 0.5. Phasing A and B alleles (bottom) enables accurate, unidirectional signal inference.
[0040] FIG. 23: Barplot of ctDNA detection (Sample BAF score, y-axis) in pre and posttreatment BC samples (n=22) using prototype plasma BAF classifier. Pretreatment plasma is detected above posttreatment (neoadjuvant therapy and surgery) samples from patients without disease recurrence in long term follow-up. X-axis is annotated with patient (NY-), plasma timepoint (P0), clinical stage (e.g., Tib), and receptor subtype (TNBC vs. HR+ / HER2-).
[0041] FIG. 24: A graph of B-allele frequency (BAF) scores across timepoints in a subject. Serial monitoring of response to neoadjuvant immunotherapy and surgery in the patient NA-47, a 61 -year-old female with stage IIIA non-small cell lung cancer (squamous histology), using population phasing to phase A and B alleles. Plasma BAF score declines from pretreatment timepoint (A) to preoperative timepoint at week 6 (D) and remains decreased at 6 months post-surgery (F). This patient did not have recurrence at 39-month follow-up.
[0042] FIG. 25: A graph of fragment (FRAG) scores across timepoints in a subject. Serial monitoring of response to neoadjuvant immunotherapy and surgery in the patient NA-45, a 77-year-old female with stage IIIA non-small cell lung cancer (adenocarcinoma), using population phasing to phase A and B alleles. Plasma FRAG score declines from-11-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 pretreatment timepoint following treatment with immunotherapy and further declines at 6 months post-surgery. This patient did not have recurrence at 44-month follow-up.
[0043] FIG. 26 depicts a block diagram of a system for selecting subjects for cancer therapies based on detection of allelic imbalances in cell-free deoxyribonucleic acid (cfDNA) in bodily fluid samples, in accordance with an illustrative embodiment.
[0044] FIGs. 27A and 27B depict block diagrams of a process for parsing genomic sequence reads of cell-free deoxyribonucleic acid (cfDNA) in bodily fluid samples to select subjects for cancer therapies based on detection of allelic imbalances, in accordance with an illustrative embodiment.
[0045] FIG. 28 depicts a block diagram of a flow of a process for determining allelic imbalances in reads of single nucleotide polymorphism (SNP) alleles in the bodily fluid samples, in accordance with an illustrative embodiment.
[0046] FIG. 29 depicts a flow diagram of a method of selecting subjects for cancer therapies based on detection of allelic imbalances in cell-free deoxyribonucleic acid (cfDNA) in bodily fluid samples, in accordance with an illustrative embodiment.
[0047] FIG. 30 depicts a block diagram of a system for selecting subjects for cancer therapies based on B-allele frequency (BAF) and fragment variation in accordance with an illustrative embodiment.
[0048] FIGs. 31A-31C depict block diagrams of a process for parsing genomic sequence reads of cell-free deoxyribonucleic acid (cfDNA) in bodily fluid samples to select subjects for cancer therapies based on B-allele frequency (BAF) and fragment variation, in accordance with an illustrative embodiment.
[0049] FIG. 32 depicts a flow diagram of a method of selecting subjects for cancer therapies based on B-allele frequency (BAF) and fragment variation in accordance with an illustrative embodiment.
[0050] FIG. 33 depicts a block diagram of a server system and a client computer system, in accordance with one or more implementations.-12-4903-2435-7751.1Atty. Dkt. No.: 115872-3343DETAILED DESCRIPTION
[0051] It is to be appreciated that certain aspects, modes, embodiments, variations and features of the present methods are described below in various levels of detail in order to provide a substantial understanding of the present technology. It is to be understood that the present disclosure is not limited to particular uses, methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.
[0052] Following below are more detailed descriptions of various concepts related to, and embodiments of, systems and methods for selecting a subject for a cancer therapy based on detection of allelic imbalances in cell-free deoxyribonucleic acid (cfDNA) in a bodily fluid sample of the subject. It should be appreciated that various concepts introduced above and discussed in greater detail below may be implemented in any of numerous ways, as the disclosed concepts are not limited to any particular manner of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes.
[0053] Section A describes the use of allelic imbalance for early detection of cancer in high depth whole-genome sequencing (WGS).
[0054] Section B describes the use of B-allele frequency (BAF) and fragment variation in high depth plasma whole-genome sequencing (WGS).
[0055] Section C describes systems and methods of selecting subjects for cancer therapies based on detection of allelic imbalances in cell-free deoxyribonucleic acid (cfDNA) in bodily fluid samples.
[0056] Section D describes systems and methods of selecting subjects for cancer therapies based on B-allele frequency (BAF) and fragment variation.
[0057] Section E describes a network environment and computing environment which may be useful for practicing various computing related embodiments described herein.Definitions
[0058] At a given SNP, the most common allele is referred to as the “major allele”; whereas the least common (or rarer) allele is referred to as the “minor allele”. The minor -13-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 allele frequency is therefore the frequency at which the minor allele of an SNP occurs within a population.
[0059] A “phase group” refers to a collection of SNP alleles grouped together from long-read sequencing of gDNA or cell-free DNA (cfDNA) or population databases (population phasing.(0060] “Phasing” corresponds to the process of assigning or associating portions of genomic sequence reads to alleles from a biological sample from a subject.
[0061] A “minor allele fraction” or “B-allele frequency” or “BAF” refer to the proportion of a specific population where the second most common allele (the “minor allele”) at a particular genetic locus is present.
[0062] A “B-allele frequency (BAF) score” or “phase score” is based on a frequency of alleles in each phase group (or fixed window size). The frequency may be over a total number of single nucleotide polymorphisms (SNPs) in a sample.
[0063] A “major allele fraction” or “A-allele frequency” refer to the proportion of a specific population where the most common allele (the “major allele”) at a particular genetic locus is present.
[0064] “Whole genome sequencing (WGS)” is a process of determining a deoxyribonucleic acid (DNA) sequence from a biological sample in a single run.
[0065] “Long read sequencing” is a DNA sequencing technique that uses long DNA fragments to produce a high-resolution view of the entire genome. Long-read sequencing can detect genetic variants that are difficult to detect with short-read sequencing methods.
[0066] Short read Sequencing” is currently the most commonly used form of nextgeneration sequencing (NGS) and has a wide range of diagnostic applications. In these types of sequencing, the genome is broken into small fragments (usually 50 to 300 bases) before being sequenced.-14-4903-2435-7751.1Atty. Dkt. No.: 115872-3343
[0067] “Fragmentomics” is the study of the structure and functions of cell-free DNA (cfDNA) fragments and their patterns. Fragmentomics involves looking at markers such as the size of the cfDNA fragment, and the pattern of DNA bases at the ends of the fragments
[0068] A “regional fragment score” refers to a metric associated with a particular cfDNA fragment based on the length of the fragment that satisfies a predetermined range.
[0069] A “combined fragment score” refers to a combination of fragment scores across different regions based on length.A. Allelic Imbalance for Early Detect of Cancer in High Depth Whole-Genome Sequencing (WGS)
[0070] Can liquid biopsy enable the early detection of cancer in high-risk patients through a novel whole-genome sequencing (WGS) paradigm that leverages signal from millions of single nucleotide polymorphisms (SNPs) altered in cancer aneuploidy? High- risk patients with a genetic predisposition to breast cancer may be considered: screening through biannual mammograms and MRIs decreases breast cancer mortality1, yet patient adherence is suboptimal2and imaging modalities suffer from poor sensitivity (mammograms) and specificity (MRIs) in young patients3. Liquid biopsy for circulating tumor DNA (ctDNA) offers the potential to noninvasively detect breast cancer below the limit of radiographic detection, yet when tumor burden is low, ctDNA is sparsely diluted among cell-free DNA (cfDNA) from noncancerous cells. The prevailing paradigm in the field involves deep targeted sequencing of a few informative sites (patient-specific or cancer driver hotspots) to overcome this sparsity of ctDNA signal. It is shown that this approach faces a fundamental sensitivity barrier when ctDNA fraction falls below 1 : 10'3, as is commonly seen in early stage cancer and other low volume cancer settings4.
[0071] To overcome the limitation of low cfDNA input, it is reasoned that in high aneuploidy tumor types like breast and ovarian cancer, WGS of plasma can enable ctDNA signal inference from large copy number variant (CNV) regions across the genome including amplification, deletion, and loss of heterozygosity (LOH) events. The innovative framework MRD-EDGE used plasma WGS to identify ctDNA in tumor fractions as low as 1 : 1 O’6in colorectal cancer and non-small cell lung cancer using matched tumor tissue from-15-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 surgical resection specimens, and successfully detected postoperative minimal residual disease (MRD) linked to disease recurrence in CRC and NSCLC.
[0072] Plasma WGS enables signal inference from CNV events across the breadth of the genome. In MRD-EDGE, a novel B-allele frequency (BAF) classifier is introduced as a radical departure from conventional deep targeted sequencing panels. Rather than focus on the sparse cancer mutations common to leading targeted panels6which present significant limitations in low tumor mutational burden solid tumors like breast cancer, the millions of heterozygous SNPs in CNV regions can be leveraged to detect allelic imbalance indicative of ctDNA shedding. In each of these CNV regions, aneuploidy alters the balance between diploid chromosomes, creating a major (overrepresented) and minor (underrepresented) allele that causes BAF, the ratio of alleles at heterozygous SNPs, to deviate from an expected diploid value of 0.50. Briefly, the tumor-informed BAF classifier(i) identifies the major allele and copy number state of heterozygous SNPs in tumor tissue,(ii) applies quality filters to remove artifact bias and mapping error in plasma WGS reads, and (iii) uses a linear regressor accounting for tumor copy number state and major allele BAF to estimate a sample-wide ctDNA tumor fraction (TF) in plasma. In silico mixing studies of this novel classifier demonstrated ultrasensitive ctDNA detection5at TFs as low as 5*10'5.
[0073] To demonstrate the clinical potential of the tumor-informed BAF framework, serial postoperative plasma samples were evaluated from an observational cohort of patients (n=18) with triple negative breast cancer (TNBC) with disease recurrence after definitive therapy (surgery combined with neoadjuvant (n=9) or adjuvant (n=9) chemotherapy). The BAF classifier demonstrated strong sensitivity for MRD in this cohort, as ctDNA was detected following treatment initiation (neoadjuvant chemotherapy or surgery) and prior to recurrence in 17 of 18 patients (94.4%). Results compare favorably to a 2023 analysis of TNBC ctDNA detection in the adjuvant setting using ddPCR in which sensitivity for relapse in serial monitoring was estimated at 58%9.
[0074] A core premise of plasma WGS is that increased depth of sequencing will enable greater ctDNA signal inference4. With costs of WGS declining dramatically11, greater depth of sequencing can enable a transformative liquid biopsy approach: the de novo detection of somatic aneuploidy in cfDNA without matched tumor tissue. Such an-16-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 advance would expand the applicability of plasma WGS to cancer screening. However, in the screening setting, there is also an absence of matched tumor tissue, and the phase of heterozygous alleles is unknown. Therefore, a second core premise to BAF-based cancer detection is the phasing of SNP alleles through long-read sequencing, which allows one allele to be matched against the other when evaluating for allelic imbalance through loss of heterozygosity or copy number amplification. Finally, novel analytical signal inference methods are needed to harvest signal in this setting, where copy number state is unknown.
[0075] The process may be as followed. First, long read sequencing of matched normal tissue (peripheral blood mononuclear cells or other normal tissue) may be used for familial phasing of heterozygous SNPs. In some embodiments, short read sequencing of the biological sample (e.g., normal tissue with gDNA or plasma sample with cfDNA) may be used for population phasing of heterozygous SNPs, instead of long read sequencing, to identify phase groups. In some embodiments, long and short sequencing of matched normal tissue may be used for familial phasing of heterozygous SNPs to identify phase groups. This is performed once to phase alleles and does not need to be repeated. Second, WGS sequencing of plasma may be used to identify allelic imbalance (discriminating an excess of one allele against another). This is performed in regular intervals in the screening setting. The sequencing may be of various depths, such as low-depth sequencing (e.g., less than 5- lOx), medium-depth sequence (e.g., between lOx and 30x), or high-depth sequence (e.g., more than 3 Ox), among others. Third, a novel analytical signal inference algorithms may be used to capture allelic imbalance attributable to ctDNA signal.
[0076] The following phasing may be used in the techniques detailed herein to determine BAF and fragmentomics metrics. In some embodiments, phase groups may be identified through population phasing of heterozygous SNPs identified through short read sequencing of plasma cfDNA. In some embodiments, phase groups may be identified through population phasing of heterozygous SNPs identified through short read sequencing of matched normal gDNA. In some embodiments, phase groups may be identified through familial phasing of heterozygous SNPs identified through long read sequencing of matched normal gDNA. In some embodiments, phase groups may be identified through familial phasing of heterozygous SNPs identified through short read and long read sequencing (both sequencing types) of matched normal gDNA.-17-4903-2435-7751.1Atty. Dkt. No.: 115872-3343
[0077] From a benchmarking perspective, the opportunity to identify allelic imbalance can be modeled as a function of WGS coverage, plasma ctDNA TF, and the number of heterozygous SNPs needed to observe a difference from a balanced allele distribution at different plasma tumor fractions (FIG. 1). By leveraging low-cost WGS to increase coverage from ~25X to 200X, the number of altered SNPs may be reduced to observe allelic imbalance by a factor of 8, enabling sensitive detection of tumor fractions as low as 5*10'5, in line with the detection capability of the tumor-informed BAF classifier. The average genome has 4.3 million heterozygous SNPs12, making detection in breast cancer, mean aneuploidy 2 gigabases13, feasible.
[0078] As a proof of concept, a pilot cohort of serial plasma timepoints from 2 patients with germline BRCA mutations and screen-detected early-stage breast cancers was evaluated. Plasma samples were sequenced to a median depth of 150-200X with low-cost Ultima Genomics11WGS. Long read sequencing from Oxford Nanopore was applied on matched normal samples. Short read WGS through Ultima Genomics was also applied to matched normal samples with a median depth of 40X. To capture ctDNA signal in this cohort, it was reasoned that LOH events are large (chromosomal or chromosomal- arm level) and localized to specific windows within the genome. In some embodiments, regional allelic imbalance can be recursively compared along each region of the genome with every other region of the genome, allowing a gaining of a comprehensive statistical summary of allelic imbalance within each individual plasma sample (labeled ‘t sum”, FIG. 18). In the patient NY-3001, with stage IIB triple-negative breast cancer, this approach sensitively tracked tumor burden from pretreatment through the neoadjuvant therapy period and through surgery (FIG. 18).
[0079] Referring now to FIG. 2A, depicted is a schematic of a haplotype-naive and haplotype-aware hidden Markov models (HMMs) to detect allelic imbalances. State transitions are dependent on copy number and phases. Referring now to FIG. 2B, depicted is a schematic of a model to integrate haplotype information obtained from populationbased phasing with allele and expression signals to detect copy number variations from single-cell ribonucleic acid (scRNA) sequences.
[0080] Referring now to FIGs. 3A and 3B, depicted are graphs of allele frequencies and depth ratios across chromosomes 1-22 in tumor tissue. For non-cancer samples, the-18-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 allele frequency is expected to be about 0.5, correlated with a roughly even split between maternal and paternal genetic material sources. Due to the presence of cancer, certain chromosomes are shown with an uneven allele frequency, such as chromosomes 3, 8, 13, 17, 18, 20, and 21, among others. This is due to loss of heterozygosity (LOH) in the affected chromosomes.
[0081] Referring now to FIG. 4, shown is a graph B-allele frequency (BAF) variance differs per sample. As shown, the graph shows that due to the BAF variance per sample, approaches as described in FIGs. 2A and 2B are poor at distinguishing allelic imbalances. Referring now to FIG. 5, depicted is a graph showing that population phasing is subject to phase switches. Population phasing may include using genetic data from a large group of individuals to infer the haplotypes (combinations of alleles at different loci on the same chromosome) present in the population. Phase switches refer to errors in the inferred haplotypes where the phasing incorrectly assigns alleles from different chromosomes. Phase switches in population phasing arise from limitations in data resolution, genetic recombination, algorithmic ambiguities, genotyping errors, population structure, and reference panel limitations, among a whole host of other factors. Long-read sequencing (with familial phasing) and population phasing can be used to detect allelic imbalance using the techniques as detailed herein.References1. Evans DG, Harkness EF, Howell A, et al. Intensive breast screening in BRCA2 mutation carriers is associated with reduced breast cancer specific and all cause mortality. Hered Cancer Clin Pract. 2016; 14:8.2. Ferreira CS, Rodrigues J, Moreira S, Ribeiro F, Longatto-Filho A. Breast cancer screening adherence rates and barriers of implementation in ethnic, cultural and religious minorities: A systematic review. Mol Clin Oncol. 2021 ; 15(1): 139.3. Mann RM, Kuhl CK, Moy L. Contrast-enhanced MRI for breast cancer screening. J Magn Reson Imaging. 2019;50(2):377-390.4. Zviran A, Schulman RC, Shah M, et al. Genome-wide cell-free DNA mutational integration enables ultra-sensitive cancer monitoring. Nat Med. 2020;26(7): 1114-1124.-19-4903-2435-7751.1Atty. Dkt. No.: 115872-33435. Widman AJ, Shah M, Ogaard N, et al. Machine learning guided signal enrichment for ultrasensitive plasma tumor burden monitoring. bioRxiv. Published online January 20, 2022:2022.01.17.476508. doi: 10.1101 / 2022.01.17.4765086. Powles T, Assaf ZJ, Davarpanah N, et al. ctDNA guiding adjuvant immunotherapy in urothelial carcinoma. Nature. Published online June 16, 2021. doi: 10.1038 / s41586-021- 03642-97. Phallen J, Sausen M, Adleff V, et al. Direct detection of early-stage cancers using circulating tumor DNA. Sci Transl Med. 2017;9(403). doi : 10.1126 / scitranslmed.aan24158. Tie J, Cohen JD, Lahouel K, et al. Circulating Tumor DNA Analysis Guiding Adjuvant Therapy in Stage II Colon Cancer. N Engl J Med. 2022;386(24):2261-2272.9. Turner NC, Swift C, Jenkins B, et al. Results of the c-TRAK TN trial: a clinical trial utilising ctDNA mutation tracking to detect molecular residual disease and trigger intervention in patients with moderate- and high-risk early-stage triple-negative breast cancer. Ann Oncol. 2023; 34(2): 200-211.10. Shaw J, Page K, Ambasger B, et al. Serial postoperative ctDNA monitoring of breast cancer recurrence. J Clin Orthod. 2022;40(16_suppl):562-562.11. Almogy G, Pratt M, Oberstrass F, et al. Cost-efficient whole genome-sequencing using novel mostly natural sequencing-by-synthesis chemistry and open fluidics platform. bioRxiv. Published online August 10, 2022:2022.05.29.493900. doi: 10.1101 / 2022.05.29.49390012. 1000 Genomes Project Consortium, Auton A, Brooks LD, et al. A global reference for human genetic variation. Nature. 2015;526(7571):68-74.13. Cancer Genome Atlas Network. Comprehensive molecular portraits of human breast tumours. Nature. 2012;490(7418):61-70.14. Adalsteinsson VA, Ha G, Freeman SS, et al. Scalable whole-exome sequencing of cell- free DNA reveals high concordance with metastatic tumors. Nat Commun. 2017;8(l): 1324.-20-4903-2435-7751.1Atty. Dkt. No.: 115872-3343B. B-Allele Frequency (BAF) and Fragment Variation in High Depth Plasma Whole-Genome Sequencing (WGS).Approach to measuring BAF in the tumor-naive (de novo) context
[0082] While high depth sequencing expands ctDNA signal inference for a single SNP site, a second novel strategy is needed to group SNPs together. In regions of cancer aneuploidy including amplifications, deletions, and copy-neutral loss-of-heterozygosity, the composition of aneuploidy favors a major, or overexpressed allele, and a minor allele. In the tumor-informed BAF approach, SNP phasing was inferred by identifying the major allele in tumor tissue, and plasma TF subsequently calculated by comparing observed allelic imbalance compared to the known aneuploidy profile in tumor. Because the major allele is unknown in the plasma-only setting, an alternative strategy was employed to phase parental (A and B) haplotypes. A second innovative strategy was therefore introduced, not previously considered in other liquid biopsy platforms: use of long-read sequencing13of normal tissue to directly phase A and B haplotypes, enabling us to confidently group SNPs by region.
[0083] While population phasing14could putatively be used for this purpose, pitfalls such as switch errors, in which phases in an individual do not match population patterns, and a lack of diversity in population groups used to build population phasing databases, prevent confident signal inference in the low TF settings needed for early BC detections. Critically, long-read phasing allows us to aggregate large swaths of neighboring SNPs to determine allelic imbalance at the regional level, as the majority of BC aneuploidy events are large (arm or chromosome level)3, providing strong statistical power for aneuploidy detection. For example, long-read phasing enables the visualization of A and B alleles in the plasma of a patient with lymph node-positive HR+ / HER2- BC. Long-read phasing is a one-time test of normal tissue that is collected in the same tube as plasma, and in subsequent screening intervals only plasma will be studied. Although population phasing is less sensitive than long sequence phasing, population phasing can be used to determine BAF and fragment variation as detailed herein. For instance, the reads generated from short read sequencing of cfDNA (e.g., from fragments with 50-300 base pairs) of a given subject may be phased via population phasing to determine an arrangement of alleles on chromosomes for the subject. Even with the lower sensitivity, the phased read can be used to detect allelic-21-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 imbalances and determine fragmentomics related metrics as detailed herein to identify whether treatment is appropriate for the subject.
[0084] Long read phasing of alleles was used from normal tissue (e.g., buffy coat) to place maternal alleles into phased blocks of various lengths. In plasma, the major allele was empirically inferred based on the mean BAF value of the phase block. Aneuploidy may be present in all regions of the genome and this can be inferred through the major allele. A series of filters was applied to reduce noise, including removing poorly mappable regions, WASP plasma and normal correction15, and coverage-based filters to remove regions of excess or sparse coverage. Poor quality SNPs may be removed according to the SNP calling algorithm (e.g., DeepVariant or GATK HaplotypeCaller), BAFs that exceed a certain value within a binomial distribution may be removed based on individual SNP coverage and an expected ‘coin flip’ BAF value of 0.5. The length of phased blocks were further filtered on, removing phased blocks that are too short to produce reliable signal (e.g., <25 contiguous phased SNPs).
[0085] Once the BAF value is observed, a statistical measure of the BAF can be applied from 0.5, such as a T test of all BAF values and the inverse in a phase block against 0.5, or a mean of the deviation according to a binomial distance from 0.5. After calculating BAF signal in a phase block in this manner, the number of SNPs in each phase block were normalized. This normalized signal was then summed across the genome to produce an aggregate BAF score. This sample-wide BAF score has strong sensitivity in in silico mixing studies (FIG. 6)Approach to measuring fragment score in the tumor-naive (de novo) context
[0086] SNPs, when phased into parental alleles through long read sequencing, an opportunity to measure the contribution of fragmentomics to cancer signal were also produced. Circulating tumor DNA fragments were previously demonstrated in regions of cancer aneuploidy have lengths that are more heterogenous than non-cancer cell free DNA (cfDNA) fragments4. However, nothing has compared allele A vs. allele B fragment size distributions in this setting.
[0087] The relationship between fragment heterogeneity in cfDNA and major allele BAF was sought to be explored. The difference in coefficient of variation (standard -22-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 deviation / mean) was found, a statistical measure of variation that unlike measures like Shannon’s entropy is largely independent of the number of fragments in a distribution, between the fragments belonging to the major and minor allele (as determined empirically by BAF), largely replicated B AF signal in a patient with lymph node positive breast cancer (FIG. 7).
[0088] At the individual SNP level, this coefficient of variation signal is not correlated to BAF signal, making the two orthogonal measures of aneuploidy signal (FIG. 8). Following this observation, further characterization of this signal was sought. In the plasma of several high burden breast cancer patients the difference in insert sizes between major and minor alleles was measured. Even in regions of aneuploidy in high burden cancer, the fragment distributions of major and minor alleles at each SNP site have considerable overlap.
[0089] Because the regions overlap, the normalized (number of fragments at each insert size / by total number of fragments in the distribution) difference in fragment insert size between major and minor alleles at SNPs within 1 megabase was plotted. In CNV regions, it was found that major alleles, which have a larger concentration of ctDNA, had more short and long fragments and fewer fragments in the median cfDNA (167-171 base pairs) distribution (FIG. 9). This relationship was not observed in postoperative plasma samples with minimal cancer burden (FIGs. 10A and 10B)
[0090] The metric fragment score was therefore designed to evaluate for this pattern of enrichment in individual SNPs. Fragment score is the positive sum of the short fragment enrichment peak (120-160 base pair insert size) minus the depleted median region (165— 225 base pairs) + the sum of fragments in the short dinucleosome fragment peak (245-315 base pairs).
[0091] Critically, fragment score holds little correlation to BAF at the individual SNP level, making the two measures orthogonal metrics for aneuploidy signal. The sample level fragment score is similar to BAF score — the same SNP filters and the same sample summary metric (normalize by the number of SNPs in each phase block and then sum this normalized signal across the genome to produce an aggregate fragment score) can be used.-23-4903-2435-7751.1Atty. Dkt. No.: 115872-3343
[0092] The sample level fragment score, in which major allele nominated the BAF, was evaluated to demonstrate ultrasensitive detection in in silico mixing studies (FIG. 12). The BAF and fragment score can be combined at the phase group level in in silico mixing studies. Clinically, the BAF score and the fragment score were used to evaluate for ctDNA burden in pre and postoperative plasma samples. Here, serial plasma samples were collected from patients with germline predisposition to breast cancer. Samples were collected prior to treatment, following neoadjuvant therapy, and following surgical resection. None of the patients recurred in long term follow up. These postoperative samples therefore serve as non cancer controls with the same SNP profiles as the pretreatment plasma samples.
[0093] The BAF classifier accurately tracked ctDNA tumor burden between pretreatment (P01), following neoadjuvant chemotherapy (P02) and after surgery (P03) (FIG. 14). Clinically, the sample level fragment score, in which major allele was nominated by BAF, was applied to 37 patients (7 pretreatment and 30 posttreatment or postoperative). This is an especially challenging dataset as 5 of the patients had stage 1 breast cancer, including stage 1 hormone receptor positive, HER2 negative breast cancer, a challenging ctDNA target even for state-of-the-art targeted panels7.
[0094] The fragment score classifier, in which major allele was nominated by BAF, had outstanding performance of pretreatment samples (positive label) over posttreatment (negative label) in an area under the receiver operating curve (ROC) analysis. (FIG. 15) Results are further shown in barplot. Remarkably, stage I HR+ / HER2- breast cancers without matched tumor tissue are detected above all postoperative controls. The three pretreatment cancer samples that are not detected beyond controls are very small (< 1cm), stage I HR+ / HER2- breast cancers, including 2 pTla breast cancers. No existing ctDNA platforms can reliably detect these cancers. (FIG. 16).
[0095] The independence of the BAF and fragment score signals is what enables observation of ctDNA signal in low TF samples. The empirically observed fragment score can similarly be used to nominate the major allele for the BAF classifier, with similarly strong performance in this extremely challenging dataset.
[0096] References-24-4903-2435-7751.1Atty. Dkt. No.: 115872-33431. Magbanua, M. J. M. et al. Circulating tumor DNA in neoadjuvant-treated breast cancer reflects response and survival. Ann. Oncol. 32, 229-239 (2021).2. Taylor, A. M. et al. Genomic and Functional Approaches to Understanding Cancer Aneuploidy. Cancer Cell 33, 676-689. e3 (2018).3. Cancer Genome Atlas Network. Comprehensive molecular portraits of human breast tumours. Nature 490, 61-70 (2012).4. Widman, A. J. et al. Ultrasensitive plasma-based monitoring of tumor burden using machine-learning-guided signal enrichment. Nat. Med. (2024) doi: 10.1038 / s41591- 024-03040-4.5. Widman, A. J. et al. Machine learning guided signal enrichment for ultrasensitive plasma tumor burden monitoring. bioRxiv 2022.01.17.476508 (2022) doi:10.1101 / 2022.01.17.476508.6. Turner, N. C. et al. Results of the c-TRAK TN trial: a clinical trial utilising ctDNA mutation tracking to detect molecular residual disease and trigger intervention in patients with moderate- and high-risk early-stage triple-negative breast cancer. Ann. Oncol. 34, 200-211 (2023).7. Shaw, J. et al. Serial postoperative ctDNA monitoring of breast cancer recurrence. J. Clin. Orthod. 40, 562-562 (2022).8. Adalsteinsson, V. A. et al. Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatic tumors. Nat. Commun. 8, 1324 (2017).9. Doebley, A.-L. et al. Griffin: Framework for clinical cancer subtyping from nucleosome profiling of cell-free DNA. bioRxiv (2021) doi: 10.1101 / 2021.08.31.21262867.10. Benjamini, Y. & Speed, T. P. Summarizing and correcting the GC content bias in high-throughput sequencing. Nucleic Acids Res. 40, e72 (2012).-25-4903-2435-7751.1Atty. Dkt. No.: 115872-334311. Almogy, G. et al. Cost-efficient whole genome-sequencing using novel mostly natural sequencing-by-synthesis chemistry and open fluidics platform. bioRxiv 2022.05.29.493900 (2022) doi: 10.1101 / 2022.05.29.493900.12. 1000 Genomes Project Consortium et al. A global reference for human genetic variation. Nature 526, 68-74 (2015).13. Maestri, S. et al. A long-read sequencing approach for direct haplotype phasing in clinical settings. Int. J. Mol. Set. 21, 9177 (2020).14. Loh, P.-R. et al. Reference-based phasing using the Haplotype Reference Consortium panel. Nat. Genet. 48, 1443-1448 (2016).15. van de Geijn, B., McVicker, G., Gilad, Y. & Pritchard, J. K. WASP: allele-specific software for robust molecular quantitative trait locus discovery. Nat. Methods 12, 1061-1063 (2015).Use Case - Liquid biopsy for circulating tumor DNA (ctDNA) to Detect Breast Cancer
[0097] Screening through biannual mammograms and MRIs decreases breast cancer mortality in patients with breast cancer1(BC), yet patient adherence is suboptimal2and imaging modalities suffer from poor sensitivity (mammograms) and specificity (MRIs) in young patients3. Furthermore, access to high quality screening programs can be challenging for underserved populations4,5. Liquid biopsy for circulating tumor DNA (ctDNA) offers the potential to noninvasively detect BC and other high aneuploidy cancer types below the limit of radiographic detection. However, when tumor burden is low, ctDNA is sparsely diluted among cell-free DNA (cfDNA) from noncancerous cells.
[0098] To overcome the limitation of low cfDNA input, the whole genome sequencing (WGS) of plasma was reasoned and can enable ctDNA signal inference from large copy number variant (CNV) regions across the genome including amplification, deletion, and loss of heterozygosity (LOH) events. A novel B-allele frequency (B AF) classifier was previously6introduced as a radical departure from conventional liquid biopsy approaches. Rather than focus on the sparse cancer mutations common to leading targeted panels7 9, which present significant limitations in BC8, instead the millions of heterozygous-26-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 single nucleotide polymorphisms (SNPs) were leveraged in CNV regions to detect allelic imbalance indicative of ctDNA shedding. In each of these CNV regions, aneuploidy alters the balance between diploid chromosomes, creating a major (overrepresented) and minor (underrepresented) allele that causes B AF, the ratio of alleles at heterozygous SNPs, to deviate from an expected diploid value of 0.50. The tumor-informed BAF classifier demonstrated strong sensitivity for minimal residual disease (MRD) following surgery in triple negative BC patients with disease recurrence after definitive treatment (surgery and chemotherapy).
[0099] In plasma WGS, increased depth of sequencing enables greater ctDNA signal inference10. With costs of WGS declining dramatically11, it is hypothesized that greater sequencing depth can enable a transformative liquid biopsy approach: the detection of cancer-associated aneuploidy in cfDNA without matched tumor tissue. Such an advance would expand the applicability of plasma WGS to BC screening, offering a noninvasive complement to radiographic imaging as part of a novel strategy for early diagnosis, and would be an ideal early detection mechanism for other high aneuploidy tumor types. The benchmarking studies demonstrate that by leveraging low-cost WGS to increase plasma coverage from ~25X to 200X and phasing haplotypes with long-read sequencing, the plasma BAF approach would have sensitivity on par with the tumor-informed BAF classifier. Proof-of-concept data demonstrates de novo detection of pretreatment BC above controls in small hormone-receptor positive (HR+), HER2-negative BCs.
[0100] Aim 1 : Develop and implement a plasma BAF classifier to identify ctDNA shedding in young patients with BC. Pretreatment plasma sequenced will be evaluated to a target depth of 200x from 50 young (age < 50) BC patients with germline predisposition to cancer, a cohort size that will enable extensive methods optimization during the early phase of the CAMS award period. It is expected that >80% of pretreatment early-stage BC plasma samples will demonstrate ctDNA shedding, which will have critical implications for noninvasive cancer screening through liquid biopsy. The clinicopathologic and genomic factors will be further characterized as associated with high ctDNA shedding in early-stage BC.
[0101] Aim 2: Expand BAF -based early detection to other high aneuploidy tumor types. The proof-of-concept data demonstrate that even small (<2 cm) breast cancers can be -27-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 detected using the plasma BAF approach. The approach for other high aneuploidy tumor types that are currently underserved by existing early detection approaches will be developed: ovarian, esophageal and pancreatic cancers, where early detection of small cancers is likely to have a transformative clinical impact. Following validation in these tumor types, machine learning will be used to classify plasma aneuploidy profiles according to underlying tumor type, leading to a novel BAF -based multi-cancer early detection (MCED) platform.
[0102] Impact: Liquid biopsy can enable the early detection of BC and other high aneuploidy cancers through a novel BAF classifier that leverages signal from millions of SNPs altered in CNV regions. As aneuploidy is a marker specific to advanced dysplasia or invasive cancer, this approach may avoid the false positives in healthy persons that have confounded other MCED (e.g., methylation-based) results. Finally, the simple WGS workflow obviates the need for custom panel generation and molecular barcodes and could drive widespread adoption of this approach. The CAMS award will ensure availability of the resources and support needed to pursue this high-risk, clinically transformative project during the transition to independence.[01031 Liquid biopsy can augment conventional BC screening strategies. BC is the second leading cause of cancer mortality for women in the USA with an estimated 310,720 new cases and 44,250 deaths in 2024, and rates of invasive BC are rising among young women (age <50)12. Despite significant advances in treatment, early detection remains a priority as prognosis closely aligns with stage at diagnosis. Current screening techniques such as mammography and MRI are effective but have limitations in sensitivity and specificity13,14and suffer from suboptimal adherence.
[0104] Liquid biopsy to identify ctDNA in plasma presents a noninvasive approach for cancer diagnosis that has the potential to improve screening participation in young patients and underserved communities. Nevertheless, even state-of-the-art, tumor-informed bespoke panels fail to detect pretreatment ctDNA in large, hormone-receptor (HR) positive, HER-2 negative breast cancers15.
[0105] BAF harnesses the breadth of plasma WGS for ultrasensitive CNV detection. Aneuploidy is a prominent hallmark of the cancer genome16and is particularly enriched in-28-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 BC, where 95% of tumors have a CNV footprint of >100 megabases17. Plasma WGS enables signal inference from CNV events across the breadth of the genome. A tumor- informed BAF classifier was previously6introduced to detect ctDNA by evaluating for allelic imbalance among SNPs in amplifications, deletions, and copy-neutral LOH regions nominated from matched tumor tissue from surgical resection or biopsy specimens (FIG. 20, panel a). Briefly, the tumor-informed BAF classifier i) identifies the major allele and copy number state of heterozygous SNPs in tumor tissue, ii) applies quality filters to remove artifact bias and mapping error in plasma WGS reads, and iii) uses a linear regressor accounting for tumor copy number state and major allele BAF to estimate a sample-wide ctDNA tumor fraction (TF) in plasma. In silico mixing studies of this novel classifier demonstrated ultrasensitive ctDNA detection18at TFs as low as 5xl0'5(FIG. 20, panel b) at 25X WGS coverage.
[0106] To demonstrate the clinical potential of the tumor-informed BAF framework, serial postoperative plasma samples from an observational cohort of patients (n=18) with triple negative breast cancer (TNBC) with disease recurrence after definitive therapy (surgery combined with neoadjuvant (n=9) or adjuvant (n=9) chemotherapy) were evaluated. The BAF classifier demonstrated strong sensitivity for MRD in this cohort, as ctDNA was detected following treatment initiation (neoadjuvant chemotherapy or surgery) and prior to recurrence in 17 of 18 patients (FIG. 21, 94.4%). Results compare favorably to a 2023 analysis of TNBC ctDNA detection in the adjuvant setting using ddPCR in which sensitivity for relapse in serial monitoring was estimated at 58%19. Following completion of definitive treatment (surgery and ACT), average lead time of ctDNA detection was 9 months and maximum lead-time was 27 months, competitive with leading bespoke panels20in TNBC despite sparse sampling at varying time points and relatively modest 25X WGS coverage. Detection of postoperative MRD with the tumor-informed BAF classifier was further linked to short recurrence-free survival in colorectal cancer and non-small lung cancer.
[0107] High-depth WGS enables ultrasensitive ctDNA detection without matched tumor tissue. Importantly, when evaluating for changes in BAF, greater sequencing depth enables more opportunities to observe allelic imbalance at each individual SNP site. While other plasma-only aneuploidy approaches such as measuring read depths21,22are prone to-29-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 regional GC bias23and mapping bias that scale with increased coverage (analog signal), BAF compares alleles A and B at the same SNP site and each incremental read improves signal inference (digital signal). The recent advent of low-cost WGS from Ultima Genomics presents an opportunity to sequence to previously unobtainable sequencing depths to radically change sensitivity for BAF signal11, enabling development of non-tumor informed plasma BAF approaches.
[0108] Aim 1 : Develop and implement a plasma-only BAF classifier to identify ctDNA shedding in young patients with BC. From a benchmarking perspective, the opportunity can be modeled to identify allelic imbalance as a function of WGS coverage, plasma ctDNA TF, and the number of heterozygous SNPs needed to observe a difference from a balanced allele distribution at different plasma tumor fractions (FIG. 20). By leveraging low-cost WGS to increase coverage from ~25X to 200X, the number of altered SNPs required to observe allelic imbalance would be reduced by a factor of 8 to roughly 1 million, enabling sensitive detection of tumor fractions as low as 5* 10"5in line with the detection capability of the tumor-informed BAF classifier at 25X WGS. The average genome has 4.3 million heterozygous SNPs24, making detection in BC, where 95% of tumors have significant aneuploidy (> 100MB)17, feasible.
[0109] While high depth sequencing expands ctDNA signal inference for a single SNP site, a second novel strategy is needed to group SNPs together. In regions of cancer aneuploidy including amplifications, deletions, and copy-neutral loss-of-heterozygosity, the composition of aneuploidy favors a major, or overexpressed allele, and a minor allele. In the tumor-informed BAF approach, SNP phasing is inferred by identifying the major allele in tumor tissue, and plasma TF is subsequently calculated by comparing observed allelic imbalance compared to the known aneuploidy profile in tumor. Because the major allele is unknown in the plasma-only setting, an alternative strategy must be employed to phase parental (A and B) haplotypes.[OHO] A second innovative strategy not previously considered in other liquid biopsy platforms has therefore been introduced: use of long-read sequencing25of normal tissue to directly phase A and B haplotypes, enabling us to confidently group SNPs by region. While population phasing26could putatively be used for this purpose, pitfalls such as switch errors, in which phases in an individual do not match population patterns, and a -30-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 lack of diversity in population groups used to build population phasing databases, prevent confident signal inference in the low TF settings needed for early BC detections. Critically, long-read phasing allows aggregation of large swaths of neighboring SNPs to determine allelic imbalance at the regional level, as the majority of BC aneuploidy events are large (arm or chromosome level)17, providing strong statistical power for aneuploidy detection. For example, long-read phasing enables the visualization of A and B alleles in the plasma of a patient with lymph node-positive HR+ / HER2- BC (FIG. 22). Long-read phasing is a onetime test of normal tissue that is collected in the same tube as plasma, and in subsequent screening intervals only plasma will be studied.
[0111] To begin the work on this ambitious problem, a pilot cohort of serial plasma timepoints from 10 patients with germline BRCA mutations and screen-detected, small, early-stage breast cancers was evaluated. Plasma samples were sequenced to a median depth of 150-200X with low-cost Ultima Genomics11WGS. Patients were treated with a mix of neoadjuvant therapy and surgery. In total, 22 plasma samples were sequenced, including 4 pretreatment plasma samples from 4 patients and 18 posttreatment plasma samples. None of the patients have documented recurrence in long term (>7 year) follow up.
[0112] Changes in the mean BAF value of heterozygous SNPs that correlated to clinical course were observed. For example, in patient NY-4787, a 29-year-old woman with BRCA1 -mutant stage Illa disease, BAF declined from pretreatment timepoint following neoadjuvant therapy and further declined after surgical resection (FIG. 14). The patient has no evidence of recurrent disease in clinical follow-up. Overall, across 22 samples, 4 pretreatment BC samples above 18 posttreatment plasma timepoints (FIG. 23) were detected, including a stage I (pTlbNO, <1 cm) ER+ / HER2- BC, an extremely challenging ctDNA detection target even for state of the art tumor-informed bespoke panels15. Notably, 3 of the 4 patients had undetectable ctDNA with a standard plasma-only WGS tool, iChorCNA21. The pilot cohort demonstrates the outstanding potential of the plasma BAF classifier to track disease burden, even in patients with very small tumors that are not detected with conventional liquid biopsy approaches, and there is dramatic room for improvement from methods optimization as outlined in this proposal.-31-4903-2435-7751.1Atty. Dkt. No.: 115872-3343
[0113] Methods: High depth WGS will be used to identify de novo ctDNA in a critical cancer screening setting: young (age < 50) patients with mammographically-evident cancer and germline predisposition to BC (e.g., BRCA mutation). In year 1 appropriate patients will be identified biobank, which contains pretreatment plasma samples from over 1,000 patients. cfDNA will be extracted and plasma WGS will be performed to a target depth of 200x from 50 young patients with mammographically detected early-stage BC, including at least 15 stage I HR+ / HER2- BCs. For the control group, plasma and matched normal tissue will be sourced from 30 young patients with high-risk germline mutations. In both groups, peripheral blood mononuclear cells will serve as matched normal samples and will undergo long-read sequencing through Oxford Nanopore to a target depth of 20-3 OX, which will allow for haplotype phasing.
[0114] As the pilot cohort expands, methods development for the plasma B AF classifier will be continued. This approach will first use quality filters to remove noise from mapping error and misidentified homozygous SNPs. Long-read phasing will subsequently be used to group regional SNPs into alleles in segments of 25 SNPs or greater. Finally, the ‘major’ allele will be empirically observed in each phased SNP group, and the statistical deviation of the observed BAF values will be characterized from an expected diploid value of 0.5. This local BAF will be aggregated across the genome to generate a sample-wide BAF score indicative of aneuploidy load and allelic imbalance, a proxy for TF. Methods development will focus on (1) understanding underlying patterns that contribute to “noise” in BAF at both the individual SNP and regional level, such as mapping error and cfDNA fragment length, (2) optimizing statistical inference in noise-filtered data, and 3) integrating additional promising features such as regional coverage depth biases and fragmentomics at individual SNP sites, both of which have been explored extensively in the tumor-informed CNV work6.
[0115] Statistical considerations. Following methods development and optimization, sample level BAF scores from pretreatment BC plasma samples will be compared to non-cancer samples as outlined above. Sensitivity for BC will be assessed against a 95% specificity threshold in the non-cancer control plasma samples. Plasma BAF classifier performance will be assessed through use of area under the receiver operating curve (AUC) for positive (pretreatment BC) and negative (non-cancer control) labels.-32-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 Lower limit of detection (LLOD) for the plasma BAF approach will be assessed through in silico mixing studies of pre- and postoperative plasma samples from the same patient.
[0116] The study will include 50 patients with pretreatment disease and 30 controls, providing sufficient power to evaluate relationships between baseline circulating tumor DNA (ctDNA) shedding and radiographic tumor volume. Appropriate statistical tests such as correlation or regression analyses will be used to assess these relationships. For comparisons across BC subtypes (HR+ / HER2-, HER2+, TNBC), analysis of variance (ANOVA) or Kruskal-Wallis tests will be employed, depending on the distribution of the data, to validate findings from existing literature such as increased ctDNA shedding observed in TNBC compared to HR+ patients15.
[0117] Expected outcomes: It is expected that the optimized BAF classifier to have LLOD on par with the tumor-informed classifier (5*10'5) in in silico mixing studies, providing adequate sensitivity to detect ctDNA shedding from a majority (>80%) of small, early-stage breast cancers. These findings will have critical implications for noninvasive cancer screening through liquid biopsy and can pave the way for a larger screening trial to complement radiographic imaging.
[0118] Pitfalls and alternatives: If ctDNA signal are not detected using the plasma BAF classifier, it is possible to sequence to higher depth or include patients with larger (e.g., node positive stage III) disease where there is strong preliminary data from outside platforms of ctDNA shedding at TF concentrations > l*10'4, the typical limit for commercial targeted panels15. As matched tumor tissue will be available in biobanks, tumor tissue can be sequenced in undetectable patients to (1) confirm ctDNA shedding, (2) compare plasma aneuploidy calls to known tumor aneuploidy profiles to assess accuracy at the regional level, and (3) compare inferred signal to TF estimation from the tumor- informed BAF classifier.
[0119] Aim 2: Expand BAF -based early detection to other highly aneuploid tumor types. Despite the extensive development of MCED tests, largely based on methylation patterns, sensitivity for early-stage disease is suboptimal in high aneuploidy tumor types including pancreatic, ovarian, and esophageal cancer27, while analytical validation of these approaches demonstrated an LLOD of roughly l*10'3in several cancer types28. Further,-33-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 positive predictive value in a leading approach is < 40%27, which limits applicability to broad populations due to expense and harms of widespread false positives and may form an even larger problem across serial repeat screenings. The plasma BAF approach has shown promise in detecting clinical stage I BC and has vast potential to increase sensitivity in other high aneuploidy tumor types. Almost all healthy tissue is diploid, and compared to methylation and single nucleotide variant (SNV) approaches which may capture false positive signal due to tissue injury29(e.g., falsely detecting lung cancer in an individual with pneumonia or chronic-obstructive pulmonary disease), aneuploidy is likely to be highly specific for invasive cancer or advanced non-invasive neoplasia.
[0120] It is therefore proposed to develop and implement the plasma BAF approach in patients with early-stage ovarian, esophageal, and pancreatic cancers. Samples will be sourced from pretreatment samples. The optimized BAF approach (Aim 1), combining high-depth WGS and long read sequencing of normal tissue, will be applied to each cancer type to establish sensitivity for early-stage disease. Similar to BC, clinical relationships such as tumor volume and extent of ctDNA shedding will be analyzed, as well as extent of baseline ctDNA shedding as a biomarker for recurrent disease.[01211 Aneuploidy profiles vary widely between solid tumors of different sites of origin, though patterns are highly conserved within individual cancer types16. Application of the BAF classifier to all cancer types (breast, esophageal, ovarian, pancreatic) will allow empirical observation of plasma CNV profiles based on BAF signal. This presents a unique opportunity to link observed aneuploidy profiles to primary disease site. Here, machine learning will be used to train a classifier to predict underlying tumor type based on regional BAF classifier outputs. If needed, this regional aneuploidy profile classifier can be pretrained with tumor BAF profiles from large WGS datasets30, and later adapted to the specific task of classifying sparse aneuploidy signal in plasma. This primary site aneuploidy classifier will transform the plasma BAF approach into a novel MCED platform with a completely different mechanism than conventional tests.
[0122] Methods: Pretreatment, early-stage cancer samples will be collected from gastrointestinal (esophageal and pancreatic) and genitourinary (ovarian) services. The goal will be to collect 40 early-stage patients from each disease type, including a mix of stage I, stage II, and stage III disease. Similar to Aim 1, performance will be evaluated through-34-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 AUC, comparing sample-wide B AF signal from the optimized BAF classifier in cancer samples (positive labels) to non-cancer control plasma samples (negative label). Notably, the non-cancer control plasma dataset (Aim 1) can be used as the negative label for each subsequent tumor type, reducing sequencing costs.
[0123] To design a primary site aneuploidy profile classifier based on regional plasma BAF patterns, machine learning algorithms will be applied such as random forests or support vector machines to the dataset. Cross-validation techniques, including k-fold cross-validation, will be employed to avoid overfitting and ensure the generalizability of the models. Similar approaches have been used in fragmentomics models31.
[0124] Statistical considerations. General sample level and cohort level plasma BAF analyses as in Aim 1. For the primary site aneuploidy profile classifier, performance of the optimized approach will be assessed through overall accuracy, precision per tumortype, recall per tumor-type, and Fl -score per tumor type.
[0125] Expected outcomes: The plasma BAF classifier can detect ctDNA shedding in each cancer type proportional to the extent of underlying aneuploidy and disease stage.As HR+ / HER2- BC is a difficult ctDNA detection target, an equal or improved sensitivity is expected for the plasma BAF classifier in esophageal, ovarian, and pancreatic cancer. The increased sensitivity of the approach may be clinically transformative in these cancers that lack established screening protocols. Ultimately, a novel MCED platform is expected to be produced with greater sensitivity and specificity than current state-of-the-art MCED approaches. These improvements, combined with rapidly declining costs of WGS and a simple laboratory workflow, are likely to drive widespread adoption of this approach.
[0126] Pitfalls and alternatives: As in Aim 1, to increase sensitivity, plasma samples can be sequenced to higher depth. Matched tumor tissue can be sequenced to help guide classifier development. Importantly, even if it fails to achieve best-in-class sensitivity for early-stage cancer with the plasma BAF classifier, the high-depth plasma WGS samples will form a unique resource that should substantially increase sensitivity for orthogonal approaches rooted in fragmentomics31or SNV detection6. The WGS data generated will be an invaluable resource for the lab and the broader liquid biopsy community.
[0127] References-35-4903-2435-7751.1Atty. Dkt. No.: 115872-33431. Evans, D. G. et al. Intensive breast screening in BRCA2 mutation carriers is associated with reduced breast cancer specific and all cause mortality. Hered. Cancer Clin. Pract. 14, 8 (2016).2. Ferreira, C. S., Rodrigues, J., Moreira, S., Ribeiro, F. & Longatto-Filho, A. Breast cancer screening adherence rates and barriers of implementation in ethnic, cultural and religious minorities: A systematic review. Mol Clin Oncol 15, 139 (2021).3. Mann, R. M., Kuhl, C. K. & Moy, L. Contrast-enhanced MRI for breast cancer screening. J. Magn. Reson. Imaging 50, 377-390 (2019).4. Ahmed, A. T. et al. Racial disparities in screening mammography in the United States: A systematic review and meta-analysis. J. Am. Coll. Radiol. 14, 157-165. e9 (2017).5. Chen, T., Kharazmi, E. & Fallah, M. Race and ethnicity-adjusted age recommendation for initiating breast cancer screening. JAMA Netw. Open 6, e238893 (2023).6. Widman, A. J. et al. Ultrasensitive plasma-based monitoring of tumor burden using machine-learning-guided signal enrichment. Nat. Med. (2024) doi: 10.1038 / s41591- 024-03040-4.7. Powles, T. et al. ctDNA guiding adjuvant immunotherapy in urothelial carcinoma. Nature (2021) doi: 10.1038 / s41586-021-03642-9.8. Phallen, J. et al. Direct detection of early-stage cancers using circulating tumor DNA. Sci. Transl. Med. 9, (2017).9. Tie, J. et al. Circulating Tumor DNA Analysis Guiding Adjuvant Therapy in Stage II Colon Cancer. N. Engl. J. Med. 386, 2261-2272 (2022).10. Zviran, A. et al. Genome-wide cell-free DNA mutational integration enables ultrasensitive cancer monitoring. Nat. Med. 26, 1114-1124 (2020).-36-4903-2435-7751.1Atty. Dkt. No.: 115872-334311. Almogy, G. et al. Cost-efficient whole genome-sequencing using novel mostly natural sequencing-by-synthesis chemistry and open fluidics platform. bioRxiv 2022.05.29.493900 (2022) doi: 10.1101 / 2022.05.29.493900.12. American Cancer Society. Breast Cancer Facts & Figures 2024-2025. Preprint at https: / / www.cancer.org / content / dam / cancer-org / research / cancer-facts-and- statistics / breast-cancer-facts-and-figures / 2024 / breast-cancer-facts-and-figures- 2024.pdf (2024).13. Holm, J. et al. Risk factors and tumor characteristics of interval cancers by mammographic density. J. Clin. Oncol. 33, 1030-1037 (2015).14. Lehman, C. D. et al. Diagnostic accuracy of digital screening mammography with and without computer-aided detection. JAMA Intern. Med. 175, 1828-1837 (2015).15. Magbanua, M. J. M. et al. Circulating tumor DNA in neoadjuvant-treated breast cancer reflects response and survival. Ann. Oncol. 32, 229-239 (2021).16. Taylor, A. M. et al. Genomic and Functional Approaches to Understanding Cancer Aneuploidy. Cancer Cell 33, 676-689. e3 (2018).17. Cancer Genome Atlas Network. Comprehensive molecular portraits of human breast tumours. Nature 490, 61-70 (2012).18. Widman, A. J. et al. Machine learning guided signal enrichment for ultrasensitive plasma tumor burden monitoring. bioRxiv 2022.01.17.476508 (2022) doi: 10.1101 / 2022.01.17.476508.19. Turner, N. C. et al. Results of the c-TRAK TN trial: a clinical trial utilising ctDNA mutation tracking to detect molecular residual disease and trigger intervention in patients with moderate- and high-risk early-stage triple-negative breast cancer. Ann. Oncol. 34, 200-211 (2023).20. Shaw, J. et al. Serial postoperative ctDNA monitoring of breast cancer recurrence. J. Clin. Orthod. 40, 562-562 (2022).-37-4903-2435-7751.1Atty. Dkt. No.: 115872-334321. Adalsteinsson, V. A. et al. Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatic tumors. Nat. Commun. 8, 1324 (2017).22. Doebley, A.-L. et al. Griffin: Framework for clinical cancer subtyping from nucleosome profiling of cell-free DNA. bioRxiv (2021) doi: 10.1101 / 2021.08.31.21262867.23. Benjamini, Y. & Speed, T. P. Summarizing and correcting the GC content bias in high-throughput sequencing. Nucleic Acids Res. 40, e72 (2012).24. 1000 Genomes Project Consortium et al. A global reference for human genetic variation. Nature 526, 68-74 (2015).25. Maestri, S. et al. A long-read sequencing approach for direct haplotype phasing in clinical settings. Int. J. Mol. Sci. 21, 9177 (2020).26. Loh, P.-R. et al. Reference-based phasing using the Haplotype Reference Consortium panel. Nat. Genet. 48, 1443-1448 (2016).27. Klein, E. A. et al. Clinical validation of a targeted methylation-based multi-cancer early detection test using an independent validation set. Ann. Oncol. 32, 1167-1177 (2021).28. Alexander, G. E. et al. Analytical validation of a multi-cancer early detection test with cancer signal origin using a cell-free DNA-based targeted methylation assay. PLoS One 18, e0283001 (2023).29. Lutfi, A., Afghan, M. K., Swed, B. & Kasi, P. M. False-positive liquid biopsy assays secondary to overlapping aberrant methylation from non-cancer disease states. Case Rep. Oncol. 16, 1536-1541 (2023).30. Gerstung, M. et al. The evolutionary history of 2,658 cancers. Nature 578, 122-128 (2020).31. Cristiano, S. et al. Genome-wide cell-free DNA fragmentation in patients with cancer. Nature 570, 385-389 (2019).-38-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 Prediction of Response to Immunotherapy in Subjects Using Population Phasing
[0128] Referring now to FIG. 24, depicted is a graph of B-Allele frequency (BAF) scores across timepoints in a subject. Serial monitoring of response to neoadjuvant immunotherapy and surgery in the patient NA-47, a 61 -year-old female with stage IIIA nonsmall cell lung cancer (squamous histology), using population phasing to phase A and B alleles. Plasma BAF score declines from pretreatment timepoint (A) to preoperative timepoint at week 6 (D) and remains decreased at 6 months post surgery (F). This patient did not have recurrence at 39 month follow up.
[0129] Referring now to FIG. 25, depicted is a graph of fragment (FRAG) scores across timepoints in a subject. Serial monitoring of response to neoadjuvant immunotherapy and surgery in the patient NA-45, a 77-year-old female with stage IIIA nonsmall cell lung cancer (adenocarcinoma), using population phasing to phase A and B alleles. Plasma FRAG score declines from pretreatment timepoint following treatment with immunotherapy and further declines at 6 months post-surgery. This patient did not have recurrence at 44 month follow up.C. Systems and Methods of Selecting Subjects for Cancer Therapies Based on Detection of Allelic Imbalances in Cell-Free Deoxyribonucleic Acid (cfDNA) in Bodily Fluid Samples.
[0130] Referring now to FIG. 26, depicted is a block diagram of a system 100 for selecting subjects for cancer therapies based on detection of allelic imbalances in cell-free deoxyribonucleic acid (cfDNA) in bodily fluid samples. In brief overview, the system 100 may include at least one data processing system 105, at least one sequencing device 110, and at least one administrative device 115, communicatively coupled with one another via at least one network 120. The data processing system 105 may include at least one dataset aggregator 125, at least one sequence phaser 130, at least one allele parser 135, at least one imbalance evaluator 140, at least one metric evaluator 145, at least one output handler 150, and at least one database 155, among others. Each of the components in the system 100 as detailed herein may be implemented using hardware (e.g., one or more processors coupled with memory), or a combination of hardware and software as detailed herein in Section E.-39-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 Each of the components in the system 100 may implement or execute the functionalities detailed herein, such as those described in Sections A, B, and D.
[0131] In further detail, the data processing system 105 may (sometimes herein generally referred to as a computing system or a server) be any computing device including one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The data processing system 105 can be in communication with the sequencing device 110, the administrative device 115, the database 155, and other devices, via the network 120. The data processing system 105 may be situated, located, or otherwise associated with at least one server group. The server group may correspond to a data center, a branch office, or a site at which one or more servers corresponding to the data processing system 105 is situated.
[0132] The data processing system 105 may include one or more subsystems, modules, or components to executing the various processes and tasks detailed herein. On the data processing system 105, the dataset aggregator 125 may acquire or obtain genomic sequence reads from the sequencing device 110. The sequence phaser 130 may perform phasing of the genomic sequence reads to identify corresponding chromosomes (e.g., paternal and maternal chromosomes). The allele parser 135 may divide the genomic sequence reads into a set of portions (or windows or phase groups). The imbalance evaluator 140 may generate a metric for each portion of genomic sequence reads as a function of the differences. The metric evaluator 145 may determine an aggregate score across the portions of genomic sequence reads for the subject. The output handler 150 may generate outputs to select cancer therapies for the subject based on the aggregate score.
[0133] The sequencing device 110 may perform genetic sequencing on a deoxyribonucleic acid (DNA) sample (e.g., including ctDNA and gDNA) taken from a subject to generate genome sequencing datasets (e.g., long read or short read sequences). The genetic sequencing carried out may be a high throughput, massively parallel sequencing technique (sometimes herein referred to as next-generation sequencing), such as long-read whole genome sequencing (WGS), pyrosequencing, Reversible dye-terminator sequencing, SOLiD sequencing, Ion semiconductor sequencing, Helioscope single molecule sequencing, among others. The gene sequencing and RNA sequence may be maintained using one or more files according to a format (e.g., FASTQ, BCL, or VCF formats). The genomic-40-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 sequences obtained via next-generation sequencing comprise a plurality of sequencing reads.
[0134] The administrative device 115 (sometimes herein referred to as an end user computing device) may be any computing device comprising one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The administrative device 115 may be in communication with the data processing system 105 and the sequencing device 110 via the network 120. The administrative device 115 may have at least one display. The administrative device 115 may be associated with an entity (e.g., a clinician) examining the subject or genomic data from the subject. The display may present information about the subject provided by the data processing system 105.
[0135] The database 155 may store and maintain various resources and data associated with the data processing system 105, the sequencing device 110, and the administrative device 115, among others. The database 155 may include a database management system (DBMS) to arrange and organize the data maintained thereon. The database 155 may be in communication with the data processing system 105, the sequencing device 110, and the administrative device 115, via the network 120. While running various operations, the data processing system 105, the sequencing device 110, and the administrative device 115 may access the database 155 to retrieve identified data therefrom. The data processing system 105, the sequencing device 110, and the administrative device 115 may also write data onto the database 155 from running such operations.
[0136] Referring now to FIGs. 27A and 27B, depicted are block diagrams of a process 200 for parsing genomic sequence reads of cell-free deoxyribonucleic acid (cfDNA) in bodily fluid samples to select subjects for cancer therapies based on detection of allelic imbalances. The process 200 may include or correspond to operations performed in the system 100 to obtain sequence reads, detect allelic imbalances from the reads, and select subjects for cancer therapies based on the detection. Starting with FIG. 27 A, under the process 200, the sequencing device 110 may execute, carry out, or otherwise perform genome sequencing on at least one bodily fluid sample from at least one subject 205. The subject 205 may be at risk of or diagnosed with cancer. The cancer may be, for example,-41-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 lung cancer, brain cancer, head and neck cancer, colon cancer, rectal cancer, uterine cancer, endometrial cancer, stomach cancer, ovarian cancer, cervical cancer, bladder cancer, pancreatic cancer, esophageal cancer, prostate cancer, renal cancer, skin cancer, or breast cancer, among others. In some embodiments, the sequencing device 110 may execute, carry out, or otherwise perform genome sequencing on normal tissue from the subject 205 (or from another individual diagnosed as not having cancer).
[0137] The bodily fluid sample taken from the subject 205 may include at least one of plasma, urine, cerebral spinal fluid, blood, or lung fluid, among others. The bodily fluid sample may contain cell-free deoxyribonucleic acid (cfDNA) (e.g., in the form of circulating tumor DNA (ctDNA)) originating from the tumor cells associated with the cancer in the subject 205. The cfDNA in the bodily fluid sample may be used to detect the presence of cancer in the subject 205, and to select cancer therapies to be administered to the subject 205. In some embodiments, the bodily fluid sample may be directly taken from the organ affected by the cancer. For example, when the cancer under examination is lung cancer, the bodily fluid sample may be lung fluids obtained from the lung. In some embodiments, the bodily fluid sample may be taken from a region outside the organ affected by the cancer. For instance, when the cancer for which the subject 205 is to be screened is stomach cancer, the bodily fluid sample may be plasma obtained from one of the arms of the subject 205. The normal tissue sample taken from the subject 205 (or from another subject) may include at least one of organ tissue, plasma, urine, cerebral spinal fluid, blood, or lung fluid, among others.
[0138] In performing the genome sequencing on the bodily fluid sample from the subject 205, the sequencing device 110 may produce, create, otherwise at least one genome sequencing dataset 210. The genome sequencing dataset 210 may include or identify at least one set of genomic sequence reads generated by sequencing of cfDNA on the bodily fluid sample. The sequence depth for the short read sequencing (e.g., for cfDNA) may range between 5-250x. In some embodiments, the sequence depth may include low-depth sequencing (e.g., less than 5-10x), medium-depth sequence (e.g., between lOx and 30x), or high-depth sequence (e.g., more than 30x), among others. The WGS used in the long-read sequencing may also use de novo assembly to combine fragments (e.g., on normal tissue) to create complete sequences, without reliance on a reference genome. The genome-42-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 sequencing dataset 210 may include genomic sequence reads of cfDNA across multiple chromosomes (e.g., chromosomes 1-22 and the sex chromosomes). In some embodiments, the genome sequencing dataset 210 may lack or omit genomic sequence reads of cfDNA corresponding to the sex chromosomes (x and y chromosomes).
[0139] In some embodiments, the genome sequencing dataset 210 may include or identify metadata associated with the set of genomic sequence reads. The metadata may include, for example, an identifier (e.g., an anonymized identifier) for the subject 205, a location identifier corresponding to an anatomic location from which the bodily fluid sample was obtained, an identifier for a type of cancer to be evaluated for in the subject 205, a time of acquisition of the set of genomic sequence reads, an indication of whether the subject 205 has been administered with a cancer therapy, a type of cancer therapy, or a time at which the cancer therapy was administered to the subject 205, among others. In some embodiments, the metadata may be provided separately from the genome sequence dataset 210. For instance, the metadata may be entered by a clinician via the administrative device 115 and provided by the administrative device to the data processing system 105 or the database 155. With the generation, the sequencing device 110 may provide, send, or otherwise transmit the genome sequencing dataset 210 to the data processing system 105 or for storage on the database 155.
[0140] The dataset aggregator 125 on the data processing system 105 may receive, identify, or otherwise obtain the genome sequencing dataset 210. In some embodiments, the dataset aggregator 125 may obtain the genome sequencing dataset 210 from the sequencing device 110 (e.g., as shown) or from the database 155. With obtaining, the dataset aggregator 125 may perform preprocessing (e.g., filtering for quality control) on the genomic sequences identified in the genome sequencing dataset 210. The preprocessing may be for quality control as detailed herein in Section A. In some embodiments, the dataset aggregator 125 may delete, erase, or otherwise remove at least a portion of the genome sequences corresponding to whole genome sequencing of cfDNA with quality scores not satisfying (e.g., less than) a threshold score, from the genome sequencing dataset 210. Conversely, the dataset aggregator 125 may keep or maintain at least a remaining portion of the genome sequences corresponding to whole genome sequencing of cfDNA with quality scores satisfying (e.g., greater than or equal to) the threshold score in the-43-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 genome sequencing dataset 210. The filtered genome sequencing dataset 210 may be used for additional processing by the data processing system 105.
[0141] In some embodiments, the dataset aggregator 125 may identify, retrieve, or otherwise obtain a reference genome sequencing dataset 210’. The reference genome sequencing dataset 210’ may be generated in a similar manner as the genome sequencing dataset 210. The reference genome sequencing dataset 210’ may be generated using the normal tissue sample from the subject 205, another subject, or a group of subjects for population phasing. The reference set of genomic sequence reads may be derived from a non-cancer control sample. In some embodiments, the non-cancer control sample may be from the same subject 205. For instance, the non-cancer control sample may be a sample from another region of the subject 205, unaffected by the cancer, benign, or otherwise lacking cfDNA, such as from benign tissue. In some embodiments, the non-cancer control sample may be from a different subject (e.g., an individual without cancer). The reference genome sequencing dataset 210’ may include or identify at least one set of reference genomic sequence reads generated by sequencing of the normal tissue. The reference genome sequencing dataset 210’ may include or identify at least one set of genomic sequence reads generated by long-read sequencing of gDNA on the normal tissue sample. In some embodiments, the genome sequencing may include long-read sequencing, shortread sequencing, or both long and short-read sequencing of gDNA in the normal tissue in accordance with whole genome sequencing (WGS). The sequence depth for the short read sequencing may range between 5-250x. In some embodiments, the sequence depth may include low-depth sequencing (e.g., less than 5-10x), medium-depth sequence (e.g., between lOx and 30x), or high-depth sequence (e.g., more than 30x), among others.
[0142] The sequence phaser 130 on the data processing system 105 may execute, carry out, or otherwise perform phasing of the set of genomic sequence reads in the genome sequencing dataset 210. The phasing may correspond to the process of assigning or associating portions of genomic sequence reads to single nucleotide polymorphism (SNP) alleles in the bodily fluid sample from the subject 205. The phasing may be used to evaluate the set of genomic sequence reads for allelic imbalances as a result of loss of heterozygosity or copy number amplification (e.g., within a given chromosome) in the cfDNA. In some embodiments, the phasing may include population phasing (e.g.,-44-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 assignment of a given allele as matched with a reference population or not matched with the reference population) or familial phasing (e.g., assignment of a given allele to maternal or paternal chromosome). From phasing, the sequence phaser 130 may create, output, or otherwise generate at least one first set of genomic sequences of SNP alleles 215 A and at least one second set of genomic sequences of SNP alleles 215B. The first set of genomic sequences of SNP alleles 215 A may correspond to a first portion of the set of genomic sequences of the genome sequencing dataset 210. The second set of genomic sequences of SNP alleles 215B may correspond to a second portion of the set of genomic sequences of the genome sequencing dataset 210.
[0143] In some embodiments, the sequence phaser 130 on the data processing system 105 may execute, carry out, or otherwise perform phasing of the set of reference genomic sequence reads in the reference genome sequencing dataset 210’. The phasing on the reference genome sequencing dataset 210’ may be similar to the phasing of the genome sequencing dataset 210 (e.g., using population phasing or familial phasing). The phasing may correspond to the process of assigning or associating portions of genomic sequence reads to single nucleotide polymorphism (SNP) alleles in the normal tissue sample. From phasing the reference genome sequence dataset 210’, the sequence phaser 130 may create, output, or otherwise generate at least one first set of genomic sequences of SNP alleles 215’A and at least one second set of genomic sequences of SNP alleles 215’B. The first set of genomic sequences of SNP alleles 215’ A may correspond to a first portion of the set of genomic sequences of the reference genome sequencing dataset 210’. The second set of genomic sequences of SNP alleles 215’B may correspond to a second portion of the set of genomic sequences of the reference genome sequencing dataset 210’. In some embodiments, the sequence phaser 130 may use the phasing of the set of reference genomic sequence reads in the reference genome sequencing dataset 210’ as a guide to phase the set of genomic sequence reads in the genome sequencing dataset 210. In some embodiments, the sequence phaser 130 may perform the population phasing on the genome sequencing dataset 210 without the reference genome sequencing dataset 210’.
[0144] In some embodiments, the sequence phaser 130 may generate, by phasing the first plurality of genomic sequence reads of the genome sequencing dataset (e.g., the genome sequencing dataset 210), (i) a first portion of the first plurality of genomic sequence-45-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 reads including a first set of genomic sequences of single nucleotide polymorphism (SNP) alleles (e.g., the second sequence of SNP alleles 215 A) in the bodily fluid sample and (ii) a second portion of the second plurality of genomic sequence reads including a second set of genomic sequences of SNP alleles (e.g., the second sequence of SNP alleles 215B) in the bodily fluid sample. The sequence phaser 130 may generate, by phasing the second plurality of genomic sequence reads of the reference sequencing dataset (e.g., the reference genome sequencing dataset 210’), (i) a first portion of the second plurality of genomic sequence reads (e.g., the first sequence of SNP alleles 215’ A) including a first set of genomic sequences of SNP alleles in the normal tissue and (ii) a second portion of the second plurality of genomic sequence reads (e.g., the second sequence of SNP alleles 215’B) including a second set of genomic sequences of SNP alleles in the normal tissue.
[0145] In some embodiments, the first set of genomic sequences of SNP alleles 215 A may correspond to a set of genomic sequences of SNP alleles (e.g., maternal) in the bodily fluid sample. The second genomic sequences of SNP alleles 215B may correspond to a set of genomic sequences of SNP alleles (e.g., paternal) in the bodily fluid sample. For example, the first set of genomic sequences of SNP alleles 215 A may correspond to the portions of the cfDNA in the bodily fluid sample that are attributable or originate from alleles. Conversely, the second set of genomic sequences of SNP alleles 215B may correspond to the portions of the cfDNA in the bodily fluid sample that are attributable or originate from other alleles. In some embodiments, the first set of genomic sequences of SNP alleles 215 A may correspond to at least a portion of the set of genomic sequences of SNP alleles. The second genomic sequences of SNP alleles 215B may correspond to at least a portion of the set of genomic sequences of SNP alleles. For instance, most (e.g., more than 90%) of the first set of genomic sequences of SNP alleles 215 A may correspond to alleles of one type (e.g., paternal alleles), and the remainder (e.g., less than 10%) of the first set of genomic sequences of SNP alleles 215A may correspond to alleles of another type (e.g., maternal alleles). This may be as a result of a hybrid parental chromosomes, as a result of cross-over events. For this reason, the alleles are often referred to as Allele A and Allele B rather than maternal or paternal alleles.
[0146] In some embodiments, the first set of genomic sequences of SNP alleles 215’ A may correspond to a set of genomic sequences of SNP alleles (e.g., maternal) in the-46-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 normal tissue sample. The second genomic sequences of SNP alleles 215’B may correspond to a set of genomic sequences of SNP alleles (e.g., paternal) in the normal tissue sample. In some embodiments, the first set of genomic sequences of SNP alleles 215’ A may correspond to at least a portion of the set of genomic sequences of SNP alleles. The second genomic sequences of SNP alleles 215’B may correspond to at least a portion of the set of genomic sequences of SNP alleles. For instance, most (e.g., more than 90%) of the first set of genomic sequences of SNP alleles 215’ A may correspond to alleles of one type (e.g., paternal alleles), and the remainder (e.g., less than 10%) of the first set of genomic sequences of SNP alleles 215’A may correspond to alleles of another type (e.g., maternal alleles).
[0147] The allele parser 135 on the data processing system 105 may partition, define, or otherwise divide each of the first set of genomic sequences of SNP alleles 215 A and the second set of genomic sequences of SNP alleles 215B for each chromosome into a set of windows 220A-N (hereinafter generally referred to as windows 220). The allele parser 135 may process and traverse through each of the first set of genomic sequences of SNP alleles 215 A and the second set of genomic sequences of SNP alleles 215B. In some embodiments, the allele parser 135 may identify, from the first set of genomic sequences of SNP alleles 215 A and the second set of genomic sequences of SNP alleles 215B, the set of windows 220. Each window 220 may correspond to a respective portion of the set of genomic sequence reads that falls within a respective chromosome. For example, the first window 220A may correspond to the portion of the first genomic sequences of SNP alleles 215 A and the second set of genomic sequences of SNP alleles 215B that are associated with the chromosome 1; the second window 220B may correspond to the respective portion of the first genomic sequences of SNP alleles 215 A and the second set of genomic sequences of SNP alleles 215B that are associated with chromosome 2, and so forth. In some embodiments, the allele parser 135 may partition, define, or otherwise divide each of the first set of genomic sequences of SNP alleles 215’ A and the second set of genomic sequences of SNP alleles 215’B for each chromosome into a set of windows 220. The allele parser 135 may process and traverse through each of the first set of genomic sequences of SNP alleles 215’A and the second set of genomic sequences of SNP alleles 215’B. Each window 220 may correspond to a respective portion of the set of genomic sequence reads that falls within a respective chromosome.-47-4903-2435-7751.1Atty. Dkt. No.: 115872-3343
[0148] In some embodiments, the allele parser 135 may compare (i) the first portion of the first plurality of genomic sequence reads (e.g., the first sequence of SNP alleles 215 A) with the first portion of the second plurality of genomic sequence reads (e.g., the first sequence of SNP alleles 215’ A) and (ii) the second portion of the first plurality of genomic sequence reads (e.g., the second sequence of SNP alleles 215B) with the second portion of the second plurality of genomic sequence reads (e.g., the second sequence of SNP alleles 215’B). From comparison, the allele parser 135 may identify or determine an alignment of the first sequence of SNP alleles 215A derived from the bodily fluid sample with the first sequence of SNP alleles 215’A derived from the normal tissue. The allele parser 135 may identify or determine an alignment of the second sequence of SNP alleles 215B derived from the bodily fluid sample with the second sequence of SNP alleles 215’B derived from the normal tissue sample. Using the alignments, the allele parser 135 may partition and form the set of windows 220 across the genome sequences. Based on the comparison, the allele parser 135 may determine or identify a plurality of phase groups (e.g., an example of windows 220) across the genome sequences. Each phase group may correspond to at least one respective window 220. With the determination of the windows 220, the allele parser 135 may discard the first sequence of SNP alleles 215’ A, the second sequence of SNP alleles 215’B, and the reference genome sequencing dataset 210’ from further processing.
[0149] In some embodiments, the allele parser 135 may divide each of the first set of genomic sequences of SNP alleles 215 A and the second set of genomic sequences of SNP alleles 215B into the set of windows 220 that are non-overlapping with one another. The allele parser 135 may perform as similar division on each of the first set of genomic sequences of SNP alleles 215’ A and the second set of genomic sequences of SNP alleles 215’B into the set of windows 220. Each window 220 may have a length of base pairs for the respective chromosome. In some embodiments, each window 220 may range between 50 to 150M base pairs for the respective chromosome. In some embodiments, the allele parser 135 may apply preprocessing (e.g., filtering for quality control) on one or more of the windows 220. The filter may be to remove artifact biases (e.g., sequencing errors, insertion or deletion errors, or amplification bias) or mapping errors (e.g., assigning of a portion of the genome sequence reads to an incorrect chromosome, poorly mapped regions, removal of SNP alleles with extreme disturbances in allele frequencies, or improper correspondence).-48-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 In some embodiments, the preprocessing may be performed prior to the division of the first set of genomic sequences of SNP alleles 215 A and the second set of genomic sequences of SNP alleles 215B into the set of windows 220. In some embodiments, the preprocessing may be performed prior to the division of the first set of genomic sequences of SNP alleles 215’ A and the second set of genomic sequences of SNP alleles 215’B into the set of windows 220.
[0150] Moving onto FIG. 27B, the imbalance evaluator 140 on the data processing system 105 may calculate, generate, or otherwise determine at least one difference in B- allele frequency (BAF) for each window 220. The difference may be between the respective portion of the set of genomic sequence reads (e.g., in the first set of genomic sequences of SNP alleles 215 A and the second set of genomic sequences of SNP alleles 215B) and a corresponding portion of the reference set of genomic sequence reads that fall within the chromosome for the window 220. BAF may be defined as the estimated number of B alleles divided by the sum of both B allele and the reference allele (A allele) at a given SNP location. In tumor tissue, a BAF of 0 may represent the genotype (A / A or A / -), 0.5 may represent (A / B) and 1 may represent (B / B or B / -). In cfDNA or other bodily fluids, the contribution of tumor-derived DNA is diluted within cfDNA from non-cancer cells.
[0151] In some embodiments, to determine the difference, the imbalance evaluator 140 may determine or identify a relative BAF value corresponding to a BAF in the portion of the first set of genomic sequence reads (e.g., in the first set of genomic sequences of SNP alleles 215 A) in relation to a BAF in the portion of the second set of genomic sequence reads (e.g., the second set of genomic sequences of SNP alleles 215B) for the chromosome associated with the window 220. For each respective portion (e.g., set of genomic sequences of SNP alleles 215 A or 215B), the imbalance evaluator 140 may calculate, identify, or otherwise determine a respective BAF (sometimes referred herein as phase score). In some embodiments, the imbalance evaluator 140 may determine, for each phase group (or window 220), (i) a first phase score based on a frequency of major alleles for each of the plurality of SNP alleles and (ii) a second phase score based on a frequency of minor alleles of the plurality of SNP alleles in the bodily fluid sample.
[0152] In some embodiments, the imbalance evaluator 140 may select or identify the major and minor alleles for the bodily fluid sample based on an adjusted BAF score.-49-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 The imbalance evaluator 140 may calculate, determine, or otherwise generate the adjusted BAF score based on a ratio of the first phase score and the second phase score. In some embodiments, the imbalance evaluator 140 may select or identify the major and minor alleles based on a fragment score (e.g., the combined fragment score as detailed herein in Section C). The identifications of the major and minor alleles may be used to detect imbalances (e.g., aneuploidy, such as amplifications, deletions, or loss of heterozygosity (LOH)).
[0153] With the identification, the imbalance evaluator 140 may determine whether the relative BAF value is positive or negative. The relative BAF value may be positive when the BAF in the portion of the genomic sequence reads for evaluation exceeds the BAF in the corresponding portion of the reference genomic sequence reads derived from the noncancer control sample. If the relative BAF value is negative, the imbalance evaluator 140 may identify the corresponding allele as a major allele. The imbalance evaluator 140 may also discard the relative BAF value and set the difference for the given window 220 to null (e.g., 0). Conversely, if the relative BAF value is positive, the imbalance evaluator 140 may use the relative BAF value as the difference for the window 220. In some embodiments, the imbalance evaluator 140 may identify the corresponding allele as a minor allele. The imbalance evaluator 140 may repeat the determination of the differences in BAFs across the set of windows 220. In some embodiments, the imbalance evaluator 140 may omit the determination of differences in BAFs for the sex chromosomes.
[0154] With the determination of the differences in BAFs, the imbalance evaluator 140 may calculate, determine, or otherwise generate at least one imbalance measure 225A- N (hereinafter generally referred to as imbalance measure 225) for each window 220. The imbalance measure 225 may be generated as a function of the difference in BAF in a given window 220 and the differences in BAFs in one or more other remaining windows 220. The imbalance measure 225 may indicate a statistical significance of the difference in BAF in a given window 220, in comparison to the differences in BAFs in other windows 220. The function may include, for example, a one-direction T-test, a bidirectional T-test, Mann- Whitney U-test, signed-rank test, Kruskal-Wallis H-test, one-way analysis of variance (ANOVA), or regression model (e.g., linear or logistic regression), among others. For example, the imbalance evaluator 140 may generate the imbalance measure 225 A for the-50-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 window 220A, as the function of the difference in BAF in the window 220A versus the differences in BAFs in the other windows 220B-N.
[0155] In some embodiments, the imbalance evaluator 140 may use a T-test (e.g., one-direction or bidirectional) on the difference in BAF in a given window 220 and the differences in BAFs in one or more other remaining windows 220 in determining the imbalance measure 225 for the window 220. In performing the one-direction T-test, the imbalance evaluator 140 may determine a value (e.g., a mean value) for the null hypothesis, using the differences in BAFs in the other windows 220. The imbalance evaluator 140 may also generate a statistic metrics (e.g., standard deviation and mean) of the difference in BAFs across the set of windows 220 to form a distribution (e.g., a normal distribution). The imbalance evaluator 140 may determine a significance value by comparing the difference in BAF in the given window 220 with the distribution. For a one-direction T test, the imbalance evaluator 140 may determine a single significance value relative to the mean. For a bidirectional T test, the imbalance evaluator 140 may determine a positive significance value and a negative significance value on both sides of the mean. The imbalance evaluator 140 may use the significance value as the imbalance measure 225 for the given window 220. The imbalance evaluator 140 may repeat the determination of the imbalance measures 225 over the set of windows 220. In some embodiments, the imbalance evaluator 140 may omit the determination of imbalance measures 225 for the sex chromosomes. In some instances, BAF values may be converted to a binomial distribution based on total coverage to standardize values with respect to varying coverage at individual SNP sites.
[0156] The metric evaluator 145 on the data processing system 105 may calculate, generate, or otherwise determine at least one aggregate score 230 for the subject 205 based on the set of imbalance measures 225. The aggregate score 230 may indicate an allelic imbalance attributable to the cfDNA in the bodily fluid sample of the subject 205. In some embodiments, the metric evaluator 145 may determine the aggregate score 230 as a function of the set of imbalance measures 225. The function may include, for example, a sum, a weighted sum, or a weighted average, among others. The imbalance measures 225 may be the sum of the set of imbalance measures 225 corresponding to the chromosomes (e.g., excluding the sex chromosomes) in the cfDNA in the bodily fluid sample of the subject 205.-51-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 In some embodiments, the metric evaluator 145 may generate a relative BAF score as a function of the first phase score and the second phase score for each phase group of the plurality of phase groups. The relative BAF score can be used as the aggregate score 230.
[0157] Based on the aggregate score 230, the metric evaluator 145 may identify or select the subject 205 for a cancer therapy. To select, the metric evaluator 145 may compare the aggregate score 230 with a threshold. The threshold may delineate, identify, or otherwise define a value for the aggregate score 230 at which to select the subject 205 for cancer therapy. In some embodiments, the threshold may also define the value for the aggregate score 230 at which to determine that subject 205 has cancer (e.g., of the type under examination). When the aggregate score 230 satisfies (e.g., greater than or equal to) the threshold, the metric evaluator 145 may select the subject 205 for the cancer therapy. In addition, the metric evaluator 145 may determine that the subject 205 has cancer. Conversely, when the aggregate score 230 does not satisfy (e.g., less than) the threshold, the metric evaluator 145 may exclude the subject 205 from the cancer therapy. In addition, the metric evaluator 145 may determine that the subject 205 does not have cancer.
[0158] In some embodiments, the metric evaluator 145 may select or identify the cancer therapy for the subject 205 from a set of candidate cancer therapies based on the aggregate score 230. The identification of the type of cancer therapy may be performed, when the aggregate score 230 satisfies the threshold. The set of candidate cancer therapies may include, for example, neoadjuvant therapy (e.g., agents to reduce or shrink tumor, such as doxorubicin, cyclophosphamide, aromatase inhibitor, or cisplatin), radiotherapy (e.g., high-energy radiation such as external beam radiation therapy (EBRT), intensity -modulated radiation therapy (IMRT), or brachytherapy), immunotherapy (e.g., immune checkpoint inhibitor, monoclonal antibodies, cytokines, or cancer vaccines), chemotherapy (e.g., alkylating agents, antimetabolites, anti -microtubule agents, topoisomerase inhibitors, or cytotoxic antibiotics), targeted therapy (e.g., tyrosine-kinase inhibitor, small molecule drug conjugate, serine / threonine kinase, or monoclonal antibody), or surgery (e.g., physical removal, such as lumpectomy, mastectomy, colectomy, prostatectomy, or lobectomy), among others, or any combination thereof. In some embodiments, each candidate therapy may correspond to a range of values for the aggregate score 230. To identify, the metric evaluator 145 may compare the aggregate score 230 with each range of values for the-52-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 corresponding candidate therapy. If the aggregate score 230 is within the range of values for a candidate therapy, the metric evaluator 145 may select the candidate cancer therapy for the subject 205 (e.g., as recommended for administration).
[0159] The metric evaluator 145 may use aggregate scores 230 across one or more administrations of cancer therapy to determine whether to select the subject 205 for additional cancer therapy. In some embodiments, the metric evaluator 145 may select or identify the subject 205 for cancer therapy based on the aggregate score 230, at least one prior aggregate score, and an identification of a cancer therapy previously administered to the subject 205. The prior aggregate score and the identification may be obtained from the metadata or the database 155. The prior aggregate score may correspond to the aggregate score determined prior to the administration of the cancer therapy. The metric evaluator 145 may calculate or determine a difference between the current aggregate score 230 relative to the prior aggregate score. The difference may be associated with an effect of the administration of the cancer therapy on the cancer cells within the subject 205. A negative difference may indicate that the previous administration of the cancer therapy had an effect on the cancer in the subject 205. Conversely, a positive difference may indicate that the previous administration of the cancer therapy had little to no effect on the cancer.
[0160] With the determination, the metric evaluator 145 may compare the difference to a threshold margin. The comparison may be based on the absolute value of the difference. The threshold margin may define, delineate, or identify a value of the difference at which to select the subject for additional cancer therapy. The threshold margin may also identify the value for the difference at which the prior cancer therapy is successful. When the absolute difference satisfies (e.g., greater than or equal to) threshold and the difference is negative, the metric evaluator 145 may exclude the subject 205 from additional cancer therapy. In addition, the metric evaluator 145 may determine that the previous administration of the cancer therapy was effective.
[0161] On the other hand, if the absolute difference does not satisfy (e.g., less than or equal to) threshold and the difference is negative or if the difference is positive, the metric evaluator 145 may select the subject 205 for additional cancer therapy. The metric evaluator 145 may determine that the previous administration of the cancer therapy was ineffective. In some embodiments, the metric evaluator 145 may select or identify the-53-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 cancer therapy for the subject 205 from a set of candidate cancer therapies based on the identification of the previous cancer therapy administered to the subject 205. The identification of the cancer therapy to administer may be in accordance with a schedule. For example, the schedule may specify that invasive surgery to remove tumorous cells, subsequent to provision of neoadjuvant therapy for the associated cancer to the subject 205.
[0162] The output handler 150 on the data processing system 105 may store and maintain an association between subject 205 (e.g., using the subject identifier) and a selection 235 (or exclusion) of the subject 205 for the cancer therapy. The association may be stored on the database 155 as one or more data structures, such as a linked list, a tree, an array, a table, a matrix, a stack, a queue, or a heap, among others. In some embodiments, the output handler 150 may store the metadata associated with the genome sequencing dataset 210, the aggregate score 230, and the set of imbalance measures 225 for the subject 205, with the association on the database 155. When the aggregate score 230 satisfies the threshold, the output handler 150 may store the association between subject 205 and the selection 235 of the cancer therapy for the subject 205 (including the identified cancer therapy). On the other hand, when the aggregate score 230 does not satisfy the threshold, the output handler 150 may store the association between subject 205 and the exclusion of the subject 205 from the cancer therapy.
[0163] In some embodiments, the output handler 150 may provide, send, or otherwise transmit at least one instruction 240 to the administrative device 115. The output handler 150 may generate the instruction 240 to include information based on the association. For instance, the information may include or identify an identifier of the subject 205, the aggregate score 230, an indication of the selection 235 (or exclusion) of the subject 205 for the cancer therapy, and the identified cancer therapy for the subject 205, among others. The instruction 240 may also include a script defining the presentation of the information on the administrative device 115. With the generation, the output handler 150 may send the instruction 240 to the administrative device 115. Upon receipt of the instruction 240, the administrative device 115 may display, render, or otherwise present the information. For instance, the administrative device 115 may display the information via a graphical user interface of an application that is interfacing with the data processing system 105. Using the information, a clinician examining the subject 205 may make further-54-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 determination as to which actions to take or which cancer therapy to administer to the subject 205 to alleviate or treat the cancer. For example, upon presentation of the information in the instruction 240, the subject 205 may be administered with the cancer therapy as identified in the selection 235.
[0164] In this manner, the data processing system 105 may detect allelic imbalance from at least one bodily fluid sample from the subject 205, without the reliance on matched tumor tissue. Since additional tissues are not used for this technique, this can reduce the invasiveness of the technique used on the subject 205 to detect allelic imbalances. In addition, the data processing system 105 may significantly lessen the consumption of amount of computer resources (e.g., processor and memory) in processing to derive useful and meaningful results regarding the genome sequencing dataset. Using the set of imbalance measures 225 and the aggregate score 230, the data processing system 105 may be able to better select subject 205 for cancer therapies, thereby increasing the effectiveness of the administration of therapies in treating the cancer. By extension, the data processing system 105 may reduce the occurrences of useless or ineffective administration of cancer therapies to the subject 205. There is an expectation that such cancer therapies selected by the data processing system 105 will better treat the cancer in the subject 205.
[0165] Referring now to FIG. 28, depicted is a block diagram of a flow of a process 300 for determining allelic imbalances in reads of single nucleotide polymorphism (SNP) alleles in the bodily fluid samples. The process 300 may be implemented by any components detailed herein, such as the system 100, such as the data processing system 105. In some embodiments, under the process 300, from a tumor copy number variation (CNV) profile 305, a portion of the genome corresponding to a deletion 310 and one or more other portions corresponding to neutral 315 may be parsed and identified. The CNV profile 305 may be evaluated to determine plasma B-allele frequencies (BAF) 325 by dividing the genomes into one or more genome windows 330A-C. Each window 330 may correspond to a particular chromosome. For each window 330, the allelic imbalance 335A-C may be determined based on the BAF in the given window 330 versus the BAFs in the all the other windows 330. Based on the allelic imbalances 335A-C (e.g., as a summation), the aggregate allelic imbalance 340 may be determined. Other methods can be used to identify aggregate allelic imbalance.-55-4903-2435-7751.1Atty. Dkt. No.: 115872-3343
[0166] Referring now to FIG. 29, depicted is a flow diagram of a method 400 of selecting subjects for cancer therapies based on detection of allelic imbalances in cell-free deoxyribonucleic acid (cfDNA) in bodily fluid samples. The method 400 can be implemented by any components detailed herein, such as the system 100, such as the data processing system 105, or the system 800. Under the method 400, a computing system may obtain a sequencing dataset of a bodily fluid sample (405). The computing system may phase genomic sequence reads (410). The computing system may compare portions of genomic sequence reads (415). The computing system may identify phase groups (420). The computing system may determine phase scores (425). The computing system may generate a relative BAF score (430). The computing system may determine whether the relative BAF score satisfies a threshold (435). If the aggregate score satisfies the threshold, the computing system may select the subject for cancer therapy (440). In contrast, if the aggregate score does not satisfy the threshold, the computing system may exclude the subject for cancer therapy (445). The computing system may provide based on the selection (450).D. Systems and Methods of Selecting Subjects for Cancer Therapies Based On B- Allele Frequency (BAF) And Fragment Variation
[0167] Referring now to FIG. 30, depicted is a block diagram of a system 500 for selecting subjects for cancer therapies based on B-allele frequency (BAF) and fragment variation. In brief overview, the system 500 may include at least one data processing system 505, at least one sequencing device 510, and at least one administrative device 515, communicatively coupled with one another via at least one network 520. The data processing system 505 may include at least one dataset aggregator 525, at least one sequence phaser 530, at least one allele parser 535, at least one fragment analyzer 540, at least one region classifier 545, at least one aggregate evaluator 550, at least one output handler 555, and at least one database 560, among others. Each of the components in the system 500 as detailed herein may be implemented using hardware (e.g., one or more processors coupled with memory), or a combination of hardware and software as detailed herein in Section E. Each of the components in the system 500 may implement or execute the functionalities detailed herein, such as those described in Sections A, B, and D.-56-4903-2435-7751.1Atty. Dkt. No.: 115872-3343
[0168] In further detail, the data processing system 505 may (sometimes herein generally referred to as a computing system or a server) be any computing device including one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The data processing system 505 can be in communication with the sequencing device 510, the administrative device 515, the database 560, and other devices, via the network 520. The data processing system 505 may be situated, located, or otherwise associated with at least one server group. The server group may correspond to a data center, a branch office, or a site at which one or more servers corresponding to the data processing system 505 is situated.
[0169] The data processing system 505 may include one or more subsystems, modules, or components to executing the various processes and tasks detailed herein. On the data processing system 505, the dataset aggregator 525 may acquire or obtain genomic sequence reads from the sequencing device 510. The sequence phaser 530 may perform phasing of the genomic sequence reads to identify corresponding chromosomes (e.g., paternal and maternal chromosomes). The allele parser 535 may divide the genomic sequence reads into a set of portions (or windows or phase groups) with fragments. The fragment analyzer 540 may identify differences in lengths between fragments across genomic sequence reads. The region classifier 545 may categorize the differences in lengths into respective regions and may generate a regional fragment score. The aggregate evaluator 550 may determine an aggregate fragment score based on the region fragment scores across the regions. The output handler 555 may generate outputs to select cancer therapies for the subject based on the aggregate fragment score.
[0170] The sequencing device 510 may perform genetic sequencing on a deoxyribonucleic acid (DNA) sample (e.g., including ctDNA and gDNA) taken from a subject to generate genome sequencing datasets (e.g., long read or short read sequences). The genetic sequencing carried out may be a high throughput, massively parallel sequencing technique (sometimes herein referred to as next-generation sequencing), such as long-read whole genome sequencing (WGS), pyrosequencing, Reversible dye-terminator sequencing, SOLiD sequencing, Ion semiconductor sequencing, Helioscope single molecule sequencing, among others. The gene sequencing and RNA sequence may be maintained using one or more files according to a format (e.g., FASTQ, BCL, or VCF formats). The genomic-57-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 sequences obtained via next-generation sequencing comprise a plurality of sequencing reads.
[0171] The administrative device 515 (sometimes herein referred to as an end user computing device) may be any computing device comprising one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The administrative device 515 may be in communication with the data processing system 505 and the sequencing device 510 via the network 520. The administrative device 515 may have at least one display. The administrative device 515 may be associated with an entity (e.g., a clinician) examining the subject or genomic data from the subject. The display may present information about the subject provided by the data processing system 505.
[0172] The database 560 may store and maintain various resources and data associated with the data processing system 505, the sequencing device 510, and the administrative device 515, among others. The database 560 may include a database management system (DBMS) to arrange and organize the data maintained thereon. The database 560 may be in communication with the data processing system 505, the sequencing device 510, and the administrative device 515, via the network 520. While running various operations, the data processing system 505, the sequencing device 510, and the administrative device 515 may access the database 560 to retrieve identified data therefrom. The data processing system 505, the sequencing device 510, and the administrative device 515 may also write data onto the database 560 from running such operations.
[0173] Referring now to FIGs. 31A-C, depicted are block diagrams of a process 600 for parsing genomic sequence reads of cell-free deoxyribonucleic acid (cfDNA) in bodily fluid samples to select subjects for cancer therapies based on B-allele frequency (BAF) and fragment variation. The process 600 may include or correspond to operations performed in the system 500 to obtain sequence reads, determine fragment scores from the reads, and select subjects for cancer therapies based on the scores. Starting with FIG. 31 A, under the process 600, the sequencing device 510 may execute, carry out, or otherwise perform genome sequencing on at least one bodily fluid sample from at least one subject 605. The subject 605 may be at risk of or diagnosed with cancer. The cancer may be, for-58-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 example, lung cancer, brain cancer, head and neck cancer, colon cancer, rectal cancer, uterine cancer, endometrial cancer, stomach cancer, ovarian cancer, cervical cancer, bladder cancer, pancreatic cancer, esophageal cancer, prostate cancer, renal cancer, skin cancer, or breast cancer, among others. In some embodiments, the sequencing device 510 may execute, carry out, or otherwise perform genome sequencing on normal tissue sample from the subject 605 (or from another individual diagnosed as not having cancer).
[0174] The bodily fluid sample taken from the subject 605 comprises at least one of plasma, urine, cerebral spinal fluid, blood, or lung fluid, among others. The bodily fluid sample may contain cell-free deoxyribonucleic acid (cfDNA) (e.g., in the form of circulating tumor DNA (ctDNA)) originating from the tumor cells associated with the cancer in the subject 605. The cfDNA in the bodily fluid sample may be used to detect the presence of cancer in the subject 605, and to select cancer therapies to be administered to the subject 605. In some embodiments, the bodily fluid sample may be directly taken from the organ affected by the cancer. For example, when the cancer under examination is lung cancer, the bodily fluid sample may be lung fluids obtained from the lung. In some embodiments, the bodily fluid sample may be taken from a region outside the organ affected by the cancer. For instance, when the cancer for which the subject 605 is to be screened is stomach cancer, the bodily fluid sample may be plasma obtained from one of the arms of the subject 605. The normal tissue sample taken from the subject 205 (or from another subject) may include at least one of organ tissue, plasma, urine, cerebral spinal fluid, blood, or lung fluid, among others.
[0175] In performing the genome sequencing on the bodily fluid sample from the subject 605, the sequencing device 510 may produce, create, otherwise at least one sequencing dataset 610. The sequencing dataset 610 may include or identify at least one set of fragments including at least one set of genomic sequence reads generated by sequencing of cfDNA on the bodily fluid sample. The sequencing dataset 610 may include or identify at least one set of fragments including a set of genomic sequence reads generated by long- read sequencing of gDNA on the normal tissue sample. Each fragment may correspond to pieces of cfDNA in the bodily fluid resulting from breakdown of chromosomes into subsections. In some embodiments, the genome sequencing may include long-read sequencing of gDNA in the normal tissue in accordance with whole genome sequencing-59-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 (WGS). The sequence depth for the short read sequencing (e.g., for cfDNA) may range between 5-250x. In some embodiments, the sequence depth may include low-depth sequencing (e.g., less than 5-10x), medium-depth sequence (e.g., between lOx and 30x), or high-depth sequence (e.g., more than 30x), among others. The WGS used in the long-read sequencing may also use de novo assembly to combine fragments (e.g., on normal tissue) to create complete sequences, without reliance on a reference genome. The sequencing dataset 610 may include genomic sequence reads of fragments of cfDNA across multiple chromosomes (e.g., chromosomes 5-22 and the sex chromosomes). In some embodiments, the sequencing dataset 610 may lack or omit genomic sequence reads of cfDNA corresponding to the sex chromosomes (x and y chromosomes).
[0176] In some embodiments, the genome sequencing dataset 610 may include or identify metadata associated with the set of fragments. The metadata may include, for example, an identifier (e.g., an anonymized identifier) for the subject 605, a location identifier corresponding to an anatomic location from which the bodily fluid sample was obtained, an identifier for a type of cancer to be evaluated for in the subject 605, a time of acquisition of the set of genomic sequence reads, an indication of whether the subject 605 has been administered with a cancer therapy, a type of cancer therapy, or a time at which the cancer therapy was administered to the subject 605, among others. In some embodiments, the metadata may be provided separately from the genome sequence dataset 610. For instance, the metadata may be entered by a clinician via the administrative device 515 and provided by the administrative device to the data processing system 505 or the database 560. With the generation, the sequencing device 510 may provide, send, or otherwise transmit the genome sequencing dataset 610 to the data processing system 505 or for storage on the database 560.
[0177] The dataset aggregator 525 on the data processing system 505 may receive, identify, or otherwise obtain the sequencing dataset 610. In some embodiments, the dataset aggregator 525 may obtain the sequencing dataset 610 from the sequencing device 510 (e.g., as shown) or from the database 560. With obtaining, the dataset aggregator 525 may perform preprocessing (e.g., filtering for quality control) on the genomic sequences identified in the genome sequencing dataset 610. The preprocessing may be for quality control as detailed herein in Section A. In some embodiments, the dataset aggregator 525-60-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 may delete, erase, or otherwise remove at least a portion of the genome sequences corresponding to whole genome sequencing of cfDNA with quality scores not satisfying (e.g., less than) a threshold score, from the genome sequencing dataset 610. Conversely, the dataset aggregator 525 may keep or maintain at least a remaining portion of the genome sequences corresponding to whole genome sequencing of cfDNA with quality scores satisfying (e.g., greater than or equal to) the threshold score in the genome sequencing dataset 610. The filtered genome sequencing dataset 610 may be used for additional processing by the data processing system 505.
[0178] In some embodiments, the dataset aggregator 525 may identify, retrieve, or otherwise obtain a reference genome sequencing dataset 610’. The reference genome sequencing dataset 610’ may be generated in a similar manner as the genome sequencing dataset 610. The reference genome sequencing dataset 610’ may be generated using the normal tissue sample from the subject 605, another subject, or a group of subjects for population phasing. The reference set of genomic sequence reads may be derived from a non-cancer control sample. In some embodiments, the non-cancer control sample may be from the same subject 605. For instance, the non-cancer control sample may be a sample from another region of the subject 605, unaffected by the cancer, benign, or otherwise lacking cfDNA, such as from benign tissue. In some embodiments, the non-cancer control sample may be from a different subject (e.g., an individual without cancer). The reference genome sequencing dataset 610’ may include or identify at least one set of fragments including at least one set of reference genomic sequence reads generated by sequencing of the normal tissue. The reference genome sequencing dataset 610’ may include or identify at least one set of fragments including at least one set of genomic sequence reads generated by long-read sequencing of gDNA on the normal tissue sample. In some embodiments, the genome sequencing may include long-read sequencing long-read sequencing, short-read sequencing, or both long and short-read sequencing of gDNA in the normal tissue in accordance with whole genome sequencing (WGS). The sequence depth for the short read sequencing may range between 5-250x. In some embodiments, the sequence depth may include low-depth sequencing (e.g., less than 5-10x), medium-depth sequence (e.g., between 50x and 30x), or high-depth sequence (e.g., more than 30x), among others.-61-4903-2435-7751.1Atty. Dkt. No.: 115872-3343
[0179] The sequence phaser 530 on the data processing system 505 may execute, carry out, or otherwise perform phasing of the set of genomic sequence reads in the sequencing dataset 610. The phasing may correspond to the process of assigning or associating portions of genomic sequence reads to single nucleotide polymorphism (SNP) alleles (e.g., major or minor alleles) in the bodily fluid sample from the subject 605. The phasing may be used to evaluate the set of genomic sequence reads for fragmentomics. In some embodiments, the phasing may include population phasing (e.g., assignment of a given allele as matched with a reference population or not matched with the reference population) or familial phasing (e.g., assignment of a given allele to maternal or paternal chromosome). The major allele may correspond to an allele that is relatively more common (e.g., for a particular subject or across a cohort or population) at a given SNP. Conversely, the minor major allele may correspond to an allele that is relatively less common (e.g., for a particular subject or across a cohort or population) at the given SNP.
[0180] In the embodiments, the major allele can be selected (e.g., via empirically observation) by evaluating for BAF differences. The major allele may be selected based on the overexpressed allele through BAF. In some embodiments, the major allele can be selected (e.g., via empirically observation) by evaluating for fragment differences. The major allele may be selected based on the overexpressed allele through fragmentomics. In some embodiments, the empiric major allele as observed through BAF can nominate the major allele for evaluating fragment scores. In some embodiments, the empiric major allele as observed through fragment score can identify the major allele for use in evaluating BAF. In some embodiments, rather than empirically observing the major allele, the major allele can be observed through known aneuploidy patterns in cancer from public databases. In such cases one can choose a prevalence cutoff for the major alleles. In some embodiments, the major allele can be inferred from tumor profiles from public databases.
[0181] In some embodiments, the major and minor alleles may be identified by the sequence phaser 630 based on frequency measures (e.g., at least one of relative BAF score, fragment scores, or length metrics as detailed herein). For instance, the sequence phaser 530 may perform phasing on the genome sequencing dataset 610 to candidate alleles. The candidate alleles may include a set of alleles that can be potentially identified as major or minor in the SNPs in the bodily fluid across various subjects for evaluation thereof. The-62-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 sequence phaser 630 may execute phasing of the genome sequencing dataset 610 to identify sets of fragments corresponding to the candidate major or minor alleles. The sequence phaser 630 may determine the frequency measures of the candidate alleles (e.g., at least one of relative BAF score, fragment scores, or length metrics as detailed herein). Using the frequency measures, the sequence phaser 630 may select or identify at least one allele associated with a corresponding set of fragments as the major allele and at least one other allele associated with another corresponding set of fragments as the minor allele. For example, the sequence phaser 630 may identify one allele with the greatest value of the frequency measure as the major allele. Conversely, the sequence phaser 630 may identify one allele with the lowest value of the frequency measure as the minor allele.
[0182] From phasing the genome sequencing dataset 610, the sequence phaser 530 may create, output, or otherwise generate at least one first set of fragments 615 A and at least one second set of fragments 615B. The first set of fragments 615 A may include a first portion of the fragments in the genome sequencing dataset 610 corresponding to a major allele at each SNP in the bodily fluid sample. The second set of fragments 615B may include a first portion of the fragments in the genome sequencing dataset 610 corresponding to a minor allele at each SNP in the bodily fluid sample. In some embodiments, the sequence phaser 530 on the data processing system 505 may execute, carry out, or otherwise perform phasing of the set of reference genomic sequence reads in the reference genome sequencing dataset 610’. The phasing on the reference genome sequencing dataset 610’ may be similar to the phasing of the genome sequencing dataset 610 (e.g., using population phasing or familial phasing). The phasing may correspond to the process of assigning or associating portions of genomic sequence reads to single nucleotide polymorphism (SNP) alleles in the normal tissue sample. From phasing the reference sequence dataset 610’, the sequence phaser 530 may create, output, or otherwise generate at least one first set of fragments 615’ A and at least one second set of fragments 615’B. The first set of fragments 615’ A may include a first portion of the fragments in the genome sequencing dataset 610’ corresponding to the major allele at each SNP in the normal tissue sample. The second set of fragments 615’B may include a first portion of the fragments in the genome sequencing dataset 610’ corresponding to the minor allele at each SNP in the normal tissue sample. In some embodiments, the sequence phaser 530 may use the phasing of the set of reference genomic sequence reads in the reference genome sequencing dataset-63-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 610’ as a guide to phase the set of genomic sequence reads in the genome sequencing dataset 610. In some embodiments, the sequence phaser 530 may perform the population phasing on the genome sequencing dataset 610 without the reference genome sequencing dataset 610’.
[0183] The allele parser 535 on the data processing system 505 may partition, define, or otherwise divide each of the first set of fragments 615 A and the second set of fragments 615B for each chromosome into a set of phase groups 620A-N (hereinafter generally referred to as phase groups 620 or windows). The allele parser 535 may process and traverse through each of the first set of fragments 615 A and the second set of fragments 615B. Each phase group 620 may correspond to a respective portion of the set of genomic sequence reads that falls within a respective chromosome (or other defined window). For example, the first phase group 620A may correspond to the portion of the first set of fragments 615 A and the second set of fragments 615B that are associated with the chromosome 5; the second phase group 620B may correspond to the respective portion of the first set of fragments 615 A and the second set of fragments 615B that are associated with chromosome 2, and so forth. In some embodiments, the allele parser 535 may partition, define, or otherwise divide each of the first set of fragments 615’A and the second set of fragments 615’B for each chromosome into the set of phase groups 620. The allele parser 535 may process and traverse through each of the first set of fragments 615’ A and the second set of fragments 615’B. Each phase group 620 may correspond to a respective portion of the set of genomic sequence reads that falls within a respective chromosome.
[0184] In some embodiments, the allele parser 535 may compare (i) the first portion of the first plurality of fragments (e.g., the first set of fragments 615 A) with the first portion of the second plurality of genomic sequence reads (e.g., the first set of fragments 615’ A) and (ii) the second portion of the first plurality of genomic sequence reads (e.g., the second set of fragments 615B) with the second portion of the second plurality of genomic sequence reads (e.g., the second set of fragments 615’B). From comparison, the allele parser 535 may identify or determine an alignment of the first set of fragments 615 A derived from the bodily fluid sample with the first set of fragments 615’A derived from the normal tissue sample. The allele parser 535 may identify or determine an alignment of the second set of fragments 615B derived from the bodily fluid sample with the second set of fragments-64-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 615’B derived from the normal tissue sample. Using the alignments, the allele parser 135 may partition and form the set of phase groups 620 across the genome sequences. Based on the comparison, the allele parser 535 may determine or identify a plurality of phase groups 620 across the genome sequences from the bodily fluid sample and normal tissue sample. Each phase group 620 may include at least one fragment (e.g., as depicted) from the set of fragments 615A, 615’A, 615B, or 615’B. With the determination of the phase group 620, the allele parser 535 may discard the first set of fragments 615’ A, the second set of fragments 615’B, and the reference genome sequencing dataset 610’ from further processing.
[0185] In some embodiments, the allele parser 535 may divide each of the first set of fragments 615 A and the second set of fragments 615B into the set of phase groups 620 that are non-overlapping with one another. The allele parser 535 may perform as similar division on each of the first set of fragments 615’A and the second set of fragments 615’B into the set of phase groups 620. Each phase group 620 may have a length of base pairs for the respective chromosome. In some embodiments, each phase group 620 may range between 50 to 550M base pairs for the respective chromosome. In some embodiments, the allele parser 535 may apply preprocessing (e.g., filtering for quality control) on one or more of the phase groups 620. The filter may be to remove artifact biases (e.g., sequencing errors, insertion or deletion errors, or amplification bias) or mapping errors (e.g., assigning of a portion of the genome sequence reads to an incorrect chromosome, poorly mapped regions, removal of SNP alleles with extreme disturbances in allele frequencies, or improper correspondence). In some embodiments, the preprocessing may be performed prior to the division of the first set of fragments 615 A and the second set of fragments 615B into the set of phase groups 620. In some embodiments, the preprocessing may be performed prior to the division of the first set of fragments 615’A and the second set of fragments 615’B into the set of phase groups 620.
[0186] Moving onto FIG. 31B, the fragment analyzer 540 executing on the data processing system 105 may may calculate, generate, or otherwise determine a set of length metrics 665 A-N (hereinafter generally referred to as length metrics 665) for the fragments (e.g., from the set of fragments 615 A or 615B corresponding to the SNPs) for each phase group 620. Each length metric 665 may correspond to a respective SNP of the set of SNPs.-65-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 Each length metric 665 may indicate or identify a difference in fragment lengths (e.g., measured in number of base pairs) between (i) the first set of fragments 615 A corresponding to a major allele at the respective SNP and (ii) the second set of fragments 615B corresponding to a minor allele at the respective SNP. To determine, for each phase group 620, the fragment analyzer 540 may select or identify one or more fragments 615 A corresponding to the major alleles at the corresponding SNPs. In addition, the fragment analyzer 540 may select or identify one or more fragments 615B corresponding to the minor alleles at the corresponding SNPs. For each fragment 615A or 615B, the fragment analyzer 540 may identify or determine a respective length. With the determination, the fragment analyzer 540 may compare the respective length of the fragment 615 A corresponding to the major allele and the respective length of the fragment 615B corresponding to the minor allele at the respective SNP. Based on the comparison at each SNP, the fragment analyzer 540 may calculate or determine the length metric 665.
[0187] In some embodiments, the fragment analyzer 540 may calculate, determine, or generate a set of B-allele frequency (BAF) relative scores 670A-N (hereinafter generally referred to as BAF relative scores 670) corresponding to the set of phase groups 620. The BAF relative scores 670 can be determined in a manner as discussed herein in Sections A-C (e.g., similar to imbalance measures 225). For each phase group 620, the fragment analyzer 540 may calculate or determine a first phase score based on a frequency of major alleles for each SNP allele in the bodily fluid sample. In addition, the fragment analyzer 540 may calculate or determine a second phase score based on a frequency of minor alleles for each SNP allele in the bodily fluid sample. With the determinations, the fragment analyzer 540 may generate the BAF relative score 670 for the phase group 620, as a function of the first phase score and the second phase score. The fragment analyzer 540 may iterate through the phase groups 620 to determine the BAF relative scores 670. In some embodiments, the fragment analyzer 540 may omit the determination of the BAF relative scores 670.
[0188] The region classifier 545 executing on the data processing system 105 may categorize, group, or otherwise classify the length metrics 665 into at least one of a set of regions 675A-N (hereinafter generally referred to regions 675) for each phase group 620. Each region 675 may be defined by a range of values for the length metrics 665. Each region 675 may be defined by a respective minimum length and a respective maximum-66-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 length. The set of regions 675 may include, for example, one or more of: (i) a short fragment enrichment peak region, (ii) a depleted median fragment region, and (iii) a short dinucleosome fragment peak region, among others.
[0189] Continuing on, the short fragment enrichment peak region may have the respective minimum length corresponding to 120 base pairs and the respective maximum length corresponding to 160 base pairs. The depleted median fragment region has the respective minimum length corresponding to 165 base pairs and the respective maximum length corresponding to 225 base pairs. The short dinucleosome fragment peak region has the respective minimum length corresponding to 245 base pairs and the respective maximum length corresponding to 315 base pairs. To classify, the region classifier 545 may compare the length metrics 665 from each phase group 620 with the minimum lengths and maximum lengths of the regions 675. From the comparison of each length metric 665, the region classifier 545 may identify the region 675 with the minimum length and maximum length in which the length metric 665 falls (e.g., within the range of base pair values). With the identification, the region classifier 545 may assign or classify the length metric 665 to the identified region 675. The region classifier 545 may iterate or traverse through the length metrics 665 to classify to at least one of the regions 675.
[0190] With the classifications, the region classifier 545 may calculate, generate, or otherwise determine a set of regional fragment scores 680A-N (hereinafter generally referred to as regional fragment scores 680) for the set of regions 675 in each phase group 620. Each regional fragment score 680 may be a combination (e.g., a sum, weighted average, or another function) of the length metrics 665 classified into the respective region 675. For each region 675, the region classifier 545 may determine the respective regional fragment score 680 based on the length metrics 665 classified into the region 675. For instance, the region classifier 545 may combine the length metrics 665 classified into the region 675 to derive a sum or a weighed sum to use as the regional fragment score 680 for the region 675. In some embodiments, the region classifier 545 may determine one or more of: at least one regional fragment score 680 corresponding to the short fragment enrichment peak region; at least one regional fragment score 680 corresponding to the depleted median fragment region; or at least one regional fragment score 680 corresponding to the short dinucleosome fragment peak region, among others.-67-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 [01911 Continuing onto FIG. 31C, the aggregate evaluator 550 executing on the data processing system 105 may calculate, determine, or otherwise generate at least one combined fragment score 685. The combined fragment score 685 may represent biologically meaningful differences in fragment size distributions (e.g., in the cfDNA of the bodily fluid samples) between major and minor alleles in regions of cancer aneuploidy. The combined fragment score 685 may be generated as a function of the set of regional fragment scores 680 for the set of regions 675 across each phase group 620. The function may use a combination (e.g., summation or weighted summation) of the regional fragment scores 680 to yield the combined fragment score 685. In some embodiments, the function may include a difference between (i) a summation of the regional fragment score 680 corresponding to the short fragment enrichment peak region and the regional fragment score 680 corresponding to the short dinucleosome fragment peak region and (ii) the regional fragment score 680 corresponding to the depleted median fragment region.
[0192] In some embodiments, the aggregate evaluator 550 may calculate, determine, or otherwise generate at least one aggregate score based on the combined fragment score 685 and the relative BAF scores 670. For instance, the aggregate evaluator 550 may use a combination (e.g., summation or weighted average) of the combined fragment score 685 and the relative BAF scores 670 to determine the aggregate score. In some embodiments, the aggregate evaluator 550 may generate at least one aggregate BAF score based on the relative BAF scores 670 across the phase group 620. The generation of the aggregate BAF score may be in a similar manner as the determination of the aggregate score 230 using the set of imbalance measures 225 across the set of windows 220, as detailed herein. The aggregate score may be used to evaluate cfDNA burden in the bodily fluid sample from the subject 605.
[0193] The output handler 555 executing on the data processing system 505 may identify or select the subject 605 for a cancer therapy based on the combined fragment score 685 (or the aggregate score). The selection of the subject 605 using the combined fragment score 685 by the output handler 555 as a candidate for administration of cancer therapy may be similar to selection of the subject 205 using the aggregate score 230 as detailed herein Section C. To select, the output handler 555 may compare the combined fragment score 685 (or the aggregate score) with a threshold. The threshold may delineate, identify, or-68-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 otherwise define a value for the combined fragment score 685 at which to select the subject 605 for cancer therapy. In some embodiments, the threshold may also define the value for the combined fragment score 685 at which to determine that subject 605 has cancer (e.g., of the type under examination). When the combined fragment score 685 satisfies (e.g., greater than or equal to) the threshold, the output handler 555 may select the subject 605 for the cancer therapy. In addition, the output handler 555 may determine that the subject 605 has cancer. In some embodiments, when the aggregate score satisfies (e.g., greater than or equal to) the threshold, the output handler 555 may select the subject 605 for the cancer therapy. Conversely, when the combined fragment score 685 does not satisfy (e.g., less than) the threshold, the output handler 555 may exclude the subject 605 from the cancer therapy. In addition, the output handler 555 may determine that the subject 605 does not have cancer. In some embodiments, when the aggregate score does not satisfy (e.g., less than) the threshold, the output handler 555 may exclude the subject 605 from selection.[0194J In some embodiments, the output handler 555 may select or identify the cancer therapy for the subject 605 from a set of candidate cancer therapies based on the combined fragment score 685 (or based on the aggregate score). The identification of the type of cancer therapy may be performed, when the combined fragment score 685 satisfies the threshold. The set of candidate cancer therapies may include, for example, neoadjuvant therapy (e.g., agents to reduce or shrink tumor, such as doxorubicin, cyclophosphamide, aromatase inhibitor, or cisplatin), radiotherapy (e.g., high-energy radiation such as external beam radiation therapy (EBRT), intensity -modulated radiation therapy (IMRT), or brachytherapy), immunotherapy (e.g., immune checkpoint inhibitor, monoclonal antibodies, cytokines, or cancer vaccines), chemotherapy (e.g., alkylating agents, antimetabolites, antimicrotubule agents, topoisomerase inhibitors, or cytotoxic antibiotics), targeted therapy (e.g., tyrosine-kinase inhibitor, small molecule drug conjugate, serine / threonine kinase, or monoclonal antibody), or surgery (e.g., physical removal, such as lumpectomy, mastectomy, colectomy, prostatectomy, or lobectomy), among others, or any combination thereof. In some embodiments, each candidate therapy may correspond to a range of values for the combined fragment score 685. To identify, the output handler 555 may compare the combined fragment score 685 (or the aggregate score) with each range of values for the corresponding candidate therapy. If the combined fragment score 685 is within the range of-69-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 values for a candidate therapy, the output handler 555 may select the candidate cancer therapy for the subject 605 (e.g., as recommended for administration).
[0195] The output handler 555 may use combined fragment scores 685 (or the aggregate scores) across one or more administrations of cancer therapy to determine whether to select the subject 605 for additional cancer therapy. In some embodiments, the output handler 555 may select or identify the subject 605 for cancer therapy based on the combined fragment score 685, at least one prior aggregate score, and an identification of a cancer therapy previously administered to the subject 605. The prior aggregate score and the identification may be obtained from the metadata or the database 560. The prior aggregate score may correspond to the aggregate score determined prior to the administration of the cancer therapy. The output handler 555 may calculate or determine a difference between the current combined fragment score 685 relative to the prior aggregate score. The difference may be associated with an effect of the administration of the cancer therapy on the cancer cells within the subject 605. A negative difference may indicate that the previous administration of the cancer therapy had an effect on the cancer in the subject 605. Conversely, a positive difference may indicate that the previous administration of the cancer therapy had little to no effect on the cancer.
[0196] With the determination, the output handler 555 may compare the difference to a threshold margin. The comparison may be based on the absolute value of the difference. The threshold margin may define, delineate, or identify a value of the difference at which to select the subject for additional cancer therapy. The threshold margin may also identify the value for the difference at which the prior cancer therapy is successful. When the absolute difference satisfies (e.g., greater than or equal to) threshold and the difference is negative, the output handler 555 may exclude the subject 605 from additional cancer therapy. In addition, the output handler 555 may determine that the previous administration of the cancer therapy was effective.
[0197] On the other hand, if the absolute difference does not satisfy (e.g., less than or equal to) threshold and the difference is negative or if the difference is positive, the output handler 555 may select the subject 605 for additional cancer therapy. The output handler 555 may determine that the previous administration of the cancer therapy was ineffective. In some embodiments, the output handler 555 may select or identify the cancer -70-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 therapy for the subject 605 from a set of candidate cancer therapies based on the identification of the previous cancer therapy administered to the subject 605. The identification of the cancer therapy to administer may be in accordance with a schedule. For example, the schedule may specify that invasive surgery to remove tumorous cells, subsequent to provision of neoadjuvant therapy for the associated cancer to the subject 605.
[0198] In this manner, the data processing system 505 may detect cancer-associated aneuploidy at extremely low tumor fractions (as low as 5* I O5) by deriving metrics on differences in fragment length distributions between major and minor alleles at SNP sites. Fragmentomics may provide a signal that is largely independent (orthogonal) to B-allele frequency scores, allowing for the combination of both metrics to improve detection accuracy. This approach may increase robustness, as fragment size variation and allelic imbalance can be detected even when one signal is weak or ambiguous. The data processing system 505 may significantly lessen the consumption in the amount of computer resources (e.g., processor and memory) in processing to derive useful and meaningful results regarding the genome sequencing data. Using the combined fragment scores 685 the data processing system 505 may be able to better select subject 605 for cancer therapies, thereby increasing the effectiveness of the administration of therapies in treating the cancer. By extension, the data processing system 505 may reduce the occurrences of useless or ineffective administration of cancer therapies to the subject 605.
[0199] Referring now to FIG. 32, depicted is a flow diagram of a method 700 of selecting subjects for cancer therapies based on B-allele frequency (BAF) and fragment variation. The method 700 can be implemented by any components detailed herein, such as the system 100, such as the data processing system 105 or 505, or the system 800. Under the method 700, a computing system may obtain a sequencing dataset of a bodily fluid sample (705). The computing system may phase fragment of minor and major alleles (710). The computing system may determine lengths between minor and major alleles (715). The computing system may classify the lengths into regions (720). The computing system may determine regional fragment score for each region (725). The computing system may generate a combined fragment score as a function of the region fragment scores (730). The computing system may generate a relative BAF score (735). The computing system may generate an aggregate score based on the BAF score and the combined fragment score-71-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 (740). The computing system may determine whether the aggregate score satisfies a threshold (745). If the aggregate score satisfies the threshold, the computing system may select the subject for cancer therapy (750). In contrast, if the aggregate score does not satisfy the threshold, the computing system may exclude the subject for cancer therapy (755). The computing system may provide instructions based on the selection (760).E. Computing and Network Environment
[0200] Various operations described herein can be implemented on computer systems. FIG. 33 shows a simplified block diagram of a representative server system 800, client computing system 814, and network 826 usable to implement certain embodiments of the present disclosure. In various embodiments, server system 800 or similar systems can implement services or servers described herein or portions thereof. Client computing system 814 or similar systems can implement clients described herein. The system 800 described herein can be similar to the server system 800. Server system 800 can have a modular design that incorporates a number of modules 802 (e.g., blades in a blade server embodiment); while two modules 802 are shown, any number can be provided. Each module 802 can include processing unit(s) 804 and local storage 806.[02011 Processing unit(s) 804 can include a single processor, which can have one or more cores, or multiple processors. In some embodiments, processing unit(s) 804 can include a general-purpose primary processor as well as one or more special-purpose coprocessors such as graphics processors, digital signal processors, or the like. In some embodiments, some or all processing units 804 can be implemented using customized circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In other embodiments, processing unit(s) 804 can execute instructions stored in local storage 806. Any type of processors in any combination can be included in processing unit(s) 804.
[0202] Local storage 806 can include volatile storage media (e.g., DRAM, SRAM, SDRAM, or the like) and / or non-volatile storage media (e.g., magnetic or optical disk, flash memory, or the like). Storage media incorporated in local storage 806 can be fixed, removable or upgradeable as desired. Local storage 806 can be physically or logically divided into various subunits such as a system memory, a read-only memory (ROM), and a-72-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 permanent storage device. The system memory can be a read-and-write memory device or a volatile read-and-write memory, such as dynamic random-access memory. The system memory can store some or all of the instructions and data that processing unit(s) 804 need at runtime. The ROM can store static data and instructions that are needed by processing unit(s) 804. The permanent storage device can be a non-volatile read-and-write memory device that can store instructions and data even when module 802 is powered down. The term “storage medium” as used herein includes any medium in which data can be stored indefinitely (subject to overwriting, electrical disturbance, power loss, or the like) and does not include carrier waves and transitory electronic signals propagating wirelessly or over wired connections.
[0203] In some embodiments, local storage 806 can store one or more software programs to be executed by processing unit(s) 804, such as an operating system and / or programs implementing various server functions such as functions of the system 100 of FIG. 26 or the system 500 of FIG. 30 or any other system described herein, or any other server(s) associated with system 800 or any other system described herein.
[0204] Software” refers generally to sequences of instructions that, when executed by processing unit(s) 804 cause server system 800 (or portions thereof) to perform various operations, thus defining one or more specific machine embodiments that execute and perform the operations of the software programs. The instructions can be stored as firmware residing in read-only memory and / or program code stored in non-volatile storage media that can be read into volatile working memory for execution by processing unit(s) 804. Software can be implemented as a single program or a collection of separate programs or program modules that interact as desired. From local storage 806 (or non-local storage described below), processing unit(s) 804 can retrieve program instructions to execute and data to process in order to execute various operations described above.
[0205] In some server systems 800, multiple modules 802 can be interconnected via a bus or other interconnect 808, forming a local area network that supports communication between modules 802 and other components of server system 800. Interconnect 808 can be implemented using various technologies including server racks, hubs, routers, etc.
[0206] A wide area network (WAN) interface 810 can provide data communication capability between the local area network (interconnect 808) and the network 826, such as-73-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 the Internet. Technologies can be used, including wired (e.g., Ethernet, IEEE 802.3 standards) and / or wireless technologies (e.g., Wi-Fi, IEEE 802.24 standards).
[0207] In some embodiments, local storage 806 is intended to provide working memory for processing unit(s) 804, providing fast access to programs and / or data to be processed while reducing traffic on interconnect 808. Storage for larger quantities of data can be provided on the local area network by one or more mass storage subsystems 812 that can be connected to interconnect 808. Mass storage subsystem 812 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network-attached storage, and the like can be used. Any data stores or other collections of data described herein as being produced, consumed, or maintained by a service or server can be stored in mass storage subsystem 812. In some embodiments, additional data storage resources may be accessible via WAN interface 810 (potentially with increased latency).
[0208] Server system 800 can operate in response to requests received via WAN interface 810. For example, one of modules 802 can implement a supervisory function and assign discrete tasks to other modules 802 in response to received requests. Work allocation techniques can be used. As requests are processed, results can be returned to the requester via WAN interface 810. Such operation can generally be automated. Further, in some embodiments, WAN interface 810 can connect multiple server systems 800 to each other, providing scalable systems capable of managing high volumes of activity. Other techniques for managing server systems and server farms (collections of server systems that cooperate) can be used, including dynamic resource allocation and reallocation.
[0209] Server system 800 can interact with various user-owned or user-operated devices via a wide-area network such as the Internet. An example of a user-operated device is the client computing system 814. Client computing system 814 can be implemented, for example, as a consumer device such as a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smart watch, eyeglasses), desktop computer, laptop computer, and so on.
[0210] For example, client computing system 814 can communicate via WAN interface 810. Client computing system 814 can include computer components such as processing unit(s) 816, storage device 818, network interface 820, user input device 822,-74-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 and user output device 824. Client computing system 814 can be a computing device implemented in a variety of form factors, such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, or the like.
[0211] Processing unit(s) 816 and storage device 818 can be similar to processing unit(s) 804 and local storage 806 described above. Suitable devices can be selected based on the demands to be placed on client computing system 814; for example, client computing system 814 can be implemented as a “thin” client with limited processing capability or as a high-powered computing device. Client computing system 814 can be provisioned with program code executable by processing unit(s) 816 to enable various interactions with server system 800.
[0212] Network interface 820 can provide a connection to the network 826, such as a wide area network (e.g., the Internet) to which WAN interface 810 of server system 800 is also connected. In various embodiments, network interface 820 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various RF data communication standards such as Wi-Fi, Bluetooth, or cellular data network standards (e.g., 4G, 8G, LTE, etc ).
[0213] User input device 822 can include any device (or devices) via which a user can provide signals to client computing system 814; client computing system 814 can interpret the signals as indicative of particular user requests or information. In various embodiments, user input device 822 can include any or all of a keyboard, touch pad, touch screen, mouse or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, and so on.
[0214] User output device 824 can include any device via which client computing system 814 can provide information to a user. For example, user output device 824 can include a display to present images generated by or delivered to client computing system 814. The display can incorporate various image generation technologies, e.g., a liquid crystal display (LCD), light-emitting diode (LED) including organic light-emitting diodes (OLED), projection system, cathode ray tube (CRT), or the like, together with supporting electronics (e.g., digital-to-analog or analog-to-digital converters, signal processors, or the like). Some embodiments can include a device such as a touchscreen that function as both-75-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 input and output device. In some embodiments, other user output devices 824 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on.
[0215] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a computer-readable storage medium. Many of the features described in this specification can be implemented as processes that are specified as a set of program instructions encoded on a computer-readable storage medium. When these program instructions are executed by one or more processing units, they cause the processing unit(s) to perform various operation indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, processing unit(s) 804 and 816 can provide various functionality for server system 800 and client computing system 814, including any of the functionality described herein as being performed by a server or client, or other functionality.
[0216] It will be appreciated that server system 800 and client computing system 814 are illustrative and that variations and modifications are possible. Computer systems used in connection with embodiments of the present disclosure can have other capabilities not specifically described here. Further, while server system 800 and client computing system 814 are described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be but need not be located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be configured to perform various operations, e.g., by programming a processor or providing appropriate control circuitry, and various blocks might or might not be reconfigurable depending on how the initial configuration is obtained. Embodiments of the present disclosure can be realized in a variety of apparatus including electronic devices implemented using any combination of circuitry and software.
[0217] While the disclosure has been described with respect to specific embodiments, one skilled in the art will recognize that numerous modifications are possible.-76-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 Embodiments of the disclosure can be realized using a variety of computer systems and communication technologies including but not limited to the specific examples described herein. Embodiments of the present disclosure can be realized using any combination of dedicated components and / or programmable processors and / or other programmable devices. The various processes described herein can be implemented on the same processor or different processors in any combination. Where components are described as being configured to perform certain operations, such configuration can be accomplished, e.g., by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, or any combination thereof. Further, while the embodiments described above may make reference to specific hardware and software components, those skilled in the art will appreciate that different combinations of hardware and / or software components may also be used and that particular operations described as being implemented in hardware might also be implemented in software or vice versa.
[0218] Computer programs incorporating various features of the present disclosure may be encoded and stored on various computer-readable storage media; suitable media include magnetic disk or tape, optical storage media such as compact disk (CD) or DVD (digital versatile disk), flash memory, and other non-transitory media. Computer-readable media encoded with the program code may be packaged with a compatible electronic device, or the program code may be provided separately from electronic devices (e.g., via Internet download or as a separately packaged computer-readable storage medium).
[0219] Thus, although the disclosure has been described with respect to specific embodiments, it will be appreciated that the disclosure is intended to cover all modifications and equivalents within the scope of the following claims.EQUIVALENTS
[0220] The present technology is not to be limited in terms of the particular embodiments described in this application, which are intended as single illustrations of individual aspects of the present technology. Many modifications and variations of this present technology can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the present technology, in addition to those enumerated herein, will be apparent-77-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the present technology. It is to be understood that this present technology is not limited to particular methods, reagents, compounds compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.
[0221] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.
[0222] As will be understood by one skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like, include the number recited and refer to ranges which can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 cells refers to groups having 1, 2, or 3 cells. Similarly, a group having 1-5 cells refers to groups having 1, 2, 3, 4, or 5 cells, and so forth.
[0223] All patents, patent applications, provisional applications, and publications referred to or cited herein are incorporated by reference in the entirety, including all figures and tables, to the extent they are not inconsistent with the explicit teachings of this specification.-78-4903-2435-7751.1
Claims
Atty. Dkt. No.: 115872-3343WHAT IS CLAIMED IS:
1. A method of selecting a subject for a cancer therapy based on detection of allelic imbalances in cell-free deoxyribonucleic acid (cfDNA) in a bodily fluid sample of the subject, comprising: obtaining, by one or more processors, for a subject at risk of or diagnosed with cancer, a first sequencing dataset comprising a first plurality of genomic sequence reads generated by sequencing of cfDNA in a bodily fluid sample from the subject generating, by the one or more processors, by phasing the first plurality of genomic sequence reads of the first sequencing dataset, (i) a first portion of the first plurality of genomic sequence reads comprising a first set of genomic sequences of single nucleotide polymorphism (SNP) alleles in the bodily fluid sample and (ii) a second portion of the second plurality of genomic sequence reads comprising a second set of genomic sequences of SNP alleles in the bodily fluid sample; identifying, by the one or more processors, from the first portion of the first plurality of genomic sequence reads and the second portion of the first plurality of genomic sequence reads, a plurality of phase groups; determining, by the one or more processors, for each phase group, (i) a first phase score based on a frequency of major alleles for each of the plurality of SNP alleles and (ii) a second phase score based on a frequency of minor alleles of the plurality of SNP alleles in the bodily fluid sample; generating, by the one or more processors, a relative BAF score as a function of the first phase score and the second phase score for each phase group of the plurality of phase groups; selecting, by the one or more processors, the subject for a cancer therapy responsive to the relative BAF score satisfying a threshold; and storing, by the one or more processors, using one or more data structures, an association between the subject and the cancer therapy.
2. The method of claim 1, further comprising: identifying, by the one or more processors, for the subject, (i) a second relative BAF score and (ii) a second cancer therapy administered to the subject; and-79-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 determining, by the one or more processors, whether the second relative BAF score is different from relative BAF score by the second threshold, and wherein selecting the subject for the cancer therapy further comprises selecting the subject for the cancer therapy based on whether the second relative BAF score is different from aggregate score by the second threshold.
3. The method of claim 1, further comprising: excluding, by the one or more processors, a second subject from the cancer therapy responsive to a second relative BAF score not satisfying the threshold; and storing, by the one or more processors, using one or more data structures, an association between the second subject and exclusion from the cancer therapy.
4. The method of claim 1, wherein obtaining the first sequencing dataset further comprises obtaining the sequencing dataset comprising the first plurality of genomic sequence reads by sequencing of cfDNA in the bodily fluid sample from the subject in accordance with whole genome sequencing (WGS) at a sequence depth, the sequence depth ranging between 5-250x.
5. The method of claim 1, further comprising identifying, by the one or more processors, at least one major allele and at least one minor allele based on an adjusted BAF score generated using a ratio of the first phase score and the second phase score.
6. The method of claim 1, further comprising identifying, by the one or more processors, from the plurality of SNP alleles, the major allele and the minor allele each of the plurality of SNP alleles based on a combined fragment score.
7. The method of claim 1, wherein selecting the subject for cancer therapy further comprises identifying the cancer therapy from a plurality of cancer therapies based on the relative BAF score, wherein the plurality of cancer therapies comprises at least one of neoadjuvant therapy, radiotherapy, immunotherapy, chemotherapy, targeted therapy, or surgery.
8. The method of claim 1, further comprising:-80-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 obtaining, by the one or more processors, a second sequencing dataset comprising a second plurality of genomic sequence reads generated by whole genome sequencing of gDNA in a normal tissue from the subject; generating, by the one or more processors, by phasing the second plurality of genomic sequence reads of the sequencing dataset, (i) a first portion of the second plurality of genomic sequence reads comprising a first set of genomic sequences of single nucleotide polymorphism (SNP) alleles in the normal tissue and (ii) a second portion of the second plurality of genomic sequence reads comprising a second set of genomic sequences of SNP alleles in the normal tissue; and wherein identifying the plurality of phases further comprises comparing (i) the first portion of the first plurality of genomic sequence reads with the first portion of the second plurality of genomic sequence reads and (ii) the second portion of the first plurality of genomic sequence reads with the second portion of the second plurality of genomic sequence reads to identify the plurality of phase groups.
9. The method of claim 1, wherein phasing further comprises at least one of population phasing or familial phasing.
10. The method of claim 1, wherein the bodily fluid sample comprises at least one of plasma, urine, cerebral spinal fluid, or lung fluid, wherein the normal tissue is obtained from the subject.
11. The method of claim 1, wherein the cancer includes at least one of lung cancer, brain cancer, head and neck cancer, colon cancer, rectal cancer, uterine cancer, endometrial cancer, stomach cancer, ovarian cancer, cervical cancer, bladder cancer, pancreatic cancer, esophageal cancer, prostate cancer, renal cancer, skin cancer, or breast cancer.
12. A system for selecting a subject for a cancer therapy based on detection of allelic imbalances in cell-free deoxyribonucleic acid (cfDNA) in a bodily fluid sample of the subject, comprising: one or more processors coupled with memory, configured to:-81-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 obtain, for a subject at risk of or diagnosed with cancer, a first sequencing dataset comprising a first plurality of genomic sequence reads generated by sequencing of cfDNA in a bodily fluid sample from the subject; generate, by phasing the first plurality of genomic sequence reads of the first sequencing dataset, (i) a first portion of the first plurality of genomic sequence reads comprising a first set of genomic sequences of single nucleotide polymorphism (SNP) alleles in the bodily fluid sample and (ii) a second portion of the second plurality of genomic sequence reads comprising a second set of genomic sequences of SNP alleles in the bodily fluid sample; identify, from the first portion of the first plurality of genomic sequence reads and the second portion of the first plurality of genomic sequence reads, a plurality of phase groups; determine, for each phase group, (i) a first phase score based on a frequency of major alleles for each of the plurality of SNP alleles and (ii) a second phase score based on a frequency of minor alleles of the plurality of SNP alleles in the bodily fluid sample; generate a relative BAF score as a function of the first phase score and the second phase score for each phase group of the plurality of phase groups; select the subject for a cancer therapy responsive to the relative BAF score satisfying a threshold; and store, using one or more data structures, an association between the subject and the cancer therapy.
13. The system of claim 12, wherein the one or more processors are further configured to: identify, for the subject, (i) a second relative BAF score and (ii) a second cancer therapy administered to the subject; determine whether the second relative BAF score is different from relative BAF score by the second threshold; and select the subject for the cancer therapy based on whether the second relative BAF score is different from relative BAF score by the second threshold.
14. The system of claim 12, wherein the one or more processors are further configured to:-82-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 exclude the second subject from a cancer therapy responsive to a second relative BAF score not satisfying the threshold; and store, using one or more data structures, an association between the second subject and exclusion from the cancer therapy.
15. The system of claim 12, wherein the one or more processors are further configured to obtain the sequencing dataset comprising the first plurality of genomic sequence reads by sequencing of cfDNA in the bodily fluid sample from the subject in accordance with whole genome sequencing (WGS) at a sequence depth, the sequence depth ranging between 5- 250x.
16. The system of claim 12, wherein the one or more processors are further configured to identify at least one major allele and at least one minor allele based on an adjusted BAF score generated using a ratio of the first phase score and the second phase score.
17. The system of claim 12, wherein the one or more processors are further configured to identify, from the plurality of SNP alleles, the major allele and the minor allele each of the plurality of SNP alleles based on a combined fragment score.
18. The system of claim 12, wherein the one or more processors are further configured to identify the cancer therapy from a plurality of cancer therapies based on the relative BAF score, wherein the plurality of cancer therapies comprises at least one of neoadjuvant therapy, radiotherapy, immunotherapy, chemotherapy, targeted therapy, or surgery.
19. The system of claim 11, wherein the one or more processors are further configured to: obtain a second sequencing dataset comprising a second plurality of genomic sequence reads generated by whole genome sequencing of gDNA in a normal tissue from the subject; generate, by phasing the second plurality of genomic sequence reads of the sequencing dataset, (i) a first portion of the second plurality of genomic sequence reads comprising a first set of genomic sequences of single nucleotide polymorphism (SNP) alleles in the normal tissue and (ii) a second portion of the second plurality of genomic-83-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 sequence reads comprising a second set of genomic sequences of SNP alleles in the normal tissue; and compare (i) the first portion of the first plurality of genomic sequence reads with the first portion of the second plurality of genomic sequence reads and (ii) the second portion of the first plurality of genomic sequence reads with the second portion of the second plurality of genomic sequence reads to identify the plurality of phase groups.
20. The system of claim 11, wherein phasing further comprises at least one of population phasing or familial phasing.
21. The system of claim 11, wherein the bodily fluid sample comprises at least one of plasma, urine, cerebral spinal fluid, or lung fluid, wherein the normal tissue is obtained from the subject.
22. The system of claim 11, wherein the cancer includes at least one of lung cancer, brain cancer, head and neck cancer, colon cancer, rectal cancer, uterine cancer, endometrial cancer, stomach cancer, ovarian cancer, cervical cancer, bladder cancer, pancreatic cancer, esophageal cancer, prostate cancer, renal cancer, skin cancer, or breast cancer.
23. A method of selecting subjects for cancer therapies based on fragment variation, comprising: obtaining, by one or more processors, for a subject at risk of or diagnosed with cancer, a sequencing dataset comprising a first plurality of fragments including a first plurality of genomic sequences generated by whole genome sequencing of cfDNA in a bodily fluid sample from the subject; generating, by the one or more processors, by phasing the first plurality of fragments comprising the first plurality of genomic sequences of the first sequencing dataset corresponding to a plurality of single nucleotide polymorphisms (SNPs), (i) a first portion of the first plurality of fragments comprising a first set of fragments corresponding to a major allele at each of the plurality of SNPs in the bodily fluid sample and (ii) a second portion of the first plurality of fragments comprising a second set of fragments corresponding to a minor allele at each of the plurality of SNPs in the bodily fluid sample;-84-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 identifying, by the one or more processors, from the first portion of the first plurality of fragments and the second portion of the first plurality of fragments, a plurality of phase groups; determining, by the one or more processors, for each phase group of the plurality of phase groups, a plurality of lengths for the plurality of fragments corresponding to the plurality of SNPs, each of the plurality of lengths identifying a difference of lengths between the first set of fragments corresponding to a major allele at a respective SNP of the plurality of SNPs and the second set of fragments corresponding to a minor allele at the respective SNP of the plurality of SNPs; classifying, by the one or more processors, for each phase group of the plurality of phase groups, the plurality of lengths into one of a plurality of regions, each of the plurality of regions defined by a respective minimum length and a respective maximum length; determining, by the one or more processors, for each respective region of the plurality of regions in each phase group of the plurality of phase groups, a regional fragment score based on the difference of lengths for each of the plurality of lengths classified into respective region; generating, by the one or more processors, a combined fragment score as a function of the regional fragment score for each respective region of the plurality of regions in each phase group of the plurality of phase groups; selecting, by the one or more processors, the subject for a cancer therapy responsive to the combined fragment score satisfying a threshold; and storing, by the one or more processors, using one or more data structures, an association between the subject and the cancer therapy.
24. The method of claim 23, further comprising: determining, by the one or more processors, for each phase group, (i) a first phase score based on a frequency of major alleles for each of the plurality of SNP alleles and (ii) a second phase score based on a frequency of minor alleles of the plurality of SNP alleles in the bodily fluid sample; generating, by the one or more processors, a relative BAF score as a function of the first phase score and the second phase score for each phase group of the plurality of phase groups.-85-4903-2435-7751.1Atty. Dkt. No.: 115872-334325. The method of claim 23, further comprising determining, by the one or more processors, an aggregate score based on the combined fragment score and the relative BAF score; and wherein selecting the subject further comprises selecting the subject responsive to the aggregate score satisfying the threshold.
26. The method of claim 23, further comprising identifying, by the one or more processors, from the plurality of SNP alleles, the major allele and the minor allele each of the plurality of SNPs based on at least one of the combined fragment score or the relative BAF score .
27. The method of claim 23, further comprising: generating, by the one or more processors, a second combined fragment score for a second subject as the function of a second regional fragment score for each respective region of a second plurality of regions, selecting, by the one or more processors, the second subject for the cancer therapy responsive to the second combined fragment score not satisfying the threshold; and storing, by the one or more processors, using the one or more data structures, an association between the second subject and exclusion from the cancer therapy.
28. The method of claim 23, wherein classifying the plurality of lengths further comprises classifying the plurality of lengths into one of the plurality of regions, the plurality of regions comprising (i) a short fragment enrichment peak region, (ii) a depleted median fragment region, and (iii) a short dinucleosome fragment peak region.
29. The method of claim 28, wherein generating the combined fragment score further comprises generating the combined fragment score as the function comprising a difference between (i) a summation of a first regional fragment score corresponding to the short fragment enrichment peak region and a second regional fragment score corresponding to the short dinucleosome fragment peak region and (ii) a third regional fragment score corresponding to the depleted median fragment region.-86-4903-2435-7751.1Atty. Dkt. No.: 115872-334330. The method of claim 28, wherein the short fragment enrichment peak region has the respective minimum length corresponding to 120 base pairs and the respective maximum length corresponding to 160 base pairs, wherein the depleted median fragment region has the respective minimum length corresponding to 165 base pairs and the respective maximum length corresponding to 225 base pairs, and wherein the short dinucleosome fragment peak region has the respective minimum length corresponding to 245 base pairs and the respective maximum length corresponding to 315 base pairs.
31. The method of claim 23, further comprising: obtaining, by the one or more processors, a second sequencing dataset comprising a second plurality of fragments including a second plurality of genomic sequences generated by whole genome sequencing of gDNA in a normal tissue from the subject; generating, by the one or more processors, by phasing the second plurality of fragments comprising the second plurality of genomic sequences of the second sequencing dataset corresponding to a plurality of single nucleotide polymorphisms (SNPs), (i) a first portion of the plurality of fragments comprising a first set of fragments corresponding to a major allele at each of the plurality of SNPs in the normal tissue and (ii) a second portion of the plurality of fragments comprising a second set of fragments corresponding to a minor allele at each of the plurality of SNPs in the normal tissue; wherein identifying the plurality of phase groups further comprises comparing (i) the first portion of the first plurality of fragments with the first portion of the second plurality of fragments and (ii) the second portion of the first plurality of fragments with the second portion of the second plurality of fragments to identify the plurality of phase groups.
32. The method of claim 23, further comprising identifying, by the one or more processors, the cancer therapy from a plurality of cancer therapies based on the combined fragment score, wherein the plurality of cancer therapies comprises at least one of neoadjuvant therapy, radiotherapy, immunotherapy, chemotherapy, targeted therapy, or surgery.-87-4903-2435-7751.1Atty. Dkt. No.: 115872-334333. The method of claim 23, wherein phasing further comprises at least one of population phasing or familial phasing.
34. The method of claim 23, wherein the bodily fluid sample comprises at least one of plasma, urine, cerebral spinal fluid, or lung fluid, and wherein the cancer includes at least one of lung cancer, brain cancer, head and neck cancer, colon cancer, rectal cancer, uterine cancer, endometrial cancer, stomach cancer, ovarian cancer, cervical cancer, bladder cancer, pancreatic cancer, esophageal cancer, prostate cancer, renal cancer, skin cancer, or breast cancer.
35. A system for selecting subjects for cancer therapies based on fragment variation, comprising: one or more processors coupled with memory, configured to: obtain, for a subject at risk of or diagnosed with cancer, a first sequencing dataset comprising a first plurality of fragments including a first plurality of genomic sequences generated by long-read sequencing of cfDNA in a bodily fluid sample from the subject; generate, by phasing the first plurality of fragments comprising the first plurality of genomic sequences of the first sequencing dataset corresponding to a plurality of single nucleotide polymorphisms (SNPs), (i) a first portion of the first plurality of fragments comprising a first set of fragments corresponding to a first allele at each of the plurality of SNPs in the bodily fluid sample and (ii) a second portion of the first plurality of fragments comprising a second set of fragments corresponding to a second allele at each of the plurality of SNPs in the bodily fluid sample; identify, from the first portion of the first plurality of fragments and the second portion of the first plurality of fragments, a plurality of phase groups; determine, for each phase group of the plurality of phase groups, a plurality of lengths for the plurality of fragments corresponding to the plurality of SNPs, each of the plurality of lengths identifying a difference of lengths between the first set of fragments corresponding to a major allele at a respective SNP of the plurality of SNPs and the second set of fragments corresponding to a minor allele at the respective SNP of the plurality of SNPs;-88-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 classify, for each phase group of the plurality of phase groups, the plurality of lengths into one of a plurality of regions, each of the plurality of regions defined by a respective minimum length and a respective maximum length; determine, for each respective region of the plurality of regions in each phase group of the plurality of phase groups, a regional fragment score based on the difference of lengths for each of the plurality of lengths classified into respective region; generate a combined fragment score as a function of the regional fragment score for each respective region of the plurality of regions in each phase group of the plurality of phase groups; select the subject for a cancer therapy responsive to the combined fragment score satisfying a threshold; and store, using one or more data structures, an association between the subject and the cancer therapy.
36. The system of claim 35, wherein the one or more processors are further configured to: determine, for each phase group, (i) a first phase score based on a frequency of major alleles for each of the plurality of SNP alleles and (ii) a second phase score based on a frequency of minor alleles of the plurality of SNP alleles in the bodily fluid sample; generate a relative BAF score as a function of the first phase score and the second phase score for each phase group of the plurality of phase groups.
37. The system of claim 35, wherein the one or more processors are further configured to: determine an aggregate score based on the combined fragment score and the relative BAF score; and select the subject responsive to the aggregate score satisfying the threshold.
38. The system of claim 35, wherein the one or more processors are further configured to identify, from the plurality of SNP alleles, the major allele and the minor allele each of the plurality of SNPs based on the relative BAF score .
39. The system of claim 35, wherein the one or more processors are further configured to:-89-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 generate a second combined fragment score for a second subject as the function of a second regional fragment score for each respective region of a second plurality of regions, select the second subject for the cancer therapy responsive to the second combined fragment score not satisfying the threshold; and store, using the one or more data structures, an association between the second subject and exclusion from the cancer therapy.
40. The system of claim 35, wherein the one or more processors are further configured to classify the plurality of lengths into one of the plurality of regions, the plurality of regions comprising (i) a short fragment enrichment peak region, (ii) a depleted median fragment region, and (iii) a short dinucleosome fragment peak region.
41. The system of claim 40, wherein the one or more processors are further configured to generate the combined fragment score as the function comprising a difference between (i) a summation of a first regional fragment score corresponding to the short fragment enrichment peak region and a second regional fragment score corresponding to the short dinucleosome fragment peak region and (ii) a third regional fragment score corresponding to the depleted median fragment region.
42. The system of claim 40, wherein the short fragment enrichment peak region has the respective minimum length corresponding to 120 base pairs and the respective maximum length corresponding to 160 base pairs, wherein the depleted median fragment region has the respective minimum length corresponding to 165 base pairs and the respective maximum length corresponding to 225 base pairs, and wherein the short dinucleosome fragment peak region has the respective minimum length corresponding to 245 base pairs and the respective maximum length corresponding to 315 base pairs.
43. The system of claim 35, wherein the one or more processors are further configured to:-90-4903-2435-7751.1Atty. Dkt. No.: 115872-3343 obtain a second sequencing dataset comprising a second plurality of fragments including a second plurality of genomic sequences generated by whole genome sequencing of gDNA in a normal tissue from the subject; generate, by phasing the second plurality of fragments comprising the second plurality of genomic sequences of the second sequencing dataset corresponding to a plurality of single nucleotide polymorphisms (SNPs), (i) a first portion of the plurality of fragments comprising a first set of fragments corresponding to a major allele at each of the plurality of SNPs in the normal tissue and (ii) a second portion of the plurality of fragments comprising a second set of fragments corresponding to a minor allele at each of the plurality of SNPs in the normal tissue; and compare (i) the first portion of the first plurality of fragments with the first portion of the second plurality of fragments and (ii) the second portion of the first plurality of fragments with the second portion of the second plurality of fragments to identify the plurality of phase groups.
44. The system of claim 35, wherein the one or more processors are further configured to identify the cancer therapy from a plurality of cancer therapies based on the combined fragment score, wherein the plurality of cancer therapies comprises at least one of neoadjuvant therapy, radiotherapy, immunotherapy, chemotherapy, targeted therapy, or surgery.
45. The system of claim 35, wherein phasing further comprises at least one of population phasing or familial phasing.
46. The system of claim 35, wherein the bodily fluid sample comprises at least one of plasma, urine, cerebral spinal fluid, or lung fluid, and wherein the cancer includes at least one of lung cancer, brain cancer, head and neck cancer, colon cancer, rectal cancer, uterine cancer, endometrial cancer, stomach cancer, ovarian cancer, cervical cancer, bladder cancer, pancreatic cancer, esophageal cancer, prostate cancer, renal cancer, skin cancer, or breast cancer.-91-4903-2435-7751.1