Determining latent variable-based fragment mix signatures of nucleic acid molecules

JP2024542031A5Pending Publication Date: 2025-11-06PERSONALIS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024525704
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-01
Filing Date
2022-10-31
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Current methods struggle to accurately detect cancer from plasma samples due to varying tumor DNA characteristics based on tumor type, stage, and shedding, and low tumor DNA concentrations, making it difficult to identify specific features like fragment length and terminal motifs.

Method used

The method involves generating fragmentomic signatures by projecting array size values and terminal motif frequencies onto latent variables using signal separation algorithms, followed by machine learning models to predict disease classification.

Benefits of technology

This approach enables accurate prediction of disease presence, even with low tumor DNA concentrations, by enhancing the detection of cancer through improved fragment length and terminal motif analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The method of predicting a classification of a disease of a subject based on a fragment mix signature may include accessing sequence data of a biological sample of the subject. The method may also include generating a set of sequence size values ​​based on the sequence data. Each sequence size value of the set may correspond to a size of an array of the sequence data. The method also includes determining a fragment mix signature amplitude of the subject by projecting the set of sequence size values ​​onto a latent variable of the fragment mix signature. The latent variable may be generated by applying one or more signal separation algorithms to other sequence size values ​​obtained from one or more reference biological samples. The method may also include generating a result by processing the fragment mix signature amplitude using a machine learning model. The result may include a classification predicting whether the subject is afflicted with a particular disease.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 274,330, entitled "Determining Gene Signatures Based On Latent Variables Of Nucleic Acid Molecules," filed November 1, 2021, the contents of which are incorporated by reference in their entirety for all purposes. [Background technology]

[0002] Next generation sequencing can be used to identify genetic characteristics of a subject. For example, whole genome sequencing can be used to reveal somatic mutations in a subject's sequence data, some of which correspond to tumor DNA. Furthermore, recent research on cell-free DNA fragmentomics characteristics, such as fragment length and terminal motifs, has expanded the possibilities for accurately diagnosing a subject using plasma samples. For example, such research has facilitated the discovery of fragment length signatures that correspond to various types of tumors. Efforts are also being made to develop the above technologies for detecting cancer in a subject.

[0003] Despite such efforts, it remains difficult to accurately detect cancer from plasma samples. One of the reasons for this difficulty is that the amount of tumor DNA in plasma samples can vary greatly depending on the type of tumor, the stage of disease progression, and the degree of release of tumor cell DNA that is accessible to the circulation (tumor "shedding"). Due to these various characteristics, it has been difficult to identify specific features of tumor DNA (e.g., fragment length, terminal motifs). Furthermore, the concentration of tumor DNA in plasma samples can be low to such an extent that accurate diagnosis of the subject becomes difficult. Summary of the Invention

[0004] In some embodiments, a method is provided for predicting a classification of a subject's disease based on a fragment mix signature based on a size distribution of nucleic acid molecules. The method may include accessing sequence data of a biological sample of the subject. The method may also include generating a set of sequence size values ​​based on the sequence data. Each sequence size value of the set may correspond to a sequence size of the sequence data. The method may also include determining a fragment mix signature amplitude of the subject by projecting the set of sequence size values ​​onto a latent variable of the fragment mix signature. The latent variable may be generated by applying one or more signal separation algorithms to other sequence size values ​​obtained from one or more reference biological samples. The method may generate a result by processing the fragment mix signature amplitude using a machine learning model. The result may include a classification predicting whether the subject is afflicted with a particular disease. The method may include outputting the result.

[0005] In some embodiments, a method for predicting a classification of a subject's disease based on a fragment mix signature based on a distribution of terminal motif frequencies is provided. The method may include accessing sequence data of a biological sample of the subject. The method may also include generating a set of terminal motif sequence data based on the sequence data. Each terminal motif sequence data of the set identifies the number or relative frequency of occurrence of nucleic acid molecules having a terminal sequence corresponding to a particular terminal motif. The method also includes determining a fragment mix signature amplitude of the subject by projecting the set of terminal motif sequence data onto a latent variable of a fragment mix signature. The latent variable may be generated by applying one or more signal separation algorithms to other terminal motif sequence data obtained from one or more reference biological samples. The method may generate a result by processing the fragment mix signature amplitude using a machine learning model. The result may include a classification predicting whether the subject is afflicted with a particular disease. The method may include outputting the result.

[0006] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods and / or some or all of the processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause the one or more data processors to perform some or all of one or more of the methods and / or some or all of the processes disclosed herein.

[0007] The terms and expressions used are used as terms of description and not as terms of limitation, and the use of such terms and expressions is not intended to exclude any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the invention as claimed has been specifically disclosed by certain embodiments and optional features, modifications and variations of the concepts disclosed herein may be left to those skilled in the art, and such modifications and variations are believed to be within the scope of the invention as defined by the appended claims.

[0008] The present disclosure will be described in conjunction with the accompanying drawings. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 shows a schematic diagram illustrating the process of generating latent variables from sequence data and using the latent variables to predict the presence of cancer, according to some embodiments. [Diagram 2] 1 includes a flowchart illustrating an example of a method for predicting a subject's disease classification based on a fragment mix signature determined based on a size distribution of nucleic acid molecules, according to some embodiments. [Diagram 3]1 shows a schematic diagram illustrating an example technique for generating array size values, according to some embodiments. [Figure 4] FIG. 1 shows an example diagram illustrating a process for generating a set of unmixed images using a signal separation algorithm, according to some embodiments. [Diagram 5] FIG. 1 shows a schematic diagram illustrating an example technique for generating a set of array-sized latent variables, according to some embodiments. [Figure 6] FIG. 1 shows a schematic diagram illustrating an exemplary technique for generating fragment-mix signatures using an independent component analysis algorithm on whole-exome sequence data, according to some embodiments. [Figure 7] 1 shows a set of latent variables generated by applying a signal separation algorithm to sequence size values ​​corresponding to biological samples of normal subjects. [Figure 8] 13 shows a set of latent variables generated by applying the signal separation algorithm to sequence size values ​​corresponding to a biological sample of another normal subject. [Figure 9] 1 shows a set of latent variables generated by applying a signal separation algorithm to sequence size values ​​corresponding to biological samples of subjects diagnosed with colorectal cancer. [Figure 10] 13 shows a set of latent variables generated by applying the signal separation algorithm to sequence size values ​​corresponding to biological samples of another subject diagnosed with colorectal cancer. [Figure 11] FIG. 1 shows a schematic illustrating a technique for predicting enrichment of DNA molecules derived from loci with different epigenetic states based on fragment mix signatures, according to some embodiments. [Figure 12] FIG. 1 shows a schematic illustrating a process of pre-processing raw sequence size distributions by projecting them onto latent variables, converting the distributions into a set of fragment mix signature amplitudes, and using the amplitudes to detect cancer-associated genetic mutations, according to some embodiments. [Figure 13]FIG. 1 shows a schematic diagram illustrating an exemplary technique for generating a set of latent variables using an independent component analysis algorithm on a hybridization capture sample, according to some embodiments. [Figure 14] FIG. 1 shows a schematic diagram illustrating an exemplary technique for monitoring cancer recurrence using handcrafted or latent variable features, according to some embodiments. [Figure 15] 1 shows a set of receiver operating characteristic (ROC) curves illustrating the level of accuracy when using fragmentmix signatures for disease classification. [Figure 16] 1 includes a flowchart illustrating an example of a method for predicting a subject's disease classification based on a fragment mix signature determined based on terminal motif occurrence frequencies of nucleic acid molecules, according to some embodiments. [Figure 17] 1 illustrates an example of a computer system for implementing some embodiments disclosed herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Fragmentomics generally refers to the analysis of cell-free DNA fragmentation patterns, including fragment sizes and terminal motifs. These fragmentation patterns can be associated with epigenetic signatures specific to tissue types and cancers. In part, differences in histones and associated regulatory proteins that control chromatin structure, as well as the interrelated gene transcription landscapes, in normal and cancer cells may manifest as differences in DNA fragment length distributions. Despite these advances, it is difficult to exploit such features for the development of effective therapeutics (e.g., neo-antigen vaccines). For example, tumor DNA levels in plasma samples are usually low, and such limited data may not provide sufficient accuracy in predicting whether a subject has cancer. In another example, somatic mutations of non-tumor origin found in cell-free DNA, including those derived from clonal hematopoiesis, can complicate variant calling and detection of tumor DNA. In yet another example, the availability of tumor sequence data in sequence knowledge databases may be inconsistent, resulting in failure to serve adequately as training data for training machine learning models to predict whether a subject has cancer.

[0011] To address these challenges, the present technology may include predicting a subject's disease classification based on a fragment mix signature. As used herein, a fragment mix signature refers to one or more signatures of size and / or terminal motif distribution of a nucleic acid molecule that can predict disease classification. For example, a fragment mix signature can represent an estimated distribution of fragment lengths (also referred to as "sizes") of sequences that match each genomic region of a set of genomic regions that can predict disease classification. Additionally or alternatively, a fragment mix signature can represent an estimated distribution of X-bp 5' and 3' sequence identities or "terminal motifs" (e.g., CCCA) that can predict disease classification.

[0012] The fragment mix signature may include one or more latent variables generated by applying a blind source separation (BSS) algorithm to the size and / or terminal motif distribution of nucleic acid molecules obtained from a reference cohort. Each latent variable identifies an estimated size and / or terminal motif distribution of variably enriched sequences across a set of genomic regions and defines a novel basis vector of sequence size and / or terminal motif data. In some examples, the one or more latent variables are determined based on sequence data obtained from a reference sample for disease diagnosis. In some examples, the one or more latent variables identify an estimated distribution of sequence reads having terminal sequences corresponding to a particular set of terminal motifs. The fragment mix signature may be used (for example) to predict whether a subject is afflicted with cancer.

[0013] The technique for predicting a subject's disease classification based on a fragment mix signature can begin by accessing sequence data of a subject's biological sample. In some examples, the sequence data identifies a plurality of nucleic acid sequences obtained by sequencing a plurality of cell-free DNA molecules of the biological sample. The plurality of cell-free DNA molecules can include circulating tumor DNA molecules. Additionally or alternatively, the sequence data can also identify sequence reads corresponding to a plurality of somatic variants detected from the biological sample. The plurality of somatic variants can be detected by aligning each sequence read of the sequence data to a reference sequence (e.g., a human reference genome).

[0014] In some examples, the reference sequences include "normal" or "healthy" sequences obtained from healthy blood cells (e.g., white blood cells), oral cells, and / or hair root cells of one or more subjects. Cells are identified as healthy in various ways, for example, when they have been previously diagnosed as not having a particular type of cancer, or when a sample can be taken from a tissue that is unlikely to contain cancerous or precancerous cells. In some examples, the normal sequences are obtained by (i) separating plasma from a buffy coat containing white blood cells and peripheral blood mononuclear cells of a blood sample, (ii) isolating DNA from the buffy coat, and (iii) sequencing the normal sequences from the isolated DNA. Exemplary techniques for sequencing normal sequences from biological samples are further described in U.S. Pat. No. 10,125,399, the contents of which are incorporated herein by reference for all purposes.

[0015] A set of sequence size values ​​(e.g., a two-dimensional matrix of sequence size values) can be generated based on the sequence data. For example, the set of sequence size values ​​can include a two-dimensional matrix of sequence size values, in which a first dimension that defines a sequence size (e.g., 30 bp) is associated with a second dimension that identifies the number of sequences corresponding to the sequence size (e.g., 50 counts). The set of sequence size values ​​can include an sequence size value representing a size of the sequence for each nucleic acid sequence of the sequence data that aligns to a corresponding genomic region of the set of genomic regions. Each sequence represented by the corresponding sequence size value can include DNA fragments within a particular size (bp) range (e.g., a size range of 60 bp to 600 bp). In some examples, the set of genomic regions is identified using a reference sequence (e.g., a human reference genome). Additionally or alternatively, the set of sequence size values ​​can be converted into a set of empirical probability mass functions, and each empirical probability function can be generated based on the sequence size value of a nucleic acid sequence that aligns to a corresponding genomic region. In some examples, size distribution data representing the set of sequence size values ​​is determined from the sequence data of the subject. For example, a set of sequence size values ​​can be converted into an empirical probability mass function (PMF). As used herein, a PMF refers to a function that estimates the probability that a discrete random variable is exactly equal to some value. The set of PMFs can be used as an input to determine a set of latent variables.

[0016] A fragment mix signature can be determined to predict disease classification using a set of sequence size values ​​of a biological sample. To determine a fragment mix signature, a set of latent variables can be first generated. As used herein, the term "latent variable" refers to an estimated distribution of sequence sizes (a pattern encoded as a numeric vector of signed weights of sequence sizes) that corresponds to one of several potential sources of variation underlying the sequence data aligned to a set of genomic regions. The set of latent variables can be determined by applying one or more signal separation algorithms to sequence size values ​​obtained from one or more reference samples. In some examples, the reference sample includes a biological sample taken from another subject who has received a diagnosis of disease (e.g., cancer, healthy). Additionally or alternatively, the reference sample can include biological samples taken from the same subject at different time points to monitor disease progression. For example, the reference sample can correspond to a biological sample of the subject taken two years prior to the time of taking the current biological sample.

[0017] The one or more signal separation algorithms may include a blind source separation algorithm, an independent component analysis algorithm, and / or a non-negative matrix factorization algorithm. As used herein, a blind source separation algorithm is a technique that separates a set of source signals (e.g., latent variables) from a set of mixed signals (e.g., sequence size values) without (or with little) information about the source signals or the mixing process. Additionally or alternatively, the one or more signal separation algorithms may be one or more other unsupervised machine learning techniques. In some examples, a subset of latent variables (also referred to as "fixed" latent variables) is selected by applying a clustering algorithm to the set of latent variables and identifying the centroid or medoid of each cluster generated by the clustering algorithm. The fixed latent variables can be used as fragment mix signatures that may represent nucleic acid molecules generated by various biological processes reflected in one or more reference samples.

[0018] Then, the set of sequence size values ​​of biological samples can be projected onto the latent variable of fragment mix signature.For example, the size distribution data determined from the set of sequence size values ​​can be projected onto a fixed latent variable to generate one or more latent variable coefficients (also called "amplitudes").Fragment mix signature amplitudes can represent the enrichment of various groups of nucleic acid molecules in biological samples.

[0019] The fragment mix signature amplitudes can be processed using a machine learning model to generate results, including a classification that predicts whether the subject is afflicted with a particular disease. Additionally or alternatively, the results can include a classification that predicts whether the subject is afflicted with a particular type of cancer, which can be used to predict a treatment for the subject or how often the subject should be treated.

[0020] In some examples, the fragment mix signatures can be used to generate results for performing other types of tasks. According to some embodiments, an example task can include measuring the enrichment of a set of protein molecules linked to an epigenetic state in the sequence size data based on identifying a particular latent variable as a signal carried by a set of independently regulated molecules based on the fragment mix signature. For example, the enrichment of a biological sample for a nucleic acid molecule known to be associated with a particular genomic region or chromosome of a particular allele can be used to predict the enrichment of chromatosome binding and silenced heterochromatin states at the corresponding cellular loci that contribute to the sequence data. In another example, the enrichment of a biological sample for a nucleic acid molecule associated with cancer in the sequence data can be used to predict whether a particular allele is derived from a cancer cell. In yet another example, the enrichment of a biological sample for a nucleic acid molecule associated with gene expression in the sequence data can be used to predict whether a particular allele that contributes to the sequence data is expressed.

[0021] Yet another task includes estimating the size variability and central tendency of nucleic acid molecules in a biological sample that are known to bind to protein molecules such as nucleosomes and transcription factors. Another task includes inferring spacing and linker DNA sizes for various proteins by comparing the size differences between peaks of one or more multimodal latent variables. Other tasks include predicting the relative abundance of short and long sequence size species of one or more multimodal latent variables, and inferring cell-free DNA degradation rates for fragmentomics signatures from various rates of change in the latent variable space.

[0022] Additionally or alternatively, fragmentomics signatures can be used for data denoising, which can be advantageous, especially when data volumes are low. Projecting raw data (e.g., size distribution data) into a latent variable space can reduce overfitting downstream of the model and improve the accuracy of statistical tests by removing noise or other data deemed uninteresting. Yet another example includes using latent variables for dimensionality reduction by projecting raw data onto several latent variables to generate several numerical amplitudes (e.g., two if plotting the data in a 2D plot). Numerical amplitudes between reads supporting wild-type and mutant alleles can be compared, thereby ruling out false positive mutations arising from PCR or sequencing errors.

[0023] Thus, some embodiments of the present disclosure provide technical advantages over conventional techniques by using signal separation and clustering algorithms to generate fragment mix signatures that represent accurate and generalizable models of independent sources of variation underlying sequence size and / or terminal motif data in biological samples. For example, blind source separation algorithms can identify hidden but independent factors from sequence data, which can be used to predict a disease or pathology of interest. In some examples, fragment mix signatures can be used to predict the properties of various nucleic acid or protein molecules, which can be used in the diagnosis or prognosis of various epigenetic conditions. As explained above, fragment mix may be able to identify unique properties of nucleic acid molecules that aid in the interpretation of disease phenotypes by understanding what the nucleic acid molecules represent.

[0024] All the aforementioned technical advantages translate into downstream benefits for computational efficiency, interpretation, generalization, and performance of supervised machine learning models. The fragment mix signature, including the fixed latent variables, can then be used to predict whether a biological sample contains tumor DNA and / or can also be used to predict a particular stage of cancer. Furthermore, the fragment mix signature can be used to accurately predict whether a subject suffers from a particular disease even when the amount of tumor DNA detected is low. Thus, the present technology facilitates accurate and reliable detection of disease in a subject by applying a machine learning model to sequence data obtained from cell-free plasma samples.

[0025] The following examples are provided to introduce certain embodiments. In the following description, for the purpose of explanation, specific details are set forth to allow the examples of the present disclosure to be fully understood. However, it will be apparent that various examples can be practiced without these specific details. For example, devices, systems, structures, assemblies, methods, and other components may be shown as components in block diagram form so as not to obscure the examples in unnecessary detail. In other examples, well-known devices, processes, systems, structures, and techniques may be shown without necessary details so as to avoid obscuring the examples. The drawings and descriptions are not intended to be limiting. The terms and expressions used in this disclosure are used as terms of description rather than limitation, and the use of such terms and expressions is not intended to exclude any equivalents of the features shown and described or portions thereof. The term "exemplary" is used herein to mean "to serve as an example, instance, or illustration." Any embodiment or design described herein as an "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs.

[0026] I. Overview of predicting disease classification of subjects based on fragmentomics signatures Paired studies of tumor biopsies and cell-free DNA plasma in subjects revealed distinct characteristics between the two corresponding somatic variant call sets. Differences between tumor and plasma somatic variant call sets pairs were found to be highly variable, potentially influenced based on tumor type, stage, heterogeneity, and time of acquisition of each sample.

[0027] The overall pattern of tumor DNA fragment lengths has been shown to be shorter compared to that of normal cell-free DNA. By enriching biological samples for known tumor variants and then examining the fragment lengths estimated from sequence reads generated from pairs of normal and enriched biological samples, techniques can be implemented to predict features that distinguish disease states. Still, it can be difficult to distinguish tumor DNA from normal DNA based solely on their size distribution. Furthermore, the amount of tumor DNA in a biological sample can vary, and conventional techniques cannot accurately predict whether a particular set of sequences is from tumor DNA. Due to this difficulty, it remains challenging to incorporate DNA fragment lengths of tumor DNA into various techniques (e.g., somatic variant calling).

[0028] To address these challenges, the techniques of the present invention may include determining a fragment mix signature based on latent variables generated by the BSS algorithm. The fragment mix signature may be used to predict the classification of a disease in a subject. The latent variables may be generated by applying one or more machine learning models (e.g., one or more unsupervised machine learning models, the BSS algorithm) to a plurality of sets of sequence size values ​​obtained from a reference sample. The reference sample may be obtained from a subject diagnosed with a disease (e.g., cancer) and / or a healthy subject. The size of each sequence may be measured to identify an array size value. Each set of the plurality of sets of sequence size values ​​may be associated with a genomic region of the set of genomic regions identified in the reference genome. The plurality of sets of sequence size values ​​may be represented as a matrix or any data structure that represents a size distribution of the sequence size values. Additionally or alternatively, each set of sequence size values ​​may be converted to an empirical probability mass function (PMF), thereby representing the plurality of sets of sequence size values ​​with a set of PMFs. Each latent variable represents a set of weighted fragment sizes associated with an independent biological process, technical artifact, or noise source. Biologically meaningful latent variables (e.g., fixed latent variables) can be selected to be used as the fragment mix signature. In some examples, a clustering algorithm is applied to the latent variables to select the fixed latent variables that form the fragment mix signature.

[0029] The fragmentmix signature amplitude of a particular subject (e.g., a subject who is / is not diagnosed with a disease) can be determined by projecting the size distribution of sequences of the particular subject's biological sample onto a set of biologically meaningful latent variables. In some examples, the fragmentmix signature amplitude can be used to predict whether the subject's sequence data contains somatic variants of known tumor origin, which facilitates cell-free variant calling and early cancer detection. In some examples, the fragmentmix signature amplitude is processed by a classification model to predict whether a person is afflicted with a particular disease (e.g., cancer).

[0030] A. Exemplary Techniques for Predicting Disease of a Subject Using Fragmentomics Signatures 1 is a schematic diagram 100 illustrating a process for generating latent variables from sequence data and using the latent variables to predict the presence of cancer, according to some embodiments. The process 100 includes a first stage 102 of (i) generating a set of latent variables using a blind source separation algorithm on a reference sample, and a second stage 104 of (ii) projecting sequence data of new samples onto the fixed latent variables.

[0031] i. Phase 1 In step 106, cell-free DNA molecules (also called "cfDNA") can be obtained from different reference samples. cfDNA can be generated as a result of several biological processes, such as apoptosis and necrosis. These cfDNAs from different sources can have different fragment size distributions, which can be separated by the BSS algorithm. In some examples, the reference sample includes a biological sample (e.g., a serum sample or a plasma sample) taken from a subject diagnosed with a disease (e.g., cancer). As a result, the latent variables determined for the above reference samples can be used as latent variables (or "fixed latent variables") to compare the size distribution of sequence data determined from another biological sample of another subject.

[0032] In step 108, a size distribution of the cell-free DNA molecules can be determined. A set of sequence size values ​​can be identified for each genomic region of the N genomic regions of the reference genome. Each sequence size value in the set of sequence size values ​​can identify a size of a corresponding sequence read aligned to the genomic region (e.g., paired end read fragment insertion size between DNA adaptor sequences, fragment length in bp). A fragment size distribution can then be generated for each reference sample for the set of sequence size values. Only sequence size values ​​between 50 and 550 bp are used in the analysis. In some examples, a PMF is generated for each set of sequence size values ​​to determine the size distribution, where each sequence size value in the set is normalized by the total number of sequences in the set.

[0033] In step 110, the size distribution of the array size values ​​may be processed by a blind source separation algorithm to generate a set of latent variables. In some examples, the set of PMFs is used as input data (e.g., a matrix with N×501 dimensions) for one or more blind source separation algorithms. In this example, the BSS algorithm included an independent component analysis (ICA) algorithm (e.g., fastICA) or a nonnegative matrix factorization (NMF) algorithm. Depending on the type of BSS algorithm, additional formatting of the input data was implemented. For example, the X matrix input to the nonnegative matrix factorization algorithm was the raw PMF described herein. In another example, the X matrix input to the independent component analysis algorithm was first mean centroided and then scaled to unit variance. The BSS algorithm was implemented multiple times until the stochastic initial state converged (e.g., a solution was reached, a minimization or maximization criterion was met, or the non-Gaussianity of the ICA was maximized).

[0034] After performing BSS for each sample, the latent variables from the J reference samples can be formed into stored clusters using a clustering algorithm (e.g., K-means clustering algorithm). The centroid or medoid of each cluster can be selected, and the centroid or medoid can be used as a fixed latent variable that can represent cfDNA from various biological processes. Additionally, or alternatively, a genome-wide PMF of sequence sizes can be identified for each of the J reference samples. This J x 501 (50-550 bp) matrix can be used as input for one or more BSS algorithms. The output latent variables from the BSS algorithms can be used directly as fixed latent variables. The fixed latent variables can be collectively used as fragmentomics signatures to predict cancer in other biological samples.

[0035] ii. Phase 2 In a second step 104, cell-free DNA molecules can be obtained from another biological sample. The other biological sample can be taken from another subject with an unknown disease diagnosis. The other biological sample can be taken from a subject who has undergone cancer treatment and surgery. In some examples, the other biological sample is taken from the same subject from which the reference sample was taken, but at a different time point within a period of time. The other biological sample can be compared to fixed latent variables of the fragment mix signature to predict the classification of the subject's cancer (for example) and / or monitor the progression of the cancer. In step 112, size distribution data of the cfDNA of the other subject can be determined. The size distribution data can include one or more PMFs. The process of determining the size distribution is described in step 108 of FIG. 1.

[0036] In step 114, the size distribution data representing the sequence sizes of the other subjects can be projected onto the fixed latent variables of the fragment mix signature to generate the fragment mix signature amplitudes of the other subjects. To calculate the fragment mix signature amplitudes, the PMF matrix representing the size distribution data of the other subjects is right-multiplied by the inverse of the latent variable matrix. Depending on the BSS algorithm used to generate the latent variables, additional formatting of the input PMF, such as scaling, whitening, etc., can be performed. Additionally or alternatively, a similarity measure (e.g., cosine similarity or Pearson correlation) can be used to measure the enrichment of each fixed latent variable in the fragment size PMF. In some examples, the fragment size PMF of the other samples can be generated from the size distribution of the sequence sizes, and the PMF can be compared or projected against the k fixed latent variables. In some examples, a 1×k feature amplitude vector is selected and used as a low-dimensional representation of the PMF for downstream analysis.

[0037] In step 116, downstream analysis can be performed to predict whether other subjects have cancer based on the fragment mix signature amplitudes of the other subjects. In some examples, the machine learning model trained using the latent variables as features is configured to perform one or more prediction tasks, including (i) predicting whether other subjects carry a cancer-associated genetic mutation, (ii) predicting whether other subjects have a disease (e.g., cancer), (iii) predicting whether other subjects have a particular type of disease (e.g., liver cancer), (iv) predicting the stage of the disease (e.g., stage IV cancer), and (v) predicting whether other subjects have recovered from a disease in response to a particular treatment.

[0038] In some embodiments, the machine learning model includes multiple models (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 machine learning models). In some examples, the trained machine learning model includes a deep neural network. Deep neural networks can be used to capture the internal structure of increasingly large and high-dimensional datasets (e.g., nucleic acid sequence data). Deep neural networks identify high-level features, improving performance over traditional statistical models, increasing interpretability, and providing further understanding of the structure of nucleic acid sequence data.

[0039] Other types of machine learning models may include one or more of gradient boosting decision trees (e.g., XGBoost framework, LightGBM framework), bagging methods, boosting methods, support vector machines, and / or random forest algorithms. For example, gradient boosting may correspond to a type of machine learning technique that can be used for regression and classification problems, as well as for generating predictive models (which may include ensembles of weak predictive models such as decision trees). In some examples, gradient boosting decision trees may include, for example, the XGBoost framework or the LightGBM framework.

[0040] A machine learning model may include hyperparameters. A hyperparameter may be a configuration that is external to the model and whose value is not estimated from data (e.g., training data, input data). In some examples, the hyperparameters are tuned, e.g., tuned to solve a particular predictive modeling problem. In some examples, the hyperparameters are used to estimate model parameters. The hyperparameters may be specified by a user. In some examples, the hyperparameters may be determined using a series of heuristic algorithms.

[0041] B. Methods for predicting a disease classification of a subject based on a fragment mix signature determined from the size distribution of nucleic acid molecules FIG. 2 includes a flowchart 200 illustrating an example of a method for determining a fragment mix signature of a subject based on a size distribution of nucleic acid molecules of a biological sample, according to some embodiments. Some of the operations described in the flowchart 200 may be performed by a computer system (e.g., computer system 1700 of FIG. 17). Although the flowchart 200 may describe the operations as a sequential process, in various embodiments, many of the operations may be performed in parallel or simultaneously. Furthermore, the order of the operations may be rearranged. The operations may have additional steps not shown in the figure. Furthermore, some embodiments of the method may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the associated tasks may be stored in a computer-readable medium, such as a storage medium.

[0042] In step 202, sequence data for a biological sample of the subject can be accessed. In some examples, the sequence data corresponds to a plurality of cell-free DNA molecules of the biological sample, where the plurality of cell-free DNA molecules includes circulating tumor DNA molecules. The sequence data can also include sequences corresponding to a plurality of somatic variants detected in the biological sample.

[0043] In step 204, a set of sequence size values ​​(e.g., a two-dimensional matrix of sequence size values) can be generated based on the sequence data. The sequence size values ​​can include an sequence size value corresponding to the size of the sequence for each sequence of the sequence data that aligns to a corresponding genomic region of the set of genomic regions. Each sequence represented by the corresponding sequence size value includes DNA fragments having a size within a certain range. In some examples, the set of genomic regions is identified using a reference sequence (e.g., a human reference genome). The set of sequence size values ​​can be an empirical probability mass function (PMF) generated based on the sequence of the sequence data that aligns to the corresponding genomic region. In some examples, a PMF is generated for the set of sequence size values ​​to determine a size distribution of the sequence data, where each sequence size value of the set is normalized by the total number of sequences in the biological sample.

[0044] In step 206, the set of sequence size values ​​can be compared or projected against latent variables of a fragment mix signature. The fragment mix signature can include one or more signatures of size distribution of nucleic acid molecules that can predict the classification of a disease. The fragment mix signature can include one or more fixed latent variables. The one or more fixed latent variables can be determined by (i) applying one or more signal separation algorithms to sequence size values ​​obtained from one or more reference samples to generate a set of latent variables, and (ii) applying a clustering algorithm to select a subset of the set of latent variables, which subset corresponds to the fixed latent variables. In some examples, the reference sample includes a biological sample (e.g., tissue, plasma sample) taken from a subject diagnosed with a disease (e.g., cancer). Additionally or alternatively, the reference sample can include biological samples obtained from the same subject at different times. In some examples, each latent variable of the latent variable set includes a histogram or weight vector representing the estimated size distribution of nucleic acid molecules generated by various biological processes. The one or more signal separation algorithms may be blind source separation algorithms, which may include an independent component analysis algorithm and / or a non-negative matrix factorization algorithm. In some examples, a clustering algorithm is applied to the latent variables generated from the different reference samples to determine a subset of the set of latent variables. The subset of latent variables may include centroids or medoids selected from the clusters of identified latent variables.

[0045] Additionally or alternatively, derivatives of the latent variables may be determined by applying transformations such as scaling, translation, averaging, frequency domain transformations (e.g., Fast Fourier Transforms), and other similar transformations.

[0046] In some examples, given that the BSS algorithm has records that estimate signals carried by physically distinct entities, the set of latent variables can predict sets of protein molecules that bind to and protect cell-free DNA (e.g., mononucleosomes, dinucleosomes, monochromatosomes, dichromatosomes, and transcription factor complexes). In some examples, the latent variables can predict potential novel nucleic acid and / or protein entities and their corresponding structures based on associated DNA fragmentation signatures.

[0047] Furthermore, for latent variables corresponding to epigenetic states, density properties such as variance and centroid tendency (e.g., peak sequence size) can predict the degree of cell-to-cell heterogeneity of the bound DNA, and therefore of the proteins enabling the epigenetic state. The relative spacing and average linker DNA size of different proteins in the latent variables may also be predicted by comparing the size difference between the peaks of the multimodal latent variables. Furthermore, the rate of cell-free DNA degradation can be predicted by measuring the relative proportions of short and long size densities (e.g., associated mononucleosomes and dinucleosomes, respectively) of the multimodal latent variables, and the rate of change from long to short sequence sizes along the curves in the latent variable space.

[0048] In step 208, one or more fragment mix signature amplitudes of the biological sample can be determined based on projecting a set of sequence size values ​​(e.g., size distribution data) onto the fragment mix signature. The amplitude of the fragment mix signature can be determined by projecting the size distribution data (e.g., PMF) onto a subset of the latent variables of the fragment mix signature.

[0049] In step 210, the fragment mix signature amplitudes can be used as inputs to a machine learning algorithm (e.g., a logistic regression classification model) to generate results. The results can predict a classification that predicts whether the subject is afflicted with a particular disease. The particular disease can include cancer. In some examples, the results can predict the progression or recurrence of a particular disease when reference samples are taken from the same subject at different time points. The results can be used to identify a treatment for the subject and / or determine how often the subject should be treated. In another example, the results can predict the presence of chromatosome binding at an allele of interest and the presence of silencing of that allele based on an algorithm that determines significant enrichment of monochromatosome and dichromatosome sequences associated with the allele. Thus, by utilizing cell-free sequence data generated from whole exome sequencing of a subject or specific genomic regions (for example), each genomic region of the set of genomic regions can be modeled as a linear mixture of allelic epigenetic states (e.g., transcribed or repressed loci), each with a different DNA sequence size signature and amplitude in the mixture.

[0050] In step 214, the results may be output. For example, the results may be displayed locally or sent to another device. The results may be output along with the subject's identifier. Process 200 then ends.

[0051] II. Sequence Data A. Subjects and Samples Nucleic acid sequence data representing a plurality of nucleic acid molecules can be obtained from a biological sample of a subject to determine a fragment mix signature based on the size of the nucleic acid molecules of the biological sample of the subject. The subject can be a human. The subject can be male or female. The subject can be a fetus, an infant, a child, an adolescent, a teenager, or an adult. The subject can be a patient of any age. For example, the subject can be a patient less than about 10 years old. For example, the subject can be a patient at least about 0, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 years old. The subject is a patient or other individual undergoing or being evaluated for a treatment regimen (e.g., cancer treatment). However, in some examples, the subject is not undergoing a treatment regimen.

[0052] In some examples, the subject may be a mammal or a non-mammal. In some examples, the subject is a mammal, for example, a human, a non-human primate (e.g., ape, monkey, chimpanzee), a cat, a dog, a rabbit, a goat, a horse, a cow, a pig, a rodent, a mouse, a SCID mouse, a rat, a guinea pig, or a sheep. In some embodiments, species variants or homologous genes of these genes are used in non-human animal models. Species variants may be heterologous genes that have maximum sequence identity and similarity of functional properties with each other. Many of such species variant human genes may be listed in the Swiss-Prot database.

[0053] Certain embodiments may include obtaining a sample from a subject, such as a human subject. In some examples, a clinical specimen is obtained from a patient. For example, blood may be drawn from a patient. Certain embodiments may include specifically detecting, profiling, or quantifying molecules (e.g., nucleic acids, DNA, RNA, etc.) in a biological sample.

[0054] The sample may be a tissue sample or a bodily fluid. In some examples, the sample is a tissue sample or an organ sample, such as a biopsy. In some examples, the sample includes cancer cells. In some examples, the sample includes cancer cells and normal cells. In some examples, the sample is a tumor biopsy. The bodily fluid may be sweat, saliva, tears, urine, blood, menstrual blood, semen, and / or spinal fluid. In some examples, the sample is a blood sample. The sample may include one or more peripheral blood lymphocytes. The sample may be a whole blood sample. The blood sample may be a peripheral blood sample. In some examples, the sample includes peripheral blood mononuclear cells (PBMCs), and in some examples, the sample includes peripheral blood lymphocytes (PBLs). The sample may be a serum sample. The sample may be a plasma sample.

[0055] The sample may be obtained using any method that can provide a sample suitable for the analytical methods described herein. The sample may be collected by non-invasive methods such as throat swabs, buccal mucosa swabs, bronchial lavage, urine collection, skin or cervical scraping, buccal swabs, saliva collection, fecal collection, menstrual blood collection, or semen collection. The sample may be collected by minimally invasive methods such as blood collection. The sample may be collected by venipuncture. In other examples, the sample is collected by invasive procedures including, but not limited to, biopsy, alveolar or pulmonary lavage, or needle aspiration. Biopsy methods include surgical biopsy, incisional biopsy, excision biopsy, punch biopsy, shaved biopsy, or skin biopsy. The sample may be a formalin-fixed section. Needle aspiration methods may further include fine needle biopsy, core needle biopsy, vacuum assisted biopsy, or large core biopsy. In some examples, multiple samples may be collected by the methods described herein to ensure a sufficient amount of biomaterial. In some instances, the sample is not obtained by biopsy.

[0056] B. Generation of sequence data In some embodiments, the sample is processed to obtain nucleic acid sequence data. A "nucleic acid" or "nucleic acid molecule" corresponds to a polymeric form of nucleotides of any length, either ribonucleotides, deoxyribonucleotides, or peptide nucleic acids (PNAs), containing purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The backbone of a polynucleotide contains sugar and phosphate groups typically found in RNA or DNA, or modified or substituted sugar or phosphate groups. A polynucleotide may contain modified nucleotides, such as methylated nucleotides and nucleotide analogs. The sequence of nucleotides may be interrupted by non-nucleotide components. Thus, the terms nucleoside, nucleotide, deoxynucleoside, and deoxynucleotide generally include analogs as described herein. These analogs are molecules that have structural features in common with natural nucleosides or nucleotides, and when incorporated into a nucleic acid or oligonucleoside sequence, are capable of hybridizing to natural nucleic acid sequences in solution. Typically, these analogs are derived from natural nucleosides and nucleotides by substituting and / or modifying the base, ribose, or phosphodiester moiety. These modifications can be customized to stabilize or destabilize hybridization or, if desired, increase the specificity of hybridization with complementary nucleic acid sequences. The nucleic acid molecule can be a DNA molecule. The nucleic acid molecule can be an RNA molecule.

[0057] Sample processing includes processing the nucleic acid sample and then sequencing the nucleic acid sample. Part or all of the biological sample may be sequenced to provide nucleic acid sequence data, which may be stored or maintained in electronic, magnetic, or optical storage. The sequence information may be analyzed with the aid of a computer processor, and the analyzed sequence information may be stored in an electronic storage. The electronic storage may contain a pool or collection of sequence information generated from the nucleic acid sample and analyzed sequence information. In some embodiments, the biological sample is taken from a subject suffering from or suspected of having cancer.

[0058] In some embodiments, nucleic acid sequence data is generated from pure tumor samples and pure normal samples. Matched cell lines can be obtained from separate sources (e.g., American Type Culture Collection). Each matched pair can include a tumor cell line and a normal cell line from the same subject. The cell lines are cultured and expanded in vitro to obtain a suitable number of cells for DNA extraction. DNA is extracted, processed, and subjected to whole-exome or whole-genome sequencing. Sequence reads can be subjected to quality control processing (e.g., by FastQC) to provide FASTQ files.

[0059] In some examples, whole genome sequencing is used to generate nucleic acid sequence data. In some examples, whole genome sequencing is used to identify variants in humans. Whole genome sequencing includes shallow sequencing (1-5x coverage) or deep sequencing (30-100x coverage) across the entire or nearly entire genome. In some examples, sequencing includes sequencing a portion of the genome. For example, the portion of the genome includes at least about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1,000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 2000, 2200, 2400, 2500, 2600, 2750 ...00, 550, 600, 650, 700 The sequence of genome may be 00, 1900, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000 or more bases or base pairs. In some examples, genome may be sequenced over 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, 9 million, 10 million or more bases or base pairs. In some examples, genome may be sequenced over the entire exome (e.g., whole exome sequencing). In some examples, deep sequencing can include obtaining multiple reads over a portion of a genome. For example, obtaining multiple reads can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000 reads, or more than 10,000 reads over a portion of a genome.

[0060] Additionally or alternatively, the nucleic acid sequence data may be generated using panel-based sequencing. Panel-based sequencing may be used to simultaneously evaluate multiple potential genetic causes of a suspected disorder, disease, or phenotype. Gene panels may be used to identify target genomic regions (to whose sequences nucleic acid molecules are aligned). In some examples, the number of target genomic regions includes at least 50, 100, 200, 500, 1000, 1500, 2000 genomic regions. Panel-based sequencing may define the footprint of the target genomic regions rather than the number of target genomic regions. For example, the footprint may range from 175 Kb to 3 Gb. In other embodiments, panel-based sequencing may target a number of genomic regions known to be biologically significant, such as genomic regions associated with variants of cancer driver genes or tumor escape genes.

[0061] In some examples, generating nucleic acid sequence data includes detecting low allele fractions by deep sequencing. In some examples, deep sequencing is performed by next generation sequencing. In some examples, deep sequencing is performed while avoiding error-prone regions. In some examples, error-prone regions may include regions near sequence duplications, regions with unusually high or low GC ratios, regions near homopolymers, dinucleotides, and trinucleotides, and regions near other short repeats. In some examples, error-prone regions may include regions that lead to DNA sequence errors (e.g., polymerase slippage in homopolymer sequences).

[0062] In some examples, generating nucleic acid sequence data includes performing one or more sequencing reactions on one or more nucleic acid molecules in the sample. Certain embodiments may include performing 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more sequencing reactions on one or more nucleic acid molecules in the sample. The sequencing reactions may be performed simultaneously, sequentially, or combinations thereof. The sequencing reactions may include whole genome sequencing, exome sequencing, or targeted sequencing of smaller panels. The sequencing reactions may include Maxim-Gilbert sequencing, chain termination, or high throughput systems. Alternatively, or in addition, the sequencing reaction may include Helioscope™ single molecule sequencing, Nanopore DNA sequencing, Lynx Therapeutics' Massively Parallel Signature Sequencing (MPSS), 454 Pyrosequencing, single molecule real-time (RNAP) sequencing, Illumina (Solexa) sequencing, SOLiD sequencing, Ion Torrent™, ion semiconductor sequencing, single molecule SMRT™ sequencing, Polony sequencing, DNA nanoball sequencing, VisiGen Biotechnologies approaches, or combinations thereof.Alternatively or additionally, the sequencing reaction may include one or more sequencing platforms, including but not limited to single molecule real-time (SMRT™) technologies, such as Genome Analyzer IIx, HiSeq, MiSeq, and NovaSeq from Illumina, PacBio RS system from Pacific Biosciences (California), and Solexa Sequencer, True Single Molecule Sequencing (tSMS™) technologies, such as HeliScope™ sequencer from Helicos Inc. (Cambridge, MA). The sequencing reaction may also include electron microscopy or chemically sensitive field effect transistor (chemFET) arrays. In some embodiments of the present disclosure, the sequencing reaction includes capillary sequencing, next generation sequencing, Sanger sequencing, sequencing by synthesis, sequencing by ligation, sequencing by hybridization, single molecule sequencing, or combinations thereof. Sequencing by synthesis includes reversible terminator sequencing, processive single molecule sequencing, sequential flow sequencing, or a combination thereof. Sequential flow sequencing includes pyrosequencing, pH-mediated sequencing, semiconductor sequencing, or a combination thereof.

[0063] In some examples, generating the nucleic acid sequence data includes performing at least one long read sequencing reaction and at least one short read sequencing reaction. The long read sequencing reaction and / or the short read sequencing reaction may be performed on at least a portion of the subset of nucleic acid molecules. The long read sequencing reaction and / or the short read sequencing reaction may be performed on at least a portion of two or more subsets of nucleic acid molecules. Both long read and short read sequencing reactions may be performed on at least a portion of one or more subsets of nucleic acid molecules.

[0064] Sequencing of one or more nucleic acid molecules or a subset thereof may include sequencing of at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,500, 2,000, 2,500, 3,000, 3,500, 4,000, 4,500, 5,000, 5,500, 6,000, 6,500, 7,000, 7,500, 8,000, 9,000, 10,000, 11,000, 12,000, 13,000, 14,000, 15,000, 16,000, 17,000, 18,000, 19,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, 30,000, 31,000, 32,000, 33,000, 34,000, 35,000, 36,000, 37,000, 38,000, 39,000, 40,000, 41,000, 42,000, 43,000, 44,000, 45,00 In one embodiment, the sequence may include 00, 8,500, 9,000, 9,500, 10,000, 25,000, 50,000, 75,000, 100,000, 250,000, 500,000, 750,000, 10,000,000, 25,000,000, 50,000,000, 100,000,000, 250,000,000, 500,000,000, 750,000,000, 1,000,000,000 or more sequencing reads.

[0065] The sequencing reaction may include at least about 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1,000, 1,500, 2, In some embodiments, the sequencing may include sequencing of 000, 2,500, 3,000, 3,500, 4,000, 4,500, 5,000, 5,500, 6,000, 6,500, 7,000, 7,500, 8,000, 8,500, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, or more bases or base pairs. The sequencing reaction may include at least about 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1,000, 1,500, 2,000, 3,500, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 11,000, 12,000, 13,000, 14,000, 15,000, 16,000, 17,000, 18,000, 19,000, 20,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, 30,000, 31,000, 32,000, 33,000, 34,000, 35,000, 36,000, 37,000, 38,000, 39,000, 40,000, This may include sequencing of 0, 2,500, 3,000, 3,500, 4,000, 4,500, 5,000, 5,500, 6,000, 6500, 7,000, 7,500, 8,000, 8,500, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000 or more contiguous bases or base pairs.

[0066] In some examples, the sequencing technique generates at least 100 reads per run, at least 200 reads per run, at least 300 reads per run, at least 400 reads per run, at least 500 reads per run, at least 600 reads per run, at least 700 reads per run, at least 800 reads per run, at least 900 reads per run, at least 1,000 reads per run, at least 5,000 reads per run, at least 10,000 reads per run, at least 50,000 reads per run, at least 100,000 reads per run, at least 500,000 reads per run, or at least 1,000,000 reads per run. Alternatively, the sequencing technique generates at least 1,500,000 reads per run, at least 2,000,000 reads per run, at least 2,500,000 reads per run, at least 3,000,000 reads per run, at least 3,500,000 reads per run, at least 4,000,000 reads per run, at least 4,500,000 reads per run, or at least 5,000,000 reads per run.

[0067] In some examples, the sequencing technique generates at least about 30 base pairs, at least about 40 base pairs, at least about 50 base pairs, at least about 60 base pairs, at least about 70 base pairs, at least about 80 base pairs, at least about 90 base pairs, at least about 100 base pairs, at least about 110, at least about 120 base pairs per read, at least about 150 base pairs, at least about 200 base pairs, at least about 250 base pairs, at least about 300 base pairs, at least about 350 base pairs, at least about 400 base pairs, at least about 450 base pairs, at least about 500 base pairs, at least about 550 base pairs, at least about 600 base pairs, at least about 700 base pairs, at least about 800 base pairs, at least about 900 base pairs, or at least about 1,000 base pairs per read. Additionally or alternatively, the sequencing technique can generate long sequencing reads. In some examples, the sequencing technique provides a sequencing readout of at least about 1,200 base pairs / read, at least about 1,500 base pairs / read, at least about 1,800 base pairs / read, at least about 2,000 base pairs / read, at least about 2,500 base pairs / read, at least about 3,000 base pairs / read, at least about 3,500 base pairs / read, at least about 4,000 base pairs / read, at least about 4,500 base pairs / read, at least about 5,000 base pairs / read, at least about 6,000 base pairs / read, at least about 7,000 base pairs / read, at least about 8,000 base pairs / read, at least about 9,000 base pairs / read, at least about 10,000 base pairs / read, at least about 11,000 base pairs / read, at least about 12,000 base pairs / read, at least about 13,000 base pairs / read, at least about 14,000 base pairs / read, at least about 15,000 base pairs / read, at least about 16,000 base pairs / read, at least about 17,000 base pairs / read, at least about 18,000 base pairs / read, at least about 19,000 base pairs / read, at least about 21,000 base pairs / read, at least about 22,000 base pairs / read, at least about 23,000 base pairs / read, at least about 24,000 base pairs / read, at least about 25,000 base pairs / read, at least about 26,000 base pairs / read, at least about 27,000 base pairs / read, at least about 28,000 base pairs / read, at least about 29,000 base pairs / read, at least about 30,000 base pairs / read, at least In some embodiments, the nucleic acid sequence may be generated in a number of ways, such as at least about 8,000 base pairs / read, at least about 9,000 base pairs / read, at least about 10,000 base pairs / read, at least about 20,000 base pairs / read, at least about 30,000 base pairs / read, at least about 40,000 base pairs / read, at least about 50,000 base pairs / read, at least about 60,000 base pairs / read, at least about 70,000 base pairs / read, at least about 80,000 base pairs / read, at least about 90,000 base pairs / read, or at least about 100,000 base pairs / read.

[0068] High-throughput sequencing system can detect sequenced nucleotides immediately after or at the time of incorporation into growing chain, i.e., detect sequences in real time or substantially real time.In some examples, high-throughput sequencing generates at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, or at least 500,000 sequence reads per hour, each read being at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 bases / read. Sequencing can be performed using a nucleic acid described herein, eg, genomic DNA, mtDNA, cDNA derived from an RNA transcript, or RNA, as a template.

[0069] III. Input data for determining latent variables A fragment mix signature can be defined to include one or more latent variables. Each latent variable can identify the estimated size and / or terminal motif distribution of variably enriched sequences across a set of genomic regions. The latent variables can represent novel independent basis vectors for sequence size and / or terminal motif data, which are then used to determine the enrichment of the fragment mix signature of interest.

[0070] A. Array size value 3 shows a schematic 300 illustrating an exemplary technique for generating sequence size values, according to some embodiments. The schematic 300 shows sequence data 305 aligned to a reference sequence 310. In some examples, the reference sequence 310 is a human reference genome (e.g., the hg19 genome). The reference sequence 310 may be divided into a set of genomic regions that includes genomic region 315. In the example shown in FIG. 3, the sequence set of the sequence data 305 may be aligned to the genomic region 315. A size distribution of the sequence set for the genomic region 315 may be identified.

[0071] A set of multiple sequence size values ​​320 can be generated based on the sequence data 305. Each set of sequence size values ​​can include, for each genomic region, a sequence size value that corresponds to a size of a corresponding sequence set. The set of sequence size values ​​can correspond to a set of integers in a box. In some examples, a set of sequence size values ​​spanning the entire genome is used. The set of sequence size values ​​can be determined from sequences aligned to a set of genomic regions (e.g., genomic regions in which somatic variants are known to be associated with cancer).

[0072] In some examples, each sequence represented by a corresponding sequence size value contains DNA fragments with sizes ranging from 50 bp to 550 bp. Sequence size values ​​outside the range of 50 to 550 bp (e.g., template fragment length values ​​in SAM / BAM files) were not used in the analysis.

[0073] In addition to or instead of a set of sequence size values, each set of sequence size values ​​can be used to generate a PMF 325. A PMF can refer to a function that gives the probability that a discrete random variable (e.g., a sequence size value) is equal to some value (e.g., 167 bp) within a particular range.

[0074] In some embodiments, a PMF325 was generated for each set of sequence size values, and each sequence size value in the set was binned at a resolution of 1 bp and then normalized by the total number of read pairs. This set of PMFs was used as input data (e.g., a matrix with N x 501 dimensions) for the BSS algorithm to identify one or more components that predict a particular disease.

[0075] B. Terminal motif A terminal motif can be identified as a terminal sequence of nucleotides of a DNA fragment, for example, a sequence of K bases at either end of the fragment. The terminal sequence can be a k-mer with a variable number of bases, such as 1, 2, 3, 4, 5, 6, 7, etc. In some examples, a terminal motif is determined by aligning sequence reads to a reference genome and identifying a nucleotide base immediately before the start position or immediately after the end position. Such bases correspond to the ends of a DNA fragment, for example, because they are identified based on the terminal sequence of the fragment.

[0076] Terminal motifs may be identified from aligned sequence reads using various techniques. In some examples, k-mer terminal motifs are constructed directly from the first k-bp of sequence at both ends of the plasma DNA molecule. For example, the first four nucleotides or the last four nucleotides of the sequenced fragment may be used. In another example, a (k-2)-mer sequence from the sequenced end of a fragment and another (k-2)-mer sequence from the genomic region adjacent to the end of the fragment are utilized to construct a k-mer terminal motif jointly. Terminal motifs of various lengths may be used, for example, 1-mer, 2-mer, 3-mer, 5-mer, 6-mer, 7-mer terminal motifs. Sequence reads may be aligned to the terminal motifs, thereby generating a set of terminal motif sequence data. The set of terminal motif sequence data may be applied to one or more signal separation algorithms to generate a corresponding set of latent variables.

[0077] The more nucleotides contained in the cell-free DNA end signature, the more specific the motif. For example, the probability of six bases being sequenced in the correct order in a genome is lower than the probability of two bases being sequenced in the correct order in a genome. Thus, the selection of the length of the end motif may depend on the sensitivity and / or specificity required for the intended use.

[0078] IV. Determining latent variables of fragment mix signatures using signal separation algorithms To determine the latent variables of the fragment mix signature, one or more signal separation algorithms can be applied to the plurality of sets of sequence size values ​​and / or terminal motif sequence data. In some examples, each latent variable of the latent variable set includes a histogram or weight vector that represents the size distribution underlying a proportion of the sequence data of the biological sample. The one or more signal separation algorithms can be blind source separation algorithms, which can use an independent component analysis algorithm and / or a non-negative matrix factorization algorithm.

[0079] A. Signal Separation Algorithm Signal separation algorithms can be constructed and used to estimate a set of source signals from a set of observed signals, where each observed signal is a mixture of source signals. Figure 4 shows an exemplary diagram 400 illustrating a signal separation algorithm applied to a set of image mixtures created from linearly added image source signals. In Figure 4, an input image 405 was generated based on a mixture of a set of image signals 410. Each image signal S i , the random coefficient k i To generate a diversity of source image amplitudes within the image blend compendium, at least one image signal may include a random noise signal that is unique for each image blend. In FIG. 4, 2000 input images 405 constructed based on a blend of four image signal sets 410 may be expressed as:

number

[0080] The input images 405 can be processed using one or more signal separation algorithms to output an image set 415 that visually approximates the image represented by the image signal set 410. The signal separation algorithms can be blind source separation algorithms. Blind source separation algorithms are called "blind" because they can estimate source signals without directly observing the source signals, but by observing their effects on measurable fundamental quantities.

[0081] The blind source separation algorithm may include a singular value decomposition (SVD) algorithm 420 or a principal component analysis (PCA) achieved by an ICA algorithm 430 (e.g., fastICA, InfoMax). Both the SVD algorithm 420 and the ICA algorithm 430 may be trained on a data set and the output may be applied to a new data set. As used herein, the SVD algorithm 420 may refer to a technique that identifies a small number of uncorrelated variables, known as principal components, from a larger data set. This technique is widely used to highlight changes and capture strong patterns within a data set (e.g., a set of pixels corresponding to an input image 405). In FIG. 4, an exemplary set of unmixed images 425 is shown as the output generated by applying the SVD algorithm 420 to the entire set of 2,000 image mixtures 405. In some examples, the input images 405 are preprocessed before applying the SVD algorithm 420.

[0082] The ICA algorithm 430 may include a fixed-point iterative scheme to find the maximum non-Gaussianity, and the ICA algorithm may then be run iteratively until the maximum non-Gaussianity value is reached. In some examples, the input images 405 are pre-processed before applying the ICA algorithm 430, including zero centroiding and scaling to unit variance. As shown in FIG. 4, exemplary unmixed image sets 435 and 440 may be shown as outputs generated by applying the ICA algorithm 430 to different amounts of the input images 405. The visual characteristics of the first unmixed image set 435 do not appear to match those of the image signal set 410, but this is due to the use of a small number (10) of image mixtures (I 1 -I 10 In contrast, the visual characteristics of the second image set 440 appear to match those of the respective image signal sets, which may be due to the use of a large number (2000) of image mixtures (I 1 -I 2000 ) is used. Furthermore, although not shown in FIG. 4, the blind source separation algorithm may include an NMF algorithm.

[0083] B. Exemplary scheme for determining latent variables of fragment mix signatures FIG. 5 shows a schematic diagram 500 illustrating an example technique for generating a set of latent variables, according to some embodiments. As shown in FIG. 5, each set of the multiple sets of array size values ​​505 can be used as an input for a BSS algorithm 510. The BSS algorithm 510 can include an ICA algorithm 515 or an NMF algorithm 520. Depending on the type of BSS algorithm, additional formats of the multiple sets of array size values ​​505 can be performed. For example, the NMF algorithm 520 can use raw PMFs corresponding to the set of array size values. In another example, the ICA algorithm 515 can use PMFs that are first mean-centroided and then scaled to unit variance. In some examples, the BSS algorithm 510 is performed multiple times until the stochastic initial conditions converge (e.g., a solution is reached, a minimization or maximization criterion is met, or the non-Gaussianity of the ICA is maximized). The BSS algorithm was performed repeatedly to ensure that properly estimated latent variables were obtained from the reproducibility across multiple runs.

[0084] Although the exemplary technique of FIG. 5 describes generating a set of latent variables from a set of sequence size values, the BSS algorithm 510 can also be applied to a set of terminal motif sequence data to generate another set of latent variables. In this case, the other set of latent variables identifies a terminal motif distribution of the sequence data. For example, the terminal motif distribution may include, for each nucleotide k-mer terminal motif (e.g., CCCA, TAAA), the number or relative frequency of occurrence of DNA fragments having terminal sequences corresponding to the k-mer terminal motif. By separating the terminal motif distribution of the sequence data into a set of independent weight vectors, a fragment mix signature can be determined.

[0085] Additionally or alternatively, the BSS algorithm 510 can be applied to generate a set of latent variables, where each latent variable can identify a probable size distribution of sequence reads having end sequences corresponding to a respective end motif.

[0086] The input to the ICA algorithm 515 is a data matrix X, and the output includes (i) an S matrix containing histogram or signed weight vector data corresponding to the set of latent variables, and (ii) an A linear mixing matrix containing estimated amplitudes of the latent variables across the set of genomic regions. Similarly, the input to the NMF algorithm 520 is a data matrix Y, and the output includes (i) a W matrix containing non-negative histogram data corresponding to the set of latent variables, and (ii) an H matrix containing estimated amplitudes of the latent variables across the set of genomic regions. These data matrices can be used to generate a set of histograms or weight vectors, which can be represented as corresponding fragment mix signature latent variables 525.

[0087] Continuing with the example shown in Figure 5, the latent variables 525 correspond to a weight vector or histogram generated by applying the BSS algorithm 510 to the set of sequence size values ​​505. The peaks and amplitudes identified in each of the latent variables 525 can reveal the size of the nucleic acid molecules of the biological sample and / or the sequence fragmentation patterns of the sequences within the sequence data. The latent variables 525 can represent a fragment mix signature of the subject.

[0088] V. Using latent variables of fragmentomics signatures to detect cancer-associated gene mutations Many cancers are caused by recurrent genetic mutations, such as BRCA and HER2 mutations in breast cancer. Accurate detection of these cancer-associated genetic mutations is therefore key to targeted therapy for cancer patients. However, noise, such as PCR or sequencing errors, can occur at various steps of an NGS experiment. These noises are particularly problematic in cfDNA samples where the allele frequency of the actual mutation is low. Previous studies have shown that cfDNA molecules from cancer tissues have different fragment size distributions compared to cfDNA molecules from healthy tissues. By examining the fragmentomics signatures of reads supporting the mutation, it is possible to confirm whether the observed mutation is cancer-associated or noise-associated.

[0089] A. Generating fragment mix signatures from a set of cfDNA whole exome sequencing (WES) samples 6 shows a schematic diagram 600 illustrating an example technique of using an independent component analysis algorithm to generate a set of latent variables for a reference signature, according to some embodiments. As shown in FIG. 6, a set 605 of cfDNA samples (n=10) can be collected. The set 605 of cfDNA samples can include four colorectal cancer cfDNA samples identified with the prefix "CRC" and six healthy donor samples identified with the prefix "PON". Whole exome sequencing (WES) can be used to sequence the set 605 of biological samples to generate sequence data from which a set of sequence size values ​​can be generated.

[0090] The sequence data can be divided into 309 uniform intervals and a fragment size distribution for each region can be generated. An ICA algorithm can then be applied to the 309×541 (60-600 bp) matrix for each sample 610 to generate an S-matrix corresponding to a set of latent variables. To select the most relevant latent variables from such a set, a clustering algorithm 615 (e.g., a k-means clustering algorithm) can be applied to the set of latent variables for all samples, and the latent variables in the set can be clustered based on their similarity. For each cluster, a centroid of the latent variables can be generated. For each subject, the latent variable most similar to the cluster centroid can be used to detect cancer mutations. For example, latent variables 620 similar to the cluster centroid can be selected from healthy cfDNA samples. In another example, latent variables 625 similar to the cluster centroid can be selected from cfDNA samples associated with colorectal cancer.

[0091] B. Comparison of latent variables between healthy donors and cancer patients Returning to Figure 6, latent variables 620 of nucleic acid molecules of a normal biological sample are compared to latent variables 625 of nucleic acid molecules of a biological sample containing tumor DNA (e.g., TP53 tumor variants). The difference between latent variables 620 and 625 can be used to predict mutations in cancer-associated genes.

[0092] Continuing with the example of FIG. 6, the comparison may include comparing the sequence size values ​​of the peaks represented in each of the latent variable sets 620 with the sequence size values ​​of the peaks represented in the corresponding latent variables of the latent variable set 625. In one example, the sequence size values ​​of the peaks represented in the latent variable set 625 (colorectal cancer samples) are approximately 10-30 base pairs smaller than the values ​​represented in the latent variable set 620 (healthy samples). Thus, the abundance of relatively short DNA molecules may predict whether a subject is afflicted with cancer. In some examples, the difference between the sequence size values ​​of the peaks of the latent variables of the cancer samples and the sequence size values ​​of the peaks of the corresponding latent variables of the healthy samples is calculated to predict a particular stage of the disease. For example, if the difference indicates a relatively small value (e.g., less than 10 bp), it may be predicted that the cancer is in an early stage. In contrast, if the difference indicates a relatively large value (e.g., more than 30 bp), it may be predicted that the cancer is in an advanced stage.

[0093] Additionally or alternatively, the comparison may include determining, for each latent variable, a density of sequence size values ​​within a predefined size range. The density of sequence size values ​​can be compared to a predefined threshold to predict whether the subject is afflicted with a particular disease. For example, the density of sequence size values ​​of latent variables within a size range of 200-300 bp can be calculated and the density compared to a threshold (e.g., 2) to predict whether the subject is afflicted with cancer.

[0094] Additionally or alternatively, comparing the latent variables may include comparing the amplitude value of each latent variable in latent variable set 620 with the amplitude value of a corresponding latent variable in latent variable set 625. In some examples, the difference between each pair of amplitude values ​​is compared to a threshold value.

[0095] C. Examples of fragment mix signatures for individual samples. 7-10 show an exemplary set of fragment mix signatures for individual samples, according to some embodiments. Each figure highlights latent variables that are repeatedly identified across samples and across multiple runs of the BSS algorithm. FIG. 7 shows a set of latent variables 700 generated by applying a signal separation algorithm to sequence size values ​​corresponding to biological samples of normal subjects. A subset of latent variables 705 can be selected from the set of latent variables 700, where the subset of latent variables 705 identifies a size distribution of sequences representative of the fragment mix signature of the corresponding biological sample. In some examples, the latent variables of the subset 705 are selected based on whether the latent variables include peaks having peak width values ​​greater than a particular threshold. Additionally or alternatively, the latent variables of the subset 705 can be selected by applying a clustering algorithm to the latent variables to generate a set of clusters, and selecting a latent variable that corresponds to the centroid of a particular cluster of the set.

[0096] 8 shows a set of latent variables 800 generated by applying a signal separation algorithm to sequence size values ​​corresponding to a biological sample of another normal subject. A subset of latent variables 805 can be selected from the set of latent variables 800, where the subset of latent variables 805 identifies a size distribution of sequences representative of the fragment mix signature of the corresponding biological sample. As shown in FIG. 8, the subset of latent variables 805 includes a subset of latent variables that share similarities with the latent variables of the fragment mix signature 705 of FIG. 7.

[0097] FIG. 9 shows a set of latent variables 900 generated by applying a signal separation algorithm to sequence size values ​​corresponding to a biological sample of a subject diagnosed with colorectal cancer. A subset of latent variables 905 can be selected from the set of latent variables 900, where the subset of latent variables 905 identifies a size distribution of sequences representative of a fragment mix signature of the corresponding biological sample. FIG. 10 shows a set of latent variables 1000 generated by applying a signal separation algorithm to sequence size values ​​corresponding to a biological sample of another subject diagnosed with colorectal cancer. A subset of latent variables 1005 can be selected from the set of latent variables 1000, where the subset of latent variables 1005 identifies a size distribution of sequences representative of a fragment mix signature of the corresponding biological sample.

[0098] D. Biological processes that affect the latent variables In some examples, the shapes of the latent variables can be used to predict the enrichment of protein molecule sets (e.g., mononucleosomes, dinucleosomes, monochromatosomes, dichromatosomes, and transcription factor complexes) in cell-free DNA sequence size data and to predict the binding of said molecule sets at corresponding cellular loci in vivo based on utilizing the latent variables as statistical objects representing independently regulated molecules that govern chromatin epigenetic states. In some examples, the latent variables are also used to predict potential novel nucleic acid and protein entities and their corresponding structures based on fragment length patterns.

[0099] For example, the enrichment of latent variables known to be associated with chromosomes in the sequence data can be used to predict the enrichment of chromatosome binding and silenced heterochromatin states at corresponding cellular loci. In another example, the enrichment of cancer-related latent variables in the sequence data can be used to predict whether a particular allele is derived from a cancer cell. In yet another example, the enrichment of gene expression-related latent variables in the sequence data can be used to predict whether a particular allele contributing to the sequence data is expressed. Additionally or alternatively, the relative abundance of cell-free DNA of shorter and longer size types in a molecule set can be predicted by comparing bimodal latent variable density peaks and inferring the rate of cell-free DNA degradation from the rate of change in the latent variable space.

[0100] FIG. 11 shows a schematic 1100 illustrating a technique for predicting enrichment of a set of protein molecules in sequence size data, according to some embodiments.

[0101] The latent variable set 1105 may include latent variables 1110a-e, each latent variable providing an estimated size distribution of nucleic acid molecules in a biological sample. For each latent variable, a set of peaks may be identified. Each peak in the peak set may be used to predict the binding of a particular protein. For example, latent variable 1110a may be represented as a unimodal size distribution with peak 1115a at a sequence size of approximately 150 bp. From such a size distribution peak, the location of nucleosome binding may be predicted if there is an abundance of sequence size data associated with genomic locations. In addition to peak 1115a, a periodic pattern of peaks with a peak width of approximately 10 bp may predict nucleosome binding by considering the effect of additional DNA degradation signals, possibly correlated with progressive digestion of individual fragments associated with the DNA helical pitch bound to nucleosomes. These peaks are believed to correlate with the behavior of nucleic acid molecules during apoptotic or necrotic cell death and the resulting release into and transit through the circulation. For example, protein-bound DNA molecules, usually associated with histones or transcription factors, selectively survive damage (eg, digestion) and are released into the circulation, while unbound DNA molecules are lost.

[0102] Continuing with the example of FIG. 11, some latent variables can predict the binding of specific proteins. For example, latent variable 1110e includes peaks 1120a-b that are represented together as a multimodal distribution, and each of peaks 1120a-b can be used to predict the binding of transcription factors and transcription factor-dinucleosome complexes when enriched in sequence size data associated with transcribed loci. Additionally, latent variable 1110e includes a third peak 1120c that has a larger peak width value. Such a peak can be used to predict the binding of transcription factors and mononucleosomes. In addition to the width of each peak 1220a-b, the roughly 10 bp superimposed periodic patterns that appear to fingerprint CTCF (e.g., as seen in ATAC-seq data) can predict cell-to-cell heterogeneity of DNA-binding proteins that control the epigenetic state of transcribed loci.

[0103] The techniques described in Figure 11 can be applied to different latent variable sets 1125. The predicted protein binding can be used to define genome-wide nucleosome footprint patterns, which can then be used to predict whether a subject is afflicted with cancer.

[0104] E. Using cancer-derived latent variables to identify cancer-associated genetic mutations The latent variables of the fragmentomics signature can be used as features to predict whether nucleic acid molecules in biological samples harbor cancer-associated mutations. For example, the TP53 gene encodes a tumor suppressor protein that controls cell division. It is one of the most frequently mutated genes in human cancers. In a CRC10063 cfDNA sample, the TP53_chr17_7578265_A>C mutation was identified. The same mutation was also found in the paired tumor sample and not in the corresponding adjacent healthy tissue or white blood cell normal samples, indicating that the reads supporting the mutation from the cfDNA are of tumor cell origin. The sequence size distribution of the reads supporting these TP53 mutations can be projected onto a fixed latent variable, generating 1 × 5 vectors that represent the amplitude or enrichment of each latent variable in the sequence size data. As a negative control, the same ICA projection was performed using reads supporting the high-quality wild-type allele. These two vectors can provide a way to quantify the difference in fragment size.

[0105] FIG. 12 shows a schematic diagram 1200 illustrating a process of preprocessing raw sequence size distributions by projecting them onto latent variables, converting the fragment size distributions into a set of amplitudes, and using the amplitudes to detect cancer-associated genetic mutations, according to some embodiments. As an initial step, BSS can be applied to sequence data obtained from cancer reference samples to select latent variables that are repeatedly identified across different reference samples. In step 1205, after determining candidate mutations in cfDNA samples, reads supporting the mutation (ALT allele) and reads supporting the wild-type allele (REF allele) can be separated. In step 1210, the sequence size distributions of the ALT allele and the REF allele can be converted into their respective PMFs and projected onto a fixed latent variable. The resulting amplitudes of the latent variables can be used to measure the quality of the mutation determination. For example, the amplitudes of the latent variables corresponding to the mutation-supporting reads from tumor tissue should be different from the amplitudes of the latent variables corresponding to the wild-type-supporting reads, while the amplitudes of the mutant reads originating from experimental noise or error should be closer to the amplitudes of the wild-type reads.

[0106] It has been shown that latent variables can be highly enriched in the sequence size distribution of circulating tumor DNA (ctDNA) in plasma samples. For example, a latent variable includes a weight vector generated by applying the ICA algorithm to sequence size data from advanced stage colorectal cancer patients. The aforementioned plasma cfDNA sequence size data associated with TP53 mutations, known to be derived from matched tumor biopsy samples, shows a narrow peak in the 140-150bp range and a broader peak in the 200-300bp range, which is very similar to the latent variable "LV5".

[0107] To develop a model for filtering candidate mutation calls, sequence size distributions corresponding to 78 loci with high-quality mutations were selected and projected onto a fixed latent variable (see step 1210). At the same time, the corresponding high-quality wild-type reads were projected using the same method. The LV1 and LV5 amplitudes for each locus were plotted. The LV amplitude plot 1215 shows that the result set corresponding to the REF allele is clustered in a first region 1220, while the result set corresponding to the ALT allele is mostly away from that cluster. In some examples, one or more cutoffs are determined based on the plot results shown in 1215. For example, the cutoff 1225 is approximately 0.4. Results above the cutoff 1225 may be considered to have somatic mutations specific to tumor DNA. As an example, the results 1230 identified from the LV amplitude plot 1215 identify a size distribution of nucleic acid molecules with known tumor-associated variants (TP53 gene variants).

[0108] VI. Cancer Prediction Using Latent Variables of Fragment Mix Signatures The latent variables represented by the fragment mix signature can also be used for early detection of cancer in a subject and monitoring the recurrence of cancer after treatment. Cancer patients usually have different size distribution of cfDNA fragments and frequency of occurrence of terminal motifs compared to healthy control subjects. Size distribution data of a particular subject can be projected onto a set of latent variables of the fragment mix signature to generate a fragment mix signature amplitude, and the fragment mix signature amplitude can then be used (e.g., via machine learning) to predict whether a particular subject is affected by a disease (e.g., cancer). For example, the fragment mix signature amplitude can be compared to the fragment mix signature amplitude of a healthy donor to predict whether the subject has an abnormality. Other types of predictions can be considered based on the type of latent variable and the type of training data used for the machine learning model. If the training data includes reference samples of patients at a particular stage of the disease, a latent variable-based model can be used to predict whether a particular subject is at a particular stage of the disease (e.g., stage IV cancer). If the training data includes reference samples with a particular type of disease, the machine learning model can be used to predict whether a subject has a particular type of disease (e.g., colorectal cancer).

[0109] A. Generation of latent variables from a set of cfDNA samples enriched by a tumor informatics panel cfDNA analysis is an emerging method for detecting minimal residual disease or cancer recurrence after treatment. However, its sensitivity is limited due to low tumor signal in plasma. In the following description, we show an example where tumor-derived DNA can be enriched by personalized hybridization capture and tumor-enriched samples can be used to derive latent variables for detecting cancer of interest.

[0110] FIG. 13 shows a schematic diagram illustrating an exemplary technique for generating a set of latent variables using an independent component analysis algorithm on hybridization capture samples according to some embodiments. In FIG. 13, triplicates of normal / tumor / plasma samples of various tumor types were used. Whole genome sequencing was performed using normal and tumor samples to detect somatic mutations in tumors. In step 1305, about 1800 somatic mutations with high frequency alleles were used to design individualized hybridization capture panels for each patient. This tumor information panel can be used to enrich tumor signals from plasma. For each panel design, hybridization capture was also performed using one or two healthy plasma samples as references.

[0111] To generate latent variables, sequence reads were mapped to the reference genome, keeping all reads within 50-550 bp. In step 1310, fragment size PMFs around each target region were generated for each plasma sample. In step 1315, the 1800x501 matrix identifying the sequence size distribution of fragments was used as input to a BSS algorithm such as ICA. In this example, ICA was performed for each plasma sample, and 10 latent variables (independent components) were retained.

[0112] After collecting latent variables from 119 healthy plasma samples and 101 patient plasma samples, we used K-means to identify similar latent variables across different samples (step 1320). We ran a clustering algorithm using different parameters and selected an 8-cluster model according to several clustering evaluation methods. The results 1325 show the centroids of each cluster, which are used as latent variables for downstream analyses (e.g., analyses to predict cancer in other subjects).

[0113] B. Detecting cancer onset or recurrence using latent variable amplitudes The following description provides an example of using latent variables to monitor cancer recurrence after surgery. FIG. 14 shows a schematic diagram illustrating an exemplary technique for monitoring cancer recurrence using handcrafted or latent variable features according to some embodiments. Hybridization capture panel design was performed according to somatic mutations of tumor samples. In step 1405, plasma samples were collected at various time points after surgery, and then hybridization capture was used to enrich the tumor samples for target cfDNA. Hybridization capture was performed similarly using healthy plasma samples as a reference. In step 1410, fragment size distributions can be extracted from both patient and healthy samples and projected onto eight fixed latent variables to monitor the possibility of cancer recurrence at various time points. In step 1415, a pre-trained machine learning model can be used to compare the patient's latent variable amplitudes with those of the controls and assign status for different time points. In step 1420, to compare the latent variable approach with other methods, we also trained a baseline model using 13 handcrafted features, including the cumulative probability of reads within 50-250 bp, the probability of reads within 160-170 bp, and the ratio of reads between 250-300 bp and 300-350 bp.

[0114] Regarding the machine learning models, two logistic regression models based on handcrafted (step 1420) or latent variable features (step 1415) were trained using a dataset of 43 healthy-control pairs and 59 healthy-patient pairs. Figure 15 shows a set of receiver operating characteristic (ROC) curves 1500 illustrating the accuracy level when using fragment mix signatures for disease classification. In Figure 15, graphs 1505 and 1510 show the receiver operating characteristic curves that distinguish the performance of the two models. The latent variable feature-based model (graph 1510) outperforms the handcrafted feature-based model (graph 1505), with the area under the training curve increasing from 95.8% to 97% and the area under the cross-validation curve (average) increasing from 93.8% to 97%. Thus, the BSS algorithm can be used to effectively extract unbiased data-driven latent variable features. Furthermore, latent variables can be used to obtain better performance than handcrafted features in cancer detection.

[0115] VII. Methods for predicting disease classification of a subject based on fragment mix signatures determined from terminal motif frequency FIG. 16 includes a flowchart 1600 illustrating an example of a method for determining a fragment mix signature of a subject based on terminal motif occurrence frequency of a nucleic acid molecule of a biological sample according to some embodiments. Some of the operations described in the flowchart 1600 may be performed by a computer system (e.g., computer system 1700 of FIG. 17). Although the flowchart 1600 may describe the operations as a sequential process, in various embodiments, many of the operations may be performed in parallel or simultaneously. Furthermore, the order of the operations may be rearranged. The operations may have additional steps not shown in the figure. Furthermore, some embodiments of the method may be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the associated tasks may be stored in a computer-readable medium, such as a storage medium.

[0116] In step 1602, sequence data for the subject's biological sample can be accessed. In some examples, the sequence data corresponds to a plurality of cell-free DNA molecules for the biological sample, the plurality of cell-free DNA molecules including circulating tumor DNA molecules. The sequence data can also include sequences corresponding to a plurality of somatic variants detected from the biological sample. Additionally, each cell-free DNA molecule can include a corresponding terminal motif. A terminal motif can be identified as a terminal sequence of nucleotides for a DNA fragment, e.g., a sequence of K bases at either end of the fragment. The terminal sequence can be a k-mer having a variable number of bases, such as 1, 2, 3, 4, 5, 6, 7, etc. In some examples, the terminal motif is determined by aligning sequence reads to a reference genome and identifying a nucleotide base immediately preceding the start position or immediately following the end position. Such bases correspond to the ends of the DNA fragment, e.g., because they are identified based on the terminal sequence of the fragment. The process of identifying terminal motifs for cell-free DNA molecules is further detailed in section III.B of this disclosure.

[0117] In step 1604, a set of terminal motif sequence data can be generated based on the sequence data. In some examples, each terminal motif sequence data of the set identifies the number or relative frequency of nucleic acid molecules having a terminal sequence corresponding to a particular terminal motif. For example, the set of terminal motif sequence data can include, for a particular tetramer terminal motif (e.g., CCCA), the number of sequence reads having that terminal motif. In some examples, the number or relative frequency of nucleic acid molecules having a particular terminal motif can be normalized (e.g., using the total number of sequence reads in the biological sample).

[0118] In step 1606, the set of terminal motif sequence data can be projected onto a latent variable of a fragment mix signature. The fragment mix signature can include one or more signatures of the distribution of terminal motif occurrence frequencies of nucleic acid molecules that can predict disease classification. The fragment mix signature can include one or more fixed latent variables. The fixed latent variables can be determined by (i) applying one or more signal separation algorithms to the terminal motif sequence data obtained from one or more reference samples to generate a set of latent variables, and (ii) applying a clustering algorithm to select a subset of the set of latent variables, which subset corresponds to the fixed latent variables. In some examples, the reference sample includes a biological sample (e.g., tissue, plasma sample) taken from a subject with a disease diagnosis (e.g., cancer). Additionally or alternatively, the reference sample can include biological samples obtained from the same subject at different time points. In some examples, each latent variable of the set of latent variables includes a histogram or weight vector representing the distribution of terminal motif occurrence frequencies of the reference samples. The one or more signal separation algorithms may be blind source separation algorithms, which may include an independent component analysis algorithm and / or a non-negative matrix factorization algorithm.

[0119] Additionally or alternatively, derivatives of the latent variables may be determined by applying transformations such as scaling, translation, averaging, frequency domain transformations (e.g., Fast Fourier Transforms), and other similar transformations.

[0120] Plasma DNA nucleases such as DFFB, DNASE1L3, and DNASE1 are involved in both the generation and removal of cfDNA. The cleavage preferences of these different nucleases may affect the frequency of occurrence of cfDNA terminal motifs. Studies have shown that the activity of plasma DNA nucleases can be modified by multiple diseases such as cancer and systemic lupus erythematosus. In some instances, the BSS algorithm has records that estimate the signals carried by physically distinct entities, so a set of latent variables can predict the activity of plasma DNA nucleases in different subjects.

[0121] Moreover, cfDNA terminal motif frequencies correspond to the accessibility of different genomic regions to DNA nucleases, and thus latent variables can predict the degree of cell-to-cell heterogeneity of bound DNA and, therefore, proteins enabling epigenetic states.

[0122] In step 1608, one or more fragment mix signature amplitudes of the biological sample can be determined based on projecting the set of terminal motif sequence data onto the set of latent variables. The fragment mix signature amplitudes can be determined by projecting the set of terminal motif sequence data (e.g., PMF) onto a subset of the set of latent variables of the reference sample. In some examples, a clustering algorithm is applied to the latent variables generated from the different reference samples to determine a subset of the set of latent variables. The subset of latent variables can include centroids or medoids selected from the clusters of the identified latent variables.

[0123] In step 1610, the fragment mix signature amplitudes can be used as inputs to a machine learning algorithm (e.g., a logistic regression classification model) to generate results. The results can predict a classification that predicts whether the subject is afflicted with a particular disease. The particular disease can include cancer. In some examples, the results can predict the progression or recurrence of a particular disease when reference samples are taken from the same subject at different times. The results can be used to identify a treatment for the subject and / or determine how often to administer a treatment to the subject. In another example, the results can predict the presence of chromatosome binding at an allele of interest and the presence of silencing of the allele based on an algorithm that determines significant enrichment of monochromatosome and dichromatosome sequences associated with the allele.

[0124] Thus, by leveraging terminal motif data generated from whole exome sequencing or specific genomic regions of interest (e.g., cfDNA mapped to each genomic region of a set of genomic regions can be modeled as a linear mixture of DNA molecules processed by different plasma DNA nucleases. The BSS algorithm is then applied to the set of mixtures to estimate the fragment mix signature of the subject.

[0125] In step 1614, the results may be output. For example, the results may be displayed locally or sent to another device. The results may be output along with the subject's identifier. Process 1600 then ends.

[0126] VIII. Diseases and Treatment Certain embodiments may include predicting, diagnosing, and / or prognosing a disease or condition state or outcome in a subject based on one or more biomedical outputs. Predicting, diagnosing, and / or prognosing a disease state or outcome in a subject may include diagnosing a disease or condition, predicting a disease or condition, predicting a stage of a disease or condition, assessing risk of a disease or condition, assessing risk of disease recurrence, assessing efficacy of a drug, assessing risk of adverse drug reactions, predicting optimal drug dosage, predicting drug resistance, or a combination thereof.

[0127] The sample disclosed herein may be from a pregnant woman. The sample may be a maternal plasma sample containing fetal nucleic acid molecules. The fetus may have chromosomal aneuploidy. Fetal aneuploidy may cause various diseases, such as Down's syndrome (trisomy 21), Patau's syndrome (trisomy 13), Edwards' syndrome (trisomy 18), etc. The fetus may have diseases caused by gene mutations or deletions, such as spinal muscular atrophy and DiGeorge syndrome.

[0128] The samples disclosed herein may be from a subject suffering from cancer. The samples may include malignant tissue, benign tissue, liquid biopsy, or a mixture thereof. The cancer may be recurrent and / or refractory cancer. Examples of cancer include, but are not limited to, sarcoma, carcinoma, lymphoma, or leukemia. In some examples, a sample containing cancer tissue is collected, but a matched normal sample is not collected. In some examples, a matched normal sample is not available. In some examples, a matched normal sample is collected (e.g., for training and testing the models disclosed herein).

[0129] Sarcomas are cancers of bone, cartilage, fat, muscle, blood vessels, or other connective or supporting tissues. Sarcomas include, but are not limited to, bone cancer, fibrosarcoma, chondrosarcoma, Ewing's sarcoma, malignant hemangioendothelioma, malignant schwannoma, bilateral vestibular schwannoma, osteosarcoma, soft tissue sarcomas (e.g., alveolar soft part sarcoma, angiosarcoma, phyllodes cystosarcoma, dermatofibrosarcoma, desmoid tumor, epithelioid sarcoma, extraskeletal osteosarcoma, fibrosarcoma, hemangiopericytoma, angiosarcoma, Kaposi's sarcoma, leiomyosarcoma, liposarcoma, lymphangiosarcoma, lymphosarcoma, malignant fibrous histiocytoma, neurofibrosarcoma, rhabdomyosarcoma, and synovial sarcoma).

[0130] Carcinomas are cancers that originate from epithelial cells that cover the surface of the body, produce hormones, and make up glands. Non-limiting examples of carcinomas include breast cancer, pancreatic cancer, lung cancer, colon cancer, colorectal cancer, rectal cancer, kidney cancer, bladder cancer, stomach cancer, prostate cancer, liver cancer, ovarian cancer, brain cancer, vaginal cancer, vulvar cancer, uterine cancer, oral cancer, penile cancer, testicular cancer, esophageal cancer, skin cancer, fallopian tube cancer, head and neck cancer, gastrointestinal stromal cancer, adenocarcinoma, skin or intraocular melanoma, anal region cancer, small intestine cancer, endocrine system cancer, thyroid cancer, parathyroid cancer, adrenal gland cancer, urethral cancer, renal pelvis cancer, ureter cancer, endometrial cancer, cervical cancer, pituitary cancer, central nervous system (CNS), primary CNS lymphoma, brain stem glioma, and spinal axis tumor. The cancer may be a skin cancer, such as basal cell carcinoma, squamous cell carcinoma, melanoma, non-melanoma, or actinic (solar) keratosis.

[0131] The cancer may be lung cancer. Lung cancer may occur in the airways that branch off the trachea that supply the lungs (bronchi) or the small air sacs of the lungs (alveoli). Lung cancer includes non-small cell lung cancer (NSCLC), small cell lung cancer, and mesothelioma. Examples of NSCLC include squamous cell carcinoma, adenocarcinoma, and large cell carcinoma. Mesothelioma may be a cancerous tumor of the lining of the lungs and chest cavity (pleura) or the abdominal lining (peritoneum). Mesothelioma may be caused by exposure to asbestos. The cancer may be a brain cancer, such as glioblastoma.

[0132] The cancer may be a central nervous system (CNS) tumor. CNS tumors may be classified as gliomas or non-gliomas. Gliomas may be malignant gliomas, high-grade gliomas, diffuse intrinsic pontine gliomas. Examples of gliomas include astrocytomas, oligodendrogliomas (or mixed oligodendrogliomas and astrocytomas elements), and ependymomas. Astrocytomas include, but are not limited to, low-grade astrocytomas, anaplastic astrocytomas, glioblastoma multiforme, pilocytic astrocytomas, multimorphic xanthoastrocytomas, and subependymal giant cell astrocytomas. Oligodendrogliomas include low-grade oligodendrogliomas (or oligoastrocytomas) and anaplastic oligodendrogliomas. Non-gliomas include meningiomas, pituitary adenomas, primary CNS lymphomas, and medulloblastomas. The cancer may be a meningioma.

[0133] The leukemia may be acute lymphocytic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, or chronic myelogenous leukemia. Additional types of leukemia include hairy cell leukemia, chronic myelomonocytic leukemia, and juvenile myelomonocytic leukemia.

[0134] Lymphoma is a cancer of lymphocytes and can arise from either B or T lymphocytes. The two main types of lymphoma are Hodgkin's lymphoma, formerly called Hodgkin's disease, and non-Hodgkin's lymphoma. Hodgkin's lymphoma is characterized by the presence of Reed-Sternberg cells. Non-Hodgkin's lymphoma is any lymphoma other than Hodgkin's lymphoma. Non-Hodgkin's lymphoma can be low-grade lymphoma and high-grade lymphoma. Non-Hodgkin's lymphomas include, but are not limited to, diffuse large B-cell lymphoma, follicular lymphoma, mucosa-associated lymphoid tissue lymphoma (MALT), small cell lymphocytic lymphoma, mantle cell lymphoma, Burkitt's lymphoma, mediastinal large B-cell lymphoma, Waldenstrom's macroglobulinemia, nodal marginal zone B-cell lymphoma (NMZL), splenic marginal zone lymphoma (SMZL), extranodal marginal zone B-cell lymphoma, intravascular large B-cell lymphoma, primary effusion lymphoma, and lymphomatoid granulomatosis.

[0135] Certain embodiments may include treating and / or preventing a disease or condition of a subject based on the one or more biomedical outputs. The one or more biomedical outputs may recommend one or more therapies. The one or more biomedical outputs may suggest, select, prescribe, recommend, or determine a course of treatment and / or prevention of the disease or condition. The one or more biomedical outputs may recommend a change or continuation of one or more therapies. The change in one or more therapies may include administration, initiation, reduction, increase, and / or cessation of one or more therapies. The one or more therapies include an anti-cancer therapy, an anti-viral therapy, an anti-bacterial therapy, an anti-fungal therapy, an immunosuppressive therapy, or a combination thereof. The one or more therapies may treat, alleviate, or prevent one or more diseases or indications.

[0136] Examples of anti-cancer therapy include, but are not limited to, surgery, chemotherapy, radiation therapy, immunotherapy / biological therapy, photodynamic therapy. Anti-cancer therapy can include chemotherapy, monoclonal antibodies (e.g., rituximab, trastuzumab), cancer vaccines (e.g., therapeutic vaccines, prophylactic vaccines), gene therapy, or a combination thereof.

[0137] IX. Computing Environment 17 illustrates an example of a computer system 1700 for implementing some embodiments disclosed herein. The computer system 1700 may include a distributed architecture, where some components (e.g., memory and processor) are part of an end-user device, and some other similar components (e.g., memory and processor) are part of a computer server. In some examples, the computer system 1700 is a computer system for determining a fragment mix signature based on a size distribution of nucleic acid molecules, and includes at least a processor 1702, a memory 1704, a storage device 1706, an input / output (I / O) peripheral 1708, a communication peripheral 1710, and an interface bus 1712. The interface bus 1712 is configured to communicate, transmit, and transfer data, control, and commands between various components of the computer system 1700. The processor 1702 may include one or more processing units, such as a CPU, a GPU, a TPU, a systolic array, or a SIMD processor. The memory 1704 and the storage devices 1706 include computer readable storage media, such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard drives, CD-ROM, optical storage devices, magnetic storage devices, electronic non-volatile computer storage devices, such as Flash memory, and other tangible storage media. Any such computer readable storage media can be configured to store instructions or program code embodying aspects of the present disclosure. The memory 1704 and the storage devices 1706 also include computer readable signal media.

[0138] A computer-readable signal medium includes a propagated data signal having computer-readable program code embodied therein. Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any combination thereof. A computer-readable signal medium includes any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transmit a program for use with computer system 1700.

[0139] Additionally, memory 1704 includes an operating system, programs, and applications. Processor 1702 is configured to execute stored instructions and includes, for example, logic processing units, microprocessors, digital signal processors, and other processors. For example, computing system 1700 may execute instructions (e.g., program code) that configure processor 1702 to perform one or more operations described herein. Program code may include, for example, code that implements analysis of sequence data and / or any other suitable application that performs one or more operations described herein. Instructions may include, for example, processor-specific instructions generated by a compiler or interpreter from code written in any suitable computer programming language, including C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, R, and ActionScript.

[0140] The program code may be stored in memory 1704 or any suitable computer readable medium and executed by processor 1702 or any other suitable processor. In some embodiments, all of the modules in the computer system for performing the various functions and processes described herein are stored in memory 1704. In additional or alternative embodiments, one or more of these modules of the computer system are stored in different memory devices in different computing systems.

[0141] The memory 1704 and / or the processor 1702 can be virtualized and hosted in another computing system, for example, in a cloud network or a data center. The I / O peripherals 1708 include user interfaces, such as keyboards, screens (e.g., touch screens), microphones, speakers, other input / output devices, as well as computing components, such as image processing units, serial ports, parallel ports, universal serial buses, and other input / output peripherals. The I / O peripherals 1708 connect to the processor 1702 through any of the ports coupled to the interface bus 1712. The communication peripherals 1710 are configured to facilitate communication between the computer system 1700 and other computing devices in a communication network, and include, for example, network interface controllers, modems, wireless and wired interface cards, antennas, and other communication peripherals. For example, the computing system 1700 can use the network interface devices of the communication peripherals 1710 to communicate with one or more other computing devices (e.g., a computing device that determines a fragment mix signature based on the size distribution of nucleic acid molecules, another computing device that generates sequence data for a biological sample of a subject) over a data network.

[0142] Although the subject matter of the present invention has been described in detail with respect to certain embodiments thereof, those skilled in the art will appreciate that, upon gaining the foregoing understanding, they may easily generate modifications, variations, and equivalents to such embodiments. It should therefore be understood that the present disclosure is presented for purposes of illustration and not limitation, and does not exclude the inclusion of modifications, variations, and / or additions to the subject matter of the present invention, such as those that would be readily apparent to those skilled in the art. Indeed, the methods and systems described herein may be embodied in other various forms, and furthermore, various omissions, substitutions, and changes in the form of the methods and systems described herein may be made without departing from the spirit of the present disclosure. The appended claims and their equivalents are intended to cover such forms or modifications as are within the scope and spirit of the present disclosure.

[0143] Unless otherwise indicated, throughout this specification, descriptions utilizing terms such as "processing," "computing," "calculating," "determining," and "determining" should be understood to refer to the operations or processes of a computing device, e.g., one or more computers, or similar electronic computing device(s), that manipulates or transforms data represented as physical electronic or magnetic quantities in the memory, registers, or other information storage, transmission, or display devices of a computing platform.

[0144] The system(s) discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provides results conditioned on one or more inputs. Suitable computing devices include general-purpose microprocessor-based computing systems that access stored software that programs or configures the computing system from a general-purpose computing device to a specialized computing device that implements one or more embodiments of the inventive subject matter. Any suitable programming, scripting, or other type of language or combination of languages ​​may be used to implement the teachings contained herein into the software used to program or configure a computing device.

[0145] Certain embodiments of the methods disclosed herein may be implemented in operation of such a computing device. The order of the blocks presented in the above examples may be changed, e.g., the blocks may be rearranged, combined, and / or divided into sub-blocks. Certain blocks or processes may be implemented in parallel.

[0146] Conditional language used herein, such as, among others, "can," "could," "might," "may," "eg," and the like, is generally intended to convey that certain examples include certain features, elements, and / or steps and other examples do not include them, unless expressly stated otherwise or understood otherwise within the context in which it is used. Thus, such conditional language is not intended to imply that features, elements, and / or steps are generally required in any way in one or more examples, or that one or more examples necessarily include logic for determining whether those features, elements, and / or steps are included in or performed in any particular example, with or without author input or direction.

[0147] The terms "comprising," "including," "having," and the like are synonymous and are used in an inclusive, open-ended manner and do not exclude additional elements, features, acts, operations, and the like. The term "or" is also used in an inclusive (and not exclusive) sense, so that, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Use of "adapted to" or "configured to" herein is intended as an open and inclusive statement that does not exclude devices adapted or configured to perform additional tasks or steps. Additionally, use of "based on" is intended to be open and inclusive in that a process, step, calculation, or other operation "based on" one or more recited conditions or values ​​may in fact be based on additional conditions or values ​​other than those recited. Similarly, use of "based at least in part on" is intended to be open and inclusive in that a process, step, calculation, or other operation that is "based at least in part on" one or more recited conditions or values ​​may in fact be based on additional conditions or values ​​other than those recited. The headings, lists, and numbering contained herein are for ease of description only and are not meant to be limiting.

[0148] The various features and processes described above may be used independently of one another or may be combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure. Furthermore, in some implementations, certain method or process blocks may be omitted. Also, the methods and processes described herein are not limited to any particular order, and the blocks or states thereof may be performed in other orders as appropriate. For example, the blocks or states described may be performed in an order other than that specifically disclosed, or multiple blocks or states may be combined into a single block or state. The example blocks or states may be performed sequentially, in parallel, or in some other manner. Blocks and states may be added to or removed from the examples of the present disclosure. Similarly, the example systems and components described herein may be configured differently than described. For example, elements may be added, removed, or rearranged compared to the examples of the present disclosure.

Claims

1. accessing sequence data of a subject's biological sample; generating a first set of array size values ​​based on the array data, wherein each array size value in the first set corresponds to a size of an array of the array data; determining a fragment mix signature amplitude for the subject by projecting the first set of sequence size values ​​onto a fragment mix signature latent variable, the latent variable being generated by applying one or more signal separation algorithms to a set(s) of other sequence size values ​​obtained from one or more reference biological samples; processing the fragmentmix signature amplitudes using a machine learning model to generate a result comprising a classification that predicts whether the subject has a particular disease; and Outputting the results A method comprising:

2. 2. The method of claim 1, wherein each latent variable of the latent variables of the fragment mix signature comprises a histogram or weight vector representing the size distribution of the set(s) of other sequence size values ​​of the one or more reference samples, and the fragment mix signature amplitude of the biological sample is determined by projecting the first set of sequence size values ​​and the set(s) of other sequence size values ​​onto each latent variable of the latent variables.

3. 2. The method of claim 1, wherein the first set of sequence size values ​​corresponds to sequences of the sequence data that align to one or more genomic regions.

4. The method of claim 1 , wherein the one or more signal separation algorithms include one or more blind source separation algorithms.

5. The method of claim 4 , wherein the one or more blind source separation algorithms further comprise an independent component analysis algorithm.

6. The method of claim 4 , wherein the one or more blind source separation algorithms further comprise a non-negative matrix factorization algorithm.

7. 2. The method of claim 1, wherein the method predicts progressive digestion of a DNA fragment associated with a nucleosome-bound DNA helix pitch from one or more graph components of a first latent variable of the latent variables.

8. 10. The method of claim 1, further comprising predicting cell-to-cell heterogeneity of a DNA-binding protein from one or more graph components of a second latent variable of the latent variables.

9. 2. The method of claim 1, wherein each sequence represented by a corresponding sequence size value comprises DNA fragments having sizes ranging from 60 bp to 600 bp.

10. 10. The method of claim 1, wherein the sequence data comprises sequences corresponding to a plurality of somatic variants detected in the biological sample.

11. The method of claim 1 , wherein the set of sequence size values ​​is further an empirical probability mass function generated based on the sequence of the sequence data.

12. 2. The method of claim 1, wherein the sequence data corresponds to a plurality of cell-free DNA molecules of the biological sample, the plurality of cell-free DNA molecules comprising circulating tumor DNA molecules.

13. The method of claim 1 , wherein the specific disease is cancer.

14. 2. The method of claim 1, wherein determining the fragment mix signature amplitude of the subject comprises projecting the first set of array size values ​​onto each of the subset of latent variables.

15. The method of claim 14 , wherein subsetting the latent variables of the fragment mix signature comprises applying a clustering algorithm to the latent variables.

16. The method of claim 14 , wherein the subset of the latent variables of the fragment mix signature is used as preprocessing training data for another machine learning model.

17. The method of claim 14 , wherein the subset of the latent variables of the fragmentmix signature is used as components in a subsequent principal component analysis.

18. generating a set of terminal motif sequence data based on the sequence data, wherein each terminal motif sequence data of the set identifies the number or relative frequency of nucleic acid molecules having a terminal sequence corresponding to a particular terminal motif; and determining latent variables of alternative fragment mix signatures by applying one or more signal separation algorithms to the set of terminal motif sequence data; The method of claim 1 further comprising:

19. 19. The method of claim 18, wherein determining the fragment mix signature amplitude of the subject comprises projecting the set of terminal motif sequence data onto the latent variables of the other fragment mix signatures.

20. one or more data processors; and A non-transitory computer-readable storage medium comprising instructions that, when executed on said one or more data processors, cause said one or more data processors to perform one or more methods disclosed herein. Including, the system.

21. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product comprising instructions configured to cause one or more data processors to perform one or more of the methods disclosed herein.