Optimization of sequencing panel assignment

The panel assignment model optimizes participant assignment to a target panel using computer models, addressing inefficiencies in current cancer detection methods by improving sensitivity and reducing the need for multiple screenings, enabling early cancer detection and treatment planning.

JP2026510669APending Publication Date: 2026-04-10GRAIL INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Current cancer detection methods are cancer-type specific, inefficient, and plagued by low detection rates and high false positives, making it difficult to detect early cancers and requiring multiple invasive screening methods, while existing genetic models struggle with extremely small tumor-to-healthy cell ratios and variability in variant coverage.

Method used

A panel assignment model optimizes participant assignment to a target panel by considering participant and variant characteristics, using computer models for cancer signal identification and quantification, enabling a single, comprehensive screening method for multiple cancer types from a single cfDNA sample.

Benefits of technology

Improves cancer detection sensitivity and reduces the number of necessary panels, allowing for more cost-effective and less invasive screening by identifying cancer signals in a single sample, enhancing early cancer detection and treatment planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510669000001_ABST
    Figure 2026510669000001_ABST
Patent Text Reader

Abstract

This disclosure relates to a method for improving sequencing panel assignments for samples from two or more individuals. The system is configured to generate sequencing panel assignments with optimized sample sets for each panel, reducing costs without compromising the detection limits of the assay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 489,826, filed on March 13, 2023, which is incorporated herein by reference.

Background Art

[0002] Cancer is the leading cause of death worldwide. The fact that cancer is usually detected at a late stage, limiting the effectiveness of treatment options for long - term survival, has pushed up the lethality rate of cancer. Current detection methods are generally cancer - type specific, that is, screening is performed individually for each cancer type. Each individual screening method is tailored to the cancer type. For example, mammography scans are used for breast cancer detection, while colonoscopy or stool tests assist in the detection of colorectal cancer. It is not possible to cross - apply each of the various screening methods to other cancer types. For example, if one individual is to be screened for three different possible cancer types, a healthcare provider may need to perform or order the performance of three different screening methods. Each of these screening methods can necessarily involve a combination of invasive and / or non - invasive techniques for the identification of tumorous growth, the taking of a biopsy from that growth, and the performance of an analysis regarding the tissue biopsy.

[0003] Furthermore, current screening methods are plagued by low detection rates or high false - positive rates. With a low detection rate, early cancers, such as those that have just started to develop, are often missed. With a high positive rate, individuals without cancer are misdiagnosed as having a positive cancer status. As a result, most screening tests are practical only when used to examine individuals with a high risk of developing the cancer being screened, and the ability to detect cancer in the general population is limited.

[0004] Novel studies have suggested that various genetic variations are involved as early markers and indicators of cancer development. For example, these studies have focused on copy number variations, small variants (including single nucleotide polymorphisms, insertions, and deletions), and methylation abnormalities. This research uses models to identify correlations between these genetic variations and cancer status. Nevertheless, even such models face several challenges. Early cancer detection is particularly challenging due to the extremely small ratio of target tumor cells to non-cancerous cells. Such extremely small ratios can be on the order of 1:1000, 1:10,000, or even 1:100,000. This creates the challenge of detecting minute cancer signals amidst healthy signals.

[0005] To investigate the tumor ratio in all cfDNA samples, targeted sequencing panels (e.g., detecting small nucleotide variants) are designed to identify somatic variants called from whole-genome sequencing of tumor biopsies from each participant. Rather than ordering one panel for each participant, sequencing data from multiple participants can be combined into a single panel, resulting in both cost reduction and simplification of work. Challenges arise when optimizing combinations of multiple participants into a single panel. A rudimentary method would be to randomly assign participants to the panel, but this method may result in suboptimal pairings of participants, such as participants with different numbers of variants (leading to insufficient density and potentially resulting in more panels than necessary), participants with variants that have different probe pull-down efficiencies (leading to large variability in variant coverage and reduced sensitivity for variants with low coverage), or participants with variants with different noise rates (similarly leading to variability in panel sensitivity and the possibility of missed variant calls). [Overview of the project] [Problems that the invention aims to solve]

[0006] This disclosure relates to addressing the issues mentioned above, and more specifically, to optimizing the assignment of participants to the target panel. The background information provided herein is intended to provide a general overview of the context of this disclosure. Unless otherwise specifically indicated herein, the material described in this section is not prior art to the claims of this application, and its inclusion in this section does not constitute prior art or an indication of prior art. [Means for solving the problem]

[0007] One or more of the present inventions described herein provide improvements in cancer detection and treatment, more particularly, optimization of participant assignment to a target panel. More specifically, the panel assignment model takes into account various participant characteristics, characteristics of the variants to be screened for each participant, characteristics of the probes targeting the variants, or a combination of some of these, when determining participant assignment to a panel and determining which variants each participant will be subjected to targeted screening. Such optimization improves the targeted sequencing process, for example, by improving the sensitivity of variant calling among participants in a single panel and by identifying the densest configuration of participants for a panel, thereby minimizing the number of panels and the associated sequencing costs.

[0008] One or more of the present inventions involve screening for cancer signals in a target cell-free deoxyribonucleic acid (cfDNA) sample. Such a cfDNA sample may contain thousands, tens of thousands, hundreds of thousands, millions, or more cfDNA fragments, and therefore the sequence reads output by the sequencer may also be of a similar magnitude, or even several times that magnitude depending on the sequencing depth of the sample. Each sequence read for a cfDNA fragment may vary in length, for example, up to 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000 bp. These next-generation sequencing techniques significantly increase the amount of fragments that can be sequenced and analyzed, and therefore enable identification by such models even if the amount of cancer signal in the sample is very small. One or more of the present inventions generally have the ability to screen for cancer, or more types of cancer, from a single sample. This improves upon traditional cancer-specific screening methods by providing a single, comprehensive screening method capable of screening various cancer types from a single cfDNA sample. For example, screening different cancers typically involved separate screening methods, each of which would involve detecting abnormal tissue growth, taking a biopsy of the detected growth, and then determining malignancy by analyzing the clinical laboratory values ​​of the biopsy.

[0009] One or more of the present inventions implement computer models for identifying and quantifying cancer signals. In one or more embodiments, the computer model may perform variant calling as features for cancer classification. The computer model may include a trained cancer classifier configured to take a feature vector generated based on called variants as input and output a cancer prediction based on the input feature vector. The cancer prediction may be a binary prediction and / or a multiclass prediction. A binary prediction may be the likelihood of cancer being present. A multiclass prediction may be the likelihood of a particular cancer species from several cancer species being evaluated. Because a cancer classifier with screening capabilities across multiple cancer species is trained, healthcare professionals can utilize a single, comprehensive screening rather than multiple different types of screening. In one or more embodiments, the cancer classifier is a machine learning model rooted in computer capabilities, which is virtually impossible for the human brain to perform. In one or more embodiments, the cancer classifier includes non-mathematical operations, including operations based on handling electronic data in the context of a computing device.

[0010] Cancer prediction can be used by healthcare providers for diagnosis, prognosis assessment, treatment personalization, treatment evaluation, and minimal residual disease detection. Obtaining such insights from minimally invasive liquid biopsy can increase screening frequency while minimizing harm to patients.

[0011] Close 1. A method for performing target variant sequencing of a sample collection, comprising: obtaining initial sequencing data for each sample describing the presence or absence of each of several gene variants in the reference genome, wherein the initial sequencing data includes sequence reads of nucleic acid fragments in the biological sample obtained from the subject; determining the number of gene variants present in the sequencing data of each sample; determining one or more characteristics of each gene variant present in the sequencing data of each sample; and determining the panel assignment for each sample to one of several target variant sequencing panels by applying a panel assignment model, wherein applying the panel assignment model ensures uniformity of panel size across all target variant sequencing panels and assignment to each target variant sequencing panel by processing each sample in the sample collection sequentially. A method comprising: determining corresponding panel assignments to optimize the uniformity of gene variant characteristics across all samples; swapping panel assignments of at least two samples by performing a swap operation to further optimize the uniformity of panel sizes across all target variant sequencing panels and the uniformity of gene variant characteristics across all samples assigned to each target variant sequencing panel; generating each target sequencing panel that includes the samples with panel assignments to the target sequencing panel and indicates an aggregated set of gene variants across all samples assigned to the target sequencing panel; and performing target variant sequencing on each target variant sequencing panel, including the assigned samples, with respect to cell-free deoxyribonucleic acid (cfDNA) samples subsequently collected from the subject.

[0012] Close 2. A method of Close 1 or any subordinate Close, wherein the initial sequencing includes whole-genome sequencing or whole-exome sequencing of each specimen.

[0013] Close 3. A method according to Close 1 or any Close dependent thereon, wherein the gene variant includes single nucleotide variants, insertion variants, and deletion variants.

[0014] Close 4. A method of Close 1 or any Close dependent thereon, wherein the gene variant includes a copy number variant.

[0015] Close 5. The method of Close 1 or any dependent Close, wherein one or more characteristics of each gene variant are selected from the group consisting of the guanine-cytosine content of the gene variant, the error rate of the targeting probe of the gene variant, the sequence depth count of the gene variant, the presence or absence of the gene variant, the mean allele frequency of the gene variant, the total number of gene variants, and the allele frequency of the true gene variant.

[0016] Close 6. A method of Close 1 or any Close dependent thereon, further comprising applying a panel assignment model to determine a panel assignment for each sample in a sample set subject to one or more hard constraints.

[0017] The method according to Close 6, wherein one or more hard constraints are selected from the group consisting of the maximum number of samples per target sequencing panel, the maximum number of samples of one type per target sequencing panel, or the total number of gene variants per target sequencing panel.

[0018] Close 8. A method by Close 1 or any dependent Close, wherein the panel assignment model includes a function that scores the uniformity of panel sizes across all target variant sequencing panels and the uniformity of gene variant characteristics across all samples assigned to each target variant sequencing panel.

[0019] The method according to Close 8, wherein determining the corresponding panel assignment is done by sequentially processing each sample in the sample set, or by applying a greedy algorithm or a dynamic programming algorithm to the function.

[0020] Close 10. A method of Close 8 or any Close dependent thereon, comprising performing a swap operation, evaluating the change in scores based on a function that evaluates the swap of at least two samples, and swapping the panel assignments of at least two samples based on whether the change in scores exceeds a threshold.

[0021] Close 11. A method of Close 1 or any subordinate Close that determines pairs of samples assigned to different target sequencing panels by repeatedly performing a swap operation.

[0022] Close 12. A method according to Close 1 or any Close dependent thereon, further comprising identifying an optimal aggregated set of gene variants across all samples based on the characteristics of gene variants in each sample assigned to the target sequencing panel for each target sequencing panel.

[0023] Close 13. The method of Close 1 or any Close dependent thereon, comprising: identifying an optimal set of gene variants for each target sequencing panel, determining for each target sequencing panel whether each specimen has a total number of gene variants exceeding a threshold per specimen; and, in response to at least one specimen having a total number of gene variants exceeding a threshold per specimen, identifying a subset of gene variants in specimens that should be included in a target sequencing panel to optimize the properties of the optimal aggregated set of gene variants.

[0024] The method according to Closure 13, wherein identifying a subset of gene variants to be included in a target sequencing panel includes including gene variants present in other specimens of the target sequencing panel.

[0025] The method according to Closure 1 or any one of the closures dependent thereon, wherein generating each target sequencing panel further includes identifying corresponding targeting probes that target an aggregated set of gene variants.

[0026] The method according to Closure 1 or any one of the closures dependent thereon, further including obtaining target sequencing data for a set of specimens from a target variant sequencing panel for a cfDNA specimen, wherein the target sequencing data includes sequence reads of cell-free nucleic acid fragments in a blood specimen; calling one or more variants present in the target sequencing data for each cfDNA specimen; determining a feature vector based on the one or more called variants; and predicting a tumor ratio in the cfDNA specimen by applying a cancer classifier to the feature vector.

[0027] A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to Closure 1 or any one of the closures dependent thereon when executed by the processor.

[0028] A system including a processor and the non-transitory computer-readable storage medium according to Closure 17.

[0029] Close 19. A target sequencing panel comprising a set of targeting probes that target aggregate gene variant sets across multiple specimens assigned to the target sequencing panel, wherein the aggregate gene variant sets and multiple specimens obtain initial sequencing data for each specimen in a specimen set describing the presence or absence of each of multiple gene variants present in a reference genome, wherein the initial sequencing data includes sequence reads of nucleic acid fragments in a biological specimen obtained from one subject; for each specimen, the number of gene variants present in the specimen sequencing data is determined; for each specimen, one or more characteristics of each gene variant present in the specimen sequencing data are determined; and by applying a panel assignment model, multiple target variant sequences are obtained for each specimen. A target sequencing panel, comprising determining a panel assignment to one of the sequencing panels, applying a panel assignment model to determine a corresponding panel assignment that optimizes the uniformity of panel size across all target variant sequencing panels and the uniformity of gene variant characteristics across all samples assigned to each target variant sequencing panel by processing each sample in the sample set sequentially, and performing a swap operation to swap the panel assignments of at least two samples to further optimize the uniformity of panel size across all target variant sequencing panels and the uniformity of gene variant characteristics across all samples assigned to each target variant sequencing panel.

[0030] Close 20. Multiple target sequencing panels, each including a set of targeting probes that target aggregate gene variant sets across multiple samples assigned to the target sequencing panel, wherein the target sequencing panel obtains sample-by-sample initial sequencing data in a sample set describing the presence or absence of each of multiple gene variants in a reference genome, wherein the initial sequencing data includes sequence reads of nucleic acid fragments in a biological sample obtained from one subject; for each sample, the number of gene variants present in the sample sequencing data is determined; for each sample, one or more characteristics of each gene variant present in the sample sequencing data are determined; and by applying a panel assignment model, multiple target batches are obtained for each sample. Multiple target sequencing panels, comprising determining a panel assignment to one of the target sequencing panels, applying a panel assignment model to determine a corresponding panel assignment that optimizes the uniformity of panel size across all target variant sequencing panels and the uniformity of gene variant characteristics across all samples assigned to each target variant sequencing panel by processing each sample in the sample set sequentially, and performing a swap operation to swap the panel assignments of at least two samples to further optimize the uniformity of panel size across all target variant sequencing panels and the uniformity of gene variant characteristics across all samples assigned to each target variant sequencing panel.

[0031] A method for improving sequencing panel assignments for samples from 2 or more individuals, comprising: obtaining sequencing data for each sample; selecting feature values ​​from the sequencing data; applying a machine learning model to determine sequencing panel assignments based on the feature values ​​from the sequencing data; and generating an optimized sequencing panel assignment that includes samples from 2 or more individuals.

[0032] Close 22. The method according to Close 21, wherein the selection step further includes determining an optimized set of feature values.

[0033] Close 23. The method according to Close 22, wherein the machine learning model is selected from a classifier model, a pre-specified algorithm, and a regression model.

[0034] Close 24. The method described in Close 23, where the machine learning model is a classifier model.

[0035] Close 25. The method according to any one of Closes 21-24, wherein applying a classifier model ranks the samples based on decreasing feature values; and applying a greedy algorithm to add the next highest ranked sample from the remaining ranked samples to a panel, wherein the panel on which the samples are rearranged contains the lowest value of the feature values.

[0036] Close 26. The method according to Close 25, further comprising assigning samples to panels by processing the samples sequentially; determining the mean of feature values ​​for each panel; swapping two samples between two different panels; and measuring the deviation of the mean feature values ​​for each of the two different panels after the swap.

[0037] Close 27. The method according to Close 26, further comprising repeating the step described in claim 26 a predetermined number of swaps, thereby generating a panel assignment based on feature values ​​from sequencing data.

[0038] Close 28. The method of Close 27, wherein the repeated step is performed until the decrease in the mean of the features falls below a threshold.

[0039] Close 29. The method according to any one of Closes 21-28, wherein the panel contains 16 specimens.

[0040] Close 30. The method according to any one of Closes 21 to 28, wherein the panel has 16 or fewer specimens.

[0041] Close 31. The method according to any one of Closes 21-30, wherein the panel includes a benchmark sample.

[0042] The method according to Close 31, wherein the panel has one or fewer benchmark samples.

[0043] Close 33. The method according to Close 32, wherein applying a classifier model includes seeding a sequencing panel based on the rank number of decreasing feature values; swapping sequencing panel assignments for two participants seeded into two different panels; measuring the decrease in the loss function after the swap; and comparing the set of sequencing panel assignments with the feature values.

[0044] Close 34. The method of Close 33, further comprising repeating over a predetermined number of steps or until the decrease in the loss function falls below a threshold.

[0045] The method according to Close 33 or 34, further comprising determining a sequencing panel assignment that satisfies Close 35.DRAWING.

[0046] Close 36. A method according to any one of Closes 21-35, wherein sequencing data is obtained from sequencing cell-free nucleic acid molecules present in biological specimens obtained from multiple individuals.

[0047] Close 37. A method according to any one of Close 21 to 36, wherein the feature value corresponds to a genomic region containing one or more of the following: cancer-related genes, mutation hotspots, and viral regions.

[0048] Close 38. The method according to any one of Closes 21-37, wherein the sequencing data includes genomic regions associated with high-signal cancer or liquefied cancer.

[0049] Close 39. The method according to any one of Close 21 to 38, wherein the feature value corresponds to a feature corresponding to one or more of the following: GC content, error rate, sequence depth count, presence or absence of variants, mean allele frequency, total number of minor variants, and allele frequency of true variants.

[0050] Close 40. A method according to one of Close 21-39, wherein the feature value is the sum of multiple feature values.

[0051] Close 41. The method described in any one of Close 21-40, wherein the feature value is a variant.

[0052] Close 42. The method according to Close 41, wherein the variant comprises one or more of the following: a single nucleotide variant, an insertion, and a deletion.

[0053] A non-temporary computer-readable medium storing one or more programs, wherein one or more programs include instructions that cause the device to perform one of the methods described in any one of the methods described in 42, including a processor, when the device is executed by the device.

[0054] Close 44. An electronic device comprising one or more processors; memory; and one or more programs, which are stored in memory and configured to be executed by one or more processors, and which include instructions for performing the method described in any one of Close 101 to 42. [Brief explanation of the drawing]

[0055] [Figure 1] This is an illustrative flowchart illustrating the entire workflow for cancer classification of specimens according to one or more embodiments. [Figure 2A]An illustrative flowchart of a nucleic acid sample sequencing apparatus according to one or more embodiments is shown. [Figure 2B] This is an example of a block diagram of an analytics system according to one or more embodiments. [Figure 3] This is a flowchart for sequencing nucleic acid samples according to one or more embodiments. [Figure 4] This is a flowchart of an optimized sequencing panel assignment according to one or more embodiments. [Figure 5] This is a flowchart of a variant calling method according to one or more embodiments. [Figure 6A] This is a flowchart for training a variant feature-based cancer classifier according to one or more embodiments. [Figure 6B] This is a flowchart for deploying a variant feature-based cancer classifier according to one or more embodiments. [Figure 7] The tumor ratio (TF) plots for 210 participants are shown. [Figure 8] The tumor ratio (TF) plots for 201 benchmark specimens are shown. [Figure 9] The tumor ratio (TF) plots for 5 benchmark specimens out of 201 cases with a 10% TF threshold, as shown in Figure 8, are illustrated. [Figure 10] The characteristics, including the number of variants, for the five benchmark samples are shown in Figures 8 and 9. [Figure 11] This histogram shows the error rate (y-axis) for each type of transformation (x-axis) during SNP analysis. [Figure 12] This bar graph shows the mean target coverage (MTC) after collapse relative to the input cfDNA yield. [Figure 13] A dot plot of MTC against input cfDNA mass is shown. [Figure 14] The box plots for post-collapse MTC for each of the conditions indicated on the X-axis are shown. [Figure 15] This graph shows the limit of detection (LoD) for the total depth from a sample taken from a single plasma tube. [Figure 16] This dot plot shows the limit of detection (LoD) modeled based on logMin tumor ratio (TF) (x axis) versus logN variants (y axis). [Figure 17] This graph shows the LoD modeled based on LogMinTF (x-axis) versus participant ratio (y-axis). [Figure 18] This shows a dot plot for the Line of Derivation (LoD) modeled based on LogMinTF (x-axis) and LogNVariants (y-axis). [Figure 19] This graph shows the LoD modeled based on LogMinTF (x-axis) and participant ratio (y-axis). [Figure 20] The bar graph shows the types and number of cancers present in 129 specimens. [Figure 21] The bar graph shows the stage of cancer present in 129 samples. [Figure 22] The plot of the log-likelihood ratio (LLR) for true calls versus noise is shown. [Figure 23] This shows a histogram of the error rate (y-axis) for each different type of transformation (x-axis) during SNP analysis. [Figure 24] A dot plot illustrating the relationship between GC content and sequencing depth is shown for each sample combination. [Figure 25] A box plot summarizing the plots from Figure 24 is shown. [Figure 26] This heatmap identifies the SNP prioritization driven by SNP type (i.e., error rate) within the LoD framework. [Figure 27A] A dot plot showing GC content (x axis) versus sequence depth is shown. [Figure 27B] This dot plot shows GC content (x-axis) versus average bag size. [Figure 28A]The box plots shown summarize the normalized depth for each sample corresponding to the GC bin on the x-axis. [Figure 28B] The table corresponding to the data in Figure 28A is shown. [Figure 29] The dot plot shows total coverage ("Depth_Raw"; x-axis) versus average_bagsize ("Average_BagSize_1_Tube"; y-axis). [Figure 30] The dot plot shows bag size versus duplex percentage. [Figure 31] This plot shows a comparison of noise rate ("value"; y-axis) and different transformation types ("snp_group"; x-axis). [Figure 32A] This graph shows the relative cost (y-axis) and expected effort-free unique coverage (EFC) for all samples. [Figure 32B] This graph shows the relative cost (y-axis) and expected effort-free unique coverage (EFC) for all samples. [Figure 33A] The optimal number of panels is shown as a graph (Figure 33A). [Figure 33B] The optimal number of panels is shown in a table (Figure 33B). [Modes for carrying out the invention]

[0056] These figures depict various embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following considerations that alternative embodiments of the structures and methods exemplified herein can be used without departing from the principles described herein.

[0057] Detailed explanation I. Overview Early detection and classification of cancer are crucial skills. Detecting cancer before it becomes symptomatic benefits everyone involved, including patients, physicians, and loved ones. For patients, early detection increases the chances of achieving beneficial outcomes; for physicians, it expands treatment pathways that can lead to beneficial outcomes; and for loved ones, it increases the likelihood of saving their friends and family from this disease.

[0058] In recent years, early cancer detection technologies have advanced in the direction of analyzing gene fragments (e.g., DNA) in a person's blood to determine whether those gene fragments originate from cancer cells. These new techniques allow physicians to identify the presence of cancer in patients that might otherwise be undetectable through conventional screening processes. For example, consider a person at high risk of breast cancer. Traditionally, this person would undergo regular mammograms by their doctor, who would use the images of their breast tissue (e.g., X-rays) to identify cancerous tissue. Unfortunately, even with the highest resolution mammograms, physicians can only identify a tumor when it reaches approximately millimeter size. This means the cancer may have been present in the person for some time, undiagnosed and untreated. For most cancers, this visual determination is typical—that is, it can only be identified when it has grown large enough to be detectable by some kind of imaging technique.

[0059] Cancer detection using analysis of gene fragments in a patient's blood, for example, mitigates this problem. To explain, as soon as cancer cells form, DNA fragments begin to detach from them into the bloodstream. This occurs when cancer cells are very few in number, before they are even visible using imaging techniques. Therefore, with the right methods, a system that analyzes DNA fragments in the bloodstream could potentially identify the presence of cancer in a given person using these detached cancer DNA fragments. More importantly, this system could potentially do so before the cancer becomes detectable using conventional cancer detection techniques.

[0060] Cancer detection based on the analysis of DNA fragments is made possible by next-generation sequencing ("NGS") techniques. NGS, in a broad sense, is a group of technologies that enable high-throughput sequencing of genetic material. As will be discussed in more detail herein, NGS broadly consists of (1) sample preparation, (2) DNA sequencing, and (3) data analysis. Sample preparation is the laboratory method required to prepare DNA fragments for sequencing, sequencing is the process of reading ordered nucleotides in the sample, and data analysis is the process of processing and analyzing the genetic information of the sequencing data to identify the presence of cancer.

[0061] While these NGS steps can help achieve early cancer detection, they also introduce their own complex and harmful problems into cancer detection. Therefore, any improvements in sample preparation, DNA sequencing, and / or data analysis, including sample preparation, algorithmic processing, and summarization or presentation of predictions or conclusions, would lead to improvements in cancer detection technology and, more generally, early cancer detection.

[0062] To explain further, for example, (1) problems introduced into sample preparation include optimizing the samples assigned to the panel, the quality of DNA samples, sample contamination, fragmentation bias, and accurate indexing. Addressing these problems can lead to better genetic data for cancer detection. Similarly, (2) problems introduced into sequencing include, for example, errors in accurate transcription of fragments (e.g., reading "A" instead of "C"), inaccurate or difficult fragment assembly and overlap, inconsistent coverage uniformity, the conflict between sequencing depth, cost, and specificity, and insufficient sequencing length. Again, addressing any of these problems can lead to improved genetic data for cancer detection.

[0063] (3) The problems in data analysis are the most troublesome and complex. The challenges that arise stem from the enormous amount of data produced by NGS sequencing techniques. The resulting gene data set is typically in the order of terabytes, and effectively and efficiently analyzing this amount of data requires a great deal of effort, both procedurally and computationally. For example, NGS sequencing analysis involves several baseline processing steps, such as aligning reads with each other, aligning and mapping reads to a reference genome, identifying and calling mutant genes, identifying and calling methylation abnormal genes, and generating functional annotations. Performing any of these processes on several terabytes of gene data is computationally expensive even with the most powerful computer architectures, and is completely impossible for the average human mind. In addition, gene sequencing data obtained from error-prone sample preparation and sequence reading processes may result in a large portion of the resulting gene data being of low quality or unusable for cancer identification. For example, large amounts of genetic data may contain contaminated samples, transcription errors, mismatched regions, and excessively occurring regions, making them unsuitable for high-precision cancer detection. Identifying and considering low-quality genetic data across the vast amounts of genetic data obtained from NGS sequencing is also procedurally and computationally demanding, and practically impossible for humans to accomplish. Overall, creating any process that leads to more efficient processing of large array sequencing data could improve cancer detection using NGS sequencing.

[0064] Finally, and perhaps most importantly, accurately identifying DNA from NGS data that provides information for identifying the presence of cancer is also difficult (especially in the context of early cancer detection). For it to be effective, algorithms are needed that, for example, correct errors generated by sample preparation and sequencing, and resolve the problems of large-scale data analysis associated with NGS techniques. In other words, the design of one or more machine learning models or other computational algorithms that realize early cancer detection based on next-generation sequencing techniques must be configured to take into account the problems that these techniques create. Some of these techniques and models are discussed below, and further detailed improvements to techniques and models at the current level of technology are considered.

[0065] Training of post-machine learning models described herein (such as contamination models, cancer classifiers, any other neural networks, and any other models referenced herein) may involve, but are not limited to, data loading operations, data storage operations, data toggling or modification operations, non-temporary computer-readable storage medium modification operations, metadata removal or data cleaning operations, data compression operations, protein structure modification operations, image modification operations, noise application operations, and noise reduction operations, and may involve, at least in part, the performance of one or more non-mathematical operations or implementation of non-mathematical functions by a machine or computing system. Thus, training of post-machine learning models described herein may be based on or involve mathematical concepts, but is not limited to merely mathematical calculations, mathematical operations, or computational acts of variables or numbers using mathematical methods.

[0066] Similarly, it must be noted that training these models described herein is virtually impossible for the human brain alone. These models are inherently complex, containing a vast number of weights and parameters linked through one or more complex functions. Training and / or deploying such models involves such a large number of operations that it is neither feasible for the human brain alone nor for the use of pen and paper. In such embodiments, the number of operations can reach hundreds, thousands, tens of thousands, hundreds of thousands, millions, billions, or even trillions. Furthermore, the training data may contain hundreds, thousands, tens of thousands, hundreds of thousands, millions, or billions of sequence reads, each sequence read potentially containing hundreds to thousands of nucleotides. Rapid processing of such enormous amounts of sequencing data is also of paramount importance for effectiveness and applicability in early cancer screening. For example, even if (hypothetically) all the necessary calculations were performed by the human brain, the time required could extend to several years (though not to several decades), at which point the advantage of sequence data analysis for early cancer detection would be lost. Therefore, such models are inevitably rooted in computer technology in their implementation and use.

[0067] IA Cancer Classification Workflow Figure 1 is an illustrative flowchart illustrating an overall workflow 100 for cancer classification of specimens according to one or more embodiments. The workflow 100 is performed by one or more entities, including, for example, healthcare providers, sequencing equipment, analytics systems, etc. The purpose of the workflow includes detecting and / or monitoring cancer in individuals. From a medical perspective, workflow 100 may serve to complement other existing cancer diagnostic tools. Workflow 100 may serve to provide routine cancer monitoring to better inform early cancer detection and / or treatment planning for individuals diagnosed with cancer. The entire workflow 100 may include additional / fewer steps than those shown in Figure 1.

[0068] A healthcare provider performs specimen collection 110. An individual undergoing cancer classification visits their healthcare provider. The healthcare provider collects a specimen for cancer classification. Examples of biological specimens include, but are not limited to, a biopsy of the target tissue, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or ascites. In one or more embodiments, specimen collection is minimally invasive or non-invasive. The specimen contains genetic material belonging to the individual, which may be extracted and sequenced for cancer classification. After specimen collection is complete, the specimen is provided to a sequencing device. Along with the specimen, the healthcare provider may collect other information about the individual, such as biological sex, age, race, smoking status, other health indicators, any prior diagnoses, etc.

[0069] A sequencing instrument performs sample sequencing 120. A laboratory clinician may perform one or more processing steps on the sample in preparation for sequencing. When ready, the clinician loads the sample into the sequencing instrument. Examples of instruments used for sequencing are further described in conjunction with Figures 2A and 2B. Generally, a sequencing instrument extracts and isolates a fragment of nucleic acid to be sequenced and determines the sequence of nucleic acid bases corresponding to that fragment. Sequencing may also include amplification of nucleic acid material. Various sequencing processes include Sanger sequencing, fragment analysis, and other next-generation sequencing techniques. Next-generation sequencing has the capability to generate high-throughput sequencing data, such as 10,000, 100,000, 1,000,000, or 10,000,000 sequence reads, with each sequence read having a length of 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, etc. Therefore, high-throughput sequencing data is of a size that is impractical for human analysis. Sequencing may be whole-genome sequencing or targeted sequencing using a targeted panel. In the context of DNA methylation, bisulfite sequencing can determine the methylation state through the conversion of unmethylated cytosine at CpG sites to bisulfite. Sample sequencing 120 generates sequences for multiple nucleic acid fragments in a sample.

[0070] The analytics system performs pre-analysis processing 130. An example analytics system is shown in Figure 2B. Pre-analysis processing 130 may include, but is not limited to, deduplication of sequence reads, determination of coverage metrics, determination of contamination in the sample, removal of contaminated fragments, calling sequencing errors, etc.

[0071] The analytics system performs one or more analyses 140. The analyses are statistical analyses or the application of one or more trained models to predict at least the cancerous status of the original individual from which the sample originates. Different genetic features may be evaluated and considered, such as CpG site methylation, single nucleotide variants (SNVs), insertions or deletions (indels), copy number variations, and other types of gene variants. Generally, analyses may include detecting sample contamination, determining the variant characteristics of the sample, determining panel assignments by applying a panel assignment model, variant calling, feature extraction from the called variants, and cancer classification. The cancer classifier 146 determines the cancer prediction by inputting the extracted features. The cancer prediction may be a label or a value. Labels may indicate a specific cancerous status; for example, a binary label may indicate the presence or absence of cancer, and a multiclass label may indicate one or more cancer types from multiple cancer types being screened. Values ​​may indicate the likelihood of a specific cancerous status, e.g., the likelihood of cancer, and / or the likelihood of a specific cancer type.

[0072] The analytics system returns a prediction of 150 to the healthcare provider. The healthcare provider can then develop or adjust a treatment plan based on the cancer prediction. Treatment optimization is further described in Section IVC. Treatment. In some embodiments, the analytics system may utilize a cancer classification workflow for prognosis determination, treatment personalization, treatment evaluation, and cancer status monitoring.

[0073] IB definition When used herein, the terms “about” or “approximately” may mean within an acceptable margin of error for a particular value as determined by those skilled in the art, such margin of error may depend in part on how the value is measured or determined, for example, on the limits of the measuring system. For example, “about” may mean within one standard deviation or more, according to the practice in the art. “About” may mean within ±20%, ±10%, ±5%, or ±1% of a given value. The terms “about” or “approximately” may mean within a certain range, within five times, or within twice a certain value. Where detailed values ​​are described in this application and claims, unless otherwise specified, the term “about” should be assumed to mean within an acceptable margin of error for those detailed values. The term “about” may have the meaning as generally understood by those skilled in the art. The term “about” may mean ±10%. The term “about” may mean ±5%.

[0074] As used herein, the terms “alternative allele” or “ALT” refer to an allele that has one or more mutations compared to, for example, a reference allele corresponding to a known gene.

[0075] As used herein, the term “alternate depth” or “AD” refers to the number of read segments in a specimen that support ALT, for example, those containing ALT mutations.

[0076] As used herein, the terms “alternate frequency” or “AF” refer to the frequency of a given ALT. AF can be obtained for a given ALT by dividing the corresponding AD of a sample by the depth of that sample.

[0077] As used herein, the term “cell-free nucleic acid” or “cfNA” refers to nucleic acid fragments originating from one or more healthy cells and / or one or more unhealthy cells (e.g., cancer cells) that circulate within an individual’s body (e.g., in the blood). The term “cell-free DNA” or “cfDNA” refers to deoxyribonucleic acid fragments circulating within an individual’s body (e.g., in the blood). In addition, cfNA or cfDNA within an individual’s body may originate from other non-human sources.

[0078] As used herein, the terms “circulating tumor DNA” or “ctDNA” mean nucleic acid fragments originating from tumor cells or other types of cancer cells that are released into the bodily fluids of an individual (e.g., blood, sweat, urine, or saliva) as a result of biological processes such as apoptosis or necrosis of dying cells, or that are actively released by living tumor cells.

[0079] As used herein, the terms “DNA fragment,” “fragment,” or “DNA molecule” may generally refer to any deoxyribonucleic acid fragment, i.e., cfDNA, gDNA, ctDNA, etc.

[0080] As used herein, the terms “genomic nucleic acid,” “genomic DNA,” or “gDNA” refer to nucleic acid molecules or deoxyribonucleic acid molecules obtained from one or more cells. In various embodiments, gDNA may be extracted from healthy cells (e.g., non-tumor cells) or from tumor cells (e.g., biopsy specimens). In some embodiments, gDNA may be extracted from cells of a blood cell lineage, such as leukocytes.

[0081] As used herein, the term “informational sequence read” refers to a sequence read that has a called variant (e.g., a copy number variant, a single nucleotide variant, an insertion, a deletion, or other type of genetic variation).

[0082] As used herein, the term “informational score” may refer to a score for a given genomic region based on sequence reads that provide information from samples within that genomic region. For example, an informational score could be a count of called variants. Informational scores are used in the context of feature development of samples for classification.

[0083] As used herein, the terms “biological specimen,” “patient specimen,” or “specimen” refer to any specimen taken from an subject, which may reflect the biological state associated with the subject, and may include cell-free DNA. A specimen may be a liquid specimen or a solid specimen (e.g., a cell or tissue specimen). A biological specimen may be a bodily fluid, such as blood, plasma, serum, urine, vaginal fluid, fluid from a hydrocele (e.g., of the testes), vaginal lavage fluid, pleural fluid, pericardial fluid, peritoneal fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple secretions, or aspirates from various parts of the body (e.g., thyroid, breast). A biological specimen may also be a fecal specimen. A biological specimen may contain any tissue or material derived from a living or dead subject. A biological specimen may be a cell-free specimen. A biological specimen may contain nucleic acids (e.g., DNA or RNA) or fragments thereof.

[0084] The term "nucleic acid" may refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or any hybrid or fragment thereof. The nucleic acids in a sample may be cell-free nucleic acids. In various embodiments, the majority of the DNA in a biological sample enriched with respect to cell-free DNA (e.g., a plasma sample obtained by a centrifugation protocol) may be cell-free (e.g., more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA may be cell-free). The biological sample may be subjected to a process that physically disrupts the tissue or cellular structure (e.g., centrifugation and / or cell lysis), thereby releasing intracellular components into a solution that may further contain enzymes, buffers, salts, surfactants, etc., which can be used to prepare a sample for analysis.

[0085] As used herein, the terms “control,” “control specimen,” “reference,” “reference specimen,” “normal,” and “normal specimen” refer to specimens from subjects that do not have a particular condition or are otherwise healthy. For example, the method disclosed herein can be performed on subjects with tumors, where the reference specimen is a specimen taken from healthy tissue of the subject. The reference specimen may be obtained from the subject or from a database. The reference may be, for example, a reference genome used to map nucleic acid fragment sequences obtained by sequencing specimens from the subject. A reference genome may refer to a haploid or diploid genome to which nucleic acid fragment sequences from biological specimens and constitutive specimens can be aligned and compared. An example of a constitutive specimen may be DNA from leukocytes obtained from the subject. For haploid genomes, there may be only one nucleotide at each locus. For diploid genomes, heterozygous loci can be identified; each heterozygous locus has two alleles, where either allele may enable matching during alignment with that locus.

[0086] As used herein, the terms “cancer” or “tumor” refer to an abnormal mass of tissue in which the growth of the mass exceeds and does not coordinate with the growth of normal tissue.

[0087] As used herein, the term “false positive” refers to a mutation that is incorrectly determined to be a true positive. Generally, false positives are more likely to occur when processing sequence reads associated with a higher mean noise rate or higher noise rate uncertainty.

[0088] As used herein, the term “genomic” refers to the characteristics of the genome of an organism. Examples of genomic characteristics include, but are not limited to, those relating to the primary nucleic acid sequences of all or part of the genome (e.g., presence or absence of nucleotide polymorphisms, indels, sequence rearrangements, mutation frequency, etc.), the copy number of one or more detailed nucleotide sequences within the genome (e.g., copy number, allele frequency ratio, monochromacy or ploidy of the entire genome, etc.), the epigenetic state of all or part of the genome (e.g., covalent nucleic acid modifications such as methylation, histone modifications, nucleosome arrangement, etc.), and the expression profile of the genome of an organism (e.g., gene expression levels, isotype expression levels, gene expression ratios, etc.).

[0089] As used herein, the term “healthy” refers to an subject possessing good health. A healthy subject may demonstrate the absence of any malignant or non-malignant disease. A “healthy individual” may have other diseases or conditions unrelated to the disease under assay that would not normally be considered “healthy.”

[0090] As used herein, the term "indel" refers to any insertion or deletion of one or more base pairs having a certain length and position (which may also be called an anchor position) within a sequence read. Insertions correspond to positive lengths, while deletions correspond to negative lengths.

[0091] As used herein, the term “methylation” refers to a modification of deoxyribonucleic acid (DNA) such that a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group to form 5-methylcytosine. More specifically, methylation tends to occur in cytosine and guanine dinucleotides, which are referred to herein as “CpG sites.” In other cases, methylation may occur in cytosine or other nucleotides that are not part of a CpG site; however, these occurrences are rare. cfDNA methylation information can be identified as hypermethylation or hypomethylation, both of which may indicate cancerous conditions.

[0092] As used synonymously herein, the terms “methylated fragment” or “nucleic acid methylated fragment” refer to a sequence of methylation states for each of several CpG sites, determined by methylation sequencing of nucleic acids (e.g., nucleic acid molecules and / or nucleic acid fragments). In a methylated fragment, the position and methylation state for each CpG site in the nucleic acid fragment are determined based on the alignment of sequence reads (e.g., obtained from nucleic acid sequencing) to a reference genome. A nucleic acid methylated fragment includes the methylation state (e.g., a methylation state vector) for each of several CpG sites, which specifies the position of the nucleic acid fragment in the reference genome (e.g., as specified by the position of the first CpG site in the nucleic acid fragment using a CpG index or another similar metric) and the number of CpG sites in the nucleic acid fragment. Alignment of sequence reads to a reference genome based on methylation sequencing of nucleic acid molecules can be performed using a CpG index. As used herein, the term “CpG index” refers to a list of individual CpG sites in a reference genome, such as the human reference genome (e.g., CpG1, CpG2, CpG3, etc.), which may be in electronic form. The CpG index further includes, for each CpG site, the corresponding genomic location in the corresponding reference genome. Thus, for each nucleic acid methylation fragment, each CpG site is indexed to a specific location in the respective reference genome, which can be determined using the CpG index.

[0093] As used herein, the term “reference genome” refers to any specific, known, sequenced, or characterized genome of any organism or virus, whether partially or completely, that may be used as a reference for sequences identified from a subject. Exemplary reference genomes used for human subjects and many other organisms are provided in online genome browsers hosted by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz (UCSC). “Genome” refers to the complete genetic information of an organism or virus as represented by nucleic acid sequences. As used herein, a reference sequence or reference genome is often a genome sequence assembled or partially assembled from one or more individuals. In some embodiments, a reference genome is a genome sequence assembled or partially assembled from one or more human individuals. A reference genome can be considered a representative example of the gene set of a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. Examples of exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).

[0094] As used herein, the terms “sequence read” or “read” refer to a nucleotide sequence produced by any sequencing process described herein or known in the art. Reads may be generated from one end of a nucleic acid fragment (“single-ended read”), or sometimes from both ends of a nucleic acid (e.g., paired-ended read, double-ended read). In some embodiments, sequence reads (e.g., single-ended or paired-ended reads) may be generated from one or both strands of a target nucleic acid fragment. The length of a sequence read is often related to the detailed sequencing technique. For example, high-throughput methods may provide sequence reads that vary in size from tens to hundreds of base pairs (bp). In some embodiments, sequence reads have an average, median, or mean length of approximately 15 bp to 900 bp (e.g., approximately 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 110 bp, 120 bp, 130, 140 bp, 150 bp, 200 bp, 450 bp, 300 bp, 350 bp, 400 bp, 450 bp, or approximately 500 bp). In some embodiments, sequence reads have an average, median, or mean length of approximately 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. Or it is the average length. For example, nanopore sequencing may provide sequence reads that can vary in size from tens to hundreds or thousands of base pairs. Illumina parallel sequencing can provide sequence reads with less variation, for example, sequence reads that can mostly be smaller than 200 bp. A sequence read (or sequencing read) can refer to sequence information (e.g., a nucleotide string) corresponding to a nucleic acid molecule. For example, a sequence read may correspond to a nucleotide string from a portion of a nucleic acid fragment (e.g., about 20 to about 150), to a nucleotide string at one or both ends of a nucleic acid fragment, or to nucleotides from the entire nucleic acid fragment.Sequence reads can be obtained by various methods, such as using sequencing techniques, using probes with hybridization arrays or capture probes, or using amplification techniques such as polymerase chain reaction (PCR) or linear or isothermal amplification using single primers.

[0095] As used herein, the terms “sensitivity” or “true positive rate” (TPR) refer to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can be characterized by the ability of an assay or method to correctly identify the proportion of a population that truly has a certain condition. For example, sensitivity can be characterized by the ability of a method to correctly identify the number of subjects in a population that has cancer. In another example, sensitivity can be characterized by the ability of a method to correctly identify one or more markers that indicate cancer.

[0096] As used herein, the terms “sequencing,” etc., generally refer to any biochemical process that can be used to determine the order of biological macromolecules such as nucleic acids or proteins. Sequencing may include next-generation sequencing, including whole-genome sequencing, whole-exome sequencing, and targeted sequencing. For example, sequencing data may include all or part of the nucleotide bases in a nucleic acid molecule, such as a DNA fragment.

[0097] As used herein, the term “sequencing depth” is used synonymously with the term “coverage” and refers to the number of times a locus is covered by consensus sequence reads corresponding to unique nucleic acid target molecules aligned with that locus; for example, sequencing depth is equal to the number of unique nucleic acid target molecules covering that locus. A locus may be as small as a nucleotide, as large as a chromosome arm, or as large as an entire genome. Sequencing depth can be expressed as “Yx,” for example, 50x, 100x, etc., where “Y” is the number of times the locus is covered by sequences corresponding to nucleic acid targets; for example, the number of times independent sequence information covering a particular locus is obtained. In some embodiments, sequencing depth corresponds to the number of genomes sequenced. Sequencing depth can also be applied to multiple loci or the entire genome, in which case Y may refer to the average number of times the locus, haploid genome, or entire genome is sequenced, respectively. When the mean depth is quoted, the actual depth for the various loci included in the dataset can range from one value to another. Ultra-deep sequencing can refer to a sequencing depth of at least 100x at a given locus.

[0098] As used herein, the term “sequencing panel” refers to a combination of sequencing data from two or more specimens (e.g., individuals).

[0099] As used herein, the terms “single nucleotide variant” or “SNV” refer to a substitution in which one nucleobase at a certain position (e.g., site) on a nucleobase sequence from an organism, such as a sequence read, is replaced by a different nucleobase. A substitution from a first nucleobase X to a second nucleobase Y can be denoted as “X>Y”. For example, an SNV from cytosine to thymine can be denoted as “C>T”.

[0100] As used herein, the term “specificity” or “true negative rate” (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to correctly identify the proportion of a population that does not truly have a particular disease. For example, specificity can characterize the ability of a method to correctly identify the number of subjects in a population that do not have cancer. In another example, specificity can characterize the ability of a method to correctly identify one or more markers that indicate cancer.

[0101] As used herein, the term “subject” means any living or non-living organism, including, but not limited to, humans (e.g., male humans, female humans, fetuses, pregnant women, children, etc.), non-human animals, plants, bacteria, fungi, or protists. Any human or non-human animal may act as a subject, including, but not limited to, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cattle), equids (e.g., horses), goats and sheep (e.g., sheep, goats), pigs (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursids (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. In some embodiments, the subject is a male or female at any stage (e.g., male, female, or child). The subjects from whom specimens are taken, or the subjects treated by any of the methods or compositions described herein, may be of any age and may be adults, infants, or children.

[0102] As used herein, the term “tissue” may refer to a group of cells that function together as a single unit. A single tissue may contain two or more types of cells. Different types of tissue may consist of different types of cells (e.g., hepatocytes, alveolar cells, or blood cells), and may also refer to tissue from different organisms (mother or fetus), or to healthy cells or tumor cells. Generally, the term “tissue” may refer to any group of cells found in the human body (e.g., cardiac tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some embodiments, the terms “tissue” or “tissue type” may be used to refer to the tissue from which cell-free nucleic acids originate. In one example, viral nucleic acid fragments may originate from blood tissue. In another example, viral nucleic acid fragments may originate from tumor tissue.

[0103] As used herein, the term “true positive” (TP) refers to an object that meets a certain condition. A “true positive” may refer to an object that has a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), localized or metastatic cancer, or a non-malignant disease. A “true positive” may refer to an object that meets a certain condition and is identified as having that condition by the assay or method of this disclosure. As used herein, the term “true negative” (TN) refers to an object that does not meet a certain condition or does not have a certain detectable condition. A “true negative” may refer to an object that does not have a disease or detectable disease, such as a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), localized or metastatic cancer, or a non-malignant disease, or an object that is otherwise healthy. A “true negative” may refer to an object that does not meet a certain condition or does not have a certain detectable condition, or is identified as not having that condition by the assay or method of this disclosure.

[0104] As used herein, the term “leukocyte DNA” or “wbcDNA” refers to nucleic acids, including chromosomal DNA, originating from leukocytes. Generally, wbcDNA is assumed to be gDNA and healthy DNA.

[0105] The terminology used herein is intended to describe specific cases only and is not intended to be limiting. When used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context specifically indicates otherwise. Furthermore, the terms “including,” “includes,” “having,” “has,” “with,” or their variations thereof, are intended to be as inclusive as the term “comprising,” insofar as they are used in any of the detailed description and / or claims.

[0106] IC Exemplary Analytics System Figure 2A is an illustrative flowchart of a nucleic acid sample sequencing apparatus according to one or more embodiments. This illustrative flowchart includes instruments such as a sequencer 220 and an analytics system 200. The sequencer 220 and the analytics system 200 can work together to perform one or more steps of the process.

[0107] In various embodiments, the sequencer 220 receives the concentrated nucleic acid sample 210. As shown in Figure 2A, the sequencer 220 may include a graphical user interface 225 that allows the user to interact with specific tasks (e.g., start sequencing or end sequencing), and another loading station 230 for loading a sequencing cartridge containing the concentrated fragment sample and / or loading the buffers necessary to perform the sequencing assay. Thus, once the user of the sequencer 220 has finished providing the necessary reagents and sequencing cartridge to the loading station 230 of the sequencer 220, the user can start sequencing by interacting with the graphical user interface 225 of the sequencer 220. After start, the sequencer 220 performs sequencing and outputs sequence reads of the concentrated fragment from the nucleic acid sample 210.

[0108] In some embodiments, the sequencer 220 is communicatively coupled to an analytics system 200. The analytics system 200 comprises several computing devices used for processing sequence reads for various applications, such as variant calling or panel assignment optimization. The sequencer 220 may provide the analytics system 200 with sequence reads in BAM file format. The analytics system 200 may be communicatively coupled to the sequencer 220 by wireless, wired, or a combination of wireless and wired communication technologies. Generally, the analytics system 200 comprises a processor and a non-temporary computer-readable storage medium that stores computer instructions causing the processor to process sequence reads or perform one or more steps of any of the methods or processes disclosed herein when the processor is running.

[0109] In some embodiments, alignment position information can be determined by aligning a sequence read with a reference genome using methods known in the art. The alignment position can generally describe the start and end positions of a region in the reference genome corresponding to the start and end nucleotide bases of a given sequence read. By generalizing the alignment position information in accordance with methylation sequencing, the first and last CpG sites included in the sequence read according to the alignment with the reference genome can be indicated. The alignment position information can further indicate the methylation status and location of all CpG sites in a given sequence read. The region in the reference genome may be associated with a gene or gene segment; thus, the analytics system 200 can label a sequence read with one or more genes that it aligns with. In one embodiment, the fragment length (or size) is determined from the start and end positions.

[0110] In various embodiments, for example, when a paired-end sequencing process is used, the sequence reads consist of a read pair denoted as R_1 and R_2. For example, the first read R_1 may be sequenced from the first end of a double-stranded DNA (dsDNA) molecule, while the second read R_2 may be sequenced from the second end of the double-stranded DNA (dsDNA). Thus, the nucleotide base pairs of the first read R_1 and the second read R_2 may be aligned to match the nucleotide bases of the reference genome (for example, in the opposite direction). The alignment position information obtained from the read pair R_1 and R_2 may include the start position in the reference genome corresponding to the end of the first read (e.g., R_1) and the end position in the reference genome corresponding to the end of the second read (e.g., R_2). In other words, the start and end positions in the reference genome may represent the expected positions of the nucleic acid fragments within the corresponding reference genome. For further analysis, output files in SAM (Sequence Alignment Map) format or BAM (Binary) format may be generated and output.

[0111] Referring now to Figure 2B, which is a block diagram of an analytics system 200 for processing DNA samples according to one embodiment. The analytics system implements one or more computing devices used for analyzing DNA samples. The analytics system 200 comprises a sequence processor 240, a sequence database 245, a model database 255, a model 250, a parameter database 265, and a score engine 260. In some embodiments, the analytics system 200 performs some or all of the processes described throughout this disclosure.

[0112] Furthermore, multiple different models 250 can be stored in the model database 255 or retrieved for use with test samples. In one example, the model is a trained cancer classifier that determines cancer predictions for test samples using feature vectors derived from information fragments. The training and use of cancer classifiers will be discussed further in Section IV, Cancer Classifiers for Determining Cancer. The analytics system 200 trains one or more models 250 and stores various trained parameters in the parameter database 265. The analytics system 200 stores the models 250 along with their functions in the model database 255.

[0113] The score engine 260 returns an output using one or more models 250 during inference. The score engine 260 accesses the models 250 in the model database 255, along with trained parameters from the parameter database 265. For each model, the score engine receives the appropriate input for that model and calculates the output based on the received input, parameters, and a function of each model on the input and output. In some use cases, the score engine 260 further calculates metrics correlated with the confidence in the output calculated from the model. In other use cases, the score engine 260 calculates other intermediate values ​​to be used by the model.

[0114] ID specimen sequencing and processing Figure 3 is an illustrative flowchart illustrating a process 300 for obtaining sequence reads by sequencing fragments of cfDNA according to one or more embodiments. The analytics system first obtains a sample 310 from an individual containing multiple cfDNA molecules. In further embodiments, process 300 may be applied to the sequencing of other types of DNA molecules. Process 300 is an embodiment of the sample sequencing 120 in Figure 1.

[0115] The analytics system can isolate each cfDNA molecule from a sample.310 A sample can be any subset of the human genome, including the whole genome. A sample can be taken from a subject who is known to have cancer, suspected to have cancer, or whose cancer status is unknown. Examples of samples include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. In some cases, a sample may be tissue or bodily fluids extracted from tissue. In some embodiments, a method for taking a blood sample (e.g., by syringe or finger prick) may be less invasive than a procedure for obtaining a tissue biopsy, which may require surgery. Examples of samples extracted include cfDNA and / or ctDNA. In healthy individuals, the human body can naturally expel cfDNA and other cellular debris. If the subject has cancer or disease, ctDNA in the extracted sample may be present at diagnostically detectable levels.

[0116] In addition, wbcDNA can be mentioned as a sample to be extracted. Extracting a nucleic acid sample may further include separating cfDNA and / or ctDNA from wbcDNA. Extraction of wbcDNA from cfDNA and / or ctDNA may be performed when the DNA is separated from the sample. In the case of a blood sample, wbcDNA is obtained from the buff coat fraction of the blood sample. Shearing of wbcDNA may yield swbcDNA fragments less than 300 base pairs long. Separating wbcDNA from cfDNA and / or ctDNA makes it possible to sequence wbcDNA independently of cfDNA and / or ctDNA. Generally, the sequencing process for wbcDNA is similar to that for cfDNA and / or ctDNA.

[0117] A sequencing library can be prepared from the converted cfDNA molecule.330 During library preparation, a unique molecular identifier (UMI) can be added to the nucleic acid molecule (e.g., DNA molecule) by adapter ligation. The UMI may be a short nucleic acid sequence (e.g., 4-10 base pairs) attached to the end of a DNA fragment (e.g., a DNA molecule fragmented by physical shear, enzymatic digestion, and / or chemical fragmentation) during adapter ligation. The UMI may also be denatured base pairs that act as a unique tag that can be used to identify sequence reads originating from a particular DNA fragment. During PCR amplification after adapter ligation, the UMI can be replicated along with the attached DNA fragment. This may provide a method for identifying sequence reads arising from fragments of the same origin in downstream analysis.

[0118] Optionally, the sequencing library may be enriched with cfDNA molecules or genomic regions that provide information about cancer status using multiple targeting probes.335 Targeting probes are short oligonucleotides that have the ability to hybridize to a specially designated cfDNA molecule or target region to enrich those fragments or regions for subsequent sequencing and analysis. Targeting probes can be used to perform targeted, high-depth analysis of a designated set of CpG sites of interest to the researcher. Targeting probes can be tiled across one or more target sequences with coverage of 1X, 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, or more than 10X. For example, a hybridization probe tiled with 2X coverage contains probes that overlap so that each portion of the target sequence hybridizes into two independent probes. Targeting probes can be tiled across one or more target sequences with coverage of less than 1X.

[0119] In one embodiment, a targeting probe is designed to enrich DNA molecules that have undergone a conversion process from unmethylated cytosine to uracil (e.g., using a bisulfite). During enrichment, the targeting probe (also referred to herein as the “probe”) can be used to target and pull down nucleic acid fragments that provide information about the presence or absence of cancer (or disease), cancer status, or cancer classification (e.g., cancer class or tissue of origin). The probe may be designed to anneal to (or hybridize with) a target (complementary) strand of DNA. The target strand may be a “plus” strand (e.g., the strand transcribed into mRNA and subsequently translated into protein) or a complementary “minus” strand. The probe may range in length from tens, hundreds, or thousands of base pairs. The probe may be designed based on a panel of targeted sequences. The probe may be designed based on a panel of target genes for analyzing detailed mutations or target regions of the genome (e.g., human or another organism) suspected to correspond to a particular cancer or other type of disease. Furthermore, the probe may cover overlapping portions of target regions.

[0120] Once preparation is complete, a sequencing library or a portion thereof can be sequenced to obtain 340 or more sequence reads. The sequence reads may be in a computer-readable digital format for processing and interpretation by computer software. Alignment position information can be determined by aligning these sequence reads with a reference genome. Alignment position information may indicate the start and end positions of a region in the reference genome corresponding to the start and end nucleotide bases of a given sequence read. Alignment position information may also include the sequence read length, which may be determined from the start and end positions. The region in the reference genome may be associated with a gene or gene segment. Sequence reads may consist of read pairs denoted as R1 and R2. For example, the first read R1 may be sequenced from the first end of a nucleic acid fragment, while the second read R2 may be sequenced from the second end of the nucleic acid fragment. Thus, the nucleotide base pairs of the first read R1 and the second read R2 may be aligned with the nucleotide bases of the reference genome (for example, in the reverse direction). The alignment position information obtained from the read pair R1 and R2 may include the start position in the reference genome corresponding to the end of the first read (e.g., R1) and the end position in the reference genome corresponding to the end of the second read (e.g., R2). In other words, the start and end positions in the reference genome may represent the expected positions of the nucleic acid fragments within the corresponding reference genome. For further analysis, output files in SAM (Sequence Alignment Map) format or BAM (Binary) format may be generated and output.

[0121] One or more alternative sequencing methods can be used to obtain sequence reads from nucleic acids in biological specimens. Such one or more sequencing methods may include, but are not limited to, any form of sequencing that can be used to obtain multiple sequence reads measured from nucleic acids (e.g., cell-free nucleic acids), including high-throughput sequencing systems such as the Roche 454 platform, Applied Biosystems SOLID platform, Helicos True single-molecule DNA sequencing technology, sequencing-bi-hybridization platforms from Affymetrix Inc., single-molecule real-time (SMRT) technology from Pacific Biosciences, sequencing-bi-synthesis platforms from 454 Life Sciences, Illumina / Solexa, and Helicos Biosciences, and sequencing-bi-ligation platforms from Applied Biosystems. ION TORRENT technology and nanopore sequencing from Life Technologies can also be used to obtain sequence reads from nucleic acids (e.g., cell-free nucleic acids) in biological specimens. Sequencing-by-synthesis and reversible terminator-based sequencing (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ 4500 (Illumina, San Diego Calif.)) can be used to obtain sequence reads from cell-free nucleic acids obtained from training specimens to form genotype datasets. Millions of cell-free nucleic acid (e.g., DNA) fragments can be sequenced in parallel. An example of this type of sequencing technique uses a flow cell containing an optically transparent slide with eight separate lanes, to which oligonucleotide anchors (e.g., adapter primers) are bound. Cell-free nucleic acid specimens may contain signals or tags to facilitate detection.Obtaining sequence reads from cell-free nucleic acids obtained from biological specimens may involve obtaining quantitative information on signals or tags through various techniques, such as flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene tip analysis, microarrays, mass spectrometry, cellular fluorescence quantitative analysis, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, field suspension, sequencing, and combinations thereof.

[0122] One or more sequencing methods may include whole-genome sequencing assays. Whole-genome sequencing assays may include physical assays that generate sequence reads for the entire genome or a substantial portion of the entire genome that can be used to determine large variability, such as copy number variations or copy number anomalies. Such physical assays may utilize whole-genome sequencing techniques or whole-exome sequencing techniques. Whole-genome sequencing assays may have an average sequencing depth of at least 1x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, at least 20x, at least 30x, or at least 40x across the genome under test. In some embodiments, the sequencing depth is approximately 30,000x. One or more sequencing methods may include targeted sequencing panel assays. A targeted sequencing panel assay may have an average sequencing depth of at least 50,000x, at least 55,000x, at least 60,000x, or at least 70,000x for the targeted gene panel. The targeted gene panel may contain 450 to 500 genes. The targeted gene panel may contain genes in the range of 500 ± 5, 500 ± 10, or 500 ± 25.

[0123] One or more sequencing methods may include paired-end sequencing. One or more sequencing methods may generate multiple sequence reads. Multiple sequence reads may have average lengths in the range of 10–700, 50–400, or 100–300. One or more sequencing methods may include methylation sequencing assays. Methylation sequencing may be i) whole-genome methylation sequencing or ii) targeted DNA methylation sequencing using multiple nucleic acid probes. For example, methylation sequencing is whole-genome bisulfite sequencing (e.g., WGBS). Methylation sequencing may be targeted DNA methylation sequencing using multiple nucleic acid probes that target the most informative regions of the methylome, a unique methylation database, and pre-prototype whole-genome and targeted sequencing assays.

[0124] II. Panel Assignment The analytics system assigns samples to a target variant sequencing panel. Optimizing panel assignment is advantageous because it improves the sensitivity of variant calling from sequencing data obtained from the target variant sequencing panel, and minimizes sequencing material and cost by making the samples densest on the panel. Furthermore, optimizing panel assignment can make variant coverage across samples in a single panel more uniform, thereby improving sensitivity in variant calling. In addition, the analytics system can group variants with similar noise rates into a single panel, preventing missed calls among variants with large noise rate variances. Finally, optimizing for denser sampling on the panel aims to minimize the number of panels required to perform target sequencing for a given sample set. Minimizing the number of panels inevitably minimizes sequencing cost and material.

[0125] Figure 4 is a flowchart of an optimized sequencing panel assignment 400 according to one or more embodiments. The analytics system 200 may perform this optimized sequencing panel assignment method. In other embodiments, the analytics system 200 functions in conjunction with a sequencer 220. For example, the analytics system 200 may perform one or more analyses on the sequencing data, while the sequencer 220 may perform one or more sequencing steps or any other physical assay steps.

[0126] The analytics system 200 obtains initial sequencing data for a sample set to be assigned to a target variant sequencing panel 410. The initial sequencing data may cover representative variants or genomic regions, for example, sequenced at a low sequencing depth. In other embodiments, the initial sequencing data may cover a majority or all of the genome, for example, whole-genome sequencing or whole-exome sequencing. The analytics system may determine characteristics for each variant or genomic region, such as guanine-cytosine content, targeting probe error rate, sequencing depth count, presence or absence of variants, mean allele frequency, total number of minor variants, and allele frequency of true variants.

[0127] The analytics system 200 determines one or more characteristics of each sample based on the initial sequencing data 420. As one characteristic, the analytics system may determine and identify candidate variants for each sample. In connection with this, the analytics system may determine the total number of candidate variants present in the sample. The analytics system may further aggregate the characteristics of each identified candidate variant under that sample. For example, the first sample has a total of 150 candidate variants, where the characteristics of each variant are listed together with the sample.

[0128] The analytics system 200 applies a panel assignment model to the characteristics of the samples 430 to determine the panel assignment for each sample. In one or more embodiments, the panel assignment model evaluates a subset of one or more characteristics. For example, the panel assignment model may determine the number of variants per sample and the GC content of the variants in the sample. As another example, the panel assignment model may determine the error rate of the variants in the sample in addition to the two characteristics above. The panel assignment model may determine the panel assignment for each sample according to an objective function (or, alternatively, a loss function). The objective function (or, alternatively, a loss function) may be heuristically defined to score the panel assignment based on one or more criteria. For example, when considering the number of variants in each sample and the GC content of the variants in the sample, the objective function may assign a higher score to uniformity of panel size across all panels (e.g., the number of variants to be considered) and a higher score to uniformity of GC content of variants within each panel. Alternatively, in the implementation of the loss function, a large variance in the uniformity of panel sizes may lead to a higher loss, and a large variance in the GC content within each panel may also lead to a higher loss. In one or more embodiments, the panel assignment model applies a greedy algorithm to determine the panel assignment for each sample. In some embodiments, the panel assignment model may use a dynamic programming algorithm to determine the panel assignment.

[0129] Criteria can also be ranked by assigning different weights to their contributions to the objective function (or, alternatively, the loss function). For example, panel homogeneity is given first priority, and therefore homogeneity receives a greater weight to the objective function compared to GC content, which is given second priority. Different weighting of criteria considered by the function can provide the possibility of modularization in panel assignment models. For example, in one run of the panel assignment model, one set of weights may prioritize some criteria, while in a second run of the same panel assignment model, a different set of weights may prioritize other criteria. This modularity allows the panel assignment model to be flexibly applied to panel assignments of different sample sets.

[0130] In a further embodiment, the panel assignment model may perform a swap operation at the time the samples are initially assigned to the panels to determine improvements in the optimization of the panel assignments. The swap operation may identify panel assignments that negatively impact the objective function and / or significantly contribute to the loss function as swapping candidates. The swap operation iteratively evaluates the change in score (objective or loss) when swapping the panel assignments of samples from their assigned panels to different panels. If the swap improves the overall optimization (maximizes the objective score and / or minimizes the loss), the swap operation makes the swap formal, i.e., selects a panel assignment for the first sample and assigns it to the second sample, and vice versa. The swap operation may iteratively evaluate and swap samples until a stopping condition is met. One stopping condition is the maximum total number of swaps. Another stopping condition is when the improvement of swaps approaches zero, i.e., when swaps no longer have a significant impact on the score. A combined stop condition may consider either the maximum total number of swaps or the swap improvement approaching zero. For example, if either condition is met, the swap operation will stop.

[0131] In one or more embodiments, the panel assignment model determines panel assignments by one or more hard constraints. Hard constraints are constraints that must not be broken; that is, the panel assignment model must satisfy hard constraints when determining panel assignments. For example, one hard constraint is the number of specimens per panel. In another example, another hard constraint is the number of one or more types of specimens per panel. In this embodiment, the specimens assigned to panels may be of different types. For example, specimens tested for tumor ratios may be considered “test specimens,” while specimens with known tumor ratios (i.e., tissue biopsy specimens titrated to the exact tumor ratio) may be considered “benchmark specimens.” One hard constraint may limit the number of benchmark specimens that can be assigned to any given panel to, for example, at most one per panel. Another hard constraint may involve placing each benchmark specimen in some number of panels, for example, exactly three panels.

[0132] In some embodiments, the analytics system 200 selects a set of core variants to be screened for one or more samples.440 In some embodiments, one or more of those samples may have more variants to be screened by the target sequencing panel than a threshold number. For example, the target sequencing panel may have an upper limit of 500 variants to evaluate for any given sample. For a sample with 600 variants, the analytics system 200 identifies a subset of those 600 variants to be screened with respect to the assigned panel. To identify this subset, the analytics system 200 may apply a selection algorithm. The selection algorithm may utilize a function that scores the variants in a sample based on the characteristics of the variants. In one or more embodiments, this function may be similar to the objective function and / or loss function of a panel assignment model. The function may determine the GC content of a variant, the error rate of the associated probe, the location of the variant in the genome, the sequencing depth of the variant, the presence or absence of the variant, the mean allele frequency, the total number of minor variants, the allele frequency of the true variant, or a combination of some of these. The selection algorithm may attempt to optimize the selected variants in the context of other samples on the panel. For example, if a variant is so frequent that it is present in a large proportion of samples on the panel, that variant may be given a high weight as having a sequencing cost-saving benefit. Such a benefit may be balanced with other factors, such as the sensitivity of variant calling.

[0133] The analytics system 200 returns a panel assignment for each sample in the sample set 450. Each panel contains the variant to be screened and the assigned sample. The panel may further contain probes that target the identified variant.

[0134] Sequencer 220 can perform target sequencing on the assigned panel 460. For example, target sequencing is further explained in Section ID, "Sample Sequencing and Processing".

[0135] III. Variant Calling The analytics system can call variants from sequencing data for a particular sample. In detail, Figure 5 is a flowchart of the workflow for determining variants of sequence reads according to one embodiment. The sequencing data may be obtained, for example, from target sequencing by a panel generated by a panel assignment model, as described in Section II, "Panel Assignment".

[0136] In some embodiments, the analytics system 200 uses workflow 500 to perform variant calling (e.g., for SNVs and / or indels) based on input sequencing data. Furthermore, the analytics system 200 can obtain input sequencing data from output files associated with nucleic acid samples prepared using workflow 100 described above. Workflow 500 includes, but is not limited to, the following steps, which will be described in relation to the components of the analytics system 200. In other embodiments, one or more steps of workflow 500 can be replaced with steps of a different process for generating variant calls using a variant call format (VCF), such as HaplotypeCaller, VarScan, Strelka, or SomaticSniper.

[0137] In step 510, the analytics system 200 collapses the aligned sequence reads of the input sequencing data. In one embodiment, collapsing sequence reads involves collapsing multiple sequence reads into a consensus sequence using UMI and, optionally, alignment position information from the sequencing data of an output file (e.g., from workflow 100 shown in Figure 1) to determine the most likely sequence or portion thereof of a nucleic acid fragment. Since the UMI is replicated through enrichment and PCR along with the ligated nucleic acid fragment, the sequence processor can determine that a particular sequence read originated from the same molecule in the nucleic acid sample. In some embodiments, sequence reads having the same or similar alignment position information (e.g., start and end positions within a threshold offset) and containing a common UMI are collapsed, and the sequence processor generates a collapsed read (also referred to herein as a consensus read) representing that nucleic acid fragment. The sequence processor designates a consensus read as “duplex” if the corresponding pair of collapsed reads have a common UMI, indicating that both the positive and negative strands of the originating nucleic acid molecule are captured; otherwise, the collapsed read is designated as “non-duplex.” In some embodiments, the sequence processor may perform other types of error correction on the sequence reads in lieu of or in addition to collapsing the sequence reads.

[0138] In step 515, the analytics system 200 stitches the collapsed reads based on the corresponding alignment position information. In some embodiments, the sequence processor determines whether the nucleobase base pairs of the first and second reads overlap in the reference genome by comparing the alignment position information between the first and second reads. In one use case, in response to a determination that the overlap between the first and second reads (e.g., of a given number of nucleobase bases) is longer than a threshold length (e.g., a threshold number of nucleobase bases), the sequence processor designates the first and second reads as “stitched”; otherwise, the collapsed reads are designated as “not stitched”. In some embodiments, the first and second reads are stitched if the overlap is longer than a threshold length and the overlap is not a sliding overlap. For example, a sliding overlap may include a homopolymer run (e.g., a single repeating nucleobase), a dinucleobase run (e.g., a 2-nucleobase sequence), or a trinucleobase run (e.g., a 3-nucleobase sequence), where the homopolymer run, dinucleobase run, or trinucleobase run has at least a threshold length of base pairs.

[0139] In step 520, the analytics system 200 assembles the reads into a pass. In some embodiments, the analytics system 200 assembles the reads to generate a directed graph, e.g., a de Bruijn graph, for the target region (e.g., a gene). The unidirectional edges of the directed graph represent sequences of k-nucleobase bases (also referred herein as "k-mers") in the target region, and these edges are connected by vertices (or nodes). The sequence processor aligns the collapsed reads with the directed graph so that any of the collapsed reads can be represented in order by a subset of edges and corresponding vertices.

[0140] In some embodiments, the analytics system 200 determines a set of parameters that describe a directed graph and processes the directed graph. In addition, the parameter set may include counts of k-mers from collapsed reads that have successfully aligned with k-mers represented by nodes or edges in the directed graph. The analytics system 200 stores the directed graph and the corresponding parameter set in, for example, an array database 245, which can be retrieved for updating the graph or generating a new graph. For example, the analytics system 200 can generate a compressed version of the directed graph based on the parameter set (for example, or modify an existing graph). In one use case, to filter out less important data in the directed graph, the array processor removes nodes or edges with counts below a threshold (e.g., "trimming" or "pruning") and retains nodes or edges with counts above the threshold.

[0141] In one embodiment, the analytics system 200 can store sequencing data (e.g., variants and normals) in the sequence database 245, which can be used to detect the presence, absence, or level of feature values ​​(e.g., GC content, error rate, sequence depth count, presence or absence of variants, mean allele frequency, total number of minor variants, and allele frequency of true variants) in a sample from a subject, and / or otherwise predict the cost associated with the variants (e.g., cost value per site and relative cost). The sequence database 245 can also store sequencing data processed by the analytics system 200, but can also store sequencing data not processed by the analytics system 200, such as sequencing data uploaded from an external source and / or obtained from an external or publicly available database in other ways.

[0142] In step 525, the analytics system 200 generates candidate variants from the paths assembled by the sequence processor 240. In one embodiment, the analytics system 200 generates candidate variants by comparing a directed graph (which may have been compressed by pruning edges or nodes in step 510) with a reference sequence of the target region of the genome. The analytics system 200 can align the edges of the directed graph with the reference sequence and record the genomic locations of mismatched edges and mismatched nucleobase bases adjacent to those edges as the locations of the candidate variants. In addition, the analytics system 200 can generate candidate variants based on the sequence depth of the target region. Specifically, the analytics system 200 may have greater confidence in identifying variants with greater sequence depth in the target region because, for example, it will resolve mismatches or other base pair variations between sequences (e.g., using redundancy) with the help of a larger number of sequence reads.

[0143] In one embodiment, the analytics system 200 generates candidate variants using a variant model that determines the expected noise rate for sequence reads from a subject. The variant model may be a Bayesian hierarchical model, however, in some embodiments, the analytics system 200 uses one or more different types of models. Furthermore, the Bayesian hierarchical model may be one of many possible model architectures that can be used to generate candidate variants, and such models are all related in that they model location-specific noise information to improve the sensitivity / specificity of variant calling. More specifically, the analytics system 200 may train the variant model using samples from healthy individuals to model the expected noise rate for each location of sequence reads.

[0144] Furthermore, the model database 255 can store multiple different models, or retrieve them for post-training application. For example, a first model can be trained to model the SNV noise rate, and a second model can be trained to model the indel noise rate. Additionally, the score engine 260 can use the parameters of the variant model to determine the likelihood of one or more true positives in the sequence reads. Based on that likelihood, the score engine 260 can determine a quality score (e.g., on a logarithmic scale). For example, the quality score could be a Phred quality score: Q = -10·log 10 P This is the case (where P is the likelihood of an incorrect candidate variant call (e.g., a false positive)).

[0145] In step 530, the score engine 260 assigns a score to the candidate variants based on the likelihood of the variant model or the corresponding true positive or quality score.

[0146] In step 535, the analytics system 200 outputs candidate variants. In some embodiments, the analytics system 200 outputs some or all of the determined candidate variants along with their corresponding scores. For example, an external downstream system or other component of the analytics system 200 may use these candidate variants and scores for a variety of applications, including, but not limited to, predicting the presence of cancer, disease, or germline mutations.

[0147] In one embodiment, candidate variants are output for both cfDNA and / or ctDNA and wbcDNA. Here, generally, candidate variants for wbcDNA are "normal," while candidate variants for cfDNA and / or ctDNA are "variants." Various detection methods and models can be used to compare variants to normals to determine whether a variant contains a signature for cancer or any other disease. In various embodiments, normals and variants may be generated using any other process, any number of samples (e.g., tumor biopsies or blood samples), or accessed from a database storing the candidate variants.

[0148] In one embodiment, the output candidate variant is used in a method described herein for generating an optimized sequencing panel assignment.

[0149] For further details regarding variant calling, see U.S. Patent Application No. 16 / 119,961, filed August 31, 2018, entitled “Identifying False Positive Variants Using a Significance Model” (both incorporated herein by reference). For further details regarding the identification of copy number aberrations or copy number variants, see U.S. Patent Application No. 15 / 853,314, filed December 22, 2017, entitled “Base Coverage Normalization and Use Thereof in Detecting Copy Number Variation” and U.S. Patent Application No. 16 / 352,214, filed March 13, 2019, entitled “Identifying Copy Number Aberrations” (both incorporated herein by reference).

[0150] IV. Cancer classifier Cancer classification involves extracting genetic features and determining cancer predictions by applying one or more models to the extracted features. The analytics system aggregates the extracted features into a feature vector, which is then input into a trained cancer prediction model, and cancer predictions are determined based on the input feature vector. Cancer predictions may include one or more labels and / or one or more values. One label may be binary and indicate the presence or absence of cancer in the test subject. Another label may be multiclass and indicate one or more specific cancer types from multiple cancer types being screened. One value may indicate the likelihood of cancer presence. Another value may indicate the likelihood of cancer absence. Yet another value may indicate a different prognosis for cancer in other cases. For example, a value may quantify cancer progression and / or malignancy. Furthermore, one embodiment of the cancer classifier may output a predicted tumor ratio in the sample. Such cancer predictions can be useful in monitoring cancer progression, evaluating the effectiveness of treatment, detecting cancer recurrence, and measuring minimal residual disease.

[0151] In some embodiments, the cancer classifier may be a post-machine learning model that includes multiple classification parameters and a function representing the relationship between a feature vector as input and a cancer prediction as output. When the feature vector is input to the function along with the classification parameters, a cancer prediction is generated. The post-machine learning model may be trained using training samples derived from individuals with known cancer diagnoses. The training samples may be divided into cohorts of various labels. For example, there may be a cohort of training samples for each cancer type. As another example, the training samples may have known tumor ratios through titration of tissue biopsy samples (also referred to as “benchmark samples”).

[0152] IV.A. Training of the cancer classifier Figure 6A is a flowchart illustrating the training process 600 of a cancer classifier according to one embodiment. The analytics system obtains multiple training samples 610, each training sample containing a set of called variants and a cancer type label. The multiple training samples may include any combination of samples from healthy individuals with a general label of “non-cancerous” and samples from subjects with a general label of “cancer” or a detailed label (e.g., “breast cancer,” “lung cancer,” etc.). The labels may also indicate tumor ratios, expressed as proportions, for example. For a given cancer type, the training samples from subjects may be referred to as a cohort of that cancer type or a cancer type cohort.

[0153] The analytics system determines a feature vector for each training sample based on the called variants in the training sample.620 This feature may be binary and may indicate the presence or absence of a variant. One feature may indicate the proportion of sequence reads in the sample's sequencing data that indicate the presence of a variant, or another feature may indicate the proportion that indicates the absence of a variant. In some embodiments, the feature may relate to the allele frequency of the variant. In some embodiments, the feature may aggregate the count or proportion of sequence reads containing the called variant in a segmented genomic region. Other types of features based on called variants may also be implemented.

[0154] Once all features have been determined for sample training, the analytics system may determine a feature vector as a vector of elements, each containing one feature. The analytics system may then normalize the features based, for example, on the sequence depth or coverage of each variant.

[0155] For further methods of feature-finding samples, see U.S. Patent Application No. 16 / 384,784, entitled "Multi-Assay Prediction Model for Cancer Detection," and U.S. Patent Application No. 16 / 579,805, entitled "Mixture Model for Targeted Sequencing" (these are incorporated by reference as a whole).

[0156] Using the feature vectors of the training sample, the analytics system can train a cancer classifier in one of several ways. In one embodiment, the analytics system trains a binary cancer classifier to distinguish between cancerous and non-cancerous samples based on the feature vectors of the training sample.620 In this way, the analytics system uses a training sample that includes both non-cancerous samples from healthy individuals and cancerous samples from subjects. Each training sample may have one of two labels: "cancer" or "non-cancerous". In this embodiment, the classifier outputs a cancer prediction indicating the likelihood of the presence or absence of cancer.

[0157] In another embodiment, the analytics system trains a multi-class cancer classifier to distinguish between many cancer types (also referred to as tissue of origin (TOO) labels).620 The cancer types may include one or more cancers, and may include non-cancer types (and may also include any further other diseases or genetic disorders, etc.). To this end, the analytics system can use cancer type cohorts, and may or may not include non-cancer type cohorts. In this multi-cancer embodiment, the cancer classifier is trained to determine cancer predictions (or more specifically TOO predictions) that include predicted values ​​for each of the cancer types under classification. The predicted values ​​may correspond to the likelihood that a given training sample (and test sample during inference) has each of the cancer types. In one embodiment, the predicted values ​​are scored between 0 and 100, where the sum of the predicted values ​​is equal to 100. For example, the cancer classifier returns cancer predictions that include predicted values ​​for breast cancer, lung cancer, and non-cancer. For example, a classifier might return a cancer prediction that the test sample has a 65% likelihood of breast cancer, a 25% likelihood of lung cancer, and a 10% likelihood of being non-cancerous. The analytics system could further evaluate the predictions to generate a prediction of the presence of one or more cancers in the sample, which may also be called a TOO prediction, indicating one or more TOO labels, e.g., a first TOO label which is the highest prediction, a second TOO label which is the second highest prediction, and so on. Continuing the above example, given these proportions, the system might determine that the sample has breast cancer, given that breast cancer has the highest likelihood.

[0158] In a third embodiment, the analytics system determines a tumor ratio based on the feature vector of a training sample by training a cancer classifier. In such an embodiment, the tumor ratio indicates the amount of tumor signal in the sample, which can, for example, serve as a surrogate indicator of cancer progression. To train such a classifier, the analytics system may train the classifier in a supervised manner, inputting the feature vector of a training sample and predicting a known tumor ratio for the training sample. A training sample with such a known tumor ratio can be referred to as a benchmark sample. The cancer classifier may be trained as a machine learning regression model.

[0159] In various embodiments, the analytics system inputs a set of training samples along with their feature vectors into the cancer classifier and trains the cancer classifier by adjusting the classification parameters so that the classifier's function accurately relates the training feature vectors to their corresponding labels (e.g., cancer status, cancer type, or tumor ratio). The analytics system may group the training samples into one or more sets of training samples for iterative batch training of the cancer classifier. After inputting all sets of training samples, including their training feature vectors, and adjusting the classification parameters, the cancer classifier may be sufficiently trained to label test samples according to their feature vectors within some margin of error. The analytics system may train the cancer classifier according to one of several methods. For example, a binary cancer classifier may be an L2 regularized logistic regression classifier trained using a log-loss function. Another example is a multi-cancer classifier, which may be a multinomial logistic regression. In practice, any type of cancer classifier may be trained using other techniques. There are many such techniques, including the use of machine learning algorithms such as kernel methods, random forest classifiers, mixture models, autocoding models, and multilayer neural networks.

[0160] Examples of classifiers include logistic regression algorithms, neural network algorithms, support vector machine algorithms, naive Bayesian algorithms, nearest neighbor algorithms, boosting tree algorithms, random forest algorithms, decision tree algorithms, multinomial logistic regression algorithms, linear models, or linear regression algorithms.

[0161] IV.B. Deployment of the Cancer Classifier When using a cancer classifier, the analytics system may obtain test samples from a target of an unknown cancer species. The analytics system can process these test samples, which consist of DNA molecules, and call variants in the genetic data of the test samples. The analytics system can determine the test feature vector to be used by the cancer classifier, following a similar principle to that considered in process 600. The analytics system can generate a feature vector for the test samples to be input into the cancer classifier. For example, the cancer classifier receives information about approximately 2,000 variants in the human genome as an input feature vector. Therefore, based on the variants called for the test samples, the analytics system can determine a test feature vector that encompasses the features of those 2,000 variants.

[0162] Next, the analytics system can input the test feature vector into a cancer classifier. The cancer classifier's function can then generate cancer predictions based on the classification parameters and test feature vectors trained in process 600. In the first method, the cancer prediction may be binary, selected from the group consisting of "cancer" or "non-cancer". In the second method, the cancer prediction is selected from a group of many cancer types and "non-cancer". In the third method, the cancer prediction indicates the tumor ratio present in the sample. In a further embodiment, the cancer prediction has a prediction value for each of many cancer types. Furthermore, the analytics system may determine that the test sample is most likely to be one of the cancer types. Following the above example, where the cancer prediction for the test sample is breast cancer with a 65% likelihood, lung cancer with a 25% likelihood, and non-cancer with a 10% likelihood, the analytics system may determine that the test sample is most likely to have breast cancer. In another example, if the cancer prediction is binary, with a 60% likelihood of being non-cancerous and a 40% likelihood of being cancerous, the analytics system determines that the test sample is most likely not to have cancer. In a further embodiment, the highest-likelihood cancer prediction may still be compared to a threshold (e.g., 40%, 50%, 60%, 70%) to call the test subject having its cancer type. If the highest-likelihood cancer prediction does not exceed that threshold, the analytics system may return an indeterminate result.

[0163] In a further embodiment, the analytics system connects a series of cancer classifiers. For example, the analytics system may input a test feature vector into a cancer classifier trained as a binary classifier in step 620 of process 600. The analytics system can receive the output of a cancer prediction. The cancer prediction may be binary, indicating whether the subject is likely to have cancer or not. In other implementations, the cancer prediction includes prediction values ​​describing the likelihood of cancer and the likelihood of not having cancer. For example, the cancer prediction may have an 85% cancer prediction value and a 15% non-cancer prediction value. The analytics system may determine that this subject is likely to have cancer. If the analytics system determines that a subject is likely to have cancer, it may input the test feature vector into a multiclass cancer classifier trained to distinguish between different cancer types, for example, as described in step 630 of process 600. The multiclass cancer classifier receives the test feature vector and returns a cancer prediction for one of several cancer types. For example, a multiclass cancer classifier provides a cancer prediction that designates the test subject as having the highest probability of having ovarian cancer. In another embodiment, the multiclass cancer classifier provides a prediction value for each of several cancer species. For example, the cancer prediction may include a 40% breast cancer prediction, a 15% colorectal cancer prediction, and a 45% liver cancer prediction. In another embodiment, the analytics system may further input test feature vectors into a cancer classifier trained to estimate tumor ratios, for example, as described in step 640 of process 600. For example, the estimated tumor ratios may be expressed as percentages, such as a 25% tumor ratio. When concatenated as a series in the multiclass cancer classifier, the estimated tumor ratios may specify different ratios for each cancer species.

[0164] According to a generalized embodiment of binary cancer classification, an analytics system can determine a cancer score for a test sample based on the test sample's sequencing data (e.g., methylation sequencing data, small variant sequencing data, other DNA sequencing data, RNA sequencing data, etc.). The analytics system can compare the test sample's cancer score against a binary threshold cutoff for predicting whether the test sample is likely to have cancer. The binary threshold cutoff can be tuned using TOO thresholding based on one or more TOO subtype classes. The analytics system can further generate feature vectors for the test sample used in a multi-class cancer classifier to determine a cancer prediction indicating one or more likely cancer types.

[0165] The classifier can be used to determine the disease status of a test subject, for example, a subject whose disease status is unknown. This method may involve obtaining a test genome data construct (e.g., test data at a single time point) in electronic format, containing values ​​for each genomic characteristic of multiple genomic characteristics of corresponding nucleic acid fragments in a biological specimen obtained from the test subject. This method may then involve applying the test genome data construct to the test classifier to determine the disease status of the test subject. The test subject may not have been previously diagnosed with the disease condition.

[0166] The classifier may be a temporary classifier that uses at least (i) a first test genome data construct generated from a first biological specimen collected from the test subject at a first time point, and (ii) a second test genome data construct generated from a second biological specimen collected from the test subject at a second time point.

[0167] A trained classifier can be used to determine the disease status of a test subject, for example, a subject whose disease status is unknown. In this case, the method may involve obtaining a test time-series dataset electronically for a test subject, where the test time-series dataset includes a corresponding test genotype data construct containing values ​​for multiple genotype characteristics of multiple corresponding nucleic acid fragments in corresponding biological samples obtained from the test subject at each of multiple time points, and an indicator value for the length of time between each pair of consecutive time points at each of multiple time points. The method may then involve applying the test genotype data construct to the test classifier to determine the disease condition status of the test subject. The test subject may not have been previously diagnosed with the disease condition.

[0168] V. Application In some embodiments, the methods, analytics systems, and / or classifiers of the present invention can be used to detect the presence of cancer, monitor the progression or recurrence of cancer, monitor treatment response or effectiveness, determine or monitor the presence of minimal residual disease (MRD), or any combination thereof. For example, as described herein, the classifier can be used to generate a probability score (e.g., 0 to 100) that describes the likelihood that a given test feature vector is from a subject with cancer. In some embodiments, whether or not a subject has cancer is determined by comparing the probability score to a threshold probability. In other embodiments, disease progression or treatment effectiveness (e.g., therapeutic effectiveness) is monitored by determining the likelihood or probability score at several different points in time (e.g., before or after treatment). In yet another embodiment, the likelihood or probability score can be used to make or influence clinical decisions (e.g., diagnosis of cancer, selection of treatment, determination of treatment effectiveness, etc.). For example, in one embodiment, if the probability score exceeds a threshold, a physician can prescribe appropriate treatment. In further embodiments, the method, analytics system, and / or classifier can be implemented to detect sources of contamination in sample processing and analysis workflows. Once contamination and / or sources of contamination are detected, the analytics system may help implement remediation measures to mitigate the contamination and its negative effects (e.g., skew results, bias the training of the classifier, etc.).

[0169] VA Panel Assignment Generation In one exemplary embodiment, the analytics system generates sequencing panel assignments with the aim of reducing sequencing costs without compromising the level of detection (LoD).

[0170] The analytics system 200 obtains sequencing data (e.g., test sequences) for a set of samples (e.g., in this case, samples that meet a set of criteria described herein). The first sequencing data may be a set of CCGA indicators, but it may also be a set of other genomic regions to be analyzed. These sequencing data are associated with several test sequences and linked to feature values ​​(e.g., GC content, error rate, sequence depth count, presence or absence of variants, mean allele frequency, total number of minor variants, and allele frequency of true variants).

[0171] The analytics system 200 selects the feature values ​​to analyze for each sample. For example, the feature values ​​may include the GC content, error rate, sequence depth count, presence or absence of variants, mean allele frequency, total number of minor variants, and allele frequency of true variants in the sequencing data. Other feature values ​​are also possible.

[0172] The analytics system 200 determines the rank of samples in a set of genome regions according to their feature values. For example, the sample with the highest feature value is ranked at the top, while the sample with the lowest feature value is ranked at the bottom.

[0173] After ranking the samples based on decreasing feature values, this method involves applying a greedy algorithm to add the next highest-ranked sample from the remaining ranked samples to the panel, such that the panel on which the samples are rearranged contains the lowest feature value.

[0174] In some embodiments, the analytics system 200 may access one or more additional feature sets and apply a machine learning model to the samples based on these additional feature sets. By doing so, the analytics system 200 can identify one or more additional feature subsets to consider when assigning the samples to a panel by applying a greedy algorithm.

[0175] The analytics system 200 generates optimized sequencing panel assignments by aggregating panels using seeding and swapping techniques. In some embodiments, the analytics system 200 processes samples sequentially to assign them to panels, determines the mean value of feature values ​​for each panel, swaps two samples between two different panels, and measures the deviation of mean feature values ​​for each of the two different panels after the swap. The sequencing panel generator includes repeating these steps a predetermined number of swaps, thereby generating panel assignments based on feature values ​​from the sequencing data. In such cases, the repeated steps are performed until the decrease in mean feature values ​​falls below a threshold.

[0176] Several filtering methods can be used to improve sequencing panel assignments. In the first example, the sequencing panel generator can only obtain feature values ​​for genomic regions that have variants in a threshold number of sequences in the sequencing data. In the second example, the sequencing panel generator can duplicate or remove duplicate genomic regions from the panel to improve detection capability. In the third example, the system administrator can remove genomic regions from the analysis. In the fourth example, the system administrator can remove samples from the sequencing panel. Finally, the sequencing panel generator can remove feature values ​​from the panel based on a feature value blacklist. The feature value blacklist may include patented feature values, feature values ​​known to cause false positives, or any other feature values ​​that may reduce the detection capability of the panel.

[0177] VAI (Vehicle-Assisted Inventory) Exemplary Panel Assignment Considerations As described herein, the analytics system 200 generates sequencing panel assignments with the aim of reducing sequencing costs without compromising the limit of detection (LoD).

[0178] As described above, the aim of the method described herein is to reduce sequencing costs without compromising the limit of detection (LoD). The type of specimen (e.g., target and control specimens used in the development of this method) is considered when determining the steps required to optimize the sequencing panel. Here, the target specimens selected for testing included one or more of the following: (i) undetectable by previous classifiers (i.e., detectability), (ii) available plasma tubes (i.e., available plasma), and (iii) if the tumor ratio estimate currently available is less than 1% (i.e., low TF estimate). As described above, the pool of possible specimens included specimens from the Circulating Cell Genome Atlas 1 study (CCGA1) and the Circulating Cell Genome Atlas 2 study (CCGA2).

[0179] (i) Applying the “detectability” filter to CCGA1 and CCGA2 samples revealed approximately 297 participants. For example, clinically evaluable participants for whom variant calling data was already available were analyzed. Specifically, participants with WGBS of tumor biopsies and WGS of cfDNA were selected from the Circulating Cell Genome Atlas 1 study (CCGA1) and the Circulating Cell Genome Atlas 2 study (CCGA2). For samples from CCGA1, 129 patients were revealed by a threshold Mscore (0.755) below the 98% specificity score cutoff. For samples from CCGA2, 196 participants were revealed by participants who were undetectable across all samples at the 99.4% specificity score cutoff (v0.5 tube 1 / tube 2, v2 lane 12 / 34). Of these 297 participants who met the detectability criteria, only 236 had sufficient plasma available for further analysis. Of these 236 participants who met the first two criteria (i.e., never having been detected before and having available plasma tubes), 210 participants had a tumor ratio estimate of less than 1% (see Figure 7). Seven participants had a TF estimate greater than 1%, and 19 participants had no TF estimate; these 26 participants were excluded. These 210 participants were used to test the methods described herein.

[0180] The benchmark sample is an artificial titration at a known tumor ratio (e.g., any of the tumor ratio values ​​described herein). In some embodiments, the estimated tumor ratio (TF) of the benchmark sample is at least 10%. In one or more embodiments, the benchmark sample can be used for benchmarking variant calling. In one embodiment, the benchmark sample can also be used for TF estimation.

[0181] The benchmark sample was selected from CCGA1 participants that included WBC WGS data. Of these 201 participants, only 5 had a tumor ratio estimate higher than 10% (see Figures 8 and 9). From these 5 samples, the 3 samples with the highest number of variant calls were selected as the benchmark sample (see Figure 10).

[0182] One of the main considerations when developing improved sequencing panel assignment methods is to find the lowest tumor ratio (e.g., detection limit) that can be detected using feature values ​​(e.g., SNPs / variants) for a given set of genomic regions, but not below it.

[0183] To determine the lowest tumor ratio (e.g., limit of detection) for each participant, simulations were performed including the expected number of all fragments and surrogate allele-containing fragments under a range of possible tumor ratios, as well as pure noise (i.e., 0% TF). The limit of detection (LoD) refers to the lowest TF that well separates the surrogate ratio distribution from the noise. For LoD, the primary parameters were the number of post-collapse fragments per target and the error rate of post-collapse fragments.

[0184] As shown in Figure 11, the error rates for different types of transformations reveal that when using variants as feature values, the type of transformation may need to be considered in variant selection. In addition, Figures 12-14 show that the amount of cfDNA should be additionally considered when thinking about LoD.

[0185] In one embodiment, variants are used as feature values. For example, sequencing data containing variant calls that pass a threshold were analyzed for each patient. Variant calls, e.g., up to N variant calls, were selected and prioritized based on the noise rate and allele ratio in the tumor ratio. For each tumor ratio (TF), simulations were performed with logspace(-7, -3, and 50): alt_rates = Tumor Fraction × allele_fractions + (1-TF) × noise_rates. Approximately 10,000 simulations were performed: (1) total post-collapse target coverage of the sample (same for all sites); (2) sample alt counts per variant according to alt_rates and total coverage; and (3) alt_frac = sum(alt counts) / sum(total counts). A noise cutoff was used to ensure good separation between the alt ratio and the noise distribution: noise_cutoff = quantile(noise_alt_fracs_0.99). The tumor ratio of LoD was identified as having TF with 95% alt_frac specimens > noise_cutoff.

[0186] Figures 16 and 17 show the results of LoD modeling for the sample (patient) and SNPs as feature values. The expected LoD was primarily driven by the number of variants. Here, approximately 300 to 400 variants were needed to obtain a LoD of less than 5E-5. Figures 18 and 19 show simulations when the variant subset (i.e., the subset of genomic regions related to those variants) was set to a maximum of 500 variants per participant. 129 out of 201 target participants had an expected LoD < 5E-5 (excluding samples with fewer than 20 variants). LoD < 5E-5 was primarily driven by the number of variants. The inflection point of TF LoD is around the LoD target. Analysis of the 129 target participants with an expected LoD < 5E-5 revealed specific cancer types (Figure 20) and stages (Figure 21).

[0187] The number of variants for each participant can be determined experimentally. In some embodiments, the determination of the number of variants for each participant includes (i) the confidence of variant calling, (ii) the site error rate, and (iii) the ease of sequencing / availability of cfDNA. In one embodiment, the number of variants for each participant is less than 500.

[0188] In one embodiment, (i) determining the confidence level of a variant call includes the log-likelihood of true call versus noise.

number

[0189] As shown in Figure 22, the change in PPV becomes slight after 16. In one embodiment, the PPV cutoff is 25 (see Figure 22).

[0190] In one embodiment, (ii) the determination of “segment error” (i.e., segment of the variant) depends at least in part on the type of transformation of the variant (e.g., SNP). Figure 23 shows the error rates for different transformation types. This shows that the error rates are higher for A>G, T>C;C>A, G>T; and C>T, G>A. The box in Figure 23 highlights the transformation types with the lowest error rates, including A>C, T>G;A>T, T>A; and C>G, G>C. These variants with the lowest error rates met the criteria for “segment error rate”.

[0191] In one embodiment, (iii) the determination of ease of sequencing / cfDNA availability depends at least in part on GC content. Those skilled in the art will understand that other factors contribute to ease of sequencing. Here, relative coverage of genomic regions is reproducible across all samples. Therefore, consistent (or inconsistent) coverage is used when determining ease of sequencing. Figure 24 shows sequencing depth (y-axis) compared to GC content for different combinations. As shown in Figure 25, the normalized depth for GC bins (binned according to percent (%) GC content) shows a series of normalized depths before and after 1 with GC content of approximately 0.4 to approximately 0.7 percent. This indicates that GC content can act as a surrogate for sequencing depth, and that GC content of 0.4 to 0.7 may be optimal to ensure sufficient sequencing depth. This suggests that GC content is a feature value that can be used in the methods described herein for improving sequencing panel assignments.

[0192] In determining the variant for each participant, the LoD estimation framework was used to take into account the error rate and relative depth. This analysis used 500 sites with the same GC content and error rate, estimated LoD, and rankings of combinations (GC, SNP type) based on the LoD estimate. Figure 26 shows that the prioritization of variant (i.e., feature value) selection is driven primarily by the SNP type (i.e., error rate), except when the GC content is high. This suggests that SNPs (e.g., excluding those with high error rates) and / or GC content can be used as feature values ​​in the method described herein.

[0193] In determining the number of feature values ​​and corresponding genomic regions to include in the analysis, a site-specific (i.e., site-specific) cost value was determined. In one embodiment, the site-specific (e.g., site-specific) cost value depends at least in part on the difference between sequencing depth and GC content (see Figures 24, 25, and 27A). Comparing depth vs. GC content (Figure 25A) with average bag size vs. GC content (Figure 27B), it is shown that lower coverage sites also have smaller bag sizes. This suggests that, for a given amount of sequencing, the relative distribution of each target varies depending on the GC content.

[0194] Further analysis of site-specific costs (i.e., costs per genomic region) is shown in Figures 28A and 28B. Here, the normalized depth of the GC bins (Figures 28A and 28B) assigns a "relative" cost measure to each genomic region. The raw depth assigned to each region is total_reads × (target_cost / su(target_costs)). This analysis suggests that a GC content of approximately 0.4–0.7 correlates with sequencing depth and, therefore, with the relative cost measured for each region. In summary, samples with a GC content of approximately 0.4–0.7 correspond to feature values ​​that can be used in the methods described herein.

[0195] In the initial variant selection criteria, which examined relative costs, emphasis was placed on selection based on a certain number of read pairs and total cost. In such cases, factors included total read pairs per variant (total_rp × (cost / assumed total cost); total coverage (cov) versus average bag size (and unique cov) (Figure 29); average bag size versus duplex % (Figure 30); duplex % versus noise rate (e.g., noise rate = duplex_perc × duplex_noise + (1-duplex_perc) × nonduplex_noise; Figure 31 (slide 13); and (1-noise rate) × unique_cov versus expected effort-free unique coverage (EFC). Then, variants were selected based on (EFC / raw RP) until the total cost was reached (Figures 32A and 32B).

[0196] The number of participants (e.g., samples) per panel can be determined experimentally. In some embodiments, determining the number of participants per panel involves sequencing costs and panel ordering costs. When a sequencing panel assignment includes a fixed number of participants, each participant is assigned to a panel, and the number of participants per panel is a trade-off between the cost of ordering more panels and the increased sequencing cost "wasted" on participants (i.e., samples) from whom the necessary information has already been obtained.

[0197] In one embodiment, the number of participants (i.e., samples) per panel minimizes the total cost as a function of the number of panels ordered (when the number of participants is constant, the number of panels is equal to the number of participants per panel). Therefore, minimizing the total cost means that

number

[0198] Thus, the cost of sequencing is

number

number

[0199] Early detection of VB cancer In some embodiments, the methods and / or classifiers of the present invention can be used to detect the presence or absence of cancer in subjects suspected of having cancer. For example, a classifier (e.g., as described in Section IV and illustrated in Section V above) can be used to determine a cancer prediction that describes the likelihood that a given test feature vector is from a subject with cancer.

[0200] In one embodiment, cancer prediction is the likelihood (e.g., scored on a scale of 0 to 100) of whether a test sample has cancer (i.e., a binary classification). Therefore, the analytics system can determine a threshold for determining whether a test subject has cancer. For example, a cancer prediction of 60 or higher may indicate that the subject has cancer. In other embodiments, cancer predictions of 65 or higher, 70 or higher, 75 or higher, 80 or higher, 85 or higher, 90 or higher, or 95 or higher may indicate that the subject has cancer. In other embodiments, cancer prediction may indicate the severity of the disease. For example, a cancer prediction of 80 may indicate a more severe or later-stage cancer compared to a cancer prediction below 80 (e.g., a probability score of 70). Similarly, an increase in cancer prediction over time (e.g., determined by classifying test feature vectors from multiple samples taken from the same subject at two or more time points) may indicate disease progression, or a decrease in cancer prediction over time may indicate successful treatment.

[0201] In another embodiment, the cancer prediction includes a number of prediction values, where each of the multiple cancer types under classification (i.e., multi-class classification) has a prediction value (e.g., scored between 0 and 100). The prediction values ​​may correspond to the likelihood that a given training sample (and, at inference, the training sample) has each of the cancer types. The analytics system may identify the cancer type with the highest prediction value and indicate that the test subject is likely to have that cancer type. In another embodiment, the analytics system further determines that the test subject is likely to have that cancer type by comparing the highest prediction value to a threshold (e.g., 50, 55, 60, 65, 70, 75, 80, 85, etc.). In another embodiment, the prediction values ​​may also indicate the severity of the disease. For example, a prediction value higher than 80 may indicate a more severe or later-stage cancer compared to a prediction value of 60. Similarly, an increase in the predicted value over time (determined, for example, by classifying test feature vectors from multiple samples taken from the same subject at two or more time points) may indicate disease progression, or a decrease in the predicted value over time may indicate treatment success.

[0202] According to aspects of the present invention, the method and system of the present invention can be trained to detect or classify multiple adaptive cancers. For example, the method, system and classifier of the present invention can be used to detect the presence of one or more, two or more, three or more, five or more, ten or more, fifteen or more, or twenty or more different types of cancer.

[0203] Examples of cancers that can be detected using the methods, systems, and classifiers of the present invention include carcinomas, lymphomas, blastomas, sarcomas, and leukemia or lymphoid malignancies. More detailed examples of such cancers include, but are not limited to, squamous cell carcinoma (e.g., squamous cell carcinoma), skin cancer, melanoma, lung cancer including small cell lung cancer, non-small cell lung cancer ("NSCLC"), lung adenocarcinoma and lung squamous cell carcinoma, peritoneal cancer, gastric cancer including gastrointestinal cancer or stomach cancer. Examples include cancers such as pancreatic cancer (e.g., pancreatic ductal adenocarcinoma), cervical cancer, ovarian cancer (e.g., high-grade serous ovarian cancer), liver cancer (e.g., hepatocellular carcinoma (HCC)), hepatoma, bladder cancer (e.g., urothelial bladder cancer), testicular (germ cell tumor) cancer, breast cancer (e.g., HER2-positive, HER2-negative, and triple-negative breast cancer), brain cancer (e.g., astrocytoma, glioma (e.g., glioblastoma)), colon cancer, rectal cancer, colorectal cancer, endometrial cancer or uterine cancer, salivary gland cancer, kidney cancer or renal cancer (e.g., renal cell carcinoma, nephroblastoma or Wilms' tumor), prostate cancer, vulvar cancer, thyroid cancer, anal cancer, penile cancer, head and neck cancer, esophageal cancer, and nasopharyngeal cancer (NPC). Further examples of cancer include, but are not limited to, retinoblastoma, theca cell tumor, masculinizing tumor, non-Hodgkin lymphoma (NHL), multiple myeloma and acute hematological malignancies, endometriosis, fibrosarcoma, choriocarcinoma, laryngeal cancer, Kaposi's sarcoma, Schwannoma, oligodendroglioma, neuroblastoma, rhabdomyosarcoma, osteosarcoma, leiomyosarcoma, and urinary tract cancer.

[0204] In some embodiments, cancer is one or more of the following: anorectal cancer, bladder cancer, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastric cancer, head and neck cancer, hepatobiliary tract cancer, leukemia, lung cancer, lymphoma, melanoma, multiple myeloma, ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer, thyroid cancer, uterine cancer, or any combination thereof.

[0205] In some embodiments, one or more cancers may be “high-signal” cancers (defined as cancers with a 5-year cancer-specific mortality rate greater than 50%), such as anorectal cancer, colorectal cancer, esophageal cancer, head and neck cancer, hepatobiliary tract cancer, lung cancer, ovarian cancer, and pancreatic cancer, as well as lymphoma and multiple myeloma. High-signal cancers tend to be more aggressive and typically exhibit above-average cell-free nucleic acid concentrations in test specimens obtained from patients.

[0206] Monitoring of VC cancer and treatment In some embodiments, cancer predictions can be determined at multiple different time points (e.g., before or after treatment) to monitor disease progression or treatment effectiveness (e.g., therapeutic efficacy). For example, the present invention includes a method involving obtaining a first sample (e.g., a first plasma cfDNA sample) from a cancer patient at a first time point, determining a first cancer prediction therefrom (as described herein), obtaining a second test sample (e.g., a second plasma cfDNA sample) from a cancer patient at a second time point, and determining a second cancer prediction therefrom (as described herein). This classification may further quantify tumor volume to determine changes over time.

[0207] In certain embodiments, the first time point is before cancer treatment (e.g., before surgical resection or therapeutic intervention), and the second time point is after cancer treatment (e.g., after surgical resection or therapeutic intervention), and the effectiveness of the treatment is monitored using a classifier. For example, if the second cancer prediction is lower than the first cancer prediction, the treatment is considered successful. However, if the second cancer prediction is higher than the first cancer prediction, the treatment is considered unsuccessful. In other embodiments, both the first and second time points are before cancer treatment (e.g., before surgical resection or therapeutic intervention). In yet another embodiment, both the first and second time points are after cancer treatment (e.g., after surgical resection or therapeutic intervention). In yet another embodiment, cfDNA samples are obtained from cancer patients at the first and second time points and analyzed to, for example, monitor cancer progression, determine whether the cancer is in remission (e.g., after treatment), monitor or detect residual disease or disease recurrence, or monitor the effectiveness of the treatment (e.g., therapeutically).

[0208] Those skilled in the art will readily recognize that test samples can be obtained from cancer patients over any desired set of time points and analyzed by the method of the present invention to monitor the patient's cancer status. In some embodiments, the first and second time points may be about 30 minutes, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or about 24 hours, about 1, 2, 3, 4, 5, 10, 15, 20, 25, or about 50 days, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months, or about 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7 The time intervals are separated by a magnitude ranging from approximately 15 minutes to a maximum of approximately 30 years, such as 0.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5, or approximately 30 years. In other embodiments, test specimens can be obtained from the patient at least once every 5 months, at least once every 6 months, at least once a year, at least once every 2 years, at least once every 3 years, at least once every 4 years, or at least once every 5 years.

[0209] VD treatment In another embodiment, cancer prediction can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, assessment of treatment effectiveness, etc.). For example, in one embodiment, if cancer prediction (e.g., for cancer or for a particular type of cancer) exceeds a threshold, a physician may prescribe appropriate treatment (e.g., surgical resection, radiotherapy, chemotherapy, and / or immunotherapy).

[0210] Using a classifier (as described herein), it is possible to determine a cancer prediction that a given sample feature vector is from a subject with cancer. In one embodiment, when the cancer prediction exceeds a threshold, an appropriate treatment (e.g., surgical resection or therapy) is prescribed. For example, in one embodiment, if the cancer prediction is 60 or higher, one or more appropriate treatments are prescribed. In another embodiment, if the cancer prediction is 65 or higher, 70 or higher, 75 or higher, 80 or higher, 85 or higher, 90 or higher, one or more appropriate treatments are prescribed. In yet another embodiment, the cancer prediction may indicate the severity of the disease. An appropriate treatment that matches the severity of the disease may then be prescribed.

[0211] In some embodiments, the treatment is one or more cancer therapy agents selected from the group consisting of chemotherapeutic agents, targeted cancer therapy agents, differentiation therapy agents, hormone therapy agents, and immunotherapy agents. For example, the treatment may be one or more chemotherapeutic agents selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeletal disruptors (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based drugs, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapy agents selected from the group consisting of signaling inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteosome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation therapy agents, including retinoids such as tretinoin, alitretinoin, and bexarotene. In some embodiments, the treatment is one or more hormonal therapy agents selected from the group consisting of anti-estrogens, aromatase inhibitors, progestins, estrogens, anti-androgens, and GnRH agonists or analogues. In one embodiment, the treatment is one or more immunotherapy agents selected from the group consisting of rituximab (RITUXAN) and alemtuzumab (CAMPATH), nonspecific immunotherapies and adjuvants such as BCG, interleukin-2 (IL-2), and interferon alfa, immunomodulatory agents such as thalidomide and lenalidomide (REVLIMID), and monoclonal antibody therapies. Selecting an appropriate cancer therapy agent based on characteristics such as tumor type, cancer stage, past exposure to cancer treatment or cancer therapy agents, and other characteristics of cancer is within the capabilities of a physician or oncologist skilled in the art.

[0212] Embodiment of the VE kit Furthermore, this specification discloses kits for carrying out the methods described above, including methods relating to cancer classifiers. The kit may include one or more collection containers for collecting specimens from individuals containing genetic material. Examples of specimens include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. Such a kit may include reagents for isolating nucleic acids from the specimens. The reagents may further include reagents for sequencing nucleic acids, including buffers and detection agents. In one or more embodiments, the kit may include one or more sequencing panels containing probes for targeting specific genomic regions, specific mutations, specific gene variants, or combinations thereof. For example, an analytics system may generate different treatment kits for different target sequencing panels as determined by a panel assignment model. In other embodiments, specimens collected by the kit are provided to a sequencing laboratory capable of sequencing the nucleic acids in the specimens using the sequencing panels.

[0213] The kit may further include instructions for using the reagents included in the kit. For example, the kit may include instructions for sample collection and extraction of nucleic acids from test samples. Exemplary instructions may include the order in which reagents should be added, the centrifugation rate to be used for isolating nucleic acids from test samples, how to amplify the nucleic acids, how to sequence the nucleic acids, or any combination thereof. The instructions may further clarify how the computing device should be operated as the analytics system 200 for the purpose of performing any step of the described method.

[0214] In addition to the components described above, the kit may include a computer-readable storage medium containing computer software for carrying out the various methods described throughout this disclosure. One possible form of these instructions is as printed information on a suitable medium or substrate, for example, on one or more sheets of paper on which the information is printed, on the packaging of the kit, in accompanying documentation. Another possible means is a computer-readable medium in which the instructions are stored in the form of computer code, for example, a diskette, CD, hard drive, or network data storage. Yet another possible means is a website address or QR code that can be used via the Internet to access the information at a removed site.

[0215] VI. Additional considerations A detailed description of the embodiments described above is provided with reference to the accompanying drawings, which illustrate specific embodiments of the present disclosure. Other embodiments having different structures and operations do not depart from the scope of the present disclosure. The terms “the present invention,” etc., are used in reference to specific specific examples of many alternative aspects or embodiments of the applicants’ invention shown herein, and neither their use nor their absence is intended to limit the scope of the applicants’ invention or the claims.

[0216] Embodiments of the present invention also relate to equipment for carrying out the operations described herein. This equipment may be specifically configured for a required purpose and / or may include a general-purpose computing device that is selectively enabled or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-temporary, tangible computer-readable storage medium or any type of medium suitable for storing electronic instructions, the medium may be coupled to a computer system bus. Furthermore, any computing system referenced herein may include a single processor or may be an architecture employing a multi-processor design for increased computing power.

[0217] Any step, operation, or process described herein as being performed by an analytics system may be performed or implemented by one or more hardware or software modules of the device, either alone or in combination with other computing devices. In one embodiment, the software module is implemented using a computer program product including a computer-readable medium containing computer program code that can be executed by a computer processor in order to perform any of the steps, operations, or processes described herein.

Claims

1. A method for performing target variant sequencing of a sample set, Obtaining initial sequencing data for each specimen that describes the presence or absence of each of several gene variants in the reference genome, wherein the initial sequencing data includes sequence reads of nucleic acid fragments present in the biological specimen obtained from the subject; For each sample, determine the number of gene variants present in the sequencing data of the sample; For each sample, determine one or more characteristics of each gene variant present in the sequencing data of the sample; The panel assignment of each sample to one of multiple target variant sequencing panels is determined by applying a panel assignment model, and the application of the panel assignment model is By processing each sample in the sample set sequentially, the corresponding panel assignment is determined to optimize the uniformity of panel size across all target variant sequencing panels and the uniformity of the gene variant characteristics across all samples assigned to each target variant sequencing panel. By performing a swap operation, the panel assignments of at least two specimens are swapped to further optimize the uniformity of panel size across all of the target variant sequencing panels and the uniformity of the gene variant characteristics across all specimens assigned to each target variant sequencing panel; and To generate each target sequencing panel that includes the specimens assigned to the target sequencing panel and that indicates the aggregated gene variant set across all the specimens assigned to the target sequencing panel. Including; and Perform the target variant sequencing on each target variant sequencing panel, including the assigned specimens, with respect to the self-free deoxyribonucleic acid (cfDNA) specimens subsequently collected from the aforementioned subject. A method that includes this.

2. The method according to claim 1 or any dependent claim, wherein the initial sequencing includes whole-genome sequencing or whole-exome sequencing of each specimen.

3. The method according to claim 1 or any dependent claim, wherein the gene variant includes single nucleotide variants, insertion variants, and deletion variants.

4. The method according to claim 1 or any dependent claim, wherein the gene variant includes a copy number variant.

5. The method according to claim 1 or any dependent claim, wherein one or more of the characteristics of each gene variant are selected from the group consisting of the guanine-cytosine content of the gene variant, the error rate of the targeting probe of the gene variant, the sequence depth count of the gene variant, the presence or absence of the gene variant, the average allele frequency of the gene variant, the total number of gene variants, and the allele frequency of the true gene variant.

6. Applying the aforementioned panel assignment model is Determining the panel assignment for each sample in the sample set according to one or more hard constraints. The method according to claim 1 or any dependent claim, further comprising:

7. The method according to claim 6, wherein one or more hard constraints are selected from the group consisting of the maximum number of samples per target sequencing panel, the maximum number of samples of one type per target sequencing panel, or the total number of gene variants per target sequencing panel.

8. The method according to claim 1 or any dependent claim, wherein the panel assignment model includes a function that scores the uniformity of panel sizes across all the target variant sequencing panels and the uniformity of the gene variant characteristics across all samples assigned to each target variant sequencing panel.

9. The corresponding panel assignment is determined by sequentially processing each sample in the aforementioned sample set. The corresponding panel assignment is determined by applying a greedy algorithm or a dynamic programming algorithm to the aforementioned function. The method according to claim 8, including the method described in claim 8.

10. Performing a swap operation Evaluating the change in score based on the function that evaluates the swap of the at least two samples; and Swapping the panel assignments of at least two samples based on the change in the score exceeding a threshold. The method according to claim 8 or any claim dependent thereon, including the method according to claim 8.

11. The method according to claim 1 or any dependent claim, wherein pairs of samples assigned to different target sequencing panels are determined by repeatedly performing the swap operation.

12. The method according to claim 1 or any dependent claim, further comprising identifying an optimal aggregated set of gene variants across all specimens based on the characteristics of the gene variants in each specimen assigned to the target sequencing panel for each target sequencing panel.

13. Identifying the optimal gene variant set for each target sequencing panel is possible for each target sequencing panel: To determine whether each sample has a total number of gene variants exceeding the threshold per sample; and In response to at least one sample having a total number of gene variants exceeding the threshold per sample, identify a subset of gene variants in the sample to be included in the target sequencing panel that optimizes the properties of the optimal aggregated gene variant set. The method according to claim 12 or any claim dependent thereon, including the method described therein.

14. The method according to claim 13, wherein identifying a subset of gene variants to be included in the target sequencing panel includes including gene variants present in other samples of the target sequencing panel.

15. The method according to claim 1 or any dependent claim, further comprising generating each target sequencing panel to identify a corresponding targeting probe that targets the aggregated gene variant set.

16. Obtaining target sequencing data for the sample set from the target variant sequencing panel relating to the cfDNA sample, wherein the target sequencing data includes sequence reads of cell-free nucleic acid fragments in the blood sample; For each cfDNA sample, Calling one or more variants present in the target sequencing data; Determining a feature vector based on one or more of the called variants; and The tumor ratio in the cfDNA sample is predicted by applying a cancer classifier to the aforementioned feature vector. The method according to claim 1 or any dependent claim, further comprising:

17. A non-temporary computer-readable storage medium that stores instructions causing the processor to perform the method according to claim 1 or any claim dependent thereon during execution by the processor.

18. Processor; and Non-temporary computer-readable storage medium according to claim 17 A system that includes this.

19. A target sequencing panel, A set of targeting probes that target aggregate gene variant sets across multiple samples assigned to the target sequencing panel, wherein the aggregate gene variant set and the multiple samples are Obtaining initial sequencing data for each sample in a sample set describing the presence or absence of each of several gene variants in a reference genome, wherein the initial sequencing data includes sequence reads of nucleic acid fragments in a biological sample obtained from one subject; For each sample, determine the number of gene variants present in the sequencing data of the sample; For each sample, determine one or more characteristics of each gene variant present in the sequencing data of the sample; By applying a panel assignment model, the panel assignment to one of multiple target variant sequencing panels is determined for each sample. Determined by, and applying the panel assignment model, By processing each sample in the sample set sequentially, the corresponding panel assignment is determined to optimize the uniformity of panel size across all target variant sequencing panels and the uniformity of the gene variant characteristics across all samples assigned to each target variant sequencing panel. By performing a swap operation, the panel assignments of at least two specimens are swapped to further optimize the uniformity of panel size across all target variant sequencing panels and the uniformity of the gene variant characteristics across all specimens assigned to each target variant sequencing panel. A target sequencing panel, including the target sequencing panel.

20. Multiple target sequencing panels, A set of targeting probes that target aggregate gene variant sets across multiple samples assigned to each target sequencing panel. Includes, The target sequencing panel, Obtaining initial sequencing data for each sample in a sample set describing the presence or absence of each of several gene variants in a reference genome, wherein the initial sequencing data includes sequence reads of nucleic acid fragments in a biological sample obtained from one subject; For each sample, determine the number of gene variants present in the sequencing data of the sample; For each sample, determine one or more characteristics of each gene variant present in the sequencing data of the sample; By applying a panel assignment model, the panel assignment to one of multiple target variant sequencing panels is determined for each sample. Determined by, and applying the panel assignment model, By processing each sample in the sample set sequentially, the corresponding panel assignment is determined to optimize the uniformity of panel size across all target variant sequencing panels and the uniformity of the gene variant characteristics across all samples assigned to each target variant sequencing panel. By performing a swap operation, the panel assignments of at least two specimens are swapped to further optimize the uniformity of panel size across all target variant sequencing panels and the uniformity of the gene variant characteristics across all specimens assigned to each target variant sequencing panel. A target sequencing panel, including the target sequencing panel.

21. A method for improving sequencing panel assignment for specimens from two or more individuals, To obtain sequencing data for each sample; Selecting feature values ​​from the aforementioned sequencing data; Applying a machine learning model that determines the sequencing panel assignment based on the feature values ​​from the sequencing data; and To generate an optimized sequencing panel assignment that includes samples from two or more individuals. A method that includes this.

22. The method according to claim 21, wherein the selection step further comprises determining a set of optimized feature values.

23. The method according to claim 22, wherein the machine learning model is selected from a classifier model, a pre-specified algorithm, and a regression model.

24. The method according to claim 23, wherein the machine learning model is a classifier model.

25. Applying the aforementioned classifier model Ranking the samples based on decreasing feature values; The greedy algorithm is applied to add the next highest ranked sample from the remaining ranked samples to the panel, such that the panel from which the samples are rearranged contains the lowest value among the feature values. The method according to any one of claims 21 to 24, including the method described in any one of claims 21 to 24.

26. Assigning the samples to a panel by processing them sequentially; Determining the average value of feature values ​​for each panel; Swapping two samples between two different panels; and After the swap, measure the deviation of the mean feature values ​​for each of the two different panels. The method according to claim 25, further comprising:

27. The method according to claim 26, further comprising repeating the steps described in claim 26 a predetermined number of swaps, thereby generating a panel assignment based on the feature values ​​from the sequencing data.

28. The method according to claim 27, wherein the repeated step is carried out until the decrease in the mean of the feature falls below a threshold.

29. The method according to any one of claims 21 to 28, wherein the panel comprises 16 specimens.

30. The method according to any one of claims 21 to 28, wherein the panel has 16 or fewer specimens.

31. The method according to any one of claims 21 to 30, wherein the panel includes benchmark samples.

32. The method according to claim 31, wherein the panel has one or fewer benchmark samples.

33. Applying the aforementioned classifier model Seeding the sequencing panel based on a decrease in the number of feature values; Swapping sequencing panel assignments for two participants seeded into two different panels; Measuring the reduction in the loss function after swapping; and Comparing the sequenced panel assignment set with the feature values. The method according to claim 32, including the method described in claim 32.

34. The method according to claim 33, further comprising repeating the process over a predetermined number of steps or until the decrease in the loss function falls below a threshold. [Request Item 35] [Number 1] The method according to claim 33 or 34, further comprising determining the sequencing panel assignment that satisfies the following conditions.

36. The method according to any one of claims 21 to 35, wherein the sequencing data is obtained from sequencing of cell-free nucleic acid molecules present in biological specimens obtained from multiple individuals.

37. The method according to any one of claims 21 to 36, wherein the feature value corresponds to a genomic region that includes one or more of the following: cancer-related genes, mutation hotspots, and viral regions.

38. The method according to any one of claims 21 to 37, wherein the sequencing data includes genomic regions associated with high-signal cancer or liquid cancer.

39. The method according to any one of claims 21 to 38, wherein the feature values ​​correspond to features corresponding to one or more of the following: GC content, error rate, sequence depth count, presence or absence of variants, mean allele frequency, total number of minor variants, and allele frequency of true variants.

40. The method according to any one of claims 21 to 39, wherein the feature value is the sum of a plurality of feature values.

41. The method according to any one of claims 21 to 40, wherein the feature value is a variant.

42. The method according to claim 41, wherein the variant includes one or more of the following: a single-nucleotide variant, an insertion, and a deletion.

43. A non-temporary computer-readable medium for storing one or more programs, wherein the one or more programs include instructions that cause an electronic device including a processor to perform the method described in any one of claims 21 to 42 when the device is executed by the device.

44. One or more processors; memory; and One or more programs, which are stored in the memory and configured to be executed by the one or more processors, include instructions for carrying out the method according to any one of claims 21 to 42. Electronic devices including