Ultra-deep sequencing for dependency identification

EP4802091A1Pending Publication Date: 2026-09-09LESHCHINER IGNATY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024809467
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-10-31
Publication Date
2026-09-09

AI Technical Summary

Technical Problem

Current technologies for identifying genetic dependencies in cancer and somatic tissues are limited by their inability to perform ultra-deep sequencing accurately and efficiently, especially in vivo, leading to high error rates and impracticality for clinical use.

Method used

The DepSeq technology employs a double-stranded polynucleotide construct with strands irreversibly attached at one region and de-hybridized at another, allowing for ultra-deep sequencing with high fidelity. This technology can be applied in vivo, enabling the direct profiling of millions of cells from patient tumor samples.

Benefits of technology

DepSeq achieves low error rates (at least 10A-6) and high sequencing depth, allowing for the accurate identification of tissue-specific cancer vulnerabilities and therapeutic windows, thereby overcoming the limitations of existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024054023_08052025_PF_FP_ABST
    Figure US2024054023_08052025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to compositions, methods, and kits for using ultra-deep sequencing to identify in vivo, ex vivo, and in vitro dependencies. Also provided are software and electronic devices for executing such methods. Such techniques can be applied to living tumors, disease tissues, blood, cell free nucleic acids, body fluids, animal samples, normal human or non-human tissues to profile mutational, epi-mutational, structural and other genomic landscapes and identify therapeutic targets toward disease treatment.
Need to check novelty before this filing date? Find Prior Art

Description

ULTRA-DEEP SEQUENCING FOR DEPENDENCY IDENTIFICATIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 546,688, filed October 31, 2023, which is incorporated herein by reference in its entirety.FIELD

[0002] The present disclosure relates to the field of sequencing and cancer, somatic tissue and disease research. Specifically, it pertains to compositions and methods for ultra-deep sequencing to identify in vivo, ex vivo and in vitro genetic dependencies.BACKGROUND

[0003] Identifying the vulnerabilities of cancer and other somatic tissues is important for developing new therapeutics, discovering new therapeutic targets, and understanding tumor and disease biology and treatment strategies. Current technologies rely on sample-prep techniques such as next-generation sequencing (NGS), PCR, and CRISPR-based approaches to process cell or tissue samples and to nominate dependencies and targets of interest.Current dependency profiling technologies also are limited to in vitro applications and cannot directly profile the genes or genetic regions that tumors or somatic tissue depend on, known as the “Achilles’ heels” of individual patient tumors, disease or somatic tissue in vivo. These methods are also costly and time-consuming, making them impractical for widespread clinical use. Measured dependencies on ex vivo tissue or models are often distinct from the dependencies of the tissue in vivo and could not be currently reliably measured. Current sequencing techniques are quite limited in their ability to accurately sequence to ultra-high depth, (often having error rates of, e.g., around 0.01% to 1%) and identify individual events developed in millions of growing cells in the tissue.

[0004] Accordingly, there exists a need for improved compositions, systems and methods for more accurate and efficient deep sequencing to allow better identification of cancer dependencies.BRIEF SUMMARY

[0005] The present disclosure addresses the limitations of existing technologies by providing compositions, methods, systems and kits for ultra-deep sequencing with high fidelity, capable of profiling millions of cells from patient tumor samples. This technology, referred to hereinas DepSeq (Dependency Sequencing), can be in vitro but also directly applied in vivo, including for example, to living tumors, disease tissues, or normal tissues, enabling the faster, more accurate identification of tissue-specific cancer vulnerabilities and therapeutic windows. DepSeq techniques, as disclosed herein, can profile the mutational landscape of millions of cells simultaneously to measure fitness effects and negative selection of specific genes.

[0006] In one aspect, the disclosure provides a construct, such as a polynucleotide construct for sequencing. The construct includes a double-stranded polynucleotide having first and second strands, each strand having first and second ends and a sequence of bases extending therebetween, and the sequence of bases in the first strand are at least partially hybridized to the sequence of bases in the second strand. A respective portion of the first strand is connected to a portion of the second strand by a non-base pairing mechanism, e.g., at least one nucleotide of the first strand is connected through a non-base pairing mechanism to at least one nucleotide of the second strand, so as to attach the two strands in a first region. This attachment can be done irreversibly, such as by covalent bonding of the two strands via a linker or other structure, as disclosed herein. It can also be done reversibly by an affinitybased linkage, a biotin or desthiobiotin linkage, an interconnected probe, a UV-, light- or temperature-cleavable linkage, or a chemically or enzymatically cleavable linkage. Other regions of the two strands are separated from each other so as to prevent hybridization of such other regions of the strands during sequencing. This separation can be done irreversibly, for example by a modification of one or more portions of the strands so as to irreversibly weaken or prevent hybridization in that region. Other mechanisms can be used to separate the strands, such as by cleavable nucleotide protection, NaOH, formamide, a base, solvent or temperature change, pH change, or a helicase. Because the two strands are separated in regions other than the site of the irreversible attachment, the separated regions can more freely be accessed by sequencing tools for more efficient and accurate sequencing. Some or even all of the non-connected portions may be irreversibly separated from each other (no longer hybridized and prevented from rehybridizing), to allow space throughout the strands for the tools used in PCR, nanopore sequencing and other tools to better access the polynucleotide regions of interest, so as to provide faster, more accurate PCR, nanopore sequencing or other sequencing techniques. In this respect, the first and second strands are configured to be attached at one or more portions, but also irreversibly de-hybridized from each other at other portions. The separation of the respective strand portions can be done, for example, by application of chemical agents, such as NaOH, formamide, or a base; ornucleotide substitution conversion such as conversion of cytosine into uracil through deamination, or an enzyme such as ADAR, TdAR, or a helicase; and other techniques as disclosed herein.

[0007] In some embodiments of the construct, the double-stranded polynucleotide is DNA, RNA a DNA-RNA hybrid, phosphorothioate nucleic acid, locked nucleic acid (LNA), morpholino nucleic acids (MNAs)or a peptide nucleic acid (PNA).

[0008] In some embodiments of the construct, the double-stranded polynucleotide is end- repaired, end-blunted, nick-repaired, ddNTP-incorporated, dA-tailed, and / or middle- base or end-tail modified with functionalized bases for crosslinking, click-chemistry, or di-sulfide.

[0009] In some embodiments of the construct, the at least one nucleotide of the first strand is irreversibly connected through a non-base pairing mechanism to at least one nucleotide of the second strand.

[0010] In some embodiments of the construct, the irreversible connection is a polylactic acid (PLA), a polyamide, a polyimide, an oligoethylene glycol or another polymer or oligomer linker, a hairpin loop, or a modified or natural nucleotide crosslinked with epoxy, cisplatin, carboplatin or derivatives, click-chemistry or disulfide bond.

[0011] In some embodiments of the construct, the at least one nucleotide of the first strand is or reversibly connected through a non-base pairing mechanism to at least one nucleotide of the second strand.

[0012] In some embodiments of the construct, the reversible connection is an affinity-based linkage, a biotin or desthiobiotin linkage, an interconnected probe, a UV-, light- or temperature-cleavable linkage, or a chemically or enzymatically cleavable linkage.

[0013] In some embodiments of the construct, the de-hybridization is irreversible.

[0014] In some embodiments of the construct, the irreversible de-hybridization is by means of nucleotide substitution conversion, an adenosine deaminase acting on RNA (ADAR) or DNA (TdAR family) enzymes, inosine, a chemical modification, conversion of cytosine into uracil by cytosine deaminase, incorporation of a non-natural or natural nucleotide or a nucleotide-analogue that facilitates de-hybridization.

[0015] In some embodiments of the construct, the de-hybridization is reversible.

[0016] In some embodiments of the construct, the reversible de-hybridization is by means of cleavable nucleotide protection, NaOH, formamide, a base, solvent or temperature change,pH change, or a helicase, or any combination thereof.

[0017] In some embodiments of the construct, the non-base pairing mechanism connects one end of the first strand with the opposite distant end of the second strand.

[0018] In some embodiments of the construct, the non-base pairing mechanism is a linker, a sugar unit connection, a phosphate connection, an affinity, or physical connection or any combination thereof.

[0019] In some embodiments of the construct, the linker is a polylactic acid (PLA), polyamide, polyimide, a hairpin loop oligoethylene glycol or another polymer or oligomer linker, or a modified or natural nucleotide that can be crosslinked with epoxy, cisplatin, carboplatin or derivatives, click-chemistry or disulfide bond.

[0020] In some embodiments of the construct, the linker is configured to bind to a solid support or a bead.

[0021] In some embodiments of the construct, the linker is configured to attach to desthiobiotin or derivatives for selection with a streptavidin bead.

[0022] In some embodiments of the construct, the construct is selected using a streptavidin bead.

[0023] In some embodiments of the construct, the process includes one or more of nucleotide substitution conversion, an adenosine deaminase acting on RNA (ADAR family) or DNA (TdAR family) enzymes, inosine, a chemical modification, or any combination thereof.

[0024] In some embodiments of the construct, the process includes one or more of conversion of cytosine into uracil by cytosine deaminase or other chemical or biochemical process, NaOH, formamide, a base, and a helicase.

[0025] In some embodiments of the construct, the process comprises a strand copying step incorporating a non-natural or natural nucleotide or a nucleotide-analogue that facilitates dehybridization.

[0026] In some embodiments of the construct, at least one end of the double-stranded polynucleotide is attached with a pair of adapters.

[0027] In some embodiments of the construct, the pair of adapters are UMI adapters or UDI adapters, or any combination thereof.

[0028] In some embodiments of the construct, the pair of adapters are further attached withPCR primers.

[0029] In some embodiments of the construct, the double-stranded or one strand of the polynucleotide or the linker is cut.

[0030] In some embodiments of the construct, the cut double-stranded polynucleotide is linearized.

[0031] In some embodiments of the construct, the double-stranded polynucleotide is circularized.

[0032] In some embodiments of the construct, the construct is configured for sequencing with a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher.

[0033] In some embodiments of the construct, the construct is configured for sequencing at an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

[0034] In some embodiments of the construct, the first and second strands are irreversibly dehybridized from each other in the second region.

[0035] In another aspect, and as alluded above, methods for sequencing polynucleotides are also contemplated. The sequencing methods provide low error rates. The disclosed methods include steps of providing a polynucleotide construct, such as any of the embodiments referenced above or elsewhere in this disclosure. The first and second strands of the polynucleotide are attached irreversibly at one or more portions (e.g., in a first region), and irreversibly de-hybridized from each other at other portions (e.g. in a second region).Irreversibly de-hybridizing the first and second strands can separate them from each other in the second region, so as to improve access of the strands in the region to sequencing tools. Sequencing is performed, such as on the second region, and thereby provides a more accurate sequence (with a lower error rate), for example a more accurate sequence readout of the desired region of the polynucleotide.

[0036] In some applications, the methods include the steps of providing a construct for sequencing, according to any of the preceding embodiments; and performing sequencing on the construct. The construct is provided in advance (or is formed at the site of usage) with an irreversible connection in a first region, comprising a non-base pairing mechanism thatconnects at least one nucleotide of the first strand and at least one nucleotide of the second strand in the first region. The construct is also provided so that a second region is prevented from hybridizing. That step of preventing hybridization can be done as disclosed herein. Optionally, the strands can also be formed with a second irreversible attachment (e.g., by a second linker, on the opposite ends of the strands or elsewhere in the second region), which is then cut prior to sequencing. This optional step can be used to preserve the structure of the construct during storage and shipping, and can be released (cut) at the site of usage.

[0037] In some embodiments of the method, the de-hybridizing is irreversible.

[0038] In some embodiments of the method, the irreversible de-hybridizing is by means of a polylactic acid (PLA), a polyamide, a polyimide, an oligoethylene glycol or another polymer or oligomer linker, a hairpin loop, or a modified or natural nucleotide crosslinked with epoxy, cisplatin, carboplatin or derivatives, click-chemistry or disulfide bond.

[0039] In some embodiments of the method, the de-hybridizing is reversible.

[0040] In some embodiments of the method, the reversible de-hybridizing is by means of an affinity-based linkage, a biotin or desthiobiotin linkage, an interconnected probe, a UV-, light- or temperature-cleavable linkage, or a chemically or enzymatically cleavable linkage.

[0041] In some embodiments of the method, the method results in a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher.

[0042] In some embodiments of the method, the method results in an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

[0043] In some embodiments of the method, the sequencing comprises using duplex sequencing.

[0044] In some embodiments of the method, the duplex sequencing comprises using duplex adapter-based sequencing, circular strand displacement polymerization reaction (CSDPR), rolling circle amplification (RCA), concatenating original duplex for error correction (CODEC), Ultima genomics duplex, or any combination thereof.

[0045] In some embodiments of the method, the sequencing comprises using single-cell sequencing.

[0046] In some embodiments of the method, the sequencing comprises using a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, PacificBiosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

[0047] In some embodiments of the method, the method further comprises obtaining sequencing data from the sequencing and identifying cancer or tissue dependencies, cancer or tissue expansion driving mutations, and / or cancer or disease resistance mutations based on the sequencing data.

[0048] In some embodiments of the method, the strands in the second region are subjected to substituted conversion of one or more nucleotides in the second region, application of adenosine deaminase acting on RNA (ADAR) and DNA (TdAR) enzyme families, inosine, a chemical modification, or any combination thereof.

[0049] In some embodiments of the method, the polynucleotide is DNA, and the dehybridization process comprises nucleotide substitution conversion, an adenosine deaminase acting on RNA (ADAR) or DNA TdAR enzyme, inosine, a chemical modification, or application of bisulfite conversion.

[0050] In some embodiments of the method, the construct is sequenced using duplex sequencing.

[0051] In some embodiments of the method, the duplex sequencing comprises using duplex adapter based sequencing, circular strand displacement polymerization reaction (CSDPR), rolling circle amplification (RCA), concatenating original duplex for error correction (CODEC), Ultima genomics duplex, or any combination thereof.

[0052] In some embodiments of the method, the construct is sequenced using single-cell sequencing.

[0053] In some embodiments of the method, the construct is subject to target capture, PCR enrichment, WES, CRISPR based enrichment, or any combination thereof.

[0054] In some embodiments of the method, the construct is sequenced on a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Pacific Biosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

[0055] In some embodiments of the method, the construct is provided with an irreversible connection in the second region, comprising a non-base pairing mechanism that connects at least one nucleotide of the second strand and at least one nucleotide of the first strand in thesecond region.

[0056] In some embodiments of the method, the method further comprises a step of applying a transposase enzyme with oligonucleotide integration or without, a restrictase, a glycosylase, or other chemical or biochemical process to cut the second irreversible connection in the second region prior to sequencing.

[0057] In some embodiments of the method, the method further comprises a step of applying a transposase enzyme with oligonucleotide integration or without, a restrictase, a glycosylase or other chemical or biochemical process to cut the linearized or circularly amplified molecule in one, two dissenting, or multiple distinct places within or outside of the original polynucleotide sequence.

[0058] In some embodiments of the method, the construct originated from a tumor or a somatic cell.

[0059] In some embodiments of the method, the method comprises identifying, from the sequencing performed on the construct, one or more genes or genomic regions that are dependencies of tumors or somatic tissue or mutations or epi-mutations depleted by the tumor.

[0060] Further provided is a method for profiling dependencies, the method comprising: providing sequence data of a sample; and identifying genetic dependencies based on the sequence data, wherein the sequence data comprise a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher, wherein the sequence data results in an error rate of at least 10A-5 or lower, at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

[0061] In some embodiments of the method, the sample is originated from a tumor or a somatic cell.

[0062] In some embodiments of the method, the method further comprises identifying, from the sequence data, one or more genes or genomic regions that are or have been dependencies of tumors or somatic tissue or mutations or epi-mutations depleted by the tumor.

[0063] In some embodiments of the method, the method further comprises identifying genomic regions of negative selection from mutations due to fitness constrains of growing cancer or non-cancer cells.

[0064] In some embodiments of the method, the method further comprises identifying genomic regions of positive selection from cancer or tissue drivers, growth drivers, and / or resistant events.

[0065] In some embodiments of the method, the sequence data are from duplex sequencing.

[0066] In some embodiments of the method, the duplex sequencing comprises using duplex adapter based sequencing, circular strand displacement polymerization reaction (CSDPR), rolling circle amplification (RCA), concatenating original duplex for error correction (CODEC), Ultima genomics duplex, or any combination thereof

[0067] In yet another aspect, the disclosure provides a kit for sequencing with low error rates, the kit comprising: a construct of any one of preceding embodiment; reagents for preparing a sequencing library; adaptors and / or linkers for duplex sequencing; and instructions for performing a sequencing method.

[0068] In some embodiments of the kit, the reagents comprise end-repair reagents, dA-tailing reagents, nick-repair reagents, ddNTPs, adapter ligation reagents, linker attachment reagents, PCR primers, genomic target-enrichment reagents, or any combination thereof.

[0069] In some embodiments of the kit, the sequencing method comprises duplex sequencing.

[0070] In some embodiments of the kit, the adaptors and / or linkers are for duplex sequencing.

[0071] In some embodiments of the kit, the sequencing method is performed on a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Element Biosciences, Pacific Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

[0072] In some embodiments of the kit, the kit comprises a de-hybridization agent comprising one or more of an agent configured to facilitate substitution conversion of a nucleotide on the first or second strand, an adenosine deaminase configured to act on RNA (ADAR) or DNA (TdAR family enzymes), inosine, a chemical modifier, or any combination thereof.

[0073] In some embodiments of the kit, the de-hybridization agent comprises or more of a compound or enzyme configured to convert a cytosine on the first or second strand into uracil, NaOH, formamide, a base, and a helicase.

[0074] In some embodiments of the kit, the de-hybridization agent comprises a non-natural or natural nucleotide or a nucleotide-analogue that facilitates de-hybridization for strand copying.

[0075] In yet another aspect, the disclosure provides a method for analyzing sequencing data with low error rates, the method comprising: (a) receiving, via an input device, a set of sequencing data associated with a cancer sample; (b) automatically inputting, via a processor, the set of sequencing data to an algorithm, wherein the algorithm is configured to output a set of mutations, genes, or genomic regions corresponding to the cancer dependencies of the cancer sample; and (c) outputting the set of mutations via an output device.

[0076] In some embodiments of the method, the set of sequencing data is obtained from a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Pacific Biosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

[0077] In some embodiments of the method, the set of sequencing data is obtained from a sequencing protocol that results in a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher.

[0078] In some embodiments of the method, the method results in an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

[0079] In some embodiments of the method, the algorithm comprises a negative binomial, Poisson model or log normal Poisson model for one or more of background selection, mutation rate identification, and methylation status quantification.

[0080] In some embodiments of the method, the algorithm utilizes local clustering, sub- clonality, copy-number information, structural variation, transcription-coupled repair, replication time information, mutation spectra and processes, or any combination thereof.

[0081] In some embodiments of the method, the algorithm is implemented using high- performance computing or a cloud system.

[0082] In some embodiments of the method, the method further comprises annotating the set of mutations with amino acid changes, mutational signature profiles, genomic elements positions and effects, or any combination thereof.

[0083] In some embodiments of the method, the method further comprises identifyinggenomic regions under positive, negative or neutral selection based on the annotation of the set of mutations.

[0084] In some embodiments of the method, the method further comprises calculating dN / dS rates, mutation effect sizes and statistical significance, copy-number variations (CNVs), or any combination thereof.

[0085] In some embodiments of the method, the method further comprises at least one of the following: determining single cell mutations; calculating statistical depletion or enrichment of a specific mutation type; identifying genomic regions of negative selection from mutation due to fitness constrains of the growing cancer or non-cancer cells; identifying genomic regions of positive selection from cancer drivers, growth drivers, and / or resistant events; summarizing information in a report; identifying cancer vulnerabilities, disease vulnerabilities and druggable targets in the genome; and profiling of selection before or after treatment from in vivo tissue, ex vivo tissue, cell lines, blood, nucleic acids, circulating nucleic acids, circulating cells, patient derived material, patient derived models, animal models, or organoids.

[0086] In some embodiments of the method, the method further comprises at least one of the following: estimating the copy -number profile of cancer based on selected reads; identifying highly mutable regions with a high-likelihood of selection signals; merging signals of methylation, fragmentation, copy-number, and mutation for accurate estimation of selection signals; selecting regions with clinical significance and / or mutations that are oncogenic or resistant; profiling using panels based on gene, gene region, non-coding region, and / or whole exome / genome screens; estimating cancer cell fraction and frequency of identified mutations to discover selection signals and changes thereof over time; and estimating selection signals for normal tissues, jointly or independently for mixtures of cells and tissue types.

[0087] Further provided is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to perform the method of any one of the preceding embodiments.

[0088] In some embodiments of the non-transitory computer-readable storage medium, the set of sequencing data is obtained from a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Pacific Biosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

[0089] In some embodiments of the non-transitory computer-readable storage medium, the set of sequencing data is obtained from a sequencing protocol that results in a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher.

[0090] In some embodiments of the non-transitory computer-readable storage medium, the method results in an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A- 8 or lower, or at least 10A-9 or lower.

[0091] In some embodiments of the non-transitory computer-readable storage medium, the algorithm comprises a negative binomial, Poisson model or log normal Poisson model for one or more of background selection, mutation rate identification, and methylation status quantification.

[0092] In some embodiments of the non-transitory computer-readable storage medium, the algorithm utilizes local clustering, sub-clonality, copy-number information, structural variation, transcription-coupled repair, replication time information, mutation spectra and processes, or any combination thereof.

[0093] In some embodiments of the non-transitory computer-readable storage medium, the algorithm is implemented using high-performance computing or a cloud system.

[0094] In some embodiments of the non-transitory computer-readable storage medium, the method further comprises annotating the set of mutations with amino acid changes, mutational signature profiles, genomic elements positions and effects, or any combination thereof.

[0095] In some embodiments of the non-transitory computer-readable storage medium, the method further comprises identifying genomic regions under positive negative or neutral selection based on the annotation of the set of mutations.

[0096] In some embodiments of the non-transitory computer-readable storage medium, the method further comprises calculating dN / dS rates, mutation effect sizes and statistical significance, copy-number variations (CNVs), or any combination thereof.

[0097] In some embodiments of the non-transitory computer-readable storage medium, the method further comprises at least one of the following: determining single cell mutations; calculating statistical depletion or enrichment of a specific mutation type; identifying genomic regions of negative selection from mutation due to fitness constrains of the growingcancer or non-cancer cells; identifying genomic regions of positive selection from cancer drivers, growth drivers, and / or resistant events; summarizing information in a report; identifying cancer vulnerabilities, disease vulnerabilities and druggable targets in the genome; and profiling of selection before or after treatment from in vivo tissue, ex vivo tissue, cell lines, blood, nucleic acids, circulating nucleic acids, circulating cells, patient derived material, patient derived models, animal models, or organoids.

[0098] In some embodiments of the non-transitory computer-readable storage medium, the method further comprises at least one of the following: estimating the copy-number profile of cancer based on selected reads; identifying highly mutable regions with a high likelihood of selection signals; merging signals of methylation, fragmentation, copy -number, and mutation for accurate estimation of selection signals; selecting regions with clinical significance and / or mutations that are oncogenic or resistant; profiling using panels based on gene, gene region, non-coding region, and / or whole exome / genome screens; estimating cancer cell fraction and frequency of identified mutations to discover selection signals and changes thereof over time; and estimating selection signals for normal tissues, jointly or independently for mixtures of cells and tissue types.

[0099] Further provided is an electronic device, comprising: a display; one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any one of the preceding embodiments.

[0100] In some embodiments of the electronic device, the set of sequencing data is obtained from a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Pacific Biosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

[0101] In some embodiments of the electronic device, the set of sequencing data is obtained from a sequencing protocol that results in a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000- fold and higher, or 1,000,000-fold and higher.

[0102] In some embodiments of the electronic device, the method results in an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

[0103] In some embodiments of the electronic device, the algorithm comprises a negative binomial, Poisson model or log normal Poisson model for one or more of background selection, mutation rate identification, and methylation status quantification.

[0104] In some embodiments of the electronic device, the algorithm utilizes local clustering, sub-clonality, copy-number information, structural variation, transcription-coupled repair, replication time information, mutation spectra and processes, or any combination thereof.

[0105] In some embodiments of the electronic device, the algorithm is implemented using high-performance computing or a cloud system.

[0106] In some embodiments of the electronic device, the method further comprises annotating the set of mutations with amino acid changes, mutational signature profiles, genomic elements positions and effects, or any combination thereof.

[0107] In some embodiments of the electronic device, the method further comprises identifying genomic regions under positive negative or neutral selection based on the annotation of the set of mutations.

[0108] In some embodiments of the electronic device, the method further comprises calculating dN / dS rates, mutation effect sizes and statistical significance, copy-number variations (CNVs), or any combination thereof.

[0109] In some embodiments of the electronic device, the method further comprises at least one of the following: determining single cell mutations; calculating statistical depletion or enrichment of a specific mutation type; identifying genomic regions of negative selection from mutation due to fitness constrains of the growing cancer or non-cancer cells; identifying genomic regions of positive selection from cancer drivers, growth drivers, and / or resistant events; summarizing information in a report; identifying cancer vulnerabilities, disease vulnerabilities and druggable targets in the genome; and profiling of selection before or after treatment from in vivo tissue, ex vivo tissue, cell lines, blood, nucleic acids, circulating nucleic acids, circulating cells, patient derived material, patient derived models, animal models, or organoids.

[0110] In some embodiments of the electronic device, the method further comprises at least one of the following: estimating the copy-number profile of cancer based on selected reads; identifying highly mutable regions with a high likelihood of selection signals; merging signals of methylation, fragmentation, copy-number, and mutation for accurate estimation of selection signals; selecting regions with clinical significance and / or mutations that areoncogenic or resistant; profiling using panels based on gene, gene region, non-coding region, and / or whole exome / genome screens; estimating cancer cell fraction and frequency of identified mutations to discover selection signals and changes thereof over time; estimating selection signals for normal tissues, jointly or independently for mixtures of cells and tissue types.BRIEF DESCRIPTION OF THE DRAWINGS

[0111] FIG. 1 shows an exemplary DepSeq process schema. DNA (or RNA) from 103- 108or more cells are collected from a sample, then prepared into a DepSeq library. The synthesis of the construct consists of steps (1-5), where 1) two hybridized polynucleotide strands are in solution, 2) the polynucleotide is pre-processed and / or repaired; then 3) a linker molecules are adapted to the hybridized polynucleotide in any of possible configurations 4) The molecule is treated with an agent to weaken hybridization 5) The molecule is optionally cut to linearize the polynucleotide (and integrate optional sequence at site of cut) . Steps 4 and 5 may be interchanged. Synthesis of the construct is followed by ultra deep sequencing (1,000- l,000,000x or deeper) and consensus of the linked reads. The highly accurate sequencing result allows single cell, or single molecule variant detection. The variants are used for direct nonsynonymous / synonymous (dN / dS) score measurement from a single sample, identifying candidate positively and negatively selected genes without additional experimentation.

[0112] FIG. 2 shows exemplary embodiments of the linker configuration and structure. DNA strands can be linked with a variety of chemical approaches, including: a) di-sulfide linkage; b) epoxy based linkage; c) DNA damaging agent such as cis-platinum; d) polylactic acid linkage; e) polynucleotide linkage; f) polynucleotide linkage with embedded UMI and UDI barcodes; g) the attachment of any linker with streptavidin to a bead that enables affinity selection; h) strong binding probes that are pre-linked in any way which allow synthesis of new, connected polynucleotides; i) a linker that connects the polynucleotide in two places, which then requires treatment of the nucleotide; the connected molecule can then be cut; j) same as i), but the linker is identical on both sides, and rolling circle amplification is used to create a linear product; k) same as j), but the linker is alternating. A cut site can be integrated into the linker or into the polynucleotide. Cutting at the site creates paired copies of the original polynucleotide.

[0113] FIG. 3 shows exemplary click chemistry linkers. Linkers can be constructed with click chemistry to connect the polynucleotide strands at any position, in one or more places.a) N3 may be connected to the polynucleotide at the positions required for linking, followed by any molecule connected to two C-C triple bonds, b) The complementary scheme where instead the C-C triple bond is connected to the polynucleotide and the connecting molecule contains two N3 groups, c) Alternative, copper-free click chemistry producing compatible results.

[0114] FIG. 4 shows an exemplary process of DepSeq for fast profiling. Treatment of the molecule is performed during sequencing with third-generation sequencing after construct assembly.

[0115] FIG. 5 shows an exemplary process for increased yield in DepSeq library. Affinity tag attachment at time of construct assembly allows for purification and high yield results.

[0116] FIG. 6 shows an exemplary treatment process to weaken hybridization, in which the construct undergoes either: a) physical process to separate the strands, such as NaOH or other base, formamide, and / or a helicase etc.; or b) chemical modification of the nucleotides to prevent their binding, such as protection of functional groups, oxidation of functional groups, bisulfite treatment, acetylation, glycosylation, etc., through nucleotide conversion by an enzymatic reaction (e.g., deamination of cytosine into uracil) adenosine deaminase acting on RNA (e.g. ADAR) and DNA (e.g., TdAR, ABE8, ABE7.10, SPACE) enzyme, inosine, a chemical treatment, or any combination thereof.

[0117] FIG. 7 shows an exemplary process of performing DepSeq on patient sample. The DepSeq protocol to identify dependencies of a sample through ultra deep sequencing.

[0118] FIG. 8. shows an exemplary process of DepSeq workflow to identify differential dependencies (treatment, location, exposure, immune pressure, tissue or model type) or longitudinal changes in selection within a patient or model by differential measurement between two or more samples.

[0119] FIG. 9 shows exemplary processes of DepSeq workflow to detect selection of epigenetic modifications of DNA and / or RNA, either with conversion of CpG nucleotides or directly through third-generation sequencing.

[0120] FIG. 10 shows exemplary results of cfDNA size distribution, a) Before linker demonstrating the primary peak near 167bp, and a secondary peak of double length at 334bp. b) After linker attachment, before affinity selection. A peak of single reads with relevant adapters is present at 351bp, and the connected peak is visible at 563bp. A small peak is visible at 978bp representing double length cfDNA. c) After linker attachment and affinityselection the dominant peak of 556bp consists of connected reads, and a secondary peak at 945 bp is present representing connected double length cfDNA.

[0121] FIG. 11 shows exemplary results of sequencing and consensus example using Illumina Sequencing. Top panel shows reads before consensus, and bottom shows the corresponding reads after consensus. Mismatches with reference are marked as solid colors, “N” bases are shaded with a hatch pattern.

[0122] FIG. 12 shows exemplary results of sequencing and consensus example using Illumina sequencing in the EGFR genomic region. Top panel shows reads before consensus, and bottom shows the corresponding reads after consensus. Mismatches with reference are marked as solid colors, “N” bases are shaded with a hatch pattern.

[0123] FIG. 13 shows exemplary results of sequencing and consensus example using Singular sequencing. Top panel shows reads before consensus, and bottom shows the corresponding reads after consensus. Mismatches with reference are marked as solid colors, “N” bases are shaded with a hatch pattern.

[0124] FIG. 14 shows exemplary results of Nanopore sequencing without conversion showing consensus. Top panel shows reads before consensus, and bottom shows the corresponding reads after consensus. Mismatches with reference are marked as solid colors, “N” bases are shaded with a hatch pattern.

[0125] FIG. 15 shows exemplary results of Nanopore sequencing with native (original) base modifications (methylation) consensus. Top panel shows reads before consensus, and bottom shows the corresponding reads after consensus. Mismatches with reference are marked as solid colors, “N” bases are shaded with a hatch pattern.

[0126] FIG. 16 shows exemplary results of Nanopore sequencing after conversion showing base modification (methylation) locations and converted bases. Mismatches with reference are marked as solid colors.

[0127] FIG. 17 shows exemplary results of relative error rates across the length of the read (a) before and (b) after removal of DNA ends that can contain higher levels of damage.

[0128] FIG. 18 shows exemplary results of mutational spectra for Nanopore long-read sequencing, (a) All 3 base contexts, (b) 3 base contexts excluding C->T and G->A events.

[0129] FIG. 19 shows exemplary results of mutational spectra for Singular sequencing, (a) All 3 base contexts, (b) 3 base contexts excluding C->T and G->A events.

[0130] FIG. 20 shows exemplary results of mutational spectra for Illumina short-read sequencing, (a) All 3 base contexts, (b) 3 base contexts excluding C->T and G->A events.

[0131] FIG. 21 shows exemplary box plots with single molecule mutational frequencies for 12 mutational contexts, (a) Oxford Nanopore long-read sequencing, (b) Singular sequencing, (c) Illumina short-read sequencing.

[0132] FIG. 22 shows exemplary DepSeq sequencing of candidate genes in cell lines with high rates of synonymous mutations (negative selection), a) RUNX2 from a HCC1395 cell line, b) HTR1B from a A431 cell line.

[0133] FIG. 23 shows exemplary DepSeq sequencing of candidate genes in a HCC1395 cell line, genes showing high rates of nonsynonymous mutations (positive selection), a) H4C4 b) FOXA1.

[0134] FIG. 24 shows exemplary DepSeq sequencing results in a A431 cell line. Plot of dN / dS by gene, showing candidate genes for positive selection (high dN / dS), no selection (dN / dS ~ 1), and negative selection (low dN / dS).

[0135] FIG. 25 shows exemplary DepSeq sequencing resulting in error rates reaching IO'7a) Read structure example for long read sequencing, b) Mutational spectra by 3 base context, c) Mutational frequencies by substitution type.DETAILED DESCRIPTION

[0136] The following description is presented to enable a person of ordinary skill in the art to make and use the various embodiments. Descriptions of specific devices, techniques, and applications are provided only as examples. Various modifications to the examples described herein will be readily apparent to those of ordinary skill in the art, and the general principles defined herein may be applied to other examples and applications without departing from the spirit and scope of the various embodiments. Thus, the various embodiments are not intended to be limited to the examples described herein and shown but are to be accorded the scope consistent with the claims.

[0137] Genetic dependency identification or screening is the process of identifying genes or genomic regions that are essential for the survival and proliferation of cells. These genes are known as genetic dependencies or simply dependencies.

[0138] The present disclosure is based on the development of a superior technology for identifying genetic dependencies with surprisingly low error rates (e.g., at least 10A-6 orlower). Termed DepSeq (Dependency Sequencing), this technology involves compositions and methods for accurate, efficient, and cost-effective identification of dependencies by using ultra-deep sequencing. This technology can be directly applied to a living tumor, disease or normal tissue, as well as any cells, animal or ex vivo models to directly profile a patient’s tumor, tissue, or models, providing an advantage over existing technologies that can only identify dependencies in vitro. FIG. 1 shows an exemplary DepSeq process schema.I. Constructs

[0139] The constructs described herein are crucial to DepSeq in identifying dependencies with ultra-low error rates. Without wishing to be bound by any theory, the disclosed constructs have two complementary strands of a nucleic acid physically linked together, which are then opened — e.g., either by strand nicking or treatment to de-hybridize — for sequencing. This enables accurate sequencing of each sequence twice directly (forward strand and reverse strand) without the use of in silico sequence matching that would otherwise increase cost and sequencing errors.

[0140] Accordingly, in one aspect, the present disclosure provides a construct for sequencing, the construct comprising: a double-stranded polynucleotide having first and second strands, each strand having first and second ends and a sequence of bases extending therebetween, wherein the sequence of bases in the first strand are at least partially hybridized to the sequence of bases in the second strand; wherein at least one nucleotide of the first strand is connected through a non-base pairing mechanism to at least one nucleotide of the second strand; and wherein the first and second strands are configured to irreversibly de-hybridize from each other when subject to a treatment that weakens or prevents hybridization.

[0141] Various double-stranded polynucleotides may be used for the constructs. In some embodiments, the double-stranded polynucleotide is DNA, RNA a DNA-RNA hybrid, phosphorothioate nucleic acid, locked nucleic acid (LNA), morpholino nucleic acids (MNAs), or a peptide nucleic acid (PNA).

[0142] The double-stranded polynucleotide may be repaired and / or pre-processed prior to connecting the two strands. In some embodiments, the double-stranded polynucleotide is end- repaired, end-blunted, nick-repaired, ddNTP-incorporated, dA-tailed, and / or middle-base or end-tail modified with functionalized bases for crosslinking, click-chemistry, or di-sulfide.

[0143] When two strands of a polynucleotide are partially hybridized, it means that certain regions of the strands are properly base-paired (hybridized) according to Watson-Crick basepairing rules (A with T / U, C with G), whereas other regions remain single-stranded and unpaired.

[0144] A non-base pairing mechanism in polynucleotides refers to an alternative way that nucleotides or polynucleotide strands can interact without following the traditional Watson- Crick base pairing rules (A-T and G-C). The connection of the two strands of the doublestranded polynucleotide through a non-base pairing mechanism may be achieved in various approaches. In some embodiments, the non-base pairing mechanism connects one end of the first strand with the distant end of the second strand (e.g., 5’ of the first strand to 5’ of the second strand, or 3’ of the first strand to 3’ of the second strand).

[0145] In some embodiments, the non-base pairing mechanism is a linker, a sugar unit connection, a phosphate connection, or any combination thereof. In some embodiments, the linker is a polyamide, polyimide, oligoethylene glycol or another polymer or oligomer linker, hairpin loop, or a modified or natural nucleotide that can be crosslinked with epoxy, cisplatin, carboplatin or derivatives, click-chemistry or disulfide bond. Further linker materials may include, e.g., alkyl cellulose, carboxyl ethyl cellulose, cellulose acetate, cellulose acetate butyrate, cellulose acetate phthalate, cellulose ethers, cellulose esters, cellulose propionate, cellulose sulphate sodium salt, cellulose triacetate, chitosan, dextran, ethyl cellulose, hydroxyalkyl celluloses, hydroxybutyl methyl cellulose, hydroxypropyl cellulose, hydroxypropyl methyl cellulose, methyl cellulose, nitro celluloses, poly- or oligo(alkylene alkylate), oligo(butylmethacrylate), oligo(caprolactone), oligo(caprolactone) / oligo(ethylene glycol) cooligomer, oligo(dioxanone), oligo(ethylmethacrylate), oligo(ethylene oxide), oligo(ethylene terephthalate), oligo(glycolic acid), oligo(glycolide), oligo(glycolide) / oligo(ethylene glycol) cooligomer, oligo(hydroxybutyrate), oligo(imides), oligo(isobutyl acrylate), oligo(isobutylmethacrylate), oligo(isodecylmethacrylate), oligo(isopropyl acrylate), oligo(lactide), oligo(lactide-co-caprolactone), oligo(lactide-co- glycolide), oligo(lactic acid), oligo(lactic acid-co-glycolic acid), oligo(lactic acid-co-glycolic acid) / oligo(ethylene glycol) cooligomer, oligo(lactide) / oligo(ethylene glycol) cooligomers, oligo(methyl acrylate), oligo(methyl methacrylate), oligo(octadecyl acrylate), oligo(orthoester), oligo(phenyl methacrylate), oligo(phosphazene), oligo(propylene oligoethylene glycol), oligo(styrene), oligo(vinyl acetate), oligo(vinyl alcohols), oligo(vinyl chloride), oligocarbonates, oligocaprolactone, oligoglycolides, oligo(hydroxybutyrate), oligomer acrylic and methacrylic esters, oligoesteramide, oligoanhydride, oligopropylene, oligo- and poly-siloxanes. FIG. 2 shows exemplary embodiments of the linker configurationand structure. FIG. 3 shows exemplary click chemistry linkers. FIG. 4 shows an exemplary process of DepSeq for fast profiling, in which two nucleotides from opposite strands are crosslinked.

[0146] The linker may be further configured to bind to a solid support, a bead, or an affinity tag to facilitate selection of the molecule. For instance, in some embodiments, the linker is configured to attach to desthiobiotin for selection with a streptavidin bead and further optional replacement desorption with biotin. FIG. 5 shows an exemplary process for increased yield in DepSeq library, in which affinity tag attachment at time of construct assembly allows for purification and high yield results. Linker attachment and affinity selection can also be used for size selection. FIG. 10 shows exemplary results of cfDNA size distribution using linker attachment and affinity selection.

[0147] Opening up the base paring of the double-stranded polynucleotide may be achieved via various means, including e.g., by treatment to weaken hybridization of, or to dehybridize, the double strands. In some embodiments, the treatment is nucleotide substitution conversion, an adenosine deaminase acting on RNA (ADAR) and DNA (TdAR) enzyme families, inosine modifications, a chemical treatment, or any combination thereof. In some embodiments, the treatment is conversion of cytosine into uracil through deamination, NaOH, formamide, a base, and / or a helicase. FIG. 6 shows an exemplary treatment process to de- hybridize / weaken hybridization. Through a process comprising of a strand copying (e.g., PCR, reverse description, RNA transcription) with a non-natural or natural nucleotide or nucleotide-analogue that facilitates de-hybridization naturally or through another process of chemical, biochemical or physical process or conversion or any combination thereof

[0148] Opening up the base paring of the double-stranded polynucleotide can also be achieved by cutting or strand nicking, e.g., using transposases or endonucleases. In some embodiments, the double-stranded polynucleotide is cut either on one strand, both strands or multiple locations of the strand. In some embodiments, the cut double-stranded polynucleotide is linearized. In some embodiments, the double-stranded polynucleotide is circularized, which can then be used as a template for rolling circle amplification (RCA) or a similar technique.

[0149] Adapters may be added to aid sequencing. In some embodiments, at least one end of the double-stranded polynucleotide is attached with a pair of adapters. In some embodiments, the pair of adapters are UMI (unique molecular identifier) adapters or UDI (unique dualindex) adapters. In some embodiments, the pair of adapters are further attached with PCR primers.II. Ultra-Deep Sequencing

[0150] Ultra-deep sequencing is another important factor contributing to the desired results of low error rates. Conventional next generation sequencing (NGS) exhibits inherent error rates between 0.01% and 1%, presenting significant challenges for accurate genomic analysis. By way of example, with circulating cell-free DNA (cfDNA) levels typically limited to 15-30 ng / mL of plasma, achieving exceptionally low error rates becomes crucial for accurate genomic data interpretation.

[0151] Accordingly, in another aspect, provided herein is a method for sequencing with low error rates, the method comprising: providing a construct of any one of the preceding embodiments; and performing sequencing on the construct. FIG. 7 shows an exemplary process of performing DepSeq on patient sample and identifying dependencies of the sample through ultra deep sequencing. FIG. 8. shows an exemplary process of DepSeq workflow to identify differential dependencies (treatment, location, exposure, immune pressure, tissue or model type) or longitudinal changes in selection within a patient or model by differential measurement between two or more samples.

[0152] DepSeq requires deep sequence depths. In some embodiments, the method results in a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher.

[0153] Such ultra-deep sequencing contributes to the desired low rates of sequencing error. In some embodiments, the method results in an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

[0154] When applicable, duplex sequencing is used to sequence the disclosed constructs. Duplex sequencing reduces error rates by validating mutations present on both DNA strands. Early duplex sequencing iterations used unique molecular identifiers (UMI) without physically linking sense and antisense strands, requiring very deep sequencing for strand identification. These early approaches improved error rates by magnitudes of up to lOOOx. Recent advancements, such as Concatenating Original Duplex for Error Correction (CODEC), enable low abundance mutation detection at reduced sequencing depth and cost, though require tedious preparation of multipart adapters and can suffer from presence of multiple DNA byproducts and are without methylation / modification detection capabilities.In some embodiments, the duplex sequencing comprises using circular strand displacement polymerization reaction (CSDPR), rolling circle amplification (RCA), CODEC, Ultima genomics duplex, or any combination thereof.

[0155] In some embodiments, the sequencing comprises the use of single-cell sequencing. Single-cell sequencing is a powerful technique used to analyze the genetic material of individual cells, providing insights into cellular diversity, gene expression patterns, and cellspecific mutations that are often masked in bulk cell populations. This technique isolates single cells and sequences the amplified DNA or RNA, enabling detailed study of variations across individual cells within a complex tissue or organism. DepSeq allows for detection of mutations, methylation and modification changes in single cells and low nucleic acid yield applications.

[0156] The ultra-deep sequencing of the constructs may be used on a variety of sequencing platforms, including e.g., next generation sequence (NGS) platforms and third generation sequence platforms. In some embodiments, the sequencing comprises using a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Pacific Biosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

[0157] In yet another aspect, the disclosure provides a kit for sequencing with low error rates, the kit comprising: a construct of any one of preceding embodiment; reagents for preparing a sequencing library; adaptors and / or linkers for duplex sequencing; and instructions for performing a sequencing method.

[0158] In some embodiments of the kit, the reagents comprise end-repair reagents, dA-tailing reagents, nick-repair reagents, ddNTPs, adapter ligation reagents, linker ligation or attachment reagents, crosslinking reagents, PCR primers, genomic target-enrichment reagents, or any combination thereof.III. Sequencing Data Analysis

[0159] In some embodiments, the method further comprises obtaining sequencing data from the sequencing platforms and identifying cancer and tissue dependencies, cancer, disease and tissue driving mutations, and / or treatment resistance mutations based on the sequencing data.

[0160] Accordingly, in yet another aspect, the disclosure provides a method for analyzing sequencing data with low error rates, the method comprising: (a) receiving, via an input device, a set of sequencing data associated with a sample; (b) automatically inputting, via aprocessor, the set of sequencing data to an algorithm, wherein the algorithm is configured to output a set of mutations, genes, or genomic regions corresponding to the cancer dependencies of the cancer sample; and (c) outputting the set of mutations via an output device. The sample can be various suitable samples, including for example, a cancer sample, a normal tissue sample, a bodily fluid or other patient material sample, a cell line or an animal model sample.

[0161] In some embodiments of the method, the algorithm comprises a negative binomial, Poisson model or log normal Poisson model for background selection, mutation rate identification, and / or methylation status quantification. FIG. 9 shows exemplary processes of DepSeq workflow to detect selection of epigenetic modifications of DNA and / or RNA, either with conversion of CpG nucleotides, C nucleotides, A>I conversion or other conversion process or directly through third-generation sequencing.

[0162] In some embodiments of the method, the algorithm utilizes local clustering, sub- clonality, copy-number information, structural variation, transcription-coupled repair, replication time information, mutation spectra and processes, or any combination thereof.

[0163] In some embodiments of the method, the algorithm is implemented using high- performance computing or a cloud system.

[0164] In some embodiments, the method further comprises annotating the set of mutations with amino acid changes, mutational signature profiles, genomic elements positions and effects, or any combination thereof.

[0165] In some embodiments, the method further comprises identifying genomic regions under positive negative or neutral selection based on the annotation of the set of mutations.

[0166] In some embodiments, the method further comprises calculating dN / dS rates, mutation effect sizes and statistical significance, copy-number variations (CNVs), or any combination thereof. The dN / dS rate (also called the nonsynonymous / synonymous substitution ratio) is a measure used in genetics to infer the selection pressure acting on protein-coding genes. It compares the rate of nonsynonymous mutations (dN) that change amino acids in a protein sequence with the rate of synonymous mutations (dS) that do not change the protein. When dN / dS < 1 : Indicates purifying (negative) selection, where deleterious mutations are selected against to preserve protein function. When dN / dS = 1 : Indicates neutral evolution, where mutations are neither selected for nor against, often seen in non-functional regions. When dN / dS > 1 : Indicates positive (adaptive) selection, wheremutations provide an advantage and are selected for, driving adaptive changes in the protein.

[0167] In some embodiments, the method further comprises at least one of the following: a) determining single cell mutations; b) calculating statistical depletion of a specific mutation type; c) identifying genomic regions of negative selection from mutation due to fitness constrains of the growing cancer or non-cancer cells; d) identifying genomic regions of positive selection from cancer drivers, growth drives, and / or resistant events; e) summarizing information in a report; f) identifying cancer vulnerabilities and druggable targets in the genome; and g) profiling of selection before or after treatment from in vivo tissue, ex vivo tissue, cell lines, blood, nucleic acids, circulating nucleic acids, circulating cells, patient derived material, patient derived models, animal models, or organoids.

[0168] In some other embodiments, the method further comprises comprising at least one of the following: a) estimating the copy-number profile of cancer based on selected reads; b) identifying highly mutable regions with a high-likelihood of selection signals; c) merging signals of methylation, fragmentation, copy-number, and mutation for accurate estimation of selection signals; d) selecting regions with clinical significance and / or mutations that are oncogenic or resistant; e) profiling using panels based on gene, gene region, non-coding region, and / or whole exome / genome screens; f) estimating cancer cell fraction and frequency of identified mutations to discover selection signals and changes thereof over time; and estimating selection signals for normal tissues, jointly or independently for mixtures of cells and tissue types.IV. Software

[0169] Where applicable, any of the aforementioned methods of present disclosure may be implemented as computer program processes that are specified as a set of instructions recorded on a non-transitory computer-readable storage medium (also referred to as a computer-readable medium-CRM).

[0170] Accordingly, in still another aspect, the present disclosure provides a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to: (a) receive, via an input device, a set of sequencing data associated with a cancer sample; (b) automatically input, via a processor, the set of sequencing data to an algorithm, wherein the algorithm is configured to output a set of mutations, genes, or genomic regions corresponding to the cancer dependencies of thecancer sample; and (c) output the set of mutations, genes, or genomic regions via an output device.

[0171] Examples of computer-readable storage media include RAM, ROM, read-only compact discs (CD-ROM), recordable compact discs (CD-R), rewritable compact discs (CD- RW), read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM), a variety of recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g, SD cards, mini-SD cards, micro-SD cards, etc.), magnetic and / or solid state hard drives, ultra-density optical discs, any other optical or magnetic media, and floppy disks. In some embodiments, the computer-readable storage medium is a solid-state device, a hard disk, a CD-ROM, or any other non-volatile computer-readable storage medium.

[0172] The computer-readable storage media can store a set of computer-executable instructions (e.g, a “computer program”) that is executable by at least one processing unit and includes sets of instructions for performing various operations.

[0173] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a standalone program or as a module, component, or subroutine, object, or other component suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subprograms or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network. Examples of computer programs or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter.

[0174] As used herein, the term “software” is meant to include firmware residing in readonly memory or applications stored in magnetic storage, which can be read into memory for processing by a processor. Also, in some implementations, multiple software aspects of the subject disclosure can be implemented as sub-parts of a larger program while remaining distinct software aspects of the subject disclosure. In some implementations, multiplesoftware aspects can also be implemented as separate programs. Any combination of separate programs that together implement a software aspect described here is within the scope of the subject disclosure. In some implementations, the software programs, when installed to operate on one or more electronic systems, define one or more specific machine implementations that execute and perform the operations of the software programs.

[0175] Any suitable machine learning models may be used with the methods of the present invention and be implemented as computer program processes that are specified as a set of instructions recorded on a computer-readable storage medium. In some embodiments, the model is a discriminative model or a generative model.V. Device

[0176] Further, any one of the preceding methods of the present disclosure may be implemented in one or more computer systems or other forms of apparatus. Examples of apparatus include but are not limited to, a computer, a tablet personal computer, a personal digital assistant, and a cellular telephone.

[0177] In some embodiments, the disclosure further provides an electronic device, comprising: a display; one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for (a) receiving, via an input device, a set of sequencing data associated with a cancer sample; (b) automatically inputting, via a processor, the set of sequencing data to an algorithm, wherein the algorithm is configured to output a set of mutations, genes, or genomic regions corresponding to the cancer dependencies of the cancer sample; and (c) outputting the set of mutations, genes, or genomic regions via an output device.

[0178] In some embodiments, the electronic device may be a server computer, a client computer, a personal computer (PC), a user device, a tablet PC, a laptop computer, a personal digital assistant (PDA), a cellular telephone, or any machine capable of executing a set of instructions, sequential or otherwise, that specify actions to be taken by that machine. In some embodiments, the electronic device may further include keyboard and pointing devices, touch devices, display devices, and network devices.

[0179] As used herein, the terms “computer”, “processor”, and “memory” all refer to electronic or other technological devices. These terms exclude people or groups of people. For the purposes of the specification, the terms “display” or “displaying” means displayingon an electronic device. As used in this specification and any claims of this application, the terms “computer readable medium” and “computer readable media” are entirely restricted to tangible, physical objects that store information in a form that is readable by a computer. These terms exclude any wireless signals, wired download signals, and any other ephemeral signals.

[0180] To provide for interaction with a user, implementations of the subject matter described in this specification can be implemented on a computer having a display device described herein for displaying information to the user and a virtual or physical keyboard and a pointing device, such as a finger, pencil, mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speed, or tactile input.VI. Applications of DepSeq

[0181] DepSeq can be used for multiple dependency screening including identification of drug combinations, screening for drug response in primary patient samples, clinical profiling of individuals for assignment to clinical trials, interrogation of cancer cell in animal models for dependencies and drug responses, changes in selection environment in tumors upon treatment, identification of immune response and immune pressures in vivo, far and close to treatment with or without immunotherapy, targeted therapy, chemotherapy, radiation and any other treatment modalities. This assay can be used in conjunction, additionally as prescreening for CRISPR and CRISPRi based assays. This assay can be used to assess treatment response and ability of tumor to evade specific treatments. Such techniques can be applied to living tumors, disease tissues, blood, cell free nucleic acids, fecal matter, body fluid, patient derived material or normal tissues to profile mutational, epi -mutational, structural and other genomic landscapes and identify therapeutic targets toward cancer disease treatment.

[0182] While a number of exemplary aspects and embodiments have been discussed above, those of skill in the art will recognize certain modifications, permutations, additions and subcombinations thereof. It is therefore intended that the following appended claims and claims hereafter introduced are interpreted to include all such modifications, permutations, additions and sub-combinations as are within their true spirit and scope.EXAMPLES

[0183] The following examples are offered to illustrate provided embodiments and are not intended to limit the scope of the present disclosure.Example 1. Development of DepSeq for In Vivo Dependency Identification

[0184] This example demonstrates the application of DepSeq for identifying dependencies / vulnerabilities of cancer tissue through ultra-deep sequencing.A, Sample Requirements and Collection Parameters

[0185] To achieve the goals of DepSeq assay and identify positive and negative selection drivers in vivo, ex vivo or in vitro, specific sample parameters have been established. The methodology requires obtaining DNA from anywhere between 100,000 to 10 million or more individual cell genomes. Based on the determination that approximately 1000 cell genomes are contained in 6 to 10 nanograms of DNA, this constitutes a requirement of anywhere from one microgram to 100 micrograms of tumor DNA. Typically, a cubic centimeter of tumor tissue may contain between a billion and 10 billion cells.

[0186] Sample types that can be utilized include:• Surgical specimens• Biopsies. FNAs• Blood collections• Urine collections• Fecal collections• Saliva collections• Body fluid collections• Cell lines• Patient-derived models• Patient-derived materials• Animals or animal models• Bacterial, viral, fungal or other organism cultures or specimens

[0187] Preservation methods can include:• Fresh• Fresh frozen• Formalin fixed. FFPE• Blood biopsy. EDTA• Methanol / Ethanol / Alcohol fixationB, Biological Considerations

[0188] The methodology utilizes selection pressures that exist in living tumors. When cells divide, they introduce mutations, with cancer cells typically dividing more frequently than normal cells. Each division typically introduces anywhere from a few to tens of new mutations. In a population of billion cells, this results in likely mutation of every single gene and genomic position multiple times.

[0189] Genomic or epigenomic mutations that lead to fitness disadvantage compared to surrounding cells result in depletion of cells with specific mutations or functional losses or changes, thus "protecting" some regions from mutations in the population. Genes or genomic regions that when mutated lead to increase of frequency of specific events are often drivers and give fitness advantage to the cells (positive selection).D. Technical Implementation

[0190] The technology provides accurate sequencing of up to 1 million copies of DNA / RNA or more or more from the same genomic location. This is achieved through sequencing technologies that allow reading both strands of DNA or RNA / DNA hybrid or dsRNA simultaneously to high fidelity (i.e., duplex sequencing) to lower the error rate to 10A-6 or lower. This can be accomplished through:• Special library preparation techniques for sequencing by synthesis technologies (Illumina, ULTIMA etc.)• Native approaches for Oxford Nanopore Technologies• PAC Bio sequencers

[0191] This method can be directly applied to:• Normal tissues or patient derived material• Body fluid• Animal or animal model• Bacterial, viral, fungal or other organism cultures or specimens

[0192] The method also enables identification of tissue-specific vulnerability that allows for "therapeutic windows" and an ability to selectively kill cancer or diseased cells. It allows for measurement of deferential selection effects on different tissues and though the aging or neurodegeneration process.

[0193] This example demonstrates that the DepSeq methodology provides a novel approachfor directly profiling cellular dependencies in patient samples, offering significant advantages over traditional CRISPR-based screening methods in terms of in vivo applicability and direct tissue analysis capabilities.Example 2. Development of High-Accuracy Sequencing Protocols

[0194] This example demonstrates the development and implementation of high-accuracy sequencing protocols for achieving ultra-low error rates.

[0195] Human genomic DNA was sheared to an average size of 250 base pairs using focused ultrasonication (Covaris S2, programmed to 200 target base pairs, 10% duty cycle, 5 intensity, 200 cycles per burst for 180 seconds). Mung Bean blunt end reaction was performed to remove single-stranded DNA ends post-ultrasonication. Mung Bean nuclease (NEB, M0250S) was diluted to 1 U pl-1 in lx Mung Bean nuclease buffer. Each reaction was done using 50 ng of genomic DNA diluted in 10 pl TE lx, 2.9 pl 10x Mung Bean nuclease buffer, 1 pl diluted Mung Bean nuclease and 16.1 pl nuclease free water (NEW). Mung Bean reaction was incubated at 30°C for 30 minutes. After the incubation, 1 pl 0.3% SDS was added to stop the reaction. DNA was purified using 77.5 pl Ampure XP beads (Beckman Coulter, A63880) and eluted in 12pl NFW.

[0196] Next, phosphorylation was performed with 10 pl DNA from the Mung Bean reaction,1.5 pl NEBuffer 4 (NEB, B7004S), 1.5 pl 10 mM ATP (ThermoFisher, R0441), 0.6 pl T4 Polynucleotide Kinase (ThermoFisher, EK0031) and 1.4 pl NFW. Phosphorylation reaction was incubated at 37°C for 30 min. Subsequently, A-tailing and DNA gap-protection was performed by adding 13 pl from the previous phosphorylation reaction to 6.2 pl NEBuffer 4,7.5 pl 1 mM dATP / ddBTP (NEB, I N0440S / GE Healthcare, 27204501), 0.75 pl Klenow fragment (3' to 5' exo-, NEB, M0212L) and 47.55 pl NFW. This reaction was then incubated at 37 °C for 30 min.

[0197] Thereafter, hairpin adapter ligation was performed using 73 pl from the previous A- tailing reaction with 37.5 pl of ligation master mix (NEB), 1.25 pl Ligation enhancer (NEB) X NFW and X adapter (add 1.5x 3’-5’ moles of DNA calculation) (IDT, example sequences: ATGACGATGCGTTCGAGCATCGUCAUT, ATGACGATGCNNNCGAGCATCGUCACT, ATGACGATCCGTACGGCATCGCCACT). Ligation reaction was incubated at 20°C for 60 minutes. DNA was purified using 105.3 pl Ampure XP beads and eluted in 52pl NFW. 50 pl of the purified DNA was then taken to forkhead adaptor ligation using 30pl ligation master mix (NEB, E7645S), Ipl ligationenhancer, 5 .1 forkhead adapter (IDT, example sequences forward - ACACTCTTTCCCTACACGACGCTCTTCCGATC*T and reverse - GATCGGAAGAGCACACGTCTGAACTCCAGTCA, with and without methylation, TruSeq or Twist Biosciences UMI adaptors) and 14 pl NFW. Samples were incubated at 20°C for 15 min. DNA was purified using 120 pl Ampure XP beads and eluted in 35 pl NFW.

[0198] Next, methylated cytosines were protected through oxidation. 3 Opl DNA from the previous reaction were added to 10 pl reconstituted TET2 5* supplement buffer (10 mM a- Ketoglutarate, 0.25 M Tris-HCl pH 8.0, 10 mM ATP), 1 pl UDP glucose, 1 pl lOOmM DTT, 1 pl T4 beta-glucosyltransferase (both catalog no. EO0831, Thermo Fisher Scientific) and 2 pl TET2. Oxidation reaction was incubated at 37°C for 60 minutes. DNA was purified with 90 pl Ampure XP beads and eluted in 32pl NFW. Then, protected cytosines were converted into uracils through deamination using 3 Opl from the previous reaction, 12.7 NFW, 17.5 4x APOBEC buffer (200 mM BisTris pH6.1, 0.4% Tween), 1.8 100 mM ATP (catalog no. A6559, Sigma), 3.5pl 100 mM MgCh (catalog no. M1028, Sigma), 2 pl APOBEC-A3A, 2.5 pl UvrD Helicase. Deamination reaction was incubated at 37°C for 90 minutes. DNA was purified with 6.5 pl DA cleanup and 91.8 pl Ampure XP beads. It was then eluted in 20pl NFW.

[0199] Purified DNA was then taken into library amplification using 20pl 2X KAPA HiFi U+ Polymerase (Roche, KK2801) and 5 pl Abclonal Unique Dual Index Primers for Illumina (Abclonal, RK21624_SetA) or other similar TruSeq compatible primers. The library was sequenced to ultra-high depth on Illumina Next seq or Novoseq sequencers. In another set of experiments the duplex DNA was selected on a targeted methylated panel (TWIST) to sequence only parts of the gene to ultra-high depth. For examples a panel selected sample ~1.5 MB panel coverage was sequenced to 300-500 thousand depth.Example 3. Consensus Calling and Error Rate Analysis

[0200] This example demonstrates an exemplary bioinformatics pipeline utilized for analyzing DepSeq data.

[0201] After performing quality control and trimming, the modified duplex reads in the FASTQ files were aligned to the hg38 reference genome using Biscuit (https: / / huishenlab.github.io / biscuit / ) as paired reads. PCR duplicates were identified and removed with Dupsifter (https: / / github.com / huishenlab / dupsifter). Reads that were misaligned, unpaired, or simplex were filtered out using SAMtools(http : / / github .com / samtool s / ) .

[0202] Consensus calling was carried out using a Python script that compares the nucleotide sequences of the two duplex reads at each position. This process considers nucleotide quality scores, read orientation, and alignment information (including CIGAR, MC, and MD tags) in relation to the reference sequence to determine the most probable nucleotide. The specific resolution rules utilized in this algorithm are outlined in Table 1.Table 1. Resolution Rules for the Consensus Calling Algorithm(a) If the mapped bases are identical and match the reference base, the consensus base is set to that value, (b) If the mapped bases are the same between the duplex reads but do not match the reference, a true mutation is identified, (c) If the base in read 1 matches the reference but there is an error in read 2, the consensus will reflect the reference base, (d) If there is an error in read 1 but read 2 matches the reference, the consensus base will be set to N. (e) If the base in read 1 is on the forward strand and represents a converted C, the consensus base will be set to C. (f) If the base in read 1 is on the reverse strand and represents a converted C, the consensus base will be set to G. (g) If the base in read 2 is on the reverse strand and represents a converted C, the consensus base will be set to G. (h) If the base in read 2 is on the forward strand and represents a converted C, the consensus base will be set to C. (i) If there are errors in both read 1 and read 2 bases, the consensus base will be set to N.

[0203] FIG. 11 shows exemplary results of sequencing and consensus example using Illumina Sequencing. Top panel shows reads before consensus, and bottom shows the corresponding reads after consensus. Mismatches with reference are marked as solid colors, “N” bases are shaded with a hatch pattern.

[0204] FIG. 12 shows exemplary results of sequencing and consensus example using Illumina sequencing in the EGFR genomic region. Top panel shows reads before consensus,and bottom shows the corresponding reads after consensus. Mismatches with reference are marked as solid colors, “N” bases are shaded with a hatch pattern.

[0205] FIG. 13 shows exemplary results of sequencing and consensus example using Singular sequencing. Top panel shows reads before consensus, and bottom shows the corresponding reads after consensus. Mismatches with reference are marked as solid colors, “N” bases are shaded with a hatch pattern.

[0206] FIG. 14 shows exemplary results of Nanopore sequencing without conversion showing consensus. Top panel shows reads before consensus, and bottom shows the corresponding reads after consensus. Mismatches with reference are marked as solid colors, “N” bases are shaded with a hatch pattern.

[0207] FIG. 15 shows exemplary results of Nanopore sequencing with native (original) base modifications (methylation) consensus. Top panel shows reads before consensus, and bottom shows the corresponding reads after consensus. Mismatches with reference are marked as solid colors, “N” bases are shaded with a hatch pattern.

[0208] FIG. 16 shows exemplary results of Nanopore sequencing after conversion showing base modification (methylation) locations and converted bases. Mismatches with reference are marked as solid colors.

[0209] Consensus reads underwent final quality control by assessing error rates across the entire length of the reads. This involved obtaining a distribution of variants based on read length (see FIG. 17). In cases where there was an excess of variants at the ends of the reads, the reads were trimmed, and the consensus calling step was repeated.

[0210] The resulting files were then sorted and indexed using SAMtools, followed by variant calling through the mpileup function in SAMtools. Variants were annotated using the GnomAD database, and germline SNPs with an allele fraction greater than 0.0001 were excluded.

[0211] Subsequently, mutation spectra plots were generated for each of the 96 mutation contexts (see FIGS. 18-20), characterized by one of six possible unique base pair changes (C:G > A:T, C:G > G:C, C:G > T: A, T: A > A:T, T: A > C:G, and T: A > G:C) and one of sixteen different trinucleotide sequence contexts and normalized by the total number of contexts.

[0212] To construct boxplots for each nucleotide in the consensus BAM file, the 3-nucleotidereference context was obtained. Subsequently, each of the 12 plots for 12 substitution subtypes (C>A, OG, OT, T>C, >. A, T>G, G>A, G>C, G>T, A>T, A>C, A>G) was generated based on the logarithmic ratios of the total number of mutations for each nucleotide in each context across all chromosomes to the overall number of contexts for that nucleotide in each chromosome. FIG. 21 shows exemplary box plots with single molecule mutational frequencies for 12 mutational contexts.

[0213] Results of this example show that DepSeq successfully identified mutation rates in many contexts that are below 10A-6. This level of accuracy allows for the detection of single molecule mutations and profiling for negative selection.Example 4. Cell Line Validation

[0214] This example demonstrate validation of DepSeq methodology using two established cancer cell lines: A431 (epidermoid carcinoma) and HCC1395 (breast cancer).

[0215] FIG. 22 shows exemplary DepSeq sequencing of candidate genes in cell lines with high rates of synonymous mutations (negative selection), a) RUNX2 from a HCC1395 cell line, b) HTR1B from a A431 cell line.

[0216] FIG. 23 shows exemplary DepSeq sequencing of candidate genes in a HCC1395 cell line, genes showing high rates of nonsynonymous mutations (positive selection), a) H4C4 b) FOXA1.

[0217] FIG. 24 shows exemplary DepSeq sequencing results in a A431 cell line. Plot of dN / dS by gene, showing candidate genes for positive selection (high dN / dS), no selection (dN / dS ~ 1), and negative selection (low dN / dS).

[0218] FIG. 25 shows exemplary DepSeq sequencing resulting in error rates reaching IO'7a) Read structure example for long read sequencing, b) Mutational spectra by 3 base context, c) Mutational frequencies by substitution type.

[0219] The data shows systematic identification of essential genes and context-specific dependencies, validating the DepSeq methodology's ability to identify cellular vulnerabilities.Example 5. Advanced DepSeq Sequencing Implementation for High-Accuracy Detection

[0220] This example demonstrates the development and implementation of specialized protocols for achieving ultra-low error rates in sequencing applications.

[0221] For high accuracy detection of modified reads in patient samples, both strands of the extracted DNA must be sequenced. Since real mutations and 5mC CpG methylation changes in the genome are present on both strands, this approach differentiates sequencing artifacts regardless of sequencing technology.A, DepSeq Protocol Development

[0222] FIG. 25 illustrates the DepSeq variant with Oxford Nanopore sequencing protocol developed for single-molecule accurate detection. The protocol demonstrates:• Single point substitution rates showing <10A-6 error in consensus reads• Several orders of magnitude improvement over current sequencing instruments• Efficient inference of read structure across sequenced samplesB, Technical Implementation

[0223] The DepSeq methodology employs multiple approaches for high-accuracy sequencing:1. Native DepSeq Sequencing:• Current PromethlON flow cells capable of native long read sequencing• Enhanced accuracy through complementary strand reading2. Enhanced Selection Methods:• Size selection of modified peak (~500bp for cfDNA)• Linker design utilizing dT-desthiobiotin / streptavidin / biotin selection• -100% enrichment of construct reads3. Alternative Duplex Methods for DepSeq for dependency identification:• CODEC sequencing• Rolling circle amplification• Ultima genomics duplex protocol

[0224] Taken together, these examples demonstrate successful implementation of the DepSeq technology for identifying cellular dependencies with unprecedented accuracy. The method achieves ultra-low error rates (10A-6 to 10A-7) and provides direct profiling capabilities in living, in vitro, and ex-vivo tissue samples. The technology offers significant advantages over traditional CRISPR-based screening methods in terms of cost, efficiency, and applicability to in vivo systems.

[0225] While the advantages and preferred embodiments of the present invention have been described hereinbefore, those skilled in the art should be understood that the above are merely several illustrative embodiments of the present invention without limiting the scope thereof, wherein various modifications, alterations or substitutions may be made to the specific components of the embodiments without departing from the spirit and scope of the invention and its claims.

Claims

CLAIMS1. A construct for sequencing, the construct comprising: a double-stranded polynucleotide having first and second strands, each strand having first and second ends and a sequence of bases extending therebetween, wherein the sequence of bases in the first strand are at least partially hybridized to the sequence of bases in the second strand; the polynucleotide having a first region comprising a connecting region, wherein at least one nucleotide of the first strand is irreversibly or reversibly connected through a non-base pairing mechanism to at least one nucleotide of the second strand; and the polynucleotide having a second region, outside the connecting region, wherein the first and second strands are configured to de-hybridize from each other when subject to a process that weakens or prevents hybridization.

2. The construct of claim 1, wherein the double-stranded polynucleotide is DNA, RNA a DNA-RNA hybrid, phosphorothioate nucleic acid, locked nucleic acid (LNA), morpholino nucleic acids (MNAs)or a peptide nucleic acid (PNA).

3. The construct of claim 1 or 2, wherein the double-stranded polynucleotide is end- repaired, end-blunted, nick-repaired, ddNTP-incorporated, dA-tailed, and / or middlebase or end-tail modified with functionalized bases for crosslinking, click-chemistry, or di-sulfide.

4. The construct of any of claims 1-3, wherein the at least one nucleotide of the first strand is irreversibly connected through a non-base pairing mechanism to at least one nucleotide of the second strand.

5. The construct of claim 4, wherein the irreversible connection is a polylactic acid (PLA), a polyamide, a polyimide, an oligoethylene glycol or another polymer or oligomer linker, a hairpin loop, or a modified or natural nucleotide crosslinked with epoxy, cisplatin, carboplatin or derivatives, click-chemistry or disulfide bond.

6. The construct of any of claims 1-3, wherein the at least one nucleotide of the first strand is or reversibly connected through a non-base pairing mechanism to at least one nucleotide of the second strand.

7. The construct of claim 6, wherein the reversible connection is an affinity-based linkage, a biotin or desthiobiotin linkage, an interconnected probe, a UV-, light- or temperature-cleavable linkage, or a chemically or enzymatically cleavable linkage.

8. The construct of any of claims 1-7, wherein the de-hybridization is irreversible.

9. The construct of claim 8, wherein the irreversible de-hybridization is by means of nucleotide substitution conversion, an adenosine deaminase acting on RNA (ADAR) or DNA (TdAR family) enzymes, inosine, a chemical modification, conversion of cytosine into uracil by cytosine deaminase, incorporation of a non-natural or natural nucleotide or a nucleotide-analogue that facilitates de-hybridization.

10. The construct of any of claims 1-7, wherein the de-hybridization is reversible.

11. The construct of claim 10, wherein the reversible de-hybridization is by means of cleavable nucleotide protection, NaOH, formamide, a base, solvent or temperature change, pH change, or a helicase, or any combination thereof.

12. The construct of any of claims 1-11, wherein the non-base pairing mechanism connects one end of the first strand with the opposite distant end of the second strand.

13. The construct of any of claims 1-12, wherein the non-base pairing mechanism is a linker, a sugar unit connection, a phosphate connection, an affinity, or physical connection or any combination thereof.

14. The construct of claim 13, wherein the linker is a polylactic acid (PLA), polyamide, polyimide, a hairpin loop, oligoethylene glycol or another polymer or oligomer linker, or a modified or natural nucleotide that can be crosslinked with epoxy, cisplatin, carboplatin or derivatives, click-chemistry or disulfide bond.

15. The construct of claim 13, wherein the linker is configured to bind to a solid support or a bead.

16. The construct of claim 13, wherein the linker is configured to attach to desthiobiotin or derivatives for selection with a streptavidin bead.

17. The construct of claim 14, wherein the construct is selected using a streptavidin bead.

18. The construct of any of claims 1-17, wherein the process includes one or more of nucleotide substitution conversion, an adenosine deaminase acting on RNA (ADAR) or DNA (TdAR family) enzymes, inosine, a chemical modification, or any combination thereof.

19. The construct of any of claims 1-18, wherein the process includes one or more of conversion of cytosine into uracil by cytosine deaminase or other chemical or biochemical process, NaOH, formamide, a base, and a helicase.

20. The construct of any of claims 1-19, wherein the process comprises a strand copying step incorporating a non-natural or natural nucleotide or a nucleotide-analogue that facilitates de-hybridization.

21. The construct of any of claims 1-20, wherein at least one end of the double-stranded polynucleotide is attached with a pair of adapters.

22. The construct of claim 21, wherein the pair of adapters are UMI adapters or UDI adapters, or any combination thereof.

23. The construct of claim 20 or 21, wherein the pair of adapters are further attached with PCR primers.

24. The construct of any of claims 1-23, wherein the double-stranded or one strand of the polynucleotide or the linker is cut.

25. The construct of claim 24, wherein the cut double-stranded polynucleotide is linearized.

26. The construct of any of claims 1-25, wherein the double-stranded polynucleotide is circularized.

27. The construct of any of claims 1-26, wherein the construct is configured for sequencing with a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher.

28. The construct of any of claims 1-27, wherein the construct is configured for sequencing at an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

29. The construct of any of claims 1-28, wherein the first and second strands are irreversibly de-hybridized from each other in the second region.

30. A method for sequencing a blood, cell, nucleic acid or tissue sample, the method comprising: a) providing a construct of any of claims 1-29; b) de-hybridizing the first and second strands so as to separate them from each other in the second region, and c) performing sequencing on the construct.

31. The method of claim 30, wherein the de-hybridizing is irreversible.

32. The method of claim 31, wherein the irreversible de-hybridizing is by means of nucleotide substitution conversion, conversion of cytosine into uracil by cytosine deaminase, an adenosine deaminase acting on RNA (ADAR) or DNA (TdAR family) enzymes, inosine modification, a chemical modification, or any combination thereof.

33. The method of claim 30, wherein the de-hybridizing is reversible.

34. The method of claim 33, wherein the reversible de-hybridizing is by means of cleavable nucleotide protection, NaOH, formamide, a base, a solvent or temperature change, a PH change, a helicase, or any combination thereof.

35. The method of any of claims 30-34, wherein the method results in a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher.

36. The method of any of claims 30-35, wherein the method results in an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

37. The method of any of claims 30-36, wherein the sequencing comprises using duplex sequencing.

38. The method of claim 37, wherein the duplex sequencing comprises using duplex adapter based sequencing, circular strand displacement polymerization reaction (CSDPR), rolling circle amplification (RCA), concatenating original duplex for error correction (CODEC), Ultima genomics duplex, or any combination thereof.

39. The method of any of claims 30-38, wherein the sequencing comprises using singlecell sequencing.

40. The method of any of claims 30-39, wherein the sequencing comprises using a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Pacific Biosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

41. The method of any of claims 30-40, further comprising obtaining sequencing data from the sequencing and identifying cancer or tissue dependencies, cancer or tissue expansion driving mutations, and / or cancer or disease resistance mutations based on the sequencing data.

42. The method of any of claims 30-41, wherein the strands in the second region are subjected to substituted conversion of one or more nucleotides in the second region, application of adenosine deaminase acting on RNA (ADAR) and DNA TdAR enzyme, inosine, a chemical modification, or any combination thereof.

43. The method of claim 42, wherein the polynucleotide is DNA, and the de-hybridization process comprises nucleotide substitution conversion, an adenosine deaminase acting on RNA (ADAR) or DNA TdAR enzyme, inosine, a chemical modification, or application of bisulfite conversion.

44. The method of any of claims 30-43, wherein the construct is sequenced using duplex sequencing.

45. The method of any of claims 30-44, wherein the duplex sequencing comprises using duplex adapter based sequencing, circular strand displacement polymerization reaction (CSDPR), rolling circle amplification (RCA), concatenating original duplex for error correction (CODEC), Ultima genomics duplex, or any combination thereof.

46. The method of any of claims 30-45, wherein the construct is sequenced using singlecell sequencing.

47. The method of any of claims 30-46, wherein the construct is subject to target capture, PCR enrichment, WES, CRISPR based enrichment, or any combination thereof.

48. The method of any of claims 30-47, wherein the construct is sequenced on a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Pacific Biosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

49. The method of any of claims 30-48, wherein the construct is provided with an irreversible connection in the second region, comprising a non-base pairing mechanism that connects at least one nucleotide of the second strand and at least one nucleotide of the first strand in the second region.

50. The method of claim 49, comprising a step of applying a transposase enzyme with oligonucleotide integration or without, a restrictase, a glycosylase, or other chemical or biochemical process to cut the second irreversible connection in the second region prior to sequencing.

51. The method of claim 50, comprising a step of applying a transposase enzyme with oligonucleotide integration or without, a restrictase, a glycosylase or other chemical or biochemical process to cut the linearized or circularly amplified molecule in one, two dissenting, or multiple distinct places within or outside of the original polynucleotide sequence.

52. The method of any of claims 30-51, wherein the construct is originated from a tumor or a somatic cell.

53. The method of claim 52, comprising identifying, from the sequencing performed on the construct, one or more genes or genomic regions that are dependencies of tumors or somatic tissue or mutations or epi-mutations depleted by the tumor.

54. A method for profiling dependencies, the method comprising: providing sequence data of a sample; and identifying genetic dependencies based on the sequence data, wherein the sequence data comprise a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher, wherein the sequence data results in an error rate of at least 10A-5 or lower, at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

55. The method of any of claim 54, wherein the sample is originated from a tumor or a somatic cell.

56. The method of claim 54 or 55, further comprising identifying, from the sequence data, one or more genes or genomic regions that are or have been dependencies of tumors or somatic tissue or mutations or epi-mutations depleted by the tumor.

57. The method of any of claims 54-56, further comprising identifying genomic regions of negative selection from mutations due to fitness constrains of growing cancer or non-cancer cells.

58. The method of any of claims 54-57, further comprising identifying genomic regions of positive selection from cancer or tissue drivers, growth drivers, and / or resistant events.

59. The method of any of claims 54-58, wherein the sequence data are from duplex sequencing.

60. The method of claim 59, wherein the duplex sequencing comprises using duplex adapter based sequencing, circular strand displacement polymerization reaction(CSDPR), rolling circle amplification (RCA), concatenating original duplex for error correction (CODEC), Ultima genomics duplex, or any combination thereof.

61. A kit for sequencing a polynucleotide , the kit comprising: a) a construct of any of claims 1-29; b) reagents for preparing a sequencing library; c) adaptors and / or linkers for construct sequencing; and d) instructions for performing a sequencing method.

62. The kit of claim 61, wherein the reagents comprise end-repair reagents, dA-tailing reagents, nick-repair reagents, ddNTPs, adapter ligation reagents, linker attachment reagents, PCR primers, genomic target-enrichment reagents, or any combination thereof.

63. The kit of claim 61 or 62, wherein the sequencing method comprises duplex sequencing.

64. The kit of any of claims 61-63, wherein the adaptors and / or linkers are for duplex sequencing.

65. The kit of any of claims 61-64, wherein the sequencing method is performed on a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Element Biosciences, Pacific Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

66. The kit of any of claims 61-65, comprising a de-hybridization agent comprising one or more of an agent configured to facilitate substitution conversion of a nucleotide on the first or second strand, an adenosine deaminase configured to act on RNA (ADAR) or DNA (TdAR family enzymes), inosine, a chemical modifier, or any combination thereof.

67. The kit of claim 66, wherein the de-hybridization agent comprises or more of a compound or enzyme configured to convert a cytosine on the first or second strand into uracil, NaOH, formamide, a base, and a helicase.

68. The kit of claim 66, wherein the de-hybridization agent comprises a non-natural or natural nucleotide or a nucleotide-analogue that facilitates de-hybridization for strand copying.

69. A method for analyzing sequencing data with low error rates, the method comprising:(a) receiving, via an input device, a set of sequencing data associated with a sample;(b) automatically inputting, via a processor, the set of sequencing data to an algorithm, wherein the algorithm is configured to output a set of mutations, genes, orgenomic regions corresponding to the dependencies of the sample; and (c) outputting the set of mutations via an output device.

70. The method of claim 69, wherein the set of sequencing data is obtained from a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Pacific Biosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

71. The method of claim 69 or 70, wherein the set of sequencing data is obtained from a sequencing protocol that results in a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher.

72. The method of any of claims 69-71, wherein the method results in an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

73. The method of any of claims 69-72, wherein the algorithm comprises a negative binomial, Poisson model or log normal Poisson model for one or more of background selection, mutation rate identification, and methylation status quantification.

74. The method of any of claims 69-73, wherein the algorithm utilizes local clustering, sub-clonality, copy-number information, structural variation, transcription-coupled repair, replication time information, mutation spectra and processes, or any combination thereof.

75. The method of any of claims 69-74, wherein the algorithm is implemented using high- performance computing or a cloud system.

76. The method of any of claims 69-75, further comprising annotating the set of mutations with amino acid changes, mutational signature profiles, genomic elements positions and effects, or any combination thereof.

77. The method of any of claims 69-76, further comprising identifying genomic regions under positive, negative or neutral selection based on the annotation of the set of mutations.

78. The method of any of claims 69-77, further comprising calculating dN / dS rates, mutation effect sizes and statistical significance, copy-number variations (CNVs), or any combination thereof.

79. The method of any of claims 69-78, comprising at least one of the following: a) determining single cell mutations; b) calculating statistical depletion or enrichment of a specific mutation type; c) identifying genomic regions of negative selection from mutation due to fitness constrains of the growing cancer or non-cancer cells;d) identifying genomic regions of positive selection from cancer drivers, growth drivers, and / or resistant events; e) summarizing information in a report; f) identifying cancer vulnerabilities, disease vulnerabilities and druggable targets in the genome; and g) profiling of selection before or after treatment from in vivo tissue, ex vivo tissue, cell lines, blood, nucleic acids, circulating nucleic acids, circulating cells, patient derived material, patient derived models, animal models, or organoids.

80. The method of any of claims 69-79, further comprising at least one of the following: a) estimating the copy-number profile of cancer based on selected reads; b) identifying highly mutable regions with a high-likelihood of selection signals; c) merging signals of methylation, fragmentation, copy-number, and mutation for accurate estimation of selection signals; d) selecting regions with clinical significance and / or mutations that are oncogenic or resistant; e) profiling using panels based on gene, gene region, non-coding region, and / or whole exome / genome screens; f) estimating cancer cell fraction and frequency of identified mutations to discover selection signals and changes thereof over time; and g) estimating selection signals for normal tissues, jointly or independently for mixtures of cells and tissue types.

81. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to perform the method of any of claims 69-80.

82. The non-transitory computer-readable storage medium of claim 81, wherein the set of sequencing data is obtained from a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Pacific Biosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

83. The non-transitory computer-readable storage medium of claim 81 or 82, wherein the set of sequencing data is obtained from a sequencing protocol that results in a sequencing depth of at least 100-fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000- fold and higher.

84. The non-transitory computer-readable storage medium of any of claims 81-83, wherein the method results in an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

85. The non-transitory computer-readable storage medium of any of claims 81-84, wherein the algorithm comprises a negative binomial, Poisson model or log normal Poisson model for one or more of background selection, mutation rate identification, and methylation status quantification.

86. The non-transitory computer-readable storage medium of any of claims 81-85, wherein the algorithm utilizes local clustering, sub-clonality, copy-number information, structural variation, transcription-coupled repair, replication time information, mutation spectra and processes, or any combination thereof.

87. The non-transitory computer-readable storage medium of any of claims 81-86, wherein the algorithm is implemented using high-performance computing or a cloud system.

88. The non-transitory computer-readable storage medium of any of claims 81-87, wherein the method further comprises annotating the set of mutations with amino acid changes, mutational signature profiles, genomic elements positions and effects, or any combination thereof.

89. The non-transitory computer-readable storage medium of any of claims 81-88, wherein the method further comprises identifying genomic regions under positive negative or neutral selection based on the annotation of the set of mutations.

90. The non-transitory computer-readable storage medium of any of claims 81-89, wherein the method further comprises calculating dN / dS rates, mutation effect sizes and statistical significance, copy-number variations (CNVs), or any combination thereof.

91. The non-transitory computer-readable storage medium of any of claims 81-90, wherein the method further comprises at least one of the following: a) determining single cell mutations; b) calculating statistical depletion or enrichment of a specific mutation type; c) identifying genomic regions of negative selection from mutation due to fitness constrains of the growing cancer or non-cancer cells; d) identifying genomic regions of positive selection from cancer drivers, growth drivers, and / or resistant events; e) summarizing information in a report; f) identifying cancer vulnerabilities, disease vulnerabilities and druggable targets in the genome; andg) profiling of selection before or after treatment from in vivo tissue, ex vivo tissue, cell lines, blood, nucleic acids, circulating nucleic acids, circulating cells, patient derived material, patient derived models, animal models, or organoids.

92. The non-transitory computer-readable storage medium of any of claims 81-91, wherein the method further comprises at least one of the following: a) estimating the copy-number profile of cancer based on selected reads; b) identifying highly mutable regions with a high likelihood of selection signals; c) merging signals of methylation, fragmentation, copy-number, and mutation for accurate estimation of selection signals; d) selecting regions with clinical significance and / or mutations that are oncogenic or resistant; e) profiling using panels based on gene, gene region, non-coding region, and / or whole exome / genome screens; f) estimating cancer cell fraction and frequency of identified mutations to discover selection signals and changes thereof over time; and g) estimating selection signals for normal tissues, jointly or independently for mixtures of cells and tissue types.

93. An electronic device, comprising: a display; one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any of claims 69-80.

94. The electronic device of claim 93, wherein the set of sequencing data is obtained from a sequencing platform selected from the group consisting of Illumina, Oxford NanoPore, Pacific Biosciences, Element Biosciences, Quantapore, Singular Genomics, Ultima Genomics, and Stratos.

95. The electronic device of claim 93 or 94, wherein the set of sequencing data is obtained from a sequencing protocol that results in a sequencing depth of at least 100- fold coverage or higher, at least 500-fold coverage or higher, at least 1000-fold coverage or higher, 100,000-fold and higher, or 1,000,000-fold and higher.

96. The electronic device of any of claims 93-95, wherein the method results in an error rate of at least 10A-6 or lower, at least 10A-7 or lower, at least 10A-8 or lower, or at least 10A-9 or lower.

97. The electronic device of any of claims 93-96, wherein the algorithm comprises a negative binomial, Poisson model or log normal Poisson model for one or more of background selection, mutation rate identification, and methylation status quantification.

98. The electronic device of any of claims 93-97, wherein the algorithm utilizes local clustering, sub-clonality, copy-number information, structural variation, transcription- coupled repair, replication time information, mutation spectra and processes, or any combination thereof.

99. The electronic device of any of claims 93-98, wherein the algorithm is implemented using high-performance computing or a cloud system.

100. The electronic device of any of claims 93-99, wherein the method further comprises annotating the set of mutations with amino acid changes, mutational signature profiles, genomic elements positions and effects, or any combination thereof.

101. The electronic device of any of claims 93-100, wherein the method further comprises identifying genomic regions under positive negative or neutral selection based on the annotation of the set of mutations.

102. The electronic device of any of claims 93-101, wherein the method further comprises calculating dN / dS rates, mutation effect sizes and statistical significance, copy-number variations (CNVs), or any combination thereof.

103. The electronic device of any of claims 93-102, wherein the method further comprises at least one of the following: a) determining single cell mutations; b) calculating statistical depletion or enrichment of a specific mutation type; c) identifying genomic regions of negative selection from mutation due to fitness constrains of the growing cancer or non-cancer cells; d) identifying genomic regions of positive selection from cancer drivers, growth drivers, and / or resistant events; e) summarizing information in a report; f) identifying cancer vulnerabilities, disease vulnerabilities and druggable targets in the genome; and g) profiling of selection before or after treatment from in vivo tissue, ex vivo tissue, cell lines, blood, nucleic acids, circulating nucleic acids, circulating cells, patient derived material, patient derived models, animal models, or organoids.

104. The electronic device of any of claims 93-103, wherein the method further comprises at least one of the following:a) estimating the copy-number profile of cancer based on selected reads; b) identifying highly mutable regions with a high likelihood of selection signals; c) merging signals of methylation, fragmentation, copy-number, and mutation for accurate estimation of selection signals; d) selecting regions with clinical significance and / or mutations that are oncogenic or resistant; e) profiling using panels based on gene, gene region, non-coding region, and / or whole exome / genome screens; f) estimating cancer cell fraction and frequency of identified mutations to discover selection signals and changes thereof over time; and g) estimating selection signals for normal tissues, jointly or independently for mixtures of cells and tissue types.