GRNA sequencing compositions and methods

The described sequencing adapter composition addresses the challenge of detecting and quantifying sequence variants in synthetic oligonucleotides by constructing NGS libraries with specific adapters, correcting errors, and normalizing bias, ensuring accurate sequence analysis for safe gene editing.

WO2025255512A1PCT designated stage Publication Date: 2025-12-11JUNO THERAPEUTICS INC
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/032717
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-07
Filing Date
2025-06-06
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing methods for analyzing synthetic oligonucleotides, such as HPLC and mass spectrometry, often fail to detect sequence variants like single base substitutions, posing a risk for off-target gene editing due to impurities in guide RNA (gRNA) samples, and there is a need for improved methods to quantify sequence variants and purity.

Method used

A sequencing adapter composition is used to construct NGS libraries from single-stranded nucleic acids, comprising 3' and 5' adapter sequences with specific configurations for ligation, allowing for reverse transcription and index primer binding, followed by paired-end sequencing to correct amplification and sequencing errors, and normalize sequence-based amplification bias.

Benefits of technology

Enables accurate detection and quantification of intended and variant sequences, correcting errors and normalizing bias, thus ensuring the safety and efficacy of gene editing processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025032717_11122025_PF_FP_ABST
    Figure US2025032717_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Sequencing adapter compositions and methods for use thereof to construct Next-Generation Sequencing (NGS) libraries from single-stranded nucleic acid molecules are described. Also described are methods for detecting and quantifying unique sequences present in a population of single-stranded template nucleic acid molecules.
Need to check novelty before this filing date? Find Prior Art

Description

GRNA SEQUENCING COMPOSITIONS AND METHODSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 657,731, filed June 7, 2024, entitled “GRNA SEQUENCING COMPOSITIONS AND METHODS,” which is herein incorporated by reference in its entirety for all purposes.REFERENCE TO AN ELECTRONIC SEQUENCE LISTING

[0002] The contents of the electronic sequence listing (217372000440SEQLIST.xml; Size: 151,724 bytes; and Date of Creation: May 29, 2025) is herein incorporated by reference in its entirety.FIELD

[0003] The present disclosure relates generally to the field of nucleic acid sequencing. More specifically, compositions and methods are described that are useful in generating Next- Generation Sequencing (NGS) libraries from single-stranded nucleic acid molecules (e.g., synthetic guide RNA (gRNA) samples) and providing quantitative assessments of sample purity and the identity and frequency of sequence variants present in the sample.BACKGROUND

[0004] In many research and clinical fields related to molecular biology, gene editing, and cellular therapy, it is desirable to obtain quantitative nucleic acid sequence information for synthetic oligonucleotides used as critical raw materials.

[0005] Custom oligonucleotide products are typically synthesized by sequential reaction cycles where each cycle incorporates one additional nucleotide onto each individual oligonucleotide molecule in parallel to generate a very large number of oligonucleotides with the same intended sequence. Due to imperfect fidelity of sequential nucleotide incorporation, some oligonucleotide molecules within the population will contain sequence errors, resulting in the presence of sequence variants that can be considered impurities of the intended oligonucleotide product. The frequency of such unintended oligonucleotide sequence variants is expected to vary depending on the synthesis process, production scale, and purification method used during manufacturing and is expected to increase with the length of the oligonucleotide.

[0006] The ability to assess the frequency of correct oligonucleotide sequences (z.e., the intended sequences) as well as to identify and quantify unintended variant sequences within an oligonucleotide product would be highly useful for many applications of custom synthetic oligonucleotides (e.g., cell therapy and gene therapy research, and clinical applications). For example, sequence purity measurements would be beneficial for applications that use guide RNA (gRNA) oligonucleotides as critical raw materials for genome editing (for instance, CRISPR-Cas targeted gene editing using Cas9 or Casl2a). Cas9 is a dual RNA-guided DNA endonuclease enzyme associated with the Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) adaptive immune system in .S'. pyogenes which evolved to interrogate and cleave foreign DNA. Cas9 unwinds double- stranded template DNA and identifies “protospacer adjacent motif’ (PAM) sequences - short, conserved sequences of 2 to 5 bp which are used by the enzyme for target recognition. If the adjacent template DNA sequence is complementary to a spacer region of the gRNA molecule, the RuvC nuclease domain of Cas9 cleaves the non-target strand of the template DNA and the HNH nuclease domain cleaves the target strand. Casl2a forms a complex with a synthetic CRISPR RNA (crRNA) and a double- stranded template DNA molecule. Here the crRNA serves as the guide RNA, the RuvC nuclease domain cleaves the non-target strand of the template DNA and the Nuc nuclease domain cleaves the target strand.

[0007] Because the genomic location of gene editing is directed by the gRNA sequence, the presence of gRNA impurities having unintended sequence(s) within a gRNA sample presents a potential concern for the safety and efficacy of the gene editing process. One such concern is that gRNA materials with impurities having unintended sequence(s) could result in off-target gene editing. Unfortunately, commonly used analytical methods for characterizing synthetic oligonucleotides, such as HPLC and mass spectrometry, often fail to detect sequence variants such as single base substitutions. Accordingly, there is a need for improved methods that provide quantitative assessments of oligonucleotide sample purity, and the identity and frequency of sequence variants within the oligonucleotide sample.BRIEF SUMMARY OF THE INVENTION

[0008] Provided herein are compositions and methods for constructing Next-Generation Sequencing (NGS) libraries from single- stranded nucleic acid molecules (e.g., synthetic guide RNA (gRNA) samples), and for detecting and quantifying unique sequences (e.g., intended sequences and / or variant sequences) present in a population of single-stranded template nucleicacid molecules. In some embodiments, the compositions and methods disclosed herein can be used to correct amplification errors. In some embodiments, the compositions and methods disclosed herein can be used to correct sequencing errors. In some embodiments, the compositions and methods disclosed herein can be used to identify or detect presence of a known or unknown sequence (e.g., a contaminant sequence). In some embodiments, the compositions and methods disclosed herein can be used to normalize for sequence-based amplification bias. In some embodiments, the disclosed compositions and methods enable accurate detection and quantification of nucleic acid molecules comprising the intended sequence in a synthetic oligonucleotide sample (e.g., the frequency of the intended sequence, the relative frequency of the intended sequence (i.e., the fraction of the total number of molecules comprising the intended sequence), or the percentage of molecules in the sample that comprise the intended sequence). In some embodiments, the disclosed compositions and methods also enable accurate detection and quantification of nucleic acid molecules comprising known or unknown variant sequence(s) in a synthetic oligonucleotide sample (e.g., the frequency of a given variant sequence, the relative frequency of a given variant sequence (i.e., the fraction of the total number of molecules comprising the given variant sequence), or the percentage of molecules in the sample that comprise a given variant sequence).

[0009] In some embodiments, provided herein is a sequencing adapter composition comprising: a) a 3’ adapter sequence configured for ligation to a 3’ end of a single-stranded oligonucleotide, the 3’ adapter sequence comprising, in a 5’ to 3’ direction: a ligatable 5’ end, a first linker sequence, a first index primer binding site sequence, and a non-ligatable 3’ end, wherein the 3’ adapter sequence is capable of hybridizing to a reverse transcription primer; and b) a 5’ adapter sequence configured for ligation to a 5’ end of the single-stranded oligonucleotide, the 5’ adapter sequence comprising, in a 5’ to 3’ direction: a non-ligatable 5’ end, a second index primer binding site sequence, a randomized sequence, a second linker sequence, and a ligatable 3’ end.

[0010] In some embodiments, the 3’ adapter sequence does not comprise a stem-loop structure sequence. In any of the embodiments herein, less than 5%, 10%, or 20% of the 3’ adapter sequence is capable of forming a secondary structure. In any of the embodiments herein, the 3’ adapter sequence may but does not need to be capable of forming a secondary structure.

[0011] In any of the embodiments herein, the ligatable 5’ end of the 3’ adapter sequence may comprise a 5 ’-adenylated nucleotide. In any of the embodiments herein, the ligatable 5’ end of the 3’ adapter sequence may comprise a 5 ’-phosphorylated nucleotide.

[0012] In any of the embodiments herein, the ligatable 3’ end of the 5’ adapter sequence may comprise a 3 ’-hydroxylated nucleotide. In any of the embodiments herein, the ligatable 5’ end of the 3’ adapter sequence may comprise a 5 ’-hydroxylated nucleotide. In any of the embodiments herein, the ligatable 3’ end of the 5’ adapter sequence may comprise a 3 ’-phosphorylated nucleotide.

[0013] In any of the embodiments herein, the non-ligatable 3’ end of the 3’ adapter sequence may comprise a 3’ modification of a 3’ nucleotide. In some embodiments, the 3’ modification is selected from the group consisting of: a 3’ amino group, a 3’ dideoxy-C modification, a 3’ inverted dT modification, a 3’ C3 spacer modification, or 3’ phosphoryl modification.

[0014] In any of the embodiments herein, the non-ligatable 5’ end of the 5’ adapter sequence may comprise a 5’ modification of a 5’ nucleotide. In some embodiments, the 5’ modification is selected from the group consisting of: a 5’ hydroxyl group, a 5’ dideoxy-C modification, 5’ inverted dideoxy-T modification, or 5’ C3 spacer modification.

[0015] In any of the embodiments herein, the 3’ adapter sequence may be at least about 80%, at least about 85%, at least about 90%, or at least about 95% identical to a sequence that is complementary to the reverse transcription primer. In any of the embodiments herein, the 3’ adapter sequence may be 100% identical to a sequence that is complementary to the reverse transcription primer.

[0016] In any of the embodiments herein, the first linker sequence may have a length of about 4 nucleotides to about 20 nucleotides. In any of the embodiments herein, the first linker sequence may have a length of 10 nucleotides.

[0017] In any of the embodiments herein, the first index primer binding site sequence may have a length of about 20 nucleotides to about 35 nucleotides. In any of the embodiments herein, the first index primer binding site sequence may have a length of about 28 nucleotides.

[0018] In any of the embodiments herein, the first index primer binding site sequence may be capable of binding to a sequencing read 2 adapter.

[0019] In any of the embodiments herein, the second index primer binding site sequence may have a length of about 20 nucleotides to about 30 nucleotides. In any of the embodiments herein, the second index primer binding site sequence may have a length of about 22 nucleotides to about 28 nucleotides. In any of the embodiments herein, the second index primer binding site sequence may have a length of about 26 nucleotides.

[0020] In any of the embodiments herein, the second index primer binding site sequence may be capable of binding to a sequencing read 1 adapter.

[0021] In any of the embodiments herein, the randomized sequence may have a length of about 10 nucleotides to about 20 nucleotides. In any of the embodiments herein, the randomized sequence may have a length of about 15 nucleotides.

[0022] In any of the embodiments herein, the second linker sequence may have a length of about 4 nucleotides to about 20 nucleotides. In any of the embodiments herein, the second linker sequence may have a length of about 7 nucleotides.

[0023] In any of the embodiments herein, the 3’ adapter sequence may have a total length of about 20 nucleotides to about 60 nucleotides. In any of the embodiments herein, the 3’ adapter sequence may have a total length of about 30 nucleotides to about 45 nucleotides. In any of the embodiments herein, the 3’ adapter sequence may have a total length of about 38 nucleotides.

[0024] In any of the embodiments herein, the 5’ adapter sequence may have a total length of about 40 nucleotides to about 60 nucleotides. In any of the embodiments herein, the 5’ adapter sequence may have a total length of about 42 nucleotides to about 55 nucleotides. In any of the embodiments herein, the 5’ adapter sequence may have a total length of about 48 nucleotides.

[0025] In any of the embodiments herein, the 3’ adapter sequence may comprise a DNA sequence.

[0026] In any of the embodiments herein, the 5’ adapter sequence may comprise an RNA sequence or a DNA sequence.

[0027] In any of the embodiments herein, the 3’ adapter sequence and the 5’ adapter sequence may but do not need to comprise substantial complementarity with each other.

[0028] In some embodiments, provided herein is a method of preparing a sequencing library for a single- stranded template nucleic acid molecule, comprising: a) contacting the single-strandednucleic acid molecule with 3’ adapter sequence disclosed herein and performing a ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule; b) performing an end modification reaction to modify a 5’ end of the first ligated nucleic acid molecule; c) contacting the first ligated nucleic acid molecule with the 5’ adapter sequence disclosed herein and performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule; d) contacting the second ligated nucleic acid molecule with a reverse transcription primer and allowing the reverse transcription primer to hybridize to the 3’ adapter sequence; e) performing a reverse transcription reaction, thereby producing a cDNA molecule comprising the template nucleic acid molecule; f) contacting the cDNA molecule with a first index primer capable of hybridizing to the first index primer binding site sequence, and with a second index primer capable of hybridizing to the second index primer binding site sequence, under conditions that allow hybridization of the first and second index primers to their respective binding site sequences; g) amplifying the cDNA molecule to produce an amplification product; and h) purifying the amplification product to produce the sequencing library for the single-stranded template nucleic acid molecule.

[0029] In some embodiments, the end modification reaction comprises a phosphorylation reaction to phosphorylate the 5’ end of the first ligated nucleic acid molecule. In some embodiments, the end modification reaction comprises an adenylation reaction to adenylate the 5’ end of the first ligated nucleic acid molecule.

[0030] In any of the embodiments herein, ligation of the 3’ adapter sequence to a 3’ end of the single- stranded template nucleic acid molecule may comprise the use of a DNA ligase or an RNA ligase. In any of the embodiments herein, ligation of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule may comprise the use of a DNA ligase or an RNA ligase. In any of the embodiments herein, the ligation may comprise the use of RtcB ligase.

[0031] In any of the embodiments herein, the method may but does not need to comprise a step of cleaving the first ligated nucleic acid molecule or the second ligated nucleic acid molecule.

[0032] In any of the embodiments herein, the single- stranded template nucleic acid molecule may be an RNA molecule. In any of the embodiments herein, the single- stranded template nucleic acid molecule may be a guide RNA (gRNA) molecule. In any of the embodiments herein, the singlestranded template nucleic acid molecule may be a small interfering RNA (siRNA) molecule.

[0033] In any of the embodiments herein, the single- stranded template nucleic acid molecule may be a DNA molecule.

[0034] In some embodiments, provided herein is a method for detecting and quantifying unique sequences in a population of single-stranded template nucleic acid molecules, comprising: a) preparing a sequencing library for the single- stranded template nucleic acid molecule using the method disclosed herein; b) sequencing the sequencing library using a paired-end sequencing method to generate a plurality of sequence reads corresponding to the single- stranded template nucleic acid molecule; and c) analyzing sequence read data corresponding to the plurality of sequence reads to achieve at least one of: (i) correction of amplification errors; (ii) correction of sequencing errors; (iii) normalization of sequence-based amplification bias; and (iv) detection and quantification of unique sequences present in the population of single-stranded template nucleic acid molecules.

[0035] In some embodiments, the population of single- stranded template nucleic acid molecules is a putatively homogenous population.

[0036] In any of the embodiments herein, analyzing sequence read data may comprise the use of one or more processors.

[0037] In any of the embodiments herein, analyzing the sequence read data may comprise aligning paired sequence reads and discarding discrepant read pairs from further analysis.

[0038] In any of the embodiments herein, analyzing the sequence read data may comprise grouping sequence reads according to unique molecular identifiers (UMIs) associated with each sequence read. In some embodiments, a unique molecular identifier (UMI) associated with each sequence read corresponds to the randomized sequence in the 5’ adapter sequence disclosed herein.

[0039] In any of the embodiments herein, the method may further comprise determining a consensus sequence for each group of sequence reads identified based on the associated unique molecular identifier (UMI).

[0040] In some embodiments, the method further comprises identifying unique sequences present in the population of single- stranded template nucleic acid molecules based on the determined consensus sequence for each group of sequence reads.

[0041] In any of the embodiments herein, the method may further comprise correcting for amplification and / or sequencing errors based on the determined consensus sequence for each group of sequence reads.

[0042] In any of the embodiments herein, the method may further comprise aligning a consensus sequence for each group of sequence reads to a reference sequence to identify variant sequences present in the population of single-stranded template nucleic acid molecules.

[0043] In some embodiments, the method further comprises determining a frequency of sequences present in the population of single-stranded template nucleic acid molecules that are identical to the reference sequence, and / or determining a variant sequence frequency for a variant sequence identified in the population of single-stranded template nucleic acid molecules.

[0044] In any of the embodiments herein, the population of single-stranded template nucleic acid molecules may comprise guide RNA (gRNA) molecules. In any of the embodiments herein, the population of single-stranded template nucleic acid molecules may comprise small interfering RNA (siRNA) molecules. In any of the embodiments herein, the population of single-stranded template nucleic acid molecules comprise DNA molecules.

[0045] In some embodiments, provided herein is a kit comprising: a) the sequencing adapter composition disclosed herein, and b) at least one of: a ligase; a buffer; enzymes and / or reagents for performing an end modification reaction to modify a 5’ end of a nucleic acid molecule; enzymes and / or reagents for performing nucleic acid amplification; enzymes and / or reagents for performing paired-end sequencing.

[0046] In some embodiments, provided herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive sequence read data for a sequencing library generated using the method disclosed herein; and analyze the sequence read data for the sequencing library to achieve at least one of: (i) correction of amplification errors; (ii) correction of sequencing errors; (iii) normalization of sequence-based amplification bias; and (iv) detection and quantification of unique sequences present in a population of single-stranded template nucleic acid molecules.

[0047] In some embodiments, analyzing the sequence read data comprises aligning paired sequence reads and discarding discrepant read pairs from further analysis.

[0048] In any of the embodiments herein, analyzing the sequence read data may comprise grouping sequence reads according to unique molecular identifiers (UMIs) associated with each sequence read. In some embodiments, a unique molecular identifier (UMI) associated with each sequence read corresponds to a randomized sequence in the 5’ adapter sequence disclosed herein. In any of the embodiments herein, the system may further comprise determining a consensus sequence for each group of sequence reads identified based on the associated unique molecular identifier (UMI).

[0049] In some embodiments, the system further comprises identifying unique sequences present in the population of single- stranded template nucleic acid molecules based on the determined consensus sequence for each group of sequence reads.

[0050] In any of the embodiments herein, the system may further comprise correcting for amplification and / or sequencing errors based on the determined consensus sequence for each group of sequence reads.

[0051] In some embodiments, the system further comprises aligning a consensus sequence for each group of sequence reads to a reference sequence to identify variant sequences present in the population of single-stranded template nucleic acid molecules. In some embodiments, the system further comprises determining a frequency of sequences present in the population of singlestranded template nucleic acid molecules that are identical to the reference sequence, and / or determining a variant sequence frequency for a variant sequence identified in the population of single- stranded template nucleic acid molecules.

[0052] In any of the embodiments herein, the single-stranded template nucleic acid molecules in the population may comprise guide RNA (gRNA) molecules. In any of the embodiments herein, the single-stranded template nucleic acid molecules in the population may comprise small interfering RNA (siRNA) molecules. In any of the embodiments herein, the single-stranded template nucleic acid molecules in the population may comprise DNA molecules.

[0053] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.INCORPORATION BY REFERENCE

[0054] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, of which:

[0056] FIG. 1 provides a non-limiting schematic illustration of a 3’ adapter sequence configured to be ligated to the 3’ end of a single- stranded template nucleic acid molecule, according to some embodiments of the compositions described herein.

[0057] FIG. 2 provides a non-limiting schematic illustration of a 5’ adapter sequence configured to be ligated to the 5’ end of a single- stranded template nucleic acid molecule, according to some embodiments of the compositions described herein.

[0058] FIG. 3 provides a non-limiting schematic illustration of a workflow for preparing a sequencing library from a single- stranded nucleic acid molecule (e.g., a synthetic gRNA molecule), according to some embodiments of the methods described herein.

[0059] FIG. 4 provides a non-limiting schematic illustration of a workflow for processing sequence read data to correct for amplification- and sequencing-based sequence errors, correct for sequence-dependent amplification bias, and identify and quantify the presence of variant sequences in a nucleic acid sample (e.g., a synthetic oligonucleotide sample), according to some embodiments of the methods described herein.

[0060] FIG. 5 provides a non-limiting example of an electropherogram trace for a purified NGS amplicon library prepared according to one embodiment of the methods described herein.

[0061] FIG. 6 provides a non-limiting example of data for the log2 fold-change in gRNA sequence frequency between read-based quantification and UMI-based quantification (y-axis) plotted as a function of UMI-based sequence frequency (x-axis) wherein each point represents an individual observed gRNA sequence.

[0062] FIG. 7 provides another non-limiting example of an electropherogram trace for a purified NGS amplicon library prepared according to one embodiment of the methods described herein.

[0063] FIG. 8 provides another non-limiting schematic illustration of a 3’ adapter sequence configured to be ligated to the 3’ end of a single-stranded template nucleic acid molecule, according to some embodiments of the compositions described herein.

[0064] FIG. 9 provides another non-limiting example of an electropherogram trace for a purified NGS amplicon library prepared according to one embodiment of the methods described herein.DETAILED DESCRIPTION

[0065] Compositions and methods for constructing Next-Generation Sequencing (NGS) libraries from single- stranded nucleic acid molecules (e.g., synthetic guide RNA (gRNA) samples), and for detecting and quantifying unique sequences (e.g., intended sequences and / or known or unknown variant sequences) present in a population of single- stranded template nucleic acid molecules are described. The disclosed compositions and methods allow analytical correction of amplification errors, correction of sequencing errors, and / or normalization of sequence read data to correct for sequence-based amplification bias. The disclosed compositions and methods enable accurate detection and quantification of nucleic acid molecules comprising the intended sequence, variant sequence(s), and / or undesired (e.g. contaminant) sequences in a nucleic acid sample, e.g., a synthetic oligonucleotide sample.

[0066] In one aspect, provided herein is a sequencing adapter composition comprising: a) a 3’ adapter sequence configured for ligation to a 3’ end of a single-stranded oligonucleotide, the 3’ adapter sequence comprising, in a 5’ to 3’ direction: a ligatable 5’ end, a first linker sequence, a first index primer binding site sequence, and a non-ligatable 3’ end, wherein the 3’ adapter sequence is capable of hybridizing to a reverse transcription primer; and b) a 5’ adapter sequence configured for ligation to a 5’ end of the single- stranded oligonucleotide, the 5’ adapter sequence comprising, in a 5’ to 3’ direction: a non-ligatable 5’ end, a second index primer binding site sequence, a randomized sequence, a second linker sequence, and a ligatable 3’ end.

[0067] In another aspect, provided herein is a sequencing adapter composition comprising: a) a 3’ adapter sequence configured for ligation to a 3’ end of a single-stranded oligonucleotide, the 3’ adapter sequence comprising, in a 5’ to 3’ direction: a ligatable 5’ end, a first linker sequence, a first index primer binding site sequence, and a non-ligatable 3’ end, wherein at least a portion of the 3’ adapter sequence is capable of hybridizing to a reverse transcription primer; and b) a 5’ adapter sequence configured for ligation to a 5’ end of the single-stranded oligonucleotide, the 5’ adapter sequence comprising, in a 5’ to 3’ direction: a non-ligatable 5’ end, a second index primer binding site sequence, and a ligatable 3’ end; wherein at least one of the 3’ adapter sequence and the 5’ adapter sequence comprises a randomized sequence.

[0068] In another aspect, provided herein is a method of preparing a sequencing library for a single- stranded template nucleic acid molecule, comprising: a) contacting the single-stranded nucleic acid molecule with any of the 3’ adapter sequences disclosed herein and performing a ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule; b) performing an end modification reaction to modify a 5’ end of the first ligated nucleic acid molecule; c) contacting the first ligated nucleic acid molecule with any of the 5’ adapter sequences disclosed herein and performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule; d) contacting the second ligated nucleic acid molecule with a reverse transcription primer and allowing the reverse transcription primer to hybridize to the 3’ adapter sequence; e) performing a reverse transcription reaction, thereby producing a cDNA molecule comprising the template nucleic acid molecule; f) contacting the cDNA molecule with a first index primer capable of hybridizing to the first index primer binding site sequence, and with a second index primer capable of hybridizing to the second index primer binding site sequence, under conditions that allow hybridization of the first and second index primers to their respective binding site sequences; g) amplifying the cDNA molecule to produce an amplification product; and h) purifying the amplification product to produce the sequencing library for the single-stranded template nucleic acid molecule.

[0069] In another aspect, provided herein is a method for detecting and quantifying variant sequences in a population of single-stranded template nucleic acid molecules, comprising: a) preparing a sequencing library for the single- stranded template nucleic acid molecule using any of the methods disclosed herein; b) sequencing the sequencing library using a paired-end sequencing method to generate a plurality of sequence reads corresponding to the single- stranded templatenucleic acid molecule; c) analyzing sequence read data corresponding to the plurality of sequence reads to achieve one or more of: (i) correction of amplification errors; (ii) correction of sequencing errors; (iii) normalization of sequence-based amplification bias; and (iv) detection and quantification of sequence variants present in the population of single-stranded template nucleic acid molecules.

[0070] In another aspect, provided herein is a kit comprising: a) any of the sequencing adapter compositions disclosed herein, and one or more of: a ligase; a buffer; enzymes and / or reagents for performing an end modification reaction to modify a 5’ end of a nucleic acid molecule; enzymes and / or reagents for performing nucleic acid amplification; enzymes and / or reagents for performing paired-end sequencing.

[0071] In yet another aspect, provided herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive sequence read data for a sequencing library generated using any of the methods disclosed herein; and analyze the sequence read data for the sequencing library to achieve one or more of: (i) correction of amplification errors; (ii) correction of sequencing errors; (iii) correction of sequencebased amplification bias; and (iv) detection and quantification of sequence variants present in the population of single-stranded template nucleic acid molecules.

[0072] The following description is presented to enable a person of ordinary skill in the art to make and use the various embodiments. Descriptions of specific devices, techniques, and applications are provided only as examples. Various modifications to the examples described herein will be readily apparent to those of ordinary skill in the art, and the general principles defined herein may be applied to other examples and applications without departing from the spirit and scope of the various embodiments. Thus, the various embodiments are not intended to be limited to the examples described herein and shown but are to be accorded the scope consistent with the claims.Definitions

[0073] Unless otherwise defined, all of the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field to which this disclosure belongs.

[0074] As used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural references unless the context clearly indicates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated, and encompasses any and all possible combinations of one or more of the associated listed items.

[0075] As used herein, the terms “includes, “including,” “comprises,” and / or “comprising” specify the presence of stated features, integers, steps, operations, elements, components, and / or units but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, units, and / or groups thereof.

[0076] Throughout this application, various parameter values may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity, and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all possible subranges as well as individual numerical values within that range, irrespective of whether a specific numerical value or specific sub-range is expressly stated. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as individual numbers within that range, for example, 1, 1.4, 2, 3, 3.6, 4, 5, 5.8, and 6. This applies regardless of the breadth of the range.

[0077] Numbers may be expressed herein as being “about” a particular value. Similarly, ranges may be expressed herein as from “about” one particular value and / or to “about” another particular value. The terms “about” and “approximately” shall generally mean an acceptable degree of error or variation for a given value or range of values, such as, for example, a degree of error or variation that is within 20 percent (%), within 15%, within 10%, or within 5% of a given value or range of values.

[0078] As used herein, a “nucleotide” comprises a “base” (alternatively, a “nucleobase” or “nitrogenous base”), a “sugar” (in particular, a five-carbon sugar, e.g., ribose or 2- deoxyribose), and a “phosphate moiety” of one or more phosphate groups e.g., a monophosphate, a diphosphate, or a triphosphate consisting of one, two, or three linked phosphates, respectively). Without the phosphate moiety, the nucleobase and the sugar compose a “nucleoside”. A nucleotide can thus also be called a nucleoside monophosphate or a nucleoside diphosphate or a nucleoside triphosphate, depending on the number of phosphate groups attached. The phosphate moiety isusually attached to the 5’ carbon of the sugar, though some nucleotides comprise phosphate moieties attached to the 2-carbon or the 3-carbon of the sugar. Nucleotides contain either a purine (in the nucleotides adenine and guanine) or a pyrimidine base (in the nucleotides cytosine, thymine, and uracil). Ribonucleotides are nucleotides in which the sugar is ribose. Deoxyribonucleotides are nucleotides in which the sugar is deoxyribose.

[0079] As used herein, a “nucleic acid” shall mean any nucleic acid molecule, including, without limitation, DNA, RNA, hybrids, and chimeras thereof. The nucleic acid bases that form nucleic acid molecules can be the bases A, C, G, T and U, as well as derivatives thereof. Derivatives of these bases are well known in the art. The term should be understood to include, as equivalents, analogs of either DNA or RNA made from nucleotide analogs. The term as used herein also encompasses cDNA, that is complementary, or copy, DNA produced from an RNA template, for example by the action of a reverse transcriptase.

[0080] As used herein, “complementary” generally refers to specific nucleotide duplexing to form canonical Watson-Crick base pairs, as is understood by those skilled in the art. However, complementary also includes base-pairing of nucleotide analogs that are capable of universal basepairing with A, T, G or C nucleotides and locked nucleic acids that enhance the thermal stability of duplexes. One skilled in the art will recognize that hybridization stringency is a determinant in the degree of match or mismatch in the duplex formed by hybridization.

[0081] A “polymerase” is an enzyme generally for joining 3'-OH 5'-triphosphate nucleotides, oligomers, and their analogs. Polymerases include, but are not limited to, DNA-dependent DNA polymerases, DNA-dependent RNA polymerases, RNA-dependent DNA polymerases, RNA- dependent RNA polymerases, T7 DNA polymerase, T3 DNA polymerase, T4 DNA polymerase, T7 RNA polymerase, T3 RNA polymerase, SP6 RNA polymerase, DNA polymerase 1, Klenow fragment, Thermophilus aquaticus (Taq) DNA polymerase, Thermus thermophilus (Tth) DNA polymerase, Vent DNA polymerase (New England Biolabs), Deep Vent DNA polymerase (New England Biolabs), Bacillus stearothermophilus (Bst) DNA polymerase, DNA Polymerase Large Fragment, Stoeffel Fragment, 9°N DNA Polymerase, 9°Nmpolymerase, Pyrococcus furiosis (Pfu) DNA Polymerase, Thermus filiformis DNA Polymerase, RepliPHI Phi29 Polymerase, Thermococcus litoralis (Tli) DNA polymerase, eukaryotic DNA polymerase beta, telomerase, Therminator (e.g., THERMINATOR I, THERMINATOR II, etc.) polymerase (New England Biolabs), KOD HiFi. DNA polymerase (Novagen), KOD1 DNA polymerase, Q-beta replicase,terminal transferase, AMV reverse transcriptase, M-MLV reverse transcriptase, Phi6 reverse transcriptase, HIV-1 reverse transcriptase, novel polymerases discovered by bioprospecting and / or molecular evolution, and polymerases cited in U.S. Pat. Appl. Pub. No. 2007 / 0048748 and in U.S. Pat. Nos. 6,329,178; 6,602,695; and 6,395,524. These polymerases include wild-type, mutant isoforms, and genetically engineered variants such as exopolymerases; polymerases with minimized, undetectable, and / or decreased 3' to 5' proofreading exonuclease activity, and other mutants, e.g., that tolerate labeled nucleotides and incorporate them into a strand of nucleic acid. In some embodiments, the polymerase is designed for use, e.g., in real-time PCR, high fidelity PCR, next-generation DNA sequencing, fast PCR, hot start PCR, crude sample PCR, robust PCR, and / or molecular diagnostics. Such enzymes are available from many commercial suppliers, e.g., Kapa Enzymes, Finnzymes, Promega, Invitrogen, Life Technologies, Thermo Scientific, Qiagen, Roche, etc.

[0082] As used herein, the term “primer” refers to an oligonucleotide, whether occurring naturally as in a purified restriction digest or produced synthetically, that is capable of acting as a point of initiation of synthesis when placed under conditions in which synthesis of a primer extension product that is complementary to a nucleic acid strand is induced, (e.g., in the presence of nucleotides and an inducing agent such as DNA polymerase and at a suitable temperature and pH). The primer is preferably single- stranded for maximum efficiency in amplification, but may alternatively be double-stranded. If double-stranded, the primer is first treated to separate its strands before being used to prepare extension products. The primer must be sufficiently long to prime the synthesis of extension products in the presence of the inducing agent.

[0083] As used herein, “index” generally refers to a distinctive or identifying mark or characteristic.

[0084] As used herein, “nucleic acid sequencing data”, “nucleic acid sequencing information”, “nucleic acid sequence”, or “nucleic acid sequencing read” denotes any information or data that is indicative of the order of the nucleotide bases (e.g. , adenine, guanine, cytosine, and thymine / uracil) in a molecule (e.g., an oligonucleotide, polynucleotide, etc.) of DNA or RNA. It should be understood that the present teachings contemplate sequence information obtained using all available varieties of techniques, platforms or technologies, including, but not limited to: capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH- based detection systems, electronic signature-based systems, etc.

[0085] As used herein, a “polynucleotide”, “nucleic acid”, or “oligonucleotide” refers to a linear a biopolymer composed of nucleotide monomers that are covalently bonded in a chain.

[0086] As used herein, an “oligonucleotide substrate” or a “substrate molecule” refers to any nucleotide sequence (e.g. , RNA or DNA), the manipulation of which may be deemed desirable for any reason by one of ordinary skill in the art. In some contexts, an “oligonucleotide substrate” or a “substrate molecule” refers to a nucleotide sequence whose nucleotide sequence is to be determined or is desired to be determined, and / or whose frequency / quantity in a population is to be determined or is desired to be determined.

[0087] As used herein, the term “library” refers to a plurality of nucleic acids, e.g., a plurality of different nucleic acids.

[0088] As used herein, the term “consensus sequence” refers to a sequence that is common to, or otherwise present in the largest fraction, of an aligned group of sequences. The consensus sequence shows the nucleotide most commonly found at each position within the nucleic acid sequences of the group of sequences.

[0089] The phrase “sequencing run” refers to any step or portion of a sequencing experiment performed to determine some information relating to at least one biomolecule (e.g., nucleic acid molecule).

[0090] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.Overview - compositions and methods for preparing NGS libraries from single-stranded nucleic acid molecules

[0091] Next-generation sequencing (NGS) is a powerful tool for generating sequence data for millions of nucleic acid molecules and can be used to assess the rate of sequence variability within a molecular population. Many methods exist for both general and specialized NGS sample preparation and data analysis; however, there is a need for a method specifically designed and optimized to prepare sequencing libraries from oligonucleotide sample substrates (such as gRNA compositions, siRNA compositions, or other single- stranded DNA or RNA oligonucleotideshaving a length of 15-500 nucleotides) and provide quantitative sequence frequency data representative of the original molecular population. The general requirements of a quantitative sequencing method for oligonucleotide samples are that the method must be able to add sequencing adapters to the 5’ and 3’ sides of the single- stranded nucleic acid and convert it to double-stranded DNA without altering the original sequence, must avoid significant molecular bias which could enrich or deplete certain oligonucleotide sequences from the original population due to differential efficiency of biochemical reactions, and must include a mechanism for distinguishing PCR and sequencing errors from true molecular variants in the original population. Current methods known to the art generally do not satisfy these requirements because they lack optimizations to maximize reaction efficiency for single- stranded oligonucleotide substrates, adapters designed to minimize molecular bias related to nucleic acid secondary structure, and / or a mechanism for correcting PCR and sequencing errors to improve accuracy of true variant identification and quantification.

[0092] In most applications of oligonucleotide sequencing, it is desirable to obtain sequence information from the entire oligonucleotide molecule, which is not compatible with the use of a sequence-specific primer that hybridizes to a portion of the substrate molecule to initiate reverse transcription or amplification or to add NGS adapter and index sequences. The use of sequencespecific primers is not suitable for sequencing molecular populations of unknown composition; it inherently biases sequencing libraries towards an expected sequence, and it causes a portion of the original sequence information to be lost due to the primer sequence replacing the original substrate molecule sequence within the primer binding site. Furthermore, in order to provide unambiguous determination of substrate molecule termini, the incorporation of non-templated nucleotides, such as a polyadenosine tail, must be avoided. These factors necessitate the use of adapters that ligate to the 5’ and 3’ ends of substrate molecules to add the necessary NGS indices and / or serve as primer binding sites outside of the original substrate sequence.

[0093] Single-stranded oligonucleotide substrates are highly prone to formation of secondary structure, which is known to substantially impact the efficiency of biochemical reactions necessary for NGS library preparation, particularly the ligation of sequencing adapters to the substrate molecule. The secondary structure of nucleic acid strands is ultimately a function of primary sequence, and differential reaction efficiency can therefore introduce a source of sequencedependent molecular bias to library preparation methods. Some substrate molecules may be successfully converted into sequencing library molecules (i.e., with full adapter and index sequence added) at a higher rate than others due to secondary structures that are more or lessfavorable for reaction completion, thus causing an overrepresentation or underrepresentation of those sequences in the final NGS data. Oligonucleotides used for certain applications may include modified nucleotides, such as those derived by 2’ oxygen methylation of terminal nucleotides in gRNAs to improve stability, which may also affect the efficiency of library preparation reactions. Strategies to mitigate the effects of differential reaction efficiency among substrate molecules in a sample population have been utilized in some methods (Song et al. (2014) Elimination of Ligation Dependent Artifacts in T4 RNA Ligase to Achieve High Efficiency and Low Bias MicroRNA Capture. PLoS ONE 9(4): e94619.; Euchs et al. (2015) Bias in Ligation-Based Small RNA Sequencing Library Construction Is Determined by Adaptor and RNA Structure. PLoS ONE 10(5): e0126049.), but such strategies typically have not been implemented in the context of library preparation from synthetic oligonucleotide substrates for quantitative NGS sequencing. There are three primary sources of error in NGS sequencing of single-stranded RNA and chimeric DNA- RNA oligonucleotide substrates: (i) incorrect copying of the template strand during reverse transcription, (ii) incorrect copying of the template strand during PCR amplification, and (iii) reporting of incorrect basecalls from the sequencer instrument. The estimated error rates for reverse transcription, PCR, and Illumina NGS sequencing are approximately 6xl0'5errors / base, 2xl0'6errors / base, and 6xl0'3errors / base, respectively (Potapov et al. Base modifications affecting RNA polymerase and reverse transcriptase fidelity. Nucleic Acids Res. 2018 Jun 20;46(l l):5753-5763. doi: 10.1093 / nar / gky341. PMID: 29750267; PMCID: PMC6009661.; Staler & Nekrutenko, Sequencing error profiles of Illumina sequencing instruments, NAR Genomics and Bioinformatics, Volume 3, Issue 1, March 2021, lqab019). These sources of error create a relatively high background of sequence variability in NGS read data, which can be difficult to distinguish from genuine variants within the original population of substrate molecules. The presence of a background error rate in NGS data is problematic for the application of performing NGS sequencing of oligonucleotide samples, such as gRNA compositions, to accurately identify and quantify variants. Strategies to correct PCR and sequencing errors and improve the overall accuracy of NGS data are implemented in some existing methods but are not currently applied in methods for sequencing of oligonucleotide substrates and detecting and quantifying variant sequences.

[0094] During sequencing library amplification, each original substrate molecule will generate multiple amplicon copies which will eventually be sequenced and consequently result in multiple sequencing reads originating from the same single substrate molecule. Due to stochasticvariability in PCR amplification and differential amplification efficiency between different template sequences, it is expected that the prevalence of PCR amplicons does not necessarily remain proportional to the prevalence of template molecules in the original population. In many instances of amplicon sequencing data, it is not possible to determine if sequencing reads with the same sequence originated from the same substrate molecule or different molecules, limiting the ability to interpret data as a representation of molecular frequency in the original population.

[0095] What current methods known to those of skill in the art generally lack is the ability to generate NGS libraries from single-stranded oligonucleotide substrates and produce quantitative data representative of the sequence frequency composition of the original sample.

[0096] One objective of the present disclosure is to provide methods that can generate NGS- compatible sequencing libraries from single- stranded RNA and chimeric DNA-RNA oligonucleotide substrates with reduced molecular bias compared with other commonly used methods.

[0097] It is a further objective of the present disclosure to provide adapter designs and methods to tag individual substrate molecules with unique molecular identifiers (UMIs) to provide a mechanism to correct PCR and sequencing errors.

[0098] It is a further objective of the present disclosure to provide adapter designs and methods to tag individual substrate molecules with UMIs to provide a mechanism for normalizing sequencing read count by UMI to reduce the impact of amplification variability on molecular quantification.

[0099] It is a further objective of the present disclosure to provide adapter designs and data analysis methods to correct PCR and sequencing errors in NGS read data to reduce the frequency of falsepositive variant sequence identification and improve the accuracy of variant sequence quantification.

[0100] The present disclosure features compositions and methods for assessing sample purity and quantifying the number of unique sequences identified in samples of single- stranded nucleic acid molecules (e.g., synthetic gRNA samples, such as synthetic Cas9 gRNA samples or synthetic Casl2a gRNA samples) that address limitations of existing approaches. The methods include methods of preparing Next-Generation Sequencing (NGS) libraries from single- stranded nucleic acid molecules. Also provided herein are methods for sequencing the libraries and processing resulting sequence data to determine the purity of single-stranded nucleic acid samples.I. Compositions for preparing NGS libraries from single-stranded nucleic acid molecules

[0101] In one aspect, provided herein is a sequencing adapter composition comprising: a) a 3’ adapter sequence configured for ligation to a 3’ end of a single-stranded oligonucleotide, the 3’ adapter sequence comprising, in a 5’ to 3’ direction: a ligatable 5’ end, a first linker sequence, a first index primer binding site sequence, and a non-ligatable 3’ end, wherein the 3’ adapter sequence is capable of hybridizing to a reverse transcription primer; and b) a 5’ adapter sequence configured for ligation to a 5’ end of the single- stranded oligonucleotide, the 5’ adapter sequence comprising, in a 5’ to 3’ direction: a non-ligatable 5’ end, a second index primer binding site sequence, a randomized sequence, a second linker sequence, and a ligatable 3’ end.

[0102] In another aspect, provided herein is a sequencing adapter composition comprising: a) a 3’ adapter sequence configured for ligation to a 3’ end of a single-stranded oligonucleotide, the 3’ adapter sequence comprising, in a 5’ to 3’ direction: a ligatable 5’ end, a first linker sequence, a first index primer binding site sequence, and a non-ligatable 3’ end, wherein at least a portion of the 3’ adapter sequence is capable of hybridizing to a reverse transcription primer; and b) a 5’ adapter sequence configured for ligation to a 5’ end of the single-stranded oligonucleotide, the 5’ adapter sequence comprising, in a 5’ to 3’ direction: a non-ligatable 5’ end, a second index primer binding site sequence, and a ligatable 3’ end; wherein at least one of the 3’ adapter sequence and the 5’ adapter sequence comprises a randomized sequence.

[0103] In some embodiments, the 3’ adapter sequence comprises a randomized nucleotide sequence (e.g., a first randomized nucleotide sequence). The randomized sequence of the 3’ adapter may improve ligation efficiency. Without wishing to be bound by any particular theory, the randomized sequence of the 3’ adapter, when located at the 5’ end, can help improve ligation efficiency by overcoming unfavorable secondary structure interactions and nucleotide compatibility between the ligation donor nucleotide (5’ terminus of the 3’ adapter molecule) and ligation acceptor nucleotide (3’ terminus of the insert molecule, e.g., a single- stranded nucleic acid molecule).

[0104] The randomized sequence of the 3’ adapter, or a portion thereof, can also be used as a unique molecular identifier (UMI). In general, UMIs are sequences of nucleotides applied to or identified in polynucleotides that may be used to distinguish individual nucleic acid molecules that are present in an initial reaction from one another. In a sequencing reaction, UMIs may be sequenced along with the insert sequences with which they are associated to determine whetherthe read sequences are derived from one source nucleic acid molecule (e.g., a single- stranded nucleic acid) or another. Commonly, multiple copies derived from a single source molecule are sequenced. In the case of sequencing by synthesis using Illumina’s sequencing technology, the source molecule may be PCR amplified before delivery to a flow cell. For purposes such as accurately measuring the frequency of a certain sequence within a source population of nucleic acid sequences, it could be helpful to group sequences with the same UMI, which indicates that they originated from the same source nucleic acid molecule. UMIs allow this grouping. Additionally, by comparing sequences having the same UMI to one another to produce a consensus sequence, amplification (e.g. PCR) error rates and sequencing error rates can be reduced. For example, a nucleotide difference that is present in some, but not most or all reads with the same UMI may be ignored as not a true representation of the source nucleic acid molecule. Accuracy can be further increased by combining approaches, such as comparing within and between UMIs, and / or between multiple different UMIs. Additional examples of UMIs and uses thereof are provided in, e.g., US20160319345, which is incorporated herein by reference.

[0105] In some embodiments, the randomized sequence of the 3’ adapter has a length of about 1 nucleotides to about 20 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 1 nucleotide to about 6 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 2 nucleotides to about 18 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 4 nucleotides to about 18 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 4 nucleotides to about 16 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 6 nucleotides to about 16 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 6 nucleotides to about 14 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 8 nucleotides to about 14 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 1 nucleotide. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 2 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 3 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 4 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 5 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 6 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a lengthof about 7 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 8 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 9 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 10 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 11 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 12 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 13 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 14 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 15 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 16 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 17 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 18 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 19 nucleotides. In some embodiments, the randomized sequence of the 3’ adapter has a length of about 20 nucleotides.

[0106] In some embodiments, the 3’ adapter sequence does not comprise a stem-loop structure sequence. In some embodiments, less than 5%, 10%, or 20% of the 3’ adapter sequence is capable of forming a secondary structure. In some embodiments, less than 20% of the 3’ adapter sequence is capable of forming a secondary structure. In some embodiments, less than 10% of the 3’ adapter sequence is capable of forming a secondary structure. In some embodiments, less than 5% of the 3’ adapter sequence is capable of forming a secondary structure. In some embodiments, the 3’ adapter sequence is not capable of forming a secondary structure. In some embodiments, the 3’ adapter sequence does not comprise sequences that are substantially complementary to each other. In some embodiments, the 3’ adapter sequence does not comprise sequences that are capable of hybridizing to each other.

[0107] In some embodiments, the 5’ adapter sequence and the 3’ adapter sequence are not complementary or not substantially complementary to each other. In some embodiments, the 5’ adapter sequence and the 3’ adapter sequence are not capable of hybridizing to each other. In some embodiments, less than 5%, 10%, or 20% of the 3’ adapter sequence is complementary to the 5’ adapter sequence. In some embodiments, less than 5%, 10%, or 20% of the 5’ adapter sequence is complementary to the 3’ adapter sequence. In some embodiments, the 5’ adapter sequence and the 3’ adapter sequence are capable of hybridizing to each other.

[0108] In some embodiments, the ligatable 5’ end of the 3’ adapter sequence comprises a 5’- adenylated nucleotide. Use of such a 5’-adenylated nucleotide as a ligation donor may increase ligation efficiency in the absence of adenosine triphosphate (ATP). In some embodiments, the ligatable 5’ end of the 3’ adapter sequence comprises a 5 ’-phosphorylated nucleotide. In some embodiments, the ligatable 3’ end of the 5’ adapter sequence comprises a 3 ’-hydroxylated nucleotide. In some embodiments, the ligatable 5’ end of the 3’ adapter sequence comprises a 5’- adenylated nucleotide and the ligatable 3’ end of the 5’ adapter sequence comprises a 3’- hydroxylated nucleotide. In some embodiments, the ligatable 5’ end of the 3’ adapter sequence comprises a 5 ’-phosphorylated nucleotide and the ligatable 3’ end of the 5’ adapter sequence comprises a 3 ’-hydroxylated nucleotide.

[0109] In some embodiments, the ligatable 5’ end of the 3’ adapter sequence comprises a 5’- hydroxylated nucleotide. In some embodiments, the ligatable 3’ end of the 5’ adapter sequence comprises a 3 ’-phosphorylated nucleotide. In some embodiments, the ligatable 5’ end of the 3’ adapter sequence comprises a 5 ’-hydroxylated nucleotide and the ligatable 3’ end of the 5’ adapter sequence comprises a 3 ’-phosphorylated nucleotide. In some embodiments, 3’ phosphorylated oligonucleotides and / or 5- hydroxylated oligonucleotides may be synthesized using modified nucleotides and solid-phase oligonucleotide synthesis techniques. In some embodiments, 5’ phosphorylation of DNA and RNA oligonucleotides may be performed post-synthesis using, for example, T4 polynucleotide kinase.

[0110] In some embodiments, the non-ligatable 3’ end of the 3’ adapter sequence comprises a 3’ modification of a 3’ nucleotide. The non-ligatable 3’ end of the 3’ adapter sequence prevents ligation at the 3’ end of the 3’ adapter sequence, thereby preventing concatenation of adapters. In some embodiments, the 3’ modification is selected from the group consisting of: a 3’ amino group, a 3’ dideoxy-C modification, a 3’ inverted dT modification, a 3’ C3 spacer modification, or 3’ phosphoryl modification. In some embodiments, the 3’ modification is a 3’ amino group.

[0111] In some embodiments, the non-ligatable 5’ end of the 5’ adapter sequence comprises a 5’ modification of a 5’ nucleotide. The non-ligatable 5’ end of the 5’ adapter sequence prevents ligation at the 5’ end of the 5’ adapter sequence, thereby preventing concatenation of adapters. In some embodiments, the 5’ modification is selected from the group consisting of: a 5’ hydroxyl group, a 5’ dideoxy-C modification, 5’ inverted dideoxy-T modification, or 5’ C3 spacer modification. In some embodiments, the 5’ modification is a 5’ hydroxyl group.

[0112] The 3’ adapter sequence of the present disclosure can serve as a binding site for a reverse transcription primer, as at least a portion of the 3’ adapter sequence is capable of hybridizing to a reverse transcription primer. In some embodiments, the 3’ adapter sequence comprises a sequence that is at least about 80%, at least about 85%, at least about 90%, or at least about 95% identical to a sequence that is complementary to the reverse transcription primer. In some embodiments, the 3’ adapter sequence comprises a sequence that is 100% identical to a sequence that is complementary to the reverse transcription primer. In some embodiments, the 3’ adapter sequence is at least about 80%, at least about 85%, at least about 90%, or at least about 95% identical to a sequence that is complementary to the reverse transcription primer. In some embodiments, the 3’ adapter sequence is 100% identical to a sequence that is complementary to the reverse transcription primer. In some embodiments, the 3’ adapter comprises a randomized sequence. In some embodiments, the non-randomized portion of the 3’ adapter sequence, or a portion thereof, is capable of hybridizing to a reverse transcription primer. In some embodiments, the nonrandomized portion of the 3’ adapter sequence is 100% identical to a sequence that is complementary to the reverse transcription primer. In some embodiments, the reverse transcription primer is not specific to any of the single- stranded nucleic acid molecules. In some embodiments, the 3’ adapter comprises a universal sequence.

[0113] The first linker included in the 3’ adapter of the present disclosure is situated between the first index primer binding site sequence and the insert sequence (generated from the singlestranded nucleic acid molecule), and will immediately precede the 3’ terminus of the insert sequence in each sequencing read. This facilitates quality control in data analysis, helps mitigate the impact of typically lower quality in the first few base calls, and allows for unambiguous determination of the termini of the original substrate molecule (z.e., the single-stranded nucleic acid molecule). In some embodiments, the first linker sequence has a length of about 4 nucleotides to about 20 nucleotides. In some embodiments, the first linker sequence has a length of about 5 nucleotides to about 18 nucleotides. In some embodiments, the first linker sequence has a length of about 5 nucleotides to about 15 nucleotides. In some embodiments, the first linker sequence has a length of about 7 nucleotides to about 14 nucleotides. In some embodiments, the first linker sequence has a length of about 8 nucleotides to about 14 nucleotides. In some embodiments, the first linker sequence has a length of about 8 nucleotides to about 12 nucleotides. In some embodiments, the first linker sequence has a length of 9 nucleotides to 11 nucleotides. In some embodiments, the first linker sequence has a length of about 4 nucleotides. In some embodiments,the first linker sequence has a length of about 5 nucleotides. In some embodiments, the first linker sequence has a length of about 6 nucleotides. In some embodiments, the first linker sequence has a length of about 7 nucleotides. In some embodiments, the first linker sequence has a length of about 8 nucleotides. In some embodiments, the first linker sequence has a length of about 9 nucleotides. In some embodiments, the first linker sequence has a length of about 10 nucleotides.In some embodiments, the first linker sequence has a length of about 11 nucleotides. In some embodiments, the first linker sequence has a length of about 12 nucleotides. In some embodiments, the first linker sequence has a length of about 13 nucleotides. In some embodiments, the first linker sequence has a length of about 14 nucleotides. In some embodiments, the first linker sequence has a length of about 15 nucleotides. In some embodiments, the first linker sequence has a length of about 16 nucleotides. In some embodiments, the first linker sequence has a length of about 17 nucleotides. In some embodiments, the first linker sequence has a length of about 18 nucleotides. In some embodiments, the first linker sequence has a length of about 19 nucleotides. In some embodiments, the first linker sequence has a length of about 20 nucleotides.

[0114] In some embodiments, the first index primer binding site sequence has a length of about 20 nucleotides to about 35 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 22 nucleotides to about 34 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 24 nucleotides to about 32 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 26 nucleotides to about 30 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 20 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 21 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 22 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 23 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 24 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 25 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 26 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 27 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 28 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 29 nucleotides. In some embodiments, the first index primer binding sitesequence has a length of about 30 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 31 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 32 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 33 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 34 nucleotides. In some embodiments, the first index primer binding site sequence has a length of about 35 nucleotides. In some embodiments, the first index primer binding site sequence is capable of binding to a sequencing read 2 adapter. In some embodiments, the first index primer binding site sequence is substantially complementary to a sequencing read 2 adapter. In some embodiments, the sequencing read 2 adapter is an Illumina i7 indexing primer. In some embodiments, the sequencing read 2 adapter is an Illumina i7 universal primer. In some embodiments, the first index primer binding site sequence is capable of binding to a sequencing read 1 adapter. In some embodiments, the first index primer binding site sequence is substantially complementary to a sequencing read 1 adapter. In some embodiments, the sequencing read 1 adapter is an Illumina i5 indexing primer. In some embodiments, the sequencing read 1 adapter is an Illumina i5 universal primer.

[0115] In some embodiments, the 3’ adapter sequence has a total length of about 20 nucleotides to about 60 nucleotides. In some embodiments, the 3’ adapter sequence has a total length of about 25 nucleotides to about 50 nucleotides. In some embodiments, the 3’ adapter sequence has a total length of about 30 nucleotides to about 45 nucleotides. In some embodiments, the 3’ adapter sequence has a total length of about 35 nucleotides to about 45 nucleotides. In some embodiments, the 3’ adapter sequence has a total length of about 35 nucleotides to about 40 nucleotides. In some embodiments, the 3’ adapter sequence has a total length of about 37 nucleotides to about 40 nucleotides. In some embodiments, the 3’ adapter sequence has a total length of about 38 nucleotides.

[0116] In some embodiments, the 3’ adapter sequence comprises a DNA sequence. In some embodiments, the 3’ adapter sequence comprises a single- stranded DNA sequence.

[0117] In some embodiments, the second index primer binding site sequence included in the 5’ adapter of the present disclosure has a length of about 20 nucleotides to about 30 nucleotides. In some embodiments, the second index primer binding site sequence included in the 5’ adapter has a length of about 22 nucleotides to about 28 nucleotides. In some embodiments, the second index primer binding site sequence included in the 5’ adapter has a length of about 24 nucleotides to 1about 28 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 20 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 21 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 22 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 23 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 24 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 25 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 26 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 27 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 28 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 29 nucleotides. In some embodiments, the second index primer binding site sequence has a length of about 30 nucleotides. In some embodiments, the second index primer binding site sequence is capable of binding to a sequencing read 1 adapter. In some embodiments, the second index primer binding site sequence is substantially complementary to a sequencing read 1 adapter. In some embodiments, the sequencing read 1 adapter is an Illumina i5 universal primer. In some embodiments, the sequencing read 1 adapter is an Illumina i5 index primer. In some embodiments, the second index primer binding site sequence is capable of binding to a sequencing read 2 adapter. In some embodiments, the second index primer binding site sequence is substantially complementary to a sequencing read 2 adapter. In some embodiments, the sequencing read 2 adapter is an Illumina i7 indexing primer. In some embodiments, the sequencing read 2 adapter is an Illumina i7 universal primer.

[0118] In some embodiments, the 5’ adapter sequence comprises a randomized sequence. The randomized sequence of the 5’ adapter, or a portion thereof, may improve ligation efficiency. The randomized sequence of the 5’ adapter, or a portion thereof, can also be used as a unique molecular identifier (UMI) similarly as described for the randomized sequence of the 3’ adapter. In embodiments where the 3’ adapter sequence comprises a first randomized sequence and the 5’ adapter sequence comprises a second randomized sequence, a combination of the first and second randomized sequences, or portions thereof, can be used as a UMI, e.g., to increase the total number of unique sequences that can be represented.

[0119] In some embodiments, the randomized sequence has a length of about 6 nucleotides to about 30 nucleotides. In some embodiments, the randomized sequence has a length of about 10nucleotides to about 20 nucleotides. In some embodiments, the randomized sequence has a length of about 12 nucleotides to about 18 nucleotides. In some embodiments, the randomized sequence has a length of about 14 nucleotides to about 18 nucleotides. In some embodiments, the randomized sequence has a length of about 14 nucleotides to about 16 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 6 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 7 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 8 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 9 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 10 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 11 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 12 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 13 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 14 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 15 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 16 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 17 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 18 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 19 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 20 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 21 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 22 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 23 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 24 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 25 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 26 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 27 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 28 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 29 nucleotides. In some embodiments, the randomized sequence of the 5’ adapter has a length of about 30 nucleotides.

[0120] The second linker included in the 5’ adapter of the present disclosure is situated between the randomized sequence and the insert sequence (generated from the single- stranded nucleic acid molecule), and will immediately precede the 5’ terminus of the insert sequence in each sequencing read. Similar to the first linker of the 3’ adapter, this facilitates quality control in data analysis, helps mitigate the impact of typically lower quality in the first few base calls, and allows for unambiguous determination of the termini of the original substrate molecule (z.e., the singlestranded nucleic acid molecule). In some embodiments, the second linker sequence has a length of about 4 nucleotides to about 20 nucleotides. In some embodiments, the second linker sequence has a length of about 6 nucleotides to about 20 nucleotides. In some embodiments, the second linker sequence has a length of about 6 nucleotides to about 18 nucleotides. In some embodiments, the second linker sequence has a length of about 8 nucleotides to about 18 nucleotides. In some embodiments, the second linker sequence has a length of about 8 nucleotides to about 16 nucleotides. In some embodiments, the second linker sequence has a length of about 10 nucleotides to about 16 nucleotides. In some embodiments, the second linker sequence has a length of about 10 nucleotides to about 14 nucleotides. In some embodiments, the second linker sequence has a length of about 10 nucleotides to about 12 nucleotides. In some embodiments, the second linker sequence has a length of about 4 nucleotides to about 20 nucleotides. In some embodiments, the second linker sequence has a length of about 4 nucleotides. In some embodiments, the second linker sequence has a length of about 5 nucleotides.some embodiments, the second linker sequence has a length of aboutnucleotides.some embodiments, the second linker sequence has a length of about 7 nucleotides. In some embodiments, the second linker sequence has a length of about 8 nucleotides. In some embodiments, the second linker sequence has a length of about 9 nucleotides. In some embodiments, the second linker sequence has a length of about 10 nucleotides.some embodiments, the second linker sequence has a lengthaboutnucleotides.some embodiments, the second linker sequence has a lengthaboutnucleotides. In some embodiments, the second linker sequence has a length of about 13 nucleotides. In some embodiments, the second linker sequence has a length of about 14 nucleotides. In some embodiments, the second linker sequence has a length of about 15 nucleotides.some embodiments, the second linker sequence has a length of aboutnucleotides.some embodiments, the second linker sequence has a length of aboutnucleotides. In some embodiments, the second linker sequence has a length of about 18 nucleotides. In someembodiments, the second linker sequence has a length of about 19 nucleotides. In some embodiments, the second linker sequence has a length of about 20 nucleotides.

[0121] In some embodiments, the 5’ adapter sequence has a total length of about 40 nucleotides to about 60 nucleotides. In some embodiments, the 5’ adapter sequence has a total length of about 42 nucleotides to about 55 nucleotides. In some embodiments, the 5’ adapter sequence has a total length of about 48 nucleotides.

[0122] In some embodiments, the 5’ adapter sequence comprises an RNA sequence or a DNA sequence. In some embodiments, the 5’ adapter sequence comprises an RNA sequence. In some embodiments, the 5’ adapter sequence comprises a single-stranded RNA sequence or a singlestranded DNA sequence. In some embodiments, the 5’ adapter sequence comprises a singlestranded RNA sequence.

[0123] An example of a 3’ adapter is illustrated in FIG. 1. The 3’ adapter shown in FIG. 1 is single-stranded DNA (“ssDNA”) and comprises, from 5’ to 3’: a ligatable 5’ end comprising a 5’- adenylated nucleotide, a first linker sequence having a length of 10 nucleotides, an Illumina i7 indexing primer binding site sequence having a length of 28 nucleotides, and a non-ligatable 3’ end comprising a 3 ’-aminated nucleotide. The 3’ adapter as a whole serves as a binding site for a reverse transcription primer. In the depicted embodiment, the total length of the 3’ adapter is 38 nucleotides.

[0124] In some embodiments, the 3’ adapter of the present disclosure comprises the nucleotide sequence set forth in SEQ ID NO: 1 or a variant thereof. In some embodiments, the 3’ adapter of the present disclosure comprises a sequence having at least about 80% similarity to the nucleotide sequence set forth in SEQ ID NO: 1. In some embodiments, the 3’ adapter of the present disclosure comprises a sequence having at least about 85% similarity to the nucleotide sequence set forth in SEQ ID NO: 1. In some embodiments, the 3’ adapter of the present disclosure comprises a sequence having at least about 90% similarity to the nucleotide sequence set forth in SEQ ID NO: 1. In some embodiments, the 3’ adapter of the present disclosure comprises a sequence having at least about 95% similarity to the nucleotide sequence set forth in SEQ ID NO: 1. In some embodiments, the 3’ adapter of the present disclosure comprises the nucleotide sequence set forth in SEQ ID NO: 1.

[0125] In some embodiments, there is provided a 3’ adapter sequence configured for ligation to a 3’ end of a single- stranded oligonucleotide, the 3’ adapter sequence comprising, in a 5’ to 3’ direction: a ligatable 5’ end, a first randomized sequence, a first linker sequence, a first index primer binding site sequence, and a non-ligatable 3’ end. An example of such a 3’ adapter is illustrated in FIG. 8. The 3’ adapter shown in FIG. 8 is single- stranded DNA (“ssDNA”) and comprises, from 5’ to 3’: a ligatable 5’ end comprising a 5 ’-adenylated nucleotide, a first randomized sequence having a length of 2 nucleotides, a first linker sequence having a length of 8 nucleotides, an Illumina i7 indexing primer binding site sequence having a length of 28 nucleotides, and a non-ligatable 3’ end comprising a 3 ’-aminated nucleotide. The non-randomized portion of the 3’ adapter (consisting of the first linker and Illumina i7 indexing primer binding site sequence) serves as a binding site for a reverse transcription primer. In the depicted embodiment, the total length of the 3’ adapter is 38 nucleotides.

[0126] In some embodiments, the 3’ adapter of the present disclosure comprises the nucleotide sequence set forth in SEQ ID NO: 106 or a variant thereof. In some embodiments, the 3’ adapter of the present disclosure comprises a sequence having at least about 80% similarity to the nucleotide sequence set forth in SEQ ID NO: 106. In some embodiments, the 3’ adapter of the present disclosure comprises a sequence having at least about 85% similarity to the nucleotide sequence set forth in SEQ ID NO: 106. In some embodiments, the 3’ adapter of the present disclosure comprises a sequence having at least about 90% similarity to the nucleotide sequence set forth in SEQ ID NO: 106. In some embodiments, the 3’ adapter of the present disclosure comprises a sequence having at least about 95% similarity to the nucleotide sequence set forth in SEQ ID NO: 106. In some embodiments, the 3’ adapter of the present disclosure comprises the nucleotide sequence set forth in SEQ ID NO: 106.

[0127] An example of a 5’ adapter is illustrated in FIG. 2. The 5’ adapter shown in FIG. 2 is single- stranded RNA (“ssRNA”) and comprises, from 5’ to 3’: a non-ligatable 5’ end comprising a 5 ’-hydroxylated nucleotide, an Illumina i5 universal primer binding site sequence having a length of 26 nucleotides, a UMI (randomized sequence) having a length of 15 nucleotides, a second linker sequence having a length of 7 nucleotides, and a ligatable 3’ end comprising a 3 ’-hydroxylated nucleotide. In the depicted embodiment, the total length of the 5’ adapter is 48 nucleotides.

[0128] In some embodiments, the 5’ adapter of the present disclosure comprises the nucleotide sequence set forth in SEQ ID NO: 2 or a variant thereof. In some embodiments, the 5’ adapter ofthe present disclosure comprises a sequence having at least about 80% similarity to the nucleotide sequence set forth in SEQ ID NO: 2. In some embodiments, the 5’ adapter of the present disclosure comprises a sequence having at least about 85% similarity to the nucleotide sequence set forth in SEQ ID NO: 2. In some embodiments, the 5’ adapter of the present disclosure comprises a sequence having at least about 90% similarity to the nucleotide sequence set forth in SEQ ID NO: 2. In some embodiments, the 5’ adapter of the present disclosure comprises a sequence having at least about 95% similarity to the nucleotide sequence set forth in SEQ ID NO: 2. In some embodiments, the 5’ adapter of the present disclosure comprises the nucleotide sequence set forth in SEQ ID NO: 2.

[0129] In some embodiments, the 5’ adapter of the present disclosure comprises the nucleotide sequence set forth in SEQ ID NO: 107 or a variant thereof. In some embodiments, the 5’ adapter of the present disclosure comprises a sequence having at least about 80% similarity to the nucleotide sequence set forth in SEQ ID NO: 107. In some embodiments, the 5’ adapter of the present disclosure comprises a sequence having at least about 85% similarity to the nucleotide sequence set forth in SEQ ID NO: 107. In some embodiments, the 5’ adapter of the present disclosure comprises a sequence having at least about 90% similarity to the nucleotide sequence set forth in SEQ ID NO: 107. In some embodiments, the 5’ adapter of the present disclosure comprises a sequence having at least about 95% similarity to the nucleotide sequence set forth in SEQ ID NO: 107. In some embodiments, the 5’ adapter of the present disclosure comprises the nucleotide sequence set forth in SEQ ID NO: 107.

[0130] In some embodiments, there is provided a sequencing adapter composition comprising: a) a 3’ adapter sequence configured for ligation to a 3’ end of a single- stranded oligonucleotide, the 3’ adapter sequence comprising, in a 5’ to 3’ direction: a ligatable 5’ end, a first randomized sequence, a first linker sequence, a first index primer binding site sequence, and a non-ligatable 3’ end, wherein a non-randomized portion of the 3’ adapter sequence is capable of hybridizing to a reverse transcription primer; and b) a 5’ adapter sequence configured for ligation to a 5’ end of the single- stranded oligonucleotide, the 5’ adapter sequence comprising, in a 5’ to 3’ direction: a non- ligatable 5’ end, a second index primer binding site sequence, a second randomized sequence, a second linker sequence, and a ligatable 3’ end. In some embodiments, the 3’ adapter sequence does not comprise a stem- loop structure sequence. In some embodiments, the ligatable 5’ end of the 3’ adapter sequence comprises a 5’-adenylated nucleotide. In some embodiments, the non- ligatable 3’ end of the 3’ adapter sequence comprises a 3 ’-aminated nucleotide. In someembodiments, the ligatable 3’ end of the 5’ adapter sequence comprises a 3 ’-hydroxylated nucleotide. In some embodiments, the non-ligatable 5’ end of the 5’ adapter sequence comprises a 5 ’-hydroxylated nucleotide. In some embodiments, the first randomized sequence has a length of 2 nucleotides. In some embodiments, the second randomized sequence has a length of about 15 nucleotides.IL Methods of preparing NGS libraries from single-stranded nucleic acid molecules

[0131] In one aspect, provided herein is a method of preparing a sequencing library for a singlestranded template nucleic acid molecule, comprising: a) contacting the single-stranded nucleic acid molecule with 3’ adapter sequence of the present disclosure and performing a ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single- stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule; b) performing an end modification reaction to modify a 5’ end of the first ligated nucleic acid molecule; c) contacting the first ligated nucleic acid molecule with the 5’ adapter sequence of the present disclosure and performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule; d) contacting the second ligated nucleic acid molecule with a reverse transcription primer and allowing the reverse transcription primer to hybridize to the 3’ adapter sequence; e) performing a reverse transcription reaction, thereby producing a cDNA molecule comprising the template nucleic acid molecule; f) contacting the cDNA molecule with a first index primer capable of hybridizing to the first index primer binding site sequence, and with a second index primer capable of hybridizing to the second index primer binding site sequence, under conditions that allow hybridization of the first and second index primers to their respective binding site sequences; g) amplifying the cDNA molecule to produce an amplification product; and h) purifying the amplification product to produce the sequencing library for the single- stranded template nucleic acid molecule. In some embodiments, the single-stranded template nucleic acid molecule is an RNA molecule. In some embodiments, the single- stranded template nucleic acid molecule is a guide RNA (gRNA) molecule. In some embodiments, the gRNA molecule is recovered from a gRNA-Cas nuclease ribonucleoprotein (RNP) by complex denaturation and nucleic acid purification. In some embodiments, the singlestranded template nucleic acid molecule is a small interfering RNA (siRNA) molecule. In some embodiments, the single- stranded template nucleic acid molecule is a prime editing guide RNA (pegRNA) molecule. In some embodiments, the single- stranded template nucleic acid molecule is a trans activating CRISPR RNA (tracrRNA) molecule. In some embodiments, the single- strandedtemplate nucleic acid molecule is a CRISPR RNA (crRNA) molecule. In some embodiments, the single- stranded template nucleic acid molecule is a DNA molecule. In some embodiments, the single- stranded template nucleic acid molecule is a DNA / RNA chimeric molecule.

[0132] In some embodiments, the end modification reaction allows the 5’ end of the first ligated nucleic acid molecule to be ligatable. In some embodiments, the end modification reaction comprises an adenylation reaction to adenylate the 5’ end of the first ligated nucleic acid molecule. In some embodiments, the end modification reaction comprises a phosphorylation reaction to phosphorylate the 5’ end of the first ligated nucleic acid molecule. In some embodiments, the 5’ phosphorylation reaction is carried out in the presence of one or more chemical additives that impact nucleotide-nucleotide and nucleotide-protein interactions. In some embodiments, the one or more chemical additives is selected from the group consisting of DMSO, formamide, TMAC, betaine, ammonium sulfate, glycerol, and polyethylene glycol.

[0133] In some embodiments, ligation of the 3’ adapter sequence to a 3’ end of the single- stranded template nucleic acid molecule comprises the use of a DNA ligase or an RNA ligase. In some embodiments, the ligase is RtcB ligase, T4 RNA Ligase 1, T4 RNA Ligase 2, a truncated version of T4 RNA Ligase 2 (e.g., T4 RNA Ligase 2 Truncated, T4 RNA Ligase 2 Truncated K227Q, T4 RNA Ligase 2 truncated KQ), Thermostable 5' App DNA / RNA Ligase. In some embodiments, the ligase is RtcB ligase. In some embodiments, the ligase is T4 RNA ligase 2 truncated.

[0134] In some embodiments, ligation of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule comprises the use of a DNA ligase or an RNA ligase. In some embodiments, the ligase is RtcB ligase, T4 RNA Ligase 1, T4 RNA Ligase 2, T4 RNA Ligase 2 Truncated, T4 RNA Ligase 2 Truncated K227Q, T4 RNA Ligase 2 truncated KQ, Thermostable 5' App DNA / RNA Ligase. In some embodiments, the ligase is RtcB ligase. In some embodiments, the ligase is T4 RNA ligase 1.

[0135] In some embodiments, the 3’ ligation reaction is carried out in the presence of one or more chemical additives, such as those that impact nucleotide-nucleotide and nucleotide-protein interactions. In some embodiments, the one or more chemical additives is selected from the group consisting of DMSO, formamide, TMAC, betaine, ammonium sulfate, glycerol, and polyethylene glycol. In some embodiments, the 3’ ligation reaction is carried out in the presence of polyethylene glycol (PEG). In some embodiments, the polyethylene glycol (PEG) is PEG 2K, 3K, 4K, 5K, 6K, 7K, 8K, 9K, 10K, UK, 12K, 13K, 14K, 15K, or 16K. In some embodiments, ligation of the 3’adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule comprises the use of a reaction buffer containing about 10% to about 25% (v / v) polyethylene glycol (PEG). In some embodiments, ligation of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule comprises the use of a reaction buffer containing about 10%, about 15%, about 20%, or about 25% (v / v) polyethylene glycol (PEG). In some embodiments, the reaction buffer contains about 25% (v / v) polyethylene glycol (PEG). In some embodiments, the reaction buffer contains about 25% (v / v) PEG 8K (PEG8000). In some embodiments, ligation of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule comprises the use of a reaction buffer containing about 5% to about 15% (v / v) glycerol. In some embodiments, ligation of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule comprises the use of a reaction buffer containing 10% (v / v) glycerol.

[0136] In some embodiments, the 3’ ligation reaction is carried out in the presence of a singlestranded DNA binding protein. Examples of single-stranded DNA binding proteins include but are not limited to Escherichia coli single- stranded DNA-binding protein (EcoSSB) and Extreme Thermostable Single-Stranded DNA Binding Protein (ET SSB).

[0137] In some embodiments, the 5’ ligation reaction contains one or more chemical additives, such as those impact nucleotide-nucleotide and nucleotide-protein interactions. In some embodiments, the one or more chemical additives is selected from the group consisting of DMSO, formamide, TMAC, betaine, ammonium sulfate, glycerol, and polyethylene glycol (PEG). In some embodiments, the 5’ ligation reaction is carried out in the presence of polyethylene glycol (PEG). In some embodiments, the polyethylene glycol (PEG) is PEG 2K, 3K, 4K, 5K, 6K, 7K, 8K, 9K, 10K, UK, 12K, 13K, 14K, 15K, or 16K. In some embodiments, ligation of the 5’ adapter sequence to a 3’ end of the single- stranded template nucleic acid molecule comprises the use of a reaction buffer containing about 10% to about 25% (v / v) polyethylene glycol (PEG). In some embodiments, ligation of the 5’ adapter sequence to a 5’ end of the single-stranded template nucleic acid molecule comprises the use of a reaction buffer containing about 10%, about 15%, about 20%, or about 25% (v / v) polyethylene glycol (PEG). In some embodiments, the reaction buffer contains about 25% (v / v) polyethylene glycol (PEG). In some embodiments, the reaction buffer contains about 25% (v / v) PEG 8K (PEG8000). In some embodiments, ligation of the 5’ adapter sequence to a 5’ end of the single-stranded template nucleic acid molecule comprises the use of a reaction buffer containing about 5% to about 15% (v / v) glycerol. In some embodiments,ligation of the 5’ adapter sequence to a 5’ end of the single-stranded template nucleic acid molecule comprises the use of a reaction buffer containing 10% (v / v) glycerol.

[0138] In some embodiments, ligation of the 3’ adapter sequence to a 3’ end of the single- stranded template nucleic acid molecule is carried out at a reaction temperature of about 4°C to about 37 °C, such as at about 4°C, about 5°C, about 10°C, about 15°C, about 20°C, about 25°C, about 30°C, about 35°C, or about 37°C. In some embodiments, ligation of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule is carried out at a reaction temperature of 25°C. In some embodiments, ligation of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule is carried out in variable reaction temperature conditions cycling between 4°C and 37°C. In some embodiments, the variable reaction temperature conditions may improve the capture of the broadest possible selection of template nucleic acid molecules by carrying out the reaction at multiple temperatures where different temperatures may capture a different subset of molecules.

[0139] In some embodiments, ligation of the 5’ adapter sequence to a 5’ end of the single- stranded template nucleic acid molecule is carried out at a reaction temperature of about 4°C to about 37 °C, such as at about 4°C, about 5°C, about 10°C, about 15°C, about 20°C, about 25°C, about 30°C, about 35°C, or about 37°C. In some embodiments, ligation of the 5’adapter sequence to a 5’ end of the single-stranded template nucleic acid molecule is carried out at a reaction temperature of 25°C. In some embodiments, ligation of the 5’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule is carried out in variable reaction temperature conditions cycling between 4°C and 37°C.

[0140] In general, reverse transcription comprises extension of an oligonucleotide primer hybridized to a target RNA (e.g., a template nucleic acid molecule) by an RNA-dependent DNA polymerase (also referred to as a “reverse transcriptase”), using the target RNA molecule as the template to produce a complementary DNA (cDNA). Examples of reverse transcriptases include, but are not limited to, retroviral reverse transcriptase (e.g., Moloney Murine Leukemia Virus (MMLV), Avian Myeloblastosis Virus (AMV) or Rous Sarcoma Virus (RSV) reverse transcriptases), Superscript I™, Superscript II™, Superscript III™, retrotransposon reverse transcriptase, hepatitis B reverse transcriptase, cauliflower mosaic virus reverse transcriptase, bacterial reverse transcriptase, and mutants, variants or derivatives thereof. In some embodiments, the reverse transcriptase is a hot-start reverse transcriptase enzyme.

[0141] The immediate product of a typical RT reaction is an RNA-DNA hybrid molecule comprising the template RNA (and optionally a TSO) hybridized to the cDNA resulting from primer extension. In some embodiments, the RNA-DNA hybrids are denatured and / or the RNA template is degraded as part of or subsequent to the RT reaction. For example, the RNA-DNA hybrids can be denatured in the presence of an enzyme that degrades RNA, such as RNase A. In some embodiments, RNA of the RNA-DNA hybrids is degraded without denaturing the complex by using an enzyme having such activity, such as RNase H. In some embodiments, the reverse transcriptase in the RT reaction comprises RNase H activity.

[0142] In some embodiments, the method does not comprise a step of cleaving the first ligated nucleic acid molecule or the second ligated nucleic acid molecule.

[0143] In some embodiments, the method further comprises one or more purification steps. In some embodiments, the method comprises one or more purification steps after the contacting in step a). In some embodiments, the method comprises one or more purification steps after the end modification reaction in step b). In some embodiments, the method comprises one or more purification steps after the contacting in step c). In some embodiments, the method comprises one or more purification steps after the contacting in step f). Purification can be performed by methods known in the art. In some embodiments, purification is performed by buffer exchange by magnetic bead purification.

[0144] An exemplary workflow of a method of preparing NGS libraries from single- stranded nucleic acid molecules is illustrated in FIG. 3. In this exemplary workflow, a single-stranded gRNA molecule is contacted with and ligated to a 3’ adapter, which comprises a 5’ adenyl modification and a 3’ amino modification. After buffer exchange by magnetic bead purification, the ligated molecule is modified by 5’ phosphorylation, after which the modified ligated molecule is contacted with and ligated to a 5’ adapter comprising a 3’ hydroxyl group and a 5’ hydroxyl group. The 5’ adapter also comprises a UMI. Another buffer exchange by magnetic bead purification is performed. Reverse transcription is then carried out using a reverse transcription primer that hybridizes to the 3’ adapter to make a cDNA. After buffer exchange by magnetic bead purification, the cDNA is contacted with indexing primers and amplified via PCR. Lastly, the amplified nucleic acids are purified to make a gRNA sequencing library.III. Methods and systems for detecting and quantifying unique variant sequences in a population of single-stranded template nuclei acid molecules

[0145] In one aspect, provided herein is a method for detecting and quantifying unique sequences in a population of single- stranded template nucleic acid molecules, the method comprising: a) preparing a sequencing library for the single- stranded template nucleic acid molecule using any of the methods described herein; b) sequencing the sequencing library using a paired-end sequencing method to generate a plurality of sequence reads corresponding to the single- stranded template nucleic acid molecule; and c) analyzing sequence read data corresponding to the plurality of sequence reads to achieve at least one of: (i) correction of amplification errors; (ii) correction of sequencing errors; (iii) normalization of sequence-based amplification bias; and (iv) detection and quantification of unique sequences present in the population of single-stranded template nucleic acid molecules. In some instances, the population of single- stranded template nucleic acid molecules is a putatively homogenous population.

[0146] In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is about 10 to about 600 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is about 15 to about 600 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is about 15 to about 500 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is about 20 to about 500 nucleotides. In some embodiments, the average length of the population of singlestranded template nucleic acid molecules is about 20 to about 400 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is about 25 to about 400 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is about 25 to about 300 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is about 30 to about 300 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is about 30 to about 200 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is about 35 to about 200 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is about 35 to about 100 nucleotides. In some embodiments, the average length of the population of singlestranded template nucleic acid molecules is about 40 to about 100 nucleotides. In someembodiments, the average length of the population of single- stranded template nucleic acid molecules is about 10 to about 20 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is about 20 to about 30 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is about 30 to about 40 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is about 40 to about 50 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is about 50 to about 60 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is about 60 to about 70 nucleotides. In some embodiments, the average length of the population of singlestranded template nucleic acid molecules is about 70 to about 80 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is about 80 to about 90 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is about 90 to about 100 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is at least 10 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is at least 20 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is at least 30 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is at least 40 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is at least 50 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is not more than 600 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is not more than 500 nucleotides. In some embodiments, the average length of the population of single-stranded template nucleic acid molecules is not more than 400 nucleotides. In some embodiments, the average length of the population of single- stranded template nucleic acid molecules is not more than 300 nucleotides.

[0147] In some instances, analyzing sequence read data comprises the use of one or more processors.

[0148] In some instances, analyzing the sequence read data comprises aligning paired sequence reads (as discussed in more detail below) and discarding discrepant read pairs from further analysis.

[0149] In some instances, as discussed elsewhere herein, analyzing the sequence read data comprises grouping sequence reads according to unique molecular identifiers (UMIs) associated with each sequence read. In some instances, for example, a unique molecular identifier (UMI) associated with each sequence read may correspond to a specific member of the randomized sequence in the 3’ adapter sequence and / or 5’ adapter sequence described above. In some instances, the method can further comprise determining a consensus sequence for each group of sequence reads identified based on the associated unique molecular identifier (UMI).

[0150] In some instances, the method may further comprise identifying unique sequences present in the population of single-stranded template nucleic acid molecules (e.g., the intended sequence in a synthetic oligonucleotide sample and / or variant sequences arising from synthesis errors) based on the determined consensus sequence for each group of sequence reads.

[0151] In some instances, the method may further comprise correcting for amplification and / or sequencing errors based on the determined consensus sequence for each group of sequence reads (e.g., by correcting or discarding variant sequences that differ from the consensus sequence for each group of sequence reads by 1 or more bases (e.g., 2, 3, or 4 bases)).

[0152] In some instances, the method may further comprise aligning a consensus sequence for each group of sequence reads to a reference sequence (e.g. , a known target sequence for a synthetic oligonucleotide samples, or a reference genomic sequence (e.g., the human genome reference sequence) for a nucleic acid sample collected from a biological source) to identify variant sequences present in the population of single-stranded template nucleic acid molecules.

[0153] In some embodiments, the length of the reference sequence is about 10 to about 600 nucleotides. In some embodiments, the length of the reference sequence is about 15 to about 600 nucleotides. In some embodiments, the length of the reference sequence is about 15 to about 500 nucleotides. In some embodiments, the length of the reference sequence is about 20 to about 500 nucleotides. In some embodiments, the length of the reference sequence is about 20 to about 400 nucleotides. In some embodiments, the length of the reference sequence is about 25 to about 400 nucleotides. In some embodiments, the length of the reference sequence is about 25 to about 300nucleotides. In some embodiments, the length of the reference sequence is about 30 to about 300 nucleotides. In some embodiments, the length of the reference sequence is about 30 to about 200 nucleotides. In some embodiments, the length of the reference sequence is about 35 to about 200 nucleotides. In some embodiments, the length of the reference sequence is about 35 to about 100 nucleotides. In some embodiments, the length of the reference sequence is about 40 to about 100 nucleotides. In some embodiments, the length of the reference sequence is about 10 to about 20 nucleotides. In some embodiments, the length of the reference sequence is about 20 to about 30 nucleotides. In some embodiments, the length of the reference sequence is about 30 to about 40 nucleotides. In some embodiments, the length of the reference sequence is about 40 to about 50 nucleotides. In some embodiments, the length of the reference sequence is about 50 to about 60 nucleotides. In some embodiments, the length of the reference sequence is about 60 to about 70 nucleotides. In some embodiments, the length of the reference sequence is about 70 to about 80 nucleotides. In some embodiments, the length of the reference sequence is about 80 to about 90 nucleotides. In some embodiments, the length of the reference sequence is about 90 to about 100 nucleotides. In some embodiments, the length of the reference sequence is at least 10 nucleotides.In some embodiments, the length of the reference sequence is at least 20 nucleotides. In some embodiments, the length of the reference sequence is at least 30 nucleotides. In some embodiments, the length of the reference sequence is at least 40 nucleotides. In some embodiments, the length of the reference sequence is at least 50 nucleotides. In some embodiments, the length of the reference sequence is not more than 600 nucleotides. In some embodiments, the length of the reference sequence is not more than 500 nucleotides. In some embodiments, the length of the reference sequence is not more than 400 nucleotides. In some embodiments, the length of the reference sequence is not more than 300 nucleotides. In some embodiments, the length of the reference sequence is about 66 nucleotides.

[0154] In some instances, the method may further comprising determining a frequency, fraction, or percentage of sequences present in the population of single- stranded template nucleic acid molecules that are identical to the reference sequence, and / or determining a variant sequence frequency, fraction, or percentage for a variant sequence (e.g., 1, 2, 3, 4, 5, or more than 5 variant sequences) identified in the population of single-stranded template nucleic acid molecules. The variant sequence may be a known variant with a known sequence, or an unknown variant with an unknown sequence.

[0155] In some instances, as discussed elsewhere herein, the population of single- stranded template nucleic acid molecules may comprise guide RNA (gRNA) molecules (e.g., Cas9 gRNA molecules or Casl2a gRNA molecules). In some instances, the population of single-stranded template nucleic acid molecules may comprise gRNA molecules recovered from gRNA-Cas nuclease ribonucleoproteins (RNPs). In some instances, the population of single-stranded template nucleic acid molecules may comprise prime editing guide RNA (pegRNA) molecule. In some instances, the population of single- stranded template nucleic acid molecules may comprise small interfering RNA (siRNA) molecules. In some instances, the population of single-stranded template nucleic acid molecules may comprise DNA molecules.

[0156] As noted above, in some instances, the analysis of the sequence read data generated for a sequencing library prepared as described elsewhere herein may comprise the use of one or more processors or computer systems. Accordingly, also disclosed herein are systems comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receive sequence read data for a sequencing library generated using any of the methods described herein; and analyze the sequence read data for the sequencing library to achieve at least one of: (i) correction of amplification errors; (ii) correction of sequencing errors; (iii) normalization of sequence-based amplification bias; and (iv) detection and quantification of unique sequences present in a population of single- stranded template nucleic acid molecules.

[0157] In some instances, analysis of the sequence read data by the system may comprise aligning paired sequence reads and discarding discrepant read pairs from further analysis.

[0158] In some instances, analysis of the sequence read data by the system may comprise grouping sequence reads according to unique molecular identifiers (UMIs) associated with each sequence read, where the unique molecular identifier (UMI) associated with each sequence read corresponds to a member of the randomized sequence in the 3’ adapter sequence and / or 5’ adapter sequence described above. In some instances, the analysis may further comprise determining a consensus sequence for each group of sequence reads identified based on the associated unique molecular identifier (UMI). In some instances, the analysis may further comprise identifying unique sequences present in the population of single- stranded template nucleic acid molecules based on the determined consensus sequence for each group of sequence reads. In some instances, theanalysis may further comprise correcting for amplification and / or sequencing errors based on the determined consensus sequence for each group of sequence reads.

[0159] In some instances, the analysis may further comprise aligning a consensus sequence for each group of sequence reads to a reference sequence to identify variant sequences present in the population of single-stranded template nucleic acid molecules. In some instances, the analysis may further comprise determining a frequency, fraction, or percentage of sequences present in the population of single-stranded template nucleic acid molecules that are identical to the reference sequence (e.g., a known target sequence for a synthetic oligonucleotide samples, or a reference genomic sequence (e.g., the human genome reference sequence) for a nucleic acid sample collected from a biological source), and / or determining a variant sequence frequency, fraction, or percentage for a variant sequence (e.g., 1, 2, 3, 4, 5, or more than 5 variant sequences) identified in the population of single-stranded template nucleic acid molecules.

[0160] An exemplary workflow for sequence read data analysis is illustrated in FIG. 4, including grouping sequence reads by UMI, collapsing reads with the same UMI into consensus sequences to correct amplification and sequencing errors, aligning consensus sequences to a reference intended sequence, providing a descriptive variant call for each unique sequence observed, and quantifying the frequency of the intended sequence and of each variant sequence observed. “UMI 1” in the illustration depicts reads containing amplification or sequencing errors which could potentially be counted as false-positive sequence variants. UMI error correction creates a consensus sequence from all reads with the same “UMI 1” sequence and corrects the errors. “UMI 2” in the illustration depicts reads containing a true variant sequence, which is confirmed by UMI consensus sequence collapsing. The number of sequencing reads in the “UMI 1” and “UMI 2” groups is different, but UMI collapsing normalizes each UMI group to a single consensus sequence to avoid quantification bias.IV. Kits for preparing sequencing libraries and / or sequencing single-stranded nuclei acid molecules

[0161] In one aspect, provided herein are kits for preparing sequencing libraries and / or sequencing single- stranded nucleic acid molecules. In some instances, for example, the kit may comprise a sequencing adapter composition as described elsewhere herein, and at least one of: a ligase (e.g., a ligase as described elsewhere herein); a buffer (e.g. , a ligation buffer, a nucleic acid amplification buffer, etc.); enzymes and / or reagents for performing an end modification reaction to modify a 5’end of a nucleic acid molecule; enzymes and / or reagents for performing nucleic acid amplification (e.g., a DNA or RNA polymerase and / or nucleotides (e.g., deoxy nucleotides or ribonucleotides)); or enzymes and / or reagents for performing paired-end sequencing (e.g., sequencing primers, fluorescently labeled nucleotides (e.g., fluorescently-labeled deoxynucleotides or fluorescently- labeled ribonucleotides)).V. CRISPR-Cas gene editing and gRNA features

[0162] CRISPR genome editing technologies utilize a Cas enzyme and a guide RNA (gRNA) to generate mutations, gene knockouts, and knock-ins at targeted genomic locations. Naturally- occurring Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) loci comprising direct repeat sequences interspaced by similarly sized variable sequences are present in prokaryotes. In prokaryotes, a combination of the repeat and variable sequences of the CRISPR loci (transcribed and processed to generate CRISPR RNA (crRNA) molecules and associated proteins (z.e., CRISPR-associated (Cas) proteins comprising both helicase and endonuclease activities) function as a primitive adaptive immune system in prokaryotes, and work in tandem to recognize and cleave invading foreign DNA. The variable sequences (spacers) are relics of previous infection events where fragments of invading DNA (protospacers) have been incorporated into the host genome at the CRISPR locus to serve as an “immunological memory”. The characterization of naturally-occurring CRISPR-Cas systems as a type of immune system in prokaryotes has led to the development of powerful gene editing tools for a variety of biotechnology applications (see, e.g., Nowak et al. (2016), “Guide RNA Engineering for Versatile Cas9 Functionality”, Nucleic Acids Research 44(20):9555-9564; and Mohr et al. (2016), “CRISPR guide RNA Design for Research Applications”, The FEBS Journal 283:3232-3238).

[0163] With the CRISPR-Cas9 system, for example, targeting of the endonuclease activity to a particular genomic locus relies on the guide RNA (gRNA) and the presence of a “protospacer adjacent motif’ (PAM) sequence at the targeted DNA segment. In some instances, the guide RNA can be a duplex structure comprising two RNA molecules: a CRISPR RNA (crRNA) molecule and a trans-activating crRNA (tracrRNA) molecule (equivalent to the endogenous duplex structures found in prokaryotes). The crRNA component (comprising a spacer sequence (about 20 nucleotides in length) and a repeat sequence (about 22 nucleotides in length)) is responsible for binding to the targeted DNA region, while the tracrRNA component is responsible for activation of the Cas9 endonuclease activity. In some instances, the guide RNA can be a single guide RNA (sgRNA) molecule comprising an artificially-engineered combination of a crRNA molecule and atracrRNA molecule that are linked by a short tetraloop structure (where the overall length of the sgRNA molecule is typically about 100 nucleotides). The tracrRNA consists of base pairs that form a stem- loop structure, thereby enabling its attachment to the Cas9 protein. The spacer region (typically 18-20 base pairs (bp) in length) of the crRNA (or sgRNA) guides the Cas9 endonuclease to a complementary target region on a template DNA molecule. Recognition of an NGG trinucleotide sequence (for Streptococcus pyogenes Cas9) at the 5’ end of a protospacer adjacent motif (PAM) sequence (a short G-rich oligonucleotide sequence downstream of the targeted DNA cleavage site) initiates complementary base pairing between the spacer sequence of the crRNA (or sgRNA) molecule and the target DNA strand opposite the PAM site, and triggers cleavage of the target DNA by the nuclease activity of Cas9. The 5’ end sequence motif of the PAM sequence differs for Cas9 proteins from different species (see, e.g., Nowak et al. (2016), ibid.).

[0164] Modifications to the spacer sequence within the gRNA can alter the targeted location for binding of the sgRNA-Cas9 catalytically-active ribonucleoprotein (RNP) complex to the target DNA, thereby enabling programmable gene editing. Design considerations for generating application-specific gRNA molecules includes: (i) defining the target DNA region or gene; (ii) specifying the Cas9 protein to be used (and the corresponding Cas9-specific PAM sequence(s) to be recognized); and, for in vitro or in vivo expression of the gRNA (as opposed to chemical synthesis of the gRNA), (iii) specifying what promoter and / or cloning strategy will be used, which may put additional constraints on the gRNA sequence and can impact the design of specific oligonucleotides to be synthesized and either used as a template for RNA production or cloned into an expression vector (see, e.g., Mohr et al. (2016), ibid.).VI. Oligonucleotide synthesis and purification methods

[0165] The disclosed sequencing adapter compositions may be synthesized by any of a variety of oligonucleotide synthesis techniques known to those of skill in the art. DNA and / or RNA oligonucleotides (including modified DNA and / or RNA oligonucleotide sequences) may be synthesized, for example, using solid-phase synthesis techniques and phosphoramidite-based methods comprising the steps of deprotection, coupling, sulfurization / oxidation, and capping to extend an oligonucleotide sequence tethered to a solid support in the 3’ to 5’ direction (see, e.g., Catani et al. (2020), “Oligonucleotides: Current Trends and Innovative Applications in the Synthesis, Characterization, and Purification”, Biotechnol. J. 15:1900226, and Beaucage (2008), “Solid-Phase Synthesis of siRNA Oligonucleotides”, Curr. Opinion in Drug Discovery & Develop. 11(2):203-216). Once synthesis is complete and the oligonucleotide sample has beencleaved from the solid support, it may be subjected to a variety of purification and / or analytical characterization techniques including, but not limited to, ion pair reversed-phase liquid chromatography, ion exchange liquid chromatography, hydrophilic interaction liquid chromatography, mixed-mode chromatography, mass spectrometry, preparative ion exchange liquid chromatography, preparative reversed-phase liquid chromatography, preparative ion pair reversed-phase liquid chromatography, etc.VII. Ligases and ligation methods

[0166] Ligation of the disclosed adapter sequence compositions to single-stranded template nucleic acid molecules may be performed using any of a variety of techniques known to those of skill in the art. In some instances, for example, enzymatic ligation of oligonucleotide sequences may be performed (e.g., using any of the ligases disclosed herein). In some instances, chemical ligation of oligonucleotide sequences may be performed (e.g., using modified oligonucleotides and Click ligation based on copper-catalyzed azide - alkyne (CuAAC) reaction (see, e.g., El- Sagheer et al. (2012), “Click Nucleic Acid Ligation: Applications in Biology and Nanotechnology”, Accounts of Chemical Research 45(8): 1258-1267).VIII. Nucleic acid amplification

[0167] In some embodiments, methods provided herein comprise amplification of DNA (e.g. cDNA derived from single-stranded RNA). A variety of amplification procedures are available, selection of which may depend on factors such as the type of sequencing platform to be used. Examples of amplification reactions include thermal cycling reactions and isothermal reactions. Typically, amplification reactions comprise primer extension reactions. General methods for primer-directed amplification of target polynucleotides are known in the art, and include without limitation, methods based on the polymerase chain reaction (PCR). Conditions that are generally favorable to the amplification of target sequences by PCR have been characterized, can be optimized at a variety of steps in the process, and may depend on characteristics of elements in the reaction, such as target type, target concentration, sequence length to be amplified, sequence of the target and / or one or more primers, primer length, primer concentration, polymerase used, reaction volume, ratio of one or more elements to one or more other elements, and others, some or all of which can be altered. In general, PCR involves the steps of denaturation of the target to be amplified (e.g., cDNA) (if double-stranded), hybridization of one or more primers to the target template, and extension of the primers by a DNA polymerase. The steps can be repeated (or“cycled”) in order to further amplify the target sequence. Steps in this process can be optimized for various outcomes, such as to enhance yield, decrease the formation of spurious products, and / or increase or decrease specificity of primer annealing. Example methods of optimization include, but are not limited to, adjustments to the type or amount of elements in the amplification reaction and / or to the conditions of a given step in the process, such as temperature at a particular step, duration of a particular step, and / or number of cycles. In some embodiments, an amplification reaction comprises a single primer extension step. In some embodiments, an amplification reaction comprises at least 5, 10, 15, 20, 25, 30, 35, 50, or more cycles. In some embodiments, an amplification reaction comprises no more than 5, 10, 15, 20, 25, 35, 50, or more cycles. Cycles can contain any suitable number of steps, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more steps. Steps can comprise any temperature or gradient of temperatures, suitable for achieving the purpose of the given step, including but not limited to, strand denaturation, primer annealing, and primer extension. Steps can be of any suitable duration, including but not limited to about, less than about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 180, 240, 300, 360, 420, 480, 540, 600, or more seconds, including indefinitely until manually interrupted.

[0168] DNA can also be amplified isothermally. SPIA is an example of a linear, isothermal amplification method, an example description of which can be found in U.S. Patent No. 6,251,639, which is incorporated herein by reference. Generally, the method includes hybridizing chimeric RNA / DNA amplification primers to the probes or target. The DNA portion of the probe is 3' to the RNA. Following hybridization of the primer to the template, the primer is extended with DNA polymerase. Subsequently, the RNA is cleaved from the composite primer with an enzyme that cleaves RNA from an RNA / DNA hybrid. Subsequently, an additional RNA / DNA chimeric primer is hybridized to the template such that the first extended primer is displaced from the target probe. The extension reaction is repeated, whereby multiple copies of the probe sequence are generated.

[0169] Oligonucleotide primers utilized in DNA amplification reactions are referred to herein generally as “amplification primers” or simply “primers.” In some embodiments, all of the primers in an amplification reaction have the same sequence, such that there is only one type of primer participating in the reaction. In some embodiments, particularly in the case of exponential amplification, the amplification reaction comprises one or more pairs of primers, wherein in one primer of the pair hybridizes to an is extended along an initial template and the second primer hybridizes to and is extended along the complementary strand of the initial template and / or theextension product of the first primer in the pair. In some embodiments, a primer is about or at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 100, 150, or more nucleotides in length, or a length between any of these. In some embodiments, a primer is less than about 150, 100, 90, 80, 70, 60, 50, 40, 30, 20, 15, or fewer nucleotides in length. In some embodiments, a primer is between 5-100, between 10-75, or between 15-50 nucleotides in length. A primer hybridizes to a DNA template to be amplified via base complementarity between the template DNA and a complementary sequence in the primer. In some embodiments, the amplification reaction comprises one or more primers in which the complementary sequence is pre-defined, such as when a particular target (e.g. one or more genes) or a particular class of targets (e.g. DNA and / or cDNA comprising defined sequence elements, such as sequence elements contained in an adaptor of the present disclosure, or reverse transcription primer) is desired to be amplified. In some embodiments, a plurality of different primers, each having a different predefined complementary sequence, are present in a single reaction. For example, an amplification reaction can comprise about or at least about 2, 5, 10, 15, 20, 25, 30, 40, 50, 75, 100, 150, 200, 300, 400, 500, or more different primers.IX. Sequencing methods

[0170] A variety of suitable sequencing techniques are available. In some embodiments, sequencing comprises massively parallel sequencing of about, or at least about 10000, 100000, 500000, 1000000, or more DNA molecules using a high-throughput sequencing by synthesis process, such as Illumina’s sequencing-by- synthesis and reversible terminator-based sequencing chemistry (e.g. as described in Bentley el al., Nature 6:53-59 (2009)). In some embodiments, particularly when cfDNA is included among the polynucleotides to be sequenced, DNA is not fragmented prior to sequencing. Typically, Illumina’s sequencing process comprises attachment of template DNA to a planar, optically transparent surface on which oligonucleotide anchors are bound. Template DNA is end- repaired to generate 5 ’-phosphorylated blunt ends, and the polymerase activity of Klenow fragment is used to add a single A base to the 3 ‘end of the blunt phosphorylated DNA. This addition prepares the DNA for ligation to oligonucleotide adapters, which optionally have an overhang of a single T base at their 3’ end to increase ligation efficiency. The adapter oligonucleotides are complementary to the flow-cell anchor oligos. Under limitingdilution conditions, adapter-modified, single-stranded template DNA is added to the flow cell and immobilized by hybridization to the anchor oligos. Attached DNA fragments are extended and bridge amplified to create an ultra-high density sequencing flow cell with hundreds of millions ofclusters, each containing about 1,000 copies of the same template. In one embodiment, the template DNA is amplified using PCR before it is subjected to cluster amplification, such as in a process described above. In some applications, the templates are sequenced using a robust four- color DNA sequencing-by-synthesis technology that employs reversible terminators with removable fluorescent dyes. High-sensitivity fluorescence detection is achieved using laser excitation and total internal reflection optics. Short sequence reads of about tens to a few hundred base pairs are aligned against a reference genome, and unique mapping of the short sequence reads to the reference genome are identified using specially developed data analysis pipeline software. After completion of the first read, the templates can be regenerated in situ to enable a second read from the opposite end of the fragments. Thus, either single-end or paired-end sequencing of the DNA fragments can be used.

[0171] Another non-limiting example sequencing process is the single molecule sequencing technology of the Helicos True Single Molecule Sequencing (tSMS) technology (e.g. as described in Harris T. D. et al., Science 320: 106-109 (2008)). In a typical tSMS process, a DNA sample is cleaved into, or otherwise provided as strands of approximately 100 to 200 nucleotides, and a poly A sequence is added to the 3’ end of each DNA strand. Each strand is labeled by the addition of a fluorescently labeled adenosine nucleotide. The DNA strands are then hybridized to a flow cell, which contains millions of oligo-T capture sites that are immobilized to the flow cell surface. In some embodiments, the templates are at a density of about 100 million templates / cm2. The flow cell is then loaded into an instrument, e.g., HeliScope™ sequencer, and a laser illuminates the surface of the flow cell, revealing the position of each template. A CCD camera can map the position of the templates on the flow cell surface. The template fluorescent label is then cleaved and washed away. The sequencing reaction begins by introducing a DNA polymerase and a fluorescently labeled nucleotide. The oligo-T nucleic acid serves as a primer. The polymerase incorporates the labeled nucleotides to the primer in a template directed manner. The polymerase and unincorporated nucleotides are removed. The templates that have directed incorporation of the fluorescently labeled nucleotide are discerned by imaging the flow cell surface. After imaging, a cleavage step removes the fluorescent label, and the process is repeated with other fluorescently labeled nucleotides until the desired read length is achieved. Sequence information is collected with each nucleotide addition step. Whole genome sequencing by single molecule sequencing technologies excludes or typically obviates PCR-based amplification in the preparation of the sequencing libraries.

[0172] Another illustrative, but non-limiting example sequencing process is pyrosequencing, such as in the 454 sequencing platform (Roche) (e.g. as described in Margulies, M. el al. Nature 437:376-380 (2005)). 454 sequencing typically involves two steps. In the first step, DNA is sheared into fragments of, or otherwise provided (e.g. as naturally occurring cfDNA molecules, or cDNA from naturally short RNA molecules) as DNA having sizes of approximately 300-800 base pairs, and the polynucleotides are blunt-ended. Oligonucleotide adapters are then ligated to the ends of the DNA. The adapters serve as primers for amplification and sequencing of the DNA. The DNA can be attached to capture beads, e.g., streptavidin-coated beads using, e.g., adapter B, which contains 5 ’-biotin tag. The DNA attached to the beads are PCR amplified within droplets of an oil-water emulsion. The result is multiple copies of clonally amplified DNA molecules on each bead. In the second step, the beads are captured in wells (e.g., picoliter- sized wells).

[0173] Pyrosequencing is performed on each DNA molecule in parallel. Addition of one or more nucleotides generates a light signal that is recorded by a CCD camera in a sequencing instrument. The signal strength is proportional to the number of nucleotides incorporated. Pyro sequencing makes use of pyrophosphate (Ppi) which is released upon nucleotide addition. Ppi is converted to ATP by ATP sulfurylase in the presence of adenosine 5’ phosphosulfate. Luciferase uses ATP to convert luciferin to oxyluciferin, and this reaction generates light that is measured and analyzed.

[0174] Further high-throughput sequencing processes are available. Non-limiting examples include sequencing by ligation technologies (e.g., SOLiD™ sequencing of Applied Biosy stems), single-molecule real-time sequencing (e.g., Pacific Biosciences sequencing platforms utilizing zero-mode wave detectors), nanopore sequencing (e.g., as described in Soni G V and Meller A. Clin Chem 53: 1996-2001 (2007)), sequencing using a chemical-sensitive field effect transistor (e.g., as described in U.S. Patent Application Publication No. 20090026082), sequencing platforms by Ion Torrent (pairing semiconductor technology with sequencing chemistry to directly translate chemically encoded information (A, C, G, T) into digital information (0, 1) on a semiconductor chip), and sequencing by hybridization.

[0175] Additional illustrative details regarding sequencing technologies can be found in, e.g., U.S.Patent Application Publication No. 2016 / 0319345, incorporated herein by reference.X. Sequence reads and sequence read alignment / assembly

[0176] The output of the base-calling process performed by NGS sequencing platforms consists of a plurality of sequence reads, e.g., the nucleotide sequences determined for all or a portion of a template nucleic acid molecule (e.g., a synthetic oligonucleotide).

[0177] In some instances, the sequence reads may comprise sequence reads of at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides or base pairs of the template nucleic acid sequences. In some instances, the sequence reads generated using the disclosed methods may comprise sequence reads of at least about 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, or more than 400 nucleotides or base pairs of the template nucleic acid sequences.

[0178] In some instances, the sequencing process (e.g., a sequencing-by- synthesis (SBS) process) may generate at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, or more than 1,000 sequencing reads per run. In some instances, the sequencing process may generate at least about 1,000, 1,500, 2,000, 2,500, 3,000, 3,500, 4,000, 4,500, 5,000, 5,500, 6,000, 6,500, 7,000, 7,500, 8,000, 8,500, 9,000, 9,500, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 106, 5 x 106, 107, or more than 107sequencing reads per run.

[0179] In some instances, the disclosed methods may comprise alignment of sequence reads and / or assembled sequences to a known reference sequence (e.g., a known target sequence for a synthetic oligonucleotide molecules). In some instances, the disclosed sequencing methods may comprise alignment of sequence reads and / or assembled sequences to a known reference sequence or consensus sequence (e.g., the GRCh38 human reference genome (Genome Reference Consortium)) from the same or a similar organism for naturally-occurring nucleic acid sequences. Alignment to a reference sequence or consensus sequence may be used to identify gaps, errors, or variants in the sequence reads and / or assembled sequences.

[0180] Any of a variety of alignment software tools known to those of skill in the art may be used to align sequence reads to a reference sequence or consensus sequence. Examples include, but are not limited to, Bowtie, BWA, mr- and mrsFAST, Novoalign, SHRiMP and SOAPv2 (see, e.g., Ruffalo et al. (2011), “Comparative Analysis of Algorithms for Next-Generation Sequencing Read Alignment”, Bioinformatics 27(20):2790-2796).

[0181] Any of a variety of bioinformatics software programs known to those of skill in the art may be used to assemble longer sequences from relatively short sequence reads. Examples include, but are not limited to, DBG2OLC (see, e.g., Ye et al. (2016), “DBG2OLC: Efficient Assembly of Large Genomes Using Long Erroneous Reads of the Third Generation Sequencing Technologies”, Scientific Reports 6:31900), SPAdes (see, e.g., Bankevich et al. (2012), “SPAdes: A New Genome Assembly Algorithm and Its Applications to Single-Cell Sequencing”, J. Computational Biol. 19(5):455-477), Sparse Assembler (see, e.g., Ye et al. (2012), “Exploiting Sparseness in de novo Genome Assembly”, BMC Bioinformatics 13(Suppl 6):S1), Fermi (see, e.g., Li (2012), “Exploring Single-Sample SNP and INDEL Calling with Whole-Genome de novo Assembly”, Bioinformatics 28(14):1838-1844), and String Graph Assembler (SGA) (see, e.g., Simpson et al. (2012), “Efficient de novo Assembly of Large Genomes Using Compressed Data Structures”, Genome Res. 22: 549-556).XL Use of UMIs for amplification and sequencing error correction

[0182] In some embodiments using UMIs, multiple sequence reads having the same UMI(s) are collapsed to obtain one or more consensus sequences, which are then used to determine the sequence of a source cDNA molecule. Multiple distinct reads may be generated from distinct instances of the same source cDNA molecule, and these reads may be compared to produce a consensus sequence. The instances may be generated by amplifying a source cDNA molecule prior to sequencing, such that distinct sequencing operations are performed on distinct amplification products, each sharing the source cDNA molecule's sequence. Of course, amplification may introduce errors such that the sequences of the distinct amplification products have differences. In the context some sequencing technologies such as an embodiment of Illumina's sequencing-by-synthesis, a source cDNA molecule or an amplification product thereof forms a cluster of DNA molecules linked to a region of a flow cell. The molecules of the cluster collectively provide a read. Typically, at least two reads are required to provide a consensus sequence. Sequencing depths of 100, 1000, and 10,000 are examples of sequencing depths useful in the disclosed embodiments for creating consensus reads for low frequencies e.g., about 1% or less). In some embodiments, nucleotides that are consistent across 100% of the reads sharing a UMI or combination of UMIs are included in the consensus sequence. In some embodiments, consensus criterion can be lower than 100%. For instance, a 90% consensus criterion may be used, which means that base pairs that exist in 90% or more of the reads in the group are included in the consensus sequence. In some embodiments, the consensus criterion may be set at about, or morethan about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, or about 100%.

[0183] In addition to amplification errors, sequencing errors can also introduce false positives for sequence variant. For instance, sequencing reads sharing the same UMI might have different sequences because of sequencing errors such as a false base call. By grouping sequencing reads sharing the same UMI together and obtaining a consensus sequence, the methods disclosed herein can also correct for sequencing errors.EXAMPLES

[0184] The following examples are included for illustrative purposes only and are not intended to limit the scope of the present disclosure.Example 1 - General Approach

[0185] With the approach described herein, NGS libraries were made from synthetic singlestranded guide RNA (gRNA) oligonucleotide molecules comprising RNA nucleotides or a mixture of DNA and RNA nucleotides.

[0186] NGS libraries were generated from manufactured gRNA samples through a series of sequence-agnostic biochemical reactions which ligate adapters onto the 5’ and 3’ ends of gRNAs and then convert adapter-ligated gRNA molecules into double-stranded DNA compatible with commercial sequencing platforms (e.g., Illumina sequencing platforms). The 3’ ends of gRNA molecules were first ligated to custom pre-adenylated single-stranded DNA adapters using RNA ligase buffer and RNA ligase enzyme (e.g., from the NEBNext Small RNA Library Prep Set for Illumina (NEB)). A buffer exchange was then performed using beads (e.g., RNAClean XP beads (Beckman-Coulter)), which bind nucleic acid strands and allow removal of buffers and enzymes. The 5’ ends of gRNA oligonucleotide substrates were then phosphorylated using polynucleotide kinase (NEB), and another buffer exchange was performed (e.g., with RNAClean XP beads). The phosphorylated 5’ ends of gRNA substrate molecules were then ligated to custom single-stranded RNA adapters, for instance using RNA ligase buffer and enzyme from the NEBNext Small RNA Library Prep Set for Illumina (NEB). Adapters used for 5’ ligation contained unique molecular identifiers (UMIs) made up of 15 randomized nucleotides which were used for single-molecule quantification and sequencing error correction in downstream analysis. Adapter-ligated gRNA molecules were then reverse transcribed into cDNA (e.g., using AMV Reverse Transcriptase(NEB)). The cDNA was purified (e.g., using SPRIselect beads (Beckman Coulter)), diluted 1:1000, and used as a template for PCR using primers that incorporate NGS adapters (e.g., commercially available adapters, such as P5 and P7 Illumina NGS adapters) to each molecule. NGS libraries were purified using beads (e.g., SPRIselect beads (Beckman Coulter)) and quantified (e.g., using the KAPA Library Quantification Kit qPCR assay (KAPA)) before loading onto a sequencer (e.g., the Illumina MiSeq) for paired-end sequencing. Sequencing reads were demultiplexed, e.g., using MiSeq Reporter (Illumina).Data analytics

[0187] To quantify the frequency of intended (fully correct) and variant gRNA molecules in each sample, demultiplexed sequencing reads in paired FASTQ files were first trimmed to remove the NGS adapter sequences (e.g., Illumina adapter sequences), for instance, using Trim Galore. The 5’ UMI in each read was extracted and moved to the read name, for instance, using UMI- tools Extract. Paired read 1 and read 2 files were merged using FLASh and read pairs containing base call disagreement between read 1 and read 2 were discarded. Reads that had identical 15 nucleotide UMI sequences and had UMI base quality scores >Q30 were grouped together. One consensus sequence representative of a single original gRNA molecule was generated from each UMI group containing 5 or more reads by collapsing grouped reads to the most frequent base call at each nucleotide position. Ambiguous consensus sequences with greater than 20% of contributing reads not matching the consensus call were discarded. Consensus reads were aligned to a reference gRNA sequence (the intended sequence) and the frequency and descriptive identity of each unique variant sequence (based on deviation from the reference sequence) in the population was determined and returned in data output files.

[0188] For the purposes of data analysis, all sequencing reads with the same UMI were considered to be derived from the same original substrate molecule due to the extremely low probability of more than one substrate molecule receiving the same UMI sequence. Consequently, each distinct UMI observed in the sequencing data was considered to represent a single original substrate molecule, enabling single-molecule quantification and the formation of consensus sequences to reduce PCR and sequencing error. Furthermore, molecular quantification based on UMI counts can greatly reduce the impact of stochastic variability in PCR amplification and molecular bias in amplification efficiency, thereby making final quantified data more representative of the original molecular population (illustrated in FIG. 6 and Table 1).Ligation adapter design

[0189] The sequence, secondary structure, and base modifications of ligation adapters used in sequencing library preparation are critical design elements that determine reaction efficiency, preference for different molecular substrates, and the structure of resultant sequencing reads. These design elements can substantially impact the overall performance of the library preparation method and the ability to generate representative and quantitative sequence data for oligonucleotide substrates. The following sections detail key design elements for the 3’ and 5’ ligation adapters utilized herein.3' adapter design

[0190] In one example, the composition of the sequencing adapter ligated to the 3’ end of the substrate molecule was a single-stranded DNA oligonucleotide 38 bases in length modified with an adenyl group on its 5’ side and an amino group on its 3’ side (schematic shown in FIG. 1). In some embodiments, a 3’ dideoxy-C modification, 3’ inverted dT modification, 3’ C3 spacer modification, or 3’ phosphoryl modification can be used as alternative options to block ligation to the 3’ end of the adapter.

[0191] The presence of a 5 ’-adenylated nucleotide on the 5’ terminus of the adapter improves efficiency as a ligation donor in the absence of ATP and limits the formation of side products during the ligation reaction. The presence of a 3 ’-aminated nucleotide on the 3’ terminus of the adapter limits side product formation by preventing concatenation of adapters during the ligation reaction. In some instances, the entire 3’ adapter sequence (e.g., a 38-nucleotide 3’ adapter sequence as described above) can serve as a binding site for a reverse transcription primer. In some instances, a portion of the 3’ adapter sequence may serve as a binding site for a reverse transcription primer. Within the 38 nucleotide sequence, the 28 nucleotides closest to the 3’ end of the adapter also serve as a binding site for an i7 Illumina indexing PCR primer and the 10 nucleotides closest to the 5’ end of the adapter also serve as a spacer sequence situated between the Illumina read 2 sequencing primer binding site and the insert sequence. Using the 3’ adapter as the reverse transcription primer binding site allows for the generation of cDNA without altering the original sequence of the template.

[0192] The inclusion of a spacer sequence in the 3’ adapter ensures that the same 10 nucleotide sequence will immediately precede the 3’ terminus of the insert sequence in each sequencing read.This facilitates quality control in data analysis, helps mitigate the impact of typically lower quality in the first few base calls, and allows for unambiguous determination of the termini of the original substrate molecule. In some embodiments, the 3’ spacer sequence is designed such that the length of the spacer sequence plus the full sequence of an expected insert are fully covered by sequencing read 2 to allow the potential for both read 1 and read 2 to cover the full insert sequence, thereby enabling an analytical method for correcting some sequencing errors by discarding read pairs with discrepancies (shown in Table 1).Table 1. Comparison of sequence quantification using different analytical steps.

[0193] Table 1 shows a comparison of key data output metrics for gRNA variant frequency quantification performed with and without UMI consensus sequence formation and deduplication. All analyses were performed using the same dataset, following the analytical steps described in each column heading. In the first column, only read trimming and quality filtering was performed. In the second column, an additional step of paired read 1 and read 2 merging was performed, which requires complete agreement on insert base calls between read 1 and read 2 and removal of all discrepant read pairs. In the third column, an additional step of UMI consensus sequence collapsing was performed. The observed frequency of intended sequence increases and the observed frequency of mismatch errors decreased as error correction was performed through read pair merging and UMI consensus sequence formation.

[0194] The sequence of the 3’ adapter was designed for low internal secondary structure and minimal complementarity to the 5’ adapter to avoid molecular interactions that are unfavorable to the ligation reaction.5’ adapter design

[0195] The composition of the sequencing adapter ligated to the 5’ end of the substrate molecule was a single-stranded RNA oligonucleotide 48 bases in length with hydroxyl groups on its 5’ and 3’ sides (schematic shown in FIG. 2).

[0196] The presence of a 3 ’-hydroxylated nucleotide on the 3’ terminus of the adapter allows it to act as a ligation acceptor when paired with a 5 ’-phosphorylated ligation donor substrate. The presence of a 5 ’-hydroxylated nucleotide on the 5’ terminus of the adapter prevents it from acting as a ligation donor, thereby avoiding concatenation of adapters during the ligation reaction. In some embodiments, a 5’ dideoxy-C modification, 5’ inverted dideoxy-T modification, or 5’ C3 spacer modification can be used as alternative options to block ligation to the 5’ end of the adapter. In some embodiments, a 3’ phosphate group on the 5’ adapter can be used to ligate to a 5’ hydroxylated substrate molecule using RtcB Ligase (New England Biolabs).

[0197] The 5’ adapter sequence included a 26-nucleotide binding site for the i5 Illumina indexing PCR primer, a 15-nucleotide string of randomized nucleotides serving as a unique molecular identifier (UMI), and a 7-nucleotide spacer sequence between the end of the UMI and the first nucleotide of the insert sequence. The string of randomized nucleotides is critical for reducing molecular bias caused by structural compatibility between ligation donor (insert molecule) and ligation acceptor (adapter) molecules. Furthermore, the 15-nucleotide string of randomized nucleotides has over one billion possible permutations, allowing for each original insert molecule to receive a unique nucleotide string that can be used as a UMI for downstream data processing. UMI labeling of original insert molecules allows for PCR and sequencing error correction and normalization of read quantification agnostic to differential amplification.

[0198] The inclusion of a spacer sequence in the 5’ adapter ensures that the same 7 nucleotide sequence will immediately precede the 5’ terminus of the insert sequence in each sequencing read. This facilitates quality control in data analysis, helps mitigate the impact of typically lower quality in the first few base calls, and allows for unambiguous determination of the termini of the original substrate molecule. In some embodiments, the UMI sequence and 5’ spacer sequence are designed such that the length of the UMI, spacer sequence, and the full sequence of an expected insert are fully covered by sequencing read 1 to allow the potential for both read 1 and read 2 to cover the full insert sequence, thereby enabling an analytical method for correcting some sequencing errors by discarding read pairs with discrepancies.Example 2 - Assessment of a Cas9 gRNA sample

[0199] This example demonstrates the application of methods in the present invention to the preparation of Illumina Next-Generation Sequencing (NGS) libraries from exemplary Cas9 gRNA samples designed as lOOmer RNA oligonucleotides which included phosphorothioate linkages and 2’ -Oxygen methylation modifications on the four nucleotides on the 5’ terminus and 4 nucleotides on the 3’ terminus. The following experiment was performed to produce representative sequencing libraries from each gRNA sample.

[0200] Combination and denaturation of the gRNA sample and the 3’ adapter was performed as follows: 2.17 pL of 3.07 pM gRNA sample and 1.33 pL of 7.5 pM 3’ adapter (SEQ ID NO: 1) were combined in a well of a 96-well plate, the plate was incubated at 70°C in a thermal cycler for 2 minutes, and the plate was then immediately transferred to a 4°C cold block for at least 5 minutes.

[0201] SEQ ID NO: 1: 5’- / 5rApp / CGAGTGTCATAGATCGGAAGAGCACACGTCTGAACTCC / 3AmMO / -3’

[0202] Ligation of the 3’ adapter to the gRNA substrate was performed as follows: 5 pL of 2x 3’ Ligation Reaction Buffer (New England Biolabs) and 1.5 pL of 3’ Ligation Enzyme Mix (New England Biolabs) were added to the plate well containing the denatured gRNA substrate and 3’ adapter, the sample was mixed by pipetting, and the plate was incubated at 16°C for 18 hours.

[0203] The 3’ adapter ligation reaction products were purified to perform a buffer exchange as follows: 12 pL of RNAClean XP bead solution (Beckman-Coulter) were added to the 3’ ligation reaction and mixed by pipetting, the mixture was incubated at room temperature for 5 minutes, the plate was transferred to a magnetic rack for 5 minutes to collect beads on the side of the well, the supernatant was removed, beads were washed 3 times with 70% ethanol (Fisher) (volume / volume in nuclease-free water (Ambion)), beads were air-dried at room temperature for 5 minutes to remove residual ethanol, beads were resuspended in 11 pL of nuclease-free water, the plate was incubated at room temperature off of the magnet for 2 minutes, the plate was transferred back to the magnet for 3 minutes to collect beads on the side of the well, and finally 9.75 pL of supernatant containing the eluted ligation product was transferred to a new well.

[0204] Phosphorylation of the 5’ terminus of the gRNA substrate was performed as follows: 9.75 pL of bead-cleaned 3’ ligation product, 1.25 pL of lOx T4 PNK Reaction Buffer (New England Biolabs), 1.25 pL lOmM ATP (New England Biolabs), and 0.25 pL T4 Polynucleotide Kinase(New England Biolabs) were combined in a well of a 96-well plate, and the plate was incubated at 37 °C for 30 minutes in a thermal cycler.

[0205] The 5’ phosphorylation reaction products were purified to perform a buffer exchange as follows: 15 pL of RNAClean XP bead solution (Beckman-Coulter) were added to the 5’ phosphorylation reaction and mixed by pipetting, the mixture was incubated at room temperature for 5 minutes, the plate was transferred to a magnetic rack for 5 minutes to collect beads on the side of the well, the supernatant was removed, beads were washed 3 times with 70% ethanol (volume / volume in nuclease-free water), beads were air-dried at room temperature for 5 minutes to remove residual ethanol, beads were resuspended in 8.2 pL of nuclease-free water (Ambion), the plate was incubated at room temperature off of the magnet for 2 minutes, the plate was transferred back to the magnet for 3 minutes to collect beads on the side of the well, and finally 6.92 pL of supernatant containing the eluted ligation product was transferred to a new well.

[0206] Combination and denaturation of the gRNA sample and the 5’ adapter was performed as follows: 6.92 pL of bead-cleaned gRNA 5’ phosphorylation product and 1.33 pL of 11.25 pM 5’ adapter (e.g. SEQ ID NO: 2) were combined in a well of a 96-well plate, the plate was incubated at 70°C in a thermal cycler for 2 minutes, and the plate was then immediately transferred to a 4°C cold block for at least 5 minutes.

[0207] SEQ ID NO: 2: 5’- rGrUrUrCrArGrArGrUrUrCrUrArCrArGrUrCrCrGrArCrGrArUrCrNrNrNrNrNrNrNrNrNrNrN rNrNrNrNr ArCrGr ArUr ArC-3 ’

[0208] Ligation of the 5’ adapter to the gRNA substrate was performed as follows: 5 pL of 2x 3’ Ligation Reaction Buffer (New England Biolabs), 0.5 pL of lOx 5’ Ligation Reaction Buffer, and 1.25 pL of 5’ Ligation Enzyme Mix (New England Biolabs) were added to the well containing the denatured gRNA substrate and 5’ adapter, the sample was mixed by pipetting, and the plate was incubated at 25 °C for 90 minutes.

[0209] Reverse transcription of the adapter-ligated gRNA substrate to generate cDNA was performed as follows: 7 pL of 5’ adapter ligation reaction, 0.5 pL of 10 pM RT primer (e.g. SEQ ID NO: 3), 4 pL of 5x First Strand Synthesis Reaction Buffer (New England Biolabs), 0.5 pL Murine RNase Inhibitor (New England Biolabs), 2 pL lOx AMV RT Buffer (New England Biolabs), 1 pL AMV Reverse Transcriptase (New England Biolabs), and 5 pL nuclease-free water(Ambion) were combined in a well of a 96-well plate, and the plate was incubated at 50°C for 30 minutes.

[0210] SEQ ID NO: 3: 5’- GGAGTTCAGACGTGTGCTCTTCCGATCTATGACACTCG-3’

[0211] The reverse transcription reaction products were purified to perform a buffer exchange as follows: 24 pL of SPRIselect bead solution (Beckman-Coulter) was added to the reverse transcription reaction and mixed by pipetting, the mixture was incubated at room temperature for 5 minutes, the plate was transferred to a magnetic rack for 5 minutes to collect beads on the side of the well, the supernatant was removed, beads were washed 2 times with 80% ethanol (volume / volume in nuclease-free water), beads were air-dried at room temperature for 5 minutes to remove residual ethanol, beads were resuspended in 25 pL of nuclease-free water (Ambion), the plate was incubated at room temperature off of the magnet for 2 minutes, the plate was transferred back to the magnet for 3 minutes to collect beads on the side of the well, and finally 22 pL of supernatant containing the eluted cDNA product was transferred to a new well.

[0212] Purified cDNA was serially diluted to a final dilution of 1:500 as follows: 5 pL of cDNA to 95 pL of nuclease-free water (Ambion) and then mixed by pipetting to create a 1:20 dilution, 5 pL of 1:20 diluted cDNA was added to 120 pL nuclease-free water (Ambion) and mixed by pipetting.

[0213] PCR amplification and Illumina indexing of cDNA derived from the gRNA substrate was performed as follows: 10 pL of 1:500 diluted cDNA, 12.5 pL of 2x Q5 Hot Start Master Mix (New England Biolabs), 1.25 pL of 10 pM i5 Universal Primer (e.g. SEQ ID NO: 4), and 1.25 pL of 10 pM i7 Indexing Primer 1 (SEQ ID NO: 5) were combined in a well of a 96-well plate and mixed by pipetting. The Ns subsequence of the indexing primer sequence indicates a variable but known sequence that can be used to demultiplex sequencing data from multiplexed sequencing runs. The reaction was carried out in a thermal cycler using the following cycle program:

[0214] 1. 98°C for 30 seconds

[0215] 2. 98°C for 10 seconds

[0216] 3. 69°C for 20 seconds

[0217] 4. 72°C for 20 seconds

[0218] 5. Repeat steps 2-4 22 times

[0219] 6. 72°C for 2 minutes

[0220] 7. Hold at 4°C

[0221] SEQ ID NO: 4: 5’-AATGATACGGCGACCACCGAGATCTACACGTTCAGAGTTCTACAGTCCGACGATC-3’

[0222] SEQ ID NO: 5: 5’-CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTGCT CTTCCGATCT-3’

[0223] The PCR products were purified and size selected to remove residual primers and adapterdimer side products as follows: 30 pL (1.2x reaction volumes) of SPRI select bead solution (Beckman-Coulter) was added to the reverse transcription reaction and mixed by pipetting, the mixture was incubated at room temperature for 5 minutes, the plate was transferred to a magnetic rack for 5 minutes to collect beads on the side of the well, the supernatant was removed, beads were washed 2 times with 80% ethanol (volume / volume in nuclease-free water), beads were airdried at room temperature for 5 minutes to remove residual ethanol, beads were resuspended in 50 pL of nuclease-free water (Ambion), the plate was incubated at room temperature off of the magnet for 2 minutes, the plate was transferred back to the magnet for 3 minutes to collect beads on the side of the well, and finally 45 pL of supernatant containing the eluted amplicon library was transferred to a new sample tube.

[0224] The purified amplicon library concentration and average amplicon size were assessed by TapeStation analysis as follows: 1 pL of purified PCR product was combined with 3 pL of TapeStation D1000 Loading Dye (Agilent) in a well of an 8-well PCR strip tube, 1 pL of TapeStation DI 000 Ladder (Agilent) was combined with 3 pL of DI 000 Loading Dye (Agilent) in a separate well of the same 8-well PCR strip tube, the contents of the strip tube were mixed by vortexing, the strip tube was placed in an Agilent TapeStation 4200 along with a TapeStation D1000 Screen Tape (Agilent), and the default D1000 program was performed. The TapeStation results were assessed to calculate library concentration and assess average library size.

[0225] FIG. 5 provides a non-limiting example of a representative TapeStation D1000 electropherogram trace for a purified NGS amplicon library prepared using the methods describedherein for a 100-nucleotide Cas9 gRNA. The peaks labeled “Lower” and “Upper” are the lower and upper size markers, respectively, used for sample size determination with TapeStation D1000 reagents. The observed average size of the final sequencing library was 268 bp, which is within a reasonable range of the expected size of 253 bp for fully indexed library molecules with a 100 bp insert.

[0226] The amplicon library was denatured and diluted for sequencing on the Illumina MiSeq instrument as follows: the purified library DNA was diluted to 4 nM in nuclease-free water (Ambion) using the concentration calculated from TapeStation analysis in the previous step, 5 pL of the 4 nM library was combined with 5 pL of 0.2 N Sodium hydroxide (Sigma-Aldrich) in a 1.5mL DNA LoBind tube (Eppendorf) and mixed by pipetting, the tube was incubated at room temperature for 5 minutes, the library was diluted to 20 pM by adding 990 pL of pre-chilled HT1 buffer (Illumina), the tube was mixed by vortexing and placed on ice, the denatured library was diluted to 13 pM by combining 390 pL of 20 pM and 210 pL of pre-chilled HT1 buffer (Illumina). PhiX control library DNA was spiked in at a 5% concentration by combining 30 pL of 12.5 pM denatured PhiX library (Illumina) with 570 pL of 13 pM gRNA library and mixing by pipetting.

[0227] The library was heat denatured as follows: the tube containing the final combined sequencing library was incubated at 96°C for 2 minutes, the tube was vortexed to mix, the tube was spun down in a microcentrifuge to collect liquid at the bottom of the tube, and the tube was immediately placed on ice for at least 5 minutes.

[0228] The final denatured library was loaded into a MiSeq v2 Reagent Kit 300-cycle (Illumina) and a paired-end 2 x 151 sequencing run was performed according to a run sample sheet provided by the user, which includes indices that identify the sample. Accordingly, multiple samples can be run simultaneously.Tables 2-5 provide summaries of the reagents, oligonucleotides, consumable materials, and instruments used to perform library preparation and sequencing.Table 2. Reagents.Table 3. Oligonucleotides.Table 4. Consumables.Table 5. Instruments.

[0229] MiSeq sequencing data was analyzed as follows: sequencing reads were demultiplexed using the indices provided in the run sample sheet. The FASTQ files were processed to trim the Illumina adapter sequences using the Trim Galore program, paired reads containing intact 5’ and 3’ adapter sequences were identified and retained using the UMI- tools Extract program following a regular expression (UMI pattern 1: “A(?P<umi_l>[ATGC]{ 15})ACGATAC”; UMI pattern 2: “AATGACACTCG”) and the 15 nucleotide UMI in each read was extracted and moved to the read name using the UMI-tools Extract program. Paired read 1 and read 2 FASTQ files were merged using the FLASh merge program and read pairs containing base call disagreement between read 1 and read 2 were discarded. Using custom software, reads that shared identical 15 nucleotide UMI sequences and had UMI base quality scores >Q30 were grouped together using a metadata label. Each UMI group containing 5 or more sequencing reads was collapsed into a single consensus sequence using the most frequent base call at each nucleotide position. Consensus sequences with greater than 20% of contributing reads not matching the consensus nucleotide call were considered ambiguous and were discarded. The output FASTQ file listed all observed gRNA sequences, and the number of times each unique sequence occurred in this list was representative of the number of original gRNA molecules in the sample that received a unique UMI. The number of occurrences of each unique gRNA sequence was used to calculate the frequency of the fully correct (intended) gRNA sequence and the frequency of each observed variant gRNA sequence.

[0230] Consensus gRNA sequences were aligned to the reference gRNA sequence (intended sequence) using the BWA MEM program with a minimum alignment score of 0. A descriptive identity label was applied to each variant including the variant type (mismatch, deletion, or insertion), position of difference from the reference sequence, and the nucleotide change (formismatch and insertion variants) or deletion length (for deletion variants). The variant ID, full variant sequence, sequence length, UMI count, and population frequency were summarized in a report table.

[0231] Table 6 shows an abridged summary of gRNA sequence frequency results for the Cas9 gRNA sample. A total of 58310 different UMI groups representative of individual original gRNA molecules were assayed, from which 1966 unique gRNA sequences were identified across a wide range of frequencies. As expected, the intended gRNA sequence was the most common sequence in the population by a large margin, observed at a frequency of 78.69%. Exceptionally rare variant gRNA sequences that were observed only once within the assayed population (0.0017% frequency) are reliably detected and counted, indicating high sensitivity for sequence variant detection.Table 6: Top 50 most frequent sequences observed for the Cas9 gRNA sample

[0232] 1Descriptive variant identity summaries (Variant ID) are reported in the format: [variant type], [first nucleotide position of deviation from reference sequence], [length of deletion or inserted / changed nucleotide identity]. Variant types include deletions (DEL), insertions (INS), and mismatches (MM).

[0233] FIG. 6 provides a non-limiting example of data for the log2 fold-change in gRNA sequence frequency between read-based quantification and UMLbased quantification (y-axis) plotted as a function of UMLbased sequence frequency (x-axis) wherein each point represents an individual observed gRNA sequence. The log2 fold-change values shown on this plot indicate the magnitude of impact that UMI normalization had on the frequency quantification for each gRNA sequence. A log2 fold-change value close to 1.0 for a gRNA sequence indicates close correspondence between read-based and UMLbased frequency quantification (e.g. diamond point). A large positive log2 fold-change value for a gRNA sequence indicates over-representation of sequencing reads for that gRNA sequence relative to the original molecular population (e.g. square point), while a large negative log2 fold-change value indicates under-representation (e.g. triangle point).Example 3 - Assessment of a Cas 12a gRNA sample

[0234] The same procedure as described in Example 2 was carried out, except that an exemplary Cas 12a gRNA sample designed as a 66mer chimeric DNA-RNA oligonucleotide was used in place of the Cas9 gRNA sample.

[0235] FIG. 7 provides a non-limiting example of a representative TapeStation electropherogram trace for a purified NGS amplicon library prepared using the methods described herein for a 66- nucleotide Cas 12a gRNA. The peaks labeled “Lower” and “Upper” are the lower and upper size markers, respectively, used for sample size determination with TapeStation D1000 reagents. The observed average size of the final sequencing library was 223 bp, which is within a reasonable range of the expected size of 219 bp for fully indexed library molecules with a 66 bp insert.

[0236] Table 7 shows an abridged summary of gRNA sequence frequency results for the Cas 12a gRNA sample. A total of 22960 different UMI groups representative of individual original gRNA molecules were assayed, from which 499 unique gRNA sequences were identified across a wide range of frequencies. As expected, the intended gRNA sequence was the most common sequence in the population by a large margin, observed at a frequency of 84.40%. Exceptionally rare variant gRNA sequences that were observed only once within the assayed population (0.0044%frequency) were reliably detected and counted, indicating high sensitivity for sequence variant detection. Comparison of this data to the data shown in Table 6 indicates that the method described in this disclosure performs similarly well for both Cas9 and Casl2a gRNAs.Table 7: Top 50 most frequent variants observed for the Casl2a gRNA sample.

[0237] 1Descriptive variant identity summaries (Variant ID) are reported in the format: [variant type], [first nucleotide position of deviation from reference sequence], [length of deletion or inserted / changed nucleotide identity]. Variant types include deletions (DEL), insertions (INS), and mismatches (MM).Example 4 - Assessment of a gRNA sample recovered from a denatured ribonucleoprotein complex

[0238] This example demonstrates the application of methods in the present invention to the preparation of Illumina NGS libraries from exemplary gRNA-Casl2a ribonucleoprotein (RNP) complex samples designed as a 66mer chimeric DNA-RNA oligonucleotide gRNA bound to a Casl2a nuclease protein. The following experiment was performed to produce representative sequencing libraries from each RNP sample.

[0239] Dilution and denaturation of RNP sample was performed as follows: the RNP sample was diluted to a final concentration of 25 pM in nuclease-free water, and 20 pL of 25 pM RNP sample was transferred into a well of a 96-well plate. The plate was sealed and incubated at a temperature of 75 °C in a thermal cycler for 20 minutes and following incubation the plate was transferred to room temperature.

[0240] The denatured Casl2a protein was removed and the gRNA component from the denatured RNP was purified as follows: 30 pL of RNAClean XP bead solution (Beckman-Coulter) were added to the denatured RNP sample and mixed by pipetting, the mixture was incubated at room temperature for 5 minutes, the plate was transferred to a magnetic rack for 5 minutes to collect beads on the side of the well, the supernatant was removed, beads were washed 3 times with 70% ethanol (Fisher) (volume / volume in nuclease-free water (Ambion)), beads were air-dried at room temperature for 5 minutes to remove residual ethanol, beads were resuspended in 20 pL of nuclease-free water, the plate was incubated at room temperature off of the magnet for 2 minutes, the plate was transferred back to the magnet for 3 minutes to collect beads on the side of the well, and finally 18 pL of supernatant containing the eluted gRNA was transferred to a new well.

[0241] Concentration measurement and dilution of purified gRNA was measured as follows: 2 pL of nuclease-free water was used as a blank measurement on a NanoDrop One Spectrophotometer instrument (Thermo), 2 pL of purified gRNA was used for concentration measurement with the RNA setting on the NanoDrop One instrument, and based on the measured concentration the gRNA was diluted to a final concentration of 3.07 pM in nuclease-free water with a final volume of 20 pL.

[0242] Following RNP denaturation, gRNA recovery, and gRNA dilution, the remainder of the procedure was carried out fully as described in Example 2. The data output format was the same as the format described in Example 2 and Example 3 and the data obtained from the experiment was highly similar to that shown in Example 3 and typical for Casl2a gRNA input material. The approach described here can be used to compare the gRNA sequence sub-population frequencies across different RNP samples and assess the impact of factors such as RNP complexing procedure, storage conditions, or Cas protein versions.Example 5 - Modifications to enhance broad capture of different nucleic acid sequences

[0243] This example demonstrates the application of modified methods to the preparation of Illumina Next-Generation Sequencing (NGS) libraries from exemplary Cas 12a gRNA samples designed as a 66mer chimeric DNA-RNA oligonucleotide. Such samples produced lower yield libraries when the method described in Examples 2 and 3 was applied. Differences in primary sequence and secondary structure of the substrate oligonucleotide may impact efficiency of the 3’ and 5’ ligation reactions and therefore lead to low sequencing library yield for some substrate sequences. Various reaction conditions can be modified to improve ligation efficiency as may be helpful for specific sequence substrates, as exemplified here.

[0244] The modified approach employed in this Example was similar to the procedures described in Example 2 and Example 3, but incorporated the following modifications: T4 RNA ligase, truncated (New England Biolabs) was used as the ligase enzyme for the 3’ ligation reaction, the 3’ ligation reaction was carried out at a temperature increasing from 4°C to 37°C over the course of 18 hours, the 3’ ligation reaction was carried out in the presence of 25% polyethylene glycol, and a 3’ adapter with two randomized nucleotides at its 5’ terminus (SEQ ID NO: 106, as shown below) was used for ligation to the gRNA substrate. Inclusion of the randomized nucleotides in the adapter sequence improved the compatibility of the ligation procedure with gRNA samples having a primary sequence that otherwise exhibited low yields.

[0245] Combination and denaturation of the gRNA sample and the 3’ adapter was performed as follows: 1.5 pL of 3.33 pM gRNA sample and 2 pL of 7.5 pM 3’ adapter (SEQ ID NO: 106) were combined in a well of a 96-well plate, the plate was incubated at 70°C in a thermal cycler for 2 minutes, and the plate was then immediately transferred to a 4°C cold block for at least 5 minutes.

[0246] SEQ ID NO: 106: 5’- / 5rApp / NNAGTGTCATAGATCGGAAGAGCACACGTCTGAACTCC / 3AmMO / -3’

[0247] Ligation of the 3’ adapter to the gRNA substrate was performed as follows: 1 pL of lOx T4 RNA Ligase Buffer (New England Biolabs), 5 pL of 50% polyethylene glycol 8000 (New England Biolabs), 0.5 pL of T4 RNA Ligase 2 truncated enzyme (New England Biolabs) were added to the plate well containing the denatured gRNA substrate and 3’ adapter, and the sample was mixed by pipetting. The reaction was carried out in a thermal cycler using the following cycle program:

[0248] 1. 4°C for 4 hours

[0249] 2. 16°C for 4 hours

[0250] 3. 25°C for 6 hours

[0251] 4. 37°C for 4 hours

[0252] The 3’ adapter ligation reaction products were purified to perform a buffer exchange as follows: 12 pL of RNAClean XP bead solution (Beckman-Coulter) were added to the 3’ ligation reaction and mixed by pipetting, the mixture was incubated at room temperature for 5 minutes, the plate was transferred to a magnetic rack for 5 minutes to collect beads on the side of the well, the supernatant was removed, beads were washed 3 times with 70% ethanol (Fisher) (volume / volume in nuclease-free water (Ambion)), beads were air-dried at room temperature for 5 minutes to remove residual ethanol, beads were resuspended in 11 pL of nuclease-free water, the plate was incubated at room temperature off of the magnet for 2 minutes, the plate was transferred back to the magnet for 3 minutes to collect beads on the side of the well, and finally 9.75 pL of supernatant containing the eluted ligation product was transferred to a new well.

[0253] Phosphorylation of the 5’ terminus of the gRNA substrate was performed as follows: 9.75 pL of bead-cleaned 3’ ligation product, 1.25 pL of lOx T4 PNK Reaction Buffer (New EnglandBiolabs), 1.25 pL lOmM ATP (New England Biolabs), and 0.25 pL T4 Polynucleotide Kinase (New England Biolabs) were combined in a well of a 96-well plate, and the plate was incubated at 37 °C for 30 minutes in a thermal cycler.

[0254] The 5’ phosphorylation reaction products were purified to perform a buffer exchange as follows: 15 pL of RNAClean XP bead solution (Beckman-Coulter) were added to the 5’ phosphorylation reaction and mixed by pipetting, the mixture was incubated at room temperature for 5 minutes, the plate was transferred to a magnetic rack for 5 minutes to collect beads on the side of the well, the supernatant was removed, beads were washed 3 times with 70% ethanol (volume / volume in nuclease-free water), beads were air-dried at room temperature for 5 minutes to remove residual ethanol, beads were resuspended in 8.2 pL of nuclease-free water (Ambion), the plate was incubated at room temperature off of the magnet for 2 minutes, the plate was transferred back to the magnet for 3 minutes to collect beads on the side of the well, and finally 6.8 pL of supernatant containing the eluted ligation product was transferred to a new well.

[0255] Combination and denaturation of the gRNA sample and the 5’ adapter was performed as follows: 5.42 pL of bead-cleaned gRNA 5’ phosphorylation product and 1.33 pL of 11.25 pM 5’ adapter (e.g. SEQ ID NO: 107) were combined in a well of a 96-well plate, the plate was incubated at 70°C in a thermal cycler for 2 minutes, and the plate was then immediately transferred to a 4°C cold block for at least 5 minutes.

[0256] SEQ ID NO: 107: 5’- / 5AmMC6 / rGrUrUrCrArGrArGrUrUrCrUrArCrArGrUrCrCrGrArCrGrArUrCrNrNrNrNrNrNrNrNrNrNrN rNrNrNrNr ArCrGr ArUr ArC-3 ’

[0257] Ligation of the 5’ adapter to the gRNA substrate was performed as follows: 1.5pL of lOx T4 RNA Ligase Buffer (New England Biolabs), 4.5 pL of 50% polyethylene glycol 8000 (New England Biolabs), 1.5 pL of 10 mM ATP (New England Biolabs), and 0.75 pL of T4 RNA Ligase 1 (New England Biolabs) were added to the well containing the denatured gRNA substrate and 5’ adapter, the sample was mixed by pipetting, and the plate was incubated at 25°C for 2 hours.

[0258] Reverse transcription of the adapter-ligated gRNA substrate to generate cDNA was performed as follows: 7 pL of 5’ adapter ligation reaction, 0.5 pL of 10 pM RT primer (e.g. SEQ ID NO: 108), 0.5 pL Murine RNase Inhibitor (New England Biolabs), 2 pL lOx AMV RT Buffer (New England Biolabs), 1 pL of 10 mM dNTP Mix (New England Biolabs), 1 pL AMV ReverseTranscriptase (New England Biolabs), and 8 pL nuclease-free water (Ambion) were combined in a well of a 96-well plate, and the plate was incubated at 50°C for 30 minutes.

[0259] SEQ ID NO: 108: 5’- GGAGTTCAGACGTGTGCTCTTCCGATCTATGACACT-3’

[0260] The reverse transcription reaction products were purified to perform a buffer exchange as follows: 24 pL of SPRIselect bead solution (Beckman-Coulter) was added to the reverse transcription reaction and mixed by pipetting, the mixture was incubated at room temperature for 5 minutes, the plate was transferred to a magnetic rack for 5 minutes to collect beads on the side of the well, the supernatant was removed, beads were washed 2 times with 80% ethanol (volume / volume in nuclease-free water), beads were air-dried at room temperature for 5 minutes to remove residual ethanol, beads were resuspended in 25 pL of nuclease-free water (Ambion), the plate was incubated at room temperature off of the magnet for 2 minutes, the plate was transferred back to the magnet for 3 minutes to collect beads on the side of the well, and finally 22 pL of supernatant containing the eluted cDNA product was transferred to a new well.

[0261] Purified cDNA was serially diluted to a final dilution of 1:500 as follows: 5 pL of cDNA to 95 pL of nuclease-free water (Ambion) and then mixed by pipetting to create a 1:20 dilution, 5 pL of 1:20 diluted cDNA was added to 120 pL nuclease-free water (Ambion) and mixed by pipetting.

[0262] PCR amplification and Illumina indexing of cDNA derived from the gRNA substrate was performed as follows: 10 pL of 1:500 diluted cDNA, 12.5 pL of 2x Q5 Hot Start Master Mix (New England Biolabs), 1.25 pL of 10 pM i5 Universal Primer (e.g., SEQ ID NO: 4), and 1.25 pL of 10 pM i7 Indexing Primer 1 (SEQ ID NO: 5) were combined in a well of a 96-well plate and mixed by pipetting. The reaction was carried out in a thermal cycler using the following cycle program:

[0263] 1. 98°C for 30 seconds

[0264] 2. 98°C for 10 seconds

[0265] 3. 69°C for 20 seconds

[0266] 4. 72°C for 20 seconds

[0267] 5. Repeat steps 2-4 22 times

[0268] 6. 72°C for 2 minutes

[0269] 7. Hold at 4°C

[0270] SEQ ID NO: 4: 5’-AATGATACGGCGACCACCGAGATCTACACGTTCAGAGTTCTACAGTCCGACGATC-3’

[0271] SEQ ID NO: 5: 5’-CAAGCAGAAGACGGCATACGAGATNNNNNNNNGTGACTGGAGTTCAGACGTGTGCT CTTCCGATCT-3’

[0272] The PCR products were purified and size selected to remove residual primers and adapterdimer side products as follows: 30 pL (1.2x reaction volumes) of SPRI select bead solution (Beckman-Coulter) was added to the reverse transcription reaction and mixed by pipetting, the mixture was incubated at room temperature for 5 minutes, the plate was transferred to a magnetic rack for 5 minutes to collect beads on the side of the well, the supernatant was removed, beads were washed 2 times with 80% ethanol (volume / volume in nuclease-free water), beads were airdried at room temperature for 5 minutes to remove residual ethanol, beads were resuspended in 50 pL of nuclease-free water (Ambion), the plate was incubated at room temperature off of the magnet for 2 minutes, the plate was transferred back to the magnet for 3 minutes to collect beads on the side of the well, and finally 45 pL of supernatant containing the eluted amplicon library was transferred to a new sample tube.

[0273] The purified amplicon library concentration and average amplicon size were assessed by TapeStation analysis as follows: 1 pL of purified PCR product was combined with 3 pL of TapeStation D1000 Loading Dye (Agilent) in a well of an 8-well PCR strip tube, 1 pL of TapeStation DI 000 Ladder (Agilent) was combined with 3 pL of DI 000 Loading Dye (Agilent) in a separate well of the same 8-well PCR strip tube, the contents of the strip tube were mixed by vortexing, the strip tube was placed in an Agilent TapeStation 4200 along with a TapeStation D1000 Screen Tape (Agilent), and the default D1000 program was performed. The TapeStation results were assessed to calculate library concentration and assess average library size.

[0274] FIG. 9 provides a non-limiting example of a representative TapeStation D1000 electropherogram trace for a purified NGS amplicon library prepared using the methods described herein for a 66-nucleotide Casl2a gRNA. The peaks labeled “Lower” and “Upper” are the lower and upper size markers, respectively, used for sample size determination with TapeStation D1000reagents. The observed average size of the final sequencing library was 223 bp, which is within a reasonable range of the expected size of 219 bp for fully indexed library molecules with a 66 bp insert.

[0275] The amplicon library was denatured and diluted for sequencing on the Illumina MiSeq instrument as follows: the purified library DNA was diluted to 4 nM in nuclease-free water (Ambion) using the concentration calculated from TapeStation analysis in the previous step, 5 pL of the 4 nM library was combined with 5 pL of 0.2 N Sodium hydroxide (Sigma- Aldrich) in a 1.5 mL DNA LoBind tube (Eppendorf) and mixed by pipetting, the tube was incubated at room temperature for 5 minutes, the library was diluted to 20 pM by adding 990 pL of pre-chilled HT1 buffer (Illumina), the tube was mixed by vortexing and placed on ice, the denatured library was diluted to 13 pM by combining 390 pL of 20 pM and 210 pL of pre-chilled HT1 buffer (Illumina). PhiX control library DNA was spiked in at a 5% concentration by combining 30 pL of 12.5 pM denatured PhiX library (Illumina) with 570 pL of 13 pM gRNA library and mixing by pipetting.

[0276] The library was heat denatured as follows: the tube containing the final combined sequencing library was incubated at 96°C for 2 minutes, the tube was vortexed to mix, the tube was spun down in a microcentrifuge to collect liquid at the bottom of the tube, and the tube was immediately placed on ice for at least 5 minutes.

[0277] The final denatured library was loaded into a MiSeq v2 Reagent Kit 300-cycle (Illumina) and a paired-end 2 x 151 sequencing run was performed according to a run sample sheet provided by the user, which includes indices that identify the sample. Accordingly, multiple samples can be run simultaneously.

[0278] Tables 8 to 11 provide summaries of the reagents, oligonucleotides, consumable materials, and instruments used to perform library preparation and sequencing.Table 8. Reagents.Table 9. Oligonucleotides.Table 10. Consumables.Table 11. Instruments.

[0279] MiSeq sequencing data was analyzed as follows: sequencing reads were demultiplexed using the indices provided in the run sample sheet. The FASTQ files were processed to trim the Illumina adapter sequences using the Trim Galore program, paired reads containing intact 5’ and 3’ adapter sequences were identified and retained using the UMI- tools Extract program following a regular expression (UMI pattern 1: “A(?P<umi_l>[ATGC]{ 15})ACGATAC”; UMI pattern 2: “AATGACACT”) and the 15 nucleotide UMI in each read was extracted and moved to the read name using the UMI-tools Extract program. Paired read 1 and read 2 FASTQ files were merged using the FLASh merge program and read pairs containing base call disagreement between read 1 and read 2 were discarded. Using custom software, reads that shared identical 15 nucleotide UMI sequences and had UMI base quality scores >Q30 were grouped together using a metadata label. Each UMI group containing 5 or more sequencing reads was collapsed into a single consensus sequence using the most frequent base call at each nucleotide position. Consensus sequences with greater than 20% of contributing reads not matching the consensus nucleotide call were considered ambiguous and were discarded. The output FASTQ file listed all observed gRNA sequences, and the number of times each unique sequence occurred in this list was representative of the number of original gRNA molecules in the sample that received a unique UMI. The number of occurrences of each unique gRNA sequence was used to calculate the frequency of the fully correct (intended) gRNA sequence and the frequency of each observed variant gRNA sequence.

[0280] Consensus gRNA sequences were aligned to the reference gRNA sequence (intended sequence) using the BWA MEM program with a minimum alignment score of 0. A descriptiveidentity label was applied to each variant including the variant type (mismatch, deletion, or insertion), position of difference from the reference sequence, and the nucleotide change (for mismatch and insertion variants) or deletion length (for deletion variants). The variant ID, full variant sequence, sequence length, UMI count, and population frequency were summarized in a report table.

[0281] Table 12 shows an abridged summary of gRNA sequence frequency results for the Cas9 gRNA sample. A total of 32929 different UMI groups representative of individual original gRNA molecules were assayed, from which 647 unique gRNA sequences were identified across a wide range of frequencies. As expected, the intended gRNA sequence was the most common sequence in the population by a large margin, observed at a frequency of 83.10%. Exceptionally rare variant gRNA sequences that were observed only once within the assayed population (0.0030% frequency) were reliably detected and counted, indicating high sensitivity for sequence variant detection.Table 12: Top 50 most frequent sequences observed for the Casl2a gRNA sample

[0282] 1Descriptive variant identity summaries (Variant ID) are reported in the format: [variant type], [first nucleotide position of deviation from reference sequence], [length of deletion or inserted / changed nucleotide identity]. Variant types include deletions (DEL), insertions (INS), and mismatches (MM).Example 6 - Utilization for ultra-sensitive detection of contaminating oligonucleotide sequences

[0283] This example demonstrates the application of modified methods to the preparation of Illumina Next-Generation Sequencing (NGS) libraries from two exemplary Casl2a gRNA samples suspected of potentially containing a very low level of an unknown contaminant gRNA sequence. With minor modifications to the method described in Examples 2, 3, and 5, the sensitivity of the assay for detecting extremely rare specific sequences within the molecular population can be greatly increased at the expense of sacrificing the ability to correct PCR and sequencing errors through collapsing of UMI groups. In this example, the indexing PCR step was modified to increase the concentration of cDNA template used and reduce the number of PCR cycles, which resulted in fewer PCR duplicate sequencing reads for each UMI and a greater number of total unique UMIs representative of unique original gRNA molecules in the final sequencing data.

[0284] While UMLbased error correction is not possible with this configuration of the method, the UMI labeling approach still provided single-molecule quantification and allowed for the precise calculation of the total number of gRNA molecules analyzed. The configuration of the method described in this example achieved extremely high sensitivity well beyond what is possible with other assays by sequencing and analyzing up to 1.2xl06original gRNA molecules. Because the data output from this method provided the full sequence of all gRNA molecules observed in the sample, the presence of contaminant gRNA sequences could be confirmed unambiguously.

[0285] The same Illumina sequencing library preparation procedure as described in Example 5 was carried out, except that 10 pL of a 1:20 dilution of cDNA was used as the template for the indexing PCR step instead of a 1:500 dilution of cDNA and a total of 18 PCR cycles were performed instead of 23 cycles of PCR. Library quantification and sequencing was performed as described in Example 5.

[0286] MiSeq sequencing data was analyzed as follows: sequencing reads were demultiplexed using the indices provided in the run sample sheet. The FASTQ files were processed to trim the Illumina adapter sequences using the Trim Galore program, paired reads containing intact 5’ and 3’ adapter sequences were identified and retained using the UMI- tools Extract program following a regular expression (UMI pattern 1: “A(?P<umi_l>[ATGC]{ 15})ACGATAC”; UMI pattern 2: “AATGACACT”) and the 15 nucleotide UMI in each read was extracted and moved to the read name using the UMI-tools Extract program. Paired read 1 and read 2 FASTQ files were merged using the FLASh merge program and basecall discrepancies were resolved by selecting the basecall with higher quality score. Using custom software, reads that shared identical 15 nucleotide UMI sequences and had UMI base quality scores >Q20 were grouped together using a metadata label. UMI groups containing more than one sequencing read were collapsed into a single consensus sequence using the most frequent base call at each nucleotide position and UMI groups containing a single sequencing read were retained unchanged. The output FASTQ file listed all observed gRNA sequences, and the number of times each unique sequence occurred in this list was representative of the number of original gRNA molecules in the sample that received a unique UMI. The total number of sequences in this output FASTQ file provides the number of individual original gRNA molecules analyzed reported in Table 13. The Cutadapt program was used to search the collapsed UMI FASTQ file for a portion of the suspected contaminant gRNA sequence and generate an output FASTQ file containing the full sequence of all reads that include all or part of the query sequence ‘GAGTCTCTCAGCTGGTACACG’ (SEQ ID NO: 159), within the program search parameters set to require at least 12 nucleotide overlap with the query sequence and allow up to 4 mismatches from the query sequence. All sequences present in the output FASTQ file were manually verified to match the suspected contaminant gRNA sequence and the number of different UMIs corresponding to the contaminant gRNA sequence in each sample was reported in Table 13.

[0287] Table 13 shows the numbers of raw sequencing reads, total unique molecules assayed, number of sequences matching the suspected contaminant gRNA, and the full sequences ofcontaminant gRNAs detected in two samples. Approximately 200,000 and 1,200,000 unique molecules were analyzed for Sample A and Sample B, respectively. Two gRNA sequences matching the suspected contaminant gRNA were observed in Sample A, while no contaminant gRNA sequences were observed in Sample B. The detected contaminant gRNA sequences were observed at a frequency of 0.00103% in the molecular population of Sample A, illustrating an extremely high sensitivity of this configuration of the method. The absence of any contaminant gRNA sequences within the nearly 1.2 million unique sequences analyzed from Sample B provides a strong indication that the sample was likely free of this specific contamination down to a very low limit of detection. The full sequences of each detected contaminant gRNA are provided to allow definitive manual verification of the presence of these sequences within the sample. The permissive settings for the query sequence search allowed for the detection of one contaminant sequence containing a one nucleotide mismatch from the contaminant gRNA reference sequence.Table 13: Summary of potential gRNA contaminant detection analysis for two Casl2a gRNA samplesEXEMPLARY EMBODIMENTS

[0288] Exemplary embodiments of the compositions, methods, systems, and kits described herein include:

[0289] Embodiment 1: A sequencing adapter composition comprising: a) a 3’ adapter sequence configured for ligation to a 3’ end of a single- stranded oligonucleotide, the 3’ adapter sequence comprising, in a 5’ to 3’ direction: a ligatable 5’ end, a first linker sequence, a first index primerbinding site sequence, and a non-ligatable 3’ end, wherein at least a portion of the 3’ adapter sequence is capable of hybridizing to a reverse transcription primer; and b) a 5’ adapter sequence configured for ligation to a 5’ end of the single- stranded oligonucleotide, the 5’ adapter sequence comprising, in a 5’ to 3’ direction: a non-ligatable 5’ end, a second index primer binding site sequence, and a ligatable 3’ end; wherein at least one of the 3’ adapter sequence and the 5’ adapter sequence comprises a randomized sequence.

[0290] Embodiment 2: The sequencing adapter composition of Embodiment 1, wherein the randomized sequence is 5’ adjacent to the first linker and / or 3’ adjacent to the second index primer binding site sequence.

[0291] Embodiment 3: The sequencing adapter composition of Embodiment 1 or 2, wherein the 5’ adapter sequence further comprises a second linker sequence.

[0292] Embodiment 4: The sequencing adapter composition of any one of Embodiments 1-3, wherein the 3’ adapter sequence comprises in a 5’ to 3’ direction: a ligatable 5’ end, a first randomized sequence, a first linker sequence, a first index primer binding site sequence, and a non- ligatable 3’ end.

[0293] Embodiment 5: The sequencing adapter composition of any one of Embodiments 1-4, wherein the 5’ adapter sequence comprises in a 5’ to 3’ direction: a non-ligatable 5’ end, a second index primer binding site sequence, a second randomized sequence, a second linker sequence, and a ligatable 3’ end.

[0294] Embodiment 6: The sequencing adapter composition of any one of Embodiments 1-5, wherein the 3’ adapter sequence does not comprise a stem- loop structure sequence.

[0295] Embodiment 7: The sequencing adapter composition of any one of Embodiments 1-6, wherein less than 5%, 10%, or 20% of the 3’ adapter sequence is capable of forming a secondary structure.

[0296] Embodiment 8: The sequencing adapter composition of any one of Embodiments 1-7, wherein the 3’ adapter sequence is not capable of forming a secondary structure.

[0297] Embodiment 9: The sequencing adapter composition of any one of Embodiments 1-8, wherein the ligatable 5’ end of the 3’ adapter sequence comprises a 5 ’-adenylated nucleotide.

[0298] Embodiment 10: The sequencing adapter composition of any one of Embodiments 1-9, wherein the ligatable 5’ end of the 3’ adapter sequence comprises a 5 ’-phosphorylated nucleotide.

[0299] Embodiment 11: The sequencing adapter composition of any one of Embodiments 1-10, wherein the ligatable 3’ end of the 5’ adapter sequence comprises a 3 ’-hydroxylated nucleotide.

[0300] Embodiment 12: The sequencing adapter composition of any one of Embodiments 1-11, wherein the ligatable 5’ end of the 3’ adapter sequence comprises a 5 ’-hydroxylated nucleotide.

[0301] Embodiment 13: The sequencing adapter composition of any one of Embodiments 1-8 and 12, wherein the ligatable 3’ end of the 5’ adapter sequence comprises a 3 ’-phosphorylated nucleotide.

[0302] Embodiment 14: The sequencing adapter composition of any one of Embodiments 1-13, wherein the non-ligatable 3’ end of the 3’ adapter sequence comprises a 3’ modification of a 3’ nucleotide.

[0303] Embodiment 15: The sequencing adapter composition of Embodiment 14, wherein the 3’ modification is selected from the group consisting of: a 3’ amino group, a 3’ dideoxy-C modification, a 3’ inverted dT modification, a 3’ C3 spacer modification, and a 3’ phosphoryl modification.

[0304] Embodiment 16: The sequencing adapter composition of any one of Embodiments 1-15, wherein the non-ligatable 5’ end of the 5’ adapter sequence comprises a 5’ modification of a 5’ nucleotide.

[0305] Embodiment 17: The sequencing adapter composition of Embodiment 16, wherein the 5’ modification is selected from the group consisting of: a 5’ hydroxyl group, a 5’ dideoxy-C modification, 5’ inverted dideoxy-T modification, and a 5’ C3 spacer modification.

[0306] Embodiment 18: The sequencing adapter composition of any one of Embodiments 1-17, wherein the 3’ adapter sequence is or comprises a sequence that is at least about 80%, at least about 85%, at least about 90%, or at least about 95% identical to a sequence that is complementary to the reverse transcription primer.

[0307] Embodiment 19: The sequencing adapter composition any one of Embodiments 1-3 and 5- 18, wherein the 3’ adapter sequence is or comprises a sequence that is 100% identical to a sequence that is complementary to the reverse transcription primer.

[0308] Embodiment 20: The sequencing adapter composition of any one of Embodiments 1-19, wherein the first linker sequence has a length of about 4 nucleotides to about 20 nucleotides.

[0309] Embodiment 21: The sequencing adapter composition of any one of Embodiments 1-20, wherein the first linker sequence has a length of 10 nucleotides.

[0310] Embodiment 22: The sequencing adapter composition of any one of Embodiments 1-21, wherein the first index primer binding site sequence has a length of about 20 nucleotides to about 35 nucleotides.

[0311] Embodiment 23: The sequencing adapter composition of any one of Embodiments 1-22, wherein the first index primer binding site sequence has a length of about 28 nucleotides.

[0312] Embodiment 24: The sequencing adapter composition of any one of Embodiments 1-23, wherein the first index primer binding site sequence is capable of binding to a sequencing read 2 adapter.

[0313] Embodiment 25: The sequencing adapter composition of any one of Embodiments 1-24, wherein the second index primer binding site sequence has a length of about 20 nucleotides to about 30 nucleotides.

[0314] Embodiment 26: The sequencing adapter composition of any one of Embodiments 1-25, wherein the second index primer binding site sequence has a length of about 22 nucleotides to about 28 nucleotides.

[0315] Embodiment 27: The sequencing adapter composition of any one of Embodiments 1-26, wherein the second index primer binding site sequence has a length of about 26 nucleotides.

[0316] Embodiment 28: The sequencing adapter composition of any one of Embodiments 1-27, wherein the second index primer binding site sequence is capable of binding to a sequencing read 1 adapter.

[0317] Embodiment 29: The sequencing adapter composition of any one of Embodiments 4-28, wherein the first randomized sequence has a length of about 1 nucleotide to about 20 nucleotides.

[0318] Embodiment 30: The sequencing adapter composition of any one of Embodiments 4-29, wherein the first randomized sequence has a length of at least 2 nucleotides.

[0319] Embodiment 31: The sequencing adapter composition of any one of Embodiments 5-30, wherein the second randomized sequence has a length of about 10 nucleotides to about 20 nucleotides.

[0320] Embodiment 32: The sequencing adapter composition of any one of Embodiments 5-31, wherein the second randomized sequence has a length of about 15 nucleotides.

[0321] Embodiment 33: The sequencing adapter composition of any one of Embodiments 3-32, wherein the second linker sequence has a length of about 4 nucleotides to about 20 nucleotides.

[0322] Embodiment 34: The sequencing adapter composition of any one of Embodiments 3-33, wherein the second linker sequence has a length of about 7 nucleotides.

[0323] Embodiment 35: The sequencing adapter composition of any one of Embodiments 1-34, wherein the 3’ adapter sequence has a total length of about 20 nucleotides to about 60 nucleotides.

[0324] Embodiment 36: The sequencing adapter composition of any one of Embodiments 1-35, wherein the 3’ adapter sequence has a total length of about 30 nucleotides to about 45 nucleotides.

[0325] Embodiment 37: The sequencing adapter composition of any one of Embodiments 1-36, wherein the 3’ adapter sequence has a total length of about 38 nucleotides.

[0326] Embodiment 38: The sequencing adapter composition of any one of Embodiments 1-37, wherein the 5’ adapter sequence has a total length of about 40 nucleotides to about 60 nucleotides.

[0327] Embodiment 39: The sequencing adapter composition of any one of Embodiments 1-38, wherein the 5’ adapter sequence has a total length of about 42 nucleotides to about 55 nucleotides.

[0328] Embodiment 40: The sequencing adapter composition of any one of Embodiments 1-39, wherein the 5’ adapter sequence has a total length of about 48 nucleotides.

[0329] Embodiment 41: The sequencing adapter composition of any one of Embodiments 1-40, wherein the 3’ adapter sequence comprises a DNA sequence.

[0330] Embodiment 42: The sequencing adapter composition of any one of Embodiments 1-41, wherein the 5’ adapter sequence comprises an RNA sequence or a DNA sequence.

[0331] Embodiment 43: The sequencing adapter composition of any one of Embodiments 1-42, wherein the 3’ adapter sequence and the 5’ adapter sequence do not comprise substantial complementarity with each other.

[0332] Embodiment 44: A method of preparing a sequencing library for a single- stranded template nucleic acid molecule, comprising: a) contacting the single- stranded nucleic acid molecule with the 3’ adapter sequence of any one of Embodiments 1-43 and performing a ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single- stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule; b) performing an end modification reaction to modify a 5’ end of the first ligated nucleic acid molecule; c) contacting the first ligated nucleic acid molecule with the 5’ adapter sequence of any one of Embodiments 1-43 and performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule; d) contacting the second ligated nucleic acid molecule with a reverse transcription primer and allowing the reverse transcription primer to hybridize to the 3’ adapter sequence; e) performing a reverse transcription reaction, thereby producing a cDNA molecule comprising the template nucleic acid molecule; f) contacting the cDNA molecule with a first index primer capable of hybridizing to the first index primer binding site sequence, and with a second index primer capable of hybridizing to the second index primer binding site sequence, under conditions that allow hybridization of the first and second index primers to their respective binding site sequences; g) amplifying the cDNA molecule to produce an amplification product; and h) purifying the amplification product to produce the sequencing library for the single-stranded template nucleic acid molecule.

[0333] Embodiment 45: The method of Embodiment 44, wherein the end modification reaction comprises a phosphorylation reaction to phosphorylate the 5’ end of the first ligated nucleic acid molecule.

[0334] Embodiment 46: The method of Embodiment 44, wherein the end modification reaction comprises an adenylation reaction to adenylate the 5’ end of the first ligated nucleic acid molecule.

[0335] Embodiment 47: The method of any one of Embodiments 44-46, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a DNA ligase or an RNA ligase.

[0336] Embodiment 48: The method of Embodiment 47, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a DNA ligase or an RNA ligase, comprises the use of a truncated version of a T4 RNA ligase 2.

[0337] Embodiment 49: The method of any one of Embodiments 44-48, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of a DNA ligase or an RNA ligase.

[0338] Embodiment 50: The method of Embodiment 49, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of RtcB ligase or T4 RNA ligase 1.

[0339] Embodiment 51: The method of any one of Embodiments 44-50, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a reaction buffer containing about 10% to about 25% (v / v) polyethylene glycol (PEG).

[0340] Embodiment 52: The method of any one of Embodiments 44-51, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a reaction buffer containing 25% (v / v) polyethylene glycol (PEG).

[0341] Embodiment 53: The method of any one of Embodiments 44-52, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of a reaction buffer containing about 10% to about 25% (v / v) polyethylene glycol (PEG).

[0342] Embodiment 54: The method of any one of Embodiments 44-53, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of a reaction buffer containing 25% (v / v) polyethylene glycol (PEG).

[0343] Embodiment 55: The method of any one of Embodiments 44-54, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a reaction buffer containing about 5% to about 15% (v / v) glycerol.

[0344] Embodiment 56: The method of any one of Embodiments 44-55, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a reaction buffer containing 10% (v / v) glycerol.

[0345] Embodiment 57: The method of any one of Embodiments 44-56, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of a reaction buffer containing about 5% to about 15% (v / v) glycerol.

[0346] Embodiment 58: The method of any one of Embodiments 44-57, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of a reaction buffer containing 10% (v / v) glycerol.

[0347] Embodiment 59: The method of any one of Embodiments 44-58, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule is carried out at a reaction temperature of about 4°C to about 37 °C.

[0348] Embodiment 60: The method of any one of Embodiments 44-59, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule is carried out at a reaction temperature of 25 °C.

[0349] Embodiment 61: The method of any one of Embodiments 44-60, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule is carried out in variable reaction temperature conditions cycling between 4°C and 37°C.

[0350] Embodiment 62: The method of any one of Embodiments 44-61, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule is carried out at a reaction temperature of about 4°C to about 37°C.

[0351] Embodiment 63: The method of any one of Embodiments 44-62, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule is carried out at a reaction temperature of 25°C.

[0352] Embodiment 64: The method of any one of Embodiments 44-63, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule is carried out in variable reaction temperature conditions cycling between 4°C and 37 °C.

[0353] Embodiment 65: The method of any one of Embodiments 44-64, wherein the method does not comprise a step of cleaving the first ligated nucleic acid molecule or the second ligated nucleic acid molecule.

[0354] Embodiment 66: The method of any one of Embodiments 44-65, wherein the singlestranded template nucleic acid molecule is an RNA molecule.

[0355] Embodiment 67: The method of any one of Embodiments 44-66, wherein the singlestranded template nucleic acid molecule is a guide RNA (gRNA) molecule.10356] Embodiment 68: The method of any one of Embodiments 44-67, wherein the singlestranded template nucleic acid molecule is a gRNA molecule recovered from a gRNA-Cas nuclease ribonucleoprotein (RNP).

[0357] Embodiment 69: The method of any one of Embodiments 44-67, wherein the single- stranded template nucleic acid molecule is a prime editing guide RNA (pegRNA) molecule.

[0358] Embodiment 70: The method of any one of Embodiments 44-66, wherein the singlestranded template nucleic acid molecule is a small interfering RNA (siRNA) molecule.

[0359] Embodiment 71: The method of any one of Embodiments 44-65, wherein the singlestranded template nucleic acid molecule is a DNA molecule.

[0360] Embodiment 72: A method for detecting and quantifying unique sequences in a population of single- stranded template nucleic acid molecules, comprising: a) preparing a sequencing library for the single- stranded template nucleic acid molecule using the method of any one of Embodiments 44-71 ; b) sequencing the sequencing library using a paired-end sequencing method to generate a plurality of sequence reads corresponding to the single- stranded template nucleic acid molecule; and c) analyzing sequence read data corresponding to the plurality of sequence reads to achieve at least one of: (i) correction of amplification errors; (ii) correction of sequencing errors; (iii) normalization of sequence-based amplification bias; and (iv) detection and quantification of unique sequences present in the population of single-stranded template nucleic acid molecules.

[0361] Embodiment 73: The method of Embodiments 72, wherein the population of singlestranded template nucleic acid molecules is a putatively homogenous population.

[0362] Embodiment 74: The method of Embodiments 72 or 73, wherein analyzing sequence read data comprises the use of one or more processors.

[0363] Embodiment 75: The method of any one of Embodiments 72-74, wherein analyzing the sequence read data comprises aligning paired sequence reads and discarding discrepant read pairs from further analysis.

[0364] Embodiment 76: The method of any one of Embodiments 72-74, wherein analyzing the sequence read data comprises grouping sequence reads according to unique molecular identifiers (UMIs) associated with each sequence read.

[0365] Embodiment 77: The method of Embodiment 76, wherein a unique molecular identifier (UMI) associated with each sequence read corresponds to the randomized sequence in at least one of the 3’ adapter sequence and the 5’ adapter sequence of any one of Embodiments 1-43.

[0366] Embodiment 78: The method of Embodiment 76, wherein a unique molecular identifier (UMI) associated with each sequence read corresponds to the second randomized sequence in the 5’ adapter sequence of any one of Embodiments 5-43.

[0367] Embodiment 79: The method of any one of Embodiments 76-78, further comprising determining a consensus sequence for each group of sequence reads identified based on the associated unique molecular identifier (UMI).

[0368] Embodiment 80: The method of Embodiment 79, further comprising identifying unique sequences present in the population of single- stranded template nucleic acid molecules based on the determined consensus sequence for each group of sequence reads.

[0369] Embodiment 81: The method of Embodiment 79 or 80, further comprising correcting for amplification and / or sequencing errors based on the determined consensus sequence for each group of sequence reads.

[0370] Embodiment 82: The method of any one of Embodiments 79-81, further comprising aligning a consensus sequence for each group of sequence reads to a reference sequence to identify variant sequences present in the population of single- stranded template nucleic acid molecules.

[0371] Embodiment 83: The method of Embodiment 82, further comprising determining a frequency of sequences present in the population of single- stranded template nucleic acid molecules that are identical to the reference sequence, and / or determining a variant sequence frequency for a variant sequence identified in the population of single-stranded template nucleic acid molecules.

[0372] Embodiment 84: The method of any one of Embodiments 72-83, wherein the population of single- stranded template nucleic acid molecules comprise guide RNA (gRNA) molecules.

[0373] Embodiment 85: The method of any one of Embodiments 72-84, wherein the population of single-stranded template nucleic acid molecules comprise gRNA molecules recovered from gRNA-Cas nuclease ribonucleoproteins (RNPs).

[0374] Embodiment 86: The method of any one of Embodiments 72-84, wherein the population of single-stranded template nucleic acid molecules comprise prime editing guide RNA (pegRNA) molecules.

[0375] Embodiment 87: The method of any one of Embodiments 72-83, wherein the population of single-stranded template nucleic acid molecules comprise small interfering RNA (siRNA) molecules.

[0376] Embodiment 88: The method of any one of Embodiments 72-83, wherein the population of single- stranded template nucleic acid molecules comprise DNA molecules.

[0377] Embodiment 89: A kit comprising: a) the sequencing adapter composition of any one of Embodiments 1-43, and b) at least one of: (i) a ligase; (ii) a buffer; (iii) enzymes and / or reagents for performing an end modification reaction to modify a 5’ end of a nucleic acid molecule; (iv) enzymes and / or reagents for performing nucleic acid amplification; and (v) enzymes and / or reagents for performing paired-end sequencing.

[0378] Embodiment 90: A system comprising: a) one or more processors; and b) a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: 1) receive sequence read data for a sequencing library generated using the method of any one of Embodiments 44-71 ; and 2) analyze the sequence read data for the sequencing library to achieve at least one of: (i) correction of amplification errors; (ii) correction of sequencing errors; (iii) normalization of sequence-based amplification bias; and (iv) detection and quantification of unique sequences present in a population of single-stranded template nucleic acid molecules.

[0379] Embodiment 91: The system of Embodiment 90, wherein analyzing the sequence read data comprises aligning paired sequence reads and discarding discrepant read pairs from further analysis.

[0380] Embodiment 92: The system of Embodiment 90 or 91, wherein analyzing the sequence read data comprises grouping sequence reads according to unique molecular identifiers (UMIs) associated with each sequence read.

[0381] Embodiment 93: The system of Embodiment 92, wherein a unique molecular identifier (UMI) associated with each sequence read corresponds to the randomized sequence in at least one of the 3’ adapter sequence and the 5’ adapter sequence of any one of Embodiments 1-43.

[0382] Embodiment 94: The method of Embodiment 93, wherein a unique molecular identifier (UMI) associated with each sequence read corresponds to the second randomized sequence in the 5’ adapter sequence of any one of Embodiments 5-43.

[0383] Embodiment 95: The system of any one of Embodiment 92-94, further comprising determining a consensus sequence for each group of sequence reads identified based on the associated unique molecular identifier (UMI).

[0384] Embodiment 96: The system of Embodiment 95, further comprising identifying unique sequences present in the population of single- stranded template nucleic acid molecules based on the determined consensus sequence for each group of sequence reads.

[0385] Embodiment 97: The system of Embodiment 95 or 96, further comprising correcting for amplification and / or sequencing errors based on the determined consensus sequence for each group of sequence reads.

[0386] Embodiment 98: The system of Embodiment 97, further comprising aligning a consensus sequence for each group of sequence reads to a reference sequence to identify variant sequences present in the population of single-stranded template nucleic acid molecules.

[0387] Embodiment 99: The system of Embodiment 98, further comprising determining a frequency of sequences present in the population of single- stranded template nucleic acid molecules that are identical to the reference sequence, and / or determining a variant sequence frequency for a variant sequence identified in the population of single-stranded template nucleic acid molecules.

[0388] Embodiment 100: The system of any one of Embodiments 90-99, wherein the singlestranded template nucleic acid molecules in the population comprise guide RNA (gRNA) molecules.

[0389] Embodiment 101: The method of any one of Embodiments 90-100, wherein the population of single-stranded template nucleic acid molecules comprise gRNA molecules recovered from gRNA-Cas nuclease ribonucleoproteins (RNPs).

[0390] Embodiment 102: The method of any one of Embodiments 90-100, wherein the population of single-stranded template nucleic acid molecules comprise prime editing guide RNA (pegRNA) molecules.

[0391] Embodiment 103: The system of any one of Embodiments 90-99, wherein the singlestranded template nucleic acid molecules in the population comprise small interfering RNA (siRNA) molecules.

[0392] Embodiment 104: The system of any one of Embodiments 90-99, wherein the singlestranded template nucleic acid molecules in the population comprise DNA molecules.

[0393] It should be understood from the foregoing that, while particular implementations of the disclosed compositions, methods, and systems have been illustrated and described, various modifications can be made thereto and are contemplated herein. It is also not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the preferable embodiments herein are not meant to be construed in a limiting sense. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. Various modifications in form and detail of the embodiments of the invention will be apparent to a person skilled in the art. It is therefore contemplated that the invention shall also cover any such modifications, variations and equivalents.

Claims

1. CLAIMSWhat is claimed is:

1. A sequencing adapter composition comprising: a) a 3’ adapter sequence configured for ligation to a 3’ end of a single- stranded oligonucleotide, the 3’ adapter sequence comprising, in a 5’ to 3’ direction: a ligatable 5’ end, a first linker sequence, a first index primer binding site sequence, and a non-ligatable 3’ end, wherein at least a portion of the 3’ adapter sequence is capable of hybridizing to a reverse transcription primer; and b) a 5’ adapter sequence configured for ligation to a 5’ end of the single- stranded oligonucleotide, the 5’ adapter sequence comprising, in a 5’ to 3’ direction: a non-ligatable 5’ end, a second index primer binding site sequence, and a ligatable 3’ end; wherein at least one of the 3’ adapter sequence and the 5’ adapter sequence comprises a randomized sequence.

2. The sequencing adapter composition of claim 1, wherein the randomized sequence is 5’ adjacent to the first linker and / or 3’ adjacent to the second index primer binding site sequence.

3. The sequencing adapter composition of claim 1 or claim 2, wherein the 5’ adapter sequence further comprises a second linker sequence.

4. The sequencing adapter composition of any one of claims 1-3, wherein the 3’ adapter sequence comprises in a 5’ to 3’ direction: a ligatable 5’ end, a first randomized sequence, a first linker sequence, a first index primer binding site sequence, and a non-ligatable 3’ end.

5. The sequencing adapter composition of any one of claims 1-4, wherein the 5’ adapter sequence comprises in a 5’ to 3’ direction: a non-ligatable 5’ end, a second index primer binding site sequence, a second randomized sequence, a second linker sequence, and a ligatable 3’ end.

6. The sequencing adapter composition of any one of claims 1-5, wherein the 3’ adapter sequence does not comprise a stem-loop structure sequence.

7. The sequencing adapter composition of any one of claims 1-6, wherein less than 5%, 10%, or 20% of the 3’ adapter sequence is capable of forming a secondary structure.

8. The sequencing adapter composition of any one of claims 1-7, wherein the 3’ adapter sequence is not capable of forming a secondary structure.

9. The sequencing adapter composition of any one of claims 1-8, wherein the ligatable 5’ end of the 3’ adapter sequence comprises a 5 ’-adenylated nucleotide.

10. The sequencing adapter composition of any one of claims 1-9, wherein the ligatable 5’ end of the 3’ adapter sequence comprises a 5 ’-phosphorylated nucleotide.

11. The sequencing adapter composition of any one of claims 1-10, wherein the ligatable 3’ end of the 5’ adapter sequence comprises a 3 ’-hydroxylated nucleotide.

12. The sequencing adapter composition of any one of claims 1-11, wherein the ligatable 5’ end of the 3’ adapter sequence comprises a 5 ’-hydroxylated nucleotide.

13. The sequencing adapter composition of any one of claims 1-8 and 12, wherein the ligatable 3’ end of the 5’ adapter sequence comprises a 3 ’-phosphorylated nucleotide.

14. The sequencing adapter composition of any one of claims 1-13, wherein the non-ligatable 3’ end of the 3’ adapter sequence comprises a 3’ modification of a 3’ nucleotide.

15. The sequencing adapter composition of claim 14, wherein the 3’ modification is selected from the group consisting of: a 3’ amino group, a 3’ dideoxy-C modification, a 3’ inverted dT modification, a 3’ C3 spacer modification, and a 3’ phosphoryl modification.

16. The sequencing adapter composition of any one of claims 1-15, wherein the non-ligatable 5’ end of the 5’ adapter sequence comprises a 5’ modification of a 5’ nucleotide.

17. The sequencing adapter composition of claim 16, wherein the 5’ modification is selected from the group consisting of: a 5’ hydroxyl group, a 5’ dideoxy-C modification, 5’ inverted dideoxy-T modification, and a 5’ C3 spacer modification.

18. The sequencing adapter composition of any one of claims 1-17, wherein the 3’ adapter sequence is or comprises a sequence that is at least about 80%, at least about 85%, at least about 90%, or at least about 95% identical to a sequence that is complementary to the reverse transcription primer.

19. The sequencing adapter composition any one of claims 1-3 and 5-18, wherein the 3’ adapter sequence is or comprises a sequence that is 100% identical to a sequence that is complementary to the reverse transcription primer.

20. The sequencing adapter composition of any one of claims 1-19, wherein the first linker sequence has a length of about 4 nucleotides to about 20 nucleotides.

21. The sequencing adapter composition of any one of claims 1-20, wherein the first linker sequence has a length of 10 nucleotides.

22. The sequencing adapter composition of any one of claims 1-21, wherein the first index primer binding site sequence has a length of about 20 nucleotides to about 35 nucleotides.

23. The sequencing adapter composition of any one of claims 1-22, wherein the first index primer binding site sequence has a length of about 28 nucleotides.

24. The sequencing adapter composition of any one of claims 1-23, wherein the first index primer binding site sequence is capable of binding to a sequencing read 2 adapter.

25. The sequencing adapter composition of any one of claims 1-24, wherein the second index primer binding site sequence has a length of about 20 nucleotides to about 30 nucleotides.

26. The sequencing adapter composition of any one of claims 1-25, wherein the second index primer binding site sequence has a length of about 22 nucleotides to about 28 nucleotides.

27. The sequencing adapter composition of any one of claims 1-26, wherein the second index primer binding site sequence has a length of about 26 nucleotides.

28. The sequencing adapter composition of any one of claims 1-27, wherein the second index primer binding site sequence is capable of binding to a sequencing read 1 adapter.

29. The sequencing adapter composition of any one of claims 4-28, wherein the first randomized sequence has a length of about 1 nucleotide to about 20 nucleotides.

30. The sequencing adapter composition of any one of claims 4-29, wherein the first randomized sequence has a length of at least 2 nucleotides.

31. The sequencing adapter composition of any one of claims 5-30, wherein the second randomized sequence has a length of about 10 nucleotides to about 20 nucleotides.

32. The sequencing adapter composition of any one of claims 5-31, wherein the second randomized sequence has a length of about 15 nucleotides.

33. The sequencing adapter composition of any one of claims 3-32, wherein the second linker sequence has a length of about 4 nucleotides to about 20 nucleotides.

34. The sequencing adapter composition of any one of claims 3-33, wherein the second linker sequence has a length of about 7 nucleotides.

35. The sequencing adapter composition of any one of claims 1-34, wherein the 3’ adapter sequence has a total length of about 20 nucleotides to about 60 nucleotides.

36. The sequencing adapter composition of any one of claims 1-35, wherein the 3’ adapter sequence has a total length of about 30 nucleotides to about 45 nucleotides.

37. The sequencing adapter composition of any one of claims 1-36, wherein the 3’ adapter sequence has a total length of about 38 nucleotides.

38. The sequencing adapter composition of any one of claims 1-37, wherein the 5’ adapter sequence has a total length of about 40 nucleotides to about 60 nucleotides.

39. The sequencing adapter composition of any one of claims 1-38, wherein the 5’ adapter sequence has a total length of about 42 nucleotides to about 55 nucleotides.

40. The sequencing adapter composition of any one of claims 1-39, wherein the 5’ adapter sequence has a total length of about 48 nucleotides.

41. The sequencing adapter composition of any one of claims 1-40, wherein the 3’ adapter sequence comprises a DNA sequence.

42. The sequencing adapter composition of any one of claims 1-41, wherein the 5’ adapter sequence comprises an RNA sequence or a DNA sequence.

43. The sequencing adapter composition of any one of claims 1-42, wherein the 3’ adapter sequence and the 5’ adapter sequence do not comprise substantial complementarity with each other.

44. A method of preparing a sequencing library for a single-stranded template nucleic acid molecule, comprising: a) contacting the single-stranded nucleic acid molecule with the 3’ adapter sequence of any one of claims 1-43 and performing a ligation reaction to ligate the 5’ end of the 3’ adaptersequence to a 3’ end of the single- stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule; b) performing an end modification reaction to modify a 5’ end of the first ligated nucleic acid molecule; c) contacting the first ligated nucleic acid molecule with the 5’ adapter sequence of any one of claims 1-43 and performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule; d) contacting the second ligated nucleic acid molecule with a reverse transcription primer and allowing the reverse transcription primer to hybridize to the 3’ adapter sequence; e) performing a reverse transcription reaction, thereby producing a cDNA molecule comprising the template nucleic acid molecule; f) contacting the cDNA molecule with a first index primer capable of hybridizing to the first index primer binding site sequence, and with a second index primer capable of hybridizing to the second index primer binding site sequence, under conditions that allow hybridization of the first and second index primers to their respective binding site sequences; g) amplifying the cDNA molecule to produce an amplification product; and h) purifying the amplification product to produce the sequencing library for the singlestranded template nucleic acid molecule.

45. The method of claim 44, wherein the end modification reaction comprises a phosphorylation reaction to phosphorylate the 5’ end of the first ligated nucleic acid molecule.

46. The method of claim 44, wherein the end modification reaction comprises an adenylation reaction to adenylate the 5’ end of the first ligated nucleic acid molecule.

47. The method of any one of claims 44-46, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a DNA ligase or an RNA ligase.

48. The method of claim 47, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a truncated version of a T4 RNA ligase 2.

49. The method of any one of claims 44-48, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of a DNA ligase or an RNA ligase.

50. The method of claim 49, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of RtcB ligase or T4 RNA ligase 1.

51. The method of any one of claims 44-50, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a reaction buffer containing about 10% to about 25% (v / v) polyethylene glycol (PEG).

52. The method of any one of claims 44-51, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a reaction buffer containing 25% (v / v) polyethylene glycol (PEG).

53. The method of any one of claims 44-52, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of a reaction buffer containing about 10% to about 25% (v / v) polyethylene glycol (PEG).

54. The method of any one of claims 44-53, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of a reaction buffer containing 25% (v / v) polyethylene glycol (PEG).

55. The method of any one of claims 44-54, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule comprises the use of a reaction buffer containing about 5% to about 15% (v / v) glycerol.

56. The method of any one of claims 44-55, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleicacid molecule to produce a first ligated nucleic acid molecule comprises the use of a reaction buffer containing 10% (v / v) glycerol.

57. The method of any one of claims 44-56, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of a reaction buffer containing about 5% to about 15% (v / v) glycerol.

58. The method of any one of claims 44-57, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule comprises the use of a reaction buffer containing 10% (v / v) glycerol.

59. The method of any one of claims 44-58, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule is carried out at a reaction temperature of about 4°C to about 37°C.

60. The method of any one of claims 44-59, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule is carried out at a reaction temperature of 25°C.

61. The method of any one of claims 44-60, wherein performing the ligation reaction to ligate the 5’ end of the 3’ adapter sequence to a 3’ end of the single-stranded template nucleic acid molecule to produce a first ligated nucleic acid molecule is carried out in variable reaction temperature conditions cycling between 4°C and 37°C.

62. The method of any one of claims 44-61, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule is carried out at a reaction temperature of about 4°C to about 37°C.

63. The method of any one of claims 44-62, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule is carried out at a reaction temperature of 25°C.

64. The method of any one of claims 44-63, wherein performing a ligation reaction to ligate the 3’ end of the 5’ adapter sequence to the 5’ end of the first ligated nucleic acid molecule to produce a second ligated nucleic acid molecule is carried out in variable reaction temperature conditions cycling between 4°C and 37°C.

65. The method of any one of claims 44-64, wherein the method does not comprise a step of cleaving the first ligated nucleic acid molecule or the second ligated nucleic acid molecule.

66. The method of any one of claims 44-65, wherein the single-stranded template nucleic acid molecule is an RNA molecule.

67. The method of any one of claims 44-66, wherein the single-stranded template nucleic acid molecule is a guide RNA (gRNA) molecule.

68. The method of any one of claims 44-67, wherein the single-stranded template nucleic acid molecule is a gRNA molecule recovered from a gRNA-Cas nuclease ribonucleoprotein (RNP).

69. The method of any one of claims 44-67, wherein the single-stranded template nucleic acid molecule is a prime editing guide RNA (pegRNA) molecule.

70. The method of any one of claims 44-66, wherein the single-stranded template nucleic acid molecule is a small interfering RNA (siRNA) molecule.

71. The method of any one of claims 44-65, wherein the single-stranded template nucleic acid molecule is a DNA molecule.

72. A method for detecting and quantifying unique sequences in a population of singlestranded template nucleic acid molecules, comprising: a) preparing a sequencing library for the single-stranded template nucleic acid molecule using the method of any one of claims 44-71 ; b) sequencing the sequencing library using a paired-end sequencing method to generate a plurality of sequence reads corresponding to the single-stranded template nucleic acid molecule; and c) analyzing sequence read data corresponding to the plurality of sequence reads to achieve at least one of:(i) correction of amplification errors;(ii) correction of sequencing errors;(iii) normalization of sequence-based amplification bias; and(iv) detection and quantification of unique sequences present in the population of single- stranded template nucleic acid molecules.

73. The method of claim 72, wherein the population of single- stranded template nucleic acid molecules is a putatively homogenous population.

74. The method of claim 72 or 73, wherein analyzing sequence read data comprises the use of one or more processors.

75. The method of any one of claims 72-74, wherein analyzing the sequence read data comprises aligning paired sequence reads and discarding discrepant read pairs from further analysis.

76. The method of any one of claims 72-74, wherein analyzing the sequence read data comprises grouping sequence reads according to unique molecular identifiers (UMIs) associated with each sequence read.

77. The method of claim 76, wherein a unique molecular identifier (UMI) associated with each sequence read corresponds to the randomized sequence in at least one of the 3’ adapter sequence and the 5’ adapter sequence of any one of claims 1-43.

78. The method of claim 76, wherein a unique molecular identifier (UMI) associated with each sequence read corresponds to the second randomized sequence in the 5’ adapter sequence of any one of claims 5-43.

79. The method of any one of claims 76-78, further comprising determining a consensus sequence for each group of sequence reads identified based on the associated unique molecular identifier (UMI).

80. The method of claim 79, further comprising identifying unique sequences present in the population of single-stranded template nucleic acid molecules based on the determined consensus sequence for each group of sequence reads.

81. The method of claim 79 or 80, further comprising correcting for amplification and / or sequencing errors based on the determined consensus sequence for each group of sequence reads.

82. The method of any one of claims 79-81, further comprising aligning a consensus sequence for each group of sequence reads to a reference sequence to identify variant sequences present in the population of single-stranded template nucleic acid molecules.

83. The method of claim 82, further comprising determining a frequency of sequences present in the population of single-stranded template nucleic acid molecules that are identical to the reference sequence, and / or determining a variant sequence frequency for a variant sequence identified in the population of single-stranded template nucleic acid molecules.

84. The method of any one of claims 72-83, wherein the population of single- stranded template nucleic acid molecules comprise guide RNA (gRNA) molecules.

85. The method of any one of claims 72-84, wherein the population of single- stranded template nucleic acid molecules comprise gRNA molecules recovered from gRNA-Cas nuclease ribonucleoproteins (RNPs).

86. The method of any one of claims 72-84, wherein the population of single- stranded template nucleic acid molecules comprise prime editing guide RNA (pegRNA) molecules.

87. The method of any one of claims 72-83, wherein the population of single- stranded template nucleic acid molecules comprise small interfering RNA (siRNA) molecules.

88. The method of any one of claims 72-83, wherein the population of single- stranded template nucleic acid molecules comprise DNA molecules.

89. A kit comprising: a) the sequencing adapter composition of any one of claims 1-43, and b) at least one of:(i) a ligase;(ii) a buffer;(iii) enzymes and / or reagents for performing an end modification reaction to modify a 5’ end of a nucleic acid molecule;(iv) enzymes and / or reagents for performing nucleic acid amplification; and(v) enzymes and / or reagents for performing paired-end sequencing.

90. A system comprising: a) one or more processors; and b) a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to:1) receive sequence read data for a sequencing library generated using the method of any one of claims 44-71 ; and2) analyze the sequence read data for the sequencing library to achieve at least one of:(i) correction of amplification errors;(ii) correction of sequencing errors;(iii) normalization of sequence-based amplification bias; and(iv) detection and quantification of unique sequences present in a population of single-stranded template nucleic acid molecules.

91. The system of claim 90, wherein analyzing the sequence read data comprises aligning paired sequence reads and discarding discrepant read pairs from further analysis.

92. The system of claim 90 or 91, wherein analyzing the sequence read data comprises grouping sequence reads according to unique molecular identifiers (UMIs) associated with each sequence read.

93. The system of claim 92, wherein a unique molecular identifier (UMI) associated with each sequence read corresponds to the randomized sequence in at least one of the 3’ adapter sequence and the 5’ adapter sequence of any one of claims 1-43.

94. The method of claim 93, wherein a unique molecular identifier (UMI) associated with each sequence read corresponds to the second randomized sequence in the 5’ adapter sequence of any one of claims 5-43.

95. The system of any one of claims 92-94, further comprising determining a consensus sequence for each group of sequence reads identified based on the associated unique molecular identifier (UMI).

96. The system of claim 95, further comprising identifying unique sequences present in the population of single-stranded template nucleic acid molecules based on the determined consensus sequence for each group of sequence reads.

97. The system of claim 95 or 96, further comprising correcting for amplification and / or sequencing errors based on the determined consensus sequence for each group of sequence reads.

98. The system of claim 97, further comprising aligning a consensus sequence for each group of sequence reads to a reference sequence to identify variant sequences present in the population of single- stranded template nucleic acid molecules.

99. The system of claim 98, further comprising determining a frequency of sequences present in the population of single-stranded template nucleic acid molecules that are identical to the reference sequence, and / or determining a variant sequence frequency for a variant sequence identified in the population of single-stranded template nucleic acid molecules.

100. The system of any one of claims 90-99, wherein the single- stranded template nucleic acid molecules in the population comprise guide RNA (gRNA) molecules.

101. The method of any one of claims 90-100, wherein the population of single- stranded template nucleic acid molecules comprise gRNA molecules recovered from gRNA-Cas nuclease ribonucleoproteins (RNPs).

102. The method of any one of claims 90-100, wherein the population of single- stranded template nucleic acid molecules comprise prime editing guide RNA (pegRNA) molecules.

103. The system of any one of claims 90-99, wherein the single- stranded template nucleic acid molecules in the population comprise small interfering RNA (siRNA) molecules.

104. The system of any one of claims 90-99, wherein the single- stranded template nucleic acid molecules in the population comprise DNA molecules.I l l

Citation Information

Patent Citations

  • Mutant polymerases for sequencing and genotyping

    US20070048748A1

  • Methods and apparatus for measuring analytes using large scale FET arrays

    US20090026082A1

  • Error suppression in sequenced DNA fragments using redundant reads with unique molecular indices (UMIS)

    US20160319345A1

  • Methods and compositions for linear isothermal amplification of polynucleotide sequences, using a RNA-DNA composite primer

    US6251639B1

  • DNA polymerase mutant having one or more mutations in the active site

    US6329178B1