Synthetic gene circuits and uses thereof
A high-throughput method using long-read and short-read sequencing screens gene circuit libraries by associating expression units with phenotypes, addressing the challenges of unpredictable interactions and molecular coupling in synthetic gene circuits, thereby reducing the number of DBTL cycles and enhancing the efficiency of circuit design.
Patent Information
- Application Number
- PCT/US2025/020024
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-18
- Filing Date
- 2025-03-14
- Publication Date
- 2025-09-18
AI Technical Summary
The design of synthetic gene circuits is challenging due to unpredictable regulatory interactions and molecular coupling, requiring numerous iterative Design, Build, Test, Learn (DBTL) cycles to achieve desired circuit behavior.
A high-throughput method using long-read and short-read sequencing to screen a gene circuit library, involving the assembly of barcoded nucleic acids, integration into cells, and phenotypic sorting to associate expression units with phenotypes, enabling efficient identification of desired circuit behaviors.
This method significantly reduces the number of DBTL cycles by providing a high-throughput approach to screen and characterize gene circuit libraries, allowing for rapid profiling of complex circuit design spaces and identifying part compositions with desired quantitative behaviors.
Smart Images

Figure US2025020024_18092025_PF_FP_ABST
Abstract
Description
SYNTHETIC GENE CIRCUITS AND U!
[0001] This application is an International Application, which claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 565,400, filed on March 14, 2024, and U.S. Provisional Patent Application No. 63 / 709,252, filed on October 18, 2024, the entire contents of each which are incorporated herein by reference in their entireties.
[0002] All patents, patent applications and publications cited herein are hereby incorporated by reference in their entirety. The disclosures of these publications in their entireties are hereby incorporated by reference into this application in order to more fully describe the state of the art as known to those skilled therein as of the date of the invention described and claimed herein.
[0003] This patent disclosure contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure as it appears in the U.S. Patent and Trademark Office patent file or records, but otherwise reserves any and all copyright rights.GOVERNMENT INTERESTS
[0004] This invention was made with government support under Grant No. N00014-21-1- 4006, awarded by the Department of Defense, Office of Naval Research, and Grant Nos. R01EB029483 and R01EB032272, awarded by the National Institutes of Health. The government has certain rights in the invention.FIELD OF THE INVENTION
[0005] This invention is directed to a high-throughput method for screening a gene circuit library. The pace of gene circuit engineering can thereby be increased by expanding the number of circuits tested in each cycle. By measuring large collections of circuits in a single experiment, such an approach could be used for rapidly profiling complex circuit design spaces to identify part compositions with desired quantitative behaviors.BACKGROUND
[0006] Synthetic gene circuits are constructed by assembling DNA-encoded genetic parts into multi-gene programs that perform computational tasks in living cells. Over the past two decades, gene circuits have emerged as important models for understanding native generegulation and have been used to create powerful biotechn control over cellular behavior. Despite this progress, the design or quaiiuiauveiy precise circuit behavior remains challenging. Regulatory interactions within a circuit must be carefully tuned, often through multiple iterative DBTL (Design, Build, Test, Learn) cycles, before part compositions that support a desired circuit behavior are identified.
[0007] Additionally, since genetic parts must work in close physical proximity to one another, as well as within a crowded intracellular environment, incidental molecular coupling can occur between parts and with host cell regulatory machinery. Because these context-dependent interactions are difficult to predict, they can confound model-driven circuit design, further extending the number of DBTL cycles required to achieve a target behavior.SUMMARY OF THE INVENTION
[0008] Aspects of the invention are directed to a high-throughput method for screening a gene circuit library. In embodiments, the method can comprise preparing an indexed circuit library comprising a plurality of barcoded-linked nucleic acids by combining a first pool of a plurality of nucleic acids with a second pool of a plurality of nucleic acids, wherein each nucleic acid in the first pool comprises one or more expression unit(s) (EU(s)), with each EU comprising at least a promoter, an open reading frame, and a terminator, and wherein each nucleic acid in the second pool comprises a barcode. In embodiments, the method can comprise providing a construct-to-barcode index using long-read sequencing of the index circuit library. In embodiments, the method can comprise providing a barcode-to-phenotype index by integrating the index circuit library into a population of cells, sorting the population of cells according to phenotype, and using short-read amplicon sequencing. In embodiments, the method can comprise comparing the construct-to-barcode index with the barcode-to-phenotype index to associate an expression unit with a phenotype, thereby screening the gene circuit library.
[0009] In embodiments, the method can further comprise assembling the first pool of a plurality of nucleic acids, assembling the second pool of a plurality of nucleic acids or both.
[0010] In embodiments, assembling can be hierarchical.
[0011] In embodiments, the barcodes can be structured.
[0012] In embodiments, the method can further comprise analyzing long-read sequencing data, thereby allowing identification of sequence variants from a single read.
[0013] In embodiments, the long-read sequencing data ca pipeline. For example, the software pipeline can identify circuit components anu uarcoue sequences from individual sequencing data reads.
[0014] In embodiments, the phenotype can comprise a gene expression output.
[0015] In embodiments, method described herein is depicted in FIG. 1.
[0016] In embodiments, the population of cells can comprise bacteria cells, yeast cells, or mammalian cells. For example, the population of cells can comprise HEK293T cells, T cells, mesenchymal stem cells (MSCs), induced pluripotent stem cells (iPSCs), or K562 cells.
[0017] In embodiments, the expression unit (EU) can further comprise a chromatin insulator, enhancer, promoter, Kozak sequence, activation domain (AD), intrinsically disordered region (IDR), 4-hydroxytamoxifen-responsive domain, a zinc finger (ZF) domain, a terminator, a spacer sequence, varied orientation, or any combination thereof.
[0018] In embodiments, the first pool of a plurality of nucleic acids can comprise a diverse pool of nucleic acids.
[0019] In embodiments, the open reading frame region can encode for a small moleculeinducible transcription factor or a fluorescent reporter element.
[0020] Aspects of the invention are further directed to a nucleic acid comprising one or more expression units (EU). In embodiments, each expression unit can comprise a promoter region, an open reading frame region, and a terminator region.
[0021] In embodiments, the nucleic acid further can comprise a chromatin insulator, enhancer, promoter, Kozak sequence, activation domain (AD), intrinsically disordered region (IDR), 4- hydroxytamoxifen-responsive domain, a zinc finger (ZF) domain, a terminator, a spacer sequence, or any combination thereof.
[0022] In embodiments, the open reading frame region can encode for a small moleculeinducible transcription factor or a fluorescent reporter element.
[0023] Aspects of the invention are further drawn to a plasmid vector comprising the nucleic acid described herein.
[0024] In embodiments, the plasmid vector can further comprise an antibiotic resistance cassette, an origin of replication, and restriction enzyme recognition sequences. For example, the restriction enzyme recognition sequences can enable assembly of sub-components into the vector.
[0025] Still further, aspects of the invention are to a cell comprising the plasmid vector described herein.
[0026] Still further, aspects of the invention are drawn to a nucleic acids described herein or a plurality of vectors descriueu Herein.
[0027] Aspects of the invention are further drawn to a drug-inducible gene circuit comprising the nucleic acid described herein.
[0028] In embodiments, the drug-inducible gene circuit described herein can comprise two or more expression units.
[0029] Still further, aspects of the invention are drawn to a high-throughput method of characterizing a gene circuit library. In embodiments, the method can comprise assembling a first pool of a plurality of nucleic acids, wherein each nucleic acid of the first pool comprises an expression unit (EU) comprising at least a promoter, an open reading frame, and a terminator. In embodiments, the method can comprise combining the first pool with a second pool of nucleic acids, wherein each nucleic acid in the second pool can comprise a barcode, thereby providing a nucleic acid comprising at least one expression unit operably linked to a barcode. In embodiments, the method can comprise performing long-read sequencing of the barcoded-linked nucleic acid, thereby providing a construct-to-barcode index. In embodiments, the method can comprise analyzing long-read Nanopore sequencing data, thereby allowing identification of sequence variants from a single read. In embodiments, the method can comprise integrating the barcoded-linked nucleic acid into a plurality of cells. In embodiments, the method can comprise sorting the plurality of cells according to phenotype and performing short-read amplicon sequencing, thereby providing a barcode-to-phenotype index, and comparing the construct-to-barcode index with the barcode-to-phenotype index, thereby linking the expression unit with a phenotype.
[0030] In embodiments, the long-read Nanopore sequencing data is analyzed using a software pipeline. For example, the software pipeline identifies circuit components and barcode sequences from individual Nanopore sequencing data reads.
[0031] In embodiments, the phenotype can comprise a gene expression output.
[0032] In embodiments, the method described herein is depicted in FIG. 1.
[0033] In embodiments, the plurality of cells can comprise bacteria cells, yeast cells, or mammalian cells. For example, the plurality of cells can comprise HEK293T cells, T cells, mesenchymal stem cells (MSCs), induced pluripotent stem cells (iPSCs), or K562 cells.
[0034] Other objects and advantages of this invention will become readily apparent from the ensuing description.BRIEF DESCRIPTION OF THE I
[0035] FIG. 1 provides schematics showing the use of CLASaic, (comoming long- anu snon- range sequencing to investigate genetic complexity) to systematically map the design space of complex genetic programs. Panel A provides a schematic showing the overview of CLASSIC. Pooled assembly of genetic parts with DNA barcode sequences yields libraries of barcode- indexed constructs of arbitrary length and complexity. Long-read nanopore sequencing is used to create an index matching construct composition to an associated barcode. In parallel, libraries introduced into cells undergo sorting or selection to bin expression phenotypes. Barcode amplicons generated for each bin are subjected to short-read NGS to quantify expression phenotype, which is then mapped to construct composition via barcode indexes. Panel B provides a schematic showing an application of CLASSIC to profile a synthetic gene circuit design space. Hierarchical golden gate assembly is used to compose libraries of multi- EU circuits with combinatorially varied part compositions and circuit designs. Sequence fragment pools for different parts categories from level 0 are combined to yield level 1 pools of promoters, open reading frames (ORF), and terminators (term.), which are then combined to yield level 2 EUs (square brackets: fragment / part pools). Barcode pools are combined with EU pools to create indexed multi-EU circuit libraries (level 3) that are integrated into HEK293T-LP cells placed at the AAVS1 locus in chromosome 19 via expression of the BxBl recombinase (top right). Library analysis by a combination of nanopore and flow-seq yields composition-to-function mapping.
[0036] FIG. 2 provides data showing that CLASSIC can quantitatively profile diverse compositions of genetic parts. Panel A shows expression unit (EU) library design space. Left: the EU library comprises combinations of parts from 3 categories (promoters, Kozak sequences, and terminators). Right: Part pools are assembled (step 1) with an mRuby ORF to generate a 384-member EU pool, which is combined with a barcoded-BFP EU pool to generate the indexed EU library. Panel B shows EU library assembly and indexing balance. Oxford nanopore sequencing was performed on the assembled library. Data were analyzed to assess library composition count (top) and composition / barcode balance [reads per composition (light grey), unique barcodes per composition (dark grey); data plotted in rank order of reads per composition] (bottom). Panel C shows EU library expression quantification. Top left: the library was flow sorted into 10 equally log-spaced bins [empty HEK293T cells (grey histogram), library (pink histogram)]. Top right: residual for FACS-measured geometric mean fluorescence values for sorting-isolated clones (n=15) plotted against CL AS SIC-derived values[ERCH (grey band), error range from clonal heterogeneity]; bRepresentative data from 4 clonal isolates (pink solid) aie snown wun corresponding CL AS SIC-computed distributions [kernel density (black line) calculated from normalized barcode read count (grey vertical bars)]; MAE, mean absolute error; AU, arbitrary fluorescence units. Panel D shows CLASSIC precision. Correlation of CL AS SIC-computed EU expression between technical replicate (top) and biological replicate (bottom) experiments. Panel E shows influence of part identity on expression. Violin plots of each part-specific distribution, ordered from strongest to weakest. Dotted line, HEK293T background expression mean. Panel F shows RF analysis of EU behavior space. Left: RF modeled values plotted against CLASSIC measurements. Grey, training data; purple, test data. Top right: 10-fold cross-validated error (y axis) vs number of learning cycles (x axis). CV, cross-validation; MSE, mean squared error. Bottom right: feature importance scores for each part category. Panel G shows analysis of part interference. Left: error [logio(RF predicted) - logio(CLASSIC-observed)] is plotted for all compositions associated with each promoter. White lines, median values; red dots, outliers (error >0.025). Right: outlier part configurations. Red shaded area highlights common terminator T8.
[0037] FIG. 3 provides data showing the use of CLASSIC to profile a synthetic gene circuit design landscape. Panel A provides a schematic showing inducible synTF circuit diagram. Left: The circuit contains two EUs: one codes for the synTF and the other is an eGFP reporter. Without inducer, the synTF localizes to the cytoplasm. Upon addition of 4-OHT (input), the synTF translocates into the nucleus to bind to and activate reporter eGFP expression (output). Right: Expression fold change is the ratio of expression levels in the presence or absence of inducer. Panel B shows synTF circuit design space. Left: SynTF diversity (arranged N-to-C term): of 4 ADs, 4 IDPs, and 3 ZF affinity mutants (high: WT, med: AC, low: AC 6x). synTF coding EU diversity: 4 constitutive promoters, 4 terminators. Reporter EU diversity: 4 BM number variants, 3 minimal core promoters. This unit contains a constant terminator. Both EUs have 3’ spacing sequences of Obp, 250bp or 500bp downstream. Right: EUs assembled from input parts and combined in 2 different 5’-to-3’ orientations, to generate a combinatorial diversity of 165,888 possible circuit variants. Panel C provides data showing the balance of circuit library assembly and indexing. Left: data for ~121k out of 166k variants (73%) were recovered by CLASSIC; compositions, light grey; barcodes, dark grey. Right: Nanopore sequencing data of the library were analyzed to assess library composition count (top) and composition / barcode balance [reads per composition (light grey), unique barcodes percomposition (dark grey); data plotted in rank order of reads ]D shows circuit library sorting and measurement. Top right: e rrr expression 01 circuit iiurary in presence (solid green) and absence (dotted green line) of 4-OHT (top, center) is shown, along with boundaries of flow sorting bins (vertical grey dotted lines); grey histogram, empty HEK293T-LP cells. Top middle: Fold-change values for 40 clonal isolates plotted against CL AS SIC-derived values. Black dots, isolates displayed at the bottom of the panel; MAE, mean average error; grey region, ERCH, error range of clonal heterogeneity AU, arbitrary fluorescence units. Top right: % similarity between basal and induced values for CLASSIC and isolates. Bottom: Flow data from 4 isolates (green dotted line, uninduced; green solid, induced) are shown with corresponding CLAS SIC-computed distributions [kernel density (black line) calculated from normalized barcode read count (grey vertical bars)]. Parts combinations corresponding to each index are shown.
[0038] FIG. 4 provides data showing analysis of CLASSIC library behavior space reveals gene circuit design rules. Panel A shows ML model of inducible synTF circuit behavior space. Left: Basal and induced CLASSIC measurements were plotted as a contour plot (contours, light to dark: 97.5%, 90%, 70%, 50% of total measurements). Highlighted regions bounded by dotted lines: low basal (<500 AU), purple arrow; high induction (>70k AU) blue; high fold change (HFC) (>25x, green). Values in the plot indicate the number of compositions in each region. Middle: RF of with 80:20 traimtest split for basal and induced CLASSIC data [training data (grey) plotted against test data (purple, basal; blue, induced). Right: Contour plot of RF modeled design space for unmeasured compositions predicted by the RF model, measured circuits, and combined data (complete design space). Inset: average model error (Manhattan distance) between the logio(CL AS SIC-derived measurements) and loglO(RF-computed values). Panel B provides data showing experimentally validating ML prediction of circuit function. Top left: Fold-change values for cell lines from unmeasured configurations (red) and measured configurations (teal) were plotted against CLAS SIC-derived values. Black dots, isolates displayed at the periphery of the panel; MAE, mean average error; grey region, ERCH, error range of clonal heterogeneity; green box, HFC behavior region. Right and bottom: Flow cytometry analysis of cell line behavior (green dotted line, basal; solid green, induced) and associated circuit compositions (below each plot). Horizontal bars indicate RF predicted fold change. AU, arbitrary fluorescence units. Panel C shows genetic part usage in highlighted regions of behavior space. Part fold enrichment is calculated by dividing observed part occurrence by expected part occurrence from a balanced library. Red text, categories with highasymmetry used for cluster analysis. Panel D shows m categories in different regions of behavior space. MI between pan categories is uenoieu uy l eu line thickness. Panel E shows clustering analysis of HFC circuit designs. High asymmetry part categories were chosen for UMAP dimensional reduction followed by K-means clustering. Bar plots denote number of part occurrences within each cluster. Panel F provides strategies for engineering synTF circuits with high fold change behavior involve combining design elements that maximize induction while limiting leaky basal expression.
[0039] FIG. 5 provides data showing 384-member library balance and barcoding analysis. Panel A provides confusion matrices showing the percentage of reads unambiguously assigned to each part (diagonal) and those ambiguously linked to two parts (off diagonal) in the level 2 (top) and level 3 (bottom) assembled libraries. Individual values representing <0.5% of the total reads are not shown. The percentage of reads in which a part identity could not be determined are shown at the bottom of the respective part confusion matrices. Panel B provides a bar graph showing total number and percentage barcodes that uniquely map to a single composition (“unique”) or multiple compositions (“non-unique") as determined from level 3 library nanopore sequencing analysis.
[0040] FIG. 6 provides data showing 384-member library short-read data processing and measurement validation. Panel A (top) provides data showing testing three approaches for aggregating data from multiple barcodes indexed to the same composition, (top right) scatter plots showing mRuby fluorescence geometric mean values (x-axis) and expression values computed using the three methods (y-axes). Shaded gray region: ERCH, MAE: Mean absolute error, (bottom) bar plots showing MAE values for all methods for expression conversion tested for all three libraries (lib. 1, lib. 2 sorts 1 and 2). Panel B provides expression histograms for all 15 clonally isolated populations measured by flow cytometry (top, pink) and corresponding barcode measurements (bottom, gray). Panel C provides data showing mean absolute error (MAE) as a function of the number of barcodes used for estimating mean expression. Gray region: standard deviation from n = 50 iterations; dotted blue line: error from clonal heterogeneity. Panel D provides data showing MAE as a function of barcode read depth. Gray region: standard deviation from n = 50 iterations, dotted blue line: error from clonal heterogeneity. Panel E provides a scatter plot showing expression levels for all barcodes from lib. 2 sorts 1 and 2.
[0041] FIG. 7 provides heat maps comparing 384-mei combination. Compositions are grouped by promoters into o mamces. ROWS represent terminators T1-T8; columns represent Kozak sequences K1-K6.
[0042] FIG. 8 provides data showing 166k-member library balance and barcoding analysis metrics. Panel A provides confusion matrices showing the percentage of reads unambiguously assigned to each genetic part (on-diagonal) and reads ambiguously linked to two parts (off- diagonal) following level 3 library assembly. Individual values representing <0.5% of the total reads are not shown. The percentage of reads in which a part identity could not be determined are shown at the bottom of the respective part confusion matrices. Panel B provides a bar graph showing number of barcodes determined to be uniquely mapped to a single composition (“unique”) or multiple compositions (“non-unique”), as determined from nanopore sequencing analysis of the level 3 library. Number of unique (BC1, BC2), and combined (BC1-BC2) pairs identified in the library.
[0043] FIG. 9 provides data showing analysis of clonally isolated cell lines used for validation of CLASSIC. Flow cytometry was used to measure GFP expression for 40 isolated clonal populations harboring integrated circuit compositions from the 166k-member library design space. Upper plots of each panel show basal (green dotted line) and induced (solid green) states with associated expression fold change (green text). Lower plots show corresponding barcode events from CLASSIC basal (gray bars) and induced (light green bars) measurements for all barcodes associated with the circuit, along with computed barcode distribution KDEs (dotted line, induced KDE; solid green lines, basal KDEs), and associated expression fold change values (grey text). Composition identity is shown below each set of plots. Black numbered circles: circuit compositions shown in Figure 3D.
[0044] FIG. 10 provides data showing part usage in circuit compositions not measured by CLASSIC. Genetic part frequencies for ~45k circuit compositions for which basal and / or induced CLASSIC measurements could not be obtained, represented as the fold enrichment over the expected part usage.
[0045] FIG. 11 provides data showing summary of cell lines constructed from the 166k-library design space to validate predictions from the RF model. Cell lines were constructed by integrating assembled circuit compositions into HEK-LP cells and tested for GFP expression by flow cytometry. Upper plots of each panel show basal (green dotted line) and induced (solid green) states with associated expression fold change (green text). Corresponding RF-predicted basal and induced GFP expression levels for each composition are represented as red (out-of-sample) or teal (in-sample) bars along with the calculated fol shown below each set of plots. Panel A provides 15 compositions 101 WHICH uasai anu / or induced CLASSIC measurements were not collected (i.e., out-of-sample). Panel B provides 81 compositions for which basal and / or induced CLASSIC measurements were collected (i.e., in-sample). Panel C provides basal and induced GFP expression levels for each constructed circuit composition overlaid onto the behavior space (clonally isolated, green; constructed out- of-sample, red; constructed in-sample, teal). High- lighted regions bounded by dotted lines: low basal (<500 AU), purple arrow; high induction (>70k AU) blue; high fold change (HFC) (>25x, green). The grey region represents the contour for 97.5% of compositions in RF- predicted behavior space (as in Fig. 4A). Panel D provides data showing comparing foldchange values for RF model predicted values (dark red) with CLASSIC (green) and isolate (from Fig. 3D) or cell line (from Fig. 4B) measurements (dark blue).
[0046] FIG. 12 provides data showing genetic part usage and coupling across different regions of behavior space for the 166k library. Panel A provides data showing genetic part frequencies for circuit compositions predicted by the RF to be low basal (purple, n = 57,505), high induced (navy blue, n = 10,493), HFC (green, n = 5,706), non-responsive (brown, n = 5,706), and constitutively high (pink, n = 5,706) regions of behavior space, as represented as the fold enrichment over the expected (i.e., equal) part usage. Red part category labels represent categories in which there is substantial asymmetry in part usage between regions. Panel B provides plots showing pairwise MI (in bits) between genetic part categories indicate the degree of coupling between part categories in different regions of the behavior space. The width of the red lines connecting part categories are proportional to the MI between the two part categories. Panel C provides plots showing pairwise relative MI between genetic part categories to show the degree of coupling between part categories in different regions of the behavior space, accounting for the different part frequencies in each region. Red labels indicate categories with high asymmetry in part usage between regions.
[0047] FIG. 13 provides data showing clustering analysis of HFC compositions. Panel A provides cluster evaluation plot for k-means clustering using the 7 part categories showing the greatest usage asymmetry (left) and UMAP projection of the three resulting clusters (right). Panel B provides data showing mapping of basal and induced GFP expression of compositions from each of the three clusters, overlaid on a contour constructed from 97.5% of the data from the RF-predicted behavior space (see Figure 4A) (grey fill). Panel C provides data showing distribution of basal (dotted line) and induced GFP expression (solid line) (bottom axis), aswell as fold change values (grey line, top axis) for each clu: values and boxes represent 25th to 75th percentiles. Panel D pioviues a uai cnan snowing me number of individually constructed compositions within each cluster.
[0048] FIG. 14 provides a schematic showing an overview of the CLASSIC platform. CLASSIC, a platform that combines long- and short-read sequencing to interrogate genetic complexity, assembles pools of genetic parts into libraries of constructs with combinatorially varied genetic compositions (left). Random association of assembled library members with short barcodes (bottom left) enables indexing of part compositions to barcodes with via long- read nanopore sequencing and data processing using custom software (center). Libraries can be introduced into cells (top center), phenotypically sorted, and then barcodes quantitatively assessed via short-read Illumina sequencing to construct a phenotype-to-barcode index (right). Indexes can then be compared to assign compositions to phenotype (bottom right). The mapping of part design space that is enabled by CLASSIC can be used to identify compositions that support target behavior, understand rules for part context dependence, or training of ML models for circuit behavior (bottom).
[0049] FIG. 15 provides a schematic showing the Sesame Street cloning system. Sesame street is a modular hierarchical cloning system designed to accommodate both arrayed and pooled cloning workflows. Levels in the hierarchy represent pooled golden-gate assembly steps use input plasmids to generate constructs with successively more complex designs, starting with genetic parts and ending with multi-gene programs. Each level uses a different Type Ils enzyme for cloning (colored lines with superimposed arrow, restriction site and directionality of cut), as well as a different set of destination vectors harboring distinct selectable markers (Crt, orange carotenoid; ccdB, negative selection gene; KanR, kanamycin resistance; AmpR, ampicillin resistance; SpectR, spectinomycin resistance). Latin and Greek letters represent distinct 4 bp overhangs.
[0050] FIG. 16 provides timelines for arrayed and pooled cloning workflows using Sesame Street. Days in the timeline are split into two slots, morning (AM) and afternoon (PM), and the activity associated with each time slot is listed under an intersecting dotted black line. The timeline for cloning individual plasmids is depicted on the left, and the timeline for constructing plasmid libraries is depicted on the right.
[0051] FIG. 17 provides a schematic showing Sesame street library diversification strategies. The platform allows library diversification to occur at different stages within the cloning hierarchy. Panel A provides a schematic showing that by diversifying fragment of promoters,libraries can be created to study the organizational basis of g a schematic showing that by introducing diversity at protein uoinams, me mouuiai oasis 01 protein function can be explored. Panel C provides a schematic showing that diversifying the design and arrangement of expression units in a multi -gene construct can enable exploration of the design rules underlying circuit function.
[0052] FIG. 18 provides data showing validation of CLASSIC barcoding strategy. Panel A provides a schematic showing multiple expression units (A through n) and a barcoded (BC1) BFP expression unit (n + 1) assembling into a barcoded (BC2) level 3 destination vector to produce a level 3 plasmid library containing a combined barcode (BC1 + BC2) pool. Panel B provides a schematic depicting two cloning methods for incorporating barcode sequences downstream of a constitutively expressed BFP gene. One method (left) uses PCR amplification of the BFP expression unit by semi-degenerate barcode-containing primer set and auto-ligation of the amplicon. The other (top right) uses golden gate-based replacement of a ccdB cassette with a semi-degenerate barcode sequence. Bottom: Bar chart representation of the number of unique barcodes generated per 105 colonies transformed using each strategy. Panel C provides data showing comparison of two methods for propagation of barcoded constructs following transformation into E. coli. One (top) uses plate-based overnight outgrowth and colony scraping, and the other (bottom) uses overnight outgrowth in liquid culture. Histograms represent the normalized barcode read distribution. Panel D provides data showing quantitative comparison of structured and assessment unstructured barcode strategies. Barcode identification for the semi-degenerate [BBA]6 pool versus a fully degenerate ION pool were compared across three base calling accuracy modes. Left: BFP EUs containing either pool were assembled into a level 3 destination vector along with an mRuby cassette driven by hEFlal or RS V, respectively, and sequenced using nanopore. Right: Stacked bars represent the proportion of nanopore reads corresponding to unambiguously identified barcodes (match, dark grey), miscalls that can be modified to yield a correct barcode (error corrected, grey), or correct barcodes that cannot be identified (error retained, light grey) for three different Guppy basecalling accuracy modes (fast, high accuracy, and super high accuracy). Panel E provides data showing Sanger sequencing validation of combined barcode assembly. Left: traces of the level 2 (BC1) and level 3 destination vector (BC2) barcode pools are shown prior to level 3 assembly. Middle: Sanger sequencing trace for the post-assembly pool with combined BC1 + BC2 pools is shown. Right: Illumina sequencing of the level 2 (BC1) pool depicting the number of unique barcodes and their abundance in the pool.
[0053] FIG. 19 provides a schematic showing the overvi sequencing analysis pipelines. Schematic depicting the CLA^iv pipeline wim canouis io me software modules used to analyze long-read nanopore (bottom) and short-read (right) sequencing data is provided.
[0054] FIG. 20 provides data showing construction and validation of HEK293T landing pad cell line. Panel A provides a timeline for construction and validation of a clonal HEK-LP cell line. Red half arrows and dotted red lines represent genotyping primers and the resulting amplicons, respectively. Panel B provides a schematic depicting the 10-day timeline for generating single-copy integrated cell lines using the HEK-LP system. Data for integration of a destination vector containing a constitutively active BFP expression unit is shown in the absence (left, middle) or presence (left, bottom) of puromycin selection. Red half arrows and dotted red lines represent genotyping primers and the resulting amplicons, respectively. Solid green histograms represent YFP expression in the HEK-LP (top) or HEK-LP + pDest cell lines (bottom), solid blue histograms represent BFP expression in the HEK-LP (top) or HEK-LP + pDest cell lines (bottom), and solid grey histograms represent YFP or BFP expression in HEK293T. Panel C provides data showing 1.5% agarose gel of amplicons (red dotted lines, S7A and S7B) resulting from genotyping PCR carried out on three cell populations: HEK293T (left), HEK-LP (middle), and HEK-LP + pDest (right). Expected amplicon sizes: AAVS1 WT = 1870 bp, LP-AAVS1 = 1451 bp, EFla- YFP = 642 bp, EFla-PuroR = 430 bp. Panel D provides flow cytometry measurements of the mCherry expression from 23 clonally expanded cell populations (light pink histograms) sorted randomly from a cell line generated from LP integration of a destination vector containing a constitutive mCherry expression unit (dark pink histogram). HEK- 293T (grey histogram) shown for comparison. Mean expression level for each population denoted by white bar. Light grey horizontal bar represents the Error Range of Clonal Heterogeneity (ERCH).
[0055] FIG. 21 provides data showing validation of flow-seq technique for CLASSIC library analysis. Panel A provides a composition of a 24-member mCherry EU library containing four promoters (hEFlal, CMV, RSV, hPGK) and six terminators (T1-T6). Panel B provides a schematic outlining an optimized flow-seq protocol for the 24-member library. Panel C provides a histogram representing the full mCherry expression distribution for the 24-member library (light red), split into 10 equally log-spaced bins (top). Cells from bins 3 and 8 (dark red) were sorted, processed, and the barcode regions were sequenced to assess pre- and post-sort processing methods for minimizing cross-bin contamination of errant transcripts. Bar chartrepresenting the fraction of barcodes identified in bins 3 processing strategy (bottom).
[0056] FIG. 22 provides data showing the overview of 384-member EU library construction. Panel A provides a timeline for 384-member library construction process from level 1 part plasmids pools to level 3 library. Panel B provides a schematic of the construction strategy for each level of the assembly process (grey boxes), with number of cells indicated for assembly transformation and plating (brown circles), colony scraping, and HEK-LP transfection (pink circles). The inputs and products for each level are represented to the right. Panel C provides a plot of read length distribution from nanopore sequencing of the level 3 library assembly. Relative proportions of each DNA species are indicated: library members (pink), recircularized level 3 destination vector (dark grey), undigested level 3 destination vector (orange), and contamination from level 1 genetic parts (light grey).
[0057] FIG. 23 provides data showing flow-sequencing of 384-member library. Panel A provides a schematic showing flow-sorting strategy for 384-member library. Panel B provides the number of cells sorted into each mRuby expression bin. Panel C provides BFP and mRuby flow sorting data. Cells were first sorted into a BFP gate (top left) and then sorted into 10 expression bins based on mRuby expression (top right) (empty HEK293T, grey). Cells from each bin were expanded and remeasured for BFP (bottom left) and mRuby (bottom right). Mean expression values for each distribution is listed.
[0058] FIG. 24 provides data comparing accuracy between mechanism-based and RF regression models. Panel A provides data comparing linear regression and an optimized bagging RF model for predicting mRuby expression from EU composition. Equations used to generate the linear regression model are shown (left). Scatter plots of CLASSIC expression level measurements (x-axis) and the regression- (middle) or RF-based (right) expression level predictions (y-axis). Panel B provides data showing analysis of part usage in EU compositions with high absolute error between CLASSIC measurements and the linear regression (top) or RF (bottom) models. Panel C provides a scatter plot of absolute error between CLASSIC measurements and the RF model (x-axis) and the absolute error between the CLASSIC measurements and the linear regression model (y-axis) (best fit line, dotted black line).
[0059] FIG. 25 provides amino acid sequences of synTF domains. A schematic depicting the synTF EU, comprising promoters (n = 4), a constant Kozak sequence, variable strength activation domains (AD, n = 4), intrinsically disordered domains (IDR, n = 4), a constant mRuby ORF fused to an E2 domain with a flexible GS-linker, flag tag (orange), variableaffinity zinc finger DNA-binding domains (DBD, n = 3), and(n = 4) is provided. Amino acid sequences for the components mat comprise me syiiuieuc transcription factor are shown. DNA sequences for all coding and non-coding elements are listed in Appendix A.
[0060] FIG. 26 provides an overview of 166K-member gene circuit library construction. Panel A provides a timeline of library construction process from level 0 part fragments to level 3 libraries. Panel B provides a schematic of the construction strategy for each level of the assembly process (grey boxes), with number of cells indicated for assembly transformation and plating (brown circles, 10 cm plates; rectangles, autoclave trays), colony scraping, and HEK- LP transfection (pink circles). The inputs and products for each level are represented to the right. Panel C provides data showing nanopore sequencing read-length distribution of level 3 166k-member inducible circuit library, with the relative proportion of each identifiable DNA product denoted: library members (pink), re-circularized level 3 destination vector (dark grey), undigested level 3 destination vector (orange), and contamination from level 1 genetic parts (light grey).
[0061] FIG. 27 provides data showing flow-sequencing of 166k-member library. Panel A provides a schematic outlining the previously optimized flow-seq protocol that was used to phenotypically assay the library in either the basal (left, brown cells and dotted green histogram) or induced (right, green cells and green filled histogram) states before pooling the samples and sequencing. Panel B provides the number of cells (in millions) sorted in each GFP expression bin for basal and induced libraries. Panel C provides BFP and GFP flow sorting data. Histograms are shown for BFP (top left) and GFP (top right) expression from the 166k- member basal (dotted line) and induced (solid fill) libraries relative to HEK293T cells (grey). Cells were gated on BFP and then were sorted into the GFP bins, expanded for 72 h with or without inducer, and re-measured. The mean expression values of expanded and re-measured distributions are listed (basal, dark blue or green italic; induced, light blue or green italic).
[0062] FIG. 28 provides data showing validation of random RF modelling. Panel A provides data comparing bagging- (top left) and boosting-based (bottom left) approaches, where decision trees are either trained in parallel (bagging) or sequentially (boosting). Comparison of basal GFP expression for CLASSIC measurements (x-axis) versus corresponding basal GFP expression level RF-predictions (y-axis) for both HPO architectures (right, top and bottom). Panel B provides a plot outlining the impact of CLASSIC data set filtering based on the number of barcodes uniquely associated with each composition, as determined by test set r2. The modelshown in the plot is a non-HPO boosting forest algorithm. Fil barcode per composition (left inset) produces a test set r2 of O.H , anu niienng 101 greater man 12 unique barcodes per composition (top inset) produces a test set r2 of 0.80. The number of compositions available for use in the training set decreases as the barcode filtering threshold increases (right inset). The dark green line represents the average test set r2 from 50 randomly generated testtrain splits, with the green band the associated standard deviation. Specific barcode filtering thresholds are represented by white dots. Panel C provides plots showing the impact of different downsampling strategies using a non-HPO boosting forest after filtering the training set for compositions with greater than 12 unique barcodes: no downsampling (left), downsampling relative to the minority class (middle), and downsampling relative to the majority class (right). Panel D provides data showing hyperparameter scanning for a boosting random forest to maximize test set r2 using a training set constructed from compositions with greater than 12 unique barcodes. Learn rate (x-axis), maximum number of decision splits (x- axis), and number of cycles (nc). Panel E provides a plot comparing the basal GFP expression (left), induced GFP expression (middle), and fold change (right) flow cytometry measurements (x-axis) with RF predictions (y-axis) for 40 clonally isolated cell populations (see Fig 3D). Panel F provides data showing r2 values from linear independent and quadratic interaction regression models on the test set and 40 clonally isolated cell populations (left). Scatter plot showing the expression from the clonal isolates (x-axis) and predicted expression from the regression models (y-axis) for the linear (right, top) and quadratic (right, bottom) models.
[0063] FIG. 29 provides data showing flow cytometry gating strategy. Gating strategy used to assess the fluorescence distributions of all libraries and isolates / constructed cell lines shown throughout the manuscript.
[0064] FIG. 30 provides data showing using CLASSIC to quantitatively profile a synthetic gene circuit design landscape. Panel A provides a diagram showing an inducible synTF circuit. Left: The circuit contains two EUs: one codes for the synTF and the other is an eGFP reporter. Without inducer, the synTF localizes to the cytoplasm. Upon addition of 4-OHT (input), the synTF translocates into the nucleus to bind to and activate reporter eGFP expression (output). Right: Expression fold change is the ratio of expression levels in the presence or absence of inducer. Panel B shows synTF circuit design space. Left: SynTF diversity (arranged N-to-C term): of 4 TAs, 4 IDPs, and 3 ZF affinity mutants (high: WT, med: AC, low: AC 6x). synTF coding EU diversity: 4 constitutive promoters, 4 terminators. Reporter EU diversity: 4 BM number variants, 3 minimal core promoters. This unit contains a constant terminator. Both EUshave 3’ spacing sequences of Obp, 250bp or 500bp down str input parts and combined in 2 different 5’-to-3’ orientations, io generate a coinuiiiaioiiai diversity of 165,888 possible circuit variants. Panel C shows the balance of circuit library assembly and indexing. Left: data for ~121k out of 166k variants (73%) were recovered by CLASSIC; compositions, light grey; barcodes, dark grey. Right: Nanopore sequencing data of the library were analyzed to assess library composition count (top) and composition / barcode balance [reads per composition (light grey), unique barcodes per composition (dark grey); data plotted in rank order of reads per composition] (bottom). Panel D provides circuit library sorting and measurement. Top right: eGFP expression of circuit library in presence (solid green) and absence (dotted green line) of 4-OHT (top, center) is shown, along with boundaries of flow sorting bins (vertical grey dotted lines); grey histogram, empty HEK293T-LP cells. Top middle: Fold-change values for 40 clonal isolates plotted against CL AS SIC-derived values. Black dots, isolates displayed at the bottom of the panel; MAE, mean average error; grey region, ERCH, error range of clonal heterogeneity AU, arbitrary fluorescence units. Top right: % similarity between basal and induced values for CLASSIC and isolates. Bottom: Flow data from 4 isolates (green dotted line, uninduced; green solid, induced) are shown with corresponding CL AS SIC-computed distributions [kernel density (black line) calculated from normalized barcode read count (grey vertical bars)]. Parts combinations corresponding to each index are shown.
[0065] FIG. 31 provides data showing ML-aided mapping of single-input circuit design space reveals gene circuit design rules. (Panel A) Left: Multi Layer Perceptron (MLP) neural network model of inducible synTF circuit behavior space. Circuit part inputs are one-hot encoded, flattened and passed through a 4-layer MLP with 2 output heads that predict basal and induced expression. Right: Description of the ML task. From the full design space (gray) of ~166k compositions, data from 121k compositions was collected (green) and used to train the ML model to re-fit circuit behavior from measured compositions and infer the unmeasured space (pink). (Panel B) Left: MLP predictions on a held out high-quality test set against basal and induced CLASSIC measurements (purple, basal; blue, induced) (r2, Pearson’s r2). Right: Contour plot of MLP modeled basal and induced expression behavior space (contours, light to dark: 97.5%, 90%, 70%, 50% of total measurements). Highlighted regions bounded by dotted lines: low basal (<500 AU), purple arrow; high induction (>70k AU) blue; high fold change (HFC) (>25x, green). Values in the plot indicate the number of compositions in each region. Panel C provides data showing experimentally validating ML prediction of circuit function.Top left: Fold-change values for cell lines from classic meas and unmeasured configurations (white filled, pink outlined) were pioueu agamsi L LAMIL - derived values. Black dots, isolates displayed at the periphery of the panel; MAE, mean average error; grey region, ERCH, error range of clonal heterogeneity; green box, HFC behavior region. Right and bottom: Flow cytometry analysis of cell line behavior (green dotted line, basal; solid green, induced) and associated circuit compositions (below each plot). Horizontal pink bars indicate ML predicted fold change. AU, arbitrary fluorescence units. Panel D provides data showing genetic part usage in highlighted regions of behavior space; left to right: low basal, high induced, HFC. Part fold enrichment is calculated by dividing observed part occurrence by expected part occurrence from a balanced library. Panel E provides data showing feature importance for all part categories based on absolute SHAP values computed in the HFC region. Larger SHAP values indicate greater stronger contribution to the model’s predictions of HFC behavior. Panel F provides data showing mutual information between part categories in different regions of behavior space. MI between part categories is denoted by red line thickness. Panel G provides data showing clustering analysis of HFC circuit designs. Part features from HFC designs were subjected to UMAP dimensional reduction followed by K-means clustering with optimal cluster number determined using the gap test. Top left: UMAP projection of HFC designs. All clusters displayed significant part occurrence similarities in the locus design (top middle), and inter cluster differences in synTF expression level, binding affinity, and activation domain preference (bottom). Right: Mechanism for achieving HFC behavior based on clustering results. Locus features are optimized to reduce basal and maximize induced expression levels, while cluster specific choices in synTF design tend to maximize induced expression. Panel H provides data showing model efficiency and expansion to new features via fine-tuning. Left: The model training set was down-sampled by removing compositions and re-training on smaller percentage of the design space. Middle left: Model performance (Pearson’s r2) on compositions in the high-quality test set, evaluated on basal (purple line) and induced (green line) expression levels as a function of percentage of total compositions removed from the training set. Dark gray box: maximum observed r2± 5%. Middle right: Expansion of the design space by fine-tuning the model on data collected from a library of compositions containing a new part feature (NFZ activation domain). Right: Model performance (test set Pearson’ s r2) for basal and induced expression as a function of increasing number of compositions used in the training set during fine-tuning. Dotted line: number of compositions needed to obtain 95% of maximum observed r2value.
[0066] FIG. 32 provides data showing ML guided explor behavior. Panel A provides data showing multi-input induciuie synir circuit uiagrain. LCH. the circuit contains three EUs: a 4-OHT inducible synTF A (navy), a GZV-inducible synTF B (teal), and a GFP reporter unit (green). As in Fig. 3A, 4-OHT (input A) binds to the ERT2 domain (light blue), allowing synTF A to translocate into the nucleus and activate expression of the GFP (output). GZV (input B) inhibits the proteolytic activity of NS3 (orange), preventing the fragmentation of the synTF B and allowing it to GFP expression (output). Right: the multiinput circuit architecture enables cells to perform Boolean logical operations, including OR (top) and AND-logic (bottom). Panel B shows the genetic components that comprise the multiinput design space, organized into synTF parts (top), coding parts (middle left), and reporter parts (middle right). Schematic of the assembly workflow, from protein domains to full-length gene circuits (bottom), where the numbers inside each plasmid or plasmid pool represent the number of variants included in the base / full design space. Panel C shows the coverage, balance, and indexing for multi-input circuit libraries. Left: circuit (light gray text, bottom) and barcode (dark gray text, top) recovery at each step of the CLASSIC workflow for the base (light blue boxes) and expansion (light gray boxes) libraries. Top right: Nanopore sequencing data of the library to assess composition counts, revealing 7% coverage in the base (light blue histogram) and 0.002% coverage of the expansion (light gray histogram) design spaces. Bottom right: Nanopore sequencing data plotted in rank order by reads per composition (base library: light blue, expansion library: light gray) and barcodes per composition (base library: dark blue, expansion library: dark gray). Panel D shows the circuit library sorting and measurement. Top left: GFP expression histograms for base (dotted line) and expansion (shaded) circuit libraries in the presence of no inducer (gray, top left), 4-OHT only (navy, top right), GZV only (orange, bottom left), or both inducers (green, bottom right), overlaid with the boundaries of flow sorting binds (vertical grey dotted lines) and empty HEK293T-LP cells (light gray histogram). Top right: CLASSIC versus ground-truth GFP expression measurements of randomly isolated compositions from the base (n = 14, left) and expansion (n = 13, right) design spaces for all four inducer conditions (no inducer: gray, 4-OHT only: navy, GZV-only: orange, both inducers: green). Bottom: Flow cytometry histograms (top) and CL AS SIC-computed distributions (bottom, kernel density (solid line) calculated from normalized barcode read count (shaded vertical bars) for one composition from the base library (dark gray fill in scatter) are shown alongside the accompanying composition ID.
[0067] FIG. 33 shows the analysis of digital logic gene cin design space. (Panel A) Left: ML model for predicting mum -input circuit uenavio s ^uasai. light gray, navy: 4-OHT only, orange: GZV only, green: both inducers) using OHE part compositions from base (navy) and expansion (dark gray) library data (left). Right: the architecture of the MLP enables exploration of large design spaces (light gray) by first training on a base library (green, top) to infer a small section of the design space (pink, middle) and then fine-tuning using additional data (green, middle) to infer the entire design space (pink, bottom). (Panel B) Left: scatter plots of the test set r2from the fine-tuned model for each of the four input conditions (basal: light gray, navy: OHT only, orange: GZV only, green: both inducers). Right: compositions from the base (dark gray) and full (light gray) design spaces are plotted as Kullback-Leibler divergences (DKL), which describe the similarity between the model-computed output behaviors and the archetypal AND and OR behaviors. Compositions to the left of the vertical dotted line or below the horizontal dotted line represent the most AND- or OR-like designs, respectively. Panel C provides a scatter plot of predicted versus groundtruth GFP expression measurements for 36 individually constructed compositions from across the behavior space (white dots: basal, navy dots: 4-OHT only, orange dots: GZV only, green dots: both inducers) (MAE = 0.32) (top). Only 1 of the 36 compositions was measured by CLASSIC. The most AND-like (thick red outline) and OR-like (thick purple outline) designs are highlighted in the scatter plot (top) and broken out below (middle and bottom, respectively). Panel D shows genetic part usage in the AND- (left) and OR-like (right) regions of design space, grouped according to the synTF (synTF A: left, synTF B: right). Part fold enrichment is calculated by dividing the observed part occurrence by expected part occurrence from a balance library. Panel E shows feature importance for all part categories based on absolute SHAP values computed in either AND-like (top) or OR-like (bottom) regions. Larger SHAP values indicate greater contribution to the model’s predictions of AND- or OR-like behavior (navy: synTF A, dark gray: synTF coding, teal: synTF B, light blue: synTF B coding, green: reporter). Panel F provides mutual information between part categories in the AND- (left) and OR-like (right) regions of behavior space, as indicated by red line thickness (navy: synTF A, dark gray: synTF coding, teal: synTF B, light blue: synTF B coding, green: reporter). Panel G provides clustering analysis of the AND-like circuit designs. Top left: UMAP projection of AND-like designs (light red: Cluster A, dark red: cluster B, blue: cluster C). Top right: schematic of the most common locus design observed in AND-like compositions (navy: synTF A, green: reporter, teal: synTF B). Bottom left: schematics of the part usage in synTF B (left) and synTFA (middle) in clusters A and B (light red: Cluster A, dark n the dominant strategy for achieving AND-like behavior (right;, DOUOIII ngni; scnemaucs or me part usage in synTF B (left) and synTF A (right) in the minority cluster (blue: cluster C). Panel H provides clustering analysis of the OR-like circuit designs. (Top left) UMAP projection of AND-like designs (blue: Cluster A, light red: Cluster B, dark red: cluster C, dark blue: cluster D). Top right: schematic of the most common locus design observed in AND-like compositions (navy: synTF A, green: reporter, teal: synTF B). Bottom: schematics of the part usage in synTF B and synTF A (left, middle) in all clusters and an overview of the strategy for achieving OR- like behavior (right) (blue: Cluster A, light red: Cluster B, dark red: cluster C, dark blue: cluster D).
[0068] FIG. 34 provides data comparing sampling approaches for navigating vast design spaces using Al. Panel A provides a schematic of two strategies for fine-tuning ML models to expand design spaces. Approach A (left) relies on holding out parts from the design space (light grey region) and experimentally measuring compositions (green region) from the sampled design space (gray dotted region) to train a base model (red region). This model is then finetuned in a second round using small sub libraries (green regions) containing new parts from the initially held-out design space to infer the full design space (light grey region). Approach B (right) involves randomly sampling compositions with no part holdout. For consistency, the same number of total circuits were sampled in both approaches (see Methods). Panel B shows simulating the experiment with data from the 166k-member library, gray: held out parts, BCs: barcodes. Panel C shows model performance for approach A (left) vs approach B (right) for basal (purple) and induced (navy) expression values. Base set: first sampling round, expanded set: second sampling round. Panel D provides data showing Approach A outperforms approach B in all metrics measured, including test set r2 values (top), training time (bottom left), and root mean squared error at the end of training (RMSE, bottom right).
[0069] FIG. 35 shows an overview of construction of the base library for multi-input circuit design space. Panel A provides a timeline of library construction process from level 0 part fragments to level 3 libraries. Panel B provides a schematic of the construction strategy for each level of the assembly process (grey boxes), with number of cells indicated for assembly transformation and plating (brown circles, 10 cm plates), colony scraping, and HEK-LP transfection (10 cm plates, pink circles). The inputs and products for each level are represented to the right. Dotted lines represent picking individual colonies (rather than scraping the entire plate) in order to intentionally limit the downstream diversity (i.e., bottleneck). Panel Cprovides nanopore sequencing read-length distribution of lev proportion of each identifiable DNA product denoted: library ineinueis tpniK , re-cn cuianzeu level 3 destination vector (orange), empty vector (dark grey), contamination from level 0-2 genetic parts (light blue), and unidentified DNA species (dark blue).
[0070] FIG. 36 provides an overview of expansion library for multi-input circuit design space. Panel A provides a timeline of library construction process from level 0 part fragments to level 3 libraries. Panel B, Left, provides a schematic of the construction strategy for each level of the assembly process (grey boxes), with number of cells indicated for assembly transformation and plating (brown circles, 10 cm plates), colony scraping, and HEK-LP transfection (pink circles). The inputs and products for each level are represented to the right. Dotted lines represent picking individual colonies (rather than scraping the entire plate) in order to intentionally limit the downstream diversity (i.e., bottleneck). Transparent green boxes represent libraries containing a new part. Right: schematized overview of fine-tuning process. Panel C provides nanopore sequencing read-length distribution of the level 3 expansion library, with the relative proportion of each identifiable DNA product denoted: library members (pink), re-circularized level 3 destination vector (orange), empty vector (dark grey), and unknown DNA species (dark blue).DETAILED DESCRIPTION OF THE INVENTION
[0071] Detailed descriptions of one or more embodiments are provided herein. It is to be understood, however, that the present invention may be embodied in various forms. Therefore, specific details disclosed herein are not to be interpreted as limiting, but rather as a basis for the claims and as a representative basis for teaching one skilled in the art to employ the present invention in any appropriate manner.
[0072] The singular forms “a”, “an” and “the” include plural reference unless the context clearly dictates otherwise. The use of the word “a” or “an” when used in conjunction with the term “comprising” in the claims and / or the specification may mean “one,” but it is also consistent with the meaning of “one or more,” “at least one,” and “one or more than one.”
[0073] Wherever any of the phrases “for example,” “such as,” “including” and the like are used herein, the phrase “and without limitation” is understood to follow unless explicitly stated otherwise. Similarly, “an example,” “exemplary” and the like are understood to be nonlimiting.
[0074] The term “substantially” allows for deviations negatively impact the intended purpose. Descriptive terms are unue sioou io ue inouineu uy the term “substantially” even if the word “substantially” is not explicitly recited.
[0075] The terms “comprising” and “including” and “having” and “involving” (and similarly “comprises”, “includes,” “has,” and “involves”) and the like are used interchangeably and have the same meaning. Specifically, each of the terms is defined consistent with the common United States patent law definition of “comprising” and is therefore interpreted to be an open term meaning “at least the following,” and is also interpreted not to exclude additional features, limitations, aspects, etc. Thus, for example, “a process involving steps a, b, and c” means that the process includes at least steps a, b and c. Wherever the terms “a” or “an” are used, “one or more” is understood, unless such interpretation is nonsensical in context.
[0076] The term “about” is used herein to mean approximately, roughly, around, or in the region of. When the term “about” is used in conjunction with a numerical range, it modifies that range by extending the boundaries above and below the numerical values set forth. In general, the term “about” is used herein to modify a numerical value above and below the stated value by a variance of 20 percent up or down (higher or lower).
[0077] The term "gene" can refer to a nucleic acid sequence comprising an open reading frame encoding a polypeptide, including both exon and (optionally) intron sequences. A "gene" can refer to coding sequence of a gene product, as well as non-coding regions of the gene product, including 5'UTR and 3'UTR regions, introns and the promoter of the gene product. A "gene", as used herein, is the segment of nucleic acid (typically DNA) that is involved in producing a polypeptide or ribonucleic acid gene product. It includes regions preceding and following the coding region (leader and trailer) as well as intervening sequences (introns) between individual coding segments (exons). This term can also include the necessary control sequences for gene expression (e.g. enhancers, silencers, promoters, terminators etc.), which may be adjacent to or distant to the relevant coding sequence, as well as the coding and / or transcribed regions encoding the gene product. These definitions generally can refer to a singlestranded molecule, but in specific embodiments will also encompass an additional strand that is partially, substantially or fully complementary to the single-stranded molecule. Thus, a nucleic acid sequence can encompass a double-stranded molecule or a double-stranded molecule that comprises one or more complementary strand(s) or "complement(s)" of a particular sequence comprising a molecule. As used herein, a single stranded nucleic acid canbe denoted by the prefix "ss", a double stranded nucleic acii stranded nucleic acid by the prefix "ts."
[0078] The term "expression" as used herein can refer to the biosynthesis of a gene product, preferably to the transcription and / or translation of a nucleotide sequence, for example an endogenous gene or a heterologous gene, in a cell. For example, in the case of a heterologous nucleic acid sequence, expression involves transcription of the heterologous nucleic acid sequence into mRNA and, optionally, the subsequent translation of mRNA into one or more polypeptides.
[0079] The terms “decrease” "reduced", "reduction", or "inhibit" as used herein can refer to a decrease by a statistically significant amount. In some embodiments, "reduce," "reduction" or "decrease" or "inhibit" can refer to a decrease by at least 10% as compared to a reference level (e.g. the absence of a given treatment) and can include, for example, a decrease by at least about 10%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or more.
[0080] The terms "increased", "increase", "enhance", or "activate" as used herein can refer to an increase by a statically significant amount. In some embodiments, the terms "increased", "increase", "enhance", or "activate" can refer to an increase of at least 10% as compared to a reference level, for example an increase of at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 70%, or at least about 80%, or at least about 90% or up to and including a 100% increase or any increase between 10-100% as compared to a reference level, or at least about a 2-fold, or at least about a 3 -fold, or at least about a 4-fold, or at least about a 5 -fold or at least about a 10- fold increase, or any increase between 2-fold and 10-fold or greater as compared to a reference level.
[0081] The term "amino acid" can refer to naturally occurring L a-amino acids or residues. The commonly used one and three letter abbreviations for naturally occurring amino acids are used herein: A=Ala: C=Cys; D=Asp; E=Glu; F=Phe; G=Gly; H=His; I=Ile; K=Lys; L=Leu; M=Met; N=Asn; P=Pro; Q=Gln; R=Arg; S=Ser; T=Thr; V=Val; W=Trp; and Y=Tyr (Lehninger, A. L., (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" can further refer to D-amino acids, retro-inverso amino acids as well as chemically modified amino acids such as amino acid analogues, naturally occurringamino acids that are not usually incorporated into proteins si synthesized compounds having properties known in the art io ue cnaracieiisuc or an ammo acid, such as P amino acids.
[0082] The term "peptide" as used herein (e.g., in the context of a synTF) can refer to a plurality of amino acids joined together in a linear or circular chain. The term oligopeptide is typically used to describe peptides having between 2 and about 50 or more amino acids. Peptides larger than about 50 amino acids are often referred to as polypeptides or proteins. For purposes of the present disclosure, however, the term "peptide" is not limited to any particular number of amino acids, and is used interchangeably with the terms "polypeptide" and "protein".
[0083] The terms "nucleic acids" and "nucleotides" can refer to naturally occurring or synthetic or artificial nucleic acid or nucleotides. The terms "nucleic acids" and "nucleotides" can comprise deoxyribonucleotides or ribonucleotides or any nucleotide analogue and polymers or hybrids thereof in either single- or double-stranded, sense or antisense form. As will also be appreciated by those in the art, many variants of a nucleic acid can be used for the same purpose as a given nucleic acid. Thus, a nucleic acid can also encompass substantially identical nucleic acids and complements thereof.
[0084] The term "nucleic acid sequence" or "oligonucleotide" or "polynucleotide" are used interchangeably herein and can refer to at least two nucleotides covalently linked together. The term "nucleic acid sequence" is also used inter-changeably herein with "gene", "cDNA", and "mRNA". As will be appreciated by those in the art, the depiction of a single nucleic acid sequence also defines the sequence of the complementary nucleic acid sequence. Thus, a nucleic acid sequence also encompasses the complementary strand of a depicted single strand.
[0085] The term "oligonucleotide" as used herein can refer to an oligomer or polymer of ribonucleic acid (RNA) or deoxyribonucleic acid (DNA) or mimetics thereof, as well as oligonucleotides having nonnaturally-occurring portions which function similarly. An oligonucleotide can include two or more nucleomonomers covalently coupled to each other by linkages (e.g., phosphodiesters) or substitute linkages.
[0086] Aspects of the invention are drawn to a platform for high-throughput construction and single cell measurement of large and complex synthetic genetic compositions. Using a custom catalogue of plasmid vectors and type Ils restriction enzyme overhangs, we demonstrate the ability to use hierarchical golden gate assembly to efficiently construct large (~10A5) pools of barcoded multi-gene DNA constructs from donor plasmids harboring genetic sub-components. Through the combinatorial assembly of pools of genetic sub-componentswith a pool of short, semi-degenerate barcode sequences, v and flexibly generate highly diverse barcoded genetic composition nuianes. inese nuraiies are sequenced using long-read sequencing (Oxford Nanopore) to generate a barcode-to- composition map. In parallel, the libraries are introduced into cells and can be assayed using new or existing high-throughput phenotypic characterization techniques, such as flow-seq or bulk RNA-sequencing, to map each barcode to a behavior. The two datasets can then be united to yield a genotype-to-phenotype index.
[0087] As described herein, aspects of the invention are further drawn to a high-throughput method of characterizing a gene circuit library. See, for example, the CLASSIC (Combining Long- And Short-range Sequencing to Investigate genetic Complexity) platform as exemplified herein. The CLASSIC platform has utility across diverse organismal contexts and application spaces. While the CLASSIC platform can be applied to engineer strong, constitutive expression cassettes and optimize the fold-change behavior of genetic switches in mammalian cells, the applications extend beyond basic research and biotechnology. For example, ongoing work in the therapeutic space involves the use of CLASSIC to develop chimeric antigen receptors (CARs) that lead to more effective and persistent immune cell phenotypes. Additionally, the CLASSIC platform can be used to generate datasets that can be used to answer complex questions about how neighboring regulatory elements, both proximal and distal, work together to coordinate gene regulatory patterns. Finally, the CLASSIC platform can be applied to other organisms, such as yeast or bacteria, for the construction and optimization of genetic circuits for bioproduction, sensing, or therapeutic applications.
[0088] Accordingly, aspects of the invention described herein are drawn to a high- throughput method for screening or characterizing a gene circuit library.
[0089] A “gene circuit” can refer to an assembly of genetic components or biological parts that can be assembled in various configurations. Gene circuits can encode RNA or proteins that enables individual cells perform specific biological functions, such as gene expression control, signal processing, mimicking regulatory networks, or responding to or interacting with each other to perform some logical functions. Gene circuit libraries can be used in synthetic biology and genetic engineering to construct custom biological systems for applications like biosensing, drug production, or bioengineering. Accordingly, in embodiments, the gene circuit can be a drug-inducible gene circuit.
[0090] As described herein, embodiments of the invention comprise preparing an indexed gene circuit library by combining a first pool of a plurality of nucleic acids with a second poolof a plurality of nucleic acids, wherein each nucleic acic expression unit (EU) comprising at least a promoter, an open reaumg name, anu a leiininaioi, and wherein each nucleic acid in the second pool comprises a barcode.
[0091] A “nucleic acid” or “nucleic acid molecule” can refer to a DNA molecule, an RNA molecule, or a molecule that has been artificially constructed. A nucleic acid molecule can be single-stranded or double-stranded. In embodiments, the nucleic acid molecule is single stranded. In embodiments, the nucleic acid molecule is double-stranded DNA. In embodiments, the nucleic acid can comprise synthetic DNA fragments generated through chemical or enzymatic synthesis.
[0092] Each nucleic acid in the first pool of the plurality of nucleic acids comprises one or more expression unit(s) (EU(S)).
[0093] An “expression unit” (EU) or a “circuit / expression unit” can refer to a functional unit of gene expression, such as a set of genetic components needed to produce a specific transcript. In embodiments, the transcript can be a nucleic acid. In embodiments, the transcript can encode protein. For example, the transcript can be a signaling molecule, a transcription factor, a receptor, a reporter molecule, and the like.
[0094] In embodiments, the transcript can comprise a synthetic molecule. A “synthetic molecule” can refer to any non-naturally occurring compound that arises from human engineering. In embodiments, the synthetic molecule can be partially synthetic or fully synthetic. In some embodiments, the synthetic molecule is produced in vitro. In other embodiments, the synthetic molecule is produced in vivo (i.e., in a living organism). In some embodiments, the synthetic molecule is produced in vivo and is subsequently isolated and purified. In some embodiments, the synthetic molecule comprises an amino acid sequence (i.e., a peptide component), a nucleic acid component, or both. The skilled artisan will recognize that any synthetic molecule can be utilized herein
[0095] For example, the synthetic molecule can be a synthetic transcription factor (sTF), which refers to an engineered DNA binding protein that targets a specific DNA sequence and can activate or repress gene expression. In this example, the gene circuit can comprise an expression unit (EU) encoding a synthetic transcription factor. The skilled artisan will recognize that any synthetic transcription factor can be utilized herein, non-limiting examples of which include MyoD, Ngnl, Ngn2, Tbx family members, or any combination thereof.
[0096] In embodiments, the transcription factor can be a small molecule inducible transcription factor. The term “small molecule-inducible transcription factor” or“chemically inducible transcription factor” can refer to expression in response to a small molecule or chemical compounu. inuuciuie gene swiicnes can allow for precise regulation of biological systems.
[0097] In embodiments, a gene circuit can comprise one or more Expression Units. For example, the gene circuit comprises one EU, two EUs, three EUs, four EUs, five EUs, six EUs, seven EUs, eight EUs, nine EUs, ten EUs, or more than ten EUs.
[0098] In embodiments, an expression unit can comprise one or more genetic elements, such as transcriptional / translational regulatory elements. In embodiments, the transcriptional / translational regulatory elements can comprise a promoter and / or a terminator, and an open reading frame. In embodiments, the open reading frame can encode a protein such as, but not limited to, a secreted therapeutic, surface modification, regulatory genetic circuit, or reporter. In embodiments, the expression unit can express a secreted therapeutic, a surface protein, a regulatory genetic circuit, a reporter protein, transcription factors, pro-apoptotic proteins, or any combination thereof.
[0099] In embodiments, an expression unit can comprise one or more genetic elements. A genetic element can refer to a distinct unit of genetic material, usually a segment of DNA, that carries genetic information and can be passed on from one generation to the next. Non-limiting examples of genetic elements can comprise a promoter or promoter region, an open reading frame region, a terminator or terminator region, a chromatin insulator, enhancer, Kozak sequence, activation domain (AD), intrinsically disordered region (IDR), 4-hydroxytamoxifen- responsive domain, a zinc finger (ZF) domain, grazoprevir-responsive domains, mRNA stability elements, a spacer sequence, transposons, regulatory sequences, insertion sequences, integrons, gene cassettes, varied orientation, or any combination thereof.
[0100] The skilled artisan will recognize the modularity of the expression units. However, without wishing to be bound, exemplary nucleic acid sequences for contemplated genetic elements are provided in Example 4 and Table 1. Further contemplated are variants of the disclosed nucleotide sequences with a sequence identity of at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% of the sequences.Herein, “sequence identity“, “sequence homology”, or “sequence similarity” can refer to the similarity between amino acid or nucleic acid sequences, which is expressed as the similarity between the sequences. Sequence identity can be frequently measured as percent identity, in which two sequences are considered more similar the higher the percentage. Homologs or variants of a polypeptide or nucleic acid molecule possess a relatively high degreeof sequence identity when aligned using standard methods, Venclovas, Methods for Sequence- Structure Alignment in Homoiogy iviouenng. iviemous anu Protocols, 55-82 (Andrew Orry and Ruben Abagyan, eds., 2012)).
[0101] The term "substantially complementary", when used herein with respect to a nucleotide sequence in relation to a reference or target nucleotide sequence, means a nucleotide sequence having a percentage of identity between the substantially complementary nucleotide sequence and the exact complementary sequence of said reference or target nucleotide sequence of at least 60%, at least 70%, at least 80% or 85%, at least 90%, at least 93%, at least 95% or 96%, at least 97% or 98%, at least 99% or 100% (the latter being equivalent to the term "identical" in this context).
[0102] The term “promoter” or “promoter region” can refer to a region of DNA upstream of a gene where relevant proteins (such as RNA polymerase and transcription factors) bind to initiate transcription of that gene. The resulting transcription produces an RNA molecule (such as mRNA). A promoter can also contain sub-regions at which regulatory proteins and molecules may bind, such as RNA polymerase and other transcription factors. Promoters may be constitutive, inducible, activatable, repressible, tissue-specific or any combination thereof. In some embodiments, the promoter is constitutive. In some embodiments, the promoter is inducible. In some embodiments, the promoter is a mammalian promoter. Promoters can be located near the transcription start sites of genes, on the same strand and upstream on the DNA. The skilled artisan will recognize that any promoter or combinations thereof can be utilized herein, including those described in Example 4 and Table 1. Non-limiting examples of promoters can comprise hEFlal, hEFla2, CMV, RSV, hPGK, SV40, sFFV, mPGK, CAG, or any combination thereof.
[0103] The term “open reading frame” or “open reading frame region” can refer to a portion of a DNA sequence that can be transcribed and / or translated. For example, the Expression Unit can comprise an open reading frame that encodes for a synthetic molecule. In some embodiments, an open reading frame does not include a stop codon. In some embodiments, an open reading frame can refer to a sequence of adjacent codons that starts with a start codon, followed by a series of codons for amino acids, and ends with a stop codon. This sequence has the potential to be translated into a polypeptide product. The term “codon” can refer to a DNA or RNA sequence of three nucleotides (a trinucleotide) that forms a unit of genomic information encoding a particular amino acid or signaling the termination of protein synthesis. A “start codon” can refer to a specific three-nucleotide sequence in messenger RNA(mRNA) that signals the beginning of protein synthesis and into a protein. A “stop codon” can refer to a trinucleotide inor messenger JVIM ( HU L A; that signals a halt to protein synthesis in the cell.
[0104] The term “terminator” or “terminator region” can refer to a region of DNA that determines the detachment of RNA polymerase from the DNA template strand, which occurs towards the end of the transcription process. The skilled artisan will recognize that any terminator or combinations thereof can be utilized herein, including those described in Example 4 and Table 1. Non-limiting examples of promoters can comprise Tl, T2, T3, T4, T5, T6, T7, T8, bGH225, hGH477, SV40, beta globin, or any combination thereof.
[0105] The term “chromatin insulator” can refer to regulatory DNA elements that create boundaries in chromatin, delineating the ranges over which other regulatory influences take effect. Chromatin insulators can be implicated in the organization of chromatin and the regulation of transcription. The skilled artisan will recognize that any terminator or combinations thereof can be utilized herein. Non-limiting examples of chromatin insulators can comprise A4, DI, or both.
[0106] The term “enhancer” can refer to a regulatory DNA sequence that, when bound by specific proteins called transcription factors, can enhance the transcription of an associated gene. The skilled artisan will recognize that any enhancer can be utilized herein, including those described in Example 4 and Table 1. For example, enhancers can comprise sequences composed of binding sites for either 10-1 synthetic zinc finger domain, 1-3 synthetic zinc finger domain, or both.
[0107] The term “Kozak sequence” can refer to set of nucleotides that serves as the initiation site where protein translation begins in eukaryotic mRNA produced from transcription. The sequence is important for the initiation of translation and for the regulation of protein production and can ensure accurate translation of the protein in terms of ribosome assembly and translation initiation. The skilled artisan will recognize that any Kozak sequence can be utilized herein, including those described in Example 4 and Table 1. Non-limiting examples of Kozak sequences can comprise KI, K2, K3, K4, K5, K6, or any combination thereof.
[0108] The term “activation domain” can refer to a region of transcription factors (TFs) that bind to coactivator complexes to activate transcription. Activation domains can refer to short regions that directly bind to coactivators. The skilled artisan will recognize that any activation domain can be utilized herein, including those described in Example 4 and Table 1.Non-limiting examples of activation domains can comprise any combination thereof.
[0109] The terms “intrinsically disordered region” or “IDR” can refer to a polypeptide segment that does not contain sufficient hydrophobic amino acids to mediate co-operative folding. IDRs contain a higher proportion of polar or charged amino acids, and thus, lack a unique three-dimensional structure entirely or in parts in their native state. The term “intrinsically disordered protein” can refer to any protein insofar as it has an intrinsically disordered region. When the intrinsically disordered protein is used for screening, the intrinsically disordered protein can be used in a naturally occurring state, or the disordered portion can be cut out and used. The disordered portion can take any form insofar as an interaction with a test substance can be detected, and the disordered sequence portion and a portion having a domain structure can be bound to each other. The skilled artisan will recognize that any intrinsically disordered region can be utilized herein, including those described in Example 4 and Table 1. Non-limiting examples of intrinsically disordered regions can comprise FusN, DDX4, HNR, SRSF, or any combination thereof.
[0110] The term “4-hydroxytamoxifen-responsive domain” can refer to a specific region within a protein that responds to the presence of 4-hydroxytamoxifen (4-OHT), a metabolite of the drug tamoxifen. This domain can be part of a larger protein, such as the estrogen receptor, which can be activated or repressed by 4-OHT. 4-OHT is known for its role in breast cancer treatment, where it acts as an anti-estrogen by binding to estrogen receptors and inhibiting their activity. 4-OHT -responsive domains can be used in various experimental setups to control gene expression.
[0111] The term "zinc finger" or "ZF" can refer to a protein having DNA binding domains that are stabilized by zinc. Zinc finger domains can bind both DNA and RNA. The individual DNA binding domains are typically referred to as "fingers." A zinc finger protein has at least one finger, typically two fingers, three fingers, or six fingers. Each finger binds from two to four base pairs of DNA, typically three or four base pairs of DNA (the "subsite"). In embodiments, the zinc finger domain can comprise a high-affinity zinc finger, a mediumaffinity zinc finger, a low-affinity zinc finger, or any combination thereof. A zinc finger protein binds to a nucleic acid sequence called a target site. Each finger typically comprises approximately 30 amino acids as a zinc-chelating, DNA-binding subdomain. Studies have demonstrated that a single zinc finger of this class consists of an alpha helix containing the two invariant histidine residues coordinated with zinc along with the two cysteine residues of asingle beta turn (see, e.g., Berg & Shi, Science 271 : 1081-108 recognize that any zinc finger protein or polypeptide can be uuuzeu Herein, iiiciuumg mose described in Example 4 and Table 1.
[0112] The term "DNA binding domain" or "DBD" can refer to an independently folded protein domain that contains at least one structural motif that recognizes double- or singlestranded DNA. A DBD can recognize a specific DNA sequence, a recognition sequence or DNA binding motif (DBM), or have a general affinity to DNA. Some DNA-binding domains may also include nucleic acids in their folded structure. Examples for DBDs include, but are not limited to, the zinc finger (ZF) domain, the helix-turn-helix motif, the basic leucine zipper (bZIP) domain, the winged helix (WH) domain, the winged helix-tum-helix (wHTH) domain, the High Mobility Group box (HMG)-box domains, White-Opaque Regulator 3 domains and oligonucleotide / oligosaccharide folding domains. The helix-tum-helix motif is commonly found in repressor proteins and is about 20 amino acids long. The zinc finger domain is generally between 23 and 28 amino acids long and is stabilized by coordinating zinc ions with regularly spaced zinc-coordinating residues (either histidines or cysteines).
[0113] The term "binding" can refer to a non-covalent interaction between macromolecules (e.g. between a DNA binding domain (DBD) and a DNA binding motif (DBM) nucleic acid target site or sequence). In some embodiments, binding will be sequence-specific, and can be specific for one or more specific nucleotides (or base pairs) (e.g., such as between the binding of the DBD and the DBM sequence). In some embodiments, not all components of a binding interaction need to be sequence-specific (e.g. non-covalent interactions with phosphate residues in a DNA backbone). Binding interactions between a DBM nucleic acid sequence and DBD may be characterized by binding affinity and / or dissociation constant (Kt). A suitable dissociation constant for a DBD of the disclosure binding to its target DBM sequence may be in the order of 1 pM or lower, 1 nM or lower, or 1 pM or lower.
[0114] The term "affinity" can refer to the strength of binding of two binding partners. Typically, as the binding affinity increases, the Kt or Kp will reduce in value. Affinity can refer to the strength of binding as described herein to DNA, RNA, and / or even proteins. In some embodiments, a binding protein is designed or selected to have sequence-specific dsDNA- binding activity. For example, the DBM site for a particular DBD is a sequence to which the DBD concerned is capable of nucleotide-specific binding. It will be appreciated, however, that depending on the amino acid sequence of a DBM, the DBD can bind to or recognize more than one target DBM sequence, although typically one sequence will be bound in preference to anyother recognized sequences, depending on the relative sf covalent interactions. Thus, in some embodiments, the DBD win uinu me preierreu sequence with high affinity and non-target sites with low affinity (or will not bind at all: "lack of affinity"). In some embodiments, the DBD can be designed or selected to have a desirable affinity for a given target (e.g., high affinity vs. low affinity sequence specific dsDNA- binding). In embodiments, high affinity binding to a DNA binding site can deter other endogenous or synthetic transcription factors with lower affinity from displacing the high affinity binding partner to the site. Thus, selecting for high affinity binding vs. low affinity binding between two binding partners can be used to fine-tune the desired modulation of gene expression and / or the reversibility of the gene expression by modulating the strength of interaction between the two binding partners. Generally, high affinity binding comprises a dissociation constant (Kt or Kp) of 1 nM or lower, 100 pM or lower; or 10 pM or lower.
[0115] The term “grazoprevir-responsive domain” can refer to specific regions within proteins that respond to the presence of grazoprevir, an antiviral drug primarily used to treat hepatitis C. These domains can be engineered to interact with grazoprevir, allowing researchers to control the activity of certain proteins or gene expression in response to the drug. In embodiments, grazoprevir-responsive domains can be used to create gene circuits where the presence of grazoprevir triggers specific cellular responses. This can include the activation or repression of target genes, enabling precise control over biological processes.
[0116] The term “mRNA stability elements” can refer to specific sequences or structural features within mRNA molecules that influence their stability and lifespan within the cell. These elements can play a crucial role in regulating gene expression by determining how long an mRNA molecule remains intact and available for translation into protein. Non-limiting examples of mRNA stability elements comprise AU-rich elements (AREs), poly(A) tail, ironresponsive elements (IREs), RNA-binding proteins (RBPs), and chemical modifications.
[0117] The term “spacer sequence” can refer to a sequence of bases of unknown function occurring in DNA and lying between sequences of transcribed DNA. In embodiments, the spacer sequence can be of various lengths, including, but not limited to, zero base pairs, 25 base pairs, 50 base pairs, 75 base pairs, 100 base pairs, 150 base pairs, 200 base pairs, 250 base pairs, 300 base pairs, 350 base pairs, 400 base pairs, 450 base pairs, 500 base pairs, 550 base pairs, 600 base pairs, 650 base pairs, 700 base pairs, 750 base pairs, 800 base pairs, 850 base pairs, 900 base pairs, 950 base pairs, 1,000 base pairs, or more than 1,000 base pairs.
[0118] The term “transposon” can refer to a DNA sequ within the genome. Transposons can move from one location io anomer, eimer wirnin me same chromosome or to a different chromosome. This movement can cause mutations, alter the cell's genetic identity, and impact genome size.
[0119] The term “regulatory sequence” can refer to a specific region of DNA that can control the expression of genes. A regulatory sequence can play a crucial role in determining when, where, and how much a gene is expressed. Non-limiting examples of regulatory sequences comprise promoters, enhancers, silencers, insulators, operators, and response elements.
[0120] The term “insertion sequence” can refer to a short DNA sequence that acts as a simple transposable element, capable of moving, or transposing, within a prokaryotic genome, leading to genome rearrangements and potentially causing mutations.
[0121] The term “integron” can refer to a genetic element that enables bacteria to capture and express new genes, particularly those related to antibiotic resistance. In embodiments, integrons capture and express genes from gene cassettes. Integrons can play a role in the spread of antibiotic resistance.
[0122] The term “gene cassette” can refer to a mobile genetic element that can comprise a single gene and a recombination site. Gene cassettes can be integrated into larger genetic structures, such as integrons, allowing bacteria to acquire and express new genes.
[0123] The term “varied orientation” can refer to the various ways that genes or genetic elements can be arranged and oriented within the genome. In embodiments, varied orientation can refer to the direction in which genes are transcribed and their relative positions to one another.
[0124] Each nucleic acid in the second pool of a plurality of nucleic acids comprises a barcode. The term “barcode” can refer to a short, unique sequence of nucleotides (DNA or RNA) attached to a sample or molecule, acting as a tag to identify and track the sample or molecule during experiments. In embodiments, a barcode can facilitate amplicon-based readout by short-read next-generation sequencing. Barcodes can be useful in high-throughput sequencing and single-cell genomics, where they can help to distinguish between different molecules and reduce errors in sequencing data.
[0125] In embodiments, the barcodes can be structured, random or semi-random. The term “structured barcode” can refer to a barcode that comprises a sequence of nucleotides that is intentionally designed to identify and / or track a sample or molecule. The term “randombarcode” can refer to a barcode that comprises a completelyThe term “semi-random barcode” can refer to a barcode tnai comprises a paniany raiiuoin sequence of nucleotides with some fixed positions.
[0126] Accordingly, embodiments of the invention comprise assembling the first pool of a plurality of nucleic acids, assembling the second pool of a plurality of nucleic acids or both. The term “assembling” can refer the creating of a mixture of DNA or RNA sequences for various purposes, including but not limited to, gene synthesis, library construction, and screening. In embodiments, the synthetic gene circuits described herein can be constructed by assembling DNA-encoded genetic parts into multi-gene programs that can perform computational tasks in living cells.
[0127] Assembling the first pool of a plurality of nucleic acids, for example, can provide a diverse pool of nucleic acids each comprising an expression unit. Herein, the term “diverse” can refer to the variety and variability within genetic material. In embodiments, the libraries of constructs can comprise diverse genetic part combinations within cells, such as multiple nucleic acids of each type (i.e., four different promoters, one of which will assemble with one of four activation domains).
[0128] In embodiments, the assembling is hierarchical. The term “hierarchical” can refer to a system or structure where elements are ranked or organized in levels, with each level subordinate to the one above it. In embodiments, hierarchical assembly can be used to compose libraries described herein.
[0129] By combining the first pool of a plurality of nucleic acids (i.e., with the EU) with the second pool of a plurality of nucleic acids (e.g., with the barcodes), an indexed circuit library comprising a plurality of barcoded-linked nucleic acids is prepared. The term “library” or “nucleic acid library” can refer to a collection of nucleic acid molecules (circular, linear, or both).
[0130] In embodiments, the library can comprise a plurality of nucleic acids as described herein, or a plurality of vectors, as described herein. In embodiments, a library can be used to identify performance-optimized synthetic genetic circuit architectures. In embodiments, a library can be expressed in cells and phenotypically separated using new or existing high- throughput phenotypic characterization techniques, such as flow-seq or bulk RNA-sequencing, to map each barcode to a behavior.
[0131] Such libraries may or may not be contained in one or more vectors. The term "vector" is used interchangeably with "plasmid" and can refer to a nucleic acid moleculecapable of transporting another nucleic acid to which it has 1 acid molecule capable of directing the expression of genes anu / or nucieic aciu sequence io which they are operatively linked. The skilled artisan will recognize that any suitable vector can be used, non-limiting examples of which are provided in Example 4 and Table 3.
[0132] The skilled artisan will recognize that such plasmid vector(s) can comprise one or more regulator elements, non-limiting examples of which can comprise an antibiotic resistance gene, an origin of replication, and restriction enzyme recognition sequences that enable assembly of sub-components into the vector.
[0133] The term “antibiotic resistance gene” can refer to a gene that allows bacteria to resist the effects of antibiotics. Non-limiting examples of an antibiotic resistance gene can comprise a puromycin resistance gene, a kanamycin resistance gene, an ampicillin resistance gene, a spectinomycin resistance gene, a zeocin resistance gene, a chloramphenicol resistance gene, or any combination thereof.
[0134] The term “origin of replication” can refer to a DNA sequence which allows initiation of replication within a plasmid by recruiting replication machinery proteins.
[0135] The term “restriction enzyme” can refer to a protein isolated from bacteria that cleaves DNA sequences at sequence-specific sites, producing DNA fragments with a known sequence at each end.
[0136] In embodiments, the barcoded-linked nucleic acids (i.e., the index circuit library) can be analyzed by long-read sequencing to provide a construct-to-barcode index. The term “long-read sequencing” can refer to a DNA sequencing technique that enables the sequencing of much longer DNA fragments than traditional short-read sequencing methods. Long-read sequencing can allow for the detection of complex structural variants that may be difficult to detect with short reads, such as large inversions, deletions, and translocations. For example, long-read sequencing techniques can comprise nanopore sequencing. The term “nanopore sequencing” can refer to a long-read sequencing technology that allows for the direct, real-time analysis of DNA or RNA molecules as they pass through a nanopore. The term “nanopore” can refer to a tiny hole embedded in a membrane.
[0137] Embodiments can comprise analyzing the long-read sequencing data or output, thereby allowing identification of sequence variants from a single read. Long-read sequencing data analysis can comprise basecalling, alignment, variant calling, and functional analysis. The term “basecalling” can refer to the conversion of the raw data from a sequencing technique into nucleotide sequences. The term “alignment” can refer the process of arranging andcomparing DNA, RNA, or protein sequences to identify regio mapping and visualization of the sequencing data. The term variant caning can reier io me identification of various genomic variants, including but not limited to, single nucleotide polymorphisms (SNPs), copy number variations (CNVs), indels (insertions / deletions), and larger structural variants. The term “functional analysis” can refer to the interpretation of sequencing data to understand the biological functions and roles of genes, proteins, or other genomic elements.
[0138] For example, the long-read sequencing data can be analyzed using a software pipeline that identifies circuit components and barcode sequences from individual sequencing data reads.
[0139] In embodiments, the barcoded-linked nucleic acids (e.g., the index circuit library) can be integrated into a population of cells and analyzed by short-read amplicon sequencing to provide a barcode-to-phenotype index. Integrating an index circuit library into a population of cells can involve introducing a collection of genetic constructs, each with unique identifiers (barcodes), into the cells. This process can allow for tracking and analyzing the behavior of individual cells or groups of cells within a larger population.
[0140] In embodiments, the population of cells can be bacteria cells, yeast cells, or mammalian cells. Non-limiting examples can include HEK293T cells, T cells, mesenchymal stem cells (MSCs), induced pluripotent stem cells (iPSCs), or K562 cells, NK cells, macrophages, human and mouse embryonic stem cells, yeast, E. coli, P. putida, cortical neurons, or primary mouse hepatocytes. In embodiments, the one or more cells can comprise a population of cells.
[0141] In embodiments, the phenotype can comprise a gene expression output. In embodiments, the gene expression output can comprise a secreted therapeutic protein, a surface protein, a regulatory circuit, a reporter protein, transcription factors, pro-apoptotic proteins, or any combination thereof.
[0142] The term “short-read amplicon sequencing” can refer to a sequencing technique used to analyze specific regions of DNA by first amplifying the regions using PCR (Polymerase Chain Reaction) and then sequencing the amplified products using short-read sequencing technologies. The term “short-read sequencing” can refer to a DNA sequencing technique that generates short DNA or RNA sequences, typically ranging from 50 to 300 base pairs in length. The term “amplicon” can refer to a specific, pre-defined segment of DNA that has been amplified by PCR.
[0143] In embodiments, the population of cells compri; acids (e.g., the index circuit library) can be sorted according io pnenoiype anu anaiyzeu using short-read amplicon sequencing. Sorting the population of cells according to phenotype can involve separating and categorizing cells based on their observable characteristics or traits. For example, the phenotype can comprise a gene expression output, protein expression output, cell size, cell shape, and cellular metabolic activity. For example, the phenotype can comprise a gene expression output.
[0144] The skilled artisan will recognize that providing the construct-to-barcode index or the barcode-to-phenotype index can be performed in any order (either sequentially, simultaneously, concurrently, etc.).
[0145] Embodiments of the invention can further comprise comparing the construct-to- barcode index with the barcode-to-phenotype index to associate an expression unit with a phenotype, thereby screening the gene circuit library. In embodiments, this process can involve linking the genetic constructs (genetic circuits) to their corresponding barcodes and then associating the barcodes with observed phenotypes. Screening a gene circuit library can involve evaluating a collection of gene circuits to identify those with desired properties or functions. The term “screening” can refer the process of testing a population or a set of samples to identify specific genetic traits, variations, or mutations.
[0146] Aspects of the invention are drawn to a genetic screening platform. In embodiments, the genetic screening platform combines long- and short-read next-generation sequencing modalities, and assembles pools of genetic parts into libraries of constructs with combinatorially varied genetic compositions. The skilled artisan will recognize that the platform described herein can be adapted to various applications, (i) cell-based therapies, for example identifying more productive CAR-T therapies or generating new therapeutic programs; (ii) development of stem cell-based cell or tissue products, such as blood or cardiac tissue; (iii) construction of more compact and effective genome editing technologies; (iv) discovery of new synthetic regulatory elements, such as large enhancers or promoters; (v) rapid optimization of multi-gene dynamic circuits with desirable properties; and (vi) Al driven gene circuit design with models trained using the data obtained from the method described herein.
[0147] Aspects of the invention are further drawn to a high-throughput method of characterizing a gene circuit library. See, for example, the CLASSIC (combining long- and short-range sequencing to investigate genetic complexity) platform as exemplified herein. As shown in FIG. 1 and FIG. 14, the method can comprise a platform that combines long- andshort-read sequencing to interrogate genetic complexity. genetic parts into libraries of constructs with combinatorially vaneu genetic compositions, in embodiments, random association of assembled library members with short barcodes can enable indexing of part compositions to barcodes with via long-read nanopore sequencing and data processing using custom software. In embodiments, libraries can be introduced into cells, phenotypically sorted, and then barcodes quantitatively assessed via short-read Illumina sequencing to construct a phenotype-to-barcode index. In embodiments, indexes can then be compared to assign compositions to phenotype. In embodiments, the mapping of part design space that is enabled by CLASSIC can be used to identify compositions that support target behavior, understand rules for part context dependence, or training of ML models for circuit behavior.
[0148] In embodiments, a first pool of genetic parts is assembled into a library of constructs with combinatorially varied genetic compositions. In embodiments, the first pool can comprise a plurality (i.e., one or more) of nucleic acids comprising an expression unit as described herein. In embodiments, the first pool can be combined with a second pool of nucleic acids (e.g., the barcode pool). In embodiments, each nucleic acid of the second pool can comprise a barcode, thereby providing a nucleic acid comprising one expression unit operably linked to at least one barcode. In embodiments, long-read sequencing of the barcoded-linked nucleic acid is can be performed, thereby providing a construct-to-barcode index. In embodiments, the long-read sequencing of the barcoded-linked nucleic acid can be analyzed, thereby allowing identification of sequence variants from a single read. In embodiments, the barcoded-liked nucleic acid can be integrated into a plurality of cells. In embodiments, the plurality of cells can be sorted according to phenotype, and short-read amplicon sequencing can be performed, thereby providing a barcode-to-phenotype index. In embodiments, the construct-to-barcode index can be compared with the barcode-to-phenotype index, thereby linking the expression unit with a phenotype.
[0149] Aspects of the invention are also directed towards kits, such as kits comprising the materials and reagents as described herein.
[0150] In one embodiment, the kit includes (a) a container that contains the reagents(s), such as that described herein, and optionally (b) informational material. The informational material can be descriptive, instructional, marketing or other material that relates to the methods described herein and / or the use of the agents for research and / or therapeutic benefit.
[0151] The informational material of the kits is not limited the informational material can include information about prouucuon or me reagents, moiecuiar weight of the reagents, concentration, date of expiration, batch or production site information, and so forth. The information can be provided in a variety of formats, include printed text, computer readable material, video recording, or audio recording, or information that provides a link or address to substantive material.
[0152] The kit can include other ingredients, such as a solvent or buffer, a stabilizer, or a preservative. When the reagents are provided in a liquid solution, the liquid solution can be an aqueous solution. When the reagents are provided as a dried form, reconstitution is by the addition of a suitable solvent. The solvent, e.g., sterile water or buffer, can optionally be provided in the kit.
[0153] The kit can include one or more containers for the reagents. In some embodiments, the kit contains separate containers, dividers or compartments for the composition and informational material. For example, the composition can be contained in a bottle, vial, or syringe, and the informational material can be contained in a plastic sleeve or packet. In other embodiments, the separate elements of the kit are contained within a single, undivided container. For example, the composition is contained in a bottle, vial or syringe that has attached thereto the informational material in the form of a label. In some embodiments, the kit includes a plurality (e.g., a pack) of individual containers, each containing one or more unit dosage forms (e.g., a dosage form described herein) of the agents. The containers can include a combination unit dosage, e.g., in a target ratio. For example, the kit includes a plurality of syringes, ampules, foil packets, blister packs, or medical devices, e.g., each containing a single combination unit dose. The containers of the kits can be air-tight, waterproof (e.g., impermeable to changes in moisture or evaporation), and / or light-tight. The kit optionally includes a device suitable for administration of the composition, e.g., a syringe or other suitable delivery device. The device can be provided pre-loaded with one or both of the agents or can be empty, but suitable for loading.Other Embodiments
[0154] While the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
[0155] The invention will be further described in the 1 limit the scope of the invention described in the claims.EXAMPLES
[0156] Examples are provided below to facilitate a more complete understanding of the invention. The following examples illustrate the exemplary modes of making and practicing the invention. However, the scope of the invention is not limited to specific embodiments disclosed in these Examples, which are for purposes of illustration only, since alternative methods can be utilized to obtain similar results.EXAMPLE 1
[0157] Sesame Street: a hierarchical, DNA assembly toolkit for genetic control in mammalian cells
[0158] We describe a composition of matter platform for high-throughput construction and single cell measurement of large and complex synthetic genetic compositions. Using a custom catalogue of plasmid vectors and type Ils restriction enzyme overhangs, we demonstrate the ability to use hierarchical golden gate assembly to efficiently construct large (~105) pools of barcoded multi-gene DNA constructs from donor plasmids harboring genetic sub-components. Through the combinatorial assembly of pools of genetic sub-components with a pool of short, semi-degenerate barcode sequences, we show the capability to quickly and flexibly generate highly diverse barcoded genetic composition libraries. These libraries are sequenced using long-read sequencing (Oxford Nanopore) to generate a barcode-to- composition map. In parallel, the libraries are introduced into cells and can be assayed using new or existing high-throughput phenotypic characterization techniques, such as flow-seq or bulk RNA-sequencing, to map each barcode to a behavior. The two datasets can then be united to yield a genotype-to-phenotype index.
[0159] Cloning Pipeline (Sesame Street): The invention described here is a toolkit of DNA sequences and type IIS restriction enzyme overhang sequences that enable efficient, hierarchical golden gate-style assembly of multi-expression unit plasmids from smaller donor plasmids. Each level of the hierarchy has a corresponding plasmid backbone sequence, comprised of a selectable marker and an origin of replication, into which genetic components of interest can be assembled. The lowest level (level 0) are sub-part plasmids and house protein sub-domains, enhancer regions, minimal promoters, insulator sequences, Kozak sequences, ormRNA stability elements. Level 0 plasmids are assembled i into level 1 ampicillin-resistant destination vectors, which contain an nuxuuyz-ccuo IUSIOII protein that is toxic in certain E. coli cloning strains and enables red-white color screening. Level 1 plasmids consist of full-length promoters, functional protein domains, and transcriptional terminators. Level 1 plasmids can then be assembled using Esp3I into level 2 kanamycin-resistant destination vectors, which contain an mRuby2-ccdB fusion protein, to form functional expression units. A set of / ?z / -cleavable overhang sequences in the level 2 destination vector determine the order in which level 2 expression unit plasmids can be assembled into a spectinomycin-resistant level 3 destination. The level 3 destination vector contains a large orange chromoprotein that is replaced during assembly and can be used for orange-white color screening. The assembly efficiency that our cloning hierarchy enables stoichiometrically-balanced construction of large level 1, level 2, and level 3 plasmid libraries from pools of donor plasmids.
[0160] Pooled Assembly Method (CLASSIC): With the ability to construct large (> 106 members) libraries of diverse sequence compositions from a discrete set of input plasmids, we identified that this strategy could be used to identify performance-optimized synthetic genetic circuit architectures. We generated a novel two-part, customized, semi -degenerate barcode sequence, half of which is contained in a level 2 plasmid pool and half of which is contained in a level 3 destination vector pool. During assembly of the final level 3 multi -unit expression arrays, the two barcode pools combine to yield a single barcode pool equal in size to the product of the two individual pools. The barcoded plasmid pool can then be sequenced using Nanopore long-read sequencing technology (or an alternative long-read sequencing platform) and analyze the resulting data using software we developed, to generate a barcode- to-composition index. The software pipeline leverages the semi -degenerate structure of the barcode pool to accurately identify barcode sequences in the Nanopore data. Importantly, the ability to assemble a large pool of barcodes with a smaller pool of compositions enables many- to-one barcode-to-composition mappings, allowing for collection of multiple measurements on a given composition. Simultaneously, the plasmid library can be expressed in cells (bacteria, yeast, mammalian cells, etc.) and phenotypically separated using any available method. To date, we have used flow sorting and growth-based assays to separate cells based on their behavior. Short-range sequencing of the barcode region from cells in each phenotypic population yields a phenotype-to-barcode index. The phenotype-to-barcode index can then beunited with the composition-to-barcode index to produce a large (< 20 kb), complex synthetic sequence compositions.
[0161] HEK293T Landing Pad cell line: To enable single-copy genomic integration of library members, we generated a custom “landing pad” cell line in HEK293T cells. We used Cas9-mediated cutting at the AAVS1 locus on chromosome 19 and homology directed repair to insert into the AAVS1 locus a custom landing pad transgene containing eYFP and a hygromycin resistance gene driven by an hEFla promoter. Between the promoter and the coding sequence is a BxBl attP recombinase recognition site, which, in the presence of the recombinase BxBl, permits irreversible genomic insertion of a destination plasmid harboring a cognate BxBl attB recognition site at the attP site.
[0162] Software: To construct a scalable and affordable platform for mapping long genetic compositions to phenotypic changes, we elected to use Nanopore long-read sequencing for building the composition-to-barcode index. Nanopore sequencing is significantly more affordable than other long-read sequencing modalities but suffers from higher base-calling error rates. To navigate higher error rates, we developed a custom software package named WIMPY, which enables high accuracy identification of circuit components and barcode sequences from individual reads.
[0163] Given the generalizable nature of this platform, the CLASSIC workflow can be applied broadly to any problem in which the successful identification of productive genetic compositions is limited by throughput and combinatorial complexity. Most immediately, we envision this platform being adapted to applications for (i) cell-based therapies, for example identifying more productive CAR-T therapies or generating new therapeutic programs, (ii) development of stem cell-based cell or tissue products, such as blood or cardiac tissue, (iii) construction of more compact and effective genome editing technologies, (iv) discovery of new synthetic regulatory elements, such as large enhancers or promoters, (v) rapid optimization of multi-gene dynamic circuits with desirable properties, and (vi) Al driven gene circuit design with models trained using the data obtained from CLASSIC experiments (as we have demonstrated).
[0164] There are several molecular cloning toolkits that current exist (Yeast Toolkit, Mammalian Toolkit, Bacterial Toolkit), however we take a unique approach in that we developed a highly stratified assembly hierarchy, which increases the time taken to assemble multi-gene constructs but significantly increases the efficiency of each step. This increase in efficiency enables pooled library assembly, which is a critically important downstreamapplication for constructing genetic libraries. While there assembly of DNA fragments or plasmids to generate a barcoue-muexeu piasimu pooi i^sucn as nested assembly approaches or PASIV), our platform enables many-to-one barcoding of genetic variants. Additionally, the novel use of two separate barcode pools that assemble in the final step to form a single barcode provides a scalable solution for exploring large genetic design spaces. Finally, our software pipeline leverages the semi -degenerate barcode pools to assay large pools of barcoded libraries and identify compositions and barcodes at a much higher efficiency than any other software pipeline available that is compatible with similar datasets.
[0165] Without wishing to be bound by theory, the platform described herein can be used with different genomic integration technologies, including but not limited to PiggyBac or Sleeping Beauty transposition systems, retroviruses, and CRISPR-based platforms, as well as non-integrating strategies such as adenovirus or baculovirus. This will enable the use of CLASSIC with less genetically tractable or manipulable cell types where generating landing pad lines is not feasible, such as primary human cells. Additionally, we intend to use and develop new assays for phenotypic assessment, such as single cell RNA sequencing and imagebased genotyping.EXAMPLE 2
[0166] Ultra-high throughput mapping of genetic design space
[0167] Abstract
[0168] Massively parallel genetic screens have been used to map sequence-to-function relationships for a variety of genetic elements. However, because these approaches only interrogate short sequences, it remains challenging to perform high throughput (HT) assays on constructs containing combinations of multiple sequence elements arranged across multi-kb length scales. Overcoming this barrier could accelerate synthetic biology; by screening diverse gene circuit designs, “composition-to-function” mappings could be created that reveal genetic part composability rules and enable rapid identification of behavior-optimized variants. Here, we introduce CLASSIC, a genetic screening platform that combines long- and short-read nextgeneration sequencing (NGS) modalities to quantitatively assess pools of constructs of arbitrary length containing diverse part compositions. We show that CLASSIC can measure expression profiles of >105gene circuit designs (from 5-20 kb) in a single experiment in human cells. The resulting datasets can be used to train ML models that accurately predict circuit behavior across expansive circuit design landscapes, revealing part composability rules thatgovern circuit performance. Our work shows that by expandi build-test-leam (DBTL) cycle, CLASSIC enhances the pace anu scaie or syiimeuc uioiogy anu establishes an experimental basis for data-driven design of complex genetic systems.
[0169] Introduction
[0170] Synthetic gene circuits are constructed by assembling DNA-encoded genetic parts into multi-gene programs that perform computational tasks in living cells1'3. Over the past two decades, gene circuits have emerged as important models for understanding native gene regulation1and have been used to create powerful biotechnologies by enabling user-defined control over cellular behavior5'7. Despite this progress, the design of quantitatively precise circuit behavior remains challenging. Regulatory interactions within a circuit must be carefully tuned, often through multiple iterative DBTL cycles, before part compositions that support a desired circuit behavior are identified8. Additionally, since genetic parts must work in close physical proximity to one another, as well as within a crowded intracellular environment9, incidental molecular coupling can occur between parts and with host cell regulatory machinery10'12. These context-dependent interactions can confound the ability to design circuits using biophysical and mechanistic modeling frameworks, further extending the number of DBTL cycles required to achieve a target behavior13.
[0171] One potential strategy for increasing the pace of gene circuit engineering is to expand the number of circuits tested in each cycle by performing HT functional screens on pooled circuit libraries. By measuring large collections of circuits in a single experiment, such an approach could be used for rapidly profiling complex circuit design spaces to identify part compositions with desired quantitative behaviors. In contrast to traditional circuit engineering approaches, collecting large amounts of circuit activity data could also facilitate the training of ML / Al models that are capable of inferring context-specific part function or more accurately predicting the behavior of untested circuit designs. HT screening approaches that utilize shortread NGS as a readout14'18have been used to generate detailed sequence-to-function mappings for multiple genetic part classes, including promoters19'23, terminators and 3’ UTRs24,25, transcription factors (TFs)26,27, nucleic acid switches28, and receptors29. However, high-depth functional profiling of libraries of DNA constructs long enough (>1 kb) to encode entire circuits can be costly and technically challenging. Methods have emerged that permit multiplexed analysis of long constructs by appending short barcode index sequences that facilitate amplicon-based readout by short-read NGS30. This approach has been used in array formats in which each construct is assembled and barcoded separately, or through nestedassembly schemes30,31, limiting library design flexibility and random barcoding and indexing via long-read single-molecuie reai-iiine ^ivuxi ; sequencing (PacBio) has been used to characterize libraries of variants ranging from 102to 104constructs, mostly for individual ORFs of up to 6 kb in length32'34.
[0172] Results
[0173] To perform indexed multiplexing for construct libraries containing combinations of genetic parts at larger scales, we devised an approach that combines long-read nanopore and short-read Illumina NGS to analyze construct libraries generated via pooled genetic part assembly (Fig. 1A). Our assembly scheme incorporates semi -random barcodes that can be used to create a construct composition-to-barcode index by nanopore sequencing. Libraries can then be introduced into cells, binned based on functional phenotype, and analyzed by short-read (Illumina) amplicon sequencing to produce a barcode-to-phenotype index. A map that matches construct composition to phenotype can then be revealed by comparing the two indices. Using this technique, which we refer to as CLASSIC (combining long- and short-range sequencing to investigate genetic complexity), it is possible to obtain high-depth phenotypic expression data for large libraries (>105) of genetic part compositions of arbitrary length using standard phenotypic selection or flow sorting experiments.
[0174] To configure CLASSIC to quantitatively profile libraries of constructs containing diverse genetic part combinations in human cells (FIG. IB, FIG. 14), we designed a custom hierarchical golden-gate35cloning scheme in which library diversity is programmed through a series of pooled DNA assembly steps (FIG. 15-17). Diversified pools of 5’, coding, and 3’ gene elements are first generated through assembly of input part fragments (level 0 to 1), and subsequently combined to create single-gene expression unit (EU) pools (level 1 to 2) (FIG. 15). The EU pools are then combined with a BFP EU plasmid pool that expresses a transcript containing semi-degenerate barcode sequences in the 3’ UTR (FIG. 18, yielding barcoded multi -EU circuit libraries (level 2 to 3). Following nanopore sequencing and data analysis (see Methods, FIG. 19), the libraries are genomically integrated at single copy into a HEK293T cell line harboring a custom “landing pad” cassette (HEK-LP)36,37(see Methods, FIG. 20). Library-integrated cells are then flow-sorted based on circuit expression output, followed by Illumina NGS to quantitate bin distributions of circuit- associated barcodes14,15(see Methods, FIG. 21). We validated this workflow with a small (n=384) expression unit (EU) library containing shuffled regulatory elements (promoter, Kozak sequence, and terminator) (FIG. 22), demonstrating that our assembly scheme yields complete, part-balanced library coverage (FIG. 23A-C) and proportional assortment of unique barcodes (FIG. 23D-E).HEK-LP library integration and flow-seq measurements wei measurements made with a set of randomly isolated, clonally expanueu variants, inese grounu truth” measurements showed excellent agreement with corresponding CLASSIC -derived values (MAE=0.07), and demonstrated that measurement precision is enhanced by mapping multiple unique barcodes to each composition. We also collected replicate libraries to demonstrate that CLASSIC is technically and biologically reproducible. Finally, we showed that the data can be used to quantify relative part contributions to expression level, as well as identify unexpected interference between parts.
[0175] To evaluate whether CLASSIC can be used to experimentally profile a complex design space for a multi-gene circuit, we constructed a library of small-molecule inducible, single-input transcriptional switches comprising >105potential unique compositions. Circuits in this library consist of two EUs (FIG. 30A): one encoding a constitutively expressed synthetic Cys2-His2 zinc-finger (ZF)-based transcription factor (synTF)38'40with appended transcriptional activation domains (TAs), and the other encoding a reporter gene harboring cognate synTF binding motifs (BMs) located upstream of a minimal promoter driving eGFP expression38. Other non-limiting examples of reporter elements comprise fluorescent reporter elements, such as blue fluorescent protein (BFP), green fluorescent protein (eGFP), yellow fluorescent protein (YFP), mRuby, mCherry, or fluorescent proteins in the infrared range, such as iRFP670 and iRFP713. Transcriptional induction occurs upon addition of the small molecule 4-hydroxytamoxifen (4-OHT), which binds to a mutant version of the human estrogen receptor (ERT2)41appended to the synTF, facilitating its translocation from the cytoplasm into the nucleus to activate reporter transcription. As we and others have demonstrated38,42'44, identifying designs that are optimized for high fold-change (HFC) (FIG. 30A, right) expression can be challenging in mammalian systems; to ensure robust expression of a transgene exclusively in the presence of inducer, basal and induced expression levels must be respectively minimized and maximized through tuning of both the transcriptional regulatory features of the locus and the molecular properties of the synTF.
[0176] To create a library that diversifies the design of this circuit architecture, we identified 10 genetic part categories that we hypothesized could play a role in circuit tuning to achieve HFC behavior. This included 4 different transcriptional activation domains (TAs)45'47, 3 ZF affinities38, and a set of 4 intrinsically disordered protein (IDP) domains48'50that have been shown to facilitate liquid-liquid phase condensation of nuclear-localized TFs48,49(FIG. 30, left). We also varied transcriptional regulatory features of the circuit, including 4 promotersand 4 terminators in the synTF coding EU, as well as 3 , numbers of BMs (2, 4, 8, or 12) in the reporter EU (FIG. 3br», ten;. / .uuiiioiiaiiy, we vaneu the spacing between the EUs (0, 250, or 500 bp) along with their 5’-to-3’ orientation (FIG. 30B, left) to yield an overall circuit design space of 165,888 compositions (FIG. 30B, right). This library was constructed in 3 steps by first assembling protein domain parts to create a level 1 synTF ORF pool, then conducting parallel assemblies to generate level 2 pools of synTF coding and reporter EUs, and finally combining EU and barcode pools into the level 3 destination vector (FIG. 30B right, FIG. 5).
[0177] Nanopore sequencing of the pooled library yielded barcode assignments for 95.3% of total compositions (FIG. 30C), with a mean of 8.4 barcodes for each circuit (FIG. 30C bottom left). We integrated the library into HEK-LP cells and sorted un-induced and 4-OHT- induced populations (8.6 and 15.7»106cells, respectively) separately into 8 bins based on eGFP fluorescence (FIG. 30D left). Following analysis by Illumina NGS, we matched a total of 121,292 (73% of design space) compositions to barcodes that were detected in the sorted populations (mean barcodes circuit'1=2.25) (FIG. 30C). We then used barcode bin distributions to compute basal, induced, and fold-change expression values for each variant. CL AS SIC- derived fold-change values demonstrated excellent overall agreement ( LEA0.15) (FIG. 30D top middle) with those of random isolates (n=40) (FIG. 30D bottom), and values for both basal and induced expression showed comparably high percent similarities (FIG. 30D top right). These results confirm that CLASSIC can generate quantitatively accurate measurements of libraries large enough to profile complex genetic design spaces.
[0178] Since an incomplete mapping of circuit design space (27% was unmeasured) could limit our ability to identify HFC-optimized compositions, we tested whether a complete mapping could be obtained by training an ML model with our data and then using it to predict the behavior of unmeasured compositions. We trained several model classes and found that multi-layer perceptron (MLP) (90: 10 train / validation split, with high-quality test set r2=0.86 for basal 0.90 for induced data) and random forest (0.84 and 0.86) models significantly outperformed biophysically-consistent (0.46 and 0.27) and unconstrained linear regression models (0.53 and 0.64), indicating that the complex, non-linear interactions between genetic parts can be more accurately captured by advanced model architectures. We selected the MLP to move forward with (FIG. 31A) due to its predictive accuracy and potential for extensibility (FIG. 31A) (see Methods). We used the MLP to impute basal and induced eGFP expression values for the 44,596 unmeasured part compositions, and adjusted values for the 121,292 CLASSIC-measured compositions to yield a complete design space analyzed the distribution of values across a 2-dimensional projection 01 uenavior space ^uasai vs. induced expression), examining the relative densities of compositions in 3 regions of interest: low basal (<500 AU), high induced (>70,000 AU), and HFC expression (>25x foldchange) (FIG. 31B, right). We observed a greater proportion of library members in the low basal than in the high induced region (~4.3x), with -50% of compositions showing fold-change values of <3x. Only a small fraction of the library (-8%) fell within the HFC region. While a comparison between MLP-modeled and CL AS SIC-measured distributions revealed similar global features, we observed high absolute error (>2) between the model predictions and CL AS SIC-measured values at the periphery of behavior space (-3% of variants, including many in the HFC region), suggesting either CLASSIC measurement errors or poor model prediction in these regions.
[0179] To experimentally assess if our model can accurately predict behavior across circuit design space, we constructed cell lines harboring individual CL AS SIC-measured (n=15) and unmeasured (n=81) compositions, including many of which were predicted to be among the highest fold-change compositions in the design space (FIG. 31C). Flow cytometry analysis of the cell lines closely corresponded with model predictions for both measured (MAE=Q.18) and unmeasured (MAE=0A9) compositions, confirming the behavior of HFC compositions predicted to be in the 50-80x range. Cell lines constructed for circuit configurations that correspond to apparent high-error CLASSIC measurements also showed strong overall agreement with the model, confirming that it performs accurate error correction for inaccurately measured configurations. These results validate the ability of our CLASSIC- trained ML model to enable gene circuit design space mapping through accurate imputation of inaccurately-sampled or unsampled compositions.
[0180] With a quantitatively accurate mapping of single-input design space in-hand, we sought to uncover part composition rules governing circuit behavior. First, we examined relative part usage frequencies in design space regions of interest (FIG. 31D). We observed distinct part usage in categories associated with the synTF protein (TA, IDP, and ZF affinity), as well as the promoter driving synTF expression, the number of BMs, the core promoter, and EU orientation. For example, VPR, a strong TA, is used extensively in the high induced region, but is nearly absent amongst low basal circuits in favor of a weaker TA, p65, while intermediate-strength VP16 and VP64 are enriched amongst HFC circuits. By the same token, lower- and medium-activity core promoters (miniTK and ybTATA) are excluded from highactivity circuits, while the low basal region has few compos core promoter), and HFC circuits utilize a mix of ybTATA anu mviviv . nrc circuits tavoieu compositions with the reporter EU in the 5’ position, an orientation that appears to minimize basal circuit expression across design space, potentially by eliminating transcriptional crosstalk from the upstream promoter.
[0181] To gain further insight into the design rules underlying HFC behavior, we used our trained MLP to compute Shapley additive explanation (SHAP) values, which provide an objective metric of each part category’s relative importance in determining target behavior (FIG. 31E)54. SHAP analysis revealed that HFC behavior is primarily influenced by part selection within six categories: TA, IDP, ZF affinity, and synTF expression and core promoters. To investigate part usage co-variance, we computed cross-category mutual information (MI) for compositions within the HFC region (FIG. 31F). This analysis revealed strong coupling among the six SHAP-important categories, indicating that their co-optimization is critical for achieving HFC behavior. To further explore this interdependence, we applied UMAP-assisted K-means clustering55on HFC compositions, identifying three clusters (FIG. 31G). Interestingly, all three had distinct synTF designs, but shared locus features: high BM valency (n=8 or 12), a strong core promoter (mCMV), and extended spacing (500 bp) between the EUs (FIG. 31G). The largest cluster (A) featured medium-activity promoters driving synTF expression (hEFlal and hPGK) and medium-strength TAs (VP64 and VP 16) paired with higher medium-affinity ZFs. Performing a dose-response analysis on a representative composition from this cluster revealed a Hill coefficient (nn) of 2.13, indicating moderate cooperativity. In contrast to cluster A, compositions in clusters B and C favored strong expression (CMV promoter), with B containing medium-affinity ZFs and a weak TA (p65), and C showing enrichment for weak-affinity ZFs and medium-strength TA (VP 16). Taken together, these observations suggest that HFC behavior is achieved via a common locus design that minimizes basal reporter expression, and synTF designs that carefully balance expression level, ZF affinity, and transcriptional activation. Taken together, these results demonstrate that analysis of CL AS SIC-enabled design space mappings can provide novel insight into how part composition supports target circuit behavior.
[0182] While our results establish that accurate design space mappings can be obtained for regimes with balanced training-to-prediction ratios, we realized that models trained with CLASSIC data would offer greater utility if they could generalize to a broader range of out-of- sample test cases. To evaluate the model's generalization capacity, we progressivelydownsampled the training data while performing validation31H, left). Surprisingly, we found that only -7% of the data was suinuieni io aumeve an of >0.8. We then examined whether the design space mapping could be expanded by incorporating additional genetic part measurement data (FIG. 31H, right). To do this, we assembled two libraries, each containing approximately 3.5k members: one incorporating a previously unmeasured TA (NFZ)56and another with a 10-GS linker replacing the IDP. Following measurement of these libraries, we fine-tuned the existing "base" model with the new data to predict the expanded design space, which now included the new parts (approximately 259k compositions). While the base model alone showed good predictions of new part behavior (r2of -0.82 for basal and 0.87, for induced conditions) the addition of the new data improved model generalization, maximizing its accuracy (r2of -0.88 for basal and 0.92, for induced conditions) with just 1,000 instances. These results show that our modeling framework is highly data-efficient and can easily and rapidly incorporate new parts and design features through fine-tuning, suggesting its potential to generalize to design spaces that are much larger than the datasets used for training.
[0183] We attempted to evaluate these capabilities for design space expansion by mapping a design space for multi-input circuits with two inducible synTFs, one regulated by 4-OHT and the other by Grazoprevir (GZV), a small molecule protease inhibitor that blocks autolytic cleavage of aNS3 protease-synTF fusion57(FIG. 32A). In this 3-EU architecture, co-regulation of the reporter locus by the 2 synTFs could potentially support Boolean logic, including AND- gate function if the synTFs cooperatively activate the reporter, and OR-gate function if they independently maximize activation. We hypothesized that this architecture could be tuned to achieve these behaviors by diversifying a set of 68 design features, including many parts from the single-input library, as well as BMs for the two synTFs arranged in four distinct patterns, resulting in an overall design space of approximately 3.4 billion compositions (FIG. 32B). To create a mapping for this space, we employed a two-step data collection strategy, first measuring a densely sampled ’’base library” space comprising a of 44-feature subset (1.1M compositions), and then obtaining data for an “expansion library” that sparsely sampled the remaining features (FIG. 32B). During base library construction, we deliberately constrained diversity to ensure high read depth (FIG. 35A and B), obtaining barcode indexes for -7% (— 112k compositions) of the 1.1M potential compositions (FIG. 32C), with a mean of 3.2 barcodes for each circuit (FIG. 32C bottom left). The expansion library was constructed by assembling 24 independent 1,500-member pools, each containing one of the expansion featuresheld constant against a random sampling of all other base libraries were integrated separately into HEK-LP cells and treaieu wun ‘t-un r JZ. V , or oom inducers. Each treatment was sorted into 8 bins based on eGFP fluorescence and isolates were obtained from the base (n=14) and expansion libraries (n=13). As before, we performed Illumina NGS to obtain expression values from both libraries (~83k and ~45k respectively for base and expansion libraries) (FIG. 31C, bottom), once again observing overall agreement between CLASSIC and isolate measurements (MAE=0.28, 0.23) (FIG. 32C).
[0184] To create a multi-input circuit design space mapping, we first trained a 5-layer MLP (FIG. 33 A) with the base library data, yielding a model with strong predictive accuracy (90: 10 train / validation split with test set r2values of 0.80 / 0.90 / 0.86 / 0.88 for basal / 4-OHT / GZV / both inducer predictions). We then extended the model to the full 3.4B-member design space by fine-tuning with expansion library data, resulting in good predictive performance (r2values of 0.79 / 0.84 / 0.84 / 0.82) (FIG. 33B, left). We found that, while only a small region of base library design space demonstrated Boolean-like behavior (DKL values of <0.12 for AND, <0.05 for OR), these regions expanded in the complete design space (FIG. 33B, right). To validate the model, we constructed 36 individual compositions spanning the behavior space, 35 of which were unmeasured, including many predicted to exhibit AND- and OR-like function (FIG. 33C). As before, we observed good agreement between measurements of individual compositions and the model (MAE = 0.32) (FIG. 33C), confirming the effectiveness of our data sampling and training approaches.
[0185] Similar to the single-input design space, we observed strong polarization in part usage within both AND- and OR-like regions (FIG. 33D). While compositions from both regions were enriched for a high number of BMs (n=8 or 12) and strong core promoters (mCMV or SCP2), AND-like compositions predominantly utilized clustered BM patterns, while OR gates used exclusively interspersed patterns. In both regions, the 4-OHT-inducible synTF (synTF A) was expressed at intermediate levels (hEFlal) while the GZV-inducible synTF (synTF B) showed stronger expression and greater overall part diversity. SHAP analysis revealed that core promoter and inter-EU spacing were common drivers of both behaviors (FIG. 33E), and OR-like behavior was impacted by more features than AND-like behavior, indicating a greater degree of feature interactivity. This was consistent with part usage coupling analysis, which showed a lower degree of interdependence amongst AND-like compositions, with strong coupling between BM pattern, both synTF expression promoters, and synTF A affinity (FIG. 33F, left). The higher inter-category coupling in the OR-like region was drivenby strong interactions between features on different synTFs right), suggesting co-optimization of expression level and activation sueiigui are impoi tain 101 tuning circuit OR-like behavior.
[0186] Clustering analysis of AND-like compositions yielded 3 major clusters. As with HFC circuits, locus design remained constant across clusters (FIG. 33G). The largest subset of AND-like compositions (clusters A and B) featured a low-affinity synTF A binding to BMs near the promoter, while synTF B with a medium-affinity ZF and a weak TA bound distally. This suggests a mechanism in which synTF B, separated from the core promoter by -110 bp, is unable to independently activate transcription but can facilitate activation by reinforcing weakly-binding synTF A, potentially through IDR-mediated interactions (HNR on synTF A, and FusN [cluster A] or HNR [cluster B] on synTF B) (FIG. 33G, right). While this “asymmetric” activation mechanism comprises the majority of AND-like composition, we observed a third, minority solution family (14%) that utilizes balanced synTF expression levels and affinities to achieve AND-like behavior. Clustering of OR-like compositions revealed 4 clusters containing designs allowing synTFs A and B to independently activate expression. This is likely facilitated by interspersed BM binding arrays, which provide both synTFs with similar access to the core promoter (FIG. 33H). Medium-affinity synTF binding is common feature of all 4 clusters, and a lack of IDRs for synTF A suggesting that synTF interaction plays a limited role in reporter activation. In summary, our mapping of multi-input circuit design space allowed for identification of non-intuitive compositional rules, and offered insights into the role that different genetic parts play in determining circuit function.
[0187] Here, we have established the feasibility of combining long- and short-read NGS modalities to perform massively parallel quantitative profiling of multi-kb length-scale genetic part compositions in human cells. CLASSIC holds potential as an approach for exploring the emergence of phenotypic function from genetic composition across a range of organizational scales and phylogenetic contexts, including for viruses58, bacterial operons59, and chromatin domains59'61. Additionally, the flexibility of our assembly workflow makes CLASSIC adaptable to a wide array of optimization tasks and phenotypic screening strategies. As we show in this piece, by enabling HT profiling of diverse genetic part compositions, our approach significantly expands the scope of inquiry for synthetic biology projects, reducing the time and cost required to identify behavior-optimized part compositions for projects that require exploration of medium- to large-scale design spaces. Though, for smaller-scale projects (<100 constructs) that can leverage well-defined design rules, traditional approaches may be morecost-effective. Furthermore, since CLASSIC leverages exte cloning and experimental analysis pipelines, we anticipate ii win increase me pace anu scaie of genetic design for diverse synthetic biology applications across a spectrum of organismal hosts, including the development of bioproduction strains62and multigenic cell therapy programs5, as well as other applications that require probing complex genetic design spaces to achieve multi-dimensional phenotypic optima.
[0188] Finally, we showed that data acquired using CLASSIC can not only train ML models to accurately make predictions for out-of-sample and edge-case circuit behaviors but, with properly crafted training strategies, can generalize to orders-of-magnitude larger genetic design spaces. Our results demonstrate the utility of using a data-driven approach to solve complex circuit design challenges, showcasing the ability to identify designs that support target circuit behavior, but also to reveal underlying design principles from across design space that may be non-intuitive and challenging to capture using biophysical modeling alone. While extensive recent work has leveraged ML approaches to develop sequence-to-function models for various classes of genetic parts28,63'65, our work serves as a starting point for developing Al-based models of gene circuit function that use part compositions as learned features. In the future, multi-modal integration of CLASSIC data with sequence-to-function data could be used to train high capacity deep-learning architectures (e.g., transformers) to enable the generative design of circuits to user-specific function. Such approaches could work in black-box fashion, without the incorporation of regulatory or biophysical priors, or synergistically with existing mechanistic frameworks to create interpretable models that provide deeper insights into genetic design66.
[0189] Methods
[0190] Plasmid library construction
[0191] Circuit libraries were cloned using a custom hierarchical golden gate30assembly scheme, which enables rapid, modular cloning of complex gene circuits starting from individual genetic part sequences (FIG. 15-17). Briefly, input DNA fragments are amplified from genomic or commercial DNA sources and cloned into kanamycin resistant sub-part entry vectors using Bpil (Thermofisher) (level O'). Sequence-verified sub-part plasmids are then used as inputs for assembly into carbenicillin-resistant entry vectors using Bsal-v2 HF (NEB) to yield genetic part plasmids (i.e., promoters, ORFs, terminators) (level 1). EUs are then constructed by assembling promoter, ORF, and terminator part plasmids into a kanamycin-resistant entry vector using Esp3I (Thermofisher) (level 2). EUs are then combined into multi-unit arrays byassembling into a carbenicillin- or spectinomycin-resistan(Thermofisher) (level 3).
[0192] In this system, barcode pools are incorporated into library assemblies at level 2 and 3 (FIG. 18). Level 2 barcoded EU pools were generated by first constructing a destination vector carrying a ccdB placeholder expression cassette downstream of the BFP stop codon. A semidegenerate 18 bp barcode oligo pool (IDT)65was polymerase-extended and cloned in place of the ccdBcassette using Bsal-v2 (NEB) in a 20 pL golden gate reaction: 2 min at 37 °C and 5 min at 16 °C for a number of cycles equal to lOx the number of input fragments, followed by a 30 min digestion step at 37 °C and subsequent sequential 15 min denaturation steps at 65 °C and 80 °C. Resulting assemblies were purified using a miniprep column (Epoch Life Sciences), electroporated (BioRad) into 100 pL NEB 10-Beta electrocompetent E. coli (NEB), and plated onto a custom 30 in x 24 in LBKan agar plate, yielding -120M colonies. Plates were grown at 37 °C for 16 h, at which time colonies were scraped into LBKan (20 mL), incubated for 30 min at 37 °C, and the plasmid library extracted via miniprep (Qiagen). For large-scale libraries level 3 barcode pools were generated in order to ensure a unique barcode-to-circuit mapping. To create level 3 barcode pools, a semi-degenerate 18 bp oligo pool (IDT) was polymerase-extended and cloned via golden gate assembly into a SpectR destination vector carrying a placeholder ccdB cassette using PaqCI (NEB) following the assembly protocol above. Resulting assemblies were transformed into 100 pL homemade chemically competent66,67Stbl3 E. coli and plated onto a 10 cm LBSpect agar plate. Plates were grown at 37 °C for 20 h to yield approximately 15,000 colonies, followed by colony scraping and plasmid DNA extraction, as described herein.
[0193] Part pools (levels 1), EU pools (level 2) and circuit libraries (level 3) were constructed by combining input plasmids at 50 fmol per part category (15 pL total volume) using the abovedescribed cycling protocol. Transformations varied by library size: level 2 and 3 assemblies for the 384-member library (FIG.2) were transformed into 100 pL of chemically competent^, coli (DH5a and Stbl3, cells / mL) and respectively plated on LBKan and LBCarb agar plates for 12-16 h (37 °C) to yield -15,000 colonies. Colonies were scraped into 13 mL LBKan or LBCarb and mini -prepped using a Qiagen kit. Level 1 and 2 assemblies for the 166k-member library (FIG. 3) were respectively transformed using 100 pL and 400 pL of chemically competent E. coli (Stbl3, cells / mL) and grown on LBCarb and LBKan agar plates at 37 °C for 12-16 h to yield -5,000 and -20,000 colonies, respectively. Level 3 166k-member assemblies were purified using a miniprep column (Epoch Life Sciences), electroporated (BioRad) into 100 pL 10-beta electrocompetent cells (NEB), and then plated on 10 15 cm LBSpect agar plates for 16-20 h (37 °C) to yield -3M colonies.
[0194] Long-read plasmid sequencing
[0195] To generate indices linking assembled constructs to men associaieuuarcoues, we used Oxford Nanopore Technology (ONT) long-read sequencing68’69. Libraries were prepared for long-read sequencing by digesting 2 pg of a level 3 plasmid library using Esp3I (Thermofisher) and purifying linearized fragments using a 0.5x volume of magnetic beads (Omega Bio-Tek) by volume. Nanopore sequencing adapters were added to the linearized DNA pool using the LSK-112 genomic ligation kit (ONT) and sequenced using on a Minion device (ONT) equipped with an RIO flow cell (FLO-MINI 12). Basecalling was performed using Guppy (ONT, super high accuracy mode) running on a GPU (Nvidia RTX 3090). Composition and barcode assignment was performed using WIMPY (what’s in my pot, y’all), a custom Matlab analysis pipeline composed of six discrete modules that converts basecalled Fastq files into a part composition array and barcode sequence for each Nanopore read (FIG. 19). Briefly, WIMPY first calls Fastqall (1) to import fastq files from Guppy and Bowtiles (2) to index the reads to a constant region in the level 3 plasmid backbone. In this study, the puromycin resistance cassette (PuroR) was designated as the constant region, and after indexing to PuroR, reads were filtered to remove incomplete assemblies and non-library fragments. WIMPY then calls Tilepin (3), which uses localized containment searching (aka tiling) to identify important genetic landmarks, such as ORFs, to define the different sections of the sequence (e.g., reporter array, synthetic transcription factor coding region, or barcoded BFP expression unit). Chophat (4) then truncates the previously identified sections into smaller arrays to aid later identification of individual parts. WIMPY then calls Viscount (5a), which determines compositions for filtered reads by identifying and assigning genetic parts through a combination of general Smith-Waterman alignments70and localized containment searches. This is done by splitting the reference sequence into 6-10bp “tiles” with a Ibp stride, after which nanopore reads are queried for the number of tiles contained for each part reference sequence. Reads containing >3% of tiles for a reference sequence are assigned while reads with more than one reference assignment are discarded. Barcode sequences for each read are concurrently determined by Barcoat (5b), in which the sequence section downstream of the BFP EU is aligned to a degenerate reference sequence using a custom alignment matrix.
[0196] Cell culture
[0197] HEK293T cells (ATCC® CRL-11268™) used in this study were cultured under humidity control at 37 °C with 5% CO2 in media containing Dulbecco’s modified Eagle medium (DMEM) with high glucose (Gibco, 12100061) supplemented with 10% Fetal Bovine Serum(FBS; GeminiBio, 900-108), 50 units / ml penicillin, 50 pg / m15070063), 2 mM L-Alanyl-L-Glutamine (Caisson labs, GLL0 eieneu io ne eanei as complete DMEM. HEK293T-LP cells with and without integrated libraries were maintained in DMEM supplemented with 50 pg / mL Hygromycin B (Sigma, H3274) and 1 pg / mL Puromycin (Sigma, P8833), respectively.
[0198] Single-copy library integration
[0199] To establish a landing pad (LP) cell line, low passage HEK293T cells were cotransfected a dual Cas9 and gRNA expression vector (pX330-T2) targeting the human AAVS1 locus71 and a Pmel-linearized (NEB) repair template comprising a constitutively expressed YFP- HygR-expression cassette with a BxBl attP recognition site flanked by 800bp homology arms for the human AAVS1 locus (pROC079) (fig. S7A). Genomic integration events were selected using 50 pg / mL Hygromycin B, and YFP+ cells were flow sorted to isolate clones (WOLF, NanoCellect). Clones were expanded to 24-well plates and tested for integration competency by integration of an attB -containing destination vector and subsequent flow cytometry and the presence of the intact LP cassette at the AAVS1 locus by PCR (FIG. 20, panel C). gDNA was extracted from each clonal population (Omega Bio-Tek) and 50 ng of gDNA was used as the template in a 25 pL reaction along with 0.25 pL of Q5 polymerase (NEB), 0.5 pL of lOmM dNTPs, 5 pL of 5X Q5 polymerase buffer, and 1 pL of each 1 mM genotyping. Primers were specific to the native AAVS1 locus, Landing Pad, or destination vector. The thermocycling protocol was 2 min at 98 °C, then 30 cycles of 98 °C for 20 s, 65 °C for 20 s, 72 °C for 1 min and then a final step of 72 °C for 5 min. All genotyping primer sequences are listed in Appendix A. We identified a clone (LP79 #36) that demonstrated BxBl -dependent, single copy integration of vectors harboring an attB recognition site and used this cell line for all following experiments.
[0200] To integrate individual constructs, 125k cells were plated into a 24-well plate one day prior to transfection. The well was co-transfected with 250 ng of a level 3 vector containing a construct of interest and 125 ng of BxBl expression plasmid (Addgene #51271) using JetPrime (VWR) (FIG. 20, panel B). pCAG-NLS-HA-Bxbl was a gift from Pawel Pelczar (Addgene plasmid #51271)72. Two days after transfection, cells were passaged into complete DMEM containing 1 pg / mL Puromycin and selected for 10 days to achieve homogenous expression. Library integration was performed by co-transfecting HEK293T-LP with 250 ng of plasmid library and 125 ng of BxBl per 200k cells (384-member library = IM cells transfected; 166k-member library = 100M cells transfected) using JetPrime. Media exchange was performed after 8 h and cells were expanded for an additional 40 h, split into five culture dishes, and grown for 6-8 daysunder Puromycin selection. Cells were passaged 1:3 and cultu puromycin selection, and then combined in complete DMEM for now sorting t I UIVI uens / mnj.
[0201] To determine the clonal variability within LP-integrated cell populations, we integrated a constitutively expressed mCherry EU following the protocol described above. We then randomly sorted 23 clones from the population (WOLF, NanoCellect) and measured their mCherry expression (FIG. 20, panel D). We calculated the MAE for this set (MAE = 0.08) and used this value to define an expected error range from clonal heterogeneity (ERCH), which represents the expected variability of cells from an integrated population. For the 384-member library, the expression ERCH was set to 0.08 on either side of the expected value, and for the 166k-member library, the fold change ERCH was calculated to be 0.16 on either side of the expected value.
[0202] Flow sorting
[0203] To prepare libraries for flow sorting, cells were lifted using TrypLE (Gibco), washed twice with PBS, and resuspended in complete DMEM (10M cells / mL). For the 384-member library, the mRuby expression distribution was sorted into 10 evenly log-spaced bins on a Sony MA900 set to purity mode (FIG. 23). The number of cells collected for each bin was proportional to the percentage of the mRuby distribution expressed in that bin, with a target of ~125k cells sorted in the most populous bin [e.g., 2,252 cells were collected for bin 1 (0.53% of library), 126,562 for bin 4 (29.78% of library)]. For the 166k-member library, cells were split and either grown in DMEM containing IpM 4-OHT (Sigma Aldrich) (“induced”) or without 4-OHT (“uninduced”) for 72 h (FIG. 27). The eGFP expression distributions for both conditions were then sorted separately into 8 evenly log-spaced bins. Cells were sorted into bins in proportion to the relative abundance of the library in each bin, with the most populous bin set at 1.5M cells. For both libraries, sorted cells were plated into 96-, 24- or 6-well plates to achieve a plating density of 10-30%, grown under puromycin selection for 2 days, passaged and washed with PBS, and grown for an additional 3 days. Cells were then lifted using TrypLE and total RNA was extracted using the Takara RNA plus kit following manufacturer’s instructions (Takara). 5 pL of total mRNA was converted to total cDNA using random hexamers following manufacturer’s instructions (Verso, Life Technologies).
[0204] Short-read (Illumina) sequencing
[0205] To prepare sorted libraries for short-read sequencing, barcode regions were PCR amplified from cDNA in a 15 pL reaction with Phantamax polymerase (Vazyme) and custom bin-specific primer sets (IDT) using the following thermocycling protocol: 2 min at 98 °C, then 12 cycles of 98 °C for 20 s, 60 °C for 10 s, 66 °C for 10 s, 72 °C for 20 s and then a finalstep of 72 °C for 2 min. The resulting amplicon pool was extracPCR step, which used 10 cycles of the thermocycling protocol for PCrc step i, auueu wuninia sequencing adapters (i5 and i7) and sample-specific sequencing barcodes using Phantamax and custom primer sets (IDT), and the amplified product was extracted from a 2% agarose gel. All primer sets used for NGS are reported in Appendix A. Purified amplicons from each bin were pooled at equimolar concentrations and sequenced using an Illumina MiSeq (kit v2, 300 cycles) or NovaSeq6000 (Sp vl .5 kit, 300 cycles) for the 384- and 166k-member libraries, respectively. Illumina data were analyzed using a custom analysis pipeline (Matlab). Briefly, fastq files are imported and split according to their sample-specific sequencing barcode. Barcode sequencing data is converted to average expression for all compositions / variants as follows: (1) Counts for each barcode in each sample / bin are normalized by the percentage of the library distribution in that bin to account for differences in read depth across bins. (2) A weighted average of number of reads in each bin for each barcode is then calculated, to assign an average expression level (using a bin —> expression conversion table) to each barcode. (3) Barcode expression levels are then cross- referenced with the Nanopore indices to assign average expression values for all barcodes associated with a given variant in a n-by-y cell array, where n is the number of library variants and y is the (variable) number of barcodes for each variant. (4) For each variant, a gaussian kernel density estimation is performed over the log-normalized expression values from all barcodes associated with each variant, using the Matlab function ksdensity. The expression value corresponding to the kernel peak is assigned as the mean expression value for each variant.
[0206] Experimental validation of CLASSIC measurements
[0207] To assess the accuracy of CLASSIC measurements for both the 384-member and 166k-member libraries, single cells were sorted from the integrated cell populations into 96- well plates (MA900) and clonally expanded from across the mRuby and eGFP expression distributions, respectively. Isolated clonal populations, which are referred to as ‘isolates’, were expanded to 24-well plates and the fluorescence distribution was measured by flow cytometry (MA900). For the 384-member library, one 96-well plate of cells was sorted based on mRuby expression and 15 isolates were taken for validation (FIG. 6, panel B). For the 166k-member library, two 96-well plates of cells were sorted based on eGFP expression from the uninduced cell library to yield 40 isolates total (FIG. 9). Following fluorescence measurement, gDNA was extracted from each cell population (Omega Bio-Tek). The genetic variant in each isolate was identified by PCR-based amplification and Nanopore sequencing of the Landing Pad locus using a forward primer in Puromycin resistance cassette and a reverse primer in the BFP open reading frame. A set of 30 different forward primers, each with a defined barcode, were paired with acommon reverse primer to amplify the genetic variants from ea variant with the isolate during Nanopore sequencing. 50 ng of guiM was useu as me lempiaie m a 20 pL reaction along with 10 pL of 2X PhantaMax Master Mix (Vazyme), 1 pL of each 1 mM primer, and brought up to 20 pL in nuclease-free H2O. The thermocycling protocol was 2 min at 98 °C, then 30 cycles of 98 °C for 20 s, 65 °C for 20 s, 72 °C for 5 min and then a final step of 72 °C for 10 min. Each PCR reaction was extracted from a 1% agarose gel and column purified using QG buffer (Qiagen), following manufacture instructions. Purified amplicons were pooled and prepared for Nanopore sequencing using LSK-112 genomic ligation kit (ONT) and sequenced using on a 19inion device (ONT) equipped with an R10 flow cell (FLO- MINI 12). Basecalling was performed using Guppy (ONT, super high accuracy mode) running on a GPU (Nvidia RTX 3090). Composition and barcode assignment was performed using WIMPY, as described herein.
[0208] Random forest regression
[0209] For the 384-member EU library, a bootstrap aggregation (bag) RF model was constructed using the matlab function templatetree (FIG. 24). The promoter, Kozak and terminator were used as categorical input variables for each EU, with log(mRuby) expression (AU) as the predicted output. The data was split 80:20 (307 EUs and 77 EUs) for training and testing, respectively. The number of trees was chosen based on 10-fold cross validated RMSE loss as a function of increasing number of trees using the Matlab function fitrensemble . The trained RF model was then validated on the test set (FIG. 2, panel F). Identification of genetic part interactions was determined by the absolute log error between predicted and observed mRuby expression.
[0210] For the 166k-member circuit library, two independent least-squares boosted RF models were used to predict basal and induced expression for experimentally unmapped circuit members using the same Matlab functions mentioned above (FIG. 28). Data used for RF training were obtained by first filtering the dataset and keeping points that had 12 or more barcode reads. The space was then split into 8 equally log-spaced regions. If the bin with the highest frequency of data contained more than 2x the number of points as the next most frequent bin, data was randomly downsampled from that bin to match the frequency of the next most frequent bin. This downsampled dataset was then randomly split 80:20 into a training and test set respectively. The inputs to the model were the 10 categorical variables representing the parameters tuned in the library (e.g., SynTF affinity, #BM etc.), while the output was the log(eGFP) expression (AU) for basal or induced expression values. Following hyperparameter optimization for number of trees, least squares boosting algorithm learn rate, and maximum decision splits, the model structuresand performance were as follows: basal, 70 trees, max numbe r2fra / n=0.88, / '2zez=0.83 and induced, 100 trees, max number or spms-o i, warn raie-v.ujo, / '2Zz'zzzz-z=0.92, / '2test=O .87.
[0211] Mutual information calculation
[0212] For a given pair of features in a region of design space, mutual information was computed using the following formula:where x and y are 2 different features of the library, and p(x, y) represents the joint frequency for specific x and y combinations divided by the total number of points in the dataset. p(x) and p(y) represent the fractional frequency of the corresponding parts by themselves.
[0213] UMAP clustering
[0214] For UMAP dimensional reduction, part composition data for the seven genetic part categories with the greatest part usage asymmetry (promoter, AD, IDR, ZF affinity, BM #, core promoter, and orientation) from the high fold change (>25-fold region) were used as inputs into the cluster evaluation function in Matlab. We carried out a gap test to identify the optimal number of clusters in the data, followed by K-means clustering using the kmeans function in Matlab to identify clusters. Cluster regions were generated by drawing a boundary around the outermost points corresponding to a particular cluster.
[0215] Data analysis and statistics
[0216] All statistical analyses and graphical displays were performed in Matlab (2022a). Pearson’s r2was used to calculate correlations between CLASSIC replicate experiments, CLASSIC expression measurements and model predictions, and individual expression measurements and their model predictions. The MAE was calculated as the mean of the absolute differences of the logic transformed expression between every pair of isolate and CLASSIC measurements. The MAE of the logic transformed fold change values was computed between CLASSIC measurements and individual measurements as well as model predictions and individual measurements.
[0217] Rationale for using pre-trained ML models to expand design space.
[0218] As we demonstrated in the previous section, we can assay large design spaces using CLASSIC and use the data to train ML models to accurately predict the behavior of unmeasured circuit compositions. We next wanted to test whether it is possible to useCLASSIC+ML to explore circuit design spaces that are ord< be experimentally measured. As a showcase for this approacn, we sei out io uuiiu a nuiaiy or multi -input circuits composed of two orthogonal synTFs that respond to 4-hydroxytam oxifen (4-OHT) and grazoprevir (GZV) to induce transcription of GFP, hereafter respectively referred to as synTF A and synTF B. Boolean logic gates have historically been challenging to design and implement in mammalian cell contexts, and design rules for stably integrated, single-copy circuits have not been reported. We asked if we could use a data-driven design approach to identify multi-input circuit architectures that show AND-gate and OR-gate behavior (FIG. 32A). Given the greater design complexity of this circuit class, identifying behavior-optimized designs requires probing of a 104- 105x larger design space. Based on our exploration of part function in the EU and single-input libraries, we hypothesized an appropriate diversity would include the following parts (FIG. 32B): 6 promoters, 5 ADs, 4 IDPs, and 6 ZF affinities 4 core promoters, 3 spacer sequences, and BM arrays composed of 3 different numbers of BMs arranged in 4 possible patterns, all in 6 possible EU orientations — an overall design space of consisting of 47 features and a total of 3,359,232,000 genetic circuits. Because this space vastly exceeds what we previously constructed and assayed using CLASSIC (~105), we set out to identify an iterative library construction and model training strategy that would allow us to effectively map the space in its entirety (FIG. 34 A).
[0219] To identify such a strategy, we used data from the 166k-member library to simulate two approaches using small datasets to learn a large space (FIG. 34A-B). Approach A consists of first generating a small library (<1% of total design space) of high-quality measurements (>10 barcodes circuit'1) using a subset of the genetic parts (11,446 / 10,368 members). A second, high-quality library is then created in which every circuit contains one part that had been ‘held out’ from the previous library. Approach B comprises two small libraries that randomly sample the total space (165,888 members). In both approaches, the model is trained on the first, “base” library, and fine-tuned on the second. We found, for both the basal and induced conditions, Approach A produced a superior test set r2after the first (0.88 vs. 0.82 and 0.87 vs. 0.81, respectively) and second (0.85 vs 0.79 and 0.88 vs 0.82, respectively) iterations (FIG. 34C-D, top). Additionally, during training, Approach A was more computationally efficient, as it converged after fewer cycles (FIG. 34D, bottom). Based on these results, we designed and constructed the multi-input libraries in accordance with Approach A.
[0220] Using CLASSIC libraries to map larger genetic design spaces.
[0221] To comprehensively map the 3.4B-member multi-i out 1-3 parts from each category to reduce the design space iroin J .HD (uie run space ) io 1.1M members (the “base library space”) (FIG. 32C). This base library was used to collect high-confidence measurements on -3-7% configurations in the base design space and train a model (the “base model”). A second library (the “expansion library”) was then constructed by adding back previously held-out parts to train the model to map the full space. For this library, each of these sub-libraries was assembled to contain only one new part, thus ensuring that part is represented in order to allowing the fine-tuned model (the “expansion model”) to more accurately infer unseen part compositions.
[0222] Construction of base multi-input library.
[0223] To construct the multi -input circuit base library, we started by removing 22 parts from the full design space (FIG. 32C) and followed the standard Sesame Street library construction scheme describedherein. For each of the inducible synTFs (A or B), level 1 ORF plasmids were generated by assembling a level 1 destination vector with level 0 plasmid pools encoding 4 ADs, 2 IDPs, and 4 DBD affinity mutants (Fig. 35A-B). Level 2 synTF EU pools for each inducible system were generated by assembling a position-specific level 2 destination vector (AB, BC, or CD) with a level 1 plasmid pool containing 3 promoters, the previously assembled level 1 open reading frame library, and 1 terminator (SV40 pA) with a single spacer length of 1 kb. Position-specific, level 2 assembly reactions were transformed into lOOpL of chemically competent DH5a E. coli and plated on a 10cm LB-Agarxan plate. Level 2 reporter plasmid pools were generated by assembling a position-specific level 2 destination vector (AB, BC, or CD) with a pool of level 1 plasmids encoding an array of either 2, 8, or 12 BMs for either the 10-1 or 1-3 ZF arranged in 4 different patterns upstream of 2 minimal promoters, an eGFP open reading frame level 1 plasmid, and 1 level 1 terminator (SV40 pA) plasmid containing a single spacer length of 1 kb. Similarly, position-specific assembly reactions were each transformed into lOOpL of chemically competent DH5a E. coli and plated on a 10cm LB- Agarx an plate.
[0224] To ensure high-depth sub-sampling of the large design space, we picked 32 individual colonies from each of the 6 level 2 synTF EU pools and 20 individual colonies from each of the 3 reporter pools (FIG. 35A-B). Individual cultures from each pool were grown overnight in 1 mL of TBKan in 96-well blocks before pooling 200pL from each culture and miniprepping to yield 9 position-specific plasmid pools. This method, which we term ‘bottlenecking’, artificially caps the maximum possible diversity of the library at 122,880 totalvariants. This value matches the upper limit of what we siCLASSIC experiments and allows collection high-quality measurements uy increasing me likelihood each variant is linked to multiple barcodes. We then generated 6 orientation-specific level 3 plasmid pools by assembling the appropriate level 2 pools with a barcoded BFP EU and a barcoded level 3 destination vector. Each reaction was transformed into 15 pL of electrocompetent NEB10B E. coli cells to yield ~200k colonies per orientation, representing approximately 10-fold more colonies than possible circuit compositions. This library was scraped, miniprepped, and sequenced across 3 MinlON flow cells (Nanopore) (FIG. 35C). Interestingly, we observe only -22% of all reads in the expected range of our library (-14-20 kb), which is significantly lower than for the 384-member EU and 166k-member libraries (FIG. 35C). We hypothesize this is a consequence of the increased size and complexity of the final assembly step for the multi-input induction circuits. However, despite lower assembly fidelity for the level 3 circuit pool, our bottlenecking strategy allowed us to maintain an approximately equal proportion of each genetic part. Additionally, the >95% of barcodes identified by nanopore sequencing were uniquely linked to one circuit composition. We then co-transfected the level 3 barcoded circuit library and a BxBl expression vector into 2»108HEK-LP cells, yielding an estimated l»106to 2»106integration events (FIG. 35B, see Methods).
[0225] Flow sorting of base multi-input library.
[0226] Following puromycin selection of the library-integrated cell population for 14 d, 2»108cells were split evenly and cultured with 1 pM 4-OHT, 20 pM GZV, both inducers, or neither inducer. After 72 h, cells were harvested and prepared for flow-seq (see Methods). As with the EU and single-input libraries, cells were gated based on size characteristics and BFP expression and sorted into 8 equally log-spaced GFP gates (FIG. 32D). Cells from each sorted bin were grown for 5 d post-sort to ensure the complete removal of free RNA species that were aberrantly sorted. All multi-input libraries show a large GFP-negative peak, which is likely a result of a larger fraction of the library representing integration-competent, incomplete assembly products that were co-transfected along with the correctly assembled library members (FIG. 32D)
[0227] Construction of expansion multi-input library.
[0228] To construct the expansion library, we re-introduced the 24 previously held-out parts (FIG. 32B, bottom). In accordance with the approach described herein and shown in FIG. 34, we designed the expansion library by randomly sampling a small design space around eachnew part. To provide high-quality, unbiased training data, u each new part (FIG. 36A, B). These libraries contained only one instance 01 eacn new pan, with the remainder of the parts sampled from the base design space. To ensure that orientation and inducible (ERT2 or NS3) contexts were sampled evenly, for each new part we constructed separate libraries for each EU position (AB, BC, or CD) and each inducer. Altogether, the addition of new parts required construction of 114 separate 250-member level 3 libraries, but expanded the learnable design space by over 3 orders of magnitude. To maximize measurement accuracy, we once again intentionally capped our total library diversity, this time at ~30k members. To ensure representative sampling around each new part while maintaining high- quality measurements, we picked 10 colonies from each level 2 new part library, grew them separately to saturation in 1 mL TBKan, pooled them, and then miniprepped the combined culture using the Star Liquid Handler (Hamilton) to produce 10-member pools (FIG. 36A, B). Similarly, to ensure only one new part was introduced in each library, for each new part we randomly sampled 5 level 2 synTF (X and / or Y) or reporter compositions by picking colonies, pooling, and miniprepping. This process allowed us to generate 250 circuit designs for each new part library, which, across 6 different orientations, provided 1,500 total instances.
[0229] As an example, to add VP 16 to the design space, we first constructed level 1 ORF libraries for the ERT2 and NS3 systems with VP16 as the only part at the AD position. This library was then assembled with the level 1 promoter and terminator pools from the original design space into a level 2 position-specific destination vector (AB, BC, or CD) to yield 6 total plasmid pools. Each plasmid pool was transformed into 20 pL of chemically competent DH5a E. coh. plated on LB-AgarKan, and grown overnight. 10 colonies were picked from each of the 6 plasmid pools, grown overnight, and combined to generate 6 position-specific pools of 10 members each. Each of the 10-member VP 16 pools was then assembled with a barcoded level 3 destination vector, a barcoded BFP EU, and the appropriate combination of 5-member ERT2, NS3, and reporter pools composed of original parts, such that VP16 was represented in both synTF pools and all 6 possible orientations.
[0230] All 114 of the 250-member level 3 new part libraries were assembled separately, pooled according to their orientation, column-purified, and transformed into 7.5 pL of electrocompetent NEB 10B E. coli. This process yielded ~100k colonies per orientation, which were scraped, miniprepped, pooled, and sequenced with two MinlON flow cells (FIG. 36C). As seen for the 3.4B-member base library, we observe -28% of all reads in the expected range of our library (-14-20 kb) (FIG. 36C). We then co-transfected the level 3 barcoded multi-inputcircuit library and a BxBl expression vector into 2»108HE] l»106- 2»106integration events (FIG. 36B, see Methods).
[0231] Flow sorting of expansion library multi-input library.
[0232] Following puromycin selection of the library-integrated cell population for 14 d, 2»108cells were split evenly and cultured with 1 pM 4-OHT, 20 pM GZV, both inducers, or neither inducer. After a 72 h induction, cells were harvested and prepared for flow-seq (see Methods). Cells were gated based on size characteristics and BFP expression and sorted into 8 equally log-spaced GFP gates. Cells from each bin were grown for 5 d post-sort to ensure complete removal of aberrantly sorted RNA. All multi-input libraries show a large GFP- negative peak, likely resulting from a larger fraction of the library being integration-competent, incomplete assembly products that were co-transfected along with the correctly assembled library members (FIG. 35-36).
[0233] Validation of base and expansion library measurement accuracy.
[0234] Illumina and Nanopore data were processed as described for the 166k-member library. We isolated and expanded up to 30 clones from each of the integrated base and expansion libraries and measured their GFP expression across 4 conditions (no inducer, OHT only, GZV only, and both inducers) via flow cytometry (FIG. 32D, right). Given the size of the 3 -gene circuits (-14-20 kb), we elected to genotype the circuit compositions in each clonal population by genomically amplifying and sequencing the barcode region. From the pool of expanded clones, we made unique barcode to composition matches for 14 and 13 clones from the base and expansion libraries, respectively. When we compared mean GFP expression values obtained via flow cytometry with CLASSIC measurements, we observed good agreement across all 4 conditions (MAEbase = 0.28, MAEexpansion = 0.23) (FIG. 32D, right).
[0235] Data analysis for base design space.
[0236] To train a ML model for the 1. IM-member base design space, we first pre-processed the data using the following pipeline: 1) GFP expression outputs from all circuit measurements across all 4 conditions (no inducer, 4-OHT only, GZV only, and both inducers) were log 10 transformed; 2) Any circuits lacking all 4 measurements were discarded; and 3) Circuit parts were one-hot encoded and flattened, with extra slots included to accommodate new parts added in the second step via fine-tuning. We selected a high a quality test set consisting of variants with >50 barcoded reads and held them aside for post-training model validation. The remaining data was split 90: 10 into training and validation sets. The validation dataset was used for monitoring training progress, hyperparameter optimization. The validation dataset was alsoused for early stopping, employing a patience paramete optimization, and 300 during final model training. During iiype paiaineiei optimization, we first co-varied the learning rate and number of layers in the model in a 2D scan, and also varied the amount of training data from ~3%-100% of the total collected data to test if we were learning in a data sparse regime, or if the dataset could be further down sampled. The scan confirmed that training with the entire dataset, along with 5 hidden layers with an intermediate learning rate of 5 x 10'3was optimal. We next varied the momentum parameter within the SDGM solver from 0 to 0.99, finding that momentum 0.9 was optimal. Finally, we varied the solver (SGDM, ADAM, RMSprop, and LFBG-S), and found that our initial choice of SGDM with the optimized momentum parameter outperformed other solvers. The model was then trained with a relaxed validation patience metric of 300 iterations for a maximum of 200 possible epochs. The final training curves showing the RMSE and model loss during training as a function of iteration (iteration = 1 batch of data, i.e., 256 data points) are described herein. The model was then evaluated on the held-out high-quality test set and achieved remarkably high r2values of 0.8-0.9 for the different induction conditions. The model also demonstrated good agreement with random isolates shown in FIG. 32D.
[0237] We used the model to predict expression levels for the entire design space, and computed KL-divergence between the expression levels across all 4 induction conditions to that of “ideal” AND and OR gates. The KL divergence between two distributions was computed using the following equation:Where q(x) is the probability distribution of the ideal gates, represented by ID vectors [0.01, 0.01, 0.01, 0.97] for ideal AND, and [0.01, 0.33, 0.33, 0.33] for ideal OR gates for no inducer, 4-OHT only, GZV only, and both inducers, respectively. Ideal OR-like configurations are represented by low KL divergence against an ideal OR gate and high divergence from ideal AND gate, while good AND gates follow the opposite trend. Computation of the two KL divergences on the entire ~1.1 million members in the design space revealed that only a small fraction of the space produced productive AND gate or OR gate designs (FIG. 33B).
[0238] Data analysis for expansion design space.
[0239] We next sought to expand model-predicted mapping to all 3.4B design space members via fine tuning. Data was processed and split into training, validation, and high qualitytest sets (see section 3.4). For the high-quality test set, the thr< with >75 barcoded reads, which was enabled by the higher reau uepm we ouiameu ror mese libraries (FIG. 32C, right). The model was then fine-tuned by initializing with the pre-trained weights from the base model and training with a learning rate of 5 • 10'3and batch size of 256. Approximately 10% of the data from the base library was also included during training to preserve base model memory as the model learns new part behavior. Model performance on the new parts increased significantly after fine tuning, with r2values increasing from -0.63- 0.68 to -0.79-0.86 for the high quality test set. The model’s pre-fine-tuning performance was expected since the sub-library members contain only one new part, while all other components of the compositions are from parts within the base library space. As such, the model is still able to predict circuit behavior reasonably well due to the amount of variance explained from parts that have already been learned. To better understand the amount of information learned during fine tuning, we grouped the points in the high quality test set by part, and observed r2values before and after fine tuning Model predictions for 21 / 22 (-95%) of the new parts improved with fine-tuning. The only exception was DDX4, where performance decreased, most likely due to there being only 8 measurements for this part in the test set. We validated model performance by individually constructing 13 circuit designs and testing their behavior. We found that, once again, the model predictions were extremely accurate despite training on an extremely small fraction of the design space (0.001%). We used the model to compute KL- divergence-based AND and OR like characteristics of the entire design space, finding the significant expansion in all directions, particularly for the OR-like behavior region (FIG. 33A).
[0240] We next constructed and tested 36 individual compositions from across the complete multi-input design space, including 1 circuit that measured in the base library (in-sample) and 35 unmeasured out-of-sample circuit (FIG. 33C). This set also included circuits from the AND-like and OR-like regions. We observed good agreement between the individually constructed cell lines and the model predictions for both in-sample (MAE = 0.21) and out-of- sample (MAE = 0.32) compositions. Altogether, these results validate our strategy of using an ML modelling framework to fine-tune a model using small auxiliary datasets containing new genetic components to predict the behavior of circuits containing combinations of new parts.
[0241] Extended multi-input library data analysis.
[0242] The mapping of multi-input circuit design space enabled by our MLP provides the opportunity to comprehensively determine what parts and features are enriched in regions of AND-like and OR-like behavior (FIG. 33). After calculating DKLAND and DKLOR values forcircuit compositions across the expanded design space, we irDKLOR < 0.005) capturing behavior space distribution tails containing ~ UK circuits iroin eacn region — a set we hypothesized would allow us to use statistical analyses to assess trends in part usage (FIG. 33B). After calculating the frequency of each feature in these regions, we observed depletion and enrichment of specific parts (FIG. 33D). Interestingly, in both regions, 13 of 30 part categories were enriched for newly added parts, suggesting that the model-expanded space provided numerous additional design strategies for high-performing circuits. We also calculated MI values, which revealed the strength of pairwise part interactions within and between EUs. We found that AND-gate designs demonstrated the strongest couplings between the following combinations: BM pattern and ZF affinity for synTF A, BM pattern and the synTF B promoter, synTF A and synTF B promoters, and the synTF B promoter and IDP.
[0243] To identify design rules for constructing AND-like and OR-like compositions, we once again computed SHAP values for 200 randomly sampled compositions within the thresholded design space regions. We found that AND-like designs are dominated by a single reporter array design - high BM valency, BMs arranged in clusters with synTF A BMs proximate to a mCMV core promoter - and are arranged in Orientation 2 (i.e., synTF Areporter synTF B) with large spacers following the synTFs and a shorter spacer at the end of the reporter. Clustering on all part features from the top ~24k compositions in the AND-like region (AND DKL < 0.12) generated 3 clusters (FIG. 33G. We assessed cluster stability as described for the single-input libraries (section 3.5) yielding similarity and rand scores of 0.89 and 0.71, respectively . Across the three clusters, we identified two distinct solution families, defined by the apparent degree of interaction between the two synTFs. The majority design solution (clusters A and B) appears to require that weak binding of synTF A is reinforced by synTF B through interactions between the HNR IDP on synTF A and an aggregation-prone IDP (HNR or FusN) on synTF B. This design consists of a low affinity synTF A containing HNR and a medium activity TA (VP 16 or VP64) coupled to a medium affinity synTF B containing low or medium activity TAs (p65 or VP 16) and either FusN (cluster A) or HNR (cluster B). In cluster A, high expression of the synTF A is balanced by low expression of the synTF B, and vice versa for cluster B. Notably, our best performing AND gate design (FIG. 33C, middle) falls into the design strategy detailed in cluster B. An alternate, though significantly smaller, solution in this behavior space (cluster C) involves the use of two medium-affinity ZFs, suggesting that both synTF A and B are equally capable of binding to the reporter array and do not require reinforcement. Furthermore, the flexibility at the IDPposition on synTF B indicates that these compositions do interactions, though other features of the synTF may still inieraci (J&UJOUO^, Z JZJZUHJ. Interestingly, this minority solution makes use of a distinct BM pattern in which synTF B binds closest to the core promoter.
[0244] When we examined compositions in the OR-like region of behavior space, we observed vastly different part usage and SHAP values (FIG. 33D, E). Clustering on all part categories yielded 4 clusters (FIG. 33H). We again assessed cluster stability, obtaining similarity and rand scores of 0.75 and 0.64, respectively. As with the AND gates, most OR- like compositions generally make use of a common reporter architecture (FIG. 33H), which is composed of 12 interspersed synTF A and synTF B BMs upstream of an SCP2 core promoter. Orientation 3 (i.e., synTF BsynTF A reporter) dominates the OR gate designs, with large spacers following the synTFs but a slightly shorter spacer after the reporter. From the four clusters, two solution families emerge, which are defined by their synTF part usage. The widespread use of medium affinity ZFs and the pairing of aggregation-prone IDPs (FusN or HNR) (28041848) with no IDP may suggest that compositions in this region of behavior space do not require optimization of interactions between the synTF A and synTF B, but instead utilize two synTFs of similar activity and expression level (FIG. 33H). Clusters A and D utilize high expression of a synTF B that contains a medium activity TA (VP 16) paired with a synTF A that is respectively expressed at low levels with a medium activity TA (VP 16) or at high levels with a medium-to-weak (NFZ, VP16, or p65) TA. The most OR-like individually constructed OR gate closely resembles the design of cluster A (FIG. 33C, bottom). Clusters B and C contain highly expressing synTF B with the most active TA, VPR, and highly expressing synTF A containing either medium (VP 16, VP64, or NFZ) or low (p65) activity TAs, respectively. We hypothesize that these OR gate solutions represent some of the best designs for their respective inducible system (4-OHT or GZV).
[0245] References Cited in this Example:1 English, M. A., Gayet, R. V. & Collins, J. J. Designing Biological Circuits: Synthetic Biology Within the Operon Model and Beyond. Annu Rev Biochem 90, 221-244, doi: 10.1146 / annurev-biochem-013118-111914 (2021).2 Mahata, B. et al. Compact engineered human transactivation modules enable potent and versatile synthetic transcriptional control. bioRxiv, 2022.2003.2021.485228, doi:10.1101 / 2022.03.21.485228 (2022).3 Slusarczyk, A. L., Lin, A. & Weiss, R. Foundations f implementation of synthetic genetic circuits. Nat Rev Genet doi: 10.1038 / nrg3227 (2012).4 Bashor, C. J. & Collins, J. J. Understanding Biological Regulation Through Synthetic Biology. Annu Rev Biophys 47, 399-423, doi: 10.1146 / annurev-biophys-070816-033903 (2018).5 Bashor, C. J., Hilton, I. B., Bandukwala, H., Smith, D. M. & Veiseh, O. Engineering the next generation of cell-based therapeutics. Nat Rev Drug Discov, doi: 10.1038 / s41573- 022-00476-6 (2022).6 Beitz, A. M., Oakes, C. G. & Galloway, K. E. Synthetic gene circuits as tools for drug discovery. Trends Biotechnol 40, 210-225, doi: 10.1016 / j.tibtech.2021.06.007 (2022).7 Kitada, T., DiAndreth, B., Teague, B. & Weiss, R. Programming gene and engineered-cell therapies with synthetic biology. Science 359, doi: 10.1126 / science.aadl067 (2018).8 Cameron, D. E., Bashor, C. J. & Collins, J. J. A brief history of synthetic biology. Nat Rev Microbiol 12, 381-390, doi:10.1038 / nrmicro3239 (2014).9 Yeung, E. et al. Biophysical Constraints Arising from Compositional Context in Synthetic Gene Networks. Cell Sy st 5, 11-24 el2, doi: 10.1016 / j . cels.2017.06.001 (2017).10 Lou, C., Stanton, B., Chen, Y. J., Munsky, B. & Voigt, C. A. Ribozyme-based insulator parts buffer synthetic circuits from genetic context. Nat Biotechnol 30, 1137-1142, doi: 10.1038 / nbt.2401 (2012).11 Muller, I. E. et al. Gene networks that compensate for crosstalk with crosstalk. Nat Commun 10, 4028, doi: 10.1038 / s41467-019-12021-y (2019).12 Prindle, A. et al. Rapid and tunable post-translational coupling of genetic circuits. Nature 508, 387-391, doi: 10.1038 / naturel3238 (2014).13 Shaw, W. M. et al. Engineering a Model Cell for Rational Tuning of GPCR Signaling. Cell 177, 782-796 e727, doi: 10.1016 / j. cell.2019.02.023 (2019).14 Kinney, J. B., Murugan, A., Callan, C. G., Jr. & Cox, E. C. Using deep sequencing to characterize the biophysical mechanism of a transcriptional regulatory sequence. Proc Natl AcadSci USA 107, 9158-9163, doi: 10.1073 / pnas.1004290107 (2010).15 Sharon, E. et al. Inferring gene regulatory logic from high-throughput measurements of thousands of systematically designed promoters. Nat Biotechnol 30, 521-530, doi: 10.1038 / nbt.2205 (2012).16 Noderer, W. L. et al. Quantitative analysis of mammalian translation initiation sites by FACS-seq. Mol Sy st Biol , 748, doi: 10.15252 / msb.20145136 (2014).17 Kosuri, S. et al. Composability of regulatory sequences controlling transcription and translation in Escherichia coli. Proc Natl Acad Sci USA 110, 14024-14029, doi: 10.1073 / pnas.1301301110 (2013).18 Arnold, C. D. et al. Genome-wide quantitative enhancer activity maps identified by STARR-seq. Science 339, 1074-1077, doi: 10.1126 / science.1232542 (2013).19 de Boer, C. G. et al. Deciphering eukaryotic gene-regulatory logic with 100 million random promoters. Nat Biotechnol 38, 56-65, doi: 10.1038 / s41587-019-0315-8 (2020).20 Taskiran, II et al. Cell-type-directed design of synthe220, doi: 10.1038 / s41586-023-06936-2 (2024).21 Gosai, S. J. et al. Machine-guided design of cell-type-targeting cis-regulatory elements. Nature 634, 1211-1220, doi: 10.1038 / s41586-024-08070-z (2024).22 Sahu, B. et al. Sequence determinants of human gene regulatory elements. Nat Genet 54, 283-294, doi: 10.1038 / s41588-021-01009-4 (2022).23 Agarwal, V. et al. Massively parallel characterization of transcriptional regulatory elements. Nature, doi:10.1038 / s41586-024-08430-9 (2025).24 Bogard, N., Linder, J., Rosenberg, A. B. & Seelig, G. A Deep Neural Network for Predicting and Engineering Alternative Polyadenylation. Cell 178, 91-106 el23, doi: 10.1016 / j. cell.2019.04.046 (2019).25 Khoroshkin, M. et al. A generative framework for enhanced cell-type specificity in rationally designed mRNAs. bioRxiv, doi: 10.1101 / 2024.12.31.630783 (2024).26 Gera, T., Jonas, F., More, R. & Barkai, N. Evolution of binding preferences among whole-genome duplicated transcription factors. Elife 11, doi: 10.7554 / eLife.73225 (2022).27 DelRosso, N. et al. Large-scale mapping and systematic mutagenesis of human transcriptional effector domains. bioRxiv, 2022.2008.2026.505496, doi: 10.1101 / 2022.08.26.505496 (2022).28 Angenent-Mari, N. M., Garruss, A. S., Soenksen, L. R., Church, G. & Collins, J. J. A deep learning approach to programmable RNA switches. Nat Commun 11, 5057, doi : 10.1038 / s41467-020- 18677- 1 (2020).29 Jones, E. M. et al. Structural and functional characterization of G protein-coupled receptors with deep mutational scanning. Elife 9, doi: 10.7554 / eLife.54895 (2020).30 Zhou, Y. et al. Encoding Genetic Circuits with DNA Barcodes Paves the Way for Machine Learning- Assisted Metabolite Biosensor Response Curve Profiling in Yeast. ACS Synth Biol 11, 977-989, doi: 10.1021 / acssynbio.lc00595 (2022).31 Wong, A. S., Choi, G. C., Cheng, A. A., Purcell, O. & Lu, T. K. Massively parallel high-order combinatorial genetics in human cells. Nat Biotechnol 33, 952-961, doi: 10.1038 / nbt.3326 (2015).32 Matreyek, K. A. et al. Multiplex assessment of protein variant abundance by massively parallel sequencing. Nat Genet 50, 874-882, doi: 10.1038 / s41588-018-0122-z (2018).33 Castellanos-Rueda, R. et al. speedingCARs: accelerating the engineering of CAR T cells by signaling domain shuffling and single-cell sequencing. Nat Commun 13, 6555, doi : 10.1038 / s41467-022-34141 -8 (2022).34 Liu, H. et al. Magic Pools: Parallel Assessment of Transposon Delivery Vectors in Bacteria. mSystems 3, doi: 10.1128 / mSystems.00143-17 (2018).35 Weber, E., Engler, C., Gruetzner, R., Werner, S. & Marillonnet, S. A modular cloning system for standardized assembly of multigene constructs. PLoS One 6, el6765, doi: 10.1371 / journal. pone.0016765 (2011).36 O'Gorman, S., Fox, D. T. & Wahl, G. M. Recombina site-specific integration in mammalian cells. Science 251, 13 doi : 10.1126 / science.1900642 (1991).37 Duportet, X. et al. A platform for rapid prototyping of synthetic gene networks in mammalian cells. Nucleic Acids Res 42, 13440-13451, doi:10.1093 / nar / gkul082 (2014).38 Khalil, A. S. et al. A synthetic biology framework for programming eukaryotic transcription functions. Cell 150, 647-658, doi : 10.1016 / j . cell.2012.05.045 (2012).39 Maeder, M. L., Thibodeau-Beganny, S., Sander, J. D., Voytas, D. F. & Joung, J. K. Oligomerized pool engineering (OPEN): an 'open-source' protocol for making customized zinc-finger arrays. Nat Protoc 4, 1471-1501, doi: 10.1038 / nprot.2009.98 (2009).40 Li, H. S. et al. Multidimensional control of therapeutic human cell function with synthetic gene circuits. Science 378, 1227-1234, doi: 10.1126 / science.ade0156 (2022).41 Feil, R., Wagner, J., Metzger, D. & Chambon, P. Regulation of Cre recombinase activity by mutated estrogen receptor ligand-binding domains. Biochem Biophys Res Commun 237, 752-757, doi: 10.1006 / bbrc.1997.7124 (1997).42 Bashor, C. J. et al. Complex signal processing in synthetic gene circuits using cooperative regulatory assemblies. Science 364, 593-597, doi: 10.1126 / science.aau8287 (2019).43 Donahue, P. S. et al. The COMET toolkit for composing customizable genetic programs in mammalian cells. Nat Commun 11, 779, doi: 10.1038 / s41467-019-14147-5 (2020).44 Muldoon, J. J. et al. Model -guided design of mammalian genetic programs. Sci Adv 7, doi: 10.1126 / sciadv.abe9375 (2021).45 Kabadi, A. M. & Gersbach, C. A. Engineering synthetic TALE and CRISPR / Cas9 transcription factors for regulating gene expression. Methods 69, 188-197, doi: 10.1016 / j.ymeth.2014.06.014 (2014).46 La Russa, M. F. & Qi, L. S. The New State of the Art: Cas9 for Gene Activation and Repression. MolCellBiol 35, 3800-3809, doi:10.1128 / MCB.00512-15 (2015).47 Sadowski, I., Ma, J., Triezenberg, S. & Ptashne, M. GAL4-VP16 is an unusually potent transcriptional activator. Nature 335, 563-564, doi: 10.1038 / 335563a0 (1988).48 Shin, Y. et al. Spatiotemporal Control of Intracellular Phase Transitions Using Light- Activated optoDroplets. Cell 168, 159-171 el l4, doi: 10.1016 / j .cell.2016.11.054 (2017).49 Schneider, N. et al. Liquid-liquid phase separation of light-inducible transcription factors increases transcription activation in mammalian cells and mice. Sci Adv 7, doi: 10.1126 / sciadv.abd3568 (2021).50 Boeynaems, S. et al. Phase Separation in Biology and Disease; Current Perspectives and Open Questions. J Mol Biol 435, 167971, doi: 10.1016 / j.jmb.2023.167971 (2023).51 Gossen, M. & Bujard, H. Tight control of gene expression in mammalian cells by tetracycline-responsive promoters. Proc Natl Acad Sci U S A 89, 5547-5551, doi: 10.1073 / pnas.89.12.5547 (1992).52 Hansen, J. et al. Transplantation of prokaryotic two-c into mammalian cells. Proc Natl Acad Sci USA 111, 15705- doi : 10.1073 / pnas.1406482111 (2014).53 Ede, C., Chen, X., Lin, M. Y. & Chen, Y. Y. Quantitative Analyses of Core Promoters Enable Precise Engineering of Regulated Gene Expression in Mammalian Cells. ACS Synth Biol 5, 395-404, doi:10.1021 / acssynbio.5b00266 (2016).54 Chen, H., Covert, I. C., Lundberg, S. M. & Lee, S.-I. Algorithms to estimate Shapley value feature attributions. Nature Machine Intelligence 5, 590-601, doi: 10.1038 / s42256-023- 00657-x (2023).55 Mclnnes, L., Healy, J. & Melville, J. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv: 1802.03426 (2018).<https: / / ui.adsabs.harvard.edu / abs / 2018arXivl80203426M>.56 Tycko, J. et al. Development of compact transcriptional effectors using high- throughput measurements in diverse contexts. Nat Biolechnol. doi: 10.1038 / s41587-024- 02442-6 (2024).57 Tague, E. P., Dotson, H. L., Tunney, S. N., Sloas, D. C. & Ngo, J. T. Chemogenetic control of gene expression and cell signaling with antiviral drugs. Nat Methods 15, 519-522, doi : 10.1038 / s41592-018-0042-y (2018).58 Wimmer, E., Mueller, S., Tumpey, T. M. & Taubenberger, J. K. Synthetic viruses: a new opportunity to understand and prevent viral disease. Nat Biotechnol 27, 1163-1172, doi : 10.1038 / nbt.1593 (2009).59 Brophy, J. A. & Voigt, C. A. Principles of genetic circuit design. Nat Methods 11, 508-520, doi : 10.1038 / nmeth.2926 (2014).60 Whyte, W. A. et al. Master transcription factors and mediator establish superenhancers at key cell identity genes. Cell 153, 307-319, doi : 10.1016 / j . cell.2013.03.035 (2013).61 Pinglay, S. et al. Synthetic regulatory reconstitution reveals principles of mammalian Hox cluster regulation. Science 377, eabk2820, doi: 10.1126 / science.abk2820 (2022).62 Voigt, C. A. Synthetic biology 2020-2030: six commercially-available products that are changing our world. Nat Commun 11, 6379, doi: 10.1038 / s41467-020-20122-2 (2020).63 Valeri, J. A. et al. Sequence-to-function deep learning frameworks for engineered riboregulators. Nat Commun 11, 5058, doi: 10.1038 / s41467-020-18676-2 (2020).64 Hollerer, S. et al. Large-scale DNA-based phenotypic recording and deep learning enable highly accurate sequence-function mapping. Nat Commun 11, 3551, doi : 10.1038 / s41467-020-17222-4 (2020).65 LaFleur, T. L., Hossain, A. & Salis, H. M. Automated model-predictive design of synthetic promoters to control transcriptional profiles in bacteria. Nat Commun 13, 5159, doi : 10.1038 / s41467-022-32829-5 (2022).66 Rai, K., Wang, Y., O’Connell, R. W., Patel, A. B. & Bashor, C. J. Using Machine Learning to Enhance and Accelerate Synthetic Biology. Current Opinion in Biomedical Engineering (2024).67 Karst, S. M. et al. High-accuracy long-read amplicon molecular identifiers with Nanopore or PacBio sequencing. 1 doi : 10.1038 / s41592-020-01041 -y (2021).68 Chung, C. T., Niemela, S. L. & Miller, R. H. One-step preparation of competent Escherichia coli: transformation and storage of bacterial cells in the same solution. Proc Natl AcadSci USA 86, 2172-2175, doi: 10.1073 / pnas.86.7.2172 (1989).69 Parrish, J. R. et al. High-throughput cloning of Campylobacter jejuni ORfs by in vivo recombination in Escherichia coli. J Proteome Res 3, 582-586, doi: 10.1021 / pr0341134 (2004).70 Currin, A. et al. Highly multiplexed, fast and accurate nanopore sequencing for verification of synthetic DNA constructs and sequence libraries. Synth Biol (Oxf) 4, ysz025, doi : 10.1093 / synbio / y sz025 (2019).71 De Coster, W., Weissensteiner, M. H. & Sedlazeck, F. J. Towards population-scale long-read sequencing. Nat Rev Genet 22, 572-587, doi:10.1038 / s41576-021-00367-3 (2021).72 Smith, T. F. & Waterman, M. S. Identification of common molecular subsequences. J Mol Biol 147, 195-197, doi:10.1016 / 0022-2836(81)90087-5 (1981).73 Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823, doi: 10.1126 / science, 1231143 (2013).74 Hermann, M. et al. Binary recombinase systems for high-resolution conditional mutagenesis. Nucleic Acids Res 42, 3894-3907, doi:10.1093 / nar / gktl361 (2014).EXAMPLE 3
[0246] CLASSIC: A platform for high throughput genetic library characterization at arbitrary length scales to map genetic design spaces.
[0247] In this study, our objective was to develop a highly accessible approach for ultra-high throughput characterization of complex libraries of genetic constructs of multi-kb length scales in mammalian cells. An extended overview of the approach that we developed, which we have named CLASSIC (combining long- and short-read sequencing for investigating genetic complexity), is depicted in FIG. 14. The CLASSIC platform uses a custom pooled library construction approach in which genetic parts are assembled in combinatorial fashion to yield pools of diversified sequence compositions from a multi-dimensional user- defined design space. Assembly involves a custom Type Ils cloning scheme, known as Sesame Street, which enables successive hierarchical pooled plasmid assemblies to create libraries of >105compositions. During the assembly process, each construct in the library receives at least one unique DNA barcode sequence that consists of a combination of variable bases and constant “tether” bases. This barcode structure facilitates faithful sequenceidentification by long-read nanopore sequencing, which is us barcode sequences to their associated construct compositions via a custom sequence analysis software package, WIMPY. Barcoded circuit libraries are introduced into cells at single copy via integration into a HEK293T cell line containing a genomic “landing pad” (LP) sequence. Following expansion, the cells containing members of the library are physically separated into subsets based on their phenotype; in this study, we use flow cytometry to sort cells into bins based on expression profile of a fluorescent protein reporter. The composition- associated barcodes expressed by each cell are assessed by short-read sequencing and analyzed by a second software package to determine expression based on barcode bin distribution. These data are then compared to the barcode-to-composition index to map expression profiles to each construct, thereby enabling large-scale, quantitative phenotypic mapping of a genetic design space.
[0248] In this study we focused on developing and validating the CLASSIC platform as a way to accelerate synthetic biology and overcome barriers to engineering synthetic gene circuits in mammalian cells. We demonstrated that CLASSIC is capable of identifying high-performing circuit compositions from amongst many potential designs, characterizing context-dependent or non-linear interactions between genetic parts, and developing data-driven models that can be used to generate comprehensive maps of circuit design spaces (FIG. 14, bottom). While addressing synthetic circuit design challenges offers an excellent showcase for the capabilities of our platform, we believe that CLASSIC is applicable to a variety of different natural contexts in which function arises from complex and context-specific regulatory relationships between genetic elements. This includes diverse transcriptional regulatory elements arrayed across a gene locus or chromatin neighborhood1, disparate open reading frames comprising a biosynthetic bacterial operon2, or multi-transgene programs introduced into cells for therapeutic applications3. Below we provide an in-depth description of CLASSIC, first detailing proof-of-concept work that allowed us to develop the platform, and then describing how we used it to conduct experiments detailed in FIGs. 2-4.
[0249] A modular, hierarchical cloning workflow that scales to high-throughput plasmid library construction.
[0250] Gene circuit engineering proceeds through successive design-build-test-learn (DBTL) cycles, a process where initial designs are assayed, used to parameterize a quantitative model, and then tuned to ultimately converge on a target behavior4. Because iteration through multiple DBTL cycles is a fundamental part of any synthetic biology project, the ability to efficiently construct DNA molecules of arbitrary sequence composition usingmolecular cloning approaches is of central importance. In recType Ils enzymes4that cut DNA at a site outside of their recognition sequence nas given use io streamlined cloning workflows for robustly combining multiple genetic part plasmids into a destination vector without the need for purification of linearized fragments5. Widespread adoption of these approaches has led to development of a number of cloning “toolkit” platforms, where standardized sets of genetic parts and vectors are used to compose systems that allow tiered construction of genetic circuits through successive, hierarchical assembly steps, resulting in rapid, robust prototyping of large, complex synthetic constructs. Such systems have been developed in yeast6, plants7, and mammalian cells8.
[0251] To facilitate systematic exploration of the genetic design spaces envisioned for CLASSIC, we developed a Type Ils toolkit for mammalian cells that is optimized for assembly of constructs in both individual and pooled formats. Our system, which we call Sesame Street (FIG. 15), has the following features that are enabling for CLASSIC: i) A modular, four-level hierarchy that allows the introduction of diversity at any scale of transgene or gene circuit organization, including the protein domain, expression unit (EU), or multi-gene expression array levels. By balancing complexity across assembly levels, this scheme maximizes efficiency while minimizing incomplete assemblies; ii) overhang sequences that are optimized for pooled assemblies that robustly yield stoichiometrically balanced plasmid libraries; iii) Vectors harboring selection markers that streamline the library cloning process, such as the bacterial toxin ccdB9and the color-selectable Crt operon10, which are excised during the assembly step to limit the occurrence of false positives during library construction.
[0252] Plasmid construction using Sesame Street begins with chemical synthesis or PCR-amplification of DNA fragment “sub-parts” that are assembled with a custom backbone fragment containing a kanamycin resistance cassette (KanR) using Bpil and transformed into E. coll and prepped to yield sub-part “Baby Bear” plasmids (e.g., protein domains, or promoter and enhancer fragments; level 0) (FIG. 15, top). The overhangs used to generate these plasmids are customizable to facilitate full sequence control over the selection of boundaries between subparts (e.g., protein domains) within a genetic part (e.g., a multi-domain protein). The resulting sub-part plasmids are then used as inputs for assembly into ccdB-containing ampicillin-resistant (AmpR) level 1 “Bert” entry vectors with Bsal. Each category of part resulting from this assembly step (i.e., promoter, ORF, or terminator) has a pre-defined set of standard overhang sequences that specify their ordered assembly to form EUs in subsequent steps, as well as modular exchange of genetic parts within the same category. Transformation and prepping theseassemblies yield level 1 genetic part plasmids. EUs are then coORF, and terminator part plasmids into a KanR level 2 “Ernie entry vector containing a uuai ccdB-mCherry2 expression cassette using Esp3I. This vector contains standard overhang sequences that accept any correctly constructed promoter, ORF, and terminator combination. Transformation subsequently yields level 2 EU plasmids. Level 0 through Level 2 plasmids can all be linearized using SapI for sequence validation by nanopore sequencing. Validated EUs can then be combined into a final “destination” “Big Bird” vector containing an ampicillin- or spectinomycin-resi stance (SpectR) marker. This occurs via BpiLmediated assembly to replace a large marker sequence encoding the Crt operon, resulting in multi-unit arrays where genes are ordered based their level 2 entry vector overhangs. Transformation and DNA prepping yields completed level 3 multi-unit regulatory programs. Level 3 vectors contain a unique Esp3I restriction site to allow sequence-level validation via nanopore sequencing.
[0253] Sesame Street provides a parallelizable end-to-end cloning platform for robustly constructing synthetic regulatory programs of arbitrary size (<20 kb) and complexity, with a standard workflow requiring approximately 10 days to go from level 0 sub-part fragments to verified level 3 plasmids (FIG. 16). Additionally, our platform enables construction of large plasmid libraries by pooled assembly of genetic species. Diversity can be introduced at any of the construction levels and successfully carried through multiple assembly steps to generate stoichiometrically balanced plasmid pools (see results below). For example, both promoters (FIG. 17, panel A) and ORFs (FIG. 17, panel B) can be diversified at the level 0 Baby Bear assembly level and propagated to the level 2 Ernie assembly level to generate diversified EU libraries. Alternatively, diversity at the EU level can be created with diverse inputs from the level 1 Bert stage, and then used for assembly of diverse gene circuit topologies at the level 3 Big Bird stage. Furthermore, a variety of different classes of vectors can be used in this level 3 destination context depending on the desired application, such as transient expression or genomic integration with CRISPR, Landing Pad, PiggyBac, or lentivirus.
[0254] S1.3 CLASSIC barcoding strategy.
[0255] CLASSIC uses short barcode sequences to map expression phenotype to specific genetic compositions. To introduce barcodes during gene circuit library assembly, we devised a two-part barcoding scheme that is implemented during EU assembly into level 3 destination vectors (see FIG. 15). The first part (BC1, 18 bp) is located within the 3’ UTR of a BFP EU that assembles with other EUs in the array, while the second part (BC2, 18 bp) is contained in the destination vector itself (FIG. 1, panel A; FIG. 18, panel A). Since the BFP EU is always placedin the final position of the multi -EU array, BC1 and BC2 assei contiguous 2-part barcode that is contained in the BFP 3 ’ UTR anu can ue ainpiineu iroin reverse transcribed mRNA for quantitative assessment by short-read NGS11. To facilitate the accurate identification of barcode sequences from individual nanopore reads, we devised a structured barcoding scheme in which a repeated combination of variable and constant bases (6 BBA repeats for BC1, 6 DDC repeats for BC2) (FIG. 18, panel A) provides anchor points; a strategy we hypothesized would enable robust alignment and improve the identification of miscalled bases. Additionally, this structure minimizes incidences of problematic sequences, such as homopolymers or enzyme recognition sites12.
[0256] The combination of two distinct barcode pools, each with up to ~5»105members, yields a potential diversity of almost 3»10nbarcode sequences. To maximize incorporation of this diversity into our working barcode pool, we explored two approaches for cloning the barcode pools into their respective vectors (FIG. 18, panel B). In one approach, we amplified vectors by PCR with a primer encoding the barcode pool followed by digestion and auto ligation, and in the second we used golden-gate assembly to replace a ccdB cassette in the vectors with the barcode pool (see Methods). For the golden gate assembly strategy, we observed greater assembly efficiency as determined by colony counts (data not shown), lower rates of incomplete assembly as determined by the ratio of colony counts on a vector-only control plate (data not shown), and greater barcode diversity as determined by short-read NGS (FIG. 18, panel B, bottom). Therefore, this method was used to generate all barcode pools used in this study. To further maximize the diversity of our barcode pools, we tested two transformation methods in DH5aE. coli (FIG. 18, panel C). In one strategy we transferred the transformed cells directly into LBKan liquid media, grew overnight (750 rpm, 37 °C), and then mini prepped the cell pellet. In a second strategy we plated the transformation onto LBxan-agar plates, grew overnight (37 °C), and then we scraped the plate followed by a 45 min outgrowth in LB Kan (750 rpm, 37 °C), and miniprepped the cell mass. Following short-read NGS of the barcode regions cloned using each strategy, we observed almost twice as many unique barcodes (6,278 vs 3,604) and a more balanced representation of barcode sequences in the plated-and-scraped sample (FIG. 18, panel C). Accordingly, this method was used to prepare all plasmid pools used in this study.
[0257] With the ability to create deep barcode pools in-hand, we sought to demonstrate that our structured barcoding strategy can improve barcode identification. We compared our ability to accurately identify either structured or random barcode sequences from nanopore sequencing data. We cloned either a semi -degenerate oligo containing 6 repeats of the nucleotidebase pattern BBA, or a random 10-base-pair barcode sequel plasmid (FIG. 18, panel D, left). We scraped approximately ~ ,vvv coiomes 101 eacn ua coue set and then assembled them into a level 3 destination vector along with a constant mRuby EU driven by either RSV (BBA) or hEFlal (ION) promoters (FIG. 18, panel D, left). -500 colonies were then scraped from each set, miniprepped, and prepared for nanopore sequencing by digestion with Esp3I (see Methods). The barcode region from the two pools was also amplified and sequenced using Illumina MiSeq (see Methods) to serve as a ground truth for comparing with barcode sequences obtained from nanopore sequencing analysis. The nanopore sequencing analysis proceeded by using Viscount and Chophat software modules (description in section S1.4) to separate the pools based on the RSV and hEFlal promoter. The barcodes were extracted from these pools using the software module Barcoat (description in section SI.4) for the BBA pool or using 5 bp constant upstream and downstream sequences for the ION pool. The assignment of barcodes from these sequences followed a 3 -step protocol: i) search for an exact match for nanopore barcode in the Illumina ground truth table. If found, assign to the “match” array (FIG. 18, panel D dark gray bars); ii) for the BBA barcodes specifically, the structure was corrected and searched again with mutations in constant regions fixed to the constant base, while non-constant mutations were converted to all possible bases (T, G or C) to generate a set of possible barcodes to search for. Any indels in these regions were also corrected so that every possible barcode generated using the combination of these corrections was placed in the set. The entire set of possible variations of this barcode was then searched in the truth table for matches. If only one match was found, the barcode was added to the “error corrected” array (FIG. 18, panel D medium gray bars), while barcodes with more than two matches were discarded; iii) For both BBA and ION barcodes, the reads were searched in the truth table allowing for one mutation or indel. To achieve this, a set of sequences was generated to contain every possible single bp mutation in each barcode and searched in the Illumina truth table. Reads containing barcodes that could not be identified after step 3 were added to an “error retained” array (FIG. 18, panel D light gray bars). Across the three available nanopore basecalling modes — fast, high accuracy, and super high accuracy settings in Guppy, the basecaller software from ONT — we observed that the structured BBA barcode sequence allowed recovery of a greater proportion of the originally unidentified barcodes and led to an increase in barcode to circuit assignment, compared to the random ION barcode (FIG. 18, panel D, right). Our data also demonstrate that barcode identification for BBA structured barcodes is more robust to lower accuracy basecalling modes.
[0258] To verify that barcode structure is maintained t we used Sanger sequencing to assess the base proportions at eacn posiuon or me muiviuuai (rio. 18, panel E, left) and combined (FIG. 18, panel E, middle) BC1 and BC2 plasmid pools. We found constant bases (“A” for BC1 and “C” for BC2) exclusively in the tethering positions and approximately equal proportions of variable bases at their intended positions. Additionally, high- depth Illumina sequencing of the BC1 pool demonstrated that our cloning method allows us to generate large pools that capture >99.9% of the possible barcode diversity (FIG. 18, panel E, right)
[0259] CLASSIC computational analysis pipelines for characterizing genetic composition and construct behavior.
[0260] To analyze the sequencing data gathered using CLASSIC, we developed a pair of computational pipelines: one for analyzing long-read nanopore sequencing to create a composition-to-barcode index, and the other to analyze short-read sequencing data to establish a phenotype-to-barcode index (FIG. 19). Advances in nanopore sequencing technology have given rise to its use in a variety of applications, including multiplexed sequencing of whole bacterial genomes and metagenomes13’14and sequencing of low-diversity pools of barcoded plasmids15. Most of these applications rely on the alignment of dozens to hundreds of reads of a single sequence composition to generate an accurate consensus sequence. Since CLASSIC involves obtaining nanopore data for large pools of barcoded plasmids at relatively low-read depth (<10 reads per unique composition), most of which potentially contain unique composition / barcode combinations, it required us to develop new analysis software. The package we created, which we named What’s in My Pot, Y’all (WIMPY), utilizes a combination of localized containment searches to identify circuit compositions and sequence alignment to assign a barcode to each read. By tiling small sub-sections of each read against a short list of potential genetic parts, WIMPY quickly and accurately identifies the genetic composition of each read, even in the presence of basecalling errors. In parallel, WIMPY aligns the barcode sub-section of each read to a semi -degenerate reference barcode sequence, leveraging the barcode structure to correctly place miscalled bases (described above). These functions allow WIMPY to construct highly accurate barcode-to-composition maps on a single-read basis, enabling high throughput (HT) library assessment at much lower read depths than is required by consensus sequence generationbased tools (FIG. 19, bottom). WIMPY includes the following modules:
[0261] Fastqall - Combines reads from all fastq files in the directory, which are by default split into 4000-read blocks by Guppy, into a single cell array while discarding sequenceheaders and information about per-base quality scores. The < array containing all the reads, where n is the read depth obtaineu iroin me naiiopore run.
[0262] Bowtiles - Indexes the reads and tethers them to a common reference point in the plasmid library, such as the C-terminus of the puroR cassette, which should be contained in all level- 3 plasmid assemblies. Reads in which puroR is not detected are discarded at this stage, while on-target reads are reconstructed so the reference sequence is located on the 5’ end of the top strand (reads from the opposite orientation are reverse complemented). The output of Bowtiles is a n»l cell array containing all reads indexed to PuroR.
[0263] Tilepin - Identifies library-specific constant landmarks in each read using a containment search (see Methods) and indexes the locations of these landmarks for later use. Examples of invariant genetic landmarks include fluorescent reporters, such as mRuby (384- member library, FIG. 2, panel A), the Ert2 domain (166k member library, FIG. 3, panel A), or the BFP fluorescent protein in the barcoded EU used in both libraries. The inputs to Tilepin are the cell array containing all the reads, and DNA sequences representing identifiable landmarks (i.e., mRuby and BFP). The output is an n»2 integer array where n is the read depth on the library, and the columns contain the position of mRuby and BFP respectively.
[0264] Chophat - Truncates reads based on landmark coordinates identified by Tilepin, yielding a n»[m + 1] cell array, where n is the total number of reads and m is the number of landmarks (e.g., mRuby, Ert2 domain, BFP etc.). The inputs to Chophat are the n»l cell array containing all nanopore reads (Bowtiles output) and the n»m integer array specifying landmark indices within each read (Tilepin output). The output is an n»[m + 1] cell array containing each nanopore read (n) sub-divided into sections that correspond to the distance from one landmark (m) to the next (m+1). In the case of all libraries used in this study, the final column of the Chophat output array for each read contains the region downstream of the BFP expression unit, which runs from genetic landmark m (BFP) to m+1 (end of the read). The outputs from Chophat are sent in parallel to Viscount and Barcoat for identity assignment and barcode sequence determination, respectively.
[0265] Viscount - uses a fast, highly specific containment search method to query a defined sub- section of each read against a shortlist of genetic part sequences to make part identity assignments. The inputs to Viscount are columns 1 through m of the n»[m + 1] cell array output from Chophat (i.e., all sequence sub-sections other than the region downstream of BFP) and a reference cell array containing the sequences of all possible genetic parts within a category (e.g., all 8 promoters used in the 384-member library) (FIG. 2, panel A). Thecontainment search works as follows: i) each reference gene in a user- defined manner. For this study we used tile lengtns or o up wun a i up smue 101 Kozak sequences, and tile lengths of 10 bp with a 2 bp stride for all other parts; ii) All reads are queried for all tiles for every reference sequence using a binary containment search (Matlab function contains),' iii) The number of tiles contained in each read is counted for all reference sequences; iv) The number of tiles assigned to each read is normalized by the number of reference tiles to account for variable reference sequence lengths, yielding the fraction of tiles contained in each read for each reference sequence; v) If a read contains more tiles than the specified threshold (user defined, typically -3-5%) for a given reference sequence, it is assigned a part identity. Reads can receive more than one part identity by meeting the threshold for multiple parts within the category; such reads are referred to as “confused” reads and are potentially a product of noisy sequencing data. The outputs from this process are part assignments for each read found to unambiguously contain just one part from the set, and a confusion matrix showing the number of reads that were assigned to one part (diagonal of the matrix) and two parts (off-diagonal elements). Within the part assignment output, reads in which a part could not be identified are assigned “0”, while confused reads are assigned “-1”. All 0 and -1 reads identified by this analysis are subsequently discarded.
[0266] Barcoat - leverages the semi-degenerate BBA / DDC structure in our barcoding scheme to identify barcodes using a custom alignment matrix with a 0 penalty for on target alignments (“B to B” or “D to D”), and -5 penalty for unexpected alignments (“A to B” or “C to D”) in the barcode region. The Barcoat input is column [m + 1] from the Chophat output, which contains the barcode region downstream of the BFP coding sequence. Sequences with a perfect alignment score and length (18 bp) are saved directly. Sequences with 1 mismatch / indel in the constant tethered bases (A for BC1, C for BC2) are corrected by adding back the constant tethered base and then saved. Sequences with a mutation / indel in the variable region are queried against the short-read sequencing data obtained from flow-seq experiments (see below). If a match is found in the region containing the mutation / indel, the sequence is corrected based on the illumina reference. Sequences with two or more mismatches / indels are discarded.
[0267] The output of Barcoat enables us to create an indexed library with composition- to-barcode look-up tables. This table also facilitates the identification of barcodes that are associated with multiple compositions during short-read sequencing analysis.
[0268] The phenotype-to-barcode mapping pipeline we developed to analyze CLASSIC flow-seq data follows a standard Illumina analysis workflow16(FIG. 19, right). Illuminalibraries are generated using a 2-step PCR-based protocol1illumina adapters as well as bin / sample-specific barcodes (4 nij io oom me 101 wa u anu reverse primers. The second PCR step uses universal primers that prime on the partial adapters to complete the adapter sequences, generating a ready-to-sequence library. Sequencing data analysis proceeds following a standard flow-seq data processing pipeline, with a few key differences. Briefly: i) For all reads, the sample / bin identities added during the PCR1 step are decoded and the sequences are extracted from read 1. The barcode and bin IDs are saved for each read at this step; ii) After determining the barcodes contained in each read, the counts on each unique barcode in each bin are determined using the Matlab functions unique and accumarray, iii) Counts for each barcode are normalized by read depth and scaled to a constant number. The outputs from this step are normalized read counts for all barcodes; iv) Uneven bin-to-bin sampling of library-containing cells during sorting are corrected by multiplying the relative fraction of sorted events from each bin by the normalized barcode read counts measured in that bin; v) A weighted average (depth- and bin-normalized read count • bin #) over the resulting read count distribution is computed to generate an “average bin” for each barcode; vi) The average bin is converted to average expression using an experimentally generated standard curve that relates expression level to bin number. This process yields an average expression mapping for every barcode sequence identified in the flow-seq pipeline. Using the barcode-to-composition map generated during the nanopore sequencing analysis, barcodes linked to the same composition can then be combined to yield an aggregate expression distribution for each circuit composition. A single expression value is then assigned by computing a kernel density estimate (KDE) on the loglO(expression) of the distribution and taking the maximum probability.
[0269] Developing a HEK293T landing pad cell line for single-copy integration of CLASSIC libraries.
[0270] In order to quantitatively compare the behavior of different circuit designs across a library generated with CLASSIC, it was necessary to develop a method for efficiently introducing constructs into human cells at single copy. While there are widely used genetic parts capable of robustly maintaining a defined copy number in microorganisms like E. coli and S. cerevisiae6, establishing copy number control for quantitatively accurate measurements in high-throughput screening studies in mammalian cells is challenging. CRISPR has emerged as a versatile tool for editing mammalian genomes19, however the inefficiency of homology- directed repair for integrating larger length-scale constructs and the associated genotypicvariability make it difficult to generate and quantitatively con using CRISPR integration. Other methods for stable genomic maiiiieiiaiice or uaiisgeiies, sucn as transposon- or retro / lentivirus-based methods, integrate at random locations with uncontrollable frequency. In recent years, the use of LPs has emerged to overcome this barrier20,21. A LP cell line is established by locus-specific CRISPR-integration of a short DNA sequence containing selection markers and bacteriophage recombinase sites, followed by clonal sorting and genotypic validation for single-copy integration. Large DNA sequences (>10kb) housed in a donor vector can then be co-transfected with a recombinase, catalyzing a recombination event that results in donor integration at the LP locus, usually accompanied by an exchange of positive or negative markers that enables selection of successful Integrants. LP cell lines are ideal for testing libraries in mammalian cells since transfection of pooled DNA constructs can be expressed across a population of cells at single-copy, all from the same genomic context ’ .
[0271] To establish a high-efficiency LP cell line for use in screening CL AS SIC-generated libraries, we integrated a custom LP sequence into the AAVS1 locus in HEK293T cells, a human cell line used for bioproduction and mammalian synthetic biology22(FIG. 20, panel A). Our LP construct contained AAVS1 flanking homology sequences, an attP site for BxBl recombinase21, a YFP-hygromycin expression cassette under control of the hEFlal promoter, and a constitutively expressed TKHSV marker for negative selection of secondary integrants (not used in this study). Selection of the LP- integrated population with hygromycin was followed by clonal sorting into a 96-well plate based on YFP fluorescence. Clones were expanded and validated by PCR genotyping to demonstrate single-copy CRISPR integration of the LP, and also tested for integration capacity using a destination vector containing a constitutively expressed BFP unit (pROC.A585) (FIG. 20, panels B, C). To test integration capacity, clonal LP cell lines (YFP+, BFP-) were plated in 24-well format and co-transfected with the destination vector and a BxBl expression plasmid. After 2 days, cells were passaged 1 :5 into a 24-well plate containing 1 pg / mL puromycin and grown for one passage (8 days) under puromycin selection, yielding a homogeneous population of BFP-expressing cells (YFP- , BFP+). In parallel, transfected cells were passaged 1 : 10 into DMEM and grown for multiple passages (8 days) to assess the efficiency of genomic integration. After 10 days in culture, we observed a small population of BFP-expressing cells, which represents the frequency of BxBl- mediated integration in the absence of puromycin selection. One cell line (out of ~20 tested)that showed ~5% integration efficiency which we refer to a the study to integrate CLASSIC libraries.
[0272] To assess the degree to which integration into our HEK293T-LP cell line yields phenotypically homogeneous cell populations, we integrated a constitutively expressed mCherry cassette (pROC. A0003), grew cells to homogeneity under puromycin selection (FIG.20, panel D), and sorted and expanded 23 clones. Flow cytometry was then used to measure mCherry fluorescence for each of the clonally-expanded populations. Values for mean expression of clonal populations (mChn) and the bulk population (mChbulk) were used in the following calculation:
[0273] The resulting MAE of 0.08 provides a quantification of expression variability between genetically identical integrants and can be used to estimate the typical range of heterogeneity expected across HEK-LP cell populations harboring integrated expression cassettes. This value, which we termed the Error Range from Clonal Heterogeneity (ERCH), is used later in the study to assess the agreement between expression values measured using CLASSIC and those measured clonal isolates (FIG. 2, panel C; FIG. 3, panel D) and constructed cell lines (FIG. 4, panel B).
[0274] Phenotypic selection of CLASSIC libraries using flow-seq.
[0275] We used flow-seq to phenotypically separate cells containing integrated library composition based on their eGFP reporter fluorescence expression profiles. To establish a CLASSIC-compatible flow-seq strategy we constructed a 24-member mCherry EU test library composed of 4 variable strength mammalian promoters (CMV, hEFlal, RSV, and hPGK) and 6 standard terminators (T1-T6) (FIG. 21, panel A). This library was assembled and barcoded (FIG. 18) such that the majority of barcodes were assigned to a single EU composition and then integrated into HEK-LP cells and expanded under puromycin selection (FIG. 20; FIG.21, panel B, top). In order to determine how cells containing LP -integrated libraries should be handled during flow-seq, we tested pre-sort washing and a 5 day post-sort outgrowth as two ways to potentially minimize mRNA cross-contamination. After 10 days of selection, library- integrated cells were lifted using TrypLE and split evenly into two groups before sorting; one group was washed twice with PBS and the other was left unwashed. 20k cells were then sorted from bins 3 and 8 for both washed and unwashed cell populations. The sorted cells from eachgroup were again split evenly into two groups: one was use and NGS, while the other was plated in 96-well format and expanu 101 J uays ^passageu once; before sequencing. We saw significant barcode overlap when total RNA was extracted immediately following the sort (52.09% vs 9.03%), which is likely the result of non-specific contaminant RNA from cells lysed and errantly sorted with cells from bins 3 and 8 (FIG. 21, panel C). We observed a significant decrease in overlap in post-sort expanded cells, likely because non-specific mRNA degrades during cell outgrowth. We also observed a decrease (9.03% vs 4.53%) in barcode overlap when cells were washed using PBS before sorting (FIG. 21, panel C). We thus settled on a protocol that includes these pre-sort wash and post-sort expansion steps for all flow-seq experiments reported in this study (FIG. 21, panels B, C).
[0276] Assessing genetic part interactions across a 384-member EU library.
[0277] In this study, we prioritized building constructs from genetic parts that are frequently used in mammalian synthetic biology to tune gene expression. Despite their common use, many of these parts have not been rigorously benchmarked, and the context-dependent interactions that exist between them have yet to be systematically studied. This makes it challenging to use them to precisely program gene circuit behavior, or subsequently re-use them across different circuit designs. In this section, we demonstrate that CLASSIC can be used to address this challenge: we constructed a 384-member mRuby EU library to: 1) characterize the expression profiles of genetic parts commonly used in mammalian synthetic biology; 2) systematically investigate the context-specific, non-linear interactions between genetic parts, and; 3) develop a machine learning (ML) model for predicting part interactions.
[0278] 384-member EU library construction.
[0279] Construction of the 384-member mRuby EU library followed the standard Sesame Street library construction scheme detailed herein, with minor adjustments. We assembled 8 promoters, 6 Kozaks, and 8 terminator sequences into level 1 entry vectors to yield 22 distinct plasmids. Since the diversity we tested was not divisible into sub-parts, we were able to bypass level 0 plasmids in the assembly of this library. Level 1 plasmids for each part category were pooled in equimolar amounts and assembled along with a level 1 plasmid encoding an mRuby ORF into a level 2 destination vector (FIG. 22, panels A, B). Transformation into 100 pL of chemically competent DH5a E. coli yielded approximately 1.5»104colonies, which were scraped and miniprepped to yield a level 2 EU library. Here, we aimed for a number of colonies that would ensure >30x coverage of the design space. The level 2 mRuby EU pool (384- members) was then assembled with a barcoded BFP EU pool (~5.31* 105members) into a non-barcoded level 3 destination vector. The assembly reaction chemically competent Stbl3 E. coli to yield approximately I .J- IU coiomes (rnj. , panel B). In these final assembly and transformation steps, the number of colonies targeted is a balance of three important factors: design space coverage, uniqueness of barcoding, and feasibility of nanopore sequencing. While high colony counts increase the likelihood of covering of the design space and also buffer against part imbalance, the probability of unique barcode-to-circuit mapping decreases as the number of colonies harvested approaches the size of the barcode pool. Additionally, larger pools of unique barcode / circuit combinations require greater nanopore sequencing depth to identify and index all genetic variants contained in the prepped DNA sample. For our 384-member library, we targeted between 10- and 100-fold more colonies than compositions, which we estimated would give complete coverage of the design space while also enabling efficient sequencing and indexing. We then co-transfected the level 3 barcoded mRuby EU library and a BxBl expression vector into 106HEK-LP cells, yielding and an estimated l»104- 5»104integration events (FIG. 22, panel B).
[0280] To assess the extent to which our modular cloning scheme permits unbiased assembly of genetic parts in a pooled reaction, we carried out nanopore sequencing on the level 3 barcoded mRuby EU library (FIG. 22, panel C). Analysis of the raw read lengths plot indicates that the plasmid pool is primarily composed of complete assembly products (86.9%), with a relatively low proportion of non-library species that could aberrantly integrate into HEK- LP, such as re-circularized empty (3.0%) and undigested (4.5%) destination vector. We then analyzed the sequencing data using the WIMPY pipeline to assign genetic compositions and barcodes to each read (FIG. 19).
[0281] Closer inspection of the relative stoichiometries of the genetic parts in the level 3 barcoded mRuby EU library revealed approximately 4-fold, 3 -fold, and 6-fold differences between the most and least frequently observed parts for the promoter, Kozak, and terminator categories, respectively (FIG. 5, panel A, bottom). Comparison with the level 2 mRuby EU library, which served as an input for the level 3 assembly, indicates a similar difference between the promoter (<3-fold), Kozak (<2-fold), and terminator (<2-fold) categories, suggesting that part balance is conserved between consecutive assembly cycles (FIG. 5, panel A, top). Critically, the degree of part balance appears to be well maintained for parts with disparate length scales (e.g., promoters ranging from 0.3 - 1.2 kb). Additionally, the diversity of our barcode pool appears to be sufficiently large for a library of this scale, with greater than 95% of barcodes uniquely mapping to a single EU composition (FIG. 5, panel B).
[0282] 384-member EU library flow sorting.
[0283] CL AS SIC-compatible flow-seq relies on extracting mioimauoii auoui genetic circuit behavior from distributions of barcoded transcripts in each sorting bin. Using our optimized flow-seq protocol (FIG. 21, panel C), we first gated cells containing our 384-member library based on size characteristics as well as BFP expression to exclude either empty cells or cells integrated with non-library constructs. We then sorted cells into 10 equally log-spaced bins based on mRuby expression. We expanded each sorted bin for 5 days, extracted total RNA which we converted to cDNA, amplified barcode regions with bin-specific and sample-specific primers, and sequenced the amplicon pool using Illumina (FIG. 23, panel A). The number of cells sorted from each bin was proportional to their relative abundance in the mRuby distribution (FIG. 23, panel B). Transcriptional coupling has been suggested to exist between adjacent chromosomally integrated EUs23. However, we observed low variance of mean BFP expression across increasing mRuby expression (CoV = 6%) suggests that the expression level of BFP barcode transcripts is decoupled from the mRuby expression level (FIG. 23, panel C)
[0284] Data processing and CLASSIC measurement validation.
[0285] Since a single genetic composition in our library can associate with multiple barcodes, we elected to combine data from multiple barcode measurements to create a more accurate representation of expression level for a composition, an approach we thought would be more robust to potential sources of noise in our data such as PCR bias, insufficient numbers of sorted cells per composition, low sequencing read depth, and sorting errors. In order to systematically assess the best approach for aggregating multiple barcode data into a single expression value, we tested three statistical methods: i) a weighted average (mean) across the barcodes; ii) picking the barcode representing the median expression value, and; iii) gaussian KDE on the raw or log transformed expression values (FIG. 6, panel A, top). Expression value measurements from our randomly sorted EU isolates (FIG. 2, panel C) served as a ground truth to assess the accuracy of each method (FIG. 6, panel A). While application of all three methods resulted in MAE values less than the ERCH, KDE demonstrated the lowest MAE (FIG. 6, panel A, bottom) and was therefore used for all downstream data analysis. As expected, we did not observe a significant difference in the MAE when using mean bin expression values gathered from the original sort gates or means values from the sorted and expanded populations (FIG. 6, panel A, bottom). Additionally, we observed low MAE for a replicate of the library (library 2) (FIG. 6, panel A). Expression distributions for each of the15 clonally isolated library members as measured by flow c;KDE-derived CLASSIC barcode distributions are shown in riu. o, panel r».
[0286] We demonstrated that the CLASSIC pipeline is largely robust to variability introduced during library construction, cellular integration, and cell sorting (FIG. 2, panel D). We hypothesize that this is partially the result of many-to-one barcode-to-composition correspondences, which reduces the effects of measurement noise by averaging across multiple measurements. We plotted the MAE between CL AS SIC-derived values and isolate measurements against the number of barcodes used to calculate expression (FIG. 6, panel C). This revealed that, as the number of barcodes increases from 1 to 10, the MAE decreases and appears to plateau at around 0.08, which we identified as the natural variability between clones. Notably, for individual barcodes, increasing read depth improves the mean absolute error up to a point (-500 reads) after which the MAE begins to increase again (FIG. 6, panel D). This is potentially the result of preferential bin-specific amplification of certain barcodes, leading to substantially higher reads depths relative to the rest of the distribution. Interestingly, the lowest MAE is achieved by aggregating multiple barcodes (FIG. 6, panel C) rather than filtering by read depth (FIG. 6, panel D), suggesting that aggregating multiple barcodes is a better strategy for estimating expression level. Additionally, expression of EU compositions from two separate sorts (FIG. 2, panel D, top) demonstrates substantially higher correlation (r2=0.99) than when comparing measurements on individual barcodes (r2=0.90) (FIG. 6, panel E). This suggests that the accuracy of expression level estimation for a given composition is significantly improved by aggregating measurements from multiple barcodes associated with that composition.
[0287] Extended 384-member library data analysis.
[0288] One longstanding goal for synthetic biology has been the creation of modular, quantitatively benchmarked parts sets that demonstrate predictable behavior across multiple circuit designs24. Development of such part sets often involves testing them in a limited number of circuit design contexts and is often accompanied by fitting with a biophysical or mechanistic model that parameterizes part function (e.g., relative transcriptional strengths for promoters, or translation initiation rates for Kozak sequences). While carefully benchmarked parts can, in principle, be used be used to compose new circuits, parts often behave differently when placed in juxtaposition with new or previously untested parts.
[0289] CL AS SIC has the potential to overcome part interoperability challenges by enabling gene circuit engineers to simultaneously gather data for part behavior across diverse part compositions. For example, besides showing that each of the three parts categories comprising our 384-memberEU library can differentially modulate mRuby expression (FIG. that complex, context-specific part interdependence is operative iniougnoui me iiuraiy i riu. / ;. While there is a clear rank order of expression activity across much of the design space we also observe many unexpected context-specific interactions (FIG. 7). For example, the hEFla2 and CMV promoters yield near-maximal expression when paired with a KI Kozak and T4 terminator, yet this Kozak / terminator pairing are relatively weak when combined with another strong promoter, hEFlal.
[0290] By enabling characterization of entire design spaces eri masse, the datasets generated by CLASSIC offer a unique opportunity to explore ML approaches as a way to model non-linear, context-specific part behavior. Though ML methods tend to lack interpretability of regulatory mechanism, abilities to generate large data sets and learn on them has already been leveraged to derive more accurate and generalizable models of complex biological behavior25,26. Therefore, ML models of circuit behavior may be able to make more accurate predictions the context-specific behavior of genetic parts. We tested this possibility with our 384-member EU library data by first fitting it to a mechanistic model that accounts for contributions of parts to transcription and translation (FIG. 24, panel A, left) This model assumes that parts of different categories make independent contributions to mRuby expression, where promoters contribute to transcriptional strength, Kozaks determine strength of translation, and terminators modulate transcript termination rate and the transcript stability. To parameterize the parts, the loglO(mRuby) expression was fit to a linear regression model with 3 categorical predictors for each of the 384 data points. The model generated an r2= 0.86. Inspection of data in the model -predicted vs. CL AS SIC-measured expression scatter plot (FIG. 24, panel A, left) revealed a sigmoidal shape; compositions with the lowest and highest expression were under-predicted by the model, while compositions with intermediate expression were over-predicted.
[0291] We next tested whether an ML-based regression algorithm could outperform the mechanistic model. We chose to use a random forest (RF) model with bootstrap aggregation (FIG. 2, panel F), a model class that pairs excellent predictive power with good interpretability27. The model demonstrated excellent predictive accuracy across all modes of the design space, with an r2= 0.96 (FIG. 24, panel A, right). Interestingly, both the mechanistic model and the RF model showed combination of strong promoters (hEFlal, CMV, or hEFla2) with weak terminators (T8) as those with the greatest difference between the predicted and measured expression values (FIG. 24, panel B). We speculate that this could be due to a transcriptional bottlenecking effect, in which transcripts are being rapidly produced by a strong promoter but ineffectively terminated, leading togreater noise and unpredictable expression profiles. Finally, v generated by the two models are weakly positively correlated (Fi ..panel ). mis suggests mat while both models are capable of predicting expression level for part combinations, the RF is far better at learning and accounting for non-linear part interactions that the mechanistic model fails to capture.
[0292] Scaling the CLASSIC workflow to map complex genetic circuit design spaces.
[0293] Small-molecule-inducible gene circuits are an important molecular tool for enabling user-defined control over gene expression for diverse biotechnological applications4. Despite their apparent functional simplicity, sets of genetic parts and design permutations that can be used to construct the circuits can comprise s large potential design space. Currently, there are limited guiding principles for identifying behavior-optimized circuit designs. To address the challenge of reliably engineering inducible switches with high fold-change behavior for use in mammalian cells, we scaled our CLASSIC pipeline to enable interrogation of design spaces on the order of 105compositions.
[0294] Construction of a 166k-member gene circuit library.
[0295] To construct the ~166k library, we scaled the methods we used previously to generate the 384-member library (FIG. 26, panel A). We first combined 13 level 0 sub-part plasmids to construct a level 1 plasmid pool encoding 48 ORF compositions. The level 1 ORF plasmid pool was then assembled with 4 promoters, 4 terminators, and 3 spacers (48 compositions) to yield an EU library containing 2,304 compositions. In parallel, we generated a level 2 GFP reporter EU pool (36 compositions) and a level 2 barcoded BFP pool (>5» 105variants). The level 2 EU libraries were then assembled into a barcoded destination vector (>104variants) and the purified assembly mixture was transformed into electrocompetent lOBeta E. coli. Nanopore sequencing of the resulting plasmid library was used to characterize: i) assembly accuracy (FIG. 26, panel C); ii) design space coverage (FIG. 3, panel B); iii) relative representation / frequency of each circuit composition in the final library pool (FIG. 8, panel A), and; iv) barcode pool diversity and uniqueness (FIG. 8, panel B). A read length histogram generated from the unfiltered sequencing data showed a prominent peak corresponding to the size of the library (10-13 kb, -65%) as well as a set of smaller peaks representing the empty / re-circularized destination vector (3 kb, -5%) and the unmodified destination vector (8.5 kb, -16%) (FIG. 26, panel C). Assessment of the sequencing data confirmed that the level 3 library contained >95% (158,038 compositions) of the overall design space, with >78.9% containing at least one unique barcode (FIG. 3, panel B). To assess whether part-specific assembly bias was introduced during library construction we determined the relative frequency ofeach level 0 and level 1 part in the final level 3 library (FIG. 8, [ part proportion of level 0 and level 1 was preserved through muiupie rounds or assemuiy, wun less than two-fold category-specific part imbalance within part categories in the final 166k-member library (FIG. 8, panel A). Additionally, the barcode pool generated by the assembly was sufficiently large to cover the design space, with more than 99% of the identified compositions associated with a unique barcode sequence (FIG. 8, panel B).
[0296] Flow sorting pipeline for 166k-member library.
[0297] Our large-scale experiment sought to assay basal and induced expression characteristics for every circuit in our library, with the goal of identifying principles underlying the design of high fold-change circuits. Following puromycin selection of the library-integrated cell population for 14 days (see FIG. 26), 2»108cells were split evenly and cultured with or without 1 pM 4-OHT, an inducer concentration commonly used for Ert2-based activation in mammalian cells that we observed did not affect HEK293T growth rates (results not shown). After a 72 h induction, cells were harvested and prepared for flow-seq (FIG. 27, panel A). As with the 384-member library, cells were gated based on size characteristics and BFP expression and sorted into 8 equally logspaced GFP gates (FIG. 27, panels B, C). Cells from each sorted bin were grown for 5 days postsort and re-measured (FIG. 27, panels B, C). We observed consistent BFP expression across multiple logs of GFP expression. The 166k-member circuit library has a larger BFP-negative population than the 384-member EU library, potentially due in part to a greater fraction of nonlibrary, integration competent DNA species co-transfected with the assembled library (FIG. 26, panel C). Similarly, cells re-grown from the induced population generally exhibit larger BFP- negative cell populations compared to uninduced, re-grown populations. This accounts for the lower mean BFP expression level in induced populations.
[0298] Data processing and validation of the 166k library.
[0299] After applying KDE data processing (FIG. 6) to Illumina data from our 166k-member library, we aggregated barcode read counts to assign mean GFP expression values to the ~121k members of the design space for which we obtained both basal and induced CLASSIC measurements. We discarded all barcodes associated with multiple circuit compositions or compositions for which only a basal or induced measurement was collected. This yielded ~273k barcodes that mapped to —12 Ik circuit compositions (FIG. 3, panel C). We isolated and expanded 40 clones from the integrated library and measured their induced GFP expression distributions via flow cytometry (FIG. 3, panel D). We observed excellent agreement between CLASSIC andindividually isolated basal expression, induced expression, and 1 0.15) for all 40 isolates (FIG. 9).
[0300] While the stoichiometry of the assembled 166k library does not appear to be biased towards any particular parts or part combinations (FIG. 8), it is critical that genomic integration into the landing pad locus also does not suffer from bias as a consequence of circuit composition. To verily this, we examined the 44,596 (27%) compositions for which complete CLASSIC data were not collected and observed no substantial part or part combination bias (FIG. 10). For example, the lowest affinity ZF, which is the most over-represented genetic part in the dropout set, is observed only at 1 ,4-fold greater than its expected frequency.
[0301] Extended 166k- member library data analysis.
[0302] We used CLASSIC to measure 121,292 compositions (-73%) of our ~166k gene circuit design space (FIG. 3, panel C). Because our library measurements reflect circuit behavior with a high degree of accuracy (FIG. 3, panel D), we hypothesized that a data set of this size could be used to train an RF model that accurately predicts basal and induced expression values for un-measured library members, and could potentially be used to refine measured values as well. Since the design, size, and coverage of this library was significantly different from the 384-member (FIG. 2, panel A) library, the RF algorithm needed to be adjusted accordingly. Following an 80:20 randomized traimtest split of the data, we tested two different classes of RFs: a standard bootstrap aggregation, or “bag” forest, and a boosting forest (FIG. 28, panel A, left). While a bagging forest had previously worked well for the 384-member library, we found that a boosting forest showed superior performance for the larger library (test r2= 0.46 for bag, 0.69 for boost) (FIG. 28, panel A, right). This may be due to the ability of the boosting forest to better capture the intrinsic expression level imbalance in our dataset, wherein most (-85%) circuit compositions have low levels of fluorescence when uninduced. Since a boosting forest is built sequentially with a weight applied at each learning cycle to compositions in poorly predicted regions of behavior space, it may be able to more effectively extract information from sparse regions of behavior space compared to a bagging forest, where all trees are trained in parallel and without weights.
[0303] We next assessed whether filtering lower confidence measurements from our dataset would enable better RF performance. We predicted that training on measured values aggregated from a higher number of barcode counts could allow for better prediction accuracy. Indeed, we observed an increase in the test r2as a function of barcoded reads per circuit (FIG. 28, panel B), which rose to a value of 0.80 when trained on circuits with 11 barcodes. While we saw a further increase in r2beyond 11 barcodes, we observed a concurrent increase in standard deviation,indicating a dependence on the specific train:test. Therefore, w filtering threshold. We further filtered our data to account for library imuaiance, testing majority anu minority class relative downsampling (FIG. 28, panel C) with the latter showing superior performance. Finally, we performed a hyperparameter optimization scan, varying the learn rate, number of learning cycles / trees, and maximum number of permitted decision splits (FIG. 28, panel D), using the test r2as the metric for optimization. We found the highest r2from the scan to be 0.83 for basal expression and 0.87 for induced expression (FIG. 4, panel A). To validate that the model was trained accurately, we predicted the expression of our 40 random isolates (FIG. 3, panel C; FIG. 9) and found the model could accurately predict basal (r2= 0.96) and induced (r2= 0.91) expression levels (FIG. 28, panel E). To determine how performance of the RF model compared to more traditional regression fitting approaches, the downsampled data set used to train the RF was trained on a pair of linear regression models where the interaction terms were either linear or quadratic (FIG. 28, panel F). We then predicted basal expression levels for the RF test set (runear2 = 0.52, rquac = 0.63) and our 40 random isolates (rimear2= 0.58, rq,ia,r = 0.81). In both cases, the RF demonstrated superior predictive power.
[0304] We next used predictions from our two RF models to select 96 circuit designs from across behavior space to construct and test, including numerous compositions from the HFC region (>25 fold-change expression, n=55), as well as circuit compositions that were not measured as part of the 121k-member CLASSIC dataset and thus represent true predictions (n=15) (FIG. 11, panels A-C; FIG. 4, panel A right inset; FIG. 4, panel B). Measurements of out-of-sample circuits showed excellent overall agreement between the measured fold-change values and the predictions by our RF models (A74 / ;=O. I6), confirming their accuracy for de novo behavior predictions. The models also accurately predicted fold-change behavior for the 81 in-sample circuits configurations (MAE=Q.16). For configurations from this set that were drawn from the periphery of design space, model predictions proved to be more accurate than our CLASSIC measurements, including for many of the HFC circuits (FIG. 11, panel D). Indeed, the model enabled accurate identification of numerous high performing circuits with fold-change values >50x. Overall, the strong correlation we observe between the RF-modeled and individually measured values across all of the 96 circuits we constructed convincingly demonstrates the robustness and accuracy of the RF model for predicting circuit behavior (FIG. 4, panel B).
[0305] To explore design principles and define how genetic part combinations contribute to circuit behavior, we applied statistical inference methods to RF-generated circuit behavior predictions for the entire design space (166k configurations). We started by identifying five regionsof the behavior space that we deemed functionally significant underlie circuit behavior: low basal expression, high induced expression, IOW uasai anu IOW muuceu, high basal and high induced, and HFC (FIG. 4, panel A). For each region we calculated genetic part frequencies to explore the role that individual components play in determining circuit behavior (FIG. 4, panel C; FIG. 12, panel A). While regions defined by strong induction are exclusively dominated by components that promote strong expression (CMV, mCMV), highly induced activity (VPR), or encourage strong, multi -valent interactions (high affinity ZF, 12 BMs), regions defined by weak induction show a tendency towards components that result in weak transcriptional strength (hPGK, RSV, mTK, ybTATA) lower activity (p65, low affinity ZF). Interestingly HFC configurations appear to strike a balance between strong and weak components and demonstrate less asymmetry in part usage than the other regions. This suggests that most of the genetic parts in this design space have the potential to produce idealized circuit behavior when deployed in the appropriate context. However, we also identified part categories that do not appear to be polarized towards a specific behavior and do not demonstrate region-dependent asymmetry (e.g., terminator and spacer sequences).
[0306] As shown in FIG. 4, panel D, we observed significant part usage coupling in the HFC region by computing the pairwise mutual information between all pairs of features in different regions of behavior space (low basal, high induced, and HFC). However, for accurate comparison of the degree of coupling across these regions, it is also important to account for differences in part category distributions. We thus computed the relative mutual information (RM1), defined as the mutual information (in bits) divided by the product of the ID entropy of the features alone. The relative mutual information for all features across the space is shown in FIG. 12, panel C, and is consistent with the MI score itself, indicating that differences in the regions are indeed due to coupling between part categories.
[0307] To identify different routes through which part combinations can yield HFC genetic circuits, we carried out dimensional reduction of the seven part categories with the greatest spread in part usage (FIG. 12, panel A, red labels). Following gap penalty evaluation and K-means clustering we identified three distinct solution families (clusters A, B, and C) (FIG. 4, panel E; FIG. 13, panel A). The three routes to HFC behavior use a similar reporter architecture (8 BMs and a high-activity core promoter) and are predominantly defined by the composition of the synTF EU (FIG. 4, panel F). The circuit compositions that comprise each cluster are distributed evenly across the HFC region, indicating that even with a limited set of parts (i.e., within a single cluster) there is the opportunity for diverse expression profiles (FIG. 13, panel B). Closer inspection of the clustersreveals slightly higher average basal and induced expression leve decrease in the average and range of fold change for this cluster reiauve io clusters A anu v (r i . 13, panel C). Interestingly, cluster A contains a greater proportion of circuits with low basal expression and high basal expression than the other two clusters, leading to the most ideal fold change characteristics of the three clusters. Of the 96 intentionally constructed circuit compositions, 61 exist within the high fold change region of behavior space. Roughly even distribution of cluster assignment across the 61 HFC cell lines confirms that these clusters are equally effective routes to productive circuit behavior (FIG. 13, panel D).
[0308] All plasmids and sequences used in this study are listed in Appendix A.. The expression level and genetic compositions of all cell lines used for validation described herein. All code used to analyze sequencing data is available at github.com / cbashorlab / CLASSIC. The background strain in this study was a HEK293T-LP cell line integrated with a plasmid encoding a puromycin resistance gene and a BFP cassette, both of which are driven by hEFlal. Example gating for the 384-member library is shown in FIG. 29.
[0309] References Cited in this Example:
[0310] 1. Hnisz, D., Day, D. S. & Young, R. A. Insulated Neighborhoods: Structural and Functional Units of Mammalian Gene Control. Cell 167, 1188-1200, doi:10.1016 / j.cell.20I6.10.024 (2016).
[0311] 2. Brophy, J. A. & Voigt, C. A. Principles of genetic circuit design. Nat Methods 11, 508-520, doi : 10.1038 / nmeth.2926 (2014).
[0312] 3. Bashor, C. J., Hilton, I. B., Bandukwala, H., Smith, D. M. & Veiseh, O. Engineering the next generation of cell-based therapeutics. Nat Rev Drug Discov 21, 655-675, doi : 10.1038 / s41573-022-00476-6 (2022).
[0313] 4. Bashor, C. J. & Collins, J. J. Understanding Biological Regulation Through Synthetic Biology. Annu Rev Biophys 47, 399-423, doi: 10.1146 / annurev-biophys-070816- 033903 (2018).
[0314] 5. Weber, E., Engler, C., Gruetzner, R., Werner, S. & Marillonnet, S. A modular cloning system for standardized assembly of multigene constructs. PLoS One 6, el6765, doi: 10.1371 / journal. pone.0016765 (2011).
[0315] 6. Lee, M. E., DeLoache, W. C., Cervantes, B. & Dueber, J. E. A Highly Characterized Yeast Toolkit for Modular, Multipart Assembly. ACS Synth Biol 4, 975-986, doi: 10.1021 / sb500366v (2015).
[0316] 7. Engler, C. et al. A golden gate modular cloniBiol3, 839- 843, doi: 10.1021 / sb4001504 (2014).
[0317] 8. Fonseca, J. P. et al. A Toolkit for Rapid Modular Construction of Biological Circuits in Mammalian Cells. ACS Synth Biol 8, 2593-2606, doi:10.1021 / acssynbio.9b00322 (2019).
[0318] 9. Dao-Thi, M. H. et al. Molecular basis of gyrase poisoning by the addiction toxin CcdB . J Mol Biol 348, 1091 - 1102, doi : 10.1016 / j .j mb .2005.03.049 (2005).
[0319] 10. To, K. Y. et al. Analysis of the gene cluster encoding carotenoid biosynthesis in Erwinia herbicola Ehol3. Microbiology (Reading) 140 ( Pt 2), 331-339, doi:10.1099 / 13500872- 140-2-331 (1994).
[0320] 11. Kinney, J. B., Murugan, A., Callan, C. G., Jr. & Cox, E. C. Using deep sequencing to characterize the biophysical mechanism of a transcriptional regulatory sequence. Proc Natl Acad Sci USA 107, 9158-9163, doi: 10.1073 / pnas.1004290107 (2010).
[0321] 12. Philpott, M. et al. Nanopore sequencing of single-cell transcriptomes with scCOLOR-seq. Nat Biotechnol 39, 1517- 1520, doi : 10.1038 / s41587-021 -00965-w (2021 ).
[0322] 13. Schmidt, K. et al. Identification of bacterial pathogens and antimicrobial resistance directly from clinical urines by nanopore-based metagenomic sequencing. J Antimicrob Chemother 72, 104-114, doi: 10.1093 / jac / dkw397 (2017).
[0323] 14. Curry, K. D. et al. Emu: species-level microbial community profiling of full- length 16S rRNA Oxford Nanopore sequencing data. Nat Methods 19, 845-853, doi : 10.1038 / s41592-022- 01520-4 (2022).
[0324] 15. Currin, A. et al. Highly multiplexed, fast and accurate nanopore sequencing for verification of synthetic DNA constructs and sequence libraries. Synth Biol (Oxf) 4, ysz025, doi : 10.1093 / synbio / y sz025 (2019).
[0325] 16. de Boer, C. G. et al. Deciphering eukaryotic gene-regulatory logic with 100 million random promoters. Nat Biotechnol 38, 56-65, doi:10.1038 / s41587-019-0315-8 (2020).
[0326] 17. Smola, M. J., Rice, G. M., Busan, S., Siegfried, N. A. & Weeks, K. M. Selective 2'-hydroxyl acylation analyzed by primer extension and mutational profiling (SHAPE-MaP) for direct, versatile and accurate RNA structure analysis. Nat Protoc 10, 1643-1669, doi : 10.1038 / nprot.2015.103 (2015).
[0327] 18. Rouches, M. V., Xu, Y., Cortes, L. B. G. & Lambert, G. A plasmid system with tunable copy number. Nat Commun 13, 3908, doi: 10.1038 / s41467-022-31422-0 (2022).
[0328] 19. Cong, L. et al. Multiplex genome engineerScience 339, 819-823, doi: 10.1126 / science, 1231143 (2013).
[0329] 20. O'Gorman, S., Fox, D. T. & Wahl, G. M. Recombinase-mediated gene activation and site- specific integration in mammalian cells. Science 251, 1351-1355, doi : 10.1126 / science.1900642 (1991).
[0330] 21. Duportet, X. et al. A platform for rapid prototyping of synthetic gene networks in mammalian cells. Nucleic Acids Res 42, 13440-13451, doi: 10.1093 / nar / gkul082 (2014).
[0331] 22. Bachhav, B., de Rossi, J., Llanos, C. D. & Segatori, L. Cell factory engineering: Challenges and opportunities for synthetic biology applications. Biotechnol Bioeng, doi: 10.1002 / bit.28365 (2023).
[0332] 23. Johnstone, C. P. & Galloway, K. E. Supercoiling-mediated feedback rapidly couples and tunes transcription. Cell Rep 41, 111492, doi: 10.1016 / j.celrep.2022.111492 (2022).
[0333] 24. Nielsen, A. A. et al. Genetic circuit design automation. Science 352, aac7341, doi : 10.1126 / science. aac7341 (2016).
[0334] 25. Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589, doi: 10.1038 / s41586-021-03819-2 (2021).
[0335] 26. Avsec, Z. et al. Effective gene expression prediction from sequence by integrating long- range interactions. Nat Methods 18, 1196-1203, doi:10.1038 / s41592-021- 01252-x (2021).
[0336] 27. Couronne, R., Probst, P. & Boulesteix, A. L. Random forest versus logistic regression: a large-scale benchmark experiment. BMC Bioinformatics 19, 270, doi: 10.1186 / S12859- 018-2264-5 (2018).EXAMPLE 4
[0337] Plasmids and sequences described herein.
[0338] Table 1. Genetic Parts described herein.
[0339] Table 2. Sesame Street Overhangs described herein.
[0340] Table 3. Plasmids described herein.atgagtgacgactgaatccggtgagaatgg aggccagccattacgctcgtcatcaaaatca cgcctgagccagtcgaaatacgcgatcgctgttaaaaggacaattacaaacaggaatcgaatg caaccggcgcaggaacactgccagcgcatcaacaatattttcacctgaatcaggatattcttcta atacctggaatgctgtttttccggggatcgcagtggtgagtaaccatgcatcatcaggagtacgg ataaaatgcttgatggtcggaagaggcataaattccgtcagccagtttagtctgaccatctcatct gtaacatcattggcaacgctacctttgccatgtttcagaaacaactctggcgcatcgggcttccca tacaagcgatagattgtcgcacctgattgcccgacattatcgcgagcccatttatacccatataaa tcagcatccatgttggaatttaatcgcggcctcgacgtttcccgttgaatatggctcataacaccc cttgtattactgtttatgtaagcagacagttttattgttcatgatgatattattttatcttgtgcaatgtaa catcagagattttgagacacggaatcgacgctcaagtcagaggtcctgacgtcgacggatcgg gagatctcccgatcccctatggtgcactctcagtacaatctgctctgatgccgcatagttaagcca gtatctgctccctgcttgtgtgttggaggtcgctgagtaGGATCCtGAAGACatGGA GTCGCtGAGACGCGTCTCaCAGTGGCAatGTCTTCcGGATCCgggtcatagctgtttcctgt gtgaaattgttatccgctcacaattccacacaacatacgagccggaagcataaagtgtaaagcct ggggtgcctaatgagtgagctaactcacattaattgcgttgcgctcactgcccgctttccagtcg ggaaacctgtcgtgccagctgcattaatgaatcggccaacgcgcggggagaggcggtttgcg tattgggcgctcttccgcttcctcgctcactgactcgctgcgctcggtcgttcggctgcggcgag cggtatcagctcactcaaaggcggtaatacggttatccacagaatcaggggataacgcaggaa agaacatgtgagcaaaaggccagcaaaaggccaggaaccgtaaaaaggccgcgttgctggc gtttttccataggctccgcccccctgacgagcatcacaaaaatcgacgctcaagtcagaggtgg cgaaacccgacaggactataaagataccaggcgtttccccctggaagctccctcgtgcgctctc ctgttccgaccctgccgcttaccggatacctgtccgcctttctcccttcgggaagcgtggcgcttt ctcatagctcacgctgtaggtatctcagttcggtgtaggtcgttcgctccaagctgggctgtgtgc acgaaccccccgttcagcccgaccgctgcgccttatccggtaactatcgtcttgagtccaaccc ggtaagacacgacttatcgccactggcagcagccactggtaacaggattagcagagcgaggt atgtaggcggtgctacagagttcttgaagtggtggcctaactacggctacactagaagaacagt atttggtatctgcgctctgctgaagccagttaccttcggaaaaagagttggtagctcttgatccgg caaacaaaccaccgctggtagcggtttttttgtttgcaagcagcagattacgcgcagaaaaaaa ggatctcaagaagatcctttgatcttttctacggggtctgacgctcagtggaacgaatgtggtaat gctctgccagtgttacaaccaattaaccaattctgattagaaaaactcatcgagcatcaaatgaaa ctgcaatttattcatatcaggattatcaataccatatttttgaaaaagccgtttctgtaatgaaggag aaaactcaccgaggcagttccataggatggcaagatcctggtatcggtctgcgattccgactcg tccaacatcaatacaacctattaatttcccctcgtcaaaaataaggttatcaagtgagaaatcacc atgagtgacgactgaatccggtgagaatggcaaaagtttatgcatttctttccagacttgttcaac aggccagccattacgctcgtcatcaaaatcactcgcatcaaccaaaccgttattcattcgtgattg cgcctgagccagtcgaaatacgcgatcgctgttaaaaggacaattacaaacaggaatcgaatg caaccggcgcaggaacactgccagcgcatcaacaatattttcacctgaatcaggatattcttcta atacctggaatgctgtttttccggggatcgcagtggtgagtaaccatgcatcatcaggagtacgg ataaaatgcttgatggtcggaagaggcataaattccgtcagccagtttagtctgaccatctcatct gtaacatcattggcaacgctacctttgccatgtttcagaaacaactctggcgcatcgggcttccca tacaagcgatagattgtcgcacctgattgcccgacattatcgcgagcccatttatacccatataaa tcagcatccatgttggaatttaatcgcggcctcgacgtttcccgttgaatatggctcataacaccc cttgtattactgtttatgtaagcagacagttttattgttcatgatgatattattttatcttgtgcaatgtaa catcagagattttgagacacggaatcgacgctcaagtcagaggtcctgacgtcgacggatcgg gagatctcccgatcccctatggtgcactctcagtacaatctgctctgatgccgcatagttaagcca gtatctgctccctgcttgtgtgttggaggtcgctgagtaGGATCCtGAAGACatGGAGTCGCtGAGACG
[0341] Table 4. Genotyping Primers described herein.
[0342] Table 5. ddPCR Primers described herein.
[0343] Table 6. Next-Generation Sequencing Primers described herein.
[0344] Table 7. Isolate Amplification Primers described herein.EQUIVALENTS
[0345] Those skilled in the art will recognize, or be able to ascertain, using no more than routine experimentation, numerous equivalents to the specific substances and procedures described herein. Such equivalents are considered to be within the scope of this invention and are covered by the following claims.
Claims
CLAIMSWhat is claimed:
1. A high-throughput method for screening a gene circuit library, the method comprising: preparing an indexed circuit library comprising a plurality of barcoded- linked nucleic acids by combining a first pool of a plurality of nucleic acids with a second pool of a plurality of nucleic acids, wherein each nucleic acid in the first pool comprises one or more expression unit(s) (EU(s)), with each EU comprising at least a promoter, an open reading frame, and a terminator, and wherein each nucleic acid in the second pool comprises a barcode; providing a construct-to-barcode index using long-read sequencing of the index circuit library; providing a barcode-to-phenotype index by integrating the index circuit library into a population of cells, sorting the population of cells according to phenotype, and using short-read amplicon sequencing; and comparing the construct-to-barcode index with the barcode-to- phenotype index to associate an expression unit with a phenotype, thereby screening the gene circuit library.
2. The method of claim 1, further comprising assembling the first pool of a plurality of nucleic acids, assembling the second pool of a plurality of nucleic acids or both.
3. The method of claim 2, wherein the assembling is hierarchical.
4. The method of claim 1, wherein the barcodes are structured.
5. The method of claim 1, further comprising analyzing long-read sequencing data, thereby allowing identification of sequence variants from a single read.
6. The method of claim 5, wherein the long-read sequencing data is analyzed using a software pipeline.
7. The method of claim 6, wherein the software pipeline identifies circuit components and barcode sequences from individual sequencing data reads.
8. The method of claim 1, wherein the phenotype comprises a gene expression output.
9. The method of claim 1 as depicted in FIG. 1.
10. The method of claim 1, wherein the population of cells comprises bacteria cells, yeast cells, or mammalian cells.
11. The method of claim 10, wherein the population of c cells, mesenchymal stem cells (MSCs), induced pluripoiem stem ceus (irsvsj, or IKJO cells.
12. The method of claim 1, wherein expression unit (EU) further comprises a chromatin insulator, enhancer, promoter, Kozak sequence, activation domain (AD), intrinsically disordered region (IDR), 4-hydroxytamoxifen-responsive domain, a zinc finger (ZF) domain, a terminator, a spacer sequence, varied orientation, or any combination thereof.
13. The method of claim 1, wherein the first pool of a plurality of nucleic acids comprises a diverse pool of nucleic acids.
14. The method of claim 1, wherein the open reading frame region encodes for a small molecule-inducible transcription factor or a fluorescent reporter element.
Citation Information
Patent Citations
RNA probe for mutation profiling and use thereof
US20240052339A1