High throughput screening of expression constructs using microcapillary arrays

The microcapillary array platform addresses inefficiencies in current screening techniques by enabling high-throughput, efficient analysis of gene expression and protein production, significantly enhancing the discovery of optimal plasmid constructs and protein sequences for biomanufacturing.

WO2026161689A1PCT designated stage Publication Date: 2026-07-30THE PENN STATE RES FOUND INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
THE PENN STATE RES FOUND INC
Filing Date
2026-01-23
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Current cell sorting and screening techniques, particularly for large-scale plasmid libraries, lack efficiency and throughput, and are impractical for analyzing clonal libraries due to inadequate spatial isolation and complicated handling methods, making them unsuitable for biomanufacturing applications.

Method used

A microcapillary array-based screening platform with an optical module and control unit for high-throughput analysis of gene expression and protein production, utilizing photoactive reporters and optical emissions for quantification and population binning, enabling efficient screening of miniature cell cultures.

Benefits of technology

The platform achieves a 102-103fold increase in screening rate compared to traditional methods, allowing for rapid identification of optimal plasmid constructs and protein sequences, and establishing genotype-phenotype correlations with high enrichment ratios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026012344_30072026_PF_FP_ABST
    Figure US2026012344_30072026_PF_FP_ABST
Patent Text Reader

Abstract

A screening platform for analyzing gene expression or evaluating protein production is provided herein. The screening platform includes an array of microcapillary tubes each configured to contain a volume of cell culture. The screening platform further includes an optical module having a light source configured to illuminate the volume of cell culture, thereby producing an optical emission; and a detector configured to detect the optical emission. The screening platform further includes a control unit communicatively connected to the optical module and configured to quantify a level of gene expression or protein production of the cell culture as a function of the optical emission; optionally, perform population binning as a function of the level of gene expression or protein production; and identify a genotype of the cell culture as a function of the quantification and, optionally, the population binning, thereby providing population-based screening of the cell culture.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 0073605-001135 HIGH THROUGHPUT SCREENING OF EXPRESSION CONSTRUCTS USING MICROCAPILLARY ARRAYS CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority of U.S. Provisional Application Serial No.63 / 749,069, filed on January 24, 2025, and entitled “High Throughput Screening of Expression Constructs Using Microcapillary Arrays”, the entirety of which is incorporated herein by reference. FIELD OF THE INVENTION

[0002] The present invention generally relates to the field of cell sorting and screening. In particular, the present invention is directed to high throughput screening of expression constructs using microcapillary arrays.REFERENCE TO SEQUENCE LISTING

[0003] This specification includes a sequence listing submitted herewith, which includes the file entitled 0073605-001135.xml having the following size: 157,463 bytes, which was created January 21, 2026, the contents of which are incorporated by reference herein.BACKGROUND

[0004] Gene expression is a complex phenomenon with numerous interlinked variables, and studying these variables to control gene expression plays an important role in bioengineering and biomanufacturing. Plasmids are often used to introduce heterologous genes into host cells, where desired protein products are expressed. Cost-effective protocols for rapid synthesis of large plasmid libraries are being developed, and techniques that can complement such libraries with multiplex screening at similar scales are highly sought after. Cloning strategies like mutagenesis and combinatorial fragment assembly enable the construction of libraries covering vast design spaces, which can inform protein structure-expression relations and the discovery of new protein sequences (e.g., with similar properties but improved expression) through screening. Moreover, algorithms for the massively parallel design of plasmid elements (e.g., promoters, ribosome binding sites (RBSs)) have opened an avenue for expansive optimization of the expression of recombinant genes.However, while cloning techniques to achieve plasmid libraries covering large design spaces exist, multiplex techniques that offer cell culture screening at similar scales are still lacking.

[0005] Cell sorting and screening techniques are important for biological research, bioengineering, and medicine, as they allow the isolation of specific clones from heterogenous populations and enable analysis of a myriad of cell phenotypes in studies involving drug and immune responses, disease progression, enzyme and biomarker engineering, and cell differentiation.Attorney Docket No. 0073605-001135 In addition to biology and medicine, screening techniques find potential in the biomanufacturing and bioprocessing industry, e.g., biopharmaceuticals, biofuel, food and agriculture, and bioremediation, all of which require protein-titer optimization.

[0006] While drastic achievements in single-cell screening have been made with microfluidic platforms, advancements in population-based screening are still lacking. Microwell plate readers and multi-parallel bioreactors are current workhorse tools for screening cell cultures in laboratories worldwide. However, these platforms are not suitable due to their lack of efficiency in analyzing large clonal libraries, inadequate throughput, or complicated handling methods. Screening libraries on microwell plates is impractical, as it requires differentiation of thousands of variants on agar plates followed by picking and inoculation. The lack of a mechanism for spatial isolation of cells makes these technologies impractical for library cloning protocols that produce mixtures of variants as final products.SUMMARY OF THE DISCLOSURE

[0007] An aspect of the present disclosure is a screening platform for analyzing gene expression or evaluating protein production. The screening platform includes an array of microcapillary tubes each configured to hold or contain a volume of cell culture (i.e., each capillary tube is configured to contain miniature cell culture). The cell culture of at least one microcapillary tube of the array of microcapillary tubes is configured to express at least a photoactive reporter. The screening platform further includes an optical module having a light source configured to illuminate the volume of cell culture, thereby producing an optical emission, and a detector configured to detect the optical emission. The screening platform further includes a control unit communicatively connected to the optical module. The control unit is configured to quantify a level of gene expression or protein production of the cell culture as a function of the optical emission. In some embodiments, the control unit is further configured to perform population binning of cell clones as a function of the level of gene expression or protein production. The control unit is further configured to identify a genotype of the cell culture as a function of the quantification and, optionally, the population binning, thereby providing population-based screening of the cell culture. In some embodiments, the screening platform includes, is communicatively connected to, or otherwise performs one or more functions as a data acquisition and analysis system.

[0008] In some embodiments of the screening platform, the array of microcapillary tubes includes an array of at least 105microcapillary tubes.Attorney Docket No. 0073605-001135

[0009] In some embodiments of the screening platform, the light source of the optical module includes a multiband illumination system across a wavelength range of at least 300 nm and no greater than 800 nm.

[0010] In some embodiments of the screening platform, the optical module is configured to perform a bright-field absorbance measurement.

[0011] In some embodiments of the screening platform, the at least a photoactive reporter includes at least a fluorescent reporter. In some embodiments, the at least a fluorescent reporter functions as a fluorescent biosensor for in vivo measurements, with substantially no toxicity to the cell culture.

[0012] In some embodiments of the screening platform, the at least a fluorescent reporter includes a plurality of fluorescent reporters. Accordingly, in such embodiments, the control unit is further configured to perform multi-reporter operon screening, such as without limitation dualreporter operon screening, as a function of the optical emission. In some embodiments of the screening platform, performing the multi-reporter operon screening includes performing a comparative analysis of ribosome binding sites. In some embodiments of the screening platform, performing the multi-reporter operon screening includes performing a comparative analysis of operon dynamics.

[0013] In some embodiments of the screening platform, identifying the genotype of the cell culture includes isolating a phenotypic trait of the cell culture as a function of the level of gene expression or protein production; and identifying the genotype of the cell culture as a function of the phenotypic trait.

[0014] In some embodiments of the screening platform, the control unit is further configured to determine a level of transcriptional activity as a function of the level of protein production.

[0015] In some embodiments, the screening platform further includes a chemical perturbation module communicatively connected to the control unit and configured to apply a controlled chemical stimulus to the cell culture, wherein the control unit is further configured to evaluate an effect of the controlled chemical stimulus on the gene expression or protein production in the cell culture. In some embodiments of the screening platform, applying the controlled chemical stimulus includes applying one or more inducers.

[0016] In some embodiments of the screening platform, the control unit is further configured to measure a half-life of mRNA as a function of the optical emission and correlate the half-life and the level of protein production.Attorney Docket No. 0073605-001135

[0017] In some embodiments of the screening platform, the control unit is further configured to generate a cell growth profile of the cell culture and determine one or more structural properties of a protein as a function of or in relation to the cell growth profile. In some embodiments of the screening platform, the one or more structural properties of the protein include hydrophobicity of the protein.

[0018] In some embodiments of the screening platform, the control unit is further configured to generate an output as a function of the level of gene expression or protein production.

[0019] In some embodiments of the screening platform, the output includes one or more of a suitable plasmid construct for protein production, a suitable promoter for protein production, a suitable 5’ untranslated region (5’ UTR) for protein production, and a suitable amino acid sequence for protein production.

[0020] In some embodiments of the screening platform, the control unit is further configured to display the output using an output device or output interface. In some embodiments, the output device or output interface is configured to generate one or more quantitative datasets identifying suitable or optimal plasmid constructs, including without limitation one or more suitable combinations of promoter(s), 5’ UTR(s), and amino acid sequence(s) that increase or maximize protein production.

[0021] In some embodiments, the screening platform is configured to screen one or more plasmid design libraries. Each plasmid design library can include a plurality of promoters with varying transcription rates, 5’ UTRs configured to modulate mRNA stability and / or having various translation initiation rates (TIRs), and / or amino acid sequence variants adapted or optimized to reduce hydrophobicity and enhance protein solubility and / or expression.

[0022] In some embodiments, the screening platform further includes an image-processing module communicatively connected to the control unit, wherein generating the output includes generating an intensity map, using the image-processing module, as a function of the optical emission.

[0023] In some embodiments of the screening platform, the output includes one or more coordinates or indices of one or more microcapillaries of interest.

[0024] In some embodiments, the screening platform further includes one or more extraction slips each configured to collect the volume of cell culture extracted from the one or more microcapillaries of interest.Attorney Docket No. 0073605-001135

[0025] In some embodiments of the screening platform, the volume of cell culture contains one or more magnetic beads and is extracted using a focused pulse-laser beam.

[0026] In some embodiments of the screening platform, one or more plasmids of the extracted volume of cell culture are further amplified and sequenced using polymerase chain reaction (PCR).

[0027] In some embodiments, wherein the screening platform provides a clone recovery mechanism having an enrichment ratio of up to 105: 1.

[0028] Another aspect of the present disclosure is a method for analyzing gene expression or evaluating protein production. The method includes applying a plurality of cell cultures to an array of microcapillary tubes each configured to contain a volume of cell culture. The cell culture in at least one microcapillary tube of the array of microcapillary tubes is configured to express at least a photoactive reporter. The method further includes measuring an optical emission of the plurality of cell cultures using an optical module having a light source and a detector. The method further includes quantifying, using a control unit communicatively connected to the optical module, a level of gene expression or protein production in the plurality of cell cultures, as a function of the optical emission. In some embodiments, the method further includes performing population binning, using the control unit, as a function of the level of gene expression or protein production. The method further includes identifying a genotype of at least one cell culture of the plurality of cell cultures, using the control unit, as a function of the quantification and, optionally, the population binning, thereby providing population-based screening of the cell culture.

[0029] In some embodiments of the method, the array of microcapillary tubes includes an array of at least 105microcapillary tubes.

[0030] In some embodiments of the method, the light source includes a multiband illumination system across a wavelength range of at least 300 nm and no greater than 800 nm.

[0031] In some embodiments, the method further includes measuring a bright-field absorbance of the volume of cell culture using the optical module.

[0032] In some embodiments of the method, the at least a photoactive reporter includes at least a fluorescent reporter.

[0033] In some embodiments of the method, the at least a fluorescent reporter includes a plurality of fluorescent reporters. Accordingly, in some embodiments, the method further includes performing multi-reporter operon screening, such as without limitation dual-reporter operon screening, using the control unit, as a function of the optical emission.Attorney Docket No. 0073605-001135

[0034] In some embodiments of the method, performing the multi-reporter operon screening includes performing a comparative analysis of ribosome binding sites. In some embodiments of the method, performing the multi-reporter operon screening includes performing a comparative analysis of operon dynamics.

[0035] In some embodiments of the method, identifying the genotype of the cell culture includes isolating a phenotypic trait of the cell culture as a function of the level of gene expression or protein production and identifying the genotype of the cell culture as a function of the phenotypic trait.

[0036] In some embodiments, the method further includes determining a level of transcriptional activity as a function of the level of protein production.

[0037] In some embodiments, the method further includes applying a controlled chemical stimulus to the cell culture, using a chemical perturbation module communicatively connected to the control unit, and evaluating an effect of the controlled chemical stimulus on the gene expression or protein production in the cell culture.

[0038] In some embodiments of the method, applying the controlled chemical stimulus includes applying one or more inducers.

[0039] In some embodiments, the method further includes measuring the half-life of mRNA, using the control unit, as a function of the optical emission, and correlating the half-life and the level of protein production using the control unit.

[0040] In some embodiments, the method further includes generating a cell growth profile of the cell culture using the control unit, and determining, using the control unit, one or more structural properties of a protein as a function of the cell growth profile. In some embodiments of the method, the one or more structural properties of the protein includes hydrophobicity of the protein.

[0041] In some embodiments, the method further includes generating an output as a function of the level of gene expression or protein production.

[0042] In some embodiments of the method, the output includes one or more of a suitable plasmid construct for protein production, a suitable promoter for protein production, a suitable 5’ untranslated region for protein production, and a suitable amino acid sequence for protein production.

[0043] In some embodiments, the method further includes displaying the output using an output device communicatively connected to the control unit.Attorney Docket No. 0073605-001135

[0044] In some embodiments of the method, generating the output includes generating an intensity map, using an image-processing module communicatively connected to the control unit, as a function of the optical emission.

[0045] In some embodiments of the method, the output includes one or more coordinates or indices of one or more microcapillaries of interest.

[0046] In some embodiments, the method further includes collecting the volume of cell culture extracted from the one or more microcapillaries of interest using one or more extraction slips.

[0047] In some embodiments of the method, the volume of cell culture contains one or more magnetic beads and is extracted using a focused pulse-laser beam.

[0048] In some embodiments, the method further includes amplifying and sequencing one or more plasmids of the extracted volume of cell culture using polymerase chain reaction.

[0049] In some embodiments, the method provides a clone recovery mechanism having an enrichment ratio of up to 105: 1.

[0050] These and other aspects and features of nonlimiting embodiments of the present invention will become apparent to those skilled in the art upon review of the following description of specific nonlimiting embodiments of the invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] For the purpose of illustrating the invention, the drawings show aspects of one or more embodiments of the invention.

[0052] FIG. 1A-1D depict aspects of an exemplary microcapillary array -based platform for screening expression constructs. (1A) The workflow aims to optimize biomanufacturability by sorting plasmid libraries. Desired clones can then undergo bio-fermentation techniques to validate titer levels and production. (IB) Three parameters studied: (i) promoters, (ii) 5’ untranslated regions, and (iii) protein sequences. (1C) Two adjacent capillaries are shown; one (encircled) is extracted by the pulsed laser while the other remains undisturbed. The contents are collected on the extraction slip (image on the right). (ID) Upon pulsed-laser excitation, the magnetic beads superheat, disrupting the meniscus and discharging the contents collected on the extraction slip beneath the microcapillary array.

[0053] FIG. 2A-2D depict aspects of an exemplary microcapillary array -based screening method. (2A) The screening methodology is illustrated with details on microarray, image-processing algorithm, and laser-based extraction. (2B) The experimental screening setup is shown, including the inverted fluorescence microscope, a collimated UV laser beam, and an extraction slip setup. (2C)Attorney Docket No. 0073605-001135 Representative microarray images with mRFP1-expressing pFTV1 plasmid in Escherichia coli after various incubation periods are shown. Over time, cells replicate and occupy full capillary faces. Scale bars = 100 pm. (2D) Average capillary intensity (total capillaries = 2,695) across all incubation times.

[0054] FIG. 3A-3H depict exemplary data based on screening of promoter library. (3 A) Construction of the mRFP1-expressing plasmid is shown. (3B) Image of an array loaded with the promoter library after 12 h incubation. (3C) Intensity maps from three replicates of the library screening, scanning approximately 68,000, 77,000, and 72,000 capillaries, respectively. Five capillaries were extracted from the top bins. (3D) Transcription rates of extracted variants, estimated from DNA / RNA sequencing counts, include only sequences with a 100% match. (3E) Images of a capillary from the low (left, norm, intensity: 0.0332 au), medium (middle, norm, intensity: 0.2429 au), and high (right, norm, intensity: 0.6266 au) bins show intensity variation. (3F) The normalized intensity map is divided into low [0, 0.1], medium [0.1, 0.316], and high [0.316, 1] bins for the binning experiment. (3G) DNA / RNA sequencing read count data of identified promoters from each bin are shown as two replicates. (3H) Distribution of transcription rates from extracted promoters in bins, with crosses representing means in box-whisker plots.

[0055] FIG. 4A-4L depict exemplary data based on screening of 5’ UTR library. (4A) Construction of the sfGFP and mRFP1-expressing plasmid is shown. (4B, 4C) Representative images of the array loaded with the UTR library for sfGFP and mRFP1 signals are shown. (4D) Intensity maps forbinning experiments for sfGFP and mRFP1 screening. (4E, 4F) Half-life and R2of fit for the decay constant are plotted for the identified 5’ UTRs. (4G) Binarized bright-field image of an array as the mask for sfGFP and mRFP1 images is depicted. (4H) Scatter plot of sfGFP and mRFP1 fluorescence intensity in each capillary. (41) Intensity map for the sfGFP / mRFP1screening is shown. (4J) Ratios of translation initiation rates of RBSs for sfGFP to mRFP1 compared across sfGFP bins and ratio-based screened clones. (4K) Intensity map for the summation-based screening is depicted. Ten capillaries were extracted from the top bins and were sequence-identified in (4I, 4K). (4L) Translation initiation rates of sfGFP and mRFP1 RBS variants identified in the arithmeticbased screening experiments. Cross represents the mean in box-whisker plots.

[0056] FIG 5A-5G depict exemplary data based on chemical perturbation of structural proteins and screening. (5 A) Plasmid map shows structural protein linked to mCherry. Coding sequences are modified for cement and reflectin proteins, keeping RBS strength comparable. HydrophobicityAttorney Docket No. 0073605-001135 scores: cement (-75.26 au), reflectin (14 au). (5B) The strategy for chemical perturbation of cells is depicted. (5C) Representative fluorescence images of cement and reflectin proteins with IPTG induction at t = 6 h, 48 h incubation. Scale bars = 100 / zm. (5D) Fluorescence intensity distribution of capillaries (top 500) at various induction conditions after 48 h incubation is plotted (-30,000 total capillaries scanned). (5E) Schematic showing how cell density affects bright-field light absorption in capillaries. (5F) Bright-field image of cement protein array t = 6 h induction) shows capillaries with different cell densities and absorbance. (5G) Bright-field absorbance distributions (top 500 capillaries) for cement and reflectin highlight better growth profiles for cement with delayed IPTG induction. The cross symbolizes the mean in box -whisker plots.

[0057] FIG. 6 depicts an exemplary laser-based extraction of a single, isolated capillary. The plot depicts the laser pulse train with 2 ms wide pulses with 2 ms separation at DAQ level of 118. The number of pulses was adjusted for efficient extraction.

[0058] FIG. 7A-7B depict exemplary data from cell growth experiment with pFTV1 plasmid. (7A) Representative fluorescence images of pFTV1 plasmid expressing mRFP1 protein in E. coli at different incubation times (constant exposure). Cell growth and protein expression increase over time can be noticed. Scale bars = 100 pm. (7B) Raw intensity map of array after each incubation period is illustrated. A total of « 8,400 capillaries were scanned (x-axis limited to 4,000 here), of which 2,695 were identified to have acquired cell(s).

[0059] FIG. 8 depicts exemplary data from Promoter library replicates, including raw intensity vs. frequency maps for the three replicates of the promoter library.

[0060] FIG. 9A-9I depict exemplary aspects of a pixel-dilation strategy for low bin extraction. Data pertaining to the promoter library are utilized here. Representative original (9 A) and binarized (9B) images of capillaries identified as singular objects (norm, intensity = 0.41, 0.60).Representative original (9C, 9F) and binarized (9D, 9G) images of capillaries identified as clusters of several objects from the medium (9C) and low (9F) bins with normalized intensities 0.11 and 0.03, resp. Upon pixel dilation, separate objects within individual capillaries are merged (9E, 9H). (91) Reduction in frequency of the lowest bin (= 7’) of the raw intensity vs. frequency map due to pixel dilation. Upon pixel-dilation, the distribution is stretched resulting in a few medium bin clones to be included in low bin. We quantified this overlap to be 3.2%, which is minimal.

[0061] FIG. 10 depicts exemplary data from screening of 5’ UTR library, including raw intensity vs. frequency maps for sfGFP and mRFP1 signals of the 5’ UTR library (total capillaries scanned « 393,000).Attorney Docket No. 0073605-001135

[0062] FIG. 11 depicts aspects of an image acquisition and image-processing strategy for the ratio-based screening. At discrete locations on the array loaded with the 5’ UTR library, bright-field, sfGFP, and mRFP1 snapshots were acquired (total locations = 121, total snapshots = 363) after 12 hr. incubation period. Bright-field images were binarized and capillary faces were identified as white objects in the images. Indices of the pixels within each capillary were recorded in a matrix. These pixels were visited in the fluorescence images, and their values were summed to estimated intensities of sfGFP and mRFP1 signals from each capillary. Lastly, ratios of the intensities were calculated for all capillaries. Brightness and contrast of fluorescence images enhanced for visualization. This algorithm can be utilized for any mathematical operation with capillary fluorescence intensities.

[0063] FIG. 12A-12B depicts aspects of summation-based screening of 5’ UTR library. (12A) Representative fluorescence signatures of sfGFP, mRFP1, and concurrent excitation of sfGFP and mRFP1 are shown. Snapshots were acquired at 300 ms exposure. Fluorescence intensities of three capillaries across the images are quantified in the table. The slight deviation from the theoretical sum is a result of protein photobleaching during imaging. (12B) Raw intensity vs. frequency maps for sfGFP+mRFP1 fluorescence of the 5’ UTR library.

[0064] FIG. 13A-13C depict representative images from the ratio experiment. Columns i and ii depict as-acquired sfGFP and mRFP1 images, and column iii depicts brightness-enhanced mRFP1 images for visualization. The ratio of sfGFP to mRFP1 intensities of encircled capillaries in 13A, 13B, and 13C are 19.95, 15.76, and 12.43, respectively.

[0065] FIG. 14A-14D depict exemplary data showing a comparison of translation initiation rates. (14A, 14B) Translation initiation rates of sfGFP and mRFP1 rates of RBS variants identified in the binning experiments. (14C) Distributions of ratios of translation initiation rates for RBS variants of sfGFP to mRFP1 identified in the mRFP1 binning experiment. (14D) mRNA half-lives and R2of fit for the decay constant are plotted for the arithmetic-based screening experiments.

[0066] FIG. 15A-15E depict exemplary data showing fluorescence signatures of SPs and viability. (15A) A representative bright-field image shows a minority of capillaries that exhibit significant growth (which appear as dark spots) while most capillaries seem empty. Upon closer inspection, cells were noticed in seemingly empty capillaries, which were either nonviable or experienced restrained growth, as some are labelled in the zoomed-in image. Fluorescence snapshots of cement w / o IPTG (15B), w / IPTG at t = 0 (15 C), and w / IPTG at t = 12 hr (15D) after 48-hour incubation period are depicted. Also, a fluorescence snapshot of reflectin w / o IPTG is shown (15E). Scale bars = 100 pm.Attorney Docket No. 0073605-001135

[0067] FIG 16A-16D depict exemplary data showing fluorescence intensity distributions for SPs. Fluorescence intensity maps for cement at 24-hour (16A) and 48-hour (16B) incubation periods are shown. Inset in (16A) shows capillary intensity for each induction state. It can be noticed that the distribution gets stretched-out with longer incubation. Top 500 capillaries with the highest intensities were utilized for analysis. Similarly, fluorescence intensity maps for reflectin are shown at 24-hour (16C) and 48-hour (16D) incubation periods. Top 500 capillaries with the highest intensities were utilized for analysis.

[0068] FIG. 17A-17H depict exemplary data from bright-field absorbance measurements. (17A, 17B) Bright-field images of cement w / o IPTG and w / IPTG at t = 12 hr show lower contrast as compared to other arrays. Bright-field absorbance scatter plots for cement (17A-17F) and reflectin (17G-17H) after 48-hour incubation are shown. The absorbance values are relative to the variable contrast offered by arrays and interference from agarose gel, which introduces errors in baseline subtraction. However, for a comparative analysis, it is safe to assume that the maximum absorbance across all arrays should be similar due to the growth medium being limited.

[0069] FIG. 18A-18C depict exemplary data pertaining to transmission characteristics of filters. The transmission characteristics of three-band (18A, 18B) and four-band filter (18C) cube sets are plotted in solid lines. The plots are overlaid with excitation and emission spectra (in dashed lines) of respective fluorescent protein measured, namely, mRFP1 (18A), sfGFP (18B), and mCherry (18C). Transmission data were obtained from Chroma, and intensity data were obtained from FBbase.

[0070] FIG. 19A-19B depict exemplary data based on bright-field absorbance calculation. (19A) Raw bright-field intensities of capillaries are plotted. The capillaries with lower intensities depict high cell count. (19B) Bright-field intensities were converted into absorbance levels with baseline correction centered around the mode of each data set. Absorbance trends appear inverted compared to bright-field intensity. In all the curves here, sorted lists of capillaries were utilized, unlike scatter plots where data points were plotted as is.DETAILED DESCRIPTION

[0071] The following description sets forth various examples along with specific details to provide a thorough understanding of claimed subject matter. It will be understood by those skilled in the art, however, that claimed subject matter may be practiced without one or more of the specific details disclosed herein. Further, in some circumstances, well-known methods, procedures, systems, and / or components have not been described in detail in order to avoid unnecessarily obscuring claimed subject matter. The illustrative embodiments described in the detailed description andAttorney Docket No. 0073605-001135 claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented here. It will be readily understood that the aspects of the present disclosure, as generally described herein, can be arranged, substituted, combined, and designed in a wide variety of different configurations, all of which are explicitly contemplated and make part of this disclosure.

[0072] The present disclosure provides a microcapillary array-based screening platform for high-throughput, multiplex screening of miniature cell cultures through fluorescent reporters. The present disclosure also provides a novel technical solution for population-based screening, a key screening modality desired for biomanufacturing, wherein miniature cell cultures of individual clones can be grown and analyzed in parallel. The clone recovery mechanism described herein offers superior (e.g., 100x) enrichment ratios compared to traditional techniques for establishing phenotype-to-genotype linkages.

[0073] Aspect of the present disclosure can be used to delineate effects of key plasmid design features, such as without limitation, promoters, 5’ untranslated regions (5’ UTR), and amino acid sequences, on protein titer. Aspects of the present disclosure can be used to identify, from a large collection of promoters with widely varying transcription rates, a small set of promoters that increase or maximize protein titer. Aspects of the present disclosure can be used to establish correlations between mRNA half-lives, controlled by 5’ UTRs, with a level of protein expression. Aspects of the present disclosure can be used to render relative analyses of multiple ribosome binding sites in operons with multi-reporter (e g., dual -reporter) imaging. Aspects of the present disclosure can be used to further elucidate the effect of structural protein hydrophobicity scores on their expression and cell growth profdes. Aspects of the present disclosure further demonstrate population binning, dual-reporter operon screening, chemical perturbation, and cell growth estimation through bright-field absorbance measurements with the screening platform described herein.

[0074] The screening platform described herein increases the screening rates by a factor of 102-103with respect to plate readers, even for analysis of single variants. The screening platform described in the present disclosure has the potential to aid in the discovery of novel protein sequences with greater biomanufacturability while conserving functionality, especially when complemented with mutagenesis and directed evolution.Screening Platform for Analyzing Gene Expression and / or Evaluating Protein Production

[0075] An objective of the present disclosure is to provide a high-throughput screening platform for analyzing gene expression and / or evaluating protein production. As used herein, a “screeningAttorney Docket No. 0073605-001135 platform” is an integrated apparatus or system configured to examine a plurality of samples, often a large number or collection of samples, and identify one or more samples of interest therefrom upon detection of certain attribute(s) or trait(s).Array of Microcapillary Tubes

[0076] The screening platform includes an array of microcapillary tubes. As used herein, an “array” is a geometrical arrangement and / or topology among a plurality of elements, typically within a substantially flat plane (i.e., a plane that is apparently flat to a naked eye). An array can be deployed in any means or layout deemed suitable or relevant by a person or ordinary skill in the art, in view of the entirety of the present disclosure, such as without limitation as a straight or curved line, zig zag, circle, ellipse, or polygon, or preferably in an oblique, rectangular / centered rectangular, square, or hexagonal lattice.

[0077] As used herein, a “microcapillary tube”, “capillary tube”, “capillary”, or “microcapillary” is a device having a tubular shape and a hollow cavity disposed inside, typically characterized by an inner diameter on the scale of a micron. A microcapillary tube typically has a width, diameter, or lateral dimension between 1 pm and several mm (e.g., in inner diameter). The length or longitudinal dimension of a microcapillary tube is not particularly limited, and a microcapillary tube can be as long as permitted by the mechanical strength of its constituent(s) and / or the instrumental capability for its preparation. As used herein, a “longitudinal” direction of a microcapillary tube is the direction that extending from one terminal of the microtube to the other, whereas a “lateral” direction is the direction that is transverse (i.e., perpendicular to) the longitudinal direction. Accordingly, a “longitudinal dimension” of a microcapillary tube is the size of the microtube along its longitudinal direction, and a “lateral dimension” of a microcapillary tube is the size of the microtube along its lateral direction. In some embodiments, the lateral dimension can be an inner diameter of the body of the microcapillary tube.

[0078] The length of the microcapillary tube can extend between its first end and its second end along a length of the body of the microcapillary tube. In some embodiments, both the first end and the second end have an opening; in other words, the microcapillary tube is configured as an open-faced capillary tube. The openings of the first and second ends can be in fluid communication with an inner channel having an inner diameter defined by the elongated body of the microcapillary tube. The inner channel can extend along the length of the body and also have an inner lateral dimension extending between opposed inner sides of the body of microtube. The inner lateral dimension can beAttorney Docket No. 0073605-001135 an inner diameter. The cross-sectional area of the microtube is not particularly limited and can be circular, oval / elliptical, or polygonal.

[0079] In some embodiments, the microcapillary tube has an inner diameter of at least 1 pm and no greater than 2 mm, preferably at least 1 pm and no greater than 100 pm, more preferably at least 5 pm and no greater than 50 pm. It should be noted that the inner diameter of the microcapillary tube can be equal to any value(s) within any of these respective ranges, including the endpoints of these ranges.

[0080] In some embodiments, the microcapillary tube has a length or longitudinal dimension of at least 10 pm, preferably at least 100 pm, more preferably at least 1 mm, more preferably at least 2 mm, further preferably at least 5 mm. It should be noted that the length or longitudinal dimension of the microcapillary tube is not particularly limited, in accordance with details described above in this disclosure.

[0081] The microcapillary tube described herein can be constructed using any type of material deemed suitable or relevant by a person of ordinary skill in the art, upon reviewing the entirety of the present disclosure. Nonlimiting examples of such materials include fused silica, glass (e.g., borosilicate glass), and plastics (e.g., polycarbonates (PC); fluoropolymers including fluorinated ethylene propylene (FEP), perfluoroalkoxy alkane (PF A), polytetrafluoroethylene (PTFE), ethylene tetrafluoroethylene (ETFE), and polyvinylidene fluoride (PVDF); polyether ether ketone (PEEK); cyclic olefin polymers / copolymers (COP / COC); polyetherimide (PEI); nylon or polyamides; poly(methyl methacrylate) (PMMA); and high-density polyethylene (HDPE)), among others.

[0082] The number or packing density of the microcapillary tubes is not particularly limited and can be adjusted based on specific use cases. In some embodiments of the screening platform, the array of microcapillary tubes includes an array of at least 105microcapillary tubes.Cell Culture

[0083] The miniature culture array system described herein supports parallel growth of individual cell clones in small-volume cultures, enabling population-based screening, in accordance with details described throughout the present disclosure. Specifically, each microcapillary tube within the array of microcapillary tubes is configured to contain a volume of cell culture (i.e., each capillary tube is configured to contain miniature cell culture). In some embodiments, an open-faced microcapillary tube is used to retain cells suspended in growth media due to surface tension.

[0084] The type of cells or cell cultures described herein is not particularly limited and can include any type of cell or cell culture deemed suitable by a person of ordinary skill in the art, uponAttorney Docket No. 0073605-001135 reviewing the entirety of the present disclosure. In some embodiments, the cell culture includes one or more immortalized cell lines, such as without limitation human cervical cancer cell lines including HeLa, CaSki, and SiHa; Human Embryonic Kidney 293 cells (HEK 293); Chinese hamster ovary cells (CHO); human breast cancer cells including MCF-7; rat pheochromocytoma cells including PC-12; etc. In some embodiments, the cell culture includes one or more types of stem cells, such as without limitation one or more of embryonic stem cells (ES), induced pluripotent stem cells (iPS), and mesenchymal stem cells (MSCs). In some embodiments, the cell culture includes a bacterial culture, such as without limitation, Escherichia coli (E. coli), Saccharomyces cerevisiae (S. cerevisiae or Baker’s Yeast), Bacillus subtilis (B. subtilis), Agrobacterium tumefaciens (A. tumefaciens), Lactobacillus, Geobacter, and / or Shewanella. In some embodiments, the cell culture includes a fungal culture, such as without limitation, Penicillium, Aspergillus, etc. In some embodiments, the cell culture includes a microalgal culture.

[0085] The cell culture is maintained at a suitable condition. As used herein, a “suitable condition” is an environmental condition or factor suitable for the growth and / or replication of a cell. In some cases, a suitable condition can vary from one type of cell to another. A suitable condition can include without limitation a suitable temperature / temperature range, a suitable pressure / pressure range, a suitable pH or pH range, a suitable ionic strength / range of ionic strength, a suitable osmotic pressure or osmolarity / range of osmotic pressure or osmolarity, a suitable concentration / concentration range of one or more nutrients, a suitable level of metabolic waste / metabolites, among others. A person of ordinary skill in the art, upon reviewing the entirety of the present disclosure, will be able to identify suitable conditions specific to one or more cells described herein.Photoactive Reporter

[0086] The cell culture of at least one microcapillary tube of the array of microcapillary tubes is configured to express at least a photoactive reporter. As used herein, a “photoactive reporter” is a chemical species, typically a protein, that is produced concurrently with a chemical species of interest and emits photons upon photoexcitation, thereby signaling or indicating the production or presence of the chemical species of interest.

[0087] In some embodiments of the screening platform, the at least a photoactive reporter includes at least a phosphorescent reporter.

[0088] In some embodiments of the screening platform, the at least a photoactive reporter includes at least a fluorescent reporter. In some embodiments, the at least a fluorescent reporterAttorney Docket No. 0073605-001135 functions as a fluorescent biosensor for in vivo measurements, with substantially no toxicity to the cell culture, thereby reducing or eliminating the use of fluorescent tags or invasive protocols.

[0089] The type of fluorescent reporter described herein is not particularly limited. In some embodiments, the fluorescent reporter includes a fluorescent protein that emits photons in the visible range of electromagnetic radiation, such as without limitation a green fluorescent protein (GFP), a yellow fluorescent protein (YFP), a blue fluorescent protein (BFP), a cyan fluorescent protein (CFP), a red fluorescent protein (RFP), among others. Specific, nonlimiting examples of such fluorescent proteins include mCherry, tdTomato, mApple, mOrange2, Citrine, mNeonGreen, sfGFP, mRFP1, etc. In some embodiments, the fluorescent reporter has an emission spectrum having a peak wavelength from 405 nm to 635 nm, covering DAPI through Cy. It should be noted that the peak wavelength of the emission spectrum described herein can be equal to any value(s) within this range, including its endpoints.

[0090] In some embodiments of the screening platform, the at least a fluorescent reporter includes a plurality of fluorescent reporters. As a nonlimiting example, the at least a fluorescent reporter includes a combination of sfGFP and mRFP1.Optical Module

[0091] The screening platform further includes an optical module including a light source configured to illuminate the volume of cell culture, thereby producing an optical emission; and a detector configured to detect the optical emission.

[0092] In some embodiments, the screening platform described herein can support fluorophores across the visible spectrum without requiring optical modifications and excitation-wavelength switching during experiments, thereby enabling concurrent screening of multiple fluorophores. The versatile image acquisition methods disclosed herein can not only facilitate the screening of singlereporter plasmids but also the comparative examination of multiple fluorophores with operons, in accordance with details described throughout the present disclosure.

[0093] In some embodiments of the screening platform, the light source of the optical module includes a multiband illumination system across a wavelength range of at least 300 nm and no greater than 800 nm, preferably from at least 380 nm to no greater than 760 nm. It should be noted that the wavelength of the light source described herein can be equal to any value(s) within any of these respective ranges, including the endpoints of these ranges.

[0094] In some embodiments of the screening platform, the optical module is configured to perform a bright-field absorbance measurement.Attorney Docket No. 0073605-001135

[0095] In some embodiments, the optical module includes, or is coupled or integrated with, a microscope, such as an inverted fluorescence microscope. As a nonlimiting example, the inverted fluorescence microscope can be equipped with a broad-spectrum fluorescence illumination system, filter cubes (FC), digital camera, motorized stage, and focus drive.Control Unit

[0096] The screening platform further includes a control unit communicatively connected to the optical module. As used herein, “communicatively connected” means connected by way of a connection, attachment, or linkage between two or more relata which allows for reception and / or transmittance of information therebetween. As a nonlimiting example, such connection can be wired or wireless, direct or indirect, and / or between two or more components, circuits, devices, systems, and / or the like, which allows for reception and / or transmittance of data and / or signal(s) therebetween.

[0097] In some embodiments, the control unit can include, be included in, or otherwise be communicatively connected to a computing device. The computing device described herein can include any type of computing device deemed suitable or relevant by a person of ordinary skill in the art, upon reviewing the entirety of the present disclosure. The computing device can include any analog or digital control circuit, such as without limitation, an operational amplifier circuit, a combinational logic circuit, a sequential logic circuit, an application-specific integrated circuit (ASIC), a field programmable gate arrays (FPGA), or the like. The computing device includes a processor and a memory communicatively connected to the processor, wherein the memory contains instructions configuring the processor to perform one or more processing steps, in accordance with details described throughout the present disclosure. Nonlimiting examples of potentially suitable computing devices include a microcontroller, a microprocessor, a digital signal processor, and / or a system on a chip. The computing device can include a single computing device operating independently, or two or more computing device operating in concert, in parallel, sequentially, or the like.

[0098] The control unit is configured to quantify a level of gene expression or protein production of the cell culture as a function of the optical emission. Such quantification can be performed at a fixed wavelength (e.g., by passing the optical emission through a narrow-band filter and measuring the fluorescence intensity at a predetermined wavelength) or over at least a portion of its full emission spectrum (e g., by integrating the emission spectrum over a wavelength range).Attorney Docket No. 0073605-001135

[0099] In some embodiments, the control unit is further configured to perform population binning of cell clones as a function of the level of gene expression or protein production; such clonal binning feature helps reveal biological correlations between gene expression and plasmid design rules.

[0100] The control unit is further configured to identify a genotype of the cell culture as a function of the quantification and, optionally, the population binning, thereby providing populationbased screening of the cell culture. In some embodiments, the screening platform includes, is communicatively connected to, or otherwise performs one or more functions as a data acquisition and analysis system, in accordance with details described throughout the present disclosure.Chemical Perturbation Module

[0101] In some embodiments, the screening platform further includes a chemical perturbation module communicatively connected to the control unit and configured to apply a controlled chemical stimulus to the cell culture, wherein the control unit is further configured to evaluate an effect of the controlled chemical stimulus on the gene expression or protein production in the cell culture. Such effect can include without limitation an environmental effect or the impact of a specific growth condition.

[0102] In some embodiments of the screening platform, applying the controlled chemical stimulus includes applying one or more inducers, such as without limitation, isopropyl β-D-1-thiogalactopyranoside (IPTG). As used herein, an “inducer” is a molecule that triggers or activates gene expression, often by binding to and inactivating a repressor protein and allowing RNA polymerase to transcribe genes, effectively “turning on” an operon. A person of ordinary skill in the art, upon reviewing the entirety of the present disclosure, would be able to identify suitable types or dosages of the controlled chemical stimulus, such as the concentration of one or more inducers, to be implemented by the chemical perturbation module.Extraction of Cell Culture

[0103] In some embodiments, the screening platform further includes one or more extraction slips each configured to collect the volume of cell culture extracted from the one or more microcapillaries of interest. In some embodiments, the extraction slip includes a micro cover glass. In some embodiments, the extraction slip is held at temperature of 28-30 °C to reduce or prevent condensation.

[0104] In some embodiments of the screening platform, the volume of cell culture contains one or more magnetic beads and is extracted using a focused pulse-laser beam. The focused pulsed-laserAttorney Docket No. 0073605-001135 beam can superheat the magnetic beads (e g., under 20 ms) settled at the bottom meniscus of the microcapillaries, causing cavitation that disrupts the meniscus and resulting in quick release of the cell culture.

[0105] In some embodiments of the screening platform, one or more plasmids of the extracted volume of cell culture are further amplified and sequenced using polymerase chain reaction (PCR). A person of ordinary skill in the art, upon reviewing the entirety of the present disclosure, would be able to identify and select suitable parameters for performing such amplification and sequencing.

[0106] As used herein, a “plasmid” is a circular, double-stranded DNA molecule. Plasmids are distinct from a cell’s chromosomal DNA and are capable of autonomous replication. Plasmids can be used as vectors for insertion, expression, and propagation of foreign genes within a cell or host organism. Such vectors can include specific sequences for an origin of replication, selectable markers, and cloning sites, enabling manipulation and study of genetic material for applications in research, biotechnology, and therapeutic development. As a nonlimiting example, the plasmid can include a pFTVl plasmid.Multi-Reporter / Dual Reporter Operon Screening

[0107] As noted above, in some embodiments, the screening platform implements a plurality of fluorescent reporters, such as without limitation, a combination of sfGFP and mRFP1. Accordingly, in such embodiments, the control unit is further configured to perform multi-reporter operon screening, such as without limitation dual-reporter operon screening, as a function of the optical emission. In some embodiments of the screening platform, performing the multi -reporter operon screening includes performing a comparative analysis of ribosome binding sites and / or operon dynamics. In some embodiments, the screening platform includes dual-reporter operon screening functionality.

[0108] As used herein, an “operon” is a functioning unit of DNA containing a cluster of genes under the control of a single promoter. It is commonly found in prokaryotes such as bacteria. These genes are transcribed together into a single messenger RNA strand and typically encode proteins that work together in a specific biological pathway. An operon can include a regulatory element, such as an operator, where an activator or repressor protein can bind to increase or inhibit transcription. An operon can include one or more regulatory genes that encode one or more such activator or repressor proteins. A nonlimiting example of operon includes the lac operon, which regulates lactose metabolism in E. coli.Attorney Docket No. 0073605-001135

[0109] As used herein, a “promoter” is a DNA sequence where an RNA polymerase binds to initiate transcription. As used herein, an “operator” is a DNA segment that regulates the transcription of adjacent genes.Phenotype-to-Genotype Linkage and Transcriptional Activity

[0110] The screening platforms described herein can establish high-efficiency phenotype-to-genotype linkages for the cell cultures by isolating and identifying specific genotypes based on observed phenotypic traits. Specifically, in some embodiments of the screening platform, identifying the genotype of the cell culture includes isolating a phenotypic trait of the cell culture as a function of the level of gene expression or protein production, and identifying the genotype of the cell culture as a function of the phenotypic trait. This process can be referred to as the “clone recovery mechanism” throughout the present disclosure.

[0111] In some embodiments of the screening platform, the control unit is further configured to determine a level of transcriptional activity (e.g., transcription rate, translation initiation rate (TIR), among others) as a function of the level of protein production (i.e., the protein titer).5’ Untranslated Region (5’ UTR) and mRNA Stability

[0112] In some embodiments of the screening platform, the control unit is further configured to measure a half-life of mRNA as a function of the optical emission and correlate the half-life and the level of protein production.

[0113] The half-life of mRNA can be a function of the identity of 5’ UTR, in accordance with details described throughout the present disclosure. For example, 5’ UTR can directly influence the decay of transcribed mRNAs within cells, thus controlling the extent of protein translation. As used herein, a “5’ untranslated region” or “5’ UTR)” is the region of an mRNA that is directly upstream from the initiation codon and regulates one or more aspects of the transcriptional and / or translational process of the mRNA. The 5’ UTR begins at the transcription start site and ends one nucleotide before the initiation sequence (typically AUG) of the coding region.Cell Growth Profile

[0114] In some embodiments of the screening platform, the control unit is further configured to generate a cell growth profile of the cell culture (e.g., via bright-field absorbance measurements described herein) and determine one or more structural properties of a protein as a function of or in relation to the cell growth profile. In some embodiments of the screening platform, the one or more structural properties of the protein includes hydrophobicity (e.g., the hydrophobicity score) of the protein. Proteins or amino acid sequences with reduced hydrophobicity can potentially enhanceAttorney Docket No. 0073605-001135 protein solubility and efficacy of protein expression. This feature can be especially useful for the quantification of stochastic variability in gene expression and growth with toxic coding sequences (CDSs) to control noise by biological factors.Output

[0115] In some embodiments of the screening platform, the control unit is further configured to generate an output as a function of the level of gene expression or protein production.

[0116] In some embodiments of the screening platform, the output includes one or more of a suitable plasmid construct for protein production, a suitable promoter for protein production a suitable 5’ UTR for protein production, and a suitable amino acid sequence for protein production.

[0117] In some embodiments of the screening platform, the control unit is further configured to display the output using an output device or output interface. In some embodiments, the output device or output interface is configured to generate quantitative datasets identifying suitable or optimal plasmid constructs, including without limitation one or more suitable combinations of promoter(s), 5’ UTR(s), and amino acid sequence(s) that increase or maximize protein production.

[0118] In some embodiments, the screening platform is configured to screen one or more plasmid design libraries. Each plasmid design library can include a plurality of promoters with varying transcription rates, a plurality of 5’ UTRs configured to modulate mRNA stability and or having various translation initiation rates, and / or a plurality of amino acid sequence variants adapted or optimized to reduce hydrophobicity and enhance protein solubility and / or protein expression.

[0119] In some embodiments, the screening platform further includes an image-processing module communicatively connected to the control unit, wherein generating the output includes generating an intensity map, using the image-processing module, as a function of the optical emission. In some embodiments of the screening platform, the output includes one or more coordinates or indices of one or more microcapillaries of interest.Enrichment Ratio

[0120] In some embodiments, wherein the screening platform provides a clone recovery mechanism having an enrichment ratio of up to 105: 1. As used herein, an “enrichment ratio” is a fold-increase in concentration and / or purity of a target cell population compared to an initial heterogeneous population.Method For Analyzing Gene Expression or Evaluating Protein Production

[0121] Another objective of the present disclosure is to provide a method for analyzing gene expression or evaluating protein production.Attorney Docket No. 0073605-001135

[0122] The method includes applying a plurality of cell cultures to an array of microcapillary tubes each configured to contain a volume of cell culture. In some embodiments, heterogenous mixtures of clones can be loaded directly onto the arrays without needing prior selection or isolation. The cell culture in at least one microcapillary tube of the array of microcapillary tubes is configured to express at least a photoactive reporter, in accordance with details described throughout the present disclosure.

[0123] The method further includes measuring an optical emission of the plurality of cell cultures using an optical module having a light source and a detector. This aspect of the method can be implemented in accordance with details described throughout the present disclosure.

[0124] The method further includes quantifying, using a control unit communicatively connected to the optical module, a level of gene expression or protein production in the plurality of cell cultures, as a function of the optical emission. This aspect of the method can be implemented in accordance with details described throughout the present disclosure.

[0125] In some embodiments, the method further includes performing population binning, using the control unit, as a function of the level of gene expression or protein production. This aspect of the method can be implemented in accordance with details described throughout the present disclosure.

[0126] The method further includes identifying a genotype of at least one cell culture of the plurality of cell cultures, using the control unit, as a function of the quantification and, optionally, the population binning, thereby providing population-based screening of the cell culture. This aspect of the method can be implemented in accordance with details described throughout the present disclosure.

[0127] In some embodiments of the method, the array of microcapillary tubes includes an array of at least 105microcapillary tubes, in accordance with details described throughout the present disclosure.

[0128] In some embodiments of the method, the light source includes a multiband illumination system across a wavelength range of at least 300 nm and no greater than 800 nm, in accordance with details described throughout the present disclosure.

[0129] In some embodiments, the method further includes measuring a bright-field absorbance of the volume of cell culture using the optical module, in accordance with details described throughout the present disclosure.Attorney Docket No. 0073605-001135

[0130] In some embodiments of the method, the at least a photoactive reporter includes at least a fluorescent reporter, in accordance with details described throughout the present disclosure.

[0131] In some embodiments of the method, the at least a fluorescent reporter includes a plurality of fluorescent reporters. Accordingly, in some embodiments, the method further includes performing multi-reporter operon screening, such as without limitation dual-reporter operon screening, using the control unit, as a function of the optical emission, in accordance with details described throughout the present disclosure.

[0132] In some embodiments of the method, performing the multi-reporter operon screening includes performing a comparative analysis of ribosome binding sites and / or operon dynamics, in accordance with details described throughout the present disclosure.

[0133] In some embodiments of the method, identifying the genotype of the cell culture includes isolating a phenotypic trait of the cell culture as a function of the level of gene expression or protein production, and identifying the genotype of the cell culture as a function of the phenotypic trait, in accordance with details described throughout the present disclosure.

[0134] In some embodiments, the method further includes determining a level of transcriptional activity as a function of the level of protein production, in accordance with details described throughout the present disclosure.

[0135] In some embodiments, the method further includes applying a controlled chemical stimulus to the cell culture, using a chemical perturbation module communicatively connected to the control unit, and evaluating an effect of the controlled chemical stimulus on the gene expression or protein production in the cell culture, in accordance with details described throughout the present disclosure.

[0136] In some embodiments of the method, applying the controlled chemical stimulus includes applying one or more inducers, in accordance with details described throughout the present disclosure.

[0137] In some embodiments, the method further includes measuring the half-life of mRNA, using the control unit, as a function of the optical emission, and correlating the half-life and the level of protein production using the control unit, in accordance with details described throughout the present disclosure.

[0138] In some embodiments, the method further includes generating a cell growth profile of the cell culture using the control unit and determining, using the control unit, one or more structural properties of a protein as a function of the cell growth profile. In some embodiments of the method,Attorney Docket No. 0073605-001135 the one or more structural properties of the protein includes hydrophobicity of the protein, in accordance with details described throughout the present disclosure.

[0139] In some embodiments, the method further includes generating an output as a function of the level of gene expression or protein production, in accordance with details described throughout the present disclosure.

[0140] In some embodiments of the method, the output includes one or more of a suitable plasmid construct for protein production, a suitable promoter for protein production, a suitable 5’ untranslated region for protein production, and a suitable amino acid sequence for protein production, in accordance with details described throughout the present disclosure.

[0141] In some embodiments, the method further includes displaying the output using an output device communicatively connected to the control unit, in accordance with details described throughout the present disclosure.

[0142] In some embodiments of the method, generating the output includes generating an intensity map, using an image-processing module communicatively connected to the control unit, as a function of the optical emission, in accordance with details described throughout the present disclosure.

[0143] In some embodiments of the method, the output includes one or more coordinates or indices of one or more microcapillaries of interest, in accordance with details described throughout the present disclosure.

[0144] In some embodiments, the method further includes collecting the volume of cell culture extracted from the one or more microcapillaries of interest using one or more extraction slips, in accordance with details described throughout the present disclosure.

[0145] In some embodiments of the method, the volume of cell culture contains one or more magnetic beads and is extracted using a focused pulse-laser beam, in accordance with details described throughout the present disclosure.

[0146] In some embodiments, the method further includes amplifying and sequencing one or more plasmids of the extracted volume of cell culture using polymerase chain reaction, in accordance with details described throughout the present disclosure.

[0147] In some embodiments, the method provides a clone recovery mechanism having an enrichment ratio of up to 105: 1, in accordance with details described throughout the present disclosure.Attorney Docket No. 0073605-001135 EXAMPLE

[0148] Cell sorting and screening techniques are useful for biological research, bioengineering, and medicine as they facilitate the isolation of specific clones from heterogeneous populations and enable the analysis of numerous cell phenotypes in studies related to drug and immune responses (Skardal, A.; Shupe, T.; Atala, A. Organoid-on-a-chip and body-on-a-chip systems for drug screening and disease modeling. Drug Discovery Today 2016, 21, 1399-1411; Broome, S.; Gilbert, W. Immunological screening method to detect specific translation products. Proc. Natl. Acad. Sci. U. S. A. 1978, 75, 2746-2749; Sharma, S.; Rao, A. RNAi screening: tips and techniques. Nat.Immunol. 2009, 10, 799-804), disease progression (Carey, A.; Edwards, D. K.; Eide, C. A.; et al. Identification of Interleukin- 1 by Functional Screening as a Key Mediator of Cellular Expansion and Disease Progression in Acute Myeloid Leukemia. Cell Rep. 2017, 18, 3204-3218), enzyme and biomarker engineering (Longwell, C. K.; Labanieh, L.; Cochran, J. R. High-throughput screening technologies for enzyme engineering. Curr. Opin. Biotechnol. 2017, 48, 196-202) and cell differentiation (Seo, J.; Shin, J. Y.; Leijten, J.; et al. High-throughput approaches for screening and analysis of cell behaviors. Biomaterials 2018, 153, 85-101). Beyond biology and medicine, screening techniques have significant potential in the biomanufacturing and bioprocessing industries, such as biopharmaceuticals, biofuels, food and agriculture, and bioremediation, which focus on protein-titer optimization (Ward, O. P. Bioprocessing; Springer US: Boston, MA, 1991; Mandenius, C.; Brundin, A. Bioprocess optimization using design-of-experiments methodology. Biotechnol. Prog. 2008, 24, 1191-1203; Overton, T. W. Recombinant protein production in bacterial hosts. Drug Discovery Today 2014, 19, 590-601). The key screening modality sought for biomanufacturing is population-based screening, where miniature cell cultures of individual clones can be grown and analyzed simultaneously. While significant strides in single-cell screening have been achieved using microfluidic platforms (Witek, M. A.; Freed, I. M.; Soper, S. A. Cell Separations and Sorting. Anal. Chem. 2020, 92, 105-131), progress in population-based screening remains inadequate. Microwell plate readers (Packer, M. S.; Liu, D. R. Methods for the directed evolution of proteins. Nat. Rev. Genet. 2015, 16, 379-394; Dorr, M.; Fibinger, M. P.; Last, D.; et al. Fully automatized high-throughput enzyme library screening using a robotic platform. Biotechnol. Bioeng. 2016, 113, 1421-1432; Lafferty, M.; Dycaico, M. J. GigaMatrixTM: An Ultra High-Throughput Tool for Accessing Biodiversity. JATA J. Assoc. Lab. Autom. 2004, 9, 200-208; Auld, D. S.; Coassin, P. A.; Coussens, N. P.et al. Microplate Selection and Recommended Practices in High-throughput Screening and Quantitative Biology. In Assay Guidance Manual,' Markossian, S.;Attorney Docket No. 0073605-001135 Grossman, A.; Arkin, M.et al., Eds.; Eli Lilly & Company and the National Center for Advancing Translational Sciences: Bethesda (MD), 2020) and multi-parallel bioreactors (Bareither, R.; Pollard, D. A review of advanced small-scale parallel bioreactor technology for accelerated process development: Current state and future need. Biotechnol. Prog. 2011, 27, 2-14; Bertaux, F.; Sosa-Carrillo, S.; Gross, V.; et al. Enhancing bioreactor arrays for automated measurements and reactive control with ReacSight. Nat. Commun. 2022, 13, No. 3363) serve as essential tools for screening cell cultures in laboratories globally. However, these platforms are not ideal due to inefficient methods of analyzing large clonal libraries, insufficient throughput, or complex handling procedures. In this context, we concentrate on a population-based screening platform and illustrate its potential in biomanufacturing and gene expression studies.

[0149] Cost-effective protocols for the rapid synthesis of large plasmid libraries are being developed, and techniques that can complement these libraries with multiplex screening at similar scales are essential. Cloning strategies such as mutagenesis and combinatorial fragment assembly facilitate the construction of libraries that cover vast design spaces, which can inform protein structure-expression relationships and aid in discovering new protein sequences (e.g., those with similar properties but improved expression) through screening. Furthermore, algorithms for the massively parallel design of plasmid elements (e.g., promoters, ribosome binding sites, RBS) have opened new avenues for extensive optimization of recombinant gene expression. While predicting expression and cell growth remains complex due to astronomical cellular interdependencies (Snoeck, S.; Guidi, C.; De Mey, M. “Metabolic burden” explained: stress symptoms and its related responses induced by (over)expression of (heterologous) proteins in Escherichia coli. Microb. Cell Fact. 2024, 23, No. 96; Carneiro, S.; Ferreira, E. C.; Rocha, I. Metabolic responses to recombinant bioprocesses in Escherichia coli. J. Biotechnol. 2013, 164, 396-408; Santos-Navarro, F. N.; Vignoni, A.; Boada, Y.; Pico, J. RBS and Promoter Strengths Determine the Cell-Growth-Dependent Protein Mass Fractions and Their Optimal Synthesis Rates. ACS Synth. Biol. 2021, 10, 3290-3303) and until machine learning models can fully support these predictions (Martiny, H.-M.; Armenteros, J. J. A.; Johansen, A. R.; Salomon, J.; Nielsen, H. Deep protein representations enable recombinant protein expression prediction. Comput. Biol. Chem. 2021, 95, No. 107596; Habibi, N.; Mohd Hashim, S. Z.; Norouzi, A.; Samian, M. R. A review of machine learning methods to predict the solubility of overexpressed recombinant proteins in Escherichia coli. BMC Bioinf. 2014, 15, No. 134; Fu, H.; Liang, Y.; Zhong, X.; et al. Codon optimization with deep learning to enhance protein expression. Sci. Rep. 2020, 10, No. 17617), high-throughput population-based screening techniques are ofAttorney Docket No. 0073605-001135 substantial importance in synthetic biology (FIG. 1). Chen and Lim et al. (Chen, B.; Lim, S.;Kannan, A.; et al. High-throughput analysis and protein engineering using microcapillary arrays. Nat. Chem. Biol. 2016, 12, 76-81) developed a microcapillary array-based technique for high-throughput protein analysis and engineering. Through directed evolution of clones and successive screening rounds on the platform, the authors discovered a new antibody, a fluorescent protein biosensor, and an enzyme with enhanced resistance to inhibitors. The technique involves the spatial isolation of clones (>105) in capillaries and the observation of cell phenotypes over extended periods using fluorescent reporters. The platform facilitates the proliferation of cells into isogenic miniature cultures and supports a population-based screening method, unlike flow cytometry or flow-assisted cell sorting (FACS). The functional activity of designed clones is indicated by fluorescence intensity, and the clones demonstrating the highest functionality can be quickly recovered using noncontact laser-based methods extraction (FIG. 1C, ID).

[0150] In this study, we developed a screening platform based on microcapillary arrays and introduced relevant features related to gene expression analysis and optimization. We incorporated a clonal binning feature to reveal biological correlations between gene expression and plasmid design rules. Our platform supports fluorophores across the visible spectrum without requiring optical modifications or excitation- wavelength switching during experiments, enabling concurrent screening of multiple fluorophores. The versatile image acquisition methods not only facilitate the screening of single-reporter plasmids but also allow for the comparative examination of multiple fluorophores with operons. In addition, in-experiment chemical perturbation of cell cultures is enabled for studying expression phenotypes versus growth conditions (e g., inducers (Rosano, G. L.; Ceccarelli, E. A. Recombinant protein expression in Escherichia coir, advances and challenges. Front.Microbiol. 2014, 5, No. 172)). Alongside fluorescence imaging, we demonstrate the assessment of cell growth profiles for individual clones through bright-field absorbance measurement (Babakhanova, G.; Zimmerman, S. M.; Pierce, L. T.; et al. Quantitative, traceable determination of cell viability using absorbance microscopy. PLoS One 2022, 17, No. e0262119). This feature is useful for quantifying stochastic variability in gene expression and growth with coding sequences (CDSs) to mitigate noise from biological factors (Elowitz, M. B.; Levine, A. J.; Siggia, E. D.; Swain, P. S. Stochastic Gene Expression in a Single Cell. Science 2002, 297, 1183-1186; McAdams, H. H.; Arkin, A. Stochastic mechanisms in gene expression. Proc. Natl. Acad. Sci. U. S. A. 1997, 94, 814-819; Raj, A.; Van Oudenaarden, A. Nature, Nurture, or Chance: Stochastic Gene Expression and Its Consequences. Cell 2008, 135, 216-226) The high-throughput platform maintains the laser-basedAttorney Docket No. 0073605-001135 mechanism for clone recovery with an enrichment ratio of up to 1:105. We utilize libraries of promoters and 5’ untranslated regions (UTRs) to define the impact of transcription rates, translation initiation rates, and mRNA stability on protein titer (FIG. 1B). Additionally, by employing the operonic structure of the 5’ UTR library, we compare the effect of individual RBSs on the simultaneous expression of two fluorophores (mRFP1 and sfGFP). We introduce a fluorescent biosensor design to enable the quantification of nonfluorescent structural proteins in vivo. Finally, we investigate stochastic variability in protein expression and cell growth with two structural proteins of varying hydrophobicity under different growth conditions.Results

[0151] Microcapillary Array-Based Cell Screening. The screening methodology and the instrument are illustrated in FIG. 2. The methodology involves the spatial isolation and proliferation of cells within a lattice of micrometer-sized capillaries (>105). The open-faced capillaries maintain the cells suspended in growth media due to surface tension. The expression characteristics of constructs are monitored using fluorescent reporters, supporting both single-reporter plasmids and multi-reporter operons. The platform features a multiband illumination system accommodating fluorophores across the visible spectrum. Fluorescence signals from the capillaries are recorded and compiled automatically, and according to the screening motivation, clones exhibiting desired fluorescence characteristics are recovered to establish phenotype-to-genotype linkages.Heterogeneous mixtures of clones are loaded directly onto the arrays without prior selection or isolation. Additionally, growth characteristics are quantified through bright-field absorbance measurements across populations, particularly benefiting single-clone studies for variability.

[0152] The laser-based recovery technique precisely extracts cells from individual capillaries with each iteration (FIG. 1C). The process uses a focused pulsed-laser beam to heat magnetic beads (for less than 20 ms) that settle at the bottom meniscus of the capillaries, creating cavitation that disrupts the meniscus for rapid release (FIG. ID). After extraction, the plasmids within the extracted cells are identified. Our strategy involved PCR amplification of constructs, sequencing, and alignment with reference sequences to identify the recovered clones. We designed vector-specific primer pairs for each library and collectively amplified all constructs from each extraction slip in a single PCR. The amplified constructs were then sequenced (Oxford Nanopore) and aligned.Combined with the “Context Aware Auto-align” algorithm (De Novo DNA) for matching thousands of sequenced templates to references in seconds, this comprehensive protocol enables quick, accurate, and efficient identification of recovered clones.Attorney Docket No. 0073605-001135

[0153] To demonstrate cell growth and protein expression in an array, we cultured an isogenic E. coli colony that contained the plasmid expressing the red fluorescent protein mRFP1 (pFTVl). FIG. 2C shows the fluorescence snapshots of the array after gradual incubation for 24 h at 37 °C. Cell growth and fluorescent protein expression are evident in these images (refer to FIG. 7A for images taken at constant exposure). In FIG. 2D, the average intensity per capillary displays a nonlinear increasing pattern throughout the incubation period (raw data are illustrated in FIG. 7B), likely due to the initial exponential growth phase. Notably, the inter-capillary variability in intensity is low (i.e., low standard deviation). The homogeneity of capillaries across the array ensures consistent cell growth (Chen, B.; Lim, S.; Kannan, A.; et al. High-throughput analysis and protein engineering using microcapillary arrays. Nat. Chem. Biol. 2016, 12, 76-81).

[0154] Transcription Rate Effects on Protein Expression in Promoter Library, Gene promoters are key design components of plasmids that regulate protein expression by modulating the rate of DNA transcription into mRNA. We screened a library of 4,351 promoters in E. coli to categorize them based on the respective mRFP l fluorescent reporter titers (FIG. 3 A). The transcription rates of the promoters vary from 0.006 au to 4731.420 au (Hossain, A.; Lopez, E.; Halper, S. M.; et al.Automated design of thousands of nonrepetitive parts for engineering stable genetic systems. Nat. Biotechnol. 2020, 38, 1466-1475), quantified using DNA and RNA sequencing read counts (Hossain, A.; Lopez, E.; Halper, S. M.; et al. Automated design of thousands of nonrepetitive parts for engineering stable genetic systems. Nat. Biotechnol. 2020, 38, 1466-1475). The strength of the RBS, indicated by its translation initiation rate (TIR), remains constant, with the transcription rate being the only variable controlled in the plasmids within the library.

[0155] In FIG. 3B, the fluorescence image of a section of an array is shown after a 12-h incubation at 37 °C. The image displays varying levels of mRFP1 expression, indicated by the presence of both dim and bright capillaries. We conducted three replicates of cell growth and screening with the library, and the normalized intensity maps for all replicates are depicted in FIG.3C (see FIG. 8 for raw intensity distributions). The trend of fluorescence intensity trends confirms consistent cell growth profiles across replicates. We extracted and identified the five best-performing promoters from the top bins in the replicates, revealing low transcription rates ranging from 5 au to 15 au (FIG. 3D). Compared to several stronger promoters (Supporting Data 1), DNA-sequencing read counts for the extracted promoters were higher, while RNA-sequencing read counts were similar. This suggests that promoters associated with high mRFP1 fluorescence levels while maintaining a faster growth rate are preferentially selected. Strong promoters exhibiting elevatedAttorney Docket No. 0073605-001135 protein expression per cell may not be optimal for protein production due to metabolic burden or toxicity (Snoeck, S.; Guidi, C.; DeMey, M. “Metabolic burden” explained: stress symptoms and its related responses induced by (over)expression of (heterologous) proteins in Escherichia coli.Microb. Cell Fact. 2024, 23, No. 96; Carneiro, S.; Ferreira, E. C.; Rocha, I. Metabolic responses to recombinant bioprocesses in Escherichia coli. J. Biotechnol. 2013, 164, 396-408; Dumon- Seignovert, L.; Cariot, G.; Vuillard, L. The toxicity of recombinant proteins in Escherichia coli: a comparison of over-expression in BL21(DE3), C41(DE3), and C43(DE3). Protein Expr. Purif. 2004, 37, 203-206) Screening the best-performing promoters also serves as a rigorous test of the platform’s reproducibility. In addition to the transcription rates of the extracted promoters being similar, two of the promoters identified by replicates 1 and 3 were the same (Supporting Data 1).

[0156] Binning of clones is performed to reveal biological correlations between design and expression, with the transcription rate being the focus in this case (Shen, Y.; Chaigne-Delalande, B.; Lee, R. W. J.; Losert, W. CytoBinning: Immunological insights from multi-dimensional data. PLoS One 2018, 13, No. e0205291). We divided the normalized mRFP1 fluorescence intensity map into three logarithmically spaced bins: low, medium, and high (FIG. 3F). A representative capillary from each bin is depicted in FIG. 3E, showing pronounced inter-capillary variance in fluorescence signatures. Intra-capillary variance also exists, which is especially critical in recovering low-bin clones. While high-intensity capillaries have all their pixels above the threshold level, dim capillaries can have several pixels with intensities below this threshold (FIG. 9A, 9C, 9F). Post binarization, this latter case may result in capillaries being identified as clusters of several disjoint regions (see FIG. 9B, 9D, 9G for comparison). Since each region is treated as a distinct entity on the intensity map, small regions can cause capillaries from higher bins to be misidentified as low-bin capillaries during sorting; that is, small regions can be false low-bin regions. To address this issue, we employed a pixel-dilation strategy (detailed in Methods) to eliminate false low-bin regions, which digitally coalesces small regions with the larger ones in capillaries (FIG. 9E, 9H). The outcome is signified by a multifold reduction in the frequency of low-bin designations due to the elimination of false low-bin regions (FIG. 91). Hence, the misidentification of low-bin clones during binning was curtailed.

[0157] FIG. 3G illustrates the correlation between DNA and RNA sequencing read counts and fluorescence intensity for promoters randomly recovered from each bin. Clearly, RNA count is positively correlated with increasing fluorescence across low to high bins. Higher RNA counts are associated with a stronger fluorescence signal and, consequently, a higher protein titer. DNA countsAttorney Docket No. 0073605-001135 also show a positive correlation, indicating that cell growth is crucial for protein titer. However, no clear trend in the transcription rate of promoters was observed to account for the DNA counts (FIG.3H). The greater variance in the low and medium bin transcription rates reflects a wider range of promoters, including both strong and weak ones. This result suggests that lower titer levels can result from weak promoters (low net RNA count) or toxic, strong promoters (low DNA count). We also link these observations (FIG. 3G, 3H) to inherent variations in protein expression (over 10×) among promoters with similar transcription rates (Hossain, A.; Lopez, E.; Halper, S. M.; et al. Automated design of thousands of nonrepetitive parts for engineering stable genetic systems. Nat.Biotechnol. 2020, 38, 1466-1475) Essentially, protein titer is closely linked to the influence of the promoter on cell growth and mRNA transcription. A promoter that maximizes mRNA translation across the colony-by optimizing the mRNA available per cell and the net cell count-is favorable for achieving high protein titers. Compared to screening the best performers, the binning algorithm required only a minimal additional time (101 s) for image processing, with recovery essentially being the rate-limiting step.

[0158] CTRNA Decay Effects on Protein Expression in 5’ UTR Library. Another critical component of the protein expression workflow is the 5’ UTR of the mRNA. This region at the 5’ end (FIG. 1B) directly impacts the decay of transcribed mRNAs within cells, thereby controlling the level of protein translation (Heck, A. M.; Wilusz, J. The Interplay between the RNA Decay and Translation Machinery in Eukaryotes. Cold Spring Harbor Perspect. Biol. 2018, 10, No. a032839). We utilized a library of 62,1205’ UTRs with varying exponential mRNA decay rates (Cetnar, D. P.; Hossain, A.; Vezeau, G. E.; Salis, H. M. Predicting synthetic mRNA stability using massively parallel kinetic measurements, biophysical modeling, and machine learning. Nat. Commun. 2024, 15, No. 9601) and demonstrated the effect of mRNA half-lives (Cetnar, D. P.; Salis, H. M. Systematic Quantification of Sequence and Structural Determinants Controlling mRNA stability in Bacterial Operons. ACS Synth. Biol. 2021, 10, 318-332) on protein titer through binning. This library of operons (FIG. 4A) encodes sfGFP and mRFP1 fluorescent reporters, with the 5’ UTRs upstream of sfGFP acting as the control variable. Furthermore, using dual reporters enabled us to sort cells based on mathematical operations performed with distinct fluorescence signals to clarify the properties of RBS for each reporter.

[0159] FIG. 4B, 4C display snapshots of sfGFP and mRFP1 fluorescence from the library in an array incubated at 37 °C, respectively. Visually, these images reveal variation in the expression of both reporters and within each reporter. This variation is quantified in the normalized intensity mapsAttorney Docket No. 0073605-001135 (FIG. 4D) and raw intensity maps (FIG. 10). While the range of sfGFP and mRFP1 intensities is similar, the frequency of low-intensity regions is significantly higher for mRFP1. We speculate that the lesser impact of highly unstable 5’ UTRs on mRFP1 translation is the likely cause. We screened the library for sfGFP and mRFP1 signals separately, following the previously discussed protocol for binning. We scanned approximately 393,000 capillaries, compared to about 92,000 for the promoter library, to provide a broader representation of the larger 5’ UTR library on the array. The mRNA half-life increases significantly as sfGFP fluorescence rises from low to high bins (FIG. 4E). 5’ UTRs identified in the high bin exhibit on average over 4 times longer half-lives compared to those in the low bin, supporting sustained translation and leading to increased protein production and fluorescence intensity. Since 5’ UTRs were designed solely to regulate sfGFP translation, with mRFP1 having an independent RBS, a similar effect on mRFP1 levels was not anticipated or observed (FIG. 4F). However, there is a positive correlation between the half-lives of sfGFP and mRFP1, which can be attributed to the operonic nature of the expression system.

[0160] In addition to regulating protein expression through mRNA decay, the 5’ UTR contains the ribosome binding site (RBS) for sfGFP, whose ribosome-recruiting strength directly influences translation. Each 5’ UTR sequence yields a unique RBS with a specific translation initiation rate (TIR). The operonic nature of the plasmid necessitates relative analysis of reporter TIRs through dual reporter imaging. We screened capillaries based on fluorescence intensity ratios and the total sums of sfGFP and mRFP1. For the ratio experiment, we gathered sfGFP-mRFP1 fluorescence intensities in each capillary from discrete snapshots taken across the array. To process the sfGFP and mRFP1 signals in each capillary, the capillaries in fluorescence images of one reporter must be aligned with those of the other. Unlike image stitching, which can cause slight misalignment, discrete imaging avoids drift and ensures precise capillary tracking across the two sets of images (detailed in FIG. 11). Additionally, we used bright-field snapshots as binary masks (FIG. 4G) for intensity quantification per capillary to eliminate errors arising from intracapillary variations. While the discrete imaging protocol can be utilized for all mathematical operations, for the summation experiment, we captured the array with concurrent excitation of the reporters, making automated image stitching feasible. Both reporters were excited simultaneously using their respective illumination LEDs at the same irradiance level (illustrated in FIG. 12). The remaining steps for screening capillaries exhibiting the highest intensity sums were identical to those for identifying the best-performing promoters.Attorney Docket No. 0073605-001135

[0161] FIG. 4H illustrates the fluorescence signatures of reporters within capillaries. The platform’s ability to reveal highly differentiated cell populations indicated that 5’ UTRs led to disproportionate protein expression by operons. FIG. 41, 4K display nonlinearly decreasing profiles of intensity ratios and sums, with only a subset of clones maximizing each operation. The mRFP1 fluorescence in capillaries with high intensity ratios (sfGFP / mRFP1) was negligible (see FIG. 13 for images). The ratio experiment resulted in the preferential recovery of clones with significantly greater relative TIRs for sfGFP (FIG. 4 J) compared to those identified regardless of mRFP1 intensities (FIG. 4E). Thus, disproportionate sfGFP expression is associated with relatively stronger sfGFP RBSs compared to mRFP1 RBSs. The summation experiment, on the other hand, identified clones with higher absolute TIRs for both reporters on average, as compared to the ratio experiment (11.4 times for sfGFP and 1.7 times for mRFP1, FIG. 4L). Furthermore, the half-lives of these 5’ UTRs were also greater (by 36.7% on average, FIG. 14). These observations align with the respective screening modalities, such as disproportionate and cumulative expression. It is also worth noting that clones with extremely high sfGFP expression could potentially appear in both context experiments.

[0162] Perturbation of Structural Proteins and Variability Analysis. Producing recombinant structural protein is challenging due to issues with self-assembly, hydrophobicity, and post-translational modifications. Structural proteins often possess complex shapes that are essential for their function. The cellular machinery of a host organism used for recombinant production may not be equipped to manage these intricate folding patterns. Likewise, the hydrophobic regions of structural proteins can lead to aggregates and inclusion bodies, which may be toxic to the host organism at high titers. FIG. 5A illustrates the design of the inducible plasmid featuring a T7 promoter and a fluorescent protein biosensor. We used the mCherry fluorescent reporter to quantify nonfluorescent structural protein titers in vivo. The four-nucleotide overlap (i.e., ATGA) between the structural protein stop codon and the mCherry start codon leads to ribosome reinitiation (Tian, T.; Salis, H. M. A predictive biophysical model of translational coupling to coordinate and control protein expression in bacterial operons. Nucleic Acids Res. 2015, 43, 7137-7151), which causes the translational coupling of mCherry and coding sequences (CDS). The additional experimental approaches discussed here include chemical perturbation of cells and measurement of growth indicators to analyze stochastic variability in recombinant gene expression. This platform is suitable for in-experiment induction, as the agarose gel can be exposed to chemicals (in this case, isopropyl β-D-1-thiogalactopyranoside (IPTG) as the inducer) that will diffuse through the matrixAttorney Docket No. 0073605-001135 and into the capillaries overtime (FIG. 5B). Instead of limiting the study to a single clone, we broadened our approach by using two proteins, namely cement and reflectin (see Supporting Data 1 for amino acid sequences and CDSs), which have distinct hydrophobicity scores (FIG. 5A, calculated using models in Wimley, W. C.; White, S. H. Experimentally determined hydrophobicity scale for proteins at membrane interfaces. Nat. Struct. Mol. Biol. 1996, 3, 842-848). This property is associated with protein folding and aggregation (Beygmoradi, A.; Homaei, A.; Hemmati, R.;Fernandes, P. Recombinant protein expression: Challenges in production and folding related matters. Int. J. Biol. Macromol. 2023, 233, No. 123407).

[0163] To quantify the stochastic variability in structural protein production, we measured titer levels in cell cultures derived from single cells using a high-throughput approach. We anticipated the fluorescence signature from approximately 30% of the total capillaries; however, for structural proteins, we found that cells grew significantly in only 5.56 to 6.67% of the expected capillaries (i.e., only 500 to 600 out of roughly 9000). This indicates that most cells in the starter culture were either nonviable or slow-growing (FIG. 15 A) and would not contribute to the amplification of the titer. Additionally, the distributions of capillary fluorescence (FIG. 16) spanned the entire intensity range, indicating that the expression was highly diffuse for both protein types (for comparison, see pFTVl distribution in the inset of FIG. 7B). The fluorescence snapshots in FIG. 5C and FIG. 15 illustrate the differences in intensities of capillaries containing the two proteins. The plots in FIG. 5D display intensities (or titers) at various induction states after a 48-h incubation. The basal titers for both proteins are higher than those at the respective induced state (t = 6 h) due to the added toxicity and metabolic burden from the T7 promoter. However, compared to cement, reflectin exhibited significantly lower protein titers regardless of induction. While the reduced expression levels in reflectin can be linked to greater hydrophobicity and vice versa for cement, studying cell growth profiles across all conditions is essential to explain titer levels comprehensively.

[0164] We did not include magnetic beads in the single-clone experiments presented in FIG. 5. This removed optical interference from magnetic beads, ensuring that the degree of light transmission depended solely on cell count (FIG. 5E, 5F). Therefore, the extent of cell growth can be estimated by relatively quantifying absorbance using capillary intensity in bright-field snapshots. FIG. 5G (raw data in FIG. 17C-17H) displays the absorbance distribution per capillary above the baseline levels of empty capillaries. Regardless of induction, 95.5% more capillaries on average achieved absorbance over 0.2 au with cement than with reflectin, which clusters at lower absorbanceAttorney Docket No. 0073605-001135 levels (87% of capillaries in the 0.1-0.2 au range on average), indicating that reflectin expression significantly inhibited growth. This observation aligns with the lower mCherry fluorescence intensities associated with reflectin.

[0165] In addition to the hydrophobicity of amino acids, the timing of induction is crucial for cell growth. On average, the number of cement capillaries that achieved an absorbance of over 0.2 au without induction and with a t = 12 h induction is 42.3% higher than under early induction conditions. The variability in protein titer was confirmed through fluorescence assays of cement at these induction time points (t = 0, 6 h, and 12 h). Cement protein exhibited high net fluorescence intensity even without IPTG due to the basal expression of the T7 reporter (FIG. 5D). The intensity decreased significantly with early inductions at t = 0 and 6 h. However, with a t = 12 h induction, we reached a protein titer comparable to the noninduced state. Thus, the separation of growth and production phases was achieved with late-stage induction. This ensured that cells had ample time for growth (in the exponential growth phase) before induced expression, thereby mitigating metabolic burden with prior growth for improved titers.Discussion

[0166] We conducted high-throughput population screening utilizing fluorescent reporters on a microcapillary array platform. This platform enables the consistent growth of spatially isolated cells within capillary channels and allows for their precise recovery. In this study, we screened over 3.9 * 105capillaries with a turnaround time of 10 min to an hour. Microwell plate readers (Auld, D. S.; Coassin, P. A.; Coussens, N. P.et al. Microplate Selection and Recommended Practices in High-throughput Screening and Quantitative Biology. In Assay Guidance Manual; Markossian, S.;Grossman, A.; Arkin, M.et al., Eds.; Eli Lilly & Company and the National Center for Advancing Translational Sciences: Bethesda (MD), 2020) and multi -parallel bioreactors (Bareither, R.; Pollard, D. A review of advanced small-scale parallel bioreactor technology for accelerated process development: Current state and future need. Biotechnol. Prog. 2011, 27, 2-14; Bertaux, F.; Sosa-Carrillo, S.; Gross, V.; et al. Enhancing bioreactor arrays for automated measurements and reactive control with ReacSight. Nat. Commun. 2022, 13, No. 3363; Carneiro, S.; Ferreira, E. C.; Rocha, I. Metabolic responses to recombinant bioprocesses in Escherichia coli. J. Biotechnol. 2013, 164, 396-408) are alternative commercial technologies that allow for the screening of cell populations. However, these technologies have limited throughput, with the number of clones per screen restricted to 103. Screening libraries on microwell plates is impractical, as it necessitates the differentiation of thousands of variants on agar plates, followed by picking and inoculation. TheAttorney Docket No. 0073605-001135 absence of an inherent mechanism for spatial isolation of cells renders these technologies unfit for library cloning protocols that yield mixtures of variants as final products. Microcapillary arrays offer screening rates that are 102to 103orders of magnitude higher, outperforming plate readers even in analyses of single variants. We demonstrated this by scanning approximately 183,000 capillaries (containing around 54,000 clones) per session, a capacity that can be extended as needed.Furthermore, the superior throughput of arrays helps reduce data noise by ensuring parallel replicates of variants.

[0167] Our platform ensures the accuracy of binning gates because, unlike FACS, binning gates are determined after screening the entire population rather than just a representative sample. In addition, we have customized the imaging and analysis protocols to facilitate the sorting of dual-reporter systems based on unique fluorescence signatures and specific mathematical functions. To our knowledge, such a modality in screening systems has yet to be introduced. Lastly, we have demonstrated in-experiment induction of structural proteins, with titer measurements in vivo and cell count estimates using bright-field absorbance.

[0168] We demonstrated the potential of arrays in optimizing biomanufacturability by screening libraries of promoters and 5’ UTRs, which are among the most critical plasmid elements. After testing a range of transcription rates on the order of 106, we found that promoters with moderate transcription rates were optimal for mRFP1 protein titer due to enhanced cell proliferation. Additionally, screening a library of 62,1205’ UTRs revealed a direct correlation between protein titer and mRNA half-life. Furthermore, we quantified the highly differentiated expression of sfGFP and mRFP1 in clones through dual-reporter imaging. We identified small subsets of clones that maximized the disproportionate sfGFP expression and cumulative sfGFP-mRFP1 expression, which correlated with the respective RBS strength. Such features may contribute to fundamental studies involving operons, such as optimizing relative protein expression, increasing overall protein levels, and analyzing elements like promoters and RBSs.

[0169] This study also focused on evaluating the variability in recombinant structural protein expression. We showed that the added toxicity of these proteins caused most cells in cultures to be either nonviable or to grow at reduced rates, resulting in broad titer distributions compared to the narrow distributions for mRFP1. The conventional strategy of in-experiment induction at various time points was replicated, demonstrating the separation of growth and production phases, which led to improved titers. Fluorescence imaging and direct cell growth measurements in capillaries using bright-field absorbance supported these experiments, which have not previously been demonstratedAttorney Docket No. 0073605-001135 with microcapillary arrays. This study also confirmed conventional knowledge about expression systems, including the inhibited growth due to the T7 induction system, the impact of the hydrophobicity score of the amino acid sequence, and the evolution of culture properties over time. Additionally, a fluorescent biosensor was designed for in vivo measurements to eliminate the need for fluorescent tags and invasive protocols for protein concentration determination. The platform has the potential to aid in discovering novel protein sequences with greater biomanufacturability while preserving functionality, especially when complemented with mutagenesis and directed evolution (Packer, M. S.; Liu, D. R. Methods for the directed evolution of proteins. Nat. Rev. Genet. 2015, 16, 379-394; Wang, T.; Badran, A. H.; Huang, T. P.; Liu, D. R. Continuous directed evolution of proteins with improved soluble expression. Nat. Chem. Biol. 2018, 14, 972-980).

[0170] Our platform facilitates the screening of various fluorescent reporters in the visible region (peak wavelengths from 405 to 635 nm, covering DAPI to Cy5), controls irradiance levels in 1% increments, allows for individual and simultaneous excitation, supports multiple image acquisition methods, and enables custom image analysis. We have demonstrated these features exclusively through the various experiments discussed above. The two main challenges of screening-namely, the creation of false low bin regions with low-fluorescing clones and drift during automated image acquisition- were addressed through image manipulation and discrete imaging, respectively. From a user-friendliness perspective, microscope environmental chambers (Yuan, Y.; Lu, F. A Flexible Chamber for Time-Lapse Live-Cell Imaging with Stimulated Raman Scattering Microscopy. J. Vis. Exp. 2022, No. 64449) can help eliminate the need to transfer arrays to incubators between successive imaging sessions.

[0171] Structural proteins self-assemble into large supramolecular constructs. However, these folded conformations are difficult to replicate inside the cell due to the post-translational modifications or chemical environment necessary for self-assembly and the varying concentrations of expressed proteins. Compared to flow cytometry, the microarray platform provides real-time data related to cell growth, which can aid in screening structural information even at low levels.However, new screening methodologies beyond optical techniques, employed in this study (e.g., thermal, acoustic, electrical), should be implemented to address this long-standing challenge for high-throughput systems and large library candidates.

[0172] In summary, our microcapillary array platform enables rapid multiplex analyses of plasmid libraries, which would take weeks of continuous effort (such as laborious agar plating and colony picking) with conventional technologies. The relevance of miniature growth chambers inAttorney Docket No. 0073605-001135 understanding microorganism behavior has been a significant concern (Van Zee, M.; de Rutte, J.; Rumyan, R.; et al. High-throughput selection of cells based on accumulated growth and division using PicoShell particles. Proc. Natl. Acad. Set. U. S. A. 2022, 119, No. e2109430119). However, micro-capillary arrays show expression-growth profiles that correlate with DNA / RNA sequencing markers derived from traditional culturing techniques. The spatial isolation of clones prevents biological crosstalk and allows users to study the unique phenotypic behaviors independently. Using this platform, we outlined the influence of several key plasmid components and demonstrated its potential to enhance the understanding of the complex phenomenon of protein expression.Methods

[0173] Instrumentation. The platform (FIG. 2D) is based on an inverted fluorescence microscope (Olympus IX73) with a multichannel LED illuminator (CoolLed pE-300Ultraor pE-400Max) The microscope was equipped with a digital camera (Olympus DP23M), motorized XY stage (Marzhauser Wetzlar Tango 3), and motorized focus drive (Marzhauser Wetzlar). The proprietary software cell Sens (Olympus) was used for imaging and stage control. The transmission characteristics of the filter cubes utilized (69302, 89401, Chroma) are illustrated in FIG. 18.Microcapillary arrays with 20 μm capillary diameter (INCOM, Inc.) were used throughout all experiments, which contain about 8 x 105capillaries in total. A 405 nm laser diode with a collimator was used for laser extraction. The laser was pulsed using a custom micro-controller. The pulse train was optimized (2 ms pulse separation, 2 ms pulse width, 4-5 in number at 118 DAQ level (power ~ 75 mW); FIG. 6) to achieve the highest extraction efficiency. The power of the ultraviolet (UV) laser beam emanating from the objective was measured using an optical meter and sensor kit (Model 843-R, Newport Corp.). The laser was directed through the main objective by a beam reflector. The objective focused the collimated laser beam into a ≈ 10 μm spot, aligned with the face of the capillary to be extracted. Although pooled extraction has been demonstrated previously (Chen, B.; Lim, S.; Kannan, A.; et al. High-throughput analysis and protein engineering using microcapillary arrays. Nat. Chem. Biol. 2016, 12, 76-81), we conducted manual extractions here to confirm the recovery of contents on slips after each iteration.

[0174] Imaging and Sorting. All snapshots were acquired with a 10x objective (UPlanFL N, NA = 0.30, Olympus). An image of the loaded area of an array was constructed by stitching snapshots (3088 x 2076 pixels, about 1482 x 996 pm) together in an automated fashion (unless stated otherwise). Snapshots were obtained through raster scanning along a predefined path in a two-dimensional (2D) plane (10,210 and 7150 capillaries per second at 100 and 200 ms exposure,Attorney Docket No. 0073605-001135 respectively), with the focus being automatically adjusted at each subsequent location. The generated image was processed with custom MATLAB code within 102seconds. The images were converted to grayscale and then binarized using the “imbinarize” function with the thresholding level defined by the “graythresh” function. Fluorescent capillaries were identified as white regions on a black background. Properties (centroid, area, and pixel indices) for all regions were obtained using the “regionprops” function. The fluorescence intensity of a region was calculated by summing the grayscale pixel values of all the comprising pixels. All regions (and their properties) were stored in a matrix and sorted for the measured intensity (or a mathematical operation). Before extraction, properties of regions of interest were analyzed, and care was taken to omit regions with areas deviating from normal substantially (e.g., coalesced regions). The coordinates of all places in the image were calibrated with the sample stage coordinates. Using the calibrated coordinates, regions of interest were swiftly aligned with the UV laser spot in cellSens for extraction. In all instances in the text, the total number of capillaries scanned refers to all the capillaries in the image frame (with or without cells).

[0175] Microcapillary Array Loading, Screening, and Sequencing. The arrays were sterilized before use. A mixture of magnetic beads (Dynabeads MyOne Silane, Invitrogen) and cells in media (Luria broth (LB)) was loaded onto the array. The small diameter makes it easier for liquid to pass through and be contained because of the surface tension of the liquid (de Gennes, P.-G.; Brochard-Wyart, F.; Quere, D. Capillarity and Wetting Phenomena.; Springer New York: New York, NY, 2004). This property is utilized to encapsulate cell cultures in capillaries, providing a controlled microenvironment for cell proliferation. To prevent capillaries from being overpopulated, we adjusted the cell concentration in the loading mixture to ensure that coverage remains below 33%. On average, one in every three capillaries acquires cells according to Poisson’s statistics (Chen, B.; Lim, S.; Kannan, A.; et al. High-throughput analysis and protein engineering using microcapillary arrays. Nat. Chem. Biol. 2016, 12, 76-81). The magnetic beads cause a shift in the intensity distribution of capillaries; however, the distribution profile remains unchanged (Chen, B.; Lim, S.; Kannan, A.; et al. High-throughput analysis and protein engineering using microcapillary arrays. Nat. Chem. Biol. 2016, 12, 76-81). The concentration of magnetic beads was kept at 14 mg / mL. After loading, the array was overlaid with an agarose gel layer (2% w / v, 1-2 mm thick) and incubated in a sealed Petri dish lined with moist wipes for a desired period. The setup of the sample stage is shown in FIG. 2B. The array rests on two standoffs on an indium titanium oxide (ITO) coated glass piece (Adafruit). The extraction slip (micro cover glasses, VWR) was inserted betweenAttorney Docket No. 0073605-001135 the array and the glass piece. ITO glass was Joule-heated to 28-30 °C temperatures by passing 48 mA current using Bio-Rad PowerPac. This prevented condensation over the extraction slip during experiments. After screening, arrays were sterilized in 70% ethanol for several hours. Arrays were cleaned thoroughly under a DI water stream, with intermediate ultrasonication for 2 min.

[0176] PCRs for plasmid identification were conducted with Q5 high-fidelity kits (New England Biolabs). After each screening experiment, extracted contents were scraped off the extraction slip using a fine micropipette tip while applying 60 μL diluted Q5 reaction buffer (with 10× Q5 buffer to nuclease-free water ratio = 10:33.5 μL) on the slip and collected in an Eppendorf tube. PCR with 43.5 μL elution was carried out for 35 cycles at designated melting temperatures of primers and extension times for amplicons. PCR products were sequenced through the Oxford Nanopore technique (Plasmidsaurus), and alignment was done using the Context Aware AutoAlign algorithm (De Nova DNA) for sequence identification.

[0177] pFTVl Cell Growth. The plasmid design can be accessed at Addgene (Addgene plasmid #63848, RRID: Addgene_63848). Isogenic colonies were picked and cultured in LB supplemented with chloramphenicol (20 μg / mL) for 5 h at 37 °C and 250 rpm. The cell culture was diluted in fresh media (5 μL in 1 mL), and 10 μL was loaded onto an array. The array was incubated at 37 °C. Bright-field, and mRFP1 fluorescence snapshots of several array locations were acquired at t = 1, 4, 8, and 24 h incubation periods. After each round of imaging, the array was incubated, and the exact locations were imaged in subsequent rounds to record cell growth and protein expression. Bright-field snapshots were binarized and used as masks for fluorescence snapshots to calculate intensity per capillary.

[0178] Binning Experiments. Fluorescence regions identified by image processing were grouped into three logarithmically spaced bins for normalized intensity (low [0, 0.1], medium [0.1, 0.316], and high [0.316, 1]).intensity min(intensity)normalized intensity = - - - - - - - - maxyintensity) min(intensity)

[0179] Each bin’s desired capillaries were extracted on a clean extraction slip. The original fluorescence image was used to screen directly for medium and high bins. However, for the low bin, the grayscale fluorescence image was dilated once before image processing using the “imdilate” function with a disc-shaped structural element of size 6. A new normalized intensity map was generated in which low bin capillaries were identified and extracted in desired numbers. Capillaries to be extracted from each bin were selected from the corresponding list randomly using the “randperm” function in MATLAB.Attorney Docket No. 0073605-001135

[0180] Promoter Library Preparation and Screening. The detailed methodology of preparation of the promoter library is mentioned in ref 29 Cells were picked from the cryostock using a micropipette tip. Cells were cultured in LB supplemented with chloramphenicol (50 μg / mL) at 37 °C and 250 rpm until an 0D600 of 0.1-0.2 was reached. The cell culture was diluted in 100 μL fresh media (as per relation: 2 μL starter culture at OD600 = 0.1) containing magnetic beads at 14 mg / mL cone. The diluted culture was loaded onto an array and incubated at 37 °C overnight. The extraction of top performers (3 replicates) and binning were performed with the library. A separate array was used for each experiment.

[0181] 5L UTR Library Preparation and Screening. The detailed methodology of preparation of the promoter library is mentioned in ref 33 The same protocol of array loading and incubation as for the promoter library was followed. Cells picked from the cryostock were used to prepare cell cultures. The volume of diluted culture loaded onto the array was increased to achieve greater array coverage. Binning of sfGFP and mRFP1 each and arithmetic-based screening were performed with the library. A separate array was used for each experiment. TIRs were calculated using the RBS calculator (denovodna. com), and the sequences are mentioned in Supporting Data 1.

[0182] Arithmetic-Based Screening. Schematic in FIG. 11 illustrates the algorithm for ratiobased screening. At several array locations, bright-field and fluorescence snapshots were acquired discretely. Bright-field snapshots were binarized, and resulting snapshots were used as binary masks for fluorescence snapshots. The mathematical operation was done on sfGFP and mRFP1 intensities for each capillary, and the cells with the greatest sfGFP to mRFP1 ratios were extracted, and the sequences of 5’ UTRs were identified.

[0183] The previous approach of raster scanning and image stitching (as for binning) was followed for the summation-based screening experiment. Our fluorescence illumination system allowed for the use of multiple LEDs simultaneously for concurrent excitation of sfGFP and mRFP1. In systems where various LEDs cannot be used simultaneously, the ratio-based screening protocol can be used while switching the mathematical operation to summation.

[0184] Structural Proteins Cloning and Screening. The plasmids were designed using the Genetic Systems Builder (all calculators are available at De Novo DNA). The algorithm first optimizes DNA sequence for efficient expression of the gene of interest. All plasmids are transformed into a pool of ssDNA oligos which was sourced commercially (Twist Bioscience). Furthermore, PCR amplification primers for each oligo type were ordered separately. Using the specific primer pair, oligos of a desired plasmid were amplified and converted into dsDNAAttorney Docket No. 0073605-001135 fragments through PCR. This PCR product was then purified and used for standard GG assembly. The assembled plasmids were transformed into E. coli BL21-DE3 strain (NEB) using the standard protocol.

[0185] Isogenic colonies of SP proteins were inoculated in LB supplemented with chloramphenicol at 28 °C overnight, and the culture was utilized to load the arrays. The culture was diluted according to the relationship mentioned earlier for libraries. The arrays were incubated at 30 °C. For t = 0 induction, IPTG was added to the culture before loading at a final concentration of 400 μM. For in-experiment inductions, 0.1 mM IPTG stock was applied to the agarose gel surface and spread gently. The volume of the gel and loaded area of the array were considered to estimate the volume of the stock solution to be added to achieve a final concentration of 400 μM in capillaries. Diffusion is a kinetic process, and we attempted to measure the diffusion rate of IPTG into the capillaries through Fourier Transform Infrared Spectroscopy of the elution from arrays after several hours. However, the IPTG concentration was too low to be detected. Earlier studies on diffusion of similar molecules in agarose43suggest that saturation in the setup (through 1.5% w / v gel) would take 2-3 h at 25 °C. For focused studies, dedicated diffusion analyses can be carried out for the molecule in question. A separate array was used for each experiment with an SP.

[0186] The discrete imaging technique was followed for screening SPs to minimize errors in fluorescence intensity measurements per capillary. For background correction, fluorescence snapshots of blank regions of arrays under the gel were acquired. After segmentation, bright -field images were acquired to act as binary masks for fluorescence images. For bright-field absorbance measurements, the irradiance of the light was reduced such that capillaries with low cell count posed significant obstruction for their detection. The bright-field images were segmented and used as binary masks to calculate bright-field intensity per capillary. Baseline correction of the distributions was performed manually using levels centered around the mode of each trend, as detailed in FIG. 19. Finally, absorbance per capillary was calculated using the following relationintensityAbsorbance = −log10- -; - baseline intensity

[0187] It is to be noted that baseline correction does not alter the scattered distribution patterns of data points and only influences the numerical values of absorbance. The capillaries were filtered for size, and the ones with face areas smaller than average were excluded from the analysis.Protein SequencesAttorney Docket No. 0073605-001135

[0188] Refl ectin and cement are structural proteins that play key roles in the iridescent coloration of cephalopods and in the adhesive properties of barnacle cement, respectively. We chose these proteins because of their significant hydrophobicity differences, which are essential for biomanufacturability (i.e., protein level expression and cell growth). Below are also the protein sequences of the reporter fluorescent proteins provided.

[0189] Reflectin Protein (SEQ ID NO: 1):MYYPERYFDMSNWQMDMQGRWMDMQGRYCSPYWYNWYGRQMYYPYQNYYWYGRW DYPGMDYSNWQMDMQGRWMDMQGRYMDPWWMNDSYYNNYYN*

[0190] Cement-1 Protein (SEQ ID NO: 2)

[0191] MSGPVKDDDEYKETREYTKEYTVDRNKGYNRKGYGDDVSAKETFLRTTEYND KDGYGSKNDGKQVKESYTREYEIDKDGYGKDDDVTGKETYRREVTVDDDGYSPRGRYGR TDAYGLNDGYVRKDAYGLRRGPGLGVYPGPGLGRKDVYPATGLGRKDVYPVTGLGRKDV YPATGLGRKDVYPATGLGRKDVYPATGLGRKDVYPVPGLGRKDVYPGPGLGRKGVYPGR GFGVPGPFW

[0192] sfGFP (SEQ ID NO: 3):

[0193] MELKQNKEQITKQNINRKGEELFTGVVPILVELDGD VNGHKF S VRGEGEGD ATN GKLTLKFICTTGKLPVPWPTLVTTLTYGVQCFARYPDHMKQHDFFKSAMPEGYVQERTISF KDDGTYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYNFNSHNVYITADKQKNGI KANFKIRHNVEDGSVQLADHYQQNTPIGDGPVLLPDNHYLSTQSVLSKDPNEKRDHMVLL EFVTAAGITHGMDELYK*

[0194] mCherry (SEQ ID NO: 4):MISKGEEDNLAIIKEFMRFKVHMEGSVNGHEFEIEGEGEGRPYEGTQTAKLKVTKGGPLPFA WDILSPQFMYGSKAYVKHPADIPDYLKLSFPEGFKWERVMNFEDGGVVTVTQDSS LQDGEFIYKVKLRGTNFP SDGP VMQKKTMGWEAS SERMYPEDGALKGEIKQRLKLKD GGHYDAEVKTTYKAKKPVQLPGAYNVNIKLDITSHNEDYTIVEQYERAEGRHSTGGMD ELYK*

[0195] mRFP1 (promoter library; SEQ ID NO: 5):MASSEDVIKEFMRFKVRMEGSVNGHEFEIEGEGEGRPYEGTQTAKLKVTKGGPLPFAWD IL SPQFQYGSKAYVKHPADIPDYLKLSFPEGFKWERVMNFEDGGVVTVTQDSSLQDGE FIYKVKLRGTNFPSDGPVMQKKTMGWEASTERMYPEDGALKGEIKMRLKLKDGGHYD AEVKTTYMAKKPVQLPGAYKTDIKLDITSHNEDYTIVEQYERAEGRHSTGA*

[0196] mRFP1 (5’ UTR library; SEQ ID NO: 6):Attorney Docket No. 0073605-001135 MLEASSEDVIKEFMRFKVRMEGSVNGHEFEIEGEGEGRPYEGTQTAKLKVTKGGPLPFA WDIL SPQFQ YGSKAYVKHP ADIPD YLKL SFPEGFKWERVMNFEDGGVVTVTQD S SLQD GEFIYKVKLRGTNFPSDGPVMQKKTMGWEASTERMYPEDGALKGEIKMRLKLKDGGH YDAEVKTTYMAKKPVQLPGAYKTDIKLDITSHNEDYTIVEQYERAEGRHSTGA*Supporting Data 1: Promoter LibraryTable 1: List of identified promotersNormalized DNA Normalized RNA Normalized TX Count Count Rate Replicate 1Promoter PromoterR1 R2 R1 R2 MEAN STDEV ID SequenceSEQ IDNO:7:ATACTAAG ATCTCAGG GTCATTGA SLP2018-2- CAACAGG13605.3 16765.1 48746.8 41056.2 10.1 0.3 54*** GGAATGCACTGTTAT AATTCTAC GCCGGTGA CGTGGACG GGACTCGA SEQ IDNO:8:TCTCATAT SLP2018-2- CTCCTGAG3550.1 4823.3 42528.9 32528.0 31.0 3.9 818 GTTTTTGACATGGACG TATGAATC TTGTATAGAttorney Docket No. 0073605-001135 TCTGACTC CGTCCCAT GAACCAGT CGCTGA SEQ IDNO:9:AAGATAA CACAAGA CCGACCTT SLP2018-2- GACACATG 185319. 125407.66760.2 74976.4 7.4 0.7 ygj * * * ACGCTAGG 0 3CAAGATAT AGTGATAG GCCGGTAA GTCAAGCT ACTGATCG SEQ IDNO:10:AAACCGA ATAAGAA ATCACCTT SLP2018-1- GACATACG8262.3 10238.9 40574.9 33430.8 13.7 0.6 684 TTAGATGTTGTTTTAT AATCGCAA ACCTGTCG GCCCGTGA GCTTAGCGReplicate 2SEQ IDNO:SLP2018-2- 11: 5376.9 6776.1 13170.7 10078.3 6.5 0.6 626GCCGTATAAttorney Docket No. 0073605-001135 TCGTAGCT TCGATTGA CATGGTAT CTTGATTT GCTTAGAA TCGGTCCC CAACCGAT AATTAGAG GTCTTA SEQ IDNO:12:CGGTAACG TGAACTCT TGCCTTGA SLP2018-2- CAACTACA15700.3 19750.1 56107.9 21945.9 7.4 3.0 154 GGACAGTATACTATA ATGTCTAA CCACTGAA CGATGCGT CTAGTTC SEQ ID NO:13:GTCTCGTA CGTCGTGA SLP2018-2- AAACTTGA5812.0 5418.6 23792.5 16752.3 12.2 0.2 81 CACTCTCTCCGCGCCG ACGTATAA TGATCCGC CAATGCAAAttorney Docket No. 0073605-001135 CGGATACG GCGGGT SEQ ID NO:14:AGAACAA TGCTACGG TTTTGTTG SLP2018-2- ACACTAGC8298.2 10877.5 42011.7 37357.0 14.3 0.5 245 GCGGCGTTAGGGTATA ATCAGGTT CCGGATGC GTTTTGCG GTCGTAG SEQ ID NO:15:GACCTATT AATTATCG AGCATTGA SLP2018-1- CAATGCCG6067.0 6885.9 29117.3 15515.2 11.5 2.5 231 TAGCGCCATGTTATAA TTACCTTC CAGTTCAG TACTTGTG GCTGTCReplicate 3SEQ ID NO:16:SLP2018-2- ACATACCT 14453.3 16344.5 72074.9 48391.0 13.2 1.3 15GACTACGC GGCCTTGAAttorney Docket No. 0073605-001135 CACGGCGC ATGGAGCC CGTTATAA TACACCCC CAATTACT GGTCTCAG GCCTAT SEQ IDNO:17:ATACTAAG ATCTCAGG GTCATTGA SLP2018-2- CAACAGG13605.3 16765.1 48746.8 41056.2 10.1 0.3 54*** GGAATGCACTGTTAT AATTCTAC GCCGGTGA CGTGGACG GGACTCGA SEQ IDNO:18:AAGATAA CACAAGA CCGACCTT SLP2018-2- GACACATG 185319. 125407.66760.2 74976.4 7.4 0.7 793 * * * ACGCTAGG 0 3CAAGATAT AGTGATAG GCCGGTAA GTCAAGCT ACTGATCG***Promoters identified in both Replicates 1 and 3.Attorney Docket No. 0073605-001135Other stronger promoters in librarySEQ IDNO:19:TCTACGTG CCCTCATA ACCATTGA SLP2018-1- CAGACTCC 160050. 138923.834.4 960.9 569.0 10.9 352 ACGGGAA 0 6ATCCTATA ATTTAAAA CCTGAGAT GATAAGA CTTCGAGT SEQ IDNO:20:AGACCATC CCGCGCTC AGATTTGA SLP2018-1- CAATGCTC78.4 83.2 11043.1 11917.2 492.4 82.5 516 GCTGCTGGCTTTATAA TTGTTGGC CTCTAGGT CGTATACG TACTAA SEQ IDNO:21:CGCCCTCC SLP2018-2- GTAAGCTC 190.8 176.4 19297.4 18852.4 361.5 67.2 1328TGTATTGA CGAAATG AATAGTGTAttorney Docket No. 0073605-001135 TGCACATA ATAGCATC CCCGTTGG TGGGATCA AAAGGAA SEQ IDNO:22:ACTCAATT TATTCTAG CAATTTGA CAACCGCT SLP2018-2- 219875. 111900.GCGAGAA 1531.7 1593.9 349.6 68.0 255 8 2CTTCTATA ATCGAAG ACCAAATT TGGAGAGT AGCAAAG A SEQ ID NO:23:GATGATAC CGGCGTCC CTAGTTGA SLP2018-2- CATGTTAT964.7 1021.2 79030.3 73685.4 263.9 25.5 162 CAGGGTATTTTTATAA TAAAGAG CCTGCATA CTTGCAAT AGGGGTT SLP2018-1- SEQ IDNO:894.6 1214.3 85759.1 67146.1 250.4 28.6 573 24:Attorney Docket No. 0073605-001135 TTAAAAAT AGCTGATC TCCCTTGA CAGGGTG ATTTAGCC GAGATATA ATTATGGT CCAGGGTT CAGAACCT GAGTCCA SEQ ID NO:25:CGCAGCTC GATACTAG ATTTTTGC SLP2018-2- AACTAGGT26.2 97.4 3000.1 3473.5 238.0 94.9 1813 AGCATTGAAGTTATAG TTTTACTC CAGTATCC TGGAGCA GTGAAGTTable 2: List of identified promotersNormalized DNA Normalized RNA Normalized TX Count Count Rate Low BinPromoter PromoterR1 R2 R1 R2 MEAN STDEV ID SequenceSEQ ID NO:SLP2018-1- 26: 5665.5 5579.9 12292.5 12947.1 7.8 1.5 530TAGTCAGGAttorney Docket No. 0073605-001135 AGTGAAG GCTTATTG ACAACAC GCAAGGA GTATACTA TAATCCCG GGCCAGG AACTTGGC AAGATCG AGG SEQ ID NO:27:GGAAGGG ACAAGAT GATGGCTT SLP2018-1- GACATGCT18362.0 25907.4 22749.3 19354.3 3.3 0.3 64 GCATCTAGGGAACTAT AATCCAAT GCCCGCGT TTCTATAC CTGTTAGT SEQ ID NO:28:TTACGTAG GCTTAGGG SLP2018-1- TCCGTTGA2975.7 3966.1 17710.2 26729.3 22.2 4.9 692 CACATTGTGAGTAGG AGTCTATA ATTAACAG CCCTAAAGAttorney Docket No. 0073605-001135 GGATGGA AGGCGTA G SEQ ID NO:29:TTTCTCTC TATTCATT ATTGTTGG SLP2018-2- TATTCACC776.0 840.2 8387.5 8551.0 36.1 4.7 1711 CTTAGATGTTCTATAC TCGGGGAC CCCTGGTT TTGATGCA GCACCG SEQ ID NO:30:TTGGGGAT AATTTCTC TCAATTGC SLP2018-2- AAGAGTA1499.4 2181.0 3714.2 4640.3 7.9 0.7 1776 AGTCCGCACTAGTATA CTTGCTGG CCACATAG TGCTCGTG GTGCACA SEQ ID NO:31:SLP2018-2- CACTCGAT 5208.1 6543.4 4215.1 5038.9 2.7 0.4 2441ATCAAACT TTGCTTCCAttorney Docket No. 0073605-001135 TAGCAATA CAGCTTAA TGCTACAA TCTTGCAC CTGACTGA GCGCCCTG AGCTAT SEQ ID NO:32:TTCTAAGC CTGCTACG TACATTGA SLP2018-2- CAGACGGT32021.2 37292.3 60828.3 46706.6 5.3 0.3 250 TAGCTTAGACGTATAA TTCACCTC CCTTCAAA GTGCTGAA TGAGAT SEQ ID NO:33:GCCCACAC TCTAAACG TCAGTTGA SLP2018-2- CAGTGAA3004.4 3313.9 53635.0 41746.1 51.2 0.7 283 AATTCCTGGGGCTATA ATCGCTTT CCGTTTTG GAGGTGG GTAGCGCTAttorney Docket No. 0073605-001135 SEQ ID NO:34:TCGGTAGA TAAGCTCT AAGTTTGA SLP2018-2- CAAGATG10039.6 12762.3 29285.1 24654.3 8.1 0.4 381 GTTGTTTAACCATATA ATTCAAAG CCGATGGG TGTGACCG GTCGCCA SEQ ID NO:35:GTAGTGTT TACCTGTC TACGTTGA SLP2018-2- CAAGCAA14215.5 17115.8 47164.8 41459.0 9.7 0.0 813 CCGGGCGGACTTTAT AGTCTCTA TCCATACC CGACGTAA TGGTACTC SEQ ID NO:36:ACATTTCC SLP2018-2- GCACGAA8067.9 9853.4 11911.9 6000.8 3.4 0.9 826 GTGTCTTGACAGAAG TTCCGCAT CAGGTTATAttorney Docket No. 0073605-001135 ACTGTAGG ACCTATAC CAAATGG GTCAATGC T SEQ ID NO:37:ACTACCCG TGGATTAC ATCATGGA SLP2018-2- CAATGCGT19815.5 24382.9 48946.5 42513.4 7.1 0.1 866 CTACTGCCGAATATAA TGGAAGTC CTAGGCTC ACTTAGCC TCGAAT SEQ ID NO:38:CATATTAC GACAAATT ACTGTTGA SLP2018-1- CAGGACCC 126,2230.5 2122.8 88914.4 166.3 1.7 306 GAGACAA 240.2CTTCTATA ATCCCTAG CCATCAGA GTATCGAA TATCCTA SEQ ID NO:SLP2018-2- 39: 18146.2 19313.6 977.5 1519.2 0.2 0.1 3403ATGTTGTAAttorney Docket No. 0073605-001135 TCGTGGCC TAGACGA ACAGGTG GACAGAA ACGGTAAT TAAAGAA AGTCCTGC GGTGCTGA TCGGGTTT AGMedium BinSEQ ID NO:40:GGGTCTTA ACCCGCGT GTCTTTGA SLP2018-1- CATTAGCC10164.0 11662.2 25736.8 24971.7 8.0 0.6 10 ATCCTACCGAATATAA TCATGGCC CTGAATAA GGATCACG TATTCG SEQ ID NO:41:ACCCTTAG AGGCATTT SLP2018-1- TTGTTTGA 5592.7 7011.1 73855.3 72166.3 39.9 1.4 310CAATGTCC GGAGACG GCTCTATA ATAACACCAttorney Docket No. 0073605-001135 CCGCGTAT GAGGGCTT CTTAGAT SEQ ID NO:42:GGACCTCA ACAGCCCC ACACTTGA SLP2018-1- CACGGAC46613.9 53842.0 93257.2 85169.5 6.1 0.3 743 ATTCAACCAACCTATA ATCCGAAC CCATTCGT GCACGTTT GAGTCGA SEQ ID NO:43:GCCGAGTA GTATCGAC GTAGTTGA CAGTCTTC SLP2018-2- 147, 128,AACAGAA 24518.6 24197.2 19.4 1.9 12 148.1 344.3CCGATATA ATAAAGG GCCAAGG ATCATTCA CGGGTCTT G SEQ IDNO:SLP2018-2- 44:6805.6 8403.6 27324.3 27979.5 12.5 0.8 1501 CGTCAAAGAAGCACCAttorney Docket No. 0073605-001135 GACAATTG CTACCCCT TGTTGAAC ATGTTACA ATCATGCC CCCATAGA TCTGATGA AGCGTTT SEQ ID NO:45:TGACCCCA TTCACGCT AGATTTGA SLP2018-2- CACTGCCC 136, 138,14845.5 17957.0 28.9 2.1 378 AACCGATT 807.9 800.1TTGTATAA TACCAGAC CGTAGCGA ATGAGCG ACGATCT SEQ ID NO:46:CCCGGCAA GTCCGAGC CTGATTGA SLP2018-2- CAATTCAT 102,14811.6 14609.0 92228.0 22.7 2.6 407 GGTAATGG 481.9CGCTATAA TGTGGGTC CGCAGAC ATACTTAC TTGGGAAAttorney Docket No. 0073605-001135 SEQ IDNO:47:GCACATGC AGATAGG GAACATTG SLP2018-2- ACAGTATC 123,92782.2 22019.9 17661.8 0.6 0.1 442 TACCCGAC 207.9CAGGTATA ATGGTTGC CCCCTGAG ACGTAGTA ATGCCAT SEQ IDNO:48:TACCCCTA ATGAAGA CCAACTTG ACAAAGC SLP2018-2- GGATAGA 9632.2 11021.5 13340.0 9964.7 3.8 0.2 444CCGTACTA TAATGAGC AGCCAAG GGCTATTG ATCATCTC TC SEQ IDNO:49:GCAAGCA SLP2018-2- 153, 166,CTCGGTGA 4765.8 4789.5 0.1 0.0 862 234.6 862.6CCGTGATG ACACGCTG TGGTCCAGAttorney Docket No. 0073605-001135 GGCATATA ATGCCGAA CCCGGAAC GTAGCATG TAAAACGHigh BinSEQ IDNO:50:GAAATGC ATTGCGTT GTTCCTTG SLP2018-1- ACAAGTCG69564.7 86108.8 96013.6 82028.8 3.9 0.1 190 CGGGACATTACGTATA ATCACACG CCATCCTC GCAGGTG ATGCATCC SEQ ID NO:51:GAGATCG GTCCTTTA TACTATTG ACATTACT SLP2018-2- 235, 198,GGGATGG 9047.2 10021.9 77.7 1.9 129 734.0 747.0GAACCTAT AATCACCA ACCTTGTA CCTCAGTA GAGTAGACAttorney Docket No. 0073605-001135 SEQ IDNO:52:GAATAAG CGCCCCTC GTGAGTTG ACTTTTAA SLP2018-2- 140, 105,CGAATTAG 25322.0 27725.0 15.7 0.4 1329 656.8 500.8CCTCGATA ATGAAGC ACCGATCA AGAACAA GGGTCAA AA SEQ ID NO:53:TTCCCCTG AAACGTAC CGTGTTGC SLP2018-2- TTATGGGA891.9 1037.7 2062.4 1558.5 6.4 0.4 2603 GCCAAAAAGTTTACC ATCTGTCC CCAATAAA AACAAAT GCATCAAA SEQ IDNO:54:TGGGTGGA SLP2018-2- CCTATGGC 8284.2 9402.4 19093.4 16290.9 6.8 0.1 328CCCCTTGA CACCCGTA CTAATCTAAttorney Docket No. 0073605-001135 AGCTATAA TAGGTGAC CGTAAATG GACAACGT TACTTT SEQ IDNO:55:AGGCGACT CGGTTTTC ATGTTTGA SLP2018-2- CACGAGTC24937.2 30332.6 54323.4 37365.9 5.6 0.7 472 TTAACCGCGAGTATAA TGTTTTGC CTCTTACC TACATGAC AATGGG SEQ IDNO:56:TCGTGCCG TACTTGTC GTCTTTGA SLP2018-2- CATACGAC 144, 188, 101,86206.4 1.9 0.1 731 TAAACGA 740.4 162.3 444.6GGAGTATC ATGTTTCA CCTTCTGT CTCACTGC TTGGTAT SEQ ID NO:SLP2018-2- 153, 134,57: 16291.8 15422.6 31.2 3.8 769 462.3 548.6ATGGATCGAttorney Docket No. 0073605-001135 CGTGCCAG GACATTGA CAATGATA AAAGAAC CTTCTATA TTGGGCGG CCCCGTGT GGTATTCT GATAGAC SEQ IDNO:58:AAGATAA CACAAGA CCGACCTT SLP2018-2- GACACATG 185, 125,66760.2 74976.4 7.4 0.7 793 ACGCTAGG 319.0 407.3CAAGATAT AGTGATAG GCCGGTAA GTCAAGCT ACTGATCGSupporting Data 2: 5’ UTR libraryTable 3: List of identified 5’ UTRs (mRFP1)Designed 5’ halfRNA RNA_ RNA_ RNA_ RNA k_ k_r2 Index UTR DNA life TO T2 T4 T8 T16 fit (min) fit SequenceLow BinSEQ ID NO:59:8753 13615 23152 8067 6293 3968 11608 1.1 0.6 1.0 TATGTTTTT GTTTTTGTTAttorney Docket No. 0073605-001135 TTCCCGCC ATATGGCG GGCCTGCT GGGTTGGT GCCGCGGG GTTTTGGG SEQ ID NO:60:TTCTTCTTC TTCTTCCCG CCATATGG20734 83608 73887 44694 39799 19029 31445 0.8 0.9 1.0 CGGGGCTC GTCCGTCTT CCCGGGTG GTTGGCGT G SEQ ID NO:61:AATTGCAC CGCCATAT GGCGGTGC37401 3492 5035 3203 2893 759 1608 0.7 1.0 1.0 AATTTTTA ACGGGAGA GCCTCGGT CAGCGCCA CC SEQ ID NO:62:GGGTCTAC52142 8799 10107 5319 5805 4640 7371 0.8 0.9 0.9 GGGGTGTG TAGGGGGT AGGGGGAAAttorney Docket No. 0073605-001135 ATCCCGCC ATATGGCG GGCCCAGG GAGACAAA TCCCGCAT TTATTGAG SEQ IDNO:63:TTCGTTTTT CTTTTTCTT TTCCCGCC10993 11395 8822 6269 4674 3140 1858 0.7 1.0 1.0 ATATGGCG GGTTTGGG GCTGCCGG CTCCCCCC GGCCCTCG SEQ ID NO:64:GGGTGGAG GGGGCTCA GCGGGTTA CAAGGGTT52889 7402 9633 3653 4139 3882 3419 1.0 0.7 0.9 ATCCCGCC ATATGGCG GGAGACGG AACCACGC CTAAGAGC GTTGAGCG SEQ IDNO:65: 137, 101, 77, 33, 80,44603 48858 0.7 1.0 1.0 GAAGCATA 382 446 306 682 361CAGAAGGAAttorney Docket No. 0073605-001135 TCAGAGGA AAGGCATT ACACGAT SEQ IDNO:66:CTACTAGA CGACTGTA45807 9385 16400 7845 5569 2563 8893 0.9 0.8 1.0 GAGCATCC CGATAAGA AGAGACGA GC SEQ IDNO:67:CCCCCCCCTCCCGCCA27678 TATGGCGG 9189 7919 9178 13592 5812 9640 0.2 2.9 0.7 GCCCCTTTC CACACTAT ATAATTTA ACTACT SEQ IDNO:68:CCCCCCAT TCCCCCAA GACCCGGC AATCCCTTT56852 7986 1928 1495 1491 1096 998 0.6 1.2 0.9 TCCCGCCA TATGGCGG GAGACGGA ACCACGCC TAAGAGCG TTGAGCGAttorney Docket No. 0073605-001135 SEQ ID NO:69:CAATGCAC CGCCATAT GGCGGTGC 197, 320, 167, 144, 802, 188,38909 0.8 0.8 1.0 ATTGCTCG 787 290 957 921 02 252ACTATATA TGTCTCAG GCCCCTGG CAMedium BinSEQ ID NO:70:AAGCGCCC ACAAGGGC GCAAAACA30076 AACCCGCC 494 611 473 331 163 397 0.7 1.1 1.0 ATATGGCG GGAAAAAG CCAAGCAC GCAAACCC GCACACAC SEQ ID NO:71:AAACCGGA AGCGGCCT TCCGAAGA50055 10636 22959 10488 11537 3437 9949 0.9 0.8 1.0 GACAAAGG ACAAGACA AGGATGGT AACGCGTT ACAttorney Docket No. 0073605-001135 SEQ ID NO:72:TTGAGCAC CGCCATAT GGCGGTGC38069 13617 27909 27397 31005 13556 25944 0.4 1.7 0.9 TCAAAAAG TCGTTTTAA GCAAGTGA GACATGAC T SEQ ID NO:73:CACACACA61011 CAGCTCGT 604 3630 5239 8315 6930 9160 0.1 8.0 0.3 TACAAATA AGGAGGTA AAA SEQ ID NO:74:TCCCACCT ACGAGAAC CCCACGTG CACCCACG 120,48217 69849 90687 76771 31415 61670 0.6 1.1 1.0 TGTAGCCG 472AAAAGTAG GTTCTAGA ATTATCAT GTTAACTA CTG SEQ ID NO:47371 75: 29365 43256 26954 22877 12706 24514 0.7 0.9 1.0 GCGCGCTCAttorney Docket No. 0073605-001135 GGGAGTTA CGGAACCC ACGGCTTG AAAGTACA CTATAGAA GGTTGA SEQ IDNO:76:CCCAATGC GCCCTTCC ATCCCCTTT TTCCCAAA56157 15844 3685 2417 2731 1778 2717 0.6 1.1 0.9 ACCCGCCA TATGGCGG GCCCCGAT ACCCTGCG CAAATCGC GCACAGC SEQ ID NO:77:ATAAGCAC AGGAAATC41897 CATCCTGT 3666 6296 3590 3297 2723 4685 0.8 0.9 0.9 GCTTATTTG GTCCATTG GTGTAGTA G SEQ IDNO:78:3778 CTGATTTTT 2012 2055 1906 1354 698 589 0.5 1.3 1.0TTCCCGCCAttorney Docket No. 0073605-001135 ATATGGCG GGGCTTGG CGCTCGGT CGTCGGGG GCGTGCTC SEQ ID NO:79:ACAGGATG TTTCCTGTC TCATTTAAT51103 AGTTCGGA 2093 4876 3515 2366 1966 2510 0.7 1.0 0.9 CCCCCTCC ACACGGAG GGAGTTTT GTGGAGGG TCTTTHigh BinSEQ ID NO:80:AAAAAAGA AAAAAGAA AAAAGAAC19879 CCGCCATA 431 725 548 579 930 476 0.5 1.4 0.5 TGGCGGGA AGACGGGG CGGCACGG GGCGGAAC CCGGC SEQ ID NO:81:44652 8170 25214 16013 13124 6760 8865 0.7 0.9 1.0 AGCAAACG TTTTTGCTCAttorney Docket No. 0073605-001135 TGACATAC AACACTGT ATGGTAAA AAGTGTAA CCAAAGGA TTTT SEQ ID NO: 82:GGGGGGGG GTGGGGGG GGGTGCCC29218 GCCATATG 454 755 446 897 75 824 0.6 1.2 0.8 GCGGGGAG GGGGGAAA TTAATAGG TTAAGATA GGT SEQ ID NO: 83:CGCATTTTT CTTTTTCTT TTCCCGCC12130 18066 7472 3558 3811 1459 3966 0.9 0.8 1.0 ATATGGCG GGTCTCGG TTTCTTCGT TTCCTGGTT CGTTCC SEQ ID NO: 84:33621 GCGCCCAC 3803 11523 11067 10118 5169 8007 0.5 1.5 0.9 AAGGGCGC GGGAGGGAAttorney Docket No. 0073605-001135 GGGCCCGC CATATGGC GGGAGGGT GGGAATTT GGGTATGT GAGATTTT A SEQ ID NO:85:CCTGTTTTT CTTTTTCTT TTCCCGCC12215 18134 39245 37434 28297 8878 16946 0.5 1.3 1.0 ATATGGCG GGTTCTTTC GTCCTGCG TCTGGGGG GTGTCCT SEQ ID NO:86:CCTAAAAA AACCCGCC16289 ATATGGCG 4187 3655 3195 2720 1903 2538 0.5 1.3 0.9 GGGAACGA AGAACCGA ACAAACCC CAACAAGA SEQ ID NO:87:CCGCGCCC32850 871 1828 3120 3387 1780 6814 0.1 5.1 0.5 ACAAGGGC GCCCACCA CCACCCCCAttorney Docket No. 0073605-001135 GCCATATG GCGGGCCT TTCATCTAT TAACTCAG GAGGAGGA GG SEQ ID NO:88:GCTAACAG ATTAGCCA AAACAGTT CAGATGAA45303 4071 18638 20280 28529 13279 24153 0.3 2.5 0.7 CTTCGAGA CTCGGGTC TCCAGATT GACACGAT AGGATACT AAT SEQ ID NO:89:GTTTTTTTT TGTTTTTTT23569 CCCGCCAT 4907 582 384 458 138 136 0.6 1.1 0.9 ATGGCGGG TTGTGGCG TTTTGGCG CTGGGGGG TCTCGG SEQ ID NO:47944 90: 25710 60033 30319 24690 13216 28233 0.9 0.8 1.0 ATCCTAACAttorney Docket No. 0073605-001135 ACCTTAGG TATGCCTA AGAAAATA GGATAGAG AGTAACAC CGGGGAGT TTTGC SEQ ID NO:91:GCGCCCAC AAGGGCGC CCTCCTCCT34182 CCTCCCGC 23765 2064 1345 1624 980 1750 0.6 1.1 0.9 CATATGGC GGGTCTTC TCTCTTTAA ACATCTCC TCACCACCTable 4: Translation initiation rates (TIR; mRFP1) corresponding to 5’ UTR sequences of Table 3 Designed 5’ UTR sfGFP RBS TIR mRFP1 RBS TIR TIR Ratio IndexSequence (au) (au) (sfGFP / mRFP1)Low BinSEQ ID NO: 59:TATGTTTTTGTTTTTG TTTTCCCGCCATATGG8753 183 14837 0.01 CGGGCCTGCTGGGTT GGTGCCGCGGGGTTT TGGG SEQ ID NO: 60:20734 TTCTTCTTCTTCTTCC 183 9502 0.02CGCCATATGGCGGGGAttorney Docket No. 0073605-001135 CTCGTCCGTCTTCCCG GGTGGTTGGCGTG SEQ ID NO: 61:AATTGCACCGCCATA37401 TGGCGGTGCAATTTTT 138 3709 0.04AACGGGAGAGCCTCG GTCAGCGCCACC SEQ ID NO: 62:GGGTCTACGGGGTGT GTAGGGGGTAGGGGG52142 AAATCCCGCCATATG 0 2683 0.00GCGGGCCCAGGGAGA CAAATCCCGCATTTAT TGAG SEQ ID NO: 63:TTCGTTTTTCTTTTTCT TTTCCCGCCATATGGC10993 14 6750 0.00GGGTTTGGGGCTGCC GGCTCCCCCCGGCCC TCG SEQ ID NO: 64:GGGTGGAGGGGGCTC AGCGGGTTACAAGGG52889 TTATCCCGCCATATGG 58 4882 0.01CGGGAGACGGAACCA CGCCTAAGAGCGTTG AGCG SEQ ID NO: 65:GAAGCATACAGAAGG44603 11023 1521 7.25ATCAGAGGAAAGGCA TTACACGATAttorney Docket No. 0073605-001135 SEQ ID NO: 66:CTACTAGACGACTGT45807 525 2170 0.24AGAGCATCCCGATAA GAAGAGACGAGC SEQ ID NO: 67:CCCCCCCCTCCCGCCA27678 TATGGCGGGCCCCTTT 2254 8448 0.27CCACACTATATAATTT AACTACT SEQ ID NO: 68:CCCCCCATTCCCCCAA GACCCGGCAATCCCT56852 TTTCCCGCCATATGGC 1519 5818 0.26GGGAGACGGAACCAC GCCTAAGAGCGTTGA GCG SEQ ID NO: 69:CAATGCACCGCCATA38909 TGGCGGTGCATTGCT 1770 3212 0.55CGACTATATATGTCTC AGGCCCCTGGCAMedium BinSEQ ID NO: 70:AAGCGCCCACAAGGG CGCAAAACAAACCCG30076 613 3594 0.17CCATATGGCGGGAAA AAGCCAAGCACGCAA ACCCGCACACAC SEQ ID NO: 71:50055 AAACCGGAAGCGGCC 13005 11480 1.13TTCCGAAGAGACAAAAttorney Docket No. 0073605-001135 GGACAAGACAAGGAT GGTAACGCGTTAC SEQ ID NO: 72:TTGAGCACCGCCATA38069 TGGCGGTGCTCAAAA 2336 10258 0.23AGTCGTTTTAAGCAA GTGAGACATGACT SEQ ID NO: 73:CACACACACAGCTCG61011 159310 1797 88.65TTACAAATAAGGAGG TAAAA SEQ ID NO: 74:TCCCACCTACGAGAA CCCCACGTGCACCCA48217 0 2032 0.00CGTGTAGCCGAAAAG TAGGTTCTAGAATTAT CATGTTAACTACTG SEQ ID NO: 75:GCGCGCTCGGGAGTT47371 ACGGAACCCACGGCT 19741 11781 1.68TGAAAGTACACTATA GAAGGTTGA SEQ ID NO: 76:CCCAATGCGCCCTTCC ATCCCCTTTTTCCCAA56157 AACCCGCCATATGGC 11 7825 0.00GGGCCCCGATACCCT GCGCAAATCGCGCAC AGC SEQ ID NO: 77:41897 ATAAGCACAGGAAAT 112 9857 0.01CCATCCTGTGCTTATTAttorney Docket No. 0073605-001135 TGGTCCATTGGTGTA GTAG SEQ ID NO: 78:CTGATTTTTTTTTTTTT TTTCCCGCCATATGGC3778 1252 7783 0.16 GGGGCTTGGCGCTCG GTCGTCGGGGGCGTG CTC SEQ ID NO: 79:ACAGGATGTTTCCTGT CTCATTTAATAGTTCG51103 8670 2452 3.54GACCCCCTCCACACG GAGGGAGTTTTGTGG AGGGTCTTTHigh BinSEQ ID NO: 80:AAAAAAGAAAAAAG AAAAAAGAACCCGCC19879 227 8448 0.03ATATGGCGGGAAGAC GGGGCGGCACGGGGC GGAACCCGGC SEQ ID NO: 81:AGCAAACGTTTTTGCT44652 CTGACATACAACACT 6309 8529 0.74GTATGGTAAAAAGTG TAACCAAAGGATTTT SEQ ID NO: 82:GGGGGGGGGTGGGGG GGGGTGCCCGCCATA29218 4082 1071 3.81TGGCGGGGAGGGGGG AAATTAATAGGTTAA GATAGGTAttorney Docket No. 0073605-001135 SEQ ID NO: 83:CGCATTTTTCTTTTTC TTTTCCCGCCATATGG12130 138 7330 0.02CGGGTCTCGGTTTCTT CGTTTCCTGGTTCGTT CC SEQ ID NO: 84:GCGCCCACAAGGGCG CGGGAGGGAGGGCCC33621 1534 2591 0.59GCCATATGGCGGGAG GGTGGGAATTTGGGT ATGTGAGATTTTA SEQ ID NO: 85:CCTGTTTTTCTTTTTCT TTTCCCGCCATATGGC12215 2038 12789 0.16GGGTTCTTTCGTCCTG CGTCTGGGGGGTGTC CT SEQ ID NO: 86:CCTAAAAAAACCCGC16289 CATATGGCGGGGAAC 2245 10258 0.22GAAGAACCGAACAAA CCCCAACAAGA SEQ ID NO: 87:CCGCGCCCACAAGGG CGCCCACCACCACCC32850 28046 1397 20.08CCGCCATATGGCGGG CCTTTCATCTATTAAC TCAGGAGGAGGAGG SEQ ID NO: 88:45303 GCTAACAGATTAGCC 11744 7656 1.53AAAACAGTTCAGATGAttorney Docket No. 0073605-001135 AACTTCGAGACTCGG GTCTCCAGATTGACA CGATAGGATACTAAT SEQ IDNO: 89:TTGTTTTTTTCCCGCC23569 1903 14837 0.13ATATGGCGGGTTGTG GCGTTTTGGCGCTGG GGGGTCTCGG SEQ IDNO: 90:ATCCTAACACCTTAG GTATGCCTAAGAAAA47944 7479 1493 5.01TAGGATAGAGAGTAA CACCGGGGAGTTTTG C SEQ ID NO: 91:GCGCCCACAAGGGCG CCCTCCTCCTCCTCCC34182 46 1019 0.05GCCATATGGCGGGTC TTCTCTCTTTAAACAT CTCCTCACCACCTable 5: List of identified 5’ UTRs (sfGFP)Designed 5’ halfRNA RNA RNA RNA RNA k_r2 Index UTR DNA k fit life _T0 _T2 _T4 _T8 _T16Sequence (min)Low BinSEQ ID NO:92:10,12736 ACTCACCC 4129 7059 5442 2760 4943 0.7 1.0 1.0075GCCATATG GCGGGAAAttorney Docket No. 0073605-001135 CCCGACG AAAACAA GGAACCC ACCGCAA SEQ ID NO:93:AGTTGCAC CGCCATAT GGCGGTG 12,37662 6041 6835 7388 4059 8585 0.8 0.9 0.9CAACTCAC 630CGACTCAT ATCGTAGG GGGCGGG TAAA SEQ ID NO:94:AGGACCA TATGGTCC25, 23, 22, 10, 23,35639 TACCTATG 8741 0.5 1.4 0.9179 512 066 993 716GTAATGG GGGAAAG GCGTATAA C SEQ ID NO:95:GGGGACC CGCCATAT48, 90, 201, 276, 172, 369,25610 GGCGGGG 0.1 13.0 0.6898 153 188 434 114 883ATTGGTTT GTTGATTG ATTAAAGT GGGATAttorney Docket No. 0073605-001135 SEQ ID NO:96:AGTCTCAA CTGCATGC ATGTCAAA TATGTTTG 45, 109, 49, 24, 19,48070 9609 1.0 0.7 1.0ACTTAACC 934 636 428 150 679TAAGTTGA TTTCAGGA ACACGTA GCAAAAA GAT SEQ ID NO:97:GACCATTT TTGTAAGT TTGTAAAC 14, 56, 97, 13, 60, 63,47002 0.1 5.1 0.8TTAGTGGT 423 486 347 5593 663 373CTTAACAA ACGGAAA TAGAGGT ACTG SEQ ID NO:98:CCCTAGTG ACCCCCAT CCCCCGAA 14,57094 4737 3120 2868 1274 3041 0.7 1.0 1.0CAACCCTT 837AACCCGCC ATATGGCG GGGCACC GCAAATTAAttorney Docket No. 0073605-001135 CAACACC ATAACCCA GA SEQ ID NO:99:ACCGAGC GCCGTATC 13, 32, 37, 38, 12, 19,60486 0.4 1.8 0.9CCACTACT 080 879 472 151 708 309GCTTACGT AAGGAGG TAATTGG SEQ ID NO:100:AAGCGCC CACAAGG GCGCAAA AAAAGAA AAAAAGC31174 1155 984 208 399 38 576 1.3 0.5 1.0CCGCCATA TGGCGGG GCAGCCC AGAAGGC AACACCC GACCAAG CCMedium BinSEQ ID NO:101:60996 CACACAC 61 1561 5116 5805 3702 4965 0.0 17.2 0.7ACAGCTG GTTTCAAAAttorney Docket No. 0073605-001135 TAAGGAG GTAAAA SEQ ID NO:102:CATCGCAC CGGGCCC GGTGCACT223, 20, 18, 10,51425 ACTTAAGT 8563 6817 0.5 1.4 0.916 866 090 103AACTGAA CACCAATA AAACCAT AGGACAC GCGTC SEQ ID NO:103:TTTGTTTG TTTGTTTG TTTGTTTG TTTGCCCG 12, 20, 19, 16, 11,22556 7248 0.5 1.4 0.9CCATATGG 470 406 420 814 599CGGGGTCC GGGCGGT GCGTGGTA GGAGGAG GAGG SEQ ID NO:104:AAAACAA17569 AACCCGCC 2897 6354 5293 5282 2239 1934 0.5 1.3 0.9ATATGGCG GGACCAC CACCAGCAttorney Docket No. 0073605-001135 GGCGAAC GGGGAAC GGAC SEQ ID NO:105:AGGGCAG GCACATGT GCCTGGTC13, 13, 135,51436 GAGGGAC 304 2550 8480 0.0 -0.1747 368 82TTATTGTG TACGATTT TTAAGGA GAAGCAG TTGCT SEQ ID NO:106:TAGCCGTC TGGACAA CAGTATCG CGTATCGG15, 32, 47, 31, 58,48583 GTGTCATA 2888 0.1 12.2 0.5846 413 972 012 844TATTGACG AGGCTAA CTAAACA GAACAGG GGGAGGG TAA SEQ ID NO:107:142733696 GGGCGCC 1112 3274 4993 6367 7398 0.1 11.4 -0.20CACAAGG GCGCGGGAttorney Docket No. 0073605-001135 GAGGGCC CGCCATAT GGCGGGG TAGTATAG TGAGTATA TAAGGAG AAGAGA SEQ ID NO:108:TAACTCTA CGCTATCA42624 TGGGCATG 1472 5387 4583 3482 1946 3654 0.6 1.2 0.9ATAATACA CAACTTAA CGGACCA CTTCT SEQ ID NO:109:CTCGCTGC TGATGAGT GATGACG GCCATCAC36, 63, 10, 78, 14,61770 GAAACTA 5802 0.1 12.3 0.2611 179 0391 324 3028CCGTACTG CGGTAGTC CAGCGCA CATTAAGG AGGTAGG CA SEQ ID NO:29, 42, 61, 38, 69,43664 110: 6513 0.1 5.7 0.5013 841 655 681 389GCATCCACAttorney Docket No. 0073605-001135 GTCCCAAC ACAATTAG GATCATTG GGACGTG GATGCAA GCAAACG ACGACTG GGAGGGG AAHigh BinSEQ ID NO:111:GCGCCCAC AAGGGCG CGGGGGT GGGGGCC 10, 18, 37, 25, 73,34993 1987 0.0 38.9 -0.1CGCCATAT 005 962 317 150 090GGCGGGG AGTTTTTG TAGTTTTA GGGGGGGA ATTTTT SEQ ID NO:112:GTACAAA AAACCCG CCATATGG15518 474 1437 3152 4515 1672 1242 0.1 6.7 0.8CGGGGAA CCCCACCG ACCGAGA GAAGGAG CAACCAttorney Docket No. 0073605-001135 SEQ ID NO:113:AAAAGAA AAGAAAA CCCGCCAT10, 13, 16,19415 ATGGCGG 2163 6982 6699 0.2 2.8 0.8173 146 771GGGACGA GGAGAAC AAGGGGG CGCAGCG GCG SEQ ID NO:114:GGGGGGG AGGGGGG CCCGCCAT18, 34, 21, 52,26257 ATGGCGG 2397 8756 0.0 29.4 0.2937 692 797 851GATTGAG AGTTAAA AAAAAGA GGAGGGA AAA SEQ ID NO:115:AACAACA ACACCCGC CATATGGC17160 349 985 3884 5117 802 711 0.1 10.1 0.5GGGACAC AGAAAGA CACACGA AGGGCAA GCAAAAttorney Docket No. 0073605-001135 SEQ ID NO:116:CGGAGCA CCGCCATA TGGCGGTG 14, 31, 54, 34, 88,39222 5987 0.0 24.7 0.2CTCCGCGT 884 862 157 742 844TTCTTAAA CAAAGGC GAGGATG GATGG SEQ ID NO:117:TACGGGC AACCGTAT GTGTAAGT ACACTTAC 12, 14, 13,45582 8377 6731 8879 0.4 1.9 0.9AATCCGCG 424 646 779CGTACGCG GCTGTATA AATAATCT AGGGTCTT AA SEQ ID NO:118:AAACTGTC CATGCATG GACACAC 11, 10,50194 1040 4954 8706 7785 0.1 7.0 0.6ATCTTAAT 509 943ACAGCGT AACGGTA CGGAGGA GGCGTAttorney Docket No. 0073605-001135 SEQ ID NO:119:GCGCCCAT ATGGGCG36813 CATCTAGC 3870 5045 3137 2764 1004 7300 0.7 0.9 0.9AGTACCAC GGGTTGG GTACAGG A SEQ ID NO:120:GTCAAGA GTCAGGCC CTACAACA GTAGGGA19, 32, 56, 23, 39,45101 TAATATAC 2534 0.1 7.0 0.7020 994 617 429 336AGTCGCG ATCGACTG AGCAGTG GAACAATT CGGAGAA GGT SEQ ID NO:121:AGCCCGCC ATATGGCG18669 GGATGGG 3093 6760 7801 6044 3084 5191 0.4 1.7 0.9AAGGGTT GATGAAG GAGTAGG TGGTAttorney Docket No. 0073605-001135 Table 6: Translation initiation rates (TIR; sfGFP) corresponding to 5’ UTR sequences of Table 5 Designed 5’ UTR sfGFP RBS TIR mRFP1 RBS TIR Ratio IndexSequence (au) TIR (au) (sfGFP / mRFP1)Low binSEQ ID NO: 92:ACTCACCCGCCATAT12736 GGCGGGAACCCGACG 12915 8448 1.53AAAACAAGGAACCCA CCGCAA SEQ ID NO: 93:AGTTGCACCGCCATA37662 TGGCGGTGCAACTCA 12764 4474 2.85CCGACTCATATCGTA GGGGGCGGGTAAA SEQ ID NO: 94:AGGACCATATGGTCC35639 TACCTATGGTAATGG 18313 4694 3.90GGGAAAGGCGTATAA C SEQ ID NO: 95:GGGGACCCGCCATAT25610 GGCGGGGATTGGTTT 0 1819 0.00GTTGATTGATTAAAG TGGGAT SEQ ID NO: 96:AGTCTCAACTGCATG CATGTCAAATATGTTT48070 1340 1461 0.92GACTTAACCTAAGTT GATTTCAGGAACACG TAGCAAAAAGAT SEQ ID NO: 97:47002 9375 746 12.57GACCATTTTTGTAAGTAttorney Docket No. 0073605-001135 TTGTAAACTTAGTGGT CTTAACAAACGGAAA TAGAGGTACTG SEQ ID NO: 98:CCCTAGTGACCCCCA TCCCCCGAACAACCC57094 TTAACCCGCCATATG 0 1128 0.00GCGGGGCACCGCAAA TTACAACACCATAAC CCAGA SEQ ID NO: 99:ACCGAGCGCCGTATC60486 2085 1413 1.48CCACTACTGCTTACGT AAGGAGGTAATTGG SEQ ID NO: 100:AAGCGCCCACAAGGG CGCAAAAAAAGAAA31174 AAAAGCCCGCCATAT 91 9285 0.01GGCGGGGCAGCCCAG AAGGCAACACCCGAC CAAGCCMedium BinSEQ ID NO: 101:CACACACACAGCTGG60996 404418 3283 123.19TTTCAAATAAGGAGG TAAAA SEQ ID NO: 102:CATCGCACCGGGCCC GGTGCACTACTTAAG51425 8201 5818 1.41TAACTGAACACCAAT AAAACCATAGGACAC GCGTCAttorney Docket No. 0073605-001135 SEQ ID NO: 103:TTTGTTTGTTTGTTTG TTTGTTTGTTTGCCCG22556 11714 2684 4.36CCATATGGCGGGGTC CGGGCGGTGCGTGGT AGGAGGAGGAGG SEQ ID NO: 104:AAAACAAAACCCGCC17569 ATATGGCGGGACCAC 5661 12393 0.46CACCAGCGGCGAACG GGGAACGGAC SEQ ID NO: 105:AGGGCAGGCACATGT GCCTGGTCGAGGGAC51436 6121 2964 2.07TTATTGTGTACGATTT TTAAGGAGAAGCAGT TGCT SEQ ID NO: 106:TAGCCGTCTGGACAA CAGTATCGCGTATCG48583 GGTGTCATATATTGA 29794 6872 4.34CGAGGCTAACTAAAC AGAACAGGGGGAGG GTAA SEQ ID NO: 107:GGGCGCCCACAAGGG CGCGGGGAGGGCCCG33696 72804 11174 6.52CCATATGGCGGGGTA GTATAGTGAGTATAT AAGGAGAAGAGA SEQ ID NO: 108:42624 5223 6165 0.85TAACTCTACGCTATCAttorney Docket No. 0073605-001135 ATGGGCATGATAATA CACAACTTAACGGAC CACTTCT SEQ ID NO: 109:CTCGCTGCTGATGAG TGATGACGGCCATCA61770 CGAAACTACCGTACT 69764 1838 37.96GCGGTAGTCCAGCGC ACATTAAGGAGGTAG GCA SEQ ID NO: 110:GCATCCACGTCCCAA CACAATTAGGATCAT43664 7863 4631 1.70TGGGACGTGGATGCA AGCAAACGACGACTG GGAGGGGAAHigh BinSEQ ID NO: 111:GCGCCCACAAGGGCG CGGGGGTGGGGGCCC34993 12776 5789 2.21GCCATATGGCGGGGA GTTTTTGTAGTTTTAG GGGGGAATTTTT SEQ ID NO: 112:GTACAAAAAACCCGC15518 CATATGGCGGGGAAC 14421 20390 0.71CCCACCGACCGAGAG AAGGAGCAACC SEQ ID NO: 113:AAAAGAAAAGAAAA19415 17939 8448 2.12CCCGCCATATGGCGG GGGACGAGGAGAACAttorney Docket No. 0073605-001135 AAGGGGGCGCAGCGG CG SEQ ID NO: 114:GGGGGGGAGGGGGG CCCGCCATATGGCGG26257 154370 10258 15.05GATTGAGAGTTAAAA AAAAGAGGAGGGAA AA SEQ ID NO: 115:AACAACAACACCCGC17160 CATATGGCGGGACAC 15333 1593 9.63AGAAAGACACACGAA GGGCAAGCAAA SEQ ID NO: 116:CGGAGCACCGCCATA39222 TGGCGGTGCTCCGCG 15390 300 51.36TTTCTTAAACAAAGG CGAGGATGGATGG SEQ ID NO: 117:TACGGGCAACCGTAT GTGTAAGTACACTTA45582 5183 4310 1.20CAATCCGCGCGTACG CGGCTGTATAAATAA TCTAGGGTCTTAA SEQ ID NO: 118:AAACTGTCCATGCAT50194 GGACACACATCTTAA 10287 1215 8.47TACAGCGTAACGGTA CGGAGGAGGCGT SEQ ID NO: 119:36813 GCGCCCATATGGGCG 4505 1711 2.63CATCTAGCAGTACCAAttorney Docket No. 0073605-001135 CGGGTTGGGTACAGG A SEQ ID NO: 120:GTCAAGAGTCAGGCC CTACAACAGTAGGGA45101 TAATATACAGTCGCG 6121 7411 0.83ATCGACTGAGCAGTG GAACAATTCGGAGAA GGT SEQ ID NO: 121:AGCCCGCCATATGGC18669 GGGATGGGAAGGGTT 9857 5241 1.88GATGAAGGAGTAGGT GGTTable 7: List of identified 5’ UTRs (sfGFP / mRFP1)halfDesigned 5’ UTR RNA RNA RNA RNA RNA k_r2 Index DNA k_fit half-life (min) k_r2_fit Sequence _T0 _T2 _T4 _T8 _T16 SEQ IDNO: 122:TTCGAAAAAC AAAAACAAAA CCCGCCATAT4850 492 1602 854 719 449 782 0.8 0.8 1.0 GGCGGGCGGC GGAGGGCCAG GAGGGAACAG CGCGAC SEQ IDNO: 123:GAGGTTTTTTT1687 2336 3275 1738 26243159 TTTTTTTTTCC 5715 0.2 4.1 0.75 1 2 9 9CGCCATATGG CGGGCCTTGTAttorney Docket No. 0073605-001135 CTTGTGCGCG GTAGGAGGAG GAGG SEQ IDNO: 124:AAGAAGAAGA AGAAGAAGAA GCCCGCCATA19047 415 1976 3596 2874 2211 619 0.2 3.7 0.9 TGGCGGGGAG CGGGCAGCCA CGAGGAGGCA GAGAACG SEQ IDNO: 125:TTTGTTTGTTT GTTTCCCGCCA1862 2917 4316 2637 507822474 TATGGCGGGC 4683 0.1 6.7 0.5 CTTTTTTCGTT 0 1 8 4 8TTGCTGAGGA GGAGGAGG SEQ ID NO: 126:GGAGGACCCG CCATATGGCG1911 4783 6342 6459 2804 400025191 GGGTGGAGTG 0.3 2.4 0.91 2 9 7 8 8TGAGGTTTTG GGGTGGAATT GT SEQ IDNO: 127:CACTCCATAT GGAGTGGTAA 1701 2552 4259 1856 382136958 3105 0.1 5.6 0.6 ACAATGGGGG 1 7 9 5 6AGGTTATAAC AGGAATAttorney Docket No. 0073605-001135 SEQ ID NO: 128:GTTGGAACCC CAACCAGTAC CGCTATCGGT1642 1930 1166 260045333 AAAGTTCACA 1372 8237 0.1 7.6 0.76 0 4 1TATGAACATA ATATCAATTA GGGAGAAAAA GT SEQ IDNO: 129:TACCCGGAAT GGTAAGCGTC ACTTATTAGTG120548923 ACTATTACCA 4074 6833 5388 2118 4777 0.8 0.9 1.06GAACGGGTAG ACACGGCGCA AAGAGGACAT ACACTable 8: Translation initiation rates (TIR; sfGFP / mRFP1) corresponding to 5’ UTR sequences of Table 7Designed 5’ UTR sfGFP RBS TIR mRFP1 RBS TIR TIR Ratio IndexSequence (au) (au) (sfGFP / mRFP1) SEQ ID NO: 122:TTCGAAAAACAAA AACAAAACCCGCC4850 ATATGGCGGGCGG 3100 740 4.19CGGAGGGCCAGGA GGGAACAGCGCGA C SEQ IDNO: 123:3159 13109 6132 2.14GAGGTTTTTTTTTTAttorney Docket No. 0073605-001135 TTTTTTTCCCGCCAT ATGGCGGGCCTTG TCTTGTGCGCGGT AGGAGGAGGAGG SEQ IDNO: 124:AAGAAGAAGAAG AAGAAGAAGCCCG19047 CCATATGGCGGGG 18550 865 21.45AGCGGGCAGCCAC GAGGAGGCAGAG AACG SEQ IDNO: 125:TTTGTTTGTTTGTT TCCCGCCATATGG22474 10280 6633 1.55CGGGCCTTTTTTCG TTTTGCTGAGGAG GAGGAGG SEQ IDNO: 126:GGAGGACCCGCCA25191 TATGGCGGGGTGG 3451 76 45.41AGTGTGAGGTTTT GGGGTGGAATTGT SEQ IDNO: 127:CACTCCATATGGA36958 GTGGTAAACAATG 24116 2386 10.11GGGGAGGTTATAA CAGGAAT SEQ IDNO: 128:GTTGGAACCCCAA45333 CCAGTACCGCTAT 50489 6404 7.88CGGTAAAGTTCAC ATATGAACATAATAttorney Docket No. 0073605-001135 ATCAATTAGGGAG AAAAAGT SEQ ID NO: 129:TACCCGGAATGGT AAGCGTCACTTAT48923 TAGTGACTATTAC 5171 101 51.20CAGAACGGGTAGA CACGGCGCAAAGA GGACATACACTable 9: List of identified 5’ UTR regions (sfGFP + mRFP1)Designed 5’ halfRNA RNA RNA RNA RNA k_r2_ Index UTR DNA k_fit lifeTO _T2 _T4 _T8 _T16 fit Sequence (min) SEQ ID NO:130:TGGGTTTTTTTTCCCG2901 CCATATGG 16558 15145 9719 7448 3089 8035 0.7 0.9 1.0 CGGGCGC GGGGGTG GCTTCCTG TCGGGTG GTTGT SEQ ID NO:131:AGCATTTT8546 TGTTTTTG 10089 13485 4994 4129 2240 4347 1.0 0.7 1.0 TTTTCCCG CCATATGG CGGGTTTGAttorney Docket No. 0073605-001135 CGGTGTG GCCCCTCC GCGTTTGC TGG SEQ ID NO:132:GCTCAAA AAGAAAA AGAAAAC CCGCCATA7615 4257 19152 26584 28743 10839 14922 0.3 2.6 0.9 TGGCGGG ACACGAG GCAGCGG ACGAGGA GGCCGAC GG SEQ ID NO:133:AGAATTTT TCTTTTTC TTTTCCCG10503 CCATATGG 742 4943 7658 9110 5021 7917 0.2 4.1 0.7 CGGGCGC TCGTTTTG TGTGTTGA GGAGGAG GAGG SEQ ID NO:134:29516 TTGCGCCC 48571 19142 13388 11309 7510 14969 0.7 1.0 0.9 ACAAGGG CGCTTTTTAttorney Docket No. 0073605-001135TCCCGCCA TATGGCG GGCGGCC GGGGGTG GCTGCTGG CGCTTCTG GG SEQ ID NO:135:ATAGAAA GCACACA43742 TTTACATG 13901 49857 45154 42865 14265 24494 0.5 1.4 0.9 CTTTAACA GTACAGG AACCGGA GGCAGAC SEQ ID NO:136:TATGTCGA TACAGTTA CTCACCTG CAAATTTC48354 AATCACG 2741 15435 22188 24829 8309 8619 0.3 2.7 0.9 AGAATAG TGACATA GTACCGTA ATAGATA GGATGAA TTA SEQ ID NO:50185 298 1450 3793 4862 2304 4290 0.1 10.8 0.8 137:Attorney Docket No. 0073605-001135 GGGGAAT ATGGATA CATATTAA GGTTGCAC AGTCAAA CTCAAAG GAGTAAA AAGAAT SEQ ID NO:138:CACACAC61236 ACTGCAC 301 823 1631 1638 2534 3625 0.0 17.9 0.0 GTTTCAAA TAAGGAG GTAAAA SEQ ID NO:139:CACACAC61087 ACACCAG 894 5269 8572 13439 12106 19749 0.1 12.2 0.1 GATACAA ATAAGGA GGTAAAATable 10: Translation initiation rates (TIR; sfGFP + mRFP1) corresponding to 5’ UTR sequences of Table 9Designed 5’ UTRIndex sfGFP RBS TIR mRFP1 RBS TIR (au)SequenceSEQ ID NO: 130:TGGGTTTTTTTTT2901 TTTTTTTCCCGCC 1093 2113 ATATGGCGGGCG CGGGGGTGGCTTAttorney Docket No. 0073605-001135 CCTGTCGGGTGGT TGT SEQ ID NO: 131:AGCATTTTTGTTT TTGTTTTCCCGCC8546 ATATGGCGGGTTT 14 2252 GCGGTGTGGCCC CTCCGCGTTTGCT GG SEQ ID NO: 132:GCTCAAAAAGAA AAAGAAAACCCG7615 CCATATGGCGGG 33343 4652 ACACGAGGCAGC GGACGAGGAGGC CGACGG SEQ ID NO: 133:AGAATTTTTCTTT TTCTTTTCCCGCC10503 ATATGGCGGGCG 17963 10258 CTCGTTTTGTGTG TTGAGGAGGAGG AGG SEQ ID NO: 134:TTGCGCCCACAA GGGCGCTTTTTTT TTTTTTTCCCGCC29516 3 6004ATATGGCGGGCG GCCGGGGGTGGC TGCTGGCGCTTCT GGGAttorney Docket No. 0073605-001135 SEQ ID NO: 135:ATAGAAAGCACA CATTTACATGCTT43742 36249 2191TAACAGTACAGG AACCGGAGGCAG AC SEQ ID NO: 136:TATGTCGATACAG TTACTCACCTGCA AATTTCAATCACG48354 7293 4329AGAATAGTGACA TAGTACCGTAATA GATAGGATGAAT TA SEQ ID NO: 137:GGGGAATATGGA TACATATTAAGGT50185 7160 10258TGCACAGTCAAA CTCAAAGGAGTA AAAAGAAT SEQ ID NO: 138:CACACACACTGC61236 515674 3082ACGTTTCAAATAA GGAGGTAAAA SEQ ID NO: 139:CACACACACACC61087 1212646 4208AGGATACAAATA AGGAGGTAAAASupporting Data 3: Structural Protein (SP) AnalysisTable 11: Plasmid element sequencesAttorney Docket No. 0073605-001135 T7 Promoter + SEQ ID NO: 140lac operator TAATACGACTCACTATAGGGGAATTGTGAGCGGATAACAATTCC SEQ ID NO: 141 ATGATAAGTAAAGGTGAAGAAGATAACCTGGCCATTATTAAAGAGT TTATGCGCTTTAAGGTGCATATGGAGGGTTCCGTTAATGGGCATGAG TTTGAGATAGAAGGTGAGGGTGAAGGACGTCCTTATGAAGGTACTC AAACAGCAAAACTGAAAGTGACGAAAGGTGGACCTCTTCCCTTTGC ATGGGATATCCTTTCACCGCAATTTATGTACGGTAGTAAAGCATACG TGAAACATCCTGCGGATATTCCGGATTATCTTAAGTTATCGTTTCCG GAGGGGTTTAAATGGGAACGCGTGATGAATTTTGAAGATGGTGGGGmCherry CDS TCGTCACAGTGACGCAGGACTCCTCTCTTCAGGACGGTGAGTTCATT TATAAAGTGAAACTCCGTGGAACGAATTTTCCTTCTGACGGTCCAGT GATGCAGAAAAAGACAATGGGATGGGAAGCTTCTTCAGAACGGATG TATCCTGAAGATGGAGCCTTAAAAGGTGAAATAAAACAGAGGCTGA AATTAAAGGATGGAGGACATTACGACGCGGAAGTGAAAACAACCTA TAAAGCAAAAAAACCTGTTCAGTTACCTGGTGCGTACAATGTGAAC ATCAAACTCGACATAACGTCTCACAACGAGGATTACACAATCGTAG AGCAGTATGAAAGAGCTGAAGGGCGACATAGTACTGGGGGTATGGA CGAACTTTACAAA SEQ ID NO: 142Terminator ACTCGAGAGATAACAGATACTTCGGTATCTGTTATCTGTTTTTTTTCAACAGATAGCCGCGTTCGCGCGGCTATCTGTTTTTTTTTable 12: Cementl SequencesSEQ IDNO: 143 MSGPVKDDDEYKETREYTKEYTVDRNKGYNRKGYGDDVSAKETFLRT TEYNDKDGYGSKNDGKQVKESYTREYEIDKDGYGKDDDVTGKETYRR Amino acid EVTVDDDGYSPRGRYGRTDAYGLNDGYVRKDAYGLRRGPGLGVYPG PGLGRKDVYPATGLGRKDVYPVTGLGRKDVYPATGLGRKDVYPATGL GRKDVYPATGLGRKDVYPVPGLGRKDVYPGPGLGRKGVYPGRGFGVP GPFW CDS* SEQ ID NO: 144Attorney Docket No. 0073605-001135 ATGTCCGGTCCGGTAAAAGATGACGATGAGTACAAAGAGACACGAG AGTACACCAAGGAATACACCGTGGATCGCAATAAAGGCTACAATAG GAAAGGCTATGGCGATGATGTGTCCGCGAAAGAAACATTTCTGAGA ACGACAGAATACAACGATAAAGATGGGTACGGCTCTAAGAACGATG GGAAACAGGTAAAAGAAAGCTATACCAGGGAATATGAGATTGATA AAGATGGTTACGGTAAAGACGATGATGTCACCGGGAAGGAAACTTA CCGGCGAGAAGTAACTGTCGACGATGATGGCTACTCTCCGAGAGGC AGATACGGTCGTACAGACGCATACGGTTTAAACGACGGCTACGTCC GAAAAGATGCGTACGGTTTGCGGAGAGGCCCCGGTCTGGGCGTGTA CCCTGGTCCCGGTTTGGGGCGTAAAGATGTGTACCCGGCCACAGGA TTGGGTCGTAAGGATGTTTACCCCGTCACCGGATTAGGAAGGAAAG ATGTATACCCTGCTACAGGATTGGGCCGTAAAGATGTATACCCGGC GACAGGCTTGGGCAGGAAAGACGTGTACCCCGCTACAGGCCTGGGT AGAAAAGATGTGTACCCTGTGCCTGGCTTGGGTAGGAAGGATGTGT ACCCCGGACCGGGTTTAGGTCGTAAAGGTGTATACCCTGGAAGGGG ATTTGGTGTTCCGGGACCCTTCTGGTTAGGTTCATCAGGTGGTTCAT CTGGAGTGTCAGGGTGGAGACTCTTTAAAAAGATTTCTGGATGA SEQ ID NO: 145RBS TAAAAAATAGAAAGTTCCCCACGATAAGAGGTTTTAT SEQ ID NO: 146 TAATACGACTCACTATAGGGGAATTGTGAGCGGATAACAATTCCTA AAAAATAGAAAGTTCCCCACGATAAGAGGTTTTATATGTCCGGTCC GGTAAAAGATGACGATGAGTACAAAGAGACACGAGAGTACACCAA GGAATACACCGTGGATCGCAATAAAGGCTACAATAGGAAAGGCTAT GGCGATGATGTGTCCGCGAAAGAAACATTTCTGAGAACGACAGAATWhole plasmid ACAACGATAAAGATGGGTACGGCTCTAAGAACGATGGGAAACAGGT AAAAGAAAGCTATACCAGGGAATATGAGATTGATAAAGATGGTTAC GGTAAAGACGATGATGTCACCGGGAAGGAAACTTACCGGCGAGAA GTAACTGTCGACGATGATGGCTACTCTCCGAGAGGCAGATACGGTC GTACAGACGCATACGGTTTAAACGACGGCTACGTCCGAAAAGATGC GTACGGTTTGCGGAGAGGCCCCGGTCTGGGCGTGTACCCTGGTCCCG GTTTGGGGCGTAAAGATGTGTACCCGGCCACAGGATTGGGTCGTAAAttorney Docket No. 0073605-001135 GGATGTTTACCCCGTCACCGGATTAGGAAGGAAAGATGTATACCCT GCTACAGGATTGGGCCGTAAAGATGTATACCCGGCGACAGGCTTGG GCAGGAAAGACGTGTACCCCGCTACAGGCCTGGGTAGAAAAGATGT GTACCCTGTGCCTGGCTTGGGTAGGAAGGATGTGTACCCCGGACCG GGTTTAGGTCGTAAAGGTGTATACCCTGGAAGGGGATTTGGTGTTCC GGGACCCTTCTGGTTAGGTTCATCAGGTGGTTCATCTGGAGTGTCAG GGTGGAGACTCTTTAAAAAGATTTCTGGATGATAAGTAAAGGTGAA GAAGATAACCTGGCCATTATTAAAGAGTTTATGCGCTTTAAGGTGCA TATGGAGGGTTCCGTTAATGGGCATGAGTTTGAGATAGAAGGTGAG GGTGAAGGACGTCCTTATGAAGGTACTCAAACAGCAAAACTGAAAG TGACGAAAGGTGGACCTCTTCCCTTTGCATGGGATATCCTTTCACCG CAATTTATGTACGGTAGTAAAGCATACGTGAAACATCCTGCGGATAT TCCGGATTATCTTAAGTTATCGTTTCCGGAGGGGTTTAAATGGGAAC GCGTGATGAATTTTGAAGATGGTGGGGTCGTCACAGTGACGCAGGA CTCCTCTCTTCAGGACGGTGAGTTCATTTATAAAGTGAAACTCCGTG GAACGAATTTTCCTTCTGACGGTCCAGTGATGCAGAAAAAGACAAT GGGATGGGAAGCTTCTTCAGAACGGATGTATCCTGAAGATGGAGCC TTAAAAGGTGAAATAAAACAGAGGCTGAAATTAAAGGATGGAGGA CATTACGACGCGGAAGTGAAAACAACCTATAAAGCAAAAAAACCTG TTCAGTTACCTGGTGCGTACAATGTGAACATCAAACTCGACATAACG TCTCACAACGAGGATTACACAATCGTAGAGCAGTATGAAAGAGCTG AAGGGCGACATAGTACTGGGGGTATGGACGAACTTTACAAATAATG AAC TC GAGAGAT AAC AGATAC TTCGGT ATCTGTT ATC TGTTTTTTTTC AACAGATAGCCGCGTTCGCGCGGCTATCTGTTTTTTTTTCTAGAGAG ATCCCGGACACCATCGAGACACAGACGGATAGAAGGCACGTACCTA ATCGGAGTGACTCTACTTGCCATGGTGCACAGGACGTTTCCTTAGGG GGTTCTCCCATGAAACCAGTGACACTTTACGATGTCGCTGAGTATGC AGGAGTGTCTTACCAAACCGTATCTCGCGTCGTTAATCAGGCGAGTC ATGTATCAGCGAAAACCCGGGAAAAGGTAGAGGCTGCCATGGCCGA ACTCAATTACATCCCCAATCGTGTTGCTCAGCAATTAGCCGGGAAGC AGAGCTTGTTAATTGGAGTAGCGACCTCTTCCTTAGCCCTCCATGCG CCTTCACAAATAGTTGCGGCAATAAAGAGTAGGGCCGATCAATTGGAttorney Docket No. 0073605-001135 GAGCTTCAGTGGTCGTCTCCATGGTCGAGCGATCTGGGGTAGAAGC GTGCAAGGCAGCGGTACACAATTTGTTAGCGCAAAGGGTTTCCGGG TTAATAATCAATTATCCATTGGACGATCAGGATGCGATAGCGGTGG AAGCTGCATGCACCAATGTTCCGGCATTGTTCTTAGATGTGAGCGAT CAAACTCCAATCAACTCTATCATTTTCTCCCACGAGGACGGGACAAG GCTCGGCGTCGAACACTTGGTTGCCCTGGGTCACCAGCAAATAGCCC TTTTAGCAGGACCACTATCGTCCGTTTCGGCTCGGTTGCGTCTAGCG GGTTGGCACAAGTACCTGACTCGCAATCAAATTCAACCTATTGCCGA ACGCGAGGGAGACTGGTCAGCGATGTCGGGTTTCCAACAAACCATG CAGATGCTGAACGAAGGTATTGTTCCCACAGCCATGTTGGTTGCTAA CGACCAAATGGCCCTCGGCGCCATGCGAGCTATAACGGAGTCGGGA TTAAGGGTAGGCGCGGACATAAGTGTCGTGGGTTACGACGATACCG AGGACAGCAGTTGTTACATTCCACCTCTCACCACGATTAAGCAGGAT TTCAGACTACTGGGACAAACATCAGTAGACAGACTCCTGCAGTTGTC GCAGGGTCAAGCGGTAAAGGGAAACCAGTTACTACCTGTTTCTTTG GTTAAGAGAAAGACTACCTTGGCTCCCAACACGCAAACGGCAAGTC CGAGAGCTTTAGCGGACAGTTTGATGCAATTGGCGCGTCAGGTAAG CCGTCTGGAGAGTGGCCAATAATAACTCGGTACCAAATTCCAGAAA AGAGACGCTGAAAAGCGTCTTTTTTCGTTTTGGTCCGTTTTTCCATAG GCTCCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAG AGGTGGCGAAACCCGACAGGACTATAAAGATACCAGGCGTTTCCCC CTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCCGCTTACC GGATACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGCTTTCTCA TAGCTCACGCTGTAGGTATCTCAGTTCGGTGTAGGTCGTTCGCTCCA AGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCC TTATCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACACGACTT ATCGCCACTGGCAGCAGCCACTGGTAACAGGATTAGCAGAGCGAGG TATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTAACTACGG CTACACTAGAAGGACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAG TTACCTTCGGAAAAAGAGTTGGTAGCTCTTGATCCGGCAAACAAAC CACCGCTGGTAGCGTTGGTTTTTTTGTTTGCAAGCAGCAGATTACGC GCAGAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTCTACGGGAttorney Docket No. 0073605-001135 GTCTTTTTGTTATCAATAAAAAAGGCCCCCCTTTCGGGAGGCCTTAT TGTTCGTCAGTTGATCGGCACGTAAGAGAAGAGATAGAAAACTTCC ACAACTAGACAAATTCTCCCCTAACACGGGAGTATTCGAGGGGGCA GGCGGATAAGGTGGTTCTAGCATGGAGAAGAAGATAACCGGATACA CAACGGTTGACATTTCACAATGGCACCGGAAAGAGCACTTTGAAGC TTTCCAAAGCGTAGCCCAATGTACGTACAACCAAACGGTACAGTTA GACATCACGGCGTTCTTAAAGACCGTTAAGAAGAACAAGCACAAGT TCTACCCAGCCTTCATTCATATCCTCGCGAGGTTAATGAATGCCCAC CCCGAATTTCGGATGGCTATGAAAGACGGGGAATTAGTCATCTGGG ATAGTGTTCATCCCTGCTACACGGTGTTCCACGAGCAGACTGAAACA TTCAGTAGCTTGTGGTCTGAGTATCATGACGATTTTAGGCAATTCTT ACATATCTATAGCCAAGATGTGGCATGTTATGGAGAGAACCTTGCAT ACTTTCCTAAAGGGTTCATAGAGAACATGTTTTTCGTGTCGGCCAAT CCATGGGTGAGCTTCACTTCCTTCGACTTAAATGTGGCAAATATGGA TAACTTCTTCGCACCGGTCTTTACCATGGGCAAGTACTACACGCAGG GTGATAAGGTTCTGATGCCCCTAGCCATCCAGGTTCATCATGCTGTA TGCGACGGATTTCATGTGGGCCGCATGCTGAATGAGCTACAGCAGT ACTGCGATGAGTGGCAAGGTGGTGCGTGATAAGTCAGTTTCACCTGT TTTACGTAAAAACCCGCTTCGGCGGGTTTTTACTTTTGGAGCTCAGC ATCGCAAGAATTCTable 13: Reflectin Q3L8V4 SequencesSEQ IDNO: 147 MYYPERYFDMSNWQMDMQGRWMDMQGRYCSPYWYNWYGRQMYYAmino acidPYQNYYWYGRWDYPGMDYSNWQMDMQGRWMDMQGRYMDPWWM NDSYYNNYYN SEQ ID NO: 148 ATGTACTACCCCGAACGGTACTTTGACATGTCGAACTGGCAGATGG ATATGCAAGGCCGTTGGATGGATATGCAGGGTCGTTATTGCAGCCCC CDS*TACTGGTACAACTGGTATGGCCGGCAGATGTATTACCCTTACCAGAA TTATTACTGGTACGGTCGTTGGGACTACCCTGGAATGGATTACAGCA ATTGGCAAATGGATATGCAGGGTAGATGGATGGATATGCAGGGACGAttorney Docket No. 0073605-001135 CTATATGGATCCATGGTGGATGAATGATTCGTATTACAATAACTACT ACAATTTAGGTTCATCAGGTGGTTCATCTGGAGTGTCAGGGTGGAGA CTCTTTAAAAAGATTTCTGGATGA SEQ IDNO: 149RBS TAAAATCACTCCGAAAGGAAAATAAGGAGCCCTAAG SEQ IDNO: 150 AGAGAGATCCCGGACACCATCGAGACACAGACGGATAGAAGGCAC GTACCTAATCGGAGTGACTCTACTTGCCATGGTGCACAGGACGTTTC CTTAGGGGGTTCTCCCATGAAACCAGTGACACTTTACGATGTCGCTG AGTATGCAGGAGTGTCTTACCAAACCGTATCTCGCGTCGTTAATCAG GCGAGTCATGTATCAGCGAAAACCCGGGAAAAGGTAGAGGCTGCCA TGGCCGAACTCAATTACATCCCCAATCGTGTTGCTCAGCAATTAGCC GGGAAGCAGAGCTTGTTAATTGGAGTAGCGACCTCTTCCTTAGCCCT CCATGCGCCTTCACAAATAGTTGCGGCAATAAAGAGTAGGGCCGAT CAATTGGGAGCTTCAGTGGTCGTCTCCATGGTCGAGCGATCTGGGGT AGAAGCGTGCAAGGCAGCGGTACACAATTTGTTAGCGCAAAGGGTT TCCGGGTTAATAATCAATTATCCATTGGACGATCAGGATGCGATAGC GGTGGAAGCTGCATGCACCAATGTTCCGGCATTGTTCTTAGATGTGAWhole plasmid GCGATCAAACTCCAATCAACTCTATCATTTTCTCCCACGAGGACGGG ACAAGGCTCGGCGTCGAACACTTGGTTGCCCTGGGTCACCAGCAAA TAGCCCTTTTAGCAGGACCACTATCGTCCGTTTCGGCTCGGTTGCGT CTAGCGGGTTGGCACAAGTACCTGACTCGCAATCAAATTCAACCTAT TGCCGAACGCGAGGGAGACTGGTCAGCGATGTCGGGTTTCCAACAA ACCATGCAGATGCTGAACGAAGGTATTGTTCCCACAGCCATGTTGGT TGCTAACGACCAAATGGCCCTCGGCGCCATGCGAGCTATAACGGAG TCGGGATTAAGGGTAGGCGCGGACATAAGTGTCGTGGGTTACGACG ATACCGAGGACAGCAGTTGTTACATTCCACCTCTCACCACGATTAAG CAGGATTTCAGACTACTGGGACAAACATCAGTAGACAGACTCCTGC AGTTGTCGCAGGGTCAAGCGGTAAAGGGAAACCAGTTACTACCTGT TTCTTTGGTTAAGAGAAAGACTACCTTGGCTCCCAACACGCAAACGG CAAGTCCGAGAGCTTTAGCGGACAGTTTGATGCAATTGGCGCGTCA GGTAAGCCGTCTGGAGAGTGGCCAATAATAACTCGGTACCAAATTCAttorney Docket No. 0073605-001135 CAGAAAAGAGACGCTGAAAAGCGTCTTTTTTCGTTTTGGTCCGTTTT TCCATAGGCTCCGCCCCCCTGACGAGCATCACAAAAATCGACGCTC AAGTCAGAGGTGGCGAAACCCGACAGGACTATAAAGATACCAGGC GTTTCCCCCTGGAAGCTCCCTCGTGCGCTCTCCTGTTCCGACCCTGCC GCTTACCGGATACCTGTCCGCCTTTCTCCCTTCGGGAAGCGTGGCGC TTTCTCATAGCTCACGCTGTAGGTATCTCAGTTCGGTGTAGGTCGTTC GCTCCAAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGC TGCGCCTTATCCGGTAACTATCGTCTTGAGTCCAACCCGGTAAGACA CGACTTATCGCCACTGGCAGCAGCCACTGGTAACAGGATTAGCAGA GCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGGCCTA ACTACGGCTACACTAGAAGGACAGTATTTGGTATCTGCGCTCTGCTG AAGCCAGTTACCTTCGGAAAAAGAGTTGGTAGCTCTTGATCCGGCA AACAAACCACCGCTGGTAGCGTTGGTTTTTTTGTTTGCAAGCAGCAG ATTACGCGCAGAAAAAAAGGATCTCAAGAAGATCCTTTGATCTTTTC TACGGGGTCTTTTTGTTATCAATAAAAAAGGCCCCCCTTTCGGGAGG CCTTATTGTTCGTCAGTTGATCGGCACGTAAGAGAAGAGATAGAAA ACTTCCACAACTAGACAAATTCTCCCCTAACACGGGAGTATTCGAGG GGGCAGGCGGATAAGGTGGTTCTAGCATGGAGAAGAAGATAACCG GATACACAACGGTTGACATTTCACAATGGCACCGGAAAGAGCACTT TGAAGCTTTCCAAAGCGTAGCCCAATGTACGTACAACCAAACGGTA CAGTTAGACATCACGGCGTTCTTAAAGACCGTTAAGAAGAACAAGC ACAAGTTCTACCCAGCCTTCATTCATATCCTCGCGAGGTTAATGAAT GCCCACCCCGAATTTCGGATGGCTATGAAAGACGGGGAATTAGTCA TCTGGGATAGTGTTCATCCCTGCTACACGGTGTTCCACGAGCAGACT GAAACATTCAGTAGCTTGTGGTCTGAGTATCATGACGATTTTAGGCA ATTCTTACATATCTATAGCCAAGATGTGGCATGTTATGGAGAGAACC TTGCATACTTTCCTAAAGGGTTCATAGAGAACATGTTTTTCGTGTCG GCCAATCCATGGGTGAGCTTCACTTCCTTCGACTTAAATGTGGCAAA TATGGATAACTTCTTCGCACCGGTCTTTACCATGGGCAAGTACTACA CGCAGGGTGATAAGGTTCTGATGCCCCTAGCCATCCAGGTTCATCAT GCTGTATGCGACGGATTTCATGTGGGCCGCATGCTGAATGAGCTACA GCAGTACTGCGATGAGTGGCAAGGTGGTGCGTGATAAGTCAGTTTCAttorney Docket No. 0073605-001135 ACCTGTTTTACGTAAAAACCCGCTTCGGCGGGTTTTTACTTTTGGAG CTCAGCATCGCAAGAATTCTAATACGACTCACTATAGGGGAATTGTG AGCGGATAACAATTCCTAAAATCACTCCGAAAGGAAAATAAGGAGC CCTAAGATGTACTACCCCGAACGGTACTTTGACATGTCGAACTGGCA GATGGATATGCAAGGCCGTTGGATGGATATGCAGGGTCGTTATTGC AGCCCCTACTGGTACAACTGGTATGGCCGGCAGATGTATTACCCTTA CCAGAATTATTACTGGTACGGTCGTTGGGACTACCCTGGAATGGATT ACAGCAATTGGCAAATGGATATGCAGGGTAGATGGATGGATATGCA GGGACGCTATATGGATCCATGGTGGATGAATGATTCGTATTACAATA ACTACTACAATTTAGGTTCATCAGGTGGTTCATCTGGAGTGTCAGGG TGGAGACTCTTTAAAAAGATTTCTGGATGATAAGTAAAGGTGAAGA AGATAACCTGGCCATTATTAAAGAGTTTATGCGCTTTAAGGTGCATA TGGAGGGTTCCGTTAATGGGCATGAGTTTGAGATAGAAGGTGAGGG TGAAGGACGTCCTTATGAAGGTACTCAAACAGCAAAACTGAAAGTG ACGAAAGGTGGACCTCTTCCCTTTGCATGGGATATCCTTTCACCGCA ATTTATGTACGGTAGTAAAGCATACGTGAAACATCCTGCGGATATTC CGGATTATCTTAAGTTATCGTTTCCGGAGGGGTTTAAATGGGAACGC GTGATGAATTTTGAAGATGGTGGGGTCGTCACAGTGACGCAGGACT CCTCTCTTCAGGACGGTGAGTTCATTTATAAAGTGAAACTCCGTGGA ACGAATTTTCCTTCTGACGGTCCAGTGATGCAGAAAAAGACAATGG GATGGGAAGCTTCTTCAGAACGGATGTATCCTGAAGATGGAGCCTT AAAAGGTGAAATAAAACAGAGGCTGAAATTAAAGGATGGAGGACA TTACGACGCGGAAGTGAAAACAACCTATAAAGCAAAAAAACCTGTT CAGTTACCTGGTGCGTACAATGTGAACATCAAACTCGACATAACGTC TCACAACGAGGATTACACAATCGTAGAGCAGTATGAAAGAGCTGAA GGGCGACATAGTACTGGGGGTATGGACGAACTTTACAAATAATGAA CTCGAGAGATAACAGATACTTCGGTATCTGTTATCTGTTTTTTTTCAA CAGATAGCCGCGTTCGCGCGGCTATCTGTTTTTTTTTTTable 14: Additional SequencesAttorney Docket No. 0073605-001135 Forward (PromoterSEQ ID NO: 151Library, SupportingTGGAACCTCTTACGTGCCGAData 1)ReverseSEQ ID NO: 152(Promoter Library,TCAGCGTCGTAGTGACCACCSupporting Data 1)ForwardSEQ ID NO: 153(5’ UTR Library,TCGGTGATGGTCCGGTTCTGSupporting Data 2)ReverseSEQ ID NO: 154(5’ UTR Library,CGTCGGACGGGAAGTTGGTASupporting Data 1)SEQ ID NO: 155 ATGGAGCTCAAACAGAACAAAGAACAGATCACCAAACAGAACA TCAACCGTAAAGGCGAAGAACTGTTTACCGGTGTGGTTCCGATT CTGGTGGAACTGGATGGTGATGTTAATGGTCATAAATTCAGCGT TCGTGGTGAAGGCGAAGGTGATGCCACGAATGGTAAACTGACC CTGAAATTTATCTGCACCACAGGTAAACTGCCGGTTCCGTGGCC GACCCTGGTTACCACCCTGACCTATGGTGTTCAGTGTTTCGCAC GTTATCCGGATCATATGAAACAGCACGATTTCTTTAAAAGCGCCsfGFP coding ATGCCGGAAGGTTATGTTCAGGAACGTACCATTAGCTTTAAAGA sequence, TGACGGCACCTATAAAACCCGTGCCGAAGTTAAATTCGAAGGC Supporting Data 2 GATACCCTGGTGAATCGTATCGAACTGAAAGGCATCGATTTTAA AGAGGATGGTAATATCCTGGGCCATAAACTGGAATATAATTTTA ACAGCCATAACGTGTATATCACCGCAGATAAACAGAAAAACGG CATTAAAGCGAACTTTAAAATCCGCCATAATGTGGAAGATGGTA GCGTTCAGCTGGCAGATCATTATCAGCAGAATACGCCGATCGGT GATGGTCCGGTTCTGCTGCCGGATAATCATTATCTGAGCACCCA GAGCGTTCTGAGTAAAGATCCGAATGAAAAACGTGATCACATG GTGCTGTTAGAGTTCGTTACCGCAGCAGGTATTACACATGGTAT GGATGAACTGTATAAATAAAttorney Docket No. 0073605-001135SEQ ID NO: 156 ATGCTCGAGGCTTCCTCCGAAGACGTTATCAAAGAGTTCATGCG TTTCAAAGTTCGTATGGAAGGTTCCGTTAACGGTCACGAGTTCG AAATCGAAGGTGAAGGTGAAGGTCGTCCGTACGAAGGTACCCA GACCGCTAAACTGAAAGTTACCAAAGGTGGTCCGCTGCCGTTCG CTTGGGACATCCTGTCCCCGCAGTTCCAGTACGGTTCCAAAGCT TACGTTAAACACCCGGCTGACATCCCGGACTACCTGAAACTGTCmRFP1 coding CTTCCCGGAAGGTTTCAAATGGGAACGTGTTATGAACTTCGAAG sequence, ACGGTGGTGTTGTTACCGTTACCCAGGACTCCTCCCTGCAAGAC Supporting Data 1 GGTGAGTTCATCTACAAAGTTAAACTGCGTGGTACCAACTTCCC GTCCGACGGTCCGGTTATGCAGAAAAAAACCATGGGTTGGGAA GCTTCCACCGAACGTATGTACCCGGAAGACGGTGCTCTGAAAG GTGAAATCAAAATGCGTCTGAAACTGAAAGACGGTGGTCACTA CGACGCTGAAGTTAAAACCACCTACATGGCTAAAAAACCGGTT CAGCTGCCGGGTGCTTACAAAACCGACATCAAACTGGACATCA CCTCCCACAACGAAGACTACACCATCGTTGAACAGTACGAACGT GCTGAAGGTCGTCACTCCACCGGTGCTTAATAA

[0197] It should be understood that modifications to the embodiments disclosed herein can be made to meet a particular set of design criteria. For instance, the number of or configuration of components or parameters may be used to meet a particular objective.

[0198] It will be apparent to those skilled in the art that numerous modifications and variations of the described examples and embodiments are possible in light of the above teachings of the disclosure. The disclosed examples and embodiments are presented for purposes of illustration only. Other alternative embodiments may include some or all the features of the various embodiments disclosed herein. For instance, it is contemplated that a particular feature described, either individually or as part of an embodiment, can be combined with other individually described features, or parts of other embodiments. The elements and acts of the various embodiments described herein can therefore be combined to provide further embodiments.

[0199] It is the intent to cover all such modifications and alternative embodiments as may come within the true scope of this invention, which is to be given the full breadth thereof. Additionally, the disclosure of a range of values is a disclosure of every numerical value within that range, including the end points. Thus, while certain exemplary embodiments of the device and methods of makingAttorney Docket No. 0073605-001135 and using the same have been discussed and illustrated herein, it is to be distinctly understood that the invention is not limited thereto but may be otherwise variously embodied and practiced within the scope of the following claims.

Claims

Attorney Docket No. 0073605-001135CLAIMS1. A screening platform for analyzing gene expression or evaluating protein production, comprising:an array of microcapillary tubes each configured to contain a volume of cell culture, the cell culture of at least one microcapillary tube of the array of microcapillary tubes configured to express at least a photoactive reporter;an optical module comprising:a light source configured to illuminate the volume of cell culture, thereby producing an optical emission; anda detector configured to detect the optical emission; anda control unit communicatively connected to the optical module, wherein the control unit is configured to:quantify a level of gene expression or protein production of the cell culture as a function of the optical emission;identify a genotype of the cell culture as a function of the quantification, thereby providing population-based screening of the cell culture.

2. The screening platform according to claim 1, wherein the array of microcapillary tubes comprises an array of at least 105microcapillary tubes.

3. The screening platform according to claim 1, wherein the light source comprises a multiband illumination system across a wavelength range of at least 300 nm and no greater than 800 nm.

4. The screening platform according to claim 1, wherein the optical module is configured to perform a bright-field absorbance measurement.

5. The screening platform according to claim 1, wherein the at least a photoactive reporter comprises at least a fluorescent reporter.

6. The screening platform according to claim 5, wherein:the at least a fluorescent reporter comprises a plurality of fluorescent reporters; and the control unit is further configured to perform multi-reporter operon screening as a function of the optical emission.

7. The screening platform according to claim 6, wherein performing the multi -reporter operon screening comprises performing a comparative analysis of ribosome binding sites or operon dynamics.Attorney Docket No. 0073605-001135 8. The screening platform according to claim 1, wherein identifying the genotype of the cell culture comprises:isolating a phenotypic trait of the cell culture as a function of the level of gene expression or protein production; andidentifying the genotype of the cell culture as a function of the phenotypic trait.

9. The screening platform according to claim 1, wherein the control unit is further configured to determine a level of transcriptional activity as a function of the level of protein production.

10. The screening platform according to claim 1, further comprising a chemical perturbation module communicatively connected to the control unit and configured to apply a controlled chemical stimulus to the cell culture, wherein the control unit is further configured to evaluate an effect of the controlled chemical stimulus on the gene expression or protein production in the cell culture.

11. The screening platform according to claim 10, wherein applying the controlled chemical stimulus comprises applying one or more inducers.

12. The screening platform according to claim 1, wherein the control unit is further configured to:measure a half-life of mRNA as a function of the optical emission; andcorrelate the half-life and the level of protein production.

13. The screening platform according to claim 1, wherein the control unit is further configured to:generate a cell growth profile of the cell culture; anddetermine one or more structural properties of a protein as a function of the cell growth profile.

14. The screening platform according to claim 13, wherein the one or more structural properties of the protein includes hydrophobicity of the protein.

15. The screening platform according to claim 1, wherein the control unit is further configured to generate an output as a function of the level of gene expression or protein production.

16. The screening platform according to claim 15, wherein the output comprises one or more of a suitable plasmid construct for protein production, a suitable promoter for protein production, a suitable 5’ untranslated region for protein production, and a suitable amino acid sequence for protein production.Attorney Docket No. 0073605-001135 17. The screening platform according to claim 15 or 16, wherein the control unit is further configured to display the output using an output device.

18. The screening platform according to claim 15 or 16, further comprising an image-processing module communicatively connected to the control unit, wherein generating the output comprises generating an intensity map, using the image-processing module, as a function of the optical emission.

19. The screening platform according to claim 15 or 16, wherein the output comprises one or more coordinates or indices of one or more microcapillaries of interest.

20. The screening platform according to claim 19, further comprising one or more extraction slips configured to collect the volume of cell culture extracted from the one or more microcapillaries of interest.

21. The screening platform according to claim 20, wherein the volume of cell culture contains one or more magnetic beads and is extracted using a focused pulse-laser beam.

22. The screening platform according to claim 21, wherein one or more plasmids of the extracted volume of cell culture are further amplified and sequenced using polymerase chain reaction.

23. The screening platform according to claim 1, wherein the screening platform provides a clone recovery mechanism having an enrichment ratio of up to 105: 1.

24. The screening platform according claim 1, wherein the control unit is further configured to:perform population binning as a function of the level of gene expression or protein production; andidentifying the genotype of the cell culture as a function of the population binning.

25. A method for analyzing gene expression or evaluating protein production, comprising:applying a plurality of cell cultures to an array of microcapillary tubes each configured to contain a volume of cell culture, the cell culture in at least one microcapillary tube of the array of microcapillary tubes configured to express at least a photoactive reporter; measuring an optical emission of the plurality of cell cultures using an optical module comprising a light source and a detector;quantifying, using a control unit communicatively connected to the optical module, a level of gene expression or protein production in the plurality of cell cultures, as a function of the optical emission; andAttorney Docket No. 0073605-001135 identifying a genotype of at least one cell culture of the plurality of cell cultures, using the control unit, as a function of the quantification, thereby providing population-based screening of the cell culture.

26. The method according to claim 25, wherein the array of microcapillary tubes comprises an array of at least 105microcapillary tubes.

27. The method according to claim 25, wherein the light source comprises a multiband illumination system across a wavelength range of at least 300 nm and no greater than 800 nm.

28. The method according to claim 25, further comprising measuring a bright-field absorbance of the volume of cell culture using the optical module.

29. The method according to claim 25, wherein the at least a photoactive reporter comprises at least a fluorescent reporter.

30. The method according to claim 29, wherein the at least a fluorescent reporter comprises a plurality of fluorescent reporters, the method further comprising performing multi-reporter operon screening, using the control unit, as a function of the optical emission.

31. The method according to claim 30, wherein performing the multi -reporter operon screening comprises performing a comparative analysis of ribosome binding sites or operon dynamics.

32. The method according to claim 25, wherein identifying the genotype of the cell culture comprises:isolating a phenotypic trait of the cell culture as a function of the level of gene expression or protein production; andidentifying the genotype of the cell culture as a function of the phenotypic trait.

33. The method according to claim 25, further comprising determining a level of transcriptional activity as a function of the level of protein production.

34. The method according to claim 25, further comprising:applying a controlled chemical stimulus to the cell culture, using a chemical perturbation module communicatively connected to the control unit; andevaluating an effect of the controlled chemical stimulus on the gene expression or protein production in the cell culture.

35. The method according to claim 34, wherein applying the controlled chemical stimulus comprises applying one or more inducers.

36. The method according to claim 25, further comprising:Attorney Docket No. 0073605-001135 measuring a half-life of mRNA, using the control unit, as a function of the optical emission;andcorrelating the half-life and the level of protein production using the control unit.

37. The method according to claim 25, further comprising:generating a cell growth profile of the cell culture using the control unit; and determining, using the control unit, one or more structural properties of a protein as a function of the cell growth profile.

38. The method according to claim 37, wherein the one or more structural properties of the protein includes hydrophobicity of the protein.

39. The method according to claim 25, further comprising generating an output as a function of the level of gene expression or protein production.

40. The method according to claim 39, wherein the output comprises one or more of a suitable plasmid construct for protein production, a suitable promoter for protein production, a suitable 5’ untranslated region for protein production, and a suitable amino acid sequence for protein production.

41. The method according to claim 39 or 40, further comprising displaying the output using an output device communicatively connected to the control unit.

42. The method according to claim 39 or 40, wherein generating the output comprises generating an intensity map, using an image-processing module communicatively connected to the control unit, as a function of the optical emission.

43. The method according to claim 39 or 40, wherein the output comprises one or more coordinates or indices of one or more microcapillaries of interest.

44. The method according to claim 43, further comprising collecting the volume of cell culture extracted from the one or more microcapillaries of interest using one or more extraction slips.

45. The method according to claim 44, wherein the volume of cell culture contains one or more magnetic beads and is extracted using a focused pulse-laser beam.

46. The method according to claim 45, further comprising amplifying and sequencing one or more plasmids of the extracted volume of cell culture using polymerase chain reaction.

47. The method according to claim 25, the method providing a clone recovery mechanism having an enrichment ratio of up to 105:1.

48. The method according to claim 25, further comprising:Attorney Docket No. 0073605-001135 performing population binning, using the control unit, as a function of the level of gene expression or protein production; andidentifying the genotype of the cell culture, using the control unit, as a function of the population binning.