Screening method based on liquid drops

By sorting droplets into multiple output channels and combining them with machine learning algorithms, the inefficiency of existing HTS methods is solved, enabling efficient screening and sequencing of large biological libraries and generating high-quality datasets for computational model training.

CN121925476APending Publication Date: 2026-04-24NOVOZYMES AS
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NOVOZYMES AS
Filing Date
2024-09-26
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing high-throughput screening (HTS) methods are inefficient and cannot efficiently screen and sequence large biological libraries, resulting in high time and cost. Furthermore, conventional droplet sorting methods cannot perform multiple bin separations in a single step, leading to initial sample loss and low screening efficiency.

Method used

Droplets were sorted into at least three output channels using a microdroplet sorter, and sorted based on the amount of screenable products. Sequencing was then performed, and machine learning algorithms were used to score and identify polynucleotides.

Benefits of technology

It improves screening efficiency, reduces costs and standard deviation, enables screening of larger libraries, generates more efficient computational models, reduces starting sample loss, identifies library members that are difficult to screen, and provides detailed polynucleotide variant studies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925476A_ABST
    Figure CN121925476A_ABST
Patent Text Reader

Abstract

The present invention relates to methods for screening biological libraries. The invention also relates to nucleic acid sequences, vectors and host cells that have been isolated and / or produced by the methods of the invention.
Need to check novelty before this filing date? Find Prior Art

Description

Background of the Invention Technical Field

[0001] This invention relates to methods for screening biological libraries using microfluidic chips. The invention also relates to nucleic acid sequences, vectors, and host cells that have been isolated and / or generated using the methods of this invention.

[0002] The method of the present invention also relates to the identification of target polynucleotides and / or cells having the desired characteristics. Droplets sorted in a chip, sometimes referred to as microdroplets or microcapsules, typically have an average diameter of about 20 micrometers and are used as compartments or tiny reaction vessels. They may contain, for example, live microbial cells that secrete enzymes. Alternatively or additionally, the droplets may contain cell extracts capable of expressing proteins encoded by the target polynucleotide. The droplets may also contain other components, such as fluorescent enzyme substrates that can reveal enzyme activity. Background Technology

[0003] Reducing the cost of ordering synthetic DNA and creating DNA libraries has enabled the production of larger libraries with high sequence diversity. However, despite the decreasing cost of DNA sequencing, characterizing sequences in a time- and cost-efficient manner remains challenging. Conventional high-throughput screening (HTS) methods using microtiter plates (MTPs) are known to be both expensive and slow. While library sorting using two-dimensional FACS / FADS methods or two-dimensional microdroplet sorters is available, these methods are inefficient because separation into multiple bins cannot be performed in a single step. For example, sorting into multiple bins must be done sequentially, requiring large starting samples, which is time-inefficient and results in significant loss of starting samples.

[0004] Therefore, improved library sorting and sequencing methods are needed. Summary of the Invention

[0005] The inventors who proposed the method unexpectedly discovered that sorting microdroplets into at least three exit channels and subsequently sequencing them offers several combined benefits compared to conventional high-throughput screening (HTS) methods:

[0006] • Efficiency: The method of this invention reduces the time and cost required for each library member.

[0007] • Improved performance: The method of the present invention improves the overall efficiency of the process.

[0008] • Robustness and accuracy: The method of this invention is as robust as conventional methods (as demonstrated in Example 3), while exhibiting reduced variability (proven by a reduced standard deviation, as shown in Example 4). The reduced standard deviation contributes to the generation of computationally better-performing models.

[0009] • Scalability: The method of this invention enables the screening of larger libraries. Screening a greater number of library members allows for the generation of more efficient models (as shown in Example 6). The method of this invention also identifies library members that are difficult or impossible to identify using other methods (as shown in Example 7).

[0010] The advantages mentioned above will be elaborated in more detail below.

[0011] Compared to conventional HTS methods, droplet microfluidics can increase the generation rate of biological screening data by approximately 1000 times at less than 1% cost. Compared to conventional two-way sorting methods, fewer library members and reagents are discarded into the bin when sorting biological libraries according to the present invention. Therefore, the method of the present invention requires fewer cells and / or library members in the starting sample without compromising readout quality. Furthermore, scoring of the target polynucleotide after sorting into three or more receiving channels allows for detailed study of the relationship between variants of the target polynucleotide and the desired effect (e.g., enzyme activity or enzyme yield). Compared to droplet sorting to two channels (two subsets) that only allows droplet separation based on a single threshold, the method of the present invention allows for higher readout resolution, i.e., the ability to distinguish multiple subsets after droplet separation based on multiple thresholds. For example, in some cases, it is not advantageous to have the highest binding activity of the peptide to the substrate or inhibitor, but rather a moderate binding activity preferred at the “optimal point.”

[0012] Variant polynucleotide sequences not only associated with positive effects (e.g., increased enzyme yield or increased enzyme activity) but also variant sequences associated with less desirable effects (e.g., decreased enzyme yield or decreased enzyme activity) are part of the detailed output and can be considered for library screening and / or mutation in further rounds.

[0013] As illustrated in the examples, the results obtained using the method of this invention unexpectedly show a strong correlation with the results of the MTP screening method. Therefore, the method of this invention allows skipping or replacing the well-established MTP screening method, which is known to require time and resources. While maintaining a strong correlation with MTP results, the method of this invention also exhibits a lower standard deviation (STD) compared to MTP. This lower STD is particularly advantageous when the obtained data is used as input data for machine learning models, as a lower STD produces high-confidence machine learning models that exhibit a more accurate and robust algorithmic form.

[0014] Advantageously, as shown in Table 1 and Example 7, the method of the present invention allows for the identification of promising library members from large libraries that would otherwise be impossible to identify using conventional screening methods, which are limited to smaller libraries.

[0015] Leveraging the power of machine learning depends on the availability of large datasets. In conventional droplet screening, only a small fraction of droplets are separated and analyzed. Therefore, for most library members, for example, when screening cellular libraries, the association between phenotype and genotype is lost. To characterize the diversity of the entire library, the method of this invention provides a strategy in which droplets are sorted into multiple (at least three) output channels (pools). The abundance of each individual sequence in each pool is then measured and a score is calculated for each polynucleotide sequence. In this way, droplet screening techniques can be used to efficiently generate extensive datasets while considering every polynucleotide sequence present in each pool.

[0016] Therefore, advantageously, scoring, identification, and / or sequencing of one or more target polynucleotides in one or more output channels can be used to train computational models to gain further insights into sequence characteristics to improve desired sequence features (e.g., increased yield and / or increased enzyme activity) and / or produce synthetic sequences with such improved features. Furthermore, when utilizing the output of the computational model, it is important to minimize the number and extent of artifacts by maintaining uniform incubation conditions. Artifacts (e.g., edge effects) typically occur in MTPs because outer pores are exposed to slightly different conditions than pores located more centrally, for example, due to heat capacity and evaporation effects. For microdroplets, these artifacts can be avoided or minimized, which significantly improves the quality of the model.

[0017] In a first aspect, the present invention relates to a method for screening biological libraries, the method comprising the following steps:

[0018] a) Provide a microfluidic device including a droplet sorter (200) having at least three output channels (301, 302, 303).

[0019] b) Provides a droplet emulsion containing a target polynucleotide library and screenable products.

[0020] c) Determine the amount of screenable products from one or more droplets in the microfluidic device.

[0021] d) The droplet sorter (200) sorts the one or more droplets into the receiving output channels of the at least three output channels (301, 302, 303), wherein the receiving output channel is determined based on the amount of screenable product per droplet, and wherein the at least three receiving output channels receive a plurality of droplets, the plurality of droplets comprising an amount of screenable product above and / or below one or more predetermined threshold levels.

[0022] e) Identify one or more target polynucleotides present in the at least three output channels (301, 302, 303), and obtain the sequence data of the one or more target polynucleotides.

[0023] f) For one or more output channels (301, 302, 303), assign a score to each of the one or more target polynucleotides, wherein the score is calculated based on the abundance of each of the one or more target polynucleotides in one of the one or more output channels (301, 302, 302).

[0024] In a second aspect, the present invention relates to a host cell whose genome contains a synthetic target polynucleotide produced in a further step h) and / or a target polynucleotide identified in step e).

[0025] In a third aspect, the present invention relates to a method for producing a target polypeptide, the method comprising culturing cells according to the second aspect under conditions conducive to the production of the polypeptide. Attached Figure Description

[0026] Figure 1 A schematic overview diagram of a microfluidic device with multiple output channels according to an embodiment of the method according to the present invention is shown.

[0027] Figure 2 The measurement response for 100,000 droplets and the thresholds used to separate them into five output channels (cells 1-5) are shown.

[0028] Figure 3 The relative abundance of 102 signal peptide variants sorted into five output channels (pools 1-5) is shown.

[0029] Figure 4 The correlation between the droplet method of the present invention and the score of MTP measurement is shown.

[0030] Figure 5 The standard deviation (σ) of MTP fermentation (A) and droplet fermentation (B) based on count and relative protein yield is shown.

[0031] Figure 6 The fraction of proline-containing sequences obtained from MTP screening is shown in (A), and the order of signal peptide sequences obtained from the same MTP screening is shown in (B).

[0032] Figure 7 The fraction of proline-containing sequences obtained from microdroplet screening is shown (A), and the order of signal peptide sequences obtained from the same microdroplet screening is shown (B).

[0033] Figure 8The correlation coefficient between predicted and observed values ​​is shown, varying depending on the size of the training data. definition

[0034] Based on this detailed description, the following definitions apply. Note that the singular forms “a / an” and “the” include plural indicators unless the context explicitly indicates otherwise.

[0035] Unless otherwise defined or explicitly indicated by the context, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0036] Biological library: The term “biological library” refers to a library containing multiple variants, including but not limited to one or more of DNA sequence variants, amino acid sequence variants, and cell variants.

[0037] Non-limiting examples of DNA sequence variants include recombinant cell libraries containing said DNA sequence variants and cell-free systems containing said DNA sequence variants. Thus, in one embodiment, the target polynucleotide library is contained within a recombinant host cell.

[0038] Non-limiting examples of DNA sequence variants include wild-type cell libraries containing natural DNA sequence variants. Thus, in one embodiment, the target polynucleotide library is contained in wild-type cells.

[0039] Non-limiting examples of amino acid variants include libraries or purified peptide variants and recombinant cell libraries expressing peptide variants.

[0040] Non-restrictive examples of cell variants include wild-type cell libraries and recombinant cell libraries containing DNA sequence variants.

[0041] cDNA: The term "cDNA" refers to a DNA molecule that can be prepared by reverse transcription from mature, spliced ​​mRNA molecules obtained from eukaryotic or prokaryotic cells. cDNA lacks intron sequences that can be present in the corresponding genomic DNA. The initial primary RNA transcript is the precursor of mRNA, which is processed through a series of steps (including splicing) to become mature, spliced ​​mRNA.

[0042] Coding sequence: The term "coding sequence" refers to a polynucleotide that directly specifies the amino acid sequence of a polypeptide. The boundaries of a coding sequence are typically defined by an open reading frame (OPG), which begins with a start codon (such as ATG, GTG, or TTG) and ends with a stop codon (such as TAA, TAG, or TGA). Coding sequences can be genomic DNA, cDNA, synthetic DNA, or a combination thereof.

[0043] Computational Model: The term "computational model" refers to a computational program or process designed to enable a computer or machine to learn from data and improve its performance on a specific task without being explicitly programmed for that task. Computational models include machine learning algorithms. These algorithms utilize statistical techniques to identify patterns, relationships, and correlations within a dataset, enabling them to make predictions, classifications, or decisions based on new or unseen data.

[0044] Examples of machine learning algorithms include, but are not limited to:

[0045] Linear Regression A basic algorithm that models the relationship between a dependent variable and one or more independent variables by fitting a linear equation to observed data points. It is commonly used for tasks such as predicting numerical values, such as predicting house prices based on factors such as area and location.

[0046] Decision Tree Decision trees are a method for making decisions based on multiple conditions using a tree structure. Each node in the tree represents a decision based on specific features, ultimately leading to a leaf node with a final prediction or classification. Decision trees are used for tasks (like classification) where algorithms determine the category of an input based on features.

[0047] Random Forest Random forests are an ensemble learning technique that combines multiple decision trees to enhance prediction accuracy and mitigate overfitting. Each tree in the forest makes a prediction, and the final result is determined by aggregating the predictions of all the trees. Random forests have been applied in various fields, such as image classification and medical diagnosis.

[0048] Random forest models are advanced machine learning algorithms used in a variety of applications, such as classification, regression, and data analysis. During its training phase, the algorithm builds an ensemble consisting of many decision trees. Notably, each decision tree is built using a subset of the training dataset and randomized classification of the input features.

[0049] The unique power of random forest models stems from their ability to integrate predictions derived from multiple decision trees. This fusion, known as "bagging," produces higher accuracy and stability compared to single trees. By mitigating overall bias and suppressing the variability exhibited by individual trees, random forest models effectively avoid overfitting. Consequently, their performance is significantly enhanced in accurately predicting novel, previously unseen data instances.

[0050] Furthermore, random forest models skillfully manage high-dimensional datasets and complex feature dependencies, making them particularly suitable for complex real-world problems. Notably, they have the ability to accommodate missing data values, ensuring consistent accuracy even when the data is partially incomplete.

[0051] Support Vector Machine (SVM): A classification algorithm that finds the optimal hyperplane for separating data points from different classes by maximizing the margins between them. SVMs are used for tasks such as text classification, image recognition, and bioinformatics.

[0052] Neural Networks Deep learning refers to complex algorithms inspired by the structure and function of biological neural networks. They consist of layers of interconnected nodes (neurons) that process and transform data. Deep learning is a subset of neural networks that involves multiple hidden layers and is used for tasks such as natural language processing, image generation, and autonomous driving.

[0053] K-means clustering A non-supervised learning algorithm used to divide a dataset into different clusters based on the similarity of data point features. It is used for market segmentation, customer profiling, and image compression.

[0054] Reinforcement learning: A method in which an algorithm learns to make a set of decisions by interacting with its environment to maximize cumulative reward. This is commonly used in robotics, game theory, and autonomous systems.

[0055] Naive Bayes: A probabilistic algorithm based on Bayes' theorem is particularly effective for text classification tasks such as spam detection and sentiment analysis.

[0056] Principal Component Analysis (PCA): A dimensionality reduction technique used to transform high-dimensional data into a low-dimensional representation while preserving as much variance as possible. PCA is used for image compression, data visualization, and feature extraction.

[0057] Gaussian Mixture Model (GMM): A probabilistic model that assumes data is generated by a mixture of several Gaussian distributions. GMMs are applied in various fields, including speech recognition, image segmentation, and anomaly detection.

[0058] Generative Adversarial Networks (GANs): A specialized class of machine learning algorithms involves two neural networks (a generator and a discriminator) competing in a process. The generator creates synthetic data instances (such as images or text) that resemble real data, while the discriminator evaluates whether a given data instance is real or generated. These two networks iteratively refine their performance, with the generator aiming to produce increasingly realistic data and the discriminator improving its ability to distinguish between real and generated data. A suitable, non-restricted example of a GAN is disclosed in WO 2024 / 133344 (Novozymes A / S).

[0059] Control Sequences: The term "control sequence" refers to a nucleic acid sequence involved in regulating the expression of polynucleotides in a particular organism, either in vivo or in vitro. Each control sequence can be native (i.e., from the same gene) or heterologous (i.e., from different genes) for the polynucleotide encoding a polypeptide, and is native or heterologous relative to each other. Such control sequences include, but are not limited to, leader sequences, polyadenylation sequences, propeptides, propeptides, signal peptides, promoters, terminators, enhancers, and transcription or translation initiator and terminator sequences. At a minimum, control sequences include promoters and transcription and translation termination signals. These control sequences may be provided with multiple linkers for the purpose of introducing specific restriction sites that facilitate the linking of control sequences to the coding regions of polynucleotides encoding polypeptides.

[0060] ddPCR: The term "droplet digital PCR" or "ddPCR" refers to an advanced molecular biology technique used for the precise analysis and quantification of nucleic acids (including DNA and RNA) within a sample. This method represents an innovation over conventional polymerase chain reaction (PCR) methods, designed to overcome the inherent limitations of traditional PCR by enabling the accurate measurement and detection of subtle changes in rare target sequences or target concentrations.

[0061] In ddPCR, the sample containing the target nucleic acid is intelligently subdivided into numerous individual droplets, each operating as an independent reaction compartment. This strategic segmentation step facilitates the isolation and amplification of the target nucleic acid, minimizing amplification bias and the possibility of interference from non-target molecules. After amplification, the droplets undergo fluorescence-based analysis to determine the presence or absence of the amplified target sequence within each individual droplet.

[0062] Droplet sorter: The term "droplet sorter" (200) refers to an arrangement within a microfluidic device that allows droplets to be sorted into three or more output channels, where sorting is based on the amount of screenable product detected in the droplet. Sorting is performed, for example, by using one or more sorting devices (e.g., electrodes or valves). The amount of screenable product in the droplet is detected by one or more sensor elements and conveyed to the sorting device, such as two or more electrodes.

[0063] Expression: The term “expression” refers to any step involved in polypeptide production, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, and secretion.

[0064] Expression vector: An expression vector is a linear or circular DNA construct containing a DNA sequence encoding a polypeptide, with the coding sequence operatively linked to a suitable control sequence that can influence the expression of the DNA in a suitable host. Such control sequences may include promoters that influence transcription, optional operon sequences that control transcription, sequences encoding suitable ribosome binding sites on mRNA, enhancers, and sequences that control the termination of transcription and translation.

[0065] Extension: The term "extension" refers to the addition of one or more amino acids to the amino and / or carboxyl termini of a polypeptide, wherein the "extended" polypeptide has enzymatic activity.

[0066] Fragment: The term "fragment" refers to a polypeptide in which one or more amino acids are missing from the amino and / or carboxyl termini of a mature polypeptide, wherein the fragment has enzymatic activity.

[0067] Fusion polypeptide: The term "fusion polypeptide" is a polypeptide in which one of the polypeptides is fused to the N-terminus and / or C-terminus of the polypeptide of the present invention. Fusion polypeptides are generated by fusing a polynucleotide encoding another polypeptide with a polynucleotide of the present invention, or by fusing two or more polynucleotides of the present invention together. Techniques for generating fusion polypeptides are known in the art and include linking the coding sequences of the polypeptides such that they conform to reading frames, and that the expression of the fusion polypeptide is under the control of the same promoter and terminator. Fusion polypeptides can also be constructed using intronomer technology, wherein the fusion polypeptide is generated post-translational (Cooper et al., 1993, EMBO J. [Journal of the European Society for Molecular Biology] 12: 2575-2583; Dawson et al., 1994, Science [Science] 266: 776-779). Fusion polypeptides may further include a cleavage site between the two polypeptides. This site is cleaved upon secretion of the fusion protein, thereby releasing both polypeptides. Examples of cleavage sites include, but are not limited to, those disclosed in the following literature: Martin et al., 2003, J. Ind. Microbiol. Biotechnol. [Journal of Industrial Microbiology and Biotechnology] 3: 568-576; Svetina et al., 2000, J. Biotechnol. [Journal of Biotechnology] 76: 245-251; Rasmussen-Wilson et al., 1997, Appl. Environ. Microbiol. [Applied and Environmental Microbiology] 63: 3488-3493; Ward et al., 1995, Biotechnology [Biotechnology] 13: 498-503; and Contreras et al., 1991, Biotechnology [Biotechnology] 9: 378-381; Eaton et al., 1986, Biochemistry [Biochemistry] 25: 505-512; Collins-Racie et al., 1995, Biotechnology 13: 982-987; Carter et al., 1989, Proteins: Structure, Function, and Genetics 6: 240-248; and Stevens, 2003, Drug Discovery World 4: 35-48.

[0068] Heterogeneous: For host cells, the term "heterogeneous" means that the polypeptide or nucleic acid is not naturally present in the host cell. For polypeptides or nucleic acids, the term "heterogeneous" means that the control sequence (e.g., promoter) of the polypeptide or nucleic acid is not naturally associated with that polypeptide or nucleic acid; that is, the control sequence comes from a gene other than the gene encoding the mature polypeptide.

[0069] Host strain or host cell: A “host strain” or “host cell” is an organism that contains the target polynucleotide. An exemplary host strain is a microbial cell (e.g., bacteria, filamentous fungi, and yeast) capable of expressing the target polypeptide and / or fermentable sugar and / or probiotic microorganisms.

[0070] A recombinant host strain or recombinant host cell refers to an organism in which an expression vector, bacteriophage, virus, or other DNA construct (including a polynucleotide encoding a target polypeptide, such as amylase) has been introduced. An exemplary recombinant host strain is a microbial cell (e.g., bacteria, filamentous fungi, and yeast) capable of expressing a target polypeptide and / or fermenting sugar. The term "host cell" includes protoplasts produced by cells.

[0071] Introduction: In the context of inserting a nucleic acid sequence into a cell, the term “introduction” means “transfection,” “conversion,” or “transduction,” as is known in the art.

[0072] Isolated: The term "isolated" means a polypeptide, nucleic acid, cell, or other specific material or component that has been separated from at least one other material or component (including, but not limited to, other proteins, nucleic acids, cells, etc.). Therefore, the isolated polypeptide, nucleic acid, cell, or other material exists in a form not found in nature. Isolated polypeptides include, but are not limited to, culture media containing secreted polypeptides expressed in host cells.

[0073] Mature polypeptide: The term “mature polypeptide” refers to a polypeptide that has been processed at its N-terminus and / or C-terminus (e.g., removal of the signal peptide) to be in its mature form.

[0074] Mature polypeptide coding sequence: The term "mature polypeptide coding sequence" refers to the polynucleotide that encodes a mature polypeptide.

[0075] Microfluidic Device: According to the present invention, the microfluidic device includes a droplet sorter (200) comprising at least three output channels (301, 302, 303). Typically, the microfluidic device also includes multiple liquid inlets and / or liquid outlets. In one embodiment, the device includes an incubation chamber (500).

[0076] Natural: The term "natural" refers to nucleic acids or polypeptides that are naturally present in host cells.

[0077] Nucleic acid: The term "nucleic acid" encompasses DNA, RNA, heteroduplexes, and synthetic molecules capable of encoding polypeptides. Nucleic acids can be single-stranded or double-stranded and can be chemically modified. The terms "nucleic acid" and "polynucleotide" are used interchangeably. Because the genetic code is degenerate, more than one codon can be used to encode a specific amino acid, and the compositions and methods of the present invention cover nucleotide sequences encoding specific amino acid sequences. Unless otherwise indicated, nucleic acid sequences are presented in a 5' to 3' orientation.

[0078] Nucleic acid constructs: The term “nucleic acid construct” refers to a single-stranded or double-stranded nucleic acid molecule that is isolated from a naturally occurring gene or modified in a way that does not originally exist in nature to contain a segment of nucleic acid or is synthesized and contains one or more control sequences that are operatively linked to the nucleic acid sequence.

[0079] Operationally linked: The term "operationally linked" means that specified components are in a relationship that allows them to function in the intended manner (including, but not limited to, juxtaposition). For example, a regulatory sequence is operationally linked to a coding sequence such that the expression of the coding sequence is under the control of the regulatory sequence.

[0080] Protease: In one respect, the target polynucleotide encodes a protease. Suitable proteases include those of bacterial, fungal, plant, viral, or animal origin, such as microbial or plant origins. Microbial origin is preferred. This includes chemically modified variants or protein-engineered variants. It can be an alkaline protease, such as a serine protease or a metalloproteinase. Serine proteases can be, for example, from the S1 family (such as trypsin) or the S8 family (such as subtilisin). Metalloproteinases can be, for example, thermophilic bacterial proteases from the M4 family or other metalloproteinases, such as those from the M5, M7, or M8 families. The serine endopeptidase hydrolyzes the substrate N-succinyl-Ala-Ala-Pro-Phe p-nitroaniline. In the context of this example, the reaction is carried out at room temperature at pH 9.0. The release of pNA results in an increase in absorbance at 405 nm, and this increase is proportional to the enzyme activity measured against standards.

[0081] Purified: The term "purified" means nucleic acids, peptides, or cells that are substantially free of other components, as determined by analytical techniques well known in the art (e.g., in electrophoretic gels, chromatographic eluates, and / or media subjected to density gradient centrifugation, where purified peptides or nucleic acids form discrete bands). Purified nucleic acids or peptides are at least about 50% pure, and typically at least about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 99.5%, about 99.6%, about 99.7%, about 99.8%, or more pure (e.g., weight percentage or molar percentage). In a relevant sense, a composition is enriched with the molecule when the concentration of the molecule increases significantly after the application of purification or enrichment techniques. The term “enrichment” refers to the presence of compounds, peptides, cells, nucleic acids, amino acids, or other specified materials or components in a composition at a relative or absolute concentration higher than that of the starting composition.

[0082] In one respect, the term "purified," as used herein, means that the polypeptide or cell is substantially free of components (especially insoluble components) from the producing organism. In another respect, the term "purified" means that the polypeptide is substantially free of insoluble components (especially insoluble components) from the natural organism from which it was obtained. In one respect, the polypeptide is separated from some soluble components of the organism from which it was recovered and the culture medium. The polypeptide can be purified (i.e., separated) by one or more of the following methods: unit operation filtration, precipitation, or chromatography.

[0083] Accordingly, peptides can be purified so that only small amounts of other proteins, particularly other peptides, are present. The term "purified" as used herein can refer to the removal of other components, particularly other proteins and most particularly other enzymes, present in the cells from which the peptide originates. A peptide can be "substantially pure," meaning it is free from other components from the organism that produced it (e.g., the host organism used to recombinantly produce the peptide). In one aspect, the peptide is at least 40% pure by weight of the total peptide material present in the formulation. In another aspect, the peptide is at least 50%, 60%, 70%, 80%, or 90% pure by weight of the total peptide material present in the formulation. As used herein, "substantially pure peptide" can mean a peptide formulation containing, by weight, at most 10%, preferably at most 8%, more preferably at most 6%, more preferably at most 5%, more preferably at most 4%, more preferably at most 3%, even more preferably at most 2%, most preferably at most 1%, and even most preferably at most 0.5% of the peptide and other peptide material associated with it, either naturally or recombinantly.

[0084] Therefore, it is preferred that the substantially pure polypeptide, based on the weight of the total polypeptide material present in the formulation, is at least 92% pure, preferably at least 94% pure, more preferably at least 95% pure, more preferably at least 96% pure, more preferably at least 97% pure, more preferably at least 98% pure, even more preferably at least 99% pure, and most preferably at least 99.5% pure. The polypeptides of the present invention are preferably in a substantially pure form (i.e., the formulation is substantially free of other polypeptide materials associated with it, either naturally or recombinantly). This can be achieved, for example, by preparing the polypeptide using well-known recombinant methods or classical purification methods.

[0085] Recombination: The term "recombination," used in its conventional sense, refers to the manipulation (e.g., cutting and rejoining) of nucleic acid sequences to form a sequence group different from that found in nature. The term recombination refers to cells, nucleic acids, polypeptides, or vectors that have been modified from their natural state. Thus, for example, recombinant cells express genes not found in their natural (non-recombinant) forms, or express natural genes at different levels or under different conditions compared to those found in nature. The term "recombination" is synonymous with "genetically modified" and "transgenic."

[0086] Recovery: The term "recovery" refers to the removal of peptides from at least one fermentation broth component selected from a list of cells, nucleic acids, or other specified materials, for example, by methods such as: harvesting peptides by peptide crystallization, by filtration (e.g., deep filtration (using filter aids or packed filter media, cloth filtration in a box filter, rotary drum filtration, drum filtration, rotary vacuum drum filtration, candle filter, horizontal leaf filter, or the like, using sheet or pad filtration in a frame or modular device) or membrane filtration (using plate filtration, modular filtration, candle filtration, microfiltration, crossflow, dynamic crossflow, or ultrafiltration in dead-end operation)), or by centrifugation (using a sedimentation centrifuge, disc stack centrifuge, hyrdo cyclone, or the like), or by precipitation of peptides and using relevant solid-liquid separation methods to harvest peptides from broth media by particle size fractionation. Recovery encompasses the isolation and / or purification of peptides.

[0087] Scoring: In the context of this invention, a score is calculated for each of one or more target polynucleotides. In one embodiment, the score is the sum of the products of the normalized relative abundance in each output channel and the sorting threshold score of the corresponding output channel. In a preferred embodiment, the score is calculated as described in Example 3.

[0088] Screenable Products: The term "screenable product" refers to a molecule detectable by the sensor device (600). Screenable products include, but are not limited to, fluorescent molecules (e.g., green fluorescent protein (GFP), mCherry, mVenus, DsRed, EGFP, Nile Red (9-(diethylamino)benzo[a]phenoxazine-5-one), fluorescent vitamins, DAPI (4',6-diamidinyl-2-phenylindole), and BIODIPY) and fluorescent molecules (e.g., fluorescent rhodamine). For example, screenable products are added to an emulsion or generated from a substrate through a process occurring in a droplet (e.g., during incubation). In another example, the screenable product is a polypeptide expressed in a droplet. In yet another example, the screenable product is a host cell in a droplet. In still another example, the screenable product comprises a light-absorbing molecule. In one embodiment, the light-absorbing molecule comprises p-nitroaniline (pNA).

[0089] For example, the amount of screenable product in a droplet can be inversely proportional to the amount of target peptide in the droplet and / or expressed by the host cell, for example, when the target peptide binds to or degrades the screenable product.

[0090] For example, the amount of screenable product in the droplet can be proportional to the amount of target peptide in the droplet and / or expressed by the host cell, for example, when the target peptide degrades the substrate leading to the formation of a screenable product, or when the screenable product (e.g., Nile Red) is incorporated into the host cell or a portion thereof (e.g., the host cell membrane or host cell wall).

[0091] Therefore, screenable products can be used, for example, as alternative indicators selected from one or more characteristics in a list of cell growth, cell division, target peptide expression, target peptide binding, target peptide stability, and target peptide activity. In some instances, more than one screenable product is present in the droplet, for example, to identify two or more distinct characteristics selected from the aforementioned features.

[0092] Sequence identity: The degree of association between two amino acid sequences or two nucleotide sequences is described by the parameter "sequence identity".

[0093] For the purposes of this invention, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. [Journal of Molecular Biology] 48: 443-453) is used to determine the sequence identity between two amino acid sequences as the output of "longest identity". This algorithm is implemented in the Niedel program using the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. [Trends in Genetics] 16: 276-277) (preferably version 6.6.0 or later). The parameters used are a vacancy opening penalty of 10, a vacancy extension penalty of 0.5, and an EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. For the Niedel program to report the longest identity, the non-brief (-nobrief) option must be specified on the command line. The Niedel-marked "longest identity" output is calculated as follows:

[0094] (Identical residues × 100) / (Alignment length - Total number of vacancies in the alignment)

[0095] For the purposes of this invention, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, ibid.) is used to determine the sequence identity between two polynucleotide sequences as the output of "longest identity," as implemented in the Niedel program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, ibid.) (preferably version 6.6.0 or later). The parameters used are a vacancy opening penalty of 10, a vacancy extension penalty of 0.5, and an EDNAFULL substitution matrix (EMBOSS version of NCBI NUC4.4). For the Niedel program to report the longest identity, the non-simplified option must be specified in the command line. The Niedel-marked "longest identity" output is calculated as follows:

[0096] (Identical deoxyribonucleotides × 100) / (Alignment length – Total number of vacancies in the alignment)

[0097] Signal peptide: A signal peptide is an amino acid sequence that attaches to the N-terminal portion of a protein and promotes its secretion outside the cell. The mature form of extracellular proteins lacks a signal peptide, which is cleaved during the secretion process.

[0098] Subsequence: The term "subsequence" refers to a polynucleotide in which one or more nucleotides are deleted from the 5' and / or 3' end of the coding sequence of a mature polypeptide; wherein the subsequence encodes a fragment with enzymatic activity.

[0099] Variant: The term "variant" refers to a polypeptide that has enzymatic activity and contains artificial mutations (i.e., substitutions, insertions (including extensions), and / or deletions (e.g., truncations)) at one or more positions. Substitution means replacing an amino acid occupying a position with a different amino acid; deletion means removing an amino acid occupying a position; and insertion means adding 1-5 amino acids (e.g., 1-3 amino acids, especially 1 amino acid) adjacent to and immediately following the amino acid occupying a position.

[0100] Wild-type: When referring to an amino acid or nucleic acid sequence, the term "wild-type" means that the amino acid or nucleic acid sequence is natural or naturally occurring. As used herein, the term "naturally occurring" refers to any substance found in nature (e.g., protein, amino acid, or nucleic acid sequences). Conversely, the term "non-naturally occurring" refers to any substance not found in nature (e.g., recombinant nucleic acid and protein sequences produced in a laboratory, or modifications of wild-type sequences). Detailed Implementation

[0101] In a first aspect, the present invention relates to a method for screening biological libraries, the method comprising the following steps:

[0102] a) Provide a microfluidic device including a droplet sorter (200) having at least three output channels (301, 302, 303).

[0103] b) Provides a droplet emulsion containing a target polynucleotide library and screenable products.

[0104] c) Determine the amount of screenable products from one or more droplets in the microfluidic device.

[0105] d) The droplet sorter (200) sorts the one or more droplets into the receiving output channels of the at least three output channels (301, 302, 303), wherein the receiving output channel is determined based on the amount of screenable product per droplet, and wherein the at least three receiving output channels receive a plurality of droplets, the plurality of droplets comprising an amount of screenable product above and / or below one or more predetermined threshold levels.

[0106] e) Identify one or more target polynucleotides present in the at least three output channels (301, 302, 303), and obtain the sequence data of the one or more target polynucleotides.

[0107] f) For one or more output channels (301, 302, 303), assign a score to each of the one or more target polynucleotides, wherein the score is calculated based on the abundance of each of the one or more target polynucleotides in one of the one or more output channels (301, 302, 302).

[0108] In one embodiment, the droplet emulsion contains one or more host cells.

[0109] In one embodiment, each host cell contains one or more target polynucleotides from the target polynucleotide library.

[0110] In one embodiment, in step b), each droplet contains at most one host cell or multiple host cells derived from the same parent host cell.

[0111] In one embodiment, in step b), each droplet contains at most one target polynucleotide.

[0112] In one embodiment, the screenable product is produced by the host cell.

[0113] In one embodiment, the screenable product is catalyzed by an enzyme, preferably an enzyme encoded by a target polynucleotide.

[0114] In one embodiment, the screenable product is encoded by one or more target polynucleotides.

[0115] In one embodiment, the screenable product provides the production of a polypeptide expressed by a host cell.

[0116] In one embodiment, the screenable product is generated by a polypeptide encoded by one or more target polynucleotides.

[0117] In one embodiment, the screenable product is a polypeptide expressed by the host cell.

[0118] In one embodiment, the screenable product is an enzyme.

[0119] In one embodiment, the enzyme is expressed by the host cell.

[0120] In one embodiment, the enzyme is selected from the list of: hydrolases, isomerases, ligases, lyases, oxidoreductases, or transferases, such as aminopeptidase, amylase, glycosylase, carboxypeptidase, catalase, cellobiase, cellulase, chitinase, keratinase, cyclodextrin glycosyltransferase, deoxyribonuclease, endoglucanase, esterase, α-galactosidase, β-galactosidase, glucosylamylase, α-glucosidase, β-glucosidase, invertase, laccase, lipase, mannosidase, polysaccharidase, oxidase, pectinase, peroxidase, phytase, polyphenol oxidase, proteolytic enzyme, ribonuclease, transglutaminase, xylanase, or β-xylosidase.

[0121] In one embodiment, the screenable product is degraded by the host cell.

[0122] In one embodiment, the screenable product is degraded by a polypeptide encoded by one or more target polynucleotides.

[0123] In one embodiment, the screenable product is degraded by a peptide expressed by the host cell.

[0124] In one embodiment, the screenable product is an enzyme substrate, preferably selected from the list of: hydrolases, isomerases, ligases, lyases, oxidoreductases, or transferases, such as aminopeptidase, amylase, glycosylase, carboxypeptidase, catalase, cellobiase, cellulase, chitinase, keratinase, cyclodextrin glycosyltransferase, deoxyribonuclease, endoglucanase, esterase, α-galactosidase, β-galactosidase, glucosylamylase, α-glucosidase, β-glucosidase, invertase, laccase, lipase, mannosidase, polysaccharidase, oxidase, pectinase, peroxidase, phytase, polyphenol oxidase, proteolytic enzyme, ribonuclease, transglutaminase, xylanase, or β-xylosidase.

[0125] In one embodiment, the screenable product is a fluorescent product.

[0126] In one embodiment, the fluorescent product is derived from a fluorescent substrate by an enzyme encoded by the target polynucleotide.

[0127] In one embodiment, the amount of screenable product is inversely proportional to one or more of cell number, cell growth, cell division, cell viability, or cell growth rate.

[0128] In one embodiment, the amount of screenable product is proportional to one or more of cell number, cell growth, cell division, cell viability, or cell growth rate.

[0129] In one embodiment, the screenable product comprises or consists of one or more host cells.

[0130] In one embodiment, the screenable product comprises or consists of substantially all host cells in the droplet.

[0131] In one embodiment, the screenable product is the product of an enzymatic reaction, preferably the product of a reaction catalyzed by an enzyme selected from the list of enzymes: hydrolases, isomerases, ligases, lyases, oxidoreductases, or transferases, such as aminopeptidase, amylase, glycosylase, carboxypeptidase, catalase, cellobiase, cellulase, chitinase, keratinase, cyclodextrin glycosyltransferase, deoxyribonuclease, endoglucanase, esterase, α-galactosidase, β-galactosidase, glucosylamylase, α-glucosidase, β-glucosidase, invertase, laccase, lipase, mannosidase, polysaccharidase, oxidase, pectinase, peroxidase, phytase, polyphenol oxidase, proteolytic enzyme, ribonuclease, transglutaminase, xylanase, or β-xylosidase.

[0132] In one embodiment, the score is proportional to the number of identical DNA sequences of the first target polynucleotide present in the output channel, for example, the score is normalized relative to that number.

[0133] In one embodiment, the score is the total number of identical DNA sequences of the first target polynucleotide present in the output channel.

[0134] In one embodiment, the score is proportional to the number of identical DNA sequences of the second target polynucleotide present in the output channel, for example, the score is normalized relative to that number.

[0135] In one embodiment, the score is the total number of identical DNA sequences of the second target polynucleotide present in the output channel.

[0136] In one embodiment, the microfluidic device includes an incubation zone (500).

[0137] In one embodiment, the incubation area (500) is located upstream of the droplet sorter (200) and / or upstream of one or more sorting devices (401, 402).

[0138] In one embodiment, the method includes incubating a droplet emulsion under conditions that allow cell growth and / or allow DNA to be transcribed from DNA into RNA and / or allow RNA to be translated into polypeptides, preferably, the incubation is performed before step c).

[0139] In one embodiment, incubation does not occur within the microfluidic chip.

[0140] In one embodiment, incubation occurs on and / or within the microfluidic device.

[0141] In one embodiment, after incubation, the cells contained in a droplet are genetically identical, i.e., these cells are derived from a parent host cell, preferably the same parent host cell.

[0142] In one embodiment, the droplet sorter includes one or more sensor elements (600), which are preferably located downstream of the incubation area (500) and / or upstream of the sorting devices (401, 402).

[0143] In one embodiment, the one or more sensor devices (600) include a fluorescence sensor.

[0144] In one embodiment, the one or more sensor devices (600) include a light-absorbing sensor.

[0145] In one embodiment, the one or more sensor devices (600) include an image sensor, such as a CMOS sensor, a CCD sensor, or a PMT sensor.

[0146] In one embodiment, the one or more sensor devices (600) include NEMS (nanoelectromechanical systems) sensors.

[0147] In one embodiment, the one or more sensor devices (600) include a mass analyzer suitable for mass spectrometry, such as a quadrupole mass analyzer, a TOF mass analyzer, an ion trap mass analyzer, an orbit trap mass analyzer, a sector magnetic field mass analyzer, a Q-TOF mass analyzer, or an FT-ICR mass analyzer.

[0148] In one embodiment, step e) includes DNA amplification of one or more target polynucleotides within each output channel.

[0149] In one embodiment, DNA amplification is performed using a PCR method.

[0150] In one embodiment, DNA amplification is performed using the ddPCR (droplet digital PCR) method.

[0151] In one embodiment, step e) includes DNA sequencing of the one or more target polynucleotides, for example, sequencing after PCR amplification or sequencing via nanopore sequencing.

[0152] In one embodiment, during step e), the one or more target polynucleotides are identified by DNA barcoding.

[0153] In one embodiment, the droplet sorter (200) includes one or more sorting devices (401, 402).

[0154] In one embodiment, the one or more sorting devices include or consist of one or more electrodes, one or more acoustic generators, one or more valves, and / or one or more pressure-controlled outlets.

[0155] In one embodiment, the one or more sorting devices include at least two electrodes.

[0156] In one embodiment, the one or more sorting devices consist of an electrode.

[0157] In one embodiment, the one or more sorting devices consist of two electrodes.

[0158] In one embodiment, the biological library contains or is composed of wild-type cells with different genotypes and / or different phenotypes.

[0159] In one embodiment, the biological library contains or is composed of recombinant cells.

[0160] In one embodiment, the biological library encodes different variants of the same target polypeptide, preferably an enzyme.

[0161] In one embodiment, the biological library encodes different signal peptide variants.

[0162] In one embodiment, the biological library encodes different promoter variants.

[0163] In one embodiment, the biological library contains different codon-optimized DNA sequences with the same amino acid sequence encoding a target polypeptide (e.g., a signal peptide and / or an enzyme).

[0164] In one embodiment, the target polynucleotide encodes the target polypeptide.

[0165] In one embodiment, the biological library contains or is composed of a variety of target polynucleotides, each of which encodes a variant of the target polypeptide.

[0166] In one embodiment, the target polynucleotide comprises a first target polynucleotide encoding a control sequence and a second target polynucleotide encoding a target polypeptide.

[0167] In one embodiment, the biological library contains or is composed of a variety of target polynucleotides, each of which encodes a variant of a control sequence.

[0168] In one embodiment, the control sequence is a promoter sequence, signal peptide, leader sequence, polyadenylation sequence, propeptide sequence, or transcription terminator.

[0169] In one embodiment, the target polynucleotide comprises a first target polynucleotide encoding a signal peptide and a second target polynucleotide encoding a target polypeptide, wherein the first target polynucleotide is operatively linked to and located upstream of the second target polynucleotide.

[0170] In one embodiment, the target polynucleotide comprises a first target polynucleotide containing a promoter sequence and a second target polynucleotide encoding a target polypeptide, wherein the first target polynucleotide is operatively linked to and located upstream of the second target polynucleotide.

[0171] In one embodiment, the biological library contains multiple variants of the same second-purpose polynucleotide and the first-purpose polynucleotide.

[0172] In one embodiment, the biological library contains multiple variants of the same first-purpose polynucleotide and second-purpose polynucleotide.

[0173] In one embodiment, the first target polynucleotide is heterologous to the second target polynucleotide.

[0174] In one embodiment, the first target polynucleotide is endogenous to the second target polynucleotide.

[0175] In one embodiment, the one or more target polynucleotides comprise a promoter, a polynucleotide encoding a signal peptide, a polynucleotide encoding a target polypeptide, or a natural host cell gene.

[0176] In one embodiment, the target polynucleotide is essentially the entire genome of the host cell.

[0177] In one embodiment, the one or more targeted polynucleotides are heterologous to the host cell.

[0178] In one embodiment, the one or more targeted polynucleotides are endogenous to the host cell.

[0179] In one embodiment, the first target polynucleotide is heterologous to the host cell.

[0180] In one embodiment, the first target polynucleotide is endogenous to the host cell.

[0181] In one embodiment, the second target polynucleotide is heterologous to the host cell.

[0182] In one embodiment, the second target polynucleotide is endogenous to the host cell.

[0183] In one embodiment, the first and second target polynucleotides are heterologous to the host cell.

[0184] In one embodiment, the first and second target polynucleotides are endogenous to the host cell.

[0185] In one embodiment, the one or more target polynucleotides encode a target polypeptide.

[0186] In one embodiment, the target polypeptide is an enzyme, nanobody, antibody, antibody fragment, fluorescent polypeptide (e.g., GFP), or α-lactalbumin.

[0187] In one embodiment, the amount of screenable product in the droplet is proportional to the amount of polypeptide encoded by the one or more target polynucleotides.

[0188] In one embodiment, the amount of screenable product in the droplet is inversely proportional to the amount of polypeptide encoded by the one or more target polynucleotides.

[0189] In one embodiment, the biological library contains at least 100 different one or more target polynucleotides, at least 200 different one or more target polynucleotides, at least 500 different one or more target polynucleotides, at least 1,000 different one or more target polynucleotides, at least 2,000 different one or more target polynucleotides, at least 3,000 different one or more target polynucleotides, at least 5,000 different one or more target polynucleotides, at least 10,000 different one or more target polynucleotides, at least 100,000 different one or more target polynucleotides, at least 1,000,000 different one or more target polynucleotides, at least 10,000,000 different one or more target polynucleotides, at least 50,000,000 different one or more target polynucleotides, or at least 100,000,000 different target polynucleotides.

[0190] In one embodiment, the biological library contains at least 100 different host cells, at least 200 different host cells, at least 500 different host cells, at least 1,000 different host cells, at least 2,000 different host cells, at least 3,000 different host cells, at least 5,000 different host cells, at least 10,000 different host cells, at least 100,000 different host cells, at least 200,000 different host cells, at least 500,000 different host cells, at least 1,000,000 different host cells, at least 5,000,000 different host cells, at least 10,000,000 different host cells, or at least 100,000,000 different host cells.

[0191] In one embodiment, the amount of screenable product in the droplet is proportional to one or more of the following: the stability of the target peptide, the transcription of the target peptide, the translation of the target peptide, the secretion of the target peptide, the yield of the target peptide, the binding strength of the target peptide to the target molecule, and the activity of the target peptide.

[0192] In one embodiment, the amount of screenable product in the droplet is inversely proportional to one or more of the following: the stability of the target peptide, the transcription of the target peptide, the translation of the target peptide, the secretion of the target peptide, the yield of the target peptide, the binding strength of the target peptide to the target molecule, and the activity of the target peptide.

[0193] In one embodiment, the amount of screenable product in the droplet is proportional to one or more of the following: cell number, host cell viability, host cell division rate, host cell growth rate, host cell size, and host cell protein secretion.

[0194] In one embodiment, the amount of screenable product in the droplet is inversely proportional to one or more of the following: cell number, host cell viability, host cell division rate, host cell growth rate, host cell size, and host cell protein secretion.

[0195] In one embodiment, the one or more droplets contain a substrate.

[0196] In one embodiment, the substrate comprises or consists of a screenable product.

[0197] In one embodiment, the substrate is a fluorescent substrate.

[0198] In one embodiment, the substrate is rhodamine, which is capable of producing fluorescence.

[0199] In one embodiment, the substrate is a fluorescent dye.

[0200] In one embodiment, the substrate is a fluorescent substrate.

[0201] In one embodiment, the substrate comprises a fluorophore (e.g., fluorescein) or fluorescein-labeled starch. In one embodiment, the substrate is Nile Red.

[0202] In one embodiment, the substrate is DAPI (4',6-diamidinyl-2-phenylindole).

[0203] In one embodiment, prior to optional incubation, each droplet contains up to 0.01 cells, up to 0.02 cells, up to 0.03 cells, up to 0.04 cells, up to 0.05 cells, up to 0.06 cells, up to 0.07 cells, up to 0.08 cells, up to 0.09 cells, up to 0.1 cells, up to 0.2 cells, up to 0.3 cells, up to 0.4 cells, up to 0.5 cells, up to 0.6 cells, or up to 0.7 cells; preferably an average occupancy of up to 0.1 cells.

[0204] In one embodiment, each droplet contains up to 0.01, up to 0.02, up to 0.03, up to 0.04, up to 0.05, up to 0.06, up to 0.07, up to 0.08, up to 0.09, up to 0.1, up to 0.2, up to 0.3, up to 0.4, up to 0.5, up to 0.6, or up to 0.7 target polynucleotides; preferably an average content of up to 0.1 target polynucleotides.

[0205] In one embodiment, droplet sorting is facilitated by an electric field generated by one or more electrodes (401, 402) adjacent to the droplet sorter.

[0206] In one embodiment, droplet sorting is facilitated by acoustic waves generated by one or more acoustic generators (401, 402) adjacent to the droplet sorter.

[0207] In one embodiment, droplet sorting is facilitated by localized pressure variations generated by one or more pressure-controlled outlets (401, 402) adjacent to the droplet sorter, for example, wherein the one or more pressure-controlled outlets are included in one or more output channels.

[0208] In one embodiment, the amount of screenable product in step c) is determined using fluorescence-based signal, absorbance, Raman spectroscopy, mass spectrometry (MS), or MALDI-MS.

[0209] In one embodiment, the relative and / or absolute amount of the screenable product / droplet is determined by one or more sensor elements (600).

[0210] In one embodiment, after step d), one or more output channels contain at least 10,000 droplets, at least 50,000 droplets, at least 100,000 droplets, at least 500,000 droplets, at least 1,000,000 droplets, at least 2,000,000 droplets, at least 5,000,000 droplets, at least 10,000,000 droplets, or at least 100,000,000 droplets.

[0211] In one embodiment, the droplet sorter includes at least four output channels, at least five output channels, at least six output channels, at least seven output channels, at least eight output channels, at least nine output channels, or at least ten output channels.

[0212] In one embodiment, the host cell is a yeast host cell, such as cells of the genera *Candida*, *Hansenula*, *Kluyveromyces*, *Pichia*, *Saccharomyces*, *Schizosaccharomyces*, or *Yarrowia*, such as *Kluyveromyces lactis*, *Saccharomyces carlsbergensis*, *Saccharomyces cerevisiae*, *Saccharomyces diastaticus*, *Saccharomyces douglasii*, *Saccharomyces kluyveri*, *Saccharomyces norbensis*, *Saccharomyces oviformis*, or *Yarrowia lipolytica*.

[0213] In one embodiment, the host cell is a filamentous fungal host cell, such as *Acremonium*, *Aspergillus*, *Aureobasidium*, *Bjerkandera*, *Ceriporiopsis*, *Chrysosporium*, *Coprinus*, *Coriolus*, *Cryptococcus*, *Filibasidium*, *Fusarium*, *Humicola*, *Magnaporthe*, *Mucor*, *Myceliophthora*, and *Neomycetes*. Cells of the genera *Neocallimastix*, *Neurospora*, *Paecilomyces*, *Penicillium*, *Phanerochaete*, *Phlebia*, *Piromyces*, *Pleurotus*, *Schizophyllum*, *Talaromyces*, *Thermoascus*, *Thielavia*, *Tolypocladium*, *Trametes*, or *Trichoderma*, especially *Aspergillus buergerianum*. Aspergillus awamori, Aspergillus foetidus, Aspergillus fumigatus, Aspergillus japonicus, Aspergillus nidulans, Aspergillus niger, Aspergillus oryzae, Bjerkanderaadusta, Ceriporiopsis aneirina, Ceriporiopsis caregiea, Ceriporiopsis gilvescens, Ceriporiopsis pannocinta, Ceriporiopsis rivulosa, Ceriporiopsis subrufa, Ceriporiopsis insectisThe following fungi are listed: *Chrysosporium subvermispora*, *Chrysosporium inops*, *Chrysosporium keratinophilum*, *Chrysosporium lucknowense*, *Chrysosporium merdarium*, *Chrysosporium pannicola*, *Chrysosporium queenslandicum*, *Chrysosporium tropicum*, *Chrysosporium zonatum*, *Coprinus cinereus*, *Coriolus hirsutus*, *Fusarium bactridioides*, *Fusarium cerealis*, *Fusarium crookwellense*, *Fusarium culmorum*, and *Fusarium graminearum*. Fusarium graminearum, Fusarium graminum, Fusarium heterosporum, Fusarium negundi, Fusarium oxysporum, Fusarium reticulatum, Fusarium roseum, Fusarium sambucinum, Fusarium sarcochroum, Fusarium sporotrichioides, Fusarium sulphureum, Fusarium torulosum, Fusarium trichothecioides, Fusarium venenatum, Humicola insolens, Humicola lanuginosa, Mucor miehei, Myceliophthora Thermophila, Neurospora crassa, Penicillium purpurogenum, PhanerochaeteThe cells of *Trichoderma chrysosporium*, *Phlebia radiata*, *Pleurotus eryngii*, *Talaromyces emersonii*, *Thielavia terrestris*, *Trametes villosa*, *Trametes versicolor*, *Trichoderma harzianum*, *Trichoderma koningii*, *Trichoderma longibrachiatum*, *Trichoderma reesei*, or *Trichoderma viride*.

[0214] In one embodiment, the host cell is a prokaryotic host cell, for example, a Gram-positive cell selected from the group consisting of: Bacillus, Clostridium, Enterococcus, Geobacillus, Lactobacillus, Lactococcus, Oceanobacillus, Staphylococcus, Streptococcus, or Streptomyces cells, or a Gram-negative bacterium selected from the group consisting of: Campylobacter, Escherichia coli (E. coli). *Clostridium*, *Flavobacterium*, *Fusobacterium*, *Helicobacter*, *Ilyobacter*, *Neisseria*, *Pseudomonas*, *Salmonella*, and *Ureaplasma* cells, such as *Bacillus alkalophilus*, *Bacillus amyloliquefaciens*, *Bacillus brevis*, *Bacillus circulans*, *Bacillus clausii*, *Bacillus scoagulans*, *Bacillus firmus*, *Bacillus lautus*, *Bacillus lentus*, and *Bacillus licheniformis*. The following bacteria are listed: *Bacillus licheniformis*, *Bacillus megaterium*, *Bacillus pumilus*, *Bacillus stearothermophilus*, *Bacillus subtilis*, *Bacillus thuringiensis*, *Streptococcus equisimilis*, *Streptococcus pyogenes*, *Streptococcus uberis*, and *Streptococcus equi subsp.*The following fungi are mentioned: *Zooepidemicus*, *Streptomyces achromogenes*, *Streptomyces avermitilis*, *Streptomyces coelicolor*, *Streptomyces griseus*, and *Streptomyces lividans*.

[0215] In one embodiment, the host cell is Bacillus subtilis.

[0216] In one embodiment, the host cell is Bacillus licheniformis.

[0217] In one embodiment, the host cell is Trichoderma reesei.

[0218] In one embodiment, the host cell is Aspergillus niger.

[0219] In one embodiment, the host cell is Aspergillus oryzae.

[0220] In one embodiment, the host cell is a genus of Bifidobacterium, such as Bifidobacterium animalis or Bifidobacterium animalis subsp. lactis.

[0221] It is conceivable that the method of the present invention can use isolated polynucleotide libraries and in vitro expression systems, as well as recombinant cells and / or wild-type cells. However, a preferred embodiment of the present invention is a setup that includes host cells, i.e., wherein the library is contained within the host cells.

[0222] In a preferred embodiment, each individual polynucleotide of the library resides in its own separate host cell. Thus, in one embodiment, each droplet in step (b) of the first aspect contains at most a single host cell, which may optionally be incubated to grow into multiple cells prior to determining the amount of screenable product in step (c).

[0223] For example, the characteristics of one or more target polynucleotides in a biological library can be qualitatively or quantitatively determined by detecting the conversion of an enzyme substrate into a detectable or quantifiable enzyme product (= screenable product). For example, adding a fluorescent enzyme substrate transforms it into a fluorescent enzyme product (screenable product) to be detected or measured.

[0224] Therefore, in a preferred embodiment of the invention, the substrate used for one or more enzymes is fluorescent, and the enzyme activity converts the fluorescent substrate into a fluorescent product (a screenable product).

[0225] Once a feature of particular interest is detected in a selected droplet or in multiple droplets collected in one or more output channels, it is necessary to identify the polynucleotide library members within the droplet. Typically, polynucleotides are identified by DNA sequencing.

[0226] When constructing a library, polynucleotides can also be equipped with identification sequence tags to act as "barcodes," thus eliminating the need for sequencing. Barcode-based identification immediately reveals the DNA sequence of the polynucleotide, thereby identifying it.

[0227] In a preferred aspect of the invention, the one or more target polynucleotides are identified in step e) by DNA sequencing of the one or more target polynucleotides.

[0228] In a first aspect of the invention, several equal portions of the solution are introduced into a droplet. The volume of the equal portions is typically much smaller than the droplet, but their size can, in principle, be equal to or even larger than the volume of the droplet. In the examples below, the equal portions are significantly smaller than the droplet. There are many methods for introducing equal portions into a droplet, or in other words, for merging or coalescing two droplets, in a microfluidic device.

[0229] For example, WO 2007 / 061448 discloses the design of a microfluidic device capable of applying an electric field to merge or coalesce two or more droplets. Another method for introducing small aliquots of aqueous liquid into droplets in a microfluidic device is called “microinjection”, and is disclosed, for example, in WO 2010 / 151776.

[0230] In the example below, the aliquot sample is introduced into the droplet by applying an electric field to merge or coalesce the aliquot sample and the droplet.

[0231] Accordingly, in a preferred embodiment of the first aspect, the aliquot sample is introduced into the droplet by applying an electric field or by injection to combine or coalesce the aliquot sample and the droplet.

[0232] Figure 1An embodiment of the invention is illustrated, wherein the apparatus includes a droplet sorter (200) having five output channels (301, 302, 303, 304, 305) and an incubation zone (500). The apparatus also includes a sensor (600) and two electrodes (401, 402). Droplets containing host cells and screenable products are shown in circles. Schematably, the amount of screenable product present in each droplet is represented by black fill. Schematably, the amount of black is proportional to the amount of screenable product present in each droplet.

[0233] The flow from the incubation zone (500) to the output channels (301, 302, 303, 304, 305) allows droplets to pass through a sensor (600) that determines the amount of screenable product in each droplet (step c). The sensor (600) conveys a certain amount of screenable product to electrodes (401, 402). Based on the information about the amount of screenable product in each droplet, an electric field is applied to the electrodes, allowing the droplets to be sorted into one of the five output channels (step d). In this example, droplets with a large amount of screenable product are sorted into the top output channel (305) and collected in pool 5, while droplets with little or no screenable product are sorted into the lowest output channel (301) and collected in pool 1. Furthermore, droplets with an intermediate amount of screenable product are sorted into the remaining three output channels (302, 303, and 304) and collected in pools 2-4. Designs with three or more output channels allow for parallel sorting to multiple output channels using multiple predetermined thresholds, with no sample volume loss.

[0234] Library design and variant sequence generation

[0235] The method of the present invention utilizes biological libraries of variants (amino acid sequences, DNA sequences, and / or host cell variants), but can also generate synthetic variant sequences based on the readout of the method.

[0236] On the one hand, synthetic sequence variants are produced by substituting, deleting, or adding one or more amino acids (for polypeptide variants) or one or more nucleotides (for DNA sequence variants).

[0237] On one hand, peptide variants are derived from mature peptides through the substitution, deletion, or addition of one or more amino acids. On the other hand, the number of amino acid substitutions, deletions, and / or insertions introduced into the peptide can be up to 15, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. Amino acid changes can be minor, i.e., conserved amino acid substitutions or insertions that do not significantly affect protein folding and / or activity; typically small deletions of 1–30 amino acids; small N-terminal or C-terminal extensions, such as methionine residues at the N-terminus; small linker peptides of up to 20–25 residues; or small extensions that facilitate purification by altering net charge or another function, such as a polyhistidine fragment, an antigenic epitope, or a binding module.

[0238] Essential amino acids in peptides can be identified using procedures known in the art, such as site-directed mutagenesis or alanine scanning mutagenesis (Cunningham and Wells, 1989, Science 244: 1081-1085). In the latter technique, a single alanine mutation is introduced at each residue in the molecule, and the enzyme activity of the resulting molecule is tested to identify the amino acid residues critical to the activity of that molecule. See also Hilton et al., 1996, J. Biol. Chem. 271: 4699-4708. The active sites of enzymes or other biological interactions can also be determined by physical analysis of the structure, such as by techniques like nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, along with mutagenesis of the amino acids at the putative contact sites. See, for example, de Vos et al., 1992, Science 255: 306-312; Smith et al., 1992, J. Mol. Biol. 224: 899-904; Wlodaver et al., 1992, FEBS Lett. 309: 59-64. The identity of essential amino acids can also be inferred from alignment with related peptides, and / or from sequence homology and conserved catalytic mechanisms with related peptides or peptide / protein families from a common ancestor (typically possessing similar three-dimensional structures, functions, and significant sequence similarities). Alternatively or additionally, protein structure prediction tools can be used for protein structure modeling to identify essential amino acids and / or active sites of peptides. See, for example, Jumper et al., 2021, “Highly accurate protein structure prediction with AlphaFold”, Nature 596: 583-589.

[0239] Using known mutagenesis, recombination, and / or tampering methods, followed by relevant screening procedures, one or more amino acid substitutions, deletions, and / or insertions can be made and tested, such as those disclosed by Reidhaar-Olson and Sauer, 1988, Science 241: 53-57; Bowie and Sauer, 1989, Proc. Natl. Acad. Sci. USA 86: 2152-2156; WO 95 / 17413; or WO 95 / 22625. Other methods that can be used include error-prone PCR, phage display (e.g., Lowman et al., 1991, Biochemistry 30: 10832-10837; US 5,223,409; WO 92 / 06204), and region-directed mutagenesis (Derbyshire et al., 1986, Gene 46: 145; Ner et al., 1988, DNA 7: 127).

[0240] Mutagenesis / recombination methods can be combined with high-throughput, automated screening methods to detect the activity of cloned, mutagenesis-encoded peptides expressed by host cells (Ness et al., 1999, Nature Biotechnology 17: 893-896). Mutagenesis-encoded DNA molecules encoding active peptides can be recovered from host cells and rapidly sequenced using standard methods in the art. These methods allow for the rapid determination of the importance of individual amino acid residues within the peptide.

[0241] On the one hand, this is achieved by replacing, deleting, or adding one or more nucleic acid-derived DNA variant sequences.

[0242] These polynucleotides can also be constructed by introducing nucleotide substitutions that do not cause a change in the amino acid sequence of the polypeptide but correspond to the codons used by the host organism intended to produce the enzyme, or by introducing nucleotide substitutions that may produce different amino acid sequences. For a general description of nucleotide substitutions, see, for example, Ford et al., 1991, Protein Expression and Purification 2: 95-107.

[0243] DNA sequences used for library design can be obtained from any genus of microorganisms.

[0244] Similarly, polypeptide sequences comprising, for example, enzymes, signal peptides, or nanobodies can be obtained from any genus of microorganisms. For the purposes of this invention, the term "obtained from," as used herein in conjunction with a given source, should mean that the polypeptide encoded by the polynucleotide is produced by that source or by a strain that has inserted the polynucleotide of this invention. In one aspect, polypeptides obtained from a given source are secreted extracellularly.

[0245] It should be understood that, for the aforementioned species, this invention covers complete and incomplete stages, as well as other taxonomic equivalents, such as asexual forms, regardless of their known species names. Those skilled in the art will readily identify the appropriate equivalents.

[0246] The probes mentioned above can be used to identify and obtain polypeptides from other sources, including microorganisms isolated from nature (e.g., soil, compost, water, etc.) or DNA samples obtained directly from natural materials (e.g., soil, compost, water, etc.). Techniques for directly isolating microorganisms and DNA from natural habitats are well known in the art. The polynucleotide encoding the polypeptide can then be obtained by similarly screening a library of genomic DNA or cDNA from another microorganism or a mixed DNA sample. Once the polynucleotide encoding the polypeptide has been detected with the probe, it can be isolated or cloned using techniques known to those skilled in the art (see, for example, Davis et al., 2012, Basic Methods in Molecular Biology, Elsevier).

[0247] Screening biological libraries containing control sequences

[0248] The present invention also relates to screening biological libraries, wherein the biological library comprises or is composed of a plurality of target polynucleotides, each target polynucleotide comprising a first target polynucleotide encoding a control sequence and a second target polynucleotide encoding a target polypeptide. In one embodiment, the biological library comprises a plurality of variants of the control sequence.

[0249] Preferably, the second target polynucleotide is operatively linked to one or more control sequences (the first target polynucleotide), which, under conditions compatible with the control sequences, guide the expression of the second target polynucleotide in a suitable host cell.

[0250] Control sequences can be manipulated in a variety of ways to provide expression of the target peptide and / or create control sequence libraries. Depending on the expression vector, manipulation of the control sequence prior to its insertion into the vector is desirable or necessary. Techniques for modifying control sequences using recombinant DNA methods are well known in the art.

[0251] promoter

[0252] The control sequence may be a promoter, i.e., a polynucleotide recognized by the host cell for expressing the polypeptide encoding the present invention. The promoter contains a transcriptional control sequence that mediates polypeptide expression. The promoter can be any polynucleotide exhibiting transcriptional activity in the host cell, including mutant promoters, truncated promoters, and heterozygous promoters, and can be a gene encoding an extracellular or intracellular polypeptide that is homologous or heterologous to the host cell.

[0253] Examples of suitable promoters for guiding the transcription of polynucleotides in bacterial host cells are described in Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Lab, NY; Davis et al., 2012, ibid.; and Song et al., 2016, PLOS One 11(7): e0158447.

[0254] Examples of suitable promoters for guiding the transcription of polynucleotides in filamentous fungal host cells are promoters obtained from Aspergillus, Fusarium, Rhizomucor, and Trichoderma cells, such as those described in: Mukherjee et al., 2013, “Trichoderma: Biology and Applications” and Schmoll and Dattenböck, 2016, “Gene Expression Systems in Fungi: Advancements and Applications”, Fungal Biology.

[0255] Examples of useful promoters for expression in yeast hosts are described in: Smolke et al., 2018, “Synthetic Biology: Parts, Devices and Applications” (Chapter 6: Constitutive and Regulated Promoters in Yeast: How to Design and Make Use of Promoters in S. cerevisiae) and Schmoll and Dattenböck, 2016, “Gene Expression Systems in Fungi: Advancements and Applications”, Fungal Biology.

[0256] Termination

[0257] The control sequence can also be a transcription terminator that is recognized by the host cell to terminate transcription. The terminator is operatively linked to the 3' end of a polynucleotide encoding a polypeptide. Any terminator that is functional in the host cell can be used in this invention.

[0258] Preferred terminators for bacterial host cells can be obtained from genes in the following: Bacillus clausii alkaline protease (aprH), Bacillus licheniformis α-amylase (amyL), and Escherichia coli ribosomal RNA (rrnB).

[0259] Preferred terminators for filamentous fungal host cells can be derived from Aspergillus or Trichoderma species, such as genes derived from Aspergillus niger glucosidase, Trichoderma reesei β-glucosidase, Trichoderma reesei cellobiase I, and Trichoderma reesei endoglucanase I, as described in the following terminators: Mukherjee et al., 2013, “Trichoderma: Biology and Applications” and Schmoll and Dattenböck, 2016, “Gene Expression Systems in Fungi: Advancements and Applications”, Fungal Biology.

[0260] Preferred terminators for yeast host cells can be obtained from genes including *Saccharomyces cerevisiae* enolase, *Saccharomyces cerevisiae* cytochrome C (CYC1), and *Saccharomyces cerevisiae* glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are described by Romanos et al., 1992, Yeast [Yeast] 8: 423-488.

[0261] mRNA stabilizers

[0262] Control sequences can also be mRNA stabilizer regions downstream of the promoter and upstream of the gene's coding sequence, which increase the expression of the gene.

[0263] Examples of suitable mRNA stabilizer regions were obtained from the Bacillus thuringiensis cryIIIA gene (WO 94 / 25612) and the Bacillus subtilis SP82 gene (Hue et al., 1995, J. Bacteriol. [Journal of Bacteriology] 177: 3465-3471).

[0264] Examples of mRNA stabilizer regions in fungal cells are described in Geisberg et al., 2014, Cell [Cell] 156(4): 812-824 and Morozov et al., 2006, Eukaryotic Cell [Eukaryotic Cell] 5(11): 1838-1846.

[0265] Leader sequence

[0266] The control sequence can also be a leader sequence, i.e., an untranslated region of mRNA that is important for translation in the host cell. The leader sequence is operatively linked to the 5' end of a polynucleotide encoding a polypeptide. Any leader sequence that is functional in the host cell can be used.

[0267] The appropriate leader sequence for bacterial host cells is described by Hambraeus et al., 2000, Microbiology 146(12): 3051-3059 and Kaberdin and Bläsi, 2006, FEMS Microbiol. Rev. 30(6): 967-979.

[0268] Preferred leader sequences for filamentous fungal host cells can be obtained from genes of Aspergillus oryzae TAKA amylase and Aspergillus nidulans triphosphate isomerase.

[0269] The appropriate leader sequence for yeast host cells can be obtained from the following genes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae α-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).

[0270] polyadenylation sequence

[0271] The control sequence can also be a polyadenylation sequence, which is a sequence operatively linked to the 3' end of a polynucleotide that is recognized by the host cell during transcription as a signal to add polyadenylation residues to the transcribed mRNA. Any polyadenylation sequence that is functional in the host cell can be used.

[0272] Preferred polyadenylated sequences for filamentous fungal host cells are obtained from the following genes: Aspergillus nidulans o-aminobenzoic acid synthase, Aspergillus niger glucosyl amylase, Aspergillus niger α-glucosidase, Aspergillus oryzae TAKA amylase, and Fusarium oxysporum trypsin-like protease.

[0273] Useful polyadenylation sequences in yeast host cells are described by Guo and Sherman, 1995, Mol. Cellular Biol. [Molecular Cell Biology] 15: 5983-5990.

[0274] signal peptide

[0275] As illustrated in the examples, the control sequence can also be a signal peptide coding region encoding a signal peptide linked to the N-terminus of a polypeptide and guiding the polypeptide into the cell's secretion pathway. The 5' end of the polynucleotide coding sequence itself may contain a signal peptide coding sequence naturally linked to the coding sequence segment of the polypeptide within the translation reading frame. Alternatively, the 5' end of the coding sequence may contain a signal peptide coding sequence that is heterologous to the coding sequence. In cases where the coding sequence does not naturally contain a signal peptide coding sequence, a heterologous signal peptide coding sequence may be required. Alternatively, a heterologous signal peptide coding sequence may simply replace the natural signal peptide coding sequence to enhance polypeptide secretion. Any signal peptide coding sequence that guides the expressed polypeptide into the host cell's secretion pathway can be used.

[0276] The effective signal peptide coding sequences of bacterial host cells are obtained from the signal peptide coding sequences of the following genes: maltose amylase produced by Bacillus NCIB 11837, Bacillus subtilis protease, Bacillus β-lactamase, Bacillus thermophilus α-amylase, Bacillus thermophilus neutral proteases (nprT, nprS, nprM), and Bacillus subtilis prsA. Other signal peptides are described by Freudl, 2018, Microbial Cell Factories 17: 52.

[0277] The effective signal peptide coding sequences for filamentous fungal host cells are obtained from the signal peptide coding sequences of the following genes: Aspergillus niger neutral amylase, Aspergillus niger glucosylase, Aspergillus oryzae TAKA amylase, Aspergillus oryzae cellulase, Aspergillus oryzae endoglucanase V, Aspergillus pubescens lipase, and Rhizopus oryzae aspartic protease, as described by Xu et al., 2018, Biotechnology Letters 40: 949-955.

[0278] Useful signal peptides in yeast host cells are acquired from genes of *Saccharomyces cerevisiae* α-factor and *Saccharomyces cerevisiae* invertase. Sequences encoding other useful signal peptides are described above by Romanos et al., 1992.

[0279] propeptide

[0280] The control sequence can also be a propeptide-coding sequence encoding the propeptide located at the N-terminus of the polypeptide. The resulting polypeptide is called a proenzyme or propeptide progenitor (or, in some cases, a zymogen). The propeptide progenitor is usually inactive and can be converted into an active polypeptide by catalytic cleavage or autocatalytic cleavage of the propeptide progenitor. Propeptide-coding sequences can be obtained from genes of the following: Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Thermophilus laccase (WO 95 / 33836), Rhizopus oryzae aspartic protease, and Saccharomyces cerevisiae α-factor.

[0281] When both a signal peptide sequence and a propeptide sequence are present, the propeptide sequence is located immediately adjacent to the N-terminus of the polypeptide, and the signal peptide sequence is located immediately adjacent to the N-terminus of the propeptide sequence. Alternatively, when both a signal peptide sequence and a propeptide sequence are present, the polypeptide may contain only a portion of the signal peptide sequence and / or only a portion of the propeptide sequence. Alternatively, the final or isolated polypeptide may comprise a mixture of a mature polypeptide and a polypeptide containing partial or full-length propeptide sequences and / or signal peptide sequences.

[0282] Regulation sequence

[0283] It is also desirable to add regulatory sequences that regulate the expression of peptides related to host cell growth. Examples of regulatory sequences are those that cause gene expression to turn on or off in response to chemical or physical stimuli, including the presence of regulatory compounds.

[0284] Regulatory sequences in prokaryotic systems include the lac, tac, and trp operon systems. In yeast, the ADH2 or GAL1 system can be used. In filamentous fungi, the *Aspergillus niger* glucosylamylase promoter, the *Aspergillus oryzae* TAKA α-amylase promoter and *Aspergillus oryzae* glucosylamylase promoter, and the *Trichoderma reesei* cellobiose hydrolase I and cellobiose hydrolase II promoters can be used. Other examples of regulatory sequences are those that allow gene amplification. In fungal systems, these regulatory sequences include dihydrofolate reductase genes amplified in the presence of methotrexate and metallothionein genes amplified with heavy metals.

[0285] transcription factors

[0286] Control sequences can also be transcription factors, which are polynucleotides encoding polynucleotide-specific DNA-binding polypeptides that control the rate of transcription of genetic information from DNA to mRNA by binding to specific polynucleotide sequences. Transcription factors can function alone and / or together with one or more other polypeptides or transcription factors in a complex by promoting or blocking the recruitment of RNA polymerase. Transcription factors are characterized by containing at least one DNA-binding domain, which is typically attached to a specific DNA sequence adjacent to the genetic element regulated by the transcription factor. Transcription factors can regulate the expression of a target protein directly (i.e., by binding to their promoter to activate the transcription of a gene encoding the target protein) or indirectly (i.e., by binding to the promoter of another transcription factor that regulates the transcription of a gene encoding the target protein) to activate the transcription of another transcription factor. Suitable transcription factors for fungal host cells are described in WO 2017 / 144177. Suitable transcription factors for prokaryotic host cells are described in Seshasayee et al., 2011, Subcellular Biochemistry 52: 7-23 and Balleza et al., 2009, FEMS Microbiol. Rev. 33(1): 133-151.

[0287] expression carrier

[0288] The method of the present invention also utilizes a recombinant expression vector containing a target polynucleotide. Various nucleotides and control sequences can be linked together to produce a recombinant expression vector, which may include one or more suitable restriction sites to allow insertion or substitution of the target polynucleotide at such sites. Alternatively, the polynucleotide can be expressed by inserting the polynucleotide or a nucleic acid construct containing the polynucleotide into a suitable vector for expression. In producing the expression vector, the coding sequence is located in the vector such that the coding sequence is operatively linked to a suitable control sequence for expression.

[0289] Recombinant expression vectors can be any vector (e.g., plasmids or viruses) that can readily undergo recombinant DNA procedures and induce polynucleotide expression. The choice of vector will typically depend on its compatibility with the host cell to which it will be introduced. Vectors can be linear or closed circular plasmids.

[0290] The vector can be a self-replicating vector, that is, a vector that exists as an extrachromosomal entity and replicates independently of chromosome replication, such as a plasmid, extrachromosomal element, mini-chromosome, or artificial chromosome. The vector can contain any means to ensure self-replication. Alternatively, the vector can be one that integrates into the genome when introduced into a host cell and replicates along with the chromosome in which it has been integrated. Furthermore, a single vector or plasmid, or two or more vectors or plasmids collectively containing the total DNA of the host cell genome to be introduced, or transposons can be used.

[0291] The vector preferably contains one or more selective markers that allow for convenient selection of cells such as transformed cells, transfected cells, and transduced cells. A selective marker is a gene whose product provides resistance to biocides or viruses, resistance to heavy metals, or prototrophic auxotrophic traits, etc.

[0292] The vector preferably contains at least one element that allows the vector to integrate into the genome of the host cell or to replicate autonomously in the cell independently of the genome.

[0293] In order to integrate into the host cell genome, the vector may depend on a polynucleotide sequence encoding a polypeptide or any other element of the vector used for integration into the genome via homologous recombination (such as homologous directed repair (HDR)) or non-homologous recombination (such as non-homologous end joining (NHEJ)).

[0294] For autonomous replication, the vector may further include an origin of replication, which enables the vector to replicate autonomously in the host cell discussed. The origin of replication can be any plasmid replicon that mediates autonomous replication and functions within the cell. The terms "origin of replication" or "plasmid replicon" refer to the polynucleotide that enables a plasmid or vector to replicate in vivo.

[0295] More than one copy of a target polynucleotide can be inserted into a host cell to increase polypeptide production. For example, two, three, four, five, or more copies can be inserted into the host cell. An increased copy number of the polynucleotide can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene along with the polynucleotide, wherein cells containing the amplified copy of the selectable marker gene and thus additional copies of the polynucleotide can be selected by culturing the cells in the presence of a suitable selectivity reagent.

[0296] host cells

[0297] In a second aspect, the present invention relates to a host cell whose genome contains a target polynucleotide sequence generated in a further step h) and / or a polynucleotide sequence identified in step e).

[0298] The present invention also relates to non-recombinant host cells, i.e., wild-type host cells. Such host cells include, but are not limited to, probiotics, for example, wherein the host cell is of the genus Bifidobacterium, such as Bifidobacterium animalis or Bifidobacterium animalis subsp. lactis.

[0299] The present invention also relates to recombinant host cells containing a target polynucleotide and / or a polynucleotide operatively linked to one or more control sequences that direct the production of a target polypeptide.

[0300] A construct or vector containing a polynucleotide is introduced into a host cell, such that the construct or vector is maintained as a chromosomal integrase or as a self-replicating extrachromosomal vector, as previously described. The choice of host cell will depend largely on the gene encoding the polypeptide and its origin. The polypeptide can be native or heterologous to the recombinant host cell. Furthermore, at least one of the one or more control sequences can be heterologous to the polynucleotide encoding the polypeptide. The recombinant host cell may contain a single copy or at least two copies, such as three, four, five or more copies, of the polynucleotide of the invention.

[0301] The host cell can be any mammalian cell that can be used to recombinantly generate the target polypeptide, such as Chinese hamster ovary cells, BHK cells, mouse cells, and HEK cells.

[0302] The host cell can be any microbial cell that can be used to recombinantly produce the target polypeptide, such as prokaryotic cells or fungal cells.

[0303] Prokaryotic host cells can be any Gram-positive or Gram-negative bacteria. Gram-positive bacteria include, but are not limited to: Bacillus spp., Bifidobacterium (e.g., BB-12®), Clostridium spp., Enterococcus spp., Bacillus spp., Lactobacillus spp., Lactococcus spp., Bacillus spp., Staphylococcus spp., Streptococcus spp., and Streptomyces spp. Gram-negative bacteria include, but are not limited to: Campylobacter spp., Escherichia coli, Flavobacterium spp., Fusobacterium spp., Helicobacter spp., Ilyobacter spp., Neisseria spp., Pseudomonas spp., Salmonella spp., and Ureaplasma spp.

[0304] The bacterial host cell can be any Bacillus genus cell, including but not limited to Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus brevis, Bacillus circulans, Bacillus croceus, Bacillus coagulans, Bacillus firmus, Bacillus lautus, Bacillus lentus, Bacillus licheniformis, Bacillus megaterium, Bacillus pumilus, Bacillus thermophilus, Bacillus subtilis, and Bacillus thuringiensis cells. In the examples, the Bacillus cells are Bacillus amyloliquefaciens, Bacillus licheniformis, and Bacillus subtilis cells.

[0305] For the purposes of this invention, Bacillus species should be defined as described in Patel and Gupta, 2020, Int. J.Syst.Evol.Microbiol. [International Journal of Systematic and Evolutionary Microbiology] 70: 406-438.

[0306] The bacterial host cell can also be any streptococcus cell, including but not limited to Streptococcus equisimilis, Streptococcus pyogenes, Streptococcus uberis, and Streptococcus equi subsp. Zooepidemicus.

[0307] The bacterial host cell can also be any Streptomyces cell, including but not limited to: Streptomyces achromogenes, Streptomyces avermitilis, Streptomyces coelicolor, Streptomyces griseus, and Streptomyces lividans cells.

[0308] Methods for introducing DNA into prokaryotic host cells are well known in the art, and any suitable method can be used, including but not limited to protoplast transformation, competent cell transformation, electroporation, conjugation, and transduction, wherein the DNA is introduced as a linearized or circular polynucleotide. Those skilled in the art will be able to readily determine, for example, a suitable method for introducing DNA into a given prokaryotic cell based on the genus. Methods for introducing DNA into prokaryotic host cells are described, for example, in Heinze et al., 2018, BMC Microbiology 18:56; Burke et al., 2001, Proc. Natl. Acad. Sci. USA 98: 6289-6294; Choi et al., 2006, J. Microbiol. Methods 64: 391-397; and Donald et al., 2013, J. Bacteriol. 195(11): 2612-2620.

[0309] The host cell can be a fungal cell. As used herein, “fungus” includes Ascomycota, Basidiomycota, Chytridiomycota, Zygomycota, Oomycota, and all mitotic fungi (as defined by Hawksworth et al. in Ainsworth and Bisby’s Dictionary of The Fungi, 8th edition, 1995, CAB International, University Press, Cambridge, UK).

[0310] Fungal cells can be transformed via processes involving protoplast-mediated transformation, Agrobacterium-mediated transformation, electroporation, gene gun methods, and shock wave-mediated transformation (as reviewed in Li et al., 2017, Microbial Cell Factories, 16: 168), and procedures described in EP 238023, Yelton et al., 1984, Proc. Natl. Acad. Sci. USA, 81: 1470-1474; Christensen et al., 1988, Bio / Technology, 6: 1419-1422; and Lubertozzi and Keasling, 2009, Biotechn. Advances, 27: 53-75. However, any method known in the art for introducing DNA into fungal host cells can be used, and the DNA can be introduced as linearized or circular polynucleotides.

[0311] The host cell for fungi can be a yeast cell. As used herein, “yeast” includes ascosporogenous yeast (Endomycetales), basidiosporogenous yeast, and yeasts belonging to the class Fungi Imperfecti (Blastomycetes). For the purposes of this invention, yeast should be defined as described in Biology and Activities of Yeast (edited by Skinner, Passmore, and Davenport, Soc. App. Bacteriol. Symposium Series No. 9, 1980).

[0312] Yeast host cells can be cells from the genera *Candida*, *Hansenula*, *Kluyveromyces*, *Pichia*, *Saccharomyces*, *Schizosaccharomyces*, or *Yarrowia*, such as *Kluyveromyces lactis*, *Saccharomyces carlsbergensis*, *Saccharomyces diastaticus*, *Saccharomyces douglasii*, *Saccharomyces kluyveri*, *Saccharomyces norbensis*, *Saccharomyces oviformis*, or *Yarrowia lipolytica*. In a preferred embodiment, the yeast host cell is a cell of the genera Pichia or Komagataella, such as Pichia pastoris cells (Komagataella phaffii).

[0313] The host cell for fungi can be a filamentous fungal cell. "Filamentous fungi" includes all filamentous forms within the phylum Eumycota and the subphylum Oomycetes (as defined by Hawksworth et al., 1995, ibid.). Filamentous fungi are generally characterized by a hyphal wall composed of chitin, cellulose, glucan, chitosan, mannan, and other complex polysaccharides. Vegetative growth occurs through hyphal elongation, and carbon metabolism is obligate aerobic. In contrast, the vegetative growth of yeasts (such as Saccharomyces cerevisiae) occurs through budding of single-celled cells, and carbon metabolism can be fermentative.

[0314] The host cells of filamentous fungi can be *Acremonium*, *Aspergillus*, *Aureobasidium*, *Bjerkandera*, *Ceriporiopsis*, *Chrysosporium*, *Coprinus*, *Coriolus*, *Cryptococcus*, *Filibasidium*, *Fusarium*, *Humicola*, *Magnaporthe*, *Mucor*, *Myceliophthora*, and *Neurospora*. Cells of the genera *Neocallimastix*, *Neurospora*, *Paecilomyces*, *Penicillium*, *Phanerochaete*, *Phlebia*, *Piromyces*, *Pleurotus*, *Schizophyllum*, *Talaromyces*, *Thermoascus*, *Thielavia*, *Tolypocladium*, *Trametes*, or *Trichoderma* are used. In a preferred embodiment, the filamentous fungal host cells are *Aspergillus*, *Trichoderma*, or *Fusarium* cells. In another preferred embodiment, the filamentous fungal host cells are *Aspergillus niger*, *Aspergillus oryzae*, *Trichoderma reesei*, or *Fusarium* cells.

[0315] For example, the host cells of filamentous fungi can be *Aspergillus bubomori*, *Aspergillus sulphureus*, *Aspergillus fumigatus*, *Aspergillus japonicus*, *Aspergillus nidus*, *Aspergillus oryzae*, *Aspergillus nidus*, *Aspergillus nidus*, *Aspergillus carnegieii*, *Aspergillus chrysogenus*, *Aspergillus pannohita*, *Aspergillus circumferentialis*, *Aspergillus microerythropus*, *Aspergillus wormi*, *Aureobasidium stenoptera*, *Aureobasidium leucis*, *Aureobasidium lucime*, *Aureobasidium stenoptera*, *Aureobasidium querceae*, *Aureobasidium tropicum*, *Aureobasidium brevicornum*, *Coprinus comatus*, *Aureobasidium triflorum*, *Fusarium solani*, *Fusarium solani*, and *Fusarium kuwai*. Fusarium spp., Fusarium graminearum, Fusarium graminearum, Fusarium heterosporum, Fusarium spp. ...

[0316] On the one hand, the host cell is isolated.

[0317] On the other hand, the host cell is purified.

[0318] Generation method

[0319] In a third aspect, the invention also relates to methods for producing a desired polypeptide, the methods comprising (a) culturing cells according to the second aspect under conditions favorable for polypeptide production; and optionally (b) recovering the polypeptide. In one aspect, the cells are Bacillus cells. In another aspect, the cells are Bacillus licheniformis cells. In another aspect, the cells are Aspergillus cells. In one aspect, the cells are Aspergillus niger cells. In another aspect, the cells are Aspergillus oryzae cells. In another aspect, the cells are Trichoderma reesei cells.

[0320] The present invention also relates to methods for producing host cell culture media, the methods comprising (a) culturing host cells according to the second aspect under conditions conducive to the production of host cells; and optionally, (b) recovering the host cells.

[0321] On one hand, the recovered cells are Bifidobacteria, such as Bifidobacterium animalis or Bifidobacterium animalis subsp. lactis.

[0322] The host cells are cultured in a nutrient medium suitable for producing peptides using methods known in the art. For example, cells can be cultured in a suitable medium and under conditions that allow for peptide expression and / or isolation by shake-flask culture or by small-scale or large-scale fermentation (including continuous, batch, fed-batch, or solid-state and / or microcarrier-based fermentation) in a laboratory or industrial fermenter. Suitable media are available from commercial suppliers or can be prepared according to publicly available compositions (e.g., in the catalogue of the U.S. Center for Type Culture Collection). If the peptide is secreted into the nutrient medium, it can be recovered directly from that medium. If the peptide is not secreted, it can be recovered from cell lysates.

[0323] Peptides can be detected using methods known in the art that are specific to peptides, including but not limited to the use of specific antibodies, enzyme product formation, enzyme substrate disappearance, or determination of the relative or specific activity of the peptide.

[0324] Peptides can be recovered from culture media using methods known in the art, including but not limited to collection, centrifugation, filtration, extraction, spray drying, evaporation, or precipitation. On one hand, whole fermentation broth containing peptides is recovered. On the other hand, cell-free fermentation broth containing peptides is recovered.

[0325] Peptides can be purified using a variety of procedures known in the art to obtain substantially pure peptides and / or peptide fragments (see, for example, Wingfield, 2015, Current Protocols in Protein Science; 80(1): 6.1.1-6.1.35; Labrou, 2014, Protein Downstream Processing, 1129: 3-10).

[0326] In terms of alternatives, peptides are not recycled.

[0327] Computational model

[0328] In one aspect, the present invention relates to methods according to the first aspect, which further include step g) training a computational model, such as a machine learning algorithm, with sequence data obtained from step e) and / or score data obtained from step f).

[0329] In one embodiment, the computational model for step g) is selected from the list of the following: linear regression, decision tree, random forest model, support vector machine (SVM), neural network, K-means clustering, Naive Bayes, Gaussian mixture model (GMM), or generative model.

[0330] In one embodiment, the computational model is performed in an electronic device to provide candidate biological sequences, the method comprising:

[0331] - Obtain input data indicating the input biological sequence;

[0332] - Determine the candidate biological sequence by applying the model to the input data; and

[0333] - Provides biological sequence data indicating the candidate biological sequence.

[0334] In one embodiment, the model is a generative model.

[0335] In one embodiment, the generative model is non-unidirectional.

[0336] In one embodiment, the input biological sequence comprises one or more target polynucleotides identified in step e).

[0337] In one embodiment, the input biological sequence is one or more of the following: an amino acid sequence of a target polypeptide, a nucleic acid sequence encoding the target polypeptide, a control sequence (e.g., an expression control sequence), and a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence).

[0338] In one embodiment, the candidate biological sequence is one or more of the following: a control sequence (e.g., an expression control sequence), a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence), an amino acid sequence of the target polypeptide, and a nucleic acid sequence encoding the target polypeptide.

[0339] In one embodiment, the input biological sequence is the amino acid sequence of the target polypeptide and / or the nucleic acid sequence encoding the target polypeptide, and the candidate biological sequence is a nucleic acid sequence that enhances compatibility with host cells.

[0340] In one embodiment, the input biological sequence is the amino acid sequence of the target polypeptide and / or the nucleic acid sequence encoding the target polypeptide, and the candidate biological sequence is a control sequence (e.g., an expression control sequence) and / or a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence).

[0341] In one embodiment, the input biological sequence is a control sequence (e.g., an expression control sequence) and / or a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence), and wherein the candidate biological sequence is a nucleic acid sequence encoding a target polypeptide.

[0342] In one embodiment, the model is a generative model.

[0343] In one embodiment, the generative model is non-unidirectional.

[0344] In one embodiment, the generative model is one or more of the following: generative adversarial network model, Wasserstein generative adversarial network model, diffusion model, and variational autoencoder.

[0345] In one embodiment, applying a non-unidirectional model to input data includes dividing the non-unidirectional model into multiple generators, wherein each of the multiple generators is configured to determine one or more candidate biological sequences based on the input data, for a subset of nucleotides and / or a subset of amino acids, and predetermined criteria.

[0346] In one embodiment, determining candidate biological sequences by applying a model to input data includes:

[0347] - Use the generator to predict the compatibility of the candidate biological sequence with the host cell; and

[0348] - Identify candidate biological sequences that have predicted compatibility that meet the predetermined criteria.

[0349] In one embodiment, the predetermined standard is based on one or more of the following:

[0350] - The proportion of nucleotide sets in the candidate biological sequence;

[0351] -The type of host cell;

[0352] -The host cell genus or species;

[0353] - The GC content of the host cell genome;

[0354] - The GC content of the candidate biological sequence; and

[0355] - Parameters related to the characteristics of the candidate biological sequence.

[0356] In one embodiment, the method includes training a model based on a training set of biological sequences, wherein the training set of biological sequences includes training data indicating one or more biological sequences associated with a host cell.

[0357] In one embodiment, the training set of biological sequences is heterologous to the genus of the host cell, preferably heterologous to one or more species of the host cell.

[0358] In one embodiment, the training data includes training input data indicating one or more of the following: an amino acid sequence of a target peptide, a nucleic acid sequence encoding the target peptide, a control sequence (e.g., an expression control sequence), and a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence).

[0359] In one embodiment, the training data includes training output data indicating one or more of the following: a control sequence (e.g., an expression control sequence), a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence), an amino acid sequence of a target peptide, and a nucleic acid sequence encoding a target peptide.

[0360] In one embodiment, training the model includes using a discriminator to predict a score indicating that a training candidate biological sequence is a reference biological sequence, taking a training set of biological sequences and training candidate biological sequences as input.

[0361] In one embodiment, the method includes obtaining experimental data related to candidate biological sequences and host cells from a test environment data repository; wherein the experimental data indicates the yield performance of the candidate biological sequence associated with the host cell.

[0362] In one embodiment, the method includes validating candidate biological sequences based on experimental data.

[0363] In one embodiment, the method includes selecting one or more generators based on experimental data.

[0364] In one embodiment, the method includes adjusting the model based on experimental data.

[0365] In one embodiment, obtaining input data indicating an input biological sequence includes obtaining input data for the input biological sequence from a database and / or memory of an electronic device.

[0366] The present invention also relates to an electronic device comprising memory circuitry, processor circuitry, and an interface, wherein the electronic device is configured to perform any method according to the present invention.

[0367] The present invention also relates to a computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by an electronic device, cause the electronic device to perform any of the methods of the present invention.

[0368] In one embodiment, the method includes an additional step h) generating one or more synthetic target polynucleotides based on the output of a computational model.

[0369] In one embodiment, one or more synthetic polynucleotides generated in step h) contain or consist of candidate biological sequences.

[0370] In one embodiment, one or more synthetic target polynucleotides generated in step h) are codon-optimized.

[0371] In one embodiment, one or more synthetic target polynucleotides generated in step h) encode a polypeptide having increased substrate binding, increased receptor binding, increased substrate specificity, increased specific activity, and / or increased stability.

[0372] In one embodiment, one or more synthetic target polynucleotides generated in step h) lead to an increase in the expression of the target polypeptide.

[0373] In one embodiment, one or more synthetic target polynucleotides generated in step h) contain a control sequence.

[0374] Due to the combinatorial nature of the problem, the chances of finding a satisfactory "partner" biological sequence using conventional screening strategies are limited. For example, finding the optimal 10-mer peptide (e.g., a peptide consisting of 10 amino acids) to use in combination with some target biological sequences would require screening 20^10 different biological sequences (if only 20 different amino acids from the standard genetic code are considered). This illustrates that clearly finding a satisfactory sequence through screening is a tedious task and may be infeasible within a reasonable timeframe, especially when biological sequences most frequently include more than 10 amino acids.

[0375] This disclosure allows for the extraction of learning from natural biological sequence pairs, for example, when the biological sequences exist in nature. For instance, biological sequence pairs are available in databases such as public and / or private databases (e.g., the National Center for Biotechnology Information (NCBI) database and / or nucleotide archives (e.g., EMBL)). It is conceivable that the extracted learning can be transferred to experimental settings. This disclosure allows for the learning of compatibility rules from natural sequence pairs provided by databases and provides the learned compatibility rules. It is conceivable that the compatibility rules can be further adapted to experimental settings.

[0376] This disclosure allows for some interaction between machine learning methods and experimental methods. Machine learning-based biological sequence analysis and experimental screening methods contribute two distinct layers of learning. For example, the first layer of learning allows for the extraction of complex biological rules that need to be followed, while the second layer of learning accumulates data specific to the experimental setting. In the disclosed techniques, for example, the extracted learning is used to provide (using deep learning methods, such generative models) a relevant subset of candidate biological sequences, which can now be used for feasible screening using experimental methods.

[0377] By applying models (e.g., generative models), the disclosed techniques unlock the potential of experimental screening methods by significantly reducing the complexity of the process of finding satisfactory "mate" biological sequences. In other words, the actual quality of candidate biological sequences is validated experimentally. In some instances, learning can be used as feedback to update the model (thus "informing" the model about the quality of the recommendations).

[0378] This disclosure provides a method for providing candidate biological sequences performed by an electronic device. In other words, the method can be a computer-implemented method.

[0379] This method includes, for example, obtaining input data indicating an input biological sequence from a biological library with different sequences. The input data may be associated with and / or represent the input biological sequence. The input data includes data representing the input biological sequence, such as data representing one or more properties of the input biological sequence. In some instances, one or more properties of the input data include one or more of the following: amino acid sequence, nucleic acid sequence, three-dimensional structure of the input biological polypeptide sequence (e.g., obtained via Alpha-Fold2), folding of the input biological sequence, and nucleic acid pairing.

[0380] This method involves applying a model (e.g., a generative model) to input data to determine candidate biological sequences. For example, candidate biological sequences may be determined for compatibility with host cells (e.g., target compatibility with a given host cell) and / or for increasing compatibility with a given host cell. In other words, for example, the model applied to the input data aims to increase one or more expression steps of a target peptide in the host cell, such as increasing or modifying one or more of the following: transcription, post-transcriptional modification, translation, post-translational modification, folding, secretion, phenotypic traits, and yield of the target peptide in the host cell. Yield can be intracellular and / or extracellular. In other words, yield can be considered as the target performance parameter to be optimized when determining candidate biological sequences. It can be noted that yield can be optimized individually or in combination via multiple steps, such as modified secretion, modified transcription, and modified translation.

[0381] In some instances, the model generates candidate biological sequences based on the input data. In some instances, candidate biological sequences may be determined based on one or more of the following: the score generated in step f), host cell data, input data, and information indicating the type of biological sequence to be identified as a candidate biological sequence.

[0382] The invention is further described by the following examples, which should not be construed as limiting the scope of the invention. Example

[0383] Example 1: Multichannel sorting and subsequent DNA sequencing

[0384] The strain library contained 102 different polynucleotides encoding signal peptides, which were ordered as synthetic DNA and fused upstream of the DNA sequence encoding a protease (encoding a serine endopeptidase). The library was transformed into Bacillus licheniformis strain MOL3320, as described in patent US 2019 / 0185847 A1. Selection was performed based on ERM. The resulting strain expressed proteases with different signal peptide variants.

[0385] The resulting strain was then subjected to sectional fermentation on a microfluidic droplet production chip in 50 pL droplets of nutrient-controlled medium in fluorinated oil (HFE 7500). The droplets were stabilized with 2 wt% fluorinated surfactant (008-fluorinated surfactant, RanBiotech). The resulting emulsion was incubated in collection vials at 37°C for 4 days.

[0386] The host cell-secreted serine endopeptidase hydrolyzes a proprietary fluorescent rhodamine substrate. Commercially available fluorescent rhodamine substrates include rhodamine 110-bis-(succinyl-L-alanyl-L-alanyl-L-prolyl-L-phenylalanylamide) (CPC Scientific Inc., San Jose, CA). The substrate was added to each droplet on the microfluidic chip, and after incubation for 4 minutes, the fluorescence response was measured, as shown in the figure. Figure 2 As shown, the release of Rhodamine 110 resulted in an increase in fluorescence at 520 nm. This increase was proportional to the enzyme activity measured relative to the standard. Therefore, the measured level of the fluorescence signal was directly related to the concentration of the protease in each droplet. Figure 1 The apparatus shown sorts each droplet into one of five output channels based on its measured fluorescence level. Here, we use two electrodes on either side of the input channel to guide the droplets to one of the five output channels via dielectrophoretic droplet sorting. After connecting the five output channels to five collection tubes and collecting at least 1000 droplets in each tube, we detach the collection tubes from the microfluidic device. The collected cell pools are named Pool 1, Pool 2, Pool 3, Pool 4, and Pool 5 (see [link to documentation]). Figure 1 The signal peptide sequence upstream of the DNA sequence encoding the protease contained in each pool is amplified by PCR and then sent for DNA sequencing.

[0387] like Figure 2 As shown, pool 1 contains empty droplets (peak value around 2500 RFU) and droplets with no activity or very weak activity. Also as... Figure 2 As shown, sorting the library into five pools allows for the study of each signal peptide variant based on the protease activity of the relevant droplets. In other words, analyzing each library member provides an analysis of the complete library without losing data about one or more library members, because after sorting, the coding sequence of each signal peptide will be sequenced in a subsequent step (see Example 2).

[0388] Example 2: Signal peptide library analysis and MTP culture

[0389] To score individual signal peptide sequences in the form of MTP, the strain was fermented for approximately 120 hours, and protease activity was measured at the end of fermentation.

[0390] For scoring individual signal peptide sequences using the droplet sorting method outlined herein, the abundance of each signal peptide sequence in each of the five pools was analyzed. The abundance of 102 signal peptide sequences in each pool can be analyzed. Figure 3 The data shows that (black = high abundance; white = low abundance).

[0391] exist Figure 3In the MTP, the signal peptide sequences are ordered from “1” to “102” according to their protease activities measured in MTP form (“0” = highest activity, “102” = lowest activity). Five droplet sorting cells are ordered according to the fluorescence threshold used for separation (cell 5 contains droplets measured with the highest fluorescence signal, and cell 1 contains droplets measured with the lowest fluorescence signal).

[0392] As from Figure 3 As can be seen, in pool 1 (lowest fluorescence signal), signal peptides were primarily identified, which exhibited very poor protease activity in the MTP form. In pool 3 (medium fluorescence signal), signal peptides exhibiting moderate protease activity in the MTP form were primarily identified. In pool 5 (highest fluorescence signal), signal peptides leading to high protease activity during the MTP period were primarily identified.

[0393] Furthermore, such as from Figure 3 It can be seen that the proportion of SP sequences with high protease activity increases from pool 1 to pool 2, from pool 2 to pool 3, and from pool 3 to pool 4, and has the highest proportion in pool 5.

[0394] SP sequences with intermediate protease activity were found in pool 2 (medium-low protease activity) and pool 4 (medium-high protease activity). As shown in this example, sorting into up to two pools (e.g., up to five pools) improves output resolution and allows identification not only of the best or worst performers but also sequences in between. This approach is particularly advantageous for signal peptides, for example, when the aim is to fine-tune the expression of a target polypeptide.

[0395] In summary, the screening method allows for the efficient screening of the entire library while dividing library members into five pools, thereby enabling detailed analysis and understanding of each library member.

[0396] Example 3: Validation relative to the MTP screening method

[0397] This example validates the results of multichannel sorting of SP libraries compared to results obtained from culturing the same SP library in MTP form (Examples 1 and 2).

[0398] A score for a given sequence is calculated based on the abundance of that sequence in each sorting pool obtained through multi-channel sorting. First, for a given sequence, the score of the corresponding read in the pool is determined by dividing the number of reads in the given sequence by the total number of reads obtained when sequencing the entire pool. This is performed for each pool generated in the experiment. Next, the relative proportion of the given sequence in each pool is calculated across all pools. Finally, the score of the given sequence is calculated by summing the products of the relative proportion in each pool and the corresponding selection threshold. Similarly, a score is obtained for each signal peptide sequence cultured in MTP form based on the protease activity exhibited by each sequence.

[0399] Scores derived from multichannel droplet sorting experiments show a high level of correlation (R0) with scores derived from measurements in the form of MTP. 2 = 0.85, Figure 4 ). Figure 4 The curve shown is plotted as the score from the microtiter plate (MTP form, y-axis) versus the score from the signal peptide sequence from the multichannel sorting (x-axis).

[0400] Due to the high correlation between the scores of the two methods, we conclude that the multichannel droplet sorting of this invention represents an improved and significantly cheaper screening method, saving time and sample volume, while providing high-resolution output when screening large biological libraries (see [link to original text]). Figure 3 ).

[0401] Example 4: The microdroplet method reduces the standard deviation of the assignment score.

[0402] This example investigated the variability of measurements taken from clonal cells cultured in the MTP or microdroplet form of this invention. The cells used in this experiment were Bacillus licheniformis cells with six copies of the secreted protease (target protein, POI).

[0403] We compared the variability of POI concentration measurements in 1 mL (96-well MTP) and 50 pL microdroplet formats. We performed parallel fermentations of single POI-producing strains and measured the standard deviation of the amount of POI released.

[0404] The results are shown in Figure 5 middle, Figure 5 A shows MTP fermentation and Figure 5 Figure B shows droplet fermentation (y-axis: count; x-axis: relative POI yield). The standard deviation (STD) for 96-well MTP is 7.3% (see Figure B). Figure 5 A), while the standard deviation for droplet form is 4.6% (see A). Figure 5(B) This result demonstrates that the droplet form exhibits significantly lower variability compared to the traditional MTP form. The reduction in STD directly impacts score calculation and leads to more accurate scoring of library members. When scoring is applied to the computational model, the more accurate scoring ultimately results in improved input data quality, allowing for the construction of high-confidence machine learning models. Therefore, models trained on data obtained from the microdroplet method of this invention exhibit higher confidence compared to models trained on MTP data.

[0405] Example 5: Processing Filtering Results Using a Computational Model

[0406] This example validates the results of multichannel microdroplet sorting of SP libraries compared to results obtained from culturing the same SP library in MTP form (Examples 1 and 2).

[0407] We specifically focus on a subset of signal peptides found in both the MTP method and the microdroplet method of this invention. Using this subset, the findings of the two methods can be compared with each other. For this subset, we rank the signal peptides according to yield, specifically at the amino acid level, i.e., excluding the contribution of the codon level to the yield. This is achieved by considering the codon distribution of each individual signal peptide and defining it as the maximum achievable yield relative to the codons, then interpreting it as the yield potential derived solely from the amino acid sequence. Here, for each given signal peptide, we apply a robust estimate of the maximum achievable yield by considering the 75th percentile of yield measurements for codon variants. We refer to this value as the "peptide-level yield".

[0408] In the following sections, we evaluate whether the ranking of “peptide level yield” in MTP and microdroplets is similar in terms of the underlying features that generate the ranking, and whether machine learning models trained with MTP or microdroplet data will lead to similar outputs.

[0409] like Figure 6 As shown, we calculate the "yield of peptide level" from data from MTP, and based on the calculated "yield_peptide_level" ( Figure 6 B; the y-axis indicates relative peptide yield), by cutting the signal peptide in half ( Figure 6 The dashed line in section B divides the signal peptides into two groups ("good" and "bad"). The "good" signal peptides are shown in... Figure 6 In the right half of B, and the "difference" signal peptide is shown in Figure 6 The left half of B. Then, we train a computational model with these "good" and "bad" sequences (subsets) separately, and ask whether there are any amino acids that are over- or under-presented relative to the sequences in other subsets within each subset. This task uses a random forest (RF) model.

[0410] The model answers the question by indicating that the amino acid proline is relatively more common in "good" sequences compared to its presence in "bad" sequences. Figure 6 A on the y-axis indicates the fraction of sequences containing proline. For example... Figure 6 As shown, there are significant differences in the fraction of signal peptides containing proline, depending on the yield of the signal peptide; approximately two-thirds of the good sequences contain at least one proline, while only about one-third of the poor sequences contain at least one proline. Therefore, the model concludes that the presence of proline in the signal peptide is a strong indicator of good expression of the studied POI.

[0411] Repeat the microdroplet method of the present invention. Figure 6 The same experiment as the MTP shown is presented. The results are shown in... Figure 7 In, among them Figure 7 In A, the y-axis indicates the fraction of proline-containing sequences as output of the computational model, and Figure 7 In B, the y-axis indicates the relative peptide yield of different SP sequences. Comparison Figure 6 The result of A and Figure 7 Based on the results of A, we conclude that the results obtained using the microdroplet method of this invention are comparable to those obtained by training a computational model using the results of the conventional MTP method.

[0412] from Figure 6 and Figure 7 From this, we can conclude that although the microdroplet screening of the present invention has a lower operating cost than MTP screening, it can provide high-quality data in a shorter time compared to the MTP method. These results further confirm the unexpected efficiency of the method of the present invention.

[0413] Example 6: The increased training data size due to microdroplets improved the performance of machine learning models.

[0414] For the task of predicting the yield of peptide levels in microdroplets (as explained above), we addressed the problem of demonstrating the importance of training data size for the performance of machine learning models.

[0415] In the following sections, we calculate the microdroplet screening results from this invention. all The “peptide level yield” of the signal peptide, rather than simply using the signal peptide that is also found in MTP, as in the previous example.

[0416] Here, 30% of the microdroplet data was reserved as the test set, and the remaining 70% was used to train the random forest regression model. Within the remaining 70% of the training data, we varied the proportion of training data allowed for the model. Each time, we trained the model on the available data and evaluated the performance of the machine learning model against the initially allocated test set by calculating the correlation coefficient between the predicted and actual observations.

[0417] The results are shown in Figure 8 In the middle. We repeated the entire process three times (to minimize the contribution from the random component of the method) and plotted all correlation coefficients ( Figure 8 The y-axis is used as a function of the actual amount of training data that the model is allowed to use in each round.

[0418] like Figure 8 As shown in the diagram, we can clearly see that the model's performance improves as we add more data for training. This indicates that data size (a property of the microdroplet method of this invention) is highly valuable for downstream machine learning methods.

[0419] Example 7: Microdroplet methods allow for the identification of superior library members.

[0420] This example compares the results of multichannel sorting of SP libraries from Examples 1 and 2 with results obtained from screening the same SP libraries in MTP form. Compared to the previous example using 5 pools, this example sorts the libraries into 7 pools. Each library member is given a unique signal peptide identifier. Each library member consists of a different DNA sequence encoding the signal peptide. Using the microdroplet method of this invention, the SP libraries are sorted into 7 pools, i.e., pools 1 through 7. Droplets with the lowest signal are sorted into pool 1, while droplets with the highest signal are sorted into pool 7. Pools 2-6 contain cells containing library members that exhibit signals below a threshold in pool 7 and above a threshold in pool 1. In other words, the signal threshold increases from pool 1 to pool 7.

[0421] Table 1 shows the results of 304 library members identified using the microdroplet method. Sequences identified in both MTP screening and microdroplet screening are marked with gray shading (e.g., SP_GAN_208). For each given sequence, Table 1 shows the number of droplets containing that sequence in each pool. For example, the sequence SP_GAN_205 appeared in 61 droplets in pool 7 and 2 droplets in pool 1. Additionally, for sequences identified using the microdroplet method, the table shows the scores calculated as described in Example 3. For sequences also identified in MTP, Table 1 shows the relative activity identified during MTP incubation. Importantly, the sequences in Table 1 are arranged in descending order of droplet score, i.e., the highest droplet score is at the top of Table 1, and the lowest droplet score is at the bottom of Table 1.

[0422] As can be seen in Table 1, the method of this invention allows for the identification of library members that would otherwise be ignored and / or undetectable using conventional methods (such as MTP). These library members are shown in Table 1 as lines with a transparent background. Lines with a gray background indicate library members that have already been identified using MTP.

[0423] As shown in Table 1, in terms of score, sequence SP_GAN_208 (marked in gray) was the best-performing library member identified in the MTP screening. During microdroplet screening, the same library members were also identified as well-performing sequences. However, the microdroplet method of this invention identified six additional library members that showed higher scores compared to SP_GAN_208; these library members were identified as SP_GAN_205, SP_GAN_206, SP_GAN_217, SP_GAN_126, SP_GAN_47, and SP_GAN_232. Therefore, the method of this invention is highly beneficial for further improving biotechnological challenges, such as by increasing the expression of POIs with novel signal peptide sequences.

[0424] Table 1. Scores obtained from SP DNA sequences after sorting into 7 pools.

[0425]

[0426] The invention described and claimed herein is not limited to the specific aspects disclosed herein, as these aspects are intended to illustrate several aspects of the invention. Any equivalent aspects are intended to be within the scope of the invention. In fact, various modifications to the invention, in addition to those shown and described herein, will become apparent to those skilled in the art from the foregoing description. Such modifications are also intended to fall within the scope of the appended claims. In case of conflict, the disclosure including the definition shall prevail.

[0427] The invention is further defined by the following numbered paragraphs:

[0428] 1. A method for screening biological libraries, the method comprising the following steps:

[0429] a) Provide a microfluidic device including a droplet sorter (200) having at least three output channels (301, 302, 303).

[0430] b) Provides a droplet emulsion containing a target polynucleotide library and screenable products.

[0431] c) Determine the amount of screenable products from one or more droplets in the microfluidic device.

[0432] d) The droplet sorter (200) sorts the one or more droplets into the receiving output channels of the at least three output channels (301, 302, 303), wherein the receiving output channel is determined based on the amount of screenable product per droplet, and wherein the at least three receiving output channels receive a plurality of droplets, the plurality of droplets comprising an amount of screenable product above and / or below one or more predetermined threshold levels.

[0433] e) Identify one or more target polynucleotides present in the at least three output channels (301, 302, 303), and obtain the sequence data of the one or more target polynucleotides.

[0434] f) For one or more output channels (301, 302, 303), assign a score to each of the one or more target polynucleotides, wherein the score is calculated based on the abundance of each of the one or more target polynucleotides in one of the one or more output channels (301, 302, 302).

[0435] 2. The method as described in paragraph 1, wherein the droplet emulsion contains one or more host cells.

[0436] 3. The method as described in paragraph 2, wherein each host cell contains one or more target polynucleotides of the target polynucleotide library.

[0437] 4. The method as described in any of the preceding paragraphs, wherein in step b), each droplet contains at most one host cell or multiple host cells derived from the same parent host cell.

[0438] 5. The method as described in any of the preceding paragraphs, wherein in step b), each droplet contains at most one target polynucleotide.

[0439] 6. The method as described in any of the preceding paragraphs, wherein the screenable product is produced by these host cells.

[0440] 6a. The method as described in any of the preceding paragraphs, wherein the screenable product is catalyzed by an enzyme, preferably the enzyme being encoded by the target polynucleotide.

[0441] 7. The method as described in any of the preceding paragraphs, wherein the screenable product is encoded by one or more target polynucleotides.

[0442] 8. The method as described in any of the preceding paragraphs, wherein the screenable product is generated by a polypeptide expressed by these host cells.

[0443] 9. The method as described in any of the preceding paragraphs, wherein the screenable product is generated by a polypeptide encoded by the one or more target polynucleotides.

[0444] 10. The method as described in any of the preceding paragraphs, wherein the screenable product is a polypeptide expressed by these host cells.

[0445] 11. The method as described in any of the preceding paragraphs, wherein the screenable product is an enzyme.

[0446] 12. The method as described in any of the preceding paragraphs, wherein the enzyme is expressed by these host cells.

[0447] 13. The method as described in any of the preceding paragraphs, wherein the enzyme is selected from the list of: hydrolases, isomerases, ligases, lyases, oxidoreductases, or transferases, such as aminopeptidase, amylase, glycosylase, carboxypeptidase, catalase, cellobiase, cellulase, chitinase, keratinase, cyclodextrin glycosyltransferase, deoxyribonuclease, endoglucanase, esterase, α-galactosidase, β-galactosidase, glucosylamylase, α-glucosidase, β-glucosidase, invertase, laccase, lipase, mannosidase, polysaccharidase, oxidase, pectinase, peroxidase, phytase, polyphenol oxidase, proteolytic enzyme, ribonuclease, transglutaminase, xylanase, or β-xylosidase.

[0448] 14. The method as described in any of the preceding paragraphs, wherein the screenable product is degraded by these host cells.

[0449] 15. The method as described in any of the preceding paragraphs, wherein the screenable product is degraded by a polypeptide encoded by the one or more target polynucleotides.

[0450] 16. The method as described in any of the preceding paragraphs, wherein the screenable product is degraded by a polypeptide expressed by these host cells.

[0451] 17. The method as described in any of the preceding paragraphs, wherein the screenable product is an enzyme substrate, preferably the enzyme is selected from the list of: hydrolases, isomerases, ligases, lyases, oxidoreductases, or transferases, such as aminopeptidase, amylase, glycosylase, carboxypeptidase, catalase, cellobiase, cellulase, chitinase, keratinase, cyclodextrin glycosyltransferase, deoxyribonuclease, endoglucanase, esterase, α-galactosidase, β-galactosidase, glucosylamylase, α-glucosidase, β-glucosidase, invertase, laccase, lipase, mannosidase, polysaccharidase, oxidase, pectinase, peroxidase, phytase, polyphenol oxidase, proteolytic enzyme, ribonuclease, transglutaminase, xylanase, or β-xylosidase.

[0452] 18. The method of any of the preceding paragraphs, wherein the screenable product is a fluorescent product.

[0453] 19. The method of any of the preceding paragraphs, wherein the fluorescent product is derived from a fluorescent substrate by an enzyme encoded by the target polynucleotide.

[0454] 20. The method as described in any of the preceding paragraphs, wherein the amount of the screenable product is inversely proportional to one or more of cell number, cell growth, cell division, cell viability, or cell growth rate.

[0455] 21. The method as described in any of the preceding paragraphs, wherein the amount of the screenable product is proportional to one or more of cell number, cell growth, cell division, cell viability, or cell growth rate.

[0456] 22. The method as described in any of the preceding paragraphs, wherein the screenable product comprises or is composed of one or more host cells.

[0457] 23. The method as described in any of the preceding paragraphs, wherein the screenable product comprises or consists of substantially all host cells in the droplet.

[0458] 24. The method as described in any of the preceding paragraphs, wherein the screenable product is the product of an enzymatic reaction, preferably the product of a reaction catalyzed by an enzyme selected from the list of enzymes: hydrolases, isomerases, ligases, lyases, oxidoreductases, or transferases, such as aminopeptidase, amylase, glycosylase, carboxypeptidase, catalase, cellobiase, cellulase, chitinase, keratinase, cyclodextrin glycosyltransferase, deoxyribonuclease, endoglucanase, esterase, α-galactosidase, β-galactosidase, glucosylamylase, α-glucosidase, β-glucosidase, invertase, laccase, lipase, mannosidase, polysaccharidase, oxidase, pectinase, peroxidase, phytase, polyphenol oxidase, proteolytic enzyme, ribonuclease, transglutaminase, xylanase, or β-xylosidase.

[0459] 25. The method as described in any of the preceding paragraphs, wherein the score is proportional to the number of identical DNA sequences of the first target polynucleotide present in the output channel, for example, the score is normalized relative to the number.

[0460] 26. The method as described in any of the preceding paragraphs, wherein the score is the total number of identical DNA sequences of the first target polynucleotide present in the output channel.

[0461] 27. The method as described in any of the preceding paragraphs, wherein the score is proportional to the number of identical DNA sequences of the second target polynucleotide present in the output channel, for example, the score is normalized relative to the number.

[0462] 28. The method as described in any of the preceding paragraphs, wherein the score is the total number of identical DNA sequences of the second target polynucleotide present in the output channel.

[0463] 29. The method as described in any of the preceding paragraphs, wherein the microfluidic device includes an incubation zone (500).

[0464] 30. The method as described in any of the preceding paragraphs, wherein the incubation area (500) is located upstream of the droplet sorter (200) and / or upstream of one or more sorting devices (401, 402).

[0465] 31. The method as described in any of the preceding paragraphs, wherein the method comprises incubating the droplet emulsion under conditions that allow cell growth and / or allow DNA to be transcribed from DNA into RNA and / or allow RNA to be translated into polypeptides, preferably, the incubation is performed prior to step c).

[0466] 32. The method as described in any of the preceding paragraphs, wherein the incubation is not performed on the microfluidic chip.

[0467] 33. The method as described in any of the preceding paragraphs, wherein the incubation is performed on and / or within the microfluidic device.

[0468] 34. The method as described in any of the preceding paragraphs, wherein after incubation, the cells contained in a droplet are genetically identical, i.e., these cells are derived from a parental host cell, preferably the same parental host cell.

[0469] 35. The method as described in any of the preceding paragraphs, wherein the droplet sorter includes one or more sensor elements (600), which are preferably located downstream of the incubation area (500) and / or upstream of the sorting device (401, 402).

[0470] 36. The method as described in any of the preceding paragraphs, wherein the one or more sensor elements (600) include a fluorescence sensor.

[0471] 37. The method as described in any of the preceding paragraphs, wherein the one or more sensor elements (600) include a light-absorbing sensor.

[0472] 38. The method as described in any of the preceding paragraphs, wherein the one or more sensor devices (600) include an image sensor, such as a CMOS sensor, a CCD sensor, or a PMT sensor.

[0473] 39. The method as described in any of the preceding paragraphs, wherein the one or more sensor elements (600) include NEMS (nanoelectromechanical systems) sensors.

[0474] 40. The method as described in any of the preceding paragraphs, wherein the one or more sensor elements (600) include a mass analyzer suitable for mass spectrometry, such as a quadrupole mass analyzer, a TOF mass analyzer, an ion trap mass analyzer, an orbit trap mass analyzer, a sector magnetic field mass analyzer, a Q-TOF mass analyzer, or an FT-ICR mass analyzer.

[0475] 41. The method as described in any of the preceding paragraphs, wherein step e) comprises DNA amplification of the one or more target polynucleotides in each output channel.

[0476] 42. The method as described in any of the preceding paragraphs, wherein the DNA amplification is a PCR method.

[0477] 43. The method as described in any of the preceding paragraphs, wherein the DNA amplification is a ddPCR (droplet digital PCR) method.

[0478] 44. The method as described in any of the preceding paragraphs, wherein step e) comprises DNA sequencing of the one or more target polynucleotides, for example, sequencing after PCR amplification, or sequencing via nanopore sequencing.

[0479] 45. The method as described in any of the preceding paragraphs, wherein during step e), the one or more target polynucleotides are identified by DNA barcoding.

[0480] 46. ​​The method as described in any of the preceding paragraphs, the method comprising step g) training a computational model, for example, a machine learning algorithm, with the sequence data obtained from step e) and / or the score data obtained from step f).

[0481] 47. The method as described in paragraph 46, wherein the computational model of step g) is selected from the list of the following: linear regression, decision tree, random forest model, support vector machine (SVM), neural network, K-means clustering, Naive Bayes, Gaussian mixture model (GMM) or generative model.

[0482] 48. The method as described in any of the preceding paragraphs, wherein the computational model is performed in an electronic device for providing candidate biological sequences, the method comprising:

[0483] - Obtain input data indicating the input biological sequence;

[0484] - The candidate biological sequence is determined by applying a model, such as a generative model, to the input data, preferably wherein the generative model is non-unidirectional; and

[0485] - Provides biological sequence data indicating the candidate biological sequence.

[0486] 49. The method as described in any of the preceding paragraphs, wherein the input biological sequence comprises one or more target polynucleotides identified in step e).

[0487] 50. The method as described in any of the preceding paragraphs, wherein the input biological sequence is one or more of the following: an amino acid sequence of a target polypeptide, a nucleic acid sequence encoding the target polypeptide, a control sequence (e.g., an expression control sequence), and a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence).

[0488] 51. The method as described in any of the preceding paragraphs, wherein the candidate biological sequence is one or more of the following: a control sequence (e.g., an expression control sequence), a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence), an amino acid sequence of a target polypeptide, and a nucleic acid sequence encoding a target polypeptide.

[0489] 52. The method as described in any of the preceding paragraphs, wherein the input biological sequence is an amino acid sequence of the target polypeptide and / or a nucleic acid sequence encoding the target polypeptide, and wherein the candidate biological sequence is a nucleic acid sequence that enhances compatibility with host cells.

[0490] 53. The method as described in any of the preceding paragraphs, wherein the input biological sequence is an amino acid sequence of a target polypeptide and / or a nucleic acid sequence encoding the target polypeptide, and wherein the candidate biological sequence is a control sequence (e.g., an expression control sequence) and / or a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence).

[0491] 54. The method as described in any of the preceding paragraphs, wherein the input biological sequence is a control sequence (e.g., an expression control sequence) and / or a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence), and wherein the candidate biological sequence is a nucleic acid sequence encoding a target polypeptide.

[0492] 55. The method as described in any of the preceding paragraphs, wherein the generative model is one or more of the following: a generative adversarial network (GAN) model, a Wasserstein generative adversarial network model, a diffusion model, and a variational autoencoder.

[0493] 56. The method as described in any of the preceding paragraphs, wherein applying the non-unidirectional generative model to the input data comprises dividing the non-unidirectional generative model into a plurality of generators, wherein each of the plurality of generators is configured to determine one or more candidate biological sequences based on the input data for a subset of nucleotides and / or a subset of amino acids and predetermined criteria.

[0494] 57. The method as described in any of the preceding paragraphs, wherein determining the candidate biological sequence by applying the model to the input data comprises:

[0495] - Use the generator to predict the compatibility of the candidate biological sequence with the host cell; and

[0496] - Identify candidate biological sequences that have predicted compatibility that meet the predetermined criteria.

[0497] 57a. The method described in paragraph 57, wherein the model is a generative model.

[0498] 58. The method as described in any of the preceding paragraphs, wherein the predetermined criterion is based on one or more of the following:

[0499] - The proportion of nucleotide sets in the candidate biological sequence;

[0500] -The type of host cell;

[0501] -The host cell genus or species;

[0502] - The GC content of the host cell genome;

[0503] - The GC content of the candidate biological sequence; and

[0504] - Parameters related to the characteristics of the candidate biological sequence.

[0505] 59. The method as described in any of the preceding paragraphs, the method comprising training the model based on a training set of biological sequences, wherein the training set of biological sequences includes training data indicating one or more biological sequences associated with the host cell.

[0506] 60. The method as described in any of the preceding paragraphs, wherein the training set of the biological sequence is heterologous to the genus of the host cell, preferably heterologous to one or more species of the host cell.

[0507] 61. The method as described in any of the preceding paragraphs, wherein the training data includes training input data indicating one or more of the following: an amino acid sequence of a target polypeptide, a nucleic acid sequence encoding the target polypeptide, a control sequence (e.g., an expression control sequence), and a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence).

[0508] 62. The method as described in any of the preceding paragraphs, wherein the training data includes training output data indicating one or more of the following: a control sequence (e.g., an expression control sequence), a nucleic acid sequence encoding a control sequence (e.g., an expression control sequence), an amino acid sequence of a target polypeptide, and a nucleic acid sequence encoding a target polypeptide.

[0509] 63. The method as described in any of the preceding paragraphs, wherein training the model includes using a discriminator as input to predict a score indicating that the training candidate biological sequence is a reference biological sequence.

[0510] 64. The method as described in any of the preceding paragraphs, the method comprising obtaining experimental data relating to the candidate biological sequence and the host cell from a test environment data repository; wherein the experimental data indicates the yield performance of the candidate biological sequence relating to the host cell.

[0511] 65. The method as described in any of the preceding paragraphs, the method comprising validating the candidate biological sequence based on the experimental data.

[0512] 66. The method as described in any of the preceding paragraphs, wherein the method includes selecting one or more generators based on the experimental data.

[0513] 67. The method as described in any of the preceding paragraphs, wherein the method includes adjusting the model based on the experimental data.

[0514] 68. The method of any of the preceding claims, wherein obtaining input data indicating an input biological sequence comprises obtaining the input data of the input biological sequence from a database and / or memory of the electronic device.

[0515] 69. An electronic device comprising memory circuitry, processor circuitry, and an interface, wherein the electronic device is configured to perform any of the methods described in any of the preceding paragraphs.

[0516] 70. A computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by an electronic device, cause the electronic device to perform any of the methods described in any of the preceding paragraphs.

[0517] 71. The method as described in any of the preceding paragraphs, the method comprising the additional step h) generating one or more synthetic target polynucleotides based on the output of the computational model.

[0518] 72. The method as described in any of the preceding paragraphs, wherein the one or more synthetic polynucleotides produced in step h) comprise or consist of a candidate biological sequence.

[0519] 73. The method as described in any of the preceding paragraphs, wherein the one or more synthetic target polynucleotides produced in step h) are codon-optimized.

[0520] 74. The method as described in any of the preceding paragraphs, wherein the one or more synthetic target polynucleotides produced in step h) encode a polypeptide having increased substrate binding, increased receptor binding, increased substrate specificity, increased specific activity, and / or increased stability.

[0521] 75. The method as described in any of the preceding paragraphs, wherein the one or more synthetic target polynucleotides produced in step h) result in an increase in the expression of the target polypeptide.

[0522] 76. The method as described in any of the preceding paragraphs, wherein the one or more synthetic target polynucleotides produced in step h) contain a control sequence.

[0523] 77. The method as described in any of the preceding paragraphs, wherein the droplet sorter (200) includes one or more sorting devices (401, 402).

[0524] 78. The method as described in any of the preceding paragraphs, wherein the one or more sorting devices comprise or consist of: one or more electrodes, one or more acoustic generators, one or more valves and / or one or more pressure-controlled outlets.

[0525] 79. The method as described in any of the preceding paragraphs, wherein the one or more sorting devices comprise at least two electrodes.

[0526] 80. The method as described in any of the preceding paragraphs, wherein the one or more sorting devices comprise an electrode.

[0527] 81. The method as described in any of the preceding paragraphs, wherein the one or more sorting devices consist of two electrodes.

[0528] 82. The method as described in any of the preceding paragraphs, wherein the biological library comprises or is composed of wild-type cells with different genotypes and / or different phenotypes.

[0529] 83. The method as described in any of the preceding paragraphs, wherein the biological library comprises or is composed of recombinant cells.

[0530] 84. The method as described in any of the preceding paragraphs, wherein the biological library encodes different variants of the same target polypeptide, preferably the target polypeptide being an enzyme.

[0531] 85. The method as described in any of the preceding paragraphs, wherein the biological library encodes different signal peptide variants.

[0532] 86. The method as described in any of the preceding paragraphs, wherein the biological library encodes different promoter variants.

[0533] 87. The method as described in any of the preceding paragraphs, wherein the biological library comprises different DNA sequences optimized by codons that encode the same amino acid sequence of a target polypeptide, such as a signal peptide and / or an enzyme.

[0534] 88. The method as described in any of the preceding paragraphs, wherein the target polynucleotide encodes the target polypeptide.

[0535] 89. The method as described in any of the preceding paragraphs, wherein the biological library comprises or is composed of a plurality of target polynucleotides, each target polynucleotide encoding a variant of a target polypeptide.

[0536] 90. The method as described in any of the preceding paragraphs, wherein the target polynucleotide comprises a first target polynucleotide encoding a control sequence and a second target polynucleotide encoding a target polypeptide.

[0537] 91. The method as described in any of the preceding paragraphs, wherein the biological library comprises or consists of a plurality of target polynucleotides, each target polynucleotide encoding a variant of a control sequence.

[0538] 92. The method as described in any of the preceding paragraphs, wherein the control sequence is a promoter sequence, signal peptide, leader sequence, polyadenylation sequence, propeptide sequence, or transcription terminator.

[0539] 93. The method of any of the preceding paragraphs, wherein the target polynucleotide comprises a first target polynucleotide encoding a signal peptide and a second target polynucleotide encoding a target polypeptide, wherein the first target polynucleotide is operatively linked to the second target polynucleotide and is located upstream of the second target polynucleotide.

[0540] 94. The method of any of the preceding paragraphs, wherein the target polynucleotide comprises a first target polynucleotide containing a promoter sequence and a second target polynucleotide encoding a target polypeptide, wherein the first target polynucleotide is operatively linked to the second target polynucleotide and is located upstream of the second target polynucleotide.

[0541] 95. The method as described in any of the preceding paragraphs, wherein the biological library comprises the same second-target polynucleotide and multiple variants of the first-target polynucleotide.

[0542] 96. The method as described in any of the preceding paragraphs, wherein the biological library comprises the same first target polynucleotide and multiple variants of the second target polynucleotide.

[0543] 97. The method as described in any of the preceding paragraphs, wherein the first target polynucleotide is heterologous to the second target polynucleotide.

[0544] 98. The method as described in any of the preceding paragraphs, wherein the first target polynucleotide is endogenous to the second target polynucleotide.

[0545] 99. The method as described in any of the preceding paragraphs, wherein the one or more target polynucleotides comprise a promoter, a polynucleotide encoding a signal peptide, a polynucleotide encoding a target polypeptide, or a natural host cell gene.

[0546] 100. The method as described in any of the preceding paragraphs, wherein the target polynucleotide is substantially the whole genome of the host cell.

[0547] 101. The method as described in any of the preceding paragraphs, wherein the one or more target polynucleotides are heterologous to the host cell.

[0548] 102. The method as described in any of the preceding paragraphs, wherein the one or more target polynucleotides are endogenous to the host cell.

[0549] 103. The method as described in any of the preceding paragraphs, wherein the first target polynucleotide is heterologous to the host cell.

[0550] 104. The method as described in any of the preceding paragraphs, wherein the first target polynucleotide is endogenous to the host cell.

[0551] 105. The method as described in any of the preceding paragraphs, wherein the second target polynucleotide is heterologous to the host cell.

[0552] 106. The method as described in any of the preceding paragraphs, wherein the second target polynucleotide is endogenous to the host cell.

[0553] 107. The method as described in any of the preceding paragraphs, wherein the first and second target polynucleotides are heterologous to the host cell.

[0554] 108. The method as described in any of the preceding paragraphs, wherein the first and second target polynucleotides are endogenous to the host cell.

[0555] 109. The method as described in any of the preceding paragraphs, wherein the one or more target polynucleotides encode a target polypeptide.

[0556] 110. The method as described in any of the preceding paragraphs, wherein the target polypeptide is an enzyme, nanobody, antibody, antibody fragment, fluorescent polypeptide (e.g., GFP) or α-lactalbumin.

[0557] 111. The method as described in any of the preceding paragraphs, wherein the amount of screenable product in the droplet is proportional to the amount of polypeptide encoded by the one or more target polynucleotides.

[0558] 112. The method as described in any of the preceding paragraphs, wherein the amount of screenable product in the droplet is inversely proportional to the amount of polypeptide encoded by the one or more target polynucleotides.

[0559] 113. The method as described in any of the preceding paragraphs, wherein the biological library comprises at least 100 different one or more target polynucleotides, at least 200 different one or more target polynucleotides, at least 500 different one or more target polynucleotides, at least 1,000 different one or more target polynucleotides, at least 2,000 different one or more target polynucleotides, at least 3,000 different one or more target polynucleotides, at least 5,000 different one or more target polynucleotides, at least 10,000 different one or more target polynucleotides, at least 100,000 different one or more target polynucleotides, at least 1,000,000 different one or more target polynucleotides, at least 10,000,000 different one or more target polynucleotides, at least 50,000,000 different one or more target polynucleotides, or at least 100,000,000 different target polynucleotides.

[0560] 114. The method as described in any of the preceding paragraphs, wherein the biological library comprises at least 100 different host cells, at least 200 different host cells, at least 500 different host cells, at least 1,000 different host cells, at least 2,000 different host cells, at least 3,000 different host cells, at least 5,000 different host cells, at least 10,000 different host cells, at least 100,000 different host cells, at least 200,000 different host cells, at least 500,000 different host cells, at least 1,000,000 different host cells, at least 5,000,000 different host cells, at least 10,000,000 different host cells, or at least 100,000,000 different host cells.

[0561] 115. The method as described in any of the preceding paragraphs, wherein the amount of screenable product in the droplet is proportional to one or more of the following: stability of the target polypeptide, transcription of the target polypeptide, translation of the target polypeptide, secretion of the target polypeptide, yield of the target polypeptide, binding strength of the target polypeptide to the target molecule, and activity of the target polypeptide.

[0562] 116. The method as described in any of the preceding paragraphs, wherein the amount of screenable product in the droplet is inversely proportional to one or more of the following: the stability of the target polypeptide, the transcription of the target polypeptide, the translation of the target polypeptide, the secretion of the target polypeptide, the yield of the target polypeptide, the binding strength of the target polypeptide to the target molecule, and the activity of the target polypeptide.

[0563] 117. The method as described in any of the preceding paragraphs, wherein the amount of screenable product in the droplet is proportional to one or more of the following: cell number, host cell viability, host cell division rate, host cell growth rate, host cell size, and host cell protein secretion.

[0564] 118. The method as described in any of the preceding paragraphs, wherein the amount of screenable product in the droplet is inversely proportional to one or more of the following: cell number, host cell viability, host cell division rate, host cell growth rate, host cell size, and host cell protein secretion.

[0565] 119. The method as described in any of the preceding paragraphs, wherein the one or more droplets contain a substrate.

[0566] 120. The method as described in any of the preceding paragraphs, wherein the substrate comprises or consists of the screenable product.

[0567] 121. The method as described in any of the preceding paragraphs, wherein the substrate is a fluorescent substrate.

[0568] 122. The method as described in any of the preceding paragraphs, wherein the substrate is rhodamine, which is capable of producing fluorescence.

[0569] 123. The method as described in any of the preceding paragraphs, wherein the substrate is a fluorescent dye.

[0570] 124. The method as described in any of the preceding paragraphs, wherein the substrate is a fluorescent substrate.

[0571] 125. The method as described in any of the preceding paragraphs, wherein the substrate comprises a fluorophore (e.g., fluorescein) or fluorescein-labeled starch.

[0572] 126. The method as described in any of the preceding paragraphs, wherein the substrate is Nile Red.

[0573] 127. The method as described in any of the preceding paragraphs, wherein the substrate is DAPI (4',6-diamidinyl-2-phenylindole).

[0574] 128. The method as described in any of the preceding paragraphs, wherein prior to the optional incubation, each droplet comprises at most 0.01 cells, at most 0.02 cells, at most 0.03 cells, at most 0.04 cells, at most 0.05 cells, at most 0.06 cells, at most 0.07 cells, at most 0.08 cells, at most 0.09 cells, at most 0.1 cells, at most 0.2 cells, at most 0.3 cells, at most 0.4 cells, at most 0.5 cells, at most 0.6 cells, or at most 0.7 cells; preferably an average occupancy of at most 0.1 cells.

[0575] 129. The method as described in any of the preceding paragraphs, wherein each droplet comprises at most 0.01, at most 0.02, at most 0.03, at most 0.04, at most 0.05, at most 0.06, at most 0.07, at most 0.08, at most 0.09, at most 0.1, at most 0.2, at most 0.3, at most 0.4, at most 0.5, at most 0.6, or at most 0.7 target polynucleotides; preferably an average occupancy of at most 0.1 target polynucleotides.

[0576] 130. The method as described in any of the preceding paragraphs, wherein the droplet sorting is facilitated by an electric field generated by one or more electrodes (401, 402) adjacent to the droplet sorter.

[0577] 131. The method as described in any of the preceding paragraphs, wherein the droplet sorting is facilitated by acoustic waves generated by one or more acoustic generators (401, 402) adjacent to the droplet sorter.

[0578] 132. The method as described in any of the preceding paragraphs, wherein droplet sorting is facilitated by local pressure changes generated by one or more pressure-controlled outlets (401, 402) adjacent to the droplet sorter, for example, wherein the one or more pressure-controlled outlets are included in one or more output channels.

[0579] 133. The method as described in any of the preceding paragraphs, wherein the amount of the screenable product in step c) is determined using fluorescence-based signal, absorbance, Raman spectroscopy, mass spectrometry (MS), or MALDI-MS.

[0580] 134. The method as described in any of the preceding paragraphs, wherein the relative and / or absolute amount of the screenable product / droplet is determined by the one or more sensor elements (600).

[0581] 135. The method as described in any of the preceding paragraphs, wherein after step d), one or more output channels contain at least 10,000 droplets, at least 50,000 droplets, at least 100,000 droplets, at least 500,000 droplets, at least 1,000,000 droplets, at least 2,000,000 droplets, at least 5,000,000 droplets, at least 10,000,000 droplets, or at least 100,000,000 droplets.

[0582] 136. The method as described in any of the preceding paragraphs, wherein the droplet sorter comprises at least four output channels, at least five output channels, at least six output channels, at least seven output channels, at least eight output channels, at least nine output channels, or at least ten output channels.

[0583] 137. The method as described in any of the preceding paragraphs, wherein the host cell is a yeast host cell, such as cells of the genera *Candida*, *Hansenula*, *Kluyveromyces*, *Pichia*, *Saccharomyces*, *Saccharomyces*, *Schizosaccharomyces*, or *Yersinia*, such as *Kluyveromyces lactis*, *Kalvatia*, *Saccharomyces cerevisiae*, *Saccharomyces sacchariformis*, *Saccharomyces douglas*, *Kluyveromyces kluyvernsis*, *Nordicia*, *Ovoyces*, or *Yersinia lipolytica*.

[0584] 138. The method as described in any of the preceding paragraphs, wherein the host cell is a filamentous fungal host cell, such as *Cladosporium*, *Aspergillus*, *Briefomus*, *Cirsium*, *Pseudomonas*, *Aureospora*, *Coprinus*, *Leptochloa*, *Cryptococcus*, *Ustilago*, *Fusarium*, *Pyrophyllus*, *Mucor*, *Hydrophyllus*, *Neurospora*, *Neurospora*, *Penicillium*, *Penicillium*, *Pseudomonas ... *Aspergillus*, *Ruminocytotrichum*, *Pleurotus*, *Schizophyllum*, *Basilaria*, *Thermophilus*, *Fusporium*, *Cyclophorus*, *Pterocaryon*, or *Trichoderma* cells, particularly *Aspergillus avocado*, *Aspergillus sulphureus*, *Aspergillus fumigatus*, *Aspergillus japonicus*, *Aspergillus nidus*, *Aspergillus niger*, *Aspergillus oryzae*, *Pterocaryon*, *Carnegiea*, *Pterocaryon .... Wax fungus, narrow-rimmed aurantiacus, keratophilic aurantiacus, Lukenowens aurantiacus, fecal aurantiacus, *Rhizopus spp.*, Queen's aurantiacus, tropical aurantiacus, brown aurantiacus, *Coprinus comatus*, *Caragana korshinskii*, *Fusarium moniliforme ... Spores, skin-colored Fusarium, Fusarium pseudocranioides, sulfur-colored Fusarium, round Fusarium, Fusarium pseudofilariae, Fusarium moniliforme, specific humic mold, soft-haired humic mold, Rhizopus spp., thermophilic filamentous mold, rough spores, Penicillium purpureum, Chlorella vulgaris, Pleurotus eryngii, Pleurotus eryngii, Clostridium emarginatum, Trichoderma harzianum, Trichoderma cornigrin, Trichoderma longibranchii, Trichoderma reesei, or green Trichoderma cells.

[0585] 139. The method as described in any of the preceding paragraphs, wherein the host cell is a prokaryotic host cell, for example, Gram-positive cells selected from the group consisting of: Bacillus spp., Clostridium spp., Enterococcus spp., Bacillus spp., Lactobacillus spp., Lactococcus spp., Bacillus spp., Staphylococcus spp., Streptococcus spp., or Streptomyces spp. cells, or Gram-negative bacteria selected from the group consisting of: Campylobacter spp., Escherichia coli, Flavobacterium spp., Fusobacterium spp., Helicobacter spp., Coptis spp., Neisseria spp., Pseudomonas spp., Salmonella spp., And Ureaplasma cells, such as alkalophilic Bacillus, amyloliquefaciens, brevis, circular Bacillus, Clausii, coagulant Bacillus, sclerosus, brilliant Bacillus, slow-release Bacillus, licheniformis, megaterium, brevis, thermophilic steatobacterium, Bacillus subtilis, Bacillus thuringiensis, Streptococcus equine, Streptococcus pyogenes, Streptococcus lactis and Streptococcus equine subsp. porphyria, non-chromogenic Streptococcus, insecticidal Streptococcus, blue Streptococcus, gray Streptococcus, and pale blue Streptococcus cells.

[0586] 140. The method as described in any of the preceding paragraphs, wherein the host cell is Bacillus subtilis.

[0587] 141. The method as described in any of the preceding paragraphs, wherein the host cell is Bacillus licheniformis.

[0588] 142. The method as described in any of the preceding paragraphs, wherein the host cell is Trichoderma reesei.

[0589] 143. The method as described in any of the preceding paragraphs, wherein the host cell is Aspergillus niger.

[0590] 144. The method as described in any of the preceding paragraphs, wherein the host cell is Aspergillus oryzae.

[0591] 145. The method as described in any of the preceding paragraphs, wherein the host cell is a Bifidobacterium, such as Bifidobacterium animalis or Bifidobacterium lactis subsp. animalis.

[0592] 146. A host cell that contains in its genome the target polynucleotide sequence generated in step h) and / or the polynucleotide sequence identified in step e).

[0593] 147. The host cell as described in any of the preceding paragraphs, wherein the host cell is isolated.

[0594] 148. The host cell as described in any of the preceding paragraphs, wherein the host cell is purified.

[0595] 149. The host cell as described in any of the preceding paragraphs, wherein the host cell contains at least two copies, such as three, four, five or more copies, of the target polynucleotide sequence.

[0596] 150. A method for producing a target polypeptide, the method comprising culturing cells as described in any of the preceding paragraphs under conditions conducive to the production of the polypeptide.

[0597] 151. The method as described in paragraph 150, further comprising recovering the polypeptide.

[0598] 152. A nucleic acid construct or expression vector comprising a target polynucleotide identified by step e) and / or a polynucleotide sequence generated in step h).

Claims

1. A method for screening biological libraries, the method comprising the following steps: a) Provide a microfluidic device including a droplet sorter (200) having at least three output channels (301, 302, 303). b) Provides a droplet emulsion containing a target polynucleotide library and screenable products. c) Determine the amount of screenable products from one or more droplets in the microfluidic device. d) The droplet sorter (200) sorts the one or more droplets into the receiving output channels of the at least three output channels (301, 302, 303), wherein the receiving output channel is determined based on the amount of screenable product per droplet, and wherein the at least three receiving output channels receive a plurality of droplets, the plurality of droplets comprising an amount of screenable product above and / or below one or more predetermined threshold levels. e) Identify one or more target polynucleotides present in the at least three output channels (301, 302, 303), and obtain the sequence data of the one or more target polynucleotides. f) For one or more output channels (301, 302, 303), assign a score to each of the one or more target polynucleotides, wherein the score is calculated based on the abundance of each of the one or more target polynucleotides in one of the one or more output channels (301, 302, 302).

2. The method of claim 1, wherein the droplet sorter comprises at least four output channels, at least five output channels, at least six output channels, at least seven output channels, at least eight output channels, at least nine output channels, or at least ten output channels.

3. The method of any one of claims 1-2, wherein the droplet emulsion comprises one or more host cells.

4. The method of any one of claims 1-3, wherein the microfluidic device includes an incubation zone (500), and wherein the method includes incubating the droplet emulsion under conditions that allow host cell growth and / or allow DNA transcription into RNA and / or allow RNA translation into polypeptides, preferably, the incubation is performed prior to step c).

5. The method of claim 4, wherein prior to the incubation, each of the one or more droplets comprises at most 0.01 cells, at most 0.02 cells, at most 0.03 cells, at most 0.04 cells, at most 0.05 cells, at most 0.06 cells, at most 0.07 cells, at most 0.08 cells, at most 0.09 cells, at most 0.1 cells, at most 0.2 cells, at most 0.3 cells, at most 0.4 cells, at most 0.5 cells, at most 0.6 cells, or at most 0.7 cells; preferably an average occupancy of at most 0.1 cells.

6. The method of any one of claims 3-5, wherein each host cell contains one or more target polynucleotides of the target polynucleotide library.

7. The method of any of the preceding claims, wherein the screenable product is catalyzed by an enzyme, preferably the enzyme being encoded by the target polynucleotide.

8. The method of any of the preceding claims, wherein the screenable product is produced by these host cells.

9. The method of any of the preceding claims, wherein the biological library comprises or is composed of a plurality of target polynucleotides, each target polynucleotide encoding a variant of a target polypeptide.

10. The method of any of the preceding claims, wherein the target polynucleotide comprises a first target polynucleotide encoding a control sequence and a second target polynucleotide encoding a target polypeptide.

11. The method as described in any of the preceding claims, the method comprising step g) training a computational model, for example, a machine learning algorithm, with the sequence data obtained from step e) and / or the score data obtained from step f).

12. The method of any of the preceding claims, wherein the computational model provides candidate biological sequences, the method comprising: - Obtain input data indicating the input biological sequence; - The candidate biological sequence is determined by applying the model to the input data; as well as - Provides biological sequence data indicating the candidate biological sequence.

13. The method of claim 12, wherein the input biological sequence comprises one or more target polynucleotides identified in step e).

14. The method as described in any of the preceding claims, the method comprising the additional step h) generating one or more synthetic target polynucleotides using the computational model.

15. A host cell whose genome contains: (i) the synthetic target polynucleotide produced in step h) as described in claim 14 and / or the target polynucleotide identified in step e) as described in claim 1; and (ii) Polynucleotides encoding the target polypeptide.

16. A method for producing a target polypeptide, the method comprising culturing the cells as described in claim 15 under conditions conducive to the production of the polypeptide.

Citation Information

Patent Citations

  • Process for the production of protein products in Aspergillus oryzae and a promoter for use in Aspergillus

    EP0238023A2

  • Improving a Microorganism by CRISPR-Inhibition

    US20190185847A1

  • Directed evolution of novel binding proteins

    US5223409A

  • Surface expression libraries of heteromeric receptors

    WO1992006204A1

  • Nucleotide sequences for the control of the expression of DNA sequences in a cellular host

    WO1994025612A2