Mapping systems and methods for mutagenicity and activity profiles

A high-throughput method for evaluating protein variants in mammalian cells using a plasmid cDNA library and viral transduction allows for comprehensive profiling of mutagenesis effects, addressing the lack of effective assays in existing technologies and enabling the discovery of functionally relevant protein variants.

JP2026515702APending Publication Date: 2026-05-19HELIGENICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
HELIGENICS INC
Filing Date
2024-04-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods lack a high-throughput assay to effectively evaluate molecular function in the context of human or mammalian cells, particularly for identifying mutagenic effects on protein activity and discovering pharmacologically relevant protein variants.

Method used

A method involving the creation of a plasmid cDNA library with unique molecular identifiers, conversion into a viral library, transduction into engineered cell lines with surface receptors and reporters, and analysis through flow cytometry and sequencing to determine the biological activity of unique protein variants.

Benefits of technology

Enables comprehensive profiling of mutagenesis effects on protein activity, facilitating the discovery and validation of unique variants with altered functions, including loss-of-function, gain-of-function, and drug sensitivity, in a high-throughput manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026515702000001_ABST
    Figure 2026515702000001_ABST
Patent Text Reader

Abstract

The methods disclosed herein utilize a high-throughput screening assay (GigaAssay) to generate a comprehensive gene activity mutagenesis (MEGA)-mutagenesis profile (Map). These methods can be used to evaluate mutagenesis on any gene under any conditions (e.g., drug therapy) using any assay in cultured mammalian cells. Therefore, the methods provided herein can be used to discover and screen dominant-negative variants and to provide a reliable solution to the problem of identifying pharmacologically active unique variants of proteins. Furthermore, these methods can be integrated into cell-based assays to investigate disease pathogenesis and test promising drugs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference This application claims the benefits of U.S. Provisional Patent Application No. 63 / 494,228, filed on April 4, 2023, and U.S. Provisional Patent Application No. 63 / 498,831, filed on April 28, 2023, and incorporates the entire contents of these by reference herein.

Summary of the Invention

Means for Solving the Problems

[0002] Overview A method for identifying mutagenic effects on protein activity, comprising: (a) obtaining a library containing a plurality of cDNAs, wherein each cDNA in the plurality encodes a unique variant(s) of a target protein, and each unique variant has one or more amino acid substitutions compared to the target protein; and (b) independently and individually incorporating one of the plurality of cDNAs into a plasmid, thereby forming a plasmid cDNA library, wherein each plasmid in the plasmid cDNA library is encoded with a unique molecular identifier (UMI) or a barcode group. Methods are disclosed herein that further include the steps of: (c) converting a plasmid cDNA library into a viral library containing virions, wherein each virion contains only one plasmid from the plasmid cDNA library, and the viral library may optionally be a lentiviral library; (d) manipulating a human or mammalian cell line to encode a receptor operably linked to a reporter system, or optionally a receptor expressed on its surface, thereby forming a manipulated human or mammalian cell line, wherein the receptor and reporter system may indicate whether one of the unique variants binds to the receptor and modulates its biological activity; and (e) transducing virions from the viral library of (c) into cells of the manipulated human or mammalian cell line of (d) so that each virion is transduced into a different cell, thereby forming a library of transduced cells. In some embodiments, the method results in the expression of a unique variant of the target protein in transduced cells of the manipulated or mammalian cell line. In some embodiments, the method further includes the step of determining a measure of the biological activity of a unique variant in transduced cells of an engineered or mammalian cell line.In some embodiments, the determination step includes calculating a measure of biological activity by computer. In some embodiments, the method further includes sorting transduced cells containing the unique variant by flow cytometry using UMI or barcode groups. In some embodiments, the method further includes sorting multiple cells or all cells from a library of transduced cells to form a flow-sorted pool(s) of UMI(s) or a flow-sorted pool(s) of barcode groups. In some embodiments, the measurement of biological activity includes detecting the binding of the unique variant(s) to a receptor operably linked to a reporter system, and determining the amount, intensity, or both of the report from the reporter system when measured by fluorescence microscopy, flow cytometry, or both. In some embodiments, the method further includes comparing the distribution of a flow-sorted pool of UMI or barcode groups of single cell populations individually expressing the unique variant(s) with the distribution of a flow-sorted pool of single cell populations expressing wild-type target protein or wild-type cDNA without one or more amino acid substitutions. In some embodiments, the measurement of the biological activity of unique variants(s) is repeated at least 100 times using UMIs, barcodes, or both. In some embodiments, each UMI or barcode independently contains approximately 12 to 32 nucleotides. In some embodiments, the viral library is transduced into cells of an engineered human or mammalian cell line with an infection multiplicity (MOI) of at least 0.1 to at least 1. In some embodiments, the MOI is 0.1. In other embodiments, the MOI prevents double insertion of virions into a single transduced cell or minimizes the transduction error rate. In some embodiments, cells of an engineered human or mammalian cell line are selected using a marker for lentiviral incorporation. In some embodiments, the receptor and reporter system is operably linked to an inducible promoter. In some embodiments, the inducible promoter drives the expression of the receptor and reporter system.In some embodiments, the inducible promoter includes a doxycycline-inducible promoter. In some embodiments, the reporter system includes a polynucleotide encoding a fluorescent protein, the fluorescent protein being GFP. In some embodiments, a single cell population is isolated using flow cytometry based on UMI and sorted into one or more bins. In some embodiments, one or more bins contain cDNAs encoding unique variants, their amplicons, or any combination thereof. In some embodiments, cDNAs encoding unique variants, their amplicons, or any combination thereof are sequenced using a sequencing method. In some embodiments, the sequencing method includes next-generation sequencing, Sanger sequencing, whole-genome sequencing, RNA sequencing, or shotgun sequencing. In some embodiments, the sequencing method is next-generation sequencing, which generates the sequence dataset. In some embodiments, the sequence data includes read depth for each unique variant. In some embodiments, each unique variant has a read depth of approximately 2,000 × to approximately 90,000 × sequencing coverage. In some embodiments, a population of single cells expressing a unique variant is compared to a population of single cells expressing a wild-type target protein using a statistical model to test specific hypotheses regarding the biological activity of each unique variant. In some embodiments, the biological activity of each unique variant includes loss-of-function variants, gain-of-function variants, variants with substantially similar activity to the wild-type target protein, drug-resistant variants, or drug-sensitive variants. In some embodiments, the method further includes a step of validating the method by comparing the biological activity of a subset of unique variants analyzed by the method with the previously determined activity of the wild-type target protein. In some embodiments, the method further includes a step of validating the method by comparing the true negatives determined by the method with true negatives determined by an independent method.In some embodiments, the method further includes a step of validating the method by comparing the biological activity of a subset of unique variants analyzed by the method with independent testing of a separate set of clones containing the unique variant(s). In some embodiments, the method further includes a step of validating the method by comparing the results of the method across one or more different samples. In some embodiments, the method further includes a step of validating the method by comparing the results of the method in two different cell lines. In some embodiments, the method further includes a step of generating an oligonucleotide sequence dataset. In some embodiments, the method further includes the step of analyzing an oligonucleotide sequence dataset using a computer, the analysis step of (a) generating a unique variant-UMI index library using long reads derived from next-generation sequencing (NGS), (b) identifying UMI counts and barcode bins by analyzing short reads of a flow-selected group, (c) calculating an activity score from the UMI-barcode read counts, including the UMI counts and barcode bins from (b), (d) evaluating the effect of the unique variant(s) on biological activity by calculating the read percentage and p-value of a reporter protein, e.g., a fluorescent protein, and (e) accurately and quantitatively evaluating the activity of the unique variant(s) relative to a previously characterized wild type, where the fluorescent protein is GFP and the biological activity includes gene activity or protein activity. In some embodiments, the transduction cell library contains approximately 10,000 to 10 million cells.In some embodiments, the transduced cell library contains approximately 10,000 cells, 20,000 cells, 30,000 cells, 40,000 cells, 50,000 cells, 60,000 cells, 70,000 cells, 80,000 cells, 90,000 cells, 100,000 cells, 150,000 cells, 200,000 cells, 300,000 cells, 400,000 cells, and 500,000 cells. A cell containing 600,000 cells, 700,000 cells, 800,000 cells, 900,000 cells, 1,000,000 cells, 2,000,000 cells, 3,000,000 cells, 4,000,000 cells, 5,000,000 cells, 6,000,000 cells, 7,000,000 cells, 8,000,000 cells, 9,000,000 cells, or 10,000,000 cells.

[0003] Furthermore, protein variants discovered by the methods disclosed herein are provided herein.

[0004] A library of transduced cells formed by the methods disclosed herein is provided herein. In some embodiments, the library of transduced cells contains approximately 10,000 to 10 million cells.

[0005] Furthermore, the present invention provides a library of isolated and purified transdextrin cells comprising approximately 10,000 to 10 million cells, wherein each isolated and purified transdextrin cell comprises a plasmid containing cDNA encoding a unique variant of the wild-type protein, and the cell surface receptor can be examined for each unique variant, such that each transdextrin cell further comprises a surface-expressed receptor operably coupled to a reporter system. In some embodiments, each isolated and purified transdextrin further comprises a unique variant on its surface.

[0006] Embedding by reference All published documents, patents, and patent applications referenced herein are incorporated by reference to the same extent as each individual published document, patent, and patent application is incorporated by reference in detail and individually.

[0007] Novel features of the present invention are described in detail in the appended claims. A better understanding of the features and advantages of this disclosure can be obtained by referring to the following detailed description and the following appended drawings, which describe exemplary embodiments utilizing the principles of this disclosure. [Brief explanation of the drawing]

[0008] [Figure 1A] Figures 1A–1D illustrate schematic diagrams for generating a gene mutation library (GML) analyzed using the high-throughput screening assay (GigaAssay) disclosed herein to produce a comprehensive gene activity mutagenesis (MEGA)-mutagenesis activity profile (Map) combined (MEGA-Map). Figure 1A illustrates the steps for creating the gene mutation library disclosed herein. Figure 1B illustrates the steps for validating the library disclosed herein. Figure 1C illustrates the high-throughput assay disclosed herein. Figure 1D illustrates the steps of the bioinformatics pipeline. [Figure 1B] Same as above. [Figure 1C] Same as above. [Figure 1D] Same as above.

[0009] [Figure 2] Figure 2 shows the variant assembly. Figure 2 shows (from left to right) the 5'LTR, TreTIGHTpro, the sequence encoding the target gene (GOI) containing the variant region corresponding to the 50-amino acid coding sequence in the context of the lentiviral expression vector, the AsiSI-Mlul restriction site on the 3'UTR corresponding to the unique molecular identifier (UMI) insertion site, and then the 3'LTR.

[0010] [Figure 3]Figure 3 shows a variant assembly. Figure 3 shows a magnified view of the variant region in the context of dsDNA. A 150 bp variant region and a 25 bp overlapping DNA sequence flanking the variant region create a 200 bp tile.

[0011] [Figure 4] Figure 4 shows a comparison of inverse PCR products and variant oligopools in preparation for HiFi DNA assembly. Inverse PCR primers (arrows) are designed to contain a 25 bp sequence that overlaps with the oligopool. The oligopool consists of up to several thousand oligos containing the desired mutation (star).

[0012] [Figure 5] Figure 5 shows T5 exonuclease activity in HiFi DNA Assembly. T5 exonuclease partially digests (chews back) the dsDNA strand from 5' to 3'. This exposes the 3' DNA strand, which anneals to the complementary sequence found in the oligopool. Note: In this example, we focus on one oligo from the oligopool.

[0013] [Figure 6] Figure 6 shows the activity of HiFi DNA polymerase and DNA ligase in HiFi DNA assembly. HiFi DNA polymerase synthesizes DNA from 5' to 3', synthesizing the reverse strand of the variant oligo and simultaneously filling the gap created after DNA strand annealing. Then, DNA ligase closes the nick to create a continuous DNA strand. Note: In this example, we focus on one oligo from the oligopool.

[0014] [Figure 7]Figure 7 represents the continuation of T5 exonuclease activity in HiFi DNA Assembly. The T5 exonuclease continues to digest the exposed ends of the DNA strands in the 5' to 3' direction, exposing the DNA and allowing the complementary strands to anneal. Note: In this example, one oligo from the oligo pool is being focused on.

[0015] [Figure 8] Figure 8 represents the continuation of HiFi DNA polymerase and DNA ligase activities in HiFi DNA Assembly. The DNA polymerase continues DNA synthesis in the 5' to 3' direction to fill the gaps, and the DNA ligase seals the nicks to create a continuous strand of DNA. The final result of a complete assembly is a recombinant vector free of traces, containing the variant of interest cloned within the variant region. Note: In this example, one oligo from the oligo pool is being focused on.

[0016] [Figure 9] Figure 9 represents the final product of HiFi DNA Assembly. The result of a complete assembly is a recombinant vector free of traces, containing the variant of interest cloned within the variant region. In this exemplary HiFi DNA assembly, one oligo from the oligo pool that results in a particular combination of variants is being focused on.

[0017] [Figure 10-1]Figure 10 represents a target DNA sequence, which is the variant region targeted by the gene mutation library. The target DNA sequence is divided into 200-nucleotide tiles. These tiles correspond to 200-mer synthetic oligos for the variant of interest. Each tile consists of a variant region (yellow) flanked by 25 nucleotides that also function as inverse PCR primer binding sites for inverse PCR of vector DNA and overlap with the vector cloning site. The tiles overlap such that the entire target region is covered. In this example, the target DNA sequence is 600 bp in length and is covered by three overlapping tiles. Each tile is processed as a separate variant assembly reaction in the expression vector using the same general strategy as above. Variant assembly results in three separate pools of assembly products, which are combined to generate a plasmid variant library (without UMI). In the 3’UTR of the vector, there are unique restriction sites (AsiSI and MluI). Using these sites, the vector DNA is linearized, and by HiFi DNA assembly, a UMI is added to the 3’UTR using an ssDNA UMI oligo, resulting in a plasmid variant library with a UMI added to each plasmid DNA molecule of the library. [Figure 10-2] Same as above.

[0018] [Figure 11] Figure 11 represents an enlarged view focusing on the 3’UTR UMI assembly region in the context of dsDNA. In this example, AsiSI and MluI are the restriction sites used for linearizing the variant plasmid library. The thick dashed line indicates the cleavage site.

[0019] [Figure 12]Figure 12 shows a comparison of linearized vector DNA using AsiSI-MluI and UMI oligos. The HiFi DNA Assembly reaction is performed using AsiSI-MluI linearized vector DNA and UMI oligos having a 25-nucleotide overlap sequence containing the cloning site. The assembly reaction is carried out in the same manner as in Figures 4-7 (above) until the assembly is complete.

[0020] [Figure 13] Figure 13 shows an example of a completed plasmid gene mutation library. The final product of the UMI assembly is a plasmid GML containing the target variant along with a unique molecular identifier for each plasmid DNA molecule in the library.

[0021] [Figure 14-1] Figure 14 shows a bioinformatics pipeline flowchart. The goal of this process is to evaluate the performance of variants identified and generated by the high-throughput screening assay (GigaAssay) disclosed herein in performance screening for GMLs that require long-read sequencing identification of libraries, resulting in a sequence dataset. To analyze the sequence dataset, including the output of the selected pool, unique variants must first be identified and mapped to 32BP barcode bins. After identification, the barcode bins are quantified and their presentation in each of the selected pools is analyzed. This allows for interpretation of the results using the method disclosed herein. [Figure 14-2] Same as above.

[0022] [Figure 15] Figure 15 shows the design of a dominant-negative GML library of candidate genes prioritized by GWAS, polygene risk score (PRS), or other omics approaches.

[0023] [Figure 16]Figure 16 shows the design of a GML library of chimeric genes containing dominant-negative versions of a single candidate gene, prioritized by GWAS, polygene risk score (PRS), or other omics approaches. [Modes for carrying out the invention]

[0024] Detailed description of the present invention Overview The methods disclosed herein utilize a high-throughput screening assay (GigaAssay) to generate a comprehensive mutagenesis (MEGA)-mutagenesis profile (Map) for gene activity. No high-throughput assay exists that broadly evaluates molecular function in the context of human or mammalian cells. Molecular function is key to understanding mechanisms, disease etiology, and therapeutic drug development. The methods disclosed herein can be used to evaluate mutagenesis to any gene under any conditions (e.g., drug therapy) by any assay in mammalian cells in culture. Phage or yeast displays, yeast 1 or 2 hybrids, DNA-coding libraries (DELs), and affinity mass spectrometry evaluate molecular interactions, a common type of function, but not interactions in living mammalian cells. The methods provided herein can be used to discover and screen dominant-negative variants. The methods provided herein provide a reliable solution to the problem of identifying unique variants of pharmacological activity. Here, the methods described herein detail an approach that can be used to prioritize lead protein variants by using the GigaAssay approach described to address this problem. Furthermore, the methods provided herein can be integrated into cell-based assays to investigate disease pathogenesis and test promising drugs.

[0025] definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art in which the methods disclosed herein pertain. Any similar or equivalent methods and materials may be used in the practice of the tests described herein, but preferred materials and methods are described herein. The following terms are used in the descriptions and claims of this disclosure:

[0026] Furthermore, it should be understood that the terms used herein are solely for the purpose of describing specific embodiments and are not intended to be limiting. The articles “a” and “an” are used herein to refer to one or more (i.e., at least one) grammatical objects of the article. For example, “an element” means one or more elements.

[0027] When "about" is used herein to refer to a measurable value, such as a quantity or period of time, it is intended to include a variation of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1% with respect to the specified value, because such variation is reasonable for carrying out the methods of this disclosure.

[0028] "Biomolecules" or "biological molecules" refer to molecules commonly found in or produced by biological organisms. In some embodiments, biological molecules include polymeric biological macromolecules (i.e., "biopolymers") having multiple subunits. Typical biomolecules include proteins, enzymes, and other polypeptides, DNA, RNA, and other polynucleotides, and may also include naturally occurring polymers, e.g., molecules that share some structural features with RNA (formed from nucleotide subunits), DNA (formed from nucleotide subunits), and peptides or polypeptides (formed from amino acid subunits), e.g., RNA analogs, DNA analogs, polypeptide analogs, peptide nucleic acids (PNAs), combinations of RNA and DNA (e.g., chimeraplasts), etc. Since any suitable biological molecule is available in this disclosure, it is not intended to limit biological molecules to any particular molecule, and include, but are not limited to, lipids, carbohydrates, or other organic molecules made up of one or more genetically encoding molecules (e.g., one or more enzymes or enzymatic pathways), etc.

[0029] The term “unique variant” or “unique variants” refers to a protein that contains one or more amino acid substitutions compared to the wild-type protein. Of particular interest in some embodiments of this disclosure are biomolecules having an active site that interacts with a reporter molecule (e.g., a receptor) to influence chemical or biological transformation, e.g., catalysis of a substrate, activation or inactivation of a biomolecule, more specifically, enzymes.

[0030] In some embodiments, "biological activity" or "activity" refers to: catalytic rate (k cat ), substrate binding affinity (K M / K D ), catalyst efficiency (k cat / K M), substrate specificity, chemoselectivity, regioselectivity, stereoselectivity, stereospecificity, ligand specificity, receptor agonism, receptor antagonistism, cofactor conversion, oxygen stability, protein expression level, solubility, thermal activity, thermal stability, pH activity, pH stability (e.g., at alkaline or acidic pH), glucose inhibition, and / or an increase or decrease in resistance to inhibitors (e.g., acetic acid, lectins, tannic acid, phenolic compounds) and proteases. Other desired activities may include altered profiles in response to specific stimuli (e.g., altered temperature and / or pH profiles). In the context of theoretical ligand design, optimization of targeted covalent inhibition (TCI) is one type of activity. In some embodiments, two or more variants screened herein act on the same substrate but differ with respect to one or more of the following activities: product formation rate, substrate-to-product conversion percentage, selectivity, and / or cofactor conversion percentage. This disclosure is not intended to limit any particular beneficial property and / or desired activity.

[0031] In some embodiments, “activity” is used to describe a more limited concept of an enzyme’s ability to catalyze the turnover of a substrate into a product. A relevant characteristic of the enzyme is its “selectivity” for a particular product, such as an enantiomer or a regioselective product. The broad definition of “activity” presented herein includes selectivity, although conventionally, selectivity may be considered distinct from enzyme activity.

[0032] "Code" refers to a specific sequence of nucleotides within a polynucleotide, such as a gene, cDNA, or mRNA, which functions as a template for the synthesis of other polymers and macromolecules in biological processes (i.e., rRNA, tRNA, and mRNA) or a specific sequence of amino acids, and which has the biological properties derived therefrom. Thus, a gene codes for a protein when the transcription and translation of the mRNA corresponding to that gene produces a protein in a cell or other biological system. Both the coding strand, whose nucleotide sequence is identical to the mRNA sequence and is usually presented in a sequence listing, and the non-coding strand, which is used as a template for the transcription of the gene or cDNA, can be said to code for the protein product or other product of that gene or cDNA.

[0033] As used herein, “endogenous” means any substance that originates from or is produced within an organism, cell, tissue, or system.

[0034] As used herein, the term “exogenous” means any substance introduced into or produced outside of an organism, cell, tissue, or system.

[0035] As used herein, the term "expression" is defined as the transcription and / or translation of a particular nucleotide sequence driven by its promoter.

[0036] An "expression vector" refers to a vector containing recombinant polynucleotides that include an expression regulatory sequence operably ligated to the nucleotide sequence to be expressed. An expression vector contains sufficient cis-acting elements for expression, and other elements for expression may be supplied by a host cell or an in vitro expression system. Expression vectors incorporating recombinant polynucleotides include all known in the art, e.g., cosmids, plasmids (e.g., naked or liposome-containing), and viruses (e.g., Sendai virus, lentivirus, retrovirus, adenovirus, and adeno-associated virus).

[0037] As used herein, "homology" refers to the subunit sequence identity between two polymer molecules, for example, two nucleic acid molecules, for example, two DNA molecules or two RNA molecules, or two polypeptide molecules. If the same monomer subunit occupies the same subunit position in both molecules, for example, if adenine occupies the same position in each of two DNA molecules, then these molecules are homologous in that position. The homology between two sequences is a positive function of the number of matching or homologous positions, for example, if half of the positions in two sequences (e.g., five positions in a polymer with a length of 10 subunits) are homologous, then the two sequences are 50% homologous, and if 90% of the positions (e.g., nine out of ten) are matching or homologous, then the two sequences are 90% homologous.

[0038] In the context of this disclosure, the following abbreviations are used for commonly existing nucleic acid bases: "A" refers to adenosine, "C" to cytosine, "G" to guanosine, "T" to thymidine, and "U" to uridine.

[0039] Unless otherwise specified, “nucleotide sequences encoding an amino acid sequence” includes all nucleotide sequences that are degenerate versions of each other and that encode the same amino acid sequence. Furthermore, the phrase “nucleotide sequences encoding a protein or RNA” may contain introns to the same extent that “nucleotide sequences encoding a protein” may contain introns in some versions.

[0040] The term "operably linked" refers to a functional linkage between a regulatory sequence that produces the expression of a heterologous nucleic acid sequence and the heterologous nucleic acid sequence itself. For example, when a first nucleic acid sequence is positioned in a functional relationship with a second nucleic acid sequence, the first nucleic acid sequence is operably linked to the second nucleic acid sequence. As an example, when a promoter affects the transcription or expression of a coding sequence, the promoter is operably linked to the coding sequence. Generally, operably linked DNA sequences are adjacent and located within the same reading frame, which is necessary to join two protein coding regions. The term "polynucleotide," as used herein, is defined as a chain of nucleotides. Furthermore, nucleic acids are polymers of nucleotides. Thus, nucleic acids and polynucleotides are interchangeable as used herein. Those skilled in the art have general knowledge that nucleic acids are polynucleotides and are hydrolyzable to monomeric "nucleotides." Monomeric nucleotides can be hydrolyzed to nucleosides. As used herein, polynucleotides include, but are not limited to, all nucleic acid sequences obtained by any means available in the art, including, but not limited to, recombinant means, i.e., cloning and synthesis means of recombinant libraries or cell genomes using conventional cloning techniques and PCR®, etc.

[0041] As used herein, the terms “peptide,” “polypeptide,” and “protein” are interchangeable and refer to compounds composed of amino acid residues covalently linked by peptide bonds. A protein or peptide must contain at least two amino acids, and there is no limit to the maximum number of amino acids that can constitute a protein or peptide sequence. A polypeptide includes any peptide or protein containing two or more amino acids linked to each other by peptide bonds. As used herein, this term refers to both short chains, also commonly called peptides, oligopeptides, and oligomers in the art, and long chains, also commonly called proteins in the art, of which many types exist. A “polypeptide” includes, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, polypeptide variants, modified polypeptides, derivatives, analogs, and fusion proteins. Polypeptides include native peptides, recombinant peptides, synthetic peptides, or combinations thereof.

[0042] As used herein, the term "promoter" is defined as a DNA sequence required to initiate the specific transcription of a polynucleotide sequence, which is recognized by the cellular synthetic mechanism or introduced synthetic mechanism.

[0043] As used herein, the term “promoter / regulatory sequence” means a nucleic acid sequence required for the expression of a gene product that is operably linked to a promoter / regulatory sequence. In some examples, this sequence may be a core promoter sequence, and in other examples, this sequence may also include enhancer sequences and other regulatory elements required for the expression of the gene product. A promoter / regulatory sequence may, for example, be a promoter / regulatory sequence that expresses a gene product in a tissue-specific manner.

[0044] A "constitutive" promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or identifying a gene product, causes the cell to produce that gene product under almost all physiological cellular conditions.

[0045] An "inducible" promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or identifying a gene product, causes the cell to produce the gene product only when a corresponding inducer is present within the cell.

[0046] As used herein, the term "SAM file" refers to a type of text file format that contains alignment information for one or more nucleotide or protein sequences mapped to one or more reference sequences. These files may also contain unmapped sequences.

[0047] As used herein, the term "BAM file" refers to a file containing alignment information of various nucleotide or protein sequences mapped to one or more reference sequences, in binary file format. BAM files are small, more efficient in software operation than SAM files, and save time and reduce costs in computer computation and storage.

[0048] The term “subject” is intended to include living organisms (e.g., mammals) capable of eliciting an immune response. “Subject” or “patient” as used herein may be human or non-human mammals. Non-human mammals include, for example, livestock and pets, such as sheep, cattle, pigs, dogs, cats, and mice. Preferably, the subject is human.

[0049] When used herein, the phrases “transcriptionally controlled” or “operatably linked” mean that the promoter is located in the correct position and orientation relative to the polynucleotide to control the initiation of transcription and polynucleotide expression by RNA polymerase.

[0050] A "vector" is a composition of substances containing isolated nucleic acids that can be used for the delivery of isolated nucleic acids into cells. Numerous vectors are known in the art and include, but are not limited to, linear polynucleotides, polynucleotides that associate with ionic or amphiphilic compounds, plasmids, and viruses. Therefore, the term "vector" includes self-replicating plasmids or viruses. Furthermore, the term should be interpreted to include non-plasmid compounds and non-viral compounds that facilitate the transfer of nucleic acids into cells, such as polylysine compounds and liposomes. Examples of viral vectors include, but are not limited to, Sendai virus vectors, adenovirus vectors, adeno-associated virus vectors, retroviral vectors, and lentiviral vectors.

[0051] Scope: Throughout this disclosure, various embodiments of this disclosure may be presented in scope form. It should be understood that the scope form is merely for convenience and brevity and should not be interpreted as an immutable limitation on the scope of this disclosure. Therefore, the scope description should be considered to have as many partial scopes as possible that are disclosed in detail, as well as the individual numbers within those scopes. For example, the scope description, e.g., 1–6, should be considered to have the partial scopes disclosed in detail, e.g., 1–3, 1–4, 1–5, 2–4, 2–6, 3–6, etc., as well as the individual numbers within those scopes, e.g., 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the width of the scope.

[0052] The methods described herein include comprehensive screening of the mutagenic effects on gene activity. In some embodiments, gene activity can be used interchangeably with protein activity. In some embodiments, the mutagenic effects of gene or protein activity are used to generate profiles.

[0053] Furthermore, the methods provided herein include generating a gene mutation library of cDNAs of unique variants. As used herein, a unique variant comprises one or more amino acid substitutions from the wild-type protein. In some embodiments, a unique variant may comprise a single amino acid substitution. In other embodiments, a unique variant may comprise at least two amino acid substitutions. In some embodiments, each unique variant is compared to the wild-type target protein. In some embodiments, the comparison is based on a statistical model that tests specific hypotheses regarding the biological activity of each unique variant. In some embodiments, specific hypotheses regarding the biological activity of each unique variant may include, but are not limited to, loss of function, gain of function, activity substantially similar to the wild-type target protein, drug resistance, or drug sensitivity. In some embodiments, specific hypotheses regarding the biological activity of each unique variant may include loss-of-function variants, gain-of-function variants, unique variants with activity substantially similar to the wild-type target protein, drug-resistant variants, or drug-sensitive variants. In some embodiments, the unique variant is a loss-of-function variant, a gain-of-function variant, a variant with substantially similar activity to the wild-type protein, a drug-resistant variant, or a drug-sensitive variant. In some embodiments, each cDNA contains a single unique variant. In some embodiments, each cDNA in the library encodes a unique variant containing an amino acid substitution in the wild-type target protein, generating a cDNA library. Furthermore, the cDNA library may contain cDNAs encoding unique variants, each containing an amino acid substitution at a different amino acid of the wild-type protein. In some embodiments, the cDNA library may contain mutations at any residue of the wild-type protein. In some examples, a cDNA library is generated, each containing a cDNA with at least one amino acid substitution in the target protein. In some embodiments, the cDNA results in the expression of a target protein containing a mutation at one or more amino acid positions.

[0054] The unique proteins disclosed herein may include wild-type proteins comprising one or more amino acid substitutions. In some embodiments, wild-type proteins may include, but are not limited to, blood coagulation proteins such as factor IX (FIX), factor VIII (FVIII), factor VIIa (FVIIa), von Willebrand factor (VWF), factor FV (FV), factor X (FX), factor XI (FXI), factor XII (FXII), thrombin (FII), protein C, protein S, tPA, PA-1, tissue factor (TF), ADAMTS 13 protease, or fragments thereof. In other embodiments, the wild-type protein may be immunoglobulins, cytokines, e.g., IL-1 alpha, IL-1 beta, IL-2, IL-3, IL-4, IL-5, IL-6, IL-11, colony-stimulating factor 1 (CSF-1), M-CSF, SCF, GM-CSF, granulocyte colony-stimulating factor (G-CSF), EPO, interferon-alpha (IFN-alpha), consensus interferon, IFN-beta, IFN-gamma, IFN-omega, IL-7, IL-8, IL-9, IL-10, IL-12, IL-13, IL-14, IL-15, IL-16, IL-17, IL-18, IL-19, IL-20, IL-21, IL-22, IL-23, IL-24, IL-31, IL-32 alpha, IL-33, thrombopoietin (TPO), angiopoietin, e.g., Ang- 1, Ang-2, Ang-4, Ang-Y, Human angiopoietin-like polypeptide ANGPTL1-7, Vitronectin, Vascular endothelial growth factor (VEGF), Angiogenin, Activin A, Activin B, Activin C, Bone morphogenetic protein-1, Bone morphogenetic protein-2, Bone morphogenetic protein-3, Bone morphogenetic protein-4, Bone morphogenetic protein-5, Bone morphogenetic protein-6, Bone morphogenetic protein-7, Bone morphogenetic protein-8, Bone morphogenetic protein-9, Bone morphogenetic protein-10, Bone morphogenetic protein-11, Bone morphogenetic protein-12, Bone morphogenetic protein-13, Bone morphogenetic protein-14, Bone morphogenetic protein-15, Bone morphogenetic protein receptor IA, Bone morphogenetic protein receptor IB, Bone morphogenetic protein receptor II, Brain-derived neurotrophic factor,Cardiotrophin-1, ciliary neutrophil factor, ciliary neutrophil receptor, crypto, cryptic, cytokine-induced neutrophil chemotactic factor 1, cytokine-induced neutrophil, chemotactic factor 2α, cytokine-induced neutrophil chemotactic factor 2β, β-endothelial growth factor, endothelin-1, epidermal growth factor, epigen, epiregulin, epithelial-derived neutrophil attractant, fibroblast growth factor 4, fibroblast growth factor 5, fibroblast growth factor 6, fibroblast growth factor 7, fibroblast growth factor 8, fibroblast growth factor 8b, fibroblast growth factor 8c, fibroblast growth factor 9, fiber Fibroblast growth factor 10, Fibroblast growth factor 11, Fibroblast growth factor 12, Fibroblast growth factor 13, Fibroblast growth factor 16, Fibroblast growth factor 17, Fibroblast growth factor 19, Fibroblast growth factor 2β, Fibroblast growth factor 21, Acid Fibroblast Growth Factor, Basic Fibroblast Growth Factor, Glial Cell Line-Derived Neurotrophic Factor Receptor α1, Glial Cell Line-Derived Neurotrophic Factor Receptor α2, Growth-Related Protein, Growth-Related Protein α, Growth-Related Protein β, Growth-Related Protein γ, Heparin-Binding Epithelial Growth Factor, Hepatocyte Growth Factor, Hepatocyte Platelet growth factor receptor, hepatocellular carcinoma growth factor, insulin-like growth factor I, insulin-like growth factor receptor, insulin-like growth factor II, insulin-like growth factor binding protein, keratinocyte growth factor, leukemia inhibitor, leukemia inhibitor receptor α, nerve growth factor, nerve growth factor receptor, neuropoietin, neuropoietin-3, neuropoietin-4, oncostatin M (OSM), placental growth factor, placental growth factor 2, platelet-derived endothelial growth factor, platelet-derived growth factor, platelet-derived growth factor A chain, blood Platelet-derived growth factor AA, platelet-derived growth factor AB, platelet-derived growth factor B chain, platelet-derived growth factor BB, platelet-derived growth factor receptor α, platelet-derived growth factor receptor β, pre-B cell growth stimulant, stem cell factor (SCF), stem cell factor receptor, TNF0, TNF1, TNF2 and other TNFs, transforming growth factor a, transforming growth factor β, transforming growth factor β1, transforming growth factor β1.2, transforming growth factor β2, transforming growth factor β3, transforming growth factor β5, latent transforming growth factor β1, transforming growth factor β-binding protein I,Transforming growth factor β-binding protein II, transforming growth factor-binding protein III, thymic stromal lymphocyte activator (TSLP), tumor necrosis factor receptor type I, tumor necrosis factor receptor type II, urokinase-type plasminogen activator receptor, vascular endothelial growth factor, or active fragments thereof, are included but not limited to these. Other non-restrictive examples of wild-type proteins include alpha-interferon, beta-interferon, and gamma-interferon, colony-stimulating factors including granulocyte colony-stimulating factor, fibroblast growth factor, platelet-derived growth factor, phospholipase-activated protein (PUP), insulin, plant proteins such as lectins and lysine, tumor necrosis factor and related alleles, soluble forms of tumor necrosis factor receptor, interleukin receptor and soluble forms of interleukin receptor, growth factors such as tissue growth factors such as TGFα or TGFβ, and epidermal growth factor, hormones, somatomedin, pigment hormones, hypothalamic-release factor, antidiuretic hormone, pro The protein comprises lactin, chorionic gonadotropins, follicle-stimulating hormone, thyroid-stimulating hormone, tissue plasminogen activator, and immunoglobulins, e.g., IgG, IgE, IgM, IgA, and IgD, galactosidase, α-galactosidase, β-galactosidase, DNase, fetuin, luteinizing hormone, estrogen, corticosteroids, insulin, albumin, lipoprotein, fetoprotein, transferrin, thrombopoietin, urokinase, DNase, integrin, thrombin, hematopoietic growth factor, leptin, glycosidase, and fragments thereof, or any fusion protein comprising any of the above proteins or fragments thereof. In some embodiments, the wild-type protein is a protein that binds to Erbb2, an Erbb2 variant, or a fragment thereof. In some cases, the wild-type protein is not a protein that binds to Erbb2, an Erbb2 variant, or a fragment thereof.

[0055] Furthermore, engineered human or mammalian cell lines encoding receptors operably linked to reporter systems, and optionally, reporters expressed on the surface, are provided herein. In some cases, the receptor and reporter system can be directed to whether one of the unique variants binds to the receptor and modulates its biological activity. In some embodiments, the receptor is a surface-expressed receptor. The term “surface-expressed receptor” refers to cell surface receptors (membrane receptors, transmembrane receptors) embedded in the cell’s plasma membrane. For example, surface-expressed receptors include cytokine receptors, chemokine receptors, interferon receptors, 5T4, A33, activin receptors, adrenomedullin receptors, AFP, AGS-5, ALK, annexin, AXL, B7-H3, B7-H4, BAGE protein, BCMA, bombesin, C33 antigen, C4.4a, type C lectin-like (receptor), CA19.9, CA-125, CADM1, CAIX, CanAg, CAR, carbonic anhydrase, caveolin-1, CCK2R, CD4, CD10, CD19, CD20, CD21, CD22, CD25, CD27, CD30, CD33, CD37, CD38, CD44, CD51, CD57, CD70, CD73, CD74, CD79a, CD79b, CD80, CEA, CEACAM, c-kit, claudin, chemokine receptor (i.e., CXCR4, CXCRS), c-Met, Cripto-1, DEC-205, Derlin-1, desmoglein-3, Dlk-1, DLL3, DS6, E-cadherin, E-ce Lectin, EAG-1, ED-B, EpCAM, EGFR, EGFRvIII, emmprin, endothelin receptor, ErbB2 / Her2, ErbB3, ErbB4, ETV6-AML, ephrin type A receptor, epiregulin, ETA, FAP-alpha, FcyR, FGFR, FOLR1, Frizzled, Fyn3, galectin, ganglioside, GCC, GD2, GD3, GloboH, glypican-3 (lypican-3), GLUTS, GPNMB, G protein-coupled receptor (i.e., GPR49), gp100, Hsp, HLA / B-raf, HLA-DR, HLA / k-ras, HLA MAG E-A3, HMW-MAA, hTERT, ICAM-3, IGF-R, IL-13-R, L1CAM, laminin receptor, LIV1, LMP2, LRP5, LRP6, MAGE protein, MART-1, melanotransferrin, mesothelin, metalloproteinase, ML-IAP, mucin, Mud, Mud 6 (CA-125), MU M1, N-cadherin, NA17, NCAM-1, nectin 4, Notch, NP-55, NRP1, NY-BR1, NY-BR62, NY-BR85, NY-ES01, PLAC1, PRLR, PRAME, prominin-1, PSMA (FOLH 1) May contain RON, SLC44A4, SLITRK6, Steap-1, Steap-2, surviving, syndecan, TAG-72, TF, TGF-p, TMPRSS2, TMEFF2, TNFR, Tn, TROP2, TRP-1, TRP-2, TWEAKR, tyrosinase, uroplakin-3, and VEGFR.

[0056] As used herein, the term “library” refers to a pool of at least two polynucleotides, cell clones, molecules, or proteins. In certain embodiments, the library is used to screen cDNA encoding unique variants. In some embodiments, the library is a plasmid cDNA library. In other embodiments, the library is a viral library. For example, a viral library may include a lentiviral library, an adenovirus library, or a retrovirus library. In other embodiments, a plasmid cDNA library of cDNA is converted into a viral library. In other embodiments, a drug library is mixed with single cell clones, and then a biological assay is used to identify which drugs produce a response in the cell clones. A plasmid cDNA library may contain 100, 1,000, 10,000, 100,000, or 1,000,000 different cDNA molecules. In other embodiments, the plasmid cDNA library contains 100, 1,000, 10,000, 100,000, or 1,000,000 cells that have been modified to express cDNA further containing a unique molecular identifier (UMI) or a barcode group, or cDNA having a UMI or barcode as RNA.

[0057] In some embodiments, each cDNA in the cDNA library is incorporated into a single plasmid such that each cDNA in the cDNA library is individually and independently incorporated into its own plasmid, forming a plasmid cDNA library. In some embodiments, each plasmid in the plasmid cDNA library contains only one unique cDNA out of several cDNAs, and each plasmid in the plasmid cDNA library further includes a unique molecular identifier (UMI) or barcode group. In some embodiments, the cDNA library is packaged into a viral vector to generate a viral library. In some embodiments, the viral library contains virions. In some embodiments, each virion contains only one viral plasmid from the cDNA library. In some embodiments, the viral library is a retroviral library, an adenovirus library, or a lentivirus library. In some embodiments, the viral library is a lentivirus library. In some embodiments, the lentivirus vector contains an expression cassette. In some embodiments, the expression cassette may contain a promoter operably ligated to a polynucleotide sequence encoding a unique variant. In some embodiments, the promoter operably ligated to the polynucleotide encoding the unique variant is inducible. In some embodiments, target cells are transduced with a lentiviral pool from a cDNA library at a low MOI such that >90% of the target cells express a single unique variant. In other embodiments, target cells are transduced with a lentiviral pool from a lentiviral library at a high MOI such that the target cells express a unique variant. The terms “MOI” or “infection multiplicity” are used according to their simple general meaning in virology and refer to the ratio of infectious material (e.g., virus) to a target (e.g., cells) in a given area or volume. In embodiments, the area or volume is assumed to be homogeneous.In some embodiments, the MOI is at least 0.001, at least 0.01, at least 0.01, at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or at least 1.

[0058] This specification presents a method for preparing a plurality of plasmid cDNA libraries from a plurality of cDNAs, comprising the step of incorporating tags into cDNAs to yield a plurality of tagged cDNA samples, wherein the cDNA in each tagged cDNA sample codes for a unique variant. In one embodiment, the tag comprises a cell-specific identifier sequence and a unique molecular identifier (UMI) sequence. In some embodiments, the tag comprises a cell-specific identifier sequence that does not include a UMI. In some embodiments, the tag comprises a barcode. The method further comprises the steps of pooling tagged cDNA samples and, optionally, generating a plurality of tagged cDNAs by amplifying the pooled cDNA samples to produce a cDNA library containing double-stranded cDNAs. In some embodiments, the cDNAs of the plurality of cDNAs further comprise a UMI or barcode group. In some embodiments, the UMI or barcode group can be immobilized on a solid support. For example, the solid support may be one or more beads. Thus, in a particular embodiment, a plurality of beads may be presented, each bead having a unique sample barcode and / or UMI sequence. In some embodiments, each cDNA in the cDNA library is identified by contacting it with one or more beads having a unique set of sample barcodes and / or UMI sequences. In some embodiments, purified nucleic acids derived from transduced cells are identified by contacting them with one or more beads having a unique set of sample barcodes and / or UMI sequences.

[0059] As used herein, the terms “transfected,” “transformed,” or “transduced” refer to the process of transferring or introducing an exogenous nucleic acid into a host cell. A “transfected,” “transformed,” or “transduced” cell is a cell that has been transfected, transformed, or transduced with cDNA encoding a unique variant. The cells include primary target cells and their offspring. In some embodiments, a library of transduced cells may contain about 100, about 1,000, about 10,000, about 100,000, about 500,000, about 1 million, about 5 million, or about 10 million transduced cells. In some embodiments, a library of transduced cells may contain about 500,000 to about 10 million cells. In some embodiments, the transduced cell library contains approximately 10,000 cells, 20,000 cells, 30,000 cells, 40,000 cells, 50,000 cells, 60,000 cells, 70,000 cells, 80,000 cells, 90,000 cells, 100,000 cells, 150,000 cells, 200,000 cells, 300,000 cells, 400,000 cells, and 500,000 cells. The cells include 600,000 cells, 700,000 cells, 800,000 cells, 900,000 cells, 1,000,000 cells, 2,000,000 cells, 3,000,000 cells, 4,000,000 cells, 5,000,000 cells, 6,000,000 cells, 7,000,000 cells, 8,000,000 cells, 9,000,000 cells, or 10,000,000 cells. In some embodiments, cells in the transduced cell library are isolated and purified. In some embodiments, the isolated and purified transdextrin library contains approximately 10,000 to 10 million transdextrin cells, and each isolated and purified transdextrin contains a plasmid containing cDNA encoding a unique variant of the wild-type protein, so that each transdextrin further contains a cell surface receptor operably coupled to a reporter system, and the cell surface receptor can be examined for each unique variant.In some cases, each isolated and purified transduced cell further contains unique variants on its surface.

[0060] In some cases, a high-throughput, optionally computerized or robotically implemented system is provided for identifying proteins. In such embodiments, the method provided herein may include a library of plasmid cDNA and transduced cells arranged in a number of compartments. With respect to plasmids, the library contained in the compartments may include plasmids for expressing cDNA encoding a unique variant disclosed herein. Such a library can be used very efficiently to transduce cells to generate a library of cells in a number of compartments, each containing a cell transduced with one vector. The library may, if necessary, be grown in packaging cells before use in cell transduction.

[0061] Libraries of transduced cells can be analyzed for the action of unique variants using machine-implemented microarray or macroarray techniques. For example, machine-implemented techniques can be used to determine large amounts of sequencing using a single "chip" for hybridization of mRNA or corresponding cDNA encoding unique variants isolated from cells. Furthermore, libraries of transduced cells may be subjected to further processing or altered conditions before analysis of their action on cellular factors. Cells and, therefore, their action on cellular factors can also be analyzed temporally. Gene or protein function can be evaluated through cell differentiation and in vivo function during culture, or after transplantation into animal models or humans or non-human primates.

[0062] Furthermore, in some embodiments, the cells used in cell-based assays may contain nucleic acids encoding reporter molecules (e.g., reporter proteins) that operatively ligate to promoters and / or enhancers responsive to the activity of a unique variant. “Promoter” refers to a regulatory region of a nucleic acid that controls the initiation and rate of the rest of the transcription of a nucleic acid sequence. Promoters are typically located at or near the transcription start site of a gene to drive the transcription of the nucleic acid sequence being regulated. In some embodiments, promoters are 100 to 1000 nucleotides long. Promoters may also contain small regions to which regulatory proteins and other molecules, such as RNA polymerase and other transcription factors, can bind. Promoters can be constitutive (e.g., CAG promoter, cytomegalovirus (CMV) promoter), inducible (also called activatable), repressive, tissue-specific, developmental stage-specific, or any combination of two or more of the aforementioned promoters. A promoter is considered “operatively ligated” when it is present in the correct functional location and orientation relative to the nucleic acid sequence being regulated (e.g., to control ("drive") the transcription initiation and / or expression of its sequence). In some embodiments, the promoter may be obtained by isolating a 5' non-coding sequence(s) that is naturally associated with the nucleic acid and located upstream of the coding region of a given nucleic acid. Such promoters are also called “endogenous” promoters.

[0063] In some embodiments, the promoter may be a constitutively active promoter (i.e., a promoter that is constitutively active / "on"), an inductive promoter (i.e., a promoter whose active / "on" or inactive / "off" state is controlled by an external stimulus, e.g., a specific temperature, compound, or protein), a spatially restricted promoter (i.e., a transcriptional regulatory element, enhancer, etc.) (e.g., a tissue-specific promoter, a cell-type-specific promoter, etc.), or a temporally restricted promoter. The suitable promoter may be derived from a virus and therefore may also be called a viral promoter, or it may be derived from any organism, including prokaryotes or eukaryotes. The suitable promoter can be used to drive expression by any RNA polymerase (e.g., pol I, pol II, pol III). Examples of promoters include, but are not limited to, the SV40 early promoter, the mouse mammary tumor virus terminal repeat (LTR) promoter, the adenovirus major late promoter (Ad MLP), the herpes simplex virus (HSV) promoter, the cytomegalovirus (CMV) promoter, e.g., the CMV very early promoter region (CMVIE), the Roussarcoma virus (RSV) promoter, the human U6 intranuclear small molecule promoter (U6), and the human H1 promoter (H1). Examples of inducible promoters include, but are not limited to, the T7 RNA polymerase promoter, the T3 RNA polymerase promoter, the isopropyl-beta-D-thiogalactopyranoside (IPTG) regulatory promoter, the lactose inducible promoter, the heat shock promoter, the tetracycline regulatory promoter (e.g., Tet-ON, Tet-OFF, etc.), the steroid regulatory promoter, the metal regulatory promoter, and the estrogen receptor regulatory promoter. Therefore, the inducible promoter can be regulated by molecules that non-restrictively include doxycycline; RNA polymerase, e.g., T7 RNA polymerase; estrogen receptor; estrogen receptor fusion, etc. In some embodiments, the promoter is a doxycycline-inducible promoter.

[0064] In some embodiments, the promoter is a spatially restricted promoter (i.e., a cell type-specific promoter, a tissue-specific promoter, etc.) such that the promoter is active (i.e., "on") in a specific subset of cells of a multicellular organism. Spatially restricted promoters may also be called enhancers, transcriptional regulatory elements, regulatory sequences, etc. Any convenient spatially restricted promoter may be used, and the selection of a suitable promoter (e.g., a brain-specific promoter, a promoter that drives expression in a subset of neurons, a promoter that drives expression in the germline, a promoter that drives expression in the lungs, a promoter that drives expression in muscle, a promoter that drives expression in pancreatic islet cells, etc.) is organism-dependent.

[0065] In some embodiments, the promoter and / or enhancer responsive to the activity of a unique variant is an expression regulatory sequence. Expression vectors and cloning vectors typically include a promoter that is recognized by the host organism and operably linked to a nucleic acid encoding a polypeptide (e.g., a reporter polypeptide). Preferably, the expression regulatory sequence is a eukaryotic cell promoter system within the vector that can transform or transfect eukaryotic host cells. Once the vector is incorporated into a suitable host, the host is maintained under conditions suitable for high levels of expression of the nucleotide sequence after T cell activation. In some embodiments, the method provides cells of a human or mammalian cell line containing a nucleic acid encoding a reporter molecule under the control of a promoter responsive to activation by a unique variant. In some embodiments, the expression reporter vector for use in eukaryotic host cells (nucleated cells derived from yeast, fungi, insects, plants, animals, humans, or other multicellular organisms) also includes sequences necessary for transcription termination and mRNA stabilization. Such sequences are generally available from the 5', sometimes 3', untranslated region of eukaryotic or viral DNA or cDNA. One useful transcription termination region is the polyadenylated region of bovine growth hormone.

[0066] In some embodiments, the method provides a vector for the expression of a unique variant in cells of a human or mammalian cell line. The vector components generally include, but are not limited to, one or more signal sequences, origins of replication, one or more marker genes, multiple cloning sites including recognition sequences for numerous restriction endonucleases, enhancer elements, promoters (e.g., enhancer elements and / or promoters responsive to activation), and transcription termination sequences. In some embodiments, the vector is a plasmid. In other embodiments, the vector is a recombinant viral genome, e.g., a recombinant lentiviral genome, a recombinant retroviral genome, or a recombinant adeno-associated virus genome. The vector, containing a cDNA library, is transduced into cells of a human or mammalian cell line.

[0067] In some embodiments, a cell-based assay is provided that detects the activity of a unique variant by contacting a population of cells with a reporter assay system responsive to the unique variant. In some embodiments, the cells are mammalian cells. Mammalian cells may include, but are not limited to, mouse, rat, hamster, or human cells. In some embodiments, the cells are human cells. In some embodiments, the cells exhibit increased biological activity. In some embodiments, the population of cells is a population of immortalized cells (e.g., an immortalized cell line).

[0068] In some embodiments, cells introduced with a cDNA library are screened for receptor activation. For example, stable clones may be isolated by limiting dilution and screened for their binding response to the receptor. In some embodiments, stable reporter T cells are screened with receptors greater than approximately 1 μg / mL, 2 μg / mL, 3 μg / mL, 4 μg / mL, 5 μg / mL, 6 μg / mL, 7 μg / mL, 8 μg / mL, 9 μg / mL, or 10 μg / mL. In some embodiments, the receptor may be a surface-based receptor, a costimulatory molecule, an accessory molecule, an immune checkpoint molecule, a member of the TNF family or TNF family receptors, a cytokine receptor, a chemokine receptor, or an adhesion molecule.

[0069] This specification provides a method for screening large quantities of compounds for activity related to specific biological functions, requiring the preparation of cell arrays for parallel processing of cells and reagents. A standard 96-well microtiter plate, measuring 86 mm × 129 mm and having 6 mm diameter wells at a 9 mm pitch, is used to accommodate current automated loading and robotic processing systems. The microplate is typically 20 mm × 30 mm and arranges cells at a pitch of approximately 500 microns with dimensions of 100–200 microns. The microplate may consist of a coplanar layer of cell-adhering material, a coplanar layer patterned with non-cell-adhering material, or a coplanar layer etched onto a three-dimensional surface of similarly patterned material. For the purposes of the following discussion, the terms “well” and “microwell” refer to locations within any structural array where cells adhere and are imaged. The microplate may also include fluid delivery channels in the spaces between wells. Smaller microplate formats improve the overall efficiency of the system by minimizing the amount of reagents prepared, stored, and processed, as well as the overall operation required for scanning. In addition, the entire surface area of ​​the microplate can be imaged more efficiently, enabling a second operating mode for the microplate reader, which will be described later in this document.

[0070] This specification provides a method for evaluating the biological activity of unique variants using fluorescent and luminescent reagents to measure the temporal and spatial distribution, content, and activity of intracellular ions, metabolites, macromolecules, and organelles. Classes of these reagents include labeling reagents that measure the distribution and number of molecules in living and fixed cells, environmental indicators that report signaling events in time and space, and fluorescent protein biosensors that measure the activity of target molecules in living cells.

[0071] The methods of this disclosure are based on the high affinity of fluorescent or luminescent molecules to specific cellular components. Affinity to specific components is governed by physical forces, such as ionic interactions, covalent bonds (including chimeric fusions with protein-based chromophores, fluorophores, and lumiphores), as well as hydrophobic interactions, electrical potential, and in some cases, simple capture within the cellular component. The luminescent probe may be a small molecule, a labeled polymer, or a genetically modified protein, and may include, but is not limited to, a green fluorescent protein chimera.

[0072] As used herein, the methods described include reporter assays. “Reporter assay,” as used herein, refers to an analytical method that enables the biological characterization of a stimulus by monitoring the induction of reporter expression in cells. Stimuli trigger the induction of intracellular signaling pathways, resulting in cellular responses that typically include the modulation of gene transcription.

[0073] In some cases, stimulation of cellular signaling pathways results in modulation of gene expression through the regulation of transcription factors and recruitment to non-coding regions upstream of DNA, which are necessary to initiate RNA transcription and trigger protein production. The regulation of gene transcription and translation in response to stimuli is required to induce most biological responses, such as cell proliferation, differentiation, survival, and immune responses. These non-coding regions of DNA, also called enhancers, contain specific sequences that are recognition elements of transcription factors, which regulate the efficiency of gene transcription and, therefore, the amount and type of proteins produced by cells in response to stimuli. In a reporter assay, a responsive enhancer element and minimal promoter are manipulated using standard molecular biological methods to drive the expression of a reporter gene. The DNA is then transfected into cells containing all mechanisms that respond specifically to the stimulus, and the level of transcription, translation, or activity of the reporter gene is measured as a surrogate measure of the biological response.

[0074] In some embodiments, the disclosed method includes a cell-based assay in which a receptor is operably linked to a reporter system. In some embodiments, the receptor is a surface-expressed receptor. In some embodiments, the reporter system includes a reporter molecule. The reporter molecule can be any molecule for which an assay can be developed to measure the amount of a molecule produced by a cell in response to a stimulus. For example, the reporter molecule may be a reporter protein encoded by a reporter gene that is responsive to a stimulus. Commonly used examples of reporter molecules include, but are not limited to, luminescent proteins, such as luciferases, which emit light as a byproduct of the catalytic action of an experimentally measurable substrate. Luciferases are a class of bioluminescent proteins derived from many sources, including firefly luciferase (from Photinus pyralis), sea pansy luciferase (Renilla reniformis), click beetle luciferase (from Pyrearinus termitilluminans), marine copepod gaussia luciferase (from Gaussia princeps), and deep-sea shrimp nanoluciferase (from Oplophorus gracilirostris). Firefly luciferase catalyzes the oxygenation of luciferin to oxyluciferin, resulting in the emission of photons, while other luciferases, such as sea pansy luciferase, emit light by catalyzing coelenterazine. The wavelengths of light emitted by various luciferase types and variants can be read using various filter systems that facilitate multiplexing. Since the amount of luminescence is proportional to the amount of luciferase expressed in the cell, the luciferase gene was used as a sensitivity reporter to evaluate the effects of stimuli that induce biological responses.

[0075] In some embodiments, the method provides a cell-based assay for detecting the unique variants disclosed herein, wherein cells of a human or mammalian cell line encode a reporter construct responsive to the biological activation of the unique variant. In some embodiments, the reporter construct comprises a luciferase. In some embodiments, the luciferase is a firefly luciferase (e.g., from Photinus pyralis), a sea pansy luciferase (e.g., from Renilla reniformis), a click beetle luciferase (e.g., from Pyrearinus termitilluminans), a marine copepod gaussia luciferase (e.g., from Gaussia princeps), and a deep-sea shrimp nanoluciferase (e.g., from Oplophorus gracilirostris). In some embodiments, the expression of the luciferase in cells of a human or mammalian cell line directs the binding activity between the unique variant and the reporter protein. In some examples, the reporter construct encodes β-glucuronidase (GUS); fluorescent proteins, such as green fluorescent protein (GFP), red fluorescent protein (RFP), blue fluorescent protein (BFP), yellow fluorescent protein (YFP), and their variants; chloramphenicol acetyltransferase (CAT); β-galactosidase; β-lactamase; or secreted alkaline phosphatase (SEAP).

[0076] In some embodiments, the method further provides a sequencing method for evaluating the mutagenic effect on gene or protein activity. In some embodiments, the sequencing method is next-generation sequencing, Sanger sequencing, whole-genome sequencing, RNA sequencing, or shotgun sequencing. In some embodiments, the sequencing method is next-generation sequencing. In some cases, next-generation sequencing produces a sequence dataset. The sequence data may include a read depth of approximately 2000 × to approximately 90,000 × sequencing coverage for each unique variant. In some embodiments, the read depth for each unique variant is at least approximately 1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10000x, 50,000x, or 100,000x. This is performed for several sequencing reactions. Sequencing reactions of fewer than 50,000 or fewer than 100,000 reads may be performed. Example read depths are approximately 1,000 to 50,000, 2,000 to 90,000, or 5,000 to 100,000 reads per locus (base position). In some embodiments, the method further includes a step of validating the method by comparing the biological activity of a subset of unique variants analyzed by the method with the previously determined activity of the wild-type target protein. In other embodiments, the method further includes a step of validating the method by comparing the true negatives determined by the method with true negatives determined by an independent method. In some cases, the method further includes a step of comparing the biological activity of a subset of unique variants analyzed by the method with independent testing of a separate set of clones containing the unique variant(s). In some embodiments, the method is further validated by comparing the results of the method across one or more different samples. In some cases, the method is further validated by comparing the results in two different cell lines.

[0077] Furthermore, the Specified Method includes a step of generating an oligonucleotide sequence dataset. In some embodiments, the Method further includes a step of analyzing the oligonucleotide sequence dataset or the analysis using a computer. In some examples, the step of analysis or interpretation includes (a) generating a unique variant-UMI index library using long reads derived from next-generation sequencing (NGS); (b) identifying UMI counts and barcode bins by analyzing short reads of a flow-selected group; (c) calculating an activity score from the UMI-barcode read counts, including the UMI counts and barcode bins of (b); (d) evaluating the effect of the unique variant(s) on biological activity by calculating the read percentages and p-values ​​of a reporter protein, e.g., a fluorescent protein; and (e) accurately and quantitatively evaluating the activity of the unique variant(s) relative to a previously characterized wild type. In some embodiments, the fluorescent protein is GFP, and the biological activity includes gene activity or protein activity. [Examples]

[0078] (Example 1) Genesis and analysis of genetic mutation libraries (GMLs) To generate a comprehensive gene activity mutation (MEGA)-mutation activity profile (Map), GMLs are analyzed using the assay system described.

[0079] Creation of a Genetic Mutation Library (GML) We designed a library of mutant cDNA molecules against target genes. This library includes the promoter, coding region, 3'UTR, minigene, and introns or any of the above parts containing splice sites.

[0080] UMI-Barcoded Variant Plasmid Library Generation Doxycycline-inducible lentiviral plasmids were constructed by PCR amplification of the TreTIGHT doxycycline-inducible promoter derived from pSSI9343 (donated by Sierra Sciences, Reno, NV) using oHI-00176 (with a 5' overhang complementary to the upstream sequence of the SnaBI restriction site in pLJM1_MCS) and oHI-00177 (with a 5' overhang complementary to the downstream sequence of BmtI in pLJM1_MCS). These plasmids were then gel-purified using the Nucleospin Gel and PCR Cleanup Kit (Macherey-Nagel, Allentown, PA). The PCR products were then mixed with SnaBI-BmtI linearized pLJM1_MCS (2:1 insert:vector molar ratio) and combined in a HiFi DNA Assembly reaction mixture containing NEBuilder HiFi DNA Assembly Master Mix (New England Biolabs, Ipswich, PA). The reaction mixture was incubated at 50°C for 15 minutes to generate a final lentiviral plasmid with multiple cloning sites downstream of the doxycycline-inducible TreTIGHT promoter.

[0081] A double-stranded (ds)DNA library containing codon-optimized full-length Erbb2 cDNA with sequences of any possible single-amino acid variants within the near-membrane domain and tyrosine kinase domain was synthesized by Twist Bioscience (San Francisco, CA). ds-DNA from each well of a 96-well plate was pooled, and the cDNA was purified using the Nucleospin Gel and PCR Cleanup Kit (Macherey-Nagel, Allentown, PA). The synthesized cDNA contains a 5' overhang sequence (BmtI region in lowercase; (5'-GGTTTAGTGAACCGTCAGATCCgctagc-3')) upstream of a Kozak-containing Erbb2 ATG start codon complementary to the upstream sequence of the BmtI region of lentiviral plasmid pHI-00104, and a 3' overhang sequence (5'-TCGATCCCGTACCGAGGAGATCTG-3') downstream of a HER2 stop codon complementary to the 3' end of the oligo (oHI-00165) containing a 32-nucleotide randomized DNA sequence. The randomized DNA sequence of oHI-00165 is cloned into the 3' untranslated region (UTR) of the final plasmid library so that all cDNA molecules are barcoded with a unique molecular identifier (UMI).

[0082] The 5' end of oHI-00165 contains an overhang sequence complementary to the downstream sequence of the MluI site of pHI-00104 (MluI site in lowercase: 5'-ATTTGTCTCGAGGTCGATTCGAATacgcgt-3'). The purified ds-DNA mutant library fragment, UMI oligo oHI-00165, and the pHI-00104 lentiviral plasmid backbone digested with BmtI-MluI were mixed in an insert:oligo:vector molar ratio of 2:5:1 and used in a HiFi DNA assembly reaction scaled up to 5× with NEBuilder HiFi DNA Assembly Master Mix (New England Biolabs, Ipswich, PA), and incubated at 50°C for 1 hour. The assembled reaction product was cleaned up using the Monarch PCR & DNA Cleanup Kit (New England Biolabs, Ipswich, PA) and dialyzed by drip dialysis onto a 0.025 μm pore size MF-Millipore Membrane filter (Millipore Sigma, #VSWPO2500, Burlington, MA).

[0083] Endura electrocompetent cells (Lucigen, Middleton, WI) were electroporated with purified and dialyzed assembly reaction mixtures, plated onto pre-warmed LB ampicillin plates, and incubated at 37°C for 16 hours. Transformants were scraped, and plasmid libraries derived from the pooled cell suspensions were isolated using the EndoFree Plasmid Mega kit (Qiagen, #12381, Germantown, MD). PacBio long-read sequencing was performed at the DNA sequencing center at Brigham Young University, Salt Lake City, UT, to sequence the full-length Erbb2 cDNA and UMI region of the GML. Using PCR (15 cycles), Q5 High Fidelity DNA Polymerase (New England Biolabs, Ipswich, PA) was used with PCR primers oHI-20-00030 (targeting the 5' untranslated region upstream of the ATG start codon in the Erbb2 gene) and oHI-00203 (targeting the 3' untranslated region downstream of the UMI site) to generate sequence datasets containing DNA fragments for long-read sequencing. Alternatively, a non-PCR-based method was used for long-read sequencing by gel purification of DNA fragments obtained from digestion of a gene mutation library with NheI-MluI restriction enzymes.

[0084] Construction of individual Erbb2 variant alleles for control Plasmid pcDNA3.1+ / C-(K)-DYK-Erbb2, containing the full-length wild-type Erbb2, was PCR amplified using Q5 High Fidelity DNA Polymerase (New England Biolabs, Ipswich, PA) with PCR primers (having a 5' overhang containing the NheI / BmtI restriction site and Kosack site upstream of the ATG start codon) and oHI-20-00009 (having a 5' overhang containing the SalI restriction site downstream of the stop codon to replace the C-terminal DYK tag). The PCR product was digested with DpnI and the A-tail was added by incubation with Taq DNA polymerase (New England Biolabs, Ipswich, PA) at 72°C for 20 minutes. The PCR product was TOPO cloned into a pCR4-TOPO vector using the TOPO TA Cloning Kit (Thermo Fisher Scientific, Hampton, NH) to create pHI-00062. Erbb2 was subcloned from pHI-00062 to the BmtI-SalI site of the intermediate plasmid pHI-001 to create pHI-00065, and then subcloned again from pHI-00065 to the BmtI-AsiSI site of the lentiviral plasmid pLJM1_MCS to create pHI-00076, in which Erbb2 is expressed by the CMV promoter.

[0085] TreTIGHT doxycycline-inducible promoters were obtained by PCR amplification of pSSI9343 (donated by Sierra Sciences, Reno, NV) using oHI-00176 (having a 5' overhang complementary to the upstream sequence of the SnaBI restriction site of pHI-00076) and oHI-00178 (having a 5' overhang complementary to the downstream sequence of BmtI of pHI-00076), and gel-purified using the Nucleospin Gel and PCR Cleanup Kit (Macherey-Nagel, Allentown, PA). The PCR products were then mixed with SnaBI-BmtI linearized pHI-00076 in a 2:1 insert:vector molar ratio in a HiFi DNA Assembly reaction mixture containing NEBuilder HiFi DNA Assembly Master Mix (New England Biolabs, Ipswich, PA), and incubated at 50°C for 15 minutes. This process generated a final lentiviral plasmid expressing wild-type Errb2 mRNA upon induction of a doxycycline-inducible TreTIGHT promoter.

[0086] To generate a YVMA Erbb2 GOF control mutant, we used forward PCR primer oHI-20-00028 (adjacent to AatII of HER2), and downstream PCR primer oHI-20-00029 containing a 12-nt YVMA indel (underlined) and a silent mutation (lowercase font) that disrupts the NdeI site of Erbb2. [ka] pHI-00062 was PCR amplified using [specified method]. The PCR product was gel-purified using the Nucleospin Gel and PCR Cleanup Kit. The PCR product was then mixed with AatII-NdeI linearized pHI-00062 in a 2:1 insert:vector molar ratio and mixed in a HiFi DNA Assembly reaction mixture containing NEBuilder HiFi DNA Assembly Master Mix (New England Biolabs, #E2621L, Ipswich, PA), and incubated at 50°C for 15 minutes to generate pHI-00067. The YVMA HER2 mutant was subcloned from pHI-00067 to the BmtI-SalI site of the intermediate plasmid pHI-00060 to create pHI-00068. Then, pHI-00068 was subcloned again to the BmtI-AsiSI site of the lentiviral plasmid pLJM1_MCS to create pHI-00077, in which YVMA HER2 is expressed by the CMV promoter.

[0087] Similar to the method described above for wild-type Erbb2, the TreTIGHT promoter was cloned into pHI-00077 to create a final lentiviral plasmid expressing YVMA HER2 by induction of the doxycycline-inducible TreTIGHT promoter. Codon-optimized Erbb2 was isolated as a clone from the GML, and the entire Erbb2 coding sequence was sequenced using Sanger sequencing and named pHI-00163. Codon-optimized wild-type Erbb2 was PCR amplified using Q5 High Fidelity DNA Polymerase (New England Biolabs, Ipswich, PA) with PCR primers oHI-00414 and oHI-00415, and then TOPO cloned into pCR-Blunt-II-TOPO using the Zero Blunt TOPO PCR Cloning Kit (Thermo Fisher Scientific, Hampton, NH). Sequence determination and confirmation by Sanger sequencing created pHI-00196.

[0088] All other HER2 controls were prepared using the QuikChange Lightning Site-Directed Mutagenesis Kit (Agilent, Santa Clara, CA) with codon-optimized wild-type Erbb2 as the template. Primers used for mutagenesis were designed using Agilent's QuikChange Primer Design tool. Mutagenesis was performed using a thermocycler with the following PCR protocol: initial denaturation at 95°C for 2 minutes; 18 cycles of denaturation at 95°C for 2 minutes, annealing at 60°C for 10 seconds, extension at 68°C for 3 minutes 45 seconds; final extension at 68°C for 5 minutes; and termination by holding at 4°C. Transformants were screened and Sanger sequencing was performed on the mutagenesis sites. Mutant Erbb2 cDNA was subcloned into the NheI-SalI site of plasmid pHI-00104 to create lentiviral plasmids expressing the mutant Erbb2 control using the doxycycline-inducible TreTIGHT promoter.

[0089] Deep sequencing of UMI-barcoded variant libraries A variant library for the coding region was subdivided into overlapping tiles, each containing single-stranded DNA of less than 300 nucleotides, which was then synthesized as a single-stranded cDNA. The library's variant collection and each synthesized cDNA were designed using a custom program.

[0090] The cDNAs in each pool were joined using overlap sequences. HiFi DNA Assembly is a reaction used to join DNA fragments with overlap sequences. During the assembly reaction, T5 exonuclease partially digests the 5' end of the DNA fragment, leaving the overlap sequence exposed to allow the complementary strand to anneal. HiFi DNA polymerase fills the gap by synthesizing DNA from 5' to 3', and DNA ligase closures the nick. This activity continues throughout the reaction period, resulting in a highly efficient method of assembling DNA. Furthermore, this reaction can be carried out using ssDNA oligos as bridges between the two exposed DNA ends, provided the oligo has an overlap DNA sequence of 25-30 bp with each end of the DNA fragment.

[0091] The ability to assemble ssDNA oligos into dsDNA fragments presupposes the use of HiFi DNA Assembly for GML generation. ssDNA oligos can take the form of synthetic oligopools, each containing thousands of different oligos with an oligo length of approximately 200 nucleotides. The oligopools are designed so that each oligo has a 25-nucleotide overlap sequence at its 5' and 3' ends, corresponding to 25 bp at the 5' and 3' ends of the dsDNA fragment from which the oligo is cloned. This allows for the introduction of desired mutations into target DNA sequences, promoter sequences, etc., corresponding to protein-coding sequences, via a 150-nucleotide (corresponding to a 50-amino acid variant region) within the oligo's center. This approach does not require PCR amplification of the oligopools. Preparation of vector DNA for variant assembly involves inverse PCR, where the entire vector sequence excluding the variant region is amplified using PCR primers, resulting in a 25 bp overlap corresponding to each oligopool.

[0092] Furthermore, the entire code sequence can be covered by dividing it into overlapping sections (or tiles) corresponding to the desired oligo lengths. Each tiled section corresponds to a separate oligo pool.

[0093] Once the variant assembly of the GML is complete, separate UMI assembly reactions can be performed on the purified variant assembly plasmid library. Therefore, the plasmid library can be linearized using restriction enzymes, and a UMI oligo with a 25-nucleotide overlap sequence for assembly can be cloned into the UMI site (e.g., 3'UTR).

[0094] Generation of a lentivirus GML library Lenti-X 293T cells were plated in 2 × 150 mm plates (12.5E6 cells / plate) and allowed to adhere for 24 hours. Approximately 40 cells were displayed per clone. Each plate was transfected with 10 μg of plasmid library (approximately 1.38E6 copies of plasmid DNA per clone), 9.5 μg of psPAX2, and 5.2 μg of pMD2.G, using Fugene HD as the transfection reagent in a ratio of Fugene 3 μL:DNA 1 μg. The culture medium was changed on the plate 24 hours after transfection. The viral supernatant was collected approximately 72 hours after transfection, filtered through a 0.45 μM filter, and flash-frozen in aliquots. The viral aliquots were thawed on ice, and viral titers were established using Lenti-X GOSTIX (Takara). Further details are listed below.

[0095] Counting and seeding of cells for transfection LentiX293T cells were grown in 10 ml of complete DMEM medium (DMEM + 10% fetal bovine serum) in a 100 mm petri dish and incubated in a CO2 incubator at 37°C with a 5% CO2 atmosphere. The cells were trypsinized, and the cell pellet was resuspended in 3 ml of complete DMEM medium. 3 million cells were seeded into a new petri dish containing 10 ml of complete DMEM medium. The petri dish was labeled with the cell passage number, passage date, and the name of the culturist. The cells were cultured in a CO2 incubator at 37°C with a 5% CO2 atmosphere until they reached 80-90% confluence. 4 million cells from the remaining cells were seeded into a 100 mm petri dish containing 10 ml of complete DMEM medium.

[0096] Simultaneous transfection and generation of lentiviral vectors For each expression construct, 0.58E6 Lenti-X 293T cells were plated into 1× wells of a 6-well plate and allowed to adhere for 24 hours. Each plate was transfected with 1.1 μg of plasmid library, 1.1 μg of psPAX2, and 0.6 μg of pMD2.G, using Fugene HD as the transfection reagent in a ratio of Fugene 3 μL:DNA 1 μg. The culture medium was changed on the plate 24 hours after transfection. The viral supernatant was collected approximately 72 hours after transfection, filtered through a 0.45 μM filter, and flash-frozen in aliquots. The viral aliquots were thawed on ice, and viral titers were established using Lenti-X GOSTIX (Takara). Cells were ready for transfection 24 hours after seeding in a 100 mm Petri dish.

[0097] Transfection mixes were prepared in 15 ml tubes for each 100 mm Petri dish containing pLjm1_Twist Tat Library 8.5 μg; pMDLG / pRRE 7.6 μg (plasmid stored in the refrigerator); pRSV / pRev 4.0 μg; and pMD2.G 4.0 μg. Sterile water was then added to bring the final volume to 613 μl. 2 M CaCl2 87 μl was also added. After mixing the plasmid, water, and 2 M CaCl2, 2 × HBS 700 μl was added dropwise to the transfection mix, and the 15 ml tubes were gently mixed in a circular motion or slow vortex. The transfection mixes were incubated for 15 minutes and then added dropwise to the 100 mm Petri dishes. The plates were incubated at 37°C for 8 hours to overnight (8–14 hours) in a CO2 incubator with a 5% CO2 atmosphere. The calcium phosphate-containing medium was removed and replaced with 7 ml of complete DMEM medium (DMEM + 10% FBS), and incubated for 48 hours in a CO2 incubator at 37°C and a 5% CO2 atmosphere. The used medium containing the lentivirus was collected from fully confluent LentiX293T transfected cells and filtered through a 0.45 μm PES filter. The lentivirus was divided into aliquots and frozen or concentrated (skip to step 3). Multiple aliquots of lentivirus were prepared in the range of a few μl to 5 ml, depending on the transduction scale, for aliquot division. For large-scale production, the master mix of the transfection mix was used in multiple petri dishes. For even larger production, 5 ml lentivirus aliquots were prepared in 15 ml tubes. For lentivirus titration, small aliquots of lentivirus medium (50-200 μl) were prepared. The lentivirus stock was then stored in a freezer at -80°C.

[0098] Lentiviral library enrichment process by ultracentrifugation A filter-sterile PBS / 20% sucrose solution was prepared for the following day. After 48 hours, the supernatant was collected. The supernatant was pre-cleared by circulating it at 3,000 rpm for 5 minutes and filtered through a 0.45 μm filter. The volume of supernatant pooled from two 100 mm Petri dishes could be 14 ml. For Q-PCR, 20 U / ml of DNAase I was added. 10 ml of filter-sterile PBS / 20% sucrose was added to the bottom of the tube, and the supernatant was gently added onto the sucrose pad of the ultracentrifuge tube. The supernatant and sucrose should not be mixed. The ultracentrifuge was operated at 35,000 rpm for 2 hours at 4°C. (Note: Ultracentrifugation is very sensitive to tube balance. The weights of the tubes should be exactly the same.) After 2 hours, when a pellet appeared, the pellet was gently aspirated and removed with a pipette. The supernatant and PBS-sucrose supernatant were discarded in a 10% bleach. Next, the virus pellet was resuspended in 1 ml of complete DMEM (10% FBS + P / S).

[0099] Lentivirus stock titer (1) Counting and seeding of cells for transfection

[0100] LentiX293T cells were cultured in 10 ml of complete DMEM medium (DMEM + 10% fetal bovine serum) in a 100 mm Petri dish and incubated in a CO2 incubator at 37°C and 5% CO2 according to SOP2.x. Cells that reached 80%–90% confluence were ready for transfection. Used DMEM medium was discarded by aspirating from the Petri dish into a waste flask or by manual disposal. Cells were trypsinized with 1.5 ml of 0.25% trypsin solution according to SOP2.x. Cells from the final step were resuspended in 3 ml of complete DMEM medium. Cells were counted according to SOP2.x. 3 million cells were seeded into a new 100 mm Petri dish containing 10 ml of complete DMEM medium. Cells that reached 80–90% confluence were ready for the next passage. From the remaining cells, 500,000 cells from 500 μl of complete DMEM medium (DMEM + 10% FBS) in a 15 ml tube were added to 4.5 ml of complete DMEM medium. 100 μl of cells were added to each well of a 96-well plate. The cell density was 10,000 cells / well.

[0101] (2) Transduction of cells

[0102] Twenty-four hours after seeding the cells into 96-well plates, the cells were ready for transduction. Small aliquots of lentivirus were thawed on ice and used to measure lentiviral titers.

[0103] Preparation of serial dilutions of lentiviruses according to Table 1 [Table 1]

[0104] The culture medium was replaced with serial dilutions of lentivirus (100 μl), and the cells were transduced for 4 hours. This was done in triplicate. After 24 hours, 100 μl of complete DMEM medium with puromycin (3 μg / ml) was added to bring the total volume to 200 μl and the final puromycin concentration to 1.5 μg / ml. The plates were then incubated at 37°C for 120 hours in a 5% CO2 incubator. The cells were observed under a microscope. Colonies at the maximum dilution were considered to calculate the infective units (IFU) / ml. IFU / ml was calculated as follows: IFU / ml = Number of colonies × Dilution factor × 10

[0105] The triplet was averaged.

[0106] (3) Freezing of concentrated stock

[0107] The stock was frozen in 1.5 ml or 2 ml cryovials labeled with the virus name, date, researcher's initials, and titer. The lentivirus stock was stored in a -80°C freezer.

[0108] (Example 2) Design and validation of specific fluorescence assays for molecular function or cellular processes. Confirmation of functional assays using positive and negative controls The fluorescence assay of the sequence-based assay was validated. This can be performed using a separate approach that employs either a fluorescent reporter, a fluorescent dye, or immunostaining with an antibody. This example describes an example based on immunostaining of phospho-Her2 in cells.

[0109] In steps 4-6, cells were manipulated and the assays were tested for specific assays using controls. Controls, including variants and generated using lentiviral plasmids encoding the controls, were transduced into reporter cell lines using a standard approach. The reporter or immunostaining activity of the controls was measured by flow cytometry, fluorescence microscopy, and / or fluorimeter. Alternatively, sequence-based reporters may be used. Controls included Erbb2-free, wild-type Erbb2, and well-characterized variants or drug-sensitive variants with loss of function, gain of function, or drug resistance. In step 6, conditions were modified to maximize the dynamic range of the reporter signal using the controls. Further details are listed below.

[0110] Cell culture maintenance All experimental cell lines were generated using HEK-293T cells expressing the TET-On 3G element. Cells were maintained in Dulbecco's Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (Cytiva, #SH30396.03HI), 20 mM HEPES (Sigma Aldrich, #H3375), 60 mg / L penicillin G (Gold Bio, #P-304-100), and 100 mg / L streptomycin sulfate (Gold Bio, #S-150-50), along with 25 mM glucose and 1 mM sodium pyruvate (Thermo Fisher Scientific, #11-995-081), and all experiments were performed.

[0111] Transduction of control cells into cell lines Cells were transduced at an MOI of 0.1 in the presence of 10 μg / mL of polyblen (Millipore, TR-1003-G). Toxin selection was performed on the cells for 7 days in a medium supplemented with 2 μg / mL of puromycin (Gold Bio, #P-600-100). Cells were frozen in aliquots of 3E6 cells / cryovial and stored in liquid nitrogen. One vial of cells was thawed for each control assay, then passaged for one phase before the assay was performed. The maximum passage number for all experiments was 3 phases.

[0112] Each expression construct or control was used in Lenti-X 293T cells, 0.6 × 10⁶ cells. 6 The cells were plated into the wells of a 6-well plate and cultured for 24 hours. Each well was transiently transfected with 1.1 μg of GML plasmid library, 1.1 μg of psPAX2, and 0.6 μg of pMD2.G using FuGENE HD in a ratio of 3 μL:1 μg of DNA. Approximately 40 cells were displayed per clone. Each plate contained 10 μg of plasmid library (approximately 1.38 × 10¹⁶ plasmid DNA per clone) as described above. 6 9.5 μg of psPAX2 and 5.2 μg of pMD2.G were transfected using FuGENE. The culture medium was changed 24 hours after transfection. Approximately 72 hours after transfection, the supernatant containing the virus was collected, filtered through a 0.45 μM filter, divided into aliquots, and stored in a freezer. The virus aliquots were thawed on ice, and the viral titer was measured using Lenti-X GOSTIX (Takara).

[0113] Individual control experiments 1E6 cells were plated in a 100 mm plate. 24 hours after plating, variant expression was stimulated by treating the cells with 25 ng / mL doxycycline (Thermo Fisher Scientific, #BP25535) for 72 hours. Since the half-life of doxycycline in the medium is 24 hours, 12.5 ng / mL doxycycline was spiked into the medium every 24 hours to maintain a concentration of 25 ng / mL through variant induction. After 72 hours, the cells were collected by trypsin treatment and then fixed in 4% formaldehyde for 15 minutes. The cells were then washed with 1×PBS and permeabilized with 90% methanol (Thermo Fisher Scientific, BP1105-1) for 16 hours. Next, the cells were washed with 1×PBS and then transferred to blocking buffer (1×PBS supplemented with 1% bovine serum albumin (Rockland Immunochemicals, #BSA-50) and 5 mM EDTA (Thermo Fisher Scientific, 15575020)) and stored at 4°C until immunostaining.

[0114] (Example 3) GML High-Throughput Screening Assay Generation of cell lines for pooled libraries Cells were transduced with a plasmid library lentivirus at an MOI of 0.1 in the presence of 10 μg / mL polyblen. The expected yield was 20 transduced cells per clone. Subsequently, the cells were subjected to toxin selection for 7 days in a medium supplemented with 2 μg / mL puromycin. The cells were frozen in aliquots of 5E6 cells / cryovial and stored in liquid nitrogen. Three vials (15E6 cells) were thawed for GigaAssay, and then grown for one passage before GigaAssay was performed. The maximum passage number was 3-4.

[0115] High-throughput screening assay In GigaAssay, 40 million cells were plated on day 1, and ultimately, approximately 400-500 million cells were collected on day 5. Cells were plated in 20 plates per GigaAssay at a density of 2 million cells / 150 mm plate in 20 mL of medium. Variant expression was stimulated 24 hours after plating by treating cells with 100 ng / mL doxycycline (Thermo Fisher Scientific, #BP25535) for 72 hours. A concentration of 100 ng / mL was maintained through variant induction by spiking the medium with 50 ng / mL doxycycline every 24 hours. After 72 hours, cells were collected by trypsin treatment and then fixed in 4% formaldehyde for 15 minutes. Next, the cells were washed with 1×PBS and then transferred to blocking buffer (1×PBS supplemented with 1% bovine serum albumin (Rockland Immunochemicals, #BSA-50) and 5 mM EDTA (Thermo Fisher Scientific, 15575020)) and stored at 4°C until immunostaining.

[0116] immunostaining The cells were fixed in 4% formaldehyde for 15 minutes. Then, after washing with PBS, the cells were permeabilized with 90% methanol (Thermo Fisher Scientific) for 16 hours. Next, after washing with PBS, the cells were transferred to blocking buffer (PBS supplemented with 1% bovine serum albumin (Rockland Immunochemicals) and 5 mM EDTA (Thermo Fisher Scientific)) and stored at 4°C until immunostaining. The cells in blocking buffer were incubated at 4°C for 1 hour with dilutions of monoclonal antibodies generated against human Her2 (rabbit mAb, Cell Signaling) and (p)Her2 (Tyr1248) (mouse mAb, Thermo Fisher Scientific). The cells were washed three times with blocking buffer and then incubated at 4°C for 1 hour with secondary antibodies conjugated to various fluorophores. For individual control assays, the following secondary antibodies were used: rabbit IgG(H+L) cross-adsorbed Alexa Fluor 647 conjugate (for Her2) (Fisher Scientific) and mouse IgG(H+L) cross-adsorbed Alexa Fluor 488 conjugate (for pHer2) (Thermo Fisher Scientific). For the GigaAssay, the following secondary antibodies were used: rabbit IgG(H+L) cross-adsorbed PE conjugate (for Her2) (Fisher Scientific) and mouse IgG(H+L) cross-adsorbed Alexa Fluor 647 conjugate (for pHer2) (Thermo Fisher Scientific) for 1 hour. Cells were washed three times with blocking buffer and then stored in blocking buffer at 4°C until FAC analysis or sorting.

[0117] Cell sorting into bins by flow cytometry Individual cells for the GigaAssay experiment were analyzed and sorted by flow cytometry. All flow cytometry analyses were performed using a Sony SH800Z flow cytometer (SONY, Tokyo, Japan). At least 10,000 events were captured for each sample. First, cells were gated for high Her2 expression (AF647), and then their pHer2 expression levels (AF488) were compared. FAC data analysis for individual control experiments was performed using FlowJo (v.10.8.1, Ashland, OR).

[0118] Cell population selection for GigaAssay was performed on FACsAria II (BD, Franklin Lakes, NJ) using a Stanford flow cytometry core. Cells were gated for high Her2 expression (PE) and then sorted into four bins according to increasing pHer2 expression levels (AF647). Based on the percentage of the total cell population, the cell population was sorted into bin 1 = 50% of cells with the lowest pHer2 expression, bin 2 = the next 20%, bin 3 = the next 20%, and bin 4 = 10% of cells expressing the highest level of pHer2. 5 × 10 6 Individual cells were sorted and placed into separate vials. After sorting, the cells were pelletized by centrifugation, washed once with PBS, and stored at -80°C.

[0119] (Example 4) Targeted NGS sequencing of UMI barcodes in gDNA Next-generation sequencing of UMI-barcoded variant libraries For next-generation sequencing, PCR primers were designed to add twisted nucleotides to the 5' and 3' ends of the targeted insert, along with the Nextera read adapter sequence, to enhance diversity during the first 18 cycles of Illumina sequencing, while flanking the targeted UMI region of the variant library. The 5' primers were a pool of 10 individual primers (oHI-00210~oHI-00219), all targeting the same site but each containing a different twisted nucleotide sequence between the PCR primer binding site at the 5' end and the Nextera read 1 adapter sequence overhang. The 3' primers were also a pool of 10 individual primers (oHI-00333~oHI-00342), each containing a different twisted nucleotide sequence between the PCR primer binding site at the 5' end and the Nextera read 2 adapter sequence overhang. gDNA extracted from the selected cell population was divided and PCR amplified with the aforementioned PCR primer pool using the NEBNext Ultra II Q5 Master Mix (New England Biolabs, #M0544L, Ipswich, PA) in a multiple reaction (without using more than 2.0 μg of DNA per reaction). This reaction involved 25 cycles of initial denaturation at 98°C for 2 minutes; denaturation at 98°C for 10 seconds, annealing at 60°C for 10 seconds, extension at 72°C for 20 seconds; final extension at 72°C for 2 minutes; and termination by holding at 4°C.

[0120] PCR products with sizes ranging from 203 to 221 bp were purified using SPRIselect Beads (Beckman, #B23317, Brea, CA) for size exclusion and selection at 0.65 × right-side beads:sample ratio and 1.0 × left-side beads:sample ratio. These were eluted with 30 μl of Buffer EB (Qiagen, Germantown, MD) and pooled for each selected cell population. 5 μL of the pooled, size-selected PCR product from the first PCR was used in a second PCR with limited cycles to add a dual index, with a P5 sequence at the 5' end and a P7 sequence at the 3' end, respectively. The PCR primers (oHI-00350 to oHI-00369) were a selection of eight primers with various i5 indices and twelve primers with various i7 indices. These were used to generate up to 96 different dual index combinations depending on the number of selected cell populations to be sequenced.

[0121] PCR was performed using KAPA HiFi HotStart ReadyMix, following the PCR protocol: initial denaturation at 98°C for 3 minutes; denaturation at 98°C for 30 seconds, annealing at 55°C for 30 seconds, extension at 72°C for 30 seconds, final extension at 72°C for 5 minutes; and termination by holding at 4°C for 6 or 8 cycles. PCR products ranging in size from 272 to 290 bp were purified using SPRIselect Beads (Beckman, Brea, CA) for size exclusion and selection at 0.8 × right-side beads:sample ratio and 1.0 × left-side beads:sample ratio, and eluted with 25 μl of Buffer EB (Qiagen, Germantown, MD). The final constructed NGS library was quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher Scientific, Hampton, NH), and the predicted size was determined using the DNA 7500 Kit for Bioanalyzer 2100 (Agilent, #Santa Clara, CA). Sequencing was performed on a NextSeq 500.

[0122] Isolation of genomic DNA (gDNA) from cells Genomic DNA is isolated from cells collected in step 3m using a Genomic DNA isolation kit (Zymogen, catalog number D4068 or Quick-DNA™ Miniprep Plus Kit) according to the manufacturer's protocol.

[0123] Cells collected in a 15 ml tube were pelletized by centrifugation at 1000 rpm for 5 minutes. The culture medium was discarded and 1 ml of PBS was added. The resuspended cells were transferred to a 1.5 ml microcentrifuge tube. The tube was centrifuged at 3000 rpm for 5 minutes, and the PBS was discarded by aspiration. The cell pellet was resuspended in 200 μl of PBS. 200 μl of BioFluid & Cell Buffer (red) and 20 μl of proteinase K were added. The cell pellet and reagents in the tube were thoroughly mixed or vortexed for 10-15 seconds, and then the tube was incubated at 55°C for 10 minutes. 1 volume of genome binding buffer was added to the digested sample and thoroughly mixed for 10-15 seconds. The DNA mixture was then transferred to a Zymo-Spin® IIC-XLR column in the collection tube. Centrifuged at ≥12,000 × g for 1 minute. The collection tube showed flow-through. 400 μl of DNA pre-wash buffer was added to the spin column in a new collection tube. Centrifuged at ≥12,000 × g for 1 minute. The collection tube was emptied, and 700 μl of g-DNA wash buffer was added to the spin column. The spin column was then centrifuged at ≥12,000 × g for 1 minute, and the collection tube was emptied after centrifugation. 200 μl of g-DNA wash buffer was added to the spin column. Centrifuged at ≥12,000 × g for 1 minute. The collection tube with flow-through was discarded, and the spin column was transferred to a clean microcentrifuge tube. ≥50 μl of DNA elution buffer was added directly to the matrix, incubated at room temperature for 5 minutes, and then centrifuged at maximum speed for 1 minute to elute the DNA.

[0124] The eluted DNA is quantified using a Nanodrop spectrophotometer.

[0125] The eluted DNA can be used immediately for molecular-based applications or stored at -20°C for screening and subsequent use.

[0126] (Example 5) Bioinformatics analysis of screening results for generating MEGA-MAPs through variant classification Each high-throughput screening assay described herein utilized long-read NGS sequencing of UMI-coding plasmid libraries and short-read sequencing of gDNA isolated from flow-sorted bins. Using these data, a pipeline was designed to call variants for each UMI-coding cDNA, and then calculate activity for each UMI-coding cDNA from the UMI frequencies of the short reads in each flow bin. The bioinformatics pipeline for GigaAssay is summarized in Figure 14: Bioinformatics Pipeline. The goal of this pipeline is to generate a UMI barcode variant map, creating an index between barcodes identified in the long-read library and the variants in the library. This index is later used to generate flow-sorted short reads, which are used to quantify the barcodes and generate activity scores.

[0127] QC scanning and lead trimming Long read data from the fastq file library was analyzed for low-quality reads and trimmed. This step is unnecessary when using PacBio CCS reads. Reads with a 4-base pair (BP) window and an average phred score ≤ 15 were identified and filtered using Trimmomatic. The following command was used for single-ended read data, and a results file containing all reads that passed the sliding window test was generated using 24 threads. The log file was saved to trimmomaticLog.txt.

number

[0128] Adapter trimming The reads were trimmed using Cutadapt so that the variant region and barcode were present within the read, with a 15BP buffer before and after the variant region. Trimming the reads creates space for sequence alignment and variant calling. A 12BP storage region (3':CGTCAGATCCGC; 5':GCGATCGCAGCG) was input to Cutadapt as the 3' and 5' boundaries to remove excess BP. Cutadapt searched these sequences using its default error tolerance, which includes a 10% missense or indel error. This setting allows for one error. If either or both of the 12BP regions were present, the BP inside and before the 3' region, and after the 5' region, were removed and copied to the output file. Fragments of reads that were neither trimmed nor excised were output to a separate file. Cutadapt detected, trimmed, and reversed reads that were inversely complementary in the output file using the --rc parameter in the following command:

number

[0129] Barcode Extraction The barcode was extracted from the read while allowing for a barcode error tolerance, and the read was demultiplexed. The UMI barcode (32bp) was extracted from the filtered read, a new 3' region (12bp) was selected using Cutadapt, and the barcode was trimmed and extracted using the following command:

number

[0130] Barcode filtering Occasionally, during filtering, the sequences align with incorrect segments of the reads, resulting in a barcode size different from what was expected. Therefore, only barcodes between 28 and 36 in length were selected using the Cutadapt awk command and included in the output file:

number

[0131] Barcode grouping UMI barcodes were grouped so that each group contained a UMI with a sequencing error. The barcodes were clustered using Starcode and the Levenshtein distance 2 with the following command:

number

[0132] demultiplexing Reads with the same UMI barcode are generated in various samples, e.g., in various flow-sorted bins. To call variants, reads with the same barcode base must be isolated and grouped. This process was completed using a custom Python script. This script uses the output above and demultiplexes the reads based on the identified barcode base. If the primary barcode length was not 32, the barcode base was not demultiplexed. If the primary barcode length was 32, the reads were demultiplexed using the following command to generate a directory and fastq file for each barcode base:

number

[0133] Variant Call A reference sequence was specified for sequence alignment and variant calling. The variant region (along with 15BP buffers on both sides) was placed within the fasta file for alignment. Reads were aligned to the reference using BWA (http: / / bio-bwa.sourceforge.net / ). The ref file was initially indexed using the following command: bwa index(IN FILE)

[0134] The file had a supporting index file that was generated in the same directory where the file resided.

[0135] To map each barcode to a specific variant, variants were called using a custom script called caller.py. The caller.py script parsed the fastq of each barcode base into a bwa mem, converted the output to a bam, sorted it using Samtools, ran Samtools mpileup, and then ran bcf tools to call unique variants. This generated a VCF derived from each barcode base and its associated reads.

[0136] The CCS read was reported to have incorrect indel call errors generated from incorrect sequencing. The indels were removed using callerNoIndel.py. The variant was called using the following command with the caller script, indexed reference files, demultiplexed directories, and the required number of each:

number

[0137] Each fastq file in a subdirectory has a corresponding VCF file that was generated for use in variant interpretation.

[0138] Indexing Barcode hopping is frequently observed in the sequencing results. A custom script quantifies the number of barcode hopping events. Using the identified core count in a subdirectory, the selected BAM files were indexed and corresponding index files were generated according to the following command:

number

[0139] Indelkol Insertion and deletion variants (indels) were identified using the Nanocaller program. The parameter settings that yielded the best performance were "--mode indels --mincov 3 --sequencing pacbio --del_threshold 0.4 --suppress_progress_bar --impute_indel_phase --enable_whatshap", but these parameters should be optimized for each library analyzed. These settings were captured in the "IndelCaller.py" program and used to perform indel calls on all subdirectories.

number

[0140] Mutant-Barcode Modeling After variant calling, a variant-barcode map was generated to interpret the short-read screening gate. A custom Python script was used to filter reads based on QC and the required read depth per barcode. When working with CCS reads, QC 30 and depth 1 are recommended. Alternatively, the ERBB2 library was processed with the lowest QC 0 and depth 1, and further filtered downstream. This variant-barcode map was generated using the following command:

number

[0141] The output is a CSV file containing all barcode bases, each variant identified for each barcode base (and whether each variant was predicted by twist), and some QC statistical data. All variant information required for short-read interpretation and statistical data is summarized and output to phenoModel.csv. This flowchart illustrates the concise statistical portion of the long-short-read pipeline.

[0142] (Example 6) Use of HiFi DNA Assembly for GML Creation The following steps describe the creation of a GML that covers a 150-nucleotide variant region based on the coding sequence of the target gene, using HiFi DNA assembly.

[0143] Variant Assembly

[0144] As shown in Figure 2, the lentiviral expression vector contains the sequence encoding the target gene (GOI) along with a variant region corresponding to the 50-amino acid coding sequence. Two sections near the AsiSI-MluI restriction site in the 3'UTR correspond to UMI (Unique Molecular Identifier) ​​insertion sites, which are later covered during UMI assembly. The variant region is 150 bp, and the 25 bp overlapping DNA sequence flanking the variant region results in a 200 bp tile (Figure 3). Figure 4 shows a comparison of the inverse PCR product and the variant oligopool in the HiFi DNA Assembly preparation. The inverse PCR primer is designed to include a 25 bp sequence that overlaps with the oligopool. The oligopool consists of up to several thousand oligos containing the desired mutation. The T5 exonuclease partially digests the dsDNA strand from 5' to 3' (Figure 5). This exposes the 3' DNA strand and allows it to anneal to the complementary sequence found in the oligopool (Note: In this example, we focus on one oligo derived from the oligopool). HiFi DNA polymerase synthesizes the DNA from 5' to 3' to synthesize the reverse strand of the variant oligo and simultaneously fills the gap created after DNA strand annealing. Next, DNA ligase closes the nick to create a continuous DNA strand (Figure 6). Then, T5 exonuclease continues digestion of the exposed end of the DNA strand from 5' to 3', exposing the DNA and allowing it to anneal to the complementary strand (Figure 7). DNA polymerase continues DNA synthesis from 5' to 3' to fill the gap, and DNA ligase closes the nick to create a continuous DNA strand. The final result of complete assembly is a traceless recombinant vector containing the target variant cloned within the variant region (Figure 8). The result of the complete assembly is a traceless recombinant vector containing the target variant cloned within the variant region (Figure 9).

[0145] HiFi DNA Assembly of GMLs requiring two or more oligopools As shown in Figure 10, the target DNA sequence (blue) is the variant region targeted by the GML. The target DNA sequence is divided into 200-nucleotide tiles. These tiles correspond to a 200-mer synthetic oligo for the desired variant. Each tile consists of a variant region flanked by 25 nucleotides that overlap with the vector cloning site, which also functions as an inverse PCR primer binding site for inverse PCR of the vector DNA. The tiles overlap to cover the entire target region. In this embodiment, the target DNA sequence is 600 bp long and is covered by three overlapping tiles. Each tile is processed as a separate variant assembly reaction product in the expression vector using the same general strategy as described above. Variant assembly yields three separate assembly product pools, which are combined to generate a plasmid variant library (UMI-free). The 3'UTR of the vector contains unique restriction sites (AsiSI and MluI). These sites are used to linearize the vector DNA, enabling HiFi DNA assembly. UMI is then added to the 3'UTR using ssDNA UMI oligos, resulting in a plasmid variant library in which UMI is added to each plasmid DNA molecule in the library.

[0146] UMI Assembly This step is the same as the variant assembly step, except for the UMI assembly, which linearizes the variant plasmid library using restriction enzymes at the 3'UTR. This linearized DNA is then used in a HiFi DNA Assembly reaction with a UMI oligo containing 25 nucleotides at the 5' and 3' ends that overlap with 25 bp DNA sequences at the 5' and 3' ends of each linearization vector DNA molecule. The UMI oligo is also designed to include 32 degenerate nucleotides that function as unique molecular identifiers for each molecule in the plasmid library. The 3'UTR UMI assembly region (Figure 11) contains two restriction sites, AsiSI and MluI, used for linearizing the variant plasmid library. The HiFi DNA Assembly reaction is carried out using the AsiSI-MluI linearization vector DNA and a UMI oligo with a 25-nucleotide overlap sequence containing the cloning site (Figure 12). The assembly reaction is carried out in the same manner as in Figures 5-8 above until assembly is complete. The final product of the UMI assembly is a plasmid GML containing the desired variant along with a unique molecular identifier for each plasmid DNA molecule in the library (Figure 13).

[0147] Considerations on using HiFi DNA Assembly for GML generation Before starting, the target DNA sequence and vector DNA sequence should not contain restriction sites that conflict with the cloning process, including a unique restriction site in the 3'UTR for UMI assembly. If any restriction sites are present in the DNA sequence, they should be removed during library design.

[0148] The assembly process requires multiple rounds of electroporation depending on the desired clonal count to achieve the desired library coverage (at least approximately 200x coverage). For example, if 200x coverage is required for 600 variants, 120,000 clones should be generated in each of these steps. The average cloning efficiency of variant assembly combined with electroporation using Endura electrocompetent cells is approximately 1 × 10^8 CFU / ug or approximately 2 × 10^6 CFU / ml for a tile of 200 bp and a vector size of approximately 11.2 kb. The average cloning efficiency of UMI assembly combined with electroporation using Endura electrocompetent cells is approximately 1.7 × 10^8 CFU / ug or approximately 3.4 × 10^6 CFU / ml for a UMI oligo of 123 nucleotides and a vector size of approximately 11.2 kb.

[0149] Experimental indexing of UMI using variant cDNA by long-read sequencing of plasmid libraries DNA digestion and cleanup for PacBio long-read sequencing

[0150] Prepare a restriction enzyme digestion product and digest 10 μg of plasmid-based gene variant library (Note: Do not exceed 1 μg of DNA per 50 μl of restriction enzyme digestion product).

[0151] The quantities are based on 1-50 μl of reactant (adjusted for every 10 μg of DNA), and the following reaction setup is used: 1 μg (plasmid-based gene variant library); up to 50 μl of nuclease-free solution; 5.0 μl of 10×CutSmart buffer; 1.0 μl of NheI-HF (30U); and 1.0 μl of MluI-HF (30U).

[0152] The reaction mixture is mixed by aspirating and discharging 10 times with a pipette, and then centrifuged for 15 seconds to pull down the liquid. After incubation at 37°C for 1 hour, the gel loading dye purple (6x) (SDS-free) is added to the total reaction volume to be loaded onto the gel. The mixture is then mixed again by aspirating and discharging 10 times with a pipette, and then centrifuged briefly to pull down the liquid.

[0153] Load the sample into each well without puncturing the sides of the wells with the tip of the pipette. Each well of the 6-well comb can hold a maximum volume of 60 μl. Load the total reaction volume onto the agarose gel(s). Then, load 6 μl of Benchtop kb DNA ladder into the empty lane in the row of wells, place the cover and electrodes on the unit, and start electrophoresis using a voltage set to 120 V. Monitor the electrophoresis periodically. Once the bands are clearly separated, take an image of the gel and you can determine the size using the provided photo viewer on a smartphone or tablet.

[0154] Annotate the gel image in LucidChart (preferred) or PowerPoint. Label the gel image with the experiment name, lane number, agarose percentage, and date. Proceed with gel extraction and cleanup of desired DNA fragments.

[0155] When using the NucleoSpin gel extraction kit, set the heat block to 50°C and 65°C. Use a clean 1.5 ml microcentrifuge tube for each sample. Add the empty weight to the Weight Calculation table (in the table section of Benchling's notebook entry). Transfer the gel to a blue LED transilluminator. Switch on the blue LED light and use a fresh razor blade to cut out the desired DNA band.

[0156] Place the excised bands into microcentrifuge tubes. Each tube corresponds to a band excised from each lane of the gel. Weigh each tube containing the gel fragments and record the total weight of the gel in the Weight Calculation table (in the table section of the Benchling notebook entry).

[0157] Gel extraction is carried out in the following steps: Add 2 volumes of Buffer NT1 to 1 volume of gel. Therefore, 200 μl of Buffer NT1 is added to 100 mg of each gel. For a >2% agarose gel, the volume of Buffer NT1 is doubled.

[0158] Heat the mixed solution at 50°C for 10 minutes or until the gel slice is completely dissolved. Mix the solution by vortexing the tube every 2-3 minutes during incubation. Place the NucleoSpin column in the provided 2 ml collection tube. Load up to 700 μl of sample onto the column at once (Note: According to the NucleoSpin manual, up to 200 mg of agarose can be dissolved in 400 μl of Buffer NT1 and loaded onto the column in one step. However, by proportionally increasing Buffer NT1 and adding multiple loading steps, it is possible to load virtually unlimited amounts of gel without clogging the column). Rotate the column at 11,000 × g for 30 seconds. After rotation, discard the flow-through and return the NucleoSpin column to the same collection tube. If necessary, load the remaining sample and repeat the centrifugation step. The loading step should be repeated for all tubes corresponding to the excised DNA bands of each plasmid library being processed, so that all excised DNA from that sample is loaded onto the same column. If the limiting digestion reaction volume for 10 μg of plasmid DNA plus the total volume of loading dye is 550 μl, loading 60–65 μl into each lane of an agarose gel (using a 6-well comb) results in two agarose gels. In the first gel, load the sample into 5 lanes (60–65 μl per lane) to obtain a DNA ladder in lane 6. In the second agarose gel, load the sample into 4 lanes (60–65 μl per lane), leaving one empty lane for the DNA ladder. This results in 9 tubes of excised DNA (each tube corresponding to each sample lane on the agarose gel). After adding Buffer NT1 and incubating at 50°C for 10 minutes, load these 9 tubes of excised DNA onto a single NucleoSpin spin column, loading up to 700 μl at a time.

[0159] After loading all sample tubes of excised DNA corresponding to the same plasmid library onto the column, add 650 μl of Buffer NT3 to the NucleoSpin column and centrifuge at 11,000 × g for 30 seconds. After centrifugation, discard the flow-through. Next, add 650 μl of Buffer NT3 to the NucleoSpin column and centrifuge at 11,000 × g for 30 seconds. After centrifugation, discard the flow-through. Add 650 μl of Buffer NT3 to the NucleoSpin column and centrifuge at 11,000 × g for 30 seconds. After centrifugation, discard the flow-through without contacting the spin column with the wash flow-through. Wipe the outside of the column with a Kimwipe, transfer the contents to a new collection tube, and centrifuge at 11,000 × g for a further 2 minutes to dry the column. Place the NucleoSpin spin column in the new collection tube and dry the column by placing it on a 65°C heat block with the lid open for 5 minutes. Next, place the NucleoSpin column in a clean 1.5 ml microcentrifuge tube. Add the desired volume (20 μl) (15-30 μl) of Buffer NE (room temperature) to the center of the NucleoSpin spin column membrane. Allow the column to stand at room temperature for 5 minutes, then centrifuge at 30 × g for 1 minute and at 11,000 × g for 1 minute. After centrifugation, add another 20 μl of Buffer NE (room temperature) to the center of the NucleoSpin spin column membrane. Next, allow the column to stand at room temperature for 1 minute, then centrifuge at 30 × g for 1 minute and at 11,000 × g for 1 minute. After centrifugation, add another 20 μl of Buffer NE (room temperature) to the center of the NucleoSpin spin column membrane. Next, allow the column to stand at room temperature for 1 minute, then centrifuge at 30 × g for 1 minute and at 11,000 × g for 1 minute. After centrifugation, discard the spin column. The eluted DNA is used in downstream applications. The DNA is then quantified, the yield calculated, and the data recorded. Final cleanup is performed using size-selective beads.

[0160] Before starting, use either AMPure XP or SPRI Select size selection beads. The size of the selected DNA fragments is ≥2,000 bp. Remove any DNA fragments smaller than the target fragment size. If using AMPure XP beads, allow them to stand at room temperature for 30 minutes before use. Prepare fresh 80% molecular-grade ethanol in appropriate volumes for all samples to be processed. Each sample requires 400 μl (excluding dead volume). Proceed with processing the samples in 1.5 ml microcentrifuge tubes. If working with >8 samples and volumes <50 μl, the samples may be transferred to 0.2 ml 96-well plates.

[0161] In this procedure, a 0.5 × bead:sample ratio is used to bind all DNA fragments ≥ approximately 2000 bp, removing smaller fragments and leaving them in the supernatant.

[0162] Vortex the size-selection beads until well dispersed, then add 0.5 μl of the well-mixed size-selection beads per μl of DNA sample. Gently pipette and drain the entire volume 10 times to mix thoroughly. Allow the DNA fragments to bind to the paramagnetic beads by incubation at room temperature for 15 minutes.

[0163] Place the plate on the magnetic stand for 2 minutes or until the supernatant becomes clear. Using a pipette, remove and discard the supernatant containing any unwanted fragments and impurities. With the tube still on the magnetic stand, add 200 μl of 80% ethanol to further remove impurities. Incubate the plate on the magnetic stand at room temperature for 30 seconds, then remove and discard the supernatant. With the tube still on the magnetic stand, add 200 μl of 80% ethanol to further remove impurities. Incubate the tube on the magnetic stand at room temperature for 30 seconds, then remove and discard the supernatant. Centrifuge the sample by microcentrifugation. After centrifugation, return the tube to the magnetic stand and remove any excess ethanol using the fine tip of a P20 multichannel pipette. Remove the plate from the magnetic stand before all ethanol is removed and the beads dry. Add 40 μL of elution buffer (10 mM Tris-HCl (pH 8.5) (i.e., Qiagen Buffer EB)) to elute the DNA, and mix thoroughly by gently aspirating and discharging with a pipette 10 times. Incubate at room temperature for 5 minutes without shaking, then place the tube on a magnetic stand for 2 minutes until the supernatant becomes clear. Then transfer the entire volume to a new microcentrifuge tube with a screw cap. If using a multichannel pipette, exchange the tip between samples or between rows or columns.

[0164] Next, the DNA is quantified, the yield is calculated, and the data is recorded.

[0165] 40 ng of DNA is run on a 1% agarose gel to ensure the presence of a dominant single band of the expected size. The DNA is then sent to a sequencing core on dry ice for PacBio sequencing.

[0166] (Example 7) Application of high-throughput screening for dominant-negative biologics As shown in Figure 15, two process strategies are used to identify dominant-negative biologics by analyzing a set or subset of reads from the gene coding region to experimentally identify one or more lead therapeutics. The high-throughput screening assays described herein are used to identify dominant-negative alleles that yield pathological assays. As is currently understood, the average human protein has 4.5 domains, and two or more domains are required for correct gene activity. Genetic variant libraries (GMLs) of WT gene variants (dominant-negative GML libraries of candidate genes prioritized by GWAS, polygenic risk score (PRS), or other omics approaches) are generated, where each variant has a deletion in a different domain of the protein but maintains the correct reading frame. One or more of these variants are expected to function correctly. These variants, having all the other domains, act as dominant-negative alleles when competing for cellular factors. Alternatively, it may polymerize with the wild-type allele into a macromolecular complex, acting as point subunits 28755412 and 19544283.

[0167] The specific high-throughput screening assays selected for these experiments are appropriate for the pathological function to be treated and are used to assess how drugs affect molecular function, cellular processes, or both. Up to 5,000 combinations are tested using the high-throughput screening assays described herein. Fluorescence readout information in cells, manipulated cells, or both is generated as the outcome. The output from this initial screening is used to identify dominant-negative variants for most of the genes under test. One or more of the identified dominant-negative variants are investigated as novel drug leads that act as antagonists to specific pathways.

[0168] (Example 8) Application of high-throughput screening for bispecific biologics As shown in Figure 16, to identify bispecific biologics, novel GMLs are constructed using one or more of the identified dominant-negative variants, and the genes containing the dominant-negative variants are fused to the wild-type (WT) or DN variant versions of other genes to create a chimeric library of all or nearly all variants using combinations of the starting gene and these variants. This GML is screened using the high-throughput screening assay described herein. Since each chimera may affect two or more reporters, reads are validated using a multi-orthogonal high-throughput screening assay. The GML also includes a control variant that is not DN.

[0169] By identifying reverse agonists or dominant-negative versions of genes using the high-throughput screening assays described herein, one or both of these can be developed into bispecific biologics. Both can replace neutralizing humanized monoclonal antibody therapies. These have the advantage of being unique molecules and, due to minimal amino acid changes, should not induce an immune response like antibody therapy. Furthermore, these therapeutics can be used as more general alternative therapies to treat patients who develop an immune response to antibody-based therapies. Examples include reverse agonists of TNF as an alternative to dominant-negative PCSK9 or Humira or related biosimilar drugs for treating hypercholesterolemia.

[0170] Preferred embodiments of the Disclosure are shown and described herein, but it will be obvious to those skilled in the art that such embodiments are provided merely as examples. Herein, numerous variations, alterations, and substitutions can be found to those skilled in the art without departing from the methods of the Disclosure. It should be understood that in the practice of the methods of the Disclosure, various substitutes for embodiments of the methods described herein may be utilized. The scope of the Disclosure is defined by the following claims, and is intended to encompass methods and structures within the scope of such claims and their equivalents.

Claims

1. A method for identifying mutagenic effects on protein activity, (a) A step of obtaining a library containing a plurality of cDNAs, wherein each cDNA in the plurality of cDNAs encodes a unique variant (may be a plurality) of a target protein, and each unique variant has one or more amino acid substitutions compared to the target protein, (b) A step of independently and individually incorporating one of the plurality of cDNAs into a plasmid, so that each of the plurality of cDNAs is individually and independently incorporated into its own plasmid and operably linked to a promoter on the plasmid, thereby forming a plasmid cDNA library, wherein each plasmid in the plasmid cDNA library further comprises a unique molecular identifier (UMI) or a barcode group. (c) A step of converting the plasmid cDNA library into a viral library containing virions, wherein each virion contains only one plasmid from the plasmid cDNA library, and the viral library may be a lentiviral library as needed. (d) A step of forming a human or mammalian cell line by manipulating a human or mammalian cell line to encode a receptor operably linked to a reporter system, or optionally a receptor expressed on its surface, wherein the receptor and the reporter system can indicate whether one of the unique variants binds to the receptor and modulates its biological activity. (e) Transduction of virions from the viral library in (c) into cells of the manipulated human or mammalian cell line in (d) so that each virion is transduced into a different cell, thereby forming a library of transduced cells. Methods that include...

2. The method according to claim 1, wherein expression of the unique variant of the target protein occurs in the transduced cells of the manipulated or mammalian cell line.

3. The method according to claim 1 or 2, further comprising the step of determining a measure of the biological activity of the unique variant in the transduced cells of the manipulated or mammalian cell line.

4. The method according to any one of claims 1 to 3, wherein the step of determining includes a step of calculating the measured value of the biological activity using a computer.

5. The method according to any one of claims 1 to 4, further comprising the step of selecting transduced cells containing the unique variant by flow cytometry using UMI or a barcode group.

6. The method according to claim 5, comprising the step of selecting a plurality of cells or all cells from the library of transduced cells to form a flow-selected pool(or more) of the UMI(or more) or a flow-selected pool(or more) of the barcode base.

7. The method according to any one of claims 3 to 6, wherein the measurement of the biological activity includes the steps of detecting the binding of the unique variant(s) to the receptor operably linked to the reporter system, and determining the amount, intensity, or both of the report from the reporter system when measured by fluorescence microscopy, flow cytometry, or both.

8. The method according to claim 5 or 7, further comprising the step of comparing the distribution of the flow-selected pool of the UMI or barcode group in a population of single cells individually expressing the unique variant(s) with the distribution of the flow-selected pool of single cells expressing a wild-type target protein or wild-type cDNA that does not have one or more amino acid substitutions.

9. The method according to any one of claims 5 to 8, wherein the measurement of the biological activity of the unique variant(s) is repeated at least 100 times using the UMI, the barcode, or both.

10. The method according to any one of claims 1 to 9, wherein each UMI or barcode independently comprises about 12 to about 32 nucleotides.

11. The method according to any one of claims 1 to 10, comprising transducing the cells of the manipulated human or mammalian cell line with the virus library at a multiple of infection (MOI) of at least 0.1 to at least 1.

12. The method according to claim 11, wherein the MOI is 0.

1.

13. The method according to claim 11 or 12, wherein the MOI prevents double insertion of virions into a single transduction cell or minimizes the transduction error rate.

14. The method according to any one of claims 1 to 13, wherein the cells of the manipulated human or mammalian cell line are selected using a marker for lentiviral integration.

15. The method according to any one of claims 1 to 14, wherein the receptor and the reporter system are operably linked to an inductive promoter.

16. The method according to claim 15, wherein the inducible promoter drives the expression of the receptor and the reporter system.

17. The method according to claim 15 or 16, wherein the inducible promoter includes a doxycycline-inducible promoter.

18. The method according to claim 16 or 17, wherein the reporter system comprises a polynucleotide encoding a fluorescent protein, and the fluorescent protein is GFP.

19. The method according to any one of claims 1 to 18, comprising isolating a single cell population based on the UMI using flow cytometry and sorting it into one or more vials.

20. The method according to claim 19, wherein one or more bottles contain the cDNA(s) encoding the unique variant(s), its amplicon(s), or any combination thereof.

21. The method according to claim 20, wherein the cDNA(s) encoding the unique variant(s), the amplicon(s) thereof, or any combination thereof is sequenced using a sequencing method.

22. The method according to claim 21, wherein the sequencing method includes next-generation sequencing, Sanger sequencing, whole-genome sequencing, RNA sequencing, or shotgun sequencing.

23. The method according to claim 21 or 22, wherein the sequencing method is next-generation sequencing, and the next-generation sequencing generates an array dataset.

24. The method according to any one of claims 19 to 23, wherein the sequence data includes the read depth for each unique variant.

25. The method according to claim 24, wherein each unique variant has a read depth of approximately 2,000 × to approximately 90,000 × sequencing coverage.

26. The method according to any one of claims 1 to 25, wherein a group of single cells expressing a unique variant is compared with a group of single cells expressing a wild-type target protein using a statistical model to test a specific hypothesis regarding the biological activity of each unique variant.

27. The method according to claim 26, wherein the biological activity of each unique variant includes a loss-of-function variant, a gain-of-function variant, a variant having activity substantially similar to that of the wild-type target protein, a drug-resistant variant, or a drug-sensitive variant.

28. The method according to any one of claims 1 to 27, further comprising the step of verifying the method by comparing the biological activity of a subset of unique variants analyzed by the method with the previously determined activity of a wild-type target protein.

29. The method according to any one of claims 1 to 28, further comprising the step of verifying the method by comparing the true negative determined by the method with the true negative determined by an independent method.

30. The method according to any one of claims 1 to 29, further comprising the step of verifying the method by comparing the biological activity of a subset of unique variants analyzed by the method with independent tests of separate sets of clones containing the unique variant(s).

31. The method according to any one of claims 1 to 30, further comprising the step of verifying the method by comparing the results of the method with one or more different samples.

32. The method according to claim 31, further comprising the step of verifying the method by comparing the results of the method in two different cell lines.

33. The method according to any one of claims 21 to 32, further comprising the step of generating an oligonucleotide sequence dataset.

34. A method further comprising the step of analyzing the oligonucleotide sequence dataset using a computer, wherein the analysis step is (a) A process of generating a unique variant-UMI index library using long reads derived from next-generation sequencing (NGS), (b) A step of identifying UMI counts and barcode bins by analyzing short reads of the flow-sorted group, (c) A step of calculating an activity score from the UMI count and the UMI-barcode read count including the barcode bin in (b), (d) A step of evaluating the effect of the unique variant(s) on biological activity by calculating the read percentage and p-value of a reporter protein, such as a fluorescent protein, (e) A step of accurately and quantitatively evaluating the activity of the unique variant(s) with respect to the previously characterized wild type. The method according to claim 33, wherein the fluorescent protein is GFP and the biological activity includes gene activity or protein activity.

35. The method according to any one of the preceding claims, wherein the library of transduced cells comprises about 10,000 to about 10 million cells.

36. The library of transduced cells contains approximately 10,000 cells, 20,000 cells, 30,000 cells, 40,000 cells, 50,000 cells, 60,000 cells, 70,000 cells, 80,000 cells, 90,000 cells, 100,000 cells, 150,000 cells, 200,000 cells, 300,000 cells, 400,000 cells, 500,000 cells, and 600,000 cells. The method according to claim 35, comprising 700,000 cells, 800,000 cells, 900,000 cells, 1,000,000 cells, 2,000,000 cells, 3,000,000 cells, 4,000,000 cells, 5,000,000 cells, 6,000,000 cells, 7,000,000 cells, 8,000,000 cells, 9,000,000 cells, or 10,000,000 cells.

37. A protein variant discovered by the method described in any one of the preceding claims.

38. A library of transduced cells formed by the method described in any one of claims 1 to 3.

39. A library of transduced cells according to claim 38, comprising approximately 10,000 to approximately 10 million cells.

40. A library of isolated and purified transdextrin cells comprising approximately 10,000 to 10 million cells, wherein each isolated and purified transdextrin cell comprises a plasmid containing cDNA encoding a unique variant of the wild-type protein, and the cell surface receptor can be examined for each unique variant, such that each transdextrin further comprises a surface-expressed receptor operably coupled to a reporter system.

41. The library according to claim 40, wherein each isolated and purified transduced cell further comprises a unique variant on its surface.