Compositions and methods for the continuous directed evolution of proteins in b cells
By leveraging B cell somatic hypermutation machinery with integration plasmids and CRISPR/Cas9, the method addresses the inefficiencies of traditional directed evolution, enabling rapid and efficient protein evolution with enhanced affinity.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-04-02
AI Technical Summary
Existing directed evolution methods for proteins are labor-intensive and time-consuming due to low mutation rates and genomic tolerance, requiring in vitro diversification and screening, which limits the efficiency of identifying biomolecules with desired functions.
Utilizing integration plasmids that harness the somatic hypermutation machinery of B cells to rapidly evolve proteins by integrating at a specific genomic locus, combined with CRISPR/Cas9 technology for targeted mutagenesis and display of proteins on the B cell surface, enabling continuous directed evolution without compromising cell viability.
Facilitates rapid and efficient generation of protein variants with enhanced affinity for target antigens through iterative mutagenesis and selection, reducing the need for extensive engineering and screening, and allowing for continuous evolution of proteins like antibodies.
Smart Images

Figure US2025048620_02042026_PF_FP_ABST
Abstract
Description
[0001] PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0002] 7950-112573-02
[0003] COMPOSITIONS AND METHODS FOR THE CONTINUOUS DIRECTED EVOLUTION OF PROTEINS IN B CELLS
[0004] CROSS REFERENCE TO RELATED APPLICATIONS
[0005] 5 This application claims the benefit of U.S. Provisional Application No. 63 / 700,867, filed September 30, 2024, which is herein incorporated by reference in its entirety.
[0006] FIELD
[0007] This disclosure concerns plasmids, kits and methods for rapidly evolving proteins of interest, such as antibodies or antibody fragments, by harnessing the somatic hypermutation (SHM) machinery of B cells.
[0008] INCORPORATION OF ELECTRONIC SEQUENCE LISTING
[0009] The electronic sequence listing, submitted herewith as an XML file named 7950-112573-02.xml (91,195 bytes), created on September 26, 2025, is herein incorporated by reference in its entirety.
[0010] 15
[0011] BACKGROUND
[0012] In the evolution of life, random mutations followed by natural selection is one of the key mechanisms that results in evolution of new functions. Typically, the mutational rates are significantly low so as not to compromise the fitness of the organism due to genome wide mutations. However, this also
[0013] 20 results in a relatively slow diversification of biomolecules and the subsequent evolution of new functions. For example, in a typical bacterial or human cell, the mutational frequency is around IO10to 10'9substitutions per base (Drake et al., Genetics 148, 1667-1686). This means that for a mutation to occur within a single gene of around 1000 bases, a cell would have to undergo more than a million rounds of cell division. In addition to this, the rate of evolution is further constrained by the number of mutations a
[0014] 25 genome can tolerate before the organism is not viable. This is especially important in the context of directed evolution experiments that rely on biomolecular diversity to evolve biomolecules (proteins, RNAs) with desired functions. To overcome the bottleneck of low mutation rates and genomic tolerance of mutations, directed evolution efforts have traditionally resorted to in vitro diversification of genes encoding corresponding biomolecules using randomized oligonucleotide sequences or error-prone polymerase chain reactions (PCR). This is followed by screening or selection methods for enriching and / or identifying variants resulting in a defined phenotype, followed by characterization to link the genotype to the phenotype. Given the process’s iterative requirements, each round of a directed evolution campaign requires generating biomolecular diversity in vitro before proceeding with screening and selection, rendering it a labor-intensive and time-consuming approach to profile the evolutionary landscape and identify biomolecules with desired
[0015] 35 properties (Hermes et al., Proc. Natl. Acad. Sci. 87, 696-700; Chen et al., Engineering new catalytic activities in enzymes. Nat. Catal. 3, 203-213; Winter et al., Making antibodies by phage display technology. PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0016] 7950-112573-02
[0017] Annu. Rev. Immunol. 12, 433^155; Xiao et al., Cold Spring Harb. Perspect. Biol. 8, a023945; Peisajovich et al., Nat. Methods 4, 991-994; Zeymer et al., Annu. Rev. Biochem. 87, 131-157).
[0018] Efforts have been made to use cell or virus-based approaches for continuous directed evolution (Morrison et al., Nat. Chem. Biol. 16, 610-619). Particularly, orthogonal replicons with high mutagenic
[0019] 5 rates have been established in model bacteria like E. coli (Tian et al., Science 383, 421-426) or model eukaryotes like yeasts (Ravikumar et al., Cell 175, 1946-1957), and in mammalian cells (Arzumanyan et al., ACS Synth. Biol. 7, 1722-1729). Genes encoding proteins of interests are encoded by the orthogonal replicons, and they are subjected to continuous mutagenesis by expressing error prone DNA polymerases that specifically replicate the orthogonal replication system and not the genomic DNA. This approach results in the continuous generation of a library of biomolecular variants within the cells without compromising their viability due to genomic mutations. The cells expressing a library of variants are screened under selection conditions to identify cells expressing variants resulting in a desired phenotype (e.g., antibiotic resistance, protein / protein binding). This step is typically followed by sequencing experiments to identify the mutant gene encoding the selected protein variant.
[0020] 15 Similar to orthogonal replicons, transcription coupled DNA modification enzymes (e.g., T7 RNA polymerases that have been fused to deaminases) have been engineered to create DNA mismatches at C:G base pairs, which upon repair result in the generation of biomolecular diversity (Moore et al., J. Am. Chem. Soc. 140, 11560-11564; Chen et al., Nat. Biotechnol. 38, 165-168; Hendel et al., Nat. Methods 78, 346- 357). In addition to these cell-based approaches, virus-based continuous evolution approaches have been
[0021] 20 developed (Esvelt et al., Nature 472, 499-503; Berman et al., J. Am. Chem. Soc. 140, 18093-18103; Jewel et al., Nat. Methods 20, 95-103; English et al., Cell 178, 748-761). In these platforms, either the viral genome or viral helper plasmids encode genes corresponding to the protein of interest. Biomolecular diversity is generated by either engineering mutagenic polymerases or by expressing mutagenic factors. These approaches typically link evolutionary outcomes (binding, enzyme activity) to viral propagation.
[0022] 25 Each of these approaches requires extensive engineering efforts and several of the continuous directed evolution efforts in mammalian cells are typically virus mediated approaches (Berman et al., J. Am. Chem. Soc. 140, 18093-18103; Jewel et al., Nat. Methods 20, 95-103; English et al., Cell 178, 748-761).
[0023] SUMMARY
[0024] Described herein are compositions (such as integration plasmids and CRISPR / Cas9 plasmids), kits, and methods for the rapid evolution of proteins of interest, such as antibodies or antibody fragments, by harnessing the inherent somatic hypermutation (SHM) machinery of B cells. Also described are plasmids and methods for displaying antibodies or antibody fragments on the surface of B cells.
[0025] Provided herein are integration plasmids capable of integrating at a specific genomic locus, such as a
[0026] 35 stable, non-immunoglobulin locus. In some aspects, the integration plasmid includes a heterologous gene operably linked to a first promoter; a SHM enhancer element upstream of the first promoter and the heterologous gene; a first nucleic acid sequence homologous to genomic sequences at the integration site, PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0027] 7950-112573-02 wherein the first nucleic acid sequence is upstream of the SHM enhancer element; and a second nucleic acid sequence homologous to genomic sequences at the integration site, wherein the second nucleic acid sequence is downstream of the heterologous gene and a polyadenylation (poly-A) signal sequence. In some examples, the heterologous gene of the integration plasmid encodes an antibody or an antibody fragment. In
[0028] 5 some examples, the heterologous gene further encodes a membrane localization signal, a transmembrane helix, and a linker between the antibody fragment and the transmembrane helix.
[0029] Also provided are kits that include an integration plasmid disclosed herein. In some aspects, the kits further include one or more of: a CRISPR / Cas9 plasmid encoding Cas9 and one or more guide RNAs that target the genomic locus; isolated B cells; B cell culture media; tissue culture flask(s); and buffer.
[0030] Further provided are isolated B cells, such as primary B cells or cells of a B cell line, that include an integration plasmid disclosed herein. The isolated B cells can be, for example, of human, avian or swine origin.
[0031] Also provided is a method of generating an antibody or antibody fragment that has higher affinity for a target antigen than a parental antibody or antibody fragment that binds the same antigen. In some
[0032] 15 aspects, the method includes transfecting isolated B cells with an integration plasmid disclosed herein and a CRISPR / Cas9 plasmid, wherein the CRISPR / Cas9 plasmid encodes Cas9 and one or more guide RNAs that target the genomic locus of the integration plasmid; enriching the B cells through antibiotic selection; and passaging the transfected B cells for multiple passages (e.g., at least three times, at least four times, at least five times, or at least six times). In some examples, the method further includes measuring affinity of the
[0033] 20 passaged B cells to the antigen. In some examples, the antibody fragment is a single-chain antibody binding fragment (Fab).
[0034] Further provided are methods of displaying an antibody or antibody fragment (such as an Fab fragment) on the surface of B cells. In some aspects, the method includes transfecting B cells with a nucleic acid molecule encoding in the 5’ to 3' direction: a membrane localization sequence; the antibody or antibody
[0035] 25 fragment; and a transmembrane helix, thereby displaying the antibody or antibody fragment on the surface of the B cells. In some examples, the nucleic acid molecule further includes a promoter, such as a human EFla promoter, and / or a coding sequence for a protein tag, such as a FLAG tag.
[0036] The foregoing and other features of this disclosure will become more apparent from the following detailed description of several aspects which proceeds with reference to the accompanying figures.
[0037] BRIEF DESCRIPTION OF THE DRAWINGS
[0038] FIGS. 1A-1E: Streamlining RAI transfection and enrichment workflow. (FIG. 1A) Transfection scheme for RAI using a two-plasmid system. (FIG. IB) Integration scheme of a homology cassette,
[0039] 35 following double-strand break (DSB) formation by encoded Cas9. (FIG. 1C) Selection scheme for transfected RAI cells containing puromycin resistance cassette, puroR. (FIG. ID) FACS density plot of PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0040] 7950-112573-02 gated eGFP+ cells in untransfected RAI and those enriched through puromycin selection. (FIG. IE) Imaging performed on enriched RAl-274-eGFP cells following puromycin selection. Scale bar indicates 100 pM.
[0041] FIGS. 2A-2G: Creation of an eGFP knockout (eGFP*) and subsequent restoration of fluorescence through SHM. (FIG. 2A) Scheme for introducing a C to T mutation into a codon in the chromophore
[0042] 5 responsible for fluorescence generation, resulting in a Thr to lie substitution. (FIG. 2B) Schematic of SHM function in RAl-proXIV-eGFP* cells. (FIG. 2C) Schematic of serial passaging and observation of emerging eGFP+ fluorescence. (FIG. 2D) Imaging of various different instances of eGFP+ cells within a single sample. (FIG. 2E) FACS density plot for eGFP+ cells, as sorted from the bulk population of an enriched RAl-274-proXIV-eGFP* sample. (FIG. 2F) Sunburst plot from NGS analysis shows introduction of various mutations across the sample. (FIG. 2G) Table detailing the most enriched nucleotide substitution obtained through PacBio NGS analysis.
[0043] FIGS. 3A-3D: Chimeric protein design of Fab display construct and function in vivo. (FIG. 3A) AlphaFold3 predicted structure of the fusion protein and visualized representations of the various components. (FIG. 3B) Schematic of Fabdisplay on the surface of RA-1 cells, utilizing fluorescent HA-
[0044] 15 mCherry and fluorescent anti-FLAG-APC mAbs to parameterize the chimeric protein into separate binding and expression variables. (FIG. 3C) Schematic of FACS experiment. (FIG. 3D) Density FACS plots of enriched RA1-CR9114 cells stained with anti-FLAG-APC and Hl-mCherry.
[0045] FIGS. 4A-4G: Workflow for the continuous directed evolution of higher-binding CR9114 variants. (FIG. 4A) Schematic for the incorporation of mutations into the gene encoding CR9114 Fab by iterative
[0046] 20 cycles of SHM. (FIG. 4B) Schematic for the iterative cycle of enriching cells, followed by diversification of the CR9114 gene through SHM that occurs during transcription, followed by selection of higher-binding variants and subsequent enrichment. (FIG. 4C) FACS density plots of the RA LSB3 gating strategy expressing (APC+) and binding (FITC+) channels, yielding a linear relationship from which higher-binding variants may be gated and sorted for growth. (FIG. 4D) FACS density plots for iterative sorting and
[0047] 25 passaging (expression vs binding) for continuously evolving RA 1-SB3 cell lines. For this verification experiment His-tagged H5 was used and FITC-conjugated anti-His antibody was used to detect binding of cells to His-tagged H5 (FIG. 4E) Bar graph representing the gate escape frequency of each iterative sort. (FIG. 4F) Histogram representing each iterative sort in terms of H5 binding. (FIG. 4G) Bar graph representing mean fluorescence intensity (F.I.) of the histogram.
[0048] FIGS. 5A-5E: Isolation of enriched CR9114 variants obtained by SHM and their experimental validation. (FIG. 5A) Experimental scheme for identification and characterizing higher-binding variants. (FIG. 5B) Heatmap showing density of mutations across the surface displayed CR9114 Fabafter 3 rounds of sorting and enrichment. (FIG. 5C) Table of read frequencies for the tested variants alongside their substituted codons, encoded amino acids, and EC50 values; variants were named based on the amino acid
[0049] 35 substitution found in the heavy chain of CR9114 and numbered based on the read frequency of their unique nucleotide substitution in the sequencing data - lower numbers indicate a higher degree enrichment. Point mutations listed for the Fab fragment also include at least one additional mutation, specifically A1601T PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0050] 7950-112573-02 and / or A1149G. (FIG. 5D) Bar chart representing the EC50 values of the 16 CR9114 heavy chain variants. (FIG. 5E) Titration curve for the two high-affinity CR9114 variants, W461R and S427P, relative to the CR9114 wild-type control.
[0051] FIG. 6: Microscopy images. Time course of selection for puromycin-resistant mutant cell lines
[0052] 5 expressing eGFP.
[0053] FIGS. 7A-7B: Genomic DNA analysis of the mutant cell lines. (FIG. 7A) PCR analysis to confirm the integration locus at the 5’-end of the integration site (expected size -2650 bp for RA1-SB1, 2860 bp for RA1-SB2, and -3900 bp for RA1-SB4). (FIG. 7B) PCR analysis to confirm the integration locus at the 3’ end of the integration site (expected size -3550 bp for RA1-SB1 and RA1-SB2, and -4500 bp for RAISES).
[0054] FIG. 8: Microscopy images for the eGFP* evolution experiments. Instances of eGFP* reversion.
[0055] FIGS. 9A-9E: Additional results from sequencing of eGFP* locus upon evolution experiments. (FIG. 9A) Optimization of PCR conditions. (FIG. 9B) Sanger sequencing chromatograms of bulk mixture that show an example of mutations. (FIG. 9C) Sunburst plot of eGFP* from cDNA sequenced using PacBio
[0056] 15 sequencing. (FIG. 9D) Histogram of nucleotide substitutions in eGFP* sequence. (FIG. 9E) Sunburst plot of total population, categorized by the quantity of mutations and which mutants were enriched (top); sunburst plot of mutated population categorized by the quantity of mutations and which mutants were enriched (bottom left); and sunburst plot of the most enriched mutants in the mutated sample, irrespective of their parent category (bottom right).
[0057] 20 FIGS. 10A-10B: FACS analysis of wild-type CR9114 binding variants of hemagglutinin. (FIG.
[0058] IOA) Density -corrected FACS dot plots for Expi 293 cells transfected with pSB3 and individual expression (left), binding (middle), and binding vs. expression (right) plots when incubated with Hl-mCherry. (FIG.
[0059] IOB) Density-corrected FACS dot plots for Expi 293 cells transfected with pSB3 and individual expression (left column), binding (middle column), and binding vs. expression (right column) plots when incubated
[0060] 25 with H3 (top row), H5 (middle row), and H7 (bottom row).
[0061] FIGS. 11A-11B: FACS sorting and detection of variant Y532F. (FIG. HA) Density-corrected FACS dot plots for enriched RA1-SB4 and their corresponding changes in the expression (anti-FLAG binding) during the course of the sort. Binding channel did not change magnitude, but expression channel shows population that decreased magnitude during the sort. (FIG. 1 IB) Chromatogram revealing the nucleotide substitution detected through its phenotype by FACS.
[0062] FIGS. 12A-12F: Mutational distribution of CR9114. (FIG. 12A) Light chain variable region. (FIG. 12B) Light chain constant region. (FIG. 12C) Sc60 linker region. (FIG. 12D) linker-FLAG-MHCI helix region (FIG. 12E) Heavy chain variable region. (FIG. 12F) Heavy chain constant region.
[0063] FIG. 13: Mutational distribution of CR9114 heavy chain, with highlighted mutants that showed
[0064] 35 lower EC50 values, relative to wild-type.
[0065] FIGS. 14A-14R: Analysis of H5 binding to evolved Fab variants. (FIG. 14A) Overlap of H5 binding of key variants. (FIGS. 14B-14R) Individual traces showing H5 binding of key variants: (FIG. 14B) PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0066] 7950-112573-02
[0067] Wile-type. (FIG. 14C) F394L. (FIG. 14D) M347V. (FIG. 14E) Q412R. (FIG. 14F) N330S. (FIG. 14G) D372E. (FIG. 14H) A475V. (FIG. 141) W461R. (FIG. 14J) V459I. (FIG. 14K) Y501H. (FIG. 14L) L477P. (FIG. 14M) S427P. (FIG. 14N) G343R. (FIG. 140) I373F. (FIG. 14P) G326D. (FIG. 14Q) F394S. (FIG. 14R) Q364R.
[0068] 5 FIGS. 15A-15B: (FIG. 15A) Titration curves comparing the binding of wild-type CR9114 and the W461R mutant towards other HA variants (Hl, H3, H7). (FIG. 15B) Bar graph representing their EC50 values.
[0069] FIG. 16: The crystal structure of CR9114, with highlighted residues of 1373 (left), M347 (middle), and Q412 (right), the most enriched heavy chain variable region variants with the lowest EC50 values.
[0070] FIGS. 17A-17D: Mutational analysis of unsorted, naive pool RAI cells engineered to express eGFP* after iterative rounds of passaging. (FIG. 17A) Phylogenetic tree showing representative nucleotide level mutations in the eGFP* locus. Mutations labelled “A” initially occur during an early passage and then accumulate associated mutations in late passages. Mutations labeled “B” are exclusive to the late passage and do not occur in the early passage. (FIG. 17B) Integral plot of occurrences over the sequence length of
[0071] 15 eGFP*. (FIG. 17C) Bar plot comparison of nucleotide mutations in the eGFP* sequence at a late passage. Mutations are plotted as the percentage of total reads / sum of all mutations. (FIG. 17D) Heatmap showing density of mutations across the eGFP* sequence after 8 passages.
[0072] FIGS. 18A-18D: Mechanistic studies that resulted in the identification of additional SHM recruiting sequences. (FIG. 18A) Genomic locus corresponding to IgHV(4-34) from proXIV-2, proXIV-2-862,
[0073] 20 proXIV-2-1417 and proXIV-2-936. (FIG. 18B) Flow cytometry analysis density plots of the transfected RAI cells. (FIG. 18C) FACS density plots of evolved cells that have either proXIV-2 obtained from IgHV(4-34) or proXIV-1 obtained from IgHV(4-55) upstream of eGFP* sequence. (FIG. 18D) Bar chart representing the average event counts per 500,000 events for each sequence at day 13. A: EFla-eGFP*, B: proXIV(4-34), which is proXIV-2, C: proXIV(4-55), which is proXIV-1, D: proXIV-2-862-EFla-eGFP*, E:
[0074] 25 proXIV-2-1417-EFla-eGFP*, F: proXIV-2-936-EFla-eGFP*.
[0075] FIGS. 19A-19E: Viral neutralization validation of CR9114 mutants of interest generated using CODE-HB. (FIG. 19A) Schematic of the hemagglutination-inhibition (HAI) assay workflow. (FIG. 19B) Hemagglutination assay for CR9114 variants that exhibited ELISA binding, with corresponding HAI titers required for inhibition. (FIG. 19C) Workflow schematic for influenza neutralization assays. (FIG. 19D) Imaging data showing neutralization of A / Puerto Rico / 1934 (H1N1) following incubation with purified variants CR9114 mAbs. (FIG. 19E) Antibody dose response experiments for viral neutralization assays demonstrating improved efficacy of W154R variant of CR9114 as compared to wild type CR9114.
[0076] FIGS. 20A-20E: Directed evolution of the lateral-patch-binding Fab 047-09_lA02 to improve binding against contemporary variants of H1N1 influenza. (FIG. 20A) Overview of 047-09_lA02 evolution
[0077] 35 using CODE-HB. (FIG. 20B) Evolutionary trajectory after iterative FACS sorting: binding to Hl / Michigan / 2015 plotted against surface expression of the Fab, illustrating an approximately linear relationship. (FIG. 20C) Heat map showing mutation density across the 047-09_lA02 Fab sequence. (FIG. PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0078] 7950-112573-02
[0079] 20D) Recunent in-frame deletions (span shown) and insertions (anchor nucleotide indicated) displayed alongside frequent single-nucleotide substitutions; reads with more than two amino-acid substitutions are also highlighted. Counts are normalized to 600,000 reads. (SEQ ID NOs. 56-58) (FIG. 20E) FACS-based
[0080] Hl / Michigan / 2015 titration comparing the wild-type 047-09_lA02 Fab with a variant containing a 7-amino-
[0081] 5 acid deletion (positions 36-42) in the sc60 linker between the light and heavy chains.
[0082] SEQUENCES
[0083] The nucleic acid and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotide bases, and single letter code for amino acids, as defined in
[0084] 10 37 C.F.R. 1.822. Only one strand of each nucleic acid sequence is shown, but the complementary strand is understood as included by any reference to the displayed strand. In the accompanying sequence listing:
[0085] SEQ ID NO: 1 is the nucleic acid sequence of a somatic hypermutation (SHM) enhancer element (proXIV-1). ttctcagagggcacagccagcatacacctcccagggtgagcccaaaagactggggcctccctcatccctttttacctatccatacaaaggcaccacccacatg
[0086] 15 caaatcctcacttaggcacccacaggaaatgactacacatttccttaaattcagggtccagctcacatgggaagtgctttctgagagtcatggacctcctgcaca agaac
[0087] SEQ ID NO: 2 is the nucleic acid sequence of plasmid pSB4. ctgtgtgaaattgttatccgctcacaattccacacaacatacgagccggaagcataaagtgtaaagcctggggtgcctaatgagtgagctaactcacattaattg
[0088] 20 cgttgcgctcactgcccgctttccagtcgggaaacctgtcgtgccagctgcattaatgaatcggccaacgcgcggggagaggcggtttgcgtattgggcgct cttccgcttcctcgctcactgactcgctgcgctcggtcgttcggctgcggcgagcggtatcagctcactcaaaggcggtaatacggttatccacagaatcagg ggataacgcaggaaagaacatgtgagcaaaaggccagcaaaaggccaggaaccgtaaaaaggccgcgttgctggcgtttttccataggctccgcccccct gacgagcatcacaaaaatcgacgctcaagtcagaggtggcgaaacccgacaggactataaagataccaggcgtttccccctggaagctccctcgtgcgctc tcctgttccgaccctgccgcttaccggatacctgtccgcctttctcccttcgggaagcgtggcgctttctcatagctcacgctgtaggtatctcagttcggtgtag
[0089] 25 gtcgttcgctccaagctgggctgtgtgcacgaaccccccgttcagcccgaccgctgcgccttatccggtaactatcgtcttgagtccaacccggtaagacacg acttatcgccactggcagcagccactggtaacaggattagcagagcgaggtatgtaggcggtgctacagagttcttgaagtggtggcctaactacggctaca ctagaagaacagtatttggtatctgcgctctgctgaagccagttaccttcggaaaaagagttggtagctcttgatccggcaaacaaaccaccgctggtagcggt ggtttttttgtttgcaagcagcagattacgcgcagaaaaaaaggatctcaagaagatcctttgatcttttctacggggtctgacgctcagtggaacgaaaactcac gttaagggattttggtcatgagattatcaaaaaggatcttcacctagatccttttaaattaaaaatgaagttttaaatcaatctaaagtatatatgagtaaacttggtct
[0090] 30 gacagttaccaatgcttaatcagtgaggcacctatctcagcgatctgtctatttcgttcatccatagttgcctgactccccgtcgtgtagataactacgatacggga gggcttaccatctggccccagtgctgcaatgataccgcgagacccacgctcaccggctccagatttatcagcaataaaccagccagccggaagggccgag cgcagaagtggtcctgcaactttatccgcctccatccagtctattaattgttgccgggaagctagagtaagtagttcgccagttaatagtttgcgcaacgttgttgc cattgctacaggcatcgtggtgtcacgctcgtcgtttggtatggcttcattcagctccggttcccaacgatcaaggcgagttacatgatcccccatgttgtgcaaa aaagcggttagctccttcggtcctccgatcgttgtcagaagtaagttggccgcagtgttatcactcatggttatggcagcactgcataattctcttactgtcatgcc
[0091] 35 atccgtaagatgcttttctgtgactggtgagtactcaaccaagtcattctgagaatagtgtatgcggcgaccgagttgctcttgcccggcgtcaatacgggataat accgcgccacatagcagaactttaaaagtgctcatcattggaaaacgttcttcggggcgaaaactctcaaggatcttaccgctgttgagatccagttcgatgtaa cccactcgtgcacccaactgatcttcagcatcttttactttcaccagcgtttctgggtgagcaaaaacaggaaggcaaaatgccgcaaaaaagggaataaggg cgacacggaaatgttgaatactcatactcttcctttttcaatattattgaagcatttatcagggttattgtctcatgagcggatacatatttgaatgtatttagaaaaata aacaaataggggttccgcgcacatttccccgaaaagtgccacctgacgcgccctgtagcggcgcattaagcgcggcgggtgtggtggttacgcgcagcgt
[0092] 40 gaccgctacacttgccagcgccctagcgcccgctcctttcgctttcttcccttcctttctcgccacgttcgccggctttccccgtcaagctctaaatcgggggctc cctttagggttccgatttagtgctttacggcacctcgaccccaaaaaacttgattagggtgatggttcacgtagtgggccatcgccctgatagacggtttttcgcc ctttgacgttggagtccacgttctttaatagtggactcttgttccaaactggaacaacactcaaccctatctcggtctattcttttgatttataagggattttgccgattt cggcctattggttaaaaaatgagctgatttaacaaaaatttaacgcgaattttaacaaaatattaacgcttacaatttccattcgccattcaggctgcgcaactgttg ggaagggcgatcggtgcgggcctcttcgctattacgccagctggcgaaagggggatgtgctgcaaggcgattaagttgggtaacgccagggttttcccagt
[0093] 45 cacgacgttgtaaaacgacggccagtgccaagctgatctatacattgaatcaatattggcaattagccatattagtcattggttatatagcataaatcaatattggc tattggccattgcatacgtttcttgtgaactcatctcgtgcctgcaagacatcaaaagtttctccaaagacatcccttcctctctgccctcatcctatatcaaggtctt agctccttgaagactgcattaatgtgtcttttgccttcatcccccccccaccctcccgcttcctctagtccttgtgtttaccaaaatacttttgctaaaacctgtatgtct agcttctgctgatactcttagcaatactcccttatttccttctggcgccatttgccaatcaccagctaatggctttgctctttgcatggtacctgtttctgtcctactagc PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0094] 7950-112573-02 ttctagtctgaaaaggaagcctttactagcagcttcaacccagtcctccttaccttgcacctattaacgttggttcaacatgaaggggaagtaagataagtggaa gaggatacacaggaaaggtggaactggattcaggtatctgtgttctgccattgacacttgggactgcgttccagcttccacctgtggaaaccagaaagcatttg ccaaaaattctaaacagagggtctatttctgaagctgaggaatcacatggagtgaatagcatgggggaaggggattactaatgaggctgaaagaaaagtacc aggaagagactgacttcatgcataatccataaagccaggttttcttgaaaggtgacctgctttcttgtattcctgcaagaaacggctcgaggttctcagagggca
[0095] 5 cagccagcatacacctcccagggtgagcccaaaagactggggcctccctcatccctttttacctatccatacaaaggcaccacccacatgcaaatcctcactt aggcacccacaggaaatgactacacatttccttaaattcagggtccagctcacatgggaagtgctttctgagagtcatggacctcctgcacaagaacggctcc ggtgcccgtcagtgggcagagcgcacatcgcccacagtccccgagaagttgtggggaggggtcggcaattgaagcggtgcctagagaaggtggcgcgg ggtaaactgggaaagtgatgtcgtgtactggctccgcctttttcccgagggtgggggagaaccgtatataagtgcagtagtcgccgtgaacgttctttttcgcaa cgggtttgccgccagaacacaggtaagtgccgtgtgtggttcccgcgggcctggcctctttacgggttatggcccttgcgtgccttgaattacttccacctggct
[0096] 10 gcagtacgtgattcttgatcccgagcttcgggttggaagtgggtgggagagttcgaggccttgcgcttaaggagccccttcgcctcgtgcttgagttgaggcct ggcctgggcgctggggccgccgcgtgcgaatctggtggcaccttcgcgcctgtctcgctgctttcgataagtctctagccatttaaaatttttgatgacctgctg cgacgctttttttctggcaagatagtcttgtaaatgcgggccaagatctgcacactggtatttcggtttttggggccgcgggcggcgacggggcccgtgcgtcc cagcgcacatgttcggcgaggcggggcctgcgagcgcggccaccgagaatcggacgggggtagtctcaagctggccggcctgctctggtgcctggcctc gcgccgccgtgtatcgccccgccctgggcggcaaggctggcccggtcggcaccagttgcgtgagcggaaagatggccgcttcccggccctgctgcagg
[0097] 15 gagctcaaaatggaggacgcggcgctcgggagagcgggcgggtgagtcacccacacaaaggaaaagggcctttccgtcctcagccgtcgcttcatgtga ctccacggagtaccgggcgccgtccaggcacctcgattagttctcgagcttttggagtacgtcgtctttaggttggggggaggggttttatgcgatggagtttcc ccacactgagtgggtggagactgaagttaggccagcttggcacttgatgtaattctccttggaatttgccctttttgagtttggatcttggttcattctcaagcctca gacagtggttcaaagtttttttcttccatttcaggtgtcgtgaggggatcgtcgaccgtacggccaccatgAAGAAGAACATAGCGTTTTTG CTCGCCTTGATGTTTGTTTTTTCAATAGCAACTAATGCCTACGCTCAATCAGCTCTCACGCAACC
[0098] 20 GCCAGCTGTTTCCGGTACCCCAGGTCAACGCGTTACCATATCATGTAGCGGCTCTGACTCAAAT ATTGGTCGCCGCTCCGTAAACTGGTACCAACAGTTTCCCGGGACAGCTCCGAAACTCCTGATCT ATAGCAACGATCAACGCCCGTCAGTGGTACCAGACAGGTTTAGTGGCTCTAAATCCGGAACAT CAGCTTCTCTCGCTATCAGCGGTCTCCAATCAGAAGACGAAGCTGAGTACTACTGCGCAGCTTG GGATGACTCCTTGAAGGGAGCTGTCTTCGGCGGGGGTACTCAGCTTACAGTGCTTGGACAGCC
[0099] 25 CAAAGCCGCCCCTTCCGTAACTCTTTTCCCCCCATCAAGCGAGGAACTTCAAGCAAATAAGGCA ACCCTTGTCTGCCTGATTAGTGATTTTTATCCGGGTGCCGTTACAGTGGCATGGAAAGCAGACT CTAGTCCAGTCAAGGCAGGAGTTGAGACAACTACGCCATCCAAGCAATCCAACAACAAATACG CTGCCTCAAGTTATTTGAGCCTTACCCCAGAGCAATGGAAAAGTCATCGCTCATATTCTTGTCA GGTCACACATGAAGGCAGCACAGTGGAAAAGACGGTTGCGCCAACGGAATGTTCCGGGGGGA
[0100] 30 GTTCCGGTAGTGGTTCCGGTTCCACGGGTACCTCCTCCTCAGGTACCGGGACTTCCGCGGGCAC AACGGGAACCAGTGCCTCTACTTCCGGCTCAGGAAGTGGAGGCGGAGGGGGGAGCGGAGGCG GTGGATCCGCGGGCGGAACCGCTACCGCTGGGGCGTCATCCGGATCTCAAGTACAGCTCGTTC AATCAGGAGCTGAGGTGAAAAAACCCGGTTCTTCTGTAAAAGTTAGCTGCAAATCATCTGGAG GCACTAGCAACAACTACGCTATAAGTTGGGTACGACAAGCTCCAGGGCAAGGTCTGGACTGGA
[0101] 35 TGGGAGGCATTTCACCGATATTCGGGAGCACAGCTTACGCTCAGAAGTTTCAGGGCAGGGTAA CCATCTCAGCTGATATTTTTTCTAACACAGCTTACATGGAACTCAATAGTCTCACATCAGAAGA CACAGCTGTGTACTTTTGTGCTAGGCATGGCAATTATTACTACTACAGTGGCATGGATGTATGG GGCCAGGGAACAACAGTTACAGTCTCTTCAGCCAGTACCAAGGGGCCCTCCGTCTTTCCTCTCG CCCCGAGTAGCAAAAGCACTTCAGGGGGGACTGCAGCACTCGGATGCCTGGTTAAAGATTACT
[0102] 40 TTCCAGAACCTGTCACCGTATCCTGGAATAGTGGCGCTCTCACTAGCGGCGTACATACGTTCCC GGCGGTCCTCCAGAGTAGCGGTCTGTATAGTTTGAGTTCTGTCGTAACGGTTCCCTCATCTTCTC TGGGTACACAGACCTATATCTGCAATGTGAATCATAAACCATCAAATACGAAGGTTGACAAAC GAGTAgctagctccggaggcagcggatccggaggcagcggaGATTACAAAGATGACGATGATAAAggctctggagcctccg gaggcagcggaggaagtgctatccccatcatgggtatcgttgctggcctggttgtccttgcagctgtagtcactggagctgcggtcgctgctgtgctgtggag
[0103] 45 aaagaagagctcagattaaagcggccgcactcctcaggtgcaggctgcctatcagaaggtggtggctggtgtggccaatgccctggctcacaaataccact gagatctttttccctctgccaaaaattatggggacatcatgaagccccttgagcatctgacttctggctaataaaggaaatttattttcattgcaatagtgtgttggaa ttttttgtgtctctcactcggaaggacatatgggagggcgatcgccagtactagtgaacctcttcgagggacctaataacttcgtatagcatacattatacgaagtt atattaagggttccggatctcgacctcgaaattctaccgggtaggggaggcgcttttcccaaggcagtctggagcatgcgctttagcagccccgctgggcact tggcgctacacaagtggcctctggcctcgcacacattccacatccaccggtaggcgccaaccggctccgttctttggtggccccttcgcgccaccttctactcc
[0104] 50 tcccctagtcaggaagttcccccccgccccgcagctcgcgtcgtgcaggacgtgacaaatggaagtagcacgtctcactagtctcgtgcagatggacagca ccgctgagcaatggaagcgggtaggcctttggggcagcggccaatagcagctttgctccttcgctttctgggctcagaggctgggaaggggtgggtccggg ggcgggctcaggggcgggctcaggggcggggcgggcgcccgaaggtcctccggaggcccggcattctgcacgcttcaaaagcgcacgtctgccgcgc tgttctcctcttcctcatctccgggcctttcgacctgcatccatctagatctcgagcagctgaagcttaccatgaccgagtacaagcccacggtgcgcctcgcca cccgcgacgacgtccccagggccgtacgcaccctcgccgccgcgttcgccgactaccccgccacgcgccacaccgtcgatccggaccgccacatcgag
[0105] 55 cgggtcaccgagctgcaagaactcttcctcacgcgcgtcgggctcgacatcggcaaggtgtgggtcgcggacgacggcgccgcggtggcggtctggacc acgccggagagcgtcgaagcgggggcggtgttcgccgagatcggcccgcgcatggccgagttgagcggttcccggctggccgcgcagcaacagatgg PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0106] 7950-112573-02 aaggcctcctggcgccgcaccggcccaaggagcccgcgtggttcctggccaccgtcggcgtctcgcccgaccaccagggcaagggtctgggcagcgcc gtcgtgctccccggagtggaggcggccgagcgcgccggggtgcccgccttcctggagacctccgcgccccgcaacctccccttctacgagcggctcggc ttcaccgtcaccgccgacgtcgaggtgcccgaaggaccgcgcacctggtgcatgacccgcaagcccggtgcctgacgcccgccccacgacccgcagcg cccgaccgaaaggagcgcacgaccccatgcatcgatgatatcagatccccgggatgcagaaattgatgatctattaaacaataaagatgtccactaaaatgg
[0107] 5 aagtttttcctgtcatactttgttaagaagggtgagaacagagtacctacattttgaatggaaggattggagctacgggggtgggggtggggtgggattagata aatgcctgctctttactgaaggctctttactattgctttatgataatgtttcatagttggatatcataatttaaacaagcaaaaccaaattaagggccagctcattcctc ccactcatgatctatagatctatagatctctcgtgggatcattgtttttctcttgattcccactttgtggttctaagtactgtggtttccaaatgtgtcagtttcatagcct gaagaacgagatcagcagcctctgttccacatacacttcattctcagtattgttttgccaagttctaattccatcagaagctggtcgagatcctaagcttggctgga cgtaaactcctcttcagacctaataacttcgtatagcatacattatacgaagttatattaagggttattgaatatgatcggaattgcggccctgacctgttggggtctt
[0108] 10 taaagctcaaggaaaaaggccatagttgatttctcctaaatcaagatagagtccaattaacttttttttttttttaagatggagtctcactctgtcgccaggctggact gcagtggcgtgatctcagctcactgcaacctctgcctcccaggttcaggcgattcttctgtctcagcctcccgagtagctgggactacaagtgtgggccaccat gcccagctagttttgtatttttagtagagatgggatttcaccatcttggctaggatggtcttgatctcttgaccttgtgattcgcctgcctctgcctcccagagtgctg agattacaggcgtgagccacagtgcccaatttccatgggttttcaagaaaaacttaacttacttaaacttcagatcacttccttaaaggaacatactgagcatagc ttggctggttagaaattaggaagactagcttgggctacatggtaagaccctgcctctacaaaaaataagaaaaaagttagtggcatgtgcctgtagttccatcta
[0109] 15 cttgggaggctgaggtgagaggatcgcttgagcccaggaggttgaagctgcagtgagccacgattgcaccactgtactccagcctgggtgacacagagtg agaccctgtctccaaacaaaattaggaagggtttcagagaggaaataaacacggccg
[0110] SEQ ID NO: 3 is the nucleic acid sequence of plasmid p276. gagggcctatttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatattagtacaaaat
[0111] 20 acgtgacgtagaaagtaataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgcttaccgtaacttgaaagtatttcgatttcttggctt tatatatcttgtggaaaggacgaaacaccgccaacaggtcagtttatacgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaaa agtggcaccgagtcggtgcttttttgttttagagctagaaatagcaagttaaaataaggctagtccgtttttagcgcgtgcgccaattctgcagacaaatggctct agagagggcctatttcccatgattccttcatatttgcatatacgatacaaggctgttagagagataattggaattaatttgactgtaaacacaaagatattagtacaa aatacgtgacgtagaaagtaataatttcttgggtagtttgcagttttaaaattatgttttaaaatggactatcatatgcttaccgtaacttgaaagtatttcgatttcttgg
[0112] 25 ctttatatatcttgtggaaaggacgaaacaccgattcctgcaagtacacggcgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga aaaagtggcaccgagtcggtgcttttttgttttagagctagaaatagcaagttaaaataaggctagtccgtttttagcgcgtgcgccaattctgcagacaaatggg gtacccgttacataacttacggtaaatggcccgcctggctgaccgcccaacgacccccgcccattgacgtcaatagtaacgccaatagggactttccattgac gtcaatgggtggagtatttacggtaaactgcccacttggcagtacatcaagtgtatcatatgccaagtacgccccctattgacgtcaatgacggtaaatggccc gcctggcattgtgcccagtacatgaccttatgggactttcctacttggcagtacatctacgtattagtcatcgctattaccatggtcgaggtgagccccacgttctg
[0113] 30 cttcactctccccatctcccccccctccccacccccaattttgtatttatttattttttaattattttgtgcagcgatgggggcgggggggggggggggggcgsscc mcgggggggggggggggggggsgsgsgssaggcggggcggggcggggcgaggggcggggcggggcgaggcggagaggtgcggcggcagcc aatcagagcggcgcgctccgaaagtttccttttatggcgaggcggcggcggcggcggccctataaaaagcgaagcgcgcggcgggcgggagtcgctgc gcgctgccttcgccccgtgccccgctccgccgccgcctcgcgccgcccgccccggctctgactgaccgcgttactcccacaggtgagcgggcgggacgg cccttctcctccgggctgtaattagctgagcaagaggtaagggtttaagggatggttggttggtggggtattaatgtttaattacctggagcacctgcctgaaatc
[0114] 35 actttttttcaggttggaccggtgccaccatggactataaggaccacgacggagactacaaggatcatgatattgattacaaagacgatgacgataagatggcc ccaaagaagaagcggaaggtcggtatccacggagtcccagcagccgacaagaagtacagcatcggcctggacatcggcaccaactctgtgggctgggc cgtgatcaccgacgagtacaaggtgcccagcaagaaattcaaggtgctgggcaacaccgaccggcacagcatcaagaagaacctgatcggagccctgct gttcgacagcggcgaaacagccgaggccacccggctgaagagaaccgccagaagaagatacaccagacggaagaaccggatctgctatctgcaagaga tcttcagcaacgagatggccaaggtggacgacagcttcttccacagactggaagagtccttcctggtggaagaggataagaagcacgagcggcaccccatc
[0115] 40 ttcggcaacatcgtggacgaggtggcctaccacgagaagtaccccaccatctaccacctgagaaagaaactggtggacagcaccgacaaggccgacctg cggctgatctatctggccctggcccacatgatcaagttccggggccacttcctgatcgagggcgacctgaaccccgacaacagcgacgtggacaagctgtt catccagctggtgcagacctacaaccagctgttcgaggaaaaccccatcaacgccagcggcgtggacgccaaggccatcctgtctgccagactgagcaag agcagacggctggaaaatctgatcgcccagctgcccggcgagaagaagaatggcctgttcggaaacctgattgccctgagcctgggcctgacccccaactt caagagcaacttcgacctggccgaggatgccaaactgcagctgagcaaggacacctacgacgacgacctggacaacctgctggcccagatcggcgacca
[0116] 45 gtacgccgacctgtttctggccgccaagaacctgtccgacgccatcctgctgagcgacatcctgagagtgaacaccgagatcaccaaggcccccctgagcg cctctatgatcaagagatacgacgagcaccaccaggacctgaccctgctgaaagctctcgtgcggcagcagctgcctgagaagtacaaagagattttcttcg accagagcaagaacggctacgccggctacattgacggcggagccagccaggaagagttctacaagttcatcaagcccatcctggaaaagatggacggca ccgaggaactgctcgtgaagctgaacagagaggacctgctgcggaagcagcggaccttcgacaacggcagcatcccccaccagatccacctgggagag ctgcacgccattctgcggcggcaggaagatttttacccattcctgaaggacaaccgggaaaagatcgagaagatcctgaccttccgcatcccctactacgtgg
[0117] 50 gccctctggccaggggaaacagcagattcgcctggatgaccagaaagagcgaggaaaccatcaccccctggaacttcgaggaagtggtggacaagggc gcttccgcccagagcttcatcgagcggatgaccaacttcgataagaacctgcccaacgagaaggtgctgcccaagcacagcctgctgtacgagtacttcacc gtgtataacgagctgaccaaagtgaaatacgtgaccgagggaatgagaaagcccgccttcctgagcggcgagcagaaaaaggccatcgtggacctgctgt tcaagaccaaccggaaagtgaccgtgaagcagctgaaagaggactacttcaagaaaatcgagtgcttcgactccgtggaaatctccggcgtggaagatcgg ttcaacgcctccctgggcacataccacgatctgctgaaaattatcaaggacaaggacttcctggacaatgaggaaaacgaggacattctggaagatatcgtgc
[0118] 55 tgaccctgacactgtttgaggacagagagatgatcgaggaacggctgaaaacctatgcccacctgttcgacgacaaagtgatgaagcagctgaagcggcg PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0119] 7950-112573-02 gagatacaccggctggggcaggctgagccggaagctgatcaacggcatccgggacaagcagtccggcaagacaatcctggatttcctgaagtccgacgg cttcgccaacagaaacttcatgcagctgatccacgacgacagcctgacctttaaagaggacatccagaaagcccaggtgtccggccagggcgatagcctgc acgagcacattgccaatctggccggcagccccgccattaagaagggcatcctgcagacagtgaaggtggtggacgagctcgtgaaagtgatgggccggc acaagcccgagaacatcgtgatcgaaatggccagagagaaccagaccacccagaagggacagaagaacagccgcgagagaatgaagcggatcgaag
[0120] 5 agggcatcaaagagctgggcagccagatcctgaaagaacaccccgtggaaaacacccagctgcagaacgagaagctgtacctgtactacctgcagaatg ggcgggatatgtacgtggaccaggaactggacatcaaccggctgtccgactacgatgtggaccatatcgtgcctcagagctttctggccgacgactccatcg acaacaaggtgctgaccagaagcgacaagaaccggggcaagagcgacaacgtgccctccgaagaggtcgtgaagaagatgaagaactactggcggca gctgctgaacgccaagctgattacccagagaaagttcgacaatctgaccaaggccgagagaggcggcctgagcgaactggataaggccggcttcatcaag agacagctggtggaaacccggcagatcacaaagcacgtggcacagatcctggactcccggatgaacactaagtacgacgagaatgacaagctgatccgg
[0121] 10 gaagtgaaagtgatcaccctgaagtccaagctggtgtccgatttccggaaggatttccagttttacaaagtgcgcgagatcaacaactaccaccacgcccacg acgcctacctgaacgccgtcgtgggaaccgccctgatcaaaaagtaccctgcgctggaaagcgagttcgtgtacggcgactacaaggtgtacgacgtgcg gaagatgatcgccaagagcgagcaggaaatcggcaaggctaccgccaagtacttcttctacagcaacatcatgaactttttcaagaccgagattaccctggc caacggcgagatccggaaggcgcctctgatcgagacaaacggcgaaaccggggagatcgtgtgggataagggccgggattttgccaccgtgcggaaag tgctgagcatgccccaagtgaatatcgtgaaaaagaccgaggtgcagacaggcggcttcagcaaagagtctatcctgcccaagaggaacagcgataagct
[0122] 15 gatcgccagaaagaaggactgggaccctaagaagtacggcggcttcgacagccccaccgtggcctattctgtgctggtggtggccaaagtggaaaagggc aagtccaagaaactgaagagtgtgaaagagctgctggggatcaccatcatggaaagaagcagcttcgagaagaatcccatcgactttctggaagccaaggg ctacaaagaagtgaaaaaggacctgatcatcaagctgcctaagtactccctgttcgagctggaaaacggccggaagagaatgctggcctctgccggcgaac tgcagaagggaaacgaactggccctgccctccaaatatgtgaacttcctgtacctggccagccactatgagaagctgaagggctcccccgaggataatgag cagaaacagctgtttgtggaacagcacaagcactacctggacgagatcatcgagcagatcagcgagttctccaagagagtgatcctggccgacgctaatctg
[0123] 20 gacaaagtgctgtccgcctacaacaagcaccgggataagcccatcagagagcaggccgagaatatcatccacctgtttaccctgaccaatctgggagcccc tgccgccttcaagtactttgacaccaccatcgaccggaagaggtacaccagcaccaaagaggtgctggacgccaccctgatccaccagagcatcaccggc ctgtacgagacacggatcgacctgtctcagctgggaggcgacaaaaggccggcggccacgaaaaaggccggccaggcaaaaaagaaaaagtaagaatt cctagagctcgctgatcagcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactg tcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattggga
[0124] 25 agagaatagcaggcatgctggggagcggccgcaggaacccctagtgatggagttggccactccctctctgcgcgctcgctcgctcactgaggccgggcg accaaaggtcgcccgacgcccgggctttgcccgggcggcctcagtgagcgagcgagcgcgcagctgcctgcaggggcgcctgatgcggtattttctcctt acgcatctgtgcggtatttcacaccgcatacgtcaaagcaaccatagtacgcgccctgtagcggcgcattaagcgcggcgggtgtggtggttacgcgcagc gtgaccgctacacttgccagcgccttagcgcccgctcctttcgctttcttcccttcctttctcgccacgttcgccggctttccccgtcaagctctaaatcgggggct ccctttagggttccgatttagtgctttacggcacctcgaccccaaaaaacttgatttgggtgatggttcacgtagtgggccatcgccctgatagacggtttttcgc
[0125] 30 cctttgacgttggagtccacgttctttaatagtggactcttgttccaaactggaacaacactcaactctatctcgggctattcttttgatttataagggattttgccgatt tcggtctattggttaaaaaatgagctgatttaacaaaaatttaacgcgaattttaacaaaatattaacgtttacaattttatggtgcactctcagtacaatctgctctga tgccgcatagttaagccagccccgacacccgccaacacccgctgacgcgccctgacgggcttgtctgctcccggcatccgcttacagacaagctgtgaccg tctccgggagctgcatgtgtcagaggttttcaccgtcatcaccgaaacgcgcgagacgaaagggcctcgtgatacgcctatttttataggttaatgtcatgataa taatggtttcttagacgtcaggtggcacttttcggggaaatgtgcgcggaacccctatttgtttatttttctaaatacattcaaatatgtatccgctcatgagacaata
[0126] 35 accctgataaatgcttcaataatattgaaaaaggaagagtatgagtattcaacatttccgtgtcgcccttattcccttttttgcggcattttgccttcctgtttttgctca cccagaaacgctggtgaaagtaaaagatgctgaagatcagttgggtgcacgagtgggttacatcgaactggatctcaacagcggtaagatccttgagagtttt cgccccgaagaacgttttccaatgatgagcacttttaaagttctgctatgtggcgcggtattatcccgtattgacgccgggcaagagcaactcggtcgccgcat acactattctcagaatgacttggttgagtactcaccagtcacagaaaagcatcttacggatggcatgacagtaagagaattatgcagtgctgccataaccatga gtgataacactgcggccaacttacttctgacaacgatcggaggaccgaaggagctaaccgcttttttgcacaacatgggggatcatgtaactcgccttgatcgt
[0127] 40 tgggaaccggagctgaatgaagccctaccaaacgacgagcgtgacaccacgatgcctgtagcaatggcaacaacgttgcgcaaactattaactggcgaact acttactctagcttcccggcaacaattaatagactggatggaggcggataaagttgcaggaccacttctgcgctcggcccttccggctggctggtttattgctga taaatctggagccggtgagcgtggaagccgcggtatcattgcagcactggggccagatggtaagccctcccgtatcgtagttatctacacgacggggagtca ggcaactatggatgaacgaaatagacagatcgctgagataggtgcctcactgattaagcattggtaactgtcagaccaagtttactcatatatactttagattgatt taaaacttcatttttaatttaaaaggatctaggtgaagatcctttttgataatctcatgaccaaaatcccttaacgtgagttttcgttccactgagcgtcagaccccgta
[0128] 45 gaaaagatcaaaggatcttcttgagatcctttttttctgcgcgtaatctgctgcttgcaaacaaaaaaaccaccgctaccagcggtggtttgtttgccggatcaag agctaccaactctttttccgaaggtaactggcttcagcagagcgcagataccaaatactgttcttctagtgtagccgtagttaggccaccacttcaagaactctgt agcaccgcctacatacctcgctctgctaatcctgttaccagtggctgctgccagtggcgataagtcgtgtcttaccgggttggactcaagacgatagttaccgg ataaggcgcagcggtcgggctgaacggggggttcgtgcacacagcccagcttggagcgaacgacctacaccgaactgagatacctacagcgtgagctat gagaaagcgccacgcttcccgaagggagaaaggcggacaggtatccggtaagcggcagggtcggaacaggagagcgcacgagggagcttccagggg
[0129] 50 gaaacgcctggtatctttatagtcctgtcgggtttcgccacctctgacttgagcgtcgatttttgtgatgctcgtcaggggggcggagcctatggaaaaacgcca gcaacgcggcctttttacggttcctggccttttgctggccttttgctcacatgt
[0130] SEQ ID NO: 4 is the nucleic acid sequence of a heterologous gene encoding Fab CR9114. atgAAGAAGAACATAGCGTTTTTGCTCGCCTTGATGTTTGTTTTTTCAATAGCAACTAATGCCTAC
[0131] 55 GCTCAATCAGCTCTCACGCAACCGCCAGCTGTTTCCGGTACCCCAGGTCAACGCGTTACCATAT PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0132] 7950-112573-02
[0133] CATGTAGCGGCTCTGACTCAAATATTGGTCGCCGCTCCGTAAACTGGTACCAACAGTTTCCCGG GACAGCTCCGAAACTCCTGATCTATAGCAACGATCAACGCCCGTCAGTGGTACCAGACAGGTT TAGTGGCTCTAAATCCGGAACATCAGCTTCTCTCGCTATCAGCGGTCTCCAATCAGAAGACGAA GCTGAGTACTACTGCGCAGCTTGGGATGACTCCTTGAAGGGAGCTGTCTTCGGCGGGGGTACTC
[0134] 5 AGCTTACAGTGCTTGGACAGCCCAAAGCCGCCCCTTCCGTAACTCTTTTCCCCCCATCAAGCGA
[0135] GGAACTTCAAGCAAATAAGGCAACCCTTGTCTGCCTGATTAGTGATTTTTATCCGGGTGCCGTT
[0136] ACAGTGGCATGGAAAGCAGACTCTAGTCCAGTCAAGGCAGGAGTTGAGACAACTACGCCATCC
[0137] AAGCAATCCAACAACAAATACGCTGCCTCAAGTTATTTGAGCCTTACCCCAGAGCAATGGAAA
[0138] AGTCATCGCTCATATTCTTGTCAGGTCACACATGAAGGCAGCACAGTGGAAAAGACGGTTGCG
[0139] 10 CCAACGGAATGTTCCGGGGGGAGTTCCGGTAGTGGTTCCGGTTCCACGGGTACCTCCTCCTCAG
[0140] GTACCGGGACTTCCGCGGGCACAACGGGAACCAGTGCCTCTACTTCCGGCTCAGGAAGTGGAG
[0141] GCGGAGGGGGGAGCGGAGGCGGTGGATCCGCGGGCGGAACCGCTACCGCTGGGGCGTCATCC
[0142] GGATCTCAAGTACAGCTCGTTCAATCAGGAGCTGAGGTGAAAAAACCCGGTTCTTCTGTAAAA
[0143] GTTAGCTGCAAATCATCTGGAGGCACTAGCAACAACTACGCTATAAGTTGGGTACGACAAGCT
[0144] 15 CCAGGGCAAGGTCTGGACTGGATGGGAGGCATTTCACCGATATTCGGGAGCACAGCTTACGCT
[0145] CAGAAGTTTCAGGGCAGGGTAACCATCTCAGCTGATATTTTTTCTAACACAGCTTACATGGAAC
[0146] TCAATAGTCTCACATCAGAAGACACAGCTGTGTACTTTTGTGCTAGGCATGGCAATTATTACTA
[0147] CTACAGTGGCATGGATGTATGGGGCCAGGGAACAACAGTTACAGTCTCTTCAGCCAGTACCAA GGGGCCCTCCGTCTTTCCTCTCGCCCCGAGTAGCAAAAGCACTTCAGGGGGGACTGCAGCACTC
[0148] 20 GGATGCCTGGTTAAAGATTACTTTCCAGAACCTGTCACCGTATCCTGGAATAGTGGCGCTCTCA CTAGCGGCGTACATACGTTCCCGGCGGTCCTCCAGAGTAGCGGTCTGTATAGTTTGAGTTCTGT
[0149] CGTAACGGTTCCCTCATCTTCTCTGGGTACACAGACCTATATCTGCAATGTGAATCATAAACCA
[0150] TCAAATACGAAGGTTGACAAACGAGTAgctagctccggaggcagcggatccggaggcagcggaGATTACAAAGAT
[0151] GACGATGATAAAggctctggagcctccggaggcagcggaggaagtgctatccccatcatgggtatcgttgctggcctggttgtccttgcagctgt
[0152] 25 agtcactggagctgcggtcgctgctgtgctgtggagaaagaagagctcagattaa
[0153] SEQ ID NO: 5 is a nucleic acid sequence encoding a secretion / membrane localization signal. atgAAGAAGAACATAGCGTTTTTGCTCGCCTTGATGTTTGTTTTTTCAATAGCAACTAATGCCTAC GCT
[0154] 30
[0155] SEQ ID NO: 6 is a nucleic acid sequence encoding the Fab CR9114 light chain.
[0156] CAATCAGCTCTCACGCAACCGCCAGCTGTTTCCGGTACCCCAGGTCAACGCGTTACCATATCAT GTAGCGGCTCTGACTCAAATATTGGTCGCCGCTCCGTAAACTGGTACCAACAGTTTCCCGGGAC AGCTCCGAAACTCCTGATCTATAGCAACGATCAACGCCCGTCAGTGGTACCAGACAGGTTTAG
[0157] 35 TGGCTCTAAATCCGGAACATCAGCTTCTCTCGCTATCAGCGGTCTCCAATCAGAAGACGAAGCT
[0158] GAGTACTACTGCGCAGCTTGGGATGACTCCTTGAAGGGAGCTGTCTTCGGCGGGGGTACTCAG
[0159] CTTACAGTGCTTGGACAGCCCAAAGCCGCCCCTTCCGTAACTCTTTTCCCCCCATCAAGCGAGG
[0160] AACTTCAAGCAAATAAGGCAACCCTTGTCTGCCTGATTAGTGATTTTTATCCGGGTGCCGTTAC AGTGGCATGGAAAGCAGACTCTAGTCCAGTCAAGGCAGGAGTTGAGACAACTACGCCATCCAA
[0161] 40 GCAATCCAACAACAAATACGCTGCCTCAAGTTATTTGAGCCTTACCCCAGAGCAATGGAAAAG TCATCGCTCATATTCTTGTCAGGTCACACATGAAGGCAGCACAGTGGAAAAGACGGTTGCGCC
[0162] AACGGAATGTTCC
[0163] SEQ ID NO: 7 is a nucleic acid sequence encoding a GS linker.
[0164] 45 GGGGGGAGTTCCGGTAGTGGTTCCGGTTCCACGGGTACCTCCTCCTCAGGTACCGGGACTTCCG CGGGCACAACGGGAACCAGTGCCTCTACTTCCGGCTCAGGAAGTGGAGGCGGAGGGGGGAGC GGAGGCGGTGGATCCGCGGGCGGAACCGCTACCGCTGGGGCGTCATCCGGATCT
[0165] SEQ ID NO: 8 is a nucleic acid sequence encoding the Fab CR9114 heavy chain.
[0166] 50 CAAGTACAGCTCGTTCAATCAGGAGCTGAGGTGAAAAAACCCGGTTCTTCTGTAAAAGTTAGC
[0167] TGCAAATCATCTGGAGGCACTAGCAACAACTACGCTATAAGTTGGGTACGACAAGCTCCAGGG CAAGGTCTGGACTGGATGGGAGGCATTTCACCGATATTCGGGAGCACAGCTTACGCTCAGAAG TTTCAGGGCAGGGTAACCATCTCAGCTGATATTTTTTCTAACACAGCTTACATGGAACTCAATA PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0168] 7950-112573-02
[0169] GTCTCACATCAGAAGACACAGCTGTGTACTTTTGTGCTAGGCATGGCAATTATTACTACTACAG TGGCATGGATGTATGGGGCCAGGGAACAACAGTTACAGTCTCTTCAGCCAGTACCAAGGGGCC CTCCGTCTTTCCTCTCGCCCCGAGTAGCAAAAGCACTTCAGGGGGGACTGCAGCACTCGGATGC CTGGTTAAAGATTACTTTCCAGAACCTGTCACCGTATCCTGGAATAGTGGCGCTCTCACTAGCG
[0170] 5 GCGTACATACGTTCCCGGCGGTCCTCCAGAGTAGCGGTCTGTATAGTTTGAGTTCTGTCGTAAC GGTTCCCTCATCTTCTCTGGGTACACAGACCTATATCTGCAATGTGAATCATAAACCATCAAAT ACGAAGGTTGACAAACGAGTA
[0171] SEQ ID NO: 9 is a nucleic acid sequence encoding a FLAG tag flanked by linker sequences.
[0172] 10 gctagctccggaggcagcggatccggaggcagcggaGATTACAAAGATGACGATGATAAAggctctggagcctccggaggcag cggaggaagtgct
[0173] SEQ ID NO: 10 is a nucleic acid sequence encoding a MHCI transmembrane helix. atccccatcatgggtatcgttgctggcctggttgtccttgcagctgtagtcactggagctgcggtcgctgctgtgctgtggagaaagaagagctcagattaa
[0174] 15
[0175] SEQ ID NO: 11 is the amino acid sequence of a chimeric protein that includes Fab CR9114, a membrane localization signal and a transmembrane helix.
[0176] MKKNIAFLLALMFVFSIATNAYAQSALTQPPAVSGTPGQRVTISCSGSDSNIGRRSVNWYQQFPGTA
[0177] PKLLIYSNDQRPSVVPDRFSGSKSGTSASLAISGLQSEDEAEYYCAAWDDSLKGAVFGGGTQLTVL
[0178] 20 GQPKAAPSVTLFPPSSEELQANKATLVCLISDFYPGAVTVAWKADSSPVKAGVETTTPSKQSNNKY
[0179] AASSYLSLTPEQWKSHRSYSCQVTHEGSTVEKTVAPTECSGGSSGSGSGSTGTSSSGTGTSAGTTGT
[0180] SASTSGSGSGGGGGSGGGGSAGGTATAGASSGSQVQLVQSGAEVKKPGSSVKVSCKSSGGTSNNY
[0181] AISWVRQAPGQGLDWMGGISPIFGSTAYAQKFQGRVTISADIFSNTAYMELNSLTSEDTAVYFCAR
[0182] HGNYYYYSGMDVWGQGTTVTVSSASTKGPSVFPLAPSSKSTSGGTAALGCLVKDYFPEPVTVSWN
[0183] 25 SGALTSGVHTFPAVLQSSGLYSLSSVVTVPSSSLGTQTYICNVNHKPSNTKVDKRVASSGGSGSGGS GDYKDDDDKGSGASGGSGGSAIPIMGIVAGLVVLAAVVTGAAVAAVLWRKKSSD
[0184] SEQ ID NO: 12 is the amino acid sequence of a secretion / membrane localization signal.
[0185] MKKNIAFLLALMFVFSIATNAYA
[0186] 30
[0187] SEQ ID NO: 13 is the amino of the Fab CR9114 light chain.
[0188] QSALTQPPAVSGTPGQRVTISCSGSDSNIGRRSVNWYQQFPGTAPKLLIYSNDQRPSVVPDRFSGSKS GTS ASLAISGLQSEDEAEYYCAAWDDSLKGAVFGGGTQLTVLGQPKAAPSVTLFPPSSEELQ ANKA TLVCLISDFYPGAVTVAWKADSSPVKAGVETTTPSKQSNNKYAASSYLSLTPEQWKSHRSYSCQVT
[0189] 35 HEGSTVEKTVAPTECS
[0190] SEQ ID NO: 14 is the amino of a GS linker.
[0191] GGSSGSGSGSTGTSSSGTGTSAGTTGTSASTSGSGSGGGGGSGGGGSAGGTATAGASSGS
[0192] 40 SEQ ID NO: 15 is the amino of the Fab CR9114 heavy chain.
[0193] QVQLVQSGAEVKKPGSSVKVSCKSSGGTSNNYAISWVRQAPGQGLDWMGGISPIFGSTAYAQKFQ GRVTISADIFSNTAYMELNSLTSEDTAVYFCARHGNYYYYSGMDVWGQGTTVTVSSASTKGPSVFP LAPSSKSTSGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGLYSLSSVVTVPSSSLG TQTYICNVNHKPSNTKVDKRV
[0194] 45
[0195] SEQ ID NO: 16 is the amino acid sequence of a FLAG tag flanked by linker sequences.
[0196] ASSGGSGSGGSGDYKDDDDKGSGASGGSGGSA
[0197] SEQ ID NO: 17 is the amino acid sequence of a MCHI transmembrane helix.
[0198] 50 IPIMGIVAGLVVLAAVVTGAAVAAVLWRKKSSD PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0199] 7950-112573-02
[0200] SEQ ID NOs: 18-36 are oligonucleotide sequences (Table 1).
[0201] SEQ ID NO: 37 is a synthesized DNA encoding the CR9114 Fab (Table 2).
[0202] SEQ ID NO: 38 is a synthesized DNA encoding linker-FLAG-MHCI helix (Table 2).
[0203] SEQ ID NOs: 39-42 are nucleic acid and amino acid sequences shown in FIG. 2A.
[0204] 5 SEQ ID NOs: 43 and 44 are nucleic acid and amino acid sequences shown in FIG. 9B. caccctcgtgaccaccctgacctacggcgtgcagtgcttcagccgc (SEQ ID NO: 43)
[0205] TLVTTLIYGVQCFSR (SEQ ID NO: 44)
[0206] SEQ ID NOs: 45 and 46 are nucleic acid and amino acid sequences shown in FIG. I IB.
[0207] TACGAAGGTTGACAAACGAGTAgctagctccggaggcagcggatccggaggcagcggaGATTACAAAGATGACGA
[0208] 10 TGATAAAggctctggagcctcc (SEQ ID NO: 45)
[0209] TKVDKRVASSGGSGSGGSGDYKDDDDKGSGAS (SEQ ID NO: 46)
[0210] SEQ ID NOs: 47-52 are amino acid sequences shown in FIGS. 12A-12F.
[0211] 15 FIG. 12A (SEQ ID NO: 47)
[0212] QSALTQPPAVSGTPGQRVTISCSGSDSNIGRRSVNWYQQFPGTAPKLLIYSNDQRPSVVPDRFSGSKS
[0213] GTSASLAISGLQSEDEAEYYCAAWDDSLKGAVFGGGTQLTV
[0214] FIG. 12B (SEQ ID NO: 48)
[0215] 20 LGQPKAAPSVTLFPPSSEELQANKATLVCLISDFYPGAVTVAWKADSSPVKAGVETTTPSKQSNNK
[0216] YAASSYLSLTPEQWKSHRSYSCQVTHEGSTVEKTVAPTECS
[0217] FIG. 12C (SEQ ID NO: 49)
[0218] GGSSGSGSGSTGTSSSGTGTSAGTTGTSASTSGSGSGGGGGSGGGGSAGGTATAGASSGS
[0219] 25
[0220] FIG. 12D (SEQ ID NO: 50)
[0221] SGGSGSGGSGDYKDDDDKGSGASGGSGGSAIPIMGIVAGLVVLAAVVTGA
[0222] FIG. 12E (SEQ ID NO: 51)
[0223] 30 QVQLVQSGAEVKKPGSSVKVSCKSSGGTSNNYAISWVRQAPGQGLDWMGGISPIFGSTAYAQKFQ
[0224] GRVTISADIFSNTAYMELNSLTSEDTAVYFCARHGNYYYYSGMDVWGQGTTVT
[0225] FIG. 12F (SEQ ID NO: 52)
[0226] VSSASTKGPSVFPLAPSSKSTSGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGLYS
[0227] 35 LSSVVTVPSSSLGTQTYICNVNHKPSNTKVDKRV
[0228] SEQ ID NO: 53 is an amino acid sequence shown in FIG. 13.
[0229] QVQLVQSGAEVKKPGSSVKVSCKSSGGTSNNYAISWVRQAPGQGLDWMGGISPIFGSTAYAQKFQ
[0230] GRVTISADIFSNTAYMELNSLTSEDTAVYFCARHGNYYYYSGMDVWGQGTTVTVSSASTKGPSVFP
[0231] 40 LAPSSKSTSGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGLYSLSSVVTVPSSSLG
[0232] TQTYICNVNHKPSNTKVDKRV
[0233] SEQ ID NO: 54 is the amino acid sequence of eGFP*
[0234] MVSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTTLIY
[0235] 45 GVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFKE
[0236] DGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGSVQLADHYQQNTPIGDGPVLLPDN
[0237] HYLSTQSALSKDPNEKRDHMVLLEFVTAAGITLGMDELYK*
[0238] SEQ ID NO: 55 is a nucleic acid sequence encoding eGFP* PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0239] 7950-112573-02 atggtgagcaagggcgaggagctgttcaccggggtggtgcccatcctggtcgagctggacggcgacgtaaacggccacaagttcagcgtgtccggcgag ggcgagggcgatgccacctacggcaagctgaccctgaagttcatctgcaccaccggcaagctgcccgtgccctggcccaccctcgtgaccaccctgatcta cggcgtgcagtgcttcagccgctaccccgaccacatgaagcagcacgacttcttcaagtccgccatgcccgaaggctacgtccaggagcgcaccatcttctt caaggacgacggcaactacaagacccgcgccgaggtgaagttcgagggcgacaccctggtgaaccgcatcgagctgaagggcatcgacttcaaggagg
[0240] 5 acggcaacatcctggggcacaagctggagtacaactacaacagccacaacgtctatatcatggccgacaagcagaagaacggcatcaaggtgaacttcaa gatccgccacaacatcgaggacggcagcgtgcagctcgccgaccactaccagcagaacacccccatcggcgacggccccgtgctgctgcccgacaacc actacctgagcacccagtccgccctgagcaaagaccccaacgagaagcgcgatcacatggtcctgctggagttcgtgaccgccgccgggatcactctcgg catggacgagctgtacaagtaa
[0241] 10 SEQ ID NO: 54 is the amino acid sequence of eGFP.
[0242] MVSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTTLT YGVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLVNRIELKGIDFK EDGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGSVQLADHYQQNTPIGDGPVLLPD NHYLSTQSALSKDPNEKRDHMVLLEFVTAAGITLGMDELYK*
[0243] 15
[0244] SEQ ID NO: 55 is a nucleic acid sequence encoding eGFP. atggtgagcaagggcgaggagctgttcaccggggtggtgcccatcctggtcgagctggacggcgacgtaaacggccacaagttcagcgtgtccggcgag ggcgagggcgatgccacctacggcaagctgaccctgaagttcatctgcaccaccggcaagctgcccgtgccctggcccaccctcgtgaccaccctgacct acggcgtgcagtgcttcagccgctaccccgaccacatgaagcagcacgacttcttcaagtccgccatgcccgaaggctacgtccaggagcgcaccatcttct
[0245] 20 tcaaggacgacggcaactacaagacccgcgccgaggtgaagttcgagggcgacaccctggtgaaccgcatcgagctgaagggcatcgacttcaaggag gacggcaacatcctggggcacaagctggagtacaactacaacagccacaacgtctatatcatggccgacaagcagaagaacggcatcaaggtgaacttca agatccgccacaacatcgaggacggcagcgtgcagctcgccgaccactaccagcagaacacccccatcggcgacggccccgtgctgctgcccgacaac cactacctgagcacccagtccgccctgagcaaagaccccaacgagaagcgcgatcacatggtcctgctggagttcgtgaccgccgccgggatcactctcg gcatggacgagctgtacaagtaa
[0246] 25
[0247] SEQ ID NOs: 56-58 are amino acid sequences shown in FIG. 20D.
[0248] SGGGGGS (SEQ ID NO: 56)
[0249] 30 GSGGGG (SEQ ID NO: 57)
[0250] VTHQGLSSPVTKSFNRGECGGSSGSGSGSTGTSSS (SEQ ID NO: 58)
[0251] SEQ ID NOs: 59-68 are nucleic acid sequences of exemplary SHM enhancer elements (see Table 5).
[0252] 35
[0253] DETAILED DESCRIPTION
[0254] I. Introduction
[0255] The adaptive immune system has evolved mechanisms to overcome the evolutionary bottleneck of low error frequencies and mutational tolerance to rapidly evolve proteins in response to antigen exposure.
[0256] 40 This is particularly observed in B cells, which have evolved genetic recombination (Tonegawa et ah, Nature 302, 575-581) and somatic hypermutation (SHM) mechanisms (Di Noia et al., Annu Rev Biochem 76, 1- 22). The SHM machinery allows B cells to introduce mutations specifically at the immunoglobulin genomic loci at a significantly higher frequency than the rest of the genome. This allows B cells to rapidly evolve new antibody sequences without compromising the fitness of the immune cells due to genome-wide mutations.
[0257] 45 This is an attractive feature for biomolecular evolution. In vivo methods have been used to evolve antibodies targeting defined epitopes (Milstein et al., Nature 305, 537-540) as well as for evolving catalytic antibodies (Schultz et al., Science 269, 1835-1842). Additionally, previous studies have attempted to generate and PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0258] 7950-112573-02 optimize antibody sequences by incorporating defined antibody encoding genes at the immunoglobulin loci of the B cells (Seo et al., Cell. Mol. Immunol. 18, 1545-1561; Nahmad et al., Nat. Commun. 11, 5851; Nahmad et al., Nat. Biotechnol. 40, 1241-1249). Such loci are genetically unstable due to recombination events. B cells have also been engineered to express custom antibodies targeting defined epitopes as
[0259] 5 potential therapeutics. Another study demonstrated that low frequency, virus-based genomic integration methods resulted in transient expression and evolution of fluorescent proteins when the corresponding gene was integrated at the immunoglobulin locus. However, a need remains for an efficient, non-virus-based method for rapidly evolving proteins of interest.
[0260] The present disclosure describes studies to investigate whether the SHM mechanisms of B cells can be hijacked and repurposed to engineer human B cell hnes as chassis for continuous directed evolution. In particular, disclosed here is a Continuous Directed Evolution platform in Human B cells (CODE-HB) that recruits and repurposes the inherent B cell SHM mechanisms to rapidly, and orthogonally, evolve proteins of interest. First, precise genome editing approaches to incorporate defined genetic elements at a stable, nonimmunoglobulin locus in human B cell line genomes were developed. This allowed for development of
[0261] 15 stable high levels of protein expression in more than 90% of the cell population. To continuously evolve reporter proteins expressed from this locus, a series of defined DNA elements were identified and added from the immunoglobulin locus upstream of the gene encoding a reporter protein to facilitate the recruitment of the B cell SHM machinery to this non-immunoglobulin stable locus. These engineered human B cell lines rapidly evolve reporter proteins (e.g., enhanced green fluorescent protein, eGFP). To adapt this platform to
[0262] 20 rapidly evolve membrane displayed proteins, a human B cell surface display platform was developed and combined with CODE-HB. Using this approach, fragment antigen binding domains (Fat>) of antibodies on the surface of human B cells were displayed and CODE-HB was used to evolve these Fab sequences for targeting avian subtypes of influenza hemagglutinin. This resulted in the identification of H5 targeting antibodies. Such antibodies are important in the context of recent zoonotic spread of H5N1 strains of avian
[0263] 25 influenza (Eisfeld et al., Nature 633(8029) :426-432, 2024). This platform can be used to evolve antibodies to escape variants of hemagglutinin. Further, it is demonstrated that evolved antibodies are more potent than the parent antibodies in viral neutralization assays. The mutational profile and breadth of this approach was comprehensively characterized by using single-molecule sequencing experiments and a broad mutational profile comprised of substitutions, deletions and insertions was observed. Given the modularity and simplicity of CODE-HB to rapidly evolve both cytoplasmic proteins as well as surface displayed proteins, this platform can be used for developing biologies as well as for proteins and enzymes directly in human cell lines. Furthermore, this engineered stable system can be used to provide new insights into the SHM mechanisms.
[0264] 35 IL Abbreviations
[0265] AAV adeno-associated virus
[0266] AID activation-induced deaminase PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0267] 7950-112573-02
[0268] BER base excision repair
[0269] CDR complementarity determining region
[0270] CODE-HB Continuous Directed Evolution platform in Human B cells CRISPR clustered regularly interspaced short palindromic repeats
[0271] 5 DSB double strand break
[0272] EC50 half-maximal effective concentration eGFP enhanced green fluorescent protein
[0273] Fab single -chain antibody binding fragment.
[0274] FACS fluorescence activated cell sorting
[0275] HA hemagglutinin
[0276] PolyA polyadenylation
[0277] MHCI major histocompatibility complex class I
[0278] SMH somatic hypermutation
[0279] 15 III. Summary of Terms
[0280] Unless otherwise noted, technical terms are used according to conventional usage. Definitions of many common terms in molecular biology may be found in Krebs et al. (eds.), Lewin ’s genes XII, published by Jones & Bartlett Learning, 2017. As used herein, the singular forms “a,” “an,” and “the,” refer to both the singular as well as plural, unless the context clearly indicates otherwise. For example, the term “an
[0281] 20 antigen” includes singular or plural antigens and can be considered equivalent to the phrase “at least one antigen.” As used herein, the term “comprises” means “includes.” It is further to be understood that any and all base sizes or amino acid sizes, and all molecular weight or molecular mass values, given for nucleic acids or polypeptides are approximate, and are provided for descriptive purposes, unless otherwise indicated. Although many methods and materials similar or equivalent to those described herein can be used, particular
[0282] 25 suitable methods and materials are described herein. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. To facilitate review of the various aspects, the following explanations of terms are provided:
[0283] Antibody and antigen binding fragment: An immunoglobulin, antigen-binding fragment, or derivative thereof, that specifically binds and recognizes an analyte (antigen) such as, but not limited to, a viral antigen, a bacterial antigen, a fungal antigen, or a tumor antigen. The term “antibody” is used herein in the broadest sense and encompasses various antibody structures, including but not limited to monoclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments, so long as they exhibit the desired antigen-binding activity.
[0284] 35 Non-limiting examples of antibodies include, for example, intact immunoglobulins and variants and fragments thereof that retain binding affinity for the antigen. Examples of antibody fragments include, but are not limited to, Fv, Fab, Fab', Fab'-SH, F(ab')2; diabodies; linear antibodies; single-chain antibody PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0285] 7950-112573-02 molecules (e.g. scFv, VHH); and multispecific antibodies formed from antibody fragments. Antibody fragments include antigen binding fragments either produced by the modification of whole antibodies or those synthesized de novo using recombinant DNA methodologies (see, e.g., Kontermann and Diibel (Eds.), Antibody Engineering, Vols. 1-2, 2nded., Springer-Verlag, 2010).
[0286] 5 Fab antibody fragments contain a monovalent antigen-binding fragment of an antibody molecule, which can be produced by digestion of whole antibody with the enzyme papain (which removes the Fc portion of the antibody) to yield an intact light chain and a portion of one heavy chain. Fab molecules can also be recombinantly produced. Fab' fragments can be obtained by treating whole antibody with pepsin, followed by reduction, to yield an intact light chain and a portion of the heavy chain; two Fab’ fragments are obtained per antibody molecule. (Fab’)2 antibody fragments can be obtained by treating whole antibody with the enzyme pepsin without subsequent reduction; F(ab’)2 is a dimer of two Fab' fragments held together by two disulfide bonds.
[0287] A single-chain antibody (scFv) is a genetically engineered molecule containing the VH and VL domains of one or more antibody(ies) linked by a suitable polypeptide linker as a genetically fused single
[0288] 15 chain molecule (see, for example, Bird et al., Science, 242(4877):423-426, 1988; Huston et al., Proc. Natl. Acad. Sci. U.S.A., 85(16):5879-5883, 1988; Ahmad etal., Clin. Dev. Immunol., 2012, doi: 10.1155 / 2012 / 980250; Marbry and Snavely, IDrugs, 13(8):543-549, 2010). The intramolecular orientation of the VH domain and the VL domain in a scFv is typically not decisive for scFvs. Thus, scFvs with both possible arrangements (VH domain-linker domain-V, domain; V domain-linker domain-Vn
[0289] 20 domain) may be used.
[0290] In a dsFv, the VH and VL have been mutated to introduce a disulfide bond to stabilize the association of the chains. Diabodies also are included, which are bivalent, bispecific antibodies in which VH and VL domains are expressed on a single polypeptide chain, but using a linker that is too short to allow for pairing between the two domains on the same chain, thereby forcing the domains to pair with complementary
[0291] 25 domains of another chain and creating two antigen binding sites (see, for example, Holliger et al., Proc. Natl. Acad. Sci. U.S.A., 90(14):6444-6448, 1993; Poljak eta / ., Structure, 2(12):1121-1123, 1994).
[0292] Antibodies also include genetically engineered forms such as chimeric antibodies (such as humanized murine or macaque antibodies) and heteroconjugate antibodies (such as bispecific antibodies).
[0293] Typically, a naturally occurring mammalian immunoglobulin has heavy (H) chains and light (L) chains interconnected by disulfide bonds. Immunoglobulin genes include the kappa, lambda, alpha, gamma, delta, epsilon and mu constant region genes, as well as the myriad immunoglobulin variable domain genes. There are two types of light chain, lambda ( ) and kappa (K). There are five main heavy chain classes (or isotypes) that determine the functional activity of a mammalian antibody molecule: IgM, IgD, IgG, IgA and IgE.
[0294] 35 Each heavy and light chain contain a constant region (or constant domain) and a variable region (or variable domain). In several aspects, the VH and VL combine to specifically bind the antigen. In additional aspects, only the VH is required. For example, naturally occurring camelid antibodies consisting of a heavy PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0295] 7950-112573-02 chain only (VRH) are functional and stable in the absence of light chain. Antibodies can also include a heterologous constant domain. For example, the antibody can include a constant domain that is different from a native constant domain, such as a constant domain including one or more modifications (such as the “LS” mutations) to increase half-life.
[0296] 5 References to “VH” or “VH” refer to the variable region of an antibody heavy chain, including that of an antigen binding fragment, such as Fv, scFv, dsFv or Fab. References to “VL” or “VL” refer to the variable domain of an antibody light chain, including that of an Fv, scFv, dsFv or Fab.
[0297] The VHand VLcontain a “framework” region interrupted by three hypervariable regions, also called “complementarity-determining regions” or “CDRs” (see, e.g., Kabat et al., Sequences of Proteins of Immunological Interest, 5thed., NIH Publication No. 91-3242, Public Health Service, National Institutes of Health, U.S. Department of Health and Human Services, 1991). The sequences of the framework regions of different light or heavy chains are relatively conserved within a species. The framework region of an antibody, that is the combined framework regions of the constituent light and heavy chains, serves to position and align the CDRs in three-dimensional space.
[0298] 15 The CDRs are primarily responsible for binding to an epitope of an antigen. The amino acid sequence boundaries of a given CDR can be readily determined using any of a number of well-known schemes, including those described by Kabat et al. (Sequences of Proteins of Immunological Interest, U.S. Department of Health and Human Services, 1991; the “Kabat” numbering scheme), Chothia et al. (see Chothia and Lesk, J Mol Biol 196:901-917, 1987; Chothia et al., Nature 342:877, 1989; and Al-Lazikani et
[0299] 20 al., JMB 273,927-948, 1997; the “Chothia” numbering scheme), Kunik et al. (see Kunik et al., PLoS Comput Biol 8:el002388, 2012; and Kunik et al., Nucleic Acids Res 40(Web Server issue):W521-524, 2012; “Paratome CDRs”) and the ImMunoGeneTics (IMGT) database (see, Lefranc, Nucleic Acids Res 29:207-9, 2001; the “IMGT” numbering scheme). The Kabat, Paratome and IMGT databases are maintained online. In addition, the AbRSA tool can be used to determine the CDR boundaries according to Kabat, IMGT or
[0300] 25 Chothia (online at aligncdr.labshare.cn / aligncdr / abrsa.php). The CDRs of each chain are typically referred to as CDR1, CDR2, and CDR3 (from the N-terminus to C-terminus), and are also typically identified by the chain in which the particular CDR is located. Thus, a VH CDR3 is the CDR3 from the VH of the antibody in which it is found, whereas a VL CDR1 is the CDR1 from the VL of the antibody in which it is found. Light chain CDRs are sometimes referred to as LCDR1, LCDR2, and LCDR3. Heavy chain CDRs are sometimes referred to as HCDR1, HCDR2, and HCDR3.
[0301] A “monoclonal antibody” is an antibody obtained from a population of substantially homogeneous antibodies, that is, the individual antibodies comprising the population are identical and / or bind the same epitope, except for possible variant antibodies, for example, containing naturally occurring mutations or arising during production of a monoclonal antibody preparation, such variants generally being present in
[0302] 35 minor amounts. In contrast to polyclonal antibody preparations, which typically include different antibodies directed against different determinants (epitopes), each monoclonal antibody of a monoclonal antibody preparation is directed against a single determinant on an antigen. Thus, the modifier “monoclonal” indicates PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0303] 7950-112573-02 the character of the antibody as being obtained from a substantially homogeneous population of antibodies, and is not to be construed as requiring production of the antibody by any particular method. For example, the monoclonal antibodies may be made by a variety of techniques, including but not limited to the hybridoma method, recombinant DNA methods, phage-display methods, and methods utilizing transgenic
[0304] 5 animals containing all or part of the human immunoglobulin loci, such methods and other exemplary methods for making monoclonal antibodies are well-known. In some examples, monoclonal antibodies are isolated from a subject. Monoclonal antibodies can have conservative amino acid substitutions which have substantially no effect on antigen binding or other immunoglobulin functions. (See, for example, Greenfield (Ed.), Antibodies: A Laboratory Manual, 2nded. New York: Cold Spring Harbor Laboratory Press, 2014.)
[0305] A “humanized” antibody or antigen binding fragment includes a human framework region and one or more CDRs from a non-human (such as a non-human primate, mouse, rat, or synthetic) antibody or antigen binding fragment. The non-human antibody or antigen binding fragment providing the CDRs is termed a “donor,” and the human antibody or antigen binding fragment providing the framework is termed an “acceptor.” Constant regions need not be present, but if they are, they can be substantially identical to
[0306] 15 human immunoglobulin constant regions, such as at least about 85-90%, such as about 95% or more identical. Hence, all parts of a humanized antibody or antigen binding fragment, except possibly the CDRs, are substantially identical to corresponding parts of natural human antibody sequences.
[0307] A “chimeric antibody” is an antibody that includes sequences derived from two different antibodies, which typically are of different species. In some examples, a chimeric antibody includes one or more CDRs
[0308] 20 and / or framework regions from one human antibody and CDRs and / or framework regions from another human antibody.
[0309] A “fully human antibody” or “human antibody” is an antibody which includes sequences from (or derived from) the human genome, and does not include sequence from another species. In some aspects, a human antibody includes CDRs, framework regions, and (if present) an Fc region from (or derived from) the
[0310] 25 human genome. Human antibodies can be identified and isolated using technologies for creating antibodies based on sequences derived from the human genome, for example by phage display or using transgenic animals (see, e.g., Barbas et al. Phage display: A Laboratory Manuel. 1sted. New York: Cold Spring Harbor Laboratory Press, 2004; Lonberg, Nat. Biotechnol., 23(9): 1117-1125, 2005; Lonberg, Curr. Opin. Immunol. 20(4):450-459, 2008).
[0311] Binding affinity: Affinity of an antibody for an antigen. In one aspect, affinity is calculated by a modification of the Scatchard method described by Frankel et al., Mol. Immunol., 16:101-106, 1979. In another aspect, binding affinity is measured by an antigen / antibody dissociation rate. In another aspect, a binding affinity is measured by a competition radioimmunoassay. In another aspect, binding affinity is measured by ELISA. In some aspects, binding affinity is measured using bio-layer interferometry (BLI)
[0312] 35 technology, such as by using the Octet system (Creative Biolabs). In other aspects, Kd is measured using a surface plasmon resonance (SPR) assay, such as by using a BIACORES-2000 or a BIACORES-3000 (BIAcore, Inc., Piscataway, N.J.). In other aspects, antibody affinity is measured by flow cytometry. An PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0313] 7950-112573-02 antibody that “specifically binds” an antigen is an antibody that binds the antigen with high affinity and does not significantly bind other unrelated antigens.
[0314] B cells (or B lymphocytes): A type of white blood cell that plays an important role in the adaptive immune system. B cells produce antibody molecules. In the context of the present disclosure, a B cell can
[0315] 5 be a primary B cell (such as a B cell isolated from a subject) or a B cell that is from a B cell line. In some aspects, the B cell is a mammalian B cell, such as a human B cell or a swine B cell. In other aspects, the B cell is an avian B cell. In particular examples, the B cell line is the human RAI cell line. Numerous B cell lines are known and available, including the following human B cell lines (dtcore.northwestern.edu / dtc-cell- line-list):
[0316] 10
[0317] Codon-optimized: A nucleic acid molecule encoding a protein can be codon-optimized for expression of the protein in a particular organism by including the codon most likely to encode a particular amino acid at each position of the sequence. Codon usage bias is the difference in the frequency of occurrence of synonymous codons (encoding the same amino acid) in coding DNA. A codon is a series of 15 three nucleotides (a triplet) that encodes a specific amino acid residue in a polypeptide chain or for the termination of translation. There are 20 different naturally-occurring amino acids, but 64 different codons (61 codons encoding for amino acids plus 3 stop codons). Thus, there is degeneracy because one amino acid can be encoded by more than one codon. A nucleic acid sequence can be optimized for expression in a particular organism (such as a human) by evaluating the codon usage bias in that organism and selecting the PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0318] 7950-112573-02 codon most likely to encode a particular amino acid. Multivariate statistical methods, such as correspondence analysis and principal component analysis, are widely used to analyze variations in codon usage. Computer programs are available to implement the statistical analyses related to codon usage, such as Codon W, GCUA, and INCA.
[0319] 5 CRISPR / Cas9 system: A prokaryotic immune system that confers resistance to foreign genetic elements, such as plasmids and phages, and provides a form of acquired immunity. CRISPR (clustered regularly interspaced short palindromic repeats) refers to DNA loci containing short repetitions of base sequences. Each repetition is followed by short segments of "spacer DNA" from previous exposures to a virus. CRISPRs are found in approximately 40% of sequenced bacteria genomes and 90% of sequenced archaea. CRISPRs are often associated with Cas genes that code for proteins related to CRISPRs. CRISPR spacers recognize and cut these exogenous genetic elements in a manner analogous to RNAi in eukaryotic organisms. The CRISPR / Cas system can be used for gene editing (adding, disrupting or changing the sequence of specific genes) and gene regulation. By delivering the Cas9 protein and appropriate guide RNAs into a cell, the organism's genome can be cut at any desired location. Cas9 is an RNA-guided DNA
[0320] 15 endonuclease enzyme that can cut DNA. Cas9 has two active cutting sites (HNH and RuvC), one for each strand of the double helix.
[0321] Guide RNA (gRNA): An RNA sequence having sufficient complementarity with a target polynucleotide sequence to hybridize with the target sequence and direct sequence-specific binding of a Cas9 to the target sequence. In some instances, the guide RNA can include modified bases or chemical
[0322] 20 modifications (e.g., see Latorre et al., Angewandte Chemie 55:3548-50, 2016). In some aspects, the degree of complementarity between a guide RNA and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or more. Optimal alignment may be determined with the use of any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the
[0323] 25 Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAST, Novoalign (Novocraft Technologies), ELAND (Illumina, San Diego, CA), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net). In some aspects, a guide RNA is about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 or 35 or more nucleotides in length.
[0324] The ability of a guide RNA to direct sequence-specific binding of a CRISPR complex to a target sequence may be assessed by any suitable assay. For example, the components of a CRISPR system sufficient to form a CRISPR complex, including the guide sequence to be tested, may be provided to a host cell having the corresponding target sequence, such as by transfection with vectors encoding the components of the CRISPR sequence, followed by an assessment of preferential cleavage within the target sequence.
[0325] 35 Similarly, cleavage of a target polynucleotide sequence may be evaluated in a test tube by providing the target sequence, components of a CRISPR complex, including the guide sequence to be tested and a control PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0326] 7950-112573-02 guide sequence different from the test guide sequence, and comparing binding or rate of cleavage at the target sequence between the test and control guide sequence reactions.
[0327] Hemagglutinin (HA): An influenza virus surface glycoprotein. HA mediates binding of the virus particle to host cells and subsequent entry of the virus into the host cell. HA also causes red blood cells to
[0328] 5 agglutinate. HA (along with NA) is one of the two major influenza virus antigenic determinants.
[0329] Heterologous: Originating from a separate genetic source or species. For example, a heterologous gene refers to a gene derived from a different source or species (such as a different source of species than the other sequences of an integration plasmid).
[0330] Homologous sequence: In the context of the present disclosure, refers to a nucleic acid sequence that has substantial (up to 100%) sequence identity to another sequence, such as genomic sequence. In some aspects, the sequence identity is at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%.
[0331] Influenza virus: A segmented, negative-strand RNA virus that belongs to the Orthomyxoviridae family. Influenza viruses are enveloped viruses. There are four types of influenza viruses, A, B, C and D.
[0332] 15 Integration plasmid: A plasmid that includes the components necessary to allow for integration of a heterologous gene into a genomic locus (such as integration into the genome of a B cell). In some aspects herein, the integration plasmid includes the heterologous gene operably linked to a promoter; a SHM enhancer element (such as the proXIV-1 sequence of SEQ ID NO: 1, or any of SEQ ID NOs: 59-68) upstream of the promoter; a first nucleic acid sequence homologous to genomic sequences at the integration
[0333] 20 site, which is upstream of the SHM enhancer element; and a second nucleic acid sequence homologous to genomic sequences at the integration site, which is downstream of the heterologous gene. In some examples, the second nucleic acid sequence is downstream of the heterologous gene and a polyadenylation (poly-A) signal sequence.
[0334] Isolated: An “isolated” or “purified” biological component (such as a nucleic acid, peptide, protein,
[0335] 25 protein complex, cell or virus) has been substantially separated, produced apart from, or purified away from other biological components in the cell or organism in which the component occurs, that is, other chromosomal and extrachromosomal DNA and RNA, proteins, and cells. In some aspects herein, an “isolated cell” refers to a cell that is not part of an organism (e.g., an isolated primary cell or a cultured cell). The term “isolated” or “purified” does not require absolute purity; rather, it is intended as a relative term. Thus, for example, an isolated biological component (such as a cell) is one in which the biological component is more enriched than the biological component is in its natural environment. Preferably, a preparation is purified such that the biological component represents at least 50%, such as at least 70%, at least 90%, at least 95%, or greater, of the total biological component content of the preparation.
[0336] Linker: A bi-functional molecule (such as a peptide) that can be used to link two molecules into
[0337] 35 one contiguous molecule, for example, to link a light chain and a heavy chain of an antibody. Non-limiting examples of peptide linkers include glycine-serine linkers, or glycine-serine-alanine-threonine linkers. The terms “conjugating,” “joining,” “bonding,” or “linking” can refer to making two molecules into one PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0338] 7950-112573-02 contiguous molecule; for example, linking two polypeptides into one contiguous polypeptide, or covalently attaching an effector molecule or detectable marker radionuclide or other molecule to a polypeptide, such as an antibody or antibody fragment. The linkage can be either by chemical or recombinant means. “Chemical means” refers to a reaction between the antibody moiety and the effector molecule such that there is a
[0339] 5 covalent bond formed between the two molecules to form one molecule.
[0340] Membrane localization signal sequence: A sequence that facilitates targeting of proteins to the plasma membrane. In some aspects herein, the membrane localization signal sequence has the amino acid sequence of SEQ ID NO: 12.
[0341] Operably linked: A first nucleic acid sequence is operably linked with a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For instance, a promoter, such as the EFl a promoter, is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence. Generally, operably linked DNA sequences are contiguous and, where necessary to join two protein-coding regions, in the same reading frame.
[0342] 15 Peptide tag (or protein tag): A heterologous peptide appended to a protein to, for example, assist with detection or purification of the protein. Exemplary peptide tags include, but are not limited to, FLAG, streptavidin, STREP-TACTIN, STREP-TAG, His (e.g., 6XHis), HBH (bacterially-derived in vivo biotinylation signaling peptide flanked by hexahistidine motifs), S-tag, glutathione S-transferase (GST), hemagglutinin (HA), maltose-binding protein (MBP), tandem affinity purification (TAP), calmodulin-
[0343] 20 binding peptide (CBP), thioredoxin (TRX), bacteriophage V5 epitope, bacteriophage T7 epitope, Myc, and VSV-G (for a review, see Kimple et al., CurrProtoc Protein Sci 73:9.9.1-9.9.23, 2013; see also blog.addgene.org / plasmids-101-protein-tags).
[0344] Promoter: An array of nucleic acid control sequences that direct transcription of a nucleic acid. A promoter includes necessary nucleic acid sequences near the start site of transcription. A promoter also
[0345] 25 optionally includes distal enhancer or repressor elements. A “constitutive promoter” is a promoter that is continuously active and is not subject to regulation by external signals or molecules. In contrast, the activity of an “inducible promoter” is regulated by an external signal or molecule (for example, a transcription factor). In some aspects herein, the promoter is the phosphoglycerol kinase (PGK) promoter or the EFla promoter. In particular examples, the EFla promoter includes an intronic enhancer.
[0346] Somatic hypermutation (SHM): A process that occurs in B cells to generate antibody diversity and high-affinity antibodies. SHM is a cellular mechanism that assists the immune system to adapt to foreign antigens. An SHM enhancer element is a sequence that functions to target SHM machinery to immunoglobulin genes in activated B cells. In some aspects of the present disclosure, the SHM enhancer element includes the sequence of proXIV-1 (SEQ ID NO: 1), or the sequence of any one of SEQ ID NOs:
[0347] 35 59-68. In the context of the present disclosure, the SHM enhancer element targets the SHM machinery to a selected non-immunoglobulin genomic locus. PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0348] 7950-112573-02
[0349] Stable genomic locus: A genomic locus that allows for stable heterologous gene expression upon insertion of the heterologous gene at the locus. Other criteria for selection of a stable genomic locus include being outside of transcriptional units and ultra-conserved regions, 50-300 kb away from the 5' end of genes, and avoiding cancer-related genes and microRNA. Exemplary stable genomic loci (also referred to as “safe
[0350] 5 harbor sites”) include those described in Pellenz et al. (Hum Gene Ther 39(7):814-828, 2019) and Aznauryan et al. (Cell Rep Methods 2(1): 100154, 2022). In some aspects herein, the stable genomic locus is the Hl 1 locus of human chromosome 22 (22ql2.2) or the AAVS1 locus of human chromosome 19q (see, e.g., Ruan et al., Sci Reports 5:14253, 2015; Hayashi et al., Sci Reports 10:21474, 2020).
[0351] Transduced, transformed, or transfected: A virus, vector, or plasmid “transduces” a cell when it transfers nucleic acid into the cell. A cell is “transformed” or “transfected” by a nucleic acid transduced into the cell when the nucleic acid molecule becomes stably replicated by the cell, either by incorporation of the nucleic acid into the cellular genome, or by episomal replication. Numerous methods of transfection can be used, such as: chemical methods (e.g., calcium-phosphate transfection), physical methods (e.g., electroporation, microinjection, particle bombardment), fusion (e.g., liposomes), receptor-mediated
[0352] 15 endocytosis (e.g., DNA-protein complexes, viral envelope / capsid-DNA complexes) and by biological infection by viruses such as recombinant viruses (Wolff, J. A., ed, Gene Therapeutics, Birkhauser, Boston, USA (1994).
[0353] IV. Compositions and Methods for Rapid Evolution of Proteins
[0354] 20 The present disclosure describes compositions (including integration plasmids and CRISPR / Cas9- encoding plasmids), kits, and methods for the rapid evolution of proteins of interest, such as antibodies or antibody fragments, by harnessing the inherent somatic hypermutation (SHM) machinery of B cells. Also described are plasmids and methods for displaying antibodies or antibody fragments on the surface of B cells.
[0355] 25 A. Integration Plasmids
[0356] Provided herein are integration plasmids capable of integrating at a specific genomic locus, such as at a specific locus of the human genome. In some aspects, the integration plasmid includes a heterologous gene operably linked to a first promoter; a SHM enhancer element upstream of the first promoter and the heterologous gene; a first nucleic acid sequence homologous to genomic sequences at the integration site, wherein the first nucleic acid sequence is upstream of the SHM enhancer element; and a second nucleic acid sequence homologous to genomic sequences at the integration site, wherein the second nucleic acid sequence is downstream of the heterologous gene. In some examples, the plasmid further includes a polyadenylation (poly -A) signal sequence following the heterologous gene (and upstream of the second nucleic acid sequence homologous to genomic sequences at the integration site).
[0357] 35 In some aspects, the genomic locus (also referred to as “the insertion site”) is a stable, nonimmunoglobulin locus. A number of stable, non-immunoglobulin loci are known and an appropriate locus can be selected by a skilled person. Exemplary stable genomic loci (also referred to as “safe harbor sites”) PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0358] 7950-112573-02 are described in Pellenz et al. (Hum Gene Ther 39(7) : 814-828, 2019) and Aznauryan et al. (Cell Rep Methods 2(1) : 100154, 2022). In some examples herein, the genomic locus is the Hl l locus of human chromosome 22. In other examples, the genomic locus is the AAVS1 locus of human chromosome 19.
[0359] The heterologous gene of the integration plasmid can be any gene encoding a protein for which
[0360] 5 rapid evolution is desired. Exemplary genes include those that encode antibodies or antibody fragments. In some aspects, the heterologous gene encodes an antibody or antibody fragment. Antibody fragments include, but are not limited to Fab, Fab, Fab', Fab'-SH, F(ab')2, Fv, scFv, and VHH antibody fragments. In some examples, the antibody fragment is a single-chain antibody binding fragment (Fab). In particular examples, the Fab includes a Fab light chain and a Fab heavy chain and a linker between the Fab light chain and the Fab heavy chain.
[0361] The antibody or antibody fragment can bind any protein of interest. In some aspects, the antibody is specific for a protein of a microorganism, such as a viral, bacterial, fungal, protozoal, or parasitic antigen. In other aspects, the antibody is specific for a tumor antigen.
[0362] In some aspects, the antibody or antibody fragment binds a viral protein, such as an antigen of an
[0363] 15 RNA virus (e.g., an antigen of a negative-sense single-stranded RNA virus, a positive-sense single-stranded RNA virus or a double-stranded RNA viruses) or a DNA virus (e.g., an antigen of a single-stranded DNA viruses or a double-stranded DNA viruses). In specific examples, the viral protein is an influenza virus protein, such as influenza virus hemagglutinin (HA) protein or neuraminidase (NA) protein.
[0364] In some aspects, the heterologous gene further encodes a membrane localization signal, a
[0365] 20 transmembrane helix, and / or a linker between the antibody / antibody fragment and the transmembrane helix.
[0366] In some examples, the nucleic acid sequence encoding the membrane localization signal is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO: 5, or includes or consists of SEQ ID NO: 5. In some examples, the amino acid sequence of the membrane localization signal is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at
[0367] 25 least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 12, or includes or consists of SEQ ID NO: 12.
[0368] In some examples, the transmembrane helix is an MHC class I (MHCI) transmembrane helix. In specific examples, the nucleic acid sequence encoding the transmembrane helix is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 10, or includes or consists of SEQ ID NO: 10. In specific examples, the amino acid sequence of the transmembrane helix is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 17, or includes or consists of SEQ ID NO: 17.
[0369] In some examples, the linker between the antibody / antibody fragment and the transmembrane helix includes a peptide tag. In specific examples, the peptide tag is a FLAG tag. In other specific examples, the
[0370] 35 peptide tag includes streptavidin, STREP-TACTIN, STREP-TAG, His (e.g., 6XHis), HBH, S-tag, GST, HA, MBP, TAP, CBP, TRX, bacteriophage V5 epitope, bacteriophage T7 epitope, Myc, or VSV-G. In particular examples, the nucleic acid sequence encoding the linker with a FLAG tag (linker-FLAG-linker) is at least PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0371] 7950-112573-02
[0372] 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 9 (or includes or consists of SEQ ID NO: 9) and / or the amino acid sequence of the linker is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 16 (or includes or consists of SEQ ID NO: 16).
[0373] 5 In some aspects, the nucleotide sequence of the heterologous gene is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 4. In some examples, the nucleotide sequence of the heterologous gene includes or consists of SEQ ID NO: 4. In some examples, the nucleic acid sequence of the heterologous gene, or any portion thereof (such as the antibody or antibody fragment coding sequence, the membrane localization sequence, the transmembrane helix coding sequence and / or the peptide tag coding sequence) is codon-optimized for expression in mammalian cells, such as human cells.
[0374] In some aspects, the nucleotide sequence of the SHM enhancer element is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 1. In some examples, the nucleotide sequence of the SHM enhancer element includes or consists of SEQ ID
[0375] 15 NO: 1. In other aspects, the nucleotide sequence of the SHM enhancer element is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to any one of SEQ ID NOs: 59-68. In some examples, the nucleotide sequence of the SHM enhancer element includes or consists of any one of SEQ ID NOs: 59-68.
[0376] In some aspects, the first promoter of the integration plasmid is the human EFla promoter. In some
[0377] 20 examples, the human EFla promoter includes a transcription start site and an intronic enhancer (the latter of which is spliced out during mRNA maturation). In other examples, a different human promoter is used, or a synthetic or evolved promoter (see, e.g., Kim et al., Nature 436(7052):876-880, 2005).
[0378] In some aspects, the first nucleic acid sequence homologous to genomic sequences at the integration site is about 200 to about 800 nucleotides in length, about 250 to about 750 nucleotides in length, about 300
[0379] 25 to about 700 nucleotides in length, about 350 to about 650 nucleotides in length, or about 400 to about 600 nucleotides in length, such as about 200, about 225, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, about 525 about 550, about 575, about 600, about 625, about 650, about 675, about 700, about 725, about 750, about 775, or about 800 nucleotides in length; and / or the second nucleic acid sequence homologous to genomic sequences at the integration site is about 200 to about 800 nucleotides in length, about 250 to about 750 nucleotides in length, about 300 to about 700 nucleotides in length, about 350 to about 650 nucleotides in length, or about 400 to about 600 nucleotides in length, such as about 200, about 225, about 250, about 275, about 300, about 325, about 350, about 375, about 400, about 425, about 450, about 475, about 500, about 525 about 550, about 575, about 600, about 625, about 650, about 675, about 700, about 725, about 750, about 775, or about 800 nucleotides
[0380] 35 in length. In particular examples, the first nucleic acid sequence homologous to genomic sequences at the integration site is about 700 nucleotides in length; and / or the second nucleic acid sequence homologous to genomic sequences at the integration site is about 700 nucleotides in length. PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0381] 7950-112573-02
[0382] In some aspects, the integration plasmid further includes an antibiotic resistance cassette. In some examples, the antibiotic resistance cassette is a puromycin resistance cassette. In some examples, the antibiotic resistance cassette includes an antibiotic resistance gene operably linked to a second promoter. In particular examples, the second promoter is a PGK promoter. In one example, the antibiotic resistance
[0383] 5 cassette is a puromycin resistance cassette and the second promoter is a PGK promoter.
[0384] In some aspects, the nucleotide sequence of the integration plasmid is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 2. In some examples, the nucleotide sequence of the integration plasmid includes or consists of SEQ ID NO: 2.
[0385] B. Kits
[0386] Further described herein are kits, such as kits that have one or more components to assist with carrying out the disclosed methods of rapidly evolving proteins of interest. In some aspects, the kits include one or more integration plasmids disclosed herein. In some aspects, the kit further includes a CRISPR / Cas9 plasmid encoding Cas9 and one or more guide RNAs that target the genomic locus of the integration plasmid; isolated B cells; B cell culture media; tissue culture flask(s); and / or buffer(s).
[0387] 15 In some aspects, the CRISPR / Cas9 plasmid includes one or more guide RNAs that target the Hl 1 locus of human chromosome 22q. In some examples, the CRISPR / Cas9 plasmid is the p276 plasmid (Addgene ID 164850). In specific examples, the nucleotide sequence of the CRISPR / Cas9 plasmid is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 3, or includes or consists of SEQ ID NO: 3. In other aspects, the CRISPR / Cas9 plasmid
[0388] 20 includes one or more guide RNAs that target the AAVS1 locus of human chromosome 19q. In yet other aspects, the CRISPR / Cas9 plasmid includes one or more guide RNAs that target another known stable nonimmunoglobulin locus (such as any one of the loci described in Pellenz et al., Hum Gene Ther 39(7):814- 828, 2019).
[0389] In some aspects of the kit, the isolated B cells are primary B cells, such as primary B cells from any
[0390] 25 mammalian species (such as human B cells or swine B cells), or primary B cells from an avian species. In other aspects, the isolated B cells are cells of a B cell line, such as a mammalian B cell line (for example, a human B cell line or a swine B cell line), or an avian B cell line. In some examples, the B cell line is the human RAI B cell line. In other examples, the B cell line is another human B cell line, such as but not limited to, the Z138, KMS 11 GFP, KMS 11 tdtomato-luc, U-2932, RPMI 8226, SU-DHL4, SU-DHL4 tdtomato, BC-1, BC1 Flue, Ramos, or Ramos tdTomato luc2 B cell line.
[0391] C. Isolated B Cells
[0392] Also described are isolated B cells that include an integration plasmid disclosed herein integrated into the genome of the B cell.
[0393] In some aspects, the isolated B cells are primary B cells. In some examples, the primary B cells are
[0394] 35 from any mammalian species or any avian species. In specific examples, the primary B cells are human B cells, swine B cells, or avian B cells. PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0395] 7950-112573-02
[0396] In other aspects, the isolated B cells are cells of a B cell line. In some examples, the B cell line is a mammalian B cell line or an avian B cells line. In specific examples, the mammalian B cell line is a human B cell line or a swine B cell line. In a particular example, the B cell line is the human RAI B cell line. In other particular examples, the B cell line is another human B cell line, such as but not limited to the Z138,
[0397] 5 KMS11 GFP, KMS11 tdtomato-luc, U-2932, RPMI 8226, SU-DHL4, SU-DHL4 tdtomato, BC-1, BC1 Flue, Ramos, or Ramos tdTomato luc2 B cell line.
[0398] In some aspects, the B cell expresses on its surface an antibody or antibody fragment encoded by the integration plasmid.
[0399] D. Methods of Generating High- ffinity Antibodies or Antibody Fragments
[0400] Further described are methods of generating an antibody or antibody fragment that has higher affinity for a target antigen than a parental antibody or antibody fragment that binds the same antigen. In some aspects, the method includes transfecting isolated B cells with an integration plasmid disclosed herein and a CRISPR / Cas9 plasmid, wherein the CRISPR / Cas9 plasmid encodes Cas9 and one or more guide RNAs that target the genomic locus of the integration plasmid; enriching the B cells through antibiotic
[0401] 15 selection; and passaging the transfected B cells for multiple passages. In some aspects, the method further includes measuring affinity of the passaged B cells to the antigen. Any method for measuring the affinity of an antibody for a target protein can be used, such as, but not limited to, ELISA, competition radioimmunoassay, BLI, Octet, surface plasmon resonance, or flow cytometry.
[0402] In some aspects, the antibody fragment is an Fab fragment. In other aspects, the antibody fragment
[0403] 20 is an Fab, Fab', Fab'-SH, Flab'Iz, Fv, scFv, or VHH.
[0404] The antibody or antibody fragment can bind any protein of interest. In some aspects, the antibody is specific for a protein of a microorganism, such as a viral, bacterial, fungal, protozoal, or parasitic antigen. In other aspects, the antibody is specific for a tumor antigen.
[0405] In some examples, the antibody or antibody fragment binds a viral protein, such as an antigen of an
[0406] 25 RNA virus (e.g., an antigen of a negative-sense single-stranded RNA virus, a positive-sense single-stranded RNA virus or a double-stranded RNA viruses) or a DNA virus (e.g., an antigen of a single-stranded DNA viruses or a double-stranded DNA viruses). In specific examples, the viral protein is an influenza virus protein, such as influenza virus hemagglutinin (HA) protein or neuraminidase (NA) protein.
[0407] In some aspects, the CRISPR / Cas9 plasmid includes one or more guide RNAs that target the Hl 1 locus of human chromosome 22q. In some examples, the CRISPR / Cas9 plasmid is the p276 plasmid (Addgene ID 164850). In specific examples, the nucleotide sequence of the CRISPR / Cas9 plasmid is at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 3, or includes or consists of SEQ ID NO: 3. In other aspects, the CRISPR / Cas9 plasmid includes one or more guide RNAs that target the AAVS1 locus of human chromosome 19q. In yet other
[0408] 35 aspects, the CRISPR / Cas9 plasmid includes one or more guide RNAs that target another known stable nonimmunoglobulin locus (such as any one of the loci described in Pellenz et al., Hum Gene Ther 39(7):814- 828, 2019). PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0409] 7950-112573-02
[0410] In some aspects, the isolated B cells are primary B cells. In some examples, the primary B cells are from any mammalian species or any avian species. In specific examples, the primary B cells are human B cells, swine B cells, or avian B cells.
[0411] In other aspects, the isolated B cells are cells of a B cell line. In some examples, the B cell line is a
[0412] 5 mammalian B cell line or an avian B cells line. In specific examples, the mammalian B cell line is a human B cell line or a swine B cell line. In a particular example, the B cell line is the human RAI B cell line. In other particular examples, the B cell line is another human B cell line, such as but not limited to the Z138, KMS11 GFP, KMS11 tdtomato-luc, U-2932, RPMI 8226, SU-DHL4, SU-DHL4 tdtomato, BC-1, BC1 Flue, Ramos, or Ramos tdTomato luc2 B cell line.
[0413] The step of enriching the B cells through antibiotic selection includes culturing the B cells in the presence of antibiotic (the antibiotic corresponding to the antibiotic resistance cassette of the integration plasmid) to select for cells that have obtained antibiotic resistance from the integration plasmid.
[0414] In some aspects of the method, the B cells are passaged at least three times, at least four times, at least five times, or at least six times. Affinity of the antibody or antibody fragment expressed by the B cell
[0415] 15 can be measured following one or more passages to evaluate whether the antibody has increased binding affinity compared to the parental antibody. Measurements of binding affinity can help determine how many times the isolated B cells should be passaged.
[0416] E. Methods of Displaying Antibodies or Antibody Fragments on B Cells
[0417] Also described herein are methods of displaying an antibody or an antibody fragment on the surface
[0418] 20 of B cells. In some aspects, the method includes transfecting B cells with a nucleic acid molecule encoding in the 5’ to 3’ direction: a membrane localization sequence; the antibody or antibody fragment; and a transmembrane helix. In some aspects, an antibody fragment is displayed on the B cells. In some examples, the antibody fragment is an Fab. In particular examples, the Fab includes a light chain and a heavy chain separated by a linker. In other aspects, a full antibody (Fab and Fc region) is displayed on the B cells.
[0419] 25 In some aspects of the method, the nucleic acid molecule further includes a promoter. In some examples, the promoter is a human EFla promoter. In particular examples, the human EFla promoter includes a transcription start site and an intronic enhancer (the latter of which is spliced out during mRNA maturation).
[0420] In some aspects, the nucleic acid molecule further encodes a protein tag. In some examples, the protein tag is a FLAG tag. In other examples, , the peptide tag includes streptavidin, STREP-TACTIN, STREP-TAG, His (e.g., 6XHis), HBH, S-tag, GST, HA, MBP, TAP, CBP, TRX, bacteriophage V5 epitope, bacteriophage T7 epitope, Myc, or VSV-G.
[0421] In some aspects of the method, the nucleic acid sequence encoding the membrane localization signal is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%
[0422] 35 identical to SEQ ID NO: 5, or includes or consists of SEQ ID NO: 5. In some aspects, the amino acid sequence of the membrane localization signal is at least 80%, at least 85%, at least 90%, at least 95%, at PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0423] 7950-112573-02 least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 12, or includes or consists of SEQ ID NO: 12.
[0424] In some aspects, the linker (separating the light chain and heavy chain of the Fab) is a glycine- serine-alanine-threonine linker. In some examples, the amino acid sequence of the glycine-serine-alanine-
[0425] 5 threonine linker comprises SEQ ID NO: 14.
[0426] In some aspects, the nucleic acid sequence encoding the Fab light chain is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO: 6. In some examples, the nucleic acid sequence encoding the Fab light chain includes or consists of SEQ ID NO: 6. In some aspects, the amino acid sequence of the Fab light chain is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO: 13. In some examples, the amino acid sequence of the Fab light chain includes or consists of SEQ ID NO: 13.
[0427] In some aspects, the nucleic acid sequence encoding the Fab heavy chain is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO: 8. In some examples, the nucleic acid sequence encoding the Fab heavy chain includes or consists of
[0428] 15 SEQ ID NO: 8. In some aspects, the amino acid sequence of the Fab heavy chain is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to SEQ ID NO: 15. In some examples, the amino acid sequence of the Fab heavy chain includes or consists of SEQ ID NO: 15.
[0429] In some aspects, the transmembrane helix is an MHC class I transmembrane helix. In some
[0430] 20 examples, the nucleic acid sequence encoding the transmembrane helix is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 10, or includes or consists of SEQ ID NO: 10. In some examples, the amino acid sequence of the transmembrane helix is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 17, or includes or consists of SEQ ID NO: 17.
[0431] 25 In some aspects, the nucleotide sequence of the nucleic acid molecule is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SEQ ID NO: 4. In some examples, the nucleotide sequence of the nucleic acid molecule includes or consists of SEQ ID NO: 4.
[0432] V. Overview of Aspects
[0433] Aspect 1. An integration plasmid capable of integrating at a specific genomic locus, comprising: a heterologous gene operably linked to a first promoter; a somatic hypermutation (SHM) enhancer element upstream of the first promoter and the
[0434] 35 heterologous gene; a first nucleic acid sequence homologous to genomic sequences at the integration site, wherein the first nucleic acid sequence is upstream of the SHM enhancer element; and PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0435] 7950-112573-02 a second nucleic acid sequence homologous to genomic sequences at the integration site, wherein the second nucleic acid sequence is downstream of the heterologous gene and a polyadenylation (poly-A) signal sequence.
[0436] 5 Aspect 2. The integration plasmid of aspect 1, wherein the genomic locus is a stable, nonimmunoglobulin locus.
[0437] Aspect 3. The integration plasmid of aspect 1 or aspect 2, wherein the genomic locus is the Hl l locus of human chromosome 22.
[0438] Aspect 4. The integration plasmid of aspect 1 or aspect 2, wherein the genomic locus is the AAVS1 locus of human chromosome 19.
[0439] Aspect 5. The integration plasmid of any one of aspects 1-4, wherein the heterologous gene
[0440] 15 encodes an antibody fragment.
[0441] Aspect 6. The integration plasmid of aspect 5, wherein the heterologous gene further encodes a membrane localization signal, a transmembrane helix, and a linker between the antibody fragment and the transmembrane helix.
[0442] 20
[0443] Aspect 7. The integration plasmid of aspect 6, wherein the linker comprises a peptide tag.
[0444] Aspect 8. The integration plasmid of any one of aspects 5-7, wherein the antibody fragment is a single-chain antibody binding fragment (Fab).
[0445] 25
[0446] Aspect 9. The integration plasmid of aspect 8, wherein the Fab comprises a Fab light chain and a Fab heavy chain and a linker between the Fab light chain and the Fab heavy chain.
[0447] Aspect 10. The integration plasmid of any one of aspects 5-9, wherein the antibody fragment binds a viral protein.
[0448] Aspect 11. The integration plasmid of aspect 10, wherein the viral protein is an influenza virus protein.
[0449] 35 Aspect 12. The integration plasmid of aspect 11, wherein the influenza virus protein is hemagglutinin (HA). PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0450] 7950-112573-02
[0451] Aspect 13. The integration plasmid of any one of aspects 1-12, wherein the nucleotide sequence of the heterologous gene is at least 85% identical to SEQ ID NO: 4.
[0452] Aspect 14. The integration plasmid of any one of aspects 1-13, wherein the nucleotide
[0453] 5 sequence of the heterologous gene comprises or consists of SEQ ID NO: 4.
[0454] Aspect 15. The integration plasmid of any one of aspects 1-14, wherein the nucleotide sequence of the SHM enhancer element is at least 85% identical to any one of SEQ ID NOs: 1 and 59-68.
[0455] Aspect 16. The integration plasmid of any one of aspects 1-15, wherein the nucleotide sequence of the SHM enhancer element comprises or consists of any one of SEQ ID NOs: 1 and 59-68.
[0456] Aspect 17. The integration plasmid of any one of aspects 1-16, wherein the first promoter is the human EFla promoter.
[0457] 15
[0458] Aspect 18. The integration plasmid of any one of aspects 1-17, wherein: the first nucleic acid sequence homologous to genomic sequences at the integration site is about 200 to about 800 nucleotides in length; and / or the second nucleic acid sequence homologous to genomic sequences at the integration site is about
[0459] 20 200 to about 800 nucleotides in length.
[0460] Aspect 19. The integration plasmid of any one of aspects 1-18, wherein: the first nucleic acid sequence homologous to genomic sequences at the integration site is about 700 nucleotides in length; and / or
[0461] 25 the second nucleic acid sequence homologous to genomic sequences at the integration site is about 700 nucleotides in length.
[0462] Aspect 20. The integration plasmid of any one of aspects 1-19, further comprising an antibiotic resistance cassette.
[0463] Aspect 21. The integration plasmid of aspect 21, wherein the antibiotic resistance cassette is a puromycin resistance cassette.
[0464] Aspect 22. The integration plasmid of aspect 20 or aspect 21, wherein the antibiotic resistance
[0465] 35 cassette comprises an antibiotic resistance gene operably linked to a second promoter. PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0466] 7950-112573-02
[0467] Aspect 23. The integration plasmid of aspect 22, wherein the second promoter is a PGK promoter.
[0468] Aspect 24. The integration plasmid of any one of aspects 1-23, wherein the nucleotide
[0469] 5 sequence of the integration plasmid is at least 95% identical to SEQ ID NO: 2.
[0470] Aspect 25. The integration plasmid of any one of aspects 1-24, wherein the nucleotide sequence of the integration plasmid comprises or consists of SEQ ID NO: 2.
[0471] Aspect 26. A kit comprising the integration plasmid of any one of aspects 1-25, and one or more of: a CRISPR / Cas9 plasmid encoding Cas9 and one or more guide RNAs that target the genomic locus; isolated B cells;
[0472] B cell culture media;
[0473] 15 tissue culture flask(s); and buffer.
[0474] Aspect 27. The kit of aspect 26, wherein the CRISPR / Cas9 plasmid is the p276 plasmid having the nucleotide sequence of SEQ ID NO: 3.
[0475] 20
[0476] Aspect 28. The kit of aspect 26 or aspect 27, wherein the isolated B cells are primary B cells.
[0477] Aspect 29. The kit of aspect 26 or aspect 27, wherein the isolated B cells are cells of a B cell line.
[0478] 25
[0479] Aspect 30. The kit of aspect 29, wherein the B cell line is the RAI B cell line.
[0480] Aspect 31. An isolated B cell comprising the integration plasmid of any one of aspects 1-25 integrated into the genome of the cell.
[0481] Aspect 32. The isolated B cell of aspect 31, which is a primary B cell.
[0482] Aspect 33. The isolated B cell of aspect 31, which is a B cell line.
[0483] 35 Aspect 34. The isolated B cell of aspect 33, wherein the B cell line is the RAI B cell line. PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0484] 7950-112573-02
[0485] Aspect 35. The isolated B cell of any one of aspects 31-34, which is a human B cell, an avian B cell, or a swine B cell.
[0486] Aspect 36. The isolated B cell of any one of aspects 31-35, wherein the B cell expresses on its
[0487] 5 surface an antibody fragment encoded by the integration plasmid.
[0488] Aspect 37. A method of generating an antibody or antibody fragment that has higher affinity for a target antigen than a parental antibody or antibody fragment that binds the same antigen, comprising: transfecting isolated B cells with the integration plasmid of any one of aspects 1-25 and a CRISPR / Cas9 plasmid, wherein the CRISPR / Cas9 plasmid encodes Cas9 and one or more guide RNAs that target the genomic locus of the integration plasmid; enriching the B cells through antibiotic selection; and passaging the transfected B cells for multiple passages.
[0489] 15 Aspect 38. The method of aspect 37, further comprising measuring affinity of the passaged B cells to the antigen.
[0490] Aspect 39. The method of aspect 37 or aspect 38, wherein the antibody fragment is a singlechain antibody binding fragment (Fab).
[0491] 20
[0492] Aspect 40. The method of any one of aspects 37-39, wherein the antibody or antibody fragment binds a viral protein.
[0493] Aspect 41. The method of aspect 40, wherein the viral protein is an influenza virus protein.
[0494] 25
[0495] Aspect 42. The method of aspect 41, wherein the influenza virus protein is hemagglutinin
[0496] (HA).
[0497] Aspect 43. The method of any one of aspects 37-42, wherein the CRISPR / Cas9 plasmid is the p276 plasmid having the nucleotide sequence of SEQ ID NO: 3.
[0498] Aspect 44. The method of any one of aspects 37-43, wherein the isolated B cells are primary B cells.
[0499] 35 Aspect 45. The method of any one of aspects 37-43, wherein the isolated B cells are cells of a B cell line. PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0500] 7950-112573-02
[0501] Aspect 46. The method of any one of aspects 37-45, wherein the B cells are human B cells, avian B cells, or swine B cells.
[0502] Aspect 47. The method of any one of aspects 37-46, wherein the B cells are passaged at least
[0503] 5 three times, at least four times, at least five times, or at least six times.
[0504] Aspect 48. A method of displaying a single-chain antibody binding fragment (Fab) on the surface of B cells, comprising transfecting B cells with a nucleic acid molecule encoding in the 5’ to 3’ direction: a membrane localization sequence; the Fab, wherein the Fab comprises a light chain and a heavy chain separated by a linker; and a transmembrane helix.
[0505] Aspect 49. The method of aspect 48, wherein the nucleic acid molecule further comprises a
[0506] 15 promoter.
[0507] Aspect 50. The method of aspect 49, wherein the promoter is a human EFla promoter.
[0508] Aspect 51. The method of any one of aspects 48-50, wherein the nucleic acid molecule further
[0509] 20 encodes a protein tag.
[0510] Aspect 52. The method of aspect 51, wherein the protein tag is a FLAG tag.
[0511] Aspect 53. The method of any one of aspects 48-52, wherein the amino acid sequence of the
[0512] 25 membrane localization sequence comprises SEQ ID NO: 12.
[0513] Aspect 54. The method of any one of aspects 48-53, wherein the linker is a glycine-serine- alanine -threonine linker.
[0514] Aspect 55. The method of aspect 54, wherein the amino acid sequence of the GS linker comprises SEQ ID NO: 14.
[0515] Aspect 56. The method of any one of aspects 48-55, wherein the amino acid sequence of the Fab light chain comprises SEQ ID NO: 13 and / or the amino acid sequence of the Fab heavy chain comprises
[0516] 35 SEQ ID NO: 15. PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0517] 7950-112573-02
[0518] Aspect 57. The method of any one of aspects 48-56, wherein the transmembrane helix is a major histocompatibility complex I (MHCI) transmembrane helix.
[0519] Aspect 58. The method of any one of aspects 48-57, wherein the amino acid sequence of the
[0520] 5 transmembrane helix comprises SEQ ID NO: 17.
[0521] Aspect 59. The method of any one of aspects 48-58, wherein the nucleotide sequence of the nucleic acid molecule is at least 95% identical to SEQ ID NO: 4.
[0522] Aspect 60. The method of any one of aspects 48-59, wherein the nucleotide sequence of the nucleic acid molecule comprises or consists of SEQ ID NO: 4.
[0523] EXAMPLES
[0524] The following examples are provided to illustrate particular features of certain aspects of the
[0525] 15 disclosure, but the scope of the claims should not be limited to those features exemplified.
[0526] Example 1: Materials and Methods
[0527] This example describes the materials and experimental procedures for the studies described in Examples 2-10.
[0528] Bacterial strains, mammalian cell lines, growth media, DNA constructs
[0529] E. coli DH5a cells were used for all transformations and bacterial cultures and grown in either LB or 2xYT medium. Expi293 cells were used for transient transfection and initial construct testing and were grown in Expi293 Expression Medium. RAI cells (ATCC CRL-1596) were used for SHM experiments and
[0530] 25 grown in RPMI 1640 Medium supplemented with 10% FBS. The oligonucleotides used in this study are listed in Table 1, the synthesized DNA fragments are listed in Table 2, the plasmid DNAs used in this study are listed in Table 3 and the B cell lines generated in this study are listed in Table 4.
[0531] Table 1: Oligonucleotides PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0532] 7950-112573-02
[0533] Table 2: Synthesized DNA fragments PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0534] 7950-112573-02
[0535] Table 3: Plasmid DNA constructs
[0536] Table 4: B cell lines generated PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0537] 7950-112573-02
[0538] Preparation of plasmid DNA samples for transfection of RAI cells
[0539] Using the QIAprep® Spin Miniprep Kit, plasmids for transfection were prepared from E. coli cultures. Initially, E. coli containing the desired plasmids were grown in 2 mL of 2x YT media, inoculated from glycerol stock. The next day, this culture was diluted 1% into 30 mL of LB media in a 125 mL
[0540] 5 Erlenmeyer flask and incubated at 37°C with 220 rpm shaking. After 24 hours, the culture was collected in 50 mL conical tube and centrifuged at 10,000 x g for 5 minutes, the supernatant was decanted, and the pellet was resuspended in 800 pL of Pl buffer with RNase. Following vortexing, 1 mL of P2 buffer was added and the mixture was lysed for 5-7 minutes with shaking at 150 rpm. Then, 1.4 mL of N3 buffer was added, the tube was inverted for mixing, and left with shaking for 3 minutes to form a precipitate. This was centrifuged at 12,000 x g for 5 minutes, and the clear supernatant was distributed among two spin columns for purification, following the kit's protocol. The DNA was then serially eluted from the two columns using 45 pL of DNase-free water, using the pass-through of the first column as the elution media for the next column.
[0541] Transfection of RAI Cells
[0542] 15 RAI cells were cultured to a density of 1.5-2.0 million cells / mL using growth media RPMI 1640 + 10% FBS. For transfection, the Minis electroporation buffer was pre-warmed by shaking at 37°C for 10 minutes. Meanwhile, 7 million cells were centrifuged at 100 x g for 10 minutes at room temperature, resuspended in 100 pL Mirus electroporation buffer, and combined with 20 pL DNA solution, containing 10 pg integration vector and 10 pg Cas9 / sgRNA plasmid. A final volume of 110 pL was then transferred to a
[0543] 20 0.2 cm electroporation cuvette.
[0544] Electroporation was performed using the BioRad XCell Electroporation system at 133V, 950 pF, infinite resistance, in a 0.2 cm cuvette, at room temperature. Post-electroporation, cells rested at room temperature for 8-10 minutes, then were resuspended in 1 mL RPMI + 10% FBS and transferred to a 6-well plate in a total volume of 3 mL RPMI + 10% FBS. Cells were then grown in this vessel for 48 hours,
[0545] 25 washed, and resuspended in 10 mL RPMI + 10% FBS media in a T025 flask, and grown for 48 hours.
[0546] Selection of RAI Cells
[0547] Transfected RAI cells, post-cultivation in T025 flasks, were passaged into fresh T025 flasks at a density of 0.5 x 106cells / mL in 10 mL volume of growth media, supplemented with 0.3 pg / mL puromycin. This selection medium was utilized to grow the cells for 72 hours. Subsequently, cell cultures were diluted to 15 mL by adding 5 mL of growth media and incubated for an additional 48 hours. Following this, cells were further diluted to 20 mL with 5 mL of growth media and again grown for 48 hours. The dilution process continued, increasing the volume to 30 mL with an addition of 10 mL growth media for another 48 hours, and finally to 40 mL with an additional 10 mL of growth media for a last 48-hour growth period. At
[0548] 35 the conclusion of the selection regimen, cells exhibiting puromycin resistance were deemed "enriched" for the integration event. These enriched cells were cryopreserved at a density of 1 x 107cells / mL. The PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0549] 7950-112573-02 cry opreservation protocol involved resuspending cells in growth media amended with 5% DMSO, followed by a 30-minute chill on dry ice before storage in the vapor phase of liquid nitrogen.
[0550] RNA isolation and cDNA preparation
[0551] 5 Total RNA from sorted cells was isolated using the PureLink™ RNA Mini Kit, following the manufacturer's instructions. Cells (5 x 106) were resuspended in 600 pL Lysis Buffer with 1% BME, vortexed for cell lysis, and homogenized via a 21-gauge needle. After adding 1.5 volumes of ethanol (96- 100%) and vortexing, the mixture was column-purified and eluted in 18 pL RNase-free water. The isolated RNA underwent reverse transcription with the Superscript III First Strand Synthesis Kit, per protocol. For this, 16 pL of the isolated RNA solution was combined with the provided set of 2 pL poly-dT primer and 2 pL dNTPs, heated at 65°C for 5 minutes, then cooled on ice for 1 minute. A mix of 2 pL reverse transcriptase, 8 pL MgCh, 4 pL DTT, 4 pL 10X RT buffer, and 2 pL RNase Out (total 20 pL) was prepared. Both mixes were combined, vortexed, and incubated for 50 minutes at 50°C, then denatured at 85°C for 5 minutes. Adding 2 pL RNase H and incubating for 20 minutes at 37°C finalized the cDNA mixture that was
[0552] 15 ready for use in PCR.
[0553] PCR Optimization and Amplification from cDNA
[0554] PCR optimization for individual cDNA samples involved initial dilution of the cDNA to 10%, which then served as a template in PCR mixtures. In these mixtures, each primer was serially diluted 2-fold
[0555] 20 from a starting concentration of 1.6 pM. For the amplification of CR9114-based amplicons, the optimization process utilized primers SB 274A and SB 212B. Similarly, eGFP-based amplicons were obtained through optimization with primers SB 193A and SB 193B.
[0556] Microscopy
[0557] 25 Images were produced by the Bio- Rad ZOE Fluorescent Cell Imager.
[0558] FACS Analysis and Sorting
[0559] To identify FLAG-tagged Fab displays, APC-conjugated Anti-FLAG antibodies (Biolegend 637308) and FITC-conjugated Anti-His antibodies (ICLLab CHIS-45A-Z) were utilized, targeting the 6xHis tag on hemagglutinins H3 and H5 (SinoBiological 40859-V08H and 11700-V08H, respectively). FACS analysis and sorting were performed on a BD FACSMelody system. Binding versus expression was quantified by establishing a linear regression between the FITC and APC signals. Cells displaying APC fluorescence two logarithms higher than untransfected controls were classified as "Expressing". High-binding cells were isolated using triangular gating, adjusted so the gate's hypotenuse was parallel to the regression line and
[0560] 35 elevated to include no more than 2% of parent gate events. For sorting, 6-8 million cells were processed at the maximum flow rate and collected into a collection tube with 3 mL of growth media, all maintained at 5°C. PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0561] 7950-112573-02
[0562] Third generation sequencing data analysis
[0563] Multi-FASTA files corresponding to PacBio sequencing data were analyzed using Python (3.10.12). Briefly, multi-FASTA files were imported using BioPython (1.59) and collapsed to unique reads and their
[0564] 5 occurrence. Unique reads with an occurrence threshold below 15 were discarded. Surviving unique reads were then aligned to the reference sequence using BioPython with the following alignment parameters: match_score = 2, mismatch_score =-3, open_gap_score= -5, extend_gap_score = -2, query _right_open_gap_score = 0, query_left_open_gap_score = 0, query_right_extend_gap_score = -1, query _left_extend_gap_score = -1. Alignments which introduced gaps in the interior of the alignment were discarded. Surviving unique reads were then used as a basis for mutational profile calculations. Data handling was done using Pandas (2.2.2). Histogram generation was done using matplotlib (3.9.0). Sunburst plots and pie charts were generated using Plotly (5.22.0)
[0565] Structural analysis and modeling
[0566] 15 Crystal structures of hemagglutinin H5 and CR9114 (PDB: 4FQI) were visualized using PyMOL. The predicted crystal structure of the chimeric CR9114 surface-displayed protein was calculated using AlphaFold3 and visualized using PyMOL.
[0567] Calculation of EC50
[0568] 20 The EC50 of surface-displayed CR9114 and its variants was determined by fitting the sigmoidal data to a dose -response curve using the Hill equation protocol in OriginPro 2024, and subsequently extracting the EC50 parameter.
[0569] Construction of plasmid DNAs
[0570] 25 PrimeStar HS polymerase, provided as a 2x Mastermix by Takara, was utilized for PCR. The integration plasmid, p274, was sourced from Addgene (ID 164851), and the Cas9 / sgRNA plasmid, p276, was also sourced from Addgene (ID 164850). Cloning on p274 was performed by extracting the proXIV sequence from RAI genomic DNA using primers SB 168A and SB 168B, which was then cloned directly upstream of the EFlex promoter via Gibson Assembly. Single-point mutagenesis to create an eGFP* variant was conducted using the KLD method with primers SB 171A and SB 17 IB. To substitute the eGFP component of the native p274 sequence with novel genes, p274 was linearized using primers SB 209A and SB 209B. Fab constructs, synthesized as gBlocks by IDT, were amplified with primers SB 274A and SB 212B and cloned into p274 using Gibson Assembly. For the purpose of Fab evolution, p274 was linearized with SB 209A and SB 275B, introducing an optional BsiWI cut site for restriction / ligation cloning. PCR
[0571] 35 bands were purified using a combination of the QIAprep® Spin Miniprep Kit and GeneJET Gel Extraction Kit. PCR products, in 30 pL reactions, were electrophoresed on a 1% agarose gel at 130V for 14 minutes. The fluorescent bands were then excised and dissolved in 400 pL of GeneJET binding buffer, incubated at PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0572] 7950-112573-02
[0573] 55°C for 10 minutes, and vortexed. To this, 300 (iL of 100% ethanol was added and the mixture was vortexed again before being transferred to a spin column. The purification process involved sequential washing with 500 u L of PB Buffer and 750 p L of PE Buffer from the QIAprep® kit, and then 750 p L of Wash Buffer from the GeneJET kit. After centrifugation to dry the spin column, DNA was eluted in DNase-
[0574] 5 free water. pSBl: pSB2 was linearized by PCR using the oligonucleotides SB 169 A / B. The proXIV element was amplified by PCR from the genome of RAI using the oligonucleotides SB 168 A / B. The amplified proXIV fragment was inserted into linearized pSB2 by Gibson assembly to afford pSBl. pSB2: p274 was linearized by PCR using the oligonucleotides SB 171 A / B, which encode the eGFP* knockout mutation. The linearized p274 was subject to blunt-end ligation by Kinase-Ligase-Dpnl (KLD), to afford the plasmid pSB2
[0575] 15 pSB3: p274 was linearized by PCR using the oligonucleotides SB 209A / SB 275B. A gBlock of the light / heavy fusion protein Fab-CR9114 (PDB: 4FQI) was codon optimized for H. sapiens and amplified in three steps to incorporate the secretion signal, using the oligonucleotides SB 246A I SB 274B, followed by GL 7B / SB 274B, lastly by SB 274 A / B. A gBlock of the linker-FLAG-MHCI helix domain of the chimera was codon optimized for H. sapiens and amplified using the oligonucleotides SB 275A / SB 252B. The two
[0576] 20 fragments were inserted into the linearized p274 by Gibson assembly to afford the plasmid pSB3. pSB4: pSBl was linearized using the oligonucleotides SB 209A / SB 275B. The complete chimeric Fab- CR9114-Linker-FLAG-MHCI helix gene was amplified by PCR of pSB3 using oligonucleotides SB 274A I SB 252B. The gene was inserted into linearized pSBl by Gibson assembly to afford the plasmid pSB4.
[0577] 25
[0578] Influenza virus propagation
[0579] Madin-Darby canine kidney (MDCK) cells (American Type Culture Collection, CCL-34) and human embryonic kidney (HEK) 293T were grown and maintained in Minimum Essential Medium supplemented with 10% fetal bovine serum (Gibco), 100 U / mL penicillin and 100 pg / mL streptomycin (Gibco), and lx GlutaMAX (Gibco) at 37°C, 5% CO2, and 95% humidity. To generate a seed stock, eight plasmids encoding each of the segments of influenza A / Puerto Rico / 8 / 1934 (H1N1) virus were transfected into HEK 293T using jetOPTIMUS® (Polyplus) following the manufacturer’s protocol. 48 hours posttransfection, media were collected as seed stock and TCID50 was tittered with MDCK cells. A T75 of full confluence MDCK cells were infected with a multiplicity of infection of 0.0001 of virus and media was
[0580] 35 collected after 24 hours. The virus was aliquoted, kept in -80°C and tittered for TCID50.
[0581] The hemagglutination inhibition assay PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0582] 7950-112573-02
[0583] 50 (il of 2-fold serial dilutions of virus in Dulbecco's PBS (DPBS) was mixed with 50 jrL 0.5% turkey RBCs (Fisher, 50-203-4867) in a round-bottom 96-well plate (Greiner Bio-One) and incubated at 4°C for 1 hour. HA titer was determined to be 128 as reciprocal of highest dilution providing full hemagglutination.
[0584] 5 25 pL of 2-fold dilutions of antibodies in DPBS were prepared with the highest and lowest final concentrations at 100 pg / mL and 0.19 pg / ml, respectively. This was mixed with 25 pL of H1N1 virus with 4 HA units in round-bottom 96-well plate and incubated for 30 minutes at room temperature. Then 0.5% turkey RBCs were added and the mixture was incubated at 4°C for one hour. All antibodies were tested in triplicate.
[0585] Microneutralization assay
[0586] MDCK cells were seeded in 96-well plates at 8,000 cells per well and incubated overnight.
[0587] Antibodies were serially diluted at 2-fold, starting from 100 pg / mL to 0.78 pg / mL, and then mixed with 100 TCID50of influenza A HINA virus in 100 pL of infection medium (Minimum Essential Medium
[0588] 15 supplemented with 10 mM HEPES, 0.125% BSA, lx GlutaMAX (Gibco) and 1 pg / mL TPCK-treated trypsin). The virus-antibody mixture was incubated at 37°C, 5% CO2 for 1 hour. MDCK cells were washed with PBS once and infected with 100 pl virus-antibody mixture for 1 hour at 37°C, 5% CO2. Then the mixture was removed and replaced with 100 pL of infection medium. The plates were kept at 37°C, 5% CO2 for 24 hours and a fluorescence-focused assay was performed. The cells were washed with PBS, replaced
[0589] 20 with 100 pL methanol and kept at -20°C for 30 minutes. Then the cells were washed twice with PBS, 3% BSA in PBS was added, and cells were incubated at room temperature for 1 hour. The cells were stained with influenza A NP antibody (D67J) conjugated with FITC (Invitrogen, cat# MAI-7322) at 1 pg / mL for two hours at room temperature. Then the cells were washed twice with PBST, one time with PBS, and stained with DAPI for 10 minutes. The plates were then washed with PBS twice and imaged with Biotek
[0590] 25 Cytation 5 for analysis. qPCR experiments
[0591] Genomic DNA (gDNA) was isolated from enriched RAl-eGFP cells using a PureLink™ Genomic DNA Mini Kit, and quantified by NanoDrop. The stock gDNA (10 ng / pL) was used to generate an 8-point, 1:2 serial dilution series for the sample reactions. Primers specific for 100-200 bp regions of both the eGFP coding sequence and the chromosome 22 homology cassette were designed using the IDT PrimerQuest™ Tool. Primer pairs were synthesized at 100 pM each and combined into a final primer mixture at 0.83 pM each. Reactions were performed in 0.2 mL PCR tubes in 10 pL total volume, using PowerUp™ SYBR™ Green Master Mix. For each reaction strip (8 wells), 10 pL of primer mix was prepared and aliquoted into
[0592] 35 seven of eight tubes in a strip. 20 pL of primer mix containing 10 ng / pL gDNA was added to the first tube of a PCR tube strip and a two-fold serial dilution was performed by transferring 10 pL from the first tube into 10 pL water in the second tube, and so on through the series. Using a P20 multichannel pipette, 5 pL of PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0593] 7950-112573-02 each diluted gDNA / primer mix was transferred into separate PCR tubes already containing 5 pL of PowerUp™ SYBR™ Green Master Mix, yielding a final reaction volume of 10 pL per well with 10 ng to 0.078 ng input gDNA. Amplicons corresponding to the eGFP target and the chromosome 22 homology cassette were PCR-amplified, gel-purified, and quantified. Each amplicon was diluted in an identical 1 :2
[0594] 5 series starting at 1 ng / pL and processed in parallel with sample reactions to generate standard curves. All qPCR reactions were performed on a QuantStudio 7 Pro Real-Time PCR System using the manufacturer’s recommended cycling conditions. Each run began with a two minute UNG incubation at 50°C, followed by a two minute polymerase activation step at 95°C. Amplification consisted of 40 cycles of denaturation at 95°C for 15 seconds and combined annealing / extension at 60°C for one minute. Quantification cycle (Cq) values were exported and plotted against logio[Standard (ng / pL)] to generate standard curves. Sample copy numbers were calculated by interpolation from the linear regression of the standards, and normalized to input gDNA.
[0595] Antibody overexpression and purification
[0596] 15 Antibody genes encoding CR9114 and its eight variants were cloned into the pVitro vector — each construct harboring a hygromycin-resistance cassette — and transformed into E. coli DH5a (pV-1 through pV-8). To maximize plasmid yield without impairing growth, bacterial cultures were maintained in 37.5-75 pg / mL hygromycin, after which plasmids were purified and prepared for mammalian expression. Expi293 cells were seeded at 0.5 x 106cells / mL in 30 mL Expi293 medium and expanded to ~6-7 x 106cells / mL,
[0597] 20 then diluted to 3 x 106cells / mL 24 hours before transfection. For each milliliter of cells, 4 pL ExpiFectamine™ 293 reagent was mixed with 46 pL Opti-MEM and incubated for 5 minutes, while separately 1 pg of plasmid DNA was diluted in 50 pL Opti-MEM; the two mixtures were then combined and incubated for an additional 20 minutes. During this interval, cells were counted and adjusted to 2-3 x 106cells / mL in 125 mL flasks, then returned to the incubator. The transfection mixtures were then added to the
[0598] 25 cultures, which were incubated for 18 hours in the incubator set to 37°C with 8% CO2. Finally, Enhancer 1 (6 pL / mL) and Enhancer 2 (60 pL / mL) were added sequentially, and cells were cultured for a further 3-5 days to allow robust antibody expression. Antibodies were purified from the clarified culture supernatant by Protein A affinity chromatography (MabSelect, Cytiva). First, expression medium was harvested from 125 mL flasks and clarified by two successive spins at 4,200 x g for 10 minutes, with the supernatant transferred to fresh tubes after each centrifugation. The clarified supernatant was then filtered through a 40 pm mesh into a chilled tube. Meanwhile, 750-1,000 pL of mAbselect resin beads were pelleted at 1,500 x g for 5 minutes to remove the 20% ethanol storage buffer, then washed twice with 12 mL binding buffer (20 mM sodium phosphate, 150 mM NaCl, pH 7). The washed resin was resuspended in 1-3 mL binding buffer and combined with the chilled supernatant, then gently rocked at 4°C for two hours to bind antibody. The resin¬
[0599] 35 antibody slurry was applied to a gravity-flow column, washed with 20 mL of binding buffer (delivered as two 10 mL aliquots), and eluted with 5 mL of elution buffer (50 mM sodium phosphate, pH 3.0). Finally, the PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0600] 7950-112573-02 pooled eluate was concentrated using a 100 kDa MWCO, 4 mL Amicon concentrator and the antibody concentration was determined by Qubit fluorometry.
[0601] Antibody binding assays using ELISA
[0602] 5 ELISA assays were carried out in Nunc Maxisorp 96-well plates. Lyophilized H5 hemagglutinin (HA) was reconstituted to 0.2 mg / mL according to the manufacturer’s instructions (Sino Biological, A / Vietnam / 1194 / 2004), then diluted 1:200 in phosphate -buffered saline (PBS). One hundred microliters of the diluted HA solution were added to each well and the plate was incubated overnight at 4°C. The following day, wells were emptied by pipetting, washed once with 410 pL PBS, and blocked with 410 pL of 1% bovine serum albumin (BSA) in PBS for one hour at room temperature. After blocking, the solution was removed and wells were washed once with 410 pL PBS. Purified monoclonal antibodies (mAbs) were prepared in a separate dilution plate by first bringing each to 400 nM in 150 pL of 1% BSA / PBS, then performing two-fold serial dilutions across the plate, leaving one well blank as a no-antibody negative control. Seventy microliters of each mAb dilution were transferred to the antigen-coated plate and incubated
[0603] 15 for two hours at 37°C. Wells were then emptied and washed four times with 410 pL PBS before addition of goat anti-human IgGl-HRP secondary antibody at a 1:3,000 dilution in 1% BSA / PBS. After a two hour incubation at 37°C, wells were washed five times with 410 pL PBS. One hundred microliters of TMB substrate (Invitrogen) were added to each well and allowed to develop for 30 minutes at room temperature, then the reaction was stopped with 50 pL of 2 M H2SO4. Absorbance at 450 nm was measured using a
[0604] 20 Cytation5 Plate Reader.
[0605] Sequencing analysis for 1A02 evolution experiments
[0606] PacBio sequencing data for the 047-09_lA02 lineage were processed in Python (v3.10) leveraging Biopython (vl.81), pandas, and matplotlib. Initially, raw reads were oriented by pairwise alignment against
[0607] 25 annotated reference sequences using Biopython’s PairwiseAligner, then subjected to high-throughput mapping with Minimap2 (v2.26). Resulting SAM files were converted to sorted, indexed BAMs via SAM tools (vl.17). For mutation profiling, unique read sequences and their abundances were imported from pre-parsed CSV tables. Reads failing to meet an occurrence threshold (>15) or deviating by more than ±10% from the expected amplicon length were discarded. Remaining high-confidence reads were globally aligned to the reference, and base-level variants — including substitutions, insertions, and deletions — were catalogued. Coding-region nucleotide changes were translated to amino acid substitutions; synonymous mutations were excluded, and nonsynonymous events were tabulated to generate position-wise mutation frequencies and comprehensive amino acid-level mutation landscapes. To interrogate mutational linkage, nonsynonymous substitutions within each read were used to construct pairwise co-occurrence matrices. The
[0608] 35 most prevalent co-occurring mutation pairs were visualized as heatmaps, and hierarchical clustering of the co-occurrence data produced Newick-format trees delineating mutational lineages. Indel events were extracted from the aligned BAM files using pysam, with each insertion or deletion annotated by genomic PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0609] 7950-112573-02 position, size, and local sequence context. Aggregate indel counts and length distributions were displayed in bar-plot form to summarize insertion and deletion patterns across the dataset.
[0610] Genomic DNA sequencing analysis
[0611] 5 Genomic DNA was harvested from RAI EGFP* cells at early (passage 3) and late (passage 10) time points using the PureLink™ Genomic DNA Kit, following the manufacturer’ s protocol. On-target editing at the EGFP* locus was assayed by PCR amplification with primers SB1 / SB2, yielding a -700 bp fragment; the ARID 1 A locus on chromosome 1 served as an off-target control and was amplified in parallel. Amplicons were purified and subjected to paired-end Illumina MiSeq sequencing (2x250 bp reads). Raw reads were mapped to either the EGFP* or ARID 1 A reference sequences via Minimap2 (v2.26) using shortread alignment presets. Resulting SAM files were converted into sorted, duplicate-marked BAMs with SAM tools (vl.17). Only properly paired, uniquely mapped reads were retained for variant calling. Substitutions, insertions, and deletions were extracted using custom Python routines built on the pysam and Biopython libraries; high-confidence single-nucleotide variants were filtered by base quality (Q > 30).
[0612] 15 Mutation counts were tallied by nucleotide position and aggregated into 10-nt bins to reveal hotspot regions. Per-base coverage depth was likewise computed and binned to evaluate uniformity and to flag any undersequenced intervals within each amplicon.
[0613] For naive pool mutational analysis, amplicon library preparation and paired-end MiSeq sequencing followed the protocols established for on- and off-target mutation profiling, with sequencing depth increased
[0614] 20 to 25 million reads per sample to maximize coverage and sensitivity. Raw reads were aligned to the respective reference amplicons via Minimap2 (v2.26) using short-read parameters. Resulting SAM files were sorted, deduplicated, and indexed with SAMtools (vl.17), and only properly paired, uniquely mapped reads were retained. Custom Python routines built on pysam and Biopython parsed single-nucleotide substitutions, insertions, and deletions; high-confidence variants were filtered by base quality (Q > 30).
[0615] 25 Mutation frequencies were aggregated into 10-nucleotide bins for both substitutions and indels at each time point, and per-base coverage was computed to verify uniform sequencing depth across the targeted loci.
[0616] Example 2: Precise genome editing in human B cell lines to express reporter proteins from a stable B cell locus
[0617] The first goal was to develop an approach that would allow incorporation of defined genetic elements at a defined stable locus in B cell lines. Most commonly, lentivirus-based approaches have been used to randomly incorporate genetic elements in the genomes of B cell lines (Moffett et al., Sci. Immunol. 4, eaax0644). In addition to this, it has been significantly challenging to transfect B cell lines using lipid- based transfections or nucleofection techniques. Recently, CRISPR / Cas9 plasmids along with adeno-
[0618] 35 associated virus (AAV) containing a recombination cassette has also been used to replace endogenously encoded antibodies with antibodies in B cells with specific antigen targeting antibodies (Goossens et al., Proc. Natl. Acad. Sci. 95, 2463-2468). For the studies herein, it was first investigated whether a relatively PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0619] 7950-112573-02 simple, plasmid DNA-based and virus-free approach could be developed to edit genomes of B cell lines with a goal to recombinantly express reporter proteins (FIGS. 1A-1C) from a defined and stable genomic locus. Particularly, it was envisioned that a two plasmid system could be used: (i) CRISPR / Cas9 plasmid DNA containing genetic elements to express guide RNA and Cas9 targeting nuclear genomes, and (ii) a separate
[0620] 5 donor plasmid DNA containing homology regions to a specific locus on the B cell genome as well as DNA sequence encoding a defined integration cassette. Briefly, donor plasmid DNA (p274) and CRISPR / Cas9 plasmid DNA (p276) (Pryzhkova et al., Tissue Eng. Regen. Med. 17, 223-235) that were compatible for integration at the safe harbor Hl 1 locus (Chr.22: 31437323 - 31439351) at chromosome 22 were used in human stem cells. Importantly, this locus on chromosome 22 is a non-immunoglobulin locus in B cells with no reported instability or recombination (Tonegawa et ah, Nature 302, 575-581; Zhu et ah, Nucleic Acids Res. 42, e34-e34). Then, DNA formulations and electroporation conditions were optimized to transfect RAI B cell lines with plasmid DNAs (see Example 1). This was followed by developing recovery and selection conditions that would allow generation of RAI mutant cell lines where a defined region of the donor plasmid DNA p274 was integrated into the defined locus within chromosome 22 of B cell lines. The
[0621] 15 integration cassette from p274 contained a gene encoding eGFP driven by the EFla promoter, a puromycin resistance cassette driven by the PGK promoter, and ~ 700 bp homology regions encoding the Hl 1 integration site. Between these homology arms, an EFla promoter was placed to drive eGFP expression, and a PGK promoter was inserted to drive puromycin resistance cassette (puroR) expression. For determining the concentration of puromycin required for the selection of cell lines with genomic integrant, puromycin
[0622] 20 levels were titrated in the growth medium. Cell viability was evaluated using trypan blue staining and microscopic examination, as well as assessing the impact on growth rates via cell counts, all aiming to find an optimal puromycin concentration that allowed for selection without irreversibly compromising cell health or inducing irreversible morphological alterations. Using this approach, it was determined that the ideal starting concentration of puromycin for effective selection of RAI mutant cell lines was 0.3 pg / rnL. The
[0623] 25 RAI cells that were transfected with p274 and p276 were recovered and selected in growth medium containing 0.3 pg / ml of puromycin for 72 hours. The volume of growth medium was increased in a stepwise manner from 10 mL to 40 mL over the course of one week. The cells were then analyzed by microscopy and fluorescence activated cell sorting (FACS), to determine the levels of eGFP expression. As shown in FIGS. 1D-1E, at least 94% of the total cells demonstrated GFP signals. Integration of the desired recombinant DNA at the integration site in chromosome 22 was confirmed by isolating the total genomic DNA, followed by PCR analysis (FIG. 7).
[0624] Example 3: Engineering human B cell lines to continuously evolve a reporter protein
[0625] Having established plasmid-based methods for precise genome editing in B cell lines, efforts were
[0626] 35 next directed towards repurposing human B cell lines for continuous directed evolution. A previous study demonstrated that viral based integration of fluorescent proteins within the IgH locus facilitates the evolution of these fluorescent proteins with shifted spectral properties (Wang et al., Proc. Natl. Acad. Sci. 101, 16745- PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0627] 7950-112573-02
[0628] 16749). A series of variants of fluorescent proteins evolved when the corresponding gene was recombined into the genomic Ig heavy-chain locus in chromosome 14 between gene IgHV7-34-l and IgHV4-34 genes.
[0629] It was hypothesized that the DNA sequence upstream of this integration site at the immunoglobulin locus could play an important role in the previous evolution experiments performed by Tsien and coworkers.
[0630] 5 The sequences upstream of this integration site were analyzed and the presence of a DNA sequence was identified just upstream of the integration site observed by Tsien and coworkers that had a putative recognition site for the Oct transcription factor (Song et al., Blood 137, 2920-2934, 2021; Wang et al., Proc. Natl. Acad. Sci. 101, 2005-2010, 2004). The sequences from upstream of IgHV4 locus were further analyzed, and the presence of relatively conserved sequences that had the oct motifs were observed (Table 5), which are suggested to be important for transcription from immunoglobulin loci (Wang et al., Proc. Natl. Acad. Sci. 101, 2005-2010, 2004). Transcription initiation at the immunoglobulin loci has been linked to SHM (Peters et al., Immunity 4, 57-65, 1996; Komori et al., Mol. Immunol. 43, 1817-1826, 2006). Based on these observations, it was investigated if addition of one of these sequences upstream of the eGFP sequence could result in SHM at this safe harbor stable Hl l locus. Particularly, a 213 base pair fragment
[0631] 15 from IgH V4-55 (henceforth termed proXIV-1; SEQ ID NO: 1) was chosen to test this hypothesis. In addition, an eGFP variant, T66I, that lacked detectable fluorescence was used as a reporter protein (eGFP*). The proXIV-1 DNA sequence was placed upstream of the gene encoding eGFP* and the EFla promoter (plasmid DNA, pSBl). Another control DNA construct was made that had eGFP* driven by EFla promoter but did not contain proXIV (pSB2). If proXIV-1 was able to successfully recruit B cell SHM machinery to
[0632] 20 this non-immunoglobulin locus where the gene encoding eGFP* was integrated, mutations would be expected to accumulate in this gene, with some of the mutations resulting in the reversion of the non- fluorescent eGFP* to fluorescent eGFP, or variants thereof (FIGS. 2A-2C).
[0633] To convert eGFP T66 to 166, the codon was modified from ACC to ATC, involving a single nucleotide change from C to T. This intentional alteration was designed to leverage the activity of
[0634] 25 activation-induced deaminase (AID), a key enzyme in SHM that deaminates cytidine residues to uridine. This mismatch initiates error-prone base excision repair (BER), leading to mutations. A common result of this process is the transition of T to C. Therefore, if this transition mutation occurs in this experiment, it would revert eGFP* back to its original eGFP form.
[0635] To evaluate the experimental approach, RAI cells were transfected with plasmid DNAs p276 and pSBl. RAI cell lines were then recovered and selected where the DNA segment containing proXIV, EFla promoter, eGFP, and the puromycin cassette was integrated into chromosome 22, creating the stable RA1- proXIV-eGFP* cell line (RA1-SB1). For control experiments, RAI cells were transfected with plasmid DNA p274-eGFP*, generating the stable control RAI -eGFP* cell line, which lacked the proXIV sequence upstream of the eGFP* sequence (RA1-SB2). For both cell lines, the integration of the DNA segment at the
[0636] 35 chromosome 22 locus was confirmed using PCR analysis (FIG. 7). Both of the cell lines were propagated for multiple passages and at each passage, the cell lines were monitored for the appearance of fluorescence signal corresponding to reverted eGFP* variant, using the plate reader. After five passages, the appearance PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0637] 7950-112573-02 of a fluorescent signal corresponding to eGFP was observed in RAl-proXIV-eGFP* but not in RAl-eGFP* cell lines (FIG. 2D and FIG. 8). Subsequent flow cytometry analysis confirmed the presence of significant green fluorescence within a subpopulation of RAl-proXIV-eGFP* cells (FIG. 2E). The cells that possessed green fluorescence were then sorted using fluorescence-activated cell sorting (FACS), the RNA from these
[0638] 5 cells was isolated and cDNA corresponding to the eGFP* transcript was generated using reverse transcriptase-PCR (RT-PCR) methods. This DNA segment was subcloned into a pUC19 vector and the subcloned gene corresponding to eGFP* was sequenced to detect mutations. In several clones that were sequenced, the reversion of 166 to T66 was observed (FIGS. 9A-9B). This suggested that the SHM recruiting sequences upstream of eGFP* sequences were able to recruit the inherent SHM machinery to a non-
[0639] 10 immunoglobulin locus of eGFP* and introduce mutations within this gene. To understand the mutational scope of this approach, deep sequencing experiments were also performed, as described in Example 4.
[0640] Table 5. Exemplary SMH recruiting sequences PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0641] 7950-112573-02 PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0642] 7950-112573-02
[0643] Example 4: Single-molecule long read sequencing experiments to determine the mutational profiles generated by this approach
[0644] To comprehensively characterize the mutational profiles generated using this approach, sequencing experiments were performed. Particularly, single molecule long read PacBio sequencing was used, which
[0645] 5 allows for characterization of the mutational profile within a single eGFP* transcript. To prepare samples for sequencing, GFP positive cells were sorted from eGFP* evolution experiments using FACS, RNA was isolated, cDNA was generated and the cDNA samples were used for sequencing. Upon analyses of sequencing reads (FIG. 2D and FIGS. 9C-9E), it was observed that the introduced mutation in eGFP (C197T) to produce eGFP*, was reverted to its original Anorogenic sequence (T197C). This reversion accounted for 58.52% of the reads, and resulted in the change from isoleucine (I) to threonine (T) at residue 66, indicating a strong selection pressure for this mutation. Additionally, the second most prevalent mutation, T197G, converted the Ile-66 residue into Ser-66, effectively recreating the chromophore tripeptide sequence of native GFP, Ser-Tyr-Gly, accounting for 10.19% of mutational reads.
[0646] Other notable mutations included G310A (9.98%) and A596G (7.71%). The G310A substitution
[0647] 15 changed aspartic acid (D) to asparagine (N) at residue 104. The A596G substitution changed asparagine to serine at residue 199. Additional substitutions, although less frequent, also contributed to the diversity of amino acid changes. For instance, C208A changed glutamine (Q) to lysine (K) at residue 70, introducing a positive charge; C127G changed leucine (L) to valine (V) at residue 43, both of which are non-polar but with different side chain structures; and T50A changed valine (V) to glycine (G) at residue 17, potentially
[0648] 20 introducing flexibility. It is also worth noting that the T197C substitution frequently occurred both independently and in conjunction with other mutations. The co-occurring mutations that are most enriched in the population alongside T197C include G310A, A596G, C208A, and C127G. Notably, T197A is the only one of these mutations that does not co-occur with T197C, as they are mutually exclusive due to both being substitutions at the same site. This observation suggests that an initial eGFP* integrant may have reverted to
[0649] 25 T197C, but over the course of passaging, it began to accumulate additional mutations such as G310A, A596G, C208A, and C127G. The independent occurrence of T197A implies it might also restore fiuorescence, though this requires further experimental validation. The abundance of these additional mutations in the sorted sample may be attributed to their neutral effect on the function of eGFP*; these mutations may neither impede nor enhance the protein's fiuorescent properties, allowing them to persist in the population without affecting the overall selection for fiuorescence. Interestingly, in some cases, T197C was accompanied by as many as three other single point mutations [C127G, T197C, C208A, A305G]. An A196G / T197C double mutation was also detected, isoleucine was mutated to alanine, a known GFP variant (Grojohann et al., Elife 1, e00248; El Khatib et al., Sci. Rep. 6, 18459). The only other A196 mutation was A196C, which was minimally detected and generates a proline residue undescribed in GFP literature
[0650] 35 (BLAST-P analysis). C198 mutations occurred at a similar frequency as A196 mutations and was only observed to mutate to adenosine. C198A most frequently mutated in conjunction with T197C mutations, resulting in additional I66T mutations. Additionally, C198A was only observed to mutate alone, resulting in PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0651] 7950-112573-02 a silent isoleucine mutation. Beyond the 197-nucleotide mutation site, additional mutations were found throughout the GFP CDS. The most prevalent of these mutations were C127G (L43G), C208A (Q71K), G310A (D104N), and A596G (N199S). These mutations primary occurred in conjunction with other mutations, primarily with 197C / G mutations. These sequencing experiments demonstrated the remarkable
[0652] 5 scope of mutations that can be accrued within the gene encoding reporter protein using this approach. This continuous directed evolution platform pipeline was termed continuous directed evolution in human B cells (CODE-HB).
[0653] Example 5: Mutational bias during eGFP* evolution experiments using CODE-HB
[0654] In order to characterize the mutational profile and biases, the mutations observed during the eGFP* evolution experiments were analyzed. First, the sequencing data was analyzed from the naive pool of unsorted cells that were passaged for multiple passages thereby resulting in the accumulation of mutations. Based on these sequencing profiles, the mutational biases for each of the bases were characterized. A broad substitution pattern was observed (FIG. 17C) that is consistent with SHM mediated mutations (Sale et al.,
[0655] 15 Philos. Trans. R. Soc. Fond. B. Biol. Sci. 356, 21-28, 2001; Grotjohann et ah, Elife 1, e00248, 2012). It was observed that cytosine (C) bases were mutated to thymine (T), guanine (G) and adenine (A) at relatively high frequency. This can be attributed to AID mediated cytidine deamination followed by repair. Similarly, it was observed that the G bases were substituted with A, T and C bases with considerable frequency. It was also observed that A bases were substituted with G and T bases at moderate to high frequency and C bases at
[0656] 20 a lower frequency. T bases were substituted at a relatively lesser frequency with either A or G bases. A plot of the cumulative occurrences of mutations is shown in FIG. 17B. This trend line exhibited sections with increases in slope, corresponding to regions with higher number of mutations. In addition to substitution mutations, the presence of deletions and insertions was also observed. In the eGFP positive sorted samples, a broad mutational profile was similarly observed. When the corresponding amino acid changes were mapped
[0657] 25 onto the protein structure, they localized to a lateral portion of the eGFP barrel, proximal to and including the chromophore. These mutational profiles, preferences and spread were similar to those previously characterized SHM in B cells and B cell lines (Zheng et ah, J. Exp. Med. 201, 1467-1478, 2005; Unniraman et al., Science 317, 1227-1230, 2007), indicating that the mutations in CODE-HB are predominantly due to B cell SHM mechanisms.
[0658] Example 6: Mechanistic characterization to elucidate the DN sequences that can be used to recruit SHM machinery
[0659] Further studies were performed to investigate if other sequences like proXIV-1 could be identified that are capable of driving mutagenesis in CODE-HB. When the sequences from upstream of the IgHV4
[0660] 35 locus were analyzed, the presence of several sequences that were similar to proXIV-1 were observed. It was investigated if the addition of one of these sequences upstream of the eGFP* sequence could result in mutagenesis in the CODE-HB system. Particularly, a 213 base pair fragment from IgH V4-34 (henceforth PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0661] 7950-112573-02 termed proXIV-2) was chosen and proXIV-1 was replaced in pSBl with proXIV-2 (resulting plasmid was named pSB5). pSB5 was transfected along with p276 in RA 1 cells to generate RA 1-SB5 cell line which expresses non-fluorescent eGFP*. Next, culture and passage this cell line was continued and analyzed by flow cytometry analysis to determine if GFP positive signals appeared over time. Remarkably, the presence
[0662] 5 of significant GFP positive events were detected (FIGS. 18 A, 18C and 18D) in the RA 1-SB5 cell line and minimal GFP positive events (if any) in RA 1-SB2 cell lines that lack either proXIV-1 or proXIV-2 sequence. This clearly suggested that similar to the proXIV-1 sequence, the proXIV-2 sequence was also able to drive mutagenesis in CODE-HB.
[0663] To understand if other DNA fragments further upstream of proXIV-2 sequence could drive mutagenesis in CODE-HB, three distinct DNA fragments were generated from an approximately 2500 base pair sequence upstream of proXIV-2— proXIV-2-862, proXIV-2-1417 and proXIV-2-936 (FIGS. 18 A-18C); each of these fragments had a potential transcription start site and proXi V-2- 862 sequence contains proXIV- 2 sequence within it. proXIV-2 was replaced in pSB5 with either of the three DNA fragments to generate plasmids pSZl, pSZ2 and pSZ3. These plasmids were individually transfected along with p276 in RA 1 cell
[0664] 15 lines to generate RA 1-SZ1, RA 1-SZ2 and RA 1-SZ3 cell lines. In each case, culture and passage of these cell lines was continued and they were analyzed by flow cytometry analysis to determine if GFP positive signals appeared over time. The presence of GFP positive signals was observed in RA 1-SZ1 and RA 1-SZ2 cell lines and minimal GFP positive signals (if any) in RA 1-SZ3 cell lines (FIGS. 18C-18D). This suggested that both proXIV-2-862 and proXIV-2-1417 were able to drive mutagenesis in CODE-HB but proXIV-2-
[0665] 20 936 was not able to drive mutagenesis in CODE-HB. Taken together, these experiments allowed for the identification of four different DNA sequences that were able to drive mutagenesis in the CODE-HB system. Among all sequences tested, the proXIV-1 sequence from IgH V4-55 consistently yielded the highest frequency of eGFP-positive events and was therefore selected for further evolution experiments.
[0666] 25 Example 7 : Developing surface display platform in B cell lines to display antibody fragments
[0667] It was investigated if CODE-HB could be used to evolve antibody fragments targeting epitopes from emerging strains of avian influenza. To do this, a B cell surface display platform was developed that allows display of antibody fragments (e.g., fragment antigen binding regions of the antibody, Fab) on the surface of RAI cells (FIGS. 3A-3D). The Fab region of a well characterized broadly neutralizing antibody, CR9114, that was previously identified from phage display libraries (Dreyfus et ah, Science 337, 1343-1348, 2012) was selected. Following a method similar to Moffett et al., a single-chain antibody binding fragment (Fab) was engineered by fusing the light and heavy chains with a 60 amino acid glycine-serine (GS) linker (Moffet et al., Sci. Immunol. 4, eaax0644, 2019). This fusion construct was further modified for surface display by adding a membrane localization sequence at the N-terminus and an MHC I transmembrane helix at the C-
[0668] 35 terminus, facilitating its stable membrane tethering for external presentation (Wieczorek et al., Front. Immunol. 8, 292, 2017). The DNA encoding this chimeric protein sequence (Fab-surface display protein) was subcloned in a pUC19 vector where the expression of the Fab-surface display protein was driven by the PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0669] 7950-112573-02
[0670] EFla promoter. Further, a FLAG-tag was inserted between the constant region of the heavy chain and the transmembrane helix to facilitate detection of construct expression independently from antigen binding. This final resulting plasmid was called pSB3. To evaluate whether this surface display platform was functional in human cell lines, Expi293 cell lines were transiently transfected with plasmid pSB3. The expression of the Fab in the transfected Expi293 cells was evaluated by binding of the APC-conjugated Anti-FLAG antibodies to these Expi293 cells. The binding of the surface displayed Fab to the reporter hemagglutinin A, Hl, was evaluated by using the fluorescent signal of a C-terminally fused mCherry. Analysis of expression and binding was performed using flow cytometry. Using this approach, it was possible to distinctly delineate expression and binding due to non-overlapping fluorescence signals corresponding to binding (mCherry) and expression (APC) (FIG. 10A). This distinction facilitated accurate and independent evaluations of both expression levels and binding efficacy, simultaneously. Furthermore, the investigation revealed that the population of CR9114 Fab, when displayed on the cell surface, maintained its functional integrity, demonstrating effective binding across a spectrum of hemagglutinin subtypes, including H3, H5, and H7 (FIG. 10B). This outcome highlights the construct's versatility and potential in recognizing a wide range of targets / epitopes from emerging strains of avian influenza.
[0671] A further study was conducted to validate that this surface display platform was functional in B cells as well. Plasmids pSB3 and p276 were transfected into B cells, and the B cells were recovered and selected as before to generate RA-SB3 cell lines. These RA-SB3 cell lines were analyzed for expression of Fab- surface display protein using APC-conjugated anti-FLAG antibody and the binding of the surface displayed Fab to Hl was evaluated using an Hl-mCherry fusion protein (FIGS. 3B-3C). Robust expression and Hl binding was observed for these RAI mutant cell lines displaying CR9114 Fab (FIG. 3D).
[0672] Example 8: Combining the B cell surface display platform with the B cell evolution platform to continuously evolve antibody fragments:
[0673] Having successfully developed a Fab surface display platform in B cell lines, the next goal was to use CODE-HB to evolve the Fab sequences. Particularly, it was investigated if the CR9114 Fab can be evolved to be a better binder for hemagglutinin A from emerging strains of avian influenza (H5) (Eisfeld et al., Nature, 1-3, 2024); this is especially important in the context of recent zoonotic spread of H5N1 strains of avian influenza virus. The gene encoding CR9114 Fab surface display protein was inserted into a B cell integration vector containing the proXIV enhancer sequence to generate plasmid pSB4 (FIG. 4A). Plasmids pSB4 and p276 were transfect into B cells, and the B cells were recovered and selected as before to generate RA1-SB4. The control plasmid pSB3 that was identical to pSB4 but lacked proXIV was also utilized. Using this integration plasmid, a control cell line RA1-SB3 was generated. Multiple iterative rounds of cell passaging were done to determine if the Fab sequences could have increased binding towards H5 (FIG. 4B). To refine the selection process for cells exhibiting increased affinity towards H5, a critical examination was undertaken to distinguish between levels of construct expression and binding efficiency. Utilizing a control cohort transfected with a vector devoid of the proXIV sequence facilitated the establishment of a normative PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0674] 7950-112573-02 regression line for binding relative to expression. This baseline enabled the definition of a "High Binding" gate for events surpassing expected binding levels based on expression alone (FIG. 3D). Cells that exceeded this benchmark within the proXIV-inclusive sample, yet not present in the control, were posited as having undergone favorable somatic hypermutations within their recombinant CR9114 construct sequences.
[0675] 5 Upon utilizing the "High Binding" gate for selection, an enrichment of approximately 40,000 cells was achieved and cultivated to saturation. Subsequent iterations of FACS sorting and analysis of this population highlighted an intriguing divergence from the control group's fluorescence profile, characterized by a decreasing mean fluorescence intensity in the expression channel, contrasted with a consistent intensity in the binding channel (FIG. 11). This unexpected trend prompted a deeper molecular investigation, leading to the isolation of RNA from the sorted population and Sanger sequencing of the CR9114 amplicon. The sequencing results revealed a prevalent mutation, specifically the conversion of Y532 to F532 (TAC to TTC), adjacent to the FLAG-tag region, suggesting that a targeted selection pressure in this area resulted in a phenotype that could be detected (FIG. 11). These observations clearly demonstrated that a phenotype can be linked to the genotype by using this approach.
[0676] 15 In addition to this population, a population of high binders as compared to the control group was observed (FIG. 4D). The enrichment process showed a marked improvement in target population isolation across successive sorting steps. Initially, the percentage of events that escaped the gating was 2.97%. This increased substantially to 94.48% after the second sort, and further to 97.07% following the third sort, demonstrating progressive refinement of the target cell population. Correspondingly, the histogram data
[0677] 20 illustrated a significant rise in the mean fluorescence intensity (F.I.) of the H5-binding channel (FITC+), with values of 361.8, 716.87, and 1155.62 in the initial, second, and third sorts, respectively (FIGS. 4E-4F). For this verification experiment, His-tagged H5 was used and FITC-conjugated anti-His antibody was used to detect binding of cells to His-tagged H5 (FIG. 4E-4F). Post-sorting, the cells were expanded, and their RNA was isolated for subsequent conversion to cDNA.
[0678] 25
[0679] Example 9: Single-molecule long read sequencing experiments to determine the mutational breadth and scope of mutational profiles generated by this approach
[0680] Cells corresponding to these high binders were sorted, the RNA was isolated, the cDNA of transcript corresponding to the Fab-surface display protein was generated and then long read single molecule sequencing was performed to determine the sequences of variant Fab (FIG. 4A). Sequencing generated a total of 642,858 reads, of which 471,562 met the expression threshold based on published PacBio error rates (15 observations) (Lang et al., Gigascience 9, giaal23, 2020).
[0681] Focusing on heavy chain variants, owing to the heavy chain-only binding mode of CR9114 to HA as indicated by crystal structure (PDB: 4CQI) (Barrack et al., Acta Crystallogr. Sect. F Struct. Biol. Commun.
[0682] 35 71, 539-546, 2015), the unique reads of the top 200 variants were shortlisted. This refinement process yielded a final set of 16 variants, encompassing both variable and constant regions of the heavy chain. Notably, these variants represented only 0.25% of the total reads. The analysis revealed a variety of PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0683] 7950-112573-02 nucleotide substitutions, including both transitions and transversions (FIGS. 5C, 5F). Approximately 22% of unique reads read sequences were discarded, as they generated insertion or deletion mutations when aligned to the reference sequence. Of these remaining sequences, 82% of reads corresponded to a single mutational profile: Al 149G and A1601T. The first mutation corresponds to a silent mutation for codon optimization of
[0684] 5 a glutamic acid residue. The second mutation corresponds to a tyrosine to phenylalanine mutation within the FLAG tag sequence (tAc:tTC Y:F, FIGS. 12 and 13). This mutation results in the change of a tyrosine residue within the FLAG tag to phenylalanine which potentially reduces antibody recognition. The subsequent multiple sequence alignment included 208 unique sequences. In addition to these mutations, a series of stacked mutations on top of the A1601T mutant were observed. For example, the A1149G / A1601T mutant, A1601T / C107A / A1149G mutant, and C107A / G119T / A1149G / A1601T mutant were observed. Beyond these variants, a variety of other mutations were present across the Fab sequence with varying frequency. A heat map of CR9114 mutational profile after multiple rounds of diversification, selection and enrichment (3 rounds of sorting and enrichment) is shown in FIG. 5F. There is a significant diversification of CR9114 sequences after multiple rounds of evolution and enrichment.
[0685] 15 Following the nucleotide-level analysis, remaining reads were translated, and protein level mutations were analyzed. Many amino acids were not mutated (FIGS. 12 and 13). Across the Fab sequence, two mutational hot-spots appeared. The first was within the L(CR9114) region specifically between amino acid positions T128-L144. Within this region, there is a marked increase in mutational frequency, with the exception of P142, which was never mutated. Mutations within this region are primarily to another single
[0686] 20 amino acid, with few residues showing mutation to two different amino acids. The second mutational hostspot is the H(CR9114) Fab region. In contrast to the L(CR9114) mutational hot-spot, within this region specific residues were marked mutated: G343R, M347V, I373F, F394L, and F394S. Beyond these mutants, a myriad of other mutations were present across the Fab sequence with varying frequency.
[0687] 25 Example 10: Validation of Fab sequences with enhanced binding
[0688] Next, Fab sequences that demonstrate enhanced binding were identified and validated. A subset of variants were selected from the sequencing data based on the mutational frequency and in some cases based on the combination of mutational frequency and the location of variation (e.g., in the CDR loops). Functional characterization of these variants (FIGS. 5, 14 and 15) involved determining their EC50 values through hemagglutinin H5 titration. Overall, the variants exhibited lower EC50 values compared to the CR9114 wild-type control, reflecting enhanced binding affinities. Among these, two variants, VH S120P and VH W154R, demonstrated the lowest EC50 values of around 4.2 nM (FIG. 4C-4D). These values were approximately three times lower than that of the CR9114 wild-type control, highlighting the superior binding affinity of these selected variants. Interestingly, the two variants with the lowest EC50 values, VH
[0689] 35 S120P and VH W154R, both contained mutations in the heavy chain constant region. This finding is particularly noteworthy as it suggests that alterations in the constant region, typically associated with structural and stability roles rather than direct antigen binding, can significantly influence the binding PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0690] 7950-112573-02 affinity of the antibody. These results underscore the importance of the constant region in modulating antibody function and highlight the potential for targeted mutations in this region to improve therapeutic antibody efficacy. The mutant VH W 154R was tested for retention of binding towards other hemagglutinin variants (Hl, H3, H5) and was found to not only retain but also exhibit more favorable EC50 values than the
[0691] 5 wild-type CR9114.
[0692] In addition to this, the presence of variants was also detected in the heavy variable region that demonstrated enhanced binding as compared to the controls (FIGS. 5C-5E). The I373F variant of the CR9114 antibody is a single point mutation within a critical complementarity-determining region (CDR) loop on the heavy chain, which interfaces directly with the hemagglutinin H5 antigen (crystal structure: 4CQI). The VH M48V variant of the CR9114 antibody is another point mutation within the beta sheet of the heavy chain variable region, a critical area for maintaining the antibody's structural integrity. The VH Q65R variant of the CR9114 antibody involved a point mutation within a flexible loop of the heavy chain variable region, situated near a hydrogen bond network in CR9114 (PDB: 4CQI). The original isoleucine at position 373 may nestle into a hydrophobic pocket on the H5 surface, stabilizing the antibody-antigen interaction
[0693] 15 through hydrophobic forces. This interaction may be crucial as it anchors the antibody to the antigen for effective neutralization of the virus. The substitution of isoleucine with phenylalanine at this position introduces a larger aromatic side chain, which may enhance stability via pi-stacking interactions with the adjacent phenylalanine at position 374 and possibly contribute to the structural stability of the CDR loop. Furthermore, the bulkier phenylalanine side chain may more effectively fill the hydrophobic pocket on the
[0694] 20 antigen surface, thereby increasing the hydrophobic interactions that are essential for strong binding affinity. This mutation may not only augment the stability of the CR9114-H5 complex but also improve the specificity of the binding; by occupying the hydrophobic pocket more completely, phenylalanine at position 373 could reduce the likelihood of off-target interactions, enhancing the selectivity of CR9114 for the H5 antigen (FIG. 16). The M347V variant of the CR9114 antibody is a single point mutation within the beta
[0695] 25 sheet of the heavy chain variable region, a critical area for maintaining the antibody's structural integrity. This region, situated between a beta strand and an alpha helix, displays a degree of torsion established by this potential axis of interaction that is also flanked on all sides by aromatic residues. The methionine at position 347 may introduce a bulky side chain, acting as a mediator between an inner beta strand tyrosine and an outer small, alpha-helical phenylalanine, potentially affecting the overall efficacy of the CR9114 variant (FIG. 16). The Q412R variant of the CR9114 antibody involves a single point mutation within a flexible loop of the heavy chain variable region, situated near a hydrogen bond network comprising residues Q6, T107, S7, and S21. According to the crystal structure (PDB: 4CQI), the amide nitrogen of Q412 is approximately 5.5 angstroms from the oxygen of T107, though the flexibility of the loop may allow it to migrate closer. Q6 is positioned 3 angstroms from T107, forming part of this network. With the substitution
[0696] 35 of glutamine to arginine at position 412, an extended carbon chain and an additional interacting nitrogen are introduced, which may enhance the stability of the hydrogen bond network. This stabilization could contribute to the upstream integrity of the CDR loops, essential for antigen binding. Consequently, it is PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0697] 7950-112573-02 speculated that the Q41 R mutation might improve the structural stability and functional efficacy of the CR9114 antibody by reinforcing the local hydrogen bond network (FIG. 16). The results herein underscore the potent application of directed somatic hypermutation for the refinement of antibody engineering, highlighting a methodology to uncover and harness these molecular insights for antibody evolution.
[0698] 5
[0699] Example 11: Testing the evolved variant of CR114 using viral neutralization assays
[0700] To determine if the evolved CR9114 variants still retained viral neutralization potential, CR9114 variants were recombinantly produced. ELISA assays were performed to verify binding to H5. Next, to determine if the antibody variants are able to neutralize the influenza virus, viral neutralization assays were performed. To quantify the capacity of CR9114 variants to block receptor engagement, hemagglutinationinhibition (HAI) assays were performed using turkey erythrocytes and influenza A / Puerto Rico / 1934 (H1N1) (FIG. 19A). Wild-type CR9114 exhibited an HALo of 50 pg / mL, consistent with its broadly neutralizing profile. Each of the heavy-chain variable -region mutants — designed to probe fine adjustments in the antigen-contact interface — yielded HALo values statistically indistinguishable from the parental antibody
[0701] 15 (mean HALo ~ 50-100 pg / mL; n = 3). In stark contrast, the constant-region substitution W154R conferred a significant four-fold enhancement in inhibitory potency: the W154R variant achieved 50% inhibition of hemagglutination at just 12.5 pg / mL (FIGS. 19B-19C). This improvement implies that modifications outside of the antigen-binding site can influence the functional avidity the ability of CR9114.
[0702] To determine whether the W461R variant’s increased HAI activity translates to stronger blockade of
[0703] 20 viral entry, viral neutralization in a cell-based infection assay was next assessed. Influenza A / Puerto Rico / 1934 virions were pre-incubated with serial dilutions of each mAb before inoculation onto MDCK monolayers. Wild-type CR9114 required concentrations above 100 nM to achieve >90% reduction in cytopathic effect, whereas the W461R variant maintained equivalent neutralizing efficacy at 10 nM — well within physiologically attainable antibody titers (FIGS. 19D-19E).
[0704] 25 Together, these data reveal that mutations in the Fc-proximal domain can markedly improve functional potency of CR9114 against H1N1. These experiments demonstrate the utility of CODE-HB for rapid evolution and engineering of high-affinity antibodies.
[0705] Example 12: Leveraging broad mutational profile of CODE-HB for evolution experiments
[0706] To showcase how the broad mutational profile of CODE-HB can be used for antibody engineering, another example of antibody evolution campaign is presented. This evolution campaign starts off with previously identified 1A02 antibody that shows high binding to influenza Hl / California / 2009 but lower binding to influenza Hl / Michigan / 2015 Hl hemagglutinins (Guthmiller et al., Immunity 53, 1230-1244.e5, 2020). To enhance the binding 1 A02 to influenza Hl / Michigan / 2015, the CODE-HB system was employed
[0707] 35 (FIG. 20A). Particularly, 1A02 Fab was displayed on the surface of the B cell lines and evolution experiments were performed to identify potent binders of Hl / Michigan / 2015 (FIG. 20B). After sequential rounds of evolution and enrichment, the 1A02 Fab transcripts were sequenced to determine key mutations PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0708] 7950-112573-02 that result in enhanced binding to Hl / Michigan / 2015. A wide range of substitution mutations was observed (FIGS. 20C-20D). In addition to this, a seven-amino-acid deletion within the sc60 linker was also observed. Additionally, several substitution mutations on top of the 7 amino acid deletion were seen (FIG. 20D), including mutations in the region corresponding to CDR loops. By constructing and assaying variants
[0709] 5 bearing each alteration alone or in combination, excising those seven residues from the 1 A02 Fabplays a crucial role in improving the binding of the 1 A02 Fab to Hl / Michigan / 2015 (FIG. 20E). The presence of these stacked variants that include deletions and substitutions in a single construct clearly demonstrates the unique ability of CODE-HB to evolve variants with a broad mutational profile comprising of substitutions, deletions and insertions; such mutational profiles typically inaccessible using other continuous directed evolution approaches.
[0710] Discussion
[0711] Directed evolution platforms have significantly advanced the ability to rapidly evolve biomolecules with defined functions. Advances in molecular biology and synthetic biology further resulted in the
[0712] 15 development of continuous directed evolution platforms where cell or virus-based systems are used. One advantage of these platforms is the ability of engineered cells to introduce biomolecular diversity in situ without requiring in vitro library generation. Another attractive feature of some of these systems is the linkage of evolutionary outcomes to the propagation of cells or viruses with selective growth advantage. Several other efforts are ongoing to build, optimize and utilize continuous directed evolution platforms in
[0713] 20 model bacteria, yeasts, mammalian cells and viruses. Several studies have extensively characterized the biochemistry and intrinsic mutation rates of SHM, and the timing of SHM in relation to phenotype evolution in B cells and B cell lines, including RA 1 cell lines (Neuberger et al., Exp. Med. 204, 7-10, 2007; Sale et al., Immunity 9, 859-869, 1998; Senigl et al., Cell Rep. 29, 3902-3915, 2019).
[0714] Despite a few previous efforts to test evolution experiments in B cells, none of these approaches
[0715] 25 have found any widespread utility. To meet this unmet need, described herein is a method for virus free, precise and efficient genome editing in human B cell lines. This is in direct contrast to all previous studies that use low efficiency viral transduction with random genomic integration. Stable high levels of protein expression in more than 90% of the cell population was achieved. A new B cell surface display platform for selection is also described. This does not limit the type of antibodies or antibody fragments that can be displayed on the surface of the B cell lines. Most importantly, stable, continuous diversification and evolution of genes from a stable, safe harbor locus in more than 90% of cells was achieved. All previous approaches evolve genes from a very low fraction of cells. Further in these approaches, evolution only occurs on an even smaller fraction of cells which have random integration event at the immunoglobulin locus. Additionally, these previous approaches are further hindered by the fact that the immunoglobulin loci
[0716] 35 can undergo significant gene deletions. Further, in this platform to continuously evolve variants, a strategy that recruits and repurposes the inherent B cell SHM mechanisms, to rapidly and orthogonally evolve PCT / US25 / 48620 30 September 2025 (30.09.2025)
[0717] 7950-112573-02 proteins of interest from a stable, safe harbor locus within human B cells was developed. This platform can be used to rapidly evolve eGFP variants.
[0718] In addition, a B cell surface display platform for displaying fragment antigen binding domains, Fab, of antibodies was also developed. Combining the Fab surface display platform with CODE-HB, Fab
[0719] 5 sequences targeting avian subtypes of influenza hemagglutinin were evolved (e.g., H5). It was also demonstrated that this approach can be used to rapidly and continuously evolve antibody sequences in human B cell lines. Particularly, this evolution experiment with matured antibody sequence CR9114 resulted in isolation of CR9114 variants that were able to bind H5 with higher avidity. Importantly, several of these evolved antibody variants retained their viral neutralization potential at similar or higher level to the starting
[0720] 10 CR9114. These evolution experiments with 1 A02 antibody resulted in the evolution of a Fab variant that had significantly more potent binding than the starting 1 A02 Fab sequence. Additionally, unlike other continuous directed evolution approaches in bacteria, yeast and viruses, this approach makes it possible to attain a broad mutational spectrum, including substitutions, deletions and insertions, and at least in one of the evolutionary campaigns this mutational profile plays an important role in evolving a desired phenotype. This clearly highlights the utility of CODE-HB for continuous directed evolution in human cells.
[0721] It will be apparent that the precise details of the methods or compositions described may be varied or modified without departing from the spirit of the described aspects of the disclosure. We claim all such modifications and variations that fall within the scope and spirit of the claims below.
[0722] 20
Claims
7950-112573-02CLAIMS1. An integration plasmid capable of integrating at a specific genomic locus, comprising: a heterologous gene operably linked to a first promoter; a somatic hypermutation (SHM) enhancer element upstream of the first promoter and the heterologous gene; a first nucleic acid sequence homologous to genomic sequences at the integration site, wherein the first nucleic acid sequence is upstream of the SHM enhancer element; and a second nucleic acid sequence homologous to genomic sequences at the integration site, wherein the second nucleic acid sequence is downstream of the heterologous gene and a polyadenylation (poly-A) signal sequence.
2. The integration plasmid of claim 1, wherein the genomic locus is a stable, nonimmunoglobulin locus.
3. The integration plasmid of claim 1, wherein the genomic locus is the Hl 1 locus of human chromosome 22 or the AAVS1 locus of human chromosome 19.
4. The integration plasmid of claim 1, wherein the heterologous gene encodes an antibody fragment.
5. The integration plasmid of claim 4, wherein the heterologous gene further encodes a membrane localization signal, a transmembrane helix, and a linker between the antibody fragment and the transmembrane helix.
6. The integration plasmid of claim 4, wherein the antibody fragment is a single-chain antibody binding fragment (Fab).
7. The integration plasmid of claim 4, wherein the antibody fragment binds a viral protein.
8. The integration plasmid of claim 7, wherein the viral protein is an influenza virus protein.
9. The integration plasmid of claim 8, wherein the influenza virus protein is hemagglutinin(HA).
10. The integration plasmid of claim 1, wherein: the nucleotide sequence of the heterologous gene is at least 85% identical to SEQ ID NO: 4;7950-112573-02 the nucleotide sequence of the SHM enhancer element is at least 85% identical to any one of SEQ ID NOs: 1 and 59-68; and / or the nucleotide sequence of the integration plasmid is at least 95% identical to SEQ ID NO: 2.
11. The integration plasmid of claim 1, wherein: the nucleotide sequence of the heterologous gene comprises or consists of SEQ ID NO: 4; the nucleotide sequence of the SHM enhancer element comprises or consists of any one of SEQ ID NOs: 1 and 59-68; and / or the nucleotide sequence of the integration plasmid comprises or consists of SEQ ID NO: 2.
12. The integration plasmid of claim 1, wherein the first promoter is the human EFl a promoter.
13. The integration plasmid of claim 1, wherein: the first nucleic acid sequence homologous to genomic sequences at the integration site is about 200 to about 800 nucleotides in length; and / or the second nucleic acid sequence homologous to genomic sequences at the integration site is about 200 to about 800 nucleotides in length.
14. The integration plasmid of claim 1, further comprising an antibiotic resistance cassette.
15. A kit comprising the integration plasmid of claim 1, and one or more of: a CRISPR / Cas9 plasmid encoding Cas9 and one or more guide RNAs that target the genomic locus; isolated B cells;B cell culture media; tissue culture flask(s); and buffer.
16. The kit of claim 15, wherein the CRISPR / Cas9 plasmid is the p276 plasmid having the nucleotide sequence of SEQ ID NO: 3.
17. The kit of claim 15, wherein the isolated B cells are primary B cells or the isolated B cells are cells of a B cell line.
18. An isolated B cell comprising the integration plasmid of claim 1 integrated into the genome of the cell.
19. The isolated B cell of claim 18, which is a primary B cell or a B cell line.7950-112573-0220. The isolated B cell of claim 18, which is a human B cell, an avian B cell, or a swine B cell.
21. The isolated B cell of claim 18, wherein the B cell expresses on its surface an antibody fragment encoded by the integration plasmid.
22. A method of generating an antibody or antibody fragment that has higher affinity for a target antigen than a parental antibody or antibody fragment that binds the same antigen, comprising: transfecting isolated B cells with the integration plasmid of claim 1 and a CRISPR / Cas9 plasmid, wherein the CRISPR / Cas9 plasmid encodes Cas9 and one or more guide RNAs that target the genomic locus of the integration plasmid; enriching the B cells through antibiotic selection; and passaging the transfected B cells for multiple passages.
23. The method of claim 22, further comprising measuring affinity of the passaged B cells to the antigen.
24. The method of claim 22, wherein the antibody fragment is a single-chain antibody binding fragment (Fab).
25. The method of claim 22, wherein the CRISPR / Cas9 plasmid is the p276 plasmid having the nucleotide sequence of SEQ ID NO: 3.
26. The method of claim 22, wherein the isolated B cells are primary B cells or the isolated B cells are cells of a B cell line.
27. The method of claim 22, wherein the B cells are human B cells, avian B cells, or swine B cells.
28. A method of displaying a single-chain antibody binding fragment (Fab) on the surface of B cells, comprising transfecting B cells with a nucleic acid molecule encoding in the 5’ to 3’ direction: a membrane localization sequence; the Fab, wherein the Fab comprises a light chain and a heavy chain separated by a linker; and a transmembrane helix.
29. The method of claim 28, wherein the amino acid sequence of the membrane localization sequence comprises SEQ ID NO: 12.7950-112573-0230. The method of claim 28, wherein the amino acid sequence of the Fab light chain comprises SEQ ID NO: 13 and / or the amino acid sequence of the Fab heavy chain comprises SEQ ID NO:
15.
31. The method of claim 28, wherein the transmembrane helix is a major histocompatibility complex I (MHCI) transmembrane helix.
32. The method of claim 28, wherein the amino acid sequence of the transmembrane helix comprises SEQ ID NO: 17.
33. The method of claim 28, wherein the nucleotide sequence of the nucleic acid molecule is at least 95% identical to SEQ ID NO: 4.
Citation Information
Patent Citations
In vivo affinity maturation scheme
US20060099611A1
Method to bioengineer designer red blood cells using gene editing and stem cell methodologies
US20200332259A1
Non-human animals having a limited lambda light chain repertoire expressed from the kappa locus and uses thereof
US20220330532A1
Anti-PD-l1 / Anti-b7-h3 multispecific antibodies and uses thereof
US20230192861A1