Transgenic rodents for cell line identification and enrichment
Patent Information
- Application Number
- JP2024520084
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-01
- Filing Date
- 2022-09-30
- Publication Date
- 2025-10-07
AI Technical Summary
Current methods for identifying and enriching cells expressing specific proteins or being at specific developmental stages, such as Ig-expressing cells, are inefficient due to limitations in cell surface marker availability and specificity, leading to low yield, purity, and contamination during enrichment.
The use of nucleic acid constructs containing a leader sequence, LoxP-Stop-LoxP cassette, affinity tag, transmembrane domain, and fluorescent reporter protein, integrated into safe harbor sites in the genome of transgenic rodents, allowing for precise identification and enrichment of cells through fluorescence-activated cell sorting (FACS) or magnetic activated cell sorting (MACS).
This approach enables high-purity isolation of target cells with minimal contamination, facilitating the production of therapeutic or diagnostic antibodies by ensuring accurate identification and sorting of cells expressing specific proteins.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to nucleic acid constructs, transgenic rodents, rodent cell lines, and methods that allow for the identification and enrichment of particular cell types, e.g., cells at a particular developmental stage, cells expressing a particular promoter, or cells expressing a particular protein, such as an antibody. [Background technology]
[0002] Identifying and enriching cells that have been modified to express specific proteins or be at a specific developmental stage is a key challenge in the development of biotherapeutics. A common workflow for enriching specific cell populations is to create a single-cell suspension, stain the cell mixture with a panel of antibodies that recognize surface markers, and separate the cells using either magnetic or flow-based methods. However, this procedure is limited by our current knowledge of cell type-specific cell surface markers, and the specificity and availability of antibodies that recognize those markers. This procedure generally results in less than ideal yields and purity of the desired cells after enrichment, a high percentage of unwanted contaminating cells, and loss of the desired cells during enrichment. For example, a common strategy to identify Ig-expressing cells is based on combining known endogenous lineage surface markers with antibody staining to detect those markers.
[0003] Antibodies commonly used to enrich for mouse Ig-expressing cells are anti-CD19, anti-CD138, and anti-Ig. However, there are differences in the expression of these three markers during B cell differentiation, and not all populations can be efficiently enriched using cell surface markers. For example, CD19 is considered a pan-B cell marker (including B cell precursors that do not express Ig), but its expression is dramatically reduced in antibody-secreting cells, so it cannot enrich for that useful population. CD138 is considered a plasma cell marker, but it is also expressed by some early stage precursor B cells that do not express Ig. Therefore, this marker will enrich for this undesirable population. During B cell development, pre-B cells begin to exhibit Ig on the cell surface after differentiation into immature B cells, so Ig markers can be used to capture this population. However, when mature B cells fully differentiate into plasma cells, the surface expression of Ig is lost. As a result, when these markers are used in magnetic-based strategies (which offer better scale and time efficiency compared to flow-based sorting) to enrich for Ig-expressing cells, the resulting enriched cell populations often contain contamination from non-Ig-expressing B cells, resulting in inefficient enrichment and loss of antibody-secreting cells.
[0004] Isolation and enrichment of cell lines expressing tissue-specific promoters is also challenging for similar reasons: tissue specificity is determined primarily by transcription factors, and cell surface markers may not be available to enrich for cell lines expressing proteins in a tissue-specific manner, or available markers may not be specific enough to provide useful enrichment. Summary of the Invention
[0005] In embodiments, the present disclosure provides a nucleic acid construct comprising a leader sequence, a LoxP-Stop-LoxP cassette, and a transmembrane reporter cassette encoding an affinity tag, a transmembrane (TM) domain, and a fluorescent reporter protein. In embodiments, the nucleic acid construct comprises single-stranded DNA, double-stranded DNA, a plasmid, or a viral vector.
[0006] In an embodiment, the nucleic acid construct further comprises a first homology arm and a second homology arm that are homologous to a first target sequence and a second target sequence, respectively, within the safe harbor site in the non-human mammal. In an embodiment, the first homology arm and the second homology arm each independently comprise from about 15 nucleotides to about 12000 nucleotides.
[0007] In embodiments of the nucleic acid construct, the safe harbor site includes the Rosa26 site on chromosome 6 of the mouse genome or the Hipp11 site on chromosome 11 of the mouse genome.
[0008] In embodiments, the nucleic acid construct further comprises a promoter. In embodiments, the promoter comprises a mammalian promoter. In embodiments, the promoter comprises a CAG, CMV, EF1a, SV40, PGK1, Ubc, or human beta-actin promoter. In embodiments, the leader sequence comprises a secretory signal peptide. In embodiments, the secretory signal peptide comprises the IL-2 leader sequence MYRMQLLSCIALSLALVTNS (SEQ ID NO:2).
[0009] In an embodiment of the nucleic acid construct, the affinity tag comprises a StrepII tag. In an embodiment, the affinity tag comprises tandem repeats of a StrepII tag. In an embodiment, the affinity tag comprises about 1 to about 18 tandem repeats of a StrepII tag with a tag linker between the repeats. In an embodiment, the affinity tag comprises 3 tandem repeats of a StrepII tag. In an embodiment, the StrepII tag comprises a peptide sequence of 8 amino acids of WSHPQFEK (SEQ ID NO: 1). In an embodiment, the transmembrane domain comprises a hydrophobic alpha helix.
[0010] In embodiments of the nucleic acid construct, the fluorescent reporter protein comprises green fluorescent protein (GFP), enhanced green fluorescent protein (EGFP), enhanced yellow fluorescent protein (EYFP), or enhanced cyan fluorescent protein (ECFP).
[0011] In embodiments, the disclosure provides a method of generating a genetically modified non-human mammalian cell, the method comprising: (a) introducing a nucleic acid construct described herein into a non-human mammalian cell; and (b) introducing a nuclease into the non-human mammalian cell, where the nuclease causes a single-stranded or double-stranded break at a safe harbor site in the genome of the non-human mammalian cell, and where the nucleic acid construct is integrated into the genome of the non-human mammalian cell at the safe harbor site by homologous recombination.
[0012] In an embodiment of the method, introducing the nuclease comprises introducing an expression construct encoding the nuclease. In an embodiment, introducing the nuclease comprises introducing an mRNA encoding the nuclease. In an embodiment, the nuclease comprises a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, or a clustered regularly interspaced short palindromic repeats (CRISPR) associated (Cas) protein and a guide RNA (gRNA). In an embodiment, the gRNA comprises a CRISPR RNA (crRNA) and a transactivating CRISPR RNA (tracrRNA) that targets the recognition site. In an embodiment, the CRISPR-Cas protein comprises Cas9.
[0013] In an embodiment of the method, the non-human mammalian cell is a rodent cell. In an embodiment, the rodent cell is a rat cell or a mouse cell. In an embodiment, the safe harbor site comprises the Rosa26 site on chromosome 6 or the Hipp11 site on chromosome 11 of the mouse genome. In an embodiment, the non-human mammalian cell is a pluripotent cell. In an embodiment, the pluripotent cell is a non-human zygote or a non-human embryonic stem (ES) cell. In an embodiment, the pluripotent cell is a mouse zygote cell or a rat zygote cell. In an embodiment, the pluripotent cell is a mouse embryonic stem (ES) cell or a rat embryonic stem (ES) cell.
[0014] In embodiments, the method further comprises isolating the genetically modified non-human mammalian cell into which the nucleic acid construct has been integrated at the safe harbor site.
[0015] In embodiments, the present disclosure provides genetically modified non-human mammalian cells produced by the methods of producing genetically modified non-human mammalian cells described herein.
[0016] In an embodiment of the method, the method further comprises injecting the isolated cell into a blastocyst to generate a transgenic non-human mammal comprising the nucleic acid construct integrated into the safe harbor site. In an embodiment, the disclosure provides a genetically modified non-human transgenic mammal generated by the method. In an embodiment, the mammal is a rodent. In an embodiment, the rodent is a rat or a mouse.
[0017] In an embodiment, the method further comprises mating the transgenic non-human mammal comprising the nucleic acid construct integrated at the safe harbor site with a transgenic non-human mammal expressing a Cre recombinase to obtain a non-human mammal having cells expressing a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein. In an embodiment, the transgenic non-human mammal comprising the nucleic acid construct integrated at the safe harbor site is a mouse comprising a nucleic acid construct integrated at the Rosa26 site, and the transgenic non-human mammal expressing the Cre recombinase is a mouse. In an embodiment, the transgenic non-human mammal comprising the nucleic acid construct integrated at the safe harbor site is a mouse comprising a nucleic acid construct integrated at the Hipp11 site, and the transgenic non-human mammal expressing the Cre recombinase is a mouse. In an embodiment, the Cre expression in the transgenic mouse is tissue specific. In an embodiment, the present disclosure provides a genetically modified non-human mammal having cells expressing a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein generated by the method.
[0018] In an embodiment, the present disclosure provides a genetically modified non-human mammalian cell comprising a genome comprising a nucleic acid construct described herein integrated into a safe harbor site. In an embodiment, the safe harbor site comprises the Rosa26 site on chromosome 26 of the mouse genome or the Hipp11 site on chromosome 11 of the mouse genome. In an embodiment, the genetically modified non-human mammalian cell is a hybridoma or an immortalized cell.
[0019] In an embodiment of the genetically modified non-human mammalian cell, the cell expresses a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein. In an embodiment, the affinity tag is expressed on the cell surface of the non-human mammalian cell. In an embodiment, the affinity tag comprises a StrepII tag. In an embodiment, the fluorescent reporter protein is exposed on the cytoplasmic surface of the non-human mammalian cell. In an embodiment, the fluorescent reporter protein comprises green fluorescent protein (GFP), enhanced green fluorescent protein (EGFP), enhanced yellow fluorescent protein (EYFP), or enhanced cyan fluorescent protein (ECFP).
[0020] In embodiments, the disclosure provides a method for isolating cells obtained from a genetically modified non-human mammal, the method comprising: (a) obtaining cells from a genetically modified non-human mammal as described herein; (b) screening the cells obtained from the genetically modified non-human mammal for expression of a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein; and (c) isolating cells expressing the fusion protein.
[0021] In an embodiment of the method for isolating cells, the cells are sorted by fluorescence activated cell sorting (FACS) or magnetic activated cell sorting (MACS). In an embodiment, the affinity tag is expressed on the cell surface of the genetically modified non-human mammalian cell. In an embodiment, the affinity tag comprises a StrepII tag. In an embodiment, the fluorescent reporter protein is exposed on the cytoplasmic surface of the non-human mammalian cell. In an embodiment, the fluorescent reporter protein comprises green fluorescent protein (GFP), enhanced green fluorescent protein (EGFP), enhanced yellow fluorescent protein (EYFP), or enhanced cyan fluorescent protein (ECFP).
[0022] In embodiments, the disclosure further provides a nucleic acid construct comprising a linker, a leader sequence, and a transmembrane reporter cassette encoding an affinity tag, a transmembrane domain, and a fluorescent reporter.
[0023] In embodiments, the nucleic acid construct comprises a single-stranded DNA, a double-stranded DNA, a plasmid, or a viral vector. In embodiments, the nucleic acid construct further comprises a first homology arm and a second homology arm that are homologous to a first target sequence and a second target sequence, respectively. In embodiments, the first target sequence is upstream of the immunoglobulin constant region site, and the second target sequence is downstream of a stop codon of the immunoglobulin constant region site. In embodiments, the immunoglobulin constant region site is an immunoglobulin light chain constant region site. In embodiments, the immunoglobulin light chain constant region site is an immunoglobulin kappa constant region site. In embodiments, the immunoglobulin light chain constant region site is an immunoglobulin lambda constant region site. In embodiments, the immunoglobulin constant region site is an immunoglobulin heavy chain constant region site. In embodiments, the immunoglobulin heavy chain constant region site is a gamma, delta, alpha, mu, or epsilon immunoglobulin heavy chain constant region site.
[0024] In an embodiment of the nucleic acid construct, the first homology arm and the second homology arm each independently comprise from about 15 nucleotides to about 12000 nucleotides. In an embodiment, the linker comprises a stop codon and an internal ribosome entry site (IRES). In an embodiment, the linker comprises a protease recognition site and a self-cleaving peptide. In an embodiment, the linker comprises a leaky stop codon (LSC) with a peptide linker, a protease recognition site, and a self-cleaving peptide. In an embodiment, the protease recognition site comprises a Furin protease recognition site. In an embodiment, the Furin protease recognition site comprises a nucleic acid sequence encoding a peptide of Arg-X-Arg-Arg. In an embodiment, X is a hydrophobic amino acid. In an embodiment, X is a hydrophilic amino acid. In an embodiment, X is lysine. In embodiments, the Furin protease recognition site comprises a nucleic acid sequence encoding a peptide of X-Arg-X-Lys-Arg-X or X-Arg-X-Arg-Arg-X. In embodiments, X is a hydrophobic amino acid. In embodiments, the hydrophobic amino acid is Gly, Ala, Ile, Leu, Met, Val, Phe, Trp, or Tyr. In embodiments, X is a hydrophilic amino acid. In embodiments, the hydrophilic amino acid is lysine. In embodiments, the self-cleaving peptide comprises a 2A self-cleaving peptide. In embodiments, the leaky stop codon comprises TGACTAG. In embodiments, the dipeptide linker comprises Leu-Gly.
[0025] In an embodiment of the nucleic acid construct, the leader sequence comprises a secretory signal peptide. In an embodiment, the secretory signal peptide comprises the IL-2 leader sequence MYRMQLLSCIALSLALVTNS (SEQ ID NO:2).
[0026] In an embodiment of the nucleic acid construct, the affinity tag comprises a StrepII tag. In an embodiment, the affinity tag comprises tandem repeats of a StrepII tag. In an embodiment, the affinity tag comprises about 1 to about 18 tandem repeats of a StrepII tag with a tag linker between the repeats. In an embodiment, the affinity tag comprises 3 tandem repeats of a StrepII tag. In an embodiment, the StrepII tag comprises a peptide sequence of 8 amino acids of WSHPQFEK (SEQ ID NO: 1). In an embodiment, the transmembrane domain comprises a hydrophobic alpha helix.
[0027] In embodiments of the nucleic acid construct, the fluorescent reporter protein comprises green fluorescent protein (GFP), enhanced green fluorescent protein (EGFP), enhanced yellow fluorescent protein (EYFP), or enhanced cyan fluorescent protein (ECFP).
[0028] In embodiments, the disclosure provides a method of generating a genetically modified non-human mammalian cell, the method comprising: (a) introducing a nucleic acid construct described herein into the non-human mammalian cell; and (b) introducing a nuclease into the non-human mammalian cell, wherein the nuclease creates a single-stranded or double-stranded break at an immunoglobulin constant region site in the genome of the non-human mammalian cell, and the nucleic acid construct is integrated into the genome of the non-human mammalian cell at the immunoglobulin constant region site by homologous recombination. In embodiments, the immunoglobulin constant region site is an immunoglobulin light chain constant region site. In embodiments, the immunoglobulin light chain constant region site is a kappa light chain constant region site. In embodiments, the immunoglobulin light chain constant region site is a lambda light chain constant region site. In embodiments, the immunoglobulin constant region site is an immunoglobulin heavy chain constant region site. In embodiments, the immunoglobulin heavy chain constant region site is a gamma, delta, alpha, mu, or epsilon immunoglobulin constant region site.
[0029] In an embodiment of the method, introducing the nuclease comprises introducing an expression construct encoding the nuclease. In an embodiment, introducing the nuclease comprises introducing an mRNA encoding the nuclease. In an embodiment, the nuclease comprises a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, or a clustered regularly interspaced short palindromic repeats (CRISPR) associated (Cas) protein and a guide RNA (gRNA). In an embodiment, the gRNA comprises a CRISPR RNA (crRNA) that targets a recognition site and a transactivating CRISPR RNA (tracrRNA). In an embodiment, the CRISPR-Cas protein comprises Cas9.
[0030] In an embodiment of the method, the non-human mammalian cell is a rodent cell. In an embodiment, the rodent cell is a rat cell or a mouse cell. In an embodiment, the non-human mammalian cell is a pluripotent cell. In an embodiment, the pluripotent cell is a non-human embryonic stem (ES) cell. In an embodiment, the pluripotent cell is a mouse embryonic stem (ES) cell or a rat embryonic stem (ES) cell.
[0031] In embodiments, the method further comprises isolating a genetically modified non-human mammalian cell into which the nucleic acid construct has been integrated at an immunoglobulin constant region site. In embodiments, the immunoglobulin constant region site is an immunoglobulin light chain constant region site. In embodiments, the immunoglobulin light chain constant region site is a kappa light chain constant region site. In embodiments, the immunoglobulin light chain constant region site is a lambda light chain constant region site. In embodiments, the immunoglobulin constant region site is an immunoglobulin heavy chain constant region site. In embodiments, the immunoglobulin heavy chain constant region site is a gamma, delta, alpha, mu, or epsilon immunoglobulin constant region site.
[0032] In embodiments, the present disclosure provides genetically modified non-human mammalian cells produced by the methods disclosed herein.
[0033] In embodiments, the method further comprises injecting the isolated cell into a blastocyst to generate a transgenic non-human mammal comprising the nucleic acid construct integrated into an immunoglobulin constant region locus. In embodiments, the immunoglobulin constant region locus is an immunoglobulin light chain constant region locus. In embodiments, the immunoglobulin light chain constant region locus is a kappa light chain constant region locus. In embodiments, the immunoglobulin light chain constant region locus is a lambda light chain constant region locus. In embodiments, the immunoglobulin constant region locus is an immunoglobulin heavy chain constant region locus. In embodiments, the immunoglobulin heavy chain constant region locus is a gamma, delta, alpha, mu, or epsilon immunoglobulin constant region locus. In embodiments, the disclosure provides a genetically modified non-human transgenic mammal generated by the method.
[0034] In embodiments, the disclosure provides a genetically modified non-human mammalian cell comprising a genome comprising a nucleic acid construct described herein integrated into an immunoglobulin constant region site. In embodiments, the genetically modified non-human mammalian cell comprises a genome comprising a nucleic acid construct described herein integrated into an immunoglobulin constant region site. In embodiments, the immunoglobulin constant region site is a light chain constant region site. In embodiments, the light chain constant region site is a kappa constant region site. In embodiments, the light chain constant region site is a lambda constant region site. In embodiments, the constant region site is a heavy chain constant region site. In embodiments, the immunoglobulin expressing cell is obtained from the immunized mammal. In embodiments, the cell is an immunoglobulin expressing cell. In embodiments, the genetically modified non-human mammalian cell expresses an immunoglobulin kappa light chain.
[0035] In embodiments of the immunoglobulin-expressing non-human mammalian cell, the cell is an immature B cell or a progeny of an immature B cell. In embodiments, the cell is a hybridoma, a stem cell, or an immortalized cell.
[0036] In an embodiment of the immunoglobulin-expressing non-human mammalian cell, the cell expresses a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein. In an embodiment, the affinity tag is expressed on the cytoplasmic surface of the non-human mammalian cell. In an embodiment, the affinity tag comprises a StrepII tag. In an embodiment, the fluorescent reporter protein is exposed on the cytoplasmic surface of the non-human mammalian cell. In an embodiment, the fluorescent reporter protein comprises a green fluorescent protein (GFP), an enhanced green fluorescent protein (EGFP), an enhanced yellow fluorescent protein (EYFP), or an enhanced cyan fluorescent protein (ECFP). In an embodiment, the fluorescent reporter protein comprises a red fluorescent protein (RFP). In an embodiment, the red fluorescent protein is a monomeric cherry (mCherry) or a tandem dimer Tomato (tdTomato). Other fluorescent proteins are known and can be used in the constructs described herein. See, for example, Li et al. (2018) "Overview of the reporter genes and reporter mouse models," Anim Models and Exp Med.1:29-35 (doi.org / 10.1002 / ame2.12008).
[0037] In an embodiment of the immunoglobulin-expressing non-human mammalian cell, expression of the fusion protein is driven by an endogenous immunoglobulin transcription regulator. In an embodiment, the endogenous immunoglobulin transcription regulator is an endogenous immunoglobulin light chain transcription regulator. In an embodiment, the endogenous immunoglobulin light chain transcription regulator comprises a promoter and other cis elements in the mouse light chain site. In an embodiment, the endogenous immunoglobulin kappa light chain transcription regulator comprises a promoter and other cis elements in the mouse light chain site. In an embodiment, the endogenous immunoglobulin lambda light chain transcription regulator comprises a promoter and other cis elements in the mouse light chain site. In an embodiment, the endogenous immunoglobulin transcription regulator is an endogenous immunoglobulin heavy chain transcription regulator. In an embodiment, the endogenous immunoglobulin heavy light chain transcription regulator comprises a promoter and other cis elements in the mouse heavy chain site.
[0038] In embodiments, the disclosure provides a method for identifying immunoglobulin-expressing cells obtained from a genetically modified non-human mammal, the method comprising: (a) obtaining cells from the genetically modified non-human mammal described herein; (b) screening the cells obtained from the genetically modified non-human mammal for expression of a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein; and (c) identifying the immunoglobulin-expressing cells based on the expression of the fusion protein.
[0039] In an embodiment of the method, the cells are screened by fluorescence activated cell sorting (FACS) or magnetic activated cell sorting (MACS). In an embodiment, the affinity tag is expressed on the cell surface of the genetically modified non-human mammalian cell. In an embodiment, the affinity tag comprises a StrepII tag. In an embodiment, the fluorescent reporter protein is exposed on the cytoplasmic surface of the non-human mammalian cell. In an embodiment, the fluorescent reporter protein comprises a green fluorescent protein (GFP), an enhanced green fluorescent protein (EGFP), an enhanced yellow fluorescent protein (EYFP), or an enhanced cyan fluorescent protein (ECFP). In an embodiment, the fluorescent reporter protein comprises a red fluorescent protein (RFP). In an embodiment, the red fluorescent protein is a monomeric cherry (mCherry) or a tandem dimer tomato (tdTomato).
[0040] In an embodiment of the method, the genetically modified non-human mammal has been immunized with an antigen of interest. In an embodiment, the immunoglobulin-expressing cells express an immunoglobulin light chain. In an embodiment, the immunoglobulin-expressing cells express an immunoglobulin kappa light chain. In an embodiment, the immunoglobulin-expressing cells express an immunoglobulin lambda light chain. In an embodiment, the immunoglobulin-expressing cells express an immunoglobulin heavy chain. In an embodiment, the immunoglobulin-expressing cells include immature B cells and their progeny.
[0041] In embodiments, the method further comprises isolating the expressed immunoglobulin from the cell obtained from the genetically modified non-human mammal. In embodiments, the disclosure provides an immunoglobulin obtained by the method.
[0042] In embodiments, the disclosure provides a method of producing a therapeutic or diagnostic immunoglobulin, the method comprising (i) cloning a variable region of an immunoglobulin described herein and (ii) producing a therapeutic or diagnostic immunoglobulin comprising the variable region obtained in (i).
[0043] In embodiments, the disclosure provides a method of producing a monoclonal antibody, the method comprising (i) obtaining an immunoglobulin-expressing cell from a genetically modified non-human mammal as described herein, (ii) immortalizing the immunoglobulin-expressing cell obtained in (i), and (iii) isolating a monoclonal antibody or a nucleic acid sequence encoding the monoclonal antibody expressed by the immortalized immunoglobulin-expressing cell. In embodiments, the method further comprises (iv) cloning the variable region of the isolated monoclonal antibody, and (v) producing a therapeutic or diagnostic antibody comprising the cloned variable region. In embodiments, the disclosure provides a therapeutic or diagnostic antibody produced by the method. [Brief description of the drawings]
[0044] [Figure 1A-C]Schematic diagram of the construction and use of one embodiment of the conditional reporter nucleic acid construct described herein. As shown in FIG. 1A, the nucleic acid construct is inserted into the safe harbor site of the ROSA26 site. In the diagram, CAGGS represents the CAG promoter, L represents the leader sequence, the LoxP-Stop-LoxP cassette is positioned as shown, STX3 represents three tandem repeats of the Strep-II tag, TM represents the transmembrane domain, and GFP represents the green fluorescent protein reporter. FIG. 1B is a schematic diagram of the cross performed between the Cre switch line and the conditional reporter line to form tissue-specific reporter mouse lines. As shown in the schematic, after crossing the conditional reporter line with the switch line, the Cre recombinase removes the stop codon in front of the reporter and turns on expression of the reporter in the nuclei of Cre-expressing cells. As a result, these cells are permanently labeled with the affinity tag on the cell surface and the fluorescent marker within the cells. FIG. 1C is a schematic diagram depicting how cells isolated from the switch reporter line of FIG. 1B are separated using FACS or MACS as described herein.
[0045] [Diagram 2] Schematic diagram of the targeting strategy for mouse / rat IgK locus. After targeting, the tag cassette is knocked in at the stop codon of the IgK gene and under the control of the IgK locus promoter (note that LK in the diagram below is a linker sequence, more details below). In the diagram, the black rectangle represents the V and J segments of the region, LK is the linker sequence, L is the leader sequence, STX3 represents three tandem repeats of the Strep-II tag, TM is the transmembrane domain, and GFP is the green fluorescent protein reporter.
[0046] [Diagram 3] FIG. 1 is a schematic diagram of the configuration of an embodiment of an IgK reporter mouse formed as described herein; and FIG. 2 is a schematic diagram of how pooled cells isolated from the mouse are separated using FACS or MACS as described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0047] The present disclosure provides embodiments of a nucleic acid construct comprising an affinity tag, a transmembrane (TM) domain, and a transmembrane reporter cassette encoding a fluorescent reporter protein. In embodiments, the nucleic acid construct is inserted into a safe harbor site or an immunoglobulin constant region site in a cell of a non-human mammal. In embodiments, when the transmembrane reporter cassette is expressed in the cell, the affinity tag is displayed on the cell surface and the fluorescent reporter protein is located inside the cell membrane. The presence of the affinity tag and the fluorescent reporter protein allows for identification, sorting, and / or isolation of cells expressing the nucleic acid construct. The present disclosure also provides embodiments of cells and non-human organisms produced using the disclosed methods, as well as embodiments of methods of modifying cells and non-human organisms using the nucleic acid construct.
[0048] A.Definition Unless otherwise defined, scientific and technical terms used herein shall have the meanings commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include the plural, and plural terms shall include the singular, e.g., "a" or "an," and plural terms shall include the plural, e.g., "one or more" or "at least one," and the term "or" may mean "and / or." The terms "including," "includes," and "included" are not limiting. Ranges provided herein, regardless of type, include all values within the particular ranges described, as well as the endpoints of the particular ranges.
[0049] The term "about" as used herein is used to modify, for example, the amount, concentration, volume, process temperature, process time, yield, flow rate, pressure, and ranges thereof of components in a composition. The term "about" refers to the variation in numerical quantity that may occur, for example, in typical measuring and handling procedures used to make a compound, composition, concentrate, or formulation, inadvertent errors in these procedures, in the manufacturing equipment used to carry out the method, in the raw materials, or in the purity of the starting materials or components, and other similar considerations. The term "about" also includes amounts that vary with age of a formulation having a particular initial concentration or mixture, and amounts that vary with mixing or processing of a formulation having a particular initial concentration or mixture. When modified with the term "about", the appended claims include such equivalents.
[0050] Generally, the nomenclature used in connection with cell and tissue culture, molecular biology, protein chemistry and oligo- or polynucleotide chemistry, and hybridization described herein, and the techniques thereof, are well known and commonly used in the art. Amino acids may be referred to herein by either their commonly known three-letter symbols or the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides may also be referred to by their commonly accepted one-letter codes.
[0051] As used herein, the terms "polypeptide" or "protein" can be used interchangeably to refer to a molecule having two or more amino acid residues bound to each other by peptide bonds. The term "polypeptide" can refer to antibodies and other non-antibody proteins. Non-antibody proteins include, but are not limited to, proteins such as enzymes, receptors, ligands for cell surface proteins, secreted proteins, and fusion proteins or fragments thereof. Polypeptides can be of scientific or commercial interest, including protein-based therapeutics.
[0052] As used herein, the terms "antibody" and "immunoglobulin" are used interchangeably and refer to a polypeptide or group of polypeptides that contain at least one binding domain formed from the folding of a polypeptide chain with a three-dimensional binding space that has an internal surface shape and charge distribution complementary to the antigenic determinant characteristics of an antigen. Naturally occurring antibodies are usually tetramers, with two pairs of polypeptide chains, each pair having one "light chain" and one "heavy chain". The variable regions of each light / heavy chain pair form the antibody binding site. Each light chain is linked to a heavy chain by one covalent disulfide bond, but the number of disulfide bonds varies depending on the immunoglobulin isotype. Each heavy and light chain has disulfide bonds at regular intervals within the chain. At one end of each heavy chain is a variable region (VH) followed by several constant regions (CH). Each light chain has a variable region (VL) at one end and a constant region (CL) at the other end, with the light chain constant region aligned with the first constant region of the heavy chain, and the light chain variable region aligned with the variable region of the heavy chain. Light chains are classified as lambda chains and kappa chains based on the amino acid sequence of the light chain constant region. Heavy chains are classified as gamma chains, delta chains, alpha chains, mu chains, or epsilon chains based on the amino acid sequence of the heavy chain constant region.
[0053] The term "antigen-binding fragment" or "immunologically active fragment" refers to a fragment of an antibody that contains at least one antigen-binding site and retains the ability to specifically bind to an antigen. Immunoglobulin molecules can be of any isotype (e.g., IgG, IgE, IgM, IgD, IgA, and IgY, etc.), subisotype (e.g., IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2, etc.), or allotype (e.g., Gm, e.g., G1m (f, z, a, or x), G2m (n), G3m (g, b, or c), Am, Em, and Km (1, 2, or 3)). Subtypes include subclasses found in non-human mammals, such as rodents, e.g., IgG1, IgG2a, IgG2b, IgG2c, and IgG3. Immunoglobulins include, but are not limited to, monoclonal antibodies (including full-length monoclonal antibodies), polyclonal antibodies, multispecific antibodies formed from at least two different epitope-binding fragments (e.g., bispecific antibodies), CDR-grafted antibodies, human antibodies, humanized antibodies, camelized antibodies, chimeric antibodies, anti-idiotypic (anti-Id) antibodies, intrabodies, and any desired antigen-binding fragment thereof, including recombinantly produced antibody fragments. Examples of antibody fragments that can be recombinantly produced include, but are not limited to, antibody fragments that contain variable heavy and light chain domains, such as single chain Fvs (scFvs), single chain antibodies, Fab fragments, Fab' fragments, F(ab')2 fragments, etc. Antibody fragments also include epitope-binding fragments or derivatives of the antibodies listed above.
[0054] The term "recombinant" refers to biological material, such as a nucleic acid or protein, that has been artificially or synthetically (i.e., non-naturally) modified or produced by human intervention. The term "recombinant antibody" refers to antibodies prepared by recombinant DNA processes, such as, for example, antibodies expressed using a recombinant expression vector transfected into a host cell, as well as antibodies isolated from a recombinant combinatorial human antibody library. In embodiments, the recombinant antibody is a recombinant human antibody, including, but not limited to, antibodies isolated from a transgenic animal carrying human immunoglobulin genes or antibodies prepared by splicing human immunoglobulin gene sequences into another DNA sequence.
[0055] As used herein, a "coding sequence" or a sequence that "encodes" a selected polypeptide refers to a nucleic acid molecule that can be transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide, e.g., in vivo, when placed under the control of appropriate regulatory sequences (or "control elements"). The boundaries of the coding sequence are usually determined by a start codon at the 5' (amino) terminus and a translation stop codon at the 3' (carboxy) terminus. Coding sequences include, but are not limited to, cDNA from viral, prokaryotic, or eukaryotic mRNA, genomic DNA sequences from viral or prokaryotic DNA, and even synthetic DNA sequences. A transcription termination sequence may also be located 3' of the coding sequence. Other "control elements" may also be associated with a coding sequence. A DNA sequence encoding a polypeptide can be optimized for expression in a selected cell by expressing the DNA copy of the desired polypeptide coding sequence with codons preferred by the selected cell.
[0056] As used herein, "encoded" refers to a nucleic acid sequence that encodes a polypeptide sequence, or a portion thereof, that includes an amino acid sequence of at least about 3 to about 5 amino acids, at least about 8 to about 10 amino acids, or at least about 15 to about 20 amino acids from the polypeptide encoded by the nucleic acid sequence. Also included are polypeptide sequences that are immunologically distinguishable from the polypeptide encoded by the sequence.
[0057] As used herein, "operably linked" refers to an arrangement of elements configured so that the components thus described perform their normal functions. Thus, a given promoter operably linked to a coding sequence (e.g., a reporter expression cassette) can cause expression of the coding sequence when the appropriate enzymes are present. A promoter or other control element need not be contiguous with a coding sequence, so long as it functions to direct its expression. For example, a promoter sequence can be considered to be "operably linked" to a coding sequence even if there is an intervening sequence between the promoter sequence and the coding sequence that is not yet transcribed.
[0058] As used herein, a "vector" is capable of introducing a gene sequence into a target cell. Typically, the terms "vector construct", "expression vector" and "gene transfer vector" refer to any nucleic acid construct that can induce expression of a gene of interest and introduce a gene sequence into a target cell. Thus, the term includes cloning, expression vehicles and integrating vectors.
[0059] As used herein, an "expression cassette" includes any nucleic acid construct capable of directing the expression of a gene / coding sequence of interest. Such cassettes may be constructed into "vectors," "vector constructs," "expression vectors," or "gene transfer vectors" for introducing the expression cassette into a target cell. Thus, the term also includes cloning and expression vehicles, as well as viral vectors.
[0060] As used herein, a "tandem repeat" is a repeat of one or more nucleotides (in a nucleic acid) or one or more amino acid residues (in a protein) where the repeats are adjacent to one another in the sequence. The tandem repeats may be contiguous (i.e., there are no other nucleotides or residues between the repeats) or the tandem repeats may be separated by one or more nucleotides or residues between the repeats.
[0061] The term "expression vector" as used herein refers to any suitable recombinant expression vector that can be used to transform or transfect a suitable host cell. The term "host cell" as used herein refers to a cell into which a recombinant expression vector has been introduced. The term "host cell" refers not only to a cell into which an expression vector has been introduced (a "parent" cell) but also to the progeny of such a cell. Since modifications may occur in the progeny due to, for example, mutations or environmental influences, the progeny cell may not be identical to the parent cell, but is still within the scope of the term "host cell."
[0062] The term "transformation" as used herein refers to the genetic change of a cell caused by the incorporation of foreign DNA. Suitable methods for transforming cells include viral infection, transfection, conjugation, protoplast fusion, electroporation, particle gun technology, calcium phosphate precipitation, direct microinjection, and the like. The choice of method generally depends on the type of cell to be transformed and the context in which the transformation is performed (i.e., in vitro, ex vivo, in vivo). A general discussion of these methods can be found in Ausubel, et al, Short Protocols in Molecular Biology, 3rd ed, Wiley & Sons, 1995.
[0063] As used herein, "immortalized cells" refers to cells of a type that normally does not proliferate indefinitely, but that have a mutation that circumvents cellular senescence, allowing them to continue dividing indefinitely. In embodiments, immortalized cell lines can be derived from tumor cell lines, or from cell lines that have been engineered to allow cells to proliferate indefinitely.
[0064] As used herein, "knock-in" refers to a transgenic cell or animal produced by genetic engineering techniques that inserts a heterologous DNA sequence at a specific genomic location. In some aspects, the heterologous DNA sequence is inserted by homologous recombination. In some aspects, the heterologous DNA sequence is inserted using the CRISPR / Cas9 system. In some aspects, the heterologous DNA sequence is inserted at a "safe harbor site." As used herein, a "safe harbor site" refers to a site in the genome that can accommodate the incorporation of new genetic material such that the new genetic element functions as predicted and does not cause changes in the host genome that pose a danger to the host cell or organism. A "knock-in" includes progeny that contain a heterologous DNA sequence in at least one allele. In embodiments, the addition of a reporter gene at this site allows for tracing the lineage of the cell.
[0065] As used herein, "heterologous" refers to a nucleic acid that does not naturally occur within a cell or animal, or a nucleic acid that is naturally present in a cell or animal but has been modified or mutated.
[0066] As used herein, "transmembrane domain" (TM domain) refers to a generally hydrophobic region of a protein that crosses the plasma membrane of a cell. In embodiments, the TM domain links an extracellular portion of a construct to an intracellular portion. In embodiments, the TM domain links an extracellular affinity tag and an intracellular fluorescent reporter protein. The TM domain may comprise a transmembrane region of a protein, a transmembrane fragment of a protein, an artificial hydrophobic sequence, or a combination thereof. In embodiments, the transmembrane domain is a type I transmembrane protein. In embodiments, the TM domain comprises one or more α-helices. In embodiments, the TM domain comprises one or more β-strands. In embodiments, the transmembrane domain comprises an IgG transmembrane domain. In embodiments, the transmembrane domain comprises a human IgG transmembrane domain. In embodiments, the transmembrane domain comprises a mouse IgG transmembrane domain.
[0067] In embodiments, the transmembrane domain comprises a mammalian transmembrane domain. In embodiments, the transmembrane domain comprises the transmembrane domain of mouse protein Tmem53, Lrtm1, or Nrg1. Although specific examples are provided herein, other transmembrane domains will be apparent to those of skill in the art and can be used in conjunction with the constructs described herein. See, e.g., Yu and Zhang (2013) "A simple method for predicting transmembrane proteins based on wavelet transform," Int.J. Biol.Sci.9(1):22-33.
[0068] Development of BB cells In embodiments, the cells identified and / or isolated using the methods described herein are B cells. B cells arise from hematopoietic stem cells (HSCs) in the bone marrow and undergo several antigen-independent developmental steps to generate immature B cells. Immature B cells express IgM on their surface (membrane IgM expression). Immature B cells migrate from the bone marrow to the spleen where they differentiate into mature naive B cells (membrane IgM and IgD expression). Some of these mature naive B cells differentiate into memory B cells. Memory B cells are long-lived, quiescent cells that are rapidly activated, proliferate, and differentiate into plasma cells to fight new infections when re-exposed to antigen. When naive or memory B cells are activated by antigen, they proliferate and differentiate into antibody-secreting cells.
[0069] Subsequently, when the cells fully mature into plasma cells, they express secretory Ig but lose surface Ig expression. In mice and rats, approximately 99% of antibody-expressing cells use Ig kappa as the light chain.
[0070] After hematopoietic stem cells commit to the B cell lineage, B cell precursors undergo a series of differentiation events to become mature B cells.
[0071] C. Homologous recombination and site-specific nucleases As described herein, the present disclosure provides a method for producing genetically modified non-human mammalian cells and organisms, the method comprising introducing a nuclease into a non-human mammalian cell, the nuclease causing a single-strand or double-strand break at a location in the genome of the modified cell. In an embodiment, repair of the single-strand or double-strand break results in integration of the nucleic acid sequence into the genome of the modified cell. In an embodiment, the integration occurs via homologous recombination.
[0072] Homologous recombination (HR): Homologous recombination allows the insertion of a target gene at a specific site in the genome of an organism (gene targeting). Creating a DNA construct containing a template that matches the genomic sequence to be targeted allows the HR process in the cell to insert the construct at the desired location. Using this method with embryonic stem cells, transgenic mice have been developed in which targeted genes have been knocked out, i.e., removed from the genome, or knocked in, i.e., added to the genome.
[0073] Gene knock-in method using double-strand break followed by HR has been described in the art, for example, in U.S. Patent Nos. 5,474,896; 5,792,632; 5,866,361; 5,948,678; 5,948,678, 5,962,327; 6,395,959; 6,238,924; and 5,830,729, each of which is incorporated herein by reference.Exemplary methodologies for homologous recombination are described in U.S. Patent Nos. 6,689,610; 6,204,061; 5,631,153; 5,627,059; 5,487,992; and 5,464,764, each of which is incorporated herein by reference.
[0074] In embodiments, the single-stranded or double-stranded break is introduced using a site-specific nuclease. Such nucleases are known in the art and examples of such nucleases are provided herein.
[0075] Zinc finger nucleases: Zinc finger nucleases have a DNA binding domain and can precisely target DNA sequences. Each zinc finger can recognize a portion of a desired DNA sequence and can therefore be modularly assembled to bind to specific sequences. The binding domain guides the cleavage of a restriction enzyme, which causes a double-stranded break in the DNA.
[0076] Transcription Activator-Like Effector Nucleases (TALENs): Transcription Activator-Like Effector Nucleases (TALENs) also contain a DNA-binding domain and a nuclease that cleaves DNA. The DNA-binding domain contains amino acid repeats that each recognize one base pair in the desired target DNA sequence. The nuclease creates a double-stranded break in the DNA.
[0077] CRISPR / Cas: Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) / CRISPR-associated proteins (Cas) is a genome editing method that includes a guide RNA complexed with a Cas protein. The guide RNA can be designed to match the desired DNA sequence by simple complementary base pairing, as opposed to the assembly of constructs required by zinc fingers or TALENs. The bound Cas creates a double-stranded break in the DNA. In an embodiment, the Cas protein includes Cas9. In an embodiment, the Cas protein includes Cas9, Cas12, Cas12a, Cas13, Cas14, or CasΦ. In an embodiment, the Cas protein includes Cas3, Cas8, Cas10, Cas11, Cas12, Cas12a, Cas13, Cas14, or CasΦ.
[0078] D. Cre recombinase Cre (Cre recombinase) is one of the tyrosine site-specific recombinases (T-SSRs) that include flippase (Flp) and D6-specific recombinase (Dre). It was discovered as a 38 kDa DNA recombinase produced from the cre (cyclization recombinase) gene of bacteriophage P1. It is a 38 kDa DNA recombinase produced by the loxP ( l ocus o f x -over, P 1) and mediates the site-specific deletion of the DNA sequence between the two loxP sites. The loxP site is a 34-bp sequence that contains two 13-bp inverted and palindromic repeats and an 8-bp core sequence.
[0079] As described further herein, appropriate insertion of loxP-flanked "stop" sequences (transcription termination elements) between the leader sequence and the reporter sequence encoding the transgene shuts off expression of the gene. After mating the conditional reporter line with the switch line, Cre recombinase removes the stop element in front of the reporter, turning on expression of the reporter in the nuclei of Cre-expressing cells. As a result, these cells are permanently labeled with an affinity tag on the cell surface and a fluorescent marker within the cells. This process is shown diagrammatically in Figure 1B.
[0080] Thousands of mouse lines have been developed in which Cre is under the control of tissue-specific promoters, meaning that Cre is only expressed in certain tissues of the mouse. By crossing different switch lines, as described further herein, different cells of interest can be labeled with reporter proteins and isolated using large-scale magnetic- or flow-based methods.
[0081] E. Rodent immunoglobulins Like humans, mice and rats have five antibody isotypes (IgA, IgD, IgE, IgG, and IgM). Each isotype has a different heavy chain. Isotypes are sometimes called classes. Naive B cells produce IgM and IgD. During B cell maturation, isotype switching allows mature B cells to produce either IgG, IgA, or IgE isotypes and subclasses. Different isotypes have different in vivo half-lives, ranging from 12 hours to 8 days.
[0082] IgA, IgD, and IgG heavy chains have a constant region with three immunoglobulin (Ig) domains. Other types of heavy chains may have a different number of immunoglobulin domains. IgE and IgM heavy chains have a constant region with four immunoglobulin domains. Each of the above isotypes of heavy chains has a membrane-bound and secreted form at the C-terminal region due to alternative splicing events during transcription. The membrane-bound mRNA contains two more exons at the C-terminus. Thus, the membrane-bound heavy chain protein is a longer protein with a transmembrane domain and a cytoplasmic C-terminal tail. All isotypes of heavy chains have a variable region with a single immunoglobulin domain.
[0083] Each light chain (either kappa or lambda) has one constant immunoglobulin domain and one variable immunoglobulin domain. In rats and mice, the ratio of kappa to lambda light chain usage is approximately 99 to 1, meaning that about 99% of antibody-expressing cells express kappa light chains. The mouse immunoglobulin kappa (kappa) light chain multigene family contains a constant region locus (Ckappa), four joining region genes, and a family of about 95 kappa variable (Vkappa) regions.
[0084] F. Transgenic Animals A "transgenic animal" is a non-human animal, usually a mammal, in which an exogenous nucleic acid sequence is present in some of its cells as an extrachromosomal element or stably integrated into the germline DNA (i.e., most or all of the genomic sequence of the cell). In embodiments herein, the transgenic animal comprises an exogenous nucleic acid introduced into the germline of such transgenic animal, for example, by genetically manipulating the host animal's embryo or embryonic stem cell, according to methods well known in the art. In embodiments herein, the transgenic animal comprises more than the nucleic acid reporter constructs described herein. In embodiments, the transgenic animal comprises one or more additional nucleic acids encoding a product produced by the transgenic animal, for example, a protein such as an enzyme or immunoglobulin, or a nucleic acid such as DNA or RNA. In a specific aspect, the methods herein provide for the generation of transgenic animals that comprise a human immunoglobulin region introduced in part with a nucleic acid encoding a reporter construct described herein.
[0085] In embodiments, the transgenic animal is a rodent, such as a mouse or a rat. In embodiments, the transgenic rodent comprises human immunoglobulin sequences and endogenous mouse immunoglobulin regions to generate partially or fully human antibodies for drug discovery purposes. Examples of such mice include those described in, for example, U.S. Patent Nos. 7,145,056; 7,064,244; 7,041,871; 6,673,986; 6,596,541; 6,570,061; 6,162,963; 6,130,364; 6,091,001; 6,023,010; 5,593,598; 5,877,397; 5,874,299; 5,814,318; 5,789,650; 5,661,016; 5,612,205; and 5,591,669, which are incorporated herein by reference. In embodiments, the transgenic rodent is a transgenic mouse whose genome contains an entire endogenous mouse immunoglobulin site variable region deleted and replaced with a modified immunoglobulin site variable region. Examples of such mice include those described, for example, in U.S. Pat. No. 10,881,084 and U.S. Patent Application Publication No. 2020 / 0190218. In embodiments, the transgenic mouse is modified to express human or partially human antibodies. In other embodiments, the transgenic mouse is modified to express canine, equine, or bovine antibodies. See U.S. Patent Application Publication Nos. 10,793,829, 2020 / 0308307, and 2021 / 0000087, and WO 2021 / 003152.
[0086] G. Cell Sorting Methods Fluorescence-activated cell sorting (FACS) is a specialized type of flow cytometry. This method allows for the sorting of a heterogeneous mixture of cells, one cell at a time, into two or more containers based on the specific light scattering and fluorescence properties of each cell. FACS is performed using cell sorting equipment designed for this technique. FACS rapidly, objectively, and quantitatively records the fluorescent signals from individual cells and physically separates cells of particular interest.
[0087] In an embodiment, FACS is generally performed as follows: A suspension of cells to be sorted is entrained in the center of a narrow, rapidly flowing stream of liquid. The stream is arranged so that the spacing between cells is large relative to its diameter. A vibration mechanism breaks the stream of cells into individual droplets. The system is adjusted so that the probability of more than one cell per droplet is low. Just before the stream splits into droplets, it passes through a fluorescence measurement station where the fluorescent properties of each cell are measured. A charging ring is placed at the exact point where the stream splits into droplets. A charge is applied to the ring based on the previous fluorescence intensity measurement, and an opposite charge is captured on the droplet as it leaves the stream. The charged droplets fall through an electrostatic deflection system that redirects them into a container based on their charge. In some systems, the charge is applied directly to the stream, and the separating droplets retain a charge of the same sign as the stream. The stream then returns to neutral after the droplets break off, and the next droplet is measured and sorted.
[0088] Magnetically Activated Cell Sorting (MACS; Miltenyi Biotech) is a method to separate cells by cell surface markers. In an embodiment, the MACS system uses superparamagnetic nanoparticles and a column. Superparamagnetic nanoparticles are nanoparticles on the order of 100 nm. The nanoparticles tag targeted cells and capture them in the column. The column is placed between permanent magnets so that the tagged cells are captured when the magnetic particle-cell complex passes through it. The magnetic nanoparticles are coated with a reagent that binds a specific marker to their surface. Cells that express the marker attach to the magnetic nanoparticles. After incubating the beads and cells, the solution is transferred to the column in a strong magnetic field. Cells that attach to the nanoparticles (expressing the marker) remain in the column, while other cells (not expressing the marker) pass through the column.
[0089] In embodiments, cells are sorted using affinity tags expressed on the cell surface. Affinity tags for this purpose are described herein. In these embodiments, cells can be sorted using affinity purification columns or resins that bind to affinity tags, using methods known in the art. As a non-limiting example, if cells express a StrepII tag on their surface, resins that bind to StrepII tags, such as Strep-Tactin® Sepharose® (IBA Lifesciences), can be used to capture cells and thus sort them.
[0090] H. Conditional Reporter Nucleic Acid Constructs In embodiments, provided herein is a nucleic acid construct comprising a leader sequence, a LoxP-Stop-LoxP cassette, and a transmembrane reporter cassette encoding an affinity tag, a transmembrane (TM) domain, and a fluorescent reporter protein. This embodiment of the nucleic acid construct may be referred to herein as a "conditional reporter nucleic acid construct."
[0091] In an embodiment of the construct, a leader sequence is present upstream of the LoxP-Stop-LoxP cassette. In an embodiment, a leader sequence is present downstream of the LoxP-Stop-LoxP cassette.
[0092] In embodiments, the LoxP-Stop-LoxP cassette comprises a stop element, e.g., any type of sequence that terminates translation or transcription. In embodiments, the stop element comprises one or more SV40 polyadenylation sequences.
[0093] In an embodiment, the LoxP-Stop-LoxP cassette comprises two LoxP sites flanking a sequence that causes termination of transcription. In an embodiment, the LoxP-Stop-LoxP cassette comprises a LoxP-flanked polyadenylation sequence that causes termination of transcription. In an embodiment, the polyadenylation signal is an SV40, hGH, BGH, or rbGlob polyadenylation signal. In an embodiment, the LoxP-Stop-LoxP cassette comprises a LoxP-flanked triple repeat of polyadenylation sequence. In an embodiment, the LoxP-Stop-LoxP cassette comprises a LoxP-flanked double repeat of polyadenylation sequence. In an embodiment, the LoxP-Stop-LoxP cassette comprises a LoxP-flanked single polyadenylation sequence.
[0094] In an embodiment, the LoxP-Stop-LoxP cassette comprises a LoxP-flanked triple repeat of the SV40 polyadenylation sequence. In an embodiment, the LoxP-Stop-LoxP cassette comprises a LoxP-flanked double repeat of the SV40 polyadenylation sequence. In an embodiment, the LoxP-Stop-LoxP cassette comprises a LoxP-flanked single SV40 polyadenylation sequence.
[0095] In an embodiment, the stop element comprises one or more stop codons that cause the termination of translation. In an embodiment, the LoxP-Stop-LoxP cassette comprises a LoxP-flanked stop codon. In an embodiment, the stop codon is TAG, TAA, or TGA.
[0096] In embodiments, the nucleic acid construct comprises a single-stranded DNA, a double-stranded DNA, a plasmid, or a viral vector. In embodiments, the nucleic acid construct is a linear DNA. In embodiments, the nucleic acid construct is a circular DNA.
[0097] In an embodiment, the nucleic acid construct further comprises a first homology arm and a second homology arm that are homologous to a first target sequence and a second target sequence in the genome of the non-human mammal. The homology region allows the nucleic acid construct to be integrated into the genome of the non-human mammal using methods described herein and known in the art. In an embodiment, the nucleic acid construct further comprises a first homology arm and a second homology arm that are homologous to a first target sequence and a second target sequence, respectively, within a safe harbor site in the non-human mammal.
[0098] In an embodiment, the first homology arm and the second homology arm each independently comprise about 15 nucleotides to about 12000 nucleotides. In an embodiment, the first homology arm and the second homology arm each independently comprise about 30 nucleotides to about 11000 nucleotides. In an embodiment, the first homology arm and the second homology arm each independently comprise about 50 nucleotides to about 10000 nucleotides. In an embodiment, the first homology arm and the second homology arm each independently comprise about 100 nucleotides to about 7500 nucleotides. In an embodiment, the first homology arm and the second homology arm each independently comprise about 200 nucleotides to about 5000 nucleotides. In an embodiment, the first homology arm and the second homology arm each independently comprise about 300 nucleotides to about 2500 nucleotides.
[0099] In embodiments, a safe harbor site is any site in the genome that can accommodate the integration of new genetic material so that the new genetic element functions as expected and does not cause changes in the host genome that pose a danger to the host cell or organism. In embodiments, the safe harbor site is a mouse safe harbor site. In embodiments, the safe harbor site is a rat safe harbor site. In embodiments, the safe harbor site comprises the Rosa26 site on chromosome 6 of the mouse genome. In embodiments, the safe harbor site comprises the Hipp11 site on chromosome 11 of the mouse genome.
[0100] In embodiments, the nucleic acid construct is expressed using an endogenous promoter. In embodiments, the nucleic acid construct is expressed using an endogenous promoter located in a safe harbor site.
[0101] In embodiments, the nucleic acid construct further comprises a promoter. In embodiments, the promoter is a mammalian constitutive promoter. In embodiments, the promoter is a human promoter. In embodiments, the promoter is a mouse promoter. In embodiments, the promoter is a rat promoter. In embodiments, the promoter is a viral promoter. In embodiments, the promoter comprises a CAG promoter. In embodiments, the promoter comprises a CAG, CMV, EF1a, SV40, PGK1, Ubc, or human beta actin promoter.
[0102] In an embodiment, the leader sequence comprises a secretory signal peptide. In an embodiment, the secretory signal peptide is an IL-2 leader sequence. In an embodiment, the secretory signal peptide is a human OSM, VSV-G mouse Ig kappa, human IgG2H, BM40, Secrecon, human IgKVIII, CD33, tPA, human chymotrypsinogen, human trypsinogen-2, human IL-12, or human serum albumin signal peptide. In an embodiment, the secretory signal peptide comprises an IL-2 leader sequence MYRMQLLSCIALSLALVTNS (SEQ ID NO: 2). One of skill in the art will appreciate that signal peptides can be predicted using algorithms known in the art, such as SignalP-5.0 (www.cbs.dtu.dk / services / SignalP / ) or SecretomeP 2.0 (www.cbs.dtu.dk / services / SecretomeP / ).
[0103] The affinity tag of the nucleic acid construct can be used to subsequently isolate or purify the protein expressed by the construct. In embodiments, the affinity tag comprises a StrepII, hexahistidine, FLAG, HA, Myc, VA, GST, β-GAL, MBP, or VSV-G tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of the tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of the tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of the tag. In embodiments, the affinity tag comprises 3 tandem repeats of the tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of the tag.
[0104] In embodiments, the affinity tag comprises a StrepII tag. In embodiments, the affinity tag comprises tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises 3 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of a StrepII tag. In an embodiment, the tandem repeat has a tag linker between the repeats. In an embodiment, the linker is a dipeptide or a tripeptide. In an embodiment, the tag linker is the dipeptide Ser-Ala.
[0105] In an embodiment, the StrepII tag comprises an eight amino acid peptide sequence of WSHPQFEK (SEQ ID NO: 1). In an embodiment, the affinity tag comprises WSHPQFEKSAWSHPQFEKSAWSHPQFEK (SEQ ID NO: 3).
[0106] The transmembrane domain encoded by the transmembrane reporter cassette allows the affinity tag to be presented on the surface of a cell expressing the nucleic acid construct. In an embodiment, the transmembrane domain comprises a hydrophobic alpha helix. In an embodiment, the transmembrane domain comprises an IgG transmembrane domain. In an embodiment, the transmembrane domain comprises a human IgG1 transmembrane domain. In an embodiment, the transmembrane domain comprises a mouse IgG transmembrane domain. In an embodiment, the transmembrane domain comprises a mouse IgG1, IgG2a, IgG2b, or IgG2c transmembrane domain. In an embodiment, the transmembrane domain comprises the transmembrane domain of mouse protein Tmem53, Lrtm1, or Nrg1.
[0107] The fluorescent reporter protein allows detection of cells expressing the nucleic acid construct. In an embodiment, the fluorescent reporter protein comprises green fluorescent protein, yellow fluorescent protein, cyan fluorescent protein, red fluorescent protein, blue fluorescent protein, red fluorescent protein, or orange fluorescent protein. In an embodiment, the fluorescent reporter protein comprises green fluorescent protein (GFP), enhanced green fluorescent protein (EGFP), enhanced yellow fluorescent protein (EYFP), or enhanced cyan fluorescent protein (ECFP). In an embodiment, the fluorescent reporter protein comprises red fluorescent protein (RFP). In an embodiment, the red fluorescent protein is monomeric cherry (mCherry) or tandem dimeric tomato (tdTomato).
[0108] In an embodiment, the nucleic acid construct is the nucleic acid construct shown diagrammatically in FIG. 1A, where CAGGS represents the CAG promoter, L represents the leader, the LoxP-Stop-LoxP cassette arranged as shown, STX3 represents three tandem repeats of the Strep-II tag, TM represents the transmembrane domain, and GFP represents the green fluorescent protein reporter.
[0109] I. Methods for generating conditional reporter modified cells and organisms In embodiments, provided herein is a method of generating a genetically modified non-human mammalian cell, the method comprising: (a) introducing a conditional reporter nucleic acid construct described herein into a non-human mammalian cell; and (b) introducing a nuclease into the non-human mammalian cell, wherein the nuclease causes a single-stranded or double-stranded break at a safe harbor site in the genome of the non-human mammalian cell, and wherein the nucleic acid construct is integrated into the genome of the non-human mammalian cell at the safe harbor site by homologous recombination.
[0110] In embodiments, the nuclease causes a double stranded break. In embodiments, the nuclease causes a single stranded break (e.g., a nick).
[0111] The nuclease can be introduced into the cell using methods known in the art. In embodiments, introducing the nuclease comprises introducing an expression construct encoding the nuclease. In embodiments, the introduction of the expression construct is by injection or electroporation. In embodiments, introducing the nuclease comprises introducing a plasmid encoding the nuclease. In embodiments, introducing the nuclease comprises introducing a viral vector encoding the nuclease. In embodiments, introducing the nuclease comprises introducing an mRNA encoding the nuclease. In embodiments, the mRNA comprises one or more modified bases. In embodiments, the mRNA is encapsulated in a lipid nanoparticle. In embodiments, the introduction of the plasmid, viral vector, or mRNA is by injection or electroporation. In embodiments, introducing the nuclease comprises directly introducing the nuclease protein into the cell. In embodiments, the direct introduction of the nuclease protein into the cell is by injection or electroporation.
[0112] In an embodiment, the nuclease is a nuclease as described herein. In an embodiment, the nuclease comprises a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, or a clustered regularly interspaced short palindromic repeats (CRISPR) associated (Cas) protein and a guide RNA (gRNA). In an embodiment, the gRNA comprises a CRISPR RNA (crRNA) and a transactivating CRISPR RNA (tracrRNA) that targets a recognition site. In an embodiment, the CRISPR-Cas protein comprises Cas9. In an embodiment, the Cas protein comprises Cas9, Cas12 Cas12a, Cas13, Cas14, or CasΦ. In an embodiment, the CRISPR-Cas protein comprises Cas3, Cas8, Cas10, Cas11, Cas12, Cas12a, Cas13, Cas14, or CasΦ.
[0113] In an embodiment, the non-human mammalian cell is from a mammal used in scientific research. In an embodiment, the non-human mammalian cell is a rodent cell. In an embodiment, the rodent cell is a rat cell or a mouse cell.
[0114] In an embodiment, a safe harbor site is any site in the genome that can accommodate the integration of new genetic material so that the new genetic element functions predictably and does not cause changes in the host genome that pose a danger to the host cell or organism. In an embodiment, the safe harbor site is a mouse safe harbor site. In an embodiment, the safe harbor site is a rat safe harbor site. In an embodiment, the safe harbor site comprises the Rosa26 site on chromosome 6. In an embodiment, the safe harbor site comprises the Hipp11 site on chromosome 11 of the mouse genome.
[0115] In an embodiment, the non-human mammalian cell is a pluripotent cell. In an embodiment, the pluripotent cell is a non-human zygote. In an embodiment, the pluripotent cell is a mouse zygote. In an embodiment, the pluripotent cell is a rat zygote. In an embodiment, the pluripotent cell is a non-human embryonic stem (ES) cell. In an embodiment, the pluripotent cell is a mouse embryonic stem (ES) cell or a rat embryonic stem (ES) cell.
[0116] In embodiments, the zygote is injected with a nucleic acid construct as described herein. In embodiments, the nucleic acid construct is injected into the pronucleus of the zygote. In embodiments, the microinjected zygote is implanted into the oviduct of a pseudo-pregnant female rodent. In embodiments, the pseudo-pregnant female rodent is a mouse. In embodiments, the pseudo-pregnant female rodent is a rat. In embodiments, the implanted zygote develops into a fetus and is born to provide a genetically modified non-human mammal. In embodiments, the genetically modified non-human mammal is a mouse. In embodiments, the genetically modified non-human mammal is a rat.
[0117] In embodiments, the method of generating a genetically modified non-human mammalian cell further comprises isolating the genetically modified non-human mammalian cell into which the nucleic acid construct has been integrated at a safe harbor site.
[0118] In embodiments, also provided herein is a genetically modified non-human mammalian cell produced by the above-described method for producing a genetically modified non-human mammalian cell. In embodiments, the non-human mammal is a rodent. In embodiments, the non-human mammal is a mouse or a rat.
[0119] In an embodiment of the method of generating a genetically modified non-human mammalian cell, the method further comprises generating a transgenic non-human mammal. In an embodiment, the method further comprises injecting the isolated cell into a blastocyst to generate a transgenic non-human mammal comprising the nucleic acid construct integrated into the safe harbor site.
[0120] In embodiments, the disclosure provides a genetically modified non-human transgenic mammal produced by the method. In embodiments, the transgenic mammal is a rodent. In embodiments, the rodent is a rat or a mouse.
[0121] In an embodiment of the method for generating a transgenic non-human mammal, the method further comprises mating the transgenic non-human mammal comprising the nucleic acid construct integrated into the safe harbor site with a transgenic non-human mammal expressing Cre recombinase to obtain a non-human mammal having cells expressing a fusion protein comprising the affinity tag, the transmembrane domain, and the fluorescent reporter protein.
[0122] In an embodiment of this method, the transgenic non-human mammal comprising a nucleic acid construct integrated at a safe harbor site is a mouse comprising a nucleic acid construct integrated at a Rosa26 site, and the transgenic non-human mammal expressing Cre recombinase is a mouse.
[0123] In an embodiment of this method, the transgenic non-human mammal comprising a nucleic acid construct integrated at a safe harbor site is a mouse comprising a nucleic acid construct integrated at a Hipp11 site, and the transgenic non-human mammal expressing Cre recombinase is a mouse.
[0124] In an embodiment, the transgenic non-human mammal expressing Cre recombinase expresses Cre recombinase under the control of a tissue-specific promoter. In an embodiment, the transgenic non-human mammal expressing Cre recombinase expresses Cre recombinase under the control of a promoter that is activated only at a specific time during cell development.
[0125] In an embodiment, the transgenic non-human mammal is a Cre switch strain mouse, for example, a Cre switch strain found in the Mouse Genome Informatics database: www.informatics.jax.org / home / recombinase. In an embodiment, the Cre switch strain mouse is a Blimp1-Cre ERT2 As a non-limiting example, the Blimp1-Cre mouse strain ERT2 can be used to cross with the above-mentioned genetically modified non-human transgenic mammals to label plasmablasts and plasma cells expressing the Blimp1 transcription factor. In an embodiment, the Cre switch strain mouse is creERT2 In an embodiment, the Jchain mouse strain is creERT2 The mouse can be crossed with the genetically modified non-human transgenic mammal described herein to more specifically label plasma cells, including all immunoglobulin isotypes.In an embodiment, the Cre switch line mouse is the Xbp1 mouse line.In an embodiment, the Cre switch line mouse is the Irf4 mouse line.
[0126] In embodiments, Cre expression in the transgenic mice is tissue specific. In embodiments, Cre expression in the transgenic mice is specific for the developmental state of the cells.
[0127] In an embodiment, a method for generating a transgenic non-human mammal, comprising mating a transgenic non-human mammal containing a nucleic acid construct integrated into a safe harbor site with a transgenic non-human mammal expressing a Cre recombinase enzyme to obtain a non-human mammal having cells expressing a fusion protein containing an affinity tag, a transmembrane domain, and a fluorescent reporter protein, is performed as represented by the schematic diagram of FIG. 1B. The transgenic non-human mammal generated by this method functions as shown in FIG. 1B. In cells in which the tissue-specific promoter is not active, the stop sequence remains in the construct, so the transgene is silent. When the tissue-specific promoter is expressed, the expressed Cre excises the stop sequence and the transgene is expressed.
[0128] In embodiments, provided herein is a genetically modified non-human mammal having cells expressing a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein produced by the methods described above.
[0129] J. Conditional Reporter Engineered Cells In embodiments, provided herein is a genetically modified non-human mammalian cell comprising a genome comprising a conditional reporter nucleic acid construct described herein integrated into a safe harbor site. In embodiments of the cell, the safe harbor site comprises the Rosa26 site on chromosome 26 of the mouse genome. In embodiments of the cell, the safe harbor site comprises the Hipp11 site on chromosome 11 of the mouse genome.
[0130] In embodiments, the cell is a hybridoma. In embodiments, the cell is a stem cell. In embodiments, the stem cell is an embryonic stem cell. In embodiments, the stem cell is an adult stem cell. In embodiments, the stem cell is an induced pluripotent stem cell. In embodiments, the stem cell is a perinatal stem cell. In embodiments, the cell is an immortalized cell.
[0131] In embodiments, the genetically modified non-human mammalian cell expresses a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein. In embodiments, the affinity tag is expressed on the cell surface of the non-human mammalian cell. Examples of affinity tags that can be expressed on the cell surface of the non-human mammalian cell are described herein.
[0132] In embodiments, the affinity tag comprises a StrepII, hexahistidine, FLAG, HA, Myc, VA, GST, β-GAL, MBP, or VSV-G tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of the tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of the tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of the tag. In embodiments, the affinity tag comprises 3 tandem repeats of the tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of the tag.
[0133] In an embodiment, the affinity tag comprises a StrepII tag. In an embodiment, the affinity tag comprises tandem repeats of a StrepII tag. In an embodiment, the affinity tag comprises about 1 to about 18 tandem repeats of a StrepII tag. In an embodiment, the affinity tag comprises about 2 to about 15 tandem repeats of a StrepII tag. In an embodiment, the affinity tag comprises about 3 to about 10 tandem repeats of a StrepII tag. In an embodiment, the affinity tag comprises 3 tandem repeats of a StrepII tag. In an embodiment, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of a StrepII tag. In an embodiment, the tandem repeat has a tag linker between the repeats. In an embodiment, the linker is a dipeptide or a tripeptide. In an embodiment, the tag linker is the dipeptide Ser-Ala.
[0134] In an embodiment, the StrepII tag comprises an eight amino acid peptide sequence of WSHPQFEK (SEQ ID NO: 1). In an embodiment, the affinity tag comprises WSHPQFEKSAWSHPQFEKSAWSHPQFEK (SEQ ID NO: 3).
[0135] In an embodiment, the fluorescent reporter protein is exposed on the cytoplasmic face of the non-human mammalian cell.
[0136] In an embodiment, the fluorescent reporter protein comprises a green fluorescent protein, a yellow fluorescent protein, a cyan fluorescent protein, a red fluorescent protein, a blue fluorescent protein, a red fluorescent protein, or an orange fluorescent protein. In an embodiment, the fluorescent reporter protein comprises a green fluorescent protein (GFP), an enhanced green fluorescent protein (EGFP), an enhanced yellow fluorescent protein (EYFP), or an enhanced cyan fluorescent protein (ECFP). In an embodiment, the fluorescent reporter protein comprises a red fluorescent protein (RFP). In an embodiment, the red fluorescent protein is a monomeric cherry (mCherry) or a tandem dimer tomato (tdTomato).
[0137] K. Methods for Isolating Conditional Reporter Modified Cells In embodiments, provided herein is a method of isolating a cell obtained from a genetically modified non-human mammal, the method comprising: (a) obtaining a cell from a genetically modified conditional reporter non-human mammal as described herein; (b) screening the cells obtained from the genetically modified non-human mammal for expression of a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein; and (c) isolating cells that express the fusion protein.
[0138] In an embodiment of the method of isolating cells, the cells are sorted by fluorescence activated cell sorting (FACS) or magnetic activated cell sorting (MACS). In an embodiment, in the method of isolating cells, the cells are screened by fluorescence activated cell sorting (FACS). In an embodiment, in the method of isolating cells, the cells are sorted by magnetic activated cell sorting (MACS). Both FACS and MACS techniques are known in the art and described elsewhere herein. In an embodiment, separation of cells using either FACS or MACS is shown diagrammatically in FIG. 1C. In an embodiment of the method of isolating cells, the cells are separated using affinity tags expressed on the surface of the cells as described herein. In an embodiment of separating cells using affinity tags, the cells are separated using affinity columns or affinity resins that bind affinity tags using methods known in the art.
[0139] In embodiments, the affinity tag is expressed on the cell surface of the genetically modified non-human mammalian cell.
[0140] In embodiments, the affinity tag comprises a StrepII, hexahistidine, FLAG, HA, Myc, VA, GST, β-GAL, MBP, or VSV-G tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of the tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of the tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of the tag. In embodiments, the affinity tag comprises 3 tandem repeats of the tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of the tag.
[0141] In embodiments, the affinity tag comprises a StrepII tag. In embodiments, the affinity tag comprises tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises 3 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of a StrepII tag. In an embodiment, the tandem repeat has a tag linker between the repeats. In an embodiment, the linker is a dipeptide or a tripeptide. In an embodiment, the tag linker is the dipeptide Ser-Ala.
[0142] In an embodiment, the StrepII tag comprises an eight amino acid peptide sequence of WSHPQFEK (SEQ ID NO: 1). In an embodiment, the affinity tag comprises WSHPQFEKSAWSHPQFEKSAWSHPQFEK (SEQ ID NO: 3).
[0143] In an embodiment, the fluorescent reporter protein is exposed on the cytoplasmic face of the non-human mammalian cell.
[0144] In an embodiment, the fluorescent reporter protein comprises a green fluorescent protein, a yellow fluorescent protein, a cyan fluorescent protein, a red fluorescent protein, a blue fluorescent protein, a red fluorescent protein, or an orange fluorescent protein. In an embodiment, the fluorescent reporter protein comprises a green fluorescent protein (GFP), an enhanced green fluorescent protein (EGFP), an enhanced yellow fluorescent protein (EYFP), or an enhanced cyan fluorescent protein (ECFP). In an embodiment, the fluorescent reporter protein comprises a red fluorescent protein (RFP). In an embodiment, the red fluorescent protein is a monomeric cherry (mCherry) or a tandem dimer tomato (tdTomato).
[0145] L. Immunoglobulin Reporter Nucleic Acid Constructs In embodiments, provided herein is a nucleic acid construct comprising a linker, a leader sequence, and a transmembrane reporter cassette encoding an affinity tag, a transmembrane domain, and a fluorescent reporter, an embodiment of which may be referred to herein as an "immunoglobulin reporter nucleic acid construct."
[0146] In embodiments, the nucleic acid construct comprises a single-stranded DNA, a double-stranded DNA, a plasmid, or a viral vector. In embodiments, the nucleic acid construct is a linear DNA. In embodiments, the nucleic acid construct is a circular DNA.
[0147] In an embodiment, the nucleic acid construct further comprises a first homology arm and a second homology arm that are homologous to a first target sequence and a second target sequence in the genome of the non-human mammal. The homology region allows the nucleic acid construct to be integrated into the genome of the non-human mammal using methods described herein and known in the art. In an embodiment, the nucleic acid construct further comprises a first homology arm and a second homology arm that are homologous to a first target sequence and a second target sequence, respectively, in an immunoglobulin site in the non-human mammal, such as an immunoglobulin variable region site or an immunoglobulin constant region site, or a site therebetween. In an embodiment, the nucleic acid construct further comprises a first homology arm and a second homology arm that are homologous to a first target sequence and a second target sequence, respectively, in an immunoglobulin constant region site in the non-human mammal.
[0148] In an embodiment, a first and a second homology arm are homologous to a first and a second target sequence, respectively, and the first and the second target sequence are adjacent to an immunoglobulin constant region site. In an embodiment, the immunoglobulin constant region site is an immunoglobulin light chain constant region site. In an embodiment, the immunoglobulin light chain constant region site is a kappa light chain constant region site. In an embodiment, the immunoglobulin light chain constant region site is a lambda light chain constant region site. In an embodiment, the immunoglobulin constant region site is an immunoglobulin heavy chain constant region site. In an embodiment, the immunoglobulin heavy chain constant region site is a gamma, delta, alpha, mu, or epsilon immunoglobulin constant region site.
[0149] In an embodiment, the first target sequence is upstream of the immunoglobulin constant region site and the second target sequence is downstream of the stop codon of the immunoglobulin constant region site. In an embodiment, the immunoglobulin constant region site is an immunoglobulin light chain constant region site. In an embodiment, the immunoglobulin light chain constant region site is an immunoglobulin kappa constant region site. In an embodiment, the immunoglobulin light chain constant region site is an immunoglobulin lambda constant region site.
[0150] In embodiments, the immunoglobulin constant region moiety is an immunoglobulin heavy chain constant region moiety. In embodiments, the immunoglobulin heavy chain constant region moiety is a gamma, delta, alpha, mu, or epsilon immunoglobulin constant region moiety.
[0151] In an embodiment, the first homology arm and the second homology arm each independently comprise about 15 nucleotides to about 12000 nucleotides. In an embodiment, the first homology arm and the second homology arm each independently comprise about 30 nucleotides to about 11000 nucleotides. In an embodiment, the first homology arm and the second homology arm each independently comprise about 50 nucleotides to about 10000 nucleotides. In an embodiment, the first homology arm and the second homology arm each independently comprise about 100 nucleotides to about 7500 nucleotides. In an embodiment, the first homology arm and the second homology arm each independently comprise about 200 nucleotides to about 5000 nucleotides. In an embodiment, the first homology arm and the second homology arm each independently comprise about 300 nucleotides to about 2500 nucleotides.
[0152] In some embodiments, the linker comprises a stop codon and an internal ribosome entry site (IRES). In embodiments, the linker comprises a protease recognition site and a self-cleaving peptide. In embodiments, the linker comprises a leaky stop codon (LSC) with a peptide linker, a protease recognition site, and a self-cleaving peptide.
[0153] Embodiments of protease recognition sites are described herein. In embodiments, the protease recognition site comprises a Furin protease recognition site. In embodiments, the Furin protease recognition site comprises a nucleic acid sequence encoding a peptide of Arg-X-Arg-Arg. In embodiments, X is a hydrophobic amino acid. In embodiments, X is a hydrophilic amino acid. In embodiments, X is lysine. In embodiments, orthodox Furin protease recognition sites comprise a nucleic acid sequence encoding a peptide X-Arg-X-Lys-Arg-X or X-Arg-X-Arg-Arg-X. In embodiments, X is a hydrophobic amino acid. In embodiments, the hydrophobic amino acid comprises Gly, Ala, Ile, Leu, Met, Val, Phe, Trp, or Tyr. In embodiments, X is a hydrophilic amino acid. In embodiments, the hydrophilic amino acid is lysine. In embodiments, the Furin protease recognition site comprises a nucleic acid sequence encoding a peptide Arg-Lys-Arg-Arg. In embodiments, the Furin protease recognition site comprises a nucleic acid sequence encoding the peptide Arg-Arg-Arg-Arg. In embodiments, the Furin protease recognition site comprises a nucleic acid sequence encoding the peptide Arg-Arg-Lys-Arg. In embodiments, the Furin protease recognition site comprises a nucleic acid sequence encoding the peptide Arg-Lys-Lys-Arg. In embodiments, the Lys residue immediately preceding the Furin protease site is deleted. In embodiments, the Furin protease recognition site is a Furin protease recognition site as described in Fang et al., Molecular Therapy 15(6):1153-1159 (2007), which is incorporated herein by reference.
[0154] In embodiments, the protease is an endoprotease. In embodiments, the protease is a mammalian endoprotease. In embodiments, the protease is an endoprotease endogenously expressed in a cell comprising the nucleic acid construct. In embodiments, the protease recognition site comprises a trypsin, chymotrypsin, elastase, thermolysin, pepsin, glutamyl endopeptidase, or neprilysin recognition site.
[0155] An embodiment of a self-cleaving peptide is described herein. In an embodiment, the self-cleaving peptide comprises a 2A self-cleaving peptide. In an embodiment, the self-cleaving peptide comprises a T2A (EGRGSLLTCGDVEENPGP; SEQ ID NO: 4), a P2A (ATNFSLLKQAGDVEENPGP; SEQ ID NO: 5), an E2A (QCTNYALLKLAGDVESNPGP; SEQ ID NO: 6), or an F2A (VKQTLNFDLLKLAGDVESNPGP; SEQ ID NO: 7) self-cleaving peptide.
[0156] Embodiments of leaky stop codons are described herein. In an embodiment, the sequence encoding the leaky stop codon comprises TGACTAG. In an embodiment, the sequence encoding the leaky stop codon comprises TGACGG. In an embodiment, the sequence encoding the leaky stop codon comprises TAGCAATTA. In an embodiment, the sequence encoding the leaky stop codon comprises TAGCAATCA. In an embodiment, the sequence encoding the leaky stop codon comprises TGACTA.
[0157] In embodiments where the linker comprises a leaky stop codon, the leaky stop codon allows some read-through of the codon to express the transmembrane reporter cassette. In embodiments, read-through transcription of the leaky codon occurs at a rate of about 5%. In embodiments, read-through transcription of the leaky codon occurs at a rate of about 1% to about 10%. In embodiments, when read-through transcription does not occur, the immunoglobulin is expressed in an endogenous format and the transmembrane reporter cassette is not expressed.
[0158] In embodiments, the linker is a peptide linker, e.g., an amino acid chain 2 to 24 residues in length. In embodiments, the peptide linker is a dipeptide linker. In embodiments, the linker is a tripeptide linker. In embodiments, the linker is 4 amino acids in length. In embodiments, the linker comprises Leu-Gly. In embodiments, the linker comprises Gly-Ser-Gly. In embodiments, the linker comprises Leu-Gly-Ser-Gly. In embodiments, the linker comprises about 4, about 5, about 6, about 7, about 8, about 9, or about 10 amino acid residues. In embodiments, the peptide linker comprises 4 to 24 amino acid residues. In embodiments, the peptide linker comprises 5 to 20 amino acid residues. In embodiments, the peptide linker comprises 7 to 15 amino acid residues.
[0159] In an embodiment, the leader sequence further comprises a secretory signal peptide. In an embodiment, the secretory signal peptide is an IL-2 leader sequence. In an embodiment, the secretory signal peptide is a human OSM, VSV-G mouse Ig kappa, human IgG2H, BM40, Secreto, human IgKVIII, CD33, tPA, human chymotrypsinogen, human trypsinogen-2, human IL-2, or human serum albumin signal peptide. In an embodiment, the secretory signal peptide comprises an IL-2 leader sequence MYRMQLLSCIALSLALVTNS (SEQ ID NO: 2). Those skilled in the art will understand that signal peptides can be predicted using algorithms known in the art, such as SignalP-5.0 (www.cbs.dtu.dk / services / SignalP / ) and SecretomeP 2.0 (www.cbs.dtu.dk / services / SecretomeP / ).
[0160] The affinity tag of the nucleic acid construct can be used to subsequently isolate or purify the protein expressed by the construct. In embodiments, the affinity tag comprises a StrepII, hexahistidine, FLAG, HA, Myc, VA, GST, β-GAL, MBP, or VSV-G tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of the tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of the tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of the tag. In embodiments, the affinity tag comprises 3 tandem repeats of the tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of the tag.
[0161] In embodiments, the affinity tag comprises a StrepII tag. In embodiments, the affinity tag comprises tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises 3 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of a StrepII tag. In an embodiment, the tandem repeat has a tag linker between the repeats. In an embodiment, the tag linker is a dipeptide or a tripeptide. In an embodiment, the tag linker is a dipeptide Ser-Ala. In an embodiment, the tag linker comprises a (G4S)2 linker (GGGGSGGGGS; SEQ ID NO: 8). In an embodiment, the tag linker is a (G4S)2 linker (GGGGSGGGGS; SEQ ID NO: 8).
[0162] In an embodiment, the StrepII tag comprises an eight amino acid peptide sequence of WSHPQFEK (SEQ ID NO: 1). In an embodiment, the affinity tag comprises WSHPQFEKSAWSHPQFEKSAWSHPQFEK (SEQ ID NO: 3).
[0163] The transmembrane domain encoded by the transmembrane reporter cassette allows the affinity tag to be displayed on the surface of a cell expressing the nucleic acid construct. In an embodiment, the transmembrane domain comprises a hydrophobic alpha helix. In an embodiment, the transmembrane domain is an IgG transmembrane domain. In an embodiment, the transmembrane domain is a human IgG1 transmembrane domain. In an embodiment, the transmembrane domain comprises a mouse IgG transmembrane domain. In an embodiment, the transmembrane domain comprises a mouse IgG1, IgG2a, IgG2b, or IgG2c transmembrane domain. In an embodiment, the transmembrane domain is the transmembrane domain of mouse protein Tmem53, Lrtm1, or Nrg1.
[0164] The fluorescent reporter protein allows detection of cells expressing the nucleic acid construct. In an embodiment, the fluorescent reporter protein comprises green fluorescent protein, yellow fluorescent protein, cyan fluorescent protein, red fluorescent protein, blue fluorescent protein, red fluorescent protein, or orange fluorescent protein. In an embodiment, the fluorescent reporter protein comprises green fluorescent protein (GFP), enhanced green fluorescent protein (EGFP), enhanced yellow fluorescent protein (EYFP), or enhanced cyan fluorescent protein (ECFP). In an embodiment, the fluorescent reporter protein comprises red fluorescent protein (RFP). In an embodiment, the red fluorescent protein is monomeric cherry (mCherry) or tandem dimeric tomato (tdTomato).
[0165] In embodiments, the nucleic acid construct is the nucleic acid construct shown diagrammatically in Figure 2 for a particular embodiment incorporated into a light chain kappa constant region, where the black rectangle represents the V and J segments of the region, LK represents the linker sequence, L represents the leader sequence, STX3 represents three tandem repeats of the Strep-II tag, TM represents the transmembrane domain, and GFP represents the green fluorescent protein reporter. In other embodiments (not shown), the nucleic acid construct is incorporated into a light chain lambda constant region or a heavy chain constant region.
[0166] M. Methods for generating immunoglobulin reporter modified cells and organisms In embodiments, provided herein is a method of generating a genetically modified non-human mammalian cell, the method comprising: (a) introducing an immunoglobulin reporter nucleic acid construct described herein into a non-human mammalian cell; and (b) introducing a nuclease into the non-human mammalian cell, wherein the nuclease causes a single-stranded or double-stranded break at an immunoglobulin constant region site in the genome of the non-human mammalian cell, and the nucleic acid construct is integrated into the immunoglobulin constant region site in the genome of the non-human mammalian cell by homologous recombination.
[0167] In embodiments, the nuclease causes a double stranded break. In embodiments, the nuclease causes a single stranded break (e.g., a nick).
[0168] The nuclease can be introduced into the cell using methods known in the art. In embodiments, introducing the nuclease comprises introducing an expression construct encoding the nuclease. In embodiments, the introduction of the expression construct is by injection or electroporation. In embodiments, introducing the nuclease comprises introducing a plasmid encoding the nuclease. In embodiments, introducing the nuclease comprises introducing a viral vector encoding the nuclease. In embodiments, introducing the nuclease comprises introducing an mRNA encoding the nuclease. In embodiments, the mRNA comprises one or more modified bases. In embodiments, the mRNA is encapsulated in a lipid nanoparticle. In embodiments, the introduction of the plasmid, viral vector or mRNA is by injection or electroporation. In embodiments, introducing the nuclease comprises directly introducing the nuclease protein into the cell. In embodiments, the direct introduction of the nuclease protein into the cell is by injection or electroporation.
[0169] In embodiments, the immunoglobulin constant region site is an immunoglobulin light chain constant region site. In embodiments, the immunoglobulin light chain constant region site is an immunoglobulin kappa constant region site. In embodiments, the immunoglobulin light chain constant region site is an immunoglobulin lambda constant region site. In some aspects, the immunoglobulin constant region site is an immunoglobulin heavy chain constant region site.
[0170] In embodiments, the immunoglobulin constant region moiety is an immunoglobulin heavy chain constant region moiety. In embodiments, the immunoglobulin heavy chain constant region moiety is a gamma, delta, alpha, mu, or epsilon immunoglobulin constant region moiety.
[0171] In an embodiment, the nuclease is a nuclease as described herein. In an embodiment, the nuclease comprises a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, or a clustered regularly interspaced short palindromic repeats (CRISPR) associated (Cas) protein and a guide RNA (gRNA). In an embodiment, the gRNA comprises a CRISPR RNA (crRNA) that targets a recognition site and a transactivating CRISPR RNA (tracrRNA). In an embodiment, the CRISPR-Cas protein comprises Cas9. In an embodiment, the Cas protein comprises Cas9, Cas12, Cas12a, Cas13, Cas14, or CasΦ. In an embodiment, the CRISPR-Cas protein comprises Cas3, Cas8, Cas10, Cas11, Cas12, Cas12a, Cas13, Cas14, or CasΦ.
[0172] In an embodiment, the non-human mammalian cell is from a mammal used in scientific research. In an embodiment, the non-human mammalian cell is a rodent cell. In an embodiment, the rodent cell is a rat cell or a mouse cell.
[0173] In an embodiment, the non-human mammalian cell is a pluripotent cell. In an embodiment, the pluripotent cell is a non-human zygote. In an embodiment, the pluripotent cell is a mouse zygote. In an embodiment, the pluripotent cell is a rat zygote. In an embodiment, the pluripotent cell is a non-human embryonic stem (ES) cell. In an embodiment, the pluripotent cell is a mouse embryonic stem (ES) cell or a rat embryonic stem (ES) cell.
[0174] In embodiments, the zygote is injected with a nucleic acid construct as described herein. In embodiments, the nucleic acid construct is injected into the pronucleus of the zygote. In embodiments, the microinjected zygote is implanted into the oviduct of a pseudo-pregnant female rodent. In embodiments, the pseudo-pregnant female rodent is a mouse. In embodiments, the pseudo-pregnant female rodent is a rat. In embodiments, the implanted zygote develops into a fetus and is born to provide a genetically modified non-human mammal. In embodiments, the genetically modified non-human mammal is a mouse. In embodiments, the genetically modified non-human mammal is a rat.
[0175] In embodiments, the method of generating a genetically modified non-human mammalian cell further comprises isolating the genetically modified non-human mammalian cell into which the nucleic acid construct has been integrated at an immunoglobulin constant region site.
[0176] In embodiments, also provided herein is a genetically modified non-human mammalian cell produced by the above-described method of producing a genetically modified non-human mammalian cell. In embodiments, the non-human mammal is a rodent.
[0177] In an embodiment of the method of generating a genetically modified non-human mammalian cell, the method further comprises generating a transgenic non-human mammal. In an embodiment, the method further comprises injecting the gene editing material into a zygote or the modified isolated cell into a blastocyst to generate a transgenic non-human mammal comprising the nucleic acid construct integrated at an immunoglobulin constant region site.
[0178] In embodiments, the disclosure provides a genetically modified non-human transgenic mammal produced by the method. In embodiments, the transgenic mammal is a rodent. In embodiments, the rodent is a rat or a mouse.
[0179] N. Immunoglobulin reporter engineered cells In embodiments, provided herein is a genetically modified non-human mammalian cell comprising a genome that includes an immunoglobulin reporter nucleic acid construct described herein integrated into an immunoglobulin constant region site.
[0180] In embodiments of the cell, the immunoglobulin constant region site is an immunoglobulin light chain constant region site. In embodiments, the immunoglobulin light chain constant region site is an immunoglobulin kappa constant region site. In embodiments, the immunoglobulin light chain constant region site is an immunoglobulin lambda constant region site. In embodiments, the immunoglobulin constant region site is an immunoglobulin heavy chain constant region site.
[0181] In embodiments of the cell, the immunoglobulin constant region locus is an immunoglobulin heavy chain constant region locus. In embodiments, the immunoglobulin heavy chain constant region locus is a gamma, delta, alpha, mu, or epsilon immunoglobulin constant region locus.
[0182] In embodiments, the immunoglobulin expressing cells are obtained from an immunized mammal. In embodiments, the immunized mammal is a rodent. In embodiments, the immunized mammal is a mouse or a rat.
[0183] In embodiments, the cell is an immunoglobulin expressing cell. In embodiments, the immunoglobulin expressing cell is an immature B cell or a progeny of an immature B cell. In embodiments, the cell is a hybridoma, a stem cell, or an immortalized cell. In embodiments, the stem cell is an embryonic stem cell. In embodiments, the stem cell is an adult stem cell. In embodiments, the stem cell is an induced pluripotent stem cell. In embodiments, the stem cell is a perinatal stem cell.
[0184] In embodiments, the genetically modified non-human mammalian cells express an immunoglobulin kappa light chain. In embodiments, the genetically modified non-human mammalian cells express an immunoglobulin lambda light chain. In embodiments, the genetically modified non-human mammalian cells express an immunoglobulin heavy chain.
[0185] In embodiments, the genetically modified non-human mammalian cell expresses a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein. In embodiments, the affinity tag is expressed on the cell surface of the non-human mammalian cell. Examples of affinity tags that can be expressed on the cell surface of the non-human mammalian cell are described herein.
[0186] In embodiments, the affinity tag comprises a StrepII, hexahistidine, FLAG, HA, Myc, VA, GST, β-GAL, MBP, or VSV-G tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of the tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of the tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of the tag. In embodiments, the affinity tag comprises 3 tandem repeats of the tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of the tag.
[0187] In embodiments, the affinity tag comprises a StrepII tag. In embodiments, the affinity tag comprises tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises 3 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of a StrepII tag. In an embodiment, the tandem repeat has a tag linker between the repeats. In an embodiment, the tag linker is a dipeptide or a tripeptide. In an embodiment, the tag linker is the dipeptide Ser-Ala.
[0188] In an embodiment, the StrepII tag comprises an eight amino acid peptide sequence of WSHPQFEK (SEQ ID NO: 1). In an embodiment, the affinity tag comprises WSHPQFEKSAWSHPQFEKSAWSHPQFEK (SEQ ID NO: 3).
[0189] In an embodiment, the fluorescent reporter protein is exposed on the cytoplasmic face of the non-human mammalian cell.
[0190] In an embodiment, the fluorescent reporter protein comprises a green fluorescent protein, a yellow fluorescent protein, a cyan fluorescent protein, a red fluorescent protein, a blue fluorescent protein, a red fluorescent protein, or an orange fluorescent protein. In an embodiment, the fluorescent reporter protein comprises a green fluorescent protein (GFP), an enhanced green fluorescent protein (EGFP), an enhanced yellow fluorescent protein (EYFP), or an enhanced cyan fluorescent protein (ECFP). In an embodiment, the fluorescent reporter protein comprises a red fluorescent protein (RFP). In an embodiment, the red fluorescent protein is a monomeric cherry (mCherry) or a tandem dimer tomato (tdTomato).
[0191] In embodiments of the cell, expression of the fusion protein is driven by an endogenous immunoglobulin transcriptional regulator. In embodiments, the endogenous immunoglobulin transcriptional regulator is an endogenous immunoglobulin light chain transcriptional regulator. In embodiments, the endogenous immunoglobulin light chain transcriptional regulator comprises a promoter and other cis elements in the mouse light chain locus. In embodiments, the endogenous immunoglobulin transcriptional regulator is an endogenous immunoglobulin heavy chain transcriptional regulator. In embodiments, the endogenous immunoglobulin heavy chain transcriptional regulator comprises a promoter and other cis elements in the mouse heavy chain locus.
[0192] O. Methods for Identifying Immunoglobulin Reporter Modified Cells In embodiments, provided herein is a method for identifying an immunoglobulin-expressing cell obtained from a genetically modified immunoglobulin reporter non-human mammal, the method comprising: (a) obtaining cells from a genetically modified immunoglobulin reporter non-human mammal as described herein; (b) screening the cells obtained from the genetically modified non-human mammal for expression of a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein; and (c) identifying immunoglobulin-expressing cells based on expression of the fusion protein;
[0193] In an embodiment, in a method of isolating cells, cells are sorted by fluorescence activated cell sorting (FACS) or magnetic activated cell sorting (MACS). In an embodiment, in a method of isolating cells, cells are screened by fluorescence activated cell sorting (FACS). In an embodiment, in a method of isolating cells, cells are sorted by magnetic activated cell sorting (MACS). Both FACS and MACS techniques are known in the art and described elsewhere herein. In an embodiment, an example of a process for obtaining cells from a rodent modified with an immunoglobulin reporter, pooling the cells, and separating the cells using either FACS or MACS is shown diagrammatically in FIG. 3. In an embodiment of a method of isolating cells, the cells are separated using an affinity tag expressed on the surface of the cells as described herein. In an embodiment of separating cells using an affinity tag, the cells are separated using an affinity column or affinity resin that binds the affinity tag using methods known in the art.
[0194] In embodiments, the affinity tag is expressed on the cell surface of the genetically modified non-human mammalian cell.
[0195] In embodiments, the affinity tag comprises a StrepII, hexahistidine, FLAG, HA, Myc, VA, GST, β-GAL, MBP, or VSV-G tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of the tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of the tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of the tag. In embodiments, the affinity tag comprises 3 tandem repeats of the tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of the tag.
[0196] In embodiments, the affinity tag comprises a StrepII tag. In embodiments, the affinity tag comprises tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 1 to about 18 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 2 to about 15 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 3 to about 10 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises 3 tandem repeats of a StrepII tag. In embodiments, the affinity tag comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 tandem repeats of a StrepII tag. In an embodiment, the tandem repeat has a tag linker between the repeats. In an embodiment, the tag linker is a dipeptide or a tripeptide. In an embodiment, the tag linker is the dipeptide Ser-Ala.
[0197] In an embodiment, the StrepII tag comprises an eight amino acid peptide sequence of WSHPQFEK (SEQ ID NO: 1). In an embodiment, the affinity tag comprises WSHPQFEKSAWSHPQFEKSAWSHPQFEK (SEQ ID NO: 3).
[0198] In an embodiment, the fluorescent reporter protein is exposed on the cytoplasmic face of the non-human mammalian cell.
[0199] In an embodiment, the fluorescent reporter protein comprises a green fluorescent protein, a yellow fluorescent protein, a cyan fluorescent protein, a red fluorescent protein, a blue fluorescent protein, a red fluorescent protein, or an orange fluorescent protein. In an embodiment, the fluorescent reporter protein comprises a green fluorescent protein (GFP), an enhanced green fluorescent protein (EGFP), an enhanced yellow fluorescent protein (EYFP), or an enhanced cyan fluorescent protein (ECFP). In an embodiment, the fluorescent reporter protein comprises a red fluorescent protein (RFP). In an embodiment, the red fluorescent protein is a monomeric cherry (mCherry) or a tandem dimer tomato (tdTomato).
[0200] In an embodiment of the method, the genetically modified non-human mammal has been immunized with an antigen of interest. In an embodiment, the immunoglobulin-expressing cells express an immunoglobulin kappa light chain. In an embodiment, the immunoglobulin-expressing cells express an immunoglobulin lambda light chain. In an embodiment, the immunoglobulin-expressing cells express an immunoglobulin heavy chain. In an embodiment, the immunoglobulin-expressing cells include immature B cells and their progeny.
[0201] In embodiments, the method further comprises isolating the expressed immunoglobulin from the cell obtained from the genetically modified non-human mammal. In embodiments, provided herein is an immunoglobulin obtained by the method.
[0202] In embodiments, provided herein is a method of producing a therapeutic or diagnostic immunoglobulin, the method comprising: (i) cloning the variable regions of the immunoglobulins disclosed herein; and (ii) producing a therapeutic or diagnostic immunoglobulin comprising the variable region obtained in (i);
[0203] In embodiments, there is also provided herein a method for producing a monoclonal antibody, the method comprising: (i) obtaining immunoglobulin-expressing cells from a genetically modified non-human mammal as disclosed herein; (ii) immortalizing the immunoglobulin-expressing cells obtained in (i); and (iii) isolating the immortalized immunoglobulin-expressing cells or the monoclonal antibodies expressed by the nucleic acid sequences encoding the monoclonal antibodies. In an embodiment, the method further comprises: (iv) cloning the variable regions of the isolated monoclonal antibodies; and (v) producing therapeutic or diagnostic antibodies containing the cloned variable regions.
[0204] In embodiments, provided herein are therapeutic or diagnostic antibodies produced by the above methods.
[0205] P. Incorporation by Reference All references cited herein, including patents, patent applications, articles, textbooks, and the like, and the references cited therein, to the extent not otherwise disclosed, are hereby incorporated by reference in their entirety. EXAMPLES
[0206] Example 1. Construction of immunoglobulin tagging cassette and targeting vector Step 1: A transmembrane labeling cassette is assembled by ligating constituent sequences to form a contiguous cassette. In this example, the labeling cassette includes a linker, a leader sequence, three repeats of a Strep-II tag with a tag linker (WSHPQFEKSAWSHPQFEKSAWSHPQFEK (SEQ ID NO: 3)), a transmembrane domain, and a sequence encoding green fluorescent protein (GFP), as shown diagrammatically in FIG.
[0207] There are three different linker options: Option 1: LSL-Furin-2A. This linker contains a leaky stop codon, a Leu-Gly linker sequence, a Furin protease recognition and cleavage site, and a 2A self-cleaving peptide. The advantage of this design is that it maintains the majority of IgK in the endogenous format. Because the leaky stop codon allows only 5% of the transcription to read through, the expression level of StrepII-tagged-GFP is only 5% of the total IgK. Therefore, with this design, the expression level of StrepII-tagged-GFP may be too low to efficiently enrich such cells using FACS / MACS in some cases. Option 2: Furin-2A. This linker contains the Furin protease recognition and cleavage site, and the 2A self-cleaving peptide. The advantage of this linker is that it ensures high expression of StrepII-tagged GFP, but this level of expression may be toxic to cells in some cases. Option 3: Stop codon-IRES (internal ribosome entry site). This linker contains an IRES followed by a stop codon, which allows the reporter gene to be transcribed as a separate protein from the immunoglobulin. The IRES provides a level of StrepII-tagged GFP expression intermediate between the two strategies above.
[0208] Step 2: The construct from step 1 is ligated into a targeting vector with homologous flanking regions that target the stop codon of the rodent Ig kappa (IgK) genomic region. Vectors that can be used for knock-in include pUC18, pUC19, and pBluescriptII KS+. This vector knocks in a StrepII tagged-GFP labelled cassette at the stop codon of IgK as shown diagrammatically in Figure 2.
[0209] Alternatively, a synthetic single-stranded DNA can be synthesized with 200-500 bp of homology-defining regions on either side of the tag cassette to form a synthetic targeting cassette, which can then be directly integrated using the CRISPR targeting system.
[0210] Example 2. In vitro evaluation of StrepII tag and GFP expression level To ensure that the expression level of StrepII tag-GFP is compatible with downstream applications, in vitro experiments with the targeting vector are performed using rodent B cell lines to evaluate the expression levels of StrepII tag and GFP along with the expression level of IgK. The vector that gives the highest IgK expression level and StrepII tag and GFP levels suitable for downstream applications is selected for further studies. Secreted antibodies are quantified by biochemical measurements such as octet. The antibodies displayed on the cell surface are detected by flow analysis. Briefly, cells are incubated with fluorescently labeled anti-immunoglobulin antibodies for 30 min at 4 °C, and the fluorescent signal is measured by a flow cytometer. To ensure that the expression of the labeling cassette at the IgK site does not interfere with the function of IgK, in vitro tests are performed using rodent B cell lines to identify the ideal linker sequence.
[0211] Example 3. Generation of IgK reporter rodent strains Mouse embryonic stem cells are transformed with a targeting vector, allowing the insertion of a tag cassette at the IgK stop codon by homologous recombination. To increase the efficiency of targeted knock-in at this IgK stop codon, the CRISPR / Cas9 system is used. sgRNA components that target the adjacent regions around the IgK stop codon in the genome are designed and synthesized, then associated with the Cas9 enzyme to form an RNP complex, which is then injected into fertilized eggs or embryonic stem cells as single-stranded DNA or vector together with a homology-directed repair (HDR) template. A donor fragment containing the tag cassette is integrated into the site after the targeted double-strand break caused by Cas9 by homologous recombination. Successfully recombined embryonic stem cells (determined by Southern blot analysis and PCR) are microinjected into blastocysts to generate transgenic mice.
[0212] mAb-expressing cells are obtained from transgenic mice to evaluate the labeling efficiency. Briefly, antibody-expressing cells are isolated using conventional flow markers by FACS. The cells are incubated with fluorescently labeled anti-immunoglobulin antibodies for 30 min at 4°C. The fluorescent signal is measured using a flow cytometer.
[0213] Example 4. Construction of conditional reporter labeling cassette and targeting vector Step 1: Transmembrane tagging cassette transgenes are assembled by ligating constituent sequences to form a continuous cassette. As shown diagrammatically in Figure 1A, the tagging cassette contains a CAG promoter, a leader sequence, a LoxP-Stop-LoxP cassette, three tandem repeats of a StrepII tag, a transmembrane domain, and green fluorescent protein (GFP).
[0214] Step 2: The construct from step 1 is ligated into a targeting vector with homologous flanking regions targeting intron 1 of the ROSA26 genomic region. A splice acceptor (SA) sequence and a DNA cassette are inserted into the Xba1 restriction enzyme site within the first intron of the ROSA26 gene. This vector knocks in a StrepII tag-GFP labeling cassette into the safe harbor ROSA26 site, as shown diagrammatically in Figure 1A.
[0215] Alternatively, a synthetic single-stranded DNA can be synthesized with 200-500 bp of homology-defining regions on either side of the tagging cassette to form a synthetic targeting cassette. This synthetic targeting cassette construct can be directly integrated into intron 1 of ROSA26 using the CRISPR targeting system.
[0216] Example 5. Generation of conditional reporter rodent strains Mouse embryonic stem cells are transformed with a transgene targeting vector, allowing the insertion of a tag cassette into intron 1 of ROSA26 by homologous recombination. Alternatively, the CRISPR / Cas9 system is a targeted knock-in into intron 1 of ROSA26. An sgRNA targeting intron 1 of ROSA26 in the genome is designed and synthesized, then assembled with Cas9 enzyme to form an RNP complex, and co-injected or electroporated with a homology-directed repair (HDR) template as single-stranded DNA or vector into fertilized eggs or embryonic stem cells. By homologous recombination, the donor fragment containing the tag cassette is integrated into the site after the targeted double-strand break caused by Cas9. Successfully recombined embryonic stem cells (determined by Southern blot analysis and PCR) are microinjected into blastocysts to generate transgenic mice.
[0217] Mice containing the integrated transgene are crossed with Cre-switch mice (a mouse strain carrying Cre under the control of a tissue-specific promoter of interest). In one experiment, the Cre-switch mice were ERT2 In another embodiment, the switch strain mouse is a Jchain mouse line that expresses Cre under the control of the Blimp1 promoter, which is expressed in plasmablasts and plasma cells. creERT2 It's a mouse.
[0218] Example 6. Generation of reporter rodent lines from zygotes The vector described in Example 1 or 3 is directly injected into a mouse zygote. The vector is microinjected into the pronucleus of the zygote (a fertilized mouse egg). The resulting embryo is implanted into the oviduct of a pseudopregnant female and allowed to develop. The embryo is expelled into the mouse oviduct and the wound is closed with wound clips. Mice are examined for the birth of live pups on days 18-21. Newborn mice are analyzed for expression of the transmembrane tagging construct using the methods described above.
[0219] array JPEG2024534688000002.jpg96161
Claims
1. A nucleic acid construct comprising a leader sequence, a LoxP-Stop-LoxP cassette, and a transmembrane reporter cassette encoding an affinity tag, a transmembrane (TM) domain, and a fluorescent reporter protein.
2. 10. The nucleic acid construct of claim 1, comprising single-stranded DNA, double-stranded DNA, a plasmid, or a viral vector.
3. 2. The nucleic acid construct of claim 1, further comprising a first homology arm and a second homology arm that are homologous to a first target sequence and a second target sequence, respectively, within a safe harbor site of a non-human mammal.
4. 4. The nucleic acid construct of claim 3, wherein the safe harbor site comprises the Rosa26 site on chromosome 6 of the mouse genome or the Hippl 1 site on chromosome 11 of the mouse genome.
5. 2. The nucleic acid construct of claim 1, wherein the leader sequence comprises a secretory signal peptide.
6. 6. The nucleic acid construct of claim 5, wherein the secretory signal peptide comprises the IL-2 leader sequence MYRMQLLSCIALSLALVTNS (SEQ ID NO: 2).
7. 2. The nucleic acid construct of claim 1, wherein the affinity tag comprises a Strep II tag.
8. (a) introducing the nucleic acid construct of any one of claims 1 to 7 into a non-human mammalian cell; (b) introducing a nuclease into a non-human mammalian cell; wherein the nuclease causes a single-strand break or a double-strand break at a safe harbor site in the genome of the non-human mammalian cell, and the nucleic acid construct is integrated into the genome of the non-human mammalian cell at the safe harbor site by homologous recombination.
9. 9. The method of claim 8, wherein introducing the nuclease comprises introducing an expression construct encoding the nuclease.
10. 9. The method of claim 8, wherein the nuclease comprises a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, or a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) protein and a guide RNA (gRNA).
11. A genetically modified non-human mammalian cell produced by the method of claim 8.
12. A genetically modified non-human mammal having cells expressing a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein, produced by the method of claim 8.
13. A genetically modified non-human mammalian cell, the genome of which comprises the nucleic acid construct of any one of claims 1 to 7 integrated into a safe harbor site.
14. (a) obtaining cells from the genetically modified non-human mammal of claim 12; (b) screening cells obtained from the genetically modified non-human mammal for expression of a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein; (c) isolating cells that express the fusion protein; and 1. A method for isolating cells obtained from a genetically modified non-human mammal, comprising:
15. 15. The method of claim 14, wherein the cells are screened by fluorescence-activated cell sorting (FACS) or magnetic-activated cell sorting (MACS).
16. 15. The method of claim 14, wherein the affinity tag is expressed on the cell surface of the genetically modified non-human mammalian cell.
17. The method of claim 14, wherein the fluorescent reporter protein is exposed on the cytoplasmic surface of the non-human mammalian cell.
18. A nucleic acid construct comprising a linker, a leader sequence, and a transmembrane reporter cassette encoding an affinity tag, a transmembrane domain, and a fluorescent reporter protein.
19. 20. The nucleic acid construct of claim 18, comprising single-stranded DNA, double-stranded DNA, a plasmid, or a viral vector.
20. 20. The nucleic acid construct of claim 19, further comprising a first homology arm and a second homology arm that are homologous to a first target sequence and a second target sequence, respectively, wherein the first and second target sequences are flanked by immunoglobulin constant region sites.
21. 20. The nucleic acid construct of claim 19, wherein the first target sequence is located upstream of the immunoglobulin constant region site and the second target sequence is located downstream of the stop codon of the immunoglobulin constant region site.
22. 19. The nucleic acid construct of claim 18, wherein the linker comprises a stop codon and an internal ribosome entry site (IRES).
23. 23. The nucleic acid construct of claim 22, wherein the linker comprises a protease recognition site and a self-cleaving peptide, and the protease recognition site comprises a Furin protease recognition site.
24. (a) introducing a nucleic acid construct according to any one of claims 18 to 23 into a non-human mammalian cell; (b) introducing a nuclease into a non-human mammalian cell; wherein the nuclease causes a single-strand break or a double-strand break at an immunoglobulin constant region site in the genome of the non-human mammalian cell, and the nucleic acid construct is integrated into the genome of the non-human mammalian cell at the immunoglobulin constant region site by homologous recombination.
25. 25. The method of claim 24, wherein introducing the nuclease comprises introducing an expression construct encoding the nuclease.
26. 25. The method of claim 24, wherein the nuclease comprises a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a meganuclease, or a clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) protein and a guide RNA (gRNA).
27. 25. A genetically modified non-human mammalian cell produced by the method of claim 24.
28. 24. A genetically modified non-human mammalian cell comprising a genome comprising the nucleic acid construct of any one of claims 18 to 23 integrated into an immunoglobulin constant region site.
29. (a) obtaining cells from the genetically modified non-human mammal of claim 28; (b) screening cells obtained from the genetically modified non-human mammal for expression of a fusion protein comprising an affinity tag, a transmembrane domain, and a fluorescent reporter protein; (c) identifying immunoglobulin-expressing cells based on expression of the fusion protein; 1. A method for identifying immunoglobulin-expressing cells obtained from a genetically modified non-human mammal, comprising:
30. 30. The method of claim 29, wherein the cells are screened by fluorescence-activated cell sorting (FACS) or magnetic-activated cell sorting (MACS).
31. 30. The method of claim 29, wherein the affinity tag is expressed on the cell surface of the genetically modified non-human mammalian cell.
32. 30. The method of claim 29, wherein the fluorescent reporter protein is exposed on the cytoplasmic surface of the non-human mammalian cell.
33. 30. The method of claim 29, further comprising isolating the immunoglobulin expressed from the cells obtained from the genetically modified non-human mammal.
34. (i) cloning the variable region of the immunoglobulin of claim 33; (ii) producing a therapeutic or diagnostic immunoglobulin comprising the variable region obtained in (i); 10. A method for producing a therapeutic or diagnostic immunoglobulin, comprising:
35. (i) obtaining immunoglobulin-expressing cells from the genetically modified non-human mammal of claim 27; (ii) immortalizing the immunoglobulin-expressing cells obtained in (i); and (iii) isolating the immortalized immunoglobulin-expressing cells or the monoclonal antibodies expressed by the nucleic acid sequences encoding the monoclonal antibodies; A method for producing a monoclonal antibody, comprising: