Screening method for signal peptides for efficient expression and secretion of heterologous polypeptides in mammalian cells - Patents.com

JP2024545273A5Pending Publication Date: 2025-12-22ヘルシンギン イリオピスト
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024536406
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-23
Filing Date
2022-12-23
Publication Date
2025-12-22

AI Technical Summary

Technical Problem

Current methods for optimizing signal peptides for recombinant protein production in mammalian cells are labor-intensive and time-consuming, limiting the number of sequences that can be tested for efficient protein secretion.

Method used

A rapid and exhaustive screening method using a physical platform and artificial intelligence to identify optimal signal peptides, involving a pool of viral expression vectors, host cell transformation, and next-generation sequencing to determine the most efficient signal peptides for protein secretion.

Benefits of technology

Enables high-level production of therapeutically relevant proteins by efficiently screening thousands of signal peptides, significantly faster and more comprehensive than traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000020_0000
    Figure 00000020_0000
  • Figure 00000020_0001
    Figure 00000020_0001
  • Figure 00000021_0000
    Figure 00000021_0000
Patent Text Reader

Abstract

The present disclosure is directed to a method for screening signal peptides for efficient expression and secretion of heterologous polypeptides in mammalian cells, comprising: providing a pool of viral expression vectors encoding a polypeptide of interest comprising a variety of candidate signal peptides, wherein each viral expression vector of the pool comprises at least a polynucleotide encoding a fusion protein, the polynucleotide comprising: (i) a promoter, (ii) a sequence encoding a signal peptide, (iii) a sequence encoding a polypeptide of interest, (iv) optionally, a sequence encoding an epitope tag, and (v) a sequence encoding a C-terminal signal peptide for glycosylphosphatidylinositol (GPI) attachment or a sequence encoding a transmembrane domain; hosting the viral vector pool such that each host cell is preferably transformed with only one viral vector from the pool on average. The method includes the steps of transforming a host cell; expressing the fusion protein in the host cell to produce a fusion protein that is anchored to the host cell surface by a GPI or a transmembrane domain of the fusion protein; contacting the host cell with a first binding reagent, preferably an antibody that specifically binds to the target polypeptide; or, if an epitope tag is present in the fusion protein, a first binding reagent that specifically binds to the epitope tag; contacting the host cell with the first binding reagent, which may optionally be labeled with a fluorescent label or other detection means; dividing the transformed host cells into at least two groups based on the fluorescent properties of each host cell; and performing next-generation sequencing on the group of host cells that show the most efficient fusion protein expression to identify the optimal signal peptide for the target polypeptide in the host cell. The present disclosure is also directed to the vector, host cell, or DNA library used in the method.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to the field of recombinant protein production. In particular, the present disclosure relates to a method for selecting a specific signal peptide from among a variety of signal peptides that efficiently secretes a protein of interest from a host cell. [Background technology]

[0002] Protein trafficking is a dynamic mechanism for the transmembrane transport and intramembrane trafficking of approximately 30% of all cellular proteins (Gemmer and Forster, 2020). This process is facilitated by the translocon, a multi-subunit protein complex present in the endoplasmic reticulum (ER) membrane. Sec61, a universally conserved heterotrimeric protein channel, forms the core of the translocon.

[0003] When ribosomes initiate secretory protein synthesis in the cytoplasm of eukaryotic cells, a sequence located at the amino terminus of the nascent polypeptide chain, the ER signal peptide (SP), directs the ribosome to the ER membrane. The SP sequence of the nascent protein is recognized by the signal recognition particle (SRP), which translocates the nascent polypeptide across the ER membrane. Thus, the SP functions as a destination code that marks the protein secretory pathway and the protein target site (Blobel and Dobberstein, 1975). Based on computational and experimental studies, the SP sequence has been divided into three characteristic regions; a hydrophilic N-terminal region containing positively charged amino acids; a hydrophobic core region; and a C-terminal region that contains a cleavage site by signal peptidases, usually containing polar amino acid residues (von Heijne, 1985).

[0004] Signal peptide-mediated transport of secretory proteins into the ER lumen has been found to be a bottleneck in the secretory pathway and is therefore a key challenge that needs to be resolved in order to achieve stable production of recombinant proteins. Although signal peptides are highly heterogeneous, many signal peptides have been shown to be functionally interchangeable among different species (Tan, Ho et al., 2002). On the other hand, different signal peptides can have very different effects on the secretion and function of the produced proteins (Kober et al., 2013). Thus, the efficiency of protein secretion can be significantly affected by the signal peptide sequence. These findings strongly suggest that signal peptide optimization is important if the goal is to produce the highest amount of recombinant proteins in mammalian systems.

[0005] Current methods for signal peptide optimization rely on laborious testing of signal peptides individually (Kober et al., 2013), which limits the number of sequences that can be tested and makes finding the optimal signal peptide impossible or prohibitively time-consuming. Summary of the Invention

[0006] In this patent application, we present a screening platform capable of rapid and efficient comprehensive screening of thousands of signal peptides to discover optimal signal peptides that enable high-level production of therapeutically relevant proteins in mammalian cells (Figures 1 and 6). The disclosure includes details of the use of the platform for signal peptide optimization of two model proteins: superfolder GFP and coagulation factor VII (FVII), a therapeutic protein used to treat hemophilia-like bleeding disorders resulting from deficiencies in factors VII or IX.

[0007] In addition to the physical signal peptide screening platform, the inventors are also designing an artificial intelligence method to predict optimal signal peptides for the production of different proteins of biotechnological importance, which is considered to be one of the areas of interest in the art, and this approach is likely to constitute an important part of the present disclosure.

[0008] In one aspect of the disclosure, there is provided a method for screening for a signal peptide that allows efficient expression and secretion of a heterologous polypeptide in a mammalian cell, the method comprising the steps of: (a) providing a pool of viral expression vectors encoding a polypeptide of interest containing a variety of candidate signal peptides; wherein each viral expression vector of the pool comprises at least a polynucleotide encoding a fusion protein; The polynucleotide comprises: (i) a promoter; (ii) a sequence encoding a signal peptide, (iii) a sequence encoding a polypeptide of interest; (iv) optionally, a sequence encoding an epitope tag; and (v) a sequence encoding a C-terminal signal peptide for glycosylphosphatidylinositol (GPI) attachment, or a sequence encoding a transmembrane domain; Including, The process of providing a pool; (b) transforming host cells with the pool of viral vectors such that each host cell is preferably transformed, on average, with only one viral vector of the pool; (c) expressing the fusion protein in a host cell to produce a fusion protein that is GPI-anchored to the host cell surface or that is anchored to the host cell surface by a transmembrane domain of the fusion protein; (d) contacting the host cell with a first binding reagent; wherein said first binding reagent is preferably an antibody that specifically binds to said polypeptide of interest; Alternatively, if the epitope tag is present in the fusion protein, contacting the fusion protein with a first binding reagent that specifically binds to the epitope tag; wherein the binding reagent is optionally labeled with a fluorescent label or other detection means; contacting the host cell with a first binding reagent; (e) optionally contacting the host cells with a second binding reagent, preferably an antibody; wherein if the first binding reagent is not a labeled binding reagent, the second binding reagent is fluorescently labeled; wherein said second binding reagent specifically binds to said first binding reagent; contacting; (f) dividing the transformed host cells into at least two groups based on the fluorescent characteristics of each host cell; and (g) performing next generation sequencing on the group of host cells exhibiting the most efficient fusion protein expression in order to identify the optimal signal peptide for the polypeptide of interest in said host cells.

[0009] In another aspect of the present disclosure, there is provided a viral expression vector comprising at least a polynucleotide encoding a fusion protein; wherein the polynucleotide comprises: (i) a promoter; (ii) a sequence encoding a signal peptide, (iii) a sequence encoding a polypeptide of interest; (iv) optionally, a sequence encoding an epitope tag; and (v) a sequence encoding a C-terminal signal peptide for glycosylphosphatidylinositol (GPI) attachment.

[0010] In a further aspect of the present disclosure: (1) a host cell comprising a viral vector according to the present disclosure; and (2) a DNA library comprising a plurality of viral vectors according to the present disclosure; Provides The vector now encodes a variety of candidate signal peptides for the polypeptide of interest. [Brief description of the drawings]

[0011] [Figure 1]Illustrative diagram of the signal peptide (SP) screening platform. (A) Schematic diagram of the fate of epitope-tagged, GPI-anchored proteins of interest (POI-EPI-GPI) with functional and non-functional signal peptides. A functional signal peptide directs the protein of interest to the ER where folding and maturation of the protein occurs. Correctly folded and processed POI-EPI-GPI is transported to the cell surface where it can be labeled with an epitope tag or an antibody that binds to the folded POI. If the signal peptide does not facilitate translocation of the POI to the ER or its maturation or folding (i.e., it is non-functional), the POI will remain there (translocation defective) or will be translocated to the cytoplasm via an ER-associated degradation pathway (folding / maturation defective) and degraded by the cytoplasmic proteasome. (B) Workflow of the signal peptide screening process. A pool of DNA oligos encoding the signal peptide is synthesized and preferably batch cloned into a lentiviral transfer plasmid encoding the POI-EPI-GPI. Human or Chinese Hamster Ovary (CHO) cells are transduced at a low MOI with the construct encoding SP-POI-EPI-GPI, such that each transduced cell contains, on average, only one copy of the single construct encoding SP-POI-EPI-GPI. Transduced cells expressing SP-POI-EPI-GPI and preferably the iRFP protein as a transduction marker are stained with an epitope tag-recognizing antibody (anti-EPI) and then sorted based on the anti-EPI signal. Alternatively, an antibody recognizing folded protein can be used for cell staining. The SP insert coding region of the genome-integrated lentiviral construct is amplified and then identified by next-generation sequencing. Comparison of the read SP sequences of cell populations showing strong and weak anti-EPI / anti-folded protein signals allows the identification of SPs that enhance POI production to the highest level. As an enrichment control before next-generation sequencing, the specific enrichment of the signal peptide in the sorting pool is evaluated by qPCR. [Diagram 2](A) Lentiviral expression cassette. Signal peptide (SP) library and variable regions 1 and 2 into which the protein of interest (POI) is cloned. (B) Schematic representation of the translated protein polypeptide that is secreted outside the expression host and displayed on the cell membrane via a GPI / TM anchor. LTR, long terminal repeat; promoter, inducible promoter; SP, signal peptide; POI, protein of interest; GPI, glycophosphatidylinositol; TM, transmembrane domain; IRES, internal ribosome entry site; transduction marker, fluorescent protein; N, amino terminus; C, carboxy terminus. (C) In this experiment, the construct contains coagulation factor VII (FVII) as the POI and the epitope tag is a 3xFLAG tag. [Diagram 3] Identification of enriched SP by qPCR / RT-PCR. In step 1, GFP-expressing CHO cells were sorted into two different tubes based on GFP expression levels (high GFP (Q2) and low GFP (Q3)). In steps 2 and 3, genomic DNA was extracted from the sorted cells and used as template in qPCR / RTPCR. Site-specific primers were used for amplification of the signal peptide region by qPCR. The low GFP cells contained the azurocidin preprotein SP. GFP, green fluorescent protein; RFP, red fluorescent protein; gDNA, genomic DNA; high GFP, cells expressing higher levels of GFP; low GFP, cells expressing low levels of GFP; qPCR, quantitative polymerase chain reaction; RT-PCR, reverse transcription PCR. [Figure 4] Specific SPs affect the expression levels of coagulation factor VII (FVII). (A) Western blot of HEK293T cells transduced with construct 1: ORI SP FVII and construct 2: IgK SP FVII (expected protein size 50 kDa) stained with anti-FLAG antibody. (B) Quantification of relative expression levels based on the average intensity of the 50 kDa band. [Diagram 5]Antibody staining of GPI-anchored FVII allows sorting of differentially stained cells. (A) FACS plot of HEK293T cells (negative control) induced with doxycycline and stained with both primary (anti-FLAG) and secondary (AlexaFluor647) antibodies. (B) FACS plot of HEK293T cells (negative control) transduced with SP library FVII, induced with doxycycline and stained with only secondary (AlexaFluor647) antibodies. (C) FACS plot of HEK293T cells transduced with SP library FVII, induced with doxycycline and stained with both primary (anti-FLAG) and secondary (AlexaFluor647) antibodies; 1.2% of cells show both the transduction marker (tRFP) and secondary antibody signal (highlighted by the dotted box in Q2). [Figure 6] 1 shows the step-by-step workflow used in the signal peptide screening platform of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] The term "signal peptide" as used herein means a peptide that is the N-terminal portion of a secretory protein that is secreted outside the cell and passes through the cell membrane. A signal peptide is usually composed of approximately 10 to 30 amino acids, and is later cleaved and removed by a protease specific to the cell membrane. Therefore, only secretory proteins are transported outside the cell. Signal peptides act as targeting signals, enabling intracellular transport mechanisms to guide proteins to specific intracellular or extracellular locations. To date, it has been known that more than 4,000 types of signal peptides exist in eukaryotic cells. DNA libraries encoding signal peptides are disclosed, for example, in WO2021045541A1 and KR20210028116A.

[0013] The term "protein of interest" or "POI" as used herein refers to a protein intended for efficient production by using a suitable host cell. The POI is preferably a protein of therapeutic or diagnostic importance, such as an antibody.

[0014] As used herein, the term "epitope tag" or EP refers to a technique of fusing a tag (typically 6-30 amino acids) to a recombinant protein by using genetic engineering to place a sequence encoding the epitope in the same open reading frame of the protein. This technique allows the detection of the tagged protein even when no antibody is available by selecting an epitope tag for which an antibody is available. By selecting an appropriate epitope tag and antibody pair, it is possible to find a combination with suitable properties for the desired experimental method (Western blot analysis, immunoprecipitation, immunochemistry, affinity purification, etc.). Preferred epitope tags can be selected from the group consisting of, for example, histidine tag (His tag), myc tag, FLAG tag, small ubiquitin-like modification tag (SUMO tag), protein C heavy chain tag (HPC tag), calmodulin-binding peptide tag (CBP tag), and hemagglutinin tag (HA tag).

[0015] The term "GPI-anchored" as used herein refers to a protein anchored to the outer surface of a eukaryotic cell by a glycosylphosphatidylinositol (GPI). These secreted proteins are anchored to the cell membrane by a GPI moiety covalently attached to the C-terminus of the protein. The GPI moiety consists of a conserved core glycan, a phosphatidylinositol, and a glycan side chain. The structure of the core glycan is EtNP-6Manα2-Manα6-(EtNP)2Manα4-GlNα6-myoIno-P-lipid (EtNP, ethanolamine phosphate; Man, mannose; GlcN, glucosamine; Ino, inositol). In the present disclosure, the protein is GPI-attached by a C-terminal signal peptide (see, e.g., EP3389682).

[0016] The term "transmembrane domain" as used herein refers to a hydrophobic alpha-helical structure that penetrates the host cell membrane. The transmembrane domain may be directly fused to the C-terminal portion of the fusion protein encoded by the vector. In certain embodiments, the transmembrane domain is derived from an integral membrane protein (e.g., a receptor, a cluster of differentiation (CD) molecule, an enzyme, a transporter, a cell adhesion molecule, etc.). In a preferred embodiment, the transmembrane domain is derived from a type I transmembrane protein, exemplified by the human VCAM-1 protein (vascular cell adhesion molecule 1). When the mature form of a type I transmembrane protein is localized to the cell membrane, the protein is anchored to the lipid membrane by the membrane stop sequence of the anchor, but has an N-terminal domain for guidance to the extracellular space. Further examples of transmembrane domains according to the present disclosure include, but are not limited to, the Tim1, Tim2, and Tim3 transmembrane domains, the FcR transmembrane domain, and the CD8a transmembrane domain. Further transmembrane domains for use in the present invention are disclosed in EP3389682.

[0017] The term "vector" as used herein refers to a nucleic acid molecule capable of mediating entry, e.g., introduction, import, etc., of another nucleic acid molecule into a cell. The introduced nucleic acid is generally linked, e.g., inserted, into the vector nucleic acid molecule. The vector can include sequences directing autonomous replication or sequences sufficient for integration into the host cell DNA. Useful vectors include, e.g., plasmids, cosmids, and viral vectors. Useful viral vectors include, e.g., replication-defective retroviruses, adenoviruses, adeno-associated viruses, and lentiviruses. As will be apparent to those skilled in the art, viral vectors may contain various viral components in addition to the nucleic acid that mediates entry of the introduced nucleic acid. Thus, the term "viral vector" may refer to either a virus or viral particle capable of introducing a nucleic acid into a cell, or the introduced nucleic acid itself.

[0018] In one embodiment, the present disclosure is directed to a method for screening a signal peptide for highly efficient expression and secretion of a heterologous polypeptide in a mammalian cell, the method comprising the steps of: (a) providing a pool of viral expression vectors encoding a polypeptide of interest containing a variety of candidate signal peptides; wherein each viral expression vector of the pool comprises at least a polynucleotide encoding a fusion protein; The polynucleotide comprises: (i) a promoter; (ii) a sequence encoding a signal peptide, (iii) a sequence encoding a polypeptide of interest; (iv) optionally, a sequence encoding an epitope tag; and (v) a sequence encoding a C-terminal signal peptide for glycosylphosphatidylinositol (GPI) attachment, or a sequence encoding a transmembrane domain; Including, The process of providing a pool; (b) transforming host cells with the pool of viral vectors such that each host cell is preferably transformed, on average, with only one viral vector of the pool; (c) expressing the fusion protein in a host cell to produce a fusion protein that is GPI-anchored to the host cell surface or that is anchored to the host cell surface by a transmembrane domain of the fusion protein; (d) contacting the host cell with a first binding reagent; wherein said first binding reagent is preferably an antibody or an aptamer that specifically binds to said polypeptide of interest; Alternatively, if the epitope tag is present in the fusion protein, contacting the fusion protein with a first binding reagent that specifically binds to the epitope tag; wherein the binding reagent is optionally labeled with a fluorescent label or other detection means; contacting the host cell with a first binding reagent; (e) optionally contacting the host cells with a second binding reagent, preferably an antibody or an aptamer; wherein if the first binding reagent is not a labeled binding reagent, the second binding reagent is fluorescently labeled; wherein said second binding reagent specifically binds to said first binding reagent; contacting; (f) dividing the transformed host cells into at least two groups based on the fluorescent properties of each host cell, preferably using fluorescence activated cell sorting (FACS); and (g) performing next generation sequencing on the group of host cells exhibiting the most efficient fusion protein expression in order to identify the optimal signal peptide for the polypeptide of interest in the host cells.

[0019] In a preferred embodiment, in step (a) of the method, said pool of viral expression vectors encoding a polypeptide of interest is prepared by the following method: synthesizing a library of oligonucleotides encoding various signal peptides and their complementary sequences; wherein said oligonucleotides, upon binding to complementary oligonucleotides in the library, form restriction enzyme recognition site overhangs at both free ends of the sequence; annealing the oligonucleotides in the library at the overhangs of the restriction enzyme recognition sites with complementary oligonucleotides present in the library to generate double-stranded DNA fragments; and combining a DNA fragment encoding a signal peptide with a polynucleotide encoding a protein of interest and ligating the double-stranded DNA fragment into a viral vector with suitable cloning sites in order to generate a pool of viral expression vectors encoding polypeptides of interest with a variety of candidate signal peptides.

[0020] In a preferred embodiment, the library of oligonucleotides encoding various signal peptides includes known signal peptides of eukaryotic, preferably mammalian, bacterial or viral proteins, modified and / or artificial sequences thereof.

[0021] In another preferred embodiment, the method of the present disclosure comprises the further steps of: (h) cloning the polynucleotide encoding the optimal signal peptide combination detected in step (g) and the target polypeptide into a second expression vector, transforming a host cell with the second vector, and producing the polypeptide in the host cell.

[0022] In a preferred embodiment, the promoter of the vector is an inducible promoter, preferably a tetracycline-regulated promoter.

[0023] In another embodiment, the present disclosure is directed to a viral expression vector comprising at least a polynucleotide encoding a fusion protein; wherein the polynucleotide is: (i) a promoter; (ii) a sequence encoding a signal peptide, (iii) a sequence encoding a polypeptide of interest; (iv) optionally, a sequence encoding an epitope tag; and (v) a sequence encoding a C-terminal signal peptide for glycosylphosphatidylinositol (GPI) attachment; Includes.

[0024] In a further embodiment, the present disclosure provides a method for producing a method for treating a cancer cell comprising: (1) a host cell comprising a vector according to the present disclosure; or (2) a DNA library comprising a plurality of viral vectors according to the present disclosure; The target is The vector now encodes a variety of candidate signal peptides for the polypeptide of interest.

[0025] The following examples illustrate the principles of the invention through one or more more specific applications, but it will be apparent to those skilled in the art that many changes in the embodiments, applications and details may be made without the use of inventive faculty and without departing from the principles and concepts of the disclosure. Accordingly, it is not intended that the invention be limited except as defined by the following claims.

[0026] The verbs "to comprise" and "to include" are used herein as open limitations that do not exclude or require the presence of unrecited features. Unless expressly stated otherwise, features recited in dependent claims are mutually combinable. Furthermore, it is to be understood that the use of "a" and "an" as singular forms throughout this specification does not exclude the presence of a plurality.

[0027] Reference throughout this specification to one embodiment or an embodiment means that the particular characteristic, structure, or feature described in relation to that embodiment is included in at least one embodiment of the present disclosure. Thus, even if the expressions "in one embodiment," "in an embodiment," or "in a preferred embodiment" are used in various places throughout this specification, they do not necessarily all refer to the same embodiment. For example, when referring to a numerical value using terms such as about or substantially, the exact numerical value is also disclosed.

[0028] Experimental section Construction of pINDUCER-based mammalian expression plasmids Lentiviral transfer plasmids that allow cell membrane-anchored expression of proteins of interest (POI) form the core of our signal peptide screening platform (Figure 1). Lentiviral transduction of expression constructs allows the isolation of cells with a single signal peptide-POI combination that exhibits high expression levels of the POI (Figure 1). By combining the use of our lentiviral constructs with large-scale gene synthesis of signal peptide libraries and the use of next-generation sequencing to identify optimal signal peptides from flow cytometry sorted pooled cell samples, a comprehensive screening of signal peptides can be achieved (Figure 1). This massively parallel signal peptide screening platform allows the identification of optimal signal sequences for protein production both more rapidly and more comprehensively compared to current signal peptide optimization methods that rely on testing signal peptides individually.

[0029] The lentiviral transfer plasmids used for genomic integration of the signal peptide-protein of interest constructs we tested (Figure 2) were based on the pINDUCER11 plasmid (Meerbrey, Hu et al. 2011) in which the original transduction marker (GFP) was changed to iRFP670 (Shcherbakova, Verkhusha 2013) or tRFP (Strack, Strongin et al. 2008) by amplifying the IRES-tRFP or IRES-iRFP670 insert by overlap extension PCR and then cloning the insert into the vector by digestion with the restriction enzymes AscI and PacI (NEB) and subsequent ligation with T4 ligase (NEB).

[0030] For signal peptide screening of superfolder-GFP SP-ubiquitin G76V (Ub(G76V)-sfGFP-GPI anchor constructs, construction was performed by overlap extension PCR from different inserts of Ub(G76V) (Dantuma, Lindsten et al. 2000), sfGFP (Costantini, Baloban et al. 2015) and GPI anchor (Rhee, Pirity et al. 2006). The constructed fusion protein insert was cloned into the original pINDUCER11 plasmid at the position of tRFP-shRNA insert.

[0031] For FVII signal peptide screening, DNA constructs containing the hamster codon-optimized inactive FVII (D302N)-mutant were ordered from Twist Biosciences as fusion protein gene fragments of SP-3xFLAG-tag-FVII-GPI or SP-HA-tag-FVII-GPI. These DNA constructs encoding the fusion proteins were cloned into the original pINDUCER11 plasmid at the position of the tRFP-shRNA insert.

[0032] Functional elements of the modified pINDUCER11 plasmid CMV promoter: a doxycycline-inducible promoter that controls the expression of a gene of interest Enhanced green fluorescent protein (eGFP): Used as the gene of interest in our proof-of-concept experiments. · Internal ribosome entry site (IRES): Allows cap-independent translation initiation. · Transactivator 3 (rtTA3): controls doxycycline induction of the CMV promoter. Under the control of the constitutive promoter hUBC. Human ubiquitin C promoter (hUBC): A constitutive promoter that controls ectopic expression of iRFP670 or tRFP. · iRFP670: Transduction of a reporter protein under the control of the constitutive promoter hUBC. tRFP: Transduction of reporter proteins under the control of the constitutive promoter hUBC SV40 gene: plasmid replication region

[0033] Cloning of signal peptide into expression plasmid To perform directional cloning, the SP insert and the circular pINDUCER vector generated by oligonucleotide annealing were digested with two restriction enzymes, MluI and NotI. The pINDUCER vector digested with the two restriction enzymes was further treated with shrimp alkaline phosphatase (rSAP), which nonspecifically catalyzes the dephosphorylation of the 5′ end; this treatment was to avoid self-ligation of the vector. However, since ligation of DNA fragments requires the 5′ phosphate group to form a phosphodiester bond, the oligonucleotides digested with the two restriction enzymes were phosphorylated by thermal cycling with T4 polynucleotide kinase (NEB) in the presence of ATP. The SP insert and the linear pINDUCER vector thus generated were covalently linked by ligation with T4 DNA ligase (NEB). The entire ligation mixture was transformed into stable3 (Thermo Fisher) chemically competent E. coli cells and grown on LB agar plates containing the antibiotic ampicillin (100 μg / ml). The resulting colonies were screened by Sanger sequencing for containing the correct SP insert.

[0034] Cell lines and their culture under standard conditions Human Embryonic Kidney 293 cells (HEK293T) (Thermo-Fisher Scientific) were cultured as adherent monolayers in DMEM containing 10% fetal bovine serum (FBS) and 0.5% L-glutamine. FreeStyle Chinese Hamster Ovary (CHO) suspension-adapted (CHO-S) cells (Thermo-Fisher Scientific) were grown in suspension culture in FreeStyle™ CHO medium (Thermo-Fisher Scientific). Both cell lines were cultured under standard conditions at 37°C and 5% CO2.

[0035] Example 1: Proof-of-concept experiment Model protein 1: Green fluorescent protein (GFP) For proof-of-concept experiments with sfGFP, the following signal peptides were selected: SP1 (interleukin 4), SP2 (serum albumin), SP3 fPrP (mut17-21), SP4 (azurocidin preprotein), SP5 (cellulase), SP6 (PrP), SP7 (Vcam), and SP8 (FCRL-1).

[0036] The above SPs were cloned into modified pINDUCER11 (see 2). Restriction enzyme sites MluI (NEB) and NotI (NEB) were utilized. The detailed cloning strategy is described herein.

[0037] For each signal peptide, oligonucleotides were synthesized containing the respective sequences encoding the signal peptide (Integrated Data Technologies). All oligonucleotides were flanked at the 5' and 3' ends by MluI and NotI restriction enzyme sites, respectively. dsDNA inserts were generated by annealing the single-stranded oligonucleotides in a thermal cycler (Bio-Rad).

[0038] Model protein 2: Coagulation factor FVII (FVII) As disclosed above, we designed a third-generation lentiviral transfer plasmid, pINDUCER11, containing the gene of interest. Two different signal peptides (Table 1), selected based on the differential effect of signal peptides on FVII expression PMID (Peng, Yu et al. 2016), were cloned into the N-terminus of the 3xFLAG tag-FVII-GPI fusion construct and the HA-tag-FVII-GPI fusion construct.

[0039] [Table 1]

[0040] (a) Preparation of lentivirus expressing a protein of interest in mammalian cells The third generation lentiviral transfer plasmid pINDUCER11 containing the gene of interest was designed as disclosed above. Other lentiviral packaging plasmids pVP157, pVP158, pVP159, and pVP160 were kindly provided by Martin Kampmann / Jonathan Weissmann (Bassik, Kampmann et al. 2013).

[0041] Lentivirus preparation was performed according to a polyethylenimine (PEI)-mediated transfection protocol according to methods described elsewhere (PMID: (Lobato-Pascual, Saether et al., 2013; Bassik, Kampmann et al., 2013). Briefly, the packaging vector was introduced into HEK293T cells using the cationic polymer PEI with Transporter5® transfection reagent (Polysciences, Germany). A total of 800-1000 μg of transfection plasmid was mixed with the packaging plasmid. The transfection reagent was diluted in PBS and mixed with the plasmid. The mixture was incubated at room temperature for 25 min. The mixture was added dropwise to the adherent HEK293T cell culture to ensure uniform distribution. After transfection, the cells were incubated for 3 days (72 h). The supernatant containing the lentivirus was collected. The supernatant was filtered using a 0.45 μ filtration assembly. The filtered lentivirus was collected, aliquoted, and stored at 4 °C.

[0042] (b) Measurement of virus titers in HEK293T cells and CHO cells For lentivirus titration, the protocol described by Singer Tiscornia et al. (2006) was used. Day 1: Seed 24-well plates with 500 μl of cells (100,000 cells per well) Day 3: Serial 4-fold dilutions were prepared from the lentivirus stock in culture medium. Virus dilution: undiluted, 1:4, 1:16, 1:64 Media was removed from wells of a 24-well plate and 250 μL of fresh media was added. Virus dilutions (20 μL) were added dropwise to the cells, gently mixed and the cells were incubated at 37°C. After 2-3 hours, an additional 250 μL of media was added to the wells. Cells were allowed to grow for 48 hours. Day 5: The medium in the wells was discarded. The cells were washed once with 150 μl of PBS. After detaching the cells with 30 μl of trypsin (0.5%), 500 μl of fresh medium was mixed. In another plate, 400 μl of fresh medium was added together with 100 μl of cells from day 3. The cells were then incubated at 37° C. for 3 days. This step was repeated three more times. Overall, the cells were washed and allowed to split for one week. These extensive washing steps were important for the removal of viral particles. Day 14: Cells were prepared for FACS analysis as described in the Flow Cytometry section below. FACS analysis was performed to measure the percentage of fluorescent reporter (GFP and mRFP) positive cells.

[0043] Viral titers (VT=TU / mL, transducing units) were calculated according to the following formula: Transducing units (TU / ml) = (N x P) / (V x D) Where: N = number of cells in each well used for infection on day 2; P = percentage of GFP / RFP positive cells (should be 10%–20%); V = virus volume used to infect each well; In this protocol, V(ml) = 20(μl) × 10-3; D = dilution factor; TU=transducing unit.

[0044] (c) Cell surface antibody staining and flow cytometry HEK293T and CHO cells were transduced with lentiviruses encoding FVII- or sfGFP-fusion proteins containing 3X-FLAG or HA tags. Transduced cells were divided twice. 24 hours before harvesting, cells were induced with 1 μg / ml doxycycline. Cells were harvested with 10 mM EDTA. Cells were pelleted by centrifugation at 2400 rpm for 5 min at room temperature and then resuspended in FACS buffer at room temperature (1xPBS, 4% FBS, 10 mM EDTA). An aliquot of sample was saved for FACS analysis, the rest was pelleted by centrifugation at 2400 rpm for 2 min at room temperature and then resuspended in 1 mL FACS buffer. Centrifugation was repeated and then the medium was removed (1 wash). Cell pellets were resuspended in 500 μl of primary antibody solution (see Table 2) and incubated for 15 min at room temperature, followed by two washes with 1 mL FACS buffer. Cells were then resuspended in 500 μl of secondary antibody solution (see Table 2) and incubated for 15 min at room temperature, washed twice with 1 mL of FACS buffer, and finally resuspended in FACS buffer. Immediately prior to FACS measurement, cells were filtered twice into FACS tubes and then analyzed on a BD LSRFortessa flow cytometer; or sorted into populations showing low and high antibody signals on a BD FACSAria II flow cytometer (see Figure 1). The flow cytometry data were analyzed with FlowJo software (Tree Star, Ashland, OR, USA).

[0045] [Table 2]

[0046] The laser and filter parameters used for the BD LSRFortessa analysis were: (1) APC-A for iRFP signal; and antibodies: goat anti-mouse Alexa 647, goat anti-rabbit Alexa Fluor 546 (2) tRFP signal was measured using PE-A (3) AlexaFluor488 for GFP signal; and antibody: PGT145 Alexa488 (4) AlexaFluor 700 for HA signal; and antibody: anti-rabbit 680RD

[0047] Identification of individual signal peptides by qPCR qPCR was used to identify individual signal peptides fused to the protein of interest (POI) from cells transduced with a mixed pool SP-POI lentiviral construct. For this purpose, gDNA from flow cytometrically sorted cells was isolated by Nucleospin Blood Quickpure or L kit (Macherey Nagel). The enrichment of constructs containing specific signal peptides in specific sorted cell populations (see Figure 1) was then shown by qPCR with signal peptide specific primers (Table 3). This qPCR was performed with a 5 μl reaction mixture containing 75 ng gDNA, 1 μM forward and reverse primers, and additionally 1 mM MgCl2. The PCR conditions were: Step 1: 95°C for 10 minutes; Step 2: 95°C for 15 seconds; Step 3: 66°C for 5 seconds; Step 4: 72°C for 10 seconds; Repeat steps 2-4 45 times. Step 5: 72°C for 5 min.

[0048] [Table 3]

[0049] Identification of signal peptides from large pools by next-generation sequencing Next generation sequencing (NGS) was used to identify all enriched signal peptides for the best expression of POIs showing high antibody signal flow cytometry sorted cell populations (see Figure 1). For this purpose, gDNA of flow cytometry sorted cells was first isolated by Nucleospin Blood Quickpure or L kit (Macherey Nagel). 2 μg of gDNA, corresponding to approximately 360000 copies of the diploid CHO cell genome, was added to a 50 μl PCR reaction mixture; the mixture contained Q5 buffer (NEB), 0.2 mM dNTPs, 0.5 mM forward and reverse primers (see Table 4), 4% DMSO, 1 U Q5 hot start polymerase (NEB), and additionally 2 mM MgCl2. PCR was then performed using the following parameters: Step 1: 98°C for 3 minutes; Step 2: 98°C for 30 seconds; Step 3: 61°C for 15 seconds; Step 4: 72°C for 15 seconds; Repeat steps 2 to 4 25 times. Step 5: 72°C for 5 min.

[0050] The amplified SP-encoding amplicons were then size-selected using AMPure XP beads (Beckman Coulter) and finally sequenced by MiSeq sequencing (Illumina). Signal peptides conferring the highest expression levels were identified by comparing read enrichment between the highest expressing 1% and the remaining 99% of transduced sorted cells.

[0051] [Table 4]

[0052] Assay for protein production by transient transformation We have demonstrated that the results of our SP screening can be applied to protein production conditions that mimic high-level protein production in the biopharmaceutical industry using ELISA assays. Here, the above assays were performed after identifying the signal peptides that conferred the highest expression levels for both FVII and sfGFP. To express the above SPs, including secreted proteins, expression plasmids encoding the proteins were used to transiently transform HEK293T cells or CHO cells, or other suitable mammalian expression host cells.

[0053] Depending on the target of interest, media containing secreted protein was typically harvested 3-5 days after transformation. Expression test samples were collected in conical tubes and centrifuged (Eppendorf) at 500xg for 5 minutes at room temperature to pellet cells. The clarified supernatant was transferred to a new tube and the amount of secreted sfGFP or FVII was quantified using a sandwich ELISA assay for the target of interest.

[0054] Briefly, 96-well ELISA plates were pre-coated with POI or its epitope tag-binding protein. The harvested clarified expression medium was then added to the coated plates. In addition, different positive and negative controls (recovered medium from commercially purified POI and mock-transduced cells, respectively) were used to confirm the functionality of the assay, in order to avoid possible false positives and to estimate its functionality based on POI expression levels and target affinity.

[0055] Typically, ELISA tests were performed following standard methods recommended by the manufacturer (e.g., Thermo-Fisher Scientific). Plate coating of the capture target was typically performed overnight at 4°C. Unbound proteins were washed off with assay buffer; washing was repeated three times, followed by quick drying by tapping the plate. A commercial blocking buffer (e.g., Thermo-Fisher Scientific products) was added to all wells and the plate was incubated for 1 h at room temperature. The blocking buffer was removed and cleared expression medium and controls were added to the wells. The washing steps were repeated as above, and the wells were treated with the appropriate labeled antibody (Thermo-Fisher Scientific). For the purpose of detecting protein-target complex signals, the plate was tapped on a paper towel and briefly air-dried, after which the labeled detection secondary antibody in blocking buffer (50 μl) was added to the wells. Excess detection antibody was washed off and the appropriate assay buffer was added. Signals were measured using an ELISA plate reader (Thermo-Fisher Scientific).

[0056] Optionally, HRP detection was used. In this case, instead of the assay buffer, HRP conjugate and HRP substrate were added in the final step, and then detection was performed using a plate reader. The results of the above method are shown in Figures 3-5. References [Table 5] JPEG2024545273000006.jpg193132JPEG2024545273000007.jpg123132

Claims

1. 1. A method for screening a signal peptide for highly efficient expression and secretion of a heterologous polypeptide in a mammalian cell, the method comprising: (a) providing a pool of viral expression vectors encoding polypeptides of interest containing various candidate signal peptides; wherein each viral expression vector of the pool comprises at least a polynucleotide encoding a fusion protein; The polynucleotide: (i) a promoter; (ii) a sequence encoding a signal peptide; (iii) a sequence encoding a polypeptide of interest; (iv) optionally, a sequence encoding an epitope tag; and (v) a sequence encoding a C-terminal signal peptide for glycosylphosphatidylinositol (GPI) attachment or a sequence encoding a transmembrane domain; Including, providing a pool; (b) transforming host cells with the pool of viral vectors such that each host cell is preferably transformed, on average, with only one viral vector from the pool; (c) expressing the fusion protein in a host cell to produce a fusion protein that is GPI-anchored to the host cell surface or that is anchored to the host cell surface by the transmembrane domain of the fusion protein; (d) contacting the host cell with a first binding reagent; wherein said first binding reagent is preferably an antibody that specifically binds to said polypeptide of interest; Alternatively, if the fusion protein contains an epitope tag, contacting the fusion protein with a first binding reagent that specifically binds to the epitope tag; wherein the binding reagent is optionally labeled with a fluorescent label or other detection means; contacting the host cell with a first binding reagent; (e) optionally contacting the host cells with a second binding reagent, preferably an antibody; wherein if the first binding reagent is not a labeled binding reagent, the second binding reagent is fluorescently labeled; wherein said second binding reagent specifically binds to said first binding reagent; contacting; (f) dividing the transformed host cells into at least two groups based on the fluorescent characteristics of each host cell; (g) performing next generation sequencing on the group of host cells that exhibit the most efficient fusion protein expression in order to identify the optimal signal peptide for the polypeptide of interest in the host cells; Including, method.

2. 10. The method according to claim 1, wherein said pool of viral expression vectors encoding a polypeptide of interest comprises: synthesizing a library of oligonucleotides encoding various signal peptides and their complementary sequences, wherein said oligonucleotides, upon binding to complementary oligonucleotides in the library, form restriction enzyme recognition site overhangs at both free ends of the sequence; annealing the oligonucleotides in the library at the overhangs of the restriction enzyme recognition sites with complementary oligonucleotides present in the library to generate double-stranded DNA fragments; and ligating the double-stranded DNA fragment into a viral vector with suitable cloning sites in order to combine the signal peptide-encoding DNA fragment with a polynucleotide encoding a protein of interest and generate a pool of viral expression vectors encoding polypeptides of interest with a variety of candidate signal peptides; A method prepared by

3. 3. The method according to claim 2, wherein the various signal peptides are selected from known signal peptides of eukaryotic, preferably mammalian, bacterial or viral proteins, modified and / or artificial sequences thereof.

4. 10. A method according to claim 1, comprising: (h) cloning the polynucleotide encoding the optimal signal peptide combination detected in step (g) and the target polypeptide into a second expression vector, transforming a host cell with the second vector, and allowing the host cell to produce the polypeptide; method.

5. 2. The method according to claim 1, wherein the host cells are cells selected from the group consisting of CHO cells, HeLa cells, HEK293 cells, BHK cells, COS7 cells, COP5 cells, A549 cells, NIH3T3 cells, MDCK cells, and WI38 cells, preferably CHO cells, or any other cell type that can be transduced by a lentivirus.

6. 10. The method according to claim 1, wherein the viral vector is a lentiviral vector.

7. 2. The method according to claim 1, wherein the promoter is an inducible promoter, preferably a tetracycline-regulated promoter, or any other suitable promoter sequence, including a constitutive promoter.

8. 2. The method according to claim 1, wherein the epitope tag is selected from the group consisting of a histidine tag (His tag), a myc tag, a FLAG tag, a small ubiquitin-like modification tag (SUMO tag), a protein C heavy chain tag (HPC tag), a calmodulin-binding peptide tag (CBP tag), and a hemagglutinin tag (HA tag), or other epitope tag or other labeling group.

9. 10. The method according to claim 1, wherein said polypeptide of interest is a therapeutic or diagnostic protein, such as a therapeutic or diagnostic antibody.

10. 10. The method according to claim 1, wherein step (f) is performed using fluorescence activated cell sorting (FACS).

11. 11. The method according to any one of claims 1 to 10, wherein each viral expression vector of the pool further comprises a polynucleotide encoding a transduction marker protein, preferably red fluorescent protein, RFP or other fluorescent protein, or a genetic marker detectable by flow cytometry.

12. 1. A viral expression vector comprising at least a polynucleotide encoding a fusion protein, the polynucleotide comprising: (i) a promoter; (ii) a sequence encoding a signal peptide; (iii) a sequence encoding a polypeptide of interest; (iv) optionally, a sequence encoding an epitope tag; and (v) a sequence encoding a C-terminal signal peptide for glycosylphosphatidylinositol (GPI) attachment or a sequence encoding a transmembrane domain; A viral expression vector comprising:

13. 13. The vector according to claim 12, wherein the inducible promoter is a tetracycline-regulated promoter or other suitable promoter sequence, including a constitutive promoter.

14. 13. The vector according to claim 12, wherein the epitope tag is selected from the group consisting of a histidine tag (His tag), a myc tag, a FLAG tag, a small ubiquitin-like modification tag (SUMO tag), a protein C heavy chain tag (HPC tag), a calmodulin-binding peptide tag (CBP tag), and a hemagglutinin tag (HA tag), or other epitope tag or other labeling group.

15. 13. The vector according to claim 12, wherein the polypeptide of interest is a therapeutic protein, such as a therapeutic antibody.

16. 13. The vector according to claim 12, wherein the vector is a lentiviral vector.

17. 13. A vector according to claim 12, further comprising a sequence encoding a transduction marker protein, preferably red fluorescent protein, RFP.

18. A host cell comprising a viral vector according to any one of claims 12 to 17.

19. 20. A host cell according to claim 18, which displays the fusion protein on its cell membrane surface, wherein the fusion protein is attached to the cell membrane by a GPI attachment or via a transmembrane domain.

20. 19. A host cell according to claim 18, wherein the host cell is selected from the group consisting of CHO cells, HeLa cells, HEK293 cells, BHK cells, COS7 cells, COP5 cells, A549 cells, NIH3T3 cells, MDCK cells and WI38 cells, preferably CHO cells, or any other cell type that can be transduced with a lentivirus.

21. A DNA library comprising a plurality of viral vectors according to any one of claims 12 to 17, wherein the vectors encode various candidate signal peptides for a polypeptide of interest.