Cell strain, and preparation method therefor and use thereof
By integrating exogenous nucleic acid fragments into high-expression fragments in CHO cells, optimizing gene sequence, the problems of differences in cell bank stability and expression efficiency in CHO cells were solved, and the effect of efficient, stable expression and reduced mismatch rate was achieved.
Patent Information
- Application Number
- PCT/CN2024/126623
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-08
- Filing Date
- 2024-10-23
- Publication Date
- 2025-05-15
AI Technical Summary
In the production process of CHO cells, the prior art uses non-targeted transgenic integration methods to cause significant differences in the stability and expression efficiency of the cell bank, and the expression results of site-specific integration still need to be improved, and the mismatch rate is high.
By site-directed integration of exogenous nucleic acid fragments into high-expression fragments in CHO cells, it is ensured that the genes encoding difficult-to-express proteins are located upstream of the genes encoding easy-to-express proteins, optimize the gene sequence to achieve high and balanced expression, and reduce the mismatch rate.
It achieves efficient and stable expression of CHO cell lines, shortens the cell line development timeline, reduces the mismatch rate, and improves production efficiency.
Smart Images

Figure CN2024126623_15052025_PF_FP_ABST
Abstract
Description
A cell line and its preparation method and application
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application number 2023114848053, filed with the Chinese Patent Office on November 8, 2023, entitled “A cell line, its preparation method and application”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of biotechnology, and specifically provides a cell line and a preparation method and application thereof. Background Art
[0004] Chinese hamster ovary (CHO) cells are an expression host used for the commercial production of recombinant proteins. Currently, a diverse portfolio of biotherapeutic drugs produced in CHO cells has been approved for marketing, including numerous monoclonal antibodies (mAbs) and other proteins. Since the advent of CHO cell culture production techniques over 30 years ago, unremitting efforts by those skilled in the art have significantly improved the productivity of CHO cells. In particular, titers of industrial antibody production in the past were typically well below 1 g / L. However, manufacturing teams are now achieving titers exceeding grams per liter, often exceeding 3 g / L. However, the current upstream cell line selection process still relies on non-targeted transgene integration to generate stably transfected CHO cell banks, from which production cell lines are established. Consequently, tedious screening within heterogeneous cell banks is required to identify clones with suitable production characteristics. Because these selected cell lines are often multi-copy, the stability of different sites within the cell line varies significantly. In contrast, site-specific integration provides a means to generate more consistent clones, reducing the cell line development timeline. Therefore, site-specific integration has become a promising strategy by which CHO cell lines can be reproducibly targeted to their preferred genomic sites for highly active and stable expression. However, in the actual use of site-specific integration, expression results still need to be improved and the problem of high mismatch rate still exists.
[0005] Summary of the Invention
[0006] The first object of the present application is to provide a cell line.
[0007] The second purpose of this application is to provide a method for preparing the above-mentioned cell line.
[0008] The third object of this application is to provide the application of the above cell line.
[0009] In order to achieve the above purpose, this application adopts the following technical solutions:
[0010] A cell line comprising a highly expressed fragment and an exogenous nucleic acid fragment integrated into the highly expressed fragment;
[0011] The highly expressed fragment comprises the nucleotide sequence shown in SEQ ID No. 1;
[0012] The exogenous nucleic acid fragment comprises a first target gene encoding a first protein fragment and a second target gene encoding a second protein fragment. The expression difficulty of the first protein fragment is higher than that of the second protein fragment. The first target gene is located upstream of the second target gene.
[0013] Furthermore, the exogenous nucleic acid fragment also includes a third target gene encoding a third protein fragment and a fourth target gene encoding a fourth protein fragment, the expression difficulty of the third protein fragment is higher than that of the fourth protein fragment, and the third target gene is located upstream of the fourth target gene, wherein the first target gene, the second target gene, the third target gene and the fourth target gene jointly encode a bispecific antibody or an antigen-binding fragment thereof.
[0014] Furthermore, the second target gene is located upstream of the third target gene.
[0015] Furthermore, the first protein fragment is a first light chain, the second protein fragment is a second light chain, the third protein fragment is a first heavy chain, and the fourth protein fragment is a second heavy chain; optionally, the protein bound to the first light chain and the first heavy chain can specifically bind to a first target; the protein bound to the second light chain and the second heavy chain can specifically bind to a second target.
[0016] Furthermore, the sequence of the first target gene is shown as SEQ ID No.2, the sequence of the third target gene is shown as SEQ ID No.3, the sequence of the second target gene is shown as SEQ ID No.4, and the sequence of the fourth target gene is shown as SEQ ID No.5.
[0017] Furthermore, the upstream of the first target gene and the second target gene respectively contain a nucleotide sequence encoding the first signal peptide shown in SEQ ID No. 6, and the upstream of the third target gene and the fourth target gene respectively contain a nucleotide sequence encoding the second signal peptide shown in SEQ ID No. 7.
[0018] Furthermore, the integration site of the exogenous nucleic acid fragment is any site within the 11th to 430th base interval of the highly expressed fragment;
[0019] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 21st to 414th base interval of the highly expressed fragment;
[0020] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 38th to 402nd base interval of the highly expressed fragment;
[0021] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 53rd to 389th base interval of the highly expressed fragment;
[0022] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 73rd to 373rd base interval of the highly expressed fragment;
[0023] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 91st to 360th base interval of the highly expressed fragment;
[0024] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 108th to 342nd base interval of the highly expressed fragment;
[0025] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 126th to 326th base interval of the highly expressed fragment;
[0026] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 143rd to 310th base interval of the highly expressed fragment;
[0027] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 160th to 295th base interval of the highly expressed fragment;
[0028] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 178th to 274th base interval of the highly expressed fragment;
[0029] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 194th to 263rd base interval of the highly expressed fragment;
[0030] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 209th to 253rd base interval of the highly expressed fragment;
[0031] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 221st to 242nd base interval of the highly expressed fragment;
[0032] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 231st to 240th base interval of the highly expressed fragment;
[0033] Optionally, the integration site of the exogenous nucleic acid fragment is the position shown in the annotation information NW_003616785.1:83044 on the CHO cell.
[0034] Furthermore, the cell line is a CHO cell.
[0035] A method for preparing a cell line comprises the following steps: integrating the exogenous nucleic acid fragment into the highly expressed fragment to obtain the cell line.
[0036] Furthermore, the preparation method comprises the following steps:
[0037] S1: An exogenous nucleic acid fragment is provided, wherein the exogenous nucleic acid fragment includes, in the 5'-3' direction, a first target gene encoding a first light chain, a second target gene encoding a second light chain, a third target gene encoding a first heavy chain, and a fourth target gene encoding a second heavy chain, wherein the expression difficulty of the first target gene is higher than that of the second target gene, and the expression difficulty of the third target gene is higher than that of the fourth target gene; and the first target gene, the second target gene, the third target gene, and the fourth target gene are each independently provided with a nucleotide sequence encoding a signal peptide upstream;
[0038] The first target gene, the second target gene, the third target gene and the fourth target gene jointly encode a bispecific antibody or an antigen-binding fragment thereof;
[0039] S2. Integrate the exogenous nucleic acid fragment into the highly expressed fragment to obtain a cell line.
[0040] Application of the above cell lines in protein expression.
[0041] Compared with the prior art, the technical effects of this application are:
[0042] The present application screens and analyzes gene sites within CHO cells to obtain insertion sites and highly expressed fragments suitable for site-specific recombination (or site-specific integration) technology. Further research reveals that in the case of multiple target proteins, when the gene encoding a protein that is relatively difficult to express is located upstream of the gene encoding a protein that is relatively easy to express, that is, by limiting the order of the target protein genes, high expression of the target protein, balanced expression of multiple target proteins, and reduced mismatch rate of the target protein can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0044] FIG1 is a flowchart of the operation of screening highly expressed fragments in Example 1;
[0045] Figure 2 is a schematic diagram of the structure of the screening site plasmid in step (1) of Example 1;
[0046] Figure 3 is a schematic diagram of the structure of the pDonor0.0-H plasmid in step (1) of Example 2;
[0047] Figure 4 is a schematic diagram of the structure of the pDonor0.0-L plasmid in step (1) of Example 2;
[0048] Figure 5 is a schematic diagram of the structure of the pLangding0.0 plasmid in step (1) of Example 2;
[0049] Figure 6 is a schematic diagram of the structure of the pLanding0.0-AL-VL-AH-VH plasmid in step (1) of Example 2;
[0050] Figure 7 is a schematic diagram of the structure of the pLanding0.0-AL-VL-VH-AH plasmid in step (1) of Example 2;
[0051] Figure 8 is a schematic diagram of the structure of the pLanding0.0-VL-AL-VH-AH plasmid in step (1) of Example 2;
[0052] Figure 9 is a schematic diagram of the structure of the pLanding0.0-VL-AL-AH-VH plasmid in step (1) of Example 2;
[0053] Figure 10 is a map of the pGBB Bxb-1 integrase plasmid in step (2) of Example 2;
[0054] Figure 11 is a diagram showing the results of SDS-PAGE analysis in step (6) of Example 2. DETAILED DESCRIPTION
[0055] definition
[0056] In this application, “further”, “further”, “particularly”, etc. are used for descriptive purposes to indicate differences in content, but should not be understood as limiting the scope of protection of the present invention.
[0057] In this application, the terms "optionally," "optional," and "optional" mean optional or dispensable, i.e., they refer to either option being selected from two parallel options: "with" or "without." If a technical solution contains multiple "optional" clauses, each "optional" clause is independent unless otherwise specified and there are no contradictions or constraints.
[0058] In this application, "plurality", "multiple", "multiple times", "multiples", etc., unless otherwise specified, refer to a quantity greater than or equal to 2. For example, "one or more" means one or more than or equal to two.
[0059] The phrase "highly expressed fragment" refers to a nucleotide sequence in the cell genome, which, when a suitable gene or construct is exogenously added (i.e., integrated) into the nucleotide sequence, exhibits a higher level of expression than other regions or sequences in the genome, particularly high expression at the protein level.
[0060] The phrase "exogenous nucleic acid fragment" refers to any DNA sequence or gene that is not found in a highly expressed fragment found in nature. For example, the "exogenous nucleic acid fragment" in a highly expressed fragment of CHO (e.g., a highly expressed fragment comprising SEQ ID No: 1) can be a hamster gene that is not found in the highly expressed fragment of a particular CHO in nature (i.e., the hamster gene is derived from another nucleotide sequence in the hamster genome), a gene from any other species (e.g., a human gene), a chimeric gene (e.g., human / mouse), or any other gene not found in nature that is present in the highly expressed fragment of the target CHO.
[0061] The phrase "difficult to express" refers to proteins that are difficult to produce. For example, in the same expression environment, the products of "difficult to express" proteins are more difficult to produce than other proteins during transcription, translation, or folding. For example, in the examples of this application, the constant region CL of the ANG-2 antibody light chain sequence and the constant region CH1 of the ANG-2 antibody heavy chain sequence are interchanged. Because artificial sequence swaps are more likely to produce mismatches, the expression difficulty of the ANG-2 antibody light chain sequence is higher than that of the VEGF antibody light chain sequence, and the expression difficulty of the ANG-2 antibody heavy chain sequence is higher than that of the VEGF antibody heavy chain sequence.
[0062] The phrase "antigen-binding fragment" includes proteins having at least one CDR and capable of selectively recognizing an antigen, ie, capable of binding an antigen with a KD of at least in the micromolar range.
[0063] The phrase "bispecific antibody" (bsAb) includes antibodies that can selectively bind to two epitopes. Bispecific antibodies generally comprise two different heavy chains, wherein each heavy chain specifically binds a different epitope—on two different molecules (e.g., antigens) or on the same molecule (e.g., on the same antigen).
[0064] Bispecific antibodies come in a variety of formats, ranging from Amgen's BiTE technology platform, consisting of only two tandem antigen-binding fragments, to IgG-like immunoglobulins containing additional domains, such as Kangfang Biopharma's already marketed candilizumab. Currently, over 20 different commercial technology platforms are available for the creation and development of bispecific antibodies, with more than a dozen already marketed and over 100 in clinical development. However, identifying a bispecific antibody-expressing cell line that meets the expected quality and is suitable for later-stage production processes remains a major pain point for the industry. Due to the unique structural characteristics of bispecific antibodies, the design structures of different platforms vary significantly. Consequently, the production process for bispecific antibodies presents significant challenges. A superior cell line serves as the starting point for the entire production process and largely determines the complexity of subsequent processes and the subsequent production costs of the marketed product. Therefore, selecting an excellent cell line is crucial for the entire pharmaceutical industry.
[0065] The phrase "site-specific integration" (also known as "site-specific recombination" or "site-specific recombination") refers to a gene targeting method used to direct the insertion or integration of a gene or nucleic acid sequence into a specific location in the genome, that is, to direct DNA to a specific site between two nucleotides in a contiguous polynucleotide chain.
[0066] The applicant has discovered a CHO nucleotide sequence that can stably and highly express a target protein, and has used it as a highly expressed fragment of the target protein. Furthermore, through site-specific integration technology, it was found that inserting (or integrating) the nucleotide sequence encoding the target protein as an exogenous nucleic acid fragment into the highly expressed fragment can greatly reduce the construction cost of the cell line while stably and highly expressing the target protein. Further research has found that in site-specific integration technology, when the nucleotide sequence encoding the target protein involves multiple nucleotide sequences, the order of arrangement of the multiple nucleotide sequences will affect the expression level of each encoded protein and also affect the mismatch rate of the target protein, thereby proposing the present application.
[0067] The present application discloses a cell line, which comprises a highly expressed fragment and an exogenous nucleic acid fragment integrated into the highly expressed fragment; the highly expressed fragment comprises the nucleotide sequence shown in SEQ ID No. 1; the exogenous nucleic acid fragment comprises a first target gene encoding a first protein fragment and a second target gene encoding a second protein fragment, the expression difficulty of the first protein fragment is higher than that of the second protein fragment, and the first target gene is located upstream of the second target gene.
[0068] In the present application, the integration site of the exogenous nucleic acid fragment can correspond to positions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, ... 444 of the highly expressed fragment.
[0069] The above-mentioned "cell line" includes any cell suitable for expressing a recombinant nucleic acid sequence and having a highly expressed fragment that allows stable integration and enhanced expression of an exogenous nucleic acid fragment. The cell includes a mammalian cell, such as a non-human animal cell, a human cell, or a cell fusion, such as a hybridoma or a quadroma. In some embodiments, the cell is a human, monkey, ape, hamster, rat, or mouse cell. In some embodiments, the cell is a mammalian cell selected from the following cells: CHO (e.g., CHO K1, etc.), COS (e.g., COS-7), retinal cells, Vero, CV1, kidney (e.g., HEK293, 293EBNA, MDCK, etc.), HeLa, HepG2, myeloma cells, tumor cells. Specifically, the cell line contains a highly expressed fragment of the nucleotide sequence shown in SEQ ID No. 1. On this basis, the exogenous nucleic acid fragment is integrated into the highly expressed fragment to obtain a cell line that stably and highly expresses the target protein. Preferably, the cell line is a CHO cell.
[0070] The exogenous nucleic acid fragment encodes the target protein. When the target protein is composed of multiple parts (for example, a single-chain antibody is mainly composed of a heavy chain variable region and a light chain variable region), the protein coding sequence that is more difficult to express is preferably set upstream of the protein coding sequence that is less difficult to express. This setting is beneficial to the balanced expression of multiple parts, the expression of the target protein and the reduction of the mismatch rate.
[0071] In some embodiments, the exogenous nucleic acid fragments comprise a first target gene encoding a first protein fragment, a second target gene encoding a second protein fragment, a third target gene encoding a third protein fragment, and a fourth target gene encoding a fourth protein fragment. The first protein fragment is more difficult to express than the second protein fragment, and the first target gene is located upstream of the second target gene; the third protein fragment is more difficult to express than the fourth protein fragment, and the third target gene is located upstream of the fourth target gene. Furthermore, the first, second, third, and fourth target genes collectively encode a bispecific antibody or antigen-binding fragment thereof.
[0072] From the above, in order to balance the expression levels of multiple parts, facilitate the expression level of the target protein and reduce the mismatch rate, the protein coding sequence with higher expression difficulty is preferably set upstream of the protein coding sequence with lower expression difficulty.
[0073] When the target protein is a bispecific antibody or its antigen-binding fragment, since the expression of the heavy chain depends on the correct folding of the light chain, and the light chain can assist in the secretion of the heavy chain, it is preferable to place the sequence encoding the light chain upstream to facilitate the production of more light chain antibody fragments, thereby further increasing the yield of the target protein. Furthermore, in the light chain or heavy chain portion, the protein coding sequence with higher expression difficulty is preferably placed upstream of the protein coding sequence with lower expression difficulty, further increasing the expression level of the bispecific antibody or its antigen-binding fragment and reducing the mismatch rate.
[0074] In some embodiments, further, the second target gene is located upstream of the third target gene.
[0075] In specific embodiments, the structure of a bispecific antibody or antigen-binding fragment thereof is as follows: the first protein fragment is a first light chain, the second protein fragment is a second light chain, the third protein fragment is a first heavy chain, and the fourth protein fragment is a second heavy chain. In some embodiments, the protein bound by the first light chain and the first heavy chain can specifically bind to a first target; and the protein bound by the second light chain and the second heavy chain can specifically bind to a second target.
[0076] In a more specific embodiment, the sequence of the first target gene is shown as SEQ ID No. 2, the sequence of the third target gene is shown as SEQ ID No. 3, the sequence of the second target gene is shown as SEQ ID No. 4, and the sequence of the fourth target gene is shown as SEQ ID No. 5. The target protein here is an ANG-2 / VEGF bispecific bivalent antibody, comprising a heavy chain sequence (encoding sequence: SEQ ID No. 3) and a light chain sequence (encoding sequence: SEQ ID No. 2) corresponding to the ANG-2 target, and a heavy chain sequence (encoding sequence: SEQ ID No. 5) and a light chain sequence (encoding sequence: SEQ ID No. 4) corresponding to the VEGF target. The constant region CL of the ANG-2 antibody light chain sequence is replaced with the constant region CH1 of the ANG-2 antibody heavy chain sequence. Therefore, the expression difficulty of the ANG-2 antibody light chain sequence (encoding sequence: SEQ ID No. 2) is higher than that of the VEGF antibody light chain sequence (encoding sequence: SEQ ID No. 4), and the expression difficulty of the ANG-2 antibody heavy chain sequence (encoding sequence: SEQ ID No. 3) is higher than that of the VEGF antibody heavy chain sequence (encoding sequence: SEQ ID No. 5). Therefore, the sequence order integrated into the highly expressed fragment is as follows: the first light chain gene (encoding sequence: SEQ ID No. 2), which is difficult to express; the second light chain gene (encoding sequence: SEQ ID No. 4), which is easy to express; and the first heavy chain gene (encoding sequence: SEQ ID No. 5). No.3), and the easily expressed second heavy chain gene (SEQ ID No.5).
[0077] In some embodiments, the upstream of the first target gene and the second target gene respectively contain a nucleotide sequence encoding the first signal peptide shown in SEQ ID No. 6, and the upstream of the third target gene and the fourth target gene respectively contain a nucleotide sequence encoding the second signal peptide shown in SEQ ID No. 7.
[0078] In a preferred embodiment, the integration site of the exogenous nucleic acid fragment is any site within the 11th to 430th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 21st to 414th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 38th to 402nd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 53rd to 389th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 73rd to 373rd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 91st to 360th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 108th to 342th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 126th to 326th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 143rd to 310th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 160th to 295th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 178th to 274th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 194th to 263rd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 209th to 253rd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 221st to 242nd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 231st to 240th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is the position shown in the annotated information of CHO cells as NW_003616785.1:83044.
[0079] The present application also discloses a method for preparing a cell line. When the cell is a CHO cell, its genome contains the highly expressed fragment of the present application, and the exogenous nucleic acid fragment is site-specifically integrated into the highly expressed fragment to obtain a cell line.
[0080] When the cell is another cell line, its genome does not contain the highly expressed fragment of the present application. The highly expressed fragment can be first stably introduced into the cell, and then the exogenous nucleic acid fragment can be site-specifically integrated into the highly expressed fragment to obtain a cell line.
[0081] In some embodiments, the preparation method comprises the following steps:
[0082] S1. Setting an exogenous nucleic acid fragment, the exogenous nucleic acid fragment includes a first target gene and a second target gene, the expression difficulty of the first target gene is higher than that of the second target gene, and the first target gene is upstream of the second target gene;
[0083] S2. Integrate the exogenous nucleic acid fragment into the highly expressed fragment to obtain a cell line.
[0084] In some embodiments, the preparation method comprises the following steps:
[0085] S1. Set exogenous nucleic acid fragments, the exogenous nucleic acid fragments include a first target gene, a second target gene, a third target gene, and a fourth target gene, the first target gene is more difficult to express than the second target gene, and the first target gene is upstream of the second target gene; the third protein fragment is more difficult to express than the fourth protein fragment, and the third target gene is upstream of the fourth target gene;
[0086] S2. Integrate the exogenous nucleic acid fragment into the highly expressed fragment to obtain a cell line.
[0087] In some embodiments, the preparation method comprises the following steps:
[0088] S1: An exogenous nucleic acid fragment is provided, wherein the exogenous nucleic acid fragment includes, in the 5'-3' direction, a first target gene encoding a first light chain, a second target gene encoding a second light chain, a third target gene encoding a first heavy chain, and a fourth target gene encoding a second heavy chain, wherein the expression difficulty of the first target gene is higher than that of the second target gene, and the expression difficulty of the third target gene is higher than that of the fourth target gene; the first target gene, the second target gene, the third target gene, and the fourth target gene are each independently provided with a nucleotide sequence encoding a signal peptide upstream;
[0089] The first target gene, the second target gene, the third target gene, and the fourth target gene collectively encode a bispecific antibody or an antigen-binding fragment thereof;
[0090] S2. Integrate the exogenous nucleic acid fragment into the highly expressed fragment to obtain a cell line.
[0091] In a specific embodiment, the preparation method comprises the following steps:
[0092] S1: setting an exogenous nucleic acid fragment, wherein the exogenous nucleic acid fragment includes, in the 5'-3' direction, a nucleotide sequence encoding SEQ ID No.6, SEQ ID No.2, a nucleotide sequence encoding SEQ ID No.6, SEQ ID No.4, a nucleotide sequence encoding SEQ ID No.7, SEQ ID No.3, a nucleotide sequence encoding SEQ ID No.7, and SEQ ID No.5;
[0093] S2. Integrate the exogenous nucleic acid fragment into the highly expressed fragment to obtain a cell line.
[0094] This application also discloses the use of the above-mentioned cell lines for protein expression. In some embodiments, the protein comprises an antigen binder. In some embodiments, the antigen binder is an antibody. In some embodiments, the antigen binder is a bispecific antibody. In some embodiments, the bispecific antibody is a bivalent antibody.
[0095] The present invention is further described below by way of examples. Unless otherwise specified, the materials in the examples were prepared according to existing methods or purchased directly from the market.
[0096] Example 1: Screening of highly expressed fragments:
[0097] Specific consumables: Neon Resuspension Buffer R (Thermo Fisher), a resuspension buffer specifically for cell electroporation; E1 Buffer (Thermo Fisher); recovery medium consisting of 80% (v / v) EX-CELL CHO Cloning Medium (Sigma-Aldrich) and 20% (v / v) EX-CELL Advanced CHO Fed-batch Medium (Sigma-Aldrich) supplemented with 1% GlutaMAX (Thermo Fisher); expansion medium and subculture medium consisting of EX-CELL Advanced CHO Fed-batch Medium supplemented with 1% GlutaMAX (Thermo Fisher); pressurized medium 1 consisting of subculture medium supplemented with G418 at a final concentration of 800 μg / ml; pressurized medium 2 consisting of subculture medium supplemented with hygromycin at a final concentration of 250 μg / ml; conditioned medium consisting of the supernatant obtained by sterile filtration after inoculation of subculture medium with CHO-K1 cells for one day; and cloning medium consisting of 75% (v / v) EX-CELL Advanced CHO Fed-batch Medium (Sigma-Aldrich). CHO Cloning Medium, 20% (v / v) conditioned medium, and 5% (v / v) ClonaCell-CHO ACF Supplement were supplemented with 1% GlutaMAX. The basal medium in the fed-batch medium was EX-CELL Advanced CHO Fed-batch Medium supplemented with 1% GlutaMAX (ThermoFisher), and the feed medium was Cell Boost 7a / 7b (HyClone).
[0098] The operation process of this embodiment is shown in Figure 1, which mainly includes the following steps:
[0099] (1) Using the green fluorescent protein gene (EGFP) as a screening marker gene and the attP sequence as a homology arm, a recombinant plasmid containing RMCE (recombinase-mediated cassette exchange) was constructed, as shown in Figure 2; the constructed plasmid was linearized, and the linear DNA was purified and recovered; CHO-K1 cells were taken, centrifuged, and the supernatant was discarded; the cells were resuspended in 100 μL of R Buffer, a resuspension buffer specifically for cell electroporation; and electroporation was performed three times.
[0100] (2) Pressurization: the pressurization reagent is G418, the pressurization concentration is 800 μg / ml, and the cells are divided into minipools (i.e., 96-well plates) for culture.
[0101] (3) Check the plate after ten days. When a large number of fluorescent cell clusters are observed, they can be enriched. The one-well-to-one-well principle should be followed during enrichment.
[0102] (4) Observe the growth of enriched cells at any time and expand them when the coverage rate reaches more than 50%.
[0103] (5) After expansion to a shake flask, the cells were passaged three times to form a stable cell pool. The formation of monoclonal cells and screening for stable fluorescence were achieved through artificial intelligence. The steps are shown in Figure 1, and the specific steps are shown in a)-h).
[0104] a) The cells in the stable cell pool are diluted to a certain concentration and then inoculated into a culture dish containing a semi-solid culture medium. During the inoculation process, the cells are ensured to be approximately evenly distributed in all positions of the culture dish. The cells are allowed to stand for about half an hour to allow the cells to settle to the bottom of the culture dish.
[0105] b) The culture dish is transferred to the electron microscope stage, and the electron microscope performs high-throughput scanning on the cells in the culture dish;
[0106] c) uploading the electron microscope scanned image to a server for artificial intelligence image analysis, the analysis process including monoclonal cell line detection, protein expression level prediction of the monoclonal cell line, protein expression level ranking, and coding and localization of the screened high protein expressing cell lines; the protein expression level prediction of the monoclonal cell line can be based on fluorescence or not, and step c is the protein expression level prediction not based on fluorescence;
[0107] Step c is based on the detection of monoclonal cell line targets using image processing technology. In this embodiment, the target detection algorithm of YOLOv8 is adopted, and the actual detection effect reaches mAP 94.2%. The target detection model for monoclonal cell lines in step c is consistent with the commonly used deep learning target detection model, so it is not described in detail. Specifically, the cell image needs to be annotated first. In order to improve the prediction accuracy of the algorithm, the bounding boxes of all monoclonal cell lines and adhesion cell lines will be marked and used as the real target bounding box (ground truth) for the loss calculation of the model output. After algorithm training, the model learns the ability to extract the border information of monoclonal cell lines. In the actual application scenario step c, the trained model can predict the border of the monoclonal cell line in the image.
[0108] Step c predicts the fluorescent protein expression level based on monoclonal cell images. In this example, the SqueezeNet deep learning network and the MSE loss function are used. Since the predicted protein expression levels in this project are ranked, the ranking result is the ultimate goal of the algorithm, so the NDCG evaluation standard of the ranking algorithm is used here. The specific calculation formula is as follows:
[0109] Where IDCG = best-ranked DCG. Specifically in this example, cell imaging predicted expression levels, sorted by predicted expression levels, and compared with the actual fluorescence value sorting results, the NDCG = 0.89. The expression level prediction model in this invention is described in detail in patent CN112037862B.
[0110] d) The coding and location information of the screened protein high-expressing cell lines are returned to the robot control software;
[0111] The cell codes and positions returned in step d, in this embodiment, return the top 100 cells in terms of predicted expression levels, with the codes ranging from 1 to 100. The position information includes the coordinates in the plane coordinate system relative to the center of the microscope (the depth of the culture medium is not considered for the time being, because all the cells to be selected are deposited to the bottom of the culture dish. In addition, during the photography process, cells that have not settled to the bottom of the culture dish will be out of focus, resulting in blurred cells, and cells with blurred images will also be excluded).
[0112] e) The robot control software automatically operates (or manually assists) the robotic arm to aspirate the selected monoclonal cell lines and transfer them to the designated well plate. This process is repeated until all high-expressing cell lines are transferred.
[0113] Step e: The process of the robot arm aspirating the target monoclonal cell line and transferring it to the well plate is described in detail in patents CN113821287B, CN113403431B, N113771030A, and CN113733087B.
[0114] f) After the cell clones in the well plate have grown to a certain number, they are transferred to a larger well plate for culture, and finally expanded to a shake flask for culture. The cell expression level is then tested to confirm that the cell line selected by the artificial intelligence is a high-yield cell line;
[0115] In this embodiment, step f is to transfer the cells to a 96-well plate for amplification and then transfer them to a shake flask for culture.
[0116] g) The selected high-yield cell lines are subcultured. Before each subculture, a diluted portion of the sample is placed in a culture dish and imaged using an electron microscope. Images of the cell lines are continuously taken for several generations. Currently, the expression of the selected cells can be estimated using a model learned from cell morphology by taking photos, or the expression of the selected cells can be fluorescently labeled.
[0117] Step g: Photographing Monoclonal Cells In this example, a Thermo Fisher Scientific M7000 electron microscope was used to photograph the cells.
[0118] h) Inputting the collected images of several generations of cell lines into an artificial intelligence algorithm, predicting the cell line's passage stability based on fluorescence intensity and / or protein expression, and selecting cell lines with high protein expression characteristics that can be stably passaged as the final candidate cell lines; the prediction of the cell line's passage stability can be based on fluorescence or on cell morphology rather than fluorescence. Step h is the prediction of protein expression based on cell morphology rather than fluorescence;
[0119] Step h predicts the cell line stability, including but not limited to using traditional image recognition algorithms to extract and predict image features using histograms, automatically extracting and predicting image features using deep learning neural networks, and other image-based prediction techniques. In this embodiment, a deep learning-based method was used to automatically extract image features and predict stability for stable and unstable cell lines, achieving an accuracy rate of 84% for predictions across different cell lines. This is described in detail in patent CN114417582A.
[0120] Wherein, step d and step e can be completed manually. In this embodiment, they are completed automatically by a robotic arm.
[0121] (6) CHO-K1 cells were used as control and the seeding density was 5×10 5 cells / ml, inoculate 30 ml of the system, count the cells on the day of inoculation, culture day 3, culture day 5, culture day 7, culture day 9, culture day 11, culture day 13, and culture day 14, and screen out monoclones with a growth rate lower than that of CHO-K1.
[0122] (7) Recover the remaining monoclonal clones and perform stable subculture after recovery. Subculture should be performed every 3 or 4 days. The cell density of the 4-day subculture is 3×10 5 cells / ml, and the cell density after 3 days was 5×10 5 cells / ml, and the fluorescence of the cell line was regularly monitored during subculture. The stability of the high-fluorescence cell line was studied for approximately 90 days using the subculture medium. GBB003 cells, which exhibited minimal fluorescence fluctuations and stable cell line growth, were confirmed to be stable high-fluorescence cells. NGS sequencing of the integration site of this cell line revealed the highly expressed fragment sequence located at the integration site as SEQ ID No. 1. Specifically, the annotation information for the integration site on the CHO gene is: NW_003616785.1:83044.
[0123] Example 2: Cell line construction
[0124] (1) Exogenous nucleic acid construction
[0125] The present application provides a method for constructing a multi-reading frame plasmid, as shown in Figure 3, wherein pDonor0.0-H is used to construct an expression reading frame vector for the heavy chain, and the target gene sequence can be inserted in the middle of the XmaI and / or HindIII restriction sites, wherein KanR provides a kanamycin resistance gene for screening in plasmid construction; the Ori segment is the plasmid replication initiation site, which is used for plasmid replication and amplification in prokaryotic cells; the CMV segment (not shown in the figure) is used to initiate the expression of the target gene; SV40 poly (A) is used to terminate the transcription termination of the target gene; NotI and other restriction sites are used to insert a complete reading frame into the pLangding0.0 plasmid for restriction enzyme cutting and connection. As shown in Figure 4 , pDonor0.0-L is used to construct a vector for the light chain expression reading frame, wherein KanR provides a kanamycin resistance gene for screening during plasmid construction; the Ori segment is the plasmid replication initiation site, used for transcriptional amplification of the plasmid in prokaryotes; the CMV segment (not shown) is used to initiate expression of the target gene; the SV40 poly(A) is used to terminate transcription of the target gene; XmaI and / or HindIII restriction sites are used to insert the target gene sequence; and other restriction sites such as NotI are used to enzyme-link and insert the complete reading frame into the pLangding0.0 plasmid. pLangding0.0 is used to accept the reading frames of pDonor0.0-H and / or pDonor0.0-L. The pLangding0.0 vector is shown in Figure 5 , wherein the Ori segment is the plasmid replication initiation site, used for transcriptional amplification of the plasmid in prokaryotes; and AmpR provides an ampicillin resistance gene for screening during plasmid construction.
[0126] The target genes in the constructed plasmid structures in this example are arranged and named as molecule No. 1, molecule No. 2, molecule No. 3, and molecule No. 4 respectively.
[0127] Molecular design No. 1 was constructed by synthetically inserting the Kozak sequence, signal peptide sequence (SEQ ID No. 6), and ANG-2 antibody light chain sequence (coding sequence SEQ ID No. 2) into the Hind III and Xma I restriction sites of the pDonor0.0-L plasmid to form the pDonor0.0-ANG-2-LC plasmid. The Kozak sequence, signal peptide sequence (SEQ ID No. 6), and VEGF antibody light chain sequence (coding sequence SEQ ID No. 4) were synthetically inserted into the Hind III and Xma I restriction sites of the pDonor0.0-L plasmid to form the pDonor0.0-VEGF-LC plasmid. The Kozak sequence, signal peptide sequence (SEQ ID No. 7), and ANG-2 antibody heavy chain sequence (coding sequence SEQ ID No. 3) were synthetically inserted into the Hind III and Xma I restriction sites of the pDonor0.0-H plasmid to form the pDonor0.0-ANG-2-HC plasmid. The Kozak sequence, signal peptide sequence (SEQ ID No. 7), and VEGF antibody heavy chain sequence (coding sequence SEQ ID No. 5) were synthetically inserted into the HindIII and XmaI restriction sites of the pDonor0.0-H plasmid to generate the pDonor0.0-VEGF-HC plasmid. The pDonor0.0-ANG-2-LC and pLanding0.0 plasmids were double-digested with NotI / AgeI, and the ANG-2-LC reading frame fragment of pDonor0.0-ANG-2-LC and the linearized pLanding0.0 fragment were recovered. These fragments were then ligated with T4 DNA ligase and transformed into DH5α chemically competent cells. Colonies were then plated onto LB ampicillin-resistant plates, washed, and harvested with LB medium. The plasmid was then extracted to generate the pLanding0.0-AL plasmid. After the pDonor0.0-VEGF-LC and pLanding0.0-AL plasmids were double-digested with BsiwI / MluI, the VEGF-LC reading frame fragment of pDonor0.0-VEGF-LC and the linearized pLanding0.0-AL fragment were recovered, and then ligated with T4 DNA ligase to transform StBl3 chemically competent cells. After colonies grew on LB ampicillin-resistant plates, the cells were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL.The pDonor0.0-ANG-2-HC and pLanding0.0-AL-VL plasmids were double-digested with NheI / XhoI, and the ANG-2-HC reading frame fragment of pDonor0.0-ANG-2-HC and the linearized pLanding0.0-AL-VL fragment were recovered. They were then ligated with T4 DNA ligase and transformed into StBl3 chemically competent cells. After colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL-AH. The pDonor0.0-VEGF-HC and pLanding0.0-AL-VL-AH plasmids were double-digested with EcoRI / KpnI, and the VEGF-HC reading frame fragment of pDonor0.0-VEGF-HC and the linearized pLanding0.0-AL-VL-AH fragment were recovered. These fragments were then ligated and transformed into StB13 chemically competent cells using T4 DNA ligase. After colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL-AH-VH. In this example, the ANG-2 antibody light chain sequence was the first target gene, the VEGF antibody light chain sequence was the second target gene, the ANG-2 antibody heavy chain sequence was the third target gene, and the VEGF antibody heavy chain sequence was the fourth target gene. A schematic diagram of the plasmids is shown in Figure 6.
[0128] The construction of molecular design No. 2 is to double-digest the pDonor0.0-ANG-2-LC and pLanding0.0 plasmids with Not I / Age I, respectively, and then recover the ANG-2-LC reading frame fragment of pDonor0.0-ANG-2-LC and the linearized pLanding0.0 fragment, and then use T4 DNA ligase to connect and transform DH5α chemically competent cells. After the colonies are grown on LB ampicillin-resistant plates, they are washed and collected with LB medium, and the plasmid is extracted to form pLanding0.0-AL. After the pDonor0.0-VEGF-LC and pLanding0.0-AL plasmids were double-digested with BsiwI / MluI, the VEGF-LC reading frame fragment of pDonor0.0-VEGF-LC and the linearized pLanding0.0-AL fragment were recovered, and then ligated with T4 DNA ligase to transform StBl3 chemically competent cells. After colonies grew on LB ampicillin-resistant plates, the cells were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL. The pDonor0.0-VEGF-HC and pLanding0.0-AL-VL plasmids were double-digested with NheI / XhoI, and the VEGF-HC reading frame fragment of pDonor0.0-VEGF-HC and the linearized pLanding0.0-AL-VL fragment were recovered. These fragments were then ligated with T4 DNA ligase and transformed into StB13 chemically competent cells. After colonies were grown on LB ampicillin-resistant plates, the cells were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL-VH. After double digestion with EcoRI / KpnI, the pDonor0.0-ANG-2-HC and pLanding0.0-AL-VL-VH plasmids were recovered from the ANG-2-HC reading frame fragment of pDonor0.0-ANG-2-HC and the linearized pLanding0.0-AL-VL-VH fragment. These fragments were then ligated and transformed into StB13 chemically competent cells using T4 DNA ligase. After colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL-VH-AH. In this example, the ANG-2 antibody light chain sequence was the first gene of interest, the VEGF antibody light chain sequence was the second gene of interest, the VEGF antibody heavy chain sequence was the third gene of interest, and the ANG-2 antibody heavy chain sequence was the fourth gene of interest. A schematic diagram of the plasmids is shown in Figure 7.
[0129] The construction of molecular design No. 3 is to double-digest the pDonor0.0-VEGF-LC and pLanding0.0 plasmids with NotI / AgeI respectively, recover the VEGF-LC reading frame fragment of pDonor0.0-VEGF-LC and the linearized pLanding0.0 fragment, and then use T4 DNA ligase to connect and transform DH5α chemically competent cells. After the colonies are grown on LB ampicillin-resistant plates, they are washed and collected with LB medium, and the plasmid is extracted to form pLanding0.0-VL. The pDonor0.0-ANG-2-LC and pLanding0.0-VL plasmids were double-digested with BsiwI / MluI, and the ANG-2-LC reading frame fragment of pDonor0.0-ANG-2-LC and the linearized pLanding0.0-VL fragment were recovered. They were then ligated with T4 DNA ligase and transformed into Stbl3 chemically competent cells. After colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL. After double digestion of the pDonor0.0-VEGF-HC and pLanding0.0-VL-AL plasmids with NheI / XhoI, the VEGF-HC reading frame fragment of pDonor0.0-VEGF-HC and the linearized pLanding0.0-VL-AL fragment were recovered. They were then ligated with T4 DNA ligase and transformed into StBl3 chemically competent cells. After colonies were grown on LB ampicillin-resistant plates, the cells were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL-VH. The pDonor0.0-ANG-2-HC and pLanding0.0-VL-AL-VH plasmids were double-digested with EcoRI / KpnI, and the ANG-2-HC reading frame fragment of pDonor0.0-ANG-2-HC and the linearized pLanding0.0-VL-AL-VH fragment were recovered. These fragments were then transformed into Stbl3 chemically competent cells using T4 DNA ligase. After colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL-VH-AH. In this example, the VEGF antibody light chain sequence was the first target gene, the ANG-2 antibody light chain sequence was the second target gene, the VEGF antibody heavy chain sequence was the third target gene, and the ANG-2 antibody heavy chain sequence was the fourth target gene. A schematic diagram of the plasmids is shown in Figure 8.
[0130] The construction of molecular design No. 4 is to double-digest the pDonor0.0-VEGF-LC and pLanding0.0 plasmids with NotI / AgeI respectively, recover the VEGF-LC reading frame fragment of pDonor0.0-VEGF-LC and the linearized pLanding0.0 fragment, and then use T4 DNA ligase to connect and transform DH5α chemically competent cells. After the colonies are grown on LB ampicillin-resistant plates, they are washed and collected with LB medium, and the plasmid is extracted to form pLanding0.0-VL. The pDonor0.0-ANG-2-LC and pLanding0.0-VL plasmids were double-digested with BsiwI / MluI, and the ANG-2-LC reading frame fragment of pDonor0.0-ANG-2-LC and the linearized pLanding0.0-VL fragment were recovered. These fragments were then ligated and transformed into StBl3 chemically competent cells using T4 DNA ligase. After colonies were grown on LB ampicillin-resistant plates, the cells were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL. The pDonor0.0-ANG-2-HC and pLanding0.0-VL-AL plasmids were double-digested with NheI / XhoI, and the ANG-2-HC reading frame fragment of pDonor0.0-ANG-2-HC and the linearized pLanding0.0-VL-AL fragment were recovered. They were then ligated with T4 DNA ligase and transformed into StBl3 chemically competent cells. After colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL-AH. After double digestion with EcoRI / KpnI, the pDonor0.0-VEGF-HC and pLanding0.0-VL-AL-AH plasmids were recovered from the VEGF-HC reading frame fragment of pDonor0.0-VEGF-HC and the linearized pLanding0.0-VL-AL-AH fragment. These fragments were then ligated and transformed into StB13 chemically competent cells using T4 DNA ligase. After colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL-AH-VH. In this example, the VEGF antibody light chain sequence was the first target gene, the ANG-2 antibody light chain sequence was the second target gene, the ANG-2 antibody heavy chain sequence was the third target gene, and the VEGF antibody heavy chain sequence was the fourth target gene. A schematic diagram of the plasmids is shown in Figure 9.
[0131] This example is used to express an ANG-2 / VEGF bispecific bivalent antibody. The antibody contains the heavy chain sequence (encoding sequence: SEQ ID No. 3) and light chain sequence (encoding sequence: SEQ ID No. 2) corresponding to the ANG-2 target, and the heavy chain sequence (encoding sequence: SEQ ID No. 5) and light chain sequence (encoding sequence: SEQ ID No. 4) corresponding to the VEGF target. All genes were synthesized by GenScript. Some information is as follows:
[0132] In the ANG-2 antibody light chain sequence, the coding sequence of the heavy chain constant region 1 (CH1) in the first light chain gene is SEQ ID No. 8.
[0133] In the ANG-2 antibody heavy chain sequence, the coding sequence of the light chain constant region (CL) within the first heavy chain gene is SEQ ID No. 9.
[0134] Wherein, a signal peptide sequence SEQ ID No.6 required for antibody secretion is provided before the antibody light chain sequence.
[0135] Wherein, a signal peptide sequence SEQ ID No. 7 required for antibody secretion is provided before the antibody heavy chain sequence.
[0136] In the ANG-2 / VEGF bispecific bivalent antibody formed after expression of the four molecules, the constant region CL of the ANG-2 antibody light chain sequence and the constant region CH1 of the ANG-2 antibody heavy chain sequence are replaced with each other. Therefore, the expression of the ANG-2 antibody light chain sequence is more difficult than that of the VEGF antibody light chain sequence, and the expression of the ANG-2 antibody heavy chain sequence is more difficult than that of the VEGF antibody heavy chain sequence.
[0137] (2) Stable high-fluorescence cell GBB003 was selected for transfection and integration verification. The recombinant plasmids expressing the ANG-2 / VEGF bispecific bivalent antibody genes in the four arrangements contained in step (1) above were amplified and linearized, and then co-transfected with Bxb-1 integrase into the stable high-fluorescence cell GBB003. The Bxb-1 integrase plasmid map is shown in FIG10 .
[0138] (3) One day after transfection, the cell pool was used to construct a minipool, and pressure screening was performed using pressurized medium 2;
[0139] (4) After 14 days, the non-fluorescent cells in the minipool were expanded and cultured in the following manner: 96-well plate, 24-well plate, and 6-well plate. When the cells were expanded to 6-well plate, 6-well batch culture was performed. 6-well batch culture was performed by inoculating 5×10 5 / ml, 2ml system, 37℃ 5% CO2 120 rpm condition culture, 6-well batch culture experimental results are shown in Table 1:
[0140] Table 1: Statistical results of expression levels in six-well batch culture of cell pools
[0141] (5) For each molecule, select the top 1 cell pool (i.e., 1-9 cell pool, 2-5 cell pool, 3-4 cell pool, 4-6 cell pool) and expand it to T125 shake flasks. After two subcultures, shake flasks are fed for 14 days. The inoculation density is 5×10 5 / ml, the inoculation system was 30ml, and the cells were counted on the day of inoculation, culture day 3, culture day 5, culture day 7, culture day 9, culture day 11, culture day 13, and culture day 14. On day 3, 3% Cell Boost 7a and 0.3% Cell Boost 7b were supplemented, and on culture day 5, culture day 7, culture day 9, culture day 11, and culture day 13, 5% Cell Boost 7a and 0.5% Cell Boost 7b were supplemented. During this period, the glucose concentration was controlled at 2-8g / L. The expression level was measured using Octect after 14 days. The expression results of the shake flask fed-batch expression are shown in Table 2.
[0142] Table 2: Statistical results of shake flask feeding expression
[0143] (6) The supernatant of the shake flask feed was purified by Protein A chromatography column and then subjected to SDS-PAGE analysis (the results are shown in Figure 11) and SEC purity analysis. The SEC purity is shown in Table 3. The SEC purity of molecule No. 1 is 83.47%, which is significantly better than other molecular designs. The ratio of each chain was analyzed by reducing CE-SDS. As shown in Table 4, the ratio of the two light chains and the two heavy chains of molecule No. 1 is closer to 1:1, which is more in line with the expected theoretical value. Comprehensive analysis shows that the purity and molecular expression of molecule No. 1 are the best, so the 1-9 cell pool of molecule No. 1 was selected for monocloning, and the 2-5, 3-4, and 4-6 cell pools were used as control groups for monocloning;
[0144] Table 3: SEC purity analysis results of cell pool
[0145] Table 4: CE-SDS purity analysis results of cell pool reduction
[0146] (7) Perform limiting dilution on cell pools 1-9, 2-5, 3-4, and 4-6, and plate a 96-well plate at a density of 0.8 cells / well. Use VIPS to photograph the cell growth status at 1 hour, 1 day, 2 days, 3 days, 5 days, 7 days, and 14 days, and count the non-fluorescent monoclones.
[0147] (8) Expand the monoclonal wells without fluorescence, and follow the method of 96-well plate-24-well plate-6-well plate-shake flask. In the process of expansion, the clones with low growth rate are eliminated. No monoclonal clones are formed in 4-6, and the cells do not grow and reproduce. Therefore, the cell pools 4-6 cannot be screened by six-well batch culture. After the 6-well is expanded, six-well batch culture is performed. The 6-well batch culture is to inoculate 5×10 5 / ml, 2ml system, 37°C 5% CO2 120 rpm conditions for 7 days, the top 5 clones expressing the best in each 6-well batch culture of each cell pool are shown in Table 5;
[0148] Table 5: Statistical results of expression levels of monoclonal six-well batch culture
[0149] (9) Select the top three cell lines with the highest expression in each cell pool and culture them in six-well batches. After two subcultures, perform shake flask feeding for 14 days. The seeding density is 5×10 5 / ml, inoculate 30ml of the system, count on the day of inoculation, culture day 3, culture day 5, culture day 7, culture day 9, culture day 11, culture day 13 and culture day 14, and supplement 3% Cell Boost 7a and 0.3% Cell Boost 7b on the 3rd day, 5% Cell Boost 7a and 0.5% Cell Boost 7b on the 5th day, culture day 7, culture day 9, culture day 11 and culture day 13, during which the glucose concentration was controlled at 2-8g / L. After 14 days, use The expression level was determined, and the results of monoclonal shake flask feeding expression are shown in Table 6;
[0150] Table 6: Statistical results of monoclonal shake flask feeding expression
[0151] The supernatant from the shake flask feed was captured using Protein A media and then analyzed by SEC. The results are shown in Table 7. The quality attributes of the monoclonal clones produced by each cell pool were relatively uniform, with the cell pool containing molecule No. 1 showing the highest SEC purity. Calculation of the target molecule yield showed that both 1-9-1 and 1-9-2 performed well, with a theoretical yield of approximately 4.7 g. The calculated results are shown in Table 8.
[0152] Table 7: Monoclonal SEC analysis results
[0153] Table 8: Statistics of relative yields of monoclonal target products
[0154] Through the above experiments, we identified a cell line that met our expectations for bispecific antibody expression. The experimental process is controllable, the workload is significantly reduced compared to traditional methods, and the antibody mismatch rate is lower. In the four ANG-2 / VEGF bivalent antibodies produced after expression, the constant region CL of the ANG-2 light chain sequence is swapped with the constant region CH1 of the ANG-2 heavy chain sequence. This makes expression of the ANG-2 light chain sequence more challenging than that of the VEGF light chain sequence, and vice versa. Furthermore, since artificial sequence swaps are more prone to mismatches, placing the ANG-2 expression reading frame at the front of the plasmid design minimizes the impact of the preceding reading frame. Since heavy chain expression in conventional antibody expression depends on the proper folding of the light chain, and the light chain assists in heavy chain secretion, placing the light chain reading frame at the front of the plasmid design allows for the production of more light chain antibody fragments, further increasing the yield of the target molecule. This explains the higher expression yield and lower mismatch rate of our design #1.
[0155] The exogenous nucleic acid fragment can be replaced. The first target gene can be replaced with a gene for encoding the first protein, the second target gene can be replaced with a gene for encoding the second fusion protein, and the expression difficulty of the first protein is higher than that of the second protein. The first target gene can be replaced with a gene for encoding a monoclonal antibody light chain, and the second target gene can be replaced with a gene for encoding a monoclonal antibody heavy chain, and the expression difficulty of the light chain is higher than that of the heavy chain. The first target gene can be replaced with a gene for encoding a monoclonal antibody heavy chain, and the second target gene can be replaced with a gene for encoding a monoclonal antibody light chain, and the expression difficulty of the heavy chain is higher than that of the light chain. There can be multiple target genes in the exogenous nucleic acid fragment, and the difficulty of expressing proteins in each target gene segment is different. The target gene with high expression difficulty is located before the target gene with low expression in the exogenous nucleic acid fragment.
[0156] Note that the above are only preferred embodiments of the present application and the technical principles employed. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the technical concept of the present application, all of which fall within the scope of protection of the present application.
[0157] Those skilled in the art will also recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the appended claims.
Claims
1. A cell line, characterized in that The cell line comprises a highly expressed fragment and an exogenous nucleic acid fragment integrated into the highly expressed fragment; The highly expressed fragment comprises the nucleotide sequence shown in SEQ ID No.1; The exogenous nucleic acid fragments include a first target gene encoding a first protein fragment and a second target gene encoding a second protein fragment. The expression difficulty of the first protein fragment is higher than that of the second protein fragment. The first target gene is located upstream of the second target gene.
2. The cell line according to claim 1, characterized in that The exogenous nucleic acid fragment further comprises a third target gene encoding a third protein fragment and a fourth target gene encoding a fourth protein fragment, the expression difficulty of the third protein fragment is higher than that of the fourth protein fragment, the third target gene is located upstream of the fourth target gene, wherein the first target gene, the second target gene, the third target gene and the fourth target gene jointly encode a bispecific antibody or an antigen-binding fragment thereof.
3. The cell line according to claim 2, characterized in that The second target gene is located upstream of the third target gene.
4. The cell line according to claim 3, characterized in that The first protein fragment is a first light chain, the second protein fragment is a second light chain, the third protein fragment is a first heavy chain, and the fourth protein fragment is a second heavy chain; optionally, the protein bound by the first light chain and the first heavy chain can specifically bind to a first target; the protein bound by the second light chain and the second heavy chain can specifically bind to a second target.
5. The cell line according to claim 4, characterized in that The sequence of the first target gene is shown as SEQ ID No.2, the sequence of the third target gene is shown as SEQ ID No.3, the sequence of the second target gene is shown as SEQ ID No.4, and the sequence of the fourth target gene is shown as SEQ ID No.
5.
6. The cell line according to claim 5, characterized in that The upstream of the first target gene and the second target gene respectively contain nucleotide sequences encoding the first signal peptide shown in SEQ ID No.6, and the upstream of the third target gene and the fourth target gene respectively contain nucleotide sequences encoding the second signal peptide shown in SEQ ID No.
7.
7. The cell line according to any one of claims 1 to 6, characterized in that The integration site of the exogenous nucleic acid fragment is any site within the 11th to 430th base interval of the highly expressed fragment; The integration site of the exogenous nucleic acid fragment is the position shown in the annotation information on the CHO cell as NW_003616785.1:83044.
8. The cell line according to claim 7, characterized in that The cell line is CHO cell.
9. A method for preparing the cell line according to any one of claims 1 to 8, characterized in that: The preparation method comprises the following steps: integrating the exogenous nucleic acid fragment into the highly expressed fragment at a fixed point to obtain the cell line.
10. The preparation method according to claim 9, characterized in that: The preparation method comprises the following steps: S1: An exogenous nucleic acid fragment is provided, wherein the exogenous nucleic acid fragment includes, in the 5'-3' direction, a first target gene encoding a first light chain, a second target gene encoding a second light chain, a third target gene encoding a first heavy chain, and a fourth target gene encoding a second heavy chain, wherein the expression difficulty of the first target gene is higher than that of the second target gene, and the expression difficulty of the third target gene is higher than that of the fourth target gene; the first target gene, the second target gene, the third target gene, and the fourth target gene are independently provided with a nucleotide sequence encoding a signal peptide upstream; The first target gene, the second target gene, the third target gene and the fourth target gene jointly encode a bispecific antibody or an antigen-binding fragment thereof; S2. The exogenous nucleic acid fragment is site-specifically integrated into the highly expressed fragment to obtain a cell line.
11. Use of the cell line according to any one of claims 1 to 8 in protein expression.
Citation Information
Patent Citations
Cell screening method and apparatus based on convolutional neural networks
CN112037862B
Robot-based cell fluid extraction control methods, devices, equipment, and storage media
CN113403431B
Control information configuration methods, devices, equipment, and media for cell manipulation robots
CN113733087B
Control method, device and equipment of cell operation robot and storage medium
CN113771030A
Robot-based methods, devices, equipment, and media for cell manipulation tasks.
CN113821287B