Cell strain as well as preparation method and application thereof
By integrating exogenous nucleic acid fragments into high-expression fragments in CHO cells, optimizing gene sequence, the problem of inconsistent cell bank stability and expression results in CHO cells was solved, and efficient and balanced protein expression and reduced mismatch rate were achieved.
Patent Information
- Application Number
- CN202311484805.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-08
- Publication Date
- 2025-05-09
AI Technical Summary
During the production process of CHO cells, the prior art uses non-targeted transgenic integration methods to cause inconsistent stability and expression results of the cell bank, and the mismatch rate is high, affecting production efficiency.
By integrating exogenous nucleic acid fragments into high-expression fragments in CHO cells, the sequence of target genes is optimized, and genes with high expression difficulty are placed upstream of genes with low expression difficulty, so as to achieve the purpose of high expression and balanced expression.
This method significantly increases the expression of the target protein, reduces the mismatch rate, shortens the time for cell line construction, and improves production efficiency.
Smart Images

Figure CN119955730A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of biotechnology, and specifically provides a cell line and a preparation method and application thereof. Background Art
[0002] Chinese hamster ovary (CHO) cells are expression hosts for commercial production of recombinant proteins. Currently, a variety of biotherapeutic drug combinations produced by CHO cells have been approved for marketing, including a variety of monoclonal antibodies (mAbs) and other proteins. Since the emergence of CHO cell culture production processes more than 30 years ago, those skilled in the art have made unremitting efforts to significantly improve the production capacity of CHO cells. In particular, the titers of industrial antibody production in the past were usually far below 1g / L. However, the production of antibodies by manufacturing teams now achieves titers above grams / liter, usually above 3g / L. However, at present, it is still customary to use non-targeted transgenic integration methods in the upstream cell line construction and screening stage to produce a stably transfected CHO cell library, from which a production cell line is established. Therefore, it is necessary to identify those clones with suitable production characteristics through tedious screening in a cell library full of heterogeneity. Since this type of screened cell lines are usually multi-copy, there are great differences in the stability of different sites of the cell line. In contrast, site-specific integration provides a means to produce more consistent clones, and this technology reduces the cell line development timeline. Therefore, site-specific integration has become a promising strategy by which CHO cell lines can be repeatedly targeted to their preferred genomic sites that can be highly active and stably expressed. However, in the actual use of site-specific integration, the expression results still need to be improved and there is still a high mismatch rate problem.
[0003] In view of this, this application is hereby filed. Summary of the invention
[0004] The first objective of the present application is to provide a cell line.
[0005] The second objective of the present application is to provide a method for preparing the above-mentioned cell line.
[0006] The third objective of the present application is to provide the application of the above cell line.
[0007] In order to achieve the above purpose, this application adopts the following technical solutions:
[0008] A cell line comprising a highly expressed fragment and an exogenous nucleic acid fragment integrated into the highly expressed fragment;
[0009] The highly expressed fragment comprises the nucleotide sequence shown in SEQ ID No.1;
[0010] The exogenous nucleic acid fragments include a first target gene encoding a first protein fragment and a second target gene encoding a second protein fragment. The expression difficulty of the first protein fragment is higher than that of the second protein fragment. The first target gene is located upstream of the second target gene.
[0011] Furthermore, the exogenous nucleic acid fragment also includes a third target gene encoding a third protein fragment and a fourth target gene encoding a fourth protein fragment, the expression difficulty of the third protein fragment is higher than that of the fourth protein fragment, and the third target gene is located upstream of the fourth target gene, wherein the first target gene, the second target gene, the third target gene and the fourth target gene jointly encode a bispecific antibody or an antigen-binding fragment thereof.
[0012] Furthermore, the second target gene is located upstream of the third target gene.
[0013] Furthermore, the first protein fragment is a first light chain, the second protein fragment is a second light chain, the third protein fragment is a first heavy chain, and the fourth protein fragment is a second heavy chain; optionally, the protein bound to the first light chain and the first heavy chain can specifically bind to a first target; the protein bound to the second light chain and the second heavy chain can specifically bind to a second target.
[0014] Furthermore, the sequence of the first target gene is shown as SEQ ID No.2, the sequence of the third target gene is shown as SEQ ID No.3, the sequence of the second target gene is shown as SEQ ID No.4, and the sequence of the fourth target gene is shown as SEQ ID No.5.
[0015] Furthermore, the upstream of the first target gene and the second target gene respectively contains a nucleotide sequence encoding the first signal peptide shown in SEQ ID No. 6, and the upstream of the third target gene and the fourth target gene respectively contains a nucleotide sequence encoding the second signal peptide shown in SEQ ID No. 7.
[0016] Furthermore, the integration site of the exogenous nucleic acid fragment is any site within the 11th to 430th base interval of the highly expressed fragment;
[0017] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 21st to 414th base interval of the highly expressed fragment;
[0018] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 38th to 402nd base interval of the highly expressed fragment;
[0019] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 53rd to 389th base interval of the highly expressed fragment;
[0020] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 73rd to 373rd base interval of the highly expressed fragment;
[0021] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 91st to 360th base interval of the highly expressed fragment;
[0022] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 108th to 342nd base interval of the highly expressed fragment;
[0023] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 126th to 326th base interval of the highly expressed fragment;
[0024] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 143rd to 310th base interval of the highly expressed fragment;
[0025] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 160th to 295th base interval of the highly expressed fragment;
[0026] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 178th to 274th base interval of the highly expressed fragment;
[0027] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 194th to 263rd base interval of the highly expressed fragment;
[0028] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 209th to 253rd base interval of the highly expressed fragment;
[0029] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 221st to 242nd base interval of the highly expressed fragment;
[0030] Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 231st to 240th base interval of the highly expressed fragment;
[0031] Optionally, the integration site of the exogenous nucleic acid fragment is the position shown in the annotation information on the CHO cell as NW_003616785.1:83044.
[0032] Furthermore, the cell line is a CHO cell.
[0033] A method for preparing a cell line comprises the following steps: integrating the exogenous nucleic acid fragment into the highly expressed fragment to obtain the cell line.
[0034] Furthermore, the preparation method comprises the following steps:
[0035] S1: An exogenous nucleic acid fragment is provided, wherein the exogenous nucleic acid fragment includes, in the 5'-3' direction, a first target gene encoding a first light chain, a second target gene encoding a second light chain, a third target gene encoding a first heavy chain, and a fourth target gene encoding a second heavy chain, wherein the expression difficulty of the first target gene is higher than that of the second target gene, and the expression difficulty of the third target gene is higher than that of the fourth target gene; the first target gene, the second target gene, the third target gene, and the fourth target gene are independently provided with a nucleotide sequence encoding a signal peptide upstream;
[0036] The first target gene, the second target gene, the third target gene and the fourth target gene jointly encode a bispecific antibody or an antigen-binding fragment thereof;
[0037] S2. The exogenous nucleic acid fragment is site-specifically integrated into the highly expressed fragment to obtain a cell line.
[0038] Application of the above cell lines in protein expression.
[0039] Compared with the prior art, the technical effects of this application are:
[0040] The present application obtains insertion sites and high expression fragments suitable for site-specific recombination (or site-specific integration) technology by screening and analyzing gene sites in CHO cells. Further research has found that in the case of multiple target proteins, when the gene encoding a protein that is relatively difficult to express is located upstream of the gene encoding a protein that is relatively easy to express, that is, by limiting the order of the target protein gene, high expression of the target protein, balanced expression of multiple target proteins, and reduction of the mismatch rate of the target protein can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0042] Figure 1 is an operational flow chart of screening highly expressed fragments in Example 1;
[0043] Figure 2 is a schematic diagram of the structure of the screening site plasmid in step (1) of Example 1;
[0044] Figure 3 is a schematic diagram of the structure of the pDonor0.0-H plasmid in step (1) of Example 2;
[0045] Figure 4 is a schematic diagram of the structure of the pDonor0.0-L plasmid in step (1) of Example 2;
[0046] Figure 5 is a schematic diagram of the structure of the pLangding0.0 plasmid in step (1) of Example 2;
[0047] Figure 6 is a schematic diagram of the structure of the pLanding0.0-AL-VL-AH-VH plasmid in step (1) of Example 2;
[0048] Figure 7 is a schematic diagram of the structure of the pLanding0.0-AL-VL-VH-AH plasmid in step (1) of Example 2;
[0049] Figure 8 is a schematic diagram of the structure of the pLanding0.0-VL-AL-VH-AH plasmid in step (1) of Example 2;
[0050] Fig. 9 is a schematic diagram of the structure of the pLanding0.0-VL-AL-AH-VH plasmid in step (1) of Example 2;
[0051] Fig.10 is the pGBB Bxb-1 integrase plasmid map in step (2) of Example 2;
[0052] Fig.11 This is a diagram of the SDS-PAGE analysis results in step (6) of Example 2. DETAILED DESCRIPTION
[0053] definition
[0054] In the present application, “further”, “furthermore”, “particularly”, etc. are used for descriptive purposes to indicate differences in content, but should not be construed as limiting the scope of protection of the present invention.
[0055] In this application, "optionally", "optional", and "optional" mean optional or dispensable, that is, any one of the two parallel schemes of "yes" or "no". If multiple "options" appear in a technical solution, unless otherwise specified and there is no contradiction or mutual restriction, each "optional" is independent.
[0056] In the present application, "plurality", "multiple", "multiple times", "multiples", etc., unless otherwise specified, refer to a number greater than 2 or equal to 2. For example, "one or more" means one or greater than or equal to two.
[0057] The phrase "highly expressed fragment" refers to a nucleotide sequence in the cell genome, which, when a suitable gene or construct is exogenously added (i.e., integrated) into the nucleotide sequence, exhibits a higher level of expression than other regions or sequences in the genome, particularly high expression at the protein level.
[0058] The phrase "exogenous nucleic acid fragment" refers to any DNA sequence or gene that is not present in the highly expressed fragment found in nature. For example, the "exogenous nucleic acid fragment" in the highly expressed fragment of CHO (e.g., the highly expressed fragment comprising the sequence of SEQ ID No: 1) can be a hamster gene that is not found in the highly expressed fragment of a particular CHO in nature (i.e., the hamster gene is from another nucleotide sequence of the hamster genome), a gene from any other species (e.g., a human gene), a chimeric gene (e.g., human / mouse), or any other gene that is not found in nature and is present in the highly expressed fragment of the target CHO.
[0059] The phrase "high difficulty in expression" refers to a protein that is difficult to produce. For example, in the same expression environment, the product of a protein with "high difficulty in expression" is more difficult to produce than other proteins in the process of transcription, translation or folding. For example, in the embodiment of the present application, the constant region CL of the ANG-2 antibody light chain sequence and the constant region CH1 in the ANG-2 antibody heavy chain sequence are replaced with each other. Since artificial sequence replacement is more likely to produce mismatches, the expression difficulty of the ANG-2 antibody light chain sequence is higher than that of the VEGF antibody light chain sequence, and the expression difficulty of the ANG-2 antibody heavy chain sequence is higher than that of the VEGF antibody heavy chain sequence.
[0060] The phrase "antigen-binding fragment" includes a protein having at least one CDR and capable of selectively recognizing an antigen, ie, capable of binding an antigen with a KD at least in the micromolar range.
[0061] The phrase "bispecific antibody" (bsAb) includes antibodies that are capable of selectively binding to two epitopes. Bispecific antibodies generally comprise two different heavy chains, wherein each heavy chain specifically binds a different epitope - on two different molecules (eg, antigens) or on the same molecule (eg, on the same antigen).
[0062] Bispecific antibodies come in many forms, such as Amgen's small protein BiTE technology platform consisting of only two antigen-binding fragments in series to IgG-like immunoglobulins containing additional domains, such as Akeso Biopharm's already-launched cardunilimab. Currently, there are more than 20 different commercial technology platforms available for the creation and development of bispecific antibodies, more than a dozen bispecific antibodies have been launched, and more than 100 bispecific antibodies are in clinical development. However, how to screen a bispecific antibody-expressing cell line that meets the expected quality and is suitable for the later production process is a pain point in the industry. Due to their unique structural characteristics, different platform design structures of bispecific antibodies vary greatly. Therefore, there are huge challenges in the production process of bispecific antibodies, and an excellent cell line is the starting point of the entire production process and largely determines the complexity of the later process and the production cost of the later product after it is launched. Therefore, screening an excellent cell line is particularly important for the entire pharmaceutical company.
[0063] The phrase "site-specific integration" (also known as "site-specific recombination" or "site-specific recombination") refers to a gene targeting method used to direct the insertion or integration of a gene or nucleic acid sequence into a specific location in the genome, that is, to direct DNA to a specific site between two nucleotides in a contiguous polynucleotide chain.
[0064] The applicant has found through research that a CHO nucleotide sequence that can stably and highly express a target protein is used as a highly expressed fragment of the target protein. Further, through site-specific integration technology, it is found that inserting (or integrating) the nucleotide sequence encoding the target protein as an exogenous nucleic acid fragment into the highly expressed fragment can greatly reduce the construction cost of the cell line, while stably and highly expressing the target protein. Further research has found that in site-specific integration technology, when the nucleotide sequence encoding the target protein involves multiple nucleotide sequences, the arrangement order of the multiple nucleotide sequences will affect the expression amount of each encoded protein, and also affect the mismatch rate of the target protein, thereby proposing this application.
[0065] The present application discloses a cell strain, which comprises a highly expressed fragment and an exogenous nucleic acid fragment integrated in the highly expressed fragment; the highly expressed fragment comprises the nucleotide sequence shown in SEQ ID No.1; the exogenous nucleic acid fragment comprises a first target gene encoding a first protein fragment and a second target gene encoding a second protein fragment, the expression difficulty of the first protein fragment is higher than that of the second protein fragment, and the first target gene is located upstream of the second target gene.
[0066] In the present application, the integration site of the exogenous nucleic acid fragment can correspond to the 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, 10th, ... 444th position of the highly expressed fragment.
[0067] The above-mentioned "cell strain" includes any cell suitable for expressing a recombinant nucleic acid sequence and having a high expression fragment that allows stable integration and enhanced expression of an exogenous nucleic acid fragment. The cell includes a mammalian cell, such as a non-human animal cell, a human cell, or a cell fusion, such as a hybridoma or a quadroma. In some embodiments, the cell is a human, monkey, ape, hamster, rat, or mouse cell. In some embodiments, the cell is a mammalian cell selected from the following cells: CHO (e.g., CHO K1, etc.), COS (e.g., COS-7), retinal cells, Vero, CV1, kidney (e.g., HEK293, 293EBNA, MDCK, etc.), HeLa, HepG2, myeloma cells, tumor cells. Specifically, the cell strain contains a high expression fragment of the nucleotide sequence shown in SEQ ID No.1, and on this basis, the exogenous nucleic acid fragment is integrated into the high expression fragment to obtain a cell strain that stably and highly expresses the target protein. Preferably, the cell strain is a CHO cell.
[0068] The exogenous nucleic acid fragment encodes the target protein. When the target protein is composed of multiple parts (for example, a single-chain antibody is mainly composed of a heavy chain variable region and a light chain variable region), the protein coding sequence with a higher expression difficulty is preferably set upstream of the protein coding sequence with a lower expression difficulty. This setting is beneficial to the balanced expression of the multiple parts, the expression of the target protein and the reduction of the mismatch rate.
[0069] In some embodiments, the exogenous nucleic acid fragment comprises a first target gene of a first protein fragment, a second target gene encoding a second protein fragment, a third target gene encoding a third protein fragment, and a fourth target gene encoding a fourth protein fragment. The first protein fragment is more difficult to express than the second protein fragment, and the first target gene is located upstream of the second target gene; the third protein fragment is more difficult to express than the fourth protein fragment, and the third target gene is located upstream of the fourth target gene. At the same time, the first target gene, the second target gene, the third target gene, and the fourth target gene jointly encode a bispecific antibody or an antigen-binding fragment thereof.
[0070] From the above, in order to balance the expression levels of multiple parts, to increase the expression level of the target protein and to reduce the mismatch rate, the protein coding sequence with higher expression difficulty is preferably set upstream of the protein coding sequence with lower expression difficulty.
[0071] When the target protein is a bispecific antibody or an antigen-binding fragment thereof, since the expression of the heavy chain depends on the correct folding of the light chain, and the light chain can assist the secretion of the heavy chain, it is preferred to place the sequence encoding the light chain upstream to facilitate the production of more light chain antibody fragments, thereby further increasing the yield of the target protein. Furthermore, in the light chain part or the heavy chain part, the protein coding sequence with a higher expression difficulty is preferably placed upstream of the protein coding sequence with a lower expression difficulty, further increasing the expression amount of the bispecific antibody or its antigen-binding fragment and reducing the mismatch rate.
[0072] In some embodiments, further, the second target gene is located upstream of the third target gene.
[0073] In a specific embodiment, the structure of the bispecific antibody or its antigen-binding fragment is: the first protein fragment is the first light chain, the second protein fragment is the second light chain, the third protein fragment is the first heavy chain, and the fourth protein fragment is the second heavy chain. In some embodiments, the protein bound by the first light chain and the first heavy chain can specifically bind to the first target; the protein bound by the second light chain and the second heavy chain can specifically bind to the second target.
[0074] In a more specific embodiment, the sequence of the first target gene is shown as SEQ ID No.2, the sequence of the third target gene is shown as SEQ ID No.3, the sequence of the second target gene is shown as SEQ ID No.4, and the sequence of the fourth target gene is shown as SEQ ID No.5. Here, the target protein is an ANG-2 / VEGF bispecific bivalent antibody, the antibody comprises a heavy chain sequence (coding sequence is SEQ ID No.3) and a light chain sequence (coding sequence is SEQ ID No.2) corresponding to the ANG-2 target, and a heavy chain sequence (coding sequence is SEQ ID No.5) and a light chain sequence (coding sequence is SEQ ID No.4) corresponding to the VEGF target, wherein the constant region CL of the ANG-2 antibody light chain sequence is replaced with the constant region CH1 in the ANG-2 antibody heavy chain sequence, so the expression difficulty of the ANG-2 antibody light chain sequence (coding sequence is SEQ ID No.2) is higher than that of the VEGF antibody light chain sequence (coding sequence is SEQ ID No.4), and the expression difficulty of the ANG-2 antibody heavy chain sequence (coding sequence is SEQ ID No.3) is higher than that of the VEGF antibody heavy chain sequence (coding sequence is SEQ ID No.5), so the sequence order integrated into the high expression fragment is the first light chain gene (SEQ ID No.2) which is difficult to express, the second light chain gene (SEQ ID No.4) which is easy to express, and the first heavy chain gene (SEQ ID No.5) which is difficult to express. No.3) and the easily expressed second heavy chain gene (SEQ ID No.5).
[0075] In some embodiments, the upstream of the first target gene and the second target gene respectively contain a nucleotide sequence encoding the first signal peptide shown in SEQ ID No. 6, and the upstream of the third target gene and the fourth target gene respectively contain a nucleotide sequence encoding the second signal peptide shown in SEQ ID No. 7.
[0076] In a preferred embodiment, the integration site of the exogenous nucleic acid fragment is any site within the 11th to 430th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 21st to 414th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 38th to 402nd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 53rd to 389th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 73rd to 373rd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 91st to 360th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 108th to 342nd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 126th to 326th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 143rd to 310th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 160th to 295th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 178th to 274th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 194th to 263rd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 209th to 253rd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 221st to 242nd base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is any site within the 231st to 240th base interval of the highly expressed fragment. In some embodiments, the integration site of the exogenous nucleic acid fragment is the position shown in the annotation information of the CHO cell, such as NW_003616785.1:83044.
[0077] The present application also discloses a method for preparing a cell line. When the cell is a CHO cell, its genome contains the highly expressed fragment of the present application, and the exogenous nucleic acid fragment is site-specifically integrated into the highly expressed fragment to obtain a cell line.
[0078] When the cell is another cell line, its genome does not contain the highly expressed fragment of the present application, and the highly expressed fragment can be first stably introduced into the cell, and then the exogenous nucleic acid fragment can be site-specifically integrated into the highly expressed fragment to obtain a cell line.
[0079] In some embodiments, the preparation method comprises the following steps:
[0080] S1. An exogenous nucleic acid fragment is provided, wherein the exogenous nucleic acid fragment includes a first target gene and a second target gene, the expression difficulty of the first target gene is higher than that of the second target gene, and the first target gene is upstream of the second target gene;
[0081] S2. Integrate the exogenous nucleic acid fragment into the highly expressed fragment to obtain a cell line.
[0082] In some embodiments, the preparation method comprises the following steps:
[0083] S1. An exogenous nucleic acid fragment is set, the exogenous nucleic acid fragment includes a first target gene, a second target gene, a third target gene and a fourth target gene, the expression difficulty of the first target gene is higher than that of the second target gene, and the first target gene is upstream of the second target gene; the expression difficulty of the third protein fragment is higher than that of the fourth protein fragment, and the third target gene is upstream of the fourth target gene;
[0084] S2. Integrate the exogenous nucleic acid fragment into the highly expressed fragment to obtain a cell line.
[0085] In some embodiments, the preparation method comprises the following steps:
[0086] S1: an exogenous nucleic acid fragment is set, wherein the exogenous nucleic acid fragment includes, in the 5'-3' direction, a first target gene encoding a first light chain, a second target gene encoding a second light chain, a third target gene encoding a first heavy chain, and a fourth target gene encoding a second heavy chain, wherein the expression difficulty of the first target gene is higher than that of the second target gene, and the expression difficulty of the third target gene is higher than that of the fourth target gene; the first target gene, the second target gene, the third target gene, and the fourth target gene are independently provided with a nucleotide sequence encoding a signal peptide upstream;
[0087] The first target gene, the second target gene, the third target gene, and the fourth target gene jointly encode a bispecific antibody or an antigen-binding fragment thereof;
[0088] S2. Integrate the exogenous nucleic acid fragment into the highly expressed fragment to obtain a cell line.
[0089] In a specific embodiment, the preparation method comprises the following steps:
[0090] S1: setting an exogenous nucleic acid fragment, wherein the exogenous nucleic acid fragment includes, in the 5'-3' direction, a nucleotide sequence encoding SEQ ID No.6, SEQ ID No.2, a nucleotide sequence encoding SEQ ID No.6, SEQ ID No.4, a nucleotide sequence encoding SEQ ID No.7, SEQ ID No.3, a nucleotide sequence encoding SEQ ID No.7, and SEQ ID No.5;
[0091] S2. Integrate the exogenous nucleic acid fragment into the highly expressed fragment to obtain a cell line.
[0092] The present application also discloses the use of the above cell lines in protein expression. In some embodiments, the protein includes an antigen binder. In some embodiments, the antigen binder is an antibody. In some embodiments, the antigen binder is a bispecific antibody. In some embodiments, the bispecific antibody is a bivalent antibody.
[0093] The present application is further described below by way of examples. Unless otherwise specified, the materials in the examples are prepared according to existing methods or purchased directly from the market.
[0094] Example 1: Screening of highly expressed fragments:
[0095] Specific consumables: Neon Resuspension Buffer R (ThermoFisher), a resuspension buffer for cell electroporation; E1 Buffer (Thermo Fisher); recovery medium is 80% (v / v) EX-CELL CHOCloning Medium (Sigma-Aldrich) and 20% (v / v) EX-CELL Advanced CHO Fed-batch Medium (Sigma-Aldrich) with 1% GlutaMAX (ThermoFisher); expansion medium and subculture medium are both EX-CELL Advanced CHO Fed-batch Medium with 1% GlutaMAX (ThermoFisher); pressurized medium 1 is subculture medium with G418 added to a final concentration of 800 μg / ml; pressurized medium 2 is subculture medium with hygromycin added to a final concentration of 250 μg / ml; conditioned medium is the supernatant obtained by sterile filtration after inoculating CHO-K1 in subculture medium for 1 day; cloning medium is 75% (v / v) EX-CELL CHO Cloning Medium, 20% (v / v) conditioned medium and 5% (v / v) ClonaCell-CHO ACF Supplement were supplemented with 1% GlutaMAX. The basal medium in the fed-batch medium was EX-CELL Advanced CHO Fed-batch Medium supplemented with 1% GlutaMAX (ThermoFisher), and the feed medium was Cell Boost 7a / 7b (HyClone).
[0096] The operation process of this embodiment is as follows Figure 1 As shown, it mainly includes the following steps:
[0097] (1) Using the green fluorescent protein gene (EGFP) as a marker gene for screening and the attP sequence as a homology arm, a recombinant plasmid containing RMCE (recombinase-mediated cassette exchange) was constructed, see Figure 2 ; After linearizing the constructed plasmid, purify and recover the linear DNA; take CHO-K1 cells, centrifuge, and discard the supernatant; resuspend the cells with 100μL of R Buffer, a special resuspension buffer for cell electroporation instrument; electroporate 3 times.
[0098] (2) Pressurization: the pressurization reagent is G418, the pressurization concentration is 800 μg / ml, and the cells are divided into minipools (i.e., 96-well plates) for culture.
[0099] (3) Check the plate after ten days. When a large number of fluorescent cell clusters are observed, they can be enriched. The one-well-to-one-well principle should be followed during enrichment.
[0100] (4) Observe the growth of the enriched cells at any time and expand them when the coverage rate reaches more than 50%.
[0101] (5) After expansion to a shake flask, the cells were passaged three times to form a stable cell pool. The formation of monoclonal cells and the screening of stable fluorescence were achieved by artificial intelligence. The steps are shown in Figure 1 The specific steps are shown in a)-h).
[0102] a) The cells in the stable cell pool are diluted to a certain concentration and then inoculated into a culture dish containing a semi-solid culture medium. During the inoculation process, the cells are ensured to be approximately evenly distributed in various positions of the culture dish. The cells are allowed to stand for about half an hour and wait for the cells to settle to the bottom of the culture dish.
[0103] b) The culture dish is transferred to the stage of an electron microscope, and the electron microscope performs high-throughput scanning on the cells in the culture dish;
[0104] c) uploading the image scanned by the electron microscope to the server for artificial intelligence image analysis, the analysis process includes monoclonal cell line detection, protein expression level prediction of monoclonal cell lines, protein expression level ranking, and coding and positioning of screened protein high-expressing cell lines; the protein expression level prediction of the monoclonal cell lines can be based on fluorescence or not, and the following step c is the protein expression level prediction not based on fluorescence;
[0105] Step c is based on the monoclonal cell line target detection of image processing technology. In this embodiment, the YOLOv8 target detection algorithm is used, and the actual detection effect reaches mAP 94.2%. The target detection model for monoclonal cell lines in step c is consistent with the commonly used deep learning target detection model, so it is not described in detail. Specifically, the cell image needs to be annotated first. In order to improve the prediction accuracy of the algorithm, the bounding box of all monoclonal cell lines and adhesion cell lines will be annotated. It is used as the real target bounding box (ground truth) for the loss calculation of the model output. After algorithm training, the model learns the ability to extract the border information of monoclonal cell lines. In the actual application scenario step c, the trained model can predict the border of the monoclonal cell line in the image.
[0106] Step c is based on the prediction of fluorescent protein expression based on monoclonal cell images. In this embodiment, the deep learning network of SqueezeNet and the MSE Loss function are used. Since the predicted protein expression in this project is not sorted, the sorting result is the ultimate goal of the algorithm, so the evaluation standard NDCG of the sorting algorithm is used here. The specific calculation formula is as follows:
[0107] Where IDCG = the best ranked DCG. Specifically in this embodiment, the cell image predicts the expression, and the predicted expression is sorted, and compared with the real fluorescence value sorting result, NDCG = 0.89. The expression prediction model in the present invention is described in detail in patent CN112037862B.
[0108] d) The coding and location information of the screened protein high-expressing cell lines are returned to the robot control software;
[0109] The cell codes and positions returned in step d, in this embodiment, return the top 100 cells in terms of expression level prediction, with the codes ranging from 1 to 100, and the position information includes the coordinates in the plane coordinate system relative to the center of the microscope (the depth of the culture medium is not considered for the time being, because all the cells to be selected are deposited to the bottom of the culture dish. In addition, during the photo shooting process, the cells that have not settled to the bottom of the culture dish will be out of focus, resulting in blurred cells, and the cells with blurred images will also be excluded).
[0110] e) The robot control software automatically operates (or manually assists) the robot arm to aspirate the screened monoclonal cell lines and transfer them to the designated well plate, and the process is repeated until all high-expressing cell lines are transferred;
[0111] Step e: The robot arm absorbs the target monoclonal cell line and transfers it to the well plate. Patents CN113821287B, CN113403431B, N113771030A, and CN113733087B all have detailed descriptions.
[0112] f) After the cell clones in the well plate are cultured to a certain number, they are transferred to a larger well plate for culture, and finally expanded to a shake flask for culture, and the cell expression level is detected to confirm that the cell line selected by artificial intelligence is a high-yield cell line;
[0113] In this embodiment, step f is to transfer to a 96-well plate for amplification culture, and then transfer to a shake flask for culture.
[0114] g) The selected high-yield cell lines are subcultured. Before each generation of cell lines is subcultured, a portion of the sample is diluted and placed in a culture dish to take images using an electron microscope. The images of several generations of cell lines are taken continuously. Currently, the expression of the selected cells can be predicted by the model learned from the cell morphology by taking pictures, and the expression of the selected cells can also be fluorescently labeled.
[0115] Step g: Photographing Monoclonal Cells In this example, a Thermo Fisher Scientific M7000 electron microscope was used to photograph the cells.
[0116] h) Inputting the collected cell line images of several generations into an artificial intelligence algorithm, predicting the cell line's stability in succession according to the fluorescence intensity and / or protein expression, and selecting cell lines with high protein expression characteristics that can be stably proliferated as the final candidate cell lines; the prediction of the cell line's stability in succession can be based on fluorescence, or can be based on cell morphology instead of fluorescence, and the following step h is the prediction of protein expression based on cell morphology instead of fluorescence;
[0117] Step h predicts the cell line stability, including but not limited to the traditional image recognition algorithm using histogram to extract image features and predict, using deep learning neural network to automatically extract image features and predict, and other image-based technical predictions. In this embodiment, a deep learning-based method is used to automatically extract image features of stable cell lines and unstable cell lines and predict stability, and the accuracy of prediction results for different cell lines reaches 84%. Patent CN114417582A is described in detail.
[0118] Wherein, step d and step e can be completed manually. In this embodiment, they are completed automatically by a robot arm.
[0119] (6) CHO-K1 cells were used as control and the inoculation density was 5×10 5 cells / ml, the inoculation system was 30ml, and the cells were counted on the day of inoculation, the 3rd day, the 5th day, the 7th day, the 9th day, the 11th day, the 13th day and the 14th day to screen out the monoclones with a growth rate lower than that of CHO-K1.
[0120] (7) Recover the remaining monoclonal clones and perform stable subculture after recovery. Subculture should be performed every 3 or 4 days. The cell density of the 4-day subculture is 3×10 5 cells / ml, and the cell density for 3-day passage was 5×10 5cells / ml, and the fluorescence value of the cell line was regularly monitored during the subculture. The above-mentioned high-fluorescence cell line was subjected to a subculture stability study for about 90 days using the subculture medium, and it was confirmed that the GBB003 cells with small fluorescence changes and stable cell line growth were stable high-fluorescence cells. The integration site of the cell line was sequenced by NGS, and the sequence of the highly expressed fragment where the integration site was located was SEQ ID No. 1. Specifically, the annotation information of the integration site on CHO is: NW_003616785.1:83044.
[0121] Example 2: Construction of cell lines
[0122] (1) Exogenous nucleic acid construction
[0123] The present application provides a method for constructing a multi-reading frame plasmid, such as Figure 3 As shown, pDonor0.0-H is used to construct the expression reading frame vector of the heavy chain, and the target gene sequence can be inserted in the middle of the XmaI and / or HindIII restriction sites, wherein KanR provides a kanamycin resistance gene for screening in plasmid construction; the Ori segment is the plasmid replication initiation site, which is used for plasmid replication and amplification in prokaryotic cells; the CMV segment (not shown in the figure) is used to initiate the expression of the target gene; SV40 poly (A) is used to terminate the transcription of the target gene; NotI and other restriction sites are used to insert the complete reading frame into the pLangding0.0 plasmid restriction enzyme cutting connection. Figure 4 As shown, pDonor0.0-L is used to construct a vector for light chain expression reading frame, wherein KanR provides a kanamycin resistance gene for screening in plasmid construction; the Ori segment is the plasmid replication initiation site, which is used for transcriptional amplification of the plasmid in prokaryotic cells; the CMV segment (not shown in the figure) is used to initiate the expression of the target gene; SV40 poly (A) is used to terminate the transcription of the target gene; XmaI and / or HindIII restriction sites are used to insert the target gene sequence; NotI and other restriction sites are used to insert the complete reading frame into the pLangding0.0 plasmid by restriction enzyme cutting and connection. pLangding0.0 is used to accept the reading frames of pDonor0.0-H and / or pDonor0.0-L. The pLangding0.0 vector is as shown in FIG. Figure 5 As shown, the Ori segment is the plasmid replication initiation site, which is used for transcriptional amplification of the plasmid in prokaryotic cells; AmpR provides an ampicillin resistance gene for screening in plasmid construction.
[0124] The target genes in the constructed plasmid structure in this example are arranged and named as molecule No. 1, molecule No. 2, molecule No. 3, and molecule No. 4 respectively.
[0125] The construction method of molecular design No. 1 is to synthesize and insert the kozak sequence, signal peptide sequence (SEQ ID No. 6) and ANG-2 antibody light chain sequence (coding sequence is SEQ ID No. 2) into the Hind III and Xma I restriction sites of the pDonor0.0-L plasmid to form the pDonor0.0-ANG-2-LC plasmid. The kozak sequence, signal peptide sequence (SEQ ID No. 6) and VEGF antibody light chain sequence (coding sequence is SEQ ID No. 4) are synthesized and inserted into the Hind III and Xma I restriction sites of the pDonor0.0-L plasmid to form the pDonor0.0-VEGF-LC plasmid. The kozak sequence, signal peptide sequence (SEQ ID No. 7) and ANG-2 antibody heavy chain sequence (coding sequence is SEQ ID No. 3) are synthesized and inserted into the Hind III and Xma I restriction sites of the pDonor0.0-H plasmid to form the pDonor0.0-ANG-2-HC plasmid. The kozak sequence, signal peptide sequence (SEQ ID No.7) and VEGF antibody heavy chain sequence (coding sequence SEQ ID No.5) were synthetically inserted into the HindIII and XmaI restriction sites of the pDonor0.0-H plasmid to form the pDonor0.0-VEGF-HC plasmid. After the pDonor0.0-ANG-2-LC and pLanding0.0 plasmids were double-digested with NotI / AgeI, the ANG-2-LC reading frame fragment of pDonor0.0-ANG-2-LC and the linearized pLanding0.0 fragment were recovered, and then ligated and transformed into DH5α chemical competent cells with T4 DNA ligase, and the colonies were grown on LB ampicillin-resistant plates, washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL. After the pDonor0.0-VEGF-LC and pLanding0.0-AL plasmids were double-digested with BsiwI / MluI, the VEGF-LC reading frame fragment of pDonor0.0-VEGF-LC and the linearized pLanding0.0-AL fragment were recovered, and then transformed into StBl3 chemically competent cells using T4 DNA ligase. After the colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL.After the pDonor0.0-ANG-2-HC and pLanding0.0-AL-VL plasmids were double-digested with NheI / XhoI, the ANG-2-HC reading frame fragment of pDonor0.0-ANG-2-HC and the linearized pLanding0.0-AL-VL fragment were recovered, and then ligated with T4 DNA ligase to transform StBl3 chemical competent cells. After the colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL-AH. After the pDonor0.0-VEGF-HC and pLanding0.0-AL-VL-AH plasmids were double-digested with EcoRI / KpnI, the VEGF-HC reading frame fragment of pDonor0.0-VEGF-HC and the linearized pLanding0.0-AL-VL-AH fragment were recovered, and then transformed into StB13 chemical competent cells using T4 DNA ligase, and the colonies were grown on LB ampicillin-resistant plates, washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL-AH-VH. In this embodiment, the ANG-2 antibody light chain sequence is the first target gene, the VEGF antibody light chain sequence is the second target gene, the ANG-2 antibody heavy chain sequence is the third target gene, and the VEGF antibody heavy chain sequence is the fourth target gene. The plasmid schematic diagram is as follows. Figure 6 shown.
[0126] The construction of molecular design No. 2 is to double-digest the pDonor0.0-ANG-2-LC and pLanding0.0 plasmids with Not I / Age I, respectively, recover the ANG-2-LC reading frame fragment of pDonor0.0-ANG-2-LC and the linearized pLanding0.0 fragment, and then connect and transform DH5α chemically competent cells with T4 DNA ligase, spread on LB ampicillin-resistant plates, grow colonies, wash and collect the bacteria with LB medium, and extract the plasmid to form pLanding0.0-AL. After the pDonor0.0-VEGF-LC and pLanding0.0-AL plasmids were double-digested with BsiwI / MluI, the VEGF-LC reading frame fragment of pDonor0.0-VEGF-LC and the linearized pLanding0.0-AL fragment were recovered, and then transformed into StBl3 chemically competent cells using T4 DNA ligase. After the colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL. After the pDonor0.0-VEGF-HC and pLanding0.0-AL-VL plasmids were double-digested with NheI / XhoI, the VEGF-HC reading frame fragment of pDonor0.0-VEGF-HC and the linearized pLanding0.0-AL-VL fragment were recovered, and then ligated with T4 DNA ligase to transform StBl3 chemically competent cells. After the colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL-VH. After the pDonor0.0-ANG-2-HC and pLanding0.0-AL-VL-VH plasmids were double-digested with EcoRI / KpnI, the ANG-2-HC reading frame fragment of pDonor0.0-ANG-2-HC and the linearized pLanding0.0-AL-VL-VH fragment were recovered, and then transformed into StB13 chemical competent cells using T4 DNA ligase, and the colonies were grown on LB ampicillin-resistant plates, washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-AL-VL-VH-AH. In this embodiment, the ANG-2 antibody light chain sequence is the first target gene, the VEGF antibody light chain sequence is the second target gene, the VEGF antibody heavy chain sequence is the third target gene, and the ANG-2 antibody heavy chain sequence is the fourth target gene. The schematic diagram of the plasmid is shown in FIG. Figure 7 shown.
[0127] The construction of molecular design No. 3 is to double-digest the pDonor0.0-VEGF-LC and pLanding0.0 plasmids with NotI / AgeI, respectively, recover the VEGF-LC reading frame fragment of pDonor0.0-VEGF-LC and the linearized pLanding0.0 fragment, and then connect and transform DH5α chemically competent cells with T4 DNA ligase, spread on LB ampicillin-resistant plates, grow colonies, wash and collect the bacteria with LB medium, and extract the plasmid to form pLanding0.0-VL. After the pDonor0.0-ANG-2-LC and pLanding0.0-VL plasmids were double-digested with BsiwI / MluI, the ANG-2-LC reading frame fragment of pDonor0.0-ANG-2-LC and the linearized pLanding0.0-VL fragment were recovered, and then ligated with T4 DNA ligase to transform Stbl3 chemical competent cells. After the colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL. After the pDonor0.0-VEGF-HC and pLanding0.0-VL-AL plasmids were double-digested with NheI / XhoI, the VEGF-HC reading frame fragment of pDonor0.0-VEGF-HC and the linearized pLanding0.0-VL-AL fragment were recovered, and then ligated with T4 DNA ligase to transform StBl3 chemically competent cells. After the colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL-VH. After the pDonor0.0-ANG-2-HC and pLanding0.0-VL-AL-VH plasmids were double-digested with EcoRI / KpnI, the ANG-2-HC reading frame fragment of pDonor0.0-ANG-2-HC and the linearized pLanding0.0-VL-AL-VH fragment were recovered, and then transformed into Stbl3 chemical competent cells using T4 DNA ligase, and the colonies were grown on LB ampicillin-resistant plates, washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL-VH-AH. In this embodiment, the VEGF antibody light chain sequence is the first target gene, the ANG-2 antibody light chain sequence is the second target gene, the VEGF antibody heavy chain sequence is the third target gene, and the ANG-2 antibody heavy chain sequence is the fourth target gene. The schematic diagram of the plasmid is shown in FIG. Figure 8 shown.
[0128] The construction of molecular design No. 4 is to double-digest the pDonor0.0-VEGF-LC and pLanding0.0 plasmids with NotI / AgeI, respectively, recover the VEGF-LC reading frame fragment of pDonor0.0-VEGF-LC and the linearized pLanding0.0 fragment, and then connect and transform DH5α chemically competent cells with T4 DNA ligase, spread on LB ampicillin-resistant plates, grow colonies, wash and collect the bacteria with LB medium, and extract the plasmid to form pLanding0.0-VL. After the pDonor0.0-ANG-2-LC and pLanding0.0-VL plasmids were double-digested with BsiwI / MluI, the ANG-2-LC reading frame fragment of pDonor0.0-ANG-2-LC and the linearized pLanding0.0-VL fragment were recovered, and then ligated with T4 DNA ligase to transform StBl3 chemical competent cells. After the colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL. After the pDonor0.0-ANG-2-HC and pLanding0.0-VL-AL plasmids were double-digested with NheI / XhoI, the ANG-2-HC reading frame fragment of pDonor0.0-ANG-2-HC and the linearized pLanding0.0-VL-AL fragment were recovered, and then ligated with T4 DNA ligase to transform StBl3 chemical competent cells. After the colonies were grown on LB ampicillin-resistant plates, they were washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL-AH. After the pDonor0.0-VEGF-HC and pLanding0.0-VL-AL-AH plasmids were double-digested with EcoRI / KpnI, the VEGF-HC reading frame fragment of pDonor0.0-VEGF-HC and the linearized pLanding0.0-VL-AL-AH fragment were recovered, and then transformed into StB13 chemical competent cells using T4 DNA ligase, and the colonies were grown on LB ampicillin-resistant plates, washed and collected with LB medium, and the plasmid was extracted to form pLanding0.0-VL-AL-AH-VH. In this embodiment, the VEGF antibody light chain sequence is the first target gene, the ANG-2 antibody light chain sequence is the second target gene, the ANG-2 antibody heavy chain sequence is the third target gene, and the VEGF antibody heavy chain sequence is the fourth target gene. The schematic diagram of the plasmid is shown in FIG. Fig. 9 shown.
[0129] This example is used to express an ANG-2 / VEGF bispecific bivalent antibody, which contains a heavy chain sequence (encoding sequence is SEQ ID No. 3) and a light chain sequence (encoding sequence is SEQ ID No. 2) corresponding to the ANG-2 target, and a heavy chain sequence (encoding sequence is SEQ ID No. 5) and a light chain sequence (encoding sequence is SEQ ID No. 4) corresponding to the VEGF target. All genes are synthesized by GenScript, and some information is as follows:
[0130] In the ANG-2 antibody light chain sequence, the coding sequence of the heavy chain constant region 1 (CH1) in the first light chain gene is SEQ ID No.8.
[0131] In the ANG-2 antibody heavy chain sequence, the coding sequence of the light chain constant region (CL) in the first heavy chain gene is SEQ ID No.9.
[0132] Wherein, a signal peptide sequence SEQ ID No.6 required for antibody secretion is provided before the antibody light chain sequence.
[0133] Wherein, a signal peptide sequence SEQ ID No.7 required for antibody secretion is provided before the antibody heavy chain sequence.
[0134] In the ANG-2 / VEGF bispecific bivalent antibody formed after the four molecules are expressed, the constant region CL of the ANG-2 antibody light chain sequence and the constant region CH1 in the ANG-2 antibody heavy chain sequence are replaced with each other, so the expression difficulty of the ANG-2 antibody light chain sequence is higher than that of the VEGF antibody light chain sequence, and the expression difficulty of the ANG-2 antibody heavy chain sequence is higher than that of the VEGF antibody heavy chain sequence.
[0135] (2) Stable high fluorescence cell GBB003 was selected for transfection and integration verification. The recombinant plasmids of the ANG-2 / VEGF bispecific bivalent antibody expression genes in the four arrangements contained in step (1) were amplified and linearized, and then co-transfected with Bxb-1 integrase into stable high fluorescence cell GBB003. The plasmid map of Bxb-1 integrase is shown in Fig.10 As shown;
[0136] (3) One day after transfection, the cell pool was used to construct a minipool and pressurized screening was performed using pressurized medium 2;
[0137] (4) After 14 days, the non-fluorescent cells in the minipool were expanded and cultured in a 96-well plate-24-well plate-6-well plate. When expanding to a 6-well plate, a 6-well batch culture was performed. The 6-well batch culture was performed by inoculating 5×10 5 / ml, 2ml system, 37℃ 5% carbon dioxide 120 rpm condition culture, 6-well batch culture experimental results are shown in Table 1:
[0138] Table 1: Statistical results of expression in six-well batch culture of cell pools
[0139] (5) For each molecule, select the TOP1 cell pool (i.e., 1-9 cell pool, 2-5 cell pool, 3-4 cell pool, 4-6 cell pool) and expand it to T125 shake flasks. After two subcultures, feed the shake flasks for 14 days. The inoculation density is 5×10 5 / ml, inoculation system 30ml, counted on the day of inoculation, culture day 3, culture day 5, culture day 7, culture day 9, culture day 11, culture day 13 and culture day 14, and supplemented with 3% Cell Boost 7a and 0.3% Cell Boost 7b on the 3rd day, supplemented with 5% Cell Boost 7a and 0.5% Cell Boost7b on the 5th day, culture day 7, culture day 9, culture day 11 and culture day 13, during which the glucose concentration was controlled at the level of 2-8g / L, and the expression level was measured by Octect after 14 days, and the results of shake flask feeding expression are shown in Table 2;
[0140] Table 2: Statistical results of shake flask feeding expression
[0141] (6) The supernatant of the shake flask feed was purified by Protein A chromatography column and analyzed by SDS-PAGE (the results are shown in Fig.11 The SEC purity is shown in Table 3. The SEC purity of molecule No. 1 is 83.47%, which is significantly better than other molecular designs. The ratio of each chain was analyzed by reducing CE-SDS. As shown in Table 4, the ratio of the two light chains and the two heavy chains of molecule No. 1 is closer to 1:1, which is more in line with the expected theoretical value. Comprehensive analysis shows that the purity and molecular expression of molecule No. 1 are optimal, so cell pools 1-9 of molecule No. 1 are selected for monocloning, and cell pools 2-5, 3-4, and 4-6 are used as control groups for monocloning;
[0142] Table 3: SEC purity analysis results of cell pool
[0143] Table 4: CE-SDS purity analysis results of cell pool reduction Molecular number Cell pool number L1(%) L2(%) H1(%) H2(%) Molecule No. 1 1-9 15.1 14 30.9 40 Molecule No. 2 2-5 0 23.9 22.8 53.3 Molecule No. 3 3-4 21.7 7.7 31.8 38.8 Molecule No. 4 4-6 15.6 12 27.2 45.1
[0144] (7) Cell pools 1-9, 2-5, 3-4, and 4-6 were diluted to a 96-well plate at a density of 0.8 cells / well. Cell growth status was photographed using VIPS at 1 hour, 1 day, 2 days, 3 days, 5 days, 7 days, and 14 days, and non-fluorescent monoclones were counted;
[0145] (8) Expand the monoclonal wells without fluorescence, and follow the method of 96-well plate-24-well plate-6-well plate-shake flask. In the process of expansion, clones with low growth rate are eliminated. No monoclonal clones are formed in 4-6, and the cells do not grow and reproduce. Therefore, cell pools 4-6 cannot be screened by six-well batch culture. After the 6-well is expanded, six-well batch culture is performed. The 6-well batch culture is to inoculate 5×10 5 / ml, 2ml system, 37°C 5% carbon dioxide 120 rpm conditions for 7 days, the top 5 monoclones expressed in 6-well batch culture of each cell pool are shown in Table 5;
[0146] Table 5: Statistical results of expression of monoclonal six-well batch culture
[0147] (9) Select the top three cell lines in each cell pool with the highest expression in six-well batch culture. After two subcultures, perform feeding in shake flasks for 14 days. The inoculation density is 5×10 5 / ml, inoculation system 30ml, count on the day of inoculation, 3rd day of culture, 5th day of culture, 7th day of culture, 9th day of culture, 11th day of culture, 13th day of culture and 14th day of culture, and supplement 3% Cell Boost 7a and 0.3% Cell Boost 7b on the 3rd day, 5% Cell Boost7a and 0.5% Cell Boost 7b on the 5th day of culture, 7th day of culture, 9th day of culture, 11th day of culture and 13th day of culture, during which the glucose concentration was controlled at 2-8g / L, and after 14 days, use The expression level was measured, and the results of monoclonal shake flask feeding expression are shown in Table 6;
[0148] Table 6: Statistical results of monoclonal shake flask feeding expression
[0149] The supernatant of the shake flask feed was captured with Protein A filler and then analyzed by SEC. The analysis results are shown in Table 7. The quality attributes of the monoclonal clones produced by each cell pool are relatively uniform, among which the cell pool of molecule No. 1 has the highest SEC purity. By calculating the yield of the target molecule, 1-9-1 and 1-9-2 both performed well, and the theoretical yield of the target molecule was about 4.7g. The calculation results are shown in Table 8.
[0150] Table 7: Monoclonal SEC analysis results
[0151] Table 8: Statistics of relative yields of monoclonal target products
[0152] Through the above experimental contents, we have screened the expected bispecific antibody expression cell lines. The experimental process is controllable, the workload is greatly reduced compared with the traditional method, and the antibody mismatch rate is lower. In the ANG-2 / VEGF bispecific bivalent antibodies formed after the expression of the four molecules, the constant region CL of the ANG-2 antibody light chain sequence and the constant region CH1 in the ANG-2 antibody heavy chain sequence are replaced with each other. Therefore, the expression difficulty of the ANG-2 antibody light chain sequence is higher than that of the VEGF antibody light chain sequence, and the expression difficulty of the ANG-2 antibody heavy chain sequence is higher than that of the VEGF antibody heavy chain sequence. In addition, since artificial sequence replacement is more likely to produce mismatches, the expression reading frame of ANG-2 is placed at the front of the plasmid design, so that it is less affected by the previous reading frame. Since the expression of the heavy chain in conventional antibody expression depends on the correct folding of the light chain, and the light chain can assist the secretion of the heavy chain, placing the light chain reading frame at the front of the plasmid design can produce more light chain antibody fragments, thereby further increasing the yield of the target molecule. Therefore, it explains why our No. 1 molecule design has a higher expression amount and a lower mismatch rate.
[0153] The exogenous nucleic acid fragment can be replaced. The first target gene can be replaced by a gene encoding a first protein, the second target gene can be replaced by a gene encoding a second fusion protein, and the expression difficulty of the first protein is higher than that of the second protein. The first target gene can be replaced by a gene encoding a monoclonal antibody light chain, and the second target gene can be replaced by a gene encoding a monoclonal antibody heavy chain, and the expression difficulty of the light chain is higher than that of the heavy chain. The first target gene can be replaced by a gene encoding a monoclonal antibody heavy chain, and the second target gene can be replaced by a gene encoding a monoclonal antibody light chain, and the expression difficulty of the heavy chain is higher than that of the light chain. There can be multiple target genes in the exogenous nucleic acid fragment, and the difficulty of expressing proteins in each target gene segment is different. The target gene with high expression difficulty is located before the target gene with low expression in the exogenous nucleic acid fragment.
[0154] Note that the above are only preferred embodiments of the present application and the technical principles used. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present application. Therefore, although the present application is described in more detail through the above embodiments, the present application is not limited to the above embodiments, and may also include more other equivalent embodiments without departing from the technical concept of the present application, all of which belong to the protection scope of the present application.
[0155] Those skilled in the art will also recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the appended claims.
Claims
1. A cell line, characterized in that The cell line comprises a highly expressed fragment and an exogenous nucleic acid fragment integrated into the highly expressed fragment; The highly expressed fragment comprises the nucleotide sequence shown in SEQ ID No.1; The exogenous nucleic acid fragments include a first target gene encoding a first protein fragment and a second target gene encoding a second protein fragment. The expression difficulty of the first protein fragment is higher than that of the second protein fragment. The first target gene is located upstream of the second target gene.
2. The cell line according to claim 1, characterized in that The exogenous nucleic acid fragment further comprises a third target gene encoding a third protein fragment and a fourth target gene encoding a fourth protein fragment, the expression difficulty of the third protein fragment is higher than that of the fourth protein fragment, the third target gene is located upstream of the fourth target gene, wherein the first target gene, the second target gene, the third target gene and the fourth target gene jointly encode a bispecific antibody or an antigen-binding fragment thereof.
3. The cell line according to claim 2, characterized in that The second target gene is located upstream of the third target gene.
4. The cell line according to claim 3, characterized in that The first protein fragment is a first light chain, the second protein fragment is a second light chain, the third protein fragment is a first heavy chain, and the fourth protein fragment is a second heavy chain; optionally, the protein bound by the first light chain and the first heavy chain can specifically bind to a first target; the protein bound by the second light chain and the second heavy chain can specifically bind to a second target.
5. The cell line according to claim 4, characterized in that The sequence of the first target gene is shown as SEQ ID No.2, the sequence of the third target gene is shown as SEQ ID No.3, the sequence of the second target gene is shown as SEQ ID No.4, and the sequence of the fourth target gene is shown as SEQ ID No.
5.
6. The cell line according to claim 5, characterized in that The upstream of the first target gene and the second target gene respectively contain nucleotide sequences encoding the first signal peptide shown in SEQ ID No.6, and the upstream of the third target gene and the fourth target gene respectively contain nucleotide sequences encoding the second signal peptide shown in SEQ ID No.
7.
7. The cell line according to any one of claims 1 to 6, characterized in that The integration site of the exogenous nucleic acid fragment is any site within the 11th to 430th base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 21st to 414th base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 38th to 402nd base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 53rd to 389th base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 73rd to 373rd base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 91st to 360th base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 108th to 342nd base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 126th to 326th base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 143rd to 310th base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 160th to 295th base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 178th to 274th base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 194th to 263rd base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 209th to 253rd base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 221st to 242nd base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is any site within the 231st to 240th base interval of the highly expressed fragment; Optionally, the integration site of the exogenous nucleic acid fragment is the position shown in the annotation information of CHO cells such as NW_003616785.1:83044.
8. The cell line according to claim 7, characterized in that The cell line is CHO cell.
9. A method for preparing the cell line according to any one of claims 1 to 8, characterized in that: The preparation method comprises the following steps: integrating the exogenous nucleic acid fragment into the highly expressed fragment at a fixed point to obtain the cell line.
10. The preparation method according to claim 9, characterized in that: The preparation method comprises the following steps: S1: An exogenous nucleic acid fragment is provided, wherein the exogenous nucleic acid fragment includes, in the 5'-3' direction, a first target gene encoding a first light chain, a second target gene encoding a second light chain, a third target gene encoding a first heavy chain, and a fourth target gene encoding a second heavy chain, wherein the expression difficulty of the first target gene is higher than that of the second target gene, and the expression difficulty of the third target gene is higher than that of the fourth target gene; the first target gene, the second target gene, the third target gene, and the fourth target gene are independently provided with a nucleotide sequence encoding a signal peptide upstream; The first target gene, the second target gene, the third target gene and the fourth target gene jointly encode a bispecific antibody or an antigen-binding fragment thereof; S2. The exogenous nucleic acid fragment is site-specifically integrated into the highly expressed fragment to obtain a cell line.
11. Use of the cell line according to any one of claims 1 to 8 in protein expression.
Citation Information
Patent Citations
Robot-based cell fluid extraction control methods, devices, equipment, and storage media
CN113403431B
Control information configuration methods, devices, equipment, and media for cell manipulation robots
CN113733087B
Robot-based methods, devices, equipment, and media for cell manipulation tasks.
CN113821287B
Cell strain stability prediction method and device, computer equipment and storage medium
CN114417582A
Cited By
Cell strain, and preparation method therefor and use thereof
EP4806990A1