Methods for estimating the productivity of a target protein
By employing three relative quantities to estimate target protein productivity, the method addresses the limitation of existing technologies in selecting high-productivity cell lines, ensuring effective protein production.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CHITOSE LAB
- Filing Date
- 2024-11-21
- Publication Date
- 2026-06-02
AI Technical Summary
Existing methods fail to estimate the productivity of a specific target protein expressed by a specific cell line, limiting the selection of high-productivity cell lines for protein production.
A method using three relative quantities - the relative amount of DNA fragment encoding the target protein, mRNA transcript, and the target protein present on the cell surface - to estimate and select high-productivity cell lines, followed by culturing these lines to produce the target protein effectively.
Enables accurate estimation of target protein productivity and selection of high-productivity cell lines, facilitating efficient production of the target protein.
Smart Images

Figure 2026090148000001_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for estimating the productivity of a target protein and the like.
Background Art
[0002] Selecting cell lines that can produce more proteins is important. For example, Japanese Patent Application Laid-Open No. 2024-25663 describes a method, system, and apparatus for estimating the expression level of a protein in a cell from a non-fluorescent stained image of the cell.
[0003] With this method, the total amount of protein expressed by a cell line can be estimated. However, it is not possible to estimate the expression level of a specific protein (target protein) expressed by a specific cell line.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] An object of this invention is to provide a method for estimating the productivity of a target protein and the like.
[0006] An object of this invention is to provide a method for selecting a cell line having excellent productivity of a target protein and the like.
Means for Solving the Problems
[0007] Basically, this invention is based on the finding of an embodiment in which the productivity of a target protein in a genetically modified cell can be estimated by using three relative amounts.
[0008] The first invention relates to a method for estimating the productivity of a target protein in a host cell. This method uses three relative quantities to estimate the productivity of a target protein in a host cell. The three relative quantities are (1) The relative amount of the DNA fragment encoding the target protein in a host cell that has been genetically modified to express the target protein, (2) The relative amount of mRNA transcript encoding the target protein in host cells, and (3) The relative amount of the target protein present on the surface of the host cell.
[0009] The second invention relates to a method for producing a target protein. This method involves creating genetically modified cells, obtaining multiple culture lines, For each culture strain, the productivity of the target protein is estimated using the method for estimating the productivity of the target protein described above. Using these results, a high-productivity strain is selected from among several culture lines that have high productivity of the target protein (high-productivity strain selection step). Subsequently, the selected high-productivity strains are cultured to obtain the target protein (target protein acquisition step). In this way, the target protein can be produced effectively. [Effects of the Invention]
[0010] This invention provides a method for estimating the productivity of a target protein by using three relative quantities.
[0011] This invention provides a method for selecting cell lines with superior productivity of a target protein by using three relative quantities. [Brief explanation of the drawing]
[0012] [Figure 1] Figure 1 is a graph showing antibody concentrations. [Figure 2]Figure 2 is a graph showing the relative quantification results of copy numbers in each clone cell. [Figure 3] Figure 3 is a graph showing the analysis results of relative quantification of mRNA in each clone cell. [Figure 4] Figure 4 is a graph showing the calculation results of Stain Index in each clone cell. [Figure 5] Figure 5 is a graph showing the bivariate relationship between antibody concentration by fed-batch culture and copy number, mRNA amount, and Stain Index. [Figure 6] Figure 6 is a graph showing the results of principal component analysis. [Figure 7] Figure 7 is a graph showing the bivariate relationship between antibody concentration by fed-batch culture.
Mode for Carrying Out the Invention
[0013] The first invention relates to a method for estimating the productivity of a target protein in a host cell. This method relates to a method of introducing a gene for expressing the target protein into a host cell and estimating a strain with a high production amount of the target protein from a plurality of host cell lines into which the gene has been introduced.
[0014] A host cell is a cell used for gene introduction. In this invention, various host cells can be employed. Examples of host cells are fungi, yeast, insect cells, animal cells (mammalian cells), and plant cells. An example of a fungus as a host cell is Escherichia coli. Examples of yeast as a host cell are Saccharomyces cerevisiae, Schizosaccharomyces pombe, alkane-assimilating yeast, and methanol-assimilating yeast. Examples of insect cells as a host cell are SF9 cells and High Five cells. Examples of animal cells as a host cell are HEK293 cells (derived from human kidney) and CHO cells (Chinese hamster ovary cells). An example of a plant cell as a host cell is cultured cells of tobacco. Among these, CHO cells are preferred as the host cell.
[0015] A host cell genetically modified to express a target protein A host cell genetically modified to express a target protein can be obtained, for example, by introducing a gene encoding the target protein into the host cell. The method of gene introduction may be appropriately selected and adjusted according to the type of host and the purpose.
[0016] Gene introduction As for the gene introduction method itself, known techniques may be appropriately employed. Examples of gene introduction are transformation, transduction, Agrobacterium-mediated method, lipofection, microinjection, and gene gun (particle gun) method. Examples of transformation are chemical transformation and electroporation. Transduction is a method of introducing a gene into a host cell using a viral vector. Examples of viral vectors used for transduction are lentiviral vectors and adenoviral vectors. Although electroporation is preferred for the gene introduction method, for example, the present invention is not limited thereto.
[0017] Target protein The target protein may be any protein produced by the host cell. Examples of the target protein are antibodies (e.g., IgG antibodies, particularly human IgG antibodies), peptides, and various therapeutic proteins. Techniques for introducing a gene to cause a host cell to produce a target protein are known. Therefore, a known method may be used to cause the host cell to produce the target protein.
[0018] This method estimates the productivity of the target protein in the host cell using three relative amounts. The three relative amounts are (1) The relative amount of a DNA fragment encoding the target protein in a host cell genetically modified to express the target protein, (2) The relative amount of a transcript of mRNA encoding the target protein in the host cell, and (3) The relative amount of the target protein present on the surface of the host cell. As demonstrated in the examples described later, the productivity of the target protein in host cells can be estimated using the three relative quantities mentioned above. Specifically, by using the three relative quantities to determine an index value for estimating the productivity of the target protein in host cells, the productivity of the target protein can be compared by comparing these index values.
[0019] (1) The relative amount of the DNA fragment encoding the target protein in a host cell that has been genetically modified to express the target protein. The relative amount of this DNA fragment can be determined by comparing the amount of the DNA fragment encoding the target protein in each of several strains of host cells obtained through gene transfer (and appropriate culture) with a reference amount. Methods for quantitatively analyzing DNA fragments are already known. Therefore, in this invention, any known method can be appropriately employed for quantitatively analyzing DNA fragments. Examples of methods for quantitatively analyzing DNA fragments include UV absorbance measurement, fluorescence measurement, real-time PCR, a combination of agarose gel electrophoresis and imaging, digital droplet PCR (ddPCR), and turbidimetry. Of these, real-time PCR is preferred. The DNA fragment may be full-length DNA encoding the target protein or a fragment thereof. The relative amount of this DNA fragment can be determined, for example, by taking a predetermined amount of host cells from each strain and using real-time PCR to determine the amount of DNA fragment encoding the target protein in the obtained host cells. The host cells used as a reference for determining the relative amount may be, for example, strains that have been conventionally able to reliably produce the target protein. This point is the same below. The obtained amount of DNA fragment or the relative amount of DNA fragment may be appropriately stored, for example, in the memory of the measurement system or a computer connected to the measurement system.
[0020] (2) The relative amount of mRNA transcripts encoding the target protein in host cells The relative amount of mRNA transcript encoding the target protein in a host cell can be determined by quantitatively analyzing the mRNA transcript of the target protein contained in the host cell using a known method and comparing it to a reference amount. Methods for quantitatively analyzing mRNA transcripts are already known. Therefore, in this invention, known methods can be appropriately adopted for quantitatively analyzing mRNA transcripts. Examples of methods for quantitatively analyzing mRNA transcripts include real-time PCR, Northern blotting, RNA sequencing, digital droplet PCR (ddPCR), and microarrays. Among these, real-time PCR is preferred. Real-time PCR is a method in which mRNA is converted to cDNA with reverse transcriptase, and the amount of mRNA is quantified in real time using a fluorescent dye or probe while being amplified by PCR using specific primers. The amount of mRNA obtained, or the relative amount of mRNA, may be appropriately stored, for example, in the memory of the measurement system or a computer connected to the measurement system.
[0021] (3) The relative amount of the target protein present on the surface of the host cell. The relative amount of the target protein present on the surface of host cells can be determined by quantitatively analyzing the amount of the target protein present on the surface of host cells using known methods and comparing it to a reference amount. Methods for quantifying the target protein present on the surface of host cells are already known. Therefore, in this invention, known methods can be appropriately employed to determine the relative amount of the target protein present on the surface of host cells. Examples of methods for quantitatively analyzing target proteins present on the surface of host cells include flow cytometry, ELISA (enzyme-linked immunosorbent assay), Western blotting, immunofluorescence staining, radiolabeling, and surface plasmon resonance (SPR). Flow cytometry is a method that measures the expression level of a specific protein on the cell surface using a fluorescently labeled antibody. In this method, cells are stained with a fluorescently labeled antibody specific to the protein, and the fluorescence intensity of individual cells is measured using a flow cytometer. Quantitative analysis can be performed from the average fluorescence intensity or the percentage of fluorescence-positive cells. Preferably, the relative amount of target protein present on the surface of host cells is a value obtained using the amount of target protein present on the surface of host cells measured using a fluorescently labeled antibody that specifically binds to the target protein. The amount of target protein obtained, or the relative amount of target protein, may be appropriately stored, for example, in the memory of the measurement system or a computer connected to the measurement system.
[0022] For example, a value (index value) is calculated to estimate the productivity of a target protein in a host cell using three relative quantities. Let the three relative quantities be x, y, and z. Then, let the index value be S. Then, for example, the index can be obtained by calculating S after finding the coefficients a, b, and c such that S = ax + by + cz. For example, a computer can store the coefficients a, b, and c. Then, when the three relative quantities x, y, and z are input, the coefficients can be read from the memory and the calculation unit can be made to calculate S. In this way, an index value (S) for a particular strain can be obtained. The obtained index value may be stored in memory in association with the strain's identification number. After calculating index values for multiple strains, the computer can compare the obtained index values to obtain information about culture strains with high productivity of the target protein (information about the identification numbers of such culture strains). Note that the formula for calculating the index value S is not limited to a linear function and can be adjusted as appropriate considering the contribution of each relative quantity.
[0023] As described above, the present invention may use a computer to estimate the productivity of a target protein in a host cell. In other words, the present invention also provides a method for estimating the productivity of a target protein in a host cell using a computer, and a device for estimating the productivity of a target protein in a host cell using a computer. Furthermore, the present invention also provides a program that causes a computer to execute a method for estimating the productivity of a target protein in a host cell, and a non-temporary information recording medium that stores such a program. Examples of non-temporary information recording media include DVDs, CDs, hard disks, USB memory sticks, and SD cards. The computer described above can be said to be a computer (a device for estimating the productivity of a target protein in a host cell) having a productivity estimation unit for estimating the productivity of a target protein in a host cell using the three relative quantities described above.
[0024] A computer has an input unit, an output unit, a control unit, an arithmetic unit, and a memory unit, and each element is connected by a bus or the like to enable the exchange of information. Computers usually handle digital information. For example, a computer may store programs or various kinds of information in its memory unit. When predetermined information is input from the input unit, the control unit reads the program stored in the memory unit. The control unit then reads the information stored in the memory unit as appropriate and transmits it to the arithmetic unit. The control unit also transmits the input information to the arithmetic unit as appropriate. The arithmetic unit performs calculations using the received information and stores it in the memory unit. The control unit reads the calculation results stored in the memory unit and outputs them from the output unit. In this way, various processes and steps are executed. Each unit and each means is responsible for executing these various processes. A computer may have a processor, and the processor may implement various functions and steps. A computer may be standalone. A computer may have some of its functions distributed between a server and terminals. In that case, it is preferable that the server and terminals can exchange information via a network such as the internet or an intranet. A computer may include a processor and memory connected to the processor. The memory may store instructions, and when executed by the processor, these instructions may cause the computer to perform various processes and function as various components. The computer may build a learning model by providing various training data and perform various calculations through machine learning. In this case, the computer may perform various analyses and interpretations using the learning model created by AI (artificial intelligence) machine learning and deep learning.
[0025] How to build a pre-trained model To estimate the productivity of a target protein in a host cell, a trained model may be constructed using training data. To construct a trained model, for example, a dataset is prepared for a cultured host cell line, using the three relative quantities described earlier (relative quantity of DNA fragments, relative quantity of mRNA transcripts, and relative quantity of the target protein) as features and the production amount (relative quantity) of the target protein as a label. These datasets are used as labeled training data with answers to train an AI model. The AI learns from these datasets and is trained to predict the correct label for new data. In this way, a trained model with the three relative quantities as features can be constructed. Furthermore, the accuracy of the trained model can be improved by inputting sample features into the trained model and feeding back the resulting label, which is the production amount (relative quantity) of the target protein (model learning). This invention also provides a trained model for estimating the productivity of a target protein in a host cell. Moreover, the computer described above may further possess a trained model for estimating the productivity of a target protein in a host cell.
[0026] The three relative quantities that constitute the features (relative quantity of DNA fragment, relative quantity of mRNA transcript, and relative quantity of target protein) may be values after dimensionality reduction. Methods for dimensionality reduction are already known. Therefore, when performing dimensionality reduction on three relative quantities that are features, any known method can be appropriately adopted. Examples of dimensionality reduction methods include principal component analysis (PCA), t-SNE (t-Distributed Stochastic Neighbor Embedding), UMAP (Uniform Manifold Approximation and Projection), self-organizing maps (SOM), and autoencoders. These are used to reduce the number of dimensions while preserving the characteristics of the data, thereby simplifying data visualization and analysis. A preferred example of a dimensionality reduction method is principal component analysis. In this case, the values after dimensionality reduction can be used as the three relative quantities that are features. A computer may have a dimensionality reduction processing unit to obtain the values after dimensionality reduction. When three relative quantities are input, the dimensionality reduction processing unit performs dimensionality reduction on one or more of the three relative quantities to obtain the three relative quantities after dimensionality reduction. The computer then inputs the three obtained relative quantities after dimensionality reduction as features into a trained model. The computer described above may also have a further dimensionality reduction processing unit. This point applies not only when building a pre-trained model, but also when using a pre-trained model.
[0027] A method for estimating the productivity of a target protein in a host cell using a pre-trained model. The computer reads, for example, the three relative quantities described earlier (the relative quantity of the DNA fragment, the relative quantity of the mRNA transcript, and the relative quantity of the target protein) from its memory unit (or the three relative quantities are input into the computer). The three relative quantities may be determined by the computer, or they may be values directly input into the computer from the measurement system. The computer inputs three relative quantities into a pre-trained model. As described above, the pre-trained model is trained using multiple datasets so that when features are input, it can output an index value related to the productivity of the target protein, which is the label. Therefore, when the three relative quantities are input into the pre-trained model, a value (index value) related to the productivity of the target protein in the host cell can be obtained. To reiterate, the three relative quantities may be the values after dimensionality reduction. The computer inputs these three relative quantities after dimensionality reduction as features into the trained model. The trained model then outputs an index value related to the productivity of the target protein, which is the label. The index values obtained in this way are stored in the memory unit as appropriate, and when evaluating the productivity of multiple stocks, they can be read out and compared by the calculation unit.
[0028] Examples of methods for estimating the productivity of a target protein in a host cell. A gene for expressing the target protein is introduced into a host cell. At this time, genes for expressing resistance proteins and genes used for analysis may also be introduced into the host cell. In this way, a host cell genetically modified to express the target protein is obtained. Genetically modified host cells may be cultured as appropriate. Multiple strains of the host cells are then obtained. At this stage, it is impossible to predict which strain (group) has a high capacity to produce the target protein. In fact, culturing all strains for about two weeks would allow for the prediction of which strains have a high capacity to produce the target protein. However, culturing all strains for two weeks is not very productive. Therefore, the relative quantities of each strain are obtained as follows to estimate the productivity of the host cells in producing the target protein.
[0029] Using the method described above, the measurement system obtains the relative amounts of a DNA fragment, mRNA transcript, and target protein from a given strain. The three obtained relative amounts are stored in the measurement system's memory, associated with strain identification information as appropriate. These three relative amounts, along with the strain identification information, are then input into a computer. The three relative amounts input into the computer may undergo dimensionality reduction processing by a dimensionality reduction processing unit. The computer then inputs the three relative amounts as features into a trained model. The trained model then outputs index values related to the productivity of the target protein, which are labels corresponding to the three input relative amounts. The output index values are stored in the memory, associated with the identification information below as appropriate. Similarly, the identification information and index values below are stored in the memory for multiple strains. The computer reads the multiple index values and has the calculation unit perform a calculation to compare the index values. The identification information of the strain that gives a high index value is then stored in the memory as the identification information of the high-productivity strain, which is a culture strain with high productivity of the target protein. The computer outputs identification information for this high-productivity strain. For example, the computer may display this identification information for the high-productivity strain on a display unit in a way that is visible to a human. In this way, the productivity of the target protein in the host cell can be estimated. Furthermore, in this manner, the computer can estimate which strain produces the highest amount of the target protein from among several host cell lines that have undergone gene transfer.
[0030] Method for producing the target protein Next, we will describe the method for producing the target protein. This method involves creating genetically modified cells and obtaining multiple culture lines. Then, for each of the obtained culture lines, the productivity of the target protein is estimated using the method for estimating the productivity of the target protein described above. Using these results, a high-productivity strain of the target protein is selected (high-productivity strain selection step). Subsequently, the selected high-productivity cell line is cultured to obtain the target protein (target protein acquisition step). The cell line culture method and protein acquisition method (including purification method) are publicly known. Therefore, the target protein should be obtained using a culture method suitable for the gene-transformed host cells and a method suitable for obtaining the target protein. In this way, the target protein can be produced effectively. The obtained target protein can then be used as appropriate according to its intended application. [Examples]
[0031] The sequence of a humanized IgG antibody encoding trastuzumab was incorporated into an expression vector containing a drug (geneticin) resistance cassette. This vector was then introduced into host CHO cells by electroporation. This introduction procedure was repeated twice. The cells from the first introduction were cultured for 7 days from the day after vector introduction in a cell line development medium containing geneticin, under static conditions of 37°C and 5% CO2 concentration. Subsequently, under the same conditions, the cells were transferred to a shaking culture with a rotation diameter of 20 mm and a shaking speed of 120 rpm, and cultured for another 7 days. The resulting cells were used as a stable expression cell line.
[0032] For the second vector-transferred cells, the cells were stained with a fluorescently labeled (FITC) anti-human IgG antibody the day after vector transfer and used as samples for cell sorting. Cell sorting was performed using a cell sorter manufactured by Sony (registered trademark). The sample gating procedure was as follows: First, doublet cells were removed to obtain single cells. Furthermore, the single cells were fractionated based on the signals of the FSC area and BSC area. Cells from these fractions were gated based on fluorescence intensity and used as seeded cells for cell sorting. The seeded cells were seeded in tens of cells per well into a 96-well culture plate lined with isolation medium containing Geneticin. The seeded cells were cultured for 14 days under static culture conditions at a temperature of 37°C and a CO2 concentration of 5%, and cells were harvested from wells where cell proliferation was confirmed to obtain cloned cells.
[0033] After harvesting, the cloned cells were divided into two cell groups. The first cell group underwent expansion culture and then production culture using a 10-day fed-batch method together with a separately prepared stable-expression cell line. After the completion of production culture, the antibody concentration in the culture medium at 10 days of culture was measured using Protein A affinity HPLC (Figure 1). In the second cell group, each cloned cell was further divided into three groups, and genomic DNA extraction, RNA extraction, and staining with fluorescent (FITC)-labeled anti-human IgG antibody were performed on each cell group. The same procedure was also performed on a stable-expression cell line as a control.
[0034] For the extracted genomic DNA, the relative copy number of the trastuzumab-coding DNA sequence in each cloned cell and stable-expressing cell line was quantified using real-time PCR with a TaqMan probe. Figure 2 shows the relative quantification results of the copy number in each cloned cell, with the copy number of the stable-expressing cell line set to 1. For the extracted RNA, the amount of mRNA transcribed from the trastuzumab-coding sequence in each cloned cell and stable-expressing cell line was analyzed relatively using 1-step RT-qPCR with a TaqMan probe. Figure 3 shows the analysis results of the relative quantification of mRNA in each cloned cell, with the mRNA amount of the stable-expressing cell line set to 1.
[0035] For cells stained with fluorescently labeled antibodies, the Stain Index was calculated using a flow cytometer manufactured by Cytek Biosciences®, comparing the fluorescence intensity of stained cells to that of unstained cells. Equation (1) was used to calculate the Stain Index. The results of the Stain Index calculation for each clonal cell are shown in Figure 4.
[0036] Stain Index = MFI(pos)-MFI(neg) / 2σ (1) MFI(pos): Mean fluorescence intensity of stained cells MFI(neg): Mean fluorescence intensity of unstained cells σ: Standard deviation
[0037] Figure 5 and Table 1 show the bivariate relationships between copy number, mRNA quantity, and Stain Index and antibody concentration obtained by fed-batch culture. The criteria for determining the correlation coefficient R were as follows: R=0~0.2: no correlation, R=0.2~0.4: weak correlation, R=0.4~0.7: moderate correlation, R=0.7~1.0: strong correlation. As a result, copy number, mRNA quantity, and Stain Index showed a moderate correlation with antibody concentration.
[0038] [Table 1]
[0039] Next, the copy number, mRNA quantity, and Stain Index values were standardized using equation (2) and converted to X'.
[0040] X' = X - μ / σ (2) X: Value before standardization μ: Mean value before standardization σ: Standard deviation before standardization
[0041] Principal component analysis (PCA) was performed using JMP statistical analysis software with standardized copy number, mRNA quantity, and Stain Index X'. The results of the PCA analysis are shown in Figure 6, and the values of the first and second principal components obtained from the PCA analysis are shown in Table 2. The bivariate relationship between the first and second principal components and antibody concentration obtained from fed-batch culture is shown in Figure 7 and Table 3, respectively. As a result, a strong correlation was observed between the first principal component and antibody concentration. Furthermore, since the principle of this method is based on phenomena that follow the universal central dogma of living organisms and extracellular secretion mechanisms, it can be applied not only to humanized IgG antibodies encoding trastuzumab, but also to antibody-expressing cells such as human antibodies and chimeric antibodies, and to all expression cells expressing recombinant proteins.
[0042] [Table 2-1]
[0043] [Table 2-2]
[0044] [Table 2-3]
[0045] [Table 3] [Industrial applicability]
[0046] This invention can be used in the bio-industry, pharmaceutical industry, and other fields because it allows for the estimation of the productivity of a target protein.
Claims
1. The relative amount of the DNA fragment encoding the target protein in a host cell genetically modified to express the target protein, The relative amount of the mRNA transcript encoding the target protein in the host cell, and The relative amount of the target protein present on the surface of the host cell A method for estimating the productivity of a target protein in the host cell using the above.
2. A method according to claim 1, wherein the host cell is a CHO cell.
3. A method according to claim 1, wherein the relative amount of the DNA fragment encoding the target protein and the relative amount of the mRNA transcript encoding the target protein in the host cell are values obtained using the amount of the DNA fragment encoding the target protein and the amount of the mRNA transcript encoding the target protein, which are determined by real-time PCR.
4. A method according to claim 1, wherein the relative amount of the target protein present on the surface of the host cell is a value determined using the amount of the target protein present on the surface of the host cell, measured using a fluorescently labeled antibody that specifically binds to the target protein.
5. The method according to claim 1, A method comprising the step of inputting the relative amounts of the DNA fragment, the relative amounts of the mRNA transcript, and the relative amounts of the target protein into a trained model.
6. The method according to claim 1, A step of obtaining dimensionality-reduced relative amount data by reducing the dimensionality of one or more of the relative amounts of the DNA fragment, the mRNA transcript, and the target protein. A method comprising the step of estimating the productivity of a target protein in the host cell using the relative quantity data after dimensionality reduction.
7. A high-productivity strain selection step, which is a step of selecting a high-productivity strain, which is a culture strain with high productivity of the target protein, using a method for estimating the productivity of the target protein according to any one of claims 1 to 6, The process includes a step of culturing the high-productivity strain and obtaining the target protein, Method for producing the target protein.