Method for producing protein, and information processing device supporting same

A method for protein production that divides and mutates domains based on fitness and functional tests, supported by an information processing device, efficiently produces proteins with enhanced functions by minimizing verification testing.

WO2026047843A1PCT designated stage Publication Date: 2026-03-05HITACHI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/030456
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-27
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing methods for mutating amino acid sequences in proteins, such as CAR-T cells, require extensive verification testing, making them cumbersome and difficult to apply effectively.

Method used

A method that divides a template protein into domains, selects target domains for mutation, evaluates mutant domains based on fitness and functional tests, and replaces target domains with useful mutants, supported by an information processing device for efficient protein production.

Benefits of technology

Enables the reliable production of proteins with improved functions while significantly reducing the verification test burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024030456_05032026_PF_FP_ABST
    Figure JP2024030456_05032026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention makes it possible to surely obtain a target protein having an improved function from a template protein with a small verification test load. The present invention sets a plurality of mutation domains in each of which a portion of the amino acid sequence for a target domain of a template protein is mutated, determines the fitness of each of the plurality of mutation domains, and evaluates each of the plurality of mutation domains on the basis of the correlation between the fitness and the result of a functional test.
Need to check novelty before this filing date? Find Prior Art

Description

Protein production method and information processing device supporting the method

[0001] The present invention relates to a method for producing a protein, and more particularly to a method for producing a protein in which a target protein having improved functionality compared to a template protein is obtained by mutating a part of the amino acid sequence of the domain based on the domain of the template protein, and an information processing device for supporting this method.

[0002] Numerous functional proteins are known to have various physiological roles and functions in the body. For example, the CAR in chimeric antigen receptor (CAR) T cells is an artificial protein designed to enable T cells to recognize specific antigens and attack target cells. In recent years, there has been a surge in the development of treatments using CAR-T cells, primarily for blood cancers. Examples include Kymriah, Yescarta, and Breyanzi, which are T cell therapies incorporating anti-CD19 chimeric antigen receptors.

[0003] In these CAR-T cell therapies, new CARs are being developed with the aim of adding additional functions to T cells. For example, Greenhalgh, JC et al., Machine learning-guided acyl-ACP reductase engineering for improved in vivo fatty alcohol production. Nat. Commun. 12, 5825 (2021) (Non-Patent Document 1) describes a method for designing chimeric proteins in which a protein is divided into domains, some domains are further broken down into blocks, and variations of the domains and blocks are combined.

[0004] Japanese Patent No. 688749 (Patent Document 1) describes experimental evaluation of a CAR in which an amino acid mutation has been introduced into a costimulatory domain derived from human 4-1BB. Furthermore, Hopf, T. et al., "Mutation effects predicted from sequence covariation. Nat. Biotechnol. 35, 128-135 (2017)" (Non-Patent Document 2) describes a method for predicting the effect of amino acid mutations on protein function.

[0005] Patent No. 688749

[0006] Greenhalgh, JC et al., Machine learning-guided acyl-ACP reductase engineering for improved in vivo fatty alcohol production. Nat. Commun. 12, 5825 (2021)Hopf, T. et al., Mutation effects predicted from sequence covariation. Nat. Biotechnol. 35, 128-135 (2017)

[0007] As described above, protein design techniques that mutate a part of the basic amino acid sequence are used to improve or enhance the function of proteins. Known methods for this include random mutation, directed evolution, and a method based on three-dimensional structure, but all of these methods require a heavy burden of verification testing, making them difficult to apply to artificial proteins such as CARs.

[0008] Therefore, an object of the present invention is to provide a protein production method that enables a target protein having improved function from a template protein to be reliably obtained with a small verification test load, and an information processing device that supports this method.

[0009] In order to achieve the above-mentioned object, the present invention provides a method for producing a protein based on a template protein, comprising: a first step of dividing the template protein into multiple domains; a second step of selecting a target domain from the multiple domains; a third step of setting multiple mutant domains by mutating part of the amino acid sequence of the target domain; a fourth step of determining the fitness of each of the multiple mutant domains based on the amino acid sequence of the mutant domain; a fifth step of conducting a functional test of each of the multiple mutant domains; a sixth step of evaluating each of the multiple mutant domains based on the correlation between the fitness and the results of the functional test; and a seventh step of selecting a useful mutant domain from the multiple mutant domains that is useful for replacing the target domain based on the results of the evaluation, wherein the method is characterized by producing a protein in which the target domain has been replaced with the useful mutant domain.

[0010] Furthermore, a second invention is an information processing device in which a processor supports the production of a protein based on a template protein based on a program stored in memory, wherein, when a user inputs the template protein, the processor divides the template protein into a plurality of domains and notifies the user of these domains, allows the user to select a target domain from the plurality of domains whose amino acid sequence is to be mutated, collects a plurality of similar amino acid sequences that are similar to the amino acid sequence of the target domain, selects a plurality of similar amino acid sequences from the collected similar amino acid sequences as mutant domains in which a portion of the amino acid sequence of the target domain is mutated, registers functional test measurement values ​​collected by the user for each of the plurality of mutant domains, evaluates each of the plurality of mutant domains based on a fitness calculated based on the amino acid sequence and a correlation between the fitness and the measurement results of the functional test, selects a useful mutant domain from the plurality of mutant domains that is useful for replacing the target domain based on the results of the evaluation, and notifies the user of the amino acid sequence in which the target domain has been replaced with the useful mutant domain.

[0011] As described above, the present invention can provide a protein production method that enables a target protein having improved function from a template protein to be reliably obtained with a small verification test load, and an information processing device that supports this method.

[0012] 1 is a process diagram of an embodiment of a protein production method according to the present invention. FIG. 2 is a block diagram showing an example of a template protein model. FIG. 3 shows an example of an amino acid sequence of target domain A and an example of a similar amino acid sequence. FIG. 4 shows an example of an amino acid sequence of target domain B and an example of a similar amino acid sequence. FIG. 5 shows an example of an amino acid sequence of target domain C and an example of a similar amino acid sequence. FIG. 6 is an example of a calculation of the fitness of an amino acid sequence related to domain A. FIG. 7 is an example of a calculation of the fitness of an amino acid sequence related to domain B. FIG. 8 is an example of a calculation of the fitness of an amino acid sequence related to domain C. FIG. 9 is an example of an example of a calculation of the fitness of a mutated domain relative to target domain A and the measured values ​​of a functional activity test. FIG. 10 is an example of an example of a calculation of the fitness of a mutated domain relative to target domain B and the measured values ​​of a functional activity test. FIG. 11 is an example of an example of a calculation of the fitness of a mutated domain relative to target domain C and the measured values ​​of a functional activity test. FIG. 12 is a hardware block diagram of an embodiment of an information processing device. FIG. 13 is a functional block diagram of a support system for supporting a protein production method. FIG. 14 is a detailed functional block diagram of a mutated domain generation module. FIG. 15 is a detailed functional block diagram of a mutated domain generation module. FIG. 16 is a template protein registration table. FIG. 17 is a domain registration table. 1 is a similar amino acid sequence collection policy setting table; 2 is a similar amino acid sequence table; 3 is a mutation domain registration table; 4 is a mutation domain analysis table; 5 is a useful mutation domain combination registration table; 6 is a correlation diagram between the cytotoxic activity and fitness of a mutation domain (target domain: hinge domain); 7 is a correlation diagram between the cytotoxic activity and fitness of a mutation domain (target domain: costimulatory domain); 8 is a graph of the cytotoxic activity of a mutation domain (target domain: hinge domain and costimulatory domain).

[0013] An embodiment of a protein production method according to the present invention will be described. The protein production method according to the present invention includes a protein design process, which is known as a means for developing pharmaceuticals and chemicals and includes a process for designing or creating proteins with specific functions. The method of the present invention is based on a domain of a template protein, and mutates a portion of the amino acid sequence of the domain to obtain a target protein with improved function compared to the template protein. Figure 1 is a flow chart of an embodiment of a protein production method.

[0014] In step S1, a template protein is determined as a starting material for the target protein. The template protein may be a known one, for example, an artificial protein that has the ability to bind to a target and whose target-binding ability can be improved by mutation of its amino acid sequence, more specifically, the above-mentioned CAR. The template protein may be an enzyme, antibody, antigen, hormone, or the like, whose amino acid sequence is known.

[0015] In step S2, the template protein is divided into domains. In known CARs, the multiple domains that make up the CAR are known. If the domain structure of the template protein is unknown, the domains may be newly defined. In step S3, one or more target domains in which part of the amino acid sequence is to be mutated are selected from all domains of the template protein.

[0016] In step S4, for each target domain, similar amino acid sequences are collected that are similar to the amino acid sequence (original sequence). This can be done, for example, by referring to sources such as scientific literature and databases. The original sequence can also be found from sources such as scientific literature and databases. Similarity refers to, for example, the number of amino acids in the target domain, the types of amino acids, and the amino acid sequence pattern, which are similar and at least maintain the performance and function of the original sequence. For example, similar amino acid sequences refer to sequences in which the number of similar amino acids is the same as the original sequence or differs only slightly, and a portion (one or several) of the similar amino acid sequence is replaced with other amino acids, or a portion of the original sequence is deleted, or a small number of amino acids is added to the similar amino acid sequence. The similar amino acid sequence may be the entire protein or polypeptide, or a portion thereof.

[0017] In step S5, the fitness of each of the multiple similar amino acid sequences is calculated. Fitness is a measure of how advantageous a specific amino acid substitution in a protein is for the survival and reproduction of an organism, and is calculated from the perspective of how frequently a specific amino acid substitution occurs during the evolutionary process. As described below, fitness was calculated using the computational model described in the aforementioned Non-Patent Document 2.

[0018] In step S6, a plurality of similar amino acid sequences are selected for each of a plurality of target domains according to fitness, superiority, etc., and a predetermined number of similar amino acid sequences are selected in descending order of fitness as mutation domains related to the target domain, thereby setting a plurality of mutation domains in which a portion of the amino acid sequence of the target domain has been mutated.

[0019] Next, in step S7, a measurement test (demonstration test) of the function of the template protein is performed for the selected multiple mutation domains. In step S8, a predetermined number of mutation domains are extracted from the multiple mutation domains as useful mutation domains that are useful for replacing the target domain, for each of the multiple target domains, based on an evaluation of the fitness and measured values ​​of functional activity for the multiple mutation domains, for example, an evaluation of the correlation between the measured value and fitness. Here, the functional test is, for example, a test of the cytotoxic activity of CAR-T cells against target cancer cells, assuming that the template protein is the above-mentioned CAR.

[0020] In step S10, the target domain is replaced with the useful mutation domain. In step S11, the amino acid sequence of the template protein in which the target domain has been replaced with the useful mutation domain is reported as the amino acid sequence of the target protein. A target protein is a protein derived from a template protein, and whose function has been improved to a predetermined level. Note that specific aspects of the functional activity test and the generation of the target protein vary depending on the attributes and type of the template protein, but specific aspects assuming that the template protein is a CAR are described below.

[0021] The steps in Fig. 1 will be explained in detail with reference to the drawings. Fig. 2A is a block diagram showing an example of a template protein model. This template protein model explains that it comprises domain A, domain B, and domain C. Regarding CARs, first-generation CARs are known to be composed of a binding domain that binds to a target antigen, a hinge domain, a T cell transmembrane domain, and an activation domain, in that order. Furthermore, second-generation CARs are known to have a costimulatory domain between the transmembrane domain and the activation domain. Several amino acid sequences are known to constitute each domain.

[0022] FIG. 2B shows the amino acid sequence A that constitutes the target domain A. 0 Here is an example. Each letter corresponds to the one-letter code of an amino acid. This will also be the case in the following explanations. 1 , A 2 , A3 ...A n Each of the amino acid sequences A and B of the target domain A is 0 Examples of multiple similar amino acid sequences are shown below. 1 , A 2 , A 3 ...A n Each of the domains A and B is denoted by a. 0 In the same position as above, the amino acid is mutated as shown in the underlined part.

[0023] FIG. 2C shows the amino acid sequence B constituting the domain of interest B. 0 An example of B is shown below. 1 , B 2 , B 3 ...B n Each of the amino acid sequences B of the target domain B 0 shows several similar amino acid sequences similar to B 1 , B 2 , B 3 ...B n Each of the above has one amino acid mutation at the position indicated by b1 or b2, as underlined.

[0024] FIG. 2D shows the amino acid sequence C that constitutes the domain of interest C. 0 Here is an example of C 1 , C 2 , C 3 ...C n Each of these is the amino acid sequence C of the target domain C. 0 Examples of several similar amino acid sequences are shown below. 1 , C 2 , C 3 ...C n Each of the above has one amino acid substitution at the position indicated by c1 or c2, as shown in the underlined part.

[0025] FIG. 3A shows the amino acid sequence of domain A. 0 and amino acid sequence A 0 Similar amino acid sequence to 1 , A 2 , A 3 ...A n ) is an example of calculating the fitness of each amino acid sequence. 0The fitness for similar amino acid sequences is "0". The fitness of similar amino acid sequences varies depending on the type of amino acid, the order of the amino acid sequence, and other differences in the form of the amino acid sequence. Here, an example of calculating fitness will be explained. For fitness calculation, "EVmutation" from Non-Patent Document 2 was used. Assume that the amino acid sequence σ is generated with the probability P(σ) given by equation (1).

[0026] Z is the partition function, and E(σ) is the energy function given by equation (2). h i is the constraint term for amino acid residue i, and J ij represents the constraint term for all amino acid residue combinations (i, j). If the original sequence is (wild) and the mutated sequence is (mutant), the energy difference between the two is expressed by equation (3).

[0027] In EVmutation, the parameters of equations (1) and (2) are determined based on a multiple sequence alignment consisting of the original sequence (wild) and the mutated sequence (mutant), and ΔE is calculated. In the present invention, ΔE is treated as fitness.

[0028] FIG. 3B shows the amino acid sequence of domain B. 0 and a similar amino acid sequence (B 1 , B 2 , B 3 ...B n ) is an example of calculating the fitness of each. 0 and a similar amino acid sequence (C 1 , C 2 , C 3 ...C n ) is an example of calculating the fitness of each.

[0029] The fitness of a similar amino acid sequence can be used to estimate the suitability of that sequence to become part of a template protein. Therefore, for each similar amino acid sequence in domains A, B, and C, mutation domains are selected as useful mutation domains in descending order of fitness. The reason for selecting a portion of the collected similar amino acid sequences as mutation domains is to limit the scope of functional activity testing and reduce the burden of verification testing. Details of the measurement test are described below.

[0030] FIG. 4A shows the target domain A (A 0 ) Mutation domain A 1 ...A k The template polypeptide is (A 0 -B 0 -C 0 ) is expressed by the target domain A (A 0 ) among the n mutation domains, k useful mutation domains A 1 ...A k is selected. Fitness X A1 ..X Ak is a useful mutation domain A 1 ...A k The measured value Y is the fitness of each template protein (A 0 -B 0 -C 0 ) is the measured value of the measured value Y A1 ...Y Ak is a useful mutation domain A 1 ...A k Each of them corresponds to the target domain A (A 0 ) is replaced, and the non-target domain is the original (B 0 , C 0 ) protein (A 1 -B 0 -C 0 ) ... (A k -B 0 -C 0 ) are the respective measured values. It can be seen that each of the multiple useful mutation domains exhibits a different fitness and measured value.

[0031] FIG. 4B shows the target domain B (B0 ) Useful mutation domain B 1 ...B k Figure 4C shows the fitness and activity test measurements for each of the target domains C (C 0 ) Useful mutation domain C 1 ...C k The fitness and activity test measurements of each of the above are shown.

[0032] Domain A 0 Mutation domain, domain B 0 The mutated domain of and domain C 0 For each of the mutation domains, a useful mutation domain is determined from each mutation domain based on a combined evaluation of the fitness and the measured value. For example, an amino acid sequence that becomes a useful mutation domain is determined based on a combined evaluation of the correlation between the fitness and the measured value and the activity value.

[0033] The method then comprises: 1 , A 2 , A 3 ), a useful mutation domain of the target domain B (e.g., B 1 , B 2 , B 3 ), and a useful mutation domain C of the target domain C (e.g., C 1 , C 2 , C 3 ) are combined between multiple target domains A, B, and C to form an amino acid sequence consisting of useful mutation domain A, useful mutation domain B, and useful mutation domain C, which are used as multiple patterns of the amino acid sequence of the target protein (Figure 4D). Based on the example described here, a total of 27 patterns of amino acid sequences are determined. Each of these multiple patterns is a candidate amino acid sequence for the target protein with improved function of the template protein. Proteins can be generated from each pattern, and the target proteins can be narrowed down based on the relative merits of the measured functional activity of each pattern.

[0034] According to the protein production method of the present invention described above, it is not necessary to perform verification tests for all amino acid substitutions, but rather verification tests can be performed on limited targets, namely, the mutant domain and candidate amino acid sequences for the target protein. This makes it possible to reliably produce a target protein with improved function from a template protein while reducing the verification test burden.

[0035] Next, an embodiment of an information processing device for supporting a protein production method will be described. Fig. 5 is a hardware block diagram of the information processing device 10. The information processing device 10 has a processor 100, a main memory 102, a network interface 104, a display interface 106, an input / output interface 108, and an auxiliary storage device 110.

[0036] The processor 100 and main memory 102 execute the program in the auxiliary storage device 110 to drive a support system 20 ( FIG. 6 ) that supports a protein production method. The support system 20 includes a mutation domain generation module 30 and a mutation domain evaluation module 40. A module is a function of the program and may be replaced with other terms such as means, unit, or element. A module may be realized by dedicated hardware such as an integrated circuit.

[0037] 7A shows a detailed functional block diagram of the mutation domain generation module 30. The template protein determination module 300 determines a template protein, which is the starting material for producing a target protein, based on input from the user, and stores the amino acid sequence ID and amino acid sequence of the template protein in a template protein registration table (FIG. 8). This table is recorded in a predetermined area of ​​the auxiliary storage device 110. The same applies to the tables described below.

[0038] The target domain determination module 302 divides the template protein into multiple domains and registers them in the domain registration table of FIG. 9. This table includes input areas for the domain name, the amino acid sequence of the domain, and whether the user wants to select it, and is notified to the user. By entering information into this table, the user can select a target domain from multiple domains in which some amino acids in the amino acid sequence will be mutated. "Y" indicates selection, and "N" indicates non-selection. FIG. 9 shows, as an example, that domains A, B, and C have been selected.

[0039] The similar amino acid sequence collection policy setting module 304 registers policies, which are conditions for collecting similar amino acid sequences that are similar to the amino acid sequence of the target domain, in a similar amino acid sequence collection policy setting table (FIG. 10) based on user input. This table is configured so that the user can register, for each target domain, the number of amino acid mutations (substitutions) for the target domain, for example, "1" or "2." Allowing the user to set the policy increases the degree of design freedom in determining the similarity range of the target domain.

[0040] The similar amino acid sequence collection module 306 refers to the policy setting table (FIG. 10), searches scientific information sources such as databases (DB), and registers the similar amino acid sequences collected for each target domain in the similar amino acid sequence table (FIG. 11). When registering the collected similar amino acid sequences in the table, the similar amino acid sequence collection module 306 registers information on amino acid mutations for each similar amino acid sequence ID in the table. The amino acid mutation information is information that indicates the positions where the collected similar amino acid sequences differ from the amino acid sequence of the target domain and the details of the mutations. This information is expressed according to the HGVS naming rules. Note that A 1 , A 2 , A 3 The "A" in ... indicates that the target domain is domain A. The same applies to B and C.

[0041] The similar amino acid sequence reading module 308 reproduces and reads a similar amino acid sequence for each similar amino acid sequence ID in the similar amino acid sequence registration table (FIG. 11) based on the information on amino acid mutations and the information on the amino acid sequence of the target domain (FIG. 9), and transmits this to the fitness calculation module 310. The fitness calculation module 310 executes calculation processing based on the fitness calculation example described above, and records the fitness of each of the collected multiple similar amino acid sequences in the similar amino acid sequence registration table (FIG. 11).

[0042] The similar amino acid sequence selection module 312 refers to the similar amino acid sequence registration table (FIG. 11), and based on user input, selects a predetermined number or a predetermined percentage of similar amino acid sequences for each target domain, primarily in descending order of fitness, and registers the selection information in the table (FIG. 11). The mutation domain determination module 314 refers to the similar amino acid sequence registration table (FIG. 11), and outputs the selected similar amino acid sequences as mutation domains.

[0043] 7B is a functional block diagram illustrating the details of the mutation domain evaluation module 40. The mutation domain input module 400 receives the output from the mutation domain determination module 314 and registers the mutation domain in the mutation domain registration table (FIG. 12). 1 ~A n are mutation domains for the target domain A, and B 1 ~B n are mutation domains relative to the target domain B, and C 1 ~C n are mutation domains for the target domain C. The mutation domain input module 400 generates a protein domain structure (FIG. 12) for each mutation domain. The domain structure indicates what domains the template protein has. For example, the coordinate ID "1" indicates the mutation domain A. 1 and the original domain B 0 , C 0The mutant domain input module 400 recreates the amino acid sequence related to the domain configuration and registers it in the mutant domain registration table (FIG. 12). Fitness is calculated for the amino acid sequence of the domain that is a combination of the mutant domain and the original domain.

[0044] The functional activity measurement processing module 402 references the mutation domain registration table (FIG. 12) and outputs it to the user. The user generates a protein based on the amino acid sequence for each sequence ID and performs a functional activity test on it. The functional activity measurement processing module 402 registers the test measurement value in the mutation domain registration table (FIG. 12) based on the input from the user. When this registration is complete, the mutation domain input module 400 notifies the mutation domain analysis module 404 of the fact.

[0045] The mutation domain analysis module 404 evaluates the mutation domains based on their fitness and functional activity measurements, and extracts a specific number of mutation domains suitable as useful mutation domains from the mutation domains (this specific number is set in the mutation domain analysis module 404 by a specific number input module 406 upon receiving input from the user) as useful domains. The details are as follows.

[0046] The mutation domain analysis module 404 analyzes the target domain by analyzing each of the multiple combinations (A l , A m , A n), the correlation coefficient (Y) between fitness and functional activity measurement value, the average (X) of the functional activity measurement values ​​of the three mutant domains, (X) * (Y) are calculated, and these are registered in the mutant domain analysis table (Figure 13). Among the combinations of multiple mutant domains for which correlation can be determined, the higher the correlation (Y), the better the balance between the suitability of the mutant domain to become a target domain and the property of improving the function of the template protein, and therefore it can be said to be preferable. On the other hand, the higher (X), the better the mutant domain is at improving the function of the template protein. Therefore, secondly, an evaluation is performed that integrates the correlation and functional activity, and the larger the product of the correlation and functional activity, the better the domain is as a different domain to replace the target domain. Therefore, for example, based on Figure 13, for target domain A, the mutant domain (A 2 , A 3 , A 4 ) are suitable as useful mutation domains. Explanation of target domains B and C will be omitted.

[0047] The mutation domain analysis table (FIG. 13) has an input area for the user to select or not select. The mutation domain analysis module 404 displays the mutation domain analysis table to the user, providing the user with an opportunity to input whether or not to select. The user will normally input a selection in an area where the (X)*(Y) value is large in each of the target domains A, B, and C, but this does not prevent the user from inputting a selection in an area where this is not the case. Note that the mutation domain analysis module 404 may determine the mutation domain combination without relying on user input.

[0048] When the mutation domain analysis module 404 selects useful mutation domains for all target domains, the useful mutation domain combination generation module 408 generates patterns of combinations of useful mutation domains between multiple target domains and registers them in a useful mutation domain combination registration table (FIG. 14). In FIG. 14, for convenience, the useful mutation domain for target domain A is designated as A 1 -A 3 The useful mutation domain for the target domain B is B 1 -B 3 The useful mutation domain for the target domain C is C1 -C 3 Therefore, there are a total of 27 patterns of combinations of useful mutation domains. Module 408 references tables (FIGS. 9 and 11), reproduces the amino acid sequence for each pattern, and registers it in table (FIG. 14). Module 408 notifies the user of all patterns of amino acid sequences based on table (FIG. 14) as a target protein candidate list 410. This allows the user to know the target protein candidates, and by conducting demonstration experiments on a limited number of target protein candidates, the target protein can be reliably and efficiently obtained based on the results.

[0049] Here, an example of a functional activity test will be described. The template protein was CAR, and the target domains were the hinge domain and costimulatory domain. Preparation of mRNA for CAR expression: A plasmid vector was constructed in which CAR, P2A, GFP, and a polyadenylation sequence were placed downstream of the spleen focus-forming virus (SFFV) promoter. DNA for in vitro transcription was prepared by PCR using this plasmid vector as a template, and mRNA was synthesized by in vitro transcription. From 5' to 3', the CAR contained a signal peptide, an anti-CD19 scFv, a CD28-derived hinge domain, a transmembrane domain and costimulatory domain, and a CD3ζ activation domain. For the hinge domain and costimulatory domain, single-mutated hinge domains and single-mutated costimulatory domains with single amino acid substitutions were designed.

[0050] Preparation of CAR-T cells: CAR expression mRNA was introduced into human CD8-positive T cells by electroporation to prepare CAR-T cells. Measurement of cytotoxic activity: Luciferase-expressing Nalm cells were seeded onto a 384-well plate, and then CAR-T cells were added to initiate co-culture. After co-culture was completed, surviving Nalm cells were quantified by luciferase assay to evaluate cytotoxic activity.

[0051] An example of the operation of the mutation domain evaluation module 40 for a target protein is described below. Single amino acid substitutions (similar amino acid sequences) were selected for each of the hinge domain and costimulatory domain (target domain) in descending order of ΔE (fitness). Sixteen mutation domains were selected for the hinge domain, and 18 mutation domains were selected for the costimulatory domain. The cytotoxicity of CAR-T cells expressing these mutation domains was then measured. Figure 15 shows a correlation diagram between the cytotoxic activity and fitness of the hinge domain / mutation domain. Figure 16 shows a correlation diagram between the cytotoxic activity and fitness of the costimulatory domain / mutation domain. Each plot represents a mutation domain. Homologous sequences of the amino acid sequences of the hinge domain and costimulatory domain of CAR were searched using the EVcouplings server (https: / / v2.evcouplings.org / ), multiple alignments were created, and ΔE was calculated. UniRef100 was used as the sequence database.

[0052] 15 , the top three mutation domains in terms of measured cytotoxic activity are H1, H9, and H15, but the correlation coefficient of the regression line 1400 between fitness and cytotoxic activity for these combinations was low. On the other hand, when a search was conducted for a combination of three mutation domains that had a high measured cytotoxic activity and a high correlation with fitness, it was found that the correlation coefficient of the regression line 1402 between fitness and cytotoxic activity for the combination of H1, H13, and H15 was high, and the product of the correlation coefficient and the measured cytotoxic activity was also high. Therefore, for the hinge domain, H1, H13, and H15 were determined to be useful mutation domains.

[0053] In Figure 16, the top three mutation domains in terms of measured cytotoxic activity were C6, C7, and C10, but the correlation coefficient of the regression line 1500 between fitness and cytotoxic activity for these combinations was low. On the other hand, when a combination of three mutation domains was searched for that had a high measured cytotoxic activity and a high correlation with fitness, it was found that the correlation coefficient of the regression line 1502 between fitness and cytotoxic activity for the combination of C7, C10, and C17 was high, and the product of the correlation coefficient and the measured cytotoxic activity was also high. Therefore, C7, C10, and C17 were determined to be useful mutation domains for the costimulatory domain.

[0054] Figure 17 is a graph showing the results of a cytotoxic activity test of CAR-T cells in which the hinge domain (target domain) of a CAR was replaced with the mutant domain shown in Figure 15 and the costimulatory domain was replaced with the mutant domain shown in Figure 16. In Figure 17, "wild type" indicates the results obtained from T cells having an original CAR in which the hinge domain and costimulatory domain were not replaced with mutant domains. "Random mutant domain" indicates the results obtained from T cells having a CAR in which the hinge domain and costimulatory domain were replaced with randomly selected mutant domains. "Top mutant domain" indicates the results obtained from T cells having a CAR in which the hinge domain was replaced with H1, H9, or H15 and the costimulatory domain was replaced with C6, C7, or C10 (measurements are averages). "Highly correlated mutant domain" indicates the results obtained from T cells having a CAR in which the hinge domain was replaced with H1, H13, or H15 and the costimulatory domain was replaced with C7, C10, or C17 (measurements are averages). As can be seen from this figure, the measurement values ​​of the "highly correlated mutation domains," which have a high product of the correlation coefficient and the measured cytotoxic activity, are significantly higher (p<0.05) than the measurement values ​​of the "top mutation domains."

[0055] The above-described embodiments are merely examples, and the present invention is not limited to these embodiments. For example, while it has been described that amino acid sequences similar to a target domain are collected from information sources such as databases, they may also be collected from virtually created sequences using the learning function of AI. Furthermore, while the above-described embodiments have focused on CAR as the template protein, the present invention is not limited to this. Furthermore, the above-described embodiments also encompass, as the subject of the invention to be protected, new proteins that are target proteins and are composed of novel mutant domains.

[0056] Furthermore, in the above-described embodiment, the number of useful mutation domains for a target domain is set to "3," but this is not limited to this and may be "4" or more. However, as the number of useful mutation domains increases, correlation can be evaluated more accurately, but there is a concern that the number of patterns of target protein candidates will increase, increasing the burden of functional confirmation tests. The number of useful mutation domains to be output can be determined by taking these factors into consideration comprehensively.

Claims

1. A method for producing a target protein based on a template protein, comprising: a first step of dividing the template protein into a plurality of domains; a second step of selecting a target domain from the plurality of domains; a third step of setting a plurality of mutant domains by mutating a portion of the amino acid sequence of the target domain; a fourth step of determining the fitness of each of the plurality of mutant domains based on the amino acid sequence of the mutant domain; a fifth step of conducting a functional test on each of the plurality of mutant domains; a sixth step of evaluating each of the plurality of mutant domains based on the correlation between the fitness and the results of the functional test; and a seventh step of selecting a useful mutant domain from the plurality of mutant domains that is useful for replacing the target domain based on the results of the evaluation, wherein the method for producing the target protein has the target domain replaced with the useful mutant domain.

2. The manufacturing method described in claim 1, wherein the third step involves collecting multiple similar amino acid sequences similar to the target domain in order to set multiple mutation domains, and selecting the mutation domain from the collected multiple similar amino acid sequences based on the fitness of the amino acid sequence.

3. The production method according to claim 1 or 2, further comprising an eighth step of reporting the amino acid sequence of the template protein in which the target domain has been replaced with the useful mutation domain as the amino acid sequence of the target protein.

4. The manufacturing method described in claim 3, wherein the second step selects a plurality of target domains, the seventh step selects a plurality of useful mutation domains for each of the plurality of target domains, replaces each of the plurality of target domains with each of the plurality of specific mutation domains, and reports a plurality of patterns as the amino acid sequence of the target protein.

5. The production method according to claim 4, wherein proteins are produced based on the amino acid sequences of each of the plurality of patterns, and the protein that gives the highest result in the functional test is determined to be the target protein.

6. The method according to claim 1, wherein the template protein is an artificial protein that has the ability to bind to a target and whose ability to bind to the target can be improved by mutation of the amino acid sequence.

7. An information processing device in which a processor supports the production of a target protein based on a template protein based on a program in memory, wherein the processor, when a user inputs the template protein, divides the template protein into a plurality of domains and notifies the user of the domains, allows the user to select a target domain from the plurality of domains whose amino acid sequence is to be mutated, collects a plurality of similar amino acid sequences that are similar to the amino acid sequence of the target domain, selects a plurality of similar amino acid sequences from the collected similar amino acid sequences as mutant domains in which part of the amino acid sequence of the target domain is mutated, registers measurement values ​​of functional tests collected by the user for each of the plurality of mutant domains, evaluates each of the plurality of mutant domains based on fitness calculated based on the amino acid sequence and the correlation between the fitness and the measurement results of the functional test, selects a useful mutant domain from the plurality of mutant domains that is useful for replacing the target domain based on the results of the evaluation, and notifies the amino acid sequence in which the target domain is replaced with the useful mutant domain.

8. The information processing device according to claim 7, wherein the processor calculates fitness based on each of the collected plurality of similar amino acid sequences, and selects the mutation domain based on the calculation result.

9. The information processing device according to claim 7 or 8, wherein the processor allows a user to input conditions for collecting the plurality of similar amino acid sequences.

10. The information processing device according to claim 7, wherein the processor selects the useful mutation domain based on an integrated evaluation of the correlation and the measurement value of the functional test.

11. The information processing device according to claim 7, wherein the processor reports the amino acid sequence of the template protein in which the target domain has been replaced with the useful mutation domain as the amino acid sequence of the target protein.

12. The information processing device of claim 11, wherein the processor selects a plurality of useful mutation domains for each of a plurality of target domains set by the user, creates a plurality of combination patterns of the useful mutation domains between the plurality of target domains in replacing each of the plurality of target domains with each of the plurality of useful mutation domains, and notifies the created plurality of combination patterns as patterns of the amino acid sequence of the target protein.