How to create libraries using machine learning
By analyzing sublibraries at various stages of biopanning and using machine learning to expand sequence space, the method addresses the limitations of conventional nucleic acid library creation, resulting in a larger and more cost-effective library of functional proteins.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TOHOKU UNIV
- Filing Date
- 2022-03-10
- Publication Date
- 2026-04-28
AI Technical Summary
Conventional methods for creating nucleic acid libraries using machine learning struggle to generate large datasets with high-quality data, often resulting in small libraries with limited sequence exploration and costly synthesis, and fail to accurately predict mutants with improved functional properties due to biased enrichment during phage infection and amplification processes.
A method involving biopanning operations to collect data from sublibraries at various stages, calculating estimated binding strengths, and using machine learning to construct a secondary library that includes sequences similar to those predicted by machine learning, thereby expanding the sequence space and reducing costs.
This approach allows for the creation of a nucleic acid library containing a larger amount of target proteins, enabling efficient improvement of functional proteins like antibodies and enzymes by accurately predicting and incorporating sequences that may not be enriched during conventional methods.
Smart Images

Figure 0007852890000018 
Figure 0007852890000019 
Figure 0007852890000020
Abstract
Description
[Technical Field]
[0001] This invention relates to a method for preparing a nucleic acid library using machine learning. More specifically, it relates to a method for preparing a nucleic acid library containing a large amount of nucleic acids encoding a target protein by using more appropriate data as machine learning data. [Background technology]
[0002] There is a widespread need to modify functional proteins such as antibodies and enzymes to improve their function. Recently, research has been progressing on how to modify protein function more efficiently using machine learning. In these studies, a mutant library of a certain size is created, the amino acid sequence and function of the mutants are experimentally measured, and this linked data is used as training data to build a machine learning model that predicts function from the sequence. Then, by using the constructed machine learning model, mutants that are predicted to have improved function are predicted.
[0003] Regarding machine learning datasets, two types are applied: datasets that directly or indirectly link amino acid sequences with functional / physical properties. In directly linked datasets, the functional / physical properties of each mutant are measured for each mutant, and these functional / physical properties are linked to the sequence of the corresponding mutant (e.g., Non-Patent Document 1). On the other hand, in indirectly linked datasets, functional / physical properties are not directly measured, and datasets are created using the number of amino acid sequence reads obtained from deep sequencing analysis as a substitute for functional / physical properties (e.g., Non-Patent Documents 2 and 3).
[0004] Direct linking amino acid sequences to functional and physical properties can result in high-quality datasets for machine learning, but creating large datasets is difficult, often remaining in size only a few dozen to a few hundred, and limiting the sequences that can be explored. On the other hand, while indirect linking results in lower data quality than directly linked datasets, it allows the use of large-sized amino acid sequence data obtained through deep sequencing analysis. Therefore, directly linked datasets are often applied when the location and number of mutant residues and the appearing amino acids are limited, while indirectly linked datasets are often applied for antibody lead molecule discovery using molecular presentation methods.
[0005] Biopanning from molecular libraries using phage display (see Figure 1A) 10 This is an effective method for obtaining antibody fragments or antibody-like molecules that exhibit target binding from a large-scale group of mutants. In recent years, a method has been reported in which sequences that are highly enriched (have a high abundance) in the library after selection are estimated to be sequences with high binding affinity based on the results of sequence analysis using next-generation sequencing (NGS), and machine learning is performed (Patent Document 1). In previous reports, machine learning has been performed using data from populations after E. coli infection (Figure 1A (v)) or after phage amplification (Figure 1A (vi)) using the phage display method (Non-Patent Document 2). However, in reality, it is often not possible to obtain a phage population containing mutants with significantly improved function (target binding ability). Furthermore, since the degree of enrichment is influenced not only by target binding ability but also by the phage infection and amplification processes of E. coli, sequences with high enrichment do not necessarily improve the desired function (Non-Patent Document 4).
[0006] When generating a certain number of top-ranked suggested sequences from machine learning predictions, the sequence diversity necessitates the synthesis of each sequence gene, which imposes cost limitations on the number of sequences that can be evaluated. Consequently, depending on the accuracy of the training data, it may not be possible to obtain sequences with the desired function. Therefore, conventional methods result in a small second library. [Prior art documents] [Patent Documents]
[0007] [Patent Document 1] US2019 / 0065677 [Non-patent literature]
[0008] [Non-Patent Document 1] Saito et al., "Machine-Learning-Guided Mutagenesis for Directed Evolution of Fluorescent Proteins" ACS Synth Biol. 2018;7(9):2014-2022 [Non-Patent Document 2] Liu et al., "Antibody complementarity determining region design using high-capacity machine learning" Bioinformatics, 2020;36(7):2126-2133 [Non-Patent Document 3] Saka et al., "Antibody design using LSTM based deep generative model from phage display library for affinity maturation" Scientific Reports, 2021;11(1):5852 [Non-Patent Document 4] Ito et al., “Application of next-generation sequencing analysis in the directed evolution for creating antibody mimic” 65th Annual Meeting of the Biophysical Society. 2021.2.25. (Boston, MA, USA) [Overview of the Initiative] [Problems that the invention aims to solve]
[0009] The object of the present invention is to provide a library containing nucleic acids encoding a target protein. In particular, it is to provide a method for obtaining a library containing the target functional molecule even from biopanning operations in which clear positive variants have not been obtained. [Means for solving the problem]
[0010] We calculated estimated binding strengths to the target using sequence data from sublibraries at various stages and evaluated their correlation with measured values for mutants. We found that by using sublibrary data from the target-binding sequence elution stage (Figure 1A (iv)), we could obtain estimated binding strengths with a high correlation to measured values even when the sequence enrichment due to selective pressure caused by target binding was smaller than the sequence enrichment caused by biased selection during phage infection and amplification of E. coli. Furthermore, we found that by combining degenerate codon design with sequence populations predicted by machine learning from indirectly linked datasets, and constructing a secondary library that includes sequences similar to those predicted by machine learning, a library containing a larger amount of the target protein can be constructed inexpensively.
[0011] In other words, the present invention relates to the following [1] to
[11] . [1] A method for preparing a nucleic acid library, 1) A step of preparing a first library consisting of mutants in which random mutations have been introduced into nucleic acid sequences encoding proteins that bind to or are to be bound to a target, using phage display. 2) A step of performing biopanning on the first library and obtaining data to be used for machine learning from the resulting sublibraries, and 3) The process includes the step of performing machine learning using the aforementioned data and obtaining a second library from the first library based on the machine learning prediction, The method wherein the data used for the machine learning includes the sequences of the mutant population included in the sub-library of the target binding sequence elution operation step, the estimated binding strength to the target, and the measured values of the binding to the target of some mutants included in the mutant population. [2] The data used for the machine learning is obtained by the following steps: i) A step of obtaining data on sequences and their occurrence frequencies for a sub-library of the target binding sequence elution operation step and a sub-library of one or more steps different from said step; ii) A step of calculating as a score indicating the estimated binding strength to the target from said occurrence frequency; iii) A step of determining, as data to be used for machine learning, said score, the measured value of the binding to the target, and the sequence data giving them, the method according to [1]. [3] The method according to [2], wherein the one or more different steps are steps selected from the group consisting of a non-specific binding sequence removal operation step, a target binding sequence selection operation step, an infection operation step with Escherichia coli, and a selected sequence amplification operation step in the same round, or steps selected from the group consisting of a non-specific binding sequence removal operation step, a target binding sequence selection operation step, a target binding sequence elution operation step, an infection operation step with Escherichia coli, and a selected sequence amplification operation step in different rounds, or both. [4] The method according to [2], wherein the score is calculated using the ratio of the occurrence frequencies between the sub-library of the target binding sequence elution operation step and the sub-library of the non-specific binding sequence removal operation step or the selected sequence amplification operation step. [5] The method according to [2], wherein the score is calculated using the ratio of the occurrence frequencies between the sub-library of the target binding sequence elution operation step and the sub-library of the non-specific binding sequence removal operation step in the same round, or using the ratio of the occurrence frequencies between the sub-library of the target binding sequence elution operation step and the sub-library of the selected sequence amplification operation step in different rounds. [6] The method according to [2], wherein the score is calculated using data of sub-libraries of 2 to 4 rounds. [7] The method according to [2], wherein the score is calculated according to any one of the following formulas 1) to 6). [Number] Here, F x,n (i) represents the abundance (number of reads of unique sequences / total number of reads in the library) of variant i in the sub-library n at the x-th round. n is n = 1: The first library n = 2: Sub-library from phages removed by non-specific binding phage removal operation n = 3: Sub-library from phages removed in the target binding sequence elution step n = 4: Sub-library from phages after the target binding sequence elution step n = 5: Sub-library from Escherichia coli after phage infection n = 6: Sub-library from phages after amplification [8] The method according to any one of [1] to [7], wherein the measured value of binding to the target is the measured value by ELISA. [9] The method according to any one of [1] to [8], wherein in step 3, by designing degenerate codons, sequences not predicted by machine learning are included in the second library.
[10] The method according to any one of [1] to [9], wherein the protein to be bound to or made to bind to the target is an antibody, antibody-like molecule, or enzyme.
[11] A method for producing an optimized protein, comprising a step of obtaining a second library according to the method described in any one of [1] to
[10] , a step of screening the second library to determine a nucleic acid sequence encoding the optimized protein, and a step of producing an optimized protein based on the nucleic acid sequence. [Advantages of the Invention]
[0012] The present invention has the following features: (1) it uses a sublibrary of the target-binding sequence elution operation stage as a phage population at an appropriate stage; (2) it creates a second library that covers a wider sequence space rather than only including the top sequences predicted by machine learning; and (3) it can be realized at low cost by again using the phage presentation method for the second library.
[0013] According to the present invention, it is possible to construct a library containing a larger amount of nucleic acids encoding the target protein. This makes it possible to efficiently improve the function of industrially useful proteins such as antibodies and enzymes. [Brief explanation of the drawing]
[0014] [Figure 1] A: An example of biopanning. B: Biopanning in Examples 1 and 2. [Figure 2] Amino acid sequence of the 2u2f protein [Figure 3] Polyclonal phage ELISA was performed using amplified phages after each round. Binding was evaluated using samples prepared from 5.0 × 10¹¹ cfu, undiluted, 5-fold diluted, and 25-fold diluted. Each sample was detected using anti-M13 phage-HRP antibody. [Figure 4] (A) Evaluation of the physical properties and function of the C6 mutant. Purification of the C6 mutant by size exclusion chromatography (arrows indicate monomer fraction). (B) Evaluation of C6 mutant binding by ELISA. (Black): Binding signal to wells immobilized with Galectin-3 via NeutrAvidin. (Gray): Binding signal to wells immobilized with NeutrAvidin alone (without Galectin-3). (C) CD spectrum measurement of the C6 mutant (gray) and wild-type 2u2f (black). [Figure 5] Percentage of reads accounted for by unique sequences in each sublibrary [Figure 6]Changes in the abundance of each unique sequence between sublibraries. The diagonal lines in the figure represent the baseline y=x. Each axis shows the logarithmic value of the abundance of the mutant in the sublibrary of interest. (A): Changes in the abundance of amplified phage from round 1 to round 2 (left), round 2 to round 3 (center), and round 3 to round 4 (right). (B): Changes in the abundance of input (amplified phage from the previous round) to output (eluted phage) in rounds 2 (left), 3 (center), and 4 (right). [Figure 7] Score calculation Fx,n: Presence rate in sublibrary n in the xth round (number of unique array reads / total number of reads in the sublibrary) [Figure 8] Changes in amino acid frequency at each residue position in rounds 2 and 3: Amino acid frequency (-1.0 - 1.0) = log2 (amino acid frequency of eluting phage (2nd) / amino acid frequency of amplification phage (1st)) [Figure 9] Amino acid frequency at each residue position in the top 10,000 sequences predicted by machine learning. [Figure 10] Clustering results of the top 10,000 sequences predicted by machine learning: (A) Number of sequences and amino acid frequency in each cluster; (B) Rank distribution of sequences included in each cluster (arrows: clusters containing the top 1,000 sequences). [Figure 11] Amino acid frequency at each residue position in the designed library (left: sequence predicted by machine learning, right: designed library) [Figure 12]Polyclonal phage ELISA using amplified phages after each round. From left to right in each graph: 5.0 x 10¹¹ cfu, 1.0 x 10¹¹ cfu, 2.0 x 10¹⁰ cfu (Target: Gal-3 (+)), 5.0 x 10¹¹ cfu, 1.0 x 10¹¹ cfu, 2.0 x 10¹⁰ cfu (Target: Gal-3 (-)) (Gal-3 (+)): Binding signal to wells immobilized with Galectin-3 via NeutrAvidin. (Gal-3 (-)): Binding signal to wells immobilized with NeutrAvidin only (without Galectin-3). [Figure 13] ELISA binding evaluation of 12 promising mutants (Gal-3 (+)): Binding signal in wells immobilized with Galectin-3 via NeutrAvidin (Gal-3 (-)): Binding signal in wells immobilized with NeutrAvidin alone (Galectin-3 absent) [Figure 14] EC50 measurement results for 1E2, 1H2, 3B5, and 4H5 mutants. [Figure 15] CD spectral measurements of wild-type 2u2f, 1H2, 1E2, 3B5, and 4H5 [Figure 16] Amino acid sequence of cAbBCII-10 and mutation site (boxed: CDR in the definition of AbM) [Figure 17] Polyclonal phage ELISA results. From left to right in each graph: 5.0 x 10¹⁰ cfu, 1.7 x 10¹⁰ cfu, 5.6 x 10⁹ cfu, 1.9 x 10⁹ cfu, 6.2 x 10⁸ cfu, 2.1 x 10⁸ cfu, 6.9 x 10⁷ cfu. (A): Binding signal to wells immobilized with Galectin-3 via NeutrAvidin. (B): Binding signal to wells immobilized with NeutrAvidin alone (without Galectin-3). [Figure 18] SEC (A) of wild-type VHH (top) and 12G mutant (bottom). Arrows: monomer, ELISA (B) (black: target molecule present, gray: target molecule absent), CD spectrum (C) results (black: wild-type VHH, gray: 12G). [Figure 19] Changes in mutant group distribution during in vitro selection (far left: initial phage; from left to right in each round: negative phage, wash phage, eluted phage, infected E. coli, amplified phage) [Figure 20] SEC (A) of wild-type VHH (top) and 738 mutant (bottom). Arrows: monomer, ELISA (B) (black: with target molecule, gray: without target molecule), CD spectrum (C) results (black: wild-type VHH, gray: 738). [Figure 21] SEC (A) and CD (B) spectra of 2G and 6C mutants (from top: WT, 738, 6C, 2G) [Figure 22] ELISA results for 2G and 6C mutants: (A) Binding signal in wells immobilized with Galectin-3 via NeutrAvidin; (B) Binding signal in wells immobilized with NeutrAvidin only (no Galectin-3); (C) Binding signal in wells immobilized with BSA (no Galectin-3); (D) ELISA results with varying concentrations of 2G and 6C mutants in wells immobilized with Galectin-3. [Modes for carrying out the invention]
[0015] This invention relates to a method for preparing nucleic acid libraries using the phage display method.
[0016] 1. Creating the initial library (the first library) First, a library is prepared using phage display, consisting of mutants in which random mutations have been introduced into the protein that "binds to or is intended to bind to the target." In this specification, this initially prepared library is referred to as the "initial library" or "first library" to distinguish it from the library enriched by machine learning. The terms "initial library" and "first library" are used interchangeably in this specification.
[0017] The "protein that binds to or is intended to bind to the target" is not particularly limited, but functional proteins that require improvement in properties, such as antibodies, antibody-like molecules, or enzymes, are preferred. Antibodies include small molecule antibodies such as VHH antibodies, Fab, F(ab') 2 This also includes antibody fragments such as scFv, diabody, and minibody. Antibody-like molecules are compounds that function by specifically binding to antigens, similar to antibodies, but are not structurally related to antibodies; they are also called antibody mimetic molecules. Examples of antibody-like molecules include afibody, afimer, afitin, alphabody, antikalin, avimer, finomer, monobody, DARPins, and nanoCLAMP.
[0018] The site where mutations are introduced ("mutation site") is selected to be a site that affects the characteristic to be optimized. "Affecting the characteristic" means that the characteristic changes or improves as a result of the change (substitution, deletion, or insertion) of amino acids in that site, especially amino acid substitution.
[0019] The selection of mutation sites, for example in the case of antibodies, involves residues including the complementarity-determining region (CDR) region, which is the antigen-recognition site, and its surrounding areas. CDRs are defined by Chothia, AbM, Kabat, Contact, etc. For antibody-like molecules of non-antibody proteins, reported mutation sites can be selected, or the mutation site can be selected based on its surface exposure or the frequency of amino acid occurrences at each residue position in homologous proteins found in nature.
[0020] Furthermore, when applying selective pressure to improve structural stability without impairing binding function, the selection of mutation sites can be carried out based on consensus engineering. "Consensus engineering" is a design based on consensus (consensus design or consensus-based engineering), an approach that enhances protein stability by modifying the protein sequence to approach a consensus sequence obtained from the alignment of numerous proteins of a specific family (e.g., Porebski and Buckle, “Consensus protein design” Protein Engineering, Design & Selection, 2016, 29(7):245-251, Steipe B., et al., J. Mol. Biol, 1994, 240(3):188-192).
[0021] Specifically, in the case of modifying enzyme function (such as improving the thermal stability of enzymes), based on the assumption that amino acid residues frequently selected in nature contribute to improved enzyme function, the frequency of amino acid occurrence at each residue position is calculated using multiple sequence alignment methods (such as ClustalW or MAFFT) for amino acid sequences of proteins belonging to the same family as the amino acid sequence of the starting protein, and the most frequently conserved amino acid residue is designated as the consensus residue. Then, each amino acid residue position of the starting protein is mutated to the consensus residue. On the other hand, with regard to antibodies, based on the assumption that the various mutations observed in germline families are due to the elimination of mutations that cause structural instability, the amino acid most frequently observed at a specific position in the alignment of immunoglobulin (Ig) variable region fragments is considered to be the most favorable amino acid for thermodynamic stability.
[0022] Consensus engineering allows for protein function modification using only amino acid sequences, without requiring knowledge of crystal structures or complex in silico calculations. However, simply substituting non-consensus amino acids with consensus residues often leads to decreased structural stability, or improved structural stability but reduced other functions (such as enzyme activity or antigen-binding activity). Therefore, selecting the appropriate residue location and the amino acid to appear in that location is crucial.
[0023] Mutation can be introduced using methods known in the field, including overlap extension PCR with primers containing degenerate codons, error-prone PCR, random primer method, inverse PCR, DNA shuffling, staggered PCR, Kunkel method, and quick-change method. Commercially available mutation introduction kits can also be used.
[0024] The library size is not particularly limited and is determined as appropriate depending on the number of mutation sites. Since there are 20 natural amino acids, for example, if there are 3 mutation sites, then 20 3 For 4 residues, it's about 8000, or 20. 4 This results in a size of approximately 160,000. The method of the present invention can be particularly suitably used when altering the function of binding to a target, especially when the mutation site consists of 7 or more residues.
[0025] 2. Acquisition of data for machine learning Next, biopanning is performed on the first library, and data to be used for machine learning is obtained from the resulting sublibraries.
[0026] "Biopanning" is a selective enrichment process for target proteins that utilizes specific binding to a target (see Figure 1A). For example, if the target protein is an antibody or antibody-like molecule, biopanning is performed to target antigen binding; if it is an enzyme, biopanning is performed to target substrate binding.
[0027] In a library population, sequences that are highly enriched (have a high abundance) through biopanning are expected to have strong binding affinity to the target. Therefore, for each mutant population (sublibrary) included in each stage of biopanning, the sequences (amino acid sequences and nucleic acid sequences) and their frequency of occurrence (number of reads for a given mutant / total number of reads in the sublibrary) are analyzed to determine the enrichment level of each sequence, which is then defined as the "estimated binding strength" to the target. This "estimated binding strength" is then scored for use in machine learning.
[0028] As mentioned above, in conventional methods, data (enrichment) of populations after selecting phages have infected E. coli (Figure 1A (v)) or after the phages have been amplified (Figure 1A (vi)) were used for machine learning. However, the frequency of appearance of populations after E. coli infection and phage amplification is biased and does not reflect the actual values. The inventors analyzed the sequences and frequency of appearance of mutant populations included in sublibraries at various stages of biopanning, scored the estimated binding strength using various calculation formulas, and compared the correlation with the actual values. As a result, they found that the data of populations after the target binding sequence elution operation (iv) showed a high correlation with the actual values. It is common in biopanning for the enrichment after target binding sequence elution to be lower than the enrichment of populations after E. coli infection and phage amplification. In this case, the enrichment of target binding is masked by the bias change that occurs during E. coli infection and phage amplification, and the enrichment due to the selection operation is not observed.
[0029] The "stages" of biopanning include, for example, the stages in each round of biopanning: the removal of nonspecific binding sequences, the selection of target binding sequences, the elution of target binding sequences, the infection of E. coli, and the amplification of selected sequences.
[0030] The data used for machine learning in this invention includes the sequences of mutant populations included in the sublibrary of the target-binding sequence elution operation step, the estimated binding strength to the target, and the measured values of binding to the target.
[0031] The data used for machine learning is acquired through the following process, for example: i) A step of obtaining data on the sequences and occurrence frequencies of mutant populations included in each step, for the biopanning target binding sequence elution step (Figure 1A (iv)) and one or more steps different from the above step. ii) A step of calculating a score indicating the estimated binding strength to the target from the frequency of occurrence (for example, normalizing it to a value between 0 and 1), iii) A step of determining the score, the measured value of binding to the target, and the sequence data that gives them to be used as data for machine learning.
[0032] The number of sequences analyzed in each sublibrary is not particularly limited, as long as meaningful training data can be provided for the artificial intelligence. (e.g., the number of sequences in the initial library used for selection) 9 A sequence is preferable, but more than 100,000 sequences are also acceptable.
[0033] In the present invention, the number of biopanning rounds is not particularly limited and is set appropriately depending on the number of target mutants and their affinity to the target. Generally, biopanning is performed in 2 or more rounds, preferably 3 or more rounds, 4 or more rounds, generally 2 to 6 rounds, and particularly 2 to 4 rounds.
[0034] The one or more different steps may be different from the target-binding sequence elution step in the same round, different from the steps in a different round, or both. Preferably, the one or more different steps are different from the target-binding sequence elution step in the same round.
[0035] Specifically, the different steps may be selected from the group consisting of a nonspecific binding sequence removal step, a target binding sequence selection step, an infection step with E. coli, and a selected sequence amplification step, all within the same round; or selected from the group consisting of a nonspecific binding sequence removal step, a target binding sequence selection step, a target binding sequence elution step, an infection step with E. coli, and a selected sequence amplification step, or both. The different steps may preferably consist of a nonspecific binding sequence removal step and / or a selected sequence amplification step, with the nonspecific binding sequence removal step being more preferred.
[0036] The score is a normalized and standardized numerical value calculated using, for example, the ratio of occurrences of sublibraries in the target-binding sequence elution stage to sublibraries in the non-specific sequence removal stage or the selected sequence amplification stage. More specifically, the score is a normalized and standardized numerical value calculated using the ratio of occurrences of sublibraries in the target-binding sequence elution stage to sublibraries in the non-specific sequence removal stage from the same round, or using the ratio of occurrences of sublibraries in the target-binding sequence elution stage to sublibraries in the selected sequence amplification stage from different rounds.
[0037] The score is calculated using data from the sublibraries of rounds 2, 3, 4, or 5, preferably rounds 2 through 4.
[0038] The score is calculated based on one of the following formulas 1) to 6).
number
[0039] The choice of function fx(i) can be determined by calculating the numerical value associated with the array using each function and then deciding based on its AUC (Area Under the Curve) value. For example, one can select an appropriate function from those that provide an AUC value of 0.5 or higher, 0.6 or higher, or 0.7 or higher.
[0040] The above scores may be further normalized as needed. For example, as in Examples 1 and 2 described later, the logarithm of the "estimated binding strength" value is used as the Enrichment Rate (ER(i)), and nScore(i) is calculated to normalize the results so that a higher ER(i) value indicates a better result.
number
[0041] In the machine learning process described later, the score value is converted to an appropriate numerical value according to the processing method used. For example, in the case of COMBO, the score is converted to a range of -1 to 0 before being used for machine learning.
[0042] The measured value of binding to the target is not particularly limited. Preferably, the measured value of binding to the target is measured by ELISA. Binding to the target can serve as an indicator of function such as affinity (binding activity), target specificity, substrate specificity, and catalytic activity. Depending on the measurement conditions, it can also serve as an indicator of structural stability, thermal stability, pH stability, aggregation, salt stability, pressure stability, reduction stability, and denaturant stability.
[0043] 3. Machine Learning In this invention, machine learning is performed using scores selected based on measured values of several mutants, along with their sequence information, as training data. Specifically, the score values obtained for a portion of the library and the corresponding mutant sequence information are used to train an artificial intelligence, which then predicts and ranks the scores of all mutants in the library. Bayesian optimization is preferred as the machine learning method.
[0044] The amino acid sequence information is input by converting it from characters to numbers (numerical vectors). Methods known in the field can be used for this purpose, such as T-scale, Z-scale, ST-scale, BLOSUM, FASGAI, MSWHIM, ProtFP, ProtFP-Feature, VHSE, Aromaphilicity, PSSM, etc. (van Westen et al., J Cheminform. 2013; 5: 41).
[0045] Bayesian optimization is a hyperparameter tuning technique, a machine learning method that finds the optimal value (maximum or minimum value) of a function whose form is unknown (a black-box function). Each candidate point is represented by a numerical vector called a descriptor. In each iteration, a machine learning model is trained using the data of candidate points evaluated so far, and the predicted value and predicted variance of the model function for the remaining candidate points are calculated using this trained model. Furthermore, a score dependent on the predicted value and predicted variance is calculated, and the candidate point with the highest score is determined as the next evaluation point, and the function evaluation is performed. The new data obtained here is added to the training data.
[0046] For "Bayesian optimization," publicly available software can be used. For example, 2DMAT (https: / / www.pasums.issp.u-tokyo.ac.jp / 2dmat / ), Common Bayesian Optimization Library (COMBO) (Ueno et al., Mater. Discov., 4, 18-21 (2016), https: / / tomoki-yamashita.github.io / CrySPY_doc / ), CrySPY (https: / / tomoki-yamashita.github.io / CrySPY_doc / ), and PHYSBO (optimization tools for PHYsics based on Bayesian Optimization) (https: / / www.pasums.issp.u-tokyo.ac.jp / physbo / ) are known, but are not limited to these. Among these, COMBO is preferred.
[0047] 4. Creating a second library Using machine learning with data from a subset of variants, artificial intelligence predicts and ranks the score values of all variants in the library. Based on the prediction results, suitable variants can be selected to create a library that is more concentrated with the target protein than the initial library. This concentrated library is referred to herein as the "second library."
[0048] If necessary, the library may be enriched more than once. That is, a second library can be created from the initial library, and then a third library can be created using the second library as the initial library. By repeating this process, enrichment is possible any number of times. The "two or more characteristics" used in the first enrichment and the characteristics used in subsequent enrichments may be the same or different. From the second enrichment onward, enrichment may be performed on two or more characteristics, or on just one characteristic.
[0049] The second library preferably includes sequences not predicted by machine learning through the design of degenerate codons. Here, the unpredicted sequences are preferably sequences similar to those predicted by machine learning.
[0050] 5. Production of optimized proteins Functional prediction through machine learning allows for the selection of mutants optimized for two or more characteristics from second, third, and subsequent libraries. The predicted mutants may then be expressed, their characteristics evaluated and confirmed, and the best one selected. For industrial applications, a smaller number of mutation sites is generally preferable. Therefore, the optimal protein (mutant) will ultimately be determined by considering both functional improvement and the number of mutations introduced. [Examples]
[0051] The present invention will be described in detail below with reference to examples, but the present invention is not limited to these examples.
[0052] [Example 1] Creation of function in antibody-like molecules Antibodies and antibody-like molecules with specific molecular recognition capabilities can be obtained through selection operations using genotype-phenotype integrated systems, such as biopanning from molecular libraries using phage display. However, it is often difficult to obtain mutants with the desired function and properties. In recent years, there have been attempts to obtain target functional molecules by using next-generation sequencers (NGS) to create indirect sequence-function linked data, treating mutants with high enrichment as high-function mutants, and performing machine learning. However, in many cases, specific mutants do not show appropriate enrichment during the selection operation, and even training data cannot be obtained. In this example, with the aim of creating antibody-like molecules, we developed a machine learning process that can obtain target functional molecules even from biopanning operations where mutants with the appropriate function and properties have not been obtained. This involved creating training data by selecting appropriate sublibraries from NGS analysis, and constructing a second library that includes sequences not predicted by machine learning from the sequence population predicted by machine learning, thereby obtaining mutants with the appropriate function and properties.
[0053] A protein in which the cysteine at position 48 of Protein Data Bank number 2u2f (SEQ ID NO: 1) was replaced with alanine was used as a scaffold protein for the antibody-like molecule, and the mutation was performed in two loop regions of the 2u2f protein (loop1: positions 11-14 (NYLN: SEQ ID NO: 2), loop2: positions 66-72 (MQLGDKK: SEQ ID NO: 3)) (Figure 2). To enable molecular recognition of 2u2f, a biopanning operation was performed targeting Galectin-3, one of the cancer markers (Figure 1B). Galectin-3 is a type of Galectin family that recognizes β-galactoside-containing glycans and is attracting attention not only as a biomarker for heart failure and cancer, but also as a novel drug target. The M13 phage presentation method was used for the selection operation. In the selection operation, an M13 phage library presenting the 2u2f mutant was first constructed. Next, a biopanning operation was performed several times, in which each cycle involved selecting and amplifying phages that exhibited target-binding activity. From the resulting phage group, several hundred types of phages were isolated, and those exhibiting target-binding activity were obtained. Furthermore, the function of promising target-binding mutants was measured even after they were cleaved from the phage, and their potential use as antibody-like molecules was evaluated.
[0054] 1. Phage library preparation and biopanning procedure PCR was performed using primers that randomize the two loop regions (loop1, 2) of 2u2f to have the same amino acid occurrence frequencies as the CDRs that appear in the human non-immune antibody library (Naive library) (Kruziki et al., “A 45-Amino-Acid Scaffold Mined from the PDB for High-Affinity Ligand Engineering,” Chemistry & Biology, 22, 946-956 (2015)). The obtained gene fragments were inserted into a pUC vector with the pIII protein of M13 phage added to the C-terminus. Escherichia coli TG-1 strain was transformed by electroporation using the obtained plasmid, and a M13 phage library of 1.0×10 9 scale was prepared.
[0055] Biopanning was performed using the prepared phage library (Figure 1B). First, the selection operation of target-binding phages was carried out. In the selection operation, negative selection was performed to remove phages that non-specifically adsorb to magnetic particles without immobilizing the target molecule using 5.0×10 11 cfu of phages (ii in Figure 1B). After that, the remaining phage solution was mixed with magnetic particles immobilized with the target Galectin-3, and phages that did not bind were washed and removed (iii in Figure 1B). Positive selection was performed to elute and recover the bound phages, obtaining a sub-library "eluted phages" (iv in Figure 1B). Next, the eluted phages were infected with Escherichia coli TG-1 strain and grown overnight on an agar medium containing ampicillin and glucose to obtain a sub-library "infected Escherichia coli" (v in Figure 1B). Furthermore, the infected Escherichia coli was cultured in a liquid medium and super-infected with helper phages to produce and amplify phages, obtaining a sub-library "amplified phages" (vi in Figure 1B). Again, using the "amplified phages", the above was repeated for a total of 4 rounds.
[0056] After the selection operation, polyclonal phage ELISA was performed using the initial library and amplified phages after each round to evaluate whether target-binding mutants were selected, and binding to Galectin-3 was assessed. The results showed an increase in signal with each round, suggesting that mutants with affinity for the target were being selected through the biopanning operation (Figure 3).
[0057] Therefore, in order to obtain mutants that exhibit target binding, monoclonal phages were prepared using 96 deep-well plates with 186 mutants each from infected E. coli after 3 and 4 rounds of processing, and binding was evaluated by phage ELISA. As a result, 52 mutant samples were obtained that showed a higher signal than the phage presenting the wild-type 2u2f and did not undergo a frameshift in the gene sequence. Among these 52 mutants, the C6 mutant (Table 1), which appeared in multiple wells, was attempted to be prepared as a protein cleaved from the phage.
[0058] [Table 1]
[0059] The C6 mutant gene, which had been inserted into a phagemid vector, was transferred to a pET vector. The resulting plasmid was used to transform E. coli BL21(DE3) strain, and after culturing, purification was performed by immobilized metal ion affinity chromatography (IMAC) and size exclusion chromatography (SEC). As a result, unlike the wild-type 2u2f without the mutation, it was expressed in various association states (Figure 4A). ELISA binding evaluation of the monomeric fraction showed that it bound not only to the target molecule Galectin-3 but also to NeutrAvidin, which was used as an anchor to immobilize Galectin-3 on the plate, indicating a lack of target specificity (Figure 4B). Furthermore, evaluation of the secondary structure of the purified protein by circular dichroism (CD) spectroscopy revealed a significant structural change compared to wild-type 2u2f, indicating that the three-dimensional structure did not maintain the native structure (Figure 4C). Therefore, while biopanning using 2u2f as a scaffold protein selected mutants with affinity for the target, it was not possible to isolate target-specific mutants.
[0060] 2. Next-Generation Sequencing Analysis (NGS) (1) DNA was extracted from the phage population or E. coli population selected in the biopanning procedure performed in (2) of 1. In addition to the "initial phage library," sublibraries such as "eluted phage," "infected E. coli," and "amplified phage" were collected as shown in (i) to (vi) in Figure 1B. The 2u2f mutant sequence fragments in each sublibrary were amplified by PCR, purified using agarose gel electrophoresis, and subjected to NGS analysis.
[0061] NGS analysis was performed using Illumia's MiSeq. The analysis employed a 2×250 pair-end sequencing method, analyzing 250 base pairs from both the 3' and 5' ends of the target DNA. After the analysis, the resulting sequence data was trimmed to remove bases with poor analysis accuracy, and then the sequences analyzed from the 3' and 5' ends were joined together (pair-end merge). The sequence data was then translated from the start codon, and sequences showing one or more substitutions, deletions, or insertions in the framework region outside the mutated loop area were removed. As a result, the number of read sequences for each sublibrary was obtained as shown in Table 2.
[0062] [Table 2]
[0063] To determine effective sublibraries for training data for machine learning, we identified rounds and operations in which mutant enrichment occurred using sequences obtained from NGS analysis. In NGS analysis, the number of sequences analyzed is called the number of reads, and unique sequences that do not overlap within the sequence set output from NGS are called unique sequences. A larger increase in the number of reads for each unique sequence compared between rounds or operations indicates stronger sequence enrichment.
[0064] To observe the rounds and operations in which sequence enrichment occurred, the proportion of each unique sequence among the sequences read by NGS was calculated and compared across sublibraries (Figure 5). As a result, enrichment of specific variants was observed from amplified phage (round 1) to eluted phage (round 2), and from amplified phage (round 2) to eluted phage (round 3). These comparisons of sublibraries represent a direct comparison of input to output in the selection operations, suggesting that the selection operation based on binding affinity worked well in rounds 2 and 3. However, significant enrichment of specific variants was also observed from eluted phage to infected E. coli in round 1, and conversely, dispersion of distribution was observed from eluted phage to infected E. coli within rounds 2, 3, and 4. This suggests that biases other than binding affinity to the target are present in the E. coli infection operation stage (v).
[0065] Next, to analyze the enrichment of each mutant resulting from the biopanning procedure, the abundance of each unique sequence was compared across sublibraries. First, the abundance of each unique sequence in each sublibrary (number of reads for the unique sequence / total number of reads in the sublibrary) was calculated, and as an analysis of enrichment between rounds, the abundance from round 1 to round 2, round 2 to round 3, and round 3 to round 4 using the infected E. coli sublibrary was compared (Figure 6A). As a result, most mutants did not show a change in abundance between rounds and were distributed around the straight line y=x, suggesting that mutant enrichment cannot be observed by comparing the output after the E. coli infection procedure across rounds. On the other hand, when comparing the abundance of each mutant from input to output in the biopanning operation in rounds 2, 3, and 4—that is, from amplified phage (round 1) to eluted phage (round 2), from amplified phage (round 2) to eluted phage (round 3), and from amplified phage (round 3) to eluted phage (round 4)—the abundance of each mutant increased from input to output, and many mutants were found to be shifted above the y=x line (Figure 6B). This suggests that by comparing rounds using the input from the previous round and the output from the current round, it will be possible to observe the enrichment of each mutant.
[0066] 3. Indirect sequencing – Creating function-linked training data The results of step 2 showed that the mutants were enriched from amplified phages to eluted phages in rounds 2 and 3. In biopanning, enrichment indicates that more molecules are bound to the antigen than other mutants. Therefore, more enriched mutants have higher binding affinity than other mutants, and the increase in their abundance from amplified phages to eluted phages can be considered as binding affinity. Furthermore, mutants that showed enrichment in different rounds can be considered to have a higher certainty of binding to the target.
[0067] Next, from the 52 samples selected from the results of the monoclonal phage ELISA in step 1, six mutants including the C6 mutant and 11 samples that were judged not to bind to the target based on the same monoclonal phage ELISA results were extracted. Using the results of these monoclonal phage ELISAs, a score value was calculated to link the sequence using the formula shown in Figure 7, and the AUC (Area Under the Curve) values were compared (Table 3). As a result, the AUC values were higher when calculated using eluted phages relative to input phages (amplified phages from the previous round), and in particular, formulas 2-2, 2-4, 2-5, and 2-6 had AUC values exceeding 0.7. In this study, formula 2-4 was used among those with an AUC value exceeding 0.7.
[0068] [Table 3]
[0069] Based on the results of 2 and 3, the enrichment rate (ER(i)) of mutant i was defined.
number
[0070] 4. Creating a prediction system using machine learning Using the above data as training data, machine learning was performed to predict the functional evaluation values of unknown mutants from their amino acid sequences. The prediction system was constructed using COMBO, a high-speed Bayesian optimization software (Ueno et al., 2016, etc.). The mutant sequence data was represented using an appropriate index or combination thereof, representing each residue as a 1 to 10-dimensional vector, as previously reported (van Westen et al., 2013).
[0071] Next, we defined the set of sequences (prediction space) to be used for predicting functional values. The size of the prediction space is determined by Ln (n=1~11), where Ln is the number of different amino acids appearing at residue position n. Prediction space = L1 × L2 × ...L11 This can be expressed as follows. The 2u2f mutant library used in this study has 11 mutation sites, so the sequence space when all 20 amino acids appear at all sites is 2.0 × 10⁻¹⁶. 14 In this study, we limited the number of amino acids appearing at each residue position, and the scale was 10 9 The prediction space was designed to achieve a certain degree.
[0072] To limit the amino acids appearing in the predicted space, the degree of amino acid enrichment at each residue position was used. Amino acids at each residue position whose frequency of appearance increased due to the biopanning operation in 1. are those that may be involved in binding at that position, while amino acids whose frequency of appearance decreased due to the selection operation are those that are not involved in binding or may inhibit binding. Therefore, the rate of change in amino acid appearance frequency from amplified phage (round 1) to eluted phage (round 2), and from amplified phage (round 2) to eluted phage (round 3), where enrichment of mutants with binding affinity was suggested, was calculated (Figure 8). Here, the frequency of appearance of a certain amino acid k at residue position m when focusing on sublibrary n is:
number
[0073] [Table 4]
[0074] 5. Narrowing down promising mutants using a prediction system The constructed prediction system was used to calculate the predicted values for all mutants contained within the sequence space where a specific amino acid (Table 4) appears at 11 residue positions (11-14, 66-72 in Figure 2). The top 10,000 predicted sequences were selected as promising mutants (Figure 9).
[0075] 6. Designing the Second Library To create a second library containing the top 10,000 sequences predicted by machine learning (5.) and perform biopanning using phage display, similar sequences were grouped together from the top 10,000 sequences predicted by machine learning. For grouping, the Basic Local Alignment Search Tool (BLAST) (Crooks et al., WebLogo: A sequence logo generator, Genome Research, 14, 1188-1190 (2004)) was used to perform pairwise alignment on all of the top 10,000 sequences, and sequences with an e-value of 0.1 or less were considered similar. At this time, the alignment was performed with settings that did not include sequence shifts (gaps). As a result, the top 10,000 sequences from machine learning were classified into nine main clusters, and each cluster was named Cluster 1 to 9 in order of the number of sequences contained within the cluster (Figure 10A). Furthermore, examining the rank distribution of amino acid sequences within each cluster, we found that among Clusters 1-9, Clusters 1, 3, 4, and 6 contained sequences that ranked in the top 1,000 of the predicted sequences, indicating that overall, a high proportion of variants with high machine learning prediction ranks were observed (Figure 10B).
[0076] Therefore, the design of phage library gene sets containing sequences from Clusters 1, 3, 4, and 6, which include mutants with high machine learning prediction ranks, was performed using degenerate codons. In each cluster, the frequency of amino acid occurrence at each residue position was calculated from the sequence population within the cluster, and degenerate codons were designed to create a group of 2u2f mutant genes in which residues with an occurrence frequency of 5% or more appear. Specifically, after determining which amino acids to be expressed, codon design was performed from the following perspectives. (i) Amino acids proposed by the prediction system (with an occurrence frequency of 5% or more) must be included. (ii) Avoid producing unnecessary amino acids as much as possible. (iii) Do not allow the TAA·TGA stop codons to appear, but do not allow the TAG stop codon to appear as much as possible.
[0077] As a result, we were able to design codons for each cluster that included the amino acids that appear at each residue position while eliminating as many extraneous amino acids as possible. However, there were also sequences that were not included in the machine learning predictions, and the proportion of target mutants in the designed library was 0.82%, 0.33%, 1.18%, and 0.18% for Clusters 1, 3, 4, and 6, respectively (Figure 11, Table 5). Although the proportion of sequences predicted by machine learning was small, we thought that it might be possible to obtain mutants with further optimized predicted sequences by using a library containing sequences similar to the predicted sequences, and therefore we prepared an M13 phage library based on this codon design.
[0078] [Table 5]
[0079] 7. Phage library preparation and second biopanning A second library was constructed using primers designed with degenerate codons, and an M13 phage library presenting the 2u2f mutant was created. 8The libraries were prepared on a large scale. This scale is more than 100 times larger than the sequence space of each library, so it can be said that a phage library containing not only the cluster sequences predicted by machine learning but also all the variants included in each library was prepared.
[0080] Next, biopanning was performed using the second phage library prepared, and polyclonal phage ELISA was performed using the amplified phage groups from each round. All clusters showed an increasing signal with each subsequent round (Figure 12). At this time, Cluster 6 enriched mutants that also bound to NeutrAvidin in the wells immobilized with NeutrAvidin alone, while the polyclonal phages from Clusters 1, 3, and 4 showed specific binding.
[0081] Therefore, 88 clones were isolated from each library's mutant group after 3 rounds, and a monoclonal phage ELISA was performed to screen for mutants that specifically bind to the target Galectin-3. A total of 63 mutants were obtained that showed specific binding to Galectin-3: 20 from Cluster 1, 14 from Cluster 3, 20 from Cluster 4, and 9 from Cluster 6. Here, the cluster number from which the mutant originated was placed at the beginning, followed by the well number of the 96-well plate from which it was obtained. For example, a mutant obtained from Cluster 1 and cultured in well E2 would be named "1E2". To narrow down candidate molecules from these 63 obtained mutants, the selected mutant genes were first transferred from phagemide vectors to pET22b vectors for protein expression. Then, mutants expressed in small-scale cultures using 96 deep-well plates were evaluated for monomeric expression using Blue Native PAGE (BN-PAGE), narrowing the selection down to 12 types. Further cultures were performed on a 500 mL scale, and 11 types of mutants were obtained as monomers by purification of the soluble fraction using IMAC and SEC. When these obtained mutants were evaluated using ELISA to see whether they bound to Galectin-3, mutants 1E2, 1H2, 3B5, and 4H5 showed dominant binding to Galectin-3 (Figure 13).
[0082] Next, to quantify the affinity of the four mutants that showed specific binding to the target Galectin-3, eight series were prepared by diluting them 2-fold from 1.5 μM, and the EC50 values were calculated from binding measurements by ELISA. As a result, the EC50 values of the 1E2, 1H2, 3B5, and 4H5 mutants were calculated. 50The chromosome values were 92.5 nM, 79.9 nM, 277.4 nM, and 200.8 nM, respectively (Figure 14). Furthermore, CD spectroscopy was performed to evaluate whether these mutants formed secondary structures. As a result, it was found that the C6 mutant obtained from wet experiments alone adopted a random coil structure (Figure 4C), while the 1H2 and 4H5 mutants obtained in this study, in particular, adopted a secondary structure similar to the wild-type 2u2f (Figure 15). Thus, from the second library designed using the results from the prediction system, mutants that could not be found in wet experiments alone, which maintain their three-dimensional structure while exhibiting specificity to the target, were obtained.
[0083] The 1E2, 1H2, 3B5, and 4H5 mutants were not included in the top 10,000 predicted sequences in machine learning. Four residues in the 1E2 mutant, three residues in the 1H2 mutant, two residues in the 3B5 mutant, and two residues in the 4H5 mutant were amino acids that did not appear in the prediction space in machine learning (Table 6; the amino acid sequences are shown in SEQ ID NOs. 6-13). Furthermore, two residues in the 3B5 mutant and one residue in the 4H5 mutant were included in the prediction space in machine learning, but did not appear in Cluster 3 and Cluster 4 after clustering. Based on these results, by including sequences similar to the top predicted sequences in machine learning in the second library, mutants with the desired function and properties could be obtained.
[0084] [Table 6]
[0085] [Example 2] Functional improvement of weakly bound molecules identified by biopanning method In genotype-phenotype integrated systems such as biopanning from molecular libraries using phage display, it is not always possible to obtain mutants with the appropriate desired function and properties. In recent years, there have been attempts to obtain molecules with the desired function by creating indirect sequence-function linked data using next-generation sequencers (NGS) to treat mutants with high enrichment as high-function mutants and performing machine learning. However, in many cases, specific mutants do not show appropriate enrichment during the selection process, and even training data cannot be obtained. In this example, as a way to create function for camel heavy chain antibody heavy chain variable region fragments (VHH), we developed a machine learning process that improves the function and properties of mutants with insufficient function and properties obtained by biopanning, using NGS analysis results as training data and incorporating machine learning.
[0086] 1. Phage library preparation and biopanning procedure Using the anti-β-lactamase camel antibody fragment cAbBCII-10 VHH (PDB ID: 3DWT (SEQ ID NO: 14)) as a scaffold protein, three CDRs defined by AbM were selected as mutation sites (39 residues) (Figure 16). PCR was performed using primers that randomized the amino acid occurrence frequency to match that of the CDRs appearing in the human non-immune antibody library (Naive library), as in Example 1. The obtained gene fragment was inserted into a pUC vector with the M13 phage pIII protein added to the C-terminus. Using the obtained plasmid, E. coli TG-1 strain was transformed by electroporation, and this transformant was used to produce 8.6 × 10⁶ molecules. 7 A large-scale M13 phage library was constructed.
[0087] Using the prepared phage library, the same biopanning procedure as in Example 1 was performed to obtain sublibraries (indicated by (i) to (vi) in Figure 1B) of "eluted phages," "infected E. coli," and "amplified phages" from rounds 1 to 4.
[0088] After the selection operation, polyclonal phage ELISA was performed using the initial library and amplified phages after each round to evaluate whether target-binding mutants were selected, and binding to Galectin-3 was assessed. As a result, the signal increased with each round (Figure 17), suggesting that mutants with affinity for the target were being selected through the biopanning operation.
[0089] Therefore, to obtain mutants exhibiting target binding, 180 clones were isolated from infected E. coli after 4 rounds of incubation. Monoclonal phages were prepared using 96-deep-well plates, and binding was evaluated by phage ELISA. As a result, five mutants were obtained that showed a signal more than three times higher than the phage displaying wild-type VHH (7B, 11E, 11D, 4H, 12G). Therefore, we attempted to prepare monomeric proteins cleaved from these five mutants.
[0090] The mutant genes inserted into the phagemide vectors of five mutants that showed binding positivity were transferred to a pRA5 vector. The resulting plasmids were used to transform E. coli BL21(DE3) strain, and after culturing, purification was performed using IMAC and SEC. For comparison, we also attempted to produce monomeric proteins from two mutants (6G, 6F) that showed negative binding in the Galectin-3 binding ELISA. As a result, only the 12G mutant was partially eluted by SEC at a monomer position similar to that of wild-type VHH, but the yield was less than 1 / 20 of that of wild-type (Figure 18A). The 12G mutant prepared as a monomer showed specific binding affinity to target Galectin-3 in ELISA (Figure 18B), but when the secondary structure of the purified protein was evaluated by CD spectroscopy, it was found that the structure was significantly different compared to wild-type VHH, and the three-dimensional structure did not maintain the native structure (Figure 18C).
[0091] 2. Next-Generation Sequencing Analysis (NGS) Similar to Example 1, NGS analysis was performed on the sublibraries (i) to (vi) in Figure 1B using Illumia's MiSeq, and the sequences in Table 10 were obtained for each sublibrary. Then, similar to Example 1, in order to observe the rounds and operations in which sequence enrichment occurred, the proportion of each unique sequence among the sequences read by NGS was calculated and compared between the sublibraries (Figure 19). As a result, similar to Example 1, it was found that the distribution change occurred more significantly during the E. coli infection and amplification operation than during the selection operation. This indicates that the influence of the distribution change due to the amplification operation needs to be removed when linking functional information. As a result, it was found that the distribution change from eluted phages to infected E. coli was more significant than the distribution change due to the selection operation, indicating that the influence of the distribution change due to the amplification operation needs to be removed when linking functional information.
[0092] [Table 7]
[0093] 3. Indirect sequencing – Creating function-linked training data Next, to analyze the degree of enrichment of each mutant caused by the biopanning procedure, the results of monoclonal phage ELISA for the five binding-positive mutants and two binding-negative mutants obtained above were used to calculate a score value linked to the sequence using the formula shown in Figure 7, and the AUC values were compared (Table 8).
[0094] [Table 8]
[0095] As a result, the AUC values were higher when calculated using eluted phages compared to phages removed by negative selection, and in particular, equations 1-3 and 1-6 showed AUC values exceeding 0.7. In this study, we used equation 1-3 among those with an AUC value exceeding 0.7.
[0096] The formula obtained by dividing the number of "eluted phages" from four rounds by the number of "negatively selected phages" was found to be the most effective at distinguishing between binding-positive and binding-negative mutants.
[0097] Based on the above results, the enrichment rate (ER(i)) of mutant i was defined.
number
[0098] 4. Searching for novel binding-positive mutants from mutant groups using clustering analysis. 4 th Using the NGS data of the post-round mutant group, we searched for mutants with amino acid sequences similar to the 12G CDR using the homology sequence search program BLAST. Clustering analysis with a threshold of 10 or less for the expected E-value during BLAST search revealed 38 12G-like mutants.
[0099] Next, we limited the protein preparation to mutants with a phage abundance ratio of 1 or higher in the "eluted phage" sublibraries of the 3rd and 4th rounds, out of the 38 12G-like mutants. As a result, one similar mutant (738, Table 12) was prepared as a monomeric protein without aggregate formation (Figure 20A), and binding evaluation by ELISA showed positive binding to the target molecule (Figure 20B). Furthermore, secondary structure evaluation by CD spectroscopy revealed that it retained a secondary structure similar to wild-type VHH (Figure 20C).
[0100] [Table 9]
[0101] 5. Creating a prediction system using machine learning Using the training data prepared in step 3, machine learning was used to predict the residue positions that contribute to improved binding affinity in the binding-positive mutant 738. The prediction system was prepared using COMBO, as in Example 1, and the mutant sequence data was also represented using an appropriate index, either a 1- to 10-dimensional vector per residue or a combination thereof, as in Example 1.
[0102] Next, the sequence set to be used to predict functional values (prediction space) is a sequence space in which mutants are introduced into the 19 amino acid sequences located at CDR3 of the 738 mutants, with a maximum of 4 residue mutations introduced into each mutant. 19 C3×20 4 = 6.2 × 10 8 We designed a prediction space for ).
[0103] 6. Designing a second library using a prediction system Using the constructed prediction system, we calculated the predicted values for all mutants contained within the sequence space represented by 19 residues in CDR3. Then, we determined that the four residue positions in CDR3 with the most mutations among the top 1,000 predicted sequences (35, 37, 38, 39) would be the mutation introduction sites for the second library (Table 13).
[0104] [Table 10]
[0105] To design a second library of genes containing amino acids that appear in at least 10 of the top 10,000 sequences predicted by the prediction system, we used degenerate codons to create a second library of genes that would appear at the four determined mutant residue positions. This design was successful, with only residue position 39 containing an excluded amino acid (R). Using primers with degenerate codons that represent a sequence space size of 648 (9×4×2×9), we performed PCR using 738 mutants as a template to construct the second library. The gene fragments from the constructed second library were inserted into a pRA5 vector, and 180 clones of E. coli BL21(DE3) transformed with the constructed plasmid were cultured on a small scale in 96-deep-well plates. The expressed mutants were evaluated for binding to Galectin-3 using ELISA. Two mutants (2G, 6C) that specifically bound to Galectin-3 were selected and cultured in 500 mL scale. Purification by IMAC and SEC revealed that both mutants could be prepared as monomers (Figure 21A), and that they formed secondary structures similar to the wild type in CD spectra (Figure 21B). Furthermore, ELISA evaluation showed that both 6C mutants bound to target Galectin-3 approximately 20 times more strongly than the 738 mutant (Figure 22). [Industrial applicability]
[0106] According to the present invention, optimized proteins can be efficiently obtained for proteins with high industrial value, such as antibodies and enzymes. This makes it easy to modify these proteins to improve their function.
[0107] All publications, patents, and patent applications cited herein are incorporated herein by reference in their entirety. [Sequence Listing Free Text]
[0108] Sequence ID 4: synthetic peptide C6 Loop 1 Sequence ID 5: synthetic peptide C6 Loop 2 Sequence ID 6: synthetic peptide 1E2 Loop 1 Sequence ID 7: synthetic peptide 1E2 Loop 2 SEQ ID NO: 8: synthetic peptide 1H2 Loop 1 Sequence ID 9: synthetic peptide 1H2 Loop 2 Sequence ID 10: synthetic peptide 3B5 Loop 1 Sequence ID 11: synthetic peptide 3B5 Loop 2 Sequence ID 12: synthetic peptide 4H5 Loop 1 SEQ ID NO: 13: synthetic peptide 4H5 Loop 2 Sequence ID 14: cAbBCII-10 VHH Sequence ID 15: CDR3 of 12G mutant Sequence ID 16: CDR3 of 738 mutant
Claims
1. A method for preparing a nucleic acid library, 1) A step of preparing a first library consisting of mutants in which random mutations have been introduced into nucleic acid sequences encoding proteins that bind to or are to be bound to a target, using phage display. 2) A step of performing biopanning on the first library and obtaining data to be used for machine learning from the resulting sublibraries, and 3) The process includes the step of performing machine learning using the data and obtaining a second library from the first library based on the machine learning prediction, The method wherein the data used for machine learning includes sequences of mutant populations included in a sublibrary of the target-binding sequence elution operation step, estimated binding strength to the target, and measured values of the binding of some mutants included in the mutant population to the target.
2. The data used for machine learning goes through the following process: i) A step of acquiring sequence and frequency data for a sublibrary of the target binding sequence elution step and one or more sublibraries of steps different from the said step. ii) A step of calculating a score indicating the estimated binding strength to the target from the frequency of occurrence, iii) The method according to claim 1, obtained by the step of determining the score, the measured value of binding to the target, and the sequence data that gives them as data to be used for machine learning.
3. The method according to claim 2, wherein one or more different steps are selected from the group consisting of a nonspecific binding sequence removal step, a target binding sequence selection step, an infection step with Escherichia coli, and a selected sequence amplification step in the same round, or are selected from the group consisting of a nonspecific binding sequence removal step, a target binding sequence selection step, a target binding sequence elution step, an infection step with Escherichia coli, and a selected sequence amplification step in different rounds, or both.
4. The method according to claim 2, wherein the score is calculated using the ratio of the frequency of occurrence of the sublibrary from the target binding sequence elution step to the sublibrary from the nonspecific binding sequence removal step or the selected sequence amplification step.
5. The method according to claim 2, wherein the score is calculated using the ratio of occurrence frequencies of sublibraries from the target-binding sequence elution stage to sublibraries from the non-specific binding sequence removal stage in the same round, or using the ratio of occurrence frequencies of sublibraries from the target-binding sequence elution stage to sublibraries from the selected sequence amplification stage in different rounds.
6. The method according to claim 2, wherein the score is calculated using data from sublibraries of rounds 2 to 4.
7. The method according to claim 2, wherein the score is calculated according to one of the following formulas 1) to 6). [Math 1] Here, F x,n (i) represents the prevalence of variant i in sublibrary n in the xth round (number of unique sequence reads / total number of reads in the sublibrary). n is n=1: The first library n=2: Sublibrary from phages removed by nonspecific binding phage removal operation n=3: Sublibrary from phages removed at the target binding sequence elution stage n=4: Sublibraries from phages after the target-binding sequence elution step. n=5: Sublime library from E. coli after phage infection n=6: Sublibraries from amplified phages
8. The method according to any one of claims 1 to 7, wherein the measured value of binding to the target is the value measured by ELISA.
9. The method according to any one of claims 1 to 8, wherein in step 3, the design of degenerate codons is used to include sequences not predicted by machine learning in the second library.
10. The method according to any one of claims 1 to 9, wherein the protein to be bound to or to be bound to the target is an antibody, an antibody-like molecule, or an enzyme.
11. A method for producing an optimized protein, A step of obtaining a second library according to the method described in any one of claims 1 to 10, The steps include screening the second library and determining the nucleic acid sequence encoding the optimized protein, and The method comprising the step of producing an optimized protein based on the nucleic acid sequence.
Citation Information
Patent Citations
Machine learning based antibody design
US20190065677A1