Information processing system, information processing method, program, and method for producing antigen-binding molecule or protein
The information processing system addresses the challenge of predicting antigen-binding molecule sequences by using sequence learning and generation units to mutate and estimate characteristic evaluations, improving sequence prediction accuracy and efficiency.
Patent Information
- Application Number
- JP2024107694
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-07
- Filing Date
- 2024-07-03
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-06-08
AI Technical Summary
There is a demand in the pharmaceutical field for predicting the sequence of desired antigen-binding molecules using machine learning models, but existing technologies are inadequate in providing accurate information on these molecules.
An information processing system and method that utilizes sequence learning and generation units to generate trained models capable of mutating building blocks in antigen-binding molecule sequences, allowing for the estimation of characteristic evaluations of virtual sequences.
Enables the provision of information on desired antigen-binding molecules or proteins, enhancing the prediction accuracy and efficiency in generating desired sequences.
Smart Images

Figure 0007757472000003 
Figure 0007757472000004 
Figure 0007757472000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system, an information processing method, a program, and a method for producing an antigen-binding molecule or protein. This application claims priority from Japanese Patent Application No. 2019-106814, filed on June 7, 2019, the contents of which are incorporated herein by reference. [Background technology]
[0002] In recent years, machine learning information processing technology has been utilized in the pharmaceutical field. For example, in the technology described in Patent Document 1, a machine learning engine is trained using affinity information representing various antibodies and the affinity of the antibodies for antigens. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2018 / 132752 Summary of the Invention [Problem to be solved by the invention]
[0004] Meanwhile, in machine learning in the pharmaceutical field, there is a demand for predicting information such as the sequence of a desired antigen-binding molecule using a trained model and providing that information.
[0005] The present invention has been made to solve the above-mentioned problems, and an object of the present invention is to provide an information processing system, an information processing method, a program, and a method for producing an antigen-binding molecule or protein, which are capable of providing information on a desired antigen-binding molecule or protein. [Means for solving the problem]
[0006] The present invention has been made to solve the above-mentioned problems, and one aspect of the present invention is an information processing system comprising: a sequence learning unit that performs machine learning based on sequence information representing sequences that include some or all of the sequences of a plurality of antigen-binding molecules, to generate a trained model that has learned features of the sequences; and a sequence generation unit that generates, based on the trained model, virtual sequence information in which at least one of the building blocks constituting the sequence represented by the sequence information has been mutated.
[0007] Another aspect of the present invention is an information processing system that includes a sequence learning unit that performs machine learning based on sequence information representing a sequence that includes part or all of the sequences of each of a plurality of proteins to generate a trained model that has learned the characteristics of the sequence, and a sequence generation unit that generates virtual sequence information based on the trained model by mutating at least one of the constituent units that make up the sequence represented by the sequence information.
[0008] Another aspect of the present invention is an information processing system comprising: a learning unit that generates a second trained model by performing machine learning based on sequence information representing sequences including some or all of the sequences of multiple antigen-binding molecules or proteins and results of characteristic evaluation of the antigen-binding molecules or proteins represented by the sequences; and an estimation unit that inputs virtual sequence information generated based on the first trained model, which is the trained model, into the second trained model and performs arithmetic processing of the second trained model, thereby estimating a predicted value of characteristic evaluation of the antigen-binding molecule or protein having a sequence represented by the input virtual sequence information.
[0009] Another aspect of the present invention is an information processing method for an information processing system, the information processing method comprising: a sequence learning process for performing machine learning based on sequence information representing sequences including some or all of the sequences of a plurality of antigen-binding molecules or proteins to generate a trained model that has learned features of the sequence information; and a sequence generation process for generating virtual sequence information, based on the trained model, by mutating at least one of the building blocks that constitute the sequence represented by the sequence information.
[0010] Another aspect of the present invention is a program for causing a computer in an information processing system to execute: a sequence learning procedure for generating a trained model by performing machine learning based on sequence information representing sequences including some or all of the sequences of multiple antigen-binding molecules or proteins; and a sequence generation procedure for generating, based on the trained model, virtual sequence information in which at least one of the constituent units constituting the sequence represented by the sequence information has been mutated.
[0011] Another aspect of the present invention is a method for producing an antigen-binding molecule or protein represented by a virtual sequence for which a predicted value of characteristic evaluation has been estimated, using the information processing system. [Effects of the Invention]
[0012] According to the present invention, information on desired antigen-binding molecules or proteins can be provided. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a schematic diagram illustrating an example of an information processing system according to a first embodiment. [Figure 2] FIG. 10 is an explanatory diagram illustrating an example of a series of panning operations according to the present embodiment. [Figure 3] FIG. 2 is a block diagram illustrating an example of a user terminal according to the present embodiment. [Figure 4] FIG. 10 is a diagram showing an example of a screen flow according to the present embodiment. [Figure 5] FIG. 2 is a block diagram showing an example of a next-generation sequencer according to the present embodiment. [Figure 6] FIG. 2 is a block diagram illustrating an example of a server according to the present embodiment. [Figure 7] FIG. 10 is a diagram showing an example of experiment information according to the present embodiment. [Figure 8] FIG. 10 is a diagram showing an example of experiment attribute information according to the present embodiment. [Figure 9] FIG. 2 is a diagram illustrating an example of a data set according to the present embodiment. [Figure 10] FIG. 10 is a diagram illustrating another example of a data set according to the present embodiment. [Figure 11] FIG. 10 is a diagram illustrating an example of a training data set according to the present embodiment. [Figure 12] FIG. 10 is a diagram showing an example of prediction target sequence information according to the present embodiment. [Figure 13] FIG. 10 is a diagram showing an example of characteristic evaluation information according to the present embodiment. [Figure 14] FIG. 10 is an explanatory diagram illustrating an example of a learning process according to the embodiment. [Figure 15] FIG. 2 is a conceptual diagram illustrating the structure of an LSTM according to the present embodiment. [Figure 16] 10 is a flowchart showing an example of the operation of a virtual array generation unit according to the present embodiment. [Figure 17] 10 is a flowchart illustrating an example of an operation of a server according to the present embodiment. [Figure 18] 10 is a flowchart showing another example of the operation of the server according to the embodiment. [Figure 19] FIG. 10 is a block diagram illustrating an example of a server according to the second embodiment. [Figure 20] FIG. 2 is a diagram illustrating an example of a data set according to the present embodiment. [Figure 21] FIG. 10 is a diagram illustrating another example of a data set according to the present embodiment. [Figure 22] FIG. 11 is a block diagram illustrating an example of a server according to the third embodiment. [Figure 23] FIG. 1 is a diagram illustrating an overview of a learning model according to the present embodiment. [Figure 24] FIG. 11 is a block diagram showing an example of a user terminal according to the fourth embodiment. [Figure 25] FIG. 2 is a block diagram illustrating an example of a server according to the present embodiment. [Figure 26] FIG. 4 is a diagram showing an example of array information according to the present embodiment. [Figure 27] FIG. 4 is a diagram illustrating an example of characteristic information according to the embodiment. [Figure 28] FIG. 2 is a diagram showing an example of a sensorgram according to the present embodiment. [Figure 29] FIG. 10 is a diagram showing an example of prediction target sequence information according to the present embodiment. [Figure 30] FIG. 10 is a diagram showing an example of evaluation result information according to the embodiment. [Figure 31] FIG. 10 is an explanatory diagram illustrating an example of a learning process according to the embodiment. [Figure 32] FIG. 10 is an explanatory diagram illustrating an example of an evaluation process according to the present embodiment. [Figure 33] 10 is a flowchart illustrating an example of an operation of a server according to the present embodiment. [Figure 34] 10 is a flowchart showing another example of the operation of the server according to the embodiment. [Figure 35] FIG. 13 is a block diagram showing an example of a server according to the fifth embodiment. [Figure 36] 10 is a flowchart illustrating an example of an operation of the information processing system according to the embodiment. [Figure 37] FIG. 2 is a block diagram illustrating an example of a hardware configuration of a server according to the embodiment. [Figure 38] FIG. 10 is a diagram showing an example of structural analysis information according to the fourth and fifth embodiments. [Figure 39] FIG. 10 is a diagram showing the relationship between the arrangement and the characteristics according to the embodiment. [Figure 40] FIG. 10 is a diagram showing the prediction accuracy of sequence characteristics according to an embodiment. [Figure 41] FIG. 10 is a diagram showing the similarity between a training sequence and a virtual sequence according to an embodiment. [Figure 42] FIG. 10 is another diagram showing the similarity between the training sequence and the virtual sequence according to the embodiment. [Figure 43A] FIG. 10 is a diagram showing the correlation between predicted and measured values of characteristics of an array according to an example. [Figure 43B] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to an embodiment. [Figure 43C] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to an embodiment. [Figure 43D] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to an embodiment. [Figure 43E] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to an embodiment. [Figure 43F] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to an embodiment. [Figure 43G] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to an embodiment. [Figure 43H] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to an embodiment. [Figure 43I] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to an embodiment. [Figure 44A] FIG. 10 is a diagram showing the correlation between predicted values and measured values of characteristics of an array according to another embodiment. [Figure 44B] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to another embodiment. [Figure 44C] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to another embodiment. [Figure 44D] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to another embodiment. [Figure 44E] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to another embodiment. [Figure 44F] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to another embodiment. [Figure 44G] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to another embodiment. [Figure 44H] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to another embodiment. [Figure 44I] FIG. 10 is a diagram showing the correlation between predicted values and actual measured values of another characteristic of an array according to another embodiment. [Figure 45] FIG. 10 is a diagram for explaining the improvement of the characteristics of the arrangement according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] <Terms, etc.> The following definitions and detailed description are provided to facilitate understanding of the disclosure described herein.
[0015] ·amino acid As used herein, amino acids are represented by one-letter or three-letter codes, or both, such as Ala / A, Leu / L, Arg / R, Lys / K, Asn / N, Met / M, Asp / D, Phe / F, Cys / C, Pro / P, Gln / Q, Ser / S, Glu / E, Thr / T, Gly / G, Trp / W, His / H, Tyr / Y, Ile / I, and Val / V.
[0016] Amino acid modification To modify amino acids in the amino acid sequence of an antigen-binding molecule, known methods such as site-directed mutagenesis (Kunkel et al. (Proc. Natl. Acad. Sci. USA (1985) 82, 488-492)) and overlap extension PCR can be appropriately used. Furthermore, several known methods can also be used to modify amino acids by substituting amino acids other than natural amino acids (Annu. Rev. Biophys. Biomol. Struct. (2006) 35, 225-249, Proc. Natl. Acad. Sci. USA (2003) 100 (11), 6353-6357). For example, a cell-free translation system (Clover Direct (Protein Express)) containing a tRNA in which a non-natural amino acid is bound to an amber suppressor tRNA complementary to the UAG codon (amber codon), a type of stop codon, can also be suitably used.
[0017] ·antigen As used herein, the structure of an "antigen" is not limited to a specific structure, as long as it contains an epitope to which an antigen-binding domain binds. In some embodiments, the antigen is a peptide, polypeptide, or protein of four or more amino acids. Examples of the above antigens include membrane molecules that are expressed on the cell membrane, and soluble molecules that are secreted extracellularly from cells.
[0018] Antigen-binding domain As used herein, the term "antigen-binding domain" may refer to any domain of any structure as long as it binds to a target antigen. Examples of such domains include the variable regions of the heavy and light chains of an antibody, a module called an A domain of approximately 35 amino acids contained in Avimer, a cell membrane protein present in vivo (International Publication Nos. WO 2004 / 044011 and WO 2005 / 040229), Adnectin (International Publication No. WO 2002 / 032925) containing the Fn3 domain, which is a domain that binds to proteins in fibronectin, a glycoprotein expressed on the cell membrane, Affibody (International Publication No. WO 1995 / 001937) using an IgG-binding domain consisting of a 58-amino acid, three-helix bundle of Protein A as a scaffold, and DARPins (Designed Ankyrin Repeats), which are regions exposed on the molecular surface of ankyrin repeats (AR) with a structure in which a 33-amino acid turn, two antiparallel helices, and a loop subunit are repeatedly stacked. Suitable examples include lipocalin molecules such as lipocalin proteins (International Publication No. WO 2002 / 020565), anticalin, which is a four-loop region supporting one side of a barrel structure in which eight highly conserved antiparallel strands twist toward the center, found in lipocalin molecules such as neutrophil gelatinase-associated lipocalin (NGAL) (International Publication No. WO 2003 / 029462), and a concave region of a parallel sheet structure within a horseshoe-shaped structure in which leucine-rich repeat (LRR) modules of the variable lymphocyte receptor (VLR), which does not have an immunoglobulin structure and is part of the adaptive immune system of jawless fish such as lampreys and hagfish, are repeatedly stacked (International Publication No. WO 2008 / 016854). Preferred examples of the antigen-binding domain of the present disclosure include antigen-binding domains comprising the variable regions of the heavy and light chains of an antibody. Preferred examples of such antigen-binding domains include "scFv (single chain Fv)," "single chain antibody," "Fv," "scFv2 (single chain Fv 2)," "Fab," and "F(ab')2."
[0019] ·Antigen-binding molecules In the present disclosure, the term "antigen-binding molecule containing an antigen-binding domain" is used in the broadest sense and specifically includes various molecular types as long as it contains an antigen-binding domain. An antigen-binding molecule may be a molecule consisting of only an antigen-binding domain, or a molecule containing an antigen-binding domain and other domains. For example, when an antigen-binding molecule is a molecule in which an antigen-binding domain is bound to an Fc region, examples include complete antibodies and antibody fragments. Antibodies may include single monoclonal antibodies (including agonist and antagonist antibodies), human antibodies, humanized antibodies, chimeric antibodies, etc. Scaffold molecules, in which existing stable α / β barrel protein structures or other three-dimensional structures are used as scaffolds (foundations), and only partial structures of these structures are compiled into libraries for constructing antigen-binding domains, are also included in the antigen-binding molecules of the present disclosure.
[0020] ·antibody As used herein, an antibody refers to an antigen-binding molecule containing a full-length immunoglobulin or a partial immunoglobulin sequence, whether naturally occurring or partially or completely synthetically produced. Antibodies can be isolated from natural sources such as plasma or serum, or from the culture supernatant of antibody-producing hybridoma cells. Alternatively, they can be partially or completely synthesized using techniques such as genetic recombination. Examples of antibodies include immunoglobulin isotypes and their isotypic subclasses. Nine known classes (isotypes) of human immunoglobulins are IgG1, IgG2, IgG3, IgG4, IgA1, IgA2, IgD, IgE, and IgM. Antibodies of the present disclosure may include IgG1, IgG2, IgG3, and IgG4 of these isotypes. Multiple allotype sequences due to genetic polymorphisms for the human IgG1, IgG2, IgG3, and IgG4 constant regions are described in "Sequences of proteins of immunological interest," NIH Publication No. 91-3242, and any of these may be used in the present disclosure. In particular, for human IgG1 sequences, the amino acid sequence at positions 356-358 (EU numbering) may be DEL or EEM. Furthermore, for the human Igκ (Kappa) constant region and the human Igλ (Lambda) constant region, multiple allotype sequences due to genetic polymorphisms are described in "Sequences of proteins of immunological interest," NIH Publication No. 91-3242, and either of these sequences may be used in the present disclosure.
[0021] EU numbering and Kabat numbering According to the method used in the present disclosure, the amino acid positions assigned to the CDRs and FRs of an antibody are defined according to Kabat (Sequences of Proteins of Immunological Interest (National Institutes of Health, Bethesda, Md., 1987 and 1991)). Herein, when the antigen-binding molecule is an antibody or an antigen-binding fragment, the amino acids in the variable region are represented according to the Kabat numbering, and the amino acids in the constant region are represented according to the EU numbering based on the Kabat amino acid positions.
[0022] Variable region The term "variable region" or "variable domain" refers to the domain of an antibody heavy or light chain that is involved in binding the antibody to an antigen. The heavy and light chain variable domains (VH and VL, respectively) of natural antibodies typically have a similar structure, with each domain containing four conserved framework regions (FR) and three hypervariable regions (HVR). (See, for example, Kindt et al., Kuby Immunology, 6th ed., W.H. Freeman and Co., page 91 (2007)). A single VH or VL domain may be sufficient to confer antigen-binding specificity. Furthermore, antibodies that bind to a specific antigen may be isolated by screening a complementary library of VL or VH domains, respectively, using a VH or VL domain from an antibody that binds to that antigen. See, e.g., Portolano et al., J. Immunol. 150:880-887 (1993); Clarkson et al., Nature 352:624-628 (1991).
[0023] Hypervariable region As used herein, the term "hypervariable region" or "HVR" refers to each region of an antibody variable domain that is hypervariable in sequence (the "complementarity determining region" or "CDR") and / or forms structurally defined loops (the "hypervariable loops") and / or contains antigen-contacting residues (the "antigen contacts"). Typically, antibodies contain six HVRs: three in the VH (H1, H2, H3) and three in the VL (L1, L2, L3). Exemplary HVRs herein include the following: (a) hypervariable loops occurring at amino acid residues 26-32 (L1), 50-52 (L2), 91-96 (L3), 26-32 (H1), 53-55 (H2), and 96-101 (H3) (Chothia and Lesk, J. Mol. Biol. 196:901-917 (1987)); (b) CDRs occurring at amino acid residues 24-34 (L1), 50-56 (L2), 89-97 (L3), 31-35b (H1), 50-65 (H2), and 95-102 (H3) (Kabat et al., Sequences of Proteins of Immunological Interest, 5th Ed. Public Health Service, National Institutes of Health, Bethesda, MD (1991)); (c) antigenic contacts occurring at amino acid residues 27c-36 (L1), 46-55 (L2), 89-96 (L3), 30-35b (H1), 47-58 (H2), and 93-101 (H3) (MacCallum et al. J. Mol. Biol. 262: 732-745 (1996)); and (d) A combination of (a), (b), and / or (c), comprising HVR amino acid residues 46-56 (L2), 47-56 (L2), 48-56 (L2), 49-56 (L2), 26-35 (H1), 26-35b (H1), 49-65 (H2), 93-102 (H3), and 94-102 (H3). Unless otherwise indicated, HVR residues and other residues in the variable domain (e.g., FR residues) are numbered herein according to Kabat et al., supra.
[0024] Framework "Framework" or "FR" refers to variable domain residues other than hypervariable region (HVR) residues. The FR of a variable domain typically consists of four FR domains: FR1, FR2, FR3, and FR4. Accordingly, the HVR and FR sequences typically appear in VH (or VL) in the following order: FR1-H1(L1)-FR2-H2(L2)-FR3-H3(L3)-FR4.
[0025] ·Fc area The Fc region comprises an amino acid sequence derived from the constant region of an antibody heavy chain. The Fc region is a portion of the antibody heavy chain constant region, spanning from the N-terminus of the hinge region of the papain cleavage site at approximately amino acid position 216 (EU numbering) to the hinge, CH2, and CH3 domains. The Fc region can be obtained from human IgG1, but is not limited to a specific IgG subclass. Preferred examples of the Fc region include Fc regions that have FcRn-binding activity in the acidic pH range, as described below. Preferred examples of the Fc region also include Fc regions that have Fcγ receptor-binding activity, as described below. Non-limiting examples of such Fc regions include the Fc regions of human IgG1 (SEQ ID NO: XX), IgG2 (SEQ ID NO: XX), IgG3 (SEQ ID NO: XX), or IgG4 (SEQ ID NO: XX).
[0026] ·Low molecular antibody The antibodies used in the present disclosure are not limited to full-length antibody molecules, but may also be minibodies or modified versions thereof. Minibodies include antibody fragments in which a portion of a full-length antibody (e.g., a whole antibody such as whole IgG) is deleted, and are not particularly limited as long as they have antigen-binding activity. The minibodies of the present disclosure are not particularly limited as long as they are a portion of a full-length antibody, but preferably contain a heavy chain variable region (VH) or / and a light chain variable region (VL). The amino acid sequence of VH or VL may be substituted, deleted, added, and / or inserted. Furthermore, portions of VH and / or VL may be deleted as long as they have antigen-binding activity. The variable regions may also be chimerized or humanized. Specific examples of antibody fragments include Fab, Fab', F(ab')2, and Fv. Specific examples of minibodies include Fab, Fab', F(ab')2, Fv, scFv (single chain Fv), diabody, sc(Fv)2 (single chain (Fv)2), etc. Multimers of these antibodies (e.g., dimers, trimers, tetramers, polymers) are also included in the minibodies of the present disclosure. Diabodies are bivalent minibodies constructed by gene fusion (Holliger et al., Proc. Natl. Acad. Sci. USA 90, 6444-6448 (1993), European Patent Publication EP404097, and PCT Publication WO1993 / 011161, etc.). Diabodies are dimers composed of two polypeptide chains. Typically, the VL and VH in each polypeptide chain are linked by a linker that is so short, for example, about five residues, that they cannot bind to each other. Because the linker between the VL and VH encoded on the same polypeptide chain is too short to form a single-chain variable region fragment, they form a dimer, and thus diabodies have two antigen-binding sites. An scFv can be obtained by linking the heavy chain variable region and light chain variable region of an antibody. In this scFv, the heavy chain variable region and light chain variable region are linked via a linker, preferably a peptide linker (Huston et al. (Proc. Natl. Acad. Sci. USA (1988) 85, 5879-5883)). The heavy chain variable region and light chain variable region in an scFv may be derived from any of the antibodies described herein. There are no particular limitations on the peptide linker that links the variable regions, and examples that can be used include any single-chain peptide consisting of approximately 3 to 25 residues, as well as the peptide linkers described below. sc(Fv)2 is a minibody made into a single chain by linking two VHs and two VLs with a linker or the like (Hudson et al. (J. Immunol. Methods (1999) 231, 177-189)). sc(Fv)2 can be prepared, for example, by linking scFvs with a linker. Furthermore, preferred antibodies have two VHs and two VLs arranged in the following order, starting from the N-terminus of the single-chain polypeptide: VH, VL, VH, VL ([VH] linker [VL] linker [VH] linker [VL]). The order of the two VHs and two VLs is not limited to the above arrangement, and they may be arranged in any order. For example, the following arrangements may be mentioned. -[VL] linker [VH] linker [VH] linker [VL] -[VH] linker [VL] linker [VL] linker [VH] -[VH] linker [VH] linker [VL] linker [VL] -[VL] linker [VL] linker [VH] linker [VH] -[VL] linker [VH] linker [VL] linker [VH] The linker used to link the antibody variable regions may be the same as the linkers described above in the section on antigen-binding molecules. For example, particularly preferred embodiments of sc(Fv)2 in the present disclosure include the following sc(Fv)2: -[VH] peptide linker (15 amino acids) [VL] peptide linker (15 amino acids) [VH] peptide linker (15 amino acids) [VL] When four antibody variable regions are linked, three linkers are usually required, and the same linkers may be used for all of them, or different linkers may be used. Such minibodies can be obtained by treating antibodies with enzymes such as papain or pepsin to generate antibody fragments, or by constructing DNA encoding these antibody fragments or minibodies, introducing it into an expression vector, and then expressing it in an appropriate host cell (see, for example, Co, MS et al., J. Immunol. (1994) 152, 2968-2976; Better, M. and Horwitz, AH, Methods Enzymol. (1989) 178, 476-496; Pluckthun, A. and Skerra, A., Methods Enzymol. (1989) 178, 497-515; Lamoyi, E., Methods Enzymol. (1986) 121, 652-663; Rousseaux, J. et al., Methods Enzymol. (1986) 121, 652-663). 663-669; Bird, RE and Walker, BW, Trends Biotechnol. (1991) 9, 132-137).
[0027] Single domain antibodies A suitable example of the antigen-binding domain of the present invention is a single-domain antibody (sdAb). As used herein, the term "single-domain antibody" refers to a domain that exhibits antigen-binding activity by itself, and its structure is not limited as much as possible. Conventional antibodies, such as IgG antibodies, exhibit antigen-binding activity when the variable region is formed by pairing VH and VL, whereas single-domain antibodies are known to be able to exhibit antigen-binding activity by their own domain structure alone, without pairing with other domains. Single-domain antibodies usually have a relatively low molecular weight and exist in the form of a monomer. Examples of single-domain antibodies include, but are not limited to, antigen-binding molecules that congenitally lack light chains, such as camelid VHHs and shark VNARs, or antibody fragments comprising all or a portion of the VH domain or all or a portion of the VL domain of an antibody. Examples of single-domain antibodies, which are antibody fragments comprising all or a portion of the VH / VL domains of an antibody, include, but are not limited to, single-domain antibodies artificially created starting from a human antibody VH or human antibody VL, such as those described in U.S. Patent No. 6,248,516 B1. In some embodiments of the present invention, a single-domain antibody has three CDRs (CDR1, CDR2, and CDR3). Single-domain antibodies can be obtained from animals capable of producing single-domain antibodies or by immunizing animals capable of producing single-domain antibodies. Examples of animals capable of producing single-domain antibodies include, but are not limited to, camelids and transgenic animals into which genes capable of producing single-domain antibodies have been introduced. Camelids include camels, llamas, alpacas, dromedaries, and guanacos. Examples of transgenic animals into which genes capable of producing single-domain antibodies have been introduced include, but are not limited to, the transgenic animals described in International Publication No. WO 2015 / 143414 and U.S. Patent Publication No. US 2011 / 0123527 A1. Humanized single-domain antibodies can also be obtained by substituting human germline sequences or sequences similar thereto for the framework sequences of single-domain antibodies obtained from animals. Humanized single-domain antibodies (e.g., humanized VHHs) are also an embodiment of the single-domain antibodies of the present invention. A "humanized single-domain antibody" refers to a chimeric single-domain antibody comprising amino acid residues from non-human CDRs and human FRs. In one embodiment, a humanized single-domain antibody has all or substantially all CDRs corresponding to those of a non-human antibody, and all or substantially all FRs corresponding to those of a human antibody. In a humanized antibody, even if some of the residues in the FRs do not correspond to those of a human antibody, this is still considered an example in which substantially all of the FRs correspond to those of a human antibody. For example, when humanizing a VHH, which is one embodiment of a single-domain antibody, some of the residues in the FRs must be changed to residues that do not correspond to those of a human antibody (C. Vincke et al., The Journal of Biological Chemistry 284, 3273-3284). Alternatively, single domain antibodies can be obtained from a polypeptide library containing single domain antibodies by ELISA, panning, or the like. Examples of polypeptide libraries containing single domain antibodies include, but are not limited to, naive antibody libraries obtained from various animals or humans (e.g., Methods in Molecular Biology 2012 911 (65-78), Biochimica et Biophysica Acta - Proteins and Proteomics 2006 1764:8 (1307-1319)), antibody libraries obtained by immunizing various animals (e.g., Journal of Applied Microbiology 2014 117:2 (528-536)), or synthetic antibody libraries prepared from antibody genes of various animals or humans (e.g., Journal of Biomolecular Screening 2016 21:1 (35-43), Journal of Biological Chemistry 2016 291:24 (12641-12657), AIDS 2016 30:11 (1691-1701)).
[0028] Library As used herein, the term "library" refers to antigen-binding molecules comprising multiple antigen-binding domains whose sequences differ from one another, and / or nucleic acids or polynucleotides encoding antigen-binding molecules comprising multiple antigen-binding domains whose sequences differ from one another. The antigen-binding molecules comprising antigen-binding domains contained in the library and / or the nucleic acids encoding the antigen-binding molecules comprising antigen-binding domains do not have a single sequence, but rather comprise multiple antigen-binding molecules whose sequences differ from one another and / or nucleic acids encoding multiple antigen-binding molecules whose sequences differ from one another. In one embodiment of the present disclosure, a fusion polypeptide can be prepared between an antigen-binding molecule of the present disclosure and a heterologous polypeptide. In some embodiments, the fusion polypeptide can be fused to at least a portion of a viral coat protein, for example, a viral coat protein selected from the group consisting of pIII, pVIII, pVII, pIX, Soc, Hoc, gpD, pVI, and mutants thereof. In one embodiment, the antigen-binding molecules of the present disclosure may be ScFv, Fab fragments, F(ab)2, or F(ab')2. In another embodiment, a library is provided that primarily comprises a plurality of fusion polypeptides, each of which has a different sequence from the other, between these antigen-binding molecules and a heterologous polypeptide. Specifically, a library is provided that primarily comprises a plurality of fusion polypeptides, each of which has a different sequence from the other, fused with these antigen-binding molecules and at least a portion of a viral coat protein, for example, selected from the group consisting of pIII, pVIII, pVII, pIX, Soc, Hoc, gpD, pVI, and variants thereof. The antigen-binding molecules of the present disclosure may further comprise a dimerization domain. In one embodiment, the dimerization domain may be located between the antibody heavy or light chain variable region and at least a portion of the viral coat protein. The dimerization domain may comprise at least one dimerization sequence and / or a sequence containing one or more cysteine residues. The dimerization domain may preferably be linked to the C-terminus of the heavy chain variable region or constant region. The dimerization domain can have various structures depending on whether the antibody variable region is produced as a fusion polypeptide component with a viral coat protein component (i.e., does not have an amber stop codon after the dimerization domain) or whether the antibody variable region is produced primarily without a viral coat protein component (e.g., has an amber stop codon after the dimerization domain). When the antibody variable region is produced primarily as a fusion polypeptide with a viral coat protein component, bivalent display is achieved by one or more disulfide bonds and / or a single dimerization sequence. In one non-limiting embodiment of the library of the present disclosure, a library of 1.2 x 10 8 A library with a diversity of 1.2 × 10 8 Examples of such a library include the above-mentioned antigen-binding molecule comprising multiple antigen-binding domains whose sequences differ from one another, or a library comprising nucleic acids encoding antigen-binding molecules comprising multiple antigen-binding domains whose sequences differ from one another. As used herein, the term "different sequences" in the phrase "antigen-binding molecules comprising multiple antigen-binding domains with different sequences" means that the sequences of the individual antigen-binding molecules in a library are different from each other. In other words, the number of different sequences in a library reflects the number of independent clones with different sequences in the library, and is sometimes referred to as the "library size." In a typical phage display library, 10 6 From 10 12 By applying known techniques such as ribosome display, the library size can be increased to 10 14 However, the actual number of phage particles used in panning selection of a phage library is usually 10 to 10,000 times larger than the library size. This excess, also called the "library equivalent number," indicates that there may be 10 to 10,000 individual clones with the same amino acid sequence. Therefore, the term "different in sequence from each other" in the present disclosure refers to the fact that the sequences of the individual antigen-binding molecules in the library, excluding the library equivalent number, are different from each other, more specifically, the number of antigen-binding molecules whose sequences differ from each other is 10 to 10,000. 6 From 10 14 molecules, preferably 10 7 From 10 12 molecules, more preferably 10 8 From 10 11 , particularly preferably 10 8 From 10 10 It means to exist. Furthermore, the term "plurality" in the description of antigen-binding molecules comprising multiple antigen-binding domains whose sequences differ from one another, or / and libraries primarily composed of nucleic acids encoding antigen-binding molecules comprising multiple antigen-binding domains whose sequences differ from one another, typically refers to a collection of two or more types of substances, such as antigen-binding molecules, fusion polypeptides, polynucleotide molecules, vectors, or viruses of the present disclosure. For example, if two or more substances differ from one another in a specific trait, this indicates that two or more types of the substance exist. An example would be a mutant in which an amino acid mutation is observed at a specific amino acid site in the amino acid sequence. For example, if there are two or more antigen-binding molecules that are substantially the same, preferably identical, except for the amino acid at a specific amino acid site, then there are multiple antigen-binding molecules. In another example, if there are two or more polynucleotide molecules that are substantially the same, preferably identical, except for the bases encoding the amino acids at a specific amino acid site, then there are multiple polynucleotide molecules.
[0029] (Definition of experimental method) [Characteristics evaluation] In one embodiment of the present disclosure, a trained model is generated by performing machine learning based on sequence information of an antigen-binding molecule and evaluation result information of the properties of the antigen-binding molecule. Non-limiting examples of the properties of the antigen-binding molecule include, but are not limited to, evaluation of the affinity, pharmacological activity, physical properties, kinetics, and safety of the antigen-binding molecule.
[0030] Affinity evaluation The method for evaluating the affinity of an antigen-binding molecule is not particularly limited, and can be evaluated by measuring the binding activity between the antigen-binding molecule and the antigen. "Binding activity" refers to the strength of the total non-covalent interactions between one or more binding sites of a molecule (e.g., an antibody) and the molecule's binding partner (e.g., an antigen). Here, "binding activity" is not strictly limited to a 1:1 interaction between members of a binding pair (e.g., an antibody and an antigen). For example, when members of a binding pair reflect a monovalent 1:1 interaction, binding activity refers to the intrinsic binding affinity ("affinity"). When members of a binding pair are capable of both monovalent and multivalent binding, binding activity is the sum of these binding forces. The binding activity of molecule X to its partner Y can generally be expressed by the dissociation constant (KD) or "amount of analyte bound per unit amount of ligand." Binding activity can be measured by conventional methods known in the art, including those described herein. Conditions other than the concentration of the target tissue-specific compound can be appropriately determined by those skilled in the art. In certain embodiments, the antigen-binding molecule provided herein is an antibody, and the binding activity of the antibody is ≦1 μM, ≦100 nM, ≦10 nM, ≦1 nM, ≦0.1 nM, ≦0.01 nM, or ≦0.001 nM (e.g., 10 -8 M or less, e.g. 10 -8 M~10 -13 M, e.g. 10 -9 M~10 -13 is the dissociation constant (KD) of the molecule (M).
[0031] In one embodiment, antibody binding activity is measured using a ligand capture method based on surface plasmon resonance analysis, e.g., a BIACORE™ T200 or BIACORE™ 4000 (GE Healthcare, Uppsala, Sweden). BIACORE™ Control Software is used to operate the instrument. In one embodiment, an amine coupling kit (GE Healthcare, Uppsala, Sweden) is used according to the supplier's instructions to immobilize a ligand capture molecule, such as an anti-tag antibody, anti-IgG antibody, or protein A, on a carboxymethyldextran-coated sensor chip (GE Healthcare, Uppsala, Sweden). The ligand capture molecule is diluted with 10 mM sodium acetate solution at an appropriate pH and injected at an appropriate flow rate and injection time. Binding activity measurements are performed using a buffer containing 0.05% polysorbate 20 (also known as Tween®-20) as the measurement buffer, at a flow rate of 10-30 μL / min, and at a temperature of preferably 25°C or 37°C. When measurements are performed by capturing an antibody as a ligand on a ligand capture molecule, the antibody is injected and captured in a desired amount, and then a serial dilution (analyte) of an antigen and / or Fc receptor prepared using the measurement buffer is injected. When measurements are performed by capturing an antigen and / or Fc receptor as a ligand on a ligand capture molecule, the antigen and / or Fc receptor is injected and captured in a desired amount, and then a serial dilution (analyte) of an antibody prepared using the measurement buffer is injected.
[0032] In one embodiment, the measurement results are analyzed using BIACORE® Evaluation Software. Kinetic parameter calculations are performed by simultaneously fitting the binding and dissociation sensorgrams using a 1:1 binding model, and the binding rate (k or ka), dissociation rate (k or k), and equilibrium dissociation constant (KD) can be calculated. When binding activity is weak, particularly when dissociation is rapid and calculation of kinetic parameters is difficult, the equilibrium dissociation constant (KD) may be calculated using a steady state model. Another parameter of binding activity that can be calculated is the "amount of analyte bound per unit amount of ligand," calculated by dividing the amount of analyte bound (RU) at a specific concentration by the amount of ligand captured (RU).
[0033] As the value of antigen-binding activity, KD (dissociation rate constant) can be used when the antigen is a soluble molecule, and apparent kD (apparent dissociation rate constant) can be used when the antigen is a membrane-type molecule. kD (dissociation rate constant) and apparent KD (apparent dissociation rate constant) can be measured by methods known to those skilled in the art, for example, using Biacore (GE Healthcare), a flow cytometer, etc.
[0034] Another aspect of property evaluation is a method for selecting antigen-binding molecules using a display library. One aspect includes panning using phage display. For example, affinity evaluation can be performed by preparing a phage library displaying multiple different antigen-binding molecules, contacting the prepared phages with a target antigen, and then washing away unbound phages to concentrate phages displaying antigen-binding molecules that interact with the target antigen. Sequences with affinity for the target antigen can be identified by analyzing the nucleic acid sequences encoding the antigen-binding molecules contained in the concentrated phages. Another aspect includes panning using mammalian cell display. For example, pharmacological activity evaluation using this display system can be performed by expressing a library containing multiple different antigen-binding molecules in target mammalian cells and varying reporter activity or the like depending on the effect the molecules have on the same cells, allowing cells carrying antigen-binding molecule genes with the desired pharmacological activity to be isolated using a flow cytometer or the like. Furthermore, as an example of physical property evaluation using this display system, a library containing multiple different antigen-binding molecules is expressed in target mammalian cells, and the expression level is stained with an antibody specific to the antigen-binding molecule, allowing cells harboring antigen-binding molecule genes that can be stably and highly expressed to be isolated using a flow cytometer, etc. Evaluation of the properties of antigen-binding molecules by panning is not limited to techniques using phages or mammalian cells, and various techniques can be used as long as they can display antigen-binding molecules, including, but not limited to, techniques in which the antigen-binding molecules are displayed on ribosomes, mRNA, viruses other than phages, and bacteria such as Escherichia coli. Another aspect of characterization is a method of obtaining antibody gene sequences from immune cells derived from an individual, or a method of obtaining antibody protein sequences from serum. For example, affinity evaluation, in which antibody gene sequences are extracted from immune cells, involves administering a target antigen protein to an individual to induce immune sensitization, and then extracting genes from immune cells that have antibody genes that bind to the target antigen, thereby making it possible to identify sequences that have affinity for the target antigen. The antigen that induces immune sensitization is not limited to the technique using the above protein, but a gene encoding the protein or a cell that expresses the protein can also be used. Furthermore, target individuals include, but are not limited to, humans, mice, rats, hamsters, rabbits, monkeys, chickens, camels, llamas, and alpacas. Furthermore, examples of methods for analyzing the nucleic acid sequences and occurrence frequencies include, but are not limited to, cloning a genetically modified organism having the nucleic acid sequence of each antigen-binding molecule and analyzing it by the Sanger method using capillary electrophoresis, and analyzing it using a next-generation sequencer. When analyzing the nucleic acid sequences, it is also possible to judge the strength of a characteristic based on the frequency of occurrence. For example, it is possible to estimate that antigen-binding molecules encoded by sequences that occur frequently after enrichment have a high characteristic, and that antigen-binding molecules encoded by sequences that occur less frequently after enrichment have a lower characteristic than antigen-binding molecules encoded by sequences that occur more frequently. Furthermore, the techniques for obtaining information on antigen-binding molecules derived from the display library or an individual can be applied to various property evaluations, and are not limited to those described above.
[0035] Pharmacological activity evaluation The pharmacological activity of an antigen-binding molecule can be evaluated by any method, including, for example, measuring the neutralizing activity, agonistic activity, or cytotoxic activity of the antigen-binding molecule. Examples of pharmacological activity evaluation include cytotoxic activity evaluation, such as antibody-dependent cell-mediated cytotoxicity (ADCC), complement-dependent cytotoxicity (CDC), T-cell-dependent cytotoxicity (TDCC), and antibody-dependent cellular phagocytosis (ADCP). CDC activity refers to cytotoxic activity mediated by the complement system. ADCC activity refers to the activity of immune cells, etc., binding to the Fc region of an antigen-binding molecule containing an antigen-binding domain that binds to a membrane-type molecule expressed on the cell membrane of the target cell via the Fcγ receptor expressed on the immune cell, causing the immune cell to injure the target cell. TDCC activity refers to the activity of a T cell to damage a target cell by bringing the target cell and the T cell into close proximity using a bi-specific antibody containing an antigen-binding domain that binds to a membrane molecule expressed on the cell membrane of the target cell and an antigen-binding domain for one of the constituent subunits of the T cell receptor (TCR) complex on the T cell, particularly an antigen-binding domain that binds to the CD3 epsilon chain. Whether an antigen-binding molecule of interest has ADCC activity, CDC activity, TDCC activity, or ADCP activity can be determined by known methods. Neutralizing activity refers to the activity of inhibiting the biological activity of a ligand, such as a virus or a toxin, that has biological activity against cells. That is, a substance with neutralizing activity refers to a substance that binds to the ligand or a receptor to which the ligand binds, and inhibits the binding of the ligand to the receptor. A receptor whose binding to a ligand is blocked by neutralizing activity is no longer able to exert biological activity mediated by the receptor. When the antigen-binding molecule is an antibody, such an antibody with neutralizing activity is generally called a neutralizing antibody, and neutralizing activity can be measured by measuring the inhibitory activity of the ligand-receptor binding. Ligands that have biological activity against cells are not limited to viruses, toxins, etc.; neutralizing activity also refers to the inhibitory activity of physiological actions induced by the binding of endogenous ligands, such as cytokines and chemokines, to receptors. Furthermore, neutralizing activity is not limited to inhibiting the binding of a ligand to a receptor; it also refers to the activity of inhibiting the function of a protein with biological activity, and an example of the function of the protein is enzymatic activity.
[0036] Physical property evaluation Methods for evaluating the physical properties of antigen-binding molecules are not particularly limited, and examples include thermal stability, chemical stability, solubility, viscosity, light stability, long-term storage stability, and nonspecific adsorption. These various physical properties can be measured by methods known to those skilled in the art. The evaluation method is not particularly limited, and stability evaluations, such as thermal stability, chemical stability, light stability, stability against mechanical stimuli, and long-term storage stability, can be performed by measuring the degradation, chemical modification, and aggregation of the antigen-binding molecules before and after the treatments targeted for the stability evaluation, such as heat treatment, exposure to a low pH environment, light exposure, mechanical stirring, and long-term storage. Non-limiting examples of measurement methods for performing such stability evaluations include, but are not limited to, chromatographic methods such as ion exchange chromatography and size exclusion chromatography, mass spectrometry, and electrophoresis, and measurements can be performed by various methods known to those skilled in the art. Other examples of physical property evaluations include, but are not limited to, evaluation of protein solubility using polyethylene glycol precipitation, evaluation of viscosity using small-angle X-ray scattering, and evaluation of nonspecific binding based on binding to extracellular matrix (ECM). Furthermore, physical properties such as protein expression level, binding to purification resins or purification ligands, and surface charge can be evaluated as long as they can be measured by methods known to those skilled in the art.
[0037] Dynamic assessment The method for evaluating the kinetics of antigen-binding molecules is not particularly limited, and can be performed by administering the antigen-binding molecule to animals such as mice, rats, monkeys, and dogs and measuring the amount of the antigen-binding molecule in the blood over time after administration, a method widely known to those skilled in the art as pharmacokinetic (PK) evaluation. In addition to methods for directly evaluating PK, the kinetic behavior of the antigen-binding molecule can also be predicted from the amino acid sequence of the antigen-binding molecule by calculating the surface charge, isoelectric point, etc. of the antigen-binding molecule using software.
[0038] Safety assessment Methods for evaluating the safety of antigen-binding molecules are not particularly limited, and examples include immunogenicity prediction tools such as ISPRI Web-Based Immunogenicity Screening (EpiVax), HLA binding assessment of fragment peptides of antigen-binding molecules, detection of T cell epitopes and evaluation of immunogenicity using MAPPs (MHC-Associated Peptide Proteomics) or T cell proliferation assessment, etc. In addition, evaluation can be performed as long as it is measurable by methods known to those skilled in the art, such as binding to rheumatoid factor (RF), evaluation of immune responses using PBMCs or whole blood, and platelet aggregation assessment.
[0039] (Definitions of terms and techniques used in machine learning (LSTM, RF)) An RNN (Recurrent Neural Network) is a neural network that connects multiple neural networks. An example of its application to peptide sequencing is given by Muller AT et al. (J Chem Inf Model. 2018 Feb 26;58(2):472-479.). LSTM (Long Short Term Memory) is a special type of RNN that has been adapted to have excellent long-term memory. GRU (Gated Recurrent Unit) is a special type of RNN that has neurons equivalent to long-term memory. GAN is an adversarial network, a machine learning method that aims for more accurate classification by using both a model that attempts to classify accurately and a model that generates deceptive samples. A VAE (Variational AutoEncoder) is a neural network that uses the same data for both input and output layers in a supervised learning process based on the calculus of variations. The Flow deep generative model is a model that learns the distribution of data based on log-likelihood through reversible variable transformation. Gaussian processes are machine learning techniques that output not only predicted values for a given input, but also the distribution of predicted values. The distribution of predicted values can be considered the reliability of the predicted value. Bayesian optimization is a technique that samples (measures) points that improve prediction accuracy based on the prediction results of the Gaussian process, updates the prediction model including new measured values, and makes more accurate predictions. Furthermore, prediction models can be updated repeatedly.
[0040] Probabilistic models (HMM, MM) A Markov Model (MM) is a model in which multiple states and the transition probabilities between them are given. The transition probability to the next state is determined only by information about the previous state. An HMM (Hidden Markov Model) is a model that is given multiple states and the transition probabilities between them, and outputs a different quantity for each state with a defined probability. The transition probability to the next state is determined only by information about the previous state.
[0041] - Amino acid recognition method (sequence information expressed as strings, numerical vectors, and physical properties) One possible input method for antibody sequence machine learning techniques is to treat the sequence as a string and input the string as a string. Another possible method is to convert the characters in the sequence into a numerical value for each position in the sequence using the physical properties of the amino acid at that position (molecular weight, charge, hydrophobicity, side chain volume, etc.). Another possible method is to convert the entire string of the sequence into a numerical value using statistics on which amino acids are likely to appear around each amino acid (Doc2Vec method). For example, when using the Doc2Vec method, the computer regards the target amino acid sequence as a sentence. The computer divides the amino acid sequence into groups of a predetermined number of amino acids (for example, three, but it can be other than three) in order from front to back according to the sequence order. For each group after division, the computer generates a string of characters representing the amino acids as a word. The computer uses the Doc2Vec method to map each amino acid sequence, where pairs of amino acids are linked together, into a vector space as a sentence where words are linked together in order. In this way, the computer may analyze the sequence using a method used for document analysis, regarding the sequence as a sentence and the pairs of amino acids in the sequence as words.
[0042] (First embodiment) Hereinafter, a first embodiment of the present invention will be described in detail with reference to the drawings. In this embodiment, in the information processing system 1 of FIG. 1 , the server 30 generates a sequence-trained model (an example of a "first trained model") by performing machine learning based on sequence information representing the sequences of antigen-binding molecules. The sequence-trained model is a trained model that learns the characteristics of the sequence represented by the input sequence information and outputs prediction target sequence information (an example of "virtual sequence information") as a result of the learning. The prediction target sequence represented by the prediction target sequence information is a virtual sequence obtained by mutating at least one of the amino acids constituting at least one of the sequences of the antigen-binding molecules used for learning. This allows the information processing system 1 to generate virtual sequences of antigen-binding molecules in which some of their amino acids have been mutated. For example, the sequence-trained model is trained using sequence information of antigen-binding molecules with desired properties. In this case, the server 30 can generate, for example, a large number of virtual sequences that are likely to have the desired properties as a virtual sequence group. In this embodiment, an example is described in which the frequency of binding to an antigen by panning using the antigen is used as characteristic evaluation. Panning will be described later.
[0043] <Information Processing System> FIG. 1 is a schematic diagram showing an example of an information processing system 1 according to the first embodiment. The information processing system 1 includes a user terminal 10, a next-generation sequencer 20, and a server 30. The user terminal 10, the next-generation sequencer 20, and the server 30 are connected via a network NW. The network NW is, for example, an information communication network such as a LAN (Local Area Network) or the Internet. The information communication network may be a wired or wireless network, or may be a network that combines various networks. The user terminal 10, the next-generation sequencer 20, and the server 30 may exchange data via storage media such as an HDD (Hard Disk Drive) or a USB memory.
[0044] The user terminal 10 is, for example, a personal computer through which a user inputs and outputs data, and may also be a portable terminal such as a tablet terminal or a smartphone. The next-generation sequencer 20 is a device that analyzes the base sequence of DNA (deoxyribonucleic acid). The server 30 is an information processing device such as a server. The server 30 performs learning using analysis result information that indicates the analysis results by the next-generation sequencer 20. The server 30 transmits output information to the user terminal 10 based on input information from the user terminal 10 and the learning results.
[0045] For example, in panning of multiple antibodies and a target antigen (an example of "affinity evaluation"), the next-generation sequencer 20 measures and analyzes each of the antibodies contained in a sample, and outputs analysis result information including sequence information representing the base sequence of each antibody and its occurrence frequency (an example of "evaluation result information of affinity evaluation"; also called the number of reads). The occurrence frequency is the proportion of the number of each sequence among the sequences of antibodies that bind to the target antigen, when the total number of sequences analyzed by the next-generation sequencer 20 (total number of reads) is used as a parameter. However, the present invention is not limited to this, and the occurrence frequency may also be the number of sequences analyzed by the next-generation sequencer 20.
[0046] The server 30 receives the analysis result information via the network NW or a storage medium, and generates (an example of acquisition) a group of sequences having desirable properties as a training data set according to the sequence information and occurrence frequency included in the analysis result information. The server 30 learns based on the training data set, and stores a sequence-trained model in which the characteristics of sequences having desirable properties have been learned. The server 30 generates, based on the stored sequence-trained model, a group of new virtual sequences having the characteristics of sequences with desirable properties as sequences to be predicted.
[0047] The server 30 then predicts a prediction score representing affinity information with the target antigen (an example of "affinity information representing affinity with the target antigen") for each of the generated prediction target sequence information representing the prediction target sequence using a characteristic prediction trained model (an example of a "second trained model") described below. The affinity information is information indicating whether an antibody binds to a target antigen (is a binding antibody) or does not bind to it (is a non-binding antibody). The server 30 transmits candidate antibody information representing candidates for antibodies that are expected to bind to the target antigen according to the predicted prediction score to the user terminal 10. The user terminal 10 displays the candidate antibody information according to the received prediction score.
[0048] This allows the information processing system 1 to generate a large amount of sequence information narrowed down to sequences that are predicted to have better characteristics than when sequence information predicting characteristics with a target antigen is randomly generated. Therefore, the information processing system 1 can provide information on a desired antibody while further reducing processing time or processing load.
[0049] <Affinity evaluation (panning)> FIG. 2 is an explanatory diagram for explaining an example of a series of panning according to this embodiment. Each repeated panning (referred to as "binding test" in the figure) in the series of pannings will be explained. Note that the mth (m=1 to M: m is a natural number) round of panning in the series of pannings will also be referred to as the mth round of panning or round m of panning. Each panning is carried out through the following four steps (P1) to (P4). (P1) Reaction of target antigen with antibody (P2) Washing of antibodies that did not bind to the target antigen (indicated as "unbound antibodies" in the figure) (P3) Elution of antibody bound to target antigen (referred to as "bound antibody" in the figure) (P4) Amplification of DNA that serves as a template for producing eluted antibodies Here, the antibodies are associated one-to-one with the above-mentioned DNA using various existing antibody display techniques.
[0050] In the first panning (round 1), a collection of multiple antibodies (hereinafter also referred to as an "antibody library") is subjected to panning. This collection in round 1 is prepared in advance. In the second and subsequent rounds, a collection of binding antibodies determined to have bound to the target antigen in the previous round (an example of "antibodies with affinity for the target antigen") is subjected to panning. In other words, in the second and subsequent rounds, a collection of non-binding antibodies determined not to have bound to the target antigen in the previous round (an example of "antibodies evaluated as having low affinity") is not subjected to panning. More specifically, the antibodies subjected to the second and subsequent rounds (binding antibodies in the immediately preceding round) are prepared using DNA amplified in the immediately preceding round. Note that the series of panning is terminated, for example, after panning has been repeated a predetermined number of times. However, the present invention is not limited to this, and the series of panning may be terminated when the number of binding antibodies becomes low or at the discretion of the experimenter.
[0051] In each panning, experimental conditions (an example of "evaluation conditions") are set. The experimental conditions are changeable conditions for the reaction between the target antigen and the antibody. The experimental conditions include the target antigen conditions, the antibody conditions, the conditions of the solution present in the reaction site, the reaction time and reaction temperature of the target antigen and the antibody, etc. The target antigen conditions include, for example, the concentration of the target antigen and molecular information of the target antigen. The concentration of the target antigen indicates the concentration of the target antigen in the reaction solution where the target antigen reacts with the antibody. The molecular information of the target antigen includes, for example, the sample name and amino acid sequence. Antibody conditions include, for example, antibody display method, antibody origin, domain type, and germline. The antibody display method refers to the antibody display method for antibodies subjected to panning. The antibody origin refers to the origin of the antibody subjected to panning, such as human, mouse, rat, hamster, rabbit, monkey, chicken, camel, llama, alpaca, or artificially synthesized. The domain type refers to, for example, heavy chain and light chain. The solution conditions are, for example, the composition of the buffer (solution). "Buffer composition" refers to the solution conditions such as the solution composition of the reaction solution and the hydrogen ion exponent (pH). The reaction time refers to the time during which the target antigen and the antibody coexist in the solution, and the reaction temperature refers to the set temperature of the solution when the target antigen and the antibody coexist in the solution. In the example of Fig. 2, the experimental condition for round 1 is condition 1, and the experimental condition for round 2 is condition 2. Experiment attribute information indicating these experimental conditions can be managed for each panning in the server 30. However, the experimental conditions may be the same for a series of rounds, in which case conditions 1, 2, ..., N, N+1, ... in Fig. 2 will be the same experimental conditions.
[0052] In each panning, after step (P3), the antibody DNA is amplified and then the base sequence is analyzed by the next-generation sequencer 20. As the analysis result, the next-generation sequencer 20 outputs sequence information indicating the antibody base sequence for each of the multiple antibodies, and antibody evaluation result information. The evaluation result information includes, for each antibody, the frequency of appearance in each round and the rate of change in frequency of appearance between rounds. The base sequence is converted into an amino acid sequence on the server 30. Here, one antibody is composed of a combination of a heavy chain (H chain) portion and a light chain (L chain) portion. In the analysis result information, the sequence information of one antibody is obtained by measuring and analyzing the amino acid sequence of the heavy chain (H chain) portion (also referred to as the "heavy chain sequence") or the amino acid sequence of the light chain (L chain) portion (also referred to as the "light chain sequence") separately. In other words, even if the next-generation sequencer 20 cannot identify the combination of the heavy chain sequence and the light chain sequence of an antibody, it can identify the heavy chain sequence and the light chain sequence. The next-generation sequencer 20 outputs separately the sequence information and evaluation result information indicating the heavy chain sequence and the sequence information and evaluation result information indicating the light chain sequence. However, the present invention is not limited to this, and the next-generation sequencer 20 may measure the heavy chain and the light chain simultaneously and output sequence information indicating the heavy chain sequence and the light chain sequence, and evaluation result information.
[0053] 2 shows that the next-generation sequencer 20 outputs heavy chain sequence A, heavy chain sequence B, light chain sequence C, and light chain sequence D as sequence information of the bound antibody as analysis result information of panning in round 1. FIG. 2 also shows that the output of the occurrence frequency A1 of heavy chain sequence A, the occurrence frequency B1 of heavy chain sequence B, the occurrence frequency C1 of light chain sequence C, and the occurrence frequency D1 of light chain sequence D as result information of panning in round 1.
[0054] In the above, the next-generation sequencer 20 determines the base sequence of the bound antibody based on the sequence information and evaluation result information obtained from a series of pannings. Note that the series of panning shown in Figure 2 may be considered as one set, and multiple series of panning may be performed for each set. For example, the target antigen is the same for all the series of panning in the set. At least one series of panning in the set is different from the other series of panning in at least one of the antibody library and the experimental conditions.
[0055] In a series of pannings, the server 30 acquires a training data set for each of panning rounds 1 to M, and performs learning based on these training data sets. Here, the training data sets include a training data set for at least one round, i.e., a training data set related to the sequence after panning round I.
[0056] In a series of pannings, panning in a certain round (e.g., the N+1th round) is performed using an antibody (an example of an antibody with affinity) that emerged in the previous round (e.g., the Nth round) of panning and the target antigen. Here, the emerged antibody may be an antibody whose appearance frequency is higher than a predetermined threshold, or an antibody whose appearance frequency is higher than a predetermined ranking. In this way, in a series of pannings, subsequent pannings are performed using target antigens against antibodies that emerged in the previous round of panning. The information processing system 1 acquires learning data sets from these pannings and performs learning.
[0057] This allows for a larger number of training datasets for frequently occurring antibodies compared to when there is no training dataset from a series of pannings. Therefore, the information processing system 1 can make the characteristics of frequently occurring antibodies more prominent. Alternatively, the information processing system 1 can narrow down the types of antibodies compared to when training from training datasets from pannings that use a large number of types of antibodies in all pannings. This allows the information processing system 1 to reduce processing time or processing load.
[0058] <User device> FIG. 3 is a block diagram showing an example of the user terminal 10 according to this embodiment. The user terminal 10 includes a communication unit 11, an input unit 12, a storage unit 13, a processing unit 14, and a display unit 15.
[0059] The communication unit 11 is a communication module that performs various communications via the network NW. The communication unit 11 performs various communications with the server 30, for example. The input unit 12 is an input device such as a keyboard or a touch panel. The input unit 12 receives input information based on a user operation. The input unit 12 outputs the received input information to the processing unit . The storage unit 13 is, for example, a storage device such as a hard disk drive, a memory, etc. The storage unit 13 stores various programs to be executed by the processing unit 14, such as firmware and application programs, as well as results of processing executed by the processing unit 14.
[0060] The processing unit 14 is a processor such as a central processing unit (CPU). The processing unit 14 transmits various information, such as input information input from the input unit 12, to the server 30 via the communication unit 11. The server 30 stores correspondence information between the input information and the output information (for example, a trained model or a table) in advance, and generates output information for the input information. The processing unit 14 receives the output information generated by the server 30 via the communication unit 11. The processing unit 14 displays the received output information on the display unit 15 (an example of output). When the storage unit 13 stores the corresponding information, the processing unit 14 may read out the corresponding information for the input information, generate output information, and cause the display unit 15 to display the output information.
[0061] The display unit 15 is, for example, a display such as an organic electroluminescence display, a liquid crystal display, etc. The display unit 15 performs display in accordance with the display information generated by the processing unit .
[0062] <Screen flow on user device> FIG. 4 is a diagram showing an example of a screen flow according to this embodiment. This diagram shows an example of a screen flow displayed by the display unit 15. Screen D11 is a screen on which the input unit 12 accepts input information. Screen D12 is a screen on which classification criteria are set, and is displayed when button BT111 is pressed. Screen D12 is a screen on which the processing unit 14 displays output information after the items on screen D11 are entered and the search button is pressed.
[0063] On the screen D11, the input unit 12 accepts, as input information, at least one of target antigen information ("target antigen" in the figure), experiment information ("experiment" in the figure), experimental antibody information ("antibody" in the figure), experiment attribute information ("experimental conditions" in the figure), classification criteria information ("classification criteria" in the figure), position of interest information ("position of interest" in the figure), and mutation information. Here, the target antigen information is information that can identify the target antigen. The target antigen information is, for example, the name of the target antigen, but may also be the antigen sequence or an antigen identifier. The experimental information is information that can identify the experiment, such as information that identifies a series of pannings (also referred to as a "panning group") or a round (one panning), or information that indicates the contents of the experiment. The experimental antibody information is information that can identify a set of antibodies that will be subjected to panning. The experimental antibody information is, for example, a name that identifies the antibody library that will be subjected to round 1 panning. However, the present invention is not limited to this, and the experimental antibody information may also be the name or amino acid sequence of one or more antibodies. The experiment attribute information is information indicating conditions that can be changed for each evaluation in panning, such as the above-mentioned experimental conditions and the infectious titer (cfu) of the eluted phage obtained from each experiment.
[0064] The classification criteria information (an example of criteria related to affinity in affinity evaluation) is information indicating classification criteria for classifying antibodies as binding antibodies or non-binding antibodies in the learning phase. In the experimental phase, sequences isolated in a property evaluation experiment, such as panning, may include sequences that do not have the desired properties. By setting the classification criteria information input by the user, the server 30 can reclassify antibodies analyzed by the next-generation sequencer 20 as binding antibodies or non-binding antibodies in the learning phase. This allows the server 30 to sometimes classify an antibody that was erroneously determined as a binding antibody in the experimental phase as a non-binding antibody in the learning phase. In this case, the information processing system 1 can improve the accuracy of binding antibodies and classification accuracy. The classification criterion information is a threshold value for the occurrence frequency or the rate of change in the occurrence frequency between rounds. These threshold values may be set for each round or for each series of panning.
[0065] The classification criterion information also includes information on multiple candidates (criterion 1, criterion 2, and criterion 3 in the figure) (also referred to as classification criterion candidate information). Three thresholds are input for each of criteria 1, 2, and 3. The thresholds can be set to the frequency of occurrence per round or the rate of change in frequency of occurrence between rounds. In the learning stage, antibodies are classified using each criterion indicated by the multiple classification criterion candidate information, and first and second trained models are generated. Of these, the server 30 selects the trained model with the highest accuracy (high reproducibility of analysis result information) from among these first and second trained models. In this way, the server 30 also verifies multiple candidates for the classification criterion for classifying bound antibodies and unbound antigens. This allows the information processing system 1 to improve classification accuracy compared to when the classification criterion is fixed.
[0066] The positional information of interest is information that indicates the position of an amino acid in an antibody. The positional information of interest is used to narrow down the sequence information to be learned to only the sequence information of a specific position in an antibody. For example, the positional information of interest is information that indicates the position of an amino acid in the variable region of an antibody that is assumed to be important for binding to a target antigen. By setting the attention position information input by the user, the server 30 narrows down the sequence information to the part of the sequence information indicated by the attention position information and performs learning (using the first and second learning models). This allows the information processing system 1 to shorten the sequence information to the attention part, thereby reducing the processing time and processing load due to learning compared to learning the entire sequence. In addition, since the attention position information can be set, learning can be performed on the part with good classification accuracy. The attention position information may be automatically set by input from a computer.
[0067] Mutation information is information indicating the position of an amino acid in an antibody. The mutation information is used to narrow down the portion of the sequence information to be changed for the antibody sequence (referred to as the "prediction target sequence") for which a prediction score is to be calculated. The mutation information is information indicating, for example, positions at which the dissociation constant was improved in other affinity evaluations, or positions of amino acids that are assumed to be important for binding to the target antigen based on structural information of the target antigen, etc. The mutation information may be a position input by a user or a position input from a computer.
[0068] On screen D12, the processing unit 14 displays, as output information, target antigen information ("target antigen" in the figure), candidate antibody information ("antibody candidate" in the figure) representing candidate antibodies that bind to the target antigen, and a predicted score indicating the degree of binding with the target antigen. The processing unit 14 displays, on the display unit 15, the candidate antibody information with the highest degree of binding (for example, the top 20) in descending order of degree. In other words, the display unit 15 outputs the candidate antibody information according to the degree of binding with the target antigen. The predicted score may be a probability of combining, the actual value of the frequency of occurrence, a value normalized by the maximum value of the frequency of occurrence, or a value obtained by performing some kind of calculation on the frequency of occurrence.
[0069] Below, we will explain the use case of the screen in Figure 4. On screen D11, the user sets at least either target antigen information or experimental information as basic settings. Setting either target antigen information or experimental information is required, but setting other items is optional. The user can specify one or more combinations of panning groups or rounds as experimental information. On screen D11, the user can set antibody information or experimental antibody information as search conditions. When search conditions are set on screen D11, candidate antibody information that meets the set search conditions is output on screen D12.
[0070] On screen D11, the user can edit or add each of a plurality of classification criteria as the setting of classification criteria information on screen D11. The user can specify each criterion and set the category (appearance frequency or change rate), number of rounds, and threshold value for the category. On the screen D11, the user can set focus position information as focus position settings. On screen D11, the user can set mutation information as a search condition for the prediction target. When mutation information is set on screen D11, screen D12 displays candidate antibody information whose amino acid sequences differ at positions indicated by the mutation information. In other words, in this case, screen D12 displays candidate antibody information whose amino acid sequences are the same except at positions indicated by the mutation information.
[0071] <Next-generation sequencer> FIG. 5 is a block diagram showing an example of the next-generation sequencer 20 according to this embodiment. The next-generation sequencer 20 includes a communication unit 21, an input unit 22, a storage unit 23, a base sequence measurement unit 24, a control unit 25, and a display unit 26.
[0072] The communication unit 21 is a communication module that performs various communications via the network NW. The communication unit 21 performs various communications, for example, with the server 30. However, the present invention is not limited to this, and the next-generation sequencer 20 may be provided with an output port that outputs data to a storage medium instead of or in addition to the communication unit 21. The input unit 22 is an input device such as a keyboard or a touch panel. The input unit 22 receives input information based on a user operation. The input unit 22 outputs the received input information to the control unit 25. The storage unit 23 is, for example, a storage device such as a hard disk drive, a memory, etc. The storage unit 23 stores various programs to be executed by the control unit 25, such as firmware and application programs, as well as results of processing executed by the control unit 25.
[0073] The base sequence measurement unit 24 is a sequencer that measures base sequences. Samples resulting from panning are placed in the base sequence measurement unit 24. The base sequence measurement unit 24 measures the base sequences contained in the placed samples in accordance with instructions from the control unit 25. The base sequence measurement unit 24 outputs the measurement results to the control unit 25.
[0074] The control unit 25 is a processor such as a central processing unit (CPU). The control unit 25 controls the next-generation sequencing by, for example, controlling the base sequence measurement unit 24 based on input from the input unit 22. The control unit 25 calculates sequence information for each antibody contained in the sample by analyzing the measurement results from the base sequence measurement unit 24. This sequence information is sequence information for the heavy chain sequence or light chain sequence of each antibody. The control unit 25 generates analysis result information (see FIGS. 7 and 8) in which the input information input from the input unit 22, the calculated sequence information, and the occurrence frequency are associated with each other, and analysis result information in which the calculated sequence information is associated with each other. Here, the input information includes, for example, a panning group ID that identifies a series of pannings, information on the target antigens used in the panning, the number of rounds indicating the number of rounds, measured antibody information indicating the antibodies measured in each round, and experimental condition information for the series of pannings. The control unit 25 transmits analysis result information for one or more pannings to the server 30 via the communication unit 21. The control unit 25 also causes the display unit 26 to display various operation screens, information input screens, various information related to the progress of next-generation sequencing, and the like. The input information may include setting information used to control the base sequence measurement unit 24. The occurrence frequency of each piece of sequence information is calculated by the server 30 using the analysis result information. However, the present invention is not limited to this, and the next-generation sequencer 20 or another computer may calculate the occurrence frequency of each piece of sequence information using the analysis result information.
[0075] The display unit 26 is, for example, an organic electroluminescence display, a liquid crystal display, etc. The display unit 26 performs display in accordance with display information generated by the control unit 25.
[0076] <server> FIG. 6 is a block diagram showing an example of the server 30 according to this embodiment. The server 30 includes a communication unit 31, a storage unit 32, and a processing unit 33.
[0077] The communication unit 31 is a communication module that performs various communications via the network NW. The communication unit 31 performs various communications with the user terminal 10 or the next-generation sequencer 20, for example. The storage unit 32 is, for example, a storage device such as a hard disk drive, a memory, etc. The storage unit 32 stores various programs to be executed by the processing unit 33, such as firmware and application programs, as well as results of processing executed by the processing unit 33. The processing unit 33 is a processor such as a central processing unit (CPU). The processing unit 33 generates output information for the input information, for example, based on the input information input from the communication unit 31 and the information stored in the storage unit 32. The communication unit 31 transmits the generated output information to the user terminal 10 via the communication unit 31.
[0078] Specifically, the processing unit 33 acquires analysis result information from the next-generation sequencer 20 via the communication unit 31 and stores it as a dataset in the storage unit 32. At this time, the processing unit 33 converts the base sequence included in the acquired information into the corresponding amino acid sequence. The processing unit 33 generates a training dataset based on the stored dataset and performs training based on the generated training dataset. For example, the processing unit 33 first selects sequence information with a desired occurrence frequency (e.g., above a threshold). The processing unit 33 generates a sequence-trained model by learning the sequence characteristics of the selected sequence information. Next, the processing unit 33 trains, as a training data set, the sequence information and binding determination information according to the occurrence frequency for each round or the rate of change in the occurrence frequency between rounds for each panning group ID or target antigen information. The processing unit 33 generates a characteristic prediction-trained model as a learning result. The processing unit 33 stores the sequence trained model that has learned the characteristics of the sequence and the characteristic prediction trained model for predicting the prediction score in the memory unit 32 as learning results.
[0079] The processing unit 33 acquires input information (e.g., target antigen information, experiment information, experimental antibody information, experiment attribute information, classification criteria information, position of interest information, and mutation information) from the user terminal 10 via the communication unit 31. The processing unit 33 generates prediction target sequence information using the sequence trained model. The processing unit 33 inputs the prediction target sequence information into the characteristic prediction trained model and outputs a prediction score. The processing unit 33 generates candidate antibody information representing candidate antibodies that bind to the target antigen according to the prediction score. The processing unit 33 transmits the generated candidate antibody information to the user terminal 10 via the communication unit 31.
[0080] <Server storage> The following describes in detail the storage unit 32. The storage unit 32 includes an experiment information storage unit 321, a dataset storage unit 322, a classification criterion storage unit 323, a learning dataset storage unit 324, a focus position information storage unit 325, a learning result storage unit 326, a mutation information storage unit 327, a sequence storage unit 328, and a characteristic evaluation information storage unit 329.
[0081] The experiment information storage unit 321 stores the experiment information (see FIG. 7) and the experiment attribute information (see FIG. 8). This information is included in the analysis result information from the next-generation sequencer 20 and is input by the processing unit 33. The dataset storage unit 322 stores, as a dataset, sequence information and evaluation result information (appearance frequency for each round and rate of change in appearance frequency between rounds) for each antibody measured in a series of panning. Here, the dataset storage unit 322 stores separately a dataset for heavy chain sequences (see FIG. 9) and a dataset for light chain sequences (see FIG. 10). These datasets are included in the analysis result information from the next-generation sequencer 20 and are input by the processing unit 33. The dataset to be input does not necessarily have to include both heavy and light chain sequences, and may include only heavy chain sequences or only light chain sequences. Alternatively, the dataset may include both heavy and light chain sequences read at the same time, or may be input as a single linked sequence.
[0082] The classification criterion storage unit 323 (an example of a criterion storage unit) stores classification criterion information. As described above, the classification criterion information includes a plurality of pieces of classification criterion candidate information. This information is included in input information from the user terminal 10 and is set by the processing unit 33. However, the present invention is not limited to this, and the classification criterion information (a plurality of pieces of classification criterion candidate information) may be set in the classification criterion storage unit 323 in advance. The learning dataset storage unit 324 stores, as a learning dataset, sequence information and binding determination information corresponding to evaluation result information (appearance frequency per round and rate of change in appearance frequency between rounds) for each antibody including a combination of a heavy chain sequence and a light chain sequence. However, the learning dataset storage unit 324 may store, as a learning dataset, sequence information for the heavy chain sequence and the light chain sequence separately, and binding determination information corresponding to the evaluation result information.
[0083] The attention position information storage unit 325 stores attention position information. This information is included in the input information from the user terminal 10 and is set by the processing unit 33. The learning result storage unit 326 stores the sequence trained model generated by the prediction target sequence generation unit PA and the characteristic prediction trained model generated by the learning unit 334. The mutation information storage unit 327 stores the mutation information. This information is included in the input information from the user terminal 10 and is set by the processing unit 33. The sequence storage unit 328 stores prediction target sequence information indicating the amino acid sequence of the prediction target sequence. This prediction target sequence information is generated and set by the processing unit 33 using a sequence-trained model. The characteristic evaluation information storage unit 329 stores, for each prediction target sequence, a prediction score predicted by the processing unit 33 using the characteristic prediction trained model in association with the prediction score.
[0084] Hereinafter, examples of the experiment information, experiment attribute information, datasets, learning datasets, and prediction target sequence information stored in the storage unit 32 will be described with reference to FIGS.
[0085] FIG. 7 is a diagram showing an example of experiment information according to this embodiment. In the example shown in this figure, the experimental information is a relational database in which the following items are associated with each other for each panning group ID that identifies a panning group: target antigen information, antibody library, dataset, experimental condition ID, round 2 experimental condition ID, and round 3 experimental condition ID. Here, the antibody library is one piece of experimental antibody information and indicates the antibody library provided in round 1. The dataset indicates a dataset file based on analysis result information for panning identified by the panning group ID. The experimental condition ID is identification information that identifies experimental attribute information and indicates the experimental conditions for round 1. The round 2 experimental condition ID and round 3 experimental condition ID are experimental condition IDs that indicate the experimental conditions for round 2 and round 3, respectively. In the example shown in this figure, in the series of pannings with the "Panning Group ID" of "P1," the "target antigen" is "Antigen 1" and the "Antibody Library" is "Library 1." It also shows that, among the dataset files for the series of pannings of "P1," the heavy chain sequence file is "H12345.csv" and the light chain sequence file is "L54321.csv." It also shows that, in the series of pannings of "P1," the experimental conditions for round 1 are "Condition 1," those for round 2 are "Condition 2," and those for round 3 are "Condition 3."
[0086] FIG. 8 is a diagram showing an example of experiment attribute information according to this embodiment. In the example shown in this figure, the experiment attribute information is a database in which the following items are associated with each experimental condition ID: antibody display method, antibody origin, target antigen concentration, buffer composition, reaction time, and reaction temperature. Note that while the database shown is a relational database, the present invention is not limited to this, and may be a text file such as a CSV file or NoSQL (the same applies below). In the example shown in this figure, the experimental condition with "Experimental Condition ID" "P1" indicates that the "Antibody Display Method" is "Phage," the "Antibody Origin" is "Mouse," the "Target Antigen Concentration" is "1 (nM)," the "Buffer Composition" is "Composition A," the "Reaction Time" is "T0," and the "Reaction Temperature" is "t1." As mentioned above, the buffer composition may be expressed as a hydrogen ion exponent.
[0087] FIG. 9 is a diagram showing an example of a data set according to this embodiment. The dataset in this figure is associated with the panning group ID "P1" and has the file name "H12345.csv." In other words, the dataset in this figure was generated from the analysis result information in the panning series "P1," and is a dataset of antibody heavy chain sequences. In the example shown in this figure, the dataset is a database in which the following items are associated with each sequence ID: antibody sequence information (H1, H2, ..., H35a, H35b, H36, ...), appearance frequency in round 1, appearance frequency in round 2, appearance frequency in round 3, rate of change (1 → 2), and rate of change (2 → 3). Here, "sequence ID" refers to an identifier for identifying the antibody sequence.
[0088] "H1," "H2," "H35a," "H35b," and "H36" represent the amino acid positions in the variable region of the antibody heavy chain based on the Kabat numbering system, with "H" indicating a heavy chain. The change rate (N→N+1) indicates the rate of change in appearance frequency between round N and round N+1, and in the example shown in this figure, is the value obtained by dividing the appearance frequency in round N+1 by the appearance frequency in round N. The change rate may also be calculated by the server 30. In the example shown in this figure, an antibody identified by the "sequence ID" "VH001" indicates that, in its amino acid sequence, the amino acid at position "H1" is "M (methionine)," the amino acid at position "H2" is "E (glutamic acid)," the amino acid at position "H35a" is "P (proline)," the amino acid at position "H35b" is "S (serine)," and the amino acid at position "H36" is "Q (glutamine)." Furthermore, the antibody identified by "VH001" indicates, as evaluation result information, that the "occurrence frequency in round 1" is "10," the "occurrence frequency in round 2" is "25," the "occurrence frequency in round 3" is "50," the "change rate (1 → 2)" is "2.50," and the "change rate (2 → 3)" is "2.00."
[0089] FIG. 10 is a diagram showing another example of a data set according to this embodiment. The dataset in this figure is associated with the panning group ID "P1" and has the file name "L54321.csv." In other words, the dataset in this figure was generated from the analysis result information in the panning series "P1," and is a dataset of antibody light chain sequences. In the example shown in this figure, the dataset is a database in which the following items are associated with each sequence ID: antibody sequence information (L1, L2, ..., L27, ...), frequency of appearance in round 1, frequency of appearance in round 2, frequency of appearance in round 3, rate of change (1 → 2), and rate of change (2 → 3).
[0090] "L1," "L2," and "L27" are pre-assigned to amino acid positions in the antibody. Each of these items indicates a position in the variable region of the antibody light chain, and the value (in alphabetical characters in the figure) represents the amino acid located at that position. Note that the data sets in Figures 10 and 9 differ depending on whether the antibody sequence information indicates a position in the variable region of the antibody heavy chain or a position in the variable region of the antibody light chain. In the example shown in this figure, an antibody identified by the "sequence ID" "VL001" indicates that, in its amino acid sequence, the amino acid at position "L1" is "M", the amino acid at position "L2" is "F (phenylalanine)", and the amino acid at position "L27" is "A (alanine)". Furthermore, the antibody identified by "VL001" indicates, as evaluation result information, that the "occurrence frequency in round 1" is "8", the "occurrence frequency in round 2" is "20", the "occurrence frequency in round 3" is "40", the "change rate (1→2)" is "2.50", and the "change rate (2→3)" is "2.00".
[0091] FIG. 11 is a diagram showing an example of a training data set according to this embodiment. The learning data set is stored for each panning loop ID and classification criterion candidate information. The learning data set in this figure is a collection of learning data sets whose panning loop ID is "P1" and whose classification criterion candidate information is "criterion 1." In the example shown in this figure, the dataset is a database in which, for each sequence ID, the following items are associated: antibody sequence information (H1, H2, . . ., H35a, H35b, H36, . . ., L1, L2, . . ., L27, . . .), appearance frequency in round 1, appearance frequency in round 2, appearance frequency in round 3, rate of change (1 → 2), rate of change (2 → 3), and binding determination information. The binding determination information indicates whether the antibody is a binding antibody or a non-binding antibody under "criterion 1." In the example shown in this figure, the antibody identified by the "Predicted target sequence ID" "VHL0001" has, in its amino acid sequence, the amino acid at position "H1" is "M," the amino acid at position "H2" is "E," the amino acid at position "H35a" is "P," the amino acid at position "H35b" is "S," the amino acid at position "H36" is "Q," the amino acid at position "L1" is "M," the amino acid at position "L2" is "F," and the amino acid at position "L27" is "A," indicating that the "binding judgment" is "binding" (binding antibody).
[0092] The server 30 generates a group of virtual sequences having the characteristics of the combined sequences as shown in FIG. 12 using a sequence-trained model that has learned the characteristics of the combined sequences. The binding sequence group is defined as described above. The sequence-trained model is a trained model that has learned which amino acids are likely to appear at which positions in a desired sequence group, and which amino acids are likely to appear at that position depending on the amino acid group preceding that position. The server 30 generates many sequences that are believed to have desirable properties based on the sequence-trained model.
[0093] FIG. 12 is a diagram showing an example of prediction target sequence information according to this embodiment. The prediction target sequence information is information indicating the prediction target sequence. In the example shown in this figure, the prediction target sequence information is a database in which antibody sequence information (H1, H2, ..., H35a, H35b, H36, ..., L1, L2, ..., L27, ...) is associated with each sequence ID. In the example shown in this figure, the antibody identified by the "Prediction Target Sequence ID" "V000001" is the antibody for which the prediction score is calculated, and its amino acid sequence shows that the amino acid at position "H1" is "M", the amino acid at position "H2" is "E", the amino acid at position "H35a" is "D (aspartic acid)", the amino acid at position "H35b" is "S", the amino acid at position "H36" is "R (arginine)", the amino acid at position "L1" is "M", the amino acid at position "L2" is "F", and the amino acid at position "L27" is "A". The example shown in this figure is the predicted target sequence information when, for example, H35a and H36 are input as mutation information. In other words, among multiple predicted antibody sequences, the amino acids differ at the positions indicated by H35a and H36 in the sequence information, but the amino acids are identical at other positions.
[0094] FIG. 13 is a diagram showing an example of characteristic evaluation information according to this embodiment. The characteristic evaluation information is information indicating the result of prediction when predicting the characteristics of a target sequence using the learning result and the target sequence. In the example shown in this figure, the characteristic evaluation information is a database in which a predicted score is associated with each sequence ID. The predicted score is information indicating the probability and strength of binding to the target antigen. In the example shown in this figure, the "prediction score" of the antibody identified by "V000001" as the "prediction target sequence ID" is blank because no prediction has been made yet.
[0095] <Server processing section> Returning to FIG. 6, the processing unit 33 will be described in detail. The processing unit 33 includes an information acquisition unit 331 , an estimation unit 332 , a classification unit 333 , a prediction target sequence generation unit PA, a learning unit 334 , a control unit 335 , and an output processing unit 336 .
[0096] The information acquisition unit 331 acquires experimental information (see FIG. 7) and experimental attribute information (see FIG. 8) from the analysis result information from the next-generation sequencer 20, and stores this information in the experimental information storage unit 321. The information acquisition unit 331 acquires sequence information from the analysis result information from the next-generation sequencer 20. The information acquisition unit 331 calculates the occurrence frequency for each piece of sequence information in the analysis result information, and generates the sequence information and occurrence frequency as a dataset. Here, the information acquisition unit 331 divides the dataset into heavy chain sequences and light chain sequences. Specifically, the information acquisition unit 331 treats the sequence information and evaluation result information indicating the heavy chain sequence as a dataset of antibodies with heavy chain sequences (see FIG. 9), associates the file with experimental information (e.g., a panning group ID), and stores the file in the dataset storage unit 322. The information acquisition unit 331 stores the sequence information indicating the light chain sequence and the evaluation result information as a dataset of the antibody light chain sequence (see FIG. 10) and associates the file with the experimental information in the dataset storage unit 322.
[0097] When the information acquisition unit 331 acquires classification standard information from the user terminal 10, it stores the classification standard information in the classification standard storage unit 323. When the information acquisition unit 331 acquires attention position information from the user terminal 10, it stores the attention position information in the attention position information storage unit 325. When the information acquisition unit 331 acquires mutation information from the user terminal 10, it stores the mutation information in the mutation information storage unit 327.
[0098] The prediction unit 332 predicts a combination of heavy chain and light chain sequences based on the number of rounds and frequency of occurrence in a series of pannings. The prediction unit 332 predicts that an antibody containing the predicted combination of heavy chain and light chain sequences exists. Specifically, the estimation unit 332 calculates, for example, a correlation coefficient of the occurrence frequency for each combination of heavy chain sequence and light chain sequence for each number of rounds. For the combination with the highest correlation coefficient, the estimation unit 332 estimates that an antibody containing the heavy chain sequence and light chain sequence of that combination exists. The estimation unit 332 calculates a correlation coefficient of the occurrence frequency for each number of rounds for each combination of heavy chain sequence and light chain sequence other than the heavy chain sequence and light chain sequence of the combination with the highest correlation coefficient, and repeats the above processing of the estimation unit 332. In this way, the estimation unit 332 estimates the combination of heavy chain sequence and light chain sequence based on the correlation of appearance frequency over multiple rounds in a series of panning. This allows the information processing system 1 to estimate the antibody (combination of heavy chain sequence and light chain sequence) with higher accuracy than when panning is not performed.
[0099] The estimation unit 332 stores the combination of a heavy chain sequence and a light chain sequence (also referred to as a "present antibody sequence") contained in an antibody estimated to exist in the dataset storage unit 322. Note that the estimation unit 332 may, for example, calculate the rate of change in the occurrence rate between rounds or the correlation coefficient of the difference in the occurrence rate instead of the correlation function of the occurrence frequency for each number of rounds. For example, the estimation unit 332 may estimate that an antibody containing a heavy chain sequence and a light chain sequence exists for heavy chain and light chain sequences for which any of the occurrence frequency for each number of rounds, the rate of change in the occurrence rate between rounds, or the correlation coefficient of the difference in the occurrence rate is the same (including approximately the same).
[0100] However, the present invention is not limited to this, and the processing unit 33 may estimate the combination of heavy chain and light chain sequences using another method. Furthermore, the processing unit 33 may not need to estimate the combination, in which case the processing unit 33 may not include the estimation unit 332. For example, the processing unit 33 may analyze only the heavy chain sequence or only the light chain sequence. In this case, the processing unit 33 may store the heavy chain sequence (or the light chain sequence) in the dataset storage unit 322 as a present antibody sequence. Furthermore, when the analysis result information includes sequence information of the entire antibody, the processing unit 33 stores the sequence information of the entire antibody included in the analysis result information as a present antibody sequence in the dataset storage unit 322. For example, when the next-generation sequencer 20 can read a heavy chain sequence and a light chain sequence at the same time, the information acquisition unit 331 acquires sequence information of a combined sequence of the heavy chain sequence and the light chain sequence as the analysis result information and stores it in the dataset storage unit 322.
[0101] The classification unit 333 reads out a plurality of pieces of classification standard candidate information from the classification standard information in the classification standard storage unit 323. The classification unit 333 classifies the antibody indicated by the present antibody sequence into a binding antibody or a non-binding antibody according to the classification standard indicated by each piece of classification standard candidate information. Specifically, the classification unit 333 determines whether the appearance frequency of the present antibody sequence per round (or the rate of change in appearance frequency between rounds) is equal to or greater than the threshold for the appearance frequency (or rate of change) in the classification standard information. If the classification unit 333 determines that the appearance frequency (or rate of change) of the present antibody sequence is equal to or greater than the threshold, it determines that the antibody represented by the present antibody sequence is a binding antibody. On the other hand, in other cases (if the appearance frequency (or rate of change) of the present antibody sequence is lower than the threshold), the classification unit 333 determines that the antibody represented by the present antibody sequence is a non-binding antibody.
[0102] Here, when thresholds are set for multiple items in certain classification standard information (see FIG. 4), if all of those items are determined to be equal to or greater than the threshold, the classification unit 333 determines that the antibody represented by the present antibody sequence is a binding antibody. In the example of FIG. 4, the classification unit 333 determines that the antibody represented by the present antibody sequence is a binding antibody if, in criterion 1, the appearance frequency in round 1 is X1 or greater, the rate of change in appearance frequency between round 1 and round 2 is Y1 or greater, and the rate of change in appearance frequency between round 1 and round 3 is Z1 or greater. In all other cases, the classification unit 333 determines that the antibody represented by the present antibody sequence is a non-binding antibody.
[0103] Note that for each round, the occurrence frequency (or rate of change) of the present antibody sequence is either the occurrence frequency of the heavy chain sequence or the occurrence frequency of the light chain sequence (for example, the occurrence frequency of the heavy chain sequence, the occurrence frequency of the light chain sequence, the minimum occurrence frequency, or the maximum occurrence frequency), but it may also be the average value of the heavy chain sequence and the light chain sequence. Furthermore, the classification unit 333 may add classification criterion candidate information. For example, the classification unit 333 adds new classification criterion candidate information by using a threshold value obtained by varying a threshold value included in other classification criterion candidate information by a predetermined value as the threshold value.
[0104] The classification unit 333 stores, for each classification criterion candidate information, the present antibody sequence determined to be a binding antibody, its evaluation result information (appearance frequency and its rate of change), and binding determination information indicating the classification result as a learning dataset in the learning dataset storage unit 324. The binding determination information indicates whether the antibody is a binding antibody or a non-binding antibody. In other words, the binding determination information is information indicating whether the antibody binds to the target antigen or not (is non-binding). The binding determination information may be a value obtained by subtracting the threshold value of the appearance frequency (or rate of change) of the classification standard information from the appearance frequency of the present antibody sequence per round (or rate of change in appearance frequency between rounds), or may be a value based on variance or standard deviation.
[0105] <Generation of predicted target sequences> The process of generating a prediction target sequence performed by the prediction target sequence generating unit PA will be described below. The prediction target sequence generation unit PA includes a sequence selection unit PA1, a sequence learning unit PA2, and a virtual sequence generation unit PA3. The sequence selection unit PA1 reads out a training data set for each classification criterion candidate information from the training data set storage unit 324. When attention information is stored in the attention position information storage unit 325, the sequence selection unit PA1 reads out the attention information. The sequence selection unit PA1 extracts sequence information at the position indicated by the attention information from the sequence information indicated in the training data set. The sequence learning unit PA2 uses a training data set including the sequence information extracted by the sequence selection unit PA1 (hereinafter also referred to as "target sequence information") and binding judgment information for the training process. Note that when attention information is not stored in the attention position information storage unit 325, the sequence learning unit PA2 sets sequence information at all positions as target sequence information and uses a training data set including the target sequence information and binding judgment information for the training process.
[0106] The sequence selection unit PA1 determines a classification criterion with high accuracy based on the results of the learning process for each classification criterion candidate information, and selects a learning data set for the classification criterion with high accuracy. The learning process in which the sequence selection unit PA1 uses LSTM (Long Short-Term Memory) as a learning model will be described in detail below. However, the present invention is not limited to this, and other learning models may be used in the learning process.
[0107] <Learning process> FIG. 14 is an explanatory diagram illustrating an example of the learning process performed by the sequence learning unit PA2 according to this embodiment. The LSTM used in the learning process consists of three layers: an input layer, an intermediate layer, and an output layer. In the example shown in Figure 14, the input layer is denoted by X1, X2, ... X M , the intermediate layers are A1, A2, ... A M , the output layer is h1, h2, h M Each input of the input layer is an amino acid at a position indicated by the attention position information among the amino acids at each position in the training dataset. Note that the attention position information is, for example, a site in an antibody, and is a consecutive position in the sequence. However, the present invention is not limited to this, and may also include non-consecutive positions. If there is no attention position information, each input of the input layer is an amino acid at all positions in the training dataset.
[0108] The t-th (t=0, 1, 2, . . . , M) hidden layer A t has an input layer X t Input information from the t-1th hidden layer A t-1 The output information from Each middle layer A t A plurality of parameters are stored for each layer. The parameters are, for example, parameters related to the processing of the input gate, input judgment gate, forget gate, and output gate present in the intermediate layer. The parameters are stored in advance in the storage unit 32. Middle Tier A t is the input layer X t Input information from hidden layer A t-1 When the output information from the output layer h is input, the output layer h is t Calculate and output the value of
[0109] When LSTM training is performed, the input layer X t is input with the amino acid information of the t-th sequence in the sequence information. The amino acid information is a vector with 20 components, each corresponding to one of the 20 types of amino acids. For example, the fourth component of the vector corresponds to the amino acid type "E." As shown in the figure, if the amino acid at position H2 in the sequence information is "E," only the fourth component is "1," and the other components are "0." Note that a command ("START" in the figure) is input to the input layer X0 to cause the hidden layer A0 to output the value of the output layer h0. This command also indicates the start of outputting the amino acid sequence. The output layer h output from LSTM t Vector values of h t is compared with the t+1th amino acid information in the sequence information. As a result of the comparison, the hidden layer A t The parameters of the output layer h M The value of is used as information indicating the end of the amino acid sequence (shown as "END" in FIG. 14).
[0110] The example shown in Figure 14 shows the data input when training the sequence identified by the sequence ID "VHL0001" in Figure 11 to the LSTM. The amino acid "M" at position H1 is input to the input layer "X1". The amino acid "E" at position H2 is input to the input layer "X2". M " is input with information ("-") indicating that there is no amino acid at position L107a. Meanwhile, during learning, the value output from output layer "h0" is compared with the amino acid "M" at position H1, and the value output from output layer "h1" is compared with the amino acid "E" at position H2. Through the above learning process, the LSTM after learning has the hidden layer A t to the output layer h t For example, the LSTM is generated for each standard, and the training data set for that standard is used to train the LSTM.
[0111] <Execution process: Virtual array generation> We will explain the execution process of outputting a virtual array using LSTM after the learning process. When this execution process is performed, the output layer of the LSTM, h t-1 Vector values of h t-1 is the input layer X t The LSTM sequentially inputs the next amino acid in the sequence to the output layer h t Here, the output layer h t-1 Vector values of h t-1 is the input layer X t If input to , the value of all components of the vector, h t-1 One amino acid is selected according to the occurrence probability of the 20 types of amino acids, and a vector is input in which the vector component corresponding to that amino acid is "1" and the other components are "0". As an example, when generating a large number of virtual sequences, the process of "selecting an amino acid according to the probability set in the model at each position to generate one sequence" is repeated a large number of times (several million to several tens of millions). In this way, the output layer h t-1 Vector values of h t-1 The vector for which the corresponding amino acid is determined to be one is referred to as the determined vector value h t-1 It is also called. fixed vector value h t-1 The amino acid represented by is the t-th amino acid information (component) in the sequence information. In other words, the determined vector value h t-1 The sequence in which the amino acids represented by are arranged in the order of t=1 to t is output as a virtual sequence. Here, the virtual sequence is a sequence that has the characteristics of the sequence information of the training dataset. In this way, a trained model that learns the characteristics of the sequence information of the training dataset and outputs a virtual sequence is called a sequence trained model.
[0112] <Execution process: Prediction of prediction score> We will now explain the execution process of outputting a prediction score using the LSTM after the learning process. LSTM is trained only with the sequence information of the training dataset whose join decision is "join". When execution processing is performed, the input layer X tFor (t ≧ 1), among the input array information, as the t-th amino acid information, the vector value x t is input. Information indicating the start of output of the amino acid sequence ("START") is input to the input layer X0. The output layer h t outputs the vector value h t .
[0113] The vector value h t represents the predicted value of the (t + 1)-th amino acid information when the t-th amino acid information is input to the input layer X t . Since the LSTM learns only from the array information for which the binding determination is "binding", this predicted value is a predicted value with a high possibility of the binding determination being "binding". Therefore, the inner product of the vector value h t-1 and the vector value x t is information for calculating the likelihood P that the binding determination for the entire array is "binding" in the t-th amino acid information. The probability P t multiplied from t = 1 to t = M, the likelihood P = P1 × P2 × P3 × ··· × P M becomes the prediction score representing the affinity with the target antigen for the input array information. Thus, a learned model that outputs a prediction score for the input array information is called a characteristic prediction learned model.
[0114] Note that in this embodiment, since the information processing system 1 uses the same LSTM for the array learned model and the characteristic prediction learned model, the learning process can be reduced as compared with the case of performing the learning process for each model individually. However, the present invention is not limited to this, and the array learned model or the characteristic prediction learned model may have different learning data sets used for learning, or may have different learning models. The characteristic prediction learned model may be a Gaussian process, and the prediction score by the Gaussian process may be the reliability of the prediction.
[0115] <Intermediate layer of LSTM> FIG. 15 is a conceptual diagram showing the structure of the LSTM according to each of the above embodiments. This figure shows a part of the LSTM in Figure 14, and the t-th hidden layer A t This figure shows an example of the internal structure of the intermediate layer A. t In the middle layer A t-1 As input information, the t-1th cell state C t-1 , and the output layer h t-1 The vector value h output from t-1 is input. In addition, the hidden layer A t has an input layer X t to vector value x t is entered.
[0116] As a parameter of LSTM, W f , b f , W i , b i , W c , b c , W o , b o is stored in the storage unit 32. t In the input C t-1 , value h t-1 , and the value x t In contrast, the parameter W stored in the storage unit 32 f , b f , W i , b i , W c , b c , W o , b o Using this, f in the following equation (1) t , i in Eq. (2) t , C~(tilde) in equation (3) t , o in Eq.(5) t is calculated. C in Equation (4) t is the calculated f t , i t , C~(tilde) t The vector value h t is C t and t This vector value h t is the output layer h t is output from
[0117]
number
[0118] Here, σ represents a sigmoid function. In the learning process, the t-1th output layer h t-1 The vector value h output from t-1 and the t-th vector value x of the array information t The result of the comparison is the vector value h t-1 and vector value x t To reduce the error of f , b f , W i , b i , W c , b c , W o , b o is updated to the new value. The updated parameters are stored in the storage unit 32. f , b f is a parameter related to the processing of the input gate. i , b i is a parameter related to the processing of the input decision gate. c , b c is a parameter related to the forget gate process. o , b o is a parameter related to the processing of the output gate.
[0119] <Selecting a trained model> Returning to FIG. 6, when performing the learning process, the sequence selection unit PA1 first reads out a learning data set for which the combination judgment is "combined" from the learning data set. The sequence selection unit PA1 divides the read learning data set into a learning data set used in the learning process and an evaluation data set used in an evaluation process for evaluating the characteristic prediction learned model. In this embodiment, there are as many characteristic prediction learning models as there are pieces of classification criterion candidate information, but for each characteristic prediction learning model, the learning process is performed using the above-mentioned learning data set, and the evaluation process is performed using the evaluation data set. The sequence selection unit PA1 then divides the learning dataset used in the learning process into a training dataset and a validation dataset. For example, the sequence selection unit PA1 divides the learning dataset into multiple groups (G1, G2, . . . GN). Each group is made to contain approximately the same number of learning datasets. The sequence selection unit PA1 selects one of the multiple groups (e.g., Gk) as the group to be used for validation (validation group). In other words, the sequence selection unit PA1 designates the learning dataset included in the validation group as the validation dataset. The sequence selection unit PA1 sets the learning datasets included in the remaining groups as the training dataset.
[0120] The sequence learning unit PA2 generates a characteristic prediction trained model by performing a learning process on the LSTM using a training dataset. After the learning process, the sequence learning unit PA2 validates the LSTM. The sequence learning unit PA2 inputs the sequence information of the validation dataset into the characteristic prediction trained model. The sequence learning unit PA2 acquires the prediction score output from the characteristic prediction trained model as characteristic estimation information. The sequence learning unit PA2 compares the characteristic estimation information with characteristic information corresponding to the input sequence information, and determines the deviation between the characteristic estimation information and the characteristic information using a predetermined method. The predetermined method is, for example, a method of calculating the mean absolute error for all data in the evaluation dataset. The method for determining the deviation between the characteristic information and the characteristic estimation information is not limited to the above-described method, and may be, for example, a method for determining the mean square error, the root mean square error, the coefficient of determination, or the like.
[0121] The sequence learning unit PA2 changes the validation group and repeats the training and validation described above. The number of repetitions is equal to the number of divided groups. Also, a group that has been classified into a validation group once will not be classified into a validation group again. That is, for example, if the learning dataset is divided into N groups G1 to GN, training and validation are performed N times. Also, each group is included in a validation group once and is used in the validation process described above. After completing N training and validation rounds, the sequence learning unit PA2 evaluates the training using the entire LSTM learning data set using the obtained N deviations. Specifically, the sequence learning unit PA2 calculates the average of the N deviations. If the calculated average is not equal to or less than a predetermined threshold, the sequence learning unit PA2 repeats the above learning process. At this time, the sequence learning unit PA2 changes the parameters for the entire hidden layer. If the calculated average is equal to or less than the predetermined threshold, the sequence learning unit PA2 determines that learning has ended and terminates training and validation.
[0122] The sequence learning unit PA2 may perform training and validation using a method other than the above-described method of permuting validation groups. For example, the sequence learning unit PA2 may not permutate validation groups. Alternatively, the sequence learning unit PA2 may permutate a group so that each group contains one training data set. In this case, the number N of groups described above is equal to the number of training data sets. In this embodiment, what outputs characteristic estimation information for input data is called a characteristic prediction trained model.
[0123] The sequence selection unit PA1 calculates the AUC (Area Under an ROC Curve) for each data set in the evaluation data set from the combination judgment information and characteristic estimation information. The sequence selection unit PA1 performs a learning process and an evaluation process for each of all classification criterion candidate information to calculate the AUC for the characteristic prediction trained model for each of the classification criterion candidate information. The sequence selection unit PA1 stores the classification criterion candidate information (also referred to as "selected classification criterion information") with the highest AUC value and the characteristic prediction trained model for that classification criterion candidate information in the training result storage unit 326, in association with the panning group ID. The sequence training unit PA2 stores at least the LSTM portion (FIG. 14) of the selected trained models generated by the sequence selection unit PA1 in the training result storage unit 326 as a sequence trained model.
[0124] The virtual sequence generation unit PA3 generates a plurality of virtual sequences using the sequence-trained model, and sets a group consisting of the generated plurality of virtual sequences as prediction target sequence information. The sequence information of the virtual sequence is sequence information in which amino acids at one or more positions of the sequence information are changed while having the characteristics of the sequence information of the training dataset associated with the selected classification criterion information.
[0125] Specifically, as a result of learning, the LSTM in Figure 14 has learned the conditional probability of which amino acid occurs at each position and with what probability as a sequence-trained model. The virtual sequence generation unit PA3 generates sequences many times according to that probability. Specifically, when generating an amino acid sequence at position 1, the virtual sequence generation unit PA3 generates amino acids for a new virtual sequence based on the learned probability of occurrence of 20 types of amino acids. At position 2, based on the learned conditional probability of occurrence P(AA2|AA1) of the amino acid (AA), it generates amino acids for a new virtual sequence depending on the amino acid at position 1. At position 3, it generates amino acids depending on the amino acids at positions 1 and 2. Thereafter, it generates amino acids for the next position sequentially based on the following formula determined by learning, to generate the entire length of a new virtual sequence. This new virtual sequence generation is executed a large number of times. For example, the conditional probability of the following formula is probability P T+1 is expressed as
[0126]
number
[0127] Here, when mutation information is stored in the mutation information storage unit 327, the sequence generation unit PA3 fixes and inputs amino acids at positions other than the mutation position (element of sequence information) indicated by the mutation information to the sequence trained model, and generates prediction target sequence information by using the amino acid output from the sequence trained model as the amino acid at the mutation position. Specifically, when the mutation position is the t-th, the sequence generation unit PA3 generates the prediction target sequence information by using the fixed vector value h t-1 By replacing the amino acid sequence with the amino acid sequence represented by the formula (I), predicted target sequence information is generated. This allows the information processing system 1 to generate predicted target sequence information in which only amino acids at positions that are likely to bind and are desired to be mutated have been changed. The sequence generating unit PA3 stores the generated prediction target sequence information in the sequence storage unit 328 (see FIG. 12).
[0128] FIG. 16 is a flowchart showing an example of the operation of the virtual array generation unit PA3 according to this embodiment. (Step S1) The virtual sequence generation unit PA3 inputs information indicating the start of an amino acid sequence to the sequence-trained model, thereby instructing the model to generate a prediction target sequence.
[0129] (Step S2) The virtual sequence generation unit PA3 outputs a vector value h0 as amino acid information from the output layer h0 of the sequence-trained model, and inputs the determined vector value h0 as a vector value x1 to the input layer X1. t-1 The amino acid information is expressed as a vector value h t-1 and outputs the determined vector value h t-1 to the vector value x t as the input layer X t The virtual array generation unit PA3 repeats inputting to the output layer h M When information indicating the end of the amino acid sequence is output from the virtual sequence generator PA3, the process is terminated. M-1 The amino acids represented by are arranged in order, and the arranged sequence is generated as a predicted target sequence.
[0130] (Step S3) The virtual sequence generating unit PA3 stores information indicating the prediction target sequence generated in step S1 in the sequence storage unit 328 as prediction target sequence information. (Step S4) The virtual sequence generation unit PA3 determines whether a termination condition is satisfied. This termination condition is a preset condition that terminates the process of generating a group of prediction target sequences. For example, the termination condition is that the number of prediction target sequences generated in step S2 is equal to or greater than a predetermined number. However, the termination condition may be other conditions, such as when a predetermined number of identical prediction target sequences are generated, or when a predetermined number of similar prediction target sequences are generated. A similar prediction target sequence is, for example, a sequence whose prediction score is equal to or greater than a threshold. Another example of a similar prediction target sequence is a sequence whose mutual distance is shorter than a predetermined value when each prediction target sequence is mapped to a vector space using the above-mentioned Doc2Vec method. If the termination condition is not satisfied (No), the virtual array generation unit PA3 returns to step S1 again to generate a prediction target sequence. On the other hand, if the termination condition is satisfied (Yes), the virtual array generation unit PA3 ends the processing of FIG. Through the above process of FIG. 16, the virtual array generation unit PA3 generates a plurality of prediction target arrays.
[0131] Returning to FIG. 6, the learning unit 334 replicates the selected trained model generated by the sequence selection unit PA1, and stores the replicated model in the learning result storage unit 326 as a characteristic prediction trained model. The learning unit 334 may perform a learning process on the LSTM of FIG. 14 using a learning dataset (FIG. 11) associated with the selection classification criterion information, thereby generating a characteristic prediction trained model. For example, the learning unit 334 may perform the learning process by supervised learning such as deep learning using sequence information and bond determination information. In this case, the learning unit 334 may perform the learning process using a learning dataset for which the bond determination is not "bond", in addition to or in place of a part of the learning dataset for which the bond determination is "bond".
[0132] The control unit 335 uses the characteristic prediction trained model to output a prediction score (likelihood) for the input sequence information. That is, the control unit 335 predicts a prediction score for the antibody of the input sequence information with the target antigen subjected to panning with the panning group ID corresponding to the characteristic prediction trained model. For example, the control unit 335 reads out sequence information to be predicted from the sequence storage unit 328. The control unit 335 inputs the read out sequence information to be predicted as input data into the characteristic prediction trained model and outputs a prediction score. The control unit 335 stores the sequence information to be predicted and the predicted prediction score as prediction target antibody information in the characteristic evaluation information storage unit 329. For example, the control unit 335 stores the prediction score corresponding to the sequence information to be predicted for the prediction target antibody information in FIG. 12.
[0133] The output processing unit 336 outputs the predicted target sequence information of the predicted target antibody information as candidate antibody information according to the prediction score of the predicted target antibody information. The candidate antibody information indicates candidate antibodies with high affinity to the target antigen. Specifically, the output processing unit 336 reads out the prediction target antibody information from the characteristic evaluation information storage unit 329 and rearranges it in descending order of prediction score. The output processing unit 336 generates the prediction target sequence information, rearranged in descending order of prediction score, as candidate antibody information. The output processing unit 336 transmits the generated candidate antibody information to the user terminal 10 via the communication unit 31 and over the network NW. Note that the output processing unit 336 transmits the prediction target antibody information to the user terminal 10, and the user terminal 10 (processing unit 14) may rearrange the received prediction target antibody information in descending order of prediction score and display it on the display unit 15.
[0134] When a target antigen or experimental information (see FIG. 4) is specified in the user terminal 10, the output processing unit 336 selects a panning group ID associated with the specified target antigen or experimental conditions from the target antigens in the experimental information (FIG. 7). When experimental conditions are specified in the user terminal 10, the output processing unit 336 selects an experimental condition ID that satisfies the specified experimental conditions (FIG. 8), and selects a panning group ID associated with the selected experimental condition ID in the experimental information (FIG. 7). The output processing unit 336 extracts the predicted antibody information (FIG. 12) corresponding to the selected panning group ID. The output processing unit 336 sorts the extracted predicted antibody information in descending order of prediction score and transmits it to the user terminal 10.
[0135] <About operation> 17 is a flowchart showing an example of the operation of the server 30 according to this embodiment. This diagram shows the operation of the server 30 in the learning stage (learning process and evaluation process).
[0136] (Step S101) The information acquisition unit 331 acquires various pieces of information from the user terminal 10. The information acquisition unit 331 stores the acquired information in the storage unit 32. After that, the process proceeds to step S102. (Step S102) The information acquisition unit 331 acquires analysis result information from the next-generation sequencer 20. The information acquisition unit 331 stores the analysis result information acquired in step S101 as a dataset in the dataset storage unit 322. Thereafter, the process proceeds to step S103.
[0137] (Step S103) The estimation unit 332 estimates an existing antibody sequence as a combination of a heavy chain sequence and a light chain sequence based on the dataset stored in step S102. The estimation unit 332 stores the estimated existing antibody sequence in the dataset storage unit 322. Then, the process proceeds to step S104. (Step S104) The classification unit 333 classifies the antibodies indicated by the present antibody sequences stored in step S103 into binding antibodies or non-binding antibodies according to the classification criteria indicated by each classification criteria candidate information. The classification unit 333 generates a training dataset including the present antibody sequences and binding determination information indicating the classification results for each classification criteria candidate information, and stores the generated training dataset in the training dataset storage unit 324. Then, proceed to step S105.
[0138] (Step S105) The prediction target sequence generation unit PA performs a learning process for each classification criterion candidate information based on the learning data set stored in step S104 to generate a characteristic prediction trained model. The prediction target sequence generation unit PA performs an evaluation process to evaluate the accuracy of the generated characteristic prediction trained model. Then, the process proceeds to step S106. (Step S106) The prediction target sequence generation unit PA selects a characteristic prediction trained model based on the evaluation result of the evaluation process in step S105 for the characteristic prediction trained models generated in step S105. The prediction target sequence generation unit PA stores the selected trained model and a sequence trained model having the same LSTM in the learning result storage unit 326. Thereafter, the operation in this figure ends.
[0139] 18 is a flowchart showing another example of the operation of the server 30 according to this embodiment. This diagram shows the operation of the server 30 in the execution stage. The execution stage refers to a stage in which the information processing system 1 performs predictions and the like using the selected trained model after training with the training dataset.
[0140] (Step S201) The prediction target sequence generation unit PA reads out the sequence trained model stored in step S106 of Fig. 17, and generates prediction target sequence information using the characteristic prediction trained model. Then, the process proceeds to step S205. (Step S202) The control unit 335 predicts a prediction score for the prediction target sequence generated in step S201, using the characteristic prediction trained model stored in step S106 of Fig. 17. Then, the process proceeds to step S203. (Step S203) The output processing unit 336 outputs the predicted target sequence information of the predicted target antibody information as candidate antibody information according to the prediction score predicted in step S202. The output candidate antibody information is displayed on the user terminal 10. Thereafter, the operation in this figure ends.
[0141] <Summary> As described above, in the information processing system 1, the sequence learning unit PA2 (an example of a "sequence learning unit") performs a learning process (an example of "machine learning") using LSTM based on multiple sequences to generate a sequence-trained model (an example of a "first trained model") that has learned the characteristics of the sequences represented by the sequence information. Here, the multiple sequences used in the learning process are antibody sequences whose binding determination result is "bound." Therefore, the sequence-trained model has learned the characteristics of sequences that are likely to be determined to bind. The virtual sequence generation unit PA3 (an example of a "sequence generation unit") generates predicted target sequence information (an example of "virtual sequence information") representing a predicted target sequence obtained by mutating at least one of the amino acids (an example of a "building block") that constitute the sequence represented by the sequence information of the antigen-binding molecule.
[0142] In this way, the information processing system 1 performs the learning process based on sequences for which the binding determination result is "binding," and therefore can predict sequences or amino acids that are likely to be determined to bind from the sequence-trained model. The multiple sequences used in the training process are sequences of the bound antibodies determined to have bound to the target antigen and sequences of antibodies used in the training process of the selected trained model with the highest AUC value calculated by the sequence selection unit PA1. However, the present invention is not limited to this, and sequences having predetermined characteristics (for example, sequences with characteristic values above or below a threshold) may be used depending on each characteristic and classification standard. Furthermore, the bound antibodies determined to have bound to the target antigen may be bound antigens determined to have bound to the target antigen as a result of panning in a first round, or may be bound antibodies subjected to panning in a second or subsequent round, or may be bound antigens determined to have bound to the target antigen as a result of panning in a second or subsequent round.
[0143] Furthermore, the sequence features that the sequence-trained model learns are features that include the positions of amino acids (an example of a "building block") in the sequence and the context between amino acids. This allows the information processing system 1 to learn the positional and contextual characteristics of the amino acid sequences of the antigen-binding molecules used for learning. In this case, the information processing system 1 can generate prediction target sequence information representing sequences having similar positional and contextual characteristics.
[0144] The virtual sequence generating unit PA3 generates predicted target sequence information by changing at least one amino acid at a site on the set sequence that is composed of one or more amino acids. This allows the information processing system 1 to generate prediction target sequence information in which the set site has been changed. For example, by setting the site to be changed, the user can know the prediction target sequence information in which the site has been changed.
[0145] The site on the designated sequence is contained in the sequence of either the heavy chain variable region, light chain variable region or constant region of the antibody. This allows the information processing system 1 to change a site contained in the sequence of any one of the antibody heavy chain variable region, light chain variable region, or constant region, and generate changed predicted target sequence information. For example, by setting a site contained in the sequence of any one of the antibody heavy chain variable region, light chain variable region, or constant region, the user can find predicted target sequence information with the site changed.
[0146] The sequence information used for learning is selected based on the results of characteristic evaluation of the antigen-binding molecule or protein whose sequence is represented by the sequence information. This allows the information processing system 1 to generate predicted target sequence information that is likely to result in similar property evaluation. For example, by setting a desired property evaluation result, the user can know the predicted target sequence information that will result in that property evaluation.
[0147] Furthermore, in the information processing system 1, the sequence selection unit PA1 (an example of a "sequence learning unit": which may be the learning unit 334) performs a learning process based on sequence information representing sequences including a part or all of each of the sequences of a plurality of antigen-binding molecules and the results of property evaluation of the antigen-binding molecules represented by the sequences, thereby generating a property prediction trained model (a "second trained model": which may be a selection trained model). The control unit 335 (an example of an "estimation unit") executes arithmetic processing of the property prediction trained model to estimate a predicted score of property evaluation for an antigen-binding molecule of a sequence represented by the input sequence information to be predicted. Furthermore, in the information processing system 1, the control unit 335 (an example of an "estimation unit") inputs the prediction target sequence information (an example of "virtual sequence information") generated based on the sequence trained model into the selected trained model, and executes arithmetic processing of the selected trained model to predict (an example of "obtain") a predicted score of property evaluation for an antigen-binding molecule of a sequence represented by the input sequence information to be predicted (an example of a "predicted value of property evaluation"). This allows the information processing system 1 to predict a predicted score for characteristic evaluation (for example, affinity) for each of the generated pieces of sequence information to be predicted.
[0148] Furthermore, the output processing unit 336 (an example of an “output unit”) performs output based on the prediction score and the prediction target sequence information according to the prediction score estimated by the control unit 335. This allows the information processing system 1 to output the predicted target sequences, for example, by giving priority to sequences with higher characteristic evaluation results.
[0149] Furthermore, in the information processing system 1, the virtual sequence generation unit PA3 generates prediction target sequence information by changing the amino acids that constitute the sequence at the mutation position on the sequence set as mutation information. This allows the information processing system 1 to set mutation positions and generate a virtual sequence in which the amino acid at that mutation position has been changed. Therefore, the information processing system 1 can reduce the number of sequence candidates compared to when all positions are mutated. Furthermore, for example, the information processing system 1 can generate a virtual sequence in which the amino acid at a mutation position that is assumed to be important for properties such as binding has been changed. Therefore, the information processing system 1 can generate a sequence in which important mutations have been performed after narrowing down the sequence candidates, thereby efficiently generating a desired sequence.
[0150] Furthermore, in the information processing system 1, the mutation position in the set sequence is included in the sequence of any one of the heavy chain variable region, light chain variable region, or constant region of the antibody. As a result, when a mutation position assumed to be important for properties such as binding is found in, for example, the heavy chain variable region or the light chain variable region of the variable region, the information processing system 1 can generate a virtual sequence in which the amino acid at that mutation position has been changed. Also, when a mutation position assumed to be important for properties such as binding is found in, for example, the constant region, the information processing system 1 can generate a virtual sequence in which the amino acid at that mutation position has been changed.
[0151] Furthermore, in the information processing system 1, the output processing unit 336 (an example of an "output unit") outputs at least one piece of sequence information to be predicted from among the multiple pieces of sequence information to be predicted input to the selected trained model, according to the prediction score. This allows the information processing system 1 to output, for example, the generated prediction target sequence information with a high prediction score with priority. Note that outputting with priority includes outputting only high-priority information, outputting high-priority information at the top, outputting high-priority information in a display mode different from low-priority information, and recommending high-priority information.
[0152] Furthermore, in the information processing system 1, the sequence selection unit PA1 (an example of a "sequence acquisition unit") selects multiple sequences according to sequence information and analysis result information or combination judgment information (an example of "evaluation result information"). The sequence learning unit PA2 generates a sequence-trained model by performing a learning process on the selected multiple sequences according to the order of the sequences. Here, the learning process using the LSTM method is machine learning that takes into account the order of the sequences. This allows the information processing system 1 to generate a virtual sequence of an antigen-binding molecule that takes into account properties that depend on the sequence order.
[0153] Furthermore, in the information processing system 1, the sequence selection unit PA1 selects multiple sequences that have analysis result information or combination determination information values higher than a predetermined value and are determined to be "combined." The sequence learning unit PA2 generates a sequence-trained model by performing a learning process on the selected multiple sequences using the sequences as input and output. Here, the learning process using the LSTM method of this embodiment is machine learning using sequences as input and output. This allows the information processing system 1 to predict and output sequences that are likely to combine. For example, in machine learning using a supervised model, if the output is not a sequence (for example, if the output is a characteristic evaluation value), it is necessary to generate a sequence by performing a calculation to obtain a sequence with a high characteristic evaluation value. In contrast, since the output of the sequence-trained model is a sequence, the information processing system 1 can immediately obtain a sequence that is likely to combine without performing any further calculations.
[0154] Furthermore, in the information processing system 1, the virtual array generation unit PA3 performs machine learning using a deep learning model. Here, the virtual array generation unit PA3 may perform the learning process using any of a recurrent neural network (RNN), a gated recurrent unit (GRU), a generative adversarial network (GAN), a variational autoencoder (VAE), or a flow deep generative model as the deep learning model instead of or in addition to the LSTM.
[0155] Furthermore, in the information processing system 1, the virtual array generation unit PA3 may perform machine learning using a probabilistic model instead of or in addition to the LSTM. Here, the virtual array generation unit PA3 performs machine learning using either a hidden Markov model (HMM) or a Markov model (MM) as the probabilistic model.
[0156] Furthermore, in the information processing system 1, the virtual sequence generation unit PA3 performs machine learning based on sequence information in which constituent units that make up a sequence are expressed as the occurrence probability of amino acids. This allows the information processing system 1 to process the constituent units of sequence information using the probabilities of a plurality of amino acid candidates, thereby providing diversity to the constituent units of sequence information.
[0157] (Second embodiment) A second embodiment of the present invention will now be described with reference to the drawings. In this embodiment, a case where candidate antibody information is output for antigen-binding molecules that bind to two or more different target molecules is described. An antigen-binding molecule that binds to two or more different target molecules means that one antigen-binding domain in the antigen-binding molecule can also bind to a target antigen other than the one target antigen. Antigen-binding molecules that bind to two or more different target molecules can be selected by evaluating the affinity with two or more different target molecules using the above-mentioned library of antigen-binding molecules. By performing affinity evaluation under conditions in which two or more different targets are present, it is possible to select antigen-binding molecules that bind when two or more different target molecules are present. As an alternative to selecting two or more different target molecules, it is also possible to select antigen-binding molecules that bind to two or more different target molecules but do not simultaneously bind to the different target molecules. Non-limiting examples of methods for selecting antigen-binding molecules that do not simultaneously bind to different target molecules include the following. Panning is performed using one target molecule (target molecule A), followed by panning using a target molecule (target molecule B) different from the target antigen, thereby enabling the selection of antigen-binding molecules that bind to target molecules A and B. Subsequently, when panning is performed on target molecule A, target B is added in excess to the panning reaction solution, and antigen-binding molecules that have been observed to have an inhibitory effect on the binding activity of target molecule A can be presumed to be antigen-binding molecules that do not simultaneously bind to target molecules A and B. The different target molecules may be different protein antigens or low-molecular-weight compounds. Furthermore, the selection of antigen-binding molecules that bind to two or more different target molecules is not limited to methods using a library of antigen-binding molecules, and any method can be used as long as it includes a plurality of different antigen-binding molecules.
[0158] The information processing system 1a according to this embodiment performs multiple sets of panning under different experimental conditions, including a series of panning in the presence of small molecules and a series of panning in the absence of small molecules. In this embodiment, the experimental condition information includes information indicating that the concentration of the target antigen is a predetermined concentration and that the buffer solution contains a predetermined concentration of small molecules (whether or not small molecules are present). The predetermined concentration is a concentration within a predetermined value or range. The schematic diagram of the information processing system 1a is the information processing system 1 (FIG. 1) of the first embodiment, in which the server 30 is replaced with a server 30a. The user terminal 10 and the next-generation sequencer 20 have the same configuration as in the first embodiment, so their explanation will be omitted. Hereinafter, the same components as in the first embodiment will be assigned the same reference numerals, and their explanation will be omitted here.
[0159] FIG. 19 is a block diagram showing an example of a server 30a according to the second embodiment. The server 30a includes a communication unit 31, a storage unit 32a, and a processing unit 33. The storage unit 32a differs from the storage unit 32 of the first embodiment (FIG. 6) in that it includes a dataset storage unit 322a and a classification standard storage unit 323a. The basic functions of the dataset storage unit 322a are the same as those of the dataset storage unit 322. The following describes the functions of the dataset storage unit 322a that differ from those of the dataset storage unit 322.
[0160] This embodiment also shows an example in which a series of three panning sets is performed. The three sets are associated with panning group IDs "P1," "P2," and "P3." Here, the experimental conditions for the panning group "P1" are conditions in which the target antigen and small molecule are each present at a predetermined concentration. The experimental conditions for "P2" are conditions in which the target antigen is absent and the small molecule is present at a predetermined concentration. The experimental conditions for "P3" are conditions in which the target antigen is present at a predetermined concentration and the small molecule is absent. In addition to the above conditions, each experimental condition includes conditions such as those shown in FIG. 8, but these are identical or approximately identical between pannings.
[0161] FIG. 20 is a diagram showing an example of a data set according to this embodiment. The dataset in this figure is associated with panning group IDs "P1, P2, P3" and has the file name "H23456.csv." In other words, the dataset in this figure was generated from the analysis result information of three sets of panning ("P1," "P2," "P3"), and is a dataset of antibody heavy chain sequences. In the example shown in this figure, the dataset is a database in which, for each sequence ID, the following items are associated: antibody sequence information, P1·occurrence frequency in round 1, P2·occurrence frequency in round 1, and P3·occurrence frequency in round 1.
[0162] In the example of Figure 20, the "antibody heavy chain sequence information" corresponding to "VH001" indicates that the amino acid at position "H1" is "M", the amino acid at position "H2" is "E", the amino acid at position "H35a" is "P", the amino acid at position "H35b" is "S", and the amino acid at position "H36" is "Q". Furthermore, the "antibody heavy chain evaluation result information" corresponding to "VH001" indicates that the "P1, round 1 occurrence frequency" is "0.516", the "P2, round 1 occurrence frequency" is "0", and the "P3, round 1 occurrence frequency" is "0.001".
[0163] FIG. 21 is a diagram showing another example of a data set according to this embodiment. The dataset in this figure is associated with panning group IDs "P1, P2, P3" and has the file name "L65432.csv." In other words, the dataset in this figure was generated from the analysis result information of three sets of panning ("P1," "P2," "P3"), and is a dataset of antibody light chain sequences. In the example shown in this figure, the dataset is a database in which, for each sequence ID, the following items are associated: antibody sequence information, P1·occurrence frequency in round 1, P2·occurrence frequency in round 1, and P3·occurrence frequency in round 1. The data sets in Figures 20 and 21 differ in that the antibody sequence information indicates positions in the variable region of the antibody heavy chain or positions in the variable region of the antibody light chain.
[0164] In the example of Figure 21, the "sequence information of antibody light chain" corresponding to "VL001" indicates that the amino acid at position "L1" is "M", the amino acid at position "L2" is "F", and the amino acid at position "L27" is "A". Furthermore, the "evaluation result information of antibody light chain" corresponding to "VL001" indicates that the "P1, round 1 appearance frequency" is "0.050", the "P2, round 1 appearance frequency" is "0", and the "P3, round 1 appearance frequency" is "0.01".
[0165] Next, the classification criterion storage unit 323a will be described. Here, the basic functions of the classification criterion storage unit 323a are the same as those of the classification criterion storage unit 323. Below, functions of the classification criterion storage unit 323a that differ from those of the classification criterion storage unit 323 will be described.
[0166] The classification criterion storage unit 323a stores classification criterion information. The classification criterion information includes multiple classification group candidate information. Three thresholds are input for each classification criterion candidate information (corresponding to criteria 1, 2, and 3 in FIG. 4). In this embodiment, the three thresholds are the occurrence frequencies (or occurrence frequency change rates) of three panning groups, and the occurrence frequencies (or change rates) of rounds with the same round number (hereinafter referred to as "round A") can be set. That is, the thresholds can be set to the occurrence frequency of P1·round A (P1A), the occurrence frequency of P2·round A (P2A), the occurrence frequency of P3·round A (P3A), and the change rates of these occurrence frequencies. For example, the criteria can be "the occurrence frequency of P1A is X4 or more, the change rate between the occurrence frequencies of P1A and P2A (the occurrence frequency of P1A / the occurrence frequency of P2A) is Y4 or more, and the change rate between the occurrence frequencies of P1A and P3A (the occurrence frequency of P1A / the occurrence frequency of P3A) is Z4 or more." Furthermore, in this embodiment, when the frequency of appearance of P2 or P3 is used as the threshold, the criteria are such that antibodies are determined to be binding antibodies if their frequency is equal to or less than the threshold. In this case, the criteria are, for example, "the frequency of appearance of P1A is X5 or more, the frequency of appearance of P2A is Y5 or less, and the frequency of appearance of P3A is Z5 or less."
[0167] By setting the classification criteria information as described above, the processing unit 33 in this embodiment can output candidate antibodies for small molecule-dependent antibodies by processing similar to that of the processing unit 33 in the first embodiment.
[0168] <Summary> As described above, in the information processing system 1a according to this embodiment, the sequence learning unit PA2 (an example of a "sequence learning unit") performs a learning process using LSTM (an example of "machine learning") based on sequence information representing a sequence including some or all of the sequences of a plurality of antigen-binding molecules, thereby generating a sequence-trained model (an example of a "first trained model") that has learned the characteristics of the sequence represented by the sequence information. Here, the binding determination information is the sequence of an antibody whose binding determination result is "bound," which is the sequence of a small molecule-dependent antibody. However, the multiple sequences used in the learning process are the occurrence frequency or the rate of change in occurrence frequency between pannings. Furthermore, the multiple sequences used in the learning process may be the sequence of a binding antibody determined to have bound to the target antigen, or the sequence of an antibody used in the learning process of a selected trained model with the highest AUC value calculated by the sequence selection unit PA1, which is the sequence of a small molecule-dependent antibody. The virtual sequence generation unit PA3 (an example of a "sequence generation unit") generates predicted target sequence information representing a predicted target sequence obtained by mutating at least one of the amino acids (an example of a "building block") constituting the sequence represented by the sequence information of the antigen-binding molecule. In this way, the information processing system 1a performs learning processing based on sequences for which the binding judgment result is "bound" when a small molecule is present, and therefore can predict sequences or amino acids that are likely to be judged to bind when a small molecule is present from the sequence-trained model.
[0169] (Third embodiment) A third embodiment of the present invention will now be described with reference to the drawings. In this embodiment, a case will be described in which a hidden Markov model is used as the characteristic prediction learning model.
[0170] The schematic diagram of the information processing system 1b shows the information processing system 1 of the first embodiment, with the server 30 replaced by a server 30b. The user terminal 10 and the next-generation sequencer 20 have the same configuration as in the first embodiment, and therefore their explanations will be omitted. Hereinafter, the same components as in the first embodiment will be assigned the same reference numerals, and their explanations will be omitted.
[0171] FIG. 22 is a block diagram showing an example of a server 30b of an information processing system 1b according to the third embodiment. The server 30b includes a communication unit 31, a storage unit 32b, and a processing unit 33b. The storage unit 32b differs from the storage unit 32 of the first embodiment (FIG. 6) in that it does not include a focus position information storage unit 325 and in that it includes a learning result storage unit 326b. The basic functions of the learning result storage unit 326b are similar to those of the learning result storage unit 326. The learning result storage unit 326b stores a different characteristic prediction learning model than the learning result storage unit 326. The processing unit 33b differs from the processing unit 33 of the first embodiment (FIG. 6) in that it includes a sequence selection unit PA1b, a sequence learning unit PA2b, and a learning unit 334b. The basic functions of the sequence selection unit PA1b, the sequence learning unit PA2b, and the learning unit 334b are similar to those of the sequence selection unit PA1, the sequence learning unit PA2, and the learning unit 334, respectively. Below, the functions of array selection unit PA1b that differ from array selection unit PA1, the functions of array learning unit PA2b that differ from array learning unit PA2, and the functions of learning unit 334b that differ from learning unit 334 will be described.
[0172] 23 is a diagram showing an overview of a characteristic prediction learning model according to this embodiment. Here, an example will be described in which a hidden Markov model is used as the characteristic prediction learning model. In this embodiment, the amino acid sequence of an antibody is considered to be a sequence of individual amino acids.
[0173] 23, each state is represented by a square, a diamond, or a circle. The transition direction between each state is indicated by an arrow. Each arrow is associated with a state transition probability from the source state to the destination state.
[0174] Each state is identified by a predetermined identifier. A state represented by a square indicates a state in which an amino acid exists at a certain position (hereinafter also referred to as an "existing state"). The identifier for the existing state is m. The subscript of the identifier indicates the position number. For example, m1 indicates a state in which an amino acid exists at the first position. Furthermore, m0 indicates the state that indicates the start of a state transition, and m M+1 is a state indicating the end of a state transition. Each existing state is associated with information indicating the occurrence probability of 20 types of amino acids in that state. In the example shown in Figure 23, the information associated with the state is shown below each existing state. A state represented by a diamond indicates the presence of a state in which an amino acid is inserted between a certain position and the next position (hereinafter also referred to as an "insertion state"). The identifier for the insertion state is i. The subscript of the identifier indicates the number of the inserted position. For example, i1 indicates a state in which an amino acid is inserted after the first position. Similarly to the presence state, the insertion state is associated with information indicating the occurrence probability of 20 types of amino acids in each state. A state represented by a circle indicates a state in which an amino acid at a certain position is deleted (hereinafter also referred to as a "deletion state"). The identifier for the deletion state is d. The subscript of the identifier indicates the number of the position where the amino acid is deleted. For example, d1 indicates a state in which the amino acid at the first position is deleted. Unlike the above two states, the deletion state is not associated with information indicating the probability of occurrence of the amino acid.
[0175] The amino acid sequence of an antibody is divided into states m0 and m M+1The state transition route is generated by listing the amino acids that appear (or are inserted) at each position during the state transition (hereinafter, the state transition method is also referred to as the "state transition route"). The state transition route includes information on the order of state transitions and the amino acids that appear in the existing state or the inserted state. When an amino acid sequence is generated by following a certain state transition route, the probability (occurrence probability) that the amino acid will be generated by following that route is calculated. Here, the occurrence probability is the product of all state transition probabilities on the state transition route and the occurrence probability of all amino acids that appear on the state transition route. A certain amino acid sequence is generated through multiple state transition routes, so the occurrence probability of a certain amino acid sequence is calculated as the sum of the occurrence probabilities of multiple state transition routes that can generate the sequence.
[0176] Next, the learning performed by the sequence learning unit PA2b will be described. The sequence selection unit PA1b reads out a learning data set for each classification criterion candidate information from the learning data set storage unit 324. The sequence selection unit PA1b uses, among the learning data sets, a learning data set whose binding determination information is a bound antibody (referred to as a "partial learning data set") for learning performed by the sequence learning unit PA2b. The sequence selection unit PA1b divides the partial learning data set into a learning data set (training data set and validation data set) and an evaluation data set, similarly to the first embodiment. The sequence training unit PA2b uses sequence information from a training data set to train, learning the occurrence probability of amino acids in the existing state or inserted state, and the transition probability between states. In this embodiment, the characteristic estimation information is the occurrence probability of an amino acid sequence. The characteristic estimation information may be a value based on the occurrence probability of an amino acid sequence, a value obtained by performing a predetermined calculation on the occurrence probability, or the like.
[0177] The sequence learning unit PA2b uses a validation data set to verify the learning results. Based on the validation data set and the learning results, the sequence learning unit PA2b inputs the amino acid sequences included in the validation data set into a hidden Markov model and calculates the likelihood of the sequence. For each amino acid sequence included in the evaluation data, the sequence learning unit PA21b calculates the difference in likelihood between the binding sequence group and the non-binding sequence group, and uses a numerical value based on this value as accuracy information.
[0178] The sequence learning unit PA2b changes the verification group and repeats the training and verification described above. This process is the same as the process in the first embodiment, so a description thereof will be omitted here. The sequence learning unit PA2b calculates the average of the accuracy information obtained for each repetition. If the calculated average value is not equal to or less than a predetermined threshold, the sequence learning unit PA2b repeats the above learning. If the calculated average value is equal to or less than a predetermined threshold, the sequence learning unit PA2b stores the learning result in the learning result storage unit 326. Note that the calculated value does not have to be the average. For example, it may be the variance or standard deviation of the accuracy information. In this embodiment, in an example where a hidden Markov model is used as the characteristic prediction learning model, a hidden Markov model that outputs characteristic estimation information for input data as a result of learning is called a characteristic prediction learned model.
[0179] The sequence selection unit PA1b selects an optimal training model from multiple characteristic prediction trained models. Specifically, the sequence selection unit PA1b calculates the AUC (Area Under an ROC Curve) for each data set in the evaluation data set from the binding judgment information and affinity information. The sequence selection unit PA1b performs a training process and an evaluation process for each of all classification criterion candidate information to calculate the AUC for each hidden Markov model of the classification criterion candidate information. The sequence selection unit PA1b stores the classification criterion candidate information with the highest AUC value (also referred to as "selected classification criterion information") and the hidden Markov model (also referred to as "selected trained model") of that classification criterion candidate information in the training result storage unit 326b, in association with the panning group ID.
[0180] The learning unit 334b duplicates the selected trained model generated by the sequence selection unit PA1b and stores it in the learning result storage unit 326b as a characteristic prediction trained model for calculating a prediction score. The learning unit 334 may perform a learning process on the hidden Markov model of FIG. 23 using a training dataset associated with the selected classification criterion information, thereby generating a characteristic prediction trained model for calculating a prediction score. The learning unit 334b may also perform a learning process using a training dataset for which the connection determination is not "connected", in addition to or in place of a part of the training dataset for which the connection determination is "connected".
[0181] <Summary> As described above, in the information processing system 1b according to this embodiment, the sequence learning unit PA2b (an example of a "sequence learning unit") performs a learning process (an example of "machine learning") using a hidden Markov model based on multiple sequences to generate a sequence-trained model (an example of a "first trained model") that has learned the characteristics of the sequence represented by the sequence information. The virtual sequence generation unit PA3 (an example of a "sequence generation unit") generates prediction target sequence information (an example of "virtual sequence information") that represents a prediction target sequence obtained by mutating at least one of the amino acids (an example of a "building block") that constitute the sequence represented by the sequence information of the antigen-binding molecule. In this way, the information processing system 1b performs learning processing based on sequences for which the binding determination result is "binding," and therefore can predict sequences or amino acids that are likely to be determined to bind from the sequence-trained model of the hidden Markov model.
[0182] Furthermore, in the information processing system 1b, the sequence selection unit PA1b (an example of a "sequence learning unit": which may be the learning unit 334b) performs a learning process based on sequence information representing sequences including a part or all of each of the sequences of a plurality of antigen-binding molecules and the results of property evaluation of the antigen-binding molecules represented by the sequences, thereby generating a property prediction trained model (a "second trained model": which may be the selection trained model).
[0183] As described above, the virtual sequence generation unit PA3 (an example of a "sequence generation unit") may use a sequence-trained model of a hidden Markov model to generate predicted target sequence information representing a virtual sequence in which at least one of the amino acids (an example of a "building block") constituting the sequence represented by the sequence information of the antigen-binding molecule has been mutated. In this case, the information processing system 1b can generate a virtual sequence with higher characteristics even in the case of machine learning using a hidden Markov model.
[0184] In addition, in the information processing system 1b, the control unit 335 (an example of an "estimation unit") inputs multiple pieces of prediction target sequence information generated by the virtual sequence generation unit PA3 into a property prediction trained model (a hidden Markov model; an example of a "second trained model") and performs calculation processing of the property prediction trained model to predict (an example of "obtain") a predicted affinity score (an example of a "predicted value of property evaluation") for each of the multiple virtual sequences. This allows the information processing system 1b to predict the predicted affinity score for each of the generated virtual sequences using a hidden Markov model.
[0185] (Fourth embodiment) A fourth embodiment of the present invention will now be described with reference to the drawings. In this embodiment, an example will be described in which multiple characteristics related to antigen-binding molecules are used for characteristic evaluation. Also, in this embodiment, a case will be described in which LSTM is used as the training model for the sequence-trained model ("first trained model"), and Random Forest is used as the training model for the characteristic prediction-trained model ("second trained model": which may also be a selection-trained model).
[0186] A schematic diagram of an information processing system 1c according to this embodiment is the same as that of the information processing system 1 of the first embodiment (see FIG. 1), except that the next-generation sequencer 20 is removed and the user terminal 10 and server 30 are replaced with a user terminal 10c and a server 30c, respectively. Hereinafter, components similar to those of the first embodiment are given the same reference numerals and will not be described again.
[0187] In this embodiment, an example will be described in which multiple characteristics related to antigen-binding molecules are used for characteristic evaluation. In the information processing system 1c, sequence information that satisfies selection conditions set in the user terminal 10c is selected for one or more pieces of characteristic information selected in the user terminal 10c. Here, the selection conditions are, for example, a condition that the characteristics indicated by the characteristic information are good (for example, a characteristic value higher than a threshold), and are conditions that can be set for each piece of characteristic information. The server 30c performs machine learning based on the selected sequence information to generate a sequence trained model (an example of a "first trained model"). This machine learning is performed using the LSTM (FIG. 14) described in the first embodiment as the training model. In this way, the sequence-trained model is trained using sequence information that satisfies the selection conditions set by the user for the characteristic information selected by the user. That is, the server 30 can use a group of sequences that satisfies the selection conditions for each piece of characteristic information as a group of sequences with desirable properties for training. In this case, the server 30 can generate, for example, a large number of virtual sequence groups that are likely to satisfy the selection conditions for each piece of characteristic information. In addition, in this embodiment, a case will be described in which a random forest is used as a learning model of a characteristic prediction trained model (an example of a "second trained model") for predicting a prediction score.
[0188] <Information Processing System> The present embodiment will be described in detail below. In the information processing system 1c, the server 30c performs learning using sequence information indicating the amino acid sequence of an antibody (an example of an antigen-binding molecule) input from the user terminal 10c and characteristic information indicating the characteristics of the antibody. The server 30c transmits output information to the user terminal 10c based on the input information from the user terminal 10c and the learning results.
[0189] For example, the user terminal 10c receives, as antibody characteristic information, characteristic information indicating characteristics related to activity against an antigen and characteristic information indicating characteristics related to the physical properties of the antibody. The user terminal 10c transmits information associating sequence information with characteristic information to the server 30c. The user terminal 10c also receives, as conditions for selecting sequence information to be used for learning, selection conditions for one or more characteristics. The user terminal 10c receives, as input, information (template sequence information, mutation condition information) necessary for generating prediction target sequence information. The user terminal 10c transmits the input information to the server 30c.
[0190] The user terminal 10c associates characteristic information of one or more characteristics with sequence information and transmits the associated information to the server 30c. For example, two pieces of characteristic information may be individually associated with the same sequence information and transmitted. In this case, if the server 30c has previously received sequence information, the server 30c associates the newly received characteristic information with the sequence information that has already been received.
[0191] The template sequence information is information indicating an amino acid sequence that serves as a template when generating prediction target sequence information. As will be described later, in this embodiment, the prediction target sequence is generated, for example, by introducing a mutation into the template amino acid sequence. In this case, the template sequence information indicates a sequence (also referred to as a "template sequence") from which the mutation is introduced. The template sequence is one of the sequences included in the sequence information associated with the characteristic information. The mutation condition information indicates the conditions for introducing mutations into a template sequence. For example, the mutation condition information includes information indicating the upper limit of the number of mutations that can be introduced into a template sequence when generating one of the prediction target sequences.
[0192] The server 30c receives the arrangement information, selection conditions, and characteristic information via the network NW or a storage medium, and stores the received information. The server 30c learns based on the sequence information that satisfies the selection conditions for each piece of characteristic information, and generates and stores a sequence-trained model. The server 30c learns based on the sequence information and the characteristic information for each characteristic, and generates and stores a characteristic-prediction-trained model.
[0193] The server 30c receives and stores information necessary to generate prediction target sequence information via the network NW or a storage medium. The server 30c generates a sequence-trained model by performing machine learning based on the stored sequence information whose characteristic information satisfies the selection conditions. The server 30c generates and stores prediction target sequence information based on the sequence-trained model. Meanwhile, the server 30c generates a trait prediction trained model by performing machine learning based on the stored sequence information and trait information. Based on the trait prediction trained model, the server 30c predicts one or more trait scores (an example of trait evaluation information) for each trait for the input prediction target sequence information. Here, the one or more trait scores are the trait scores of the trait information used in the selection conditions. The server 30c transmits candidate antibody information indicating candidates for antibodies that are expected to bind to the target antigen according to the predicted one or more property scores to the user terminal 10c. The user terminal 10c displays the candidate antibody information according to the received characteristic score.
[0194] This allows the information processing system 1c to present antigen-binding molecule candidates that take into account the properties that will be considered when developing them as pharmaceuticals, compared to when there is no information on the properties of the antigen-binding molecules, thereby enabling the information processing system 1c to present desired antigen-binding molecule information.
[0195] <User device> FIG. 24 is a block diagram showing an example of a user terminal 10c according to this embodiment. The user terminal 10c includes a communication unit 11, an input unit 12, a storage unit 13, a processing unit 14c, and a display unit 15. The basic functions of the processing unit 14c are similar to those of the processing unit 14 (FIG. 3) of the first embodiment. The following describes the functions of the processing unit 14c that differ from the processing unit 14.
[0196] The processing unit 14c transmits various types of information, such as input information (e.g., one or more pieces of characteristic information, sequence information, template sequence information, and mutation condition information) input from the input unit 12, to the server 30c via the communication unit 11c. The server 30c pre-stores correspondence information between the input information and the output information (e.g., a trained model or a table) and generates output information for the input information. The processing unit 14c receives the output information generated by the server 30c via the communication unit 11c. The processing unit 14c displays the received output information on the display unit 15 (an example of output). When the storage unit 13 stores the corresponding information, the processing unit 14c may read out the corresponding information for the input information, generate output information, and cause the display unit 15 to display the output information.
[0197] <server> FIG. 25 is a block diagram showing an example of the server 30c according to this embodiment. The server 30c includes a communication unit 31c, a storage unit 32c, and a processing unit 33c. The basic functions of each are similar to those of the server 30 (FIG. 6) of the first embodiment. Below, functions of the communication unit 31c, the storage unit 32c, and the processing unit 33c that differ from those of the communication unit 31c, the storage unit 32c, and the processing unit 33 will be described.
[0198] The communication unit 31c is a communication module that performs various communications via the network NW. The communication unit 31c performs various communications with, for example, the user terminal 10c.
[0199] <Server storage> The storage unit 32c includes a learning dataset storage unit 324c, a learning result storage unit 326c, a mutation information storage unit 327c, a sequence storage unit 328c, and a characteristic evaluation information storage unit 329c.
[0200] The learning dataset storage unit 324c stores sequence information (see FIG. 26) and characteristic information (see FIGS. 27 and 28). These pieces of information are included in input information from the user terminal 10c and are input by the processing unit 33c. The learning result storage unit 326c stores the sequence learned model and the characteristic prediction learned model as the learning results of the learning unit 334c. The mutation information storage unit 327c stores mutation information. The mutation information is, for example, template sequence information and mutation condition information included in input information from the user terminal 10c, and is input by the processing unit 33c. The mutation information also includes mutation position information indicating mutation positions that is the result of processing by the prediction target sequence generation unit PAc. Details of the processing by the prediction target sequence generation unit PAc will be described later. However, the present invention is not limited to this, and the mutation information may be set in advance in the mutation information storage unit 327c. The characteristic evaluation information storage unit 329c stores characteristic evaluation information (see FIG. 30) in which the prediction score predicted by the processing unit 33c using the characteristic prediction trained model is associated with each prediction target sequence (see FIG. 29).
[0201] Examples of the sequence information, characteristic information, prediction target sequence information, and characteristic evaluation information stored in the storage unit 32c will be described below with reference to FIGS.
[0202] FIG. 26 is a diagram showing an example of sequence information according to this embodiment. In the example shown in this figure, the sequence information corresponds to each item of sequence ID and antibody sequence information (H1, H2, ..., H35a, H35b, H36, ..., H113, L1, L2, ..., L107, L107a). Here, "sequence ID" indicates an identifier for identifying the antibody sequence. "H1," "H2," "H35a," "H35b," "H36," "H113," "L1," "L2," "L107," and "L107a" are pre-assigned to amino acid positions in the antibody.
[0203] In the example shown in this figure, the antibody identified by the "Sequence ID" "S000001" has the following amino acid sequence: amino acid at position "H1" is "M", amino acid at position "H2" is "E", amino acid at position "H35a" is "P", amino acid at position "H35b" is "S", amino acid at position "H36" is "Q", amino acid at position "H113" is "K (lysine)", amino acid at position "L1" is "M", amino acid at position "L2" is "F", amino acid at position "L107" is "I (isoleucine)", and amino acid at position "L107a" is "- (none)".
[0204] FIG. 27 is a diagram showing an example of characteristic information according to this embodiment. In the example shown in this figure, the characteristic information corresponds to each item of sequence ID, KD, expression level, self-polymerization, sensorgram, and structural information. Here, "KD" indicates the dissociation constant of the antibody. "Expression level" indicates the expression level when the antibody is created. "Self-polymerization" indicates the degree of self-polymerization of the antibody. "Sensorgram" indicates data (sensorgram) showing the results of measuring the interaction between the target antigen and the antibody using SPR (surface plasmon resonance).
[0205] In the example shown in this figure, the characteristics of the sequence identified by the "sequence ID" "S000001" are "KD" of "1.00E-08", "expression level" of "3.20E-01", "self-polymerization" of "9.92E-01", and the "sensorgram" is the data shown in "SG000001.jpg". Depending on the sequence, characteristic information may not be present for all characteristic items (KD, expression level, self-polymerization, sensorgram, etc.) For example, in the example shown in this figure, the sequence information identified by "Sequence ID" "S000002" does not have characteristic information for the "sensorgram" item.
[0206] FIG. 28 is a diagram showing an example of a sensorgram according to this embodiment. The sensorgram in FIG. 28 is an example of data shown in the sensorgram item of the characteristic information in FIG. The example shown in this figure shows the results of SPR measurements of the interaction between a target antigen and an antibody under three different time periods. In this figure, the horizontal axis represents elapsed time, and the vertical axis represents binding strength. Time period 1 represents the time period during which the reaction between the antibody and antigen takes place in a neutral reaction solution until binding reaches saturation. Time period 2 represents the time period after binding reaches saturation, during which binding is maintained without changing the solution conditions. Time period 3 represents the time period during which the bound antibody and antigen begin to dissociate after the pH of the reaction solution is changed to acidic.
[0207] FIG. 29 is a diagram showing an example of prediction target sequence information according to this embodiment. In the example shown in this figure, the prediction target sequence information corresponds to each item of the prediction target sequence ID and antibody sequence information (H1, H2, ..., H35a, H35b, H36, ..., H113, L1, L2, ..., L107, L107a). The prediction target sequence ID is an identifier for identifying the prediction target sequence. Furthermore, the antibody sequence information is the same as the antibody sequence information in Figure 26. In the example shown in this figure, the antibody identified by the "Predicted Target Sequence ID" "V000001" has the following amino acid sequence: amino acid at position "H1" is "M", amino acid at position "H2" is "E", amino acid at position "H35a" is "D (aspartic acid)", amino acid at position "H35b" is "S", amino acid at position "H36" is "R (arginine)", amino acid at position "H113" is "K", amino acid at position "L1" is "M", amino acid at position "L2" is "F", amino acid at position "L107" is "I", and amino acid at position "L107a" is "- (none)".
[0208] FIG. 30 is a diagram showing an example of evaluation result information according to this embodiment. In the example shown in this figure, the evaluation result information corresponds to each item of predicted target sequence ID, KD, expression level, self-polymerization, and sensorgram. Here, "KD" indicates the evaluation result of the dissociation constant of the antibody. "Expression level" indicates the evaluation result of the expression level when the antibody was created. "Self-polymerization" indicates the evaluation result of the degree of self-polymerization of the antibody. "Sensorgram" indicates the evaluation result of the sensorgram regarding the interaction between the target antigen and the antibody. In the example shown in this figure, the evaluation results for the sequence identified by "V000001" as the "prediction target sequence ID" show that the "KD" is "1.12E-08", the "expression level" is "8.70E-01", and the "self-polymerization" is "9.87E-01". Furthermore, since the evaluation results for the "sensorgram" have not yet been obtained, each item is blank. The evaluation value of a sensorgram is a value obtained by scoring the graph shape of the sensorgram. For example, the evaluation score is the sum of the values of a term that gives a higher score the higher the maximum value of the graph, and a term that gives a higher score the more rapidly the graph decays after a certain time.
[0209] <Server processing section> Returning to FIG. 25, the processing unit 33c will be described in detail. The processing unit 33c is a processor such as a central processing unit (CPU). The processing unit 33c generates output information for the input information based on, for example, input information input from the communication unit 31c and information stored in the storage unit 32c. The processing unit 33c transmits the generated output information to the user terminal 10c via the communication unit 31c.
[0210] Specifically, the processing unit 33c acquires sequence information indicating the amino acid sequence of the antibody and characteristic information indicating the characteristics of the antibody from the user terminal 10c via the communication unit 31c, and stores the acquired information in the memory unit 32c as a learning dataset.
[0211] Then, the processing unit 33c selects a training dataset having characteristic information of the characteristic selected by the user. The processing unit 33c performs training on the selected training dataset based on the sequence information and the characteristic information of the characteristic, and generates a characteristic prediction trained model. The processing unit 33c stores the generated sequence trained model and characteristic prediction trained model in the storage unit 32c.
[0212] The processing unit 33c receives input information from the user terminal 10c via the communication unit 31c. The input information may be, for example, mutation information. Based on the input information, the information stored in the memory unit 32c, and the learning results, the processing unit 33c generates candidate antibody information representing candidate antibodies that bind to the target antigen according to the degree of binding to the target antigen. The processing unit 33c transmits the generated candidate antibody information to the user terminal 10c via the communication unit 31c.
[0213] <Configuration of the processing unit> In FIG. 25, the processing unit 33c includes an information acquisition unit 331c, a prediction target sequence generation unit PAc, a learning unit 334c, a control unit 335c, and an output processing unit 336c. The information acquiring unit 331c acquires sequence information (see FIG. 26) and characteristic information (see FIG. 27) from the information received from the user terminal 10c, and stores this information in the learning dataset storage unit 324c. The information acquiring unit 331c also acquires mutation information from the information received from the user terminal 10c, and stores it in the mutation information storage unit 327c.
[0214] The prediction target sequence generation unit PAc and the learning unit 334c acquire information associating sequence information with one or more pieces of characteristic information from the learning dataset storage unit 324c. Note that the prediction target sequence generation unit PAc and the learning unit 334c do not acquire information about sequence information for which characteristic information of that characteristic does not exist. The prediction target sequence generation unit PAc and the learning unit 334c use the acquired information as a learning dataset in the learning process for learning sequence characteristics, the learning process for predicting a prediction score, and the evaluation process, respectively. Here, in the learning process for learning sequence characteristics, sequence information that satisfies the selection conditions for each characteristic is used.
[0215] <Learning process for generating prediction target sequences> The prediction target sequence generation unit PAc selects a training data set. As a training process for generating a prediction target sequence (a training process for learning sequence features), the prediction target sequence generation unit PAc performs the training process based on sequence information of the selected training data set using the LSTM described in the first embodiment as a training model. Here, for example, with regard to the training dataset, the user selects one or more characteristics from the characteristics of the characteristic evaluation of the antigen-binding molecule and sets selection conditions for each characteristic. Examples of characteristic evaluation include affinity evaluation, pharmacological activity evaluation, physical property evaluation, kinetic evaluation, and safety evaluation of the original binding molecule. For example, an example of a characteristic for affinity evaluation is binding activity, and an example of a characteristic for pharmacological activity evaluation is pharmacological activity. Examples of characteristics for physical property evaluation include thermal stability, chemical stability, solubility, viscosity, photostability, long-term storage stability, and nonspecific adsorption. The sequence selection unit PA1 selects, from the training dataset, those whose characteristic information satisfies the selection conditions. The prediction target sequence generation unit PAc generates a sequence-trained model by performing a learning process for learning sequence features based on the sequence information of the selected training dataset.
[0216] For example, the user may select one or more characteristics from any one or a combination of affinity evaluation, pharmacological activity evaluation, physical property evaluation, kinetic evaluation, or safety evaluation, and set selection conditions for each characteristic. In this case, the prediction target sequence generation unit PAc selects, from the training dataset, those whose characteristic information satisfies the selection conditions for any one or a combination of affinity evaluation, pharmacological activity evaluation, physical property evaluation, kinetic evaluation, or safety evaluation. The prediction target sequence generation unit PAc generates a sequence-trained model by performing a learning process to learn the characteristics of the sequence based on the sequence information of the selected training dataset.
[0217] [Selection Criteria] The selection conditions that can be set for each characteristic are, for example, the following conditions. When the property is avidity, the selection criterion is high binding strength (e.g., KD (dissociation rate constant) equal to or greater than a threshold value). Binding strength, as described above, is the strength of the sum of non-covalent interactions between one or more binding sites of a molecule (e.g., an antibody) and the molecule's binding partner (e.g., an antigen).
[0218] When the property is stability, such as thermal stability or chemical stability, of an antigen-binding molecule, the selection criterion is high stability (above a threshold). Stability varies depending on the evaluation method used to measure stability, but can be determined using the denaturation midpoint (Tm), which is an indicator of thermal stability; a high Tm is presumed to indicate high thermal stability. Furthermore, stability can be evaluated by measuring the degradation, chemical modification, and aggregation of the antigen-antibody molecule before and after treatments such as heat treatment, exposure to a low pH environment, light exposure, mechanical stirring, and long-term storage, which are the objectives of stability evaluation. Low degradation, chemical modification, and aggregation of the antigen-binding molecule are presumed to indicate high stability. When the property is nonspecific binding evaluation based on binding to extra cellular matrix (ECM), the selection criterion is low binding strength to ECM (below a threshold). Low binding strength to ECM is presumed to indicate low nonspecific binding. The protein expression level can be measured by introducing a gene encoding the antigen-binding molecule into expressing cells, culturing the expressing cells for a certain period of time, and then measuring the concentration of the antigen-binding molecule in the culture supernatant. A high concentration of the antigen-binding molecule in the culture supernatant is presumed to be a "high expression level." In this case, the selection criterion is that the concentration be high (above a threshold value).
[0219] <Learning process for predicting prediction scores> The learning unit 334c performs a learning process for predicting a predicted score, using the sequence information of the learning data set as an input variable and the feature information as an output variable. Specifically, the learning unit 334c selects a training dataset having a value of the characteristic information of the characteristic selected by the user. The learning unit 334c learns based on the sequence information and characteristic information of the selected training dataset, and generates a characteristic prediction trained model. Note that the sequence information used to generate the characteristic prediction trained model includes sequence information in which each characteristic does not satisfy the selection conditions. This allows the characteristic prediction trained model to accurately predict prediction scores even when each characteristic does not satisfy the selection conditions and the characteristic is poor (e.g., low characteristic value). However, the sequence information used to generate the characteristic prediction trained model does not need to include sequence information in which each characteristic does not satisfy the selection conditions.
[0220] The prediction target sequence generation unit PAc performs machine learning using LSTM as the learning model of the sequence trained model. Meanwhile, the learning unit 334c performs machine learning using Random Forest as the learning model of the characteristic prediction trained model. In this way, the learning models used by the prediction target sequence generation unit PAc and the learning unit 334c may be of different types.
[0221] The learning process and evaluation process in which the learning unit 334c uses a random forest as a learning model for the characteristic prediction trained model will be described in detail below. However, the present invention is not limited to this, and other learning models may be used as the characteristic prediction learning model for the learning process and evaluation process.
[0222] <Learning and evaluation processes> Fig. 31 is an explanatory diagram illustrating an example of the learning process according to this embodiment. Fig. 31 shows an example of the learning process for generating a characteristic prediction trained model, and illustrates an example of the learning process when characteristic information can be expressed numerically. In FIG. 31, the training data set is indicated by hatched circles. The learning unit 334c extracts a predetermined number (e.g., 100) of training data sets (hereinafter also referred to as "subsets") from the training data set DS, for example, randomly. The learning unit 334c repeats this extraction a predetermined number of times (e.g., K times) to generate K subsets SS1 to SSK. The learning unit 334c generates a decision tree Tree (k=1 to K) for each subset SSk (k=1 to K). Here, the learning unit 334c sets each piece of target sequence information (sequence elements), i.e., information on the amino acids at each position, as independent variables. The learning unit 334c sets the characteristic information as dependent variables. The learning unit 334c generates a decision tree using a subset of the set independent and dependent variables (an example of "learning").
[0223] In Figure 31, a decision tree is made up of multiple nodes (white circles) and edges (arrows) that connect the nodes. When two nodes are connected by an arrow, the node at the start of the arrow is called the parent node, and the node at the end of the arrow is called the child node. Each node has at most one parent node. A node that does not have a parent node is called a root node. In the execution stage, input data to be predicted (antibody sequence information to be predicted) is classified according to the direction of the arrows from the root node to any node that does not have a child node (hereinafter also referred to as a "leaf node"). Which arrow each piece of input data passes through depends on the criterion associated with each node. The criterion is associated with the independent variable, that is, the amino acid information at each position, as a result of learning. For example, the criterion for a certain node is the criterion for the information of the amino acid at the position "H95a" in the sequence information; if this amino acid is L (leucine), proceed to the right arrow, and if I (isoleucine), proceed to the left arrow.
[0224] Since a leaf node does not have a next child node, no judgment criteria are associated with it. Each leaf node is associated with characteristic estimation information for the antibody indicated by the input data that has passed through the node. The characteristic estimation information is the result of estimation of characteristic information for the input data that has reached each leaf node. In each decision tree, the characteristic estimation information associated with the leaf node is determined by the subset SSk. The characteristic estimation information is calculated based on the training data set classified into each leaf node in the learning stage. For example, the characteristic estimation information is a value based on the statistical value of the characteristic information included in the training data set that reaches each leaf node. The characteristic estimation information is, for example, the average value of the characteristic information. Note that the characteristic estimation information may be the maximum value, minimum value, etc. of the characteristic information, or a value calculated based on the average value, maximum value, minimum value, standard deviation, etc.
[0225] Next, the learning unit 334c performs an evaluation process on the generated decision trees Tree1 to TreeK. Fig. 32 is an explanatory diagram illustrating an example of the evaluation process according to this embodiment. Similar to Fig. 31, Fig. 32 shows an example of the learning process when the characteristic information can be expressed numerically. The learning unit 334c selects one or more learning data sets (also referred to as "evaluation data sets") that are not included in any of the subsets SS1 to SSK from the learning data set DS. Fig. 32 is a diagram showing the case where two evaluation data sets TD are selected.
[0226] The learning unit 334c inputs the sequence information of each data set in the evaluation data set TD into each of the decision trees Tree1 to TreeK and acquires K pieces of characteristic estimation information T1 to TK. The learning unit 334c calculates a representative value based on the characteristic estimation information T1 to TK as characteristic evaluation information. The representative value is, for example, the average value of the characteristic estimation information T1 to TK, but the present invention is not limited to this and may be a maximum or minimum value. The learning unit 334c compares the characteristic information and the characteristic estimation information for each data set in the evaluation data set TD and determines the deviation between the characteristic information and the characteristic estimation information using a predetermined method. An example of the predetermined method is a method of calculating the mean absolute error for all data in the evaluation data set TD. The learning unit 334c determines whether the deviation between the characteristic information and the characteristic estimation information determined using the predetermined method falls within a predetermined range. If the deviation falls within the predetermined range, the learning unit 334c determines that learning has ended and terminates the learning process and the evaluation process. If the deviation does not fall within a predetermined range, the learning unit 334c performs the learning again.
[0227] The method for determining the deviation between the characteristic information and the characteristic estimation information is not limited to the above-described method, and may be, for example, a method for determining the mean square error, the root mean square error, the coefficient of determination, or the like. In this embodiment, when characteristic information can be expressed numerically, a set of decision trees Tree1 to TreeK that outputs characteristic evaluation information based on characteristic estimation information T1 to TK for input data is called a trained model.
[0228] In the learning for generating a trained characteristic prediction model, when the characteristic information is an image, the learning unit 334c extracts feature amounts of the image according to a method previously stored in the storage unit 32c, and performs the above-described processing using the extracted feature amounts as characteristic information. Criteria for determining whether the extracted feature amounts are superior are also previously stored in the storage unit 32c. The feature amounts are extracted and the criteria are determined using, for example, a known machine learning method. Furthermore, when the image is a graph, the slope of the graph, the coefficient of the function that indicates the graph's approximate curve, etc. are calculated as feature quantities. For example, in the case of a graph such as that shown in FIG. 28, the approximate curve for each time period is calculated, and the coefficient of the function that indicates the approximate curve is obtained as feature quantities. For example, in the example shown in this figure, the slope of the curve when approximated to a straight line is used as the feature quantity for time period 1 and time period 2. Furthermore, for time period 3, the half-life of the curve is used as the feature quantity by approximating it to an exponential function.
[0229] After learning, the prediction target sequence generating unit PAc and the learning unit 334c store the sequence trained model and the characteristic prediction trained model in the learning result storage unit 326c, respectively.
[0230] <Generation of predicted target sequences> The process of generating a prediction target sequence performed by the prediction target sequence generating unit PAc will be described below. The prediction target sequence generating unit PAc includes a sequence selecting unit PA1c, a sequence learning unit PA2c, and a virtual sequence generating unit PA3c. The sequence selection unit PA1c acquires sequence information of an antibody that exhibits superior characteristics to an antibody of a template sequence based on the information stored in the training dataset storage unit 324c and the template sequence information stored in the mutation information storage unit 327c during training to generate a sequence-trained model. For example, the sequence selection unit PA1c reads template sequence information from the mutation information storage unit 327c. The sequence selection unit PA1c also selects sequence information that matches the template sequence information from the training dataset storage unit 324c. The sequence selection unit PA1c acquires characteristic information (hereinafter also referred to as "reference characteristic information") associated with the selected sequence information. As a selection condition, the sequence selection unit PA1c selects characteristic information (hereinafter also referred to as "improved characteristic information") that exhibits superior characteristics to the reference characteristic information from the characteristic information stored in the training dataset storage unit 324c. Here, "exhibiting superior characteristics" means that the characteristic information is superior to the reference characteristic information for a specified characteristic item. The criteria for determining whether or not the characteristic information is superior follow the criteria previously stored in the storage unit 32c. The sequence selection unit PA1c stores the sequence information corresponding to the improved property information in the mutation information storage unit 327c. However, in this example, the sequence selection unit PA1c selects a group of sequences with characteristic values better than the template, but it may also select a group of sequences with characteristic values α times better than the template, or the top N (N is a natural number) sequence(s) with characteristic values, and use these as sequence information corresponding to the improved characteristic information.
[0231] The sequence learning unit PA2c acquires sequence information corresponding to the improved characteristic information from the mutation information storage unit 327c. The sequence learning unit PA2c performs machine learning using the acquired sequence information. The learning process is the same as in the first embodiment, so a description thereof will be omitted here. As a result of the machine learning, the sequence learning unit PA2c generates a sequence-trained model. The virtual array generation unit PA3c generates a prediction target sequence using the sequence-trained model generated by the sequence learning unit PA2c. The method for generating a prediction target sequence is also similar to that in the first embodiment, and therefore description thereof will be omitted here (see FIG. 16). The virtual array generation unit PA3c stores the generated prediction target sequence information in the sequence storage unit 328c.
[0232] <Prediction score prediction> The control unit 335c reads out the prediction target sequence information from the sequence storage unit 328c. The control unit 335c inputs the read out prediction target sequence information as input data into the characteristic prediction trained model and outputs a prediction score. The control unit 335c stores the prediction target sequence information and the predicted prediction score as characteristic evaluation information in the characteristic evaluation information storage unit 329c. For example, the control unit 335c stores the prediction score corresponding to the prediction target sequence information for the prediction target sequence information in FIG. 29. Note that the characteristic prediction trained model may be a Gaussian process, and the prediction score based on the Gaussian process may be the reliability of the prediction. The control unit 335c can also obtain characteristic evaluation information for characteristics that are estimated using the output characteristics but are not included in the learning process. Such characteristics include, for example, information indicating the viscosity of an antibody or the possibility of humanizing the antibody. The method for estimating such characteristic evaluation information is predetermined and stored in the storage unit 32c.
[0233] The output processing unit 336c outputs the sequence information of the prediction target as candidate antibody information according to the prediction score of the sequence information of the prediction target. The candidate antibody information indicates candidate antibodies with high characteristics. Specifically, for example, the output processing unit 336c reads out the characteristic evaluation information from the characteristic evaluation information storage unit 329c and ranks the information in descending order of prediction score. Since there are multiple prediction scores for each characteristic, there are multiple ranking results for each characteristic. The output processing unit 336c aggregates the results of the ranking for each characteristic of the sequence information to be predicted and ranks all the characteristics. The ranking for all the characteristics is performed, for example, based on the average value of the ranking for each characteristic. Note that the method for ranking all the characteristics is not limited to the method described above. For example, the sum of the rankings for each characteristic or the sum of predetermined scores corresponding to the rankings may be used. Furthermore, weighting for each characteristic may be performed to determine candidate antibody information. In other words, by weighting a characteristic to be emphasized, sequences with excellent characteristics are more likely to be determined as candidate antibody information.
[0234] The output processing unit 336c generates the prediction target sequence information, sorted in descending order of overall characteristics, as candidate antibody information. The output processing unit 336c transmits the generated candidate antibody information to the user terminal 10c via the communication unit 31c and the network NW. The output processing unit 336c transmits the prediction target antibody information to the user terminal 10c, and the user terminal 10c (processing unit 14c) may rearrange the received prediction target antibody information using the method described above and display it on the display unit 15.
[0235] <About operation> 33 is a flowchart showing an example of the operation of the server 30c according to this embodiment. This diagram shows the operation of the server 30c in the learning stage (learning process and evaluation process).
[0236] (Step S301) The information acquisition unit 331c acquires various pieces of information from the user terminal 10c. The information acquisition unit 331c stores the acquired information in the storage unit 32c. After that, the process proceeds to step S311. The processes of steps S311 to S313 are performed for each selection condition. These processes constitute the learning process S31 for generating a sequence to be predicted. (Step S311) The sequence selection unit PA1c selects sequence information corresponding to the improvement characteristic information as sequence information for the training data sets that satisfy the selection conditions from among the training data sets stored in step S301. Then, the process proceeds to step S312.
[0237] (Step S312) The sequence learning unit PA2c performs a learning process to generate a prediction target sequence using the sequence information selected in step S311. Then, the process proceeds to step S313. (Step S313) The sequence learning unit PA2c stores the sequence learned model generated by the learning process of step S312 in the learning result storage unit 326c. Then, the process proceeds to step S321. The processes of steps S321 to S323 are performed for each of one or more characteristics. These processes constitute the learning process S32 for predicting a predicted score.
[0238] (Step S321) The learning unit 334c selects the sequence information and characteristic information of a learning data set having one or more characteristic values (characteristics of the improvement characteristic information) of the selection conditions from the learning data sets stored in step S301. Then, the process proceeds to step S322. (Step S322) The learning unit 334c performs a learning process to predict a predicted score for each of one or more characteristics using the learning data set selected in step S321. Then, the process proceeds to step S323. (Step S323) The learning unit 334c stores the characteristic prediction trained model generated by the learning process in step S322 in the learning result storage unit 326c. Thereafter, the operation in this figure ends.
[0239] 34 is a flowchart showing another example of the operation of the server 30c according to this embodiment. This diagram shows the operation of the server 30 in the execution stage. The execution stage is a stage in which the information processing system 1c, after learning with a learning dataset, generates a sequence to be predicted using a sequence-trained model and predicts a prediction score using a characteristic prediction-trained model.
[0240] (Step S401) The virtual sequence generation unit PA3c generates prediction target sequence information using the sequence-trained model generated in step S313 of Fig. 33. The prediction target sequence generation unit PA3c stores the generated prediction target sequence information in the sequence storage unit 328c. Then, the process proceeds to step S402.
[0241] (Step S402) The control unit 335c predicts a prediction score for the prediction target sequence information generated in step S401, using the characteristic prediction trained model generated in step S323 of Figure 33. Then, the process proceeds to step S403. (Step S403) The output processing unit 336c ranks all of the characteristics according to one or more prediction scores predicted in step S402. Based on the ranking of all of the characteristics, the output processing unit 336c outputs the prediction target sequence information as candidate antibody information. The output candidate antibody information is displayed on the user terminal 10c. Thereafter, the operation of this figure ends.
[0242] As described above, in the information processing system 1c according to this embodiment, the array learning unit PA2c (an example of an "array learning unit") performs a learning process using LSTM (an example of "machine learning") based on multiple arrays to generate an array-trained model (an example of a "first trained model") that has learned the characteristics of the array represented by the array information. Here, the multiple arrays used in the learning process are arrays that satisfy the selection conditions set for each piece of characteristic information selected by the user. The virtual sequence generation unit PA3c generates predicted target sequence information (an example of "virtual sequence information") representing a virtual sequence in which at least one of the amino acids (an example of a "building block") constituting the sequence represented by the sequence information of the antigen-binding molecule is mutated. This allows the information processing system 1c to generate a virtual arrangement that is highly likely to satisfy the set selection conditions for each piece of selected characteristic information.
[0243] Furthermore, in the information processing system 1c, the control unit 335 (an example of an "estimation unit") inputs the prediction target sequence information generated based on the sequence-trained model into a selected trained model (a decision tree; an example of a "second trained model") and performs arithmetic processing of the selected trained model to predict (an example of "obtain") a predicted score of characteristic evaluation (an example of a "predicted value of characteristic evaluation") for the antigen-binding molecule of the sequence represented by the input prediction target sequence information. As a result, the information processing system 1c can use the decision tree to predict the predicted score of the characteristic evaluation (for example, affinity) for each of the generated pieces of sequence information to be predicted.
[0244] (Fifth embodiment) Hereinafter, a fifth embodiment of the present invention will be described with reference to the drawings. In this embodiment, the information processing system 1d actually measures the characteristics of the antibody of the predicted prediction target sequence (e.g., the antibody indicated by the candidate antibody information). The information processing system 1d performs further machine learning on the sequence trained model or the characteristic prediction trained model based on sequence information representing the antibody sequence and characteristic information indicating the measured characteristics. The information processing system 1d generates new prediction target sequence information using the sequence trained model that has undergone further machine learning. In this embodiment, the information processing system 1d repeats this series of processes. That is, in this embodiment, the information processing system 1d repeats a cycle consisting of learning based on a training dataset, generating a prediction target sequence based on the learning result, measuring using an antibody indicated by the candidate sequence information, and adding the measurement result and sequence information to the training dataset.
[0245] The schematic diagram of an information processing system 1d according to this embodiment is the same as that of the information processing system 1c according to the fourth embodiment, except that the server 30c is replaced with a server 30d. The user terminal 10d has the same configuration as that of the fourth embodiment, and therefore its description will be omitted. Hereinafter, the same components as those of the fourth embodiment will be assigned the same reference numerals, and their description will be omitted.
[0246] FIG. 35 is a block diagram showing an example of a server 30d according to the fifth embodiment. The server 30d includes a communication unit 31c, a storage unit 32d, and a processing unit 33d.
[0247] The storage unit 32d is different from the storage unit 32c of the fourth embodiment (FIG. 25) in that a training data set storage unit 324d is different from the training data set storage unit 324c. Here, the basic functions of the training data set storage unit 324d are the same as those of the training data set storage unit 324c. Below, the functions of the training data set storage unit 324d that differ from those of the training data set storage unit 324c will be described.
[0248] The learning dataset storage unit 324d stores sequence information and characteristic information. These pieces of information are included in input information from the user terminal 10c and input by the processing unit 33d. Here, the sequence information includes cycle number information indicating the cycle number at which the prediction target sequence generation unit PAd made a prediction.
[0249] <Configuration of the processing unit> Compared to the processing unit 33c of the fourth embodiment (FIG. 25), the processing unit 33d of FIG. 35 differs from the learning unit 334d, prediction target sequence generation unit PAd, and control unit 335d in that the learning unit 334d, prediction target sequence generation unit PAd, and control unit 335d are different from the learning unit 334c, prediction target sequence generation unit PAc, and control unit 335c, respectively. Here, the basic functions of the learning unit 334d, prediction target sequence generation unit PAd, and control unit 335d are similar to those of the learning unit 334c, prediction target sequence generation unit PAc, and control unit 335c, respectively. Below, the functions of the learning unit 334d, prediction target sequence generation unit PAd, and control unit 335d that differ from the learning unit 334c, prediction target sequence generation unit PAc, and control unit 335c, respectively, will be described.
[0250] The prediction target sequence generation unit PAd and the learning unit 334d acquire sequence information and one or more pieces of characteristic information from the learning dataset storage unit 324d. The prediction target sequence generation unit PAd and the learning unit 334d do not acquire sequence information for which characteristic information of the corresponding characteristic does not exist. The prediction target sequence generation unit PAd and the learning unit 334d use a portion of the acquired information as a learning dataset for the learning process for learning sequence characteristics, the learning process for predicting a prediction score, and the evaluation process, respectively. Here, in the learning process for learning sequence characteristics, sequence information that satisfies the selection conditions for each characteristic is used.
[0251] <Operation of information processing system 1d> Fig. 36 is a flowchart showing an example of the operation of the information processing system 1d according to this embodiment. In Fig. 36, the same processes as those in Figs. 33 and 34 are denoted by the same reference numerals. (Step S301) The information acquisition unit 331c acquires various pieces of information from the user terminal 10c. The information acquisition unit 331c stores the acquired information in the storage unit 32d. Then, the process proceeds to step S31. (Step S31) The sequence selection unit PA1d and the sequence selection unit PA1d generate a sequence-trained model by performing the learning process S31 (the process of steps S311 to S313 performed for each selection condition) for generating a prediction target sequence in Fig. 33. Then, the process proceeds to step S32. (Step S32) The learning unit 334d generates a characteristic prediction trained model by performing the learning process S32 (processing steps S321 to S323 performed for each of one or more characteristics) for predicting the predicted score in Fig. 33. Then, the process proceeds to step S401.
[0252] (Step S401) The virtual array generation unit PA3d generates prediction target sequence information using the sequence-trained model generated in step S31. The virtual array generation unit PA3d stores the generated prediction target sequence information in the array storage unit 328c. Then, the process proceeds to step S402. (Step S402) The control unit 335d predicts a prediction score for the prediction target sequence information generated in step S401, using the characteristic prediction trained model generated in step S323 of Figure 33. Then, the process proceeds to step S403. (Step S403) The output processing unit 336c ranks all the characteristics according to one or more prediction scores predicted in step S402. Based on the ranking of all the characteristics, the output processing unit 336c outputs the prediction target sequence information as candidate antibody information. The output candidate antibody information is displayed on the user terminal 10c.
[0253] (Step S501) The control unit 335d determines whether or not to perform additional property evaluation, for example, in response to input from the user terminal 10c. In the additional property evaluation, panning of multiple antibodies and target antigens is further performed, and the analysis result information is output from the next-generation sequencer 20. In the additional property evaluation, an antibody library containing antibody candidates indicated by the candidate antibody information is preferably used. If it is determined that additional property evaluation is to be performed (Yes), proceed to step S502; if it is determined that additional property evaluation is not to be performed (No), end the operation of this figure. (Step S502) Additional characteristic evaluation is performed, and as a result, the next-generation sequencer 20 outputs analysis result information. This analysis result information preferably includes, as antibodies, the antibody candidates indicated by the candidate antibody information. Then, proceed to step S503.
[0254] (Step S503) The information acquisition unit 331d acquires various information from the user terminal 10c. The information acquisition unit 331c adds the acquired information to the information stored in step S301 or the previous step S503 and stores it in the storage unit 32d. This information includes sequence information and one or more pieces of characteristic information as the analysis result information output in step S502. Then, the process proceeds to step S504. (Step S504) The control unit 335d determines whether to perform additional learning processing, for example, in response to an input from the user terminal 10c. If it is determined that additional learning processing is to be performed (Yes), the process returns to step S31, and if it is determined that additional learning processing is not to be performed (No), the operation in this figure ends.
[0255] Here, the sequence selection unit PA1d and the learning unit 334d each select information to be included in the learning data set from the acquired information as follows: The sequence selection unit PA1d and the learning unit 334d acquire the upper limit number of learning data sets from the storage unit 32d. The upper limit number is, for example, the number of learning data sets used for learning in the first cycle. The upper limit number is predetermined for each learning process (learning process for learning sequence features, learning process for predicting prediction scores), that is, for each trained model (sequence trained model, characteristic prediction trained model), and is stored in the storage unit 32d. The sequence selection unit PA1d and the learning unit 334d perform training using training data sets from at least two different cycles, with the number of training data sets being equal to or less than the upper limit. In other words, the ratio of the training data set from the previous cycle to the total training data set is reduced. This allows the information processing system 1d to gradually reduce the influence of the characteristic evaluation from the previous cycle while reflecting the characteristic evaluation from the most recent panning. In this case, the information processing system 1d may be able to converge the prediction target sequence and prediction score output from the trained model without causing significant divergence.
[0256] For example, the sequence selection unit PA1d and the learning unit 334d each refer to sequence information and preferentially acquire, as the training data set, sequences generated in a cycle closest to the current cycle. When including all sequences generated in a certain cycle (say, the Mth cycle) as a training data set would exceed an upper limit, the sequence selection unit PA1d and the learning unit 334d each select sequences to include in the training data set and sequences to exclude. The sequence selection unit PA1d and the learning unit 334d each rank the sequences generated in the Mth cycle, for example, based on characteristic information corresponding to the sequences. The ranking method may be, for example, to rank each sequence by characteristic and then calculate the average of the ranks. Note that the ranking method is not limited to this. The sequence selection unit PA1d and the learning unit 334d may select a training data set for each cycle using the same method, or one may use a training data set selected by the other. The sequence selection unit PA1d and the learning unit 334d each acquire, based on the ranking results, sequences with superior characteristics as sequences to be included in the learning data set. The learning unit 334d performs the above process until the number of learning data sets reaches the upper limit.
[0257] The sequence learning unit PA2d and the learning unit 334d each perform a learning process based on the acquired learning data set. The learning process is the same as the method described in the fourth embodiment, so a description thereof will be omitted here. The learning unit 334d stores the learning results in the learning result storage unit 326c. The selection conditions may be different for each cycle, and the user may be able to select characteristics for each cycle and set the selection conditions for each selected characteristic. For example, the selection conditions for a later cycle may be stricter (e.g., higher characteristic values) or looser (e.g., lower characteristic values) than the selection conditions for a previous cycle. Furthermore, the characteristics of the selection conditions for a later cycle may differ in part or in whole from the characteristics of the selection conditions for a previous cycle.
[0258] The prediction target sequence generation unit PAd generates prediction target sequence information. At this time, the prediction target sequence generation unit PAd generates part of the prediction target sequence information based on information acquired from the mutation information storage unit 327c. This method is similar to the method of the fourth embodiment, and therefore, a description thereof will be omitted here.
[0259] The prediction target sequence generation unit PAd may associate each piece of generated prediction target sequence information with the generation date and time or the number of cycles at which the prediction target sequence information was generated. In this case, the output processing unit 336c or the user terminal 10 may output each piece of prediction target sequence information in the order in which it was generated, or in the order of the cycles in which it was generated, or for each cycle. The output processing unit 336c or the user terminal 10 may also classify the prediction target sequence information by generation date or number of cycles, and output each piece of prediction target sequence information in a different format for each classification. For example, the user terminal 10 may display each piece of prediction target sequence information generated in the latest cycle with a character string or image (e.g., "NEW") indicating that it is new. Furthermore, the prediction target sequence generating unit PAd may generate a sequence that has not been included in the previous cycles as a prediction target sequence. Here, an example using Bayesian optimization will be described.
[0260] Bayesian optimization is a method for finding the maximum value of a function with an unknown shape (such as antibody affinity in this case). The input during training is each antibody sequence and its affinity. After training, when a new hypothetical sequence is input as test data, an acquisition function is output. The acquisition function is, for example, the maximum affinity range of the sequence, estimated based on the uncertainty estimated from the data up to that point. Possible acquisition functions include the upper confidence bound and expected improvement. Sequences with a high acquisition function can be selected and proposed for experimentation. Saito et al., ACS Synth Biol. 2018 Sep 21;7(9):2014-2022, is an example, although not for antibodies. Available algorithms include GP-UCB and Thompson sampling.
[0261] As described above, in the information processing system 1d according to this embodiment, the output processing unit 336c (an example of an "output unit") outputs at least one piece of prediction target sequence information (candidate antibody information) from among the multiple pieces of prediction target sequence information input into the selected trained model (an example of a "second trained model"), according to the prediction score. Based on the prediction target sequence information output by the output processing unit 336c, additional characteristic evaluation is performed on the antibody indicated by the prediction target sequence information, and the analysis result information is stored as a training dataset. The sequence learning unit PA2d (an example of a "sequence learning unit") performs further machine learning based on the prediction target sequence information output by the output processing unit 336c, thereby generating a new sequence-trained model. The learning unit 334d (an example of a "learning unit") performs further machine learning based on the prediction target sequence information output by the output processing unit 336c and the binding determination information of the antigen-binding molecule represented by the prediction target sequence information (an example of "evaluation result information of characteristic evaluation"), thereby generating a selection trained model (or a characteristic prediction trained model: a "second trained model"). This allows the information processing system 1d to perform further machine learning using highly specific prediction target sequence information and its binding determination information. Furthermore, the information processing system 1d may be able to increase the proportion or number of highly binding sequence information in the training dataset. In this case, the information processing system 1d can generate virtual sequences with even higher characteristics. Furthermore, since the information processing system 1d sets the mutation position, it can generate sequence information of a virtual sequence with higher characteristics from among the sequence information in which the amino acid at the mutation position has been mutated. In this case, the information processing system 1d may be able to converge the sequence information in which the amino acid at the mutation position has been mutated to sequence information of a virtual sequence with higher characteristics.
[0262] (Hardware configuration) FIG. 37 is a block diagram showing an example of the hardware configuration of the server 30 according to the embodiment. The server 30 includes a CPU 901, a storage medium interface unit 902, a storage medium 903, an input unit 904, an output unit 905, a ROM 906, a RAM 907, an auxiliary storage unit 908, and an interface unit 909. The CPU 901, the storage medium interface unit 902, the storage medium 903, the input unit 904, the output unit 905, the ROM 906, the RAM 907, the auxiliary storage unit 908, and the interface unit 909 are connected to each other via a bus. The CPU 901 referred to here refers to a processor in general, and includes not only a device called a CPU in the narrow sense, but also, for example, a GPU, a DSP, etc. The CPU 901 referred to here is not limited to being realized by a single processor, but may be realized by combining multiple processors of the same or different types.
[0263] The CPU 901 controls the server 30 by reading and executing programs stored in the auxiliary storage unit 908, ROM 906, and RAM 907, and by reading various data stored in the auxiliary storage unit 908, ROM 906, and RAM 907 and writing the various data to the auxiliary storage unit 908 and RAM 907. The CPU 901 also reads various data stored in the storage medium 903 via the storage medium interface unit 902 and writes the various data to the storage medium 903. The storage medium 903 is a portable storage medium such as a magneto-optical disk, a flexible disk, or a flash memory, and stores various data. The storage medium interface unit 902 is an interface for reading and writing data from and to the storage medium 903 .
[0264] The input unit 904 is an input device such as a mouse, a keyboard, a touch panel, a volume control button, a power button, a setting button, and an infrared receiver. The output unit 905 is an output device such as a display unit and a speaker. The ROM 906 and RAM 907 store programs for operating the various functional units of the server 30 and various data. The auxiliary storage unit 908 is a hard disk drive, flash memory, or the like, and stores programs for operating each functional unit of the server 30 and various data. The interface unit 909 has a communication interface and is connected to the network NW via wireless or wired communication.
[0265] For example, the processing unit 33 in the functional configuration of the server 30 in Fig. 6 corresponds to the CPU 901 in the hardware configuration shown in Fig. 35. Also, for example, the storage unit 32 in the functional configuration of the server 30 in Fig. 6 corresponds to the ROM 906, RAM 907, or auxiliary storage unit 908, or any combination thereof, in the hardware configuration shown in Fig. 35. Also, for example, the communication unit 31 in the functional configuration of the server 30 in Fig. 6 corresponds to the interface unit 909 in the hardware configuration shown in Fig. 35.
[0266] Furthermore, the user terminal 10 and the next-generation sequencer 20 also have similar hardware configurations, so explanations of the hardware configurations of the user terminal 10 and the next-generation sequencer 20 will be omitted here.
[0267] In the first to third embodiments described above, an example was described in which the server 30 (30a, 30b) outputs candidate antibody information indicating candidate antibodies having affinity with a target antigen. The server 30 (30a, 30b) may output experimental conditions that provide the best evaluation result information for an antibody having a certain sequence, based on the information stored using the above-described method. The server 30 (30a, 30b) has a learning model for each round with different experimental conditions using the above-described method. Therefore, the server 30 (30a, 30b) also acquires the experimental conditions as a learning dataset. The server 30 (30a, 30b) associates the experimental conditions, learning model, and evaluation result information for each sequence. The server 30 (30a, 30b) performs learning using an existing learning model based on the associated information. At this time, the server 30 (30a, 30b) sets the learning model and evaluation result information as independent variables and the experimental conditions as dependent variables. The server 30 (30a, 30b) determines experimental conditions for outputting the best evaluation result information based on the learning results and the input sequence information, and outputs the information. This allows the experimental conditions for panning to be optimized, and the characteristics of antibodies evaluated to have higher affinity to be more pronounced. In this way, the information processing system 1 (1a, 1b) can reduce processing time or processing load. Therefore, the information processing system 1 (1a, 1b) can provide desired antibody information.
[0268] Furthermore, in the above-described first to third embodiments, examples have been described in which a prediction target sequence is generated based on sequence information and mutation information of a binding sequence, but the method for generating a prediction target sequence is not limited to this. For example, the prediction target sequence may be generated based on sequence information of a binding sequence. In this case, the prediction target sequence generation unit PA (the same applies to PAb, PAc, and PAd; the same applies below) determines the occurrence probability of an amino acid for each position in the binding sequence. For each position, the prediction target sequence generation unit PA determines information on amino acids whose occurrence probability is equal to or greater than a predetermined threshold. The prediction target sequence generation unit PA generates a prediction target sequence by applying one of the above amino acids to each position.
[0269] In the first to third embodiments, examples have been described in which learning is performed using learning data sets obtained from multiple pannings, but this is not limiting. Learning may also be performed using a learning data set obtained from a single panning. In this case, the classification criterion information is, for example, a threshold value for the appearance frequency in the panning.
[0270] Furthermore, in the first to third embodiments described above, the next-generation sequencer 20 separately outputs sequence information of the antibody heavy chain and sequence information of the antibody light chain, and the server 30 estimates the combination of the antibody heavy chain and the antibody light chain, but the present invention is not limited to this. For example, if the next-generation sequencer 20 can acquire sequence information including the antibody heavy chain and the light chain, the combination of the antibody heavy chain and the antibody light chain has already been determined, and therefore the above-mentioned combination estimation process is not performed.
[0271] Furthermore, in the first to third embodiments described above, the heavy chain sequence (or light chain sequence) may be stored as a present antibody sequence in the dataset storage unit 322. For example, the server 30 may store the heavy chain sequence (or light chain sequence) as a present antibody sequence in the dataset storage unit 322. In this case, the server 30 performs a learning process based on the heavy chain sequence (or light chain sequence) to generate a sequence-trained model. The server 30 performs a learning process based on the heavy chain sequence (or light chain sequence) and the occurrence probability to generate a characteristic prediction-trained model. The server 30 generates prediction target sequence information representing a heavy chain sequence (or a light chain sequence) using the sequence trained model. The server 30 inputs each of the generated prediction target sequence information into the property prediction trained model, and predicts a prediction score for each prediction target sequence information representing a heavy chain sequence (or a light chain sequence). Meanwhile, the server 30 estimates a light chain sequence (or heavy chain sequence) that pairs with a heavy chain sequence (or light chain sequence) represented by the sequence information to be predicted. The estimation may be performed by the method performed by the estimation unit 332 described above, or by another method. The light chain sequence (or heavy chain sequence) that pairs with the heavy chain sequence (or light chain sequence) may be selected by the user. The server 30 may generate candidate antibody information by combining the heavy chain sequence (or light chain sequence) represented by the sequence information to be predicted with the light chain sequence (or heavy chain sequence) that is estimated to pair with the heavy chain sequence (or light chain sequence).
[0272] In the first to third embodiments described above, examples have been described in which learning is performed based on all sequences acquired from the next-generation sequencer 20 in each panning, but the present invention is not limited to this. For example, based on the acquired sequence information, the sequence information may be classified into multiple clusters, and learning may be performed for each cluster.
[0273] In the first to third embodiments, the information processing system has been described as using panning as an example of affinity evaluation. However, the present invention is not limited to this, and the affinity evaluation may be any evaluation method other than panning as long as it evaluates the affinity between a target antigen and an antibody. In the above-described first to third embodiments, the information processing system has been described as using the frequency of appearance of antibodies as an example of evaluation result information. However, the present invention is not limited to this, and the evaluation result information may include the frequency of appearance of each sequence for each panning, the frequency of appearance of amino acids at each position for each panning, etc.
[0274] In the first to third embodiments described above, the sequence selection unit PA1 (the same applies to PA1b, PA1c, and PA1d; hereinafter, these will be collectively referred to as PA1) and the learning unit 334 (the same applies to 334b, 334c, and 334d; hereinafter, these will be collectively referred to as 334) determine the selected classification standard information and the selected trained model based on the calculated AUC when performing the evaluation process. However, this is not limited to this. For example, the determination may be based on the correlation between the dissociation constant (KD) and affinity information. Specifically, first, a sequence with a known KD is prepared in advance as an evaluation dataset (the known KD is also referred to as a "known KD"). Furthermore, the storage unit 32 (32a, 32b) stores correspondence information between affinity information and KD. The sequence selection unit PA1 and the learning unit 334 calculate affinity information for each sequence using the evaluation dataset. The sequence selection unit PA1 and the learning unit 334 convert the affinity information into a KD based on the calculated affinity information and correspondence information between the affinity information and the KD (the converted KD is also referred to as the "calculated KD"). The sequence selection unit PA1 and the learning unit 334 calculate a correlation coefficient for the pair of the calculated KD and the known KD for each trained model. The sequence selection unit PA1 and the learning unit 334 store the classification criterion candidate information with the highest correlation coefficient and the trained model corresponding to that classification criterion candidate information as the selected classification criterion information and the selected trained model, respectively, in the learning result storage unit 326, in association with a panning group ID.
[0275] In addition, in each of the above embodiments, examples have been described in which the sequence selection unit PA1, the sequence learning unit PA2, and the learning unit 334 perform the learning process using LSTM or a hidden Markov model, but this is not limiting. For example, the sequence selection unit PA1, the sequence learning unit PA2, and the learning unit 334 may use the random forest described in the fourth and fifth embodiments. For example, the sequence learning unit PA2 uses a learning model for unsupervised learning, allowing the virtual sequence generation unit PA3 to quickly generate many sequences. Specifically, the sequence learning unit PA2 generates a sequence-trained model that classifies sequence information based on characteristic values through unsupervised learning. The virtual sequence generation unit PA3 can generate a prediction target sequence with similar characteristics by generating a sequence that belongs to the same classification as the sequence. In the case of a model with instructions, the sequence learning unit PA2 performs machine learning using a dataset of sequence information and the results of property evaluation. In this case, the virtual sequence generation unit PA3 inputs the generated sequence information into the model with instructions, selects sequence information with a high property evaluation value output from the model with instructions, and sets it as the sequence information to be predicted. Here, for example, the virtual sequence generation unit PA3 may randomly generate amino acids for each component and generate sequence information indicating a sequence in which the generated amino acids are arranged as sequence information to be input into the model with instructions. Furthermore, sequence information with a high property evaluation value may be sequence information with a property evaluation value higher than a threshold value, or may be sequence information with a high property evaluation value.
[0276] In the fourth and fifth embodiments, the server 30c (30d) treats sequence information as a string of characters representing the amino acids of an antibody and performs learning and prediction. However, the sequence information is not limited to this. For example, the server 30c (30d) may convert the amino acid sequence of an antibody into a set of physical properties of the individual amino acids that make up the sequence. That is, the server 30c (30d) receives a character string of an amino acid sequence from the user terminal 10c. The server 30c (30d) converts the character string into physical properties based on information associating the character string with the physical properties.
[0277] Here, the physical properties of an amino acid are numerical values indicating the physicochemical or biochemical properties of the amino acid, such as those registered in AAindex. Specifically, the physical properties of an amino acid include the amino acid's volume, number of atoms, side chain length, surface area, charge, hydrophobicity, the region where it frequently appears in proteins (interior or surface), the type of secondary structure it readily adopts, the angle at which β-strands are formed, the energy change upon dissolution in water, the melting point, the heat capacity, and NMR data. Server 30c (30d) acquires a predetermined combination of physical properties (also referred to as a "position physical property group") for each individual amino acid as sequence information. In other words, the sequence information is information (a physical property group) that combines position physical properties for the number of amino acids constituting the antibody.
[0278] The information regarding the combination of physical properties (information indicating which combination of physical properties to use) may be stored in advance in the server 30c (30d), or may be information input from the user terminal 10c and stored in the server 30c (30d). Also, the combination of physical properties may be different for each characteristic. Furthermore, information on physical properties corresponding to the individual amino acids constituting the amino acid sequence does not have to be stored in the server 30c (30d). For example, it may be stored in the user terminal 10c and converted into a group of physical properties in advance. In this case, the server 30c (30d) receives the converted group of physical properties as sequence information. Furthermore, for example, information on physical properties corresponding to the individual amino acids constituting the amino acid sequence may be acquired by the server 30c (30d) or the user terminal 10c from the network NW. Furthermore, the server 30c (30d) may condense the amino acid sequence using a predetermined method before converting it into physical property quantities, such as a method using an autocorrelation function.
[0279] When the control unit 335c predicts the characteristic score of the prediction target sequence, the server 30c (30d) similarly converts the amino acid sequence into a characteristic quantity once. The control unit 335c predicts the characteristic score based on the converted characteristic quantity and the learning result.
[0280] The server 30c (30d) may also estimate a three-dimensional structure from the amino acid sequence and use information based on the estimated three-dimensional structure (also referred to as "structural information") as sequence information. The structural information is information indicating hydrophobic regions, positively charged regions, and negatively charged regions. This information may be expressed three-dimensionally or projected and expressed two-dimensionally.
[0281] FIG. 38 shows an example of the structure information. The example shown in this figure is a spherical projection of the surface properties of an antibody molecule (hydrophobic regions, positively charged regions, negatively charged regions) based on the results of analytical calculations of the three-dimensional structure of an antibody. The center indicates the surface that binds to the antigen. In this figure, regions indicated by 1 to 6 indicate hydrophobic regions on the surface. Regions indicated by 7 to 9 indicate positively charged regions on the surface. Region indicated by 10 indicates a negatively charged region on the surface. The server 30c (30d) extracts features from the amino acid structural information using a predetermined method. For example, the features are information indicating the position of properties on the surface of an antibody molecule, the size of a region, etc. The server 30c (30d) performs learning based on the extracted features and characteristic information. Note that the structural information may be estimated in advance by the user terminal 10c. Alternatively, the sequence information may be transmitted to the network NW, and the corresponding structural information may be received by the user terminal 10c or the server 30c from the network NW.
[0282] Furthermore, in the fourth and fifth embodiments described above, examples have been used in which characteristic information for an antigen is used, but this is not limiting. For example, characteristic information for other antigens may be used. This characteristic information is characteristic information relating to the physical properties of an antibody.
[0283] In the fourth and fifth embodiments described above, the characteristics (characteristic information) used for learning and prediction are not limited, but the characteristics (characteristic information) to be used may be limited. For example, this information may be input by the user of the user terminal 10c and transmitted to the server 30c (30d).
[0284] In the fourth and fifth embodiments described above, examples have been described in which learning is performed for each characteristic and a trained model is created, but this is not limiting. For example, learning may be performed for multiple characteristics collectively to create one trained model. In this case, characteristic evaluation information regarding multiple characteristics is output from the trained model.
[0285] In addition, in each of the above-described embodiments, an example has been described in which the prediction target sequence generation unit PA generates a prediction target sequence using LSTM, but this is not limited to this. For example, the prediction target sequence generation unit PA may identify a position to introduce a mutation based on the acquired sequence information. This method will be described below.
[0286] In the first to third embodiments described above, the prediction target sequence generation unit PA reads out a training data set whose bond determination is "bond" from the training data set (FIG. 11) associated with the selected classification criterion information. The prediction target sequence generation unit PA generates prediction target sequence information by changing amino acids at one or more positions from the sequence information of the read training data set. However, the present invention is not limited to this, and the prediction target sequence generation unit PA may also generate prediction target sequence information randomly. In addition, when mutation information is stored in the mutation information storage unit 327, the prediction target sequence generation unit PA generates prediction target sequence information by changing the amino acid at the position (element of sequence information) indicated by the mutation information from the sequence information of the read learning dataset. This allows the information processing system 1 to generate predicted target sequence information in which only amino acids at positions that are likely to bind and are desired to be mutated have been changed. The prediction target sequence generating unit PA stores the generated prediction target sequence information in the sequence storage unit 328.
[0287] In the fourth and fifth embodiments described above, the sequence selection unit PA1c (PA1d) and the sequence learning unit PA2c (PA2d) acquire sequence information of the sequence indicated by the improvement characteristic information from the sequence information stored in the learning dataset storage unit 324c (324d). The sequence selection unit PA1, the sequence learning unit PA2, and the learning unit 334 compare the amino acid sequence of the acquired sequence information with the amino acid sequence of the template sequence information, and identify mutation positions in the amino acid sequence. The sequence selection unit PA1, the sequence learning unit PA2, and the learning unit 334 store mutation position information indicating the identified mutation positions as one piece of mutation information in the mutation information storage unit 327c.
[0288] The prediction target sequence generation unit PAc (PA1d) generates prediction target sequence information based on template sequence information and mutation information. For example, the sequence selection unit PA1c (PA1d) reads template sequence information, mutation position information, and mutation condition information from the mutation information storage unit 327c. The sequence selection unit PA1c (PA1d) determines whether the number of mutation positions indicated by the mutation position information is greater than the upper limit number of mutations indicated by the mutation condition information. If the number of mutation positions is equal to or less than the upper limit number, the sequence selection unit PA1c (PA1d) determines all mutation positions as mutation introduction sites. If the number of mutation positions is greater than the upper limit number, the sequence selection unit PAc (PA1d) randomly selects the upper limit number of mutation introduction positions from the mutation positions. The sequence selection unit PA1c (PA1d) generates prediction target sequence information by changing the amino acids indicated by the mutation introduction positions from the template sequence information. The sequence selection unit PA1c (PA1d) stores the generated prediction target sequence information in the sequence storage unit 328c.
[0289] In the above-described method, the server 30c determines the mutation introduction position, but this is not limiting. For example, a user of the user terminal 10c may input mutation information indicating a specific mutation introduction position and transmit it to the server 30c (30d). In this case, the server 30c (30d) stores the received mutation information in the mutation information storage unit 327c. The server 30c (30d) also determines the mutation introduction position so as to include the mutation introduction position.
[0290] In the fourth and fifth embodiments described above, the server 30c (30d) determines the candidate antibody information based on the ranking for each characteristic when outputting the candidate antibody information. However, this is not limiting. For example, the user terminal 10c may determine the candidate antibody information using multiple characteristics. In this case, the server 30c transmits the ranking results for each characteristic (the predicted target sequence and the ranking results for each characteristic) to the user terminal 10c.
[0291] In the fourth and fifth embodiments, the server 30c (30d) generates a prediction target sequence, but the present invention is not limited to this. For example, a prediction target sequence may be input by a user of the user terminal 10c and transmitted to the server 30c (30d).
[0292] In the fourth and fifth embodiments described above, the server 30c (30d) pre-stores parameters related to the entire intermediate layer during LSTM training and changes the parameters appropriately during re-training. However, this is not limiting. For example, each parameter may be input by the user of the user terminal 10c. In this case, for example, the server 30c (30d) transmits information requesting parameters related to the entire intermediate layer to the user terminal 10c as needed. The user terminal 10c displays the received information on the display unit 15 and transmits information input by the user of the user terminal 10c to the server 30c (30d). The server 30c (30d) determines the parameters based on the received information.
[0293] In the fourth and fifth embodiments, the sequence information and characteristic information used in the training data set are determined based on the number of cycles. However, this is not limiting. For example, sequence information and characteristic information with superior specific characteristics may be prioritized as the training data set. In this case, the server 30c stores information indicating the emphasized characteristics in the storage unit 32c or receives it from the user terminal 10c. Furthermore, for example, the server 30c may determine the sequence information and characteristic information to be used in the training dataset based on information indicating the date and time when the sequence information and characteristic information were acquired. Specifically, the server 30c acquires a predetermined upper limit of the sequence information and characteristic information to be used in the training dataset, starting with the most recent acquired date and time. The acquired date and time here may be the date and time when the server 30c acquired each piece of information, or the date and time when the characteristic information was actually acquired by performing a measurement or the like. Furthermore, for example, the server 30c may perform learning based on all the sequence information and characteristic information acquired up to that point. At this time, the server 30c may perform weighting according to the number of cycles. In other words, the server 30c may place more importance on sequence information and corresponding characteristic information generated in more recent cycles.
[0294] Although the above-described embodiments have been described using antibodies as an example of antigen-binding molecules, the term "antigen-binding molecule" is not limited to this. That is, the term "antigen-binding molecule" is used in the broadest sense. Specifically, the term "antigen-binding molecule" encompasses various molecular forms as long as it exhibits antigen-binding activity. For example, when the antigen-binding molecule is a molecule in which an antigen-binding domain and an Fc region are bound together, examples include complete antibodies and antibody fragments. Examples of antibodies include single monoclonal antibodies (including agonist and antagonist antibodies), human antibodies, humanized antibodies, chimeric antibodies, and the like. When an antibody fragment is used, preferred examples include antigen-binding domains and antigen-binding fragments (e.g., VHH, Fab, F(ab')2, scFv, and Fv). The antigen-binding molecules of the present disclosure also include scaffold molecules in which existing stable α / β barrel protein structures and other three-dimensional structures are used as scaffolds, and only partial structures of these structures are compiled into libraries for constructing antigen-binding domains.
[0295] Furthermore, in the above-described first to third embodiments, an example has been described in which the user terminal 10, the next-generation sequencer 20, and the server 30 are connected via a network NW, but this is not limited thereto. For example, the server 30 and the next-generation sequencer 20 may be the same. Furthermore, the user terminal 10 and the server 30 may be the same. Furthermore, an example has been described in which the server 30 converts a base sequence acquired from the next-generation sequencer into an amino acid sequence, but this is not limited thereto. For example, a conversion device that performs a process of converting a base sequence into an amino acid sequence may be located outside the server 30. In this case, the conversion device converts the base sequence of information input from the next-generation sequencer 20 into an amino acid sequence and outputs the converted information to the server 30. Furthermore, the next-generation sequencer 20 is not limited to a next-generation sequencer. For example, it may be another sequencer.
[0296] In the fourth and fifth embodiments described above, the user terminal 10c and the server 30c (30d) are connected via a network NW, but this is not limiting. For example, the server 30c and the user terminal 10c may be the same. For example, a device that performs the learning process of the server 30c (30d) may be provided outside the server 30c (30d). For example, a device that performs the prediction process of the server 30c (30d) may be provided outside the server 30c (30d).
[0297] In each of the above embodiments, the prediction target sequence generating unit PA and the learning unit 334 calculate the definite vector value h t-1 For example, the prediction target sequence generation unit PA and the learning unit 334 output an amino acid with an appearance probability of 40% so that it is selected 4 times out of 10 times. This allows the predicted target sequence generation unit PA and the learning unit 334 to produce a variety of amino acids at each position, compared to when selecting, for example, one amino acid with the maximum value, and to output a variety of predicted target sequences or candidate antibody information.
[0298] In each of the above-described embodiments, the prediction target sequence generation unit PA and the learning unit 334 may be provided in separate devices. Also, the sequence selection unit PA1, sequence learning unit PA2 (similar to PA2b, PA2c, and PA2d; hereinafter, these will be collectively referred to as PA2), and virtual sequence generation unit PA3 (similar to PA3b, PA3c, and PA3d; hereinafter, these will be collectively referred to as PA3) may each be provided in separate devices. In addition, in each of the above-described embodiments, the server 30 (30a, 30b, 30c, and 30d are also included; hereinafter, these will be collectively referred to as 30) may not include the classification unit 333, and the sequence selection unit PA1 may not generate an LSTM. In this case, the sequence selection unit PA1 selects sequence information from the training data set whose characteristic values satisfy predetermined conditions. The sequence learning unit PA2 generates a sequence-trained model by performing a learning process based on the sequence information selected by the sequence selection unit PA1. In addition, the server 30 may not include the learning unit 334b. In this case, the server 30 generates, as candidate antibody information, part or all of the prediction target sequence generated by the virtual sequence generation unit PA3. The output processing unit 336 outputs the generated candidate antibody information. In addition, the server 30 may not include the estimation unit 332.
[0299] In the above embodiment, the significance of generating a group of virtual sequences by machine learning and predicting values such as predicted scores for them is that there are sequences that do not express well in phage display experiments and are difficult to perform binding experiments on. This has the advantage that evaluation of sequences that are difficult to perform experiments on can be performed by computer.
[0300] In the above embodiment, the significance of generating a group of virtual sequences using LSTM is that simple enumeration can result in an explosion of combinations, making it impossible to handle even a high-performance computer. For example, if a total of 20 positions in Hch and Lch were comprehensively assigned to 19 types of amino acids, there would be an enormous number of combinations (19 to the 20th power = 3.76 × 10 to the 25th power), and it would be extremely difficult to evaluate all of them using a computer. For this reason, in the above embodiment, it is important to use LSTM to learn sequences with "good properties" and generate a group of sequences that are "likely to have good properties."
[0301] [Example] An example according to the first embodiment will be described below. Sequence information was obtained when a phage display library was panned against a certain antigen K. The sequence information was included in analysis result information obtained by analysis using a next-generation sequencer (NGS). The sequence information in the analysis result information is after panning, and is therefore sequence information of a group of antibodies that bind to the antigen K. The sequence group indicated by the sequence information after this panning was used to perform a learning process on the LSTM, and the LSTM after the learning process was used to generate a group of virtual sequences that are likely to bind as prediction target sequences.
[0302] FIG. 39 is a diagram showing the relationship between sequences and properties according to this example. The horizontal axis of FIG. 39 represents the type of sequence, and the vertical axis represents the property value. The property value is expressed as the negative common logarithm (-log 10 (KD)). Figure 39 plots the relationship between each sequence and characteristic values for the "ML top" sequence group and the "NGS top" sequence group. The "ML top" sequence group is a group of sequences whose likelihood P (prediction score) is within the top 10 of the prediction target sequences generated using LSTM after the learning process. The "NGS top" sequence group is a group of sequences whose appearance frequency is within the top 10 of the analysis result information from the next-generation sequencer. Figure 39 shows the characteristic values (-log 10 Box plots are shown for (KD).
[0303] Comparing the "ML top" sequence group with the "NGS top" sequence group, the "ML top" sequence group has a higher characteristic value than the "NGS top" sequence group. In other words, it can be seen that the predicted target sequence has a stronger binding ability than the sequence analyzed by the next-generation sequencer. In this way, by using the LSTM after the learning process, the server 30 was able to generate a virtual sequence group (predicted target sequence) with a stronger binding ability than the sequence analyzed by the next-generation sequencer 20 (the sequence used in the learning process). It was also found that scoring predicted antibody sequences using the likelihood P is effective.
[0304] 40 is a diagram showing the prediction accuracy of sequence characteristics according to this example. The vertical axis of FIG. 40 indicates the predicted value of affinity, which is the negative common logarithm (-log 10 (P)) (In Figure 40, -log 10 The horizontal axis of Figure 40 shows the actual measured affinity values, which are expressed as the negative common logarithm (-log 10 (KD)). In Figure 40, predicted values (vertical axis) and actual measured values (horizontal axis) are plotted for some of the sequences to be predicted, which were generated using the LSTM after the learning process. In this figure, the absolute value of the correlation coefficient between the predicted and actual measured values was 0.576. Also, as this figure shows, the larger the likelihood P (negative common logarithm (-log 10 Since likelihood P is the likelihood of binding to a sequence with high binding ability, the lower the value on the vertical axis, the stronger the affinity (the higher the value on the horizontal axis). In other words, the information processing system 1 (server 30) achieved high prediction accuracy in predicting sequences with strong binding ability. For sequences with high binding ability, a group of virtual sequences (sequences to be predicted) is generated using likelihood P as an index, making it possible to predict sequences with high binding ability.
[0305] In Figure 40, the dashed line "NGS freq top" indicates the dissociation constant (KD) of the most frequently occurring sequence obtained from the next-generation sequencer after panning. The dotted line "Control" indicates the dissociation constant (KD) of the template sequence used to create a group of sequences for panning. As shown in this figure, the information processing system 1 is able to generate, using LSTM, a sequence with stronger binding than the most enriched sequence after panning as a predicted target sequence. As described above, the LSTM according to the first embodiment is able to generate a group of virtual sequences with strong binding ability, and the scoring of predicted antibody sequences using the likelihood P is effective.
[0306] An example of the third embodiment will be described below. FIG. 41 is a diagram showing the similarity between the training sequence and the virtual sequence according to this example. The training sequence is a sequence used for training the LSTM, and is a sequence of the training data set. The vertical axis of FIG. 41 is the negative common logarithm of the dissociation rate (acidic koff) when the pH of the reaction solution is acidic. In FIG. 41, the common logarithm (-log koff) of the dissociation rate when the pH is 5.8 (acidic) is 10 The horizontal axis of Figure 41 is the negative common logarithm (-log KD) of the dissociation constant (neutral KD) when the pH of the reaction solution is neutral. 10 (KD)).
[0307] Figure 41(a) plots the characteristic values (measured values) for each training sequence. These sequences are -log 10 (KD)>9 and log 10 The plot shows 251 sequences that satisfy (koff)<2. Neutral KD and acidic koff are actual measured values. Figure 41(b) plots the characteristic values (predicted values) for each sequence generated by LSTM. Figure 41(b) plots a group of 1,000 hypothetical sequences output from a trained model when machine learning (LSTM) was performed using the training dataset shown in Figure 41(a). The neutral KD and acidic koff are predicted values. Figure 41(c) plots a group of 1,000 new hypothetical sequences generated by listing the mutated residues contained in the sequence shown in Figure 41(a) and randomly shuffling and combining them. The neutral KD and acidic koff are predicted values.
[0308] Comparing Figure 41(a) and (c), Figure 41(c) generated a sequence with a significantly different acidic koff range from the training sequence in Figure 41(a). This is thought to be because shuffling destroys the combination of mutations that "synergistically improve neutral KD and acidic koff," resulting in only one mutation being adopted. On the other hand, Figure 41(b) generated a sequence with a similar acidic koff range to the training sequence in Figure 41(a) compared to Figure 41(c). In this way, the hypothetical sequence group (prediction target sequence) using LSTM more strongly reflects the characteristics of the training sequence compared to when the sequence at the mutation position is randomly changed.
[0309] Figure 42 is another diagram showing the similarity between training sequences and virtual sequences according to this embodiment. The horizontal axis of Figure 42 represents the first principal component in principal component analysis, and the vertical axis represents the second principal component. In the principal component analysis, each sequence was mapped to a vector space using the Doc2Vec method. Here, the model used in the Doc2Vec method was a model that had undergone a learning process using sequences from Uniprot (http: / / www.uniprot.org / ), a protein sequence database. Figure 42(a) plots the values of each principal component for each training sequence, which is the same sequence as the sequence plotted in Figure 41(a). Figure 42(b) plots the values of each principal component for each sequence generated by the LSTM. These sequences are the same sequences as those plotted in Figure 41(b). Figure 42(c) plots the values of each principal component for sequences in which the sequences at the mutation positions were randomly altered. These sequences are the same sequences as those plotted in Figure 41(c).
[0310] In Figures 42(a) to (c), the numerical vectors are vectors that indicate the characteristics of actual amino acid sequences, so similar amino acid sequences are expressed as similar vectors. Comparing Figures 42(a) and (b) shows that the values of each sequence are closer than those in Figure 42(c). In other words, it can be seen that each sequence generated by LSTM is close to the amino acid sequence used for training. On the other hand,...
Claims
1. a sequence generation unit that, when sequence information on an amino acid sequence of an antigen-binding molecule and mutation information indicating at least one mutation position on the amino acid sequence are input, inputs the sequence information and the mutation information into a sequence-trained model that has been generated by machine learning to generate prediction target sequence information representing a hypothetical amino acid sequence in which each amino acid at the at least one mutation position is mutated, thereby generating the prediction target sequence information; a prediction unit that inputs the sequence information to be predicted into a property prediction trained model obtained by machine learning based on sequence information on amino acid sequences of a plurality of antigen-binding molecules, first property information indicating a first type of property, and second property information indicating a second type of property for each of the plurality of antigen-binding molecules, and performs arithmetic processing of the property prediction trained model to predict first property evaluation information indicating the first type of property of the sequence information to be predicted and second property evaluation information indicating the second type of property of the sequence information to be predicted; An information processing system comprising:
2. the characteristic prediction trained model includes a first trained model that is a training result based on the sequence information and the first characteristic information, and a second trained model that is a training result based on the sequence information and the second characteristic information, the prediction unit predicts the first characteristic evaluation information based on the first trained model, and predicts the second characteristic evaluation information based on the second trained model; The information processing system according to claim 1 .
3. the first characteristic information is activity information indicating the activity between a target antigen and the antigen-binding molecule; the second characteristic information is physical property information indicating a physical property of the antigen-binding molecule; 3. The information processing system according to claim 1.
4. The first type of characteristic and the second type of characteristic are information on at least two of the affinity, pharmacological activity, physical properties, kinetics, and safety of the antigen-binding molecule. The information processing system according to claim 1 .
5. The first characteristic information and the second characteristic information are at least two pieces of information selected from the group consisting of binding activity information indicating the binding activity between a target antigen and the antigen-binding molecule, pharmacological activity information indicating the pharmacological activity between the target antigen and the antigen-binding molecule, and stability information indicating the stability of the antigen-binding molecule. The information processing system according to any one of claims 1 to 4.
6. the first characteristic information and the second characteristic information are a plurality of types of binding activity information, a plurality of types of pharmacological activity information, or a plurality of types of stability information; The information processing system according to claim 5 .
7. the prediction target sequence information includes at least one of character string information representing the amino acid sequence of the antigen-binding molecule, physical property amount information indicating physical property amounts of amino acids included in the amino acid sequence of the antigen-binding molecule, and three-dimensional structure information indicating three-dimensional structural characteristics based on the amino acid sequence of the antigen-binding molecule; The information processing system according to any one of claims 1 to 6.
8. the sequence generation unit generates new prediction target sequence information using a sequence trained model generated by further machine learning based on the prediction target sequence information generated based on the sequence trained model; the prediction unit inputs the new sequence information to be predicted into the property prediction trained model generated by further machine learning based on the results of property evaluation of the first property and / or the second property for the antigen-binding molecule represented by the virtual amino acid sequence represented by the sequence information to be predicted, and estimates a predicted value of property evaluation for the antigen-binding molecule of the sequence represented by the input new sequence information to be predicted. The information processing system according to any one of claims 1 to 7.
9. The sequence-trained model learns features of a first sequence represented by sequence information based on the sequence information about the antigen-binding molecule or protein in which one or more properties selected from a plurality of properties in property evaluation of the antigen-binding molecule satisfy selection conditions for the properties. The information processing system according to any one of claims 1 to 8.
10. an output unit that outputs the prediction target sequence information in descending order of the predicted value estimated by the prediction unit; The information processing system according to claim 1 , comprising:
11. the array generation unit outputs a plurality of pieces of virtual array information; the output unit outputs at least one piece of prediction target sequence information from the plurality of pieces of prediction target sequence information input to the characteristic prediction trained model in descending order of the prediction value. The information processing system according to claim 10.
12. a sequence learning unit that generates the sequence-trained model and a characteristic prediction learning unit that generates the characteristic prediction-trained model The information processing system according to any one of claims 1 to 11.
13. The sequence learning unit performs the machine learning using a deep learning model or a probabilistic model. The information processing system according to claim 12.
14. the sequence learning unit performs the machine learning using a deep learning model; The machine learning is performed using any one of a long short-term memory (LSTM), a recurrent neural network (RNN), a gated recurrent unit (GRU), a generative adversarial network (GAN), or a variational autoencoder (VAE) as the deep learning model. The information processing system according to claim 13.
15. the sequence learning unit performs the machine learning using a probabilistic model, and the machine learning is performed using either a hidden Markov model (HMM) or a Markov model (MM) as the probabilistic model. The information processing system according to claim 14.
16. the sequence learning unit performs the machine learning based on the sequence information expressed as a character string, a numerical vector, or a physical property of a constituent unit that constitutes the sequence; 16. The information processing system according to any one of claims 12 to 15.
17. The antigen-binding molecule is a protein, an antibody, or a peptide.
17. An information processing system according to any one of claims 1 to 16.
18. An information processing method in an information processing system, comprising: a sequence generation step of inputting sequence information on the amino acid sequence of an antigen-binding molecule and mutation information indicating at least one mutation position on the amino acid sequence into a sequence-trained model generated by machine learning so as to generate prediction target sequence information representing a hypothetical amino acid sequence in which each amino acid at the at least one mutation position is mutated, thereby generating the prediction target sequence information; an estimation process of inputting the sequence information to be predicted into a property prediction trained model obtained by machine learning based on sequence information on amino acid sequences of a plurality of antigen-binding molecules, first property information indicating a first type of property, and second property information indicating a second type of property for each of the plurality of antigen-binding molecules, and performing arithmetic processing of the property prediction trained model to predict first property evaluation information indicating the first type of property of the sequence information to be predicted and second property evaluation information indicating the second type of property of the sequence information to be predicted; An information processing method comprising:
19. To the computer of the information processing system, a sequence generation step of inputting sequence information on the amino acid sequence of an antigen-binding molecule and mutation information indicating at least one mutation position on the amino acid sequence into a sequence-trained model generated by machine learning so as to generate prediction target sequence information representing a hypothetical amino acid sequence in which each amino acid at the at least one mutation position is mutated, thereby generating the prediction target sequence information; an estimation procedure of inputting the sequence information to be predicted into a property prediction trained model obtained by machine learning based on sequence information on amino acid sequences of a plurality of antigen-binding molecules, first property information indicating a first type of property, and second property information indicating a second type of property for each of the plurality of antigen-binding molecules, and performing arithmetic processing of the property prediction trained model to predict first property evaluation information indicating the first type of property of the sequence information to be predicted and second property evaluation information indicating the second type of property of the sequence information to be predicted; A program that executes the following.
Citation Information
Patent Citations
Insilico generation and selection of protein libraries
JP2005526518A
Machine learning based antibody design
US20190065677A1
Machine learning based antibody design
WO2018132752A1