Information processing device, method for operating information processing device, and program for operating information processing device

By acquiring the amino acid sequence structure information of antibodies and setting feature weights, and using machine learning models to predict property information, the problem of low accuracy in antibody property information prediction in existing technologies is solved, achieving high-precision property information prediction and supporting the stability and efficacy of biopharmaceuticals.

CN121794754APending Publication Date: 2026-04-03FUJIFILM CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Current technologies cannot accurately predict the properties of general proteins, especially the properties of antibodies, such as the optimal hydrogen ion index.

Method used

By obtaining structural information about the amino acid sequence unique to proteins, weights of feature quantities are assigned and applied to machine learning models to predict property information.

Benefits of technology

It enables high-precision prediction of antibody properties, especially accurate prediction of the optimal hydrogen ion index, supporting the stability and efficacy of biopharmaceuticals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121794754A_ABST
    Figure CN121794754A_ABST
Patent Text Reader

Abstract

Provided is an information processing device provided with a processor that performs: a process for acquiring structural information on an amino acid sequence unique to a protein and a characteristic amount of an amino acid residue constituting the amino acid sequence; setting a weight of the feature quantity according to the structure information; and causing the machine learning model to predict property information indicating the property of the protein by applying the feature amount to which the weight has been set to the machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing device, an operating method of the information processing device, and an operating procedure of the information processing device. Background Technology

[0002] In recent years, biopharmaceuticals, peptide drugs, and nucleic acid drugs have received much attention due to their high efficacy and few side effects. For example, biopharmaceuticals may use proteins such as interferon and antibodies as active ingredients. Regarding such proteins, such as antibodies, property information indicating their properties varies depending on the type, such as the hydrogen ion index (pH value) that best maintains their quality and allows for the acquisition of the highest activity level.

[0003] Japanese Patent Application Publication No. 2021-158973 discloses a technique that uses a machine learning model to predict the thermal stability (whether the heat resistance temperature is higher than a set temperature, etc.) of protein variants with amino acid substitution mutations in their amino acid sequences. In order to improve the accuracy of thermal stability prediction, Japanese Patent Application Publication No. 2021-158973 selectively applies characteristic quantities of amino acid residues in local regions of protein variants centered on the amino acid substitution mutation (type and number of amino acid residues in the local region, size of amino acid residues in the local region, angle of side chains, hydrophobicity, etc.) to the machine learning model. Summary of the Invention

[0004] The technical problem to be solved by the invention The technology described in Japanese Patent Application Publication No. 2021-158973 targets protein variants. Therefore, it cannot be used for general proteins that are not variants.

[0005] One embodiment of the present invention provides an information processing device, an operating method for the information processing device, and an operating program for the information processing device that can predict property information representing the properties of proteins with high accuracy.

[0006] means for solving technical problems The information processing apparatus of the present invention includes a processor that performs the following processing: acquiring structural information of a protein-specific amino acid sequence and feature quantities of amino acid residues constituting the amino acid sequence; setting weights for the feature quantities based on the structural information; and applying the weighted feature quantities to a machine learning model to enable the machine learning model to predict property information representing the properties of the protein.

[0007] The structural information is preferably segment region information representing the functional segment region of amino acid residues.

[0008] Preferably, the protein is an antibody, and the segment region includes a complementarity-determining region. The processor performs the following processing: the weight of the feature amount of amino acid residues in the segment region represented by the segment region information that is a complementarity-determining region is set to be higher than the weight of the feature amount of amino acid residues in the segment region information that is not a complementarity-determining region.

[0009] The structural information is preferably the three-dimensional structural information of the amino acid sequence.

[0010] The preferred three-dimensional structure information is three-dimensional shape information representing the three-dimensional shape of the amino acid sequence.

[0011] Preferably, the three-dimensional shape includes a ring shape and a linear shape, and the processor performs the following processing: the weight of the feature quantity of the amino acid residue whose three-dimensional shape is a ring shape is set to be higher than the weight of the feature quantity of the amino acid residue whose three-dimensional shape is a linear shape is represented by the three-dimensional shape information.

[0012] The processor is preferably a feature quantity or machine learning model targeting amino acid residues, and performs preprocessing according to weights.

[0013] Pretreatment is preferably a weighted average calculated based on the weights of the characteristic amounts of amino acid residues.

[0014] Preprocessing is preferably a process of setting initial coefficients in the machine learning model based on the weights.

[0015] Property information is used to select the optimal hydrogen ion index for proteins.

[0016] The preferred protein is an antibody.

[0017] The operation method of the information processing device of the present invention includes the following steps: acquiring structural information of the amino acid sequence specific to the protein and feature quantities of the amino acid residues constituting the amino acid sequence; setting weights of the feature quantities according to the structural information; and applying the feature quantities with the set weights to a machine learning model to enable the machine learning model to predict property information representing the properties of the protein.

[0018] The operating program of the information processing device of the present invention causes a computer to perform processing including the following steps: acquiring structural information of the amino acid sequence specific to the protein and feature quantities of the amino acid residues constituting the amino acid sequence; setting weights for the feature quantities based on the structural information; and applying the feature quantities with the set weights to a machine learning model to enable the machine learning model to predict property information representing the properties of the protein.

[0019] Invention Effects According to the technology of the present invention, an information processing device, an operating method of the information processing device, and an operating program of the information processing device can be provided that are capable of predicting property information representing the properties of proteins with high accuracy. Attached Figure Description

[0020] Figure 1 This is a diagram showing the information processing server and user terminals.

[0021] Figure 2 This is a diagram showing the basic structure of an antibody.

[0022] Figure 3 This is a block diagram showing the computer that constitutes the information processing server and user terminals.

[0023] Figure 4 This is a block diagram showing the processing unit of the CPU in an information processing server.

[0024] Figure 5 It is a diagram showing structural information.

[0025] Figure 6 This is a diagram showing the processing of the feature extraction unit and the feature set.

[0026] Figure 7 This is a diagram showing the settings for auxiliary information.

[0027] Figure 8 This is a diagram showing the processing of the weight setting unit.

[0028] Figure 9 This is a graph showing the weighted characteristic set.

[0029] Figure 10 This is a diagram showing the integrated feature quantities.

[0030] Figure 11 This is a diagram illustrating the processing of the prediction unit.

[0031] Figure 12 This is a diagram illustrating the neural network that constitutes the property information prediction model.

[0032] Figure 13 This is a block diagram showing the processing unit of the CPU in a user terminal.

[0033] Figure 14 This is a diagram showing the information input screen.

[0034] Figure 15 This is a diagram showing the prediction results.

[0035] Figure 16 This is a flowchart illustrating the processing sequence of the information processing server.

[0036] Figure 17This is another example of a diagram illustrating structural information.

[0037] Figure 18 This is another example of setting auxiliary information.

[0038] Figure 19 This is another example of a diagram illustrating structural information.

[0039] Figure 20 This is another example of setting auxiliary information.

[0040] Figure 21 This is a diagram showing the processing of the preprocessing unit in the second embodiment. Detailed Implementation

[0041] [First Implementation] As an example, such as Figure 1 As shown, the information processing server 10 is connected to the user terminal 11 via network 12. The information processing server 10 is an example of an "information processing device" according to the technology of this invention.

[0042] User terminal 11 is located at a pharmaceutical company developing biopharmaceuticals or an organization commissioned by a pharmaceutical company to develop biopharmaceuticals, i.e., a Contract Research Organization (CRO). User terminal 11 is operated by user U who participates in the development of biopharmaceuticals at the pharmaceutical company or CRO (hereinafter collectively referred to as the pharmaceutical facility). Network 12 is, for example, the Internet or a public communication network such as a WAN (Wide Area Network). Additionally, in Figure 1 In this case, only one user terminal 11 is connected to the information processing server 10, but in reality, multiple user terminals 11 from multiple pharmaceutical facilities are connected to the information processing server 10.

[0043] User terminal 11 sends a prediction request 15 to information processing server 10. Prediction request 15 is a request to cause information processing server 10 to predict property information 17 representing the properties of antibody 16 as an active ingredient in a biopharmaceutical. Prediction request 15 includes amino acid sequence information 18 of antibody 16. Amino acid sequence information 18 is determined experimentally and is obtained through user terminal 11's input device 30B (reference 10B). Figure 13 The input is made by [the user terminal]. Antibody 16 is an example of a "protein" involved in the technology of this invention. In addition, although the illustration is omitted, the prediction request 15 also includes a terminal ID (Identification Data) for uniquely identifying the user terminal 11 that sent the prediction request 15.

[0044] Regarding amino acid sequence information 18, the order of peptide bonds of the amino acid residues constituting antibody 16 is recorded using letter abbreviations representing the amino acid residues, from the amino terminus to the carboxyl terminus. Since there are approximately 450 amino acid residues constituting antibody 16, the letters in amino acid sequence information 18 are also arranged in a continuous sequence of approximately 450 residues. For example, "E" represents glutamic acid, "L" represents leucine, and "G" represents glycine. Such a sequence of amino acid residues is also called the primary structure.

[0045] Upon receiving prediction request 15, information processing server 10 predicts property information 17 based on amino acid sequence information 18. Here, property information 17 is the hydrogen ion index of the preservation solution for antibody 16 that best maintains the quality of antibody 16, i.e., the optimal hydrogen ion index for antibody 16. Information processing server 10 distributes property information 17 to user terminal 11, the source of prediction request 15. Upon receiving property information 17, user terminal 11 displays property information 17 on display 29B of user terminal 11 (see reference). Figure 13 In the section, user U is provided with information about the nature of the information (17).

[0046] As an example, such as Figure 2 As shown, antibody 16 essentially has four polypeptide chains: two identical heavy chains (HC) and two identical light chains (LC). Antibody 16 is a structure formed by these two heavy chains (HC) and two light chains (LC) bonded together by disulfide bonds (DB), and it has a Y-shaped structure that is symmetrical from left to right.

[0047] The heavy chain HC consists of a variable domain VH (Heavy Chain) and constant domains CH1, CH2, and CH3. The light chain LC consists of a variable domain VL (Light Chain) and a constant domain CL (Light Chain). The constant domains CH1 and CH2 of the heavy chain HC are connected by a hinge region HR. The variable domains VH and VL are collectively referred to as the variable region. The constant domains CH1 to CH3 and CL are collectively referred to as the constant region. The variable domains VH and VL, and the constant domains CH1 to CH3 and CL are examples of the "segment regions" involved in the technology of this invention.

[0048] The variable domain VH includes three complementarity-determining regions (CDR-H, Heavy Chain): CDR-H1, CDR-H2, and CDR-H3. Similarly, the variable domain VL also includes three complementarity-determining regions (CDR-L, Light Chain): CDR-L1, CDR-L2, and CDR-L3. These CDR-H1–CDR-H3 and CDR-L1–CDR-L3 are the binding sites with the antigen and are also referred to as the hypervariable region. The CDR-H1–CDR-H3 and CDR-L1–CDR-L3 are also examples of the "segment regions" involved in the technology of this invention.

[0049] The region comprised of variable domains VH and VL and constant domains CH1 and CL is the fragment antigen-binding region Fabr. The region comprised of constant domains CH2 and CH3 and a portion of the hinge region HR is the crystallizable fragment region FcR. The region comprised of variable domains VH and VL is the fragment variable region FvR. It is well known that the fragment antigen-binding region Fabr and the crystallizable fragment region FcR can be produced by papain digestion. Similarly, the fragment variable region FvR can be produced by pepsin digestion.

[0050] As an example, such as Figure 3 As shown, the computers constituting the information processing server 10 and the user terminal 11 have essentially the same structure, including a storage device 25, a memory 26, a CPU (Central Processing Unit) 27, a communication unit 28, a display 29, and an input device 30. They are connected to each other via a bus 31.

[0051] Storage device 25 is a hard disk drive built into the computer constituting information processing server 10 and user terminal 11, or connected via cable or network. Alternatively, storage device 25 is a disk array assembled by connecting multiple hard disk drives in parallel. Storage device 25 stores control programs such as operating systems, various application programs (hereinafter referred to as APs), and various data associated with these programs. Alternatively, solid-state drives can be used instead of hard disk drives.

[0052] Memory 26 is an operational memory used by CPU 27 for processing. CPU 27 loads the program stored in storage device 25 into memory 26 and executes the program accordingly. Thus, CPU 27 centrally controls the various parts of the computer. CPU 27 is an example of a "processor" according to the technology of this invention. Alternatively, memory 26 may be built into CPU 27.

[0053] The communication unit 28 controls the transmission of various information with external devices. The display 29 shows various screens. These screens are operable via a GUI (Graphical User Interface). The computer constituting the information processing server 10 and the user terminal 11 receives operation commands from the input device 30 through these screens. The input device 30 includes a keyboard, mouse, touch panel, and microphone for voice input.

[0054] In addition, in the following description, the components of the computer constituting the information processing server 10 (storage device 25 and CPU 27) are marked with reference numeral "A", and the components of the computer constituting the user terminal 11 (storage device 25, CPU 27, display 29 and input device 30) are marked with reference numeral "B" for distinction.

[0055] As an example, such as Figure 4 As shown, the storage device 25A of the information processing server 10 stores an operating program 35. The operating program 35 is an application program (AP) used to enable the computer to function as the information processing server 10. That is, the operating program 35 is an example of an "operating program for an information processing device" according to the technology of this invention. The storage device 25A also stores a feature extraction model 36, setting auxiliary information 37, and a property information prediction model 38. The property information prediction model 38 is an example of a "machine learning model" according to the technology of this invention.

[0056] When the program 35 is started, the CPU 27 and memory 26 of the computer constituting the information processing server 10 cooperate to function as the receiving unit 40, the read / write (hereinafter referred to as RW) control unit 41, the structural information export unit 42, the feature extraction unit 43, the weight setting unit 44, the preprocessing unit 45, the prediction unit 46, and the screen distribution control unit 47.

[0057] The receiving unit 40 receives various requests from the user terminal 11. In particular, the receiving unit 40 receives a prediction request 15 from the user terminal 11. As described above, the prediction request 15 includes amino acid sequence information 18. Therefore, the receiving unit 40 obtains the amino acid sequence information 18 by receiving the prediction request 15. The receiving unit 40 outputs the amino acid sequence information 18 to the RW control unit 41. Furthermore, the receiving unit 40 outputs the terminal ID of the user terminal 11 (not shown) to the screen distribution control unit 47.

[0058] The RW control unit 41 controls the storage and retrieval of various data in the storage device 25A. For example, the RW control unit 41 stores the amino acid sequence information 18 from the receiving unit 40 in the storage device 25A. Furthermore, the RW control unit 41 reads the amino acid sequence information 18 from the storage device 25A and outputs the read amino acid sequence information 18 to the structure information derivation unit 42 and the feature extraction unit 43.

[0059] The RW control unit 41 reads the feature extraction model 36 from the storage device 25A and outputs the read feature extraction model 36 to the feature extraction unit 43. Furthermore, the RW control unit 41 reads the setting auxiliary information 37 from the storage device 25A and outputs the read setting auxiliary information 37 to the weight setting unit 44. Additionally, the RW control unit 41 reads the property information prediction model 38 from the storage device 25A and outputs the read property information prediction model 38 to the prediction unit 46.

[0060] The structure information export unit 42 exports the structure information 50 of the amino acid sequence specific to the antibody 16 based on the amino acid sequence information 18. The structure information export unit 42 outputs the structure information 50 to the weight setting unit 44.

[0061] The feature extraction unit 43 applies the amino acid sequence information 18 to the feature extraction model 36 to extract a set of feature quantities of each amino acid residue constituting the amino acid sequence of the antibody 16, namely the feature quantity group 51. The feature extraction unit 43 outputs the feature quantity group 51 to the weight setting unit 44.

[0062] The weight setting unit 44 refers to the setting auxiliary information 37 and sets the weight of each feature quantity constituting the feature quantity group 51 according to the structure information 50. The weight setting unit 44 registers the set weights in the feature quantity group 51 and uses the feature quantity group 51 as a weighted feature quantity group 51W. The weight setting unit 44 outputs the weighted feature quantity group 51W to the preprocessing unit 45.

[0063] The preprocessing unit 45 performs preprocessing on the features registered in the weighted feature group 51W according to their weights, thereby generating an integrated feature 52. The preprocessing unit 45 outputs the integrated feature 52 to the prediction unit 46.

[0064] The prediction unit 46 applies the integrated feature quantity 52 to the property information prediction model 38, causing the property information prediction model 38 to predict the property information 17. The prediction unit 46 outputs the property information 17 to the screen distribution control unit 47.

[0065] The screen distribution control unit 47 controls the distribution of various screens to the user terminal 11. Specifically, the screen distribution control unit 47 distributes various screens, for example, in the form of screen data created using a markup language such as XML (Extensible Markup Language) for webpage distribution, to the user terminal 11 of the various requesting source. At this time, the screen distribution control unit 47 determines the user terminal 11 of the various requesting source based on the terminal ID from the receiving unit 40. Among the various screens is an information input screen 75 for inputting amino acid sequence information 18 (see reference). Figure 14 ) and the prediction result display screen 80 for displaying property information 17 (reference) Figure 15 Alternatively, it can be used in place of XML, such as JSON (Javascript Object Notation).

[0066] As an example, such as Figure 5 As shown, structural information 50 refers to information on the functional regions of each amino acid residue in the amino acid sequence constituting antibody 16 and the three-dimensional shape (hereinafter referred to as three-dimensional shape) of the three-dimensional structure forming the amino acid sequence. That is, structural information 50 is an example of "regional information" and "three-dimensional shape information" involved in the technology of this invention, and is also an example of "three-dimensional structural information".

[0067] Within the segment area, there are registered Figure 2 The variable structural domains VH and VL, constant structural domains CH1~CH3 and CL, or complementary determinant regions CDR-H1~CDR-H3 and CDR-L1~CDR-L3 are shown. In the three-dimensional shape, either a linear shape or a toroidal shape is registered.

[0068] The structure information extraction unit 42, for example, determines the segment regions of antibodies 16 whose segment regions are unknown by performing multiple alignments with a set of antibodies 16 whose segment regions are known. In addition, the segment regions can be obtained not only by analyzing antibodies 16 using known techniques such as X-ray crystal structure analysis, but also by obtaining them from academic papers or databases such as the Protein Data Bank.

[0069] Furthermore, the structure information export unit 42 inputs, for example, the amino acid sequence information 18 into "AlphaFold2," a protein three-dimensional structure prediction program developed by DeepMind. Then, it outputs the three-dimensional coordinates of each amino acid residue from "AlphaFold2" and determines the three-dimensional shape based on these coordinates.

[0070] As an example, such as Figure 6 As shown, the feature extraction unit 43 inputs the amino acid sequence information 18 into the feature extraction model 36, thereby outputting the feature set 51 from the feature extraction model 36. The feature extraction model 36 is, for example, a neural network encoder or a machine learning model such as a support vector machine (SVM). The learning of the feature extraction model 36 can be performed in the information processing server 10 or in a device other than the information processing server 10. Furthermore, the learning of the feature extraction model 36 can continue after it has been run.

[0071] Among the features, there are multiple types, including feature A, feature B, feature C, ..., feature ZZZ. Multiple feature types are registered for each amino acid residue in the amino acid sequence constituting antibody 16. For example, in amino acid residue "E" of No. 1, ZA1 is registered as feature A, ZB1 as feature B, ZC1 as feature C, ..., and ZZZZ1 as feature ZZZ. Thus, each amino acid residue can be represented by multiple feature types. These multiple feature types can be considered multidimensional vector data. The number of feature types (the dimension of the feature vector) ranges from hundreds to thousands. In addition to machine learning models such as feature extraction model 36, molecular dynamics (MD) methods can also be used to extract feature types.

[0072] As an example, such as Figure 7As shown, the auxiliary information 37 is information that registers weights to be set for each segment region and its combination with the 3D shape. The segment region is divided into two groups: the complementary decision region (CDR) or the region other than the complementary decision region (CDR). These two groups are further subdivided based on whether the 3D shape is linear or toroidal. The weight is 2 when the segment region is the complementary decision region (CDR) and the 3D shape is linear; the weight is 2.5 when the 3D shape is toroidal. The weight is 1 when the segment region is a region other than the complementary decision region (CDR) and the 3D shape is linear; the weight is 1.5 when the 3D shape is toroidal. The weight is higher when the segment region is the complementary decision region (CDR) compared to when the segment region is a region other than the complementary decision region (CDR). Furthermore, the weight is higher when the 3D shape is toroidal compared to when the 3D shape is linear.

[0073] As an example, such as Figure 8 As shown, the weight setting unit 44 refers to the setting auxiliary information 37 and sets weights for each amino acid residue based on the combination of the segment regions and three-dimensional shapes of the structural information 50. More specifically, the weight setting unit 44 sets a higher weight for the characteristic amount of amino acid residues whose segment regions are complementarity-determining regions (CDRs) than for the characteristic amount of amino acid residues whose segment regions are not CDRs. Furthermore, the weight setting unit 44 sets a higher weight for the characteristic amount of amino acid residues with a ring-shaped three-dimensional shape than for the characteristic amount of amino acid residues with a linear three-dimensional shape. For example, the segment region of amino acid residue "Q" in No. 3 is a constant structural domain CH1, and its three-dimensional shape is linear; therefore, the weight setting unit 44 sets a weight of 1. And, for example, the segment region of amino acid residue "S" in No. 100 is a complementarity-determining region CDR-L1, and its three-dimensional shape is ring-shaped; therefore, the weight setting unit 44 sets a weight of 2.5.

[0074] As an example, such as Figure 9 As shown, the weighted feature set 51W is a feature set with weight terms added to the feature set 51.

[0075] As an example, such as Figure 10As shown, in the preprocessing unit 45, as a preprocessing step before applying the feature quantities to the property information prediction model 38, a weighted average is calculated based on the weights of each feature quantity, and this average is used as the integrated feature quantity 52. ​​Specifically, the preprocessing unit 45 calculates the weighted average ZA based on the weights of feature quantity A (ZA1, ZA2, ZA3, ...), the weighted average ZB based on the weights of feature quantity B (ZB1, ZB2, ZB3, ...), the weighted average ZC based on the weights of feature quantity C (ZC1, ZC2, ZC3, ...), and the weighted average ZZZZ based on the weights of feature quantity ZZZ (ZZZ1, ZZZZ2, ZZZZ3, ...). Then, the set of weighted averages ZA, ZB, ZC, ..., ZZZZ is output as the integrated feature quantity 52.

[0076] As an example, such as Figure 11 As shown, the prediction unit 46 inputs the integrated feature quantity 52 into the property information prediction model 38, thereby outputting property information 17 from the property information prediction model 38.

[0077] As an example, such as Figure 12 As shown, the property information prediction model 38 is constructed from a neural network 60. As is well known, the neural network 60 has an input layer 61, intermediate layers (also called hidden layers) 62, and an output layer 63. These input layers 61, intermediate layers 62, and output layers 63 each have multiple nodes ND. Coefficients representing the binding strength of each node ND are set between the nodes ND of the input layer 61 and the nodes ND of the intermediate layer 62, between the nodes ND within the intermediate layer 62, and between the nodes ND of the intermediate layer 62 and the nodes ND of the output layer 63. In the nodes ND of the output layer 63, appropriate activation functions such as linear functions and ReLU (Rectified Linear Unit) functions are set.

[0078] In each node ND of the input layer 61, the integrated feature quantities 52, ZA, ZB, ZC, ..., ZZZZ, are input. Furthermore, property information 17 is output from the node ND of the output layer 63. Additionally, the property information prediction model 38 is not limited to the illustrated neural network 60, but can also be other machine learning models such as decision trees, gradient boosting decision trees, random forests, support vector machines, and Naive Bayes.

[0079] The learning of the property information prediction model 38 can be performed in the information processing server 10 or in a device other than the information processing server 10. Furthermore, the learning of the property information prediction model 38 can continue after it has been run.

[0080] As an example, such as Figure 13As shown, the prediction AP 70 is stored in the storage device 25B of the user terminal 11. The prediction AP 70 is installed on the user terminal 11 by the user U. The prediction AP 70 is an AP used to predict the property information 17 of the antibody 16. When the prediction AP 70 is activated, the CPU 27B of the user terminal 11, together with the memory 26, etc., functions as a browser control unit 72. The browser control unit 72 controls the operation of the dedicated web browser of the prediction AP 70.

[0081] The browser control unit 72 reproduces various screens based on various screen data from the information processing server 10, and displays the reproduced screens on the display 29B. Furthermore, the browser control unit 72 receives various operation commands input by the user U from the input device 30B through these screens. The browser control unit 72 sends various requests based on the operation commands, including prediction requests 15, to the information processing server 10.

[0082] With AP70 prediction enabled, under the control of the browser control unit 72, as an example, such as Figure 14 The information input screen 75 shown is displayed on the monitor 29B. The information input screen 75 includes an input box 76 for amino acid sequence information 18. The input box 76 allows the recording of amino acid sequence information 18 or the dragging and dropping of a file containing amino acid sequence information 18.

[0083] After user U enters the desired amino acid sequence information 18 in input box 76, they select the prediction button 77. When the prediction button 77 is selected, the browser control unit 72 generates a prediction request 15 including the amino acid sequence information 18 entered in input box 76, and sends the generated prediction request 15 to the information processing server 10.

[0084] Furthermore, in the case where the property information 17 is predicted in the information processing server 10, under the control of the browser control unit 72, as an example, Figure 15 The prediction result display screen 80 is displayed on the monitor 29B. The prediction result display screen 80 displays property information 17; more precisely, it displays a message indicating property information 17. Thus, property information 17 is presented to the user U in the form of screen data distribution.

[0085] At the top of the prediction result display screen 80, there is an amino acid sequence information display button 81. When the amino acid sequence information display button 81 is selected, the display screen of amino acid sequence information 18 is displayed in a pop-up window.

[0086] Furthermore, at the bottom of the prediction result display screen 80, there are a save button 82 and an OK button 83. When the save button 82 is selected, the property information 17 and amino acid sequence information 18 are associated with and stored in the storage device 25B of the user terminal 11. When the OK button 83 is selected, the prediction result display screen 80 disappears.

[0087] Next, refer to Figure 16 The flowchart below explains the effects resulting from the above configuration. First, when program 35 starts running in information processing server 10, as... Figure 4 As shown, the CPU 27A of the information processing server 10 functions as a receiving unit 40, an RW control unit 41, a structural information export unit 42, a feature extraction unit 43, a weight setting unit 44, a preprocessing unit 45, a prediction unit 46, and a screen distribution control unit 47. Furthermore, when the AP70 is predicted to start in the user terminal 11, as... Figure 13 As shown, the CPU 27B of the user terminal 11 functions as the browser control unit 72.

[0088] On the display 29B of the user terminal 11, under the control of the browser control unit 72, the following is displayed: Figure 14 The information input screen 75 is shown. In the information input screen 75, when user U enters the desired amino acid sequence information 18 in the input box 76 and selects the prediction button 77, a prediction request 15 is sent from the browser control unit 72 to the information processing server 10. For example... Figure 1 As shown, the prediction request 15 includes amino acid sequence information 18 and the terminal ID of the user terminal 11, etc.

[0089] In the information processing server 10, the prediction request 15 is received by the receiving unit 40 ("Yes" in step ST100). The amino acid sequence information 18 of the prediction request 15 is output from the receiving unit 40 to the RW control unit 41, and stored in the storage device 25A under the control of the RW control unit 41 (step ST110).

[0090] The amino acid sequence information 18 is read from the storage device 25A via the RW control unit 41 (step ST120). The amino acid sequence information 18 is output from the RW control unit 41 to the structure information derivation unit 42 and the feature extraction unit 43.

[0091] In the structure information derivation section 42, based on the amino acid sequence information 18, the structure information is derived. Figure 5 The structural information 50 shown is in step ST130. The structural information 50 is output from the structural information export unit 42 to the weight setting unit 44.

[0092] And, as Figure 6As shown, in the feature extraction unit 43, the amino acid sequence information 18 is input into the feature extraction model 36, thereby outputting the feature set 51 from the feature extraction model 36 (step ST140). The feature set 51 is output from the feature extraction unit 43 to the weight setting unit 44.

[0093] Next, as Figure 8 As shown, in the weight setting unit 44, the weights of the feature quantities are set with reference to the setting auxiliary information 37 and according to the structure information 50 (step ST150). In the weight setting unit 44, a weighted feature quantity group 51W including the set weight items is generated. The weighted feature quantity group 51W is output from the weight setting unit 44 to the preprocessing unit 45.

[0094] like Figure 10 As shown, in the preprocessing unit 45, as preprocessing, a weighted average is calculated based on the weights of each feature quantity, thereby generating an integrated feature quantity 52 (step ST160). The integrated feature quantity 52 is output from the preprocessing unit 45 to the prediction unit 46.

[0095] like Figure 11 As shown, in the prediction unit 46, the integrated feature quantity 52 is input into the property information prediction model 38, thereby outputting property information 17 from the property information prediction model 38 (step ST170). The property information 17 is output from the prediction unit 46 to the screen distribution control unit 47.

[0096] The screen distribution control unit 47 generates screen data for the prediction result display screen 80, which includes a message representing the nature information 17. Under the control of the screen distribution control unit 47, the screen data for the prediction result display screen 80 is distributed to the user terminal 11, the source of the prediction request 15 (step ST180).

[0097] The display 29B on user terminal 11 shows: Figure 15 The prediction results are shown on screen 80. User U views the prediction results on screen 80 and confirms the optimal hydrogen ion index for property information 17.

[0098] As explained above, the CPU27A of the information processing server 10 includes a structure information export unit 42, a feature extraction unit 43, a weight setting unit 44, and a prediction unit 46. The structure information export unit 42 obtains structure information 50 of the amino acid sequence specific to the antibody 16 through export. The feature extraction unit 43 obtains feature quantities by extracting the amino acid residues constituting the amino acid sequence of the antibody 16. The weight setting unit 44 sets the weights of the feature quantities based on the structure information 50. The prediction unit 46 applies the weighted feature quantities to the property information prediction model 38, causing the property information prediction model 38 to predict property information 17 representing the properties of the antibody 16.

[0099] By assigning weights to feature quantities based on the structural information 50 of antibody 16, the feature quantities of important amino acid residues in the structure of antibody 16 that are considered closely related to its properties can be utilized by predicting property information 17. Therefore, property information 17 representing the properties of antibody 16 can be predicted with high accuracy.

[0100] Structural information 50 is information that can be derived regardless of whether antibody 16 is a variant. Therefore, the technology of the present invention, as described in Japanese Patent Application Publication No. 2021-158973, is not limited to variants, but can also be applied to ordinary antibody 16 that is not a variant.

[0101] like Figure 5 As shown, structural information 50 represents the segmental region information indicating the functional segmental region of amino acid residues. Then, the protein is antibody 16, and the segmental region includes the complementarity-determining region (CDR). (As shown...) Figure 7 and Figure 8 As shown, the weight setting unit 44 sets the weight of the characteristic amount of amino acid residues in the segment region that is a complementarity-determining region (CDR) to be higher than the weight of the characteristic amount of amino acid residues in the segment region that is not a complementarity-determining region (CDR).

[0102] As described above, the complementarity-determining region (CDR) is the binding site to the antigen and is the region that determines the function of antibody 16. Furthermore, the CDR is generally exposed to a greater extent in solvents such as the preservation solution of antibody 16. Therefore, by setting the structural information 50 to represent a segment region including the CDR, and by setting the weight of the characteristic amounts of amino acid residues in the segment region that are part of the CDR to be higher than the weight of the characteristic amounts of amino acid residues in the segment region that are not part of the CDR, the property information 17 can be predicted with higher accuracy. This effect is particularly enhanced when the property information 17 is information related to the preservation solution of antibody 16, such as the optimal hydrogen ion index of antibody 16.

[0103] And, as Figure 5 As shown, structural information 50 is the three-dimensional structural information of the amino acid sequence, which represents the three-dimensional shape of the amino acid sequence. The three-dimensional shape includes both ring-shaped and linear shapes. For example... Figure 7 and Figure 8 As shown, the weight setting unit 44 sets the weight of the feature quantity of amino acid residues with a three-dimensional ring shape to be higher than the weight of the feature quantity of amino acid residues with a three-dimensional linear shape.

[0104] When the three-dimensional shape is ring-shaped, the amino acid residue is generally more exposed in the solvent compared to when the three-dimensional shape is linear. Therefore, by setting the structural information 50 to represent the three-dimensional shape of amino acid sequences including both ring-shaped and linear shapes, and by setting the weight of the feature values ​​of amino acid residues with ring-shaped three-dimensional shapes to be higher than the weight of the feature values ​​of amino acid residues with linear three-dimensional shapes, the property information 17 can be predicted with higher accuracy. This effect is particularly enhanced when the property information 17 relates to the preservation solution of antibody 16, such as the optimal hydrogen ion index of antibody 16.

[0105] The pretreatment unit 45 performs pretreatment based on the characteristic amounts of amino acid residues and according to weights. Specifically, such as... Figure 10 As shown, the preprocessing is a weighted average calculated based on the weights of the characteristic quantities of amino acid residues. Therefore, the characteristic quantities reflecting the weights can be applied to the property information prediction model 38.

[0106] The optimal hydrogen ion index is crucial for maintaining the stable quality of antibody 16. Therefore, by using property information 17 as the optimal hydrogen ion index for antibody 16, it is possible to maintain the stable quality of antibody 16, thereby maximizing the therapeutic efficacy of biopharmaceuticals using antibody 16 as the active ingredient.

[0107] Biopharmaceuticals containing antibody 16 as a protein are called antibody drugs, and they are used not only in the treatment of chronic diseases such as cancer, diabetes, and rheumatoid arthritis, but also widely in the treatment of rare diseases such as hemophilia and Crohn's disease. Therefore, this example of using a protein as antibody 16 can help to stably provide antibody drugs for the treatment of a wide range of diseases.

[0108] Three-dimensional structural information is not limited to the illustrated three-dimensional shape information. For example, such as... Figure 17 As shown in structural information 90, it can also be the three-dimensional coordinates of each amino acid residue and the centroid coordinates of antibody 16. The three-dimensional coordinates of each amino acid residue and the centroid coordinates of antibody 16 are obtained by inputting the amino acid sequence information 18 into "AlphaFold2".

[0109] In this case, the weight setting unit 44 refers to, as an example, Figure 18The weights are set using the setting auxiliary information 92 shown. Setting auxiliary information 92 replaces the three-dimensional shape item of setting auxiliary information 37, and includes an item with a distance from the centroid. The distance from the centroid can be calculated based on the three-dimensional coordinates of each amino acid residue and the centroid coordinates of antibody 16. Similar to setting auxiliary information 37, the segment region is divided into two groups: a complementarity-determining region (CDR) or a region excluding the CDR. These two groups are further subdivided based on whether the distance from the centroid is less than or above a threshold. When the segment region is a CDR, the weight is 2 if the distance from the centroid is less than the threshold, and 2.5 if the distance is above the threshold. When the segment region is a region excluding the CDR and the distance from the centroid is less than the threshold, the weight is 1, and 1.5 if the distance is above the threshold. The weight is set higher when the distance from the centroid is above the threshold compared to when the distance is less than the threshold.

[0110] The greater the distance from the centroid, the more the amino acid residue is generally exposed in the solvent. Therefore, by setting the stereostructure information as the three-dimensional coordinates of each amino acid residue and the centroid coordinates of antibody 16, and by setting the weight of the feature values ​​of amino acid residues at a distance of more than a threshold from the centroid to a higher weight than that of amino acid residues at a distance of less than a threshold from the centroid, property information 17 can be predicted with higher accuracy. This effect is particularly enhanced when property information 17 relates to the preservation solution of antibody 16, such as the optimal hydrogen ion index.

[0111] Furthermore, as an example, such as Figure 19 As shown in structural information 95, the three-dimensional structural information can also be the solvent accessible surface area (SASA) of each amino acid residue. The solvent accessible surface area is calculated based on the three-dimensional coordinates of each amino acid residue.

[0112] In this case, the weight setting unit 44 refers to, for example, as... Figure 20 The weights are set using the setting auxiliary information 97 shown. Setting auxiliary information 97 sets the weight to 0 for solvent-accessible surface areas less than a threshold and sets the weight to 1 for solvent-accessible surface areas greater than or equal to the threshold. The characteristic values ​​of amino acid residues with a weight set to 0 are not included in the calculation of the pretreatment weighted average. From this example, it can also be understood that the weights set by the weight setting unit 44 can include 0.

[0113] The larger the solvent-accessible surface area, the more of the amino acid residue is exposed in the solvent. Therefore, by setting the stereostructure information as information about the solvent-accessible surface area, and by assigning a higher weight to the feature values ​​of amino acid residues with a solvent-accessible surface area above a threshold than to the feature values ​​of amino acid residues with a solvent-accessible surface area below a threshold, property information 17 can be predicted with greater accuracy. This effect is particularly enhanced when property information 17 relates to the preservation solution of antibody 16, such as the optimal hydrogen ion index.

[0114] [Second Implementation] As an example, such as Figure 21 As shown, the preprocessing unit 100 of the second embodiment performs preprocessing on the property information prediction model 101 before learning according to the weights, and uses the property information prediction model 101 as the preprocessed property information prediction model 101A. In the property information prediction model 38 of the first embodiment, each feature of the integrated feature quantity 52 is input at each node ND of the input layer 61, but in the property information prediction model 101 of this second embodiment, each feature of the feature quantity group 51 is input at each node ND of the input layer 61.

[0115] In the preprocessing unit 100, as preprocessing, initial coefficients are set between each node ND of the pre-learned property information prediction model 101 according to the weights. Specifically, in order to make the features with relatively high weights more important during processing, the initial coefficients between the nodes ND related to the features with relatively high weights are set to be higher than the coefficients between the nodes ND related to the features with relatively low weights. By learning the preprocessed property information prediction model 101A with the initial coefficients set in this way, the property information prediction model 101X used in the prediction unit 46 is obtained.

[0116] Thus, in the second embodiment, preprocessing involves setting initial coefficients in the property information prediction model 101 according to the weights. Therefore, the preprocessed property information prediction model 101A can be made to converge with a small amount of teacher data, and a property information prediction model 101X with relatively high prediction accuracy can be easily obtained.

[0117] Alternatively, the structure information can be exported from a device different from the information processing server 10 and sent to the information processing server 10 from another device via a network 12 or the like. Similarly, the feature values ​​can also be extracted from a device different from the information processing server 10 and sent to the information processing server 10 from another device via a network 12 or the like.

[0118] As characteristic quantities of amino acid residues, spatial aggregation propensity (SAP) and spatial charge map (SCM) can also be used.

[0119] Property information 17 is not limited to the exemplified optimal hydrogen ion index. It can also include the viscosity of the preservation solution for antibody 16 that best maintains the quality of antibody 16, the type or concentration of additives added to the preservation solution such as L-arginine hydrochloride and refined white sugar. Furthermore, it can also include indicators indicating the degree to which antibody 16 easily aggregates, indicators indicating the hydrophobicity of antibody 16, and the toxicity of antibody 16.

[0120] Proteins are not limited to the example antibody 16. They can also be peptides, nucleic acids, etc. Furthermore, they can be cytokines (interferon, interleukin, etc.) or hormones (insulin, glucagon, follicle-stimulating hormone, erythropoietin, etc.), growth factors (IGF (Insulin-Like Growth Factor)-1, bFGF (Basic Fibroblast Growth Factor), etc.), coagulation factors (Factors 7, 8, 9, etc.), enzymes (lysosomal enzymes, DNA (deoxyribonucleic acid) degrading enzymes, etc.), Fc (Fragment Crystallizable) fusion proteins, receptors, albumins, and protein vaccines. In addition, as antibodies, they also include bispecific antibodies, antibody-drug conjugates, low-molecular-weight antibodies, and glycan-modified antibodies.

[0121] The property information 17 itself can be distributed from the information processing server 10 to the user terminal 11, instead of distributing the screen data of the prediction result display screen 80, which includes the message representing the property information 17, to the user terminal 11.

[0122] The way in which user U can view the property information 17 is not limited to the example prediction result display screen 80. Printed materials containing the property information 17 can be provided to user U, or an email containing the property information 17 can be sent to user U's mobile terminal.

[0123] The information processing server 10 can be located in each pharmaceutical facility or in a data center separate from the pharmaceutical facility. Furthermore, the user terminal 11 can also perform some or all of the functions of each processing unit 40-47 of the information processing server 10.

[0124] The hardware structure of the computer constituting the information processing server 10 according to the technology of this invention can be modified in various ways. For example, to improve processing power and reliability, the information processing server 10 can be composed of multiple computers that are hardware-separated. For example, the functions of the receiving unit 40, RW control unit 41, structure information export unit 42 and feature extraction unit 43, as well as the functions of the weight setting unit 44, preprocessing unit 45, prediction unit 46 and screen distribution control unit 47, can be distributed to two computers. In this case, the information processing server 10 is composed of two computers.

[0125] In this way, the hardware structure of the computer in the information processing server 10 can be appropriately modified according to the performance requirements such as processing power, security, and reliability. Moreover, not only the hardware, but also the APs such as the running program 35 and the prediction AP 70 can be redundantly configured or distributed across multiple storage devices for the purpose of ensuring security and reliability.

[0126] In the above embodiments, for example, as the hardware structure of the processing units that perform various processes, such as the receiving unit 40, the RW control unit 41, the structure information export unit 42, the feature extraction unit 43, the weight setting unit 44, the preprocessing units 45 and 100, the prediction unit 46, the screen distribution control unit 47, and the browser control unit 72, various processors as shown below can be used. Regarding the various processors, as mentioned above, in addition to general-purpose processors such as CPU 27A and 27B that execute software (running program 35 and predicting AP70) and function as various processing units, processors that can have their circuit structure changed after manufacturing, such as FPGA (Field Programmable Gate Array), i.e., Programmable Logic Device (PLD), and processors with circuit structures specially designed for performing specific processes, such as ASIC (Application Specific Integrated Circuit), i.e., dedicated circuits, are also included.

[0127] A processing unit can consist of one of these various processors, or it can consist of a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs and / or a combination of a CPU and an FPGA). Furthermore, a single processor can also constitute multiple processing units.

[0128] As examples of a single processor comprising multiple processing units, firstly, there is the following approach: As exemplified by client and server computers, a processor is composed of a combination of one or more CPUs and software, and this processor functions as multiple processing units. Secondly, there is the following approach: As exemplified by System-on-Chip (SoC), a processor that implements the entire system functionality, including multiple processing units, is used through a single integrated circuit (IC) chip. In this way, various processing units utilize one or more of the aforementioned processors to form the hardware structure.

[0129] Moreover, the hardware architecture of these various processors, more specifically, can utilize circuits composed of circuit elements such as semiconductor components.

[0130] Based on the above records, one can master the techniques described in the following notes.

[0131] [Note 1] An information processing device comprising a processor, The processor performs the following processing: Obtain structural information of the protein-specific amino acid sequence and the characteristic amounts of the amino acid residues constituting the amino acid sequence; The weights of the feature quantities are set according to the structural information; and By applying the feature quantities with the set weights to a machine learning model, the machine learning model can predict property information representing the properties of the protein.

[0132] [Note 2] According to the information processing apparatus described in Appendix 1, wherein, The structural information refers to the segment region information representing the functional segment region of the amino acid residue.

[0133] [Note 3] According to the information processing apparatus described in Appendix 2, wherein, The protein in question is an antibody. The segment region includes a complementarity-determining region. The processor performs the following processing: The weight of the feature quantity of the amino acid residue in the segment region represented by the segment region information that is the complementarity determining region is set to be higher than the weight of the feature quantity of the amino acid residue in the segment region represented by the segment region information that is not the complementarity determining region.

[0134] [Note 4] The information processing apparatus according to any one of appendices 1 to 3, wherein, The structural information refers to the three-dimensional structural information of the amino acid sequence.

[0135] [Note 5] According to the information processing apparatus described in Appendix 4, wherein... The three-dimensional structural information is three-dimensional shape information representing the three-dimensional shape of the amino acid sequence.

[0136] [Note 6] According to the information processing apparatus described in Appendix 5, wherein... The three-dimensional shapes include ring shapes and linear shapes. The processor performs the following processing: The weight of the feature quantity of the amino acid residue whose three-dimensional shape is the ring shape, as represented by the three-dimensional shape information, is set to be higher than the weight of the feature quantity of the amino acid residue whose three-dimensional shape is the linear shape, as represented by the three-dimensional shape information.

[0137] [Note 7] The information processing apparatus according to any one of appendices 1 to 6, wherein, The processor performs the following processing: preprocessing based on the weights for the feature values ​​of the amino acid residues or the machine learning model.

[0138] [Note 8] According to the information processing apparatus described in Appendix 7, wherein... The pretreatment is a weighted average calculated based on the weights of the characteristic quantities of the amino acid residues.

[0139] [Note 9] According to the information processing apparatus described in Appendix 7, wherein... The preprocessing is the process of setting initial coefficients in the machine learning model according to the weights.

[0140] [Note 10] The information processing apparatus according to any one of appendices 1 to 9, wherein, The property information refers to the optimal hydrogen ion index of the protein.

[0141] [Note 11] The information processing apparatus according to any one of appendices 1 to 10, wherein, The protein in question is an antibody.

[0142] The technology of the present invention can also be appropriately combined with the various embodiments and / or variations described above. Furthermore, it is not limited to the embodiments described above; various structures can be adopted as long as they do not depart from the spirit of the invention. Moreover, the technology of the present invention relates not only to programs, but also to storage media for non-transitory storage of programs and computer program products including programs.

[0143] The foregoing descriptions and illustrations are detailed explanations of the parts involved in the technology of this invention, and are merely one example of the technology of this invention. For example, the descriptions related to the above-described structure, function, effect, and effect are examples of the structure, function, effect, and effect of the parts involved in the technology of this invention. Therefore, without departing from the technical spirit of this invention, unnecessary parts can be deleted from the foregoing descriptions and illustrations, or new elements can be added or replaced. Furthermore, to avoid complications and to facilitate understanding of the parts involved in the technology of this invention, descriptions related to technical common sense that are not particularly necessary to explain in terms of how the technology of this invention can be implemented have been omitted from the foregoing descriptions and illustrations.

[0144] In this specification, "A and / or B" has the same meaning as "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B.

[0145] Furthermore, in this specification, when more than three items are connected by "and / or", the same consideration as "A and / or B" can be applied.

[0146] All documents, patent applications and technical standards described in this specification are incorporated herein by reference to the same extent as the specific documents, patent applications and technical standards described therein.

Claims

1. An information processing device comprising a processor, The processor performs the following processing: Obtain structural information of the protein-specific amino acid sequence and the characteristic amounts of the amino acid residues constituting the amino acid sequence; The weights of the feature quantities are set according to the structural information; and By applying the feature quantities with the set weights to a machine learning model, the machine learning model can predict property information representing the properties of the protein.

2. The information processing apparatus according to claim 1, wherein, The structural information refers to the segment region information representing the functional segment region of the amino acid residue.

3. The information processing apparatus according to claim 2, wherein, The protein in question is an antibody. The segment region includes a complementarity-determining region. The processor performs the following processing: The weight of the feature quantity of the amino acid residue in the segment region represented by the segment region information that is the complementarity determining region is set to be higher than the weight of the feature quantity of the amino acid residue in the segment region represented by the segment region information that is not the complementarity determining region.

4. The information processing apparatus according to claim 1, wherein, The structural information refers to the three-dimensional structural information of the amino acid sequence.

5. The information processing apparatus according to claim 4, wherein, The three-dimensional structural information is three-dimensional shape information representing the three-dimensional shape of the amino acid sequence.

6. The information processing apparatus according to claim 5, wherein, The three-dimensional shapes include ring shapes and linear shapes. The processor performs the following processing: The weight of the feature quantity of the amino acid residue whose three-dimensional shape is the ring shape, as represented by the three-dimensional shape information, is set to be higher than the weight of the feature quantity of the amino acid residue whose three-dimensional shape is the linear shape, as represented by the three-dimensional shape information.

7. The information processing apparatus according to claim 1, wherein, The processor performs the following processing: for the feature values ​​of the amino acid residues or the machine learning model, it performs preprocessing corresponding to the weights.

8. The information processing apparatus according to claim 7, wherein, The pretreatment is a weighted average calculated based on the weights of the characteristic quantities of the amino acid residues.

9. The information processing apparatus according to claim 7, wherein, The preprocessing involves setting initial coefficients for the machine learning model that correspond to the weights.

10. The information processing apparatus according to claim 1, wherein, The property information refers to the optimal hydrogen ion index of the protein.

11. The information processing apparatus according to claim 1, wherein, The protein in question is an antibody.

12. A method for operating an information processing device, comprising the following steps: Obtain structural information of the protein-specific amino acid sequence and the characteristic amounts of the amino acid residues constituting the amino acid sequence; The weights of the feature quantities are set according to the structural information; and By applying the feature quantities with the set weights to a machine learning model, the machine learning model can predict property information representing the properties of the protein.

13. An operating program for an information processing apparatus, the operating program of the information processing apparatus causing a computer to perform processing including the following steps: Obtain structural information of the protein-specific amino acid sequence and the characteristic amounts of the amino acid residues constituting the amino acid sequence; The weights of the feature quantities are set according to the structural information; and By applying the feature quantities with the set weights to a machine learning model, the machine learning model can predict property information representing the properties of the protein.

Citation Information

Patent Citations

  • Model for predicting thermal stability of mutant and use thereof

    JP2021158973A