Information processing systems, information processing methods, and programs
The information processing system accurately predicts lipid binding affinity of proteins by converting amino acid sequences into feature vectors and applying regression models, addressing the challenge of inaccurate lipid binding estimation in existing systems and facilitating the design of proteins with desired lipid binding properties.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- THE UNIV OF TOKYO
- Filing Date
- 2024-10-03
- Publication Date
- 2026-04-15
AI Technical Summary
Existing systems lack the ability to accurately estimate the binding ability of proteins to lipids.
An information processing system that utilizes a machine learning model to predict lipid binding affinity by converting amino acid sequences into feature vectors and applying regression models to estimate lipid binding strength, allowing for the identification of amino acid sequences with high lipid binding affinity.
The system achieves high accuracy in estimating lipid binding affinity of proteins, enabling the design of proteins with targeted lipid binding properties.
Smart Images

Figure 2026065368000001_ABST
Abstract
Description
Technical Field
[0004] ,
[0006] , , , , , ,
[0005] , , ,
[0007] , , ,
[0001] The present invention relates to an information processing system, an information processing method, and a program.
Background Art
[0002] Patent Document 1 discloses a system that searches for an amino acid sequence of a protein using a machine learning model.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the above system, there is no guarantee that the binding ability to lipids can be accurately estimated as a property of the protein.
[0005] In view of the above circumstances, the present invention aims to provide an information processing system and the like that can accurately estimate the binding ability to lipids.
Means for Solving the Problems
[0008] [Figure 1] This is a diagram showing the configuration of Information Processing System 1. [Figure 2] This is a block diagram showing the hardware configuration of the information processing device 10. [Figure 3] This is a block diagram showing the hardware configuration of user terminal 20. [Figure 4] This is a block diagram showing the functions realized by the information processing device 10 (control unit 11) and the user terminal 20 (control unit 21). [Figure 5] This is a schematic diagram illustrating an example of the process for estimating lipid binding strength by the estimation unit 112. [Figure 6] This graph shows the relationship between the binding affinity to phospholipids predicted by the prediction unit 112 and the measured value of that binding affinity for multiple proteins (amino acid sequences). [Figure 7] This graph shows the binding affinity to specific lipids between a protein with a randomly mutated amino acid sequence (nanobody N1) and a protein with an amino acid sequence searched by the search unit 113 (nanobody N2). [Figure 8] This graph shows the binding strengths of four proteins ("Ev44", "Ev53", "Ev56", and "Ev66") that have specific binding affinity and were searched for by the search unit 113. [Figure 9] This is an activity diagram showing an example of the flow of information processing (amino acid sequence search processing) performed by information processing system 1. [Modes for carrying out the invention]
[0009] Embodiments of the present invention will be described below with reference to the drawings. The various features shown in the embodiments below can be combined with each other.
[0010] Incidentally, the program for implementing the software appearing in one embodiment may be provided as a non-transitory computer-readable medium, or it may be provided as a downloadable medium from an external server, or it may be provided so that the program is launched on an external computer and its functions are realized on a client terminal (so-called cloud computing).
[0011] Furthermore, in various information processing according to one embodiment, an input and an output corresponding to the input can be realized. Here, as long as an output is obtained as a result of the input, the form of the information referenced in such information processing (hereinafter referred to as "reference information") is not limited. The reference information may be, for example, rule-based information such as a database, a lookup table, or a predetermined function (including a decision formula such as a regression equation constructed by a statistical method), or a pre-trained model that has learned the correlation between input and output in advance, or a large-scale language model that can output a desired result by inputting a prompt.
[0012] Furthermore, in one embodiment, "part" may include, for example, hardware resources implemented by a circuit in a broad sense, and the information processing of software that can be specifically realized by these hardware resources. Also, in one embodiment, various types of information are handled, and this information can be represented, for example, by the physical values of signal values representing voltage and current, the high or low values of signal values as a set of binary bits composed of 0s or 1s, or by quantum superposition (so-called qubits), and communication and calculations can be performed on a circuit in a broad sense.
[0013] Furthermore, a circuit in a broad sense is a circuit realized by appropriately combining at least a circuit, circuitry, a processor, a memory, etc. The processor may be a general-purpose processor or a dedicated circuit. That is, it includes an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.
[0014] 1. Hardware Configuration In this section, the hardware configuration will be described.
[0015] <Information Processing System 1> FIG. 1 is a configuration diagram showing an information processing system 1. The information processing system 1 includes a communication line 2, an information processing device 10, and a plurality of user terminals 20. The information processing device 10 and the user terminals 20 are configured to be communicable through the communication line 2. The connection between the information processing device 10 and the user terminals 20 may be wired or wireless.
[0016] In one embodiment of the information processing system 1, the information processing system 1 consists of one or more devices or components. Hereinafter, these components will be described.
[0017] <Information Processing Device 10> FIG. 2 is a block diagram showing the hardware configuration of the information processing apparatus 10. As shown in FIG. 2, the information processing apparatus 10 includes a control unit 11, a storage unit 12, a communication unit 13, and a communication bus 14. The control unit 11, the storage unit 12, and the communication unit 13 are electrically connected inside the information processing apparatus 10 via the communication bus 14.
[0018] <Control Unit 11> The control unit 11 performs processing and control of the overall operations related to the information processing apparatus 10. The control unit 11 is, for example, a Central Processing Unit (CPU). The control unit 11 realizes various functions related to the information processing apparatus 10 by reading a predetermined program stored in the storage unit 12. That is, the information processing by software stored in the storage unit 12 is specifically realized by the control unit 11, which is an example of hardware, and can be executed as each functional unit included in the control unit 11. These will be described in more detail in the next section. Note that the control unit 11 is not limited to being single, and the information processing apparatus 10 may have a plurality of control units 11 for each function. Further, the information processing apparatus 10 may have a configuration combining these.
[0019] <Storage Unit 12> The storage unit 12 stores various information defined by the foregoing description. This can be implemented, for example, as a storage device such as a Solid State Drive (SSD) that stores various programs and the like related to the information processing apparatus 10 executed by the control unit 11, or as a memory such as a Random Access Memory (RAM) that stores temporarily necessary information (arguments, arrays, etc.) related to the calculation of programs. The storage unit 12 stores various programs, variables, etc. related to the information processing apparatus 10 executed by the control unit 11.
[0020] <Communication Unit 13> The communication unit 13 preferably uses wired communication methods such as USB, IEEE1394, Thunderbolt®, and wired LAN network communication, but may also include wireless LAN network communication, mobile communication such as LTE / 5G, and Bluetooth® communication as needed. In other words, it is more preferable to implement it as a collection of these multiple communication methods. That is, the information processing device 10 may communicate various information from the outside via the communication unit 13 and the network.
[0021] The information processing device 10 may be on-premise or in a cloud-based configuration. In the case of a cloud-based information processing device 10, for example, the above-mentioned functions and processing may be provided in the form of SaaS (Software as a Service) or cloud computing.
[0022] <User terminal 20> The user terminal 20 is an information processing terminal used by a user who utilizes the information processing functions provided by the information processing device 10. Figure 3 is a block diagram showing the hardware configuration of the user terminal 20. As shown in Figure 3, the user terminal 20 comprises a control unit 21, a storage unit 22, a communication unit 23, an input unit 24, an output unit 25, and a communication bus 26. The control unit 21, storage unit 22, communication unit 23, input unit 24, and output unit 25 are electrically connected within the user terminal 20 via the communication bus 26. The descriptions of the control unit 21, storage unit 22, and communication unit 23 are the same as the descriptions of each part of the information processing device 10 and are therefore omitted.
[0023] <Input section 24> The input unit 24 receives action input from the user. The action input is transmitted as a command signal to the control unit 21 via the communication bus 26. The control unit 21 can perform predetermined controls or calculations based on the transmitted command signal as needed. The input unit 24 may be included in the casing of the user terminal 20 or it may be external. For example, the input unit 24 may be implemented as a touch panel integrated with the output unit 25. When the input unit 24 is implemented as a touch panel, the user can input tap actions, swipe actions, etc. to the input unit 24. Instead of a touch panel, the input unit 24 can be a switch button, mouse, trackpad, QWERTY keyboard, etc.
[0024] <Output section 25> The output unit 25 displays a graphical user interface (GUI) screen that allows the user to take action. The output unit 25 may be included in the casing of the user terminal 20 or it may be an external component. Specifically, the output unit 25 can be implemented as a display device such as a CRT display, liquid crystal display, organic EL display, or plasma display. It is preferable that these display devices be used in accordance with the type of user terminal 20.
[0025] 2. Functional Configuration This section describes the functional configuration of this embodiment. Information processing by software stored in the memory unit 12 is specifically realized by the control unit 11, which is an example of hardware, and can be executed as each functional unit included in the control unit 11 (the processor provided by the information processing system 1).
[0026] Figure 4 is a block diagram showing the functions realized by the information processing device 10 (control unit 11) and the user terminal 20 (control unit 21).
[0027] As shown in Figure 4A, the information processing device 10 (control unit 11) comprises an acquisition unit 111, an inference unit 112, a search unit 113, and an artificial intelligence unit 120. As shown in Figure 4B, the user terminal 20 (control unit 21) comprises a display unit 211 and an operation acquisition unit 212.
[0028] The acquisition unit 111 is configured to acquire the amino acid sequence to be predicted (the amino acid sequence of a protein whose binding affinity to lipids is predicted). The acquisition unit 111 acquires the amino acid sequence to be predicted by, for example, receiving input of the amino acid sequence from the user terminal 20. Alternatively, the acquisition unit 111 may acquire the amino acid sequence to be predicted from a storage location (network address) of the amino acid sequence data, after receiving a specification of the storage location (network address) from the user terminal 20. Furthermore, the acquisition unit 111 may acquire a mutant amino acid sequence created by the search unit 113, described later, as the amino acid sequence to be predicted.
[0029] The number of amino acids in each amino acid sequence acquired by the acquisition unit 111 (the number of amino acids in one sequence) is not particularly limited, and is, for example, 10 or more and 3000 or less.
[0030] <Guessing part 112> The prediction unit 112 is configured to predict the lipid binding affinity (hereinafter referred to as "lipid binding affinity") of a protein having the amino acid sequence to be predicted (hereinafter referred to as "predicted protein") based on the amino acid sequence to be predicted and reference information acquired by the acquisition unit 111.
[0031] Lipid binding strength is typically a value that indicates the strength (activity) of a lipid's binding to a particular lipid. Alternatively, lipid binding strength may also indicate the strength of the binding to a lipid group, which is a grouping of lipids. Lipid groups can include, for example, those grouped by chemical structure (simple lipids, complex lipids, derived lipids), those grouped by fatty acid type, and those grouped by the polar head of the lipid molecule. Furthermore, "strength of binding strength" may include statistical values such as maximum, minimum, mean, and median values.
[0032] Examples of proteins to be predicted include antibody mimetics such as nanobodies, monobodies, and affibodies, antibodies, proteins contained in the lipid-binding domain region of naturally occurring proteins, and artificial proteins designed using computers. The prediction unit 112 is preferably configured to predict the lipid-binding affinity of nanobodies. This makes it possible to design or search for nanobodies with high binding affinity to specific lipids.
[0033] The reference information includes the correlation between amino acid sequences and lipid binding strength, constructed based on measured values of protein binding strength to lipids. The reference information is stored, for example, in the memory unit 12. The reference information is an estimator constructed to output lipid binding strength from an amino acid sequence. The reference information may also include, for example, tables, functions, simple algorithms, etc., that show the correlation between feature quantities (e.g., vector data) obtained by information extraction from or information transformation of amino acid sequences and lipid binding strength. The correlations included in the reference information can be constructed, for example, by statistically analyzing data recording lipid binding strength measured from actual proteins.
[0034] The reference information may include a binding force prediction model that has been pre-trained to take an amino acid sequence as input and output lipid binding force. In this case, the prediction unit 112 inputs the amino acid sequence into the binding force prediction model included in the artificial intelligence unit 120, and causes the binding force prediction model to output the lipid binding force. The binding force prediction model is constructed, for example, by machine learning using records that associate the amino acid sequence of a certain protein with the measured value of the lipid binding force of this protein as training data. In the binding force prediction model, parameters calculated and tuned through learning constitute the correlation of the reference information.
[0035] The reference information may include first reference information and second reference information. The first reference information includes the correlation between the amino acid sequence and the feature vector. The second reference information includes the correlation between the feature vector and the lipid binding strength. In this case, the prediction unit 112 obtains the feature vector based on the amino acid sequence to be predicted and the first reference information, and predicts the lipid binding strength based on the obtained feature vector and the second reference information. This allows for the construction of the first reference information for obtaining the feature vector and the construction of the second reference information for predicting the lipid binding strength from the feature vector to be separated, thereby improving the accuracy of lipid binding strength prediction and reducing costs in the construction and updating of the information processing system 1.
[0036] <1st reference information> The first reference information includes, for example, tables, functions, simple algorithms, etc., for obtaining feature vectors by extracting information from amino acid sequences or by transforming information from amino acid sequences. The "correlation between amino acid sequences and feature vectors" in the first reference information may include not only relationships that directly represent the relationship between the two, but also indirect relationships that mathematically link the two (for example, laws that mechanically convert amino acid sequences into feature vectors, or parameters that define such laws).
[0037] Furthermore, the first reference information may include a deep learning model trained to generate feature vectors representing protein characteristics from amino acid sequences. In this case, the prediction unit 112 inputs the amino acid sequence to be predicted into the deep learning model of the artificial intelligence unit 120 and obtains feature vectors from the deep learning model. This makes it possible to obtain feature vectors of the amino acid sequence to be predicted based on pattern analysis of a large amount of amino acid sequences. As a result, the accuracy of lipid binding strength prediction is improved.
[0038] The "protein features" output by deep learning models include, for example, secondary structure, tertiary structure, structural stability, and hydrophobicity. Deep learning models generate feature vectors from amino acid sequences that directly or indirectly represent these features.
[0039] The feature vectors obtained from a deep learning model may be the output information of the deep learning model itself, or they may be intermediate information generated as a byproduct when the deep learning model generates the output information. The dimension of the feature vectors may be on the order of, for example, 1000 dimensions.
[0040] As for deep learning models, for example, so-called protein language models (pLMs) can be used. Examples of pLMs include UniRep® and ESM.
[0041] A typical deep learning model uses a large-scale protein amino acid sequence database as training data, and the parameters of the neural network are tuned to predict the next amino acid to be placed in a given amino acid sequence (or the amino acid to be placed in a masked portion of the sequence). These tuned parameters correspond to the "correlation between amino acid sequences and feature vectors" in the first reference information.
[0042] In such deep learning models, the input amino acid sequence is converted into a high-dimensional feature vector in order to obtain the desired output (predicted amino acid sequence). In other words, the deep learning model generates an embedding vector from the amino acid sequence. More specifically, the deep learning model uses, for example, a Transformer (encoder) to embed the amino acid sequence into a feature vector, and then generates the predicted amino acid sequence from the feature vector.
[0043] <Second reference information> The second reference information includes, for example, tables, functions, and simple algorithms that define the correlation between feature vectors and lipid binding forces.
[0044] The second reference information should ideally be constructed using machine learning data (training data) that combines feature vectors obtained from the amino acid sequence of the training protein with measured values of lipid binding strength in the training protein. This allows for the construction of a second reference information with relatively low cost and high prediction accuracy.
[0045] Machine learning data is prepared, for example, using the following procedure: The binding affinity of the data acquisition protein prepared for training is measured using an appropriate assay to target lipids. The amino acid sequence of the data acquisition protein is converted into a feature vector using first reference information (e.g., a deep learning model), and this feature vector is linked to the measured binding affinity. This results in machine learning data that combines the input data (feature vectors) with the ground truth data (binding affinity). Using this machine learning data, the parameters of the regression model (regression equation or algorithm) included in the second reference information are adjusted so that when a feature vector is input, an estimated value of the lipid binding affinity of the corresponding protein is output.
[0046] Examples of regression equations constructed using machine learning data include linear regression equations. Examples of algorithms constructed using machine learning data include random forests and XGBoost.
[0047] The second reference information may include a deep learning model trained to infer lipid binding forces from feature vectors. In this case, the inference unit 112 inputs the feature vectors to the deep learning model of the artificial intelligence unit 120 and obtains an estimated value of the lipid binding force from the deep learning model.
[0048] The data used to construct the second reference information (measured values of lipid binding strength) is significantly less than the data used to construct the first reference information (protein amino acid sequence data). For example, the number of data points used to construct the first reference information is on the order of 100,000 to 10 billion, while the number of data points used to construct the second reference information is on the order of 10 to 100,000.
[0049] Figure 5 is a schematic diagram showing an example of the lipid binding strength estimation process by the estimation unit 112. First, the estimation unit 112 inputs the amino acid sequence AC to be estimated into the deep learning model PM1. Next, the estimation unit 112 inputs the feature vector CV generated by the deep learning model into the regression model PM2. The estimation unit 112 obtains the lipid binding strength LB output by the regression model PM2 as the estimation result.
[0050] The prediction unit 112 may predict the binding affinity of the target protein to each of several types of lipids, or to each of several lipid groups. This makes it possible to simultaneously predict the binding affinity to each of several types of lipids for a single amino acid sequence input.
[0051] For example, the prediction unit 112 converts the amino acid sequence into a feature vector by inputting it into a deep learning model, and inputs this feature vector into multiple regression models prepared for each type of lipid to obtain an estimated value of the lipid binding strength for each type of lipid. Alternatively, the prediction unit 112 may accept the selection of lipids for which the binding strength is to be estimated from the user terminal 20, and input the feature vector into a regression model that estimates the binding strength to the selected lipids.
[0052] As described above, the estimation unit 112 estimates lipid binding strength using a common first reference information that is independent of the type of lipid, and multiple second reference information prepared for each type of lipid. Each of the multiple second reference information is constructed based on data of the lipid binding strength of proteins measured for each type of lipid. By making the first reference information common to all types of lipids in this way, a model can be constructed (added) to estimate the binding strength of new lipids for which the binding strength had not been estimated before, by constructing only the second reference information without having to reconstruct the first reference information. Therefore, the construction of estimation models for new lipids can be carried out relatively easily and at low cost.
[0053] Figure 6 is a graph showing the relationship between the phospholipid binding strength predicted by the prediction unit 112 and the measured values of that binding strength for several proteins (amino acid sequences). The horizontal axis of the four graphs in Figure 6 shows the predicted values of the binding strength to PI(3,5)P2 (phosphatidylinositol 3,5-bisphosphate), PI(3,4)P2 (phosphatidylinositol 3,4-bisphosphate), PI(4,5)P2 (phosphatidylinositol 4,5-bisphosphate), and PS (phosphatidylserine), respectively, and the vertical axis shows the measured values of the binding strength to each phosphate. As shown in each graph in Figure 6, a high correlation is observed between the predicted values of lipid binding strength by the prediction unit 112 and the measured values of lipid binding strength.
[0054] <Search section 113> The search unit 113 is configured to search for amino acid sequences with high lipid binding affinity by mutating the amino acid sequence to be predicted while referring to the lipid binding affinity predicted by the prediction unit 112.
[0055] Specifically, the search unit 113 first causes the prediction unit 112 to predict the lipid binding strength of a mutant amino acid sequence, which is a mutation of a portion of the initial amino acid sequence (the amino acid sequence input by the user terminal 20). The mutation of the amino acid sequence is performed, for example, by randomly replacing several randomly selected amino acids in the amino acid sequence with other amino acids.
[0056] Next, the search unit 113 records the mutated amino acid sequence and the estimated lipid binding strength. The search unit 113 also determines whether to maintain the preceding mutation in the mutated amino acid sequence based on the estimated lipid binding strength. Specifically, the search unit 113 maintains the mutation if the lipid binding strength satisfies predetermined conditions. The predetermined conditions for maintaining the mutation are set, for example, by the difference between the lipid binding strength of the amino acid sequence before the mutation and the lipid binding strength of the amino acid sequence after the mutation. Specifically, the mutation is maintained if the value obtained by subtracting the lipid binding strength of the amino acid sequence after the mutation from the lipid binding strength of the amino acid sequence before the mutation is greater than or equal to a threshold. Such a threshold may be a predetermined fixed value, but it is preferable that it be a value (function) that gradually increases according to the depth of the search (number of mutations). In the initial stages of the search, the threshold may be a negative value. This makes it possible to search for the maximum value of lipid binding strength in the entire search region, rather than the so-called local maximum value of lipid binding strength.
[0057] If the mutation is maintained (i.e., lipid binding affinity is improved or maintained), the search unit 113 introduces another mutation into the mutated amino acid sequence. If the mutation is not maintained (i.e., lipid binding affinity is reduced), the search unit 113 cancels the previous mutation in the mutated amino acid sequence (i.e., returns the substituted amino acid to the amino acid before the mutation) and introduces another mutation into the mutated amino acid sequence. After introducing the new mutation, the search unit 113 again has the prediction unit 112 predict the lipid binding affinity of the mutated amino acid sequence.
[0058] The search unit 113 repeats the process of introducing mutations into the amino acid sequence and the process of predicting lipid binding strength until the search termination conditions are met. Examples of search termination conditions include obtaining an amino acid sequence with a lipid binding strength exceeding a predetermined target value, reaching a predetermined number of searches (number of mutations), or the elapsed of a predetermined time.
[0059] Figure 7 is a graph showing the binding affinity to specific lipids between a protein with a randomly mutated amino acid sequence (nanobody N1) and a protein with an amino acid sequence searched by the search unit 113 (nanobody N2). The vertical axis of Figure 7 shows the estimated value of lipid binding affinity by the prediction unit 112, and the horizontal axis shows the distance from known nanobodies (degree of mutation). As shown in Figure 7, some nanobodies N2 searched by the search unit 113 have lipid binding affinity that cannot be reached by randomly mutated nanobodies N1. Furthermore, nanobodies N2 searched by the search unit 113 have novel sequences that are far removed from similar sequences of known nanobodies.
[0060] Furthermore, the search unit 113 can search for amino acid sequences that exhibit specific binding affinity, such as binding to a target lipid (e.g., PI(3,5)P2) but not to other lipids (e.g., PS). In other words, the search unit 113 can search for proteins that have high binding affinity to the target lipid and low binding affinity to other lipids. In contrast, random mutations as described above make it difficult to obtain amino acid sequences that exhibit such specific binding affinity.
[0061] Figure 8 is a graph showing the binding affinity of four proteins ("Ev44", "Ev53", "Ev56", and "Ev66") that exhibit specific binding affinity, as identified by the search unit 113. The vertical axis of the four graphs in Figure 8 represents the magnitude of lipid binding affinity. As shown in Figure 8, Ev44 does not bind to PS but has high binding affinity to PI(3,4)P2. Ev53 also does not bind to PS but has high binding affinity to PI(3,5)P2. Ev56 has low binding affinity to PS but high binding affinity to PI(3,5)P2. Ev66, like Ev53, does not bind to PS but has high binding affinity to PI(3,5)P2. However, Ev66 has lower binding affinity to PI(4,5)P2 and PI(3,4)P2 than Ev53, and its specificity to PI(3,5)P2 is stronger than that of Ev53.
[0062] <Artificial Intelligence Department 120> The artificial intelligence unit 120 is configured to receive input from each functional unit and return the instructed output. The artificial intelligence used by each functional unit of the information processing device 10 may be common to all units, or it may be prepared individually for each functional unit.
[0063] The artificial intelligence unit 120 may include learning models such as a Transformer or a Recurrent Neural Network (RNN), including a generative AI.
[0064] Furthermore, specific machine learning algorithms used to construct learning models include nearest neighbors, naive Bayes, decision trees, support vector machines, and deep learning using neural networks. The artificial intelligence unit 120 can apply the above algorithms as appropriate.
[0065] The artificial intelligence unit 120 may have a trained model constructed by a learning method such as supervised learning, unsupervised learning, or self-supervised learning. In supervised learning, machine learning is performed using training data (machine learning data). Training data consists of pairs of input data and output data (correct answer data) for training. Furthermore, the language model may not only be one trained for a specific task, but also a general-purpose model that can be used universally for a wide range of tasks. The trained model included in the artificial intelligence unit 120 can undergo additional training as transfer learning or fine-tuning.
[0066] <Display section 211> The display unit 211 of the user terminal 20 is configured to display the screen indicated by the screen data transmitted from the information processing device 10 on the output unit 25.
[0067] <Operation acquisition unit 212> The operation acquisition unit 212 of the user terminal 20 is configured to accept operations from the user of the user terminal 20.
[0068] 3. Information Processing Methods This section describes the information processing method of the information processing device 10. In this information processing method, each part of the information processing device 10 is executed by a computer as a step.
[0069] Specifically, this information processing method comprises an acquisition step, an inference step, and a search step. In the acquisition step, the amino acid sequence to be inferred is acquired. In the inference step, the lipid binding affinity of a protein having the amino acid sequence to be inferred is inferred based on the amino acid sequence to be inferred and reference information. In the search step, amino acid sequences with high lipid binding affinity are searched for by mutating the amino acid sequence to be inferred, while referring to the lipid binding affinity.
[0070] Figure 9 is an activity diagram showing an example of the flow of information processing (amino acid sequence search processing) performed by information processing system 1. The information processing will be explained below in accordance with each activity in this activity diagram.
[0071] The amino acid sequence search process begins with the user inputting an amino acid sequence. The user inputs the amino acid sequence on the user terminal 20 (Activity A110). The information processing device 10 acquires the amino acid sequence input on the user terminal 20 as the initial amino acid sequence (Activity A120).
[0072] After obtaining the initial amino acid sequence, the information processing device 10 estimates the lipid binding affinity of the initial amino acid sequence (Activity A130). Subsequently, the information processing device 10 creates a mutant amino acid sequence by mutating a part of the initial amino acid sequence (Activity A140). After creating the mutant amino acid sequence, the information processing device 10 estimates the lipid binding affinity of the mutant amino acid sequence (Activity A150).
[0073] After estimating the lipid binding affinity of the mutated amino acid sequence, the information processing device 10 determines whether the termination conditions for the amino acid sequence search have been met (Activity A160). If the termination conditions are met, the information processing device 10 outputs the search results (the searched amino acid sequence and its lipid binding affinity) to the user terminal 20 (Activity A170). As a result, the search results are displayed on the user terminal 20 (Activity A180).
[0074] On the other hand, if the termination condition is not met in Activity A160, the information processing device 10 determines whether the lipid binding strength of the immediately predicted mutated amino acid sequence has improved compared to the lipid binding strength of the amino acid sequence immediately before the mutation (Activity A190). If the lipid binding strength has improved or is maintained (for example, if the lipid binding strength is above a threshold), the information processing device 10 maintains the immediately preceding mutation (Activity A200). On the other hand, if the lipid binding strength has decreased (for example, if the lipid binding strength is below a threshold), the information processing device 10 cancels the immediately preceding mutation (Activity A210).
[0075] After deciding whether to maintain or cancel the mutation, the information processing device 10, in activity A140, creates a new mutant amino acid sequence by adding another mutation to a portion of the amino acid sequence in which the mutation was maintained or canceled. If the mutation is maintained, yet another mutation is added to the mutant amino acid sequence. The information processing device 10, in activity A150, estimates the lipid binding affinity of the newly created mutant amino acid sequence. The information processing device 10 repeats activities A140 and A150, and either activity A200 or activity A210, until the search termination conditions are met.
[0076] 4. Effect The function of this embodiment can be summarized as follows: By specifying an amino acid sequence, the binding affinity of a protein having that amino acid sequence to lipids can be estimated with high accuracy.
[0077] Although embodiments of the present invention have been described above, the present invention is not limited thereto and can be modified as appropriate without departing from the technical spirit of the invention.
[0078] 5. Others In the above embodiment, the information processing device 10 performed various storage and control functions, but instead of the information processing device 10, multiple external devices may be used. That is, various information and programs may be distributed and stored across multiple external devices using blockchain technology or the like.
[0079] The embodiments of this model are not limited to the information processing system 1, but may also be an information processing method or a program. The information processing method comprises each step executed by the information processing system 1. The program causes a computer to execute each step of the information processing system 1.
[0080] At least one of the devices included in the information processing system 1 may be located outside the country in which the functions of the information processing system 1 are performed.
[0081] The information processing system 1 may consist only of the information processing device 10. In other words, the information processing system 1 does not necessarily have to include a user terminal 20.
[0082] The information processing system 1 does not necessarily have to include a search unit 113. For example, the information processing system 1 may only have the function of estimating the lipid binding affinity of a specified amino acid sequence.
[0083] The product may be provided in any of the following embodiments.
[0084] (1) An information processing system comprising a processor configured to perform the following steps by reading a program, wherein in the acquisition step, an amino acid sequence to be predicted is acquired, and in the prediction step, the binding force of a protein having the amino acid sequence to be predicted to a lipid is predicted based on the amino acid sequence to be predicted and reference information, wherein the reference information includes a correlation between the amino acid sequence and the binding force, constructed based on measured values of the binding force.
[0085] (2) An information processing system as described in (1) above, wherein the reference information includes first reference information and second reference information, the first reference information includes a correlation between an amino acid sequence and a feature vector, the second reference information includes a correlation between the feature vector and the binding force, and in the estimation step, the system obtains the feature vector based on the amino acid sequence to be estimated and the first reference information, and estimates the binding force based on the feature vector and the second reference information.
[0086] (3) An information processing system described in (2) above, wherein the second reference information is constructed using machine learning data that combines the feature vector obtained from the amino acid sequence of the learning protein and the measured value of the binding force in the learning protein.
[0087] (4) An information processing system according to (2) or (3) above, wherein the first reference information includes a deep learning model trained to generate a feature vector representing the characteristics of a protein from an amino acid sequence, and in the prediction step, the amino acid sequence to be predicted is input to the deep learning model, and the feature vector is obtained from the deep learning model.
[0088] (5) An information processing system according to any one of (1) to (4) above, wherein the estimation step estimates the binding affinity of the protein to each of the multiple types of lipids or each of the multiple lipid groups.
[0089] (6) An information processing method comprising each step performed by the information processing system described in any one of (1) to (5) above.
[0090] (7) A program that causes a computer to perform each step of the information processing system described in any one of (1) to (5) above. Of course, this is not always the case. [Explanation of Symbols]
[0091] 1: Information Processing System 2: Communication lines 10: Information Processing Device 11: Control Unit 12: Storage section 13: Communications Department 14: Communications bus 20: User terminal 21: Control Unit 22: Storage section 23: Communications Department 24: Input section 25: Output section 26: Communications bus 111: Acquisition Department 112: Guessing part 113: Search Department 120: Artificial Intelligence Department 211: Display section 212: Operation acquisition section PM1: Deep learning model PM2: Regression Model
Claims
1. An information processing system, A processor configured to perform the following steps by reading a program, In the acquisition step, the amino acid sequence to be predicted is obtained. In the prediction step, the binding affinity of the protein having the predicted amino acid sequence to lipids is predicted based on the predicted amino acid sequence and reference information. Here, the reference information is an information processing system that includes a correlation between amino acid sequences and the binding force, constructed based on measured values of the binding force.
2. In the information processing system described in claim 1, The aforementioned reference information includes first reference information and second reference information, The first reference information includes the correlation between the amino acid sequence and the feature vector. The second reference information includes the correlation between the feature vector and the coupling force, An information processing system that, in the estimation step, obtains the feature vector based on the amino acid sequence to be estimated and the first reference information, and estimates the binding force based on the feature vector and the second reference information.
3. In the information processing system described in claim 2, The second reference information is an information processing system constructed using machine learning data that combines the feature vector obtained from the amino acid sequence of the learning protein with the measured value of the binding force in the learning protein.
4. In the information processing system described in claim 2, The aforementioned first reference information includes a deep learning model trained to generate feature vectors representing protein features from amino acid sequences. An information processing system that, in the prediction step, inputs the amino acid sequence to be predicted into the deep learning model and obtains the feature vector from the deep learning model.
5. In the information processing system described in claim 1, The aforementioned estimation step involves an information processing system that estimates the binding affinity of the protein to each of several types of lipids or each of several groups of lipids.
6. Information processing method, An information processing method comprising each step performed by the information processing system according to any one of claims 1 to 5.
7. It is a program, A program for causing a computer to perform each step of the information processing system described in any one of claims 1 to 5.
Citation Information
Patent Citations
Information processing system, information processing method, program, and method for producing antigen-binding molecule or protein
JP7516368B2