Pharmaceutical product development assistance device, operation method for pharmaceutical product development assistance device, and operation program for pharmaceutical product development assistance device

The pharmaceutical development support device improves the accuracy of predicting biopharmaceutical storage solution formulations by converting amino acid sequence information into vector data and using machine learning to derive suitable formulations based on three-dimensional structures.

WO2025142130A1PCT designated stage expired Publication Date: 2025-07-03FUJIFILM CORP
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/039307
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-25
Filing Date
2024-11-05
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing methods for determining the formulation of storage solutions for biopharmaceuticals are time-consuming and lack accuracy due to not considering the three-dimensional structure of amino acids, leading to insufficient prediction of suitable formulations.

Method used

A pharmaceutical development support device that uses a processor to convert target sequence information into vector data reflecting the three-dimensional structure of amino acids, derive target formulation information using a machine learning model trained on reference data, and present this information to improve prediction accuracy.

Benefits of technology

Enhances the prediction accuracy of storage solution formulations by considering the three-dimensional structure of amino acids, allowing for more precise formulation determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024039307_03072025_PF_FP_ABST
    Figure JP2024039307_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a pharmaceutical product development assistance device comprising a processor, wherein the processor acquires subject sequence information indicating the sequence of amino acids forming an amino acid-derived substance contained in a subject pharmaceutical product, converts the subject sequence information into subject vector data reflecting the three-dimensional structure of the amino acids, derives subject formulation information indicating a formulation of a storage solution suitable for the subject pharmaceutical product on the basis of the subject vector data, and presents the subject formulation information to a user.
Need to check novelty before this filing date? Find Prior Art

Description

Drug development support device, drug development support device operation method, and drug development support device operation program

[0001] The technology of the present disclosure relates to a drug development support device, an operating method for a drug development support device, and an operating program for a drug development support device.

[0002] Recently, pharmaceuticals such as biopharmaceuticals, peptide drugs, and nucleic acid drugs have been attracting attention due to their high efficacy and minimal side effects. For example, biopharmaceuticals use proteins such as interferon and antibodies as their active ingredients. Biopharmaceuticals are stored in a preservative solution. To maintain the stable quality of biopharmaceuticals, it is important to ensure that the preservative solution formulation (also known as formulation) is suitable for biopharmaceuticals. The preservative solution is composed of a base buffer solution and additives such as salts, sugars, amino acids, and surfactants. The preservative solution formulation includes, for example, the type of buffer solution and the presence or absence of salts, sugars, amino acids, and surfactants.

[0003] Conventionally, when determining the formulation of a preservation solution, multiple types of preservation solutions are prepared by varying the type of buffer and the combination of salts, sugars, amino acids, and surfactants, and then the storage stability of each preservation solution is confirmed through actual testing. However, this process is very time-consuming. Therefore, for example, International Publication No. 2019 / 070517 proposes a technology for predicting the formulation of a preservation solution suitable for a biopharmaceutical based on the physical properties of the protein, which is the active ingredient of the biopharmaceutical.

[0004] The amino acids that make up proteins have complex three-dimensional structures (secondary, tertiary, and quaternary structures), and these three-dimensional structures vary from protein to protein. Even proteins with roughly the same amino acid sequence (primary structure) often have completely different three-dimensional structures. These differences in three-dimensional structure affect the formulation of an appropriate preservation solution. However, WO 2019 / 070517 does not take into account the three-dimensional structure of amino acids. As a result, the accuracy of predicting the formulation of a preservation solution may be insufficient.

[0005] One embodiment of the technology of the present disclosure provides a drug development support device, an operating method for the drug development support device, and an operating program for the drug development support device that can improve the prediction accuracy of preservative solution prescriptions.

[0006] The pharmaceutical development support device disclosed herein includes a processor, which acquires target sequence information representing the sequence of amino acids that constitute an amino acid-derived substance contained in a target pharmaceutical, converts the target sequence information into target vector data that reflects the three-dimensional structure of the amino acids, derives target prescription information representing a prescription for a preservative solution suitable for the target pharmaceutical based on the target vector data, and presents the target prescription information to a user.

[0007] The processor preferably converts the target sequence information into target vector data using a language model that treats the target sequence information as a sentence in natural language processing.

[0008] The target prescription information is preferably vector data representing the type of buffer that forms the base of the preservative solution suitable for the target drug, as well as the presence or absence of salts, sugars, amino acids, and surfactants contained in the preservative solution suitable for the target drug.

[0009] It is preferable that the processor derives the target prescription information using a machine learning model trained with training data consisting of a pair of reference vector data converted from reference sequence information representing the sequence of amino acids that make up an amino acid-derived substance contained in a reference drug other than the target drug, and reference prescription information representing a prescription of a preservative solution suitable for the reference drug.

[0010] It is preferable that the reference prescription information is clustered into one of a plurality of clusters, the machine learning model outputs one of the plurality of clusters in response to input of target vector data, and the processor derives the reference prescription information belonging to the cluster output by the machine learning model as the target prescription information.

[0011] The processor preferably derives the target prescription information based on the appearance frequency of the reference prescription information belonging to the cluster output by the machine learning model.

[0012] Preferably, the processor accepts a user's specification of a deriving condition for the target prescription information, and derives the target prescription information based on the deriving condition.

[0013] The derivation condition is preferably a number for deriving the target prescription information.

[0014] The extraction conditions are preferably at least one of the type of buffer solution that is the base of the preservation solution and the presence or absence of an additive contained in the preservation solution.

[0015] The reference prescription information is preferably vector data representing the type of buffer that forms the base of a storage solution suitable for the reference drug, as well as the presence or absence of salts, sugars, amino acids, and surfactants contained in the storage solution suitable for the reference drug.

[0016] Preferably the substance is a protein.

[0017] Preferably, the protein is an antibody.

[0018] The method of operating the drug development support device disclosed herein includes obtaining target sequence information representing the sequence of amino acids that constitute an amino acid-derived substance contained in the target drug, converting the target sequence information into target vector data that reflects the three-dimensional structure of the amino acids, deriving target prescription information representing a prescription for a preservative solution suitable for the target drug based on the target vector data, and presenting the target prescription information to a user.

[0019] The operating program of the pharmaceutical development support device disclosed herein causes a computer to execute processes including obtaining target sequence information representing the sequence of amino acids that constitute amino acid-derived substances contained in the target pharmaceutical, converting the target sequence information into target vector data that reflects the three-dimensional structure of the amino acids, deriving target prescription information representing a prescription for a preservative solution suitable for the target pharmaceutical based on the target vector data, and presenting the target prescription information to a user.

[0020] According to the technology of the present disclosure, it is possible to provide a drug development support device, a method for operating a drug development support device, and an operating program for a drug development support device that are capable of improving the prediction accuracy of preservative solution prescriptions.

[0021] 1 is a diagram showing a drug development support system. FIG. 1 is a diagram showing the formation of a biopharmaceutical. FIG. 2 is a diagram showing target sequence information. FIG. 3 is a diagram showing target prescription information. FIG. 4 is a block diagram showing computers constituting a drug development support device and a user terminal. FIG. 5 is a block diagram showing a processing unit of a CPU of the drug development support device. FIG. 6 is a diagram showing the formation of a conversion model. FIG. 7 is a diagram showing the processing of the conversion unit. FIG. 8 is a diagram showing the formation of learning data for a prediction model. FIG. 9 is a diagram showing reference prescription information. FIG. 10 is a diagram showing clustering of reference prescription information and cluster information. FIG. 11 is a diagram showing generation of appearance frequency information from cluster information. FIG. 12 is a diagram showing processing in the learning phase of a prediction model. FIG. 13 is a diagram showing the processing of a derivation unit. FIG. 14 is a block diagram showing a processing unit of a CPU of a user terminal. FIG. 15 is a diagram showing a target sequence information input screen. FIG. 16 is a diagram showing a prediction result display screen. FIG. 17 is a flowchart showing the processing procedure of the drug development support device. FIG. 18 is a diagram showing a 2_1 embodiment in which the number of target prescription information to be derived is specified as a derivation condition. FIG. 19 is a diagram showing the processing of the derivation unit in the 2_1 embodiment. FIG. 19 is a diagram showing a 2_2 embodiment in which the type of buffer solution is specified as a derivation condition. FIG. 19 is a diagram showing the processing of the derivation unit in the 2_2 embodiment.

[0022] As shown in FIG. 1 , a pharmaceutical development support system 10 supports the development of a biopharmaceutical 11 and includes a pharmaceutical development support device 12 and a user terminal 13. The pharmaceutical development support device 12 and the user terminal 13 are connected via a network 14. The user terminal 13 is installed at a pharmaceutical company developing the biopharmaceutical 11 or at an organization contracted by a pharmaceutical company to develop the biopharmaceutical 11, i.e., a contract research organization (CRO). The user terminal 13 is operated by a user U involved in the development of the biopharmaceutical 11 at the pharmaceutical company or contract research organization (hereinafter collectively referred to as a pharmaceutical facility). The network 14 is, for example, a wide area network (WAN) such as the Internet or a public communication network. While only one user terminal 13 is connected to the pharmaceutical development support device 12 in FIG. 1 , in reality, multiple user terminals 13 at multiple pharmaceutical facilities are connected to the pharmaceutical development support device 12.

[0023] The user terminal 13 transmits a prediction request 15 to the drug development support device 12. The prediction request 15 is a request to have the drug development support device 12 predict information that will be useful in promoting the development of the biopharmaceutical 11.

[0024] 2, the biopharmaceutical 11 is a mixed solution of an antibody 16, which is an active ingredient, and a preservative solution 17. The antibody 16 is an example of an "amino acid-derived substance" according to the technology of the present disclosure.

[0025] The preservation solution 17 is composed of a base buffer solution 18 and additives 19 added to the buffer solution 18. The additives 19 are salts, sugars, amino acids, and surfactants. The preservation solution 17 is prepared by referring to prescription information 20, which specifies the type of buffer solution 18 and the presence or absence of salts, sugars, amino acids, and surfactants.

[0026] There are several types of buffer solutions 18, including phosphate buffer solutions, acetate buffer solutions, and citrate buffer solutions. Examples of salts include sodium chloride, potassium chloride, sodium acetate, and ammonium sulfate. Examples of sugars include pure sugars such as refined white sugar, sucrose, and trehalose, and sugar alcohols such as glycerin and sorbitol. Examples of amino acids include proline, arginine, glutamine, and histidine. Examples of surfactants include polysorbate 20 and polysorbate 80.

[0027] Returning to FIG. 1 , the prediction request 15 includes target sequence information 21T. The target sequence information 21T is information representing the sequence of amino acids constituting the antibody 16 contained in the target drug 11T. The target sequence information 21T is identified through experiments. Here, the target drug 11T is a biopharmaceutical 11 currently being developed by a user U at a pharmaceutical facility, and is a biopharmaceutical 11 for which the formulation of an appropriate preservative solution 17 is unknown. Although not shown in the figure, the prediction request 15 also includes a terminal ID (identification data) and the like for uniquely identifying the user terminal 13 that sent the prediction request 15.

[0028] When the prediction request 15 is received, the pharmaceutical development support device 12 predicts target prescription information 20T that represents a prescription for a preservative solution 17 suitable for the target pharmaceutical 11T, as information useful for promoting the development of the biopharmaceutical 11. The pharmaceutical development support device 12 then distributes the target prescription information 20T to the user terminal 13 that sent the prediction request 15. When the target prescription information 20T is received, the user terminal 13 makes the target prescription information 20T available for viewing by the user U.

[0029] As an example, as shown in FIG. 3 , the target sequence information 21T includes a target drug ID for uniquely identifying the target drug 11T. The target sequence information 21T describes the order of peptide bonds of the amino acids constituting the antibody 16 contained in the target drug 11T, from the amino terminus to the carboxyl terminus, using single-letter alphabetical abbreviations representing the amino acids. Since there are approximately 450 amino acids constituting the antibody 16, the target sequence information 21T also contains a string of approximately 450 letters. Examples of abbreviations include "E" for glutamic acid, "L" for leucine, and "G" for glycine. Such amino acid sequences are also called primary structures.

[0030] As shown in FIG. 4 as an example, the target prescription information 20T includes the target drug ID, just like the target sequence information 21T. The target prescription information 20T is information representing the type of buffer solution 18 and the presence or absence of salt, sugar, amino acid, and surfactant as multidimensional (here, 15-dimensional) vector data consisting of a series of binary digits "0" and "1." The type of buffer solution 18 is represented by a 4-bit binary digit such as "0001." Salt has the fields "yes" and "no." If the addition of salt is appropriate, "1" is registered for "yes" and "0" for "no." In contrast, if the addition of salt is not appropriate, "0" is registered for "yes" and "1" is registered for "no." Sugar has the fields "yes (sugar)," "yes (sugar alcohol)," and "no." If the addition of sugar is appropriate, "1" is registered for "yes (sugar)," and "0" is registered for all other fields. If the addition of sugar alcohols is appropriate, a "1" is registered for "Yes (Sugar Alcohols)" and a "0" is registered for all other items. On the other hand, if the addition of sugars and sugar alcohols is not appropriate, a "1" is registered for "No" and a "0" is registered for all other items.

[0031] Amino acids have the items "Yes (acidic)", "Yes (neutral)", "Yes (alkaline)", and "No". If the addition of an acidic amino acid is appropriate, "1" is registered for "Yes (acidic)", and "0" for other items. If the addition of a neutral amino acid is appropriate, "1" is registered for "Yes (neutral)", and "0" for other items. Furthermore, if the addition of an alkaline amino acid is appropriate, "1" is registered for "Yes (alkaline)", and "0" for other items. In contrast, if the addition of an amino acid is not appropriate, "1" is registered for "No" and "0" for other items. Surfactants have the items "Yes" and "Absent". If the addition of a surfactant is appropriate, "1" is registered for "Yes" and "0" for "Absent". In contrast, if the addition of a surfactant is not appropriate, "0" is registered for "Yes" and "Absent".

[0032] 5, the computers that make up the pharmaceutical development support device 12 and the user terminal 13 basically have the same configuration, and include a storage 25, a memory 26, a CPU (Central Processing Unit) 27, a communication unit 28, a display 29, and an input device 30. These are interconnected via a bus line 31.

[0033] The storage 25 is a hard disk drive built into the computer that constitutes the pharmaceutical development support device 12 and the user terminal 13, or connected via a cable or network. Alternatively, the storage 25 is a disk array consisting of multiple hard disk drives. The storage 25 stores control programs such as an operating system, various application programs (hereinafter referred to as APs (Application Programs)), and various data associated with these programs. Note that a solid state drive may be used instead of a hard disk drive.

[0034] The memory 26 is a work memory for the CPU 27 to execute processing. The CPU 27 loads programs stored in the storage 25 into the memory 26 and executes processing in accordance with the programs. In this way, the CPU 27 comprehensively controls each part of the computer. The CPU 27 is an example of a "processor" according to the technology of the present disclosure. The memory 26 may be built into the CPU 27.

[0035] The communication unit 28 is a network interface that controls the transmission of various information via the network 14, etc. The display 29 displays various screens. The various screens are equipped with operation functions using a GUI (Graphical User Interface). The computers that make up the pharmaceutical development support device 12 and the user terminal 13 accept input of operation instructions from an input device 30 via the various screens. The input device 30 is a keyboard, a mouse, a touch panel, a microphone for voice input, etc.

[0036] In the following explanation, the parts of the computer that make up the pharmaceutical development support device 12 (storage 25 and CPU 27) are distinguished by adding the suffix "A" to their reference symbols, and the parts of the computer that make up the user terminal 13 (storage 25, CPU 27, display 29, and input device 30) are distinguished by adding the suffix "B" to their reference symbols.

[0037] As an example, as shown in Figure 6, an operating program 35 is stored in storage 25A of drug development support device 12. Operating program 35 is an AP for causing a computer to function as drug development support device 12. In other words, operating program 35 is an example of an "operating program of a drug development support device" according to the technology of the present disclosure. Storage 25A also stores a conversion model 36, a prediction model 37, and appearance frequency information 38, etc. The prediction model 37 is an example of a "machine learning model" according to the technology of the present disclosure.

[0038] When the operating program 35 is started, the CPU 27A of the computer constituting the pharmaceutical development support device 12 works in cooperation with the memory 26 and the like to function as a request receiving unit 40, a read / write (hereinafter abbreviated as RW (Read Write)) control unit 41, a conversion unit 42, a derivation unit 43, and a screen distribution control unit 44.

[0039] The request receiving unit 40 receives various requests from the user terminal 13. In particular, the request receiving unit 40 receives a prediction request 15 from the user terminal 13. The prediction request 15 includes target sequence information 21T as described above. Therefore, by receiving the prediction request 15, the request receiving unit 40 acquires the target sequence information 21T. When the prediction request 15 is received, the request receiving unit 40 outputs the target sequence information 21T included in the prediction request 15 to the RW control unit 41. Furthermore, the request receiving unit 40 outputs the terminal ID of the user terminal 13 included in the prediction request 15 to the screen distribution control unit 44.

[0040] The RW control unit 41 controls the storage of various data in the storage 25A and the reading of various data from the storage 25A. For example, the RW control unit 41 stores the target sequence information 21T from the request receiving unit 40 in the storage 25A. The RW control unit 41 also reads the target sequence information 21T from the storage 25A and outputs the read target sequence information 21T to the conversion unit 42.

[0041] Furthermore, the RW control unit 41 reads the conversion model 36 from the storage 25A and outputs the read conversion model 36 to the conversion unit 42. The RW control unit 41 also reads the prediction model 37 and the occurrence frequency information 38 from the storage 25A and outputs the read prediction model 37 and the occurrence frequency information 38 to the derivation unit 43.

[0042] The conversion unit 42 converts the target array information 21T into target vector data 50T using the conversion model 36. The conversion unit 42 outputs the target vector data 50T to the derivation unit 43.

[0043] The derivation unit 43 derives the target prescription information 20T corresponding to the target vector data 50T using the prediction model 37 and the appearance frequency information 38. The derivation unit 43 outputs the target prescription information 20T to the screen delivery control unit 44.

[0044] The screen delivery control unit 44 controls the delivery of various screens to the user terminal 13. Specifically, the screen delivery control unit 44 delivers and outputs various screens to the user terminal 13 that has sent the various requests in the form of screen data for web delivery created using a markup language such as XML (Extensible Markup Language). At this time, the screen delivery control unit 44 identifies the user terminal 13 that has sent the various requests based on the terminal ID from the request receiving unit 40. The various screens include a subject sequence information input screen 80 (see FIG. 17) for inputting subject sequence information 21T and a prediction result display screen 85 (see FIG. 18) for displaying subject prescription information 20T. Note that other data description languages, such as JSON (Javascript (registered trademark) Object Notation), may be used instead of XML.

[0045] As an example, as shown in FIG. 7 , a vectorization unit 56 that is part of a language model 55 is diverted into the conversion model 36. The language model 55 is a machine learning model originally used in the field of natural language processing (NLP). The language model 55 handles sequence information 21 of amino acids that constitute an antibody 16 as a sentence. The language model 55 is based on, for example, BERT (Bidirectional Encoder Representations from Transformers) that uses a transformer encoder.

[0046] The vectorization unit 56 converts the sequence information 21 input to the language model 55 into vector data 50. This vector data 50 is multidimensional data consisting of a series of real values ​​between 0 and 1, for example (see FIGS. 8 and 9). The number of dimensions of the vector data 50 is, for example, 512, 1024, or 2048. The vector data 50 reflects not only the sequence (primary structure) of the amino acids that make up the antibody 16, but also the three-dimensional structure (secondary structure, tertiary structure, and quaternary structure). The vectorization unit 56 outputs the vector data 50 to the processing unit 57. The processing unit 57 outputs a processing result 58 based on the vector data 50.

[0047] The language model 55 is generated by pre-training a pre-trained language model 55A and then fine-tuning a pre-trained language model 55B. The pre-training includes MLM (Masked Language Modeling) and NSP (Next Sentence Prediction). MLM is a learning method that masks a portion of the amino acid sequence of the sequence information 21 and predicts which amino acid will be inserted in the masked portion, a so-called fill-in-the-blank problem. NSP is a learning method that determines whether there is a correlation between the amino acid sequence information 21 that constitutes two different antibodies 16. Fine-tuning is learning according to the desired processing task to be performed by the processing unit 57. The processing task may, for example, output an optimum value for the hydrogen ion exponent (pH (Potential of Hydrogen) value) of the preservation solution 17 or an optimum value for the temperature of the preservation solution 17 as the processing result 58 in response to the input of the sequence information 21. The vectorization unit 56 of the language model 55 generated in this manner is stored in the storage 25A as the conversion model 36. The pre-training and fine-tuning may be performed in the drug development support device 12 or in a device separate from the drug development support device 12. Furthermore, the pre-training and fine-tuning may be continued even after the conversion model 36 is stored in the storage 25A. Alternatively, only the pre-training may be performed without fine-tuning. Furthermore, NSP may not be performed.

[0048] 8, the conversion unit 42 inputs the target sequence information 21T to the conversion model 36 and outputs target vector data 50T from the conversion model 36. The target vector data 50T includes the target drug ID, just like the target sequence information 21T.

[0049] The training of the prediction model 37 will be described below. First, as shown in FIG. 9 as an example, training data 60 for the prediction model 37 is generated from a reference drug 11R. The reference drug 11R is a biopharmaceutical 11 that is different from the target drug 11T and was developed earlier than the target drug 11T, and the formulation of a suitable preservative solution 17 for this biopharmaceutical 11 is known. Reference sequence information 21R representing the amino acid sequence constituting the antibody 16 contained in the reference drug 11R is converted into reference vector data 50R using a conversion model 36. Then, a pair of this reference vector data 50R and reference prescription information 20R representing the formulation of a preservative solution 17 suitable for the reference drug 11R is stored in a database 61 as training data 60. The database 61 stores multiple training data 60 generated from multiple reference drugs 11R.

[0050] The reference sequence information 21R and the reference vector data 50R include a reference drug ID for uniquely identifying the reference drug 11R. Like the target sequence information 21T, the reference sequence information 21R describes the order of peptide bonds of the amino acids constituting the antibody 16 contained in the reference drug 11R, from the amino terminus to the carboxyl terminus, using single-letter alphabetic abbreviations representing the amino acids. Similarly to the target vector data 50T, the reference vector data 50R is composed of a series of real values, for example, between 0 and 1, that reflect the three-dimensional structure of the amino acids constituting the antibody 16 contained in the reference drug 11R.

[0051] 10, the reference prescription information 20R includes a reference drug ID, similar to the reference sequence information 21R, etc. Similarly to the target prescription information 20T, the reference prescription information 20R is information that represents the type of buffer solution 18 and the presence or absence of salt, sugar, amino acid, and surfactant in the form of vector data consisting of a series of binary digits "0" and "1."

[0052] After collecting the learning data 60, a clustering process is performed on a plurality of pieces of reference prescription information 20R to define the cluster to which each piece of reference prescription information 20R belongs, as shown in FIG. 11 as an example. FIG. 11 shows an example in which a plurality of pieces of reference prescription information 20R are clustered into three clusters: cluster 1, cluster 2, and cluster 3. The clustering process can be performed using the k-means method or hierarchical clustering. As a result of this clustering process, cluster information 65 is generated. The cluster information 65 is information in which the cluster to which each reference drug ID and each piece of reference prescription information 20R belongs is registered. For ease of explanation, in FIG. 11, the dimension of the vector space 66 in which the reference prescription information 20R is plotted is two-dimensional, having axes D1 and D2. However, the actual dimension of the vector space 66 is 512 dimensions, as described above. Instead of the reference prescription information 20R, the clustering process may be performed on the reference vector data 50R. Furthermore, clustering processing may be performed on vector data obtained by combining the reference prescription information 20R and the reference vector data 50R.

[0053] After the cluster processing, as shown in FIG. 12 as an example, appearance frequency information 38 is generated based on the cluster information 65. More specifically, the number of appearances of the reference prescription information 20R in each cluster is counted. Then, for each cluster, the reference prescription information 20R is arranged in descending order of appearance frequency to generate appearance frequency information 38. The appearance frequencies "1", "2", ... in the appearance frequency information 38 represent the order of appearance frequency. For example, the reference prescription information 20R with the highest number of appearances and the highest appearance frequency in cluster 1 is "001010100...". Furthermore, the reference prescription information 20R with the second highest number of appearances and the second highest appearance frequency in cluster 2 is "001101010...". The appearance frequency information 38 generated in this manner is stored in the storage 25A.

[0054] As an example, as shown in FIG. 13 , training of the prediction model 37 uses training data 60A, which is derived from training data 60 and is composed of pairs of reference vector data 50R and clusters. Specifically, the reference vector data 50R is input to the prediction model 37, which then outputs a training cluster prediction result 70L. The training cluster prediction result 70L is a prediction result of a cluster to which the reference prescription information 20R corresponding to the original reference drug 11R of the reference vector data 50R belongs. Based on the training cluster prediction result 70L and the clusters of the training data 60A, a loss calculation is performed for the prediction model 37 using a loss function. Then, various coefficients of the prediction model 37 (such as the filter coefficients of the convolution layer) are updated according to the results of the loss calculation, and the prediction model 37 is updated according to the update setting.

[0055] During training of the prediction model 37, the above series of processes, including input of the reference vector data 50R to the prediction model 37, output of the training cluster prediction results 70L from the prediction model 37, loss calculation, update setting, and update of the prediction model 37, are repeated while the training data 60A is exchanged. The repetition of the above series of processes is terminated when the prediction accuracy of the training cluster prediction results 70L reaches a predetermined set level. The prediction model 37 whose prediction accuracy has reached the set level is stored in the storage 25A. Note that training may be terminated after the above series of processes have been repeated a set number of times, regardless of the prediction accuracy of the training cluster prediction results 70L. Training of the prediction model 37 may be performed in the pharmaceutical development support device 12 or in a device separate from the pharmaceutical development support device 12. Training of the prediction model 37 may also be continued after the prediction model 37 is stored in the storage 25A.

[0056] 14, the derivation unit 43 inputs target vector data 50T into a prediction model 37 and outputs a cluster prediction result 70 from the prediction model 37. The cluster prediction result 70 is a prediction result of a cluster to which target prescription information 20T corresponding to the original target drug 11T of the target vector data 50T belongs. The prediction model 37 is constructed using a machine learning algorithm such as a support vector machine (SVM) or XGBoost.

[0057] As an example, as shown in FIG. 15 , the derivation unit 43 derives, as the target prescription information 20T, the reference prescription information 20R having the highest appearance frequency among the reference prescription information 20R belonging to the cluster of the cluster prediction result 70. FIG. 15 illustrates a case in which the cluster of the cluster prediction result 70 is cluster 1, and the reference prescription information 20R of "001010100..." having the highest appearance frequency in cluster 1 in the appearance frequency information 38 is derived as the target prescription information 20T. In this way, the derivation unit 43 derives, as the target prescription information 20T, the reference prescription information 20R belonging to the cluster output by the prediction model 37. Furthermore, the derivation unit 43 derives the target prescription information 20T based on the appearance frequency of the reference prescription information 20R belonging to the cluster output by the prediction model 37.

[0058] As an example, as shown in FIG. 16 , a prediction AP 75 is stored in the storage 25B of the user terminal 13. The prediction AP 75 is installed in the user terminal 13 by the user U. The prediction AP 75 is an AP for predicting information regarding the prescription of a preservative solution 17 suitable for a target drug 11T, i.e., target prescription information 20T. When the prediction AP 75 is launched, the CPU 27B of the user terminal 13 functions as a browser control unit 77 in cooperation with the memory 26 and the like. The browser control unit 77 controls the operation of a dedicated web browser for the prediction AP 75.

[0059] The browser control unit 77 reproduces various screens based on various screen data from the drug development support device 12 and displays the reproduced various screens on the display 29B. The browser control unit 77 also accepts various operation instructions input by the user U from the input device 30B via the various screens. The browser control unit 77 transmits various requests, including a prediction request 15, to the drug development support device 12 in response to the operation instructions.

[0060] When the prediction AP 75 is started, a subject sequence information input screen 80, as shown in Fig. 17 as an example, is displayed on the display 29B under the control of the browser control unit 77. The subject sequence information input screen 80 is provided with an input box 81 for subject sequence information 21T. In the input box 81, the subject sequence information 21T can be written or a file of the subject sequence information 21T can be dropped.

[0061] The user U inputs the desired target sequence information 21T into the input box 81 and then selects the prediction button 82. When the prediction button 82 is selected, the browser control unit 77 generates a prediction request 15 including the target sequence information 21T input into the input box 81 and transmits the generated prediction request 15 to the drug development support device 12.

[0062] Furthermore, when the drug development support device 12 predicts the target prescription information 20T, a prediction result display screen 85 shown in Fig. 18 as an example is displayed on the display 29B under the control of the browser control unit 77. The prediction result display screen 85 displays the target prescription information 20T, more specifically, a table 86 showing the contents of the target prescription information 20T. In this way, the target prescription information 20T is presented to the user U in the form of distributed screen data.

[0063] A target sequence information display button 87 is provided at the top of the prediction result display screen 85. When the target sequence information display button 87 is selected, a display screen for the target sequence information 21T is popped up. Furthermore, a save button 88 and an OK button 89 are provided at the bottom of the prediction result display screen 85. When the save button 88 is selected, the target prescription information 20T is associated with the target drug ID and stored in the storage 25B. When the OK button 89 is selected, the display on the prediction result display screen 85 is cleared.

[0064] Next, the operation of the above configuration will be described with reference to the flowchart shown in Fig. 19 as an example. When the operating program 35 is started in the drug development support device 12, the CPU 27A functions as the request receiving unit 40, the RW control unit 41, the conversion unit 42, the derivation unit 43, and the screen distribution control unit 44, as shown in Fig. 6. When the prediction AP 75 is started in the user terminal 13, the CPU 27B functions as the browser control unit 77, as shown in Fig. 16.

[0065] 17 is displayed on the display 29B of the user terminal 13 under the control of the browser control unit 77. When the user U inputs desired target sequence information 21T into the input box 81 on the target sequence information input screen 80 and selects the Predict button 82, a prediction request 15 is sent from the browser control unit 77 to the drug development support device 12. As shown in FIG. 1, the prediction request 15 includes the target sequence information 21T, the terminal ID of the user terminal 13, and the like.

[0066] In the pharmaceutical development support device 12, the request receiving unit 40 receives the prediction request 15, thereby acquiring the target sequence information 21T included in the prediction request 15 (YES in step ST100). The target sequence information 21T included in the prediction request 15 is output from the request receiving unit 40 to the RW control unit 41 and stored in the storage 25A under the control of the RW control unit 41 (step ST110). In addition, the terminal ID of the user terminal 13 included in the prediction request 15 is output from the request receiving unit 40 to the screen distribution control unit 44.

[0067] The target array information 21T is read from the storage 25A by the RW control unit 41 (step ST120). The target array information 21T is output from the RW control unit 41 to the conversion unit 42. The RW control unit 41 also reads the conversion model 36 from the storage 25A and outputs the read conversion model 36 to the conversion unit 42.

[0068] 8, in the conversion unit 42, the target array information 21T is input to the conversion model 36. As a result, the target vector data 50T is output from the conversion model 36. In this way, the target array information 21T is converted into the target vector data 50T using the conversion model 36 (step ST130). The target vector data 50T is output from the conversion unit 42 to the derivation unit 43.

[0069] The RW control unit 41 reads the prediction model 37 and the appearance frequency information 38 from the storage 25A, and outputs the read prediction model 37 and the appearance frequency information 38 to the derivation unit 43. Then, as shown in FIG. 14 , the derivation unit 43 inputs the target vector data 50T to the prediction model 37. As a result, the prediction model 37 outputs a cluster prediction result 70 (step ST140). Next, as shown in FIG. 15 , the derivation unit 43 derives the reference prescription information 20R with the highest appearance frequency, which belongs to the cluster of the cluster prediction result 70, as the target prescription information 20T (step ST150). The target prescription information 20T is output from the derivation unit 43 to the screen delivery control unit 44.

[0070] 18 is generated based on the target prescription information 20T by the screen delivery control unit 44. The screen data of the prediction result display screen 85 is delivered to the user terminal 13 that sent the prediction request 15 under the control of the screen delivery control unit 44 (step ST160).

[0071] In the user terminal 13, under the control of the browser control unit 77, the screen data of the prediction result display screen 85 is reproduced, and the reproduced prediction result display screen 85 is displayed on the display 29B. In this way, the target prescription information 20T is presented to the user U.

[0072] As described above, the CPU 27A of the drug development support device 12 includes a request receiving unit 40, a conversion unit 42, a derivation unit 43, and a screen distribution control unit 44. The request receiving unit 40 acquires target sequence information 21T by receiving a prediction request 15. The conversion unit 42 converts the target sequence information 21T into target vector data 50T reflecting the three-dimensional structure of amino acids. The derivation unit 43 derives target prescription information 20T representing a prescription of a preservative solution 17 suitable for the target drug 11T based on the target vector data 50T. The screen distribution control unit 44 presents the target prescription information 20T to the user U by distributing screen data of a prediction result display screen 85 including the target prescription information 20T to the user terminal 13. Because the target prescription information 20T is derived taking into account the three-dimensional structure of amino acids, it is possible to improve the prediction accuracy of the prescription of the preservative solution 17.

[0073] 8, the conversion unit 42 converts the subject sequence information 21T into subject vector data 50T using a conversion model 36, which is part of a language model 55 that treats the subject sequence information 21T as a sentence in natural language processing. This makes it possible to easily obtain subject vector data 50T that accurately reflects the three-dimensional structures of amino acids.

[0074] 4, the target prescription information 20T is vector data representing the type of buffer solution 18 that serves as the base of the preservation solution 17 suitable for the target drug 11T, as well as the presence or absence of salts, sugars, amino acids, and surfactants contained in the preservation solution 17 suitable for the target drug 11T. According to past knowledge, it is easy to prepare the preservation solution 17 suitable for the target drug 11T as long as the type of buffer solution 18 and the presence or absence of salts, sugars, amino acids, and surfactants are known. For this reason, the target prescription information 20T can be the minimum necessary information, excluding information that is not particularly useful for preparing the preservation solution 17.

[0075] As shown in Figures 9, 13, 14, and 15, the derivation unit 43 derives the target prescription information 20T using a prediction model 37 trained with training data 60A derived from training data 60 consisting of a pair of reference vector data 50R and reference prescription information 20R. The reference vector data 50R is data converted from reference sequence information 21R representing the amino acid sequence constituting the antibody 16 contained in a reference drug 11R that is separate from the target drug 11T. The reference prescription information 20R is information representing a prescription for a preservative solution 17 suitable for the reference drug 11R. This makes it possible to easily derive the target prescription information 20T with high prediction accuracy.

[0076] As shown in Fig. 11 , the reference prescription information 20R is clustered into one of a plurality of clusters. As shown in Fig. 14 , the prediction model 37 outputs one of the plurality of clusters as a cluster prediction result 70 in response to input of target vector data 50T. As shown in Fig. 15 , the derivation unit 43 derives the reference prescription information 20R belonging to the cluster output by the prediction model 37 as target prescription information 20T.

[0077] The number of combinations of buffer solution types 18 and the presence or absence of salts, sugars, amino acids, and surfactants is extremely large. Predicting a single combination from this extremely large number of combinations is difficult, and even if a prediction is possible, the accuracy of the prediction may be low. In contrast, the number of clusters is at least less than the number of combinations of buffer solution types 18 and the presence or absence of salts, sugars, amino acids, and surfactants, and it may be possible to narrow it down to a few, as in this example. Therefore, predicting clusters rather than predicting combinations of buffer solution types 18 and the presence or absence of salts, sugars, amino acids, and surfactants can improve the accuracy of the prediction of the target prescription information 20T. If there is no risk of low prediction accuracy, a prediction model 37 may be used that derives the target prescription information 20T itself, rather than the cluster prediction result 70, based on the input of the target vector data 50T.

[0078] 15 , the derivation unit 43 derives the target prescription information 20T based on the appearance frequency of the reference prescription information 20R belonging to the cluster output by the prediction model 37. Therefore, it is possible to derive the target prescription information 20T with high validity.

[0079] The reference prescription information 20R is vector data that indicates the type of buffer solution 18 that serves as the base of the preservation solution 17 suitable for the reference drug 11R, and the presence or absence of salts, sugars, amino acids, and surfactants contained in the preservation solution 17 suitable for the reference drug 11R. Therefore, similar to the case of the target prescription information 20T, the reference prescription information 20R can be the minimum necessary information that excludes information that is not particularly useful for preparing the preservation solution 17.

[0080] Biopharmaceuticals 11 containing antibodies 16 as proteins are called antibody drugs and are widely used in the treatment of chronic diseases such as cancer, diabetes, and rheumatoid arthritis, as well as rare diseases such as hemophilia and Crohn's disease. Therefore, this example, in which the substance is a protein and the protein is an antibody 16, can further promote the development of antibody drugs that are widely used in the treatment of various diseases.

[0081] In the above-described first embodiment, an example has been given in which the reference prescription information 20R having the highest appearance frequency and belonging to the cluster of the cluster prediction result 70 is derived as the target prescription information 20T, but this is not limiting. For example, three pieces of reference prescription information 20R having the first to third highest appearance frequencies and belonging to the cluster of the cluster prediction result 70 may be derived as the target prescription information 20T. Alternatively, all of the reference prescription information 20R belonging to the cluster of the cluster prediction result 70 may be derived as the target prescription information 20T. Furthermore, as in the following second_1 and second_2 embodiments, the target prescription information 20T may be derived based on a derivation condition specified by the user U.

[0082] [2_1 Embodiment] As shown in FIG. 20 as an example, in the 2_1 embodiment, the request receiving unit 40 receives a derivation condition setting request 95 from the user terminal 13. The derivation condition setting request 95 includes a derivation condition 96 that specifies the number of target prescription information 20T to be derived. The derivation condition 96 is input by the user U via the input device 30B. The request receiving unit 40 outputs the derivation condition 96 included in the derivation condition setting request 95 to the derivation unit 43. FIG. 20 illustrates an example in which "3" is specified as the derivation condition 96.

[0083] In this case, as shown in FIG. 21 as an example, the derivation unit 43 derives three pieces of reference prescription information 20R with the first to third highest appearance frequencies that belong to the cluster of the cluster prediction result 70 as target prescription information 20T.

[0084] 22 shows an example of a case where the request receiving unit 40 receives a derivation condition setting request 100 from the user terminal 13. The derivation condition setting request 100 includes a derivation condition 101 that specifies the type of buffer solution 18 and the presence or absence of salt, sugar, amino acid, and surfactant. The derivation condition 101 is input by the user U via the input device 30B. The request receiving unit 40 outputs the derivation condition 101 included in the derivation condition setting request 100 to the derivation unit 43. FIG. 22 illustrates an example where the type of buffer solution 18, i.e., "phosphate buffer solution (0010)," is specified as the derivation condition 101.

[0085] In this case, as an example, as shown in FIG. 23, the deriving unit 43 derives the reference prescription information 20R, which belongs to the cluster of the cluster prediction result 70 and in which the buffer solution 18 is a phosphate buffer solution, as the target prescription information 20T.

[0086] As described above, in the 2_1 embodiment and the 2_2 embodiment, the request receiving unit 40 receives designation by the user U of the derivation condition 96 or 101 of the target prescription information 20T. The derivation unit 43 derives the target prescription information 20T based on the derivation condition 96 or 101. Therefore, it is possible to derive the target prescription information 20T that meets the intention of the user U.

[0087] In the second_1 embodiment, the derivation condition 96 is the number of pieces of target prescription information 20T to be derived. Therefore, it is possible to derive a number of pieces of target prescription information 20T that match the user U's intentions. In the second_2 embodiment, the derivation condition 101 is at least one of the type of buffer solution 18 and the presence or absence of salt, sugar, amino acid, and surfactant. Therefore, if the user U has some knowledge about the prescription of the preservative solution 17 suitable for the target drug 11T, such as the fact that the buffer solution 18 can only be a phosphate buffer, it is possible to derive target prescription information 20T that matches the user U's knowledge. In other words, it is possible to eliminate target prescription information 20T that does not match the user U's knowledge.

[0088] In the second embodiment, three randomly selected reference prescription information 20R belonging to a cluster of the cluster prediction result 70 may be derived as the target prescription information 20T, regardless of the frequency of appearance. In the second embodiment, a complex derivation condition 101 may be specified, such as specifying the type of buffer solution and the presence or absence of salt in the derivation condition 101. Furthermore, the second embodiment and the second embodiment may be implemented in combination.

[0089] The protein is not limited to the exemplified antibody 16. It may also be a cytokine (interferon, interleukin, etc.), a hormone (insulin, glucagon, follicle-stimulating hormone, erythropoietin, etc.), a growth factor (IGF (insulin-like growth factor)-1, bFGF (basic fibroblast growth factor), etc.), a blood coagulation factor (factor 7, factor 8, factor 9, etc.), an enzyme (lysosomal enzyme, DNA (deoxyribonucleic acid) degrading enzyme, etc.), an Fc (fragment crystalline) fusion protein, a receptor, albumin, or a protein vaccine. The antibody 16 also includes bispecific antibodies, antibody-drug conjugates, low molecular weight antibodies, sugar chain modified antibodies, and the like.

[0090] Furthermore, substances derived from amino acids are not limited to proteins, but may be peptides, nucleic acids, etc. Therefore, pharmaceuticals are not limited to biopharmaceuticals 11 that require biotechnology such as genetic recombination technology and cell culture technology, but may also be peptide pharmaceuticals, nucleic acid pharmaceuticals, etc. that can be produced using only chemical synthesis technology without requiring biotechnology.

[0091] The drug development support device 12 may be installed in a pharmaceutical manufacturing facility, or may be installed in a data center independent of the pharmaceutical manufacturing facility.

[0092] Instead of delivering screen data of the prediction result display screen 85 including the target prescription information 20T to the user terminal 13, the target prescription information 20T itself may be delivered to the user terminal 13. In this case, in the user terminal 13, under the control of the browser control unit 77, the prediction result display screen 85 is generated based on the target prescription information 20T, and the prediction result display screen 85 is displayed on the display 29B.

[0093] The method of presenting the target prescription information 20T to the user U is not limited to the example of delivering screen data. The target prescription information 20T may be presented to the user U by printing it on a paper medium, or by attaching it to an email and sending it to the user terminal 13.

[0094] The hardware configuration of the computer constituting the drug development support device 12 according to the technology of the present disclosure can be modified in various ways. For example, the drug development support device 12 can be configured with multiple computers separated as hardware in order to improve processing power and reliability. For example, the functions of the request receiving unit 40 and the RW control unit 41, and the functions of the conversion unit 42, derivation unit 43, and screen distribution control unit 44 can be distributed and performed by two computers. In this case, the drug development support device 12 is configured with two computers. In addition, some or all of the functions of the drug development support device 12 may be performed by the user terminal 13.

[0095] In this way, the hardware configuration of the computer of the pharmaceutical development support device 12 can be changed as appropriate depending on the required performance, such as processing power, safety, and reliability. Furthermore, not only the hardware, but also APs such as the operating program 35 can be duplicated or stored in multiple storage devices in order to ensure safety and reliability.

[0096] In each of the above embodiments, the hardware structure of the processing unit that executes various processes, such as the request receiving unit 40, the RW control unit 41, the conversion unit 42, the derivation unit 43, the screen delivery control unit 44, and the browser control unit 77, can be any of the various processors listed below. As described above, the various processors include the CPUs 27A and 27B, which are general-purpose processors that execute software (the operating program 35 and the prediction AP 75) and function as various processing units, as well as programmable logic devices (PLDs) that are processors whose circuit configuration can be changed after manufacture, such as a field programmable gate array (FPGA), and dedicated electrical circuits that are processors having a circuit configuration designed specifically for executing specific processing, such as an application specific integrated circuit (ASIC).

[0097] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs and / or a combination of a CPU and an FPGA).Furthermore, multiple processing units may be configured with a single processor.

[0098] Examples of configuring multiple processing units with a single processor include, first, a form in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple processing units. Second, a form in which a processor is used to realize the functions of the entire system including multiple processing units with a single IC (Integrated Circuit) chip, as typified by systems on chips (SoCs). In this way, various processing units are configured using one or more of the above-mentioned various processors as a hardware structure.

[0099] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit (circuitry) that combines circuit elements such as semiconductor elements.

[0100] From the above description, the technology described in the following supplementary paragraphs can be understood.

[0101] [Supplementary Item 1] A drug development support device comprising a processor that acquires target sequence information representing a sequence of amino acids that constitute an amino acid-derived substance contained in a target drug, converts the target sequence information into target vector data reflecting the three-dimensional structure of the amino acids, derives target prescription information representing a prescription of a preservation solution suitable for the target drug based on the target vector data, and presents the target prescription information to a user. [Supplementary Item 2] The drug development support device of Supplementary Item 1, wherein the processor converts the target sequence information into the target vector data using a language model that treats the target sequence information as a sentence in natural language processing. [Supplementary Item 3] The drug development support device of Supplementary Item 1 or Supplementary Item 2, wherein the target prescription information is vector data representing the type of buffer solution that forms the base of the preservation solution suitable for the target drug, and the presence or absence of salt, sugar, amino acids, and surfactants contained in the preservation solution suitable for the target drug. [Supplementary Item 4] The drug development support device according to any one of Supplementary Item 1 to Supplementary Item 3, wherein the processor derives the target prescription information using a machine learning model trained with training data consisting of a set of reference vector data converted from reference sequence information representing an amino acid sequence constituting an amino acid-derived substance contained in a reference drug other than the target drug, and reference prescription information representing a prescription of a preservative solution suitable for the reference drug. [Supplementary Item 5] The drug development support device according to Supplementary Item 4, wherein the reference prescription information is clustered into one of a plurality of clusters, the machine learning model outputs one of the plurality of clusters in response to input of the target vector data, and the processor derives the reference prescription information belonging to the cluster output by the machine learning model as the target prescription information. [Supplementary Item 6] The drug development support device according to Supplementary Item 5, wherein the processor derives the target prescription information based on the appearance frequency of the reference prescription information belonging to the cluster output by the machine learning model. [Supplementary Item 7] The drug development support device according to Supplementary Item 5, wherein the processor accepts a specification by the user of a derivation condition for the target prescription information, and derives the target prescription information based on the derivation condition.[Supplementary Item 8] The drug development support device according to Supplementary Item 7, wherein the derivation condition is a number for deriving the target prescription information. [Supplementary Item 9] The drug development support device according to Supplementary Item 7 or Supplementary Item 8, wherein the derivation condition is at least one of the type of buffer solution that serves as a base for a preservation solution and the presence or absence of an additive contained in the preservation solution. [Supplementary Item 10] The drug development support device according to any one of Supplementary Item 4 to Supplementary Item 9, wherein the reference prescription information is vector data representing the type of buffer solution that serves as a base for a preservation solution suitable for the reference drug, and the presence or absence of salts, sugars, amino acids, and surfactants contained in the preservation solution suitable for the reference drug. [Supplementary Item 11] The drug development support device according to any one of Supplementary Item 1 to Supplementary Item 10, wherein the substance is a protein. [Supplementary Item 12] The drug development support device according to Supplementary Item 11, wherein the protein is an antibody.

[0102] The technology of the present disclosure can be appropriately combined with the various embodiments and / or various modified examples described above. Furthermore, it is not limited to the above embodiments, and various configurations can be adopted without departing from the spirit of the present disclosure. Furthermore, the technology of the present disclosure extends not only to programs, but also to storage media that non-temporarily store programs, and computer program products that include programs.

[0103] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0104] In this specification, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."

[0105] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. A pharmaceutical development support device comprising a processor, wherein the processor acquires target sequence information representing the sequence of amino acids constituting an amino acid-derived substance contained in a target pharmaceutical, converts the target sequence information into target vector data reflecting the three-dimensional structure of the amino acids, derives target prescription information representing a formulation of a storage solution suitable for the target pharmaceutical based on the target vector data, and presents the target prescription information to a user.

2. The pharmaceutical development support device according to claim 1, wherein the processor converts the target sequence information into the target vector data using a language model that treats the target sequence information as a sentence in natural language processing.

3. The pharmaceutical development support device according to claim 1, wherein the target prescription information is vector data representing the type of buffer solution that is the basis of the storage solution suitable for the target pharmaceutical, and the presence or absence of salts, sugars, amino acids, and surfactants contained in the storage solution suitable for the target pharmaceutical.

4. The pharmaceutical development support device according to claim 1, wherein the processor derives the target prescription information using a machine learning model trained with learning data composed of a set of reference vector data converted from reference sequence information representing the sequence of amino acids constituting an amino acid-derived substance contained in a reference pharmaceutical different from the target pharmaceutical, and reference prescription information representing a formulation of a storage solution suitable for the reference pharmaceutical.

5. The reference prescription information is clustered into one of a plurality of clusters, the machine learning model outputs one of the plurality of clusters in response to an input of the target vector data, and the processor derives, as the target prescription information, the reference prescription information belonging to the cluster output by the machine learning model. The pharmaceutical development support device according to claim 4.

6. The pharmaceutical development support device according to claim 5, wherein the processor derives the target prescription information based on the frequency of occurrence of the reference prescription information belonging to the cluster output by the machine learning model.

7. The pharmaceutical development support device according to claim 5, wherein the processor receives a designation by the user of the derivation conditions for the target prescription information, and derives the target prescription information based on the derivation conditions.

8. The pharmaceutical development support device according to claim 7, wherein the derivation conditions are the number of target prescription information to be derived.

9. The pharmaceutical development support device according to claim 7, wherein the derivation condition is at least one of the type of buffer solution that is the base of the storage solution and the presence or absence of additives contained in the storage solution.

10. The pharmaceutical development support device according to claim 4, wherein the reference prescription information is vector data representing the type of buffer solution that is the base of the storage solution suitable for the reference pharmaceutical product, and the presence or absence of salts, sugars, amino acids, and surfactants contained in the storage solution suitable for the reference pharmaceutical product.

11. The pharmaceutical development support device according to claim 1, wherein the substance is a protein.

12. The pharmaceutical development support device according to claim 11, wherein the protein is an antibody.

13. A method of operating a pharmaceutical development support device, comprising: obtaining target sequence information representing the amino acid sequence of the amino acid-derived substance contained in the target pharmaceutical product; converting the target sequence information into target vector data reflecting the three-dimensional structure of the amino acid; deriving target prescription information representing the prescription of the storage solution suitable for the target pharmaceutical product based on the target vector data; and presenting the target prescription information to the user.

14. A program for operating a pharmaceutical development support device, causing a computer to execute a process including: obtaining target sequence information representing the amino acid sequence of the amino acid-derived substance contained in the target pharmaceutical product; converting the target sequence information into target vector data reflecting the three-dimensional structure of the amino acid; deriving target prescription information representing the prescription of the storage solution suitable for the target pharmaceutical product based on the target vector data; and presenting the target prescription information to the user.

Citation Information

Patent Citations

  • Protein-protein interaction prediction model driven by artificial intelligence

    CN116721693A

  • Protein language model training method, electronic equipment, computer readable medium and program product

    CN116959571A

  • Peptide language model-based bitter peptide prediction method

    CN117153246A

  • Systems and methods for automated biologic development determinations

    WO2019070517A1

  • Pharmaceutical assistance device, operation method for pharmaceutical assistance device, and operation program for pharmaceutical assistance device

    WO2023053891A1