Protein language model processing device, chimeric protein analysis device, and protein language model processing method

By expanding chimeric protein sequence data with homologous sequences and re-training the model, the technology addresses the challenge of limited chimeric data in protein language models, enhancing prediction accuracy and therapeutic drug development for CAR-T cell therapy.

WO2026013857A1PCT designated stage Publication Date: 2026-01-15HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/025133
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

The complexity of biological experiments and the lack of open databases for chimeric proteins hinder the efficient fine-tuning of protein language models, particularly for CAR-T cell therapy, as existing models are trained on natural protein databases without sufficient chimeric sequence data.

Method used

A technology that generates a chimeric protein language model by expanding known chimeric sequence data with homologous sequences, using a threshold for similarity, and re-training the existing protein language model to enhance function prediction performance.

Benefits of technology

Enables efficient fine-tuning of protein language models to accurately predict the functions of chimeric proteins without extensive experimentation, facilitating the development of therapeutic drugs with high therapeutic effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024025133_15012026_PF_FP_ABST
    Figure JP2024025133_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure proposes a protein language model processing device that generates a chimeric protein language model on the basis of an existing protein language model in order to efficiently and highly accurately achieve the fine tuning of a protein language model. The protein language model processing device executes: processing for referring to a given natural protein database and generating, with respect to known sequence data representing known chimeric proteins, an extended sequence data group which represents a plurality of closely-related proteins that share homology equal to or higher than a predetermined threshold value with each other; and processing for generating a chimeric protein language model by re-training an existing protein language model by using the extended sequence data group as training data (see FIG. 1).
Need to check novelty before this filing date? Find Prior Art

Description

Protein language model processing device, chimeric protein analysis device, and protein language model processing method

[0001] The present disclosure relates to a protein language model processing device, a chimeric protein analysis device, and a protein language model processing method.

[0002] In recent years, CAR-T cell therapy has attracted attention as one type of ex vivo gene therapy. CAR-T cells are T cells that express chimeric antigen receptors (hereinafter referred to as CAR molecules) designed to recognize specific antigens and kill target cells. CAR molecules are proteins composed of amino acid sequences with specific functions, and play roles such as recognizing target antigens and transmitting intracellular signals. In recent years, attempts have been made to design the sequence of CAR molecules with the aim of enhancing or improving the durability of the therapeutic effects of CAR-T cell therapy. The sequence design of such CAR molecules aims to obtain highly functional sequences by, for example, learning information obtained from the amino acid sequence of the CAR molecule and the associated cellular functions, and then designing new sequences based on the learning results. More specifically, we will examine whether the therapeutic effects of CAR-T cell therapy can be enhanced by introducing mutations into the sequence of an existing CAR molecule.

[0003] Recently, protein language models (pLMs), which are trained using vast amounts of amino acid sequence data, have been attracting attention as one approach to extracting information from amino acid sequences. When using pLMs, they are first pre-trained using a database of natural proteins, and then additional training (fine-tuning) is performed using the protein sequences to be analyzed (thousands to hundreds of thousands). For fine-tuning, experimentally verified protein sequences or related sequences contained in open databases of natural proteins are often used. For example, Non-Patent Document 1 describes fine-tuning a protein language model using homologous sequences.

[0004] “Low-N protein engineering with data-efficient deep learning,” Surojit Biswas, Grigory Khimulya, Ethan C. Alley, Kevin M. Esvelt and George M. Church, Nature Methods volume 18, pages389-396 (2021)

[0005] However, biological experiments using proteins are generally complex, making it difficult to evaluate the functions of a large number of protein sequences at once. Furthermore, unlike natural proteins, there are no open databases for chimeric proteins, making it difficult to obtain sufficient sequence data for additional training of protein language models. Therefore, a technology to compensate for the lack of sequence data for chimeric proteins is needed. In light of this situation, the present disclosure proposes a technology for efficiently fine-tuning protein language models.

[0006] In order to solve the above-mentioned problems, the present disclosure proposes, as an example, a protein language model processing device that generates a chimeric protein language model based on an existing protein language model, the protein language model processing device comprising: a storage device that stores a program for training the existing protein language model; and a processor that reads the program from the storage device and executes a process of generating the chimeric protein language model based on the program, wherein the processor executes the following processes: a process of acquiring known sequence data representing a known chimeric protein; a process of referring to a given natural protein database, and generating, with respect to the known sequence data, an extended sequence data group representing a plurality of related proteins that have a homology of a given threshold or more; and a process of generating the chimeric protein language model by re-training the existing protein language model using the extended sequence data group as training data.

[0007] Further features related to the present disclosure will become apparent from the description of this specification and the accompanying drawings. Also, aspects of the present disclosure are achieved and realized by the elements and combinations of various elements and the aspects of the following detailed description and the appended claims. The description of this specification is merely exemplary and does not limit the scope or application of the claims of the present disclosure in any way.

[0008] The techniques of the present disclosure enable efficient fine-tuning of protein language models.

[0009] 1 is a diagram for explaining the basic concept of re-learning (fine learning) of a protein language model using extended CAR sequence data. It is a diagram showing an example of the configuration of a language model re-learning process / chimeric protein analysis processing device (also referred to as an information processing device) 200 according to this embodiment. It is a diagram showing the overall sequence of protein language model additional learning (re-learning: fine tuning). It is a diagram for explaining the procedure of generating an extended sequence used to train an existing protein language model from a known CAR sequence (chimeric protein sequence, chimeric amino acid sequence) according to this embodiment. It is a diagram showing an example of the configuration of an extended sequence generation GUI (Graphical User Interface) 501 used when generating an extended sequence group from a known chimeric protein sequence (CAR sequence) according to this embodiment. It is a diagram showing the data structures of input data and output data and their relationship. It is a flowchart for explaining an overview of the re-learning process of a protein language model using an extended sequence. It is a flowchart for explaining the details of the re-learning process of a protein language model using an extended sequence and the prediction process for outputting activity (predicted value) using a trained language model. It is a diagram showing the experimental results of a plurality of mutant sequences obtained by introducing amino acid mutations into a known chimeric protein sequence (sequence ID = FMC63-28Z). FIG. 8 is a diagram visualizing the extended sequences corresponding to each threshold (0.2, 0.4, 0.6, and 0.8 are set as the criteria (thresholds) of the standardized Levenshtein distance, which indicates similarity). It also shows the scores (prediction results) using an existing protein language model without additional learning, and the scores (classification results) obtained using the correlation learning model generated by training the existing protein language model using the extended sequences (FIG. 8) generated based on each threshold (four thresholds) and the training chimeric protein sequence data shown in FIG. 7 .

[0010] In the following embodiments, when necessary for convenience, the description will be divided into multiple sections or embodiments, but unless otherwise specified, they are not unrelated to each other, and one is a partial or complete modification, detail, supplementary explanation, etc. of the other. Furthermore, in the following embodiments, when the number of elements (including the number, numerical value, amount, range, etc.) is mentioned, it is not limited to that specific number, and may be more or less than the specific number, unless otherwise specified or when it is clearly limited to a specific number in principle.

[0011] Furthermore, in the following embodiments, it goes without saying that the components (including element steps, etc.) are not necessarily essential unless otherwise specified or considered to be clearly essential in principle. Similarly, in the following embodiments, when referring to the shape, positional relationship, etc. of components, etc., it is intended to include those that are substantially similar or similar to the shape, etc., unless otherwise specified or considered to be clearly not essential in principle. The same applies to the above-mentioned numerical values ​​and ranges. In all drawings used to explain the embodiments, the same components are generally designated by the same reference numerals, and repeated explanations thereof will be omitted.

[0012] <Basic Concept of Protein Language Model Retraining> FIG. 1 is a diagram for explaining the basic concept of protein language model retraining (fine-tuning) using extended CAR sequence data.

[0013] As described above, a problem with CAR sequence analysis is the limited amount of CAR sequence data due to the complexity of the experiments. Generally, CAR sequence analysis involves retraining using an existing trained protein language model 102. This trained protein language model 102 is a deep learning-based language model trained using a natural protein sequence database (a large-scale open protein database) 101, and is used to analyze protein sequences. When protein sequence information is input, this protein language model converts it into numerical information (numerical information taking into account the amino acid sequence) and outputs it.

[0014] However, the natural protein sequence database 101 does not include chimeric sequence data such as CAR molecules. This makes it difficult to train using an existing language model, as is the case with fluorescent proteins, enzymes, etc. To solve this problem, it is desirable to build a trained protein language model 102 using the natural protein sequence database 101, add the CAR sequence, and train it again.

[0015] However, only a small number of CAR sequences have been reported at present, and it is not easy to obtain them through experiments (ideally, it would be good to obtain many chimeric sequences whose effects have been verified through experiments and use them for learning, but the process of evaluating chimeric sequences through experiments is itself difficult.) Furthermore, it is not enough to simply create a chimeric structure (chimeric sequence) using an existing protein; in reality, there is no point in using a chimeric sequence for additional learning unless it has a stronger therapeutic effect than conventional chimeric sequences.

[0016] Therefore, in this embodiment, sequences for retraining the trained protein language model 102 are increased without conducting experiments based on a known chimeric protein sequence (chimeric sequence) 103 having a killing function. That is, homologous sequence data of the known chimeric protein sequence 103 is generated, and the chimeric sequence data is expanded (generation of an expanded sequence group 104). Note that, although a homologous sequence is given as an example of a related protein, the sequence does not have to be a homologous sequence. One feature of this embodiment is that a threshold (tolerance) for similarity is introduced to generate homologous sequences (extended sequence group 104) that meet the conditions. In this embodiment, such an expanded chimeric sequence is used to fine-train an existing protein language model to generate a retrained protein language model 105, thereby improving and strengthening the function prediction performance of the CAR language model.

[0017] <Configuration Example of Language Model Re-learning Process / Chimeric Protein Analysis Processing Device> Figure 2 is a diagram showing a configuration example of a language model re-learning process / chimeric protein analysis processing device (also simply referred to as an "information processing device") 200 according to this embodiment. The information processing device 200 shown in Figure 2 that executes the language model re-learning process and the chimeric protein analysis process can be configured by a computer, and includes, for example, a processor 201 such as a CPU or MPU, a storage device 202, an input device 203 such as a keyboard, mouse, or touch panel, an output device 204 such as a display, printer, or speaker, and a communication device 205 such as a communication IF (interface) for communicating with the outside.

[0018] The storage device 202 holds a related protein database 2021 that stores, as related proteins, the extended sequence group 104 and homologous sequences of various protein domain sequences obtained when generating the extended sequence group 104, a language model database 2022 that stores existing protein language models and language models that have been fine-tuned by re-learning (e.g., a composite language model including a CAR language model and a protein language model), and a program for executing a language model re-learning process / chimeric protein analysis process (a program corresponding to the flowcharts in FIGS. 6A and 6B ). Furthermore, in the present embodiment and examples, a homologous sequence is given as an example of a related protein, but the related protein is not limited to a homologous sequence and may be another sequence.

[0019] The processor 201 reads the above program from the storage device 202, expands it in its internal memory (not shown), and generates an extraction unit 2011, a generation unit 2012, a learning unit 2013, and a prediction unit 2014. The extraction unit 2011, for example, defines multiple domains from a known chimeric protein sequence (original sequence), searches (extracts) sequences homologous to each of the defined domains from the natural protein sequence database 101, and stores them in the related protein database 2021. The generation unit 2012 generates multiple homologous sequences similar to the original sequence as extended sequence groups based on the homologous sequence groups corresponding to each domain of the known chimeric protein sequence (original sequence), and stores these in the related protein database 2021. Note that, in order to be an extended sequence, certain conditions must be satisfied, which will be described later.

[0020] The learning unit 2013 re-learns the existing protein language model (learns the CAR sequence pattern) using the extended sequence generated by the generation unit 2012, and generates a language model (protein language model + CAR language model) to which a CAR language model with an activated CAR sequence prediction function has been added.

[0021] The prediction unit 2014 applies the chimeric protein sequence (CAR sequence) to be analyzed to the retrained (fine-tuned) protein language model, and predicts the function of the chimeric protein sequence. An example of the predicted function may be an activity value, which is a value indicating the ability to kill target cells.

[0022] The input device 203 is a device used by an operator (user) to input various commands, data, information, etc., and to add protein language models, related protein data, etc. The output device 204 is a device that displays or prints out predicted activity values, etc. obtained by applying the generated extended sequence or the retrained protein language model to the target CAR sequence. The communication device 205 is a device that accesses an external database via an external network (not shown) to acquire data, and transmits protein language model data and predicted activity values ​​obtained by retraining (additional training) to an external information terminal, etc.

[0023] <Overall Sequence of Protein Language Model Additional Learning> FIG. 3 is a diagram showing the overall sequence of protein language model additional learning (relearning: fine tuning).

[0024] (i) Sequence SQ301 The operator (user) inputs information on a known chimeric amino acid sequence (chimeric protein sequence, CAR sequence) using the input device 203. The sequence data can be input, for example, by reading a sequence data file. This input chimeric amino acid sequence is the source sequence (original sequence) from which the extended sequence is generated.

[0025] (ii) Sequence SQ302 The processor 201 executes a protein assay for a known chimeric amino acid sequence and obtains an activity value (fitness).

[0026] (iii) Sequence SQ303: The processor 201 acquires a biological activity value of the known chimeric amino acid sequence based on the protein quantification results of SQ302. The acquired biological activity value may be stored in, for example, the storage device 202. The biological activity value may be the killing ability of a chimeric antigen receptor generated from the chimeric amino acid sequence against target cells, or the stability of the chimeric antigen receptor in the body. Other examples of the biological activity value may include physical property values ​​related to the three-dimensional structure of the chimeric antigen receptor, the proliferation ability, stem cell ability, and cytokine expression level of CAR-T cells into which the chimeric antigen receptor has been introduced.

[0027] (iv) Sequence SQ304 The processor 201 (extraction unit 2011 and generation unit 2012) generates an extended sequence group based on a known chimeric amino acid sequence in order to perform correlation analysis between a new chimeric amino acid sequence group and its activity value. The generation of an extended sequence group will be described in detail later.

[0028] (v) Sequence SQ305 The processor 201 (generation unit 2012) stores the extended sequence group generated in sequence SQ304 in the storage device 202. Design sequence groups 1 to 3 store protein sequences related to known chimeric amino acid sequences (e.g., homologous sequences).

[0029] (vi) Sequence SQ306: The processor 201 (learning unit 2013) uses the extended sequence group 104 stored in the storage device 202 to relearn (fine-tuning: additionally learn what pattern the CAR sequence has) the amino acid sequence pattern of the chimeric protein sequence (CAR sequence) for the existing trained protein language model 102. Note that this relearning can be performed using, for example, deep learning.

[0030] (vii) Sequences SQ307 and SQ308 The processor 201 (learning unit 2013) applies the known chimeric amino acid sequence (extended source sequence) to the retrained protein language model obtained in sequence SQ306, and outputs a feature representation (a numerical value corresponding to the activity value: predicted activity value).

[0031] The processor 201 (learning unit 2013) then compares the feature expression with the bioactivity value (actual measured value) obtained in sequence SQ303. If the difference between the value of the feature expression and the bioactivity value is equal to or less than a preset threshold, the processor 201 (learning unit 2013) terminates the re-learning (additional learning), but if the difference between the value of the feature expression and the bioactivity value is greater than the threshold, the processor 201 (learning unit 2013) continues the additional learning again using a different extended array (sequence SQ307).

[0032] In this way, a language model (CAR language model+protein language model) that enables prediction of the function of a CAR sequence is constructed.

[0033] <Details of the procedure for generating an extended sequence> Figure 4 is a diagram for explaining the procedure for generating an extended sequence used to train an existing protein language model from a known CAR sequence (chimeric protein sequence, chimeric amino acid sequence) according to this embodiment.

[0034] When data of a chimeric protein sequence is input to the information processing device 200, the processor 201 (extraction unit 2011) determines (defines) each domain (domain A, domain B, domain C, domain D, ...) that constitutes the chimeric protein sequence. The domains required for the chimeric protein sequence (CAR sequence) can be specified in advance for each function possessed by the sequence. For example, a domain sequence that penetrates a cell membrane or a domain sequence that transmits a signal within a cell can be mentioned.

[0035] Next, the processor 201 (extraction unit 2011) searches for homologous sequences (similar sequences) for each determined domain from the natural protein sequence database (large-scale open protein database) 101. Through this search, multiple homologous sequences are extracted as a homologous sequence group for each domain.

[0036] The processor 201 (generation unit 2012) selects, from the group of homologous sequences for each domain, sequences whose length is 0.5 to 1.5 times the sequence length of each domain included in the original chimeric protein sequence as similar sequences. The processor 201 (generation unit 2012) then classifies the group of sequences homologous to each original domain sequence using the distance between the original chimeric protein sequence and the similar sequence as the similarity. Here, the similarity is a value based on a threshold value (see FIG. 5A ) input by the operator (user). For example, if the operator inputs a threshold value of 0.2, homologous sequences with a similarity (distance from the original domain sequence) of up to 0.2 are extracted.

[0037] Furthermore, the processor 201 (generation unit 2012) randomly selects one homologous sequence from each domain and connects them together to generate an extended sequence group.

[0038] <GUI for Generating Extended Sequences> FIG. 5A is a diagram showing an example of the configuration of an extended sequence generation GUI (Graphical User Interface) 501 used when generating an extended sequence group from a known chimeric protein sequence (CAR sequence) according to this embodiment.

[0039] The extension sequence generation GUI 501 includes, as its components, an extension source sequence information display field 502 that displays information on a known chimeric protein sequence (extension source sequence) input by the operator, a constituent domain information display field 503 that displays information on each domain constituting the extension source sequence determined by the extraction unit 2011 based on the input extension source sequence, a threshold input field 504 in which the operator inputs a similarity threshold for classifying domain sequence groups, and an execute button 505. Furthermore, the extension source sequence information display field 502 includes the sequence ID, sequence length, and sequence data of the extension source sequence. The constituent domain information display field 503 includes a domain name and sequence data of the corresponding domain.

[0040] 5B is a diagram showing the data structures of input data and output data and their relationship. The input data of the information processing device 200 that executes the language model re-learning process and the chimeric protein analysis process is the above-mentioned known chimeric protein sequence (extension source sequence) data 510, and CAR language model analysis / adjustment input data 512 that includes training chimeric protein sequence data 5121 and verification chimeric protein sequence data 5122. In addition, the output data of the information processing device 200 is chimeric protein sequence data 513 including activity values ​​(predicted values) obtained by re-training the trained protein language model 102 using an extended sequence group 511 constructed by randomly connecting homologous domain groups generated from known chimeric protein sequence (extension source sequence) data 510, and further applying verification chimeric protein sequence data 5122 to an adjusted protein language model (a correlation learning model between chimeric protein sequence and activity: hereinafter referred to as the "correlation learning model") trained to be more suitable for activity value prediction using training chimeric protein sequence data 5121.

[0041] The training chimeric protein sequence data 5121 is sequence data obtained by modifying a portion of the chimeric protein sequence (extension source sequence) through mutation. In the training chimeric protein sequence data 5121, underlined and bolded portions indicate portions that have been modified from the extension source sequence through mutation. The training chimeric protein sequence data 5121 does not need to match the extension sequence group 511, but rather represents the extension source sequence to which one to several mutations have been inserted (mutated). The activity (measured value) indicates a biological activity value obtained by experimenting with a sequence obtained by mutating the extension source sequence. The sequence and activity (measured value) obtained by mutating the extension source sequence can be used to analyze the CAR language model in the retrained (additionally trained) protein language model 105, allowing for better adaptation to the activity value prediction of the chimeric protein.

[0042] The verification chimeric protein sequence data 5122 is sequence data to be input into the analyzed and adjusted protein language model to predict an activity value, and is sequence data used to verify the analyzed and adjusted protein language model. This verification sequence is data generated from known chimeric protein sequence (extension source sequence) data 510, and is sequence data different from the training chimeric protein sequence data 5121. The verification chimeric protein sequence data 5122 may be generated randomly by deep learning, or may be generated by further mutating sequences from the training chimeric protein sequence data 5121 that have an activity (actual measured value) equal to or greater than a predetermined value.

[0043] The chimeric protein sequence data 513 including activity values ​​(predicted values) is sequence data including activity (predicted values) obtained by inputting the verification chimeric protein sequence data 5122 into the correlation learning model. The activity (predicted values) corresponding to each sequence can be used to determine which sequences should be subjected to actual experiments (assays). For example, it can be determined that an experiment (assay) should be performed on a sequence having a higher activity (predicted value) than other sequences.

[0044] <Outline of the training process for a protein language model using an extended sequence and the prediction process for outputting activity (predicted value) using a trained language model> FIG. 6A is a flowchart for explaining the outline of the re-training process for a protein language model using an extended sequence.

[0045] (i) Step S611 The extraction unit 2011 of the processor 201 receives a known chimeric protein sequence (hereinafter also referred to as "extension source sequence") 103 designated by the operator.

[0046] (ii) Step S612 The extraction unit 2011 and generation unit 2012 of the processor 201 generate the extended array group 104 from the extension source array 103 according to the procedure shown in FIG.

[0047] (iii) Step S613 The learning unit 2013 of the processor 201 uses the extended sequence group 104 generated in step S612 to re-learn (for example, using deep learning) the chimeric protein sequence (CAR sequence) pattern in the existing trained protein language model 102, thereby constructing a re-trained protein language model 105.

[0048] <Details of the re-training process of the protein language model using the extended sequence and the prediction process of outputting activity (predicted value) using the trained language model> Figure 6B is a flowchart for explaining the details of the re-training process of the protein language model using the extended sequence and the prediction process of outputting activity (predicted value) using the trained language model. Note that Figure 6B is a flowchart that includes predicting biological activity values ​​of both chimeric protein sequence data and normal (non-chimeric) protein sequences.

[0049] (i) Step S621 The extraction unit 2011 of the processor 201 acquires, for example, sequence data input by an operator or selected by the operator from a sequence list.

[0050] (ii) Step S622: The extraction unit 2011 determines whether the acquired sequence data includes chimeric protein sequence data. The operator (user) knows whether the input data includes chimeric protein sequence data. Therefore, the operator inputs information regarding the presence or absence of chimeric protein sequence data via the input device 203, and the extraction unit 2011 can determine whether the sequence data includes chimeric protein sequence data based on the input. If the acquired sequence data includes chimeric protein sequence data (YES in step S622), the process proceeds to step S623. If the acquired sequence data does not include chimeric protein sequence data (NO in step S622), the process proceeds to step S626.

[0051] (iii) Step S623 The extraction unit 2011 and the generation unit 2012 generate the extended array group 104 from the extension source array 103 according to the procedure shown in FIG.

[0052] (iv) Step S624: The learning unit 2013 of the processor 201 retrains (additionally trains) the existing trained protein language model 102 using the extended sequence group 104 generated in step S623. The learning unit 2013 may further train the retrained language model using the training chimeric protein sequence data 5121 and the verification chimeric protein sequence data 5122, as described in FIG. 5B, to construct a correlation learning model (a correlation learning model between chimeric protein sequence and activity).

[0053] (v) Step S625 The prediction unit 2014 of the processor 201 applies the retrained protein language model 105 or the correlation learning model acquired in step S603 to the chimeric protein sequence data to be analyzed to predict a biological activity value.

[0054] (vi) Step S626 The prediction unit 2014 applies an existing trained protein language model to the input protein sequence data (excluding chimeric sequences) to predict a biological activity value.

[0055] (vii) Step S627: The prediction unit 2014 outputs the biological activity value (predicted value) acquired in step S625 or step S626. The function of the input data can be predicted from this biological activity value (predicted value). As a result, even if the function of a chimeric protein is unknown, the present invention can accurately predict the function. Furthermore, the present invention can predict the function of a chimeric protein without conducting extensive experiments on the chimeric protein. As a result, the present invention can efficiently develop therapeutic drugs using chimeric antigen receptors with high therapeutic effects.

[0056] Examples of this embodiment will be described with reference to Figures 7 to 9. (i) Figure 7 shows experimental results for multiple mutant sequences obtained by mutating a known chimeric protein sequence (sequence ID = FMC63-28Z). As shown in Figure 7, 100 (100 types) mutant sequences (the horizontal axis of Figure 7: the number of the chimeric sequence obtained by mutation) were created from the known chimeric protein sequence, and assays were performed on each of them to determine the cytotoxicity value (biological activity value) indicating the extent to which specific target cancer cells were killed.

[0057] According to this experiment, 50 of the 100 mutant sequences had higher cytotoxic activity values ​​than the original sequence (sequence ID = FMC63-28Z), and 50 had lower cytotoxic activity values. These mutant sequences correspond to the training chimeric protein sequence data 5121 shown in Figure 5B, and by using this to further train the retrained protein language model, it becomes possible to construct a protein language model (the above-mentioned correlation learning model) that is more suitable for predicting biological activity values.

[0058] (ii) Figure 8 is a diagram visualizing the extended sequences corresponding to each threshold (0.2, 0.4, 0.6, and 0.8 are set as the criteria (thresholds) of the standardized Levenshtein distance, which indicates similarity). Specifically, Figure 8 shows a graph in which 10,000 extended sequences are selected for each threshold from a large number of extended sequences generated based on the source sequences (known chimeric protein sequences) when additionally training an existing protein language model, and these are converted into numerical vectors and visualized in two dimensions (using the dimensionality reduction method UMAP).

[0059] 8, it can be seen that the larger the threshold (standardized Levenshtein distance) is set, the more the expanded sequence changes from the original sequence. As such, it was confirmed that the larger the threshold is set, the more the plot of the expanded sequence varies, and therefore more diverse expanded sequences are generated.

[0060] (iii) Figure 9 shows the scores (prediction results) of an existing protein language model without additional training, and the scores (classification results) obtained using the correlation learning model generated by training an existing protein language model using the extended sequences (Figure 8) generated based on each threshold (four thresholds) and the training chimeric protein sequence data shown in Figure 7 to generate the correlation learning model. Here, the score indicates the AUC score, and a higher value indicates better classification performance (two-group classification). Also, Figure 9 compares prediction results using the existing protein language model ESM-2, with the number of parameters set to 8 million and 150 million, respectively. Here, the parameter specifically refers to the learnable weight included in the protein language model. In Figure 9, p is the p-value, which indicates the degree of difference between two data sets (the result for None and the result for each threshold). The p-value is an indicator that a smaller value indicates a statistically significant difference, making it possible to determine whether there is a difference in the experimental data. In FIG. 9, NS stands for Not Significant, which means that there is no significant difference.

[0061] Figure 9 shows that classification performance improves with additional training compared to without additional training. In this experiment, with ESM-2 (number of parameters: 8 million), significant differences were observed for thresholds between 0.2 and 0.8 compared to without additional training, indicating improved predictive performance. Furthermore, with ESM-2 (number of parameters: 150 million), significant differences were observed for a threshold of 0.2. These results demonstrate that the relationship between the standardized Levenshtein distance and the number of parameters contributes to predictive performance. However, this is limited to this case, and a large number of parameters does not necessarily result in poor results. In general, the more parameters there are, the better the numerical information that can be obtained.

[0062] Figure 9 shows that, at least when two protein language models with different numbers of parameters are tested, the tendency of the classification results changes depending on the sequence used, and therefore, the process of setting a threshold and selecting a sequence can be expected to improve classification performance. others

[0063] The functions of this embodiment and each example can also be realized by software program code. In this case, a storage medium on which the program code is recorded is provided to a system or device, and the computer (or CPU or MPU) of the system or device reads the program code stored in the storage medium. In this case, the program code itself read from the storage medium realizes the functions of the above-mentioned embodiments, and the program code itself and the storage medium on which it is stored constitute the present disclosure. Examples of storage media for providing such program code include flexible disks, CD-ROMs, DVD-ROMs, hard disks, optical disks, magneto-optical disks, CD-Rs, magnetic tapes, non-volatile memory cards, and ROMs.

[0064] Furthermore, an operating system (OS) running on a computer may perform some or all of the actual processing based on instructions in the program code, and the functions of the above-described embodiments may be realized by this processing.Furthermore, after the program code is read from a storage medium and written to memory on the computer, a CPU of the computer may perform some or all of the actual processing based on instructions in the program code, and the functions of the above-described embodiments may be realized by this processing.

[0065] Furthermore, the program code of the software that realizes the functions of the embodiments and each example may be distributed via a network and stored in a storage means such as a hard disk or memory of the system or device, or in a storage medium such as a CD-RW or CD-R, so that when used, the computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage means or storage medium.

[0066] The processes and techniques described herein are not inherently related to any specific device and can be implemented by a combination of components. Various types of general-purpose devices can also be added. A dedicated device may be constructed to perform the functions of this embodiment and each example. Various functions can also be formed by appropriately combining multiple components disclosed in this embodiment and each example. For example, some components may be omitted from all the components shown in the embodiment and each example, or components from different examples may be appropriately combined.

[0067] Although specific examples are described in this disclosure, they are in all respects for the purpose of explanation (understanding the technology of the present disclosure) and not for the purpose of limitation. Those skilled in the art will recognize that there are many combinations of hardware, software, and firmware suitable for implementing the technology of the present disclosure. For example, the software described can be implemented in a wide variety of programming or scripting languages, such as assembler, C / C++, Perl, Shell, PHP, Java (registered trademark), etc.

[0068] Furthermore, in the above-described embodiment, the control lines and information lines are those that are considered necessary for the explanation, and not all control lines and information lines in the product are necessarily shown. All components may be interconnected.

[0069] In addition, other implementations of the present disclosure will be apparent to those skilled in the art from consideration of the present embodiments and examples. The specification and examples are exemplary only, with the scope and spirit of the technology of the present disclosure being indicated by the following claims.

[0070] 101 Natural protein sequence database (open large-scale protein database) 102 Trained protein language model 103 Known chimeric protein sequence (known chimeric sequence) 104, 511 Extended sequence group 105 Retrained protein language model 200 Language model retraining process / chimeric protein analysis processing device (information processing device) 201 Processor 2011 Extraction unit 2012 Generation unit 2013 Learning unit 2014 Prediction unit 202 Storage device 2021 Related protein database 2022 Language model database 203 Input device 204 Output device 205 Communication device 501 Extended sequence generation GUI 502 Extension source sequence information display field 503 Constituent domain information display field 504 Threshold input field 505 Execute button 510 Known chimeric protein sequence (extension source sequence) data 512 CAR language model analysis / adjustment input data 5121 Chimeric protein sequence data for learning 5122 Chimeric protein sequence data for validation 513 Chimeric protein sequence data including activity value (predicted value)

Claims

1. A protein language model processing device that generates a chimeric protein language model based on an existing protein language model, comprising: a storage device that stores a program for training the existing protein language model; and a processor that reads the program from the storage device and executes a process of generating the chimeric protein language model based on the program, wherein the processor executes the following processes: a process of acquiring known sequence data representing a known chimeric protein; a process of referring to a given natural protein database and generating an extended sequence data group representing a plurality of related proteins that have a homology to the known sequence data of a given threshold or more; and a process of generating the chimeric protein language model by re-training the existing protein language model using the extended sequence data group as training data.

2. A protein language model processing device according to claim 1, wherein the process of generating the extended sequence data group comprises: a process of defining each domain of the known sequence data; a process of extracting, for each of the domains, domain sequences of the related proteins having a homology equal to or greater than the threshold value from the given natural protein database; and a process of generating the extended sequence data group by selecting and combining one of the multiple domain sequences for each of the domains.

3. A protein language model processing device according to claim 2, wherein the processor extracts from the given natural protein database domain sequences that are 0.5 to 1.5 times the sequence length of each domain in the known sequence data.

4. A protein language model processing device as set forth in claim 1, wherein the processor further executes a process of further training the chimeric protein language model using training chimeric protein sequence data including a plurality of mutant sequence data obtained by inserting at least one mutation into the known sequence data and a plurality of measured activity values ​​corresponding to the plurality of mutant sequence data, and generating a correlation training model between the chimeric protein and the activity value.

5. A protein language model processing device according to claim 4, wherein the training chimeric protein sequence data includes the plurality of mutant sequence data that exhibit, as the measured activity value, a value higher than the measured activity value of the known chimeric protein.

6. A chimeric protein analysis apparatus that predicts the activity value of a target chimeric protein in a target cell using a chimeric protein language model, comprising: a storage device that stores the chimeric protein language model; and a processor that generates a predicted activity value of the target chimeric protein by applying the chimeric protein language model read from the storage device to sequence data of the target chimeric protein, wherein the chimeric protein language model is a related protein that has a homology equal to or greater than a given threshold with known sequence data representing a known chimeric protein, and is a language model that has been generated by re-learning an existing protein language model using a group of extended sequence data representing a plurality of related proteins as training data; and the processor outputs the predicted activity value of the target chimeric protein via an output device.

7. A chimeric protein analysis device according to claim 6, wherein the chimeric protein language model is a language model generated by extracting, for each domain of the known sequence data, domain sequences of the related proteins having a homology equal to or greater than the threshold value from a given natural protein database, and re-learning the existing protein language model using the extended sequence data group generated by selecting and combining one from the multiple domain sequences for each domain as the training data.

8. A chimeric protein analysis device as claimed in claim 6, wherein the processor generates a predicted activity value of the target chimeric protein by applying a chimeric protein and activity value correlation learning model to the sequence data of the target chimeric protein, the chimeric protein being generated by further learning the chimeric protein language model using training chimeric protein sequence data, instead of the chimeric protein language model, which training chimeric protein sequence data includes a plurality of mutant sequence data obtained by inserting at least one mutation into the known sequence data and a plurality of measured activity values ​​corresponding to the plurality of mutant sequence data.

9. A protein language model processing method for generating a chimeric protein language model based on an existing protein language model, comprising: a processor reading, from a storage device, a program for training the existing protein language model; and the processor generating the chimeric protein language model based on the program, wherein generating the chimeric protein language model comprises: acquiring known sequence data representing a known chimeric protein; referring to a given natural protein database, generating an extended sequence data group representing a plurality of related proteins that have a homology to the known sequence data of a given threshold or more; and generating the chimeric protein language model by re-training the existing protein language model using the extended sequence data group as training data.

10. A protein language model processing method according to claim 9, wherein generating the extended sequence data group comprises: defining each domain of the known sequence data; extracting, for each domain, domain sequences of related proteins having a homology equal to or greater than the threshold value from the given natural protein database; and generating the extended sequence data group by selecting and combining one domain sequence from the plurality of domain sequences for each domain.

11. A protein language model processing method according to claim 10, wherein the processor extracts from the given natural protein database domain sequences that are 0.5 to 1.5 times the sequence length of each domain in the known sequence data.

12. A protein language model processing method according to claim 9, wherein generating the chimeric protein language model further comprises further training the chimeric protein language model using training chimeric protein sequence data including a plurality of mutant sequence data obtained by inserting at least one mutation into the known sequence data and a plurality of measured activity values ​​corresponding to the plurality of mutant sequence data, to generate a correlation training model between the chimeric protein and the activity value.

13. A protein language model processing method according to claim 12, wherein the training chimeric protein sequence data includes the plurality of mutant sequence data that show, as the measured activity value, a value higher than the measured activity value of the known chimeric protein.

Citation Information

Patent Citations

  • System and method for discovering predicted site-specific protein phosphorylation candidates

    JP2018206374A

  • Systems and methods for predicting proteins

    US20220375539A1

  • Immunogen selection

    WO2022235853A1