Model training method, chimeric antigen receptor optimization method, device and electronic equipment

By using the protein language model ESM to predict the PCP value of protein amino acid sequences, the problem of inability to batch process multiple sequences and automatically provide optimization strategies in the prior art is solved, and rapid automated optimization of CAR-T cells is achieved, which improves R&D efficiency and efficacy.

CN120220813APending Publication Date: 2025-06-27SHANGHAI TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311818104.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art cannot batch process multiple protein amino acid sequences and automatically provide optimization strategies, resulting in poor persistence and depletion of CAR-T cells.

Method used

Through the model training method, the protein language model ESM is used to predict the PCP value of protein amino acid sequence, which simplifies the calculation process of converting protein 3D structure into PCP value, and realizes a rapid and automated optimization strategy.

Benefits of technology

It greatly simplifies the calculation of PCP value, improves the calculation efficiency, and can quickly and automatically batch optimize the amino acid sequence to be processed, improving the research and development efficiency and efficacy of CAR-T cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220813A_ABST
    Figure CN120220813A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, a chimeric antigen receptor optimization method and device and electronic equipment, and the method comprises the steps: obtaining sample data, the sample data comprises an amino acid sequence of a protein and a PCP value corresponding to the amino acid sequence, and the PCP value is used for representing the number of amino acids in a positive charge plaque on the surface of the protein; and inputting the sample data into the model for training to obtain a trained prediction model. According to the chimeric antigen receptor optimization method and device, the optimized amino acid sequence for designing the chimeric antigen receptor is provided based on the prediction model trained by the model training method, so that the condition that the PCP value of the protein needs to be converted into a 3D structure for calculation is avoided; according to the method, the amino acid sequence of the to-be-optimized protein can be quickly and automatically optimized in batches based on the PCP value of the protein, the candidate amino acid sequence in the proper PCP value range is obtained, the candidate amino acid sequence is used for designing CAR-T, the research and development process of the CAR-T is greatly accelerated, and the research and development efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and particularly to a model training method, a chimeric antigen receptor optimization method, a device, and an electronic device. Background Art

[0002] Chimeric Antigen Receptor T-Cell (CAR-T cell) immunotherapy is a tumor immunotherapy that enables T cells to recognize and kill tumor cells in a patient by expressing a chimeric antigen receptor (CAR) on the T cells. Although several marketed CAR-T drugs have shown significant efficacy in hematological tumors, their efficacy in treating solid tumors has been poor. One important reason is the poor persistence and easy exhaustion of CAR-T cell function.

[0003] Research has found that the tonic signal of CAR-T cells plays a key role in controlling their cell adaptability. A low level of tonic signal leads to poor persistence of CAR-T cells, while an excessive tonic signal leads to CAR-T cell exhaustion. Recent research has found that the electrostatic interaction mediated by positively charged patches (PCP) on the surface of the CAR receptor is an important mechanism for the generation of tonic signal, and the PCP value is used to characterize the number of amino acids in the positively charged patch on the protein surface. CAR-T designers can adjust the PCP value to change the intensity of the tonic signal, thereby rationally designing the structure of the CAR, greatly accelerating the R & D efficiency and process of CAR-T, and promoting the development and progress of the field of CAR-T treatment research for tumors, especially solid tumors.

[0004] However, the current methods for calculating the PCP value and the method for optimizing the CAR structure using the PCP value are very complex. It is necessary to obtain the 3D (3-Dimensions) structure of the protein based on the amino acid sequence, and then calculate the PCP value based on this protein 3D structure. After manually selecting amino acids for mutation to optimize the CAR, it is necessary to recalculate the protein 3D structure corresponding to the mutated amino acid sequence and then calculate the PCP value. Therefore, the calculation time is long, multiple sequences cannot be calculated in batches, and an optimization strategy cannot be automatically provided. Therefore, more automated and fast calculation methods and optimization strategies are needed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the defect in the prior art that multiple sequences cannot be processed in batches and an optimization strategy cannot be automatically provided, and to provide a model training method, a chimeric antigen receptor optimization method, and a device.

[0006] The present invention solves the above technical problems through the following technical solutions: A model training method includes:

[0007] Obtain sample data, where the sample data includes the amino acid sequence of a protein and its corresponding PCP value, and the PCP value is used to characterize the number of amino acids in the positively charged patches on the surface of the protein;

[0008] Input the sample data into a model for training to obtain a trained prediction model, and the model is used to predict the corresponding PCP value according to the amino acid sequence of the protein.

[0009] Preferably, the obtaining of the sample data includes:

[0010] Obtain a protein 3D structure data set;

[0011] Calculate the PCP value of each protein 3D structure data in the protein 3D structure data set;

[0012] Screen the protein 3D structure data set to obtain a candidate protein 3D structure data set;

[0013] Match the amino acid sequence corresponding to the protein 3D structure data in the candidate protein 3D structure data set with the corresponding PCP value to form the sample data.

[0014] Preferably, the calculating of the PCP value of each protein 3D structure data in the protein 3D structure data set includes:

[0015] Divide the protein 3D structure into multiple grid regions according to a preset region size;

[0016] Calculate the electrostatic potential of each grid region;

[0017] Screen the grid regions located on the protein surface and with a positive electrostatic potential to obtain the first surface grid regions;

[0018] Connect the adjacent grid regions in the first surface grid regions in sequence to obtain multiple second surface grid regions;

[0019] Obtain the largest three of the second surface grid regions, and calculate the number of overlapping amino acids respectively;

[0020] Add the number of amino acids corresponding to the largest three of the second surface grid regions as the PCP value of the protein 3D structure data.

[0021] Preferably, the screening of the protein 3D structure data set to obtain a candidate protein 3D structure data set includes:

[0022] Screen for single-stranded structures, and ensure that both the PCP value and the sequence length are within the threshold range. After removing duplicate sequences, the candidate protein 3D structure dataset is obtained.

[0023] Preferably, the model is the pre-trained protein language model ESM (Evolutionary Scale Modeling).

[0024] Input the sample data into the model for training to obtain a trained prediction model. Specifically, it includes:

[0025] Randomly divide the sample data into a training set and a test set according to a preset ratio.

[0026] Use the amino acid sequence in the sample data of the training set as the input and the PCP value as the output to train the model and fine-tune the parameter weights of the model.

[0027] On the other hand, the present invention provides a method for optimizing a chimeric antigen receptor.

[0028] Input the amino acid sequence to be optimized into the prediction model to obtain its corresponding predicted PCP value.

[0029] Based on the predicted PCP value, perform mutation regulation on the amino acid sequence to be optimized to obtain a candidate amino acid sequence.

[0030] Input the candidate amino acid sequence into the prediction model to obtain its corresponding predicted PCP value.

[0031] Display the candidate amino acid sequence and its corresponding predicted PCP value.

[0032] Among them, the prediction model is trained by the model training method described in any one of the above.

[0033] Preferably, the display of the candidate amino acid sequence and its corresponding predicted PCP value includes:

[0034] Display the candidate amino acid sequence and its corresponding predicted PCP value in tabular form.

[0035] And / or

[0036] Display the predicted PCP value corresponding to the candidate amino acid sequence in the form of a density distribution map.

[0037] On the other hand, the present invention provides a model training device, including:

[0038] A sample acquisition module for acquiring sample data, where the sample data includes the amino acid sequence of a protein and its corresponding PCP value, and the PCP value is used to characterize the number of amino acids within the positively charged patches on the surface of the protein;

[0039] A training module for inputting the sample data into a model for training to obtain a trained prediction model, where the model is used to predict the corresponding PCP value based on the amino acid sequence of a protein.

[0040] On the other hand, the present invention provides a chimeric antigen receptor optimization device, including,

[0041] A sequence acquisition module for acquiring the amino acid sequence to be optimized;

[0042] A prediction module for inputting the amino acid sequence to be optimized into the prediction model to obtain its corresponding predicted PCP value;

[0043] A mutation regulation module for mutating and regulating the amino acid sequence to be optimized based on the predicted PCP value to obtain a candidate amino acid sequence;

[0044] A second prediction module for inputting the candidate amino acid sequence into the prediction model to obtain its corresponding predicted PCP value;

[0045] A display module for displaying the candidate amino acid sequence and its corresponding predicted PCP value;

[0046] Wherein, the prediction model is trained by the model training method described in any one of the above.

[0047] In yet another aspect, the present invention provides an electronic device, including:

[0048] At least one processor; and

[0049] A memory communicatively connected to the at least one processor; wherein,

[0050] The memory stores instructions for the at least one processor to execute, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of the above.

[0051] In yet another aspect, the present invention provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in any one of the above.

[0052] The positive and progressive effects of the present invention are as follows: The model training method, chimeric antigen receptor optimization method, device and electronic device of the present invention train a protein language model to directly predict the PCP value based on the amino acid sequence of the protein, avoiding the need to convert the amino acid sequence of the protein into a 3D structure and then calculate the PCP value, greatly simplifying the calculation method of the PCP value, improving the calculation efficiency, and being able to quickly, automatically and batch optimize the amino acid sequences to be optimized to obtain candidate amino acid sequences within a suitable PCP value range for the design of CAR-T, greatly accelerating the R & D process of CAR-T and improving the R & D efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic flowchart of the model training method provided in Embodiment 1 of the present invention;

[0054] Figure 2 It is a schematic flowchart of step S11 in Embodiment 1 of the present invention;

[0055] Figure 3 It is a schematic flowchart of step S112 in Embodiment 1 of the present invention;

[0056] Figure 4 It is a schematic flowchart of step S12 in Embodiment 1 of the present invention;

[0057] Figure 5 It is a schematic diagram of the effect of the training set of the prediction model trained by the model training method in Embodiment 1 of the present invention;

[0058] Figure 6 It is a schematic diagram of the effect of the validation set of the prediction model trained by the model training method in Embodiment 1 of the present invention

[0059] Figure 7 It is a schematic flowchart of the chimeric antigen receptor optimization method provided in Embodiment 2 of the present invention;

[0060] Figure 8 It is a schematic diagram showing the candidate amino acid sequences of the chimeric antigen receptor optimization method provided in Embodiment 2 of the present invention;

[0061] Figure 9 It is a schematic diagram showing another candidate amino acid sequence of the chimeric antigen receptor optimization method provided in Embodiment 2 of the present invention;

[0062] Figure 10a It is a schematic diagram of the basal signals of the amino acid sequence to be optimized before optimization and the candidate amino acid sequence after optimization by the optimization method provided in Embodiment 2 of the present invention;

[0063] Figure 10b is Figure 10aSchematic diagram of the proliferation ability of CAR-T cells after applying the amino acid sequence to be optimized and the optimized candidate amino acid sequence in it;

[0064] Figure 11 Schematic framework diagram of the model training device provided in Embodiment 3 of the present invention;

[0065] Figure 12 Schematic framework diagram of the chimeric antigen receptor optimization device provided in Embodiment 4 of the present invention;

[0066] Figure 13 Schematic block diagram of the exemplary electronic device 500 provided in Embodiment 5 of the present invention. Detailed implementation manners

[0067] The present invention will be further illustrated by way of examples below, but the present invention is not limited to the scope of the described examples.

[0068] Embodiment 1

[0069] Current research has found that the electrostatic interaction mediated by positively charged patches (PCP) on the surface of chimeric antigen receptors (CAR) is an important mechanism for the basal signal of CAR-T cells. CAR-T designers can change the basal signal intensity by adjusting the PCP value. For example, by adjusting the mutation of individual amino acids in the amino acid sequence of the protein, the PCP value of the protein can be adjusted, thereby rationally designing the structure of the chimeric antigen receptor CAR.

[0070] The PCP value is used to characterize the number of amino acids in the positively charged patches on the protein surface. Specifically, in the calculation, it can be defined as the sum of the number of amino acids overlapping in the three largest positively charged patches on the protein surface.

[0071] This embodiment provides a model training method, as Figure 1 shown, including,

[0072] S11, obtaining sample data, the sample data including the amino acid sequence of the protein and its corresponding PCP value, the PCP value being used to characterize the number of amino acids in the positively charged patches on the protein surface;

[0073] S12, inputting the sample data into the model for training to obtain a trained prediction model, the model being used to predict the corresponding PCP value according to the amino acid sequence of the protein.

[0074] A protein can be represented by its three-dimensional (3D) structure, that is, showing its true three-dimensional structure, or it can also be represented by an amino acid sequence. Both of these representation methods are commonly used for proteins. Usually, the steps of converting the amino acid sequence of a protein into its corresponding 3D structure are complex, with a large amount of calculation and time consumption. And calculating the PCP value of a protein requires the 3D structure of the protein. Therefore, it brings inconvenience to the calculation of the PCP value. The model training method provided by the embodiments of the present invention is used to directly predict the corresponding PCP value according to the amino acid sequence. The PCP value can be directly obtained by inputting the amino acid sequence and according to the output of the prediction model, greatly reducing the amount of calculation and improving the convenience of obtaining the PCP value.

[0075] The sample data includes the amino acid sequence of a protein and its corresponding PCP value.

[0076] Preferably, as Figure 2 shown, step S11 of obtaining sample data specifically includes

[0077] S111, obtaining a protein 3D structure data set;

[0078] S112, calculating the PCP value of each protein 3D structure data in the protein 3D structure data set;

[0079] S113, screening the protein 3D structure data set to obtain a candidate protein 3D structure data set;

[0080] S114, matching the amino acid sequence corresponding to the protein 3D structure data in the candidate protein 3D structure data set with the corresponding PCP value to form the sample data.

[0081] In step S111, to obtain a protein 3D structure data set, it can be by downloading a PDB file or a FASTA file from an existing protein database. For example, it can be the experimentally obtained protein 3D structure data from the RCSB PDB database, and / or the predicted protein 3D structure data from the Uniprot or Alphafold database, thereby forming a protein 3D structure data set.

[0082] In step S112, calculating the PCP value corresponding to each protein 3D structure data in the protein structure data set is for screening the protein 3D structure data set and forming sample data. Specifically, as Figure 3 shown, this step includes

[0083] S1121, divide the 3D protein structure into multiple grid regions according to a preset region size; preferably, in this step, a cube with an edge length of 1 angstrom can be used as the preset region size to divide the 3D protein structure into multiple grid regions.

[0084] In some embodiments, in this step, all non-standard residues can also be removed first, and hydrogen atoms can be added to the 3D protein structure to facilitate the calculation of the electrostatic potential of each subsequent grid region. Preferably, in this step, Pdb-tools can be used to delete all HETATM records in the PDB file that record non-standard residues, and the pdb2pqr30 software can be used to add hydrogen atoms.

[0085] S1122, calculate the electrostatic potential of each grid region; specifically, the electrostatic potential of each grid region can be calculated by the APBS software.

[0086] S1123, screen the grid regions located on the protein surface and with a positive electrostatic potential to obtain multiple first surface grid regions; specifically, the electrostatic potential 2kT / e can be used as a threshold to filter out the grid regions with a positive electrostatic potential, and the DMS software (a software for calculating the molecular surface) based on the Lee and Richards algorithm can be used to calculate the protein surface accessibility, so as to obtain the grid regions falling on the protein surface, ignoring all non-surface grid regions, and then obtain multiple first surface grid regions.

[0087] S1124, connect the adjacent grid regions in the first surface grid regions in sequence to obtain multiple second surface grid regions; in this step, using the KD-tree algorithm, search for multiple first surface grid regions, connect the adjacent grid regions in multiple first surface grid regions to each other in sequence, and then obtain multiple independent and separate second surface grid regions. The sizes of the second surface grid regions are different, that is, the number of grid regions included in each second surface grid region is different. This second surface grid region is the positive charge patch in the 3D protein structure.

[0088] S1125, obtain the largest three second surface grid regions and calculate the number of overlapping amino acids respectively; in this step, the largest three second surface grid regions are the three second surface grid regions with the largest number of grid regions. Calculate the number of overlapping amino acids respectively, specifically, calculate the sum of the number of amino acids in this second surface grid region and the number of amino acids within 1 angstrom of its periphery. This method can also be obtained by using the KD-tree algorithm for searching.

[0089] S1126, add the number of amino acids corresponding to the largest three second surface grid regions as the PCP value of the 3D protein structure data.

[0090] In step S113, the protein 3D structure dataset is screened to obtain a candidate protein 3D structure dataset, which includes screening for single-stranded structures, and both the PCP value and the sequence length are within the threshold range. After removing duplicate sequences, the candidate protein 3D structure dataset is obtained.

[0091] For the threshold range of the PCP value, a suitable PCP value can be selected. For example, after processing the PCP value through a normal distribution, values greater than 10 are removed. That is, after normal distribution processing, the PCP Z = log2PCP, and all protein 3D structure data corresponding to PCP values greater than 10 are removed. Z greater than 10 are removed.

[0092] For the threshold range of the sequence length, it can be selected as 100 - 400 to obtain a candidate protein 3D structure dataset with a suitable length and structure for model training.

[0093] Match the amino acid sequences corresponding to the protein 3D structure data in the candidate protein 3D structure dataset with the corresponding PCP values to form the sample data. Specifically, for each protein 3D structure data in the candidate protein 3D structure dataset, the corresponding amino acid sequence is paired one-to-one with its corresponding PCP value to form the sample data.

[0094] In step S12, the model is a pre-trained protein language model ESM (Evolutionary Scale Modeling), preferably, the model ESM2 is selected. The sample data is input into the model for training to obtain a trained prediction model. The model is used to predict the corresponding PCP value according to the amino acid sequence of the protein, as Figure 4 shown, specifically including,

[0095] S121, randomly divide the sample data into a training set and a test set according to a preset ratio; preferably, the preset ratio is 7:3;

[0096] S122, use the amino acid sequences in the sample data of the training set as the input and the PCP values as the output to train the model and fine-tune the parameter weights of the model.

[0097] The protein language model ESM (Evolutionary Scale Modeling) is a method that uses deep learning technology to predict protein structure and function. ESM trains an autoregressive neural network on a large-scale protein sequence database to learn the evolutionary laws of proteins and the relationships between sequence-structure-function. Given a protein sequence, ESM can generate its corresponding hidden vector, representing the characteristics of its structure and function, and can also use the hidden vector for various downstream tasks, such as structure prediction, function annotation, interaction analysis, etc.

[0098] Preferably, the model selects the ESM2-8M model, which has 6 Transformer layers, 8 million parameters, and the amino acid embedding vector length is 320 dimensions. To adapt to the PCP prediction task, a fully connected layer is added at the end of the pre-trained model to make the output dimension 1 for fine-tuning the regression task. Fine-tuning the parameter weights of the model means continuing to update its weights based on a specific task on the basis of the pre-trained model weights of ESM2, that is, predicting the PCP value in this embodiment. The fine-tuning training process is trained for a total of 15 epochs. The AdamW optimizer is used to update the parameter weights of the model, where β1 = 0.9, β2 = 0.999, ε = 10-8, the initial learning rate is set to 0.0002, and the weight decay constant is 0.001. The batch size during training is 64, that is, there are 64 sequences in each small batch. The loss function of the fine-tuning task is the mean squared error (MSE):

[0099]

[0100] where N is the size of a small batch, y i is the label of the sequence, that is, the PCP value corresponding to the amino acid sequence in the test sample, is the model prediction value, that is, the PCP value corresponding to the amino acid sequence predicted by the model.

[0101] As Figure 5 shown, it is a schematic diagram of the effect of the test set of the prediction model trained by the model training method of this embodiment, where the X-axis is the PCP value calculated based on the protein 3D structure, and the Y-axis is the value predicted by using the model based on the sequence.

[0102] As Figure 6 shown, it is a schematic diagram of the effect of the validation set of the prediction model trained by the model training method of this embodiment, where the X-axis is the PCP value calculated based on the protein 3D structure, and the Y-axis is the value predicted by using the model based on the sequence.

[0103] The model training method of this embodiment trains a protein language model to directly predict the PCP value based on the amino acid sequence of a protein, avoiding the need to convert the amino acid sequence of the protein into a 3D structure before calculating the PCP value, greatly simplifying the method of calculating the PCP value, improving the calculation efficiency, and reducing the operation processing time.

[0104] Example 2

[0105] As Figure 7 shown, this embodiment discloses a chimeric antigen receptor optimization method, including,

[0106] S21, input the amino acid sequence to be optimized into the prediction model to obtain its corresponding predicted PCP value;

[0107] S22, perform mutation regulation on the amino acid sequence to be optimized based on the predicted PCP value to obtain a candidate amino acid sequence;

[0108] S23, input the candidate amino acid sequence into the prediction model to obtain its corresponding predicted PCP value;

[0109] S24, display the candidate amino acid sequence and its corresponding predicted PCP value;

[0110] wherein, the prediction model is trained by the model training method in the above-mentioned Embodiment 1.

[0111] Select the amino acid sequence to be optimized. Based on the above prediction model, the PCP value corresponding to the amino acid sequence to be optimized can be predicted. Perform mutation regulation on the amino acid sequence to be optimized based on the predicted PCP value to obtain a candidate amino acid sequence, that is, by mutating individual amino acids in the amino acid sequence to be optimized to obtain a candidate amino acid sequence. For example, if you want to increase the PCP value, mutate some amino acids Q to K; if you want to decrease the PCP value, you can mutate some amino acids K to Q. When selecting the amino acids for this mutation regulation, Q amino acids or K amino acids that are relatively far from the Complementarity determining regions (CDRs) should be selected.

[0112] Input the candidate amino acid sequence into the prediction model to obtain its corresponding predicted PCP value; based on the prediction model, the PCP values corresponding to the candidate amino acid sequences can be obtained batchwise and quickly.

[0113] Displaying the candidate amino acid sequence and its corresponding predicted PCP value may include displaying the candidate amino acid sequence and its corresponding predicted PCP value in tabular form;

[0114] and / or,

[0115] The predicted PCP values corresponding to the candidate amino acid sequences are presented in the form of a density distribution map.

[0116] The candidate amino acid sequences and their corresponding predicted PCP values shown above are the optimization strategies recommended by the CAR-T optimization method.

[0117] As Figure 8 shown, it is a schematic diagram of the candidate amino acid sequences presented in tabular form. As Figure 9 shown, it is the candidate amino acid sequences presented in the form of the PCP distribution. The two vertical lines therein are the optimal interval 46 - 56 of the PCP value.

[0118] The following is a specific implementation manner for specifically optimizing a chimeric antigen receptor by selecting an amino acid sequence to be optimized. In this implementation manner, a VHH (variable domain of heavy chain of heavy-chain antibody) sequence is selected as the sequence to be optimized (named A9-CCL1). The amino acid sequence to be optimized is specifically:

[0119] EVQLVESGGGLVQPGGSLRLSCAASGIIFSIYSMGWFRQAPGKGREL VAAISRRGSYTYYPDSVEGRFTISRDNAKRMVYLQMNSLRAEDTAVYYC AAARVWSTWRSRRDYNYWGQGTQVTVSS.

[0120] Through the optimization algorithm of this example, the PCP values and basal signal values of this sequence and its optimized candidate sequences M1, M2, M3, and M4 are shown respectively. As Figure 10a shown, the amino acid sequence to be optimized is represented by WT, and the PCP value is 70. The user can choose to mutate some K amino acids in the FR region to Q amino acids to adjust its corresponding PCP value to the optimal interval 46 - 56 shown in the figure, obtaining four candidate amino acid sequences M1, M2, M3, and M4.

[0121] After constructing the VHH of the optimized mutant A9-CLL1 into a CAR and expressing it in primary T cells, then detecting its basal signal level and calculating the tonic signaling index, it is confirmed that the VHH mutant calculated and optimized by this method can effectively regulate the basal signal level of the CAR. The PCP values and basal signal values corresponding to each candidate amino acid sequence are shown in Figure 10a . The abscissa is the PCP value, and the ordinate is the basal signal value.

[0122] As shown Figure 10b in the figure, the abscissa in the figure is the PCP value, and the ordinate represents the proliferation ability of the corresponding CAR-T cells. The CAR-T cells are repeatedly stimulated with the target cells THP-1 expressing the CLL1 protein, and their proliferation ability in the continuous presence of tumor antigens is detected, further confirming that regulating the basal signal of the CAR to an appropriate level is most conducive to the continuous proliferation and anti-tumor activity of CAR-T cells.

[0123] This chimeric antigen receptor optimization method, using the prediction model trained in Example 1, can quickly and automatically optimize the amino acid sequence to be optimized, obtain a candidate amino acid sequence within a suitable PCP value range for the design of CAR-T, greatly accelerating the R & D process of CAR-T and improving the R & D efficiency.

[0124] Example 3

[0125] This example provides a model training device, as shown Figure 11 in the figure, including

[0126] A sample acquisition module 10 for acquiring sample data, where the sample data includes the amino acid sequence of a protein and its corresponding PCP value, and the PCP value is used to characterize the number of amino acids in the positively charged patch on the surface of the protein;

[0127] A training module 20 for inputting the sample data into a model for training to obtain a trained prediction model, and the model is used to predict the corresponding PCP value according to the amino acid sequence of the protein.

[0128] The sample data includes the amino acid sequence of the protein and its corresponding PCP value.

[0129] The sample acquisition module 10 further includes

[0130] An acquisition module 11 for acquiring a protein 3D structure data set; acquiring a protein 3D structure data set can be by downloading a PDB file or a FASTA file through an existing protein database, for example, it can be experimental protein 3D structure data from the RCSB PDB database, and / or protein 3D structure data predicted from the Uniprot or Alphafold database, thereby forming a protein 3D structure data set.

[0131] A calculation module 12 for calculating the PCP value of each protein 3D structure data in the protein 3D structure data set; this calculation module 12 specifically includes

[0132] A partitioning module 121 for partitioning the protein 3D structure into a plurality of grid regions according to a preset region size; preferably, in this step, a side length threshold of 1 angstrom can be used as the preset region to partition the protein 3D structure into a plurality of grid regions; in some embodiments, in this step, all non-standard residues can also be removed first, and hydrogen atoms can be added to the protein 3D structure to facilitate the calculation of the electrostatic potential of each subsequent grid region. Preferably, in this step, Pdb-tools can be used to delete all HETATM records in the PDB file that record non-standard residues, and the pdb2pqr30 software can be used to add hydrogen atoms.

[0133] An electrostatic potential calculation module 122 for calculating the electrostatic potential of each grid region; specifically, the electrostatic potential of each grid region can be calculated by the APBS software.

[0134] A first screening module 123 for screening the grid regions located on the protein surface and having a positive electrostatic potential to obtain a plurality of first surface grid regions; specifically, the electrostatic potential 2kT / e can be used as a threshold to filter out the grid regions with positive electrostatic potential, and the DMS software based on the Lee and Richards algorithm can be used to calculate the protein surface accessibility, so as to obtain the grid regions falling on the protein surface, ignoring all non-surface grid regions, and then obtaining a plurality of first surface grid regions.

[0135] A connection module 124 for sequentially connecting the adjacent grid regions in the first surface grid regions to obtain a plurality of second surface grid regions; in this step, using the KD tree algorithm, the plurality of first surface grid regions are searched, and the adjacent grid regions in the plurality of first surface grid regions are sequentially connected to each other, and then a plurality of independent and separate second surface grid regions are obtained. The sizes of the second surface grid regions are different, that is, the number of grid regions included in each second surface grid region is different. This second surface grid region is the positive charge patch in the protein 3D structure.

[0136] A determination module 125 for obtaining the largest three of the second surface grid regions and respectively calculating the number of overlapping amino acids; in this step, the largest three of the second surface grid regions are the three second surface grid regions with the largest number of grid regions, and the number of overlapping amino acids is respectively calculated. Specifically, it is to calculate the sum of the number of amino acids in the second surface grid region and the number of amino acids within 1 angstrom of its periphery. This step can also be obtained by using the KD tree algorithm for searching.

[0137] The second screening module 13 is used to screen the protein 3D structure dataset to obtain a candidate protein 3D structure dataset; screen the protein 3D structure dataset to obtain a candidate protein 3D structure dataset; including screening for a single-stranded structure, and both the PCP value and the sequence length are within the threshold range. After removing duplicate sequences, the candidate protein 3D structure dataset is obtained.

[0138] The threshold range of the PCP value can select a suitable PCP value. For example, after processing the PCP value through a normal distribution, values greater than 10 are removed. That is, after the normal distribution processing, the PCP Z = log2PCP, and remove the PCP Z corresponding to all protein 3D structure data with PCP values greater than 10.

[0139] The threshold range of the sequence length can be selected as 100 - 400 to obtain a candidate protein 3D structure dataset with a suitable length and structure for model training.

[0140] The matching module 14 is used to match the amino acid sequence corresponding to the protein 3D structure data in the candidate protein 3D structure dataset with the corresponding PCP value to form the sample data; specifically, the amino acid sequence corresponding to each protein 3D structure data in the candidate protein 3D structure dataset is corresponded one by one with its corresponding PCP value to form the sample data.

[0141] The training module 20 selects the model as the pre-trained protein language model ESM (Evolutionary Scale Modeling), preferably, selects this model as ESM2. Input the sample data into the model for training to obtain a trained prediction model. The model is used to predict the corresponding PCP value according to the amino acid sequence of the protein, and fine-tune the parameters of the model.

[0142] The protein language model ESM (Evolutionary Scale Modeling) is a method that uses deep learning technology to predict protein structure and function. ESM trains an autoregressive neural network on a large-scale protein sequence database to learn the evolutionary laws of proteins and the relationships between sequence-structure-function. ESM can generate the corresponding hidden vector for a given protein sequence, representing the characteristics of its structure and function, and can also use the hidden vector for various downstream tasks, such as structure prediction, function annotation, interaction analysis, etc.

[0143] It should be noted that the working principle of the model training device in this embodiment is the same as that of the model training method in the above embodiment 1, so it will not be elaborated here.

[0144] The model training device of this embodiment is used to train a protein language model to directly predict the PCP value based on the amino acid sequence of a protein, avoiding the need to convert the amino acid sequence of the protein into a 3D structure before calculating the PCP value, greatly simplifying the method of calculating the PCP value and improving the calculation efficiency.

[0145] Example 4

[0146] As Figure 12 shown, the chimeric antigen receptor optimization device disclosed in this embodiment includes,

[0147] The first prediction module 21 is used to input the amino acid sequence to be optimized into the prediction model to obtain its corresponding predicted PCP value;

[0148] The adjustment module 22 is used to perform mutation adjustment on the amino acid sequence to be optimized based on the predicted PCP value to obtain a candidate amino acid sequence;

[0149] The second prediction module 23 inputs the candidate amino acid sequence into the prediction model to obtain its corresponding predicted PCP value;

[0150] The display module 24 displays the candidate amino acid sequence and its corresponding predicted PCP value.

[0151] Among them, the prediction model is trained by the model training method in the above-mentioned Embodiment 1.

[0152] Select the amino acid sequence to be optimized. Based on the above-mentioned prediction model, the PCP value corresponding to the amino acid sequence to be optimized can be predicted. Perform mutation adjustment on the amino acid sequence to be optimized based on the predicted PCP value to obtain a candidate amino acid sequence, that is, by mutating individual amino acids in the amino acid sequence to be optimized to obtain a candidate amino acid sequence. For example, if you want to increase the PCP value, mutate some amino acids Q to K; if you want to decrease the PCP value, you can mutate some amino acids K to Q. When selecting the amino acids for this mutation adjustment, Q amino acids or K amino acids that are relatively far from the Complementarity determining regions (CDRs) should be selected.

[0153] The second prediction module 23 inputs the candidate amino acid sequence into the prediction model to obtain its corresponding predicted PCP value; based on the prediction model, the PCP values corresponding to the candidate amino acid sequences can be obtained batchwise and quickly.

[0154] The display of the candidate amino acid sequence and its corresponding predicted PCP value may include displaying the candidate amino acid sequence and its corresponding predicted PCP value in tabular form;

[0155] and / or,

[0156] displaying the predicted PCP value corresponding to the candidate amino acid sequence in the form of a density distribution map.

[0157] The candidate amino acid sequence and its corresponding predicted PCP value shown above are the optimization strategies recommended by the CAR-T optimization method.

[0158] The chimeric antigen receptor optimization device, using the prediction model trained in Example 1, can quickly and automatically optimize the amino acid sequence to be optimized, obtain candidate amino acid sequences within a suitable PCP value range for the design of CAR-T, and greatly accelerate the R & D efficiency and process of CAR-T.

[0159] Example 5

[0160] According to Example 5 of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product. Figure 13 FIG. shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement the embodiments of the present invention. The device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 into a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0161] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0162] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the model training method or the chimeric antigen receptor optimization method. For example, in some embodiments, the model training method or the chimeric antigen receptor optimization method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the methods described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the aforementioned model training method or chimeric antigen receptor optimization method in any other suitable manner (e.g., by means of firmware).

[0163] Those skilled in the art will appreciate that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0164] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0165] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0166] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that this is only by way of example, and the protection scope of the present invention is defined by the appended claims. Without departing from the principle and essence of the present invention, those skilled in the art can make various changes or modifications to these embodiments, but such changes and modifications all fall within the protection scope of the present invention.

[0167] Although the specific embodiments of the present invention have been described above, those skilled in the art should understand that this is only by way of example, and the protection scope of the present invention is defined by the appended claims. Without departing from the principle and essence of the present invention, those skilled in the art can make various changes or modifications to these embodiments, but such changes and modifications all fall within the protection scope of the present invention.

Claims

1. A model training method, characterized in that, Including, Obtaining sample data, where the sample data includes the amino acid sequence of a protein and its corresponding PCP value, and the PCP value is used to characterize the number of amino acids within the positively charged patches on the protein surface; Inputting the sample data into a model for training to obtain a trained prediction model, where the model is used to predict the corresponding PCP value based on the amino acid sequence of the protein.

2. The model training method according to claim 1, characterized in that, The obtaining of the sample data includes, Obtaining a protein 3D structure data set; Calculating the PCP value of each protein 3D structure data in the protein 3D structure data set; Screening the protein 3D structure data set to obtain a candidate protein 3D structure data set; Matching the amino acid sequence corresponding to the protein 3D structure data in the candidate protein 3D structure data set with the corresponding PCP value to form the sample data.

3. The model training method according to claim 2, wherein The calculating of the PCP value of each protein 3D structure data in the protein 3D structure data set includes, Dividing the protein 3D structure into multiple grid regions according to a preset region size; Calculating the electrostatic potential of each grid region; Screening the grid regions located on the protein surface and with a positive electrostatic potential to obtain the first surface grid regions; Sequentially connecting the adjacent grid regions in the first surface grid regions to obtain multiple second surface grid regions; Obtaining the largest three of the second surface grid regions and respectively calculating the number of overlapping amino acids; Adding up the number of amino acids corresponding to the largest three of the second surface grid regions as the PCP value of the protein 3D structure data.

4. The model training method according to claim 2, characterized in that The screening of the protein 3D structure data set to obtain a candidate protein 3D structure data set includes, Screening for single-chain structures, and when both the PCP value and the sequence length are within a threshold range, removing duplicate sequences to obtain the candidate protein 3D structure data set.

5. The model training method according to claim 1, wherein The model is the pre-trained protein language model ESM, Inputting the sample data into the model for training to obtain a trained prediction model, specifically including, Randomly dividing the sample data into a training set and a test set according to a preset ratio; Using the amino acid sequence in the sample data of the training set as the input and the PCP value as the output to train the model and fine-tune the parameter weights of the model.

6. A method for optimizing a chimeric antigen receptor, characterized in that, Inputting the amino acid sequence to be optimized into the prediction model to obtain its corresponding predicted PCP value; Performing mutation regulation on the amino acid sequence to be optimized based on the predicted PCP value to obtain a candidate amino acid sequence; Inputting the candidate amino acid sequence into the prediction model to obtain its corresponding predicted PCP value; Displaying the candidate amino acid sequence and its corresponding predicted PCP value; Wherein, the prediction model is trained by the model training method according to any one of claims 1-5 above.

7. The chimeric antigen receptor optimization method according to claim 6, wherein The displaying of the candidate amino acid sequence and its corresponding predicted PCP value includes, Displaying the candidate amino acid sequence and its corresponding predicted PCP value in tabular form; And / or, Displaying the predicted PCP value corresponding to the candidate amino acid sequence in the form of a density distribution map.

8. A model training device, characterized in that, Comprising, a sample acquisition module for acquiring sample data, where the sample data includes the amino acid sequence of a protein and its corresponding PCP value, and the PCP value is used to characterize the number of amino acids within the positively charged patches on the surface of the protein; a training module for inputting the sample data into a model for training to obtain a trained prediction model, where the model is used to predict the corresponding PCP value according to the amino acid sequence of a protein.

9. A chimeric antigen receptor optimization device, characterized in that, Comprising, a sequence acquisition module for acquiring the amino acid sequence to be optimized; a prediction module for inputting the amino acid sequence to be optimized into the prediction model to obtain its corresponding predicted PCP value; a mutation regulation module for performing mutation regulation on the amino acid sequence to be optimized based on the predicted PCP value to obtain a candidate amino acid sequence; a second prediction module for inputting the candidate amino acid sequence into the prediction model to obtain its corresponding predicted PCP value; a display module for displaying the candidate amino acid sequence and its corresponding predicted PCP value; wherein, the prediction model is trained by the model training method described in any one of claims 1-5 above.

10. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions for execution by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any one of claims 1-5 or claims 6-7.