CAR-T cell phenotype prediction method, model training method, equipment and storage medium

The intrinsic characteristics and CAR structure of CAR-T cells are fusion predicted through neural network models, which solves the problem of low accuracy in the prediction of CAR-T cells in the prior art, and improves the accuracy of prediction and the effect of chimeric antigen receptor design.

CN120048342APending Publication Date: 2025-05-27BOE TECHNOLOGY GROUP CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311586782.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has low accuracy in predicting CAR-T cell phenotypes.

Method used

By entering the intrinsic feature data and target CAR structure into the trained phenotype prediction model, the neural network structure of the encoding layer, fusion layer and prediction layer can be used to obtain the hidden layer characteristics of the intrinsic features of the multi-dimensional tumor and the semantic feature sequence of the target CAR structure, and fuse it to predict the CAR-T cell phenotype.

Benefits of technology

It improves the prediction accuracy of CAR-T cell phenotype, enhances the accuracy of screening the intracellular signaling domain of chimeric antigen receptors, and helps design chimeric antigen receptors with better efficacy in tumor cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048342A_ABST
    Figure CN120048342A_ABST
Patent Text Reader

Abstract

The invention discloses a chimeric antigen receptor (CAR) T cell phenotype prediction method, a model training method, equipment and a storage medium, and belongs to the field of biomedicines.The method comprises the steps that intrinsic characteristic data and a target CAR structure are input into a trained phenotype prediction model, the internal characteristic data comprises CAR-T cells formed by transfecting CAR to T cells and multi-dimensional tumor internal characteristics embodied by tumor cells, the CAR has a target CAR structure, and the phenotype prediction model comprises a coding layer, a fusion layer and a prediction layer; acquiring hidden layer features of the internal feature data by the coding layer, and acquiring a semantic feature sequence of the target CAR structure by the coding layer; fusing the semantic feature sequence and the hidden layer feature by a fusion layer to obtain a fused feature; and inputting the fusion features into a prediction layer to predict the CAR-T cell phenotype of the target CAR structure. The technical problem that the accuracy of CAR-T cell phenotype prediction is not high is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of biomedical technologies, and in particular, to a method for predicting CAR-T cell phenotypes, a method for training a model, a device, and a storage medium. Background Art

[0002] CAR-T (Chimeric Antigen Receptor T-Cell Immunotherapy) immunotherapy is a new type of precision targeted therapy for treating tumors. In CAR-T therapy, immune cells from the patient's own body or allogeneic donors are genetically transformed and then infused into the patient's body to bind to tumor cells and destroy them. A CAR (Chimeric Antigen Receptor) is an artificial receptor molecule manufactured by genetic engineering technology, which can endow immune effector cells with specificity for a target antigen epitope, thereby enhancing the functions of immune effector cells in recognizing antigen signals and activation.

[0003] The phenotype of CAR-T cells is the performance characteristics and functions of transfecting chimeric antigen receptors into T cells. Predicting the CAR-T cell phenotype can help optimize the design and improvement of the signal domain combination of CARs. However, the current prediction accuracy of the CAR-T cell phenotype is not high. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method for predicting CAR-T cell phenotypes, a method for training a model, a device, and a storage medium, so as to solve the technical problem of low prediction accuracy of CAR-T cell phenotypes in the prior art.

[0005] In a first aspect of the present disclosure, a method for predicting CAR-T cell phenotypes is provided, including: inputting intrinsic feature data and a target CAR structure into a trained phenotype prediction model, where the intrinsic feature data includes multi-dimensional tumor intrinsic features reflected by CAR-T cells formed by transfecting CARs into T cells and tumor cells, the CAR has the target CAR structure, and the phenotype prediction model includes an encoding layer, a fusion layer, and a prediction layer; obtaining hidden layer features of the intrinsic feature data by the encoding layer, and obtaining a semantic feature sequence of the target CAR structure by the encoding layer; fusing the semantic feature sequence and the hidden layer features by the fusion layer to obtain fusion features; and inputting the fusion features into the prediction layer to predict the CAR-T cell phenotype of the target CAR structure.

[0006] In combination with the first aspect, in some embodiments, the encoding layer includes a pre-trained language model. Obtaining the semantic feature sequence of the target CAR structure by the encoding layer includes: performing one-hot encoding on each motif sequence in the target CAR structure respectively to obtain one-hot encoding vectors of each motif sequence in the target CAR structure, where the target CAR structure includes a plurality of motif sequences; encoding the one-hot encoding vectors of each motif sequence in the target CAR structure through the pre-trained language model to obtain embedding vectors of each motif sequence in the target CAR structure, and the semantic feature sequence includes the embedding vectors of each motif sequence in the target CAR structure, and the pre-trained language model is pre-trained on a protein sequence corpus.

[0007] In combination with the first aspect, in some embodiments, the encoding layer further includes a first forward neural network having a plurality of hidden layers. Obtaining the hidden layer features of the intrinsic feature data by the encoding layer includes: performing normalization processing on the multi-dimensional tumor intrinsic features respectively to correspondingly obtain a plurality of standard feature values; non-linearly combining and representing the plurality of standard feature values by the first forward neural network to obtain the hidden layer features.

[0008] In combination with the first aspect, in some embodiments, the intrinsic feature data includes at least any two-dimensional tumor intrinsic features as follows: the initial quantitative relationship between CD4-T cells and CD8-T cells, the initial quantity of CAR-T cells, the antigen affinity of CAR-T cells, the antigen density of tumor cells, and the tumor burden signal.

[0009] In combination with the first aspect, in some embodiments, before obtaining the semantic feature sequence of the target CAR structure by the encoding layer, it further includes: obtaining K motif sequences from a signal domain library, where K is an integer greater than 1, and the signal domain library includes a variety of activation motifs and a variety of co-stimulatory motifs; setting the K motif sequences between a first motif sequence and a second motif sequence to obtain the target CAR structure.

[0010] In combination with the first aspect, in some embodiments, the fusion layer fuses the semantic feature sequence and the hidden layer features to obtain fusion features, including: performing attention weighting on the semantic feature sequence and the hidden layer features to obtain a weighted sum result; performing layer normalization on the sum result of the weighted sum result and the semantic feature sequence to obtain a first output result; inputting the first output result into a second forward neural network to obtain a second output result; performing layer normalization on the sum result of the first output result and the second output result to obtain the fusion features.

[0011] In combination with the first aspect, in some embodiments, the prediction layer includes a third forward neural network. The step of inputting the fused feature into the prediction layer to predict the CAR-T cell phenotype of the target CAR structure includes: inputting the fused feature into the third forward neural network, and performing regression prediction through the third forward neural network to obtain the CAR-T cell phenotype of the target CAR structure.

[0012] In combination with the first aspect, in some embodiments, it further includes the step of pre-training the phenotype prediction model: obtaining a training sample set, the training sample set includes a plurality of training sample pairs, each training sample pair is composed of a CAR structure sample and an intrinsic feature sample, and the intrinsic feature sample of each training sample pair includes the CAR-T cells formed by transfecting a chimeric antigen receptor into T cells and the multi-dimensional tumor intrinsic features reflected by tumor cells, and the structure of the chimeric antigen receptor is the same as the CAR structure sample corresponding to the intrinsic feature sample; training the constructed neural network model based on the training sample set to obtain the trained phenotype prediction model, wherein an encoding layer, a fusion layer and a prediction layer are constructed in the neural network model.

[0013] In combination with the first aspect, in some embodiments, the training sample set at least includes CAR structure samples of a target type, and the signal domain of the CAR structure samples of the target type includes at least one activation motif and at least one co-stimulatory motif.

[0014] In combination with the first aspect, in some embodiments, the obtaining of the training sample set includes: obtaining a plurality of CAR structure samples, wherein the obtaining of each CAR structure sample includes: obtaining K motif sequences from the signal domain library, and setting the K motif sequences between a first motif sequence and a second motif sequence to obtain the CAR structure sample, where K is an integer greater than 1; for each CAR structure sample among the plurality of CAR structure samples, obtaining the corresponding CAR structure sample of the CAR structure sample.

[0015] In a second aspect of the present disclosure, a method for training a phenotype prediction model is provided, including: obtaining a training sample set, the training sample set includes a plurality of training sample pairs, each training sample pair is composed of a CAR structure sample and an intrinsic feature sample, and the intrinsic feature sample of each training sample pair includes the CAR-T cells formed by transfecting a chimeric antigen receptor into T cells and the multi-dimensional tumor intrinsic features reflected by tumor cells, and the structure of the chimeric antigen receptor is the same as the CAR structure sample corresponding to the intrinsic feature sample; training the constructed neural network model based on the training sample set to obtain a phenotype prediction model for predicting the CAR-T cell phenotype, wherein an encoding layer, a fusion layer and a prediction layer are constructed in the neural network model.

[0016] In combination with the second aspect, in some embodiments, the training sample set at least includes CAR structure samples of a target type, and the signal domain of the CAR structure samples of the target type includes at least one activation motif and at least one co-stimulatory motif.

[0017] In combination with the second aspect, in some embodiments, the obtaining of the training sample set includes: obtaining a plurality of CAR structure samples, wherein the obtaining of each CAR structure sample includes: obtaining K motif sequences from the signal domain library, and setting the K motif sequences between a first motif sequence and a second motif sequence to obtain the CAR structure sample, where K is an integer greater than 1; for each CAR structure sample among the plurality of CAR structure samples, obtaining the corresponding CAR structure sample of the CAR structure sample.

[0018] In a third aspect of the present disclosure, an electronic device is provided, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute to implement the CAR-T cell phenotype prediction method according to any one of the embodiments of the first aspect, or is configured to execute to implement the training method of the phenotype prediction model according to any one of the embodiments of the second aspect.

[0019] In a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute and implement the CAR-T cell phenotype prediction method according to any one of the embodiments of the first aspect, or to implement the training method of the phenotype prediction model according to any one of the embodiments of the second aspect.

[0020] One or more technical solutions provided in the embodiments of the present disclosure have at least the following technical effects or advantages:

[0021] In the embodiments of the present disclosure, the hidden layer features of the multi-dimensional tumor intrinsic features are fused with the semantic feature sequences of the target CAR structure, and the phenotype of the CAR-T cell with the target CAR structure is predicted based on the fused features. Since the multi-dimensional tumor intrinsic features are added, the influence of the intracellular domain of the target CAR structure on the function of the CAR-T cell can be captured more accurately during the prediction process. Therefore, the phenotype of the chimeric antigen receptor transfected into T cells and transformed into CAR-T cells with the target CAR structure can be predicted more accurately. Furthermore, the accuracy of screening the intracellular signal domain of the chimeric antigen receptor is improved, which helps to design a chimeric antigen receptor with better curative effect on tumor cells. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0023] Figure 1 The flowchart of the CAR-T cell phenotype prediction method provided by some embodiments of the present disclosure is shown;

[0024] Figure 2 The model structure of the phenotype prediction model in some embodiments of the present disclosure is shown;

[0025] Figure 3 The flowchart of the training method of the phenotype prediction model in some embodiments of the present disclosure is shown;

[0026] Figure 4 The schematic structural diagram of an electronic device in some embodiments of the present disclosure is shown. Detailed implementation manners

[0027] To better understand the above technical solutions, the following will describe the above technical solutions in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0028] First, it should be noted that the term "and / or" appearing in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the preceding and following associated objects.

[0029] Figure 1 The flowchart of the CAR-T cell phenotype prediction method provided by some embodiments of the present disclosure is shown. As Figure 1 shown, the embodiments of the present disclosure provide a CAR-T cell phenotype prediction method, including the following steps S101 to S104.

[0030] In step S101: Input the intrinsic feature data and the target CAR structure into the trained phenotype prediction model. The intrinsic feature data includes the multi-dimensional tumor intrinsic features reflected by the CAR-T cells formed by transfecting CAR into T cells and tumor cells, and the CAR has the target CAR structure.

[0031] Figure 2 The model structure of the phenotype prediction model in some embodiments of the present disclosure is shown. As Figure 2 shown, the phenotype prediction model used in the embodiments of the present disclosure includes an encoding layer, a fusion layer, and a prediction layer.

[0032] The CAR-T cell phenotype is related not only to the signal domain structure of CAR-T cells, but also to the tumor-intrinsic characteristics in multiple dimensions. Therefore, it is also necessary to obtain multi-dimensional tumor-intrinsic characteristics before step S101. The obtained multi-dimensional tumor-intrinsic characteristics may include at least two-dimensional tumor-intrinsic characteristics among the following tumor-intrinsic characteristics in each dimension: the initial number of CAR-T cells, the initial number relationship between CD4-T cells and CD8-T cells, the antigen affinity of the CAR, the antigen density of tumor cells, and the tumor burden signal. The above-mentioned tumor-intrinsic characteristics in each dimension include external and internal factors that affect the efficacy of CAR-T cells in addition to the CAR structure.

[0033] In some embodiments, after transfecting a chimeric antigen receptor with a target CAR structure into T cells, the tumor-intrinsic characteristics in each dimension belonging to internal factors can be obtained from the engineered CAR-T cells: the initial number of CAR-T cells, the initial number relationship between CD4-T cells and CD8-T cells, and the antigen affinity of the CAR. The tumor-intrinsic characteristics in each dimension belonging to external factors can be obtained from the collected tumor cells: the antigen density of tumor cells and / or the tumor burden signal, where the tumor burden signal refers to the number of tumor cells, the tumor size, or the total amount of cancer lesions.

[0034] In some embodiments, the initial number relationship between CD4-T cells and CD8-T cells can be characterized by the CAR-T cell ratio of CD4:CD8. CD4 cells (i.e., CD4-T cells) refer to T lymphocytes with CD4+T molecules on their surfaces, and their function is to send counter-information to the immune system when a virus invades, playing a pivotal role in immune response. The function of CD8 cells (i.e., CD8-T cells) is to inhibit and fight the virus after receiving the counter-information. The CAR-T cell ratio of CD4:CD8, as an index of immune regulation, has a normal value of about 1.4 - 2.0. If the ratio is > 2.0 or < 1.4, it indicates a disorder in cellular immune function.

[0035] In other embodiments, the initial number relationship between CD4-T cells and CD8-T cells can be directly characterized by the initial number of CD4-T cells and the initial number of CD8-T cells.

[0036] It should be noted that the initial number of CAR-T cells, the initial number of CD4-T cells, and the initial number of CD8-T cells respectively refer to the number of CAR-T cells, CD4-T cells, and CD8-T cells after transfecting T cells with a CAR having a target CAR structure to transform them into CAR-T cells and then amplifying them. The target CAR structure is the structure of a chimeric antigen receptor designed to recognize tumor cells and activate T cells. What needs to be predicted is the phenotype of the CAR-T cells formed by transfecting the chimeric antigen receptor with the target CAR structure onto T cells.

[0037] In some embodiments, the T cells can be autologous or allogeneic, that is, derived from the tumor patient himself or a healthy donor. That is, T cells need to be isolated from the peripheral blood of the tumor patient or a healthy donor. The tumor cells are also derived from the tumor patient himself. The CAR with the target CAR structure can be transfected into the T cells isolated from the peripheral blood to transform the T cells into CAR-T cells. After detecting the initial number of CAR-T cells, the initial number of CD4-T cells, the initial number of CD8-T cells, the antigen affinity of the CAR-T cells, and the antigen density on the tumor cells, the CAR-T cell ratio of CD4:CD4 is calculated based on the detected initial number of CD4-T cells and the initial number of CD8-T cells.

[0038] In some other embodiments, the characteristic value ranges can be preset in advance for each dimension of the tumor intrinsic characteristics. For each dimension of the tumor intrinsic characteristics, the tumor intrinsic characteristics of this dimension are obtained from within the characteristic value range of this dimension: the initial number of CAR-T cells, the CAR-T cell ratio of CD4:CD8, the antigen affinity of the CAR-T cells, and the antigen density on the tumor cells. Among them, for tumor cells of different tumor types, the set characteristic value ranges for each dimension of the tumor intrinsic characteristics can be different.

[0039] Before performing step S101, it may further include the step of obtaining the target CAR structure. It should be noted that the structure of the chimeric antigen receptor includes an antigen recognition domain, a transmembrane domain (TM, transmembrane domain), and a signal domain belonging to the intracellular domain. The antigen recognition domain is an extracellular domain, which is a scFv (Single-Chain Fragment Variable) for recognizing specific antigens. The signal domain includes a co-stimulatory domain and an activation domain. The co-stimulatory domain consists of multiple co-stimulatory motifs, which are short peptides that can bind to specific downstream signaling proteins. The motif sequence of this peptide signal is the basic building block that controls the output of most signal receptors. When the chimeric antigen receptor is stimulated, the motif queue composed of motifs recruits a group of signaling proteins, shaping a unique cellular response. For example, the 4-1BB co-stimulatory domain in the chimeric antigen receptor contains a binding motif for TRAF (tumor necrosis factor receptor-associated factors) signaling proteins, resulting in an increase in T cell memory and persistence; the CD28 co-stimulatory domain, which contains a binding motif for phosphatidylinositol-3-kinase (PI3K). Therefore, the motif sequence can be considered as "words" used to form phenotypic "sentences" and convey them through the signal domain.

[0040] The target CAR structure is a combination of multiple motif sequences. Among them, the first motif sequence located at the beginning of the target CAR structure and the second motif sequence located at the end of the target CAR structure are fixed. The first motif sequence belongs to the antigen recognition domain. Since CD19 is a cell surface protein expressed at all stages of B cell maturation and is considered an ideal target for treating malignant tumors, the first motif sequence can, but is not limited to, adopt the CD19 motif. The second motif sequence belongs to the ITAM (Immunoreceptor tyrosine-based activation motif) sequence of the signal domain. The first motif sequence can, but is not limited to, adopt CD3z. CD3z is an immunoreceptor tyrosine-based activation motif containing the recruitment kinase ZAP70. The signal domain is composed of multiple signal motifs, showing great heterogeneity and determining the performance after T cell modification. In T cell receptor signal transduction, the ITAM sequence of CD3z plays a key role and can initiate signal transduction to drive the cytotoxic function of T cells.

[0041] In some embodiments, K motif sequences can be obtained from a pre-constructed signal domain library, where K is an integer greater than 1, and the signal domain library includes multiple co-stimulatory motifs and multiple activation motifs; the obtained K motif sequences are set between the first motif sequence and the second motif sequence to obtain the target CAR structure.

[0042] It can be understood that the value of K is determined according to the signal domain structure of the chimeric antigen receptor to be designed. For example, the value of K can be 2, 3, or 4 or even more. Since there are multiple co-stimulatory motifs and multiple activation motifs in the signal domain library, it is beneficial to design more categories of CAR signal domain combinations. According to the signal domain structure to be designed, the K motif sequences obtained from the signal domain library can include only co-stimulatory motifs, or can include both co-stimulatory motifs and activation motifs at the same time. The obtained K motif sequences are randomly set between the first motif sequence and the second motif sequence, or set between the first motif sequence and the second motif sequence according to the requirements of the structural design to obtain the target CAR structure.

[0043] According to different application scenarios, the method of obtaining K motif sequences from the signal domain library can be different: K motif sequences can be randomly obtained from the signal domain library; at least one co-stimulatory motif can be randomly obtained from the signal domain library and at least one activation motif can be randomly obtained to obtain K motif sequences; K motif sequences can be specified; one or more of the K motif sequences can be specified, and the remaining motif sequences are randomly obtained from the signal domain library.

[0044] Both co-stimulatory motifs and activation motifs are amino acid sequences composed of multiple amino acids. Some of the co-stimulatory motifs and some of the activation motifs in the signal domain library can be referred to as shown in Table 1 below:

[0045] Table 1. Some co-stimulatory motifs and some activation motifs in the signal domain library

[0046]

[0047]

[0048] Taking K = 3 as an example below, the target CAR structure is composed of five basic sequence combinations: M0 + M1 + M2 + M3 + M4. The first position M0 uses the first motif sequence CD19, and the last position M4 uses the second motif sequence CD3z. It is also necessary to obtain 3 motif sequences from the pre-constructed signal domain library, and the obtained 3 motif sequences are randomly or sequentially set at positions M1, M2, and M3. For example: obtain the motif sequences m 2 , m 3 , m 4, correspondingly set at positions M1, M2, and M3 to obtain the target CAR structure: CD19 + m 2 +m 3 +m 4 +CD3z; for another example: obtain the motif sequence m from the signal domain library 3 、m 4 、m 5 , randomly set at positions M1, M2, and M3 to obtain the target CAR structure: CD19 + m 5 +m 3 +m 4 +CD3z; for another example: randomly select m from the signal domain library 1 、m 6 、m 9 , randomly set at positions M1, M2, and M3 to obtain the target CAR structure: CD19 + m 6 +m 9 +m 1 +CD3z。

[0049] In step S102: obtain the hidden layer features of the intrinsic feature data from the encoding layer, and obtain the semantic feature sequence of the target CAR structure from the encoding layer.

[0050] It should be noted that the multi-dimensional tumor intrinsic features are real values belonging to different dimensions, while the target CAR structure belongs to sequence data. The sub-models for extracting features from the intrinsic feature data and the target CAR structure should be different. Therefore, as Figure 2 shown, in some embodiments, the encoding layer of the phenotype prediction model includes a first encoding sub-model for extracting features from the intrinsic feature data and a second encoding sub-model for extracting features from the target CAR structure.

[0051] In some embodiments, each motif sequence in the target CAR structure can be one-hot encoded by the second encoding sub-model of the encoding layer to obtain a one-hot encoded vector (hereinafter simply referred to as the onehot vector) for each motif sequence in the target CAR structure. The one-hot encoded vectors corresponding to each motif sequence in the target CAR structure are directly set in order to obtain the semantic feature sequence of the target CAR structure. The semantic feature sequence of the target CAR structure obtained thereby is represented by onehot vectors. Since each motif sequence is an amino acid sequence composed of multiple amino acids, a vocabulary {A, R, N,..., V} is established based on 20 amino acids. Each amino acid is represented by a 20-dimensional onehot vector, that is, the position index of the amino acid in the vocabulary is 1, and the rest of the position indexes are 0. For example, the 20-dimensional onehot vector of alanine A is [1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], and the onehot vector representation of arginine R is [0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0].

[0052] For each motif sequence in the target CAR structure, after converting each amino acid in the motif sequence into a 20-dimensional onehot vector according to the vocabulary, L 20-dimensional onehot vectors are concatenated in the amino acid order of the motif sequence to obtain an L*20-dimensional onehot vector of the motif sequence, where L is the sequence length of the motif sequence, that is, the number of amino acids in the motif sequence. For example, the m1 motif sequence DYHNPGYLVVLPDSTP in Table 1 above can be one-hot encoded to obtain a 320 (16*20)-dimensional onehot vector.

[0053] In some other embodiments, the second encoding sub-model may further include a pre-trained language model. After one-hot encoding each motif sequence in the target CAR structure to obtain a one-hot encoded vector for each motif sequence in the target CAR structure, the pre-trained language model is used to encode the one-hot encoded vector of each motif sequence in the target CAR structure to obtain an embedding vector for each motif sequence in the target CAR structure. The semantic feature sequence includes the embedding vectors of each motif sequence in the target CAR structure. It can be understood that the pre-trained language model is pre-trained on a protein sequence corpus.

[0054] It should be noted that the embedding vector of each motif sequence in the target CAR structure is a low-dimensional dense embedding vector of a specified size. The embedding vector is a low-dimensional and high-density feature representation of each motif sequence. Compared with directly using one-hot vectors to represent the semantic feature sequence of the target CAR structure, by introducing low-dimensional and high-density embedding vector representation, the problem of directly using a simple and sparse one-hot encoding method to represent the semantic feature sequence of the target CAR structure is avoided, and the problem of sparse motif semantic features is solved. For example, through the multi-layer network mapping of a pre-trained language model, an embedding vector h of 512 dimensions can be output sequence , which is a low-dimensional dense and high-level feature with the semantic information of the amino acid sequence.

[0055] In some embodiments, the pre-trained language model can be a ProtBert model. The ProtBert model is a Bert (Bidirectional Encoder Representations from Transformers) model pre-trained on a protein sequence corpus. Therefore, the ProtBert model is a multi-layer Transformer encoder. Each input vector of the ProtBert model mainly includes three parts, namely, token embedding, segment embedding, and position embedding.

[0056] In some embodiments, the first encoding sub-model of the encoding layer includes a first forward neural network. The first forward neural network has an input layer and multiple hidden layers. Since the multi-dimensional tumor intrinsic features are multiple real values belonging to different dimensions, the first forward neural network can use a multi-layer MLP (Multi-Layer Perceptron) without the need to use a more complex neural network, reducing the complexity of the entire model.

[0057] Since the tumor intrinsic features in different dimensions of the intrinsic feature data are real values with different dimensions, for example: the initial number of CAR-T cells is 1561, the ratio of CD4:CD8 CAR-T cells is 90:10, the antigen affinity of CAR-T cells is 1e-6, and the antigen density on tumor cells is 100. It can be seen that the feature values of these dimensions are all real values but with different dimensions. Therefore, in step S102, obtaining the hidden layer features of the intrinsic feature data by the encoding layer may include: respectively performing normalization processing on the multi-dimensional tumor intrinsic features to correspondingly obtain multiple standard feature values. Among them, performing normalization processing on the multi-dimensional tumor intrinsic features may be to unify the multi-dimensional tumor intrinsic features into standard feature values between [0, 1]; inputting the obtained multiple standard feature values into the first forward neural network, and the first forward neural network performs non-linear combination representation on the multiple standard feature values to obtain the hidden layer features of the multi-dimensional tumor intrinsic features. The activation function of the first forward neural network adopts a non-linear activation function (such as: ReLU function) to realize the non-linear combination representation of the first forward neural network on the multiple standard feature values.

[0058] In some embodiments, the first forward neural network may adopt a multi-layer perceptron (MLP), that is, a forward neural network with an input layer and at least one hidden layer, and output the hidden layer features after being mapped through two hidden layers to the fusion layer. The size of each hidden layer can be specified. Taking the first forward neural network having two hidden layers as an example, the size of the first hidden layer can be 16, and the size of the second hidden layer can be 32. Then, the multiple standard feature values input into the first forward neural network can obtain a 32-dimensional hidden layer vector h after being mapped through two hidden layers tumor That is, the hidden layer features are obtained.

[0059] In step S103: The fusion layer fuses the semantic feature sequence and the hidden layer features to obtain the fusion features.

[0060] In some embodiments, the fusion layer may fuse the semantic feature sequence and the hidden layer features based on the attention mechanism to obtain the fusion features.

[0061] In some embodiments, the fusion layer may employ a single Transformer model. The Transformer model is a model based on the attention mechanism (AM). By using an encoder block of the Transformer model, it is possible to encode the semantic feature sequence and the hidden layer features based on the attention mechanism to encode them into a fixed-length fusion feature vector. This fusion feature vector contains all the input information while minimizing redundant information as much as possible, that is, the fusion feature is obtained. This process is shown in the following formula (1):

[0062] h fusion = Tansformer(h sequence , h tumor ) (1)

[0063] Where h sequence and h tumor are the inputs of the Transformer model, h sequence is the semantic feature sequence of the target CAR structure, h tumor is the hidden layer feature; h fusion is the fusion feature vector output by the Transformer model.

[0064] In some embodiments, the process of fusing the semantic feature sequence and the hidden layer features through the Transformer model of the fusion layer may include:

[0065] Performing attention weighting on the semantic feature sequence and the hidden layer features to obtain a weighted sum result, and performing layer normalization on the sum result of the weighted sum result and the semantic feature sequence to obtain a first output result. This process can be referred to the following formulas (2), (3), and (4). Assign h sequence to the Q vector, assign h tumor to the K vector and the V vector, and then perform attention addition on the Q vector, the K vector, and the V vector; add the attention weighting result to the Q vector, and perform layer normalization on the obtained sum result to obtain the first output result H:

[0066] Q = h sequence (2)

[0067] K = V = h tumor (3)

[0068] H = LayerNorm(Q + MultiheadAttention(Q, K, V)) (4)

[0069] Among them, Q, K, and V respectively represent the query vector, key vector, and value vector. MultiheadAttention represents the attention weighting of the Q vector, K vector, and V vector, and LayerNorm is layer normalization.

[0070] Input the first output result into the second feed-forward neural network to obtain the second output result; perform layer normalization on the sum result of the first output result and the second output result to obtain the fused feature, which can be specifically referred to as shown in formula (5) below:

[0071] Transformer(Q, K) = LayerNorm(H + FeedForward(H)) (5)

[0072] Among them, FeedForward(H) represents the processing of the first output result by the second feed-forward neural network, and LayerNorm represents layer normalization.

[0073] In some embodiments, the second feed-forward neural network is a part of the Transformer model in the fusion layer. The second feed-forward neural network has only one layer, and the activation function of the first feed-forward neural network uses a non-linear activation function (such as: ReLU function).

[0074] The Transformer model in the fusion layer can adopt the multi-head attention mechanism or the self-attention mechanism. In some embodiments, using the multi-head attention mechanism for the attention weighting of the Q vector, K vector, and V vector can effectively improve the expression ability of the Transformer model in the fusion layer. At the same time, it can also enable the Transformer model to learn more diverse and complex features in the semantic feature sequence and hidden layer features. Under the multi-head attention mechanism, the input data will be divided into multiple heads, and each head performs independent calculations to obtain different outputs. The outputs of these heads are finally concatenated together. This process can be referred to as shown in formulas (6) and (7) below:

[0075] MultiHeadAttention(Q, K, V) = Concat(head 1 ,..., head h )W O (6)

[0076] h represents the number of heads, which is a hyperparameter of the model and can be defined by oneself. For example, it can be defined as 4. head i represents the output of the i-th head, and W O is the output transformation matrix. The output head i of each head can be expressed as:

[0077]

[0078] Among them, They are respectively the transformation matrices of the query, key, and value of the i-th head.

[0079] In step S104: Input the fused feature into the prediction layer to predict the CAR-T cell phenotype of the target CAR structure.

[0080] In some embodiments, the prediction layer includes a third forward neural network. Input the fused feature into the third forward neural network, and perform regression prediction through the third forward neural network to predict the phenotype of the CAR-T cell with the target CAR structure. Among them, the predicted CAR-T cell phenotype includes predicting the cytotoxicity and / or stemness of the CAR-T cell.

[0081] In some embodiments, the third forward neural network can adopt a perceptron model with at least two layers. All layers except the last layer use the linear activation function linear, and the last layer does not use an activation function to achieve regression prediction.

[0082] Based on the same inventive concept, the embodiments of the present disclosure provide a training method for a phenotype prediction model, Figure 3 which shows the flowchart of the training method for the phenotype prediction model in some embodiments of the present disclosure, as Figure 3 shown, including the following steps S201 to S202.

[0083] In step S201: Obtain a training sample set. The training sample set includes multiple training sample pairs. Each training sample pair is composed of a CAR structure sample and an intrinsic feature sample. The intrinsic feature sample of each training sample pair includes the multi-dimensional tumor intrinsic features reflected by the CAR-T cells formed by transfecting the chimeric antigen receptor into T cells and tumor cells. The structure of the chimeric antigen receptor is the same as the CAR structure sample corresponding to the intrinsic feature sample.

[0084] In some embodiments, the training sample pairs in the training sample set and their true values of cytotoxicity and proliferation ability can be exemplified as shown in Table 2 below:

[0085] Table 2. True values of training sample pairs and their cytotoxicity and proliferation ability

[0086]

[0087] In some embodiments, the training sample set at least includes CAR structure samples of a target type, wherein the CAR structure samples of the target type are samples whose signal domain includes at least one activation motif and at least one co-stimulatory motif. In addition to the training sample pairs with CAR structure samples of the target type, the training sample set may further include other types of training sample pairs: the signal domain of the CAR structure samples of the other types of training sample pairs only has co-stimulatory motifs. By increasing the training sample pairs with CAR structure samples of the target type, the sample diversity is increased, which is beneficial to improving the generality of the trained phenotype prediction model, so that the trained phenotype prediction model can be applied to predict the phenotypes of CAR-T cells of more signal domain categories.

[0088] In some embodiments, in step S202: the neural network model constructed based on the training sample set is trained to obtain a phenotype prediction model for predicting the phenotype of CAR-T cells, wherein an encoding layer, a fusion layer, and a prediction layer are constructed in the neural network model.

[0089] In some embodiments, the step of obtaining the training sample set may include: obtaining a plurality of CAR structure samples, and for each CAR structure sample in the plurality of CAR structure samples, obtaining a corresponding intrinsic feature sample based on the CAR structure sample to form a training sample pair.

[0090] The obtaining of the CAR structure sample in each training sample pair may include: obtaining K motif sequences from a pre-constructed signal domain library, and setting the obtained K motif sequences between a first motif sequence and a second motif sequence to obtain the CAR structure sample of the training sample pair. The signal domain library includes a variety of activation motifs and a variety of co-stimulatory motifs, and K is an integer greater than 1.

[0091] Each of the intrinsic feature samples in the obtained training sample pairs may include at least two of the following dimensions of tumor intrinsic features: the initial number of CAR-T cells, the relationship between the number of CD4-T cells and the initial number of CD8-T cells, the antigen affinity of CAR-T cells, the antigen density on tumor cells, and the tumor burden signal. Each dimension of tumor intrinsic features that are internal factors in the intrinsic feature samples can be obtained from the CAR-T cells formed by transfecting a chimeric antigen receptor with a signal domain structure in the CAR structure sample into T cells: the initial number of CAR-T cells, the relationship between the number of CD4-T cells and the initial number of CD8-T cells, the antigen affinity of CAR-T cells, and each dimension of tumor intrinsic features that are external factors can be obtained from tumor cells of a certain tumor type: the antigen density on tumor cells and / or the tumor burden signal. It should be noted that the types of tumor cells used for different intrinsic feature samples can be different to achieve sample diversity, thereby improving the generality of the obtained phenotype prediction model.

[0092] In some embodiments, K motif sequences are randomly obtained from the signal domain library each time to enable the CAR structure samples with various category signal domain combinations to be included in the formed training sample set.

[0093] By using the above training sample set to train the entire neural network model, the training of the encoding layer, the fusion layer, and the prediction layer can be achieved. Among them, if the encoding layer includes a pre-trained language model, then through the training of the entire neural network model, the fine-tuning of the pre-trained language model in the encoding layer is achieved. The loss function used for training the entire neural network model can adopt MSE (mean-square error), as shown in the following formula (8):

[0094] loss mse (Y true ,Y pred )=1 / N∑(Y true -Y pred ) 2 (8)

[0095] Among them, Y true represents the true value, Y pred represents the predicted value output by the neural network model, and N represents the number of samples.

[0096] Such as Figure 1As shown, if the encoding layer includes a first encoding sub-model and a second encoding sub-model, the pre-trained language model ProtBert in the second encoding sub-model is composed of 30 Transformer layers and is pre-trained with a masked language model (MLM) on a protein sequence corpus, that is, a part of the words (amino acids) in the sequence are masked, and then the model is trained to predict the masked words (amino acids) based on the remaining words (amino acids). Through this self-supervised training, ProtBert can learn the language information of protein sequences and generate semantic feature representations of protein sequences. Thus, an open-source and pre-trained ProtBert model is obtained.

[0097] As Figure 1 shown, since the sequence lengths of different motif sequences are inconsistent, the sequence lengths of the formed CAR structure samples will be inconsistent. The CAR structure sample with the maximum sequence length in the training sample set can also be determined, and the sequence lengths of the remaining CAR structure samples can be padded to the maximum sequence length by filling with 0s, so that the sequence lengths of all CAR structure samples in the training sample set are the same. Taking the CAR structure sample composed of 5 motif sequences as an example, it is padded to the maximum sequence length max(M0 + M1 + M2 + M3 + M4)).

[0098] It should be noted that more implementation details of the training method of the phenotype prediction model in the embodiments of the present disclosure can be referred to the embodiments of the CAR-T cell phenotype prediction method described above. For the sake of brevity of the specification, it will not be elaborated here.

[0099] Based on the same inventive concept, the embodiments of the present disclosure also provide an electronic device. Figure 4 shows a schematic structural diagram of the electronic device in some embodiments of the present disclosure. As Figure 4 shown, the electronic device includes a processor 404; a memory 402 for storing executable instructions of the processor 404; wherein, the processor 404 is configured to execute to implement the CAR-T cell phenotype prediction method described in any of the above embodiments, or is configured to execute to implement the training method of the phenotype prediction model described in any of the above embodiments.

[0100] Among them, in Figure 4Among them, a bus architecture (represented by bus 400), bus 400 may include any number of interconnected buses and bridges. Bus 400 links together various circuits including one or more processors represented by processor 402 and a memory represented by memory 404. Bus 400 may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art, and thus will not be further described herein. Bus interface 405 provides an interface between bus 400 and receiver 401 and transmitter 403. Receiver 401 and transmitter 403 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 402 is responsible for managing bus 400 and general processing, while memory 404 may be used to store data used by processor 402 when performing operations.

[0101] Based on the same inventive concept, embodiments of the present disclosure provide a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the CAR-T cell phenotype prediction method described in any of the above embodiments, or can implement the training method of the phenotype prediction model described in any of the above embodiments.

[0102] Those skilled in the art should understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0103] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0104] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the function specified in one or more processes and / or blocks Figure 1 in the flowchart Figure 1 representation, in one or more blocks or boxes.

[0105] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more processes and / or blocks Figure 1 in the flowchart Figure 1 representation, in one or more blocks or boxes.

[0106] Although the preferred embodiments of the present disclosure have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present disclosure.

[0107] Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure is also intended to include these modifications and variations.

Claims

1. A method for predicting the phenotype of CAR-T cells, characterized in that, it includes: Inputting the intrinsic feature data and the target CAR structure into a pre-trained phenotype prediction model, where the intrinsic feature data includes multi-dimensional tumor intrinsic features reflected by CAR-T cells formed by transfecting CAR into T cells and tumor cells, the CAR has the target CAR structure, and the phenotype prediction model includes an encoding layer, a fusion layer, and a prediction layer; Obtaining the hidden layer features of the intrinsic feature data by the encoding layer, and obtaining the semantic feature sequence of the target CAR structure by the encoding layer; Fusing the semantic feature sequence and the hidden layer features by the fusion layer to obtain fusion features; Inputting the fusion features into the prediction layer to predict the phenotype of CAR-T cells with the target CAR structure.

2. The method according to claim 1, characterized in that, the encoding layer includes a pre-trained language model, and obtaining the semantic feature sequence of the target CAR structure by the encoding layer includes: Performing one-hot encoding on each motif sequence in the target CAR structure respectively to obtain one-hot encoding vectors of each motif sequence in the target CAR structure, and the target CAR structure includes multiple motif sequences; Encoding the one-hot encoding vectors of each motif sequence in the target CAR structure by the pre-trained language model to obtain embedding vectors of each motif sequence in the target CAR structure, the semantic feature sequence includes the embedding vectors of each motif sequence in the target CAR structure, and the pre-trained language model is pre-trained on a protein sequence corpus.

3. The method according to claim 2, characterized in that, the encoding layer further includes a first forward neural network with multiple hidden layers, and obtaining the hidden layer features of the intrinsic feature data by the encoding layer includes: Performing standardization processing on the multi-dimensional tumor intrinsic features respectively to obtain multiple standard feature values; Performing non-linear combination representation on the multiple standard feature values by the first forward neural network to obtain the hidden layer features.

4. The method according to claim 3, characterized in that, the intrinsic feature data includes at least any two-dimensional tumor intrinsic features: the initial quantity relationship between CD4-T cells and CD8-T cells, the initial quantity of CAR-T cells, the antigen affinity of CAR-T cells, the antigen density of tumor cells, and the tumor burden signal.

5. The method according to claim 1, characterized in that, before obtaining the semantic feature sequence of the target CAR structure by the encoding layer, it further includes: Obtaining K motif sequences from a signal domain library, where K is an integer greater than 1, and the signal domain library includes various activation motifs and various co-stimulatory motifs; Setting the K motif sequences between a first motif sequence and a second motif sequence to obtain the target CAR structure.

6. The method according to claim 1, characterized in that, fusing the semantic feature sequence and the hidden layer features by the fusion layer to obtain fusion features, includes: Perform attention weighting on the semantic feature sequence and the hidden layer features to obtain a weighted sum result; Perform layer normalization on the weighted sum result and the sum result of the semantic feature sequence to obtain a first output result; Input the first output result into a second forward neural network to obtain a second output result; Perform layer normalization on the sum result of the first output result and the second output result to obtain the fusion features.

7. The method according to claim 1, wherein, the prediction layer includes a third forward neural network, and inputting the fusion features into the prediction layer to predict the CAR-T cell phenotype of the target CAR structure includes: Inputting the fusion features into the third forward neural network, and performing regression prediction through the third forward neural network to obtain the CAR-T cell phenotype of the target CAR structure.

8. The method according to any one of claims 1-7, wherein, it further includes the step of pre-training the phenotype prediction model: Obtain a training sample set, the training sample set includes a plurality of training sample pairs, each training sample pair is composed of a CAR structure sample and an intrinsic feature sample, the intrinsic feature sample of each training sample pair includes the multi-dimensional tumor intrinsic features reflected by the CAR-T cells formed by transfecting chimeric antigen receptors into T cells and tumor cells, and the structure of the chimeric antigen receptor is the same as the CAR structure sample corresponding to the intrinsic feature sample; Train the constructed neural network model based on the training sample set to obtain the trained phenotype prediction model, wherein, an encoding layer, a fusion layer and a prediction layer are constructed in the neural network model.

9. The method according to claim 8, wherein, the training sample set at least includes CAR structure samples of a target type, and the signal domain of the CAR structure samples of the target type includes at least one activation motif and at least one co-stimulatory motif.

10. The method according to claim 9, wherein, the obtaining of the training sample set includes: Obtain a plurality of CAR structure samples, wherein, the obtaining of each CAR structure sample includes: obtaining K motif sequences from the signal domain library, and setting the K motif sequences between a first motif sequence and a second motif sequence to obtain the CAR structure sample, K is an integer greater than 1; For each CAR structure sample among the plurality of CAR structure samples, obtain the corresponding CAR structure sample of the CAR structure sample.

11. A training method for a phenotype prediction model, wherein, it includes: Obtain a training sample set, the training sample set includes a plurality of training sample pairs, each training sample pair is composed of a CAR structure sample and an intrinsic feature sample, the intrinsic feature sample of each training sample pair includes the multi-dimensional tumor intrinsic features reflected by the CAR-T cells formed by transfecting chimeric antigen receptors into T cells and tumor cells, and the structure of the chimeric antigen receptor is the same as the CAR structure sample corresponding to the intrinsic feature sample; Train the constructed neural network model based on the training sample set to obtain a phenotype prediction model for predicting the phenotype of CAR-T cells. Among them, an encoding layer, a fusion layer, and a prediction layer are constructed in the neural network model.

12. The method according to claim 11, characterized in that the training sample set at least includes CAR structure samples of a target type, and the signal domain of the CAR structure samples of the target type includes at least one activation motif and at least one co-stimulatory motif.

13. The method according to claim 12, characterized in that the obtaining of the training sample set includes: obtaining a plurality of CAR structure samples, wherein the obtaining of each CAR structure sample includes: obtaining K motif sequences from the signal domain library, and setting the K motif sequences between a first motif sequence and a second motif sequence to obtain the CAR structure sample, where K is an integer greater than 1; for each CAR structure sample among the plurality of CAR structure samples, obtaining the corresponding CAR structure sample of the CAR structure sample.

14. An electronic device, characterized in that it includes: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute to implement the CAR-T cell phenotype prediction method according to any one of claims 1 to 10, or is configured to execute to implement the training method of the phenotype prediction model according to any one of claims 11 to 13.

15. A non-transitory computer-readable storage medium, characterized in that when the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute to implement the CAR-T cell phenotype prediction method according to any one of claims 1 to 10, or can implement the training method of the phenotype prediction model according to any one of claims 11 to 13.

Citation Information

Patent Citations

  • Methods for the design and optimisation of chimeric antigen receptors (CARS)

    WO2023089309A1