INKT cell Th1 type agonist dosization prediction model based on pre-training model and priori knowledge
By using a pre-trained model and prior knowledge-based iNKT cell Th1 agonist dosing prediction model, the problem of quantitative correlation analysis between the molecular structural parameters of glycolipid antigens and the degree of Th1 polarization was solved, enabling precise regulation and efficient evaluation of iNKT cell therapy and providing a reliable drug design tool.
Patent Information
- Application Number
- CN202510577103.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies lack systematic quantitative correlation analysis between the molecular structural parameters of glycolipid antigens and the degree of Th1 polarization, leading to the development of novel iNKT cell activators relying on empirical screening and neglecting synthetic and economic evaluations, resulting in difficulties in synthesis or high costs in clinical applications.
A pre-trained model and prior knowledge-based iNKT cell Th1 agonist dose prediction model was adopted. Through molecular docking simulation and small sample transfer learning, the interaction relationship between glycolipid antigen molecules and iNKT cell target proteins was constructed. By combining multilayer perceptron and graph neural network for feature fusion, the quantitative assessment of Th1 immune response bias was achieved.
It provides quantitative evaluation metrics, improves the precision regulation efficiency of iNKT cell therapy, solves the training problem of small sample datasets, has the ability to explain molecular mechanisms of action, and has good scalability and adaptability.
Smart Images

Figure CN120954510A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a dosing prediction model for Th1 agonists in iNKT cells based on a pre-trained model and prior knowledge, which belongs to the technical field of data processing. Background Technology
[0002] In recent years, iNKT cells have become a research hotspot in tumor immunotherapy due to their unique CD1d-dependent antigen recognition mechanism and pleiotropic immunomodulatory function. However, their core agonist, the glycolipid antigen molecule KRN7000, faces significant challenges in clinical application. Its non-selective induction of a mixed Th1 / Th2 response may lead to an imbalance in the immune response, weakening the anti-tumor efficacy.
[0003] Although numerous studies have explored how different glycolipid antigen derivatives can partially enhance Th1 susceptibility through chemical modification, current technologies lack systematic and quantitative correlation analysis between molecular structural parameters and molecular-protein interactions and the degree of Th1 polarization. This results in the development of novel activators still relying on empirical screening. Furthermore, existing molecular design methods often neglect the assessment of the synthetic feasibility and cost-effectiveness of new molecules, which may lead to problems such as inability to synthesize or excessive costs in practical applications.
[0004] In view of this, this invention proposes a pre-trained model and prior knowledge-based iNKT cell Th1 agonist dosing prediction model. Through molecular docking simulation technology and small sample transfer learning model, the quantitative structure-activity relationship of the Th1 agonist tendency of different glycolipid antigen molecules to activate iNKT cells was explored. This has important scientific significance and clinical application value for achieving precise regulation of iNKT cell therapy. Summary of the Invention
[0005] To address the shortcomings of the above technologies, the purpose of this invention is to provide a model for quantitatively assessing the degree of Th1-type immune response bias in iNKT cells activated by glycolipid antigen molecules.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] A quantitative prediction model for the effects of iNKT cell-promoting Th1 agonists based on a pre-trained model and prior knowledge of molecular docking includes the following steps:
[0008] (1) Dataset Construction: Molecular structure and bioactivity data of glycolipid antigen molecule KRN7000 and its derivatives were collected and organized. The correlation coefficient between the collected molecular structures and Th1 immune response bias was calculated to ensure a strong correlation between the collected molecular structures and Th1 immune response bias data. The molecular descriptors and physicochemical properties of the collected glycolipid antigen molecule KRN7000 and its derivatives were calculated using RDKit as a supplement to the dataset.
[0009] (2) Screening of key amino acid residues in the target protein group of activated iNKT cell immune response: The amino acids around the docking pocket of the target protein group of activated iNKT cell by glycolipid antigen molecules were calculated as candidate amino acids, and the relevant amino acids that have been reported in the literature were selected as key amino acids.
[0010] (3) Molecular docking based on AutoDock4GPU and interaction analysis based on PLIP library: For the collected and sorted molecular structures of glycolipid antigen molecules KRN7000 and its derivatives, the AutoDock4GPU software was used to perform batch and repeated molecular docking with the target protein group of activated iNKT cells. The interaction relationship between the molecular docking results and key amino acid residues was analyzed using the PLIP library to construct the original data features.
[0011] (4) Construction of a small sample transfer learning model for Th1 immune response based on pre-trained model and prior knowledge: The physicochemical properties of glycolipid antigen molecules and the data obtained from step (3) are used as the original data features, and the bioactivity data of glycolipid antigen molecules are used as data labels. They are input into the pre-trained model and the prior knowledge model respectively, and the regression model is trained after feature fusion.
[0012] Further, in step (1), the correlation coefficient is calculated by using the Geminimol model to calculate the correlation coefficients between molecular structure and Th1-type immune response bias, namely the Pearson coefficient and the Spearman coefficient.
[0013] The formula for calculating the Pearson coefficient is as follows:
[0014]
[0015] in, Indicates the first A pair of observations for each molecule. express and The average value, This refers to the sample size. Pearson's value reflects the linear correlation between the sample's molecular structure and bioactivity data.
[0016] The formula for calculating the Spearman coefficient is:
[0017]
[0018] in, Indicates the first Bioactivity data of individual molecules The ranking difference between them This represents the sample size. Spearman's value reflects the non-linear correlation between the sample size and the bioactivity data.
[0019] The dataset was constructed by collecting and organizing data on the molecular structure, physicochemical properties, and Th1 / Th2 cytokine content induced by glycolipid antigen molecule KRN7000 and its derivatives. The Th1 / Th2 cytokine content data included the relative ratio of the peak concentration of IFN-γ, a Th1 immune response cytokine, between glycolipid antigen molecules and KRN7000, and the relative ratio of the secretion levels of Th1 / Th2 cytokines induced by glycolipid antigen molecules to those induced by KRN7000.
[0020] Further, in step (2), the candidate amino acid residues are obtained by BioPython calculations. They are the binding pockets of the target protein group (CD1d protein and iTCR protein complex) and glycolipid molecular antigens during iNKT cell immune activation. Amino acid residues within 5 Å of the docking pocket position are selected as candidate amino acid residues. Among the amino acid residues, amino acid residues that have been reported in the literature are selected as key amino acid residues.
[0021] Furthermore, in step (3), the molecular docking work based on the AutoDock4GPU software requires converting the SMILES string of the molecular structure into a PDBQT format file with three-dimensional structural information and atomic charge information through OpenBabel. Under the same conditions, molecular docking is performed on all collected molecules in batches and repeatedly through the grid parameter file. In step (3), the interaction relationship of key amino acid residues is analyzed based on the PLIP library. The analysis scope includes hydrogen bond donors, hydrogen bond acceptors, hydrogen bond strength, halogen bond donors, halogen bond acceptors, halogen bond strength, hydrophobic interactions, π-ion interactions, π stacking interactions, number of hydrogen bonds, number of halogen bonds, number of hydrophobic interactions, number of π-ion interactions, and number of π stacking interactions.
[0022] Further, in step (4), the interaction between the pre-trained model Geminimol and the target protein group of iNKT cells activated by glycolipid antigen molecules is used as prior knowledge to train the regression model. The regression model is mainly composed of features extracted by the pre-trained model Geminimol through unsupervised learning and features extracted by the prior knowledge through unsupervised learning, which are then fused together. The prior knowledge consists of two parts: the three-dimensional coordinate features of glycolipid antigen molecules after repeated molecular docking through voxels and graph neural networks in unsupervised learning, and the molecular descriptors and physicochemical properties of glycolipid antigen molecules through multilayer perceptron in unsupervised learning. The features obtained through unsupervised learning are fused through an attention mechanism, then reduced in dimensionality through principal component analysis and input into a convolutional neural network. Finally, they are mapped to the bioactivity data labels of glycolipid antigen molecules through a fully connected layer. The accuracy of the obtained model is verified by 5-fold cross-validation.
[0023] Furthermore, the aforementioned voxel and graph neural network unsupervised learning process extracts the 3D coordinate features of glycolipid antigen molecules after multiple repeated molecular dockings. During this process, the docking results of the glycolipid antigen molecules are used to extract features from the molecular conformation through an adaptive voxel generator. This generator employs a dynamic neighbor search strategy, using the KD-tree algorithm to identify the nearest neighbors (up to 6 neighbors) of each atom within a 3 Å cutoff distance, and weights the contributions of neighboring atoms using a Gaussian weighting function. For each atom, basic features including charge information, atom coordinates, and atom type are extracted, and combined with the statistical features (mean and standard deviation) of its neighboring atoms, a local environment descriptor is constructed.
[0024] Furthermore, the feature fusion employs Gaussian radial basis functions for feature smoothing mapping and combines an attention mechanism for feature extraction.
[0025] Compared with existing technologies, the beneficial effects of this invention are as follows: By introducing prior knowledge of iNKT cell Th1 agonists obtained through molecular docking into the pre-trained model, and through transfer learning, its application in the development of iNKT cell tumor immunotherapy drugs can be specialized, providing quantitative evaluation indicators and achieving more efficient evaluation of the effects of Th1-type immune response-promoting glycolipid antigens. Specifically:
[0026] First, by introducing the pre-trained model GeminiMol to implement a transfer learning strategy, the challenge of training on small datasets is effectively addressed. Second, the attention-based feature fusion method adaptively learns the importance of different feature modalities, achieving dynamic weight allocation for features. Third, by integrating molecular docking information and protein-protein interaction features, the model gains the ability to explain molecular mechanisms of action. Finally, the modular design of this framework provides excellent scalability and adaptability, allowing for flexible adjustments or additions of new feature extraction modules based on specific needs. This systematic analytical approach not only provides a reliable computational tool for predicting the activity of glycolipid antigen molecules but also offers a valuable solution for other similar drug design tasks. Attached Figure Description
[0027] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0028] Figure 1 The distribution of the original data and the distribution of the standardized data are shown in the graph.
[0029] Figure 2 To extract the variation graph of the loss function of the fully connected neural network for molecular docking;
[0030] Figure 3 The training loss curve and validation index curve of the model when the dependent variable label (Y) is the relative ratio of Th1 cytokine IFN-γ;
[0031] Figure 4 The training loss curve and validation index curve of the model when the dependent variable label (Y) is the relative ratio of the Th1 / Th2 immune response ratio. Detailed Implementation
[0032] This invention provides a dosing prediction model for Th1 agonist dosing in iNKT cells based on a pre-trained model and prior knowledge. To make the objectives and technical solutions of this invention clearer and more explicit, the invention is further explained in detail below. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0033] Example 1
[0034] This example demonstrates an iNKT cell Th1 agonist dosing prediction model based on a pre-trained model and prior knowledge. It utilizes molecular docking to obtain the interaction relationships between glycolipid antigen molecules and target proteins associated with iNKT cell activation. The molecular descriptors and physicochemical properties of glycolipid antigen molecules are calculated using RDKit. Three unsupervised models (Geminimol molecular pre-trained model, graph neural network model, and multilayer perceptron) are combined to extract molecular features and molecular-protein interaction features. After feature fusion using an attention mechanism, an iNKT cell Th1 agonist dosing prediction model is established, and the accuracy of the model is verified using 5-fold cross-validation. The specific steps are as follows:
[0035] (1) Dataset construction:
[0036] First, we collected and organized the structural and corresponding bioactivity data of 155 different KRN7000 and its glycolipid derivatives from the literature. The bioactivity data refers to the raw data after standardization and outlier removal of the induced Th1 / Th2 cytokine levels, such as... Figure 1 As shown; the Th1 / Th2 cytokine content data include: the relative ratio of the peak concentration ratio of IFN-γ, a Th1-type immune response cytokine, to glycolipid antigen molecules (relative to KRN7000), and the relative ratio of the secretion content of Th1 / Th2 cytokines induced by glycolipid antigen molecules (relative to KRN7000). Chemical descriptors for each molecule were calculated from the SMILES string using RDKit, including: Morgan fingerprint (Morgen_fingerprint), molecular weight (MW), lipid-water partition coefficient (LogP), topological polar surface area (TPSA), drug-likeness (QED), number of hydrogen bond donors (HBD_count), number of hydrogen bond acceptors (HBA_count), and number of rotatable bonds (RotBonds). Using the aforementioned bioactivity data as the dependent variable (Y), the SMILES string of the small molecules as the independent variable (X1), and the descriptor calculated from RDKit as the independent variable (X2), a small-sample regression model is constructed. The Pearson and Spearman coefficients between the independent variable X1 and the dependent variable Y are calculated using a Geminimol pre-trained model. Outlier samples are removed to ensure that either coefficient reaches above 0.8, thereby guaranteeing a strong linear or non-linear correlation between the sample molecules and the bioactivity data. The formula for calculating the Pearson coefficient is as follows:
[0037]
[0038] in, Indicates the first A pair of observations for each molecule. express and The average value, This refers to the sample size. Pearson's value reflects the linear correlation between the sample's molecular structure and bioactivity data.
[0039] The formula for calculating the Spearman coefficient is:
[0040]
[0041] in, Indicates the first Bioactivity data of individual molecules The ranking difference between them This represents the sample size. Spearman's value reflects the non-linear correlation between the sample size and the bioactivity data.
[0042] (2) Screening of key amino acid residues in the target proteome of iNKT cell immune responses:
[0043] BioPython was used to calculate the amino acid residues within a 5 Å radius surrounding the binding site of glycolipid antigen molecules as candidate amino acid residues. These binding sites are docking pockets formed by the target protein group (CD1d-iTCR protein complex, PDBID: 3HUJ) that activates iNKT cells. Intersection amino acid residues between reported amino acid residues and candidate amino acid residues were selected as key amino acid residues. The obtained amino acid residues are shown in Table 1.
[0044] Table 1
[0045]
[0046] (3) Molecular docking workflow based on AutoDock4GPU and interaction analysis workflow based on PLIP library:
[0047] AutoDock4GPU was used to perform batch and repeated molecular docking on the small molecule SMILES strings after outlier removal. Since glycolipid antigen molecules have large molecular weights and a large number of rotatable bonds, the hyperparameters for molecular docking were modified to: nrun=50, lsmet=ad, heurmax=24000000. Ten molecular dockings were performed on each small molecule sample using these hyperparameters. The three-dimensional coordinates and atomic charges of the optimal conformation from the ten docking results were used as independent variables (X3) for subsequent construction of a small sample regression model.
[0048] For each optimal conformation, PLIP analysis was used to analyze its interactions with key amino acid residues. The analyzed interactions included: hydrogen bond donors, hydrogen bond acceptors, hydrogen bond strength, halogen bond donors, halogen bond acceptors, halogen bond strength, hydrophobic interactions, π-ion interactions, π-stacking interactions, number of hydrogen bonds, number of halogen bonds, number of hydrophobic interactions, number of π-ion interactions, and number of π-stacking interactions. The average interaction between each key amino acid residue and the glycolipid antigen molecule in 10 molecular dockings was used as the independent variable (X4) for subsequent construction of a small-sample regression model.
[0049] (4) Construction of a small-sample transfer learning model for Th1 immune response based on pre-trained models and prior knowledge:
[0050] This step employs a multimodal feature fusion approach to predict the bioactivity of glycolipid antigen molecules. A deep learning-based multimodal feature extraction and fusion framework is constructed, comprising three parallel feature extraction modules: a pre-trained model, GeminiMol, for processing the SMILES representation of molecules; a voxel grid-based graph neural network (GNN) for capturing the three-dimensional structural information of molecules; and a multilayer perceptron (MLP) for integrating the physicochemical properties and protein-protein interaction information of molecules. In this step, the model integrates prior knowledge through molecular spatial conformation information encoded by the GNN and molecular-protein interaction patterns encoded by the MLP network.
[0051] In this model, the SMIELS string of glycolipid antigen molecules is used as the independent variable (X1) and input into the GeminiMol model to extract the chemical structure features of the molecules. The three-dimensional coordinates and atomic charges of the glycolipid antigen molecules obtained from molecular docking are used as the independent variables (X3) and input into the voxel grid neural network (GNN) to process the three-dimensional spatial information of the molecules. The descriptors calculated by RDKit are used as the independent variables (X2), and the interaction relationships between the glycolipid antigen molecules obtained from molecular docking and key amino acid residues are used as the independent variables (X4) and input into the multilayer perceptron (MLP) network to integrate the molecular descriptors and protein interaction information. The loss function changes during the training process of the MLP to extract molecular features as follows: Figure 2 As shown.
[0052] The voxel grid-based graph neural network designed in this step extracts features from molecular conformations through an adaptive voxel generator. This generator employs a dynamic neighbor search strategy, using the KD-tree algorithm to identify the nearest neighbors (up to 6 neighbors) of each atom within a 3 Å cutoff distance, and constructs a local environment descriptor by weighting the contributions of neighboring atoms using a Gaussian weighting function.
[0053] This step achieves feature fusion using an attention mechanism across the three sets of multimodal features, adaptively learning the importance weights of different feature modalities. Specifically, feature fusion employs a Gaussian radial basis function for smooth feature mapping, combined with an attention mechanism for feature extraction. The Gaussian kernel function... Calculate the influence weights of atoms on surrounding spatial points, where For distance, The smoothing parameter is adjustable. The fused features are dimensionality reduced by principal component analysis and processed by a convolutional neural network, ultimately mapped to bioactivity data (Y).
[0054] Ablation experiments were conducted to test the performance of the Geminimol model and the Geminimol+prior knowledge model (which incorporates prior knowledge) on two types of biological activity data (the relative ratio of Th1 cytokine IFN-γ concentration and the relative ratio of Th1 / Th2 immune response). The changes in the loss function during training and the validation curves are shown below. Figure 3 and Figure 4 As shown, the model training results are presented in Tables 2 and 3:
[0055] The ratio model of KRN7000 derivative to KRN7000 for the Th1 cytokine IFN-γ is shown in Table 2:
[0056] Table 2
[0057]
[0058] For the Th1 / Th2 immune response ratio, the model for the ratio of KRN7000 derivatives to KRN7000 is as follows:
[0059] Table 3
[0060]
[0061] Through the above steps, a multimodal molecular feature extraction and activity prediction framework was constructed. This framework can effectively integrate multi-dimensional information such as the chemical structure, three-dimensional conformation, physicochemical properties, and protein-protein interactions of molecules. This multimodal feature fusion method not only overcomes the limitations of traditional single feature representation but also fully utilizes prior knowledge of molecule-target interactions, thereby providing more accurate and reliable bioactivity predictions.
[0062] The workflow designed in this example has several significant advantages: First, by introducing the pre-trained model GeminiMol to implement a transfer learning strategy, it effectively solves the challenge of training on small datasets. Second, the attention-based feature fusion method can adaptively learn the importance of different feature modalities, achieving dynamic weight allocation of features. Third, by integrating molecular docking information and protein-protein interaction features, the model possesses the ability to explain molecular mechanisms of action. Finally, the modular design of this framework gives it good scalability and adaptability, allowing for flexible adjustments or additions of new feature extraction modules according to specific needs. This systematic analytical approach not only provides a reliable computational tool for predicting the activity of glycolipid antigen molecules but also offers a valuable solution for other similar drug design tasks.
[0063] In summary, this invention proposes a quantitative prediction model for the effects of iNKT cell Th1 agonists using a pre-trained model and prior knowledge. Through an improved computational model, it achieves efficient quantitative prediction of the Th1 immune response tendency of glycolipid antigen molecules.
[0064] The specific exemplary embodiments described above are for illustrative and explanatory purposes, and are merely specific implementations of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A quantitative prediction model for the effect of iNKT cell Th1-type agonists based on a pre-trained model and prior knowledge of molecular docking, characterized in that, Includes the following steps: (1) Dataset construction: Collect and organize molecular structure and bioactivity data of glycolipid antigen molecule KRN7000 and its derivatives; calculate the correlation coefficient between the collected molecular structure and Th1 immune response bias; use RDKit to calculate the molecular descriptor and physicochemical properties of collected glycolipid antigen molecule KRN7000 and its derivatives as a supplement to the dataset. (2) Screening of key amino acid residues in the target protein group of activated iNKT cell immune response: The amino acids around the docking pocket of the target protein group of activated iNKT cell by glycolipid antigen molecules were calculated as candidate amino acids, and the relevant amino acids reported in the literature were selected as key amino acids. (3) Molecular docking based on AutoDock4GPU and interaction analysis based on PLIP library: For the collected and sorted molecular structures of glycolipid antigen molecules KRN7000 and its derivatives, the AutoDock4GPU software was used to perform batch and repeated molecular docking with the target protein group of activated iNKT cells. The interaction relationship between the molecular docking results and key amino acid residues was analyzed using the PLIP library to construct the original data features. (4) Construction of a small sample transfer learning model for Th1 immune response based on pre-trained model and prior knowledge: The physicochemical properties of glycolipid antigen molecules and the interaction relationship obtained in step (3) are used as the original data features, and the bioactivity data of glycolipid antigen molecules are used as data labels. They are input into the pre-trained model and the prior knowledge model respectively, and the regression model is trained after feature fusion.
2. The quantitative prediction model for the effect of iNKT cell Th1-type agonists based on a pre-trained model and prior knowledge of molecular docking, as described in claim 1, is characterized in that... In step (1), the correlation coefficient is calculated by using the Geminimol model to calculate the correlation coefficients between molecular structure and Th1-type immune response bias, namely the Pearson coefficient and the Spearman coefficient.
3. The quantitative prediction model for the effect of iNKT cell Th1-type agonists based on a pre-trained model and prior knowledge of molecular docking, as described in claim 1, is characterized in that... In step (1), the bioactivity data refers to the content data of Th1 / Th2 cytokines induced by the molecular structure; the Th1 / Th2 cytokine content data includes: the relative ratio of the peak concentration of IFN-γ, a Th1 immune response cytokine, to glycolipid antigen molecules relative to KRN7000, and the relative ratio of the secretion content of Th1 / Th2 cytokines induced by glycolipid antigen molecules relative to KRN7000.
4. The quantitative prediction model for the effect of iNKT cell Th1-type agonists based on a pre-trained model and prior knowledge of molecular docking, as described in claim 1, is characterized in that... In step (1), the descriptor includes Morgan fingerprint, molecular weight, lipid-water partition coefficient, topological polar surface area, drug-likeness, number of hydrogen bond donors, number of hydrogen bond acceptors, and number of rotatable bonds.
5. The quantitative prediction model for the effect of iNKT cell Th1-type agonists based on a pre-trained model and prior knowledge of molecular docking, as described in claim 1, is characterized in that... In step (2), the candidate amino acid residues are calculated using BioPython and are the binding pockets of the target protein group and glycolipid molecule antigens during iNKT cell immune activation. Amino acid residues within 5 Å of the docking pocket position are selected as candidate amino acid residues; the target protein group is the CD1d protein and iTCR protein complex.
6. The quantitative prediction model for the effect of iNKT cell Th1-type agonists based on a pre-trained model and prior knowledge of molecular docking, as described in claim 1, is characterized in that... In step (3), the molecular docking work based on AutoDock4GPU converts the SMILES string of molecular structure into a PDBQT format file with three-dimensional structural information and atomic charge information through OpenBabel. Under the same conditions, molecular docking is performed on all collected molecules in batches and repeatedly through the mesh parameter file.
7. The quantitative prediction model for the effect of iNKT cell Th1-type agonists based on a pre-trained model and prior knowledge of molecular docking, as described in claim 1, is characterized in that... In step (3), the interaction relationships of key amino acid residues are analyzed based on the PLIP library. The analysis scope includes hydrogen bond donors, hydrogen bond acceptors, hydrogen bond strength, halogen bond donors, halogen bond acceptors, halogen bond strength, hydrophobic interactions, π-ion interactions, π stacking interactions, number of hydrogen bonds, number of halogen bonds, number of hydrophobic interactions, number of π-ion interactions, and number of π stacking interactions.
8. The quantitative prediction model for the effect of iNKT cell Th1-type agonists based on a pre-trained model and prior knowledge of molecular docking, as described in claim 1, is characterized in that... In step (4), the regression model is trained using the interaction between the pre-trained model Geminimol and the target protein group of iNKT cells activated by glycolipid antigen molecules as prior knowledge; The regression model is composed of features extracted by the pre-trained model Geminimol through unsupervised learning and features extracted by prior knowledge through unsupervised learning, which are then fused together. The prior knowledge consists of two parts: the three-dimensional coordinate features of glycolipid antigen molecules after repeated molecular docking in batches through unsupervised learning using voxels and graph neural networks, and the molecular descriptors and physicochemical properties of glycolipid antigen molecules through unsupervised learning using multilayer perceptrons. The features obtained through unsupervised learning are fused through an attention mechanism, then reduced in dimensionality through principal component analysis and input into a convolutional neural network. Finally, they are mapped to the bioactivity data labels of glycolipid antigen molecules through a fully connected layer.
9. The quantitative prediction model for the effect of iNKT cell Th1-type agonists based on a pre-trained model and prior knowledge of molecular docking, as described in claim 8, is characterized in that: The three-dimensional coordinate features of glycolipid antigen molecules after multiple repeated molecular dockings through unsupervised learning using voxels and graph neural networks are as follows: The multiple docking results of glycolipid antigen molecules are used to extract features from the molecular conformation through an adaptive voxel generator. This generator adopts a dynamic neighbor search strategy, uses the KD tree algorithm to identify the nearest neighbor of each atom within a 3 Å cutoff distance, and weights the contribution of neighboring atoms through a Gaussian weight function. For each atom, basic features including charge information, atom coordinates, and atom type were extracted, and combined with the statistical features of its neighboring atoms: mean and standard deviation, to construct a local environment descriptor.
10. The quantitative prediction model for the effect of iNKT cell Th1-type agonists based on a pre-trained model and prior knowledge of molecular docking, as described in claim 8, is characterized in that... The feature fusion employs Gaussian radial basis functions for feature smoothing mapping and combines an attention mechanism for feature extraction.