A machine learning-based nucleoside derivative gel-forming ability prediction model
The nucleoside derivative gelation ability prediction model constructed by machine learning solves the problem that existing technologies cannot accurately predict the formation of hydrogels by nucleoside derivatives, achieves a high success rate in prediction, and provides a tool for designing nucleoside derivative hydrogels.
Patent Information
- Application Number
- CN202310947817.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-28
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2043-07-28
AI Technical Summary
Existing technologies struggle to accurately predict whether nucleoside derivatives can form hydrogels, primarily due to limited understanding of the relationship between nucleoside structure and gelation ability.
Machine learning methods were used to construct a predictive model for the gelation ability of nucleoside derivatives through feature extraction and logistic regression algorithms. Using features such as two-dimensional matrix descriptors and edge adjacency index, 24 molecular descriptors were selected and hyperparameters were optimized to establish the optimal model.
It achieved accurate prediction of the gel-forming ability of nucleoside derivatives with a success rate of 83.33%, providing a tool for the design of nucleoside derivative hydrogels.
Smart Images

Figure CN116741306B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of computer prediction system, and particularly relates to a nucleoside derivative gel-forming ability prediction model based on machine learning. BACKGROUND
[0002] In recent years, nucleoside supramolecular hydrogels have attracted increasing attention in the biomedical field due to their various non-covalent interactions, unique properties and excellent biocompatibility. Since it was reported in 1910 that concentrated solutions of guanylic acid could form gels, the research of nucleoside-based hydrogels has made great progress and has been used in various applications, including drug delivery, biosensors and tissue engineering. For example, Lehn et al. showed that guanosine derivatives are stable supramolecular hydrogels in the presence of metal cations, and provide highly selective and controllable release of bioactive substances, making them an attractive option for drug delivery. Davis et al. designed long-acting guanosine-borate hydrogels for sustained drug release.
[0003] The main challenge in this field is to predict whether a nucleoside derivative can form a hydrogel. Design suggestions for nucleoside derivative hydrogels are often proposed, but gelators are usually discovered accidentally or synthesized by modifying existing gelators. This fundamental limitation is due to the current human knowledge of the relationship between nucleoside structure and gelation ability.
[0004] Machine learning (ML) is a powerful technology that allows models to automatically learn from data and continuously improve their performance over time, helping to automate and optimize processes and improve decision-making.
[0005] The advantage of ML in predicting the ability of molecular hydrogels to form is that it can learn the structure-property relationship of molecules through high-dimensional data, which can be used to mine new gelators. In recent years, there has been significant progress in the application of ML in hydrogels. For example, Li et al. developed ML models to learn the correlation between chemical features and dipeptide hydrogen formation ability of peptoid-like molecules through algorithms such as Random forest (RF), Gradient Boosting (GB) and Logistic Regression (LR). However, due to the complexity of the self-assembly process of nucleoside derivatives to form hydrogels, there is currently no model that can accurately predict nucleoside hydrogels. SUMMARY
[0006] The present application aims to provide a nucleoside derivative gel-forming ability prediction model based on machine learning.
[0007] The application provides a method for constructing a nucleoside derivative gel-forming ability prediction model based on machine learning, comprising the following steps:
[0008] (1) Feature extraction: collect nucleoside derivatives, convert by simplified molecular-input line-entry system, select 8 types of molecular descriptors of nucleoside derivatives as features, and the 8 types of molecular descriptors are two-dimensional matrix descriptor, edge adjacency index, p_vsa type descriptor, two-dimensional atom pair, two-dimensional autocorrelation, atomic center fragment, functional group count and drug group descriptor;
[0009] (2) Model construction: the aforementioned features are trained by using a logistic regression algorithm to obtain a nucleoside derivative gel-forming ability prediction model.
[0010] Further, in step (1), the following 24 types of molecular descriptors are used as features: CATS2D_06_DL, B09[O-O], P_VSA_charge_7, H-052, CATS2D_03_DL, nN(CO)2, CATS2D_04_AA, CATS2D_05_DA, C-016, F07[N-O], VE1sign_Dz(v), CATS2D_05_DL, F05[N-N], P_VSA_charge_4, F10[O-O], MATS3p, VE3sign_D / Dt, SM10_AEA(dm), SpDiam_AEA(ed), VE1sign_B(p), GATS6i, SpMAD_EA(ri), GATS7s and CATS2D_09_DA.
[0011] Further, in step (2), the parameters of the logistic regression algorithm are as follows: 'logreg_c': 0.030490410588918396, 'l1_ratio': 0.9111079659846881,'max_iter': 85.
[0012] The application also provides a nucleoside derivative gel-forming ability prediction model constructed by the above method.
[0013] The application also provides a nucleoside derivative gel-forming ability prediction method,
[0014] 1) converting the nucleoside derivative to be detected by simplified molecular-input line-entry system, selecting 8 types of molecular descriptors of nucleoside derivatives as features, and the 8 types of molecular descriptors are two-dimensional matrix descriptor, edge adjacency index, p_vsa type descriptor, two-dimensional atom pair, two-dimensional autocorrelation, atomic center fragment, functional group count and drug group descriptor;
[0015] 2) input the descriptor results of the nucleoside derivative obtained in step 1 into the prediction model as described above;
[0016] 3) obtain the prediction results of the gelation ability of the nucleoside derivative.
[0017] Further, in step (1), the following 24 molecular descriptors are used as features: CATS2D_06_DL, B09[O-O], P_VSA_charge_7, H-052, CATS2D_03_DL, nN(CO)2, CATS2D_04_AA, CATS2D_05_DA, C-016, F07[N-O], VE1sign_Dz(v), CATS2D_05_DL, F05[N-N], P_VSA_charge_4, F10[O-O], MATS3p, VE3sign_D / Dt, SM10_AEA(dm), SpDiam_AEA(ed), VE1sign_B(p), GATS6i, SpMAD_EA(ri), GATS7s, and CATS2D_09_DA.
[0018] Further, in step (2), the parameter selection of the logistic regression algorithm is as follows: 'logreg_c': 0.030490410588918396, 'l1_ratio': 0.9111079659846881,'max_iter': 85.
[0019] The present application also provides a system for implementing the method of the nucleoside derivative gelation ability prediction model according to any one of the above, characterized in that it comprises
[0020] An input module for inputting the descriptor of the nucleoside derivative to be detected;
[0021] A calculation module comprising the model and the comparison module as described above, wherein the comparison module is used to compare the descriptor of the nucleoside derivative to be detected with the model;
[0022] An output module for outputting the final prediction results of the gelation performance of the nucleoside derivative.
[0023] Further, the descriptors include CATS2D_06_DL, B09[O-O], P_VSA_charge_7, H-052, CATS2D_03_DL, nN(CO)2, CATS2D_04_AA, CATS2D_05_DA, C-016, F07[N-O], VE1sign_Dz(v), CATS2D_05_DL, F05[N-N], P_VSA_charge_4, F10[O-O], MATS3p, VE3sign_D / Dt, SM10_AEA(dm), SpDiam_AEA(ed), VE1sign_B(p), GATS6i, SpMAD_EA(ri), GATS7s, and CATS2D_09_DA.
[0024] The application further provides a computer readable storage medium, which stores a computer program for implementing the nucleoside derivative gelation capacity prediction method.
[0025] The nucleoside derivative with known gelation capacity includes nucleoside derivatives capable of forming gels and nucleoside derivatives incapable of forming gels.
[0026] The application successfully establishes an optimal ML model for predicting the hydrogel formation capacity of nucleoside derivatives based on feature selection, hyperparameter optimization and algorithm comparison. The model can accurately predict whether the nucleoside derivative has gelation capacity. First, 4175 molecular descriptors and four kinds of fingerprints are used to convert the data set of 71 nucleoside derivatives with related gelation capacity information into a feature matrix, and through feature selection and hyperparameter optimization, four kinds of classifier algorithms are used to predict the hydrogel formation capacity, and an optimal model with an accuracy of 71% is screened out. Second, the optimal ML model is used to predict the hydrogel formation capacity of 7257 publicly disclosed nucleoside derivatives. The model can accurately predict whether the nucleoside derivative has gelation capacity. The model of the application can effectively predict the gelation capacity of the nucleoside derivative to be detected. Twelve nucleoside gel agents with greater possibility are selected, and their hydrogel formation capacity is verified through experiments. Among them, 10 nucleoside derivatives can form hydrogels, and the success rate of forming hydrogels is 83.33%, indicating that the nucleoside derivative gelation capacity prediction model of the application has high prediction accuracy. The above findings show that the machine model provides a tool for predicting nucleoside derivatives with hydrogel formation capacity in the future.
[0027] Obviously, according to the above content of the application, according to the ordinary technical knowledge and common means in the art, other various forms of modifications, replacements or changes can be made without departing from the above basic technical ideas of the application.
[0028] The above summary of the present application will be further elaborated in the following detailed description in the form of examples. However, it should not be understood that the above summary of the present application is limited to the following examples. Any technology realized based on the above summary of the present application falls within the scope of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 For descriptor and fingerprint screening: (a) flow chart of molecular descriptor screening; (b) results of rank sum test; (c) correlation heat map of 40 descriptors.
[0030] Figure 2 For evaluation index and feature importance of different models: (a) scatter plot of AUC (area under the curve) distribution and test accuracy of all models; (b) AUC curve of LR, decision tree (DT), random forest (RF), and extreme gradient boosting (XGBoost) four models; (c) results of feature importance of 24 descriptors on LR.
[0031] Figure 3To experimentally verify the novel nucleoside hydrogels predicted by the model: (a) This invention selected the top 10% of nucleoside derivatives with the highest predicted hydrogel formation probability, and selected 12 nucleoside derivatives in a uniform manner. (b). The results of biological experiments on 12 nucleoside derivatives with high gelation probability showed that 10 nucleoside derivatives could form hydrogels, that is, there were 10 true gelling agents. Gel (+) indicates that hydrogels can be formed, and Gel (-) indicates that hydrogels cannot be formed. The 12 nucleoside derivatives are (1) 1-[3,4-Dihydroxy-5-(hydroxymethyl)oxolan-2-yl]-1,3,5-triazinane-2,4,6-trione, i.e., DTT; (2) xanthosine, i.e., XTS; (3) guanine 5'-monophosphate, i.e., GMP; (4) inosine 5'-monophosphate, i.e., IMP; (5) 5-fluorouridine, i.e., 5-FUR; (6) 8-aminoguanosine, i.e., 8-AG; (7) 2'-deoxyguanosine 5'-monophosphate, i.e., dGMP; (8) 8-hydroxyguanosine, ie 8-OHG; (9) 8-azaguanosine, ie 8-azaG, (10) inosine-5'-carboxylic acid, ie I-5'-CA; (11) 2'-amino-2'-deoxyguanosine, ie 2'-NH2-dG, (12) 2'-O-Methylguanosine, namely 2'-OMe-dG.
[0032] Figure 4 The experiment was conducted to verify the results of inverting the vials of the novel nucleoside hydrogel predicted by the model, namely, the gel-forming ability of the nucleoside derivatives (1-4).
[0033] Figure 5 The experiment was conducted to verify the results of inverting the vial of the novel nucleoside hydrogel predicted by the model, namely, the gel-forming ability of the nucleoside derivatives (5-8).
[0034] Figure 6 The experiment was conducted to verify the results of inverting the vial of the novel nucleoside hydrogel predicted by the model, namely, the gel-forming ability of the nucleoside derivatives (9-12). Detailed Implementation
[0035] The raw materials and equipment used in this invention are all known products, obtained by purchasing commercially available products.
[0036] LR algorithm: Logistic Regression (LR).
[0037] Simplified Molecular Linear Input Specification (SMILES) is a specification that explicitly describes molecular structures using ASCII strings.
[0038] Example 1
[0039] I. Experimental Methods
[0040] (a) Data collection and feature selection:
[0041] 1. Data Collection
[0042] This invention collected 71 nucleoside derivatives with relevant gelation information by searching the literature. The 71 nucleoside derivatives were divided into two groups according to their hydrogel-forming ability: gelling agent group (n=38) and non-gelling agent group (n=33). Those that can form gels are in the gelling agent group, and those that cannot form gels are in the non-gelling agent group.
[0043] 2. Construct the initial feature matrix
[0044] For the subsequently constructed ML model, all molecules are converted into the Simplified Molecular Input Line Entry System (SMILES). SMILES is used to calculate commonly used molecular descriptors and four types of molecular fingerprints, thereby transforming the molecular structure information of the above 71 nucleoside derivatives into a feature matrix that can be used for model training.
[0045] 4175 molecular descriptors and four molecular fingerprints were calculated for 71 nucleoside derivatives, resulting in five feature matrices representing the molecular structure information of the nucleoside derivatives for subsequent model construction. The 4175 descriptors for the 71 nucleoside derivatives were calculated using the alvaDesc software (version 2.0.12). The four most commonly used molecular fingerprints for the 71 nucleoside derivatives—ECFP4 fingerprint, ECFP6 fingerprint, AtomPair fingerprint, and Topological Torsion fingerprint—were calculated using the RDKit package.
[0046] 3. Feature Filtering
[0047] To avoid overfitting and improve model accuracy, this invention employs a three-stage screening process, including rank-sum test, Pearson correlation coefficient, and recursive feature elimination (RFE). First, the rank-sum test identifies 144 descriptors that show significant differences between the gel and non-gel groups (P < 0.05).Figure 1 (b) Next, the 144 descriptors were normalized and then further filtered using Pearson's correlation coefficient. Molecular descriptors with high correlations (r < 0.8) were excluded. Figure 1 c) After obtaining 40 descriptors, the correlation heatmap results are as follows: Figure 1 As shown in c. Finally, the number of descriptors required for various algorithms to build the optimal model is found through recursive feature elimination.
[0048] 4. Algorithm selection and model determination
[0049] This invention proposes four commonly used machine learning algorithms to construct prediction models: Logistic Regression (LR), Decision Tree (DT), Regression Randomization (RF), and Extreme Gradient Boosting (XGBoost). Utilizing the previously mentioned four molecular fingerprints and four different molecular descriptors before and after the three-step feature selection, a total of eight feature matrices are used as model inputs. Combining the four machine learning algorithms, a total of 32 feature matrices are constructed for building the machine learning model. The rows of the feature matrices represent 71 nucleoside derivatives, and the columns represent the calculated molecular fingerprints or molecular descriptors. Molecular fingerprints have a fixed 2048 columns. Molecular descriptors have different numbers of columns depending on the feature selection. Since the recursive feature elimination results of each machine learning algorithm are different, Descriptors-REF also has four different column numbers: LR (24 descriptors), XGBoost (16 descriptors), DT (30 descriptors), and RF (37 descriptors). Specific details are shown in Table 1.
[0050]
[0051] 5. Model hyperparameter optimization
[0052] After features and labels are created, different ML models are trained and evaluated.
[0053] The model was built using Scikit-learn (Version 1.1.1, including DT, RF, and LR) and xgboost (Version 1.7.2), and hyperparameter optimization and recursive feature elimination (RFE) were performed using the Optuna package (Version 3.0.5). All statistical analyses in this invention were performed using Python (Version 3.9.12).
[0054] The optimization scope is as follows:
[0055] ①DT
[0056] max_depth: 3-5, step=1 (meaning the value range is 3-5, changing by 1 each time, i.e., taking values 3, 4, 5)
[0057] max_features : 10-20, step=1
[0058] min_samples_split:2-25, step=1
[0059] Final value
[0060] 'max_depth': 3, 'max_features': 14, 'min_samples_split': 2
[0061] ②LR
[0062] logreg_c: 1e-3, 1e3, log=True (indicates continuous logarithmic values)
[0063] l1_ratio: 0.1,1, log=False
[0064] max_iter: 100,2000, step=1
[0065] Final value
[0066] 'logreg_c': 0.030490410588918396, 'l1_ratio': 0.9111079659846881, 'max_iter': 85
[0067] ③RF
[0068] n_estimators: 100,1000, step=1
[0069] max_depth: 5, 20, step=1
[0070] max_features: 5, 30, step=1
[0071] min_impurity_decrease: 0,5, log=False
[0072] 'n_estimators': 913, 'max_depth': 8, 'max_features': 27, 'min_impurity_decrease': 0.004905793398417582
[0073] ④Xgboost
[0074] Lambda: 1e-3, 10, log=False
[0075] Alpha: 1e-3, 10, log=False
[0076] colsample_bytree: 0.3,1.0, step=0.1
[0077] Subsample: 0.4, 1.0, step=0.1
[0078] learning_rate: 0.0001, 0.2, step=0.005
[0079] n_estimators: 50,1000, step=1
[0080] Final value
[0081] 'lambda': 0.052739889359645756, 'alpha': 0.0030422324066177544, 'colsample_bytree': 0.9000000000000001, 'subsample': 0.5, 'learning_rate':0.15009999999999998, 'n_estimators': 330
[0082] Evaluation metrics included accuracy, area under the curve (AUC), and F1 score. Ten independent five-fold cross-validations were performed on hyperparameter optimization, recursive feature elimination (RFE), and the calculation of evaluation metrics. Model performance was evaluated before and after each step of the screening process using different numbers of molecular descriptors and fingerprints.
[0083] (II) Model selection and feature importance analysis:
[0084] The final model built using the RFE descriptor performs slightly better than other ML models. Figure 2 a, Figure 2 b). The LR model based on 24 molecular descriptors (Table 2) after RFE (red) was considered the optimal model (accuracy, 0.71%, 95% confidence interval (CI), 0.69–0.73; f1 score, 0.78, 95% CI 0.76–0.80; AUC, 0.84, 95% CI, 0.82–0.86).
[0085] The key feature descriptor categories for the optimal model LR mainly focus on 2D matrix-based descriptors, edge adjacency indices, P_VSA-like descriptors, 2D atom pairs, 2D autocorrelations, atom-centered fragments, functional group counts, and pharmcophore descriptors (e.g., ...). Figure 2 c and Table 1).
[0086] It is noteworthy that five molecular descriptors have an importance exceeding 0.01 in the LR model (CATS2D_06_DL, CATS2D_03_DL, P_VSA_charge_7, B09[OO], H-052, ...). Figure 2 c) The above five molecular descriptors are mainly related to hydrogen bonds, molecular polarity, and lipophilicity, which are key descriptors of the properties of gelling agents.
[0087] By constructing predictive models for 71 nucleoside derivatives using predefined molecular descriptors, this invention demonstrates that molecular dynamics (ML) can indeed help predict the hydrogel-forming ability of nucleoside derivatives, a novel discovery in this field.
[0088] Table 2 24 Molecular Descriptors
[0089]
[0090] The following experimental examples demonstrate the predictive effectiveness of the nucleoside gelling ability model of this invention.
[0091] Experimental Example 1: The gelling ability of nucleoside derivatives predicted by the model of this invention.
[0092] 1. Experimental Methods
[0093] To test the reliability of the ML model of this invention, all available nucleoside derivatives (n=7257) were screened from organic small molecule bioactivity data (PubChem database), and there are currently no reports on hydrogel-forming ability.
[0094] The predicted probabilities of the optimal model (LR with 24 descriptors) determined in Example 1 were ranked according to probability, and those with a predicted probability greater than 50% were considered potential gelling agents. Based on previous experience and knowledge of the hydrogel-forming ability of nucleoside derivatives, this invention selected 12 guanosine derivatives from the top 10% of the structures with the highest predicted probabilities for verification. Figures 3-6 These 12 nucleoside derivatives are (1) 1-[3,4-Dihydroxy-5-(hydroxymethyl)oxolan-2-yl]-1,3,5-triazinane-2,4,6-trione, i.e., DTT; (2) xanthosine, i.e., XTS; (3) guanine 5'-monophosphate, i.e., GMP; (4) inosine 5'-monophosphate, i.e., IMP; (5) 5-fluorouridine, i.e., 5-FUR; (6) 8-aminoguanosine, i.e., 8-AG; (7) 2'-deoxyguanosine 5'-monophosphate, i.e., dGMP; (8) 8-hydroxyguanosine, i.e., 8-OHG; (9) 8-azaguanosine, i.e., 8-azaG; (10) inosine-5'-carboxylic acid, i.e., I-5'-CA; (11) 2'-amino-2'-deoxyguanosine, that is, 2'-NH2-dG, (12) 2'-O-Methylguanosine, that is, 2'-OMe-dG.
[0095] 2. Experimental Results
[0096] The results showed that 10 nucleoside derivatives (1, 3, 4, 6, 7, 8, 9, 10, 11, 12) could form hydrogels, while the other two (2, 5) did not form hydrogels. Figures 3-6The success rate of hydrogel formation was 83.33% (10 / 12). Specifically, 1, 3, 4, 7, 10, and 12 formed hydrogels in the presence of AgNO3. 6 and 8 self-assembled into hydrogels in boric acid and Tris solutions, as well as NaB(OH)4 and KB(OH)4 solutions. 9 formed a hydrogel in boric acid and Tris solutions, as well as silver nitrate solution. 11 could self-assemble into a hydrogel in potassium chloride, sodium chloride, NaB(OH)4, and KB(OH)4 solutions. Of the 10 gelling nucleoside derivatives, 8 (1, 6, 7, 8, 9, 10, 11, and 12) have not yet been reported as gelling agents.
[0097] Experimental results show that the modeling method of the present invention can effectively predict the gelling ability of the nucleoside derivatives to be tested.
[0098] In summary, this invention provides an optimal machine learning (ML) model for predicting the hydrogel-forming ability of nucleoside derivatives based on feature selection, hyperparameter optimization, and algorithm comparison. First, the feature matrices (i.e., 4175 molecular descriptors and 4 fingerprints for 71 nucleoside derivatives with gel-forming information) were calculated using alvaDesc software (version 2.0.12) and the Rdkit package. Through feature selection and hyperparameter optimization, four classifier algorithms were used to predict hydrogel-forming ability, and the optimal model with an accuracy of 71% was selected. Second, the optimal ML model was applied to predict the hydrogel-forming ability of 7257 publicly disclosed nucleoside derivatives. Twelve nucleoside gelling agents with high probability were selected, and their hydrogel-forming ability was verified experimentally. Ten of these nucleoside derivatives were found to form hydrogels, with a success rate of 83.33%, indicating that the nucleoside derivative gel-forming ability prediction model of this invention has high prediction accuracy. These findings demonstrate that the machine learning model of this invention provides a tool for predicting nucleoside derivatives with hydrogel-forming ability.
Claims
1. A method for constructing a machine learning-based model to predict the gelling ability of nucleoside derivatives, characterized in that: Includes the following steps: (1) Feature extraction: Nucleoside derivatives were collected and converted by a simplified molecular linear input system. Eight types of molecular descriptors of nucleoside derivatives were selected as features. The eight types of molecular descriptors are two-dimensional matrix descriptor, edge adjacency index, p_vsa descriptor, two-dimensional atom pair, two-dimensional autocorrelation, atomic center fragment, functional group count and drug group descriptor. The conversion process includes: converting all nucleoside derivative molecules into a standard simplified molecular linear input system, using the standard simplified molecular linear input system to calculate commonly used molecular descriptors and four types of molecular fingerprints to obtain an initial feature matrix. The four types of molecular fingerprints are ECFP4 fingerprint, ECFP6 fingerprint, AtomPair fingerprint, and Topological Torsion fingerprint. The selection of features involved a screening process, which included: rank-sum test, Pearson correlation coefficient, and recursive feature elimination. First, the rank-sum test was used to screen out 144 descriptors that showed significant differences between the gelling agent group and the non-gelling agent group. The gelling agent group consisted of nucleoside derivatives that could form gels, while the non-gelling agent group consisted of nucleoside derivatives that could not form gels. Next, the 144 descriptors were normalized, and then further screened using the Pearson correlation coefficient to eliminate highly correlated molecular descriptors, resulting in 40 descriptors. Finally, recursive feature elimination was used to find the number of descriptors required for various algorithms to construct the optimal model. (2) Model construction: The aforementioned features are trained using a logistic regression algorithm to obtain a predictive model for the gelling ability of nucleoside derivatives, which includes the following: The feature matrix is obtained using four molecular fingerprints and different molecular descriptors before and after screening, and is used as the input of the model. The feature matrix is further constructed by combining the logistic regression algorithm to build a machine learning model. The rows of the feature matrix represent nucleoside derivatives and the columns represent the calculated molecular fingerprints or molecular descriptors.
2. The method according to claim 1, characterized in that: In step (1), the following 24 molecular descriptors are used as features: CATS2D_06_DL, B09[OO], P_VSA_charge_7, H-052, CATS2D_03_DL, nN(CO)2, CATS2D_04_AA, CATS2D_05_DA, C-016, F07[NO], VE1sign_Dz(v), CATS2D_05_DL, F05[NN], P_VSA_charge_4, F10[OO], MATS3p, VE3sign_D / Dt, SM10_AEA(dm), SpDiam_AEA(ed), VE1sign_B(p), GATS6i, SpMAD_EA(ri), GATS7s and CATS2D_09_DA.
3. The method according to claim 1 or 2, characterized in that: In step (2), the parameters of the logistic regression algorithm are selected as follows: 'logreg_c': 0.030490410588918396, 'l1_ratio': 0.9111079659846881, 'max_iter':
85.
4. A method for predicting the gelling ability of nucleoside derivatives, characterized in that: 1) The nucleoside derivatives to be tested are converted by a simplified molecular linear input system, and eight types of molecular descriptors of nucleoside derivatives are selected as features. The eight types of molecular descriptors are two-dimensional matrix descriptor, edge adjacency index, p_vsa type descriptor, two-dimensional atom pair, two-dimensional autocorrelation, atomic center fragment, functional group count and drug group descriptor. 2) Input the nucleoside derivative descriptor results obtained in step 1 into the prediction model, which is constructed according to the method described in any one of claims 1-3; 3) Obtain the predicted results of the gelling ability of nucleoside derivatives.
5. The method according to claim 4, characterized in that: In step (1), the following 24 molecular descriptors are used as features: CATS2D_06_DL, B09[OO], P_VSA_charge_7, H-052, CATS2D_03_DL, nN(CO)2, CATS2D_04_AA, CATS2D_05_DA, C-016, F07[NO], VE1sign_Dz(v), CATS2D_05_DL, F05[NN], P_VSA_charge_4, F10[OO], MATS3p, VE3sign_D / Dt, SM10_AEA(dm), SpDiam_AEA(ed), VE1sign_B(p), GATS6i, SpMAD_EA(ri), GATS7s and CATS2D_09_DA.
6. The method according to claim 4 or 5, characterized in that: In step (2), the parameters of the logistic regression algorithm are selected as follows: 'logreg_c': 0.030490410588918396, 'l1_ratio': 0.9111079659846881, 'max_iter':
85.
7. A system for implementing the method for predicting the gelling ability of nucleoside derivatives according to any one of claims 4-6, characterized in that, include The input module is used to input the descriptor of the nucleoside derivative to be tested; The calculation module includes the prediction model and the comparison module, wherein the comparison module is used to compare the descriptor of the nucleoside derivative to be detected with the model. The output module is used to output the predicted results of the final gelling properties of the nucleoside derivative.
8. The system according to claim 7, characterized in that: The descriptors include CATS2D_06_DL, B09[OO], P_VSA_charge_7, H-052, CATS2D_03_DL, nN(CO)2, CATS2D_04_AA, CATS2D_05_DA, C-016, F07[NO], VE1sign_Dz(v), CATS2D_05_DL, F05[NN], P_VSA_charge_4, F10[OO], MATS3p, VE3sign_D / Dt, SM10_AEA(dm), SpDiam_AEA(ed), VE1sign_B(p), GATS6i, SpMAD_EA(ri), GATS7s, and CATS2D_09_DA.
9. A computer-readable storage medium, characterized in that: It contains a computer program for implementing the method for predicting the gelling ability of nucleoside derivatives according to any one of claims 4-6.