Breast cancer candidate drug classification prediction method based on AFNN and ADE

By optimizing the molecular descriptor using the adaptive gradient descent algorithm (AFNN) and the ADE algorithm, and combining it with a probabilistic neural network, the stability and scalability issues of anti-breast cancer drug prediction in existing technologies were resolved. This resulted in more accurate prediction of compound biological activity and ADMET properties, thus optimizing drug screening performance.

CN117238395BActive Publication Date: 2026-04-24HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAIYIN INSTITUTE OF TECHNOLOGY
Filing Date
2023-09-04
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing techniques for predicting the bioactivity of anti-breast cancer drug candidates and classifying their properties using the ADMET algorithm suffer from insufficient stability and limited scalability. Simple genetic algorithms or particle swarm optimization algorithms are prone to getting stuck in local optima and struggle to obtain the optimal molecular descriptors and their values.

Method used

An adaptive fuzzy neural network (AFNN) was optimized using an adaptive gradient descent algorithm and combined with a probabilistic neural network (PNN) to construct a pIC50 prediction model for compound bioactivity and an ADMET property classification model. The values ​​of the compound molecular descriptors were then optimized using an adaptive differential evolution algorithm (ADE).

Benefits of technology

This improved the accuracy of predicting compound bioactivity and the precision of ADMET property classification, ensuring that compounds have better antagonistic effects against ERα while also possessing better ADMET properties, thus optimizing the drug screening process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117238395B_ABST
    Figure CN117238395B_ABST
Patent Text Reader

Abstract

The application discloses a breast cancer candidate drug classification prediction method based on AFNN and ADE, and comprises the following steps: performing feature selection on known experimental data by using a random forest RF algorithm, and selecting molecular descriptors with influence; constructing a compound biological activity pIC 50 prediction model based on AFNN; performing parameter optimization on the compound biological activity pIC 50 prediction model based on AFNN by using an adaptive gradient descent algorithm; establishing an ADME / T property classification model based on a probability neural network; and processing a constraint optimization problem by using an adaptive differential evolution algorithm ADE to find the optimal value corresponding to the compound molecular descriptor. The application can find the optimal value corresponding to the compound molecular descriptor, so that the compound has better antagonistic effect on ER alpha activity and has ADME / T properties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to drug classification and prediction methods in the fields of medical and artificial intelligence technologies, and more particularly to a method for classifying and predicting candidate drugs for breast cancer based on AFNN and ADE. Background Technology

[0002] In recent years, with the accelerated pace of life, the incidence of breast cancer in women has been increasing year by year. Therefore, improving the cure rate of breast cancer is urgent. Currently, commonly used treatments for breast cancer include surgery, radiotherapy, biological therapy, and targeted therapy. Surgery and radiotherapy are extremely expensive and may cause serious sequelae for patients; in comparison, targeted drug therapy has fewer side effects and is often the best choice for patients. Therefore, accelerating the research and development of targeted drugs for breast cancer is crucial.

[0003] Conventional methods for developing anti-breast cancer drugs include high-throughput screening and virtual screening of compounds. Due to advancements in computer information technology and artificial intelligence, the organic combination of experimental methods and intelligent information processing can accelerate drug development and reduce costs, attracting widespread attention from researchers. Intelligent information processing requires using molecular descriptors as independent variables and bioactivity values ​​as dependent variables to establish predictive models that can predict the activity of unknown compounds, thereby discovering novel compounds effective in inhibiting ERα.

[0004] For the estrogen receptor alpha (ERα) subtype, a therapeutic target for breast cancer, we provided bioactivity data for 1974 compounds. Based on the provided ERα antagonist information (1974 compound samples, each with 729 molecular descriptor variables, 1 bioactivity data point, and 5 ADMET property data points), we constructed quantitative prediction models for compound bioactivity and classification prediction models for ADMET properties, thus providing predictive services for simultaneously optimizing the bioactivity and ADMET properties of ERα antagonists. ADME primarily refers to the pharmacokinetic properties of a compound, considering only five ADME properties: 1) Caco-2 permeability, which measures the compound's ability to be absorbed by the human body; 2) Cytochrome P450 (CYP) 3A4 isoform (CYP3A4), a major metabolic enzyme in the human body, which measures the compound's metabolic stability; 3) Human Ether-a-go-go Related Gene (hERG), which measures the compound's cardiotoxicity; 4) Human Oral Bioavailability (HOB), which measures the proportion of a drug absorbed into the bloodstream after entering the human body; and 5) Micronucleus assay (MN), a method for detecting whether a compound has genotoxicity.

[0005] Among them, the bioactivity value of the compound against ERα (using IC50) 50 The value is an experimentally determined value, in nM. A smaller value indicates greater biological activity and greater effectiveness in inhibiting ERα activity. pIC 50 It is IC 50 The negative logarithm of the pIC value, which is typically positively correlated with biological activity, i.e. 50 A higher value indicates higher biological activity; pIC was used. 50 This is used to represent the bioactivity value.

[0006] The most similar prior art to this invention includes bioactivity prediction of anti-breast cancer drug candidates and ADMET property classification prediction technology. The function of this technology is to quantitatively predict the bioactivity of ERα antagonists of anti-breast cancer drug candidates, and to classify the five properties of compounds: Caco-2, CYP3A4, hERG, HOB, and MN.

[0007] The closest invention is "CN114627978A" "A Classification and Prediction Method for Anti-Breast Cancer Candidate Drugs Based on R-CNN-GA", which discloses a classification and prediction method for anti-breast cancer candidate drugs based on R-CNN-GA. After dimensionality reduction by RFE (Recursive Feature Elimination) and RF (Random Forest), a prediction and classification model based on CNN_FC combination is used to improve the accuracy of existing model prediction and multi-classification. Finally, it is optimized by combining genetic algorithm and SPSS software to obtain the optimal molecular descriptor and its value.

[0008] CN114334033A, entitled "Screening Method, System, and Terminal for Molecular Descriptors of Anti-Breast Cancer Drug Candidates," discloses a screening method, system, and terminal for molecular descriptors of anti-breast cancer drug candidates. The method involves acquiring bioactivity data of multiple compounds against ERα, with each compound configured with multiple molecular descriptors, resulting in a set of independent variables composed of these descriptors. A preliminary screening model is established based on the LASSO regression method. This model is then used to reduce the dimensionality of the independent variable set, yielding a preliminary set of variables with non-zero coefficients. Finally, a variable selection model is established based on the random forest recursive feature elimination method. This model iteratively selects features from the preliminary set of variables to obtain the optimal combination of feature variables with the highest classification accuracy.

[0009] CN114613512A, entitled "A Method, Apparatus, Equipment, and Storage Medium for Screening Anti-Breast Cancer Candidate Drugs," discloses a method, apparatus, equipment, and storage medium for screening anti-breast cancer candidate drugs. The method involves processing anti-breast cancer candidate drug usage data according to a preset method to obtain anti-breast cancer candidate drug usage characteristic data; establishing an initial anti-breast cancer candidate drug screening neural network model; training, validating, and testing the initial anti-breast cancer candidate drug screening neural network model based on the anti-breast cancer candidate drug usage characteristic data to obtain a target anti-breast cancer candidate drug screening neural network model; and inputting the anti-breast cancer candidate drug data to be screened into the target anti-breast cancer candidate drug screening neural network model to screen the anti-breast cancer candidate drugs.

[0010] CN115831375A, "A Method for Constructing an Efficacy Prediction Model for Anti-Breast Cancer Candidate Drugs," discloses a method for constructing an efficacy prediction model for anti-breast cancer candidate drugs. First, given molecular descriptor data is preprocessed to remove molecular descriptors that are ineffective in predicting the biological activity of the compounds. Then, random forest is used to screen out molecular descriptors that significantly affect biological activity. Next, a quantitative prediction model for the ERα biological activity of different compounds is established based on a BP neural network, and IC50 analysis is performed on the compounds. 50 and pIC 50Value prediction is performed first; then, using the best-performing BP neural network model with molecular descriptors, a classification prediction model for compounds is constructed based on ADMET data. The constructed classification prediction model is used to predict the Caco-2, CYP3A4, hERG, HOB, and MN results of the compounds. Finally, a combination of genetic algorithm and neural network model is used to select molecular descriptors, which not only ensures that the compounds have better biological activity in inhibiting ERα but also better ADEMT properties, thereby optimizing the prediction effect and leading to the selection of superior anti-breast cancer drug candidates.

[0011] Existing techniques for predicting the bioactivity of anti-breast cancer drug candidates and classifying their properties using ADMET models, while employing neural network models, lack stability and scalability. Simple genetic algorithms or particle swarm optimization algorithms are prone to getting trapped in local optima for discrete optimization problems, hindering the acquisition of optimal molecular descriptors and their values. Summary of the Invention

[0012] Objective of the Invention: To overcome the shortcomings of the prior art, the objective of this invention is to provide a classification and prediction method for anti-breast cancer drug candidates based on AFNN and ADE. Firstly, AFNN is optimized using an adaptive gradient descent algorithm, and a compound bioactivity pIC based on AFNN is constructed. 50 The invention first establishes a predictive model; then, it builds an ADMET property classification model based on a probabilistic neural network; finally, it uses the Adaptive Differential Evolutionary Algorithm (ADE) to handle constrained drug problems. This invention can quickly find the optimal values ​​for the molecular descriptors of compounds, enabling the compounds to exhibit better antagonistic effects on ERα activity while also possessing better ADMET properties.

[0013] Technical solution: The present invention provides a classification and prediction method for anti-breast cancer drug candidates based on AFNN and ADE, comprising the following steps:

[0014] (1) To address the problem of excessive features in the original dataset, the Random Forest (RF) algorithm is used to select features from the known experimental data and identify multiple influential molecular descriptors.

[0015] (2) Constructing AFNN-based compound bioactivity pIC 50 Predictive models utilize adaptive gradient descent algorithms to predict the bioactivity pIC of compounds based on AFNN. 50 The parameters of the prediction model are optimized to improve the model's prediction performance;

[0016] (3) Construct an ADMET property classification model based on a probabilistic neural network (PNN);

[0017] (4) The adaptive differential evolution algorithm (ADE) is used to process the constrained optimization problem and find the optimal value corresponding to the compound molecule descriptor.

[0018] In step (1), the Random Forest (RF) algorithm sorts the 729 molecular descriptors in descending order of importance, i.e., the average decreasing Gini index, and selects the top 20 molecular descriptors that have the most significant impact on biological activity as the main variables.

[0019] The specific implementation method of the Random Forest (RF) algorithm is as follows: Multiple decision trees are constructed to determine the initial drug characteristic data for candidate anti-breast cancer drugs. In the Random Forest RF classifier, the input data consists of 729 molecular descriptors and label data (i.e., biological activity) for 1947 compounds. The output data consists of the 20 molecular descriptors that have the most significant influence on biological activity. First, the data is standardized to initialize the random forest. The random forest is then trained using the label data, and feature importance and feature importance ranking indices are calculated. The average feature importance and the average feature importance ranking index are calculated N times. If the average feature importance ranking index remains unchanged after three rounds, the feature importance ranking index is returned. Features are selected based on the feature importance ranking index.

[0020] In step (2), considering the high computational accuracy of the adaptive fuzzy neural network (AFNN), a compound bioactivity pIC based on AFNN was established using 20 selected molecular descriptors. 50 Predictive model; using samples 1–1774 from 1974 original samples, an AFNN-based pIC of compound bioactivity was trained. 50 A prediction model was developed; its performance was then validated using 30 random samples. The mean squared error analysis was used to assess the model's fit and verify its rationality.

[0021] AFNN-based compound bioactivity pIC 50 The method for building the prediction model is as follows:

[0022] (1) Data processing

[0023] We screened 729 molecular descriptors corresponding to 50 compounds, and normalized the 20 selected molecular descriptors. Each row was used as a training set, and the AFNN bioactivity prediction model was trained using samples 1 to 1774 of the 1974 data sets. The model performance was then validated using 30 random samples.

[0024] (2) Designing the network layer structure: This invention uses a Mamdani-type fuzzy neural network for modeling.

[0025] In the modeling process of neural networks, gradient descent algorithm with adaptive learning rate is used to optimize the model parameters to improve the model's prediction performance. The update formulas for the center, width, and weights of AFNN are as follows:

[0026]

[0027]

[0028]

[0029] Among them, c ij σ is the center of the membership function. ij η is the width of the membership function, w is the connection weight of the system, η is the learning rate, and K is the system training error.

[0030] Because traditional FNNs have a fixed learning rate, they are difficult to adapt to the modeling needs of complex systems. This invention proposes an adaptive learning rate adjustment strategy as shown in the following equation:

[0031]

[0032] Among them, K t Let K be the system training error at time t. t-1 Let be the training error at time (t-1). When the error moves in the direction of decreasing, the learning rate increases; when the error moves in the direction of increasing, the learning rate decreases.

[0033] In step (2), the compound was subjected to pIC. 50 The prediction method for the values ​​is as follows: First, 20 molecular descriptors selected using the Random Forest method are normalized; each row is used as an input training set, and samples 1 to 1774 from the 1974 data sets are used as the training set; 30 random samples are used as the test set to verify the model performance. The parameters for RF are: 500 decision trees; the parameters for AFNN are: 30 rules, a maximum training step size of 1000 steps, and a learning rate of 0.01.

[0034] The validation samples showed an RMSE of 0.2914 and a MAPE of 4.14%, indicating that the model accuracy meets the requirements for modeling the activity of anti-breast cancer drug candidates. The final output obtained is pIC. 50 The predictive model for the values ​​was used to perform pIC analysis on 50 compounds. 50 Prediction of values.

[0035] In step (3), considering that probabilistic neural networks (PNNs) are commonly used for sample data classification, an ADMET property classification model based on a probabilistic neural network (PNN) was established using 20 selected molecular descriptors. The ADMET property classification model was trained using samples 1 to 1774 from 1974 original samples. The model performance was then validated using 30 random samples. The accuracy of the model was also analyzed.

[0036] The method for establishing the ADMET property classification model based on probabilistic neural networks (PNN) is as follows:

[0037] The 20 molecular descriptors selected by the Random Forest (RF) method in step (1) were normalized; each row was used as a set of input training data, and samples 1 to 1774 of the 1974 sets of data were used to train the ADMET property classification model based on PNN; and then 30 random samples were used to verify the model performance.

[0038] Based on the prediction results, a probabilistic neural network (PNN) was finally selected to construct an ADMET property classification prediction model for the compounds. The model was used to make corresponding predictions for 50 compounds, and the final prediction results were obtained.

[0039] In step (4), in order to find the optimal value corresponding to the compound molecule descriptor so that the compound antagonizes ERα activity while possessing ADMET properties, this invention utilizes Adaptive Differential Evolution (ADE) to handle the constrained optimization problem and find the optimal value corresponding to the compound molecule descriptor. The method process is as follows:

[0040] The compound bioactivity pIC based on AFNN given in step (2) 50 The prediction model and step (3) use the ADMET property classification model based on probabilistic neural network (PNN) to construct the compound fitness function; the bioactivity value pIC50 obtained by the compound bioactivity pIC50 prediction model based on AFNN is assumed to be ym; the property classification values ​​obtained by constructing the ADMET property classification model of compound using probabilistic neural network are summed, and the sum is assumed to be yc; the fitness function y of compound sample is constructed: y = ym + yc; the adaptive differential evolution algorithm (ADE) is used to optimize the fitness.

[0041] The optimization process using the Adaptive Differential Evolution (ADE) algorithm is as follows:

[0042] First, the parameters of ADE are set as follows: scaling factor F is 0.3, maximum number of iterations is 500, population size is 30, crossover probability is 0.5, and individual vector dimension is 20. Next, the population is initialized, followed by a mutation, crossover, and selection loop until the termination condition (the set number of iterations) is met. The optimal solution is then output, representing the desired value of the molecular descriptor that has the greatest impact on the compound, enabling the compound to inhibit ERα while also exhibiting ADMET properties.

[0043] Traditional differential evolution (DE) uses a fixed scaling factor F during the fitness optimization process, which is unsuitable for practical problems. If the scaling factor is too large, local optima may occur; if it is too small, population diversity may decrease, leading to premature convergence. To address this issue, the scaling factor F is improved as follows:

[0044] F = F0·2 Q (5)

[0045]

[0046] Where F0 is the scaling factor, G m Let F be the maximum number of iterations and G be the current generation. In the early stages of fitness optimization, F = 2F0, and the scaling factor is relatively large, ensuring the diversity of the population. As the number of iterations increases, the value of F gradually decreases until it approaches F0, which is beneficial for obtaining the optimal solution.

[0047] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0048] (1) This invention relates to the bioactivity pIC of compounds based on AFNN. 50 In the process of modeling the prediction model, the adaptive gradient descent algorithm is used to optimize the parameters of the prediction model in order to improve the prediction effect.

[0049] (2) In the optimization process using the Adaptive Differential Evolution (ADE) algorithm, the traditional differential evolution algorithm has a fixed scaling factor F value during the fitness optimization process, which cannot adapt to the actual problems being dealt with. If the scaling factor is too large, local optima are likely to occur; if the scaling factor is too small, the diversity of the population itself will decrease, and premature convergence may occur. This invention improves the scaling factor F to ensure the diversity of the population itself. As the number of iterations increases, the value of F gradually decreases until it approaches F0, which is beneficial for obtaining the optimal solution. Attached Figure Description

[0050] Figure 1 This is a flowchart of the anti-breast cancer candidate drug classification and prediction method based on AFNN and ADE of the present invention;

[0051] Figure 2 This is the fitting curve between the true value and the predicted value of the verification sample in this invention;

[0052] Figure 3 The prediction error curve of the verification sample for this invention;

[0053] Figure 4 The image shows the prediction results of the probabilistic neural network (PNN) model trained for Caco-2 classification in this invention.

[0054] Figure 5 The prediction error diagram after training the PNN model for Caco-2 classification in this invention;

[0055] Figure 6 This is the fitness curve of the compound using the Adaptive Differential Evolutionary Algorithm (ADE) of this invention. Detailed Implementation

[0056] like Figure 1 As shown, this invention selects molecular descriptors that enhance the bioactivity of compounds in inhibiting ERα and exhibit better ADMET properties. For accuracy, variables are sorted, and the 20 molecular descriptors with the greatest impact on bioactivity are selected. Then, an AFNN-based pIC (bioactivity index) model of compound bioactivity is established using these selected molecular descriptors. 50 Predictive models and ADMET property classification models based on probabilistic neural networks.

[0057] The present invention provides a method for classifying and predicting candidate drugs for breast cancer based on AFNN and ADE, comprising the following steps:

[0058] (1) Main feature extraction: The Random Forest (RF) algorithm was used to sort the importance of 729 molecular descriptors in descending order, i.e., the average Gini index, and the top 20 with the most significant impact on biological activity were selected as the main variables.

[0059] The specific implementation method of the Random Forest (RF) algorithm is as follows: Multiple decision trees are constructed to determine the initial drug characteristic data for anti-breast cancer candidate drugs. In the Random Forest RF classifier, the input data consists of 729 molecular descriptors and label data (i.e., biological activity) for 1947 compounds. The output data consists of the 20 molecular descriptors that have the most significant influence on biological activity. First, the input data is standardized to initialize the random forest. The random forest is then trained using the label data, and feature importance and feature importance ranking index are calculated. N average feature importance and N average feature importance ranking index are calculated. If the average feature importance ranking index remains unchanged after three rounds, the output feature importance ranking index is returned. Features are selected based on the feature importance ranking index, i.e., the 20 molecular descriptors with the most significant influence on biological activity are output.

[0060] (2) Considering the high computational accuracy of the adaptive fuzzy neural network (AFNN), a compound bioactivity pIC based on AFNN was established using 20 selected molecular descriptors. 50 Predictive model; using samples 1–1774 from 1974 original samples, an AFNN-based pIC of compound bioactivity was trained. 50 A prediction model was developed; its performance was then validated using 30 random samples. The mean squared error analysis was used to assess the model's fit and verify its rationality.

[0061] Constructing AFNN-based compound bioactivity pIC 50 The predictive model includes the following steps:

[0062] (1) Data processing

[0063] We screened 729 molecular descriptors corresponding to 50 compounds, and normalized the selected 20 molecular descriptors. Each row was used as a training set, and the AFNN bioactivity prediction model was trained using samples 1 to 1774 of the 1974 data sets. The model performance was then validated using 30 random samples.

[0064] (2) Designing the network layer structure: This invention selects a Mamdani-type fuzzy neural network for modeling.

[0065] In AFNN-based compounds, bioactivity pIC 50 In the modeling process of the prediction model, gradient descent algorithm with adaptive learning rate is used to optimize the model parameters to improve the model's prediction performance. The update formulas for the center, width, and weights of the membership function of AFNN are as follows:

[0066]

[0067]

[0068]

[0069] Among them, c ij σ is the center of the membership function. ij η is the width of the membership function, w is the connection weight of the system, η is the learning rate, and K is the system training error.

[0070] Because traditional FNNs have a fixed learning rate, they are difficult to adapt to the modeling needs of complex systems. This invention employs an adaptive learning rate adjustment strategy as shown in the following formula:

[0071]

[0072] Among them, K t Let K be the system training error at time t. t-1 Let be the training error at time (t-1). When the error moves in the direction of decreasing, the learning rate increases; when the error moves in the direction of increasing, the learning rate decreases.

[0073] In step (2), the compound undergoes pIC. 50 The prediction method for the values ​​is as follows: First, 20 molecular descriptors selected using the Random Forest method are normalized; each row is used as an input training set, and samples 1 to 1774 from the 1974 data sets are used as the training set; 30 random samples are used as the test set to verify the model performance. The parameters for RF are: 500 decision trees; the parameters for AFNN are: 30 rules, a maximum training step size of 1000 steps, and a learning rate of 0.01.

[0074] The validation samples showed an RMSE of 0.2914 and a MAPE of 4.14%, indicating that the model accuracy meets the requirements for modeling the activity of anti-breast cancer drug candidates. The final output obtained is pIC. 50 The predictive model for the values ​​was used to perform pIC analysis on 50 compounds. 50 Prediction of values.

[0075] In step (3), an ADMET property classification model based on PNN was established using 20 selected molecular descriptors; the ADMET property classification model was trained using samples 1 to 1774 from 1974 original samples; and the model performance was verified using 30 random samples.

[0076] The method for establishing the ADMET property classification model is as follows:

[0077] The 20 molecular descriptors selected by the Random Forest (RF) method in step (1) were normalized. Each row was used as a training set, and samples 1 to 1774 of the 1974 data sets were used to train an ADMET property classification model based on Probabilistic Neural Network (PNN). The model performance was then validated using 30 random samples. Based on the prediction results, a PNN model was finally selected to construct an ADMET property classification prediction model for the compounds. The model was used to make predictions for 50 compounds, and the final prediction results were obtained.

[0078] In step (4), in order to quickly find the optimal value corresponding to the compound molecule descriptor, so that the compound has better antagonistic effect on ERα activity and better ADMET properties, this invention uses the Adaptive Differential Evolutionary Algorithm (ADE) to process the constrained optimization problem and find the optimal value corresponding to the compound molecule descriptor. The specific method is as follows:

[0079] The compound bioactivity pIC based on AFNN given in step (2) 50 The prediction model and the probabilistic neural network given in step (3) are used to construct a classification model of the compound's ADMET properties to build the compound fitness function; the compound bioactivity pIC based on AFNN is used. 50 Bioactivity value pIC obtained from the prediction model 50 Let ym be the property classification value obtained by the ADMET property classification model based on probabilistic neural network (PNN). The sum of these values ​​is assumed to be yc. Construct the fitness function y for the compound samples: y = ym + yc. Use the ADE algorithm to optimize the fitness.

[0080] The optimization process of the Adaptive Differential Evolutionary Algorithm (ADE) is as follows:

[0081] First, the parameters of the ADE are set as follows: scaling factor F is 0.3, maximum number of iterations is 500, population size is 30, crossover probability is 0.5, and individual vector dimension is 20. Next, the population is initialized, followed by a mutation, crossover, and selection loop until the termination condition (the set number of iterations) is met. The optimal solution is then output, representing the desired value of the molecular descriptor that has the greatest impact on the compound, resulting in better inhibition of ERα and improved ADMET properties.

[0082] Traditional differential evolution, with its fixed scaling factor F during fitness optimization, is ill-suited for practical problems. If the scaling factor is too large, local optima emerge; if it is too small, population diversity decreases, leading to premature convergence. To address this issue, this invention improves the scaling factor F as follows:

[0083] F = F0·2Q (5)

[0084]

[0085] Where F0 is the scaling factor, G m Let F be the maximum number of iterations and G be the current generation. In the early stages of fitness optimization, F = 2F0, and the scaling factor is relatively large, ensuring the diversity of the population. As the number of iterations increases, the value of F gradually decreases until it approaches F0, which is beneficial for obtaining the optimal solution.

[0086] In this embodiment, the RF algorithm sorted the 729 molecular descriptors in descending order of importance (mean decreasing Gini index), and selected the top 20 descriptors that had the most significant impact on biological activity as the primary variables. The screening results of the molecular descriptors are shown in Table 1.

[0087] Table 1. Molecular descriptors with the most significant impact

[0088]

[0089] In establishing the bioactivity pIC of AFNN compounds 50 The prediction model was trained using samples 1 to 1774. As the number of training steps increased, the RMSE value decreased, and the model's training accuracy improved. The model performance was validated using 30 randomly selected samples. The true values ​​and model predictions of the validation samples were compared as follows: Figure 2 As shown, the prediction error is as follows Figure 3 As shown, the present invention utilizes AFNN-based compound bioactivity pIC 50 The prediction model exhibits high accuracy. Experimental results show that the RMSE for the training samples is 0.5974 and the MAPE is 7.27%, while the RMSE for the validation samples is 0.2914 and the MAPE is 4.14%. The model accuracy meets the requirements for modeling the activity of anti-breast cancer drug candidates.

[0090] Ten sets of test sample data were randomly selected, and the compound IC 50 and pIC 50 The predicted values ​​are shown in Table 2.

[0091] Table 2 Compound IC 50 and pIC 50 Predicted value

[0092] Compound numbering <![CDATA[IC 50 / nM]]> <![CDATA[pIC 50 / nM]]> 1 52.751 7.278 2 26.64 7.574 3 29.932 7.524 4 24.486 7.611 5 24.968 7.603 6 41.965 7.377 7 23.343 7.632 8 27.329 7.563 9 19.151 7.718 10 77.822 7.109

[0093] The prediction results of the Caco-2 validation samples using the ADMET property classification model based on a probabilistic neural network (PNN) are shown in the figure below. Figure 4 As shown, the prediction error diagram is as follows: Figure 5 As shown.

[0094] Experiments show that the ADMET property classification model based on a probabilistic neural network (PNN) trained using samples 1-1774 for Caco-2 classification, and validated with 30 random samples, shows that the error between the predicted and actual values ​​is 0 for most samples. This indicates that the PNN classification results are consistent with the true classification, demonstrating the high classification accuracy of the PNN model. The accuracy of the ADMET five property classification models is shown in Table 3.

[0095] Table 3. Accuracy of ADMET Five-Property Classification Model

[0096]

[0097] Experimental data show that, using sample data from 1775 to 1794 to verify the model accuracy, the five classification models have strong classification ability and high accuracy.

[0098] Ten sets of test sample data were randomly selected, and their five properties of ADMET were classified. The classification results are shown in Table 4.

[0099] Table 4. ADMET Classification of Compound Test Samples

[0100] serial number Caco-2 CYP3A4 hERG HOB MN 1 2 2 1 1 2 2 2 2 1 1 2 3 1 2 1 1 2 4 1 2 2 1 2 5 1 2 2 1 2 6 1 2 2 2 2 7 1 2 2 2 2 8 1 2 2 2 2 9 1 2 2 2 2 10 1 2 2 2 2

[0101] The fitness change curve of the ADE algorithm used to handle the constrained optimization problem is shown in the figure. Figure 6 As shown in Table 5, after 500 iterations, the ADE algorithm finds the optimal solution to the constrained optimization problem. The values ​​of the compound score descriptors corresponding to the optimal solution at this point are shown in Table 5.

[0102] Table 5 shows the optimal values ​​for compound molecule descriptors.

[0103] variable optimal value variable optimal value <![CDATA[X1]]> 35.937 <![CDATA[X 11 ]]> 2.748 <![CDATA[X2]]> 15.492 <![CDATA[X 12 ]]> 7.710 <![CDATA[X3]]> 3.921 <![CDATA[X 13 ]]> 631.7 <![CDATA[X4]]> 87.785 <![CDATA[X 14 ]]> 0.212 <![CDATA[X5]]> 0.393 <![CDATA[X 15 ]]> 10.64 <![CDATA[X6]]> 0.722 <![CDATA[X 16 ]]> 7.691 <![CDATA[X7]]> -0.273 <![CDATA[X 17 ]]> 13.37 <![CDATA[X8]]> 5.377 <![CDATA[X 18 ]]> 425.9 <![CDATA[X9]]> 1.427 <![CDATA[X 19 ]]> 101.3 <![CDATA[X 10 ]]> 4.672 <![CDATA[X 20 ]]> -1.479

Claims

1. A method for classifying and predicting candidate drugs for breast cancer based on AFNN and ADE, characterized in that: Includes the following steps: (1) The Random Forest (RF) algorithm is used to select features from the known experimental data to identify influential molecular descriptors; the process is as follows: In step (1), the data is first standardized, the random forest is initialized, the random forest is trained using the labeled data, and the feature importance and feature importance ranking index are calculated; the Nth average feature importance and Nth average feature importance ranking index are calculated; and features are selected according to the feature importance ranking index. If the average feature importance sorting index has not changed over multiple rounds, then return the feature importance sorting index; (2) Constructing AFNN-based compound bioactivity pIC 50 Predictive models utilize adaptive gradient descent algorithms to predict the bioactivity pIC of compounds based on AFNN. 50 The parameters of the prediction model are optimized. AFNN-based compound bioactivity pIC 50 The establishment of a predictive model includes the following steps: (2.1) Data processing: Normalize the molecular descriptors corresponding to the compounds and use each row as a set of input training data; (2.2) The network layer structure was designed, and a Mamdani-type fuzzy neural network was used for modeling. Sample training was conducted to model the bioactivity of compounds based on AFNN. 50 Predictive models; (2.3) Root mean square error and mean absolute percentage error are selected as performance indicators; (3) Establish an ADMET property classification model based on a probabilistic neural network; train the ADMET property classification model using samples, and then verify the performance of the ADMET property classification model using random samples; (4) The constrained optimization problem is handled using the Adaptive Differential Evolutionary Algorithm (ADE) to find the optimal value corresponding to the compound molecule descriptor; the process is as follows: The bioactivity value pIC50 obtained by the compound bioactivity pIC50 prediction model based on AFNN is assumed to be ym; the property classification values ​​obtained by constructing a classification model of compound ADMET properties using a probabilistic neural network are summed, and the sum is assumed to be yc; the fitness function y of the compound samples is constructed: y = ym + yc; the fitness of the compound samples is optimized using the adaptive differential evolution algorithm (ADE); first, the parameters of ADE are set, then the population is initialized, and then mutation, crossover, and selection are performed in a loop until the termination condition is met, i.e., the set number of iterations, and the optimal solution is output. When using the adaptive differential evolution algorithm to optimize the fitness of compound samples, the scaling factor F is improved as follows: F=F0·2 Q (5) Where F0 is the scaling factor, G m G represents the maximum number of iterations, and G represents the current generation.

Citation Information

Patent Citations

  • Method, system and terminal for screening anti-breast cancer candidate drug molecule descriptors

    CN114334033A

  • Anti-breast cancer candidate drug screening method, device and equipment and storage medium

    CN114613512A

  • Single agent anti-pd-l1 and pd-l2 dual binding antibodies and methods of use

    CN104936982A

  • Method for constructing effect prediction model of anti-breast cancer candidate drug

    CN115831375A