A model robustness-based quantitative structure-activity relationship model construction method

By optimizing the modeling data and molecular descriptors of the QSAR local model using SVR and PSO algorithms, and combining them with a consistency modeling method, the robustness and predictive performance issues of the QSAR model in predicting liver metabolic clearance rate were resolved, achieving broader and more reliable predictive results.

CN115527619BActive Publication Date: 2025-12-12UNIV OF SCI & TECH LIAONING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211121917.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-15
Publication Date
2025-12-12
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

Existing QSAR models suffer from poor robustness and difficulty in improving predictive performance in predicting liver metabolic clearance. Local models based on the mechanism of action are difficult to apply, while local models based on chemical structure similarity lack robustness guarantees.

Method used

We employ Support Vector Machine Regression (SVR) combined with Particle Swarm Optimization (PSO) to optimize modeling data and molecular descriptors, constructing a robust QSAR local model. Furthermore, we combine the local models using a consistency modeling method to form a consistent model, thereby improving prediction performance and expanding the application scope.

Benefits of technology

This improves the robustness and predictive performance of the QSAR model, expands its application scope, and provides a more reliable tool for predicting liver metabolic clearance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115527619B_ABST
    Figure CN115527619B_ABST
Patent Text Reader

Abstract

A quantitative structure-activity relationship model construction method based on model robustness relates to a quantitative structure-activity relationship model construction method, in particular to a quantitative structure-activity relationship local model construction method based on model robustness and a consistency model construction method of the quantitative structure-activity relationship local model, taking a compound liver metabolic clearance rate prediction model as an object, belonging to the field of chemical information science and bioinformatics. In view of the fact that the current global modeling method is difficult to cope with the complex modeling compounds of the liver metabolic clearance rate prediction model, the model robustness is poor, and the prediction performance is difficult to further improve, the QSAR local model based on the activity mechanism is difficult to be applied in practice, and the QSAR local model based on the chemical structure similarity lacks the technical means to guarantee the model robustness, the application provides a QSAR local model construction method based on model robustness, and a support vector machine regression technology is used to establish a QSAR local model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a method for constructing a quantitative structure-activity relationship (QSAR) model, in particular to a method for constructing a QSAR local model and a consistency model based on model robustness, with a compound liver metabolic clearance rate prediction model as the object, and belongs to the field of chemical informatics and bioinformatics. TECHNICAL BACKGROUND

[0002] Metabolic clearance rate refers to how many units of volume of body fluid can be cleared by the body in a unit of time, and is one of the important parameters reflecting the elimination of drugs from the body, and is also a very important pharmacokinetic parameter in drug research. Predicting the in vivo metabolic clearance rate of a drug through in vitro data, and further calculating the bioavailability and half-life of the drug, has very important reference value for determining the first dose and frequency of drug administration in clinical practice. The liver is the main metabolic organ of drugs, and the liver metabolic clearance rate model has been successfully used in the study of in vivo liver metabolism of drugs. Liver metabolic clearance rate models based on rat liver, mouse liver and human liver microsomes have been reported. Establishing an accurate liver metabolic clearance rate prediction model is of great significance in drug design, pharmacodynamics and drug safety.

[0003] The metabolic clearance rate model essentially studies the relationship between the microstructure of a drug molecule (compound) and its metabolic clearance rate (macro activity of the compound). QSAR technology can construct a relationship model between the activity and structure of a compound by using mathematical statistics, machine learning and other methods based on the relationship between the microstructure parameters (such as molecular descriptors, molecular fingerprints, etc.) of known compounds and their macro activity. Without biochemical experiments, QSAR technology can use computer models to quickly evaluate the biological activity of compounds in the virtual design stage or newly synthesized compounds, and can reveal the internal relationship between the microstructure of a compound and its macro activity from the molecular level, and can provide theoretical guidance for designing compounds with the desired properties of chemists. At present, QSAR technology plays an increasingly important role in computer-aided drug design, effectively reducing the time and money consumption caused by a large number of animal experiments and preclinical experiments in the drug development process, and improving the efficiency and saving the development cost of drug development.

[0004] In the prediction of drug liver clearance rate by QSAR technology, with the development of time, more and more experimental data are accumulated, and the data used for QSAR modeling research are more and more abundant. Although the abundance of data enables researchers to establish a QSAR model containing more comprehensive information, the applicable scope of the model will also be larger, but the chemical mechanism of the compounds with similar liver metabolic clearance rate characteristics is not completely the same. In the face of a large amount of modeling data, it has become more and more difficult to establish a liver metabolic clearance rate QSAR model with strong robustness, good prediction performance and wide application scope.

[0005] According to the OECD principles (the basic principles that need to be followed in QSAR modeling established by the International Organization for Economic Cooperation and Development), appropriate measures of goodness-of-fit, robustness and predictivity are important guarantees for the effectiveness of the model. Only the QSAR model that passes the internal and external validation can be applied in practice.

[0006] For internal validation of the model, a useful QSAR model generally requires that the fitting coefficient R 2 >0.6, the cross-validation coefficient R 2 cv >0.5; for external validation of the model, it is generally required that the validation coefficient R T 2 >0.6. And the good robustness of the model (the greater the value of R 2 cv , the closer to 1, the better the robustness of the model) is a necessary condition for establishing a standardized and good performance QSAR model. Increasing the R 2 cv value of the model is conducive to establishing a more reliable QSAR model.

[0007] Currently, there are three main ways to establish QSAR model: global modeling (for a modeling dataset, the established QSAR model uses all compounds in the dataset, and only one model is established), local modeling (for a modeling dataset, several local models are established by using part of the modeling data), and consistency modeling (several QSAR sub-models are combined to form a consistency model, and the output of the consistency model is obtained by combining the outputs of each sub-model). Local modeling only uses part of the modeling data to establish local models, which can overcome the shortcomings of global modeling methods by using certain data selection methods, and therefore has received widespread attention. At present, local modeling methods based on activity mechanism and local modeling methods based on chemical structure similarity (such as k-nearest neighbor model (kNN), lazy learning method, and ALL-QSAR) have been gradually developed.

[0008] Although these local modeling methods can improve the prediction performance of the model to some extent, there are still certain limitations. Local models based on activity mechanism need to be divided into several groups according to the mechanism of the compounds producing such activity before being established, and then local models are established for each group. However, due to the complexity of chemical mechanism, the mechanism of the compounds produced by the new synthesis or virtual design stage may not be clear, so it is difficult to select a suitable local model for predicting the activity of the predicted compounds when the model is applied, which will lead to difficulties in model application. Local models based on chemical structure similarity divide the modeling dataset into several groups with similar chemical structures before the model is established, and then local models are established for each group. The similarity of the chemical structure of the compounds in the local model is measured by the distance information (such as Euclidean distance, Mahalanobis distance, and geodesic distance) between the molecular descriptors representing the structure of the compounds. Different molecular descriptors and different similarity measurement methods will obtain different local modeling datasets, and thus obtain different local models. Ensuring that the cross-validation coefficient R 2 cv >0.5 is a prerequisite for QSAR model to pass the effectiveness verification, and the data division method based on the similarity of the chemical structure of the compounds used in the current local modeling technology is essentially an unsupervised data division method, and there is no effective measure to ensure that the cross-validation coefficient R 2 cv >0.5.

[0009] Consistency modeling is to combine sub-models with different structures together, and its main idea is to combine several weak learners into a strong learner, so as to improve the performance of the model. The consistency model can more comprehensively describe the characteristics of the compounds in the whole data set, can reduce the influence of the compounds with large structural differences in the modeling data set on the model, can obtain more comprehensive molecular structure information of the whole modeling data set, has stronger robustness, better prediction ability and wider application range, and can effectively improve the prediction accuracy and reliability of the model.

[0010] A QSAR model with strong robustness, good prediction performance and wide application range is the basis for its reliable prediction of liver metabolic clearance rate. The local modeling method based on model robustness proposed in the present application selects appropriate modeling data and descriptors in the modeling data set and appropriate model hyperparameters through an optimization algorithm to construct a QSAR local model with high cross-validation coefficient, which can ensure the effectiveness of the local model and promote the improvement of the model performance. The combination of the local models constructed under the restriction of the application domain of each local model into a consistency model can further improve the performance of the model and expand its application range. SUMMARY

[0011] The present application provides a method for constructing a QSAR local model based on model robustness and establishing a consistency model for predicting the liver metabolic clearance rate of compounds.

[0012] The technical problems to be solved by the present application are:

[0013] In view of the problems that the global modeling method is difficult to cope with the complex modeling compounds of the liver metabolic clearance rate prediction model, resulting in poor model robustness and difficult to further improve the prediction performance, the QSAR local model based on the activity mechanism is difficult to be applied in practice, and the QSAR local model based on the similarity of chemical structure lacks technical means to ensure the robustness of the model, the present application proposes a QSAR local model construction method based on model robustness to improve the robustness of the local model and ensure the effectiveness of the model; a local model consistency modeling method based on the restriction of the application domain is proposed to improve the prediction performance of the model and expand the application range of the model.

[0014] The technical scheme of the present application is as follows:

[0015] The present application uses support vector machine regression (support vector regression, hereinafter referred to as: SVR) to establish a QSAR local model, and uses particle swarm optimization (Particle swarm optimization, hereinafter referred to as: PSO) algorithm to select the five-fold cross-validation coefficient (R2 cv5 The reciprocal of the inverse of the function f(x) = 1 / x takes the minimum value as the optimization objective. The joint selection of the compounds in the modeling dataset, the molecular descriptors representing the structure parameters of the compounds and the hyperparameters of the SVR model is performed so as to obtain a suitable modeling data subset and molecular descriptors under suitable model parameters, thereby obtaining a QSAR local model with strong robustness and good reliability. The obtained local models are used as submodels, and a QSAR consistency model is constructed under the limitation of the application domain of each local model, so as to further improve the prediction performance of the model and expand the application range thereof.

[0016] The construction of the QSAR local model and the consistency model thereof and the verification of the model performance include the following steps:

[0017] Step (1): Selecting a compound set for establishing a QSAR model of liver metabolic clearance rate, calculating the descriptors of the compounds by using descriptor calculation software, and constructing a modeling dataset.

[0018] Step (2): Descriptor cleaning;

[0019] Step (3): Splitting of the modeling dataset;

[0020] Step (4): Selection of key descriptors;

[0021] Step (5): Construction of a local model based on the robustness of the model;

[0022] Step (6): Analysis of the application domain of each local model;

[0023] Step (7): Effectiveness verification of each local model, and calculation of the coverage rate of each local model passing the effectiveness verification (i.e., the ratio of the data within the application domain of each local model to the test dataset for the test dataset);

[0024] Step (8): Construction of a consistency model;

[0025] Step (9): Evaluation of the prediction performance of the consistency model, and comparison of the consistency model with the local model and the global model.

[0026] In the step (1), for the selected modeling compound set A, the molecular descriptors of each compound in the set A are calculated by using molecular descriptor calculation software, and the structure of the compound is numerically represented; then, the molecular descriptor values of each compound are one-to-one corresponding to the liver metabolic clearance rate values of the compounds to construct an initial modeling dataset B.

[0027] The descriptor cleaning of step (2) is to simplify the model structure and remove the descriptors with null values, redundant descriptors and unimportant descriptors in the modeling data set B; the specific cleaning method of the descriptors is as follows: operating on the data set B, firstly, all the descriptors with constant values or null values are deleted; at the same time, if it is found that there are two descriptors that are pair-correlated (correlation coefficient > 0.85), the descriptor with smaller correlation with the liver metabolic clearance rate value is deleted; finally, the simplified modeling data set C is obtained.

[0028] The step (3) completes the segmentation of the modeling data set; in order to verify the model performance, the modeling data set C needs to be divided into independent initial training data set D and initial test data set E; the specific segmentation method is as follows: randomly selecting 80% of the data in the data set C as the initial training data set D, and the remaining 20% of the data as the initial test data set E.

[0029] The step (4) is the selection of key descriptors, which is as follows: firstly, the mean decrease in impurity (MDI) method in the literature (Wang Y K, Chen X B. A joint optimization QSAR model of fathead minnow acute toxicity based on a radial basis function neural network and its consensus modeling [J]. RSC Advances, 2020, 10, 21292) is used to sort the descriptors in the initial training data set D according to the importance, and then the non-important descriptors with score value less than 0.01 are removed from the initial training data set D and the initial test data set E, respectively, to form the final training data set D1 and the test data set E1; the data set D1 is used to construct the QSAR model, and the data set E1 is used to verify the performance of the built model.

[0030] The step (5) needs to establish multiple local models, in order to obtain structure-differentiated local models, 80% of the data in the modeling data set D1 is randomly selected before each local model is established, to form the local modeling data set D2.

[0031] The step (5) establishes local models through SVR technology, and the particle swarm (PSO) algorithm is used to select appropriate data and descriptors in the model training data set D2, and the hyperparameters of the SVR model are optimized. When optimizing, the five-fold cross-validation coefficient (R 2 cv5) is taken as the optimization target to obtain optimized model hyperparameters, appropriate data subsets and descriptors, so as to improve the robustness of the model. Joint selection of modeling data, molecular descriptors and model hyperparameters; selection of an optimized data subset D 3opt and determination of the hyperparameters of the SVR model to ensure that the established local model can obtain good robustness.

[0032] The SVR technology used in the step (5) is a mature modeling tool in the field of machine learning. The nu_SVR type SVR algorithm is used in the present application, and the three hyperparameters affecting the performance of the algorithm are a penalty factor C, a regulation parameter nu and a radial

[0033] width of a base function.

[0034] The five-fold cross-validation coefficient of the model calculated in the step (5) is a mature model robustness determination method in statistics. The greater the value of the five-fold cross-validation coefficient of the model, the stronger the robustness of the model. The five-fold cross-validation coefficient R 2 cv5 The expression of the five-fold cross-validation coefficient R

[0035]

[0036] where y cv (i) is the actual activity value of the i-th cross-validation compound in the five-fold cross-validation process of the SVR model, is the model predicted activity value of the i-th cross-validation compound. is the average value of the actual activities of all compounds participating in the cross-validation, N cv is the number of compounds participating in the cross-validation.

[0037] The PSO algorithm used for model robustness optimization in the step (5) is a mature population optimization algorithm in the field of intelligent optimization. The parameters that need to be set before the PSO algorithm is used are: population size M, particle dimension L, maximum number of iterations T of the algorithm, learning factors c1 and c2, and inertia weight ω.

[0038] The particle coding method of the PSO algorithm used in the step (5) is a real number and binary hybrid coding to adapt to the joint selection of the modeling data, the molecular descriptors and the model parameters. For a data set D2 containing K pieces of modeling data (i.e. K compounds) and V descriptors, the dimension of the particle of the PSO algorithm is L (L = L1 + L2 + L3). Herein, L1 = ceil (K / 10), L2 = ceil (V / 10), and L3 is the number of the to-be-determined hyperparameters of the SVR model. The ceil () represents the rounding up of an integer. For the PSO algorithm, the present application needs to define an initial population containing M particles, each of which has a dimension of L. The first L1 dimensions of each particle in the population are used to encode the to-be-selected modeling data of the local model in the data set D2, and the value range of each particle is a real number between 0 and 1023. Each dimension can encode 10 to-be-selected data, which is converted into a ten-bit binary number after rounding down. The value 0 (binary representation '0000000000') indicates that none of the 10 data is selected by the local model, and 333 (binary representation '0101001101') indicates that the 2nd, 4th, 7th, 8th and 10th data are selected by the local model. The middle L2 dimensions of each particle in the population are used to encode the to-be-selected molecular descriptors of the local model in the data set D2, and the value range of each particle is also a real number between 0 and 1023. Each variable can also encode 10 to-be-selected descriptors. The last L3 dimensions of each particle in the population are used to encode the hyperparameters of the SVR model for constructing the local model, and the value range is also selected as a real number between 0 and 1023.

[0039] During the establishment of the local model in the step (5), the operation method of each particle in the population of the PSO algorithm is as follows:

[0040] 1) Take the first L1 dimensions of the particle, round down the integer first, then decode each dimension into binary, and select the modeling data in the data set D2 according to the decoding result.

[0041] 2) Take the middle L2 dimensions of the particle, round the integer first, then decode each dimension into binary, and select the molecular descriptors in the data set D2 according to the decoding result.

[0042] 3) Take the last L3 dimensions of the particle, divide each dimension by 100 to obtain the hyperparameters of the SVR model.

[0043] 4) After the selection of the data and the descriptors, a data subset D3 containing K1 (K1 ≤ K) pieces of modeling data and V1 (V1 ≤ V) descriptors is obtained, and then the SVR model is constructed by using the data subset D3.

[0044] The step (5) is to optimize the selected modeling data, descriptors and SVR model hyperparameters of the local model by PSO algorithm to improve the robustness of the model, and the steps of establishing a local model are as follows:

[0045] (a) Set the initial parameters c1, c2 and ω of PSO algorithm, and randomly generate an initial particle group containing M particles (each particle contains L-dimensional variables).

[0046] (b) For each particle, obtain a data subset D3 for establishing a local model from the data set D2 by binary decoding, and obtain 3 hyperparameters of the SVR model by decoding. Construct the SVR model (the input variables of the SVR model are the descriptors in the data subset D3, and the output variables are the liver metabolic clearance rate values of the compounds in the data subset D3) with the obtained data subset D3 and the hyperparameters of the SVR model, and calculate the five-fold cross-validation coefficient to obtain the fitness value of each particle.

[0047] (c) After the fitness evaluation of the QSAR local model established based on the modeling data subset and the model hyperparameters obtained by decoding all particles is completed, record the individual optimal solution, global optimal solution and current optimal fitness value of the PSO algorithm, then update the particle group according to the PSO algorithm formula, return to (b) to search for better particles, and continue until the algorithm reaches the maximum iteration number T.

[0048] (d) After the PSO algorithm search is completed, output the data subset D 3opt (the data and descriptors are optimized and selected) and the model hyperparameters corresponding to the minimum fitness value according to the decoding results of the optimal particle. 3opt and the model hyperparameters obtained by the PSO algorithm to construct a local QSAR model.

[0049] A plurality of local models are established by the above method until the data contained in these local models can cover more than 95% of the data in the training data set D1;

[0050] The step (6) completes the model application domain analysis of the plurality of local models established in step (5). The application domain of each sub-model is calculated by the IWD method of the literature (Wang Y, Chen X. QSPR model for Caco-2 cell permeability prediction using a combination of HQPSO and dual-RBF neural network [J]. RSC Advances, 2020, 10: 42938-42952) to determine the reliable application range of each local model.

[0051] The step (7) completes the validation of the effectiveness of each local model. It is required to calculate the five-fold cross-validation coefficient R 2 cv5 of each local model according to formula (1); to calculate the fitting coefficient R 3opt of each local model to the data set D 2 according to formula (2); and to calculate the prediction performance index R 2 T of each local model to the test data set E1 within the application of its model according to formula (3).

[0052] If a local model simultaneously satisfies R 2 > 0.6, R 2 cv5 > 0.5 and R 2 T > 0.6, it is considered to be effective.

[0053]

[0054] where y t (i) is the actual activity value of the i-th compound in the training data, is the model predicted output value of the i-th compound in the training data. is the average value of the activity values of all compounds in the training data, N L is the number of compounds contained in the model training data.

[0055]

[0056] where y t (i) is the actual activity value of the i-th compound in the test data, is the model predicted output value of the i-th compound in the test data. is the average value of the actual activity values of all test compounds in the test data, N t is the number of compounds contained in the test data.

[0057] The step (8) completes the consistency model construction. The constructed local model which passes the validity verification is used as a sub-model of the consistency model, and is combined in parallel to form the consistency model. For a compound to be predicted a, firstly, the application domain of a is determined according to the application domain of each local model determined in step (6), and it is judged which local model the application domain of a falls into. If a falls into the application domain of a local model, it is considered that the prediction result of the local model for a is reliable, otherwise, it is unreliable. Finally, the reliable prediction results of all local models for a are averaged to obtain the final prediction result of the consistency model for a. If a cannot fall into the application domain of any local model, the consistency model cannot reliably predict the compound.

[0058] The method for evaluating the prediction performance of the consistency model in step (9) is as follows: the consistency model established in step (8) is used to predict the metabolic clearance rate values of the compounds in the test data set E1, and the prediction performance of the consistency model for the reliable prediction results is evaluated by formula (2).

[0059] The comparative analysis of the model in step (9) mainly compares the prediction performance and model coverage range of the consistency model and each local model, and compares the consistency model and the traditional global model, and draws a conclusion whether the modeling method proposed in the application can improve the performance of the model.

[0060] Advantages:

[0061] The application uses the compound data with known liver metabolic clearance rate, adopts the local modeling method based on the robustness of the model to construct the QSAR local model, can effectively improve the robustness of each local model, and ensures the effectiveness of each local model. Compared with each local model, the construction of the consistency model can further improve the overall prediction ability of the model and obtain a wider model application range, and provides an effective tool for better and more reliable prediction of the liver metabolic clearance rate of the candidate drug. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 is a flowchart of the implementation process of the application. Figure 1 Figure 2 is a flowchart of the implementation process of the application.

[0063] Figure 3 is a schematic diagram of the PSO algorithm based on the robustness of the model for optimizing the selection of modeling data, descriptors and SVR hypermodel parameters when constructing a local model. Figure 2 Figure 4 is a flowchart of the PSO algorithm for optimizing the local model.

[0064] Figure 3 Figure 5 is a flowchart of the PSO algorithm for optimizing the local model.

[0065] Figure 6 is a flowchart of the PSO algorithm for optimizing the local model. Figure 4 Figure 7 is an implementation flowchart of the consistency model for predicting the liver metabolic clearance rate of a compound.​ DETAILED DESCRIPTION

[0066] The following examples will facilitate further understanding of the present application by those of ordinary skill in the art, but do not limit the protection scope of the present application in any form.

[0067] The present embodiment uses human liver microsomal metabolic clearance rate data to establish a QSAR prediction model of liver metabolic clearance rate. The human liver microsomal metabolic clearance rate of the compounds used in the model is derived from published literature (Wenzel J, Matter H, Schmidt F. Predictive Multitask Deep Neural Network Models for ADME-Tox Properties: Learning from Large Data Sets [J]. J. Chem. Inf. Model. 2019, 59, 1253-1268). The data in the literature is collected from the CHEMBL database, and contains 5348 compound molecules of various types, which is a large and complex data set. In order to eliminate the influence of data dimension on the prediction performance of the model, the present application takes the Log(CL) form of the metabolic clearance rate CL value of each compound as the prediction endpoint of the model. Part of the information of the compounds is shown in Table 1.

[0068] Table 1 Partial human liver microsomal metabolic clearance rate data

[0069]

[0070] The establishment and verification of the QSAR local model and its consistency model include the following steps:

[0071] (1) Construct the modeling data set. For the selected set A containing 5348 compounds, calculate the 1-2D molecular descriptors of each compound molecule using PaDEL descriptor calculation software according to its SMILES code (as shown in Table 1). Calculate 1544 molecular descriptors for each compound, then correspond the descriptor values of each compound with its human liver microsomal metabolic clearance rate data, take the descriptors as the input of the QSAR model, and take the metabolic clearance rate Log(CL) as the output of the QSAR model, to form the modeling data set B.

[0072] (2) Descriptor cleaning. Operate on the data set B, first delete all constant or null descriptors. At the same time, if it is found that there are two descriptors that are pair-correlated (correlation coefficient > 0.85), delete the descriptor with less correlation with the metabolic clearance rate value. Finally obtain the simplified modeling data set C (the data set C after data cleaning contains 978 descriptors).

[0073] (3) Modeling dataset splitting. Randomly select 80% of the data in dataset C as the initial training dataset D, and the remaining 20% ​​of the data as the initial test dataset E.

[0074] (4) Key Descriptor Selection. First, the descriptors in the initial training dataset D are sorted by importance using the Mean Impurity Decrease (MDI) method. Then, non-important descriptors with a score value less than 0.01 are removed from both the initial training dataset D and the initial test dataset E, resulting in the final modeling dataset D1 and test dataset E1. After key descriptor selection, datasets D1 and E1 each contain 117 identical key descriptors. Dataset D1 is used to build the QSAR model, and dataset E1 is used to validate the performance of the built model.

[0075] (5) Construction of multiple local models based on model robustness. Before building each local model, 80% of the data in dataset D1 is randomly selected to construct dataset D2. The PSO algorithm is used to jointly optimize and select the data, descriptors, and hyperparameters of the SVR model in dataset D2. Then, the finally selected subset of data D1 is used. 3opt A QSAR local model based on SVR technology was established using the three hyperparameters of the SVR model. When optimizing the local model, the parameters of the PSO algorithm were set as follows: population size M = 30, particle dimension L = (ceil(5348*80% / 10) + ceil(117 / 10) + 3) = 443, maximum number of iterations T = 5000, learning factors c1 = c2 = 2, and inertia weight ω = 1.0 - 0.7t / T (t is the current iteration number of the algorithm). A total of 75 local models were established, and the data contained in these 75 local models can cover 95.27% of the training dataset D1.

[0076] (6) Complete the application domain analysis of the local models. Calculate the application domain of each local model using the IWD method to determine the reliable application range of each local model.

[0077] (7) Perform validity verification and model coverage calculation for each local model. Calculate the R-value for each local model. 2 R 2 cv5 and R 2 T The results show that a total of 67 local models satisfy R. 2 >0.6, R 2 cv5 >0.5 and R 2 T These models, with an R value > 0.6, passed the validity validation. These 67 models have R... 2 The values ​​of R are distributed between [0.72, 0.81].2 cv5 The values ​​of R are distributed between [0.61, 0.67]. 2 T The values ​​are distributed between [0.63, 0.68]. Statistical results show that the coverage of each local model that passed the validity validation is between [63.3%, 71.2%]. This means that for the local model with the widest application domain, at most 71.2% of the test compounds in the test dataset E1 fall within the application domain of this model, and can provide reliable prediction results. This indicates that the application scope of these local models is relatively narrow.

[0078] (8) Construction of the consistency model. The 67 local models that passed validity verification were used as sub-models of the consistency model and combined in parallel to form the consistency model. For the compound to be predicted, a was first classified according to the application domain of each local model to determine which local model's application domain a fell into. If a fell into the application domain of a local model, the prediction result of that local model for a was considered reliable; otherwise, it was unreliable. Finally, the reliable prediction results of all local models for a were averaged to obtain the final prediction result of the consistency model for a. If a could not fall into the application domain of any local model, the consistency model could not reliably predict the compound.

[0079] The predictive performance of the consistency model was evaluated and compared. The performance of the consistency model was assessed using the test dataset E1. The results show that the consistency model can provide reliable predictions for 98.27% of the compounds in E1, meaning that 98.97% of the tested compounds fall within the application domain of at least one of the 67 valid local models. The predictive performance metric R of the consistency model is [not specified]. 2 T =0.73. Compared with the local models, the consistency model has improved prediction performance and a wide range of applications, capable of reliably predicting the vast majority of compounds in the test dataset E1. For further comparative analysis, a global QSAR model was built using all modeling data from the modeling dataset D1. When building the global model, data and descriptor selection was no longer performed; only the PSO algorithm (with the same parameters as the local models for fair comparison) was used to optimize the three hyperparameters of SVR. Statistical results show that the global model R... 2 =0.71, R 2 cv5 =0.51, R 2 T =0.65, the model coverage is 97.47%. The five-fold cross-validation coefficient R0 of the global model is... 2 cv5Significantly lower than each local model and has approached to meet the model validity verification index R 2 cv5 The lower limit of 0.5, the robustness of the model is poor; The prediction performance of the global model is significantly lower than that of the consistency model; At the same time, the coverage rate of the global model is also smaller than that of the consistency model. This shows that the construction method of the local model and its consistency model based on the robustness of the model can not only improve the robustness of the model, but also improve the prediction performance and application range of the model.

Claims

1. A method for constructing a quantitative structure-activity relationship model based on model robustness, characterized in that: For the selected liver metabolic clearance rate modeling data set, the QSAR local model is established by support vector machine regression, and the five-fold cross-validation coefficient of the local model is optimized by particle swarm optimization algorithm, and the minimum value is taken as the optimization target, and the compound, molecular descriptor representing the structure of the compound and the hyperparameter of the model in the modeling data set are selected, so as to establish a QSAR local model with strong robustness and good reliability; The local model consistency modeling method based on the application domain restriction of the model is used to combine the established multiple local models into a consistent model under the restriction of the application domain of each local model, improve the prediction performance of the model, and expand the application range of the model; The method comprises the following specific steps: Step 1): selecting a compound set for establishing a liver metabolic clearance rate QSAR model, calculating the descriptors of the compounds by descriptor calculation software, and constructing a modeling data set; Step 2): descriptor cleaning; Step 3): modeling data set segmentation; Step 4): key descriptor selection; Step 5): local model construction based on model robustness; Step 6): application domain analysis of each local model; Step 7): effectiveness verification of each local model, and calculation of the coverage rate of each local model passing the effectiveness verification; Step 8): consistency model construction; Step 9): consistency model prediction performance evaluation, comparison analysis of the consistency model with the local model and the traditional global model; The step 5) needs to establish a plurality of local models, in order to obtain structural differentiation of local models, before each local model is established, 80% of data in the modeling data set D1 is randomly selected to constitute a modeling data set D2 of the local model; through the PSO optimization algorithm, the reciprocal of the five-fold cross-validation coefficient of the local model is taken as the optimization target to obtain the minimum value, and the joint selection of the modeling data, the molecular descriptor and the model hyperparameter is carried out; according to the optimization result of the PSO algorithm, the optimized data subset D 3opt is selected from the data set D2, and the hyperparameter of the SVR model is determined, so that the established local model can obtain good robustness; The step 8) completes the construction of the consistency model; the local model passing the effectiveness verification constructed in step (5) is used as a sub-model of the consistency model, and the sub-models are connected in parallel to form the consistency model; for a to be predicted compound a, firstly, the application domain of a is judged according to the application domain of each local model, and it is judged which local model the application domain of a falls into; if a falls into the application domain of a local model, it is considered that the prediction result of the local model for a is reliable, otherwise it is unreliable; finally, the reliable prediction results of all local models for a are averaged to obtain the final prediction result of the consistency model for a; if a cannot fall into the application domain of any local model, the consistency model cannot reliably predict the compound.

2. The model robustness-based quantitative structure-activity relationship model construction method according to claim 1, characterized in that: The step 5) establishes multiple optimized local models by particle swarm optimization algorithm until the data contained in the local models can cover more than 195% of the data in the training data set D.