Method for constructing permeability prediction model, compound screening method and electronic device

By constructing a permeability prediction model that combines basic and mechanistic descriptors, the problems of high computational cost and insufficient accuracy of existing permeability prediction tools are solved, enabling rapid and accurate compound screening and structure optimization guidance.

CN121938480BActive Publication Date: 2026-06-19PEKING UNIV INST OF ADVANCED AGRI SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEKING UNIV INST OF ADVANCED AGRI SCI
Filing Date
2026-03-30
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing technologies lack rapid and reliable tools for predicting penetration in the early stages of drug discovery, especially when dealing with massive compound libraries. Existing methods are computationally expensive, lack accuracy, or cannot provide clear guidance for structure optimization.

Method used

A permeability prediction model is constructed. A preliminary screening model is trained using basic descriptors for rapid filtering, and a diagnostic model is trained using mechanistic descriptors to provide guidance for permeability prediction and structure optimization. This includes the use of basic descriptors such as lipid-water partition coefficient, topological polar surface area, and molecular weight, and mechanistic descriptors such as distribution coefficient, polar surface area, van der Waals surface area, and number of hydrogen bond donors.

Benefits of technology

It achieves a balance between computational efficiency, prediction accuracy, and decision guidance, enabling rapid screening of highly permeable compounds and providing clear structural optimization suggestions, while reducing computational load and trial-and-error costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121938480B_ABST
    Figure CN121938480B_ABST
Patent Text Reader

Abstract

This invention relates to the field of computational chemistry, specifically to a method for constructing a permeability prediction model, a compound screening method, and an electronic device. The method for constructing the permeability prediction model includes: training a primary screening model using a first dataset containing basic descriptors; training a diagnostic model using a second dataset containing mechanistic descriptors; and associating and storing or integrating the primary screening model and the diagnostic model to obtain the permeability prediction model. The permeability prediction model is configured to: call the primary screening model to filter the library of compounds to be tested to obtain a candidate subset; and call the diagnostic model to predict the permeability of the candidate subset. The permeability prediction model combines the high-throughput screening efficiency of the primary screening model with the high accuracy and interpretability advantages of the diagnostic model, significantly improving prediction accuracy and revealing the deep correlation between molecular structure features and permeability, providing a theoretical basis for targeted optimization of compounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computational chemistry, specifically to a method for constructing a permeability prediction model, a compound screening method, and an electronic device. Background Technology

[0002] The Caco-2 cell monolayer permeability assay is the gold standard for evaluating the intestinal absorption potential of drugs in vitro. In the early stages of drug discovery, facing a vast library of compounds, there is an urgent need for a rapid and reliable predictive tool for large-scale initial screening.

[0003] Among related technologies, prediction methods based on molecular dynamics simulations or complex quantum chemical calculations, while highly accurate, take hours to days per calculation, making them unsuitable for initial screening of compound libraries with tens of thousands or even millions of compounds, and resulting in high computational costs. Machine learning-based prediction models, while improving computational speed, generally suffer from "black box" characteristics. These models typically only output permeability values ​​or classification labels; when a compound is predicted to have poor permeability, they cannot clearly indicate which molecular structural feature should be modified first, limiting their guiding value. Simple empirical rules (such as the "five rules of drug-likeness") or models based on a very small number of descriptors, while fast in prediction, generally lack sufficient accuracy and reliability, resulting in a high false screening rate, making them unreliable screening tools. Summary of the Invention

[0004] This invention provides a method for constructing a permeability prediction model, a compound screening method, and an electronic device, to offer a permeability prediction solution that achieves a balance between computational efficiency, prediction accuracy, and decision guidance.

[0005] In a first aspect, the present invention provides a method for constructing a permeability prediction model, comprising the following steps: obtaining a first dataset, wherein the first dataset includes multiple first compounds, the permeability of each first compound, and a basic descriptor for each first compound; training a preset first model using the first dataset to obtain a preliminary screening model; obtaining a second dataset, wherein the second dataset includes multiple second compounds, the permeability of each second compound, and a mechanism descriptor for each second compound; training a preset second model using the second dataset to obtain a diagnostic model; associating and storing or integrating the preliminary screening model and the diagnostic model to obtain a permeability prediction model; wherein the permeability prediction model is configured to: call the preliminary screening model to filter the test compound library to obtain a candidate subset, and call the diagnostic model to predict the permeability of the candidate subset.

[0006] The method for constructing a permeability prediction model provided by this invention first trains a preliminary screening model using basic descriptors. This model can quickly filter massive compound libraries, ensuring the computational efficiency of high-throughput screening. Then, a diagnostic model is trained by introducing a second dataset containing mechanism descriptors. Utilizing the three-dimensional structural information contained in the mechanism descriptors, not only is the accuracy of permeability prediction significantly improved, but the deep correlation between molecular structural features and permeability can also be resolved, thus endowing the model with excellent interpretability and providing a clear theoretical basis for subsequent targeted optimization of drug molecules. Therefore, this invention provides a permeability prediction model that balances screening speed, accuracy, and guidance value.

[0007] In some alternative implementations, the basic descriptor includes at least one of the following: lipid-water partition coefficient, topological polar surface area, and molecular weight; the mechanistic descriptor includes at least one of the following: distribution coefficient, polar surface area, van der Waals surface area, number of hydrogen bond donors, and polarizability.

[0008] This implementation utilizes lipid-water partition coefficients, topological polar surface areas, and / or molecular weights to construct a preliminary screening model, enabling rapid and low-cost calculations based on the fundamental physicochemical properties of molecules. By introducing distribution coefficients, polar surface areas, van der Waals surface areas, the number of hydrogen bond donors, and / or polarizability to train the diagnostic model, it directly maps the core physical mechanisms of molecular transmembrane permeation. The distribution coefficient accurately reflects the influence of the molecular dissociation state on passive diffusion under physiological conditions. Van der Waals and polar surface areas jointly determine the spatial occupancy and solvation barrier of molecules in the phospholipid bilayer. The number of hydrogen bond donors quantifies the ability of molecules to form transmembrane hydrogen bonds, while polarizability characterizes the ability of molecules to deform their electron clouds in a heterogeneous electric field within the cell membrane. This descriptor selection strategy not only improves the model's predictive accuracy but also endows it with excellent interpretability.

[0009] Secondly, the present invention also provides a compound screening method, comprising the following steps: obtaining a library of compounds to be tested; using a permeability prediction model trained based on the above method to predict the permeability of the library of compounds to be tested; and screening a set of target compounds from the library of compounds to be tested based on the results of the permeability prediction.

[0010] The compound screening method provided by this invention achieves an organic unity of computational efficiency, prediction accuracy, and decision guidance by using the permeability prediction model obtained from the above training to predict the permeability of the test compound library.

[0011] In some optional implementations, using the permeability prediction model trained based on the above method to predict the permeability of the test compound library includes: obtaining the basic descriptors of each compound in the test compound library; inputting the basic descriptors of each compound in the test compound library into the initial screening model of the permeability prediction model to obtain the first permeability prediction value of each compound in the test compound library; screening out compounds from the test compound library whose first permeability prediction value is greater than a preset first threshold to obtain a candidate subset; obtaining the mechanism descriptors of each compound in the candidate subset; inputting the mechanism descriptors of each compound in the candidate subset into the diagnostic model of the permeability prediction model to obtain the second permeability prediction value of each compound in the candidate subset; and screening out compounds from the candidate subset whose second permeability prediction value is greater than a preset second threshold to obtain the target compound set.

[0012] This implementation first uses a preliminary screening model to quickly pre-filter the test compound library, efficiently removing compounds with significantly substandard permeability, thereby greatly reducing the subsequent computational load. Then, for the candidate subset retained by the preliminary screening, a mechanism descriptor and diagnostic model with high computational complexity but higher accuracy are introduced for in-depth evaluation, thereby enabling accurate identification and screening of the target compound set from the candidate subset.

[0013] In some optional embodiments, after obtaining the target compound set, the following steps are also included: obtaining the topological polar surface area of ​​each compound in the target compound set; screening out compounds from the target compound set whose topological polar surface area is less than a preset third threshold to obtain a modified target compound set.

[0014] This implementation method, by introducing an independent screening step based on topological polarity surface area, can effectively correct the prediction bias caused by the "parameter compensation" effect (such as high lipid-water partition coefficient masking high polarity surface area) in the initial screening model, and accurately eliminate those pseudo-dominant molecules that, although predicted to have high scores, actually have excessive polarity and poor permeability; thus ensuring that the compounds in the target compound set not only have high permeability potential, but also conform to the drug-like structural rules.

[0015] In some optional embodiments, the compound screening method further includes the following steps: identifying the compounds to be analyzed in the test compound library; obtaining the regression coefficients of molecular descriptors in the permeability prediction model, screening out molecular descriptors with negative regression coefficients to obtain target molecular descriptors; wherein the molecular descriptors include basic descriptors and / or mechanism descriptors; obtaining the parameter values ​​corresponding to the target molecular descriptors in the compounds to be analyzed; calculating the individual contribution value of the target molecular descriptors based on the parameter values ​​and the corresponding regression coefficients; and determining the key limiting factors restricting the permeability of the compounds to be analyzed based on the individual contribution values ​​of all target molecular descriptors in the compounds to be analyzed.

[0016] This implementation method, by calculating the individual contribution value of each molecular descriptor, can intuitively quantify and locate the shortcomings that hinder the permeability of compounds, such as excessively high polar surface area or inappropriate lipid-water partition coefficient, thereby providing a clear direction for molecular structure optimization and improving the efficiency and success rate of lead compound optimization.

[0017] In some optional implementations, after identifying the key limiting factors restricting the permeability of the compound to be analyzed, the steps further include: generating structural modification suggestions based on the key limiting factors restricting the permeability of the compound to be analyzed; modifying the structure of the compound to be analyzed based on the structural modification suggestions to construct a corresponding optimized molecule; and using a permeability prediction model to calculate the predicted permeability values ​​of the compound to be analyzed and the optimized molecule, respectively, to obtain a comparison of the predicted permeability of the compound to be analyzed before and after structural modification.

[0018] This implementation provides structural modification suggestions for permeability bottlenecks, which can pre-verify the modification effect in a virtual environment and select the optimal optimization path without actual synthesis, thereby significantly reducing trial and error costs.

[0019] Thirdly, the present invention also provides an electronic device, including a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the method for constructing a permeability prediction model according to the first aspect or any corresponding embodiment and / or the method for screening compounds according to the second aspect or any corresponding embodiment.

[0020] Fourthly, the present invention also provides a computer-readable storage medium storing computer instructions, the computer instructions being used to cause a computer to perform the method for constructing a permeability prediction model according to the first aspect or any corresponding embodiment thereof and / or the method for screening compounds according to the second aspect or any corresponding embodiment thereof.

[0021] Fifthly, the present invention also provides a computer program product, including computer instructions for causing a computer to execute the method for constructing a permeability prediction model according to the first aspect or any corresponding embodiment thereof and / or the method for screening compounds according to the second aspect or any corresponding embodiment thereof. Attached Figure Description

[0022] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention;

[0024] Figure 2 This is a flowchart of a method for constructing a penetration prediction model according to an embodiment of the present invention;

[0025] Figure 3 This is a first flowchart of a compound screening method according to an embodiment of the present invention;

[0026] Figure 4 This is a second flowchart of the compound screening method according to an embodiment of the present invention;

[0027] Figure 5 This is a third flowchart of the compound screening method according to embodiments of the present invention;

[0028] Figure 6 This is a structural block diagram of a permeability prediction model construction device according to an embodiment of the present invention;

[0029] Figure 7 This is a structural block diagram of a compound screening device according to an embodiment of the present invention;

[0030] Figure 8 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0033] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0034] As an optional application scenario of this invention, such as Figure 1As shown, the application scenario includes at least one terminal device and at least one server. Figure 1 The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.

[0035] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.

[0036] According to an embodiment of the present invention, a method for constructing a penetration prediction model is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0037] This embodiment provides a method for constructing a penetration prediction model, which can be used on mobile terminals such as mobile phones and tablets. Figure 2 This is a flowchart of a method for constructing a penetration prediction model according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:

[0038] Step S201: Obtain a first dataset, which includes multiple first compounds, the permeability of each first compound, and the basic descriptor of each first compound.

[0039] The permeability of the first compound can be considered as the apparent permeability coefficient of Caco-2 cells.

[0040] The basic descriptor includes at least one of the following: lipid-water partition coefficient LogP, topological polar surface area TPSA, and molecular weight MW. LogP represents the basic lipid solubility of the first compound and is the main driving force for passive diffusion; TPSA represents the polarity of the molecule and is the main resistance to penetrating lipid cell membranes; and MW is a proxy variable for molecular size.

[0041] Basic descriptors are available at virtually no cost. Specifically, the cheminformatics toolkit RDKit can be used to perform molecular structure analysis and feature extraction on the first compound to obtain the lipid-water partition coefficient LogP, topological polar surface area TPSA, and molecular weight MW; alternatively, they can be obtained by querying existing chemical databases or scientific literature.

[0042] Step S202: Train the preset first model using the first dataset to obtain the initial screening model.

[0043] For example, the first model can be: Caco = α0 + α1·LogP + α2·TPSA + α3·MW, where α0 is a constant, and α1, α2, and α3 represent the regression coefficients of the lipid-water partition coefficient LogP, the topological polar surface area TPSA, and the molecular weight MW, respectively.

[0044] Step S203: Obtain a second dataset, which includes multiple second compounds, the permeability of each second compound, and the mechanism descriptor of each second compound.

[0045] The first dataset is identical to, partially identical to, or completely different from the second dataset. The permeability of the second compound can be represented by the apparent permeability coefficient of Caco-2 cells.

[0046] The mechanism descriptor includes at least one of the following: distribution coefficient LogD, polar surface area PSA, van der Waals surface area VDW_Area, number of hydrogen bond donors HBD, and polarizability.

[0047] Mechanism descriptors can directly map the core physical mechanisms of molecular transmembrane permeation. Specifically, the distribution coefficient LogD accurately reflects the influence of the molecular dissociation state on passive diffusion under physiological conditions; van der Waals surface area (VDW_Area) and polar surface area (PSA) jointly determine the spatial occupancy and solvation barrier of the molecule in the phospholipid bilayer; the number of hydrogen bond donors (HBD) quantifies the molecule's ability to form transmembrane hydrogen bonds; and polarizability characterizes the molecule's ability to deform its electron cloud in a heterogeneous electric field within the cell membrane. This descriptor selection strategy not only improves the model's predictive accuracy but also endows it with excellent interpretability.

[0048] Specifically, the molecular simulation software MOE can be used to construct the three-dimensional structure of the second compound and perform energy minimization to obtain the distribution coefficient LogD (pH=7.4), polar surface area PSA, van der Waals surface area VDW_Area, and polarizability; the cheminformatics toolkit RDKit can be used to perform molecular structure analysis and feature extraction on the second compound to obtain the number of hydrogen bond donors HBD.

[0049] Step S204: Train the pre-set second model using the second dataset to obtain the diagnostic model.

[0050] For example, the second model can be: Caco = β0 + β1·LogD + β2·PSA + β3·VDW_Area + β4·HBD + β5·Polarizability. Here, β0 is a constant, and β1, β2, β3, β4, and β5 represent the regression coefficients of the distribution coefficient LogD, polar surface area PSA, van der Waals surface area VDW_Area, hydrogen bond donor number HBD, and polarizability, respectively.

[0051] Step S205: Link and store or integrate the primary screening model and the diagnostic model to obtain the permeability prediction model; wherein the permeability prediction model is configured to: call the primary screening model to filter the test compound library to obtain a candidate subset, and call the diagnostic model to perform permeability prediction on the candidate subset.

[0052] The method for constructing the permeability prediction model provided in this embodiment first uses basic descriptors to train a preliminary screening model, which can quickly filter a massive compound library, ensuring the computational efficiency of high-throughput screening. Then, a diagnostic model is trained by introducing a second dataset containing mechanism descriptors. Utilizing the three-dimensional structural information contained in the mechanism descriptors, not only is the accuracy of permeability prediction significantly improved, but the deep correlation between molecular structural features and permeability can also be resolved, thus endowing the model with excellent interpretability and providing a clear theoretical basis for subsequent targeted optimization of drug molecules. In summary, this embodiment provides a permeability prediction model that balances screening speed, accuracy, and guidance value.

[0053] To illustrate more clearly the construction method of the permeability prediction model of the present invention, a specific example 1 is given, which includes the following steps:

[0054] 1. Data Sources and Preprocessing

[0055] The Caco-2 permeability data used in this example comes from public datasets (such as the ChEMBL database) and internal experimental data from collaborating laboratories, collecting permeability values ​​(unit: nm / s) for 1200 compounds. After data cleaning, samples with duplicate structures or abnormal permeability values ​​were removed, ultimately retaining 1050 compounds as the modeling dataset. These were randomly divided into a training set (840 compounds) and a test set (210 compounds) in an 8:2 ratio.

[0056] 2. Descriptor Calculation

[0057] Basic descriptors: Calculate LogP, TPSA, MW, and HBD using RDKit.

[0058] Mechanism descriptor: After minimizing the energy of the compound using MOE, LogD (pH=7.4), PSA, VDW_Area, and Polarizability were calculated.

[0059] 1. Model building and validation

[0060] (1) Initial screening model

[0061] Using the multiple linear regression method, with LogP, TPSA, and MW as independent variables and Caco-2 as the dependent variable, the initial screening model was constructed as follows: Caco = -120.5 + 35.2·LogP - 2.8·TPSA + 0.3·MW.

[0062] The model, after rigorous validation, demonstrated excellent goodness of fit and robustness. The coefficient of determination R0 on the training set... 2 The coefficient of determination R for the test set is 0.851. 2 The R² value was 0.845, showing a high degree of consistency and close to 1, demonstrating that the initial screening model maintained stable explanatory power when fitting known data and predicting unknown samples, without overfitting. The root mean square error (RMSE) was controlled at 12.3 nm / s, indicating a small average deviation between the predicted and experimental values. Furthermore, the R² value of the 10-fold cross-validation was [not specified in the original text]. 2 The p-value was 0.842, which is highly consistent with the results on both the training and test sets, further confirming the model's good generalization ability. Regarding statistical significance, the p-values ​​for the three independent variables, LogP, TPSA, and MW, were all less than 0.01, passing the high-confidence significance test. This statistically confirms a significant linear correlation between these three descriptors and Caco-2 permeability, rather than random noise interference.

[0063] The initial screening model integrates the relationship between Caco-2 permeability and the most basic and readily available molecular descriptors. In ultra-early virtual screening, when faced with a massive library of compounds and only simple 2D structural information, this model can quickly and cost-effectively conduct a preliminary assessment of permeability potential.

[0064] (2) First diagnostic model

[0065] The first diagnostic model was constructed using multiple linear regression with LogD, VDW_Area, and PSA as independent variables and Caco-2 value as the dependent variable: Caco = -85.7 + 30.8·LogD - 0.25·VDW_Area - 1.9·PSA.

[0066] The coefficient of determination (R²) for the training set was 0.896, and for the test set it was 0.892; the root mean square error (RMSE) was 9.8 nm / s. The first diagnostic model incorporates physiologically relevant LogD and three-dimensional structure descriptors, more closely resembling biological experimental conditions, and is suitable for guiding structure modification in the lead compound optimization stage.

[0067] Based on the positive and negative directions of the regression coefficients in the first diagnostic model, the specific direction of structural modification can be clearly defined. Specifically, the regression coefficient of LogD is positive (+30.8), indicating that the lipophilicity of the molecule is positively correlated with its permeability. Therefore, in the optimization of lead compounds, the LogD value should be consciously increased by introducing hydrophobic groups or reducing polarity. On the other hand, the regression coefficients of VDW_Area and PSA are both negative (-0.25 and -1.9, respectively), indicating that excessively large molecular volume or excessively strong polarity will significantly hinder permeation. Therefore, the overall size of the molecule needs to be strictly controlled to maintain a compact structure.

[0068] (3) Second diagnostic model

[0069] A second diagnostic model was constructed using multiple linear regression with LogD, Polarizability, VDW_Area, PSA, and HBD as independent variables and Caco-2 value as the dependent variable: Caco = -52.4 + 28.5·LogD + 1.2·Polarizability - 0.18·VDW_Area - 1.5·PSA - 8.1·HBD.

[0070] The coefficient of determination (R²) for the training set was 0.931; the coefficient of determination (R²) for the test set was 0.928; and the root mean square error (RMSE) was 7.6 nm / s.

[0071] The absolute values ​​of the regression coefficients based on the second diagnostic model can clearly define the influence of each factor on intestinal permeability. The coefficient for lipid solubility (LogD) has the largest absolute value (|28.5|), followed by the number of hydrogen bond donors (HBD) (|28.5|). 8.1), these two constitute the primary determinant of permeability; in contrast, the absolute values ​​of the coefficients of polar surface area (PSA) and van der Waals surface area (VDW_Area) are smaller, belonging to secondary moderating factors, while polarizability has the weakest influence, belonging to subtle influencing factors. The positive or negative sign of the regression coefficients further indicates the specific direction of structural modification. LogD (+28.5) and polarizability (+1.2) show a positive correlation, indicating that improving these two indicators can effectively enhance permeability; while HBD ( 8.1), PSA ( 1.5) and VDW_Area ( The values ​​of 0.18 are all negative, which means that excessively high values ​​of these parameters will significantly hinder drug absorption.

[0072] In summary, to achieve high oral absorption, molecular design must follow the principles of "high lipid solubility, few hydrogen bond donors, low polarity, and compact structure." That is, while prioritizing maximizing lipid solubility, it is crucial to minimize hydrogen bond donors to avoid severe osmotic penalty, and to appropriately control the size and polarity of the molecule.

[0073] It's important to note that the first and second diagnostic models are not progressive but rather two independent choices based on different input conditions. The first diagnostic model requires only three descriptors: LogD, VDW_Area, and PSA, making it suitable for scenarios where 3D structural information is incomplete or where rapid evaluation is needed. The second diagnostic model further introduces Polarizability and HBD, which, while requiring more sophisticated input data, provides more refined and comprehensive penetration predictions and optimization guidance. Therefore, in practical applications, the choice between the two should be made based on the completeness of the available data and the specific needs of the optimization stage.

[0074] This embodiment provides a compound screening method that can be used on mobile terminals, such as mobile phones and tablets. Figure 3 This is a first flowchart of a compound screening method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps:

[0075] Step S301: Obtain the library of compounds to be tested.

[0076] Specifically, a test compound library refers to a digital collection containing a large number of molecular structures to be evaluated. These libraries come from a wide range of sources, including physical compound inventories accumulated within pharmaceutical companies, commercial virtual compound databases, and theoretical molecular libraries generated through combinatorial chemistry methods. This compound library is a fundamental resource for computer-aided drug design, providing a vast pool of candidate molecules for subsequent high-throughput virtual screening.

[0077] Step S302: Use the permeability prediction model trained based on the above-mentioned permeability prediction model construction method to predict the permeability of the test compound library, and select the target compound set from the test compound library according to the permeability prediction results.

[0078] As mentioned above, the permeability prediction model includes a primary screening model and a diagnostic model. The primary screening model is used to quickly predict the permeability of the test compound library, while the diagnostic model is used to accurately quantify the candidate compounds after primary screening.

[0079] The target compound set includes one or more compounds, which refers to the set of compounds with high intestinal permeability potential obtained after screening by a permeability prediction model.

[0080] The compound screening method provided in this embodiment achieves an organic unity of computational efficiency, prediction accuracy, and decision guidance by using the permeability prediction model trained above to predict the permeability of the test compound library.

[0081] This embodiment provides a compound screening method that can be used on mobile terminals, such as mobile phones and tablets. Figure 4 This is a second flowchart of the compound screening method according to an embodiment of the present invention, as shown below. Figure 4 As shown, the process includes the following steps:

[0082] Step S401: Obtain the library of compounds to be tested.

[0083] Step S402: Use the permeability prediction model trained based on the above method to predict the permeability of the test compound library.

[0084] In one optional implementation, the permeability prediction model trained based on the above method is used to predict the permeability of the test compound library, including the following steps S4021 to S4026.

[0085] Step S4021: Obtain the basic descriptors of each compound in the test compound library.

[0086] Step S4022: Input the basic descriptors of each compound in the test compound library into the initial screening model of the permeability prediction model to obtain the first permeability prediction value of each compound in the test compound library.

[0087] Step S4023: Select compounds from the test compound library whose first permeability prediction value is greater than the preset first threshold to obtain a candidate subset.

[0088] In other words, a lightweight initial screening model is used to quickly score and rank the library of compounds to be tested. A first threshold is set to filter out compounds with extremely poor predicted permeability (such as the bottom 90%), retaining the advantageous candidate subset. For example, the first threshold can be 80 nm / s.

[0089] Step S4024: Obtain the mechanism descriptors for each compound in the candidate subset.

[0090] Step S4025: Input the mechanism descriptors of each compound in the candidate subset into the diagnostic model of the permeability prediction model to obtain the second permeability prediction value of each compound in the candidate subset.

[0091] Step S4026: Select compounds from the candidate subset whose second permeability prediction value is greater than the preset second threshold to obtain the target compound set.

[0092] In other words, for the candidate subset retained in the initial screening, a more accurate diagnostic model is used for in-depth evaluation, and a second threshold is set to further eliminate molecules with permeation defects. For example, the second threshold can be 100 nm / s.

[0093] Step S403: Based on the results of the permeability prediction, select the target compound set from the test compound library.

[0094] Step S404: Obtain the topological polar surface area of ​​each compound in the target compound set.

[0095] This is because the scoring mechanism of the initial screening model has a "parameter compensation" effect, meaning that an excessively high LogP value may mask the negative impact of an excessively high TPSA value on permeability, leading to an artificially high model prediction. Although these molecules meet the requirement that the first permeability prediction value is greater than the first threshold, they often have potential risks such as low solubility and high metabolic clearance in actual drug development. Therefore, they need to be identified and eliminated through independent structural feature screening.

[0096] Step S405: Select compounds from the target compound set whose topological polar surface area is less than a preset third threshold to obtain the corrected target compound set.

[0097] For example, the third threshold could be 40 Ų.

[0098] By screening compounds from the target compound set whose topological polar surface area is less than a preset third threshold, "structural correction" of the target compound set is achieved. This effectively filters out potentially undesirable molecules caused by imbalances in physicochemical properties (high LogP compensates for high TPSA), ensuring that the final modified target compounds have high permeability while their molecular structure conforms to the basic rules of drug-like properties, thus reducing the risk of failure in subsequent research and development.

[0099] The compound screening method provided in this embodiment first uses a preliminary screening model to quickly filter the test compound library to obtain a candidate subset, significantly reducing the computational load. Then, a high-precision diagnostic model is introduced to perform in-depth evaluation of the candidate subset, ensuring the reliability of the prediction results. Finally, topological polar surface area is introduced as an independent structural correction index to effectively identify and eliminate non-ideal molecules. This progressive screening mechanism not only significantly improves the discovery efficiency of highly permeable compounds but also ensures the druggability of the final target compound set through multi-dimensional quality control.

[0100] This embodiment provides a compound screening method that can be used on mobile terminals, such as mobile phones and tablets. Figure 5 This is a third flowchart of the compound screening method according to embodiments of the present invention, such as... Figure 5 As shown, the process includes the following steps:

[0101] Step S501: Obtain the library of compounds to be tested.

[0102] Step S502: Use the permeability prediction model trained based on the above-mentioned permeability prediction model construction method to predict the permeability of the test compound library, and select the target compound set from the test compound library based on the permeability prediction results.

[0103] Step S503: Identify the compounds to be analyzed in the test compound library.

[0104] The compound to be analyzed can be a specific molecule specified by human intervention, or a representative sample automatically extracted from the compound library; it can be a preferred molecule in the target compound set, or a compound whose first permeability prediction value is less than or equal to a preset first threshold and / or a compound whose second permeability prediction value is less than or equal to a preset second threshold (i.e., a compound eliminated by a screening model or diagnostic model).

[0105] Step S504: Obtain the regression coefficients of molecular descriptors in the permeability prediction model, screen out molecular descriptors with negative regression coefficients, and obtain the target molecular descriptors, wherein the molecular descriptors include basic descriptors and / or mechanism descriptors.

[0106] The regression coefficients of each molecular descriptor in the permeability prediction model refer to the coefficients that the permeability prediction model learns and determines during the training process, and are used to quantify the degree and direction of the influence of each physicochemical parameter on the permeability prediction value.

[0107] For example, for the initial screening model Caco = -120.5 + 35.2·LogP -2.8·TPSA + 0.3·MW, +35.2, -2.8, and +0.3 are the regression coefficients of LogP, TPSA, and MW, respectively, and TPSA is the target molecule descriptor.

[0108] For the diagnostic model Caco = -85.7 + 30.8·LogD - 0.25·VDW_Area - 1.9·PSA, +30.8, -0.25, and -1.9 are the regression coefficients of LogD, VDW_Area, and PSA, respectively, and VDW_Area and PSA are the target molecule descriptors.

[0109] Step S505: Obtain the parameter values ​​corresponding to the target molecule descriptor in the compound to be analyzed.

[0110] Specifically, the parameter values ​​corresponding to each target molecule descriptor can be calculated using cheminformatics tools (such as RDKit, Mordred, etc.); or they can be obtained by querying existing chemical databases (such as PubChem, ChEMBL, Reaxys, etc.) or scientific literature.

[0111] Step S506: Based on the parameter values ​​and the corresponding regression coefficients, calculate the individual contribution value of the target molecule descriptor.

[0112] Specifically, for any target molecule descriptor, the regression coefficient of the target molecule descriptor is multiplied by the parameter value in the compound to be analyzed that corresponds to the target molecule descriptor to obtain the single contribution value of the target molecule descriptor.

[0113] Step S507: Based on the individual contribution values ​​of all target molecule descriptors in the compound to be analyzed, determine the key limiting factors that restrict the permeability of the compound to be analyzed.

[0114] Specifically, the process first aggregates the individual contribution values ​​of all target molecule descriptors to obtain the total contribution value. Then, it calculates the proportion of each target molecule descriptor's individual contribution value within the total contribution value, thereby quantifying the relative importance of each target molecule descriptor and identifying key limiting factors. As an example, target molecule descriptors with a proportion exceeding a preset fourth threshold can be identified as key limiting factors. For example, the fourth threshold could be 85%.

[0115] In one optional implementation, after identifying the key limiting factors restricting the permeability of the compound to be analyzed, the method further includes the following steps: generating structural modification suggestions based on the key limiting factors restricting the permeability of the compound to be analyzed; modifying the structure of the compound to be analyzed based on the structural modification suggestions to construct a corresponding optimized molecule; and using a permeability prediction model to calculate the predicted permeability values ​​of the compound to be analyzed and the optimized molecule, respectively, to obtain a comparison of the predicted permeability of the compound to be analyzed before and after structural modification.

[0116] In other words, this implementation method, after identifying key limiting factors, does not stop at qualitative analysis but further constructs a closed-loop feedback mechanism of "diagnosis-optimization-validation." By generating specific structural modification suggestions for each compound under analysis based on its specific key limiting factors, and constructing a virtual optimized molecule accordingly, the system can use a permeability prediction model to compare and evaluate the molecule before and after modification. This simulation and validation process not only intuitively quantifies the permeability gain brought about by structural modification but also effectively predicts the feasibility of optimization strategies, thereby providing researchers with data-driven, precise guidance, significantly reducing trial-and-error costs, and accelerating the screening and development process of highly permeable lead compounds.

[0117] In one alternative implementation, after identifying the key limiting factors that restrict the permeability of the analyte, the method further includes the following steps: obtaining multiple key limiting factors for the analyte; statistically analyzing the most frequently occurring key limiting factors among the multiple analyte compounds, and calculating the average contribution of the most frequently occurring key limiting factors among the multiple analyte compounds.

[0118] This allows for precise identification of common bottlenecks restricting overall penetration performance and quantification of their average contribution, thus providing clear, targeted improvement directions for compound library optimization. Compared to traditional one-by-one analysis strategies, this method significantly improves R&D efficiency, avoids repetitive optimization work caused by neglecting common issues, and reduces the bias of subjective judgment through a data-driven approach. Furthermore, it presents the most frequently occurring key limiting factors and their average contribution across multiple analyzed compounds in the report.

[0119] In some alternative implementations, after identifying the key limiting factors restricting the permeability of the compound to be analyzed, the steps further include: obtaining the individual contribution values ​​of all target molecule descriptors of the compound to be analyzed; sorting all target molecule descriptors according to their individual contribution values ​​to generate a contribution distribution map.

[0120] This allows for a clear visual representation of the influence trends of each target molecule descriptor on the permeability of compounds, helping researchers quickly identify the core molecular features that affect permeability and avoid overlooking potentially important influencing factors.

[0121] In summary, the compound screening method provided in this embodiment first uses a lightweight initial screening model to quickly score and rank the complete virtual compound library, sets a permeability threshold, and filters out compounds with extremely poor predicted permeability (such as the bottom 90%), retaining a candidate subset. Second, for the retained candidate subset, more refined descriptors are calculated, and an interpretable diagnostic model is used for precise quantitative prediction. Simultaneously, contribution analysis of the model output is used to generate customized structural optimization suggestions for each candidate compound. Furthermore, for compounds in the target compound set deemed promising by the diagnostic model, more sophisticated but time-consuming computational methods or in vitro experiments can be used for final verification. The permeability prediction model acts as an efficient pre-filter, ensuring that expensive, high-precision computational resources are concentrated only on the most promising molecules.

[0122] This invention enables the optimal allocation of computing resources and information depth, processing massive amounts of data at the lowest cost, and focusing limited high-cost computing resources (used to obtain descriptors such as LogD) and the analytical effort required for optimization design on the most valuable subset of compounds, thereby doubling the overall R&D efficiency.

[0123] This embodiment also provides a construction apparatus for a penetration prediction model, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0124] This embodiment provides an apparatus for constructing a permeability prediction model, such as... Figure 6 As shown, it includes:

[0125] The first acquisition module 601 is used to acquire a first dataset, wherein the first dataset includes multiple first compounds, the permeability of each first compound, and the basic descriptor of each first compound.

[0126] The first training module 602 is used to train the preset first model using the first dataset to obtain the initial screening model.

[0127] The second acquisition module 603 is used to acquire a second dataset, wherein the second dataset includes multiple second compounds, the permeability of each second compound, and the mechanism descriptor of each second compound.

[0128] The second training module 604 is used to train the pre-set second model using the second dataset to obtain the diagnostic model.

[0129] The combination module 605 is used to associate and store or integrate the primary screening model and the diagnostic model to obtain the permeability prediction model; wherein the permeability prediction model is configured to: call the primary screening model to filter the test compound library to obtain a candidate subset, and call the diagnostic model to perform permeability prediction on the candidate subset.

[0130] In some alternative implementations, the basic descriptor includes at least one of the following: lipid-water partition coefficient, topological polar surface area, and molecular weight; the mechanistic descriptor includes at least one of the following: distribution coefficient, polar surface area, van der Waals surface area, number of hydrogen bond donors, and polarizability.

[0131] The apparatus for constructing a permeability prediction model provided in this embodiment of the invention can execute the method for constructing a permeability prediction model provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as in the corresponding embodiments described above, and will not be repeated here.

[0132] This embodiment also provides a compound screening device, which is used to implement the above embodiments and preferred embodiments. Details that have been described will not be repeated here.

[0133] This embodiment provides a compound screening device, such as... Figure 7 As shown, it includes:

[0134] The third acquisition module 701 is used to acquire the library of compounds to be tested.

[0135] The prediction module 702 is used to perform permeability prediction on the test compound library using the permeability prediction model trained based on the above method, and to select the target compound set from the test compound library based on the permeability prediction results.

[0136] In some optional implementations, the prediction module 702 is specifically used for: acquiring the basic descriptors of each compound in the test compound library; inputting the basic descriptors of each compound in the test compound library into the initial screening model of the permeability prediction model to obtain the first permeability prediction value of each compound in the test compound library; screening out compounds from the test compound library whose first permeability prediction value is greater than a preset first threshold to obtain a candidate subset; acquiring the mechanism descriptors of each compound in the candidate subset; inputting the mechanism descriptors of each compound in the candidate subset into the diagnostic model of the permeability prediction model to obtain the second permeability prediction value of each compound in the candidate subset; and screening out compounds from the candidate subset whose second permeability prediction value is greater than a preset second threshold to obtain a target compound set.

[0137] In some optional implementations, after obtaining the target compound set, the prediction module 702 is further configured to: obtain the topological polar surface area of ​​each compound in the target compound set; and select compounds from the target compound set whose topological polar surface area is less than a preset third threshold to obtain a corrected target compound set.

[0138] In some optional embodiments, the compound screening device further includes an analysis module. The analysis module is specifically used for: identifying the compounds to be analyzed in the target compound library; obtaining the regression coefficients of molecular descriptors in the permeability prediction model, screening out molecular descriptors with negative regression coefficients to obtain target molecular descriptors, wherein the molecular descriptors include basic descriptors and / or mechanistic descriptors; obtaining the parameter values ​​corresponding to the target molecular descriptors in the target compounds; calculating the individual contribution value of the target molecular descriptors based on the parameter values ​​and the corresponding regression coefficients; and determining the key limiting factors restricting the permeability of the target compounds based on the individual contribution values ​​of all target molecular descriptors in the target compounds.

[0139] In some optional implementations, after identifying the key limiting factors restricting the permeability of the compound to be analyzed, the analysis module is further configured to: generate structural modification suggestions based on the key limiting factors restricting the permeability of the compound to be analyzed; modify the structure of the compound to be analyzed based on the structural modification suggestions to construct a corresponding optimized molecule; and use a permeability prediction model to calculate the predicted permeability values ​​of the compound to be analyzed and the optimized molecule respectively, so as to obtain a comparison of the predicted permeability of the compound to be analyzed before and after structural modification.

[0140] The compound screening apparatus provided in this embodiment of the invention can execute the compound screening method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0141] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0142] The following is a detailed reference. Figure 8 This diagram illustrates a suitable structural schematic for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 801, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 802 or a program loaded from memory 808 into random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device. The processor 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0143] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0144] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a memory 808, or installed from a ROM 802. When the computer program is executed by the processor 801, it performs the functions defined in the method for constructing a permeability prediction model and / or the compound screening method of the embodiments of the present invention.

[0145] Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention.

[0146] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the method for constructing the permeability prediction model and / or the compound screening method shown in the above embodiments is implemented.

[0147] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0148] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for constructing a penetration prediction model, characterized in that, The method includes: Obtain a first dataset, wherein the first dataset includes a plurality of first compounds, the permeability of each first compound, and a basic descriptor for each first compound; wherein the basic descriptor includes at least one of the following: lipid-water partition coefficient, topological polar surface area, and molecular weight; The first dataset is used to train the preset first model to obtain the initial screening model; Obtain a second dataset, wherein the second dataset includes multiple second compounds, the permeability of each second compound, and a mechanism descriptor for each second compound; wherein the mechanism descriptor includes at least one of the following: distribution coefficient, polar surface area, van der Waals surface area, number of hydrogen bond donors, and polarizability; The second dataset is used to train the pre-defined second model to obtain the diagnostic model; The primary screening model and the diagnostic model are associated, stored, or integrated to obtain the permeability prediction model; wherein the permeability prediction model is configured to: call the primary screening model to filter the library of compounds to be tested to obtain a candidate subset, and call the diagnostic model to perform permeability prediction on the candidate subset; The process of calling the initial screening model to filter the test compound library to obtain a candidate subset, and calling the diagnostic model to predict the penetration of the candidate subset includes: Obtain the basic descriptors of each compound in the test compound library; The basic descriptors of each compound in the test compound library are input into the initial screening model of the permeability prediction model to obtain the first permeability prediction value of each compound in the test compound library; From the compound library to be tested, compounds whose first permeability prediction value is greater than a preset first threshold are selected to obtain a candidate subset; Obtain the mechanism descriptors of each compound in the candidate subset; The mechanism descriptors of each compound in the candidate subset are input into the diagnostic model of the permeability prediction model to obtain the second permeability prediction value of each compound in the candidate subset. Compounds whose second permeability prediction value is greater than a preset second threshold are selected from the candidate subset to obtain the target compound set.

2. A method for screening compounds, characterized in that, The method includes: Obtain the library of compounds to be tested; The permeability prediction model trained based on the method described in claim 1 is used to predict the permeability of the test compound library, and a set of target compounds is selected from the test compound library based on the results of the permeability prediction.

3. The method according to claim 2, characterized in that, After obtaining the target compound set, the process further includes: Obtain the topological polar surface area of ​​each compound in the target compound set; From the target compound set, compounds with a topological polar surface area less than a preset third threshold are selected to obtain a modified target compound set.

4. The method according to claim 2, characterized in that, Also includes: Identify the compounds to be analyzed in the aforementioned compound library; Obtain the regression coefficients of molecular descriptors in the permeability prediction model, filter out molecular descriptors with negative regression coefficients, and obtain the target molecular descriptors, wherein the molecular descriptors include basic descriptors and / or mechanism descriptors; Obtain the parameter values ​​in the compound to be analyzed that correspond to the descriptor of the target molecule; Based on the parameter values ​​and the corresponding regression coefficients, the individual contribution value of the target molecule descriptor is calculated. Based on the individual contribution values ​​of all the target molecule descriptors in the compound to be analyzed, the key limiting factors restricting the permeability of the compound to be analyzed are determined.

5. The method according to claim 4, characterized in that, After identifying the key limiting factors restricting the permeability of the analyte, the following is also included: Structural modification suggestions are generated based on the key limiting factors restricting the permeability of the compound to be analyzed; Based on the proposed structural modification, the compound to be analyzed was structurally modified to construct an optimized molecule. Using the permeability prediction model, the predicted permeability values ​​of the compound to be analyzed and the optimized molecule are calculated respectively, and the permeability prediction comparison results of the compound to be analyzed before and after structural modification are obtained.

6. An electronic device, characterized in that, include: A memory and a processor are interconnected, the memory storing computer instructions, and the processor executing the computer instructions to perform the method for constructing the permeability prediction model according to claim 1 and / or the compound screening method according to any one of claims 2 to 5.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the method for constructing the permeability prediction model according to claim 1 and / or the compound screening method according to any one of claims 2 to 5.

8. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the method for constructing the permeability prediction model of claim 1 and / or the compound screening method of any one of claims 2 to 5.

Citation Information

Patent Citations

  • Permeability determination method and device, electronic equipment and storage medium

    CN118548037A

  • Method for screening potential functional raw materials of cosmetics by establishing mathematical model based on molecular docking and physical and chemical properties of substances

    CN119993323A