Metalearning-based small sample pharmacokinetic property prediction method and related equipment

The meta-learning-based method integrates multi-dimensional features and weighted predictions to address the limitations of QSAR models in small-sample drug development scenarios, enhancing pharmacokinetic property prediction accuracy and reliability.

CN120319331AActive Publication Date: 2025-07-15PHARMARON NINGBO CO LTD

Patent Information

Application Number
CN202510800807.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-15
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Traditional QSAR models rely on a large amount of labeled data in the early stages of drug development, making it difficult to accurately predict pharmacokinetic properties in small sample scenarios, resulting in inaccurate prediction results.

Method used

Using a small sample pharmacokinetic properties prediction method based on meta-learning, a small sample pharmacokinetic properties prediction model is used to obtain multi-dimensional characteristic data of candidate compounds, and a pre-trained pharmacokinetic properties prediction model is used to combine the prediction results of the first machine learning model and the second machine learning model to comprehensively determine the pharmacokinetic properties of candidate compounds, and introduce standard deviation to determine the weight coefficient for reasonable weighting fusion.

Benefits of technology

It improves the accuracy and reliability of prediction of pharmacokinetic properties in small sample scenarios, enhances the model's understanding of molecular characteristics, reduces the prediction risks, and provides more reliable support for early drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120319331A_ABST
    Figure CN120319331A_ABST
Patent Text Reader

Abstract

The invention discloses a meta-learning-based small sample pharmacokinetic property prediction method and related equipment, and relates to the field of pharmacokinetics. The method comprises the following steps: acquiring multi-dimensional feature data of candidate compounds, wherein the multi-dimensional feature data comprises atomic-scale features, molecular descriptors and molecular structure features; inputting the multi-dimensional feature data into a preset pharmacokinetic property prediction model to obtain a preliminary prediction result, the pharmacokinetic property prediction model being pre-trained through meta-learning based on small samples, the model comprising a first machine learning model and a second machine learning model, the preliminary prediction result comprises a first prediction result output by the first machine learning model and a second prediction result output by the second machine learning model; and determining a pharmacokinetic property prediction result of the candidate compound according to the first prediction result and the second prediction result. According to the method, the problem of low accuracy of a pharmacokinetic property prediction result can be relieved in a small sample scene of early drug research and development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pharmacokinetics, and particularly to a method and related device for predicting pharmacokinetic properties based on meta-learning with small samples. Background Art

[0002] In the field of drug research and development, the evaluation of pharmacokinetic properties is a key link in determining whether a candidate compound can enter clinical trials. The ADMET characteristics involved are directly related to the bioavailability, target tissue concentration, and potential safety risks of drugs in the human body. Among them, the ADMET characteristics mainly include characteristics such as absorption, distribution, metabolism, excretion, and toxicity.

[0003] Traditional experimental evaluation methods rely on a large number of in vitro experimental tests, with a long experimental period and high costs. With the rapid development of computer-aided prediction technology, machine learning and artificial intelligence technologies are increasingly used in the prediction of drug pharmacokinetic properties. In related technologies, currently, the mainly used is the conventional QSAR (Quantitative Structure-Activity Relationship) model.

[0004] However, the QSAR model relies on a large amount of labeled data and is difficult to handle small-sample scenarios in early drug development, resulting in relatively large deficiencies in the accuracy of the prediction results of related technologies for pharmacokinetic properties. Summary of the Invention

[0005] Aiming at the above technical problems and deficiencies, the purpose of the present invention is to provide a method and related device for predicting pharmacokinetic properties based on meta-learning with small samples, which can alleviate the problem of low accuracy of the prediction results of pharmacokinetic properties in small-sample scenarios of early drug development.

[0006] To achieve the above purpose, in a first aspect, the present invention provides a method for predicting pharmacokinetic properties based on meta-learning with small samples, including: obtaining multi-dimensional feature data of a candidate compound, where the multi-dimensional feature data includes atomic-level features, molecular descriptors, and molecular structure features; inputting the multi-dimensional feature data into a preset pharmacokinetic property prediction model to obtain a preliminary prediction result, where the pharmacokinetic property prediction model is pre-trained based on small samples through meta-learning, the pharmacokinetic property prediction model includes a first machine learning model and a second machine learning model, and the preliminary prediction result includes a first prediction result output by the first machine learning model and a second prediction result output by the second machine model; determining the pharmacokinetic property prediction result of the candidate compound according to the first prediction result and the second prediction result.

[0007] The present invention adopts the above method, and by obtaining multi-dimensional feature data of candidate compounds covering atomic-level features, molecular descriptors, and molecular structure features, it provides rich information for prediction. These data are input into a pharmacokinetic property prediction model pre-trained by meta-learning based on small samples. This model includes a first machine learning model and a second machine learning model, which output a first prediction result and a second prediction result respectively. By virtue of the advantages of meta-learning in small samples, the model can effectively mine data features and rules. Finally, by integrating the two prediction results, fully integrating the prediction information of different models, reducing the error caused by the limitations of a single model, the pharmacokinetic property prediction result of the candidate compound can be determined more accurately, and the accuracy of prediction in the small sample scenario is significantly improved.

[0008] Optionally, in some embodiments, determining the pharmacokinetic property prediction result of the candidate compound according to the first prediction result and the second prediction result includes: obtaining the first standard deviation of the first prediction result and the second standard deviation of the second prediction result; determining a weight coefficient based on the first standard deviation and the second standard deviation; Determining the pharmacokinetic property prediction result according to the first prediction result, the second prediction result, and the weight coefficient.

[0009] Adopting the technical solution of the above embodiment, by introducing the standard deviation to quantify the uncertainty of the first prediction result and the second prediction result, a reasonable weight coefficient can be assigned to each prediction result. This method can effectively combine the advantages of the two machine learning models, improving the accuracy and reliability of the pharmacokinetic property prediction. By determining the weight coefficient based on the standard deviation, it can be ensured that the model with lower uncertainty in the prediction results has a greater influence, thereby enhancing the credibility of the final prediction result.

[0010] Optionally, in some embodiments, the weight coefficient includes a first weight coefficient and a second weight coefficient; determining the weight coefficient based on the first standard deviation and the second standard deviation includes: determining a first confidence parameter according to the first standard deviation and a second confidence parameter according to the second standard deviation; respectively normalizing the first confidence parameter and the second confidence parameter to obtain the first weight coefficient and the second weight coefficient.

[0011] Adopting the technical solution of the above embodiment, it clarifies the method for determining the weight coefficient. By converting the standard deviation into a confidence parameter and performing normalization processing, the calculation of the weight coefficient is made more scientific and reasonable. This normalization processing method based on confidence can effectively balance the prediction results of the two models, ensuring that the prediction stability of each model is fully considered during the fusion process. This method not only improves the accuracy of the prediction results but also enhances the adaptability and robustness of the model on different data sets.

[0012] Optionally, in some embodiments, the training process of the pharmacokinetic property prediction model includes: obtaining sample data containing a plurality of multi-dimensional feature data and experimentally determined pharmacokinetic property data; dividing the sample data into a training set and a validation set; based on the training set, training a machine learning model through a meta-learning pre-training mechanism to obtain a pharmacokinetic property prediction model; validating the output result of the pharmacokinetic property prediction model through the validation set to evaluate the performance of the pharmacokinetic property prediction model.

[0013] Adopting the technical solution of the above embodiment, a systematic training method for a pharmacokinetic property prediction model is provided. By obtaining sample data and dividing it into a training set and a validation set, it can ensure that the model makes full use of the data during the training process, and at the same time objectively evaluate the model performance through the validation set. The training process based on the meta-learning pre-training mechanism enables the model to quickly adapt to the small-sample scenario, improves the generalization ability and prediction performance of the model, and provides a reliable model basis for pharmacokinetic property prediction.

[0014] Optionally, in some embodiments, based on the training set, training a machine learning model through a meta-learning pre-training mechanism to obtain a pharmacokinetic property prediction model includes: dividing the training set into a support set and a query set; pre-training a preset deep neural network through a meta-learning mechanism based on the support set to obtain a pre-trained neural network; adjusting the pre-trained neural network according to the query set to obtain a first machine learning model; in the case of training a preset gradient boosting decision model through the training set to obtain a second machine learning model, constructing a pharmacokinetic property prediction model based on the first machine learning model and the second machine learning model.

[0015] Adopting the technical solution of the above embodiment, it provides how to construct a pharmacokinetic property prediction model through a meta-learning mechanism. By dividing the training set into a support set and a query set, it can effectively simulate the small-sample learning scenario, enabling the pre-trained deep neural network to quickly adapt to a small amount of data. Combining the training of the gradient boosting decision tree model, the constructed integrated prediction model can give full play to the advantages of the deep neural network and the decision tree model, improve the accuracy and robustness of the prediction, and provide a powerful tool for pharmacokinetic property prediction in drug research and development.

[0016] Optionally, in some embodiments, inputting the multi-dimensional feature data into a preset pharmacokinetic property prediction model to obtain a preliminary prediction result includes: performing feature fusion processing on the multi-dimensional feature data through a multi-head attention mechanism to obtain high-order molecular representation data; inputting the high-order molecular representation data into the pharmacokinetic property prediction model to obtain a preliminary prediction result.

[0017] Adopting the technical solution of the above embodiment and introducing the multi-head attention mechanism to perform feature fusion processing on multi-dimensional feature data can effectively integrate information from multiple aspects such as atomic-level features, molecular descriptors, and molecular structure features. By generating high-order molecular representation data, this method can capture the complex relationships between molecular features and provide a richer information basis for pharmacokinetic property prediction. This not only improves the accuracy of the prediction model but also enhances the model's ability to understand and utilize molecular features.

[0018] Optionally, in some embodiments, through the multi-head attention mechanism, multi-dimensional feature data is subjected to feature fusion processing to obtain high-order molecular representation data, including: respectively inputting the multi-dimensional feature data into a linear projection layer to obtain query vectors, key vectors, and value vectors; generating a similarity matrix based on the query vectors and the key vectors; generating an attention weight matrix based on the similarity matrix; performing weighted summation on the value vectors according to the attention weight matrix to obtain multiple subspace feature vectors; and performing multi-layer non-linear transformation on the subspace feature vectors through the ReLU function to obtain high-order molecular representation data.

[0019] Adopting the technical solution of the above embodiment, the specific implementation steps of the multi-head attention mechanism are elaborated in detail. From generating query vectors, key vectors, and value vectors in the linear projection layer, to performing feature fusion through the similarity matrix and the attention weight matrix, and finally performing non-linear transformation using the ReLU function, this series of operations can effectively extract and integrate key information in multi-dimensional feature data. The high-order molecular representation data obtained in this way can not only reflect the complex characteristics of molecules but also improve the performance and generalization ability of the pharmacokinetic property prediction model.

[0020] In a second aspect, an embodiment of the present invention provides an electronic device, including: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation manner in the first aspect or the second aspect.

[0021] In a third aspect, the present invention provides a computer-readable storage medium, including instructions, when the above instructions run on the above electronic device, enabling the above electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation manner in the first aspect or the second aspect.

[0022] In a fourth aspect, the present invention provides a computer program product containing instructions, when the above computer program product runs on the above electronic device, enabling the above electronic device to execute the method described in the first aspect or the second aspect, and any possible implementation manner in the first aspect or the second aspect.

[0023] Understandably, the electronic device provided in the second aspect above, the storage medium provided in the third aspect, and the computer program product provided in the fourth aspect are all used to execute the method provided by the present invention. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method and will not be elaborated here.

[0024] One or more technical solutions provided by the present invention have at least the following technical effects or advantages: 1. Solving the problem of small-sample prediction accuracy: By combining multi-dimensional feature data, meta-learning mechanism, and multi-model fusion strategy, the present invention effectively solves the problem of low prediction accuracy of pharmacokinetic properties in small-sample scenarios during early drug development. Specifically, collecting multi-dimensional feature data provides comprehensive information for the model, meta-learning enables the model to quickly adapt to and extract feature patterns from small-sample data, and multi-model fusion reduces the bias of a single model, improving prediction accuracy and stability, providing reliable support for early drug development.

[0025] 2. Improving the reliability of prediction results: By determining the weight coefficient according to the standard deviation of the prediction results, the present invention realizes the reasonable weighted fusion of the prediction results of different machine learning models, significantly improving the reliability and credibility of the prediction results of pharmacokinetic properties. Specifically, using the standard deviations of the first prediction result and the second prediction result, the first confidence parameter and the second confidence parameter are determined respectively, and the corresponding weight coefficients are obtained through normalization processing to ensure that the prediction stability of each model is fully considered during the prediction process, making the final prediction result more credible.

[0026] 3. Enhancing the model's ability to understand molecular features: By adopting the multi-head attention mechanism to perform feature fusion processing on multi-dimensional feature data and generating high-order molecular representation data, the present invention enhances the model's ability to understand and utilize molecular features. Specifically, the multi-head attention mechanism enables the model to capture the complex relationships between molecular features through the generation and interaction of query vectors, key vectors, and value vectors, and further extracts high-order features through multi-layer non-linear transformation using the ReLU function, improving the model's comprehensive understanding ability of molecular features and the performance of pharmacokinetic property prediction. Description of the Drawings

[0027] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments that conform to the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings: Figure 1It is a flowchart of a small-sample pharmacokinetic property prediction method based on meta-learning according to an embodiment of the present invention; Figure 2 It is an implementation flowchart of small-sample pharmacokinetic property prediction based on meta-learning according to an embodiment of the present invention; Figure 3 It is a flowchart of the meta-learning mechanism according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the architecture of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0028] The terms used in the following embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. As used in the specification of the present invention, the singular forms "a", "an", "above-mentioned", "the" and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present invention refers to any or all possible combinations including one or more of the listed items.

[0029] Hereinafter, the terms "first" and "second" are only for descriptive purposes, and cannot be understood as implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0030] It should also be noted that, unless otherwise clearly specified and limited, terms such as "set" and "connect" in the embodiments of the present invention should be understood in a broad sense. For example, "connect" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two elements; it can be a wired communication connection or a wireless communication connection. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations. The embodiments of the present invention will be specifically described below.

[0031] Traditional QSAR models face significant limitations in the early stage of drug development. The fact that they rely on a large amount of labeled data for training conflicts with the reality of scarce experimental data in the small-sample scenario, resulting in the model being difficult to fully learn the hidden laws of complex pharmacokinetic properties (such as metabolic rate, bioavailability); at the same time, the extraction of molecular features by traditional methods is limited to a single dimension (such as only relying on molecular descriptors or two-dimensional structure information), and it is unable to effectively integrate atomic-level local features and molecular fingerprint global features, causing the loss of key information and further weakening the prediction accuracy.

[0032] Therefore, the present invention provides a technical solution for predicting pharmacokinetic properties with few samples based on meta-learning, which can effectively cope with the scenario of few samples with extremely limited experimental data in the early stage of drug research and development. Through the meta-learning framework, cross-task pre-training is performed on the first machine learning model (such as a deep neural network) and the second machine learning model (such as a gradient boosting tree), enabling them to quickly capture the cross-dimensional correlation laws among atomic-level features, molecular descriptors, and molecular structure features on a small number of support set samples, and making up for the problem of insufficient feature learning caused by insufficient data in traditional QSAR models. At the same time, by combining the first prediction result and the second prediction result, the prediction fluctuations caused by the deviation of a single model due to few samples can be better suppressed, and adaptive compensation for prediction deviation can be achieved in the scenario of few samples with scarce data, thereby improving the accuracy and stability of pharmacokinetic property prediction.

[0033] The following will specifically describe Figure 1 A method for predicting pharmacokinetic properties with few samples based on meta-learning provided by an embodiment of the present invention includes the following steps: Step 101, obtain multi-dimensional feature data of a candidate compound.

[0034] Among them, the multi-dimensional feature data includes atomic-level features, molecular descriptors, and molecular structure features.

[0035] The atomic-level features include but are not limited to atomic number, connectivity, formal charge, hybridization type, and aromaticity.

[0036] Specifically, the atomic number represents the number of protons in the atomic nucleus, determining the element type and its chemical properties; the connectivity refers to the number of other atoms directly connected to the atom in the molecule, reflecting the chemical bond distribution state; the formal charge is the charge distribution value of the atom in the molecule, used to describe the contribution of electron distribution to stability; the hybridization type describes the mixing mode of atomic orbitals (such as sp³, sp²), determining the molecular spatial configuration and reactivity; the aromaticity characterizes the delocalization state of π electrons in the cyclic conjugated system, affecting the molecular metabolic stability and binding characteristics.

[0037] In this embodiment, the atomic-level features can be obtained through the RDKit (an open-source cheminformatics toolkit that provides functions for molecular structure parsing, descriptor calculation, and fingerprint generation) tool, parsing the SMILES (Simplified Molecular Input Line Entry System) string or molecular structure file of the candidate compound, and extracting the atomic number (number of protons), connectivity (number of adjacent atoms), formal charge (charge distribution value), hybridization type (such as sp 3 、sp 2Properties such as orbital configuration and aromaticity (whether participating in the conjugated ring system) are used to form atomic-level features.

[0038] Molecular descriptors include, but are not limited to, topological polar surface area, number of hydrogen bond donors, number of hydrogen bond acceptors, molecular weight, number of rotatable bonds, number of ring structures, maximum atomic partial charge, minimum atomic partial charge, and total number of heavy atoms (non-hydrogen atoms).

[0039] Specifically, the topological polar surface area (TPSA) quantifies the total sum of polar regions on the molecular surface and predicts the ease of a compound passing through the biological membrane barrier; the number of hydrogen bond donors is the number of acidic hydrogen atoms (such as -OH, -NH) in the molecule that can provide hydrogen bonds, which affects the binding strength between the drug and the target; the number of hydrogen bond acceptors refers to the number of atoms with lone pairs of electrons (such as O, N) in the molecule that can accept hydrogen bonds, which determines solubility and transport efficiency; the molecular weight is the sum of the relative atomic masses of all atoms in the molecule, which is related to the diffusion rate and in vivo distribution of the drug; the number of rotatable bonds characterizes the number of single bonds in the molecule that can rotate freely, reflecting the conformational flexibility and metabolic stability of the molecule; the number of ring structures describes the number of closed-ring structures in the molecule, which affects molecular rigidity and the spatial arrangement of pharmacophores; the maximum atomic partial charge represents the positive extreme value of the atomic charge distribution in the molecule, revealing the reactivity of the electron-rich region; the minimum atomic partial charge is the negative extreme value of the atomic charge distribution, indicating the tendency of nucleophilic attack in the electron-deficient region; the total number of heavy atoms is the number of all atoms in the molecule except hydrogen, which measures the complexity of the molecular structure and the feasibility of synthesis.

[0040] In this embodiment, the acquisition of molecular descriptors also relies on the RDKit program. Built-in calculation functions are called to batch generate physicochemical parameters such as topological polar surface area (TPSA), number of hydrogen bond donors (-OH / -NH group numbers), number of hydrogen bond acceptors (number of lone pairs of electrons of O / N atoms), molecular weight (sum of atomic weights), number of rotatable bonds (single bond freedom), number of ring structures (number of closed rings), maximum / minimum atomic partial charges (extremes of charge distribution), and total number of heavy atoms (non-hydrogen atom count).

[0041] Molecular structure features include ECFP (Extended Connectivity Fingerprint) fingerprints and MACCS (Molecular ACCess System) fingerprints based on substructure matching.

[0042] Specifically, the ECFP fingerprint is a molecular characterization method based on the enumeration of atomic neighborhood substructures. By iteratively expanding the chemical environment of each atom (for example, when the radius = 3, it covers adjacent atoms within the range of triple bonds), each substructure is hashed into a binary bit string of a fixed length, which is used to describe local structural features such as functional groups and ring systems in the molecule. Its core advantage is the ability to dynamically capture molecular diversity without predefining substructures.

[0043] The MACCS fingerprint is a molecular fingerprint method based on multiple predefined key pharmacophore substructures. It characterizes the presence of specific functional groups (such as benzene rings, carboxylic acid groups) in a molecule through binary coding, and is used to rapidly evaluate the structural characteristics related to the binding of drug molecules to targets and metabolism.

[0044] In this embodiment, the acquisition of molecular structure characteristics uses the fingerprint generation module of RDKit. Based on the ECFP algorithm (expansion radius = 3, bit length = 2048), it dynamically encodes the local substructures of the molecule, and at the same time matches 166 predefined pharmacophore substructures of MACCS (such as benzene rings, carboxylic acid groups) to generate binary fingerprints to represent the global topology and key functional group information.

[0045] Step 102, input the multi-dimensional feature data into a preset pharmacokinetic property prediction model to obtain a preliminary prediction result; among them, the pharmacokinetic property prediction model is pre-trained by meta-learning based on a small sample. The pharmacokinetic property prediction model includes a first machine learning model and a second machine learning model. The preliminary prediction result includes a first prediction result output by the first machine learning model and a second prediction result output by the second machine model.

[0046] In this embodiment, the selection of the first machine learning model and the second machine learning model is crucial. Different types of machine learning algorithms are suitable for different types of data and prediction tasks. For example, the decision tree model is widely used for its strong interpretability and good ability to handle non-linear relationships; the support vector machine performs well in handling high-dimensional data and can effectively find the optimal classification hyperplane between data; while the neural network has powerful learning ability in handling complex feature interactions and large-scale data.

[0047] When constructing the pharmacokinetic property prediction model, the characteristics of the multi-dimensional feature data of the candidate compound and the specific pharmacokinetic property prediction target can be considered comprehensively, and factors such as the accuracy, interpretability, and computational complexity of the model can be taken into account to select a suitable algorithm to construct the first machine learning model and the second machine learning model. For example, if the feature data has obvious hierarchical structure and non-linear relationship, the neural network may be a good choice; while if a preliminary prediction result with good interpretability is needed quickly, the decision tree model may be more suitable.

[0048] After determining the model type, the pre-training process of the model is based on small samples and implemented through meta-learning. The core idea of meta-learning is to enable the model to quickly adapt to new similar tasks, especially in the case of small samples, by learning multiple related tasks. Specifically, in the pre-training stage, a large amount of data on different small sample tasks related to pharmacokinetic property prediction is collected. These data may involve different types of compounds and their corresponding pharmacokinetic properties (such as kinetic solubility, plasma protein binding rate data of different drugs, etc.). Then, these task data are input into a prediction model composed of the first and second machine learning models. By optimizing the model parameters, the model can quickly learn general feature patterns and rules on these multiple tasks. For example, in each iteration process, the model will first be trained on a part of the tasks to obtain a preliminary model update; then it will be verified on another part of the tasks, and the model parameters will be adjusted according to the verification results, enabling the model to effectively transfer and share knowledge between multiple tasks. After multiple such iterative trainings, the model gradually learns how to use limited small samples to quickly and accurately predict pharmacokinetic properties.

[0049] In this step, after collecting the multi-dimensional feature data of the candidate compound and the experimentally determined pharmacokinetic property data, these data need to be preprocessed to meet the input requirements of the model. The preprocessing process may include steps such as data cleaning (removing noise, outliers, etc.), feature scaling (scaling feature data with different dimensions to the same scale), and feature encoding (converting categorical features into numerical features). For example, for the functional group information in the molecular structure features, it may need to be converted into the one-hot encoding form; for numerical features such as atomic electronegativity in the atomic-level features, normalization processing is required. After completing the data preprocessing, the processed multi-dimensional feature data is input into the pre-trained pharmacokinetic property prediction model.

[0050] After the model receives the input data, the first machine learning model and the second machine learning model will process and analyze the data in parallel. Each model has its unique parameters and structure, capturing the relationship between the feature data and the pharmacokinetic properties from different perspectives and levels. For example, the first machine learning model may focus on mining rules from the local information of atomic-level features and molecular structure features, while the second machine learning model may pay more attention to the correlation between the overall physicochemical properties reflected by molecular descriptors and the pharmacokinetic properties.

[0051] After their respective calculation and reasoning processes, the first machine learning model outputs the first prediction result, and the second machine learning model outputs the second prediction result. These two prediction results respectively represent the preliminary judgments of the models on the pharmacokinetic properties of the candidate compound based on their own parameters and structures, and there may be differences in terms of numerical values, trends, or predictions of certain specific attributes.

[0052] The generation of the preliminary prediction results is a key step for the pharmacokinetic property prediction model to quickly make judgments in small-sample scenarios. Through the pre-training of meta-learning, the model can extract effective information from limited small-sample data and use two different machine learning models to make predictions from multiple perspectives, providing a basis for further determining the prediction results of the pharmacokinetic properties of the candidate compound.

[0053] Step 103: Determine the prediction result of the pharmacokinetic properties of the candidate compound according to the first prediction result and the second prediction result.

[0054] Specifically, various methods can be used to integrate the first prediction result and the second prediction result. One way is to take the average of the two prediction results as the final prediction result. Or, a weighted average of the prediction results of the two models can be performed, where the weights can be determined based on the performance of the models on the validation set, and a better-performing model is given a higher weight. In addition, the differences between the two prediction results can be used to decide whether further experimental verification is needed or whether the model parameters need to be adjusted.

[0055] More complex methods may include constructing a fusion model that takes the first prediction result and the second prediction result as input features and learns how to better combine these two prediction results through training to obtain a more accurate prediction result of the pharmacokinetic properties. For example, ensemble learning methods such as random forest or gradient boosting tree can be used to ensemble the two prediction results. In classification problems, a voting mechanism can be adopted. If the prediction results of the two models are the same, it is directly determined as the property result; if the results are different, other auxiliary information can be combined for further analysis.

[0056] In practical applications, which method to choose depends on the specific data, the performance of the models, and the requirements of the prediction task. The ultimate goal is to utilize the prediction results of the two models to improve the accuracy and reliability of the prediction and provide more accurate guidance for drug development.

[0057] In the embodiment of the present invention, the above-mentioned small-sample pharmacokinetic property prediction method based on meta-learning is adopted. By obtaining multi-dimensional feature data of candidate compounds, which comprehensively covers atomic-level features, molecular descriptors, and molecular structure features, it provides a rich information basis for the prediction of pharmacokinetic properties. Then, these multi-dimensional feature data are input into a preset pharmacokinetic property prediction model, which is pre-trained through meta-learning on small samples and can effectively address the challenge of limited data volume in early drug development. The model includes a first machine learning model and a second machine learning model. By respectively outputting a first prediction result and a second prediction result, and then integrating these two prediction results to determine the prediction result of the pharmacokinetic properties of the candidate compound, the accuracy and reliability of the prediction are effectively improved.

[0058] Compared with the traditional single-model prediction, this method based on meta-learning and multi-model fusion in this embodiment can better capture and integrate information from different feature dimensions in the case of small samples, reduce the prediction risk, provide more powerful support for the early screening and decision-making in drug development, accelerate the drug development process, and reduce the development cost.

[0059] In some embodiments, step 102 may specifically include the following steps: (1) Through the multi-head attention mechanism, perform feature fusion processing on the multi-dimensional feature data to obtain high-order molecular representation data.

[0060] Specifically, first perform feature fusion processing through the multi-head attention mechanism. The multi-head attention mechanism is a powerful tool that can capture the complex relationships between different features. Especially when dealing with molecular data, it can capture the interactions between atoms, molecular descriptors, and structural features. Specifically, the multi-dimensional feature data are organized into a vector sequence, and each vector represents a feature dimension. The multi-head attention mechanism calculates in parallel through multiple attention heads. Each attention head independently calculates the query, key, and value matrices, and then recombines the information of different feature dimensions through attention scores. These attention scores represent the correlation intensity between various features, enabling the model to automatically learn which feature combinations are more important for the prediction of pharmacokinetic properties. After being processed by the multi-head attention mechanism, the obtained high-order molecular representation data integrates the comprehensive information of multi-dimensional features and provides a richer representation for subsequent prediction.

[0061] Among them, the high-order molecular characterization data is a comprehensive feature vector generated by fusing atomic-level features (such as hybridization type), molecular descriptors (such as TPSA), and structural fingerprints (such as ECFP) through a multi-head attention mechanism. It encodes the cross-dimensional interactions within the molecule (such as the association between hydrophobic groups and metabolic enzyme binding sites) and non-linear combination rules, and is used as the input of the prediction model to improve the accuracy of pharmacokinetic property analysis.

[0062] In some embodiments, this step may specifically include the following steps: (1-1) Input the multi-dimensional feature data into the linear projection layer respectively to obtain query vectors, key vectors, and value vectors.

[0063] Among them, the linear projection layer is a neural network layer used to linearly transform the input data into different dimensional spaces. It realizes the linear combination and projection of the input features by learning the product of the input data and the weight matrix plus the bias term. In fields such as natural language processing and computer vision, the linear projection layer is often used to map high-dimensional data to low-dimensional spaces to reduce complexity and extract key features, or project low-dimensional features to high-dimensional spaces to enhance feature expression capabilities. The linear projection layer can achieve linear transformations of features in different subspaces, facilitating subsequent calculation of the correlation between features and feature fusion.

[0064] This embodiment can be implemented through matrix multiplication. Each feature dimension data is multiplied by the corresponding weight matrix and then added with the bias term, thereby linearly transforming the original feature data into different semantic spaces to obtain query vectors (Q), key vectors (K), and value vectors (V), which are used for subsequent attention calculation processes.

[0065] (1-2) Generate a similarity matrix based on the query vectors and the key vectors.

[0066] Specifically, the query vector matrix and the key vector matrix can be multiplied to obtain a matrix with the size of the number of heads multiplied by the feature dimension. Each element in it represents the dot product similarity between the query vector and the key vector, reflecting the degree of correlation between different features.

[0067] For example, perform the matrix multiplication of the query vector Q and the key vector K, calculate the original similarity score matrix S = QK^T, and then divide S by the scaling factor √d_k (d_k = 64) to obtain the scaled similarity matrix S_scaled, avoiding the vanishing Softmax gradient caused by the large numerical value of the dot product result due to high-dimensional features.

[0068] (1-3) Generate an attention weight matrix based on the similarity matrix.

[0069] Specifically, each element of the similarity matrix is scaled by dividing it by the square root of the key vector dimension, and then the softmax function is applied to normalize the elements in each row into a probability distribution form. The resulting matrix is the attention weight matrix, and the element values represent the degree to which each feature should be focused on during the feature fusion process.

[0070] For example, apply the Softmax function to S_scaled row by row to convert the numerical values in each row into an attention weight matrix A in the form of a probability distribution, so that the sum of the attention degrees of each feature position to other positions is 1, highlighting key feature interactions (such as the strong correlation between the number of hydrogen bond donors and the polar surface area).

[0071] (1-4) Weighted sum the value vectors according to the attention weight matrix to obtain multiple subspace feature vectors.

[0072] Specifically, perform a weighted sum operation on the value vectors according to the attention weight matrix, multiply the attention weight matrix and the value vector matrix, and obtain multiple subspace feature vectors. Each vector is a comprehensive feature representation that fuses multi-dimensional feature data in the corresponding subspace. These subspace feature vectors retain the valuable feature combination information in the input feature data for the pharmacokinetic property prediction task.

[0073] For example, multiply the attention weight matrix A by the value vector V to obtain the output head_i = AV_i of each attention head. Subsequently, concatenate the outputs of 4 heads along the feature dimension (4×256 = 1024 dimensions), and then fuse the multi-head information through a fully connected layer (dimension 1024→1024).

[0074] (1-5) Based on each subspace feature vector, perform multi-layer non-linear transformation through the ReLU (Linear rectification function) function to obtain high-order molecular representation data.

[0075] Specifically, multi-layer non-linear transformation can be performed through the ReLU function in a multi-layer perceptron (MLP). The ReLU function can introduce non-linear factors, enabling the model to learn complex patterns in the data. After multi-layer non-linear transformation, the output feature vector is the high-order molecular representation data, which integrates the comprehensive information of multi-dimensional feature data and can be more effectively used for the prediction and analysis of pharmacokinetic properties.

[0076] For example, the subspace feature vectors are input into a three-layer fully connected network (1024→512→256→1), the ReLU activation function is used to introduce non-linearity, and the original feature information is retained through residual connections and layer normalization (LayerNorm). Finally, high-order molecular representation data is output, which encodes the cross-modal synergy of atomic-physical-chemical-structural features (such as the spatial matching pattern between hydrophobic groups and metabolic enzyme binding sites).

[0077] In some embodiments, the ReLU function specifically includes:

[0078]

[0079]

[0080] Among them, x i is the input value of the i th subspace feature vector.

[0081] e task represents the task embedding vector, which is used to encode the semantic information of the current pharmacokinetic task (such as metabolic rate prediction).

[0082] s j is the subspace context vector, which represents the molecular structure correlation features extracted by the j th attention head.

[0083] W α represents a learnable parameter matrix, which is used to project the concatenated feature vectors into the scaling factor space.

[0084] W β represents a learnable parameter matrix, which is used to project the concatenated feature vectors into the offset factor space.

[0085] α ij represents the dynamic scaling factor, which is generated by the Sigmoid function and controls the activation intensity of the input features (range 0~1).

[0086] β ij represents the dynamic offset factor, which is generated by the hyperbolic tangent function and adjusts the activation threshold (range -1~1).

[0087] σ represents the Sigmoid activation function, which is used to map the input into a probability distribution form of scaling weights.

[0088] tanhRepresents the hyperbolic tangent function, which is used to generate non-linear offsets to enhance the model's expressive power.

[0089] || denotes vector concatenation.

[0090] This embodiment adopts the above-mentioned ReLU function, through the task embedding vector e task , enabling the activation function to adaptively adjust the slope ( α ij ) and offset ( β ij ) for different pharmacokinetic tasks (such as absorption, metabolism, toxicity). For example, in the metabolism task, enhance the activation intensity of features related to hydrophobic groups ( α ij increase), and in the toxicity task, suppress the interference of specific atomic charges ( β ij decrease).

[0091] Moreover, the subspace context vector s j is introduced to capture the interaction patterns of different subspaces (such as ECFP, MACCS, atomic features) in multi-head attention, and strengthen the information transmission of key subspaces during activation. For example, if the MACCS subspace detects a benzene ring structure (a key feature of toxicity), the α ij of the corresponding activation channel will automatically increase to 0.9 to amplify signal propagation.

[0092] Meanwhile, combined with the offset term β ij and Sigmoid gating, the problem of neuron death in traditional ReLU is avoided, ensuring the gradient stability of sparse features (such as ECFP binary fingerprints).

[0093] When designing the above-mentioned ReLU function in this embodiment, the following aspects are considered: 1. Task adaptability: By introducing the task embedding vector e task , the activation function can perceive the current pharmacokinetic task (such as absorption, metabolism, toxicity), and dynamically adjust parameters to adapt to the different requirements of different tasks for molecular features (such as the metabolism task requires strengthening hydrophobic features, and the toxicity task requires attention to charge distribution).

[0094] 2. Subspace cooperation: Concatenate the subspace context vector s j with the input features x i , enabling the activation process to integrate the cross-dimensional correlations captured by the multi-head attention mechanism (such as the synergistic effect between ECFP fingerprint bits and atomic charges), and enhancing the analytical ability of complex molecular structures.

[0095] 3. Dynamic gating mechanism: Use Sigmoid to generate a scaling factor α ij , to achieve soft constraints on the feature activation intensity and avoid information loss caused by the hard truncation of the traditional ReLU; at the same time, use the tanh function to generate an offset factor β ij , introduce non-linear threshold adjustment, and improve the model's expression ability for sparse features (such as MACCS binary fingerprints).

[0096] 4. Anti-overfitting design: Through the projection matrices W α and W β , map high-dimensional features to a low-dimensional parameter space, suppress small-sample noise interference, and at the same time, the parameter generation process is jointly constrained by task embedding and subspace context to avoid being dominated by a single feature.

[0097] 5. Gradient stability: The smooth gradient characteristics of the Sigmoid function and the tanh function can alleviate the problem that the gradient of the traditional ReLU is zero in the negative region, and ensure the stable update of the parameters W α and W β during backpropagation.

[0098] 6. Biological interpretability: α ij and β ij can be traced back to molecular structure features (such as α ij high weights corresponding to the logP-related subspace), assisting researchers in understanding the degree of attention of the model to specific pharmacophores.

[0099] The ReLU function in this embodiment is regulated by both task-driven and data-driven methods, taking into account prediction accuracy and generalization ability in small-sample scenarios. Experimental verification shows that its error is reduced by 28.5% and the convergence speed is increased by 40% compared with the standard ReLU in cross-task pharmacokinetic prediction.

[0100] (2)Input the high-order molecular characterization data into the pharmacokinetic property prediction model to obtain a preliminary prediction result.

[0101] Among them, the pharmacokinetic property prediction model is usually a deep learning model, such as a deep neural network. The high-order molecular characterization data is used as the input of the model and is further processed and abstracted through the multi-layer neural network structure inside the model. These layers include fully connected layers, activation function layers, and possible regularization layers (such as Dropout layers), etc., for extracting higher-level feature representations and finally mapping to the predicted values of pharmacokinetic properties.

[0102] Specifically, the high-order molecular characterization data are respectively input into a first machine learning model (deep neural network DNN) and a second machine learning model (XGBoost): DNN extracts non-linear features layer by layer through three fully connected layers and a ReLU activation function, and outputs pharmacokinetic properties (such as the predicted value of human plasma protein binding rate); XGBoost constructs a gradient boosting tree based on the high-order characterization data, analyzes the feature importance through the tree node splitting rule (such as the contribution weight of ECFP fingerprint bits to metabolism prediction), generates a regression-type prediction result (such as the clearance value), and finally, the two types of prediction results are used as preliminary results and input into the dynamic fusion module to obtain the preliminary prediction result.

[0103] In some embodiments, step 103 specifically includes: S31, obtaining a first standard deviation of the first prediction result and a second standard deviation of the second prediction result.

[0104] For the first prediction result, the first standard deviation is obtained by calculating the deviation degree of each index feature value in the first prediction result from the average value. Similarly, the second standard deviation of the second prediction result is calculated.

[0105] A large standard deviation means that the prediction result fluctuates greatly and the uncertainty is high; on the contrary, a small standard deviation indicates that the predicted values are relatively concentrated and the result is more reliable.

[0106] Specifically, for example, if the first machine learning model uses a deep neural network (DNN) and the second machine learning model uses a gradient boosting decision model (XGBoost), by statistically analyzing the output value distribution of the deep neural network in multiple predictions or using the built-in confidence estimation function of the model, the standard deviation σ1 (the first standard deviation) of its prediction result is calculated, which reflects the uncertainty of the model's prediction of the current compound (such as obtaining the prediction value fluctuation range based on Bootstrap sampling or Dropout Monte Carlo simulation).

[0107] At the same time, for the XGBoost model, by using the information gain during its tree node splitting or the historical statistics based on the test set residuals, the standard deviation σ2 (the second standard deviation) of the second prediction result is derived, which characterizes the sensitivity of the tree model to the current input features.

[0108] S32, determining the weight coefficient based on the first standard deviation and the second standard deviation.

[0109] Specifically, the corresponding relationship between each standard deviation and the weight coefficient can be pre-constructed, for example, constructing a look-up table of the standard deviation and the weight coefficient, and the weight coefficient can be quickly determined by looking up the table.

[0110] The reciprocal of the standard deviation can also be used as the basis for the weights. First, calculate the reciprocals of the first standard deviation and the second standard deviation respectively, and then normalize these two reciprocals so that their sum is 1, thereby obtaining two weight coefficients.

[0111] For example, if the reciprocal of the first standard deviation is a and the reciprocal of the second standard deviation is b, the weight coefficients are a / (a + b) and b / (a + b) respectively. In this way, the prediction result with a smaller standard deviation will obtain a larger weight coefficient because the corresponding model prediction is more stable and reliable.

[0112] In some embodiments, this step may specifically include: S321. Determine a first confidence parameter according to the first standard deviation and a second confidence parameter according to the second standard deviation.

[0113] Specifically, the following formula can be used for calculation: C1 = 1 / (σ1 + ε), C2 = 1 / (σ2 + ε); Where, C1 represents the first confidence parameter, σ1 represents the first standard deviation, ε is a smoothing term used to prevent the denominator from being 0. C2 represents the second confidence parameter, and σ2 represents the second standard deviation.

[0114] In this embodiment, ε generally takes a relatively small value, such as 10 -6 .

[0115] S322. Normalize the first confidence parameter and the second confidence parameter respectively to obtain a first weight coefficient and a second weight coefficient.

[0116] Specifically, the first confidence parameter can be divided by the sum of the first confidence parameter and the second confidence parameter to obtain the first weight coefficient; the second confidence parameter is divided by the sum of the first confidence parameter and the second confidence parameter to obtain the second weight coefficient.

[0117] The specific calculation formula is as follows: w1 = C1 / (C1 + C2), w2 = C2 / (C1 + C2); Where, w1 represents the first weight coefficient and w2 represents the second weight coefficient.

[0118] S33. Determine the pharmacokinetic property prediction result according to the first prediction result, the second prediction result and the weight coefficients.

[0119] Specifically, the first prediction result and the second prediction result can be multiplied by their corresponding weight coefficients respectively and then added together to obtain the final pharmacokinetic property prediction result.

[0120] Assume that the first prediction result is P1, the weight coefficient is w1, the second prediction result is P2, and the weight coefficient is w2. Then the final prediction result of the pharmacokinetic property P = w1×P1 + w2×P2.

[0121] This method can integrate the prediction advantages of the two models, making the prediction result more accurately reflect the actual pharmacokinetic properties of the candidate compound.

[0122] In this way, the prediction information of the two machine learning models can be fully utilized, improving the accuracy and reliability of the prediction result and providing more powerful support for drug research and development.

[0123] In some embodiments, the training process of the pharmacokinetic property prediction model includes: S201, obtain sample data including a plurality of multi-dimensional feature data and experimentally determined pharmacokinetic property data.

[0124] Among them, the experimentally determined pharmacokinetic property data refers to the specific numerical values or classification results obtained through prior experiments, reflecting the processes and related properties of drug absorption, distribution, metabolism, excretion, etc. in the body. These data include but are not limited to key indicators such as plasma protein binding rate, solubility, clearance rate, half-life, permeability, etc.

[0125] Specifically, this step first needs to collect sample data including a plurality of multi-dimensional features and experimentally determined pharmacokinetic property data from a variety of reliable data sources. The multi-dimensional feature data covers atomic-level features of the candidate compound (such as atomic number, electronegativity, hybridization type, etc.), molecular descriptors (such as lipophilic-hydrophilic partition coefficient, topological polar surface area, number of hydrogen bond donors and acceptors, molecular weight, etc.), and molecular structure features (such as functional groups, number of ring structures, etc.). These feature data are calculated and extracted through open-source tools such as RDKit.

[0126] At the same time, the experimentally determined pharmacokinetic property data (such as the specific numerical values or classification results of absorption, metabolism, excretion, etc. properties) need to be matched one by one with the corresponding multi-dimensional feature data to form complete sample data. These data usually come from existing drug research and development databases, scientific literature, or laboratory experimental results.

[0127] By integrating these multi-dimensional features and experimental data, the sample data can provide rich learning materials for the model, helping the model understand the complex relationships between different features and pharmacokinetic properties, so as to achieve accurate prediction.

[0128] In order to ensure the quality and applicability of the data, the collected data needs to be strictly screened and preprocessed, samples containing errors or missing values need to be removed, and the data needs to be standardized and normalized to eliminate the differences in dimensions and numerical ranges between different features, making the data more suitable for the subsequent model training process.

[0129] S202, dividing the sample data into a training set and a validation set.

[0130] Specifically, a random sampling method can be used to divide the data set to ensure that the sample distribution in the training set and the validation set can roughly reflect the characteristic distribution of the overall sample data. The division ratio needs to be determined based on the total amount of sample data and specific research needs. A common division ratio is 80% of the samples as training sets and 20% of the samples as validation sets. However, when the amount of sample data is small, it may be necessary to use methods such as cross-validation, such as k-fold cross-validation, to divide the data set into k subsets, and take turns using k-1 subsets as training sets and the remaining 1 subset as validation sets, and perform multiple training and validation processes, so as to make better use of limited sample data and obtain more reliable model performance evaluation results.

[0131] Before dividing the data set, the data can also be stratified and sampled to ensure that both the training set and the validation set contain samples of different categories or with different characteristics, so as to improve the representativeness of the model and the accuracy of the evaluation of various situations in the validation stage. For example, the proportion distribution of the training set and the validation set is kept consistent according to the ADMET task category (such as absorption, metabolism, toxicity, etc.), and cross-validation grouping (such as 5-Fold) is enabled for small sample tasks (such as data volume <50) to ensure that each task has representative samples in the training and validation stages.

[0132] S203, based on the training set, the machine learning model is trained through a meta-learning pre-training mechanism to obtain a pharmacokinetic property prediction model.

[0133] Among them, the key to the meta-learning pre-training mechanism is to use multiple related tasks or data sets to train the model, so that it can learn more general and more transferable feature representations and knowledge, so as to quickly adapt to new tasks with small samples and achieve better prediction performance.

[0134] In actual operation, a series of tasks or data sets related to the prediction of pharmacokinetic properties are first prepared. These tasks may involve different types of pharmacokinetic properties (such as absorption, metabolism, toxicity, etc.) or different types of compound data.

[0135] Then, a machine learning model (such as a deep neural network, a decision tree model, etc.) is trained on these multiple tasks. By optimizing the model's parameters, the model can achieve good prediction results on each task.

[0136] During the training process, appropriate optimization algorithms (such as stochastic gradient descent and its variants) and loss functions (such as mean squared error, cross-entropy, etc.) are adopted. The model parameters are adjusted according to the difference between the prediction results of the model on the training set and the true labels.

[0137] At the same time, to prevent the model from overfitting, some regularization techniques (such as L1 / L2 regularization, Dropout, etc.) and early stopping strategies may be adopted.

[0138] After multiple iterative trainings, the model gradually learns the general feature representations and rules that can adapt to various pharmacokinetic property prediction tasks. Finally, a pharmacokinetic property prediction model pre-trained by meta-learning is obtained. This model has the ability to quickly adapt to new tasks and make accurate predictions in small-sample scenarios.

[0139] Specifically, this embodiment can be implemented based on the MAML (Model-Agnostic Meta-Learning) framework: The training set is divided into a meta-training task set. Each task contains a support set (5 samples) and a query set (15 samples). In the training of the first machine learning model (such as a DNN model), it quickly adapts to new tasks through the inner loop (gradient update on the support set), and the outer loop (meta-loss optimization on the query set) improves cross-task generalization; at the same time, the second machine learning model (such as an XGBoost model) adopts task-specific hyperparameter search (such as learning rate, tree depth) and feature importance weighting to enhance the parsing ability of structured descriptors.

[0140] In some embodiments, step S203 may specifically include: S2031, divide the training set into a support set and a query set.

[0141] Among them, the support set is equivalent to the "examples" for the model to learn, which is used to let the deep neural network learn the general features and patterns of pharmacokinetic property prediction. The query set is used to verify and adjust the model during pre-training to ensure that the model can generalize well on new and unseen data. When dividing, the ratio of the support set and the query set is usually determined according to the complexity of the task and the amount of data. Generally, methods such as random sampling or stratified sampling can be used to ensure the representativeness of the two parts of the data.

[0142] For example, 70% of the training samples can be assigned to the support set and 30% to the query set, ensuring that each subset contains compound samples with different pharmacokinetic characteristics.

[0143] A task-driven partitioning strategy can also be adopted: classify the pharmacokinetic tasks (such as absorption, metabolism, toxicity prediction) in the training set according to property labels, randomly select 5-10 samples in each task as the support set (for the model to quickly adapt), and the remaining samples as the query set (for generalization evaluation). Stratified sampling is used to ensure that the support set of each task covers key molecular subclasses (such as benzene ring-containing compounds, nitrogen heterocyclic compounds) to avoid the failure of meta-learning due to structural biases.

[0144] S2032. Pre-train a preset deep neural network through a meta-learning mechanism based on the support set to obtain a pre-trained neural network.

[0145] In this embodiment, the core of meta-learning is to let the model learn how to quickly adapt to new tasks or data distributions. Specifically, during implementation, first initialize the parameters of the deep neural network, and then perform multiple iterative trainings on the support set. In each iteration, the model performs forward propagation based on the samples in the support set, calculates the loss function (such as mean square error or cross entropy), and updates the network parameters through backpropagation and optimization algorithms (such as AdamW or SGD). The difference is that meta-learning processes multiple related tasks or data subsets simultaneously, enabling the model to learn general feature representations across tasks. For example, training can be performed on multiple tasks with different pharmacokinetic properties, allowing the network to learn how to quickly extract effective information from limited samples. After multiple rounds of iteration, the model gradually converges to obtain a pre-trained deep neural network that has the ability to quickly adapt to new tasks.

[0146] Specifically, meta-pre-training can be performed on a deep neural network (DNN) based on the MAML framework: in the inner loop stage, perform a small number of steps (such as 3-5 steps) of gradient descent on the support set of each task to update the DNN parameters to quickly fit the current task. In the outer loop stage, calculate the meta-loss across tasks (such as mean square error) on the query set, and optimize the initial parameters of the DNN through second-order derivatives to enable it to have the ability to quickly adapt across ADMET tasks.

[0147] In some embodiments, MAML adjusts the model parameters through gradient descent with a fixed learning rate on the support set. When the structural differences of the support set samples are large (such as the simultaneous presence of lipophilic molecules and polar molecules), a single learning rate is difficult to balance the adaptation speeds of different structures, resulting in overfitting or underfitting.

[0148] Therefore, based on the above method, this embodiment gives a new solution, skipping the gradient iteration and directly generating dynamic parameters through the molecular structure to solve the bottleneck of the single learning rate caused by structural differences.

[0149] In step S2032, it may specifically further include the following steps: (1) Input the molecular data of the support set: Receive the molecular structure (SMILES or molecular graph) of the support set and pharmacokinetic labels.

[0150] (2) Extract task feature embeddings: Encode the topological relationships between molecular atoms and bonds through a graph attention network (GAT) to generate task-level feature vectors. Among them, the task-level feature vectors are used to characterize the distribution of the task chemical space.

[0151] (3) Dynamic parameter generation and processing: Input the task feature vectors into a multi-layer perceptron (MLP) to generate a model parameter increment matrix, whose dimension is consistent with the global model parameters.

[0152] (4) Calculate task-specific parameters: Add the model parameter increment matrix and the meta-learned global model parameters element-wise to obtain task-specific parameters adapted to the current support set structure.

[0153] Specifically, first extract the topological features of the support set molecular structure through a graph attention network to generate task embedding vectors, and then input them into a hierarchical attention module to calculate the task-level global attention (identifying the core requirements of the current pharmacokinetic task) and the parameter-level local attention (evaluating the sensitivity of each parameter to molecular structure differences) respectively. Multiply the two to generate a dynamic gating vector; this gating vector acts on the parameter increment matrix in a weight scaling manner, preferentially strengthening the update of key parameters related to molecular hydrophobicity and hydrogen bond networks, while suppressing the interference of low-correlation parameters such as atomic numbers (the noise weight decays by 60%). Finally, through residual connection, the gated and scaled increments are superimposed on the global parameters to obtain task-specific parameters, realizing "structure-driven" precise parameter adaptation.

[0154] (5) Adapt the model to predict the query set: Load the task-specific parameters into a deep neural network (DNN) to generate an adapted model that can be directly used for query set prediction.

[0155] Then, step S2033 can be entered.

[0156] In this embodiment, the topological and chemical bond features of the support set molecular structure are extracted through a graph attention network (GAT) to generate task-level embedding vectors to encode the characteristics of the current task's chemical space distribution. Furthermore, a dynamic parameter generation network (DPGN) is used to directly map molecular structure information into model parameter increments, replacing the gradient iteration update mechanism with a fixed learning rate in conventional meta-learning.

[0157] This improvement enables the model to adaptively adjust the parameter generation direction according to the structural differences in the support set samples (such as the mixture of lipophilic molecules and polar molecules), and complete the precise adaptation to complex structure tasks in a single step. This not only avoids the local optimum problem caused by a single learning rate in the gradient descent method (such as overfitting high-hydrophobic samples or underfitting polar features), but also enhances the model's ability to capture the common laws across tasks (such as the non-linear relationship between logP and metabolic rate) through molecular-level structure perception.

[0158] In this embodiment, the parameter optimization process of traditional meta-learning is transformed from "iterative trial and error" to "structure-driven generation", which significantly improves the model adaptation speed (the training time can be reduced by more than 60%) and prediction stability (the cross-task error fluctuation can be reduced by 45%) in the drug development scenario for small-sample tasks, providing an efficient and robust computational tool for early drug molecule screening, effectively shortening the R & D cycle and reducing the experimental cost.

[0159] S2033. Adjust the pre-trained neural network according to the query set to obtain the first machine learning model.

[0160] Among them, the goal of adjustment is to refine the model parameters to make them more accurately applicable to the current pharmacokinetic property prediction task. Specifically, input the samples in the query set into the pre-trained deep neural network, and calculate the error between the prediction result and the true label. According to this error, use the backpropagation algorithm to fine-tune the network parameters again, but the learning rate at this time is usually lower than that in the pre-training stage to avoid destroying the general features that have been learned.

[0161] At the same time, techniques such as early stopping or regularization (such as L2 regularization) can also be used to prevent overfitting. After this process, the obtained first machine learning model has better adaptability and prediction performance on the tasks represented by the support set and the query set, and can more accurately predict the pharmacokinetic properties of candidate compounds.

[0162] Specifically, the pre-trained DNN can be fine-tuned through the query set: freeze the underlying feature extraction layer, only adjust the weight parameters of the fully connected layer, adopt the cosine annealing learning rate strategy (initial learning rate 0.001, 50 epochs) to balance the convergence speed and stability, and at the same time introduce label smoothing regularization (smoothing factor = 0.1) to alleviate small-sample overfitting, and finally output the first machine learning model (DNN).

[0163] S2034. In the case of training the preset gradient boosting decision model through the training set to obtain the second machine learning model, construct a pharmacokinetic property prediction model based on the first machine learning model and the second machine learning model.

[0164] Specifically, the preset gradient boosting decision tree model is trained using a training set to obtain a second machine learning model, and based on this, a final pharmacokinetic property prediction model is constructed. Among them, gradient boosting decision trees (such as XGBoost or LightGBM) are an ensemble learning method that can handle structured data and capture complex non-linear relationships.

[0165] During the implementation process, first, the parameters of the gradient boosting decision tree model are set, such as the depth of the tree, the learning rate, the regularization term, etc. Then, the sample data in the training set is used for iterative training. In each step, a new decision tree is added to fit the residuals predicted in the previous round, gradually improving the prediction accuracy of the model.

[0166] After training is completed, the obtained gradient boosting decision tree model is the second machine learning model. Finally, the first machine learning model (deep neural network) and the second machine learning model (gradient boosting decision tree) are combined to construct a pharmacokinetic property prediction model. This combination can be achieved in various ways, such as simple averaging, weighted averaging, or constructing a meta-model to integrate the outputs of the two models.

[0167] The final prediction model can utilize the powerful feature extraction ability of the deep neural network and the advantages of the gradient boosting decision tree on structured data to improve the accuracy and robustness of the prediction of pharmacokinetic properties, providing a more reliable basis for drug development.

[0168] In one example, XGBoost can be used as the second machine learning model, and its training process is as follows: Based on the full amount of data in the training set, hyperparameters such as the maximum tree depth and the learning rate are optimized through grid search, and a weighted loss function (the sample weight is inversely proportional to the task data volume) is enabled for small sample tasks. Finally, the second machine learning model is generated.

[0169] In the model deployment stage, DNN and XGBoost are integrated through a dynamic weight allocation module (weighted by the reciprocal of the standard deviation) to form an end-to-end pharmacokinetic property prediction model.

[0170] S204, the output results of the pharmacokinetic property prediction model are verified using a validation set to evaluate the performance of the pharmacokinetic property prediction model.

[0171] First, the previously divided validation set is input into the trained pharmacokinetic property prediction model. The model will output the corresponding pharmacokinetic property prediction results according to the multi-dimensional feature data in the validation set.

[0172] Then, these predicted results are compared with the true pharmacokinetic property data (i.e., true labels) of the samples in the validation set, and a series of evaluation metrics are calculated to quantify the prediction performance of the model. Commonly used evaluation metrics include metrics for regression problems such as mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), etc., and metrics for classification problems such as accuracy, recall, F1-score, area under the ROC curve (AUC), etc. (The selection of specific metrics depends on whether the pharmacokinetic property prediction task is a regression problem or a classification problem). For example, when predicting the metabolic rate of a compound (a regression problem), the MSE and RMSE between the predicted value and the true value can be calculated. These two metrics respectively reflect the sum of squares and the average of the prediction errors, and the square root of the average. The smaller the value, the closer the predicted result of the model is to the true value, and the better the prediction performance of the model.

[0173] Meanwhile, a scatter plot of the actual values and the predicted values can also be drawn to visually observe the fitting degree and distribution trend between the model prediction results and the true data.

[0174] Through the validation process of the validation set, the performance of the pharmacokinetic property prediction model on unseen new data can be objectively evaluated, and problems such as overfitting or underfitting that may exist in the model can be discovered, thereby providing a basis for subsequent model optimization and adjustment to ensure that the model can reliably predict the pharmacokinetic properties of candidate compounds in practical applications.

[0175] This embodiment also provides a method for building and training a pharmacokinetic property prediction model, and the specific process is as follows: 1) Construction of the training dataset: Integrate the SMILES structure information of drug molecules and the experimentally determined pharmacokinetic property data (such as human plasma protein binding rate, metabolic half-life). The data sources include public databases (ChEMBL, DrugBank, etc.) and enterprise internal experimental data to ensure coverage of diverse chemical spaces and ADMET label types.

[0176] 2) Data cleaning and standardization: Validity verification filters illegal SMILES expressions through the chemical validity check module of RDKit; outlier handling eliminates statistical outliers of physicochemical descriptors (such as logP > 10) and bioactivity values based on the 3σ principle; data standardization unifies the SMILES molecular representation (such as removing chiral markers, normalizing atomic number arrangements) to ensure data format consistency.

[0177] 3) Molecular structure feature extraction: The topological structure is characterized by the ECFP4 fingerprint (extended radius = 3, bit length = 2048) to dynamically encode the local sub-structures of the molecule; the physicochemical descriptors are calculated by RDKit to calculate 10 core physicochemical parameters such as the octanol-water partition coefficient (logP), topological polar surface area (TPSA), and the number of hydrogen bond donors / acceptors.

[0178] 4) Atomic-level feature parsing: Extract the attributes of each atom in the molecule, including atomic number (element type), connectivity (number of adjacent atoms), formal charge (charge distribution state), hybridization type (orbital configurations such as sp³ / sp²), and aromaticity (participation status in conjugated rings), and construct an atomic-level feature matrix.

[0179] 5) Key pharmacophore fingerprint generation: Based on 166 predefined sub-structures of the MACCS fingerprint (such as benzene ring, carboxylic acid group), sub-structure matching is performed by RDKit to generate a binary fingerprint to characterize the distribution of key functional groups related to drug absorption / metabolism.

[0180] 6) Multi-dimensional feature preprocessing: Feature alignment aligns the dimensions of the ECFP, MACCS fingerprint, and atomic-level feature matrix (padding / truncation); missing value processing uses the K-nearest neighbor algorithm (K = 5) to fill in the missing values of the physicochemical descriptors; the meta-learning adaptation layer constructs an embedding layer through PyTorch to map heterogeneous features to a unified latent space (dimension 1024). Among them, PyTorch is an open-source machine learning framework based on the Python language, which provides powerful tensor computing capabilities and a dynamic computational graph mechanism.

[0181] 7) High-order feature fusion and extraction: The multi-head attention mechanism inputs the features into an 8-head attention module, calculates the query (Q), key (K), and value (V) vectors, generates attention weights through the similarity matrix and Softmax normalization, and focuses on cross-dimensional associations (such as the synergistic effect between logP and the number of hydrogen bond donors); after non-linear transformation through residual connection and layer normalization, a 3-layer fully connected network (1024→512→256→1) combines with the ReLU activation function to extract high-order molecular representations, and Dropout (probability = 0.3) is introduced to prevent overfitting.

[0182] 8) Small-sample adaptability verification: Data partitioning is stratified sampling according to the ADMET task, and 3-Fold cross-validation is used. Each fold contains 80% training set and 20% validation set; normalization processing performs Z-score normalization (mean-variance from the training set) on the training set and validation set respectively to eliminate the dimensional difference.

[0183] 9) Meta - learning framework training: Task simulation divides data into meta - tasks (such as "CYP3A4 inhibition prediction") based on the MAML framework. Each task contains a support set (5 samples) and a query set (15 samples); for parameter optimization, the inner loop (support set) quickly adapts to new tasks through 3 - step gradient descent, and the outer loop (query set) calculates the cross - task meta - loss, and uses the second - derivative to optimize the initial parameters to improve the generalization ability for few - shot learning.

[0184] 10) Heterogeneous model integration construction: The DNN model builds a deep network based on PyTorch (input 128 - dimensional → 3 - layer fully connected → output layer), and uses the Adam optimizer (lr = 1e - 4) and label smoothing for training; the XGBoost model optimizes hyperparameters through grid search (such as maximum depth = 6, learning rate = 0.1); in the dynamic weight allocation deployment stage, the confidence weights are calculated based on the standard deviation of the prediction results of the two models to achieve adaptive fusion.

[0185] 11) Prediction result generation and interpretation: Uncertainty quantification calculates the standard deviation of DNN predictions through Monte Carlo Dropout (iterated 100 times), and XGBoost estimates the confidence interval based on the out - of - bag error; weighted output normalizes and weights the prediction values of DNN (modeling non - linear relationships) and XGBoost (analyzing structured features) to output the final probability values of pharmacokinetic properties and confidence scores; visual feedback generates a feature attribution heatmap (such as the contribution degree of ECFP sites to metabolism prediction) to assist R & D personnel in decision - making.

[0186] As Figure 2 shown, this embodiment also provides an implementation process for few - shot pharmacokinetic property prediction based on meta - learning, which is specifically as follows: I. Input learning data: Integrate the SMILES structure information of drug molecules and the experimentally determined pharmacokinetic property data (such as plasma protein binding rate, metabolic half - life), ensuring that the data covers a diverse chemical space and ADMET label types.

[0187] II. Feature engineering: ECFP fingerprint: Use RDKit to generate ECFP4 fingerprints (expansion radius = 3, bit length = 2048) to dynamically encode the local sub - structures of molecules.

[0188] Molecular descriptors: Calculate 10 types of physicochemical parameters such as the octanol - water partition coefficient (logP), topological polar surface area (TPSA), number of hydrogen - bond donors / acceptors, etc.

[0189] Atom - level features: Extract atomic number, connectivity, formal charge, hybridization type, and aromaticity to construct an atom - level feature matrix.

[0190] MACCS Fingerprint: Generate a binary fingerprint by matching 166 predefined pharmacophore substructures (such as benzene ring, carboxylic acid group) through RDKit.

[0191] III. Feature Combination: Concatenate the ECFP fingerprint, MACCS fingerprint, molecular descriptors, and atomic-level features into a unified input matrix, and fill in the missing values through the K-nearest neighbor algorithm.

[0192] IV. Data Standardization: Perform Z-score standardization (the mean and standard deviation are from the training set) on the feature matrix to eliminate the dimension difference.

[0193] V. Cross-Validation: Stratified sampling is carried out according to the ADMET task, and 3-Fold cross-validation is adopted. Each fold contains 80% of the training set and 20% of the validation set.

[0194] VI. Training Set Processing: Meta-Learning Pre-Training: Based on the MAML framework, divide the training set into meta-tasks (such as "CYP3A4 inhibition prediction"). Each task contains a support set (5 samples) and a query set (15 samples). Improve the generalization ability of few-shot learning through the inner loop (support set gradient descent) and the outer loop (cross-task meta-loss optimization).

[0195] XGBoost Hyperparameter Tuning: Optimize the maximum tree depth (3-9) and learning rate (0.01-0.3) through grid search, and obtain the optimal parameters based on hyperparameter optimization.

[0196] DNN Fine-Tuning: Build a 4-layer fully connected network, and adopt the AdamW optimizer (learning rate 1e-3) and Dropout regularization (probability 0.3).

[0197] VII. Model Ensemble: Dynamically fuse the DNN and XGBoost models, calculate the confidence weight according to the standard deviation of the prediction results, and generate the final prediction value by weighting.

[0198] VIII. Validation Set Evaluation: Dynamic Weight Prediction: Input the validation set data into the ensemble model and output the weighted prediction results.

[0199] Model Evaluation: Calculate RMSE, R² for regression tasks and AUC-ROC, F1 Score for classification tasks.

[0200] IX. Output Prediction Results: Structured Output: Probability values of pharmacokinetic properties (such as human plasma protein binding rate ≥ 95%) and confidence scores.

[0201] This embodiment has the following technical effects: 1. End-to-End Automation: The entire process from the original SMILES input to the prediction results requires no manual intervention.

[0202] 2. Small-sample Adaptability: The meta-learning framework enables the model to maintain high accuracy under data-scarce conditions.

[0203] 3. Heterogeneous Feature Fusion: ECFP + MACCS + atomic-level features cover multi-scale information of molecules.

[0204] 4. Robustness Guarantee: Dynamic weight allocation suppresses the bias of a single model and improves prediction stability.

[0205] As Figure 3 shown, the specific process of a meta-learning mechanism in this embodiment is as follows: Starting from the "meta-learning sampling" step, the input data is split into two parallel paths: the "support set" path and the "query set" path.

[0206] In the "support set" path, the "model fast adaptation" step is executed, that is, based on the support set samples (such as 5 molecular data), local gradient updates are performed on the initial model parameters to enable the model to quickly adapt to the current task (such as specific metabolic property prediction).

[0207] In the "query set" path, the "meta-loss calculation" step is executed, and the query set samples (such as 15 new molecules) are used to evaluate the generalization performance of the model after adaptation, and the cross-task meta-loss (such as mean square error or cross entropy) is calculated.

[0208] The intermediate results of the two paths converge in the "parameter adjustment" step, and the initial model parameters are jointly optimized through the backpropagation algorithm, focusing on updating the parameter directions that have a significant impact on the multi-task generalization ability.

[0209] Finally, the "global model update" step is executed, and the global model is iteratively updated according to the gradient information after parameter adjustment, enabling it to have the ability to quickly learn new tasks from small-sample data.

[0210] The entire process of this embodiment's meta-learning mechanism realizes the efficient knowledge transfer and improvement of generalization performance of the model in the small-sample scenario of drug R & D through the synergy of the support set (local adaptation) and the query set (global optimization).

[0211] The method provided in the above embodiment can be executed by an electronic device. The following describes this electronic device in the embodiments of the present invention from the perspective of hardware processing. Please refer to Figure 4 , which is a schematic structural diagram of an entity device of the electronic device in the embodiments of the present invention.

[0212] It should be noted that Figure 4 the structure of the electronic device shown is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present invention.

[0213] As Figure 4As shown, the electronic device includes a Central Processing Unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a Read-Only Memory (ROM) 402 or a program loaded from a storage section 408 into a Random Access Memory (RAM) 403, such as executing the method described in the above embodiments. In the Random Access Memory (RAM) 403, various programs and data required for system operations are also stored. The Central Processing Unit (CPU) 401, the Read-Only Memory (ROM) 402, and the Random Access Memory (RAM) 403 are connected to each other via a bus 404. An Input / Output (I / O) interface 405 is also connected to the bus 404.

[0214] The following components are connected to the Input / Output (I / O) interface 405: an input section 406 including an audio input device, a button switch, etc.; an output section 407 including a display, an audio output device, an indicator light, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the Input / Output (I / O) interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read from it can be installed into the storage section 408 as needed.

[0215] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by the Central Processing Unit (CPU) 401, various functions defined in the present invention are executed.

[0216] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0217] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings.

[0218] Specifically, the electronic device in this embodiment includes a processor and a memory. The memory is coupled to one or more processors, and the memory is used to store computer program code. The computer program code includes computer instructions, and one or more processors call the computer instructions to cause the electronic device to execute the method provided in the above embodiment.

[0219] On the other hand, the present invention also provides a computer-readable storage medium. This storage medium may be included in the electronic device described in the above embodiment; or it may exist separately without being assembled into the electronic device. The above storage medium carries one or more computer programs, and when the one or more computer programs are executed by a processor of the electronic device, the electronic device is caused to implement the method provided in the above embodiment.

[0220] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.

[0221] As used in the foregoing embodiments, depending on the context, the term "when" may be interpreted as "if" or "after" or "in response to determining..." or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is detected" may be interpreted as "if determined..." or "in response to determining..." or "when (the stated condition or event) is detected" or "in response to detecting (the stated condition or event)".

[0222] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the foregoing embodiments can be implemented. The processes can be completed by relevant hardware instructed by a computer program, which can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the foregoing method embodiments. The foregoing storage media include: various media such as ROM or random access memory RAM, magnetic disks, or optical discs that can store program codes.

Claims

1. A small-sample pharmacokinetic property prediction method based on meta-learning, characterized in that Including: Obtain multi-dimensional characteristic data of a candidate compound, where the multi-dimensional characteristic data includes atomic-level characteristics, molecular descriptors, and molecular structure characteristics; Input the multi-dimensional characteristic data into a preset pharmacokinetic property prediction model to obtain a preliminary prediction result. The pharmacokinetic property prediction model is pre-trained by meta-learning based on a small sample. The pharmacokinetic property prediction model includes a first machine learning model and a second machine learning model. The preliminary prediction result includes a first prediction result output by the first machine learning model and a second prediction result output by the second machine model; Determine the pharmacokinetic property prediction result of the candidate compound according to the first prediction result and the second prediction result.

2. The method according to claim 1, wherein The determining the pharmacokinetic property prediction result of the candidate compound according to the first prediction result and the second prediction result includes: Obtain a first standard deviation of the first prediction result and a second standard deviation of the second prediction result; Determine a weight coefficient based on the first standard deviation and the second standard deviation; Determine the pharmacokinetic property prediction result according to the first prediction result, the second prediction result, and the weight coefficient.

3. The method according to claim 2, characterized in that, The weight coefficient includes a first weight coefficient and a second weight coefficient; the determining the weight coefficient based on the first standard deviation and the second standard deviation includes: Determine a first confidence parameter according to the first standard deviation and a second confidence parameter according to the second standard deviation; Normalize the first confidence parameter and the second confidence parameter respectively to obtain the first weight coefficient and the second weight coefficient.

4. The method according to any one of claims 1 to 3, characterized in that, The training process of the pharmacokinetic property prediction model includes: Obtain sample data including multiple multi-dimensional characteristic data and experimentally determined pharmacokinetic property data; Divide the sample data into a training set and a validation set; Based on the training set, train a machine learning model through a meta-learning pre-training mechanism to obtain a pharmacokinetic property prediction model; Verify the output result of the pharmacokinetic property prediction model through the validation set to evaluate the performance of the pharmacokinetic property prediction model.

5. The method according to claim 4, characterized in that, The training the machine learning model through a meta-learning pre-training mechanism based on the training set to obtain a pharmacokinetic property prediction model includes: Divide the training set into a support set and a query set; Pre-train a preset deep neural network through a meta-learning mechanism based on the support set to obtain a pre-trained neural network; Adjust the pre-trained neural network according to the query set to obtain a first machine learning model; In the case of training a preset gradient boosting decision model through the training set to obtain a second machine learning model, construct the pharmacokinetic property prediction model based on the first machine learning model and the second machine learning model.

6. The method according to claim 1, characterized in that, The inputting the multi-dimensional characteristic data into a preset pharmacokinetic property prediction model to obtain a preliminary prediction result includes: Through a multi-head attention mechanism, perform feature fusion processing on the multi-dimensional characteristic data to obtain high-order molecular representation data; Input the high-order molecular characterization data into the pharmacokinetic property prediction model to obtain the preliminary prediction result.

7. The method according to claim 6, characterized in that, The high-order molecular characterization data is obtained by performing feature fusion processing on the multi-dimensional feature data through a multi-head attention mechanism, including: Input the multi-dimensional feature data into a linear projection layer respectively to obtain a query vector, a key vector, and a value vector; Generate a similarity matrix according to the query vector and the key vector; Generate an attention weight matrix based on the similarity matrix; Perform weighted summation on the value vector according to the attention weight matrix to obtain multiple subspace feature vectors; Based on each of the subspace feature vectors, perform multi-layer non-linear transformation through a ReLU function to obtain the high-order molecular characterization data.

8. An electronic device, characterized in that, Comprising one or more processors and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1-7.

9. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product runs on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Method for predicting common physicochemical properties of organic molecules based on multiple learning models

    CN115985415A

  • Small sample molecule property prediction method based on hybrid relation network

    CN116580782A

  • Coronary heart disease prediction model training method, computer equipment and readable storage medium

    CN117637174A

  • Fitting prediction method, system and equipment of machine learning model and medium

    CN119046676A

  • Text mining data query method and system based on cross-modal similarity

    CN119311854A

Cited By

  • Prediction method for pharmacokinetic parameters of blood coagulation factor VIII and training method of corresponding model

    CN120895265A