Compound database labeling method, system and equipment based on machine learning

By building a machine learning-based compound database annotation system, the inaccuracy and inefficiency of molecular docking software selection when compound structures are diversified are solved, and efficient automatic labeling and recommendation of compound databases are achieved, which is suitable for high-throughput drug screening and AI-driven drug discovery.

CN120452615APending Publication Date: 2025-08-08CHONGQING MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510543805.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing technology lacks systematicity and repeatability in the selection of molecular docking software, resulting in inaccurate prediction results when compound structures are diversified, and traditional methods are inefficient when dealing with large-scale compound sets, making it difficult to meet the needs of high-throughput, automated and intelligent drug research and development.

Method used

By extracting compound characteristics and performance data from the compound database, using machine learning algorithms and multiple classification algorithms to build a molecular docking software selection model, generate recommended molecular docking software that is suitable for compounds, and deploy the model to the compound database for batch annotation, achieving automated and efficient software recommendations.

Benefits of technology

It realizes efficient and accurate automatic labeling of large-scale compounds, has high accuracy and efficiency in software recommendations, and is systematically scalable, suitable for high-throughput drug screening and AI-driven drug discovery platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452615A_ABST
    Figure CN120452615A_ABST
Patent Text Reader

Abstract

The invention provides a compound database labeling method, system and equipment based on machine learning, and the method comprises the steps: extracting the characteristics of a compound from a compound database, and obtaining the performance data of the compound in various molecular docking software; based on a machine learning algorithm and multiple classification algorithms, a molecular docking software selection model is constructed according to the compound characteristics and performance data, and the molecular docking software selection model is used for generating recommended molecular docking software matched with the compound; and deploying the molecular docking software selection model into a target compound database, carrying out batch labeling on compounds in the target compound database, and generating recommended molecular docking software and corresponding confidence coefficients. According to the method, automatic labeling of large-scale compounds can be achieved, the accuracy and efficiency of software recommendation are high, system expandability is achieved, and universality is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to a compound database annotation method, system, electronic device and computer-readable storage medium based on machine learning. Background Art

[0002] In drug discovery research, molecular docking technology has become an important tool for predicting interactions between small molecules and protein targets, and is widely used in key processes such as virtual screening and candidate drug optimization. The core of this technology lies in computationally predicting the binding conformation and affinity between small molecules and target proteins, thereby screening for potentially active candidate compounds. To this end, researchers typically utilize a range of molecular docking software, such as AutoDock Vina, Glide, GOLD, and MOE, for simulation and evaluation. These tools each have unique characteristics in terms of docking algorithms, scoring functions, and flexible processing capabilities, making them suitable for different docking tasks.

[0003] However, existing docking software selection techniques suffer from several key flaws, severely limiting the accuracy of docking results and the efficiency of virtual screening. First, the selection of molecular docking software currently relies primarily on researchers' empirical judgment, lacking systematicity and reproducibility. This human-driven selection mechanism can easily lead to inappropriate software selection when faced with highly diverse compound structures, thus compromising the reliability of prediction results.

[0004] Secondly, while current mainstream compound databases (such as PubChem, ChEMBL, and ZINC) provide a wealth of structural and bioactivity data, they lack dynamic annotation capabilities for docking software compatibility. Users cannot obtain docking software recommendations for specific compounds or compound classes based on the information in the database, which greatly limits the value of intelligent applications of the data.

[0005] Furthermore, different docking software programs exhibit significant differences in algorithm design and scoring systems, leading to significant performance variations when handling different types of ligands. For example, GOLD excels at docking flexible molecules, while AutoDockVina is more efficient when handling large molecular libraries. Therefore, when faced with a structurally diverse set of candidate molecules, a single software program often fails to achieve both accuracy and efficiency. There is an urgent need for intelligent strategies that automatically match the optimal tool based on molecular characteristics.

[0006] Finally, with the increasing application of virtual screening in drug development, screening tasks involving hundreds of thousands or even millions of compounds are becoming increasingly frequent. Traditional methods are inefficient when dealing with such large sets of compounds, and manual testing of multiple docking software solutions is extremely time-consuming, making it difficult to meet the high-throughput, automated, and intelligent requirements of modern drug development.

[0007] Therefore, there is an urgent need to develop a compound database annotation method that is oriented towards diverse compound structures and can achieve efficient and accurate recommendations for molecular docking software for compounds. Summary of the Invention

[0008] Based on this, it is necessary to provide a compound database annotation method, system, electronic device and computer-readable storage medium based on machine learning to address the above technical problems.

[0009] A machine learning-based compound database annotation method comprises the following steps: extracting compound features from the compound database and obtaining performance data of the compounds in a variety of molecular docking software; constructing a molecular docking software selection model based on the compound features and performance data based on a machine learning algorithm and a variety of classification algorithms, wherein the molecular docking software selection model is used to generate recommended molecular docking software that is compatible with the compounds; and deploying the molecular docking software selection model to a target compound database, batch-annotating the compounds in the target compound database, and generating recommended molecular docking software and corresponding confidence levels.

[0010] In one embodiment, the extracting of compound characteristics and obtaining performance data of the compound in docking software, and obtaining performance data of the compound in multiple molecular docking software, include: linking multiple compound data sets, extracting the physicochemical characteristics of the compound characterized by physicochemical descriptors and topological structure characteristics characterized by molecular fingerprints, the physicochemical descriptors including molecular weight, number of rotatable bonds, topological polar surface area, number of hydrogen bond acceptors / donors, lipid-water partition coefficient, ratio of hybridized carbon atoms, number of hybridized nitrogen atoms, Hall-Kier alpha value, number of aromatic rings, number of aliphatic rings and number of saturated rings; obtaining target information of all compounds, docking the compounds using multiple molecular docking software and obtaining corresponding performance data.

[0011] In one embodiment, the molecular docking software selection model is constructed based on the compound characteristics and performance data based on the machine learning algorithm and multiple classification algorithms, including: feature engineering processing of the physicochemical descriptors and molecular fingerprint characterizations, and labeling the molecular docking software; based on the machine learning algorithm and multiple classification algorithms, a multi-output classifier is used as the overall algorithm framework, a key parameter network is set with multiple classification algorithms, the processed compound characteristics and the labeled molecular docking software are input, and a parameter grid search under cross-validation is performed to obtain the best multi-output classification model as the molecular docking software selection model.

[0012] In one embodiment, the feature engineering processing of the physicochemical descriptors and molecular fingerprint representations includes: normalizing the physicochemical descriptors to the minimum or maximum extreme values so that the value range of the physicochemical descriptors is between 0 and 1, and matching them with the corresponding molecular fingerprint representations.

[0013] In one embodiment, the multiple classification algorithms include a random forest algorithm, a support vector machine, or a K-nearest neighbor algorithm.

[0014] In one embodiment, the molecular docking software selection model is deployed into a target compound database, compounds in the target compound database are batch-labeled, and recommended molecular docking software and corresponding confidence levels are generated, including: encapsulating the molecular docking software selection model and embedding it into the target compound database; batch-selecting compounds in the target compound database that have not been software-labeled and performing standardization processing, wherein the standardization processing includes desalting, charge neutralization, and standardizing molecular structure; using a distributed computing framework and a Celery task queue to allocate computing tasks, and performing batch labeling processing on the compounds in the target compound database; predicting the compatibility of the molecular docking software of the compound according to the molecular docking software selection model, obtaining a prediction probability as a confidence level, using the molecular docking software with the highest confidence level as the recommended molecular docking software, and outputting the recommended molecular docking software and the corresponding confidence level.

[0015] In one embodiment, the further embodiment further includes: dynamically updating the annotation content of the compound database when the molecular docking software selection model version is updated or new compounds are added or modified in the compound database.

[0016] A machine learning-based compound database annotation system is used to implement the machine learning-based compound database annotation method described above, comprising: a compound feature extraction module, used to extract compound features from the compound database and obtain performance data of the compounds in multiple molecular docking software; a software selection model construction module, used to construct a molecular docking software selection model based on the compound features and performance data based on a machine learning algorithm and multiple classification algorithms, the molecular docking software selection model being used to generate recommended molecular docking software compatible with the compounds; and a compound batch annotation module, used to deploy the molecular docking software selection model into a target compound database, batch-annotate the compounds in the target compound database, and generate recommended molecular docking software and corresponding confidence levels.

[0017] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the machine learning-based compound database annotation method described in each of the above embodiments are implemented.

[0018] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the machine learning-based compound database annotation method described in each of the above embodiments.

[0019] Compared with the existing technology, the advantages and beneficial effects of the present invention are: by extracting compound features from the compound database and obtaining the performance data of the compounds in multiple molecular docking software to facilitate subsequent model construction, a molecular docking software selection model is constructed based on the compound features and performance data based on machine learning algorithms and multiple classification algorithms, which is used to generate recommended molecular docking software that is compatible with the compounds, thereby achieving efficient and batch automated labeling and software recommendation. A high-accuracy molecular docking software selection model is obtained by combining machine learning algorithms and multiple classification algorithms, and the molecular docking software selection model is deployed to the target compound database, and the compounds therein are batch labeled to generate recommended molecular docking software and corresponding confidence levels, thereby achieving automatic labeling of large-scale compounds. The software recommendation has high accuracy and efficiency, and has system scalability and strong versatility. It is suitable for high-throughput drug screening, AI-driven drug discovery platforms, and drug development workflows of scientific research institutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A schematic diagram of a process for annotating a compound database based on machine learning in one embodiment;

[0021] Figure 2 A schematic diagram of the process of constructing, evaluating, and outputting the final results of a molecular docking software selection model in one embodiment;

[0022] Figure 3 Schematic diagram of the performance of three different basic classification algorithms in one embodiment;

[0023] Figure 4 A schematic diagram of the structure of a compound database annotation system based on machine learning in one embodiment;

[0024] Figure 5 FIG. 1 is a schematic diagram of the internal structure of an electronic device in one embodiment. DETAILED DESCRIPTION

[0025] Before describing the specific embodiments of the present invention, the overall concept of the present invention is described as follows:

[0026] This invention is primarily based on the molecular docking process. Currently, molecular docking software recommendations for compound databases rely primarily on manual judgment or simple rule-based matching methods. These methods are inefficient, lack systematicity, and are difficult to adapt to the structural characteristics of diverse compounds. These traditional methods fail to fully consider the complex relationship between multidimensional molecular characteristics and docking software performance, resulting in insufficient accuracy and versatility in their recommendations.

[0027] In response to the above technical bottlenecks, the present invention proposes a compound database annotation method based on machine learning. By extracting compound features from the compound database and obtaining performance data of the compounds in multiple molecular docking software to facilitate subsequent model construction, a molecular docking software selection model is constructed based on the compound features and performance data based on machine learning algorithms and multiple classification algorithms. The model is used to generate recommended molecular docking software that is compatible with the compounds, thereby achieving efficient and batch automated annotation and software recommendation. The machine learning algorithm and multiple classification algorithms are combined to obtain a high-accuracy molecular docking software selection model, which is deployed to the target compound database, and the compounds therein are batch annotated to generate recommended molecular docking software and corresponding confidence levels, thereby achieving automatic annotation of large-scale compounds. The software recommendation has high accuracy and efficiency, and has system scalability and strong versatility. It is suitable for high-throughput drug screening, AI-driven drug discovery platforms, and drug development workflows of scientific research institutions.

[0028] After introducing the overall concept of the present invention, in order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0029] In one embodiment, Figure 1 As shown, a compound database annotation method based on machine learning is provided, comprising the following steps:

[0030] Step S110 , extracting compound features from a compound database and obtaining performance data of the compound in a variety of molecular docking software.

[0031] Specifically, in order to predict the most suitable molecular docking software for a compound, it is necessary to extract multi-dimensional compound features to better characterize the compound. At the same time, it is necessary to obtain docking performance data of different molecular docking software for the compound, and train subsequent models based on the compound features and performance data.

[0032] Among them, step S110 includes: linking multiple compound databases, extracting the physicochemical characteristics of the compounds represented by physicochemical descriptors and topological structural characteristics represented by molecular fingerprints, the physicochemical descriptors include molecular weight, number of rotatable bonds, topological polar surface area, number of hydrogen bond acceptors / donors, lipid-water partition coefficient, hybridized carbon atom ratio, number of hybridized nitrogen atoms, Hall-Kier alpha value, number of aromatic rings, number of aliphatic rings and number of saturated rings; obtaining target information of all compounds, docking the compounds using multiple molecular docking software, and obtaining corresponding performance data.

[0033] Specifically, a data communication protocol is established between the database access interface and the model calculation engine; the model service API is configured to receive database query requests, multiple compound databases are linked, and the crystal structure data of target-compound molecules existing in the PDBbind database are queried and cleaned according to the molecular weight standard of the compound molecules to determine the total data set of more than 4,000 molecules.

[0034] Using open-source cheminformatics packages such as RDki, we obtained physicochemical descriptors and 1024-bit ECFP6 fingerprints of the compound molecules, with a minimum fingerprint length of 1024 bits. We then calculated information about these descriptors and docking success rates. We then docked these molecules using various molecular docking software packages using their default parameters, and calculated docking success rates to generate performance data for the compounds relative to the various docking software packages.

[0035] Among them, the physicochemical descriptors include at least five of the following: molecular weight, number of rotatable bonds, topological polar surface area, number of hydrogen bond acceptors, number of hydrogen bond donors, and lipid-water partition coefficient; the physicochemical descriptors also include: at least one of the proportion of hybridized carbon atoms, number of hybridized nitrogen atoms, and Hall-Kier alpha value; and ring structure characteristics: at least two of the following: number of aromatic rings, number of aliphatic rings, and number of saturated rings.

[0036] Step S120 , based on machine learning algorithms and multiple classification algorithms, a molecular docking software selection model is constructed according to compound characteristics and performance data. The molecular docking software selection model is used to generate recommended molecular docking software that is compatible with the compound.

[0037] Specifically, if Figure 2 As shown, the physicochemical descriptors and molecular fingerprint representations of the compounds in the data set are encoded and processed, and a molecular docking software selection model is constructed based on the compound characteristics and performance data based on machine learning algorithms and multiple classification algorithms, thereby realizing the selection of molecular docking software that is compatible with the compound, so as to obtain the recommended molecular docking software with the highest degree of compatibility with the compound.

[0038] Among them, step S120 includes: performing feature engineering processing on the physicochemical descriptors and molecular fingerprint representations, and labeling the molecular docking software; based on the machine learning algorithm and multiple classification algorithms, a multi-output classifier is used as the overall algorithm framework, and a key parameter network is set with multiple classification algorithms. The processed compound features and the labeled molecular docking software are input, and a parameter grid search is performed under cross-validation to obtain the best multi-output classification model as the molecular docking software selection model.

[0039] Among them, various classification algorithms include random forest algorithm, support vector machine or K nearest neighbor algorithm.

[0040] Specifically, the physicochemical descriptors and molecular fingerprint representations are processed through feature engineering and the docking software is labeled to facilitate subsequent model construction. Through refined feature engineering and algorithm optimization, high-precision modeling of compound structure-software compatibility is achieved; a multi-output classifier is used as the overall algorithm framework, and a variety of classification algorithms, such as random forest, K nearest neighbor and support vector machine, are used as basic classification algorithms. Key parameter networks are set for different algorithms and parameter grid search is performed under 5-fold cross-validation. The best model obtained by each algorithm is determined by comparing the area under the curve of its working characteristic curve and the accuracy-recall rate joint curve to determine the final algorithm model, such as Figure 3 As shown, the best multi-output classification model was obtained, which was used as the selection model for molecular docking software.

[0041] The feature engineering processing of the physicochemical descriptors and molecular fingerprint representations includes: normalizing the physicochemical descriptors to their minimum or maximum extreme values so that the value range of the physicochemical descriptors is between 0 and 1, and matching them with the corresponding molecular fingerprint representations.

[0042] Specifically, the physicochemical descriptors of the compounds are normalized to their minimum and maximum extreme values so that the value range of each physicochemical descriptor before input into the model training is within 0 to 1, matching its molecular fingerprint descriptor, thereby realizing the processing of compound characteristics to facilitate model construction.

[0043] Step S130 , deploying the molecular docking software selection model to the target compound database, batch-labeling the compounds in the target compound database, and generating recommended molecular docking software and corresponding confidence levels.

[0044] Specifically, the machine learning model is packaged and embedded into the target compound database, and the final packaging format of the model is determined based on the format of the different databases. Molecules that have not been software-annotated are batch-selected from the target database and standardized before being batch-annotated. The annotation content includes the recommendation of the most suitable molecular docking software for the compound and its confidence level, thus achieving automated annotation of the compound database and rapid selection of molecular docking software for compounds.

[0045] Among them, step S130 includes: encapsulating the molecular docking software selection model and embedding it into the target compound database; batch-selecting compounds that have not been software-annotated in the target compound database for standardization, and the standardization processing includes desalting, neutralizing charges, and standardizing molecular structures; through a distributed computing framework, using the Celery task queue to allocate computing tasks, and batch-annotating the compounds in the target compound database; predicting the adaptability of the molecular docking software of the compound according to the molecular docking software selection model, obtaining the prediction probability as the confidence, and using the molecular docking software with the highest confidence as the recommended molecular docking software, and outputting the recommended molecular docking software and the corresponding confidence.

[0046] Specifically, the resulting molecular docking software selection model is serialized into formats such as .pkl and .joblib and encapsulated as a RESTful API using the Flask framework, providing an interface that accepts molecular string input in the form of SMILES. A new table, software_recommendations, is created within the target compound database, containing fields such as compound_id (primary key), smiles, recommended_software, confidence_score, and update_time to display the model's prediction of the most suitable molecular docking software for the molecule. Regular expressions are also used to verify the accuracy of the input SMILES to ensure stable model operation.

[0047] Compounds not yet annotated by software are batch-selected from the database. Chemical informatics packages such as RDkit are used to standardize the molecules by removing salt, neutralizing charges, and generating standardized SMILES. In distributed computing, the Celery task queue is used to allocate computational tasks, performing batch annotation operations on compound molecules in the database. The number of molecules processed in each batch is determined based on the size of the database. Multiprocessing.Pool implements multi-process parallelism. During annotation, the molecular docking software selection model is used to predict the compatibility of the compound's molecular docking software. The predicted probability is used as the confidence level, and the confidence levels of all molecular docking software are compared. The molecular docking software with the highest confidence level is recommended as the molecular docking software for the compound. The recommended molecular docking software and the corresponding confidence level are output to complete compound annotation, thus achieving batch updates of molecules not yet annotated by software in the database. A retry mechanism is added for annotation failures, and annotation failures are output to the run log.

[0048] In one embodiment, the method further includes dynamically updating the annotation content of the compound database when the molecular docking software selection model version is updated or new compounds are added or modified in the compound database.

[0049] Specifically, the compounds in the compound database may have their compound information updated with daily use and collection, so it is necessary to establish a dynamic update mechanism for the software annotation results of the compounds in the compound database.

[0050] The trigger conditions for dynamic updates are set to be dynamically updated when the model version is updated and when the compound database content is added or modified. By monitoring the model version or data changes, the automatic annotation dynamic update is triggered to ensure the validity of the results. The control system can trace the change history of the annotation results and support retrospective analysis of the evolution of the annotation of specific compounds. For newly added compounds, the new compounds are annotated with recommended software and the confidence is calculated according to step S130. At the same time, the dynamic update mode includes two update strategies: incremental update and full re-run. The incremental update is used to meet the increase in software annotation results after the information of the compounds in the compound database is changed, and the full re-run is used to meet the update of large-scale compound annotation results in the database after the model is improved when incremental learning is set later, so as to achieve the purpose of maintaining high credibility and high practicality of the annotation results.

[0051] In addition, to improve the practicality and adaptability of the model, a user-involved feedback verification mechanism and an incremental learning process based on this can also be designed, including the design of data quality monitoring after incremental learning, such as the design of the feedback interface; feedback information reliability review; incremental learning algorithm design and data quality monitoring, etc.

[0052] Specifically, when researchers submit correction suggestions through the review interface, the system initiates an incremental learning process: First, the reliability of the manual annotations is verified, meaning users can upload experimental verification results for review. After verification and calibration, the confirmed valid data is added to the training set, and the model is optimized using a small-batch update strategy, retaining 20% of the original training data as an indicator for the new training process. The model version history is also recorded. By establishing an annotation quality monitoring system, when a significant decrease in annotation accuracy for a specific class of compounds is detected, targeted model retraining can be automatically initiated or information can be provided to database administrators to select a model version callback.

[0053] In this embodiment, compound features are extracted from a compound database, and performance data of the compounds in a variety of molecular docking software are obtained to facilitate subsequent model construction. A molecular docking software selection model is constructed based on the compound features and performance data based on machine learning algorithms and a variety of classification algorithms. The model is used to generate recommended molecular docking software that is compatible with the compounds, thereby achieving efficient and batch automated labeling and software recommendation. A highly accurate molecular docking software selection model is obtained by combining machine learning algorithms and a variety of classification algorithms. The molecular docking software selection model is deployed to a target compound database, and the compounds therein are batch labeled to generate recommended molecular docking software and corresponding confidence levels, thereby achieving automatic labeling of large-scale compounds. The software recommendations are accurate and efficient, and have system scalability and strong versatility. They are suitable for high-throughput drug screening, AI-driven drug discovery platforms, and drug development workflows of scientific research institutions.

[0054] like Figure 4 As shown, a compound database annotation system 40 based on machine learning is provided, which is used to implement the above-mentioned compound database annotation method based on machine learning, including: a compound feature extraction module 41, a software selection model construction module 42 and a compound batch annotation module 43, wherein:

[0055] The compound feature extraction module 41 is used to extract compound features from the compound database and obtain performance data of the compound in various molecular docking software;

[0056] A software selection model construction module 42 is used to construct a molecular docking software selection model based on the characteristics and performance data of the compound based on a machine learning algorithm and multiple classification algorithms. The molecular docking software selection model is used to generate a recommended molecular docking software that is compatible with the compound;

[0057] The compound batch annotation module 43 is used to deploy the molecular docking software selection model to the target compound database, batch-annotate the compounds in the target compound database, and generate recommended molecular docking software and corresponding confidence levels.

[0058] In one embodiment, the compound feature extraction module 41 is specifically used to: link multiple compound databases, extract the physicochemical characteristics of the compounds represented by physicochemical descriptors and the topological structural characteristics represented by molecular fingerprints, the physicochemical descriptors include molecular weight, number of rotatable bonds, topological polar surface area, number of hydrogen bond acceptors / donors, lipid-water partition coefficient, hybridized carbon atom ratio, number of hybridized nitrogen atoms, Hall-Kier alpha value, number of aromatic rings, number of aliphatic rings and number of saturated rings; obtain target information of all compounds, use multiple molecular docking software to dock the compounds, and obtain corresponding performance data.

[0059] In one embodiment, the software selection model construction module 42 is specifically used to: perform feature engineering processing on physicochemical descriptors and molecular fingerprint representations, and label the molecular docking software; based on machine learning algorithms and multiple classification algorithms, a multi-output classifier is used as the overall algorithm framework, and a key parameter network is set with multiple classification algorithms. The processed compound features and labeled molecular docking software are input, and a parameter grid search is performed under cross-validation to obtain the best multi-output classification model as the molecular docking software selection model.

[0060] In one embodiment, the compound batch annotation module 43 is specifically used to: encapsulate the molecular docking software selection model and embed it into the target compound database; batch select compounds that have not been software-annotated in the target compound database for standardization, and the standardization includes desalting, neutralizing charges, and standardizing molecular structures; through a distributed computing framework, use the Celery task queue to allocate computing tasks and perform batch annotation processing on the compounds in the target compound database; predict the adaptability of the molecular docking software of the compound according to the molecular docking software selection model, obtain the prediction probability as the confidence, and use the molecular docking software with the highest confidence as the recommended molecular docking software, and output the recommended molecular docking software and the corresponding confidence.

[0061] In the above embodiment, the system can realize efficient and batch automated annotation and molecular docking software recommendation, taking into account prediction accuracy, computational efficiency and system scalability, significantly improving the management efficiency of the compound database, and is suitable for high-throughput drug screening, AI-driven drug discovery platforms and drug development workflows of scientific research institutions, with significant practical value and industrialization prospects.

[0062] In one embodiment, an electronic device is provided. The device may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the device is used to provide computing and control capabilities. The memory of the device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the device is used to store configuration templates and can also be used to store target web page data. The network interface of the device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a compound database annotation method based on machine learning is implemented.

[0063] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the device to which the solution of the present application is applied. The specific device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0064] In one embodiment, a computer-readable storage medium may also be provided, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method as described in the aforementioned embodiment. The computer may be part of the above-mentioned machine learning-based compound database annotation system.

[0065] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0066] Obviously, those skilled in the art should understand that the modules or steps of the present invention described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices. Alternatively, they can be implemented using program codes executable by the computing device, so that they can be stored in a computer storage medium (ROM / RAM, magnetic disk, optical disk) and executed by the computing device. In some cases, the steps shown or described can be performed in a different order than herein, or they can be made into individual integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module for implementation. Therefore, the present invention is not limited to any specific combination of hardware and software.

[0067] The above content is a further detailed description of the present invention in conjunction with specific embodiments, and the specific implementation of the present invention cannot be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A compound database annotation method based on machine learning, characterized in that: The following steps are involved: Extract compound features from compound databases and obtain compound performance data in various molecular docking software; Based on machine learning algorithms and multiple classification algorithms, a molecular docking software selection model is constructed according to the characteristics and performance data of the compound, and the molecular docking software selection model is used to generate recommended molecular docking software that is compatible with the compound; The molecular docking software selection model is deployed into a target compound database, compounds in the target compound database are batch-annotated, and recommended molecular docking software and corresponding confidence levels are generated.

2. The method for annotating a compound database based on machine learning according to claim 1, characterized in that: The extraction of compound features and acquisition of compound performance data in docking software and acquisition of compound performance data in various molecular docking software include: Link multiple compound data sets to extract the compounds' physicochemical characteristics represented by physicochemical descriptors and topological structural features represented by molecular fingerprints. The physicochemical descriptors include molecular weight, number of rotatable bonds, topological polar surface area, number of hydrogen bond acceptors / donors, lipid-water partition coefficient, ratio of hybridized carbon atoms, number of hybridized nitrogen atoms, Hall-Kier alpha value, number of aromatic rings, number of aliphatic rings, and number of saturated rings. Obtain target information of all compounds, use a variety of molecular docking software to dock the compounds and obtain corresponding performance data.

3. The method for annotating a compound database based on machine learning according to claim 2, wherein: The molecular docking software selection model is constructed based on the machine learning algorithm and multiple classification algorithms according to the compound characteristics and performance data, including: Performing feature engineering on the physicochemical descriptors and molecular fingerprint representations, and labeling the molecular docking software; Based on machine learning algorithms and multiple classification algorithms, a multi-output classifier is used as the overall algorithm framework. The key parameter network is set with multiple classification algorithms. The processed compound features and labeled molecular docking software are input, and a parameter grid search is performed under cross-validation to obtain the best multi-output classification model as the molecular docking software selection model.

4. The method for annotating a compound database based on machine learning according to claim 3, wherein: The feature engineering processing of the physicochemical descriptors and molecular fingerprint representations includes: The physical descriptor is normalized to its minimum or maximum extreme value so that the value range of the physical descriptor is between 0 and 1, and is matched with the corresponding molecular fingerprint representation.

5. The method for annotating a compound database based on machine learning according to claim 3, wherein: The multiple classification algorithms include a random forest algorithm, a support vector machine, or a K-nearest neighbor algorithm.

6. The method for annotating a compound database based on machine learning according to claim 1, characterized in that: The molecular docking software selection model is deployed into the target compound database, compounds in the target compound database are batch-annotated, and recommended molecular docking software and corresponding confidence levels are generated, including: Encapsulating the molecular docking software selection model and embedding it into the target compound database; Compounds that have not been annotated by the software are batch-selected in the target compound database for standardization, wherein the standardization includes desalting, charge neutralization, and molecular structure standardization; Through the distributed computing framework, the Celery task queue is used to allocate computing tasks and batch-annotate the compounds in the target compound database; The molecular docking software of the compound is adapted according to the molecular docking software selection model to obtain a prediction probability as a confidence level, the molecular docking software with the highest confidence level is used as a recommended molecular docking software, and the recommended molecular docking software and the corresponding confidence level are output.

7. The method for annotating a compound database based on machine learning according to claim 1, characterized in that: Also includes: When the molecular docking software selection model version is updated or new compounds appear or are modified in the compound database, the annotation content of the compound database is dynamically updated.

8. A compound database annotation system based on machine learning, characterized in that: A method for annotating a compound database based on machine learning according to any one of claims 1 to 7, comprising: Compound feature extraction module, used to extract compound features from the compound database and obtain compound performance data in various molecular docking software; a software selection model construction module for constructing a molecular docking software selection model based on the characteristics and performance data of the compound based on a machine learning algorithm and a plurality of classification algorithms, wherein the molecular docking software selection model is used to generate a recommended molecular docking software that is compatible with the compound; The compound batch annotation module is used to deploy the molecular docking software selection model into the target compound database, batch annotate the compounds in the target compound database, and generate recommended molecular docking software and corresponding confidence levels.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.