Drug molecule screening method and system fusing molecular docking and machine learning

By integrating molecular docking and machine learning methods, the problems of low efficiency and insufficient accuracy of traditional drug molecule screening technology have been solved, efficient, automated and intelligent drug molecule screening has been achieved, and more accurate predictions of drug molecule activity have been provided.

CN120673910APending Publication Date: 2025-09-19SUZHOU GUANGZHIJI DATA TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510845853.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional drug molecular screening technology is inefficient and inaccurate, and has difficulty processing complex data. It often uses a single screening method, such as relying solely on molecular docking or machine learning, lacks deep predictive capabilities, and has a low degree of automation.

Method used

A drug molecule screening method that integrates molecular docking and machine learning. Through data preparation, molecular docking simulation, feature extraction, machine learning model training and evaluation, it can simulate the binding process and predict features between drug molecules and target proteins, and optimize screening with machine learning models.

Benefits of technology

It improves the accuracy and efficiency of drug molecule screening, realizes full process automation and intelligence, reduces manual operations, shortens the R&D cycle, and provides rich structural and data features for screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673910A_ABST
    Figure CN120673910A_ABST
Patent Text Reader

Abstract

The invention discloses a drug molecule screening method and system fusing molecular docking and machine learning, and relates to the field of drug molecule screening, and the method comprises the following operation steps: S1, data preparation; s2, carrying out molecular docking simulation; s3, feature extraction; s4, machine learning model training; s5, evaluating and optimizing the model; and S6, drug molecule screening. According to the drug molecule screening method and system fusing molecular docking and machine learning, molecular docking and machine learning are deeply fused, the combination process of drug molecules and target protein can be simulated based on the three-dimensional structure of the drug molecules and the target protein, the intermolecular interaction mode is visually displayed, the combination capacity is quantified through a scoring function, and the screening efficiency is improved. A visual structure level basis is provided for screening, the drug molecular activity can be evaluated more accurately based on the basis, and the probability of screening high-potential candidate drugs is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of drug molecule screening, and in particular to a drug molecule screening method and system integrating molecular docking and machine learning. Background Art

[0002] In the field of drug research and development, traditional drug molecule screening technology faces many bottlenecks, such as low efficiency, insufficient accuracy, and difficulty in processing complex data. At the same time, single screening methods are often used, such as relying solely on molecular docking technology to evaluate the binding ability of drug molecules and targets only from a structural level, lacking in-depth prediction of drug molecule activity; or simply using machine learning technology, but due to incomplete data characteristics, the model training effect is not good, and the traditional drug molecule screening system has a low degree of automation, and many links require manual operation.

[0003] Therefore, it is necessary to propose a drug molecule screening method and system that integrates molecular docking and machine learning to solve the above problems. Summary of the Invention

[0004] The main purpose of the present invention is to provide a drug molecule screening method and system that integrates molecular docking and machine learning, which can effectively solve the problems in the background technology.

[0005] To achieve the above object, the technical solution adopted by the present invention is: A drug molecule screening method integrating molecular docking and machine learning, comprising the following steps: S1: Data preparation: collecting molecular structural data of drugs with known activity and three-dimensional structural data of corresponding target proteins, cleaning and standardizing them to ensure data quality and avoid affecting the accuracy of screening results due to data problems; S2: Molecular docking simulation: docking the prepared drug molecules with the target protein, and calculating the binding mode and binding energy between the drug molecules and the target protein. By simulating the binding process of the drug molecules and the target protein, the binding ability of the drug molecules and the target is preliminarily evaluated, and drug molecules with potential binding possibility are screened; S3: Feature extraction: extract key features from the molecular docking results, quantify and encode these features to form feature vectors for subsequent model training and prediction; S4: Machine learning model training: Select an appropriate machine learning algorithm, use the extracted feature vectors and corresponding drug molecule activity data as a training set, and input them into the model for training. This will enable the model to predict drug activity based on its characteristics, thereby achieving efficient screening of drug molecules. S5: Model evaluation and optimization: Use a portion of untrained data as a test set to evaluate the trained machine learning model to ensure the model has high reliability and predictive ability, providing an accurate basis for subsequent drug molecule screening; S6: Drug molecule screening: The drug molecule data to be screened are subjected to molecular docking and feature extraction according to the above steps, and then input into the trained and optimized machine learning model. The model predicts the activity of the drug molecules. Based on the prediction results, the drug molecules are ranked according to their activity to screen out drug molecules with high activity potential.

[0006] Preferably, the step S1 specifically includes the following steps: S101: Collect drug molecular structure data: Download small molecule compound data with clear activity annotations from multiple public databases, including simplified molecular linear input specifications and international chemical identifiers of drug molecules, and collect drug molecular structure and activity data already available in the internal databases of drug R&D companies; S102: Collect target protein three-dimensional structural data: Retrieve protein structures related to the research target from the protein database. For target proteins whose structures have not yet been experimentally determined, use homology modeling methods to construct their three-dimensional structural models using proteins with high sequence similarity and known structures as templates; S103: Clean the collected data.

[0007] Preferably, the step S2 specifically includes the following steps: S201: Define the docking region. Define the search space for docking based on the functional sites and known active pockets of the target protein. For enzyme targets, the area within a certain range around the active center is set as the docking region. For receptor targets, the docking region is determined based on the position and size of the ligand binding site. S202: Select a docking algorithm, including rigid docking, semi-flexible docking, and flexible docking algorithms; S203: Setting the scoring function: Selecting an appropriate scoring function to evaluate the binding ability of the drug molecule to the target protein, including the Vina scoring function of AutoDockVina and the XP scoring function of Glide; S204: docking the prepared drug molecules with the target protein, and performing molecular docking on the standardized drug molecule structure file and the target protein structure file. At this point, docking calculations can be performed on each drug molecule and target protein one by one; S205: After the docking calculation is completed, the results are analyzed to screen out drug molecules with potential binding ability. The binding mode between the drug molecule and the target protein is checked, including the position and orientation of the drug molecule in the active pocket, and the type of interaction with the key amino acid residues. Based on this, the rationality of the binding between the drug molecule and the target is preliminarily judged. According to the binding energy value calculated by the scoring function, a threshold is set to screen out drug molecules with lower binding energy as candidate molecules, while excluding molecules with unreasonable binding modes.

[0008] Preferably, the step S3 specifically includes the following steps: S301: Molecular structure feature extraction, including the extraction of functional group features and topological structure features of drug molecules; S302: Extracting docking result features: extracting binding-related features from the molecular docking results, including binding energy features and binding site features; S303: Physical and chemical property feature extraction, calculating various physical and chemical property features of drug molecules, including basic properties, lipophilicity characteristics and solubility characteristics; S304: Integrate the various features extracted above, and perform unified coding to form a feature vector.

[0009] Preferably, the step S4 specifically includes the following steps: S401: Algorithm selection: Based on the task requirements and data characteristics of drug molecule screening, select appropriate machine learning algorithms to build models, including classification algorithms and regression algorithms; S402: Build a model using the required machine learning algorithm, divide the extracted feature vectors and the corresponding drug molecule activity data into a training set and a test set in a ratio of 7:3, and set the initial parameters of the model based on the characteristics of the selected algorithm; S403: Input the training set data into the constructed model for training.

[0010] Preferably, the step S5 specifically includes the following steps: S501: Model evaluation: Use the test set data to comprehensively evaluate the trained machine learning model, including classification model evaluation indicators: evaluate the model based on the calculation of accuracy, precision, recall rate, and F1 value indicators; It also includes regression model evaluation indicators: including calculation of mean square error, root mean square error, mean absolute error, and coefficient of determination to evaluate the model; S502: Based on the evaluation results, the model is optimized and improved, including parameter adjustment, feature engineering optimization, and model fusion.

[0011] Preferably, the step S6 specifically includes the following steps: S601: Preprocessing the drug molecule data to be screened, including converting the structural data of the drug molecules to be screened into a unified format and performing structural normalization to ensure the chemical rationality and consistency of the molecular structure, obtaining the corresponding target protein three-dimensional structure, and performing hydrogenation, dehydration, and removal of irrelevant ligands to ensure that the target protein structure meets the requirements of molecular docking; S602: Molecular docking: Using the same docking parameters and algorithms as S4, the drug molecules to be screened are docked with the target protein to obtain the binding mode and binding energy of each drug molecule with the target protein; S603: Feature extraction: According to the feature extraction method of S3, the same type of features are extracted from the molecular docking results and the drug molecular structure, and then normalized and encoded to form a feature vector; S604: The extracted feature vector is input into the trained and optimized machine learning model to predict the activity of drug molecules. According to the prediction results, all drug molecules to be screened are sorted. For the regression model, the predicted activity values ​​are sorted from large to small to screen out drug molecules with higher activity values. For the classification model, drug molecules predicted to be active are screened out and further sorted according to the model's prediction probability, ultimately obtaining a set of drug molecules with high activity potential.

[0012] A drug molecule screening system that integrates molecular docking and machine learning includes a data management module, a molecular docking module, a feature extraction module, a machine learning module, a model evaluation and optimization module, and a drug molecule screening and execution module.

[0013] Preferably, the data management module is used to collect, clean and standardize drug molecule and target protein data, eliminate redundant and erroneous information, unify the data format, and provide a high-quality and standardized data foundation for subsequent screening; The molecular docking module simulates the binding process between drug molecules and target proteins based on setting parameters, executing docking tasks and analyzing results, and screens out candidate drug molecules with potential binding ability.

[0014] Preferably, the feature extraction module is used to extract key features from molecular structure, docking results and physicochemical properties, integrate and encode them into feature vectors, and convert molecular information into data that can be processed by machine learning models; The machine learning module is used to select algorithms, build and train models, learn the relationship between features and drug activity, predict the activity of drug molecules, and provide model support for screening; The model evaluation and optimization module is used to evaluate model performance, adjust parameters based on the results, optimize feature engineering or fusion models, improve model accuracy and stability, and ensure screening reliability; The drug molecule screening and execution module is used to perform preprocessing, molecular docking and feature extraction on the drug molecules to be screened, use the optimized model to predict and rank the activity, and output high-potential candidate drug molecules.

[0015] Compared with the existing technology, the present invention provides a drug molecule screening method and system that integrates molecular docking and machine learning, which has the following beneficial effects: 1. This drug molecule screening method and system that integrates molecular docking and machine learning deeply integrates molecular docking and machine learning. Molecular docking technology can simulate the binding process of drug molecules and target proteins based on their three-dimensional structures, intuitively display the intermolecular interaction pattern, and quantify the binding ability through a scoring function, providing an intuitive structural level basis for screening. Machine learning technology, on the other hand, mines potential patterns from large amounts of data and builds accurate prediction models by learning the characteristics of known active drug molecules. The two work together. Molecular docking provides machine learning with rich structural features and binding information, and machine learning optimizes the screening model based on this data, avoiding the limitations of a single technology. Compared with traditional single screening methods, it can more accurately evaluate the activity of drug molecules and increase the probability of screening high-potential drug candidates.

[0016] 2. This drug molecule screening method and system that integrates molecular docking and machine learning realizes the automation and intelligence of the entire drug molecule screening process. The data management module automatically completes data collection, cleaning and standardization; the molecular docking module automatically performs docking tasks and analyzes the results according to preset parameters; the machine learning module can automatically complete model training, evaluation and optimization; at the same time, it quickly processes new data and outputs screening results. This automated process greatly reduces manual operations and reduces human errors. At the same time, through the application of intelligent algorithms, it can quickly process massive data, greatly improving screening efficiency and shortening the drug development cycle compared to traditional manual or semi-manual screening methods.

[0017] 3. This drug molecule screening method and system that integrates molecular docking and machine learning can collect data from multiple data sources and perform in-depth cleaning and standardization. During the data collection stage, it not only covers public databases but also can access internal enterprise data. The feature extraction module can further extract drug molecule features from multiple dimensions, including molecular structure, docking results, physicochemical properties, etc., providing rich and valuable information for machine learning, enabling the model to fully learn the characteristics of drug molecules. Compared with the existing technology with insufficient data processing and single feature extraction, it can more comprehensively reflect the essence of drug molecules. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0019] In order to make the technical means, creative features, objectives and effects achieved by the present invention easier to understand, the present invention is further described below in conjunction with specific implementation methods. Example

[0020] like Figure 1 As shown, a drug molecule screening method integrating molecular docking and machine learning includes the following steps: S1: Data preparation: Collect the molecular structure data of drugs with known activity and the three-dimensional structure data of the corresponding target proteins, clean and standardize them to ensure data quality and avoid affecting the accuracy of screening results due to data problems. The specific steps include the following: S101: Collect drug molecular structure data: Download small molecule compound data with clear activity annotations from multiple public databases, including simplified molecular linear input specifications and international chemical identifiers of drug molecules, and collect drug molecular structure and activity data already available in the internal databases of drug R&D companies; S102: Collect target protein three-dimensional structural data: Retrieve protein structures related to the research target from the protein database. For target proteins whose structures have not yet been experimentally determined, use homology modeling methods to construct their three-dimensional structural models using proteins with high sequence similarity and known structures as templates; S103: Clean the collected data.

[0021] S2: Molecular docking simulation: docking the prepared drug molecules with the target protein, and calculating the binding mode and binding energy between the drug molecules and the target protein. By simulating the binding process of the drug molecules and the target protein, the binding ability of the drug molecules and the target is preliminarily evaluated, and drug molecules with potential binding possibility are screened. The specific operation steps include the following; S201: Define the docking region. Based on the functional sites and known active pockets of the target protein, define the docking search space using drug molecule prediction and screening software. For enzyme targets, the area within a certain range around the active center is set as the docking region. For receptor targets, the docking region is determined based on the location and size of the ligand binding site. S202: Select a docking algorithm, including rigid docking, semi-flexible docking, and flexible docking algorithms; S203: Setting the scoring function: Selecting an appropriate scoring function to evaluate the binding ability of the drug molecule to the target protein, including the Vina scoring function of AutoDockVina and the XP scoring function of Glide; S204: docking the prepared drug molecules with the target proteins. The standardized drug molecule structure files and target protein structure files are imported into the molecular docking module of the drug molecule prediction and screening software. At this point, docking calculations can be performed for each drug molecule and target protein one by one. S205: After the docking calculation is completed, the results are analyzed to screen out drug molecules with potential binding ability. The binding mode between the drug molecule and the target protein is checked, including the position and orientation of the drug molecule in the active pocket, and the type of interaction with the key amino acid residues. Based on this, the rationality of the binding between the drug molecule and the target is preliminarily judged. According to the binding energy value calculated by the scoring function, a threshold is set to screen out drug molecules with lower binding energy as candidate molecules, while excluding molecules with unreasonable binding modes.

[0022] S3: Feature extraction: extract key features from the molecular docking results, quantify and encode these features to form feature vectors for subsequent model training and prediction. The specific steps include the following: S301: Molecular structure feature extraction, including the extraction of functional group features and topological structure features of drug molecules; S302: Extracting docking result features: extracting binding-related features from the molecular docking results, including binding energy features and binding site features; S303: Physical and chemical property feature extraction, calculating various physical and chemical property features of drug molecules, including basic properties, lipophilicity characteristics and solubility characteristics; S304: Integrate the various features extracted above, and perform unified coding to form a feature vector.

[0023] S4: Machine learning model training: Select an appropriate machine learning algorithm, use the extracted feature vectors and the corresponding drug molecule activity data as a training set, and input them into the model for training. This allows the model to predict drug activity based on its molecular characteristics, thereby achieving efficient screening of drug molecules. The specific steps include the following: S401: Algorithm selection: Based on the task requirements and data characteristics of drug molecule screening, select appropriate machine learning algorithms to build models, including classification algorithms and regression algorithms; S402: Build a model using the required machine learning algorithm, divide the extracted feature vectors and the corresponding drug molecule activity data into a training set and a test set in a ratio of 7:3, and set the initial parameters of the model based on the characteristics of the selected algorithm; S403: Input the training set data into the constructed model for training.

[0024] S5: Model evaluation and optimization: Use a portion of untrained data as a test set to evaluate the trained machine learning model to ensure the model has high reliability and predictive ability, providing an accurate basis for subsequent drug molecule screening. The specific steps include the following: S501: Model evaluation: Use the test set data to comprehensively evaluate the trained machine learning model, including classification model evaluation indicators: evaluate the model based on the calculation of accuracy, precision, recall rate, and F1 value indicators; It also includes regression model evaluation indicators: including calculation of mean square error, root mean square error, mean absolute error, and coefficient of determination to evaluate the model; S502: Based on the evaluation results, the model is optimized and improved, including parameter adjustment, feature engineering optimization, and model fusion.

[0025] S6: Drug molecule screening: The drug molecule data to be screened are subjected to molecular docking and feature extraction according to the above steps. The data are then input into the trained and optimized machine learning model. The model predicts the activity of the drug molecules. Based on the prediction results, the drug molecules are ranked according to their activity to screen out drug molecules with high activity potential. The specific steps include the following: S601: Preprocessing the drug molecule data to be screened, including converting the structural data of the drug molecules to be screened into a unified format and performing structural normalization to ensure the chemical rationality and consistency of the molecular structure, obtaining the corresponding target protein three-dimensional structure, and performing hydrogenation, dehydration, and removal of irrelevant ligands to ensure that the target protein structure meets the requirements of molecular docking; S602: Molecular docking: Using the same docking parameters and algorithms as S4, the drug molecules to be screened are docked with the target protein to obtain the binding mode and binding energy of each drug molecule with the target protein; S603: Feature extraction: According to the feature extraction method of S3, the same type of features are extracted from the molecular docking results and the drug molecular structure, and then normalized and encoded to form a feature vector; S604: The extracted feature vector is input into the trained and optimized machine learning model to predict the activity of drug molecules. According to the prediction results, all drug molecules to be screened are sorted. For the regression model, the predicted activity values ​​are sorted from large to small to screen out drug molecules with higher activity values. For the classification model, drug molecules predicted to be active are screened out and further sorted according to the model's prediction probability, ultimately obtaining a set of drug molecules with high activity potential. Example

[0026] A drug molecule screening system that integrates molecular docking and machine learning, including a data management module, a molecular docking module, a feature extraction module, a machine learning module, a model evaluation and optimization module, and a drug molecule screening and execution module; The data management module is used to collect, clean, and standardize drug molecule and target protein data, eliminate redundant and erroneous information, unify data formats, and provide a high-quality, standardized data foundation for subsequent screening; The molecular docking module simulates the binding process between drug molecules and target proteins by setting parameters, executing docking tasks, and analyzing results, thereby screening candidate drug molecules with potential binding ability. The feature extraction module is used to extract key features from molecular structure, docking results, and physicochemical properties, integrate and encode them into feature vectors, and convert molecular information into data that can be processed by machine learning models; The machine learning module is used to select algorithms, build and train models, learn the relationship between features and drug activity, predict drug molecular activity, and provide model support for screening; The model evaluation and optimization module is used to evaluate model performance, adjust parameters based on the results, optimize feature engineering or fusion models, improve model accuracy and stability, and ensure screening reliability; The drug molecule screening and execution module is used to preprocess, molecular docking and feature extraction of drug molecules to be screened, use the optimized model to predict activity and rank them, and output high-potential candidate drug molecules.

[0027] It should be noted that the present invention is a drug molecule screening method and system that integrates molecular docking and machine learning.

[0028] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A drug molecule screening method integrating molecular docking and machine learning, characterized by: The following steps are included: S1: Data preparation: collecting molecular structural data of drugs with known activity and three-dimensional structural data of corresponding target proteins, cleaning and standardizing them to ensure data quality and avoid affecting the accuracy of screening results due to data problems; S2: Molecular docking simulation: docking the prepared drug molecules with the target protein, and calculating the binding mode and binding energy between the drug molecules and the target protein. By simulating the binding process of the drug molecules and the target protein, the binding ability of the drug molecules and the target is preliminarily evaluated, and drug molecules with potential binding possibility are screened; S3: Feature extraction: extract key features from the molecular docking results, quantify and encode these features to form feature vectors for subsequent model training and prediction; S4: Machine learning model training: Select an appropriate machine learning algorithm, use the extracted feature vectors and corresponding drug molecule activity data as a training set, and input them into the model for training. This will enable the model to predict drug activity based on its characteristics, thereby achieving efficient screening of drug molecules. S5: Model evaluation and optimization: Use a portion of untrained data as a test set to evaluate the trained machine learning model to ensure the model has high reliability and predictive ability, providing an accurate basis for subsequent drug molecule screening; S6: Drug molecule screening: The drug molecule data to be screened are subjected to molecular docking and feature extraction according to the above steps, and then input into the trained and optimized machine learning model. The model predicts the activity of the drug molecules. Based on the prediction results, the drug molecules are ranked according to their activity to screen out drug molecules with high activity potential.

2. The method for drug molecule screening integrating molecular docking and machine learning according to claim 1, characterized in that: Said S1 specifically includes the following steps: S101: Collect drug molecular structure data: Download small molecule compound data with clear activity annotations from multiple public databases, including simplified molecular linear input specifications and international chemical identifiers of drug molecules, and collect drug molecular structure and activity data already available in the internal databases of drug R&D companies; S102: Collect target protein three-dimensional structural data: Retrieve protein structures related to the research target from the protein database. For target proteins whose structures have not yet been experimentally determined, use homology modeling methods to construct their three-dimensional structural models using proteins with high sequence similarity and known structures as templates; S103: Clean the collected data.

3. The drug molecule screening method integrating molecular docking and machine learning according to claim 1, characterized in that: Said S2 specifically includes the following steps: S201: Define the docking region. Define the search space for docking based on the functional sites and known active pockets of the target protein. For enzyme targets, the area within a certain range around the active center is set as the docking region. For receptor targets, the docking region is determined based on the position and size of the ligand binding site. S202: Select a docking algorithm, including rigid docking, semi-flexible docking, and flexible docking algorithms; S203: Setting the scoring function: Selecting an appropriate scoring function to evaluate the binding ability of the drug molecule to the target protein, including the Vina scoring function of AutoDockVina and the XP scoring function of Glide; S204: docking the prepared drug molecules with the target protein, and performing molecular docking on the standardized drug molecule structure file and the target protein structure file. At this point, docking calculations can be performed on each drug molecule and target protein one by one; S205: After the docking calculation is completed, the results are analyzed to screen out drug molecules with potential binding ability. The binding mode between the drug molecule and the target protein is checked, including the position and orientation of the drug molecule in the active pocket, and the type of interaction with the key amino acid residues. Based on this, the rationality of the binding between the drug molecule and the target is preliminarily judged. According to the binding energy value calculated by the scoring function, a threshold is set to screen out drug molecules with lower binding energy as candidate molecules, while excluding molecules with unreasonable binding modes.

4. The drug molecule screening method integrating molecular docking and machine learning according to claim 1, characterized in that: The S3 specifically includes the following steps: S301: Molecular structure feature extraction, including the extraction of functional group features and topological structure features of drug molecules; S302: Extracting docking result features: extracting binding-related features from the molecular docking results, including binding energy features and binding site features; S303: Physical and chemical property feature extraction, calculating various physical and chemical property features of drug molecules, including basic properties, lipophilicity characteristics and solubility characteristics; S304: Integrate the various features extracted above, and perform unified coding to form a feature vector.

5. The drug molecule screening method integrating molecular docking and machine learning according to claim 1, characterized in that: The S4 specifically includes the following steps: S401: Algorithm selection: Based on the task requirements and data characteristics of drug molecule screening, select appropriate machine learning algorithms to build models, including classification algorithms and regression algorithms; S402: Build a model using the required machine learning algorithm, divide the extracted feature vectors and the corresponding drug molecule activity data into a training set and a test set in a ratio of 7:3, and set the initial parameters of the model based on the characteristics of the selected algorithm; S403: Input the training set data into the constructed model for training.

6. The drug molecule screening method integrating molecular docking and machine learning according to claim 1, characterized in that: The S5 specifically includes the following steps: S501: Model evaluation: Use the test set data to comprehensively evaluate the trained machine learning model, including classification model evaluation indicators: evaluate the model based on the calculation of accuracy, precision, recall rate, and F1 value indicators; It also includes regression model evaluation indicators: including calculation of mean square error, root mean square error, mean absolute error, and coefficient of determination to evaluate the model; S502: Based on the evaluation results, the model is optimized and improved, including parameter adjustment, feature engineering optimization, and model fusion.

7. The method for drug molecule screening integrating molecular docking and machine learning according to claim 1, characterized in that: The S6 specifically includes the following steps: S601: Preprocessing the drug molecule data to be screened, including converting the structural data of the drug molecules to be screened into a unified format and performing structural normalization to ensure the chemical rationality and consistency of the molecular structure, obtaining the corresponding target protein three-dimensional structure, and performing hydrogenation, dehydration, and removal of irrelevant ligands to ensure that the target protein structure meets the requirements of molecular docking; S602: Molecular docking: Using the same docking parameters and algorithms as S4, the drug molecules to be screened are docked with the target protein to obtain the binding mode and binding energy of each drug molecule with the target protein; S603: Feature extraction: According to the feature extraction method of S3, the same type of features are extracted from the molecular docking results and the drug molecular structure, and then normalized and encoded to form a feature vector; S604: The extracted feature vector is input into the trained and optimized machine learning model to predict the activity of drug molecules. According to the prediction results, all drug molecules to be screened are sorted. For the regression model, the predicted activity values ​​are sorted from large to small to screen out drug molecules with higher activity values. For the classification model, drug molecules predicted to be active are screened out and further sorted according to the model's prediction probability, ultimately obtaining a set of drug molecules with high activity potential.

8. A drug molecule screening system that integrates molecular docking and machine learning, using a drug molecule screening method that integrates molecular docking and machine learning as described in any one of claims 1 to 7, comprising a data management module, a molecular docking module, a feature extraction module, a machine learning module, a model evaluation and optimization module, and a drug molecule screening and execution module.

9. The drug molecule screening system integrating molecular docking and machine learning according to claim 8, characterized in that: The data management module is used to collect, clean and standardize drug molecule and target protein data, eliminate redundant and erroneous information, unify data formats, and provide a high-quality and standardized data foundation for subsequent screening; The molecular docking module simulates the binding process between drug molecules and target proteins based on setting parameters, executing docking tasks and analyzing results, and screens out candidate drug molecules with potential binding ability.

10. The drug molecule screening system integrating molecular docking and machine learning according to claim 8, characterized in that: The feature extraction module is used to extract key features from molecular structure, docking results and physicochemical properties, integrate and encode them into feature vectors, and convert molecular information into data that can be processed by machine learning models; The machine learning module is used to select algorithms, build and train models, learn the relationship between features and drug activity, predict the activity of drug molecules, and provide model support for screening; The model evaluation and optimization module is used to evaluate model performance, adjust parameters based on the results, optimize feature engineering or fusion models, improve model accuracy and stability, and ensure screening reliability; The drug molecule screening and execution module is used to perform preprocessing, molecular docking and feature extraction on the drug molecules to be screened, use the optimized model to predict and rank the activity, and output high-potential candidate drug molecules.