Medicine comprehensive evaluation system and method based on machine learning and expert database
Through the comprehensive drug evaluation system combining machine learning and expert databases, traditional drugs have been solved, with long research and development cycles, high cost and poor prediction effects, and efficient integration and interpretability prediction of multimodal data are achieved, adapting to the clinical environment and protecting data privacy, improving the efficiency and accuracy of new drug research and development.
Patent Information
- Application Number
- CN202510365366.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional drugs have a long development cycle and high cost. Machine learning prediction methods have poor prediction effects on rare diseases or new target drugs. They lack systematic integration of field expert knowledge, insufficient fusion of multimodal data, and insufficient interpretability of prediction results.
The comprehensive drug evaluation system based on machine learning and expert database is adopted, including multi-source data processing module, expert knowledge base, prediction model cluster, dynamic optimization module and interpretability output module. Through GraphCNN, L1 regularized logistic regression, SMOTE-ENN algorithm, deep neural network, integrated learning model and symbolic reasoning engine, multimodal data fusion and dynamic optimization are achieved to provide interpretable prediction results.
It significantly improves the prediction accuracy and efficiency of new drug research and development, especially in small samples or new drug scenarios, which can learn a wide range of knowledge from limited data, optimize the expert knowledge base, improve prediction accuracy and interpretability, reduce the risk of misjudgment, adapt to the clinical environment and protect data privacy.
Smart Images

Figure CN120260970A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of drug research and development and artificial intelligence, and specifically relates to a drug comprehensive evaluation system and method combining machine learning algorithms and expert knowledge bases, which is applicable to new drug research and development, drug repositioning, and optimization of personalized treatment plans. Background Art
[0002] Traditional drug research and development has problems such as a long cycle (average 10 - 15 years) and high cost (about $2.6 billion / drug). The machine learning prediction methods in the prior art have the following defects:
[0003] 1. Over - reliance on the quality of training data, with poor prediction effects for drugs for rare diseases or new target drugs;
[0004] 2. Lack of systematic integration of domain expert knowledge;
[0005] 3. Insufficient multi - modal data fusion (molecular structure, genomics, clinical data, etc.);
[0006] 4. Insufficient interpretability of prediction results. Summary of the Invention
[0007] The present invention aims to solve the above problems of the prior art. A drug comprehensive evaluation system and method based on machine learning and an expert library are proposed. The technical solutions of the present invention are as follows:
[0008] A drug comprehensive evaluation system based on machine learning and an expert library, comprising:
[0009] A multi - source data processing module for processing and standardizing drug molecular structures, genomic features, and clinical data; an expert knowledge base containing drug attributes, disease characteristics, and clinical guidelines; a prediction model cluster composed of a deep neural network, an ensemble learning model, and a symbolic reasoning engine for predicting the IC50 value, clinical response rate, and adverse reaction probability of drugs; a dynamic optimization module for uncertainty quantification of the model, incremental learning of the training set, and dynamic update of the knowledge base; an interpretability output module for interpreting the prediction results and visualizing the output.
[0010] Further, the multi - source data processing module further includes: a molecular structure parser for converting the molecular structure into the SMILES or InChI format and performing feature encoding through GraphCNN; a genomic feature extractor for screening key SNP sites using L1 - regularized logistic regression; a clinical data standardization unit for dealing with the class imbalance problem using the SMOTE - ENN algorithm to ensure the comprehensiveness and accuracy of the data.
[0011] Furthermore, the dynamic optimization module further includes: a transfer learning unit for the adaptive adjustment of cross-domain models; an active learning controller that selects the most informative samples through an uncertainty sampling strategy to update the model, improving the model learning efficiency and prediction accuracy.
[0012] Furthermore, the deep neural network in the prediction model cluster adopts CNN and Transformer architectures, where CNN performs 3D convolutional encoding on molecular graph data, and Transformer is used to process gene sequence information; the ensemble learning model includes XGBoost and Random Forest for integrating the prediction results of multi-source data; the symbolic reasoning engine is used to combine model predictions with the rules in the expert knowledge base to enhance the clinical reliability of the predictions.
[0013] Furthermore, the interpretability output module further includes: a SHAP value calculator for quantifying the contribution of features to the prediction results; a knowledge graph visualization unit that improves the transparency and understandability of the model output by querying the relationship between the knowledge graph and the prediction results, facilitating clinical decision-making and scientific research analysis.
[0014] A drug efficacy prediction method based on the system according to any one of the above, comprising the following steps:
[0015] Data preprocessing stage:
[0016] - Molecular structure encoding: Processing molecular graph data using GraphCNN with a 3D convolutional kernel (5×5×5);
[0017] - Genome feature selection: Using L1-regularized logistic regression to screen key SNP loci;
[0018] - Clinical data cleaning: Applying the SMOTE-ENN algorithm to handle the class imbalance problem;
[0019] Model training stage:
[0020] - Multi-task learning framework: Simultaneously predicting the IC50 value, the clinical response rate ORR, and the probability of adverse reactions;
[0021] - Knowledge distillation mechanism: Using expert rules to constrain the model output space;
[0022] - Dynamic weight allocation: Adjusting the contribution degree of each data source through a gating network
[0023] Prediction optimization stage:
[0024] - Uncertainty quantification: Using Monte Carlo Dropout to estimate the prediction confidence interval;
[0025] - Knowledge base retrieval: Matching historical cases based on Jaccard similarity;
[0026] - Result correction: Applying the Mamdani type of fuzzy logic rules to adjust the original predicted value.
[0027] The advantages and beneficial effects of the present invention are as follows:
[0028] The innovation points of the present invention and their corresponding beneficial effects are as follows:
[0029] Innovation point 1: Two-way knowledge flow mechanism
[0030] - Innovation description: In traditional machine learning frameworks, expert rules are usually only used for constraints in the model training stage. However, the present invention innovatively designs a two-way knowledge flow mechanism, enabling expert knowledge to not only guide the model during training but also feedback to the knowledge base after prediction, forming the automatic discovery and integration of new knowledge.
[0031] - Beneficial effects: Significantly improved the prediction accuracy of the model. Especially in the scenarios of small samples or new drugs, the model can learn more extensive knowledge from limited training data. At the same time, through the prediction feedback mechanism, the expert knowledge base is continuously optimized and enriched, realizing the dynamic update and growth of knowledge.
[0032] - Reason not easily thought of: The two-way knowledge flow mechanism breaks the one-way dependence mode between traditional AI models and expert rules, requiring designers to not only deeply understand the machine learning framework but also have an in-depth understanding of the construction and maintenance of the knowledge base. Achieving this mechanism requires complex balancing among algorithm design, knowledge representation, and data governance.
[0033] Innovation point 2: Heterogeneous data fusion method
[0034] - Innovation description: Proposing a multi-modal feature alignment algorithm based on tensor decomposition, which can efficiently handle the fusion problem of heterogeneous data such as drug molecular structures, genomic features, and clinical data, ensuring that the model can extract comprehensive and complementary feature information from different data sources.
[0035] - Beneficial effects: Verified on the Tox21 dataset, the system of the present invention can significantly improve the prediction accuracy. Especially in drug safety assessment, through the comprehensive analysis of multi-modal data, the prediction error is reduced, and the ability to discover rare adverse reactions is improved.
[0036] - Reason not easily thought of: The fusion of heterogeneous data has always been a technical challenge, especially when the data types are extremely different, such as the fusion of structured data and unstructured data, and quantitative data and qualitative data. This innovation point requires in-depth understanding of the characteristics of each data type and algorithmic skills for handling highly complex data structures.
[0037] Innovation Point 3: Dynamic Credibility Assessment System
[0038] - Innovation Description: Integrating triple safeguards of statistical tests, domain shift detection, and expert verification to dynamically evaluate the prediction credibility of the model, ensuring that the model can adjust in a timely manner when facing new data or domain changes and maintaining high prediction quality.
[0039] - Beneficial Effects: When facing the challenge of model drift, the dynamic credibility assessment system can identify and correct model biases in a timely manner, thereby maintaining the stability and reliability of prediction results. Especially in clinical applications, this system can reduce the risk of misjudgment caused by outdated models or data biases, improving the safety and effectiveness of patient treatment plans.
[0040] - Reason not easily thought of: The implementation of the dynamic credibility assessment system requires in-depth monitoring and understanding of the model training process, and at the same time introducing multi-disciplinary evaluation criteria (such as statistics, clinical medicine, bioinformatics), which requires inventors to have cross-domain knowledge and technology integration capabilities and is not easy to conceive in a single-disciplinary background.
[0041] Innovation Point 4: Clinical Adaptability Design
[0042] - Innovation Description: The system is designed with interfaces that conform to the HL7 FHIR (Fast Healthcare Interoperability Resources) standard to ensure seamless docking with hospital information systems; at the same time, data governance follows HIPAA (Health Insurance Portability and Accountability Act) regulations to protect patient privacy.
[0043] - Beneficial Effects: The system of the present invention can be easily integrated into existing medical IT infrastructure, reducing the deployment costs and risks for hospitals and pharmaceutical companies. At the same time, HIPAA-compliant data governance strategies ensure the legality of data processing and enhance user trust in the system.
[0044] - Reason not easily thought of: Clinical adaptability design not only requires technical compatibility and security, but also in-depth understanding of medical industry standards. Designers need to have both technical development capabilities and knowledge of industry regulations, which poses a relatively high threshold in cross-disciplinary innovation.
[0045] Innovation Point 5: Adaptive Rule Engine and Model Iteration
[0046] - Innovation Description: By introducing an adaptive rule engine, the system can apply expert rules for real-time adjustment during the prediction process. At the same time, the model iteration mechanism allows the system to self-update and optimize according to new data without manual retraining of the model.
[0047] - Beneficial effects: The adaptive rule engine ensures the clinical rationality and safety of the prediction results, while the model iteration mechanism improves the long-term adaptability and prediction performance of the system, reduces the maintenance cost, and enhances the overall value and market competitiveness of the system.
[0048] - Reasons for not being easily conceived: The design of the adaptive rule engine and model iteration requires finding a balance between the flexibility of the model and the stability of the rules, which requires designers to have a high degree of innovation awareness and problem-solving ability, and at the same time have an in-depth understanding of the periodic optimization of model training and the working principle of the rule engine, and is not easily generated naturally under the conventional thinking mode. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a preferred embodiment provided by the present invention DETAILED DESCRIPTION OF THE INVENTION
[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0051] The technical solution for the present invention to solve the above technical problems is:
[0052] Preferably, as Figure 1 shown, a comprehensive drug evaluation system based on machine learning and an expert library, which includes:
[0053] A multi-source data processing module for processing and standardizing drug molecular structures, genomic features, and clinical data; an expert knowledge base containing drug attributes, disease characteristics, and clinical guidelines; a prediction model cluster composed of a deep neural network, an ensemble learning model, and a symbolic inference engine for predicting the IC50 value, clinical response rate, and adverse reaction probability of drugs; a dynamic optimization module for uncertainty quantification of the model, incremental learning of the training set, and dynamic update of the knowledge base; an interpretability output module for interpreting the prediction results and visualizing the output.
[0054] Further, the multi-source data processing module further includes: a molecular structure parser for converting the molecular structure into SMILES or InChI format and performing feature encoding through GraphCNN; a genomic feature extractor for screening key SNP sites using L1-regularized logistic regression; a clinical data standardization unit for processing the class imbalance problem using the SMOTE-ENN algorithm to ensure the comprehensiveness and accuracy of the data.
[0055] Furthermore, the dynamic optimization module further includes: a transfer learning unit for the adaptive adjustment of cross-domain models; an active learning controller that selects the most informative samples through an uncertainty sampling strategy to update the model, improving the model learning efficiency and prediction accuracy.
[0056] Furthermore, the deep neural network in the prediction model cluster adopts CNN and Transformer architectures, where CNN performs 3D convolutional encoding on molecular graph data, and Transformer is used to process gene sequence information; the ensemble learning model includes XGBoost and Random Forest for integrating the prediction results of multi-source data; the symbolic inference engine is used to combine the model prediction with the rules in the expert knowledge base to enhance the clinical reliability of the prediction.
[0057] Furthermore, the interpretability output module further includes: a SHAP value calculator for quantifying the contribution of features to the prediction result; a knowledge graph visualization unit that improves the transparency and understandability of the model output by querying the relationship between the knowledge graph and the prediction result, facilitating clinical decision-making and scientific research analysis.
[0058] A drug efficacy prediction method based on the system according to any one of the above, comprising the following steps:
[0059] Data preprocessing stage:
[0060] - Molecular structure encoding: Processing molecular graph data using GraphCNN with a 3D convolutional kernel (5×5×5);
[0061] - Genome feature selection: Using L1-regularized logistic regression to screen key SNP loci;
[0062] - Clinical data cleaning: Applying the SMOTE-ENN algorithm to handle the class imbalance problem;
[0063] Model training stage:
[0064] - Multi-task learning framework: Simultaneously predicting the IC50 value, the clinical response rate ORR, and the probability of adverse reactions;
[0065] - Knowledge distillation mechanism: Using expert rules to constrain the model output space;
[0066] - Dynamic weight allocation: Adjusting the contribution degree of each data source through a gating network
[0067] Prediction optimization stage:
[0068] - Uncertainty quantification: Using Monte Carlo Dropout to estimate the prediction confidence interval;
[0069] - Knowledge base retrieval: Match historical cases based on Jaccard similarity;
[0070] - Result correction: Apply the Mamdani type of fuzzy logic rules to adjust the original predicted value.
[0071] Furthermore, the multi-source data processing module further includes:
[0072] - Molecular structure encoding: In addition to using GraphCNN, the GraphAttention Network (GAT) applying the self-attention mechanism should also be considered to capture high-level features in the molecular graph and enhance the model's understanding depth of the molecular structure.
[0073] - Genome feature selection: Deep learning feature selection techniques, such as the deep LSTM network, can be introduced to identify time series features in genomic data and improve the ability to capture dynamic changes in gene expression.
[0074] - Clinical data cleaning: In addition to SMOTE-ENN, AutoML tools such as the automatic feature engineering function of H2O can also be used to automatically detect and process outliers and missing values in the data and enable automated feature selection.
[0075] Expert knowledge base
[0076] - Knowledge graph construction: Use knowledge graph construction tools, such as Google's Knowledge Graph or IBM's Watson Discovery, combined with NLP technology to automatically extract disease-related knowledge from a large number of medical literatures, automatically construct and update the knowledge graph, and ensure the timeliness and comprehensiveness of the knowledge base.
[0077] Prediction model cluster
[0078] - Model fusion strategy: In addition to the gating network adjusting the contribution degree of data sources, a Softmax fusion layer of deep learning can also be introduced to enable the model to automatically learn the importance of different data sources during prediction and further improve the flexibility and accuracy of prediction.
[0079] - Uncertainty quantification: In addition to Monte Carlo Dropout, the Bayesian Neural Network (BNN) can be considered to estimate the uncertainty of model parameters using Bayesian inference and provide a more reliable confidence interval estimate.
[0080] Dynamic optimization module
[0081] - Knowledge base retrieval: Adopt deep information retrieval techniques such as the BERT model to conduct in-depth queries on the knowledge graph, enhance the efficiency and accuracy of matching historical cases, and at the same time consider introducing semantic similarity calculation methods to improve the matching quality.
[0082] - Result correction: In addition to fuzzy logic rule adjustment, the decision tree model of experts can be combined for secondary correction of the results, introducing the intuition and experience of experts to improve the clinical applicability of the prediction results.
[0083] Interpretability output module
[0084] - Prediction report generation: Utilize natural language generation technology (NLG) to automatically generate prediction reports that comply with the FAIR principles (findable, accessible, interoperable, reusable), facilitating non-technical background doctors and pharmacists to understand the prediction results.
[0085] - User interface design: Design an intuitive user interface, integrating prediction results, feature importance charts, knowledge graph visualization, and expert rule explanations, enabling users to quickly understand the logic and data sources behind the prediction, and enhancing the transparency and credibility of clinical decisions. Specific embodiments
[0087] Embodiment 1: Efficacy prediction of the integration of deep learning and knowledge graph
[0088] System design
[0089] 1. Data preprocessing module:
[0090] - Molecular structure encoding: Use the RDKit tool to preprocess the compound, extract its SMILES format, and then encode it through GraphCNN (such as GIN, Graph Isomorphism Network). Considering the complex interactions between molecules, a 3D convolutional kernel is used for feature capture, with a radius of 3 and a bit number of 1024 to obtain richer structural information.
[0091] - Disease genome feature extraction: Collect RNA-seq data from databases such as the International Cancer Genome Consortium (ICGC) and The Cancer Genome Atlas (TCGA), use PCA to reduce the dimension to 500 principal components, and at the same time combine single-cell sequencing data to capture cell heterogeneity and enhance the model's understanding of disease heterogeneity.
[0092] - Clinical data cleaning: Deeply clean the clinical data, use SMOTE-ENN to handle the class imbalance problem, and introduce a deep learning-based anomaly detection mechanism to identify and handle the abnormal points in the data.
[0093] 2. Expert knowledge base construction:
[0094] - Construct a knowledge graph containing 150,000 entities and 2.3 million relationships using the Neo4j graph database. The entities include drugs, diseases, genes, clinical trials, etc., and the relationships cover drug-gene interactions, disease-gene associations, drug-disease efficacy, etc.
[0095] - Adopt automatic text mining technology to extract expert rules and knowledge from databases such as PubMed and ClinicalTrials.gov, such as drug metabolic pathways, disease phenotypes, drug-disease adaptability, etc. After being verified by domain experts, they are incorporated into the knowledge base.
[0096] 3. Model construction and training:
[0097] - Deep neural network: Use a deep neural network with GraphCNN and Transformer architectures. GraphCNN processes molecular structure information, and Transformer processes disease genomic data. The features of both are combined through the BilinearFusion mechanism.
[0098] - Knowledge distillation: Introduce a knowledge distillation mechanism during model training. Use the rules in the expert knowledge base as guidance to constrain the model output, avoid the model overfitting to the training data, and improve the generalization ability of the model.
[0099] - Dynamic weight allocation: Design a gating network to automatically adjust the contribution degrees of GraphCNN, Transformer, and knowledge distillation, ensuring that the model can adaptively learn from multi-source data.
[0100] 4. Dynamic optimization and interpretable output:
[0101] - Uncertainty quantification: Use Monte Carlo Dropout and Bayesian Neural Network techniques to quantify the uncertainty of predictions and provide the confidence interval of the prediction results.
[0102] - Knowledge graph visualization: Integrate the Cypher query language and visualization tools such as D3.js to display the association network between drugs, diseases, and genes, enhancing the interpretability of the prediction results.
[0103] - Result correction: Introduce dynamic correction of fuzzy logic and expert rules to ensure that the prediction results match clinical practice and improve the clinical value of the predictions.
[0104] Beneficial effects:
[0105] -Improved prediction accuracy: On various drug and disease datasets such as Tox21 and GDSC, the prediction accuracy of the model has been significantly improved. Especially in the scenario of new drug research and development, the ability to predict the efficacy of unknown drugs has been enhanced, accelerating the drug discovery process.
[0106] -Improved data and computational efficiency: Through efficient data preprocessing and model design, the time and computational resources required to process large datasets are reduced, enabling the model to be quickly trained and predicted in a limited computational environment.
[0107] Example 2: Prediction of antibiotic efficacy based on ensemble learning and expert rules
[0108] System design
[0109] 1. Feature engineering:
[0110] -Calculation of molecular descriptors: Use the DRAGON tool to calculate 2D chemical descriptors of drugs and ChemAxon to calculate 3D chemical descriptors, with a total of more than 2000-dimensional features.
[0111] -Extraction of clinical features: Adopt ICD-11 coding, combined with the Charlson Comorbidity Index (CCI) and patient biomarker data, to construct a comprehensive clinical feature set.
[0112] 2. Model architecture and training:
[0113] -Ensemble learning model: Design an ensemble learning framework composed of XGBoost, LightGBM, and CatBoost. Each model is independently trained, and then the final prediction is made through the Stacking integration technique.
[0114] -Expert rule engine: Use the Drools rule engine to correct the model prediction results according to the rules provided by domain experts, such as antibiotic resistance, patient allergy history, and organ function status.
[0115] 3. Rule application and dynamic update:
[0116] -Clinical dose adjustment rule: Automatically adjust the antibiotic dose according to the patient's GFR (glomerular filtration rate) and drug classification to ensure the accuracy and safety of personalized medication.
[0117] -Model drift detection and incremental learning: Implement the Page-Hinkley test to regularly check the model performance. Once a performance decline is detected, update the training set through the Reservoir Sampling technique and retrain the model to adapt to new clinical data.
[0118] Beneficial effects:
[0119] - Optimization of personalized treatment plans: By combining the specific clinical characteristics of patients and the chemical descriptors of drugs, the system can provide personalized antibiotic treatment plans for each patient, reducing the risk of antibiotic abuse and drug resistance.
[0120] - Support for regulatory decision-making: Provide real-time assessment tools for drug regulatory authorities on the safety and effectiveness of drugs, helping decision-makers promptly understand the clinical manifestations of antibiotics and enhancing the scientific nature and timeliness of decision-making.
[0121] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions.
[0122] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising said element.
[0123] The above embodiments should be understood as being only for illustrative purposes of the present invention and not for limiting the protection scope of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and such equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A comprehensive drug evaluation system based on machine learning and an expert library, characterized in that, Including: A multi-source data processing module for processing and standardizing drug molecular structures, genomic features, and clinical data; An expert knowledge base containing drug properties, disease characteristics, and clinical guidelines; A prediction model cluster consisting of a deep neural network, an ensemble learning model, and a symbolic reasoning engine for predicting the IC50 value, clinical response rate, and adverse reaction probability of drugs; A dynamic optimization module for uncertainty quantification of the model, incremental learning of the training set, and dynamic update of the knowledge base; An interpretability output module for interpreting the prediction results and visualizing the output.
2. The drug comprehensive evaluation system based on machine learning and expert library according to claim 1, characterized in that The multi-source data processing module further includes: A molecular structure parser for converting the molecular structure into SMILES or InChI format and performing feature encoding through GraphCNN; A genomic feature extractor for screening key SNP sites using L1-regularized logistic regression; A clinical data standardization unit for processing class imbalance problems using the SMOTE-ENN algorithm to ensure the comprehensiveness and accuracy of the data.
3. The drug comprehensive evaluation system based on machine learning and expert database according to claim 1, characterized in that The dynamic optimization module further includes: A transfer learning unit for adaptive adjustment of cross-domain models; An active learning controller for selecting the most informative samples through an uncertainty sampling strategy to update the model and improve the model learning efficiency and prediction accuracy.
4. The drug comprehensive evaluation system based on machine learning and expert library according to claim 1, characterized in that The deep neural network in the prediction model cluster adopts CNN and Transformer architectures, where CNN performs 3D convolutional encoding on molecular graph data and Transformer is used to process gene sequence information; The ensemble learning model includes XGBoost and Random Forest for integrating the prediction results of multi-source data; The symbolic reasoning engine for combining model predictions with rules in the expert knowledge base to enhance the clinical reliability of the predictions.
5. The drug comprehensive evaluation system based on machine learning and expert database according to claim 1, characterized in that The interpretability output module further includes: A SHAP value calculator for quantifying the contribution of features to the prediction results; A knowledge graph visualization unit for improving the transparency and understandability of the model output by querying the relationship between the knowledge graph and the prediction results, facilitating clinical decision-making and scientific research analysis.
6. A method for predicting drug efficacy based on the system according to any one of claims 1-5, characterized in that, Including the following steps: Data preprocessing stage: - Molecular structure encoding: Processing molecular graph data using GraphCNN with a 3D convolutional kernel (5×5×5); - Genomic feature selection: Screening key SNP sites using L1-regularized logistic regression; - Clinical data cleaning: Processing class imbalance problems using the SMOTE-ENN algorithm; Model training stage: - Multi-task learning framework: Simultaneously predicting the IC50 value, clinical response rate ORR, and adverse reaction probability; - Knowledge distillation mechanism: Using expert rules to constrain the model output space; - Dynamic weight allocation: Adjusting the contribution degree of each data source through a gating network Prediction optimization stage: - Uncertainty quantification: Estimating the prediction confidence interval using Monte Carlo Dropout; - Knowledge base retrieval: Matching historical cases based on Jaccard similarity; - Result correction: Adjusting the original prediction value using the Mamdani type of fuzzy logic rules.
Citation Information
Cited By
Road surface comprehensive service performance evaluation method, device, equipment and storage medium
CN121031358A