Embedding-Based Compound-Viral Protein Prediction Framework
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The traditional process of developing new compounds to treat emerging viral diseases like COVID-19 is time-consuming and costly, and there is a need for a more efficient method to identify existing drugs that can inhibit viral proteins effectively.
Innovation Solution
A consensus framework of in-silico embedding-based modeling techniques using SMILES strings, Morgan Fingerprints, and amino acid sequences to predict compound-viral protein interactions, which allows for cost-effective and time-efficient identification of candidate compounds with high activity against viral proteins.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional compound discovery process is used, then new compounds can be developed to treat viral diseases, but the process takes 15 years and costs $2-3 billion
Solution Approach 1:
The patent performs preliminary in-silico screening and prediction of compound-viral protein interactions before actual experimental testing. By using machine learning models to predict activity values and identify candidate compounds in advance, the system eliminates the need for lengthy traditional discovery processes while maintaining reliability through validated prediction frameworks.
Solution Approach 2:
The patent creates computational models that replicate and simulate compound-protein interactions without requiring physical experimentation. By copying the biological interaction process into a digital framework using molecular docking and machine learning, the system achieves rapid prediction without the time and cost constraints of traditional experimental development.
2Reliability
If traditional compound discovery process is used, then new compounds can be developed to treat viral diseases, but the process costs $2-3 billion
Solution Approach 1:
The patent replaces expensive physical experimental resources with computational models that simulate compound-viral protein interactions. By copying the biological screening process into a digital framework, the system eliminates the need for costly laboratory resources, reagents, and infrastructure while maintaining predictive accuracy through validated machine learning algorithms.
Solution Approach 2:
The patent uses inexpensive computational resources and pre-existing data frameworks to perform compound screening. Instead of investing millions in traditional discovery infrastructure, the system leverages open-access databases, publicly available protein structures, and cost-effective machine learning models to achieve reliable compound identification.
3Measurement precision
If molecular docking with high-quality three-dimensional crystal structures is used, then accurate binding affinity predictions can be obtained, but the process requires complex structural data and annotation information
Solution Approach 1:
The patent changes the input parameters from requiring high-quality three-dimensional crystal structures to accepting simplified molecular representations like SMILES strings and Morgan Fingerprints. By transforming the structural data requirements into more flexible parameter formats, the system maintains prediction accuracy while eliminating the complexity of obtaining and processing high-resolution crystallographic data.
Solution Approach 2:
The patent replaces the need for expensive, complex structural data collection and processing infrastructure with simpler, more accessible molecular representations. By using lightweight computational methods that process simplified input formats, the system achieves binding affinity predictions without requiring sophisticated structural biology resources.
Data Source
AI summary
A global effort is underway to identify compounds to treat emerging virus infections, such as COVID-19. Since de novo compound design is an extremely long, time-consuming, and expensive process, efforts are underway to discover existing compounds that can be repurposed for COVID-19 and new viral diseases. The present invention discloses a machine learning representation framework that uses deep learning-induced vector embeddings of compounds and viral proteins as features to predict compound-viral protein activity. The prediction model uses a consensus framework to rank approved compounds against viral proteins of interest.


