Embedding-Based Compound-Viral Protein Prediction Framework

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The traditional process of developing new compounds to treat emerging viral diseases like COVID-19 is time-consuming and costly, and there is a need for a more efficient method to identify existing drugs that can inhibit viral proteins effectively.

Innovation Solution

A consensus framework of in-silico embedding-based modeling techniques using SMILES strings, Morgan Fingerprints, and amino acid sequences to predict compound-viral protein interactions, which allows for cost-effective and time-efficient identification of candidate compounds with high activity against viral proteins.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional compound discovery process is used, then new compounds can be developed to treat viral diseases, but the process takes 15 years and costs $2-3 billion

Engineering Contradiction:
Improveeffectiveness of treatmentVSAvoidtime required for compound development
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary in-silico screening and prediction of compound-viral protein interactions before actual experimental testing. By using machine learning models to predict activity values and identify candidate compounds in advance, the system eliminates the need for lengthy traditional discovery processes while maintaining reliability through validated prediction frameworks.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates computational models that replicate and simulate compound-protein interactions without requiring physical experimentation. By copying the biological interaction process into a digital framework using molecular docking and machine learning, the system achieves rapid prediction without the time and cost constraints of traditional experimental development.

Inventive Principle:
Principle #26Copying

2Reliability

If traditional compound discovery process is used, then new compounds can be developed to treat viral diseases, but the process costs $2-3 billion

Engineering Contradiction:
Improveeffectiveness of treatmentVSAvoidresource consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent replaces expensive physical experimental resources with computational models that simulate compound-viral protein interactions. By copying the biological screening process into a digital framework, the system eliminates the need for costly laboratory resources, reagents, and infrastructure while maintaining predictive accuracy through validated machine learning algorithms.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent uses inexpensive computational resources and pre-existing data frameworks to perform compound screening. Instead of investing millions in traditional discovery infrastructure, the system leverages open-access databases, publicly available protein structures, and cost-effective machine learning models to achieve reliable compound identification.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Measurement precision

If molecular docking with high-quality three-dimensional crystal structures is used, then accurate binding affinity predictions can be obtained, but the process requires complex structural data and annotation information

Engineering Contradiction:
Improvebinding affinity prediction accuracyVSAvoidstructural data requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the input parameters from requiring high-quality three-dimensional crystal structures to accepting simplified molecular representations like SMILES strings and Morgan Fingerprints. By transforming the structural data requirements into more flexible parameter formats, the system maintains prediction accuracy while eliminating the complexity of obtaining and processing high-resolution crystallographic data.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the need for expensive, complex structural data collection and processing infrastructure with simpler, more accessible molecular representations. By using lightweight computational methods that process simplified input formats, the system achieves binding affinity predictions without requiring sophisticated structural biology resources.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Data Source

PatentUS20220392567A1Modelling framework for embedding-based predictions for compound-viral protein activity
Publication Date: 2022.12.08 HAMAD BIN KHALIFA UNIVERSITY
  • US20220392567A1 patent drawing
  • US20220392567A1 patent drawing
  • US20220392567A1 patent drawing

AI summary

A global effort is underway to identify compounds to treat emerging virus infections, such as COVID-19. Since de novo compound design is an extremely long, time-consuming, and expensive process, efforts are underway to discover existing compounds that can be repurposed for COVID-19 and new viral diseases. The present invention discloses a machine learning representation framework that uses deep learning-induced vector embeddings of compounds and viral proteins as features to predict compound-viral protein activity. The prediction model uses a consensus framework to rank approved compounds against viral proteins of interest.