Graph Neural Network for Biologic Binding Interface Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current software and computational tools, including artificial intelligence and machine learning, have limited success in predicting the complex properties and molecular behavior of biologics, which are essential for drug discovery and development, due to their intricate structures.

Innovation Solution

The use of graph-based neural networks to predict amino acid sequences and structures at the binding interface of custom biologics by receiving input graph representations of target molecules and partially defined biologics, generating predictions for unknown amino acid sequences and types to enhance binding capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If traditional computational tools and machine learning are used to predict biologic properties, then the prediction process can be automated, but the accuracy and reliability of predictions remain insufficient due to the complexity of biologic structures

Engineering Contradiction:
Improveprediction process automationVSAvoidprediction accuracy
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent replaces traditional machine learning approaches with a diffusion probabilistic model that uses physics-informed neural networks to simulate molecular dynamics and protein folding processes. This substitution of computational mechanics with physics-based modeling improves prediction reliability while maintaining automation, directly addressing the contradiction between automated prediction and accurate results for complex biologics

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If in vitro testing, in vivo testing, and clinical trials are conducted for new drug candidates, then comprehensive safety and efficacy data are obtained, but the time and cost requirements become enormous

Engineering Contradiction:
Improvesafety and efficacy data qualityVSAvoiddrug development time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs in silico design and prediction of biologic properties before physical experiments are conducted. The diffusion probabilistic model predicts amino acid sequences, binding interfaces, and molecular behavior in advance, allowing researchers to prioritize the most promising candidates for in vitro and in vivo testing. This preliminary computational screening reduces the number of candidates requiring extensive experimental validation, thereby reducing overall development time while maintaining data quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates virtual copies of biologic molecules and their interactions through computational modeling. The diffusion model generates predicted three-dimensional structures and binding interfaces that serve as digital twins for experimental validation. This copying approach allows comprehensive testing of multiple variants in silico before committing resources to physical experiments, reducing both time and cost while preserving reliability through subsequent experimental verification

Inventive Principle:
Principle #26Copying

3Measurement precision

If the complete amino acid sequence of a biologic is known, then accurate structure prediction and binding interface identification are possible, but obtaining complete sequence information for partially defined biologics remains challenging

Engineering Contradiction:
Improvesequence determination accuracyVSAvoidsequence determination complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and focuses specifically on predicting only the binding interface region amino acid sequences rather than determining the complete protein sequence. The diffusion probabilistic model is trained to predict amino acid types at binding interfaces using graph representations that highlight interface residues. This extraction approach reduces the complexity of sequence determination by concentrating computational resources on the most critical region for binding function, while maintaining high precision for interface prediction

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20240038337A1Systems and methods for artificial intelligence-based prediction of amino acid sequences
Publication Date: 2024.02.01 PYTHIA LABS INC
  • US20240038337A1 patent drawing
  • US20240038337A1 patent drawing
  • US20240038337A1 patent drawing

AI summary

Presented herein are systems and methods for prediction of protein sequences, such as interfaces and/or other portions of custom biologics, e.g., for binding to target molecules. In certain embodiments, technologies described herein utilize graph-based neural networks to predict portions of protein/peptide structures of a custom biologic (e.g., a protein and/or peptide) that is being designed.