Graph Neural Network for Protein Interface Amino Acid Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current software and computational tools, including artificial intelligence and machine learning, have limited success in predicting the complex properties and molecular behavior of biologics, which are essential for drug discovery and development, due to their intricate structures.

Innovation Solution

The use of graph-based neural networks to predict amino acid sequences and structures at protein interfaces, allowing for the design of custom biologics that can bind effectively to target molecules by generating structural predictions for amino acid types and sequences within interface regions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional software and computational tools are used to predict properties of biologics, then the prediction process can be performed computationally, but the prediction accuracy is insufficient due to the complex structure of biologics

Engineering Contradiction:
Improveprediction accuracyVSAvoidstructure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the biologic structure into interface regions and non-interface regions. The graph neural network model specifically targets the interface regions for prediction, breaking down the complex overall structure into manageable predictive components. This segmentation allows the model to focus computational resources on the most critical binding regions, improving prediction accuracy without requiring the model to process the entire complex structure with equal detail.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different levels of predictive detail to different regions of the biologic structure. Interface regions, which are critical for binding activity, receive focused prediction attention with detailed amino acid sequence and structure prediction. Non-interface regions are handled with less computational detail. This local quality approach allocates predictive precision where it matters most, resolving the contradiction between overall accuracy and structural complexity.

Inventive Principle:
Principle #3Local quality

2Reliability

If in vitro testing, in vivo testing, and clinical trials are conducted for new drug candidates, then the safety and efficacy can be validated, but the time and cost required are enormous

Engineering Contradiction:
Improvevalidation reliabilityVSAvoiddevelopment time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary computational prediction of biologic structure and binding properties before physical experimentation. The graph neural network model predicts amino acid sequences, interface structures, and binding affinities in silico, providing preliminary data that can guide and prioritize subsequent in vitro and in vivo testing. This preliminary computational action reduces the number of candidates that need extensive experimental validation, thereby reducing overall development time while maintaining reliability through staged validation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates computational copies and models of biologic structures and their interactions with target molecules. Instead of immediately testing every possible biologic variant physically, the system generates virtual models and predictions that can be evaluated computationally. These computational copies allow for rapid screening and comparison of multiple variants before committing to expensive and time-consuming wet lab experiments, reducing development time while preserving validation reliability through iterative refinement.

Inventive Principle:
Principle #26Copying

3Strength

If the amino acid sequence and structure at the binding interface are customized for binding to a target molecule, then the binding affinity can be improved, but the complexity of design and prediction increases

Engineering Contradiction:
Improvebinding affinityVSAvoiddesign complexity
Core Design Contradiction:
StrengthVSDevice complexity

Solution Approach 1:

The patent systematically varies and optimizes parameters at the binding interface, such as amino acid types, sequences, and structural conformations. The graph neural network model evaluates multiple parameter configurations to identify combinations that maximize binding affinity. By focusing parameter optimization specifically on interface regions rather than the entire biologic structure, the system achieves improved binding affinity while managing design complexity through targeted parameter exploration.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces the graph neural network model as an intermediary between the desired binding affinity outcome and the actual biologic design. Rather than directly and complexly manipulating all structural parameters to achieve binding optimization, the model serves as an intelligent intermediary that translates binding requirements into specific amino acid sequence and structure predictions. This intermediary approach simplifies the design process by providing guided predictions based on learned patterns from training data, reducing design complexity while achieving improved binding affinity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240096444A1Systems and methods for artificial intelligence-based prediction of amino acid sequences at a binding interface
Publication Date: 2024.03.21 PYTHIA LABS INC
  • US20240096444A1 patent drawing
  • US20240096444A1 patent drawing
  • US20240096444A1 patent drawing

AI summary

Presented herein are systems and methods for prediction of protein interfaces for binding to target molecules. In certain embodiments, technologies described herein utilize graph-based neural networks to predict portions of protein/peptide structures that are located at an interface of custom biologic (e.g., a protein and/or peptide) that is being designed for binding to a target molecule, such as another protein or peptide. In certain embodiments, graph-based neural network models described herein may receive, as input, a representation (e.g., a graph representation) of a complex comprising a target and a partially-defined custom biologic. Portions of the partially-defined custom biologic may be known, while other portions, such an amino acid sequence and/or particular amino acid types at certain locations of an interface, are unknown and/or to be customized for binding to a particular target. A graph-based neural network model as described herein may then, based on the received input, generate predictions of likely acid sequences and/or types of particular amino acids at the unknown portions. These predictions can then be used to determine (e.g., fill in) amino acid sequences and/or structures to complete the custom biologic.