Off-Target Protein Identification Using Whole-Sequence Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for identifying off-target proteins are computationally inefficient and require prior knowledge of pocket sequences and 3D structures, limiting their effectiveness and applicability.

Innovation Solution

A method that compares whole protein sequences using multiple sequence alignment and residue matching, eliminating the need for prior knowledge of pocket locations and reducing the search space by identifying similar overall protein sequences, thereby improving computational efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If pocket-to-pocket comparison is performed across the human proteome, then off-target identification accuracy is improved, but computational cost increases significantly

Engineering Contradiction:
Improveoff-target identification accuracyVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The patent segments the protein comparison task into two distinct phases: (1) whole protein sequence comparison to identify candidate proteins with overall sequence similarity, and (2) pocket sequence comparison restricted only to these candidates. This segmentation reduces the computational burden by avoiding exhaustive pocket-to-pocket comparisons across all 20,000 human proteins, while maintaining identification accuracy through the two-stage filtering approach.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by performing whole protein sequence comparison as a pre-filtering step before conducting the computationally intensive pocket sequence analysis. By identifying candidate proteins with overall sequence resemblance first, the method prepares a reduced subset of proteins for detailed pocket comparison, thereby reducing computational cost while preserving accuracy.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If pocket sequence comparison is performed, then off-target identification reliability is improved, but the method requires prior knowledge of pocket sequences which limits applicability

Engineering Contradiction:
Improveoff-target identification reliabilityVSAvoidmethod applicability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent inverts the traditional approach by performing whole protein sequence comparison first (instead of starting with pocket sequences) to identify candidate proteins. This inversion allows the method to operate without requiring prior pocket sequence knowledge for the query protein, while still enabling reliable off-target identification through subsequent pocket comparison on the candidate set.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent uses preliminary whole protein sequence comparison to identify candidate proteins before performing pocket sequence analysis. This preliminary action enables the method to work with proteins where pocket sequences are not预先 known, as the candidate identification phase only requires overall sequence information, thereby improving method applicability while maintaining reliability through the subsequent pocket comparison step.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If whole protein sequence comparison is performed first, then computational efficiency is improved, but the search space remains large

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidsearch space size
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent extracts candidate proteins from the large human proteome database by performing whole protein sequence comparison and applying a similarity threshold filter. This extraction process removes proteins that do not meet the sequence similarity criterion, reducing the search space from approximately 20,000 proteins to a much smaller candidate set that proceeds to the detailed pocket comparison stage, thereby improving computational efficiency.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP4600965A1Method for identifying off-target proteins
Publication Date: 2025.08.13 BENEVOLENTAI TECH LTD
  • EP4600965A1 patent drawingFigure 1
  • EP4600965A1 patent drawingFigure 2
  • EP4600965A1 patent drawingFigure 3

AI summary

A computer-implemented method for identifying off-target proteins is disclosed. The method comprises: receiving an indication of a drug target, wherein the drug target is a first protein comprising residues of interest for targeting; receiving data indicative of a whole protein sequence corresponding to the first protein; comparing the whole protein sequence of the first protein against a protein sequence database to identify whole protein sequences of other proteins having a threshold level of sequence resemblance to the whole protein sequence of the first protein; performing multiple sequence alignment on the whole protein sequences of the other proteins with respect to the whole protein sequence of the first protein; identifying residues within each of the aligned whole protein sequences of the other proteins which positionally correspond with the residues of interest in the whole protein sequence of the first protein; determining a measure of similarity between the first protein and each respective other protein; and identifying one or more of the other proteins as off-target proteins with respect to the drug target based on the measures of similarity.