Protein Domain Identification via 3D Structure Point Cloud Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting protein domains based on amino acid sequences are limited by low accuracy due to the high time complexity of multi-sequence alignment and the assumption of collinearity in homologous sequences, which is often violated, especially in proteins with low sequence identity.

Innovation Solution

A method and system that utilize a point cloud segmentation model based on dynamic graph convolutional neural networks to identify protein domains from three-dimensional structure images, integrating global and local structural features to improve accuracy and overcome sequence alignment errors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If amino acid sequence alignment methods are used to identify protein domains, then the identification can be performed based on sequence information, but the accuracy drops rapidly when amino acid sequence consistency is lower than a certain critical point

Engineering Contradiction:
Improveprotein domain identification accuracyVSAvoidamino acid sequence consistency
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent transitions from one-dimensional sequence alignment to three-dimensional structural analysis. Instead of comparing linear amino acid sequences, the method uses 3D coordinate information of atoms in protein domains to identify structural similarities. This dimensional transformation allows accurate identification of remote homologous proteins whose sequences have diverged but whose 3D structures remain conserved.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent replaces the mechanical sequence alignment process with a computational 3D structural comparison system. By using coordinate-based distance calculations and dynamic time warping algorithms on 3D atomic positions, the system achieves accurate domain identification without relying on sequence collinearity assumptions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If multi-sequence alignment calculation is performed to ensure accurate protein domain identification, then the identification quality improves, but the time complexity becomes very high

Engineering Contradiction:
Improveprotein domain identification accuracyVSAvoidcalculation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts only the essential 3D coordinate information of atoms from the complete protein sequence data. By focusing on the spatial positions of atoms in potential domain regions rather than performing comprehensive multi-sequence alignment, the method achieves accurate identification with significantly reduced computational time and complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates simplified 3D coordinate representations (copies) of protein domain structures that capture the essential structural features without containing the full complexity of the original sequences. These coordinate-based models enable rapid comparison and identification while preserving the structural information necessary for accurate domain detection.

Inventive Principle:
Principle #26Copying

3Ease of operation

If the assumption of collinearity in homologous protein sequences is made to simplify alignment, then the calculation process becomes easier, but this assumption is often violated in the real world leading to identification errors

Engineering Contradiction:
Improvesequence alignment processVSAvoididentification reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

Instead of assuming sequence collinearity and trying to align sequences linearly, the patent inverts the approach by directly comparing 3D structural coordinates without assuming any linear correspondence. This inversion eliminates the collinearity assumption problem by working in the structural domain where evolutionary conservation is more apparent.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent changes the fundamental parameters used for comparison from amino acid sequence identities to three-dimensional atomic coordinate distances. By transforming the data representation from sequential to spatial parameters, the method achieves reliable identification of homologous domains even when sequence collinearity is violated.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11908140B1Method and system for identifying protein domain based on protein three-dimensional structure image
Publication Date: 2024.02.20 ZHEJIANG LAB
  • US11908140B1 patent drawing
  • US11908140B1 patent drawing
  • US11908140B1 patent drawing

AI summary

Disclosed is a method and system for identifying a protein domain based on a protein three-dimensional structure image. According to the present application, the protein domain is identified based on a structure similarity, the identification errors and omissions of the protein domain caused by protein multi-sequence alignment errors when sequence consistency is not high can be effectively solved. According to the present application, the point cloud segmentation model based on the dynamic graph convolutional neural network is constructed, and by integrating global structural features and local structural features, segmentation of the protein domain and acquisition of semantic labels of the protein domain can be completed at the same time.