Protein Structure Prediction Using Sequence Inversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting the three-dimensional structure of proteins lack accuracy, particularly in predicting the tertiary structure, contact maps, and distance maps, often resulting in uneven error distribution near the ends of the protein sequence.

Innovation Solution

An information processing apparatus and method that acquires genome sequence information, inverts the sequence, and uses machine learning models to predict protein information, integrating predictions from both original and inverted sequences to enhance accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning is used to predict protein three-dimensional structure from sequence information, then prediction capability is achieved, but prediction accuracy is insufficient

Engineering Contradiction:
Improveprediction accuracyVSAvoiderror distribution uniformity
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies inversion by creating inverted sequence information where the amino acid sequence is reversed. The machine learning model processes both the original sequence and the inverted sequence, allowing the system to capture directional dependencies in the protein sequence that would be missed by single-direction processing. This dual-directional approach addresses the error distribution issues by providing complementary information from both sequence orientations.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent merges the predictions from both the original sequence information and the inverted sequence information by integrating their outputs through the machine learning model. This combination allows the system to leverage the strengths of both sequence orientations and produce a more accurate and reliable protein structure prediction that balances errors from both directions.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If sequence information is processed in a single direction, then processing is simple, but prediction accuracy is reduced

Engineering Contradiction:
Improveprediction accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements inversion by reversing the amino acid sequence to create inverted sequence information. This allows the machine learning model to process the sequence in both original and reversed directions, capturing bidirectional dependencies that enhance prediction accuracy while maintaining a relatively simple processing architecture.

Inventive Principle:
Principle #13The other way round (Inversion)

3Reliability

If only original sequence information is used, then input processing is simple, but error distribution is uneven

Engineering Contradiction:
Improveerror distribution uniformityVSAvoidsequence directionality information
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent addresses information loss by creating inverted sequence information that preserves the sequence structure while reversing the direction. This inversion recovers directional information that would be lost in single-direction processing, allowing the machine learning model to access both forward and backward sequence contexts to achieve more uniform error distribution.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS20240013863A1Information processing apparatus, information processing method, and program
Publication Date: 2024.01.11 SONY GROUP CORP
  • US20240013863A1 patent drawing
  • US20240013863A1 patent drawing
  • US20240013863A1 patent drawing

AI summary

An information processing apparatus according to an embodiment of the present technology includes: an acquisition unit; an inversion unit; and a generation unit. The acquisition unit acquires sequence information relating to a genome sequence. The inversion unit generates, on the basis of the sequence information, inversion information in which the sequence is inverted. The generation unit generates, on the basis of the inversion information, protein information relating to a protein. In this information processing apparatus, sequence information relating to a genome sequence is acquired by the acquisition unit. Further, inversion information in which the sequence is inverted is generated by the inversion unit on the basis of the sequence information. Further, protein information relating to a protein is generated by the generation unit on the basis of the inversion information. As a result, it is possible to predict information relating to a protein with high accuracy.