Protein Structure Prediction Using Sequence Inversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for predicting the three-dimensional structure of proteins lack accuracy, particularly in predicting the tertiary structure, contact maps, and distance maps, often resulting in uneven error distribution near the ends of the protein sequence.
Innovation Solution
An information processing apparatus and method that acquires genome sequence information, inverts the sequence, and uses machine learning models to predict protein information, integrating predictions from both original and inverted sequences to enhance accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine learning is used to predict protein three-dimensional structure from sequence information, then prediction capability is achieved, but prediction accuracy is insufficient
Solution Approach 1:
The patent applies inversion by creating inverted sequence information where the amino acid sequence is reversed. The machine learning model processes both the original sequence and the inverted sequence, allowing the system to capture directional dependencies in the protein sequence that would be missed by single-direction processing. This dual-directional approach addresses the error distribution issues by providing complementary information from both sequence orientations.
Solution Approach 2:
The patent merges the predictions from both the original sequence information and the inverted sequence information by integrating their outputs through the machine learning model. This combination allows the system to leverage the strengths of both sequence orientations and produce a more accurate and reliable protein structure prediction that balances errors from both directions.
2Measurement precision
If sequence information is processed in a single direction, then processing is simple, but prediction accuracy is reduced
Solution Approach 1:
The patent implements inversion by reversing the amino acid sequence to create inverted sequence information. This allows the machine learning model to process the sequence in both original and reversed directions, capturing bidirectional dependencies that enhance prediction accuracy while maintaining a relatively simple processing architecture.
3Reliability
If only original sequence information is used, then input processing is simple, but error distribution is uneven
Solution Approach 1:
The patent addresses information loss by creating inverted sequence information that preserves the sequence structure while reversing the direction. This inversion recovers directional information that would be lost in single-direction processing, allowing the machine learning model to access both forward and backward sequence contexts to achieve more uniform error distribution.
Data Source
AI summary
An information processing apparatus according to an embodiment of the present technology includes: an acquisition unit; an inversion unit; and a generation unit. The acquisition unit acquires sequence information relating to a genome sequence. The inversion unit generates, on the basis of the sequence information, inversion information in which the sequence is inverted. The generation unit generates, on the basis of the inversion information, protein information relating to a protein. In this information processing apparatus, sequence information relating to a genome sequence is acquired by the acquisition unit. Further, inversion information in which the sequence is inverted is generated by the inversion unit on the basis of the sequence information. Further, protein information relating to a protein is generated by the generation unit on the basis of the inversion information. As a result, it is possible to predict information relating to a protein with high accuracy.


