Automated Protein Structure Prediction from Cryo-EM Density Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for determining protein structures from cryo-electron microscopy data are inefficient and require extensive manual processing, as they can only determine fragments of protein complexes or require extensive manual steps, limiting the throughput and speed of vaccine and medicine development.
Innovation Solution
A fully automated software tool using deep convolutional neural networks to predict protein structures from cryo-EM density maps, which includes pre-processing voxel data, using neural networks to label voxels with atom types, backbone structure, secondary structures, and amino acid types, and aligning amino acid sequences to determine the complete molecular structure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual fitting of atoms is used to determine protein structure from cryo-EM density maps, then structural accuracy can be achieved, but the process requires enormous manual effort and is virtually impossible for larger structures with several thousand atoms
Solution Approach 1:
The patent replaces the manual mechanical process of fitting atoms with an automated computational system using deep learning neural networks. The system automatically predicts atomic structures from cryo-EM density maps without requiring manual intervention, thereby maintaining structural accuracy while eliminating the enormous manual effort previously required.
Solution Approach 2:
The system enables self-service by allowing the computational algorithm to automatically determine protein structures from cryo-EM data without human intervention. The neural network models independently perform the structure determination task, making the process autonomous and scalable to large structures.
2Extent of automation
If existing automated tools such as Rosetta, MAINMAST, and Phenix are used, then some automation is achieved, but they determine only fragments of protein complexes or require extensive manual processing steps
Solution Approach 1:
The patent creates a universal system that can handle complete protein complexes of various sizes and complexities in a single automated workflow. Unlike existing tools that are limited to fragments or require manual steps, this system universally processes entire protein complexes from cryo-EM density maps end-to-end, determining both backbone and side-chain structures automatically.
Solution Approach 2:
The system merges multiple previously separate steps (backbone determination, side-chain modeling, structure refinement) into a single integrated automated pipeline. By combining these functions into one unified system, it achieves complete automation and significantly increases throughput for determining structures of large protein complexes.
3Reliability
If existing tools are used to determine complete protein complex structures, then comprehensive structural information is obtained, but extensive manual processing steps are required which reduce efficiency
Solution Approach 1:
The system performs preliminary actions by pre-training neural network models on large datasets of protein structures before actual structure determination. This preliminary training enables the models to rapidly predict structures during inference, significantly reducing the computational time required while maintaining complete and accurate structure determination for protein complexes.
Data Source
AI summary
In some embodiments, a method of determining a molecular structure of a protein is provided. A computing system receives voxel data representing electron density obtained via cryo-electron microscopy. The computing system uses one or more neural networks to predict one or more likelihoods for each voxel. The computing system determines a backbone structure based on the predicted likelihoods. The computing system maps amino acid sequences to the backbone structure based on the predicted likelihoods. Mapping the amino acid sequences to the backbone structure based on the predicted likelihoods includes conducting an alignment technique that uses a reward function and a gap penalty. The computing system determines locations of carbon, nitrogen, and oxygen atoms within the backbone structure based on the predicted likelihoods. The computing system determines side-chain atoms based on the predicted backbone structure and the amino acid sequences to complete the molecular structure.


