CNN-Based Transcription Factor Activation Domain Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods lack effective means to experimentally detect or computationally predict the full diversity of wild-type activation domains in transcription factors, which are crucial for gene regulation and are poorly conserved in sequence but highly conserved in function.

Innovation Solution

A method utilizing a convolutional neural network (CNN) trained with functional activation domain data from one organism to identify activation domains in another organism, incorporating peptide sequence, secondary structure, disorder, and activity, and involving a library of nucleic acid molecules with DNA-binding and potential activation domains to screen for functional activation domains in cells.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If direct screening of wild-type sequences is performed, then activation domains can be identified experimentally, but only modest numbers of ADs are identified at coarse resolution and the full diversity is not captured

Engineering Contradiction:
Improveidentification accuracy of activation domainsVSAvoidnumber of activation domains identified
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the protein sequence into multiple windows or regions, each analyzed independently for activation domain characteristics. This allows comprehensive scanning of the entire sequence to identify multiple ADs at fine resolution, overcoming the limitation of coarse resolution in direct screening methods.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a computational prediction dimension by training machine learning models on known activation domain data. This adds a new dimension to the identification process, enabling fine-resolution prediction across the entire sequence simultaneously, thereby identifying the full diversity of ADs without the productivity limitations of experimental screening.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If computational prediction methods are used, then prediction speed is improved, but accuracy is insufficient due to high diversity and poor conservation of AD sequences

Engineering Contradiction:
Improveprediction speedVSAvoidprediction accuracy of activation domains
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the parameters used for prediction by incorporating multiple features including amino acid composition, sequence motifs, structural properties, and evolutionary conservation. This multi-parameter approach improves prediction accuracy despite the high diversity and poor conservation of AD sequences, while maintaining computational speed.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a composite prediction model that integrates multiple types of information (sequence data, structural data, evolutionary data) similar to how composite materials combine different properties. This composite approach enables accurate prediction of diverse activation domains by leveraging complementary information from multiple sources.

Inventive Principle:
Principle #40Composite materials

3Measurement precision

If fine-resolution identification is achieved, then the full diversity of activation domains can be captured, but the complexity of the method increases

Engineering Contradiction:
Improveresolution of activation domain identificationVSAvoidcomplexity of identification method
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses computational models that copy or simulate the characteristics of known activation domains to predict new ones. By training on experimentally verified ADs and applying the learned patterns to predict fine-resolution boundaries in new sequences, the method achieves high resolution without the experimental complexity of direct screening.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent develops a universal computational framework that can identify activation domains across different organisms and protein types. This multi-functional approach achieves fine-resolution identification throughout the sequence simultaneously, avoiding the need for multiple separate experiments or complex procedures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11732381B2Systems and methods to identify transcription factor activation domains and uses thereof
Publication Date: 2023.08.22 THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
  • US11732381B2 patent drawing
  • US11732381B2 patent drawing
  • US11732381B2 patent drawing

AI summary

Embodiments herein describe systems and methods to identify transcription factor activation domains and uses thereof. Many embodiments obtain activation measurements of tiles or segments of known transcription factors in an organism. Further embodiments train a machine learning model, such as a convolutional neural network, to identify transcription factors and activation domains in other organisms of the same or different species. Such methods and systems can be used for industrial, medical, and research purposes.