DNA Sequence Alignment via Hydrogen Bond Pattern Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods focus on specific nucleobases for protein-DNA binding, potentially overlooking essential hydrogen bonds, and do not address how information is shared between multiple DNA sequences that bind the same protein.

Innovation Solution

A computer-implemented method that converts nucleotide sequences into arrays of hydrogen bond donors, acceptors, and methyl groups, aligns these patterns to identify shared information among multiple DNA sequences, and verifies the alignment using crystal or NMR structures to determine the consensus pattern used by proteins for binding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If methods focus on specific nucleobases for protein-DNA binding analysis, then the analysis is simpler and more direct, but essential hydrogen bonds may be overlooked and information sharing between multiple DNA sequences cannot be identified

Engineering Contradiction:
Improvesimplicity of analysisVSAvoidmissed hydrogen bonds and shared information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent segments the DNA sequence analysis into distinct functional components: hydrogen bond donors, hydrogen bond acceptors, and methyl groups. Each nucleotide is decomposed into these interaction elements, allowing systematic analysis of protein-DNA binding without overlooking essential interactions. This segmentation enables the method to capture both individual hydrogen bonds and shared information patterns across multiple binding sites.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from analyzing nucleobases in one dimension (sequence level) to analyzing hydrogen bonding interactions in a second dimension (interaction level). By mapping nucleotides to their hydrogen bond donor/acceptor/methyl group characteristics, the method adds an interaction dimension that reveals shared information between different DNA sequences binding to the same protein, overcoming the limitation of base-pair-only analysis.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of manufacture

If traditional base pair analysis is used, then the methodology is well-established and easier to implement, but it cannot identify how information is shared between multiple DNA sequences binding to the same protein

Engineering Contradiction:
Improvemethodology implementationVSAvoidshared information between sequences
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The patent creates a universal hydrogen bond pattern representation that can be applied to analyze multiple different DNA sequences binding to the same protein. The hydrogen bond donor/acceptor/methyl group framework serves multiple functions: it identifies individual binding interactions, reveals shared information across binding sites, and enables consensus sequence determination. This multi-functional approach overcomes the limitation of traditional base pair analysis while remaining implementable through systematic pattern matching.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses crystal structures and NMR structures as reference copies to verify the hydrogen bond patterns identified through sequence analysis. By comparing computational predictions against experimentally determined structures, the method validates the shared information identification and ensures accuracy. This copying and verification approach maintains methodological rigor while enabling new insights into information sharing between sequences.

Inventive Principle:
Principle #26Copying

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach allows for the identification of essential hydrogen bonds and shared information between DNA sequences, enhancing the understanding of protein-DNA interactions and enabling the design of novel DNA-binding proteins with increased specificity and applicability in gene therapies.

Implementation Method 1

The present disclosure is based on the development of an algorithm that converts a nucleotide sequence into an array of hydrogen bond donors and acceptors and methyl groups

Methodology Applied
Scientific EffectHydrogen bonding: Chemical Bonding

Implementation Method 2

obtaining the crystal or NMR structures for the various protein-DNA complexes from the publicly available Protein Data Bank and verifying through the crystal and NMR structures that the maintained bonds in the alignment are indeed used by the protein for binding and recognition

Methodology Applied
Scientific EffectX-ray crystallography: X-Ray

Implementation Method 3

obtaining the crystal or NMR structures for the various protein-DNA complexes from the publicly available Protein Data Bank and verifying through the crystal and NMR structures that the maintained bonds in the alignment are indeed used by the protein for binding and recognition

Methodology Applied
Scientific EffectNMR spectroscopy: Electron Paramagnetic Resonance

Data Source

PatentUS20240177800A1Alignment and comparison of genetic information
Publication Date: 2024.05.30 GEORGE MASON UNIVERSITY
  • US20240177800A1 patent drawing
  • US20240177800A1 patent drawing
  • US20240177800A1 patent drawing

AI summary

The present disclosure generally relates to methods and systems for identifying shared information between different DNA sequences, that bind the same protein, based on alignment of major groove hydrogen bonding between the different sequences. Such methods may be useful for designing novel DNA binding proteins, as well as identification of novel DNA protein binding consensus sequences, for use in gene therapies, treatment of diseases or disorders resulting from aberrant gene expression and/or cell proliferation, as well as pathogenic infections.