Oligonucleotide Probe Design for Comprehensive Genome Coverage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genomic analysis methods, such as oligonucleotide arrays, are limited by the need for known probe sequences and often fail to represent the entire genome, leading to inefficient analysis, especially for cDNA arrays and genome-wide screening, as many designed oligonucleotides may not be present in the interrogated population.
Innovation Solution
Development of a method using a plurality of nucleic acid molecules that hybridize specifically to sequences in a genome, with specific length and sequence identity criteria, allowing for comprehensive analysis of complex genomes, including mammalian genomes, using microarray technology and algorithms like the Burrows-Wheeler Transform for efficient word counting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If oligonucleotide arrays are used for genomic analysis, then high resolution genetic changes can be detected, but the method requires prior knowledge of probe sequences and cannot represent the entire genome
Solution Approach 1:
The genome is divided into many short subsequences (oligonucleotides) of 15-30 nucleotides. These segmented pieces are distributed across the array, with each probe representing a specific genomic region. This segmentation allows comprehensive genome coverage while maintaining high resolution detection capabilities.
Solution Approach 2:
The method performs preliminary computational analysis to identify unique subsequences across the entire genome before array construction. Algorithms scan the genome sequence to select probe regions that will provide optimal coverage, ensuring that the array is pre-configured to represent the complete genome rather than relying on prior knowledge of specific genes of interest.
2Measurement precision
If cDNA arrays are used to interrogate gene expression, then specific genes can be analyzed, but only a limited set of genes is covered
Solution Approach 1:
The oligonucleotide array design provides universal coverage of the entire genome, making the array applicable to any gene or genomic region of interest. Unlike cDNA arrays limited to pre-selected genes, this universal platform can interrogate any gene expression or genomic variation across the complete genome, serving multiple analytical functions simultaneously.
3Adaptability or versatility
If many oligonucleotides are designed for array coverage, then genome-wide screening is attempted, but many oligonucleotides are unrepresented in the interrogated population resulting in inefficient analysis
Solution Approach 1:
The method creates a computational model (virtual representation) of the genome that identifies which subsequences will be present in the interrogated population. This virtual copy allows prediction of probe representation before actual hybridization, enabling optimization of the array design to maximize the proportion of probes that will have target sequences in the sample.
Solution Approach 2:
The algorithm adjusts key parameters including oligonucleotide length (15-30 nucleotides), complexity threshold (0.001%-70% of genome), and sequence identity requirements (≥90%) to optimize both genome coverage and analysis efficiency. By tuning these parameters, the method ensures that the majority of probes on the array will have corresponding targets in the interrogated population.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables accurate and efficient analysis of complex genomes by providing a high-resolution image of genetic changes, including copy number variation and methylation status, with improved hybridization kinetics and reduced cross-hybridization, facilitating the detection of polymorphisms and genomic rearrangements.
Implementation Method 1
each of the nucleic acid molecules hybridizes specifically to a sequence in a genome
Data Source
AI summary
The invention provides oligonucleotide probes that can be used to hybridize to a representation of nucleic acid sequences. Compositions containing the probes such as microarrays are also provided. The invention also provides methods of using these probes and compositions in therapeutic, diagnostic, and research applications. Systems and methods for using a word counting algorithm that can quickly and accurately count the number of times a particular string of characters (i.e., nucleotides) appears in a nucleotide sequence (e.g., a genome) are provided. This algorithm can be used to identify the oligonucleotide probes of the invention. The algorithm uses a transform of a genome and an auxiliary data structure to count the number of times a particular word occurs in the genome.


