Bicycle Gene Identification via Structural Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods struggle to identify highly divergent bicycle gene homologs due to their extreme sequence divergence, which limits functional inferences and evolutionary studies.
Innovation Solution
A sequence-independent method using a logistic regression classifier based on gene structure features, such as exon sizes and intron positions, to identify bicycle genes without relying on sequence similarity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If sequence-search methods are used to identify homologous genes, then identification process is simple and fast, but highly divergent bicycle gene homologs cannot be detected due to extreme sequence divergence
Solution Approach 1:
The patent replaces sequence-based identification methods with a gene structure-based classification system. Instead of using sequence similarity (the mechanical system of sequence searching), the invention uses structural features such as exon-intron architecture, gene length distributions, and synteny patterns to identify bicycle genes. This substitution allows detection of highly divergent homologs that sequence methods miss, as the structural features are conserved even when sequences diverge significantly.
Solution Approach 2:
The patent changes the identification parameters from sequence similarity metrics (e.g., BLAST E-values, identity percentages) to structural parameters (e.g., exon count, intron positions, gene length ranges). By transforming the detection space from sequence domain to structural domain, the method can identify bicycle genes across distant taxa where sequence divergence would traditionally prevent detection.
2Adaptability or versatility
If sequence divergence is high in bicycle genes, then functional adaptation and evolutionary innovation are enhanced, but homology detection becomes undetectable by conventional methods
Solution Approach 1:
The patent inverts the traditional homology detection approach. Instead of asking 'do these sequences look similar?' and concluding 'not homologous' when they don't, the method asks 'do these genes share conserved structural features?' and identifies homologs based on structural homology rather than sequence homology. This inversion allows detection of bicycle gene homologs that have diverged so much in sequence that conventional methods would reject them, while the structural conservation confirms their homological relationship.
3Measurement precision
If gene structure features are used for identification, then detection of divergent homologs is improved, but the identification process becomes more complex and time-consuming
Solution Approach 1:
The patent performs preliminary classification of genes into categories (e.g., unicycles, bicycles, tricycles, megacycles) based on structural features before detailed analysis. By pre-sorting genes into structural categories using rules such as exon count thresholds, intron position patterns, and gene length ranges, the method rapidly identifies potential bicycle gene candidates without needing to exhaustively search all possible sequence alignments. This preliminary structural triage significantly reduces the time required for comprehensive homology detection.
Data Source
AI summary
A method of identifying a bicycle gene involves determining for a candidate gene a series of gene structure-based predictor variables, and applying a bicycle gene classifier including the predictor variables to determine whether the candidate gene is identified as a bicycle gene. The gene structure-based predictor variables can be selected from the following: (i) total gene length (base pair, bp); (ii) total length (bp) of coding exons; (iii) first coding exon length (bp); (iv) last coding exon length (bp); (v) number of internal exons in phase 0); (vi) number of internal exons in phase 1; (vii) number of internal exons in phase 2; and (viii) mean internal exon length (bp).


