Alignment-Free K-mer Mutation Screening in Crop Genomics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for high-throughput screening of mutations in crop plants are labor-intensive, prone to false positives, and limited by the ability to measure visual traits, especially when using forward genetic screens, and face challenges in distinguishing sequencing errors from actual mutations in large populations.
Innovation Solution
A method utilizing alignment-free sequence analysis involving k-mer analysis to identify mutations in target sequences by pooling genomic DNA, amplifying, sequencing, and decomposing reads into k-mers to determine wild-type and mutant counts, allowing for the identification of statistically significant mutations and their carriers in a population.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If alignment-based sequence analysis is used to identify mutations in large populations, then measurement precision may be improved, but device complexity and loss of time increase due to labor-intensive processes
Solution Approach 1:
The patent extracts the alignment step from the sequence analysis process, using k-mer frequency analysis instead. By removing the computationally intensive alignment requirement, the method achieves rapid mutation detection in large populations without sacrificing accuracy, directly resolving the time-loss contradiction
Solution Approach 2:
The patent replaces the mechanical alignment process with a computational k-mer frequency approach. This substitution eliminates the labor-intensive nature of traditional alignment-based methods while maintaining mutation detection precision, addressing both time loss and complexity issues
2Reliability
If forward genetic screens are used to identify mutations, then reliability of trait identification may be improved, but device complexity and loss of time worsen due to inability to measure visual traits efficiently
Solution Approach 1:
The patent inverts the traditional forward genetic screen approach by using sequence-based identification without relying on visual trait measurement. This inversion maintains reliability in identifying true mutations while eliminating the complexity associated with phenotypic screening and visual assessment
3Productivity
If traditional mutation screening methods are used in large populations, then productivity may be improved through high-throughput approaches, but measurement precision worsens due to false positives
Solution Approach 1:
The patent introduces k-mer frequency analysis as an intermediary between sequencing and mutation identification. This intermediary step provides statistical validation by comparing observed k-mer frequencies against expected frequencies, enabling high-throughput screening while filtering out false positives through statistical significance testing
4Loss of time
If alignment-free sequence analysis using k-mer analysis is implemented, then loss of time and device complexity are reduced, but measurement precision may worsen due to difficulty in distinguishing sequencing errors from actual mutations
Solution Approach 1:
The patent implements feedback through statistical comparison of observed versus expected k-mer frequencies. By continuously comparing actual sequencing data against the expected distribution based on reference sequences, the method provides self-validation that distinguishes true mutations from sequencing errors, maintaining precision while achieving rapid analysis
Data Source
AI summary
The present invention provides methods for isolation of a member of a population which has one or more mutation(s) in one or more target sequence(s) in a population. The method may comprise the steps of: (a) pooling genomic DNA isolated from each member of the population in one or more dimensions; (b) amplifying the one or more target sequence(s) in the pooled genomic DNA, wherein optionally the amplification products are pooled; (c) sequencing the amplified products or obtaining the sequence reads for the amplified products, wherein, optionally, sequencing is by pair-end sequencing and further comprises merging the paired-end reads into composite read(s); (d) identifying the mutation(s) based on alignment-free sequence analysis of sequencing data, optionally by k-mer analysis and (e) identifying individual member(s) of the population comprising the one or more identified mutations in the target sequences, optionally by high-resolution DNA melting (HRM).

