Whole Genome Re-sequencing for Transgenic Plant Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for identifying transgenic or gene-edited plants and insertion sites using whole genome re-sequencing data are limited by long experimental times, low throughput, and high false negative and false positive rates, particularly due to the lack of consideration for sequence recombination and homology issues.
Innovation Solution
A method utilizing bioinformatics analysis to determine transgenic or gene editing events by aligning paired-end sequencing data with known or unknown expression vector sequences, accounting for backbone sequences and sequence rearrangements, and using specific criteria to identify accurate insertion sites and copy numbers through whole genome re-sequencing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional methods (in situ hybridization, real-time fluorescence quantitative PCR, chromosome walking technology, fluorescence in situ hybridization) are used to detect transgenic or exogenous gene fragments, then identification accuracy can be achieved, but the experimental period becomes long and throughput remains low
Solution Approach 1:
The patent replaces traditional mechanical and chemical laboratory methods (in situ hybridization, chromosome walking, fluorescence quantitative PCR) with a bioinformatics-based computational approach. By using whole genome re-sequencing data and aligning it with vector sequences through automated algorithms, the method eliminates time-consuming wet lab procedures while maintaining identification accuracy through rigorous sequence matching and statistical validation.
Solution Approach 2:
The patent creates a digital copy of the genome sequence through whole genome re-sequencing, then performs in silico analysis by aligning this digital copy with known vector sequences. This copying approach allows multiple analyses to be performed simultaneously on the same data without additional experimental time, thereby reducing the overall experimental period while maintaining comprehensive identification capability.
2Measurement precision
If traditional detection methods are used, then specific insertion sites can be identified, but the throughput remains low and multiple separate experiments are required
Solution Approach 1:
The patent merges multiple detection functions into a single integrated bioinformatics pipeline. By combining whole genome re-sequencing with comprehensive alignment algorithms that simultaneously identify insertion sites, copy numbers, and vector boundaries, the method achieves high throughput without sacrificing identification precision. The unified approach processes all detection tasks in parallel through computational analysis rather than sequential experimental steps.
Solution Approach 2:
The patent creates a universal detection system that can identify various types of transgenic events (insertions, deletions, modifications) and provide comprehensive information (insertion sites, copy numbers, vector sequences) through a single method. This multi-functional approach increases throughput by eliminating the need for multiple specialized experiments while maintaining precise identification across different detection targets.
3Quantity of substance
If existing whole genome re-sequencing methods are used that map reads to wild-type genome sequence, then comprehensive genome analysis can be performed, but the process becomes tedious, time-consuming and laborious
Solution Approach 1:
The patent extracts and separates the specific task of transgenic event detection from comprehensive whole genome analysis. By directly aligning re-sequencing reads with vector sequences rather than requiring mapping to the entire wild-type genome first, the method isolates the critical detection step, making it simpler and faster to operate while maintaining comprehensive analysis capability through targeted sequence comparison.
Solution Approach 2:
The patent performs preliminary preparation by having the vector sequence ready for direct comparison before actual detection. This preliminary action allows the detection process to begin immediately with direct alignment of re-sequencing data against the known vector sequence, eliminating the time-consuming step of mapping to wild-type genome and subsequently identifying deviations, thereby simplifying operations without reducing comprehensiveness.
4Device complexity
If existing methods do not consider recombination or rearrangement of sequences around insertion sites, then analysis simplicity is maintained, but sensitivity is reduced and false negative rate increases
Solution Approach 1:
The patent introduces dynamic analysis capabilities that adapt to different genomic contexts around insertion sites. The alignment algorithm dynamically adjusts to detect recombination events, rearrangements, and atypical insertion patterns by comparing reads against both wild-type and vector sequences simultaneously. This dynamic approach maintains analytical simplicity through automated adaptive processing while significantly improving detection sensitivity and reducing false negatives.
5Reliability
If vector sequence homology with host genome is not considered, then false positives are reduced, but when homology exists, false positive rate increases
Solution Approach 1:
The patent implements feedback mechanisms where the alignment algorithm continuously evaluates match quality and contextual information. When homology between vector and host genome is detected, the system provides feedback to adjust alignment parameters, require additional validation criteria, or increase confidence thresholds. This feedback loop maintains reliability by controlling false positives while preserving versatility to correctly identify true transgenic events even in the presence of sequence homology.
Data Source
AI summary
Provided is a method for quickly identifying clean transgenic or gene-edited plants and insertion sites by using whole genome re-sequencing data, comprising: extracting genomic DNA; obtaining paired-end sequencing data of a plant whole genome; determining whether an expression vector sequence containing T-DNA sequence, inserted into a plant is known; determining whether transgenic events or gene editing events exist in a plant to be tested, and whether a backbone sequence transfer event occurs; determining the insertion site of the T-DNA sequence. The method combines means of bioinformatic analysis to identify whether there is a transgenic or gene editing event when the expression vector is known or unknown; when the expression vector is known, it not only quickly and accurately provides accurate positioning, direction, copy number and flanking sequence information of the target sequence inserted into genome, but also allows to determine whether there is a backbone sequence inserted into the genome.


