Whole Genome Re-sequencing for Transgenic Plant Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for identifying transgenic or gene-edited plants and insertion sites using whole genome re-sequencing data are limited by long experimental times, low throughput, and high false negative and false positive rates, particularly due to the lack of consideration for sequence recombination and homology issues.

Innovation Solution

A method utilizing bioinformatics analysis to determine transgenic or gene editing events by aligning paired-end sequencing data with known or unknown expression vector sequences, accounting for backbone sequences and sequence rearrangements, and using specific criteria to identify accurate insertion sites and copy numbers through whole genome re-sequencing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional methods (in situ hybridization, real-time fluorescence quantitative PCR, chromosome walking technology, fluorescence in situ hybridization) are used to detect transgenic or exogenous gene fragments, then identification accuracy can be achieved, but the experimental period becomes long and throughput remains low

Engineering Contradiction:
Improveidentification accuracyVSAvoidexperimental period
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces traditional mechanical and chemical laboratory methods (in situ hybridization, chromosome walking, fluorescence quantitative PCR) with a bioinformatics-based computational approach. By using whole genome re-sequencing data and aligning it with vector sequences through automated algorithms, the method eliminates time-consuming wet lab procedures while maintaining identification accuracy through rigorous sequence matching and statistical validation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a digital copy of the genome sequence through whole genome re-sequencing, then performs in silico analysis by aligning this digital copy with known vector sequences. This copying approach allows multiple analyses to be performed simultaneously on the same data without additional experimental time, thereby reducing the overall experimental period while maintaining comprehensive identification capability.

Inventive Principle:
Principle #26Copying

2Measurement precision

If traditional detection methods are used, then specific insertion sites can be identified, but the throughput remains low and multiple separate experiments are required

Engineering Contradiction:
Improveinsertion site identificationVSAvoidthroughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent merges multiple detection functions into a single integrated bioinformatics pipeline. By combining whole genome re-sequencing with comprehensive alignment algorithms that simultaneously identify insertion sites, copy numbers, and vector boundaries, the method achieves high throughput without sacrificing identification precision. The unified approach processes all detection tasks in parallel through computational analysis rather than sequential experimental steps.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal detection system that can identify various types of transgenic events (insertions, deletions, modifications) and provide comprehensive information (insertion sites, copy numbers, vector sequences) through a single method. This multi-functional approach increases throughput by eliminating the need for multiple specialized experiments while maintaining precise identification across different detection targets.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If existing whole genome re-sequencing methods are used that map reads to wild-type genome sequence, then comprehensive genome analysis can be performed, but the process becomes tedious, time-consuming and laborious

Engineering Contradiction:
Improvegenome analysis comprehensivenessVSAvoidoperational simplicity
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent extracts and separates the specific task of transgenic event detection from comprehensive whole genome analysis. By directly aligning re-sequencing reads with vector sequences rather than requiring mapping to the entire wild-type genome first, the method isolates the critical detection step, making it simpler and faster to operate while maintaining comprehensive analysis capability through targeted sequence comparison.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary preparation by having the vector sequence ready for direct comparison before actual detection. This preliminary action allows the detection process to begin immediately with direct alignment of re-sequencing data against the known vector sequence, eliminating the time-consuming step of mapping to wild-type genome and subsequently identifying deviations, thereby simplifying operations without reducing comprehensiveness.

Inventive Principle:
Principle #10Preliminary action

4Device complexity

If existing methods do not consider recombination or rearrangement of sequences around insertion sites, then analysis simplicity is maintained, but sensitivity is reduced and false negative rate increases

Engineering Contradiction:
Improveanalysis simplicityVSAvoiddetection sensitivity
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces dynamic analysis capabilities that adapt to different genomic contexts around insertion sites. The alignment algorithm dynamically adjusts to detect recombination events, rearrangements, and atypical insertion patterns by comparing reads against both wild-type and vector sequences simultaneously. This dynamic approach maintains analytical simplicity through automated adaptive processing while significantly improving detection sensitivity and reducing false negatives.

Inventive Principle:
Principle #15Dynamics

5Reliability

If vector sequence homology with host genome is not considered, then false positives are reduced, but when homology exists, false positive rate increases

Engineering Contradiction:
Improvefalse positive controlVSAvoidhandling of homologous sequences
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements feedback mechanisms where the alignment algorithm continuously evaluates match quality and contextual information. When homology between vector and host genome is detected, the system provides feedback to adjust alignment parameters, require additional validation criteria, or increase confidence thresholds. This feedback loop maintains reliability by controlling false positives while preserving versatility to correctly identify true transgenic events even in the presence of sequence homology.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20220205034A1Method for quickly identifying clean transgenic or gene-edited plants and insertion sites by using whole genome re-sequencing data
Publication Date: 2022.06.30 ZHEJIANG UNIV
  • US20220205034A1 patent drawing
  • US20220205034A1 patent drawing
  • US20220205034A1 patent drawing

AI summary

Provided is a method for quickly identifying clean transgenic or gene-edited plants and insertion sites by using whole genome re-sequencing data, comprising: extracting genomic DNA; obtaining paired-end sequencing data of a plant whole genome; determining whether an expression vector sequence containing T-DNA sequence, inserted into a plant is known; determining whether transgenic events or gene editing events exist in a plant to be tested, and whether a backbone sequence transfer event occurs; determining the insertion site of the T-DNA sequence. The method combines means of bioinformatic analysis to identify whether there is a transgenic or gene editing event when the expression vector is known or unknown; when the expression vector is known, it not only quickly and accurately provides accurate positioning, direction, copy number and flanking sequence information of the target sequence inserted into genome, but also allows to determine whether there is a backbone sequence inserted into the genome.