A method for rapid identification of exogenous insertion sequences based on whole genome sequencing data
By analyzing whole-genome sequencing data, we can quickly identify the exogenous insertion characteristics of transgenic or gene-edited materials, solving the problems of long detection cycles and low throughput in existing technologies, and achieving efficient and accurate identification and risk assessment of exogenous sequences.
Patent Information
- Application Number
- CN202211715491.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-12-29
AI Technical Summary
Existing technologies for identifying exogenous insertion sequences in transgenic or gene-edited materials suffer from problems such as long detection cycles, low throughput, time-consuming and labor-intensive processes, high costs, and limited results, making it impossible to accurately assess insertion sites, copy numbers, and flanking sequences.
Using whole-genome sequencing data, this method rapidly identifies the insertion site, direction, copy number, flanking sequences, and gene function of exogenous insertion sequences through quality control, preliminary alignment, homology detection, precise insertion site identification, and primer design, combined with bioinformatics analysis. The accuracy of the results is then evaluated.
It enables rapid and accurate identification of insertion features of exogenous sequences, eliminates the false positive effect caused by homology, provides intuitive visualization, reduces experimental time and cost, and improves the comprehensiveness and understandability of results, making it suitable for safety risk assessment of transgenic plants and animals.
Smart Images

Figure CN116189772B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of food safety, genetic engineering, biological breeding, biotechnology, biomedicine, and bioinformatics analysis, and in particular to a method for rapidly identifying transgenic or gene-edited materials based on whole-genome resequencing data, including transgenic insertion sites, insertion sequences, insertion copy numbers, flanking sequences, and other transgenic insertion characteristics. Background Technology
[0002] Genes are the smallest units of heredity, and every trait of an organism is determined by its own genes. Since the dawn of agricultural civilization, humans have continuously domesticated plants and animals to obtain cultivated or domesticated varieties with superior traits. Compared to their wild plant and animal ancestors, all cultivated crops (such as rice, wheat, and vegetables) or livestock (such as pigs, cattle, and chickens) have undergone extensive changes at the phenotypic and genetic levels. Using genetic engineering and transgenic technology, the genetic material of different organisms can be artificially processed and manipulated according to human desires (such as disease and pest resistance, stress resistance, high yield, and high quality). Specific genetic material is then combined with an appropriate vector and transferred into the target organism or cell, allowing for stable expression and the production of the desired characteristics. Furthermore, this exogenous genetic species can be stably inherited by the next generation.
[0003] The continued growth of the global population has created a tremendous demand for sustainable food production. Since the large-scale commercial cultivation of transgenic crops began in 1996, the global planting area of genetically modified crops has grown rapidly, and genetically modified agricultural products that have undergone safety assessments have become staple foods for people worldwide. my country has also established a complete transgenic technology system covering the entire chain, including gene cloning, genetic transformation, variety breeding, and safety evaluation. In addition, transgenic animals play an important role in promoting the development of biomedicine, such as the preparation of animal models of human diseases and xenotransplantation.
[0004] The insertion of foreign genes mostly exhibits random integration, and this randomness leads to unpredictable expression and unstable phenotypes. Therefore, the identification of transgenic components is a necessary condition for the biosafety assessment of transgenic plants and animals, including the insertion site of the foreign sequence, the characteristics of the inserted sequence, and the flanking sequences of the insertion site.
[0005] Traditional methods for detecting exogenous sequence fragments include:
[0006] (1) In situ hybridization (ISH): to obtain the integrity of the introduced foreign sequence, the approximate copy number, etc.
[0007] (2) Real-time quantitative PCR (RT-PCR): accurately obtains the copy number of the transgene insertion sequence;
[0008] (3) Chromosome walking technology: to obtain precise transgene insertion sites;
[0009] (4) Fluorescence in situ hybridization (FISH) technology: to obtain the location of transgene insertion and integration into the chromosome;
[0010] (5) Southern Blot hybridization technique: to obtain the copy number of the transgenic insertion sequence.
[0011] The above methods are all used to detect the insertion position, copy number and flanking sequences of foreign sequences, but these methods have long experimental cycles, low throughput, are time-consuming and labor-intensive, and have high experimental costs. At the same time, the content identified is relatively simple.
[0012] With the rapid development of high-throughput sequencing technology, species reference genome sequences are becoming increasingly refined and improved. Based on whole-genome resequencing data and the comparison differences between genome sequences, a new method has emerged for rapidly identifying transgenic or gene-edited materials and identifying the insertion characteristics of foreign sequences.
[0013] The invention patent with authorization announcement number CN105631242B discloses a method for identifying transgenic events using whole-genome sequencing data. It maps sequencing data to a reference genome and expression vector sequence, and uses a three-step method—preliminary positioning, coarse localization, and fine allelic identification—to determine the location, orientation, copy number, flanking sequences, and heterozygous state of the foreign gene insertion into the genome. However, this patent invention primarily focuses on the statistical results of mapping sequencing data to the genome and expression vector sequences, failing to consider false positives caused by homology and splicing alignment, as well as the impact of recombination disorders in sequences adjacent to the foreign insertion site on the method's sensitivity.
[0014] Patent CN103270175B discloses a method and system for detecting the insertion site of a transgenic exogenous fragment. The method includes: determining a transgenic unilateral short fragment and a genomic unilateral short fragment by comparison, and determining the insertion site of the exogenous fragment in the genomic sequence based on the intersection of the two fragments. However, this patent uses a unilateral fragment sequence to determine the transgenic gene insertion location, and cannot evaluate the flanking sequences of the insertion site, the actual inserted sequence, copy number, and other characteristics of the inserted sequence.
[0015] The invention patent with authorization announcement number WO2021047363A1 discloses a method for rapidly identifying transgenic or gene-edited materials and their insertion sites using whole-genome resequencing data. This method can determine whether a transgenic or gene-editing event has occurred, whether the expression vector is known or unknown, and accurately identify the insertion site, flanking sequences, copy number, and other information when the expression vector is known. However, this invention patent only maps the resequencing data to the expression vector sequence or sequences in a vector database, significantly reducing the accuracy of the mapping.
[0016] Furthermore, none of the aforementioned invention patents take into account the accuracy of known exogenous vector sequences, contamination of sequencing data other than exogenous transfer, the reliability of alignment results when the vector sequence and genome sequence are partially homologous, and whether or not the backbone sequence is inserted. Moreover, they cannot intuitively obtain the effects of exogenous sequence insertion on sequence deletions and gene protein functions near the insertion point. Summary of the Invention
[0017] The technical problem to be solved by this invention is to provide a method for rapid identification of exogenous inserted sequences based on whole-genome sequencing data. Its purpose is to detect the insertion of exogenous sequences more simply and quickly, and to provide more accurate and intuitive information on the precise insertion site, insertion direction, copy number, flanking sequences, and gene function in the insertion site and surrounding area of exogenous sequences (including target sequences and other sequences such as vector backbone sequences). At the same time, it will also check the consistency of exogenous sequences and the accuracy of resequencing data.
[0018] This invention relates to the following definitions of terms:
[0019] Exogenous sequences: Exogenous plasmids, vector sequences, etc., used in transgenic or gene-editing processes that are not from the species genome or endogenous organelles.
[0020] Inserted fragment: The sequence portion transferred into the material sample, including the entire exogenous sequence or a portion of the exogenous sequence.
[0021] Assembly sequence: The sequence of the insertion sequence and its flanking parts obtained through second-generation data and second-generation assembly software.
[0022] The present invention solves the above-mentioned technical problems by adopting the following technical solutions:
[0023] A method for rapid identification of exogenous inserted sequences based on whole-genome sequencing data includes the following steps:
[0024] (1) Sequencing data quality control steps: Quality control is performed on the whole genome second-generation paired-end high-throughput sequencing data of the transgenic or gene-edited material to be tested;
[0025] (2) Preliminary data comparison step: The quality-controlled sequencing data is compared and mapped with the wild-type reference genome and exogenous sequence of the species of the transgenic or gene-edited material sample to be tested, and the comparison output results are obtained; the exogenous sequence is the expression vector sequence carrying the exogenous gene and whose sequence is known;
[0026] (3) Preliminary comparison result screening steps: Check and confirm whether there are reads that reveal abnormalities in the provided exogenous sequences. The reads are characterized as paired reads, with one end aligned to the exogenous sequence and the other end without alignment results. Extract the read sequences of the paired reads with no alignment results at one end and compare them with the NT library. Statistically check the proportion of exogenous microorganisms in the comparison results and extract the relevant sequence information of the exogenous microorganisms with the highest content.
[0027] (4) Secondary alignment result evaluation steps: If there is an exogenous microbial sequence in step (3) above, use it and the wild-type reference genome of the species of the transgenic or gene-edited material sample to be tested and the exogenous sequence as reference sequences, perform alignment mapping on the second-generation sequencing reads, evaluate the alignment rate, coverage and average coverage depth information of the whole genome range and the overall range of the exogenous sequence, and draw the depth distribution map of the overall range of the exogenous sequence.
[0028] (5) Homology detection steps: The exogenous sequence is compared and mapped with the wild-type reference genome sequence of the species of the transgenic or gene-edited material sample to be tested, the homology between the exogenous sequence and the species genome sequence is evaluated, and the homology characteristics are determined for subsequent filtering of false positive results caused by homology.
[0029] (6) Preliminary insertion region detection steps: Based on the condition that one end of the paired-end sequencing reads is aligned with the foreign sequence and the other end is aligned with the species reference genome sequence, the insertion region is initially located;
[0030] (7) Precise insertion site identification steps: Accurately locate the insertion site region based on the comparison of the shear breakpoint;
[0031] (8) Insertion and flanking sequence assembly steps: The sequence genes of the obtained precise insertion sites and flanking regions are evaluated, and the consistency between the assembled sequence and the foreign sequence is assessed;
[0032] (9) The impact of insertion sites on gene acquisition steps: The impact of insertion sites on the original sequence gene at the sequence level is shown;
[0033] (10) Insertion site primer design steps: Using primer design software, design PCR amplification primers for the insertion site region.
[0034] As one of the preferred embodiments of the present invention, in step (1), quality control includes sequencing quality control and data contamination checks.
[0035] As one of the preferred embodiments of the present invention, in step (3), reads with one end aligned to the exogenous source and the other end without alignment results are extracted and aligned to the NT library to check for contamination by exogenous microorganisms, thereby assessing the accuracy of the exogenous sequence; if so, the sequence of the microorganism is added to the reference sequence and used together with the sequencing reads for alignment to improve the accuracy of the alignment results.
[0036] As one of the preferred embodiments of the present invention, in step (5), the homology features include: homology location region, region length, and degree of consistency.
[0037] As one of the preferred embodiments of the present invention, in step (7), when performing precise detection on the preliminary region obtained in step (6), the filtering conditions include: filtering for homologous sequence influence, obtaining high-quality alignment results through alignment quality values, and ensuring coverage accuracy through read sequencing depth.
[0038] As one of the preferred embodiments of the present invention, in step (7), a novel display method is adopted to provide an intuitive and user-friendly visualization of the alignment of reads near the insertion site.
[0039] As one of the preferred embodiments of the present invention, in step (9), the genes near the insertion site region and their effects on the genes are displayed to intuitively understand the functional impact of exogenous insertion on the original genes.
[0040] As one of the preferred embodiments of the present invention, a primer design step is added in step (10) to facilitate subsequent experimental verification and reduce experimental time.
[0041] Furthermore, the method for rapidly identifying exogenous inserted sequences based on whole-genome sequencing data includes the following specific steps:
[0042] (1) DNA extraction: High-quality genomic DNA of transgenic or gene-edited material samples to be tested is obtained by using conventional DNA extraction methods to separate and purify the DNA.
[0043] (2) Library construction and sequencing: The extracted genomic DNA was randomly fragmented to create a 300-500bp library. The whole genome paired-end sequencing data of the sample to be tested was obtained using the PE sequencing strategy and BGI Genomics or Illumina gene sequencer.
[0044] (3) Sequencing data quality control: The whole genome resequencing data obtained in step (2) was subjected to quality control using FASTP software to remove low-quality, PCR duplicate, and short reads, and retain high-quality clean reads; the data filtering method is as follows:
[0045] A. Filter out reads containing connector sequences;
[0046] B. Remove duplicated reads caused by PCR amplification, etc.
[0047] C. When the amount of N at one end of a single-end sequencing read exceeds 10% of the length of that read, the paired reads need to be removed.
[0048] D. When the number of low-quality (<=5) bases at one end of a single-end sequencing read exceeds 50% of the length of that read, the paired reads need to be removed.
[0049] E. The minimum length of the pruned reads is 50 bp.
[0050] (4) Sequencing data contamination detection: 20,000 reads were randomly extracted from the clean data obtained in step (3) above. The randomly extracted read sequences were aligned to the NCBI NT (Nucleotide Database) database using NCBI BLAST+blastn software. The alignment results were filtered according to max_target_seqs=1, E-value<=1e-5, Align length>=100bp, and Identity>=85% to check whether the sequencing data was contaminated.
[0051] (5) Preliminary data comparison and inspection:
[0052] 5.1) Use bwa software to map the clean datat obtained in step (3) to the wild-type genome sequence of the foreign sequence and the wild-type genome sequence of the sample to be tested to obtain the alignment result file; use samtools software to analyze the alignment result. If there are more than 5 paired reads, one of which is aligned to the foreign sequence and the other read is not aligned, it indicates that the sequencing data contains other unknown sequences related to the foreign sequence.
[0053] 5.2) Using NCBI BLAST+blastn software, the single-end read sequences obtained in step 5.1) above were aligned to the NCBI NT (Nucleotide Database). The alignment results were filtered according to max_target_seqs=1, E-value<=1e-5, Align length>=100bp, and Identity>=85%. The possible source information of the reads was statistically checked, and the source species sequences were added to the reference sequence in real time to re-align and map the original sequencing data using bwa, so as to improve the accuracy and precision of the alignment results.
[0054] (6) Statistical evaluation of alignment results: The SAMTools software was used to statistically analyze the alignment rate, coverage, and coverage depth of the whole genome and the overall exogenous sequence. At the same time, a depth distribution map of the exogenous sequence was drawn to visually show the region of exogenous sequence transfer (such as the entire expression vector exogenous sequence being transferred, or partially transferred, or only the target gene being transferred) and the sequencing coverage depth. This depth can indicate the copy number of the inserted sequence.
[0055] (7) Homology detection between exogenous sequences and species genome sequences: The exogenous sequences were compared and mapped with the genome sequences of the samples to be tested using NCBI BLAST+blastn software. The comparison results were filtered according to max_target_seqs=1, E-value<=1e-5, Align length>=100bp, and Identity>=85% to evaluate the homology between the exogenous sequences and the species genome sequences. Homology characteristics such as homology location region, region length, and degree of consistency were determined. The homology between the exogenous sequences and the species genome will affect the accuracy of the sequencing data comparison and mapping.
[0056] (5) Reads screening: Paired-end sequencing reads that fall into the following four categories in the alignment results:
[0057] a) First type: Part of the reads from either end aligns with the reference genome sequence, and the other part aligns with the foreign sequence;
[0058] b) Second type: One end of the reads is aligned with the reference genome sequence, and the other end of the reads is aligned with the foreign sequence;
[0059] c) Third type: Both ends are aligned with foreign sequences;
[0060] d) Category 4: Reads (type 1 / type 2) that are related to foreign sequences overlap and both ends are aligned to the reference genome sequence.
[0061] These four types of reads can be used to assemble insertion sequences (type 1 / type 2 / type 3), extend and assemble flanking sequences of the insertion site (type 1 / type 2 / type 4). Types 1 and 2 reads can be used to accurately locate the insertion site; see details below. Figure 1 exhibit.
[0062] (9) Precise Insertion Site: Reads that align to a foreign sequence at one end and a species genome sequence at the other end are extracted. These reads are filtered based on the number of reads they support covering greater than 5. Reads within homologous regions to the foreign sequence and low-confidence alignments (reads shorter than 50 bp due to clipping) are removed. The approximate insertion region is obtained by calculating the actual number of supported reads. The candidate regions and nearby read alignments obtained using the above methods are visualized. The insertion chromosome and precise insertion site are determined based on the clipping breakpoint location.
[0063] (10) Insertion and flanking sequence assembly: Extract the four types of paired reads described in step (8) within 500 bp before and after the insertion site obtained in step (9). Use SPAdes software to automatically determine the optimal kmer size according to k-mer sizes = auto, and perform sequence assembly of the insertion fragment and its flanking segments. The preliminary assembly results are filtered according to length = 100 bp to obtain the final flanking and insertion sequence assembly results.
[0064] (11) Assembly sequence evaluation:
[0065] 11.1) Align the second-generation whole-genome resequencing data to a reference sequence obtained by merging the assembled sequence with the species' whole-genome sequence, and calculate its coverage and sequencing depth. Using NCBI BLAST+blastn software, align the assembled sequence to the exogenous sequence. Filter the alignment results according to max_target_seqs=1, E-value<=1e-5, Align length>=100bp, and Identity>=85%, and plot circos diagrams to evaluate the consistency between the assembled and exogenous sequences, including consistency in sequencing depth distribution and sequence consistency.
[0066] 11.2) Mask out the regions in the assembled sequence that are identical to the exogenous sequence, and align the masked sequence to the species reference genome. Based on max_target_seqs=1, E-value<=1e-5, Align length>=100bp, and Identity>=85%, filter the alignment results to obtain the location of the flanking sequence.
[0067] (12) Gene information in the insertion site region: Based on the precise insertion site obtained in step (9), obtain all gene information within a 2k bp region upstream and downstream of the insertion site, including but not limited to gene location, gene name, gene function, and original information of the gene where the insertion site is located.
[0068] (13) Primer design: Based on the precise insertion site obtained in step (9), use Primer3 software to design the pre- and post-insertion primer sequences to assist in the subsequent verification by first-generation sequencing using PCR technology, and obtain the actual flanking and insertion sequences.
[0069] The advantages of this invention compared to the prior art are:
[0070] (1) This invention assesses the homology between exogenous vector sequences and species genome sequences, and eliminates the false positive effect caused by multiple alignment mapping due to homology;
[0071] (2) This invention takes into account the accuracy of resequencing data and provides indications of clearly identified sources of contamination;
[0072] (3) Considering the accuracy of known exogenous vector sequences, if the actual sequenced exogenous sequence is found to differ from the known exogenous vector sequence information, the specific abnormal sequence region and its possible source will be indicated.
[0073] (4) The reliability of the comparison results is evaluated in this invention. It is believed that the comparison mapping results with a length of less than 50bp due to double-end clipping are unreliable, and the existence of false positives caused by this may be eliminated.
[0074] (5) This invention takes into account the possibility that the backbone sequence of the exogenous expression vector, other than the target gene, may be inserted into the genome of a species, and at the same time evaluates the insertion events and insertion characteristics of the entire exogenous expression vector sequence, including the target exogenous gene.
[0075] (6) This invention brings a brand-new reading alignment visualization display, which provides an intuitive and user-friendly visualization display of information such as reading alignment, depth distribution, and shear breakpoints in the region near the insertion site, making it easier to understand;
[0076] (7) This invention provides a statistical display of information such as gene distance, location, and function within a 2kb region before and after the insertion site, which facilitates a direct prediction of the impact of the inserted fragment.
[0077] (8) This invention integrates primer design software, and while identifying the insertion site, it designs primers targeting the insertion site region, which facilitates subsequent PCR experiments to verify the insertion site.
[0078] In summary, this invention combines bioinformatics analysis methods to rapidly identify the insertion characteristics of exogenous sequences given known exogenous sequence information and the wild-type genome sequence information of the sample species. These characteristics include: the insertion site, insertion sequence information, insertion direction, copy number, upstream and downstream flanking sequences of the insertion site, genes involved in the upstream and downstream regions of the insertion site and their related information, and integrate primer design results for easy amplification and verification of the insertion site sequence. It also provides a novel visualization of read alignment mapping results, offering an intuitive and user-friendly display of read alignment distribution, depth distribution, and splicing breakpoints in the region near the insertion site. Compared with traditional experimental detection techniques, this invention is less time-consuming, reproducible, and presents more comprehensive transgenic insertion characteristics. Compared with other patented methods, it is more accurate and intuitive, with more user-friendly results, and is more conducive to the safety risk assessment of transgenic plants and animals, contributing to the development and application of transgenic or gene-editing technologies. Attached Figure Description
[0079] Figure 1 This invention illustrates the four types of reads involved in flanking and insert sequence assembly (in the figure, type 1: a portion of the reads from either end aligns to the reference genome sequence, and the other portion aligns to the insert sequence; type 2: one end of the reads aligns to the reference genome sequence, and the other end of the reads aligns to the insert sequence; type 3: both ends align to the insert sequence; type 4: reads related to the insert sequence (type 1 / type 2) overlap, and both ends align to the reference genome sequence).
[0080] Figure 2 This is a flowchart illustrating the overall steps of the method for rapidly identifying exogenous inserted sequences based on whole-genome sequencing data in Example 1.
[0081] Figure 3 The process steps involved in Example 1 include resequencing data quality control assessment, homology assessment, data alignment and insertion site detection, flanking and insertion sequence assembly, and assembled sequence consistency assessment.
[0082] Figure 4 This is a sequencing depth distribution map of the exogenous sequences in Example 1 (in the figure, the dashed line is the average sequencing depth value of the whole genome, and the bar chart is the sequencing depth distribution of the exogenous sequences).
[0083] Figure 5This is a display of the distribution of reads near the candidate insertion site in Example 1 (in the figure: blue and green blocks represent read1 and read2 aligned to the same sequence, respectively; orange and red blocks represent read1 and read2 aligned to different sequences, respectively; brown blocks represent PE aligned reads aligned to different sequences, with only one side of the read in the currently displayed region; black blocks represent clip alignments of the read itself; gray blocks represent paired reads, with one aligned to the current position and the other not aligned across the entire reference genome).
[0084] Figure 6 This is a display of the read mapping distribution in the region where the actual sequencing data of the exogenous sequence is inconsistent with the provided sequence in Example 1 (in the figure: blue and green blocks represent read1 and read2 aligned to the same sequence, respectively; orange and red blocks represent read1 and read2 aligned to different sequences, respectively; brown blocks represent PE aligned reads aligned to different sequences, with only one side of the read in the currently displayed region; black blocks represent clip alignment of the read itself; gray blocks represent paired reads, one of which is aligned to the current position, and the other is not aligned in the entire reference genome).
[0085] Figure 7 This is a display of the consistency results between the flanking and inserted sequences and the known exogenous sequences in Example 1 (the circle diagram shows, from the outside in, the sequence and length, coverage depth, collinearity, and other information). Detailed Implementation
[0086] The embodiments of the present invention are described in detail below. These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operation processes. However, the scope of protection of the present invention is not limited to the following embodiments.
[0087] Example 1
[0088] This embodiment describes a method for rapidly identifying exogenous insertion characteristics (precise insertion site, insertion direction, insertion copy number, and flanking sequences) in transgenic or gene-edited materials based on whole-genome resequencing data. See [link to relevant documentation]. Figure 2 , Figure 3 It includes the following steps:
[0089] (1) DNA extraction: Genomic DNA of high-quality transgenic Arabidopsis samples was isolated and purified using conventional DNA extraction methods.
[0090] (2) Library construction and sequencing: The extracted genomic DNA was randomly fragmented to create a 300-500bp library. The whole genome paired-end sequencing data of the Arabidopsis thaliana sample to be transgenic was obtained by using the PE sequencing strategy and the MGI DNBSEQ T7 high-throughput sequencer manufactured by BGI Genomics.
[0091] (3) Sequencing data quality control: Use the FASTP software to perform quality control on the raw sequencing data obtained in step (2) above, remove low-quality, PCR duplicates, short sequences and other reads, and retain high-quality clean reads.
[0092] (4) Sequencing data contamination detection: 20,000 reads were randomly extracted from the clean data obtained in step (3) above. Using NCBI BLAST+blastn software, the randomly extracted reads were aligned to the NCBI NT (Nucleotide Database). The filtering criteria were set to max_target_seqs=1, E-value=1e-5, and -max_target_seqs=1 to filter the alignment results. The sequencing data was checked for contamination. The statistical results are shown in Table 1 below. More than 95% of the reads were aligned to the sequences of the detected species Arabidopsis thaliana and its closely related species, indicating that the original sequencing data was not contaminated.
[0093] Table 1. Comparison of sequencing data from transgenic or gene-edited samples with NT library.
[0094] Species / Organelles Number of entries compared to this entry Total number of comparisons Comparison percentage (%) Arabidopsis 16,220 19,918 81.43 Arabidopsis chloroplast 2,404 19,918 12.07 Arabidopsis mitochondria 384 19,918 1.93 Turritis mitochondria 212 19,918 1.06 Raphanus 181 19,918 0.91 Capsella mitochondria 56 19,918 0.28 Raphanus mitochondria 48 19,918 0.24 Erigeron chloroplast 37 19,918 0.19 Ipomoea 31 19,918 0.16 Boechera mitochondria 30 19,918 0.15 Aesculus chloroplast 20 19,918 0.10
[0095] (5) Preliminary comparison and inspection of sequencing data:
[0096] 5.1) Use bwa software to align the high-quality whole genome resequencing data obtained in step (3) to the exogenous sequence and the Arabidopsis wild-type genome sequence to obtain the alignment result file; use samtools software to extract and analyze the alignment result. There are more than 5 paired reads, one of which is aligned to the exogenous sequence, and the other read has no alignment result, indicating that the sequencing data contains other unknown sequences related to the exogenous sequence.
[0097] 5.2) Using NCBI BLAST+blastn software, the single-end read sequences obtained in step 5.1) were aligned to the NCBI NT (Nucleotide Database). The alignment results were filtered according to max_target_seqs=1, E-value<=1e-5, Align length>=100bp, and Identity>=85%. Statistical checks revealed that some reads may have originated from bacteria with GenBank number CP049753.1. The genome sequence of this bacterium was added to the reference sequence and the sequencing data was re-aligned using bwa to improve the accuracy and precision of the alignment results. The alignment sequence information involved in this transgenic Arabidopsis material is shown in Table 2 below.
[0098] Table 1 shows the reference sequence information used for the alignment.
[0099] Sequence source Serial ID Number of items Total length (bp) N50(bp) GC content (%) Foreign seq 563 1 13,280 13,280 51.07 Bacteria CP049753.1 1 4,075,803 4,075,803 39.02 Genome - 7 133,788,422 26,150,454 36.35
[0100] (6) Statistical evaluation of alignment results: The SAMTools software was used to statistically analyze the alignment rate, coverage, and depth of the entire genome and the overall exogenous sequence. Simultaneously, a depth distribution map of the exogenous sequence was plotted to visually display the transfection region of the exogenous sequence and the coverage sequencing depth, i.e., the inserted copy number. Detailed results can be found in [link to detailed results]. Figure 4 .
[0101] (7) Homology detection between exogenous sequences and species genome sequences: The exogenous sequences were compared and mapped with the Arabidopsis wild-type genome sequence using NCBI BLAST+blastn software. The comparison results were filtered according to max_target_seqs=1, E-value<=1e-5, Align length>=100bp, and Identity>=85% to evaluate the homology between the exogenous sequences and the Arabidopsis wild-type genome sequence. Homology characteristics such as homology location region, region length, and degree of consistency were determined, as shown in Table 3 below. The exogenous sequences showed obvious homology with three regions on two chromosomes of the Arabidopsis genome. This homology would affect the comparison results of the actual sequencing reads within the three regions.
[0102] Table 2. Statistical results of homology comparison between exogenous sequences and reference genome.
[0103]
[0104]
[0105] (8) Use samtools software to extract reads that are aligned with foreign sequences at one end and with species genome sequences at the other end. Filter reads according to the number of reads supported by more than 5. At the same time, delete read alignment results in regions homologous to foreign sequences and low-confidence alignment results caused by clipping that result in read alignment length less than 50bp. The actual number of reads supported is obtained to obtain the approximate insertion region, as shown in Table 4 below. On chromosome number 4, the region from 2,953,894bp to 2,955,099bp was filtered out due to homology.
[0106] Table 4 shows the preliminary insertion site information obtained based on read alignment.
[0107]
[0108] (9) Visualize the comparison of reads between the candidate regions and nearby regions obtained in step (8) above, such as... Figure 5 The image shows the read coverage at candidate insertion site 1: 29,147,192bp-29,148,819bp. Figure 6 The diagram shows the reads mapping distribution in regions where the actual sequencing data of the exogenous sequence is inconsistent with the provided sequence. It shows that a sequence is missing from 8,291bp to 8,299bp in the exogenous sequence. The missing sequence may be consistent with the sequence in the region of 1,531,752bp to 1,533,100bp of the sequence numbered GenBank: CP049753.1. Based on the clipping breakpoint, the inserted chromosome and the precise insertion site are obtained. At the same time, all gene information within a 2kbp range upstream and downstream of the insertion site is obtained, including but not limited to gene location, gene name, gene function, and the original information of the gene where the insertion site is located, as shown in Table 5 below.
[0109] Table 5. Information on candidate insertion sites
[0110]
[0111] (10) Insertion and flanking sequence assembly: Extract the four types of paired reads within a 500bp range before and after the insertion site obtained in step (9) above. Use SPAdes software to automatically determine the optimal kmer size according to k-mer sizes=auto, and perform sequence assembly of the insertion fragment and its flanking segments. The preliminary assembly results are filtered according to length=100bp to obtain the final flanking and insertion sequence assembly results.
[0112] (11) Assembly sequence evaluation:
[0113] 11.1) The second-generation whole-genome resequencing data were aligned to a reference sequence obtained by merging the assembled sequence with the Arabidopsis wild-type whole-genome sequence, and its coverage and sequencing depth were statistically analyzed. The assembled sequence was aligned to the exogenous sequence using NCBI BLAST+blastn software. Alignment results were filtered according to max_target_seqs=1, E-value<=1e-5, Align length>=100bp, and Identity>=85%, and circos plots were generated to evaluate the consistency between the assembled and exogenous sequences, including consistency in sequencing depth distribution and sequence consistency. Specific results are shown in [link to results]. Figure 7 The diagram shows that Foreign represents the foreign sequence, assembly* represents the assembled sequence, and the diagram from the outside in represents the sequence, coverage depth, and collinearity.
[0114] 11.2) Mask out the regions in the assembled sequence that are identical to the exogenous sequence, and align the masked sequence to the species reference genome. Based on max_target_seqs=1, E-value<=1e-5, Align length>=100bp, and Identity>=85%, filter the alignment results to obtain the location of the flanking sequence.
[0115] (12) Primer design: Based on the precise insertion site obtained in step (9), use Primer3 software to design the pre- and post-insertion primer sequences to assist in the verification by subsequent PCR technology for first-generation sequencing, and obtain the actual flanking and insertion sequences.
[0116] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for rapid identification of exogenous inserted sequences based on whole-genome sequencing data, characterized in that, Includes the following steps: (1) Sequencing data quality control steps: The whole genome second-generation paired-end high-throughput sequencing data of the transgenic or gene-edited material to be tested are subjected to quality control; the quality control includes sequencing quality control and data contamination check; (2) Preliminary data comparison step: The quality-controlled sequencing data is compared and mapped with the wild-type reference genome and exogenous sequence of the species of the transgenic or gene-edited material sample to be tested, and the comparison output results are obtained; the exogenous sequence is the expression vector sequence carrying the exogenous gene and whose sequence is known; (3) Preliminary comparison result screening steps: Check and confirm whether there are reads that reveal abnormalities in the provided exogenous sequences. The reads are characterized as paired reads, with one end aligned to the exogenous sequence and the other end without alignment results; extract the read sequences of the paired reads with no alignment results at one end and compare them with the NT library, count the comparison results to check the proportion of exogenous microorganisms, and extract the relevant sequence information of the exogenous microorganisms with the highest content; (4) Secondary alignment result evaluation steps: If there is an exogenous microbial sequence in step (3) above, use it and the wild-type reference genome of the species of the transgenic or gene-edited material sample to be tested and the exogenous sequence as reference sequences, perform alignment mapping on the second-generation sequencing reads, evaluate the alignment rate, coverage and average coverage depth information of the whole genome range and the overall range of the exogenous sequence, and draw the depth distribution map of the overall range of the exogenous sequence; (5) Homology detection steps: The exogenous sequence is compared and mapped with the wild-type reference genome sequence of the species of the transgenic or gene-edited material sample to be tested, the homology between the exogenous sequence and the species genome sequence is evaluated, and the homology characteristics are determined for subsequent filtering of false positive results caused by homology. Among them, the alignment results showed the following four types of paired-end sequencing reads: Category 1: Part of the reads from either end aligns with the reference genome sequence, while the other part aligns with the foreign sequence; The second type: one end of the reads is aligned with the reference genome sequence, and the other end of the reads is aligned with the foreign sequence; Category 3: Both ends align with the exogenous sequence; Category 4: Reads that overlap with exogenous sequences, with both ends aligned to the reference genome sequence; (6) Preliminary insertion region detection steps: Based on the condition that one end of the paired-end sequencing reads is aligned with the foreign sequence and the other end is aligned with the species reference genome sequence, the insertion region is initially located; (7) Precise insertion site identification steps: The insertion site region is precisely located based on the alignment of the cutting breakpoint; When performing precise detection on the preliminary region obtained in step (6), the filtering conditions include: filtering for homologous sequence influence, obtaining high-quality alignment results through alignment quality value, and ensuring coverage accuracy through read sequencing depth; The specific operation includes: extracting a class of reads that "align to a foreign sequence at one end and to a species genome sequence at the other end", filtering according to the number of reads supported by the reads being greater than 5, and deleting the read alignment results in regions homologous to the foreign sequence and the low-confidence alignment results where the read alignment length is less than 50bp due to clip alignment, and obtaining the actual number of reads supported to obtain the approximate insertion region; Visualize the read alignment of the candidate region and the nearby region obtained by the above methods, and obtain the inserted chromosome and precise insertion site based on the location of the clip cutting breakpoint; (8) Insertion and flanking sequence assembly steps: Assemble the sequences of the precise insertion site and flanking region obtained from the evaluation, and evaluate the consistency between the assembled sequence and the exogenous sequence; (9) Steps for obtaining genes affected by insertion site: Based on the precise insertion site obtained in step (7), obtain all gene information within a 2kbp region upstream and downstream of the insertion site, and demonstrate the impact of the insertion site on the original sequence gene at the sequence level; (10) Insertion site primer design steps: Based on the precise insertion site obtained in step (7), use Primer3 software to design the pre- and post-insertion primer sequences to assist in subsequent first-generation sequencing verification using PCR technology, and obtain the actual flanking and insertion sequences.
2. The method for rapid identification of exogenous inserted sequences based on whole-genome sequencing data according to claim 1, characterized in that, In step (3), reads with one end aligned to the exogenous source and the other end without alignment are extracted and aligned to the NT library to check for contamination by exogenous microorganisms, thereby assessing the accuracy of the exogenous sequence; if so, the sequence of the microorganism is added to the reference sequence and used together with the sequencing reads for alignment to improve the accuracy of the alignment results.
3. The method for rapid identification of exogenous inserted sequences based on whole-genome sequencing data according to claim 1, characterized in that, In step (5), the homology features include: homology location region, region length, and degree of consistency.
4. The method for rapid identification of exogenous inserted sequences based on whole-genome sequencing data according to claim 1, characterized in that, In step (8), four types of paired reads within a 500bp range before and after the insertion site obtained in step (7) are extracted. The optimal kmer size is automatically determined using SPAdes software based on k-mer sizes=auto. The sequence assembly of the insertion fragment and its flanking segments is performed. The preliminary assembly results are filtered according to length=100bp to obtain the final flanking and insertion sequence assembly results.
Citation Information
Patent Citations
Method and system for detecting the insertion sites of transgenic foreign fragments
CN103270175B
A method for identifying transgenic events using whole-genome sequencing data
CN105631242B
Method for using whole genome re-sequencing data to quickly identify transgenic or gene editing material and insertion sites thereof
WO2021047363A1
Method for rapidly identifying transgenic or gene editing material and insertion site thereof by using whole genome resequencing data
CN110556165A
Method for identifying transgenic event based on high-throughput sequencing and probe enrichment
CN113957130A