Data analysis method for identifying transgenic event by using targeted sequencing

By combining targeted sequencing with bioinformatics analysis, designing targeted sequencing targets and constructing reference genomes, the problem that traditional PCR detection technology is difficult to simultaneously detect multi-gene and multi-variety genetically modified crops has been solved, and high-precision and high-flexibility genetically modified event detection has been achieved.

CN120656529APending Publication Date: 2025-09-16CHANGSHA BAIAOYUN DATA TECH CO LTD +2
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510743276.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Traditional PCR detection technology is difficult to simultaneously and efficiently detect multiple genes and varieties of genetically modified crops, and has low flexibility. Existing detection methods lack accuracy and sensitivity for genetically modified events.

Method used

Design targeted sequencing targets, combine targeted sequencing with bioinformatics analysis, and use specific probes and primer sets to detect multiple transgenic events with high precision and specificity. Construct a reference genome containing the inserts of all transgenic events, perform sequence alignment and filtering quality control, and determine the type and transformation status of the transgenic event.

Benefits of technology

It achieves efficient and accurate detection of multiple transgenic events, does not rely on specific detection kits, is suitable for the identification of transgenic events of any species, and improves the sensitivity and flexibility of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656529A_ABST
    Figure CN120656529A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-transgenic event data analysis and identification method based on targeted sequencing, which comprises the following steps: designing targeted sequencing targets according to different transgenic events and crop varieties, and identifying the transgenic events by bioinformatics analysis of sequencing results. And high-precision and high-specificity detection can be carried out on various transgenic events and various important elements contained in the transgenic events at the same time. According to the invention, a specific detection kit is not needed, and corresponding probes and primer groups / or primer groups are designed according to transgenic events needing to be identified and reference genomes of corresponding species; by using the analysis process introduced by the invention, the sequencing data can be used for identifying and detecting a plurality of transgenic events on the detection material at the same time. The reference genome, the probe and the primer group / or the primer group and the identification method designed according to the invention can be used for identifying any transgenic event of any species.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gene sequencing, and in particular relates to a data analysis method for identifying transgenic events using targeted sequencing. Background Art

[0002] The diversity and complexity of genetically modified crops pose challenges to the detection of genetically modified products. Ensuring rapid and accurate genetic testing is a core challenge that the current genetically modified testing industry must address. Currently, genetically modified crop detection technologies fall into two main categories: nucleic acid-based detection techniques, which detect the presence of exogenous gene insertions; and protein-based enzymatic and immunological detection techniques, which detect protein products expressed by exogenous genes. Because nucleic acids, especially DNA, are highly stable and resistant to degradation during processing, nucleic acid detection technology is the most widely used in the detection of genetically modified ingredients. With the increasing number of genetically modified crops and products, traditional PCR detection technology can no longer meet the requirements for simultaneous detection of multiple genes and varieties. Multiplex PCR-based detection techniques capture less sequence information, and the accuracy and sensitivity of identifying genetically modified events rely on primer design and optimized experimental conditions. Furthermore, they rely on established detection kits, resulting in low flexibility. Summary of the Invention

[0003] To address the requirement of simultaneous detection of multiple genes and multiple varieties, the present invention proposes a method for data analysis and identification of multiple transgenic events based on targeted sequencing. Targeted sequencing targets are designed according to different transgenic events and crop varieties. Through bioinformatics analysis of the sequencing results, multiple transgenic events and the various important elements they contain can be detected simultaneously with high precision and specificity.

[0004] A data analysis method for identifying transgenic events using targeted sequencing comprises the following steps:

[0005] S1. Design probe and primer sets / primer sets for identifying target interval sequences of multiple transgenic events of the species to be tested, with each target interval sequence of the transgenic event having a corresponding probe and primer set / primer set;

[0006] S2. Constructing a reference sequence for a transgenic event: Connect the insert sequence and its upstream and downstream sequences of 300-500 bases for each transgenic event of the species to be tested via N arbitrary nucleotides to form a single reference sequence. This is then added to the reference genome sequence of the species as a scaffold to obtain a new reference genome containing the insert sequences of all transgenic events. The target intervals identifiable by each probe and primer set / or primer set in step S1 are located at their physical locations in the newly constructed reference genome sequence and a bed file is created. The target interval refers to the corresponding interval in the scaffold after the target interval sequences identifiable by all probes and primer sets / or primer sets for a transgenic event are combined.

[0007] S3. Targeted sequencing of target interval sequences for multiple transgenic events of the species to be tested;

[0008] S4. performing sequence alignment between the targeted sequencing results obtained in step S3 and the reference genome;

[0009] S5. Filter the alignment results for quality control, remove reads based on MAPQ and insert size, filter reads with multiple alignment sites, and eliminate interference caused by large spacing between alignment sites in paired-end reads.

[0010] S6. Calculate the depth and coverage of the target interval of the targeted sequencing based on the target location .bed file;

[0011] S7. Determine the type and transformation status of the transgenic event based on the combination of target interval coverage.

[0012] The probe and primer set / primer set used to detect the target interval sequence of each transgenic event in step S1 is divided into three parts, wherein:

[0013] The first part of the probe and primer set / primer set is a probe and primer set / primer set for identifying the element sequence in the transgenic event, and is used to determine whether the important elements of the transgenic event, such as the exogenous gene, have been successfully introduced;

[0014] The second part of the probe and primer set / primer set is used to identify the sequence at the junction of the transgenic event insertion fragment and the genome and its upstream and downstream sequences, and is used to determine whether the transgenic event has been successfully inserted into the genome. The second part of the probe and primer set / primer set for identifying each transgenic event includes: 5gene, 5gene-cross, 5gene-downflank, 3gene, 3gene-cross, 3gene-upflank;

[0015] 5gene represents the 5'end of the genome of the species to be tested of the inserted fragment of the transformation event, 5gene-cross represents the junction between the genome of the species to be tested and the inserted fragment at the 5'end, and 5gene-downflank represents the downstream of the 5'end junction. Similarly, 3gene represents the 3'end of the genome of the species to be tested of the inserted fragment of the transformation event, 3gene-cross represents the junction between the genome of the species to be tested and the inserted fragment at the 3'end, and 3gene-upflank represents the upstream of the 3'end junction.

[0016] The third part, probe and primer set / or primer set genome, is a probe and primer set / or primer set for identifying the original genomic sequence replaced by the inserted fragment of the transgenic event, and is used to determine whether the transgenic event has insertion in all homologous chromosomes, that is, to determine whether the transformation event is a homozygous genotype or a heterozygous genotype.

[0017] The number of arbitrary nucleotides N in step S2 is greater than twice the length of sequencing reads.

[0018] If 5gene-cross has overlapping regions with 5gene and 5gene-downflank, the region unique to 5gene-cross will be extracted as the core region and named 5gene-cross-unique to replace 5gene-cross;

[0019] If 3gene-cross has overlapping regions with 3gene-upflank and / or 3gene, the region unique to 3gene-cross will be extracted as the core region and named 3gene-cross-unique to replace 3gene-cross;

[0020] If the third part of the probe and primer set / or the primer set genome has overlapping regions with the 5gene-genome and / or the 3gene-genome, the region unique to the genome is extracted as the core region and named genome-unique to replace the genome.

[0021] Step S7 includes: if the similarity between the 5gene-cross, 3gene-cross, and genome sequences is greater than or equal to the similarity threshold, the target interval sequence coverage is required to reach 100%, and the average depth of the target interval sequence is greater than or equal to 10X, then the target interval sequence is considered to exist;

[0022] If the similarity between 5gene-cross, 3gene-cross, and genome sequences is lower than the similarity threshold, the target interval sequence coverage is required to be greater than or equal to 90%, and the average depth of the target interval sequence is greater than or equal to 5X, then the target interval sequence is considered to exist;

[0023] The similarity threshold was set as the minimum ratio of sequencing read length to 5gene-cross, 3gene-cross, and genome length.

[0024] In step S7, when the target interval sequence corresponding to the second part probe and primer set / or primer set of a transgenic event: 5gene-cross or its core region 5gene-cross-unique, 3gene-cross or its core region 3gene-cross-uniqu exists and the target interval sequence corresponding to the third part probe and primer set / or primer set genome or its core region genome-unique does not exist, it can be determined that the sample has the transgenic event and is present in all homologous chromosomes.

[0025] Furthermore, in step S7, if the target interval sequence corresponding to the probe and primer set / primer set at the junction of the transgenic event insert and the genome does not exist and the probe and primer set / primer set / primer set for the portion of the original genome sequence replaced by the transgenic event insert exists, then the material does not contain this transgenic event;

[0026] If the target interval sequence corresponding to the probe and primer set / or primer set at the junction of the transgenic event insert and the genome exists, the target interval sequence corresponding to the probe and primer set / or primer set of the portion of the original genome sequence replaced by the transgenic event insert does not exist, and the target interval sequence corresponding to the probe and primer set / or primer set of the important transgenic element exists, then this transgenic event is present in the material and the transgenic event insert is present on all homologous chromosomes;

[0027] If the target interval sequence corresponding to the probe and primer set / or primer set at the junction of the inserted fragment of a transgenic event and the genome exists, the target interval sequence corresponding to the probe and primer set / or primer set of the original genomic sequence portion replaced by the inserted fragment of the transgenic event exists, and the target interval sequence corresponding to the probe and primer set / or primer set of the important transgenic element exists, then this transgenic event exists in the material but the inserted fragment of the transgenic event exists on some homologous chromosomes.

[0028] The species to be tested is corn, and the multiple transgenic events include: LP007-1, ND207, DBN9501, DBN9858, LP026-2, nCX-1, and DBN9936.

[0029] The beneficial effects of the present invention are as follows: no specific detection kit is required. Corresponding probes and primer sets / primer sets for identifying target interval sequences are designed based on the transgenic event to be identified and the reference genome of the corresponding species. After targeted sequencing, the analysis process introduced in the invention can be used to use the sequencing data to simultaneously identify and detect multiple transgenic events in the test material.

[0030] The reference genome and probe and primer sets / or primer sets for identifying target interval sequences designed according to the present invention, and the identification methods can be used to identify any transgenic event of any species. The present invention is not limited to the use disclosed in the specification for detecting the transgenic corn event LP026-2 in corn plants. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a coverage map of the target interval of a corn sample detected by the method of the present invention;

[0032] Figure 2 is a flow chart of the method of the present invention;

[0033] Figure 3 An example of a primer set identifying region within the full-length sequence of the genome of a transgenic event. DETAILED DESCRIPTION

[0034] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0035] The primer set in this embodiment can also be replaced by a probe and primer set.

[0036] The primer set and probe in this embodiment use specific sequences to detect or amplify small fragments of DNA or RNA molecules. The primer set and probe are designed to be complementary to the target DNA or RNA sequence.

[0037] Primers are short oligonucleotides used to specifically recognize and bind to target DNA sequences in experiments such as DNA replication, amplification, and sequencing. Primers typically consist of 20-30 bases and should be complementary or partially complementary to the target DNA sequence. They guide amplification reactions and selectively detect target sequences during the experiment.

[0038] A probe is a specific oligonucleotide sequence consisting of a reporter fluorophore and a quencher fluorophore used in fluorescent probe assays. The probe typically consists of 20-30 bases and can specifically bind to a target DNA or RNA sequence. After the probe and target sequence are complementary and paired, the 5'-3' exonuclease activity of the Taq enzyme cleaves and degrades the probe during PCR amplification, separating the reporter fluorophore from the quencher fluorophore. The probe then emits fluorescence under specific light excitation. As the number of cycles increases, the amplified target gene fragment increases exponentially. By measuring the corresponding fluorescence signal intensity as it changes with amplification in real time, the Ct value is determined. By using several standards with known template concentrations as controls, the copy number of the target gene in the sample being tested can be determined.

[0039] Taking the identification of LP026-2 transgenic corn as an example: A data analysis method for identifying transgenic events using targeted sequencing includes the following steps:

[0040] 1. A primer set was designed to identify target region sequences of multiple transgenic events in maize. The transgenic events that can be identified include: LP007-1, ND207, DBN9501, DBN9858, LP026-2, nCX-1, and DBN9936.

[0041] LP007-1 is a transgenic corn developed by Longping Biotechnology (Hainan) Co., Ltd., which is characterized by resistance to feeding damage by lepidopteran pests and tolerance to the phytotoxic effects of agricultural herbicides containing glyphosate;

[0042] ND207 is a transgenic corn developed by Lai Jinsheng's team at China Agricultural University. It is highly resistant to major lepidopteran pests of corn, including corn borer, armyworm, cotton bollworm, and peach borer, and can effectively control the fall armyworm.

[0043] DBN9501 is a genetically modified corn developed by Da Bei Nong Biotechnology Co., Ltd., which is characterized by good resistance to lepidopteran insects and good tolerance to glufosinate herbicide.

[0044] DBN9858, a genetically modified corn developed by Da Bei Nong Biotechnology Co., Ltd., is characterized by good tolerance to glyphosate and glufosinate herbicides;

[0045] LP026-2 is a transgenic corn event developed by Longping Biotechnology (Hainan) Co., Ltd., characterized by good resistance to lepidopteran pests and tolerance to agricultural herbicides containing glyphosate;

[0046] nCX-1, a herbicide-tolerant corn developed by Hangzhou Ruifeng Biotechnology Co., Ltd., contains two herbicide-resistant genes, CdP450 and cp4epsps, and has good tolerance to the herbicides flazasulfuron, nicosulfuron, dimethoate, 2,4-D, and glyphosate;

[0047] DBN9936 is a genetically modified corn variety developed by the Chinese company Beijing Da Bei Nong Biotechnology Co., Ltd. This variety exhibits good resistance to major lepidopteran pests, such as the corn borer, oriental armyworm, peach borer, and cotton bollworm, and is also tolerant to glyphosate herbicide.

[0048] The primer sets used to detect each of the above transgenic events are divided into three partial primer sets:

[0049] The first part of the primer set is a primer set for identifying the element sequence in the transgenic event. It is used to determine whether the important elements of the transgenic event, such as the exogenous gene, have been successfully introduced. The first part of the primer set is named according to the element name:

[0050] Figure 1 The first part of the primer set for the transgenic event LP007-1 includes cCry1Ab, cCry2Ab, cEPSPS, cVip3Aa, and tNOS. The sequence information of the corresponding elements can be found in patent number 202110112168.1;

[0051] The first part of the primer set for transgenic event ND207 includes cBar, cCry1Ab, cCry2Ab, and tNOS. The sequence information of the corresponding elements can be extracted in patent number 202011220206.7 based on the full-length sequence and element interval provided in the patent;

[0052] The first part of the primer set for transgenic event DBN9501 includes cPAT, cVip3Aa19, and tNOS. The sequence information of the corresponding elements can be extracted from patent number 201910280088.X based on the full-length sequence and element interval provided in the patent;

[0053] The first part of the primer set for transgenic event DBN9858 includes cEPSPS, cPAT, and tNOS. The sequence information of the corresponding elements can be extracted from patent number 201510219911.8 based on the full-length sequence and element interval provided in the patent;

[0054] The first primer set for transgenic event LP026-2 is not available, while the second and third primer sets can be extracted from Patent No. 202211161744.2 based on the element intervals provided in the patent;

[0055] The first part of the primer set for the transgenic event nCX-1 includes cCP4EPSPS and cN-Z1. The sequence information of the corresponding elements can be extracted from patent number 202210538032.1 based on the full-length sequence and element interval provided in the patent;

[0056] The first part of the primer set for transgenic event DBN9936 includes cCry1Ab, cEPSPS, and tNOS. The sequence information of the corresponding elements can be extracted from patent number 201510220034.6 based on the full-length sequence and element interval provided in the patent;

[0057] cCry1Ab: It is an element sequence derived from Bacillus thuringiensis. The protein it encodes has resistance to lepidopteran insects and can control pests such as corn borer, cotton bollworm, and fall armyworm.

[0058] cCry2Ab: It is an element sequence derived from Bacillus thuringiensis, which has resistance to lepidopteran insects and provides broad-spectrum anti-insect protection together with cCry1Ab.

[0059] cEPSPS: Used to provide resistance to glyphosate herbicides. This gene encodes an enzyme that can resist the inhibitory effects of glyphosate, allowing transgenic plants to grow in the presence of glyphosate.

[0060] cVip3Aa: is an insecticidal protein gene derived from Bacillus thuringiensis. The protein it encodes is toxic to certain insects and can provide insect resistance to transgenic plants.

[0061] tNOS: usually refers to the terminator of the nopaline synthase gene, which is used as a transcription termination signal in transgenic construction to ensure that the mRNA after gene expression can be correctly terminated.

[0062] cBar: is a glyphosate-resistant element sequence that encodes a protein that can tolerate glyphosate herbicides, enabling transgenic corn to resist the phytotoxicity caused by glyphosate, so that when glyphosate herbicides are used in the field, it will not affect the growth of corn.

[0063] cVip3Aa19: confers resistance to lepidopteran insects.

[0064] cCP4EPSPS: is a gene that confers resistance to glyphosate herbicide.

[0065] p35S is a promoter sequence derived from Cauliflower Mosaic Virus (CaMV). As it is a common element in transgenic events, it has low specificity. Figure 1 The heat map shown has the information about the presence of this component filtered out.

[0066] The second part of the primer set is used to identify the sequence at the junction of the transgenic event insert and the genome and its upstream and downstream sequences. It is used to determine whether the transgenic event has been successfully inserted into the genome. The second part of the primer set for identifying each transgenic event includes: 5gene, 5gene-cross, 5gene-downflank, 3gene, 3gene-cross, 3gene-upflank;

[0067] 5gene represents the 5' end of the maize genome of the inserted fragment of the transformation event, 5gene-cross represents the junction between the maize genome and the inserted fragment at the 5' end, and 5gene-downflank represents the downstream of the 5' end junction. Similarly, 3gene represents the 3' end of the maize genome of the inserted fragment of the transformation event, 3gene-cross represents the junction between the maize genome and the inserted fragment at the 3' end, and 3gene-upflank represents the upstream of the 3' end junction. Figure 1 In the expression "5gene-genome", it indicates whether the genomic sequence at the 5' end of the inserted fragment exists in the original genome, and "3gene-genome" indicates whether the genomic sequence at the 3' end of the inserted fragment exists in the original genome. The sequences of "5gene-genome / 3gene-genome" and "5gene / 3gene" are the same, and "5gene / 3gene" are the positions on the constructed contig.

[0068] like Figure 3 SEQ ID NO.5 is a full-length sequence of the genome of a transgenic event, and the inserted fragment of the transgenic event is Figure 3 The first part of the primer set is used to identify the sequence between bNLB and bNRB in the insert. Different elements have different functions, including the coding region, promoter region and terminator region of the gene. Different transformation events may have common elements. The second part of the primer set corresponds to Figure 3 SEQ ID NO.1 and SEQ ID NO.2 in .

[0069] The third primer set is a primer set for identifying the original genomic sequence replaced by the inserted fragment of the transgenic event, and is used to determine whether the transgenic event has inserted into all homologous chromosomes. The third primer set is named genome.

[0070] The first and second parts of the sequence information can be obtained from the patent documents of the transgenic event, while the third sequence needs to be located by locating the genomic sequences connected upstream and downstream of the transgenic event insert. After obtaining the target sequence, primers need to be designed based on the length of the targeted sequencing reads. The PCR product length should be less than twice the read length, and the specificity of each primer pair should be guaranteed, that is, each primer pair can only produce a single PCR product.

[0071] Based on the genome sequences at both ends of the insertion fragment provided in the above patent, the reference genome was BLASTed to obtain the position of the insertion event in the genome, as shown in Table 1 below. The reference genome is maize B73v5 version, GenBank number GCA_902167145.1

[0072] Table 1 Location of transgenic events in the genome

[0073] transgene chr start end LP007-1 chr3 23243158 23243209 ND207 chr3 181367242 181367272 DBN9501 chr3 227681332 227681413 DBN9858 chr4 182375301 182375347 LP026-2 chr5 31559513 31559523 nCX-1 chr7 176476146 176484018 DBN9936 chr9 8370571 8370663

[0074] 2. Using LP026 transgenic corn as the experimental material, targeted sequencing of target interval sequences was performed on 10 corn materials WF1~8, WF365 and WF366. Among them, the 8 corn sample materials WF1~8 were known to have no other transgenic events except LP026, and WF365 and WF366 were non-transgenic materials.

[0075] 3. The insert fragment of each transgenic event and its upstream and downstream 300 to 500 bases are taken as a whole sequence. The whole sequences of all different transgenic events are connected with 500 Ns (N represents any nucleotide. Several N linker sequences are used to keep a certain distance between different transgenic events to ensure accurate comparison). This forms a contig formed by the whole sequence, forming a separate reference sequence. This reference sequence represents the insert fragments of all transgenic events and their upstream and downstream sequences. This separate reference sequence is added to the maize B73v5 reference genome. The final result is a reference genome containing all transgenic events. This genome can be used to compare the sequencing data of experimental materials and perform genomic analysis of experimental materials to determine whether the transgenic events exist in the individual genomes of the experimental materials.

[0076] 4. Create a bed file based on the physical location of the target interval sequence identified by each primer set in step 1 in the reference genome newly constructed in step 4.

[0077] A bed file is a text file format used to store information about genomic regions, including chromosome names, the start and end locations of the region, and other information. In this step, the locations of all target sequences (i.e., the target interval sequences identified by the primer set designed in Step 1) in the newly constructed reference genome from Step 3 are organized into a .bed file. This file can be used to guide subsequent data analysis, such as determining which regions in the sequencing data correspond to specific transgenic events.

[0078] When making a bed file, if 5gene-cross has overlapping regions with 5gene and 5gene-downflank in the second and third primer sets, the region unique to 5gene-cross will be extracted separately and named 5gene-cross-unique, and the position represented by it will be added to the bed file. This region is the core region of 5gene-cross, which includes the genome of the species to be tested and the junction of the insert fragment; similarly, if 3gene-cross has overlapping regions with 3gene-upflank or 3gene, the region unique to 3gene-cross will be extracted separately as the core region and named 3gene-cross-unique, and the position represented by it will be added to the bed file; if the third primer set genome has overlapping regions with 5gene-genome and 3gene-genome, the region unique to genome will be extracted separately as the core region and named genome-unique, and the positions of these unique regions will be added to the bed file.

[0079] 5. Compare the targeted sequencing results obtained in step 2 with the reference genome added to the constructed contig.

[0080] 6. Filter and quality control the alignment results by comparing alignment quality (MAQ) and read insert length. Alignment quality can filter reads with multiple alignment sites, while insert length filtering can eliminate interference caused by large spacing between paired-end read alignment sites. For example, reads with a MAPQ less than 1 and an insert size greater than 300 are removed to remove reads with multiple alignment sites and reads with pair reads that are too far apart.

[0081] In this step, the reads obtained from the targeted sequencing are aligned with the reference genome constructed in step 3, and reads with a MAPQ less than 1 are removed to reduce false positive results in subsequent analysis.

[0082] In this step, the two paired reads should originate from the two ends of the same DNA fragment. The insert size refers to the distance between the two reads in the genome. If the insert size is too large, it may indicate that the two reads do not actually originate from the same DNA fragment or that there are large deletions or duplications in the DNA fragment. This step sets a reasonable insert size threshold (e.g., 300bp, twice the sequencing read length) to remove these anomalous reads.

[0083] 7. Calculate the depth and coverage of each target interval in the targeted sequencing results based on the target location .bed file.

[0084] The target interval refers to the interval corresponding to the target interval sequence combination in the Scaffold that can be identified by all primer sets of a transgenic event.

[0085] In this step, depth refers to the average number of times each base is covered by sequencing reads in a specific region of the genome. The higher the depth, the more thoroughly the region is covered.

[0086] Coverage refers to the proportion of the area covered by sequencing reads in the target interval.

[0087] Scaffold refers to an intermediate product in the genome assembly process. During the genome sequencing and assembly process, due to technical limitations, it is usually impossible to obtain the complete sequence of the entire genome at one time. Therefore, the shorter DNA fragments obtained by sequencing (called reads) are assembled into longer continuous sequences. These continuous sequences are called contigs. When multiple contigs are connected through overlapping regions to form a larger sequence framework, this framework is called a scaffold. A scaffold can contain one or more contigs, as well as gaps between them. These gaps represent regions of DNA sequence that have not yet been determined.

[0088] 8. Dynamically adjust the threshold for judging whether the target interval sequence exists based on the characteristics of the primer set of the transgenic event, the accuracy requirements of the experiment, and other factors. For example, when the sequence similarity of important element intervals (such as 5gene-cross, 3gene-cross, genome) in different transgenic events is greater than or equal to the similarity threshold, in order to more accurately judge whether the primer set sequence exists, if the target interval coverage reaches 100% and the average depth of the target interval reaches 10X or above, the target interval is considered to exist. Based on this condition, the coverage of each sample in each target interval is integrated and a heat map is drawn, such as Figure 1As shown, red represents that the coverage does not reach 100%, and blue represents that the coverage reaches 100%, and it is believed that the sequence in this interval really exists;

[0089] When the sequence similarity of important element intervals (such as 5gene-cross, 3gene-cross, and genome) in different transgenic events is lower than the similarity threshold, the target interval coverage is required to reach 90% or above, and the average depth of the target interval is required to reach 5X or above. This can not only ensure a certain level of accuracy, but also improve detection efficiency and reduce unnecessary resource consumption. By comprehensively considering these factors, the target interval coverage threshold can be flexibly adjusted to ensure more accurate judgment of the presence or absence of the target interval in different situations.

[0090] Among them, in order to more accurately determine whether the target interval sequence exists, the sequence similarity threshold of the important element interval is set as the minimum value of the ratio of the sequencing read length to the element interval length of the transgenic event. For example, the ratio of the sequencing read length to the 5gene-cross, 3gene-cross, and genome length is compared. If the minimum ratio is equal to 60%, then when the sequence similarity of the important element interval is greater than or equal to 60%, the similarity between the sequences is considered high. If the target interval sequence coverage reaches 100% and the average depth of the target interval sequence reaches 10X or above, the target interval sequence is considered to exist. If the sequence similarity of the important element interval is less than 60%, the similarity between the sequences is considered low. The target interval sequence coverage is required to reach 90% or above and the average depth of the target interval sequence reaches 5X to be considered to exist.

[0091] 9. According to the judgment method of the present invention, only when the target interval sequence corresponding to 5gene-cross of material WF1~8 (representing the connection part between the 5' end corn genome and the inserted fragment) or its core area 5gene-cross-unique, the target interval sequence corresponding to 3gene-cross (representing the connection part between the 3' end corn genome and the inserted fragment) or its core area 3gene-cross-unique exists and the target interval sequence corresponding to genome (the original genome sequence replaced by the inserted fragment of the transgenic event) or its core area genome-unique does not exist, can it be judged that the sample has the transgenic event and exists on all homologous chromosomes. Figure 1 The heat map shows that only LP026 meets this condition, and no other transformation events exist. The results are in line with expectations.

[0092] Furthermore, step 9 can further determine whether the inserted fragment of the transformation event exists in all homologous chromosomes or part of the homologous chromosomes:

[0093] If the target interval sequence corresponding to the primer set at the junction of the transgenic event insert and the genome does not exist, and the target interval sequence corresponding to the primer set of the part of the original genome sequence replaced by the transgenic event insert exists, it is considered that there is no such transformation event in the material;

[0094] If the target interval sequence corresponding to the primer set at the junction of the transgenic event insert and the genome exists, the target interval sequence corresponding to the primer set of the part of the original genome sequence replaced by the transgenic event insert does not exist, and the target interval sequence corresponding to the primer set of the important transgenic element exists, then it is considered that this transformation event exists in the material and the transformation event insert is present on all homologous chromosomes;

[0095] If the target interval sequence corresponding to the primer set at the junction of the transgenic event insert and the genome exists, the target interval sequence corresponding to the primer set of the part of the original genome sequence replaced by the transgenic event insert exists, and the target interval sequence corresponding to the primer set of the important transgenic element exists, then it is considered that this transformation event exists in the material, but the transformation event insert exists on some homologous chromosomes.

Claims

1. A data analysis method for identifying transgenic events using targeted sequencing, characterized in that: The following steps are involved: S1. Design probe and primer sets / primer sets for identifying target interval sequences of multiple transgenic events of the species to be tested, with each target interval sequence of each transgenic event having a corresponding probe or primer set; S2. Constructing a reference sequence for a transgenic event: The insert sequence and its upstream and downstream sequences of 300-500 bases for each transgenic event of the species to be tested are linked together via N arbitrary nucleotides to form a single reference sequence. This is then added to the reference genome sequence of the species as a scaffold to generate a new reference genome containing the insert sequences of all transgenic events. The target intervals identifiable by each probe and primer set / or primer set in step S1 are located at their physical locations in the newly constructed reference genome sequence and a bed file is generated. The target interval refers to the corresponding interval in the scaffold after the target interval sequences identifiable by all probes or primer sets for a transgenic event are combined. S3. Targeted sequencing of target interval sequences for multiple transgenic events of the species to be tested; S4. performing sequence alignment between the targeted sequencing results obtained in step S3 and the reference genome; S5. Perform quality control on the alignment results by filtering reads based on alignment quality (MAPQ) and insert size, filtering reads with multiple alignment sites, and eliminating interference caused by large spacing between alignment sites in paired-end reads. S6. Calculate the depth and coverage of each target interval in the targeted sequencing results based on the target location .bed file; S7. Determine the type and transformation status of the transgenic event based on the combination of target interval coverage.

2. The data analysis method for identifying transgenic events using targeted sequencing according to claim 1, characterized in that: The probe and primer set / primer set used to detect the target interval sequence of each transgenic event in step S1 is divided into three parts, wherein: The first part of the probe and primer set / primer set is a probe and primer set / primer set for identifying the element sequence in the transgenic event, and is used to determine whether the important elements of the transgenic event, such as the exogenous gene, have been successfully introduced; The second part of the probe and primer set / primer set is used to identify the sequence at the junction of the transgenic event insertion fragment and the genome and its upstream and downstream sequences, and is used to determine whether the transgenic event has been successfully inserted into the genome. The second part of the probe and primer set / primer set for identifying each transgenic event includes: 5gene, 5gene-cross, 5gene-downflank, 3gene, 3gene-cross, 3gene-upflank; 5gene represents the 5'end of the genome of the species to be tested of the inserted fragment of the transformation event, 5gene-cross represents the junction between the genome of the species to be tested and the inserted fragment at the 5'end, and 5gene-downflank represents the downstream of the 5'end junction. Similarly, 3gene represents the 3'end of the genome of the species to be tested of the inserted fragment of the transformation event, 3gene-cross represents the junction between the genome of the species to be tested and the inserted fragment at the 3'end, and 3gene-upflank represents the upstream of the 3'end junction. The third part, probe and primer set / or primer set genome, is a probe and primer set / or primer set for identifying the original genome sequence replaced by the inserted fragment of the transgenic event, and is used to determine whether the transgenic event has inserted into all homologous chromosomes.

3. The data analysis method for identifying transgenic events using targeted sequencing according to claim 1, characterized in that: The number of arbitrary nucleotides N in step S2 is greater than twice the length of the sequencing reads.

4. The data analysis method for identifying transgenic events using targeted sequencing according to claim 2, characterized in that: If 5gene-cross has overlapping regions with 5gene and 5gene-downflank, the region unique to 5gene-cross will be extracted as the core region and named 5gene-cross-unique to replace 5gene-cross; If 3gene-cross has overlapping regions with 3gene-upflank and / or 3gene, the region unique to 3gene-cross will be extracted as the core region and named 3gene-cross-unique to replace 3gene-cross; If the third part of the probe and primer set / or the primer set genome has overlapping regions with the 5gene-genome and / or the 3gene-genome, the region unique to the genome is extracted as the core region and named genome-unique to replace the genome.

5. The data analysis method for identifying transgenic events using targeted sequencing according to claim 1, characterized in that: The two paired reads in step S5 should come from the two ends of the same DNA fragment.

6. The data analysis method for identifying transgenic events using targeted sequencing according to claim 1, characterized in that: Step S7 includes: if the similarity between the 5gene-cross, 3gene-cross, and genome sequences is greater than or equal to the similarity threshold, the target interval sequence coverage is required to reach 100%, and the average depth of the target interval sequence is greater than or equal to 10X, then the target interval sequence is considered to exist; If the similarity between 5gene-cross, 3gene-cross, and genome sequences is lower than the similarity threshold, the target interval sequence coverage is required to be greater than or equal to 90%, and the average depth of the target interval sequence is greater than or equal to 5X, then the target interval sequence is considered to exist; The similarity threshold was set as the minimum ratio of sequencing read length to 5gene-cross, 3gene-cross, and genome length.

7. The data analysis method for identifying transgenic events using targeted sequencing according to claim 1, characterized in that: In step S7, when the target interval sequence corresponding to the second part probe and primer set / or primer set of a transgenic event: 5gene-cross or its core region 5gene-cross-unique, 3gene-cross or its core region 3gene-cross-uniqu exists and the target interval sequence corresponding to the third part probe and primer set / or primer set genome or its core region genome-unique does not exist, it can be determined that the sample has the transgenic event and is present in all homologous chromosomes.

8. The data analysis method for identifying transgenic events using targeted sequencing according to claim 1, characterized in that: In step S7, if the target interval sequence corresponding to the probe and primer set / primer set at the junction of the transgenic event insert and the genome does not exist and the probe and primer set / primer set / primer set for the portion of the original genome sequence replaced by the transgenic event insert exists, then the transgenic event is not present in the material; If the target interval sequence corresponding to the probe and primer set / or primer set at the junction of the transgenic event insert and the genome exists, the target interval sequence corresponding to the probe and primer set / or primer set of the portion of the original genome sequence replaced by the transgenic event insert does not exist, and the target interval sequence corresponding to the probe and primer set / or primer set of the important transgenic element exists, then this transgenic event is present in the material and the transgenic event insert is present on all homologous chromosomes; If the target interval sequence corresponding to the probe and primer set / or primer set at the junction of the inserted fragment of a transgenic event and the genome exists, the target interval sequence corresponding to the probe and primer set / or primer set of the original genomic sequence portion replaced by the inserted fragment of the transgenic event exists, and the target interval sequence corresponding to the probe and primer set / or primer set of the important transgenic element exists, then this transgenic event exists in the material but the inserted fragment of the transgenic event exists on some homologous chromosomes.

9. The data analysis method for identifying transgenic events using targeted sequencing according to claim 1, characterized in that: The species to be tested is corn, and the multiple transgenic events include: LP007-1, ND207, DBN9501, DBN9858, LP026-2, nCX-1, and DBN9936.

Citation Information

Patent Citations

  • Nucleic acid sequence for detecting corn plant DBN9936, and detection method thereof

    CN104830847A

  • Nucleotide sequence and detection method for detecting herbicide-tolerant maize plant DBN9858

    CN104878095A

  • Nucleic acid sequence and detection methods for detecting corn plant DBN9501

    CN109868273A

  • Corn event 2A-7 and identification method thereof

    CN112280743A

  • Genetically modified corn incident LP007-1 and its detection method

    CN112852801B