Method for identifying exogenous gene insertion information through three-generation sequencing
By using high-throughput third-generation sequencing technology and IGV software for comparison, the complexity and inaccuracy of foreign gene insertion detection have been solved, enabling accurate identification of foreign gene insertion sites in transgenic plants, ensuring biosafety and product differentiation.
Patent Information
- Application Number
- CN202511144574.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-18
AI Technical Summary
Existing methods for detecting exogenous gene insertion are complex, time-consuming, and susceptible to human error, making it difficult to accurately identify the insertion site and copy number of exogenous genes in transgenic plants, thus affecting biosafety assessment and product differentiation.
Using high-throughput third-generation sequencing technology, combined with ONT sequencing and IGV software alignment, the insertion frequency, location, and copy number of exogenous gene fragments in the recipient genome were identified through sequencing depth analysis and insertion site identification, and the relationship with coding genes near the insertion site was also considered.
It enables precise identification of exogenous gene insertion sites in transgenic plants, assesses their safety, ensures genetic stability, provides intellectual property protection, meets regulatory requirements, and distinguishes different transgenic products.
Smart Images

Figure CN120966971A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of plant transgenic technology, and particularly relates to a method for identifying an insertion site and copy number of an exogenous gene fragment in a transgenic plant based on high-throughput third-generation sequencing. BACKGROUND
[0002] Plant transgenic technology plays a very important role and significance in crop and forest tree breeding. First, plant transgenic technology can break through the species barrier, introduce excellent genes from different species into plants, and greatly widen the range of available gene resources. Second, compared with traditional breeding techniques, transgenic breeding can more accurately improve specific traits, quickly obtain transgenic plants with desired traits by directly introducing target genes into recipient plants, thereby significantly shortening the breeding cycle and accelerating the cultivation and promotion speed of excellent varieties. Finally, plant varieties cultivated by transgenic technology, such as insect-resistant, disease-resistant, stress-resistant, and herbicide-tolerant varieties, can maintain good growth and biomass under reduced pesticide and fertilizer inputs, thereby reducing costs and pollution to soil, water, and air, improving economic efficiency and resource utilization efficiency, and being conducive to the sustainable development of ecological environment and biodiversity.
[0003] Biological safety detection of transgenic plants is an important link to ensure the safety of their commercial application, and detection of the insertion site of exogenous genes is of great significance for safety evaluation and commercial application of transgenic plants.
[0004] Traditional methods for detecting the insertion of exogenous genes into the recipient genome are mainly PCR-based detection techniques, such as inverse PCR, adapter PCR, and TAIL-PCR. These methods obtain the flanking sequences adjacent to the exogenous DNA by specific PCR reaction strategies combined with DNA sequencing technology. However, these methods have limitations such as complex operation, long time-consuming, high false positive rate, etc. Southern hybridization technology can also be used to detect the copy number and approximate position of the insertion site of exogenous genes, but it requires long operation time, high quality and quantity of sample DNA, is easily disturbed by human operation, and provides limited information. At present, the method of whole-genome individual resequencing is a relatively ideal method for detecting the insertion site of exogenous genes. By high-throughput resequencing of the genome of transgenic plants, the sequencing data are compared and analyzed with the wild-type genome sequence, and the insertion site, copy number, and sequence changes around the insertion site of the exogenous gene are determined by bioinformatics methods. In particular, the third-generation sequencing technology can obtain longer sequencing fragments, and can more accurately analyze the insertion site, integration mode, and structure of the exogenous DNA. SUMMARY
[0005] The present application takes the plants of different strains of transgenic poplar as materials, carries out high-throughput third-generation ONT sequencing, compares with the exogenous target gene fragment and the genome of the receptor Populus alba x P.glandulosa, and obtains the copy number and insertion site of the exogenous gene fragment inserted into the receptor genome by using the comparison of sequencing depth and the comparison result of the exogenous.
[0006] In order to achieve the above-mentioned purpose, the present application provides the following technical scheme:
[0007] The present application provides a method for identifying the insertion site and copy number of an exogenous gene fragment in a transgenic plant based on high-throughput third-generation sequencing, which comprises the following steps:
[0008] (1) extracting DNA from the plant sample to be tested; (2) high-throughput third-generation ONT library sequencing; (3) integrating the exogenous target gene fragment and the receptor genome as a reference genome; (4) sequence alignment and identification of the sequencing fragments and the reference genome; 1) identifying the insertion frequency of the exogenous gene fragment; 2) identifying the insertion site of the exogenous target gene fragment into the receptor genome; 3) identifying the position of the insertion site of the exogenous target gene fragment in the coding gene of the receptor genome.
[0009] In the present application, the plant includes poplar.
[0010] In the present application, the specific operation of identifying the insertion frequency of the exogenous gene fragment in step 1) is as follows:
[0011] The sequencing depth of the sequencing reads aligned with the exogenous target gene fragment in the reference genome sequence and the average sequencing depth of the receptor 84K genome are counted respectively, and the insertion frequency of the exogenous gene fragment is identified;
[0012] When the alignment sequencing depth of the exogenous target gene fragment is ≤0.3, it indicates that there is no insertion of the exogenous target gene fragment on the receptor chromosome;
[0013] When the alignment sequencing depth of the exogenous target gene fragment is 0.7 times to 1.3 times of the average sequencing depth of the receptor genome, it indicates that there is one insertion of the exogenous target gene fragment on the receptor chromosome;
[0014] When the alignment sequencing depth of the exogenous target gene fragment is 1.7 times to 2.3 times of the average sequencing depth of the receptor genome, it indicates that there are two insertions of the exogenous target gene fragment on the receptor chromosome;
[0015] By analogy, the frequency of the insertion of the exogenous target gene fragment on the receptor chromosome can be determined;
[0016] In actual operation, when the alignment sequencing depth of the exogenous target gene fragment is >0.3 times and <0.7 times, or >1.3 times and <1.7 times, too many miscellaneous items are mixed in, and the insertion situation needs to be analyzed in combination with the specific insertion site.
[0017] In the present application, the specific operation of the step 2) identifying the insertion of the exogenous target gene fragment into the receptor genome site is as follows:
[0018] (1) Using IGV software, the alignment information of the reads aligned to the end of the exogenous target gene fragment and clipped is used to provide reference information for identifying the insertion of the exogenous target gene fragment into the receptor 84K genome;
[0019] (2) According to the alignment information stored in the sam / bam file, the perl script is used to identify and extract the reads aligned to the end of the exogenous target gene fragment and clipped, and the insertion site of the exogenous target gene fragment into the receptor genome is judged according to the extracted alignment information;
[0020] (3) According to the insertion site information of the receptor genome obtained in (2), the corresponding insertion point position is searched in the receptor genome by using IGV software, and if there is an obvious gap or depth reduction near the site, it can be judged as the starting site or interval of the insertion of the exogenous target gene fragment.
[0021] In the present application, the specific operation of the step 3) identifying that the insertion site of the exogenous target gene fragment is located in the position of the coding gene in the receptor genome is as follows:
[0022] Based on the identification of the insertion site of the exogenous target gene fragment according to claim 1, the coding genes within 5000bp upstream and downstream of the insertion site are extracted as candidate genes to determine the relationship between the insertion site and the adjacent coding gene on the receptor genome; and by the direction of the insertion of the exogenous target gene fragment into the receptor genome and the transcription direction of the receptor genome and the coding gene, it is determined that the insertion site is located upstream or downstream of the coding gene.
[0023] The present application also provides the application of the above-mentioned identification method in evaluating the safety of transgenic plants.
[0024] The present application also provides the application of the above-mentioned identification method in exploring gene expression and function.
[0025] The present application also provides the application of the above-mentioned identification method in screening transgenic lines with stable inheritance.
[0026] The present application also provides the application of the above-mentioned identification method in distinguishing different transgenic products by detecting the insertion site of the exogenous DNA.
[0027] The identification method provided by the application can effectively evaluate the safety of transgenic plants. The insertion site of exogenous DNA in the recipient genome may trigger unexpected safety of the transgenic plant, such as causing silencing or activating negative regulators of some functional genes. Therefore, accurate detection of the insertion site and its flanking sequence helps to evaluate its impact on the plant genome, and further judge whether the transgenic plant has potential safety risks, which is an important basis for safety evaluation. Secondly, gene expression and function can be studied. The insertion site of the exogenous gene may affect the expression level and expression pattern of the exogenous gene. Through the determination of the insertion site, the interaction between the exogenous gene and the plant genome can be better understood, and the regulation mechanism of this interaction on the expression and function of the exogenous gene can provide an important basis for in-depth study of gene function. Meanwhile, the application can also screen transgenic lines with stable inheritance. In the process of plant transgenesis, the insertion site and copy number of the exogenous gene will affect its genetic stability in the offspring. Clarifying the insertion site helps to screen transgenic lines that can stably inherit the exogenous gene, and provides a reliable material basis for the cultivation and application of transgenic plants. In addition, the application can provide a basis for intellectual property protection. For each transgenic plant transformant, the insertion site of the exogenous DNA is unique. The detection method established by taking the exogenous vector and the connection region sequence of the recipient genome as the target has transformant specificity, and the unique insertion site and event-specific information can be used as an important identifier for transgenic plant varieties, for distinguishing different transgenic products, and protecting the legal rights of the researchers. Finally, it complies with the requirements of regulations and supervision. Accurate detection of the insertion site of the exogenous gene is an important link to meet the regulatory requirements, safety evaluation and approval of transgenic products, and helps to ensure the legal, safe and sustainable application of transgenic technology. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 ONT sequencing data before and after filtering; A and B: ONT-Cas:EXPA1-9 strain; C and D: ONT-Cas:EXPA1-12 strain; A, C: distribution graph before filtering of sequencing data; B, D: distribution graph after filtering of sequencing data; A: A base ratio; T: T base ratio; C: C base ratio; G: G base ratio; N: missing base ratio; Qmean: average base quality; Q5: reads ratio with quality value greater than Q5;
[0029] Figure 2 Comparison of sequencing reads of different transgenic strains with exogenous target gene fragments and recipient genome sequencing depth; Genome: 84K genome sequence of the recipient; Cas9_EXPA1: exogenous target gene fragment sequence;
[0030] Figure 3Figure 6 is an IGV visualization plot of the clipped reads in ONT-Cas:EXPA1-9 aligned to the exogenous gene fragment of interest in the genome of the recipient 84K; the exogenous gene fragment of interest is aligned to the insertion site region of chromosome Chr06B, B: insertion site region of chromosome Chr19A; a: Cas:EXPA1-12 transgenic line, b: Cas:EXPA1-9 transgenic line.
[0031] Figure 4 Figure 7 is an IGV visualization plot of the clipped reads in ONT-Cas:EXPA1-12 aligned to the insertion site in the genome of the recipient 84K; the exogenous gene fragment of interest is aligned to the insertion site region of chromosome Chr04B; a: Cas:EXPA1-9 transgenic line, b: Cas:EXPA1-12 transgenic line.
[0032] Figure 5 Figure 7 is an IGV visualization plot of the clipped reads in ONT-Cas:EXPA1-12 aligned to the insertion site in the genome of the recipient 84K; the exogenous gene fragment of interest is aligned to the insertion site region of chromosome Chr04B; a: Cas:EXPA1-9 transgenic line, b: Cas:EXPA1-12 transgenic line. DETAILED DESCRIPTION
[0033] In the following, unless otherwise specified, the reagents or instruments used are used according to the conventional experimental conditions, and unless otherwise specified, the reagents or instruments used are used according to the conditions recommended in the manufacturer's instructions. Unless otherwise specified, the reagents or instruments used are conventional products that can be obtained commercially.
[0034] In the present application, because the introduced fragments of the gene editing are the same, the sequences of the exogenous gene fragment of interest in the two gene editing lines ONT-Cas:EXPA1-9 and ONT-Cas:EXPA1-12 are consistent.
[0035] The present application is further described below in conjunction with specific experiments.
[0036] EXAMPLE
[0037] 1. DNA extraction and library sequencing
[0038] 1.1 Test sample material
[0039] The present method takes the pagEXPA1 gene editing Populus alba × P. tremula strain (ONT-Cas:EXPA1-9 and ONT-Cas:EXPA1-12) as the test object, extracts high-quality genomic DNA from the leaves of each strain, and uses Nanodrop, Qubit and 0.35% agarose gel electrophoresis to detect the purity, concentration and integrity of the extracted DNA, i.e. the DNA concentration needs to be higher than 150 ng / μL, OD 260 / 280 between 1.8 and 2.2, OD 260 / 230Between 1.8-2.5, the total amount of DNA needs to meet the total amount requirement calculated by the data amount, generally 8ug DNA can produce 50G data.
[0040] 1.2 Nanopore library construction sequencing
[0041] After the sample genomic DNA is detected to be qualified, Nanopore library construction sequencing is performed according to the standard method provided by Oxford Nanopore Technologies (ONT) company. First, the genomic DNA is broken into an average of about 8kb using gTube, the required fragment size is screened using Blue Pippin automatic nucleic acid fragment recovery system, and the nucleic acid fragment damage repair and end repair plus A are completed using NEBNext FFPEDNA Repair Mix and NEBNext Ultra II End Repair / dA-Tailing Module; the connection of sequencing adapters and library construction are completed using SQK-LSK109 kit, and the constructed library is first subjected to library quality inspection, and the qualified library is subjected to sequencing using the third generation Naonopore sequencing platform. The sequencing amount of each sample is about 15Gb, and the A-B two sets of genomes of silver gland poplar 84K (about 890M in size) are used as reference genomes, and the sequencing depth is about 15X. The raw data format obtained by high-throughput sequencing is binary fast5 format containing all the original sequencing signals, and a single read corresponds to a single fast5 file. After base calling by MinKNOW software, the fast5 format data is converted to fastq format for subsequent analysis.
[0042] 2. Sequencing data filtering
[0043] Because there may be some reads with low quality values or high adapter ratios in the data, these reads may bring some false alignments, so they need to be removed or trimmed. The raw data obtained by sequencing is filtered using TGSFilter software (v 1.10) (https: / / github.com / HuiyangYu / TGSFilter), and the main parameters for filtering are (-x ont-5-3-l 1000-q 10), and the default values are used for other parameters. The reads that do not meet the conditions will be filtered out as low-quality reads. Among them:
[0044] -x: specify the sequencing platform, the sequencing platform used in the present application is ONT sequencing platform, so ont is used;
[0045] -5: reads are truncated from the 5' end, and sequences with poor quality are truncated;
[0046] -3: reads are truncated from 3' end, sequences with low quality are truncated;
[0047] -l: after various low quality truncation, only reads with length greater than or equal to the value are retained and output, this patent only retains reads greater than or equal to 1000 bp;
[0048] -q: quality value filtering, only reads with average quality value greater than or equal to the value are retained, this project uses 10.
[0049] In order to intuitively show the quality of sequencing data, after the completion of raw data filtering, the modified version of fqcheck software (https: / / github.com / hewm2008 / fqcheck, v2.08) is used to count the content of A, T, C, G and N and the base quality value of each sequencing cycle before and after filtering. After the completion of statistics, the R software (https: / / www.r-project.org / , v4.3) Ggplot2 package (https: / / github.com / tidyverse / ggplot2, v3.5.1) is used to visualize the statistical results. Through the base quality and base composition distribution of the three generations of sequencing data before and after filtering (Fig. 1), it can be seen that the content ratio of N in the filtered data is very low, almost 0; A-T and C-G are obviously complementary; the base content of the front end of reads tends to be stable after filtering, indicating that the sequencing quality after filtering is significantly improved. Figure 1
[0050] 3. Integration of reference genome
[0051] In the linux environment, the cat command is used to integrate the exogenous target gene fragment sequence (definition see Appendix 1:3; as shown in SEQ ID No. 1, Cas9_EXPA1, 7161 bp) containing the exogenous expression vector (definition see Appendix 1:1) and the target gene (definition see Appendix 1:2) into the recipient 84K genome (definition see Appendix 1:4) sequence file (https: / / figshare.com / articles / dataset / 84K_genome_zip / 12369209) to obtain a merged sequence file, which is used as the reference genome sequence (definition see Appendix 1:5) for subsequent sequencing sequence alignment analysis.
[0052] 4. Alignment of sequencing fragments and reference genome
[0053] In order to identify the insertion site of the exogenous target gene fragment on the genome through the position of the related reads, the filtered sequencing reads were aligned to the integrated reference genome sequence using the minimap2 software in a Linux system. The main process and steps of the alignment are as follows:
[0054] 1) The filtered read sequence was aligned to the reference genome sequence using the minimap2 software with the sequence file of the integrated reference genome of the receptor 84K and the exogenous target gene fragment as the reference.
[0055] 2) The bam file saving the alignment information was output using the samtools software, the sequence was sorted according to the physical position using the sort function, and the index function was used to construct an index for the sorted bam file.
[0056] Application Example 1
[0057] Identification of the copy number of the exogenous target gene fragment inserted into the receptor 84K genome
[0058] In a Linux system, the Pandepth software was used to calculate the sequencing depth of the exogenous target gene fragment aligned to the reference genome sequence and the average sequencing depth of the receptor 84K genome according to the alignment data. The average sequencing depth calculation formula is: the number of bases aligned to the chromosome / the length of the chromosome. By comparing the alignment sequencing depth of the exogenous target gene fragment with the average sequencing depth of the receptor 84K genome, the frequency of the insertion of the exogenous target gene fragment into the receptor genome, i.e., the copy number of the exogenous target gene, can be determined.
[0059] 1) If the alignment sequencing depth of the exogenous target gene fragment is about 0 (usually ≤0.3), it indicates that there is no insertion of the exogenous target gene fragment into the receptor chromosome.
[0060] 2) If the alignment sequencing depth of the exogenous target gene fragment is basically consistent with the average sequencing depth of the receptor genome (i.e., 0.7 times to 1.3 times), it indicates that there is only one insertion of the exogenous target gene fragment into the receptor chromosome.
[0061] 3) If the alignment sequencing depth of the exogenous target gene fragment is about twice the average sequencing depth of the receptor genome (i.e., 1.7 times to 2.7 times), it indicates that there are two insertions of the exogenous target gene fragment into the receptor chromosome.
[0062] By analogy, the frequency of the insertion of the exogenous target gene fragment into the receptor chromosome can be determined.
[0063] The application compares sequencing depths of two strains ONT-Cas:EXPA1-9 and ONT-Cas:EXPA1-12 by sequencing depth comparison Figure 2 It can be seen that (1) the sequencing depth of ONT-Cas:EXPA1-12 on the exogenous target gene fragment (Cas9_EXPA1) is close to the sequencing depth of the receptor 84K genome (Genome), and the ratio is 0.81, which can be inferred that the exogenous target gene may have only one insertion in the genome of the ONT-Cas:EXPA1-12 strain; (2) The sequencing depth of ONT-Cas:EXPA1-9 on the exogenous target gene fragment is 1.88 times the sequencing depth of the receptor 84K genome, which shows that the genome of the ONT-Cas:EXPA1-9 strain may have two insertions.
[0064] Application Example 2
[0065] Identification of the insertion site of the exogenous target gene fragment into the receptor genome
[0066] (1) Identification of the insertion site of the exogenous target gene fragment into the receptor genome
[0067] In the windows system, the IGV (Integrative Genomics Viewer) software is used to import the bam file of the sequencing fragment and the reference genome. Then, the exogenous target gene fragment (such as SEQ ID No. 1, Cas_EXPA1, 7161bp) is selected as a contig, and the visualization result of the sequencing reads sequence alignment with the exogenous target gene fragment is displayed Figure 3 In the IGV view, clicking the reads with Clipped (the end of the reads is marked with a red line) will pop up the information about the alignment of the reads with the exogenous target gene fragment and the alignment with the receptor 84K genome, which can provide information for identifying the insertion site of the exogenous target gene fragment into the receptor 84K genome. According to this information, the position of the matereads is input in the position input box of IGV, and the accurate insertion site of the exogenous target gene fragment can be manually confirmed near the position.
[0068] The work efficiency will be very low by manually extracting the alignment information of reads on the IGV interface. Therefore, according to the regularity and alignment information of mate pair alignment in the sam / bam file (definition see Appendix 1: 6), in the linux system, using the self-compiled perl script (Appendix 2: 1, main flow of recognition of exogenous insertion site and 2, recognition of reads related to insertion site and extraction of alignment information), the following reads are recognized and the related information is extracted: the reads split mapped (including soft and hard clipped) are recognized from the CIGAR information aligned to the exogenous target gene fragment, that is, part of the reads are aligned to the exogenous target gene fragment, and the other part is aligned to the 84K genome; and the alignment position of the part aligned to the 84K genome is found from the SA information recorded in the bam file as a candidate insertion position, combined with the condition of multiple such reads and the IGV diagram information, to judge the possible insertion position.
[0069] Taking ONT-Cas:EXPA1-9 strain as an example, according to the distribution of clipped reads, the basic information of the above-mentioned read type is extracted, as shown in Table 1. The exogenous target gene fragment may be inserted between 11,046,294 and 11,086,965 on the Chr06B chromosome of ONT-Cas:EXPA1-9 strain and between 11,413,364 and 11,456,478 on the Chr19A chromosome. As can be seen from Table 2, the 6 reads aligned to the end of the recipient genome are concentrated at 11,069,945 and 11,070,446 on the Chr06B chromosome, and the 7 reads aligned to the start of the recipient genome are concentrated at 11,070,450 and 11,070,452. According to the physical position of the chromosome, it is speculated that the insertion site of the exogenous target gene on the Chr06B chromosome may be Chr06B:11,070,447-11,070,449. Similarly, the insertion site of the exogenous target gene on the Chr19A chromosome may be Chr19A:11,432,810-11,432,831. In order to judge the accuracy, verification needs to be carried out on the corresponding position of the recipient 84K genome.
[0070] Table 1 Basic situation of reads aligned to the recipient genome recognized at both ends of the exogenous target gene fragment in ONT-Cas:EXPA1-9 strain
[0071]
[0072]
[0073]
[0074] Note: NO.: Number of reads; readID: ID code of reads; split mapped: type of soft-clip and hard-clip reads; Tdna: name of exogenous target gene insert; Tdna Len: length of exogenous target gene insert; Tdna Start: start position of read alignment to exogenous target gene fragment; Tdna End: end position of read alignment to exogenous target gene fragment; Aln2nd Chr: chromosome of read alignment to recipient 84K genome; Aln2nd Start: start site of read alignment to recipient chromosome; Aln2nd End: end site of read alignment to recipient chromosome.
[0075] (2) Verification of insertion site on recipient genome
[0076] According to the information of read alignment to the end of exogenous target gene fragment, the position of matereads was input in the position input box of IGV software, and the corresponding alignment site was searched. If there is an obvious gap or depth reduction near the site, it can be judged as the start site or interval of exogenous target gene fragment insertion. ONT-Cas:EXPA1-9 strain has a 3bp small gap at the position of 11,070,447-11,070,449 of chromosome Chr06B, and also shows an insertion of 6186bp at the position. ONT-Cas:EXPA1-12 strain does not have a deletion at the corresponding position, indicating that there is a cleavage and insertion of exogenous target gene at the site (A). There is also an obvious gap at the position of 11,432,810-11,432,831 of Chr19A, and ONT-Cas:EXPA1-12 strain does not have a deletion at the corresponding position, indicating that this region is the insertion position of exogenous target gene fragment (B). Figure 4 Figure 4
[0077] In summary, there are two different insertion sites in Cas:EXPA1-9 strain, which are located at Chr06B:11,070,447-11,070,449 and Chr19A:11,432,810-11,432,831, respectively.
[0078] (3) Identification of exogenous target gene insertion site in ONT-Cas:EXPA1-12 strain
[0079] Based on the distribution of clipped reads in the ONT-Cas:EXPA1-12 strain, basic information of the identified reads was extracted, as shown in Table 2. The table shows that these reads mainly aligned to positions 16,549,663-16,583,927 of the receptor 84K genome Chr04B. Nearly half of the reads aligned to the start position of the receptor 84K genome at 16,562,672, which is larger than the end position of other reads at 16,562,664. IGV visualization also shows gaps of varying lengths between 16,562,594 and 16,562,689, but the region with the most gaps is 16,562,665-16,562,671. Figure 5 Therefore, it can be determined that there is only one insertion site for the exogenous target gene fragment in the ONT-Cas:EXPA1-12 strain, namely Chr04B:16,562,665-16,562,671.
[0080] Table 2. Basic information on the alignment of reads identified at both ends of the exogenous target fragment with the recipient genome in the ONT-Cas:EXPA1-12 strain.
[0081]
[0082]
[0083] Note: NO.: Read number; readID: Read ID; split mapped: Type of soft-clip and hard-clip reads; Tdna: Name of the foreign target gene insertion fragment; Tdna Len: Length of the foreign target gene insertion fragment; Tdna Start: Start position of the read alignment to the foreign target gene fragment; Tdna End: End position of the read alignment to the foreign target gene fragment; Aln2nd Chr: Read alignment to the recipient 84K chromosome; Aln2ndStart: Start site of the read alignment to the recipient chromosome; Aln2nd End: End site of the read alignment to the recipient chromosome.
[0084] Application Example 3
[0085] The insertion site of the exogenous target gene fragment in different transgenic lines is located at the coding gene position.
[0086] By identifying the exogenous target gene fragment insertion site of the above PagEXPA1 gene editing strains ONT-Cas_EXPA1-9 and ONT-Cas_EXPA1-12, using self-programming (Appendix 2: 3, extraction of genes and flanking sequences near the insertion site), the genes within 5000 bp upstream and downstream of the insertion site were extracted as candidate genes to determine the relationship between the insertion site position and the coding gene, so as to predict whether the insertion of the exogenous target gene fragment has an unintended impact on the coding gene. As can be seen from Table 3, there are two insertion sites in the ONT-Cas_EXPA1-9 strain, the Chr06B:11,070,447-11,070,449 insertion site is located 1664 bp upstream of the coding gene Pag.B06G001223.1, and the Chr19A:11,432,810-11,432,831 insertion site is located 62 bp upstream of the coding gene Pag.A19G000860.1, the transcription direction of the two insertion sites is the same as that of the recipient genome, but opposite to that of the coding gene. There is only one insertion site in the ONT-Cas:EXPA1-12 strain, which is Chr04B:16,562,665-16,562,671, located 439 bp downstream of the coding gene Pag.B04G001464.1, and the direction of the exogenous target gene fragment is consistent with that of the recipient genome and the transcript.
[0087] Table 3 Site information of exogenous target fragment inserted into the recipient genome in different transgenic strains
[0088]
[0089]
[0090] Note: ID: ID code of reads; TDNA: name of exogenous target gene insertion fragment; Chr: chromosome of the recipient 84K genome to which the exogenous target gene fragment insertion site is aligned; Start: start site of the recipient chromosome to which the exogenous target gene fragment insertion site is aligned; End: end site of the recipient chromosome to which the exogenous target gene fragment insertion site is aligned; Len: length of the fragment deleted by the exogenous target gene fragment insertion site; Ori: direction of the insertion of the exogenous target gene fragment into the recipient genome, (+) indicates that the transcription direction of the inserted exogenous target gene fragment is consistent with the direction of the recipient genome, (-) indicates that the transcription direction of the inserted exogenous target gene fragment is opposite to the direction of the recipient genome; geneRegion: position of the exogenous target gene fragment insertion site in a coding gene, + / - indicates that the transcription direction of the coding gene is consistent / contrary to the direction of the genome, left indicates that the insertion site is located at the 5' end of the coding gene genome position, and right indicates that the insertion site is located at the 3' end of the coding gene genome position.
[0091] In summary, the insertion site and copy number of the exogenous target gene fragment inserted into the recipient genome of the poplar can be identified by high-throughput third-generation sequencing. Meanwhile, the insertion site can be accurately located to a specific chromosome and coding gene position, thereby providing information support for whether the insertion of the exogenous target gene fragment has an unintended effect on the coding gene.
[0092] The present application is not limited to the above-mentioned embodiments, and the present application will cover the range described in the technical solutions, and various modifications and equivalent changes within the scope of the claims, and any modification or improvement of the present application made by those skilled in the art without departing from the technical solutions of the present application shall fall within the scope of the present application.
[0093] Appendix 1: Related term explanation
[0094] 1. Exogenous vector: part of the expression vector sequence introduced by the transgene, including the part of the vector inserted into the plant (LBT-DNA to RB T-DNA).
[0095] 2. Target gene: exogenous target gene sequence introduced by the transgene.
[0096] 3. Exogenous target gene fragment: sequence of the expression vector and the target gene introduced by the transgene.
[0097] 4. Recipient 84K genome: sequence of the genome of the transgenic recipient 84K.
[0098] 5. Reference genome: sequence of the exogenous target gene fragment and the recipient 84K genome.
[0099] 6. Definition of sam / bam files:
[0100] SAM (Sequence Alignment / Map) format is a standard text format used for storing alignment results of high-throughput sequencing data. It was developed by Heng Li et al. to provide a unified and detailed representation of alignments between short sequence reads and a reference genome. SAM format is widely used in bioinformatics, especially in processing and analyzing high-throughput sequencing data. Here are some key features of SAM format:
[0101] 1) Header: The beginning part of a SAM file contains a header, which includes metadata about the reference sequence, program version, library information, etc. The header starts with '@' character.
[0102] 2) Alignment Section: Following the header is a series of alignment sections, each representing an alignment result of a read to the reference genome.
[0103] 3) Column Format: Each alignment section consists of 11 basic fields separated by tabs, which contain basic information about the read, such as QNAME (Query Template Name), FLAG (Flag, indicating alignment-related information), RNAME (Reference Sequence Name), POS (Position), MAPQ (Mapping Quality), CIGAR (CIGAR String), RNEXT (Next Reference Sequence Name), PNEXT (Paired Read Position on Reference), TLEN (Template Length), SEQ (Read Sequence), QUAL (Alignment Quality).
[0104] 4) Optional Fields: In addition to the 11 basic fields, SAM format supports optional fields, which provide additional information such as pairing information, quality scores, raw sequencing data, etc.
[0105] 5) CIGAR String: CIGAR string is a compact representation of the alignment relationship between reads and reference sequences, including matches, mismatches, insertions, deletions, and soft / hard clips, etc.
[0106] Appendix 2: Perl program code script for foreign insertion site identification and flanking sequence extraction
[0107] 1. Main process of foreign insertion site identification
[0108]
[0109]
[0110]
[0111] 2. Extraction of identification and alignment information of reads associated with insertion sites
[0112]
[0113]
[0114]
[0115]
[0116] 3. Extraction of genes and flanking sequences near insertion sites
[0117]
[0118]
[0119]
[0120]
[0121]
Claims
1. A method for identifying the insertion site and copy number of exogenous gene fragments in transgenic plants based on high-throughput third-generation sequencing, characterized in that, The identification method includes the following steps: (1) Extract DNA from the plant sample to be tested; (2) High-throughput third-generation ONT library preparation and sequencing; (3) Integrate exogenous target gene fragments and receptor genomes as reference genomes; (4) Sequence alignment and identification of sequencing fragments with the reference genome; 1) Identify the insertion frequency of exogenous gene fragments; 2) Identify the insertion site of the exogenous target gene fragment into the recipient genome; 3) Identify the location of the exogenous target gene fragment insertion site in the recipient genome that encodes the gene.
2. The identification method according to claim 1, characterized in that, The plants mentioned include poplar trees.
3. The identification method according to claim 1, characterized in that, The specific steps for identifying the insertion frequency of exogenous gene fragments in step 1) are as follows: Sequencing depths of sequencing reads aligned to exogenous target gene fragments in the reference genome sequence and average sequencing depth of the receptor 84K genome were statistically analyzed to identify the insertion frequency of exogenous gene fragments. When the alignment and sequencing depth of the exogenous target gene fragment is ≤0.3, it indicates that there is no insertion of the exogenous target gene fragment on the recipient chromosome. When the sequencing depth of the exogenous target gene fragment is 0.7 to 1.3 times that of the average sequencing depth of the recipient genome, it indicates that an insertion of the exogenous target gene fragment has occurred on the recipient chromosome. When the sequencing depth of the exogenous target gene fragment is 1.7 to 2.3 times that of the average sequencing depth of the recipient genome, it indicates that two insertions of the exogenous target gene fragment have occurred on the recipient chromosome.
4. The identification method according to claim 1, characterized in that, The specific procedures for identifying the insertion site of the exogenous target gene fragment into the recipient genome in step 2) are as follows: (1) Using IGV software, by comparing the alignment information of reads that are clipped to the end of the exogenous target gene fragment, reference information is provided for identifying the location of the exogenous target gene fragment inserted into the 84K genome of the receptor. (2) Using the Perl script based on the alignment information stored in the sam / bam file, identify and extract the reads that are aligned to the end of the exogenous target gene fragment and have been cut, and determine the insertion site of the exogenous target gene fragment into the recipient genome based on the extracted alignment information. (3) Based on the insertion site information of the recipient genome obtained in (2), the corresponding insertion site is searched in the recipient genome using IGV software. If there is an obvious gap or a decrease in depth near the site, it can be determined as the starting site or interval of the insertion of the exogenous target gene fragment.
5. The identification method according to claim 1, characterized in that, The specific steps for step 3) of identifying the location of the exogenous target gene fragment insertion site within the coding gene of the recipient genome are as follows: Based on the identification of the insertion site of the exogenous target gene fragment as described in claim 1, coding genes within 5000 bp upstream and downstream of the insertion site are extracted as candidate genes to determine the relationship between the insertion site and the neighboring coding genes on the recipient genome; and the insertion site is located upstream or downstream of the coding gene by the direction of insertion of the exogenous target gene fragment into the recipient genome and the transcription direction of the recipient genome and the coding gene.
6. The application of the identification method according to any one of claims 1 to 5 in assessing the safety of transgenic plants.
7. The application of the identification method according to any one of claims 1 to 5 in the investigation of gene expression and function.
8. The application of the identification method according to any one of claims 1 to 5 in screening for stable genetically inherited transgenic lines.
9. The identification method according to any one of claims 1 to 5 is used in distinguishing different genetically modified products by detecting the insertion site of exogenous DNA.
Citation Information
Patent Citations
Method and device for detecting exogenous insertion information based on third-generation gene sequencing data
CN115620810A
Whole genome structure variation identification method based on three-generation sequencing
CN115831222A
Method for efficiently identifying exogenous DNA insertion site in genome by using high-throughput sequencing technology
CN119876357A