Identification and application of insertion information of exogenous gene of transgenic plant by next-generation sequencing
By comparing high-throughput second-generation sequencing technology with transgenic poplar plants, the problems of high cost and complex operation of third-generation sequencing were solved. This enabled high-precision identification of foreign gene insertion sites and copy numbers in transgenic plants, supporting biosafety assessment and gene function analysis.
Patent Information
- Application Number
- CN202511144576.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies for detecting foreign gene insertion sites in transgenic plants suffer from high costs, high error rates, difficult data processing, and low throughput with third-generation sequencing, while traditional methods are complex to operate and prone to false positives, making it difficult to achieve efficient and accurate identification of insertion sites and copy numbers.
High-throughput next-generation sequencing technology was used to sequence transgenic poplar plants, and the sequences were compared with the exogenous target gene fragment and the genome of the recipient, *Populus simonii* 84K. The insertion frequency and location of the exogenous gene fragment were identified by comparing sequencing depth and the comparison results. The insertion sites were analyzed using IGV software and Perl scripts.
It achieves high-precision and low-cost identification of exogenous gene insertion sites and copy numbers, and can quickly and accurately determine the insertion frequency and location of exogenous genes in the recipient genome, supporting the biosafety assessment and gene expression function analysis of transgenic plants.
Smart Images

Figure CN120966972A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of plant transgenics, and particularly relates to identification and application of exogenous gene insertion information of transgenic plants by second-generation sequencing. BACKGROUND
[0002] Plant transgenic technology is of great significance to crop and forest tree breeding. This technology not only breaks the species barrier and introduces excellent genes from different sources, greatly expanding the gene resource library. Compared with traditional breeding, it can also precisely improve specific traits. Plant transgenic technology can quickly obtain ideal plants by directly introducing target genes, significantly shorten the breeding cycle, and accelerate the promotion of new varieties. At the same time, the cultivation of transgenic varieties with insect resistance, disease resistance, stress resistance and herbicide tolerance can maintain growth and yield under the condition of reducing pesticide and fertilizer input, reduce cost and environmental pollution, improve economic benefit and resource utilization efficiency, and help protect the ecology and biodiversity.
[0003] To ensure the safety of commercial application of transgenic plants, biological safety detection is crucial, and the detection of exogenous gene insertion site is a key link. Traditional detection methods such as inverse PCR, adapter PCR and TAIL-PCR obtain flanking sequences through special strategies, but the operation is complex, time-consuming and prone to false positives. Southern blotting can detect copy number and approximate location, but it requires high-quality samples, is cumbersome to operate, and provides limited information. Currently, whole-genome resequencing is a more ideal solution. By high-throughput sequencing of transgenic plants and comparing with the wild-type genome, the precise insertion site, copy number and surrounding sequence changes of the exogenous gene can be analyzed by bioinformatics. Third-generation sequencing technology can more accurately analyze the insertion site, integration mode and structure due to its long read length advantage. However, third-generation sequencing has problems such as high cost, relatively high error rate, difficult data processing, and relatively low throughput, which will greatly increase the cost of using third-generation sequencing methods to identify the insertion site and copy number of the exogenous target gene in transgenic materials. SUMMARY
[0004] To solve the above problems, the present application uses transgenic different strains of poplar plants as materials for high-throughput second-generation sequencing. By comparing the sequencing depth of the alignment and the exogenous alignment results with the exogenous target gene fragment and the genome of the receptor Populus alba x P. glandulosa 84K, the frequency and location of the insertion of the exogenous target gene fragment on the receptor chromosome are obtained. Not only does this effectively avoid the defects of high cost and increased technical difficulty of third-generation sequencing, but it also has the advantages of high sequencing accuracy, high throughput and mature data analysis methods, making it an ideal method for large-scale identification of the insertion site and copy number of the exogenous target gene in transgenic materials.
[0005] To achieve the above purpose, the present application provides the following technical solutions:
[0006] The application provides a method for identifying an exogenous gene insertion site of a transgenic plant based on second-generation sequencing, which comprises the following steps:
[0007] (1) extracting DNA of a plant sample to be tested;
[0008] (2) high-throughput second-generation novaseq 6000 library sequencing, and filtering original sequencing data;
[0009] (3) integrating an exogenous target gene fragment and a receptor genome as a reference genome;
[0010] (4) sequence alignment and identification of filtered sequencing fragments and the integrated reference genome;
[0011] 1) identifying the frequency of insertion of an exogenous target gene fragment into a receptor chromosome;
[0012] 2) identifying the site of insertion of an exogenous target gene fragment into a receptor genome;
[0013] 3) identifying the position of an insertion site of an exogenous target gene fragment in a coding gene in a receptor genome.
[0014] In the application, the plant in step (1) is a poplar.
[0015] In the application, the main parameters for filtering the original sequencing data in step (2) are: -q 20; -u 40; -n 15; -l 50; -w 4; and the sample sequencing depth is 5X.
[0016] In the application, the specific operation of step 1) for identifying the frequency of insertion of an exogenous target gene fragment into a receptor chromosome is as follows:
[0017] The sequencing depth of the sequencing data aligned with the exogenous target gene fragment and the average sequencing depth of the receptor genome are counted respectively;
[0018] The calculation method of the average sequencing depth is: the number of bases aligned to the chromosome / the length of the chromosome;
[0019] When the aligned sequencing depth of the exogenous target gene fragment is ≤0.3 times, it indicates that there is no insertion of the exogenous target gene fragment into the receptor chromosome;
[0020] When the aligned sequencing depth of the exogenous target gene fragment is 0.7 times-1.3 times of the average sequencing depth of the receptor genome, it indicates that there is one insertion of the exogenous target gene fragment into the receptor chromosome;
[0021] When the alignment sequencing depth of the exogenous target gene fragment is 1.7 times to 2.3 times of the average sequencing depth of the receptor genome, it indicates that the insertion of the exogenous target gene fragment occurs twice on the receptor chromosome;
[0022] By analogy, the frequency of the insertion of the exogenous target gene fragment on the receptor chromosome can be determined;
[0023] In actual operation, when the alignment sequencing depth of the exogenous target gene fragment is greater than 0.3 times and less than 0.7 times, or greater than 1.3 times and less than 1.7 times, too many miscellaneous items are mixed, and the insertion condition needs to be analyzed in combination with the specific insertion site.
[0024] In the present application, the specific operation of the step 2) for identifying the site of the insertion of the exogenous target gene fragment into the receptor genome is as follows:
[0025] (1) The IGV software is used to provide reference information for identifying the position of the insertion of the exogenous target gene fragment into the receptor 84K genome by aligning the alignment information of the reads with different color marks at the 5' and 3' ends of the exogenous target gene fragment;
[0026] (2) According to the alignment rule of the mate pair, the perl script is used to identify the following two types of reads and extract relevant information according to the alignment information stored in the sam / bam file in the linux system:
[0027] 1) Type A: According to the RNAME information and RNEXT information in the bam file, the pair reads located at one end of the exogenous target gene fragment and at the other end of the receptor genome are identified, and the positions thereof are listed;
[0028] 2) Type B: The split mapped reads are identified from the CIGAR information, that is, the reads having a part aligned to the exogenous target gene fragment and another part aligned to the receptor genome; and the alignment position of the part aligned to the receptor genome is recorded from the Supplementary Alignments recorded in the bam file.
[0029] (3) According to the extracted alignment information, the insertion site of the exogenous target gene fragment into the receptor genome can be preliminarily determined, and the corresponding insertion point position in the receptor genome is searched by using the IGV software. If there is an obvious gap or depth reduction near the site, it can be determined as the starting site or interval of the insertion of the exogenous target gene fragment.
[0030] In the present application, the specific operation of the step 3) for identifying the position of the insertion site of the exogenous target gene fragment in the coding gene of the receptor genome is as follows:
[0031] The sequence within 2000bp upstream and downstream of the insertion site is extracted, the position of the insertion site in the coding gene is found through sequence alignment, so as to determine the relationship between the insertion site and the adjacent coding gene on the recipient genome; and the direction of the insertion of the exogenous target gene fragment into the recipient genome and the transcription direction of the recipient genome and the coding gene are determined, so as to determine whether the insertion site is upstream or downstream of the coding gene.
[0032] The application also provides application of the above-mentioned identification method in evaluation of biological safety of transgenic plants.
[0033] The application also provides application of the above-mentioned identification method in exploration of gene expression and function.
[0034] The application also provides application of the above-mentioned identification method in screening of transgenic products with stable heredity.
[0035] The application also provides application of the above-mentioned identification method in distinguishing different transgenic products by detecting the insertion site of the exogenous DNA.
[0036] Because the sequencing fragment of the second-generation sequencing is short, it brings great challenges in the aspects of later sequencing data alignment and insertion site analysis. In the application, the plants of different strains of transgenic poplar are taken as materials, high-throughput second-generation sequencing is carried out, the sequencing reads are aligned with the exogenous target gene fragment and the recipient Populus alba x P.glandulosa genome, the copy number of the exogenous target gene fragment inserted into the recipient genome is obtained by comparing the alignment sequencing depth, and the accurate position of the exogenous target gene inserted into the recipient genome is found through the end alignment information of the exogenous target gene fragment. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 It is a base quality and base composition distribution graph of data before and after filtering of high-throughput second-generation sequencing data; in the figure, A and B: non-transgenic control 84K; C and D: Cas:EXPA1-9 strain; E and F: Cas:EXPA1-12 strain. A, C, E: distribution graph before filtering of sequencing data; B, D, F: distribution graph after filtering of sequencing data. A: A base proportion; T: T base proportion; C: C base proportion; G: G base proportion; N: missing base proportion; Qmean: average base quality; Q20: proportion of reads with a quality value greater than Q20;
[0038] Figure 2 It is a sequencing depth comparison graph of sequencing reads of different transgenic strains with the exogenous target gene fragment and the recipient genome; in the figure, Genome represents the sequence of the recipient 84K genome; Cas9_EXPA1 represents the sequence of the exogenous target gene fragment;
[0039] Figure 3 Figure 2 is an IGV visualization diagram of the Mate pair alignment to the exogenous target gene fragment in the Cas:EXPA1-9 strain; Figure a: the non-transgenic control plant CK; Figure b: the Cas:EXPA1-9 strain. The red arrow points to a text box on the light green mate reads alignment, the other end of which is at a certain position on Chr19A; the blue arrow points to a text box on the grayish blue mate reads alignment, the other end of which is at a certain position on Chr06B;
[0040] Figure 4 Figure 3 is an IGV visualization diagram of the insertion site of the Mate pair alignment to the recipient genome in the Cas:EXPA1-9 strain; Figure a: the non-transgenic control plant CK; Figure b: the Cas:EXPA1-9 strain. The exogenous target gene fragment is aligned to the insertion site of A: chromosome Chr19A, and B: chromosome Chr06B.
[0041] Figure 5 Figure 4 is an IGV visualization diagram of the Mate pair alignment to the exogenous target gene fragment in the Cas:EXPA1-12 strain; Figure a: the non-transgenic control plant CK; Figure b: the Cas:EXPA1-12 strain.
[0042] Figure 6 Figure 5 is an IGV visualization diagram of the insertion site of the Mate pair alignment to the recipient genome in the Cas:EXPA1-12 strain; Figure a: the non-transgenic control plant CK; Figure b: the Cas:EXPA1-12 strain. The exogenous target gene fragment is aligned to the insertion site of A: chromosome Chr04B, and B: chromosome Chr12B: 12,445,705. DETAILED DESCRIPTION
[0043] In the following, if a specific technique or condition is not specified for the reagent or instrument used, it is performed according to the conventional experimental condition, and if the manufacturer's instruction is not specified, it is performed according to the condition suggested in the instruction. If the manufacturer of the reagent or instrument used is not specified, it is a conventional product that can be obtained by purchase.
[0044] In the present application, because the introduced fragments of gene editing are the same, the sequence of the exogenous target gene fragment in the Cas:EXPA1-9 and Cas:EXPA1-12 gene editing strains is consistent.
[0045] EXAMPLE
[0046] 1 DNA extraction and library sequencing
[0047] (1) Test sample material
[0048] The method takes pagEXPA1 gene edited aspen strains (Cas:EXPA1-9 and Cas:EXPA1-12) as test objects, and non-transgenic aspen 84K plants (CK) as control samples. The leaf DNA of each strain plant is extracted, and qualified by Nanodrop or Qubit detection, that is, the DNA concentration is higher than 150 ng / μL, OD 260 / 280 Between 1.8 and 2.2, OD 260 / 230 Between 1.8 and 2.5.
[0049] (2) Library sequencing
[0050] The sample genomic DNA is fragmented by mechanical disruption such as ultrasonic wave, and then the fragmented DNA is purified, end-repaired, 3' end A-tailed, and ligated with sequencing adapters. Then, the fragment size is selected by agarose gel electrophoresis, PCR amplified to form a sequencing library, and the constructed library is first subjected to library quality inspection. The qualified library is subjected to sequencing. The PE150 double-end sequencing method of the second-generation illumina novaseq 6000 platform is used, and the sequencing sample amount of each sample is about 5Gbp. The A-B two sets of genomes of aspen 84K (about 890M in size) are used as reference genomes, and the average sequencing depth of the two pagEXPA1 gene edited aspen strains is about 5x. The raw image data file obtained by high-throughput sequencing is converted into raw sequencing sequences (Raw Data) by base recognition analysis, and the results are stored in FASTQ (abbreviated as fq) file format, which contains sequence information of sequencing sequences (Reads) and corresponding sequencing quality information.
[0051] 2. Sequencing data filtering
[0052] Because there may be some reads with low quality values or high adapter ratios in the data, these reads may bring some false alignments, so they need to be removed or trimmed. The raw data obtained by sequencing is filtered by fastp software (v 0.23.4), and the main parameters for filtering are (-q 20-u 40-n 15-l 50-w 4). Other parameters use default values, and reads that do not meet the conditions will be filtered out as low-quality reads. Among them:
[0053] -q: set the threshold standard for low-quality. In this method, 20 is used, that is, bases with quality values less than 20 are considered to be low-quality, which is commonly referred to as Q20;
[0054] -u: percentage of low-quality bases. In this method, the default value 40 is used, which means 40%, that is, if a 150bp read contains more than 60 low-quality bases, it will be discarded, and as soon as one read does not meet the condition, it will be discarded.
[0055] -n: filter reads with too many N bases, if the N base content is greater than n, this read will be discarded in pairs. This method uses 15, that is, more than 10% of the 150bp read is removed;
[0056] -l: after adaptor trimming, N base trimming, low quality base trimming, reads with a length less than this length are discarded. This method uses 50, that is, fragments less than 50bp will be filtered out;
[0057] -w: the number of threads used when analyzing, does not affect the result, only affects the speed of quality control.
[0058] In order to intuitively show the quality of sequencing data, after the completion of raw data filtering, the modified version of fqcheck software (https: / / github.com / hewm2008 / fqcheck, v2.08) is used to count the content of A, T, C, G and N of each sequencing cycle before and after filtering, as well as the base quality value, and the R software (https: / / www.r-project.org / , v4.3) Ggplot2 package (https: / / github.com / tidyverse / ggplot2, v3.5.1) is used to visualize the statistical results. Taking the pagEXPA1 gene edited 2 Populus alba × P. tremula strains and non-transgenic Populus alba × P. tremula 84K plant samples as examples, the base quality and base composition distribution of the second generation resequencing data before and after filtering are shown. Figure 1 According to the information in the figure, it can be judged that the proportion of N content in the filtered data is very low, almost 0; the average base quality value is about 40, except for the first 10 bp of the read, A-T and C-G are obviously complementary, and more than 95% of the base quality values are above 20. It shows that the data quality value is high, and subsequent analysis can be carried out.
[0059] 3Integration of reference genome
[0060] In the linux environment, the cat command is used to integrate the exogenous target gene fragment sequence (definition see Appendix 1:3; such as SEQ ID No. 1, Cas9_EXPA1, 7161bp) containing the exogenous vector (definition see Appendix 1:1) and the target gene (definition see Appendix 1:2) into the recipient 84K genome (definition see Appendix 1:4) sequence file (https: / / figshare.com / articles / dataset / 84K_genome_zip / 12369209) to obtain the merged sequence file, which is used as the reference genome sequence (definition see Appendix 1:5) for subsequent sequencing sequence alignment analysis.
[0061] 4 Alignment of sequencing fragments to reference genome
[0062] To identify the insertion position of the exogenous gene fragment using the position of the related reads on the genome, the filtered sequencing reads were aligned to the integrated reference genome sequence in a Linux system using the bwa software (v0.7.18). The main process and steps of the alignment are as follows:
[0063] 1) The index function of the bwa software was used to index the reference sequence, and the filtered read sequence was aligned to the indexed reference genome sequence by the mem function;
[0064] 2) The fixmate function of the samtools software (v1.19) was used to repair the mate pair information, and the bam file saving the alignment information was output. The sort function was used to sort the sequence according to the physical position, and the index function was used to index the sorted bam file;
[0065] 3) The markdup function of the samtools software was used to mark the PCR and optical repeated reads, and then the view function of the samtools software was used to filter out the reads marked as repeated to obtain the final alignment result.
[0066] Application Example 1
[0067] Identification of the copy number of the insertion of the exogenous gene fragment into the recipient genome
[0068] In a Linux system, the Pandepth software (v2.25) was used to count the sequencing depth of the sequencing reads aligned to the exogenous gene fragment and the average sequencing depth of the 84K recipient genome according to the alignment results after removing the PCR and optical repeats.
[0069] For each sequence, the average sequencing depth calculation formula is: the number of bases aligned to the chromosome / the length of the chromosome. By comparing the alignment sequencing depth of the exogenous gene fragment and the average sequencing depth of the 84K recipient genome, the frequency of the insertion of the exogenous gene fragment into the recipient genome, i.e., the copy number of the exogenous gene, can be determined:
[0070] 1) If the alignment sequencing depth of the exogenous gene fragment is close to 0 (usually ≤0.3), it indicates that there is no insertion of the exogenous gene fragment into the recipient chromosome;
[0071] 2) When the alignment sequencing depth of the exogenous target gene fragment is 0.7-1.3 times of the average sequencing depth of the receptor genome, it indicates that one insertion of the exogenous target gene fragment occurs on the receptor chromosome;
[0072] 3) When the alignment sequencing depth of the exogenous target gene fragment is between 1.7-2.3 times of the average sequencing depth of the receptor genome, it indicates that two insertions of the exogenous target gene fragment occur on the receptor chromosome.
[0073] By analogy, the frequency of insertion of the exogenous target gene fragment on the receptor chromosome can be determined.
[0074] The present application compares the sequencing depth of Cas:EXPA1-9 and Cas:EXPA1-12 with that of the non-transgenic control plant respectively. Figure 2 ).
[0075] It can be seen that:
[0076] (1) The alignment sequencing depth of the exogenous target gene fragment (Cas9_EXPA1) on the non-transgenic control 84K is almost 0, indicating that there is no insertion of the exogenous target gene fragment on its chromosome.
[0077] (2) The alignment sequencing depth of the exogenous target gene fragment (Cas9_EXPA1) on Cas:EXPA1-12 is 0.88 times of the receptor 84K genome (Genome), which can be inferred that there is one insertion of the exogenous target gene in the genome of Cas:EXPA1-12 strain.
[0078] (3) The alignment sequencing depth of the exogenous target gene fragment on Cas:EXPA1-9 is 2.62 times of the sequencing depth of the receptor genome, indicating that there may be at least two insertions in the genome of Cas:EXPA1-9 strain.
[0079] Application Example 2
[0080] Identification of the insertion site of the exogenous target gene fragment into the receptor genome
[0081] Since there may be two insertion sites of the exogenous gene in the Cas:EXPA1-9 transgenic strain, we take it as an example to illustrate the identification process of the insertion site of the exogenous target gene fragment into the receptor genome.
[0082] (1) Extraction of end alignment information of the exogenous target gene fragment
[0083] Under Windows, the BAM file used for aligning the sequencing reads with the reference genome in Example 4 was imported into IGV (Integrative Genomics Viewer) software. Next, the exogenous target gene fragment (as shown in SEQ ID No. 1, Cas_EXPA1, 7161bp) was selected as a contig, and the visualization results of the sequence alignment of the sequencing reads with the exogenous target gene fragment were displayed. From... Figure 3 As can be seen, almost no sequencing reads in the non-transgenic control plant (CK) matched the exogenous target gene fragment. Figure 3 a), while Cas:EXPA1-9 contains a large number of sequencing reads that can be matched ( Figure 3 (b) This indicates that the Cas:EXPA1-9 strain contains the insertion of a foreign target gene. According to the IGV coloring rules for mate pair alignments, reads aligned to different chromosomes are marked with different colors. Cas:EXPA1-9 shows reads of different colors at different positions along the entire foreign target fragment, indicating that some reads may align to both the foreign target gene fragment sequence and the recipient 84K genome sequence. However, the primary objective of this invention is to identify the insertion site of the foreign target gene fragment into the recipient genome; therefore, we only need to focus on the mate pair at the 5' and 3' ends of the foreign target gene fragment. Figure 3 As shown by the red arrow, the other end of the light green reads on Cas_EXPA1 aligns to a position in Chr19A; as shown by the blue arrow, the purple reads align to a position in Chr06B. Based on this information, in the IGV position input box, enter the location of the mate reads aligned to the recipient genome, and manually confirm the accurate insertion position near that location.
[0084] Manually extracting read alignment information from the IGV interface would be inefficient. Therefore, based on the rules of mate pair alignment, we used a self-written Perl script on a Linux system (see Appendix 1: 1. Main process for identifying exogenous insertion sites and 2. Identification and extraction of alignment information for reads related to insertion sites) to identify the following two types of reads and extract relevant information based on the alignment information stored in the sam / bam file:
[0085] (1) Type A: Based on the RNAME and RNEXT information in the bam file, identify pair reads with one end located in the exogenous target gene fragment and the other end located in the recipient genome, and list their positions;
[0086] (2) Type B: reads that are split mapped, i.e. part of the read is mapped to the exogenous gene fragment and the other part is mapped to the recipient genome; and the mapping position of the part mapped to the recipient genome is recorded from Supplementary Alignments recorded in the bam file.
[0087] The basic information of the above reads types recognized by Cas:EXPA1-9 strain was extracted, as shown in Table 1. However, since the purpose of the present application is to identify the insertion site of the exogenous gene fragment into the recipient genome, and the sequencing fragment length of the second-generation resequencing is about 150 bp, only the reads that can be aligned to the 5' end and 3' end of the exogenous gene fragment within about 150 bp are selected for information collection (Table 2). As can be seen from the information in Table 2, the extracted reads types include Type A and Type B, that is, one end of these reads is mainly aligned to the exogenous gene fragment, and the other end is aligned to the recipient 84K genome and mainly aligned to the vicinity of Chr19A:114322809 and Chr06B:11070450, indicating that the exogenous gene fragment may have two different insertion points, respectively located on Chr19A and Chr06B chromosomes. From the information of the reads aligned to Chr19A, it can be seen that the 5' end and 3' end of the exogenous gene fragment can be aligned, indicating that the insertion site of the exogenous gene fragment can be found by the double-end alignment information of the exogenous gene fragment, and the reads aligned to Chr06B only find alignment information at the 3' end of the exogenous gene fragment, so it is necessary to verify the corresponding position on the recipient genome.
[0088] Table 1 Basic information of different reads types recognized on the exogenous gene fragment in Cas:EXPA1-9 strain
[0089]
[0090]
[0091] Note: readID: ID code of reads; Type: type of reads, type A is type A reads, type B is type B reads; Tdna: name of exogenous target gene fragment; TdnaStart: start position of read alignment to exogenous target gene fragment; TdnaEnd: end position of read alignment to exogenous target gene fragment; TdnaCigar: CIGAR information in alignment result file (bam file), wherein xM represents x bases aligned to exogenous target gene fragment, and yS represents y bases aligned to receptor genome; MateChr: chromosome position of read alignment to receptor genome; MateStart: start position of read alignment to receptor chromosome; Aln2nd Chr: receptor chromosome to which the folded part of read is aligned if it belongs to type B; Aln2ndStart: start position of the folded part of read alignment to receptor chromosome if it belongs to type B; Aln2ndEnd: end position of the folded part of read alignment to receptor chromosome if it belongs to type B; Aln2ndCigar: CIGAR information of the folded part of read alignment if it belongs to type B, wherein mM represents x bases aligned to the region of receptor chromosome, and nS represents y bases aligned to other regions of receptor chromosome or exogenous target gene fragment.
[0092] Table 2 Basic situation of different types of reads recognized at both ends of exogenous target fragment in Cas:EXPA1-9 strain
[0093]
[0094]
[0095] Note: readID: ID code of the read; Type: type of the read, typeA is type A reads, typeB is type B reads; Tdna: name of the exogenous target gene fragment; TdnaStart: start position of the read alignment to the exogenous target gene fragment; TdnaEnd: end position of the read alignment to the exogenous target gene fragment; TdnaCigar: CIGAR information in the alignment result file (bam file), where xM represents x bases aligned to the exogenous target gene fragment, and yS represents y bases aligned to the recipient genome; MateChr: chromosomal position of the read alignment to the recipient genome; MateStart: start position of the read alignment to the recipient chromosome; Aln2nd Chr: If the read belongs to type B, the folded portion is aligned to the recipient chromosome; Aln2ndStart: If the read belongs to type B, the folded portion is aligned to the start site of the recipient chromosome; Aln2ndEnd: If the read belongs to type B, the folded portion is aligned to the end site of the recipient chromosome; Aln2ndCigar: If the read belongs to type B, the CIGAR information of the folded portion, where mM represents x bases aligned to the recipient chromosome in that region, and nS represents y bases aligned to other regions of the recipient chromosome or exogenous target gene fragments.
[0096] (2) Identification of insertion sites on the receptor genome
[0097] Based on the information from the end-aligned reads of the exogenous target gene fragment, the location of the mate reads is entered in the location input box of the IGV software to search for the corresponding alignment site. A clear cut or decreased depth near this site indicates the insertion site or region of the exogenous target gene fragment. The results showed a clear notch at positions 11,432,810-11,432,943 on chromosome Chr19A, while the non-transgenic control CK did not have a notch at this position. This indicates that exogenous target gene splicing and insertion occurred in this region, making it an insertion site for the exogenous target gene. Figure 4 A). There is also a noticeable small notch at positions 11,070,447-11,070,449 of Chr06B, indicating that this position is also an insertion site for a foreign target gene. Figure 4 B).
[0098] Therefore, there are two insertion sites for exogenous target gene fragments in the Cas:EXPA1-9 strain, located at Chr19A:11,432,810-11,432,943 and Chr06B:11,070,447-11,070,449, respectively.
[0099] (3) Identification of the insertion site of the Cas:EXPA1-12 strain
[0100] Following the method used to locate insertion sites in the Cas:EXPA1-9 line, the insertion sites in the Cas:EXPA1-12 line were located. Similarly, under a Windows system, IGV software was used to import the BAM files of the non-transgenic CK and Cas:EXPA1-12 sequencing reads aligned with the reference genome. First, the exogenous target gene fragment (Cas_EXPA1, 7161bp) was selected as a contig, and the alignment results of the Cas:EXPA1-12 sequencing reads were displayed. Figure 5 In the non-transgenic control (CK) plants, almost no sequencing reads matched the exogenous target gene fragment. Figure 5 a), while Cas:EXPA1-12( Figure 5 b) shows a large number of sequencing reads that can be matched, and there are mate pairs of different colors at the 5' and 3' ends, indicating that there is an insertion of a foreign target gene in the Cas:EXPA1-12 line.
[0101] Using a self-developed Perl script (Appendix 2: 1. Main workflow for identifying exogenous insertion sites and 2. Extraction of identification and alignment information of reads related to insertion sites), basic information of reads with Type A or Type B characteristics in the Cas:EXPA1-12 strain was extracted. All extracted reads contained both Type A and Type B types and were primarily located at the 5' and 3' ends of the exogenous fragment, as shown in Table 3. The sites of Cas:EXPA1-12 aligned to the recipient genome were mainly concentrated on chromosomes Chr04B and Chr12B. A relatively large number of reads were found near chromosomes 16,562,664 and 16,562,672 on Chr04B, indicating a possible exogenous fragment insertion at these locations. Fragments near chromosomes 12,445,705 on Chr12B were fewer, suggesting that these are likely not insertion sites. Subsequent verification at the corresponding location on the receptor genome demonstrated that Cas:EXPA1-12 has a gap at 16,562,665-16,562,671 of Chr04B. Figure 6 A), while there is no gap or depth reduction near 12,445,705 in Chr12B ( Figure 6 B) indicates that Cas:EXPA1-12 has only one insertion site for the exogenous target gene, which is Chr04B:16,562,665-16,562,671.
[0102] Table 3 Basic information of different types of reads recognized at both ends of the exogenous target fragment in Cas:EXPA1-12 strains
[0103]
[0104]
[0105] Note: readID: ID code of reads; Type: type of reads, typeA is type A reads, typeB is type B reads; Tdna: name of exogenous target gene fragment; TdnaStart: start position of read alignment to the exogenous target gene fragment; TdnaEnd: end position of read alignment to the exogenous target gene fragment; TdnaCigar: CIGAR information in the alignment result file (bam file), wherein xM represents x bases aligned to the exogenous target gene fragment, and yS represents y bases aligned to the receptor genome; Mate Chr: chromosome position of read alignment to the receptor genome; Mate Start: start position of read alignment to the receptor chromosome; Aln2nd Chr: receptor chromosome to which the folded part of read is aligned if it belongs to typeB; Aln2ndStart: start position of the folded part of read alignment to the receptor chromosome if it belongs to typeB; Aln2ndEnd: end position of the folded part of read alignment to the receptor chromosome if it belongs to typeB; Aln2ndCigar: CIGAR information of the folded part of read alignment if it belongs to typeB, wherein mM represents x bases aligned to the receptor chromosome in this region, and nS represents y bases aligned to other regions of the receptor chromosome or the exogenous target gene fragment.
[0106] Example 3
[0107] Gene position of insertion of exogenous target gene fragment into receptor genome
[0108] By identifying the insertion site of the exogenous gene fragment in the above-mentioned PagEXPA1 gene editing strains Cas_EXPA1-9 and Cas_EXPA1-12, the sequences within 2000 bp upstream and downstream of the insertion site were extracted by using the self-programmed perl program (Appendix 2: 3, extraction of genes and flanking sequences near the insertion site), and the position of the coding gene where the insertion site is located was found by sequence alignment, such as whether it is located in the CDS or UTR region of the gene, or the non-coding region between genes, so as to determine whether the insertion of the exogenous gene will affect the functional genes near the insertion site. As can be seen from Table 4, there are two insertion sites in the Cas_EXPA1-9 strain, wherein the insertion site Chr06B:11,070,447-11,070,449 is located at the upstream 1664 bp position of the Pag.B06G001223.1 transcript, the direction of the exogenous gene fragment is the same as that of the receptor genome, and is opposite to that of the transcript; the insertion site Chr19A:11,432,810-11,432,831 is located at the upstream 62 bp position of the Pag.A19G000860.1 transcript, the direction of the exogenous gene fragment is the same as that of the receptor genome, but is opposite to that of the transcript. There is only one insertion site in the Cas:EXPA1-12 strain, which is Chr04B:16,562,665-16,562,671, and it is located at the downstream 439 bp position of the Pag.B04G001464.1 transcript, the direction of the exogenous gene fragment is the same as that of the receptor genome and the transcript.
[0109] Table 4 Site information of the insertion of the exogenous gene fragment into the receptor genome in different transgenic strains
[0110]
[0111] Note: ID: ID code of reads; TDNA: name of the inserted exogenous gene fragment; Chr: chromosome of the receptor 84K genome to which the insertion site of the exogenous gene fragment is aligned; Start: start site of the insertion site of the exogenous gene fragment aligned to the receptor chromosome; End: end site of the insertion site of the exogenous gene fragment aligned to the receptor chromosome; Len: length of the fragment deleted by the insertion site of the exogenous gene fragment; Ori: direction of the insertion of the exogenous gene fragment into the receptor genome, (+) indicates that the insertion direction of the exogenous gene fragment is consistent with that of the receptor genome, (-) indicates that the insertion direction of the exogenous gene fragment is opposite to that of the receptor genome; geneRegion: position of the transcript where the insertion site of the exogenous gene fragment is located, + / - indicates that the transcript is consistent / opposite to the genome, left indicates that the insertion site is located upstream of the genomic position of the transcript, and right indicates that the insertion site is located downstream of the genomic position of the transcript.
[0112] In summary, the insertion site and copy number of the exogenous target gene fragment inserted into the receptor genome can be identified by high-throughput second-generation sequencing. Meanwhile, the insertion site can be accurately located to the specific chromosome and gene transcript position, thereby providing information support for whether the insertion of the exogenous target gene fragment has an unintended impact on the coding gene.
[0113] In addition, the sequencing depth of the second-generation sequencing of the application is 5X, and the sequencing reads can basically cover the entire genome. Through the alignment depth of the reads and the average alignment depth of the receptor genome, the frequency of the insertion of the exogenous target gene fragment into the receptor genome, i.e., the copy number of the exogenous target gene, can also be determined. However, the second-generation sequencing fragment is short, and in the actual alignment process, there may be a case that only one end of the exogenous target gene fragment can find alignment information, and the corresponding position on the receptor genome needs to be verified. Therefore, the sequencing depth can be increased to increase the number of aligned reads in the exogenous target gene fragment, so as to obtain more alignment information. However, increasing the sequencing depth will also increase the sequencing cost, and for a large genome or a large sample size, the cost will increase exponentially, and the cost performance is not high. For only finding the copy number or insertion site of the exogenous target gene, the second-generation sequencing at 5 times the depth of the reference genome is sufficient.
[0114] The application is not limited to the above-mentioned embodiments, and the application will cover the range described in the technical solutions and various modifications and equivalent changes in the scope of the claims. Any modification or improvement of the application made by those skilled in the art without departing from the technical solutions of the application belongs to the scope of the application claimed.
[0115] Appendix 1: Related term explanation
[0116] 1. Exogenous vector: part of the sequence of the expression vector introduced by the transgene, including the part of the vector inserted into the plant body (LBT-DNA to RB T-DNA).
[0117] 2. Target gene: sequence of the exogenous target gene introduced by the transgene.
[0118] 3. Exogenous target gene fragment: sequence of the expression vector and the target gene introduced by the transgene.
[0119] 4. Receptor 84K genome: sequence of the genome of the transgenic receptor 84K.
[0120] 5. Reference genome: sequence of the integrated exogenous target gene fragment and receptor 84K genome.
[0121] 6. Definition of sam / bam file:
[0122] The SAM (Sequence Alignment / Map) format is a standard text format for storing alignment results of high-throughput sequencing data. It was developed by Heng Li et al. to provide a unified and detailed representation of alignments between short sequence reads and a reference genome. The SAM format is widely used in the field of bioinformatics, especially when dealing with and analyzing high-throughput sequencing data. Here are some key features of the SAM format:
[0123] 1) Header: The beginning part of a SAM file contains a header, which includes metadata about the reference sequence, program version, library information, etc. The header starts with a '@' character.
[0124] 2) Alignment Section: Following the header is a series of alignment sections, each representing an alignment result of a read to the reference genome.
[0125] 3) Column Format: Each alignment section consists of 11 basic fields separated by tabs, which contain basic information about the read, such as QNAME (Query Template Name), FLAG (Flag, indicating alignment-related information), RNAME (Reference Sequence Name), POS (Position), MAPQ (Mapping Quality), CIGAR (CIGAR String), RNEXT (Next Reference Sequence Name), PNEXT (Paired Read Position on Reference), TLEN (Template Length), SEQ (Read Sequence), QUAL (Alignment Quality).
[0126] 4) Optional Fields: In addition to the 11 basic fields, the SAM format supports optional fields, which provide additional information such as pairing information, quality scores, raw sequencing data, etc.
[0127] 5) CIGAR String: The CIGAR string is a compact representation of the alignment relationship between the read and the reference sequence, including matches, mismatches, insertions, deletions, and soft / hard clips, etc.
[0128] Appendix 2: Perl program code script for foreign insertion site identification and flanking sequence extraction
[0129] 1. Main process of foreign insertion site identification
[0130]
[0131]
[0132]
[0133] 2. Identification of reads associated with insertion sites and extraction of alignment information
[0134]
[0135]
[0136]
[0137]
[0138] 3. Extraction of genes and flanking sequences near insertion sites
[0139]
[0140]
[0141]
[0142]
[0143]
[0144]
Claims
1. A method for identifying the insertion site and copy number of a foreign gene in transgenic plants based on next-generation sequencing, characterized in that, The identification method includes the following steps: (1) Extract DNA from the plant sample to be tested; (2) High-throughput second-generation Novaseq 6000 library construction and sequencing, filtering raw sequencing data; (3) Integrate exogenous target gene fragments and receptor genomes as reference genomes; (4) Alignment and identification of filtered sequencing fragments with the integrated reference genome; 1) To identify the frequency of insertion of exogenous target gene fragments on the recipient chromosome; 2) Identify the insertion site of the exogenous target gene fragment into the recipient genome; 3) Identify the location of the exogenous target gene fragment insertion site in the recipient genome that encodes the gene.
2. The identification method according to claim 1, characterized in that, The plant used in step (1) is a poplar.
3. The identification method according to claim 1, characterized in that, The main parameters for filtering the raw sequencing data in step (2) are: -q 20; -u 40; -n 15; -l 50; -w 4; and the sample sequencing depth is 5X.
4. The identification method according to claim 1, characterized in that, The specific steps for identifying the frequency of exogenous target gene fragment insertion on the recipient chromosome in step 1) are as follows: Sequencing depths compared to exogenous target gene fragments and the average sequencing depth of the recipient genome were statistically analyzed. The average sequencing depth is calculated as: the number of bases aligned to the chromosome / the length of the chromosome; When the alignment and sequencing depth of the exogenous target gene fragment is ≤0.3, it indicates that there is no insertion of the exogenous target gene fragment on the recipient chromosome. When the sequencing depth of the exogenous target gene fragment is 0.7 to 1.3 times that of the average sequencing depth of the recipient genome, it indicates that an insertion of the exogenous target gene fragment has occurred on the recipient chromosome. When the sequencing depth of the exogenous target gene fragment is 1.7 to 2.3 times that of the average sequencing depth of the recipient genome, it indicates that two insertions of the exogenous target gene fragment have occurred on the recipient chromosome.
5. The identification method according to claim 1, characterized in that, The specific steps for identifying the insertion site of the exogenous target gene fragment into the recipient genome in step 2) are as follows: (1) Using IGV software, by comparing the alignment information of reads that are aligned to the 5' and 3' ends of the exogenous target gene fragment and are marked with different colors, reference information is provided for identifying the location of the exogenous target gene fragment inserted into the 84K genome of the receptor. (2) Based on the rules of mate pair alignment in the Linux system, a Perl script is used to identify the following two types of reads and extract relevant information according to the alignment information stored in the sam / bam file: 1) Type A: Based on the RNAME and RNEXT information in the BAM file, identify pair reads with one end located in the exogenous target gene fragment and the other end located in the recipient genome, and list their positions; 2) Type B: Identify split mapped reads from CIGAR information, i.e., some are aligned to the exogenous target gene fragment, and the other part is aligned to the recipient genome reads; And from the SupplementaryAlignments recorded in the bam file, record the alignment position of the part that is aligned to the receptor genome; (3) Based on the extracted comparison information, the insertion site of the exogenous target gene fragment into the recipient genome can be preliminarily determined. Using IGV software, the corresponding insertion site is searched in the recipient genome. If there is an obvious gap or a decrease in depth near the site, it can be determined as the starting site or interval of the exogenous target gene fragment insertion.
6. The identification method according to claim 1, characterized in that, The specific steps for step 3) of identifying the location of the exogenous target gene fragment insertion site within the coding gene of the recipient genome are as follows: The sequence within 2000 bp upstream and downstream of the insertion site is extracted, and the position of the coding gene at the insertion site is found by sequence alignment to determine the relationship between the insertion site and the neighboring coding genes on the recipient genome. The direction of insertion of the exogenous target gene fragment into the recipient genome and the transcription direction of the recipient genome and the coding gene are used to determine whether the insertion site is upstream or downstream of the coding gene.
7. The application of the identification method according to any one of claims 1 to 5 in assessing the biosafety of transgenic plants.
8. The application of the identification method according to any one of claims 1 to 5 in investigating gene expression and function.
9. The application of the identification method according to any one of claims 1 to 5 in screening for genetically modified products with stable inheritance.
10. The identification method according to any one of claims 1 to 5 is used in distinguishing different genetically modified products by detecting the insertion site of exogenous DNA.
Citation Information
Patent Citations
Method for determining insertion site of transgenic line by re-sequencing technique
CN108034706A
Method for rapidly identifying transgenic or gene editing material and insertion site thereof by using whole genome resequencing data
CN110556165A
Method for rapidly identifying exogenous insertion sequence based on whole genome sequencing data
CN116189772A
Method for efficiently identifying exogenous DNA insertion site in genome by using high-throughput sequencing technology
CN119876357A