Method and system for gene editing material testing
By using PCR amplification and SuperDecode software processing, the cost and complexity issues of large-scale gene-edited sample detection have been resolved, enabling rapid and low-cost genotyping identification, and it is applicable to Windows, macOS, and Linux systems.
Patent Information
- Application Number
- PCT/CN2025/091692
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-19
- Filing Date
- 2025-04-28
- Publication Date
- 2026-02-26
AI Technical Summary
Existing technologies lack rapid, low-cost methods for detecting large volumes of gene-edited samples, and require complex bioinformatics analysis and computer programming capabilities, making it difficult to meet the needs of genotyping for large-scale samples.
Multiple specific recognition sequences are constructed using PCR amplification-based methods. The gene fragments are sequenced and mutation analyzed using SuperDecode software, including the processing of Sanger, NGS, and TGS results. An operating interface is provided for Windows, macOS, and Linux systems.
It enables rapid library construction and mutation analysis for large batches of samples, can detect all types of mutations, improves detection efficiency, simplifies the operation process, and is suitable for non-professionals.
Smart Images

Figure CN2025091692_26022026_PF_FP_ABST
Abstract
Description
Method and system for gene editing material detection TECHNICAL FIELD
[0001] The present application relates to the technical field of gene editing result detection, and particularly relates to a method for gene editing material detection. BACKGROUND
[0002] The detection of gene editing results mainly includes amplicon detection based on Sanger sequencing and amplicon detection based on next-generation sequencing (NGS). Sanger sequencing technology can detect fragments with a length of 700 to 1000 bp, and has high accuracy, so Sanger sequencing is currently the gold standard for genotype detection. At the same time, with the development of bioinformatics, various mutation analysis software based on Sanger sequencing have also been developed, such as TIDE, ICE, DSDecode, etc. However, Sanger sequencing can only detect one sample at a time, and for large-scale sample genotype identification, its efficiency is low, the sequencing cost is high, and for complex mutation types and polyploid samples, Sanger sequencing is difficult to detect the results, and cannot meet the various needs of current scientific research on editing material genotype detection. In 2006, the development of NGS technology solved the problem of low throughput and inability to detect complex mutation types of Sanger sequencing. Its principle is to break the DNA fragments, add adapters, and fix them on a chip, then add dNTP containing fluorescent signals for amplification, and detect the fluorescent signals to determine the genotype of the sample. Long fragment sequencing, also known as third-generation sequencing (TGS), mainly includes single molecule fluorescent sequencing technology and nanopore single molecule sequencing technology, which can detect fragments of more than 10 kb, and cover sequencing of large fragments of genome, so it can detect and analyze experimental results such as large fragment deletion and tiled deletion. Even though long fragment sequencing may produce various errors during sequencing, it can be corrected by sequencing depth, so that the accuracy of the sequencing results is more than 99%. For long fragment sequencing technology, it is currently mainly used for sample genome assembly, methylation detection, genome structure variation detection, etc., and has not been fully applied to gene editing mutation type analysis.
[0003] At present, various sequencing technologies are very mature, and the shortcomings of various technologies can be mutually complemented, and the cost is also lower and lower, which provides convenience for gene editing detection. However, there is currently a lack of a method for building a library for a large number of samples, so that the detection of a large number of gene editing samples requires high cost and time. At the same time, analyzing the sequencing results requires certain bioinformatics analysis ability and certain computer programming ability, and it is necessary to learn the use methods of different software, which will consume a lot of time. SUMMARY
[0004] The present application aims to provide a method for detecting gene editing materials, which can quickly detect the results of gene editing of a large number of samples, detect all types of mutations in gene editing materials in all aspects, and facilitate researchers who are not proficient in using various computer programming languages to analyze sequencing results in a short time.
[0005] A method for detecting gene editing materials, comprising:
[0006] Constructing a plurality of specific recognition sequences;
[0007] Sequencing gene fragments according to the specific recognition sequences;
[0008] Using SuperDecode software to analyze mutations in the sequencing results.
[0009] Preferably, the constructing a plurality of specific recognition sequences comprises:
[0010] Adding position-barcode, plate-barcode and Library-barcode specific recognition sequence libraries to sample amplicons;
[0011] The position-barcode sequence library includes 192 specific position-barcode sequences, and each of the specific recognition sequences for NGS library construction and TGS library construction includes 96 specific recognition sequences;
[0012] The plate-barcode sequence library includes 102 specific plate-barcode sequences, 96 of which are used for NGS library construction, and 6 of which are used for TGS library construction;
[0013] The Library-barcode sequence library includes 20 specific Library-barcode sequences, all of which are used for NGS library construction.
[0014] Preferably, the sequencing gene fragments according to the specific recognition sequences comprises:
[0015] Sanger sequencing, NGS and TGS are performed on the gene fragments.
[0016] Preferably, the using SuperDecode software to analyze mutations in the sequencing results comprises analyzing mutations in Sanger sequencing results, specifically:
[0017] Obtaining first project information of the current mutation analysis according to user input files and settings;
[0018] Read the overlapping peaks in the Sanger sequencing results by using the sequencing overlapping peak graph decoding principle of DsD (Degenerate Sequence Decoding);
[0019] Based on the sequence analysis software, find the corresponding sequence in the wild type sequence, analyze the mutations of the Sanger sequencing results of the amplicon, and read the variation types of the target sequence of the sample.
[0020] Preferably, the mutation analysis of the sequencing results by using the SuperDecode software includes mutation analysis of NGS results, specifically:
[0021] Obtain the second project information of the current mutation analysis according to the user input files and settings;
[0022] Merge pairs of reads in paired-end sequencing into complete sequences, and split the library based on the specific identification sequence of the sample;
[0023] Correspond the reads to a specific sample, align the reads of the sample with the reference sequence by a sequence alignment method, and obtain the variation information of the sample.
[0024] Preferably, the mutation analysis of the sequencing results by using the SuperDecode software includes mutation analysis of TGS results, specifically:
[0025] Obtain the third project information of the current mutation analysis according to the user input files and settings;
[0026] Split the library based on the specific identification sequence of the sample, and correspond the reads to a specific sample;
[0027] Extract the sequence common to all reads of a specific sample to remove random mutations caused by sequencing, and finally align the reads of the sample with the reference sequence by a sequence alignment method to obtain the variation information of the sample.
[0028] Preferably, the first project information of the current mutation analysis according to the user input files and settings includes:
[0029] The first project information includes: amplicon wild type sequence information, sample amplicon Sanger sequencing result file, and optional first setting condition.
[0030] Preferably, the second project information of the current mutation analysis according to the user input files and settings includes:
[0031] The second project information includes: amplicon wild type sequence information, targeted sequence information of gene editing, sample NGS result information, specific recognition sequence information corresponding to each sample, and optional second setting conditions;
[0032] The optional second setting conditions are settings of mutation analysis parameters, including:
[0033] Sample analysis mode, selected from diploid, polyploid and low frequency;
[0034] Output threshold, used for setting the threshold of result output;
[0035] Running thread number, used for setting the number of threads used by the program to run, and adjusting the running speed of HiDecode decoding.
[0036] Preferably, the third project information of the current mutation analysis according to the user input file and settings includes:
[0037] The third project information includes: amplicon wild type sequence information, targeted sequence information of gene editing, sample TGS result information, specific recognition sequence information corresponding to each sample, and optional third setting conditions;
[0038] The optional third setting conditions are settings of mutation analysis parameters, including:
[0039] Sample analysis mode, selected from diploid, polyploid and low frequency;
[0040] Output threshold, used for setting the threshold of result output;
[0041] Running thread number, used for setting the number of threads used by the program to run, and adjusting the running speed of LaDecode decoding.
[0042] A system for gene editing material detection, comprising:
[0043] A data acquisition module for constructing a plurality of specific recognition sequences;
[0044] A sequencing module for sequencing gene fragments according to the specific recognition sequences;
[0045] An analysis module for performing mutation analysis on the sequencing results using SuperDecode software.
[0046] The application has the advantages that: 1. The application is based on a PCR amplification method, and the construction of an amplicon of gene editing material and the introduction of a specific recognition sequence realize NGS and TGS library construction of a large number of samples quickly and at low cost; 2. The application develops a software SuperDecode for gene editing material detection, which can analyze mutations of Sanger, NGS and TGS results commonly used at present, and can detect all types of mutations in the sample, including large fragment deletion, sequence insertion, single base mutation, chimeric mutation and the like; 3. The SuperDecode developed by the application is a Window platform software, compared with a web version tool, the SuperDecode can quickly read a large number of sample sequencing result files and perform mutation analysis. For the tool of the Linux system, the SuperDecode provides a convenient and simple operation interface, which is convenient for researchers to use, and greatly improves the efficiency of gene editing material detection. BRIEF DESCRIPTION OF DRAWINGS
[0047] Fig. 1 is a flow chart of a method for gene editing material detection according to the application;
[0048] Fig. 2 is a flow chart of NGS library construction of a large number of samples according to the application;
[0049] Fig. 3 is a result chart of NGS sample library construction according to the application;
[0050] Fig. 4 is a result chart of NGS sample detection by HiDecode according to the application;
[0051] Fig. 5 is a result chart of rice protoplast sample detection by using HiDecode according to the application;
[0052] Fig. 6 is a flow chart of TGS sample library construction according to the application;
[0053] Fig. 7 is a result chart of TGS sample detection by using LaDecode according to the application;
[0054] Fig. 8 is a result chart of rice Ehd1 promoter tiled deletion detection based on LaDecode mutation analysis according to the application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0056] It should be noted that all directionality indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present application are only used to explain the relative position relationship, movement condition, etc. between components in a certain posture (as shown in the drawings), and if the certain posture changes, the directionality indications will also change accordingly.
[0057] In addition, the descriptions involving "first", "second", etc. in the present application are only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" can explicitly or implicitly include at least one of the features. In addition, the technical solutions of various embodiments can be combined with each other, but it must be based on the fact that a person skilled in the art can realize it, and when the combination of technical solutions contradicts each other or cannot be realized, it should be considered that the combination of technical solutions does not exist and is not within the protection scope required by the present application.
[0058] At present, various sequencing technologies are very mature, and the shortcomings of various technologies can be mutually complemented, and the cost is also lower and lower, which provides convenience for gene editing detection. However, at present, there is no method for library construction of a large number of samples, so that the detection of a large number of gene editing samples needs to spend high cost and time. At the same time, biological information analysis ability and certain computer programming ability are needed to analyze the sequencing results, and the use method of different software needs to be learned, which will consume a lot of time.
[0059] The present application is based on the method of PCR amplification, the construction of the amplified subsequence of the gene editing material, and the introduction of specific recognition sequence, which realizes the library construction of NGS and TGS of a large number of samples quickly and at low cost; the present application develops a software SuperDecode for gene editing material detection, which can analyze mutations of Sanger, NGS and TGS results commonly used at present, and can detect all types of mutations in the sample, including large fragment deletion, sequence insertion, single base mutation, chimeric mutation, etc.; the SuperDecode developed by the present application is a Window platform software, compared with the web version tool, the SuperDecode can quickly read a large number of sample sequencing result files and perform mutation analysis. For the tool of Linux system, the SuperDecode provides a convenient and simple operation interface, which is convenient for researchers to use, and greatly improves the efficiency of gene editing material detection.
[0060] Sequence information
[0061] Embodiment 1
[0062] A method for gene editing material detection, referring to FIG. 1, comprising: A method for gene editing material detection, referring to FIG. 1, comprising:
[0063] S100, constructing multiple specific recognition sequences;
[0064] S200, sequencing the gene fragments according to the specific recognition sequences;
[0065] S300, using SuperDecode software to analyze mutations of the sequencing results.
[0066] In order to quickly detect the gene editing results of a large number of samples, comprehensively detect all types of mutations in gene editing materials, and facilitate researchers who are not skilled in using various computer programming languages to analyze sequencing results in a short time, the application provides a NGS amplicon library construction method suitable for 9,216 samples, a TGS amplicon library construction method suitable for 576 samples, and a use method of SuperDecode software suitable for mutation analysis of Sanger sequencing, NGS and TGS results. SperDecode is a software for Windows, macOS and Linux systems, which can detect all types of variations of samples, and adds sequencing result quality control function to make the detection results more reliable.
[0067] Preferably, referring to Figures 2 and 6, S100, constructing multiple specific recognition sequences includes:
[0068] Adding position-barcode, late-barcode and Library-barcode specific recognition sequence library to sample amplicon;
[0069] The position-barcode sequence library includes 192 specific position-barcode sequences, and there are 96 specific recognition sequences for NGS library construction and TGS library construction;
[0070] The plate-barcode sequence library includes 102 specific plate-barcode sequences, of which 96 are used for NGS library construction and 6 are used for TGS library construction;
[0071] The Library-barcode sequence library includes 20 specific Library-barcode sequences, all of which are used for NGS library construction.
[0072] Library construction method for NGS detection of a large number of samples
[0073] The application can specifically recognize up to 184,320 NGS amplicon samples by adding three specific recognition sequences of position-barcode, plate-barcode and Library-barcode to sample amplicon.
[0074] Specifically, the NGS library of the present application contains 96 specific position-barcode sequences, which can correspond to the positions of A1 to H12 on a 96-well PCR plate.
[0075] Specifically, the NGS library of the present application contains 96 specific plate-barcode sequences, which can correspond to different 96-well PCR plates.
[0076] Specifically, the NGS library of the present application contains 20 specific Library-barcode sequences, which can correspond to different libraries.
[0077] TGS library construction method for detecting a large number of samples
[0078] In the first aspect, by adding two specific recognition sequences of position-barcode and plate-barcode to the sample amplicon, specific recognition of up to 576 amplicon samples can be achieved.
[0079] Specifically, the TGS library of the present application contains 96 specific position-barcode sequences, which can correspond to the positions of A1 to H12 on a 96-well PCR plate.
[0080] Specifically, the TGS library of the present application contains 6 specific plate-barcode sequences, which can correspond to different 96-well PCR plates.
[0081] Preferably, S200, sequencing the gene fragments according to the specific recognition sequences comprises:
[0082] Sanger sequencing, NGS, and TGS are performed on the gene fragments.
[0083] Preferably, S300, performing mutation analysis on the sequencing results using SuperDecode software comprises performing mutation analysis on the Sanger sequencing results, specifically:
[0084] According to the user input file and settings, the first item information of the current mutation analysis is obtained.
[0085] Using the sequencing overlap peak decoding principle of DsD, the overlap peaks in the Sanger sequencing results are read.
[0086] Based on the sequence analysis software, the corresponding sequence in the wild type sequence is searched, and the mutation analysis of the Sanger sequencing results of the amplicon is performed to read the variation type of the target sample at the target sequence.
[0087] A module DsDecodeMS for decoding Sanger sequencing results, the analysis steps of which include:
[0088] Obtaining the project information for the current analysis according to the files and settings input by the user, specifically including amplicon wild-type sequence information, sample amplicon Sanger sequencing result files, and optional setting conditions;
[0089] Using the sequencing overlap peak graph decoding principle of DsD, reading the overlap peaks in the Sanger sequencing results, and based on the sequence analysis software, finding the corresponding sequence in the wild-type sequence, thereby performing mutation analysis on the Sanger sequencing results of the amplicon, and reading the variation type of the target sample at the target sequence.
[0090] Specifically, the amplicon wild-type sequence is the gene sequence of the sample before gene editing.
[0091] Specifically, the optional setting conditions are the settings of mutation analysis parameters, including:
[0092] Target sequence, used for quickly retrieving the mutation of the target sample at a specific position;
[0093] Fluorescence signal threshold, site sequencing confidence value, and fluorescence signal value of a certain base greater than the confidence value is effective sequencing result, the default value is 0.25;
[0094] Length of anchor sequence and merged sequence, sequence length upstream and downstream of the mutation site;
[0095] End sequence removal settings, based on the setting ratio or Richard Mott algorithm to remove low-quality end sequences, to remove the influence of low-quality sequences on mutation analysis and improve the accuracy of mutation analysis.
[0096] Optionally, the input sample amplicon Sanger sequencing file can detect sequencing quality in the peak graph display area, and the sequence retrieval function can be used to quickly locate to the target sequence of the sample in the peak graph display area.
[0097] Preferably, referring to FIGS. 4 and 5, S300, using SuperDecode software to perform mutation analysis on the sequencing results includes performing mutation analysis on NGS results, specifically:
[0098] Obtaining the second project information for the current mutation analysis according to the files and settings input by the user;
[0099] Merging pairs of reads in paired-end sequencing to complete the sequence, and splitting the library based on the specific recognition sequence of the sample;
[0100] According to the sequence alignment method, the reads of the sample are aligned with the reference sequence to obtain the variation information of the sample.
[0101] According to the user input file and setting, the third project information for the current analysis is obtained, specifically including the amplicon wild type sequence information, the targeted sequence information of gene editing, the sample NGS result information, the specific recognition sequence information corresponding to each sample, and the optional setting conditions. The HiDecode will combine the paired reads in the paired-end sequencing to complete the complete sequence, split the library based on the specific recognition sequence of the sample, correspond the reads to a specific sample, and then align the reads of the sample with the reference sequence through the sequence alignment method to obtain the sample variation information.
[0102] Specifically, the amplicon wild type sequence is the DNA sequence of the sample before gene editing.
[0103] Specifically, the sample NGS result information is the second-generation sequencing result of the sample, usually the paired-end sequencing result, including R1 and R2 two sequencing files.
[0104] Specifically, the targeted sequence information of gene editing is the DNA sequence targeted by the gene editing tool.
[0105] Specifically, the specific sequence corresponding to the sample is the position-barcode sequence and plate-barcode sequence used in the NGS library construction, which is used by HiDecode to correspond the result to a specific sample.
[0106] Preferably, referring to FIGS. 7 and 8, S300, the mutation analysis of the sequencing result by using the SuperDecode software includes the mutation analysis of the TGS result, specifically:
[0107] According to the user input file and setting, the third project information for the current analysis is obtained;
[0108] Based on the specific recognition sequence of the sample, the library is split, and the reads are corresponded to a specific sample;
[0109] Extract the sequence common to all reads of a specific sample to remove random mutations caused by sequencing, and finally align the reads of the sample with the reference sequence by using the alignment method to obtain the variation information of the sample.
[0110] Specifically, the amplicon wild type sequence is the DNA sequence of the sample before gene editing.
[0111] Specifically, the targeted sequence information of gene editing is the DNA sequence targeted by the gene editing tool.
[0112] Specifically, the sample corresponding specific sequence is the position-barcode sequence and the plate-barcode sequence used in the TGS sample library construction, which is used by LaDecode to correspond the results to a specific sample.
[0113] Preferably, the first project information for the current analysis is obtained according to the files and settings input by the user, including:
[0114] The first project information includes: amplicon wild type sequence information, sample amplicon Sanger sequencing result file, and optional first setting condition.
[0115] Preferably, the second project information for the current mutation analysis is obtained according to the files and settings input by the user, including:
[0116] The second project information includes: amplicon wild type sequence information, targeted sequence information of gene editing, sample NGS result information, specific recognition sequence information corresponding to each sample, and optional second setting condition.
[0117] The optional second setting condition is the setting of mutation analysis parameters, including:
[0118] Sample analysis mode, including diploid, polyploid and low frequency;
[0119] Output threshold, used to set the threshold of result output;
[0120] Running thread number, used to set the thread number used by the program to run, to adjust the running speed of HiDecode decoding.
[0121] Sample analysis mode, which is located in the Model setting option of HiDecode, and can select 3 modes of diploid, polyploid and low frequency. Among them, diploid represents diploid mode, which is used to detect the gene editing results of diploid samples, such as rice materials. The result output in this mode is the genotype of alleles on a pair of chromosomes. The polyploid mode is used to detect the gene editing results of polyploid samples. The low frequency mode is mainly used to detect the gene editing results of cell lines, such as rice protoplasts, callus and other materials;
[0122] Output threshold, which is located in the Threshold setting option in HiDecode, is used to set the threshold of the result output, i.e. the percentage content of a certain genotype in the result is greater than the set threshold, and the genotype is considered as the reliable result of the sample detection result. The genotype below the threshold is caused by errors in PCR amplification or sequencing. The default output threshold is 0.25 in the diploid mode, 0.1 in the polyploid mode, and 0.0025 in the low frequency mode.
[0123] Running thread number, which is located in the CPU thread setting option in HiDecode, can be used to set the number of threads used by the program to run, and adjust the running speed of HiDecode decoding.
[0124] Preferably, the third project information for the current analysis is obtained according to the user input file and settings, and includes:
[0125] The third project information includes: amplicon wild type sequence information, gene editing targeted sequence information, sample TGS result information, specific recognition sequence information corresponding to each sample, and optional third setting conditions;
[0126] The optional third setting conditions are the settings of mutation analysis parameters, including:
[0127] The sample analysis mode is selected from diploid, polyploid and low frequency;
[0128] The output threshold is used to set the threshold of the result output;
[0129] The running thread number is used to set the number of threads used by the program to run, and adjust the running speed of LaDecode decoding.
[0130] The project information for the current mutation analysis is obtained according to the user input file and settings, and specifically includes amplicon wild type sequence information, gene editing targeted sequence information, sample TGS result information, specific recognition sequence information corresponding to each sample, and optional setting conditions. LaDecode will split the library based on the specific recognition sequence of the sample, map the reads to the specific sample, then extract the common sequences in all reads of the specific sample to remove random mutations caused by sequencing, and finally use sequence alignment method to align the reads of the sample with the reference sequence to obtain sample variation information.
[0131] Specifically, the amplicon wild type sequence is the DNA sequence of the sample before gene editing.
[0132] Specifically, the targeted sequence information of the gene editing is a DNA sequence targeted by a gene editing tool.
[0133] Specifically, the sample corresponding specific sequence is a position-barcode sequence and a plate-barcode sequence used in the TGS sample library construction, which is used by LaDecode to correspond the results to a specific sample.
[0134] Specifically, the selectable setting condition is a setting of a mutation analysis parameter, including:
[0135] A sample analysis mode, which is selected from diploid, polyploid and low frequency modes in the Model setting option of LaDecode. The diploid mode is used for detecting gene editing results of a diploid sample, such as a rice material, and the output result of the mode is a genotype of alleles on a pair of chromosomes. The polyploid mode is used for detecting gene editing results of a polyploid sample. The low frequency mode is mainly used for detecting gene editing results of a cell line, such as a rice protoplast and a callus material;
[0136] An output threshold, which is set in the Threshold setting option of LaDecode, is used for setting a threshold of the result output, that is, a percentage content of a certain genotype in the result is greater than the set threshold, and the genotype is considered as a reliable result in the sample detection result. A genotype lower than the threshold is caused by an error in PCR amplification or sequencing. The result output threshold in the diploid mode is 0.25, the threshold in the polyploid mode is 0.1, and the threshold in the low frequency mode is 0.0025.
[0137] A running thread number, which is set in the CPU thread setting option of LaDecode, is used for setting a thread number used for program running, and adjusting a running speed of LaDecode.
[0138] Embodiment 2
[0139] A system for detecting a gene editing material, comprising:
[0140] A data acquisition module is used for constructing a plurality of specific recognition sequences.
[0141] A sequencing module is used for sequencing a gene fragment according to the specific recognition sequence.
[0142] An analysis module is used for performing mutation analysis on the sequencing result by using SuperDecode software.
[0143] Part of the reagents used in the examples are as follows:
[0144] Manitol, NaCl, MES, KCl, MgCl2·6H2O, PEG4000, Bovine Serum Albumin and CaCl2·2H2O were purchased from Sigma company;
[0145] MES (0.2M): 3.9048g of MES was weighed, and sterilized double distilled water was used to make up to 100mL. After adjusting pH to 5.7 using 1M KOH, it was filtered using a 0.22μm filter.
[0146] CaCl2·2H2O (1M): 14.701g of CaCl2·2H2O was weighed, and sterilized double distilled water was used to make up to 100mL. After filtering using a 0.22μm filter.
[0147] KCl (2M): 14.91g of KCl was weighed, and sterilized double distilled water was used to make up to 100mL. After filtering using a 0.22μm filter.
[0148] MgCl2·6H2O (0.5M): 10.165g of MgCl2·6H2O was weighed, and sterilized double distilled water was used to make up to 100mL. After filtering using a 0.22μm filter.
[0149] MMG: 9.11g of Manitol was weighed, 3mL of MgCl2·6H2O (0.5M), 1mL of MES (200mM) were measured, and mixed. After making up to 100mL with sterilized double distilled water and filtering using a 0.22μm filter.
[0150] W5: 0.9g of NaCl, 1.838g of CaCl2·2H2O were weighed, 0.25mL of KCl (2M), 1mL of MES (0.2M) were measured, and mixed. After making up to 100mL with sterilized double distilled water and filtering using a 0.22μm filter.
[0151] PEG4000: 40g of PEG4000, 3.644g of Manitol were weighed, 10mL of CaCl2·2H2O (1M) was measured, and mixed. After making up to 100mL with sterilized double distilled water and filtering using a 0.22μm filter.
[0152] Enzyme Solution: Take 10.932 g of Manitol, 0.1 g of Bovine Serum Albumin, 5 mL of MES (0.2 M), 1.5 g of Cellulase RS, 0.75 g of Macerozyme R10, mix and use sterilized double distilled water to make up to 100 mL, use 0.22 μm filter membrane to filter after adjusting pH to 5.7 with KOH.
[0153] Example 1. NGS library construction of large batch samples
[0154] In this example, the NGS library construction was performed on the rice editing material of OTUB1 gene (SEQ ID NO. 1), and the selected gene sequence was as follows:
[0155] CTCCTTTATTGGTGCTTGATCTACAACTGGTGTTTTACTTTTTTACAAAAAAATGTAATCTCCTTGCAGTGCACTCAAATTATTGCAACCTCCTTCCTTATGTTCCCACCCTCATTATTTTCAGATATTCATTGATCAGCTGGAAAGTGTTCTGCAGGGACATGAATCCTCCATAGGGTAAATATCCTAGAGTTATATTTGTATCCTTAATGCATATGACCAATAATCATGTATTAACAACAAGCAATTTTTGTAATTGTTTATAAAGTATGGCATGTCCATCATAAATGTTTTCCTTCTGTAGTGAATCTATTTTGTTTTCCTG (underlined region is the target region).
[0156] The gene editing material was subjected to amplicon construction, and specific primers were designed at the positions 80-160 bp upstream and downstream of the detection site. The sample amplicon specific amplification primers designed in this example were as follows:
[0157] Forward primer T#F (SEQ ID NO. 2): 5'-ctcggagtgatcgcacTTATTGCAACCTCCTTCCT-3' (concentration of 10 μM)
[0158] Reverse primer T#R (SEQ ID NO. 3): 5'-ctgagaggctggatggATGATGGACATGCCATAC-3' (concentration of 10 μM)
[0159] The amplicon primer sequences used in the second round / third round PCR reaction of this example were as follows:
[0160] Plate-01 (SEQ ID NO. 4) (10 μM): (Plate-01 specific recognition sequence in bold)
[0161] OTUB1-Pos N -F (SEQ ID NO. 5) (10 μM): (OTUB1-Pos N specific recognition sequence in italic)
[0162] Lib-F (SEQ ID NO. 6) (10 μM): AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATCTGTCA
[0163] Lib-R (SEQ ID NO. 7) (10 μM): (underlined dotted line is library specific recognition sequence)
[0164] The NGS sample library of the present application is shown in Figure 2, which can be divided into two-step method and three-step method, wherein the two-step method steps are:
[0165] The sample number detected in the present embodiment is 10, and the sample to be detected is divided into 96-well PCR holes by using a multi-channel pipettor or a common pipettor. The PCR total reaction solution is prepared according to the number of samples to be detected, and the PCR reaction system of each hole is 15 μL. The content of each component in the system is: 2x Taq mix 7.5 μL, forward primer T#F and reverse primer T#R each 0.3 μl, double distilled water 5.9 μL. The PCR reaction system is divided into 96-well PCR holes by using a multi-channel pipettor or a common pipettor, and then the template DNA (about 0.5 μL-1 μL) of the above divided sample is taken by using a 96-well needle sampler and added to each PCR reaction system. Perform PCR reaction, and the PCR reaction program is: 94℃ pre-denaturation for 3 min, 95℃ denaturation for 15 s, 55℃ annealing for 20 s, 72℃ extension for 20 s, a total of 26-28 cycles, and 72℃ final extension for 2 min.
[0166] The sequence of the fragment amplified after the first round of PCR reaction in the present embodiment is:
[0167] ctcggagtgatcgcacTTATTGCAACCTCCTTCCTTATGTTCCCACCCTCATTATTTTCAGATATTCATTGATCAGCTGGAAAGTGTTCTGCAGGGACATGAATCCTCCATAGGGTAAATATCCTAGAGTTATATTTGTATCCTTAATGCATATGACCAATAATCATGTATTAACAACAAGCAATTTTTGTAATTGTTTATAAAGTATGGCATGTCCATCATccatccagcctctcag (underlined region is the targeting region)
[0168] The first round of PCR reaction products were diluted 10 times with double distilled water. 96 different OTUB1-Pos N - F primers (10 μM) were diluted to 1 μM with 0.3x TE and then dispensed into 96-well PCR plates. Then, according to the number of samples to be detected, the total PCR reaction solution was prepared, and the PCR reaction system in each well was 15 μL, and the contents of each component in the system were as follows: 2x Taq mix 7.5 μL, Plate-01 (10 μM) 0.08 μL (added to the reaction system with a final concentration of 0.05 μM), Lib-F (10 μM) and Lib-R (10 μM) each 0.4 μL (added to the reaction system with a final concentration of 0.26 μM), double distilled water 4.62 μL. The PCR reaction system was dispensed into each well of the 96-well PCR plate using a multichannel pipette or a general pipette, and then the above diluted 10 times first round PCR products were spotted using a 96-well needle sampler (about 0.5-1 μL) and added to each PCR reaction system. PCR reaction was performed, and the PCR reaction program was as follows: 94°C denaturation for 15 s, 55°C annealing for 20 s, 63°C extension for 3 s, 70°C extension for 10 s, 63°C extension for 3 s, 70°C extension for 5 s, a total of 8 cycles, 94°C denaturation for 15 s, 63°C annealing / extension for 10 s, 70°C extension for 10 s, 63°C annealing / extension for 5 s, 70°C annealing / extension for 5 s, a total of 10-12 cycles, 68°C final extension for 1 min. N - F primers (about 1 μL) were added to each PCR reaction system, and the above diluted 10 times first round PCR products (about 0.5-1 μL) were added to each PCR reaction system using a 96-well needle sampler. PCR reaction was performed, and the PCR reaction program was as follows: 94°C denaturation for 15 s, 55°C annealing for 20 s, 63°C extension for 3 s, 70°C extension for 10 s, 63°C extension for 3 s, 70°C extension for 5 s, a total of 8 cycles, 94°C denaturation for 15 s, 63°C annealing / extension for 10 s, 70°C extension for 10 s, 63°C annealing / extension for 5 s, 70°C annealing / extension for 5 s, a total of 10-12 cycles, 68°C final extension for 1 min.
[0169] The sequence of the second round of PCR reaction amplification products of this example is as follows: (underlined solid line region is the target region, italic part is the Pos N - F specific recognition sequence, bold font is Plate-01 specific recognition sequence, underlined dotted line part is library specific recognition sequence)
[0170] Each 5 μL of the second round of PCR reaction product was mixed in a 1.5 mL-5 mL centrifuge tube, 200 μL of the mixed product was taken and subjected to 1% agarose gel electrophoresis, and the target fragment was recovered and sent to a company for NGS.
[0171] The three-step method for building a NGS sample library of the present application is as follows:
[0172] The number of samples detected in this embodiment was 10. The samples to be detected were divided into 96-well PCR holes using a multichannel pipettor or a common pipettor. The PCR total reaction solution was prepared according to the number of samples to be detected, and the PCR reaction system in each hole was 15 μL, and the content of each component in the system was as follows: 2 × Taq mix 7.5 μL, 0.3 μL of forward primer T#F and reverse primer T#R, and 5.9 μL of double distilled water. The PCR reaction system was divided into 96-well PCR holes using a multichannel pipettor or a common pipettor, and then the template DNA (about 0.5 μL-1 μL) of the above divided sample was taken using a 96-well needle sampler and added to each PCR reaction system. PCR reaction was performed, and the PCR reaction program was as follows: 94°C pre-denaturation for 3 min, 95°C denaturation for 15 s, 55°C annealing for 20 s, 72°C extension for 20 s, a total of 26-28 cycles, and 72°C final extension for 2 min.
[0173] The sequence of the amplified fragment after the first round of PCR reaction in this embodiment was as follows:
[0174] ctcggagtgatcgcacTTATTGCAACCTCCTTCCTTATGTTCCCACCCTCATTATTTTCAGATATTCATTGATCAGCTGGAAAGTGTTCTGCAGGGACATGAATCCTCCATAGGGTAAATATCCTAGAGTTATATTTGTATCCTTAATGCATATGACCAATAATCATGTATTAACAACAAGCAATTTTTGTAATTGTTTATAAAGTATGGCATGTCCATCATccatccagcctctcag (underlined region is the target region)
[0175] The first round of PCR reaction product was diluted 10 times with double distilled water. 96 different Pos N- F primer (10 μM) was diluted to 3 μM with 0.3x TE and then aliquoted into 96-well PCR plates. Then, PCR total reaction solution was prepared according to the number of detection samples, and each well of PCR reaction system was 15 μL, and the content of each component in the system was as follows: 2x Taq mix 7.5 μL, Plate-01 (10 μM) 0.3 μL, double distilled water 5.9 μL. The PCR reaction system was aliquoted into each well of the 96-well PCR plate using a multichannel pipette or a common pipette, and then the above aliquoted Pos N - F primer (about 1 μL) was added to each PCR reaction system, and the above diluted 10 times first round PCR product (about 0.5 μL-1 μL) was added to each PCR reaction system using a 96-well needle sampler. PCR reaction was performed, and the PCR reaction program was as follows: 95℃ denaturation for 20s, 55℃ annealing for 20s, 72℃ extension for 25s, a total of 10 cycles, 94℃ denaturation for 20s, 60℃ annealing for 20s, 72℃ extension for 25s, a total of 10 cycles, 72℃ terminal extension for 2min.
[0176] The sequence of the fragment after the second round of PCR amplification in this example was as follows: (The underlined solid line region is the target region, and the italic part is OTUB1-Pos N - F specific recognition sequence, bold font is Plate-01 specific recognition sequence)
[0177] The above second round PCR reaction product was mixed with 2 μL (about 60-100 samples of second round PCR reaction amplification product were mixed as a group), and PCR total reaction solution was prepared according to the number of mixed groups, and each reaction system was 25 μL, and the components were as follows: 2x Taq mix 12.5 μL, Lib-F (10 μM) 0.5 μL, Lib-R (10 μM) 0.5 μL, double distilled water 10 μL. The total reaction solution was aliquoted into a PCR plate, and 1.5 μL of the second round PCR reaction mixture was added, and PCR reaction was performed. The PCR reaction program was as follows: 95℃ denaturation for 20s, 52℃ annealing for 20s, 72℃ extension for 30s, a total of 7 cycles, 96℃ denaturation for 20s, 63℃ annealing for 20s, 72℃ extension for 30s, a total of 8 cycles, 72℃ terminal extension for 2min.
[0178] The results of the third round of PCR amplification in this example are shown in Figure 3, and the sequence of the amplified fragment is as follows: (The underlined solid line region is the target region, and the italic part is OTUB1-Pos N - F specific recognition sequence, bold font is Plate-01 specific recognition sequence, underlined dotted part is library specific recognition sequence)
[0179] The third round of PCR reaction product was mixed in 15 μL of 1.5 mL-5 mL centrifuge tube, 200 μL of mixed product was taken, 1% agarose gel electrophoresis was used, the target fragment was recovered, and the company was sent for NGS.
[0180] Example 2. Detection of rice gene editing material results by HiDecode
[0181] The NGS results in Example 1 of the application were detected using the HiDecode module in SuperDecode, and the steps were as follows:
[0182] (1) Upload the required files, including the wild type sequence file of the gene, the target sequence file (optional), and the NGS result file.
[0183] (2) Set the fragment recognition sequence information, including OTUB1-Pos-A01 to OTUB1-Pos-A10 specific recognition sequence and Plate-01 specific recognition sequence.
[0184] (3) Set the mutation analysis parameters, including analysis mode, which can be divided into polyploid, diploid and low frequency mutation type sample, the material used in this embodiment is rice, which is diploid, so the parameter is selected as diploid; The output threshold is 0.25 for diploid material by default, 0.1 for polyploid, and 0.0025 for cell line sample. When the percentage of reads is greater than the threshold, it is recorded as the effective mutation of the sample; The number of threads can be adjusted to adjust the running speed of HiDecode.
[0185] (4) Click the Run key to perform mutation analysis.
[0186] The results are shown in Figure 3. Among the 10 samples detected by HiDecode, 1 sample had a wild type (wt) genotype, i.e. no editing occurred; 2 samples had heterozygous mutations (het), i.e. one strand was wild type and the other strand was edited to produce mutations; 7 samples had heterozygous mutations (bia), i.e. both strands of the sample were edited to produce variations, and the variation types were inconsistent.
[0187] Detection of rice protoplast editing results by HiDecode
[0188] In this embodiment, LbCas12a was used to edit the protoplast of rice, and the selected DNA template sequence (Protoplast-Target, SEQ ID NO. 8) was:
[0189] TAAAGAAGATTAATTAGAAAAATAGCCAAACGATTTGTAATATGCAACGGAGTGAGTAGAAGTAATCGCCCAGCCTCGCCAACGAGGCAACGAGACCCGTAATGCAACGATCGCATCTGCGTTTCAGGCGTCAGCCATGGCGTCTGCAGAGATGCCTGGATTGTCTCCGCAAGATCTGATCCATTTCATCTCCTTCTAGAAGCACAAGCGCCGCTCGGTATAAAGGCAGACGCATTGTCACAAATAGCTGCAGTGCACCAGAGTCACAGAAACACATCACACATTCGTGAGCTCAGCTTAGCCATGGATAACGCC (underlined solid line region is the target region)
[0190] The selected target sequence (LbCas12a-Target, SEQ ID NO. 9) is: ATCTCCTTCTAGAAGCACAAGCGC. The specific steps are as follows:
[0191] (1) The rice codon-optimized LbCas12a is loaded into the 1300 vector, and the crRNA is co-loaded into the 1300 vector containing the LbCas12a gene. The crRNA sequence (LbCas12a-crRNA, SEQ ID NO. 10) is: AATTTCTACTAAGTGTAGATATCTCCTTCTAGAAGCACAAGCGC (the underlined region is the target region)
[0192] (2) The rice seed husks are removed and placed in a sterilized small triangular flask. Add 75% ethanol and shake well for 1 min, then wash with sterilized double distilled water for 3 times. Add an appropriate amount of 2% sodium hypochlorite, place in a dark environment, 28°C, 150 rpm for 30 min, wash with sterilized double distilled water for 5 times, and then disinfect with an appropriate amount of 2% sodium hypochlorite. Use sterilized tweezers to put the seeds into 1 / 2MS medium, and cultivate in a dark environment for 13-15 days.
[0193] (3) The root and leaf parts of the seedlings are cut off, the leaf sheath part is reserved, the leaf sheath is cut into about 0.5mm small pieces using a sharp knife, and is placed in W5 solution, avoiding light for 30 min. After removing the W5 solution, an appropriate amount of Enzyme Solution solution is added, and is shaken in the dark at 28°C and 50 rpm for 4h.
[0194] (4) Remove Enzyme Solution solution, use W5 solution to wash leaf sheath tissue, and use 300 mesh and 400 mesh sieves to filter the washing liquid, collect the washing liquid in a 50 mL centrifuge tube, centrifuge at 150 g for 5 min, remove the supernatant, wash the precipitated cells with W5, centrifuge at 150 g for 5 min, remove the supernatant, and suspend the cells with 2 mL of W5. Centrifuge at 200 g for 10 min, suspend the cells with MMG, and adjust the concentration to 10 5 mL.
[0195] (5) Take 10 μg of the vector in 100 μL of the suspended cells, mix gently, add 110 μL of PEG4000, mix, and stand at 23°C in the dark for 20 min.
[0196] (6) Add 900 μL of W5 solution to wash, centrifuge at 300 g for 10 min, remove the supernatant, suspend the cells with 900 μL of W5 solution again, and incubate in the dark for 60 h.
[0197] (7) Extract DNA from the protoplasts after 60 h of culture, use the NGS library construction method in Example 1 to construct the NGS library of the protoplast sample, and send it to the company for sequencing.
[0198] (8) In the HiDecode module of SuperDecode, upload the required files, including the gene wild-type sequence file, the target sequence file (optional), and the NGS result file. Set the fragment recognition sequence information, including the Pos-A01 specific recognition sequence and the Plate-01 specific recognition sequence. Set the mutation analysis parameters, and the selected mode in this example is the low-frequency mutation type sample, that is, the mutation type with a percentage of reads greater than 0.0025. The results are shown in FIG. 4. After LbCas12a treatment, about 35.7% of the cells are wild type, and the remaining cells all show large fragment deletion, which is consistent with the mutation type after LbCas12a editing, indicating that HiDecode can quickly and accurately verify the editing results in protoplasts.
[0199] TGS library construction of large quantities of samples
[0200] In this example, the sample amplicon library of TGS was constructed for 22 Ehd1 gene promoter tiled deletion rice editing materials. The Ehd1 promoter wild-type sequence (Ehd1 promoter) is shown as SEQ ID NO. 11. The specific primers were designed at the positions 200 bp upstream of the first detection site and 200 bp downstream of the last detection site. The specific amplification primers of the sample amplicon designed in this example are as follows:
[0201] Forward primer Target-F (SEQ ID NO. 12): 5'-atcgcctggctccacgctccgagttTCCCTTCATTCCGATGAGGCCCACGATT-3' (10 μM)
[0202] Reverse primer Target-R (SEQ ID NO. 13): 5'-cctggctccacgctccgagttCCCTTGTAGCTGCACTTCAGAAGTAAATCTTCC-3' (10 μM)
[0203] Plate-02 (SEQ ID NO. 14) (10 μM): (Plate-02 specific recognition sequence in bold)
[0204] Ehd1-Pos N -F (SEQ ID NO. 15) (10 μM): (Ehd1-Pos specific recognition sequence in italic) N -F (SEQ ID NO. 15) (10 μM):
[0205] The TGS mass sample library construction process of the application is shown in FIG. 6, which is mainly divided into two PCR reactions, and the main steps are as follows:
[0206] (1) The sample to be detected is divided into 96-well PCR holes using a multichannel pipettor or a common pipettor. The total PCR reaction solution is prepared according to the number of samples to be detected, and the PCR reaction system of each hole is 15 μL. Each PCR system component is: 7.5 μL of 2x Phanta Max Buffer, 0.3 μL of dNTP, 0.3 μL of Phanta Max Polymerase, 0.6 μL of forward primer Target-F and 0.6 μL of reverse primer Target-R, and 4.7 μL of double distilled water. The PCR reaction system is divided into 96-well PCR holes using a multichannel pipettor or a common pipettor, and then the DNA (about 0.5-1 μL) of the above-mentioned sample to be detected is added to each PCR reaction system using a 96-well needle sampler, and a PCR amplification reaction is performed. The PCR amplification program is: 94°C pre-denaturation for 3 min, 94°C denaturation for 15 s, 58°C annealing for 15 s, 72°C extension for 2 min, a total of 30 cycles, and 72°C final extension for 5 min.
[0207] (2) The first round of PCR reaction product is diluted ten times. 96 different Ehd1-Pos N- F primer (10 μM) was diluted to 3 μM using 0.3x TE and then aliquoted into 96-well PCR plates. Then, PCR total reaction solution was prepared according to the number of detection samples, and each well of the PCR reaction system was 15 μL, and the content of each component in the system was: 2x Phanta Max Buffer 7.5 μL, 0.3 μL of dNTP, 0.3 μL of Phanta Max Polymerase, 0.6 μL of Plate-02 primer, and 4.3 μL of double-distilled water. The PCR reaction system was aliquoted into each well of the 96-well PCR plate using a multichannel pipette or a common pipette, and then the above aliquoted Ehd1-Pos N - F primer (about 1 μL) was added to each PCR reaction system, and the above diluted 10 times first round PCR product (about 0.5 μL-1 μL) was added to each PCR reaction system using a 96-well needle sampler. PCR reaction was performed, and the PCR reaction program was: 94°C pre-denaturation for 3 min, 94°C denaturation for 15 s, 58°C annealing for 15 s, 72°C extension for 2 min, a total of 30 cycles, and 72°C final extension for 5 min.
[0208] 5 μL of the second round PCR reaction product was mixed in a 1.5 mL-5 mL centrifuge tube, 200 μL of the mixed product was taken, and 1% agarose gel electrophoresis was performed. The target fragment was recovered and sent to the company for TGS library construction and sequencing.
[0209] LaDecode was used to detect the tiled deletion
[0210] The TGS results of the tiled deletion of the Ehd1 promoter sequence (Ehd1 promoter, SEQ ID NO. 11) in Example 5 of the application were detected using the LaDecode module in SuperDecode, and the steps were as follows:
[0211] (1) Upload the required files, including the wild-type sequence file of the gene, the target sequence file (optional), and the TGS result file.
[0212] (2) Set the fragment recognition sequence information, including Ehd1-Pos N - F specific recognition sequence and Plate-02 specific recognition sequence.
[0213] (3) Set mutation analysis parameters, including sample analysis mode, which can be divided into polyploidy, diploidy and low frequency mutation analysis mode, the material used in this embodiment is rice, which is diploid, so the parameter is selected as diploid; the output threshold value, the default value of the diploid material is 0.25, the default value of the polyploidy is 0.1, and the default value of the low frequency mutation sample is 0.0025; when the percentage content of reads is greater than the threshold value, it is recorded as the effective mutation of the sample; the number of threads can change the running speed of LaDecode by adjusting the number of threads.
[0214] (4) Click the Run key to perform mutation analysis.
[0215] The results are shown in Figures 7 and 8.
[0216] The present application is based on the method of PCR amplification, the construction of amplicon of gene editing material, and the introduction of specific recognition sequence, which realizes the NGS and TGS library construction of a large number of samples quickly and at low cost; the present application develops a software SuperDecode for gene editing material detection, which can perform mutation analysis on the results of Sanger, NGS and TGS commonly used at present, and can detect all types of mutations in the sample, including large fragment deletion, sequence insertion, single base mutation, chimeric mutation, etc.; the SuperDecode developed by the present application is a Window platform software, which can quickly read a large number of sample sequencing result files and perform mutation analysis compared with web version tools. For Linux system tools, SuperDecode provides a convenient and simple operation interface, which is convenient for researchers to use, and greatly improves the efficiency of gene editing material detection.
[0217] The above description is only a specific embodiment of the present application, which enables those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features applied herein.
Claims
1. A method for detection of gene editing material, characterized in that, The application relates to a method for analyzing mutations in a gene fragment, comprising the following steps: constructing multiple specific recognition sequences; sequencing the gene fragment according to the specific recognition sequences; analyzing mutations in the sequencing results by using SuperDecode software.
2. The method for detecting gene editing material according to claim 1, wherein, The step of constructing multiple specific recognition sequences comprises the following steps: adding position-barcode, plate-barcode and Library-barcode specific recognition sequence libraries to sample amplicons; the position-barcode sequence library comprises 192 specific position-barcode sequences, 96 specific recognition sequences for NGS library construction and 96 specific recognition sequences for TGS library construction; the plate-barcode sequence library comprises 102 specific plate-barcode sequences, 96 specific recognition sequences for NGS library construction and 6 specific recognition sequences for TGS library construction; the Library-barcode sequence library comprises 20 specific Library-barcode sequences, all of which are used for NGS library construction.
3. The method for detecting gene editing material according to claim 1, wherein, The step of sequencing the gene fragment according to the specific recognition sequences comprises the following steps: Sanger sequencing, NGS sequencing and third-generation TGS of the gene fragment.
4. The method for detecting gene editing material according to claim 1, wherein, The step of analyzing mutations in the sequencing results by using SuperDecode software comprises the following steps of analyzing mutations in Sanger sequencing results: obtaining first project information of current mutation analysis according to user input files and settings; reading overlapping wave peaks in the Sanger sequencing results by using the sequencing overlapping wave peak decoding principle of DSD; decoding the Sanger sequencing results of the amplicon, reading the variation type of the target sequence of the sample, based on the sequence analysis software searching for the corresponding sequence in the wild type sequence.
5. The method for detecting gene editing material according to claim 1, wherein, The step of analyzing mutations in the sequencing results by using SuperDecode software comprises the following steps of analyzing mutations in NGS results: obtaining second project information of current mutation analysis according to user input files and settings; merging and completing paired reads in double-end sequencing into complete sequences, and splitting the library based on the specific recognition sequence of the sample; corresponding the reads to a specific sample, and obtaining the variation information of the sample by aligning the reads of the sample with a reference sequence.
6. The method for detecting gene editing material according to claim 1, wherein, The step of analyzing mutations in the sequencing results by using SuperDecode software comprises the following steps of analyzing mutations in TGS results: obtaining third project information of current mutation analysis according to user input files and settings; splitting the library based on the specific recognition sequence of the sample, and corresponding the reads to a specific sample; extracting sequences common to all reads of the specific sample to remove random mutations caused by sequencing, and finally aligning the reads of the sample with a reference sequence to obtain sample variation information.
7. The method for detecting gene editing material according to claim 4, wherein, The step of obtaining first project information of current mutation analysis according to user input files and settings comprises the following steps: the first project information comprises wild type sequence information of the amplicon, a Sanger sequencing result file of the sample amplicon and optional first setting conditions.
8. The method for detecting gene editing material according to claim 5, wherein, The second project information of the current mutation analysis is obtained according to the file and the setting input by the user, and includes: The second project information includes: amplicon wild type sequence information, gene editing targeted sequence information, sample NGS result information, specific recognition sequence information corresponding to each sample, and optional second setting conditions; The optional second setting conditions are settings of mutation analysis parameters, including: Sample analysis mode, selected from diploid, polyploid and low frequency; Output threshold, used for setting the threshold of result output; Running thread number, used for setting the thread number used by the program to run, and adjusting the running speed of HiDecode decoding.
9. The method for detecting gene editing material according to claim 5, wherein, The third project information of the current mutation analysis is obtained according to the file and the setting input by the user, and includes: The third project information includes: amplicon wild type sequence information, gene editing targeted sequence information, sample TGS result information, specific recognition sequence information corresponding to each sample, and optional third setting conditions; The optional third setting conditions are settings of mutation analysis parameters, including: Sample analysis mode, selected from diploid, polyploid and low frequency; Output threshold, used for setting the threshold of result output; Running thread number, used for setting the thread number used by the program to run, and adjusting the running speed of LaDecode decoding.
10. A system for gene editing material detection, comprising: It includes: A data acquisition module for constructing a plurality of specific recognition sequences; A sequencing module for sequencing gene fragments according to the specific recognition sequences; An analysis module for performing mutation analysis on the sequencing results by using SuperDecode software.
Citation Information
Patent Citations
Coding PCR second-generation sequencing library establishing method, kit and detection method
CN108504649A
Method for quickly and precisely determining gene editing mutation situations and application of method
CN111676276A
Method and system for detecting gene editing material
CN119252331A
Method of detecting mutation
WO2005010184A1