Primer set, kit and application thereof for detecting IDS gene mutation

Through the combination method of multiple long fragment PCR amplification and long read sequencing, a specific primer set was designed to solve the problem that the prior art cannot comprehensively detect multiple mutations in the IDS type II pathogenic gene of mucopolysaccharide storage disease, and achieve rapid and accurate detection of multiple mutations in the IDS gene.

CN119530378BActive Publication Date: 2025-05-06BERRYGENOMICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510104376.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art cannot comprehensively, accurately and rapidly detect multiple mutations in the IDS type II pathogenic gene of mucopolysaccharide storage disorder, especially complex gene rearrangements and large fragment deletions.

Method used

A combination of multiple long fragment PCR amplification and long read sequencing was used to design a specific primer set covering the full length of the IDS gene, the full length of the pseudogene IDS2 and related regions to achieve comprehensive detection of multiple mutations in the IDS gene.

Benefits of technology

Comprehensive, accurate and rapid detection of point mutations, small fragment insertions/deletions, large fragment insertions/deletions and structural variations of the IDS gene are achieved, covering the large fragment deletion in the region of about 1 Mb of the proximal centromere of the IDS gene to about 1 Mb of the centromere distal to the IDS2 gene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention provides a detection IDS Primer set, kit and application thereof for gene mutation. The primer set provided by the present invention comprises a primer pair of IDS-1-F and IDS-1-R, a primer pair of IDS-2-F and IDS-2-R, and a primer pair of IDS2-F and IDS2-R, and further comprises an IDS-Gap Mix primer set. The present invention combines the characteristics of long-fragment PCR amplification and long-read sequencing platform to achieve comprehensive, accurate, rapid and high-throughput simultaneous detection. IDS Multiple mutations of genes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of gene detection technology, and more particularly to a method for detecting a gene using a long-read sequencing platform. IDS Primer sets and kits for gene mutation and their applications. Background Art

[0002] Mucopolysaccharidosis type II (MPS II, OMIM 309900), also known as Hunter syndrome, is a rare X-linked recessive hereditary lysosomal storage disease. IDS Gene mutations lead to defective activity of iduronate-2-sulfatase (I2S), which in turn causes a breakdown disorder of glycosaminoglycans (GAGs). GAGs that are not completely broken down are pathologically stored in the lysosomes of various tissues and organs of the body, which in turn causes structural and functional abnormalities of multiple organs and systems. Clinical manifestations include multi-system involvement, such as developmental delay, skeletal abnormalities, and heart and liver damage.

[0003] IDS The gene is located on chromosome Xq28, is about 24 kb long, contains 9 exons and 8 introns, and encodes 550 amino acids. About 20 kb distal to the centromere, there is a IDS2 pseudogene. IDS2 Contains two IDS The first homologous region is homologous to IDS Exons 2, 3 and introns 2, 3 of the gene have high homology, and the second homologous region is IDS Intron 7 of the gene has a high degree of homology, thus easily causing complex rearrangements between true and false genes.

[0004] Currently, 881 gene variants have been reported in the Human Gene Variation Database (HGMD). IDS Gene mutations include point mutations, splice site mutations, small or large insertion / deletions, and gene rearrangements, among which point mutations are the most common, large insertion / deletions account for about 9%, and gene rearrangements account for about 3%. Most of the reported gene rearrangements are IDS Genes and their pseudogenes IDS2 Intrachromosomal homologous recombination between IDS Intron 7 of the gene IDS2 Inversions occur in genomic regions between homologous regions in the genome without obvious deletions or insertions. IDS Genes and their pseudogenes IDS2 Multiple homologous recombination may occur simultaneously, resulting inIDS Gene fragments showed multiple inversions, large deletions, or partial pseudogene insertions. In addition, large deletions unrelated to homologous recombination were occasionally found. These large deletions involved IDS All or part of a gene and extend to adjacent genes.

[0005] The diagnosis of MPS II is mainly based on clinical manifestations, imaging examinations, and various laboratory test results, such as urine GAG ​​detection, I2S enzyme activity detection, and IDS Gene mutation detection, etc. Among them, the qualitative and quantitative detection of urine GAG ​​is helpful for the preliminary screening of MPS II, and the detection of I2S enzyme activity is of great significance for the diagnosis of MPS II. However, both methods have certain defects. First, they cannot accurately diagnose female carriers because female carriers may have normal I2S enzyme activity related to non-random inactivation of X chromosomes; second, they cannot directly obtain IDS The specific mutation site of the gene.

[0006] IDS Gene mutation detection is an important basis for diagnosing MPS II and determining carriers. Common detection methods include first-generation sequencing (Sanger sequencing) and second-generation sequencing (NGS). Both first-generation sequencing and second-generation sequencing can detect IDS Gene point mutations and small fragment insertions / deletions, but the first-generation sequencing has a low throughput and cannot detect unknown mutations; compared with the first-generation sequencing, the second-generation sequencing has a higher sequencing throughput and a higher positive detection rate, but due to the limited read length, it cannot detect IDS Gene rearrangement cannot effectively distinguish true from false genes. In addition, whole exome sequencing (WES) in the second-generation sequencing technology does not detect gene introns, so it is easy to miss gene variants located in deep intronic regions. Multiplex ligation-dependent probe amplification (MLPA) can detect IDS However, this method is cumbersome to operate, cannot detect a large number of samples at the same time, and has the risk of missing unknown mutations, and cannot accurately determine the exact recombination site of large fragment insertion / deletion. IDS Pseudogenes or other types of complex rearrangements usually require a combination of the above methods to detect and confirm the recombination sites.

[0007] Currently, methods such as biochemical testing, first-generation sequencing or second-generation sequencing, and MLPA can achieve IDS Gene point mutation and large fragment deletion detection, but there are still the following limitations:

[0008] 1. It is impossible to achieve simultaneous detection in the same system IDSAll types of gene mutations, especially complex gene rearrangements;

[0009] 2. Biochemical testing cannot be used to detect female carriers and cannot identify the specific mutation type;

[0010] 3. The first-generation sequencing throughput is low; the second-generation sequencing cannot detect due to the limited read length IDS Gene rearrangement cannot effectively distinguish true and false genes, and comprehensive testing must be combined with different library construction methods. It is basically impossible to test a single gene, and the operation is cumbersome and costly;

[0011] 4. MLPA can only detect a few known point mutations and large insertions / deletions, but there is a risk of missing unknown mutations. It is also unable to determine the exact recombination site of large insertions / deletions, and the operation is cumbersome.

[0012] Therefore, it is necessary to develop a comprehensive, accurate and rapid detection IDS New technologies for gene mutations can make up for the shortcomings of existing methods and have important clinical and scientific research significance. Summary of the invention

[0013] The purpose of the present invention is to solve the problems of incomplete coverage of the pathogenic gene detection of mucopolysaccharidosis type II, missed detection in the specific NGS detection process, and inability of NGS technology to detect gene rearrangements. The present invention provides a detection method based on multiple long-fragment PCR amplification and long-read sequencing. IDS Full-length gene, pseudogene IDS2 Full length and IDS The centromere of the gene is about 1 Mb to IDS2 A specific primer set for the approximately 1 Mb region distal to the centromere of the gene was used to identify the pathogenic gene of mucopolysaccharidosis type II. IDS The method provided by the present invention combines the advantages of multiple long-fragment PCR amplification and long-read sequencing to achieve comprehensive, accurate and rapid simultaneous detection of MPS II pathogenic genes in multiple samples. IDS All mutation types are relevant.

[0014] According to a first aspect of the present invention, there is provided a method for detecting IDS A primer set for gene mutation, the primer set comprising the following primers:

[0015] (1) a primer pair of IDS-1-F and IDS-1-R, wherein IDS-1-F is selected from at least one of the primers shown in SEQ ID NO: 1 and SEQ ID NO: 2, and IDS-1-R is selected from at least one of the primers shown in SEQ ID NO: 4 and SEQ ID NO: 5;

[0016] (2) a primer pair of IDS-2-F and IDS-2-R, wherein IDS-2-F is selected from at least one of the primers shown in SEQ ID NO: 8 and SEQ ID NO: 9, and IDS-2-R is selected from at least one of the primers shown in SEQ ID NO: 11 and SEQ ID NO: 12; and

[0017] (3) a primer pair of IDS2-F and IDS2-R, wherein IDS2-F is selected from at least one of the primers shown in SEQ ID NO: 13 and SEQ ID NO: 15, and IDS2-R is selected from at least one of the primers shown in SEQ ID NO: 16 and SEQ ID NO: 18.

[0018] In one embodiment, the primer set further comprises an IDS-Gap Mix primer set, wherein the IDS-Gap Mix primer set comprises primers shown as SEQ ID NOs: 19-218.

[0019] According to a second aspect of the present invention, there is provided a method for detecting IDS A kit for detecting gene mutation, comprising the primer set according to the first aspect above.

[0020] According to a third aspect of the present invention, there is provided a method for detecting IDS The method for gene mutation comprises the following steps:

[0021] (1) Obtain samples from subjects;

[0022] (2) using the primer set according to the first aspect or the kit according to the second aspect to amplify the IDS Gene fragment, obtain the target fragment amplification product;

[0023] (3) constructing a long-read sequencing library based on the amplified product;

[0024] (4) performing long-read sequencing based on the long-read sequencing library to obtain a long-read sequencing result;

[0025] (5) Analyze and determine based on the long read sequencing results IDS Gene mutation.

[0026] According to a fourth aspect of the present invention, there is provided a method for detecting IDS The gene mutation system includes the following modules:

[0027] (1) Amplification module: used to amplify the sample from the subject by long-range PCR using the primer set according to the first aspect or the kit according to the second aspect. IDS Gene fragment, obtain the target fragment amplification product;

[0028] (2) Library construction module: used to construct a long-read sequencing library based on the amplification product;

[0029] (3) Sequencing module: used to perform long-read sequencing based on the long-read sequencing library to obtain long-read sequencing results;

[0030] (4) Mutation analysis module: used to analyze and determine based on long-read sequencing results IDS Gene mutation.

[0031] In one embodiment, the long-range PCR is performed in a single reaction tube.

[0032] According to a fifth aspect of the present invention, there is provided an application of the primer set according to the first aspect of the present invention or the kit according to the second aspect of the present invention in any of the following aspects:

[0033] 1) Detection of IDS gene mutations;

[0034] 2) Detection of IDS gene rearrangement;

[0035] 3) Genetic diagnosis of mucopolysaccharidosis type II.

[0036] According to a sixth aspect of the present invention, there is provided use of the primer set according to the first aspect of the present invention or the kit according to the second aspect of the present invention in the preparation of any of the following products:

[0037] 1) Products for detecting IDS gene mutations;

[0038] 2) Products for detecting IDS gene rearrangements;

[0039] 3) Products used for genetic diagnosis of mucopolysaccharidosis type II.

[0040] The method provided by the present invention is based on a specific combination of multiple long-fragment PCR amplification and long-read sequencing, which can achieve high specificity, accuracy and rapidity in the simultaneous detection of mucopolysaccharidosis type II pathogenic genes in multiple samples. IDS Various related mutations.

[0041] The excellent technical effects of the method and kit of the present invention are mainly in the following aspects:

[0042] (1) Wide detection range. The present invention can simultaneously detect IDS Point mutations, small insertion / deletions, structural variations, and IDS The centromere of the gene is about 1 Mb to IDS2 A large deletion was found within a region approximately 1 Mb distal to the centromere of the gene.

[0043] (2) Detection of multiple mutation types with a single kit. Traditional methods require a detection system to be set up for each mutation type, while the present invention achieves simultaneous detection of multiple mutations in one reaction system, including IDS Point mutations, small insertion / deletions, structural variations, and IDS The centromere of the gene is about 1 Mb to IDS2 A large deletion was found in the approximately 1Mb region distal to the centromere of the gene.

[0044] (3) Low false positive and missed detection rates. Compared with the currently commonly used biochemical detection for MPS II, the method disclosed in the present invention can accurately detect wild-type and mutant chromosomes in female carriers; compared with the molecular diagnosis of MPS II, where first-generation sequencing cannot be applied on a large scale and has a small detection range, second-generation sequencing is more suitable for IDS The process of gene mutation detection is more complicated, and the judgment of structural variation is unclear. MLPA can only detect known point mutations and large fragment deletions. The process is complicated and cannot be applied on a large scale. The method involved in the present invention directly amplifies IDS Genes and pseudogenes IDS2 The complete sequence of the gene, as well as the unknown large deletion fragments, can detect all the above mutation types at the same time, and the judgment is more accurate and comprehensive.

[0045] (4) Sample diversification. The template used for PCR can be peripheral blood, dried blood spots or extracted genomic DNA, or human cell lines or other specific tissues.

[0046] (5) High-throughput detection. Long-read sequencing can use 384 types of DNA barcode adapters, and more types of DNA barcode adapters can be designed as needed. Alternatively, a dual DNA barcode system with primers carrying DNA barcodes and adapters carrying DNA barcodes can be used to achieve more DNA barcode combinations. The high-throughput characteristics of the long-read sequencing platform determine that high-throughput sample detection can be achieved.

[0047] (6) High accuracy. PacBio's dumbbell-shaped library can be read multiple times during sequencing, and the base accuracy of the sequencing results after correction is greater than 99%. In addition, PacBio sequencing errors are random, and the base accuracy is greater than 99.9% after correction through sequencing depth. Therefore, gene mutations within the primer detection range can be accurately read.

[0048] (7) Flexible detection time. The Nanopore platform can generate data within minutes, and data analysis can be started within minutes or hours depending on the actual data volume. When the detection timeliness requirements are relatively high, the Nanopore platform has a time advantage. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 : IDS Schematic diagram of the design of primer sets for gene detection.

[0050] Figure 2 : Agarose gel electrophoresis results of the sample target fragments obtained by long-range PCR amplification using different primer combinations.

[0051] Figure 3 : Representative IDS PacBio sequencing results of gene exon point mutation samples.

[0052] Figure 4 : Representative IDS PacBio sequencing results of gene intron point mutation samples.

[0053] Figure 5 : Representative IDS PacBio sequencing results of intragenic duplicate samples.

[0054] Figure 6 : Representative IDS PacBio sequencing results of intragenic deletion samples.

[0055] Figure 7 : Representative IDS PacBio sequencing results of gene inversion samples.

[0056] Figure 8 : PacBio sequencing results of representative MPS II-related 8 kb large fragment deletion samples.

[0057] Figure 9 : PacBio sequencing results of a representative MPS II-related 210 kb large deletion sample. DETAILED DESCRIPTION

[0058] As described above, commonly used IDSGenetic variation detection methods, such as Sanger sequencing and high-throughput sequencing, still have problems such as incomplete coverage of pathogenic gene detection, missed detections in specific NGS testing processes, and inability of NGS testing to detect gene rearrangements, which lead to missed detections and false detections in clinical practice.

[0059] In order to at least partially solve the above problems and one or more of other potential problems, the present invention provides a method based on multiple long-fragment PCR amplification and long-read sequencing to detect the pathogenic gene of mucopolysaccharidosis type II. IDS Multiple related mutations. Multiple long-range PCR amplification is achieved in a single reaction tube for simultaneous detection IDS Point mutations, small insertion / deletions, structural variations, and IDS The centromere is about 1 Mb to IDS2 The large fragment in the approximately 1 Mb region distal to the centromere of the gene is missing. Combined with the characteristics of the long-read sequencing platform, the present invention can achieve comprehensive, accurate, rapid and high-throughput simultaneous detection of multiple IDS Gene mutation. The method of the present invention is simple to operate, and the quality of multiple long-fragment PCR and long-read sequencing library is reliable and highly reproducible, which is conducive to the application of long-read sequencing technology in clinical detection.

[0060] Embodiments of the present invention will be described in more detail below. However, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, which are instead provided to provide a more thorough and complete understanding of the present invention. It should also be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.

[0061] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as those generally understood by those of ordinary skill in the art to which the present invention belongs. When there is a contradiction, the definition in this specification shall prevail. In addition, unless the context otherwise requires, the term in the singular form shall include the plural form, and the term in the plural form shall include the singular form. More specifically, as used in this specification and the appended claims, unless the context otherwise clearly indicates, the singular forms "a", "one" and "the" include plural indicators. In the present invention, unless otherwise specified, the use of "or" means "and / or". In addition, the use of the term "comprising" and other forms (such as "including" and "containing") is not restrictive. In addition, the range provided in the specification and the appended claims includes all values ​​between the endpoints and the endpoints.

[0062] Generally, terms relating to cell and tissue culture, molecular biology, immunology, microbiology, genetics and protein and nucleic acid chemistry and hybridization described herein and their techniques are those well known and commonly used in the art. Unless otherwise indicated, the methods and techniques of the present invention are generally performed according to conventional methods well known in the art and as described in various general and more specific references cited and discussed throughout this specification. See, e.g., Sambrook J. & Russell D. Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2000); Abbas et al., Cellular and Molecular Immunology, 6th ed., WB Saunders Company (2010); Harlow and Lane Using Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (1998); Ausubel et al., Short Protocols in Molecular Biology: A Compendium of Methods from Current Protocols in Molecular Biology, Wiley, John & Sons, Inc. (2002); and Coligan et al., Short Protocols in Protein Science, Wiley, John & Sons, Inc. (2003). In addition, any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0063] Definition

[0064] Further, in the following specification, reference will be made to a number of expressions, which are defined as having the following meanings.

[0065] As used herein, the term "genome" or "genomic DNA" refers to the complete genetic information of an organism or virus expressed as a nucleic acid sequence. Preferably, the genome or genomic DNA refers to the chromosomal DNA of the cell nucleus.

[0066] As used herein, the term "reference genome" refers to any specific known genome sequence, whether partial or complete, of any organism or virus that can be used to reference an identified sequence from a subject. For example, exemplary reference genomes for human subjects, as well as many other organisms, can be found at the National Center for Biotechnology Information (NCBI). For example, a reference genome can be a human reference genome, such as GRCh37 or GRCh38.

[0067] As used herein, the term “ IDS "Gene" has the general meaning in the art and refers to the gene encoding iduronate-2-sulfatase (NP_000193.1). In humans, IDS The gene ID of the gene is 3423, and it is located at positions 149,476,988 to 149,505,306 of chromosome X NC_000023.11, using GRCh38 as the reference genome.

[0068] As used herein, the term "flanking" refers to a DNA sequence extending on either side of a particular DNA sequence, locus, or gene. The flanking region may be "upstream" (i.e., 5' end) or downstream (i.e., 3' end) of a particular DNA sequence, locus, or gene. Unless otherwise specified, the length of the flanking region is not specifically limited. For example, the flanking region may contain several base pairs, tens of base pairs, hundreds of base pairs, thousands of base pairs, or tens of thousands of base pairs.

[0069] As used herein, the term "primer" can be a natural or synthetic oligonucleotide that can serve as a starting point for nucleic acid synthesis after forming a duplex with a polynucleotide template and extending from its 3' end along the template to form an extended duplex. The nucleotides added to the 3' end of the extension product during the extension process are determined by the sequence of the template polynucleotide. Primers can usually be extended by a polymerase such as a nucleic acid polymerase. In the present invention, the term "primer pair" refers to a group of forward primers and reverse primers that can hybridize with the double strands of a target DNA molecule or the flanking regions at both ends of a nucleotide sequence to be amplified in a target DNA molecule and can initiate amplification of a target DNA molecule or a nucleotide sequence to be amplified. The term "primer set" refers to a combination containing at least two primer pairs. The term "primer range" or "primer coverage" refers to a nucleic acid sequence region between the 3' end hybridization site of a forward primer and the 3' end hybridization site of a reverse primer in a template polynucleotide, including all base sequences located within the region.

[0070] As used herein, the term "mutation" or "variation" refers to a change in a nucleotide sequence compared to a reference genome. Mutations may involve entire chromosomes (e.g., aneuploidy). Mutations may involve large fragments of DNA (e.g., structural variation). Mutations may involve small fragments of DNA. Examples of mutations involving small fragments of DNA include, for example, point mutations or single nucleotide polymorphisms, multinucleotide polymorphisms, insertions (e.g., adding one or more nucleotides at a locus), multinucleotide changes, deletions (e.g., losing one or more nucleotides at a locus), and inversions (e.g., reversal of one or more nucleotide sequences).

[0071] As used herein, the term "point mutation" refers to a substitution, insertion or deletion of a single nucleotide at a specific position in the genome.

[0072] As used herein, the term "small insertion / deletion" or "InDel" refers to the insertion or deletion of a nucleotide fragment less than 50 base pairs in length in a nucleotide sequence.

[0073] As used herein, the term "structural variation" refers to the structural differences of chromosomes with a length of more than 50 base pairs in a chromosome. Structural variation can be a deletion, duplication, copy number variation, insertion, inversion, translocation, or a combination thereof. Structural variation can also be a change in the arrangement or order of genomic fragments compared to a reference genome, such as gene rearrangement. According to certain embodiments, structural variation affects a sequence length of at least about 50 bases, preferably at least about 100 bases, more preferably at least about 1Kb (=1000 bases). According to certain embodiments, structural variation affects a sequence length of at most 300Mb (megabase=1,000,000 bases), for example at most 30Mb, for example at most 3Mb.

[0074] As used herein, the term "large insertion / deletion" refers to the insertion or deletion of a nucleotide fragment of more than 100 base pairs in length in a nucleotide sequence. IDS Large insertions or deletions found in genes can be thousands, tens of thousands, or hundreds of thousands of base pairs in length.

[0075] As used herein, the term "deep intron" refers to intronic regions that are more than 100 base pairs away from exon-intron boundaries.

[0076] As used herein, the term "gene rearrangement" generally refers to a change in gene expression products and / or gene transcription patterns by rearrangement of gene coding sequences. Gene rearrangement is a repair process of double-strand breaks in DNA, in which complex conversion-type movements of repeat units occur within or between genes. In some embodiments, "complex rearrangement" refers to a gene rearrangement event involving more than two breakpoints.

[0077] As used herein, the term "long read sequencing" (also referred to as "third generation sequencing") generally refers to any sequencing method that can generate a significantly longer sequencing read length (e.g., >1,000bp) than second generation sequencing. In one embodiment, the method provided herein involves the use of long read sequencing. Non-limiting examples of long read sequencing systems include systems developed by Pacific Biosciences, Oxford Nanopore Technology, Quantapore, Stratos, and Helicos. In some cases, the long read sequencing method is single molecule real-time (Single Molecule Real-Time, SMRT) sequencing developed by Pacific Biosciences. In other cases, the long read sequencing method is nanopore sequencing developed by Oxford Nanopore Technology. In some cases, long read sequencing covers any long read sequencing method or system (e.g., third generation sequencing method or system) currently being developed or developed in the future.

[0078] As used herein, the terms "DNA barcode", "barcode", "DNA tag", "tag" or "Barcode" can be a known sequence used to associate a polynucleotide fragment with an input polynucleotide or a target polynucleotide from which it is produced. The barcode sequence can be a sequence of synthetic nucleotides or natural nucleotides. The barcode sequence can be contained within a linker sequence so that the barcode sequence is included in the sequencing read. The length of each barcode sequence may include at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more nucleotides. In some cases, the barcode sequences can have sufficient length and can be sufficiently different from each other to allow the samples associated with them to be identified based on the barcode sequence. In some cases, the barcode sequence is used to label and subsequently identify the "original" nucleic acid molecule (a nucleic acid molecule present in a sample from a subject). In some cases, a barcode sequence or a combination of barcode sequences is used in conjunction with endogenous sequence information to identify the original nucleic acid molecule.

[0079] As used herein, the term "biological sample" or "sample" encompasses any sample obtained from a biological source. As a non-limiting example, a biological sample may include blood, amniotic fluid, serum, plasma, liquid or tissue biopsy samples, joint fluid, sweat, saliva, urine, feces, cerebrospinal fluid, ascites, pleural effusion, bile, pancreatic fluid, epidermal samples, skin samples, cheek swabs, sperm, gametes, blastocyst cells, cultured cells, bone marrow samples and / or chorionic villi. Convenient biological samples can be obtained by, for example, scraping cells from the buccal cavity surface. The term "biological sample" encompasses samples that have been processed to release nucleic acids or proteins or otherwise make them available for detection as described herein. The term "biological sample" also includes cell-free nucleic acids that may be present in a sample (e.g., plasma or amniotic fluid). For example, a biological sample may include cDNA obtained by reverse transcription of RNA from cells in a biological sample. Biological samples may be obtained from life stages, such as fetuses, youth, adults, etc. Fixed or frozen tissues may also be used.

[0080] Primer set

[0081] In order to at least partially solve the above problems and one or more of other potential problems, the first exemplary embodiment of the present invention proposes a method for detecting IDS A primer set for gene mutation, the primer set comprising the following primers:

[0082] (1) For amplification IDS A pair of IDS-1-F and IDS-1-R primers for exon 7 to exon 9 of a gene and the flanking introns at both ends thereof, wherein the IDS-1-F is selected from at least one of the primers shown in SEQ ID NO: 1 and SEQ ID NO: 2, and the IDS-1-R is selected from at least one of the primers shown in SEQ ID NO: 4 and SEQ ID NO: 5;

[0083] (2) For amplification IDS A pair of IDS-2-F and IDS-2-R primers for exon 1 to exon 6 of a gene and introns flanking both ends thereof, wherein the IDS-2-F is selected from at least one of the primers shown in SEQ ID NO: 8 and SEQ ID NO: 9, and the IDS-2-R is selected from at least one of the primers shown in SEQ ID NO: 11 and SEQ ID NO: 12; and

[0084] (3) For amplification IDS2The invention further comprises a pair of IDS2-F and IDS2-R primers, wherein the IDS2-F is selected from at least one of the primers shown in SEQ ID NO: 13 and SEQ ID NO: 15, and the IDS2-R is selected from at least one of the primers shown in SEQ ID NO: 16 and SEQ ID NO: 18.

[0085] In a preferred embodiment, the method for detecting IDS The primer set for gene mutation contains the following primers:

[0086] (1) a primer pair of IDS-1-F and IDS-1-R, wherein the primer pair of IDS-1-F and IDS-1-R is selected from at least one of the primer pair of SEQ ID NO: 1 and SEQ ID NO: 4 and the primer pair of SEQ ID NO: 2 and SEQ ID NO: 5;

[0087] (2) a primer pair of IDS-2-F and IDS-2-R, wherein the primer pair of IDS-2-F and IDS-2-R is selected from at least one of the primer pair of SEQ ID NO: 8 and SEQ ID NO: 11 and the primer pair of SEQ ID NO: 9 and SEQ ID NO: 12; and

[0088] (3) a primer pair of IDS2-F and IDS2-R, wherein the primer pair of IDS2-F and IDS2-R is selected from at least one of the primer pair of SEQ ID NO: 13 and SEQ ID NO: 16 and the primer pair of SEQ ID NO: 15 and SEQ ID NO: 18.

[0089] In a more preferred embodiment, the method for detecting IDS The primer set for gene mutation contains the following primers:

[0090] (1) a primer pair of IDS-1-F and IDS-1-R, wherein the primer pair of IDS-1-F and IDS-1-R are respectively shown as SEQ ID NO: 1 and SEQ ID NO: 4 or respectively shown as SEQ ID NO: 2 and SEQ ID NO: 5;

[0091] (2) a primer pair of IDS-2-F and IDS-2-R, wherein the primer pair of IDS-2-F and IDS-2-R are respectively shown as SEQ ID NO: 8 and SEQ ID NO: 11 or respectively shown as SEQ ID NO: 9 and SEQ ID NO: 12; and

[0092] (3) IDS2-F and IDS2-R primer pairs, wherein the IDS2-F and IDS2-R primer pairs are respectively shown as SEQ ID NO: 13 and SEQ ID NO: 16 or respectively shown as SEQ ID NO: 15 and SEQ ID NO: 18.

[0093] In one embodiment, the primer set provided by the present invention can cover IDS Full-length genes and pseudogenes IDS2 full length.

[0094] In one embodiment, the IDS Gene mutations include one or more of the following: IDS Point mutations, small insertion / deletion and structural variations of genes.

[0095] In one embodiment, the primer set provided by the present invention can be used to detect the following IDS One or more of the following genetic mutations: IDS Point mutations, small insertion / deletion and structural variations of genes.

[0096] In one embodiment, the primer set provided by the present invention further comprises a primer for amplification IDS The centromere is about 1 Mb to IDS2 An IDS-Gap Mix primer set for a large deletion in a region about 1 Mb distal to the centromere of a gene, wherein the IDS-Gap Mix primer set comprises primers as shown in SEQ ID NOs: 19-218.

[0097] In one embodiment, the primer set provided by the present invention can also cover IDS The centromere is about 1 Mb to IDS2 The approximately 1 Mb region distal to the gene centromere.

[0098] In one embodiment, the IDS Gene mutations further include IDS The centromere is about 1 Mb to IDS2 A large deletion was found within a region approximately 1 Mb distal to the centromere of the gene.

[0099] In one embodiment, the primer set of the present invention can also be used to detect IDS The centromere is about 1 Mb to IDS2 A large deletion was found within a region approximately 1 Mb distal to the centromere of the gene.

[0100] In a preferred embodiment, the present invention provides a method for detecting IDS The primer set for gene mutation contains the following primers:

[0101] (1) a primer pair of IDS-1-F and IDS-1-R for amplifying exon 7 to exon 9 and the flanking introns at both ends thereof, wherein the IDS-1-F is selected from at least one of the primers shown in SEQ ID NO: 1 and SEQ ID NO: 2, and the IDS-1-R is selected from at least one of the primers shown in SEQ ID NO: 4 and SEQ ID NO: 5;

[0102] (2) a primer pair of IDS-2-F and IDS-2-R for amplifying exon 1 to exon 6 and the flanking introns at both ends thereof, wherein the IDS-2-F is selected from at least one of the primers shown in SEQ ID NO: 8 and SEQ ID NO: 9, and the IDS-2-R is selected from at least one of the primers shown in SEQ ID NO: 11 and SEQ ID NO: 12;

[0103] (3) For amplification IDS2 a primer pair of IDS2-F and IDS2-R, wherein the IDS2-F is selected from at least one of the primers shown in SEQ ID NO: 13 and SEQ ID NO: 15, and the IDS2-R is selected from at least one of the primers shown in SEQ ID NO: 16 and SEQ ID NO: 18; and

[0104] (4) For amplification IDS The centromere is about 1 Mb to IDS2 An IDS-Gap Mix primer set for large deletions within a region of about 1 Mb distal to the centromere of a gene, wherein the IDS-Gap Mix comprises primers as shown in SEQ ID NOs: 19-218.

[0105] In a more preferred embodiment, the method for detecting IDS The primer set for gene mutation contains the following primers:

[0106] (1) a primer pair of IDS-1-F and IDS-1-R, wherein the primer pair of IDS-1-F and IDS-1-R is selected from at least one of the primer pair of SEQ ID NO: 1 and SEQ ID NO: 4 and the primer pair of SEQ ID NO: 2 and SEQ ID NO: 5;

[0107] (2) a primer pair of IDS-2-F and IDS-2-R, wherein the primer pair of IDS-2-F and IDS-2-R is selected from at least one of the primer pair of SEQ ID NO: 8 and SEQ ID NO: 11 and the primer pair of SEQ ID NO: 9 and SEQ ID NO: 12;

[0108] (3) a primer pair of IDS2-F and IDS2-R, wherein the primer pair of IDS2-F and IDS2-R is selected from at least one of the primer pair of SEQ ID NO: 13 and SEQ ID NO: 16 and the primer pair of SEQ ID NO: 15 and SEQ ID NO: 18; and

[0109] (4) An IDS-Gap-Mix primer set, wherein the IDS-Gap-Mix comprises primers as shown in SEQ ID NOs: 19-218.

[0110] In a more preferred embodiment, the method for detecting IDS The primer set for gene mutation contains the following primers:

[0111] (1) a primer pair of IDS-1-F and IDS-1-R, wherein the primer pair of IDS-1-F and IDS-1-R are respectively shown as SEQ ID NO: 1 and SEQ ID NO: 4 or respectively shown as SEQ ID NO: 2 and SEQ ID NO: 5;

[0112] (2) a primer pair of IDS-2-F and IDS-2-R, wherein the primer pair of IDS-2-F and IDS-2-R are respectively shown as SEQ ID NO: 8 and SEQ ID NO: 11 or respectively shown as SEQ ID NO: 9 and SEQ ID NO: 12;

[0113] (3) a primer pair of IDS2-F and IDS2-R, wherein the primer pair of IDS2-F and IDS2-R are respectively shown as SEQ ID NO: 13 and SEQ ID NO: 16 or respectively shown as SEQ ID NO: 15 and SEQ ID NO: 18; and

[0114] (4) An IDS-Gap-Mix primer set, wherein the IDS-Gap-Mix comprises primers as shown in SEQ ID NOs: 19-218.

[0115] In a more preferred embodiment, the method for detecting IDS The primer set for gene mutation contains the following primers:

[0116] (1) a primer pair of IDS-1-F and IDS-1-R, wherein the primer pair of IDS-1-F and IDS-1-R are shown in SEQ ID NO: 1 and SEQ ID NO: 4, respectively;

[0117] (2) a primer pair of IDS-2-F and IDS-2-R, wherein the primer pair of IDS-2-F and IDS-2-R are shown in SEQ ID NO: 8 and SEQ ID NO: 11, respectively;

[0118] (3) a primer pair of IDS2-F and IDS2-R, wherein the primer pair of IDS2-F and IDS2-R are shown in SEQ ID NO: 13 and SEQ ID NO: 16, respectively; and

[0119] (4) An IDS-Gap-Mix primer set, wherein the IDS-Gap-Mix comprises primers as shown in SEQ ID NOs: 19-218.

[0120] In one embodiment, the primer set provided by the present invention can cover IDS Full-length gene, pseudogene IDS2 Full length and IDS The centromere of the gene is about 1 Mb to IDS2 The approximately 1 Mb region distal to the gene centromere.

[0121] In one embodiment, the primer sets of the present invention can be used to detect the following IDS One or more of the following genetic mutations: IDS Point mutations, small fragment insertions / deletions, structural variations and IDS The centromere of the gene is about 1 Mb to IDS2 A large deletion was found within a region approximately 1 Mb distal to the centromere of the gene.

[0122] The primer set provided by the present invention can amplify the complete sequence within the primer range on the genome, wherein the IDS-1-F / R primer pair can amplify IDS The gene exon 7 to exon 9 and the flanking introns at both ends of the fragment are about 14 kb in length, and the fragment contains the pseudogene in the intron 7 region. IDS2 Homologous regions; IDS-2-F / R primer pair can amplify IDS The gene exon 1 to exon 6 and the flanking introns at both ends of the fragment are about 15 kb in length, and the fragment contains the pseudogene in the exon 2 to intron 3 region. IDS2 Homologous regions; IDS2-F / R primer pair can amplify IDS2 The gene region contains a total of about 15 kb fragments. The amplified products of the above primer set can be stably detected by long-read sequencing IDS The present invention can detect point mutations, small fragment insertion / deletion and structural variation of genes. IDSThe two fragments of the gene can not only accurately detect different types of rearrangements, but also detect complex rearrangements with multiple recombination. In addition, the primer set provided by the present invention further includes an IDS-Gap Mix primer set, which can be used to detect IDS Gene-related large fragment deletion. The overall primer design idea of ​​the primer set can be referred to Figure 1 , the specific sequence information can be found in Table 1 below.

[0123] Table 1:

[0124]

[0125] In a preferred embodiment, if there is a SNP at the hybridization site between any primer in the primer set and the genomic DNA, a primer containing a degenerate base at the SNP site is used. Those skilled in the art should be familiar with the design and synthesis of degenerate base primers.

[0126] In a preferred embodiment, in order to detect multiple samples simultaneously, the primers in the primer set can also be added with a DNA barcode at its 5' end for distinguishing different samples. The DNA barcode can be a different oligonucleotide sequence with a length of 5-50 nt, which can be used to distinguish different samples and realize simultaneous detection of multiple samples. The 5' end DNA barcodes of the forward primer and the reverse primer in the primer set can be symmetrical or asymmetrical, and those skilled in the art can select as needed. In a specific embodiment, when a symmetrical DNA barcode is used, the DNA barcode at the 5' end of the reverse primer should be reverse complementary to the DNA barcode at the 5' end of the forward primer.

[0127] The primer set of the present invention can be used for multiple long-fragment PCR amplification in a single reaction system. IDS Genes and pseudogenes IDS2 The complete sequence of the gene, as well as unknown large missing fragments.

[0128] The primer set of the present invention can be used to amplify all mutation types within the primer range by multiple long-fragment PCR in a single reaction system. IDS The relevant pathogenic fragments are further combined with the subsequent long-read sequencing platform to simultaneously detect all mutation types of the target fragment within the primer range.

[0129] In one embodiment, IDS Gene mutations include at least IDSPoint mutations, small insertion / deletions, structural variations of genes, and IDS The centromere of the gene is about 1 Mb to IDS2 A large deletion in the region of about 1 Mb distal to the centromere of a gene. IDS Point mutations, small insertion / deletions, structural variations of genes, and IDS The centromere of the gene is about 1 Mb to IDS2 The large deletions in the approximately 1 Mb region distal to the centromere of the gene include not only the large deletions that can be found in the commonly used human gene mutation databases and existing literature, but also the large deletions in the approximately 1 Mb region distal to the centromere of the gene, including the large deletions that can be found in the commonly used human gene mutation databases and existing literature. IDS Gene mutations also include unknown mutations in the above regions, which are not specifically limited here. The commonly used human gene mutation databases include but are not limited to the Human Gene Mutation Database (HGMD, https: / / www.hgmd.cf.ac.uk / ac / index.php), the Leiden Open Variation Database (LOVD, https: / / www.lovd.nl / 3.0 / home) and the Human Genetic Variation Database (ClinVar, https: / / www.ncbi.nlm.nih.gov / clinvar / ).

[0130] Kit

[0131] In order to at least partially solve the above problems and one or more of other potential problems, a second exemplary embodiment of the present invention is provided for detecting IDS A kit for gene mutation comprises at least the primer set provided according to the first exemplary embodiment of the present invention.

[0132] In one embodiment, in order to use the primer set to achieve amplification of the target fragment, the kit may also include a reagent for long-fragment PCR amplification. Specifically, the reagent for long-fragment PCR amplification is used in combination with the primer set provided in the first exemplary embodiment of the present invention to achieve amplification of the target fragment.

[0133] In a specific embodiment, the reagents for long-fragment PCR amplification include at least one of a DNA polymerase, dNTPs and a reaction buffer.

[0134] The basic principle of multiplex PCR (Multiplex Polymerase Chain Reaction) is the same as that of conventional PCR. The difference is that two or more pairs of primers are added to the multiplex PCR reaction system at the same time. These primers complement each other with specific regions on the template DNA, so that multiple target nucleic acid fragments can be amplified simultaneously in the same reaction.

[0135] In a preferred embodiment, the primer sets of the present invention are mixed and provided in a single reaction tube to obtain target fragments by multiple long-fragment PCR; preferably, the primer sets are mixed and provided in a single reaction tube in an equimolar ratio.

[0136] In one embodiment, those skilled in the art may also choose whether to purify the amplified product obtained by long-fragment PCR amplification according to actual conditions. In a specific embodiment, in order to purify the amplified product obtained by long-fragment PCR amplification, the kit may also include a reagent for DNA purification, such as a DNA purification reagent based on magnetic beads, such as DNA purification magnetic beads; or a DNA purification reagent based on a silica membrane adsorption column, such as a DNA adsorption column.

[0137] In one embodiment, in order to achieve multiple amplification products obtained by amplification with the above primer set, IDS Gene mutations are detected simultaneously, and the kit may also include reagents for constructing a long-read sequencing library.

[0138] Specifically, the reagent for constructing a long-read sequencing library is used to construct a long-read sequencing library for the amplified product obtained by amplification with the above primer set to generate a library suitable for long-read sequencing. By performing long-read sequencing on the long-read sequencing library, the long-read sequencing result of the amplified product can be obtained, thereby analyzing the amplified product. IDS Type of gene mutation.

[0139] Specifically, the long-read sequencing can be single-molecule real-time (SMRT) sequencing from Pacific Biosciences, Inc. Single-molecule real-time sequencing is a long-read sequencing technology that identifies bases based on different fluorescent signals generated when different nucleotides bind to the sequence during nucleic acid molecule synthesis.

[0140] In one embodiment, in order to achieve single-molecule real-time sequencing of the amplified product, the kit may further include reagents for constructing a single-molecule real-time sequencing library. Specifically, the reagents for constructing a single-molecule real-time sequencing library include at least one of a damage repair reagent, an end repair reagent, a label ligation reagent, a DNA purification reagent, a linker ligation reagent, and an exonuclease.

[0141] In a specific embodiment, the adapter ligation reagent for constructing a single molecule real-time sequencing library comprises at least a PacBio hairpin adapter. In a specific embodiment, the sequence of the PacBio blunt-end hairpin adapter is

[0142] 5'- / 5Phos / ATCTCTCTCTTTTCCTCCTCCTCCGTTGTTGTTGTTGAGAGAGAT-3' (SEQ ID NO: 219), the PacBio sticky end hairpin adapter sequence is

[0143] 5'- / 5Phos / ATCTCTCTCTTTTCCTCCTCCTCCGTTGTTGTTGTTGAGAGAGATT-3' (SEQ IDNO: 220)

[0144] In a preferred embodiment, the PacBio hairpin adapter is annealed to form a stem-loop structure adapter adapter. In order to detect multiple samples simultaneously, the PacBio hairpin adapter can also add a DNA barcode at the end of its stem for distinguishing different samples. The DNA barcode can be a different oligonucleotide sequence with a length of 5-50 nt, which can be used to distinguish different samples and realize the simultaneous detection of multiple samples. For example, in a specific embodiment, the sequence of the PacBio blunt-end hairpin adapter with a DNA barcode is

[0145] 5'- / 5Phos / (Barcode)ATCTCTCTCTTTTCCTCCTCCTCCGTTGTTGTTGTTGAGAGAGAT(Reverse Complement Barcode)-3'.

[0146] In one embodiment, the DNA barcode is a DNA barcode designed by PacBio or a self-designed DNA barcode, which can be selected by those skilled in the art as needed.

[0147] Specifically, the long-read sequencing may also be nanopore sequencing from Oxford Nanopore Technology, Inc. Nanopore sequencing technology is a long-read sequencing technology that identifies bases based on the difference in electrical signals generated when different nucleotides pass through a nanopore.

[0148] In one embodiment, in order to realize nanopore sequencing of the amplified product, the kit may further include reagents for constructing a nanopore sequencing library. Specifically, the reagents for constructing a nanopore sequencing library include at least one of a damage repair reagent, an end repair reagent, a label ligation reagent, a DNA purification reagent, and a linker ligation reagent.

[0149] In a specific embodiment, the adapter ligation reagent for constructing a nanopore sequencing library comprises at least a nanopore universal sequencing adapter, and the nanopore universal sequencing adapter comprises a polynucleotide sequencing adapter and a motor protein loaded to the polynucleotide adapter.

[0150] In a preferred embodiment, in order to detect multiple samples simultaneously, the label connection reagent for constructing the nanopore sequencing library includes at least a DNA barcode fragment or a transposase complex containing a DNA barcode. The DNA barcode can be a different oligonucleotide sequence with a length of 5-50 nt, so that it can be used to distinguish different samples and realize simultaneous detection of multiple samples. In one embodiment, the DNA barcode is a DNA barcode designed by ONT or a self-designed DNA barcode, and those skilled in the art can choose as needed.

[0151] In one embodiment, the kit proposed by the present invention may also optionally include conventional kits known in the art according to the needs of those skilled in the art. Those skilled in the art may select the reagent composition in the kit according to actual needs, and the present invention does not limit this.

[0152] The kit of the present invention can be used to amplify mutation types including all primers in a single reaction system by multiple long-fragment PCR. IDS The relevant pathogenic fragments are further combined with the subsequent long-read sequencing platform to simultaneously detect the mutation types of all gene fragments within the primer range.

[0153] In one embodiment, the IDS Gene mutations include at least IDS Point mutations, small insertion / deletions, structural variations of genes, and IDS The centromere of the gene is about 1 Mb to IDS2 A large deletion in the region of about 1 Mb distal to the centromere of a gene. IDS Point mutations, small insertion / deletions, structural variations of genes, and IDS The centromere of the gene is about 1 Mb to IDS2 The large deletions in the approximately 1 Mb region distal to the centromere of the gene include not only the large deletions that can be found in the commonly used human gene mutation databases and existing literature, but also the large deletions in the approximately 1 Mb region distal to the centromere of the gene, including the large deletions that can be found in the commonly used human gene mutation databases and existing literature. IDSGene mutations also include unknown mutations within the above-mentioned regions, which are not specifically limited here.

[0154] Detection method

[0155] In order to at least partially solve the above problems and one or more of other potential problems, the third exemplary embodiment of the present invention proposes a method for detecting IDS The method for gene mutation comprises the following steps:

[0156] (1) Obtain samples from subjects;

[0157] (2) using the primer set provided according to the first exemplary embodiment or the kit provided according to the second exemplary embodiment to amplify the DNA in the sample by long-range PCR. IDS Gene fragment, obtain the target fragment amplification product;

[0158] (3) constructing a long-read sequencing library based on the amplified product;

[0159] (4) performing long-read sequencing based on the long-read sequencing library to obtain a long-read sequencing result;

[0160] (5) Analyze and determine based on the long read sequencing results IDS Type of gene mutation.

[0161] In one embodiment, the sample in step (1) is selected from a biological sample or genomic DNA extracted from the biological sample. Preferably, the biological sample can be a cultured cell line, blood, tissue, amniotic fluid, chorionic villi, gametes, blastocyst cells, joint fluid, urine, sweat, saliva, feces, cerebrospinal fluid, ascites, pleural effusion, bile or pancreatic fluid, etc., which are not specifically limited here.

[0162] In a preferred embodiment, the long-range PCR in step (2) can be completed in a single reaction system.

[0163] In one embodiment, after step (2), the amplification product can be preliminarily analyzed by electrophoresis. IDS When the gene has no structural variation, the amplification product should contain the product obtained by the IDS-1-F / R primer pair. IDS The gene exon 7 to exon 9 and the flanking introns at both ends of the fragment are about 14 kb in length; IDS-2-F / R primer pair amplified IDS The gene exon 1 to exon 6 and the flanking introns at both ends of the fragment are about 15 kb in length; IDS2-F / R primer pair amplified IDS2 The gene region contains a total of about 15 kb fragments. IDSGene structural variation occurs, for example, when the sample IDS When a large gene fragment is deleted, the primer pairs upstream and downstream of the deletion site will amplify IDS A truncated fragment of a gene; IDS When a large fragment of a gene is inserted, the primer pairs upstream and downstream of the insertion site will amplify the IDS Extension of a gene.

[0164] In one embodiment, after step (2), a person skilled in the art may also choose whether to purify the amplification product obtained by long-fragment PCR amplification according to actual conditions.

[0165] In one embodiment, the construction of the long-read sequencing library in step (3) includes optional steps and means for constructing a long-read sequencing library known in the art, including but not limited to damage repair, end repair, label ligation, adapter ligation, purification and exonuclease digestion, etc.

[0166] In a specific embodiment, the long-read sequencing library in step (3) is suitable for a single-molecule real-time sequencing platform. Typically, constructing a single-molecule real-time sequencing library comprises:

[0167] (3.1) performing DNA damage repair and end repair on the target fragment obtained in step (2), and optionally, adding base A to the 3' end of the end-repaired target fragment;

[0168] (3.2) Connect the two ends of the repaired target fragment obtained in step (3.1) to the PacBio hairpin adapter;

[0169] (3.3) Digest the ligation product obtained in step (3.2) with an exonuclease;

[0170] (3.4) Purify the digestion product obtained in step (3.3) to obtain a single molecule real-time sequencing library.

[0171] In one embodiment, the connection of the PacBio hairpin adapter in step (3.2) can be performed by blunt-end connection or sticky-end connection.

[0172] In a specific embodiment, the PacBio blunt-ended hairpin adapter sequence is

[0173] 5'- / 5Phos / ATCTCTCTCTTTTCCTCCTCCTCCGTTGTTGTTGTTGAGAGAGAT-3' (SEQ ID NO: 219), the PacBio sticky end hairpin adapter sequence is

[0174] 5'- / 5Phos / ATCTCTCTCTTTTCCTCCTCCTCCGTTGTTGTTGTTGAGAGAGATT-3' (SEQ IDNO: 220).

[0175] In a preferred embodiment, the PacBio hairpin adapter is annealed to form a stem-loop structure adapter adapter. In order to detect multiple samples simultaneously, the PacBio hairpin adapter can also add a DNA barcode at the end of its stem for distinguishing different samples. The DNA barcode can be a different oligonucleotide sequence with a length of 5-50 nt, which can be used to distinguish different samples and realize the simultaneous detection of multiple samples. For example, in a specific embodiment, the sequence of the PacBio blunt-end hairpin adapter with a DNA barcode is

[0176] 5'- / 5Phos / (Barcode)ATCTCTCTCTTTTCCTCCTCCTCCGTTGTTGTTGTTGAGAGAGAT(Reverse Complement Barcode)-3'.

[0177] In one embodiment, the DNA barcode is a DNA barcode designed by PacBio or a self-designed DNA barcode, which can be selected by those skilled in the art as needed.

[0178] In a specific embodiment, the long-read sequencing library in step (3) is suitable for a nanopore sequencing platform. Typically, constructing a nanopore sequencing library comprises:

[0179] (3.1) performing DNA damage repair and end repair on the target fragment obtained in step (2), and optionally, adding base A to the 3' end of the end-repaired target fragment;

[0180] (3.2) connecting the two ends of the repaired target fragment obtained in step (3.1) to the nanopore universal sequencing adapter;

[0181] (3.3) Purify the ligation product obtained in step (3.2) to obtain a nanopore sequencing library.

[0182] In one embodiment, the ligation of the nanopore universal sequencing adapter in step (3.2) can be performed by blunt-end ligation or sticky-end ligation.

[0183] In a specific embodiment, the nanopore universal sequencing adapter comprises a polynucleotide sequencing adapter and a motor protein loaded to the polynucleotide adapter.

[0184] In a preferred embodiment, in order to detect multiple samples simultaneously, before step (3.2), DNA barcodes are connected to both ends of the repaired target fragment obtained in step (3.1). The DNA barcodes can be different oligonucleotide sequences with a length of 5-50 nt, so that they can be used to distinguish different samples and realize simultaneous detection of multiple samples. In one embodiment, the DNA barcode is a DNA barcode designed by ONT or a self-designed DNA barcode, and those skilled in the art can choose according to their needs.

[0185] In one embodiment, the long-read sequencing described in step (4) is selected from single-molecule real-time sequencing of PacBio or nanopore sequencing of ONT, and matches the long-read sequencing library constructed in step (3).

[0186] In one embodiment, step (5) analyzes and determines based on the long read sequencing results. IDS Types of gene mutations include:

[0187] (5.1) Perform basic quality control and filtering on long-read sequencing results to obtain valid sequencing results;

[0188] (5.2) Compare the effective sequencing results with the reference genome to obtain the comparison results;

[0189] (5.3) According to the comparison results, we can obtain IDS Gene mutation.

[0190] In one embodiment, the basic quality control and filtering of the long-read sequencing results described in step (5.1) can use the optional steps and means for basic quality control and filtering of long-read sequencing results known in the art, including but not limited to removing sequencing adapter sequences, low-quality bases, contaminating sequences, and sample splitting.

[0191] In one embodiment, the alignment of the valid sequencing results with the reference genome in step (5.2) can use optional alignment steps and means known in the art, and the reference genome can be selected from an optional reference genome known in the art, such as GRCh37 or GRCh38.

[0192] In one embodiment, since the long-read sequencing used in the present application can detect the length of the complete amplified fragment of the amplified product, the change in the nucleotide sequence of the amplified fragment compared with the reference genome can be directly obtained in step (5.3) according to the comparison result obtained in step (5.2), thereby obtaining IDS Gene mutation. Specifically, step (5.3) IDS Gene mutations include at least IDSPoint mutations, small insertion / deletions, structural variations of genes, and IDS The centromere of the gene is about 1 Mb to IDS2 A large deletion in the region of about 1 Mb distal to the centromere of a gene. IDS Point mutations, small insertion / deletions, structural variations of genes, and IDS The centromere of the gene is about 1 Mb to IDS2 The large deletions in the approximately 1 Mb region distal to the centromere of the gene include not only the large deletions that can be found in the commonly used human gene mutation databases and existing literature, but also the large deletions in the approximately 1 Mb region distal to the centromere of the gene, including the large deletions that can be found in the commonly used human gene mutation databases and existing literature. IDS Gene mutations also include unknown mutations within the above-mentioned regions, which are not specifically limited here.

[0193] The method described in the present invention is based on a specific combination of multiple long-fragment PCR and long-read sequencing, which can achieve high specificity, accuracy and rapidity in the simultaneous detection of MPS II pathogenic genes in multiple samples. IDS Various related mutations.

[0194] Detection system

[0195] In order to at least partially solve the above problems and one or more of other potential problems, the fourth exemplary embodiment of the present invention proposes a method for detecting IDS The gene mutation system includes the following modules:

[0196] (1) an amplification module: used to amplify the sample from the subject by long-range PCR using the primer set provided according to the first exemplary embodiment or the kit provided according to the second exemplary embodiment. IDS Gene fragment, obtain the target fragment amplification product;

[0197] (2) Library construction module: used to construct a long-read sequencing library based on the amplification product;

[0198] (3) Sequencing module: used to perform long-read sequencing based on the long-read sequencing library to obtain long-read sequencing results;

[0199] (4) Mutation analysis module: used to analyze and determine based on the long read sequencing results IDS Type of gene mutation.

[0200] In one embodiment, the sample is selected from a biological sample or genomic DNA extracted from the biological sample. Preferably, the biological sample can be a cultured cell line, blood, tissue, amniotic fluid, chorionic villi, gametes, blastocyst cells, joint fluid, urine, sweat, saliva, feces, cerebrospinal fluid, ascites, pleural effusion, bile or pancreatic fluid, etc., which are not specifically limited here.

[0201] In one embodiment, the long-range PCR is performed in a single reaction tube.

[0202] In one embodiment, the long read sequencing is selected from single molecule real-time sequencing or nanopore sequencing.

[0203] In one embodiment, the mutation analysis module can be used to determine the following IDS One or more of the following genetic mutations: IDS Point mutations, small insertion / deletions, structural variations of genes, and IDS The centromere of the gene is about 1 Mb to IDS2 A large deletion was found within a region approximately 1 Mb distal to the centromere of the gene.

[0204] It should be noted that the system of the present invention is specifically used to detect IDS Gene mutation, wherein each module can refer to existing equipment, for example, the amplification module can refer to the existing PCR amplification instrument, the library construction module can refer to the existing laboratory automated library construction workstation, the sequencing module can refer to the existing long-read sequencing device, and the mutation analysis module can refer to the existing computing equipment that can run bioinformatics processing and analysis software. The above modules are not specifically limited here.

[0205] Furthermore, the system of the present invention can be further expanded with other functional modules according to the needs of technical personnel in this field. For example, the system of the present invention can further include a genomic DNA extraction module for extracting genomic DNA from samples. The module can refer to the existing laboratory automated nucleic acid extraction workstation. The above-mentioned other expandable functional modules are not specifically limited here.

[0206] The present invention is used to detect IDS The connection between the modules in the gene mutation system can refer to the existing laboratory automation devices, and is not specifically limited here. It can be understood that the key to this application is to organically integrate the existing scattered devices or components together to achieve simultaneous detection. IDS The purpose of multiple gene mutations.

[0207] It is understood that the detection of this application IDS The gene mutation system actually uses various modules to implement the detection of this application IDS Therefore, the preferred technical solutions of each module in the system of the present invention, such as the primers used for multiple long-fragment PCR amplification, the method for constructing a long-read sequencing library, the method for determining the mutation type, etc., can refer to the detection method provided in the third exemplary embodiment of the present invention. IDS The methods of gene mutation will not be repeated here.

[0208] Application

[0209] In order to at least partially solve the above problems and one or more of other potential problems, the fifth exemplary embodiment of the present invention proposes the use of the primer set provided according to the first exemplary embodiment of the present invention or the kit provided according to the second exemplary embodiment of the present invention in any of the following aspects:

[0210] 1) Detection of IDS gene mutations;

[0211] 2) Detection of IDS gene rearrangement;

[0212] 3) Genetic diagnosis of mucopolysaccharidosis type II.

[0213] In order to at least partially solve the above problems and one or more of other potential problems, the sixth exemplary embodiment of the present invention proposes the use of the primer set provided according to the first exemplary embodiment of the present invention or the kit provided according to the second exemplary embodiment of the present invention in preparing any of the following products:

[0214] 1) Products for detecting IDS gene mutations;

[0215] 2) Products for detecting IDS gene rearrangements;

[0216] 3) Products used for genetic diagnosis of mucopolysaccharidosis type II.

[0217] In one embodiment, the IDS Gene mutations include at least IDS Point mutations, small insertion / deletions, structural variations of genes, and IDS The centromere is about 1 Mb to IDS2 A large deletion in the region of about 1 Mb distal to the centromere of a gene. IDS Point mutations, small insertion / deletions, structural variations of genes, and IDS The centromere is about 1 Mb to IDS2 The large deletions in the approximately 1 Mb region distal to the centromere of the gene include not only the large deletions that can be found in the commonly used human gene mutation databases and existing literature, but also the large deletions in the approximately 1 Mb region distal to the centromere of the gene, including the large deletions that can be found in the commonly used human gene mutation databases and existing literature. IDS Gene mutations also include unknown mutations within the above-mentioned regions, which are not specifically limited here.

[0218] Preferred embodiment

[0219] The following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. The following description is exemplary and not intended to limit the present invention, and any other similar situations also fall within the protection scope of the present invention.

[0220] Example 1: Amplification using different primer combinations provided by the present invention IDS Gene fragment

[0221] The reaction system was prepared according to Table 2 below, using IDS-1-F1 / R1 (SEQ ID NO: 1 and SEQ ID NO: 4), IDS-1-F2 / R2 (SEQ ID NO: 2 and SEQ ID NO: 5), IDS-1-F3 / R3 (SEQ ID NO: 3 and SEQ ID NO: 6), IDS-2-F1 / R1 (SEQ ID NO: 7 and SEQ ID NO: 10), IDS-2-F2 / R2 (SEQ ID NO: 8 and SEQ ID NO: 11), IDS-2-F3 / R3 (SEQ ID NO: 9 and SEQ ID NO: 12), IDS2-F1 / R1 (SEQ ID NO: 13 and SEQ ID NO: 16), IDS2-F2 / R2 (SEQ ID NO: 14 and SEQ ID NO: 17) or IDS2-F3 / R3 (SEQ ID NO: 15 and SEQ ID NO: 16). 18) Primer pairs to amplify target fragments in peripheral blood, dried blood spots or genomic DNA samples:

[0222] Table 2:

[0223]

[0224] On the PCR instrument, pre-amplification was performed according to the conditions shown in Table 3 below:

[0225] Table 3:

[0226]

[0227] After amplification, 5 μL of amplified product was taken from each sample and electrophoresed on a 1% DNA agarose gel. Figure 2 As shown, the primer pairs IDS-1-F1 / R1, IDS-1-F2 / R2, IDS-2-F2 / R2, IDS-2-F3 / R3, IDS2-F1 / R1 and IDS2-F3 / R3 can effectively amplify the target fragment.

[0228] Example 2: Construction of PacBio sequencing library using the multiple long-fragment PCR method of the present invention

[0229] Step 1: Multiplex long fragment PCR amplification

[0230] The reaction system was prepared according to Table 4 below, and the primer sets IDS-1-F1 / R1, IDS-2-F2 / R2, IDS2-F1 / R1 and IDS-Gap Mix that can effectively amplify the target fragment in Example 1 were used to amplify different MPS II related IDS Amplification of peripheral blood samples for gene mutations:

[0231] Table 4:

[0232]

[0233] On the PCR instrument, pre-amplification was performed according to the conditions shown in Table 5 below:

[0234] Table 5:

[0235]

[0236] After amplification, the amplified product was placed in a centrifuge at 10,000 rpm for 20 minutes. After centrifugation, the tube was left to stand horizontally and 4 μL of the supernatant was added to a new centrifuge tube.

[0237] Step 2: Construction of PacBio sequencing library

[0238] Prepare the library construction reaction system according to Table 6 below:

[0239] Table 6:

[0240]

[0241] The reaction was carried out in a PCR instrument under the following conditions: 37ºC for 20 min; 25ºC for 15 min; 65ºC for 10 min. After the reaction was completed, 0.5μL Exonuclease III (NEB, Cat#M0206L) and 0.5μL Exonuclease VII (NEB, Cat#M0379L) were added and the reaction was continued at 37ºC for 1 hour. After the reaction was completed, 0.6x Ampure PB magnetic beads (PacBio, Cat#100-265-900) were used for purification twice according to the manufacturer's instructions, and finally the DNA was eluted with 10μL Elution Buffer. The obtained DNA eluate is the PacBio sequencing library of the target DNA fragment. The DNA concentration was measured on the Qubit 3 Fluoromter (ThermoFisher, Cat#Q33216) using Qubit dsDNA HS reagent (ThermoFisher, Cat# Q32851). When there are PacBio sequencing libraries from multiple samples, equal amounts of PacBio sequencing libraries can be mixed together to prepare a mixed library.

[0242] Step 3: PacBio sequencing and analysis on the machine

[0243] According to the total concentration and molar concentration of the library, an appropriate volume of the library was reacted with the binding reagent (PacBio, Cat#101-820-200) and sequencing primer (PacBio, Cat#100-970-100) to prepare the final library that can be used in the machine. Representative sequencing results are shown in Figures 3 to 9 As shown, Figure 3 for IDS IGV schematic diagram of the sequencing results of gene exon point mutation samples, Figure 4 for IDS IGV schematic diagram of the sequencing results of gene intron point mutation samples. Figure 5 for IDS IGV schematic diagram of intragenic duplicate sample sequencing results. Figure 6 for IDS IGV schematic diagram of the sequencing results of intragenic deletion samples. Figure 7 For IDS Gene rearrangement related IDS IGV schematic diagram of the sequencing results of gene inversion samples. Figure 8 and Figure 9 For different regions IDS IGV schematic diagram of sequencing results of samples with large gene deletions.

[0244] Embodiment 3: IDS Detection and verification of gene mutations

[0245] Peripheral blood genomic DNA of 110 subjects was collected as validation samples, and the method (and kit) of the present invention was used to simultaneously detect IDS As a control, we also used IDS Exon Sanger sequencing and MLPA or WES detection IDS The results obtained by the method of the present invention were compared with those obtained by the control method, and the results are shown in Table 7. Among them, in 94 samples, the results of the two methods were completely consistent. In addition, in 16 samples, the method of the present invention detected additional mutations related to IDS Gene rearrangement-related mutations, including 15 samples IDS Gene intron 7 inversion; 1 sample involved IDS The gene showed a complex rearrangement with intron 7 and exon 3 inversions, and these results were confirmed to be correct by customized sequencing.

[0246] Table 7:

[0247]

[0248] Therefore, the results detected by the method of the present invention are compared with IDS Comparison between the exon Sanger sequencing method and the MLPA or WES method showed that both specificity and sensitivity reached 100%.

[0249] It should be noted that the aforementioned examples are merely illustrative and are used to explain some of the features of the present invention. The attached claims are intended to claim the widest possible range that can be imagined, and the embodiments presented herein are only descriptions of selected implementation methods based on the combination of all possible embodiments. Therefore, the applicant's intention is that the attached claims are not limited by the selection of examples that illustrate the features of the present invention. As used in the claims, the term "including" and its semantic variants logically also include different and changing terms, such as but not limited to "essentially composed of" or "composed of". When necessary, some numerical ranges are provided, and these ranges also include sub-ranges therebetween. Changes in these ranges are also self-evident to those skilled in the art, and should not be considered to be donated to the public, and these changes should also be interpreted as being covered by the attached claims where possible. For example, the reaction reagents, reaction conditions, etc. involved in the construction of multiple long-fragment PCR reactions and long-read sequencing libraries can be adjusted and changed accordingly according to specific needs. Furthermore, advances in technology will result in possible equivalents or sub-alternatives that are not presently contemplated due to inaccuracies in the language, and such variations should also be construed to be covered by the appended claims, where possible.

Claims

1. For detection IDS A primer set for gene mutation, the primer set comprising the following primers: (1) a primer pair of IDS-1-F and IDS-1-R, wherein the primer pair of IDS-1-F and IDS-1-R is selected from at least one of the primer pair of SEQ ID NO: 1 and SEQ ID NO: 4 and the primer pair of SEQ ID NO: 2 and SEQ ID NO: 5; (2) a primer pair of IDS-2-F and IDS-2-R, wherein the primer pair of IDS-2-F and IDS-2-R is selected from at least one of the primer pair of SEQ ID NO: 8 and SEQ ID NO: 11 and the primer pair of SEQ ID NO: 9 and SEQ ID NO: 12; (3) a primer pair of IDS2-F and IDS2-R, wherein the primer pair of IDS2-F and IDS2-R is selected from at least one of the primer pair of SEQ ID NO: 13 and SEQ ID NO: 16 and the primer pair of SEQ ID NO: 15 and SEQ ID NO: 18; and (4) An IDS-Gap Mix primer set, wherein the IDS-Gap Mix primer set comprises primers as shown in SEQ ID NOs: 19-218.

2. The primer set according to claim 1, IDS Gene mutations include one or more of the following: IDS Point mutations, small insertion / deletion and structural variations of genes.

3. The primer set according to claim 1, IDS Gene mutations further include IDS 1Mb proximal to the centromere of the gene IDS2 A large deletion was found within the 1 Mb region distal to the centromere of the gene. The primer set according to claim 1 , wherein the 5′ end of the primer in the primer set is further connected to a DNA barcode for distinguishing different samples.

5. Used for detection IDS A kit for detecting gene mutation, comprising the primer set according to any one of claims 1 to 4.

6. The kit according to claim 5, further comprising reagents for long-fragment PCR amplification, reagents for DNA purification and / or reagents for constructing a long-read sequencing library.

7. The kit according to claim 6, wherein the reagents for long-fragment PCR amplification include at least one of a DNA polymerase, dNTPs and a reaction buffer.

8. According to the kit according to claim 6, the reagents for constructing a long-read sequencing library include at least one of a damage repair reagent, an end repair reagent, a label ligation reagent, a DNA purification reagent, a linker ligation reagent and an exonuclease.

9. The kit according to claim 6, wherein the long-read sequencing library is suitable for a single-molecule real-time sequencing platform or a nanopore sequencing platform.

10. The kit according to claim 5, wherein the primer set is provided mixed in a single reaction tube.

11. The kit according to any one of claims 5 to 10, wherein IDS Gene mutations include one or more of the following: IDS Point mutations, small insertion / deletions, structural variations of genes, and IDS 1 Mb proximal to the centromere of the gene IDS2 A large deletion was found within the 1 Mb region distal to the centromere of the gene.

12. For detection IDS The gene mutation system includes the following modules: (1) Amplification module: used to amplify a sample from a subject by long-range PCR using the primer set of any one of claims 1 to 4 or the kit of any one of claims 5 to 11. IDS Gene fragment, obtain the target fragment amplification product; (2) Library construction module: used to construct a long-read sequencing library based on the amplification product; (3) Sequencing module: used to perform long-read sequencing based on the long-read sequencing library to obtain long-read sequencing results; (4) Mutation analysis module: used to analyze and determine based on the long read sequencing results IDS Type of gene mutation.

13. The system according to claim 12, wherein the subject-derived sample is selected from a biological sample or a genomic DNA extracted from the biological sample. The system according to claim 12 , wherein the long-range PCR is performed in a single reaction tube.

15. The system according to claim 12, wherein the long read sequencing is selected from single molecule real-time sequencing or nanopore sequencing.

16. The system according to any one of claims 12-15, wherein the mutation analysis module can be used to determine one or more of the following IDS Gene mutations: IDS Point mutations, small insertion / deletions, structural variations of genes, and IDS 1 Mb proximal to the centromere of the gene IDS2 A large deletion was found within the 1 Mb region distal to the centromere of the gene.

17. Use of the primer set according to any one of claims 1 to 4 or the kit according to any one of claims 5 to 11 in the preparation of any of the following products: 1) For detection IDS Products of genetic mutations; 2) For detection IDS Products of gene rearrangement; 3) Products used for genetic diagnosis of mucopolysaccharidosis type II.

18. The use according to claim 17, wherein IDS Gene mutations include one or more of the following: IDS Point mutations, small insertion / deletions, structural variations of genes, and IDS 1 Mb proximal to the centromere of the gene IDS2 A large deletion was found within the 1 Mb region distal to the centromere of the gene.

Citation Information

Patent Citations

  • Method, primer and kit for detecting various mutations of pigment incontinence IKBKG gene

    CN117625778A

  • Primer group and kit for detecting various mutations of human embryo alpha-thalassemia and application of primer group and kit

    CN118853875A