A new-born genetic disease screening gene combination, probe combination, kit and preparation method and device thereof
By providing probe sets and kits for newborn genetic disease screening, combined with bioinformatics analysis, the problem of low efficiency in newborn genetic disease screening has been solved, enabling rapid and accurate genetic disease detection and personalized medication plans, thereby improving the health and quality of life of newborns.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CAPITALBIO GENOMICS
- Filing Date
- 2024-12-05
- Publication Date
- 2026-05-19
AI Technical Summary
Current technologies are insufficient for efficiently screening newborns for genetic diseases, especially single-gene genetic diseases, and lack individualized treatment plans based on drug responses, resulting in low diagnostic and treatment efficiency and increasing the health and economic burden on affected children.
This invention provides a probe set and kit for newborn genetic disease screening, which includes a probe combination targeting 279 single-gene genetic diseases, 331 high-frequency gene loci, and 74 safe drug use gene loci. Combined with bioinformatics analysis, it enables rapid and accurate detection of newborn genetic diseases and drug interpretation.
It enables rapid detection of more than 40,000 pathogenic mutations in 279 genetic disease-related genes, providing genetic identification and drug interpretation, improving detection efficiency and the safety and effectiveness of medication regimens, and reducing the incidence of adverse drug reactions.
Smart Images

Figure CN120060462B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of gene detection, specifically relating to a gene combination, probe combination, reagent kit, and preparation method and device for newborn genetic disease screening. Background Technology
[0002] Statistics show that the incidence of birth defects in my country is 5.6%, and about 70% of these defects often present no symptoms in the early stages, but once they develop, they can be life-threatening or cause permanent damage to intelligence and physical health. While the incidence of individual diseases is low, the overall incidence is high, exceeding 1 / 100, far exceeding the incidence of Down syndrome. In my country's first batch of rare disease catalogs published in 2018, single-gene genetic diseases accounted for the majority. The average incidence of each of these diseases in the global population is between 1 / 10,000 and 1 / 100,000, ranging from one in several thousand newborns to one in hundreds of thousands. However, for certain regions or specific populations from certain areas, the average incidence of some single-gene genetic diseases may be as high as 1% to 2%. For example, in some ethnic minorities in Southeast Asia and southern Yunnan, the incidence of thalassemia and glucose-6-phosphate dehydrogenase deficiency is over 2%. Most single-gene disorders are fatal, cause birth defects, or lead to disability. Less than 10% have effective treatments, but these are expensive. Even if the child survives, they often suffer lifelong disabilities or intellectual impairments. Therefore, single-gene disorders not only seriously harm the health of patients but also impose a heavy emotional and economic burden on families and society, becoming a significant obstacle affecting the quality of my country's newborn population and children's health. Therefore, newborn screening for some manageable congenital and hereditary diseases is necessary for early detection and diagnosis, which helps improve newborn survival rates and the quality of life for affected children. Newborn screening is therefore crucial.
[0003] Newborn screening (NBS) refers to specialized examinations conducted during the neonatal period to detect congenital and hereditary diseases that seriously endanger the health of newborns, providing early diagnosis and assisting in maternal and infant treatment. (Technicals Specimens for Newborn Disease Screening 2010 edition. Accessed 1 Dec 2021) [1] Newborn screening has a history of over 60 years. Technological advancements have greatly promoted its development, resulting in significant improvements in both the types of diseases screened and screening efficiency. [2] Compared with traditional newborn screening, NGS-based genetic screening has a wider scope and directly targets the cause of disease, making it suitable for large-scale disease screening and subsequent reproductive guidance, which is conducive to precision diagnosis and treatment and the prevention and control of genetic diseases.
[0004] A genetic ID card is an identity document created using genetic technology, identifying a number of fixed polymorphic gene loci. The combination of these gene loci is unique, providing distinctive genetic information that can serve as reliable evidence in the process of identifying and recovering a lost child.
[0005] Pharmacogenomics studies the impact of genetic variations and expression levels of genes involved in drug metabolism, transport, and drug target mechanisms on drug response. It explains individualized differences in drug response at the genetic level, thereby improving drug efficacy, reducing side effects, saving medical costs, and ultimately achieving personalized medicine. Pharmacogenomics is a rapidly developing and highly regarded research field in recent years, and an important component of precision medicine. For patients, it not only improves the safety and effectiveness of drug therapy, enabling doctors to develop more precise individualized medication plans, but also significantly reduces the incidence of adverse drug reactions. This technology can predict drug responses, avoid drug toxicity, and optimize drug dosage, playing a crucial role, especially in the treatment of mental illnesses and other special diseases. Furthermore, pharmacogenomics helps optimize long-term treatment plans, improve chronic disease management, and thus enhance patients' quality of life.
[0006] Therefore, it is necessary to develop a more effective test kit for comprehensive screening of genetic diseases in newborns.
[0007] References:
[0008] [1]Technical specifications for newborn disease screening 2010 edition.Accessed 1 Dec 2021.
[0009] [2]Zhao ZY. Progress in international neonatal disease screening.Chin J Child Health Care. 2012;20:193–5. Summary of the Invention
[0010] The first aspect of the present invention is to provide a probe set for newborn genetic disease screening.
[0011] The second objective of this invention is to provide a kit for screening newborn genetic diseases.
[0012] The third objective of this invention is to provide a newborn genetic disease screening system.
[0013] The fourth aspect of this invention is to provide an apparatus for screening newborns for genetic diseases.
[0014] The fifth aspect of this invention is to provide an electronic device.
[0015] The sixth aspect of this invention is to provide a storage medium.
[0016] To achieve the above-mentioned objectives of this invention, the technical solution adopted by this invention is as follows:
[0017] In a first aspect, the present invention provides a primer set for newborn genetic disease screening, wherein the specific sequence of the probe set is as follows:
[0018] http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF%E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88%E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97.pdf
[0019] The probe sequence file has been timestamped and certified. The timestamp certificate can be found at: http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF%E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88%E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97_TSA%E8%AF%81%E4%B9%A6.jpg. All experiments and data in this application were completed before the application date.
[0020] In some embodiments of the present invention, the specific sequences of the probe group are shown in SEQ ID NO: 5~14913.
[0021] The probe set includes a combination of probes designed for 3908 regions, comprising a newborn screening gene combination targeting 279 single-gene genetic diseases, 331 high-frequency gene loci, and 74 safe medication gene loci. It should be noted that the probe capture regions are divided into three categories of loci: loci for the 279 single-gene genetic diseases, high-frequency loci, and safe medication loci. There is overlap between the single-gene genetic disease loci and the safe medication detection loci; therefore, there is also overlap in the corresponding capture regions. After deduplication, there are actually 3908 loci.
[0022] A second aspect of the present invention provides a kit for newborn genetic disease screening, the kit containing a probe set of the first aspect of the present invention.
[0023] In some embodiments of the present invention, the kit further comprises thalassemia large fragment deletion detection reagents, including LA Taq DNA Polymerase, KAPA 5×GC buffer, dNTP stock, etc.
[0024] In some embodiments of the present invention, the kit further includes a library preparation hybridization reagent and hybridization capture magnetic beads.
[0025] A third aspect of the present invention provides the use of the probe set of the first aspect of the present invention or the kit of the second aspect of the present invention in the preparation of products for detecting genetic diseases in newborns.
[0026] A fourth aspect of the present invention provides a newborn genetic disease screening system, the screening system comprising: an indicator input module and a single-gene genetic disease gene assessment module;
[0027] The indicator input module includes at least obtaining sequencing information of the target gene in the subject's blood and / or tissue samples using the probe set of the first aspect of the present invention or the kit of the second aspect of the present invention;
[0028] The single-gene genetic disease assessment module includes at least performing bioinformatics analysis on the sequencing results obtained from the indicator input module to obtain variation results and outputting the assessment results of the subject.
[0029] The screening flowchart of the newborn genetic disease screening system is as follows: Figure 1 As shown.
[0030] In some embodiments of the present invention, the indicator input module includes obtaining sequencing information of the target gene in the subject's blood and / or tissue samples using the following methods:
[0031] (1) DNA sample collection and extraction: Collect the samples to be tested and extract genomic DNA;
[0032] (2) DNA fragmentation / end repair / dA tail: The genomic DNA is digested with enzymes, repaired at the ends, and dA tails are added;
[0033] (3) Adapter ligation: The DNA fragments obtained in the previous step are ligated with adapters to obtain ligation products, and the products are purified;
[0034] (4) Pre-hybridization amplification: The purified product from step (3) is subjected to PCR amplification, enrichment, and purification to obtain a DNA library for hybridization capture.
[0035] (5) Probe capture: The DNA library is blocked before hybridization, and hybridization capture is performed using the probe set described in the first aspect of the present invention;
[0036] (6) Elution and target region enrichment: T1 magnetic beads are incubated with probe capture products. After incubation, elution is performed. The eluted and rehydrated products are then subjected to PCR amplification to obtain the enriched target sequence and the product to be sequenced.
[0037] (7) Sequencing: The product to be sequenced is subjected to next-generation sequencing.
[0038] In some embodiments of the present invention, the adapter in step 3) is a Y-type adapter, which includes a first sequence and a second sequence, the first sequence being shown in SEQ ID NO:1 and the second sequence being shown in SEQ ID NO:2. -s- indicates thiomodification, -p- indicates phosphorylation modification, and NNNNN(N) indicates a UMI molecular tag of 5-6 nucleotides in length.
[0039] In some embodiments of the present invention, in steps 3) to 4), the purification is performed using magnetic beads. In some embodiments of the present invention, the 25 μL reaction system for PCR amplification enrichment comprises: 9–12 μL of purified product, 10–15 μL of PCR amplification enzyme mixture, and 1–3 μL of tag primer mixture.
[0040] In some embodiments of the present invention, the forward sequence of the tag primers amplified in step 4) is shown in SEQ ID NO:3, and the reverse sequence is shown in SEQ ID NO:4; Index1 and Index2 are tag primers, consisting of 8 specific base sequences, used to identify, for example: GCCTATCA, CTTGGATG, TCACAGCA, TCCTACCT. -s- indicates thiomodification, [i5] represents the 8 bp i5 index sequence, and [i7] represents the 8 bp i7 index sequence.
[0041] In some embodiments of the present invention, the PCR amplification and enrichment reaction conditions are: pre-denaturation at 97–99°C for 05–2 min; denaturation at 95–98°C for 5–30 sec, annealing at 58–62°C for 25–35 sec, extension at 70–72°C for 25–35 sec, for 5–7 cycles; final extension at 70–72°C for 4–5 min; and holding at 3–5°C.
[0042] In some embodiments of the present invention, the PCR amplification system in step 6) is as follows: 35-37 μL of elution and rehydration product; 8-11 μL of 5×PCR amplification mixture; 1-2 μL of capture amplification primers; 1-2 μL of DNA polymerase; and 0.5-1 μL of dNTP mixture.
[0043] In some embodiments of the present invention, the PCR amplification reaction conditions in step 6) are as follows: pre-denaturation at 97-99°C for 2-3 min; denaturation at 97-99°C for 20-30 sec, annealing at 55-57°C for 20-30 sec, extension at 70-72°C for 1-2 min, for 10-12 cycles; final extension at 70-72°C for 8-12 min; and holding at 3-5°C.
[0044] In some embodiments of the present invention, in step (7), the second-generation sequencing is performed using the Illumina second-generation sequencing platform NovaSeq6000 or other domestic second-generation sequencing platforms for PE150 sequencing.
[0045] In some embodiments of the present invention, the bioinformatics analysis sequencing results include the following steps:
[0046] 1) Data quality control and cleaning: Use FastP software to perform data quality control and connector removal operations;
[0047] 2) Molecular tag (UMI) identification and sequence alignment: After labeling the UMIs with GATK, the preprocessed reads were aligned with the human reference genome hg19 using BWA software;
[0048] 3) Data quality correction: Remove low-quality reads (including unmatched, poorly matched, multiple-matched reads, and repetitive sequences caused by PCR amplification bias).
[0049] 4) Variance detection: GATK software was used for variant detection, followed by false positive site identification and filtering.
[0050] 5) Variant annotation: Using Annovar software, the sites where mutations occur are annotated to obtain information such as their genetic function, population carrier frequency, and disease relevance.
[0051] 6) CNV detection: CNVkit software is used to perform gene CNV detection. CNVkit calculates the copy number of each region in each sample and counts the copy number of the region where the large fragment deletion of thalassemia is located. At the same time, it divides the thalassemia gene region into small window regions and counts the average sequencing depth. With normal reference as background, the average depth will be reduced if there is a missing region in the sample. In summary, it can identify whether there is a missing region in the sample.
[0052] In some embodiments of the present invention, the method for detecting thalassemia-related variants includes the following specific verification experimental steps:
[0053] 1) The DNA sample used for detection was subjected to PCR amplification. The amplification experimental system was as follows: KAPA 5×GC buffer, 1.4 μL; 10 mM dNTP stock, 0.2 μL; LA Taq DNA Polymerase, 0.1 μL; Primer pool (10 pM), 0.3 μL; DNA template (30 ng) and NF-H2O / ddH2O total 8 μL; the amplification reaction system was as follows: 94℃ pre-denaturation for 5 min; 95℃ denaturation for 30 sec, 60℃ annealing for 75 sec, 72℃ extension for 90 sec, 35 cycles; 72℃ final extension for 5 min; 4℃ incubation.
[0054] 2) Perform PCR amplification products on 1.5% agarose gel electrophoresis and analyze the electrophoresis results.
[0055] In a fifth aspect, the present invention provides an apparatus for detecting genes of single-gene hereditary diseases. The apparatus can utilize the probe set of the first aspect of the present invention and the reagent kit of the second aspect of the present invention to construct a sequencing library of the sample to be tested, and perform bioinformatics analysis on the data generated after sequencing to obtain the variation results.
[0056] In some embodiments of the present invention, the bioinformatics analysis includes the following steps:
[0057] 1) Data quality control and cleaning: Use FastP software to perform data quality control and connector removal operations;
[0058] 2) Molecular tag (UMI) identification and sequence alignment: After labeling the UMIs with GATK, the preprocessed reads were aligned with the human reference genome hg19 using BWA software;
[0059] 3) Data quality correction: Remove low-quality reads (including unmatched, poorly matched, multiple-matched reads, and repetitive sequences caused by PCR amplification bias).
[0060] 4) Variance detection: GATK software was used for variant detection, followed by false positive site identification and filtering.
[0061] 5) Variant annotation: Using Annovar software, the sites where mutations occur are annotated to obtain information such as their genetic function, population carrier frequency, and disease relevance.
[0062] 6) CNV detection: CNVkit software is used to perform gene CNV detection. CNVkit calculates the copy number of each region in each sample and counts the copy number of the region where the large fragment deletion of thalassemia is located. At the same time, it divides the thalassemia gene region into small window regions and counts the average sequencing depth. With normal reference as background, the average depth will be reduced if there is a missing region in the sample. In summary, it can identify whether there is a missing region in the sample.
[0063] In some embodiments of the present invention, subsequent analysis is performed based on the results of variant detection in device 4) for newborn single-gene genetic disease gene detection. The specific analysis steps are as follows:
[0064] 1) Based on the known high-frequency gene locus combinations, the detection results of high-frequency gene loci are statistically analyzed through the mutation results. If no high-frequency gene locus is detected, it is considered to be wild type. Finally, the detection results of the high-frequency gene locus combinations serve as the genetic identity card.
[0065] 2) Based on the known combinations of safe drug use gene loci, the detection results of safe drug use gene loci are statistically analyzed. If no loci are detected, the gene is considered to be wild-type.
[0066] 3) Based on the test results and the drug interpretation database, obtain relevant drug interpretation information, including recommendations on drug risk level and treatment dosage.
[0067] A sixth aspect of the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the screening process of the screening system of the fourth aspect of the present invention.
[0068] A seventh aspect of the invention provides a storage medium storing processor-executable instructions, which, when executed by a processor, are used to perform the screening process of the screening system of the fourth aspect of the invention.
[0069] The beneficial effects of this invention are:
[0070] The kit provided in this invention contains capture probes targeting clearly pathogenic genes, high-frequency gene loci, and safe drug use gene loci. It can detect more than 40,000 pathogenic mutations on 279 genetic disease-related genes at once, enabling rapid and accurate understanding of the pathogenic mutation status of the examinee. It can also detect 331 high-frequency gene loci and 74 safe drug use gene loci at once, providing a unique genetic identity and drug interpretation, helping clinicians to develop safer and more effective medication plans for examinees.
[0071] The detection device provided in this invention can use bioinformatics analysis methods to automatically analyze and organize sequencing data obtained using the kits in this invention, assisting clinicians in targeted variant interpretation, improving work efficiency, and providing the necessary conditions for the large-scale application of newborn screening. Attached Figure Description
[0072] The present invention will be further described below with reference to the accompanying drawings and embodiments, wherein:
[0073] Figure 1 This is the overall flowchart of the present invention.
[0074] Figure 2 This is a diagram showing the average depth of the window for hybrid samples in section 4.2.
[0075] Figure 3 This is a graph showing the average depth of the sample window for --SEA / -α3.7. Detailed Implementation
[0076] The following will describe the concept and technical effects of the present invention clearly and completely with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.
[0077] Overall process of the present invention Figure 1 As shown.
[0078] In the embodiments of the present invention, the specific information of the adapter sequence and specific primer tag sequence used is shown in Table 1 (SEQ ID NO:1~SEQ ID NO:4); in the embodiments of the present invention, the DNA library construction kit (catalog number: 12205) of Yeasen was used for library construction; based on the probe capture region of the single-gene hereditary disease-related pathogenic gene provided by the present invention, the corresponding probe can be synthesized according to conventional methods. The synthesis of the probe can be completed by simply providing the information of the capture region to a professional probe design and synthesis company.
[0079] Table 1. Connector sequences and specific primer tag sequences
[0080]
[0081] Note: In the table s Indicates thiomodification. p Indicates phosphorylation modification, [i5] represents an 8bp i5 index sequence, [i7] represents an 8bp i7 index sequence, and NNNNN(N) represents 5 A 6-nucleotide-long UMI molecular tag.
[0082] Example 1: Detection probe set of 279 genes for 309 single-gene genetic diseases
[0083] After reviewing extensive literature and databases, the inventors identified 309 genetic diseases with clearly defined phenotypes, where timely intervention and treatment can reduce complications and prevent disease progression, along with their related genes (see Table 2). Through high-throughput sequencing and related experimental verification, 3536 probe capture regions were designed to detect pathogenic gene loci for single-gene genetic diseases. Using probes in combination with these capture regions can detect pathogenic gene mutations in over 40,000 single-gene genetic diseases. The designed probe set exhibits high specificity, high sensitivity, simplicity, and cost-effectiveness, facilitating the promotion and clinical application of newborn screening for single-gene genetic diseases. The probe sequence information and probe site can be found at http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF%E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88%E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97.pdf.
[0084] Table 2 Single-gene recessive genetic diseases and their related genes
[0085]
[0086] Example 2: 331 high-frequency gene loci and 330 gene detection probe sets
[0087] The inventors, after consulting the East Asian population frequencies (EAS) database of The 1000 Genomes Project, identified 331 gene loci with high population frequencies and their associated population frequencies (see Table 3). Through high-throughput sequencing experiments and related verification, 329 probe capture regions were designed to detect high-frequency gene loci (a gene may have multiple loci; a capture region can capture the base information of a segment of the genome, which may contain multiple loci; if the loci are close together, a single capture region can capture information from multiple loci). The combination of detection results for high-frequency gene loci provides a unique genetic identifier. The probe sequence information and probe capture sites can be found at http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF%E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88%E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97.pdf.
[0088] Table 3. High-frequency gene loci and their population frequencies
[0089]
[0090] Example 3: 74 safe medication gene loci and 49 gene detection probe sets
[0091] After reviewing a large amount of literature and the PharmGKB database, the inventors identified 74 safe medication gene loci, 32 genes, and their associated detection drugs (see Table 4). Through high-throughput sequencing experiments and related verification, 49 probe capture regions were designed for detecting these safe medication gene loci. Drug interpretation is obtained through the detection results of gene loci and haplotypes, helping clinicians develop safer and more effective personalized medication plans for patients. The probe sequence information and probe capture sites can be found at http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF%E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88%E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97.pdf.
[0092] Table 4. Gene loci for safe medication and related detection drugs
[0093]
[0094] Example 4: Kit for Newborn Single-Gene Genetic Disease Detection
[0095] This embodiment exemplarily provides a newborn genetic disease gene detection kit based on reversible strand termination sequencing. It is worth noting that single-gene genetic disease gene detection kits developed using the probe combination of this invention, based on other sequencing platforms or not based on sequencing, should also be within the scope of protection of this application.
[0096] Newborn genetic disease gene testing kit, including:
[0097] 1) Hybridization capture probe;
[0098] 2) Library construction and hybridization reagents: The library construction and hybridization kit in this embodiment is used for sequencing library construction.
[0099] 3) Hybrid trapping magnetic beads;
[0100] 4) Thalassemia Large Deletion Detection Reagent (GAP-PCR): This example uses the GAP-PCR reagent. Other gel running reagents can be prepared independently, such as agarose, 1 kb ladder plus, and nucleic acid staining agents.
[0101] Note: The following reagents can be prepared by yourself: extraction reagents, commercially available nucleic acid extraction and purification reagents, such as the nucleic acid extraction or purification reagent produced by Dongguan Bio-Muhua Gene Technology Co., Ltd., product number: S10040; purification magnetic beads, such as Beckman's Agencourt AMPure XP Kit.
[0102] The experimental system and reagents listed above are as follows: KAPA 5×GC buffer, 1.4 μL; 10 mM dNTPstock, 0.2 μL; LA Taq DNA Polymerase, 0.1 μL; Primer pool (10 pM), 0.3 μL; DNA template (30 ng) and NF-H2O / ddH2O total 8 μL; the amplification reaction system is as follows: 94℃ pre-denaturation for 5 min; 95℃ denaturation for 30 sec, 60℃ annealing for 75 sec, 72℃ extension for 90 sec, 35 cycles; 72℃ final extension for 5 min; 4℃ incubation.
[0103] Example 5: Library construction and sequencing of samples using a newborn genetic disease gene detection kit.
[0104] In this embodiment, an exemplary method is provided for library construction and sequencing of samples using the kit from Example 4, the steps of which are as follows:
[0105] (1) Sample collection and extraction
[0106] Genomic DNA was extracted from selected peripheral blood samples and dried blood spot samples using nucleic acid extraction or purification reagents (it is recommended to use nucleic acid extraction or purification reagents produced by Dongguan Bio-Muhua Gene Technology Co., Ltd., catalog number: S10040). The purity and concentration of the extracted genomic DNA were tested using Nanodrop2000, and the integrity of the extracted genomic DNA was verified by agarose gel electrophoresis.
[0107] (2) DNA fragmentation, end repair and dA tail addition: The genomic DNA was digested with enzymes, repaired at the ends and dA tails were added. The 30 μL reaction system was: 200 ng of DNA sample to be tested, 5 μL of fragmentation end repair and A tail addition buffer, 2.5 μL of fragmentation end repair and A tail addition enzyme, and enzyme-free water as the balance. The reaction conditions were set as follows: 4℃ for 1 min, 30℃ for 12 min, 65℃ for 20 min, 4℃, Hold.
[0108] (3) Adapter ligation and purification: Adapter ligation was performed directly after the previous reaction. The reaction system consisted of 15 μL of rapid ligation buffer, 2.5 μL of rapid maturation ligase, and 2.5 μL of adapter. The reaction conditions were set as follows: 20℃ for 15 min, 4℃, Hold. The ligation product was purified using magnetic beads. The specific steps were as follows: 50 μL of ligation product was placed in a centrifuge tube, 30 μL of magnetic beads were added, and after standing for 5 min, the tube was placed on a magnetic rack. The supernatant was removed, and the magnetic beads were retained. The tube was washed twice with 75% anhydrous ethanol, 11 μL of nuclease-free water was added, and the supernatant was transferred to a new centrifuge tube. The purified product was then used for the next step of amplification.
[0109] The adapter is a Y-type adapter, comprising: a first sequence (as shown in SEQ ID NO. 1); and a second sequence (as shown in SEQ ID NO. 2). Pre-hybridization amplification: The purified product in step (3) is enriched by PCR amplification. The PCR amplification enrichment system is: 10 μL of purified product, 12.5 μL of PCR amplification enzyme mixture, and 2.5 μL of tag primer mixture.
[0110] (4) The forward sequence of the tag primers is shown in SEQ ID NO. 3, and the reverse sequence is shown in SEQ ID NO. 4; Index1 and Index2 are tag primers, consisting of 8 specific base sequences used to identify different samples, such as: GCCTATCA, CTTGGATG, TCACAGCA, TCCTACCT. The PCR amplification and enrichment reaction conditions are: 98℃ pre-denaturation for 1 min; 98℃ denaturation for 10 sec, 60℃ annealing for 30 sec, 72℃ extension for 30 sec, cycled 4 times; 72℃ final extension for 5 min; and 4℃ incubation.
[0111] (5) Probe capture: The constructed DNA library is pooled and blocked, and a probe mixture is prepared; the probe mixture is hybridized and captured with the blocked DNA library; the hybridization and capture reaction system is shown in Table 5.
[0112] Table 5 Hybridization Capture Reaction System
[0113]
[0114] The probe mixture was prepared as follows: 2.25 uL of nuclease-free water, 0.3 uL of RNase inhibitor, and 1 uL of probe; prepared on ice.
[0115] The reaction conditions for hybridization capture are shown in Table 6.
[0116] Table 6 Hybridization capture reaction conditions
[0117]
[0118] (6) Elution and target region enrichment: Streptavidin T1 magnetic beads were incubated with the probe capture product for 60 min. After incubation, elution was performed, and the eluted and rehydrated product was subjected to PCR amplification to obtain the enriched target sequence and the product to be sequenced. The amplification reaction system was as follows: 36.5 μL of eluted and rehydrated product; 10 μL of 5×PCR amplification mixture; 1 μL of amplification primers after capture; 1 μL of DNA polymerase; 0.5 μL of dNTP mixture.
[0119] The PCR amplification and enrichment reaction conditions were as follows: 98℃ pre-denaturation for 2 min; 98℃ denaturation for 30 sec, 57℃ annealing for 30 sec, 72℃ extension for 1 min, 10 cycles; 72℃ final extension for 10 min; and incubation at 4℃.
[0120] (7) Sequencing and analysis: The product to be sequenced is subjected to second-generation sequencing, and the sequencing results are analyzed using bioinformatics.
[0121] Example 6: Gap-PCR experiment on samples using a newborn genetic disease gene detection kit.
[0122] In this embodiment, an exemplary method for performing gap-PCR experiments on samples using the kit from Example 4 is provided, and the steps are as follows:
[0123] Sample collection and extraction; collect samples to be tested and extract genomic DNA;
[0124] DNA dilution; DNA concentration was measured and diluted to 10 ng / μL;
[0125] Gap-PCR amplification; simultaneously prepare a 1.5% agarose gel for electrophoresis.
[0126] Electrophoresis: 150V, 40min electrophoresis and interpretation of results based on 3 positive controls.
[0127] The experimental system and reagents listed above are as follows: KAPA 5×GC buffer, 1.4 μL; 10 mM dNTPstock, 0.2 μL; LA Taq DNA Polymerase, 0.1 μL; Primer pool (10 pM), 0.3 μL; DNA template (30 ng) and NF-H2O / ddH2O total 8 μL; the amplification reaction system is as follows: 94℃ pre-denaturation for 5 min; 95℃ denaturation for 30 sec, 60℃ annealing for 75 sec, 72℃ extension for 90 sec, 35 cycles; 72℃ final extension for 5 min; 4℃ incubation.
[0128] Example 7: Accuracy assessment of detection of pathogenic loci, high-frequency gene loci, and safe medication gene loci in newborns with genetic diseases.
[0129] This embodiment, following the method described in Example 5, performed library construction and sequencing on 18 positive reference samples and 5 negative reference samples (positive and negative samples that have undergone clinical testing and validation by our company). Using our self-developed analysis workflow, we obtained mutation information for over 40,000 single-gene genetic disease gene pathogenic sites in each sample. The detection results were compared with known theoretical results to assess the accuracy of the detection. The results are shown in Table 7. It can be seen that the detection results of both the positive reference samples and negative controls are consistent with the theoretical genotypes, indicating that this kit has excellent detection accuracy.
[0130] Table 7. Comparison of Detection Results of Reference Samples at Pathogenic Loci of Single-Gene Hereditary Diseases
[0131]
[0132] Example 8: Accuracy assessment of high-frequency gene loci detection
[0133] This embodiment follows the method described in Example 5, performing library construction and sequencing on three reference samples (positive and negative samples that have undergone clinical testing and validation by our company). Using our self-developed analysis workflow, mutation information for 331 high-frequency gene loci in each sample was obtained. The detection results were compared with known theoretical results to evaluate the accuracy of the detection. The results are shown in Table 8. It can be seen that the detection results for the reference samples are consistent with the theoretical genotypes, indicating that this kit has excellent detection accuracy.
[0134] Table 8. Comparison of Detection Results of Reference Samples at High-Frequency Gene Loci
[0135]
[0136] Example 9: Accuracy Assessment of Detection of Gene Loci for Safe Medication Use
[0137] This embodiment, following the method described in Example 5, performed library construction and sequencing on three reference samples (positive and negative samples that have undergone clinical testing and validation by our company). Using our self-developed analysis workflow, we obtained mutation information for 74 safe drug use gene loci in each sample. The detection results were compared with known theoretical results to assess the accuracy of the detection. The results are shown in Table 9. It can be seen that the detection results for the reference samples are consistent with the theoretical genotypes, indicating that this kit has excellent detection accuracy.
[0138] Table 9. Compliance of Reference Sample Detection Results at Safe Medication Gene Loci
[0139]
[0140] Example 10: Evaluation of the effectiveness of newborn genetic disease screening results
[0141] This embodiment, following the method described in Example 5, performed library construction and sequencing on four families (positive and negative samples that have undergone clinical testing and validation by our company). Using our self-developed analysis workflow, we obtained mutation information for over 40,000 single-gene genetic disease detection loci in each sample. The test results from the families were then compiled to assess the accuracy of the screening results. The results are shown in Table 10. It can be seen that the test results of all family samples conform to the laws of inheritance, indicating that this kit has excellent effectiveness in screening for single-gene genetic diseases. Early detection and diagnosis of genetic diseases helps improve newborn survival rates and the quality of life for affected children.
[0142] Table 10. Compliance of test results for family samples
[0143]
[0144] Note: Variance classification: P (Pathogenic), LP (Likely pathogenic), P / LP (Pathogenic / Likely pathogenic, Pathogenic or Likely pathogenic).
[0145] Example 11: Effectiveness assessment of high-frequency gene loci
[0146] This embodiment follows the method described in Example 5, performing library construction and sequencing on three replicates of each of the three reference samples (positive and negative samples that have undergone clinical testing and validation by our company). Using a self-developed analysis workflow, mutation information for 331 high-frequency gene loci in each sample was obtained. The consistency of the repeatability test results for each reference sample was evaluated. When the inconsistency rate was less than 0.1%, the samples were considered to be from the same sample, thus assessing the effectiveness of the test. The results are shown in Table 11. It can be seen that the inconsistency rate of the repeatability test results for the reference samples was 0%, indicating that this kit has excellent effectiveness in combining high-frequency gene loci. Because the combination of gene loci is unique, it can be used to prove specific genetic information.
[0147] Table 11. Reproducibility of Detection Results of Reference Samples at High-Frequency Gene Loci
[0148]
[0149] Example 12: Evaluation of the effectiveness of detecting gene loci for safe drug use
[0150] This embodiment, following the method described in Example 5, performed library construction and sequencing on three reference samples (positive and negative samples that have undergone clinical testing and validation). Using a self-developed analysis workflow, mutation information for 74 safe drug use gene loci in each sample was obtained. The test results were combined with a drug interpretation library to compile drug treatment recommendations, thereby evaluating the effectiveness of the test. The results are shown in Table 12. It can be seen that the risk drug interpretation of the reference samples indicated relevant drug interpretations, demonstrating that this kit has excellent detection effectiveness.
[0151] Table 12 Reference Drug Risk Recommendations
[0152]
[0153] Example 13 Performance Evaluation of Neonatal Genetic Disease Gene Probe Detection
[0154] In this embodiment, the library was constructed and sequenced for two standards (positive and negative samples that have been clinically tested and verified by our company) according to the method described in Example 5. Using our self-developed analysis process, the evaluation variant set was finally integrated based on the true set of site variants in the capture interval of the standards and the Clinvar filtering site, totaling 218,106 sites.
[0155] To evaluate probe detection performance, the number of true positives (TP), true negatives (TP), false positives (FP), and false negatives (FN) at the evaluation mutation concentration sites in the standard samples were counted. Key evaluation metrics included sensitivity (SEN), specificity (SPE), positive predictive value (PPV), and F1 score (F1), calculated using the following formula:
[0156] SEN = TP / (TP + FN) * 100%
[0157] SPE = TN / (FP + TN) * 100%
[0158] PPV = TP / (TP + FP) * 100%
[0159] F1 = 2 * PPV * SEN / ( PPV + SEN)
[0160] As can be seen from the pre-filtering data in Table 14, the detection performance for some sites is not very good. This embodiment, through observation of the variant AF, found that filtering with an appropriate AF threshold can effectively improve PPV.
[0161] This kit uses ROC curves to determine the filtering threshold for apoptosis (AF). The ROC curve is plotted with the false positive rate (1-SPE) at different thresholds on the x-axis and the sensitivity (SEN) at different thresholds on the y-axis. The maximum value of the Youden index corresponds to the optimal diagnostic cutoff value of this method, which can be used as the optimal classification threshold. The formula for calculating the Youden index (J) is:
[0162] J = SEN + SPE – 1.
[0163] Table 13. Statistics of Youden's Index on ROC Curve
[0164]
[0165] As shown in Table 13, the Youden index (J) reaches its maximum value when AF=0.12, indicating the optimal model result. The optimal classification threshold for AF is 0.12. Filtering the unfiltered data in Table 14 based on the optimal classification threshold for AF shows improvements in specificity and F1 score, while sensitivity slightly decreases. Furthermore, Table 15 shows that within the final detection range of over 40,000 sites, the F1 score reaches 1.000, indicating that the kit's performance is at its best. This demonstrates that the probes in this kit possess excellent detection performance.
[0166] Table 14. Performance differences of the variance results of the evaluation variant set before and after filtering with AF≥0.12 parameter.
[0167]
[0168] Table 15 Performance of 47,354 detection ranges within the detection scope
[0169]
[0170] Example 14: Detection performance of large thalassemia deletions
[0171] This embodiment follows the method described in Example 5, performing library construction and sequencing on --SEA / -3.7 samples, 4.2 heterozygous samples, and three normal reference samples. The detection results were obtained using a self-developed analysis workflow. This embodiment uses two methods to detect large deletions in thalassemia: 1. Analyzing the detection results using cnvkit software; 2. Establishing small windows for the thalassemia gene region sequences and calculating the average window depth. Using the detection results of the normal reference samples as background references, the two samples with large deletions were analyzed. The detection performance was evaluated by combining the above methods.
[0172] The results from the cnvkit software are shown in Table 16. It can be seen that the --SEA / -3.7 sample has a CN < 0.1 in the 3.7 region and < 0.5 in the SEA region, thus the result is --SEA / -3.7; the 4.2 heterozygous sample has a CN of approximately 0.5 in the 4.2 region, thus the result is 4.2 heterozygous. The average depth of the small window in the thalassemia gene region is shown in... Figure 1 , Figure 2 As shown, the average depth of the --SEA / -3.7 sample and the heterozygous sample 4.2 in the corresponding region is significantly lower than that of the normal reference sample, indicating the presence of a deficiency in this region. All the above test results are consistent with the theoretical results, demonstrating excellent performance in the detection of thalassemia.
[0173] Table 16 shows the results of the cnvkit software.
[0174]
[0175] Example 15 Performance of Detection of Large Deletion Gap-PCR in Thalassemia
[0176] This embodiment follows the method described in Example 6, performing gap-PCR amplification on 18 reference samples with large fragment deletions and 7 normal reference samples (positive and negative samples that have been clinically validated by our company). The results of the control samples and the reference samples were interpreted based on their electrophoresis results. The detection results were compared with known theoretical results to evaluate the detection performance. The results are shown in Table 17. It can be seen that the detection results of all reference samples are consistent with the theoretical results, indicating that the gap-PCR detection performance for thalassemia is excellent.
[0177] Table 17 Compliance of Test Results for Reference Samples
[0178]
[0179] In summary, this invention discloses a probe combination designed to capture 3908 regions, which integrate 279 genes related to single-gene genetic diseases, 331 high-frequency gene loci, and 74 safe medication gene loci (a capture region refers to a segment of base information that may capture multiple target loci; a capture region with loci close together can capture information from multiple loci). It also discloses a kit and device based on this probe combination for newborn screening of single-gene genetic diseases. This product covers a large number of genes, has high detection accuracy, and can efficiently achieve large-scale sample testing, overcoming the problems of low throughput and limited detection range in existing methods. It provides reliable data support for further research, directly targeting the causes of newborn genetic diseases, and is suitable for large-scale disease screening and subsequent reproductive guidance. It is beneficial for precision diagnosis and treatment and the early prevention and control of genetic diseases, possessing extremely high practicality and application prospects.
Claims
1. A kit for newborn genetic disease screening, characterized in that, The kit contains a probe set for newborn genetic disease screening; The detection sites of the probe set are single-gene genetic disease gene combinations, high-frequency gene sites, and safe drug use gene sites; The gene combinations of the monogenic genetic diseases are GAA, GLA, GALC, F9, GALT, PAH, GBA, ABCG8, CYP27A1, IVD, COL1A1, ATP7B, FBN1, HEXB, ATP7A, HBB, PTPN11, DMD, DBT, BCKDHA, BCKDHB, ABCG5, MEFV, ARG1, FAH, SMPD1, WT1, BTD, RB1, PYGM, GCDH, TAFAZZIN, GALK1, ALDH7A1, PRODH, ALDH4A1, SMN1, SLC26A4, TSC1, TSC2, TAT, NPC2, LYST, PYGL, SLC37A4, G6PC1, GBE1, WAS, SLC25A13, ETHE1, ALPL, MYO7A, SLC22A5, TCN2, SCN1A, CYBB, FANCA, FANCC, LAMA3, LDLR, TG, HPD, HYAL1, ARSB, CTNS, AGL, PHKB, PHKG2, ETFA, ETFB, ETFDH, SUGCT, ALDOB, USH1C, PCDH15, HEXA, ACAT1, GALE, PCDH19, TPO, DUOX2, GAMT, GUSB, HLCS, PHKA2, PAX3, ARSA, USH2A, ADGRV1, AGXT, GRHPR, HOGA1, SLC12A3, IDUA, ALDH5A1, INSR, ABCC8, KCNJll, HADH, MMUT, DUOXA2, PROP1, OAT, HGSNAT, SLC39A4, LIPA, RMRP, MITF, SOX10, PGM1, BCL11B, OPA3, AUH, DNAJC19, PEX1, PEX6, CD40LG, BTK, IL2RG, MT-RNR1, MLYCD, GSDME, MYO1SA, TMC1, TMPRSS3, OTOF, GH1, GHRHR, DLD, CPT1A, CYP27B1, HSD3B7, ACAD8, PIGA, ACADS, PCCA, HBA, MT-ND5, PCCB, HBA2, COL1A2, MT-TL1, ADA, ACVRL1, BMPR1B, BMPR2, CAV1, ENG, KCNK3, SMAD9, TBX4, CLPB, JAG1, GLB1, ANOS1, FGFR1, PROKR2, CHD7, MT-ND4, NAGS, CPS1, MYH9, GJB3, CDH23, GJB2, TECTA, ILDRl, MARVELD2, LRTOMT, GNMT, ACSF3, OTC, SLC2A1, SLC25A20,CPT2、VDR、STAR、HADHA、HADHB、COQ4、ACADM、ELANE、XIAP、HSD17B10、COL4A5、PHEX、SH2D1 A、KRT5、ACADVL、G6PD、KRT14、ACADSB、PTS、GCH1、QDPR、PCBD1、ATL1、SPAST、CLCN1、TTPA、 PEX12、PEX2、PEX10、PEX26、LAMB3、ATP8B1、ABCB11、ABCB4、ADK、GM2A、MTHFR、COL7A1、REE P1、SPG11、MAT1A、BCAT2、NPC1、DNAJC12、NADK2、COL4A4、TH、MT-TH、MCCC1、MCCC2、COL4A3 THBD, TSH, PAX8, TSHB, THRA, AMT, LAMC2, PRF1, UNC13D, MCEE, HMGCS2, UGT1A1, HMGCL, RPL11, CBS, SPR, RPS26, FOXP3, THRB, ESPN, MMAA, MMAB, MMADHC, LMBRD1, ABCD4, HCFC1, CYP 17A1、GALNS、AHCY、SUCLA2、SUCLG1、SGSH、NAGLU、SERAC1、MTRR、MTR、CYP11B1、HSD3B2、JA K3、SLC52A3、SLC52A2、NR0B1、IL7R、MMACHC、C3、RAG1、PRDX1、CD46、RAG2、CFB、CFI、DGKE;、 The high-frequency gene loci are rs769901, rs9376523, rs2131379, rs4084233, rs10754872, rs6931199, rs2098432, rs11082612, rs4654390, rs3127178, rs1895983, rs7241484, rs12066442, rs28418962, rs4765461, rs8085547, rs9425128, rs7797255, rs9506919, rs11876372, rs473027, rs9918562, rs471283, and rs510415. rs1199668, rs778723, rs6563522, rs12956479, rs978346, rs10486867, rs790534, rs2581648, rs4970784, rs10252204, rs7327223, rs374625, rs4786 , rs1158492, rs1413111, rs2436493, rs6672099, rs10249015, rs4591014, rs57167556, rs3012195, rs1858827, rs1411139, rs4804522, rs4844585, rs 2970478, rs7982833, rs55969070, rs12751210, rs846930, rs4885089, rs1961562, rs753383, rs7778877, rs1333081, rs7508025, rs1559472, rs11978 304, rs9556519, rs2358963, rs212756, rs292606, rs9523304, rs34876435, rs10198193, rs16882516, rs6492840, rs10425026, rs4672484, rs886674, rs9284217, rs7254716, rs7600680, rs1968853, rs10135319, rs67258954, rs3771786, rs8180910, rs7142453, rs892055, rs11900333, rs1463362, rs1 0143161, rs7252127, rs584811, rs7825817, rs2153751, rs338598, rs11675245, rs11787342, rs2024809, rs371690, rs6730627, rs973623, rs9323355,rs601037、rs843436、rs16919692、rs1054218、rs10411961、rs10176993、rs10808747、rs7152091、rs3818331、rs2709544、rs1900080、rs10143492、rs556274、rs1641382、rs1096312、rs7158881、rs6086259、rs2623033、rs6997147、rs2998316、rs6081157、rs355894、rs4961023、rs234601、rs6035687、rs4972722、rs58278856、rs2403039、rs169225、rs11678036、rs8079、rs8030727、rs34355888、rs7564327、rs4876384、rs7167214、rs2425202、rs10931830、rs56926470、rs2381817、rs2294952、rs2078429、rs4350016、rs489964、rs2425486、rs360834、rs10811723、rs2292465、rs3859583、rs6551088、rs1887221、rs4300583、rs926672、rs445938、rs10971956、rs1550327、rs4558646、rs12494480、rs10125303、rs7498094、rs212600、rs12497462、rs7038645、rs2682925、rs380146、rs2290600、rs4330746、rs2008262、rs4809448、rs7642443、rs4742667、rs11635236、rs1034344、rs9861011、rs10818906、rs4439767、rs2824285、rs7623854、rs11244203、rs12597653、rs9974998、rs11709958、rs1242986、rs724614、rs61228441、rs9886989、rs10734049、rs7193079、rs2828997、rs10938839、rs2431062、rs26763、rs8132458、rs4235294、rs769009、rs6565375、rs974607、rs4407564、rs7098033、rs7204877、rs2833118、rs4320169、rs7900524、rs1345400、rs35835205、rs60716049、rs2393989、rs9940844、rs2298678、rs980363、rs12771265、rs7186429、rs2003624、rs1980187、rs1259590、rs8063990、rs9636909、rs4585306、rs10881958、rs16970800、rs10439673、rs35638846、rs7916334、rs6564299、rs375886、rs6844049、rs11192326、rs9891529、rs2282527、rs6857303、rs3903866、rs7215079、rs2150457、rs2358906、rs2164781、rs2074276、rs401416、rs1393581、rs7123041、rs12943366、rs75766、rs1919508、rs7104864、rs7225149、rs6003582、rs1422925、rs7119733、rs225196、rs1883278、rs16889179、rs4752894、rs712039、rs5762870、rs10078531、rs1945232、rs2301647、rs5997898、rs276595、rs592271、rs199453、rs1076912、rs462647、rs1249528、rs7216847、rs4821733、rs446219、rs7929510、rs1985749、rs9611528、rs1559054、rs317187、rs2252814、rs926328、rs6881764、rs986246、rs56152251、rs2281537、rs2914156、rs1218923、rs1222599、rs6530478、rs10066380、rs1632023、rs6501801、rs1884689、rs6859953、rs2279013、rs236037、rs1264012、rs10067081、rs10790913、rs9807436、rs4546820、rs7715167、rs1820608、rs12457503、rs10521499、rs6928796、rs2433651、rs4447532, rs5932451, rs6935447, rs2882855, rs948456, rs11575897, rs594008, rs11182150, rs4800924, rs773764555, rs998828, rs640081, rs12604660, rs769299826, rs2815715, rs4759294, rs1493911, rs34402762, rs7356837, rs1599750, rs7230616, rs35284970, rs13219667, rs2335882, rs1787623, rs2071394, rs6902551, rs7961750, rs307087, rs754688229, rs4895801, rs10777027 and rs788433; The safe medication gene loci are ABCB1: c.3435T>C, ABCB1: c.2677T>A, ABCB1: c.2677T>G, ADD1: c.1378G>T, ADD1: c.1378G>A, ADRB1: c.1165G>C, ADRB2: c.46A>G, AGTR1: c.*86A>C, ALDH2: c.1510G>A, ANKK1: c.2137G>A, APOE: c.388T>C, APOE: c.526C>T, APOE: c.526C>G, CHIA: c.304G>A / C, COMT: c.472G>A, CYP2B6: c.516G>T, CYP2C19: c.636G>A, CYP2C19: c.681G>A, CYP2C9: c.430C>T, CYP2C9: c.1075A>C, CYP2D6: g.4180G>C, CYP2D6: g.4172C>G, CYP2D6: g.4172C>T, CYP2D6: g.2988G>A, CYP2D6: g.2850C>T, CYP2D6: g.1846G>A, CYP2D6: g.1758G>T, CYP2D6: g.1758G>A, CYP2D6: g.997C>G, CYP2D6: g.997C>T, CYP2D6: g.984A>G, CYP2D6: g.100C>T, CYP4F2: c.1297G>A, EPHX1: c.337T>C, EPHX1: c.416A>G, UGT1A4: c.142T>G, UGT1A4: c.142T>A, G6PD: c.1388G>A, G6PD: c.1376G>G, G6PD: c.1376G>T, G6PD: c.1376G>A, G6PD: c.1360C>T, G6PD: c.1024C>T, G6PD: c.1004C>T, G6PD: c.871G>A, G6PD: c.563C>T, G6PD: c.519C>T, G6PD: c.196T>A, G6PD: c.95A>G, ITPA: c.94C>G, ITPA: c.94C>A, ITPA: c.124+21A>C, NAT2: c.282C>T, NAT2: c.341T>C, NAT2: c.481C>T, NAT2: c.590G>A, NAT2: c.803G>A, NAT2: c.857G>A, NUDT15: c.52G>A, NUDT15: c.415C>T, NUDT15: c.416G>A, OPRM1: c.118A>G, POLG: c.1399G>A, PPARG: c.34C>G, SCN2A: c.56G>A, SCN2A: c.971-32A>G, SLC22A1: c.1222A>G, SLC22A1: c.1222A>C, SLC22A2: c.808T>G, SLCO1B1: c.521T>C, STXBP1: c.922A>T, TPMT: c.719A>G / C, UGT1A1: c.211G>A and UGT2B15: c.253T>G;. The sequence of the probe group is shown in SEQ ID NO: 5~14913.
2. The reagent kit according to claim 1, characterized in that: The kit includes a reagent for detecting large fragment deletions in thalassemia.
3. The use of the kit according to any one of claims 1 to 2 in the preparation of products for newborn genetic disease screening.
4. A newborn genetic disease screening system, characterized in that: The screening system includes: an indicator input module and a newborn genetic disease assessment module; The indicator input module includes at least the following steps: using the probe set in the kit of claim 1 or 2 to construct a library, hybridize and capture the subject sample, and then sequence the sample to obtain the sequencing information of the target gene in the subject sample; The newborn genetic disease assessment module includes at least the following: performing bioinformatics analysis on the sequencing results obtained from the indicator input module to obtain variation results, and outputting the assessment results of the subject.
5. The screening system according to claim 4, characterized in that: The input module includes methods for obtaining sequencing information of the target gene in the subject sample using the following methods: Using the kit described in claim 1 or 2, genomic DNA is extracted from the test sample, DNA is fragmented by enzyme digestion, adapter ligation, and library construction is performed. The constructed library is then hybridized, captured, enriched, and sequenced for analysis.
6. The screening system according to claim 5, characterized in that: The connector is a Y-type connector, which includes a first sequence and a second sequence, the first sequence being shown in SEQ ID NO: 1 and the second sequence being shown in SEQ ID NO:
2.
7. A device for newborn genetic disease screening, wherein the device can use the probe set in the kit of any one of claims 1 to 2 to construct a library for hybridization and capture of the sample to be tested, followed by sequencing, and perform bioinformatics analysis on the data generated after sequencing to obtain the mutation results.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the screening process of the screening system according to any one of claims 4 to 6.
9. A storage medium storing processor-executable instructions, characterized in that, The processor-executable instructions, when executed by the processor, are used to perform the screening process of the screening system according to any one of claims 4 to 6.