Newborn genetic disease screening gene combination, probe combination, kit and preparation method and device thereof

By designing a combination of probes and kits covering a variety of single-gene genetic diseases and high-frequency gene loci, the problem of small detection range and insufficient accuracy in the prior art is solved, and rapid and accurate screening of neonatal genetic diseases is achieved, which is suitable for large-scale disease screening.

CN120060462AActive Publication Date: 2025-05-30CAPITALBIO GENOMICS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411778021.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-30
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

The existing neonatal screening technology is difficult to detect multiple single-gene genetic diseases at the same time, and the detection range and accuracy are insufficient, which cannot meet the needs of large-scale disease screening.

Method used

A probe combination was designed, covering 279 single-gene genetic diseases, 331 high-frequency gene loci and 74 safe drug-use gene loci. Through this probe combination and kit, a comprehensive screening of neonatal genetic diseases was achieved.

Benefits of technology

It has achieved one-time detection of more than 40,000 pathogenic mutations in 279 genetic disease-related genes, providing fast and accurate screening results, and can formulate individualized drug regimens for clinical practice to improve the efficiency and accuracy of neonatal screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005172009440000041
    Figure BDA0005172009440000041
  • Figure BDA0005172009440000052
    Figure BDA0005172009440000052
  • Figure BDA0005172009440000061
    Figure BDA0005172009440000061
Patent Text Reader

Abstract

The invention belongs to the field of gene detection, and particularly relates to a neonatal genetic disease screening gene combination, a probe combination, a kit and a preparation method and device thereof. The kit provided by the invention comprises capture probes aiming at definite pathogenic genes, high-frequency gene loci and safe medication gene loci, more than 40,000 pathogenic mutations on 279 genes related to genetic diseases can be detected at one time, and the pathogenic mutation condition of a subject can be quickly and accurately known; 331 high-frequency gene loci and 74 safe drug use gene loci can be detected at a time, a unique gene identity card is provided, drug interpretation is provided, and a safer and more effective drug use scheme can be formulated for a subject clinically. The detection device provided by the invention can be used for automatically analyzing and sorting sequencing data, assisting clinical workers in performing mutation interpretation in a targeted manner, improving the working efficiency and providing necessary conditions for large-scale application of neonatal screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of gene detection, and particularly relates to a gene combination, a probe combination, a kit and a preparation method and device thereof for neonatal genetic disease screening. Background Art

[0002] According to statistics, the incidence rate of birth defects in China is 5.6%, and about 70% of them often show no obvious symptoms in the early stage. Once the disease occurs, it will endanger life or cause permanent damage to intelligence and the body. The incidence rate of single diseases is low, but the comprehensive incidence rate is high, exceeding 1 / 100, far exceeding the incidence rate of fetal trisomy 21. Among the "First Batch of Rare Disease Catalogs in China" announced in 2018, monogenic genetic diseases accounted for the majority. The average incidence rate of each of them in the global population is between 1 / 10,000 and 1 / 100,000. In some cases, it occurs in one out of thousands of newborns, and in other cases, it occurs in one out of hundreds of thousands of newborns. However, for a local area or a specific population from certain regions, the average incidence rate of some monogenic genetic diseases may be as high as 1% - 2%. For example, among some ethnic minorities in Southeast Asia and southern Yunnan, the incidence rates of thalassemia and glucose-6-phosphate dehydrogenase deficiency are both over 2%. Most monogenic diseases can cause death, deformity or disability. There are only less than 10% of them with effective therapeutic drugs, but the treatment costs are expensive. Even if the children survive, most of them are permanently disabled or have intellectual disabilities. Therefore, monogenic diseases not only cause serious harm to the health of patients, but also bring heavy mental and economic burdens to families and society, and have become an important obstacle seriously affecting the quality of the newborn population and children's health in China. Therefore, it is necessary to screen newborns for some congenital and hereditary diseases that can be intervened, detect and diagnose them early, which helps to improve the survival rate of newborns and improve the quality of life of children. Therefore, newborn screening is of great importance.

[0003] Newborn screening (NBS) refers to a special examination for congenital and hereditary diseases that seriously endanger the health of newborns during the neonatal period, providing early diagnosis and assisting in maternal and child treatment (Technical specifications for newborn disease screening 2010 edition. Accessed 1 Dec 2021 [1] Newborn screening has a history of more than 60 years. The progress of technology has greatly promoted the development of newborn screening, and both the screening diseases and the screening efficiency have been greatly improved. [2] Compared with traditional newborn screening, NGS-based gene screening has a wide range and directly points to the cause, which is suitable for large-scale disease screening and subsequent fertility guidance, and is conducive to precise diagnosis and treatment and the prevention and control of genetic diseases moving forward.

[0004] A gene ID card is an identity document made using gene technology, which selects several fixed polymorphic gene loci for identification. The combination of these gene loci is unique and is a unique gene information certificate, which can be used as a reliable basis in the process of finding and identifying lost children.

[0005] Pharmacogenomics studies the effects of genetic variations and changes in the expression levels of genes involved in drug metabolism, transport, and drug action targets in the body on drug responses. It explains the reasons for individual differences in drug responses at the gene level, thereby improving the efficacy of drugs, reducing side effects, saving medical costs, and ultimately achieving individualized drug treatment. Pharmacogenomics is a hot research field that has developed rapidly and received much attention in recent years and is an important part of precision medicine. For the subjects being tested, it not only improves the safety and effectiveness of drug treatment, enabling doctors to formulate more precise individualized medication plans, but also significantly reduces the incidence of drug adverse reactions. This technology can predict drug responses, avoid drug toxicity, optimize drug doses, and plays a key role especially in the treatment of special diseases such as mental illnesses. In addition, pharmacogenomics helps to optimize long-term treatment plans and improve the management of chronic diseases, thereby enhancing the quality of life of patients.

[0006] Therefore, it is necessary to develop a detection kit with better effects for comprehensive screening of neonatal genetic diseases.

[0007] References:

[0008] [1]Technical specifcations for newborn disease screening 2010edi-tion.Accessed 1Dec 2021。

[0009] [2]Zhao ZY.Progress in international neonatal disease screening.ChinJ Child Health Care.2012;20:193–5。 Summary of the Invention

[0010] The purpose of the first aspect of the present invention is to provide a probe set for neonatal genetic disease screening.

[0011] The purpose of the second aspect of the present invention is to provide a kit for neonatal genetic disease screening.

[0012] The purpose of the third aspect of the present invention is to provide a neonatal genetic disease screening system.

[0013] The purpose of the fourth aspect of the present invention is to provide a device for neonatal genetic disease screening.

[0014] The objective of the fifth aspect of the present invention is to provide an electronic device.

[0015] The objective of the sixth aspect of the present invention is to provide a storage medium.

[0016] In order to achieve the above objectives of the present invention, the technical solutions adopted by the present invention are as follows:

[0017] The first aspect of the present invention provides a primer set for neonatal genetic disease screening. The specific sequences of the probe set are as follows:

[0018] http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF% E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88 E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97.pdf .

[0019] The probe sequence file has been notarized with a timestamp. The timestamp certificate is as follows:

[0020] http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF% E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88 E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97_TSA%E8%AF%81%E4%B9% A6.jpg . All experiments and data of this application were completed before the filing date.

[0021] In the primer set, it includes: a probe combination designed for 3908 regions composed of a neonatal screening gene combination for 279 single-gene genetic diseases, 331 high-frequency gene loci, and 74 safe medication gene locus combinations. It should be noted that: the probe capture regions are divided into three categories of loci: including loci for 279 single-gene genetic diseases, high-frequency loci, and safe medication loci. Among them, there are overlaps between the loci for single-gene genetic diseases and those for safe medication detection. Therefore, there are also overlapping capture regions. After deduplication, there are actually 3908 loci.

[0022] The second aspect of the present invention provides a kit for neonatal genetic disease screening, and the kit contains the probe set of the first aspect of the present invention.

[0023] In some embodiments of the present invention, the kit further includes thalassemia large fragment deletion detection reagents, including, for example, LA Taq DNA Polymerase, KAPA 5×GC buffer, dNTP stock, etc.

[0024] In some embodiments of the present invention, the kit further includes library construction hybridization reagents and hybridization capture magnetic beads.

[0025] The third aspect of the present invention provides the application of the probe set of the first aspect of the present invention or the kit of the second aspect of the present invention in the preparation of products for detecting neonatal genetic diseases.

[0026] The fourth aspect of the present invention provides a neonatal genetic disease screening system, which includes an index input module and a monogenic genetic disease gene evaluation module;

[0027] The index input module at least includes obtaining sequencing information of target genes in a subject's blood and / or tissue sample by using the probe set of the first aspect of the present invention or the kit of the second aspect of the present invention;

[0028] The monogenic genetic disease evaluation module includes at least performing bioinformatics analysis on the sequencing results obtained by the index input module to obtain variant results and outputting the evaluation results of the subject.

[0029] The screening flowchart of the neonatal genetic disease screening system is as Figure 1 shown.

[0030] In some embodiments of the present invention, the index input module includes obtaining sequencing information of target genes in a subject's blood and / or tissue sample by using the following method:

[0031] (1) DNA sample collection and extraction: Collect the sample to be tested and extract genomic DNA;

[0032] (2) DNA fragmentation / end repair / dA tailing: Digest the genomic DNA, perform end repair, and add dA tail;

[0033] (3) Adapter ligation: Ligate adapters to the DNA fragments obtained in the previous step to obtain a ligation product, and purify the product;

[0034] (4) Pre-hybridization amplification: Perform PCR amplification enrichment and purification on the purified product in step (3) to obtain a DNA library for hybridization capture

[0035] (5) Probe capture: Block the DNA library before hybridization, and perform hybridization capture by using the probe set described in the first aspect of the present invention;

[0036] (6) Elution and target region enrichment: Incubate and bind T1 magnetic beads with the probe capture product, perform elution after the incubation ends, perform PCR amplification on the eluted and re-dissolved product to obtain an enriched target sequence, and obtain the product to be sequenced;

[0037] (7) Sequencing: Perform next-generation sequencing on the product to be sequenced.

[0038] In some embodiments of the present invention, the joint in step 3) is a Y-shaped joint, which includes a first sequence and a second sequence. The first sequence is as shown in SEQ ID NO:1, and the second sequence is as shown in SEQ ID NO:2. -s- represents a thiol modification, -p- represents a phosphorylation modification, and NNNNN(N) represents a UMI molecular tag with a length of 5-6 nucleotides.

[0039] In some embodiments of the present invention, in steps 3) to 4), the purification is magnetic bead purification. In some embodiments of the present invention, the 25 μL reaction system for PCR amplification enrichment is: 9-12 μL of purified product, 10-15 μL of PCR amplification enzyme mixture; and 1-3 μL of tag primer mixture.

[0040] In some embodiments of the present invention, the forward sequence of the amplified tag primer in step 4) is as shown in SEQ ID NO:3, and the reverse sequence is as shown in SEQ ID NO:4; Index1 and Index2 are tag primers, which are composed of 8 specific base sequences and are used to identify, for example, they can be: GCCTATCA, CTTGGATG, TCACAGCA, TCCTACCT. -s- represents a thiol modification, [i5] represents an 8bp i5 Index sequence, and [i7] represents an 8bp i7 Index sequence.

[0041] In some embodiments of the present invention, the reaction conditions for PCR amplification enrichment are pre-denaturation at 97-99 °C for 0.5-2 min; denaturation at 95-98 °C for 5-30 sec, annealing at 58-62 °C for 25-35 sec, extension at 70-72 °C for 25-35 sec, cycling 5-7 times; final extension at 70-72 °C for 4-5 min; incubation at 3-5 °C.

[0042] In some embodiments of the present invention, the system for PCR amplification in step 6) is: 35-37 μL of elution and reconstitution product; 8-11 μL of 5×PCR amplification mixture; 1-2 μL of post-capture amplification primer; 1-2 μL of DNA polymerase; 0.5-1 μL of dNTP mixture;

[0043] In some embodiments of the present invention, the reaction conditions for PCR amplification in step 6) are pre-denaturation at 97-99 °C for 2-3 min; denaturation at 97-99 °C for 20-30 sec, annealing at 55-57 °C for 20-30 sec, extension at 70-72 °C for 1-2 min, cycling 10-12 times; final extension at 70-72 °C for 8-12 min; incubation at 3-5 °C.

[0044] In some embodiments of the present invention, in step (7), the second-generation sequencing is performed using the Illumina second-generation sequencing platform NovaSeq6000 or other domestic second-generation sequencing platforms for PE150 sequencing;

[0045] In some embodiments of the present invention, the bioinformatics analysis of the sequencing results includes the following steps:

[0046] 1) Data quality control and cleaning: Use the fastp software for data quality control and adapter removal operations;

[0047] 2) Molecular tag (UMI) identification and sequence alignment: After using GATK to label the UMI, use the BWA software to align the preprocessed reads with the human reference genome hg19;

[0048] 3) Data quality correction: Remove low-quality reads (including unmapped, poorly mapped, multiply mapped reads, repetitive sequences caused by PCR amplification bias, etc.);

[0049] 4) Variant detection: Use the GATK software for variant detection, and then identify and filter false positive sites

[0050] 5) Variant annotation: Use the annovar software to annotate the sites where variants occur to obtain information such as their genetic functions, population carrier frequencies, and disease correlations.

[0051] 6) CNV detection: Use the CNVkit software for gene CNV detection. CNVkit will calculate the copy number of each region in each sample and count the copy number of the region where the large fragment deletion of thalassemia is located; at the same time, divide the thalassemia gene region into small window regions and statistically analyze the average depth of sequencing. Taking the normal reference sample as the background, the average depth of the sample in the deletion region will decrease. In summary, it can be identified whether the sample has a deletion situation.

[0052] In some embodiments of the present invention, the detection method for thalassemia-related variants; the specific verification experimental steps are as follows:

[0053] 1) The sample DNA for detection is subjected to PCR amplification. The amplification experimental system is as follows: KAPA 5×GC buffer, 1.4 μL; 10 mM dNTP stock, 0.2 μL; LATaq DNA Polymerase, 0.1 μL; Primer pool (10 pM), 0.3 μL; DNA template (30 ng) and NF-H2O / ddH2O in total 8 μL. The amplification reaction system is as follows: pre-denaturation at 94°C for 5 min; denaturation at 95°C for 30 sec, annealing at 60°C for 75 sec, extension at 72°C for 90 sec, for 35 cycles; final extension at 72°C for 5 min; incubation at 4°C.

[0054] 2) The PCR amplification products are subjected to 1.5% agarose gel electrophoresis experiment, and the electrophoresis results are analyzed and judged.

[0055] In the fifth aspect of the present invention, a device for single-gene genetic disease gene detection is provided. The device can use the probe set of the first aspect of the present invention and the kit of the second aspect of the present invention to construct a sequencing library for the sample to be tested, and perform bioinformatics analysis on the data generated after sequencing to obtain a variant result.

[0056] In some embodiments of the present invention, the bioinformatics analysis includes the following steps:

[0057] 1) Data quality control and cleaning: Use the fastp software for data quality control and adapter removal and other operations;

[0058] 2) Molecular tag (UMI) identification and sequence alignment: After using GATK to label UMI, use the BWA software to align the preprocessed reads with the human reference genome hg19;

[0059] 3) Data quality correction: Remove low-quality reads (including unmapped, poorly mapped, multiply mapped reads, repeat sequences caused by PCR amplification bias, etc.);

[0060] 4) Variant detection: Use the GATK software for variant detection, and then identify and filter false positive sites

[0061] 5) Variant annotation: Use the annovar software to annotate the sites where variants occur to obtain information such as their genetic functions, population carrier frequencies, and disease correlations.

[0062] 6) CNV detection: Use the CNVkit software to perform gene CNV detection. CNVkit calculates the copy number of each region in each sample and statistically analyzes the copy number of the regions where large deletions of thalassemia are located. At the same time, the average depth of sequencing is statistically analyzed for small window regions divided in the thalassemia gene region. Taking the normal reference as the background, the average depth of the sample in the deletion region will decrease. Based on the above, it can be identified whether the sample has a deletion situation.

[0063] In some embodiments of the present invention, subsequent analysis is performed based on the results of 4) variant detection in a device for neonatal monogenic genetic disease gene detection. The specific analysis steps are as follows:

[0064] 1) According to the known combination of high-frequency gene loci, the detection results of high-frequency gene loci are statistically analyzed through the variant results. If not detected, it is considered wild-type. Finally, the detection results of the high-frequency gene locus combination are used as the gene ID card.

[0065] 2) According to the known combination of safe medication gene loci, the detection results of safe medication gene loci are statistically analyzed. If not detected, it is considered wild-type;

[0066] 3) According to the detection results and in combination with the drug interpretation library, drug interpretation-related information is obtained, including suggestions such as the degree of drug risk and treatment dosage.

[0067] In the sixth aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the screening process of the screening system in the fourth aspect of the present invention is realized.

[0068] In the seventh aspect of the present invention, a storage medium is provided, in which instructions executable by a processor are stored. The instructions executable by the processor are used to execute the screening process of the screening system in the fourth aspect of the present invention when executed by the processor.

[0069] The beneficial effects of the present invention are:

[0070] The kit provided in the present invention contains capture probes for clearly pathogenic genes, high-frequency gene loci, and safe medication gene loci, and can detect more than 40,000 pathogenic mutations on 279 genes related to genetic diseases at one time, realizing a rapid and accurate understanding of the pathogenic mutation situation of the tested person; it can detect 331 high-frequency gene loci and 74 safe medication gene loci at one time, provide a unique gene ID card and provide drug interpretation, helping clinicians formulate a safer and more effective medication plan for the tested person.

[0071] The detection device provided in the present invention can automatically analyze and sort the sequencing data obtained by using the kit in the present invention by using bioinformatics analysis methods, assist clinicians in interpreting variations in a targeted manner, improve work efficiency, and provide necessary conditions for the large-scale application of neonatal screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] The present invention will be further described below in conjunction with the drawings and embodiments, where:

[0073] Figure 1 is the overall flowchart of the present invention.

[0074] Figure 2 is the display diagram of the average depth of the 4.2 heterozygous sample window.

[0075] Figure 3 is the display diagram of the average depth of the --SEA / -α3.7 sample window. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] The concept and technical effects of the present invention will be clearly and completely described below in conjunction with the embodiments to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative efforts shall fall within the scope of protection of the present invention.

[0077] The overall process of the present invention Figure 1 is shown as follows.

[0078] In the embodiments of the present invention, the specific information of the adapter sequences and specific primer tag sequences used is shown in Table 1 (SEQ ID NO: 1 to SEQ ID NO: 4); in the embodiments of the present invention, the DNA library construction kit of Yeasen Company (product number: 12205) is used for library construction; based on the probe capture regions of the pathogenic genes related to single-gene genetic diseases provided by the present invention, the corresponding probes can be synthesized according to conventional methods. The synthesis of the probes only needs to provide the information of the capture regions to a professional probe design and synthesis company to complete.

[0079] Table 1 Adapter sequences and specific primer tag sequences

[0080]

[0081]

[0082] Note: In the table, -s- represents thiol modification, -p- represents phosphorylation modification, [i5] represents the 8bp i5 Index sequence, [i7] represents the 8bp i7 Index sequence, and NNNNN(N) represents the UMI molecular tag with a length of 5-6 nucleotides.

[0083] Example 1: Probe set for detecting 279 genes related to 309 single-gene genetic diseases

[0084] After consulting a large number of literatures and databases, the inventors obtained 309 genetic diseases, single-gene genetic diseases with clear phenotypes that can reduce the occurrence of complications and avoid the deterioration of the condition through timely intervention and treatment and their related genes (see Table 2 for details). Through high-throughput sequencing experiments screening and related experimental verification, 3536 probe capture regions were designed to detect the pathogenic gene loci of single-gene genetic diseases. The probes jointly used in the capture regions can detect pathogenic gene mutations of more than 40,000 single-gene genetic diseases. The designed probe set has strong specificity, high sensitivity, is simple and economical, and is conducive to the popularization and clinical application of neonatal screening for single-gene genetic diseases. The probe sequence information and probe sites can be found at http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF%E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88%E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97.pdf.

[0085] Table 2 Single-gene recessive genetic diseases and their related genes

[0086]

[0087]

[0088]

[0089]

[0090]

[0091] Example 2: Probe set for detecting 330 genes at 331 high-frequency gene loci

[0092] The inventor consulted the East Asian population frequency (EAS) in the 1000 Genomes Project database and obtained 331 gene loci with relatively high population frequencies and their associated population frequencies (see Table 3 for details). Through high-throughput sequencing experiments screening and related experimental verification, 329 probe capture regions were designed to detect high-frequency gene loci (a gene has multiple loci. The capture region can capture the base information of a section of the region, and there can be multiple loci on this section of the region. If the loci are relatively close, one capture region can capture the information of multiple loci). The detection results combination of high-frequency gene loci provides a unique gene identification card. The probe sequence information and the probe capture loci can be found at http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF%E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88%E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97.pdf.

[0093] Table 3 High-frequency gene loci and their population frequencies

[0094]

[0095]

[0096]

[0097] Example 3 49-gene detection probe set for 74 safe medication gene loci

[0098] The inventor consulted a large number of literatures and the PharmGKB database and obtained 74 safe medication gene loci, 32 genes and their related detected drugs (see Table 4 for details). Through high-throughput sequencing experiments screening and related experimental verification, 49 probe capture regions were designed to detect safe medication gene loci. Drug interpretation is obtained through the detection results of gene loci and haplotypes, which helps clinicians to develop a more safe and effective individualized medication plan for the tested subjects. The probe sequence information and the probe capture loci can be found at http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF%E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88%E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97.pdf.

[0099] Table 4 Safe Medication Gene Loci and Their Related Test Drugs

[0100]

[0101]

[0102] Kit for Neonatal Monogenic Genetic Disease Gene Detection in Example 4

[0103] In this example, an exemplary kit for neonatal genetic disease gene detection based on reversible chain termination sequencing method is provided. It should be noted that the kits for single-gene genetic disease gene detection developed based on other sequencing platforms or non-sequencing-based using the probe combination of the present invention should also be within the protection scope of this application.

[0104] Kit for neonatal genetic disease gene detection, comprising:

[0105] 1) Hybrid capture probe;

[0106] 2) Library construction hybridization reagent: The library construction hybridization kit in this example is used for sequencing library construction.

[0107] 3) Hybrid capture magnetic beads;

[0108] 4) Thalassemia large deletion detection reagent (gap-PCR): The gap-PCR reagent in this example. The remaining electrophoresis reagents can be prepared by yourself. Such as agarose, 1kb ladder plus, nucleic acid staining agent, etc.

[0109] Note: The following reagents can be prepared by yourself: extraction reagent, commercially available nucleic acid extraction and purification reagent, such as the nucleic acid extraction or purification reagent produced by Dongguan Boao Muhua Gene Technology Co., Ltd., product number: S10040; purification magnetic beads, such as Beckman's Agencourt AMPure XP Kit.

[0110] The experimental systems and experimental reagents listed above. KAPA 5×GC buffer, 1.4 μL; 10 mM dNTPstock, 0.2 μL; LA Taq DNA Polymerase, 0.1 μL; Primer pool (10 pM), 0.3 μL; DNA template (30 ng) and NF-H2O / ddH2O total 8 μL; The amplification reaction system is as follows: pre-denaturation at 94°C for 5 min; denaturation at 95°C for 30 sec, annealing at 60°C for 75 sec, extension at 72°C for 90 sec, cycle 35 times; final extension at 72°C for 5 min; incubation at 4°C.

[0111] Example 5 Library Construction and Sequencing of Samples Using the Kit for Neonatal Genetic Disease Gene Detection

[0112] In this embodiment, an exemplary specific method for library construction and sequencing of a sample using the kit in Embodiment 4 is provided as follows:

[0113] (1) Sample collection and extraction

[0114] The selected peripheral blood sample and dried blood spot sample are subjected to genomic DNA extraction using a nucleic acid extraction or purification reagent (it is recommended to use the nucleic acid extraction or purification reagent produced by Dongguan Boao Muhua Gene Technology Co., Ltd., product number: S10040). The extracted genomic DNA is subjected to purity detection and preliminary concentration assessment using Nanodrop2000, and then the integrity of the extracted genomic DNA is verified by agarose gel electrophoresis.

[0115] (2) DNA fragmentation, end repair and A addition: The genomic DNA is digested with enzymes, end-repaired, and dA tails are added. The 30 μL reaction system is: 200 ng of the DNA of the sample to be tested, 5 μL of fragmentation end-repair and A buffer, 2.5 μL of fragmentation end-repair and A enzyme, and the balance is nuclease-free water; The reaction conditions are set as follows: 4 °C for 1 min, 30 °C for 12 min, 65 °C for 20 min, 4 °C, Hold.

[0116] (3) Adapter ligation and purification: After the previous reaction is completed, adapter ligation is directly carried out. The reaction system is: 15 μL of rapid ligation buffer, 2.5 μL of rapid ligation enzyme, and 2.5 μL of adapter. The reaction conditions are set as follows: 20 °C for 15 min, 4 °C, Hold. The ligation product is purified using purification magnetic beads. The specific steps are as follows: Take 50 μL of the ligation product in a centrifuge tube, add 30 μL of magnetic beads, let it stand for 5 min, then place it on a magnetic rack, discard the supernatant and retain the magnetic beads, wash twice with 75% absolute ethanol, add 11 μL of Nuclease-free Water, aspirate the supernatant into a new centrifuge tube, and the purified product proceeds to the next amplification.

[0117] Among them, the adapter is a Y-shaped adapter, including: a first sequence (as shown in SEQ ID NO.1); and a second sequence (as shown in SEQ ID NO.2). Pre-hybridization amplification: The purified product in step (3) is subjected to PCR amplification enrichment. The system for PCR amplification enrichment is: 10 μL of the purified product, 12.5 μL of the PCR amplification enzyme mixture, and 2.5 μL of the tag primer mixture.

[0118] (4) The forward sequence of the tag primer is as shown in SEQ ID NO.3, and the reverse sequence is as shown in SEQ ID NO.4; Index1 and Index2 are tag primers, which consist of 8 specific base sequences and are used to identify different samples. For example, they can be: GCCTATCA, CTTGGATG, TCACAGCA, TCCTACCT. The reaction conditions for PCR amplification enrichment are: pre-denaturation at 98°C for 1 min; denaturation at 98°C for 10 sec, annealing at 60°C for 30 sec, extension at 72°C for 30 sec, for 4 cycles; final extension at 72°C for 5 min; incubation at 4°C.

[0119] (5) Probe capture: Pool and block the constructed DNA library, and prepare the probe mixture; hybridize and capture the probe mixture with the blocked DNA library; the hybridization capture reaction system is shown in Table 5.

[0120] Table 5 Hybridization capture reaction system

[0121]

[0122] The probe mixture is prepared as follows: nuclease-free water 2.25 uL, RNase inhibitor 0.3 uL, probe 1 uL; prepared on ice.

[0123] The reaction conditions for hybridization capture are shown in Table 6.

[0124] Table 6 Hybridization capture reaction conditions

[0125]

[0126] (6) Elution and target region enrichment: Incubate and bind streptavidin T1 magnetic beads with the probe capture product for 60 min. After incubation, perform elution, and perform PCR amplification on the eluted and re-dissolved product to obtain the enriched target sequence and the product to be sequenced; the amplification reaction system is: eluted and re-dissolved product 36.5 μL; 5×PCR amplification mixture 10 μL; post-capture amplification primer 1 μL; DNA polymerase 1 μL; dNTP mixture 0.5 μL;

[0127] The reaction conditions for the PCR amplification enrichment are: pre-denaturation at 98°C for 2 min; denaturation at 98°C for 30 sec, annealing at 57°C for 30 sec, extension at 72°C for 1 min, for 10 cycles; final extension at 72°C for 10 min; incubation at 4°C.

[0128] (7) Sequencing and analysis: Perform second-generation sequencing on the product to be sequenced, and analyze the sequencing results using bioinformatics.

[0129] Example 6 Use the neonatal genetic disease gene detection kit to perform gap-PCR experiments on samples

[0130] In this embodiment, an exemplary specific method for performing a gap-PCR experiment on a sample using the kit in Embodiment 4 is provided as follows:

[0131] 1) Sample collection and extraction: Collect the sample to be tested and extract genomic DNA;

[0132] 2) DNA dilution: Measure the concentration of DNA and dilute it to 10 ng / μL;

[0133] 3) gap-PCR amplification: While performing gap-PCR amplification, prepare a 1.5% agarose electrophoresis gel for standby;

[0134] 4) Electrophoresis: Perform electrophoresis at 150 V for 40 min and interpret the results according to 3 positive control products.

[0135] The experimental system and experimental reagents listed above. KAPA 5×GC buffer, 1.4 μL; 10 mM dNTP stock, 0.2 μL; LA Taq DNA Polymerase, 0.1 μL; Primer pool (10 pM), 0.3 μL; DNA template (30 ng) and NF-H2O / ddH2O total 8 μL; The amplification reaction system is as follows: Pre-denature at 94°C for 5 min; Denature at 95°C for 30 sec, anneal at 60°C for 75 sec, extend at 72°C for 90 sec, cycle 35 times; Finally extend at 72°C for 5 min; Keep at 4°C.

[0136] Evaluation of the detection accuracy of pathogenic gene loci, high-frequency gene loci, and safe drug use gene loci in neonatal genetic diseases in Example 7

[0137] In this embodiment, according to the method described in Embodiment 5, library construction and sequencing are performed on 18 positive reference products and 5 negative reference products (positive and negative samples that have been clinically tested and verified by our company), and a self-developed analysis process is used to obtain the mutation information of more than 40,000 pathogenic gene loci of single-gene genetic diseases in each sample. The detection results are compared with the known theoretical results to evaluate the detection accuracy. The results are shown in Table 7. It can be seen that the detection results of the positive reference products and the negative control products are all consistent with the theoretical genotypes, indicating that this kit has excellent detection accuracy.

[0138] Table 7 Compliance of the detection results of reference products at pathogenic gene loci of single-gene genetic diseases

[0139]

[0140]

[0141] Evaluation of the Detection Accuracy of High-Frequency Gene Loci in Example 8

[0142] In this example, according to the method described in Example 5, libraries were constructed and sequenced for 3 reference samples (positive and negative samples that have been clinically tested and verified by our company). Using the self-developed analysis process, mutation information of 331 high-frequency gene loci in each sample was obtained. The detection results were compared with the known theoretical results to evaluate the detection accuracy. As shown in Table 8, it can be seen that the detection results of the reference samples are all consistent with the theoretical genotypes, indicating that this kit has excellent detection accuracy.

[0143] Table 8. Compliance of the Detection Results of Reference Samples at High-Frequency Gene Loci

[0144]

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152] Evaluation of the Detection Accuracy of Safe Medication Gene Loci in Example 9

[0153] In this example, according to the method described in Example 5, libraries were constructed and sequenced for 3 reference samples (positive and negative samples that have been clinically tested and verified by our company). Using the self-developed analysis process, mutation information of 74 safe medication gene loci in each sample was obtained. The detection results were compared with the known theoretical results to evaluate the detection accuracy. As shown in Table 9, it can be seen that the detection results of the reference samples are all consistent with the theoretical genotypes, indicating that this kit has excellent detection accuracy.

[0154] Table 9. Compliance of the Detection Results of Reference Samples at Safe Medication Gene Loci

[0155]

[0156]

[0157] Example 10: Evaluation of the Effectiveness of Newborn Genetic Disease Gene Screening Results

[0158] In this example, according to the method described in Example 5, libraries were constructed and sequencing was performed on 4 families (positive and negative samples that have been clinically tested and verified by our company). Using the self-developed analysis process, mutation information of more than 40,000 single-gene genetic disease gene detection sites in each sample was obtained. The detection results of the families were sorted out to evaluate the accuracy of the screening results. As shown in Table 10, it can be seen that the detection results of all family samples conform to the genetic laws, indicating that this kit has excellent effectiveness in screening single-gene genetic disease genes. Early detection and diagnosis of genetic diseases help improve the survival rate of newborns and improve the quality of life of children.

[0159] Table 10 Conformity of detection results of family samples

[0160]

[0161] Note: Variant classification: P (Pathogenic, pathogenic), LP (Likely pathogenic, likely pathogenic), P / LP (Pathogenic / Likely pathogenic, pathogenic or likely pathogenic).

[0162] Example 11 Evaluation of the effectiveness of high-frequency gene loci

[0163] In this example, according to the method described in Example 5, libraries were constructed and sequencing was performed on 3 reference products (positive and negative samples that have been clinically tested and verified by our company) with 3 replicates each. Using the self-developed analysis process, mutation information of 331 high-frequency gene loci in each sample was obtained. The consistency of the repeatability detection results of each reference product was evaluated. When the inconsistency rate is less than 0.1%, it can be determined as the same sample, so as to evaluate the effectiveness of the detection. As shown in Table 11, it can be seen that the inconsistency rate of the repeatability detection results of the reference products is 0%, indicating that this kit has excellent effectiveness in the combination of high-frequency gene loci. Since the combination of gene loci is unique, it can be used to prove unique gene information.

[0164] Table 11 Repeatability of detection results of reference products at high-frequency gene loci

[0165] Item G01 G02 G03 Qualification Pass Sites 331 331 331 Number of Identical Sites 331 331 331 Number of Non-Identical Sites 0 0 0 Non-Identity Rate 0% 0% 0%

[0166] Example 12 Evaluation of the detection effectiveness of safe medication gene loci

[0167] In this example, according to the method described in Example 5, libraries were constructed and sequencing was performed on 3 reference samples (positive and negative samples that have been clinically tested and verified by our company). Using a self-developed analysis process, mutation information of 74 safe medication gene loci in each sample was obtained. The test results were combined with a drug interpretation library to sort out drug treatment suggestions, thereby evaluating the effectiveness of the test. As shown in Table 12, it can be seen that the risk drug interpretation of the reference samples prompted relevant drug interpretations, indicating that this kit has excellent test effectiveness.

[0168] Table 12 Risk Drug Recommendations for Reference Samples

[0169]

[0170]

[0171] Example 13 Performance Evaluation of Gene Probes for Neonatal Genetic Diseases

[0172] In this example, according to the method described in Example 5, libraries were constructed and sequencing was performed on 2 standard samples (positive and negative samples that have been clinically tested and verified by our company). Using a self-developed analysis process, a total of 218,106 loci including the true set of locus variations and Clinvar-filtered loci within the capture interval of the standard samples were finally integrated into an evaluation variant set.

[0173] To evaluate the probe detection performance, the numbers of true positives (TP), true negatives (TN), false positives (FP), false negatives (FN), etc. in the evaluation variant set of the standard sample were counted. The main evaluation indicators include sensitivity (SEN), specificity (SPE), positive predictive value (PPV), F1_score (F1), etc., and the calculation formulas are as follows:

[0174] SEN = TP / (TP + FN) * 100%

[0175] SPE = TN / (FP + TN) * 100%

[0176] PPV = TP / (TP + FP) * 100%

[0177] F1 = 2 * PPV * SEN / (PPV + SEN)

[0178] From the data before filtering in Table 14, it can be seen that the detection performance for some loci is not very good. In this example, by observing the variant AF, it was found that by appropriately filtering the AF threshold, the PPV can be effectively improved.

[0179] This kit determines the filtration threshold of AF through the ROC curve. The ROC curve is a curve with the false positive rate (1 - SPE) at different thresholds as the abscissa and the sensitivity (SEN) at different thresholds as the ordinate. The maximum value of the Youden index corresponds to the optimal diagnostic critical value of this method and can be used as the optimal classification threshold. The calculation formula for the Youden index (J) is:

[0180] J = SEN + SPE – 1.

[0181] Table 13. Statistical results of the Youden index of the ROC curve

[0182]

[0183]

[0184] As can be seen from Table 13, when AF = 0.12, the Youden index (J) reaches the maximum value, the model result is optimal, and the optimal classification threshold of AF is 0.12. Filter the data before filtration in Table 14 according to the optimal classification threshold of AF. The filtered data shows an improvement in specificity and F1_score, while the sensitivity slightly decreases. At the same time, as can be seen from Table 15, among more than 40,000 loci within the final detection range of the kit, the F1_score reaches 1.000, indicating that the performance of the kit reaches the best state. This shows that the probes of this kit have excellent detection performance.

[0185] Table 14. Performance differences of the variant results in the evaluation variant set before and after filtration with the parameter AF ≥ 0.12

[0186]

[0187] Table 15. Performance of 47,354 detection ranges within the detection range

[0188] Sample PPV SEN SPE F1 S1 100.00% 100.00% 100.00% 1.0000 S2 100.00% 100.00% 100.00% 1.0000

[0189] Example 14. Detection performance of thalassemia major deletions

[0190] In this example, according to the method described in Example 5, libraries were constructed and sequenced for --SEA / -3.7 samples, 4.2 heterozygous samples, and 3 normal reference products, and the detection results were obtained using the self-developed analysis process. In this example, two methods were used to detect thalassemia major deletions: 1. Analyze the detection results using the cnvkit software; 2. Establish small windows for the sequences in the thalassemia gene region and calculate the average depth of the windows. Use the detection results of the normal reference products as the background reference to analyze 2 samples with large fragment deletions. Combining the above, the detection performance was evaluated.

[0191] The results of the cnvkit software are shown in Table 16. It can be seen that for the --SEA / -3.7 sample, the copy number (CN) in the 3.7 region is <0.1, and in the SEA region is <0.5, and the judgment result is --SEA / -3.7; for the 4.2 heterozygous sample, the CN in the 4.2 region is about 0.5, and the judgment result is 4.2 heterozygous. The average depth results of the small windows in the thalassemia gene region are as Figure 1 and Figure 2 shown. It can be seen that the average depth of the --SEA / -3.7 sample and the 4.2 heterozygous sample in the corresponding region is significantly lower than that of the relative normal reference, indicating a deletion in this region. The above detection results are all consistent with the theoretical results, indicating that the detection performance of thalassemia is extremely excellent.

[0192] Table 16 Display of the results of the cnvkit software

[0193] Missing Samples CN in Region 4.2 CN in Region 3.7 CN in SEA Region Judgment Result Whether it Conforms --SEA / -3.7 0.312376982 0.077664564 0.379467247 --SEA / -3.7 Conforms 4.2 Heterozygous 0.585625371 0.868311849 0.861465216 4.2 Heterozygous Conforms Normal Sample 1 0.996828341 0.993110674 0.984584702 No Missing Detected Conforms Normal Sample 2 0.998680237 1.00213872 0.982358102 No Missing Detected Conforms Normal Sample 3 1.000518515 0.99147959 0.983171837 No Missing Detected Conforms

[0194] Example 15 Detection performance of gap-PCR for large deletions in thalassemia

[0195] In this example, according to the method described in Example 6, gap-PCR amplification was performed on 18 reference products with large fragment deletions and 7 normal reference products (positive and negative samples that have been clinically tested and verified by our company). The results were interpreted based on the control product results and the electrophoresis results of the reference products. The detection results were compared with the known theoretical results to evaluate the detection performance. The results are shown in Table 17. It can be seen that the detection results of all reference products are consistent with the theoretical results, indicating that the detection performance of gap-PCR for thalassemia is extremely excellent.

[0196] Table 17 Compliance of the detection results of the reference products

[0197]

[0198]

[0199] In summary, the present invention discloses a probe combination designed by combining 279 genes related to single-gene genetic diseases, 331 high-frequency gene loci, and 74 gene loci for safe drug use into 3,908 regions (the capture region refers to capturing the base information of a region, and there may be multiple target sites on this region. A capture region with closely spaced sites can capture the information of multiple sites). In addition, a kit and a device based on this probe combination and applicable to the neonatal screening of single-gene genetic diseases are also disclosed. This product covers a large number of genes, has high detection accuracy, can efficiently detect a large number of samples, overcomes the problems of low throughput and small detection range in the existing methods, provides reliable data support for further research, directly points to the causes of neonatal genetic diseases, is suitable for large-scale disease screening and subsequent fertility guidance, is conducive to precise diagnosis and treatment and the prevention and control of genetic diseases moving forward, and has extremely high practicality and application prospects.

Claims

1. A probe set for screening genetic diseases in newborns, characterized in that: The specific sequence of the probe set is shown in: http: / / 58.252.61.49:8003 / probe / %E6%96%B0%E7%94%9F%E5%84%BF%E9%81%97%E4%BC%A0%E7%97%85%E7%AD%9B%E6%9F%A5V1%E7%89%88%E6%8E%A2%E9%92%88%E5%BA%8F%E5%88%97.pdf.

2. A kit for screening genetic diseases in newborns, characterized in that: The kit contains the probe set according to claim 1.

3. The kit according to claim 2, characterized in that: The kit comprises a thalassemia large fragment deletion detection reagent.

4. Use of the probe set according to claim 1 or the kit according to any one of claims 2 to 3 in preparing a product for screening newborn genetic diseases.

5. A newborn genetic disease screening system, characterized in that: The screening system includes: an index input module and a neonatal genetic disease assessment module; The index input module at least includes obtaining the sequencing information of the target gene in the subject sample using the probe set according to claim 1 or the kit according to claim 2 or 3; The neonatal genetic disease assessment module includes at least performing bioinformatics analysis on the sequencing results obtained by the indicator input module to obtain variation results and output assessment results of the subject.

6. The screening system according to claim 5, characterized in that: The input module includes obtaining the sequencing information of the target gene in the subject sample using the following method: The kit according to claim 3 is used to extract genomic DNA from the test sample, fragment the DNA by enzyme digestion, connect the adapter, and construct a library, and then sequence and analyze the constructed library after hybridization capture enrichment.

7. The screening system according to claim 6, characterized in that: The connector is a Y-shaped connector, and the Y-shaped connector includes a first sequence and a second sequence, wherein the first sequence is shown as SEQ ID NO: 1, and the second sequence is shown as SEQ ID NO:

2.

8. A device for screening genetic diseases in newborns, which can use the probe group described in claim 1 and the kit described in any one of claims 2 to 3 to construct a sequencing library for a sample to be tested, and perform bioinformatics analysis on the data generated after sequencing to obtain variation results.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the screening process of the screening system according to any one of claims 5 to 7 is implemented.

10. A storage medium storing instructions executable by a processor, characterized in that: The processor-executable instructions are used to perform the screening process of the screening system according to any one of claims 5 to 7 when executed by the processor.

Citation Information

Patent Citations

  • RNA probe capable of detecting multiple neonatal hereditary diseases and gene screening kit

    CN104673925A

  • Detection probe group and screening method for 50 hereditary disease genes of newborns

    CN107937513A

  • Carrier screening gene combination, probe combination and kit for single-gene hereditary disease as well as preparation method and device of carrier screening gene combination and probe combination

    CN118562946A

  • Kit for detecting chromosome aneuploidy and single gene mutations and use

    WO2023240755A1