Primer group, kit and method for detecting STRC gene variation and application

By combining long-chain PCR and third-generation sequencing, a specific primer combination was designed for STRC gene detection, which solved the problems of low specificity and insufficient accuracy in existing technologies, and achieved efficient, rapid and low-cost STRC gene variant detection.

CN121852532APending Publication Date: 2026-04-14SOOCHOW UNIV AFFILIATED CHILDRENS HOSPITAL +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-30
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies suffer from low specificity and insufficient accuracy in detecting STRC gene variants. In particular, when dealing with highly homologous STRC genes and STRCP1 pseudogenes, it is difficult to distinguish variant types and cis-trans relationships, and the detection cost and time are relatively high.

Method used

Using a method based on long-chain PCR and third-generation sequencing, specific primer combinations were designed for targeted enrichment. Combined with third-generation sequencing technology, comprehensive detection of the full length of the STRC gene and its upstream and downstream regions was achieved. This method can distinguish copy number variations and non-allelic homologous recombination and is applicable to multiple sequencing platforms.

Benefits of technology

It enables comprehensive, accurate, and rapid detection of STRC gene variants, reducing the probability of false positives and false negatives. It has a wide detection range, short cycle time, and low cost, and can distinguish between cis-trans relationships and chain effects of variants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121852532A_ABST
    Figure CN121852532A_ABST
Patent Text Reader

Abstract

The invention discloses a primer group, a kit and a method for detecting STRC gene variation and application. The primer group comprises one or more of primer groups 1-8. Targeted enrichment based on the primer group is combined with three-generation sequencing, so that single nucleotide variation, insertion / deletion, exon copy number variation and structural variation caused by non-allelic homologous recombination of the STRC gene can be comprehensively and accurately detected, the cis-trans relationship between the variations can be defined, and an efficient tool is provided for diagnosis of hereditary hearing loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biodetection technology, specifically to primer sets, kits, methods, and applications for detecting STRC gene variations. Background Technology

[0002] Hearing loss is one of the most common sensory impairments. Hearing impairment not only means loss of hearing but is often accompanied by speech disorders and cognitive developmental delays, severely impacting patients' daily lives and social participation. Approximately 60% of congenital hearing loss is caused by genetic factors. More than 200 deafness genes are currently known, with the most common being GJB2 (congenital deafness), GJB3 (acquired high-frequency hearing loss), SLC26A4 (large vestibular aqueduct syndrome), and mtDNA12S rRNA (drug-induced hearing loss). Among these, GJB2 gene mutations are the most common cause of congenital non-syndromic hearing loss. Studies show that pathogenic variants of the STRC gene are also considered a major cause of hereditary hearing loss.

[0003] The STRC gene, located on chromosome 15q15.3, has been extensively studied in clinical practice and is considered a relatively common genetic cause of non-syndromic hearing loss. Therefore, STRC gene testing is of significant practical importance. It not only aids in the early diagnosis and differential diagnosis of hearing loss patients and their families but also provides crucial guidance for genetic screening and selective reproduction. Furthermore, it can be used for screening and diagnosing neonatal hereditary hearing loss, enabling early detection, early intervention, and early treatment.

[0004] However, a highly homologous pseudogene, STRCP1, exists downstream of the STRC gene. This pseudogene shares 98.9% homology with the full-length STRC sequence, and an even higher 99.6% homology in its coding sequence. Crucially, exons 1 through 13, 15, 17, 18, 21, 27, and 29 are completely identical to the STRC gene coding sequence (100% homology). This extremely high degree of sequence similarity severely interferes with the specificity of STRC gene detection, leading to significant challenges of low specificity and insufficient accuracy under current mainstream technologies. Summary of the Invention

[0005] Therefore, it is necessary to provide primer sets, kits, methods, and applications for detecting SRC gene variations.

[0006] The first aspect of this application provides a primer set for detecting SRC gene variants, said primer set comprising one or more of the following primer sets:

[0007] Primer set 1: Upstream and downstream primers with sequences as shown in SEQ ID NO: 1 and SEQ ID NO: 2, respectively;

[0008] Primer set 2: Upstream and downstream primers with sequences as shown in SEQ ID NO: 1 and SEQ ID NO: 3, respectively;

[0009] Primer set 3: Upstream and downstream primers with sequences as shown in SEQ ID NO: 4 and SEQ ID NO: 5, respectively;

[0010] Primer set 4: Upstream and downstream primers with sequences as shown in SEQ ID NO: 6 and SEQ ID NO: 7, respectively;

[0011] Primer set 5: Upstream and downstream primers with sequences as shown in SEQ ID NO: 8 and SEQ ID NO: 9, respectively;

[0012] Primer set 6: Upstream and downstream primers with sequences as shown in SEQ ID NO: 10 and SEQ ID NO: 11, respectively;

[0013] Primer set 7: Upstream and downstream primers with sequences as shown in SEQ ID NO: 12 and SEQ ID NO: 9, respectively;

[0014] Primer set 8: Upstream and downstream primers with sequences as shown in SEQ ID NO: 13 and SEQ ID NO: 14, respectively.

[0015] In some implementations, the 5' end of each primer in the primer set also includes a barcode.

[0016] In some embodiments, the STRC gene variants detected by the primer set include one or more of the following: single nucleotide variants, insertions or deletions, exon deletions / duplications, and copy number variations on the STRC gene.

[0017] In some embodiments, the copy number variation includes copy number variation resulting from non-allelic homologous recombination of the STRC gene and the pseudogene STRCP1, and / or copy number variation resulting from non-allelic homologous recombination of the upstream gene CKMT1B and the CKMT1A gene of the STRC gene.

[0018] The second aspect of this application provides the use of the primer set for detecting STRC gene variations described in the first aspect of this application in the preparation of products for detecting STRC gene variations.

[0019] A third aspect of this application provides a kit for detecting STRC gene variants, the kit comprising the primer set for detecting STRC gene variants described in the first aspect of this application.

[0020] In some embodiments, the kit further includes one or more of nucleic acid extraction reagents, DNA polymerase, PCR buffer, and dNTPs.

[0021] In some embodiments, the DNA polymerase includes one or more of Phanta Max enzyme, Phanta Flash enzyme, and KOD FXNeo enzyme.

[0022] The fourth aspect of this application provides a method for detecting STRC gene variations, comprising the following steps:

[0023] Extract genomic DNA from the sample to be tested;

[0024] Using the genomic DNA as a template, PCR amplification was performed using the primer set described in the first aspect of this application or the kit described in the third aspect of this application to prepare one round of amplification products;

[0025] Equal volumes of amplification products from each barcode round were mixed, purified with magnetic beads or recovered with Pippin HT, and used to construct a third-generation sequencing library; and

[0026] The third-generation sequencing library was sequenced.

[0027] In some implementations, the PCR amplification procedure includes:

[0028] Pre-denaturation at 94℃ for 2 minutes;

[0029] Denaturation at 98℃ for 10 seconds, annealing and extension at 71℃ for 8 minutes, 2-4 cycles;

[0030] Denaturation at 98℃ for 10 seconds, annealing and extension at 70℃ for 8 minutes, 6-8 cycles;

[0031] Denaturation at 98℃ for 10 seconds, annealing and extension at 68℃ for 8 minutes, 19-21 cycles.

[0032] By utilizing the aforementioned primer set for targeted enrichment combined with third-generation sequencing, we can comprehensively, accurately, and rapidly detect single nucleotide variants (SNVs), insertions / deletions (INDELs), and exon copy number variations (CNVs) in the SRC gene. We can clearly distinguish whether copy number abnormalities are caused by non-allelic homologous recombination (NAHR) between the SRC gene and the pseudogene STRCP1, or between the upstream genes CKMT1B and CKMT1A. Furthermore, we can effectively differentiate between cis-trans relationships of variants. This method offers broad detection range, short cycle time, high specificity, low cost, and compatibility with multiple sequencing platforms. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments and examples of this application, and to more completely understand this application and its beneficial effects, the accompanying drawings used in the description of the embodiments or examples will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of this application. Those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0034] Figure 1 The above is the result of agarose gel electrophoresis detection of GAP-PCR amplification products in one embodiment of this application, where M: DL15000; lanes 1 and 3: samples to be verified; lanes 2 and 4: normal control samples; lanes 1 and 2 are amplification products of primer set LF3+ LR3, and lanes 3 and 4 are amplification products of primer set LF5+ LR3.

[0035] Figure 2 This is a Sanger sequencing result of the primer set LF3+LR3 amplification product in one embodiment of this application. The red arrow indicates the base difference between exon 10 of the upstream gene CKMT1B of the STRC gene and its homologous sequence.

[0036] Figure 3 This is a Sanger sequencing result of the primer set LF5+LR3 amplification product in one embodiment of this application. The red arrow indicates the base difference between exon 10 of the upstream gene CKMT1B of the STRC gene and its homologous sequence.

[0037] Figure 4 This is a Sanger sequencing result diagram of exon 28 of the STRC gene in one embodiment of this application, where the red arrows indicate the base differences between exon 28 of the STRC gene and its homologous sequences.

[0038] Figure 5 This is a Sanger sequencing result diagram of exon 26 of the STRC gene in one embodiment of this application, where the red arrows indicate the base differences between exon 26 of the STRC gene and its homologous sequences.

[0039] Figure 6 This is a Sanger sequencing result diagram of exon 26 of the STRC gene in one embodiment of this application, where the red arrows indicate the base differences between exon 26 of the STRC gene and its homologous sequences.

[0040] Figure 7 This is a Sanger sequencing result of exon 25 of the STRC gene in one embodiment of this application, where the red arrows indicate the base differences between exon 25 of the STRC gene and its homologous sequences.

[0041] Figure 8 This is an IGV map of the patient's STRC gene targeted third-generation sequencing results in one embodiment of this application. Detailed Implementation

[0042] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings. Preferred embodiments of this application are shown in the drawings. However, this application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of this application.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0044] In this application, "optionally," "optionally," and "optional" mean that something is optional, that is, it means that it is selected from either "with" or "without." If there are multiple "optional" entries in a technical solution, unless otherwise specified, and there are no contradictions or mutual constraints, each "optional" entry shall be independent.

[0045] In this application, terms such as "preferred," "better," "more suitable," and "ideal" are merely used to describe implementation methods or embodiments that achieve better results, and should be understood not to limit the scope of protection of this application.

[0046] The terms “having,” “containing,” “comprising,” and “including” as used in this application are synonyms and are inclusive or open-ended, not excluding additional, uncited members or features. Members or features include, for example, materials or components, structures, elements, instruments, etc.; non-limiting examples of members or features include actions, conditions under which actions occur, timing, states, etc.

[0047] In this application, the technical features or solutions described in open-ended language include both closed-ended technical features or solutions consisting of the listed contents and open-ended technical features or solutions that include the listed contents.

[0048] In this application, if the unit of a data range is only followed by the right endpoint, it means that the units of the left and right endpoints are the same.

[0049] In this application, where the method flow involves multiple steps, unless otherwise explicitly stated herein, there is no strict order restriction on the execution of these steps; they can be executed in any order other than those described. Moreover, any step may include multiple sub-steps or multiple stages, which are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or simultaneously with other steps or parts of the sub-steps or stages of other steps.

[0050] In this application, the exemplary descriptions such as "in some implementations (or embodiments)" and "in one implementation (or embodiment)" may cover, but are not limited to, the following meanings: these solutions can be combined with other solutions in a suitable manner to form new technical solutions.

[0051] In this application, the terms "first aspect," "second aspect," "first peptide," "second peptide," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or quantity, nor should they be construed as implicitly indicating the importance or quantity of the indicated technical features. Moreover, "first," "second," etc., serve only as a non-exhaustive enumeration and should be understood not to constitute a closed limitation on quantity.

[0052] In this application, when numerical intervals (i.e., numerical ranges) are mentioned, unless otherwise specified, the distribution of selectable numerical values ​​within the numerical interval is considered continuous, and includes the two endpoints of the numerical interval (i.e., the minimum and maximum values), as well as every numerical value between these two endpoints. Unless otherwise specified, when a numerical interval refers only to integers within that numerical interval, it includes the two endpoint integers of the numerical range, as well as every integer between the two endpoints, which is equivalent to directly listing every integer. When multiple numerical ranges are provided to describe features or characteristics, these numerical ranges can be merged. In other words, unless otherwise specified, the numerical ranges disclosed herein should be understood to include any and all subranges included therein. The "numerical value" in the numerical interval can be any quantitative value, such as a number, percentage, ratio, etc. The term "numerical interval" can be broadly included to include numerical interval types such as percentage intervals, ratio intervals, and proportion intervals.

[0053] In this application, the terms "room temperature" or "normal temperature" generally refer to 4°C to 35°C, for example, 20°C ± 5°C. In some embodiments of this application, "room temperature" or "normal temperature" refers to 10°C to 30°C. In some embodiments of this application, "room temperature" or "normal temperature" refers to 20°C to 30°C.

[0054] Currently, clinical STRC gene testing mainly employs high-throughput sequencing or multiplex ligation probe amplification (MLPA) technology, or long-chain PCR to fully amplify the true gene followed by nested PCR and conventional Sanger sequencing. While high-throughput sequencing can effectively detect SNVs and INDELs of individual genes, indicating whether gene copy number is abnormal, it often suffers from low specificity and insufficient accuracy for highly homologous genes, requiring validation using other methodologies, which additionally increases testing costs and experimental time. Furthermore, while CN116453588A discloses a method for detecting STRC gene copy number variations based on whole-genome sequencing, it can only determine the copy number of a few specific exons of the STRC gene, while the copy number of other exons remains unclear. Multiplex ligation probe amplification (MLPA) technology, although capable of simultaneously detecting the copy number of exons in both the STRC and pseudogene STRCP1, similarly only determines the copy number of a few specific exons of the STRC gene, leaving the copy number of other exons undetermined, and also cannot determine whether true and pseudogene fusion has occurred. Vona et al. (2015) provided a set of long-chain PCR and nested PCR primers for detecting the STRC gene, but the detection range is limited to SNVs and INDELs, and cannot detect the copy number of STRC gene exons. If non-allelic homologous recombination (NAHR) occurs between the STRC gene and the pseudogene STRCP1, false positive or false negative results are highly likely. Moreover, when two or more mutations exist simultaneously at the STRC gene locus, existing methods cannot distinguish between cis and trans mutations, or whether the mutations have a linkage effect. Therefore, there is an urgent need to develop an efficient, accurate, and rapid method for detecting STRC gene variants to overcome the problems of low specificity, inaccurate results, and the need for additional detection costs and experimental time in existing STRC gene detection methods.

[0055] With the continuous development of sequencing technology, third-generation sequencing (TGS) technology has begun to emerge, and this sequencing method is now being used for actual DNA sequencing.

[0056] In this application, "non-allelic homologous recombination" refers to recombination that occurs between two highly homologous DNA sequences located at different positions in the genome.

[0057] In this application, "specific sequence" refers to a DNA sequence that exists only in the target gene (such as the STRC gene) and is absent or differs from its homologous sequence (such as the pseudogene STRCP1), and can be used to design specific primers to achieve specific amplification of the target region.

[0058] In this application, "homologous sequence" refers to a DNA sequence that is highly similar to the target gene sequence but located at a different position in the genome, such as the STRC gene and the pseudogene STRCP1.

[0059] Based on this, the embodiments of this application at least provide primer sets, kits, methods and applications for detecting STRC gene variations.

[0060] In some embodiments, this application provides a method for STRC gene detection based on long-chain PCR amplification and third-generation sequencing. Its innovations compared to traditional detection methods are as follows:

[0061] 1. Wide detection range - By using long-chain PCR to target and enrich the full length of the SRC gene and several upstream and downstream regions and prepare third-generation sequencing libraries, it can not only achieve comprehensive, accurate and rapid detection of SNVs, INDELs, exon copy number of SRC gene, as well as whether non-allelic homologous recombination (NAHR) occurs between SRC gene and pseudogene STRCP1 and the size of recombination region, but also detect abnormal SRC gene copy number caused by non-allelic homologous recombination (NAHR) between upstream genes CKMT1B and CKMT1A.

[0062] 2. It can clearly distinguish between cis and trans mutations - when two or more mutations exist at the same time on the STRC gene locus, it can effectively distinguish between cis and trans mutations, and whether the mutation has a linkage effect;

[0063] 3. Short detection cycle and low cost; it can simultaneously detect all types of variations within the same system.

[0064] 4. High accuracy, effectively reducing the probability of false detections and missed detections;

[0065] 5. The detection and sequencing platform is flexible, capable of sequencing on Pacific Biosciences' PacBio platform, ONT's Nanopore platform, or other self-developed third-generation sequencing platforms.

[0066] In a first aspect of this application, a primer set for detecting SRC gene variants is provided, comprising one or more of the following primer sets:

[0067] Primer set 1: Upstream and downstream primers with sequences as shown in SEQ ID NO: 1 and SEQ ID NO: 2, respectively;

[0068] Primer set 2: Upstream and downstream primers with sequences as shown in SEQ ID NO: 1 and SEQ ID NO: 3, respectively;

[0069] Primer set 3: Upstream and downstream primers with sequences as shown in SEQ ID NO: 4 and SEQ ID NO: 5, respectively;

[0070] Primer set 4: Upstream and downstream primers with sequences as shown in SEQ ID NO: 6 and SEQ ID NO: 7, respectively;

[0071] Primer set 5: Upstream and downstream primers with sequences as shown in SEQ ID NO: 8 and SEQ ID NO: 9, respectively;

[0072] Primer set 6: Upstream and downstream primers with sequences as shown in SEQ ID NO: 10 and SEQ ID NO: 11, respectively;

[0073] Primer set 7: Upstream and downstream primers with sequences as shown in SEQ ID NO: 12 and SEQ ID NO: 9, respectively;

[0074] Primer set 8: Upstream and downstream primers with sequences as shown in SEQ ID NO: 13 and SEQ ID NO: 14, respectively.

[0075] It should be noted that, since true and false genes are highly homologous, in order to reduce the risk of missed or false detection, the primers provided in this application adopt a shingled design. In addition, the designed primer combination preferably has the following characteristics: if the upstream primer sequence is a specific sequence or contains a differential base at the 3' end that distinguishes homologous sequences, then the downstream primer sequence is a homologous sequence; if the upstream primer sequence is a homologous sequence, then the downstream primer sequence is a specific sequence or contains a differential base at the 3' end that distinguishes homologous sequences.

[0076] Understandably, the upstream primers in primer sets 1-7 are all located on the CKMT1B gene, although they are all based on 3 , Differential bases at the ends distinguish homologous sequences, but these different bases are also SNP sites. If a flat amplification pattern is used instead of a shingled pattern, encountering SNP sites will cause amplification failure, leading to inaccurate results for the entire amplicon assay. Therefore, using a shingled pattern can reduce the risk of amplification failure due to primer 3. , The risk of false positives or false negatives due to interference from terminal SNPs is high. Secondly, subsequent testing has revealed that sometimes non-allelic homologous recombination (homozygous) occurs between the CKMT1B and CKMT1A genes, but the STRC gene copy number is not abnormal. If a flat primer pair is used, it will inevitably lead to incorrect conclusions. In addition, one primer in the primer pair distinguishes homologous regions while the other does not. If non-allelic homologous recombination occurs within the amplicon range, it can be effectively amplified. This effectively reduces the number of targeted enrichment primer pairs and also effectively avoids the situation in Comparative Example 1 where Vona et al. used specific primers.

[0077] For example, the amplicon lengths generated by the primer sets provided above are all greater than 13kb.

[0078] For example, the primer set provided above can achieve full coverage of the target region with only one PCR amplification in the same reaction system.

[0079] In some embodiments, the STRC gene variants detected by the primer set include one or more of the following: single nucleotide variants, insertions or deletions, exon deletions / duplications, and copy number variations on the STRC gene.

[0080] In some implementations, copy number variations include copy number variations resulting from non-allelic homologous recombination between the STRC gene and the pseudogene STRCP1, and / or copy number variations resulting from non-allelic homologous recombination between the upstream gene CKMT1B and the CKMT1A gene of the STRC gene.

[0081] In a second aspect of this application, the use of the above-described primer set for detecting STRC gene variations in the preparation of products for detecting STRC gene variations is provided.

[0082] In a third aspect of this application, a kit for detecting STRC gene variants is provided, comprising the primer set for detecting STRC gene variants described above.

[0083] In some implementations, the kit also includes one or more of nucleic acid extraction reagents, DNA polymerase, PCR buffer, and dNTPs.

[0084] In some embodiments, the DNA polymerase includes one or more of Phanta Max enzyme, Phanta Flash enzyme, and KOD FX Neo enzyme. Further, the DNA polymerase is KOD FX Neo enzyme.

[0085] In a fourth aspect of this application, a method for detecting STRC gene variants is provided, comprising the following steps:

[0086] S100: Extract genomic DNA from the sample to be tested;

[0087] S200: Using genomic DNA as a template, perform PCR amplification using the primer set or the kit described above to prepare one round of amplification products;

[0088] S300: Mix equal volumes of amplification products from each round with the same barcode, perform magnetic bead purification or Pippin HT recovery, and construct a third-generation sequencing library; and

[0089] S400: Sequencing of third-generation sequencing libraries.

[0090] In some implementations, gradient annealing is used in targeted PCR amplification to improve the specificity and efficiency of targeted enrichment products and reduce the enrichment of non-specific products.

[0091] In some implementations, the PCR amplification procedure includes:

[0092] Pre-denaturation at 94℃ for 2 minutes;

[0093] Denaturation at 98℃ for 10 seconds, annealing and extension at 71℃ for 8 minutes, 2-4 cycles;

[0094] Denaturation at 98℃ for 10 seconds, annealing and extension at 70℃ for 8 minutes, 6-8 cycles;

[0095] Denaturation at 98℃ for 10 seconds, annealing and extension at 68℃ for 8 minutes, 19-21 cycles.

[0096] In some implementations, the working concentration of each primer in the PCR amplification system is 10 μM.

[0097] In some implementations, the sample to be tested includes dried blood spots, peripheral blood, semen, oral mucosal epithelial cells or cultured cell lines and saliva.

[0098] In some implementations, the genome of the sample to be tested can be obtained using commercially available nucleic acid extraction or purification reagents, or it can be obtained using conventional DNA extraction methods in the art.

[0099] For example, the detection method provided above can detect SNVs, INDELs, and CNVs of the STRC gene, and can also detect whether non-allelic homologous recombination (NAHR) exists between the STRC gene and the pseudogene STRCP1, as well as the size of the recombination region. Some examples are provided below.

[0100] The embodiments of this application will be described in detail below with reference to examples. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of this application. For experimental methods in the following embodiments where conditions are not specified, reference should be made to the guidelines given in this application, or to experimental manuals or conventional conditions in the art, or to the conditions recommended by the manufacturer, or to experimental methods known in the art.

[0101] In the following examples, the measurement parameters of the raw material components may have slight deviations within the weighing accuracy range unless otherwise specified. Temperature and time parameters are subject to acceptable deviations due to instrument testing accuracy or operational precision.

[0102] The materials and methods involved in the following embodiments mainly include:

[0103] 1. Primers

[0104] This application relates to a targeted enrichment primer set for STRC gene detection. This primer set uses the full-length STRC gene sequence (NC_000015.9:g.43891761_43910998) and several upstream and downstream sequences for specific targeted primer design. The specific targeted primers are shown in Table 1 below:

[0105] Table 1

[0106]

[0107] 2. This application relates to a targeted enrichment primer set for SRC gene detection. Since true and false genes are highly homologous, to reduce the risk of missed or false detections, the primers employ a shingled design. Preferably, in the designed primer combination, if the upstream primer sequence is a specific sequence or contains a differential base at its 3' end that distinguishes homologous sequences, then the downstream primer sequence is a homologous sequence; if the upstream primer sequence is a homologous sequence, then the downstream primer sequence is a specific sequence or contains a differential base at its 3' end that distinguishes homologous sequences.

[0108] 3. This application relates to a set of targeted enrichment primers for STRC gene detection, which can achieve full coverage of the target region with only one PCR amplification in the same reaction system.

[0109] 4. This application relates to a targeted enrichment primer set for STRC gene detection, wherein the PCR amplification DNA polymerase can be Phanta Max Super-fidelity DNA Polymerase, Phanta Flash Super-Fidelity DNA Polymerase, or KOD Fx neo. KOD Fx neo is preferred.

[0110] 5. This application relates to a targeted enrichment primer set for SRC gene detection. In order to improve the specificity and efficiency of the targeted enrichment products and reduce the enrichment of non-specific products, gradient annealing is used for targeted PCR amplification.

[0111] 6. This application relates to a set of targeted enrichment primers for STRC gene detection, with amplicon lengths all greater than 13kb.

[0112] 7. This application relates to a targeted enrichment primer set for STRC gene detection, wherein each primer in the primer set also contains a barcode sequence at its 5' end.

[0113] 8. STRC gene detection targeted enrichment products can be used to construct third-generation sequencing libraries through end repair, adapter ligation, enzyme digestion and purification.

[0114] 9. STRC gene detection targeted third-generation libraries can be sequenced using either Pacific Biosciences' PacBio platform, Oxford Nanopore Technologies' (ONT) Nanopore platform, or other self-developed third-generation sequencing platforms.

[0115] 10. STRC gene-targeted third-generation sequencing can detect SNVs, INDELs, and CNVs of the STRC gene, as well as the presence of non-allelic homologous recombination (NAHR) between the STRC gene and the pseudogene STRCP1, and the size of the recombination region.

[0116] 11. STRC gene-targeted third-generation sequencing can determine whether copy number abnormalities are caused by non-allelic homologous recombination (NAHR) when detecting the copy number of exons or the entire STRC gene.

[0117] 12. STRC gene-targeted third-generation sequencing can detect abnormal STRC gene copy number caused by non-allelic homologous recombination (NAHR) between the upstream genes CKMT1B and CKMT1A.

[0118] 13. STRC gene targeted third-generation sequencing detection can effectively distinguish between cis and trans mutations when two or more mutations exist in the STRC gene at the same time, and whether the mutation has a linkage effect.

[0119] Details of the preparation experiments in this application:

[0120] Sample requirements: dried blood spots, peripheral blood, semen, oral mucosal epithelial cells, cultured cell lines, and saliva, etc.

[0121] The specific experimental operation procedure for this application is as follows:

[0122] 1. STRC gene targeted enrichment

[0123] a) Perform STRC gene-targeted PCR amplification using gDNA as a template. Prepare the amplification reaction mix according to Table 2 below on an icebox:

[0124] Table 2

[0125]

[0126] b) Vortex oscillation mixing, instantaneous centrifugation.

[0127] c) Place the PCR tubes on the PCR instrument, select "105℃" for the "heat cap", tighten the heat cap, and set the reaction parameters according to Table 3 below:

[0128] Table 3

[0129]

[0130] 2. STRC gene targeted enrichment pooling and purification

[0131] Pooling was performed on the same sample (carrying the same barcode sequence) with different targeted enrichment products in the same amount. The mixed targeted enrichment products were vortexed and then briefly centrifuged. AMPure PB magnetic beads (Pacific Biosciences) were diluted to 35% (v / v) using Elution Buffer for later use.

[0132] 1) Add 3.1x 35% (v / v) AMPure PB magnetic beads to a centrifuge tube containing the mixed targeted enrichment product and incubate at room temperature for 20 min;

[0133] 2) Place the centrifuge tubes on the magnetic rack and let them stand until the magnetic beads are completely attracted. Discard the supernatant.

[0134] 3) Add 300 µL of freshly prepared 80% ethanol, let stand for 30 seconds, and discard the supernatant;

[0135] 4) Repeat the previous step;

[0136] 5) Centrifuge briefly, then place on a magnetic rack and use a 10μL pipette to remove any residual ethanol at the bottom (be careful not to pick up the magnetic beads).

[0137] 6) Keep the centrifuge tubes on the magnetic rack at room temperature to allow the residual ethanol to evaporate completely (be careful not to dry them excessively).

[0138] 7) Remove the centrifuge tube from the magnetic rack, add 17.5 μL of nuclease-free water, resuspend the magnetic beads, and let stand at room temperature for 5 min;

[0139] 8) Place the centrifuge tube back on the magnetic rack for 2 minutes until the liquid is clear, then transfer 15 μL of the elution product to a new centrifuge tube;

[0140] 9) Take 1 μL of the elution product and quantify it using the Qubit dsDNA HS Assay Kit.

[0141] 3. Library construction and sequencing

[0142] The purified SRC gene targeted enrichment product samples carrying different barcode sequences were mixed in equal amounts (total mass > 1000 ng, total volume ≤ 46 μL), vortexed, and then purified using SMRTbell. ®Using the prep kit 3.0 (Pacific Biosciences), end repair and A-tail addition, adapter ligation and purification were performed according to its operating procedures. Enzyme digestion and purification were then carried out to prepare the third-generation library. Sequencing was performed according to the PacBio Revio recommended operating procedures and reagents.

[0143] Example 1

[0144] A patient experienced hearing loss, primarily in the high frequencies, which was first noticed around elementary school age. In 2008, they began wearing hearing aids (bilateral), with an average hearing loss of approximately 45 dB. Whole-exome sequencing (WES) suggested a possible heterozygous deletion of exons 14 to 28 of the STRC gene, particularly exons 25 and 26, which may have been homozygous deletions. However, due to the presence of pseudogenes in the STRC gene, with a coding sequence homology as high as 99.6%, the presence of pseudogenes could not be definitively confirmed, requiring further verification. Therefore, the primer set discovered in this study (LF3+LR3, LF5+LR3) was used for GAP-PCR verification according to the STRC gene targeting enrichment system and PCR amplification procedure. The GAP-PCR amplification products were detected using 0.5% agarose gel electrophoresis, and the electrophoresis pattern is shown below. Figure 1 As shown.

[0145] Agarose gel electrophoresis revealed unique bands of consistent fragment size in both the test sample and the control sample, with no bands showing differences in fragment size. Therefore, Sanger sequencing was used to further validate the amplification products of the test sample. The specific results are as follows:

[0146] Sanger sequencing using LF5:

[0147] The exon 10 sequence of the upstream gene CKMT1B from the STRC gene differs from the CKMT1A gene sequence by only two bases. In CKMT1B, the difference is between base C and base T, while in CKMT1A, it is between base A and base T. The LF3+LR3 and LF5+LR3 Sanger sequencing results are as follows: Figure 2 and Figure 3 As shown.

[0148] According to the Sanger sequencing results, the amplicon was the target sequence.

[0149] Sanger sequencing was performed using the sequencing SF primers (ACCTTGCTGTTCTGGGCTCTCCTTT (SEQ ID NO 15)). The homology of some exons of the STRC gene and its pseudogene is shown in Table 4 below:

[0150] Table 4. Homology information of STRC gene exons 25-29

[0151]

[0152] Sanger sequencing results as follows Figures 4-7 As shown.

[0153] Based on the Sanger sequencing results, some sequences were found to be the target sequences of the upstream gene CKMT1B of the STRC gene, while others were pseudogene sequences of the STRC gene. Based on this, it is speculated that the STRC gene may have undergone non-allelic homologous recombination (NAHR) with its pseudogene.

[0154] Furthermore, the patient's mother experienced bilateral hearing loss around elementary school age, with an average hearing loss of approximately 55 dB. Therefore, using the primer set discovered in this study (LF+LR, LF+LR-1, LF1+LR1, LF2+LR2, LF3+LR3, LF4+LR4, LF5+LR3, LF6+LR6), and following the experimental procedure described in this application, third-generation STRC gene targeting was performed on both the patient and her mother. The test results are as follows:

[0155] 1) Based on the targeted enrichment results, it was found that:

[0156] The patient showed amplified enriched products with all primer sets, while the patient's mother showed targeted enriched products only with the LF+LR, LF+LR-1, and LF6+LR6 primer sets, and no targeted enriched products with other primer sets.

[0157] 2) Based on the results of third-generation sequencing:

[0158] a) One DNA strand of the patient has a deletion of exons 25-29 of the STRC gene, and this deletion is due to an abnormal copy number caused by non-allelic homologous recombination (NAHR) between the STRC gene and its pseudogene; the other DNA strand has a deletion of exons 5-29 of the STRC gene, and this deletion is due to an abnormal copy number of the entire CKMT1B gene and exons 5-29 of the STRC gene caused by non-allelic homologous recombination (NAHR) between the STRC gene and its upstream CKMT1B gene and the CKMT1A and STRCP1 genes.

[0159] b) The patient's mother had a homozygous deletion of exons 5-29 of the STRC gene. This deletion was caused by non-allelic homologous recombination (NAHR) between the STRC gene and its upstream CKMT1B gene and the CKMT1A and STRCP1 genes, resulting in abnormal copy numbers of the entire CKMT1B gene and exons 5-29 of the STRC gene. Consequently, no enriched products were found in LF1+LR1, LF2+LR2, LF3+LR3, LF4+LR4, and LF5+LR3.

[0160] The patient's STRC gene-targeted third-generation sequencing IGV results are as follows: Figure 8 As shown, the embodiments of this application can not only effectively detect STRC gene mutations, but also clarify the mutation type, cause, and extent.

[0161] Comparative Example 1:

[0162] Vona et al. (2015) used two pairs of long PCR primers for the enrichment and detection of the full-length STRC gene; Cheng et al. (2024) used four pairs of long PCR primers combined with third-generation sequencing for STRC gene detection. The specific primer information is shown in Table 5 below:

[0163] Table 5 Primer Information List

[0164]

[0165] The two primer sets used by Vona et al. (2015) could only specifically amplify the STRC gene, while the four primer sets used by Cheng et al. (2024) lacked the ability to distinguish between true and false genes, and could amplify both the STRC gene and its pseudogenes simultaneously. In Example 1, when the primer set provided by Vona et al. (2015) was used to test the patient's mother's sample, effective amplification was not achieved, resulting in detection failure. Although targeted enrichment products were present in the patient's sample, only one DNA strand could be amplified, leading to false negative results and insufficient accuracy. When the primer set provided by Cheng et al. (2024) was used to test the same patient's mother's sample, although amplification products were present, three of the primer pairs amplified pseudogene sequences, which could only indicate that the STRC gene copy number was abnormal, but could not determine the specific range of the abnormality or whether large-scale non-allelic homologous recombination (NAHR) had occurred, so the detection results were also inaccurate. Similar problems existed in the testing of patient samples. Furthermore, if the STRC gene undergoes non-allelic homologous recombination (NAHR) with its pseudogene but does not cause an abnormal copy number of the STRC gene, the primer set of Cheng et al. (2024) will produce false negative results.

[0166] Example 2

[0167] The patient, aged 30, experienced hearing impairment. Following a panel test for hereditary deafness, additional testing was conducted for the STRC gene, which is associated with hearing impairment. The primer set and detection system described in this application were used for targeted enrichment, third-generation library construction, and sequencing. The quality control data for the third-generation sequencing are shown in Table 6 below.

[0168] Table 6 Quality Control of STRC Gene-Targeted Third-Generation Sequencing Data

[0169]

[0170] Note: Sequences: Total number of sequences; Bases: Total base data; Min: Shortest sequence length; Max: Longest sequence length; Average: Average sequence length; N50: The length of the sequence that reaches 50% of the genome length when the sequences are added together from longest to shortest; (C+G)s %: GC content.

[0171] The test results showed that non-allelic homologous recombination (NAHR) occurred between the STRC gene and its pseudogene in the sample, resulting in the deletion of exons 16 to 18 of the STRC gene, with a copy number of 0.

[0172] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0173] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims, and the specification and drawings can be used to interpret the content of the claims.

Claims

1. A primer set for detecting STRC gene variations, characterized in that, The primer set comprises one or more of the following primer sets: Primer set 1: Upstream and downstream primers with sequences as shown in SEQ ID NO: 1 and SEQ ID NO: 2, respectively; Primer set 2: Upstream and downstream primers with sequences as shown in SEQ ID NO: 1 and SEQ ID NO: 3, respectively; Primer set 3: Upstream and downstream primers with sequences as shown in SEQ ID NO: 4 and SEQ ID NO: 5, respectively; Primer set 4: Upstream and downstream primers with sequences as shown in SEQ ID NO: 6 and SEQ ID NO: 7, respectively; Primer set 5: Upstream and downstream primers with sequences as shown in SEQ ID NO: 8 and SEQ ID NO: 9, respectively; Primer set 6: Upstream and downstream primers with sequences as shown in SEQ ID NO: 10 and SEQ ID NO: 11, respectively; Primer set 7: Upstream and downstream primers with sequences as shown in SEQ ID NO: 12 and SEQ ID NO: 9, respectively; Primer set 8: Upstream and downstream primers with sequences as shown in SEQ ID NO: 13 and SEQ ID NO: 14, respectively.

2. The primer set for detecting STRC gene variations as described in claim 1, characterized in that, The 5' end of each primer in the primer set also contains a barcode.

3. The primer set for detecting STRC gene variations as described in claim 1 or 2, characterized in that, The primer set detects one or more of the following STRC gene variations: single nucleotide variants, insertions or deletions, exon deletions / duplications, and copy number variations on the STRC gene.

4. The primer set for detecting STRC gene variations as described in claim 3, characterized in that, The copy number variations include copy number variations caused by non-allelic homologous recombination between the STRC gene and the pseudogene STRCP1, and / or copy number variations caused by non-allelic homologous recombination between the upstream gene CKMT1B and the gene CKMT1A of the STRC gene.

5. The use of the primer set for detecting STRC gene variations as described in any one of claims 1 to 4 in the preparation of products for detecting STRC gene variations.

6. A kit for detecting STRC gene variants, characterized in that, The kit includes the primer set for detecting STRC gene variations as described in any one of claims 1 to 4.

7. The kit for detecting STRC gene variations as described in claim 6, characterized in that, The kit also includes one or more of the following: nucleic acid extraction reagents, DNA polymerase, PCR buffer, and dNTPs.

8. The kit for detecting STRC gene variations as described in claim 7, characterized in that, The DNA polymerase includes one or more of Phanta Max enzyme, Phanta Flash enzyme, and KOD FX Neo enzyme.

9. A method for detecting STRC gene variations, characterized in that, Includes the following steps: Extract genomic DNA from the sample to be tested; Using the genomic DNA as a template, PCR amplification was performed using the primer set described in any one of claims 1 to 4 or the kit described in any one of claims 6 to 8 to prepare one round of amplification products; Equal amounts of amplification products from the same barcode round were mixed, purified with magnetic beads or recovered with Pippin HT, and used to construct a third-generation sequencing library. as well as The third-generation sequencing library was sequenced.

10. The method for detecting STRC gene variations as described in claim 9, characterized in that, The PCR amplification procedure includes: Pre-denaturation at 94℃ for 2 minutes; Denaturation at 98℃ for 10 seconds, annealing and extension at 71℃ for 8 minutes, 2-4 cycles; Denaturation at 98℃ for 10 seconds, annealing and extension at 70℃ for 8 minutes, 6-8 cycles; Denaturation at 98℃ for 10 seconds, annealing and extension at 68℃ for 8 minutes, 19-21 cycles.