Detection probe combination and sequencing analysis method for dynamic mutation STR sites

By designing a specific probe combination and combining it with third-generation sequencing technology, the problems of low throughput and high cost of existing STR locus detection technology have been solved, and efficient and accurate screening of multiple STR loci has been achieved, which is suitable for STR locus detection in various sample types.

CN116042610BActive Publication Date: 2025-09-16WUHAN HOPE GRP MEDICAL LAB CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310177083.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-09-16
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

Existing STR locus detection technology has low throughput, high cost, and long cycle, making it difficult to achieve efficient and low-cost screening of multiple samples. Conventional methods cannot accurately detect STR loci with high copy numbers, limiting its application in clinical screening.

Method used

A probe combination was designed to cover specific STR loci and their upstream and downstream regions. The probe length is 50-130 bp, the GC content is 40%-60%, there is no hairpin structure, and the average interval is 1000 bp. Combined with third-generation sequencing technology, the probe combination is hybridized with the DNA library and then subjected to third-generation sequencing analysis to achieve efficient screening of multiple STR loci.

Benefits of technology

It achieves efficient and low-cost screening for 58 STR loci-related diseases, improves detection efficiency and accuracy, overcomes the limitations of traditional methods, provides full coverage of STR loci screening, and is suitable for a variety of sample types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GDA0005535185440000091
    Figure GDA0005535185440000091
  • Figure GDA0005535185440000101
    Figure GDA0005535185440000101
  • Figure GDA0005535185440000111
    Figure GDA0005535185440000111
Patent Text Reader

Abstract

The present invention relates to the field of dynamic mutation disease diagnosis, and specifically to a detection probe combination and sequencing analysis method for dynamic mutation STR loci. The present invention discloses a probe combination and detection method for detecting pathogenic STR loci and potential pathogenic STR loci. The present invention utilizes probe capture amplification technology in conjunction with third-generation long-read sequencing technology, and specifically designs calculation parameters for the number of STR sequence repeats, thereby achieving full coverage screening of pathogenic STR loci. Test results show that the screening results are accurate and reliable, with high detection efficiency. Multiple STR pathogenic loci can be screened simultaneously, improving the detection efficiency of STR loci and the clinical accessibility of large-scale screening. This overcomes the shortcomings of traditional detection technologies, such as a small number of target sites, high cost, and incomplete analysis, and provides convenience for the study of STR loci for dynamic mutation diseases. The method has significant clinical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of dynamic mutation disease diagnosis, and in particular to a detection probe combination and sequencing analysis method for dynamic mutation STR sites. Background Art

[0002] The human genome contains over one million annotated tandem repeats, and tandem duplications and expansions at over 56 loci have been linked to disease. The clinical phenotypes of diseases caused by different STR loci are highly similar, and it is currently impossible to directly identify the specific pathogenic STR loci based on clinical manifestations. This requires individual STR locus verification, which poses a significant challenge to clinical disease research.

[0003] The insufficient throughput of existing detection technologies limits the effective detection of tandem repeat sites. Conventional Southern hybridization technology is time-consuming and labor-intensive, can only detect one site at a time, has low throughput, and the experimental materials are radioactive, which is detrimental to the health of the inspector. The commonly used capillary electrophoresis technology can also only detect one site at a time. Detecting just 56 known sites requires 56 capillary electrophoresis experiments. If 56 known sites are tested on multiple samples, the workload will be even greater and the cycle will be longer. In addition, some STR loci have extremely high pathogenic thresholds, such as SCA36, which will only cause disease when the number of repeat copies reaches 650 or more. Such a high number of repeat copies exceeds the upper limit of the accuracy range of capillary electrophoresis technology, so capillary electrophoresis technology cannot be used for accurate detection. The above defects have limited the promotion and application of capillary electrophoresis technology in the field of clinical screening.

[0004] While conventional second-generation NGS sequencing offers high throughput, its short sequences mean that once STR loci extend beyond 150bp, they exceed the NGS read length, resulting in the inability to accurately detect many STR loci. Third-generation whole-genome sequencing can accurately detect STR loci, but its extensive data volume, long analysis cycles, and high sequencing costs limit its widespread application in clinical screening.

[0005] STR loci can be accurately detected through third-generation sequencing of long-range PCR products, but this also faces the limitations of PCR itself. When multiplex PCR amplifies multiple loci, primers interfere with each other, reducing the accuracy of the analysis results. Single-site PCR has low throughput and can only detect one locus at a time. Testing just 56 known loci requires 56 long-range PCR experiments. The large workload and long cycle time limit its widespread application in clinical screening. Currently, the clinical screening field urgently needs a low-cost, high-throughput, and fast-cycle screening technology that fully covers pathogenic STR loci and potential pathogenic STR loci to accelerate clinical research on such diseases. Summary of the Invention

[0006] In view of this, the technical problem to be solved by the present invention is to provide a detection probe combination and sequencing analysis method for dynamically mutated STR sites.

[0007] The present invention provides a probe combination that meets the following requirements I) to III):

[0008] I), targeting the following animal-derived STR loci: SCA1, SCA2, SCA3, SCA6, SCA7, SCA8, SCA10, SCA12, SCA17, SCA31, SCA36, SCA37, SCA50, CANVAS, DRPLA, FRDA, GDPAG, FAME1, FAME2, FAME3, FAME4, FAME6, FAME7, EPM1A, OPDM1, OPDM2, OPDM3, OPDM4, DM1, DM2, SBMA, FTDALS1, FECD3HD, HDL1, HDL2, DIP2B, FXS, SPD1, CCHS, CBL, HFG, BPES, XLMR, EIF4A3, TBX1, ZIC3, NIPA1, POLG, THAP11, OPML1, HPE5, AFF2, DBQD2, RUNX2, ARX, HMN, DIP2B;

[0009] II) Design a region extending 3 kb upstream and downstream of the STR site shown in I); the average interval between the sites covered by each two adjacent probes is 500 bp to 2 kb;

[0010] III) The probe length is 50-130 bp, the GC content is 40%-60%, and there is no hairpin structure;

[0011] Further,

[0012] The STR loci are derived from humans;

[0013] The average interval between the sites covered by each two adjacent probes was 1000 bp;

[0014] The probe is 120 bp in length.

[0015] Furthermore, it includes at least one of the following sequences:

[0016] chr1:1437017~1437137;chr1:1438571~1438691;chr1:57829425~57829545;chr1:57830517~57830637;chr1:57831425~57831545;chr1:57832425~57832545;chr1:57833629~57833749;chr1:57834466~57834586;chr1:57835480~57835600;chr1:145206532~145206652;chr1:145207288~145207408;chr1:145208311~145208431;chr1:145209962~145210082;chr1:145210288~145210408;chr1:145211300~145211420;chr1:145212288~145212408;chr1:145213274~145213394;chr1:145214286~145214406;chr1:145215290~145215410;chr1:145216288~145216408;chr1:145217288~145217408;chr1:145218909~145219029;chr1:145219293~145219413;chr1:145220689~145220809;chr1:145221288~145221408;chr1:145222318~145222438;chr1:145223625~145223745;chr1:145224470~145224590;chr1:145225356~145225476;chr1:145226285~145226405;chr1:145227211~145227331;chr1:145228288~145228408;chr1:145229159~145229279;chr1:145230409~145230529;chr1:145231373~145231493;chr1:145232288~145232408;chr1:145233325~145233445;chr1:145234335~145234455;chr1:145235288~145235408;chr1:145236288~145236408;chr1:145237288~145237408;chr1:145238344~145238464;chr1:145239591~145239711;chr1:145240288~145240408;chr1:145241317~145241437;chr1:145242272~145242392;chr1:145243288~145243408;chr1:145244250~145244370;chr1:145245889~145246009;chr1:145246316~145246436;chr1:145247260~145247380;chr1:145248519~145248639;chr1:145249596~145249716;chr1:145250288~145250408;chr1:145251391~145251511;chr1:145252319~145252439;chr1:145253288~145253408;chr1:145254301~145254421;chr1:145255810~145255930;chr1:145256210~145256330;chr1:145257372~145257492;chr1:145258288~145258408;chr1:145259291~145259411;chr1:145260288~145260408;chr1:145261344~145261464;chr1:145262288~145262408;chr1:145263288~145263408;chr1:145264228~145264348;chr1:145265288~145265408;chr1:145266356~145266476;chr1:145267349~145267469;chr1:145268264~145268384;chr1:145269425~145269545;chr1:145270604~145270724;chr1:145271482~145271602;chr1:145272507~145272627;chr1:145273288~145273408;chr1:145274288~145274408;chr1:145275187~145275307;chr1:145277244~145277364;chr1:145278529~145278649;chr1:145279233~145279353;chr1:145280229~145280349;chr1:145281288~145281408;chr1:145282265~145282385;chr1:145283304~145283424;chr1:145284349~145284469;chr1:145285944~145286064;chr10:79823547~79823667;chr10:79825547~79825667;chr10:79826550~79826670;chr10:79827547~79827667;chr10:79828547~79828667;chr11:119073928~119074048;chr11:119074888~119075008;chr11:119075687~119075807;chr11:119076506~119076626;chr11:119077687~119077807;chr11:119078592~119078712;chr11:119079637~119079757;chr12:7042630~7042750;chr12:7043579~7043699;chr12:7044579~7044699;chr12:7045579~7045699;chr12:7046579~7046699;chr12:7047579~7047699;chr12:7048510~7048630;chr12:50895378~50895498;chr12:50896429~50896549;chr12:50897539~50897659;chr12:50898467~50898587;chr12:50899467~50899587;chr12:50900441~50900561;chr12:50901689~50901809;chr12:112034117~112034237;chr12:112034810~112034930;chr12:112035459~112035579;chr12:112036459~112036579;chr12:112037496~112037616;chr12:112038304~112038424;chr12:112039975~112040095;chr12:124015026~124015146;chr12:124015955~124016075;chr12:124016955~124017075;chr12:124017955~124018075;chr12:124018964~124019084;chr12:124020906~124021026;chr13:70710164~70710284;chr13:70711074~70711194;chr13:70712270~70712390;chr13:70713264~70713384;chr13:70714209~70714329;chr13:70715019~70715139;chr13:70716828~70716948;chr13:100634917~100635037;chr13:100635424~100635544;chr13:100636395~100636515;chr13:100637892~100638012;chr13:100638360~100638480;chr13:100639363~100639483;chr13:100640395~100640515;chr14:23787878~23787998;chr14:23788567~23788687;chr14:23789366~23789486;chr14:23790366~23790486;chr14:23791366~23791486;chr14:23792367~23792487;chr14:23793375~23793495;chr14:92534044~92534164;chr14:92535187~92535307;chr14:92536138~92536258;chr14:92537151~92537271;chr14:92538046~92538166;chr14:92539185~92539305;chr14:92540707~92540827;chr15:23083375~23083495;chr15:23084360~23084480;chr15:23085096~23085216;chr15:23086586~23086706;chr15:23087428~23087548;chr15:23088540~23088660;chr15:23089134~23089254;chr15:8987349 3~89873613;chr15:89874493~89874613;chr15:89875641~89875761;chr 15:89876938~89877058;chr15:89877458~89877578;chr15:89878493~89878613;chr15:89879496~89879616;chr16:17561442~17561562;chr16: 17562442~17562562;chr16:17563442~17563562;chr16:17564860~17564980;chr16:17565431~17565551;chr16:17566693~17566813;chr16:175 67442~17567562;chr16:24621476~24621596;chr16:24622522~24622642;chr16:24623476~24623596;chr16:24624507~24624627;chr16:246254 76~24625596;chr16:24626477~24626597;chr16:24627778~24627898;chr16:66518300~66518420;chr16:66519300~66519420;chr16:66520440~ 66520560;chr16:66521300~66521420;chr16:66522197~66522317;chr16:66523300~66523420;chr16:66524361~66524481;chr16:66525300~665 25420;chr16:66526641~66526761;chr16:66527302~66527422;chr16:87634582~87634702;chr16:87635536~87635656;chr16:87636641~876367 61;chr16:87637637~87637757;chr16:87638401~87638521;chr16:8763 9587~87639707;chr16:87640582~87640702;chr17:43972331~43972451;<h2 style=";text-align:left;direction:ltr">chr17:43973102~43973222;chr17:43973963~43974083;chr17:7811737 2~78117492;chr17:78118501~78118621;chr17:78119501~78119621;chr 17:78120564~78120684;chr17:78121812~78121932;chr17:78123544~7 8123664;chr18:53250101;53250221;chr18:53251093;53251213;chr18: 53252504~53252624;chr18:53253093~53253213;chr18:53254173~5325 4293;chr18:53254977;53255097;chr18:53256077~53256197;chr19:146 03539~14603659;chr19:14605171~14605291;chr19:14605540~14605660;chr19:14606416~14606536;chr19:14607540~14607660;chr19:146085 31~14608651;chr19:14610091~14610211;chr19:46270432~46270552;chr19:46271036~46271156;chr19:46272083~46272203;chr19:46273172~ 46273292;chr19:46273994~46274114;chr19:46275164~46275284;chr19:46276164~46276284;chr2:96859118~96859238;chr2:96860436~96860 556;chr2:96861118~96861238;chr2:96863669~96863789;chr2:968641 76~96864296;chr2:96865515;96865635;chr2:176954477;176954597;ch r2:176955477~176955597;chr2:176956466~176956586;chr2:17695798 1~176958101;chr2:176958477~176958597;chr2:176959477~176959597;<h2 style=";text-align:left;direction:ltr">chr2:176960343~176960463;chr2:191742293~191742413;chr2:191743 309~191743429;chr2:191744553~191744673;chr2:191745311~19174543 1;chr2:191746517;191746637;chr2:191747623~191747743;chr2:19174 8438~191748558;chr20:2630212~2630332;chr20:2631057~2631177;chr 20:2632336~2632456;chr20:2632913~2633033;chr20:2634073~2634193;chr20:2635073~2635193;chr20:2636073~2636193;chr20:4696828~46 96948;chr20:4697550~4697670;chr20:4698550~4698670;chr20:4700083~4700203;chr20:4700603~4700723;chr20:4701550~4701670;chr21:45 193012~45193132;chr21:45194012~45194132;chr21:45195159~451952 79;chr21:45196580;45196700;chr21:45198104~45198224;chr21:45198 831~45198951;chr22:19750934~19751054;chr22:19752022~19752142;chr22:19752871~19752991;chr22:19754659~19754779;chr22:19755065~ 19755185;chr22:19755978~19756098;chr22:19756944~19757064;chr22:46188154~46188274;chr22:46188895~46189015;chr22:46190055~461 90175;chr22:46190940~46191060;chr22:46193019~46193139;chr22:46193999~46194119;chr3:63895345~63895465;chr3:63895954~63896074;chr3:63897101~63897221;chr3:63898160~63898280;chr3:63899047~63899167;chr3:63900072~63900192;chr3:63901033~63901153;chr3:128888127~128888247;chr3:128889116~128889236;chr3:128890166~128890286;chr3:128891674~128891794;chr3:128892164~128892284;chr3:128892977~128893097;chr3:128894625~128894745;chr3:138661554~138661674;chr3:138662695~138662815;chr3:138663585~138663705;chr3:138665214~138665334;chr3:138665947~138666067;chr3:138666554~138666674;chr3:138667554~138667674;chr3:183426978~183427098;chr3:183428103~183428223;chr3:183428637~183428757;chr3:183429465~183429585;chr3:183430452~183430572;chr3:183432258~183432378;chr3:183432552~183432672;chr4:3073276~3073396;chr4:3074147~3074267;chr4:3075208~3075328;chr4:3076348~3076468;chr4:3077306~3077426;chr4:3078306~3078426;chr4:3079614~3079734;chr4:39346895~39347015;chr4:39348328~39348448;chr4:39348855~39348975;chr4:39349588~39349708;chr4:39350975~39351095;chr4:39351744~39351864;chr4:39352928~39353048;chr4:41744792~41744912;chr4:41745680~41745800;chr4:41746680~41746800;chr4:41747645~41747765;chr4:41748680~41 748800;chr4:41749680~41749800;chr4:41750680~41750800;chr4:1602 60394~160260514;chr4:160261891~160262011;chr4:160262428~160262 548;chr4:160263243~160263363;chr4:160264394~160264514;chr4:1602 65404~160265524;chr4:160266394~160266514;chr5:10353157~10353277;chr5:10354326~10354446;chr5:10355199~10355319;chr5:10356106~10356226;chr5:10357157~10357277;chr5:10358196~10358316;chr5:10359358~10359478;chr5:146254976~146255096;chr5:146255977~14625609 7;chr5:146256977~146257097;chr5:146257981~146258101;chr5:146258977~146259097;chr5:146259981~146260101;chr5:146260821~146260941;chr6:16324738~16324858;chr6:16325613~16325733;chr6:16326590~16326710;chr6:16328037~16328157;chr6:16328580~16328700;chr6:16 329580~16329700;chr6:16330486~16330606;chr6:45387183~45387303;chr6:45388188~45388308;chr6:45389183~45389303;chr6:45390183~45390303;chr6:45391575~45391695;chr6:45392183~45392303;chr6:45393170~45393290;chr6:170868101~170868221;chr6:170870383~170870503;<h2 style=";text-align:left;direction:ltr">chr6:170870720~170870840;chr6:170872059~170872179;chr6:1708727 20~170872840;chr6:170873621;170873741;chr7:27236173~27236293;ch r7:27237257~27237377;chr7:27238210~27238330;chr7:27239609~27239729;chr7:27240235~27240355;chr7:27241235~27241355;chr7:2724223 5;27242355;chr8:105597960~105598080;chr8:105598875~105598995;c hr8:105599845~105599965;chr8:105601005~105601125;chr8:105601964 ~105602084;chr8:105602938~105603058;chr8:105603879~105603999;c hr8:119375733~119375853;chr8:119376775~119376895;chr8:119377654 ~119377774;chr8:119379141~119379261;chr8:119379842~119379962;chr8:119380644~119380764;chr9:27546775~27546895;chr9:27547092~27 547212;chr9:27548219~27548339;chr9:27549241~27549361;chr9:27550378~27550498;chr9:27551218~27551338;chr9:27554809~27554929;chr 9:27555213~27555333;chr9:27556214~27556334;chr9:27557237~27557357;chr9:27558215~27558335;chr9:27559214~27559334;chr9:27560614 ~27560734;chr9:27561104~27561224;chr9:27562166~27562286;chr9:27563250~27563370;chr9:27564363~27564483;chr9:27565107~27565227;chr9:27566673~27566793;chr9:27567075~27567195;chr9:27568317~27568437;chr9:27569625~27569745;chr9:27570214~27570334; 71188~27571308;chr9:27572214~27572334;chr9:27573214~27573334;c hr9:27574174~27574294;chr9:27575145~27575265;chr9:71648882~716 49002;chr9:71650239~71650359;chr9:71650951~71651071;chr9:71651882~71652002;chr9:71652913~71653033;chr9:71654234~71654354;chr9:71655073~71655193;chrX:25028401~25028521;chrX:25029326~25029446;chrX:25030401~25030521;chrX:25031960~25032080;chrX:250324 01~25032521;chrX:25033401~25033521;chrX:25034304~25034424;chr X:66761880~66762000;chrX:66762880~66763000;chrX:66764001~66764 121;chrX:66764880~66765000;chrX:66765880~66766000;chrX:6676688 0~66767000;chrX:66767849~66767969;chrX:136645674~136645794;chr X:136646671~136646791;chrX:136647617~136647737;chrX:136648516 ~136648636;chrX:136649671~136649791;chrX:136650672~136650792;c hrX:136651671~136651791;chrX:139583174~139583294;chrX:13958417 4~139584294;chrX:139585174~139585294;chrX:139586509~139586629;chrX:139587168~139587288; chrX:139588131~139588251; chrX:139589174~139589294; chrX:146990323~146990443; chrX:1 46991262~146991382; chrX:146992245~146992365; chrX:146993084~146993204; chrX:146994262~146994382; chrX:1469954 46~146995566; chrX:146996880~146997000; chrX:147578884~147579004; chrX:147579829~147579949; chrX:147580829~147 580949; chrX:147581766~147581886; chrX:147582834~147582954; chrX:147583641~147583761; chrX:147584829~147584949. ;

[0017] The probe of the present invention includes the forward sequence of the sequence as described above, and also includes the reverse complementary sequence of the sequence as described above, both of which can be used as probe sequences to capture STR target sequences.

[0018] In the present invention, the probes are designed with appropriate intervals based on the human reference genome version Hg19, and the probe density and coverage uniformity of the capture target area are maximized based on the size of the DNA fragments during the sample DNA library construction process; in some specific embodiments of the present invention, the probes are designed with average intervals of 250bp, 500bp, 1K, 2K, 3K, and 4K. The results show that the probes designed with an average interval of 1K have both a relatively small number of probes and relatively good coverage uniformity; the coverage uniformity of the probes designed with an average interval of 1K is significantly better than that of the probes designed with an interval greater than 1K, and the coverage uniformity of the probes designed with an average interval of 1K is very close to that of the probes designed with an average interval of less than 1K. Therefore, the probes designed with an average interval of 1K have the best cost-effectiveness.

[0019] In the present invention, the average probe spacing is 1K, including probes designed with spacings of about 700 to 1400 bp. In a specific embodiment of the present invention, the probe spacing designed for a single STR locus can be about 700 bp or about 1400 bp, and the probe spacing can be any value between 700 bp and 1400 bp.

[0020] In the present invention, the probes have a base number between 50 and 130 bp, a GC content of 40% to 60%, and no hairpin structure. Furthermore, the probes are preferably 120 bp in length. Blast comparisons are performed with databases such as NCBI to remove nonspecific molecular probes and prioritize highly specific probes. Therefore, the probe combination of the present invention is the optimal probe combination obtained after screening for the best cost-effectiveness, highest coverage, and best specificity for detecting and screening related diseases caused by 58 STR loci.

[0021] Each probe in the probe combination of the present invention is connected to a label, and the label is selected from at least one of biotin, avidin, antibiotin protein, antibody or chemical coupling agent; the label facilitates signal amplification and luminescence detection after the probe specifically binds to the DNA fragment.

[0022] The present invention provides a method for sequencing STR loci, which comprises the following steps:

[0023] Step 1: The genomic DNA in the sample to be tested is sheared, end-repaired, adapter-added, and amplified to obtain a DNA library;

[0024] Step 2: After blocking, the probe combination of the present invention is hybridized with the DNA library to obtain a pre-hybridization library, and the pre-hybridization library is amplified to obtain a hybridization library;

[0025] Step 3: Perform third-generation sequencing on the hybridization library and analyze the data to obtain STR locus sequence information.

[0026] Further,

[0027] In step 1 of the detection method,

[0028] The size of the interrupted gene fragment is 4.5 to 5.5 kb;

[0029] The linker includes a linker A and a linker B, wherein the nucleotide sequence of linker A is shown in SEQ ID NO: 1, and the nucleotide sequence of linker B is shown in SEQ ID NO: 2;

[0030] The amplification primers include Primer-F and Primer-R. The nucleotide sequence of Primer-F is shown in SEQ ID NO: 3, and the nucleotide sequence of Primer-R is shown in SEQ ID NO: 4.

[0031] In step 2 of the detection method,

[0032] The blocking primers include IndexABlock and IndexB Block, the nucleotide sequence of IndexA Block is shown in SEQ ID NO: 5, and the nucleotide sequence of IndexB Block is shown in SEQ ID NO: 6;

[0033] The amplification primers include Primer-F and Primer-R, the nucleotide sequence of Primer-F is shown in SEQ ID NO: 3, and the nucleotide sequence of Primer-R is shown in SEQ ID NO: 4;

[0034] Amplification system includes: 0M~5M Betaine, 0.1μM~1.0μM Primer F, 0.1μM~0.5μM PrimerR, 0.5mM~4.0mM Mg 2+ , 8 vol% dNTP Mixture, polymerase, polymerase buffer, pre-hybridization library; the concentration of each dNTP in the dNTPMixture is 0.5 mM to 5 mM; the polymerase is selected from at least one of EXTaq DNA polymerase, Pfu DNA polymerase, rTaq DNA polymerase and / or PrimeSTAR GXL polymerase;

[0035] The amplification program included: 98°C for 1 min; 98°C for 15 sec, 58°C to 64°C for 15 sec, 68°C for 6 min, 18 cycles; 68°C for 5 min.

[0036] Preferably,

[0037] The amplification system is: 1M Betaine, 0.5μM PrimerF, 0.5μM PrimerR, 2mM Mg 2+ , 8 vol% dNTPMixture, 2 vol% PrimeSTAR GXL DNA Polymerase, 20 vol% 5× PrimeSTAR GXL Buffer, 50 vol% prehybridization library; the concentration of each dNTP in dNTPMixture was 2.5 mM;

[0038] The amplification program included: 98°C for 1 min; 98°C for 15 sec, 62°C for 15 sec, 68°C for 6 min, 18 cycles; 68°C for 5 min.

[0039] The present invention adjusts parameters related to the amplification system for hybridization library construction; in some embodiments of the present invention, the primer concentration is 0.1 μM to 1.0 μM; the Betaine concentration is 0 M to 5 M; the DNA polymerase is selected from EXTaq DNA polymerase, Pfu DNA polymerase, rTaq DNA polymerase, and PrimeSTAR GXL polymerase; the magnesium ion concentration is 0.5 mM to 4.0 mM; and the annealing temperature is 58° C. to 64° C. Experimental results show that when the primer concentration is 0.5 μM, the Betaine concentration is 1 M, the DNA polymerase is PrimeSTAR GXL polymerase, the magnesium ion concentration is 2 mM, and the annealing temperature is 62° C., the amplification effect of the hybridization library is best.

[0040] In step 3 of the detection method,

[0041] The method for analyzing the data is: after extracting the STR region sequence data, a Gaussian mixture model with a smaller AIC value calculated when the parameter n_component is 1 or 2 is selected to determine the number of STR repetitions.

[0042] The criteria for determining the number of STR repeats are:

[0043] If the n_component of the selected model is 1, the peak average value of the model is directly returned as the number of repetitions of the sample;

[0044] If the n_component of the selected model is 2, the number of repetitions near the peak average of the model and with a read number greater than or equal to the minimum read number threshold of 2 is selected as the number of repetitions of the STR in the sample.

[0045] In the present invention, the data analysis software used was GrandSTR. STR locus tandem repeat copy number detection revealed that the results of the bioinformatics software of the present invention were 100% consistent with the theoretical repeat value. However, the results of the Straglr, repeatHMM, tandem-genotypes, and PACMONSTR software all showed discrepancies with the theoretical repeat value in multiple samples. This indicates that the results of the bioinformatics software of the present invention are superior to those of the Straglr, repeatHMM, tandem-genotypes, and PACMONSTR software.

[0046] In the present invention, the third-generation sequencing is performed using at least one of the platforms including PacBio Sequel, PromethION, MinION, and GridION;

[0047] Furthermore, the present invention can be performed using both the PacBio Sequel platform and the PromethION.

[0048] In the present invention, the sample to be tested includes but is not limited to at least one of blood, blood spots, semen, semen spots, bones, hair, saliva, saliva spots, sweat, amniotic fluid or a mixture thereof.

[0049] The present invention provides the use of at least one of the following in preparing a screening kit for STR loci-related diseases:

[0050] 1) The probe combination of the present invention;

[0051] II), the detection method of the present invention.

[0052] In the present invention, the STR loci include but are not limited to at least one of SCA1, SCA2, SCA3, SCA6, SCA7, SCA8, SCA10, SCA12, SCA17, SCA31, SCA3, SCA37, SCA50, CANVAS, DRPLA, FRDA, GDPAG, FAME1, FAME2, FAME3, FAME4, FAME6, FAME7, EPM1A, OPDM1, OPDM2, OPDM3, OPDM4, DM1, DM2, SBMA, FTDALS1, FECD3HD, HDL1, HDL2, DIP2B, FXS, SPD1, CCHS, CBLHFG, BPES, XLMR, EIF4A3, TBX1, ZIC3, NIPA1, POLG, THAP11, OPML1, HPE5, AFF2, DBQD2, RUNX2, ARX, HMN and / or DIP2B.

[0053] Diseases caused by STR site mutations include but are not limited to spinocerebellar ataxia type 1, spinocerebellar ataxia type 2, spinocerebellar ataxia type 3, spinocerebellar ataxia type 6, spinocerebellar ataxia type 7, spinocerebellar ataxia type 8, spinocerebellar ataxia type 10, spinocerebellar ataxia type 12, spinocerebellar ataxia type 17, spinocerebellar ataxia type 31, spinocerebellar ataxia type 36, spinocerebellar ataxia type 37, polyglutamine neurological disease, cerebellar ataxia - Neuropathy and vestibular dysreflexia syndrome, dentatorubral nucleus and globus pallidus Lewy body atrophy, Friedreich's ataxia type 1, global developmental delay - progressive ataxia and elevated glutamine, familial adult myoclonic epilepsy type 1, familial adult myoclonic epilepsy type 2, familial adult myoclonic epilepsy type 3, FAME4 familial myoclonic epilepsy type 4, familial adult myoclonic epilepsy type 6, familial adult myoclonic epilepsy type 7, epilepsy - progressive myoclonus 1A, distal oculopharyngeal myopathy, neuro Intranuclear inclusion disease, myotonic dystrophy type 2, myotonic dystrophy type 1, oculopharyngeal muscular dystrophy, spinal bulbar muscular atrophy, frontotemporal intellectual disability, amyotrophic lateral sclerosis type 1, Fuchs corneal endothelial dystrophy type 3, OPML1 oculopharyngeal disease with leukoencephalopathy type 1, Huntington's disease, Huntington's disease-like type 2, intellectual disability, FRA12A type, X-linked intellectual disability type 109 / fragile X syndrome, Desbuquois dysplasia type 2, syndactyly type 1, congenital At least one of central hypoventilation syndrome type 1, Jacobsen syndrome, brachydactyly and clavicular dysplasia BCCD, hand, foot and genital syndrome HFGS, blepharophimosis congenitalis syndrome BPES, holoprosencephaly type 5 HPE5, early infantile epileptic encephalopathy type 1 EIEE1, isolated growth hormone deficiency with mental retardation MRGH, HDL1, HMN, RCPS, TOF, VACTERLX, ALS–susceptibility, and / or POLG.

[0054] The present invention provides a product for STR locus detection, which includes one or more of the probe combination of the present invention, a solid phase support, an adapter sequence, an adapter blocking sequence, primers for binding to the adapter sequence and amplifying nucleic acid fragments, a DNA extraction system, a PCR reaction buffer, nuclease-free water, a DNA polymerase, a molecular weight marker, a target sequence eluent, an end-repair enzyme, an end-repair buffer, and a DNA ligase.

[0055] In some specific embodiments of the present invention, the linker sequence includes linker A and linker B, the sequence of linker A is shown in SEQ ID NO: 1; the sequence of linker B is shown in SEQ ID NO: 2; the linker blocking sequence includes IndexA Block and IndexB Block, the sequence of IndexABlock is shown in SEQ ID NO: 5, and the sequence of IndexABlock is shown in SEQ ID NO: 6.

[0056] In some specific embodiments of the present invention, the primers include Primer-F and Primer-R, the sequence of Primer-F is shown in SEQ ID NO: 3, and the sequence of Primer-R is shown in SEQ ID NO: 4.

[0057] The solid phase carrier of the present invention is enriched particles, and the enriched particles are coated with biotin or avidin;

[0058] Furthermore, the enrichment particles are magnetic beads, and the capture occurs on a solid phase carrier.

[0059] The present invention discloses a probe combination and detection method for detecting pathogenic STR loci and potential pathogenic STR loci. The present invention utilizes probe capture amplification technology in combination with third-generation long-read sequencing technology, and specifically designs calculation parameters for the number of STR sequence repeats, thereby achieving full coverage screening of pathogenic STR loci. Test results show that the screening results are accurate and reliable, the detection efficiency is high, and multiple STR pathogenic loci can be screened simultaneously, thereby improving the detection efficiency of STR loci and the clinical accessibility of large-scale screening. The present invention overcomes the shortcomings of traditional detection technologies such as a small number of target sites, high cost, and incomplete analysis, provides convenience for the study of STR loci for dynamic mutation diseases, and has significant clinical promotion and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 Shows probe density and coverage uniformity of capture target area;

[0061] Figure 2 The data calculation process is shown. DETAILED DESCRIPTION

[0062] The present invention provides a detection probe combination and sequencing analysis method for dynamically mutated STR loci. Those skilled in the art can draw upon the disclosure herein and appropriately modify the process parameters to implement the method. It is particularly important to note that all similar substitutions and modifications readily apparent to those skilled in the art are considered encompassed by the present invention. The methods and applications of the present invention have been described through preferred embodiments. It is readily apparent that those skilled in the art can modify, adapt, and combine the methods and applications herein to implement and apply the technology of the present invention without departing from the content, spirit, and scope of the present invention.

[0063] Dynamic mutations, also known as unstable oligonucleotide repeats (STRs), are commonly found in trinucleotide repeats such as GCA, GCT, and GCC. Some sites may contain TTTTA pentanucleotide repeats or other oligonucleotide repeats. These mutations are called dynamic mutations because the number of tandem repeats increases with each generation. When the number of tandem repeats reaches a certain threshold, disease can occur, often with premature onset, resulting in more severe symptoms in offspring. Dynamic mutations are a new type of genetic mutation that can cause human genetic diseases.

[0064] This invention combines a carefully designed long-fragment capture amplification technology for STR loci with third-generation long-read sequencing technology to achieve low-cost, high-throughput, and short-cycle screening for comprehensive coverage of pathogenic and potentially pathogenic STR loci. It can simultaneously detect 56 known pathogenic STR loci and potential pathogenic STR loci. Barcodes can be added to simultaneously test all 56 known loci on multiple individuals, making it fast and efficient, suitable for large-scale clinical sample screening. The 56 known pathogenic STR loci and two potential pathogenic STR loci are HMN and DIP2B, as shown in Table 1.

[0065] Table 1. 58 pathogenic STR loci

[0066] SCA1 SCA2 SCA3 SCA6 SCA7 SCA8 SCA10 SCA12 SCA17 SCA31 SCA36 SCA37 SCA50 CANVAS DRPLA FRDA GDPAG FAME1 FAME2 FAME3 FAME4 FAME6 FAME7 EPM1A OPDM1 OPDM2 OPDM3 OPDM4 DM1 DM2 SBMA FTDALS1 FECD3 HD HDL1 HDL2 DIP2B FXS SPD1 CCHS CBL HFG BPES XLMR eIF4A3 TBX1 ZIC3 NIPA1 POLG THAP11 OPML1 HPE5 AFF2 DBQD2 RUNX2 ARX HMN DIP2B

[0067] The locations of the known pathogenic STR loci described in Table 1 and the diseases they cause are shown in Table 2:

[0068] Table 2. Information on 58 known pathogenic STR loci

[0069]

[0070]

[0071] This method designs capture probes to cover 56 known loci and 2 potential pathogenic STR loci, and further uses third-generation long-read sequencing technology to achieve rapid detection of STR loci. The specific operation is as follows

[0072] A STR locus detection kit and method based on third-generation targeted capture sequencing, comprising the following steps:

[0073] 1) Construction of a genomic DNA library: Genomic DNA is extracted from the sample to be tested and fragmented into fragments of 2 to 10 kb, preferably about 5 kb. The ends of the fragmented double-stranded DNA fragments are repaired and then ligated with adapter sequences to obtain a DNA library.

[0074] 2) Pre-amplification of the DNA library: using primers targeting the adapter sequence, the DNA fragments with adapters in the above DNA library are subjected to a first PCR amplification to obtain a pre-amplified DNA library;

[0075] 3) Hybridization capture of target sequences: After denaturing the double-stranded DNA fragments in the pre-amplified DNA library, the target sequences are hybridized and captured using detection probes;

[0076] 4) Target sequence elution and second PCR amplification: The target sequence captured in the previous step is eluted and the eluted target sequence is subjected to a second PCR amplification;

[0077] 5) constructing a third-generation sequencing library using the product of the second PCR amplification;

[0078] 6) Sequencing the third-generation library using a third-generation sequencer;

[0079] 7) Perform bioinformatics analysis on the third-generation sequencing data to obtain the nucleic acid molecule sequence associated with the STR locus, and further determine the specific repeat copy number of the STR locus in the sample to be tested.

[0080] The test materials used in the present invention are all common commercial products and can be purchased in the market.

[0081] The present invention will be further described below in conjunction with the embodiments:

[0082] Example 1 STR site detection by third-generation targeted capture sequencing

[0083] 1. Preparation of STR site hybridization capture probes

[0084] 1.1 STR site hybridization capture probe design

[0085] For the STR site, 3Kb is extended upstream and downstream respectively, and a probe is designed every 500bp to 2K, preferably a probe is designed every 1K. (The sequencing read length of the third-generation sequencer is long. For example, the sequencing length of Pacbio's sequel or Oxford Nanopore's PromethION can reach tens of K. The genomic DNA in the sample to be tested is broken into fragments of about 5K in size. A probe is designed every 1K. On average, each DNA fragment can be bound to 5 probes, which is sufficient to effectively capture the target fragment). The number of bases of the designed probe is between 50 and 130bp, the GC content value is moderate, and there is no hairpin structure. It is further Blast-matched with databases such as NCBI to ensure the specificity of the probe. Preferably, the number of bases of the designed probe is 120bp. The STR site reference sequence is the human reference genome version Hg19, and the designed probe sequence is the "STR site probe". The probe sequence is shown in Table 3:

[0086]

[0087]

[0088]

[0089]

[0090]

[0091] 1.2 Probe density and capture target area coverage uniformity test

[0092] The percentage of sequences with an average coverage depth of 0.5× is a common parameter for evaluating panel coverage uniformity and is used to examine the limitations of low coverage depth. If more than 90% of the target regions show an average coverage depth of more than 0.5×, the panel has good coverage uniformity. To this end, we evaluated the coverage uniformity of the target regions with probes designed with an average interval of 250bp, 500bp, 1K, 2K, 3K, and 4K. Figure 1 As shown, the coverage uniformity of probes designed with an average spacing of 1K is significantly better than that of probes designed with spacing greater than 1K. The coverage uniformity of probes designed with an average spacing of 1K is very close to that of probes designed with spacing less than 1K. The probes designed with an average spacing of 1K combine relatively few probes with relatively good coverage uniformity. Therefore, the probes designed with an average spacing of 1K have the best price-performance ratio.

[0093] 2. Linker sequence

[0094] The linker is composed of sequence A and sequence B.

[0095] A sequence is:

[0096] 5'~pGATCGGAAGAGCACACGTCTGAACTNNNNNNNNNNACCCACGTCCGCCATA CTTG~3' (SEQ ID: AdaptorpartA; SEQ ID NO: 1);

[0097] The B sequence is:

[0098] 5'~CTTGGAGAACACCCCAAAGGAGATNNNNNNNNNNNNNNNNAGTTCAGACGTGTGCTCTTCCGATCT / 3SpC3 / -3'(SEQ ID: Adaptorpart B; SEQ ID NO: 2);

[0099] In the above sequence, p represents phosphorylation modification; N represents any one of the four bases A / G / C / T; / 3SpC3 / represents phosphorothioate bond modification; the A sequence NNNNNNNNNN and the B sequence NNNNNNNNNNNNNNNN are used to identify libraries constructed from different samples.

[0100] Furthermore, the specific operation of preparing a working linker from the above linker sequence A and sequence B is as follows:

[0101] 1) Dissolve the primer powder;

[0102] 2) Mix adapter A and adapter B in equal volumes in a PCR tube;

[0103] 3) Use the PCR program: 95°C for 3 minutes; slowly cool to 25°C over 1 hour, decreasing the temperature at 0.05°C / s;

[0104] A working joint is obtained.

[0105] 3. Sample genomic DNA extraction and quality assessment

[0106] Genomic DNA was extracted from 0.5–1 mL of peripheral blood using the Blood Genomic DNA Mini Kit (Beijing Kangwei Century Biotechnology Co., Ltd., peripheral blood). Genomic integrity was assessed by agarose gel electrophoresis, and genomic concentration was determined using Qubit (Life Technologies, USA). DNA samples with relatively intact genomes and a concentration of ≥25 ng / μl were selected for library construction.

[0107] 4. Construction of Capture Library

[0108] 4.1 Genomic DNA Shearing

[0109] Genomic DNA was fragmented using a Covaris E220 non-contact ultrasonic disruptor with the following parameters: Duration 600 sec, Peak Incident Power 3 watts, Duty Factor 20%, and Cycles per Burst 1000.

[0110] Fragmented genomic DNA was obtained by purification using 0.6 volumes of AMPure PB magnetic beads.

[0111] 4.2 Sample end repair and connector addition

[0112] 4.2.1 End repair, adding "A tail"

[0113] The NEB Next Ultra End Repair / dA-Tailing Module kit was used.

[0114] Table 4 A-tailed system

[0115]

[0116]

[0117] Table 5 Reaction conditions of the A-tailing system

[0118]

[0119] 4.2.2 Connector connection

[0120] The NEB Next Ultra II Ligation Module kit was used.

[0121] Table 6 Connector connection system

[0122] Reagent components Volume (μl) NEBNextUltraIILigtionMasterMix 30 NEBNextLigationEnhancer 1 Adapter (working connector) 2.5 End repair products 60 Total volume 93.5

[0123] The reaction was carried out in a PCR instrument at 20°C for 15 min. Subsequently, the library solution was purified using 0.6 volumes of AMPure XP magnetic beads.

[0124] 4.3 Pre-amplification of the library

[0125] Table 7 Library pre-amplification system

[0126] Reagent components Volume (μl) Library solution 25 5× PrimeSTARGXL Buffer 10 dNTPMixture 4 Primer 2 PrimeSTARGXLDNA Polymerase 1 Nuclease-free Water 8 Total volume 50

[0127] Among them, Primer-F and Primer-R are primers for library pre-amplification, specifically:

[0128] Primer-F: 5'-CAAGTATGGCGGACGTGGGT-3', SEQ ID NO: 3;

[0129] Primer-R: 5'-CTTGGAGAACACCCCAAAGGA-3', SEQ ID NO: 4.

[0130] Table 8 Pre-amplification reaction conditions of the library

[0131] step temperature Number of cycles time Genome degeneration 98℃ 1 1min transsexual 98℃ 15sec extend 68℃ 7 10min extend 68℃ 1 10min save 10℃ 1 ......

[0132] PCR products were recovered using 0.6 volumes of AMPure XP magnetic beads.

[0133] 4.4 Hybridization capture target sequence

[0134] 4.4.1 Capture Preparation

[0135] Equal amounts of samples to be captured were mixed, with a total DNA volume of 1 μg.

[0136] Add 4 μL of Index blocking reagent (IndexABlock, Index B Block) to block the adapter A sequence and B sequence used in library construction, respectively, to prevent the two adapter sequences on the library from hybridizing with the probe.

[0137] The IndexABlock sequence is: 5'-CAAGTATGGCGGACGTGGGTNNNNNNNNNNAGTTCAGACGTGTGCTCTTCCGATC-3', SEQ ID NO: 5;

[0138] The Index B Block sequence is: 5'-AGATCGGAAGAGCACACGTCTGAACTNNNNNNNNNNNNNNNNATCTCCTTTGGGGTGTTCTCCAAG-3', SEQ ID NO: 6;

[0139] In the above sequence, N represents any one of the four bases A / G / C / T, wherein IndexABlock is reverse complementary to the linker A sequence; IndexB Block is reverse complementary to the linker B sequence.

[0140] Mix by shaking, then concentrate under vacuum at 60℃ with the lid open to dryness;

[0141] 4.4.2 Capture

[0142] Prepare hybridization buffer according to the xGen Lockdown Reagent Kit (IDT, Cat. No. 1072281).

[0143] Table 9 Hybridization buffer system

[0144] Reagent components Volume (μl) 2xHybridizationBuffer 8.5 HybridizationBufferEnhancer 2.7 Nuclease~freeWater 1.8 Total volume 13

[0145] Add the hybridization buffer in Table 9 to the evaporated sample, shake to mix, centrifuge briefly, and incubate on a PCR instrument at 95°C for 10 min.

[0146] After the reaction is complete, centrifuge and transfer all 13 μL of the sample to a PCR tube containing 4 μL of capture probe. Mix thoroughly by pipetting and incubate at 65°C in a PCR machine for 16 hours.

[0147] 4.5 Hybridization and Elution

[0148] 4.5.1 Prepare elution reagent according to the xGen Lockdown Reagent Kit (IDT, Cat. No. 1072281):

[0149] Table 10 Elution Reagents

[0150] Reagent components Buffer volume (μl) Water volume (μl) 2×BeadWashBuffer(vial7) 250 250 10×WashBufferI(vial1) 30 270 10×WashBufferII(vial2) 20 180 10×WashBufferIII(vial3) 20 180 10×StringentWashBuffer(vial4) 40 360

[0151] 4.5.2 Capture Magnetic Bead Processing:

[0152] a) Pipette 50 μl of M280 magnetic beads into a centrifuge tube. Place the tube on a magnetic rack and wait for the supernatant to completely clear. Discard the supernatant.

[0153] b) Add 200 μl 1× BeadWash Buffer (vial 7), mix well, centrifuge, place on a magnetic rack and wait for the supernatant to completely clarify, then discard the supernatant.

[0154] c) Repeat step b once.

[0155] d) Add 100 μl of 1× BeadWash Buffer (vial 7), mix well, transfer to a PCR tube, place on a magnetic rack until the supernatant is completely clear, and then discard the supernatant.

[0156] 4.5.3 Binding of M280 magnetic beads and target library:

[0157] a) Add 17 μl of hybridization system to the PCR tube containing magnetic beads, pipette up and down slowly to mix, and incubate at 65°C for 45 minutes in a PCR instrument.

[0158] 4.5.4 Elution:

[0159] a) Add 100 μl of 65°C 1× Wash Buffer I (vial 1) and mix thoroughly. Place on a magnetic rack and wait until the supernatant is completely clear, then discard.

[0160] b) Add 200 μl of 65°C 1× Stringent Wash Buffer (vial 4), mix well, and transfer to a preheated large centrifuge tube. Incubate at 65°C for 5 minutes. Place on a magnetic stand and allow the supernatant to completely clear, then discard.

[0161] c) Add 200 μl of room temperature 1× Wash Buffer I (vial 1) and mix thoroughly. Place on a magnetic rack and wait until the supernatant is completely clear, then discard.

[0162] d) Add 200 μl of room temperature 1× Wash Buffer II (vial 2) and mix thoroughly. Place on a magnetic rack and wait until the supernatant is completely clear, then discard.

[0163] e) Add 200 μl of room temperature 1× Wash Buffer III (vial 3) and mix thoroughly. Place on a magnetic rack and allow the supernatant to completely clear, then discard.

[0164] f) Briefly centrifuge and remove any remaining liquid using a 10 μl pipette tip.

[0165] g) Use 25 μl of Nuclease-free Water to suspend the magnetic beads to obtain an M280 magnetic bead probe suspension, which is stored at -20°C.

[0166] 4.6 Optimization of hybridization library amplification reaction system and reaction conditions

[0167] 4.6.1 Reaction System and Reaction Conditions The basic system and conditions for optimization are as follows:

[0168] Take 25 μL of the M280 magnetic bead probe suspension obtained in the previous step per library and perform the following steps in a 200 ml PCR tube.

[0169] Table 11 Hybridization library amplification reaction system

[0170] Reagent components Volume (μl) Nuclease~freeWater 3 5× PrimeSTARGXL Buffer 10 dNTPMixture 4 Betaine 5 Primer 2 PrimeSTARGXLDNA Polymerase 1 M280 magnetic bead probe suspension 25 Total volume 50

[0171] The primers described in the table above are consistent with the primers used for library pre-amplification above:

[0172] Primer F: 5'~CAAGTATGGCGGACGTGGGT~3' (SEQ ID NO: 3),

[0173] Primer~R: 5'~CTTGGAGAACACCCCAAAGGA~3' (SEQ ID NO: 4).

[0174] Table 12 Hybridization library amplification reaction program

[0175]

[0176] (1) Optimization of primer concentration: When other reaction conditions remain the same, the concentrations of Primer-F and Primer-R were adjusted to 0.1 μM, 0.2 μM, 0.4 μM, 0.5 μM, 0.6 μM, 0.8 μM, and 1.0 μM, respectively. The reaction conditions were the same as those for the basic reaction. The results showed that the best amplification effect was achieved when the primer concentration was 0.5 μM, confirming that the optimal primer concentration in the reaction system was 0.5 μM.

[0177] (2) Optimization of Betaine Concentration: While maintaining the same reaction conditions, the concentration of Betaine was adjusted to 0 M, 0.1 M, 0.3 M, 1 M, 3 M, and 5 M. The reaction conditions were the same as those for the basic reaction. The results showed that a 1 M Betaine concentration produced the best amplification effect, confirming that the optimal Betaine concentration in the reaction system was 1 M.

[0178] (3) Enzyme Optimization: While other reaction system conditions remained the same, different DNA polymerases were used: EX Taq DNA polymerase, Pfu DNA polymerase, rTaq DNA polymerase, and PrimeSTAR GXL polymerase. Each reaction system contained a different DNA polymerase, and the concentrations of each DNA polymerase were the same. The reaction conditions were the same as those for the basic reaction. The results showed that the PrimeSTAR GXL polymerase reaction system had the best amplification effect, confirming that PrimeSTAR GXL polymerase was the optimal DNA polymerase.

[0179] (4) Optimization of magnesium ion concentration: While other reaction conditions remained the same, the magnesium ion concentration in the reaction system was adjusted to 0.5 mM, 1.0 mM, 2.0 mM, 3.0 mM, and 4.0 mM, respectively. The reaction conditions were the same as those in the basic reaction. The results showed that the amplification effect was best when the magnesium ion concentration was 2 mM, confirming that the optimal magnesium ion concentration in the reaction system was 2 mM.

[0180] (5) Optimization of annealing temperature: The reaction system was a basic reaction system. Under the same conditions as other reactions, the annealing temperature was adjusted to perform gradient PCR (annealing temperatures were 58°C, 60°C, 62°C, and 64°C). The results showed that the amplification effect of the reaction system with an annealing temperature of 62°C was the best, and the optimal annealing temperature was determined to be 62°C.

[0181] It is thus obtained that the above basic reaction system is the optimal reaction system, and the basic reaction conditions are the optimal reaction conditions.

[0182] After the reaction is completed, use 0.6 times the volume of AMPure XP magnetic beads for purification, and finally elute the magnetic beads with TE solution, transfer to a new centrifuge tube, and proceed to the next step, or store at ~80°C for later use.

[0183] 4.7 Library Quality Control

[0184] The library was tested by Qbuit and Agilent 2100 Bioanalyzer, and the qualified library (main peak above 2K was judged to be qualified) was used for the next step of third-generation sequencing library construction.

[0185] 5. Construction and sequencing of third-generation sequencing libraries

[0186] 5.1 Sequencing using Pacbio Sequal

[0187] 5.1.1 Use library construction kit to build library

[0188] 5.1.1.1 Repair the mixed library

[0189] Table 13 Repair solution preparation

[0190] Reagent components Volume (μl) DNA library 35 ATP high 1.5 <![CDATA[NAD + ]]> 1.5 dNTP 2 Repair buffer 7 Repair enzymes 3 Total volume 50

[0191] After the repair solution is prepared, mix it thoroughly, centrifuge it, and place it in a PCR thermal cycler to perform the repair reaction. The specific conditions are as follows:

[0192] Table 14 Repair reaction procedure

[0193] Reaction steps Reaction temperature Reaction time Repair reaction 37℃ 30min save 4℃ ......

[0194] 5.1.1.2 Purification

[0195] PB magnetic beads with a volume of 0.45 times that of the sample were used for purification, and the beads were finally eluted with double-distilled water and stored in a -20°C refrigerator.

[0196] 5.1.1.3 Connector connection

[0197] Table 15 Connection solution system

[0198]

[0199]

[0200] Mix well, centrifuge, and perform the ligation reaction in a PCR thermal cycler. The specific conditions are as follows:

[0201] Table 16 Ligation reaction program

[0202] Reaction steps Reaction temperature Reaction time Connector connection 25℃ 15h save 4℃ ......

[0203] 5.1.1.4 Purification

[0204] PB magnetic beads with a volume of 0.45 times that of the sample were used for purification, and the beads were finally eluted with double-distilled water and stored in a -20°C refrigerator.

[0205] 5.1.2 Primer annealing and binding reaction

[0206] The sequencing was performed according to the standard operation method of Pacbio sequal instrument.

[0207] 5.1.3 Third-generation sequencing

[0208] Sequencing was performed according to the standard operating procedures of the Pacbio sequal instrument.

[0209] 5.2 Sequencing using Oxford Nanopore PromethION

[0210] 5.2.1 Sample end repair

[0211] Remove DNA, place on ice, add NEB End Repair / A-Tail Ligation Reagent, and mix thoroughly. Incubate at 20°C for 40 minutes and then at 65°C for 20 minutes.

[0212] 5.2.2 DNA purification

[0213] Add 1×AMPure beads to DNA, incubate at room temperature for 15 minutes, adsorb on a magnetic stand at room temperature for 5 minutes, and discard the supernatant.

[0214] Add 80% ethanol, adsorb on a magnetic rack, discard the supernatant, and repeat once. Allow to dry at room temperature.

[0215] Add Ultra Pure Water and elute by pipetting at 37°C.

[0216] Let it stand on the magnetic rack for 5 minutes, and then aspirate the supernatant to obtain the purified DNA.

[0217] 5.2.3 Connector connection

[0218] Add NEB T4 DNA Rapid Ligation Buffer, NEB T4 DNA Rapid Ligase, and adapters, mix well, and incubate at 20°C for 20 min.

[0219] 5.2.4 DNA purification

[0220] Add 0.8× AMPure beads to the PCR product, incubate at room temperature for 5 minutes, adsorb on a magnetic rack at room temperature for 2 minutes, and discard the supernatant. Add 200 μl of SFB, mix thoroughly by pipetting, adsorb on a magnetic rack, and discard the supernatant. Repeat once. Add 15 μl of EB and elute by pipetting. Let stand on the magnetic rack for 5 minutes, then aspirate the supernatant to obtain the purified DNA.

[0221] 5.2.5 Priming and Loading Operations

[0222] The sequencing was performed according to the standard Oxford Nanopore PromethION sequencing method.

[0223] 5.2.6 Third-generation sequencing

[0224] Sequencing was performed according to the standard operating procedures of the Oxford Nanopore PromethION instrument.

[0225] 6. Bioinformatics Analysis

[0226] 6.1 Splitting sample data

[0227] If the data is a mixed test of multiple samples, the data of each sample is split according to the label sequence of each sample.

[0228] 6.2 Data Filtering

[0229] Generate high-quality data based on the data characteristics of each platform.

[0230] The ONT platform filters out data with low quality values. The Pacbio platform uses official software to generate high-quality consensus sequences.

[0231] 6.3 Data Evaluation

[0232] High-quality data are aligned to the human whole genome reference sequence, using a unique alignment (i.e., a single sequence can only be uniquely aligned to one position in the genome).

[0233] Count the number of sequences in the target region and evaluate the capture efficiency, coverage depth, and coverage completeness of the target region.

[0234] If the data is unqualified (the qualified standard is: sequencing coverage of the target region, 30-fold coverage is required to be no less than 95%), retest the data according to the situation.

[0235] 6.4 STR locus tandem repeat copy number detection

[0236] The number of tandem repeat copies of STR loci in the target region was detected, and STR loci with repeat copy numbers reaching the pathogenic threshold were screened out. The bioinformatics analysis software used in the data calculation process was the independently developed GrandSTR. The data calculation process of this software is as follows: Figure 2 .

[0237] 6.4.1 Extracting STR Region Sequences from Reads: Based on the bam format alignment file, extract the STR region sequence from the corresponding read by extending 30 bp on both sides of the STR region coordinates in the reference sequence. If the bam format alignment file does not extract the STR region sequence, re-align the reference sequence with the read sequence and extract the STR region sequence from the read.

[0238] 6.4.2 Calculating the Number of STR Sequence Repeats in Reads: If a motif change is considered, recalculate the number of repeat units in the extracted STR sequences to see if they are consistent with the expected repeat units. If the repeat units are consistent, the motif is considered unchanged. Align the STR sequence with a template sequence containing multiple copies of the STR repeat unit. Adjust the boundaries, remove poorly aligned sequences at both ends of the alignment, or extend sequences with good alignment to both ends, and calculate similarity and the number of repeats.

[0239] 6.4.3 Calculating Peak Repeat Counts Using a Gaussian Mixture Model: Based on the number of repeats of STR sequences in multiple reads, a Gaussian mixture model with the n_component parameter set to 1 or 2 is used to fit the repeat count distribution. The Akaike information criterion (AIC) is then calculated, and the model with the smaller AIC is selected. If the selected model has an n_component parameter of 1, the peak average value of the model is directly returned as the repeat count of the sample. If the selected model has an n_component parameter of 2, the repeat count of the STR in the sample is selected as the number of repeats near the peak average value of the model and with a read count greater than or equal to the minimum read count threshold (2).

[0240] 6.5 Validation of mutations.

[0241] The sites with STR expansion screened out will be verified by other means.

[0242] This example examined 20 families diagnosed with dynamic mutation diseases through clinical phenotypic observation, totaling 60 samples. STR loci were tested and the results of the present invention were compared with those of second-generation sequencing. Any discrepancies were verified with first-generation sequencing. During the testing process, neither the experimenter nor the data analyst was aware of the sample's genotype or phenotype to ensure the credibility of the results. The test results are shown in Table 17:

[0243] Table 17 Test results

[0244]

[0245]

[0246]

[0247] Note: N / A indicates not tested. Southern hybridization validation has a resolution error of ±50 bp. Therefore, Southern hybridization validation values ​​within the range of this method after deducting this error are considered consistent with the results of this method.

[0248] We also compared STR locus tandem repeat copy number detection software. This test data was compared using existing STR analysis tools, including Straglr, repeatHMM, tandem-genotypes, and PACMONSTR, with our independently developed GrandSTR. Repeat values ​​between 1 and 100 were considered consistent within a ±2% deviation from the theoretical value; values ​​between 101 and 300 were considered consistent within a ±5% deviation from the theoretical value; and values ​​above 300 were considered consistent within a ±15% deviation from the theoretical value. The comparison results are as follows:

[0249] Table 18 Comparison of STR locus tandem repeat copy number detection software

[0250]

[0251]

[0252] The above comparison results show that the results of the bioinformatics software GrandSTR of the present invention are 100% consistent with the theoretical repeat value. The results of the Straglr, repeatHMM, tandem-genotypes, and PACMONSTR software all have multiple samples that do not match the theoretical repeat value. This shows that the results of the bioinformatics software GrandSTR of the present invention are superior to those of the Straglr, repeatHMM, tandem-genotypes, and PACMONSTR software.

[0253] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A probe combination, characterized in that: Probes target the following STR loci: SCA1, SCA2, SCA3, SCA6, SCA7, SCA8, SCA10, SCA12, SCA17, SCA31, SCA36, SCA37, SCA50, CANVAS, DRPLA, FRDA, GDPAG, FAME1, FAME2, FAME3, FAME4, FAME6, FAME7, EPM1A, OPDM1, OPDM2, OPDM3, OPDM4, DM1, DM2, SBMA, FTDALS1, FECD3 HD, HDL1, HDL2, DIP2B, FXS, SPD1, CCHS, CBL, HFG, BPES, XLMR, EIF4A3, TBX1, ZIC3, NIPA1, POLG, THAP11, OPML1, HPE5, AFF2, DBQD2, RUNX2, ARX, HMN, DIP2B; The STR loci are derived from humans, the STR loci reference sequence is the human reference genome version Hg19, and the probe combination includes the sequences shown in Table 3.

2. The probe combination according to claim 1, characterized in that Each probe in the probe combination is connected to a label, and the label is selected from at least one of biotin, avidin, avidin, an antibody or a chemical coupling agent.

3. A method for sequencing STR loci, characterized in that: The following steps are involved: Step 1: The genomic DNA in the sample to be tested is sheared, end-repaired, adapter-added, and amplified to obtain a DNA library; Step 2, the probe combination according to claim 1 or 2 is blocked and hybridized with the DNA library to obtain a pre-hybridized library, and the pre-hybridized library is amplified to obtain a hybridized library; Step 3: Perform third-generation sequencing on the hybridization library and analyze the data to obtain STR locus sequence information.

4. The sequencing method according to claim 3, wherein In step 1, The size of the interrupted gene fragment is 4.5 to 5.5 kb; The linker includes a linker A and a linker B, wherein the nucleotide sequence of linker A is shown in SEQ ID NO: 1, and the nucleotide sequence of linker B is shown in SEQ ID NO: 2; The amplification primers include Primer-F and Primer-R. The nucleotide sequence of Primer-F is shown in SEQ ID NO: 3, and the nucleotide sequence of Primer-R is shown in SEQ ID NO:

4.

5. The sequencing method according to claim 3 or 4, characterized in that In step 2, The blocking primers include Index A Block and Index B Block. The nucleotide sequence of Index A Block is shown in SEQ ID NO: 5, and the nucleotide sequence of Index B Block is shown in SEQ ID NO:

6. The amplification primers include Primer-F and Primer-R, the nucleotide sequence of Primer-F is shown in SEQ ID NO: 3, and the nucleotide sequence of Primer-R is shown in SEQ ID NO: 4; Amplification system includes: 0M~5M Betaine, 0.1μM~1.0μM Primer F, 0.1μM~0.5μM Primer R, 0.5mM~4.0mM Mg 2+ , 8 vol% dNTP Mixture, polymerase, polymerase buffer, pre-hybridization library; the concentration of each dNTP in the dNTPMixture is 0.5 mM to 5 mM; the polymerase is selected from at least one of EX Taq DNA polymerase, Pfu DNA polymerase, rTaq DNA polymerase and / or PrimeSTAR GXL polymerase; The amplification program included: 98°C for 1 min; 98°C for 15 sec, 58°C to 64°C for 15 sec, 68°C for 6 min, 18 cycles; 68°C for 5 min.

6. The sequencing method according to claim 5, characterized in that In the step 2, The amplification system is: 1M Betaine, 0.5μM Primer F, 0.5μM Primer R, 2mM Mg 2+ , 8vol% dNTPMixture, 2vol% PrimeSTAR GXL DNA Polymerase, 20vol% 5× PrimeSTAR GXL Buffer, 50 vol% prehybridization library; The concentration of each dNTP in dNTP Mixture is 2.5 mM; The amplification program included: 98°C for 1 min; 98°C for 15 sec, 62°C for 15 sec, 68°C for 6 min, 18 cycles; 68℃5min.

7. The sequencing method according to claim 3, 4 or 6, characterized in that: In step 3, the method for analyzing the data is: after extracting the STR region sequence data, a Gaussian mixture model with a smaller AIC value calculated when the parameter n_component is 1 or 2 is selected to determine the number of STR repeats.

8. The sequencing method according to claim 5, characterized in that In step 3, the method for analyzing the data is: after extracting the STR region sequence data, a Gaussian mixture model with a smaller AIC value calculated when the parameter n_component is 1 or 2 is selected to determine the number of STR repeats.

9. The sequencing method according to claim 7, characterized in that The criteria for determining the number of STR repeats are: If the n_component of the selected model is 1, the peak average value of the model is directly returned as the number of repetitions of the sample; If the n_component of the selected model is 2, the number of repetitions near the peak average of the model and with a read number greater than or equal to the minimum read number threshold of 2 is selected as the number of repetitions of the STR in the sample.

10. The sequencing method according to claim 8, characterized in that The criterion for determining the number of STR repetitions is: if the n_component of the selected model is 1, the peak average value of the model is directly returned as the number of sample repetitions; If the n_component of the selected model is 2, the number of repetitions near the peak average of the model and with a read number greater than or equal to the minimum read number threshold of 2 is selected as the number of repetitions of the STR in the sample.

11. Use of the probe combination according to claim 1 or 2 in preparing a kit for screening STR locus-related diseases.

12. A product for detecting STR loci, comprising the probe combination of claim 1 or 2, and one or more of a solid support, an adapter sequence, an adapter blocking sequence, primers for binding to the adapter sequence and amplifying nucleic acid fragments, a DNA extraction system, a PCR reaction buffer, nuclease-free water, a DNA polymerase, a molecular weight marker, a target sequence eluent, an end-repair enzyme, an end-repair buffer, and a DNA ligase.

Citation Information

Patent Citations

  • Personalized delivery vector-based immunotherapy and uses thereof

    CN107847611A

  • Method for detecting short tandem repeat expansion and genotyping, electronic equipment and storage medium

    CN115240770A