Primer combinations, kits and systems for detecting variations in the SLC25A13 IVS16 region
By designing primer compositions and constructing multiplex PCR targeted sequencing libraries, and combining them with decision tree algorithms from machine learning models, the high cost of third-generation sequencing technology was solved, enabling efficient and accurate detection of SLC25A13 IVS16 region variants under second-generation sequencing technology.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SUZHOU SMK GENE TECH LTD
- Filing Date
- 2022-10-17
- Publication Date
- 2026-04-10
AI Technical Summary
The use of third-generation long-fragment sequencing technology to detect variants in the SLC25A13 IVS16 region is costly and difficult to achieve accurate detection.
Primer compositions were designed for amplification of the SLC25A13 IVS16 region using next-generation sequencing technology, and combined with multiplex PCR targeted sequencing library construction and machine learning models, a decision tree algorithm was used for variant detection.
This technology enables stable and accurate detection of SLC25A13 intron insertion variants using next-generation sequencing, reducing detection costs and improving the speed and accuracy of detection.
Smart Images

Figure CN115725720B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of gene detection, and particularly to a primer combination, a kit and a system for detecting SLC25A13 IVS16 region variation. BACKGROUND
[0002] Citrin deficiency (CD) is an autosomal recessive genetic disease, the pathogenic gene SLC25A13 of which is located on chromosome 7q21.3, and the encoded protein is called citrin.
[0003] IVS16ins3kb is a common pathogenic variation of SLC25A13 gene. Since the inserted sequence fragment is relatively long, it is difficult to accurately detect the variation by means of probe capture and other means. At present, the insertion fragment can be completely amplified by using three-generation long fragment sequencing technology, but the three-generation sequencing technology has the problems of high cost and high application price. SUMMARY
[0004] The embodiment of the present application provides a primer combination for amplifying the SLC25A13 IVS16 region under the second-generation sequencing technology, so as to solve the technical problem of high cost in the prior art using the three-generation long fragment sequencing technology.
[0005] The primer combination is for the SLC25A13 IVS16 region, and includes an upstream primer and a downstream primer. The sequence of the upstream primer is as follows: AAACTGGGGTGAGGATCGAAATACACGAGC. The downstream primer includes a mutant downstream primer and a wild type downstream primer. The sequence of the mutant downstream primer is as follows: GCCCGAACCCCTTTCCACTGCCAACACCTC. The sequence of the wild type downstream primer is as follows: CTGGCCAAACCACTTACAGCGGAGTGATAG.
[0006] Further, the primer combination further includes at least one control primer combination.
[0007] For the amplification region CG0175.ACAD9.NM_014049.5.exon7_control_ACAD9, the sequence of the upstream primer is as follows: ATAGGGGTTTGGTTTTCTCCAAAGTC, and the sequence of the downstream primer is as follows: CGCGCACAGGAGCTACTT.
[0008] For the amplification region CG0539.CYP11B1.NM_000497.3.exon8_control, the upstream primer sequence is as follows: CTCTCAGCTCGCCGCTTAC, and the downstream primer sequence is as follows: GACATGGTCCCATCCAGCAC;
[0009] For the amplification region CG0095.INSRR.NM_014215.3.exon2_control_NTRK1_exon2, the upstream primer sequence is as follows: TCCTGATGCCTAGCTTAAGGGAGTC, and the downstream primer sequence is as follows: GCATTGGGGGAAATGATCCAAATG;
[0010] For the amplification region CG0335.SIL1.NM_001037633.2.exon5_control_A001, the upstream primer sequence is as follows: TCTGTGCTCTGGGAGAGAAGTAAA, and the downstream primer sequence is as follows: GAGACTGACATGCAGATCATGGTACG;
[0011] For the amplification region CG0336.SIL1.NM_001037633.2.exon5_control_A002, the upstream primer sequence is as follows: CAGCAATCTTCTCTTCCAAACTGGAGC, and the downstream primer sequence is as follows: CCATGGTAGACCACAGACATCTTGGGC;
[0012] For the amplification region CG0593.STIM1.NM_001277961.1.exon10_control_A001, the upstream primer sequence is as follows: AAGTCCATGCCTGCAGTTCTCTT, and the downstream primer sequence is as follows: ATCCACGTCGTCAGTCATGATGAAG;
[0013] For the amplification region CG0594.STIM1.NM_001277961.1.exon10_control_A002, the upstream primer sequence is as follows: AAGTCCATGCCTGCAGTTCTCTT, and the downstream primer sequence is as follows: AAAGGCTCCTTCCTTCATCCCCGC;
[0014] For the amplification region CG0740.EDNRB.NM_001201397.1.exon3_control, the upstream primer sequence is as follows: GGAAACACTTCTGAGTGGCATTTATTTA, and the downstream primer sequence is as follows: TGAGTAAAATGAGCCATCTTTTAAGGGTCA;
[0015] For the amplification region CG0174.IQCB1.NM_001023570.exon3_control, the upstream primer sequence is as follows: GTAATACTGATATGGTACAGAAGCTTCATACCAA, and the downstream primer sequence is as follows: GTTAGGGGAGAAAAATCACAAACCTTCA.
[0016] Another embodiment of the present application provides a use of the above-mentioned primer composition in the preparation of a kit for detecting a variation in the SLC25A13 IVS16 region.
[0017] Another embodiment of the present application provides a kit comprising the above-mentioned primer composition.
[0018] Further, the kit further comprises a DNA polymerase and a reaction buffer.
[0019] A third embodiment of the present application provides a system for detecting a variation in the SLC25A13 IVS16 region, comprising the following modules:
[0020] A collection module for obtaining a sample of a subject;
[0021] An amplification module for performing PCR amplification on the sample;
[0022] A library construction module for constructing a multiplex PCR targeted sequencing library;
[0023] A sequencing module for sequencing and analysis;
[0024] The amplification module uses the primer composition of claim 1 or 2, and the sequencing module uses a second-generation sequencing technology.
[0025] Further, the sequencing module comprises a pre-trained machine learning model, and the sample data after PCR amplification and construction of a multiplex PCR targeted sequencing library is input into the machine learning model to obtain an analysis result corresponding to the sample data.
[0026] Further, the pre-trained machine learning model is trained by dividing historical samples into a training set and a test set in a ratio of 7:3.
[0027] Further, the machine learning model adopts a decision tree algorithm for classification, and the decision tree algorithm is constructed based on four parameters of total sequencing depth, wild type sequence number, mutant sequence number and variant allele frequency of each sample.
[0028] Further, the decision tree algorithm has the following standards respectively:
[0029] The decision standard for being positive is to meet the following conditions:
[0030] Total sequencing depth > 50X, and mutant sequence number ≥ 10, and mutant proportion ≥ 10%;
[0031] The decision standard for being negative is to meet the following conditions:
[0032] Total sequencing depth ≥ 50X, and mutant sequence number < 10;
[0033] The decision standard for being indeterminate is to meet the following conditions:
[0034] Total sequencing depth ≤ 50X.
[0035] In the embodiments of the present application, a new primer composition is provided, which can stably and accurately detect SLC25A13 intron insertion variation under the second-generation sequencing technology, and the detection of special variation is carried out by using the advantages of high-throughput sequencing, bioinformatics algorithm and decision tree for threshold judgment, so as to ensure the rapidity and accuracy of the detection. BRIEF DESCRIPTION OF DRAWINGS
[0036] The accompanying drawings, which form a part of this application, are intended to provide further understanding of the application and are incorporated herein in their entirety, and the illustrative embodiments thereof and their description serve the purpose of explanations rather than limiting the present application. In the drawings:
[0037] Figure 1 is a flow chart of a method for detecting SLC25A13 IVS16 region variation according to the embodiments of the present application. DETAILED DESCRIPTION
[0038] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0039] It should be noted that the steps shown in the flow chart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flow chart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.
[0040] The embodiment of the present application provides a primer composition for SLC25A13 IVS16 region, which is used for amplifying the SLC25A13 IVS16 region under a second-generation sequencing technology, so as to solve the technical problem of high cost in the prior art using a third-generation long fragment sequencing technology.
[0041] Specifically, the primer composition of the embodiment comprises an upstream primer and a downstream primer, the sequence of the upstream primer is as follows: AAACTGGGGTGAGGATCGAAATACACGAGC, and the downstream primer comprises a mutant downstream primer and a wild type downstream primer, the sequence of the mutant downstream primer is as follows: GCCCGAACCCCTTTCCACTGCCAACACCTC, and the sequence of the wild type downstream primer is as follows: CTGGCCAAACCACTTACAGCGGAGTGATAG.
[0042] Targeted sequencing technology can enrich the genomic region of interest for sequencing, and the single sample sequencing data output is less and the analysis speed is faster, so that the advantages of NGS technology can be more economically and efficiently exerted, and the targeted sequencing technology is widely applied to many fields such as clinical detection and health screening. In addition, targeted sequencing can perform deep sequencing on the target region, thereby increasing the detection sensitivity and accuracy of genetic variations in the target region.
[0043] The method of targeted sequencing mainly includes two kinds: hybrid capture sequencing and amplicon sequencing. Amplicon sequencing is a technology for designing PCR primers to amplify and enrich the target region of interest and then performing sequencing. It is usually suitable for detecting tens to thousands of sites, or regions below several tens of kb. Hybrid capture sequencing currently mainly uses liquid hybrid capture sequencing, that is, based on the principle of base complementary pairing, nucleic acid probes are designed and synthesized, and DNA library is hybridized and enriched in the target region based on a liquid environment, and then sequenced. However, the liquid hybrid capture operation is difficult, the operation time is long, and it is easily affected by the probe capture efficiency, so the amplicon sequencing is more suitable for non-professional technical personnel to operate. Multiplex PCR as a method for quickly constructing a targeted sequencing library plays an increasingly important role in the field of clinical gene detection and research due to its high efficiency, systematicness, and economic simplicity.
[0044] In order to ensure normal amplification or only amplification of target detection point mutations, the embodiment also provides a control primer combination for amplifying other regions, including at least one of the following control primer combinations:
[0045] For amplification region CG0175.ACAD9.NM_014049.5.exon7_control_ACAD9, the upstream primer sequence is as follows: ATAGGGGTTTGGTTTTCTCCAAAGTC, and the downstream primer sequence is as follows: CGCGCACAGGAGCTACTT;
[0046] For amplification region CG0539.CYP11B1.NM_000497.3.exon8_control, the upstream primer sequence is as follows: CTCTCAGCTCGCCGCTTAC, and the downstream primer sequence is as follows: GACATGGTCCCATCCAGCAC;
[0047] For amplification region CG0095.INSRR.NM_014215.3.exon2_control_NTRK1_exon2, the upstream primer sequence is as follows: TCCTGATGCCTAGCTTAAGGGAGTC, and the downstream primer sequence is as follows: GCATTGGGGGAAATGATCCAAATG;
[0048] For amplification region CG0335.SIL1.NM_001037633.2.exon5_control_A001, the upstream primer sequence is as follows: TCTGTGCTCTGGGAGAGAAGTAAA, and the downstream primer sequence is as follows: GAGACTGACATGCAGATCATGGTACG;
[0049] For amplification region CG0336.SIL1.NM_001037633.2.exon5_control_A002, the upstream primer sequence is as follows: CAGCAATCTTCTCTTCCAAACTGGAGC, and the downstream primer sequence is as follows: CCATGGTAGACCACAGACATCTTGGGC;
[0050] For amplification region CG0593.STIM1.NM_001277961.1.exon10_control_A001, the upstream primer sequence is as follows: AAGTCCATGCCTGCAGTTCTCTT, and the downstream primer sequence is as follows: ATCCACGTCGTCAGTCATGATGAAG;
[0051] For amplification region CG0594.STIM1.NM_001277961.1.exon10_control_A002, the upstream primer sequence is as follows: AAGTCCATGCCTGCAGTTCTCTT, and the downstream primer sequence is as follows: AAAGGCTCCTTCCTTCATCCCCGC;
[0052] For the amplification region CG0740.EDNRB.NM_001201397.1.exon3_control, the upstream primer sequence is as follows: GGAAACACTTCTGAGTGGCATTTATTTA, and the downstream primer sequence is as follows: TGAGTAAAATGAGCCATCTTTTAAGGGTCA;
[0053] For the amplification region CG0174.IQCB1.NM_001023570.exon3_control, the upstream primer sequence is as follows: GTAATACTGATATGGTACAGAAGCTTCATACCAA, and the downstream primer sequence is as follows: GTTAGGGGAGAAAAATCACAAACCTTCA.
[0054] The above-mentioned primer composition is applied as the content of the kit alone or together with one or more of the primer compositions of the control group to prepare a kit for detecting the SLC25A13 IVS16 region variation, and such a kit can realize the detection of the SLC25A13 IVS16 region variation by using the second-generation sequencing technology instead of the third-generation sequencing technology.
[0055] In these embodiments, the kit usually further comprises a DNA polymerase and a reaction buffer.
[0056] Since the variation insertion fragment is long, only part of the fragment is captured by the probe, and the detection effect of the variation is affected due to the limitation of sequencing depth and the like, and the multiplex amplicon sequencing can specifically amplify the variation fragment, therefore, in some embodiments of the application, bioinformatics means is supplemented to effectively avoid the randomness of the amplified fragment. Therefore, in some embodiments, a system for detecting the SLC25A13 IVS16 region variation is provided, comprising the following modules:
[0057] The collection module is used to obtain a sample of a subject.
[0058] The amplification module is used to perform PCR amplification on the sample by using the primer composition in the above-mentioned embodiments, comprising constructing an amplification region file and calculating the depth of each amplicon according to the amplification region.
[0059] The library construction module is used to construct a multiplex PCR targeted sequencing library.
[0060] The sequencing module is used to perform sequencing by using the second-generation sequencing technology and analysis.
[0061] The system composed of the above-mentioned modules is used to execute a method for detecting the SLC25A13 IVS16 region variation, comprising:
[0062] Step S101, collecting a sample of a subject;
[0063] Step S102, performing PCR amplification on the sample using the primer composition in the above embodiments, including constructing an amplification region file and calculating the depth of each amplicon according to the amplification region;
[0064] Step S103, constructing a multiplex PCR targeted sequencing library;
[0065] Step S104, sequencing using next-generation sequencing technology and analysis.
[0066] In some embodiments, the sequencing module comprises a pre-trained machine learning model, and by inputting the sample data after PCR amplification and construction of a multiplex PCR targeted sequencing library into the machine learning model, an analysis result corresponding to the sample data is obtained.
[0067] In some embodiments, the pre-trained machine learning model is trained by dividing historical samples into a training set and a test set in a ratio of 7:3.
[0068] In some embodiments, the machine learning model uses a decision tree algorithm for classification, and the decision tree algorithm is constructed based on four parameters of total sequencing depth, number of wild-type sequences, number of mutant sequences, and variant allele frequency of each sample.
[0069] In some embodiments, the decision tree algorithm has the following criteria:
[0070] The decision criteria for a positive determination are as follows:
[0071] Total sequencing depth > 50X, and number of mutant sequences ≥ 10, and mutation ratio ≥ 10%;
[0072] The decision criteria for a negative determination are as follows:
[0073] Total sequencing depth ≥ 50X, and number of mutant sequences < 10;
[0074] The decision criteria for an indeterminate determination are as follows:
[0075] Total sequencing depth ≤ 50X.
[0076] The bioinformatics method described above analyzes a large amount of historical sample data to obtain a unified decision criterion, which is the key to accurately and quickly determining mutations based on the results of amplification. Therefore, the establishment process of this set of criteria is specifically introduced, including the following steps:
[0077] First, the historical sample sequencing data FASTQ is subjected to basic quality control, and the data quality control includes data quality Q20>90%, Q30>85%. Here, Q20 and Q30 have the following meanings: each base in the sequencing data has a corresponding quality value, and the quality value is Q20, the probability of incorrect identification is 1%, i.e. the error rate is 1%, or the correct rate is 99%; the quality value is Q30, the probability of incorrect identification is 0.1%, i.e. the error rate is 0.1%, or the correct rate is 99.9%.
[0078] Then, the data after quality control is subjected to bwa mem acceleration algorithm in sentieon software (NGS gene data analysis acceleration software) to obtain the unsorted BAM file after alignment, and the final sorted.bam file, i.e. the alignment data, is obtained after sorting according to the genome coordinates.
[0079] Specifically, the following steps are included:
[0080] According to the primer design file, the start and end coordinates of the amplicon are counted.
[0081] The primer design file includes the start and end coordinates of the forward primer and the start and end coordinates of the reverse primer. The primer amplification region file is constructed by the start coordinates of the forward primer and the end coordinates of the reverse primer.
[0082] The amplicon position information, wild type and mutant sequencing sequence number of the test sample are obtained.
[0083] According to the historical test sample sorted.bam file obtained above, the paird-end sequence alignment information of each pair of primers is used for subsequent statistics.
[0084] The start and end positions of the sequence alignment to the reference genome are the 5' end coordinates of the forward primer and the 5' end coordinates of the reverse primer of the primer amplification.
[0085] By designing the start and end coordinates of the forward and reverse primers of the wild type primer, the number of sequencing sequences amplified by each pair of primers can be calculated according to the 5' end coordinates of the forward primer and the 5' end coordinates of the reverse primer. The coordinates of the sequencing sequence (the position information of the amplicon corresponding to the sequencing sequence read in the reference genome) are compared with the primer coordinates (the position information of the primer designed for the known target amplicon in the primer design file). The number of wild type sequences is counted according to the fact that the left end position of the left read in the pair reads of the test sample in the SLC25A13 IVS16 region is consistent with the 5' end coordinates of the forward primer of the wild type primer, and the right end termination position of the right read is consistent with the 5' end coordinates of the reverse primer of the wild type primer. Then, the number of sequencing sequences corresponding to the wild type amplicon is obtained.
[0086] The mutant amplicon sequence is part of the SLC25A13 gene and the other part is the inserted sequence in the pair reads. Therefore, the sequences of the two positions that the sequence can be aligned to are extracted, and the sequence matching is performed according to the mutant primer designed in the first part, allowing 2 base mismatches, matching the left end start sequence of the pair sequence with the forward primer, and matching the right end end sequence of the pair sequence with the reverse primer. The extracted sequence is counted, and the number of mutant amplicon sequences is obtained.
[0087] The historical samples are selected, and the number of wild-type and mutant amplicons of the samples is calculated using the above calculation method.
[0088] The mutation ratio and other characteristics of the mutant and wild-type in the historical samples are summarized to determine the classification threshold.
[0089] The number of mutant and wild-type in the historical positive samples and negative samples is calculated respectively. The mutation ratio calculation formula is: mutant sequence number / (mutant sequence number+wild-type sequence number). According to the mutation ratio calculated by the above formula of the historical negative and positive samples, and according to the three characteristics of the total sequencing depth of the historical samples in this region (mutant sequence number+wild-type sequence number), mutant sequence number and mutation ratio, a suitable threshold is selected to determine the sample result. The sample characteristic information is shown in Table 1:
[0090] Table 1
[0091] Examples Depth Sample wild-type sequence number Sample mutant sequence number Variant allele frequency S1 1000 500 500 0.5 S2 2000 2000 0 0 … … … … …
[0092] According to the historical sample results, the classic classification algorithm of decision tree is used to obtain the best threshold of the final feature in judging the positive result and the negative result. According to the historical results, the training set and the test set are divided according to the ratio of 7 to 3, and the corresponding decision tree is constructed based on the decision tree algorithm in the above four parameters. The final result is as follows: when the sample parameter mutant sequence number >=10, and the mutation ratio >=0.01, the sample can be judged as a positive result.
[0093] In the threshold selection process, it is found that when the total depth of the sample is less than 50X, due to the specificity problem of PCR amplification, it is impossible to determine whether the region is effectively amplified, so the result is determined as undetermined, which can be determined by experimental means whether other variations occur in the region to cause no amplification, or re-sequencing. Therefore, when the total sequencing depth of the sample in the region is less than 50X (including 50X), the sample result is undetermined.
[0094] Therefore, the final sample judgment method is as follows:
[0095] The total sequencing depth, the number of mutant sequences, the wild-type mutation data and the mutation ratio of the sample to be detected are obtained, when the total sequencing depth is less than or equal to 50X, the sample result is determined as undetermined; when the total sequencing depth is greater than or equal to 50X, if the number of mutant sequences is less than 10, the sample result is determined as negative; when the total sequencing depth is greater than 50X, the number of mutant sequences is greater than or equal to 10, if the mutation ratio is greater than or equal to 0.01, the sample result is determined as positive, otherwise as negative.
[0096] In the embodiment, an electronic device is also provided, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the method in the above embodiment.
[0097] These computer programs can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows or blocks or multiple flows or blocks can be implemented by different modules. Figure 1 The functions specified in one or more flows or blocks or multiple flows or blocks can be implemented by different modules.
[0098] The above programs can be run in a processor, or can also be stored in a memory (or called computer readable medium), the computer readable medium includes permanent and non-permanent, removable and non-removable media, which can be realized by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or any other non-transmission medium that can be used to store information accessible by a computing device. According to the definition in this paper, computer readable medium does not include transitory computer readable medium, such as modulated data signal and carrier wave.
[0099] Next, the use process of the above-mentioned system for detecting SLC25A13 IVS16 region variation is described through a specific embodiment.
[0100] After the sample is sequenced by PCR library construction, the PCR primer pool contains primer pairs that amplify wild-type sequences and primer pairs that amplify mutant sequences. After high-throughput sequencing, FASTQ data is obtained, and then the following steps are taken:
[0101] 1) Quality control, remove sequencing adapters and low-quality bases or sequences, and evaluate the quality of the raw data; data quality control includes, data quality Q20> 90%, Q30> 85%.
[0102] 2) Alignment, use BWA software to map the sequence information to the human reference genome after the first step of processing the FASTQ file, and sort the obtained BAM file to obtain the final BAM file.
[0103] 3) Depth calculation, according to the designed primer information and the BAM file obtained in the second step, the total depth of the sample in the SLC25A13 IVS16 region, the number of wild-type sequences and the number of mutant sequences are counted, and the corresponding mutation ratio is calculated;
[0104] 4) Collect historical sample information: after the historical sample is subjected to the above quality control, alignment and depth calculation, the total depth, the number of wild-type sequences, the number of mutant sequences and the mutation ratio in the region of the negative sample and the positive sample are summarized. The calculation method of the mutation ratio is: the number of mutant sequences / (the number of mutant sequences+the number of wild-type sequences).
[0105] 5) According to the results of the historical samples, use the decision tree (decision tree, model tree structure, a classic binary classification algorithm) classic classification algorithm to obtain the best threshold value of the final feature in judging positive and negative results. According to the historical results, divide them into training set and test set according to the ratio of 7:3, and construct the corresponding decision tree based on the decision tree algorithm in the above 4 parameters. The final result is as follows: when the sample parameter mutant sequence number> = 10, and the mutation ratio> = 0.01, the sample can be judged as a positive result.
[0106] At the same time, according to the specificity problem of PCR amplification, when the total sequencing depth is less than or equal to 50, it is impossible to determine whether the region is effectively amplified, so the result is determined as undetermined, which can be determined by experimental means whether other variations occur in the region to cause no amplification, or need to be re-sequenced.
[0107] Test example
[0108] Seven samples were selected, and each sample was tested three times.
[0109] The DNA of the test sample is extracted, broken, repaired, and then amplified by adding PCR primers. After library construction, sequencing is performed. The obtained data is FASTQ data.
[0110] After quality control, low-quality bases and adapter sequences are removed, and data quality statistics are performed, including data volume of 1.5G or more, average sequencing depth of 3000X, data quality Q20>90%, and Q30>85%.
[0111] After quality control and removal of low-quality bases and adapter sequences, the FASTQ sequence is mapped to the reference genome using the BWA module in the sentieon software. The alignment information of each sequence in FASTQ is obtained, and the BAM file storing the alignment information is obtained.
[0112] The number of wild-type sequences and mutant sequences of the test sample is calculated according to the BAM file of the sample and the wild-type primer sequence information and mutant primer sequence information of the amplified SLC25A13 IVS16 region. The number of wild-type sequences is calculated as follows: the left end position of the left sequence in the paired sequence of the test sample in the SLC25A13 IVS16 region is consistent with the 5' end coordinate of the wild-type primer forward primer, and the right end termination position of the right sequence is consistent with the 5' end coordinate of the wild-type primer reverse primer. Count. The total depth, wild-type sequence number, mutant sequence number, and mutation ratio in this region are determined according to the threshold value obtained in step 5. The number of mutant sequences is calculated as follows: the sequences of the two positions that the sequence may align to are extracted, and the mutant primer designed in the first part is used for sequence matching, allowing 2 base mismatches. The left end start sequence of the paired sequence matches the forward primer, and the right end termination sequence of the paired sequence matches the reverse primer. Count the extracted sequences to obtain the number of mutant amplicon sequences.
[0113] The results of 7 samples with 3 repeats are as follows:
[0114]
[0115]
[0116] The results of the test samples are determined by the above results and the threshold value obtained from the historical sample summary, and the determination results are as follows:
[0117]
[0118]
[0119] The sample and repeat results are consistent with the true results 100%.
[0120] Therefore, the embodiment is based on target region amplification high-throughput sequencing, in addition to simple operation and controllable cost, compared with hybrid capture sequencing, the mutation type of special long fragment insertion can be detected more accurately and efficiently.
[0121] According to the high-throughput data obtained by the PCR primer design, the sequence number of each amplicon is calculated according to the primer amplification coordinates of the wild type amplicon and the sequence characteristics of the mutant amplicon, and the historical samples are used as a training set to evaluate the characteristics of each sample and determine the final detection threshold. The IVS16ins3kb mutation result of SLC25A13 in the sample to be detected is detected, and the effect is considerable and the accuracy is high.
[0122] The above is only an embodiment of the present application and is not intended to limit the present application. Various changes and modifications can be made to the present application by those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the scope of the claims of the present application.
Claims
1. A system for detecting a SLC25A13 IVS16 region variation, characterized by, The method comprises the following modules: a collection module for collecting a sample of a subject; an amplification module for performing PCR amplification on the sample; a library construction module for constructing a multiplex PCR targeted sequencing library; a sequencing module for sequencing and analysis by using a second-generation sequencing technology; The amplification module uses the following primer combination: The primer combination used in amplification includes an upstream primer and a downstream primer for the SLC25A13 IVS16 region, the sequence of the upstream primer is as follows: 5'-AAACTGGGGTGAGGATCGAAATACACGAGC-3'; the downstream primer includes a mutant downstream primer and a wild-type downstream primer, the sequence of the mutant downstream primer is as follows: 5'-GCCCGAACCCCTTTCCACTGCCAACACCTC-3', and the sequence of the wild-type downstream primer is as follows: 5'-CTGGCCAAACCACTTACAGCGGAGTGATAG-3'; The primer combination at least further includes at least one control primer combination: for the amplification region CG0175.ACAD9.NM_014049.5.EXON7_CONTROL_ACAD9, the sequence of the upstream primer is as follows: 5'-ATAGGGGTTTGGTTTTCTCCAAAGTC-3', and the sequence of the downstream primer is as follows: 5'-CGCGCACAGGAGCTACTT-3'; for the amplification region CG0539.CYP11B1.NM_000497.3.EXON8_CONTROL, the sequence of the upstream primer is as follows: 5'-CTCTCAGCTCGCCGCTTAC-3', and the sequence of the downstream primer is as follows: 5'-GACATGGTCCCATCCAGCAC-3'; for the amplification region CG0095.INSRR.NM_014215.3.EXON2_CONTROL_NTRK1_EXON2, the sequence of the upstream primer is as follows: 5'-TCCTGATGCCTAGCTTAAGGGAGTC-3', and the sequence of the downstream primer is as follows: 5'-GCATTGGGGGAAATGATCCAAATG-3'; for the amplification region CG0335.SIL1.NM_001037633.2.EXON5_CONTROL_A001, the sequence of the upstream primer is as follows: 5'-TCTGTGCTCTGGGAGAGAAGTAAA-3', and the sequence of the downstream primer is as follows: 5'-GAGACTGACATGCAGATCATGGTACG-3'; For the amplification region CG0336.SIL1.NM_001037633.2.EXON5_CONTROL_A002, the upstream primer sequence is as follows: 5'-CAGCAATCTTCTCTTCCAAACTGGAGC-3', and the downstream primer sequence is as follows: 5'-CCATGGTAGACCACAGACATCTTGGGC-3'; For the amplification region CG0593.STIM1.NM_001277961.1.EXON10_CONTROL_A001, the upstream primer sequence is as follows: 5'-AAGTCCATGCCTGCAGTTCTCTT, and the downstream primer sequence is as follows: 5'-ATCCACGTCGTCAGTCATGATGAAG-3'; For the amplification region CG0594.STIM1.NM_001277961.1.EXON10_CONTROL_A002, the upstream primer sequence is as follows: 5'-AAGTCCATGCCTGCAGTTCTCTT-3', and the downstream primer sequence is as follows: 5'-AAAGGCTCCTTCCTTCATCCCCGC-3'; For the amplification region CG0740.EDNRB.NM_001201397.1.EXON3_CONTROL, the upstream primer sequence is as follows: 5'-GGAAACACTTCTGAGTGGCATTTATTTA-3', and the downstream primer sequence is as follows: 5'-TGAGTAAAATGAGCCATCTTTTAAGGGTCA-3'; For the amplification region CG0174.IQCB1.NM_001023570.EXON3_CONTROL, the upstream primer sequence is as follows: 5'-GTAATACTGATATGGTACAGAAGCTTCATACCAA-3', and the downstream primer sequence is as follows: 5'-GTTAGGGGAGAAAAATCACAAACCTTCA-3'; The sequencing module comprises a pre-trained machine learning model, and the analysis result corresponding to the sample data is obtained by inputting the sample data after PCR amplification and construction of a multiplex PCR target sequencing library into the machine learning model; The machine learning model adopts a decision tree algorithm for classification, which is constructed based on four parameters of total sequencing depth, wild type sequence number, mutant sequence number and variant allele frequency of each sample, and the classification standard of the decision tree algorithm is as follows: The decision standard for being positive is that the following conditions are met: Total sequencing depth > 50X, and mutant sequence number ≥ 10, and mutant proportion ≥ 10%, wherein the mutant proportion calculation formula is: mutant sequence number / (mutant sequence number+wild type sequence number); The decision standard for being negative is that the following conditions are met: Total sequencing depth ≥ 50X, and mutant sequence number < 10; The decision standard for being indeterminate is that the following conditions are met: The total sequencing depth is less than or equal to 50X.
2. The system of claim 1, wherein: The pre-trained machine learning model is trained by dividing historical samples into a training set and a test set in a ratio of 7:3.
Citation Information
Patent Citations
Primer set and kit for detecting four mutations of SLC25A13 gene and application thereof
CN109628587A
Machine learning platform for generating risk models
US20210375392A1