Method, device, equipment and medium for detecting point mutations in multiplex PCR sequencing

By splitting and merging the amplicons and sample sequencing sequences in multiplex PCR sequencing, the problem of false negative results in multiplex PCR sequencing was solved, and more accurate point mutation detection was achieved.

CN115662512BActive Publication Date: 2025-09-16SUZHOU SMK GENE TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211364223.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2025-09-16
Estimated Expiration
2042-11-02

AI Technical Summary

Technical Problem

False negative results are prone to occur when detecting point mutations in multiplex PCR sequencing, especially when there is overlap between amplicons, leading to inaccurate test results.

Method used

Multiple amplicons are divided into amplicon sets with no overlap, and sample sequencing sequences are divided into corresponding sets. Mutation detection is performed separately, and false negative results are avoided through splitting and merging analysis.

Benefits of technology

The stability and accuracy of the test results are improved, the false negative and false positive rates are reduced, and the reliability of the test results is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115662512B_ABST
    Figure CN115662512B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, equipment and medium for detecting point mutations in multiplex PCR sequencing, wherein the method is performed by dividing multiple amplicons into multiple amplicon sets; wherein each amplicon set includes at least one amplicon, and there is no interval overlap between the amplicons in the same amplicon set; dividing the sample sequencing sequence into multiple sample sequencing sequence sets; wherein the sample sequencing sequence set corresponds to the amplicon set one-to-one, and the sample sequencing sequence set and the amplicon set with a corresponding relationship satisfy that the sample sequencing sequence in the sample sequencing sequence set belongs to the amplicon in the amplicon set; and performing mutation detection on the multiple sample sequencing sequence sets respectively. This application solves the problem of false negative results that are prone to occur in point mutation detection in multiplex PCR sequencing in the prior art, thereby improving the stability and accuracy of the detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of gene detection, and in particular to a method, device, equipment and medium for detecting point mutations based on multiplex PCR sequencing. Background Art

[0002] Targeted sequencing technology can enrich and sequence genomic regions of interest. This produces low-quality sequencing data for a single sample and allows for rapid analysis. Therefore, it can more cost-effectively leverage the advantages of NGS technology and is widely used in numerous fields, including clinical testing and health screening. Multiplex PCR sequencing, a component of targeted sequencing, utilizes multiple amplicon sequencing, a technique that uses multiple PCR primers designed to amplify and enrich the target region of interest before sequencing. Multiplex PCR, as a method for rapidly constructing targeted sequencing libraries, is playing an increasingly important role in current clinical genetic testing and research due to its efficiency, systematic nature, and economical simplicity.

[0003] Point mutations are simple single-base substitutions, while indel mutations are the deletion or insertion of one or more bases, which introduce gaps in alignment. During the design process of multiplex PCR, overlap between amplicons is unavoidable, including partial overlap between amplified regions and overlap between amplified regions and primers. This overlap can lead to the following problems:

[0004] When the target mutation is in the overlapping part of the amplification regions of two amplicons, during conventional point mutation detection, the site is assessed for mutation based on the number of sequences supporting the mutated base and the total depth. If a gene dropout problem occurs in one of the amplicons, resulting in ineffective amplification of the sequence containing the target mutation, and the amplification efficiency of this amplicon is higher (i.e., the number of sequences amplified by this amplicon is greater), then according to conventional detection ideas, the mutation frequency of the target base at this site will be considered consistent with the background noise, resulting in missed false negative results.

[0005] When the target mutation is located in the amplification region of one amplicon and in the primer region of another amplicon, there will also be a false negative result when the two amplicons are amplified unevenly, resulting in a low site frequency and missed detection.

[0006] In the existing technology, there is no good solution to the problem caused by the above-mentioned overlapping situation, which makes the probability of false negative test results uncontrollable. Summary of the Invention

[0007] The embodiments of the present application provide a method, apparatus, device and medium for detecting point mutations in multiplex PCR sequencing, which is used to solve the problem that false negative results are prone to occur when detecting point mutations in multiplex PCR sequencing in the prior art.

[0008] According to one aspect of the present application, a method for detecting point mutations in multiplex PCR sequencing is provided, wherein a plurality of amplicons are divided into a plurality of amplicon sets; wherein each amplicon set includes at least one amplicon, and there is no overlap between any two amplicons in the same amplicon set;

[0009] Dividing the sample sequencing sequences into a plurality of sample sequencing sequence sets; wherein the sample sequencing sequence sets correspond to the amplicon sets in a one-to-one manner, and the sample sequencing sequence sets and the amplicon sets having a corresponding relationship satisfy that the sample sequencing sequences in the sample sequencing sequence sets belong to the amplicons in the amplicon set;

[0010] The plurality of sample sequencing sequence sets are respectively subjected to mutation detection.

[0011] Furthermore, dividing the multiple amplicons into multiple amplicon sets includes: each amplicon set includes one amplicon.

[0012] Furthermore, dividing the multiple amplicons into multiple amplicon sets includes: sorting the multiple amplicons according to the priorities of three dimensions: chromosome number, starting position of the forward primer in the amplicon, and ending position of the reverse primer in the amplicon, from large to small, and sorting the data of each dimension from small to large;

[0013] All the sorted amplicons are traversed and the traversed amplicons are divided into the amplicon sets. It is determined whether there is an interval overlap between the current amplicon and the amplicon set where the existing amplicons exist according to a fixed comparison order. If there is an interval overlap, the current amplicon is grouped into the first amplicon set that has no interval overlap with the amplicon; otherwise, the current amplicon is grouped into an amplicon set where no amplicon exists.

[0014] Furthermore, the sample sequencing sequences in the sample sequencing sequence set belong to the amplicons in the amplicon set, including: the difference between the starting position of the left end sequence of the sample sequencing sequence and the starting position of the forward primer of the amplicon in the amplicon set, and the difference between the ending position of the right end sequence of the same sample sequencing sequence and the ending position of the reverse primer in the same amplicon in the amplicon set are both within a preset difference range; or, the left end sequence and the right end sequence of the sample sequencing sequence are both located in a certain amplicon region in the amplicon set and the proportion of the overlapping region between them and the certain amplicon to the total length of the certain amplicon is greater than a preset proportion.

[0015] Furthermore, the amplicon set is a BED file.

[0016] Furthermore, the sample sequencing sequence set is a BAM file.

[0017] Furthermore, after performing mutation detection on the plurality of sample sequencing sequence sets respectively, the method further comprises:

[0018] The results of each mutation test are filtered to remove false positive sites;

[0019] Combine the multiple results and obtain the test results according to the following different situations:

[0020] When a locus is detected in only one amplicon, or the genotype is consistent in multiple amplicons, the test result is no change in genotype;

[0021] When a locus is detected in multiple amplicons and the genotypes are inconsistent, and a 0 / 1 heterozygous genotype appears, the test result is based on the 0 / 1 heterozygous genotype;

[0022] When a locus is detected in multiple amplicons and the genotypes are inconsistent, there is no heterozygous genotype, and at least one 0 / 0 wild type and one 1 / 1 homozygous type are detected, the test result is detected according to the 0 / 1 heterozygous genotype.

[0023] A second aspect of the present application provides a device for detecting point mutations in multiplex PCR sequencing, comprising:

[0024] A first partitioning module is configured to partition the plurality of amplicons into a plurality of amplicon sets, wherein each amplicon set includes at least one amplicon, and no two amplicons in the same amplicon set have overlapping intervals;

[0025] A second partitioning module is configured to partition the sample sequencing sequence into a plurality of sample sequencing sequence sets; wherein the sample sequencing sequence sets correspond to the amplicon sets in a one-to-one manner, and the sample sequencing sequence sets and the amplicon sets having a corresponding relationship satisfy that the sample sequencing sequences in the sample sequencing sequence sets belong to the amplicons in the amplicon set;

[0026] The mutation detection module is used to perform mutation detection on the multiple sample sequencing sequence sets respectively.

[0027] According to a third aspect of the present application, an electronic device is provided, comprising a processor, a memory, and a computer program stored in the memory, wherein the computer program is configured to execute the method described in the first aspect when executed by the processor.

[0028] A fourth aspect of the present application provides a storage medium having a computer program stored thereon, wherein the computer program is used to execute the method described in the first aspect above.

[0029] In an embodiment of the present application, a method for detecting point mutations in multiplex PCR sequencing is adopted, wherein multiple amplicons are divided into multiple amplicon sets; wherein each amplicon set includes at least one amplicon, and there is no interval overlap between the amplicons in the same amplicon set; the sample sequencing sequence is divided into multiple sample sequencing sequence sets; wherein the sample sequencing sequence set corresponds to the amplicon set one-to-one, and the sample sequencing sequence set and the amplicon set with a corresponding relationship satisfy that the sample sequencing sequence in the sample sequencing sequence set belongs to the amplicon in the amplicon set; and the multiple sample sequencing sequence sets are subjected to mutation detection respectively. This application solves the problem of false negative results that are prone to occur in point mutation detection in multiplex PCR sequencing in the prior art, thereby improving the stability and accuracy of the detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0031] Figure 1 This is a flow chart of a method for detecting point mutations in multiplex PCR sequencing according to an embodiment of the present application;

[0032] Figure 2 is a schematic diagram of the relationship between the amplicon and the target amplification region according to an embodiment of the present application;

[0033] Figure 3 is a schematic diagram of the division of amplicons into different amplicon combinations according to an embodiment of the present application;

[0034] Figure 4 is a schematic diagram of dividing the sample sequencing sequences into multiple sample sequencing sequence sets according to an embodiment of the present application;

[0035] Figure 5 Schematic diagram of performing mutation detection, filtering, and merging on a set of sequencing sequences of multiple samples according to an embodiment of the present application. DETAILED DESCRIPTION

[0036] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0037] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0038] Multiplex PCR testing is a common genetic testing method. Its main process includes obtaining subject samples, then PCR amplifying the samples, constructing multiplex PCR targeted sequencing libraries and sequencing them on a machine, and performing mutation detection and analysis based on the high-throughput sequencing data of the obtained samples.

[0039] The embodiments of the present invention provide a method for detecting point mutations in multiplex PCR sequencing. Based on traditional multiplex PCR detection methods, the method improves the post-sequencing mutation detection and analysis process. It cleverly splits high-throughput sequencing data for separate detection and subsequent combined analysis. This method eliminates the unavoidable false negatives that can occur when the target mutation is in the overlapping region of two amplicons, or when the target mutation is in the amplification region of one amplicon while simultaneously in the primer region of another amplicon, which are often caused by overall detection.

[0040] like Figure 1 As shown, the method for detecting point mutations in multiplex PCR sequencing comprises the following steps:

[0041] Step S102: Divide the multiple amplicons into multiple amplicon sets; wherein each amplicon set includes at least one amplicon, and there is no overlap between any two amplicons in the same amplicon set.

[0042] The overlap of no interval between the two amplicons can be judged by comparing the coordinates of the beginning and the end of the amplicons. In the primer design file, the starting coordinates and the ending coordinates of the forward primer and the starting coordinates and the ending coordinates of the reverse primer of each target amplicon are recorded. The primer amplification region file is constructed with the starting coordinates of the forward primer and the ending coordinates of the reverse primer. The target amplification region of each target amplicon is the 3' coordinate of the forward primer and the 3' coordinate of the reverse primer, and the region of the amplicon is the 5' end coordinate of the forward primer and the 5' end coordinate of the reverse primer.

[0043] Step S104: Divide the sample sequencing sequence into multiple sample sequencing sequence sets; wherein the sample sequencing sequence sets correspond one-to-one to the amplicon sets, and the sample sequencing sequence sets and the amplicon sets having a corresponding relationship satisfy that the sample sequencing sequences in the sample sequencing sequence sets belong to the amplicons in the amplicon set.

[0044] The sample sequencing sequences that have been re-divided according to the above corresponding relationship no longer have the situation where the target mutation is located in the overlapping part of the amplification regions of two amplicons or the target mutation is located in the amplification region of one amplicon and at the same time in the primer region of another amplicon. Therefore, the false negative situation caused by the above two situations can be overcome.

[0045] Step S106: performing mutation detection on each of the plurality of sample sequencing sequence sets.

[0046] By using the method in the above embodiment, starting from the position of the amplicons, overlapping amplicons are separated and not placed in the same set. Then, the sequencing sequences are split based on the relationship between the amplicons and the sequencing sequences. Ultimately, sequences belonging to the same sequencing sequence set are split together, and sequences in different amplification regions do not overlap, thereby avoiding the occurrence of false negative results caused by primer region or primer dropout.

[0047] like Figure 2 As shown in the figure, a splitting method is illustrated. The topmost strip in the figure is the target amplification region, and four pairs of primers are designed to amplify this region. According to the three dimensions of chromosome number, the starting position of the forward primer in the amplicon, and the ending position of the reverse primer in the amplicon, the data of each dimension are sorted from large to small. The resulting amplicons are numbered 1, 2, 3, and 4. To achieve no overlap between amplicons, according to the simplest splitting logic, each amplicon can be treated as an amplicon set. Specifically, during the operation, a BED file is generated for each amplicon. Then, when the sample sequencing sequence is subsequently split, a sample sequencing sequence set is generated corresponding to each BED file.

[0048] In certain preferred embodiments, the division of the plurality of amplicons into a plurality of amplicon sets is performed by employing the following method to divide the amplicons so as to form a plurality of amplicon sets that do not overlap with each other and have the least number of amplicon sets in order to minimize the number of amplicon sets. The method specifically comprises the following steps:

[0049] Step S102-1, sorting the multiple amplicons according to the priorities of three dimensions: chromosome number, starting position of the forward primer in the amplicon, and ending position of the reverse primer in the amplicon, from large to small, and sorting the data of each dimension from small to large;

[0050] Step S102-2: traverse all the sorted amplicons and divide the traversed amplicons into the amplicon sets, and determine whether the current amplicon has interval overlap with the amplicon sets in which existing amplicons exist according to a fixed comparison order. If there is interval overlap, the current amplicon is grouped into the first amplicon set that has no interval overlap with the amplicon; otherwise, the current amplicon is grouped into an amplicon set in which no amplicon exists.

[0051] like Figure 3 As shown in the figure, the splitting method in the above preferred embodiment is schematically illustrated. Figure 2When splitting, the number of sets should be kept to a minimum to avoid the disadvantages of long analysis time and redundant and complicated processing procedures when the number of amplicons is large. In the specific operation, the BED file should be split according to the fact that there is no overlap between the amplification regions and the number of files after splitting is the least, such as Figure 3 As shown, these four amplicons are Figure 2 The same as in the number sequence, the adjacent two amplicons have partial overlap, that is, No. 1 overlaps with No. 2, No. 2 overlaps with No. 3, and No. 3 overlaps with No. 4. Therefore, according to Figure 3 The splitting method is to classify No. 1 and No. 3 into the same set, and No. 2 and No. 4 into the same set. In this way, the 4 amplicons are split into 2 files, Figure 2 The number of split method files in has been reduced by 50%.

[0052] The BED file formed after the splitting is used as a reference for splitting the sample sequencing sequence set in the subsequent step S104. Before the above step S104, the high-throughput sequencing data of the sample is processed by the following steps.

[0053] First, basic quality control is performed on the sample sequencing data. Data quality control includes ensuring that the data volume is qualified, the average sequencing depth meets the requirements, and the data quality meets Q20>90% and Q30>85%.

[0054] The meanings of Q20 and Q30 here are as follows: each base in the sequencing data has a corresponding quality value. If the quality value is Q20, the probability of misidentification is 1%, that is, the error rate is 1%, or the accuracy is 99%; if the quality value is Q30, the probability of misidentification is 0.1%, that is, the error rate is 0.1%, or the accuracy is 99.9%.

[0055] Then, use sentieon software (NGS gene data analysis acceleration software) to process the fastq file that has passed quality control according to the BWA module of the software, map the sequence in the fastq file with the reference genome, and obtain the alignment data BAM file. Then use the module in the sentieon software to sort the aligned BAM file and perform other subsequent processing to obtain the final BAM file.

[0056] Obtained BAM file, then according to S104 according to the BED file after splitting, BAM file is split, and splitting process mainly needs to judge that the sequencing sequence recorded in the BAM file belongs to the amplicon in which BED file.In theory, sequencing sequence is aligned to the starting position and the ending position of reference genome and is the 5 ' end coordinate of the forward primer and the 5 ' end coordinate of the reverse primer of primer amplification.Left end sequence is aligned to the starting position of reference genome and is consistent with the 5 ' end coordinate of the forward primer, and right end sequence is aligned to the ending position of reference genome and is consistent with the 5 ' end coordinate of the reverse primer, i.e., judges that this sequence belongs to this amplicon.But in actual amplification result, may cause alignment position to have deviation due to mispairing or other reasons, therefore, in some preferred embodiment, in above-mentioned step S104, satisfy one of following conditions and all think that the sample sequencing sequence in described sample sequencing sequence set belongs to the amplicon in described amplicon set.

[0057] Condition 1: The difference between the starting position of the left end sequence of the sample sequencing sequence and the starting position of the forward primer of the amplicon in the amplicon set, and the difference between the ending position of the right end sequence of the same sample sequencing sequence and the ending position of the reverse primer in the same amplicon in the amplicon set are both within the preset difference range.

[0058] Condition 2: Both the left-end sequence and the right-end sequence of the sample sequencing sequence are located within a certain amplicon region in the amplicon set, and the ratio of the overlapping region between the left-end sequence and the right-end sequence and the right-end sequence of the sample sequencing sequence to the total length of the certain amplicon is greater than a preset ratio.

[0059] like Figure 4 As shown, according to the above correspondence, the split BAM files are sorted and index files are created. The number of split BAM files is consistent with the number of split BED files.

[0060] like Figure 5 As shown, in some preferred embodiments, the specific process of step S106 includes three steps: mutation detection, mutation filtering and mutation merging.

[0061] First, mutation detection is performed by analyzing and detecting the BAM files after the obtained sequencing sequence set of each sample, that is, the split BAM files. This step can be performed using a variety of point mutation detection modules, including but not limited to GATK, commercial software Sentieon, freebayes, and other software. In this example, the DNAscope module in Sentieon software is used as an example to illustrate the point mutation detection process. The DNAscope module in Sentieon is used with the parameters resample_depth 100000, min_map_qual 15, and trim_soft_clip for point mutation detection. After each BAM file is analyzed by this module, a corresponding VCF file can be obtained to record the mutation information.

[0062] Mutation filtering is then performed. Because base errors can be introduced during the multiplex PCR amplification process, these erroneous sequences can be replicated and amplified during the subsequent PCR and sequencing processes, resulting in a certain proportion of false-positive sites in the results. Furthermore, during the analysis of multiplex PCR data, deduplication cannot be performed to avoid false positives caused by PCR errors. Deduplication during analysis involves removing only one sequence from the sequenced sequences that have been aligned to the same genomic location. Sequences belonging to the same amplicon in multiplex PCR sequencing data are consistently located.

[0063] Therefore, it is necessary to process each site detected in the VCF, select NA12878 and NA24385 for multiple sequencing to obtain data as standards to analyze the characteristics of false positive sites and set the original mutation sites obtained by the detection module. Analyze the total depth of the site in the BAM file, and count the number of bases with a quality greater than or equal to 20 and a quality less than 20 among the bases supporting the mutation based on the base quality of 20, as well as the original mutation ratio, and the number of bases with a base quality greater than or equal to 20 / total depth to obtain the filtered mutation threshold as a feature. Use the true set provided by the standard to compare and obtain the label of each site (true positive, false positive). After the site is statistically analyzed as above, the decision tree model is used to classify the results to select the final filtering features and thresholds.

[0064] Through the analysis of the standard results, when the number of high-quality bases supporting the mutation is >10 and the number of high-quality bases supporting the mutation / total depth is >0.1, false positive sites can be effectively filtered and true positive sites can be ensured not to be filtered.

[0065] Finally, each VCF after filtration is merged. Since multiple PCR has a tripping phenomenon, the tripping situation needs to be considered during merging. When a site is detected in multiple amplicons, for example, one amplicon detects 0 / 0 and another amplicon detects 0 / 1. Due to the tripping phenomenon, the 0 / 0 situation is due to a point mutation on the amplicon primer and is on the same chromosome as the target mutation, then a tripping phenomenon will occur, resulting in only amplifying the wild-type sequence. Similarly, if there is a point mutation on the primer and it is not on the same chromosome as the target mutation, it will result in only amplifying the mutant sequence, and the test result will be a homozygous mutation. Therefore, when merging, it is preferred to judge that the heterozygous mutation is a true mutation. When there is no heterozygous mutation in the test result, and the test results between the amplicons are 0 / 0 and 1 / 1, it is possible that the two amplicons have a tripping phenomenon at the same time, so the result will also be output as a heterozygous mutation. Therefore, in this embodiment, after filtering the results after each mutation detection to remove false positive sites, the multiple results are merged and the test results are obtained according to the following different situations:

[0066] When a locus is detected in only one amplicon, or the genotype is consistent in multiple amplicons, the test result is no change in genotype;

[0067] When a locus is detected in multiple amplicons and the genotypes are inconsistent, and a 0 / 1 heterozygous genotype appears, the test result is based on the 0 / 1 heterozygous genotype;

[0068] When a locus is detected in multiple amplicons and the genotypes are inconsistent, there is no heterozygous genotype, and at least one 0 / 0 wild type and one 1 / 1 homozygous type are detected, the test result is detected according to the 0 / 1 heterozygous genotype.

[0069] In the above description, 0 / 0 represents the wild type, which represents the genotype of a normal person; 0 / 1 represents the heterozygous mutant type, which means that a mutation occurs at this position on one chromosome and the site on the other chromosome is normal; 1 / 1 represents the homozygous mutant type, which means that a mutation occurs at this position on both chromosomes.

[0070] The above embodiments effectively avoid false negatives by detecting point mutations by splitting amplicons, and reduce false positive rates by filtering the results based on the quality of the mutated bases.

[0071] Another embodiment of the present application provides an electronic device, including a processor, a memory, and a computer program stored in the memory, wherein the computer program is configured to execute the method for detecting point mutations in multiplex PCR sequencing described in the above embodiment when executed by the processor.

[0072] Another embodiment of the present application provides a storage medium having a computer program stored thereon, wherein the computer program is used to execute a method for detecting point mutations based on multiplex PCR sequencing as described in another embodiment of the present application.

[0073] These computer programs can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes can be implemented by different modules corresponding to different steps. In an embodiment of the present application, the modules corresponding to the method for detecting point mutations in multiplex PCR sequencing constitute an apparatus for detecting point mutations in multiplex PCR sequencing, which includes the following modules:

[0074] The first division module is used to divide the multiple amplicons into multiple amplicon sets; wherein each amplicon set includes at least one amplicon, and there is no overlap between any two amplicons in the same amplicon set.

[0075] The second division module is used to divide the sample sequencing sequence into multiple sample sequencing sequence sets; wherein the sample sequencing sequence sets correspond to the amplicon sets one-to-one, and the sample sequencing sequence sets and the amplicon sets with a corresponding relationship satisfy that the sample sequencing sequences in the sample sequencing sequence sets belong to the amplicons in the amplicon set.

[0076] The mutation detection module is used to perform mutation detection on the multiple sample sequencing sequence sets respectively.

[0077] Through the cooperation between the above modules, the embodiment of the present application can reduce the detection of false negative points during multiple PCR point mutation detection, making the detection results more stable and reliable.

[0078] The above program can be executed in a processor or stored in a memory (or computer-readable medium), which includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0079] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for detecting point mutations in multiplex PCR sequencing, characterized in that: Divide multiple amplicons into multiple amplicon sets, and split them according to whether there is no overlap between the amplified regions and the number of files after splitting is the least, including: The plurality of amplicons are sorted from large to small according to the priorities of three dimensions, namely, chromosome number, the starting position of the forward primer in the amplicon, and the ending position of the reverse primer in the amplicon, and the data of each dimension are sorted from small to large; all the sorted amplicons are traversed and the traversed amplicons are divided into the amplicon sets, and whether there is an interval overlap between the current amplicon and the amplicon sets where existing amplicons exist is determined according to a fixed comparison order; if there is an interval overlap, the current amplicon is grouped into the first amplicon set that has no interval overlap with the amplicon; otherwise, the current amplicon is grouped into an amplicon set where no amplicon exists; wherein each amplicon set includes at least one amplicon, and there is no interval overlap between any two amplicons in the same amplicon set; The sample sequencing sequences are divided into a plurality of sample sequencing sequence sets; wherein the sample sequencing sequence sets correspond to the amplicon sets in a one-to-one manner, and the sample sequencing sequence sets and the amplicon sets having a corresponding relationship satisfy that the sample sequencing sequences in the sample sequencing sequence sets belong to the amplicons in the amplicon set, including: The difference between the starting position of the left end sequence of the sample sequencing sequence and the starting position of the forward primer of the amplicon in the amplicon set, and the difference between the ending position of the right end sequence of the same sample sequencing sequence and the ending position of the reverse primer in the same amplicon in the amplicon set are both within a preset difference range, or the left end sequence and the right end sequence of the sample sequencing sequence are both located in a certain amplicon region in the amplicon set and the proportion of the overlapping region between the left end sequence and the certain amplicon to the total length of the certain amplicon is greater than a preset proportion; Perform mutation detection on the plurality of sample sequencing sequence sets respectively; After performing mutation detection on the plurality of sample sequencing sequence sets respectively, the method includes: The results of each mutation test are filtered to remove false positive sites; Combine multiple results and obtain test results according to the following different situations: When a locus is detected in only one amplicon, or the genotype is consistent in multiple amplicons, the test result is no change in genotype; When a locus is detected in multiple amplicons and the genotypes are inconsistent, and a 0 / 1 heterozygous genotype appears, the test result is based on the 0 / 1 heterozygous genotype; When a locus is detected in multiple amplicons and the genotypes are inconsistent, there is no heterozygous genotype, and at least one 0 / 0 wild type and one 1 / 1 homozygous type are detected, the test result is detected according to the 0 / 1 heterozygous genotype.

2. The method according to claim 1, characterized in that The amplicon set is a BED file.

3. The method according to claim 1, characterized in that The sample sequencing sequence set is a BAM file.

4. An electronic device, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory, wherein the computer program is configured to execute the method according to any one of claims 1 to 3 when executed by the processor.

5. A storage medium, characterized in that: A computer program is stored thereon, and the computer program is used to execute the method described in any one of claims 1 to 3.