Method and apparatus for determining true-positive mutation and false-positive mutation using start-end pair

By forming start-end pair groups and verifying mutation results based on specific characteristics, the method effectively distinguishes true and false positives in genetic sequencing, addressing the limitations of UMI-based methods and enhancing detection sensitivity.

WO2025249622A1PCT designated stage Publication Date: 2025-12-04DXOME CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/007691
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-27
Filing Date
2024-06-05
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing methods for identifying genetic mutations in sequencing processes, such as those using Unique Molecular Identifiers (UMIs), incur additional costs and complexity, and fail to reliably distinguish between true-positive and false-positive mutations, leading to low detection sensitivity in liquid biopsies.

Method used

A method and device that form multiple start-end pair groups, compare these groups to determine true-positive or false-positive mutations, and verify the results based on characteristics like read count, sense and antisense matches, base quality, mapping quality, and homopolymer length, without using UMIs, thereby reducing time and cost.

Benefits of technology

This approach enhances the reliability of mutation identification by accurately distinguishing between true and false positives, reducing the time and cost associated with UMI-based methods, and improving detection sensitivity in genetic sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024007691_04122025_PF_FP_ABST
    Figure KR2024007691_04122025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a method for determining a genetic mutation as a true-positive mutation or a false-positive mutation, the method comprising: (1) a step for forming a plurality of start-end pair groups; (2) a step for determining the genetic mutation as a true-positive mutation candidate or a false-positive mutation through a comparison of the plurality of start-end pair groups; and (3) a step for determining whether the genetic mutation is a true-positive mutation by verifying the result of determining the true-positive mutation candidate on the basis of the characteristics of the start-end pair groups.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for determining true positive and false positive mutations using a sequence combination

[0001] The present invention relates to a method and device for identifying true and false mutations using a sequence combination. More specifically, the present invention proposes a method for identifying mutations occurring during sequencing without using artificial sequences such as Unique Molecular Identifiers (UMIs), thereby enhancing the reliability of the identification results.

[0002] The general definition of genetic testing is the analysis of DNA, RNA, chromosomes, and metabolites to detect changes in the genome associated with hereditary diseases. Genetic testing can be performed by directly analyzing the DNA and RNA that make up genes (direct testing), indirectly examining the genotype inherited along with the disease gene (linkage analysis), or through biochemical testing or cytogenetic testing that analyzes metabolites. In a narrow sense, genetic testing is often used to refer to DNA testing, but since it analyzes at the molecular level of DNA, the more accurate term is molecular genetic testing.

[0003] Typically, a detection sensitivity of 0.01% to 1% is required to analyze ctDNA from cancer patients. Many liquid biopsy companies are conducting research and development to improve detection sensitivity, but actual detection sensitivity remains low due to false-positive signals generated during the biopsy process. Therefore, technologies capable of distinguishing false-positive signals are needed to improve detection sensitivity.

[0004] Meanwhile, next-generation sequencing (NGS) using UMIs attaches artificial sequences to both ends of the sample DNA, groups them together, and analyzes them by UMI. This method assumes that the same UMI group originated from the same DNA, performs repeated sequencing, and considers a true positive if all reads contain the same mutation. This method requires additional UMI sequences to be artificially synthesized, and additional UMI sequences to be read during sequencing, which incurs additional costs, and has the disadvantage of making the algorithm for UMI sequence analysis more complex.

[0005] To solve the above shortcomings, the applicant filed Republic of Korea Application No. 10-2022-0054570, but the invention presented a theoretical model that can distinguish between true-positive and false-positive mutations without using UMI sequences, but there was a problem that the reliability of the results of determining true-positive mutations could not be guaranteed.

[0006] The challenge that our center seeks to address includes reducing the time and cost incurred due to UMIs, such as having to separately synthesize UMI sequences and additionally read UMI sequences, by using UMIs in existing liquid biopsy tests using next-generation sequencing.

[0007] The above tasks are merely examples, and there may be additional tasks that a person skilled in the art can understand within the scope of this application.

[0008] The first aspect of the present invention relates to a method for determining a genetic mutation as a true positive or false positive mutation.

[0009] (1) Step of forming multiple start-end pair groups

[0010] (2) A step of determining the genetic mutation as a true positive mutation candidate or a false positive mutation through comparison between the above multiple seed combination groups, and

[0011] (3) A method is provided including a step of verifying the result of determining the positive mutation candidate based on the characteristics of the above-mentioned seed combination group to determine whether or not a positive mutation is present.

[0012] The second aspect of the present invention provides a device for determining a genetic mutation as a true positive mutation or a false positive mutation, comprising: a forming unit for forming a plurality of start-end pair groups; a determining unit for determining the genetic mutation as a true positive mutation candidate or a false positive mutation through comparison between the plurality of start-end pair groups; and a verification unit for verifying the true positive mutation candidate determination result based on the characteristics of the start-end pair groups to determine whether the genetic mutation is a true positive mutation.

[0013] The above means are only examples, and there may be additional means of solving problems that can be understood by a person skilled in the art within the scope of the present invention.

[0014] The effect according to the present invention includes not only effectively reducing the time and cost incurred due to UMI in the sequencing process, but also increasing the reliability of the results of determining true-positive or false-positive mutations.

[0015] The above effects are merely examples, and additional effects may exist that can be understood by a person skilled in the art within the scope of the present invention.

[0016] Figure 1 is a flowchart showing a method according to the present invention.

[0017] Figure 2 is a block diagram showing a device according to the present invention.

[0018] Figure 3 is an example of score calculation for the 'number of leads', which is a characteristic of the combination of start and end points according to the present invention.

[0019] Figure 4a shows the results of identifying true positive mutations according to the method of the present invention for a library (WT) produced using the applicant's product, the Pan100 panel.

[0020] Figure 4b shows the results of identifying true positive mutations according to the method of the present invention for libraries (AF1, AF2) produced using the applicant's product, the Pan100 panel.

[0021] Figure 4c is a diagram summarizing the results of the identification of a library produced using the Pan100 panel, a product of the applicant.

[0022] Figure 5a shows the results of determining true positive mutations according to the method of the present invention for a library (WT) produced using the TMB500 panel, a product of the present applicant.

[0023] Figure 5b shows the results of identifying true positive mutations according to the method of the present invention for libraries (AF1, AF2) produced using the applicant's product, the TMB500 panel.

[0024] Figure 5c is a diagram summarizing the results of the identification of a library produced using the TMB500 panel, a product of the applicant.

[0025] Below, with reference to the attached drawings, embodiments of the present invention are described in detail to facilitate easy implementation by those skilled in the art. However, the present invention can be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, irrelevant parts have been omitted for clarity, and similar reference numerals have been used throughout the specification to indicate similar elements.

[0026] Throughout this specification, whenever a part is said to 'include' a component, this does not mean that it excludes other components, but rather that it may include other components, unless otherwise specifically stated.

[0027] Throughout this specification, the terms “step of” or “step of” do not mean “step for”.

[0028] Throughout this specification, the term 'combination(s) thereof' included in the expressions in the Makushi format means one or more mixtures or combinations selected from the group consisting of the components described in the expressions in the Makushi format, and means including one or more selected from the group consisting of said components.

[0029] Throughout this specification, references to 'A and / or B' mean 'A or B, or A and B.'

[0030] Throughout this specification, "ctDNA" refers to circulating tumor DNA, primarily fragmented DNA derived from tumors that circulates in the bloodstream. ctDNA isolated from plasma provides a specific and sensitive biomarker that contains a variety of information, including diagnosis, early detection of tumors, prognosis, and disease-free survival. It can also be used to predict cancer treatment resistance and likelihood.

[0031] Throughout this specification, "genetic mutation" means a change in genetic information due to a change in the base sequence that constitutes a gene, resulting in a change in hereditary trait. In the context of this specification, "genetic mutation" may refer to a mutation in a gene sequence. During gene sequencing processes such as NGS, mutations in gene sequences can occur for various reasons. In such cases, it becomes difficult to determine whether the mutation was present in the gene prior to sequencing or whether it was newly introduced during the sequencing process.

[0032] Throughout this specification, 'false positive mutation' refers to a mutation that occurs due to various causes during the gene sequencing process, which may be a PCR error.

[0033] Throughout this specification, the term "true positive mutation" refers to a mutation that existed in a gene before sequencing, i.e., a mutation that existed originally in the gene. Furthermore, the term "true positive mutation candidate" refers to a mutation that has not yet been confirmed, but is likely to be a true positive mutation.

[0034] Throughout this specification, a "start-end pair" refers to a combination of genetic reads that have the same starting and ending points when mapped to a reference sequence. For example, when multiple DNA fragments (genes) are fragmented during the sequencing process, the starting and ending points are theoretically random. A start-end pair utilizes the extremely low probability that any two fragmented genes will have the same starting and ending points.

[0035] Throughout this specification, "base quality" refers to information indicating whether a specific base in a read produces a sufficiently accurate signal on the sequencing equipment. Base quality typically has a value between 0 and 64.

[0036] Throughout this specification, 'mapping quality' is a numerical value indicating the degree to which a sequencing read is uniquely mapped to a specific location in the step of mapping the sequencing read to the standard sequence, and is generally expressed as a value between 0 and 60.

[0037] Throughout this specification, 'repeat (homopolymer)' refers to a region where a specific base is repeated, and the greater the number of repeated bases, the more likely it is that errors will occur in sequencing equipment or in mapping algorithms.

[0038] The first aspect of the present invention relates to a method for determining a genetic mutation as a true positive or false positive mutation.

[0039] (1) Step of forming multiple start-end pair groups

[0040] (2) A step of determining the genetic mutation as a true positive mutation candidate or a false positive mutation through comparison between the above multiple seed combination groups, and

[0041] (3) A method is provided including a step of verifying the result of determining whether a true positive mutation is present by determining whether the true positive mutation is present based on the characteristics of the above-mentioned seed combination group (Fig. 1).

[0042] In one specific example of the present invention, a group of seed combinations means a group of seed combinations classified by seed combinations having the same starting point and ending point.

[0043] In one specific example of the present invention, a true positive mutation candidate may be determined based on whether the mutation is present in all the seed combinations within two or more seed combination groups. In another specific example of the present invention, a false positive mutation may be determined based on whether the mutation is present in some seed combinations within one or more seed combination groups.

[0044] In one specific example of the present invention, the characteristics of the sequence combination group may include at least one of the number of reads, sense and antisense matches, base quality, mapping quality, and number of repeats (homopolymers).

[0045] In one specific example of this application, the verification of the positive mutation candidate identification result may be performed by summing the scores for each of the multiple target groups, calculated using a scoring method based on the importance of each feature. Alternatively, a weighted scoring method may be used for verification, and other similar methods may be used according to the user's convenience.

[0046] The second aspect of the present invention provides a device for determining a genetic mutation as a true positive mutation or a false positive mutation, comprising a forming unit (10) for forming a plurality of start-end pair groups, a determining unit (20) for determining the genetic mutation as a true positive mutation candidate or a false positive mutation through comparison between the plurality of start-end pair groups, and a verification unit (30) for verifying the true positive mutation candidate determination result based on the characteristics of the start-end pair groups to determine whether the genetic mutation is a true positive mutation (Fig. 2).

[0047] In one specific example of the present invention, the determination unit (20) may determine the genetic mutation as a true positive mutation candidate based on whether the mutation exists in all the seed combinations within two or more seed combination groups.

[0048] In one specific example of the present invention, the determination unit (20) may determine the genetic mutation as a false positive mutation based on whether the mutation exists in some of the sequence combinations within one or more sequence combination groups or whether the mutation exists in all leads within one sequence combination group.

[0049] In one specific example of the present invention, the characteristics of the sequence combination group may include at least one of the number of reads, sense and antisense matches, base quality, mapping quality, and number of repeats (homopolymers).

[0050] In one specific example of the present invention, the verification unit (30) may verify the mutation verification result by the sum of scores for each of the plurality of test groups calculated through a point addition or point subtraction method according to the importance of each feature.

[0051] The second aspect of the invention shares the same technical features as the first aspect.

[0052] Here, we explain in more detail the criteria for determining false-positive mutations in this study. For example, assuming that no replication errors occurred during the gene replication process, all sequences within a given sequence group should have the same sequence. Therefore, if sequences with and without mutations exist within a given sequence group, this can be considered a false-positive mutation resulting from an error during replication. In other words, mutations present in some sequences within one or more sequence groups can be considered false-positive mutations.

[0053] Meanwhile, a mutation present in all pairs of siblings within a group of two or more siblings could be either a true positive mutation or a false positive mutation. Assuming an extreme scenario, if only one sibling pair exists within each sibling group, a mutation present in all pairs of siblings cannot be confidently considered a true positive mutation.

[0054] Therefore, the present invention provides a method for determining the above mutation as a true positive mutation candidate, and then verifying whether the result of determining the above true positive mutation candidate is actually likely to be a true positive mutation by considering the characteristics of each seed combination group.

[0055] In addition, this institute aims to provide an additional method to quantify the probability of a true-positive mutation by calculating a score for a candidate true-positive mutation using a weighted matrix based on the characteristics of each seed combination group, either by adding or subtracting points. It is expected that introducing a method to determine a mutation satisfying a specific score as a true-positive mutation using the above scores will significantly increase the reliability of the discrimination results. Since the probability of discrimination error may vary depending on the equipment or experimental environment used for sequencing, the importance of each feature may not be uniformly determined according to the criteria described in this specification. Most preferably, it can be determined through machine learning using multiple discrimination data, but is not necessarily limited thereto.

[0056] (1) Number of reads (Fig. 3)

[0057] As explained in the extreme example above, if the number of sequences (i.e., leads) contained in a particular sequence is significantly small, the credibility of the results determined based on criteria such as "mutations present in all sequences within the sequence group" will be reduced. Conversely, as the number of leads increases, the credibility of the results increases.

[0058] Therefore, the weighted matrix (W) according to the number of reads (ReadCount(x)) r ) can be configured to calculate the points according to the number of leads of a specific combination of seeds.

[0059] According to one embodiment of the present invention, a Read count weighted matrix (W r ) is a series of vectors that are progressively increasing positive numbers, and when the number of reads is n, the nth weight value (W r [n]) can be substituted. This weight can be linear, but non-linear is generally more appropriate.

[0060] (2) Whether sense and antisense match

[0061] A sequence is composed of genetic reads, which are composed of a sense strand and an antisense strand. Normally, the sense and antisense strands bind complementarily. However, if the sense and antisense strands fail to bind complementarily, this suggests the possibility of a false-positive mutation. Therefore, the presence of a sense and antisense match can be used to verify the results of a true-positive mutation, as follows.

[0062] 1) If all reads have mutations and sense and antisense match (x): A score (value for concordant sense and antisense; Vc) is given, which can be any natural number belonging to the positive range.

[0063] 2) If there is a mutation in only one strand of sense or antisense, resulting in sense and antisense discordance (y): A negative score is given (value for discordant sense and antisense; Vd), which includes a method of making the score negative.

[0064] 3) If only one strand, either sense or antisense, is sequenced and sense and antisense cannot be matched (z): A score (value for single-strand-only groups; Vs) is assigned, but not higher than the score in 1) (Vs≤Vc). The score can be any natural number within the positive range.

[0065] (3) Base quality

[0066] Base quality is information that records whether a specific base in a read produced a sufficiently accurate signal on the sequencing equipment. Base quality typically ranges from 0 to 64, with higher values ​​indicating greater replication accuracy and a lower likelihood of replication errors. Conversely, lower values ​​indicate a higher likelihood of replication errors.

[0067] Therefore, base quality is also a factor in determining whether a true mutation is present or not, in the base quality weight matrix (W b ) can be set and used for the point addition or deduction method according to the present invention.

[0068] Points added or subtracted by weight can be set to positive or negative, or they can be designed to be multiplied by a value between 0 and 1, so that points are subtracted when the value is closer to 0 and full points are given when the value is closer to 1. This weight matrix can be linear, but a non-linear one is generally more appropriate.

[0069] (4) Mapping quality

[0070] Mapping quality is a numerical value that indicates the degree to which a sequencing read is uniquely mapped to a specific location in the step of mapping the sequencing read to the standard sequence. It is usually expressed as a value between 0 and 60, and the higher the value, the higher the accuracy of replication.

[0071] Conversely, a low score means that similar sequences to the corresponding read exist in multiple places in the standard sequence, and false positives may occur due to inaccurate mapping, which may result in a mutation being read as present when it does not exist.

[0072] Therefore, base quality is also a factor in determining whether a true mutation is present or not in the mapping quality weight matrix (W m ) can be set and used for the point addition or deduction method according to the present invention.

[0073] Points added or subtracted by weight can be set to positive or negative, or they can be designed to be multiplied by a value between 0 and 1, so that points are subtracted when the value is closer to 0 and full points are given when the value is closer to 1. This weight matrix can be linear, but a non-linear one is generally more appropriate.

[0074] (5) Number of repeats (homopolymer)

[0075] A repeat (homopolymer) is a region where a specific base is repeated. The more repeating bases there are, the more likely it is that errors will occur in sequencing equipment or in mapping algorithms.

[0076] Therefore, the base quality is also a factor in determining whether a true mutation is present or not, depending on the number of repetitions, and the homopolymer length weight matrix (W h ) can be set and used for the point addition or deduction method according to the present invention.

[0077] Points added or subtracted by weight can be set to positive or negative, or they can be designed to be multiplied by a value between 0 and 1, so that points are subtracted when the value is closer to 0 and full points are given when the value is closer to 1. This weight matrix can be linear, but a non-linear one is generally more appropriate.

[0078] By selecting the above features according to the situation, a true positive mutation score calculation formula can be devised as shown below. The formula below is merely an example and does not limit the scope of this application.

[0079] [Example 1]

[0080]

[0081] [Example 2]

[0082]

[0083] Hereinafter, implementation examples and embodiments of the present invention will be described in detail with reference to the attached drawings. However, the present invention may not be limited to these implementation examples and embodiments and drawings.

[0084] 1. Standard materials used for verification

[0085] Samples with 40 clinically relevant mutations in 28 genes were tested using Seracare's standard material, which had been validated by digital PCR.

[0086] 1) Seraseq® ctDNA Mutation Mix v2 WT (#0710-0144, 10ng / ㎕)

[0087] 2) Seraseq® ctDNA Mutation Mix v2 AF1% (#0710-0140, 10ng / ㎕)

[0088] 3) Seraseq® ctDNA Mutation Mix v2 AF2% (#0710-0139, 10ng / ㎕)

[0089] 2. Creating a library

[0090] The library was created using the following two products of the applicant.

[0091] 1) Pan100 panel

[0092] Use of the applicant's Probe Mix (Pan100) (#PM5901096), Hybridization Buffer (#HR01096), Wash Buffer (#WB01096), and Cleanup Bead (#CB01096)

[0093] 2) TMB500 panel

[0094] Use of the applicant's Probe Mix (TMB500) (#PM7101096), Hybridization Buffer (#HR01096), Wash Buffer (#WB01096), and Cleanup Bead (#CB01096)

[0095] 3. Verification process

[0096] When the identification method according to this invention was applied to the produced library, it was verified whether the true positive mutation identification result was consistent with the mutation rate of the already known standard material (WT-wild type, AF1, AF2%).

[0097] In this embodiment, verification was performed using the following formula, but the formula below does not limit the scope of the present application.

[0098]

[0099] The scoring method for each characteristic of the group of seed combinations is as follows.

[0100] 1) Number of reads

[0101] Readcount number W r 10.720.930.9840.9950.999>61

[0102] For example, if the Readcount number is 1, W r [Readcount(x i or y i or z k )] = 0.7. 2) Whether sense and antisense match Vc = 3 (point coefficient when matched)

[0103] V d = 1 (point deduction coefficient for mismatch)

[0104] V s = 1 (points added when only one direction of sense / antisense is sequenced and match cannot be determined)

[0105] For example, in the above equation, for a group of sequences where sense and antisense match, V c = Calculate as 3.

[0106] 3) Base quality

[0107] Base QualityW b00.015-100.0211-150.0516-200.121-250.526-300.8>311

[0108] For example, if Base Quality is 0, W b [BaseQuality(x i or y i or z k )] = 0.01.

[0109] 4) Mapping quality

[0110] Mapping QualityW m 00.011-100.0211-200.0521-300.131-400.241-500.551-601

[0111] For example, if Mapping Quality is 0, W m [MappingQuality(x i or y i or z k )] = 0.01.

[0112] 5) Number of repeats (homopolymer)

[0113] Number of homopolymers W h 112131415160.9970.9880.9590.9100.7110.5120.2130.1140.08150.07160.06170.05180.04190.03200.01

[0114] For example, if the number of homopolymers is 1, W h [HomopolymerLength(x i or y i or z k )] = 1.

[0115] 4. Verification Results

[0116] Gene and mutation sequence corresponding to each number Gene list Mutation list 1 AKT 1 c. 49G > A2 APC c. 4666_4667 ins A3 APC c. 4348 C > T 4 ATM c. 1058_1059 ins GT 5 BRAF c. 1799 T > A6 CTNN B 1 c. 121A > G7 EGFR c. 2236_2250 ins GT 158 EGFR c. 2310_2311 ins GT 9 EGFR c.2573T>G10EGFRc.2369C>T11ERBB2c.2324_2325ins1212FGFR3c.746C>G13FLT3c.2 503G>T14GNA11c.626A>T15GNAQc.626A>C16GNASc.601C>T17IDH1c.394C>T18JAK2c.1 849G>T19KITc.2447A>T20KRASc.35G>A21NRASc.182A>G22PDGFRAc.1694_1695insA2 3PDGFRAc.2525A>T24PIK3CAc.3204_3205insA25PIK3CAc.1633G>A26PIK3CAc.3140A> G27PTENc.800delA28PTENc.741_742insA29RETc.2753T>C30SMAD4c.1394_1395insT 31TP53c.723delC32TP53c.263delC33TP53c.524G>A34TP53c.818G>A35TP53c.743G>A

[0117] WT: wild type (no mutation)

[0118] AF2: allele frequency 2%

[0119] AF1: allele frequency 1%

[0120] 1) Pan100 (Fig. 4a, Fig. 4b, Fig. 4c)

[0121] Figures 4a to 4c show the validation results (true-positive mutation candidate scores) for each library. Table 6 summarizes the scores above.

[0122] Referring to each drawing and table, it can be seen that the scores are highest in the order of WT, AF1, and AF2. Higher scores indicate a higher likelihood that the true-positive mutation candidates discovered in each seed combination group are actually true-positive mutations, confirming that the method described herein can accurately identify true-positive mutations.

[0123] Verification results for Pan100Pan100WTAF1AF2Variant scoreAverage2.5133.8583.58Variant scoreMAX5.8071.88152.23Variant scoreMIN2.508.8024.30

[0124] 2) TMB500 (Fig. 5a, Fig. 5b, Fig. 5c)

[0125] Figures 5a to 5c show the validation results (true-positive mutation candidate scores) for each library. Table 7 summarizes the scores. Referring to each figure and table, it can be seen that the scores are high in the order of WT, AF1, and AF2. A high score indicates that the true-positive mutation candidate discovered in each seed combination group is likely to be an actual true-positive mutation, confirming that the method according to the present invention can accurately identify true-positive mutations.

[0126] Verification results for Pan100TMB500WTAF1AF2Variant score Average2.5134.6284.91Variant score MAX5.9675.48158.17Variant score MIN2.509.0021.12

[0127] [Explanation of symbols]

[0128] 1: Discriminant device

[0129] 10: Formation

[0130] 20: Discrimination section

[0131] 30: Verification Department

Claims

1. In a method for determining a genetic mutation as a true positive mutation or a false positive mutation, (1) A step of forming a group of multiple start-end pairs; (2) a step of determining the genetic mutation as a true positive mutation candidate or a false positive mutation through comparison between the plurality of seed combination groups; and (3) A method including a step of verifying the result of determining the positive mutation candidate based on the characteristics of the above-mentioned seed combination group to determine whether or not it is a positive mutation.

2. In paragraph 1, A method wherein the above-mentioned positive mutation candidate is determined based on whether the mutation exists in all the seed combinations within two or more seed combination groups.

3. In paragraph 1, A method in which the above false positive mutation is determined based on whether the mutation exists in some of the seed combinations within one or more seed combination groups.

4. In paragraph 1, The characteristics of the above-mentioned group of seedlings are: A method comprising at least one of the number of reads, sense and antisense matches, base quality, mapping quality, and number of repeats (homopolymers).

5. In paragraph 4, A method in which the verification of the above-mentioned positive mutation candidate determination result is verified by the sum of the scores for each of the multiple seed groups calculated through a point addition or point subtraction method according to the importance of the characteristics of each seed combination group.

6. In a device for determining a genetic mutation as a true positive mutation or a false positive mutation, A forming unit that forms a group of multiple start-end pairs; A determination unit that determines the genetic mutation as a true positive mutation candidate or a false positive mutation through comparison between the plurality of seed combination groups; and A device including a verification unit that verifies the positive mutation candidate identification result based on the characteristics of the above-mentioned seed combination group to determine whether or not a positive mutation is present.

7. In paragraph 6, A device wherein the above-mentioned determination unit determines the genetic mutation as a true positive mutation candidate based on whether the mutation exists in all the seed combinations within two or more seed combination groups.

8. In paragraph 6, A device in which the above-mentioned determination unit determines whether the false positive mutation is a mutation existing in some of the seed combinations within one or more seed combination groups.

9. In paragraph 6, The characteristics of the above-mentioned group of seedlings are: A device comprising at least one of the number of reads, sense and antisense matches, base quality, mapping quality, and number of repeats (homopolymers).

10. In paragraph 9, The above verification unit verifies the result of determining the positive mutation candidate by the sum of the scores for each of the plurality of test groups calculated through a point addition or point subtraction method according to the importance of each feature.

Citation Information

Patent Citations

  • Roll pressing apparatus for manufacturing electrode

    KR1020240068574A

  • Manufacturing method of functional liquid manure

    KR102376447B1

  • Sensor module for electrostatic charge measurement and system for electrostatic charge monitoring using the same

    KR102504492B1