A method and apparatus for verifying a target site
Patent Information
- Application Number
- CN202211429378.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-15
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2042-11-15
AI Technical Summary
当前能够读取一代测序结果并显示峰图的软件有Bioedit、Chromas、SnapGene、Geno me compiler等,但是通过手动方式在测序结果中寻找目标位点的方法比较困难,且耗费精力;同时,在某些情况下,若待验证目标位点的附近也有少量的突变存在,则加大了定位目标位点的难度
[0027] According to the above embodiments, a method and apparatus for verifying target sites utilize multiple seed sequences of various lengths to match the sequencing results, significantly increasing the probability of locating the true target site.
Smart Images

Figure CN115910201B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, and more specifically to a method and apparatus for verifying target sites. Background Technology
[0002] Next-generation sequencing (NGS) is currently the most widely used sequencing protocol, offering advantages such as high throughput and fast sequencing speed, meeting the demands of large-scale gene sequencing in current scientific research and medical fields. Analysis of NGS data often yields rich information for research or disease diagnosis. Variance detection is a crucial step in analyzing sequencing results, providing information on all mutations occurring in the sequenced sample. After obtaining mutation information, for certain mutations closely related to diseases and recorded in databases, NGS can be used for secondary validation.
[0003] First-generation sequencing is characterized by low throughput, long read lengths, and high accuracy, making it ideal for validating target sites. After sequencing the target region containing the target site, the presence of mutations or other abnormalities at the target site can be determined from the sequencing results. Currently, software capable of reading first-generation sequencing results and displaying peak plots includes Bioedit, Chromas, SnapGene, and Genome Compiler. However, manually locating target sites in sequencing results is difficult and time-consuming. Furthermore, in some cases, the presence of a small number of mutations near the target site further complicates the localization of the target site. Summary of the Invention
[0004] According to a first aspect, in one embodiment, a method for verifying a target site is provided, comprising:
[0005] The reference genome sequence acquisition step includes obtaining a sequence containing the base positions corresponding to the target site from the reference genome based on the position of the target site to be verified in the reference genome;
[0006] The reference genome seed sequence extraction step includes extracting the first N sets of seed sequences from the reference genome that are closest to the base position corresponding to the target site to be verified in the reference genome, according to the preset seed sequence length;
[0007] The seed sequence finding step of the sequencing results includes comparing the N sets of seed sequences extracted from the reference genome with the first-generation sequencing result sequences of the sample to be tested, and searching for seed sequences in the first-generation sequencing result sequences; if a seed sequence is found in the first-generation sequencing result sequences, it is determined that the match is successful; if none of the seed sequences appear in the first-generation sequencing result sequences, it is determined that the match is unsuccessful.
[0008] The temporary target site determination step includes obtaining the successfully matched seed sequence from the first-generation sequencing results, and predicting the position of the target site in the first-generation sequencing results by the relative relationship between the seed sequence in the reference genome and the target site to be verified, which is the temporary target site.
[0009] In one embodiment, the method further includes a double sequence alignment step, which includes extracting a sequence containing the temporary target site and covering M bases upstream and downstream of it from the first-generation sequencing result sequence according to the position of the temporary target site in the first-generation sequencing result sequence, and performing double sequence alignment with the sequence extracted from the reference genome in the reference genome sequence acquisition step.
[0010] The determination step includes judging the alignment status and outputting the matching result based on the double sequence alignment results.
[0011] In one embodiment, the method further includes a location determination step, which includes determining the location of the target site in the first-generation sequencing result sequence based on the results of double sequence alignments that are determined to be accurately locatable in the determination step.
[0012] In one embodiment, a plotting step is also included, in which a result map is plotted based on the mutation sites identified in the location determination step and the first-generation sequencing result sequence.
[0013] According to a second aspect, in one embodiment, an apparatus for verifying a target site is provided, comprising:
[0014] The reference genome sequence acquisition module is used to acquire a sequence containing the base positions corresponding to the target site from the reference genome based on the position of the target site to be verified in the reference genome;
[0015] The reference genome seed sequence extraction module is used to extract the top N seed sequences from the reference genome that are closest to the base position corresponding to the target site to be verified in the reference genome, according to a preset seed sequence length.
[0016] The sequencing result seed sequence finding module is used to compare the N sets of seed sequences extracted from the reference genome with the first-generation sequencing result sequences of the sample to be tested, and to search for seed sequences in the first-generation sequencing result sequences; if a seed sequence is found in the first-generation sequencing result sequences, it is determined that the match is successful; if none of the seed sequences appear in the first-generation sequencing result sequences, it is determined that the match is unsuccessful.
[0017] The temporary target site determination module is used to obtain the successfully matched seed sequence in the first-generation sequencing result sequence. By using the relative relationship between the seed sequence in the reference genome and the target site to be verified, the position of the target site in the first-generation sequencing result sequence is predicted, which is the temporary target site.
[0018] In one embodiment, a double sequence alignment module is further included, which is used to extract a sequence containing the temporary target site and covering M bases upstream and downstream of it from the first-generation sequencing result sequence according to the position of the temporary target site in the first-generation sequencing result sequence, and to perform double sequence alignment with the sequence extracted from the reference genome in the reference genome sequence acquisition step.
[0019] The judgment module is used to determine the alignment status and output the matching result based on the double sequence alignment results.
[0020] In one embodiment, a location determination module is also included, which is used to determine the location of the target site in the first-generation sequencing result sequence based on the results of double sequence alignments that are determined to be accurately locatable in the determination step.
[0021] In one embodiment, in the location determination module, the base position in the first-generation sequencing result sequence that corresponds to the position of the target site to be verified in the reference genome sequence in the double sequence alignment result is the target site position located by this module.
[0022] In one embodiment, a plotting module is also included, which is used to plot the results based on the target site location determined by the location determination module and the first-generation sequencing result sequence.
[0023] According to a third aspect, in one embodiment, an apparatus for verifying a target site is provided, comprising:
[0024] Memory, used to store programs;
[0025] A processor for implementing a method as described in any of the first aspects by executing a program stored in the memory.
[0026] According to a fourth aspect, in one embodiment, a computer-readable storage medium is provided, the medium storing a program that can be executed by a processor to implement the method as described in any of the first aspects.
[0027] According to the above embodiments, a method and apparatus for verifying target sites utilize multiple seed sequences of various lengths to match the sequencing results, significantly increasing the probability of locating the true target site.
[0028] In one embodiment, using a double sequence alignment method to locate the target site is more accurate than manual methods or methods that infer the location of the target site by a fixed distance from the target site.
[0029] In one embodiment, a threshold is set for the double sequence alignment matching results. If the threshold is not reached, the image cannot be generated. This allows for early determination of poor sequencing results without having to perform subsequent steps, thus avoiding wasting unnecessary effort.
[0030] In one embodiment, the present invention can automatically plot the sequencing peak diagram of the sequencing results containing the target site (e.g., the mutation to be verified) and a sequence upstream and downstream of it, and label the target site, eliminating the trouble of manually finding and labeling the target site.
[0031] In one embodiment, automatic mapping is essential and meaningful for the validation of large volumes of target site (e.g., mutation) data. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of the process for plotting peaks of mutation sites in one embodiment of a first-generation sequencing algorithm.
[0033] Figure 2 This is the result of a two-sequence alignment in one embodiment. Detailed Implementation
[0034] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings. In the following embodiments, many details are described to facilitate a better understanding of the present application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other materials or methods. In some cases, certain operations related to the present application are not shown or described in the specification. This is to avoid obscuring the core parts of the present application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.
[0035] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.
[0036] The serial numbers assigned to components in this article, such as "first" and "second", are used only to distinguish the objects being described and have no sequential or technical meaning.
[0037] According to a first aspect, in one embodiment, a method for verifying a target site is provided, comprising:
[0038] The reference genome sequence acquisition step includes obtaining a sequence containing the base positions corresponding to the target site from the reference genome based on the position of the target site to be verified in the reference genome;
[0039] The reference genome seed sequence extraction step includes extracting the first N sets of seed sequences from the reference genome that are closest to the base position corresponding to the target site to be verified in the reference genome, according to the preset seed sequence length;
[0040] The seed sequence finding step of the sequencing results includes comparing the N sets of seed sequences extracted from the reference genome with the first-generation sequencing result sequences of the sample to be tested, and searching for seed sequences in the first-generation sequencing result sequences; if a seed sequence is found in the first-generation sequencing result sequences, it is determined that the match is successful; if none of the seed sequences appear in the first-generation sequencing result sequences, it is determined that the match is unsuccessful.
[0041] The temporary target site determination step includes obtaining the successfully matched seed sequence from the first-generation sequencing results, and predicting the position of the target site in the first-generation sequencing results by the relative relationship between the seed sequence in the reference genome and the target site to be verified, which is the temporary target site.
[0042] In one embodiment, the method further includes a double sequence alignment step, which includes extracting a sequence containing the temporary target site and covering M bases upstream and downstream of it from the first-generation sequencing result sequence according to the position of the temporary target site in the first-generation sequencing result sequence, and performing double sequence alignment with the sequence extracted from the reference genome in the reference genome sequence acquisition step.
[0043] The determination step includes judging the alignment status and outputting the matching result based on the double sequence alignment results.
[0044] In one embodiment, during the reference genome sequence acquisition step, each set of seed sequences has two sequences, one of which is located upstream of the target site and the other is located downstream of the target site.
[0045] In one embodiment, during the reference genome sequence acquisition step, in each set of seed sequences, the base closest to the target site in the two sequences is equidistant from the target site. For example, if the target site is located at base position 1001 in the reference genome sequence, and the first set of seed sequences is located in the range [991-1000, 1002-1011] on the reference genome, then in the seed sequence upstream of the target site, the base closest to the target site is 1000, and in the seed sequence downstream of the target site, the base closest to the target site is 1002. The difference between 1000 and 1001 is 1, and the difference between 1002 and 1001 is also 1, meaning they are equidistant. For example, if the target site is at base position 1001 in the reference genome sequence, and the first set of seed sequences is located in the range [981-990, 1012-1021] in the reference genome, then in the seed sequence upstream of the target site, the base position closest to the target site is 990, and in the seed sequence downstream of the target site, the base position closest to the target site is 1012. The difference between 990 and 1001 is 11, and the difference between 1012 and 1001 is also 11, meaning they are equidistant.
[0046] In one embodiment, in the reference genome sequence acquisition step, the sequence containing the target site is a sequence containing X bases upstream and downstream of the target site, where X ≥ 10.
[0047] In one embodiment, X and M are equal, and the sequence lengths are equal here, which facilitates sequence alignment in subsequent steps.
[0048] In one embodiment, X is 10 to 50, preferably 20 to 40, and more preferably 30.
[0049] In one embodiment, X includes, but is not limited to, 10, 20, 30, 40, and 50.
[0050] In one embodiment, M ≥ 10 in the double sequence alignment step. If the value of M is too small, the sequence is too short, resulting in lower accuracy of the sequence alignment results.
[0051] In one embodiment, in the double sequence alignment step, M ≥ 30.
[0052] In one embodiment, in the double sequence alignment step, M includes, but is not limited to, 10, 20, 30, 40, and 50.
[0053] In one embodiment, in the reference genome seed sequence extraction step, the preset seed sequence length is ≥6nt.
[0054] In one embodiment, in the reference genome seed sequence extraction step, the preset seed sequence length is ≥10 nt.
[0055] In one embodiment, the preset seed sequence length in the reference genome seed sequence extraction step includes, but is not limited to, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 20nt, 30nt, 40nt, and 50nt.
[0056] In one embodiment, during the reference genome seed sequence extraction step, the preset seed sequence length is >6 nt, preferably 6–20 nt, and more preferably 10 nt. During matching, a shorter seed sequence is easier to match in the sequencing results, but the sequence nonspecificity will be higher; therefore, the minimum seed sequence length is set to 6 nt.
[0057] In one embodiment, N ≥ 1 in the reference genome seed sequence extraction step.
[0058] In one embodiment, N ≥ 2 in the reference genome seed sequence extraction step.
[0059] In one embodiment, N ≥ 3 in the reference genome seed sequence extraction step.
[0060] In one embodiment, in the reference genome seed sequence extraction step, N is 1 to 10, preferably 1 to 5.
[0061] In one embodiment, in the reference genome seed sequence extraction step, N includes, but is not limited to, 1, 2, 3, 4, and 5. Too many seed sequence sets, or sequences too far from the target site, can easily lead to inaccurate localization.
[0062] In one embodiment, during the reference genome seed sequence extraction step, all sequences in each group of seed sequences are of equal length.
[0063] In one embodiment, during the reference genome seed sequence extraction step, when N≥2, the various subsequences upstream of the base position corresponding to the target site in the reference genome sequence are sequentially adjacent.
[0064] In one embodiment, during the reference genome seed sequence extraction step, when N≥2, the various subsequences downstream of the target site in the reference genome sequence are sequentially adjacent. For example, when the number of mutated bases is 1, the mutation position is 1001, and the seed sequence length is 10nt, the three most recent seed sequences extracted are bases from the reference genome located in the intervals [991-1000, 1002-1011], [981-990, 1012-1021], and [971-980, 1022-1031]. The three seed sequences upstream of the mutation site are [971-980], [981-990], and [991-1000], which are adjacent to each other. The three seed sequences downstream of the mutation site are [1002-1011], [1012-1021], and [1022-1031], which are also adjacent to each other.
[0065] In one embodiment, during the seed sequence search step of the sequencing results, if a match fails and the seed sequence length is greater than 6 nt, the preset length of the seed sequence is reduced by 1 nt before the reference genome seed sequence extraction step and the seed sequence search step of the sequencing results are re-executed; if the seed sequence length is less than or equal to 6 nt when a match fails, it is determined that plotting cannot be performed and the program is terminated.
[0066] In one embodiment, during the seed sequence search step of the sequencing results, if a seed sequence appears multiple times in the sequencing result sequence during matching, it is also determined to be a matching failure.
[0067] In one embodiment, the comparison result is obtained based on the similarity rate in the determination step.
[0068] In one embodiment, the similarity rate refers to the proportion of matching bases to the total number of bases in any sequence participating in the sequence alignment. The sequence participating in the alignment refers to a sequence extracted from the first-generation sequencing results that contains a temporary target site and covers M bases upstream and downstream of it.
[0069] In one embodiment, the total number of bases in the sequence includes the number of empty positions added after alignment.
[0070] In one embodiment, in the determination step, the comparison result is obtained based on the relationship between the similarity rate and the preset threshold.
[0071] In one embodiment, in the determination step, if the similarity rate is greater than or equal to a preset threshold, then the two sequences are determined to be matched and the target site can be accurately located.
[0072] If the similarity rate is less than the preset threshold, the two sequences are determined to be dissimilar, the target site location is inaccurate, and the plot cannot be drawn.
[0073] In one embodiment, the preset threshold in the determination step can be set according to the actual situation, for example, it can be 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, etc.
[0074] In one embodiment, the method further includes a location determination step, which includes determining the location of the target site in the first-generation sequencing result sequence based on the results of double sequence alignments that are determined to be accurately locatable in the determination step.
[0075] In one embodiment, in the location determination step, the base corresponding to the base at the target site to be verified in the double sequence alignment result is the target site in the first-generation sequencing result sequence.
[0076] In one embodiment, the reference genome sequence acquisition step includes a target site, which may include a mutation site or a specific site.
[0077] In one embodiment, when the target site is a mutation site, in the location determination step, the base corresponding to the base at the mutation site to be verified in the reference genome in the double sequence alignment result is the location where the mutation occurs in the first-generation sequencing result sequence.
[0078] In one embodiment, when the target site is a mutation site, in the location determination step, the gaps or mismatches that appear in the double sequence alignment results indicate that a mutation has occurred.
[0079] In one embodiment, a plotting step is also included, in which a result map is plotted based on the mutation sites identified in the location determination step and the first-generation sequencing result sequence.
[0080] In one embodiment, in the seed sequence search step of the sequencing results, the sample to be tested is derived from a human or animal, preferably a human.
[0081] In one embodiment, in the seed sequence search step of the sequencing results, the sample to be tested is a tissue sample or a body fluid sample.
[0082] According to a second aspect, in one embodiment, an apparatus for verifying a target site is provided, comprising:
[0083] The reference genome sequence acquisition module is used to acquire a sequence containing the base positions corresponding to the target site from the reference genome based on the position of the target site to be verified in the reference genome;
[0084] The reference genome seed sequence extraction module is used to extract the top N seed sequences from the reference genome that are closest to the base position corresponding to the target site to be verified in the reference genome, according to a preset seed sequence length.
[0085] The sequencing result seed sequence finding module is used to compare the N sets of seed sequences extracted from the reference genome with the first-generation sequencing result sequences of the sample to be tested, and to search for seed sequences in the first-generation sequencing result sequences; if a seed sequence is found in the first-generation sequencing result sequences, it is determined that the match is successful; if none of the seed sequences appear in the first-generation sequencing result sequences, it is determined that the match is unsuccessful.
[0086] The temporary target site determination module is used to obtain the successfully matched seed sequence in the first-generation sequencing result sequence. By using the relative relationship between the seed sequence in the reference genome and the target site to be verified, the position of the target site in the first-generation sequencing result sequence is predicted, which is the temporary target site.
[0087] In one embodiment, a double sequence alignment module is further included, which is used to extract a sequence containing the temporary target site and covering M bases upstream and downstream of it from the first-generation sequencing result sequence according to the position of the temporary target site in the first-generation sequencing result sequence, and to perform double sequence alignment with the sequence extracted from the reference genome in the reference genome sequence acquisition step.
[0088] The judgment module is used to determine the alignment status and output the matching result based on the double sequence alignment results.
[0089] In one embodiment, a location determination module is further included, which is used to determine the location of the target site in the first-generation sequencing result sequence based on the result of double sequence alignment determined by the determination module to be accurately locatable.
[0090] In one embodiment, in the location determination module, the base position in the first-generation sequencing result sequence that corresponds to the position of the target site to be verified in the reference genome sequence in the double sequence alignment result is the target site position located by this module.
[0091] In one embodiment, a plotting module is also included, which is used to plot the results based on the target site location determined by the location determination module and the first-generation sequencing result sequence.
[0092] According to a third aspect, in one embodiment, an apparatus for verifying a target site is provided, comprising:
[0093] Memory, used to store programs;
[0094] A processor for implementing a method as described in any of the first aspects by executing a program stored in the memory.
[0095] According to a fourth aspect, in one embodiment, a computer-readable storage medium is provided, the medium storing a program that can be executed by a processor to implement the method as described in any of the first aspects.
[0096] To address the challenge of locating mutation sites and to present the target mutation site more conveniently and intuitively, this invention proposes a method for automatically plotting the mutation site to be verified and its upstream and downstream base sequences in first-generation sequencing results. The plotted results visually demonstrate the occurrence of the mutation. This invention utilizes sequence matching and alignment methods to locate the target mutation in the first-generation sequencing results. By setting a threshold for the proportion of matching bases, the location of the mutation site to be labeled can be located more accurately, avoiding the need for manual location of mutation sites in sequencing results. It also solves the problem of difficulty in locating mutation sites when there are other sporadic mutations in the sequencing results.
[0097] In one embodiment, the present invention can be divided into two parts: the first part is to locate the mutation site in the result sequence, and the second part is to plot the mutation site based on its location. The first part is the most crucial; the plotting part simply involves using a tool to directly plot the mutation based on the sequencing results and the location of the mutation. In the mutation site location part, a sequence containing the mutation site is obtained from the reference genome, and multiple sequences are extracted from this sequence according to a set seed sequence length as seed sequences for matching in the sequencing results. Once a seed sequence is found in the sequencing results, the location of the mutation site in the sequencing results can be roughly confirmed based on the relative position of the seed sequence and the mutation site in the reference genome. To more accurately locate the mutation site, a double sequence alignment method is used to compare the reference genome sequence containing the mutation site with the sequencing result sequence. The alignment results can accurately and scientifically reflect the location of the mutation site in the sequencing results. Simultaneously, to make the results more scientifically reasonable, a similarity rate threshold is set. Only when the two sequences reach the similarity rate threshold is the located mutation site considered accurate; therefore, only alignment results above the threshold can be plotted.
[0098] In one embodiment, after obtaining the mutation position in the sequencing results through the first part, the position information is passed to the plotting part. The plotting part can automatically draw a sequencing peak diagram centered on the mutation to be verified, containing 10 bases upstream and downstream.
[0099] In one embodiment, the most crucial part of this invention is locating the mutation site in the sequencing results. This method accomplishes this task by using sequence alignment. Sequence alignment can identify mutations such as gaps, mismatches, and incorrect matches to a certain extent, making it an accurate method for locating mutation sites.
[0100] In one embodiment, the method of the present invention includes a reference genome sequence acquisition step, a reference genome seed sequence extraction step, a sequencing result seed sequence search step, a temporary target site determination step, a step of acquiring the target segment sequence containing the temporary target site in the sequencing result, and a step of performing a double sequence alignment between the target segment and the corresponding sequence of the reference genome. Since the position of the mutation site to be verified in the sequence from the reference genome is determined, the mutation site can be located relatively accurately based on the new position of the mutation site to be verified in the alignment result. The present invention can very accurately and automatically draw sequencing peak diagrams containing the target site to be verified and its upstream and downstream sequences in the sequencing result, and annotate the target site, eliminating the trouble of manually finding and annotating the target site.
[0101] Example 1
[0102] Figure 1 The diagram shown is a flowchart of the process for plotting mutation site peaks in first-generation sequencing in this embodiment. The specific steps in this embodiment are as follows (the steps are shown below in order from top to bottom in the flowchart):
[0103] The first step is to input the program. You need to input at least the result file of the first generation sequencing, the reference genome file corresponding to the sequencing sample, and the folder where the results are stored.
[0104] The second step involves the program obtaining a sequence of bases from the reference genome containing the base position corresponding to the mutation site, based on the location of the mutation to be verified. This sequence contains 30 bases upstream and downstream of the mutation site and will be used for subsequent double sequence alignment.
[0105] The third step is to set the initial default seed sequence length to 10nt.
[0106] The fourth step involves extracting the top three seed sequences closest to the corresponding base position of the mutation to be verified in the reference genome, according to the currently set seed sequence length. These three seed sequences are sequentially adjacent to each other (for example, when the number of mutated bases is 1, the mutation position is 1001, and the seed sequence length is 10 nt, the three closest seed sequences extracted are bases from the reference genome located in the intervals [991-1000, 1002-1011], [981-990, 1012-1021], and [971-980, 1022-1031], respectively). The distance between these sequences and the corresponding base position of the mutation site to be verified in the reference genome is fixed.
[0107] Step 5 involves comparing the three sets of seed sequences obtained in the previous step with the sequencing results, searching for the seed sequence within the sequencing results. If a seed sequence is found in the sequencing results, the match is considered successful; if none of the six seed sequences appear in the sequencing results, this round of matching is considered a failure. Only after a successful match in this step can the process proceed to step 6. If this round of matching fails, and the seed sequence length is greater than 6 nt, the seed sequence length is reduced by 1 nt, and the program is restarted from step 4. If the seed sequence length is less than or equal to 5 nt when matching fails, an explanation file is output indicating that no graph can be generated, and the program terminates. During matching, the shorter the seed sequence length, the higher the sequence nonspecificity; therefore, the minimum seed sequence length is set to 6 nt. It is important to note that to avoid sequence nonspecificity, if a seed sequence appears multiple times in the sequencing results, the match is also considered a failure.
[0108] The sixth step, if the previous step is a successful match, can roughly infer the location of the mutation site in the sequencing results by using the relative relationship between the seed sequence in the previously recorded reference genome and the mutation site to be verified. However, the location obtained at this time is not completely accurate and can only be considered as a provisional mutation site.
[0109] Step 7: After obtaining the temporarily confirmed mutation site (i.e., the provisional mutation site, also known as the temporary mutation site) in the previous step, extract a sequencing result sequence containing the temporary mutation site and covering 30 bases upstream and downstream of it using its position in the first-generation sequencing result sequence. Perform double sequence alignment with the corresponding sequence previously extracted from the reference genome.
[0110] Step 8: After double sequence alignment, the effectiveness of the alignment can be judged based on the proportion of matched bases to the total number of bases in the sequence (including any gaps added after alignment) (i.e., the similarity rate). In this embodiment, the threshold is set to 75%. If the proportion of matched bases (i.e., the similarity rate) is greater than or equal to 75%, it indicates that the two sequences are relatively well-matched, and the mutation site can be located relatively accurately. If it is less than 75%, it indicates that the two sequences are not similar enough, and the location of the found mutation site is largely inaccurate. In this case, an explanatory document is directly output to explain why a graph cannot be drawn due to the low matching similarity. In this embodiment, to verify the TGG→GG mutation (as indicated by the arrow), the results obtained during double sequence alignment are as follows: Figure 2 As shown, multiple deletion mutations in the sample resulted in gaps during alignment. However, in this embodiment, almost all bases except for the gaps matched, resulting in an 82% similarity between the two sequences. Under such high similarity alignment conditions, mutation sites in the sample sequences can be accurately located.
[0111] The ninth step is to determine the location of the mutation site to be verified in the sequencing results based on the results of the double sequence alignment. In the alignment results, the bases from the first-generation sequencing results that match the bases at the mutation site to be verified in the reference genome are the locations where the mutation appears in the sequencing results. Some gaps or mismatches in the alignment indicate that the mutation has occurred.
[0112] Step 10: The plotting program generates a result graph based on the identified mutation sites and sequencing results.
[0113] In one embodiment, the present invention can automatically generate sequencing peak diagrams of sequencing results containing the mutation to be verified and its upstream and downstream sequences, and label the target mutation, eliminating the hassle of manually finding and labeling mutation sites. Furthermore, automatic graph generation is essential and meaningful for the verification of large amounts of mutation data.
[0114] In one embodiment, matching multiple seed sequences of varying lengths in the sequencing results significantly increases the probability of locating the true mutation site.
[0115] In one embodiment, using double sequence alignment to locate mutation sites is more accurate and convincing than manual methods or methods that infer the location of mutation sites by a fixed distance from the mutation site.
[0116] In one embodiment, a threshold is set for the double sequence alignment matching results. If the threshold is not reached, the image cannot be generated. This allows for early judgment in the event of poor sequencing results without wasting extra effort.
[0117] In one embodiment, a method is provided for locating mutation sites in first-generation sequencing results and plotting their mutation status based on sequence alignment. The core of this method is the mutation site location part. After locating the mutation site, the results are plotted according to a plotting program. In the mutation site location part, firstly, multiple adjacent seed sequences upstream and downstream of the mutation site to be verified are extracted from the genome. The mutation location is roughly located in the sequencing results based on the positions of the successfully matched seed sequences. After locating the temporary mutation site, a sequence containing the temporary mutation site with equal upstream and downstream base numbers is double-aligned with the corresponding sequence in the reference genome. After double-alignment, the location of the mutation in the sequencing results can be determined based on the position of the mutation to be verified after sequence alignment.
[0118] In one embodiment, sequence alignment can also be extended to the alignment of the entire sequencing result with a general segment on the corresponding chromosome of the reference genome. However, since it is uncertain whether the sequencing result corresponds to the specific location in the reference genome, the similarity of such alignment results will not be very high, and it is relatively inaccurate.
[0119] In one embodiment, the temporary mutation site located after seed sequence matching can also be roughly used as the mutation site in the sequencing results. However, since there may be certain mutations near the mutation site, especially if there are deletions or additions of bases near its position, the location of the mutation site may be misplaced.
[0120] In one embodiment, the number of seed sequences involved in this invention can also be set to more per round, moving further upstream and downstream of the mutation site. However, since the distance from the mutation site is too far, the localization may be biased and inaccurate.
[0121] In one embodiment, the method constructed by the present invention can not only be used to view mutations, but also, with simple modifications, can be used to view mutations at specific sites in certain situations.
[0122] Those skilled in the art will understand that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or external hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.
[0123] The above examples illustrate the present invention only to aid in understanding it and are not intended to limit the scope of the invention. Those skilled in the art can make various simple deductions, modifications, or substitutions based on the ideas of this invention.
Claims
1. A method for verifying target sites, characterized in that, include: The reference genome sequence acquisition step includes obtaining a sequence containing the base positions corresponding to the target site from the reference genome based on the position of the target site to be verified in the reference genome; The reference genome seed sequence extraction step includes extracting the top N seed sequences from the reference genome that are closest to the base position corresponding to the target site to be verified in the reference genome, according to a preset seed sequence length; the preset seed sequence length is ≥6nt; N≥3; The seed sequence finding step for sequencing results includes comparing the N sets of seed sequences extracted from the reference genome with the first-generation sequencing results of the sample to be tested, and searching for seed sequences in the first-generation sequencing results. If a seed sequence is found in the first-generation sequencing results, it is considered a successful match. If none of the seed sequences appear in the first-generation sequencing results, it is considered a failed match. If a match fails and the seed sequence length is >6 nt, the preset length of the seed sequence is reduced by 1 nt and the reference genome seed sequence extraction step and the sequencing result seed sequence finding step are repeated. If the seed sequence length is ≤5 nt when a match fails, it is determined that plotting cannot be performed and the program ends. Furthermore, if a seed sequence appears multiple times in the sequencing results during matching, it is also considered a failed match. The temporary target site determination step includes obtaining the successfully matched seed sequence from the first-generation sequencing results, and predicting the position of the target site in the first-generation sequencing results by the relative relationship between the seed sequence in the reference genome and the target site to be verified, which is the temporary target site; The double sequence alignment step includes extracting a sequence containing the temporary target site and covering M bases upstream and downstream of it from the first-generation sequencing result sequence based on the position of the temporary target site in the sequence, and performing double sequence alignment with the sequence extracted from the reference genome in the reference genome sequence acquisition step. M≥10; The determination step involves obtaining the alignment result based on the relationship between the similarity rate and a preset threshold, judging the alignment status, and outputting the matching result. The similarity rate refers to the proportion of the number of matched bases to the total number of bases in the sequence. If the similarity rate is greater than or equal to the preset threshold, the two sequences are determined to be matched, and the base position corresponding to the target site can be accurately located. If the similarity rate is less than the preset threshold, the two sequences are determined to be dissimilar, the base positions of the target sites are inaccurate, and the graph cannot be drawn.
2. The method as described in claim 1, characterized in that, In the step of obtaining the reference genome sequence, the sequence containing the target site is a sequence containing X bases upstream and downstream of the corresponding base position of the target site.
3. The method as described in claim 2, characterized in that, X is equal to M.
4. The method as described in claim 2, characterized in that, In the reference genome sequence acquisition step, X ≥ 10.
5. The method as described in claim 2, characterized in that, In the reference genome sequence acquisition step, X is 10~50.
6. The method as described in claim 1, characterized in that, In the reference genome seed sequence extraction step, the preset seed sequence length is ≥10 nt.
7. The method as described in claim 1, characterized in that, The total number of bases in the sequence includes the number of missing numbers after alignment.
8. The method as described in claim 1, characterized in that, It also includes a location determination step, which involves determining the location of the target site in the first-generation sequencing result sequence based on the results of double sequence alignments that are determined to be accurately localizable in the determination step.
9. The method as described in claim 8, characterized in that, In the location determination step, the bases in the double sequence alignment results that correspond to the bases at the target site to be verified in the reference genome are the target sites in the first-generation sequencing results.
10. The method as described in claim 8, characterized in that, The reference genome sequence acquisition steps include target sites, which may include mutation sites or specific sites.
11. The method as described in claim 10, characterized in that, When the target site is a mutation site, in the location determination step, the base corresponding to the base at the mutation site to be verified in the reference genome in the double sequence alignment results is the location where the mutation occurs in the first-generation sequencing result sequence.
12. The method as described in claim 10, characterized in that, When the target site is a mutation site, the gaps or mismatches that appear in the double sequence alignment results during the location determination step indicate that a mutation has occurred.
13. The method as described in claim 8, characterized in that, It also includes a plotting step, which uses the target site location determined in the location determination step and the first-generation sequencing result sequence to draw a result map.
14. The method as described in claim 1, characterized in that, In the seed sequence search step of sequencing results, the sample to be tested is derived from a human or an animal.
15. The method as described in claim 14, characterized in that, The samples to be tested were derived from humans.
16. The method as described in claim 1, characterized in that, In the seed sequence search step of sequencing results, the sample to be tested is a tissue sample or a body fluid sample.
17. An apparatus for verifying a target site, characterized in that, include: The reference genome sequence acquisition module is used to acquire a sequence containing the base positions corresponding to the target site from the reference genome based on the position of the target site to be verified in the reference genome; The reference genome seed sequence extraction module is used to extract the top N seed sequences from the reference genome that are closest to the base position corresponding to the target site to be verified in the reference genome, according to a preset seed sequence length. The preset seed sequence length is ≥6nt; N≥3; The sequencing result seed sequence finding module is used to compare the N sets of seed sequences extracted from the reference genome with the first-generation sequencing result sequences of the sample to be tested, and to find seed sequences in the first-generation sequencing result sequences; If a seed sequence is found in the first-generation sequencing results, the match is considered successful. If none of the seed sequences appear in the first-generation sequencing results, the match is considered unsuccessful. If the match fails and the seed sequence length is >6 nt, the preset seed sequence length is reduced by 1 nt before the reference genome seed sequence extraction step and the sequencing result seed sequence search step are repeated. If the seed sequence length is ≤5 nt when the match fails, the plot cannot be generated and the program ends. Furthermore, if a seed sequence appears multiple times in the sequencing results during the matching process, the match is also considered unsuccessful. M≥10; The temporary target site determination module is used to obtain the successfully matched seed sequence in the first-generation sequencing result sequence. By using the relative relationship between the seed sequence in the reference genome and the target site to be verified, the position of the target site in the first-generation sequencing result sequence is predicted, which is the temporary target site. The double sequence alignment module is used to extract a sequence containing the temporary target site and covering M bases upstream and downstream of it from the first-generation sequencing result sequence based on the position of the temporary target site in the sequence. The extracted sequence is then double-aligned with the sequence extracted from the reference genome in the reference genome sequence acquisition step. The determination module is used to determine the alignment status and output the matching result based on the double sequence alignment results; to obtain the alignment result based on the relationship between the similarity rate and the preset threshold, determine the alignment status and output the matching result; the similarity rate refers to the proportion of the number of matched bases to the total number of bases in the sequence; if the similarity rate is ≥ the preset threshold, the two sequences are determined to be matched, and the base position corresponding to the target site can be accurately located; If the similarity rate is less than the preset threshold, the two sequences are determined to be dissimilar, the base positions of the target sites are inaccurate, and the graph cannot be drawn.
18. The apparatus as claimed in claim 17, characterized in that, It also includes a location determination module, which is used to determine the location of the target site in the first-generation sequencing result sequence based on the results of double sequence alignments that are determined to be accurately locatable in the determination module.
19. The apparatus as claimed in claim 18, characterized in that, In the location determination module, the base position in the first-generation sequencing result sequence that corresponds to the position of the target site to be verified in the reference genome sequence in the double sequence alignment result is the target site position located by this module.
20. The apparatus as claimed in claim 17, characterized in that, It also includes a plotting module, which is used to plot the results based on the target site location determined by the location determination module and the first-generation sequencing result sequence.
21. An apparatus for verifying target sites, characterized in that, include: Memory, used to store programs; A processor for implementing the method as described in any one of claims 1 to 16 by executing a program stored in the memory.
22. A computer-readable storage medium, characterized in that, The medium stores a program that can be executed by a processor to implement the method as described in any one of claims 1 to 16.
Citation Information
Patent Citations
Method and device for detecting designated location point of variation
CN106909806A
Information processing system, mutation detection system, storage medium, and information processing method
US20210158896A1