A method and system for gene mutation analysis
After two-color 2+2 fuzzy sequencing, the reference sub-gene sequence was selected and split and aligned, the error comparison problem caused by the encoding method was solved, and the accuracy of gene variant analysis was improved.
Patent Information
- Application Number
- CN202310728147.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-19
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-06-19
AI Technical Summary
When detecting SNV, the two-color 2+2 fuzzy sequencing is affected by the encoding method and is prone to incorrect gene mutation types, resulting in inaccurate gene mutation analysis results.
After two-color 2+2 fuzzy sequencing, the reference sub-gene sequence is selected from the reference gene sequence according to the first alignment results, and the fuzzy sequence and the reference sub-gene sequence to be analyzed are split according to the second base combination, and further comparative is obtained to obtain the second and third alignment results, and finally merge to obtain the gene alignment results.
Effectively exclude the influence of coding methods, the accuracy of two-color 2+2 fuzzy sequencing to detect the relative positional relationship of SNV changing the base site map is improved, and the accuracy of gene variant types is ensured.
Smart Images

Figure CN116758990B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of gene sequencing technology, and in particular, to a gene variant analysis method and system. Background Art
[0002] Fuzzy sequencing is a high-information-efficiency sequencing method that can quickly detect various gene variants. A common fuzzy sequencing method is: two-color 2+2 fuzzy sequencing. That is, during the sequencing of a nucleic acid sample, two reaction solutions are used, and each reaction solution contains nucleotide substrate molecules of two different bases, and the two nucleotides are respectively labeled with fluorescent groups of different colors; the nucleotide substrate molecules in one reaction solution can be complementary to two bases on the nucleotide sequence to be detected, and the nucleotides in the other reaction solution are complementary to the other two bases on the nucleotide sequence to be detected. The two reaction solutions are cyclically added, and the fluorescence signal is detected after each reaction is completed to obtain a sequencing signal. After encoding the sequencing signal through a preset encoding method, a fuzzy sequence of the nucleic acid sample is obtained, and it can be aligned to the encoded reference gene sequence by using conventional bioinformatics software, thereby detecting gene variants.
[0003] The inventors found that when detecting SNV (single-nucleotide variation) by two-color 2+2 fuzzy sequencing, affected by the encoding method, when the relative positional relationship of the mapped sites of bases is changed by gene variants, incorrect alignment results are likely to occur, resulting in incorrect gene variant types, making the results of gene variant analysis inaccurate. Summary of the Invention
[0004] In view of this, embodiments of the present disclosure provide a gene variant analysis method and system, which can improve the accuracy in the case where two-color 2+2 fuzzy sequencing detects the relative positional relationship of the mapped sites of bases changed by SNV.
[0005] In a first aspect, embodiments of the present disclosure provide a gene variant analysis method, adopting the following technical solution:
[0006] The gene variant analysis method includes:
[0007] Perform two-color 2+2 fuzzy sequencing on a nucleic acid sample according to a first base combination to obtain a sequencing signal of the nucleic acid sample;
[0008] Encode the sequencing signal to obtain a fuzzy sequence to be analyzed, and encode a reference genome to obtain a reference gene sequence;
[0009] Align the fuzzy sequence to be analyzed with the reference gene sequence to obtain a first alignment result;
[0010] Select a reference sub-gene sequence from the reference gene sequence according to the first alignment result, where the base positions included in the reference sub-gene sequence overlap at least partially with the base positions included in the first alignment result;
[0011] Split the fuzzy sequence to be analyzed according to the second base combination to obtain a first fuzzy semi-sequence to be analyzed and a second fuzzy semi-sequence to be analyzed;
[0012] Split the reference sub-gene sequence according to the second base combination to obtain a first reference sub-gene semi-sequence and a second reference sub-gene semi-sequence;
[0013] Align the first fuzzy semi-sequence to be analyzed with the corresponding first reference sub-gene semi-sequence to obtain a second alignment result, and align the second fuzzy semi-sequence to be analyzed with the corresponding second reference sub-gene semi-sequence to obtain a third alignment result;
[0014] Obtain a gene alignment result according to the second alignment result and the third alignment result.
[0015] Optionally, the base positions included in the first alignment result are the a-th to b-th bases of the reference gene sequence;
[0016] The step of selecting a reference sub-gene sequence from the reference gene sequence according to the first alignment result includes: selecting the c-th to d-th bases from the reference gene sequence as the reference sub-gene sequence; where a ≤ c ≤ b, or c ≤ a ≤ d.
[0017] Optionally, the values of c and d satisfy: covering the positions of known variations on the fuzzy sequence to be analyzed, and / or covering the part with higher quality or more reliable part in the first alignment result.
[0018] Optionally, the base positions included in the first alignment result are the a-th to b-th bases of the reference gene sequence;
[0019] The step of selecting a reference sub-gene sequence from the reference gene sequence according to the first alignment result includes: selecting the c-th to d-th bases from the reference gene sequence as the reference sub-gene sequence; where c ≤ a and b ≤ d.
[0020] Optionally, the base combinations include MK, RY, and WS; the first base combination is one of the base combinations, and the second base combination is one of the other two base combinations; where M represents bases A, C; K represents bases T, G; R represents bases A, G; Y represents bases C, T; W represents bases A, T; S represents bases C, G.
[0021] Optionally, the gene mutation analysis method further includes: obtaining the site mapping of the first reference sub-gene half-sequence and the second reference sub-gene half-sequence;
[0022] The second alignment result includes the base correspondence between the first fuzzy half-sequence to be analyzed and the first reference sub-gene half-sequence, and the corresponding site mapping;
[0023] The third alignment result includes the base correspondence between the second fuzzy half-sequence to be analyzed and the second reference sub-gene half-sequence, and the corresponding site mapping.
[0024] Optionally, obtaining the gene alignment result according to the second alignment result and the third alignment result includes:
[0025] According to the site mapping in the second alignment result and the site mapping in the third alignment result, merge the second alignment result and the third alignment result in a preset merging manner to obtain the gene alignment result.
[0026] Optionally, the preset merging manner includes:
[0027] Initialize the gene alignment result to be empty;
[0028] Find the target base with the smallest site mapping in the second alignment result and the third alignment result, and write the target base and its alignment result in the second alignment result or the third alignment result into the gene alignment result;
[0029] Delete the target base from the alignment result where the target base is located;
[0030] Judge whether there is an insertion after the alignment result where the target base is located. If so, write the insertion into the gene alignment result as well;
[0031] Delete the insertion after the target base from the alignment result where the target base is located;
[0032] Return to execute the step of finding the target base with the smallest site mapping in the second alignment result and the third alignment result until all bases in the second alignment result and the third alignment result have been written into the gene alignment result.
[0033] Optionally, the gene mutation analysis method further includes: re-extracting a new fuzzy sequence to be analyzed and a new reference sub-gene sequence from the gene alignment result; aligning the new fuzzy sequence to be analyzed and the new reference sub-gene sequence to obtain a final alignment result.
[0034] Second aspect, embodiments of the present disclosure further provide a gene mutation analysis system, adopting the following technical solutions:
[0035] The gene mutation analysis system includes:
[0036] A sequencing module, configured to perform two-color 2+2 ambiguous sequencing on a nucleic acid sample according to a first base combination, and obtain sequencing signals of the nucleic acid sample;
[0037] An encoding module, configured to encode the sequencing signals to obtain an ambiguous sequence to be analyzed, and encode a reference genome to obtain a reference gene sequence;
[0038] A first alignment module, configured to align the ambiguous sequence to be analyzed with the reference gene sequence to obtain a first alignment result;
[0039] A subsequence selection module, configured to select a reference sub-gene sequence from the reference gene sequence according to the first alignment result, where the base positions included in the reference sub-gene sequence at least partially overlap with the base positions included in the first alignment result;
[0040] A sequence splitting module, configured to split the ambiguous sequence to be analyzed according to a second base combination to obtain a first ambiguous semi-sequence to be analyzed and a second ambiguous semi-sequence to be analyzed, and split the reference sub-gene sequence according to the second base combination to obtain a first reference sub-gene semi-sequence and a second reference sub-gene semi-sequence;
[0041] A second alignment module, configured to align the first ambiguous semi-sequence to be analyzed with the corresponding first reference sub-gene semi-sequence to obtain a second alignment result, and align the second ambiguous semi-sequence to be analyzed with the corresponding second reference sub-gene semi-sequence to obtain a third alignment result;
[0042] A result integration module, configured to obtain a gene alignment result according to the second alignment result and the third alignment result.
[0043] Optionally, the gene mutation analysis system further includes: a third alignment module, configured to re-extract a new ambiguous sequence to be analyzed and a new reference sub-gene sequence from the gene alignment result; align the new ambiguous sequence to be analyzed and the new reference sub-gene sequence to obtain a final alignment result.
[0044] The embodiments of the present disclosure provide a gene mutation analysis method and system. In the gene mutation analysis method, after performing two-color 2+2 fuzzy sequencing based on the first base combination, aligning the fuzzy sequence to be analyzed with the reference gene sequence to obtain the first alignment result, according to the first alignment result, a reference sub-gene sequence is selected from the reference gene sequence, and according to the second base combination, the fuzzy sequence to be analyzed is split to obtain the first fuzzy semi-sequence to be analyzed and the second fuzzy semi-sequence to be analyzed. The reference sub-gene sequence is split to obtain the first reference sub-gene semi-sequence and the second reference sub-gene semi-sequence. Further, the first fuzzy semi-sequence to be analyzed is aligned with the corresponding first reference sub-gene semi-sequence to obtain the second alignment result, and the second fuzzy semi-sequence to be analyzed is aligned with the corresponding second reference sub-gene semi-sequence to obtain the third alignment result. Finally, according to the second alignment result and the third alignment result, the gene alignment result is obtained. By further aligning to obtain the second alignment result and the third alignment result, as well as the integration of the second alignment result and the third alignment result, it is possible to effectively exclude the possible incorrect alignment results in the first alignment result caused by the influence of the coding method, and thus it is possible to obtain the correct gene mutation type, improving the accuracy in the case of detecting the relative position relationship of the sites where the SNV changes the bases by two-color 2+2 fuzzy sequencing.
[0045] The above description is only an overview of the technical solution of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present disclosure more obvious and understandable, the following preferred embodiments are specifically given and described in detail in conjunction with the drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a flowchart of the gene mutation analysis method provided by the embodiments of the present disclosure;
[0048] Figure 2 It is a schematic block diagram of the principle of the gene mutation analysis system provided by the embodiments of the present disclosure;
[0049] Figure 3 It is a statistical chart of the gene mutation detection results of high-throughput sequencing in the prior art;
[0050] Figure 4 It is a statistical chart of the gene mutation detection results of high-throughput sequencing provided by the embodiments of the present disclosure. Detailed Implementation Modes
[0051] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0052] It should be clear that the following uses specific specific examples to illustrate the implementation modes of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific implementation modes. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.
[0053] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.
[0054] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0055] An example of the situation where the gene variation mentioned in the background art changes the relative positional relationship of the site mapping of bases is as follows:
[0056] Example 1: When detecting C>T by dual-color MK fuzzy sequencing, the DNA sequence of the wild-type sequence is AACAA, the site mapping is 12345, the encoded coding sequence is AAAAC, and its corresponding site mapping is 12453. After the C>T variation occurs, the DNA sequence of the mutant sequence is AATAA, its corresponding site mapping is 12345, the encoded coding sequence is AATAA, and its corresponding site mapping is 12345. The gene variation changes the relative positional relationship of the site mapping of bases in the encoded coding sequence.
[0057] Example 2: When detecting C>A by dual-color MK fuzzy sequencing, the DNA sequence of the wild-type sequence is ACACA, the site mapping is 12345, the encoded coding sequence is AAACC, and its corresponding site mapping is 13524. After the C>A mutation occurs, the DNA sequence of the mutant sequence is AAACA, its corresponding site mapping is 12345, the encoded coding sequence is AAAAC, and its corresponding site mapping is 12354. The gene mutation changes the relative position relationship of the site mapping of the bases in the encoded coding sequence.
[0058] Example 3: When detecting deletions by dual-color MK fuzzy sequencing, the DNA sequence of the wild-type sequence is ACGAC, the site mapping is 12345, the encoded coding sequence is ACGAC, and its corresponding site mapping is 12345. After the deletion mutation occurs, the DNA sequence of the mutant sequence is ACAC, its corresponding site mapping is 1234, the encoded coding sequence is AACC, and its corresponding site mapping is 1324. The gene mutation changes the relative position relationship of the site mapping of the bases in the encoded coding sequence.
[0059] The embodiments of the present disclosure provide a gene mutation analysis method. Specifically, as Figure 1 shown, the gene mutation analysis method includes:
[0060] Step S1: Perform dual-color 2+2 fuzzy sequencing on the nucleic acid sample according to the first base combination to obtain the sequencing signal of the nucleic acid sample.
[0061] Among them, the first base combination is one of MK, RY, and WS. Among them, M represents bases A and C; K represents bases T and G; R represents bases A and G; Y represents bases C and T; W represents bases A and T; S represents bases C and G.
[0062] Optionally, performing dual-color 2+2 fuzzy sequencing on the nucleic acid sample according to the first base combination to obtain the sequencing signal of the nucleic acid sample includes: performing dual-color 2+2 fuzzy sequencing on the nucleic acid sample using a nucleotide substrate molecule with a fluorescent group modified with 5'-terminal polyphosphate having fluorescence switching properties to obtain the sequencing signal. Among them, "fluorescence switching properties" means that the fluorescence signal after the sequencing reaction changes significantly compared with that before the sequencing reaction. For a more detailed description of the above content, reference can be made to patents such as CN201510212789.1 and CN201510212788.7.
[0063] Specifically, 2 + 2 fuzzy sequencing means that in the sequencing reaction, two reaction solutions are used, and each reaction solution contains nucleotide substrate molecules of two different bases; the nucleotide substrate molecules in one reaction solution can be complementary to two bases on the nucleotide sequence to be measured, and the nucleotides in the other reaction solution are complementary to the other two bases on the nucleotide sequence to be measured. First, the nucleotide sequence fragment to be measured can be fixed in the reaction chamber, then a reaction solution is introduced, and then an enzyme is used to release the fluorophore on the nucleotide substrate with a fluorescence switching property, thereby causing fluorescence switching; then the second reaction solution is introduced; an enzyme is used to release the fluorophore on the nucleotide substrate with a fluorescence switching property, thereby causing fluorescence switching; the two reaction solutions are added cyclically, and the coding information of the nucleotide substrate to be measured is obtained through the fluorescence information. Dual-color 2 + 2 fuzzy sequencing means that the two nucleotides contained in each reaction solution are respectively labeled with fluorophores of different colors.
[0064] Optionally, the process of dual-color 2 + 2 fuzzy sequencing is as follows:
[0065] The sequencing reaction mixture includes two sets of sequencing reaction solutions. Each set of sequencing reaction solutions contains nucleotide substrates of two bases. Each nucleotide substrate can contain a small number of fluorescently labeled non-terminal terminating nucleotides and a large number of unlabeled non-terminal terminating (natural) nucleotides. The fluorescent labels of the two nucleotide substrates are different. Taking the two-color MK sequencing as an example, the sequencing reaction includes two sets of sequencing reaction solutions. The nucleic acid sample sequence (clone cluster) is exposed to the sequencing reaction mixture, and polymerase extension is carried out to incorporate 0, 1, or more nucleotide bases (A, C) into each growing strand. Unreacted substrates are washed away. Then, the surface of the chip is scanned to measure the fluorescence level of each clone cluster. Then, all fluorescent labels are cleaved and washed. After that, it enters the sequencing reaction mixture including another two nucleotide substrates (G, T) for the sequencing reaction, and the two sets of sequencing reaction solutions cycle in. In a single sequencing cycle, the total amount of labeled nucleotides bound to the clone cluster is linearly proportional to the length of the corresponding homologous polymer in the template, and most of the individually synthesized DNA strands remain unlabeled, even at long template homopolymer lengths, and most of the scarring effects of reversible termination chemistry are eliminated. For a more detailed description of the above two-color 2+2 ambiguous sequencing, see the article Almogy, G., Pratt, M. et al., Cost-efficient whole genome-sequencing using novel mostly natural sequencing-by-synthesis chemistry and open fluidics platform, bioRxiv 2022.05.29.493900; doi: https: / / doi.org / 10.1101 / 2022.05.29.493900.
[0066] Optionally, taking the first base combination as MK, before sequencing, the nucleic acid sample to be analyzed is fixed. During the two-color 2+2 ambiguous sequencing process, there are a total of two sets of reaction solutions. One set of reaction solutions contains nucleotide substrate molecules of two bases A and C, or contains nucleotide substrate molecules of two bases T and G. The other set of reaction solutions contains nucleotide substrate molecules of two bases T and G, or contains nucleotide substrate molecules of two bases A and C. The two sets of reaction solutions are cycled in (for example, MKMKMK… or MKMKMMK…, etc.). The added nucleotide substrate molecules react with the fixed nucleotide fragments. After each reaction, since the nucleotide substrate molecules containing different bases are modified with different fluorescent groups, different colors are presented. According to the color and quantity of the fluorescence signals detected after each reaction, the sequencing signals can be obtained.
[0067] When the first base combination is RY, in the two-color 2+2 fuzzy sequencing process, there are a total of two groups of reaction solutions. One group of reaction solutions contains nucleotide substrate molecules of two bases, A and G, or contains nucleotide substrate molecules of two bases, C and T. The other group of reaction solutions contains nucleotide substrate molecules of two bases, C and T, or contains nucleotide substrate molecules of two bases, A and G. The two groups of reaction solutions are added cyclically. When the first base combination is WS, in the two-color 2+2 fuzzy sequencing process, there are a total of two groups of reaction solutions. One group of reaction solutions contains nucleotide substrate molecules of two bases, A and T, or contains nucleotide substrate molecules of two bases, C and G. The other group of reaction solutions contains nucleotide substrate molecules of two bases, C and G, or contains nucleotide substrate molecules of two bases, A and T. The two groups of reaction solutions are added cyclically.
[0068] Exemplarily, taking the sequencing of CACTAA by two-color MK fuzzy sequencing as an example, the sequencing reaction order is KMK. Combining different fluorescence colors, the obtained sequencing signals are: (1A+2C, 0G+1T, 2A+0C).
[0069] Step S2: Encode the sequencing signals to obtain the fuzzy sequence to be analyzed, and encode the reference genome to obtain the reference gene sequence.
[0070] Optionally, the specific coding method used for encoding satisfies that the sequencing signals or the reference genome are reverse complemented after encoding, or reverse complemented first and then encoded, and the obtained results are the same.
[0071] When encoding the reference genome to obtain the reference gene sequence, first divide the reference genome into several substrings in order. Each substring only contains the bases corresponding to this two-color 2+2 fuzzy sequencing (for example, under two-color MK fuzzy sequencing, each substring is only composed of A and / or C or only composed of G and / or T). Then, each substring is rearranged in ascending order according to the alphabetical order, and the rearranged substrings are connected in order to form a new string, and the encoding process can be completed.
[0072] When encoding the sequencing signals to obtain the fuzzy sequence to be analyzed, the sequencing signals have been divided into several substrings in order. Each substring only contains the bases corresponding to this two-color 2+2 fuzzy sequencing (for example, under two-color MK fuzzy sequencing, each substring is only composed of A and / or C or only composed of G and / or T). Then, only each substring needs to be rearranged in ascending order according to the alphabetical order, and the rearranged substrings are connected in order to form a new string, and the encoding process can be completed. Exemplarily, the sequencing signals are (2A+0C, 3G+1T, 1A+2C, 0G+1T), and after rearrangement, it is AAGGGTACCT.
[0073] Optionally, after rearranging the bases of the reference genome, the positions corresponding to each base before rearrangement can be recorded, i.e., absolute site mapping. Exemplarily, the original sequence is ACAGTCCTG (sequence number 123456789), after segmentation it is: ACA, GT, CC, TG (sequence number 123,45,67,89), after rearrangement it is AACGTCCGT (sequence number 132456798), and the final sequence number 132456798 is the absolute site mapping.
[0074] Step S3: Align the fuzzy sequence to be analyzed with the reference gene sequence to obtain a first alignment result.
[0075] Optionally, aligning the fuzzy sequence to be analyzed with the reference gene sequence to obtain a first alignment result includes: aligning the fuzzy sequence to be analyzed with the reference gene sequence through algorithms such as BWA, bowtie2, SOAP, or Smith-Waterman to obtain a first alignment result. Exemplarily, the first alignment result includes base positions (which base in the reference gene sequence the fuzzy sequence to be analyzed corresponds to).
[0076] Step S4: According to the first alignment result, select a reference sub-gene sequence from the reference gene sequence, where the base positions included in the reference sub-gene sequence at least partially overlap with the base positions included in the first alignment result.
[0077] Optionally, when the first alignment result has different situations, the selected reference sub-gene sequence can also be adjusted adaptively.
[0078] In one example, the base positions included in the first alignment result are the a-th to b-th bases of the reference gene sequence, and the c-th to d-th bases are selected from the reference gene sequence as the reference sub-gene sequence, where a ≤ c ≤ b, or c ≤ a ≤ d, so that the base positions included in the reference sub-gene sequence at least partially overlap with the base positions included in the first alignment result.
[0079] In the above value-taking method, although in individual cases there may be a situation where the values of c and d do not cover the variant base positions, resulting in the inability to detect variants in subsequent alignment steps, in the application scenario of high-throughput sequencing, it will not affect the overall detection result.
[0080] Among them, the optimal values of c and d satisfy: covering the positions of known mutations on the fuzzy sequence to be analyzed, and / or covering the parts with higher quality or more reliable parts in the first alignment result. Known mutations refer to mutations that may exist in the nucleic acid sample to be detected according to the type of nucleic acid sample detected and relevant field knowledge, or mutations whose presence or absence has great reference significance for clinical diagnosis. For example, the mutations recorded in the NCCN (National Comprehensive Cancer Network) guidelines have important guiding effects on the diagnosis, typing, treatment, and prognosis of tumors. Another example is that in the detection of minimal residual disease, whole exome sequencing is first performed on the tumor tissue removed from the patient during surgery to obtain a set of candidate mutations, and then targeted sequencing is performed on the cell-free DNA (cfDNA) in the patient's peripheral blood. At this time, the candidate mutations are the known mutations during targeted sequencing. The parts with higher quality or more reliable parts refer to, including but not limited to: 1. The remaining parts after removing clips; 2. The parts with shorter lengths of homologous polymers and / or binary copolymers (that is, selecting the parts with lengths of homologous polymers and / or binary copolymers less than 8, and according to actual needs, further selecting the parts with lengths of homologous polymers and / or binary copolymers less than 3, less than 4, less than 5, less than 6, or less than 7); 3. The parts with higher base quality values of the measured sequence (that is, the parts where the average or median of the base quality values of the measured sequence is greater than 15, and according to actual needs, further selecting the parts where the average or median of the base quality values of the measured sequence is greater than 20, greater than 25, or greater than 30).
[0081] In another example, the base positions included in the first alignment result are the a-th to b-th bases of the reference gene sequence, and the c-th to d-th bases are selected from the reference gene sequence as the reference sub-gene sequence, where c ≤ a and b ≤ d, so that the base positions included in the reference sub-gene sequence completely cover the base positions included in the first alignment result.
[0082] Step S5: According to the second base combination, split the fuzzy sequence to be analyzed to obtain a first fuzzy semi-sequence to be analyzed and a second fuzzy semi-sequence to be analyzed.
[0083] Exemplarily, when the first base combination is MK, the second base combination is RY or WS; when the first base combination is RY, the second base combination is MK or WS; when the first base combination is WS, the second base combination is MK or RY.
[0084] Taking the first base combination as MK and the second base combination as RY as an example, after splitting the fuzzy sequence GGTCCGACGCACTAACGAGGTAACCGGCGCGACG to be analyzed, the first fuzzy semi-sequence to be analyzed GGGAGAAAGAGGAAGGGGAG and the second fuzzy semi-sequence to be analyzed TCCCCCTCTCCCCC are obtained.
[0085] Step S6: According to the second base combination, split the reference sub-gene sequence to obtain a first reference sub-gene semi-sequence and a second reference sub-gene semi-sequence.
[0086] Taking the first base combination as MK and the second base combination as RY as an example, after splitting the reference sub-gene sequence GGTCCGACGAAACCCCGAGGTAACCGGCGCGACG, the first reference sub-gene semi-sequence GGGAGAAAGAGGAAGGGGAG and the second reference sub-gene semi-sequence TCCCCCCCTCCCCC are obtained.
[0087] The specific order of the above steps S5 and S6 can be adjusted according to actual needs. Step S5 can be executed first and step S6 later, or step S6 can be executed first and step S5 later, or the two can be executed simultaneously.
[0088] Optionally, the gene mutation analysis method further includes: obtaining the site mapping of the first reference sub-gene semi-sequence and the second reference sub-gene semi-sequence. The above site mapping can be an absolute site mapping or a relative site mapping. Among them, if the starting site of the reference sub-gene semi-sequence is x0, the relative site mapping of the first reference sub-gene semi-sequence is the value obtained by subtracting the starting site x0 from the absolute site mapping of the first reference sub-gene semi-sequence, and the relative site mapping of the second reference sub-gene semi-sequence is the value obtained by subtracting the starting site x0 from the absolute site mapping of the second reference sub-gene semi-sequence. If it is a relative site mapping here, the subsequent alignment calculation operations can be appropriately simplified.
[0089] Step S7: Align the first fuzzy semi-sequence to be analyzed with the corresponding first reference sub-gene semi-sequence to obtain a second alignment result, and align the second fuzzy semi-sequence to be analyzed with the corresponding second reference sub-gene semi-sequence to obtain a third alignment result.
[0090] Exemplarily, on the premise of obtaining the site mapping of the first reference sub-gene half-sequence and the second reference sub-gene half-sequence, the second alignment result includes the base correspondence between the first fuzzy half-sequence to be analyzed and the first reference sub-gene half-sequence, as well as the corresponding site mapping; the third alignment result includes the base correspondence between the second fuzzy half-sequence to be analyzed and the second reference sub-gene half-sequence, as well as the corresponding site mapping. The specific forms of the second alignment result and the third alignment result will be described in detail in the subsequent embodiments.
[0091] Step S8, obtain a gene alignment result according to the second alignment result and the third alignment result.
[0092] Optionally, on the premise that the second alignment result includes the base correspondence between the first fuzzy half-sequence to be analyzed and the first reference sub-gene half-sequence, as well as the corresponding site mapping, and the third alignment result includes the base correspondence between the second fuzzy half-sequence to be analyzed and the second reference sub-gene half-sequence, as well as the corresponding site mapping, obtaining a gene alignment result according to the second alignment result and the third alignment result includes: according to the site mapping in the second alignment result and the site mapping in the third alignment result, merge the second alignment result and the third alignment result in a preset merging manner to obtain a gene alignment result. The specific form of merging the second alignment result and the third alignment result will be described in detail in the subsequent embodiments.
[0093] Exemplarily, the preset merging manner includes:
[0094] Initialize the gene alignment result to be empty;
[0095] Find the target base with the smallest site mapping in the second alignment result and the third alignment result, and write the target base and its alignment result in the second alignment result or the third alignment result into the gene alignment result;
[0096] Delete the target base from the alignment result where the target base is located;
[0097] Judge whether there is an insertion after the alignment result where the target base is located. If so, write the insertion into the gene alignment result as well;
[0098] Delete the insertion after the target base from the alignment result where the target base is located;
[0099] Return to execute the step of finding the target base with the smallest site mapping in the second alignment result and the third alignment result until all bases in the second alignment result and the third alignment result have been written into the gene alignment result.
[0100] After completing the sequence alignment and obtaining the gene alignment result, the following analysis methods can be further used: calculating the allele frequency of the nucleic acid sequence at each locus in a specific region of the genome from the gene alignment result, and / or searching for the nucleic acid sequence aligned to the reference genome from the gene alignment result and identifying the parts that are different from the reference genome. The above analysis can be performed using one or more bioinformatics software, including but not limited to bioinformatics software such as bcftools, MuTect, MuTect2, GATK, VarDict, VarScan, DeepVariant, etc.
[0101] In this gene variant analysis method, after performing two-color 2+2 fuzzy sequencing according to the first base combination, aligning the fuzzy sequence to be analyzed with the reference gene sequence to obtain the first alignment result, the reference sub-gene sequence is further selected from the reference gene sequence according to the first alignment result, and the fuzzy sequence to be analyzed is split according to the second base combination to obtain the first fuzzy semi-sequence to be analyzed and the second fuzzy semi-sequence to be analyzed. The reference sub-gene sequence is split to obtain the first reference sub-gene semi-sequence and the second reference sub-gene semi-sequence. Further, the first fuzzy semi-sequence to be analyzed is aligned with the corresponding first reference sub-gene semi-sequence to obtain the second alignment result, and the second fuzzy semi-sequence to be analyzed is aligned with the corresponding second reference sub-gene semi-sequence to obtain the third alignment result. Finally, the gene alignment result is obtained according to the second alignment result and the third alignment result. Through further alignment to obtain the second alignment result and the third alignment result, as well as the integration of the second alignment result and the third alignment result, it is possible to effectively exclude the possible misalignment results in the first alignment result caused by the influence of the coding method, and thus obtain the correct gene variant type, improving the accuracy in the case of mapping the relative position relationship of the sites where the SNV changes the base in two-color 2+2 fuzzy sequencing.
[0102] Optionally, the gene variant analysis method in the embodiments of the present disclosure further includes: re-extracting a new fuzzy sequence to be analyzed and a new reference sub-gene sequence from the gene alignment result; and aligning the new fuzzy sequence to be analyzed and the new reference sub-gene sequence to obtain the final alignment result. Through the above steps, the gene variant analysis method in the embodiments of the present disclosure can obtain accurate detection results even when the relative position relationship of the site mapping of the SNV gene variant does not change the base. In the actual application process, if the gene variant type and the influence result of the relative position relationship of the site mapping of the base are known, the process can be directly ended after obtaining the gene alignment result, or after obtaining the final alignment result. If the gene variant type and the influence result of the relative position relationship of the site mapping of the base are unknown, it is possible to determine whether further alignment is necessary according to the variant situation identified in the gene alignment result.
[0103] Among them, an example of the case where the relative positional relationship of the site mapping of the gene mutation that does not change the base is as follows:
[0104] Example 1: When detecting C>T by dual-color MK fuzzy sequencing, the DNA sequence of the wild-type sequence is ACTT, the site mapping is 1234, the encoded coding sequence is ACTT, and its corresponding site mapping is 1234. After the C>T mutation occurs, the DNA sequence of the mutant sequence is ATTT, its corresponding site mapping is 1234, the encoded coding sequence is ATTT, and its corresponding site mapping is 1234. The gene mutation does not change the relative positional relationship of the site mapping of the bases in the encoded coding sequence.
[0105] Example 2: When detecting C>A by dual-color MK fuzzy sequencing, the DNA sequence of the wild-type sequence is ACTT, the site mapping is 1234, the encoded coding sequence is ACTT, and its corresponding site mapping is 1234. After the C>A mutation occurs, the DNA sequence of the mutant sequence is AATT, its corresponding site mapping is 1234, the encoded coding sequence is AATT, and its corresponding site mapping is 1234. The gene mutation does not change the relative positional relationship of the site mapping of the bases in the encoded coding sequence.
[0106] Example 3: When detecting deletion by dual-color MK fuzzy sequencing, the DNA sequence of the wild-type sequence is AAGCC, the site mapping is 12345, the encoded coding sequence is AAGCC, and its corresponding site mapping is 12345. After the deletion mutation occurs, the DNA sequence of the mutant sequence is AACC, its corresponding site mapping is 1234, the encoded coding sequence is AACC, and its corresponding site mapping is 1234. The gene mutation does not change the relative positional relationship of the site mapping of the bases in the encoded coding sequence.
[0107] In addition, the embodiments of the present disclosure also provide a gene mutation analysis system. Specifically, as Figure 2 shown, the gene mutation analysis system includes:
[0108] A sequencing module 10 for performing dual-color 2+2 fuzzy sequencing on a nucleic acid sample according to a first base combination to obtain a sequencing signal of the nucleic acid sample;
[0109] An encoding module 20 for encoding the sequencing signal to obtain a fuzzy sequence to be analyzed and encoding a reference genome to obtain a reference gene sequence;
[0110] A first alignment module 30 for aligning the fuzzy sequence to be analyzed with the reference gene sequence to obtain a first alignment result;
[0111] A subsequence selection module 40, configured to select a reference sub-gene sequence from a reference gene sequence according to a first alignment result, where the base positions included in the reference sub-gene sequence overlap at least partially with the base positions included in the first alignment result;
[0112] A sequence splitting module 50, configured to split an analyzed fuzzy sequence according to a second base combination to obtain a first analyzed fuzzy semi-sequence and a second analyzed fuzzy semi-sequence, and split the reference sub-gene sequence according to the second base combination to obtain a first reference sub-gene semi-sequence and a second reference sub-gene semi-sequence;
[0113] A second alignment module 60, configured to align the first analyzed fuzzy semi-sequence with the corresponding first reference sub-gene semi-sequence to obtain a second alignment result, and align the second analyzed fuzzy semi-sequence with the corresponding second reference sub-gene semi-sequence to obtain a third alignment result;
[0114] A result integration module 70, configured to obtain a gene alignment result according to the second alignment result and the third alignment result.
[0115] Optionally, the result integration module 70 is specifically configured to merge the second alignment result and the third alignment result in a preset merging manner according to the site mapping in the second alignment result and the site mapping in the third alignment result to obtain a gene alignment result.
[0116] Optionally, the gene mutation analysis system further includes: a third alignment module, configured to re-extract a new analyzed fuzzy sequence and a new reference sub-gene sequence from the gene alignment result; and align the new analyzed fuzzy sequence and the new reference sub-gene sequence to obtain a final alignment result.
[0117] In the embodiments of the present disclosure, the specific details of each step are applicable to the corresponding module, and will not be elaborated here.
[0118] Embodiment 1
[0119] Taking the λ phage of New England Biolabs with a C>T substitution mutation as an example, the actual situation is as follows:
[0120]
[0121] Performing dual-color MK fuzzy sequencing on the λ phage of New England Biolabs to obtain at least one sequencing signal, where both A and G are labeled with a red fluorescent group, and both C and T are labeled with a green fluorescent group. Encoding one sequencing signal according to a preset encoding rule to obtain an analyzed fuzzy sequence, and the first alignment result after aligning to the reference gene sequence of the λ phage is:
[0122]
[0123] In the prior art mentioned in the background art, the above first comparison result is the final comparison result obtained by detecting SNV through two-color 2+2 fuzzy sequencing. It mistakenly believes that a complex mutation of AAACCC>CACTAA has occurred, which is obviously deviated from the actual situation and there is misidentification.
[0124] In the technical solution of the embodiment of the present disclosure, a reference sub-gene sequence of 5598-5635bp of the reference gene sequence (that is, the above-mentioned part of the reference gene sequence) TCGGTCCGACGAAACCCCGAGGTAACCGGCGCGACGGT is further taken. This reference sub-gene sequence completely covers the bases corresponding to the fuzzy sequence to be analyzed, and there is a certain margin before and after. The absolute sites of this reference sub-gene sequence are mapped to 5598, 5599, 5601, 5602, 5600, 5603, 5604, 5605, 5607, 5606, 5608, 5610, 5613, 5615, 5609, 5611, 5612, 5614, 5616, 5617, 5618, 5620, 5619, 5622, 5624, 5621, 5623, 5625, 5626, 5627, 5628, 5629, 5630, 5632, 5631, 5633, 5635, 5634. The relative site mapping obtained by subtracting 5598 from the absolute site mapping is 0, 1, 3, 4, 2, 5, 6, 7, 9, 8, 10, 12, 15, 17, 11, 13, 14, 16, 18, 19, 20, 22, 21, 24, 26, 23, 25, 27, 28, 29, 30, 31, 32, 34, 33, 35, 37, 36.
[0125] The fuzzy sequence to be analyzed is split according to the base combination RY to obtain a first fuzzy semi-sequence to be analyzed GGGAGAAAGAGGAAGGGGAG and a second fuzzy semi-sequence to be analyzed CCCCCTCTCCCCC. The reference sub-gene sequence is split according to the base combination RY to obtain a first reference sub-gene semi-sequence GGGAGAAAGAGGAAGGGGAGG and a second reference sub-gene semi-sequence TCTCCCCCCCTCCCCCT.
[0126] The first fuzzy semi-sequence to be analyzed is compared with the corresponding first reference sub-gene semi-sequence to obtain a second comparison result, and the second fuzzy semi-sequence to be analyzed is compared with the corresponding second reference sub-gene semi-sequence to obtain a third comparison result. The second comparison result and the third comparison result are shown in Table 1.
[0127] Table 1
[0128]
[0129]
[0130] According to the relative site mapping in the second alignment result and the relative site mapping in the third alignment result, the second alignment result and the third alignment result are merged according to a preset merging method, and the obtained gene alignment result is shown in Table 2.
[0131] Table 2
[0132]
[0133]
[0134] YY+231715P
[0135]
[0136] It can be seen from the gene alignment result that there is a substitution mutation of C>T in the fuzzy sequence to be analyzed relative to the reference gene sequence, which is consistent with the actual situation, that is, the 5612th base changes from C to T.
[0137] Example 2
[0138] Taking the λ phage of New England Biolabs with a substitution mutation of C>A as an example, two-color MK fuzzy sequencing is performed on the λ phage of New England Biolabs to obtain at least one sequencing signal, where both A and G are labeled with red fluorescent groups, and both C and T are labeled with green fluorescent groups.
[0139] A sequencing signal is encoded according to a preset coding rule to obtain a fuzzy sequence to be analyzed. The first alignment result after aligning to the reference gene sequence of the λ phage is:
[0140]
[0141] Take the subsequence of 5598-5635bp of the reference sequence
[0142] TCGGTCCGACGAAACCCCGAGGTAACCGGCGCGACGGT. The absolute site mapping of this subsequence is 5598, 5599, 5601, 5602, 5600, 5603, 5604, 5605, 5607, 5606, 5608, 5610, 5613, 5615, 5609, 5611, 5612, 5614, 5616, 5617, 5618, 5620, 5619, 5622, 5624, 5621, 5623, 5625, 5626, 5627, 5628, 5629, 5630, 5632, 5631, 5633, 5635, 5634. The relative site mapping obtained by subtracting 5598 from the absolute site mapping is 0, 1, 3, 4, 2, 5, 6, 7, 9, 8, 10, 12, 15, 17, 11, 13, 14, 16, 18, 19, 20, 22, 21, 24, 26, 23, 25, 27, 28, 29, 30, 31, 32, 34, 33, 35, 37, 36.
[0143] The fuzzy sequence to be analyzed is split according to the base combination RY to obtain the first fuzzy semi-sequence to be analyzed GGGAGAAAAGAGGAAGGGGAG and the second fuzzy semi-sequence to be analyzed TCCCCCCTCCCCC. The reference sub-gene sequence is split according to the base combination RY to obtain the first reference sub-gene semi-sequence GGGAGAAAGAGGAAGGGGAGG and the second reference sub-gene semi-sequence
[0144] TCTCCCCCCCTCCCCCT.
[0145] The first fuzzy semi-sequence to be analyzed is aligned with the corresponding first reference sub-gene semi-sequence to obtain the second alignment result. The second fuzzy semi-sequence to be analyzed is aligned with the corresponding second reference sub-gene semi-sequence to obtain the third alignment result. The second alignment result and the third alignment result are shown in Table 3.
[0146] Table 3
[0147]
[0148]
[0149] According to the relative site mapping in the second alignment result and the relative site mapping in the third alignment result, the second alignment result and the third alignment result are merged according to the preset merging method, and the gene alignment result obtained is shown in Table 4.
[0150] Table 4
[0151]
[0152]
[0153]
[0154] It can be seen that there is an insertion and a deletion in the gene alignment result.
[0155] At this time, in the above gene alignment result, it is necessary to re-extract the new fuzzy sequence to be analyzed TGGCCGCAGCACCAAAGAGTGCACAGGCGCGCAG, and, the new reference sub-gene sequence TCTGGCCGCAGCACCACAGAGTGCACAGGCGCGCAGTG. Compare the new fuzzy sequence to be analyzed TGGCCGCAGCACCAAAGAGTGCACAGGCGCGCAG and the new reference sub-gene sequence TCTGGCCGCAGCACCACAGAGTGCACAGGCGCGCAGTG to obtain the final alignment result shown below, so as to accurately reflect the C>A mutation.
[0156]
[0157] Example 3
[0158] During high-throughput sequencing, a sequencing library of the genomic DNA of λ phage from New England Biolabs is constructed and sequenced using two-color MK fuzzy sequencing, where both A and G are labeled with red fluorescent groups, and both C and T are labeled with green fluorescent groups. In the prior art, the sequencing signals are aligned to the encoded reference genome after encoding to obtain bam file 1. Using the method described in the embodiments of the present disclosure (Steps S1-Step S8), bam file 2 is obtained.
[0159] Open bam file 1 and bam file 2 in the IGV software (integrated genome viewer), where the reference genome selected for bam file 1 is the encoded λ phage genome, and the reference genome selected for bam file 2 is the non-encoded λ phage genome.
[0160] As Figure 3 shown, near the 9996th base site of the genome, a large number of sequences in bam file 1 show two mutations, A>T and C>A; in a small number of sequences, there is also a T>G mutation. And in bam file 2, as Figure 4As shown, a large number of sequences at the 9996th base of the genome showed a C>T mutation, and the existing T>G mutations in a small number of sequences were also corrected. In bam file 2, the C>T mutation frequency at the 9996th base of the genome was 100%, significantly higher than other sites. Therefore, it can be reported that a C>T mutation was detected here.
[0161] The basic principles of the present disclosure have been described above in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the specific details disclosed above are only for illustrative and facilitating understanding purposes and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0162] In the present disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used here refer to the word "and / or" and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used here refers to the phrase "such as but not limited to" and can be used interchangeably with each other.
[0163] In addition, as used herein, the "or" used in the listing of items starting with "at least one" indicates a separate listing, so that for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the term "exemplary" does not mean that the described examples are preferred or better than other examples.
[0164] It should also be noted that in the systems and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0165] Various changes, substitutions, and alterations to the technology described herein can be made without departing from the teachings defined by the appended claims. Additionally, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Accordingly, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.
[0166] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0167] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to the form disclosed herein. Although numerous example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A method for gene mutation analysis, characterized in that, Comprising: Performing dual-color 2+2 fuzzy sequencing on a nucleic acid sample according to a first base combination to obtain sequencing signals of the nucleic acid sample; Encoding the sequencing signals to obtain a fuzzy sequence to be analyzed, and encoding a reference genome to obtain a reference gene sequence; Comparing the fuzzy sequence to be analyzed with the reference gene sequence to obtain a first comparison result; Selecting a reference sub-gene sequence from the reference gene sequence according to the first comparison result, wherein the base positions included in the reference sub-gene sequence at least partially overlap with the base positions included in the first comparison result; Splitting the fuzzy sequence to be analyzed according to a second base combination to obtain a first fuzzy semi-sequence to be analyzed and a second fuzzy semi-sequence to be analyzed; Splitting the reference sub-gene sequence according to the second base combination to obtain a first reference sub-gene semi-sequence and a second reference sub-gene semi-sequence; Comparing the first fuzzy semi-sequence to be analyzed with the corresponding first reference sub-gene semi-sequence to obtain a second comparison result, and comparing the second fuzzy semi-sequence to be analyzed with the corresponding second reference sub-gene semi-sequence to obtain a third comparison result; Obtaining a gene comparison result according to the second comparison result and the third comparison result; The base combinations include MK, RY, and WS; the first base combination is one of the base combinations, and the second base combination is one of the other two of the base combinations; wherein, M represents bases A, C; K represents bases T, G; R represents bases A, G; Y represents bases C, T; W represents bases A, T; S represents bases C, G.
2. The gene mutation analysis method according to claim 1, characterized in that The base positions included in the first comparison result are the a-th to b-th bases of the reference gene sequence; The selecting a reference sub-gene sequence from the reference gene sequence according to the first comparison result includes: selecting the c-th to d-th bases from the reference gene sequence as the reference sub-gene sequence; wherein, a≤c≤b, or, c≤a≤d.
3. The gene mutation analysis method according to claim 2, wherein The values of c and d satisfy: covering the positions of known variations on the fuzzy sequence to be analyzed, and / or, covering the higher-quality part or the more reliable part in the first comparison result; the higher-quality part or the more reliable part refers to any one of the following: the part remaining after removing clip from the reference gene sequence; the part of the reference gene sequence where the length of homologous polymers and / or binary copolymers is less than 8; the part of the reference gene sequence where the average value or median of base quality values is greater than 15.
4. The gene mutation analysis method according to any one of claims 1 to 3, characterized in that, The base positions included in the first comparison result are the a-th to b-th bases of the reference gene sequence; The selecting a reference sub-gene sequence from the reference gene sequence according to the first comparison result includes: selecting the c-th to d-th bases from the reference gene sequence as the reference sub-gene sequence; wherein, c≤a, and b≤d.
5. The gene mutation analysis method according to claim 1, wherein Also comprising: Obtaining the site mapping of the first reference sub-gene semi-sequence and the second reference sub-gene semi-sequence; The second comparison result includes the base correspondence between the first fuzzy semi-sequence to be analyzed and the first reference sub-gene semi-sequence, and the corresponding site mapping; The third comparison result includes the base correspondence between the second fuzzy semi-sequence to be analyzed and the second reference sub-gene semi-sequence, as well as the corresponding site mapping; Obtaining a gene comparison result according to the second comparison result and the third comparison result includes: According to the site mapping in the second comparison result and the site mapping in the third comparison result, the second comparison result and the third comparison result are merged in a preset merging manner to obtain the gene comparison result.
6. The gene mutation analysis method according to claim 5, characterized in that, The preset merging manner includes: Initializing the gene comparison result to be empty; Finding the target base with the smallest site mapping in the second comparison result and the third comparison result, and writing the target base and its comparison result in the second comparison result or the third comparison result into the gene comparison result; Deleting the target base from the comparison result where the target base is located; Judging whether there is an insertion after the comparison result where the target base is located. If so, writing the insertion into the gene comparison result as well; Deleting the insertion after the target base from the comparison result where the target base is located; Returning to execute the step of finding the target base with the smallest site mapping in the second comparison result and the third comparison result until all bases in the second comparison result and the third comparison result have been written into the gene comparison result.
7. The gene mutation analysis method according to claim 6, wherein It also includes: Re-extracting a new fuzzy sequence to be analyzed and a new reference sub-gene sequence from the gene comparison result; Comparing the new fuzzy sequence to be analyzed and the new reference sub-gene sequence to obtain a final comparison result.
8. A gene mutation analysis system, characterized in that, It includes: A sequencing module for performing two-color 2+2 fuzzy sequencing on a nucleic acid sample according to a first base combination to obtain a sequencing signal of the nucleic acid sample; An encoding module for encoding the sequencing signal to obtain a fuzzy sequence to be analyzed, and encoding a reference genome to obtain a reference gene sequence; A first comparison module for comparing the fuzzy sequence to be analyzed with the reference gene sequence to obtain a first comparison result; A sub-sequence selection module for selecting a reference sub-gene sequence from the reference gene sequence according to the first comparison result, where the base positions included in the reference sub-gene sequence overlap at least partially with the base positions included in the first comparison result; A sequence splitting module for splitting the fuzzy sequence to be analyzed according to a second base combination to obtain a first fuzzy semi-sequence to be analyzed and a second fuzzy semi-sequence to be analyzed, and splitting the reference sub-gene sequence according to the second base combination to obtain a first reference sub-gene semi-sequence and a second reference sub-gene semi-sequence; A second comparison module for comparing the first fuzzy semi-sequence to be analyzed with the corresponding first reference sub-gene semi-sequence to obtain a second comparison result, and comparing the second fuzzy semi-sequence to be analyzed with the corresponding second reference sub-gene semi-sequence to obtain a third comparison result; A result integration module for obtaining a gene comparison result according to the second comparison result and the third comparison result; The base combinations include MK, RY, and WS; the first base combination is one of the base combinations, and the second base combination is one of the other two of the base combinations; wherein, M represents bases A and C; K represents bases T and G; R represents bases A and G; Y represents bases C and T; W represents bases A and T; S represents bases C and G.
9. The gene mutation analysis system according to claim 8, wherein It further includes: A third comparison module, configured to re-extract a new fuzzy sequence to be analyzed and a new reference sub-gene sequence from the gene comparison result; Compare the new fuzzy sequence to be analyzed and the new reference sub-gene sequence to obtain a final comparison result.
Citation Information
Patent Citations
Methods for long-range sequence analysis of nucleic acids
CN101072882A
Pair of polypeptides for specifically identifying sheep KRT25 gene, as well as encoding gene and application thereof
CN104250297A