Gene mutation analysis, query sequence establishment method, system, device and medium

CN120600108BActive Publication Date: 2026-09-08CHANGSHA KINGMED MEDICAL DIAGNOSTICS INST +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510684521.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2026-09-08
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

*.ab1文件包含峰图信息,*.seq文件是由每个位置峰值最高的信号所代表的碱基组成的序列,虽然可以用记事本打开,但是不能体现出次峰信息

Benefits of technology

[0041]The technical solution provided by this invention has the following advantages and effects: First and second selected bases are determined based on the peak data of forward and reverse sequencing data, resulting in corresponding forward and reverse selected sequences. The merged forward and reverse selected sequences are then compared with the original forward and reverse sequences, respectively, to obtain first and second analysis results. Based on the comparison results of the first and second analysis results, base substitutions in the original sequence are controlled to reduce the impact of improper sequencing operations or sample contamination on the accuracy of the query sequence, preventing the omission of positive mutations or the occurrence of false positive mutations due to inaccurate query sequences during mutation analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600108B_ABST
    Figure CN120600108B_ABST
Patent Text Reader

Abstract

The application discloses a gene mutation analysis method, system, device and storage medium, and technical scheme points thereof are as follows: reading forward and reverse sequencing data, obtaining first peak value data and second peak value data of each base, and determining a first selected base and a second selected base based on the first peak value data and the second peak value data, and obtaining a forward selected sequence and a forward undetermined position; performing alignment analysis on the merged forward selected sequence and reverse selected sequence, the forward original sequence and the reverse original sequence, and controlling base replacement of the forward original sequence and the reverse original sequence based on the alignment analysis result, and obtaining a target forward sequencing sequence and a target reverse sequencing sequence; and judging whether to merge the forward original sequence and the reverse original sequence on the merged target forward sequencing sequence and the target reverse sequencing sequence based on the base replacement condition. The application can reduce the influence of improper operation or sample pollution on the accuracy of a query sequence in sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical technology, specifically relating to a method, system, device, and medium for gene mutation analysis and sequence lookup. Background Technology

[0002] Currently, the main method for detecting gene mutations is to sequence genes using Sanger sequencing technology. Sanger sequencing (first-generation DNA sequencing technology) was the first method applied to human genome sequencing and is known for its high accuracy; it remains the "gold standard" of existing sequencing technologies. The simple workflow, ease of use, and relatively low cost of Sanger sequencing make it uniquely advantageous in specific applications, such as single-gene genetic disease detection, precise screening of susceptibility genes, and mutation verification.

[0003] Sanger sequencing data is stored in *.seq and *.ab1 files. The *.ab1 file contains peak plot information, while the *.seq file is a sequence composed of the bases represented by the highest signal at each position. Although it can be opened with Notepad, it does not show secondary peak information. When improper sequencing operations or sample contamination result in multiple peaks at a single position, or even when the true peak height is lower than the contaminating peak, using the *.seq file as the query sequence will fail to determine the true base information at that position, leading to inaccurate query sequences. This can cause missed positive mutations or false positives during subsequent gene mutation detection. Summary of the Invention

[0004] The purpose of this invention is to provide a method, system, device and medium for gene mutation analysis and query sequence establishment, which reduces the impact of improper sequencing operations or sample contamination on the accuracy of query sequences, and prevents the omission of positive mutations or the occurrence of false positive mutations due to inaccurate query sequences in mutation analysis.

[0005] The first aspect of this invention provides a method for establishing a query sequence, comprising:

[0006] Read the bidirectional sequencing data to obtain the forward raw sequence and the first peak data for each base at each position, and obtain the reverse raw sequence and the second peak data for each base at each position;

[0007] Based on the first peak data, the first selected base at each position is determined to obtain the forward selected sequence and the forward undetermined position. Based on the second peak data, the second selected base at each position is determined to obtain the reverse selected sequence and the reverse undetermined position.

[0008] The forward selection sequence and the reverse selection sequence are merged to obtain the query sequence. The query sequence, the forward original sequence, and the reverse original sequence are compared and analyzed to obtain the target comparison result.

[0009] Based on the target alignment results, the bases at the forward undetermined positions of the forward original sequence are replaced to obtain the target forward sequencing sequence. Based on the target alignment results, the bases at the reverse undetermined positions of the reverse original sequence are replaced to obtain the target reverse sequencing sequence.

[0010] The target forward sequencing sequence and the target reverse sequencing sequence are merged to obtain the first target query sequence. Based on the base substitution, it is determined whether to merge the forward original sequence and the reverse original sequence on the first target query sequence to obtain the second target query sequence.

[0011] In some embodiments, a selected base at each position is determined based on peak data to obtain a selected sequence and a pending position. When the peak data is a first peak data, the selected base is a first selected base, the selected sequence is a forward selected sequence, and the pending position is a forward pending position. When the peak data is a second peak data, the selected base is a second selected base, the selected sequence is a reverse selected sequence, and the pending position is a reverse pending position, including:

[0012] Divide the peak value of the second highest peak at each location by the peak value of the highest peak to obtain the peak ratio.

[0013] Determine whether the peak ratio is greater than or equal to the ratio threshold. If yes, the base corresponding to the second peak is selected as the base; otherwise, the base corresponding to the highest peak is selected as the base.

[0014] The selected sequence is obtained by arranging all the selected bases in order of position, and the positions with a peak ratio greater than or equal to the ratio threshold are designated as undetermined positions.

[0015] In some embodiments, the step of performing comparative analysis on the query sequence, the forward original sequence, and the reverse original sequence to obtain the target comparison result includes:

[0016] The query sequence is compared with the forward original sequence to obtain the first alignment result; the query sequence is compared with the reverse original sequence to obtain the second alignment result.

[0017] The first comparison result is obtained by analyzing the first comparison result, the second comparison result is obtained by analyzing the second comparison result, and the target comparison result is obtained by comparing the first comparison result and the second comparison result.

[0018] In some embodiments, the analysis results are obtained by analyzing the comparison results. When the comparison result is a first comparison result, the analysis result is a first analysis result; when the comparison result is a second comparison result, the analysis result is a second analysis result, including:

[0019] The comparison results are sorted, and an index is created on the sorted comparison results to obtain the comparison index data.

[0020] Anomaly analysis was performed on the comparison index data to obtain the original analysis results;

[0021] The positions that meet the preset filtering conditions are selected from the original analysis results to obtain the analysis results. The preset filtering conditions include: both alleles at the same position are homozygous and both have the same base type as the query sequence; the sequence coverage of the position is greater than or equal to a first preset threshold; and the probability of an anomaly at the position is greater than or equal to a second preset threshold.

[0022] In some embodiments, the comparison of the first analysis result and the second analysis result to obtain the target comparison result includes:

[0023] The first and second analysis results at the same site are compared. If the bases in the first and second analysis results are complementary, the target alignment result with the second peak being a true signal is obtained. If the bases in the first and second analysis results are not complementary, the target alignment result with the second peak being a false signal is obtained.

[0024] In some embodiments, the step of controlling the substitution of bases at the undetermined positions of the forward original sequence according to the target alignment result to obtain the target forward sequencing sequence includes:

[0025] If the target alignment result corresponding to the positive undetermined position is a true signal of the second peak, the bases at the positive undetermined position are replaced with the bases represented by the second peak to obtain the target positive sequencing sequence.

[0026] The step of controlling the substitution of bases at the undetermined reverse positions of the reversed original sequence based on the target alignment result to obtain the target reverse sequencing sequence includes:

[0027] If the target alignment result corresponding to the reverse undetermined position is a true signal at the second peak, the bases at the reverse undetermined position are replaced with the bases represented by the second peak to obtain the target reverse sequencing sequence.

[0028] A second aspect of the present invention provides a method for gene mutation analysis, comprising:

[0029] Select the corresponding reference sequence based on the sequencing species;

[0030] The second target query sequence is established using the query sequence establishment method described above for the sequencing data of the sequenced species;

[0031] The second target query sequence and the reference sequence are compared to obtain comparison data;

[0032] All mutation sites were obtained by analyzing the alignment data.

[0033] A third aspect of the present invention provides a query sequence establishment system, comprising:

[0034] The data reading module is used to read bidirectional sequencing data, obtain the forward raw sequence and the first peak data of each base at each position, and obtain the reverse raw sequence and the second peak data of each base at each position;

[0035] The base selection module is used to determine the first selected base at each position based on the first peak data to obtain the forward selected sequence and the forward undetermined position, and to determine the second selected base at each position based on the second peak data to obtain the reverse selected sequence and the reverse undetermined position;

[0036] The analysis and comparison module is used to merge the forward selected sequence and the reverse selected sequence to obtain the query sequence, and to perform comparison and analysis on the query sequence, the forward original sequence, and the reverse original sequence to obtain the target comparison result;

[0037] The base substitution module is used to control the substitution of bases at the forward undetermined positions of the forward original sequence according to the target alignment result, to obtain the target forward sequencing sequence, and to control the substitution of bases at the reverse undetermined positions of the reverse original sequence according to the target alignment result, to obtain the target reverse sequencing sequence.

[0038] The sequence merging module is used to merge the target forward sequencing sequence and the target reverse sequencing sequence to obtain a first target query sequence. Based on the base substitution situation, it is determined whether to merge the forward original sequence and the reverse original sequence on the first target query sequence to obtain a second target query sequence.

[0039] A fourth aspect of the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the query sequence establishment method described above.

[0040] The fifth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the query sequence establishment method described above.

[0041] The technical solution provided by this invention has the following advantages and effects: First and second selected bases are determined based on the peak data of forward and reverse sequencing data, resulting in corresponding forward and reverse selected sequences. The merged forward and reverse selected sequences are then compared with the original forward and reverse sequences, respectively, to obtain first and second analysis results. Based on the comparison results of the first and second analysis results, base substitutions in the original sequence are controlled to reduce the impact of improper sequencing operations or sample contamination on the accuracy of the query sequence, preventing the omission of positive mutations or the occurrence of false positive mutations due to inaccurate query sequences during mutation analysis. Attached Figure Description

[0042] Figure 1 This is a flowchart illustrating the query sequence establishment method provided by the present invention;

[0043] Figure 2 It refers to the mutation site information detected using the gene mutation analysis method provided by this invention.

[0044] Figure 3 This is a structural block diagram of the query sequence establishment system provided by the present invention;

[0045] Figure 4 This is an internal structural diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0046] To facilitate understanding of the present invention, specific embodiments of the present invention will be described in more detail below with reference to the accompanying drawings.

[0047] Unless otherwise specified or defined, the terms "first," "second," etc., used in this document are for distinguishing names only and do not represent a specific number or order.

[0048] Unless otherwise stated or defined, the term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.

[0049] It should be noted that in this article, "fixed to" or "connected to" can mean directly fixed to or connected to a component, or indirectly fixed to or connected to a component.

[0050] like Figure 1 As shown, this embodiment provides a method for establishing a query sequence, including the following steps S1 to S5:

[0051] Step S1: Read the bidirectional sequencing data to obtain the forward original sequence and the first peak data of each base at each position, and to obtain the reverse original sequence and the second peak data of each base at each position.

[0052] In practical applications, bidirectional sequencing data is obtained by inputting *.seq and *.ab1 files from Sanger sequencing data. *.seq is a plain text file (FASTA format) containing only the base sequence of the sequencing results, while *.ab1 is a binary file directly generated by the sequencer, containing raw sequencing data (such as electrophoresis peaks, base sequences, and quality values). The bidirectional sequencing data includes: forward raw sequencing file F.ab1, forward raw sequence file F.seq, reverse raw sequencing file R.ab1, and reverse raw sequence file R.seq. The `readsangerseq` function of the R package `sangerseqR` is used to read F.ab1, outputting the first peak data for each base at each position. Reading F.seq yields the forward raw sequence. Similarly, the `readsangerseq` function of the R package `sangerseqR` is used to read R.ab1, outputting the second peak data for each base at each position. Reading R.seq yields the reverse raw sequence.

[0053] Step S2: Determine the first selected base at each position based on the first peak data to obtain the forward selected sequence and the forward undetermined position; determine the second selected base at each position based on the second peak data to obtain the reverse selected sequence and the reverse undetermined position.

[0054] Specifically, based on the peak data, selected bases at each position are determined to obtain selected sequences and undetermined positions. When the peak data is a first peak data point, the selected base is a first selected base, the selected sequence is a forward selected sequence, and the undetermined position is a forward undetermined position. When the peak data is a second peak data point, the selected base is a second selected base, the selected sequence is a reverse selected sequence, and the undetermined position is a reverse undetermined position, including:

[0055] Divide the peak value of the second highest peak at each location by the peak value of the highest peak to obtain the peak ratio.

[0056] Determine whether the peak ratio is greater than or equal to the ratio threshold. If yes, the base corresponding to the second peak is selected as the base; otherwise, the base corresponding to the highest peak is selected as the base.

[0057] The selected sequence is obtained by arranging all the selected bases in order of position, and the positions with a peak ratio greater than or equal to the ratio threshold are designated as undetermined positions.

[0058] In practical applications, each position corresponds to four bases: A, T, C, and G. Each base at each position has a corresponding peak data. Under normal circumstances, each position has only one obvious peak. Under abnormal circumstances, each position has at least two peaks of similar height, indicating that improper operation or sample contamination during sequencing has led to multiple peaks at one position, or even that the true peak height is lower than the contaminated peak. This method takes the highest peak at each position as the highest peak and the second highest peak at each position as the second highest peak. The ratio threshold is set to 0.2. 0.2 is obtained from testing data from multiple samples to prevent the value from being too high and missing positive sites, and the value from being too low and having many false positive sites. When the peak ratio at each position is greater than or equal to 0.2, the second peak is considered to be a true signal. In this case, the base corresponding to the second peak is selected as the selected base to facilitate further analysis of that position. This reduces the impact of improper sequencing operation or sample contamination on the accuracy of the query sequence and prevents the omission of positive mutations or the occurrence of false positive mutations due to inaccurate query sequences in mutation analysis. By designating positions where the peak ratio is greater than or equal to the ratio threshold as undetermined positions, it becomes easier to replace the bases at these undetermined positions in the future.

[0059] Step S3: Merge the forward selected sequence and the reverse selected sequence to obtain the query sequence. Perform comparative analysis on the query sequence, the forward original sequence, and the reverse original sequence to obtain the target comparison result.

[0060] Specifically, the comparison analysis of the query sequence, the forward original sequence, and the reverse original sequence to obtain the target comparison result includes:

[0061] The query sequence is compared with the forward original sequence to obtain the first alignment result; the query sequence is compared with the reverse original sequence to obtain the second alignment result.

[0062] The first comparison result is obtained by analyzing the first comparison result, the second comparison result is obtained by analyzing the second comparison result, and the target comparison result is obtained by comparing the first comparison result and the second comparison result.

[0063] In practical applications, the tool bwa is used to align the query sequence with the forward original sequence, and also with the reverse original sequence, obtaining first and second alignment results. These first and second alignment results are then analyzed to obtain corresponding first and second analysis results. The analysis results include the base changes between the original and query sequences at the same site. The base changes in the forward and reverse alignments are then compared. Based on the comparison results of the base changes in the forward and reverse alignments, it is determined whether the secondary peak represents a true signal. Since the highest peak in the forward and reverse original sequences is considered the true signal, while the forward and reverse selected sequences replace the bases represented by the secondary peak, which may represent the true signal, by merging the forward and reverse selected sequences and then aligning them to the forward or reverse original sequence, base substitution is performed at positions where a secondary peak simultaneously appears as a true signal in both forward and reverse sequencing, thereby improving substitution accuracy.

[0064] Specifically, the analysis results are obtained by analyzing the comparison results. If the comparison result is a first comparison result, the analysis result is the first analysis result; if the comparison result is a second comparison result, the analysis result is the second analysis result, including:

[0065] The comparison results are sorted, and an index is created on the sorted comparison results to obtain the comparison index data.

[0066] Anomaly analysis was performed on the comparison index data to obtain the original analysis results;

[0067] The positions that meet the preset filtering conditions are selected from the original analysis results to obtain the analysis results. The preset filtering conditions include: both alleles at the same position are homozygous and both have the same base type as the query sequence; the sequence coverage of the position is greater than or equal to a first preset threshold; and the probability of an anomaly at the position is greater than or equal to a second preset threshold.

[0068] In practical applications, the alignment results are sorted according to the positional order of the bases. The samtools tool is used to create an index of the sorted alignment results, resulting in alignment index data. The bcftools tool is then used to analyze the index data to obtain the raw analysis results. To improve the accuracy of the analysis results, preset filtering conditions are used to filter the raw analysis results, selecting positions that meet the preset filtering criteria. In the preset filtering conditions, both alleles are homozygous and have the same base type as the query sequence, which can be represented as GT = 1 / 1. GT represents the type of the two alleles carried by the site, 0 represents the base type of the original sequence, 1 represents the base type of the query sequence, and 1 / 1 represents a homozygous sequence where both alleles have the same base type as the query sequence. The sequence coverage of the position is greater than or equal to the first preset threshold, which is set to 2, and can be represented as DP>=2. The two filtering conditions, GT = 1 / 1 and DP>=2, ensure that both the forward and reverse selected sequences in the query sequence are aligned with the reference sequence and have the same base type as the query sequence. The probability of anomaly at the position is greater than or equal to the second preset threshold, where QUAL represents the probability of anomaly, i.e., the possibility of variation at the site. The higher the score, the more reliable the site is considered. The second preset threshold is set to 20, which can be represented as QUAL>=20, ensuring the reliability of the selected positions.

[0069] Specifically, the comparison of the first analysis result and the second analysis result to obtain the target comparison result includes:

[0070] The first and second analysis results at the same site are compared. If the bases in the first and second analysis results are complementary, the target alignment result with the second peak being a true signal is obtained. If the bases in the first and second analysis results are not complementary, the target alignment result with the second peak being a false signal is obtained.

[0071] In practical applications, if the first analysis result at site A is a change from base A to base C, and the second analysis result is a change from base T to base G, it indicates that the bases in the first and second analysis results at site A are complementary, and the secondary peak is a true signal. If the bases in the first and second analysis results are not complementary, such as at site B, the first analysis result is a change from base A to base C, and the second analysis result is a change from base T to base A, it indicates that the secondary peak at site B is a false signal. This prevents the omission of positive mutations or the occurrence of false positive mutations due to improper sequencing operations or sample contamination.

[0072] Step S4: Based on the target alignment result, control the substitution of bases at the forward undetermined positions of the forward original sequence to obtain the target forward sequencing sequence. Based on the target alignment result, control the substitution of bases at the reverse undetermined positions of the reverse original sequence to obtain the target reverse sequencing sequence.

[0073] Specifically, the step of controlling the substitution of bases at the forward undetermined positions of the forward original sequence according to the target alignment result to obtain the target forward sequencing sequence includes: when the target alignment result corresponding to the forward undetermined position is a true signal (second peak), replacing the bases at the forward undetermined position with the bases represented by the second peak to obtain the target forward sequencing sequence. The step of controlling the substitution of bases at the reverse undetermined positions of the reverse original sequence according to the target alignment result to obtain the target reverse sequencing sequence includes: when the target alignment result corresponding to the reverse undetermined position is a true signal (second peak), replacing the bases at the reverse undetermined position with the bases represented by the second peak to obtain the target reverse sequencing sequence.

[0074] In practical applications, if the first analysis result at site A shows a change from base A to base C, and the second analysis result shows a change from base T to base G, and the secondary peak is a true signal, then the forward undetermined position represents the location of site A in the forward original sequence, where the original base is A. This base A is replaced with base C. The reverse undetermined position represents the location of site A in the reverse original sequence, where the original base is T. This base T is replaced with base G. If the target alignment result shows a false signal for the secondary peak, such as the secondary peak corresponding to site B, then the forward undetermined position represents the location of site B in the forward original sequence, and the original base at that position is retained. The reverse undetermined position represents the location of site B in the reverse original sequence, and the original base at that position is retained. By determining whether the secondary peak is a true signal, and if so, replacing the corresponding bases in the forward and reverse original sequences when the secondary peak is a true signal, we can prevent missing positive mutations or causing false positive mutations due to improper sequencing operations or sample contamination.

[0075] Step S5: Merge the target forward sequencing sequence and the target reverse sequencing sequence to obtain the first target query sequence. Determine whether to merge the forward original sequence and the reverse original sequence on the first target query sequence based on the base substitution situation to obtain the second target query sequence.

[0076] In practical applications, the target forward and reverse sequencing sequences are merged to obtain the first target query sequence. If base substitutions are performed at the target positions, the original forward and reverse sequences need to be merged into the first target query sequence, which is then used as the second target query sequence. Even after substituting bases at positions where a secondary peak might be a true signal, it cannot be guaranteed that the resulting secondary peak is a true signal, or that the highest peak is necessarily contamination or a sequencing error. Therefore, the original sequences are merged for analysis to reduce sequencing errors. If base substitutions are not performed at the target positions, there is no need to merge the original forward and reverse sequences into the first target query sequence; the first target query sequence without merging the original forward and reverse sequences is used as the second target query sequence. By merging the forward and reverse sequences, it is easier to perform forward and reverse sequencing of the second target query sequence based on the reference sequence, improving coverage integrity and sequencing data accuracy.

[0077] This embodiment uses Demo as an example to establish the query sequence for the 7 exons of the GLA gene:

[0078] (1) Read the Sanger sequencing data to obtain the forward raw sequence and the first peak data for each base at each position; read the reverse sequencing data to obtain the reverse raw sequence and the second peak data for each base at each position:

[0079] 1.1 Obtain Sanger sequencing data of the GLA gene from sample Demo. The sample names of the sequencing data for the 7 exons are shown in the table below:

[0080] EXON1 EX1_F.ab1, EX1_F.seq EX1_R.ab1, EX1_R.seq EXON2 EX2_F.ab1, EX2_F.seq EX2_R.ab1, EX2_R.seq EXON3 EX3_F.ab1, EX3_F.seq EX3_R.ab1, EX3_R.seq EXON4 EX4_F.ab1, EX4_F.seq EX4_R.ab1, EX4_R.seq EXON5-EXON6 EX56_F.ab1, EX56_F.seq EX56_R.ab1, EX56_R.seq EXON7 EX7_F.ab1, EX7_F.seq EX7_R.ab1, EX7_R.seq

[0081] 1.2 The `readsangerseq` function of the R package `sangerseqR` is used to read the forward sequencing result files EX1_F.ab1, EX2_F.ab1, EX3_F.ab1, EX4_F.ab1, EX56_F.ab1, and EX7_F.ab1 for each exon, respectively, and outputs the peak data for each base at each position to obtain the first peak data for each exon. The `readsangerseq` function of the R package `sangerseqR` is used to read the reverse sequencing result files EX1_R.ab1, EX2_R.ab1, EX3_R.ab1, EX4_R.ab1, EX56_R.ab1, and EX7_R.ab1 for each exon, respectively, and outputs the peak data for each base at each position to obtain the second peak data for each exon.

[0082] (2) Divide the peak value of the second-highest peak at each position by the peak value of the highest peak to obtain the peak ratio. If the peak ratio is greater than or equal to 0.2, the second-highest peak is considered to be a true signal. When the peak ratio is less than 0.2, the base represented by the highest peak is selected as the selected base. When the peak ratio is greater than or equal to 0.2, i.e., it may be a true signal, the base represented by the second-highest peak is selected as the selected base. The selected bases at each position are sequentially sorted to form a complete sequence, thus obtaining the forward and reverse selection sequences in FASTA format. The results are as follows:

[0083] Positive selection sequence:

[0084] EX1_F_secondaryseq

[0085] GCTCTCGGTGATGGTCTGCCCCTGAGGTTACTCTTATAAGCCCAGGTTACCCGCGGAGATTTATGCTGTCCGGTTACCGTGACAATGCAGCTGAGGAACCCAGAACTACATCTGGGCTGCGCGCTTGCGCTTCGCTTCCTGGCCCTCGTTTCCTGGGACATCCCTGGGGCTAGAGCACTGGACAATGGATTGGCAAGGACGCCTACCATGGGCTGGCTGCACTG GGAGCGCTTCATGTGCAACCTTGACTGCCAGGAAGAGCCAGATTCCTGCATCAGGTATCAGATATTGGGTACTCCCTTCCCTTTGCTTTTCCATGTGTTTGGGTGTGTTTGGGGAACTGGAGAGTCTCAACGGGAACAGTTGAGCCCGAGGGAGAGCTCCCCCACCCGACTCTGCTGCTGCTTTTTTATCCCCAGCAAACTGTCCCGAATCAGGGGTCAAGTTG

[0086] EX2_F_secondaryseq

[0087] CAGTCATGCTGTCTGCTGATCTGACTCTTGACCTCAGGTGATCCACCCGCCTCGGCCTCCCAAAGTGTTGGGATTACAGGCGTGAGCCACCACGCCCGGCCATGAGGGCTGTTTCTAAACAAGCTTCTGTACAGAAGTGCTTACAGTCCTCTGAATGAACAAGAACATTATCTATAAACTCACATAATTAGCTAGCTGGCGAATCCCATGAGGAAAGCGCTGAGGGTCTGCCTGAAGTCTGCCTTCTGAATCTCTTTGGGGAGCCATCCAACAGTCATCAATGCAGAGGTACTCATAACCTGCATCCTTCCAGCCTTCTGAGACCATGAGCTCTGCCATCTCCATGAAGAGCTTCTCACTGAAAGAGAAATTCCAATAATCATTACAATTCATTAAATGAACACTTAGGTACCTCCCATTTATTAGGCACCTTGGGATTTCAGGTCATACCTGTTTCCTGA

[0088] >EX3_F_secondaryseq

[0089] CGTATGTCTCTCGTTCTGCTACCTCACGATTGTGCTTCTACTATGGTGACTCTTATTCCTCCCTCTCATTTCAGGTTCACAGCAAAGGACTGAAGCTAGGGATTTATGCAGATGTTGGAAATAAAACCTGCGCAGGCTTCCCTGGGAGTTTTGGATACTACGACATTGATGCCCAGACCTTTGCTGACTGGGGAGTAGATCTGCTAAAATTTGATGGTTGTTACTGTGACAGTTTGGAAAATTTGGCAGATGGTAATGTTTCATTCCAGAGATTTAGCCACAAAGGAAAGAACTTTGAGGCCATGGTAGCTGAGCCAAAGAACCAATCTTCGGTCATAGTGGTTTCCTGA

[0090] >EX4_F_secondaryseq

[0091] AGCAGCTGATCTTTCTTTCTCTTATTTTACCCATTGTTTTCTCATACAGGTTATAAGCACATGTCCTTGGCCCTGAATAGGACTGGCAGAAGCATTGTGTACTCCTGTGAGTGGCCTCTTTATATGTGGCCCTTTCAAAAGGTGAGATAGTGAGCCCAGAATCCAATAGAACTGTACTGATAGATAGAACTTGACAACAAAGGAAACCAAGGTCTCCTTCAAAGTCCAACGTTACTTACTATCATCCTACCATCTCTCCCAGGTTCCAACCACTTCTCACCATCCCCACTGGGTCATAGCTTGTTTCCTG

[0092] >EX56_F_secondaryseq

[0093] CTGAGACGTTCATCTGTAACGCGTYATAGGCTACTAGTGGCTCCTTTATACTGGTTTCTTCTCGCAAGGATGTTACTATAAAGAAGACAGAAGAGGCNTATCTGTTTTCACAGCCCAATTATACAGAAATCCGACAGTACTGCAATCACTGGCGAAATTTTGCTGACATTGATGATTCCTGGAAAAGTATAAAGAG TATCTTGGACTGGACATCTTTTAACCAGGAGAGAATTGTTGATGTTGCTGGACCAGGGGGTTGGAATGACCCAGATATGGTAAAAACTTGAGCCCTCCTTGTTCAAGACCCTGCGGTAGGCTTGTTTCCTATTTTGACATTCAAGGTAAATACAGGTAAAGTTCCTGGGAGGAGGCTTTATGTGAGAGTACTTAGA GCAGGATGCTGTGGAAAGTGGTTTCTCCATATGGGTCATCTAGGTAACTTTAAGAATGTTTCCTCCTCTCTTGTTTGAATTATTTCATTCTTTTTCTCAGTTAGTGATTGGCAACTTTGGCCTCAGCTGGAATCAGCAAGTAACTCAGATGGCCCTCTGGGCTATCATGGCTGCTCCTTTATTCATGTCTAATGACCTCCGACACATCAGCCCTCAAGCCAAAGCTCTCCTTCAGGATAAGGACGTAATTGCCATCAATCAGGACCCCTTGGGCAAGCAAGGGTACCAGCTTAGACAGGTAAATAAGAGTATATATTTTAAGATGGCTTTATATACCCAATACCAACTTTGTCTTGGGCCTAATCTATTTTGGTCTAACCCTGGTTTCCTGA

[0094] >EX7_F_secondaryseq

[0095] AGACGTCCAGTCCGACTCATGTGATCTGCCCGCCTCAGCCTCCCAAAGTGCTGGGATTACAGGCATGAGCCACCTAGCCTTGAGCTTTTAAAGTGAATGGAGAAAAAGGTGGACAGGAAGTAGTAGTTGGCAATAAAATAAACATTTTAAAGTAAGTCTTTTAATGACATCTGCATTGTATTTTCTAGCTGAAGCAAAACAGTGCCTGTGGGATTTATGTGACTTCTTAACCTTGAAGTCCATTCATAGAACCCTAGCTTCCTTTTCACAGGGAGGAGCTGTGTGATGAAGCAGGCAGGATTACAGGCCACTCCTTTACCCAGGGAAGCAACTGCGATGGTATAAGAGCGAGGTCCACCAATCTCCTGCCGGTTTATCATAGCTACAGCCCAGGCTAAGCCTGAGAGAGGTCGTTCCCACACTTCAAAGTTGTCTCCCTGAAAAACCAAGAAAGTGTGATTGCTTAGCAACTAGTGATAAGTGGCCCTGTTAGTTTGGCATTCATTCTGGTCAAGCTGGGGTTTCCGGGA

[0096] Reverse selected sequence:

[0097] >EX1_R_secondaryseq

[0098] ATGAGTGCTGGCGAACATAGCAGCAGCAGAGTCGGGTGGGGGAGCTCTCCCTCGGGCTCAACTGTTCCCGTTGAGACTCTCCAGTTCCCCAAACACACCCAAACACATGGAAAAGCAAAGGGAAGGGAGTACCCAATATCTGATACCTGATGCAGGAATCTGGCTCTTCCTGGCAGTCAAGGTTGCACATGAAGCGCTCCCAGTGCAGCCAGCCCATGGTAGGCGTCCTTGCCAATCCATTGTCCAGTGCTCTAGCCCCAGGGATGTCCCAGGAAACGAGGGCCAGGAAGCGAAGCGCAAGCGCGCAGCCCAGATGTAGTTCTGGGTTCCTCAGCTGCATTGTCACGGTAACCGGACAGCATAAATTTCCGCGGGTAACCTGGGCTTTTAAGATTAACCTCAGGGGCGGACCAATCACCGATGACTTATTAAATAAACTGGCCGCGGTTTTACAA

[0099] >EX2_R_secondaryseq

[0100] CTGTGCTATAGTGGGAGGTACCTACGTGTTCATTTAATGAATTGTAATGATTATTGGAATTTCTCTTTCAGTGAGAAGCTCTTCATGGAGATGGCAGAGCTCATGGTCTCAGAAGGCTGGAAGGATGCAGGTTATGAGTACCTCTGCATTGATGACTGTTGGATGGCTCCCCAAAGAGATTCAGAAGGCAGACTTCAGGCAGACCCTCAGCGCTTTCCTCATGGGATTCGCCAGCTAGCTAATTATGTGAGTTTATAGATAATGTTCTTGTTCATTCAGAGGACTGTAAGCACTTCTGTACAGAAGCTTGTTTAGAAACAGCCCTCATGGCCGGGCGTGGTGGCTCACGCCTGTAATCCCAACACTTTGGGAGGCCGAGGCGGGTGGATCACCTGAGGTCAAGAGTTCAAGACCAGCCTGGCCAACATGGTGAAACCCCAACTCACTGGCCGTCGTTTTACAA

[0101] >EX3_R_secondaryseq

[0102] GTTCTGCTCGCTACATGGCCTCTCAGTTCTTTCCTTTGTGGCTAAATCTCTGGAATGAAACATTACCATCTGCCAAATTTTCCAAACTGTCACAGTAACAACCATCAAATTTTAGCAGATCTACTCCCCAGTCAGCAAAGGTCTGGGCATCAATGTCGTAGTATCCAAAACT CCCAGGGAAGCCTGCGCAGGTTTTATTTCCAACATCTGCATAAATCCCTAGCTTCAGTCCTTTGCTGTGAACCTGAAATGAGAGGGAGGAAAAGAGTCACCATTGTAGAAGCACAATCGTGAGGTAGCAGAAAGAGAGAACCATTCCAGGCTGACTGGCCGTCGTTTTACAA

[0103] >EX4_R_secondaryseq

[0104] AGTACTGAAGTGCTGTACTGGGAGAGTGGTAGGATGATAGTAAGTAACGTTGGACTTTGAAGGAGACCTTGGTTTCCTTTGTTGTCAAGTTCTATCTATCAGTACAGTTCTATTGGATTCTGGGCTCACTATCTCACCTTTTGAAAGGGCCACATATAAAGAGGCCACTCACAGGAGTACACAATGCTTCTGCCAGTCCTATTCAGGGCCAAGGACATGTGCTTATAACCTGTATGAGAAAACAATGGGTAAAATAAGGGAAAGAAATGAATTTCCAGCTGGGGCTATATATAGTACTGGCCGTCGTTTTACAA

[0105] >EX56_R_secondaryseq

[0106] GATGCAGACAGTTGGTATTGGGTATATAAAGCCATCTTAAAATATATACTCTTATTTACCTGTCTAAGCTGGTACCCTTGCTTGCCCAAGGGGTCCTGATTGATGGCAATTACGTCCTTATCCTGAAGGAGAGCTTTGGCTTGAGGGCTGATGTGTCGGAGGTCATTAGACATGAATAAAGGAGCAGCCATGATAGCCCAGAGGGCCATCTGAGTTACTTGCTGATTCCAGCTGAGGCCAAAGTTGCCAATCACTAACTGAGAAAAAGAATGAAATAATTCAAACAAGAGAGGAGGAAACATTCTTAAAGTTACCTAGATGACCCATATGGAGAAACCACTTTCCACAGCATCCTGCTCTAAGTACTCTCACATAAAGCCTCCTCCCAGGAACTTTACCTGTATTTACCTTGAATGTCAAAATAGGAAACAAGCCTACCGCAGGGTCTTGAACAAGGAGGGCTCAAGTTTTTACCATATCTGGGTCATTCCAACCCCCTGGTCCAGCAACATCAACAATTCTCTCCTGGTTAAAAGATGTCCAGTCCAAGATACTCTTTATACTTTTCCAGGAATCATCAATGTCAGCAAAATTTCGCCAGTGATTGCAGTACTGTCGGATTTCTGTATAATTGGGCTGTGAAAACAGATACGACTCTTCTGTTTACTTTCTACTAACATCCTTGTGAGATGAAAACAGTTTAAAGGAGGCACTTGTAGCCTTCTCTTGAGTTTACAGATTGAACGTCTCCATAAGGAGGTCTAACTGGCGGCCCCTTTTTAAA

[0107] >EX7_R_secondaryseq

[0108] AGCAGCTAGCGGTCCCTTATCACTAGTTGCTAGCAATCACACTTTCTTGGTTTTTCAGGGAGACAACTTTGAAGTGTGGGAACGACCTCTCTCAGGCTTAGCCTGGGCTGTAGCTATGATAAACCGGCAGGAGA TTGGTGGACCTCGCTCTTATACCATCGCAGTTGCTTCCCTGGGTAAAGGAGTGGCCTGTAATCCTGCCTGCTTCATCACACAGCTCCTCCCTGTGAAAAGGAAGCTAGGGTTCTATGAATGGACTTCAAGGTTA AGAAGTCACATAAATCCCACAGGCACTGTTTTGCTTCAGCTAGAAAATACAATGCAGATGTCATTAAAAGACTTACTTTAAAATGTTTATTTTATTGCCAACTACTTCCTGTCCACCTTTTTCTCCATTCA CTTTAAAAGCTCAAGGCTAGGTGGCTCATGCCTGTAATCCCAGCACTTTGGGAGGCTGAGGCGGGCAGATCACCTGAGGTCGGGACTTTGAGACCCGCCTGGACACTGGCCGTCCGTTTTACAAGGTGGTTGGGG

[0109] (3) Merge the forward and reverse selected sequences of each exon as the query sequence, and perform comparative analysis on the query sequence, forward original sequence, and reverse original sequence of each exon to obtain the target alignment result:

[0110] 3.1 Each exon uses its corresponding forward-aligned original sequence as a reference sequence. First, the reference sequence and query sequence are aligned using the bwa tool to obtain the aligned BAM file, which is the first alignment result. Then, the samtools tool is used to create an index on the first alignment result, obtaining the first index data. The bcftools tool is then used to analyze the first index data to obtain the first raw analysis result. Finally, the first raw analysis result is filtered using preset filtering conditions to obtain the first analysis result, as shown in the table below:

[0111]

[0112] In the table, CHROM represents the analysis result number corresponding to the exon, POS represents the base position in the selected sequence (i.e., the position to be determined), REF represents the base in the reference sequence, AIL represents the base in the selected sequence, QUAL represents the probability of anomaly, INFO is used to store additional information, such as coverage depth and mutation frequency, etc.; FORMAT represents the variant site format, which corresponds one-to-one with the contents of the last column of the table. Different formats are separated by ":". In this example, GT (genotype) is 1 / 1, and PL (the probability of three different genotypes, 0 / 0, 0 / 1, and 1 / 1) are 120, 6, and 0, respectively. The smaller the value, the greater the probability of the corresponding genotype.

[0113] 3.2 Each exon uses its corresponding inverted original sequence as a reference sequence. First, the reference and query sequences are aligned using the bwa tool to obtain the aligned BAM file, which is the second alignment result. Then, the samtools tool is used to create an index on the second alignment result, obtaining the second index data. The bcftools tool is then used to analyze the second index data to obtain the second original analysis result. Finally, the second original analysis result is filtered using preset filtering conditions to obtain the second analysis result, as follows:

[0114]

[0115]

[0116] 3.3 For the first analysis result, check the second analysis result for complementary bases. In exon 7 (EXON7), base 460 changes from G to A in forward sequencing and base 38 changes from C to T in reverse sequencing. These are complementary, so the secondary peak at this position is considered the true signal, and the target alignment result with the secondary peak as the true signal is obtained.

[0117] (4) Based on the target alignment results, using the forward original sequence EX7_F.seq of exon 7 as a reference, the base at position 460 was replaced with the base A, representing the second peak, to obtain the target forward sequencing sequence EX7_F_secondarySeq of exon 7. Using the reverse original sequence EX7_R.seq of exon 7 as a reference, the base at position 38 was replaced with the base T, representing the second peak, to obtain the target reverse sequencing sequence EX7_R_secondarySeq of exon 7. The results are as follows:

[0118] EX7_F_secondarySeq

[0119] GGGCTTCCAGTCCGACCTCAGGTGATCTGCCCGCCTCAGCCTCCCAAAGTGCTGGGATTACAGGCATGAGCCACCTAGCCTTGAGCTTTTAAAGTGAATGGAGAAAAAGGTGGACAGGAAGTAGTAGTTGGCAATAAAATAAACATTTTAAAGTAAGTCTTTTAATGACATCTGCATTGTATTTTCTAGCTGAAGCAAAACAGTGCCTGTGGGATTTATGTGACTTCTTAACCTTGAAGTCCATTCATAGAACCCTAGCTTCCTTTTCACAGGGAGGAGCTGTGTGATGAAGCAGGCAGGATTACAGGCCACTCCTTTACCCAGGGAAGCAACTGCGATGGTATAAGAGCGAGGTCCACCAATCTCCTGCCGGTTTATCATAGCTACAGCCCAGGCTAAGCCTGAGAGAGGTCGTTCCCACACTTCAAAGTTGTCTCCCTGAAAAACCAAGAAAGTGTGATTGCTTAGCAACTAGTGATAAGTGGCCCTGTTAGTTTGGCATTCATTCTGGTCAAGGGGGGGTTTCCGGGAAT

[0120] >EX7_R_secondarySeq

[0121] AGCAGCTAGAGGGCCCTTATCACTAGTTGCTAAGCAATCACACTTTCTTGGTTTTTCAGGGAGACAACTTTGAAGTGTGGGAACGACCTCTCTCAGGCTTAGCCTGGGCTGTAGCTATGATAAACCGGCAGGAGA TTGGTGGACCTCGCTCTTATACCATCGCAGTTGCTTCCCTGGGTAAAGGAGTGGCCTGTAATCCTGCCTGCTTCATCACACAGCTCCTCCCTGTGAAAAGGAAGCTAGGGTTCTATGAATGGACTTCAAGGTTAA GAAGTCACATAAATCCCACAGGCACTGTTTTGCTTCAGCTAGAAAATACAATGCAGATGTCATTAAAAGACTTACTTTAAAATGTTTATTTTATTGCCAACTACTTCCTGTCCACCTTTTTCTCCATTCACT TTAAAAGCTCAAGGCTAGGTGGCTCATGCCTGTAATCCCAGCACTTTGGGAGGCTGAGGCGGGCAGATCACCTGAGGTCGGGACTTTGAGACCCGCCTGGACACTGGCCGCCCGTTTTACAAACTAGCCTGCCCAA

[0122] If the original sequence of exons 1 to 6 is the same as the selected sequence, the original sequence will be used as the target sequencing sequence.

[0123] (5) The target forward sequencing sequence and the target reverse sequencing sequence of exon 1 are merged to obtain EX1.fasta; the target forward sequencing sequence and the target reverse sequencing sequence of exon 2 are merged to obtain EX2.fasta; the target forward sequencing sequence and the target reverse sequencing sequence of exon 3 are merged to obtain EX3.fasta; the target forward sequencing sequence and the target reverse sequencing sequence of exon 4 are merged to obtain EX4.fasta; the target forward sequencing sequence and the target reverse sequencing sequence of exons 5 and 6 are merged to obtain EX56.fasta; and the forward original sequence, reverse original sequence, target forward sequencing sequence and target reverse sequencing sequence of exons 5 and 6 are merged to obtain EX7.fasta.

[0124] This invention also provides a method for gene mutation analysis, comprising:

[0125] Select the corresponding reference sequence based on the sequencing species;

[0126] The second target query sequence is established using the query sequence establishment method described above for the sequencing data of the sequenced species;

[0127] The second target query sequence and the reference sequence are compared to obtain comparison data;

[0128] All mutation sites were obtained by analyzing the alignment data.

[0129] In practical applications, taking Demo as an example, the above-described query sequence establishment method is used to establish query sequences for the seven exons of the GLA gene, obtaining the corresponding first target query sequences. Using the hg19 version of the human genome as a reference sequence, the tool bwa is used to align the reference sequence and the first target query sequence, obtaining the aligned BAM file. Then, the tool samtools is used to create an index on the BAM file, obtaining an index file. The tool bcftools is then used to analyze the index file to obtain all mutation sites.

[0130] The refGeneWithVer, dbSNP, gnomAD, ClinVar, and HGMD databases were installed using Annovar software. Bioinformatics annotation was performed on all mutation sites using Annovar software and the downloaded databases, thus obtaining detailed biological significance of the mutation sites. The results are as follows: Figure 2 As shown.

[0131] like Figure 3 As shown, this embodiment of the invention also provides a query sequence establishment system, including:

[0132] The data reading module 10 is used to read bidirectional sequencing data, obtain the forward raw sequence and the first peak data of each base at each position, and obtain the reverse raw sequence and the second peak data of each base at each position;

[0133] The base selection module 20 is used to determine the first selected base at each position based on the first peak data to obtain the forward selected sequence and the forward undetermined position, and to determine the second selected base at each position based on the second peak data to obtain the reverse selected sequence and the reverse undetermined position.

[0134] The analysis and comparison module 30 is used to merge the forward selected sequence and the reverse selected sequence to obtain the query sequence, and to perform comparison and analysis on the query sequence, the forward original sequence and the reverse original sequence to obtain the target comparison result;

[0135] The base substitution module 40 is used to control the substitution of bases at the forward undetermined positions of the forward original sequence according to the target alignment result to obtain the target forward sequencing sequence, and to control the substitution of bases at the reverse undetermined positions of the reverse original sequence according to the target alignment result to obtain the target reverse sequencing sequence.

[0136] The sequence merging module 50 is used to merge the target forward sequencing sequence and the target reverse sequencing sequence to obtain a first target query sequence. Based on the base substitution situation, it is determined whether to merge the forward original sequence and the reverse original sequence on the first target query sequence to obtain a second target query sequence.

[0137] Specific limitations regarding the query sequence establishment system can be found in the limitations on the query sequence establishment method described above, and will not be repeated here. Each module of the aforementioned query sequence establishment system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules and units can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0138] like Figure 4 As shown, an embodiment of the present invention discloses a computer device, including a memory and a processor, wherein the memory stores a computer program;

[0139] The computer device can be a server, and its internal structure diagram can be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the query sequence establishment method described in the above embodiments.

[0140] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0141] This invention also discloses a computer-readable storage medium storing a computer program that causes a computer to execute the query sequence establishment method described in the above embodiments.

[0142] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0143] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for establishing a query sequence, characterized in that, include: Read the bidirectional sequencing data to obtain the forward raw sequence and the first peak data for each base at each position, and obtain the reverse raw sequence and the second peak data for each base at each position; Based on the first peak data, the first selected base at each position is determined to obtain the forward selected sequence and the forward undetermined position. Based on the second peak data, the second selected base at each position is determined to obtain the reverse selected sequence and the reverse undetermined position. The forward selection sequence and the reverse selection sequence are merged to obtain the query sequence. The query sequence, the forward original sequence, and the reverse original sequence are compared and analyzed to obtain the target comparison result. Based on the target alignment results, the bases at the forward undetermined positions of the forward original sequence are replaced to obtain the target forward sequencing sequence. Based on the target alignment results, the bases at the reverse undetermined positions of the reverse original sequence are replaced to obtain the target reverse sequencing sequence. The target forward sequencing sequence and the target reverse sequencing sequence are merged to obtain the first target query sequence. Based on the base substitution, it is determined whether to merge the forward original sequence and the reverse original sequence on the first target query sequence to obtain the second target query sequence. The comparison analysis of the query sequence, the forward original sequence, and the reverse original sequence to obtain the target comparison result includes: The query sequence is compared with the forward original sequence to obtain the first alignment result; the query sequence is compared with the reverse original sequence to obtain the second alignment result. Analyze the first comparison result to obtain the first analysis result, analyze the second comparison result to obtain the second analysis result, and compare the first analysis result and the second analysis result to obtain the target comparison result; The comparison of the first analysis result and the second analysis result yields the target comparison result, including: The first and second analysis results at the same site are compared. If the bases in the first and second analysis results are complementary, the target alignment result with the second peak being a true signal is obtained. If the bases in the first and second analysis results are not complementary, the target alignment result with the second peak being a false signal is obtained.

2. The query sequence establishment method as described in claim 1, characterized in that, Based on the peak data, selected bases are determined at each position to obtain selected sequences and undetermined positions. When the peak data is the first peak data, the selected base is the first selected base, the selected sequence is the forward selected sequence, and the undetermined position is the forward undetermined position. When the peak data is the second peak data, the selected base is the second selected base, the selected sequence is the reverse selected sequence, and the undetermined position is the reverse undetermined position, including: Divide the peak value of the second highest peak at each location by the peak value of the highest peak to obtain the peak ratio. Determine whether the peak ratio is greater than or equal to the ratio threshold. If yes, the base corresponding to the second peak is selected as the base; otherwise, the base corresponding to the highest peak is selected as the base. The selected sequence is obtained by arranging all the selected bases in order of position, and the positions with a peak ratio greater than or equal to the ratio threshold are designated as undetermined positions.

3. The query sequence establishment method as described in claim 1, characterized in that, The analysis results are obtained by analyzing and comparing the comparison results. If the comparison result is a first comparison result, the analysis result is the first analysis result; if the comparison result is a second comparison result, the analysis result is the second analysis result, including: The comparison results are sorted, and an index is created on the sorted comparison results to obtain the comparison index data. Anomaly analysis was performed on the comparison index data to obtain the original analysis results; The positions that meet the preset filtering conditions are selected from the original analysis results to obtain the analysis results. The preset filtering conditions include: both alleles at the same position are homozygous and both have the same base type as the query sequence; the sequence coverage of the position is greater than or equal to a first preset threshold; and the probability of an anomaly at the position is greater than or equal to a second preset threshold.

4. The query sequence establishment method as described in claim 1, characterized in that, The step of controlling the substitution of bases at the undetermined positions in the forward original sequence based on the target alignment result to obtain the target forward sequencing sequence includes: If the target alignment result corresponding to the positive undetermined position is a true signal of the second peak, the bases at the positive undetermined position are replaced with the bases represented by the second peak to obtain the target positive sequencing sequence. The step of controlling the substitution of bases at the undetermined reverse positions of the reversed original sequence based on the target alignment result to obtain the target reverse sequencing sequence includes: If the target alignment result corresponding to the reverse undetermined position is a true signal at the second peak, the bases at the reverse undetermined position are replaced with the bases represented by the second peak to obtain the target reverse sequencing sequence.

5. A gene mutation analysis method, characterized in that, include: Select the corresponding reference sequence based on the sequencing species; A second target query sequence is established using the query sequence establishment method described in any one of claims 1 to 4 on the sequencing data of the sequenced species; The second target query sequence and the reference sequence are compared to obtain comparison data; All mutation sites were obtained by analyzing the alignment data.

6. A query sequence establishment system, characterized in that, include: The data reading module is used to read bidirectional sequencing data, obtain the forward raw sequence and the first peak data of each base at each position, and obtain the reverse raw sequence and the second peak data of each base at each position; The base selection module is used to determine the first selected base at each position based on the first peak data to obtain the forward selected sequence and the forward undetermined position, and to determine the second selected base at each position based on the second peak data to obtain the reverse selected sequence and the reverse undetermined position; The analysis and comparison module is used to merge the forward selected sequence and the reverse selected sequence to obtain the query sequence, and to perform comparison and analysis on the query sequence, the forward original sequence, and the reverse original sequence to obtain the target comparison result; The base substitution module is used to control the substitution of bases at the forward undetermined positions of the forward original sequence according to the target alignment result, to obtain the target forward sequencing sequence, and to control the substitution of bases at the reverse undetermined positions of the reverse original sequence according to the target alignment result, to obtain the target reverse sequencing sequence. The sequence merging module is used to merge the target forward sequencing sequence and the target reverse sequencing sequence to obtain a first target query sequence. Based on the base substitution, it is determined whether to merge the forward original sequence and the reverse original sequence on the first target query sequence to obtain a second target query sequence. The comparison analysis of the query sequence, the forward original sequence, and the reverse original sequence to obtain the target comparison result includes: The query sequence is compared with the forward original sequence to obtain the first alignment result; the query sequence is compared with the reverse original sequence to obtain the second alignment result. Analyze the first comparison result to obtain the first analysis result, analyze the second comparison result to obtain the second analysis result, and compare the first analysis result and the second analysis result to obtain the target comparison result; The comparison of the first analysis result and the second analysis result yields the target comparison result, including: The first and second analysis results at the same site are compared. If the bases in the first and second analysis results are complementary, the target alignment result with the second peak being a true signal is obtained. If the bases in the first and second analysis results are not complementary, the target alignment result with the second peak being a false signal is obtained.

7. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-4.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Probe for detecting mutation in jak2 gene and use thereof

    CN101622363A

  • SNP marker related to Sujiang pig production traits and preparation method and application thereof

    CN111793697A