Method, apparatus, device and storage medium for detecting dynamic mutation of gene

The method addresses the inefficiencies of TP-PCR by using two-generation sequencing data processing to automate dynamic mutation detection, reducing costs and time while improving efficiency.

CN115762638BActive Publication Date: 2025-07-15TIANJIN KINGMED CENT FOR CLINICAL CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211475497.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-23
Publication Date
2025-07-15
Estimated Expiration
2042-11-23

AI Technical Summary

Technical Problem

When detecting dynamic mutations of genes in the prior art, there are problems such as high experimental cost, long time and low detection efficiency.

Method used

A dynamic mutation detection method of genes is adopted, and BAM files are compared by obtaining sample comparison of target second-generation sequencing data, and reads are extracted and processed based on the preset dynamic mutation position definition file, including one-way splitting, file correction and sequence alignment, and detection results are generated.

Benefits of technology

Automatic dynamic mutation detection based on second-generation sequencing data is realized, which reduces costs, shortens detection time, and improves detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762638B_ABST
    Figure CN115762638B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method, device, equipment and storage medium for detecting dynamic mutations of genes. The method includes: extracting reads covered by each dynamic mutation from a sample alignment BAM file to obtain a covered BAM file; successively performing file processing, file correction, primary sequence alignment and secondary sequence alignment on both the first forward Fasta file and the first reverse Fasta file extracted from the covered BAM file to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file and a second reverse Fasta file; determining a target detection result according to the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file and the reverse corrected Fasta file extracted from the covered BAM file. Thus, the cost is low, the time consumption is short, and the detection efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital medical technologies, and particularly to a method, device, equipment and storage medium for detecting dynamic mutations of genes. Background Art

[0002] Dynamic mutation is an unstable state of genetic material caused by amplification during meiosis or somatic mitosis. Dynamic mutation is a new type of gene mutation that causes human genetic diseases. It is the abnormal amplification of nucleotide repeat sequences in the coding or non-coding regions of genes, and is involved in many diseases of the nervous system. Its pathogenesis is clearly related to the number of repeats of these repeat sequences. Most of such sequence variations cannot be detected by Sanger sequencing (first-generation sequencing) technology. Currently, the conventional detection method is TP-PCR (polymerase chain reaction), which has high experimental costs, long time consumption and low detection efficiency. Summary of the Invention

[0003] Based on this, in view of the technical problems of high experimental costs, long time consumption and low detection efficiency when using TP-PCR to detect dynamic mutations of genes, a method, device, equipment and storage medium for detecting dynamic mutations of genes are proposed.

[0004] The present application proposes a method for detecting dynamic mutations of genes, and the method includes:

[0005] Obtain a sample alignment BAM file corresponding to target next-generation sequencing data;

[0006] Based on a preset dynamic mutation position definition file, extract the reads covered by each dynamic mutation from the sample alignment BAM file respectively to obtain a covered BAM file;

[0007] Perform unidirectional splitting and sequence information extraction on the covered BAM file to obtain a forward BAM file, a first forward Fasta file, a reverse BAM file and a first reverse Fasta file;

[0008] Perform file processing, file correction, primary sequence alignment and secondary sequence alignment on both the first forward Fasta file and the first reverse Fasta file in sequence to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file and a second reverse Fasta file;

[0009] Generate a target detection result according to the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file and the reverse corrected Fasta file.

[0010] Further, before the step of obtaining the sample alignment BAM file corresponding to the target next-generation sequencing data, it includes:

[0011] Obtain the target genome;

[0012] Use next-generation sequencing technology and a preset sequencing range to sequence the target genome to obtain the target next-generation sequencing data;

[0013] Perform FastQC quality control, BWA alignment, and GATK analysis on the target next-generation sequencing data in sequence to obtain the sample alignment BAM file.

[0014] Further, before the step of respectively extracting the reads covered by each dynamic mutation from the sample alignment BAM file based on a preset dynamic mutation position definition file to obtain the covered BAM file, it also includes:

[0015] Respectively define Bed files according to the positions of the genes corresponding to each of the dynamic mutations in a preset reference genome to obtain the dynamic mutation position definition file;

[0016] Among them, the value range of the dynamic mutation includes: UGT1A1 mutation, NOTCH2NLC mutation, AR mutation, ATXN1 mutation, ATXN7 mutation, TBP mutation, ATN1 mutation, ATXN2 mutation, ATXN3 mutation, CACNA1A mutation, and JPH3 mutation.

[0017] Further, the step of respectively performing file processing, file correction, primary sequence alignment, and secondary sequence alignment on the first forward Fasta file and the first reverse Fasta file to obtain the forward corrected Fasta file, the second forward Fasta file, the reverse corrected Fasta file, and the second reverse Fasta file includes:

[0018] Perform N-base removal, read name uniqueness processing, and deletion of reads smaller than a preset base number on the first forward Fasta file in sequence to obtain the forward processed Fasta file, and perform N-base removal, read name uniqueness processing, and deletion of reads smaller than a preset base number on the first reverse Fasta file in sequence to obtain the reverse processed Fasta file;

[0019] Perform file correction on the forward processed Fasta file to obtain the forward corrected Fasta file, and perform file correction on the reverse processed Fasta file to obtain the reverse corrected Fasta file;

[0020] Perform primary sequence alignment and secondary sequence alignment on the forward corrected Fasta file in sequence to obtain the second forward Fasta file, and perform primary sequence alignment and secondary sequence alignment on the reverse corrected Fasta file in sequence to obtain the second reverse Fasta file.

[0021] Further, the steps of performing file correction on the forward processed Fasta file to obtain the forward corrected Fasta file and performing file correction on the reverse processed Fasta file to obtain the reverse corrected Fasta file include:

[0022] Grab the reads covered by each dynamic mutation and the unsequenced reads in the forward processed Fasta file to obtain individual single-gene forward data;

[0023] Combine and process the single-gene forward data corresponding to the ATXN1 mutation and the single-gene forward data corresponding to the ATXN3 mutation to obtain the forward data to be analyzed;

[0024] Obtain the first gene feature sequence corresponding to the ATXN1 mutation and the second gene feature sequence corresponding to the ATXN3 mutation;

[0025] According to the first gene feature sequence, search for the covered reads and unsequenced reads in the forward data to be analyzed to obtain the first forward grouped data, and according to the second gene feature sequence, search for the covered reads and unsequenced reads in the forward data to be analyzed to obtain the second forward grouped data;

[0026] Use the first forward grouped data, the second forward grouped data, and each of the single-gene forward data other than the forward data to be analyzed as the forward corrected Fasta file;

[0027] Grab the reads covered by each dynamic mutation and the unsequenced reads in the reverse processed Fasta file to obtain individual single-gene reverse data;

[0028] Combine and process the single-gene reverse data corresponding to the ATXN1 mutation and the single-gene reverse data corresponding to the ATXN3 mutation to obtain the forward data to be analyzed;

[0029] According to the first gene feature sequence, search for the covered reads and unsequenced reads in the reverse data to be analyzed to obtain the first reverse grouped data, and according to the second gene feature sequence, search for the covered reads and unsequenced reads in the reverse data to be analyzed to obtain the second reverse grouped data;

[0030] Use each of the single-gene reverse data other than the first reverse grouped data, the second reverse grouped data, and the forward data to be analyzed as the reverse-corrected Fasta file.

[0031] Further, the step of performing a first-level sequence alignment and a second-level sequence alignment on the forward-corrected Fasta file in sequence to obtain the second forward Fasta file, and performing a first-level sequence alignment and a second-level sequence alignment on the reverse-corrected Fasta file in sequence to obtain the second reverse Fasta file includes:

[0032] Obtain each preset first motif, where each of the first motifs is data extracted according to the characteristics of dynamic mutations;

[0033] Align each of the first motifs with the forward-corrected Fasta file respectively, and remove the sequences at both ends of the repeated sequences in the reads of the forward-corrected Fasta file based on the alignment position information to obtain a first-level forward alignment result;

[0034] Align each of the first motifs with the reverse-corrected Fasta file in sequence, and remove the sequences at both ends of the repeated sequences in the reads of the forward-corrected Fasta file based on the alignment position information to obtain a first-level reverse alignment result;

[0035] Obtain each preset second motif, where each of the second motifs is data extracted according to the characteristics of dynamic mutations, and the length of the second motif is greater than the length of the first motif;

[0036] Align each of the second motifs with the first-level forward alignment result in sequence, and remove the sequences at both ends of the repeated sequences in the reads of the first-level forward alignment result based on the alignment position information to obtain the second forward Fasta file;

[0037] Align each of the second motifs with the first-level reverse alignment result in sequence, and remove the sequences at both ends of the repeated sequences in the reads of the first-level reverse alignment result based on the alignment position information to obtain the second reverse Fasta file.

[0038] Further, the step of generating a detection result according to the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward-corrected Fasta file, and the reverse-corrected Fasta file to obtain a target detection result includes:

[0039] Execute a preset data summarization script to summarize the data of the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file to obtain a summary result;

[0040] Adopt a preset detection result generation rule to generate a detection result from the summary result to obtain the target detection result;

[0041] Among them, the summary result includes: the name of the dynamic mutation, the name of the sequencing reads, the direction of the sequencing reads, the trimmed reads sequence, the 10 bp bases on the left side of the original reads, the 10 bp bases on the right side of the original reads, the length of the trimmed reads sequence, the detected number of repeats without spacer sequences, the Levenshtein distance score, whether the left side of the trimmed reads sequence is sequenced through, and whether the right side of the trimmed reads sequence is sequenced through. The name of the sequencing reads and the trimmed reads sequence are both data extracted from the second forward Fasta file or the second reverse Fasta file. The direction of the sequencing reads is data extracted from the forward BAM file or the reverse BAM file. The 10 bp bases on the left side of the original reads and the 10 bp bases on the right side of the original reads are both data extracted from the forward corrected Fasta file or the reverse corrected Fasta file. The length of the trimmed reads sequence and the detected number of repeats without spacer sequences are both data calculated from the trimmed reads sequence. The Levenshtein distance score is a score calculated based on the trimmed reads sequence and a preset reference sequence. Whether the left side of the trimmed reads sequence is sequenced through is data determined according to the 10 bp bases on the left side of the original reads. Whether the right side of the trimmed reads sequence is sequenced through is data determined according to the 10 bp bases on the right side of the original reads;

[0042] The target detection result includes: the name of the dynamic mutation, the reference value of the mutation repeat count, the number of original reads, the number of reads supporting the mutation, the number of reads not sequenced through with more than 40 repeats, and the detection conclusion. The number of original reads, the number of reads supporting the mutation, and the number of reads not sequenced through with more than 40 repeats are all data calculated based on the summary result. The detection conclusion is data calculated based on the reference value of the mutation repeat count, the number of original reads, the number of reads supporting the mutation, and the number of reads not sequenced through with more than 40 repeats.

[0043] This application also proposes a device for detecting dynamic mutations of genes, and the device includes:

[0044] A data acquisition module for acquiring a sample alignment BAM file corresponding to target next-generation sequencing data;

[0045] A covered BAM file determination module for extracting reads covered by each dynamic mutation from the sample alignment BAM file respectively based on a preset dynamic mutation position definition file to obtain a covered BAM file;

[0046] A splitting module for unidirectionally splitting the covered BAM file and extracting sequence information to obtain a forward BAM file, a first forward Fasta file, a reverse BAM file, and a first reverse Fasta file;

[0047] A processing module for sequentially performing file processing, file correction, primary sequence alignment, and secondary sequence alignment on both the first forward Fasta file and the first reverse Fasta file to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file, and a second reverse Fasta file;

[0048] A detection result generation module for generating a detection result according to the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file to obtain a target detection result.

[0049] This application also provides a computer device, including a memory and a processor. When a computer program stored in the memory is executed by the processor, the processor performs the following steps:

[0050] Acquire a sample alignment BAM file corresponding to target next-generation sequencing data;

[0051] Based on a preset dynamic mutation position definition file, extract reads covered by each dynamic mutation from the sample alignment BAM file respectively to obtain a covered BAM file;

[0052] Unidirectionally split the covered BAM file and extract sequence information to obtain a forward BAM file, a first forward Fasta file, a reverse BAM file, and a first reverse Fasta file;

[0053] Sequentially perform file processing, file correction, primary sequence alignment, and secondary sequence alignment on both the first forward Fasta file and the first reverse Fasta file to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file, and a second reverse Fasta file;

[0054] Generate detection results based on the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file to obtain the target detection results.

[0055] This application also proposes a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to perform the following steps:

[0056] Obtain a sample alignment BAM file corresponding to the target next-generation sequencing data;

[0057] Based on a preset dynamic mutation position definition file, extract the reads covered by each dynamic mutation from the sample alignment BAM file respectively to obtain a covered BAM file;

[0058] Perform unidirectional splitting and sequence information extraction on the covered BAM file to obtain a forward BAM file, a first forward Fasta file, a reverse BAM file, and a first reverse Fasta file;

[0059] Perform file processing, file correction, primary sequence alignment, and secondary sequence alignment on both the first forward Fasta file and the first reverse Fasta file in sequence to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file, and a second reverse Fasta file;

[0060] Generate detection results based on the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file to obtain the target detection results.

[0061] The method for detecting dynamic mutations of genes in this application realizes automatic detection of dynamic mutations based on next-generation sequencing data. Compared with TP-PCR, next-generation sequencing has low cost and short time consumption. Also, since each dynamic mutation is detected at one time, the detection efficiency is improved, thus avoiding the technical problems of high experimental cost, long time consumption, and low detection efficiency when using TP-PCR to detect dynamic mutations of genes. Description of the Drawings

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0063] Wherein:

[0064] Figure 1 is a flowchart of a method for detecting dynamic mutations of genes in an embodiment;

[0065] Figure 2 is a structural block diagram of a device for detecting dynamic mutations of genes in an embodiment;

[0066] Figure 3 is a structural block diagram of a computer device in an embodiment. Specific Embodiments

[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0068] As Figure 1 shown, in an embodiment, a method for detecting dynamic mutations of genes is provided. This method can be applied to a terminal or a server. In this embodiment, it is exemplified by being applied to a terminal. The method for detecting dynamic mutations of genes specifically includes the following steps:

[0069] S1: Obtain a sample alignment BAM file corresponding to the target next-generation sequencing data;

[0070] Specifically, a sample alignment BAM file corresponding to the target next-generation sequencing data input by the user can be obtained, or a sample alignment BAM file corresponding to the target next-generation sequencing data can be obtained from a storage space, or a sample alignment BAM file corresponding to the target next-generation sequencing data can be obtained from a third-party application.

[0071] Perform whole-genome or whole-exome next-generation sequencing on the genome of the target living body, and use the data obtained from the next-generation sequencing as the target next-generation sequencing data. The target living body can be a human or an animal.

[0072] Perform quality control, alignment with a reference genome, and correction of the alignment results on the target next-generation sequencing data in sequence, and use the corrected data as the sample alignment BAM file.

[0073] Next-generation sequencing is also known as second-generation DNA sequencing technology or next-generation sequencing technology.

[0074] The sample alignment BAM file is a BAM file. BAM is the binary file of SAM, which has a smaller storage space, and many downstream analysis tools use the BAM format.

[0075] S2: Based on a preset dynamic mutation position definition file, extract the reads covered by each dynamic mutation from the sample alignment BAM file respectively to obtain a covered BAM file;

[0076] Specifically, the dynamic mutation position definition file is a file defined based on the positions of dynamic mutations in the reference genome; based on the preset dynamic mutation position definition file, extract the reads covered by each dynamic mutation from the sample alignment BAM file respectively, and use all the data corresponding to the extracted reads in the sample alignment BAM file as the covered BAM file.

[0077] The value range of the dynamic mutations includes but is not limited to: UGT1A1 mutation, NOTCH2NLC mutation, AR mutation, ATXN1 mutation, ATXN7 mutation, TBP mutation, ATN1 mutation, ATXN2 mutation, ATXN3 mutation, CACNA1A mutation, and JPH3 mutation. The UGT1A1 mutation is a mutation of UGT1A1; the NOTCH2NLC mutation is a mutation of NOTCH2NLC; the AR mutation is a mutation of AR; the ATXN1 mutation is a mutation of ATXN1; the ATXN7 mutation is a mutation of ATXN7; the TBP mutation is a mutation of TBP; the ATN1 mutation is a mutation of ATN1; the ATXN2 mutation is a mutation of ATXN2; the ATXN3 mutation is a mutation of ATXN3; the CACNA1A mutation is a mutation of CACNA1A; the JPH3 mutation is a mutation of JPH3. UGT1A1, NOTCH2NLC, AR, ATXN1, ATXN7, TBP, ATN1, ATXN2, ATXN3, CACNA1A, and JPH3 are genes with common dynamic mutations in clinical practice.

[0078] Optionally, use the samtools software to extract the reads covered by each dynamic mutation from the sample alignment BAM file respectively based on the preset dynamic mutation position definition file to obtain a covered BAM file.

[0079] Samtools is a collection of tools for operating on sam and bam files (usually generated by short sequence alignment tools such as bwa, bowtie2, hisat2, tophat2, etc., and the specific format can be viewed by entering "SAM" in the message box), and contains many commands.

[0080] S3: Perform unidirectional splitting and sequence information extraction on the covered BAM file to obtain a forward BAM file, a first forward Fasta file, a reverse BAM file, and a first reverse Fasta file;

[0081] Specifically, the covered BAM file is split forward to obtain a forward BAM file, and sequence information is extracted from the forward BAM file to form a Fasta file, obtaining a first forward Fasta file; the covered BAM file is split backward to obtain a reverse BAM file, and sequence information is extracted from the reverse BAM file to form a Fasta file, obtaining a first reverse Fasta file. Thus, splitting is achieved according to the read direction of next-generation sequencing.

[0082] reads, that is, sequence fragments.

[0083] Fasta file, a file in Fasta format. The Fasta format is a text-based format for representing nucleic acid sequences or polypeptide sequences. In it, nucleic acids or amino acids are each represented by a single letter, and it allows adding a sequence name and annotation before the sequence. This format has become a standard in the field of bioinformatics.

[0084] S4: The first forward Fasta file and the first reverse Fasta file are each successively subjected to file processing, file correction, primary sequence alignment, and secondary sequence alignment to obtain a forward-corrected Fasta file, a second forward Fasta file, a reverse-corrected Fasta file, and a second reverse Fasta file;

[0085] Specifically, the first forward Fasta file is successively subjected to file processing, file correction, primary sequence alignment, and secondary sequence alignment to obtain a forward-corrected Fasta file and a second forward Fasta file; the first reverse Fasta file is successively subjected to file processing, file correction, primary sequence alignment, and secondary sequence alignment to obtain a reverse-corrected Fasta file and a second reverse Fasta file. Among them, file processing is used to remove noises that affect the detection results, such as N bases. File correction is used to eliminate errors caused by different repetition times of ATXN1 mutations and ATXN3 mutations during the alignment process, such as the errors that occur during BWA alignment. Primary sequence alignment is used to remove the sequences at both ends of the repeated sequences in the reads. Secondary sequence alignment is used to remove the sequences at both ends of the repeated sequences in the reads. The length of the alignment reference for secondary sequence alignment is longer than that of the alignment reference for primary sequence alignment.

[0086] S5: Based on the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward-corrected Fasta file, and the reverse-corrected Fasta file, the target detection result is generated.

[0087] Specifically, based on a preset data summarization rule, summarize the data of the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file, and then generate a detection result based on a preset detection result generation rule for the summarized result, and use the generated detection result as the target detection result.

[0088] This embodiment realizes automatic detection of dynamic mutations based on second-generation sequencing data. Compared with TP-PCR, the second-generation sequencing has low cost and short time consumption. Moreover, since each dynamic mutation is detected at one time, the detection efficiency is improved, thus avoiding the technical problems of high experimental cost, long time consumption, and low detection efficiency when using TP-PCR to detect dynamic mutations of genes.

[0089] In one embodiment, before the step of obtaining the sample alignment BAM file corresponding to the target second-generation sequencing data, it includes:

[0090] S11: Obtain a target genome;

[0091] Specifically, a target genome input by a user can be obtained, or a target genome can be obtained from a storage space, or a target genome can be obtained from a third-party application.

[0092] The target genome is the genome of a target living body.

[0093] S12: Sequence the target genome using second-generation sequencing technology and a preset sequencing range to obtain the target second-generation sequencing data;

[0094] Specifically, sequence the target genome using second-generation sequencing technology and a preset sequencing range, and use the data obtained by second-generation sequencing as the target second-generation sequencing data. The sequencing range is the whole genome or the whole exome.

[0095] S13: Perform FastQC quality control, BWA alignment, and GATK analysis on the target second-generation sequencing data in sequence to obtain the sample alignment BAM file.

[0096] Specifically, perform FastQC quality control on the target second-generation sequencing data in sequence, perform BWA alignment on the quality control result and the reference genome corresponding to the target genome, perform GATK analysis on the result of the BWA alignment, and use the analysis result as the sample alignment BAM file.

[0097] FastQC quality control is used to delete low-quality data. FastQC quality control is performed using the FastQC software. FastQC is a Java-based software that can quickly evaluate the quality of sequencing data.

[0098] The result of BWA alignment is a BAM file, which contains: information about bases, alignment rate (how much can be aligned), alignment length, alignment positions (including start or end positions), and statistical scores. BWA mainly aligns reads to large genomes, and its main function is sequence alignment.

[0099] GATK analysis is used to perform variant analysis and correction on the results of BWA alignment using GATK, such as correction of point mutations. GATK, whose full English name is Genome Analysis Toolkit, is a set of software for processing high-throughput sequencing data. GATK was initially designed for analyzing human whole-exome and whole-genome data, and now it can also be applied to other species with continuous development.

[0100] The reference genome corresponding to the target genome is the reference genome corresponding to the living body type corresponding to the target living body. The living body type can be human or animal.

[0101] In this embodiment, the target genome is sequenced using second-generation sequencing technology and a preset sequencing range to obtain the target second-generation sequencing data, and the target second-generation sequencing data is successively subjected to FastQC quality control, BWA alignment, and GATK analysis, thereby improving the quality of the data used for dynamic mutation detection and the accuracy of dynamic mutation detection.

[0102] In one embodiment, before the step of obtaining the covered BAM file by respectively extracting the reads covered by each dynamic mutation from the sample alignment BAM file based on the preset dynamic mutation position definition file, it further includes:

[0103] S21: respectively define Bed files according to the positions of the genes corresponding to each of the dynamic mutations in the preset reference genome to obtain the dynamic mutation position definition file;

[0104] Among them, the value range of the dynamic mutation includes: UGT1A1 mutation, NOTCH2NLC mutation, AR mutation, ATXN1 mutation, ATXN7 mutation, TBP mutation, ATN1 mutation, ATXN2 mutation, ATXN3 mutation, CACNA1A mutation, and JPH3 mutation.

[0105] Specifically, Bed files are defined respectively according to the positions of each of UGT1A1, NOTCH2NLC, AR, ATXN1, ATXN7, TBP, ATN1, ATXN2, ATXN3, CACNA1A, and JPH3 in a preset reference genome, and the defined Bed files are used as the dynamic mutation position definition files.

[0106] A Bed file is a file in Bed format. The full name of a Bed format file is Browser Extensible Data, which displays annotation information by specifying the content of each line.

[0107] In this embodiment, Bed files are defined respectively according to the positions of the genes corresponding to each of the dynamic mutations in a preset reference genome to obtain the dynamic mutation position definition files, thereby providing a basis for determining the coverage BAM files based on the dynamic mutation position definition files.

[0108] In one embodiment, the steps of sequentially performing file processing, file correction, primary sequence alignment, and secondary sequence alignment on the first forward Fasta file and the first reverse Fasta file to obtain a forward-corrected Fasta file, a second forward Fasta file, a reverse-corrected Fasta file, and a second reverse Fasta file include:

[0109] S41: Sequentially perform N-base removal, read name uniqueness processing, and deletion of reads with a base number less than a preset base number on the first forward Fasta file to obtain a forward-processed Fasta file, and sequentially perform N-base removal, read name uniqueness processing, and deletion of reads with a base number less than a preset base number on the first reverse Fasta file to obtain a reverse-processed Fasta file;

[0110] An N-base is a base that cannot be recognized as ACGT in second-generation sequencing. The ACGT of an organism is a nitrogenous base, and there are 5 types of bases in total: cytosine (abbreviated as C), guanine (G), adenine (A), thymine (T, specific to DNA), and uracil (U, specific to RNA).

[0111] Read name uniqueness processing means renaming the same read names in the first forward Fasta file after removing N-bases to achieve read name uniqueness, thereby avoiding errors in detection results.

[0112] After read name uniqueness processing, delete reads with a base number less than the preset base number.

[0113] Optionally, the value range of the preset number of bases is from 40 bp to 80 bp.

[0114] Optionally, the preset number of bases is 50 bp. That is to say, the preset number of bases is 50 bases.

[0115] S42: Perform file correction on the forward-processed Fasta file to obtain the forward-corrected Fasta file, and perform file correction on the reverse-processed Fasta file to obtain the reverse-corrected Fasta file;

[0116] File correction is used to eliminate the problem of incorrect coverage during the alignment process caused by different repetition times of the two genes with dynamic mutations, namely ATXN1 and ATXN3.

[0117] S43: Perform primary sequence alignment and secondary sequence alignment on the forward-corrected Fasta file in sequence to obtain the second forward Fasta file, and perform primary sequence alignment and secondary sequence alignment on the reverse-corrected Fasta file in sequence to obtain the second reverse Fasta file.

[0118] Specifically, perform primary sequence alignment on the forward-corrected Fasta file, and perform secondary sequence alignment based on the result of the primary sequence alignment to obtain the second forward Fasta file; perform primary sequence alignment on the reverse-corrected Fasta file, and perform secondary sequence alignment based on the result of the primary sequence alignment to obtain the second reverse Fasta file. Thus, the clipping of the Fasta file is achieved.

[0119] This embodiment improves the accuracy of the detection result through file processing, file correction, primary sequence alignment, and secondary sequence alignment.

[0120] In one embodiment, the steps of performing file correction on the forward-processed Fasta file to obtain the forward-corrected Fasta file, and performing file correction on the reverse-processed Fasta file to obtain the reverse-corrected Fasta file include:

[0121] S421: Grab the reads covered by each dynamic mutation and the unsequenced reads from the forward-processed Fasta file to obtain individual single-gene forward data;

[0122] Specifically, grab the reads covered by each dynamic mutation and the unsequenced reads from the forward-processed Fasta file, and take the covered reads and unsequenced reads grabbed for the same dynamic mutation as the single-gene forward data corresponding to this dynamic mutation. That is to say, each dynamic mutation corresponds to one single-gene forward data.

[0123] Unsequenced reads, that is, highly repetitive bases are detected in the sequence but the characteristic sequence cannot be recognized.

[0124] S422: Combine the single-gene forward data corresponding to the ATXN1 mutation and the single-gene forward data corresponding to the ATXN3 mutation to obtain the forward data to be analyzed;

[0125] Specifically, splice the single-gene forward data corresponding to the ATXN1 mutation and the single-gene forward data corresponding to the ATXN3 mutation, and use the spliced data as the forward data to be analyzed.

[0126] S423: Obtain the first gene characteristic sequence corresponding to the ATXN1 mutation and the second gene characteristic sequence corresponding to the ATXN3 mutation;

[0127] The first gene characteristic sequence is a characteristic sequence extracted according to the characteristics of the gene corresponding to the ATXN1 mutation.

[0128] The second gene characteristic sequence is a characteristic sequence extracted according to the characteristics of the gene corresponding to the ATXN3 mutation.

[0129] S424: According to the first gene characteristic sequence, search for the covered reads and unsequenced reads from the forward data to be analyzed to obtain the first forward grouped data. According to the second gene characteristic sequence, search for the covered reads and unsequenced reads from the forward data to be analyzed to obtain the second forward grouped data;

[0130] Specifically, according to the first gene characteristic sequence, search for the covered reads and unsequenced reads from the forward data to be analyzed, and use all the searched reads as the first forward grouped data. According to the second gene characteristic sequence, search for the covered reads and unsequenced reads from the forward data to be analyzed, and use all the searched reads as the second forward grouped data, thus realizing re-calibration and grouping during the calibration process.

[0131] S425: Use each single-gene forward data other than the first forward grouped data, the second forward grouped data, and the forward data to be analyzed as the forward-calibrated Fasta file;

[0132] S426: Grab the reads covered by each dynamic mutation and unsequenced reads from the reverse-processed Fasta file to obtain each single-gene reverse data;

[0133] Specifically, for the reverse-processed Fasta file, reads covered by each dynamic mutation and unsequenced reads are captured, and the covered reads and unsequenced reads captured for the same dynamic mutation are used as the single-gene reverse data corresponding to this dynamic mutation. That is, each dynamic mutation corresponds to a single-gene reverse data.

[0134] S427: Merge the single-gene reverse data corresponding to the ATXN1 mutation and the single-gene reverse data corresponding to the ATXN3 mutation to obtain the forward data to be analyzed;

[0135] Specifically, splice the single-gene reverse data corresponding to the ATXN1 mutation and the single-gene reverse data corresponding to the ATXN3 mutation, and use the spliced data as the forward data to be analyzed.

[0136] S428: According to the first gene feature sequence, search for covered reads and unsequenced reads from the reverse data to be analyzed to obtain the first reverse grouped data, and according to the second gene feature sequence, search for covered reads and unsequenced reads from the reverse data to be analyzed to obtain the second reverse grouped data;

[0137] Specifically, according to the first gene feature sequence, search for covered reads and unsequenced reads from the reverse data to be analyzed, and use all the searched reads as the first reverse grouped data. According to the second gene feature sequence, search for covered reads and unsequenced reads from the reverse data to be analyzed, and use all the searched reads as the second reverse grouped data.

[0138] S429: Use each single-gene reverse data other than the first reverse grouped data, the second reverse grouped data, and the forward data to be analyzed as the reverse-corrected Fasta file.

[0139] Since the genes corresponding to the ATXN1 mutation and the genes corresponding to the ATXN3 mutation are both repetitive, and the number of repetitions is different, highly repetitive reads will have alignment errors during the bioinformatics alignment analysis, thus causing an incorrect impact during the alignment process. For example, errors occur during BWA alignment. To solve this problem, in this embodiment, each dynamic mutation is grouped first, and then the groups corresponding to the ATXN1 mutation and the groups corresponding to the ATXN3 mutation are corrected, providing a basis for ensuring the accuracy of subsequent detection results.

[0140] In one embodiment, the steps of sequentially performing primary sequence alignment and secondary sequence alignment on the forward-corrected Fasta file to obtain the second forward Fasta file, and sequentially performing primary sequence alignment and secondary sequence alignment on the reverse-corrected Fasta file to obtain the second reverse Fasta file include:

[0141] S431: Obtain each preset first motif, where each of the first motifs is data extracted according to the characteristics of dynamic mutation;

[0142] Optionally, the first motif corresponding to AR, ATXN1, ATXN7, TBP, ATN1, ATXN2, ATXN3, CACNA1A is "CAGCAGCAG"; the first motif corresponding to UGT1A1 is "TATA"; the first motif corresponding to JPH3 is "CTGCTGCTG"; the first motif corresponding to NOTCH2NLC is "CGGCGG".

[0143] S432: Align each of the first motifs with the forward-corrected Fasta file respectively, and remove the sequences at both ends of the repeated sequences in the reads of the forward-corrected Fasta file based on the alignment position information to obtain a primary forward alignment result;

[0144] Specifically, align each of the first motifs with the forward-corrected Fasta file respectively, then according to the position information in the alignment result, remove the sequences at both ends of the repeated sequences in the reads of the forward-corrected Fasta file, and use the remaining data in the forward-corrected Fasta file as the primary forward alignment result, where the sequence alignment is fuzzy positioning alignment information.

[0145] S433: Align each of the first motifs with the reverse-corrected Fasta file respectively, and remove the sequences at both ends of the repeated sequences in the reads of the forward-corrected Fasta file based on the alignment position information to obtain a primary reverse alignment result;

[0146] Specifically, align each of the first motifs with the reverse-corrected Fasta file respectively, then according to the position information in the alignment result, remove the sequences at both ends of the repeated sequences in the reads of the reverse-corrected Fasta file, and use the remaining data in the reverse-corrected Fasta file as the primary reverse alignment result.

[0147] S434: Obtain each preset second motif, where each of the second motifs is data extracted according to the characteristics of dynamic mutations, and the length of the second motif is greater than the length of the first motif;

[0148] Optionally, the second motifs corresponding to AR, ATXN1, ATXN7, TBP, ATN1, ATXN2, ATXN3, CACNA1A are "CAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAGCAG; the second motif corresponding to UGT1A1 is "TATATATATATATA"; the second motif corresponding to JPH3 is "CTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTGCTG; the second motif corresponding to NOTCH2NLC is "CGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGG.

[0149] S435: Align each of the second motifs with the first-level forward alignment result in sequence, and based on the alignment position information, remove the sequences at both ends of the repeated sequences in the reads of the first-level forward alignment result to obtain the second forward Fasta file;

[0150] Specifically, align each of the second motifs with the first-level forward alignment result in sequence, then remove the sequences at both ends of the repeated sequences in the reads of the first-level forward alignment result according to the position information in the alignment result, and use the remaining data in the first-level forward alignment result as the second forward Fasta file.

[0151] S436: Align each of the second motifs with the first-level reverse alignment result in sequence, and based on the alignment position information, remove the sequences at both ends of the repeated sequences in the reads of the first-level reverse alignment result to obtain the second reverse Fasta file.

[0152] Specifically, align each of the second motifs with the first-level reverse alignment result in sequence, then remove the sequences at both ends of the repeated sequences in the reads of the first-level reverse alignment result according to the position information in the alignment result, and use the remaining data in the first-level reverse alignment result as the second reverse Fasta file.

[0153] The length of the second motif in this embodiment is greater than the length of the first motif, realizing first performing a primary sequence alignment with a relatively short length and then performing a primary sequence alignment with a relatively long length, improving the accuracy of the generated detection results.

[0154] In one embodiment, the step of generating a target detection result according to the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file includes:

[0155] S51: Execute a preset data summarization script to summarize the data of the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file to obtain a summarization result;

[0156] S52: Use a preset detection result generation rule to generate a detection result from the summarization result to obtain the target detection result;

[0157] Among them, the summarization result includes: the name of the dynamic mutation, the name of the sequencing reads, the direction of the sequencing reads, the trimmed reads sequence, the 10 bp bases on the left side of the original reads, the 10 bp bases on the right side of the original reads, the length of the trimmed reads sequence, the detected repeat count without the spacer sequence, the Levenshtein distance score, whether the left side of the trimmed reads sequence is sequenced through, whether the right side of the trimmed reads sequence is sequenced through. The name of the sequencing reads and the trimmed reads sequence are both data extracted from the second forward Fasta file or the second reverse Fasta file. The direction of the sequencing reads is data extracted from the forward BAM file or the reverse BAM file. The 10 bp bases on the left side of the original reads and the 10 bp bases on the right side of the original reads are both data extracted from the forward corrected Fasta file or the reverse corrected Fasta file. The length of the trimmed reads sequence and the detected repeat count without the spacer sequence are both data calculated from the trimmed reads sequence. The Levenshtein distance score is a score calculated according to the trimmed reads sequence and a preset reference sequence. Whether the left side of the trimmed reads sequence is sequenced through is data determined according to the 10 bp bases on the left side of the original reads. Whether the right side of the trimmed reads sequence is sequenced through is data determined according to the 10 bp bases on the right side of the original reads;

[0158] The target detection results include: the name of the dynamic mutation, the reference value of the mutation repeat count, the number of original reads, the number of reads supporting the mutation, the number of unsequenced reads with more than 40 repeats, and the detection conclusion. The number of original reads, the number of reads supporting the mutation, and the number of unsequenced reads with more than 40 repeats are all data calculated based on the summarized results. The detection conclusion is data calculated based on the reference value of the mutation repeat count, the number of original reads, the number of reads supporting the mutation, and the number of unsequenced reads with more than 40 repeats.

[0159] The programming language for developing the data summarization script can be selected from existing computations, such as Java, which is not limited here.

[0160] Specifically, the summarized result is in tabular form. Therefore, execute a preset data summarization script to summarize the data of the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file to generate a table, and use this table as the summarized result. Among them, the field names of the header of the summarized result include: the name of the dynamic mutation, the name of the sequenced reads, the direction of the sequenced reads, the trimmed reads sequence, the 10 bp bases on the left side of the original reads, the 10 bp bases on the right side of the original reads, the length of the trimmed reads sequence, the detected repeat count without the spacer sequence, the Levenshtein distance score, whether the left side of the trimmed reads sequence is sequenced through, and whether the right side of the trimmed reads sequence is sequenced through.

[0161] The target detection result is in tabular form. Therefore, adopt a preset detection result generation rule to generate the detection result from the summarized result to form a table, and use this table as the target detection result. Among them, the field names of the header of the summarized result include: the name of the dynamic mutation, the reference value of the mutation repeat count, the number of original reads, the number of reads supporting the mutation, the number of unsequenced reads with more than 40 repeats, and the detection conclusion.

[0162] The value of the detection conclusion is normal or "mutation + the number of mutated reads". For example, the detection conclusion corresponding to UGT1A1 is mutation(3), that is, there are 3 mutated reads.

[0163] The reference value of the mutation repeat count is the reference value of the mutation repeat count preset for each dynamic mutation.

[0164] The target detection results include the detection situation and mutation information of each dynamic mutation in the target second-generation sequencing data.

[0165] This embodiment realizes automatic data aggregation and automatic generation of detection results, improving the automation level of this application and the detection efficiency.

[0166] As Figure 2 shown, this application also proposes a device for detecting dynamic mutations of genes, and the device includes:

[0167] A data acquisition module 801, configured to acquire a sample alignment BAM file corresponding to target next-generation sequencing data;

[0168] A covered BAM file determination module 802, configured to extract reads covered by each dynamic mutation from the sample alignment BAM file respectively based on a preset dynamic mutation position definition file, to obtain a covered BAM file;

[0169] A splitting module 803, configured to perform unidirectional splitting and sequence information extraction on the covered BAM file, to obtain a forward BAM file, a first forward Fasta file, a reverse BAM file, and a first reverse Fasta file;

[0170] A processing module 804, configured to sequentially perform file processing, file correction, primary sequence alignment, and secondary sequence alignment on both the first forward Fasta file and the first reverse Fasta file, to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file, and a second reverse Fasta file;

[0171] A detection result generation module 805, configured to generate a detection result according to the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file, to obtain a target detection result.

[0172] This embodiment realizes automatic detection of dynamic mutations based on next-generation sequencing data. Compared with TP-PCR, next-generation sequencing has low cost and short time consumption. Also, because various dynamic mutations are detected at one time, the detection efficiency is improved, thus avoiding the technical problems of high experimental cost, long time consumption, and low detection efficiency when using TP-PCR to detect dynamic mutations of genes.

[0173] Figure 3 shows the internal structure diagram of a computer device in an embodiment. The computer device can specifically be a terminal or a server. As Figure 3As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement a method for detecting dynamic mutations of genes. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the method for detecting dynamic mutations of genes. Those skilled in the art can understand that Figure 3 The structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0174] In one embodiment, a computer device is proposed, including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the following steps:

[0175] Obtain the sample alignment BAM file corresponding to the target next-generation sequencing data;

[0176] Based on a preset dynamic mutation position definition file, extract the reads covered by each dynamic mutation from the sample alignment BAM file respectively to obtain a covered BAM file;

[0177] Perform unidirectional splitting and sequence information extraction on the covered BAM file to obtain a forward BAM file, a first forward Fasta file, a reverse BAM file, and a first reverse Fasta file;

[0178] Perform file processing, file correction, primary sequence alignment, and secondary sequence alignment on both the first forward Fasta file and the first reverse Fasta file in sequence to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file, and a second reverse Fasta file;

[0179] Generate a detection result according to the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file to obtain a target detection result.

[0180] This embodiment realizes automatic detection of dynamic mutations based on next-generation sequencing data. Compared with TP-PCR, next-generation sequencing has low cost and short time consumption. Moreover, since each dynamic mutation is detected at one time, the detection efficiency is improved, thus avoiding the technical problems of high experimental cost, long time consumption, and low detection efficiency when using TP-PCR to detect dynamic mutations of genes.

[0181] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor, causes the processor to perform the following steps:

[0182] Obtain a sample alignment BAM file corresponding to the target next-generation sequencing data;

[0183] Based on a preset dynamic mutation position definition file, extract the reads covered by each dynamic mutation from the sample alignment BAM file respectively to obtain a covered BAM file;

[0184] Perform unidirectional splitting and sequence information extraction on the covered BAM file to obtain a forward BAM file, a first forward Fasta file, a reverse BAM file, and a first reverse Fasta file;

[0185] Perform file processing, file correction, primary sequence alignment, and secondary sequence alignment on both the first forward Fasta file and the first reverse Fasta file in sequence to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file, and a second reverse Fasta file;

[0186] Generate a detection result according to the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file to obtain a target detection result.

[0187] This embodiment realizes automatic detection of dynamic mutations based on next-generation sequencing data. Compared with TP-PCR, next-generation sequencing has low cost and short time consumption. Moreover, since each dynamic mutation is detected at one time, the detection efficiency is improved, thus avoiding the technical problems of high experimental cost, long time consumption, and low detection efficiency when using TP-PCR to detect dynamic mutations of genes.

[0188] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0189] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0190] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for detecting dynamic mutations of a gene, the method comprising: Obtaining a sample alignment BAM file corresponding to target next-generation sequencing data; Based on a preset dynamic mutation position definition file, extracting reads covered by each dynamic mutation from the sample alignment BAM file respectively to obtain a covered BAM file; Performing unidirectional splitting and sequence information extraction on the covered BAM file to obtain a forward BAM file, a first forward Fasta file, a reverse BAM file, and a first reverse Fasta file; Successively performing file processing, file correction, primary sequence alignment, and secondary sequence alignment on both the first forward Fasta file and the first reverse Fasta file to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file, and a second reverse Fasta file; Generating a detection result according to the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file to obtain a target detection result; Wherein, before the step of obtaining the sample alignment BAM file corresponding to the target next-generation sequencing data, it includes: Obtaining a target genome; Sequencing the target genome using next-generation sequencing technology and a preset sequencing range to obtain the target next-generation sequencing data; Successively performing FastQC quality control, BWA alignment, and GATK analysis on the target next-generation sequencing data to obtain the sample alignment BAM file; Wherein, the step of successively performing file processing, file correction, primary sequence alignment, and secondary sequence alignment on both the first forward Fasta file and the first reverse Fasta file to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file, and a second reverse Fasta file includes: Successively performing N-base removal, read name uniqueness processing, and deletion of reads smaller than a preset base number on the first forward Fasta file to obtain a forward processed Fasta file, and successively performing N-base removal, read name uniqueness processing, and deletion of reads smaller than a preset base number on the first reverse Fasta file to obtain a reverse processed Fasta file; Performing file correction on the forward processed Fasta file to obtain the forward corrected Fasta file, and performing file correction on the reverse processed Fasta file to obtain the reverse corrected Fasta file; Successively performing primary sequence alignment and secondary sequence alignment on the forward corrected Fasta file to obtain the second forward Fasta file, and successively performing primary sequence alignment and secondary sequence alignment on the reverse corrected Fasta file to obtain the second reverse Fasta file.

2. The method for detecting dynamic mutations of a gene according to claim 1, characterized in that, Before the step of obtaining the covered BAM file by extracting the reads covered by each dynamic mutation from the sample alignment BAM file respectively based on the preset dynamic mutation position definition file, the following steps are further included: Defining a Bed file according to the positions of the genes corresponding to each of the dynamic mutations in the preset reference genome respectively to obtain the dynamic mutation position definition file; Among them, the value range of the dynamic mutations includes: UGT1A1 mutation, NOTCH2NLC mutation, AR mutation, ATXN1 mutation, ATXN7 mutation, TBP mutation, ATN1 mutation, ATXN2 mutation, ATXN3 mutation, CACNA1A mutation, and JPH3 mutation.

3. The method for detecting dynamic mutations of the gene according to claim 1, wherein The step of performing file correction on the forward processed Fasta file to obtain the forward corrected Fasta file and performing file correction on the reverse processed Fasta file to obtain the reverse corrected Fasta file includes: Grabbing the reads covered by each dynamic mutation and the unsequenced reads from the forward processed Fasta file to obtain individual single-gene forward data; Performing a merging process on the single-gene forward data corresponding to the ATXN1 mutation and the single-gene forward data corresponding to the ATXN3 mutation to obtain the forward data to be analyzed; Obtaining the first gene characteristic sequence corresponding to the ATXN1 mutation and the second gene characteristic sequence corresponding to the ATXN3 mutation; Searching for the covered reads and the unsequenced reads from the forward data to be analyzed according to the first gene characteristic sequence to obtain the first forward grouped data, and searching for the covered reads and the unsequenced reads from the forward data to be analyzed according to the second gene characteristic sequence to obtain the second forward grouped data; Taking the first forward grouped data, the second forward grouped data, and the individual single-gene forward data other than the forward data to be analyzed as the forward corrected Fasta file; Grabbing the reads covered by each dynamic mutation and the unsequenced reads from the reverse processed Fasta file to obtain individual single-gene reverse data; Performing a merging process on the single-gene reverse data corresponding to the ATXN1 mutation and the single-gene reverse data corresponding to the ATXN3 mutation to obtain the forward data to be analyzed; Searching for the covered reads and the unsequenced reads from the reverse data to be analyzed according to the first gene characteristic sequence to obtain the first reverse grouped data, and searching for the covered reads and the unsequenced reads from the reverse data to be analyzed according to the second gene characteristic sequence to obtain the second reverse grouped data; Taking the first reverse grouped data, the second reverse grouped data, and the individual single-gene reverse data other than the forward data to be analyzed as the reverse corrected Fasta file.

4. The method for detecting dynamic mutation of a gene according to claim 1, wherein The steps of performing primary sequence alignment and secondary sequence alignment on the forward corrected Fasta file in sequence to obtain the second forward Fasta file, and performing primary sequence alignment and secondary sequence alignment on the reverse corrected Fasta file in sequence to obtain the second reverse Fasta file include: Obtain each preset first motif, where each of the first motifs is data extracted according to the characteristics of dynamic mutation; Align each of the first motifs with the forward corrected Fasta file respectively, and remove the sequences at both ends of the repeated sequences in the reads of the forward corrected Fasta file based on the alignment position information to obtain a primary forward alignment result; Align each of the first motifs with the reverse corrected Fasta file in sequence respectively, and remove the sequences at both ends of the repeated sequences in the reads of the forward corrected Fasta file based on the alignment position information to obtain a primary reverse alignment result; Obtain each preset second motif, where each of the second motifs is data extracted according to the characteristics of dynamic mutation, and the length of the second motif is greater than the length of the first motif; Align each of the second motifs with the primary forward alignment result in sequence respectively, and remove the sequences at both ends of the repeated sequences in the reads of the primary forward alignment result based on the alignment position information to obtain the second forward Fasta file; Align each of the second motifs with the primary reverse alignment result in sequence respectively, and remove the sequences at both ends of the repeated sequences in the reads of the primary reverse alignment result based on the alignment position information to obtain the second reverse Fasta file.

5. The method for detecting dynamic mutations of a gene according to claim 1, wherein The steps of generating a target detection result according to the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file include: Execute a preset data summarization script to summarize the data of the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file to obtain a summarization result; Use a preset detection result generation rule to generate a detection result from the summarization result to obtain the target detection result; Among them, the summary results include: the name of the dynamic mutation, the name of the sequencing reads, the direction of the sequencing reads, the trimmed reads sequence, the 10 bp bases on the left side of the original reads, the 10 bp bases on the right side of the original reads, the length of the trimmed reads sequence, the detected number of repeats without spacer sequences, the Levenshtein distance score, whether the left side of the trimmed reads sequence is sequenced through, and whether the right side of the trimmed reads sequence is sequenced through. The name of the sequencing reads and the trimmed reads sequence are both data extracted from the second forward Fasta file or the second reverse Fasta file. The direction of the sequencing reads is data extracted from the forward BAM file or the reverse BAM file. The 10 bp bases on the left side of the original reads and the 10 bp bases on the right side of the original reads are both data extracted from the forward corrected Fasta file or the reverse corrected Fasta file. The length of the trimmed reads sequence and the detected number of repeats without spacer sequences are both data calculated from the trimmed reads sequence. The Levenshtein distance score is a score calculated based on the trimmed reads sequence and a preset reference sequence. Whether the left side of the trimmed reads sequence is sequenced through is data determined based on the 10 bp bases on the left side of the original reads. Whether the right side of the trimmed reads sequence is sequenced through is data determined based on the 10 bp bases on the right side of the original reads; The target detection results include: the name of the dynamic mutation, the reference value of the mutation repeat count, the number of original reads, the number of reads supporting the variation, the number of reads not sequenced through with more than 40 repeats, and the detection conclusion. The number of original reads, the number of reads supporting the variation, and the number of reads not sequenced through with more than 40 repeats are all data calculated based on the summary results. The detection conclusion is data calculated based on the reference value of the mutation repeat count, the number of original reads, the number of reads supporting the variation, and the number of reads not sequenced through with more than 40 repeats.

6. A device for detecting dynamic mutations of a gene, characterized in that, A dynamic mutation detection device for a gene according to any one of claims 1 to 5, the device comprising: A data acquisition module for acquiring a sample alignment BAM file corresponding to the target next-generation sequencing data; A covered BAM file determination module for extracting the reads covered by each dynamic mutation from the sample alignment BAM file respectively based on a preset dynamic mutation position definition file to obtain a covered BAM file; A splitting module for performing unidirectional splitting and sequence information extraction on the covered BAM file to obtain a forward BAM file, a first forward Fasta file, a reverse BAM file, and a first reverse Fasta file; A processing module, configured to perform file processing, file correction, primary sequence alignment, and secondary sequence alignment on the first forward Fasta file and the first reverse Fasta file in sequence, to obtain a forward corrected Fasta file, a second forward Fasta file, a reverse corrected Fasta file, and a second reverse Fasta file; A detection result generation module, configured to generate a detection result based on the forward BAM file, the reverse BAM file, the second forward Fasta file, the second reverse Fasta file, the forward corrected Fasta file, and the reverse corrected Fasta file, to obtain a target detection result.

7. A computer-readable storage medium storing a computer program, which when executed by a processor, causes the processor to execute the steps of the method according to any one of claims 1 to 5.

8. A computer device comprising a memory and a processor, the memory storing a computer program, which when executed by the processor, causes the processor to execute the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Tumor somatic cell mutation site detection method and device

    CN111180010A