Polynucleotide joint variation detection method, device, terminal and medium

By standardizing the processing of genomic variation data and screening for quality control indicators, the problem of detecting and annotating polynucleotide combined variations was solved, ensuring accurate detection and reasonable reporting of similar mutations in molecular diagnostics and improving diagnostic accuracy.

CN117352051BActive Publication Date: 2026-02-06SHANGHAI RIGEN BIOTECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310850851.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-11
Publication Date
2026-02-06
Estimated Expiration
2043-07-11

AI Technical Summary

Technical Problem

In existing technologies, polynucleotide combined mutations have not been correctly detected and annotated in molecular diagnostics, leading to erroneous diagnostic results. In particular, adjacent amino acid mutations may be misjudged as independent mutations, affecting the accuracy of clinical diagnosis.

Method used

The system employs a detection module, a standardization module, an annotation module, a quality assessment module, and a quality control module. It utilizes variant detection and annotation tools to standardize genomic variant data, screen for polynucleotide joint variants, and conduct quality control and hazard screening based on quality control indicators such as SOR values ​​to ensure the reasonable detection and annotation of mutations on the same chromosome.

Benefits of technology

This enables the reasonable detection and accurate annotation of similar mutations, avoiding missed diagnoses and false positives, and improving the accuracy of molecular diagnostics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117352051B_ABST
    Figure CN117352051B_ABST
Patent Text Reader

Abstract

The application provides a polynucleotide joint variation detection method, device, terminal and medium, the method comprises the following steps: according to the sequencing data aligned to the reference genome, using a variation detection tool to detect potential polynucleotide joint variations and save them in a genomic variation data file; filter out SNV and Indel that do not belong to polynucleotide joint variation occurred in single base site in the genomic variation data file, and use a variation annotation tool to annotate the polynucleotide joint variation in the filtered genomic variation data file; extract the relevant annotation results of the concerned transcript and integrate the depth information and frequency information of the polynucleotide joint variation, and evaluate the quality of the polynucleotide joint variation based on one or more preset variation quality control indicators; filter out mutations that do not meet the quality control standards in sequencing depth, mutation frequency, genotype quality value and SOR value, and on this basis, screen harmful polynucleotide joint variations that affect gene coding and have low population frequency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biomedical technology, in particular to a multiple nucleotide variants detection method and device, terminal and medium. BACKGROUND

[0002] With the advent of the era of precision medicine, accurate detection and annotation of genetic variants is a basic prerequisite for auxiliary diagnosis and clinical medication. Multiple nucleotide variants (MNV) is a clinically and biologically important type of genetic variant, which is defined as two or more adjacent variants present on the same haplotype of an individual.

[0003] According to the gene mutation naming rules of the Human Genome Society (HGVS) and the latest proposal of SVD-WG010, multiple nucleotide mutations located in a haplotype should be jointly detected and reported if they are within a range of one base or affect the coding of adjacent 2 or more amino acids.

[0004] However, in the current mainstream second-generation sequencing (NGS) data analysis process (such as GATK bestpractices), multiple nucleotide variants are usually not correctly detected and annotated. For example, the codon GCT encoding Ala (alanine) has a two-base joint mutation, changing to TTT, resulting in Ala (alanine) being replaced by Phe (phenylalanine). If the conventional mutation detection process is used, two mutations may be detected: the codon GCT is mutated to TCT, resulting in Ala (alanine) being replaced by Ser (serine); and the codon GCT is mutated to GTT, resulting in Ala (alanine) being replaced by Val (valine). False mutation detection will lead to false annotation results, which may eventually lead to missed diagnosis or false positive diagnosis cases. In addition, adjacent amino acid mutations on the same haplotype may also have clinical significance. The 600-601 amino acids of the BRAF gene are commonly replaced by an aspartic acid, which may have carcinogenic properties in melanoma, thyroid cancer and lung cancer.

[0005] Therefore, there is an urgent need in the art for an effective detection tool that can reasonably detect similar mutations in molecular diagnosis, distinguish whether they occur on the same chromosome, and reasonably annotate and report. SUMMARY

[0006] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a multiple nucleotide variant detection method and device, terminal and medium, which solves the technical problem of how to provide an effective detection tool that can reasonably detect similar mutations in molecular diagnosis, distinguish whether they occur on the same chromosome, and reasonably annotate and report.

[0007] To achieve the above object and other related objects, the first aspect of the present application provides a method for detecting combined polynucleotide variations, comprising: detecting potential combined polynucleotide variations in a genomic variation data file using a variation detection tool according to sequencing data aligned with a reference genome; standardizing the mutation format in the genomic variation data file; filtering out SNV and Indel at single base sites not belonging to combined polynucleotide variations in the genomic variation data file, and annotating the combined polynucleotide variations in the filtered genomic variation data file using a variation annotation tool; extracting relevant annotation results of the transcripts of interest, integrating the depth information and frequency information of the combined polynucleotide variations, and evaluating the quality of the combined polynucleotide variations based on one or more preset variation quality control indicators; the variation quality control indicators at least include SOR value for evaluating whether the mutation has strand bias; filtering out mutations that do not meet the quality control standards of sequencing depth, mutation frequency, genotype quality value and SOR value, to control the quality of the combined polynucleotide variations, so as to screen harmful combined polynucleotide variations that affect gene coding and have low population frequency.

[0008] In some embodiments of the first aspect of the present application, the way of detecting the combined polynucleotide variations affecting the coding of the same amino acid comprises: setting the detection distance of the combined polynucleotide variations within 1 base pair; and the way of detecting the potential combined polynucleotide variations affecting the coding of adjacent amino acids comprises: setting the detection distance of the combined polynucleotide variations within 4 base pairs.

[0009] In some embodiments of the first aspect of the present application, the method further comprises: standardizing the mutation format in the genomic variation data file using bcftools and vcfallelicprimitives tools.

[0010] In some embodiments of the first aspect of the present application, the information of the combined polynucleotide variations in the filtered genomic variation data file annotated using the variation annotation tool comprises any one of the following information or a combination of multiple information: HGVS mutation format information at DNA level and protein level, gene information, transcript information, mutation type information, population frequency information, harmfulness prediction information, disease database record information.

[0011] In some embodiments of the first aspect of the present application, the calculation method of the SOR value for evaluating whether the mutation has strand bias comprises:

[0012]

[0013] R = (refFw / refRv) / (altFw / altRv);

[0014]

[0015]

[0016] wherein refFw represents the number of DNA sequence fragments supporting the reference genotype, aligned to the positive strand; refRv represents the number of DNA sequence fragments supporting the reference genotype, aligned to the negative strand; altFw represents the number of DNA sequence fragments supporting the mutant genotype, aligned to the positive strand; and altRv represents the number of DNA sequence fragments supporting the mutant genotype, aligned to the negative strand.

[0017] In some embodiments of the first aspect of the present application, the filtering of the mutations that do not meet the quality control criteria includes identifying as the somatic mutations that do not meet the quality control criteria the variants that meet any one of the following conditions: sequencing depth < 50X, mutant frequency < 1%, genotype quality value < 50, and SOR value > 3.

[0018] In some embodiments of the first aspect of the present application, the method further comprises, after the quality control of the polynucleotide joint variants, performing the following: for the polynucleotide joint variants that pass the quality control, outputting the report after evaluating the harmfulness thereof.

[0019] In some embodiments of the first aspect of the present application, for the polynucleotide joint variants that potentially affect the coding of the same or adjacent amino acids, or are located at the splicing sites, the mutations with the population frequency in the genomic aggregation database lower than a preset threshold are screened and reported.

[0020] To achieve the above object and other related objects, the second aspect of the present application provides a polynucleotide joint variation detection device, comprising: a detection module configured to detect potential polynucleotide joint variations in a genomic variation data file using a variation detection tool according to sequencing data aligned with a reference genome; a standardization module configured to standardize mutation formats in the genomic variation data file; an annotation module configured to filter SNV and Indel at single base sites not belonging to polynucleotide joint variations in the genomic variation data file after identifying the SNV and Indel through a script, and to annotate polynucleotide joint variations in the filtered genomic variation data file using a variation annotation tool; a quality assessment module configured to extract relevant annotation results of a transcript of interest, integrate depth information and frequency information of polynucleotide joint variations, and assess the quality of polynucleotide joint variations based on one or more preset variation quality control indicators; the variation quality control indicators at least include a SOR value used to assess whether the mutation has strand bias; a quality control module configured to filter somatic mutations that do not meet the quality control standards in terms of sequencing depth, mutation frequency, genotype quality value and SOR value, so as to control the quality of polynucleotide joint variations; and a harmfulness screening module configured to screen harmful polynucleotide joint variations that affect gene coding and have low population frequency.

[0021] To achieve the above object and other related objects, the third aspect of the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the polynucleotide joint variation detection method.

[0022] To achieve the above object and other related objects, the fourth aspect of the present application provides an electronic terminal, comprising: a processor and a memory; the memory is configured to store a computer program, and the processor is configured to execute the computer program stored in the memory, so that the terminal executes the polynucleotide joint variation detection method.

[0023] As described above, the polynucleotide joint variation detection method, device, terminal and medium of the present application have the following beneficial effects: the technical solution of the present application can reasonably detect similar mutations, distinguish whether they occur on the same chromosome, and reasonably annotate and report. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 A flowchart of a polynucleotide joint variation detection method according to an embodiment of the present application is shown.

[0025] Figure 2A A selection logic diagram of MNV detection range according to an embodiment of the present application is shown.

[0026] Figure 2BA BAM visualization diagram showing two mutations of the same chromosome in an embodiment of the present application.

[0027] Figure 3 A structural diagram of a polynucleotide joint variation detection device in an embodiment of the present application.

[0028] Figure 4 A structural diagram of an electronic terminal in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The above and other advantages and features of the present application will become apparent from the following description of the embodiments, taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the application. This description is given for the sake of example and the details are not intended to limit the present application. The present application can be implemented in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the application to those skilled in the art. In the drawings, like reference numerals indicate like elements through the several views. The drawings provided herein are for illustration purposes only and, therefore, the drawings are not necessarily made to scale.

[0030] It should be noted that in the following description, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration various embodiments of the present application. It is to be understood that other embodiments can be utilized and that mechanical, electrical, and structural changes can be made without departing from the spirit and scope of the present application. The following detailed description is not to be taken in a limiting sense, and the scope of the embodiments of the present application are defined only by the claims. The summary of the present application is not intended to limit the scope of the application. Spatially relative terms, such as "upper", "lower", "left", "right", "below", "above", "front", "back", and the like, can be used herein for ease of description to describe one element or feature's relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientations depicted in the figures. For example, if the device described herein is turned over or flipped, elements described as "below" or "beneath" other elements or features would then be oriented "above" or "over" the other elements or features. The device can be otherwise oriented (e.g., rotated 90° or at other

[0031] In the present application, unless specifically defined otherwise, the terms "mount", "connect", "connection", "fixed", "fixedly", and the like, are used broadly and encompass both direct and indirect mounting, mechanical and electrical connections, fixed and removable connections, and the like, as appropriate for the context in which the terms are used. Unless otherwise specifically explained, the terms "mount", "connect", "connection", "fixed", "fixedly", and the like, should not be construed as requiring direct mounting, mechanical or electrical connections, or the like.

[0032] Also, as used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including", when used herein, specify the presence of stated features, operations, elements, components, items, and / or objects, but do not preclude the presence or addition of one or more other features, operations, elements, components, items, and / or objects. The terms "or" and "and / or" as used herein are to be interpreted as inclusive, i.e., as meaning one or any combination of items. Thus, "A, B or C" or "A, B and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B and C". An exception to this definition will occur only when a combination of elements, functions, or operations is in some way inherently mutually exclusive.

[0033] To solve the problems in the above background art, the present application provides a multiple nucleotide variants detection method, device, terminal and medium, which can reasonably detect similar mutations, distinguish whether they occur on the same chromosome, and make reasonable annotation and reporting in molecular diagnosis. Meanwhile, in order to make the purpose, technical scheme and advantages of the present application more clear and explicit, the technical scheme of the embodiments of the present application will be further described in detail in conjunction with the following embodiments and the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the application.

[0034] Before the present application is further described, the terms and phrases involved in the embodiments of the present application are explained, which are applicable to the following explanations:

[0035] <1> MNV (Multiple Nucleotide Variants): multiple nucleotide variants, a clinically and biologically important genetic variation type, defined as two or more adjacent variations present on the same haplotype in an individual.

[0036] <2> BAM (Binary Alignment Map): a binary format alignment file, usually converted from a SAM (Sequence Alignment Map) file. The BAM file contains the alignment results of the original DNA sequence and the measured sequence, and is used for alignment of genomic sequences and storage of sequencing data.

[0037] <3> VCF (Variant Call Format): a text file for storing genomic variation data, a file containing variation information obtained after detecting variations in a SAM or BAM file.

[0038] <4>HGVS (Human Genome Variation Society): Human Genome Variation Society, which has established a systematic gene mutation naming method, is the currently recognized naming rule in the academic field.

[0039] <5>Transcript: One or more mature mRNAs that can be used to encode proteins formed by transcription of a gene.

[0040] <6>Reads: Read length refers to the base sequence obtained by single sequencing of a sequencer, that is, a sequence of ATCGGGTA and the like.

[0041] <7>SNV (Stable Nuclear Variant): A biologically active ribonucleic acid that is a special ribonucleic acid that naturally occurs in variation, and this type of variation is stable and does not change. SNV is a stable variation that can regulate gene expression and affect gene function.

[0042] <8>Indel (Insertion and Deletioni): Refers to the insertion or deletion of a small fragment sequence at a certain position in the genome.

[0043] The embodiment of the present application provides a polynucleotide combined variation detection method, a polynucleotide combined variation detection method system, and a storage medium for storing an executable program for implementing the polynucleotide combined variation detection method. In terms of the implementation of the polynucleotide combined variation detection method, the embodiment of the present application will illustrate an exemplary implementation scenario of polynucleotide combined variation detection.

[0044] As shown in Figure 1 , a flowchart of a polynucleotide combined variation detection method in the embodiment of the present application is shown. The polynucleotide combined variation detection method in the present embodiment mainly includes the following steps:

[0045] Step S1: According to the sequencing data aligned to the reference genome, use the variation detection tool to detect potential polynucleotide combined variations and store them in the genomic variation data file.

[0046] In the embodiment of the present application, according to the sequencing data aligned to the reference genome in the BAM file, the combined mutation on a chromosome is detected using the variation detection tool, so as to detect the polynucleotide combined variation that potentially affects the coding of the same amino acid or the coding of adjacent amino acids, so as to obtain the corresponding VCF file.

[0047] More preferably, the manner of detecting potential influence on the coding of the same amino acid comprises the following: setting the detection distance of the polynucleotide joint variation to within 1 base pair; the manner of detecting potential influence on the coding of adjacent amino acids comprises the following: setting the detection distance of the polynucleotide joint variation to within 4 base pairs. Figure 2A The selection logic of the MNV detection range is illustrated as follows: the nucleotides arranged in a row are C (deoxy cytosine nucleotide), A (deoxy adenine nucleotide), G (deoxy guanosine nucleotide), C (deoxy cytosine nucleotide), A (deoxy adenine nucleotide), and G (deoxy guanosine nucleotide). The distance between each nucleotide and the adjacent nucleotide is 0 bp (i.e. 0 base pairs). When the distance between two nucleotides is less than or equal to 1 bp, it may affect the coding of the same amino acid, as shown by the first C and the first G in the figure. When the distance between two nucleotides is less than or equal to 4 bp, it may affect the coding of adjacent amino acids, as shown by the first C and the second G in the figure.

[0048] It should be understood that the BAM (Binary Alignment Map) file is a binary format alignment file, hereinafter referred to as the BAM file, which is usually converted from a SAM (Sequence Alignment Map) file. The BAM file contains the alignment results of the original DNA sequence and the measured sequence, and is used for the alignment of genomic sequences and the storage of sequencing data. The BAM file mainly consists of four parts: a file header, an alignment text, an annotation record and a sequence record; the file header describes the source and structure of the BAM file, the alignment text records the alignment results of the original DNA sequence and the measured sequence, and the annotation record and the sequence record contain the annotation information and the original sequence information of the alignment text, respectively.

[0049] The VCF (Variant Call Format) file is a text file used to save genomic variant data, referred to as a genomic variant data file, which is used to store the variant information obtained after detecting the variants of the SAM or BAM file.

[0050] In the embodiments of the present application, the variant detection tool includes but is not limited to the following variant detection software: freebayes software and GATK software, etc., and the embodiments of the present application are not limited.

[0051] Step S2: standardizing the mutation format in the genomic variant data file.

[0052] In the embodiments of the present application, the mutation format in the VCF file can be standardized by using a standardization tool such as the bcftools tool and the vcfallelicprimitives tool.

[0053] The corresponding description is as follows: bcftools is a software tool set for operating and processing VCF files, which can be used for filtering, annotation, statistics and visualization of SNPs and Indels. SNP (single nucleotide polymorphism) refers to DNA sequence polymorphism caused by variation of a single nucleotide at the genetic level, including single base conversion and transversion; Indel is the general term for small insertion and deletion. The standardization parameters of the flow are as follows:

[0054] 1) "bcftools norm-m-file.vcf>step1.vcf", the specific function is to left-align and normalize Indel coordinates; check whether the ref (reference genotype) is consistent with the reference genome; split multiple alleles at the same position into multiple rows.

[0055] 2) "vcfallelicprimitives-mkg step1.vcf>step2.vcf", the specific function is to ensure the correctness of genotype, sequencing depth and other information when splitting multiple alleles at the same position.

[0056] Step S3: The SNV and Indel of the single base site of the polynucleotide joint variation in the genomic variation data file are identified by script and filtered, and the polynucleotide joint variation in the filtered genomic variation data file is annotated by using the variation annotation tool.

[0057] More preferably, the script identification and filtering method of SNV and Indel of the single base site of the polynucleotide joint variation includes: when the length of ref (reference genotype) and alt (mutant genotype) is 1bp, it is identified as SNV or Indel, and then it is deleted.

[0058] In the embodiments of the present application, the variant annotation tool includes but is not limited to VEP annotation tool, SnpEff annotation tool, Annovar annotation tool, Oncotator annotation tool, etc. Taking the VEP annotation tool as an example, 1 human genome contains nearly 3500000 SNV mutations and 1000 copy number variations, about 20000-25000 of which are in coding regions, 10000 of which are amino acid coding changes, and only 50-100 of which are protein truncation or loss of function. The VEP (Variant Effect Predictor) annotation tool can be used to annotate and analyze most types of genomic variations in the coding and non-coding regions of the genome. The VEP (Variant Effect Predictor) annotation tool annotates two major genomic variations: (1) sequence variations with specific and explicit changes (including SNV, insertion, deletion, multiple base pair substitution, microsatellite and tandem repeat, etc.); (2) larger structural variants (length greater than 50 nucleotides), including structural variants with changes in DNA copy number or insertion and deletion.

[0059] In the embodiments of the present application, the information of the polynucleotide joint variation in the filtered genomic variation data file annotated by the variant annotation tool includes but is not limited to: HGVS mutation format information at the DNA level and the protein level, the gene information, the transcript information, the mutation type information, the population frequency information, the hazard prediction information, the disease database record information, etc.

[0060] Step S4: Extracting the relevant annotation results of the concerned transcript and integrating the depth information and the frequency information of the polynucleotide joint variation, and evaluating the quality of the polynucleotide joint variation based on one or more preset variant quality control indicators; the variant quality control indicators at least include SOR value used to evaluate whether the mutation has strand bias.

[0061] Specifically, the SOR (Strand Odds Ratio) value is a description of strand specificity, which can well evaluate whether the mutation has strand bias, and the calculation method is as follows:

[0062]

[0063] R=(refFw / refRv) / (altFw / altRv); equation (2)

[0064]

[0065]

[0066] Wherein, refFw represents the number of DNA sequence fragments supporting the reference genotype and aligned to the positive strand; refRv represents the number of DNA sequence fragments supporting the reference genotype and aligned to the negative strand; altFw represents the number of DNA sequence fragments supporting the mutant genotype and aligned to the positive strand; and altRv represents the number of DNA sequence fragments supporting the mutant genotype and aligned to the negative strand. The greater the SOR value, the greater the strand preference of the mutation.

[0067] It should be understood that double-stranded complementary DNA is divided into a positive strand and a negative strand, the positive strand is a forwards strand, and the negative strand is a reverse strand; some genes are defined on the forwards strand, meaning that the sequence of the transcript corresponding to the gene is exactly the same as the base sequence of the forwards strand from 5' to 3'; and some genes are defined on the reverse strand, meaning that the sequence of the transcript corresponding to the genes is exactly the same as the sequence of the reverse strand from 5' to 3'. The number of DNA sequence fragments, also referred to as the number of reads, refers to the number of DNA sequence fragments read by a sequencing instrument during the process of gene sequencing, and this number is very important for the quality and accuracy of gene sequencing, and directly affects the understanding and analysis of the genome.

[0068] Step S5: filtering mutations that do not meet the quality control standards of sequencing depth, mutant frequency, genotype quality value and SOR value, to perform quality control on the polynucleotide joint variation.

[0069] Exemplarily, the somatic mutations that do not meet the quality control standards in the present application include variations that meet any one of the following conditions:

[0070] Condition 1) sequencing depth < 50X;

[0071] Condition 2) mutant frequency VAF < 1%;

[0072] Condition 3) genotype quality value QUAL < 50;

[0073] Condition 4) SOR (Strand Odds Ratio) value > 3.

[0074] It should be noted that different types of mutations have corresponding different quality control standards, and the above four conditions are the quality control standards set for somatic mutations, and another set of filtering standards can be set for other types of mutations (such as germline mutations), and the embodiments of the present application are not limited specifically.

[0075] Further, the polynucleotide joint variation detection method provided by the embodiments of the present application further performs the following after performing the above step S5: for the polynucleotide joint variation that passes the quality control, outputting a report after evaluating the harmfulness thereof.

[0076] Further, for MNVs that potentially affect the same or adjacent amino acid coding, or located at splice site, report mutations with population frequency lower than a pre-set threshold in the genomic aggregate database. For example, report MNVs that passed quality control for further evaluation of their deleteriousness; for MNVs that potentially affect the same or adjacent amino acid coding, or located at splice site, report mutations with population frequency lower than 1% in gnomAD.

[0077] For the convenience of those skilled in the art, the following will be described in combination with Figure 2B and a specific variant detection embodiment:

[0078] The conventional mutation detection process is directed to Figure 2B The sample shown in FIG. 1 can detect two mutations in Table 1, one SNP and one Indel, which respectively cause the lysine (K) at position 601 of the BRAF gene to be replaced by asparagine (N), and the valine (V) at position 600 and the lysine (K) at position 601 to be replaced by a glutamic acid (E). Since the two mutations together affect the lysine (K) at position 601, the results of separate detection and annotation are inaccurate.

[0079] The MNV resulting from the detection process of the present application is shown in Table 2, which causes the valine (V) at position 600 and the lysine (K) at position 601 to be replaced by an asparagine (D). And from the SOR, the MNV has no strand bias, the sequencing depth and the mutation frequency meet the pre-set quality control threshold. The genotype quality value extracted from the VCF file is also higher than the pre-set quality control threshold (not shown here). In addition, the mutation frequency of the MNV is consistent with the frequency of the two single-point mutations, so it can be judged that the two mutations occur in the same chromosome and belong to a haplotype, and should be reported jointly.

[0080] Table 1: Results of the conventional mutation detection process

[0081] Chr Chr7 Chr7 Start 140453132 140453134 End 140453132 140453136 Ref T TCA Alt A - DP 31536 31536 AD 12398 12398 VAF 39.31% 39.31% SOR 0.69 0.69 Gene BRAF BRAF Hgvsc c.1803A>T c.1799_1801del Hgvsp p.K601N p.V600_K601delinsE Mutation_Type Missense_variant Inframe_deletion gnomAD_AF 0 -

[0082] Wherein:

[0083] 1) Chr: Chromosome number;

[0084] 2) Start: Mutation located at the start position of the reference genome;

[0085] 3) End: Mutation located at the end position of the reference genome;

[0086] 4) Ref: Reference genotype;

[0087] 5) Alt: Mutation genotype;

[0088] 6) DP: Sequencing depth (number of gene reads) at the location of the mutation;

[0089] 7) AD: Number of gene reads supporting the mutation;

[0090] 8) VAF: Mutation frequency, i.e., AD / DP;

[0091] 9) SOR: Chain bias indicator;

[0092] 10) Gene: The gene containing the mutation;

[0093] 11) Hgvsc: A mutation annotation format at the DNA level defined by the Human Genome Society;

[0094] 12) Hgvsp: A mutation annotation format at the protein level defined by the Human Genome Society;

[0095] 13) Mutation_Type: Classified according to the location of the mutation on the gene and its effect on gene function;

[0096] 14) gnomAD_AF: Frequency of mutations annotated in the gnomAD database. "-" indicates that the mutation is not recorded in the database.

[0097] Table 2

[0098] Chr Chr7 Start 140453132 End 140453136 Ref TTTCA Alt AT DP 31555 AD 12399 VAF 39.29% SOR 0.69 Gene BRAF Hgvsc c.1799_1803delinsAT Hgvsp p.V600_K601delinsD Mutation_Type Protein_altering_variant gnomAD_AF -

[0099] like Figure 3 The diagram shows a schematic representation of a polynucleotide combined variant detection device according to an embodiment of the present invention. The detection device 300 in this embodiment includes: a detection module 301, a standardization module 302, an annotation module 303, a quality assessment module 304, a quality control module 305, and a hazard screening module 306.

[0100] The detection module 301 is used to detect potential polynucleotide joint variations using a variation detection tool based on the sequencing data of the comparison reference genome and store them in the genome variation data file.

[0101] The standardization module 302 is used to standardize the mutation format in the genomic variation data file.

[0102] The annotation module 303 is used to filter out SNVs and Indels that do not belong to polynucleotide joint variations in the genome variation data file after being identified by a script, and to annotate the polynucleotide joint variations in the filtered genome variation data file using a variation annotation tool.

[0103] The quality assessment module 304 is used to extract relevant annotation results of the transcripts of interest and integrate the depth and frequency information of polynucleotide joint variants, and to assess the quality of polynucleotide joint variants based on one or more preset variant quality control indicators; the variant quality control indicators include at least the SOR value used to assess whether the mutation has strand bias.

[0104] The quality control module 305 is used to filter somatic mutations that do not meet the quality control standards in terms of sequencing depth, mutation frequency, genotype quality value, and SOR value, so as to perform quality control on polynucleotide combined variants.

[0105] The hazard screening module 306 is used to screen for harmful polynucleotide combined variants that affect gene coding and have a low frequency in the population.

[0106] It should be noted that the polynucleotide joint variant detection device provided in the above embodiments is only illustrated by the division of the above-described program modules when performing polynucleotide joint variant detection. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the polynucleotide joint variant detection device and the polynucleotide joint variant detection method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0107] The polynucleotide joint variant detection method provided in this invention can be implemented on the terminal side or the server side. For the hardware structure of the polynucleotide joint variant detection terminal, please refer to [link to relevant documentation]. Figure 4 This is a schematic diagram of an optional hardware structure of an electronic terminal 400 provided in an embodiment of the present invention. The terminal 400 can be a mobile phone, computer device, tablet device, personal digital processing device, factory back-end processing device, etc. The electronic terminal 400 includes: at least one processor 401, a memory 402, at least one network interface 404, and a user interface 406. The various components in the device are coupled together through a bus system 405. It is understood that the bus system 405 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 405 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 4 The general will label all buses as bus systems.

[0108] The user interface 406 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.

[0109] It is to be understood that the memory 402 can be volatile or nonvolatile memory, or both. The nonvolatile memory can be read-only memory (ROM), programmable ROM (PROM), which is used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the present application is intended to include, but not be limited to, these and any other suitable type of memory.

[0110] The memory 402 in the embodiments of the present application is configured to store various types of data to support the operation of the electronic terminal 400. Examples of the data include any executable programs for operating on the electronic terminal 400, such as an operating system 4021 and an application program 4022. The operating system 4021 contains various system programs, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks. The application program 4022 can contain various application programs, such as a media player (MediaPlayer), a browser (Browser), and the like, for implementing various application services. The method for detecting joint variation of polynucleotides provided in the embodiments of the present application can be included in the application program 4022.

[0111] The method disclosed in the embodiments of the present application can be applied to the processor 401 or implemented by the processor 401. The processor 401 can be an integrated circuit chip having a processing capability. In the implementation process, each step of the above method can be completed by an integrated logic circuit or an instruction in the form of software in the processor 401. The processor 401 described above can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The processor 401 can implement or execute the disclosed methods, steps, and logic block diagrams in the embodiments of the present application. The general-purpose processor 401 can be a microprocessor or any conventional processor, etc. The steps of the method for optimizing accessories provided in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory. The processor reads the information in the memory and combines the hardware to complete the steps of the above method.

[0112] In exemplary embodiments, the electronic terminal 400 can be implemented by one or more of an application specific integrated circuit (ASIC), a DSP, a programmable logic device (PLD), a complex programmable logic device (CPLD), and a field programmable logic device (FPLD) for performing the aforementioned methods.

[0113] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by computer program related hardware. The aforementioned computer program can be stored in a computer readable storage medium. The program executes the steps of the above-mentioned method embodiments when executed; and the aforementioned storage medium includes ROM, RAM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage device, flash memory, U disk, mobile hard disk, or any other medium that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection can be appropriately referred to as a computer readable medium. For example, if the instructions are sent from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology such as infrared, radio and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology such as infrared, radio and microwave is included in the definition of the medium. However, it should be understood that the computer readable storage medium and the data storage medium do not include connections, carriers, signals or other transitory media, but are intended to be directed to non-transitory, tangible storage media. As used in the application, magnetic disks and optical disks include compact disks (CD), laser disks, optical disks, digital versatile disks (DVD), floppy disks and Blu-ray disks, wherein magnetic disks typically magnetically copy data, and optical disks optically copy data with a laser.

[0114] In the embodiments provided in the present application, the computer readable storage medium can include read only memory, random access memory, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage device, flash memory, U disk, mobile hard disk, or any other medium that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection can be appropriately referred to as a computer readable medium. For example, if the instructions are sent from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology such as infrared, radio and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology such as infrared, radio and microwave is included in the definition of the medium. However, it should be understood that the computer readable storage medium and the data storage medium do not include connections, carriers, signals or other transitory media, but are intended to be directed to non-transitory, tangible storage media. As used in the application, magnetic disks and optical disks include compact disks (CD), laser disks, optical disks, digital versatile disks (DVD), floppy disks and Blu-ray disks, wherein magnetic disks typically magnetically copy data, and optical disks optically copy data with a laser.

[0115] In summary, the present application provides a polynucleotide combined mutation detection method, device, terminal and medium. The technical scheme of the present application can reasonably detect similar mutations, distinguish whether they occur on the same chromosome, and perform reasonable annotation and reporting in molecular diagnosis. Therefore, the present application effectively overcomes the shortcomings of the prior art and has high industrial utilization value.

[0116] The above embodiments are only illustrative of the principles of the present application and their effects, and are not intended to limit the present application. Any modification or change made by any person skilled in the art without departing from the spirit and scope of the present application shall be covered by the claims of the present application.

Claims

1. A method for detecting polynucleotide combined variants, characterized in that, include: Based on the sequencing data of the reference genome, potential polynucleotide joint variants were detected using variant detection tools and stored in the genome variant data file; The mutation format in the genomic variation data file is standardized. SNVs and Indels occurring at single base sites that do not belong to polynucleotide joint variations in the genome variation data file are identified and filtered by a script, and polynucleotide joint variations in the filtered genome variation data file are annotated using a variation annotation tool. The relevant annotation results of the transcripts of interest are extracted and the depth and frequency information of polynucleotide combined variants are integrated. The quality of polynucleotide combined variants is evaluated based on one or more preset variant quality control indicators. The variant quality control indicators include at least the SOR value used to assess whether the mutation has strand bias. Mutations that do not meet the quality control standards in terms of sequencing depth, mutation frequency, genotype quality value, and SOR value are filtered out in order to control the quality of polynucleotide combined variants and screen for harmful polynucleotide combined variants that affect gene coding and have low population frequency. The methods for using variant detection tools to detect potential polynucleotide joint variants include: using variant detection tools to detect joint mutations on a chromosome, thereby detecting polynucleotide joint variants that may affect the coding of the same amino acid or adjacent amino acids; methods for detecting polynucleotide joint variants that may affect the coding of the same amino acid include: setting the detection distance of polynucleotide joint variants to within 1 base pair; methods for detecting polynucleotide joint variants that may affect the coding of adjacent amino acids include: setting the detection distance of polynucleotide joint variants to within 4 base pairs. The filtering of mutations that do not meet the quality control standards in terms of sequencing depth, mutation frequency, genotype quality value, and SOR value includes identifying somatic mutations that meet any of the following conditions as not meeting the quality control standards: sequencing depth < 50X, mutation frequency < 1%, genotype quality value < 50, and SOR value > 3.

2. The method for detecting polynucleotide combined variants according to claim 1, characterized in that, Also includes: The mutation formats in the genomic variation data files were standardized using the bcftools and vcfallelicprimitives tools.

3. The method for detecting polynucleotide combined variants according to claim 1, characterized in that, The information used to annotate polynucleotide joint variants in the filtered genomic variant data file using a variant annotation tool includes any one or a combination of the following: HGVS mutation format information at the DNA and protein levels, the gene information, transcript information, mutation type information, population frequency information, hazard prediction information, and disease database record information.

4. The method for detecting polynucleotide combined variants according to claim 1, characterized in that, The method for calculating the SOR value used to assess whether a mutation exhibits chain bias includes: R=(refFw / refRv) / (altFw / altRv); Wherein, refFw represents the number of DNA sequence fragments aligned to the positive strand that support the reference genotype; refRv represents the number of DNA sequence fragments aligned to the negative strand that support the reference genotype; altFw represents the number of DNA sequence fragments aligned to the positive strand that support the mutant genotype; and altRv represents the number of DNA sequence fragments aligned to the negative strand that support the mutant genotype.

5. The method for detecting polynucleotide combined variants according to claim 1, characterized in that, The method also includes performing the following after quality control of polynucleotide combined variants: for polynucleotide combined variants that pass quality control, outputting a report after assessing their hazard.

6. The method for detecting polynucleotide combined variants according to claim 5, characterized in that, For polynucleotide joint variations that potentially affect the same or adjacent amino acid codes, or are located at splice sites, mutations with a population frequency below a preset threshold in the genome aggregation database are screened and reported.

7. A multinucleotide combined variant detection device, characterized in that, include: The detection module is used to detect potential polynucleotide joint variants using mutation detection tools based on sequencing data from the reference genome and store them in the genome variant data file; A standardization module is used to standardize the mutation format in the genome variation data file; The annotation module is used to filter out SNVs and Indels occurring at single base sites that do not belong to polynucleotide joint variations in the genome variation data file, and to annotate the polynucleotide joint variations in the filtered genome variation data file using a variation annotation tool. The quality assessment module is used to extract relevant annotation results of transcripts of interest and integrate the depth and frequency information of polynucleotide joint variants, and to assess the quality of polynucleotide joint variants based on one or more preset variant quality control indicators. The mutation quality control indicators include at least the SOR value used to assess whether the mutation has chain bias; The quality control module is used to filter somatic mutations that do not meet the quality control standards in terms of sequencing depth, mutation frequency, genotype quality value, and SOR value, so as to perform quality control on polynucleotide combined variants. The hazard screening module is used to screen for harmful polynucleotide combined variants that affect gene coding and have low population frequency; The methods for using variant detection tools to detect potential polynucleotide joint variants include: using variant detection tools to detect joint mutations on a chromosome, thereby detecting polynucleotide joint variants that may affect the coding of the same amino acid or adjacent amino acids; methods for detecting polynucleotide joint variants that may affect the coding of the same amino acid include: setting the detection distance of polynucleotide joint variants to within 1 base pair; methods for detecting polynucleotide joint variants that may affect the coding of adjacent amino acids include: setting the detection distance of polynucleotide joint variants to within 4 base pairs. The filtering of mutations that do not meet the quality control standards in terms of sequencing depth, mutation frequency, genotype quality value, and SOR value includes identifying somatic mutations that meet any of the following conditions as not meeting the quality control standards: sequencing depth < 50X, mutation frequency < 1%, genotype quality value < 50, and SOR value > 3.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the polynucleotide joint variant detection method according to any one of claims 1 to 6.

9. An electronic terminal, characterized in that, include: Processor and memory; The memory is used to store computer programs; The processor is configured to execute the computer program stored in the memory to cause the terminal to perform the polynucleotide combined variant detection method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • High-throughput sequencing data analysis methods and devices

    CN109767810B

  • A method and apparatus for detecting mutations

    CN114596918B

  • A method and apparatus for detecting adjacent polynucleotide variations

    CN114974416B