Site detection limitation prediction method and device in whole exome sequencing, equipment and medium

By obtaining the coordinate information of target variant sites and reference genomes from the Clinvar database and genomic database, multi-dimensional sequence features are extracted. Combined with whole-exome capture probe files for simulated sequence alignment, the limitations of site detection in whole-exome sequencing are predicted, solving the problem of detection blind spots in whole-exome sequencing and improving the accuracy and efficiency of single-gene disease diagnosis.

CN121171351BActive Publication Date: 2026-02-24CIPHERGENE BEIJING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511699219.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-24
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

Existing whole-exome sequencing technology has blind spots in the diagnosis of single-gene diseases, and pathogenic variants are easily missed. In particular, pathogenic variants included in the OMIM and ClinVar databases are systematically missed. The detection sensitivity of exon margins and high GC regions is low, and there is a lack of systematic evaluation tools, resulting in insufficient diagnostic accuracy and a high risk of missed or misdiagnosed diagnoses.

Method used

By obtaining the coordinate information of the target variant site and the reference genome from the Clinvar database and the genome database, multi-dimensional sequence features are extracted. The distance value of the capture probe region is determined by combining the whole exome capture probe file. Simulated sequence sampling and alignment are performed to obtain the alignment quality value and consistency ratio, and the detection limitations of the target variant site are predicted.

Benefits of technology

Identifying variant sites with different detection risks before whole-exome sequencing can reduce the risk of missed and false detections, improve detection accuracy and stability, and reduce the cost and time extension caused by post-analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121171351B_ABST
    Figure CN121171351B_ABST
Patent Text Reader

Abstract

The application discloses a site detection limitation prediction method and device in whole exome sequencing, equipment and medium, relates to the field of biology and precision medicine high-throughput sequencing and variation detection technology, and comprises the following steps: extracting target sequence and original sequence from a reference genome based on coordinate information of a target variation site, obtaining multi-dimensional sequence characteristics of the target sequence; determining the distance value between each target variation site and the capture probe region, and counting the coverage of each target variation site according to the distance value; mutating the original sequence, simulating sampling of the mutated sequence, obtaining each simulated sequence, comparing the position of each simulated sequence with the reference genome, and obtaining the alignment quality value and alignment position consistency proportion of each simulated sequence; and predicting the detection limitation of the target variation site based on the multi-dimensional sequence characteristics, the coverage, the alignment quality value and the alignment position consistency proportion. The detection limitation of the site in the whole exome sequencing can be accurately predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of high-throughput sequencing and variant detection technology in biology and precision medicine, and particularly to methods, devices, equipment and media for predicting the limitations of site detection in whole exome sequencing. Background Technology

[0002] Next-Generation Sequencing (NGS) has propelled whole-exome sequencing (WES) to become a first-line diagnostic tool for monogenic diseases. However, due to its inherent limitations and the complexity of the genome, some known pathogenic variants are difficult to reliably capture. Technically, NGS suffers from short read lengths, PCR (Polymerase Chain Reaction) bias, and probe capture is susceptible to interference from various factors. Furthermore, bioinformatics algorithms have high false positive and false negative rates for some variants. At the genomic level, the human genome contains numerous homologous sequences, repetitive regions, and polymorphic regions, further exacerbating alignment uncertainty and potentially leaving some known pathogenic sites in the detection blind spot of conventional WES.

[0003] While WES is currently a primary molecular diagnostic tool for monogenic diseases with a positive rate of 25%–40%, it still has significant problems: First, pathogenic abnormalities recorded in OMIM (Online Mendelian Inheritance in Man) and the ClinVar database are sometimes missed by the system and require subsequent verification; second, the detection sensitivity of specific regions such as exon margins and high GC (Guanine-Cytosine) regions is significantly low, making it easy to miss pathogenic variants and affecting diagnostic accuracy.

[0004] More importantly, there is currently a lack of publicly available and systematic assessment tools, making it impossible to quantify whether specific pathogenic variants will be missed by standard procedures before WES testing. This reactive diagnostic approach not only prolongs the patient's diagnostic cycle and delays disease intervention, but also increases the risk of missed or misdiagnosed cases, failing to meet the clinical demand for accurate and efficient diagnosis and hindering the full realization of the advantages of WES.

[0005] In summary, accurately predicting the limitations of site detection in whole exome sequencing is a problem that needs to be solved in this field. Summary of the Invention

[0006] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for predicting site detection limitations in whole-exome sequencing, accurately predicting site detection limitations in whole-exome sequencing. The specific solution is as follows:

[0007] In a first aspect, this application discloses a method for predicting site detection limitations in whole-exome sequencing, including:

[0008] The coordinates of the target variant site and the reference genome were obtained from the Clinvar database and the genome database, respectively.

[0009] Based on the coordinate information, the target sequence is extracted from the reference genome, and the multi-dimensional sequence features of the target sequence are obtained;

[0010] Determine the distance between each target mutation site and the capture probe region determined based on the whole-exon capture probe file, and statistically analyze the coverage of each target mutation site based on the distance values;

[0011] Based on the coordinate information, the original sequence extracted from the reference genome is mutated to obtain a mutated sequence. The mutated sequence is then sampled to obtain simulated sequences. Each simulated sequence is then aligned with the reference genome to obtain the alignment quality value and the proportion of alignment position consistency of each simulated sequence.

[0012] The detection limitations of predicting the target variant site based on the multi-dimensional sequence features, the coverage, the alignment quality value, and the alignment position consistency ratio.

[0013] Optionally, extracting the target sequence from the reference genome based on the coordinate information includes:

[0014] Based on the coordinate information, a first preset number of base pairs, a second preset number of base pairs, a third preset number of base pairs, and a fourth preset number of base pairs are extracted from the upstream and downstream of the target mutation site from the reference genome to obtain a first target sequence, a second target sequence, a third target sequence, and a fourth target sequence with different sequence lengths.

[0015] Optionally, obtaining the multi-dimensional sequence features of the target sequence includes:

[0016] Determine the GC content of the first target sequence, and determine the GC content level of the first target sequence based on the GC content;

[0017] Determine the base frequency entropy of the second target sequence, and determine the base complexity level of the second target sequence based on the base frequency entropy;

[0018] Determine the base repetition type and number of repetitions of the third target sequence, and determine the base repetition level of the third target sequence based on the base repetition type and the number of repetitions;

[0019] A whole-genome sequence scan and alignment is performed on the fourth target sequence to identify similar homologous regions in the fourth target sequence, and the homology level of the fourth target sequence is determined based on the similar homologous regions.

[0020] Optionally, the step of calculating the coverage of each target mutation site based on the distance value includes:

[0021] If the distance between the target mutation site and the capture probe region is greater than a first preset distance threshold, then the coverage status of the target mutation site is determined to be a preset uncaptured category.

[0022] If the distance between the target mutation site and the capture probe region is greater than a second preset distance threshold but not greater than a first preset distance threshold, then the coverage status of the target mutation site is determined to be a preset incomplete capture category.

[0023] If the distance between the target mutation site and the capture probe region is not greater than a second preset distance threshold, then the coverage status of the target mutation site is determined to be a preset complete capture category.

[0024] Optionally, the step of performing simulated sampling on the mutated sequence to obtain simulated sequences includes:

[0025] The mutant sequence is simulated and sampled according to a fifth preset number of base pairs and a preset number of sampling times to obtain a simulated Fastq file containing each simulated sequence; wherein, the simulated Fastq file also includes original position information for marking each base pair in the simulated sequence in the original sequence.

[0026] Optionally, the step of aligning each of the simulated sequences with the reference genome to obtain the alignment quality value and the proportion of alignment position consistency for each of the simulated sequences includes:

[0027] The simulated sequences are aligned with the reference genome using a genome alignment tool to generate a simulated alignment BAM file containing the alignment quality value and actual alignment position information of each simulated sequence.

[0028] Based on the comparison results between the actual comparison location information and the original location information, a matching sequence is determined from the simulated sequence;

[0029] The alignment position consistency ratio is determined based on the number of aligned sequences and the number of simulated sequences.

[0030] Optionally, the detection limitations of predicting the target variant site based on the multi-dimensional sequence features, the coverage status, the alignment quality value, and the alignment position consistency ratio include:

[0031] The target correct alignment level of the target variant site is determined based on the relationship between the alignment quality value and each preset quality value threshold.

[0032] The target reliability level of the target variant site is determined based on the relationship between the consistency ratio of the comparison locations and each preset ratio threshold; wherein, the target reliability level characterizes the reliability of the comparison results;

[0033] The detection limitations of the target variant sites are predicted based on the multi-dimensional sequence features, the coverage status, the target correct alignment level, and the target reliability level.

[0034] Secondly, this application discloses a site detection limitation prediction device in whole-exome sequencing, comprising:

[0035] The signal acquisition module is used to obtain the coordinate information of the target variant site and the reference genome from the Clinvar database and the genome database, respectively.

[0036] The feature acquisition module is used to extract the target sequence from the reference genome based on the coordinate information, and to acquire the multi-dimensional sequence features of the target sequence;

[0037] The coverage status acquisition module is used to determine the distance value between each of the target mutation sites and the capture probe region determined based on the whole exon capture probe file, and to calculate the coverage status of each of the target mutation sites based on the distance value.

[0038] The position alignment module is used to mutate the original sequence extracted from the reference genome based on the coordinate information to obtain a mutated sequence, to perform simulated sampling on the mutated sequence to obtain each simulated sequence, and to perform position alignment of each simulated sequence with the reference genome to obtain the alignment quality value and the proportion of alignment position consistency of each simulated sequence.

[0039] The limitation prediction module is used to predict the detection limitations of the target variant site based on the multi-dimensional sequence features, the coverage status, the alignment quality value, and the alignment position consistency ratio.

[0040] Thirdly, this application discloses an electronic device, including:

[0041] Memory, used to store computer programs;

[0042] A processor is used to execute the computer program to implement the steps of the aforementioned disclosed method for predicting site detection limitations in whole-exome sequencing.

[0043] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned disclosed method for predicting the limitations of site detection in whole-exome sequencing.

[0044] The beneficial effects of this application are as follows: This application obtains the coordinate information of the target variant site and the reference genome from the Clinvar database and the genome database, respectively; extracts the target sequence from the reference genome based on the coordinate information, and obtains the multi-dimensional sequence features of the target sequence; determines the distance value between each target variant site and the capture probe region determined based on the whole-exome capture probe file, and statistically analyzes the coverage of each target variant site based on the distance value; mutates the original sequence extracted from the reference genome based on the coordinate information to obtain the mutated sequence, performs simulated sampling on the mutated sequence to obtain each simulated sequence, and compares the position of each simulated sequence with the reference genome to obtain the alignment quality value and the alignment position consistency ratio of each simulated sequence; and predicts the detection limitations of the target variant site based on the multi-dimensional sequence features, the coverage, the alignment quality value, and the alignment position consistency ratio. Therefore, this application obtains the coordinates of target variant sites and reference genomes from the Clinvar database and genome database, extracts target sequences and acquires multi-dimensional sequence features, determines the coverage of target variant sites by combining whole-exome capture probe files, obtains alignment quality values ​​and alignment position consistency ratios by simulating mutant sequence sampling and comparing with the reference genome, and finally predicts the detection limitations of target variant sites based on these indicators. Before actual whole-exome sequencing, by integrating multi-dimensional information such as sequence features, coverage, alignment quality, and consistency ratios, it can identify target variant sites with different detection risks in advance, thereby planning auxiliary detection methods for variant sites with different risks in advance. This effectively reduces the risk of missed or false detections caused by the characteristics of the site itself or technical limitations in actual sequencing, improves the accuracy and stability of whole-exome sequencing for the detection of clinically pathogenic variant sites, and reduces the increased costs and extended cycles caused by post-analysis. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0046] Figure 1 This is a flowchart of a method for predicting site detection limitations in whole-exome sequencing disclosed in this application;

[0047] Figure 2 This is a schematic diagram illustrating a specific method for obtaining multi-dimensional sequence features as disclosed in this application;

[0048] Figure 3 This is a schematic diagram of a specific capture coverage assessment disclosed in this application;

[0049] Figure 4 This is a specific position comparison diagram disclosed in this application;

[0050] Figure 5 This is a schematic diagram illustrating a specific detection limitation assessment disclosed in this application;

[0051] Figure 6 This is a schematic diagram illustrating the specific site detection limitation prediction disclosed in this application;

[0052] Figure 7 This is a schematic diagram of the structure of a site detection limitation prediction device in whole exome sequencing disclosed in this application;

[0053] Figure 8 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0055] High-throughput sequencing has propelled whole-exome sequencing (WES) to become a first-line diagnostic tool for monogenic diseases. However, due to its inherent limitations and the complexity of the genome, some known pathogenic variants are difficult to reliably capture. Technically, NGS suffers from short read lengths, PCR bias, and probe capture is susceptible to interference from various factors; furthermore, bioinformatics algorithms have high false positive and false negative rates for some variants. At the genomic level, the human genome contains numerous homologous sequences, repetitive regions, and polymorphic regions, further exacerbating alignment uncertainties and potentially leaving some known pathogenic sites in the detection blind spot of conventional WES.

[0056] While WES is currently a primary molecular diagnostic tool for monogenic diseases with a positive rate of 25%–40%, it still has significant problems: First, pathogenic abnormalities included in the OMIM and ClinVar databases are missed by the system and require subsequent verification; second, the detection sensitivity of specific regions such as exon margins and high GC regions is significantly low, making it easy to miss pathogenic variants and affecting diagnostic accuracy.

[0057] More importantly, there is currently a lack of publicly available and systematic assessment tools, making it impossible to quantify whether specific pathogenic variants will be missed by standard procedures before WES testing. This reactive diagnostic approach not only prolongs the patient's diagnostic cycle and delays disease intervention, but also increases the risk of missed or misdiagnosed cases, failing to meet the clinical demand for accurate and efficient diagnosis and hindering the full realization of the advantages of WES.

[0058] Therefore, this application provides a corresponding scheme for predicting site detection limitations in whole-exome sequencing, which can accurately predict site detection limitations in whole-exome sequencing.

[0059] See Figure 1 As shown in the embodiments of this application, a method for predicting site detection limitations in whole-exome sequencing is disclosed, including:

[0060] Step S11: Obtain the coordinate information of the target variant site and the reference genome from the Clinvar database and the genome database, respectively.

[0061] The Clinvar dataset contains a large number of variant sites. These variant sites are rated and screened as pathogenic variant sites and similarly pathogenic variant sites. These two types of variants are the focus of clinical gene testing and analysis. Therefore, this embodiment identifies these two types of variant sites as target variant sites. Specifically, the vcf variant files and variant rating information of the Clinvar dataset are obtained to obtain the coordinate information of the target variant sites.

[0062] To obtain a reference genome, the target genome version must first be determined before downloading it from a genome database. The reference genome is the basic framework for genome-related analysis. Only by using a suitable reference genome can the coordinates of the target variant site be accurately located, the accuracy of sequence feature analysis be ensured, and the comparability of analysis results between different experiments and different software be guaranteed. This avoids analysis errors caused by inconsistencies in genome versions or sources, and provides a reliable sequence basis for subsequent pathogenic variant detection and risk assessment.

[0063] Step S12: Extract the target sequence from the reference genome based on the coordinate information, and obtain the multi-dimensional sequence features of the target sequence.

[0064] In this embodiment, the extraction of the target sequence from the reference genome based on the coordinate information includes: extracting a first predetermined number of base pairs, a second predetermined number of base pairs, a third predetermined number of base pairs, and a fourth predetermined number of base pairs upstream and downstream of the target variant site from the reference genome based on the coordinate information, to obtain a first target sequence, a second target sequence, a third target sequence, and a fourth target sequence with different sequence lengths. Based on the coordinate information of the target variant site, the upstream and downstream sequences of the target variant site are extracted from the reference genome to generate the first target sequence, the second target sequence, the third target sequence, and the fourth target sequence; for example, extracting 50 bp (base pairs) upstream and downstream of the target variant site generates a 101 bp first target sequence, extracting 20 bp upstream and downstream of the target variant site generates a 51 bp second target sequence, extracting 50 bp downstream of the target variant site generates a 101 bp third target sequence, and extracting 50 bp upstream and downstream of the target variant site generates a 101 bp fourth target sequence.

[0065] In this embodiment, obtaining the multi-dimensional sequence features of the target sequence includes: determining the GC content of the first target sequence, and determining the GC content level of the first target sequence based on the GC content; determining the base frequency entropy of the second target sequence, and determining the base complexity level of the second target sequence based on the base frequency entropy; determining the base repeat type and repeat number of the third target sequence, and determining the base repeat level of the third target sequence based on the base repeat type and repeat number; performing a whole-genome sequence scan and alignment on the fourth target sequence to determine similar homologous regions in the fourth target sequence, and determining the homology level of the fourth target sequence based on the similar homologous regions.

[0066] For example Figure 2 The diagram illustrates a specific multi-dimensional sequence feature acquisition method. It determines the GC content of the first target sequence. The GC content is the proportion of the sum of the number of G bases (guanine) and C bases (cytosine) in the first target sequence to its length (Len). The calculation formula is as follows: Where G represents the total number of G bases (guanine) in the sequence, and C represents the total number of C bases (cytosine) in the sequence; the GC content level of the first target sequence is determined based on the GC content. Specifically, the GC content level is divided as follows: 1) Abnormal GC content level: GC content not less than 70% or GC content not greater than 25%; 2) Deviation GC content level: GC content is between [60%, 70%), or GC content is between (25%, 35%); 3) Normal GC content level: GC content is between (35%, 60%).

[0067] The base frequency entropy (H) of the second target sequence is determined. Base frequency entropy is the most commonly used diversity measure in biological sequences (DNA, RNA, protein) in information theory. It can assess the complexity of the sequence's base composition and reflects the degree of balance in the distribution of the four bases (A, T, C, G) in the second target sequence. It is quantified statistically by calculating the base frequency entropy, with a value distribution range of [0, 2]. A base frequency entropy of 2 indicates the most balanced distribution of the four bases and the highest complexity, while a base frequency entropy of 0 indicates the presence of only one base, the most unbalanced base distribution, and the lowest complexity. The calculation formula is as follows: Where i represents the four bases A, T, C, and G, and p i This indicates the frequency of base i in the sequence (by convention). The base complexity level of the second target sequence is determined based on the base frequency entropy. Specifically, the base complexity level is determined as follows: 1) Low complexity level: base frequency entropy H is not greater than 0.5; 2) Medium complexity level: base frequency entropy H is between (0.5, 1.5); 3) High complexity level: base frequency entropy H is greater than 1.5.

[0068] The type and number of base repetitions of the third target sequence are determined. Base repetitions exist in the human reference genome, including single-base repetitions and multiple-base repetitions. For example, the CFTR gene region contains TTTTTTT, with a single T base forming a repeating unit, repeated 7 times. Another example is the TGTGTGTGTGTGTGTGTGTGTG region in the CFTR gene, with T and G bases forming repeating units, repeated 11 times. Further, the third target sequence is determined based on the type and number of base repetitions. The base repetition levels are specifically classified as follows: 1) High repetition level of single base: the repetition type is single base repetition, and the number of repetition units is not less than 12; 2) Medium repetition level of single base: the repetition type is single base repetition, and the number of repetition units is between [6, 12); 3) High repetition level of multiple bases: the repetition type is multiple base repetition, and the number of repetition units is not less than 15; 4) Medium repetition level of multiple bases: the repetition type is multiple base repetition, and the number of repetition units is between [6, 15).

[0069] A whole-genome sequence scan and alignment were performed on the fourth target sequence to determine similar homologous regions based on a similarity score. Specifically, after aligning two base sequences, a similarity score was calculated between them using the formula: similarity = 100. (matches+repMatches) / (matches+repMatches+misMatches), where matches represents the number of identical base pairs, repMatches represents the number of base pairs marked as repeat but still counted as "positive", and misMatches represents the number of base mismatches. Further, the homology level of the fourth target sequence is determined based on similar homologous regions. The specific classification of homology levels is as follows: 1) High homology level: There are genomic regions that can match similarity, i.e., similar homologous regions, and the length of the similar homologous regions is not less than 95 bp, and the sequence similarity in the similar homologous regions is not less than 98%; 2) Moderate homology level: The length of the extracted similar homologous regions is not less than 95 bp, and the sequence similarity in the similar homologous regions is not less than 95%; 3) Low homology level: If the above conditions are not met, it is classified as low homology.

[0070] The multi-dimensional sequence features of the target sequence include GC content and GC content level, base frequency entropy and base complexity level, base repetition type and repetition count and base repetition level, similar homologous regions and homology level. The classification of multi-dimensional sequence features is shown in Table 1.

[0071] Table 1. Multidimensional Sequence Feature Classification Table

[0072]

[0073] Step S13: Determine the distance value between each target mutation site and the capture probe region determined based on the whole exon capture probe file, and calculate the coverage status of each target mutation site according to the distance value.

[0074] In this embodiment, the step of calculating the coverage status of each target variant site based on the distance value includes: if the distance value between the current target variant site and the capture probe region is greater than a first preset distance threshold, then the coverage status of the current target variant site is determined to be a preset uncaptured category; if the distance value between the current target variant site and the capture probe region is greater than a second preset distance threshold but not greater than the first preset distance threshold, then the coverage status of the current target variant site is determined to be a preset incompletely captured category; if the distance value between the current target variant site and the capture probe region is not greater than the second preset distance threshold, then the coverage status of the current target variant site is determined to be a preset fully captured category.

[0075] Based on the Browser Extensible Data (BED) file, the capture probe regions are determined, and the distance values ​​between each target variant site and the capture probe region are calculated to assess whether each target variant site can be effectively captured and covered by the probe. Figure 3 The diagram illustrates a specific capture coverage assessment method as follows: 1) If the distance between the current target variant site and the capture probe region is greater than a first preset distance threshold, the coverage status of the current target variant site is determined to be a preset uncaptured category. Specifically, the first preset distance threshold can be 100 bp; that is, if the distance between the current target variant site and the capture probe region is greater than 100 bp, the coverage status of the current target variant site is determined to be a preset uncaptured category. 2) If the distance between the current target variant site and the capture probe region is greater than a second preset distance threshold but not greater than the first preset distance threshold, the coverage status of the current target variant site is determined to be a preset incompletely captured category. Specifically, the first preset distance threshold can be 100 bp, and the second preset distance threshold is 20 bp; that is, if the distance between the current target variant site and the capture probe region is greater than 20 bp but not greater than 100 bp, the coverage status of the current target variant site is determined to be a preset incompletely captured category. 3) If the distance between the current target mutation site and the capture probe region is not greater than the second preset distance threshold, the coverage of the current target mutation site is determined to be the preset complete capture category; wherein, the second preset distance threshold is 20bp, that is, if the distance between the current target mutation site and the capture probe region is not greater than 20bp, the coverage of the current target mutation site is determined to be the preset complete capture category.

[0076] Step S14: Based on the coordinate information, the original sequence extracted from the reference genome is mutated to obtain a mutated sequence. The mutated sequence is then simulated to obtain simulated sequences. Each simulated sequence is then compared with the reference genome to obtain the alignment quality value and the proportion of alignment position consistency of each simulated sequence.

[0077] In this embodiment, the step of simulating sampling the mutant sequence to obtain each simulated sequence includes: simulating sampling the mutant sequence according to a fifth preset number of base pairs and a preset number of sampling times to obtain a simulated Fastq file containing each simulated sequence; wherein, the simulated Fastq file further includes original position information for marking each base pair in the simulated sequence corresponding to the original sequence.

[0078] For example Figure 4The diagram illustrates a specific alignment process. Based on coordinate information, a mutated sequence is extracted from a reference genome and mutated to obtain a mutated sequence. Specifically, 90 bp upstream and downstream of the target mutation site is extracted from the reference genome according to its coordinates, generating a 181 bp original sequence. The bases at the target mutation site in the original sequence are then replaced with the mutated bases, resulting in the mutated sequence. Simulated sampling is performed on the 181 bp original sequence, with a preset sampling number of 100 times. The extracted sequence length is a fifth preset number of base pairs, specifically 100 bp, resulting in a simulated Fastq file containing each simulated sequence. This simulated Fastq file also includes original position information in the original sequence used to label each base pair in the simulated sequence for subsequent alignment consistency analysis.

[0079] In this embodiment, the step of aligning each of the simulated sequences with the reference genome to obtain the alignment quality value and alignment position consistency ratio of each simulated sequence includes: using a genome alignment tool to align each of the simulated sequences with the reference genome to generate a simulated alignment BAM file containing the alignment quality value and actual alignment position information of each simulated sequence; determining alignment-consistent sequences from the simulated sequences based on the alignment results of the actual alignment position information and the original position information; and determining the alignment position consistency ratio based on the number of alignment-consistent sequences and the number of simulated sequences.

[0080] Genome alignment tools, such as coordinate genome alignment software for the target variant, are used to align each simulated sequence with a reference genome. This generates a simulated alignment BAM file (Binary Alignment Map) containing the Mapping Quality (MAPQ) value for each simulated sequence and the actual alignment position information. The MAPQ value measures the probability that a sequencing read is correctly aligned to its current position on the reference genome, expressed using the Phred scale. The calculation formula is: MAPQ = -10 * log10(P{position error}), typically ranging from 0 to 60. Here, P{position error} represents the probability that the sequencing read is incorrectly aligned to its current position on the reference genome. When MAPQ = 30, it means P{position error} = 0.1%; when MAPQ = 10, it means P{position error} = 10%; and when MAPQ = 0, it means the genome alignment tool is completely uncertain, or it may even be a random hit. Based on the alignment results between the actual and original alignment positions, aligned sequences are identified from the simulated sequences. These simulated sequences are then aligned to the reference genome, and the similarity between the actual and original alignment positions is statistically analyzed. Simulated sequences that match the original alignment positions are defined as aligned sequences. The alignment position consistency percentage is determined based on the number of aligned sequences and the number of simulated sequences, i.e., alignment position consistency percentage = number of aligned sequence entries / total number of simulated sequencing entries. 100%.

[0081] Step S15: Based on the multi-dimensional sequence features, the coverage status, the alignment quality value, and the alignment position consistency ratio, predict the detection limitations of the target variant site.

[0082] In this embodiment, the prediction of the detection limitations of the target variant site based on the multi-dimensional sequence features, the coverage status, the alignment quality value, and the alignment position consistency ratio includes: determining the target correct alignment level of the target variant site according to the relationship between the alignment quality value and each preset quality value threshold; determining the target reliability level of the target variant site according to the relationship between the alignment position consistency ratio and each preset ratio threshold; wherein, the target reliability level characterizes the reliability of the alignment result; and predicting the detection limitations of the target variant site based on the multi-dimensional sequence features, the coverage status, the target correct alignment level, and the target reliability level.

[0083] After obtaining the alignment quality value, it is also necessary to determine the target correct alignment level of the target variant site. The target correct alignment level of the target variant site is determined based on the relationship between the alignment quality value and each preset quality value threshold. Specifically, it can be: 1) Invalid alignment level: the proportion of aligned sequences with a MAPQ value not greater than 0 is not less than 80%; 2) Low-quality alignment level: the proportion of aligned sequences with a MAPQ value not greater than 20 is not less than 50%; 3) High-quality alignment level: if neither of the above two conditions is met, it is considered a high-quality alignment. After obtaining the alignment position consistency ratio, it is also necessary to determine the target reliability level of the target variant site. A higher target reliability level indicates higher reliability of the alignment result. The target reliability level of the target variant site is determined based on the relationship between the alignment position consistency ratio and each preset ratio threshold. The levels, from highest to lowest, are: 1) Inconsistent level: the proportion of aligned sequences is not greater than 30%; 2) Partially consistent level: the proportion of aligned sequences is between (30%) and (90%); 3) Completely consistent level: the proportion of aligned sequences is not less than 90%. The classification methods for target accuracy matching level and target reliability level are shown in Table 2:

[0084] Table 2. Alignment Feature Classification Table

[0085]

[0086] For example Figure 5 The diagram illustrates a specific assessment of detection limitations. Based on various features and level information, and combined with a variant detection risk genotyping table, a whole-exome detection risk assessment is performed on each target variant, i.e., a detection limitation assessment. The variant detection risk genotyping table is shown in Table 3.

[0087] Table 3. Risk Typing Table for Variant Detection

[0088]

[0089] The risk levels of the tests are classified into the following three categories:

[0090] 1) High-risk level: High-risk level indicates that the target variant cannot be detected by whole-exome sequencing and requires auxiliary detection methods or the development of targeted algorithms for effective detection; 2) Medium-risk level: Medium-risk level indicates that the target variant cannot be reliably detected by whole-exome sequencing and requires the development of targeted detection algorithms based on the different situations of different variants for stable detection; 3) Low-risk level: Low-risk level indicates that it can be reliably detected through conventional variant detection procedures.

[0091] Furthermore, by combining case studies of risk assessment of target variant sites, the risk of detecting pathogenic sites in whole-exome sequencing was evaluated. These cases covered different genes and related diseases, and were analyzed based on feature statistics (GC content, base frequency entropy H, single / multiple base repeat counts, homology, alignment quality value, alignment consistency, probe distance IDTV1) and feature classification dimensions. Finally, the risk levels of each variant site were classified as moderate risk, high risk, etc., as shown in Table 4.

[0092] Table 4 Examples of Risk Assessment for Target Variants

[0093]

[0094]

[0095] The beneficial effects of this application are as follows: This application obtains the coordinate information of the target variant site and the reference genome from the Clinvar database and the genome database, respectively; extracts the target sequence from the reference genome based on the coordinate information, and obtains the multi-dimensional sequence features of the target sequence; determines the distance value between each target variant site and the capture probe region determined based on the whole-exome capture probe file, and statistically analyzes the coverage of each target variant site based on the distance value; mutates the original sequence extracted from the reference genome based on the coordinate information to obtain the mutated sequence, performs simulated sampling on the mutated sequence to obtain each simulated sequence, and compares the position of each simulated sequence with the reference genome to obtain the alignment quality value and the alignment position consistency ratio of each simulated sequence; and predicts the detection limitations of the target variant site based on the multi-dimensional sequence features, the coverage, the alignment quality value, and the alignment position consistency ratio. Therefore, this application obtains the coordinates of target variant sites and reference genomes from the Clinvar database and genome database, extracts target sequences and acquires multi-dimensional sequence features, determines the coverage of target variant sites by combining whole-exome capture probe files, obtains alignment quality values ​​and alignment position consistency ratios by simulating mutant sequence sampling and comparing with the reference genome, and finally predicts the detection limitations of target variant sites based on these indicators. Before actual whole-exome sequencing, by integrating multi-dimensional information such as sequence features, coverage, alignment quality, and consistency ratios, it can identify target variant sites with different detection risks in advance, thereby planning auxiliary detection methods for variant sites with different risks in advance. This effectively reduces the risk of missed or false detections caused by the characteristics of the site itself or technical limitations in actual sequencing, improves the accuracy and stability of whole-exome sequencing for the detection of clinically pathogenic variant sites, and reduces the increased costs and extended cycles caused by post-analysis.

[0096] The following is based on Figure 6Taking a specific example of a site detection limitation prediction diagram, this application will be explained accordingly. 1) Target variant sites, i.e., pathogenic variant sites and suspected pathogenic variant sites, are obtained from the Clinvar database, and the coordinate information of the target variant sites is extracted. 2) Based on the coordinate information, the target sequence is extracted from the reference genome, and multi-dimensional sequence features of the target sequence are obtained. The multi-dimensional sequence features include features obtained from site region base composition analysis and features obtained from homology analysis. Features obtained from site region base composition analysis include GC content and GC content level, base frequency entropy and base complexity level, base repeat type and repeat count and base repeat level. Features obtained from homology analysis are similar homologous regions and homology level. 3) The distance value between each target variant site and the capture probe region determined based on the whole-exon capture probe file is determined, and the coverage of each target variant site is statistically analyzed based on the distance value. 4) Based on coordinate information, the original sequence extracted from the reference genome is mutated to obtain mutated sequences. The mutated sequences are then sampled to obtain simulated Fastq files containing each simulated sequence. These simulated Fastq files also include the original position information of each base pair in the simulated sequence corresponding to its corresponding position in the original sequence. Each simulated sequence is aligned with the reference genome to obtain its alignment quality value and alignment position consistency percentage. 5) The target correct alignment level of the target variant site is determined based on the relationship between the alignment quality value and each preset quality value threshold; the target reliability level of the target variant site is determined based on the relationship between the alignment position consistency percentage and each preset percentage threshold. The target reliability level characterizes the reliability of the alignment results. 6) The detection limitations of the target variant site are predicted based on multi-dimensional sequence features, coverage, target correct alignment level, and target reliability level.

[0097] See Figure 7 As shown in the embodiment of this application, a site detection limitation prediction device in whole exome sequencing is disclosed, comprising:

[0098] The signal acquisition module 11 is used to acquire the coordinate information of the target variant site and the reference genome from the Clinvar database and the genome database, respectively.

[0099] The feature acquisition module 12 is used to extract the target sequence from the reference genome based on the coordinate information, and to acquire the multi-dimensional sequence features of the target sequence;

[0100] The coverage status acquisition module 13 is used to determine the distance value between each of the target mutation sites and the capture probe region determined based on the whole exon capture probe file, and to calculate the coverage status of each of the target mutation sites based on the distance value.

[0101] The position alignment module 14 is used to mutate the original sequence extracted from the reference genome based on the coordinate information to obtain a mutated sequence, perform simulated sampling on the mutated sequence to obtain each simulated sequence, and perform position alignment between each simulated sequence and the reference genome to obtain the alignment quality value and the alignment position consistency ratio of each simulated sequence.

[0102] The limitation prediction module 15 is used to predict the detection limitation of the target variant site based on the multi-dimensional sequence features, the coverage status, the alignment quality value, and the alignment position consistency ratio.

[0103] Furthermore, embodiments of this application also provide an electronic device. Figure 8 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0104] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the site detection limitation prediction method for whole-exome sequencing performed by the electronic device disclosed in any of the foregoing embodiments.

[0105] In this embodiment, the power supply 23 is used to provide operating voltage for various hardware devices on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0106] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0107] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored on it include operating system 221, computer program 222 and data 223, etc., and the storage method can be temporary storage or permanent storage.

[0108] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system can be Windows, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the site detection limitation prediction method for whole-exome sequencing executed by the electronic device as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.

[0109] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned method for predicting site detection limitations in whole-exome sequencing. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0110] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0111] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly in hardware, software modules executed by a processor, or a combination of both. The software module may be located in random access memory (RAM), memory, read-only memory (ROM), electrically programmable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), register, hard disk, removable disk, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the art.

[0112] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0113] The foregoing has provided a detailed description of the method, apparatus, equipment, and medium for predicting site detection limitations in whole-exome sequencing provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only intended to help understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A method for predicting site detection limitations in whole-exome sequencing, characterized in that, include: The coordinates of the target variant site and the reference genome were obtained from the Clinvar database and the genome database, respectively. Based on the coordinate information, the target sequence is extracted from the reference genome, and the multi-dimensional sequence features of the target sequence are obtained; Determine the distance between each target mutation site and the capture probe region determined based on the whole-exon capture probe file, and statistically analyze the coverage of each target mutation site based on the distance values; Based on the coordinate information, the original sequence extracted from the reference genome is mutated to obtain a mutated sequence. The mutated sequence is then sampled to obtain simulated sequences. Each simulated sequence is then aligned with the reference genome to obtain the alignment quality value and the proportion of alignment position consistency of each simulated sequence. The detection limitations of predicting the target variant site based on the multi-dimensional sequence features, the coverage, the alignment quality value, and the alignment position consistency ratio.

2. The method for predicting site detection limitations in whole-exome sequencing according to claim 1, characterized in that, The extraction of the target sequence from the reference genome based on the coordinate information includes: Based on the coordinate information, a first preset number of base pairs, a second preset number of base pairs, a third preset number of base pairs, and a fourth preset number of base pairs are extracted from the upstream and downstream of the target mutation site from the reference genome to obtain a first target sequence, a second target sequence, a third target sequence, and a fourth target sequence with different sequence lengths.

3. The method for predicting site detection limitations in whole-exome sequencing according to claim 2, characterized in that, The process of obtaining the multi-dimensional sequence features of the target sequence includes: Determine the GC content of the first target sequence, and determine the GC content level of the first target sequence based on the GC content; Determine the base frequency entropy of the second target sequence, and determine the base complexity level of the second target sequence based on the base frequency entropy; Determine the base repetition type and number of repetitions of the third target sequence, and determine the base repetition level of the third target sequence based on the base repetition type and the number of repetitions; A whole-genome sequence scan and alignment is performed on the fourth target sequence to identify similar homologous regions in the fourth target sequence, and the homology level of the fourth target sequence is determined based on the similar homologous regions.

4. The method for predicting site detection limitations in whole-exome sequencing according to claim 1, characterized in that, The step of calculating the coverage status of each target mutation site based on the distance value includes: If the distance between the target mutation site and the capture probe region is greater than a first preset distance threshold, then the coverage status of the target mutation site is determined to be a preset uncaptured category. If the distance between the target mutation site and the capture probe region is greater than a second preset distance threshold but not greater than a first preset distance threshold, then the coverage status of the target mutation site is determined to be a preset incomplete capture category. If the distance between the target mutation site and the capture probe region is not greater than a second preset distance threshold, then the coverage status of the target mutation site is determined to be a preset complete capture category.

5. The method for predicting site detection limitations in whole-exome sequencing according to any one of claims 1 to 4, characterized in that, The simulated sampling of the mutated sequence to obtain simulated sequences includes: The mutant sequence is simulated and sampled according to a fifth preset number of base pairs and a preset number of sampling times to obtain a simulated Fastq file containing each simulated sequence; wherein, the simulated Fastq file also includes original position information for marking each base pair in the simulated sequence in the original sequence.

6. The method for predicting site detection limitations in whole-exome sequencing according to claim 5, characterized in that, The step of aligning each of the simulated sequences with the reference genome to obtain the alignment quality value and the percentage of alignment position consistency for each simulated sequence includes: The simulated sequences are aligned with the reference genome using a genome alignment tool to generate a simulated alignment BAM file containing the alignment quality value and actual alignment position information of each simulated sequence. Based on the comparison results between the actual comparison location information and the original location information, a matching sequence is determined from the simulated sequence; The alignment position consistency ratio is determined based on the number of aligned sequences and the number of simulated sequences.

7. The method for predicting site detection limitations in whole-exome sequencing according to claim 6, characterized in that, The limitations of predicting the target variant site based on the multi-dimensional sequence features, the coverage, the alignment quality value, and the alignment position consistency ratio include: The target correct alignment level of the target variant site is determined based on the relationship between the alignment quality value and each preset quality value threshold. The target reliability level of the target variant site is determined based on the relationship between the consistency ratio of the comparison locations and each preset ratio threshold; wherein, the target reliability level characterizes the reliability of the comparison results; The detection limitations of the target variant sites are predicted based on the multi-dimensional sequence features, the coverage status, the target correct alignment level, and the target reliability level.

8. A device for predicting site detection limitations in whole-exome sequencing, characterized in that, include: The signal acquisition module is used to obtain the coordinate information of the target variant site and the reference genome from the Clinvar database and the genome database, respectively. The feature acquisition module is used to extract the target sequence from the reference genome based on the coordinate information, and to acquire the multi-dimensional sequence features of the target sequence; The coverage status acquisition module is used to determine the distance value between each of the target mutation sites and the capture probe region determined based on the whole exon capture probe file, and to calculate the coverage status of each of the target mutation sites based on the distance value. The position alignment module is used to mutate the original sequence extracted from the reference genome based on the coordinate information to obtain a mutated sequence, to perform simulated sampling on the mutated sequence to obtain each simulated sequence, and to perform position alignment of each simulated sequence with the reference genome to obtain the alignment quality value and the proportion of alignment position consistency of each simulated sequence. The limitation prediction module is used to predict the detection limitations of the target variant site based on the multi-dimensional sequence features, the coverage status, the alignment quality value, and the alignment position consistency ratio.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the method for predicting site detection limitations in whole-exome sequencing as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when executed by a processor, the computer program implements the steps of the method for predicting site detection limitations in whole exome sequencing as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Tumor mutation load detection method and device and storage medium

    CN109033749A

  • A data set and method for high throughput sequencing somatic mutation detection performance evaluation

    CN111916152A