Genome deletion type structure variation accuracy evaluation method and device based on site depth

By calculating genomic locus depth and depth feature analysis, combined with multidimensional sequencing evidence, the system automatically assesses deletion-type structural variations, solving the problem of high false positive rates in existing technologies and achieving efficient and accurate detection of structural variations.

CN121034401APending Publication Date: 2025-11-28BEIJING NOVOGENE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511581014.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies for detecting deletion-type structural variations have a high false positive rate, resulting in low reliability of test results, low efficiency of manual verification, and poor reproducibility of results.

Method used

By calculating the alignment depth of each point in the genome, candidate gene fragments are identified using structural variation detection tools, and in-depth feature analysis is performed. The accuracy of deletion-type structural variations is automatically evaluated by combining site depth difference analysis, piecewise linear regression, and support of soft-splitting reads.

Benefits of technology

It significantly improves the accuracy and efficiency of detecting deletion-type structural variations, reduces the false positive rate, and enhances the reliability and consistency of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034401A_ABST
    Figure CN121034401A_ABST
Patent Text Reader

Abstract

The invention discloses a genome deletion type structure variation accuracy evaluation method and device based on site depth. The method comprises the steps that the comparison depth of each site on a genome is calculated through comparison data processing software, comparison depth data are obtained, and the genome is an object needing deletion type structure variation recognition; carrying out genome deletion variation recognition on the genome by using a structural variation detection tool to obtain candidate gene segments with genome deletion variation on the genome; performing depth feature analysis on the candidate gene segments according to the comparison depth data to obtain a depth feature analysis result; target gene segments in the candidate gene segments are determined according to the depth feature analysis result, and the target gene segments are gene segments with deletion type structure variation in the candidate gene segments. According to the invention, the technical problem of low reliability of a detection result caused by high false positive rate of a deletion type structure variation detection mode in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and in particular, to a method and device for evaluating the accuracy of genomic deletion-type structural variations based on site depth. BACKGROUND

[0002] Genomic structural variations (SV) are an important part of genomic research. Deletion variations, as one of the structural variations, have attracted widespread attention due to their key role in evolution, disease mechanisms, and individual differences. With the development of sequencing technology, especially the popularization of third-generation sequencing technology, researchers can obtain higher precision and longer read genomic sequences, which greatly improves the ability to detect large-scale genomic variations.

[0003] However, although SV detection tools are constantly emerging (such as cuteSV, svim, pbsv, Sniffles, etc.), and to some extent, they have improved the efficiency of SV identification, these tools generally face the problem of high false positive rate. On the one hand, due to the complexity of structural variations themselves, such as breakpoint uncertainty, the influence of repetitive sequences, etc., it is difficult for automated detection tools to perfectly distinguish between true variations and false signals. On the other hand, due to differences in different sequencing technologies and sample processing steps, even the most advanced algorithms can produce a certain proportion of false positive detection results.

[0004] In order to make up for the shortcomings of automated tools, IGV (Integrative Genomics Viewer) and other visualization software are widely used for manual verification in current research and practice. Although this method can effectively reduce false positives and ensure the accuracy of detection results, it also has obvious limitations:

[0005] 1) Low efficiency: The process of manual verification requires professional bioinformatics analysts to review the alignment of SV regions one by one, including read distribution, depth mutation, breakpoint characteristics, etc., which is extremely time-consuming when dealing with large-scale data sets.

[0006] 2) Subjective bias: Differences in personal experience, preferences, and judgment standards may lead to differences in interpretation of the same SV, affecting the consistency and repeatability of the results.

[0007] Therefore, the existing technology has the problems of high false positive rate of deletion-type structural variation detection, low efficiency of manual verification, and poor reliability of results. An automated scoring system based on site depth and multi-evidence fusion is provided to significantly improve the efficiency and accuracy of deletion-type SV screening and promote the standardization and automation process of genomic structural variation research.

[0008] At present, no effective solution has been proposed for the above problems. SUMMARY

[0009] The embodiment of the present application provides a site depth-based genomic deletion type structural variation accuracy evaluation method and device, so as to at least solve the technical problem of high false positive rate of deletion type structural variation detection mode in related technologies, resulting in low reliability of detection results.

[0010] According to an aspect of the embodiment of the present application, a site depth-based genomic deletion type structural variation accuracy evaluation method is provided, comprising: calculating the alignment depth of each site on the genome by using alignment data processing software to obtain alignment depth data, wherein the genome is an object to be identified for deletion type structural variation; identifying the genomic deletion variation of the genome by using a structural variation detection tool to obtain a candidate gene fragment with the genomic deletion variation on the genome; performing depth feature analysis on the candidate gene fragment according to the alignment depth data to obtain a depth feature analysis result; determining a target gene fragment in the candidate gene fragment according to the depth feature analysis result, wherein the target gene fragment is a gene fragment with deletion type structural variation in the candidate gene fragment.

[0011] Optionally, the alignment data processing software is Samtools, and the alignment depth of each site on the genome is calculated by using the alignment data processing software to obtain alignment depth data, comprising: calculating the alignment depth of each site on the genome by using the depth parameter of the Samtools to obtain the alignment depth data.

[0012] Optionally, the genomic deletion variation of the genome is identified by using a structural variation detection tool to obtain a candidate gene fragment with the genomic deletion variation on the genome, comprising: finding a read breakpoint of the genome by using the structural variation detection tool to obtain a read breakpoint finding result; searching for a mismatched paired read in the genome by using the structural variation detection tool to obtain a mismatched read result; and obtaining the candidate gene fragment according to the read breakpoint finding result and the mismatched read result.

[0013] Optionally, the candidate gene fragment is subjected to a depth feature analysis according to the aligned depth data to obtain a depth feature analysis result, including: determining a start position and an end position of the candidate gene fragment; determining first depth data of a first predetermined position and second depth data of a second predetermined position according to the aligned depth data, wherein the first predetermined position is a position corresponding to a plurality of sites in an upstream direction of the start position, and the second predetermined position is a position corresponding to a plurality of sites in a downstream direction of the end position; performing depth comparison on each of the first depth data to obtain a first depth alignment result, and performing depth comparison on each of the second depth data to obtain a second depth alignment result; and obtaining the depth feature analysis result according to the first depth alignment result and the second depth alignment result.

[0014] Optionally, the depth comparison on each of the first depth data to obtain a first depth alignment result, and the depth comparison on each of the second depth data to obtain a second depth alignment result, includes: performing rank sum test on each of the first depth data to obtain a first rank sum test value, and comparing the first rank sum test value with a first predetermined value and a second predetermined value to obtain the first depth alignment result, wherein the first predetermined value is greater than the second predetermined value; performing rank sum test on each of the second depth data to obtain a second rank sum test value, and comparing the second rank sum test value with a third predetermined value and a fourth predetermined value to obtain the second depth alignment result, wherein the third predetermined value is greater than the fourth predetermined value.

[0015] Optionally, the depth feature analysis result is obtained according to the first depth alignment result and the second depth alignment result, including: determining that there is a significant depth change at the start position when the first depth alignment result indicates that the first rank sum test value is greater than the first predetermined value or the first rank sum test value is less than the second predetermined value; and determining that there is the significant depth change at the end position when the second depth alignment result indicates that the second rank sum test value is greater than the third predetermined value or the second rank sum test value is less than the fourth predetermined value.

[0016] Optionally, the candidate gene fragment is subjected to depth feature analysis according to the aligned depth data, to obtain a depth feature analysis result, including: determining a start position and an end position of the candidate gene fragment; determining first depth data of a first predetermined position and second depth data of a second predetermined position according to the aligned depth data, wherein the first predetermined position is a position corresponding to a plurality of sites in an upstream direction of the start position, and the second predetermined position is a position corresponding to a plurality of sites in a downstream direction of the end position; performing linear regression analysis on the first depth data and the second depth data, to obtain a regression slope of the start position and the end position within a predetermined range; and obtaining the depth feature analysis result according to the regression slope.

[0017] Optionally, the candidate gene fragment is subjected to depth feature analysis according to the aligned depth data, to obtain a depth feature analysis result, including: determining a start position and an end position of the candidate gene fragment; determining third depth data of each site within a first site interval corresponding to the start position and the end position according to the aligned depth data; determining a coefficient of variation of the site interval according to the third depth data; and determining the depth feature analysis result according to the coefficient of variation.

[0018] Optionally, the candidate gene fragment is subjected to depth feature analysis according to the aligned depth data, to obtain a depth feature analysis result, including: determining a start position and an end position of the candidate gene fragment; determining whether there is a soft clipping read within a second site interval corresponding to the start position and the end position, to obtain a detection result; and obtaining the depth feature analysis result according to the detection result.

[0019] Optionally, a target gene fragment in the candidate gene fragment is determined according to the depth feature analysis result, including: determining a score corresponding to each dimension index in the depth feature analysis result and a weight of the each dimension index; performing weighted summation on the depth feature analysis result according to the score and the weight, to obtain a total score of the depth feature analysis result; and selecting a gene fragment with a total score not less than a predetermined score from the candidate gene fragment as the target gene fragment.

[0020] According to another aspect of the embodiments of the present application, there is also provided a device for evaluating accuracy of a site depth-based genomic deletion type structural variation, comprising: a computing module configured to calculate alignment depth of each site on a genome by using a data processing software to obtain alignment depth data, wherein the genome is an object to which deletion type structural variation identification is to be performed; an identifying module configured to identify genomic deletion variation of the genome by using a structural variation detection tool to obtain a candidate gene fragment in which the genomic deletion variation exists on the genome; an analyzing module configured to perform depth feature analysis on the candidate gene fragment according to the alignment depth data to obtain a depth feature analysis result; and a determining module configured to determine a target gene fragment in the candidate gene fragment according to the depth feature analysis result, wherein the target gene fragment is a gene fragment of the deletion type structural variation in the candidate gene fragment.

[0021] Optionally, the computing module comprises a computing unit configured to calculate the alignment depth of each site on the genome by using a depth parameter of the Samtools to obtain the alignment depth data.

[0022] Optionally, the identifying module comprises a searching unit configured to search for a read breakpoint of the genome by using the structural variation detection tool to obtain a read breakpoint searching result; a searching unit configured to search for a pair of mismatched reads in the genome by using the structural variation detection tool to obtain a mismatched read result; and a first processing unit configured to obtain the candidate gene fragment according to the read breakpoint searching result and the mismatched read result.

[0023] Optionally, the analyzing module comprises a first determining unit configured to determine a start position and an end position of the candidate gene fragment; a second determining unit configured to determine first depth data of a first predetermined position and second depth data of a second predetermined position according to the alignment depth data, wherein the first predetermined position is a position corresponding to a plurality of sites in an upstream direction of the start position, and the second predetermined position is a position corresponding to a plurality of sites in a downstream direction of the end position; a comparing unit configured to compare the first depth data to obtain a first depth comparison result, and compare the second depth data to obtain a second depth comparison result; and a second processing unit configured to obtain the depth feature analysis result according to the first depth comparison result and the second depth comparison result.

[0024] Optionally, the comparison unit comprises: a first test subunit, configured to perform a rank-sum test on each of the first depth data to obtain a first rank-sum test value, and compare the first rank-sum test value with a first predetermined value and a second predetermined value to obtain the first depth comparison result, wherein the first predetermined value is greater than the second predetermined value; and a second test subunit, configured to perform a rank-sum test on each of the second depth data to obtain a second rank-sum test value, and compare the second rank-sum test value with a third predetermined value and a fourth predetermined value to obtain the second depth comparison result, wherein the third predetermined value is greater than the fourth predetermined value.

[0025] Optionally, the second processing unit comprises: a first determination subunit, configured to determine that the start position has a significant depth change when the first depth comparison result indicates that the first rank-sum test value is greater than the first predetermined value or the first rank-sum test value is less than the second predetermined value; and a second determination subunit, configured to determine that the end position has the significant depth change when the second depth comparison result indicates that the second rank-sum test value is greater than the third predetermined value or the second rank-sum test value is less than the fourth predetermined value.

[0026] Optionally, the analysis module comprises: a third determination unit, configured to determine a start position and an end position of the candidate gene fragment; a fourth determination unit, configured to determine first depth data of a first predetermined position and second depth data of a second predetermined position according to the compared depth data, wherein the first predetermined position is a position corresponding to a plurality of sites in an upstream direction of the start position, and the second predetermined position is a position corresponding to a plurality of sites in a downstream direction of the end position; an analysis unit, configured to perform linear regression analysis on the first depth data and the second depth data to obtain a regression slope of the start position and the end position within a predetermined range; and a third processing unit, configured to obtain the depth feature analysis result according to the regression slope.

[0027] Optionally, the analysis module comprises: a fifth determination unit, configured to determine a start position and an end position of the candidate gene fragment; a sixth determination unit, configured to determine third depth data of each site in a first site interval corresponding to the start position and the end position according to the compared depth data; a seventh determination unit, configured to determine a coefficient of variation of the site interval according to the third depth data; and an eighth determination unit, configured to determine the depth feature analysis result according to the coefficient of variation.

[0028] Optionally, the analysis module comprises: a ninth determination unit configured to determine a start position and an end position of the candidate gene fragment; a fourth processing unit configured to detect whether there is a soft-clipped read in a second site interval corresponding to the start position and the end position, and obtain a detection result; and a fifth processing unit configured to obtain a depth feature analysis result according to the detection result.

[0029] Optionally, the determination module comprises: a tenth determination unit configured to determine a score corresponding to each dimension index in the depth feature analysis result and a weight of the each dimension index; a summation unit configured to perform weighted summation on the depth feature analysis result according to the score and the weight, and obtain a total score of the depth feature analysis result; and a selection unit configured to select a gene fragment with a total score not less than a predetermined score from the candidate gene fragments as the target gene fragment.

[0030] According to another aspect of the embodiments of the present application, a computer readable storage medium is also provided, which comprises a stored program, wherein the program performs the site depth-based genomic deletion-type structural variation accuracy evaluation method according to any one of the preceding embodiments.

[0031] According to another aspect of the embodiments of the present application, a processor is also provided, which is used to run a program, wherein the program performs the site depth-based genomic deletion-type structural variation accuracy evaluation method according to any one of the preceding embodiments when running.

[0032] According to another aspect of the embodiments of the present application, a computer program product is also provided, which comprises computer instructions, and the computer instructions perform the site depth-based genomic deletion-type structural variation accuracy evaluation method according to any one of the preceding embodiments when executed by a processor.

[0033] In the embodiment of the present application, the alignment depth of each site on the genome is calculated by comparing data processing software to obtain alignment depth data, wherein the genome is the object to be identified for deletion type structural variation; the structural variation detection tool is used to identify the genomic deletion variation of the genome to obtain the candidate gene fragment with genomic deletion variation on the genome; the depth feature analysis of the candidate gene fragment is performed according to the alignment depth data to obtain the depth feature analysis result; and the target gene fragment in the candidate gene fragment is determined according to the depth feature analysis result, wherein the target gene fragment is the gene fragment of the deletion type structural variation in the candidate gene fragment. By using the depth information of the third generation sequencing data and the multi-dimensional sequencing evidence fusion, including site depth difference analysis, segmented linear regression, SV region depth consistency test and support of soft sheared reads, the purpose of automatically evaluating the quality of deletion type structural variation is achieved, thereby realizing the technical effect of significantly improving the SV detection accuracy and efficiency, and further solving the technical problems of high false positive rate of deletion type structural variation detection mode in the related art, resulting in low reliability of the detection result. BRIEF DESCRIPTION OF DRAWINGS

[0034] The accompanying drawings, which are included to provide a further understanding of the present application and constitute a part of this application, illustrate certain illustrative embodiments of the present application and together with the description serve to explain the present application. In the drawings:

[0035] Figure 1 is a hardware structure block diagram of a mobile terminal based on a site depth-based genomic deletion type structural variation accuracy evaluation method according to an embodiment of the present application;

[0036] Figure 2 is a flowchart of a site depth-based genomic deletion type structural variation accuracy evaluation method according to an embodiment of the present application;

[0037] Figure 3 is a flowchart of an optional site depth-based genomic deletion type structural variation accuracy evaluation method according to an embodiment of the present application;

[0038] Figure 4 is a schematic diagram of a site depth-based genomic deletion type structural variation accuracy evaluation device according to an embodiment of the present application.

[0039] Among the above drawings, the following reference signs are included:

[0040] 102, processor; 104, memory; 106, transmission device; 108, input and output device. DETAILED DESCRIPTION

[0041] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0042] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0043] As described in the background section, existing methods for detecting deletion structural variations suffer from high false positive rates, leading to low reliability of test results. This invention provides a method and apparatus for assessing the accuracy of site-depth genomic deletion structural variations, a computer-readable storage medium, a processor, and a computer program product.

[0044] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0045] The methods and embodiments provided in this invention can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure diagram of a mobile terminal for a site-depth-based method for assessing the accuracy of genomic deletion structural variations, according to an embodiment of the present invention. Figure 1 As shown, a mobile terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the mobile terminal described above. For example, the mobile terminal may also include components that are more...Figure 1 more or less components than those shown in the figures, or configured differently from those shown in the figures. Figure 1

[0046] The memory 104 is used to store computer programs, such as software programs of application software and modules, for example, the computer program corresponding to the method for evaluating accuracy of deletion-type structural variations of a genome based on site depth according to the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the method described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, and the remote memory can be connected to the mobile terminal through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The transmission device 106 is used to receive or send data via a network. The specific examples of the network can include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet in a wireless manner.

[0047] According to the embodiments of the present application, a method embodiment of the method for evaluating accuracy of deletion-type structural variations of a genome based on site depth is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0048] Figure 2 is a flowchart of the method for evaluating accuracy of deletion-type structural variations of a genome based on site depth according to the embodiments of the present application, as shown in Figure 2 the method includes the following steps:

[0049] In step S202, the alignment depth of each site on the genome is calculated by using a data processing software to obtain alignment depth data, wherein the genome is an object to be identified for deletion-type structural variations.

[0050] ​In this embodiment, by calculating the sequencing depth of each site, the number of times each site is covered by sequencing reads in the whole genome range can be obtained, which is a direct basis for evaluating whether the genomic region is affected by deletion and other structural variations. By obtaining detailed sequencing depth data, necessary quantitative information can be provided for subsequent identification and verification of deletion-type structural variations. The significant feature of deletion-type structural variations is that the sequencing depth in the variation region is significantly reduced. By comparing the sequencing depths upstream and downstream, the low-depth region can be preliminarily locked as a possible deletion position, providing data support for subsequent depth difference significance test and segmented linear regression analysis.

[0051] In step S204, a structural variation detection tool is used to identify genomic deletion variations in the genome to obtain candidate gene fragments with genomic deletion variations on the genome.

[0052] In this embodiment, the structural variation detection tool can quickly screen out possible deletion variation intervals in the genome based on read alignment information. By using abnormal patterns in high-throughput sequencing data, candidate positions of deletion variations can be preliminarily locked, and target regions for subsequent depth analysis can be provided. By using the structural variation detection tool, massive sequencing data can be automatically processed, and thousands of candidate variations can be efficiently identified, avoiding the heavy work of checking each site. By using the structural variation detection tool to preliminarily identify the genome, candidate positions of deletion variations can be effectively screened and located, providing accurate targets and data basis for subsequent in-depth analysis based on site depth, thereby improving the efficiency and accuracy of genome deletion-type structural variation identification as a whole.

[0053] It should be noted that the above-mentioned structural variation detection tool refers to a bioinformatics software specially designed for identifying and analyzing genomic structural variations (Structural Variations, SVs) from high-throughput sequencing data, which can process a large amount of sequencing read data and identify various types of structural variations. The above-mentioned genomic deletion variation refers to the deletion of a certain region of DNA sequence in the individual or population genome compared with the reference genome, which can be small-scale, such as deletion of a few base pairs, or large-scale (deletion of not less than 50 bp). Large-scale DNA sequence deletion is usually referred to as deletion-type variation (Deletion SV) in structural variation (Structural Variation, SV). Genomic structural variation generally refers to large-scale sequence changes and position relationship changes in the genome.

[0054] Optionally, the structural variation detection tool can include, but is not limited to, CuteSV, SVIM, Sniffles, SVIM, etc.; and the types of structural variations can include, but are not limited to, long fragment sequence deletions (Deletions) with a length of 50 bp or more, insertions (Insertions), tandem repeats (Tandem Repeat), chromosomal inversions (Inversion), sequence translocations between chromosomes (Translocations), duplications (Duplications), copy number variations (Copy Number Variations, CNVs), and more complex chimeric variations, etc.

[0055] In step S206, depth feature analysis is performed on the candidate gene fragments according to the alignment depth data, and a depth feature analysis result is obtained.

[0056] In this embodiment, by carefully investigating the selected candidate variation region, the true characteristics of the suspected deletion variation on the genome can be revealed. The depth feature analysis can provide quantitative basis for the identification of deletion structural variations by in-depth analysis of the sequencing depth data of the candidate gene fragments. Deletion structural variations often exhibit regions with reduced sequencing depth. Through depth feature analysis, it can be directly verified whether the candidate variation region has a depth drop phenomenon, thereby enhancing the objectivity, accuracy and reliability of the variation detection, and providing tool support for structural variation analysis in large-scale genomic research.

[0057] In step S208, a target gene fragment in the candidate gene fragments is determined according to the depth feature analysis result, wherein the target gene fragment is a gene fragment of the deletion structural variation in the candidate gene fragments.

[0058] In this embodiment, by carefully analyzing the depth features of the candidate regions, the true deletion variations can be accurately distinguished from the depth changes caused by non-variation factors, thereby realizing accurate locking of the target gene fragments. The analysis based on the depth features can effectively filter out false positive variation signals caused by alignment errors, sequencing errors or other technical problems in the data processing process, reduce the false positive rate of subsequent analysis, and through determination of the target gene fragments, further analysis of the true variations can be focused on without the need for screening in a large number of candidate fragments, thereby saving time and resources and accelerating the research process.

[0059] From the above, in the embodiment of the application, the alignment depth of each site on the genome is calculated by comparing data processing software to obtain alignment depth data, wherein the genome is the object to be identified for deletion type structural variation; the structural variation detection tool is used to identify the genomic deletion variation of the genome to obtain the candidate gene fragment with the genomic deletion variation on the genome; the depth feature analysis of the candidate gene fragment is performed according to the alignment depth data to obtain the depth feature analysis result; the target gene fragment in the candidate gene fragment is determined according to the depth feature analysis result, wherein the target gene fragment is the gene fragment of the deletion type structural variation in the candidate gene fragment, and the depth information of the third generation sequencing data and the multi-dimensional sequencing evidence fusion are used, including site depth difference analysis, segmented linear regression, SV region depth consistency test and support degree of soft cutting reads, to achieve the purpose of automatically evaluating the quality of the deletion type structural variation, thereby realizing the technical effect of significantly improving the SV detection accuracy and efficiency.

[0060] Through the above technical solution provided by the embodiment of the application, the technical problem that the deletion type structural variation detection method in the related art has a high false positive rate and leads to low reliability of the detection result is solved.

[0061] According to the above embodiment of the application, the alignment data processing software is Samtools, the alignment depth of each site on the genome is calculated by the alignment data processing software to obtain alignment depth data, including: the alignment depth of each site on the genome is calculated by the depth parameter of Samtools to obtain the alignment depth data.

[0062] In this embodiment, the depth parameter of Samtools can accurately calculate the coverage of each site, providing basic data for subsequent depth feature analysis, and using Samtools for depth calculation can ensure the consistency and reliability of the data, which is beneficial to accurately evaluate the authenticity of the candidate variation.

[0063] It should be noted that Samtools is an open-source software widely used in the field of bioinformatics, mainly used for efficient analysis and management of high-throughput sequencing (HTS) data, capable of operating sequence alignment mapping (SAM) and binary sequence alignment mapping (BAM) format files, which are standard storage formats for genomic sequencing data. The sequencing depth of each position on the genome is calculated using the Samtools software depth parameter. Different versions of Samtools mainly include function enhancement, performance optimization, bug fixing, and support for new data formats and analysis methods. Version 1.17 can also be used. Version 1.17 (i.e., Samtools v1.17) contains a series of updates and optimizations for the software, enabling it to more effectively handle large sequencing data sets, especially in high-throughput sequencing (HTS) data analysis.

[0064] Optionally, the above-mentioned Samtools v1.17 provides a variety of tools and commands, which can include but are not limited to: alignment file operations, sequence depth calculation, variant calling, sequence alignment quality control, variant annotation and filtering, data compression and storage, and multi-thread processing, etc.

[0065] According to the above embodiment of the present application, the structural variation detection tool is used to identify the genomic deletion variation of the genome, and the candidate gene fragment with the genomic deletion variation on the genome is obtained, including: finding the read breakpoint of the genome by the structural variation detection tool to obtain the read breakpoint finding result; searching for the mismatched paired reads in the genome by the structural variation detection tool to obtain the mismatched read result; obtaining the candidate gene fragment according to the read breakpoint finding result and the mismatched read result.

[0066] In this embodiment, by using the structural variation detection tool, potential deletion variations in the genome can be efficiently identified, and by analyzing the breakpoints and pairing of reads, possible structural variations can be identified, and the deletion variation region on the genome can be quickly located, providing candidate fragments for subsequent depth feature analysis.

[0067] Specifically, for each candidate structural variation, especially deletion (Deletion), the interval from the start position (SV_start) to the end position (SV_end) of the structural variation, i.e., the read breakpoint (Breakpoints), is calculated. In structural variation detection, the core steps for calculating variation coordinates by parsing standard VCF (Variant Call Format) format files are as follows:

[0068] 1) Chromosome identification: directly extract the CHROM field (first column) as the chromosome number where the SV is located;

[0069] 2) Start position determination: the value of the POS field (second column) is taken as the SV start coordinate SV_start;

[0070] 3) Variant length calculation: the SV length is calculated based on the length difference between the fourth column reference sequence REF and the fifth column variant sequence ALT;

[0071] 4) Termination site derivation: the termination coordinate is determined by linear translation (start site + variant length).

[0072] It should be noted that the above VCF file is a widely used format for storing and exchanging variant information in individual or population genomes; the above CHROM field is the first field in the VCF file, which is used to identify the chromosome or sequence carrying the variation; the above POS field is the second field in the VCF file, which is used to record the start position of the variation on the reference genome; the above REF represents the reference base or reference sequence at the POS start position on the specified CHROM chromosome or sequence, and the content of the REF field can reflect the normal or wild type sequence at the position in the reference genome; the above ALT is used to record the variant sequence relative to the reference sequence, and the content of the ALT field describes the alternative base or sequence observed at the given POS position.

[0073] According to the above embodiment of the application, the depth feature analysis of the candidate gene fragment is performed according to the alignment depth data, and the depth feature analysis result is obtained, including: determining the start position and the end position of the candidate gene fragment; determining the first depth data of the first predetermined position and the second depth data of the second predetermined position according to the alignment depth data, wherein the first predetermined position is a position corresponding to a plurality of sites in the upstream direction of the start position, and the second predetermined position is a position corresponding to a plurality of sites in the downstream direction of the end position; performing depth comparison on each first depth data to obtain a first depth alignment result, and performing depth comparison on the second depth data to obtain a second depth alignment result; obtaining the depth feature analysis result according to the first depth alignment result and the second depth alignment result.

[0074] In this embodiment, the depth feature analysis is performed by comparing the sequencing depth of the candidate variation region and the adjacent region to evaluate the authenticity of the variation. A true deletion variation region usually shows a significant difference in depth from the surrounding genomic sites, and the depth feature analysis can effectively distinguish between true deletion variations and false positive results, thereby improving the accuracy of variation detection.

[0075] Specifically, for the SV start position (SV_start) and the SV end position (SV_end), the depth of the 10 bp before (upstream) and the 10 bp after (downstream) is calculated, that is, the first depth data of the first predetermined position and the second depth data of the second predetermined position. Based on the obtained sequencing depth of each site, the sequencing depth (depth) of each base in the upstream 10 bp (SV_start / SV_end before) and the downstream 10 bp (SV_start / SV_end after) of the start position (SV_start) and the end position (SV_end) of each structural variation is extracted.

[0076] wherein, for the SV_start position: the upstream 10 bp (Upstream) is defined as the interval [SV_start-10, SV_start-1], including the 10 base positions immediately before SV_start, and SV_start itself is not in this upstream interval; the downstream 10 bp (Downstream) is defined as the interval [SV_start+1, SV_start+10], including the 10 base positions immediately after SV_start, and SV_start itself is not in this downstream interval. For the SV_end position: the upstream 10 bp (Upstream) is defined as the interval [SV_end-10, SV_end-1], including the 10 base positions immediately before SV_end, and SV_end itself is not in this upstream interval; the downstream 10 bp (Downstream) is defined as the interval [SV_end+1, SV_end+10], including the 10 base positions immediately after SV_end, and SV_end itself is not in this downstream interval.

[0077] In the above embodiment of the present application, the first depth data is compared to obtain the first depth comparison result, and the second depth data is compared to obtain the second depth comparison result, including: performing rank sum test on each first depth data to obtain a first rank sum test value, and comparing the first rank sum test value with a first predetermined value and a second predetermined value to obtain the first depth comparison result, wherein the first predetermined value is greater than the second predetermined value; performing rank sum test on each second depth data to obtain a second rank sum test value, and comparing the second rank sum test value with a third predetermined value and a fourth predetermined value to obtain the second depth comparison result, wherein the third predetermined value is greater than the fourth predetermined value.

[0078] In this embodiment, by performing rank sum test on the upstream and downstream sites of the candidate variation region, the significance of the depth change can be evaluated, so as to judge the authenticity of the variation, and the variation region with significant depth change can be accurately identified through the rank sum test, thereby improving the reliability of the variation detection.

[0079] It should be noted that the above rank sum test is a non-parametric statistical method, which is suitable for comparing the central tendency of two independent samples.

[0080] Specifically, the Mann-Whitney U rank sum test (non-parametric test) is used to compare the depths before and after SV_start, i.e., [SV_start-10, SV_start-1] and [SV_start+1, SV_start+10], and the depths before and after SV_end, i.e., [SV_end-10, SV_end-1] and [SV_end+1, SV_end+10], and the calculation formula is as follows:

[0081]

[0082]

[0083]

[0084] wherein, and are the sample sizes of the two groups, and are the rank sums of the two groups of data, is the first rank sum test value, is the second rank sum test value, is the test statistic, the value range of .

[0085] It should be noted that the above Mann-Whitney U test, also known as Mann-Whitney-Wilcoxon test or Wilcoxon rank sum test, is a non-parametric statistical test method, which is used to compare whether the central tendency of two independent sample populations is the same, and is particularly suitable for the case where the data does not meet the normal distribution assumption or the data is in the form of ranking; the physical meaning of the above test statistic U value in the Mann-Whitney rank sum test reflects the relative position relationship of the two groups of data, and its essence is to measure the strength of the observation value of one sample being superior to the observation value of another sample; the U value is maximum / minimum, and the P value is very small, indicating that the position difference between the groups is significant; the U value is close to the central value (i.e., U=0), and the P value is very large, indicating that there is no evidence to show the difference between the groups.

[0086] ​In the above embodiment of the present application, the depth feature analysis result is obtained according to the first depth alignment result and the second depth alignment result, including: when the first rank sum test value is greater than the first predetermined value or the first rank sum test value is less than the second predetermined value, it is determined that the start position has a significant depth change; and when the second rank sum test value is greater than the third predetermined value or the second rank sum test value is less than the fourth predetermined value, it is determined that the end position has a significant depth change.

[0087] In this embodiment, by setting the predetermined value, the depth change can be divided into significant and non-significant two categories, which is convenient for subsequent variation screening. The significant depth change is usually associated with the real deletion variation, and the non-significant change may be a false positive or sequencing noise. Through the significance judgment of the depth change, the false positive variation can be effectively filtered out, and the efficiency and accuracy of the variation screening are improved.

[0088] Figure 3 is a flowchart of an optional site depth-based genomic deletion type structural variation accuracy evaluation method according to an embodiment of the present application, as shown in Figure 3 When the P value is less than or equal to 0.05, it is considered that the end point has a significant depth change, supporting the authenticity of the SV. When the P values of the start position and the end position are significant, 1 point is obtained. When the P values of the start position and the end position are significant at any one end, 0.6 points are obtained. When the P values of the start position and the end position are not significant, 0 points are obtained.

[0089] According to the above embodiment of the present application, the depth feature analysis of the candidate gene fragment is performed according to the alignment depth data, and the depth feature analysis result is obtained, including: determining the start position and the end position of the candidate gene fragment; determining the first depth data of the first predetermined position and the second depth data of the second predetermined position according to the alignment depth data, wherein the first predetermined position is a position corresponding to a plurality of sites in the upstream direction of the start position, and the second predetermined position is a position corresponding to a plurality of sites in the downstream direction of the end position; performing linear regression analysis on the first depth data and the second depth data to obtain the regression slope of the start position and the end position within a predetermined range; and obtaining the depth feature analysis result according to the regression slope.

[0090] In this embodiment, the linear regression analysis can reveal the change trend of the depth data, which is helpful for accurately positioning the variation breakpoint. By analyzing the linear relationship of the depth data, the position of the depth mutation, i.e. the start and end points of the deletion variation, can be identified. Through the linear regression analysis, the breakpoint of the depth change can be accurately identified, and the accuracy of the variation positioning is improved.

[0091] Specifically, a segmented linear regression is fitted in the interval ([SV_start-10, SV_start+10] and [SV_end-10, SV_end+10]) to find the optimal breakpoint. If the regression slope around SV_start and SV_end changes significantly (such as a sharp drop in depth) with a P value ≤ 0.05, the authenticity of the SV is supported. If the starting and ending positions are both significant, 1 point is scored; one end is significant, 0.6 points are scored; and neither is significant, 0 points are scored. The segmented linear regression calculation formula is as follows:

[0092]

[0093] wherein, represents a genomic position coordinate (independent variable), represents a sequencing depth (dependent variable) at a position is a candidate breakpoint position, which is each position tested in the window in the script, is an indicator function, is a global intercept term, is a baseline slope, is an additional slope change at the breakpoint, which is a key test parameter, is an error term.

[0094] It should be noted that the breakpoint detection model adopted is based on a segmented linear regression method, and a statsmodels package in Python language is used to realize statistical modeling and significance test. By establishing a depth change model before and after the target site, the statistical significance of the position as a depth change breakpoint is quantitatively evaluated. For a depth data sequence of length n in a given window, the system evaluates all possible positions from index 1 to n-1, and excludes the head and tail to avoid boundary effects. For each candidate position c, the two-tailed test p of the regression coefficient of the interaction term is extracted by least squares regression fitting model. The optimal breakpoint is selected by sorting all calculated p values, and the position meeting the global minimum value condition is selected as the optimal breakpoint. When multiple breakpoints are detected in the window, the position with the smallest global p value is selected preferentially.

[0095] According to the above embodiment of the application, the depth feature analysis result is obtained by performing depth feature analysis on the candidate gene fragment according to the alignment depth data, including: determining the start position and the end position of the candidate gene fragment; determining the third depth data of each position in the first position interval corresponding to the start position and the end position according to the alignment depth data; determining the coefficient of variation of the position interval according to the third depth data; and determining the depth feature analysis result according to the coefficient of variation.

[0096] ​In this embodiment, the depth of the real missing variation region is generally consistent, and the coefficient of variation is small, while the depth of the non-real variation region may fluctuate greatly, and the coefficient of variation is large. By calculating the coefficient of variation, the depth consistency of the candidate variation region can be evaluated, and the real missing variation is further screened out.

[0097] It should be noted that the above coefficient of variation is an important index for measuring the fluctuation degree of data, and can be used to evaluate the depth consistency of the candidate variation region.

[0098] Specifically, the coefficient of variation CV value is calculated by calculating the mean and variance of the depth of the missing variation [SV_start, SV_end] interval site. The smaller the CV value, the more uniform the interval depth is. When the CV value is less than or equal to 20%, 1 point is assigned; when the CV value is greater than 20%, 0 point is assigned. The coefficient of variation CV value calculation formula is as follows:

[0099]

[0100] Wherein, The standard deviation is represented by σ, The mean is represented by μ, The coefficient of variation CV value is represented by CV.

[0101] According to the above embodiment of the present application, the depth feature analysis of the candidate gene fragment is performed according to the alignment depth data, and the depth feature analysis result is obtained, including: determining the start position and the end position of the candidate gene fragment; detecting whether there is a soft-clipped read in the second site interval corresponding to the start position and the end position to obtain a detection result; and obtaining the depth feature analysis result according to the detection result.

[0102] In this embodiment, the existence of soft-clipped reads can be used as direct evidence to support the missing variation. Soft-clipped reads refer to a part of the reads that are not aligned to the reference genome, which is usually caused by the missing variation in the region. By detecting soft-clipped reads, additional supporting evidence can be provided to enhance the authenticity and credibility of the candidate variation.

[0103] Specifically, the pysam software package is used to count the soft-clipped reads in the SV interval. When there are soft-clipped reads in the [SV_start-5, SV_start+5] and [SV_end-5, SV_end+5] intervals, the credibility score of the SV is increased. When there are soft-clipped reads at both ends, the final result is increased by 0.1 points. When there is soft-clipped read at one end, the final result is increased by 0.06 points.

[0104] Note that pysam is a Python interface library for efficiently accessing and manipulating various bioinformatics data formats, especially those commonly used in genomic and genetic research.

[0105] According to the above embodiment of the present application, the target gene fragment in the candidate gene fragment is determined according to the depth feature analysis result, which comprises: determining the score corresponding to each dimension index and the weight of each dimension index in the depth feature analysis result; weighting and summing the depth feature analysis result according to the score and the weight to obtain the total score of the depth feature analysis result; and selecting a gene fragment with a total score not less than a predetermined score from the candidate gene fragments as the target gene fragment.

[0106] In this embodiment, by setting the weight and the threshold, the multiple characteristics of the candidate variation can be comprehensively evaluated, and automatic variation screening is realized. The weight can reflect the importance of different characteristics to the authenticity of the variation, and the total score can comprehensively reflect the information of all characteristics, which is used to judge whether the candidate variation is real. By weighting and summing, high-credibility deletion variations can be automatically screened, which can greatly reduce the workload of manual verification.

[0107] As shown in Figure 3 , four dimension weights of depth difference significance test, segmented linear regression, internal consistency test and interval breakpoint reads support are set to evaluate the final SV authenticity. The first three dimension weights are set to [0.5, 0.4, 0.1], and the interval breakpoint reads support is additionally scored. The full score is 1 point. When the total score is not less than 0.6, it is considered to be a real SV.

[0108] Taking the initial identification of whole genome structural variations by using multiple structural variation detection tools based on the target species' s third-generation sequencing data (PacBio platform) as an example, this embodiment is described. The structural variation detection tools include cuteSV v2.0.2, svim v1.4.2, pbsv v2.10.0 and Sniffles v2.0.7. The SURVIVOR v1.0.7 software is used to merge the raw SV sites output by the above tools across tools to generate a comprehensive non-redundant SV set. On the basis of the merging result, specific screening is performed on the deletion-type variation regions in the genome. In order to improve the reliability of variation authenticity judgment, a quality evaluation system containing multiple dimension score indicators is designed, which is performed on each candidate deletion-type SV interval. The quality evaluation system is as follows:

[0109] 1) Sequencing depth difference significance test: test whether the depth difference between the candidate deletion interval and the adjacent control interval is significant;

[0110] 2) Piecewise linear regression analysis: analyze the depth distribution pattern to precisely define the missing breakpoint;

[0111] 3) Internal consistency test: calculate the coefficient of variation CV value of the candidate missing interval;

[0112] 4) Interval breakpoint reads support verification: check whether there is reads evidence supporting the SV breakpoint position.

[0113] According to the above scoring system, each deletion type SV obtains a comprehensive confidence score. The above technical solutions provided in the embodiments of the present application will be described in combination with Table 1, wherein Table 1 shows the score results of four indicators of part of the structural variations:

[0114] Table 1

[0115] Structural variation Start position Wilcoxon rank test P-value End position Wilcoxon rank test P-value Start position piecewise linear regression P-value End position piecewise linear regression P-value Coefficient of variation (%) Start position number of softclipped reads End position number of softclipped reads Wilcoxon rank test score Piecewise linear regression score Coefficient of variation score Softclipped reads score Final score Determine if SV is real based on score Determine if SV is real using IGV SV1 1.00 1.00 0.08 1.00 18.15 0.00 0.00 0.00 0.00 1.00 0.00 0.10 False False SV2 2.52E-06 1.00 6.07E-06 8.90E-06 143.75 0.00 1.00 0.60 1.00 0.00 0.06 0.76 True True SV3 6.30E-08 6.30E-08 2.13E-08 6.23E-09 749.11 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV4 6.30E-08 6.30E-08 9.23E-09 5.89E-10 81.39 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV5 1.12E-06 0.01 1.07E-12 1.00 1048.02 0.00 0.00 1.00 0.60 0.00 0.00 0.74 True True SV6 6.30E-08 0.28 3.65E-07 3.65E-07 0.00 0.00 0.00 0.60 1.00 1.00 0.00 0.8 True True SV7 6.30E-08 6.30E-08 1.26E-07 3.65E-07 118.95 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV8 6.30E-08 6.30E-08 3.65E-07 2.31E-07 0.00 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV9 6.30E-08 2.75E-07 1.22E-08 1.22E-08 111.20 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV10 6.30E-08 2.75E-07 3.24E-11 6.74E-10 14.11 1.00 1.00 1.00 1.00 1.00 0.10 1.00 True True SV11 6.30E-08 1.12E-06 2.95E-10 1.32E-11 12.29 1.00 0.00 1.00 1.00 1.00 0.06 1.00 True True SV12 6.30E-08 4.26E-06 3.65E-07 3.65E-07 4.81 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV13 0.28 2.75E-07 0.00 1.22E-08 9.25 0.00 0.00 0.60 1.00 1.00 0.00 0.8 True True SV14 6.30E-08 6.30E-08 4.62E-08 3.65E-07 6.17 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV15 6.30E-08 2.75E-07 1.22E-08 1.22E-08 7.77 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV16 6.30E-08 6.30E-08 3.65E-07 3.65E-07 0.00 0.00 2.00 1.00 1.00 0.00 0.06 0.96 True True SV17 6.30E-08 2.87E-06 1.22E-08 1.50E-14 912.85 1.00 0.00 1.00 1.00 0.00 0.06 0.96 True True SV18 0.42 1.00 0.11 1.00 4.58 0.00 0.00 0.00 0.00 1.00 0.00 0.10 False False SV19 6.30E-08 2.75E-07 1.22E-08 1.22E-08 854.40 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV20 6.30E-08 2.75E-07 1.22E-08 1.22E-08 3108.05 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV21 6.30E-08 6.30E-08 1.86E-08 3.47E-07 2408.32 0.00 6.00 1.00 1.00 0.00 0.06 0.96 True False SV22 6.30E-08 2.75E-07 1.22E-08 1.22E-08 40.53 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV23 6.30E-08 2.75E-07 1.22E-08 1.22E-08 1473.09 0.00 1.00 1.00 1.00 0.00 0.06 0.96 True True SV24 6.30E-08 2.75E-07 1.33E-09 1.22E-08 26.14 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV25 6.30E-08 1.12E-06 4.90E-08 3.65E-07 73.37 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV26 6.30E-08 6.30E-08 3.65E-07 4.19E-07 65.66 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV27 0.01 0.79 1.22E-08 0.00 3.37 0.00 0.00 0.60 1.00 1.00 0.00 0.8 True False SV28 6.30E-08 2.75E-07 1.22E-08 3.04E-09 1926.14 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV29 6.30E-08 2.75E-07 1.22E-08 1.22E-08 8.19 0.00 1.00 1.00 1.00 1.00 0.06 1.00 True True SV30 6.30E-08 6.30E-08 3.65E-07 3.65E-07 66.61 0.00 1.00 1.00 1.00 0.00 0.06 0.96 True True SV31 0.15 1.00 1.00 1.00 34.52 0.00 0.00 0.00 0.00 0.00 0.00 0.00 False False SV32 6.30E-08 6.30E-08 3.65E-07 2.30E-07 1.98 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV33 1.12E-06 1.00 3.17E-13 6.47E-05 20.69 0.00 0.00 0.60 1.00 0.00 0.00 0.70 True True SV34 1.00 4.96E-05 0.02 1.11E-12 19.86 0.00 0.00 0.60 1.00 1.00 0.00 0.8 True True SV35 1.00 1.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 False False SV36 3.28E-06 7.90E-08 5.33E-06 3.15E-09 21.55 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV37 0.59 1.00 1.00 1.00 57.95 0.00 0.00 0.00 0.00 0.00 0.00 0.00 False False SV38 6.30E-08 6.30E-08 1.16E-08 1.24E-07 782.08 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV39 6.30E-08 6.30E-08 1.97E-05 3.65E-07 46.15 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV40 1.00 0.10 0.01 4.18E-08 19.23 1.00 0.00 0.00 1.00 1.00 0.06 0.56 False False SV41 6.30E-08 2.56E-07 1.22E-08 5.97E-08 9.22 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV42 6.30E-08 6.30E-08 3.65E-07 3.65E-07 0.00 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV43 6.30E-08 6.30E-08 3.65E-07 3.38E-09 10.58 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV44 0.15 1.00 0.00 1.00 19.43 0.00 0.00 0.00 0.60 1.00 0.00 0.34 False False SV45 6.30E-08 6.30E-08 3.65E-07 3.65E-07 0.00 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV46 6.30E-08 2.75E-07 1.22E-08 1.22E-08 1200.00 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV47 6.30E-08 6.30E-08 3.65E-07 1.98E-07 1.62 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV48 0.02 0.30 1.45E-05 1.00 59.15 0.00 0.00 0.60 0.60 0.00 0.00 0.54 False False SV49 6.30E-08 2.75E-07 1.22E-08 1.22E-08 25.78 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV50 6.30E-08 6.30E-08 3.65E-07 3.65E-07 0.00 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV51 6.30E-08 2.75E-07 1.22E-08 1.22E-08 19.87 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV52 1.00 1.50E-05 1.00 2.67E-05 81.87 0.00 0.00 0.60 0.60 0.00 0.00 0.54 False False SV53 6.30E-08 6.30E-08 3.65E-07 3.65E-07 0.00 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV54 6.30E-08 2.75E-07 1.22E-08 1.22E-08 24.11 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV55 1.00 1.00 1.00 1.00 0.00 0.00 0.00 0.00 0.00 1.00 0.00 0.10 False False SV56 6.30E-08 2.75E-07 1.22E-08 1.22E-08 17.67 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV57 8.07E-06 2.75E-07 6.58E-09 1.22E-08 14.22 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV58 0.00 2.75E-07 2.57E-06 1.30E-09 4.59 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV59 6.30E-08 9.13E-07 1.22E-08 1.04E-09 14.34 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV60 0.59 0.42 1.00 6.34E-05 18.91 0.00 0.00 0.00 0.60 1.00 0.00 0.34 False False SV61 6.30E-08 2.75E-07 1.98E-07 3.65E-07 26.63 2.00 0.00 1.00 1.00 0.00 0.06 0.96 True True SV62 6.30E-08 2.75E-07 1.22E-08 1.22E-08 9.67 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV63 6.30E-08 1.05E-06 1.22E-08 1.27E-09 10.23 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV64 6.30E-08 2.75E-07 2.19E-07 1.22E-08 90.05 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV65 1.12E-06 0.18 1.22E-08 3.37E-06 4.02 0.00 0.00 0.60 1.00 1.00 0.00 0.8 True True SV66 1.00 0.03 1.00 1.00 9.40 0.00 0.00 0.60 0.00 1.00 0.00 0.40 False False SV67 6.30E-08 2.75E-07 1.46E-08 1.22E-08 2.54 0.00 1.00 1.00 1.00 1.00 0.06 1.00 True True SV68 6.80E-08 1.29E-06 2.16E-08 0.00 178.75 1.00 0.00 1.00 1.00 0.00 0.06 0.96 True True SV69 0.79 0.28 1.00 6.70E-05 7.45 0.00 0.00 0.00 0.60 1.00 0.00 0.34 False True SV70 0.00 2.75E-07 1.22E-08 4.11E-09 216.38 0.00 1.00 1.00 1.00 0.00 0.06 0.96 True True SV71 6.30E-08 4.96E-05 3.65E-07 3.65E-07 0.00 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV72 1.23E-07 0.63 2.81E-05 0.01 30.45 0.00 0.00 0.60 1.00 0.00 0.00 0.70 True True SV73 6.30E-08 6.30E-08 3.65E-07 3.65E-07 0.00 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV74 6.30E-08 2.52E-06 2.27E-10 2.00E-05 67.98 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV75 0.00 0.00 0.08 0.08 140.68 0.00 0.00 1.00 0.00 0.00 0.00 0.50 False False SV76 1.00 1.00 5.47E-06 1.22E-08 12.98 0.00 0.00 0.00 1.00 1.00 0.00 0.50 False False SV77 6.30E-08 2.75E-07 1.22E-08 1.22E-08 1907.88 2.00 0.00 1.00 1.00 0.00 0.06 0.96 True True SV78 6.30E-08 6.30E-08 1.10E-08 3.65E-07 10.02 1.00 0.00 1.00 1.00 1.00 0.06 1.00 True True SV79 6.30E-08 6.30E-08 3.65E-07 6.12E-07 13.82 0.00 2.00 1.00 1.00 1.00 0.06 1.00 True True SV80 6.30E-08 9.17E-08 1.22E-08 9.25E-10 138.68 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV81 6.30E-08 6.30E-08 3.65E-07 3.65E-07 0.78 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV82 2.75E-07 6.30E-08 0.00 3.65E-07 34.99 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV83 0.01 2.75E-07 1.22E-08 1.22E-08 33.52 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV84 1.50E-05 2.75E-07 1.22E-08 1.22E-08 17.56 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV85 6.30E-08 2.75E-07 1.22E-08 1.22E-08 20.57 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV86 6.30E-08 2.75E-07 1.22E-08 1.22E-08 1421.27 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV87 1.00 1.00 1.00 1.00 21.66 0.00 0.00 0.00 0.00 0.00 0.00 0.00 False False SV88 0.00 6.30E-08 0.02 2.42E-07 4.59 0.00 2.00 1.00 1.00 1.00 0.06 1.00 True True SV89 6.30E-08 2.75E-07 1.22E-08 1.22E-08 2685.14 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV90 2.75E-07 1.50E-05 4.21E-05 1.22E-08 2858.32 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV91 6.80E-08 6.30E-08 2.84E-08 1.46E-06 436.18 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV92 6.30E-08 2.75E-07 0.00 1.22E-08 38.72 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV93 6.30E-08 6.30E-08 3.65E-07 3.65E-07 0.00 0.00 0.00 1.00 1.00 1.00 0.00 1.00 True True SV94 6.30E-08 2.75E-07 1.22E-08 1.22E-08 5448.85 0.00 1.00 1.00 1.00 0.00 0.06 0.96 True True SV95 0.00 4.54E-07 1.00 9.64E-07 334.59 0.00 0.00 1.00 0.60 0.00 0.00 0.74 True True SV96 1.00 0.79 1.00 3.60E-05 4.32 0.00 0.00 0.00 0.60 1.00 0.00 0.34 False False SV97 6.30E-08 6.30E-08 1.97E-07 3.65E-07 93.68 0.00 0.00 1.00 1.00 0.00 0.00 0.90 True True SV98 6.30E-08 1.00 1.22E-08 1.00 36.71 0.00 0.00 0.60 0.60 0.00 0.00 0.54 False False SV99 0.79 1.00 1.00 1.00 5.83 0.00 0.00 0.00 0.00 1.00 0.00 0.10 False False SV100 1.00 0.79 6.08E-05 0.08 32.33 8.00 0.00 0.00 0.60 0.00 0.06 0.30 False False

[0116] As shown in Table 1 above, the score results of each indicator of the structural variation are recorded, including the starting position rank sum test P value, the ending position rank sum test P value, the starting position piecewise linear regression P value, the ending position piecewise linear regression P value, the coefficient of variation, the number of soft clipped reads at the starting position, the number of soft clipped reads at the ending position, the rank sum test score, the piecewise linear regression score, the coefficient of variation score, the soft clipped reads score, the final score, the judgment of whether the SV is real according to the score, and the IGV verification of whether the SV is real.

[0117] To verify the effectiveness of the scoring system, 100 candidate deletion type structural variation sites were selected by random sampling method for independent verification, and the manual review results visualized by IGV were used as the standard reference (Gold Standard). By comparing the scoring threshold (≥0.6) with the verification results, the maximum effectiveness of the threshold in judging the authenticity of SV was determined.

[0118] The above technical solutions provided in the embodiments of the present application will be described in combination with Table 2, wherein Table 2 shows the constructed confusion matrix:

[0119] Table 2

[0120]

[0121] As shown in Table 2 above, the independent verification results of 100 candidate deletion type SV sites are recorded, the confusion matrix (Confusion Matrix) is constructed, and the following evaluation indicators are calculated:

[0122] True positive (True Positive, TP): the number of real SVs correctly identified by the model;

[0123] True Negative (TN): the number of non-SV sites correctly excluded by the model;

[0124] False Positive (FP): the number of non-target sites misjudged as SV by the model;

[0125] False Negative (FN): the number of real SVs missed by the model;

[0126] According to the above evaluation indicators, the core performance indicators are defined as follows:

[0127] Accuracy reflects the overall discriminant ability of the model, and the calculation formula is: ;

[0128] Precision (Precision) represents the reliability of the prediction result, and the calculation formula is: ;

[0129] Recall / Sensitivity (Recall / Sensitivity) evaluates the coverage rate of the model to the real SV, and the calculation formula is: ;

[0130] F1 Score is the harmonic mean of precision and recall, and the calculation formula is: ;

[0131] The calculation results of each index are derived from the confusion matrix shown in Table 2 above, which quantitatively evaluates the discriminant performance of the SV detection scheme:

[0132] Accuracy = (78 + 19) / (78 + 19 + 2 + 1) = 0.97;

[0133] Precision = 78 / (78 + 2) = 0.98;

[0134] Recall = 78 / (78 + 1) = 0.99;

[0135] F1 Score = 2 * (0.98 * 0.99) / (0.98 + 0.99) = 0.98.

[0136] According to the above data analysis, the application of the missing structural variation screening process based on multi-tool integration and multi-dimensional scoring can improve the detection reliability and screening efficiency of missing structural variations when processing genomic data, and the accuracy, precision, recall and F1 score are all above 0.95. This process can efficiently and reliably screen missing structural variations in species genomes.

[0137] From the above, in the embodiment of the application, based on the alignment result of the third-generation sequencing data (such as PacBio or ONT), the depth of each site on the genome is calculated using Samtools (v1.17) software, the interval of the deletion type structural variation is extracted from the genomic structural variation result, the authenticity of the structural variation is comprehensively evaluated according to the significance test of the depth difference before and after the SV region, the segmented linear regression, the SV internal consistency test, and the interval breakpoint reads support dimension setting weight, an automatic scoring system is constructed, which is used to evaluate the credibility of SV, and the SV quality scoring method based on the alignment depth and the fusion of multiple evidences has important significance for promoting the standardization and automation of genomic structural variation research, and can construct an automatic scoring system based on the depth difference of the sites before and after the SV, the reads alignment characteristics of the breakpoint region, and the sequencing depth consistency of the SV region, thereby reducing manual intervention and improving the reliability of SV detection.

[0138] The technical scheme provided by the above embodiment of the application solves the following technical problems: 1) the genomic structural variation detection and merging tool has a high false positive rate, the structural variation detection verification link is low in efficiency, professional personnel need to be invested to perform time-consuming and tedious visual inspection, time and effort are wasted, subjective bias is easily introduced, in a large-scale sequencing project, the number of SVs to be verified is huge, the efficiency of manual screening is extremely low, which seriously restricts the scalability and standardization of SV research; 2) subjective bias is introduced, the evaluation result is easily affected by the experience and judgment of the analyst, which reduces the reliability and repeatability of the result, and seriously hinders the standardization of SV data integration between different studies, limits the scale, efficiency and scalability of SV research, and has the following beneficial effects: 1) a quantitative confidence score (Quality Score) is automatically generated by comprehensively analyzing multiple objective indicators including the alignment depth difference of the sites before and after the SV region, the number of multiple supporting reads (such as soft clipped reads), and the local sequence characteristics (depth consistency of the structural variation region); 2) the dependence on manual verification is significantly reduced, the verification efficiency is greatly improved, the subjective factor interference is eliminated, and the overall reliability and accuracy of the SV detection result are finally improved.

[0139] That is, the above technical scheme provided by the embodiment of the application comprehensively considers multiple factors, constructs an automatic scoring system, and deeply evaluates the reliability of the deletion type SV, thereby solving the problem that the structural variation detection verification link highly depends on the visual tool (IGV) for manual checking, reducing the manual verification link, improving the efficiency of verifying the accuracy of SV, promoting the scaling, standardization and automation process of genomic structural variation research, and laying a reliable data foundation for subsequent biological function analysis and clinical application (such as disease association analysis and precise diagnosis), which is of great significance.

[0140] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all described as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and a necessary general hardware platform, and of course it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, an optical disk) and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the method of each embodiment of the present application.

[0142] According to the embodiments of the present application, a site depth-based genomic deletion-type structural variation accuracy evaluation device for implementing the site depth-based genomic deletion-type structural variation accuracy evaluation method is also provided, Figure 4 which is a schematic diagram of the site depth-based genomic deletion-type structural variation accuracy evaluation device according to the embodiments of the present application, as shown in Figure 4 the device includes a calculation module 401, an identification module 403, an analysis module 405, and a determination module 407. The site depth-based genomic deletion-type structural variation accuracy evaluation device will be described below.

[0143] The calculation module 401 is configured to calculate the alignment depth of each site on the genome by using a data processing software to obtain alignment depth data, wherein the genome is an object to be identified for deletion-type structural variation.

[0144] The identification module 403 is configured to identify the genomic deletion variation of the genome by using a structural variation detection tool to obtain candidate gene fragments with genomic deletion variation on the genome.

[0145] The analysis module 405 is configured to perform depth feature analysis on the candidate gene fragments according to the alignment depth data to obtain a depth feature analysis result.

[0146] The determination module 407 is configured to determine a target gene fragment in the candidate gene fragments according to the depth feature analysis result, wherein the target gene fragment is a gene fragment with deletion-type structural variation in the candidate gene fragments.

[0147] It should be noted that the above calculation module 401, identification module 403, analysis module 405 and determination module 407 correspond to steps S202 to S208 in the above embodiment, and the four modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in the above embodiment.

[0148] As can be seen from the above, in the scheme described in the above embodiment of the present application, first, the calculation module can calculate the alignment depth of each site on the genome by comparing the data processing software to obtain alignment depth data, wherein the genome is the object to be identified for deletion type structural variation; second, the identification module can identify the genomic deletion variation of the genome using a structural variation detection tool to obtain a candidate gene fragment with genomic deletion variation on the genome; third, the analysis module can analyze the depth characteristics of the candidate gene fragment according to the alignment depth data to obtain a depth characteristic analysis result; and finally, the determination module can determine a target gene fragment in the candidate gene fragment according to the depth characteristic analysis result, wherein the target gene fragment is a gene fragment with deletion type structural variation in the candidate gene fragment. By using the depth information of third-generation sequencing data and multi-dimensional sequencing evidence fusion, including site depth difference analysis, segmented linear regression, SV region depth consistency test and support degree of soft sheared reads, the purpose of automatically evaluating the quality of deletion type structural variation is achieved, thereby realizing the technical effect of significantly improving the accuracy and efficiency of SV detection.

[0149] Through the above technical scheme provided by the embodiment of the present application, the technical problem of high false positive rate in the deletion type structural variation detection method in the related art, which leads to low reliability of the detection result, is solved.

[0150] In an optional embodiment, the calculation module comprises a calculation unit configured to calculate the alignment depth of each site on the genome by using the depth parameter of Samtools to obtain alignment depth data.

[0151] In an optional embodiment, the identification module comprises a finding unit configured to find the read breakpoint of the genome by using the structural variation detection tool to obtain a read breakpoint finding result; a searching unit configured to search for mismatched paired reads in the genome by using the structural variation detection tool to obtain a mismatched read result; and a first processing unit configured to obtain a candidate gene fragment according to the read breakpoint finding result and the mismatched read result.

[0152] In an alternative embodiment, the analysis module comprises: a first determining unit configured to determine the start position and the end position of the candidate gene fragment; a second determining unit configured to determine, according to the aligned depth data, first depth data of a first predetermined position and second depth data of a second predetermined position, wherein the first predetermined position is a position corresponding to a plurality of sites in an upstream direction of the start position, and the second predetermined position is a position corresponding to a plurality of sites in a downstream direction of the end position; a comparison unit configured to perform depth comparison on each of the first depth data to obtain a first depth comparison result, and perform depth comparison on each of the second depth data to obtain a second depth comparison result; and a second processing unit configured to obtain a depth feature analysis result according to the first depth comparison result and the second depth comparison result.

[0153] In an alternative embodiment, the comparison unit comprises: a first inspection subunit configured to perform rank sum inspection on each of the first depth data to obtain a first rank sum inspection value, and compare the first rank sum inspection value with a first predetermined value and a second predetermined value to obtain a first depth comparison result, wherein the first predetermined value is greater than the second predetermined value; and a second inspection subunit configured to perform rank sum inspection on each of the second depth data to obtain a second rank sum inspection value, and compare the second rank sum inspection value with a third predetermined value and a fourth predetermined value to obtain a second depth comparison result, wherein the third predetermined value is greater than the fourth predetermined value.

[0154] In an alternative embodiment, the second processing unit comprises: a first determining subunit configured to determine that there is a significant depth change at the start position when the first depth comparison result indicates that the first rank sum inspection value is greater than the first predetermined value or the first rank sum inspection value is less than the second predetermined value; and a second determining subunit configured to determine that there is a significant depth change at the end position when the second depth comparison result indicates that the second rank sum inspection value is greater than the third predetermined value or the second rank sum inspection value is less than the fourth predetermined value.

[0155] In an alternative embodiment, the analysis module comprises: a third determining unit configured to determine the start position and the end position of the candidate gene fragment; a fourth determining unit configured to determine, according to the aligned depth data, first depth data of a first predetermined position and second depth data of a second predetermined position, wherein the first predetermined position is a position corresponding to a plurality of sites in an upstream direction of the start position, and the second predetermined position is a position corresponding to a plurality of sites in a downstream direction of the end position; an analysis unit configured to perform linear regression analysis on the first depth data and the second depth data to obtain a regression slope of the start position and the end position within a predetermined range; and a third processing unit configured to obtain a depth feature analysis result according to the regression slope.

[0156] In an optional embodiment, the analyzing module comprises: a fifth determining unit configured to determine the start position and the end position of the candidate gene fragment; a sixth determining unit configured to determine third depth data of each site in the first site interval corresponding to the start position and the end position according to the alignment depth data; a seventh determining unit configured to determine a variation coefficient of the site interval according to the third depth data; and an eighth determining unit configured to determine the depth feature analysis result according to the variation coefficient.

[0157] In an optional embodiment, the analyzing module comprises: a ninth determining unit configured to determine the start position and the end position of the candidate gene fragment; a fourth processing unit configured to determine whether there is a soft-clipped read in the second site interval corresponding to the start position and the end position, and obtain a detection result; and a fifth processing unit configured to obtain the depth feature analysis result according to the detection result.

[0158] In an optional embodiment, the determining module comprises: a tenth determining unit configured to determine a score corresponding to each dimension index in the depth feature analysis result and a weight of each dimension index; a summing unit configured to perform weighted summation on the depth feature analysis result according to the score and the weight, and obtain a total score of the depth feature analysis result; and a selecting unit configured to select a gene fragment with a total score not less than a predetermined score from the candidate gene fragments as a target gene fragment.

[0159] According to another aspect of the embodiments of the present application, a processor is also provided, which is configured to run a program, wherein the program is configured to perform the site depth-based genomic deletion-type structural variation accuracy evaluation method according to any one of the above embodiments when the program is running.

[0160] According to another aspect of the embodiments of the present application, a computer program product is also provided, which comprises computer instructions configured to perform the site depth-based genomic deletion-type structural variation accuracy evaluation method according to any one of the above embodiments when the computer instructions are executed by a processor.

[0161] According to another aspect of the embodiments of the present application, a computer readable storage medium is also provided, which comprises a stored program, wherein the program is configured to perform the site depth-based genomic deletion-type structural variation accuracy evaluation method according to any one of the above embodiments.

[0162] Optionally, in the present embodiment, the computer readable storage medium can be located in any one of computer terminals in a computer terminal group in a computer network, or in any one of communication devices in a communication device group.

[0163] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: obtaining alignment depth data of each site on the genome by calculating the alignment depth of each site on the genome through the data processing software, wherein the genome is an object to which deletion type structural variation identification is required; obtaining a candidate gene fragment with a genomic deletion variation on the genome by performing genomic deletion variation identification on the genome using a structural variation detection tool; obtaining a depth feature analysis result by performing depth feature analysis on the candidate gene fragment according to the alignment depth data; and determining a target gene fragment in the candidate gene fragment according to the depth feature analysis result, wherein the target gene fragment is a gene fragment with a deletion type structural variation in the candidate gene fragment.

[0164] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: obtaining alignment depth data of each site on the genome by calculating the alignment depth of each site on the genome through the depth parameter of Samtools.

[0165] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: obtaining a read breakpoint finding result by finding read breakpoints on the genome through the structural variation detection tool; obtaining a mismatched read result by searching for mismatched paired reads in the genome through the structural variation detection tool; and obtaining a candidate gene fragment according to the read breakpoint finding result and the mismatched read result.

[0166] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: determining a start position and an end position of the candidate gene fragment; determining first depth data of a first predetermined position and second depth data of a second predetermined position according to the alignment depth data, wherein the first predetermined position is a position corresponding to a plurality of sites in an upstream direction of the start position, and the second predetermined position is a position corresponding to a plurality of sites in a downstream direction of the end position; performing depth comparison on each first depth data to obtain a first depth comparison result, and performing depth comparison on the second depth data to obtain a second depth comparison result; and obtaining a depth feature analysis result according to the first depth comparison result and the second depth comparison result.

[0167] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: performing a rank-sum test on each first depth data to obtain a first rank-sum test value, and comparing the first rank-sum test value with a first predetermined value and a second predetermined value to obtain a first depth comparison result, wherein the first predetermined value is greater than the second predetermined value; performing a rank-sum test on each second depth data to obtain a second rank-sum test value, and comparing the second rank-sum test value with a third predetermined value and a fourth predetermined value to obtain a second depth comparison result, wherein the third predetermined value is greater than the fourth predetermined value.

[0168] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: when the first depth comparison result indicates that the first rank sum test value is greater than the first predetermined value or the first rank sum test value is less than the second predetermined value, determine that there is a significant depth change at the starting position; when the second depth comparison result indicates that the second rank sum test value is greater than the third predetermined value or the second rank sum test value is less than the fourth predetermined value, determine that there is a significant depth change at the ending position.

[0169] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining the start and end positions of the candidate gene fragment; determining first depth data for a first predetermined position and second depth data for a second predetermined position based on the alignment depth data, wherein the first predetermined position is the position corresponding to multiple sites in the upstream direction of the start position, and the second predetermined position is the position corresponding to multiple sites in the downstream direction of the end position; performing linear regression analysis on the first depth data and the second depth data to obtain the regression slope of the start and end positions within a predetermined range; and obtaining the depth feature analysis result based on the regression slope.

[0170] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining the start and end positions of the candidate gene fragment; determining the third depth data of each point within the first point interval corresponding to the start and end positions based on the alignment depth data; determining the coefficient of variation of the site interval based on the third depth data; and determining the depth feature analysis result based on the coefficient of variation.

[0171] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining the start and end positions of the candidate gene fragment; detecting whether a soft-splitting read exists within a second site interval corresponding to the start and end positions, and obtaining a detection result; and obtaining a deep feature analysis result based on the detection result.

[0172] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: determining the score corresponding to each dimension index in the deep feature analysis result and the weight of each dimension index; performing weighted summation on the deep feature analysis result according to the score and the weight to obtain a total score of the deep feature analysis result; and selecting a gene fragment with a total score not less than a predetermined score from the candidate gene fragments as the target gene fragment.

[0173] The serial numbers of the above embodiments of the application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0174] In the above embodiments of the application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0175] In several embodiments provided in the present application, it should be understood that the disclosed technical contents can be implemented by other manners. Among them, the above-described device embodiments are only schematic, for example, the division of units can be a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, which can be electrical or other forms.

[0176] The technical features of the above embodiments can be combined arbitrarily, and in order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0177] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Part or all of the units can be selected to achieve the purpose of the embodiment according to actual needs.

[0178] In addition, each functional unit in each embodiment of the application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of software functional unit.

[0179] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0180] The above is only the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. A method for assessing the accuracy of site-depth-based genomic deletion structural variations, characterized in that, include: The alignment depth of each point on the genome is calculated by alignment data processing software to obtain alignment depth data, wherein the genome is the object that needs to be identified for deletion-type structural variations. The genome is analyzed using a structural variation detection tool to identify genomic deletion variations, thereby obtaining candidate gene fragments on the genome that contain the genomic deletion variations. Based on the alignment depth data, the candidate gene fragments are subjected to depth feature analysis to obtain the depth feature analysis results; The target gene fragment in the candidate gene fragment is determined based on the deep feature analysis results, wherein the target gene fragment is a gene fragment with deletion-type structural variation in the candidate gene fragment.

2. The method for assessing the accuracy of site-depth-based genomic deletion structural variations according to claim 1, characterized in that, The alignment data processing software is Samtools. The alignment depth of each point on the genome is calculated using this software to obtain alignment depth data, including: The alignment depth data is obtained by calculating the alignment depth of each point on the genome using the depth parameter of Samtools.

3. The method for assessing the accuracy of site-depth-based genomic deletion structural variations according to claim 1, characterized in that, Using structural variation detection tools, the genome is used to identify genomic deletion variations, resulting in candidate gene fragments containing these genomic deletion variations, including: The structural variation detection tool was used to locate read breakpoints in the genome, and the results of the read breakpoint location were obtained. The structural variation detection tool is used to search for mismatched paired reads in the genome to obtain the results of mismatched reads; The candidate gene fragments are obtained based on the results of the read breakpoint search and the results of the mismatched reads.

4. The method for assessing the accuracy of site-depth-based genomic deletion structural variations according to claim 1, characterized in that, Based on the alignment depth data, a deep feature analysis is performed on the candidate gene fragments to obtain the deep feature analysis results, including: Determine the start and end positions of the candidate gene fragments; The first depth data of the first predetermined position and the second depth data of the second predetermined position are determined based on the comparison depth data, wherein the first predetermined position is the position corresponding to multiple points in the upstream direction of the starting position, and the second predetermined position is the position corresponding to multiple points in the downstream direction of the ending position. The first depth data are compared to obtain the first depth comparison result, and the second depth data are compared to obtain the second depth comparison result. The depth feature analysis results are obtained based on the first depth comparison result and the second depth comparison result.

5. The method for assessing the accuracy of site-depth-based genomic deletion structural variations according to claim 4, characterized in that, A first depth comparison result is obtained by performing depth comparison on each of the first depth data, and a second depth comparison result is obtained by performing depth comparison on the second depth data, including: Perform a rank-sum test on each of the first depth data to obtain a first rank-sum test value, and compare the first rank-sum test value with a first predetermined value and a second predetermined value to obtain the first depth comparison result, wherein the first predetermined value is greater than the second predetermined value; A rank-sum test is performed on each of the second depth data to obtain a second rank-sum test value. The second rank-sum test value is then compared with a third predetermined value and a fourth predetermined value to obtain the second depth comparison result, wherein the third predetermined value is greater than the fourth predetermined value.

6. The method for assessing the accuracy of site-depth-based genomic deletion structural variations according to claim 5, characterized in that, The depth feature analysis results are obtained based on the first depth comparison result and the second depth comparison result, including: When the first depth comparison result indicates that the first rank sum test value is greater than the first predetermined value or the first rank sum test value is less than the second predetermined value, it is determined that there is a significant depth change at the starting position. When the second depth comparison result indicates that the second rank sum test value is greater than the third predetermined value or the second rank sum test value is less than the fourth predetermined value, it is determined that there is a significant depth change at the termination position.

7. The method for assessing the accuracy of site-depth-based genomic deletion structural variations according to claim 1, characterized in that, Based on the alignment depth data, a deep feature analysis is performed on the candidate gene fragments to obtain the deep feature analysis results, including: Determine the start and end positions of the candidate gene fragments; The first depth data of the first predetermined position and the second depth data of the second predetermined position are determined based on the comparison depth data, wherein the first predetermined position is the position corresponding to multiple points in the upstream direction of the starting position, and the second predetermined position is the position corresponding to multiple points in the downstream direction of the ending position. Linear regression analysis is performed on the first depth data and the second depth data to obtain the regression slope of the starting position and the ending position within a predetermined range; The depth feature analysis results are obtained based on the regression slope.

8. The method for assessing the accuracy of site-depth-based genomic deletion structural variations according to claim 1, characterized in that, Based on the alignment depth data, a deep feature analysis is performed on the candidate gene fragments to obtain the deep feature analysis results, including: Determine the start and end positions of the candidate gene fragments; Based on the comparison depth data, determine the third depth data of each point within the first point interval corresponding to the start position and the end position; The coefficient of variation of the site interval is determined based on the third depth data; The depth feature analysis result is determined based on the coefficient of variation.

9. The method for assessing the accuracy of site-depth-based genomic deletion structural variations according to claim 1, characterized in that, Based on the alignment depth data, a deep feature analysis is performed on the candidate gene fragments to obtain the deep feature analysis results, including: Determine the start and end positions of the candidate gene fragments; The detection result is obtained by detecting whether a soft-sheared read segment exists within the second site interval corresponding to the start position and the end position; The depth feature analysis results are obtained based on the detection results.

10. The method for assessing the accuracy of site-depth-based genomic deletion structural variations according to any one of claims 1 to 9, characterized in that, Based on the deep feature analysis results, the target gene fragment among the candidate gene fragments is determined, including: Determine the scores and weights of each dimension index in the deep feature analysis results; The deep feature analysis results are weighted and summed according to the scores and weights to obtain the total score of the deep feature analysis results. Gene fragments with a total score not less than a predetermined score are selected from the candidate gene fragments as the target gene fragments.

11. A device for assessing the accuracy of site-depth-based genomic deletion structural variations, characterized in that, include: The calculation module is used to calculate the alignment depth of each point on the genome through alignment data processing software to obtain alignment depth data, wherein the genome is the object that needs to be identified for deletion-type structural variations; The identification module is used to identify genomic deletion variants in the genome using a structural variation detection tool, and to obtain candidate gene fragments on the genome that contain the genomic deletion variants. The analysis module is used to perform deep feature analysis on the candidate gene fragments based on the alignment depth data, and obtain deep feature analysis results; The determination module is used to determine the target gene fragment among the candidate gene fragments based on the deep feature analysis results, wherein the target gene fragment is a gene fragment with deletion-type structural variation among the candidate gene fragments.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program executes the site-depth-based method for assessing the accuracy of genomic deletion structural variations as described in any one of claims 1 to 10.

13. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they perform the method for assessing the accuracy of site-depth-based genomic deletion structural variations as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method for detecting copy number variations by genome sequencing fragments

    CN104603284A

  • Method and system for accurately analyzing DMD gene structural variation breakpoint

    CN107368708A

  • Method and device for analyzing gene

    WO2016208827A1