Low-coverage hi-fi data structural variant calling integration method and system
Patent Information
- Application Number
- CN202610639663.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-09-01
AI Technical Summary
然而,现有方法通常依赖较充分的读段支持或较完整的组装结果,在低覆盖度HiFi数据下容易出现结构变异漏检的情形,整体召回率较低;基于读段比对的方法更适合中小尺度变异,基于组装的方法更适合大尺度变异,现有技术难以兼顾不同尺度结构变异的检测性能;现有集成方式较为简单,对不同检测结果通常采用简单并集、交集或固定阈值合并,难以充分利用不同方法之间的互补信息,此外,现有集成方法在提高检出率时容易引入较多假阳性,而在提高准确率时又容易降低真实结构变异的检出能力
Smart Images

Figure CN122676907A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an integrated method and system for detecting structural variations in low-coverage HiFi data, belonging to the field of gene structural variation detection technology. Background Technology
[0002] Structural variants (SVs) are rearrangement events in the genome, typically longer than 50 bp, including deletions, insertions, duplications, inversions, translocations, and complex rearrangements. Structural variants significantly impact gene function regulation, genome structural stability, and phenotypic formation; therefore, their accurate detection is crucial in genomics, biomedicine, and genetic breeding. With the development of long-read sequencing technologies, especially the application of High-Fidelity Long-Read Sequencing (HiFi), the detection capability of structural variants has improved. However, high-coverage sequencing is costly, limiting its application in large-scale samples and low-cost scenarios. While low-coverage HiFi sequencing offers a significant cost advantage, its recall, accuracy, and stability for structural variant detection under these conditions are not high. Currently, structural variant detection methods mainly include read alignment-based methods and assembly alignment-based methods. These two methods have different detection principles, are complementary, and are suitable for various adaptive scenarios. However, existing methods typically rely on sufficient read support or complete assembly results, which can easily lead to missed detections of structural variations in low-coverage HiFi data, resulting in low overall recall. Read alignment-based methods are more suitable for small- to medium-scale variations, while assembly-based methods are better suited for large-scale variations. Existing technologies struggle to balance detection performance across different scales of structural variations. Current integration methods are relatively simple, typically employing simple union, intersection, or fixed threshold merging of different detection results, making it difficult to fully utilize the complementary information between different methods. Furthermore, existing integration methods tend to introduce more false positives when improving detection rate, while reducing the detection capability of true structural variations when improving accuracy. Therefore, effectively integrating the detection advantages of both methods under low-coverage HiFi data conditions has become a pressing technical problem to be solved in this field. Summary of the Invention
[0003] To overcome the aforementioned deficiencies of the prior art, this invention provides an integrated method and system for detecting structural variations in low-coverage HiFi data, which can effectively improve the recall rate of structural variation detection, reduce false negatives, and maintain high detection stability under low-coverage HiFi data conditions.
[0004] The technical solution adopted in this invention is: an integrated method for detecting structural variations in low-coverage HiFi data, comprising the following steps: S1: The low-coverage HiFi reads were aligned with the reference genome using assembly alignment and read alignment methods respectively, to obtain assembly alignment detection results and read alignment detection results; S2: Standardize and process all test results using a unified format; S3: Classify structural variations in the normalized detection results based on the type and length of the variation. S4: Cluster and determine the consistency of structural variations, identify the SV clusters retained in the integration results, and integrate different types of structural variations into multiple structural variation results based on their respective detection results integration order; S5: Perform final consistency integration on multiple structural variation results to obtain the final structural variation result.
[0005] Preferably, in step S1, during assembly alignment, two assembly results are generated based on the reference-guided assembly method, and the assembly results are compared with the reference genome to identify standard structural variation types from the alignment results.
[0006] Preferred standard structural variation types include deletion (DEL), insertion (INS), duplication (DUP), inversion (INV), and translocation (TRA).
[0007] Preferably, reads within each chromosome bin are assembled separately to generate chromosome contigs, and the contigs of all chromosome bins are merged to generate the final genome assembly result.
[0008] Preferably, in step S1, during read alignment, a set of detection tools for read alignment is preset. After aligning the HiFi read to the reference genome, each detection tool in the set is used to perform SV calls to obtain a set of read alignment detection results.
[0009] Preferably, in S2, the unified format normalization process includes unified variant type marking (SV type representation), unified breakpoint coordinate representation (breakpoint coordinate definition), unified length calculation method (length definition), and removal of records with obvious format abnormalities.
[0010] Preferably, in S3, insertion and repetitive structural variants with a length ≥10kbp are classified into the large-scale insertion and repetitive structural variant category (large-scale INS / DUP category), and the remaining result variants are uniformly classified into other structural variant categories (other SVs category).
[0011] Preferably, in S4, structural variants from different tools and paradigms are clustered by breakpoints, and a maximum allowable breakpoint offset distance is defined. When any two structural variants are identical in chromosome and variant type and the breakpoint offset does not exceed the maximum allowable breakpoint offset distance, they are classified into the same SV cluster.
[0012] Furthermore, the number of supporting tools for each SV cluster is counted, and the minimum number of supporting tools threshold is used as the consistency criterion. For any SV cluster, if the number of supporting tools is greater than or equal to the minimum number of supporting tools threshold, it is retained as the integration result; if the number of supporting tools is less than the minimum number of supporting tools threshold, it is not retained.
[0013] Preferably, in step S4, the order of integrating the detection results is the order of integrating the assembly comparison detection results and integrating the read segment comparison detection results.
[0014] Preferably, in step S4, the following three complementary integration methods are used to integrate different categories of structural variations: (1) For large-scale insertion and repetitive structural variations, the assembly comparison detection is the main method. First, the large-scale insertion and repetitive structural variations detected by all read comparisons are integrated. Then, the integration results are integrated with the large-scale insertion and repetitive structural variations detected by the two assembly methods to ensure that the final structural variation results contain at least the structural consistency support of the assembly comparison detection. (2) For other structural variations, the reading segment comparison detection is the main method. First, the other structural variations obtained by the comparison detection of the two assembly methods are integrated. Then, the integration results are combined with the other structural variations obtained by the reading segment comparison detection. By applying a strict consistency threshold to the reading segment comparison detection results, the accuracy and breakpoint stability of small and medium scale structural variations are improved. (3) For other structural variations, the assembly comparison detection is the main method. First, the other structural variations obtained by all read segment comparison detection are integrated, and then the integration result is integrated with the other structural variations obtained by the comparison detection of the two assembly methods. This is used to supplement structural variations that are not supported in read segment comparison but have stable structural signals in assembly comparison.
[0015] The structural variation detection integration system for low-coverage HiFi data employs any of the structural variation detection integration methods for low-coverage HiFi data disclosed in this invention to detect and integrate structural variations, including: Structural variation detection module: This module is used to align low-coverage HiFi reads with the reference genome using assembly alignment and read alignment methods, respectively, to obtain assembly alignment detection results and read alignment detection results. Standardization module: Used to standardize all test results into a uniform format; Structural Variation Classification Module: Used to classify structural variations in the normalized detection results based on variation type and length; The structural variation result integration module is used to cluster and determine the consistency of structural variations, identify the SV clusters retained in the integration results, and integrate different categories of structural variations into multiple structural variation results based on their respective detection results integration order. Finally, it performs a final consistency integration of multiple structural variation results to obtain the final structural variation result.
[0016] Preferably, the structural variation result integration module includes: Large-scale insertion and repetitive structural variation result integration module: It is used to detect large-scale insertion and repetitive structural variations by focusing on assembly comparison. First, it integrates the large-scale insertion and repetitive structural variations detected by all read comparisons, and then performs a secondary integration with the large-scale insertion and repetitive structural variations detected by the two assembly methods. The first module for integrating other structural variation results is used to integrate other structural variations obtained by comparison and detection of two assembly methods, with the reading segment comparison detection as the main focus. Then, the integration result is combined with other structural variations obtained by comparison and detection of multiple reading segments. The second module is for integrating other structural variation results. It is mainly based on assembly comparison detection. It first integrates other structural variations obtained from the comparison detection of all read segments, and then integrates the integrated results with other structural variations obtained from the comparison detection of the two assembly methods. Final structural variation result integration module: This module integrates the structural variation results from the large-scale insertion and repetitive structural variation result integration module, the first other structural variation result integration module, and the second other structural variation result integration module to achieve final consistency and obtain the final structural variation result.
[0017] The beneficial effects of this invention are as follows: By fusing read alignment and reference-guided assembly alignment evidence, this invention can simultaneously improve the detection effect of structural variations at both small and large scales. It performs differentiated integration of multi-source detection results, unlike simple union, intersection, or fixed threshold merging. It can more effectively utilize the complementary information between read alignment-based detection methods and assembly alignment-based detection methods. Optimized for the characteristics of low-coverage HiFi data, it can achieve effective integrated detection of structural variations of different types and scales. Thus, it can improve the recall rate of structural variation detection, reduce false negatives, and maintain high detection stability under low-coverage HiFi data conditions. It effectively solves the problems of decreased sensitivity of structural variation detection and imbalance of detection capabilities for structural variations at different scales under low-coverage HiFi sequencing conditions. Attached Figure Description
[0018] Figure 1 This is a flowchart of the integrated method for detecting structural variations in low-coverage HiFi data according to the present invention. Detailed Implementation
[0019] See Figure 1 This invention discloses an integrated method for detecting structural variations in low-coverage HiFi data. It is a hierarchical integrated detection method that combines read segment alignment with reference guidance assembly evidence, comprising the following steps: S1: Detection of multi-source structural variations: Low-coverage HiFi reads were simultaneously aligned with the reference genome using both assembly alignment and read alignment methods to obtain assembly alignment detection results and read alignment detection results.
[0020] in, S101: Acquisition of assembly alignment and detection results: Let the set of reference-guided assembly methods be... Low-coverage HiFi data were assembled using the two assembly methods described above. The assembly results were then compared with a reference genome using a whole-genome alignment. Five types of standard structural variations were identified from the assembly-reference alignment results, denoted as follows: and The five standard structural variations include deletion (DEL), insertion (INS), duplication (DUP), inversion (INV), and translocation (TRA).
[0021] S102: Obtaining the segment comparison and detection results: Let the set of detection tools based on segment comparison be . ,in Indicates the first A detection tool, , This indicates the total number of tools. After aligning the HiFi reads to the reference genome, the sets were used respectively. The various detection tools in the dataset are called using SV to obtain a set of detection results. ,in Representation tools The output set of structural variations.
[0022] S2: Standardized processing of test results: All test results undergo standardized format processing before entering the integration process. Standardized format processing includes standardized variant type labeling (SV type representation), standardized breakpoint coordinate representation (breakpoint coordinate definition), standardized length calculation method (length definition), and removal of records with obvious format abnormalities.
[0023] S3: Structural variation classification (or stratification): Based on the type and length of the variation, the normalized detection results are classified into structural variations, and the sets of large-scale insertion and repetition structural variations are defined as follows: , in, This represents the set of all structural variations in the SV detection results; Represents a set A structural variation event; Indicates variation In a VCF record, the type field's value belongs to... The timestamps represent insertion and repetition, respectively. Indicates variation The length of structural variation given in the VCF file; This represents the length threshold used to distinguish large-scale structural variations, taken as... .
[0024] The remaining SVs are classified as other structural variations: , The above classifications are applied to the results of two types of assembly alignment detection and the results of SV detection based on read segment alignment.
[0025] S4: Low-Coverage HiFi Data Integration Structural variations are clustered and their consistency is determined to identify the SV clusters retained in the integration results. Structural variations of different categories are integrated into multiple structural variation results based on their respective detection results and integration order.
[0026] During the integration phase, SURVIVOR is used as a unified tool for SV clustering and consistency determination. Let the set of SVs to be integrated be... , Indicates the first The SV result set output by each detection tool This represents the total number of tools involved in the integration. Any two structural variations... and The conditions for being considered to belong to the same SV cluster are: they are consistent in chromosome and SV type if and only if the breakpoint location satisfies: , in, and These represent the starting breakpoint positions of the two, respectively. and These represent the termination breakpoint positions of the two, respectively. This indicates the maximum allowed breakpoint offset distance.
[0027] For any SV cluster The number of tools it supports is defined as: , in, This represents a structural variation cluster formed by clustering rules. This indicates the number of detection tools that support this cluster; This is the serial number of the testing tool. Indicates the first A set of structural variation results output by each detection tool; Represents a specific structural variation event; symbol Indicates "existence"; condition Indicates variation Belongs to tools The result set Indicates variation Classified into clusters middle. This indicates that all sets of results satisfying the condition "there exists at least one variant belonging to the cluster" are considered to be clusters. A set of tool numbers; symbols This represents the number of elements in the set, therefore This is the total number of tools that support this SV cluster; This represents the minimum number of supported tools threshold, when When this occurs, it indicates that the cluster has obtained no less than The tools supported by the tool are therefore retained as the final integration result; otherwise, they are not retained.
[0028] Based on the above definition, the following three complementary integration methods are used to integrate structural variations in the pen holder category: (1) For large-scale insertions and repetitive SVs, the assembly and comparison detection is the main approach. First, low-threshold integration is performed on all large-scale insertions and repetitive SVs detected by the read segment comparison. Then, the integration results are further integrated with the large-scale insertion and repetitive SVs obtained from the two assembly methods. This ensures that the final retained structural variation results contain at least one piece of evidence from assembly alignment detection; (2) For other SVs, the main focus is on segment comparison and detection. First, the other SVs obtained by comparison and detection of the two assembly methods are integrated. Then, the integration results are combined with other SVs obtained from multiple read segment comparisons and detections. (default value is 5) to improve the consistency and reliability of small and medium-scale SVs; (3) For other SVs, the assembly comparison and detection is the main focus. First, all other SVs obtained from the comparison and detection of all read segments are integrated. Then, the integration result is integrated with other SVs obtained by comparing and detecting the two assembly methods. ), used to supplement structural variations in reading segments that are unstable but structurally consistent during assembly and comparison.
[0029] S5: Final structural variation results integration: The three SV results obtained in S4 are eventually unified and integrated. ), to obtain the final structural variation result set If a structural variation meets the corresponding consistency criterion in any integration method in S4, it is included in the final structural variation result set.
[0030] By explicitly defining the dominant relationship of evidence in different integration methods, this invention enables balanced detection of large-scale and small-to-medium-scale structural variations under low-coverage HiFi sequencing conditions.
[0031] This invention also discloses a structural variation detection and integration system for low-coverage HiFi data, which uses any of the structural variation detection and integration methods for low-coverage HiFi data disclosed in this invention to detect and integrate structural variations, including: Structural variation detection module: This module is used to align low-coverage HiFi reads with the reference genome using assembly alignment and read alignment methods, respectively, to obtain assembly alignment detection results and read alignment detection results. Standardization module: Used to standardize all test results into a uniform format; Structural Variation Classification Module: Used to classify structural variations in the normalized detection results based on variation type and length; The structural variation result integration module is used to cluster and determine the consistency of structural variations, identify the SV clusters retained in the integration results, and integrate different categories of structural variations into multiple structural variation results based on their respective detection results integration order. Finally, it performs a final consistency integration of multiple structural variation results to obtain the final structural variation result.
[0032] The structural variation result integration module preferably includes: Large-scale insertion and repetitive structural variation result integration module: It is used to detect large-scale insertion and repetitive structural variations by focusing on assembly comparison. First, it integrates the large-scale insertion and repetitive structural variations detected by all read comparisons, and then performs a secondary integration with the large-scale insertion and repetitive structural variations detected by the two assembly methods. The first module for integrating other structural variation results is used to integrate other structural variations obtained by comparison and detection of two assembly methods, with the reading segment comparison detection as the main focus. Then, the integration result is combined with other structural variations obtained by comparison and detection of multiple reading segments. The second module is for integrating other structural variation results. It is mainly based on assembly comparison detection. It first integrates other structural variations obtained from the comparison detection of all read segments, and then integrates the integrated results with other structural variations obtained from the comparison detection of the two assembly methods. If a module exists Belonging to the above three modules, making Then structural variation It was included in the final results.
[0033] Final structural variation result integration module: This module integrates the structural variation results from the large-scale insertion and repetitive structural variation result integration module, the first other structural variation result integration module, and the second other structural variation result integration module to achieve final consistency and obtain the final structural variation result.
[0034] Unless otherwise specified or further limited to one preferred or optional technical means being another, the preferred and optional technical means disclosed in this invention can be arbitrarily combined to form several different technical solutions.
Claims
1. An integrated method for structural variant calling of low-coverage HiFi data, characterized in that Includes the following steps: S1: The low-coverage HiFi reads were aligned with the reference genome using assembly alignment and read alignment methods respectively, to obtain assembly alignment detection results and read alignment detection results; S2: Standardize and process all test results using a unified format; S3: Classify structural variations in the normalized detection results based on the type and length of the variation. S4: Cluster and determine the consistency of structural variations, identify the SV clusters retained in the integration results, and integrate different types of structural variations into multiple structural variation results based on their respective detection results integration order; S5: Perform final consistency integration on multiple structural variation results to obtain the final structural variation result.
2. The integrated method for detecting structural variations in low-coverage HiFi data according to claim 1, characterized in that... In S1, during assembly alignment, two assembly results are generated based on the reference-guided assembly method. The assembly results are then compared with the reference genome to identify standard structural variant types from the alignment results.
3. The integrated method for detecting structural variations in low-coverage HiFi data according to claim 2, characterized in that... Reads within each chromosome bin are assembled individually to generate chromosome contigs. The contigs of all chromosome bins are then merged to generate the final genome assembly.
4. The integrated method for detecting structural variations in low-coverage HiFi data according to claim 1, characterized in that... In S1, during read alignment, a set of detection tools for read alignment is preset. After the HiFi read is aligned to the reference genome, each detection tool in the set of detection tools is used to perform SV calls to obtain a set of read alignment detection results.
5. The integrated method for detecting structural variations in low-coverage HiFi data according to claim 1, characterized in that... In S2, the unified format normalization process includes unified mutation type marking, unified breakpoint coordinate representation, unified length calculation method, and removal of records with obvious format abnormalities.
6. The integrated method for detecting structural variations in low-coverage HiFi data according to claim 1, characterized in that... In S3, insertion and repetition structural variants with a length ≥10kbp are classified as large-scale insertion and repetition structural variants, while other result variants are uniformly classified as other structural variants.
7. The integrated method for detecting structural variations in low-coverage HiFi data according to claim 1, characterized in that... In S4, structural variants from different tools and paradigms are clustered by breakpoints. The maximum allowable breakpoint offset distance is defined. When any two structural variants are identical in chromosome and variant type and the breakpoint offset does not exceed the maximum allowable breakpoint offset distance, they are classified into the same SV cluster.
8. The integrated method for detecting structural variations in low-coverage HiFi data according to claim 7, characterized in that... The number of supporting tools for each SV cluster is counted, and the minimum number of supporting tools is used as the consistency criterion. For any SV cluster, if the number of supporting tools is greater than or equal to the minimum number of supporting tools, it is retained as the integration result; if the number of supporting tools is less than the minimum number of supporting tools, it is not retained.
9. The integrated method for detecting structural variations in low-coverage HiFi data according to claim 1, characterized in that... In S4, the order of integrating the detection results is the order of integrating the assembly comparison detection results and integrating the read segment comparison detection results.
10. An integrated system for detecting structural variations in low-coverage HiFi data, characterized in that... The structural variation detection and integration method for low-coverage HiFi data according to any one of claims 1-9 is used to detect and integrate structural variations, including: Structural variation detection module: This module is used to align low-coverage HiFi reads with the reference genome using assembly alignment and read alignment methods, respectively, to obtain assembly alignment detection results and read alignment detection results. Standardization module: Used to standardize all test results into a uniform format; Structural Variation Classification Module: Used to classify structural variations in the normalized detection results based on variation type and length; The structural variation result integration module is used to cluster and determine the consistency of structural variations, identify the SV clusters retained in the integration results, and integrate different categories of structural variations into multiple structural variation results based on their respective detection results integration order. Finally, it performs a final consistency integration of multiple structural variation results to obtain the final structural variation result.