A method and system for constructing a robust benchmark set of structural variations of the human genome

CN122598778BActive Publication Date: 2026-09-11YANTAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611079505.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-11
Estimated Expiration
2046-07-21

AI Technical Summary

Technical Problem

另外,现有的基准集构建方法也存在一些局限性,因为它没有考虑碎片化变异,即多个小变异实际对应同一个大变异的情况,这可能导致构建的基准集中出现因表示形式差异导致的冗余问题

Benefits of technology

本发明提供的一种人类基因组结构变异稳健基准集的构建方法,首先通过多个变异检测工具对测序数据并行处理,并对每个结构变异数据集进行标准化预处理,显著提升数据处理的规范性与可比性,为后续融合奠定了低噪声、高覆盖的数据基础;然后,采用双向验证匹配策略,在两两数据集之间精准捕捉变异匹配关系,并针对等位基因标记的变异组执行大小比过滤与同聚物剔除,从而有效保留了真实的等位多样性,大幅提高了基准集的完整性和碎片化变异的识别准确率;接着,利用多个独立修正变异数据集对初步基准集进行评测,提取共有假阳性予以补充、共有假阴性予以剔除,实现基准集的自纠偏与质量闭环优化,得到修正基准集,显著降低假阳性和假阴性率;最后,将来自不同数据或不同检测方法的多个修正基准集进行合并,充分融合各类技术路线的互补优势,最终得到覆盖更全面、稳定性更强、代表性更均衡的高质量稳健基准集,为结构变异检测工具的精准评估提供了可靠的“金标准”。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598778B_ABST
    Figure CN122598778B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of variant benchmark set construction in bioinformatics, in particular to a method and system for constructing a human genome structural variant robust benchmark set; the method of the present application firstly processes sequencing data in parallel through multiple variant detection tools, adopts a bidirectional verification matching strategy to accurately capture the matching relationship of variants between each pair of data sets, and performs size ratio filtering and homopolymer removal on the variant group marked with alleles, combines with the optimal variant to form a preliminary benchmark set; then, the preliminary benchmark set is evaluated by using multiple independent revised variant data sets, common false positives are extracted for supplementation and common false negatives are extracted for removal, thereby significantly reducing the false positive and false negative rates; finally, the revised benchmark sets obtained from multiple different data sources and / or different structural variant detection methods are combined to obtain a structural variant robust benchmark set, which provides a reliable "gold standard" for accurate evaluation of variant detection tools.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics variant benchmark set construction technology, specifically a method and system for constructing a robust benchmark set of human genome structural variants. Background Technology

[0002] State-of-the-art sequencing technologies have enabled the study of challenging regions of the human genome. Genomic structural variants (SVs) can be detected with increasing accuracy, resolution, and comprehensiveness, expanding the scope of variant benchmark datasets, which are widely used in routine research and clinical practice. However, missing variants and potential false predictions affect almost all published benchmark datasets. For example, the GIAB consortium developed a benchmark dataset based on widely agreed and available Afro-Jewish pedigrees (HG002 / NA24385) from the Personal Genome Project. Comparing this benchmark dataset with VCF files (user sets) generated by variant detection tools revealed a large number of false positives and false negatives. Further examination and verification using the igv tool showed that the false negatives present in the benchmark dataset were errors in the dataset itself, and many false positives in the user set were genuine variants that did not exist in the benchmark dataset. Furthermore, existing benchmark dataset construction methods have limitations because they do not consider fragmented variants—that is, multiple small variants actually corresponding to the same large variant—which may lead to redundancy issues in the constructed benchmark dataset due to differences in representation. Secondly, the loss of some alleles means that the benchmark set only captures a portion of the variants, which directly weakens the "gold standard" value of the benchmark set. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for constructing a robust benchmark set of human genome structural variations.

[0004] The technical solution of this invention is as follows: A method for constructing a robust benchmark set of human genome structural variations includes the following operations: S1. Sequencing data is processed by multiple variant detection tools to obtain multiple structural variant datasets. Each structural variant dataset is preprocessed, including adding a unique index identifier to each variant, dividing the data into subsets according to chromosomes, and determining the search region for each variant based on variant type and breakpoint location, resulting in multiple preprocessed variant datasets. S2. Perform bidirectional validation matching on each pair of multiple preprocessed variant datasets, summarize all comprehensive matching relationships to obtain a comprehensive matching relationship set; in the comprehensive matching relationship set, extract the variants that are supported by a preset number of detection tools, and the corresponding matching variants to form supporting variant matching groups; for supporting variant matching groups of variants with allele markers, perform variant size ratio filtering and identical base filtering, and store all remaining variants in the allele array; obtain the optimal variants corresponding to the supporting variant matching groups of all variants without allele markers, form the optimal variant array, add the allele array, and obtain the preliminary benchmark set; S3. Evaluate the preliminary baseline set using multiple corrected variant datasets to obtain a false negative set and a false positive set. Add the common false positives in the false positive set to the preliminary baseline set and remove the common false negatives in the false negative set from the preliminary baseline set to obtain the corrected baseline set. S4. Merge the modified benchmark sets obtained from multiple different data sources and / or different structural variation detection methods to obtain a structural variation robust benchmark set.

[0005] In S2, the bidirectional validation matching operation is as follows: Two preprocessed mutation datasets are used as the user set and the baseline set, respectively. When the starting position of the mutation in the baseline set is within the search region of the mutation in the user set, it is determined whether the mutations overlap. If so, bidirectional local joint analysis and component matching analysis are performed. If component matching analysis fails, regular matching judgment is performed to obtain the current round's matching result, which is stored in the previous round's matching array to obtain the current round's matching array. If not, component matching analysis is performed based on the previous round's matching array to obtain the current round's matching result, which is stored in the previous round's matching array to obtain the current round's matching array. This process continues until all mutations in the user set and the baseline set have undergone bidirectional validation matching, at which point the final round's matching array is output, yielding the comprehensive matching relationship.

[0006] Bidirectional local joint analysis includes local joint analysis from the user set to the reference set and local joint analysis from the reference set to the user set. The operation of local joint analysis from the user set to the reference set is as follows: based on the variant type and the proximity of the breakpoint, variants of the same type in the neighborhood are classified into a candidate variant set; the variants in the candidate variant set are combined, and a first sequence consistency assessment is performed. If the sequence consistency assessment conditions are met, a positive matching relationship is obtained. Local joint analysis from the reference set to the user set adopts the same operation steps as local joint analysis from the user set to the reference set, only the roles of the analysis objects of the user set and the reference set are interchanged; the results of local joint analysis from the reference set to the user set are obtained as a reverse matching relationship, which, together with the positive matching relationship, forms a successful bidirectional local joint analysis result.

[0007] In the operation of classifying neighboring variants of the same type into the candidate variant set, the first and second candidate sets are initialized. For insertion variants, starting from the current insertion variant, all insertion variants of the same type are traversed backward and added to the first candidate set in turn, until the difference between the starting position of the insertion variant and the starting position of the first insertion variant in the first candidate set is greater than the first distance threshold, thus obtaining the insertion candidate variant set. For missing variants, the ending position of the current missing variant is recorded, and the same type of missing variants are traversed backward. If the difference between the starting position of the missing variant to be added to the second candidate set and the recorded ending position is not greater than the preset second distance threshold, the corresponding missing variant is added to the second candidate set, and the ending position record value is updated with the ending position of the corresponding missing variant. This process is repeated until the condition is no longer met, thus obtaining the missing candidate variant set. The insertion candidate variant set and the missing candidate variant set together form the candidate variant set.

[0008] The variants in the candidate variant set are combined in order of breakpoint position. If the original alignment BAM file is not provided, all variant combinations are enumerated, and variant combinations whose size ratio with the reference set is less than the size ratio threshold are filtered out. If the original alignment BAM file is provided, the original variant signal of the read segment corresponding to each variant combination is scanned, and variant combinations not on the same read segment are filtered out by haplotype partitioning. Combinations whose size ratio is less than the size ratio threshold are also filtered out. Combinations with overlapping original variant coordinates are filtered out, and the combination with the closest size to the variant in the reference set is selected from the remaining combinations as the optimal merging candidate for sequence consistency evaluation.

[0009] In the first sequence consistency assessment process, for the optimal merge candidate obtained after combination, the upstream and downstream reference sequences of the coverage area and the coverage area of ​​the variant in the baseline set are extracted to construct the context sequence and calculate the sequence consistency. If the sequence consistency is greater than or equal to the sequence consistency threshold, all variants in the optimal merge candidate are determined to be locally joint positive, the corresponding variants in the baseline set are marked as true positives, all variant indices in the combination are returned, and the bidirectional index relationship is recorded in the two-dimensional array storing the variant matching relationship to obtain the positive matching relationship.

[0010] The optimal mutation in S2 is selected based on mutation length and overlap.

[0011] A system for constructing a robust benchmark set of human genome structural variations, used to implement the above-mentioned method for constructing a robust benchmark set of human genome structural variations, includes: The preprocessed variant dataset generation module is used to process sequencing data through multiple variant detection tools to obtain multiple structural variant datasets. Each structural variant dataset is preprocessed, including adding a unique index identifier to each variant, dividing the data into subsets by chromosome, and determining the search region for each variant based on variant type and breakpoint location, resulting in multiple preprocessed variant datasets. The preliminary benchmark set generation module is used to perform bidirectional validation matching on multiple preprocessed variant datasets, summarize all comprehensive matching relationships, and obtain a comprehensive matching relationship set. In the comprehensive matching relationship set, variants supported by a preset number of detection tools and their corresponding matching variants are extracted to form supporting variant matching groups. For supporting variant matching groups of variants with allele markers, variant size ratio filtering and identical base filtering are performed, and all remaining variants are stored in the allele array. The optimal variants corresponding to the supporting variant matching groups of all variants without allele markers are obtained to form the optimal variant array, and the allele array is added to obtain the preliminary benchmark set. The modified baseline set generation module is used to evaluate the preliminary baseline set using multiple modified variant datasets to obtain a false negative set and a false positive set. Common false positives in the false positive set are added to the preliminary baseline set, and common false negatives in the false negative set are removed from the preliminary baseline set to obtain the modified baseline set. The structural variation robust benchmark set generation module is used to merge multiple modified benchmark sets obtained from different data sources and / or different structural variation detection methods to obtain a structural variation robust benchmark set.

[0012] An apparatus for constructing a robust benchmark set of human genome structural variations includes a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the above-described method for constructing a robust benchmark set of human genome structural variations.

[0013] A computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the above-described method for constructing a robust benchmark set of human genome structural variations.

[0014] The beneficial effects of this invention are as follows: This invention provides a method for constructing a robust benchmark set for human genome structural variations. First, sequencing data is processed in parallel using multiple variation detection tools, and each structural variation dataset undergoes standardized preprocessing, significantly improving the standardization and comparability of data processing and laying a low-noise, high-coverage data foundation for subsequent fusion. Then, a bidirectional validation matching strategy is employed to accurately capture variation matching relationships between pairs of datasets, and size ratio filtering and homopolymer removal are performed on allele marker variant groups, effectively preserving true allelic diversity and significantly improving the integrity of the benchmark set and the accuracy of identifying fragmented variations. Next, the preliminary benchmark set is evaluated using multiple independent corrected variation datasets, extracting shared false positives for supplementation and removing shared false negatives, achieving self-correction and quality closed-loop optimization of the benchmark set to obtain a corrected benchmark set, significantly reducing false positive and false negative rates. Finally, multiple corrected benchmark sets from different data or different detection methods are merged, fully integrating the complementary advantages of various technical approaches, ultimately obtaining a high-quality robust benchmark set with more comprehensive coverage, stronger stability, and more balanced representativeness, providing a reliable "gold standard" for the accurate evaluation of structural variation detection tools.

[0015] This invention provides a method for constructing a robust benchmark set of human genome structural variations, which overcomes the limitations of existing benchmark data resources and construction methods, including the lack of high-quality benchmark sets, failure to consider fragmented variations, and loss of some alleles. Through technological innovation, this invention improves the quality of the variation benchmark dataset, providing a reliable reference standard for performance evaluation and algorithm optimization of SV detection tools. This will contribute to research in clinical medicine, genetics, and biomedicine. Experimental verification shows that this invention's method exhibits significant advantages and is expected to have a positive impact on clinical medicine, genetics, and biomedicine. Attached Figure Description

[0016] The solutions and advantages of this application will become clear to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of the invention.

[0017] In the attached diagram: Figure 1 This is a schematic diagram illustrating the selection of the optimal mutation in this embodiment. Figure 2 This is a schematic diagram illustrating the correction of the initial reference set in this embodiment. Figure 3 This is a schematic diagram illustrating the merging of multiple correction reference sets in this embodiment. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the exemplary embodiments of this application clearer, the technical solutions in the exemplary embodiments of this application are described clearly and completely below. Obviously, the described exemplary embodiments are only some embodiments of this application, and not all embodiments.

[0019] This embodiment provides a method for constructing a robust benchmark set of human genome structural variations, including the following operations: S1. Sequencing data is processed by multiple variant detection tools to obtain multiple structural variant datasets. Each structural variant dataset is preprocessed, including adding a unique index identifier to each variant, dividing the data into subsets according to chromosomes, and determining the search region for each variant based on variant type and breakpoint location, resulting in multiple preprocessed variant datasets. S2. Perform bidirectional validation matching on each pair of multiple preprocessed variant datasets, summarize all comprehensive matching relationships to obtain a comprehensive matching relationship set; in the comprehensive matching relationship set, extract the variants that are supported by a preset number of detection tools, and the corresponding matching variants to form supporting variant matching groups; for supporting variant matching groups of variants with allele markers, perform variant size ratio filtering and identical base filtering, and store all remaining variants in the allele array; obtain the optimal variants corresponding to the supporting variant matching groups of all variants without allele markers, form the optimal variant array, add the allele array, and obtain the preliminary benchmark set; S3. Evaluate the preliminary baseline set using multiple corrected variant datasets to obtain a false negative set and a false positive set. Add the common false positives in the false positive set to the preliminary baseline set and remove the common false negatives in the false negative set from the preliminary baseline set to obtain the corrected baseline set. S4. Merge the modified benchmark sets obtained from multiple different data sources and / or different structural variation detection methods to obtain a structural variation robust benchmark set; The specific steps are detailed below.

[0020] S1. Sequencing data is processed by multiple variant detection tools to obtain multiple structural variant datasets. Each structural variant dataset is preprocessed, including adding a unique index identifier to each variant, dividing the data into subsets according to chromosomes, and determining the search region for each variant based on variant type and breakpoint location, resulting in multiple preprocessed variant datasets.

[0021] First, the sequencing data is processed using multiple variant detection tools to obtain multiple structural variant datasets. Specifically, for PacBio CCS data or ONT ultra-long read data, read alignment-based methods are used, including but not limited to using minimap2 to align the data to the reference genome and generate corresponding BAM files. These BAM files are then used as input, and read alignment-based variant detection tools such as ASVCLR, and / or cuteSV, and / or DeBreak, and / or kled, and / or Sniffles2, and / or SVDSS are used to independently call up variants and output VCF format variant datasets, resulting in multiple structural variant datasets. Alternatively, for PacBio CCS data or HiC data, de novo assembly methods are used, including but not limited to using hifiasm to complete de novo genome assembly. After aligning the assembled contigs with the reference genome, structural variant detection tools such as svim-asm and paftools.js are used to generate VCF format variant datasets, resulting in multiple structural variant datasets.

[0022] Then, each structural variation dataset is preprocessed, including adding a unique index identifier to each variation, partitioning the data into subsets by chromosome, and determining the search region for each variation based on variation type and breakpoint location, resulting in multiple preprocessed variation datasets. Specifically, the multiple input structural variation datasets are traversed, and a globally unique index identifier is assigned to all variations in each dataset to accurately locate each variation in subsequent matching processes. Simultaneously, a two-dimensional array is initialized to store the matching index for each variation, i.e., which variations match this variation. The row size of the two-dimensional array is the number of all variations, and the column size is the number of matching indices. The multiple input structural variation datasets are further divided into smaller datasets by chromosome and stored in a three-dimensional array. The depth of the three-dimensional array is the number of chromosomes, the row size is the number of input variation datasets, and the column size is variable length, corresponding to the number of variations on a specific chromosome in a given variation dataset. This allows for multi-threading to speed up processing based on the depth. Next, the search region for each variation during matching is determined, where the starting position of the variation is assumed to be... start_pos The end position is end_pos The size of the mutation is svlen The search region varies depending on the mutation type. When the mutation type is insertion (INS), the search region range is as follows: , When the mutation type is any type other than INS, the search area is as follows: , When the starting position of a mutation is not within the search area, the mutation does not need to proceed to the next matching step.

[0023] S2. Perform bidirectional validation matching on each pair of multiple preprocessed variant datasets, summarize all comprehensive matching relationships to obtain a comprehensive matching relationship set; in the comprehensive matching relationship set, extract the variants that are supported by the number of preset detection tools, and the corresponding matching variants to form supporting variant matching groups; for supporting variant matching groups of variants with allele markers, perform variant size ratio filtering and identical base filtering, and store all remaining variants in the allele array; obtain the optimal variants corresponding to the supporting variant matching groups of all variants without allele markers, form the optimal variant array, add the allele array, and obtain the preliminary benchmark set.

[0024] First, multiple preprocessed mutation datasets are paired and bidirectionally validated and matched. All matching relationships are then summarized to obtain a set of matching relationships.

[0025] The bidirectional validation matching operation is as follows: Two preprocessed mutation datasets to be matched are designated as the user set and the baseline set, respectively. When the starting position of the mutation in the baseline set is within the search region of the mutation in the user set, it is determined whether the mutations overlap. If they do, bidirectional local joint analysis and component matching analysis are performed. If component matching analysis fails, a regular matching judgment is performed to obtain the current round's matching result, which is then stored in the previous round's matching array to obtain the current round's matching array. If they do not overlap, component matching analysis is performed based on the previous round's matching array to obtain the current round's matching result, which is then stored in the previous round's matching array to obtain the current round's matching array. For example, if a is a 100bp del mutation, b is a 98bp del mutation, and c is a 99bp del mutation, when determining whether mutation b and mutation c overlap, it is found that they do not overlap. Based on the previous matching array obtained from the conventional matching judgment, it is known that variant a matches variant b and variant c. When variant b and variant c perform component matching analysis, variant a becomes the bridge between variant b and variant c. b and c successfully perform component matching and the matching relationship between b and c is saved. This continues until all variants in the user set and the benchmark set have completed bidirectional verification matching, and the last round of matching array is output to obtain the comprehensive matching relationship.

[0026] Bidirectional local joint analysis includes local joint analysis from the user set to the reference set, and local joint analysis after the reference set and user set are swapped. Local joint analysis from the reference set to the user set adopts the same operation steps as local joint analysis from the user set to the reference set, only the analysis object roles of the user set and the reference set are interchanged. It is only necessary to treat the user set as the "reference set" and the reference set as the "user set". Through bidirectional verification of forward matching and reverse matching, it can handle the fragmented situation where multiple small mutations correspond to a large mutation. It significantly alleviates the limitation of traditional mutation matching methods that only consider one-to-one matching and ignore many-to-one fragmented mutation correspondence, and improves the robustness of matching between mutations.

[0027] Taking the local joint analysis from the user set to the reference set as an example, the operation of the local joint analysis is as follows: based on the mutation type and the proximity of the breakpoint, the mutations of the same type in the neighborhood are classified into the candidate mutation set; the mutations in the candidate mutation set are combined to perform the first sequence consistency evaluation. If the sequence consistency evaluation conditions are met, a positive matching relationship is obtained; if the sequence consistency evaluation conditions are not met, the positive matching relationship is returned as empty.

[0028] Performing a local joint analysis from the baseline set to the user set, and using the analysis results as a reverse matching relationship, the operation for performing a local joint analysis from the baseline set to the user set is the same as above, and will not be repeated here for the sake of brevity. The forward and reverse matching relationships form a bidirectional local joint successful analysis result, which is used to perform component matching analysis.

[0029] In the operation of classifying neighboring variants of the same type into the candidate variant set, the first and second candidate sets are initialized. For insertion variants, starting from the current insertion variant, all similar insertion variants are traversed backward and added to the first candidate set sequentially until the difference between the starting position of the insertion variant and the starting position of the first insertion variant in the first candidate set is greater than a first distance threshold, preferably 500, thus obtaining the insertion candidate variant set. For missing variants, the endPos_update variable is used to record the end position of the current missing variant. Similar missing variants are traversed backward. If the difference between the starting position of the missing variant to be added to the second candidate set and the recorded end position is not greater than a preset second distance threshold, preferably 400, then the corresponding missing variant is added to the second candidate set, and the end position record value is updated with the end position of the corresponding missing variant. This process is repeated until the condition is no longer met, thus obtaining the missing candidate variant set. The insertion candidate variant set and the missing candidate variant set together form the candidate variant set.

[0030] When combining variants in the candidate variant set, the variants in the candidate variant set are combined in order of breakpoint position. If the original alignment BAM file is not provided, all variant combinations are enumerated, and variant combinations with a size ratio of less than the first size ratio threshold of the variant in the baseline set are filtered out. The first size ratio threshold of the variant is preferably 0.9. If the original alignment BAM file is provided, the original variant signal of the read corresponding to each variant combination is scanned. Variants not on the same read are filtered out by haplotype division, and combinations with a size ratio of less than the first size ratio threshold are filtered out. Combinations with overlapping original variant coordinates are further filtered out to reduce the risk of erroneous merging of multi-allelic variants. Furthermore, the combination with the size closest to the variant in the baseline set is selected from the remaining combinations as the optimal merging candidate.

[0031] In the first sequence consistency assessment, for the optimal merge candidate, upstream and downstream reference sequences are extracted from the coverage area and the variant coverage area of ​​the baseline set to construct a shared context sequence, and sequence consistency is calculated. If the sequence consistency is greater than or equal to the sequence consistency threshold (preferably 0.7), then all variants in the optimal merge candidate are determined to be locally combined positive, denoted as... LP_user This is used to indicate all smaller variants in the optimal matching cluster of the user set, with the corresponding baseline variants marked as true positives, denoted as TP. It also returns the indices of all variants in this combination, and stores the indices of all variants in this combination at the positions of the baseline variant indices in the two-dimensional array storing the variant matching relationships. This achieves bidirectional indexing in the two-dimensional array storing variant matching relationships, indicating that all smaller variants in the variant combination match with the larger variants in the corresponding baseline set, resulting in a positive matching relationship. If sequence consistency is less than the sequence consistency threshold, the positive local joint analysis fails, and all variants in the optimal merge candidate are marked as false positives, denoted as TP. FP The corresponding baseline set variation is marked as a false negative, denoted as . FN The returned mutation combinations are empty, meaning the returned positive matching relationships are empty.

[0032] The successful component matching analysis operation is as follows: In the bidirectional local joint successful analysis result and the previous round of matching array, or in the previous round of matching array, if a first mutation and a third mutation exist, indirect matching is performed through a second mutation. In this case, the first and third mutations also match, forming a component. The corresponding mutation indices are recorded, and the component matching analysis result is obtained and stored in the previous round of matching array as the current round of matching array. Specifically, when determining whether two mutations match, by default only the two mutations themselves are considered. However, if one mutation matches both of the two mutations to be matched, then no matching judgment is needed; the two mutations to be matched can be directly considered a match. For example, when mutation a and mutation c do not meet the matching condition, mutation b is introduced. a and b match successfully, and b and c also match successfully. At this time, b can act as a "bridge" between a and c, forming the abc component. Mutations within the component are determined to be the same mutation that matches each other. Using pairs...<int,int> Storing the indices of matched mutations is similar to storing the edge relationships between two mutations, and using a set to store all pairs containing a given mutation index is similar to storing all edge relationships for that mutation. Therefore, determining whether two mutations match can also be done by checking if they are within the same component. If they are, it indicates that there is an edge relationship between a certain mutation and both of them, which avoids performing a matching operation. Sometimes, it can even successfully match two mutations that would not match under normal matching operations, improving the robustness of mutation matching.

[0033] If component matching analysis fails, a regular matching judgment is executed. This is used to handle direct matching between single variants in the bidirectional local joint successful analysis results. In the regular matching judgment, if two variants satisfy the judgment conditions corresponding to the first adaptive variant size ratio judgment, type judgment, and second sequence consistency assessment, then the match is successful, and the current round's matching result is obtained and stored in the previous round's matching array. After a successful match, the index of the successfully matched variant is stored in the corresponding variant index position in the two-dimensional array storing variant matching information. Additionally, a pair indicating the edge relationship between the two variants needs to be stored in the corresponding position in the set collection to obtain the regular matching judgment result.

[0034] For the first adaptive mutation size ratio judgment, let the lengths of the two mutations be respectively... and Take the larger value First adaptive size ratio threshold Segmented dynamic setting is adopted: Calculate the actual size ratio If the actual size ratio If the actual size is greater than 1, then the match is successful. If not, the match will fail.

[0035] For type determination, if two mutations have the same mutation type, the match is successful; otherwise, the match fails. A relatively lenient strategy is adopted here, assuming that mutations of insertion and repetition types can be determined through type determination.

[0036] For the second sequence consistency assessment, the sequence consistency of the two variants is calculated. If the sequence consistency is greater than or equal to the sequence consistency threshold, the match is successful; if the sequence consistency is less than the sequence consistency threshold, the match fails.

[0037] After a user set mutation completes its matching operation with all reference set mutations in the current search area, it continues to process the next user set mutation, re-divides its search area, and repeats this process until all user set mutations have been processed.

[0038] Then, in the comprehensive matching set, the variants supported by the preset number of detection tools (i.e., variants that can be detected by the preset number of detection tools) and their corresponding matching variants are extracted to form a supporting variant matching group. Since multiple structural variant datasets come from different variant detection tools, and the final target dataset has at least n methods supporting the variants, the number of preset tools n can be selected according to actual needs. In particular, when n=1, that is, the variants in the constructed dataset have at least 1 method supporting them, its function is equivalent to performing a merging operation of multiple variant datasets. The two-dimensional array storing variant matching information is traversed, and a set is used to store the variant index to indicate that the corresponding variant has been processed. When n>1, it is determined whether the variant has n methods supporting it. If not, the next variant is judged; if the condition is met, it is checked whether the index corresponding to the variant is in the set. If it is in the set, the next variant is judged; otherwise, the variant index matching the variant and the variant index matching these indices are stored in the set, and further operations are performed on the variant index matching the variant and the variant index matching these indices.

[0039] Next, for the supporting variant matching groups of variants with allele markers, variant size ratio filtering and identical base filtering are performed, and all remaining variants are stored in the allele array. Allele markers are generated during the routine matching process. When two variants are of the same type, if the difference between the start position (start_pos) of the variant in the user set and the previous or next variant is within 100 bp, an allele marker is added to the variant in the user set. Simultaneously, the same operation is performed on the variant in the reference set; if the difference between the start position (start_pos) of the variant and the previous or next variant is within 100 bp, an allele marker is added to the variant in the reference set.

[0040] The variant size ratio filtering operation can be achieved by deleting two variants whose size ratio is greater than the corresponding second adaptive size ratio threshold. Identical base filtering can be achieved by deleting two variants whose sequences contain only the same base. Specifically, to retain alleles, variants with allele markers and their matching variants are first stored in a one-dimensional array. Then, based on the index of each variant, its dataset is determined, and these variants are stored in a two-dimensional array with a row size equal to the number of all variant datasets. nThe column size is the number of variant indices from the corresponding dataset. Next, two variants with a size ratio greater than the corresponding second adaptive size ratio threshold are filtered out, as are variants whose sequences are all the same base (e.g., both sequences are all A bases). After filtering, the variants from each dataset are sequentially examined in the two-dimensional array until an allele is found in a variant from a dataset and the number of alleles is greater than 1. This requires that if a variant dataset is obtained from an allele-sensitive tool, its input order should be considered before inputting the dataset. If the number of alleles obtained is greater than 1, then optimal variant selection is not necessary; the next variant is processed directly, and the alleles retained in this step are stored in the allele array.

[0041] Second adaptive size ratio threshold The calculation formula is as follows: , The lengths of the two variants are respectively and Take the larger value .

[0042] Subsequently, the optimal variants corresponding to the supporting variant matching groups of all variants without allele markers are obtained, forming the optimal variant array. The optimal variants are selected based on variant length and overlap; see [link to relevant documentation]. Figure 1 The process begins by obtaining the maximum length (max_len) of all variants within the currently supported variant matching group, filtering out variants with a length < 0.7 × max_len. A length difference threshold is dynamically adjusted to identify the variant group with the largest number of variants whose length differences are within the threshold range, and these are stored in a candidate optimal variant array. If the candidate optimal variant array has a size of 1, the variant corresponding to the variant index is the selected optimal variant. If the candidate optimal variant array has a size other than 1, the total overlap between each variant and other variants is calculated, and the variant with the highest total overlap is selected as the optimal variant. If multiple variants have the same total overlap, the longest variant among them is selected as the optimal representative variant. After selection, the next variant is processed, and the optimal variant selected in this step is stored in the optimal variant array.

[0043] Finally, all variants in the allele array are added to the optimal variant array and sorted to form a preliminary benchmark set.

[0044] S3. Evaluate the preliminary baseline set using multiple corrected variant datasets to obtain a false negative set and a false positive set. Add the common false positives from the false positive set to the preliminary baseline set, and remove the common false negatives from the false negative set in the preliminary baseline set to obtain the corrected baseline set.

[0045] To obtain a higher quality benchmark set, the initial benchmark set needs to be corrected. The correction process is as follows: Figure 2 As shown. First, it is proposed that there are a total of FN and shared FP The concept is that when multiple corrected mutation datasets are input and evaluated against a baseline set, multiple results will be obtained. FN Dataset (false negative set) and FP The dataset (false positive set) contains a total of FN In q indivual FN The variations present in the dataset total FP In q’ (2<= q’ <= q )indivual FP Variations present in the dataset q’ The quantity can be entered according to requirements. This is achieved by sharing the total... FN (False negatives) are removed from the revised baseline set, and the common ones are removed. FP (False positives) are added to the corrected baseline set, thus correcting the baseline set. The corrected variant dataset is different from the structural variant dataset; it is an external dataset. The evaluation operation described above is preferably implemented through bidirectional local joint analysis in S2's bidirectional validation matching, which can consider the matching of multiple small variants with a large variant, improving efficiency compared to traditional evaluation methods. FN and FP The quality of the dataset.

[0046] S4. Merge the modified benchmark sets obtained from multiple different data sources and / or different structural variation detection methods to obtain a structural variation robust benchmark set.

[0047] Using multiple different datasets, such as PacBio CCS datasets, ONT ultra-long read datasets, HiC datasets, and / or different structural variant detection methods, such as read alignment-based methods and genome assembly methods, to obtain multiple high-quality benchmark sets, a merging operation can be performed to obtain a structural variant robust benchmark set with a larger number of variants. The merging operation flowchart is shown below. Figure 3 As shown.

[0048] To verify the effectiveness of the method in this embodiment, the following experiment was conducted.

[0049] To verify the superiority of the method in this embodiment, a comparative experiment was conducted between the benchmark set obtained by the method in this embodiment and the mainstream benchmark set HG002_SVs_Tier1_v0.6. This experiment demonstrates the performance of different SVs detection methods on the two benchmark sets (minimap2 was selected as the alignment tool, and GRCh37 was used as the human reference genome). The experimental results are shown in Tables 1 and 2. It can be found that the benchmark constructed by the method in this embodiment can significantly improve the recall, precision, and F1 score when used for testing.

[0050] Table 1 Performance benchmark test of structural variation detection method (HG002_SVs_Tier1_v0.6 benchmark set) .

[0051] Table 2 Performance benchmark test of structural variation detection method (benchmark set obtained by the method in this embodiment: asvbm_create_bench) .

[0052] This embodiment provides a method for constructing a robust benchmark set for human genome structural variations. First, sequencing data is processed in parallel using multiple variation detection tools, and each structural variation dataset undergoes standardized preprocessing, significantly improving the standardization and comparability of data processing and laying a low-noise, high-coverage data foundation for subsequent fusion. Then, a bidirectional validation matching strategy is employed to accurately capture variation matching relationships between pairs of datasets, and size ratio filtering and homopolymer removal are performed on allele marker variant groups, effectively preserving true allelic diversity and significantly improving the integrity of the benchmark set and the accuracy of fragmented variation identification. Next, the preliminary benchmark set is evaluated using multiple independent corrected variation datasets, extracting shared false positives for supplementation and removing shared false negatives, achieving self-correction and quality closed-loop optimization of the benchmark set, resulting in a corrected benchmark set that significantly reduces false positive and false negative rates. Finally, multiple corrected benchmark sets from different data or different detection methods are merged, fully integrating the complementary advantages of various technical approaches, ultimately obtaining a high-quality robust benchmark set with more comprehensive coverage, stronger stability, and more balanced representativeness, providing a reliable "gold standard" for the accurate evaluation of structural variation detection tools.

[0053] This embodiment provides a method for constructing a robust benchmark set of human genome structural variations. It integrates variation datasets from various assembly and variation invocation methods to form the benchmark set. Bidirectional validation matching considers the merging of multiple small variations with a large variation, and innovatively designs bidirectional local joint analysis, effectively solving the technical bottleneck of traditional methods' difficulty in capturing the correlation between adjacent variations. This avoids redundancy issues caused by differences in representation formats within the benchmark set. Furthermore, to further improve the quality of the benchmark set, a variation dataset correction method based on shared FN / FP screening across multiple datasets is proposed. This innovatively introduces a "multi-tool consensus validation" logic: using the results of multiple mainstream variation detection tools as a reference, variations consistently identified as false negatives (FN) by multiple tools are removed from the benchmark set, while variations consistently identified as false positives (FP) by multiple tools are added to the benchmark set. This achieves bidirectional optimization of the benchmark set by "removing false positives and retaining true positives" and "completing omissions," thereby improving the quality of the benchmark set.

[0054] This embodiment provides a method for constructing a robust benchmark set of human genome structural variations, which overcomes the limitations of existing benchmark data resources and construction methods, including the lack of high-quality benchmark sets, failure to consider fragmented variations, and loss of some alleles. Through technological innovation, this embodiment improves the quality of the variation benchmark dataset, providing a reliable reference standard for performance evaluation and algorithm optimization of variation detection tools. This will contribute to research in clinical medicine, genetics, and biomedicine. Experimental verification shows that the method of this invention exhibits significant advantages and is expected to have a positive impact on clinical medicine, genetics, and biomedicine.

[0055] The robust benchmark set of structural variations constructed in this embodiment can provide a highly reliable reference standard: the variations have extremely high accuracy and can be directly used as the "standard answer" for evaluating other test results; the robust benchmark set of structural variations constructed in this embodiment covers variations of different types, sizes and genomic regions, avoiding evaluation bias caused by a single reference standard and ensuring the comprehensiveness of the evaluation.

[0056] The robust benchmark set of structural variations constructed in this embodiment can objectively evaluate the performance of variation detection tools: it supports quantitative evaluation of the performance of different variation detection tools based on key indicators such as recall (whether true variations can be detected) and precision (whether the detected variations are true). This helps researchers and developers quickly identify the strengths and weaknesses of tools.

[0057] The structurally robust benchmark set constructed in this embodiment enables unified construction of multi-sample benchmarks: avoiding the need to develop separate construction processes for different samples, reducing redundant R&D costs, and improving benchmark set production efficiency. It ensures consistency in the construction logic and validation standards of different sample benchmark sets, eliminating the problem of uneven benchmark quality caused by methodological differences.

[0058] The robust structural variation benchmark set constructed in this embodiment supports cross-sample commonality and difference analysis: it allows for the comparison of common patterns in variation characteristics among different samples based on a unified method, such as the general distribution patterns of specific variation types. It can accurately capture sample-specific differences, providing a reliable reference for population genetics research.

[0059] The structural variation robust benchmark set constructed in this embodiment can improve the universality of the evaluation system: the constructed multi-sample benchmarks can jointly form a "benchmark set matrix", which supports the evaluation of the stability of the detection tool process under different sample backgrounds, avoiding the limitation that the tool only performs well on a single sample.

[0060] The robust structural variation benchmark set constructed in this embodiment can promote the standardization of variation analysis procedures: providing a unified evaluation basis for different research teams, reducing differences in results caused by inconsistent reference standards, and promoting the reproducibility of research results. It can be used to optimize key steps in the variation analysis process, such as data filtering and variation detection parameter adjustment, thereby improving the overall reliability of the analysis.

[0061] The robust benchmark set of structural variations constructed in this embodiment can support validation in clinical and research applications. In the clinical field, it can be used to verify the accuracy of variation detection results related to genetic diseases, reducing the impact of false positives or false negatives on diagnostic decisions. In the research field, it provides a fundamental guarantee for the discovery and functional analysis of variations in genomics research, ensuring that subsequent research is based on reliable variation data.

[0062] This embodiment also provides a system for constructing a robust benchmark set of human genome structural variations, used to implement the above-described method for constructing a robust benchmark set of human genome structural variations, including: The preprocessed variant dataset generation module is used to process sequencing data through multiple variant detection tools to obtain multiple structural variant datasets. Each structural variant dataset is preprocessed, including adding a unique index identifier to each variant, dividing the data into subsets by chromosome, and determining the search region for each variant based on variant type and breakpoint location, resulting in multiple preprocessed variant datasets. The preliminary benchmark set generation module is used to perform bidirectional validation matching on multiple preprocessed variant datasets, summarize all comprehensive matching relationships, and obtain a comprehensive matching relationship set. In the comprehensive matching relationship set, variants supported by a preset number of detection tools and their corresponding matching variants are extracted to form supporting variant matching groups. For supporting variant matching groups of variants with allele markers, variant size ratio filtering and identical base filtering are performed, and all remaining variants are stored in the allele array. The optimal variants corresponding to the supporting variant matching groups of all variants without allele markers are obtained to form the optimal variant array, and the allele array is added to obtain the preliminary benchmark set. The modified baseline set generation module is used to evaluate the preliminary baseline set using multiple modified variant datasets to obtain a false negative set and a false positive set. Common false positives in the false positive set are added to the preliminary baseline set, and common false negatives in the false negative set are removed from the preliminary baseline set to obtain the modified baseline set. The structural variation robust benchmark set generation module is used to merge multiple modified benchmark sets obtained from different data sources and / or different structural variation detection methods to obtain a structural variation robust benchmark set.

[0063] This embodiment also provides an apparatus for constructing a robust benchmark set of human genome structural variations, including a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the above-described method for constructing a robust benchmark set of human genome structural variations.

[0064] This embodiment also provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the above-described method for constructing a robust benchmark set of human genome structural variations.

[0065] While exemplary embodiments of the invention have been described herein, many other variations or modifications conforming to the principles of the invention can be directly determined or derived from the disclosure of this invention without departing from its spirit and scope. Therefore, the scope of the invention should be understood and recognized to cover all such other variations or modifications.

Claims

1. A method for constructing a robust benchmark set of human genome structural variations, characterized in that, This includes the following operations: S1. Sequencing data is processed by multiple variant detection tools to obtain multiple structural variant datasets. Each structural variant dataset is preprocessed, including adding a unique index identifier to each variant, dividing the data into subsets according to chromosomes, and determining the search region for each variant based on variant type and breakpoint location, resulting in multiple preprocessed variant datasets. S2. Perform bidirectional verification matching on each pair of multiple preprocessed mutation datasets, and summarize all comprehensive matching relationships to obtain a comprehensive matching relationship set; In the comprehensive matching set, the variants that are supported by a number of preset detection tools and the corresponding matching variants are extracted to form a supporting variant matching group; For the supporting variant matching groups of variants with allele markers, perform variant size ratio filtering and same base filtering, and store all remaining variants in the allele array; obtain the optimal variants corresponding to the supporting variant matching groups of all variants without allele markers, form the optimal variant array, add the allele array, and obtain the preliminary benchmark set. S3. Evaluate the preliminary baseline set using multiple corrected variant datasets to obtain a false negative set and a false positive set. Add the common false positives in the false positive set to the preliminary baseline set and remove the common false negatives in the false negative set from the preliminary baseline set to obtain the corrected baseline set. S4. Merge the modified benchmark sets obtained from multiple different data sources and / or different structural variation detection methods to obtain a structural variation robust benchmark set.

2. The method for constructing a robust benchmark set of human genome structural variations according to claim 1, characterized in that, In S2, the bidirectional verification matching operation is as follows: take two preprocessed mutation datasets as the user set and the reference set, respectively; when the starting position of the mutation in the reference set is in the search area of ​​the mutation in the user set, determine whether the mutations overlap. If so, perform bidirectional local joint analysis and component matching analysis. If component matching analysis fails, perform regular matching judgment, obtain the matching result of the current round, and store it in the matching array of the previous round to obtain the matching array of the current round. If not, perform component matching analysis based on the previous round's matching array to obtain the current round's matching results, store them in the previous round's matching array, and obtain the current round's matching array; until all mutations in the user set and the baseline set have undergone bidirectional verification matching, output the last round's matching array to obtain the comprehensive matching relationship.

3. The method for constructing a robust benchmark set of human genome structural variations according to claim 2, characterized in that, Bidirectional local joint analysis includes local joint analysis from the user set to the reference set, and local joint analysis from the reference set to the user set; The operation of local joint analysis from user set to reference set is as follows: based on the variant type and breakpoint proximity, variants of the same type in the neighborhood are classified into candidate variant set; the variants in the candidate variant set are combined to perform the first sequence consistency evaluation; if the sequence consistency evaluation conditions are met, a positive matching relationship is obtained. The local joint analysis from the benchmark set to the user set follows the same steps as the local joint analysis from the user set to the benchmark set, only the roles of the analysis objects of the user set and the benchmark set are interchanged; the results of the local joint analysis from the benchmark set to the user set are obtained as a reverse matching relationship, which, together with the forward matching relationship, forms a bidirectional successful local joint analysis result.

4. The method for constructing a robust benchmark set of human genome structural variations according to claim 3, characterized in that, In the operation of grouping similar variants in the neighborhood into the candidate variant set, the first candidate set and the second candidate set are initialized. For insertion mutations, starting from the current insertion type mutation, all insertion mutations of the same type are traversed backward and added to the first candidate set in turn, until the difference between the starting position of the insertion mutation and the starting position of the first insertion mutation in the first candidate set is greater than the first distance threshold, and then the insertion candidate mutation set is obtained. For missing variants, record the end position of the current missing variant, traverse the same type of missing variants backward, if the difference between the start position of the missing variant to be added to the second candidate set and the recorded end position is not greater than the preset second distance threshold, then add the corresponding missing variant to the second candidate set, and update the end position record value with the end position of the corresponding missing variant, repeat until the condition is not met, and obtain the missing candidate variant set. The candidate variant set is formed by inserting candidate variant sets and missing candidate variant sets.

5. The method for constructing a robust benchmark set of human genome structural variations according to claim 3, characterized in that, The variants in the candidate variant set are combined in order of breakpoint location. If the original alignment BAM file is not provided, enumerate all variant combinations and filter variant combinations whose size ratio with the baseline set is less than the size ratio threshold. If the original alignment BAM file is provided, the original variant signal of the read segment corresponding to each variant combination is scanned, and variant combinations that are not on the same read segment are filtered by haplotype division, and combinations with a variant size ratio < variant size ratio threshold are filtered. Combinations with overlapping original coordinates of the variants are filtered out. From the remaining combinations, the combination whose variant size is closest to that of the baseline set is selected as the optimal merging candidate for sequence consistency evaluation.

6. The method for constructing a robust benchmark set of human genome structural variations according to claim 3, characterized in that, In the first sequence consistency assessment process, for the optimal merge candidate obtained after combination, the upstream and downstream reference sequences of the coverage area and the coverage area of ​​the variant in the baseline set are extracted to construct the context sequence and calculate the sequence consistency. If the sequence consistency is greater than or equal to the sequence consistency threshold, all variants in the optimal merge candidate are determined to be locally joint positive, the corresponding variants in the baseline set are marked as true positives, all variant indices in the combination are returned, and the bidirectional index relationship is recorded in the two-dimensional array storing the variant matching relationship to obtain the positive matching relationship.

7. The method for constructing a robust benchmark set of human genome structural variations according to claim 1, characterized in that, In S2, the optimal mutation is selected based on the mutation length and overlap.

8. A system for constructing a robust benchmark set of human genome structural variations, used to implement the method for constructing a robust benchmark set of human genome structural variations as described in claim 1, characterized in that, include: The preprocessing variant dataset generation module is used to process sequencing data through multiple variant detection tools to obtain multiple structural variant datasets; Each structural variant dataset is preprocessed, including adding a unique index identifier to each variant, dividing the data into subsets by chromosome, and determining the search region for each variant based on the variant type and breakpoint location, resulting in multiple preprocessed variant datasets. The preliminary baseline set generation module is used to perform bidirectional validation matching on each pair of multiple preprocessed variant datasets, summarize all comprehensive matching relationships, and obtain a comprehensive matching relationship set. In the comprehensive matching set, the variants that are supported by a number of preset detection tools and the corresponding matching variants are extracted to form a supporting variant matching group; For the supporting variant matching groups of variants with allele markers, perform variant size ratio filtering and same base filtering, and store all remaining variants in the allele array; obtain the optimal variants corresponding to the supporting variant matching groups of all variants without allele markers, form the optimal variant array, add the allele array, and obtain the preliminary benchmark set. The modified baseline set generation module is used to evaluate the preliminary baseline set using multiple modified variant datasets to obtain a false negative set and a false positive set. Common false positives in the false positive set are added to the preliminary baseline set, and common false negatives in the false negative set are removed from the preliminary baseline set to obtain the modified baseline set. The structural variation robust benchmark set generation module is used to merge multiple modified benchmark sets obtained from different data sources and / or different structural variation detection methods to obtain a structural variation robust benchmark set.

9. A device for constructing a robust benchmark set of human genome structural variations, characterized in that, It includes a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the method for constructing a robust benchmark set of human genome structural variations as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the method for constructing a robust benchmark set of human genome structural variations as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Evaluation method and system for genome structure variation detection based on reference set

    CN118645152A

  • Long-read-length data robust variation detection method and system based on data quality improvement

    CN122177206A