Structural variation recognition method and system based on generic genome map

By using a pan-genome mapping-based structural variation identification method, which expands the mapping using third-generation sequencing data and combines it with a deep learning model for intelligent filtering, the shortcomings of existing technologies in identifying unknown or individual-specific structural variations are overcome, and efficient and accurate structural variation analysis is achieved.

CN121884945AActive Publication Date: 2026-04-17XIANGYA HOSPITAL CENT SOUTH UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIANGYA HOSPITAL CENT SOUTH UNIV
Filing Date
2026-03-18
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing pangenome mapping technologies lack specificity in identifying unknown or individual-specific structural variations, lack dynamic coupling between construction and sample analysis, lack adaptive filtering strategies, fail to fully reflect the complex characteristics of structural variations, and lack quantitative modeling of the credibility of structural variations, resulting in insufficient identification accuracy and sensitivity.

Method used

A structural variation identification method based on pan-genome mapping is adopted. The map is expanded by third-generation sequencing data, and a deep learning model is used to evaluate and classify the credibility of candidate structural variations. Intelligent filtering is performed using map comparison features to achieve efficient capture of complex and individual-specific variations.

Benefits of technology

It improves the sensitivity and accuracy of structural variation identification, enhances the coverage of unknown or individual-specific variations, reduces reference bias, and improves the consistency and traceability of analysis results. It is suitable for stable analysis under different sequencing conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884945A_ABST
    Figure CN121884945A_ABST
Patent Text Reader

Abstract

The invention discloses a structural variation recognition method and system based on a generic genome map. The generic genome map is expanded / enhanced by using third-generation sequencing data; receiving input sequencing data, comparing the sequencing data with the expanded / enhanced generic genome map to obtain a map reference comparison result, identifying structural variation in the sequencing data based on the map reference comparison result, and generating a candidate structural variation set; and performing credibility evaluation and classification on the candidate structure variation set by adopting a pre-trained deep learning model to obtain a filtered high-credibility structure variation set. According to the method, on the basis of utilizing a standard generic genome map, high-quality three-generation sequencing data owned by a user is introduced to dynamically construct or expand the map, so that a reference structure can cover more unknown or individual specific structure variations, the sensitivity and the accuracy of structure variation recognition are improved, and the application prospect is wide. And the capturing capability on complex and individual specific variation is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene sequence structural variation detection technology, and in particular to a method and system for identifying structural variations based on pan-genome maps. Background Technology

[0002] With the rapid development of sequencing technology, next-generation sequencing (NGS) and third-generation sequencing (TGS) are increasingly widely used in genome research. Genome analysis based on short-read sequencing data (such as Illumina) and high-precision long-read sequencing data (such as PacBio HiFi and ONT) provides important tools for revealing complex genetic variations. Among these, structural variants (SVs) have become a key focus of precision medicine research due to their core regulatory role in genetic diseases and complex traits. Traditional linear reference gene sets (such as GRCh38) exhibit significant reference bias due to their single-sequence nature. When dealing with highly repetitive or complex rearranged regions, this often leads to inaccurate breakpoint localization, and the consistency between different tools such as Sniffles, SVIM, and Manta is poor, making it difficult to meet the high-precision requirements of clinical practice.

[0003] To overcome the limitations of linear reference genomes in representing population genetic diversity, pangenome mapping technology has been proposed in recent years. This technology uses a graph structure to simultaneously represent the sequence and variation information of multiple individuals, allowing different allelic genotypes to coexist in the same reference frame as nodes and edges. This reduces reference bias to some extent and improves the alignment capability of complex regions.

[0004] Existing methods for identifying structural variations based on pan-genome maps typically employ the following technical approach: first, a genome map is constructed based on an existing set of variations; then, sequencing data is mapped to the map structure; finally, potential structural variations are inferred based on the path selection, branch support, or path deviation of sequencing reads in the map.

[0005] Although the above methods have significant advantages over linear reference genomes in terms of alignment, they still have the following shortcomings in practical applications:

[0006] 1. It is highly dependent on existing variation information.

[0007] Most current pangenome maps are constructed based on existing variant databases or known structural variants, and their identification capabilities are largely limited by the types and range of variants pre-included in the map. For unknown, individual-specific, or population-specific structural variants, existing methods often lack sufficient specificity.

[0008] 2. There is a lack of dynamic coupling mechanism between map construction and sample analysis.

[0009] Most existing genome maps are statically constructed, making them difficult to expand flexibly once generated. They also make it difficult to dynamically integrate users' own third-generation sequencing data or high-quality structural variation information into the maps, thus limiting the application value of pan-genome in personalized analysis and specific research scenarios.

[0010] 3. Graph methods mainly focus on representation and recognition.

[0011] Existing technologies mostly use pan-genome maps as tools for alignment and variant identification. Their output still relies on traditional post-processing workflows for screening and analysis, lacking systematic integration with subsequent variant filtering, functional annotation, and multi-omics analysis.

[0012] Therefore, although existing pangenome mapping technology has made progress in mitigating linear reference bias, it has not yet formed a complete technical system for structural variations from identification to application.

[0013] In the post-identification screening stage, traditional methods mainly rely on preset hard thresholds (such as the number of supported reads, alignment quality, etc.). With the surge in the number of candidate variants, their limitations have become increasingly apparent.

[0014] 1. Difficulty in characterizing the multidimensional and complex features of structural variations

[0015] Structural variations typically involve features such as breakpoint uncertainty, complex repetitive structures, and multi-path support. A single or limited set of threshold rules can hardly fully reflect the authenticity and reliability of the variations.

[0016] 2. Unable to fully utilize the structural information generated by map comparison

[0017] The large amount of information generated during graph comparison, such as path selection probability, branch node support, and graph structure complexity, is difficult to effectively integrate and utilize using traditional rules.

[0018] 3. Filtering strategies lack adaptability and generalization ability.

[0019] Filtering methods based on fixed rules are highly sensitive to sequencing depth, sample type, and data source. They often require repeated parameter adjustments in different research scenarios and have poor stability.

[0020] 4. Lack of quantitative modeling for the uncertainty of structural variations

[0021] Existing methods mostly output results in a binary decision form of "keep or eliminate", failing to systematically model and quantify the credibility of structural variations.

[0022] Therefore, with the increasing number of pan-genome structures and complex structural variations, traditional rule-based filtering methods are no longer sufficient to meet the needs of high-precision structural variation analysis, and there is an urgent need to introduce intelligent filtering strategies that can automatically learn complex features. Summary of the Invention

[0023] The technical problem to be solved by the present invention is to provide a method and system for identifying structural variations based on pan-genome mapping, which addresses the shortcomings of the existing technology, improves the sensitivity and accuracy of structural variation identification, and enhances the ability to capture complex and individual-specific variations.

[0024] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a method for identifying structural variations based on pan-genome maps, comprising the following steps:

[0025] S1. Expand / enhance the pan-genome map using third-generation sequencing data;

[0026] S2. Receive the input sequencing data, compare the sequencing data with the expanded / enhanced pan-genome map to obtain the graph reference alignment result, and based on the graph reference alignment result, identify structural variations in the sequencing data and generate a candidate structural variation set.

[0027] S3. A pre-trained deep learning model is used to evaluate and classify the credibility of the candidate structural variant set to obtain a filtered structural variant set.

[0028] This invention uses a pan-genome map as a unified reference framework to replace the traditional linear reference genome, effectively reducing reference bias in structural variant identification and subsequent analysis, and improving the accuracy and completeness of identifying complex structural variants. This invention also supports the dynamic construction or expansion of the map using user-owned high-quality third-generation sequencing data, based on the standard pan-genome map. This allows the reference structure to cover more unknown or individual-specific structural variants, thereby improving the sensitivity and accuracy of structural variant identification and enhancing the ability to capture complex and individual-specific variants.

[0029] The specific implementation process of step S1 includes:

[0030] Determine whether the quality of the third-generation sequencing data meets the requirements for map construction or expansion. If so, use a map construction strategy based on genome assembly to expand the pan-genome map, or use a structural variation identification method based on reference genome or pan-genome alignment to generate map increment information; otherwise, directly use the original pan-genome map as the expanded pan-genome map.

[0031] The specific implementation process for determining whether the quality of third-generation sequencing data meets the requirements for map construction or expansion, and for using a genome assembly-based map construction strategy to expand the pan-genome map, or for using a structural variation identification method based on reference genome or pan-genome alignment to generate incremental map information, includes:

[0032] When the coverage depth of PacBio HiFi data is not less than 30×, meets the sequencing accuracy assessment, and the read length N50 value is not less than 20kb, a map construction strategy based on genome assembly is used to expand the pan-genome map; when the coverage depth of PacBio HiFi data is less than 30×, a structural variation identification method based on reference genome or pan-genome alignment is used to generate map increment information.

[0033] When the coverage depth of PacBio CLR or ONT data is not less than 50×, meets the sequencing accuracy assessment, and the read length N50 value is not less than 20kb, a map construction strategy based on genome assembly is used to expand the pan-genome map; otherwise, a structural variation identification method based on reference genome or pan-genome alignment is used to generate incremental map information.

[0034] The specific implementation process of step S2 includes:

[0035] Preprocess the sequencing data;

[0036] The preprocessed sequencing data is mapped onto the expanded / enhanced pangenome map to generate the path representation of the sequencing reads in the map;

[0037] By utilizing the path offset, breakage, or cross-branch alignment behavior of sequencing reads in the graph, we can identify structural variations such as insertions, deletions, inversions, and duplications. By combining the coverage changes of nodes and edges, path support strength, and graph topological differences, we can generate a candidate set of structural variations.

[0038] The specific implementation process of step S3 includes:

[0039] Extract core features from the candidate set of structural variants, use the extracted core features as input to a pre-trained deep learning model, and output the probability value and feature contribution vector of each candidate variant.

[0040] If a candidate mutation is located in the core functional area and the probability value of the candidate mutation is greater than the first set threshold, it is determined to pass; if a candidate mutation is located in the highly repetitive and complex area and the probability value of the candidate mutation is greater than the second set threshold, it is determined to pass; wherein, the second threshold is greater than the first threshold.

[0041] If the site deletion rate of a candidate variant that has been approved exceeds a set threshold, it is marked as unreliable.

[0042] Using the aforementioned feature contribution vector, specific filtering attributions are labeled for each candidate variant that was not determined to pass.

[0043] If the input data received in step S2 is third-generation sequencing data, the following judgment is also made: if the number of valid reads supporting the candidate variant is lower than the set threshold, it is determined to fail.

[0044] The core features include basic quality features, supporting signal features, topological complexity features, path consistency, genotype features, and structural variation types; the basic quality features include quality score, genotype quality value, and site deletion rate; the supporting signal features include sequencing depth, allele depth, and positive / negative strand support bias; the topological complexity features include the number of graph path bifurcations where the variation is located, the average alignment quality of reads in that region, and whether it is located in a highly repetitive region; the structural variation types include insertions, deletions, inversions, duplications, and translocations.

[0045] The feature contribution vector includes key factor labels, which include: variant reads cannot form a closed and continuous logical flow in the haplotype path in the graph; false positives of inversion; variants located in known highly repetitive regions and accompanied by multi-site alignment signals; the number of graph node bifurcations exceeds a preset threshold, resulting in an explosion of the alignment path search space; there is a statistically significant difference between the insertion / deletion length and the physical step size after Read alignment; and the average alignment quality of the variant breakpoint and its flanking regions is lower than the safety threshold.

[0046] As an inventive concept, the present invention also provides a structural variation identification system based on pan-genome mapping, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the steps of the above method.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0048] (1) This invention uses pan-genome maps as a unified reference framework to replace traditional linear reference genomes, which can effectively reduce reference bias in the identification of structural variations and subsequent analysis, and improve the accuracy and completeness of the identification of complex structural variations.

[0049] (2) This invention supports the dynamic construction or expansion of the map by introducing user-owned high-quality third-generation sequencing data on the basis of standard pan-genome maps, so that the reference structure can cover more unknown or individual-specific structural variations, thereby improving the sensitivity and accuracy of structural variation identification and enhancing the ability to capture complex and individual-specific variations.

[0050] (3) The present invention integrates the construction or retrieval of pan-genome maps, sequencing data alignment, structural variation identification and intelligent filtering into a single technical process, avoiding the problem of the separation of analysis steps in the prior art, and improving the consistency, traceability and reproducibility of the analysis results.

[0051] (4) This invention supports both second-generation sequencing data and third-generation sequencing data as input, and adopts an appropriate analysis strategy according to the characteristics of different sequencing data, so that structural variation identification and analysis can be carried out stably under different sequencing conditions, which significantly improves the versatility and practical application value of the method.

[0052] (5) The present invention uses machine learning or deep learning methods to intelligently filter structural variation candidates, which can effectively distinguish between real variations and noise when there are many candidate variations and complex structures, reduce the false positive rate, and improve the reliability of structural variation results.

[0053] (6) The present invention adopts a modular design architecture, and each functional module can be operated independently or used in combination. It is suitable for large-scale sample analysis and different application scenarios, and has good scalability and engineering feasibility. Attached Figure Description

[0054] Figure 1 A flowchart for pan-genome map selection and management;

[0055] Figure 2 This is a flowchart of the intelligent filtering process for structural variations. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] Given the numerous shortcomings of existing technologies in the identification and filtering of structural variations, particularly in the ability of next-generation sequencing data to identify complex structural variations and the dynamic utilization of pan-genome maps, the embodiments of this invention aim to solve the following technical problems:

[0058] To address the problem that existing structural variation identification methods based on second-generation sequencing data heavily rely on linear reference genomes, and that existing pan-genome maps largely depend on existing variation information, making it difficult to effectively cover unknown or individual-specific structural variations, thus limiting the ability of second-generation sequencing data to identify complex structural variations, this paper proposes a method that utilizes existing pan-genome maps or selectively introduces high-quality third-generation sequencing data from users to construct or dynamically expand the maps. This aims to reduce reference bias and improve the detection rate and accuracy of second-generation sequencing data in structural variation identification.

[0059] To address the problem that most existing pan-genome maps are statically constructed, with the map construction process being disconnected from the sample-level structural variation identification and subsequent analysis processes, this paper proposes a structural variation analysis technique that organically couples the pan-genome map construction or update process with the structural variation identification and subsequent analysis processes, thereby achieving synergistic optimization between map construction and sample analysis.

[0060] To address the problem that traditional rule-based filtering methods struggle to effectively distinguish between genuine structural variations and sequencing noise or alignment errors in structural variation identification results based on pan-genome maps due to the large number and complex types of candidate variations, this paper proposes an intelligent structural variation filtering method based on machine learning or deep learning models. This method enables automated and refined evaluation of the authenticity and reliability of candidate structural variations.

[0061] Through the above-mentioned technical means, the embodiments of the present invention aim to solve the key technical bottlenecks in the process of structural variation identification and filtering, and achieve efficient identification and highly reliable screening of structural variations.

[0062] Example 1

[0063] This invention provides an integrated system or tool for structural variant identification and intelligent filtering based on pan-genome mapping. Its overall structure adopts a modular design, comprising: a pan-genome mapping selection and management module, a sequencing data alignment and structural variant identification module, and a structural variant intelligent filtering module. These modules are connected through standardized data interfaces, forming a complete analysis workflow from raw sequencing data input to the output of a high-confidence set of structural variants. The system in this embodiment includes the following modules:

[0064] The pan-genome map selection and management module provides a graphical reference basis for subsequent analyses. This module allows users to select existing standard pan-genome maps based on their analytical needs. Pan-genome maps can be obtained from public databases or pre-constructed graphical reference genomes.

[0065] In addition, this module also allows users to use their own high-quality third-generation sequencing data (such as PacBio HiFi) to expand or enhance existing pan-genome maps. By introducing new structural variation information and sequence paths, it can improve the map's coverage of individual-specific or low-frequency structural variations, providing a more comprehensive reference structure for subsequent variation identification based on second-generation sequencing data.

[0066] The sequencing data alignment and structural variation identification module is used to receive sequencing data input by the user and complete the integrated processing of read alignment and structural variation identification based on the pan-genome map.

[0067] This module primarily uses second-generation sequencing data as the analysis object (third-generation sequencing data can also be used as input), comparing the sequencing data with a pan-genome map selected or constructed by the user. Unlike traditional alignment methods based on a linear reference genome, this module performs read mapping based on a graph reference structure. This allows for more efficient use of the information contained in the second-generation sequencing data in complex genomic regions, multi-copy regions, and regions with structural rearrangements, thereby reducing reference bias and improving alignment accuracy.

[0068] After completing the alignment, this module further identifies structural variations in the samples based on the graph reference alignment results, generating a candidate structural variation set. The identified structural variations include, but are not limited to, insertions, deletions, inversions, duplications, and other complex structural variation types. This module focuses on optimizing the limitations of second-generation sequencing data in identifying complex structural variations by utilizing the multi-path information contained in the pan-genome map to improve the ability to identify breakpoints and variation forms of complex structural variations.

[0069] By introducing pan-genome maps constructed or enhanced by third-generation sequencing data, this module can significantly improve the probability and accuracy of identifying potential structural variants in second-generation sequencing data without directly relying on third-generation sequencing data as the detection target, thereby obtaining a more comprehensive and reliable set of candidate structural variants.

[0070] The intelligent filtering module for structural variations is used to evaluate the credibility and classify the set of candidate structural variations. This module automates the analysis of each candidate structural variation based on a pre-built machine learning or deep learning model.

[0071] The model (optional XGBoost / RandomForest / LSTM / CNN / Transformer, etc., training drives filtering accuracy) comprehensively utilizes map alignment features, sequencing depth features, breakpoint support information, quality values, genotype and its score information, structural variant type features and other statistical or structural features to perform multi-dimensional discrimination of structural variants, thereby distinguishing true structural variants from false positive results caused by sequencing noise or alignment errors, and outputting a high-quality set of structural variants after intelligent filtering.

[0072] The workflow of the pan-genome map selection and management module is as follows: Figure 1 As shown. This module provides functions for selecting, constructing, expanding, and managing human pan-genome maps, and serves as the fundamental data framework for structural variation identification and analysis. Its workflow includes the following steps:

[0073] Step 1: The user starts the analysis task, selects an existing standard graph or uploads a custom graph;

[0074] When a user initiates an analysis task, the system provides the user with a built-in standard human pan-genome map library for structural variation identification and analysis. The standard pan-genome map includes publicly released versions of the human pan-genome such as CPC (Complete Pan-Genome Consortium) and HPRC (Human Pangenome Reference Consortium).

[0075] Users can select a specific version of the pangenome map according to their analysis needs or upload a pangenome map and index file that they have already constructed. The system will automatically load the corresponding map file and index information during the task initialization phase and record the name, version number and release time of the selected map for subsequent result tracing and reproduction.

[0076] Step 2: Triggering steps for graph construction or expansion

[0077] When a user selects a basic pan-genome map in an analysis task and optionally inputs their own third-generation sequencing data file, the system automatically receives the third-generation sequencing data. The data format is limited to FASTQ or the raw BAM file downloaded from the sequencing machine.

[0078] Before performing pan-genome map construction or expansion, the system first performs pre-quality verification on the received third-generation sequencing data to determine whether it meets the preset quality requirements for map construction or expansion. Then, it selects a map construction strategy to ensure the accuracy and completeness of the newly added map structure.

[0079] (I) Identification and quality assessment of third-generation sequencing data types

[0080] The system first identifies the type of the third-generation sequencing data input by the user, and then performs automated quality assessment based on the corresponding sequencing platform. The quality assessment includes at least the following:

[0081] Genome coverage depth assessment:

[0082] For long-length sequencing data, such as PacBio HiFi, PacBio CLR, or Oxford Nanopore Technologies (ONT) data, assess their average whole-genome coverage depth;

[0083] Read length statistics assessment:

[0084] Calculate the read length N50 value, preferably not less than 20 kb, more preferably not less than 30 kb, to ensure that the read segment can cross complex structural regions;

[0085] Sequencing accuracy assessment:

[0086] For PacBio HiFi data, the single read length accuracy is preferably no less than 99.9% (Q30).

[0087] For ONT data, the raw sequencing accuracy is preferably no less than 98%, and error correction or consensus generation processes are used to meet the accuracy requirements for map construction.

[0088] (ii) Automatic determination of map construction or expansion strategies

[0089] Based on the type and quality assessment results of the third-generation sequencing data, the system automatically determines the applicable map construction or expansion path, and the specific rules include, but are not limited to:

[0090] PacBio HiFi data processing rules:

[0091] When the coverage depth of PacBio HiFi data is not less than 30× (i.e., 30 times), meets the sequencing accuracy assessment, and N50 is not less than 20 kb, the system allows the use of a map construction strategy based on genome assembly.

[0092] When the coverage depth of PacBio HiFi data is less than 30×, the system generates incremental map information by identifying structural variations based on the reference genome or existing pan-genome alignment, instead of performing complete genome assembly.

[0093] PacBio CLR or ONT data processing rules:

[0094] When the coverage depth of PacBio CLR or ONT data is not less than 50×, meets the sequencing accuracy assessment, and N50 is not less than 20 kb, the system adopts a map construction strategy based on genome assembly.

[0095] When the coverage depth does not reach the above threshold, the system uses a structural variation identification method based on the reference genome or existing pan-genome alignment to generate incremental map information.

[0096] Handling cases where the construction conditions are not met:

[0097] When the overall quality of third-generation sequencing data does not meet the minimum requirements for map construction or expansion, the system will not trigger the map construction or expansion process, but will directly use the basic pan-genome map selected by the user as the reference map for subsequent structural variation identification.

[0098] Step 3: Execution of the graph construction strategy and graph integration

[0099] After completing the quality assessment of the third-generation sequencing data input by the user and confirming that the conditions for map construction or expansion are met and the map construction strategy is selected, the system performs map construction, map integration and optimization processing to generate a user-defined enhanced pan-genome map for subsequent structural variation identification.

[0100] (a) Basic quality control of user-input third-generation sequencing data

[0101] In response to the technical characteristics of different third-generation sequencing platforms, this system adopts a raw data filtering and quality control strategy based on the differences in sequencing platforms and chemical systems, so as to maximize the reliability of downstream analysis while ensuring data accuracy.

[0102] 1. PacBio platform third-generation sequencing data filtering strategy

[0103] For PacBio sequencing platforms, including Sequel II, Sequel IIe, and Revio systems, the raw data are analyzed for consistency using the CCS algorithm in the official pbbioconda to construct circular consensus sequences (CCS).

[0104] Since the HiFi data generated by the aforementioned platforms has achieved unified high-precision modeling at the algorithm level, its read segment accuracy has been consistently above Q20, and its quality distribution is highly consistent. Therefore, this system no longer performs additional base quality value filtering on the HiFi data.

[0105] Based on this, this system uses length filtering as the core step in PacBio HiFi data filtering:

[0106] The quality of the generated HiFi reads was statistically analyzed and evaluated using tools such as NanoPlot.

[0107] Depending on the specific analysis objective (such as haplotype assembly, structural variation detection, or recombination event analysis), users can customize the minimum read length threshold (e.g., ≥2 kb or ≥3 kb), with a default of 1 kb;

[0108] Use tools such as seqkit to remove short reads that are below the target length threshold.

[0109] By employing the above methods, while maintaining the high accuracy advantage of HiFi data, we can further improve the coverage of complex genomic regions and the continuity of assembly of reads.

[0110] 2. Data Filtering Strategy for Third-Generation Sequencing on the Oxford Nanopore Platform

[0111] For the Oxford Nanopore Technologies (ONT) platform, this system employs a stratified and differentiated data filtering strategy based on different sequencing chips and chemical versions.

[0112] (1) Data for the chemical system in R9.4.1

[0113] For R9.4.1 standard single-stranded sequencing data, read accuracy and length distribution exhibit some fluctuation. This system employs a combined length and quality filtering strategy under these conditions:

[0114] Use tools such as NanoFilt to set the minimum read segment length threshold (e.g., ≥1 kb), the default is 1 kb;

[0115] Simultaneously set the minimum average quality value threshold (e.g., Q7–Q10), with Q7 as the default.

[0116] Trim the ends of the reads to reduce the interference of low-quality bases on subsequent analyses.

[0117] This strategy effectively controls the risk of sequencing error accumulation while ensuring the number of valid reads.

[0118] (2) Data for R10.4.1 and R10.4.1 Duplex chemical systems

[0119] For R10.4.1 and the newer R10.4.1 Duplex chemistry systems, this system primarily utilizes its high-precision sequencing capabilities:

[0120] By processing the raw data using Dorado's duplex mode, a dual-chain consensus read segment jointly constructed by the positive and negative chains is extracted. The resulting duplex read segments have significantly improved accuracy, approaching or even reaching the level of HiFi data.

[0121] In this case, the system only performs a length-based filtering strategy (e.g., ≥1 kb) on the Duplex read segment, with 1 kb as the default, and no further quality value filtering is performed, thereby improving the effective data utilization rate while fully preserving high-precision information.

[0122] (II) Execution of the map construction strategy

[0123] 1. A strategy for constructing a genetic map based on reference gene alignment to identify genetic variations.

[0124] The system will align third-generation sequencing reads, which are subject to quality control, to a linear reference genome. By analyzing the alignment results, it will identify newly added structural paths or structural variation information and convert the information into new nodes and edge structures in a pan-genome map.

[0125] (1) Dual-Track Variant Calling

[0126] To balance the comprehensiveness and accuracy of the mutations, a strategy of parallel assembly and comparison is adopted:

[0127] Assembly-based: First, double haploid assembly is performed using hifiasm to obtain two haplotype contigs; then, the assembly results are compared with the reference genome T2T using PAV to obtain the set of variants based on assembly.

[0128] Alignment track (Mapping-based): The quality-controlled third-generation sequencing reads were aligned to the T2T reference genome using minimap2. Variants were detected using variant recognition software such as Sniffles2, PBSV, SVIM, and cuteSV to obtain a set of variants based on the alignment.

[0129] Mutation Consensus: The mutations obtained from two methods are merged using Truvari, Jasmine, or SURVIVOR. The merging strategy is based on the assembled mutation results, and if the mutation results of any of the comparison software are supported, they are retained, resulting in a comprehensive and accurate mutation set.

[0130] (2) Graph Augmentation

[0131] Background mutation acquisition: Download the official population mutation VCF from HPRC or CPC;

[0132] Mutation pool merging: First, perform bcftools normalization on all VCFs (uniform left alignment, split multi-allelic loci into single rows); then, use bcftools merge to merge the VCF set obtained in the previous step with the population variation VCFs;

[0133] Graph Construction: Using the vg construct command, all SNPs, InDels, and SVs in the VCF are transformed into branching paths in the graph, generating the graph structure.

[0134] (3) Graph structure optimization and graph index establishment

[0135] To support efficient comparison of second-generation short-read-long data, a specific index must be built.

[0136] Graph structure optimization: Use vg mod -X 100 to prune the graph and remove overly complex "bubbles" to prevent computational explosion during alignment.

[0137] Building the Giraffe index: vg Giraffe is an alignment tool specifically optimized for geographic maps. It requires using vgautoindex to generate .gbz (graphics), .min (minimized index), and .dist (distance index) files.

[0138] 2. Genome assembly-based map construction strategies

[0139] The system performs individual-level or population-level genome assembly on third-generation sequencing data that has passed quality control to generate continuous assembled sequences or haplotype sequences, and performs structural analysis on the assembly results, mapping them into nodes and edge structures in a pan-genome map.

[0140] (1) High-quality haplotype assembly

[0141] The system will automatically match the assembly pattern based on the data type provided by the user. The core process is as follows:

[0142] ① Input data: PacBio HiFi Reads (required); Parent's 2nd generation data or Hi-C data (optional).

[0143] ② Assembly process:

[0144] Data binning / phasing: If parental data is available, k-mer frequency statistics are performed using yak to identify family-specific characteristics; if Hi-C data is available, phasing is performed using hifiasm's embedded Hi-C ensemble algorithm.

[0145] Core assembly: The hifiasm algorithm is invoked to construct a double haplotype map by identifying highly accurate heterozygous differences between HiFi reads, and to separate two sets of chromosome-level contigs representing the father and mother (or haplotypes 1 and 2).

[0146] Output results: Two independent, non-collapsed, high-quality haplotype assembly sequences (Fasta) are generated.

[0147] (2) Selection of assembly-based pangenome construction strategies and map construction

[0148] For assembly-based pangenome construction, Minigraph, Minigraph-Cactus (MC), and PGGB represent three different design philosophies: Minigraph is a "lightweight backbone" strategy that captures only structural variations >50bp, resulting in extremely fast construction speed and a simple graph structure; Minigraph-Cactus, on the other hand, fills in all SNPs and small InDels on the backbone through full sequence alignment, achieving "full information coverage," and is currently the standard solution that balances accuracy and effect; while PGGB adopts a "completely symmetrical" no-reference alignment approach, without bias towards any single genome, and can most realistically reproduce the complex evolutionary relationships within a species, but it consumes the most computational resources.

[0149] ① Strategy A: Minigraph (SV-first skeleton method)

[0150] Determine the base: Select a high-quality genome as the initial node, such as T2T-CHM13 (default) or GRCh38, and arrange the haplotype sequences obtained in the previous step according to the sample order.

[0151] Sequence alignment and merging: Minigraph is used to iteratively align the haplotype sequences of subsequent samples into a graph constructed from the initial reference genome (such as T2T).

[0152] Mutation filtering: Add new nodes and edges to the graph only when structural mutation paths >50bp are identified.

[0153] Iterative increment: Input the haplotype sequences of the remaining samples sequentially.

[0154] Output: A lightweight pan-genome map containing only SVs.

[0155] ② Strategy B: Minigraph-Cactus / MC (Full Information Fusion Method - Default Process)

[0156] Constructing the SV skeleton: Running minigraph to obtain a GFA graph containing large-scale mutations.

[0157] Call the Cactus alignment engine to perform whole-genome multiple sequence alignment (MSA).

[0158] The abPOA algorithm was used to backfill all SNPs and short insertions / deletions (InDels) into the graph, ensuring that the graph has base-level accuracy.

[0159] Output: A GFA graph containing the full spectrum from SNPs to large SVs, with paths.

[0160] ③ Strategy C: PGGB (Symmetric No-Reference Construction Method)

[0161] Use wfmash to perform an all-to-all comparison of all haplotype sequences.

[0162] Use seqwish to induce the conversion of overlapping information generated by alignment into a mutation map.

[0163] Use smoothxg to perform MSA smoothing on local parts of the image to eliminate alignment noise and optimize path topology.

[0164] Output: A highly symmetric pangenome map that is completely independent of reference genome bias.

[0165] Users should choose based on their research objectives and computing resources: If the focus is on rapidly analyzing large-scale structural variations in large populations (such as hundreds of samples) and high computational efficiency is required, choose Minigraph; if the goal is to provide high-precision "navigation" for second-generation data, requiring full-base-scale variation detection and pursuing standardization and robustness of the workflow, Minigraph-Cactus (MC) is the preferred option; if there are significant genetic differences between samples (such as across species or highly diverse populations), and the goal is to completely eliminate reference genome bias, and sufficient memory resources are available, then PGGB is recommended.

[0166] (3) Graph structure optimization and graph index establishment

[0167] To support efficient comparison of second-generation short-read-long data, a specific index must be built.

[0168] Graph structure optimization: Use vg mod -X 100 to prune the graph and remove overly complex "bubbles" to prevent computational explosion during alignment.

[0169] Building the Giraffe index: vg Giraffe is an alignment tool specifically optimized for geographic maps. It requires using vgautoindex to generate .gbz (graphics), .min (minimized index), and .dist (distance index) files.

[0170] Step 4: Map Management and Version Recording Steps

[0171] The system provides unified management of constructed or expanded pan-genome maps, generates new map version numbers, and records map source information, characteristics of the third-generation sequencing data used for construction or expansion, construction strategy type, and time information. The system supports parallel storage and retrieval of multiple pan-genome versions, enabling different analysis tasks to be run with different map versions, thereby ensuring the traceability and reproducibility of structural variation identification results.

[0172] Technological advantages

[0173] 1. Supports dynamic expansion of the atlas using high-quality third-generation data, enhancing the coverage of unknown or individual-specific structural variations;

[0174] 2. Quality control and construction strategies are optional, ensuring high accuracy of the spectral data and stability for downstream analysis;

[0175] 3. Built-in version management and logging improve system reliability and the traceability of analysis results;

[0176] 4. Provide a standardized, high-quality spectral foundation for downstream modules to ensure the accuracy and consistency of structural variation identification.

[0177] The sequencing data alignment and structural variation identification module receives user-input sequencing data and, based on a selected or constructed pan-genome map, performs map alignment and structural variation identification and initial screening. The module supports second-generation or third-generation sequencing data as independent input data types and employs appropriate data quality control strategies based on the characteristics of different sequencing data to improve the accuracy and stability of identifying complex structural variations. Its workflow includes the following steps:

[0178] Step 1: Sequencing Data Input and Preprocessing

[0179] This step is used to receive sequencing data input by the user and to implement differentiated quality control strategies based on the data type.

[0180] 1. The data access and classification management system supports the access of raw data (FASTQ / BAM) from multiple mainstream sequencing platforms and automatically identifies reads based on their characteristics:

[0181] (1) Second-generation short read data: high-throughput, high-accuracy short read data from Illumina, BGI and other sources.

[0182] (2) Three generations of long read length data: PacBio HiFi, PacBio CLR or Oxford Nanopore.

[0183] 2. The differentiated quality control strategy system executes the following preprocessing criteria based on the nature of the second-generation and third-generation data:

[0184] (1) Second-generation sequencing data quality control standards:

[0185] Base quality filtering: Remove reads with a Q20 (99% accuracy) ratio below 80%; remove bases with a quality value below 20 at the end of the read.

[0186] Read length filtering: Filter extremely short reads with a length of less than 50 bp to reduce the risk of multiple-mapping during map alignment.

[0187] Adapter sequence processing: rigorously identify and remove sequencing adapters and excessively repetitive adapter contamination sequences.

[0188] (2) Quality control standards for third-generation sequencing data:

[0189] For PacBio HiFi: Prioritize filtering read segments with a QV below 20 (i.e., accuracy below 99%). Due to HiFi's high accuracy, the system preferentially retains read segments of 10 kb or more to ensure its ability to traverse regions with complex structural variations.

[0190] For Nanopore (ONT) / CLR: Perform mean quality filtering (typically requiring Mean QV > 7-10); employ a length-weighted strategy to remove short segments less than 1000 bp in length and prioritize the retention of ultra-long reads to enhance the support strength of structural breakpoints.

[0191] Low complexity filtering: For highly repetitive regions in the pan-genome map, identify and label reads with extremely low sequence complexity (such as long A / T sequences) to prevent false positive signals during map alignment.

[0192] 3. Basic Quality Indicator Statistics: After preprocessing, the system generates a quality report in real time, including:

[0193] (1) Coverage distribution: Calculate the effective physical coverage depth relative to the pan-genome map path.

[0194] (2) Length distribution map: Statistical analysis of the N50 index of three generations of data to assess its potential for detecting large-scale SV.

[0195] (3) GC content distribution: to assess whether there is obvious bias in sequencing.

[0196] 4. Data Validation and Transmission Logic

[0197] (1) Consistency verification: For quality-controlled data directly input by users, the system ensures data integrity through MD5 verification and checks the FASTQ format specification.

[0198] (2) Traceability management: The system automatically records the Read loss rate before and after quality control and generates a unique metadata identifier (Metadata ID) to ensure that every identified structural variation can be traced back to the original sequencing physical signal.

[0199] Step 2: Sequencing data alignment based on pan-genome map

[0200] After completing the sequencing data preprocessing, the system performs map alignment on the sequencing data based on the selected or constructed pan-genome map. The system uses graph structure reference rather than non-linear reference for alignment, mapping the quality-controlled second- or third-generation sequencing data to the pan-genome map and generating the path representation of the sequencing reads in the map.

[0201] 1. Automatic adaptation of comparison mode

[0202] Based on the data type identified in step 1, the system automatically calls the corresponding map comparison engine:

[0203] (1) Second-generation data alignment (based on Haplotype guidance): The vg giraffe alignment algorithm is used. Principle: Using the haplotype path index in the pan-genome map, short reads are preferentially aligned to the most likely true path in the map.

[0204] (2) Three-generation data alignment (based on long segment alignment): GraphAligner or vg mpmap algorithm is used. Principle: Taking advantage of the characteristic that the three-generation read segments span multiple mutated nodes, the optimal non-linear alignment path is found in the graph topology.

[0205] 2. Real-time loading of the atlas index

[0206] Before the comparison begins, the system automatically mounts the "three-in-one" index file generated during the map optimization process:

[0207] (1) GBZ Index: Used for quick access to graph topology and haplotype information.

[0208] (2) Minimizer index: Performs rapid location of Seeds (seed sequences) to achieve the first round of coarse screening in the alignment process.

[0209] (3) Distance index: calculates the distance constraints between nodes, and assists the algorithm in selecting the best alignment path with the highest score among multiple branch paths through dynamic programming.

[0210] 3. Sequencing data alignment based on map alignment engine and map index file.

[0211] Step 3: Candidate structural variation identification and preliminary filtering

[0212] The system performs structural variant candidate identification based on the alignment results of sequencing data on the pan-genome map.

[0213] Structural variation identification is based on graph structural changes and alignment path differences. Its identification principle includes: using the path offset, breakage or cross-branch alignment behavior of sequencing reads in the graph to identify structural variations such as insertion, deletion, inversion, and duplication; and combining the coverage changes of nodes and edges, path support strength and graph structural topological differences to generate a candidate set of structural variations.

[0214] A unified graph structure analysis logic is used for both second- and third-generation sequencing data, with parameter adaptation only applied to support thresholds and evidence integration methods. Through this approach, the system achieves stable identification of complex structural variations without relying on a single linear reference. Then, bcftools is used to perform preliminary filtering on the obtained VCF files. The preliminary filtering criteria are: the FILTER column must meet the PASS condition, and only biallelic variants are considered.

[0215] Finally, the system passes the generated set of candidate structural variations to the subsequent intelligent filtering and annotation module for further refined screening and biological interpretation.

[0216] Technological advantages

[0217] 1. Supports independent access and analysis of second- or third-generation sequencing data, adapting to different experimental designs and data conditions;

[0218] 2. Alignment based on pan-genome maps effectively reduces linear reference bias and improves the ability to identify variations in complex regions;

[0219] 3. By identifying structural variations through graph structure paths and topological information, the ability to detect complex and atypical structural variations is enhanced;

[0220] 4. Provides a high-coverage, high-sensitivity set of structural variation candidates for subsequent machine learning / deep learning filtering modules.

[0221] The intelligent structural variant filtering module is used to automatically and intelligently filter and evaluate the reliability of candidate structural variants output by the structural variant identification module. Addressing the characteristics of a large number of candidate structural variants, complex types, and high noise levels in pan-genome maps, this module introduces machine learning or deep learning models to perform multi-dimensional feature modeling and scoring for each candidate structural variant, thereby distinguishing between true structural variants and false positive results introduced by sequencing errors, alignment anomalies, or graph structural complexity. Its specific working process includes the following steps:

[0222] Step 1: Extraction of structural variation features

[0223] The intelligent structural variation filtering module receives a set of candidate structural variations output by the structural variation identification module. For each candidate structural variation, the intelligent structural variation filtering module automatically extracts the core features affecting the variation confidence from the candidate VCF files. These features include the following categories:

[0224] 1. Basic quality characteristics: including quality score (QUAL), genotype quality value (GQ), and locus deletion rate.

[0225] 2. Supporting signal characteristics: Extract sequencing depth (DP), allele depth (AD, i.e., the number of readings supporting the variant), allele frequency (AF), and strand bias.

[0226] 3. Topological complexity features: Extract graph-specific indicators, such as the number of graph path branches where the mutation is located, the average alignment quality (MAPQ) of Read in this region, and whether it is located in a highly repetitive region (Repeat Masker).

[0227] 4. Path Consistency: For the mutation information obtained from graph generalization, a unique calculation is performed to determine whether all Reads supporting the mutation follow the same Haplotype path in the graph;

[0228] 5. Genotype characteristics: Extract genotype (GT) information and analyze whether it conforms to the distribution pattern of diploids.

[0229] 6. Structural variation types, including insertion (INS), deletion (DEL), inversion (INV), duplication (DUP), and translocation (TRA).

[0230] Step 2: Intelligent scoring or classification of structural variations

[0231] This step utilizes a pre-trained machine learning model to perform nonlinear mapping and weighted calculations on the multidimensional feature matrix extracted in step 1. Its core logic lies in addressing the high complexity and typological diversity of SVs in pan-genome maps by introducing domain prior knowledge to differentially weight the feature space. This resolves the technical shortcomings of traditional filtering algorithms, which suffer from extremely high false positive rates in complex graph topological regions and specific variant types (such as inversions and large fragment insertions). Discrete statistical indicators are transformed into standardized confidence scores, which are then used to perform automated filtering and attribution labeling.

[0232] 1. Building a machine learning model (training phase)

[0233] Before the system runs, a machine learning model is pre-built, and the building process includes:

[0234] (1) Gold standard alignment: The ground truth set published by the Genome in a Bottle Consortium (GIAB) and the Human Pangenome Reference Consortium (HPRC) was used as the positive sample, and the false positive signal of the high-noise region was used as the negative sample.

[0235] (2) Model training: An ensemble learning algorithm (such as XGBoost) is used as the core engine. The above 6 types of features are used as input vectors for supervised learning, and the contribution weight of each feature to the authenticity of the mutation is automatically learned.

[0236] (3) Domain Knowledge Enhancement: Type-specific penalty weights are introduced into the training loss function. To achieve higher-dimensional filtering accuracy, the model not only learns general quality metrics during training, but also constructs the following deep weighted decision mechanism by introducing domain heuristics, taking into account the complex topological environment unique to pan-genome maps and the physical signal differences of different structural variation types:

[0237] Topology and Path Weights: In the pan-genomic context, "path consistency" and "number of graph path bifurcations" are set as high-order weight features during model building. If a candidate variant is supported by reads, but its path logic seriously conflicts with the existing haplotype topology in the graph, the decision tree inside the model will perform a penalty deduction.

[0238] Differential type weighted scoring: The model applies a "specificity penalty" to different SV types. For large fragment insertions: The model automatically increases the weight of "Flanking Sequence Alignment Quality (MAPQ)". If there are many multi-site alignments on both sides of the inserted sequence, the model will produce a lower P-value. For inversions: The model focuses on monitoring "positive and negative strand signal mutation points". If there is a lack of bidirectional read support at the breakpoint (strand bias imbalance), the model's probability of identifying it as a false positive increases significantly.

[0239] 2. Model Application and Quantization Result Output

[0240] The system inputs the real-time feature vector extracted from the current test sample in step 1 into the model already constructed above, and performs the following automated operations:

[0241] (1) Probability score generation: The model receives 6 types of feature vector inputs, calculates them through an internal decision tree cluster, and finally outputs a P∈[0,1] for each candidate variant. This value represents the predicted probability that the variant is a "True Positive".

[0242] P --> 1: This means that the variant performs perfectly in terms of depth (DP), genotype consistency (GT / AB), and graph path topology.

[0243] P --> 0: This usually indicates that the variant has obvious chain bias, low alignment quality (MAPQ), or is located in an extremely complex repetitive sequence region.

[0244] (2) Generate feature contribution vectors: The model synchronously outputs labels of key factors leading to low scores, which are used to explain the underlying bioinformatics causes of low scores and assist in the refined classification in step 3. The label dictionary includes:

[0245] Low Path Consistency (LPC): Extremely low path consistency, indicating that variant reads cannot form a closed and continuous logical flow in haplotype paths in the graph.

[0246] High Strand Bias (HSB): Extreme imbalance in the alignment of positive and negative strands, often seen in false positives of inversion (INV).

[0247] RepMaskerConflict (RMC): The variant is located in a known highly repetitive region and is accompanied by multi-site alignment signals.

[0248] AbnormalInsertSize (AIS): The insertion / missing length is statistically significantly different from the physical step size after Read alignment.

[0249] LowMAPQCluster (LMC): The average alignment quality (MAPQ) of the mutation breakpoint and its flanking regions is below the safety threshold.

[0250] Step 3: Threshold-based structural variation filtering

[0251] This step receives the probability score P and feature contribution vector output from step 2, and combines them with the prior physical properties of the genomic region to implement dynamic threshold interception and classification labeling, thereby realizing the transformation from "probability value" to "deterministic filtering result".

[0252] 1. Dynamic Threshold Decision Logic Based on Regional Attributes: The system does not use a fixed threshold, but instead calls different decision operators based on the genomic coordinates of the variant.

[0253] (1) Core functional areas (such as CDS / Exon): The “sensitivity priority” strategy is adopted. The judgment logic is: if P > 0.5, it is initially judged as PASS, ensuring that candidate variants in important functional areas are not missed.

[0254] (2) Highly repetitive and complex regions (Repeat Masker): For regions with high topological complexity, a "specificity priority" strategy is adopted. The judgment logic is: only when P > 0.85 is it judged as PASS, so as to eliminate algorithm noise caused by redundancy in the graph structure.

[0255] 2. Hard Filters Secondary Verification: Based on the intelligent scoring, the system performs physical filtering with a "one-vote veto" system. Even if the P-value in step 2 is high, blocking will still be performed if any of the following physical conditions are met:

[0256] (1) Missing rate filtering: If the missing rate of a site exceeds the set threshold (e.g., > 20%), it will be marked as unreliable even if the score is acceptable.

[0257] (2) Very low depth filtering: For third-generation sequencing data, if the number of effective reads (AD) supporting the variant is too low, interception is performed directly.

[0258] 3. Result Classification Labeling and Conflict Reduction: The system extracts the feature contribution vector from step 2 and labels each non-PASS site with a specific filtering attribution.

[0259] (1) PASS: High-quality variants that conform to diploid distribution and meet the P-value criteria, with consistent pathways.

[0260] (2) Non-PASS: LPC, HSB, RMC, AIS, LMC.

[0261] Step 4: Output a high-quality set of structural variations

[0262] The high-quality structural variants that have undergone intelligent filtering are integrated into a structural variant result set, and then normalized and output to subsequent modules.

[0263] 1. Output format: Generate a standard VCF v4.2 format file and record the model scoring results in the INFO field.

[0264] 2. Statistical Summary: Simultaneously outputs SV type distribution plots, length distribution histograms, and statistics on common / specific variations among samples, providing high-quality basic data for subsequent population genetic analysis or clinical applications.

[0265] Technological advantages

[0266] 1. By leveraging multidimensional features and intelligent models, the false positive rate of traditional rule-based filtering is significantly reduced;

[0267] 2. The numerical threshold is configurable to adapt to different sequencing depths and experimental conditions;

[0268] 3. Supports continuous scoring output, facilitating subsequent GWAS, QTL, and other analyses by weighting according to reliability;

[0269] 4. Deeply coupled with pan-genome mapping to improve the ability to distinguish complex structural variations.

[0270] Example 2

[0271] Embodiment 2 of the present invention provides a system corresponding to Embodiment 1 above, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program in the memory to implement the steps of the method of Embodiment 1 above.

[0272] In some implementations, the memory may be high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk storage device.

[0273] In other implementations, the processor can be any type of general-purpose processor, such as a central processing unit (CPU) or a digital signal processor (DSP), and there is no limitation here.

[0274] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0275] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for identifying structural variations based on pan-genome map, characterized by, Includes the following steps: S1. Expand / enhance the pan-genome map using third-generation sequencing data; S2. Receive the input sequencing data, compare the sequencing data with the expanded / enhanced pan-genome map to obtain the graph reference alignment result, and based on the graph reference alignment result, identify structural variations in the sequencing data and generate a candidate structural variation set. S3. A pre-trained deep learning model is used to evaluate and classify the credibility of the candidate structural variant set to obtain a filtered structural variant set.

2. The pan-genome graph-based structural variant calling method of claim 1, wherein, The specific implementation process of step S1 includes: Determine whether the quality of the third-generation sequencing data meets the requirements for map construction or expansion. If so, use a map construction strategy based on genome assembly to expand the pan-genome map, or use a structural variation identification method based on reference genome or pan-genome alignment to generate map increment information; otherwise, directly use the original pan-genome map as the expanded pan-genome map.

3. The structural variation identification method based on pan-genome mapping according to claim 2, characterized in that, The specific implementation process for determining whether the quality of third-generation sequencing data meets the requirements for map construction or expansion, and for using a genome assembly-based map construction strategy to expand the pan-genome map, or for using a structural variation identification method based on reference genome or pan-genome alignment to generate incremental map information, includes: When the coverage depth of PacBio HiFi data is not less than 30×, meets the sequencing accuracy assessment, and the read length N50 value is not less than 20kb, a map construction strategy based on genome assembly is used to expand the pan-genome map; when the coverage depth of PacBio HiFi data is less than 30×, a structural variation identification method based on reference genome or pan-genome alignment is used to generate incremental map information. When the coverage depth of PacBio CLR or ONT data is not less than 50×, meets the sequencing accuracy assessment, and the read length N50 value is not less than 20kb, a map construction strategy based on genome assembly is used to expand the pan-genome map; otherwise, a structural variation identification method based on reference genome or pan-genome alignment is used to generate incremental map information.

4. The structural variation identification method based on pan-genome mapping according to claim 1, characterized in that, The specific implementation process of step S2 includes: Preprocess the sequencing data; The preprocessed sequencing data is mapped onto the expanded / enhanced pangenome map to generate the path representation of the sequencing reads in the map; By utilizing the path offset, breakage, or cross-branch alignment behavior of sequencing reads in the graph, we can identify structural variations such as insertions, deletions, inversions, and duplications. By combining the coverage changes of nodes and edges, path support strength, and graph topological differences, we can generate a candidate set of structural variations.

5. The structural variation identification method based on pan-genome mapping according to claim 1, characterized in that, The specific implementation process of step S3 includes: The core features are extracted from the candidate set of structural variants, and the extracted core features are used as input to a pre-trained deep learning model to output the probability value and feature contribution vector of each candidate variant. If a candidate mutation is located in the core functional area and the probability value of the candidate mutation is greater than the first set threshold, it is determined to pass; if a candidate mutation is located in the highly repetitive and complex area and the probability value of the candidate mutation is greater than the second set threshold, it is determined to pass; wherein, the second threshold is greater than the first threshold. If the site deletion rate of a candidate variant that is deemed to have passed exceeds a set threshold, it is deemed to have failed. Using the aforementioned feature contribution vector, specific filtering attributions are labeled for each candidate variant that was not determined to pass.

6. The structural variation identification method based on pan-genome mapping according to claim 5, characterized in that, If the input data received in step S2 is third-generation sequencing data, the following judgment is also made: if the number of valid reads supporting the candidate variant is lower than the set threshold, it is determined to fail.

7. The structural variation identification method based on pan-genome mapping according to claim 5, characterized in that, The core features include basic quality features, supporting signal features, topological complexity features, path consistency, genotype features, and structural variation types; the basic quality features include quality score, genotype quality value, and site deletion rate; the supporting signal features include sequencing depth, allele depth, and positive / negative strand support bias. The topological complexity features include the number of graph path branches where the mutation occurs, the average alignment quality of Read in that region, and whether it is located in a highly repetitive region; the structural mutation types include insertion, deletion, inversion, duplication, and translocation.

8. The structural variation identification method based on pan-genome mapping according to claim 5, characterized in that, The feature contribution vector includes key factor labels, which include: variant reads cannot form a closed and continuous logical flow in the haplotype path in the graph; false positives of inversion; variants located in known highly repetitive regions and accompanied by multi-site alignment signals; the number of graph node bifurcations exceeds a preset threshold, resulting in an explosion of the alignment path search space; there is a statistically significant difference between the insertion / deletion length and the physical step size after Read alignment; and the average alignment quality of the variant breakpoint and its flanking regions is lower than the safety threshold.

9. A structural variation identification system based on pan-genome mapping, comprising a memory, a processor, and a computer program stored in the memory; characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Variation detection and typing method, system and equipment based on third-generation sequencing data and generic genome graph structure and medium

    CN118692559A

  • Method for carrying out strain level classification on metagenome data based on generic genome graph

    CN118866126A

  • Identification quantification and genome visualization method for integron-drug resistance gene cassette in environmental sample

    CN119132409A

  • Method for identifying fish genome structure variation based on graph genome and application

    CN119724338A

  • Structural variation detection algorithm, system and equipment based on third-generation sequencing data and generic genome and medium

    CN120048341A