Structure variation analysis-oriented genome data visualization system and method
The genome data visualization system solves the problem of displaying read span structural variations and coverage anomalies in graphical pan-genomes, and realizes the integrated display of multi-level information and efficient population variation analysis.
Patent Information
- Application Number
- CN202510921948.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies cannot directly view how the reads of each sample in a population cross or support specific structural variations on the graph structure, and lack tools to intuitively display areas of abnormal coverage.
This paper presents a genomic data visualization system for structural variation analysis, including modules for data input, coordinate transformation, indexing, graph-alignment extraction, functional annotation visualization, coverage analysis, and read visualization. It achieves effective connection between graphical pan-genome and linear reference genome, and displays the functional impact, in-depth support, read evidence, and population frequency of structural variations through multi-level information integration.
It enables interactive graphical display of segment comparison results, improves operational intuitiveness and compatibility, allows observation of multi-level information on structural variations on a unified platform, and enhances the efficiency and accuracy of structural variation analysis at the population scale.
Smart Images

Figure CN120977393A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, and in particular to a genome data visualization system and method for structural variation analysis. Background Technology
[0002] With the development of high-throughput sequencing and the deepening of research on genetic variation, the analysis of genomic structural variations (SVs) has received increasing attention. Structural variations include large-segment insertions, deletions, inversions, and translocations, and their impact on the genome and phenotype is more complex than that of single-base mutations. Traditional variation analysis typically uses a single reference genome for alignment and variation detection, but this linear reference framework has biases: missing sequences in the reference genome lead to incomplete alignments, failing to fully represent the presence / deletion variations (PAVs) in the population. To overcome the limitations of a single reference, the concept of pan-genomes has been proposed in recent years, which aggregates all genomic sequences of a species or population to construct a reference genome containing all genetic variations. In particular, through graphical data structures, graphical pan-genomes can use graph nodes to represent sequence segments and edges to represent the connections between different sequences, thereby organically linking reference sequences with variant sequences to comprehensively display the genetic diversity within a species.
[0003] However, the application of graphical pangenomes still faces many challenges. Among these, the visualization of graphical pangenome data is a major bottleneck. Most existing bioinformatics tools only support linear reference genomes and lack support for complex graph structures, necessitating the development of more specialized algorithms and tools for downstream analysis of graphical pangenomes. In particular, a mature solution for the correspondence between graph coordinates and linear reference coordinates is still lacking. Due to the lack of a unified coordinate mapping system, it is currently difficult to accurately locate fragments in the graph back to their original genome positions. This imperfection in the coordinate system makes it difficult to associate graph structures with traditional genome annotation, further hindering the intuitive interpretation of graphical pangenomes.
[0004] Furthermore, existing methods have significant shortcomings in the visualization and analysis of structural variations. On the one hand, traditional genome browsing tools (such as IGV) are primarily designed for linear reference sequences and cannot directly handle pan-genome graph structures with complex branching, thus failing to display the alignment details of reads on the graph structure. On the other hand, there is a lack of tools that support the visualization of structural variations at the population level. Specifically, for alignment results from multiple individuals, there is a lack of interactive visualizations to help researchers understand the distribution frequency of variations within the population and the read support. This limitation hinders researchers from gaining a deeper understanding of structural variations and their functional impact.
[0005] Currently, while a few studies have begun exploring visualization tools for graphical pangenomes, they still cannot fully meet the aforementioned needs. For example, a web-based interactive visualization framework (VRPG) exists, which provides efficient and intuitive support for users to explore and annotate graphical pangenomes in a linear genomic coordinate system by projecting the graphical pangenome onto a linear reference coordinate system. This tool has functions such as highlighting specific genomic paths in the graphical pangenome, representing node copy numbers, and providing sequence alignment queries. It can also display the graphical pangenome alongside the feature annotations of the traditional reference genome, thus achieving a seamless connection between the graphical structure and the linear genomic perspective. However, such tools mainly focus on reference sequence projection and annotation integration, and do not yet provide complete solutions for the visualization of read alignment information and population variation frequency analysis. Specifically, it is still not possible to directly view how reads of each sample in the population cross or support specific structural variations on the graphical structure, and there is a lack of intuitive display of regions with abnormal coverage. Summary of the Invention
[0006] This invention provides a genomic data visualization system and method for structural variation analysis, which solves the problems in the prior art that it is impossible to directly view how the reads of each sample in the population cross or support specific structural variations on the graph structure, and that it is impossible to intuitively display regions with abnormal coverage.
[0007] On one hand, embodiments of the present invention provide a genome data visualization system for structural variation analysis, comprising: The data input module is used to import structural data, read alignment data, and genome function annotation data from graphical pangenome maps. The coordinate transformation module is used to map the coordinates of the graph structure to the coordinates of the linear reference genome; An indexing module is used to create an index between the graph structure and the annotation file of the linear reference genome; The image-comparison and extraction module is used to read the comparison data and extract the comparison information of the corresponding target area. The functional annotation visualization module is used to display the genes, transcripts and transposon elements in the corresponding target region on a graphical genome according to the annotation file; The coverage analysis module is used to statistically calculate the coverage depth of the read segments along the sequence position of the target region; The segment visualization module is used to display segment comparison details within the target area on the user interface; The population structure variation frequency visualization module is used to statistically analyze the occurrence frequency of each structural variation in the target region in the sample population based on comparison data from multiple samples, and to display the population frequency of different variations in the form of a visualization chart. The data input module is data-connected to the coordinate transformation module, the index module, and the graph-comparison extraction module; the coordinate transformation module, the index module, and the graph-comparison extraction module are respectively data-connected to the functional annotation visualization module, the coverage analysis module, and the segment visualization module; the functional annotation visualization module and the coverage analysis module are jointly data-connected to the population structure variation frequency visualization module; the coverage analysis module and the segment visualization module are jointly data-connected to the population structure variation frequency visualization module.
[0008] In one possible implementation, the coordinate transformation module utilizes pre-stored reference path mapping information and node coordinate system to precisely locate the node set of the graphical pan-genome to the linear reference genome; the coordinate transformation module synchronously updates the corresponding linear reference coordinates as the version of the graph structure genome is updated.
[0009] In one possible implementation, the indexing module associates the nodes and edges of the graph structure with annotation information for genes, transcripts, and transposable elements. The indexing module is used to quickly query the functional and structural annotations of a specified graph region.
[0010] In one possible implementation, the graph-alignment extraction module is further configured to identify branch nodes and mutation paths in the graph structure that involve the corresponding target region, and the graph-alignment extraction module is further configured to obtain the read sequence supporting each path and its alignment details; The graph-alignment extraction module uses the VG variant graph tool to query the GAM format graph alignment file. When the input alignment data is in BAM format, it filters reads by the coordinates of the linear reference genome and combines the coordinate mapping of the graph structure to extract the set of aligned reads for the target region.
[0011] In one possible implementation, the functional annotation visualization module is also used to provide functional annotations for regions of structural variation.
[0012] In one possible implementation, the coverage analysis module generates a coverage distribution map based on the coverage depth of the read segment, and the coverage distribution map is used to identify the variation hotspot intervals of abnormal coverage changes.
[0013] In one possible implementation, the segment visualization module further includes: Draw the segments aligned with their comparison positions; Highlight the comparison features of mismatches, insertions, missing segments, and read segments at the branches of the graph structure.
[0014] On the other hand, embodiments of the present invention provide a method for visualizing genomic data for structural variation analysis, including: Import graphical pangenome data, read alignment files, and genome function annotation files through the data input module; Specify a target genomic region, and use the coordinate transformation module to perform graph-to-linear coordinate transformation to map the target genomic region to the corresponding region in the pan-genome graph; The index pre-built by the index module is used to retrieve gene, transcript, and transposon element annotation information within the region; The graph-alignment extraction module is invoked to process the read segment alignment data and obtain read segment alignment information on the region and its branch mutation paths. The distribution of genes, transcripts, and transposon elements within the region is plotted based on the annotation information from the functional annotation visualization module. The coverage analysis module calculates the coverage depth of the read segments along the sequence in the region and generates a coverage curve, marking segments with abnormal coverage. The read segment visualization module displays a detailed view of the read segment comparison within the region, intuitively presenting the comparison characteristics of the read segments at the mutation sites; The frequency of structural variations in the region in different samples is statistically analyzed using the population structure variation frequency visualization module, and the population frequency distribution of each variation is displayed in the form of a chart.
[0015] The genomic data visualization system and method for structural variation analysis disclosed in this invention have the following advantages: (1) The alignment results of reads on the graph are presented in an interactive graphical manner, which overcomes the shortcomings of existing linearized genome browsers that cannot display graph structure alignment information.
[0016] (2) Through coordinate transformation and indexing mechanisms, the graph genome is effectively linked with the traditional linear reference genome. Users can specify regions of interest using familiar reference coordinates, and the system automatically maps them to the graph structure and synchronizes the relevant functional annotations, thereby greatly improving the intuitiveness and compatibility of the operation.
[0017] (3) Integrate multi-level information such as functional annotation, coverage depth, read alignment, and population frequency into the same analysis interface. Users can simultaneously observe the functional impact of structural variations (gene / transposon element annotation), depth support (coverage map), read evidence (alignment view), and population statistics (variation frequency) on a unified platform without switching between different tools, and gain a comprehensive and in-depth understanding of the variations.
[0018] (4) The difference in the frequency of structural variation among populations helps to distinguish between high-frequency and rare variations, greatly improving the efficiency and accuracy of analyzing structural variation at the population scale. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A data flow structure diagram of a genome data visualization system for structural variation analysis provided in this application embodiment; Figure 2 This is a schematic diagram of a genome data visualization method for structural variation analysis provided in an embodiment of this application. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Figure 1 This is a flowchart illustrating a genome data visualization system for structural variation analysis provided in an embodiment of the present invention. The embodiment of the present invention provides a genome data visualization system for structural variation analysis, comprising: The data input module is used to import structural data, read alignment data, and genome function annotation data from graphical pangenome maps. The coordinate transformation module is used to map the coordinates of the graph structure to the coordinates of the linear reference genome; An indexing module is used to create an index between the graph structure and the annotation file of the linear reference genome; The image-comparison and extraction module is used to read the comparison data and extract the comparison information of the corresponding target area. The functional annotation visualization module is used to display the genes, transcripts and transposon elements in the corresponding target region on a graphical genome according to the annotation file; The coverage analysis module is used to statistically calculate the coverage depth of the read segments along the sequence position of the target region; The segment visualization module is used to display segment comparison details within the target area on the user interface; The population structure variation frequency visualization module is used to statistically analyze the occurrence frequency of each structural variation in the target region in the sample population based on comparison data from multiple samples, and to display the population frequency of different variations in the form of a visualization chart. The data input module is data-connected to the coordinate transformation module, the index module, and the graph-comparison extraction module; the coordinate transformation module, the index module, and the graph-comparison extraction module are respectively data-connected to the functional annotation visualization module, the coverage analysis module, and the segment visualization module; the functional annotation visualization module and the coverage analysis module are jointly data-connected to the population structure variation frequency visualization module; the coverage analysis module and the segment visualization module are jointly data-connected to the population structure variation frequency visualization module.
[0023] The coordinate transformation module utilizes pre-stored reference path mapping information and node coordinate system to accurately locate the node set of the graphical pan-genome to the linear reference genome; the coordinate transformation module synchronously updates the corresponding linear reference coordinates as the version of the graph structure genome is updated.
[0024] The indexing module associates the nodes and edges of the graph structure with annotation information of genes, transcripts, and transposable elements. The indexing module is used to quickly query the functional and structural annotations of a specified graph region.
[0025] The graph-alignment extraction module is also used to identify branch nodes and mutation paths in the graph structure that involve the corresponding target region. The graph-alignment extraction module is also used to obtain the read sequence that supports each path and its alignment details. The graph-alignment extraction module uses the VG variant graph tool to query the GAM format graph alignment file. When the input alignment data is in BAM format, it filters reads by the coordinates of the linear reference genome and combines the coordinate mapping of the graph structure to extract the set of aligned reads for the target region.
[0026] The functional annotation visualization module is also used to provide functional annotations for structural variation regions.
[0027] The coverage analysis module generates a coverage distribution map based on the coverage depth of the read segment. The coverage distribution map is used to identify the variation hotspot intervals of abnormal coverage changes.
[0028] The segment visualization module also includes: Draw the segments aligned with their comparison positions; Highlighting variations such as mismatches, insertions, and missing segments, as well as the comparison features of read segments at graph structure branches.
[0029] For example, such as Figure 1 As shown, the structural variation visualization analysis tool of this application includes multiple functional modules, which work sequentially according to a workflow to collaboratively achieve the visualization analysis of pan-genome structural variations. This embodiment describes the specific working process of each module and the interaction between them.
[0030] First, the data input module reads the user-provided graphical pan-genome data and related files. Specifically, the pan-genome map can be in GFA / rGFA format or the internal map format of the VG tool, describing the relationship between the reference genome sequence and various structural variation sequence fragments. Read alignment data can be either alignment files for linear references (BAM / CRAM format) or alignment files for graphical references (GAM / GAF format); in this embodiment, the GAM format generated by the VG tool is preferred to obtain accurate alignment information of reads on the map. In addition, functional annotation data, such as gene and transposon element annotation files (e.g., GFF3 format), are also loaded through the data input module. The data input module performs integrity checks and format parsing on the above-mentioned data types and loads the graphical structured genome into memory for later use.
[0031] Next, the coordinate transformation module comes into play when the user specifies the analysis region. When the user specifies the genomic region of interest (e.g., a range of chromosome positions in the reference genome) using linear reference coordinates, the coordinate transformation module finds the corresponding part of that region in the graph structure. The construction of the graphical pan-genome preserves information about the reference sequence paths (e.g., Minigraph records the path of each input genome when constructing the graph). The coordinate transformation module uses this path information, along with pre-established node coordinate indices, to map the input linear coordinates to the set of nodes and edges in the graph. For reference sequence regions spanning multiple nodes, the module locates the start and end nodes and selects all possible graph paths that traverse that linear region. Conversely, when the user selects a subgraph or node branch in the graph structure, the module can also convert it back to the corresponding linear reference coordinate range. Through this bidirectional mapping mechanism, the coordinate transformation module bridges the gap between the graphical pan-genome and the linear reference, allowing users to easily find target sites and information in both coordinate systems.
[0032] After coordinate mapping is completed, the system uses an index module to quickly retrieve functional annotation information for the target region. The index module establishes a mapping relationship between the graph structure and genome annotations during the initialization phase. One feasible implementation is to record the coordinate range of the reference sequence or the corresponding gene number for each node in the graph, or conversely, to store a list of corresponding graph node IDs for each annotation record (such as the start and end positions of genes). With this index, given a range in the graph structure, the index module can immediately find annotation entries for genes, exons, transcripts, or transposable elements covered by that region. Thus, when a user selects a target region, the system can quickly summarize the functional elements within that region, preparing for subsequent visualization.
[0033] The graph-alignment extraction module is the core mechanism for extracting alignment information from read segments. This module receives the target region definition from the coordinate transformation module and calls VG tools or other graph alignment processing libraries to filter and extract the aligned read segment data. For linear alignment data in BAM format, the module can also extract read segments within the linear coordinate interval (e.g., using samtools to extract the alignment of that region) and then infer the mapping positions of these read segments in the graph structure based on the coordinate transformation results. The graph-alignment extraction module stores the extracted read segment list and its alignment location results in a memory data structure and notifies the downstream visualization module to update it.
[0034] Once the alignment data is ready, the functional annotation visualization module and the coverage analysis module will start in parallel to generate an analysis view. The functional annotation visualization module reads the annotation information provided by the index module and the reference interval boundaries provided by the coordinate transformation module, and plots a schematic diagram of the distribution of genomic functional elements within the region on the user interface. Simultaneously, the coverage analysis module uses the read alignment information provided by the graph-alignment extraction module to calculate the depth coverage of the region. Specifically, the module uses the reference coordinates as a reference, divides the target interval into continuous windows (the window size can be a single base or a larger step size depending on the analysis requirements), and counts the number of covered reads at each position. Based on this, a line graph or histogram of coverage variation with genomic position is plotted. Through this graph, users can easily identify abnormal regions: for example, if a region has almost no read coverage, it may indicate the presence of deletion variants in some individuals; conversely, if there is a surge in coverage, it may represent copy number differences or insertions. The coverage analysis module also combines information on variant branches to mark and highlight these abnormal intervals, indicating to users that these regions are potential variant hotspots, helping users locate structural variant regions of interest.
[0035] In addition to the visual view, this application also provides a detailed display of alignment details. After receiving the alignment set of reads from the target region, the read visualization module generates a local read alignment view on the interface. This module can call tools like samtools or graphsamtools, which are specifically extended for graph structures, to align and draw reads within the selected region according to their alignment positions. In implementation, the system can select a reference main path (or any user-specified reference path) as a baseline, arranging reads on vertically stacked tracks according to their alignment positions on that path, with each read represented by a single line. When discrepancies exist in read alignments, the module uses intuitive symbols to mark them: for example, different colored letters represent mismatched bases, gaps indicate missing sequences relative to the reference in the read, and small inserted arrows or extra boxes indicate inserted sequences in the read. For alignment interruptions occurring at graph branches, the module can display a fork marker on the reference track, indicating the node where the read transitions from the main path to the variant branch.
[0036] For population studies, focusing on the frequency of variations across different individuals is equally crucial. The population structural variation frequency visualization module utilizes alignment results from multiple samples to statistically analyze each type of structural variation within the target region. Assume the user-imported read alignment data contains data from multiple samples (which can be distinguished by labels or separate storage). The frequency statistics module first identifies the dominant variation type or variation path present in the target region (e.g., an approximately 1kb insertion sequence as a branch path, or a missing reference sequence in some samples forming a missing path). Then, the module checks each sample's alignment to see if it supports the variation: for example, if several reads in a sample's read alignment enter the insertion branch path, the sample is considered to carry the insertion variation; conversely, if all reads in the sample align continuously along the main reference path without showing an insertion, the sample is considered to lack the insertion variation. By statistically analyzing the support for each variation across all samples using a similar method, the module calculates the population frequency of the variation (the proportion of samples supporting the variation out of the total number of samples). Finally, the frequency visualization module presents the results in an intuitive format, such as bar charts or pie charts: each structural variation corresponds to a legend, indicating its percentage frequency in the population, or multiple parallel bars represent the occurrence of different variations in individual samples. This visualization allows users to quickly distinguish which structural variations are common (high-frequency variations, possibly species-shared structures or core sequence differences) and which are rare or unique (low-frequency variations, possibly associated with specific subgroups or traits). This provides important reference for subsequent population genetic analysis and association analysis.
[0037] Through the collaborative work of the above modules, this application provides a complete graphical pan-genome structural variation analysis workflow. When using this tool, users only need to follow these steps: import data → select target region → view visualization results. For example, researchers want to analyze the structural variations near a known tolerance-related gene in a crop. They input the coordinate range of the gene on the reference genome into the tool, and the coordinate transformation module locates it on the pan-genome map. The system then extracts reads from that region for alignment: the results show an insertion sequence of approximately 500 bp at the second exon of the gene (represented as a branching path in the graph structure). The functional annotation view indicates that the insertion is located in an intron region within the gene, while the coverage analysis plot shows that at the location corresponding to the insertion, the read coverage decreases in most samples, but several samples show additional read coverage peaks. The read view further confirms that in samples carrying the insertion, some reads branch into the insertion sequence at this location and align, while the reads in other samples directly match the upstream and downstream exon junctions across the insertion location. The population frequency plot indicates that this insertion variation exists in approximately 30% of the samples.
[0038] Figure 2 This invention provides a flowchart of a genome data visualization system for structural variation analysis, as illustrated in an embodiment of the invention. The invention also provides a genome data visualization method for structural variation analysis, comprising: Import graphical pangenome data, read alignment files, and genome function annotation files through the data input module; Specify a target genomic region, and use the coordinate transformation module to perform graph-to-linear coordinate transformation to map the target genomic region to the corresponding region in the pan-genome graph; The index pre-built by the index module is used to retrieve gene, transcript, and transposon element annotation information within the region; The graph-alignment extraction module is invoked to process the read segment alignment data and obtain read segment alignment information on the region and its branch mutation paths. The distribution of genes, transcripts, and transposon elements within the region is plotted based on the annotation information from the functional annotation visualization module. The coverage analysis module calculates the coverage depth of the read segments along the sequence in the region and generates a coverage curve, marking segments with abnormal coverage. The read segment visualization module displays a detailed view of the read segment comparison within the region, intuitively presenting the comparison characteristics of the read segments at the mutation sites; The frequency of structural variations in the region in different samples is statistically analyzed using the population structure variation frequency visualization module, and the population frequency distribution of each variation is displayed in the form of a chart.
[0039] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0040] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A genome data visualization system for structural variation analysis, characterized in that, include: The data input module is used to import structural data, read alignment data, and genome function annotation data from graphical pangenome maps. The coordinate transformation module is used to map the coordinates of the graph structure to the coordinates of the linear reference genome; An indexing module is used to create an index between the graph structure and the annotation file of the linear reference genome; The image-comparison and extraction module is used to read the comparison data and extract the comparison information of the corresponding target area. The functional annotation visualization module is used to display the genes, transcripts and transposon elements in the corresponding target region on a graphical genome according to the annotation file; The coverage analysis module is used to statistically calculate the coverage depth of the read segments along the sequence position of the target region; The segment visualization module is used to display segment comparison details within the target area on the user interface; The population structure variation frequency visualization module is used to statistically analyze the occurrence frequency of each structural variation in the target region in the sample population based on comparison data from multiple samples, and to display the population frequency of different variations in the form of a visualization chart. The data input module is data-connected to the coordinate transformation module, the index module, and the graph-comparison extraction module; the coordinate transformation module, the index module, and the graph-comparison extraction module are respectively data-connected to the functional annotation visualization module, the coverage analysis module, and the segment visualization module; the functional annotation visualization module and the coverage analysis module are jointly data-connected to the population structure variation frequency visualization module; the coverage analysis module and the segment visualization module are jointly data-connected to the population structure variation frequency visualization module.
2. The genome data visualization system for structural variation analysis according to claim 1, characterized in that, The coordinate transformation module utilizes pre-stored reference path mapping information and node coordinate system to accurately locate the node set of the graphical pan-genome to the linear reference genome; the coordinate transformation module synchronously updates the corresponding linear reference coordinates as the version of the graph structure genome is updated.
3. The genome data visualization system for structural variation analysis according to claim 1, characterized in that, The indexing module associates the nodes and edges of the graph structure with annotation information of genes, transcripts, and transposable elements. The indexing module is used to quickly query the functional and structural annotations of a specified graph region.
4. The genome data visualization system for structural variation analysis according to claim 1, characterized in that, The graph-alignment extraction module is also used to identify branch nodes and mutation paths in the graph structure that involve the corresponding target region. The graph-alignment extraction module is also used to obtain the read sequence that supports each path and its alignment details. The graph-alignment extraction module uses the VG variant graph tool to query the GAM format graph alignment file. When the input alignment data is in BAM format, it filters reads by the coordinates of the linear reference genome and combines the coordinate mapping of the graph structure to extract the set of aligned reads for the target region.
5. A genome data visualization system for structural variation analysis according to claim 1, characterized in that, The functional annotation visualization module is also used to provide functional annotations for structural variation regions.
6. A genome data visualization system for structural variation analysis according to claim 1, characterized in that, The coverage analysis module generates a coverage distribution map based on the coverage depth of the read segment. The coverage distribution map is used to identify the variation hotspot intervals of abnormal coverage changes.
7. A genome data visualization system for structural variation analysis according to claim 1, characterized in that, The segment visualization module also includes: Draw the segments aligned with their comparison positions; Highlight mismatches, insertions, missing variations, and alignment features of read segments at graph structure branches.
8. A method for visualizing genomic data for structural variation analysis, characterized in that, include: Import graphical pangenome data, read alignment files, and genome function annotation files through the data input module; Specify a target genomic region, and use the coordinate transformation module to perform graph-to-linear coordinate transformation to map the target genomic region to the corresponding region in the pan-genome graph; The index pre-built by the index module is used to retrieve gene, transcript, and transposon element annotation information within the region; The graph-alignment extraction module is invoked to process the read segment alignment data and obtain read segment alignment information on the region and its branch mutation paths. The distribution of genes, transcripts, and transposon elements within the region is plotted based on the annotation information from the functional annotation visualization module. The coverage analysis module calculates the coverage depth of the read segments along the sequence in the region and generates a coverage curve, marking segments with abnormal coverage. The read segment visualization module displays a detailed view of the read segment comparison within the region, intuitively presenting the comparison characteristics of the read segments at the mutation sites; The frequency of structural variations in the region in different samples is statistically analyzed using the population structure variation frequency visualization module, and the population frequency distribution of each variation is displayed in the form of a chart.