Somatic Structural Variant Detection Method Based on a Hybrid Model of Single-Molecule Long-Read Sequencing Maps
Local graph genomes were constructed through single-molecular long read and long sequencing technology and partial order alignment graph algorithm, and cluster analysis was performed in combination with multi-category mixed models, which solved the alignment error and genomic heterogeneity problems of short read and long sequencing in somatic structural variation detection, achieving higher accuracy and coverage variation detection.
Patent Information
- Application Number
- CN202510120346.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-01-25
AI Technical Summary
Existing short-read and long-sequencing technologies are difficult to accurately detect somatic structural variations, especially in the case of large fragment variations, complex regions and genomic heterogeneity, which have problems with alignment errors and insufficient detection accuracy.
Single-molecule long read and long sequencing technology is used, combined with partially ordered alignment graph algorithm (sPOA) and multi-category hybrid model, somatic cell structural variant categories are identified and generated through local graph genome construction and clustering analysis, and consensus sequences are generated.
It significantly improves the detection accuracy and coverage of somatic structural variation, can accurately identify low-frequency and heterogeneous variations, reduce errors, and provide more comprehensive genomic information.
Smart Images

Figure CN120032711B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of genetics and bioinformatics, and particularly to a method for detecting somatic structural variations based on single-molecule long-read sequences. Background Art
[0002] Somatic Structural Variants (somatic SVs) refer to structural variations that occur in the somatic cell genome, usually including types such as insertions, deletions, duplications, and transversions. These variations often have important impacts on aspects such as an individual's health status, disease occurrence, and tumor development. Different from germline variations, somatic variations only exist in some cells and usually do not have heredity. The accurate detection of somatic structural variations is crucial for fields such as tumor genomics research, disease diagnosis, and the discovery of therapeutic targets.
[0003] However, the accurate detection of somatic structural variations faces many challenges, mainly due to low genotype frequencies, genomic heterogeneity, and limitations of existing sequencing technologies. Most existing methods for detecting somatic structural variations rely on short-read sequencing technologies, which have the following defects when analyzing somatic structural variations:
[0004] Limitations of short-read sequencing: Short-read sequencing technologies can only generate relatively short sequence fragments (usually 50 - 150 bases), and there are significant limitations for detecting somatic structural variations, especially those involving large fragments (such as large-scale deletions or insertions, etc.). This is because short reads cannot span the variant regions, making it difficult to fully analyze the structural variations.
[0005] Alignment errors: Short-read sequencing relies on sequence alignment to align the sequencing fragments with the reference genome to determine the variant positions. However, due to the existence of genomic heterogeneity and low-complexity sequences, errors may occur during the alignment process, leading to misjudgment of structural variations. Especially in complex regions (such as repetitive sequences or low-complexity sequences), the alignment accuracy is low, which may result in the loss or misidentification of variant information.
[0006] Genomic heterogeneity: Somatic structural variations usually only exist in specific cell populations, so there is genomic heterogeneity among different cell populations. This heterogeneity makes it impossible to accurately distinguish the source of variations through traditional sequence alignment methods, thus affecting the detection accuracy of variations.
[0007] To overcome the limitations of short-read sequencing, single molecular long-read sequencing technology, as an emerging high-throughput sequencing technology, has been widely applied in genomics research. Different from short-read sequencing, single molecular long-read sequencing can generate longer read lengths (ranging from 1,000 to 20,000 bases or longer), enabling it to span long segments of structural variation regions, significantly improving the accuracy and integrity of somatic structural variation detection.
[0008] A significant advantage of single molecular long-read sequencing is that it does not rely on the PCR amplification process, thus avoiding the potential biases that may occur during amplification. Long-read sequencing can directly read the original DNA fragments, providing more reliable data than short-read sequencing, especially when analyzing complex genomic regions (such as highly repetitive sequence regions). In addition, when performing genome assembly and de novo sequencing, long-read sequencing can provide more complete genomic information, which has important technical advantages for the detection of somatic structural variations.
[0009] Although single molecular long-read sequencing provides longer sequence read lengths, which helps to improve the detection accuracy of somatic structural variations, existing somatic structural variation detection methods based on long-read sequencing still have certain deficiencies. Most of the existing mainstream methods rely on aligning read sequences with the reference genome sequence and inferring structural variations by analyzing alignment breakpoints and sequence depth information. However, due to the heterogeneity of the genome, low-complexity sequences, and the errors of long-read sequencing itself, traditional alignment methods still cannot accurately detect complex somatic structural variations completely. Especially in the detection of low-frequency variations or heterogeneous variations between different cell populations, errors are still likely to occur.
[0010] In addition, most existing algorithms rely on alignment breakpoint information and ignore other sequence features, such as the integrity and heterogeneity of local sequences, which leads to insufficient detection ability when facing complex structural variations. How to improve the accurate detection of somatic structural variations in the presence of genomic heterogeneity and low-complexity sequences remains a major challenge in current genomics research.
[0011] Therefore, we urgently need to design a somatic structural variation detection method based on single molecular long-read sequence graphs to solve the above problems. Summary of the Invention
[0012] On the one hand, the present invention proposes a somatic structural variation detection method based on single molecular long-read sequence graphs, including the following steps:
[0013] Step 1: Extract all read - length local sequence information within the target structural variation interval. The target interval includes the reference genome sequence, the tumor sample sequence, and the normal control sample sequence. The read - length sequences are extracted from the target variation interval and its flanking intervals (default is 50bp) using the pysam package;
[0014] Step 2: Use the partial order alignment graph algorithm (sPOA) to construct a local graph genome. Align the extracted sequence information with the reference genome sequence to generate a local graph genome, and display the relationships between all aligned sequences through multiple sequence alignment;
[0015] Step 3: Convert the multiple sequence alignment results into multi - category variables. Specifically, encode the five categories of A, T, C, G, and gap in the alignment, and perform clustering analysis on the multi - category variables through a sequence graph mixture model to identify sequence features from different sources;
[0016] Step 4: According to the clustering analysis results, identify somatic structural variation categories that only originate from tumor samples;
[0017] Step 5: Generate consensus sequences for each category including the somatic structural variation groups according to the read - length clustering results, as the optimal result of the local genome.
[0018] As a preferred technical solution of the present invention, the local sequence extraction process in Step 1 includes using the pysam package to extract the spanning read - length sequence information of the reference genome, tumor samples, and normal control samples from the target variation interval and its flanking intervals. The read - length sequences contain complete information spanning the target variation region.
[0019] As a preferred technical solution of the present invention, the partial order alignment graph algorithm (sPOA) generates a local graph genome by aligning the reference genome sequence and all extracted read - length sequences. This graph shows the relative position relationships between the aligned sequences and provides a basis for subsequent multiple sequence alignment.
[0020] As a preferred technical solution of the present invention, the multi - category variable encoding process includes encoding the five categories of A, T, C, G, and gap that appear in the alignment as integer values from 0 to 4 respectively, and using the encoded sequence information as the input data for the mixture model.
[0021] As a preferred technical solution of the present invention, the sequence graph mixture model is a multi - category mixture model, which uses the expectation - maximization (EM) algorithm for clustering analysis and uses the Bayesian information criterion (BIC) to select the optimal number of clustering categories, thereby optimizing the clustering results and identifying read - length sequences from different sources.
[0022] As a preferred technical solution of the present invention, the clustering analysis results are used to identify and extract somatic structural variation categories from tumor samples, and the somatic structural variation sequences in the tumor samples are confirmed only by the clustering results identified in the tumor samples.
[0023] As a preferred technical solution of the present invention, the local graph genome in step 2 is obtained by performing a partial order alignment graph algorithm on the reference genome sequence and all the extracted read sequences. This graph genome shows the alignment relationships between the sequences, thus providing the necessary information for subsequent analysis steps.
[0024] As a preferred technical solution of the present invention, the sequence graph mixture model in step 3 is used to perform multi-class clustering on the encoded alignment results and optimize the clustering results through the Expectation-Maximization (EM) algorithm, so as to identify the tumor sample features related to somatic structural variations.
[0025] As a preferred technical solution of the present invention, the consensus sequence generation in step 5 is generated by respectively merging single-molecule long read sequences of different categories as the sequences of different components of the local graph genome. Among them, the category composed of only tumor reads is identified as the somatic structural variation sequence, ensuring the accuracy of the finally generated variant region sequence.
[0026] The present invention also provides a somatic structural variation detection system based on a single-molecule long read sequence graph mixture model, including:
[0027] Local sequence extraction module: used to extract all read local sequence information within the target variant interval, including the reference genome sequence, tumor sample sequence, and normal control sample sequence;
[0028] Partial order alignment graph construction module: used to construct a local graph genome through the partial order alignment graph algorithm (sPOA) and perform multiple sequence alignments, integrating all sequence alignment results into the same graph;
[0029] Sequence graph mixture model module: used to convert the alignment results into multi-class variables and perform clustering analysis through the mixture model, outputting clustering results of different categories;
[0030] Variant identification module: used to identify and extract somatic structural variation categories that only originate from tumor samples according to the clustering results and generate a consensus sequence.
[0031] Beneficial effects: The somatic structural variation detection method based on single-molecule long-read sequence graphs proposed by the present invention can effectively solve the influence of alignment errors and low-complexity sequences on detection accuracy in the prior art. By using the partial order alignment graph (sPOA) to construct a local graph genome, the errors generated by traditional alignment methods in complex genomic regions are avoided, thereby enabling more accurate variant detection results. Especially when dealing with samples with low-frequency variants and genomic heterogeneity, the present invention can significantly improve the detection accuracy and provide more reliable technical support for genomic research on tumors and other diseases.
[0032] As the core of the present invention, single-molecule long-read sequencing technology can generate longer reads and span structural variations in larger regions. Due to the read length limitation of traditional short-read sequencing methods, it is difficult to detect variants spanning large fragments, resulting in a serious restriction on the detection accuracy of variants. Long-read sequencing not only avoids these limitations in traditional methods but also can directly obtain more complete genomic information, providing stronger technical support for the comprehensive analysis of somatic structural variations. Through this technology, the present invention can cover a wider range of variant information when detecting structural variations in tumors and complex diseases, improving the detection coverage and accuracy.
[0033] The present invention can effectively handle genomic heterogeneity in somatic cells through a multi-class mixture model combined with the expectation-maximization (EM) algorithm. Somatic variants usually only exist in specific cell populations, and the genomes of different cell populations may vary significantly. By performing clustering analysis based on multiple sequence alignments and selecting the optimal number of clustering categories in combination with the Bayesian information criterion (BIC), variant sequences from different sources can be accurately identified. This method can not only reduce misjudgments due to genomic heterogeneity but also improve the stability and reliability of variant identification. Especially in samples with high heterogeneity, it shows stronger adaptability and accuracy than the prior art. Brief description of the drawings
[0034] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 It is a flow chart of the somatic structural variation detection method based on the single-molecule long-read sequence graph mixture model;
[0036] Figure 2 It is a schematic diagram showing the principle of step 1 of the somatic structural variation detection method based on the single-molecule long-read sequence graph mixture model;
[0037] Figure 3 Schematic diagram showing the principle of step 2 of the somatic structural variation detection method based on the single-molecule long-read sequence graph hybrid model
[0038] Figure 4 Schematic diagram showing the principle of step 3 of the somatic structural variation detection method based on the single-molecule long-read sequence graph hybrid model
[0039] Figure 5 Schematic diagram showing the principle of steps 4-5 of the somatic structural variation detection method based on the single-molecule long-read sequence graph hybrid model. Detailed implementation manner
[0040] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] The following combines the attached Figures 1-5 , and describes in detail the specific implementation manner of the present invention.
[0042] The present invention relates to a somatic structural variation detection method based on a single-molecule long-read sequence graph, which combines the advantages of a partial order alignment graph, a hybrid model, and an expectation maximization algorithm, can effectively analyze somatic structural variations, and solves the problems brought by alignment errors, low-complexity sequences, and genomic heterogeneity in the prior art.
[0043] Specifically, the implementation manner of the present invention includes the following steps: extracting local sequence information from the target structural variation interval, constructing a local graph genome using a partial order alignment graph and performing multiple sequence alignment, clustering and analyzing the alignment results through a hybrid model, identifying somatic variation categories, and generating a consensus sequence. The technical implementation of each step will be described in detail through specific embodiments.
[0044] Step 1: Local sequence extraction: The purpose of this step is to extract complete read sequence information from the target structural variation interval and its flanking regions. The extracted sequences include the reference genome sequence, the tumor sample sequence, and the normal control sample sequence.
[0045] Implementation details: Use the pysam package to extract sequences from the target variation interval and its flanking intervals, with the default flanking interval being 50bp. This package can extract sequence fragments from the genomes of each sample according to the reference genome coordinates. These sequences will be used as inputs in the subsequent steps.
[0046] Step 2: Local graph genome construction: In this process, we will construct a local graph genome through the semi-global partial order alignment (sPOA) algorithm and align all the extracted read sequences to show the relationships between sequences.
[0047] Implementation details: Use the sPOA algorithm to align the reference genome sequence of the target variant interval with the extracted tumor sample sequence and normal control sample sequence to generate a local graph genome. This graph can retain the order and positional relationships between various sequences during the alignment process and provide a basis for subsequent analysis steps.
[0048] Step 3: Multiclass variable encoding and mixture model clustering analysis: The purpose of this step is to convert the alignment results into multiclass variables and perform clustering analysis through a sequence graph mixture model.
[0049] Implementation details: For the local graph genome in Step 2, first perform integer encoding (0 - 4 respectively) on the five categories of A, T, C, G, and gap that appear in the alignment. Then, perform clustering analysis through a mixture model. Use a multiclass mixture model to optimize the multiclass variables through the expectation-maximization (EM) algorithm, and finally select the number of classes with the best Bayesian information criterion (BIC). The results of the clustering analysis will be used to identify the sequence characteristics of different sample sources.
[0050] Step 4: Somatic structural variant class identification: Through the clustering analysis results, we can identify the somatic structural variant classes composed only of the read lengths from tumor samples.
[0051] Implementation details: According to the clustering analysis results obtained in Step 3, identify the somatic structural variant classes composed only of the read lengths from tumor samples. At this time, the classes composed only of the read lengths from tumor samples will be determined as somatic structural variant classes for subsequent consensus sequence research.
[0052] Step 5: Consensus sequence generation: This step aims to generate consensus sequences for different classes of single-molecule long read sequences as the optimal results of different sequence components of the local genome.
[0053] Implementation details: Through the single-molecule long read clustering relationship, generate consensus sequences for different classes of read lengths in the variant region. This consensus sequence will be used as the final result to optimize the sequences of different genomic components in the target variant interval. Among them, the consensus sequence of the class composed only of the read lengths from tumor samples will be output as the somatic structural variant sequence.
[0054] Example 1: Detection of structural variants in tumor samples;
[0055] Background: We used the genomic data from a certain cancer patient and applied the algorithm of the present invention to detect somatic structural variants.
[0056] Step 1: Local sequence extraction: Use the pysam package to extract the sequence information of the target variant interval and its 50bp flanking regions from the reference genome and the genomes of tumor samples. At this time, the genome sequences of the reference genome and the tumor samples have been obtained through long-read sequencing technology.
[0057] Step 2: Local graph genome construction: Align the reference genome and the tumor sample sequences using the sPOA algorithm to generate a local graph genome. This graph shows the alignment relationship between the tumor sample and the reference genome in the target variant interval.
[0058] Step 3: Multi-class variable encoding and mixture model clustering analysis: During the alignment process, the five categories of A, T, C, G, and gap are respectively encoded as 0 - 4. Use the EM algorithm to perform clustering analysis on these encoded data, and select the Bayesian Information Criterion (BIC) as the basis to select the optimal number of clustering categories, and finally obtain the sequence categories of somatic structural variations in the tumor sample.
[0059] Step 4: Somatic structural variation category identification: Through the clustering analysis results, we can accurately identify the somatic structural variation categories in the tumor sample. Among them, the categories supported only by tumor reads are identified as somatic structural variation sequences.
[0060] Step 5: Consensus sequence generation: Generate a consensus sequence based on the variant sequences identified from the tumor sample. This consensus sequence, as the optimized result of the local genome, can provide a more accurate somatic structural variation sequence.
[0061] Example 2: Detection of somatic structural variations in different patient samples;
[0062] Background: To understand the universality of the present invention, we used genomic data from different tumor patients for experiments to detect somatic structural variations in different patient samples.
[0063] Step 1: Local sequence extraction: Use the pysam package to extract the sequences of the target variant interval and its flanking regions from the genomic data of each patient. The reference genome sequences and tumor sample sequences of each patient will be obtained through long-read sequencing technology.
[0064] Step 2: Local graph genome construction: Align the reference genome and the tumor sample sequences of all patient samples using the sPOA algorithm to generate a local graph genome, showing the variant regions in the reference genome and the patient samples.
[0065] Step 3: Multi-category variable encoding and mixture model clustering analysis: Encode the categories of A, T, C, G, and gap that appear in the alignment as 0-4 respectively, and use the EM algorithm to perform clustering analysis on the alignment data of all samples. Select the optimal number of clusters through BIC, and finally identify the sequence features related to somatic structural variations.
[0066] Step 4: Identification of somatic structural variation categories: According to the clustering results of each patient, identify the somatic structural variation categories in each sample, and extract the corresponding variant sequences from the tumor samples of each patient.
[0067] Step 5: Consensus sequence generation: Generate a consensus sequence for the variant regions in the tumor samples of each patient, optimize the genomic sequences of somatic structural variations of each patient, and provide support for subsequent clinical research and diagnosis.
[0068] Through the above embodiments, the somatic structural variation detection method based on single-molecule long-read sequence maps of the present invention can effectively improve the accuracy of somatic structural variation detection, especially in the case of genomic heterogeneity and low-complexity sequences. The method of the present invention can be applied not only to the detection of somatic structural variations of a single patient, but also to the batch detection of multiple patient samples, and has good universality and practical application value.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A somatic structural variation detection method based on a single-molecule long-read sequence graph mixture model, characterized in that It includes the following steps: Step 1: Extract all read - length local sequence information within the target structural variant interval. The target interval includes the reference genome sequence, the tumor sample sequence, and the normal control sample sequence. The read - length sequences are extracted from the target variant interval and its flanking intervals using the pysam package; Step 2: Use the partial order alignment graph algorithm to construct a local graph genome. Align the extracted sequence information with the reference genome sequence to jointly construct a partial order alignment graph, and display the relationships between all aligned sequences through multiple sequence alignment; Step 3: Convert the multiple sequence alignment results into multi - category variables. Specifically, it includes encoding the five categories of A, T, C, G, and gap in the alignment, and performing clustering analysis on the multi - category variables through a sequence graph mixture model to identify sequence characteristics from different sources; Step 4: According to the clustering analysis results, identify the cluster composed only of read - lengths from the tumor sample as the somatic structural variant cluster; Step 5: Generate consensus sequences for each category including the somatic structural variant cluster according to the read - length clustering results, as the optimal result of the local genome, in order to be used as the optimal result of the local genome.
2. The somatic structural variation detection method based on a single-molecule long-read sequence map according to claim 1, wherein The local sequence extraction process in Step 1 includes using the pysam package to extract the spanning read - length sequence information of the reference genome, tumor sample, and normal control sample from the target variant interval and its flanking intervals. The read - length sequences contain complete information spanning the target variant region.
3. The somatic structural variation detection method based on single-molecule long-read sequence map according to claim 1, wherein The partial order alignment graph algorithm generates a local graph genome by aligning the reference genome sequence and all extracted read - length sequences. This graph shows the relative positional relationships between the aligned sequences and provides a basis for subsequent multiple sequence alignment.
4. The somatic structural variation detection method based on a single-molecule long-read sequence map according to claim 1, wherein The multi - category variable encoding process includes encoding the five categories of A, T, C, G, and gap that appear in the alignment as integer values from 0 to 4 respectively, and using the encoded sequence information as the input data for the mixture model.
5. The somatic structural variation detection method based on the single-molecule long-read sequence map according to claim 1, wherein The sequence graph mixture model is a multi - category mixture model. It uses the expectation - maximization algorithm for clustering analysis and uses the Bayesian information criterion to select the optimal number of clustering categories, thereby optimizing the clustering results and identifying read - length sequences from different sources.
6. The somatic structural variation detection method based on single-molecule long-read sequence map according to claim 1, wherein The clustering analysis results are used to identify and extract the somatic structural variant categories from the tumor sample. The somatic structural variant sequences in the tumor sample are confirmed only through the clustering results identified in the tumor sample.
7. The somatic structural variation detection method based on the single-molecule long-read sequence map according to claim 1, wherein The local graph genome in Step 2 is obtained by using the partial order alignment graph algorithm to align the reference genome sequence with all extracted read - length sequences. This graph genome shows the alignment relationships between the sequences, thus providing necessary information for subsequent analysis steps.
8. The somatic structural variation detection method based on single-molecule long-read sequence map according to claim 1, wherein The sequence graph mixture model in Step 3 is used to perform multi - category clustering on the encoded alignment results and optimize the clustering results through the expectation - maximization algorithm, so as to identify the tumor read - length sequences related to somatic structural variants from them.
9. The somatic structural variation detection method based on a single-molecule long-read sequence map according to claim 1, wherein The consensus sequence in step 5 is generated by respectively merging single-molecule long-read sequences of different categories as sequences of different components of the local graph genome; the category supported only by tumor reads is identified as the somatic structural variation sequence, and on this basis, the local genomic sequence is optimized to make the finally generated variant region sequence accurate and error-free.
10. A somatic structural variant detection system based on a single-molecule long-read sequence map, characterized in that, Including: Local sequence extraction module: used to extract all read local sequence information within the target variant interval, including reference genome sequence, tumor sample sequence, and normal control sample sequence; Partial order alignment graph construction module: used to construct a local graph genome through the partial order alignment graph algorithm (sPOA) and output it in the form of multiple sequence alignment to obtain the alignment results of all sequences; Sequence graph mixture model module: used to convert the alignment results into multi-category variables and perform clustering analysis through the mixture model, and output the clustering results of different categories; Variant identification module: used to identify and extract the somatic structural variation categories that only originate from tumor samples according to the clustering results and generate a consensus sequence.