Genome structure variation detection method, system, equipment and medium
Patent Information
- Application Number
- CN202280102664.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-07-15
AI Technical Summary
In existing technologies, second-generation and third-generation sequencing technologies have limitations in detecting genomic structural variations (SVs), especially third-generation sequencing reads which are generally between 500 bp and 50 kb, making it difficult to detect SVs larger than 50 kb.
By obtaining the sequencing sequence of the sample to be tested, single base variation information analysis is performed, haplotype typing and assembly are carried out, and genome structural variation analysis is performed using the haplotype sequence. Combining the advantages of third-generation and second-generation sequencing, accurate haplotype typing and assembly are performed to obtain continuous haplotype sequences of more than 1MB.
This results in more complete and accurate SV detection, covering almost the entire SV size range, providing higher continuity and accuracy, and making the detection results more precise.
Smart Images

Figure CN120322562A_ABST
Abstract
Description
Method, system, device and medium for detecting genome structural variation Technical Field
[0001] The present invention relates to the field of genes, and in particular to a method, system, equipment and medium for detecting genomic structural variation. Background Art
[0002] Genomic variations are generally divided into three categories: SNP (Single Nucleotide Polymorphisms), INDEL (Insertion and Deletion, referring to the insertion or deletion of a small sequence fragment at a certain position in the genome) and SV (Structural variation).
[0003] Existing technologies, such as second-generation sequencing (NGS), can effectively detect SNP and INDEL variants. However, due to the limited read length of NGS, the detection results of SV variants obtained using NGS are inaccurate. The development of third-generation sequencing (NGS) has improved the detection of SV variants to a certain extent, but it still faces the following challenges: Since NGS read lengths generally range from 500bp to 50kb, the size of the SV that can be detected is often limited to less than 50kb.
[0004] Summary of the Invention
[0005] The technical problem to be solved by the present invention is to overcome the defect of limited size range of SV detection and provide a method, system, equipment and medium for detecting genomic structural variation.
[0006] The present invention solves the above technical problems through the following technical solutions:
[0007] In a first aspect, a method for detecting genomic structural variation is provided, the method comprising:
[0008] Obtaining a plurality of sequencing sequences of a sample to be tested;
[0009] Performing haplotype typing on the sequencing sequence according to the single base variation information in each sequencing sequence;
[0010] Sequencing sequences belonging to the same haplotype are assembled to obtain haplotype sequences, and genomic structural variation analysis is performed based on the haplotype sequences.
[0011] Optionally, performing haplotype typing on the sequenced sequences according to the single-base variation information of each sequenced sequence comprises:
[0012] Obtain single-base variation information in each sequencing sequence;
[0013] Determine the haplotype information that matches the single-base variation information, and type the sequencing sequence to the corresponding haplotype, wherein the haplotype information is obtained by performing linkage analysis on the single-base variation information of all sequencing sequences.
[0014] Optionally, obtain single-base variation information for each sequencing sequence, including:
[0015] The sequencing sequence is aligned with the reference genome, and all base information in the sequencing sequence that is inconsistent with the reference genome is determined as single base variation information of the sequencing sequence.
[0016] Optionally, determining haplotype information that matches the single-base variation information includes:
[0017] Perform linkage strength analysis on the single-base variation information in the sample to be tested by seed extension to determine the haplotype information of the sample to be tested;
[0018] The single-base variation information of the sequencing sequence is matched with the haplotype information to determine the haplotype to which the sequencing sequence belongs.
[0019] Optionally, sequencing sequences belonging to the same haplotype are assembled to obtain haplotype sequences, comprising:
[0020] Sequencing sequences belonging to the same haplotype are locally assembled to obtain haplotype sequences.
[0021] Optionally, genomic structural variation analysis is performed based on haplotype sequences, including:
[0022] The haplotype sequence is aligned with a reference genome, and whether genomic structural variation occurs in the haplotype sequence is determined based on the result of the sequence alignment.
[0023] Optionally, obtaining a sequencing sequence of a sample to be tested includes:
[0024] The sequencing sequence of the sample to be detected is obtained by performing long-read sequencing on the sample to be detected.
[0025] In a second aspect, a system for detecting genomic structural variation is provided, the system comprising:
[0026] An acquisition module is used to obtain a sequencing sequence of a sample to be tested, where the number of the sequencing sequences is multiple;
[0027] A typing module, configured to perform haplotype typing on the sequencing sequence according to the single-base variation information in each sequencing sequence;
[0028] The analysis module is used to assemble sequencing sequences belonging to the same haplotype to obtain haplotype sequences, and perform genomic structural variation analysis based on the haplotype sequences.
[0029] In a third aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for detecting genomic structural variations described in the first aspect is implemented.
[0030] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for detecting genomic structural variation described in the first aspect is implemented.
[0031] The positive and progressive effects of the present invention are that haplotype typing can be used to accurately perform haplotype typing on sequencing sequences before SV detection, sequencing sequences from different haplotype sources can be typed, and subsequent SV detection can be performed based on the haplotype sequence. On the one hand, the assembled haplotype sequence can accurately display the source of the detected SV. On the other hand, the assembled haplotype sequence has higher continuity and accuracy than the sequencing sequence itself, is often more than 1MB in length, and can cover almost the entire SV size range, making the SV detection results more complete and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG1 is a flow chart of a method for detecting genomic structural variation provided by an exemplary embodiment of the present invention;
[0033] FIG2 is a flowchart of step S11 of a method for detecting gene structural variation according to an exemplary embodiment of the present invention;
[0034] FIG3 is a schematic diagram of haplotype information of a method for detecting gene structural variation provided by an exemplary embodiment of the present invention;
[0035] FIG4 is a flowchart of step S112 of a method for detecting gene structural variation according to an exemplary embodiment of the present invention;
[0036] FIG5 is a schematic diagram of a seed extension method for detecting gene structure variation provided by an exemplary embodiment of the present invention;
[0037] FIG6 is a schematic diagram of a method for detecting genomic structural variation provided by an exemplary embodiment of the present invention;
[0038] FIG7 is a schematic diagram of SNPs after haplotype typing in a method for detecting genomic structural variation provided by an exemplary embodiment of the present invention;
[0039] FIG8 is a module diagram of a system for detecting genomic structural variation according to an exemplary embodiment of the present invention;
[0040] FIG9 is a structural diagram of a base electronic device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0041] The present invention is further described below by way of exemplary embodiments, but the present invention is not limited to the scope of the embodiments.
[0042] Prior to this example, a brief description of genomic structural variation is given. Genomic structural variation is referred to as SV in this example. SV includes insertion or deletion of long fragment sequences with a length of more than 50 bp, chromosomal inversion, sequence translocation within or between chromosomes, copy number variation, and some more complex forms of variation.
[0043] In addition, the single base variation information in this embodiment is SNP.
[0044] FIG1 is a method for detecting genomic structural variation provided by an exemplary embodiment of the present invention. Referring to FIG1 , the method includes:
[0045] S10. Obtain a sequencing sequence of a sample to be tested, where the number of sequencing sequences is multiple.
[0046] Among them, the sequencing sequence is obtained by long-read sequencing of the sample to be tested.
[0047] In one embodiment, the sequencing sequence can be obtained by performing third-generation sequencing on the sample to be tested. The longest read length of the third-generation sequencing can reach 50kb, and the average read length is 15 to 20kb. Since the sequencing sequence obtained by the third-generation sequencing of SV has better integrity, the read length of the sequencing sequence obtained by the third-generation sequencing can better support the read length requirements of SV detection.
[0048] In one embodiment, the sequencing sequence can be obtained by performing large-fragment DNA library construction technology such as stLFR (Single tube Long Fragment Read) or 10X Genomics (a sequencing technology) on the sample to be tested. It is a special second-generation sequencing, and the read length of the sequencing sequence can also reach 10kb to 300kb.
[0049] In one embodiment, the sequencing sequence can also be obtained by performing second-generation sequencing and third-generation sequencing on the sample to be detected at the same time. Among them, the sequencing sequence obtained based on the second-generation sequencing can be used to determine the haplotype of the sample to be detected, and the sequencing sequence obtained by the third-generation sequencing is used to perform haplotype typing on the sequencing sequence and assemble the haplotype sequence. The sequencing sequence obtained by the third-generation sequencing can provide the demand for long read length for SV detection, and the sequencing sequence obtained by the second-generation sequencing can provide higher accuracy for sequence comparison and SV detection by performing precision correction on the third-generation sequencing sequence.
[0050] S11. Perform haplotype typing on the sequencing sequences based on the single-base variation information in each sequencing sequence.
[0051] In one embodiment, referring to FIG2 , step S11 specifically includes:
[0052] S111. Obtain single-base variation information in each sequencing sequence.
[0053] In one embodiment, the sequencing sequence is aligned with the corresponding sites of the reference genome, and all bases in the sequencing sequence that are inconsistent with the reference genome are determined as single base variation information of the sequencing sequence, i.e., SNPs. Among them, SNPs can include heterozygous SNPs. For example, the sequencing sequence is AGTCTTAG, and the corresponding site of the reference genome is AGGCTTCG. When performing sequence alignment, it can be seen that the base types of the third and seventh sites do not match. At this time, the base information T of the third site and the base information A of the seventh site in the sequencing sequence can be determined as SNPs.
[0054] S112. Determine the haplotype information that matches the single-base variation information, and type the sequencing sequence into the corresponding haplotype. The haplotype information is obtained by performing linkage analysis on the single-base variation information of all sequencing sequences.
[0055] Here, haplotype information refers to the combination of alleles for a group of related SNPs located on a chromosome or in a certain region, and haplotype refers to the source of a group of haplotype information. Referring to Figure 3, taking a diploid genome as an example, the diploid genome includes the paternal haplotype, designated Dad, and the maternal haplotype, designated Mom. The paternal haplotype information includes GTCCA, while the maternal haplotype includes TAGTG.
[0056] In one embodiment, referring to FIG4 , step S112 specifically includes:
[0057] S112-1. Perform linkage strength analysis on the single-base variation information in the sample to be tested by seed extension to determine the haplotype information of the sample to be tested.
[0058] In one embodiment, the haplotype information of the sample to be tested is determined by performing SNP linkage strength analysis on all sequencing sequences in the sample to be tested. The haplotype information includes all SNPs belonging to the same haplotype in the sample to be tested. The haplotype information is the arrangement of SNPs with a linkage relationship.
[0059] Taking Figure 5 as an example, the determination of haplotype information by seed extension is further explained:
[0060] After aligning the sequence with the reference genome, four SNPs were identified: A1 / G1, G2 / C2, A3 / T3, and T4 / C4. The first pair of bases, A1 and G1, was used as a seed pair. The linkage strength between the seed and the other SNPs was calculated to identify the SNP with the strongest linkage strength. This site was then incorporated into the seed before the next extension. The remaining SNPs were continuously incorporated into the seed until no SNPs could be found on the same long DNA fragment as any other site on the seed, ultimately yielding the haplotype information for AGTC and GCAT.
[0061] The linkage relationship between each SNP is provided by a long DNA fragment. By performing linkage strength analysis on the SNPs of all sequenced sequences, multiple haplotype information of the sample to be tested, such as AGTC and GCAT, and the corresponding haplotypes, can be obtained.
[0062] In one embodiment, the sequencing sequence that determines that haplotype information is extended by seed can be the sequencing sequence that three generations of sequencing obtain, or it can be the sequencing sequence that second-generation sequencing such as stLFR or 10X Genomics obtain. Generally speaking, the sequencing sequence obtained by three generations of sequencing can complete the determination of haplotype information in sample to be detected, but due to the read length of the sequencing sequence of stLFR can reach about 300kb at most, the read length of the sequencing sequence that three generations of sequencing obtain is 50kb, means that stLFR can span longer pure and region (i.e., male and female identical region), and the sequencing sequence that second-generation sequencing obtains can provide support for the haplotype typing of the sequencing sequence that three generations of sequencing obtain. Make the continuity and accuracy of each haplotype sequence obtained higher.
[0063] S112-2. Match the single-base variation information of the sequencing sequence with the haplotype information to determine the haplotype to which the sequencing sequence belongs.
[0064] The single-base variation information obtained in step S111 is mapped to the corresponding haplotype information, that is, the SNP of each sequencing sequence is mapped to the corresponding SNP in the haplotype information. When the SNP of the sequencing sequence is successfully mapped to the corresponding SNP in the corresponding haplotype information, the sequencing sequence is considered to belong to the haplotype corresponding to the haplotype information. For example, if the haplotype information includes AGTCGTTTCGTTTAA, and the single-base variation information obtained from the sequencing sequence includes CGTTTCGT, the single-base variation information in the sequencing sequence can be successfully mapped to the CGTTTCGT segment in the haplotype. At this time, it can be considered that the sequencing sequence matches the haplotype information and belongs to the haplotype.
[0065] In this embodiment, by accurately performing haplotype typing on the sequencing sequences before SV detection, the haplotype to which each sequencing sequence belongs can be determined, and the sequencing sequences can be assembled according to the haplotype to obtain a haplotype sequence for subsequent SV detection. SVs obtained from the same haplotype sequence detection can be marked as coming from the same haplotype to determine the source of the SV, which can provide a reliable and stable genetic detection basis for the analysis of biological genetic diversity and clinical medical diagnosis and analysis, such as cancer.
[0066] S12. Assemble the sequencing sequences belonging to the same haplotype to obtain haplotype sequences, and perform genomic structural variation analysis based on the haplotype sequences.
[0067] In one embodiment, the sequencing sequences included in each haplotype can be obtained according to the preceding steps. Since the sequencing sequences can be obtained based on third-generation sequencing or second-generation sequencing, the sequencing sequences of the same haplotype may only include sequencing sequences obtained by third-generation sequencing, or may also include sequencing sequences obtained by second-generation sequencing and sequencing sequences obtained by third-generation sequencing.
[0068] In one embodiment, assembling the sequencing sequence in step S12 specifically includes:
[0069] Sequencing sequences belonging to the same haplotype are locally assembled to obtain haplotype sequences.
[0070] Among them, the haplotype sequence obtained by local assembly is longer and more accurate than the original sequencing sequence.
[0071] In one embodiment, when sequencing sequence promptly comprises three generations of sequencing sequences and second generation sequencing sequences, first the sequencing sequence obtained by three generations of sequencing is carried out local assembly and obtains haplotype sequence, then by second generation sequencing sequence, the assembling result of three generations of sequencing sequences is adjusted, obtain the haplotype sequence that accuracy is higher.Particularly, the sequencing sequence obtained by three generations of sequencing is assembled and obtains haplotype sequence, when also comprising the sequencing sequence that second generation sequencing obtains in the sequencing sequence of haplotype, the sequencing sequence that described second generation sequencing is obtained is matched to the corresponding position in the haplotype sequence and analyzed.For example, the base in the haplotype sequence corresponding position is AGTTCTG, and the base in the sequencing sequence that second generation sequencing obtains is AGTTCAG, now, the sequencing sequence that the haplotype sequence corresponding position can be revised according to the sequencing that second generation sequencing obtains, so that the haplotype sequence result is more accurate.
[0072] In one embodiment, the sequencing sequences are assembled using Celera (a long-read sequence assembly tool) or Canu (a long-read sequence assembly tool), and the assembly is based on the OLC principle, namely overlap-layout-consensus (overlap-arrangement-generate consistent sequence) idea.
[0073] In one embodiment, performing genomic structural variation analysis based on haplotype sequences in step S12 specifically includes:
[0074] The haplotype sequence is aligned with the corresponding position of the reference genome to obtain a BAM format alignment file. The svim tool is then used to compare the files and analyze the genomic structural variation to obtain multiple types of detection results. Finally, by screening and filtering the results, a more comprehensive and accurate SV detection result of the genomic haplotype can be obtained.
[0075] Furthermore, because SV detection is based on assembled haplotype sequences, which can contain multiple consecutive sequenced sequences, more complete SV detection results can be obtained. Furthermore, for individual SV detection, SV breakpoint information can be obtained, i.e., the start and end positions of an SV. However, due to limitations in sequencing accuracy, existing technologies may not be able to display a complete SV, meaning that SV breakpoint information cannot be obtained.
[0076] In this embodiment, the sequencing sequence of each haplotype can be assembled into one or more continuous haplotype sequences exceeding 1MB in length through local assembly. Since the length and accuracy of the continuous haplotype sequences are significantly improved relative to the original sequencing sequence, the SV results obtained by aligning them to the reference genome will be more accurate. In addition, the haplotype sequences obtained by the local assembly strategy can span multiple SVs, so the scope of SV detection in this embodiment can be expanded to truly large SVs, no longer subject to the read length of the sequencing sequence itself.
[0077] The present embodiment is further described below through a specific implementation method, see FIG6 :
[0078] In step 1, the international standard HG002 was used as the test sample, and conventional third-generation sequencing (40X) and second-generation sequencing based on the stLFR library (60X) were performed respectively, with a total base amount of 300GB and a total depth of 100X. After aligning the second-generation sequencing sequence based on the stLFR library to the reference genome using bwa (a sequence alignment tool), an effective coverage rate of 99.95% was obtained. Conventional variation detection was further performed to filter out 2.52 million heterozygous SNPs contained in the test sample for subsequent variation typing. These heterozygous SNPs were haplotyped according to seed extension, and the median length of the fragments after typing was as long as 11.35MB.
[0079] In step 2, we use minimap2 (a sequence alignment tool) to continue aligning the three-generation sequencing sequences to the reference genome to obtain the heterozygous SNPs contained in each sequencing sequence. Based on the haplotype information obtained in step 1, each sequencing sequence is mapped to the correct haplotype. There are two different haplotypes in the same chromosome region, and they each have an independent sequence pool. In order to verify the haplotype typing effect of the sequencing sequences obtained by three-generation sequencing, we aligned the sequencing sequences that have completed haplotype typing to the reference genome respectively to obtain typing of the two haplotypes, see Figure 7, each line in Figure 7 is a sequencing sequence, and the dark-colored marked part is a heterozygous SNP. It can be seen that the heterozygous SNP is specifically distributed on one of the haplotypes, and no mosaicism occurs, indicating that the typing effect of the side sequence obtained by three-generation sequencing has reached a satisfactory level.
[0080] In step 3, we performed a local assembly of the sequences in the sequence pool for each haplotype, without affecting each other. The total assembly length was 5.4GB (2.6GB*2), with a genome coverage of 96.4% and an assembly N50 of 1.1MB.
[0081] Step 4. Based on the assembled haplotype sequences, we again use minimap2 to align them to the reference genome, and then use the svim tool to detect the preliminary SV detection results, and then perform redundancy removal to obtain the final SV data set. The SV data set is compared with the HG002 international standard set to evaluate its accuracy (precision), sensitivity (recall) and F1 score. See the table below. It can be seen that the accuracy of SV detection in this embodiment is greatly improved compared to Sniffles. Among them, sensitivity = the actual number of SVs detected in the standard set / the total number of SVs in the standard set, accuracy = the correct number of SVs detected / the total number of SVs detected,
[0082]
[0083] FIG8 is a schematic diagram of a system for detecting genomic structural variation according to an exemplary embodiment of the present invention. Referring to FIG8 , the system includes:
[0084] An acquisition module 81 is used to acquire a sequencing sequence of a sample to be detected, where the number of the sequencing sequences is multiple;
[0085] A typing module 82 is used to perform haplotype typing on the sequencing sequence according to the single-base variation information in each sequencing sequence;
[0086] The analysis module 83 is used to assemble sequencing sequences belonging to the same haplotype to obtain haplotype sequences, and perform genome structural variation analysis based on the haplotype sequences.
[0087] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is only illustrative, and the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without expending creative work.
[0088] Figure 9 is a schematic diagram of the structure of an electronic device provided in this embodiment. Figure 9 shows a block diagram of an exemplary electronic device 90 suitable for implementing the embodiments of the present invention. The electronic device 90 shown in Figure 9 is merely an example and should not limit the functionality or scope of use of the embodiments of the present invention.
[0089] As shown in FIG9 , electronic device 90 may be a general-purpose computing device, such as a server device. Components of electronic device 90 may include, but are not limited to, at least one processor 91, at least one memory 92, and a bus 93 connecting various system components (including memory 92 and processor 91).
[0090] The bus 93 includes a data bus, an address bus, and a control bus.
[0091] The memory 92 may include a volatile memory, such as a random access memory (RAM) 921 and / or a cache memory 922 , and may further include a read-only memory (ROM) 923 .
[0092] The memory 92 may also include a program tool 925 having a set (at least one) of program modules 924, such program modules 924 including but not limited to: an operating system, one or more application programs, other program modules and program data, each of which or some combination may include an implementation of a network environment.
[0093] The processor 91 executes computer programs stored in the memory 92 to perform various functional applications and data processing.
[0094] The electronic device 90 can also communicate with one or more external devices 94 (e.g., a keyboard, pointing device, etc.). Such communication can occur via an input / output (I / O) interface 95. Furthermore, the electronic device 90 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 96. The network adapter 96 communicates with other modules of the electronic device 90 via a bus 93. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device 90, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.
[0095] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the present application, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0096] An exemplary embodiment of the present invention is a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for detecting genomic structural variation of the above embodiment.
[0097] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0098] In a possible implementation, the present invention may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the method for detecting genomic structural variation.
[0099] The program code for executing the present invention may be written in any combination of one or more programming languages, and the program code may be executed entirely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on a remote device, or entirely on the remote device.
[0100] Although specific embodiments of the present invention have been described above, those skilled in the art will appreciate that these are merely illustrative and that the scope of the present invention is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, and such changes and modifications are intended to fall within the scope of the present invention.
Claims
1. A method for detecting genomic structural variation, characterized in that: The method comprises: Obtaining a sequencing sequence of a sample to be detected, wherein the number of the sequencing sequences is multiple; Performing haplotype typing on the sequencing sequences according to the single base variation information in each sequencing sequence; The sequencing sequences belonging to the same haplotype are assembled to obtain haplotype sequences, and genomic structural variation analysis is performed based on the haplotype sequences.
2. The method for detecting genomic structural variation according to claim 1, characterized in that: Performing haplotype typing on the sequencing sequence according to the single base variation information of each sequencing sequence comprises: Obtain single-base variation information in each sequencing sequence; Determine the haplotype information that matches the single-base variation information, and type the sequencing sequence to the corresponding haplotype, wherein the haplotype information is obtained by performing linkage analysis on the single-base variation information of all sequencing sequences.
3. The method for detecting genomic structural variation according to claim 2, characterized in that: Obtain single-base variation information for each sequencing sequence, including: The sequencing sequence is compared with the reference genome, and all base information in the sequencing sequence that is inconsistent with the reference genome is determined as single base variation information of the sequencing sequence.
4. The method for detecting genomic structural variation according to claim 2, characterized in that: Determining haplotype information matching the single base variation information includes: Perform linkage strength analysis on single-base variation information in the sample to be tested by seed extension to determine the haplotype information of the sample to be tested; The single-base variation information of the sequencing sequence is matched with the haplotype information to determine the haplotype to which the sequencing sequence belongs.
5. The method for detecting genomic structural variation according to claim 1, characterized in that: Sequencing sequences belonging to the same haplotype are assembled to obtain haplotype sequences, including: Sequencing sequences belonging to the same haplotype are locally assembled to obtain haplotype sequences.
6. The method for detecting genomic structural variation according to claim 1, characterized in that: Analysis of genomic structural variation based on haplotype sequences, including: The haplotype sequence is compared with a reference genome, and whether the haplotype sequence has a genomic structural variation is determined based on the result of the sequence comparison.
7. The method for detecting genomic structural variation according to claim 1, characterized in that: Obtain the sequencing sequence of the sample to be tested, including: The sequencing sequence of the sample to be detected is obtained by performing long-read sequencing on the sample to be detected.
8. A system for detecting genome structural variation, characterized in that: The system comprises: An acquisition module is used to acquire a sequencing sequence of a sample to be detected, where the number of the sequencing sequences is multiple; A typing module, used for performing haplotype typing on the sequencing sequence according to the single base variation information in each sequencing sequence; The analysis module is used to assemble sequencing sequences belonging to the same haplotype to obtain haplotype sequences, and perform genome structural variation analysis based on the haplotype sequences.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for detecting genomic structural variation according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for detecting genomic structural variation according to any one of claims 1 to 7 is implemented.